{
  "id": 129270,
  "title": "How to speed up training phase?",
  "url": "/competitions/bengaliai-cv19/discussion/129270",
  "author_name": "",
  "post_date": "2020-02-06T18:09:36.775390Z",
  "votes": 7,
  "comment_count": 17,
  "views": 0,
  "content": "<p>My current setup is:\n<code>\nresolution: 128x128\nmodel: se_resnext50_32x4d\naugs: mixup + cutmix\nmachine: tesla p100\nbatch_size: 256\nfp16: True\n</code></p>\n\n<p>One epoch takes 12 mins (11 train + 1 valid). That means if I want to train for ~80 epochs I need to brake my pipeline into two kernel commits: each commit performing 40 epochs with duration approximately 8:30 hours.</p>\n\n<p>An this is just a single fold model with <code>128x128</code> images.</p>\n\n<p>For example, <code>efficientnet-b0</code> takes about 13 min/epoch with <code>224x224</code> images.</p>\n\n<p>My <code>__get_item__</code> basically doing this:</p>\n\n<p><code>\nimage_path = os.path.join(self.data_folder, image_id + '.png')\nimage = cv2.imread(image_path, 0)\nimage = cv2.cvtColor(image, cv2.COLOR_GRAY2BGR)\nif self.transform:\n    image = self.transform(image=image)['image']\nimage = np.transpose(image, (2, 0, 1)).astype(np.float32)\n</code></p>\n\n<p><code>transform</code> here is just <code>albumentations.Normalize</code>. </p>\n\n<p>I thought to save all processed (i.e. after <code>__get_item__</code>) images in <code>npy</code> format and load them right away during training. Unfortunately, one <code>npy</code> array takes approx <code>600kb</code> vs <code>2-3kb</code> original png, so it is not feasible to apply this method.</p>\n\n<p>I also tried using <code>npz</code> format, but the timings got even worse.</p>\n\n<p>Do you have any advice for me?</p>",
  "messages": [
    {
      "id": "738582",
      "postDate": "02/06/2020 18:09:36",
      "content": "<p>My current setup is:\n<code>\nresolution: 128x128\nmodel: se_resnext50_32x4d\naugs: mixup + cutmix\nmachine: tesla p100\nbatch_size: 256\nfp16: True\n</code></p>\n\n<p>One epoch takes 12 mins (11 train + 1 valid). That means if I want to train for ~80 epochs I need to brake my pipeline into two kernel commits: each commit performing 40 epochs with duration approximately 8:30 hours.</p>\n\n<p>An this is just a single fold model with <code>128x128</code> images.</p>\n\n<p>For example, <code>efficientnet-b0</code> takes about 13 min/epoch with <code>224x224</code> images.</p>\n\n<p>My <code>__get_item__</code> basically doing this:</p>\n\n<p><code>\nimage_path = os.path.join(self.data_folder, image_id + '.png')\nimage = cv2.imread(image_path, 0)\nimage = cv2.cvtColor(image, cv2.COLOR_GRAY2BGR)\nif self.transform:\n    image = self.transform(image=image)['image']\nimage = np.transpose(image, (2, 0, 1)).astype(np.float32)\n</code></p>\n\n<p><code>transform</code> here is just <code>albumentations.Normalize</code>. </p>\n\n<p>I thought to save all processed (i.e. after <code>__get_item__</code>) images in <code>npy</code> format and load them right away during training. Unfortunately, one <code>npy</code> array takes approx <code>600kb</code> vs <code>2-3kb</code> original png, so it is not feasible to apply this method.</p>\n\n<p>I also tried using <code>npz</code> format, but the timings got even worse.</p>\n\n<p>Do you have any advice for me?</p>",
      "rawMarkdown": "My current setup is:\n```\nresolution: 128x128\nmodel: se_resnext50_32x4d\naugs: mixup + cutmix\nmachine: tesla p100\nbatch_size: 256\nfp16: True\n```\n\nOne epoch takes 12 mins (11 train + 1 valid). That means if I want to train for ~80 epochs I need to brake my pipeline into two kernel commits: each commit performing 40 epochs with duration approximately 8:30 hours.\n\nAn this is just a single fold model with `128x128` images.\n\nFor example, `efficientnet-b0` takes about 13 min/epoch with `224x224` images.\n\nMy `__get_item__` basically doing this:\n\n```\nimage_path = os.path.join(self.data_folder, image_id + '.png')\nimage = cv2.imread(image_path, 0)\nimage = cv2.cvtColor(image, cv2.COLOR_GRAY2BGR)\nif self.transform:\n    image = self.transform(image=image)['image']\nimage = np.transpose(image, (2, 0, 1)).astype(np.float32)\n```\n\n`transform` here is just `albumentations.Normalize`. \n\nI thought to save all processed (i.e. after `__get_item__`) images in `npy` format and load them right away during training. Unfortunately, one `npy` array takes approx `600kb` vs `2-3kb` original png, so it is not feasible to apply this method.\n\nI also tried using `npz` format, but the timings got even worse.\n\nDo you have any advice for me?",
      "votes": null
    },
    {
      "id": "738672",
      "postDate": "02/06/2020 21:14:04",
      "content": "<p>Are you using half precision training ? if not check out <code>apex</code> library. You can fit almost double amount of images in your batch. </p>",
      "rawMarkdown": "Are you using half precision training ? if not check out `apex` library. You can fit almost double amount of images in your batch.",
      "votes": null
    },
    {
      "id": "738679",
      "postDate": "02/06/2020 21:27:43",
      "content": "<p>Yep, I do. My current batch size for <code>se_resnext</code> is 256. I might push it even higher for smaller models like <code>b0</code>, but from my experiments, there was no difference between 256 and 512 for <code>b0</code>, but I might be mistaken, so will check it out again.</p>\n\n<p>I guess bigger batch size can be beneficial for training with <code>FocalLoss</code> or <code>ReducedFocalLoss</code> because it essentially lowers number of calls of <code>focal_loss_with_logits</code>, which you call for each of 168 + 11 + 7 classes, but for <code>CrossEntropyLoss</code> it seems like I hit some limit.</p>",
      "rawMarkdown": "Yep, I do. My current batch size for `se_resnext` is 256. I might push it even higher for smaller models like `b0`, but from my experiments, there was no difference between 256 and 512 for `b0`, but I might be mistaken, so will check it out again.\n\nI guess bigger batch size can be beneficial for training with `FocalLoss` or `ReducedFocalLoss` because it essentially lowers number of calls of `focal_loss_with_logits`, which you call for each of 168 + 11 + 7 classes, but for `CrossEntropyLoss` it seems like I hit some limit.",
      "votes": null
    },
    {
      "id": "738694",
      "postDate": "02/06/2020 21:54:39",
      "content": "<p>The bottleneck is (probably) the Dataloader. Try to load all of the train data into memory. Currently, you are loading (and preprocessing) all of the images 80 (#epochs) times.\nTips:\n- use one channel (for storage) only and convert it to 3 channels in  <code>__get_item__</code>. \n- store as <code>np.uint8</code> (floats won't fit)\n- If you are using a big model, load the model+weights first, move to CUDA, and then load all of the images.\n- If you still have memory issues, then try this training method: load one (or two) parquet into memory like above; train 80 epochs, then load the next parquet (or next two parquets if you have enough ram) and train 80 more epochs</p>",
      "rawMarkdown": "The bottleneck is (probably) the Dataloader. Try to load all of the train data into memory. Currently, you are loading (and preprocessing) all of the images 80 (#epochs) times.\nTips:\n- use one channel (for storage) only and convert it to 3 channels in  `__get_item__`. \n- store as `np.uint8` (floats won't fit)\n- If you are using a big model, load the model+weights first, move to CUDA, and then load all of the images.\n- If you still have memory issues, then try this training method: load one (or two) parquet into memory like above; train 80 epochs, then load the next parquet (or next two parquets if you have enough ram) and train 80 more epochs",
      "votes": null
    },
    {
      "id": "738726",
      "postDate": "02/06/2020 23:22:06",
      "content": "<p><a href=\"/lightnezzofbeing\">@lightnezzofbeing</a> How batch_size affects the accuracy ? I'm warried that larger batch_size makes LB score worse.</p>",
      "rawMarkdown": "lightnezzofbeing How batch_size affects the accuracy ? I'm warried that larger batch_size makes LB score worse.",
      "votes": null
    },
    {
      "id": "738842",
      "postDate": "02/07/2020 04:47:49",
      "content": "<p><a href=\"/toshik\">@toshik</a> I have also found same issue! Although I am not sure why this happens.</p>",
      "rawMarkdown": "toshik I have also found same issue! Although I am not sure why this happens.",
      "votes": null
    },
    {
      "id": "739141",
      "postDate": "02/07/2020 13:12:45",
      "content": "<p>Hi <a href=\"/toshik\">@toshik</a> <br>\nI think what you are saying is true. However, for me 128 and 256 almost had no difference in score, but I suspect smaller batch size can yield a better result.</p>",
      "rawMarkdown": "Hi @toshik  \nI think what you are saying is true. However, for me 128 and 256 almost had no difference in score, but I suspect smaller batch size can yield a better result.",
      "votes": null
    },
    {
      "id": "739143",
      "postDate": "02/07/2020 13:13:23",
      "content": "<p>Thanks! Will give it a try!</p>",
      "rawMarkdown": "Thanks! Will give it a try!",
      "votes": null
    },
    {
      "id": "739200",
      "postDate": "02/07/2020 14:52:17",
      "content": "<p>just did quick test of batch size (same validation set across experiment):</p>\n\n<p><code>\nmodel: se_resnext_50\nimg_size: 128\ntrain: 100 epoch\n</code></p>\n\n<p>```\nbs: 1024\ncv: 0.9861688871</p>\n\n<p>bs: 256\ncv: 0.982144479</p>\n\n<p>bs: 128\ncv: 0.981938083\n```</p>\n\n<p>In my setup batch size worsen the performance. But It might be different depending on your optimizers and restrains. </p>",
      "rawMarkdown": "just did quick test of batch size (same validation set across experiment):\n\n```\nmodel: se_resnext_50\nimg_size: 128\ntrain: 100 epoch\n```\n\n\n```\nbs: 1024\ncv: 0.9861688871\n\nbs: 256\ncv: 0.982144479\n\nbs: 128\ncv: 0.981938083\n```\n\nIn my setup batch size worsen the performance. But It might be different depending on your optimizers and restrains.",
      "votes": null
    },
    {
      "id": "739212",
      "postDate": "02/07/2020 15:05:58",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Thank you for your information ! Larger batch_size increased you score so much !</p>",
      "rawMarkdown": "drhabib Thank you for your information ! Larger batch_size increased you score so much !",
      "votes": null
    },
    {
      "id": "739213",
      "postDate": "02/07/2020 15:06:56",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> did you use the same LR?</p>",
      "rawMarkdown": "drhabib did you use the same LR?",
      "votes": null
    },
    {
      "id": "739216",
      "postDate": "02/07/2020 15:11:04",
      "content": "<p>yes! I use same <code>LR</code>,  <code>ADAM</code>,  and same <code>WD</code> and <code>one cycle policy</code>.</p>",
      "rawMarkdown": "yes! I use same `LR`,  `ADAM`,  and same `WD` and `one cycle policy`.",
      "votes": null
    },
    {
      "id": "739218",
      "postDate": "02/07/2020 15:12:35",
      "content": "<p>for the bs 1024, I had to use google cloud, for 256 and 128 I used my local computer. I made sure the environment is the same... </p>",
      "rawMarkdown": "for the bs 1024, I had to use google cloud, for 256 and 128 I used my local computer. I made sure the environment is the same...",
      "votes": null
    },
    {
      "id": "739439",
      "postDate": "02/07/2020 20:50:48",
      "content": "<p>Thanks for sharing <a href=\"/drhabib\">@drhabib</a>. I'm very impressed how you are able to run a \"quick test\" like this 😄  Curious to know how long it takes you to train 100 epochs for each of these runs. I've mated out my batch size at 64 on a single 1080ti - it takes roughly 45 minutes per epoch (my current run has being going for 54 hours and is at 75 epochs). I'm running on 128x128 images.</p>\n\n<p>I haven't implemented apex for half precision. Maybe that's where I should focus for speed up.</p>",
      "rawMarkdown": "Thanks for sharing @drhabib. I'm very impressed how you are able to run a \"quick test\" like this 😄  Curious to know how long it takes you to train 100 epochs for each of these runs. I've mated out my batch size at 64 on a single 1080ti - it takes roughly 45 minutes per epoch (my current run has being going for 54 hours and is at 75 epochs). I'm running on 128x128 images.\n\nI haven't implemented apex for half precision. Maybe that's where I should focus for speed up.",
      "votes": null
    },
    {
      "id": "739447",
      "postDate": "02/07/2020 20:56:52",
      "content": "<p><a href=\"/robikscube\">@robikscube</a> If you are using pytorch then apex can be set up with just 3 lines of code.</p>",
      "rawMarkdown": "robikscube If you are using pytorch then apex can be set up with just 3 lines of code.",
      "votes": null
    },
    {
      "id": "739844",
      "postDate": "02/08/2020 13:44:56",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> wow interesting results. I read from papers that larger batch size would hurt the performance and I'm surprised in your case this even helped improve the score. Thanks for your info.</p>",
      "rawMarkdown": "drhabib wow interesting results. I read from papers that larger batch size would hurt the performance and I'm surprised in your case this even helped improve the score. Thanks for your info.",
      "votes": null
    },
    {
      "id": "739943",
      "postDate": "02/08/2020 15:40:34",
      "content": "<p><a href=\"/robikscube\">@robikscube</a> \nFor big batch size experiment I used google cloud because I have some credit left. On 4 <code>GPUs</code> it takes 5 min to finish one epoch. </p>\n\n<p>In my local machine it takes 20 min to finish epoch =) </p>\n\n<p><a href=\"/syoya1997\">@syoya1997</a> \n<code>Batch size</code> is something very tricky...  its another hyperparamater that can be influenced by your <code>activation function</code>, <code>weight decay</code> or <code>optimizers</code> and 'training schedule'.</p>\n\n<p><code>1- Activation Function</code> <br>\nin this paper (<a href=\"https://arxiv.org/pdf/1908.08681.pdf\">https://arxiv.org/pdf/1908.08681.pdf</a>) they compare activation function vs batch size. And you can see that for <code>ReLU</code> there is a quite some variation in terms of batch size and overall accuracy. And <code>Mish</code> with <code>Swish</code> are quite stable. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbddb87a2d2682be06180d1468e8ff644%2FScreen%20Shot%202020-02-08%20at%2010.17.49%20AM.png?generation=1581175219596572&amp;alt=media\" alt=\"\"></p>\n\n<p><code>2 - training schedule , optimizers and weight decay</code>\nIf you are using <code>one cycle learning</code> rate larger batch size can lead to better accuracy (<a href=\"https://arxiv.org/pdf/1803.09820.pdf\">https://arxiv.org/pdf/1803.09820.pdf</a>). Below is the image from this paper and you can see I get very similar results. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F761a2d7ed2c99df70822ab8dff2e0bdb%2FScreen%20Shot%202020-02-08%20at%2010.24.56%20AM.png?generation=1581175542433165&amp;alt=media\" alt=\"\"></p>\n\n<p>In contrary in this paper <a href=\"https://arxiv.org/pdf/1911.04252v2.pdf\">https://arxiv.org/pdf/1911.04252v2.pdf</a> they use to train with <code>SGD</code>. And they observed that batch size has very little influence on final accuracy <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F875bcceb8c8f5b89919cd725f66c4e94%2FScreen%20Shot%202020-02-08%20at%2010.37.30%20AM.png?generation=1581176290476008&amp;alt=media\" alt=\"\"></p>\n\n<p>But at the end of the day in normal word <code>batch size</code> doesn't matter. Since here in Kaggle we are fighting for small advantages it make sense to play with it =) </p>",
      "rawMarkdown": "robikscube \nFor big batch size experiment I used google cloud because I have some credit left. On 4 `GPUs` it takes 5 min to finish one epoch. \n\nIn my local machine it takes 20 min to finish epoch =) \n\n@syoya1997 \n`Batch size` is something very tricky...  its another hyperparamater that can be influenced by your `activation function`, `weight decay` or `optimizers` and 'training schedule'.\n\n`1- Activation Function `  \nin this paper (https://arxiv.org/pdf/1908.08681.pdf) they compare activation function vs batch size. And you can see that for `ReLU` there is a quite some variation in terms of batch size and overall accuracy. And `Mish` with `Swish` are quite stable. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbddb87a2d2682be06180d1468e8ff644%2FScreen%20Shot%202020-02-08%20at%2010.17.49%20AM.png?generation=1581175219596572&amp;alt=media)\n\n`2 - training schedule , optimizers and weight decay`\nIf you are using `one cycle learning` rate larger batch size can lead to better accuracy (https://arxiv.org/pdf/1803.09820.pdf). Below is the image from this paper and you can see I get very similar results. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F761a2d7ed2c99df70822ab8dff2e0bdb%2FScreen%20Shot%202020-02-08%20at%2010.24.56%20AM.png?generation=1581175542433165&amp;alt=media)\n\nIn contrary in this paper https://arxiv.org/pdf/1911.04252v2.pdf they use to train with `SGD`. And they observed that batch size has very little influence on final accuracy  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F875bcceb8c8f5b89919cd725f66c4e94%2FScreen%20Shot%202020-02-08%20at%2010.37.30%20AM.png?generation=1581176290476008&amp;alt=media)\n\nBut at the end of the day in normal word `batch size ` doesn't matter. Since here in Kaggle we are fighting for small advantages it make sense to play with it =)",
      "votes": null
    },
    {
      "id": "739969",
      "postDate": "02/08/2020 16:43:00",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Thanks for your sharing and it's really great to learn from you! My previous research topic is actually about those deep learning optimizers and as far as I know, larger batch size would reduce the noise in gradients under the central limit theorem assumption, which leads to poor generalization performance as the model cannot escape the sharp minima. Gradient noise help converge to flat minima and we can increase the learning to compensate for the reduced gradient noise with larger batch size. And thanks again for pointing out the paper about the activation function, I have read about the effects of final loss function on optimization but not the activation function. I would definitely take a time to read that.</p>\n\n<p>And for schedulers such as CLR, SGDR and OneCycle, they sometimes would have a better result on generalization performance. Imho it's also trying to increase the learning rate to help converge to flat minima in a more flexible and automatic way. But I don't know whether OneCycle really works? It was rejected in ICLR2018 for its irreproducibility in other tasks, although they claim the <code>super convergence</code> in their work.</p>\n\n<p>PS: I think maybe the 128 batch size is not a perfect setting with the given learning rate while bs 1024 suits that lr better in your experiment, which leads to performance improvement.</p>",
      "rawMarkdown": "drhabib Thanks for your sharing and it's really great to learn from you! My previous research topic is actually about those deep learning optimizers and as far as I know, larger batch size would reduce the noise in gradients under the central limit theorem assumption, which leads to poor generalization performance as the model cannot escape the sharp minima. Gradient noise help converge to flat minima and we can increase the learning to compensate for the reduced gradient noise with larger batch size. And thanks again for pointing out the paper about the activation function, I have read about the effects of final loss function on optimization but not the activation function. I would definitely take a time to read that.\n\nAnd for schedulers such as CLR, SGDR and OneCycle, they sometimes would have a better result on generalization performance. Imho it's also trying to increase the learning rate to help converge to flat minima in a more flexible and automatic way. But I don't know whether OneCycle really works? It was rejected in ICLR2018 for its irreproducibility in other tasks, although they claim the `super convergence` in their work.\n\nPS: I think maybe the 128 batch size is not a perfect setting with the given learning rate while bs 1024 suits that lr better in your experiment, which leads to performance improvement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 738672,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "02/06/2020 21:14:04",
      "content": "<p>Are you using half precision training ? if not check out <code>apex</code> library. You can fit almost double amount of images in your batch. </p>",
      "votes": null,
      "replies": [
        {
          "id": 738679,
          "author_name": "lightnezzofbeing",
          "author_url": "",
          "post_date": "02/06/2020 21:27:43",
          "content": "<p>Yep, I do. My current batch size for <code>se_resnext</code> is 256. I might push it even higher for smaller models like <code>b0</code>, but from my experiments, there was no difference between 256 and 512 for <code>b0</code>, but I might be mistaken, so will check it out again.</p>\n\n<p>I guess bigger batch size can be beneficial for training with <code>FocalLoss</code> or <code>ReducedFocalLoss</code> because it essentially lowers number of calls of <code>focal_loss_with_logits</code>, which you call for each of 168 + 11 + 7 classes, but for <code>CrossEntropyLoss</code> it seems like I hit some limit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738726,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/06/2020 23:22:06",
          "content": "<p><a href=\"/lightnezzofbeing\">@lightnezzofbeing</a> How batch_size affects the accuracy ? I'm warried that larger batch_size makes LB score worse.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738842,
          "author_name": "mahtabshaan",
          "author_url": "",
          "post_date": "02/07/2020 04:47:49",
          "content": "<p><a href=\"/toshik\">@toshik</a> I have also found same issue! Although I am not sure why this happens.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739141,
          "author_name": "lightnezzofbeing",
          "author_url": "",
          "post_date": "02/07/2020 13:12:45",
          "content": "<p>Hi <a href=\"/toshik\">@toshik</a> <br>\nI think what you are saying is true. However, for me 128 and 256 almost had no difference in score, but I suspect smaller batch size can yield a better result.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739200,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/07/2020 14:52:17",
          "content": "<p>just did quick test of batch size (same validation set across experiment):</p>\n\n<p><code>\nmodel: se_resnext_50\nimg_size: 128\ntrain: 100 epoch\n</code></p>\n\n<p>```\nbs: 1024\ncv: 0.9861688871</p>\n\n<p>bs: 256\ncv: 0.982144479</p>\n\n<p>bs: 128\ncv: 0.981938083\n```</p>\n\n<p>In my setup batch size worsen the performance. But It might be different depending on your optimizers and restrains. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739212,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/07/2020 15:05:58",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thank you for your information ! Larger batch_size increased you score so much !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739213,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/07/2020 15:06:56",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> did you use the same LR?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739216,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/07/2020 15:11:04",
          "content": "<p>yes! I use same <code>LR</code>,  <code>ADAM</code>,  and same <code>WD</code> and <code>one cycle policy</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739218,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/07/2020 15:12:35",
          "content": "<p>for the bs 1024, I had to use google cloud, for 256 and 128 I used my local computer. I made sure the environment is the same... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739439,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "02/07/2020 20:50:48",
          "content": "<p>Thanks for sharing <a href=\"/drhabib\">@drhabib</a>. I'm very impressed how you are able to run a \"quick test\" like this 😄  Curious to know how long it takes you to train 100 epochs for each of these runs. I've mated out my batch size at 64 on a single 1080ti - it takes roughly 45 minutes per epoch (my current run has being going for 54 hours and is at 75 epochs). I'm running on 128x128 images.</p>\n\n<p>I haven't implemented apex for half precision. Maybe that's where I should focus for speed up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739447,
          "author_name": "lightnezzofbeing",
          "author_url": "",
          "post_date": "02/07/2020 20:56:52",
          "content": "<p><a href=\"/robikscube\">@robikscube</a> If you are using pytorch then apex can be set up with just 3 lines of code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739844,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/08/2020 13:44:56",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> wow interesting results. I read from papers that larger batch size would hurt the performance and I'm surprised in your case this even helped improve the score. Thanks for your info.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739943,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/08/2020 15:40:34",
          "content": "<p><a href=\"/robikscube\">@robikscube</a> \nFor big batch size experiment I used google cloud because I have some credit left. On 4 <code>GPUs</code> it takes 5 min to finish one epoch. </p>\n\n<p>In my local machine it takes 20 min to finish epoch =) </p>\n\n<p><a href=\"/syoya1997\">@syoya1997</a> \n<code>Batch size</code> is something very tricky...  its another hyperparamater that can be influenced by your <code>activation function</code>, <code>weight decay</code> or <code>optimizers</code> and 'training schedule'.</p>\n\n<p><code>1- Activation Function</code> <br>\nin this paper (<a href=\"https://arxiv.org/pdf/1908.08681.pdf\">https://arxiv.org/pdf/1908.08681.pdf</a>) they compare activation function vs batch size. And you can see that for <code>ReLU</code> there is a quite some variation in terms of batch size and overall accuracy. And <code>Mish</code> with <code>Swish</code> are quite stable. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbddb87a2d2682be06180d1468e8ff644%2FScreen%20Shot%202020-02-08%20at%2010.17.49%20AM.png?generation=1581175219596572&amp;alt=media\" alt=\"\"></p>\n\n<p><code>2 - training schedule , optimizers and weight decay</code>\nIf you are using <code>one cycle learning</code> rate larger batch size can lead to better accuracy (<a href=\"https://arxiv.org/pdf/1803.09820.pdf\">https://arxiv.org/pdf/1803.09820.pdf</a>). Below is the image from this paper and you can see I get very similar results. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F761a2d7ed2c99df70822ab8dff2e0bdb%2FScreen%20Shot%202020-02-08%20at%2010.24.56%20AM.png?generation=1581175542433165&amp;alt=media\" alt=\"\"></p>\n\n<p>In contrary in this paper <a href=\"https://arxiv.org/pdf/1911.04252v2.pdf\">https://arxiv.org/pdf/1911.04252v2.pdf</a> they use to train with <code>SGD</code>. And they observed that batch size has very little influence on final accuracy <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F875bcceb8c8f5b89919cd725f66c4e94%2FScreen%20Shot%202020-02-08%20at%2010.37.30%20AM.png?generation=1581176290476008&amp;alt=media\" alt=\"\"></p>\n\n<p>But at the end of the day in normal word <code>batch size</code> doesn't matter. Since here in Kaggle we are fighting for small advantages it make sense to play with it =) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739969,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/08/2020 16:43:00",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thanks for your sharing and it's really great to learn from you! My previous research topic is actually about those deep learning optimizers and as far as I know, larger batch size would reduce the noise in gradients under the central limit theorem assumption, which leads to poor generalization performance as the model cannot escape the sharp minima. Gradient noise help converge to flat minima and we can increase the learning to compensate for the reduced gradient noise with larger batch size. And thanks again for pointing out the paper about the activation function, I have read about the effects of final loss function on optimization but not the activation function. I would definitely take a time to read that.</p>\n\n<p>And for schedulers such as CLR, SGDR and OneCycle, they sometimes would have a better result on generalization performance. Imho it's also trying to increase the learning rate to help converge to flat minima in a more flexible and automatic way. But I don't know whether OneCycle really works? It was rejected in ICLR2018 for its irreproducibility in other tasks, although they claim the <code>super convergence</code> in their work.</p>\n\n<p>PS: I think maybe the 128 batch size is not a perfect setting with the given learning rate while bs 1024 suits that lr better in your experiment, which leads to performance improvement.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 738694,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "02/06/2020 21:54:39",
      "content": "<p>The bottleneck is (probably) the Dataloader. Try to load all of the train data into memory. Currently, you are loading (and preprocessing) all of the images 80 (#epochs) times.\nTips:\n- use one channel (for storage) only and convert it to 3 channels in  <code>__get_item__</code>. \n- store as <code>np.uint8</code> (floats won't fit)\n- If you are using a big model, load the model+weights first, move to CUDA, and then load all of the images.\n- If you still have memory issues, then try this training method: load one (or two) parquet into memory like above; train 80 epochs, then load the next parquet (or next two parquets if you have enough ram) and train 80 more epochs</p>",
      "votes": null,
      "replies": [
        {
          "id": 739143,
          "author_name": "lightnezzofbeing",
          "author_url": "",
          "post_date": "02/07/2020 13:13:23",
          "content": "<p>Thanks! Will give it a try!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "738582": "My current setup is:\n```\nresolution: 128x128\nmodel: se_resnext50_32x4d\naugs: mixup + cutmix\nmachine: tesla p100\nbatch_size: 256\nfp16: True\n```\n\nOne epoch takes 12 mins (11 train + 1 valid). That means if I want to train for ~80 epochs I need to brake my pipeline into two kernel commits: each commit performing 40 epochs with duration approximately 8:30 hours.\n\nAn this is just a single fold model with `128x128` images.\n\nFor example, `efficientnet-b0` takes about 13 min/epoch with `224x224` images.\n\nMy `__get_item__` basically doing this:\n\n```\nimage_path = os.path.join(self.data_folder, image_id + '.png')\nimage = cv2.imread(image_path, 0)\nimage = cv2.cvtColor(image, cv2.COLOR_GRAY2BGR)\nif self.transform:\n    image = self.transform(image=image)['image']\nimage = np.transpose(image, (2, 0, 1)).astype(np.float32)\n```\n\n`transform` here is just `albumentations.Normalize`. \n\nI thought to save all processed (i.e. after `__get_item__`) images in `npy` format and load them right away during training. Unfortunately, one `npy` array takes approx `600kb` vs `2-3kb` original png, so it is not feasible to apply this method.\n\nI also tried using `npz` format, but the timings got even worse.\n\nDo you have any advice for me?",
    "738672": "Are you using half precision training ? if not check out `apex` library. You can fit almost double amount of images in your batch.",
    "738679": "Yep, I do. My current batch size for `se_resnext` is 256. I might push it even higher for smaller models like `b0`, but from my experiments, there was no difference between 256 and 512 for `b0`, but I might be mistaken, so will check it out again.\n\nI guess bigger batch size can be beneficial for training with `FocalLoss` or `ReducedFocalLoss` because it essentially lowers number of calls of `focal_loss_with_logits`, which you call for each of 168 + 11 + 7 classes, but for `CrossEntropyLoss` it seems like I hit some limit.",
    "738694": "The bottleneck is (probably) the Dataloader. Try to load all of the train data into memory. Currently, you are loading (and preprocessing) all of the images 80 (#epochs) times.\nTips:\n- use one channel (for storage) only and convert it to 3 channels in  `__get_item__`. \n- store as `np.uint8` (floats won't fit)\n- If you are using a big model, load the model+weights first, move to CUDA, and then load all of the images.\n- If you still have memory issues, then try this training method: load one (or two) parquet into memory like above; train 80 epochs, then load the next parquet (or next two parquets if you have enough ram) and train 80 more epochs",
    "738726": "lightnezzofbeing How batch_size affects the accuracy ? I'm warried that larger batch_size makes LB score worse.",
    "738842": "toshik I have also found same issue! Although I am not sure why this happens.",
    "739141": "Hi @toshik  \nI think what you are saying is true. However, for me 128 and 256 almost had no difference in score, but I suspect smaller batch size can yield a better result.",
    "739143": "Thanks! Will give it a try!",
    "739200": "just did quick test of batch size (same validation set across experiment):\n\n```\nmodel: se_resnext_50\nimg_size: 128\ntrain: 100 epoch\n```\n\n\n```\nbs: 1024\ncv: 0.9861688871\n\nbs: 256\ncv: 0.982144479\n\nbs: 128\ncv: 0.981938083\n```\n\nIn my setup batch size worsen the performance. But It might be different depending on your optimizers and restrains.",
    "739212": "drhabib Thank you for your information ! Larger batch_size increased you score so much !",
    "739213": "drhabib did you use the same LR?",
    "739216": "yes! I use same `LR`,  `ADAM`,  and same `WD` and `one cycle policy`.",
    "739218": "for the bs 1024, I had to use google cloud, for 256 and 128 I used my local computer. I made sure the environment is the same...",
    "739439": "Thanks for sharing @drhabib. I'm very impressed how you are able to run a \"quick test\" like this 😄  Curious to know how long it takes you to train 100 epochs for each of these runs. I've mated out my batch size at 64 on a single 1080ti - it takes roughly 45 minutes per epoch (my current run has being going for 54 hours and is at 75 epochs). I'm running on 128x128 images.\n\nI haven't implemented apex for half precision. Maybe that's where I should focus for speed up.",
    "739447": "robikscube If you are using pytorch then apex can be set up with just 3 lines of code.",
    "739844": "drhabib wow interesting results. I read from papers that larger batch size would hurt the performance and I'm surprised in your case this even helped improve the score. Thanks for your info.",
    "739943": "robikscube \nFor big batch size experiment I used google cloud because I have some credit left. On 4 `GPUs` it takes 5 min to finish one epoch. \n\nIn my local machine it takes 20 min to finish epoch =) \n\n@syoya1997 \n`Batch size` is something very tricky...  its another hyperparamater that can be influenced by your `activation function`, `weight decay` or `optimizers` and 'training schedule'.\n\n`1- Activation Function `  \nin this paper (https://arxiv.org/pdf/1908.08681.pdf) they compare activation function vs batch size. And you can see that for `ReLU` there is a quite some variation in terms of batch size and overall accuracy. And `Mish` with `Swish` are quite stable. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbddb87a2d2682be06180d1468e8ff644%2FScreen%20Shot%202020-02-08%20at%2010.17.49%20AM.png?generation=1581175219596572&amp;alt=media)\n\n`2 - training schedule , optimizers and weight decay`\nIf you are using `one cycle learning` rate larger batch size can lead to better accuracy (https://arxiv.org/pdf/1803.09820.pdf). Below is the image from this paper and you can see I get very similar results. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F761a2d7ed2c99df70822ab8dff2e0bdb%2FScreen%20Shot%202020-02-08%20at%2010.24.56%20AM.png?generation=1581175542433165&amp;alt=media)\n\nIn contrary in this paper https://arxiv.org/pdf/1911.04252v2.pdf they use to train with `SGD`. And they observed that batch size has very little influence on final accuracy  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F875bcceb8c8f5b89919cd725f66c4e94%2FScreen%20Shot%202020-02-08%20at%2010.37.30%20AM.png?generation=1581176290476008&amp;alt=media)\n\nBut at the end of the day in normal word `batch size ` doesn't matter. Since here in Kaggle we are fighting for small advantages it make sense to play with it =)",
    "739969": "drhabib Thanks for your sharing and it's really great to learn from you! My previous research topic is actually about those deep learning optimizers and as far as I know, larger batch size would reduce the noise in gradients under the central limit theorem assumption, which leads to poor generalization performance as the model cannot escape the sharp minima. Gradient noise help converge to flat minima and we can increase the learning to compensate for the reduced gradient noise with larger batch size. And thanks again for pointing out the paper about the activation function, I have read about the effects of final loss function on optimization but not the activation function. I would definitely take a time to read that.\n\nAnd for schedulers such as CLR, SGDR and OneCycle, they sometimes would have a better result on generalization performance. Imho it's also trying to increase the learning rate to help converge to flat minima in a more flexible and automatic way. But I don't know whether OneCycle really works? It was rejected in ICLR2018 for its irreproducibility in other tasks, although they claim the `super convergence` in their work.\n\nPS: I think maybe the 128 batch size is not a perfect setting with the given learning rate while bs 1024 suits that lr better in your experiment, which leads to performance improvement."
  },
  "source": "meta"
}