{
  "id": 173117,
  "title": "pytorch and 384x384",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173117",
  "author_name": "",
  "post_date": "2020-08-07T22:23:21.394408600Z",
  "votes": 2,
  "comment_count": 35,
  "views": 0,
  "content": "<p>I am trying to train 384x384 model on TPU with pytorch.<br>\nI was able to train 256x256 models on TPU and 384x384 on GPU on Kaggle.<br>\nWith 384 and TPU I always see: \"Your notebook tried to allocate more memory than is available.\"</p>\n<p>What is your batch size for pytorch?</p>",
  "messages": [
    {
      "id": "962186",
      "postDate": "08/07/2020 22:23:21",
      "content": "<p>I am trying to train 384x384 model on TPU with pytorch.<br>\nI was able to train 256x256 models on TPU and 384x384 on GPU on Kaggle.<br>\nWith 384 and TPU I always see: \"Your notebook tried to allocate more memory than is available.\"</p>\n<p>What is your batch size for pytorch?</p>",
      "rawMarkdown": "I am trying to train 384x384 model on TPU with pytorch.\nI was able to train 256x256 models on TPU and 384x384 on GPU on Kaggle.\nWith 384 and TPU I always see: \"Your notebook tried to allocate more memory than is available.\"\n\nWhat is your batch size for pytorch?",
      "votes": null
    },
    {
      "id": "962510",
      "postDate": "08/08/2020 07:21:05",
      "content": "<p>4 for b6 384 on 1080ti :) takes 1 hour for 1 epoch, and the result is no better than b0 256! </p>",
      "rawMarkdown": "4 for b6 384 on 1080ti :) takes 1 hour for 1 epoch, and the result is no better than b0 256!",
      "votes": null
    },
    {
      "id": "962694",
      "postDate": "08/08/2020 11:01:13",
      "content": "<p>b0 vs b6 results topic is something I don't want to discuss right now, but what I mean is 384x384 in pytorch on Kaggle kernels, I wonder maybe in this competition tensorflow works much better because data can be organized in more TPU-friendly way and with pytorch some limitations appear, people discuss 512x512 or more and I can't do 384x384 b3 on Kaggle kernel</p>",
      "rawMarkdown": "b0 vs b6 results topic is something I don't want to discuss right now, but what I mean is 384x384 in pytorch on Kaggle kernels, I wonder maybe in this competition tensorflow works much better because data can be organized in more TPU-friendly way and with pytorch some limitations appear, people discuss 512x512 or more and I can't do 384x384 b3 on Kaggle kernel",
      "votes": null
    },
    {
      "id": "962851",
      "postDate": "08/08/2020 13:41:32",
      "content": "<p>You should run your notebook in interactive mode. Then watch your notebook and see where it crashes. It may be during train, inference, or preprocess etc. This will help track down the cause.</p>",
      "rawMarkdown": "You should run your notebook in interactive mode. Then watch your notebook and see where it crashes. It may be during train, inference, or preprocess etc. This will help track down the cause.",
      "votes": null
    },
    {
      "id": "962879",
      "postDate": "08/08/2020 14:11:33",
      "content": "<p>I will try doing that on Colab, because starting TPU kernel on Kaggle takes minutes (it always installing environment) and this is last week of competition.</p>",
      "rawMarkdown": "I will try doing that on Colab, because starting TPU kernel on Kaggle takes minutes (it always installing environment) and this is last week of competition.",
      "votes": null
    },
    {
      "id": "964025",
      "postDate": "08/09/2020 14:21:34",
      "content": "<p>It might be easier to help if you post more info in here. <br>\nI used to have the same problem  using tensorflow but solved reducing the size of batch. <br>\nYou should look at your BATCH size , number of replica, GPU/TPU. </p>\n<p>Also if you train on GPU it might vary as depending on the size of the GPU the allowed batch size would vary. <br>\n(on colab you can get GPU with 8, 12 and 16 GB of memory) </p>",
      "rawMarkdown": "It might be easier to help if you post more info in here. \nI used to have the same problem  using tensorflow but solved reducing the size of batch. \nYou should look at your BATCH size , number of replica, GPU/TPU. \n\nAlso if you train on GPU it might vary as depending on the size of the GPU the allowed batch size would vary. \n(on colab you can get GPU with 8, 12 and 16 GB of memory)",
      "votes": null
    },
    {
      "id": "964043",
      "postDate": "08/09/2020 14:41:36",
      "content": "<p>Yes, that was my question - what batch size do you use on pytorch 384 on Kaggle, but as you can see pytorch is not very popular in this competition. I tried multiple times on Kaggle where I have same TPU on Colab it's random so can't compare to Kaggle.</p>\n<p>For now my greatest score comes from 256x256 trained on Kaggle GPU. </p>",
      "rawMarkdown": "Yes, that was my question - what batch size do you use on pytorch 384 on Kaggle, but as you can see pytorch is not very popular in this competition. I tried multiple times on Kaggle where I have same TPU on Colab it's random so can't compare to Kaggle.\n\nFor now my greatest score comes from 256x256 trained on Kaggle GPU.",
      "votes": null
    },
    {
      "id": "964091",
      "postDate": "08/09/2020 15:25:25",
      "content": "<p>great news</p>\n<pre><code>loaders_colab_384_b3 = {\n    \"train_batch_size\": 32,\n    \"valid_batch_size\": 32,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 575.604625, 'lr': 0.00010451362770768117}\n</code></pre>",
      "rawMarkdown": "great news\n\n```\nloaders_colab_384_b3 = {\n    \"train_batch_size\": 32,\n    \"valid_batch_size\": 32,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 575.604625, 'lr': 0.00010451362770768117}\n```",
      "votes": null
    },
    {
      "id": "964226",
      "postDate": "08/09/2020 17:11:22",
      "content": "<pre><code>Loaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: /content/drive/My Drive/KaggleLogs/session_08_09_1550/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 580.95475, 'lr': 0.00010451362770768117}\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8099368432721928, 'valid_score': 0.8418544260890589, 'train_loss': 0.08016642610127724, 'valid_loss': 0.07164809737792786, 'duration': 578.784786, 'lr': 0.00028071281016403625}\nepoch 2\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8037564002885837, 'valid_score': 0.8713453936691895, 'train_loss': 0.08079370206372652, 'valid_loss': 0.07496459415730308, 'duration': 585.470988, 'lr': 0.0005212340121216189}\nepoch 3\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8214261310940373, 'valid_score': 0.8729646759710843, 'train_loss': 0.07874054459416624, 'valid_loss': 0.06729021120159065, 'duration': 590.757463, 'lr': 0.0007614235032516977}\nepoch 4\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8226242061975689, 'valid_score': 0.8948503343788581, 'train_loss': 0.07774241340425443, 'valid_loss': 0.06472354305360246, 'duration': 585.140571, 'lr': 0.0009367167193456758}\nepoch 5\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8470639649744557, 'valid_score': 0.8916386835150915, 'train_loss': 0.07461578585322512, 'valid_loss': 0.06983104202019817, 'duration': 590.309407, 'lr': 0.000999998790010987}\nepoch 6\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8504571331704445, 'valid_score': 0.8800448928558782, 'train_loss': 0.07359173024157228, 'valid_loss': 0.07041596415288308, 'duration': 586.578736, 'lr': 0.0009872180553045955}\nepoch 7\ntrain...\nvalid...\n</code></pre>",
      "rawMarkdown": "```\n\nLoaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: /content/drive/My Drive/KaggleLogs/session_08_09_1550/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 580.95475, 'lr': 0.00010451362770768117}\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8099368432721928, 'valid_score': 0.8418544260890589, 'train_loss': 0.08016642610127724, 'valid_loss': 0.07164809737792786, 'duration': 578.784786, 'lr': 0.00028071281016403625}\nepoch 2\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8037564002885837, 'valid_score': 0.8713453936691895, 'train_loss': 0.08079370206372652, 'valid_loss': 0.07496459415730308, 'duration': 585.470988, 'lr': 0.0005212340121216189}\nepoch 3\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8214261310940373, 'valid_score': 0.8729646759710843, 'train_loss': 0.07874054459416624, 'valid_loss': 0.06729021120159065, 'duration': 590.757463, 'lr': 0.0007614235032516977}\nepoch 4\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8226242061975689, 'valid_score': 0.8948503343788581, 'train_loss': 0.07774241340425443, 'valid_loss': 0.06472354305360246, 'duration': 585.140571, 'lr': 0.0009367167193456758}\nepoch 5\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8470639649744557, 'valid_score': 0.8916386835150915, 'train_loss': 0.07461578585322512, 'valid_loss': 0.06983104202019817, 'duration': 590.309407, 'lr': 0.000999998790010987}\nepoch 6\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8504571331704445, 'valid_score': 0.8800448928558782, 'train_loss': 0.07359173024157228, 'valid_loss': 0.07041596415288308, 'duration': 586.578736, 'lr': 0.0009872180553045955}\nepoch 7\ntrain...\nvalid...\n\n\n```",
      "votes": null
    },
    {
      "id": "964230",
      "postDate": "08/09/2020 17:12:24",
      "content": "<pre><code>Exception                                 Traceback (most recent call last)\n\nipython-input-1-1dd80295c7c1 in module()\n     56 \n     57 xmp.spawn(_mp_fn, args=(config,), nprocs=8,\n---&gt; 58           start_method='fork')\n\n2 frames\n\n/usr/local/lib/python3.6/dist-packages/torch/multiprocessing/spawn.py in join(self, timeout)\n    106                 raise Exception(\n    107                     \"process %d terminated with signal %s\" %\n--&gt; 108                     (error_index, name)\n    109                 )\n    110             else:\n\nException: process 7 terminated with signal SIGKILL\n</code></pre>",
      "rawMarkdown": "```\nException                                 Traceback (most recent call last)\n\nipython-input-1-1dd80295c7c1 in module()\n     56 \n     57 xmp.spawn(_mp_fn, args=(config,), nprocs=8,\n---&gt; 58           start_method='fork')\n\n2 frames\n\n/usr/local/lib/python3.6/dist-packages/torch/multiprocessing/spawn.py in join(self, timeout)\n    106                 raise Exception(\n    107                     \"process %d terminated with signal %s\" %\n--&gt; 108                     (error_index, name)\n    109                 )\n    110             else:\n\nException: process 7 terminated with signal SIGKILL\n```",
      "votes": null
    },
    {
      "id": "964233",
      "postDate": "08/09/2020 17:13:02",
      "content": "<p>…then crash in epoch 7</p>",
      "rawMarkdown": "...then crash in epoch 7",
      "votes": null
    },
    {
      "id": "964250",
      "postDate": "08/09/2020 17:25:33",
      "content": "<p>trying now:</p>\n<pre><code>loaders_colab_384_b3 = {\n    \"train_batch_size\": 24,\n    \"valid_batch_size\": 24,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5995671565614407, 'valid_score': 0.8349309187834001, 'train_loss': 0.178136943988479, 'valid_loss': 0.07360311624320115, 'duration': 783.181116, 'lr': 0.00010446322538335191}\n</code></pre>\n<p>however it's interesting that it crashed in valid part, in all cases I could use larger valid batch size than train batch size, and code stays the same, the new part is only larger image size</p>",
      "rawMarkdown": "trying now:\n\n```\nloaders_colab_384_b3 = {\n    \"train_batch_size\": 24,\n    \"valid_batch_size\": 24,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5995671565614407, 'valid_score': 0.8349309187834001, 'train_loss': 0.178136943988479, 'valid_loss': 0.07360311624320115, 'duration': 783.181116, 'lr': 0.00010446322538335191}\n```\n\nhowever it's interesting that it crashed in valid part, in all cases I could use larger valid batch size than train batch size, and code stays the same, the new part is only larger image size",
      "votes": null
    },
    {
      "id": "964411",
      "postDate": "08/09/2020 20:41:05",
      "content": "<p>OK, tested 20 epochs on Colab and 2 epochs on Kaggle, looks like it does work with 24/24</p>\n<pre><code>Loaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: ./session_08_09_2004/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.6345018158930966, 'valid_score': 0.8512216785867384, 'train_loss': 0.15346827832700075, 'valid_loss': 0.07005745299379615, 'duration': 553.14053, 'lr': 0.0008052050401102377}\nsaving model\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8085982473311003, 'valid_score': 0.8889353661115494, 'train_loss': 0.07864298790214679, 'valid_loss': 0.0652915458293522, 'duration': 416.570716, 'lr': 7.307237815654868e-08}\nsaving model\n</code></pre>\n<p>so I am able to calculate 20 epochs within 3 hours of Kaggle TPU time</p>",
      "rawMarkdown": "OK, tested 20 epochs on Colab and 2 epochs on Kaggle, looks like it does work with 24/24\n\n```\nLoaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: ./session_08_09_2004/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.6345018158930966, 'valid_score': 0.8512216785867384, 'train_loss': 0.15346827832700075, 'valid_loss': 0.07005745299379615, 'duration': 553.14053, 'lr': 0.0008052050401102377}\nsaving model\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8085982473311003, 'valid_score': 0.8889353661115494, 'train_loss': 0.07864298790214679, 'valid_loss': 0.0652915458293522, 'duration': 416.570716, 'lr': 7.307237815654868e-08}\nsaving model\n```\n\nso I am able to calculate 20 epochs within 3 hours of Kaggle TPU time",
      "votes": null
    },
    {
      "id": "964439",
      "postDate": "08/09/2020 21:56:44",
      "content": "<p>I don't know much about Kaggle kernels/TPU, or how much memory you all get.  But yes, with validation you are not calculating gradients so. you can use ALOT more in your batch size typically.</p>\n\n<p>I have 4 GPU's and I typically run the following for Image Size of 384</p>\n\n<p>B1: train batch size: 25\n      validation batch size:  512 \nB4 train batch size: 14\n      validation batch size: 256.</p>\n\n<p>You need to make SURE you are running your model in inference mode when doing validation!!!!  For example, in PyTorch, we put the model into inference mode by doing <code>model.eval()</code>.  What this does is it stops things like <code>BatchNorm</code> and <code>DropOut</code>.  Additionally, we also wrap our validation routine in a <code>with torch.no_grad()</code> , which stops the calculation of gradients.  If you are not doing this, you will need to use the SAME batch size for train and validation, which seems to be your situation.  Validation should be ALOT quicker than train, and you should be able to use at least a batch size of 10-20x train.  Just make sure you are setting things proper for your framework.</p>",
      "rawMarkdown": "I don't know much about Kaggle kernels/TPU, or how much memory you all get.  But yes, with validation you are not calculating gradients so. you can use ALOT more in your batch size typically.\n\nI have 4 GPU's and I typically run the following for Image Size of 384\n\nB1: train batch size: 25\n      validation batch size:  512 \nB4 train batch size: 14\n      validation batch size: 256.\n\nYou need to make SURE you are running your model in inference mode when doing validation!!!!  For example, in PyTorch, we put the model into inference mode by doing `model.eval()`.  What this does is it stops things like `BatchNorm` and `DropOut`.  Additionally, we also wrap our validation routine in a `with torch.no_grad()` , which stops the calculation of gradients.  If you are not doing this, you will need to use the SAME batch size for train and validation, which seems to be your situation.  Validation should be ALOT quicker than train, and you should be able to use at least a batch size of 10-20x train.  Just make sure you are setting things proper for your framework.",
      "votes": null
    },
    {
      "id": "964487",
      "postDate": "08/09/2020 23:55:08",
      "content": "<p>What framework are you using? Why is <code>num_workers</code> set to 0? That's <strong>never</strong> the optimal value.</p>",
      "rawMarkdown": "What framework are you using? Why is `num_workers` set to 0? That's **never** the optimal value.",
      "votes": null
    },
    {
      "id": "964489",
      "postDate": "08/09/2020 23:57:30",
      "content": "<p>You can't just change model complexity and image size and leave everything else the same and expect a better result.  If you change to a more complex model, with larger images, your batch size will change, and therefore your learning rate should change at a minimum.</p>",
      "rawMarkdown": "You can't just change model complexity and image size and leave everything else the same and expect a better result.  If you change to a more complex model, with larger images, your batch size will change, and therefore your learning rate should change at a minimum.",
      "votes": null
    },
    {
      "id": "964628",
      "postDate": "08/10/2020 04:11:51",
      "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> <br>\nDo you know what is the dependence between batch size and model size? Is it possible to quickly estimate what is the \"good\" batch size for particular model and img size? </p>",
      "rawMarkdown": "brianfeeny \nDo you know what is the dependence between batch size and model size? Is it possible to quickly estimate what is the \"good\" batch size for particular model and img size?",
      "votes": null
    },
    {
      "id": "964630",
      "postDate": "08/10/2020 04:12:43",
      "content": "<p>As a matter of fact I also have quite good results on th 256 images. This is good size for the experimentation.</p>",
      "rawMarkdown": "As a matter of fact I also have quite good results on th 256 images. This is good size for the experimentation.",
      "votes": null
    },
    {
      "id": "964631",
      "postDate": "08/10/2020 04:13:31",
      "content": "<p>It  seems there is some memory leak when using pytorch? </p>",
      "rawMarkdown": "It  seems there is some memory leak when using pytorch?",
      "votes": null
    },
    {
      "id": "964695",
      "postDate": "08/10/2020 05:20:06",
      "content": "<p><a href=\"/janidziak\">@janidziak</a> just try stuff.  I have seen people do formulas but I always just set my batch to the max I can and then do some experimentation.  Using a cyclic scheduler or a \"learning rate finder\" is a good way to find something to start with.</p>",
      "rawMarkdown": "janidziak just try stuff.  I have seen people do formulas but I always just set my batch to the max I can and then do some experimentation.  Using a cyclic scheduler or a \"learning rate finder\" is a good way to find something to start with.",
      "votes": null
    },
    {
      "id": "964751",
      "postDate": "08/10/2020 06:28:00",
      "content": "<p>The main bottleneck for Pytorch Kaggle Kernel TPU is insufficient CPU cores and RAM when the data loader is being processed. A minimum requirement of at least 16 cores and high RAM in order to maximize the potential of TPU in Pytorch.<br>\nI would advise training a model with image size 256x256 if you intend to train a model in Kaggle TPU kernel</p>",
      "rawMarkdown": "The main bottleneck for Pytorch Kaggle Kernel TPU is insufficient CPU cores and RAM when the data loader is being processed. A minimum requirement of at least 16 cores and high RAM in order to maximize the potential of TPU in Pytorch.\nI would advise training a model with image size 256x256 if you intend to train a model in Kaggle TPU kernel",
      "votes": null
    },
    {
      "id": "964837",
      "postDate": "08/10/2020 08:01:32",
      "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> <br>\njest I think it is implemented correctly:</p>\n<pre><code>train_result = train_step(config_step)\n(...)\nwith torch.no_grad():\n     valid_result = train_step(config_step)\n</code></pre>\n<pre><code>    if (train_mode):\n        model.train() \n    else:\n        model.eval()         \n</code></pre>\n<p>so I wonder how can you use 256 valid when my code fails for 32</p>",
      "rawMarkdown": "brianfeeny \njest I think it is implemented correctly:\n\n```\ntrain_result = train_step(config_step)\n(...)\nwith torch.no_grad():\n     valid_result = train_step(config_step)\n\n```\n\n```\n    if (train_mode):\n        model.train() \n    else:\n        model.eval()         \n\n```\n\nso I wonder how can you use 256 valid when my code fails for 32",
      "votes": null
    },
    {
      "id": "964841",
      "postDate": "08/10/2020 08:02:39",
      "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> <br>\ndo you use TPU? I found that 0 is optimal for TPU</p>",
      "rawMarkdown": "brianfeeny \ndo you use TPU? I found that 0 is optimal for TPU",
      "votes": null
    },
    {
      "id": "964849",
      "postDate": "08/10/2020 08:08:56",
      "content": "<p>Hard to say why, but obviously something is wrong. I assume you have two dataloaders and are trying to set a different batch_size in each?</p>",
      "rawMarkdown": "Hard to say why, but obviously something is wrong. I assume you have two dataloaders and are trying to set a different batch_size in each?",
      "votes": null
    },
    {
      "id": "964854",
      "postDate": "08/10/2020 08:10:02",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> I don't use TPU, I use a GPU in my lab.  num_workers has (or should have) nothing to do with TPU vs GPU, it has to do with your CPU, as num_workers are parallel dataloading processes which load data into CPU memory.</p>",
      "rawMarkdown": "jacekpoplawski I don't use TPU, I use a GPU in my lab.  num_workers has (or should have) nothing to do with TPU vs GPU, it has to do with your CPU, as num_workers are parallel dataloading processes which load data into CPU memory.",
      "votes": null
    },
    {
      "id": "964869",
      "postDate": "08/10/2020 08:22:13",
      "content": "<p>As I said I always set valid batches higher, sometimes 2x higher and they were working. Only problem appears on image size 384 on Kaggle TPU and Colab TPU.</p>",
      "rawMarkdown": "As I said I always set valid batches higher, sometimes 2x higher and they were working. Only problem appears on image size 384 on Kaggle TPU and Colab TPU.",
      "votes": null
    },
    {
      "id": "965106",
      "postDate": "08/10/2020 11:43:22",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  Jacek, do you use mixed precision in Pytorch? I  do not use TPU, so can say for Colab GPU only.\nWith mixed precision I managed to run 384x384 B6 model with batch size 24. Without mixed precision, max batch size I was able to train the same model with was 12.</p>",
      "rawMarkdown": "jacekpoplawski  Jacek, do you use mixed precision in Pytorch? I  do not use TPU, so can say for Colab GPU only.\nWith mixed precision I managed to run 384x384 B6 model with batch size 24. Without mixed precision, max batch size I was able to train the same model with was 12.",
      "votes": null
    },
    {
      "id": "965122",
      "postDate": "08/10/2020 11:57:54",
      "content": "<p>I don't have problem on GPU, no I still was not able to learn how to use mixed precision with pretrained models.</p>",
      "rawMarkdown": "I don't have problem on GPU, no I still was not able to learn how to use mixed precision with pretrained models.",
      "votes": null
    },
    {
      "id": "965141",
      "postDate": "08/10/2020 12:10:41",
      "content": "<p>Well, mixed precision could help you with TPU issue also I guess.</p>\n\n<p>Below is a simplified code from Abhishek's repo there I added mixed precision support for GPU.</p>\n\n<p>```\nclass Engine:\n    @staticmethod\n    def train(\n        data_loader,\n        model,\n        optimizer,\n        device,\n        scheduler=None,\n        accumulation_steps=1,\n        fp16=True, # fp16 = True - if we want to apply mixed precision\n    ):</p>\n\n<pre><code>    losses = AverageMeter()\n    final_predictions = []\n    model.train()\n    if accumulation_steps &amp;gt; 1:\n        optimizer.zero_grad()\n\n    if fp16:\n      scaler = torch.cuda.amp.GradScaler()\n\n    for b_idx, data in enumerate(data_loader):\n        for key, value in data.items():\n            data[key] = value.to(device)\n        if accumulation_steps == 1 and b_idx == 0:\n            optimizer.zero_grad()\n        if fp16:    \n            with torch.cuda.amp.autocast():    \n                predictions, loss = model(**data)\n        else:\n            predictions, loss = model(**data)      \n        predictions = predictions.detach().cpu()\n        final_predictions.append(predictions) \n\n        with torch.set_grad_enabled(True):\n            if fp16:\n                scaler.scale(loss).backward()                   \n            else:\n                loss.backward()\n            if (b_idx + 1) % accumulation_steps == 0:\n                if fp16:\n                    scaler.step(optimizer)\n                    scaler.update()\n                else:     \n                    optimizer.step()\n                if scheduler is not None:\n                     scheduler.step()\n                if b_idx &amp;gt; 0:\n                    optimizer.zero_grad()\n\n        losses.update(loss.item(), data_loader.batch_size)\n\n    return final_predictions, losses.avg\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Well, mixed precision could help you with TPU issue also I guess.\n\nBelow is a simplified code from Abhishek's repo there I added mixed precision support for GPU.\n\n```\nclass Engine:\n    @staticmethod\n    def train(\n        data_loader,\n        model,\n        optimizer,\n        device,\n        scheduler=None,\n        accumulation_steps=1,\n        fp16=True, # fp16 = True - if we want to apply mixed precision\n    ):\n\n        losses = AverageMeter()\n        final_predictions = []\n        model.train()\n        if accumulation_steps &gt; 1:\n            optimizer.zero_grad()\n\n        if fp16:\n          scaler = torch.cuda.amp.GradScaler()\n\n        for b_idx, data in enumerate(data_loader):\n            for key, value in data.items():\n                data[key] = value.to(device)\n            if accumulation_steps == 1 and b_idx == 0:\n                optimizer.zero_grad()\n            if fp16:    \n                with torch.cuda.amp.autocast():    \n                    predictions, loss = model(**data)\n            else:\n                predictions, loss = model(**data)      \n            predictions = predictions.detach().cpu()\n            final_predictions.append(predictions) \n\n            with torch.set_grad_enabled(True):\n                if fp16:\n                    scaler.scale(loss).backward()                   \n                else:\n                    loss.backward()\n                if (b_idx + 1) % accumulation_steps == 0:\n                    if fp16:\n                        scaler.step(optimizer)\n                        scaler.update()\n                    else:     \n                        optimizer.step()\n                    if scheduler is not None:\n                         scheduler.step()\n                    if b_idx &gt; 0:\n                        optimizer.zero_grad()\n\n            losses.update(loss.item(), data_loader.batch_size)\n\n        return final_predictions, losses.avg\n```",
      "votes": null
    },
    {
      "id": "965144",
      "postDate": "08/10/2020 12:12:38",
      "content": "<p>can you explain what happens with weights during mixed precision? can you just use efficientnet weights and everything works?</p>",
      "rawMarkdown": "can you explain what happens with weights during mixed precision? can you just use efficientnet weights and everything works?",
      "votes": null
    },
    {
      "id": "965171",
      "postDate": "08/10/2020 12:35:16",
      "content": "<p>I ran 2 tests, with and without mixed precision. \n1. CV\\LB with mixed precision ON are almost identical to those with mixed precision OFF.\n2. I ran in on P100 that doesn't support enhanced mixed precision matrices multiplication like more modern GPUs do. But P100 does support memory optimization. So you'll reduce your memory 2x.\n3. speaking of weights, mixed precision optimization uses loss scaling - to preserve very small gradient values, it scales loss value (and therefore gradient values) up and later scales everything back down.</p>\n\n<blockquote>\n  <p>Maintain a master copy of weights in FP32\n  For each iteration:\n  Make an FP16 copy of the weights\n  Forward propagation (FP16 weights and activations)\n  Multiply the resulting loss with the scaling factor S\n  Backward propagation (FP16 weights, activations, and their gradients)\n  Multiply the weight gradient with 1/S\n  Complete the weight update (including gradient clipping, etc.)</p>\n</blockquote>",
      "rawMarkdown": "I ran 2 tests, with and without mixed precision. \n1. CV\\LB with mixed precision ON are almost identical to those with mixed precision OFF.\n2. I ran in on P100 that doesn't support enhanced mixed precision matrices multiplication like more modern GPUs do. But P100 does support memory optimization. So you'll reduce your memory 2x.\n3. speaking of weights, mixed precision optimization uses loss scaling - to preserve very small gradient values, it scales loss value (and therefore gradient values) up and later scales everything back down.\n\n&gt;Maintain a master copy of weights in FP32\nFor each iteration:\nMake an FP16 copy of the weights\nForward propagation (FP16 weights and activations)\nMultiply the resulting loss with the scaling factor S\nBackward propagation (FP16 weights, activations, and their gradients)\nMultiply the weight gradient with 1/S\nComplete the weight update (including gradient clipping, etc.)",
      "votes": null
    },
    {
      "id": "965198",
      "postDate": "08/10/2020 13:06:29",
      "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> <br>\nah so weights are in original 32-bit</p>\n<p>what about BatchNormalization issue? do you need to do anything with efficientnet?</p>",
      "rawMarkdown": "dunklerwald \nah so weights are in original 32-bit\n\nwhat about BatchNormalization issue? do you need to do anything with efficientnet?",
      "votes": null
    },
    {
      "id": "965220",
      "postDate": "08/10/2020 13:29:07",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> maybe I would need to do something but in fact I do nothing:-) and all layers are unfreezed.</p>",
      "rawMarkdown": "jacekpoplawski maybe I would need to do something but in fact I do nothing:-) and all layers are unfreezed.",
      "votes": null
    },
    {
      "id": "965237",
      "postDate": "08/10/2020 13:40:40",
      "content": "<p>I read that you need to replace BN with different implementation, but you said you did nothing and it works, right? What are your results? I assume CV/LB is same, so what about time of epoch (normal vs mixed)?</p>",
      "rawMarkdown": "I read that you need to replace BN with different implementation, but you said you did nothing and it works, right? What are your results? I assume CV/LB is same, so what about time of epoch (normal vs mixed)?",
      "votes": null
    },
    {
      "id": "965273",
      "postDate": "08/10/2020 14:10:16",
      "content": "<p>Right, I did nothing with BN. CV/LB are the same. Time-wise, i didn't observe any significant difference sadly. But must admit, I didn't track timing accurately. As I said earlier, most likely, that's because I was running it on Google Colab P100 which is not able to take advantage of optimized mixed precision matrices multiplications. If you run mixed precision  on Kaggle TPU v3, you should expect substantial time savings.</p>",
      "rawMarkdown": "Right, I did nothing with BN. CV/LB are the same. Time-wise, i didn't observe any significant difference sadly. But must admit, I didn't track timing accurately. As I said earlier, most likely, that's because I was running it on Google Colab P100 which is not able to take advantage of optimized mixed precision matrices multiplications. If you run mixed precision  on Kaggle TPU v3, you should expect substantial time savings.",
      "votes": null
    },
    {
      "id": "968611",
      "postDate": "08/13/2020 06:25:00",
      "content": "<p><a href=\"https://www.kaggle.com/sig220932\" target=\"_blank\">@sig220932</a> thanks for sharing that. i tired reduce lr and it helped for large models. thanks again. </p>",
      "rawMarkdown": "sig220932 thanks for sharing that. i tired reduce lr and it helped for large models. thanks again.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 962851,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/08/2020 13:41:32",
      "content": "<p>You should run your notebook in interactive mode. Then watch your notebook and see where it crashes. It may be during train, inference, or preprocess etc. This will help track down the cause.</p>",
      "votes": null,
      "replies": [
        {
          "id": 962879,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/08/2020 14:11:33",
          "content": "<p>I will try doing that on Colab, because starting TPU kernel on Kaggle takes minutes (it always installing environment) and this is last week of competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964091,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/09/2020 15:25:25",
          "content": "<p>great news</p>\n<pre><code>loaders_colab_384_b3 = {\n    \"train_batch_size\": 32,\n    \"valid_batch_size\": 32,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 575.604625, 'lr': 0.00010451362770768117}\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964233,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/09/2020 17:13:02",
          "content": "<p>…then crash in epoch 7</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964487,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/09/2020 23:55:08",
          "content": "<p>What framework are you using? Why is <code>num_workers</code> set to 0? That's <strong>never</strong> the optimal value.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964841,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 08:02:39",
          "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> <br>\ndo you use TPU? I found that 0 is optimal for TPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964854,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 08:10:02",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> I don't use TPU, I use a GPU in my lab.  num_workers has (or should have) nothing to do with TPU vs GPU, it has to do with your CPU, as num_workers are parallel dataloading processes which load data into CPU memory.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 964025,
      "author_name": "janidziak",
      "author_url": "",
      "post_date": "08/09/2020 14:21:34",
      "content": "<p>It might be easier to help if you post more info in here. <br>\nI used to have the same problem  using tensorflow but solved reducing the size of batch. <br>\nYou should look at your BATCH size , number of replica, GPU/TPU. </p>\n<p>Also if you train on GPU it might vary as depending on the size of the GPU the allowed batch size would vary. <br>\n(on colab you can get GPU with 8, 12 and 16 GB of memory) </p>",
      "votes": null,
      "replies": [
        {
          "id": 964043,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/09/2020 14:41:36",
          "content": "<p>Yes, that was my question - what batch size do you use on pytorch 384 on Kaggle, but as you can see pytorch is not very popular in this competition. I tried multiple times on Kaggle where I have same TPU on Colab it's random so can't compare to Kaggle.</p>\n<p>For now my greatest score comes from 256x256 trained on Kaggle GPU. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964630,
          "author_name": "janidziak",
          "author_url": "",
          "post_date": "08/10/2020 04:12:43",
          "content": "<p>As a matter of fact I also have quite good results on th 256 images. This is good size for the experimentation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 964226,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "08/09/2020 17:11:22",
      "content": "<pre><code>Loaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: /content/drive/My Drive/KaggleLogs/session_08_09_1550/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 580.95475, 'lr': 0.00010451362770768117}\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8099368432721928, 'valid_score': 0.8418544260890589, 'train_loss': 0.08016642610127724, 'valid_loss': 0.07164809737792786, 'duration': 578.784786, 'lr': 0.00028071281016403625}\nepoch 2\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8037564002885837, 'valid_score': 0.8713453936691895, 'train_loss': 0.08079370206372652, 'valid_loss': 0.07496459415730308, 'duration': 585.470988, 'lr': 0.0005212340121216189}\nepoch 3\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8214261310940373, 'valid_score': 0.8729646759710843, 'train_loss': 0.07874054459416624, 'valid_loss': 0.06729021120159065, 'duration': 590.757463, 'lr': 0.0007614235032516977}\nepoch 4\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8226242061975689, 'valid_score': 0.8948503343788581, 'train_loss': 0.07774241340425443, 'valid_loss': 0.06472354305360246, 'duration': 585.140571, 'lr': 0.0009367167193456758}\nepoch 5\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8470639649744557, 'valid_score': 0.8916386835150915, 'train_loss': 0.07461578585322512, 'valid_loss': 0.06983104202019817, 'duration': 590.309407, 'lr': 0.000999998790010987}\nepoch 6\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8504571331704445, 'valid_score': 0.8800448928558782, 'train_loss': 0.07359173024157228, 'valid_loss': 0.07041596415288308, 'duration': 586.578736, 'lr': 0.0009872180553045955}\nepoch 7\ntrain...\nvalid...\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 964230,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/09/2020 17:12:24",
          "content": "<pre><code>Exception                                 Traceback (most recent call last)\n\nipython-input-1-1dd80295c7c1 in module()\n     56 \n     57 xmp.spawn(_mp_fn, args=(config,), nprocs=8,\n---&gt; 58           start_method='fork')\n\n2 frames\n\n/usr/local/lib/python3.6/dist-packages/torch/multiprocessing/spawn.py in join(self, timeout)\n    106                 raise Exception(\n    107                     \"process %d terminated with signal %s\" %\n--&gt; 108                     (error_index, name)\n    109                 )\n    110             else:\n\nException: process 7 terminated with signal SIGKILL\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964631,
          "author_name": "janidziak",
          "author_url": "",
          "post_date": "08/10/2020 04:13:31",
          "content": "<p>It  seems there is some memory leak when using pytorch? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 964250,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "08/09/2020 17:25:33",
      "content": "<p>trying now:</p>\n<pre><code>loaders_colab_384_b3 = {\n    \"train_batch_size\": 24,\n    \"valid_batch_size\": 24,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5995671565614407, 'valid_score': 0.8349309187834001, 'train_loss': 0.178136943988479, 'valid_loss': 0.07360311624320115, 'duration': 783.181116, 'lr': 0.00010446322538335191}\n</code></pre>\n<p>however it's interesting that it crashed in valid part, in all cases I could use larger valid batch size than train batch size, and code stays the same, the new part is only larger image size</p>",
      "votes": null,
      "replies": [
        {
          "id": 964411,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/09/2020 20:41:05",
          "content": "<p>OK, tested 20 epochs on Colab and 2 epochs on Kaggle, looks like it does work with 24/24</p>\n<pre><code>Loaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: ./session_08_09_2004/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.6345018158930966, 'valid_score': 0.8512216785867384, 'train_loss': 0.15346827832700075, 'valid_loss': 0.07005745299379615, 'duration': 553.14053, 'lr': 0.0008052050401102377}\nsaving model\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8085982473311003, 'valid_score': 0.8889353661115494, 'train_loss': 0.07864298790214679, 'valid_loss': 0.0652915458293522, 'duration': 416.570716, 'lr': 7.307237815654868e-08}\nsaving model\n</code></pre>\n<p>so I am able to calculate 20 epochs within 3 hours of Kaggle TPU time</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964439,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/09/2020 21:56:44",
          "content": "<p>I don't know much about Kaggle kernels/TPU, or how much memory you all get.  But yes, with validation you are not calculating gradients so. you can use ALOT more in your batch size typically.</p>\n\n<p>I have 4 GPU's and I typically run the following for Image Size of 384</p>\n\n<p>B1: train batch size: 25\n      validation batch size:  512 \nB4 train batch size: 14\n      validation batch size: 256.</p>\n\n<p>You need to make SURE you are running your model in inference mode when doing validation!!!!  For example, in PyTorch, we put the model into inference mode by doing <code>model.eval()</code>.  What this does is it stops things like <code>BatchNorm</code> and <code>DropOut</code>.  Additionally, we also wrap our validation routine in a <code>with torch.no_grad()</code> , which stops the calculation of gradients.  If you are not doing this, you will need to use the SAME batch size for train and validation, which seems to be your situation.  Validation should be ALOT quicker than train, and you should be able to use at least a batch size of 10-20x train.  Just make sure you are setting things proper for your framework.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964628,
          "author_name": "janidziak",
          "author_url": "",
          "post_date": "08/10/2020 04:11:51",
          "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> <br>\nDo you know what is the dependence between batch size and model size? Is it possible to quickly estimate what is the \"good\" batch size for particular model and img size? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964695,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 05:20:06",
          "content": "<p><a href=\"/janidziak\">@janidziak</a> just try stuff.  I have seen people do formulas but I always just set my batch to the max I can and then do some experimentation.  Using a cyclic scheduler or a \"learning rate finder\" is a good way to find something to start with.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964837,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 08:01:32",
          "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> <br>\njest I think it is implemented correctly:</p>\n<pre><code>train_result = train_step(config_step)\n(...)\nwith torch.no_grad():\n     valid_result = train_step(config_step)\n</code></pre>\n<pre><code>    if (train_mode):\n        model.train() \n    else:\n        model.eval()         \n</code></pre>\n<p>so I wonder how can you use 256 valid when my code fails for 32</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964849,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/10/2020 08:08:56",
          "content": "<p>Hard to say why, but obviously something is wrong. I assume you have two dataloaders and are trying to set a different batch_size in each?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964869,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 08:22:13",
          "content": "<p>As I said I always set valid batches higher, sometimes 2x higher and they were working. Only problem appears on image size 384 on Kaggle TPU and Colab TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 964751,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "08/10/2020 06:28:00",
      "content": "<p>The main bottleneck for Pytorch Kaggle Kernel TPU is insufficient CPU cores and RAM when the data loader is being processed. A minimum requirement of at least 16 cores and high RAM in order to maximize the potential of TPU in Pytorch.<br>\nI would advise training a model with image size 256x256 if you intend to train a model in Kaggle TPU kernel</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962510,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "08/08/2020 07:21:05",
      "content": "<p>4 for b6 384 on 1080ti :) takes 1 hour for 1 epoch, and the result is no better than b0 256! </p>",
      "votes": null,
      "replies": [
        {
          "id": 962694,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/08/2020 11:01:13",
          "content": "<p>b0 vs b6 results topic is something I don't want to discuss right now, but what I mean is 384x384 in pytorch on Kaggle kernels, I wonder maybe in this competition tensorflow works much better because data can be organized in more TPU-friendly way and with pytorch some limitations appear, people discuss 512x512 or more and I can't do 384x384 b3 on Kaggle kernel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964489,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/09/2020 23:57:30",
          "content": "<p>You can't just change model complexity and image size and leave everything else the same and expect a better result.  If you change to a more complex model, with larger images, your batch size will change, and therefore your learning rate should change at a minimum.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 968611,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "08/13/2020 06:25:00",
          "content": "<p><a href=\"https://www.kaggle.com/sig220932\" target=\"_blank\">@sig220932</a> thanks for sharing that. i tired reduce lr and it helped for large models. thanks again. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 965106,
      "author_name": "dunklerwald",
      "author_url": "",
      "post_date": "08/10/2020 11:43:22",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  Jacek, do you use mixed precision in Pytorch? I  do not use TPU, so can say for Colab GPU only.\nWith mixed precision I managed to run 384x384 B6 model with batch size 24. Without mixed precision, max batch size I was able to train the same model with was 12.</p>",
      "votes": null,
      "replies": [
        {
          "id": 965122,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 11:57:54",
          "content": "<p>I don't have problem on GPU, no I still was not able to learn how to use mixed precision with pretrained models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965141,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/10/2020 12:10:41",
          "content": "<p>Well, mixed precision could help you with TPU issue also I guess.</p>\n\n<p>Below is a simplified code from Abhishek's repo there I added mixed precision support for GPU.</p>\n\n<p>```\nclass Engine:\n    @staticmethod\n    def train(\n        data_loader,\n        model,\n        optimizer,\n        device,\n        scheduler=None,\n        accumulation_steps=1,\n        fp16=True, # fp16 = True - if we want to apply mixed precision\n    ):</p>\n\n<pre><code>    losses = AverageMeter()\n    final_predictions = []\n    model.train()\n    if accumulation_steps &amp;gt; 1:\n        optimizer.zero_grad()\n\n    if fp16:\n      scaler = torch.cuda.amp.GradScaler()\n\n    for b_idx, data in enumerate(data_loader):\n        for key, value in data.items():\n            data[key] = value.to(device)\n        if accumulation_steps == 1 and b_idx == 0:\n            optimizer.zero_grad()\n        if fp16:    \n            with torch.cuda.amp.autocast():    \n                predictions, loss = model(**data)\n        else:\n            predictions, loss = model(**data)      \n        predictions = predictions.detach().cpu()\n        final_predictions.append(predictions) \n\n        with torch.set_grad_enabled(True):\n            if fp16:\n                scaler.scale(loss).backward()                   \n            else:\n                loss.backward()\n            if (b_idx + 1) % accumulation_steps == 0:\n                if fp16:\n                    scaler.step(optimizer)\n                    scaler.update()\n                else:     \n                    optimizer.step()\n                if scheduler is not None:\n                     scheduler.step()\n                if b_idx &amp;gt; 0:\n                    optimizer.zero_grad()\n\n        losses.update(loss.item(), data_loader.batch_size)\n\n    return final_predictions, losses.avg\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965144,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 12:12:38",
          "content": "<p>can you explain what happens with weights during mixed precision? can you just use efficientnet weights and everything works?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965171,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/10/2020 12:35:16",
          "content": "<p>I ran 2 tests, with and without mixed precision. \n1. CV\\LB with mixed precision ON are almost identical to those with mixed precision OFF.\n2. I ran in on P100 that doesn't support enhanced mixed precision matrices multiplication like more modern GPUs do. But P100 does support memory optimization. So you'll reduce your memory 2x.\n3. speaking of weights, mixed precision optimization uses loss scaling - to preserve very small gradient values, it scales loss value (and therefore gradient values) up and later scales everything back down.</p>\n\n<blockquote>\n  <p>Maintain a master copy of weights in FP32\n  For each iteration:\n  Make an FP16 copy of the weights\n  Forward propagation (FP16 weights and activations)\n  Multiply the resulting loss with the scaling factor S\n  Backward propagation (FP16 weights, activations, and their gradients)\n  Multiply the weight gradient with 1/S\n  Complete the weight update (including gradient clipping, etc.)</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965198,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 13:06:29",
          "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> <br>\nah so weights are in original 32-bit</p>\n<p>what about BatchNormalization issue? do you need to do anything with efficientnet?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965220,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/10/2020 13:29:07",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> maybe I would need to do something but in fact I do nothing:-) and all layers are unfreezed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965237,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "08/10/2020 13:40:40",
          "content": "<p>I read that you need to replace BN with different implementation, but you said you did nothing and it works, right? What are your results? I assume CV/LB is same, so what about time of epoch (normal vs mixed)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965273,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/10/2020 14:10:16",
          "content": "<p>Right, I did nothing with BN. CV/LB are the same. Time-wise, i didn't observe any significant difference sadly. But must admit, I didn't track timing accurately. As I said earlier, most likely, that's because I was running it on Google Colab P100 which is not able to take advantage of optimized mixed precision matrices multiplications. If you run mixed precision  on Kaggle TPU v3, you should expect substantial time savings.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "962186": "I am trying to train 384x384 model on TPU with pytorch.\nI was able to train 256x256 models on TPU and 384x384 on GPU on Kaggle.\nWith 384 and TPU I always see: \"Your notebook tried to allocate more memory than is available.\"\n\nWhat is your batch size for pytorch?",
    "962510": "4 for b6 384 on 1080ti :) takes 1 hour for 1 epoch, and the result is no better than b0 256!",
    "962694": "b0 vs b6 results topic is something I don't want to discuss right now, but what I mean is 384x384 in pytorch on Kaggle kernels, I wonder maybe in this competition tensorflow works much better because data can be organized in more TPU-friendly way and with pytorch some limitations appear, people discuss 512x512 or more and I can't do 384x384 b3 on Kaggle kernel",
    "962851": "You should run your notebook in interactive mode. Then watch your notebook and see where it crashes. It may be during train, inference, or preprocess etc. This will help track down the cause.",
    "962879": "I will try doing that on Colab, because starting TPU kernel on Kaggle takes minutes (it always installing environment) and this is last week of competition.",
    "964025": "It might be easier to help if you post more info in here. \nI used to have the same problem  using tensorflow but solved reducing the size of batch. \nYou should look at your BATCH size , number of replica, GPU/TPU. \n\nAlso if you train on GPU it might vary as depending on the size of the GPU the allowed batch size would vary. \n(on colab you can get GPU with 8, 12 and 16 GB of memory)",
    "964043": "Yes, that was my question - what batch size do you use on pytorch 384 on Kaggle, but as you can see pytorch is not very popular in this competition. I tried multiple times on Kaggle where I have same TPU on Colab it's random so can't compare to Kaggle.\n\nFor now my greatest score comes from 256x256 trained on Kaggle GPU.",
    "964091": "great news\n\n```\nloaders_colab_384_b3 = {\n    \"train_batch_size\": 32,\n    \"valid_batch_size\": 32,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 575.604625, 'lr': 0.00010451362770768117}\n```",
    "964226": "```\n\nLoaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: /content/drive/My Drive/KaggleLogs/session_08_09_1550/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5694665201443196, 'valid_score': 0.7578702308377641, 'train_loss': 0.20312086954699063, 'valid_loss': 0.08859821680056698, 'duration': 580.95475, 'lr': 0.00010451362770768117}\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8099368432721928, 'valid_score': 0.8418544260890589, 'train_loss': 0.08016642610127724, 'valid_loss': 0.07164809737792786, 'duration': 578.784786, 'lr': 0.00028071281016403625}\nepoch 2\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8037564002885837, 'valid_score': 0.8713453936691895, 'train_loss': 0.08079370206372652, 'valid_loss': 0.07496459415730308, 'duration': 585.470988, 'lr': 0.0005212340121216189}\nepoch 3\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8214261310940373, 'valid_score': 0.8729646759710843, 'train_loss': 0.07874054459416624, 'valid_loss': 0.06729021120159065, 'duration': 590.757463, 'lr': 0.0007614235032516977}\nepoch 4\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8226242061975689, 'valid_score': 0.8948503343788581, 'train_loss': 0.07774241340425443, 'valid_loss': 0.06472354305360246, 'duration': 585.140571, 'lr': 0.0009367167193456758}\nepoch 5\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8470639649744557, 'valid_score': 0.8916386835150915, 'train_loss': 0.07461578585322512, 'valid_loss': 0.06983104202019817, 'duration': 590.309407, 'lr': 0.000999998790010987}\nepoch 6\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8504571331704445, 'valid_score': 0.8800448928558782, 'train_loss': 0.07359173024157228, 'valid_loss': 0.07041596415288308, 'duration': 586.578736, 'lr': 0.0009872180553045955}\nepoch 7\ntrain...\nvalid...\n\n\n```",
    "964230": "```\nException                                 Traceback (most recent call last)\n\nipython-input-1-1dd80295c7c1 in module()\n     56 \n     57 xmp.spawn(_mp_fn, args=(config,), nprocs=8,\n---&gt; 58           start_method='fork')\n\n2 frames\n\n/usr/local/lib/python3.6/dist-packages/torch/multiprocessing/spawn.py in join(self, timeout)\n    106                 raise Exception(\n    107                     \"process %d terminated with signal %s\" %\n--&gt; 108                     (error_index, name)\n    109                 )\n    110             else:\n\nException: process 7 terminated with signal SIGKILL\n```",
    "964233": "...then crash in epoch 7",
    "964250": "trying now:\n\n```\nloaders_colab_384_b3 = {\n    \"train_batch_size\": 24,\n    \"valid_batch_size\": 24,\n    \"train_num_workers\": 0,\n    \"valid_num_workers\": 0\n    }\n\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.5995671565614407, 'valid_score': 0.8349309187834001, 'train_loss': 0.178136943988479, 'valid_loss': 0.07360311624320115, 'duration': 783.181116, 'lr': 0.00010446322538335191}\n```\n\nhowever it's interesting that it crashed in valid part, in all cases I could use larger valid batch size than train batch size, and code stays the same, the new part is only larger image size",
    "964411": "OK, tested 20 epochs on Colab and 2 epochs on Kaggle, looks like it does work with 24/24\n\n```\nLoaded pretrained weights for efficientnet-b3\nrunning system on TPU\ncreated directory: ./session_08_09_2004/\ntrain_csv (25903, 13) valid_csv (6798, 13)\nepoch 0\ntrain...\nvalid...\nepoch done\n{'train_score': 0.6345018158930966, 'valid_score': 0.8512216785867384, 'train_loss': 0.15346827832700075, 'valid_loss': 0.07005745299379615, 'duration': 553.14053, 'lr': 0.0008052050401102377}\nsaving model\nepoch 1\ntrain...\nvalid...\nepoch done\n{'train_score': 0.8085982473311003, 'valid_score': 0.8889353661115494, 'train_loss': 0.07864298790214679, 'valid_loss': 0.0652915458293522, 'duration': 416.570716, 'lr': 7.307237815654868e-08}\nsaving model\n```\n\nso I am able to calculate 20 epochs within 3 hours of Kaggle TPU time",
    "964439": "I don't know much about Kaggle kernels/TPU, or how much memory you all get.  But yes, with validation you are not calculating gradients so. you can use ALOT more in your batch size typically.\n\nI have 4 GPU's and I typically run the following for Image Size of 384\n\nB1: train batch size: 25\n      validation batch size:  512 \nB4 train batch size: 14\n      validation batch size: 256.\n\nYou need to make SURE you are running your model in inference mode when doing validation!!!!  For example, in PyTorch, we put the model into inference mode by doing `model.eval()`.  What this does is it stops things like `BatchNorm` and `DropOut`.  Additionally, we also wrap our validation routine in a `with torch.no_grad()` , which stops the calculation of gradients.  If you are not doing this, you will need to use the SAME batch size for train and validation, which seems to be your situation.  Validation should be ALOT quicker than train, and you should be able to use at least a batch size of 10-20x train.  Just make sure you are setting things proper for your framework.",
    "964487": "What framework are you using? Why is `num_workers` set to 0? That's **never** the optimal value.",
    "964489": "You can't just change model complexity and image size and leave everything else the same and expect a better result.  If you change to a more complex model, with larger images, your batch size will change, and therefore your learning rate should change at a minimum.",
    "964628": "brianfeeny \nDo you know what is the dependence between batch size and model size? Is it possible to quickly estimate what is the \"good\" batch size for particular model and img size?",
    "964630": "As a matter of fact I also have quite good results on th 256 images. This is good size for the experimentation.",
    "964631": "It  seems there is some memory leak when using pytorch?",
    "964695": "janidziak just try stuff.  I have seen people do formulas but I always just set my batch to the max I can and then do some experimentation.  Using a cyclic scheduler or a \"learning rate finder\" is a good way to find something to start with.",
    "964751": "The main bottleneck for Pytorch Kaggle Kernel TPU is insufficient CPU cores and RAM when the data loader is being processed. A minimum requirement of at least 16 cores and high RAM in order to maximize the potential of TPU in Pytorch.\nI would advise training a model with image size 256x256 if you intend to train a model in Kaggle TPU kernel",
    "964837": "brianfeeny \njest I think it is implemented correctly:\n\n```\ntrain_result = train_step(config_step)\n(...)\nwith torch.no_grad():\n     valid_result = train_step(config_step)\n\n```\n\n```\n    if (train_mode):\n        model.train() \n    else:\n        model.eval()         \n\n```\n\nso I wonder how can you use 256 valid when my code fails for 32",
    "964841": "brianfeeny \ndo you use TPU? I found that 0 is optimal for TPU",
    "964849": "Hard to say why, but obviously something is wrong. I assume you have two dataloaders and are trying to set a different batch_size in each?",
    "964854": "jacekpoplawski I don't use TPU, I use a GPU in my lab.  num_workers has (or should have) nothing to do with TPU vs GPU, it has to do with your CPU, as num_workers are parallel dataloading processes which load data into CPU memory.",
    "964869": "As I said I always set valid batches higher, sometimes 2x higher and they were working. Only problem appears on image size 384 on Kaggle TPU and Colab TPU.",
    "965106": "jacekpoplawski  Jacek, do you use mixed precision in Pytorch? I  do not use TPU, so can say for Colab GPU only.\nWith mixed precision I managed to run 384x384 B6 model with batch size 24. Without mixed precision, max batch size I was able to train the same model with was 12.",
    "965122": "I don't have problem on GPU, no I still was not able to learn how to use mixed precision with pretrained models.",
    "965141": "Well, mixed precision could help you with TPU issue also I guess.\n\nBelow is a simplified code from Abhishek's repo there I added mixed precision support for GPU.\n\n```\nclass Engine:\n    @staticmethod\n    def train(\n        data_loader,\n        model,\n        optimizer,\n        device,\n        scheduler=None,\n        accumulation_steps=1,\n        fp16=True, # fp16 = True - if we want to apply mixed precision\n    ):\n\n        losses = AverageMeter()\n        final_predictions = []\n        model.train()\n        if accumulation_steps &gt; 1:\n            optimizer.zero_grad()\n\n        if fp16:\n          scaler = torch.cuda.amp.GradScaler()\n\n        for b_idx, data in enumerate(data_loader):\n            for key, value in data.items():\n                data[key] = value.to(device)\n            if accumulation_steps == 1 and b_idx == 0:\n                optimizer.zero_grad()\n            if fp16:    \n                with torch.cuda.amp.autocast():    \n                    predictions, loss = model(**data)\n            else:\n                predictions, loss = model(**data)      \n            predictions = predictions.detach().cpu()\n            final_predictions.append(predictions) \n\n            with torch.set_grad_enabled(True):\n                if fp16:\n                    scaler.scale(loss).backward()                   \n                else:\n                    loss.backward()\n                if (b_idx + 1) % accumulation_steps == 0:\n                    if fp16:\n                        scaler.step(optimizer)\n                        scaler.update()\n                    else:     \n                        optimizer.step()\n                    if scheduler is not None:\n                         scheduler.step()\n                    if b_idx &gt; 0:\n                        optimizer.zero_grad()\n\n            losses.update(loss.item(), data_loader.batch_size)\n\n        return final_predictions, losses.avg\n```",
    "965144": "can you explain what happens with weights during mixed precision? can you just use efficientnet weights and everything works?",
    "965171": "I ran 2 tests, with and without mixed precision. \n1. CV\\LB with mixed precision ON are almost identical to those with mixed precision OFF.\n2. I ran in on P100 that doesn't support enhanced mixed precision matrices multiplication like more modern GPUs do. But P100 does support memory optimization. So you'll reduce your memory 2x.\n3. speaking of weights, mixed precision optimization uses loss scaling - to preserve very small gradient values, it scales loss value (and therefore gradient values) up and later scales everything back down.\n\n&gt;Maintain a master copy of weights in FP32\nFor each iteration:\nMake an FP16 copy of the weights\nForward propagation (FP16 weights and activations)\nMultiply the resulting loss with the scaling factor S\nBackward propagation (FP16 weights, activations, and their gradients)\nMultiply the weight gradient with 1/S\nComplete the weight update (including gradient clipping, etc.)",
    "965198": "dunklerwald \nah so weights are in original 32-bit\n\nwhat about BatchNormalization issue? do you need to do anything with efficientnet?",
    "965220": "jacekpoplawski maybe I would need to do something but in fact I do nothing:-) and all layers are unfreezed.",
    "965237": "I read that you need to replace BN with different implementation, but you said you did nothing and it works, right? What are your results? I assume CV/LB is same, so what about time of epoch (normal vs mixed)?",
    "965273": "Right, I did nothing with BN. CV/LB are the same. Time-wise, i didn't observe any significant difference sadly. But must admit, I didn't track timing accurately. As I said earlier, most likely, that's because I was running it on Google Colab P100 which is not able to take advantage of optimized mixed precision matrices multiplications. If you run mixed precision  on Kaggle TPU v3, you should expect substantial time savings.",
    "968611": "sig220932 thanks for sharing that. i tired reduce lr and it helped for large models. thanks again."
  },
  "source": "meta"
}