{
  "id": 161713,
  "title": "Need more runtime per run for TPUs!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/161713",
  "author_name": "",
  "post_date": "2020-06-25T21:51:48.345520700Z",
  "votes": 5,
  "comment_count": 15,
  "views": 0,
  "content": "<p>From what I see with my experiments and best scoring public kernels, I think there is a need for more execution time per runtime for TPUs. Theoretically, the bigger the EfficientNet model, the higher the image size as input should be (even if it's not a requirement).</p>\n\n<p>This <a href=\"https://www.kaggle.com/ragnar123/efficientnet-x-384\">kernel</a> whose best version scores 0.936 public can barely perform a 5-fold CV with only a B3 model, and when I'm training a B5 model with 20 epochs and image size 384x384 (instead of 456x456 suggested) from this kernel, my model is clearly not trained enough.</p>\n\n<p>This <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">kernel</a> shows that training 12 epochs with image size 224x244 without CV takes half of the 3-hour runtime for one execution.</p>\n\n<p>One of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime and is again clearly not trained enough.</p>\n\n<p>I could give dozens of examples like the previous one to illustrate that we clearly need more than 3-hour runtime per execution to get models trained properly. To cite <a href=\"http://karpathy.github.io/2019/04/25/recipe/\">Andrej Karpathy's recipe for training neural networks</a>:</p>\n\n<p>&gt; leave it training. I’ve often seen people tempted to stop the model training when the validation loss seems to be leveling off. In my experience networks keep training for unintuitively long time. One time I accidentally left a model training during the winter break and when I got back in January it was SOTA (“state of the art”).</p>\n\n<p>This is especially true for EfficientNets which are clearly SOTA models until now (unless I missed something in the last weeks).</p>\n\n<p>I would like to know if other competitors also feel limited by this 3-hour per run limit for TPUs (maybe you can give some examples of experiments that would have required more training time), and if Kaggle plans to increase the time limit for TPUs to 9 hours as for CPUs and GPUs. I do think it's necessary to have properly trained models and avoid wasting time to have models training as close as possible to the 3-hour limit.</p>",
  "messages": [
    {
      "id": "902038",
      "postDate": "06/25/2020 21:51:48",
      "content": "<p>From what I see with my experiments and best scoring public kernels, I think there is a need for more execution time per runtime for TPUs. Theoretically, the bigger the EfficientNet model, the higher the image size as input should be (even if it's not a requirement).</p>\n\n<p>This <a href=\"https://www.kaggle.com/ragnar123/efficientnet-x-384\">kernel</a> whose best version scores 0.936 public can barely perform a 5-fold CV with only a B3 model, and when I'm training a B5 model with 20 epochs and image size 384x384 (instead of 456x456 suggested) from this kernel, my model is clearly not trained enough.</p>\n\n<p>This <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">kernel</a> shows that training 12 epochs with image size 224x244 without CV takes half of the 3-hour runtime for one execution.</p>\n\n<p>One of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime and is again clearly not trained enough.</p>\n\n<p>I could give dozens of examples like the previous one to illustrate that we clearly need more than 3-hour runtime per execution to get models trained properly. To cite <a href=\"http://karpathy.github.io/2019/04/25/recipe/\">Andrej Karpathy's recipe for training neural networks</a>:</p>\n\n<p>&gt; leave it training. I’ve often seen people tempted to stop the model training when the validation loss seems to be leveling off. In my experience networks keep training for unintuitively long time. One time I accidentally left a model training during the winter break and when I got back in January it was SOTA (“state of the art”).</p>\n\n<p>This is especially true for EfficientNets which are clearly SOTA models until now (unless I missed something in the last weeks).</p>\n\n<p>I would like to know if other competitors also feel limited by this 3-hour per run limit for TPUs (maybe you can give some examples of experiments that would have required more training time), and if Kaggle plans to increase the time limit for TPUs to 9 hours as for CPUs and GPUs. I do think it's necessary to have properly trained models and avoid wasting time to have models training as close as possible to the 3-hour limit.</p>",
      "rawMarkdown": "From what I see with my experiments and best scoring public kernels, I think there is a need for more execution time per runtime for TPUs. Theoretically, the bigger the EfficientNet model, the higher the image size as input should be (even if it's not a requirement).\n\nThis [kernel](https://www.kaggle.com/ragnar123/efficientnet-x-384) whose best version scores 0.936 public can barely perform a 5-fold CV with only a B3 model, and when I'm training a B5 model with 20 epochs and image size 384x384 (instead of 456x456 suggested) from this kernel, my model is clearly not trained enough.\n\nThis [kernel](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once) shows that training 12 epochs with image size 224x244 without CV takes half of the 3-hour runtime for one execution.\n\nOne of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime and is again clearly not trained enough.\n\nI could give dozens of examples like the previous one to illustrate that we clearly need more than 3-hour runtime per execution to get models trained properly. To cite [Andrej Karpathy's recipe for training neural networks](http://karpathy.github.io/2019/04/25/recipe/):\n\n&gt; leave it training. I’ve often seen people tempted to stop the model training when the validation loss seems to be leveling off. In my experience networks keep training for unintuitively long time. One time I accidentally left a model training during the winter break and when I got back in January it was SOTA (“state of the art”).\n\nThis is especially true for EfficientNets which are clearly SOTA models until now (unless I missed something in the last weeks).\n\nI would like to know if other competitors also feel limited by this 3-hour per run limit for TPUs (maybe you can give some examples of experiments that would have required more training time), and if Kaggle plans to increase the time limit for TPUs to 9 hours as for CPUs and GPUs. I do think it's necessary to have properly trained models and avoid wasting time to have models training as close as possible to the 3-hour limit.",
      "votes": null
    },
    {
      "id": "902187",
      "postDate": "06/26/2020 01:15:55",
      "content": "<p>You might consider moving to Colab. The TPU is not as powerful as the one offered by Kaggle but it is decent enough. And you can train your model for 12 hours if you are using the free version and for 24 hours if you pay $10/month for Colab Pro. Also, you can run multiple models simultaneously (I ran up to 5 TPU models, but maybe you can run even more -- I have never checked the limit).</p>\n\n<p>UPDATE: But  don't get me wrong -- I am with you. It would be great if we were able to run TPU's on Kaggle for longer that 3 hours. We would be able to process larger batch sizes, train on larger images (e.g. 1024x1024) and use the most advanced model (e.g. EfficientNet B7). </p>",
      "rawMarkdown": "You might consider moving to Colab. The TPU is not as powerful as the one offered by Kaggle but it is decent enough. And you can train your model for 12 hours if you are using the free version and for 24 hours if you pay $10/month for Colab Pro. Also, you can run multiple models simultaneously (I ran up to 5 TPU models, but maybe you can run even more -- I have never checked the limit).\n\nUPDATE: But  don't get me wrong -- I am with you. It would be great if we were able to run TPU's on Kaggle for longer that 3 hours. We would be able to process larger batch sizes, train on larger images (e.g. 1024x1024) and use the most advanced model (e.g. EfficientNet B7).",
      "votes": null
    },
    {
      "id": "902382",
      "postDate": "06/26/2020 05:40:39",
      "content": "<p>Hi <a href=\"/mika30\">@mika30</a>, <strong>actually it is possible to run 5 fold EfficientNetB7 with image size 512x512 using Kaggle Notebooks with TPU!</strong>\nAs you mentioned:\n&gt; One of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime</p>\n\n<p>Therefore you could just simply run different folds in different session (confirmed by my experiments: 2 folds per session) and write submission for each fold. However, by doing so, you should not just blend linearly your submissions (until you have different initial weights on i.e. last dense layer). One of the options is to use rank blend method published by <a href=\"/ragnar123\">@ragnar123</a> <a href=\"https://www.kaggle.com/ragnar123/rank-then-blend\">here</a></p>\n\n<p>Moreover, running  single fold EfficientNetB7 with image size 1024x1024 took Kaggle Notebooks with TPU about 2,5 hour (with 10 epochs). This 5-CV 'monster' will 'eat' about 12,5 hours of your TPU quota, but worth to try :)</p>",
      "rawMarkdown": "Hi @mika30, **actually it is possible to run 5 fold EfficientNetB7 with image size 512x512 using Kaggle Notebooks with TPU!**\nAs you mentioned:\n&gt; One of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime\n\nTherefore you could just simply run different folds in different session (confirmed by my experiments: 2 folds per session) and write submission for each fold. However, by doing so, you should not just blend linearly your submissions (until you have different initial weights on i.e. last dense layer). One of the options is to use rank blend method published by @ragnar123 [here](https://www.kaggle.com/ragnar123/rank-then-blend)\n\nMoreover, running  single fold EfficientNetB7 with image size 1024x1024 took Kaggle Notebooks with TPU about 2,5 hour (with 10 epochs). This 5-CV 'monster' will 'eat' about 12,5 hours of your TPU quota, but worth to try :)",
      "votes": null
    },
    {
      "id": "902408",
      "postDate": "06/26/2020 06:02:55",
      "content": "<p>You can overcome TPU/GPU time limits by Save and Load/Restore <a href=\"https://www.tensorflow.org/guide/checkpoint\"><strong>Checkpoints</strong></a>\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455\">ModelCheckpoint</a> will work with a TPU if the model_dir is local and the checkpoint format is .h5\nSee example <a href=\"https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline/notebook\">Kernel</a></p>",
      "rawMarkdown": "You can overcome TPU/GPU time limits by Save and Load/Restore [**Checkpoints**](https://www.tensorflow.org/guide/checkpoint)\n[ModelCheckpoint](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455) will work with a TPU if the model_dir is local and the checkpoint format is .h5\nSee example [Kernel](https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline/notebook)",
      "votes": null
    },
    {
      "id": "903499",
      "postDate": "06/26/2020 22:04:37",
      "content": "<p>Good to know you find Colab is decent enough, I had read that there was a noticeable difference and I wasn't sure it was worth to try Colab... If considering the 12-hour limit, we'd need Colab TPUs to be no more than 4 times less \"powerful\" than Kaggle TPUs. I'll probably at least try it in the coming weeks!</p>\n\n<p>No worries about your update, I got it, thank you for your advice! 👍 </p>",
      "rawMarkdown": "Good to know you find Colab is decent enough, I had read that there was a noticeable difference and I wasn't sure it was worth to try Colab... If considering the 12-hour limit, we'd need Colab TPUs to be no more than 4 times less \"powerful\" than Kaggle TPUs. I'll probably at least try it in the coming weeks!\n\nNo worries about your update, I got it, thank you for your advice! 👍",
      "votes": null
    },
    {
      "id": "903504",
      "postDate": "06/26/2020 22:15:42",
      "content": "<p>Thank you for your advice, indeed we can do that but I'm not 100% sure we can also restore data generators to the same state. Results obtained with multi-core TPUs are not reproducible, and I think at least breaking the data generator flow can hurt the training.</p>\n\n<p>Maybe it could also improve with some luck, but for now, when I tried to save and load as you described, I've almost never reached a better score afterward. However, I'd need further investigation to be sure I did everything correctly.</p>",
      "rawMarkdown": "Thank you for your advice, indeed we can do that but I'm not 100% sure we can also restore data generators to the same state. Results obtained with multi-core TPUs are not reproducible, and I think at least breaking the data generator flow can hurt the training.\n\nMaybe it could also improve with some luck, but for now, when I tried to save and load as you described, I've almost never reached a better score afterward. However, I'd need further investigation to be sure I did everything correctly.",
      "votes": null
    },
    {
      "id": "903509",
      "postDate": "06/26/2020 22:28:30",
      "content": "<p>Hi <a href=\"/wrrosa\">@wrrosa</a>, thanks for your advice, this is actually what I began to do right after launching this discussion, in order to use my quota before today's reset! I also spotted Ragnar's kernel, thanks for mentioning it.</p>\n\n<p>Interesting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? I'm currently running one fold on 384 x 384 and 768 x 768 to see the results.</p>\n\n<p>Also, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs... Otherwise, I think more runtime is needed per single run... 😉 </p>\n\n<p>If we had at least 6 hours per single run, maybe it would be worth to use 1 week of TPU quota by running 5 folds x 6 hours of an EfficientNet B7 😄 </p>",
      "rawMarkdown": "Hi @wrrosa, thanks for your advice, this is actually what I began to do right after launching this discussion, in order to use my quota before today's reset! I also spotted Ragnar's kernel, thanks for mentioning it.\n\nInteresting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? I'm currently running one fold on 384 x 384 and 768 x 768 to see the results.\n\nAlso, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs... Otherwise, I think more runtime is needed per single run... 😉 \n\nIf we had at least 6 hours per single run, maybe it would be worth to use 1 week of TPU quota by running 5 folds x 6 hours of an EfficientNet B7 😄",
      "votes": null
    },
    {
      "id": "903803",
      "postDate": "06/27/2020 05:46:43",
      "content": "<p><a href=\"/mika30\">@mika30</a> In USA, you have $10/month cheap <a href=\"https://colab.research.google.com/signup\">Colab PRO</a> option, where you can run your training session continuously for longer periods (24-hr sessions).</p>",
      "rawMarkdown": "mika30 In USA, you have $10/month cheap [Colab PRO](https://colab.research.google.com/signup) option, where you can run your training session continuously for longer periods (24-hr sessions).",
      "votes": null
    },
    {
      "id": "904710",
      "postDate": "06/27/2020 20:35:26",
      "content": "<p>&gt; Interesting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? </p>\n\n<p>One folded 1024x1024 was only time-testing run, so it scores under 0.9. And 600x600 I haven't tried yet - but worth to try. In fact before this, one should public a dataset containing resized tfrecords, without this gcs with tpu won't recognize images. In this competition, it is amazing for me, that blending nets of different input image sizes are giving better score than i.e 512x512. It proves, that smaller nns with smaller input size images 'can see' some different things and most valuable for our predictions. Nice stuff!</p>\n\n<p>&gt; Also, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs… Otherwise, I think more runtime is needed per single run… 😉 </p>\n\n<p>Sure it's not enough. However, it's seems that <code>model.save_weights()</code> and <code>model.load_weights()</code> are working well outside <code>strategy.scope()</code> - it could be nice workaround to run tpu monster enetb7 1024 on kaggle notebooks. Anyway, relocate to Colab is coming :)</p>",
      "rawMarkdown": "&gt; Interesting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? \n\nOne folded 1024x1024 was only time-testing run, so it scores under 0.9. And 600x600 I haven't tried yet - but worth to try. In fact before this, one should public a dataset containing resized tfrecords, without this gcs with tpu won't recognize images. In this competition, it is amazing for me, that blending nets of different input image sizes are giving better score than i.e 512x512. It proves, that smaller nns with smaller input size images 'can see' some different things and most valuable for our predictions. Nice stuff!\n\n&gt; Also, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs… Otherwise, I think more runtime is needed per single run… 😉 \n\nSure it's not enough. However, it's seems that `model.save_weights()` and `model.load_weights() ` are working well outside `strategy.scope()` - it could be nice workaround to run tpu monster enetb7 1024 on kaggle notebooks. Anyway, relocate to Colab is coming :)",
      "votes": null
    },
    {
      "id": "905087",
      "postDate": "06/28/2020 08:33:20",
      "content": "<p>yes but as you wrote</p>\n\n<p>\"For now, Colab Pro is only available in the US.\" </p>\n\n<p>very disappointing</p>",
      "rawMarkdown": "yes but as you wrote\n\n\"For now, Colab Pro is only available in the US.\" \n\nvery disappointing",
      "votes": null
    },
    {
      "id": "905187",
      "postDate": "06/28/2020 10:47:58",
      "content": "<p>It's available in Europe</p>\n\n<p>I use it for a few months</p>",
      "rawMarkdown": "It's available in Europe\n\nI use it for a few months",
      "votes": null
    },
    {
      "id": "905204",
      "postDate": "06/28/2020 10:59:15",
      "content": "<p><a href=\"/serigne\">@serigne</a> do you mean this statement from the website is incorrect?</p>",
      "rawMarkdown": "serigne do you mean this statement from the website is incorrect?",
      "votes": null
    },
    {
      "id": "905217",
      "postDate": "06/28/2020 11:12:38",
      "content": "<p>I think so.</p>\n\n<p>I use Colab Pro from France since the beginning.</p>",
      "rawMarkdown": "I think so.\n\nI use Colab Pro from France since the beginning.",
      "votes": null
    },
    {
      "id": "905285",
      "postDate": "06/28/2020 12:09:16",
      "content": "<p>Please see the FAQ at the bottom of the Colab Pro page:\n<code>\nWhere is Colab Pro available?\nFor now, Colab Pro is only available in the US.\n</code>\nRegistration page also specifically asks for your US state and pincode.\nYour account may be banned if they catch you using it outside US.</p>",
      "rawMarkdown": "Please see the FAQ at the bottom of the Colab Pro page:\n```\nWhere is Colab Pro available?\nFor now, Colab Pro is only available in the US.\n```\nRegistration page also specifically asks for your US state and pincode.\nYour account may be banned if they catch you using it outside US.",
      "votes": null
    },
    {
      "id": "905299",
      "postDate": "06/28/2020 12:23:29",
      "content": "<p>My account will never be banned. </p>\n\n<p>I am a premium Google Drive User for years, got unlimited colab GPU use even without colab PRO and Google  knows my exact address in France which I used during my registration. </p>\n\n<p>I use colab Pro just for unlimited TPU.</p>",
      "rawMarkdown": "My account will never be banned. \n\nI am a premium Google Drive User for years, got unlimited colab GPU use even without colab PRO and Google  knows my exact address in France which I used during my registration. \n\nI use colab Pro just for unlimited TPU.",
      "votes": null
    },
    {
      "id": "909775",
      "postDate": "06/30/2020 20:17:31",
      "content": "<p>tf.data.Dataset is essentially stateless between epochs. Things should be fine if you stop on an epoch boundary. The one thing to take care of is your learning rate schedule though. When you restart training, you have to set the <code>initial_epoch=</code> parameter in <code>model.fit()</code>.</p>",
      "rawMarkdown": "tf.data.Dataset is essentially stateless between epochs. Things should be fine if you stop on an epoch boundary. The one thing to take care of is your learning rate schedule though. When you restart training, you have to set the `initial_epoch=` parameter in `model.fit()`.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 902187,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "06/26/2020 01:15:55",
      "content": "<p>You might consider moving to Colab. The TPU is not as powerful as the one offered by Kaggle but it is decent enough. And you can train your model for 12 hours if you are using the free version and for 24 hours if you pay $10/month for Colab Pro. Also, you can run multiple models simultaneously (I ran up to 5 TPU models, but maybe you can run even more -- I have never checked the limit).</p>\n\n<p>UPDATE: But  don't get me wrong -- I am with you. It would be great if we were able to run TPU's on Kaggle for longer that 3 hours. We would be able to process larger batch sizes, train on larger images (e.g. 1024x1024) and use the most advanced model (e.g. EfficientNet B7). </p>",
      "votes": null,
      "replies": [
        {
          "id": 903499,
          "author_name": "mika30",
          "author_url": "",
          "post_date": "06/26/2020 22:04:37",
          "content": "<p>Good to know you find Colab is decent enough, I had read that there was a noticeable difference and I wasn't sure it was worth to try Colab... If considering the 12-hour limit, we'd need Colab TPUs to be no more than 4 times less \"powerful\" than Kaggle TPUs. I'll probably at least try it in the coming weeks!</p>\n\n<p>No worries about your update, I got it, thank you for your advice! 👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 902382,
      "author_name": "wrrosa",
      "author_url": "",
      "post_date": "06/26/2020 05:40:39",
      "content": "<p>Hi <a href=\"/mika30\">@mika30</a>, <strong>actually it is possible to run 5 fold EfficientNetB7 with image size 512x512 using Kaggle Notebooks with TPU!</strong>\nAs you mentioned:\n&gt; One of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime</p>\n\n<p>Therefore you could just simply run different folds in different session (confirmed by my experiments: 2 folds per session) and write submission for each fold. However, by doing so, you should not just blend linearly your submissions (until you have different initial weights on i.e. last dense layer). One of the options is to use rank blend method published by <a href=\"/ragnar123\">@ragnar123</a> <a href=\"https://www.kaggle.com/ragnar123/rank-then-blend\">here</a></p>\n\n<p>Moreover, running  single fold EfficientNetB7 with image size 1024x1024 took Kaggle Notebooks with TPU about 2,5 hour (with 10 epochs). This 5-CV 'monster' will 'eat' about 12,5 hours of your TPU quota, but worth to try :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 903509,
          "author_name": "mika30",
          "author_url": "",
          "post_date": "06/26/2020 22:28:30",
          "content": "<p>Hi <a href=\"/wrrosa\">@wrrosa</a>, thanks for your advice, this is actually what I began to do right after launching this discussion, in order to use my quota before today's reset! I also spotted Ragnar's kernel, thanks for mentioning it.</p>\n\n<p>Interesting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? I'm currently running one fold on 384 x 384 and 768 x 768 to see the results.</p>\n\n<p>Also, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs... Otherwise, I think more runtime is needed per single run... 😉 </p>\n\n<p>If we had at least 6 hours per single run, maybe it would be worth to use 1 week of TPU quota by running 5 folds x 6 hours of an EfficientNet B7 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904710,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "06/27/2020 20:35:26",
          "content": "<p>&gt; Interesting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? </p>\n\n<p>One folded 1024x1024 was only time-testing run, so it scores under 0.9. And 600x600 I haven't tried yet - but worth to try. In fact before this, one should public a dataset containing resized tfrecords, without this gcs with tpu won't recognize images. In this competition, it is amazing for me, that blending nets of different input image sizes are giving better score than i.e 512x512. It proves, that smaller nns with smaller input size images 'can see' some different things and most valuable for our predictions. Nice stuff!</p>\n\n<p>&gt; Also, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs… Otherwise, I think more runtime is needed per single run… 😉 </p>\n\n<p>Sure it's not enough. However, it's seems that <code>model.save_weights()</code> and <code>model.load_weights()</code> are working well outside <code>strategy.scope()</code> - it could be nice workaround to run tpu monster enetb7 1024 on kaggle notebooks. Anyway, relocate to Colab is coming :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 902408,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "06/26/2020 06:02:55",
      "content": "<p>You can overcome TPU/GPU time limits by Save and Load/Restore <a href=\"https://www.tensorflow.org/guide/checkpoint\"><strong>Checkpoints</strong></a>\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455\">ModelCheckpoint</a> will work with a TPU if the model_dir is local and the checkpoint format is .h5\nSee example <a href=\"https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline/notebook\">Kernel</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 903504,
          "author_name": "mika30",
          "author_url": "",
          "post_date": "06/26/2020 22:15:42",
          "content": "<p>Thank you for your advice, indeed we can do that but I'm not 100% sure we can also restore data generators to the same state. Results obtained with multi-core TPUs are not reproducible, and I think at least breaking the data generator flow can hurt the training.</p>\n\n<p>Maybe it could also improve with some luck, but for now, when I tried to save and load as you described, I've almost never reached a better score afterward. However, I'd need further investigation to be sure I did everything correctly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 909775,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "06/30/2020 20:17:31",
          "content": "<p>tf.data.Dataset is essentially stateless between epochs. Things should be fine if you stop on an epoch boundary. The one thing to take care of is your learning rate schedule though. When you restart training, you have to set the <code>initial_epoch=</code> parameter in <code>model.fit()</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903803,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "06/27/2020 05:46:43",
      "content": "<p><a href=\"/mika30\">@mika30</a> In USA, you have $10/month cheap <a href=\"https://colab.research.google.com/signup\">Colab PRO</a> option, where you can run your training session continuously for longer periods (24-hr sessions).</p>",
      "votes": null,
      "replies": [
        {
          "id": 905087,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "06/28/2020 08:33:20",
          "content": "<p>yes but as you wrote</p>\n\n<p>\"For now, Colab Pro is only available in the US.\" </p>\n\n<p>very disappointing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 905187,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "06/28/2020 10:47:58",
          "content": "<p>It's available in Europe</p>\n\n<p>I use it for a few months</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 905204,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "06/28/2020 10:59:15",
          "content": "<p><a href=\"/serigne\">@serigne</a> do you mean this statement from the website is incorrect?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 905217,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "06/28/2020 11:12:38",
          "content": "<p>I think so.</p>\n\n<p>I use Colab Pro from France since the beginning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 905285,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "06/28/2020 12:09:16",
          "content": "<p>Please see the FAQ at the bottom of the Colab Pro page:\n<code>\nWhere is Colab Pro available?\nFor now, Colab Pro is only available in the US.\n</code>\nRegistration page also specifically asks for your US state and pincode.\nYour account may be banned if they catch you using it outside US.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 905299,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "06/28/2020 12:23:29",
          "content": "<p>My account will never be banned. </p>\n\n<p>I am a premium Google Drive User for years, got unlimited colab GPU use even without colab PRO and Google  knows my exact address in France which I used during my registration. </p>\n\n<p>I use colab Pro just for unlimited TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "902038": "From what I see with my experiments and best scoring public kernels, I think there is a need for more execution time per runtime for TPUs. Theoretically, the bigger the EfficientNet model, the higher the image size as input should be (even if it's not a requirement).\n\nThis [kernel](https://www.kaggle.com/ragnar123/efficientnet-x-384) whose best version scores 0.936 public can barely perform a 5-fold CV with only a B3 model, and when I'm training a B5 model with 20 epochs and image size 384x384 (instead of 456x456 suggested) from this kernel, my model is clearly not trained enough.\n\nThis [kernel](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once) shows that training 12 epochs with image size 224x244 without CV takes half of the 3-hour runtime for one execution.\n\nOne of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime and is again clearly not trained enough.\n\nI could give dozens of examples like the previous one to illustrate that we clearly need more than 3-hour runtime per execution to get models trained properly. To cite [Andrej Karpathy's recipe for training neural networks](http://karpathy.github.io/2019/04/25/recipe/):\n\n&gt; leave it training. I’ve often seen people tempted to stop the model training when the validation loss seems to be leveling off. In my experience networks keep training for unintuitively long time. One time I accidentally left a model training during the winter break and when I got back in January it was SOTA (“state of the art”).\n\nThis is especially true for EfficientNets which are clearly SOTA models until now (unless I missed something in the last weeks).\n\nI would like to know if other competitors also feel limited by this 3-hour per run limit for TPUs (maybe you can give some examples of experiments that would have required more training time), and if Kaggle plans to increase the time limit for TPUs to 9 hours as for CPUs and GPUs. I do think it's necessary to have properly trained models and avoid wasting time to have models training as close as possible to the 3-hour limit.",
    "902187": "You might consider moving to Colab. The TPU is not as powerful as the one offered by Kaggle but it is decent enough. And you can train your model for 12 hours if you are using the free version and for 24 hours if you pay $10/month for Colab Pro. Also, you can run multiple models simultaneously (I ran up to 5 TPU models, but maybe you can run even more -- I have never checked the limit).\n\nUPDATE: But  don't get me wrong -- I am with you. It would be great if we were able to run TPU's on Kaggle for longer that 3 hours. We would be able to process larger batch sizes, train on larger images (e.g. 1024x1024) and use the most advanced model (e.g. EfficientNet B7).",
    "902382": "Hi @mika30, **actually it is possible to run 5 fold EfficientNetB7 with image size 512x512 using Kaggle Notebooks with TPU!**\nAs you mentioned:\n&gt; One of my experiments with a B7 model, 20 epochs, image size 600x600 and without CV takes 2-hour runtime\n\nTherefore you could just simply run different folds in different session (confirmed by my experiments: 2 folds per session) and write submission for each fold. However, by doing so, you should not just blend linearly your submissions (until you have different initial weights on i.e. last dense layer). One of the options is to use rank blend method published by @ragnar123 [here](https://www.kaggle.com/ragnar123/rank-then-blend)\n\nMoreover, running  single fold EfficientNetB7 with image size 1024x1024 took Kaggle Notebooks with TPU about 2,5 hour (with 10 epochs). This 5-CV 'monster' will 'eat' about 12,5 hours of your TPU quota, but worth to try :)",
    "902408": "You can overcome TPU/GPU time limits by Save and Load/Restore [**Checkpoints**](https://www.tensorflow.org/guide/checkpoint)\n[ModelCheckpoint](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132215#757455) will work with a TPU if the model_dir is local and the checkpoint format is .h5\nSee example [Kernel](https://www.kaggle.com/dimitreoliveira/flower-classification-with-tpus-eda-and-baseline/notebook)",
    "903499": "Good to know you find Colab is decent enough, I had read that there was a noticeable difference and I wasn't sure it was worth to try Colab... If considering the 12-hour limit, we'd need Colab TPUs to be no more than 4 times less \"powerful\" than Kaggle TPUs. I'll probably at least try it in the coming weeks!\n\nNo worries about your update, I got it, thank you for your advice! 👍",
    "903504": "Thank you for your advice, indeed we can do that but I'm not 100% sure we can also restore data generators to the same state. Results obtained with multi-core TPUs are not reproducible, and I think at least breaking the data generator flow can hurt the training.\n\nMaybe it could also improve with some luck, but for now, when I tried to save and load as you described, I've almost never reached a better score afterward. However, I'd need further investigation to be sure I did everything correctly.",
    "903509": "Hi @wrrosa, thanks for your advice, this is actually what I began to do right after launching this discussion, in order to use my quota before today's reset! I also spotted Ragnar's kernel, thanks for mentioning it.\n\nInteresting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? I'm currently running one fold on 384 x 384 and 768 x 768 to see the results.\n\nAlso, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs... Otherwise, I think more runtime is needed per single run... 😉 \n\nIf we had at least 6 hours per single run, maybe it would be worth to use 1 week of TPU quota by running 5 folds x 6 hours of an EfficientNet B7 😄",
    "903803": "mika30 In USA, you have $10/month cheap [Colab PRO](https://colab.research.google.com/signup) option, where you can run your training session continuously for longer periods (24-hr sessions).",
    "904710": "&gt; Interesting that you tried 1024 x 1024, did you also try recommended 600 x 600 for B7 and if so, did one of both experiments give significantly better results? \n\nOne folded 1024x1024 was only time-testing run, so it scores under 0.9. And 600x600 I haven't tried yet - but worth to try. In fact before this, one should public a dataset containing resized tfrecords, without this gcs with tpu won't recognize images. In this competition, it is amazing for me, that blending nets of different input image sizes are giving better score than i.e 512x512. It proves, that smaller nns with smaller input size images 'can see' some different things and most valuable for our predictions. Nice stuff!\n\n&gt; Also, I think it would be worth trying more than 10 epochs for your B7 1024 x 1024 1-fold experiment, I'd be surprised if you tell me you think it's been trained enough with only 10 epochs… Otherwise, I think more runtime is needed per single run… 😉 \n\nSure it's not enough. However, it's seems that `model.save_weights()` and `model.load_weights() ` are working well outside `strategy.scope()` - it could be nice workaround to run tpu monster enetb7 1024 on kaggle notebooks. Anyway, relocate to Colab is coming :)",
    "905087": "yes but as you wrote\n\n\"For now, Colab Pro is only available in the US.\" \n\nvery disappointing",
    "905187": "It's available in Europe\n\nI use it for a few months",
    "905204": "serigne do you mean this statement from the website is incorrect?",
    "905217": "I think so.\n\nI use Colab Pro from France since the beginning.",
    "905285": "Please see the FAQ at the bottom of the Colab Pro page:\n```\nWhere is Colab Pro available?\nFor now, Colab Pro is only available in the US.\n```\nRegistration page also specifically asks for your US state and pincode.\nYour account may be banned if they catch you using it outside US.",
    "905299": "My account will never be banned. \n\nI am a premium Google Drive User for years, got unlimited colab GPU use even without colab PRO and Google  knows my exact address in France which I used during my registration. \n\nI use colab Pro just for unlimited TPU.",
    "909775": "tf.data.Dataset is essentially stateless between epochs. Things should be fine if you stop on an epoch boundary. The one thing to take care of is your learning rate schedule though. When you restart training, you have to set the `initial_epoch=` parameter in `model.fit()`."
  },
  "source": "meta"
}