{
  "id": 156263,
  "title": "Tips to make the most out of a kaggle GPU/TPU session",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156263",
  "author_name": "",
  "post_date": "2020-06-05T07:45:12.567522300Z",
  "votes": 26,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I am starting this topic because this is my first kaggle competition and the challenge for me here is more how to train complex models with a limited computation power.</p>\n\n<p>I would like to share 2 interesting topics shared by <a href=\"/maxlenormand\">@maxlenormand</a> that helped me a lot to organize my workflow:\n<strong>1-</strong> <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/137024\">Debrief - Thoughts on computer vision with limited hardware </a>\n<strong>2-</strong> <a href=\"https://towardsdatascience.com/making-the-most-out-of-limited-hardware-for-computer-vision-kaggle-competitions-a8894511026e\">Making the most out of limited hardware for Computer Vision Kaggle competitions</a>.</p>\n\n<p>In order to increase the score I am trying more augmentations, deeper architectures or simply increase the image size. But any relatively small change leads to exceeding the allowed 9 hours GPU session <em>(Already 4 timeouts exceeded 2x 9H GPU, 2 x 3H TPU).</em> I found a solution to interrupt the training before the timeout is exceeded by <a href=\"https://www.kaggle.com/danmoller/make-best-use-of-a-kernel-s-limited-uptime-keras\">creating a custom callback in Keras </a>. It helps at least to publish the notebook and see how your training went.</p>\n\n<p>I usually split the training and inference into 2 notebooks but the training time itself requires +9 hours GPU.</p>\n\n<p>Is it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs <strong>on several 9H GPU sessions</strong> by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session? <em>(The only way I know of is loading the weights for prediction and not to retake the training).</em></p>\n\n<h2>Solution (Update):</h2>\n\n<p><a href=\"https://pytorch.org/tutorials/beginner/saving_loading_models.html\">Saving &amp; Loading a General Checkpoint to Resume Training with Pytorch\n</a></p>\n\n<p>Saving the model:\n<code>\ntorch.save({\n            'epoch': epoch,\n            'model_state_dict': model.state_dict(),\n            'optimizer_state_dict': optimizer.state_dict(),\n            'loss': loss,\n            ...\n            }, PATH)\n</code></p>\n\n<p>Loading the model to resume training: \n```\nmodel = TheModelClass(*args, **kwargs)\noptimizer = TheOptimizerClass(*args, **kwargs)</p>\n\n<p>checkpoint = torch.load(PATH)\nmodel.load_state_dict(checkpoint['model_state_dict'])\noptimizer.load_state_dict(checkpoint['optimizer_state_dict'])\nepoch = checkpoint['epoch']\nloss = checkpoint['loss']</p>\n\n<p>model.train()\n```\nI am still facing some issues with TensorFlow. I am trying to find a solution to how to train a saved model on a TPU.</p>\n\n<p><strong>Any tips or recommendations on how to make the most out of the kaggle GPU/TPU sessions are welcome :)</strong></p>",
  "messages": [
    {
      "id": "874673",
      "postDate": "06/05/2020 07:45:12",
      "content": "<p>I am starting this topic because this is my first kaggle competition and the challenge for me here is more how to train complex models with a limited computation power.</p>\n\n<p>I would like to share 2 interesting topics shared by <a href=\"/maxlenormand\">@maxlenormand</a> that helped me a lot to organize my workflow:\n<strong>1-</strong> <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/137024\">Debrief - Thoughts on computer vision with limited hardware </a>\n<strong>2-</strong> <a href=\"https://towardsdatascience.com/making-the-most-out-of-limited-hardware-for-computer-vision-kaggle-competitions-a8894511026e\">Making the most out of limited hardware for Computer Vision Kaggle competitions</a>.</p>\n\n<p>In order to increase the score I am trying more augmentations, deeper architectures or simply increase the image size. But any relatively small change leads to exceeding the allowed 9 hours GPU session <em>(Already 4 timeouts exceeded 2x 9H GPU, 2 x 3H TPU).</em> I found a solution to interrupt the training before the timeout is exceeded by <a href=\"https://www.kaggle.com/danmoller/make-best-use-of-a-kernel-s-limited-uptime-keras\">creating a custom callback in Keras </a>. It helps at least to publish the notebook and see how your training went.</p>\n\n<p>I usually split the training and inference into 2 notebooks but the training time itself requires +9 hours GPU.</p>\n\n<p>Is it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs <strong>on several 9H GPU sessions</strong> by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session? <em>(The only way I know of is loading the weights for prediction and not to retake the training).</em></p>\n\n<h2>Solution (Update):</h2>\n\n<p><a href=\"https://pytorch.org/tutorials/beginner/saving_loading_models.html\">Saving &amp; Loading a General Checkpoint to Resume Training with Pytorch\n</a></p>\n\n<p>Saving the model:\n<code>\ntorch.save({\n            'epoch': epoch,\n            'model_state_dict': model.state_dict(),\n            'optimizer_state_dict': optimizer.state_dict(),\n            'loss': loss,\n            ...\n            }, PATH)\n</code></p>\n\n<p>Loading the model to resume training: \n```\nmodel = TheModelClass(*args, **kwargs)\noptimizer = TheOptimizerClass(*args, **kwargs)</p>\n\n<p>checkpoint = torch.load(PATH)\nmodel.load_state_dict(checkpoint['model_state_dict'])\noptimizer.load_state_dict(checkpoint['optimizer_state_dict'])\nepoch = checkpoint['epoch']\nloss = checkpoint['loss']</p>\n\n<p>model.train()\n```\nI am still facing some issues with TensorFlow. I am trying to find a solution to how to train a saved model on a TPU.</p>\n\n<p><strong>Any tips or recommendations on how to make the most out of the kaggle GPU/TPU sessions are welcome :)</strong></p>",
      "rawMarkdown": "I am starting this topic because this is my first kaggle competition and the challenge for me here is more how to train complex models with a limited computation power.\n\nI would like to share 2 interesting topics shared by @maxlenormand that helped me a lot to organize my workflow:\n**1-** [Debrief - Thoughts on computer vision with limited hardware ](https://www.kaggle.com/c/bengaliai-cv19/discussion/137024)\n**2-** [Making the most out of limited hardware for Computer Vision Kaggle competitions](https://towardsdatascience.com/making-the-most-out-of-limited-hardware-for-computer-vision-kaggle-competitions-a8894511026e).\n\nIn order to increase the score I am trying more augmentations, deeper architectures or simply increase the image size. But any relatively small change leads to exceeding the allowed 9 hours GPU session *(Already 4 timeouts exceeded 2x 9H GPU, 2 x 3H TPU).* I found a solution to interrupt the training before the timeout is exceeded by [creating a custom callback in Keras ](https://www.kaggle.com/danmoller/make-best-use-of-a-kernel-s-limited-uptime-keras). It helps at least to publish the notebook and see how your training went.\n\nI usually split the training and inference into 2 notebooks but the training time itself requires +9 hours GPU.\n\nIs it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs **on several 9H GPU sessions** by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session? *(The only way I know of is loading the weights for prediction and not to retake the training).*\n\n## Solution (Update):\n\n[Saving &amp; Loading a General Checkpoint to Resume Training with Pytorch\n](https://pytorch.org/tutorials/beginner/saving_loading_models.html)\n\nSaving the model:\n```\ntorch.save({\n            'epoch': epoch,\n            'model_state_dict': model.state_dict(),\n            'optimizer_state_dict': optimizer.state_dict(),\n            'loss': loss,\n            ...\n            }, PATH)\n```\n\nLoading the model to resume training: \n```\nmodel = TheModelClass(*args, **kwargs)\noptimizer = TheOptimizerClass(*args, **kwargs)\n\ncheckpoint = torch.load(PATH)\nmodel.load_state_dict(checkpoint['model_state_dict'])\noptimizer.load_state_dict(checkpoint['optimizer_state_dict'])\nepoch = checkpoint['epoch']\nloss = checkpoint['loss']\n\n\nmodel.train()\n```\nI am still facing some issues with TensorFlow. I am trying to find a solution to how to train a saved model on a TPU.\n\n\n**Any tips or recommendations on how to make the most out of the kaggle GPU/TPU sessions are welcome :)**",
      "votes": null
    },
    {
      "id": "875087",
      "postDate": "06/05/2020 14:09:20",
      "content": "<p>Hey! Thanks for mentioning me! :)</p>\n\n<p>I've actually done this idea of splitting the training among multiple sessions, so short answer, yes that's totally possible. Now that I think about it, I forgot to put that in my article!</p>\n\n<p>I was saving the epoch number of the last completed epoch, and created a parameter to pick up training from that epoch (given also some weights to load and pick up from).</p>\n\n<p>This requires a bit of work, and you might loose incomplete training for a given a epoch, but it allows you to go to a lot more epochs overall than just what would fit in a 9h session.</p>\n\n<p>I'd be very interested in having some feedback on your side if you come up with some interesting ideas :)</p>",
      "rawMarkdown": "Hey! Thanks for mentioning me! :)\n\nI've actually done this idea of splitting the training among multiple sessions, so short answer, yes that's totally possible. Now that I think about it, I forgot to put that in my article!\n\nI was saving the epoch number of the last completed epoch, and created a parameter to pick up training from that epoch (given also some weights to load and pick up from).\n\nThis requires a bit of work, and you might loose incomplete training for a given a epoch, but it allows you to go to a lot more epochs overall than just what would fit in a 9h session.\n\nI'd be very interested in having some feedback on your side if you come up with some interesting ideas :)",
      "votes": null
    },
    {
      "id": "875112",
      "postDate": "06/05/2020 14:25:15",
      "content": "<p>&gt; Is it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs on several 9H GPU sessions by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session?</p>\n\n<p>Yes. I did this in Flower Comp to train a model in Kaggle notebooks for 27 hours. And I did this in Molecule comp to train a model in Kaggle notebooks for 180 hours! </p>\n\n<p>Just train for some epochs in one notebook and then save the weights <code>model.save_weights('weights.h5')</code> and remember where your learning schedule is. Then start a new notebook, build your model <code>model = build_model()</code>, then load the old weights <code>model.load_weights('weights.h5')</code>, then set your learning schedule to where it left off. And then continue training some more epochs with <code>model.fit()</code>. Then save, reload, repeat, over and over.</p>",
      "rawMarkdown": "&gt; Is it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs on several 9H GPU sessions by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session?\n\nYes. I did this in Flower Comp to train a model in Kaggle notebooks for 27 hours. And I did this in Molecule comp to train a model in Kaggle notebooks for 180 hours! \n\nJust train for some epochs in one notebook and then save the weights `model.save_weights('weights.h5')` and remember where your learning schedule is. Then start a new notebook, build your model `model = build_model()`, then load the old weights `model.load_weights('weights.h5')`, then set your learning schedule to where it left off. And then continue training some more epochs with `model.fit()`. Then save, reload, repeat, over and over.",
      "votes": null
    },
    {
      "id": "875310",
      "postDate": "06/05/2020 16:38:59",
      "content": "<blockquote>\n  <blockquote>\n    <p><strong>Chris Deotte wrote:</strong></p>\n    \n    <p>Just train for some epochs in one notebook and then save the weights <code>model.save_weights('weights.h5')</code> and remember where your learning schedule is. </p>\n  </blockquote>\n</blockquote>\n\n<p>You can print the last learning rate with this , if ever the training output is truncated</p>\n\n<p><code>\nfrom tensorflow.keras.backend import eval\nprint(eval(model.optimizer.lr))\n</code></p>",
      "rawMarkdown": "&gt;&gt; **Chris Deotte wrote:**\n \n&gt;&gt; Just train for some epochs in one notebook and then save the weights `model.save_weights('weights.h5')` and remember where your learning schedule is. \n\n\nYou can print the last learning rate with this , if ever the training output is truncated\n\n```\nfrom tensorflow.keras.backend import eval\nprint(eval(model.optimizer.lr))\n```",
      "votes": null
    },
    {
      "id": "875846",
      "postDate": "06/06/2020 07:51:06",
      "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> and <a href=\"/serigne\">@serigne</a>! It's such a relief to know that it's possible to train the same model on different sessions.</p>\n\n<p>I had some issues with locating the weights with <code>model= build_model()</code>. So I saved the whole model <code>model.save('weights.H5')</code>instead of <code>model.save_weights('weights.H5')</code></p>\n\n<p>```\nsaved_model_path='/kaggle/input/ef7-weights/efficientNetB7.h5'</p>\n\n<p>with strategy.scope():\n    model = tf.keras.models.load_model(saved_model_path)</p>\n\n<p>model.compile(loss=tf.keras.losses.binary_crossentropy,\n                optimizer='adam',\n                metrics=['binary_crossentropy'])\n```</p>\n\n<p>There is something wrong. Is it necessary to save it in tf format <code>(save_format=\"tf\")</code>and upload it to GCS? How did you deal with this problem in the Flower Comp?</p>",
      "rawMarkdown": "Thank you @cdeotte and @serigne! It's such a relief to know that it's possible to train the same model on different sessions.\n\nI had some issues with locating the weights with `model= build_model()`. So I saved the whole model `model.save('weights.H5')`instead of `model.save_weights('weights.H5')`\n\n\n```\nsaved_model_path='/kaggle/input/ef7-weights/efficientNetB7.h5'\n\nwith strategy.scope():\n    model = tf.keras.models.load_model(saved_model_path)\n\nmodel.compile(loss=tf.keras.losses.binary_crossentropy,\n                optimizer='adam',\n                metrics=['binary_crossentropy'])\n```\n\nThere is something wrong. Is it necessary to save it in tf format `(save_format=\"tf\")`and upload it to GCS? How did you deal with this problem in the Flower Comp?",
      "votes": null
    },
    {
      "id": "875852",
      "postDate": "06/06/2020 07:55:54",
      "content": "<p>Thanks <a href=\"/maxlenormand\">@maxlenormand</a> for your feedback. I will definitely give you some feedback if I come up with any new tricks. Also, I would like to discuss your technique to send a message/report of your training to your slack channel. I am still trying to figure out how to implement it :)</p>",
      "rawMarkdown": "Thanks @maxlenormand for your feedback. I will definitely give you some feedback if I come up with any new tricks. Also, I would like to discuss your technique to send a message/report of your training to your slack channel. I am still trying to figure out how to implement it :)",
      "votes": null
    },
    {
      "id": "876996",
      "postDate": "06/07/2020 08:22:52",
      "content": "<p>Sure, feel free to ask, here, or a PM :)</p>",
      "rawMarkdown": "Sure, feel free to ask, here, or a PM :)",
      "votes": null
    },
    {
      "id": "877497",
      "postDate": "06/07/2020 16:55:02",
      "content": "<p>Guys, who has experience with Colab TPU? I like, how it is working, but don't understand how to feed data properly - simple loading them into Google Cloud doesn't help. Do I have to transfer them to GCS Backets (which is not free), or there are other ways?</p>",
      "rawMarkdown": "Guys, who has experience with Colab TPU? I like, how it is working, but don't understand how to feed data properly - simple loading them into Google Cloud doesn't help. Do I have to transfer them to GCS Backets (which is not free), or there are other ways?",
      "votes": null
    },
    {
      "id": "877559",
      "postDate": "06/07/2020 17:36:56",
      "content": "<p>I’m using it extensively. You don’t have to move anything, just get the location of the files from printing it out in a Kaggle note book and then copying it to your colab notebook. There are a few lines you will need to add and a few to comment out but it’s no problem. Unfortunately I’m at work in a lunch break on my phone so I can’t easily cut and paste a lot of help. I won’t be home for 5 hours . I’ll give you a better answer then if you still need it.</p>",
      "rawMarkdown": "I’m using it extensively. You don’t have to move anything, just get the location of the files from printing it out in a Kaggle note book and then copying it to your colab notebook. There are a few lines you will need to add and a few to comment out but it’s no problem. Unfortunately I’m at work in a lunch break on my phone so I can’t easily cut and paste a lot of help. I won’t be home for 5 hours . I’ll give you a better answer then if you still need it.",
      "votes": null
    },
    {
      "id": "877747",
      "postDate": "06/07/2020 23:22:44",
      "content": "<p>Ok Serge,</p>\n\n<p>I'm assuming that you already have a Kaggle notebook that uses TPU for this competition. If not there are a few public ones that you could use as a starting point such as this one by Ajay Kumar.\n<a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">Melanoma TPU EfficientNet B5_dense_head</a></p>\n\n<p>to get the Google data path you print out the result of,\n```</p>\n\n<h1>Data access</h1>\n\n<p>GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\nprint(GCS_PATH)\n<code>``\ngiving:\n</code>gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1`</p>\n\n<p>Now, this may change over time and you may have to redo this step to find a new link.\nIf you want to use other data sources you do a similar thing. Add the data source to your Kaggle notebook and construct a GCS_PATH as you did for the main data.\nAdded 2 other data sources|</p>\n\n<p>```\nGCS_PATH2 = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH3 = KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images')</p>\n\n<p>```\n and printed out their Google address.</p>\n\n<p><code>print(GCS_PATH2)</code>\ngs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569</p>\n\n<p><code>print(GCS_PATH3)</code>\ngs://kds-8e5c826f476888a71d6db98e2dc6ceb491e14202931bdf249c60dc97</p>\n\n<p>Please do this for yourself as the paths here may have changed.</p>\n\n<p>Now for the Google Colab part of the story...\nDownload your Kaggle notebook as a Jupyter Notebook and upload it into Colab.</p>\n\n<p>You'll need to install and load the google file system at the beginning of your code.</p>\n\n<p><code>\n!pip install gcsfs\nimport gcsfs\n</code></p>\n\n<p>You won't need this while in Colab so comment it out.\n<code>#from kaggle_datasets import KaggleDatasets</code></p>\n\n<p>So now instead of getting Kaggle to get the Google address of the data, you can comment that out.</p>\n\n<p>```</p>\n\n<h1>Data access</h1>\n\n<h1>GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')</h1>\n\n<p>```</p>\n\n<p>and just use the info you have obtained from the Kaggle notebook.</p>\n\n<p><code>GCS_PATH = 'gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1' #original data</code></p>\n\n<p>Same again if you are using other data sources such as these. (the paths may be wrong now so you need to get them yourself)</p>\n\n<p>```\nGCS_PATH2 = 'gs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569'  #original data at 256 + metadata</p>\n\n<h1>GCS_PATH3 = 'gs://kds-9e42a1abc97f350d6a801f68f97667cf54db33a6ccbea9b3757df69a'  #original data at 512 + metadata</h1>\n\n<p>```\nIf you use the original data your paths to the train and submission CSV files will be.</p>\n\n<p>```\ntrain = pd.read_csv(GCS_PATH + '/train.csv')\nsub = pd.read_csv(GCS_PATH + '/sample_submission.csv')</p>\n\n<p>```\nThe paths to the original data TFRecords will be. (as it is in a Kaggle notebook)</p>\n\n<p>```\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/test*.tfrec')</p>\n\n<p>```\nor if you use one of the sets graciously made by Chris it will be more like. (as it would be in a Kaggle notebook though the path to the CSV files will still be the original GCS_PATH)</p>\n\n<p><code>\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/test*.tfrec')\n</code></p>\n\n<p>That should get you started. <br>\nI hope it works for you. I have used the same basic method for using Colab TPUs for both TFRecords and JPEGS  using TPU.</p>\n\n<p>Make sure you actually turn on the TPU accelerator using Edit / Notebook Settings.</p>\n\n<p>Good luck!</p>\n\n<p>P.S. THe Colab TPUs aren't as \"powerful' as those on Kaggle so you will most likely have to dial back your Batch size a bit so that you won't run out of memory.</p>",
      "rawMarkdown": "Ok Serge,\n\nI'm assuming that you already have a Kaggle notebook that uses TPU for this competition. If not there are a few public ones that you could use as a starting point such as this one by Ajay Kumar.\n[Melanoma TPU EfficientNet B5_dense_head](https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head)\n\nto get the Google data path you print out the result of,\n```\n# Data access\nGCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\nprint(GCS_PATH)\n```\ngiving:\n`gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1`\n\nNow, this may change over time and you may have to redo this step to find a new link.\nIf you want to use other data sources you do a similar thing. Add the data source to your Kaggle notebook and construct a GCS_PATH as you did for the main data.\nAdded 2 other data sources|\n\n```\nGCS_PATH2 = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH3 = KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images')\n\n```\n and printed out their Google address.\n\n`print(GCS_PATH2)`\ngs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569\n\n`print(GCS_PATH3)`\ngs://kds-8e5c826f476888a71d6db98e2dc6ceb491e14202931bdf249c60dc97\n\nPlease do this for yourself as the paths here may have changed.\n\nNow for the Google Colab part of the story...\nDownload your Kaggle notebook as a Jupyter Notebook and upload it into Colab.\n\nYou'll need to install and load the google file system at the beginning of your code.\n\n```\n!pip install gcsfs\nimport gcsfs\n```\n\nYou won't need this while in Colab so comment it out.\n`#from kaggle_datasets import KaggleDatasets`\n\nSo now instead of getting Kaggle to get the Google address of the data, you can comment that out.\n\n```\n# Data access\n#GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\n```\n\nand just use the info you have obtained from the Kaggle notebook.\n\n`GCS_PATH = 'gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1' #original data`\n\nSame again if you are using other data sources such as these. (the paths may be wrong now so you need to get them yourself)\n\n```\nGCS_PATH2 = 'gs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569'  #original data at 256 + metadata\n\n#GCS_PATH3 = 'gs://kds-9e42a1abc97f350d6a801f68f97667cf54db33a6ccbea9b3757df69a'  #original data at 512 + metadata\n```\nIf you use the original data your paths to the train and submission CSV files will be.\n\n```\ntrain = pd.read_csv(GCS_PATH + '/train.csv')\nsub = pd.read_csv(GCS_PATH + '/sample_submission.csv')\n\n```\nThe paths to the original data TFRecords will be. (as it is in a Kaggle notebook)\n\n```\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/test*.tfrec')\n\n```\nor if you use one of the sets graciously made by Chris it will be more like. (as it would be in a Kaggle notebook though the path to the CSV files will still be the original GCS_PATH)\n\n```\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/test*.tfrec')\n```\n\nThat should get you started.  \nI hope it works for you. I have used the same basic method for using Colab TPUs for both TFRecords and JPEGS  using TPU.\n\nMake sure you actually turn on the TPU accelerator using Edit / Notebook Settings.\n\nGood luck!\n\nP.S. THe Colab TPUs aren't as \"powerful' as those on Kaggle so you will most likely have to dial back your Batch size a bit so that you won't run out of memory.",
      "votes": null
    },
    {
      "id": "878326",
      "postDate": "06/08/2020 12:55:50",
      "content": "<p>IMHO  saving only model weights is not sufficent for complete  continuity of trainng process. It is also necessary to save and restore optimizer weights. During bengali comp i used the next TF construction:</p>\n\n<p>```\nopt=tf.keras.optimizers.Adam(learning_rate=LR_MAX)\ncheckpoint =  tf.train.Checkpoint(latest_epoch=tf.Variable(0), optimizer=opt, model=model)\nchkp_manager =  tf.train.CheckpointManager(checkpoint, CHKP_DIR, max_to_keep=3)</p>\n\n<p>chkp_restore_status=checkpoint.restore(chkp_manager.latest_checkpoint)</p>\n\n<p>if chkp_manager.latest_checkpoint:\n  print(\"Restored from {}\".format(chkp_manager.latest_checkpoint))\n  print(f\"The Latest epoch {checkpoint.latest_epoch.numpy()}\")\nelse:\n  print(\"Initializing from scratch.\")</p>\n\n<p>model.compile(optimizer=opt, loss=loss_func)\ninitial_epoch = checkpoint.latest_epoch.numpy()</p>\n\n<p>history=model.fit(datagenerator,\n                  epochs=EPOCHS,\n                  steps_per_epoch=spe,\n                  callbacks=[my_callback(initial_epoch=initial_epoch, chkp_manager=chkp_manager, ... )],\n                  verbose=1)\n```</p>",
      "rawMarkdown": "IMHO  saving only model weights is not sufficent for complete  continuity of trainng process. It is also necessary to save and restore optimizer weights. During bengali comp i used the next TF construction:\n\n```\nopt=tf.keras.optimizers.Adam(learning_rate=LR_MAX)\ncheckpoint =  tf.train.Checkpoint(latest_epoch=tf.Variable(0), optimizer=opt, model=model)\nchkp_manager =  tf.train.CheckpointManager(checkpoint, CHKP_DIR, max_to_keep=3)\n\nchkp_restore_status=checkpoint.restore(chkp_manager.latest_checkpoint)\n\nif chkp_manager.latest_checkpoint:\n  print(\"Restored from {}\".format(chkp_manager.latest_checkpoint))\n  print(f\"The Latest epoch {checkpoint.latest_epoch.numpy()}\")\nelse:\n  print(\"Initializing from scratch.\")\n\nmodel.compile(optimizer=opt, loss=loss_func)\ninitial_epoch = checkpoint.latest_epoch.numpy()\n\nhistory=model.fit(datagenerator,\n                  epochs=EPOCHS,\n                  steps_per_epoch=spe,\n                  callbacks=[my_callback(initial_epoch=initial_epoch, chkp_manager=chkp_manager, ... )],\n                  verbose=1)\n```",
      "votes": null
    },
    {
      "id": "878340",
      "postDate": "06/08/2020 13:06:18",
      "content": "<p>Yes! It will be more correct to save the whole model, but if your model contains custom objects (for example spicific activation functions or custom layers) load_model dont work correctly.</p>",
      "rawMarkdown": "Yes! It will be more correct to save the whole model, but if your model contains custom objects (for example spicific activation functions or custom layers) load_model dont work correctly.",
      "votes": null
    },
    {
      "id": "881231",
      "postDate": "06/10/2020 20:51:55",
      "content": "<p>Thanks! \nI want to try Colab, because when I've used all memory it gives me more, then more again - up to 30 G. Plus I am not that limited by time running.</p>",
      "rawMarkdown": "Thanks! \nI want to try Colab, because when I've used all memory it gives me more, then more again - up to 30 G. Plus I am not that limited by time running.",
      "votes": null
    },
    {
      "id": "881780",
      "postDate": "06/11/2020 11:08:50",
      "content": "<p>You might be interested in resource-efficient convolutions, splitting up 2D Conv and channel mixing.\n<a href=\"https://towardsdatascience.com/a-basic-introduction-to-separable-convolutions-b99ec3102728\">https://towardsdatascience.com/a-basic-introduction-to-separable-convolutions-b99ec3102728</a>\nIf you want to use transfer learning, some tweaks to your model will be necessary!</p>\n\n<p>In fact, it is one of the reasons why EfficientNet is so efficient ;) so in fact, adopting EfficientNet is also related to this topic, but too obvious...</p>",
      "rawMarkdown": "You might be interested in resource-efficient convolutions, splitting up 2D Conv and channel mixing.\nhttps://towardsdatascience.com/a-basic-introduction-to-separable-convolutions-b99ec3102728\nIf you want to use transfer learning, some tweaks to your model will be necessary!\n\nIn fact, it is one of the reasons why EfficientNet is so efficient ;) so in fact, adopting EfficientNet is also related to this topic, but too obvious...",
      "votes": null
    },
    {
      "id": "889929",
      "postDate": "06/17/2020 08:07:16",
      "content": "<p>Thank you <a href=\"/ovdnnest\">@ovdnnest</a> for sharing this great article! Did you try to change all the normal convolutions to depthwise convolutions? Please share you experience :) Doing this reduces the number of parameters, hence the training time but don't you think we will have a Parameters-Loss trade-off? Less parameters = Higher loss?</p>",
      "rawMarkdown": "Thank you @ovdnnest for sharing this great article! Did you try to change all the normal convolutions to depthwise convolutions? Please share you experience :) Doing this reduces the number of parameters, hence the training time but don't you think we will have a Parameters-Loss trade-off? Less parameters = Higher loss?",
      "votes": null
    },
    {
      "id": "890429",
      "postDate": "06/17/2020 14:03:03",
      "content": "<p>The 2D Conv and channel mixing is already separated in EfficientNet. For other models, I have not tried adapting the architecture. It was more a theoretical tip ;)\nAnd there is no direct relationship between the number of parameters and loss, as you probably know. In the traditional bias-variance theory you need to increase your amount until you reach the interpolation regime. However recently double-descent is observed in some deep NN tasks, meaning that very large models generalize better after a larger amount of epochs (in contrast to bias-variance theory). If you are interested in that, check this out: <a href=\"https://arxiv.org/abs/1912.02292\">https://arxiv.org/abs/1912.02292</a></p>\n\n<p>Short answer: it's possible your model generalizes better with fewer parameters or needs to have more parameters to be able to overfit the training data first. This is something you have to determine in practice, there is no direct relation.</p>",
      "rawMarkdown": "The 2D Conv and channel mixing is already separated in EfficientNet. For other models, I have not tried adapting the architecture. It was more a theoretical tip ;)\nAnd there is no direct relationship between the number of parameters and loss, as you probably know. In the traditional bias-variance theory you need to increase your amount until you reach the interpolation regime. However recently double-descent is observed in some deep NN tasks, meaning that very large models generalize better after a larger amount of epochs (in contrast to bias-variance theory). If you are interested in that, check this out: https://arxiv.org/abs/1912.02292\n\nShort answer: it's possible your model generalizes better with fewer parameters or needs to have more parameters to be able to overfit the training data first. This is something you have to determine in practice, there is no direct relation.",
      "votes": null
    },
    {
      "id": "896769",
      "postDate": "06/22/2020 12:25:14",
      "content": "<p>Hi <a href=\"/mutantspore\">@mutantspore</a>  Thank you for this. I am trying to open JPG files in colab and I am missing something.  For e.g: this file exists  - \n<code>!gsutil ls gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg</code> </p>\n\n<p>However this fails  - <code>PIL.Image.open('gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg')</code>  with FileNotFoundError: [Errno 2] No such file or directory:  - </p>\n\n<p>I think I am not suppose to pass this path directly, Any workarounds ? If you have a moment please let me know.  Thank you </p>",
      "rawMarkdown": "Hi @mutantspore  Thank you for this. I am trying to open JPG files in colab and I am missing something.  For e.g: this file exists  - \n`!gsutil ls gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg` \n\nHowever this fails  - `PIL.Image.open('gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg')`  with FileNotFoundError: [Errno 2] No such file or directory:  - \n\nI think I am not suppose to pass this path directly, Any workarounds ? If you have a moment please let me know.  Thank you",
      "votes": null
    },
    {
      "id": "912923",
      "postDate": "07/02/2020 20:21:15",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<blockquote>\n  <p>Just train for some epochs in one notebook and then save the weights <code>model.save_weights('weights.h5')</code> and remember where your learning schedule is. Then start a new notebook, build your model <code>model = build_model()</code>, then load the old weights <code>model.load_weights('weights.h5')</code>, then set your learning schedule to where it left off. And then continue training some more epochs with <code>model.fit()</code>. Then save, reload, repeat, over and over.</p>\n</blockquote>\n\n<p>Can you kindly tell me how you load weights in second notebook after saving from 1st notebook? As right now I just know that I need to first save the weights/model from 1st notebook and then download the trained weights/model from the notebook and then open second notebook and upload weights/model to this notebook to start further training. Is there some other shortcut and easy way to do this ?</p>\n\n<p>Thanks in advance.</p>",
      "rawMarkdown": "cdeotte \n\n&gt; Just train for some epochs in one notebook and then save the weights `model.save_weights('weights.h5')` and remember where your learning schedule is. Then start a new notebook, build your model `model = build_model()`, then load the old weights `model.load_weights('weights.h5')`, then set your learning schedule to where it left off. And then continue training some more epochs with `model.fit()`. Then save, reload, repeat, over and over.\n\nCan you kindly tell me how you load weights in second notebook after saving from 1st notebook? As right now I just know that I need to first save the weights/model from 1st notebook and then download the trained weights/model from the notebook and then open second notebook and upload weights/model to this notebook to start further training. Is there some other shortcut and easy way to do this ?\n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "914950",
      "postDate": "07/04/2020 11:27:27",
      "content": "<p><a href=\"/rashmibanthia\">@rashmibanthia</a> did you find some solution of the above problem. I am also facing the same issue while reading <code>jpg</code>or <code>png</code>file but csv are working fine in my case.</p>",
      "rawMarkdown": "rashmibanthia did you find some solution of the above problem. I am also facing the same issue while reading `jpg `or `png `file but csv are working fine in my case.",
      "votes": null
    },
    {
      "id": "914978",
      "postDate": "07/04/2020 11:55:57",
      "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a>  Not really, I tried using my own GCS bucket and using <code>bucket.get_blob(jpgfile)</code>. It worked but it was too slow.  I am using pytorch, the best I could find is to use <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">this</a> external data and unzip it on colab or use .npy files. If you are using TF you should be able to use TF Records.</p>",
      "rawMarkdown": "abdurrehman245  Not really, I tried using my own GCS bucket and using `bucket.get_blob(jpgfile)`. It worked but it was too slow.  I am using pytorch, the best I could find is to use [this](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg) external data and unzip it on colab or use .npy files. If you are using TF you should be able to use TF Records.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 875087,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "06/05/2020 14:09:20",
      "content": "<p>Hey! Thanks for mentioning me! :)</p>\n\n<p>I've actually done this idea of splitting the training among multiple sessions, so short answer, yes that's totally possible. Now that I think about it, I forgot to put that in my article!</p>\n\n<p>I was saving the epoch number of the last completed epoch, and created a parameter to pick up training from that epoch (given also some weights to load and pick up from).</p>\n\n<p>This requires a bit of work, and you might loose incomplete training for a given a epoch, but it allows you to go to a lot more epochs overall than just what would fit in a 9h session.</p>\n\n<p>I'd be very interested in having some feedback on your side if you come up with some interesting ideas :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 875852,
          "author_name": "amiiiney",
          "author_url": "",
          "post_date": "06/06/2020 07:55:54",
          "content": "<p>Thanks <a href=\"/maxlenormand\">@maxlenormand</a> for your feedback. I will definitely give you some feedback if I come up with any new tricks. Also, I would like to discuss your technique to send a message/report of your training to your slack channel. I am still trying to figure out how to implement it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 876996,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "06/07/2020 08:22:52",
          "content": "<p>Sure, feel free to ask, here, or a PM :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 875112,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/05/2020 14:25:15",
      "content": "<p>&gt; Is it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs on several 9H GPU sessions by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session?</p>\n\n<p>Yes. I did this in Flower Comp to train a model in Kaggle notebooks for 27 hours. And I did this in Molecule comp to train a model in Kaggle notebooks for 180 hours! </p>\n\n<p>Just train for some epochs in one notebook and then save the weights <code>model.save_weights('weights.h5')</code> and remember where your learning schedule is. Then start a new notebook, build your model <code>model = build_model()</code>, then load the old weights <code>model.load_weights('weights.h5')</code>, then set your learning schedule to where it left off. And then continue training some more epochs with <code>model.fit()</code>. Then save, reload, repeat, over and over.</p>",
      "votes": null,
      "replies": [
        {
          "id": 875310,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "06/05/2020 16:38:59",
          "content": "<blockquote>\n  <blockquote>\n    <p><strong>Chris Deotte wrote:</strong></p>\n    \n    <p>Just train for some epochs in one notebook and then save the weights <code>model.save_weights('weights.h5')</code> and remember where your learning schedule is. </p>\n  </blockquote>\n</blockquote>\n\n<p>You can print the last learning rate with this , if ever the training output is truncated</p>\n\n<p><code>\nfrom tensorflow.keras.backend import eval\nprint(eval(model.optimizer.lr))\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 875846,
          "author_name": "amiiiney",
          "author_url": "",
          "post_date": "06/06/2020 07:51:06",
          "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a> and <a href=\"/serigne\">@serigne</a>! It's such a relief to know that it's possible to train the same model on different sessions.</p>\n\n<p>I had some issues with locating the weights with <code>model= build_model()</code>. So I saved the whole model <code>model.save('weights.H5')</code>instead of <code>model.save_weights('weights.H5')</code></p>\n\n<p>```\nsaved_model_path='/kaggle/input/ef7-weights/efficientNetB7.h5'</p>\n\n<p>with strategy.scope():\n    model = tf.keras.models.load_model(saved_model_path)</p>\n\n<p>model.compile(loss=tf.keras.losses.binary_crossentropy,\n                optimizer='adam',\n                metrics=['binary_crossentropy'])\n```</p>\n\n<p>There is something wrong. Is it necessary to save it in tf format <code>(save_format=\"tf\")</code>and upload it to GCS? How did you deal with this problem in the Flower Comp?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 878326,
          "author_name": "andreyzotov",
          "author_url": "",
          "post_date": "06/08/2020 12:55:50",
          "content": "<p>IMHO  saving only model weights is not sufficent for complete  continuity of trainng process. It is also necessary to save and restore optimizer weights. During bengali comp i used the next TF construction:</p>\n\n<p>```\nopt=tf.keras.optimizers.Adam(learning_rate=LR_MAX)\ncheckpoint =  tf.train.Checkpoint(latest_epoch=tf.Variable(0), optimizer=opt, model=model)\nchkp_manager =  tf.train.CheckpointManager(checkpoint, CHKP_DIR, max_to_keep=3)</p>\n\n<p>chkp_restore_status=checkpoint.restore(chkp_manager.latest_checkpoint)</p>\n\n<p>if chkp_manager.latest_checkpoint:\n  print(\"Restored from {}\".format(chkp_manager.latest_checkpoint))\n  print(f\"The Latest epoch {checkpoint.latest_epoch.numpy()}\")\nelse:\n  print(\"Initializing from scratch.\")</p>\n\n<p>model.compile(optimizer=opt, loss=loss_func)\ninitial_epoch = checkpoint.latest_epoch.numpy()</p>\n\n<p>history=model.fit(datagenerator,\n                  epochs=EPOCHS,\n                  steps_per_epoch=spe,\n                  callbacks=[my_callback(initial_epoch=initial_epoch, chkp_manager=chkp_manager, ... )],\n                  verbose=1)\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 878340,
          "author_name": "andreyzotov",
          "author_url": "",
          "post_date": "06/08/2020 13:06:18",
          "content": "<p>Yes! It will be more correct to save the whole model, but if your model contains custom objects (for example spicific activation functions or custom layers) load_model dont work correctly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912923,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "07/02/2020 20:21:15",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<blockquote>\n  <p>Just train for some epochs in one notebook and then save the weights <code>model.save_weights('weights.h5')</code> and remember where your learning schedule is. Then start a new notebook, build your model <code>model = build_model()</code>, then load the old weights <code>model.load_weights('weights.h5')</code>, then set your learning schedule to where it left off. And then continue training some more epochs with <code>model.fit()</code>. Then save, reload, repeat, over and over.</p>\n</blockquote>\n\n<p>Can you kindly tell me how you load weights in second notebook after saving from 1st notebook? As right now I just know that I need to first save the weights/model from 1st notebook and then download the trained weights/model from the notebook and then open second notebook and upload weights/model to this notebook to start further training. Is there some other shortcut and easy way to do this ?</p>\n\n<p>Thanks in advance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 877497,
      "author_name": "serg132003",
      "author_url": "",
      "post_date": "06/07/2020 16:55:02",
      "content": "<p>Guys, who has experience with Colab TPU? I like, how it is working, but don't understand how to feed data properly - simple loading them into Google Cloud doesn't help. Do I have to transfer them to GCS Backets (which is not free), or there are other ways?</p>",
      "votes": null,
      "replies": [
        {
          "id": 877559,
          "author_name": "mutantspore",
          "author_url": "",
          "post_date": "06/07/2020 17:36:56",
          "content": "<p>I’m using it extensively. You don’t have to move anything, just get the location of the files from printing it out in a Kaggle note book and then copying it to your colab notebook. There are a few lines you will need to add and a few to comment out but it’s no problem. Unfortunately I’m at work in a lunch break on my phone so I can’t easily cut and paste a lot of help. I won’t be home for 5 hours . I’ll give you a better answer then if you still need it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 877747,
          "author_name": "mutantspore",
          "author_url": "",
          "post_date": "06/07/2020 23:22:44",
          "content": "<p>Ok Serge,</p>\n\n<p>I'm assuming that you already have a Kaggle notebook that uses TPU for this competition. If not there are a few public ones that you could use as a starting point such as this one by Ajay Kumar.\n<a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">Melanoma TPU EfficientNet B5_dense_head</a></p>\n\n<p>to get the Google data path you print out the result of,\n```</p>\n\n<h1>Data access</h1>\n\n<p>GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\nprint(GCS_PATH)\n<code>``\ngiving:\n</code>gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1`</p>\n\n<p>Now, this may change over time and you may have to redo this step to find a new link.\nIf you want to use other data sources you do a similar thing. Add the data source to your Kaggle notebook and construct a GCS_PATH as you did for the main data.\nAdded 2 other data sources|</p>\n\n<p>```\nGCS_PATH2 = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH3 = KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images')</p>\n\n<p>```\n and printed out their Google address.</p>\n\n<p><code>print(GCS_PATH2)</code>\ngs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569</p>\n\n<p><code>print(GCS_PATH3)</code>\ngs://kds-8e5c826f476888a71d6db98e2dc6ceb491e14202931bdf249c60dc97</p>\n\n<p>Please do this for yourself as the paths here may have changed.</p>\n\n<p>Now for the Google Colab part of the story...\nDownload your Kaggle notebook as a Jupyter Notebook and upload it into Colab.</p>\n\n<p>You'll need to install and load the google file system at the beginning of your code.</p>\n\n<p><code>\n!pip install gcsfs\nimport gcsfs\n</code></p>\n\n<p>You won't need this while in Colab so comment it out.\n<code>#from kaggle_datasets import KaggleDatasets</code></p>\n\n<p>So now instead of getting Kaggle to get the Google address of the data, you can comment that out.</p>\n\n<p>```</p>\n\n<h1>Data access</h1>\n\n<h1>GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')</h1>\n\n<p>```</p>\n\n<p>and just use the info you have obtained from the Kaggle notebook.</p>\n\n<p><code>GCS_PATH = 'gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1' #original data</code></p>\n\n<p>Same again if you are using other data sources such as these. (the paths may be wrong now so you need to get them yourself)</p>\n\n<p>```\nGCS_PATH2 = 'gs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569'  #original data at 256 + metadata</p>\n\n<h1>GCS_PATH3 = 'gs://kds-9e42a1abc97f350d6a801f68f97667cf54db33a6ccbea9b3757df69a'  #original data at 512 + metadata</h1>\n\n<p>```\nIf you use the original data your paths to the train and submission CSV files will be.</p>\n\n<p>```\ntrain = pd.read_csv(GCS_PATH + '/train.csv')\nsub = pd.read_csv(GCS_PATH + '/sample_submission.csv')</p>\n\n<p>```\nThe paths to the original data TFRecords will be. (as it is in a Kaggle notebook)</p>\n\n<p>```\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/test*.tfrec')</p>\n\n<p>```\nor if you use one of the sets graciously made by Chris it will be more like. (as it would be in a Kaggle notebook though the path to the CSV files will still be the original GCS_PATH)</p>\n\n<p><code>\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/test*.tfrec')\n</code></p>\n\n<p>That should get you started. <br>\nI hope it works for you. I have used the same basic method for using Colab TPUs for both TFRecords and JPEGS  using TPU.</p>\n\n<p>Make sure you actually turn on the TPU accelerator using Edit / Notebook Settings.</p>\n\n<p>Good luck!</p>\n\n<p>P.S. THe Colab TPUs aren't as \"powerful' as those on Kaggle so you will most likely have to dial back your Batch size a bit so that you won't run out of memory.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 881231,
          "author_name": "serg132003",
          "author_url": "",
          "post_date": "06/10/2020 20:51:55",
          "content": "<p>Thanks! \nI want to try Colab, because when I've used all memory it gives me more, then more again - up to 30 G. Plus I am not that limited by time running.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 896769,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "06/22/2020 12:25:14",
          "content": "<p>Hi <a href=\"/mutantspore\">@mutantspore</a>  Thank you for this. I am trying to open JPG files in colab and I am missing something.  For e.g: this file exists  - \n<code>!gsutil ls gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg</code> </p>\n\n<p>However this fails  - <code>PIL.Image.open('gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg')</code>  with FileNotFoundError: [Errno 2] No such file or directory:  - </p>\n\n<p>I think I am not suppose to pass this path directly, Any workarounds ? If you have a moment please let me know.  Thank you </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914950,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "07/04/2020 11:27:27",
          "content": "<p><a href=\"/rashmibanthia\">@rashmibanthia</a> did you find some solution of the above problem. I am also facing the same issue while reading <code>jpg</code>or <code>png</code>file but csv are working fine in my case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914978,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "07/04/2020 11:55:57",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a>  Not really, I tried using my own GCS bucket and using <code>bucket.get_blob(jpgfile)</code>. It worked but it was too slow.  I am using pytorch, the best I could find is to use <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">this</a> external data and unzip it on colab or use .npy files. If you are using TF you should be able to use TF Records.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 881780,
      "author_name": "ovdnnest",
      "author_url": "",
      "post_date": "06/11/2020 11:08:50",
      "content": "<p>You might be interested in resource-efficient convolutions, splitting up 2D Conv and channel mixing.\n<a href=\"https://towardsdatascience.com/a-basic-introduction-to-separable-convolutions-b99ec3102728\">https://towardsdatascience.com/a-basic-introduction-to-separable-convolutions-b99ec3102728</a>\nIf you want to use transfer learning, some tweaks to your model will be necessary!</p>\n\n<p>In fact, it is one of the reasons why EfficientNet is so efficient ;) so in fact, adopting EfficientNet is also related to this topic, but too obvious...</p>",
      "votes": null,
      "replies": [
        {
          "id": 889929,
          "author_name": "amiiiney",
          "author_url": "",
          "post_date": "06/17/2020 08:07:16",
          "content": "<p>Thank you <a href=\"/ovdnnest\">@ovdnnest</a> for sharing this great article! Did you try to change all the normal convolutions to depthwise convolutions? Please share you experience :) Doing this reduces the number of parameters, hence the training time but don't you think we will have a Parameters-Loss trade-off? Less parameters = Higher loss?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 890429,
          "author_name": "ovdnnest",
          "author_url": "",
          "post_date": "06/17/2020 14:03:03",
          "content": "<p>The 2D Conv and channel mixing is already separated in EfficientNet. For other models, I have not tried adapting the architecture. It was more a theoretical tip ;)\nAnd there is no direct relationship between the number of parameters and loss, as you probably know. In the traditional bias-variance theory you need to increase your amount until you reach the interpolation regime. However recently double-descent is observed in some deep NN tasks, meaning that very large models generalize better after a larger amount of epochs (in contrast to bias-variance theory). If you are interested in that, check this out: <a href=\"https://arxiv.org/abs/1912.02292\">https://arxiv.org/abs/1912.02292</a></p>\n\n<p>Short answer: it's possible your model generalizes better with fewer parameters or needs to have more parameters to be able to overfit the training data first. This is something you have to determine in practice, there is no direct relation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "874673": "I am starting this topic because this is my first kaggle competition and the challenge for me here is more how to train complex models with a limited computation power.\n\nI would like to share 2 interesting topics shared by @maxlenormand that helped me a lot to organize my workflow:\n**1-** [Debrief - Thoughts on computer vision with limited hardware ](https://www.kaggle.com/c/bengaliai-cv19/discussion/137024)\n**2-** [Making the most out of limited hardware for Computer Vision Kaggle competitions](https://towardsdatascience.com/making-the-most-out-of-limited-hardware-for-computer-vision-kaggle-competitions-a8894511026e).\n\nIn order to increase the score I am trying more augmentations, deeper architectures or simply increase the image size. But any relatively small change leads to exceeding the allowed 9 hours GPU session *(Already 4 timeouts exceeded 2x 9H GPU, 2 x 3H TPU).* I found a solution to interrupt the training before the timeout is exceeded by [creating a custom callback in Keras ](https://www.kaggle.com/danmoller/make-best-use-of-a-kernel-s-limited-uptime-keras). It helps at least to publish the notebook and see how your training went.\n\nI usually split the training and inference into 2 notebooks but the training time itself requires +9 hours GPU.\n\nIs it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs **on several 9H GPU sessions** by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session? *(The only way I know of is loading the weights for prediction and not to retake the training).*\n\n## Solution (Update):\n\n[Saving &amp; Loading a General Checkpoint to Resume Training with Pytorch\n](https://pytorch.org/tutorials/beginner/saving_loading_models.html)\n\nSaving the model:\n```\ntorch.save({\n            'epoch': epoch,\n            'model_state_dict': model.state_dict(),\n            'optimizer_state_dict': optimizer.state_dict(),\n            'loss': loss,\n            ...\n            }, PATH)\n```\n\nLoading the model to resume training: \n```\nmodel = TheModelClass(*args, **kwargs)\noptimizer = TheOptimizerClass(*args, **kwargs)\n\ncheckpoint = torch.load(PATH)\nmodel.load_state_dict(checkpoint['model_state_dict'])\noptimizer.load_state_dict(checkpoint['optimizer_state_dict'])\nepoch = checkpoint['epoch']\nloss = checkpoint['loss']\n\n\nmodel.train()\n```\nI am still facing some issues with TensorFlow. I am trying to find a solution to how to train a saved model on a TPU.\n\n\n**Any tips or recommendations on how to make the most out of the kaggle GPU/TPU sessions are welcome :)**",
    "875087": "Hey! Thanks for mentioning me! :)\n\nI've actually done this idea of splitting the training among multiple sessions, so short answer, yes that's totally possible. Now that I think about it, I forgot to put that in my article!\n\nI was saving the epoch number of the last completed epoch, and created a parameter to pick up training from that epoch (given also some weights to load and pick up from).\n\nThis requires a bit of work, and you might loose incomplete training for a given a epoch, but it allows you to go to a lot more epochs overall than just what would fit in a 9h session.\n\nI'd be very interested in having some feedback on your side if you come up with some interesting ideas :)",
    "875112": "&gt; Is it possible to train an efficientNet B7 or an SEResNeXt101 on +70 epochs on several 9H GPU sessions by saving the weights and reloading them in every sessions so that the training is picked up where the model has left off in the previous session?\n\nYes. I did this in Flower Comp to train a model in Kaggle notebooks for 27 hours. And I did this in Molecule comp to train a model in Kaggle notebooks for 180 hours! \n\nJust train for some epochs in one notebook and then save the weights `model.save_weights('weights.h5')` and remember where your learning schedule is. Then start a new notebook, build your model `model = build_model()`, then load the old weights `model.load_weights('weights.h5')`, then set your learning schedule to where it left off. And then continue training some more epochs with `model.fit()`. Then save, reload, repeat, over and over.",
    "875310": "&gt;&gt; **Chris Deotte wrote:**\n \n&gt;&gt; Just train for some epochs in one notebook and then save the weights `model.save_weights('weights.h5')` and remember where your learning schedule is. \n\n\nYou can print the last learning rate with this , if ever the training output is truncated\n\n```\nfrom tensorflow.keras.backend import eval\nprint(eval(model.optimizer.lr))\n```",
    "875846": "Thank you @cdeotte and @serigne! It's such a relief to know that it's possible to train the same model on different sessions.\n\nI had some issues with locating the weights with `model= build_model()`. So I saved the whole model `model.save('weights.H5')`instead of `model.save_weights('weights.H5')`\n\n\n```\nsaved_model_path='/kaggle/input/ef7-weights/efficientNetB7.h5'\n\nwith strategy.scope():\n    model = tf.keras.models.load_model(saved_model_path)\n\nmodel.compile(loss=tf.keras.losses.binary_crossentropy,\n                optimizer='adam',\n                metrics=['binary_crossentropy'])\n```\n\nThere is something wrong. Is it necessary to save it in tf format `(save_format=\"tf\")`and upload it to GCS? How did you deal with this problem in the Flower Comp?",
    "875852": "Thanks @maxlenormand for your feedback. I will definitely give you some feedback if I come up with any new tricks. Also, I would like to discuss your technique to send a message/report of your training to your slack channel. I am still trying to figure out how to implement it :)",
    "876996": "Sure, feel free to ask, here, or a PM :)",
    "877497": "Guys, who has experience with Colab TPU? I like, how it is working, but don't understand how to feed data properly - simple loading them into Google Cloud doesn't help. Do I have to transfer them to GCS Backets (which is not free), or there are other ways?",
    "877559": "I’m using it extensively. You don’t have to move anything, just get the location of the files from printing it out in a Kaggle note book and then copying it to your colab notebook. There are a few lines you will need to add and a few to comment out but it’s no problem. Unfortunately I’m at work in a lunch break on my phone so I can’t easily cut and paste a lot of help. I won’t be home for 5 hours . I’ll give you a better answer then if you still need it.",
    "877747": "Ok Serge,\n\nI'm assuming that you already have a Kaggle notebook that uses TPU for this competition. If not there are a few public ones that you could use as a starting point such as this one by Ajay Kumar.\n[Melanoma TPU EfficientNet B5_dense_head](https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head)\n\nto get the Google data path you print out the result of,\n```\n# Data access\nGCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\nprint(GCS_PATH)\n```\ngiving:\n`gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1`\n\nNow, this may change over time and you may have to redo this step to find a new link.\nIf you want to use other data sources you do a similar thing. Add the data source to your Kaggle notebook and construct a GCS_PATH as you did for the main data.\nAdded 2 other data sources|\n\n```\nGCS_PATH2 = KaggleDatasets().get_gcs_path('melanoma-256x256')\nGCS_PATH3 = KaggleDatasets().get_gcs_path('512x512-melanoma-tfrecords-70k-images')\n\n```\n and printed out their Google address.\n\n`print(GCS_PATH2)`\ngs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569\n\n`print(GCS_PATH3)`\ngs://kds-8e5c826f476888a71d6db98e2dc6ceb491e14202931bdf249c60dc97\n\nPlease do this for yourself as the paths here may have changed.\n\nNow for the Google Colab part of the story...\nDownload your Kaggle notebook as a Jupyter Notebook and upload it into Colab.\n\nYou'll need to install and load the google file system at the beginning of your code.\n\n```\n!pip install gcsfs\nimport gcsfs\n```\n\nYou won't need this while in Colab so comment it out.\n`#from kaggle_datasets import KaggleDatasets`\n\nSo now instead of getting Kaggle to get the Google address of the data, you can comment that out.\n\n```\n# Data access\n#GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\n```\n\nand just use the info you have obtained from the Kaggle notebook.\n\n`GCS_PATH = 'gs://kds-c89313da1d85616eec461ab327fed61e1335defb486fb7729cf897b1' #original data`\n\nSame again if you are using other data sources such as these. (the paths may be wrong now so you need to get them yourself)\n\n```\nGCS_PATH2 = 'gs://kds-7d41d8cacb09841542d26c8be854b7f8b04a61076082210f89a83569'  #original data at 256 + metadata\n\n#GCS_PATH3 = 'gs://kds-9e42a1abc97f350d6a801f68f97667cf54db33a6ccbea9b3757df69a'  #original data at 512 + metadata\n```\nIf you use the original data your paths to the train and submission CSV files will be.\n\n```\ntrain = pd.read_csv(GCS_PATH + '/train.csv')\nsub = pd.read_csv(GCS_PATH + '/sample_submission.csv')\n\n```\nThe paths to the original data TFRecords will be. (as it is in a Kaggle notebook)\n\n```\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/test*.tfrec')\n\n```\nor if you use one of the sets graciously made by Chris it will be more like. (as it would be in a Kaggle notebook though the path to the CSV files will still be the original GCS_PATH)\n\n```\nTRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH2 + '/test*.tfrec')\n```\n\nThat should get you started.  \nI hope it works for you. I have used the same basic method for using Colab TPUs for both TFRecords and JPEGS  using TPU.\n\nMake sure you actually turn on the TPU accelerator using Edit / Notebook Settings.\n\nGood luck!\n\nP.S. THe Colab TPUs aren't as \"powerful' as those on Kaggle so you will most likely have to dial back your Batch size a bit so that you won't run out of memory.",
    "878326": "IMHO  saving only model weights is not sufficent for complete  continuity of trainng process. It is also necessary to save and restore optimizer weights. During bengali comp i used the next TF construction:\n\n```\nopt=tf.keras.optimizers.Adam(learning_rate=LR_MAX)\ncheckpoint =  tf.train.Checkpoint(latest_epoch=tf.Variable(0), optimizer=opt, model=model)\nchkp_manager =  tf.train.CheckpointManager(checkpoint, CHKP_DIR, max_to_keep=3)\n\nchkp_restore_status=checkpoint.restore(chkp_manager.latest_checkpoint)\n\nif chkp_manager.latest_checkpoint:\n  print(\"Restored from {}\".format(chkp_manager.latest_checkpoint))\n  print(f\"The Latest epoch {checkpoint.latest_epoch.numpy()}\")\nelse:\n  print(\"Initializing from scratch.\")\n\nmodel.compile(optimizer=opt, loss=loss_func)\ninitial_epoch = checkpoint.latest_epoch.numpy()\n\nhistory=model.fit(datagenerator,\n                  epochs=EPOCHS,\n                  steps_per_epoch=spe,\n                  callbacks=[my_callback(initial_epoch=initial_epoch, chkp_manager=chkp_manager, ... )],\n                  verbose=1)\n```",
    "878340": "Yes! It will be more correct to save the whole model, but if your model contains custom objects (for example spicific activation functions or custom layers) load_model dont work correctly.",
    "881231": "Thanks! \nI want to try Colab, because when I've used all memory it gives me more, then more again - up to 30 G. Plus I am not that limited by time running.",
    "881780": "You might be interested in resource-efficient convolutions, splitting up 2D Conv and channel mixing.\nhttps://towardsdatascience.com/a-basic-introduction-to-separable-convolutions-b99ec3102728\nIf you want to use transfer learning, some tweaks to your model will be necessary!\n\nIn fact, it is one of the reasons why EfficientNet is so efficient ;) so in fact, adopting EfficientNet is also related to this topic, but too obvious...",
    "889929": "Thank you @ovdnnest for sharing this great article! Did you try to change all the normal convolutions to depthwise convolutions? Please share you experience :) Doing this reduces the number of parameters, hence the training time but don't you think we will have a Parameters-Loss trade-off? Less parameters = Higher loss?",
    "890429": "The 2D Conv and channel mixing is already separated in EfficientNet. For other models, I have not tried adapting the architecture. It was more a theoretical tip ;)\nAnd there is no direct relationship between the number of parameters and loss, as you probably know. In the traditional bias-variance theory you need to increase your amount until you reach the interpolation regime. However recently double-descent is observed in some deep NN tasks, meaning that very large models generalize better after a larger amount of epochs (in contrast to bias-variance theory). If you are interested in that, check this out: https://arxiv.org/abs/1912.02292\n\nShort answer: it's possible your model generalizes better with fewer parameters or needs to have more parameters to be able to overfit the training data first. This is something you have to determine in practice, there is no direct relation.",
    "896769": "Hi @mutantspore  Thank you for this. I am trying to open JPG files in colab and I am missing something.  For e.g: this file exists  - \n`!gsutil ls gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg` \n\nHowever this fails  - `PIL.Image.open('gs://kds-f213033556301d1e1f16f757181fe87fa38c6bfc138aeb515577cc09/jpeg/train/ISIC_3854244.jpg')`  with FileNotFoundError: [Errno 2] No such file or directory:  - \n\nI think I am not suppose to pass this path directly, Any workarounds ? If you have a moment please let me know.  Thank you",
    "912923": "cdeotte \n\n&gt; Just train for some epochs in one notebook and then save the weights `model.save_weights('weights.h5')` and remember where your learning schedule is. Then start a new notebook, build your model `model = build_model()`, then load the old weights `model.load_weights('weights.h5')`, then set your learning schedule to where it left off. And then continue training some more epochs with `model.fit()`. Then save, reload, repeat, over and over.\n\nCan you kindly tell me how you load weights in second notebook after saving from 1st notebook? As right now I just know that I need to first save the weights/model from 1st notebook and then download the trained weights/model from the notebook and then open second notebook and upload weights/model to this notebook to start further training. Is there some other shortcut and easy way to do this ?\n\nThanks in advance.",
    "914950": "rashmibanthia did you find some solution of the above problem. I am also facing the same issue while reading `jpg `or `png `file but csv are working fine in my case.",
    "914978": "abdurrehman245  Not really, I tried using my own GCS bucket and using `bucket.get_blob(jpgfile)`. It worked but it was too slow.  I am using pytorch, the best I could find is to use [this](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg) external data and unzip it on colab or use .npy files. If you are using TF you should be able to use TF Records."
  },
  "source": "meta"
}