{
  "id": 215486,
  "title": "Slow training on kaggle kernel ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215486",
  "author_name": "",
  "post_date": "2021-01-30T06:08:55.493764900Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello guys<br>\nAs you all can see this competition has large dataset so training on kaggle kernel is very slow it is taking 4-5 mins per epoch in my case and 1-2 min when i under sample data so i am curious to know if i have to train offline or i can speed up training on kaggle kernel. </p>",
  "messages": [
    {
      "id": "1177199",
      "postDate": "01/30/2021 06:08:55",
      "content": "<p>Hello guys<br>\nAs you all can see this competition has large dataset so training on kaggle kernel is very slow it is taking 4-5 mins per epoch in my case and 1-2 min when i under sample data so i am curious to know if i have to train offline or i can speed up training on kaggle kernel. </p>",
      "rawMarkdown": "Hello guys\nAs you all can see this competition has large dataset so training on kaggle kernel is very slow it is taking 4-5 mins per epoch in my case and 1-2 min when i under sample data so i am curious to know if i have to train offline or i can speed up training on kaggle kernel.",
      "votes": null
    },
    {
      "id": "1177356",
      "postDate": "01/30/2021 08:19:31",
      "content": "<h3>Hello!</h3>\n<p>To my mind going offline is definitely not an option unless you have a monster PC.<br>\nI'd rather suggest switching to TPU as an accelerator and using TFRecords instead of Jpegs to significantly boost up your training (e.g. with TPUv3 on and TFRecords training EfficientNetB4 takes just 90 seconds per epoch). Consider looking through this great <strong><a href=\"https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease\" target=\"_blank\">community notebook</a></strong> for a quick start.</p>\n<h3>Update</h3>\n<p>There are also some other options to speed up the training process like large-batch and low precision training (e.g. switching from float32 to float16 on Nvidia V100 results in 7 times more TFlops) proposed in <strong><a href=\"https://arxiv.org/abs/1812.01187\" target=\"_blank\">this paper</a></strong>, but I haven't tried them out.</p>",
      "rawMarkdown": "### Hello! \n\nTo my mind going offline is definitely not an option unless you have a monster PC.\nI'd rather suggest switching to TPU as an accelerator and using TFRecords instead of Jpegs to significantly boost up your training (e.g. with TPUv3 on and TFRecords training EfficientNetB4 takes just 90 seconds per epoch). Consider looking through this great **[community notebook](https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease)** for a quick start.\n\n### Update\nThere are also some other options to speed up the training process like large-batch and low precision training (e.g. switching from float32 to float16 on Nvidia V100 results in 7 times more TFlops) proposed in **[this paper](https://arxiv.org/abs/1812.01187)**, but I haven't tried them out.",
      "votes": null
    },
    {
      "id": "1177432",
      "postDate": "01/30/2021 09:27:16",
      "content": "<p>Hello, If you use keras/tensorflow, you can use mixed precision in your code to significantly decrease training times and increase batch size as well for faster epochs. <br>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/mixed_precision/experimental/Policy\" target=\"_blank\">Here for more info</a></p>\n<p>Another option to try out might be to reduce your image size and during your model selection be careful not to add too many dense layers that might increase parameters unnecessarily.</p>\n<p>And the most important yet mundane thing to do is to enable a GPU/TPU accelerator in the settings as per your requirements, it really seems trivial but often forgotten when starting a new notebook.(I really hope you enabled GPU in settings.)</p>\n<p>As for the offline training, Do that after you exhaust your weekly GPU quota.</p>",
      "rawMarkdown": "Hello, If you use keras/tensorflow, you can use mixed precision in your code to significantly decrease training times and increase batch size as well for faster epochs. \n[Here for more info](https://www.tensorflow.org/api_docs/python/tf/keras/mixed_precision/experimental/Policy)\n\nAnother option to try out might be to reduce your image size and during your model selection be careful not to add too many dense layers that might increase parameters unnecessarily.\n\nAnd the most important yet mundane thing to do is to enable a GPU/TPU accelerator in the settings as per your requirements, it really seems trivial but often forgotten when starting a new notebook.(I really hope you enabled GPU in settings.)\n\nAs for the offline training, Do that after you exhaust your weekly GPU quota.",
      "votes": null
    },
    {
      "id": "1177462",
      "postDate": "01/30/2021 09:51:39",
      "content": "<p>Since we have no idea what your code looks like, pretty general answers are the best you will get.  </p>\n<ol>\n<li>I have a top shelf PC with dual GPU - my model takes 37 minutes per epoch.  In general for the same code my PC tends to be a bit faster than Kaggle since I have more GPU and more CPU cores but its no magic cure.  </li>\n<li>Batch size needs to be as big as you can get it.   Have your run tests to get the largest batch?</li>\n<li>Augmentation slows things down - lots of efficient ways to do this and lots of bad ways.</li>\n<li>GPU or TPU for sure until you use up your quota for the week.</li>\n<li>tfrecords faster.</li>\n</ol>\n<blockquote>\n  <p>4-5 mins is not bad IMO.   </p>\n</blockquote>",
      "rawMarkdown": "Since we have no idea what your code looks like, pretty general answers are the best you will get.  \n\n1.  I have a top shelf PC with dual GPU - my model takes 37 minutes per epoch.  In general for the same code my PC tends to be a bit faster than Kaggle since I have more GPU and more CPU cores but its no magic cure.  \n2.  Batch size needs to be as big as you can get it.   Have your run tests to get the largest batch?\n3.  Augmentation slows things down - lots of efficient ways to do this and lots of bad ways.\n4.  GPU or TPU for sure until you use up your quota for the week.\n5.  tfrecords faster.\n\n\n\n\n> 4-5 mins is not bad IMO.",
      "votes": null
    },
    {
      "id": "1177572",
      "postDate": "01/30/2021 11:37:26",
      "content": "<p>It take 3 minutes per epoch and I am using tfrecords and tpu to run a 512 * 512 image. So 4-5 minutes per epoch is a good one. You can decrease the size of image to 256 if you want to try few things and once you have done with the experimentation can again switch to 512 size. Using offline is going to take much more time and you do not have an option to use tpu offline.</p>",
      "rawMarkdown": "It take 3 minutes per epoch and I am using tfrecords and tpu to run a 512 * 512 image. So 4-5 minutes per epoch is a good one. You can decrease the size of image to 256 if you want to try few things and once you have done with the experimentation can again switch to 512 size. Using offline is going to take much more time and you do not have an option to use tpu offline.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1177356,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "01/30/2021 08:19:31",
      "content": "<h3>Hello!</h3>\n<p>To my mind going offline is definitely not an option unless you have a monster PC.<br>\nI'd rather suggest switching to TPU as an accelerator and using TFRecords instead of Jpegs to significantly boost up your training (e.g. with TPUv3 on and TFRecords training EfficientNetB4 takes just 90 seconds per epoch). Consider looking through this great <strong><a href=\"https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease\" target=\"_blank\">community notebook</a></strong> for a quick start.</p>\n<h3>Update</h3>\n<p>There are also some other options to speed up the training process like large-batch and low precision training (e.g. switching from float32 to float16 on Nvidia V100 results in 7 times more TFlops) proposed in <strong><a href=\"https://arxiv.org/abs/1812.01187\" target=\"_blank\">this paper</a></strong>, but I haven't tried them out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1177432,
      "author_name": "mohneesh7",
      "author_url": "",
      "post_date": "01/30/2021 09:27:16",
      "content": "<p>Hello, If you use keras/tensorflow, you can use mixed precision in your code to significantly decrease training times and increase batch size as well for faster epochs. <br>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/mixed_precision/experimental/Policy\" target=\"_blank\">Here for more info</a></p>\n<p>Another option to try out might be to reduce your image size and during your model selection be careful not to add too many dense layers that might increase parameters unnecessarily.</p>\n<p>And the most important yet mundane thing to do is to enable a GPU/TPU accelerator in the settings as per your requirements, it really seems trivial but often forgotten when starting a new notebook.(I really hope you enabled GPU in settings.)</p>\n<p>As for the offline training, Do that after you exhaust your weekly GPU quota.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1177462,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/30/2021 09:51:39",
      "content": "<p>Since we have no idea what your code looks like, pretty general answers are the best you will get.  </p>\n<ol>\n<li>I have a top shelf PC with dual GPU - my model takes 37 minutes per epoch.  In general for the same code my PC tends to be a bit faster than Kaggle since I have more GPU and more CPU cores but its no magic cure.  </li>\n<li>Batch size needs to be as big as you can get it.   Have your run tests to get the largest batch?</li>\n<li>Augmentation slows things down - lots of efficient ways to do this and lots of bad ways.</li>\n<li>GPU or TPU for sure until you use up your quota for the week.</li>\n<li>tfrecords faster.</li>\n</ol>\n<blockquote>\n  <p>4-5 mins is not bad IMO.   </p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1177572,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "01/30/2021 11:37:26",
      "content": "<p>It take 3 minutes per epoch and I am using tfrecords and tpu to run a 512 * 512 image. So 4-5 minutes per epoch is a good one. You can decrease the size of image to 256 if you want to try few things and once you have done with the experimentation can again switch to 512 size. Using offline is going to take much more time and you do not have an option to use tpu offline.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1177199": "Hello guys\nAs you all can see this competition has large dataset so training on kaggle kernel is very slow it is taking 4-5 mins per epoch in my case and 1-2 min when i under sample data so i am curious to know if i have to train offline or i can speed up training on kaggle kernel.",
    "1177356": "### Hello! \n\nTo my mind going offline is definitely not an option unless you have a monster PC.\nI'd rather suggest switching to TPU as an accelerator and using TFRecords instead of Jpegs to significantly boost up your training (e.g. with TPUv3 on and TFRecords training EfficientNetB4 takes just 90 seconds per epoch). Consider looking through this great **[community notebook](https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease)** for a quick start.\n\n### Update\nThere are also some other options to speed up the training process like large-batch and low precision training (e.g. switching from float32 to float16 on Nvidia V100 results in 7 times more TFlops) proposed in **[this paper](https://arxiv.org/abs/1812.01187)**, but I haven't tried them out.",
    "1177432": "Hello, If you use keras/tensorflow, you can use mixed precision in your code to significantly decrease training times and increase batch size as well for faster epochs. \n[Here for more info](https://www.tensorflow.org/api_docs/python/tf/keras/mixed_precision/experimental/Policy)\n\nAnother option to try out might be to reduce your image size and during your model selection be careful not to add too many dense layers that might increase parameters unnecessarily.\n\nAnd the most important yet mundane thing to do is to enable a GPU/TPU accelerator in the settings as per your requirements, it really seems trivial but often forgotten when starting a new notebook.(I really hope you enabled GPU in settings.)\n\nAs for the offline training, Do that after you exhaust your weekly GPU quota.",
    "1177462": "Since we have no idea what your code looks like, pretty general answers are the best you will get.  \n\n1.  I have a top shelf PC with dual GPU - my model takes 37 minutes per epoch.  In general for the same code my PC tends to be a bit faster than Kaggle since I have more GPU and more CPU cores but its no magic cure.  \n2.  Batch size needs to be as big as you can get it.   Have your run tests to get the largest batch?\n3.  Augmentation slows things down - lots of efficient ways to do this and lots of bad ways.\n4.  GPU or TPU for sure until you use up your quota for the week.\n5.  tfrecords faster.\n\n\n\n\n> 4-5 mins is not bad IMO.",
    "1177572": "It take 3 minutes per epoch and I am using tfrecords and tpu to run a 512 * 512 image. So 4-5 minutes per epoch is a good one. You can decrease the size of image to 256 if you want to try few things and once you have done with the experimentation can again switch to 512 size. Using offline is going to take much more time and you do not have an option to use tpu offline."
  },
  "source": "meta"
}