{
  "id": 218064,
  "title": "How do you guys do 5-fold CV on kaggle notebooks?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/218064",
  "author_name": "",
  "post_date": "2021-02-09T06:23:36.013546600Z",
  "votes": -1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am trying to do 5fold CV on kaggle notebooks, it stops midway after 9hrs, I only get 2 folds done in that time.<br>\nI use efficientnetb4, img size 512, batch size 8(or 2 sometimes).</p>\n<p>It just takes too much time even with mixed-precision training.</p>",
  "messages": [
    {
      "id": "1192497",
      "postDate": "02/09/2021 06:23:36",
      "content": "<p>I am trying to do 5fold CV on kaggle notebooks, it stops midway after 9hrs, I only get 2 folds done in that time.<br>\nI use efficientnetb4, img size 512, batch size 8(or 2 sometimes).</p>\n<p>It just takes too much time even with mixed-precision training.</p>",
      "rawMarkdown": "I am trying to do 5fold CV on kaggle notebooks, it stops midway after 9hrs, I only get 2 folds done in that time.\nI use efficientnetb4, img size 512, batch size 8(or 2 sometimes).\n\nIt just takes too much time even with mixed-precision training.",
      "votes": null
    },
    {
      "id": "1192510",
      "postDate": "02/09/2021 06:36:55",
      "content": "<p>Hello!</p>\n<p>There may be several reasons for such a long runtime. Consider the following, top to bottom:</p>\n<ol>\n<li>Switch from GPU to TPU</li>\n<li>Use TFRecords instead of Jpegs</li>\n<li>Use custom training loop</li>\n</ol>\n<p>(1-2) can be found in this <strong><a href=\"https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease\" target=\"_blank\">community notebook</a></strong>, (3) can be found in this <strong><a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">notebook by DimitreOliveira</a></strong>.<br>\nWith all the above done training EfficientNetB4 (<code>image_size=(512, 512)</code>, <code>batch_size=128</code>, i.e. 16 samples per core) should take 2 minutes per epoch, up to 4 hours total on 5 KFold.</p>",
      "rawMarkdown": "Hello!\n\nThere may be several reasons for such a long runtime. Consider the following, top to bottom:\n1. Switch from GPU to TPU\n2. Use TFRecords instead of Jpegs\n3. Use custom training loop\n\n(1-2) can be found in this **[community notebook](https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease)**, (3) can be found in this **[notebook by DimitreOliveira](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods)**.\nWith all the above done training EfficientNetB4 (`image_size=(512, 512)`, `batch_size=128`, i.e. 16 samples per core) should take 2 minutes per epoch, up to 4 hours total on 5 KFold.",
      "votes": null
    },
    {
      "id": "1192514",
      "postDate": "02/09/2021 06:42:12",
      "content": "<p>Woah! that's cool. I don't really know much about using TPU's, but I will try to learn it as this is quite an improvement over GPU.<br>\nDo I have to change my whole code to be able to use TPU or just the Training loop part?</p>",
      "rawMarkdown": "Woah! that's cool. I don't really know much about using TPU's, but I will try to learn it as this is quite an improvement over GPU.\nDo I have to change my whole code to be able to use TPU or just the Training loop part?",
      "votes": null
    },
    {
      "id": "1192524",
      "postDate": "02/09/2021 06:54:26",
      "content": "<p>You can turn on TPU anytime without changing anything in your code, but this won't give any notable difference, because the code should be optimized first. Alternatively, you can optimize your code so that it runs faster even without TPU.</p>\n<p>Until you decide to go for a custom training loop, the most important part to optimize is data workflow. In the both notebooks mentioned above you will see <code>tf.data.TFRecordDataset</code> class is utilized instead of more straightforward <code>ImageDataGenerator</code> class (which you probably used a lot). </p>\n<p>And, of course, don't forget to parallellize your model wrapping it by <code>with strategy.scope(): model = get_model(params)</code> statement when running on TPU.</p>",
      "rawMarkdown": "You can turn on TPU anytime without changing anything in your code, but this won't give any notable difference, because the code should be optimized first. Alternatively, you can optimize your code so that it runs faster even without TPU.\n\nUntil you decide to go for a custom training loop, the most important part to optimize is data workflow. In the both notebooks mentioned above you will see `tf.data.TFRecordDataset` class is utilized instead of more straightforward `ImageDataGenerator` class (which you probably used a lot). \n\nAnd, of course, don't forget to parallellize your model wrapping it by `with strategy.scope(): model = get_model(params)` statement when running on TPU.",
      "votes": null
    },
    {
      "id": "1192863",
      "postDate": "02/09/2021 10:33:22",
      "content": "<p>Some tips to reduce time:</p>\n<ul>\n<li>Do augmentation in another notebook. Make a dataset of those augmented images and feed those into your training notebook. Demo <a href=\"https://www.kaggle.com/adityakane/cassava-stratified-k-folds-tfrecords-maker\" target=\"_blank\">here</a>. </li>\n<li>Use TPUs if you can</li>\n<li>Use <code>nadam</code> optimizer, and reduce epochs to 5 or 7</li>\n<li>Batch size: 16 or 32 for 299x299 images</li>\n<li>Use tf.data module and <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> if ypu are using TensorFlow.</li>\n</ul>\n<p>I have implemented these in my notebook <a href=\"https://www.kaggle.com/adityakane/cassava-gpu-preprocessed-trainer\" target=\"_blank\">here</a>. You can take inspiration from it, and please feel free to give suggestions.</p>\n<p>Hope this helps.</p>",
      "rawMarkdown": "Some tips to reduce time:\n\n *  Do augmentation in another notebook. Make a dataset of those augmented images and feed those into your training notebook. Demo [here](https://www.kaggle.com/adityakane/cassava-stratified-k-folds-tfrecords-maker). \n *  Use TPUs if you can\n *  Use `nadam` optimizer, and reduce epochs to 5 or 7\n * Batch size: 16 or 32 for 299x299 images\n * Use tf.data module and @tf.function if ypu are using TensorFlow.\n\nI have implemented these in my notebook [here](https://www.kaggle.com/adityakane/cassava-gpu-preprocessed-trainer). You can take inspiration from it, and please feel free to give suggestions.\n\nHope this helps.",
      "votes": null
    },
    {
      "id": "1193168",
      "postDate": "02/09/2021 13:49:32",
      "content": "<p>Train 1 fold at a time<br>\nYou don't need to train all the folds in one go<br>\nYou can also use gradient accumulation to increase your batch size</p>",
      "rawMarkdown": "Train 1 fold at a time\nYou don't need to train all the folds in one go\nYou can also use gradient accumulation to increase your batch size",
      "votes": null
    },
    {
      "id": "1193579",
      "postDate": "02/09/2021 18:17:20",
      "content": "<p>Thank you very much I will try it out.</p>",
      "rawMarkdown": "Thank you very much I will try it out.",
      "votes": null
    },
    {
      "id": "1193783",
      "postDate": "02/09/2021 21:18:09",
      "content": "<p>Hi,</p>\n<p>Your input preparation might be inefficient. Because, I did experiments with like exactly your parameters except I had gradient accumulation as well. If you can share your preparation script as well, we can take a look. </p>\n<p>Regards</p>",
      "rawMarkdown": "Hi,\n\nYour input preparation might be inefficient. Because, I did experiments with like exactly your parameters except I had gradient accumulation as well. If you can share your preparation script as well, we can take a look. \n\nRegards",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1192510,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "02/09/2021 06:36:55",
      "content": "<p>Hello!</p>\n<p>There may be several reasons for such a long runtime. Consider the following, top to bottom:</p>\n<ol>\n<li>Switch from GPU to TPU</li>\n<li>Use TFRecords instead of Jpegs</li>\n<li>Use custom training loop</li>\n</ol>\n<p>(1-2) can be found in this <strong><a href=\"https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease\" target=\"_blank\">community notebook</a></strong>, (3) can be found in this <strong><a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods\" target=\"_blank\">notebook by DimitreOliveira</a></strong>.<br>\nWith all the above done training EfficientNetB4 (<code>image_size=(512, 512)</code>, <code>batch_size=128</code>, i.e. 16 samples per core) should take 2 minutes per epoch, up to 4 hours total on 5 KFold.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1192514,
          "author_name": "mohneesh7",
          "author_url": "",
          "post_date": "02/09/2021 06:42:12",
          "content": "<p>Woah! that's cool. I don't really know much about using TPU's, but I will try to learn it as this is quite an improvement over GPU.<br>\nDo I have to change my whole code to be able to use TPU or just the Training loop part?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1192524,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "02/09/2021 06:54:26",
          "content": "<p>You can turn on TPU anytime without changing anything in your code, but this won't give any notable difference, because the code should be optimized first. Alternatively, you can optimize your code so that it runs faster even without TPU.</p>\n<p>Until you decide to go for a custom training loop, the most important part to optimize is data workflow. In the both notebooks mentioned above you will see <code>tf.data.TFRecordDataset</code> class is utilized instead of more straightforward <code>ImageDataGenerator</code> class (which you probably used a lot). </p>\n<p>And, of course, don't forget to parallellize your model wrapping it by <code>with strategy.scope(): model = get_model(params)</code> statement when running on TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1192863,
      "author_name": "adityakane",
      "author_url": "",
      "post_date": "02/09/2021 10:33:22",
      "content": "<p>Some tips to reduce time:</p>\n<ul>\n<li>Do augmentation in another notebook. Make a dataset of those augmented images and feed those into your training notebook. Demo <a href=\"https://www.kaggle.com/adityakane/cassava-stratified-k-folds-tfrecords-maker\" target=\"_blank\">here</a>. </li>\n<li>Use TPUs if you can</li>\n<li>Use <code>nadam</code> optimizer, and reduce epochs to 5 or 7</li>\n<li>Batch size: 16 or 32 for 299x299 images</li>\n<li>Use tf.data module and <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> if ypu are using TensorFlow.</li>\n</ul>\n<p>I have implemented these in my notebook <a href=\"https://www.kaggle.com/adityakane/cassava-gpu-preprocessed-trainer\" target=\"_blank\">here</a>. You can take inspiration from it, and please feel free to give suggestions.</p>\n<p>Hope this helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1193168,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "02/09/2021 13:49:32",
      "content": "<p>Train 1 fold at a time<br>\nYou don't need to train all the folds in one go<br>\nYou can also use gradient accumulation to increase your batch size</p>",
      "votes": null,
      "replies": [
        {
          "id": 1193579,
          "author_name": "mohneesh7",
          "author_url": "",
          "post_date": "02/09/2021 18:17:20",
          "content": "<p>Thank you very much I will try it out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1193783,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/09/2021 21:18:09",
      "content": "<p>Hi,</p>\n<p>Your input preparation might be inefficient. Because, I did experiments with like exactly your parameters except I had gradient accumulation as well. If you can share your preparation script as well, we can take a look. </p>\n<p>Regards</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1192497": "I am trying to do 5fold CV on kaggle notebooks, it stops midway after 9hrs, I only get 2 folds done in that time.\nI use efficientnetb4, img size 512, batch size 8(or 2 sometimes).\n\nIt just takes too much time even with mixed-precision training.",
    "1192510": "Hello!\n\nThere may be several reasons for such a long runtime. Consider the following, top to bottom:\n1. Switch from GPU to TPU\n2. Use TFRecords instead of Jpegs\n3. Use custom training loop\n\n(1-2) can be found in this **[community notebook](https://www.kaggle.com/jessemostipak/getting-started-tpus-cassava-leaf-disease)**, (3) can be found in this **[notebook by DimitreOliveira](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods)**.\nWith all the above done training EfficientNetB4 (`image_size=(512, 512)`, `batch_size=128`, i.e. 16 samples per core) should take 2 minutes per epoch, up to 4 hours total on 5 KFold.",
    "1192514": "Woah! that's cool. I don't really know much about using TPU's, but I will try to learn it as this is quite an improvement over GPU.\nDo I have to change my whole code to be able to use TPU or just the Training loop part?",
    "1192524": "You can turn on TPU anytime without changing anything in your code, but this won't give any notable difference, because the code should be optimized first. Alternatively, you can optimize your code so that it runs faster even without TPU.\n\nUntil you decide to go for a custom training loop, the most important part to optimize is data workflow. In the both notebooks mentioned above you will see `tf.data.TFRecordDataset` class is utilized instead of more straightforward `ImageDataGenerator` class (which you probably used a lot). \n\nAnd, of course, don't forget to parallellize your model wrapping it by `with strategy.scope(): model = get_model(params)` statement when running on TPU.",
    "1192863": "Some tips to reduce time:\n\n *  Do augmentation in another notebook. Make a dataset of those augmented images and feed those into your training notebook. Demo [here](https://www.kaggle.com/adityakane/cassava-stratified-k-folds-tfrecords-maker). \n *  Use TPUs if you can\n *  Use `nadam` optimizer, and reduce epochs to 5 or 7\n * Batch size: 16 or 32 for 299x299 images\n * Use tf.data module and @tf.function if ypu are using TensorFlow.\n\nI have implemented these in my notebook [here](https://www.kaggle.com/adityakane/cassava-gpu-preprocessed-trainer). You can take inspiration from it, and please feel free to give suggestions.\n\nHope this helps.",
    "1193168": "Train 1 fold at a time\nYou don't need to train all the folds in one go\nYou can also use gradient accumulation to increase your batch size",
    "1193579": "Thank you very much I will try it out.",
    "1193783": "Hi,\n\nYour input preparation might be inefficient. Because, I did experiments with like exactly your parameters except I had gradient accumulation as well. If you can share your preparation script as well, we can take a look. \n\nRegards"
  },
  "source": "meta"
}