{
  "id": 203545,
  "title": "Save GPU time !! 5x Speedup using ds.cache() and mixed precision (tf and Keras)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/203545",
  "author_name": "",
  "post_date": "2020-12-15T16:53:28.698239500Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Chronology of my implementation:</p>\n<p>Running time of one epoch:</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>Execution time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Keras + Image Data Generator</td>\n<td>825s</td>\n</tr>\n<tr>\n<td>tf.data + keras</td>\n<td>456s</td>\n</tr>\n<tr>\n<td>tf.data + ds.cache() + Mixed Precision Training</td>\n<td>240s</td>\n</tr>\n</tbody>\n</table>\n<p>I was able to train EfficientNetB3 with input size 300x300 on batch size 32 within an hour</p>\n<h2>What was the problem with ds.cache() ?</h2>\n<p>The problem was I was saving cache in /kaggle/working directory which has a size limit of 20 GB it should be saved in /kaggle directory.</p>\n<h4>Tip 1:</h4>\n<p>You can double your batch size in mixed precision setup to further reduce the training time and half the time taken for 1 epoch.</p>\n<h4>Tip 2:</h4>\n<p>You can achieve even 5x if you have bigger ram and instead of storing cache on disk you store cache on RAM.</p>\n<p>All my kernels are public!</p>",
  "messages": [
    {
      "id": "1113746",
      "postDate": "12/15/2020 16:53:28",
      "content": "<p>Chronology of my implementation:</p>\n<p>Running time of one epoch:</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>Execution time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Keras + Image Data Generator</td>\n<td>825s</td>\n</tr>\n<tr>\n<td>tf.data + keras</td>\n<td>456s</td>\n</tr>\n<tr>\n<td>tf.data + ds.cache() + Mixed Precision Training</td>\n<td>240s</td>\n</tr>\n</tbody>\n</table>\n<p>I was able to train EfficientNetB3 with input size 300x300 on batch size 32 within an hour</p>\n<h2>What was the problem with ds.cache() ?</h2>\n<p>The problem was I was saving cache in /kaggle/working directory which has a size limit of 20 GB it should be saved in /kaggle directory.</p>\n<h4>Tip 1:</h4>\n<p>You can double your batch size in mixed precision setup to further reduce the training time and half the time taken for 1 epoch.</p>\n<h4>Tip 2:</h4>\n<p>You can achieve even 5x if you have bigger ram and instead of storing cache on disk you store cache on RAM.</p>\n<p>All my kernels are public!</p>",
      "rawMarkdown": "Chronology of my implementation:\n\nRunning time of one epoch:\n\n| Strategy  | Execution time |\n| --- | --- |\n| Keras + Image Data Generator | 825s  |\n| tf.data + keras | 456s |\n| tf.data + ds.cache() + Mixed Precision Training | 240s |\n\n\nI was able to train EfficientNetB3 with input size 300x300 on batch size 32 within an hour\n\n\n## What was the problem with ds.cache() ?\n\nThe problem was I was saving cache in /kaggle/working directory which has a size limit of 20 GB it should be saved in /kaggle directory.\n\n#### Tip 1:\nYou can double your batch size in mixed precision setup to further reduce the training time and half the time taken for 1 epoch.\n\n#### Tip 2:\nYou can achieve even 5x if you have bigger ram and instead of storing cache on disk you store cache on RAM.\n\nAll my kernels are public!",
      "votes": null
    },
    {
      "id": "1116620",
      "postDate": "12/17/2020 10:31:42",
      "content": "<p>Very nice <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a>, I'm following this strategy too! But I'm training on TPU.</p>",
      "rawMarkdown": "Very nice @harveenchadha, I'm following this strategy too! But I'm training on TPU.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1116620,
      "author_name": "alvarole",
      "author_url": "",
      "post_date": "12/17/2020 10:31:42",
      "content": "<p>Very nice <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a>, I'm following this strategy too! But I'm training on TPU.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1113746": "Chronology of my implementation:\n\nRunning time of one epoch:\n\n| Strategy  | Execution time |\n| --- | --- |\n| Keras + Image Data Generator | 825s  |\n| tf.data + keras | 456s |\n| tf.data + ds.cache() + Mixed Precision Training | 240s |\n\n\nI was able to train EfficientNetB3 with input size 300x300 on batch size 32 within an hour\n\n\n## What was the problem with ds.cache() ?\n\nThe problem was I was saving cache in /kaggle/working directory which has a size limit of 20 GB it should be saved in /kaggle directory.\n\n#### Tip 1:\nYou can double your batch size in mixed precision setup to further reduce the training time and half the time taken for 1 epoch.\n\n#### Tip 2:\nYou can achieve even 5x if you have bigger ram and instead of storing cache on disk you store cache on RAM.\n\nAll my kernels are public!",
    "1116620": "Very nice @harveenchadha, I'm following this strategy too! But I'm training on TPU."
  },
  "source": "meta"
}