{
  "id": 203390,
  "title": "Tip to reduce GPU training time by up to 3x using Tensorflow caching",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/203390",
  "author_name": "",
  "post_date": "2020-12-15T04:29:29.737507200Z",
  "votes": 57,
  "comment_count": 14,
  "views": 0,
  "content": "<p>When training with TPUs using <code>tf.data.Dataset</code>, you can automatically cache the training images in memory by calling <code>dataset.cache()</code>. However, this method is not very convenient if you try to run the same code on CPU, since you will quickly run out of memory. </p>\n<p>Rather than iteratively resizing and saving the images as pickle/numpy arrays, you can simply specify a local directory where you want to cache your data, e.g. <code>dataset.cache(\"/kaggle/tf_cache\")</code>. This will avoid writing custom pre-processing code, as well as eliminate the CPU bottleneck when reading the images as JPEGs and resizing them. </p>\n<p>In my <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b3-gpu-starter\" target=\"_blank\">GPU starter notebook</a>, the training time dropped from 1895s for the first epoch to 640s for the subsequent epochs. This is not quite as dramatic as <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b7-tpu-training\" target=\"_blank\">TPU's time reduction</a> (which dropped from 1191s to to 265s), but it ensures a smaller EfficientNet model can be trained in one or two hours.</p>",
  "messages": [
    {
      "id": "1112977",
      "postDate": "12/15/2020 04:29:29",
      "content": "<p>When training with TPUs using <code>tf.data.Dataset</code>, you can automatically cache the training images in memory by calling <code>dataset.cache()</code>. However, this method is not very convenient if you try to run the same code on CPU, since you will quickly run out of memory. </p>\n<p>Rather than iteratively resizing and saving the images as pickle/numpy arrays, you can simply specify a local directory where you want to cache your data, e.g. <code>dataset.cache(\"/kaggle/tf_cache\")</code>. This will avoid writing custom pre-processing code, as well as eliminate the CPU bottleneck when reading the images as JPEGs and resizing them. </p>\n<p>In my <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b3-gpu-starter\" target=\"_blank\">GPU starter notebook</a>, the training time dropped from 1895s for the first epoch to 640s for the subsequent epochs. This is not quite as dramatic as <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b7-tpu-training\" target=\"_blank\">TPU's time reduction</a> (which dropped from 1191s to to 265s), but it ensures a smaller EfficientNet model can be trained in one or two hours.</p>",
      "rawMarkdown": "When training with TPUs using `tf.data.Dataset`, you can automatically cache the training images in memory by calling `dataset.cache()`. However, this method is not very convenient if you try to run the same code on CPU, since you will quickly run out of memory. \n\nRather than iteratively resizing and saving the images as pickle/numpy arrays, you can simply specify a local directory where you want to cache your data, e.g. `dataset.cache(\"/kaggle/tf_cache\")`. This will avoid writing custom pre-processing code, as well as eliminate the CPU bottleneck when reading the images as JPEGs and resizing them. \n\nIn my [GPU starter notebook](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b3-gpu-starter), the training time dropped from 1895s for the first epoch to 640s for the subsequent epochs. This is not quite as dramatic as [TPU's time reduction](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b7-tpu-training) (which dropped from 1191s to to 265s), but it ensures a smaller EfficientNet model can be trained in one or two hours.",
      "votes": null
    },
    {
      "id": "1113016",
      "postDate": "12/15/2020 05:34:14",
      "content": "<p>Great trick. Thanks!</p>",
      "rawMarkdown": "Great trick. Thanks!",
      "votes": null
    },
    {
      "id": "1113556",
      "postDate": "12/15/2020 14:51:18",
      "content": "<p>I have tried in previous competitions and these are my findings:</p>\n<ol>\n<li>If you just use ds.cache() it would store results in RAM and in this case kernel will go out of memory.</li>\n<li>If you use ds.cache('kaggle.tfcache') the cache would be stored in disk and eventually the disk would go out of space.</li>\n</ol>\n<p>Did you encounter this issue?</p>",
      "rawMarkdown": "I have tried in previous competitions and these are my findings:\n\n1. If you just use ds.cache() it would store results in RAM and in this case kernel will go out of memory.\n2. If you use ds.cache('kaggle.tfcache') the cache would be stored in disk and eventually the disk would go out of space.\n \n\nDid you encounter this issue?",
      "votes": null
    },
    {
      "id": "1113568",
      "postDate": "12/15/2020 14:57:46",
      "content": "<p>You have to store inside /kaggle rather than /kaggle/working, since the latter will be saved as output and is limited to 20gb (and the former has much more space)</p>",
      "rawMarkdown": "You have to store inside /kaggle rather than /kaggle/working, since the latter will be saved as output and is limited to 20gb (and the former has much more space)",
      "votes": null
    },
    {
      "id": "1113584",
      "postDate": "12/15/2020 15:03:31",
      "content": "<p>Yes that is correct I was using the /kaggle/working directory. Thank you for this tip! It will save a lot of my gpu time!</p>",
      "rawMarkdown": "Yes that is correct I was using the /kaggle/working directory. Thank you for this tip! It will save a lot of my gpu time!",
      "votes": null
    },
    {
      "id": "1113638",
      "postDate": "12/15/2020 15:51:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a>, thanks for sharing your valuable insight. The \"GPU starter notebook\" is a nice work of arts. Thanks again for sharing.</p>",
      "rawMarkdown": "Hi @xhlulu, thanks for sharing your valuable insight. The \"GPU starter notebook\" is a nice work of arts. Thanks again for sharing.",
      "votes": null
    },
    {
      "id": "1114047",
      "postDate": "12/15/2020 23:53:40",
      "content": "<p>Hey! Thanks for sharing, would this work for data augmented images as well? If i call the cache method with a data augmentation function maped to the dataset, let's say three times, will it generate three sets of augmented images?</p>",
      "rawMarkdown": "Hey! Thanks for sharing, would this work for data augmented images as well? If i call the cache method with a data augmentation function maped to the dataset, let's say three times, will it generate three sets of augmented images?",
      "votes": null
    },
    {
      "id": "1115107",
      "postDate": "12/16/2020 01:49:18",
      "content": "<p>I'm assuming order matters for tensorflow, in which case I think it will depend on when exactly you cache the file. What I do is that I cache it before I call augment, so the augmentation will still be random.</p>",
      "rawMarkdown": "I'm assuming order matters for tensorflow, in which case I think it will depend on when exactly you cache the file. What I do is that I cache it before I call augment, so the augmentation will still be random.",
      "votes": null
    },
    {
      "id": "1115156",
      "postDate": "12/16/2020 03:24:04",
      "content": "<p>This is a great trick to use. Is there any such thing to use in PyTorch as well?</p>",
      "rawMarkdown": "This is a great trick to use. Is there any such thing to use in PyTorch as well?",
      "votes": null
    },
    {
      "id": "1115244",
      "postDate": "12/16/2020 05:56:14",
      "content": "<p>Awesome! For pytorch : <code>loader.dataset.set_use_cache(True)</code></p>",
      "rawMarkdown": "Awesome! For pytorch : `loader.dataset.set_use_cache(True)`",
      "votes": null
    },
    {
      "id": "1115271",
      "postDate": "12/16/2020 06:42:39",
      "content": "<p>Just to confirm:</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>Time taken</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>without cache</td>\n<td>2728s</td>\n</tr>\n<tr>\n<td>with cache</td>\n<td>646s</td>\n</tr>\n</tbody>\n</table>\n<p>I will update results with the mixed precision as well.</p>",
      "rawMarkdown": "Just to confirm:\n\n| Strategy  | Time taken  |\n| --- | --- |\n| without cache | 2728s  |\n| with cache | 646s  |\n\nI will update results with the mixed precision as well.",
      "votes": null
    },
    {
      "id": "1115585",
      "postDate": "12/16/2020 11:51:07",
      "content": "<p>Great stuff!</p>",
      "rawMarkdown": "Great stuff!",
      "votes": null
    },
    {
      "id": "1116869",
      "postDate": "12/17/2020 14:24:56",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> for sharing. It saves a lot of GPU time. But, probably, there is a limit to using kaggle disk. On 'HubMAP…' competition, I got the error below</p>\n<p>after using </p>\n<pre><code>os.makedirs('/kaggle/tf_cache')\n....\ndataset.cache('/kaggle/tf_cache')\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F877700%2Ffd51cbde9717caff294d0e20c7a36e50%2Ferror_tmp.png?generation=1608215745737036&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thank you @xhlulu for sharing. It saves a lot of GPU time. But, probably, there is a limit to using kaggle disk. On 'HubMAP...' competition, I got the error below\n\n after using \n     \n    os.makedirs('/kaggle/tf_cache')\n    ....\n    dataset.cache('/kaggle/tf_cache')\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F877700%2Ffd51cbde9717caff294d0e20c7a36e50%2Ferror_tmp.png?generation=1608215745737036&alt=media)",
      "votes": null
    },
    {
      "id": "1117683",
      "postDate": "12/18/2020 10:38:14",
      "content": "<p>Thanks for the info.</p>",
      "rawMarkdown": "Thanks for the info.",
      "votes": null
    },
    {
      "id": "1140272",
      "postDate": "01/05/2021 22:35:48",
      "content": "<p><a href=\"https://www.kaggle.com/alincijov\" target=\"_blank\">@alincijov</a> please how can i use this line of code in pytorch because when i try it it gives error<br>\n loader.dataset.set_use_cache(True)</p>",
      "rawMarkdown": "alincijov please how can i use this line of code in pytorch because when i try it it gives error\n loader.dataset.set_use_cache(True)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1113016,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/15/2020 05:34:14",
      "content": "<p>Great trick. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1113556,
      "author_name": "harveenchadha",
      "author_url": "",
      "post_date": "12/15/2020 14:51:18",
      "content": "<p>I have tried in previous competitions and these are my findings:</p>\n<ol>\n<li>If you just use ds.cache() it would store results in RAM and in this case kernel will go out of memory.</li>\n<li>If you use ds.cache('kaggle.tfcache') the cache would be stored in disk and eventually the disk would go out of space.</li>\n</ol>\n<p>Did you encounter this issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1113568,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "12/15/2020 14:57:46",
          "content": "<p>You have to store inside /kaggle rather than /kaggle/working, since the latter will be saved as output and is limited to 20gb (and the former has much more space)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1113584,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "12/15/2020 15:03:31",
          "content": "<p>Yes that is correct I was using the /kaggle/working directory. Thank you for this tip! It will save a lot of my gpu time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1113638,
      "author_name": "puzuwe",
      "author_url": "",
      "post_date": "12/15/2020 15:51:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a>, thanks for sharing your valuable insight. The \"GPU starter notebook\" is a nice work of arts. Thanks again for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1114047,
      "author_name": "capiru",
      "author_url": "",
      "post_date": "12/15/2020 23:53:40",
      "content": "<p>Hey! Thanks for sharing, would this work for data augmented images as well? If i call the cache method with a data augmentation function maped to the dataset, let's say three times, will it generate three sets of augmented images?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1115107,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "12/16/2020 01:49:18",
          "content": "<p>I'm assuming order matters for tensorflow, in which case I think it will depend on when exactly you cache the file. What I do is that I cache it before I call augment, so the augmentation will still be random.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115156,
      "author_name": "sovitrath",
      "author_url": "",
      "post_date": "12/16/2020 03:24:04",
      "content": "<p>This is a great trick to use. Is there any such thing to use in PyTorch as well?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115244,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "12/16/2020 05:56:14",
      "content": "<p>Awesome! For pytorch : <code>loader.dataset.set_use_cache(True)</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115271,
      "author_name": "harveenchadha",
      "author_url": "",
      "post_date": "12/16/2020 06:42:39",
      "content": "<p>Just to confirm:</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>Time taken</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>without cache</td>\n<td>2728s</td>\n</tr>\n<tr>\n<td>with cache</td>\n<td>646s</td>\n</tr>\n</tbody>\n</table>\n<p>I will update results with the mixed precision as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1115585,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "12/16/2020 11:51:07",
      "content": "<p>Great stuff!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1116869,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "12/17/2020 14:24:56",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> for sharing. It saves a lot of GPU time. But, probably, there is a limit to using kaggle disk. On 'HubMAP…' competition, I got the error below</p>\n<p>after using </p>\n<pre><code>os.makedirs('/kaggle/tf_cache')\n....\ndataset.cache('/kaggle/tf_cache')\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F877700%2Ffd51cbde9717caff294d0e20c7a36e50%2Ferror_tmp.png?generation=1608215745737036&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1117683,
      "author_name": "ashikm96",
      "author_url": "",
      "post_date": "12/18/2020 10:38:14",
      "content": "<p>Thanks for the info.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1140272,
      "author_name": "mohamed3abdelrazik",
      "author_url": "",
      "post_date": "01/05/2021 22:35:48",
      "content": "<p><a href=\"https://www.kaggle.com/alincijov\" target=\"_blank\">@alincijov</a> please how can i use this line of code in pytorch because when i try it it gives error<br>\n loader.dataset.set_use_cache(True)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1112977": "When training with TPUs using `tf.data.Dataset`, you can automatically cache the training images in memory by calling `dataset.cache()`. However, this method is not very convenient if you try to run the same code on CPU, since you will quickly run out of memory. \n\nRather than iteratively resizing and saving the images as pickle/numpy arrays, you can simply specify a local directory where you want to cache your data, e.g. `dataset.cache(\"/kaggle/tf_cache\")`. This will avoid writing custom pre-processing code, as well as eliminate the CPU bottleneck when reading the images as JPEGs and resizing them. \n\nIn my [GPU starter notebook](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b3-gpu-starter), the training time dropped from 1895s for the first epoch to 640s for the subsequent epochs. This is not quite as dramatic as [TPU's time reduction](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-b7-tpu-training) (which dropped from 1191s to to 265s), but it ensures a smaller EfficientNet model can be trained in one or two hours.",
    "1113016": "Great trick. Thanks!",
    "1113556": "I have tried in previous competitions and these are my findings:\n\n1. If you just use ds.cache() it would store results in RAM and in this case kernel will go out of memory.\n2. If you use ds.cache('kaggle.tfcache') the cache would be stored in disk and eventually the disk would go out of space.\n \n\nDid you encounter this issue?",
    "1113568": "You have to store inside /kaggle rather than /kaggle/working, since the latter will be saved as output and is limited to 20gb (and the former has much more space)",
    "1113584": "Yes that is correct I was using the /kaggle/working directory. Thank you for this tip! It will save a lot of my gpu time!",
    "1113638": "Hi @xhlulu, thanks for sharing your valuable insight. The \"GPU starter notebook\" is a nice work of arts. Thanks again for sharing.",
    "1114047": "Hey! Thanks for sharing, would this work for data augmented images as well? If i call the cache method with a data augmentation function maped to the dataset, let's say three times, will it generate three sets of augmented images?",
    "1115107": "I'm assuming order matters for tensorflow, in which case I think it will depend on when exactly you cache the file. What I do is that I cache it before I call augment, so the augmentation will still be random.",
    "1115156": "This is a great trick to use. Is there any such thing to use in PyTorch as well?",
    "1115244": "Awesome! For pytorch : `loader.dataset.set_use_cache(True)`",
    "1115271": "Just to confirm:\n\n| Strategy  | Time taken  |\n| --- | --- |\n| without cache | 2728s  |\n| with cache | 646s  |\n\nI will update results with the mixed precision as well.",
    "1115585": "Great stuff!",
    "1116869": "Thank you @xhlulu for sharing. It saves a lot of GPU time. But, probably, there is a limit to using kaggle disk. On 'HubMAP...' competition, I got the error below\n\n after using \n     \n    os.makedirs('/kaggle/tf_cache')\n    ....\n    dataset.cache('/kaggle/tf_cache')\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F877700%2Ffd51cbde9717caff294d0e20c7a36e50%2Ferror_tmp.png?generation=1608215745737036&alt=media)",
    "1117683": "Thanks for the info.",
    "1140272": "alincijov please how can i use this line of code in pytorch because when i try it it gives error\n loader.dataset.set_use_cache(True)"
  },
  "source": "meta"
}