{
  "id": 172511,
  "title": "CPU vs TPU usage",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172511",
  "author_name": "",
  "post_date": "2020-08-05T10:03:40.134149200Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I use the popular notebook made by <a href=\"/cdeotte\">@cdeotte</a> with image size 256*256 and TPU.</p>\n\n<p>I am training a lot of models and came across something I do not fully understand.</p>\n\n<p>Some element of the hardware(CPU/GPU/TPU) is always at its maximum capacity, right? Because if that's not the case the training could go faster by using the hardware's maximum capacity. </p>\n\n<p>The feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.</p>\n\n<p>However, if I increase the size of my network (B4--&gt; B6) the training time increases. This would strain the TPU more, but as this is not the bottleneck it should not matter for my training time.</p>\n\n<p>This strikes me as odd as the same images are preloaded. Can anybody explain this?</p>",
  "messages": [
    {
      "id": "959064",
      "postDate": "08/05/2020 10:03:40",
      "content": "<p>I use the popular notebook made by <a href=\"/cdeotte\">@cdeotte</a> with image size 256*256 and TPU.</p>\n\n<p>I am training a lot of models and came across something I do not fully understand.</p>\n\n<p>Some element of the hardware(CPU/GPU/TPU) is always at its maximum capacity, right? Because if that's not the case the training could go faster by using the hardware's maximum capacity. </p>\n\n<p>The feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.</p>\n\n<p>However, if I increase the size of my network (B4--&gt; B6) the training time increases. This would strain the TPU more, but as this is not the bottleneck it should not matter for my training time.</p>\n\n<p>This strikes me as odd as the same images are preloaded. Can anybody explain this?</p>",
      "rawMarkdown": "I use the popular notebook made by @cdeotte with image size 256*256 and TPU.\n\nI am training a lot of models and came across something I do not fully understand.\n\nSome element of the hardware(CPU/GPU/TPU) is always at its maximum capacity, right? Because if that's not the case the training could go faster by using the hardware's maximum capacity. \n\nThe feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.\n\nHowever, if I increase the size of my network (B4--&gt; B6) the training time increases. This would strain the TPU more, but as this is not the bottleneck it should not matter for my training time.\n\nThis strikes me as odd as the same images are preloaded. Can anybody explain this?",
      "votes": null
    },
    {
      "id": "959146",
      "postDate": "08/05/2020 11:38:21",
      "content": "<p>When a Kaggle Kernel runs, the job of data processing, creating datasets, Data Loaders, and iterating through data Loaders are all done by CPU. The task of traversing through your model, updating parameters is done by GPU/TPU. While GPU/TPU is training the model on a batch, CPU is busy preparing the next batch to train. So if CPU is the bottleneck, then GPU/TPU is just waiting for the CPU to prepare the next batch. </p>\n\n<p>So CPU must be never the bottleneck so that you make the best use of the time of expensive resources, i.e. GPU/TPU. </p>",
      "rawMarkdown": "When a Kaggle Kernel runs, the job of data processing, creating datasets, Data Loaders, and iterating through data Loaders are all done by CPU. The task of traversing through your model, updating parameters is done by GPU/TPU. While GPU/TPU is training the model on a batch, CPU is busy preparing the next batch to train. So if CPU is the bottleneck, then GPU/TPU is just waiting for the CPU to prepare the next batch. \n\nSo CPU must be never the bottleneck so that you make the best use of the time of expensive resources, i.e. GPU/TPU.",
      "votes": null
    },
    {
      "id": "959258",
      "postDate": "08/05/2020 12:57:44",
      "content": "<p>Allright, thanks this makes sense.</p>",
      "rawMarkdown": "Allright, thanks this makes sense.",
      "votes": null
    },
    {
      "id": "959353",
      "postDate": "08/05/2020 14:21:56",
      "content": "<blockquote>\n  <p>The feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.</p>\n</blockquote>\n\n<p>In my popular notebook \"Triple Stratified KFold with TFRecords\" <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>, the bottleneck is the GPU/TPU. If you disable all augmentation, the training time will not go faster because the CPU is not the bottleneck. The CPU is performing augmentation and providing images faster than the GPU/TPU can train.</p>",
      "rawMarkdown": "&gt; The feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.\n\nIn my popular notebook \"Triple Stratified KFold with TFRecords\" [here][1], the bottleneck is the GPU/TPU. If you disable all augmentation, the training time will not go faster because the CPU is not the bottleneck. The CPU is performing augmentation and providing images faster than the GPU/TPU can train.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": null
    },
    {
      "id": "959781",
      "postDate": "08/05/2020 22:10:44",
      "content": "<p>This explains what I am witnessing, thanks. </p>",
      "rawMarkdown": "This explains what I am witnessing, thanks.",
      "votes": null
    },
    {
      "id": "959979",
      "postDate": "08/06/2020 04:13:57",
      "content": "<p>In case you ever find that CPU is becoming a bottleneck (say due to lots of preprocessing), you may explore DALI library by Nvidia, which aims to perform data loading, augmentation, preprocessing all in GPU. I haven't tried it though. <a href=\"https://github.com/NVIDIA/DALI\">https://github.com/NVIDIA/DALI</a></p>",
      "rawMarkdown": "In case you ever find that CPU is becoming a bottleneck (say due to lots of preprocessing), you may explore DALI library by Nvidia, which aims to perform data loading, augmentation, preprocessing all in GPU. I haven't tried it though. https://github.com/NVIDIA/DALI",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 959146,
      "author_name": "krisho007",
      "author_url": "",
      "post_date": "08/05/2020 11:38:21",
      "content": "<p>When a Kaggle Kernel runs, the job of data processing, creating datasets, Data Loaders, and iterating through data Loaders are all done by CPU. The task of traversing through your model, updating parameters is done by GPU/TPU. While GPU/TPU is training the model on a batch, CPU is busy preparing the next batch to train. So if CPU is the bottleneck, then GPU/TPU is just waiting for the CPU to prepare the next batch. </p>\n\n<p>So CPU must be never the bottleneck so that you make the best use of the time of expensive resources, i.e. GPU/TPU. </p>",
      "votes": null,
      "replies": [
        {
          "id": 959258,
          "author_name": "lukereijnen",
          "author_url": "",
          "post_date": "08/05/2020 12:57:44",
          "content": "<p>Allright, thanks this makes sense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959979,
          "author_name": "krisho007",
          "author_url": "",
          "post_date": "08/06/2020 04:13:57",
          "content": "<p>In case you ever find that CPU is becoming a bottleneck (say due to lots of preprocessing), you may explore DALI library by Nvidia, which aims to perform data loading, augmentation, preprocessing all in GPU. I haven't tried it though. <a href=\"https://github.com/NVIDIA/DALI\">https://github.com/NVIDIA/DALI</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 959353,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/05/2020 14:21:56",
      "content": "<blockquote>\n  <p>The feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.</p>\n</blockquote>\n\n<p>In my popular notebook \"Triple Stratified KFold with TFRecords\" <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a>, the bottleneck is the GPU/TPU. If you disable all augmentation, the training time will not go faster because the CPU is not the bottleneck. The CPU is performing augmentation and providing images faster than the GPU/TPU can train.</p>",
      "votes": null,
      "replies": [
        {
          "id": 959781,
          "author_name": "lukereijnen",
          "author_url": "",
          "post_date": "08/05/2020 22:10:44",
          "content": "<p>This explains what I am witnessing, thanks. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "959064": "I use the popular notebook made by @cdeotte with image size 256*256 and TPU.\n\nI am training a lot of models and came across something I do not fully understand.\n\nSome element of the hardware(CPU/GPU/TPU) is always at its maximum capacity, right? Because if that's not the case the training could go faster by using the hardware's maximum capacity. \n\nThe feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.\n\nHowever, if I increase the size of my network (B4--&gt; B6) the training time increases. This would strain the TPU more, but as this is not the bottleneck it should not matter for my training time.\n\nThis strikes me as odd as the same images are preloaded. Can anybody explain this?",
    "959146": "When a Kaggle Kernel runs, the job of data processing, creating datasets, Data Loaders, and iterating through data Loaders are all done by CPU. The task of traversing through your model, updating parameters is done by GPU/TPU. While GPU/TPU is training the model on a batch, CPU is busy preparing the next batch to train. So if CPU is the bottleneck, then GPU/TPU is just waiting for the CPU to prepare the next batch. \n\nSo CPU must be never the bottleneck so that you make the best use of the time of expensive resources, i.e. GPU/TPU.",
    "959258": "Allright, thanks this makes sense.",
    "959353": "&gt; The feedback Kaggle gives me is that CPU is almost always at its maximum capacity meaning that the loading of the images it most likely the bottleneck in speed.\n\nIn my popular notebook \"Triple Stratified KFold with TFRecords\" [here][1], the bottleneck is the GPU/TPU. If you disable all augmentation, the training time will not go faster because the CPU is not the bottleneck. The CPU is performing augmentation and providing images faster than the GPU/TPU can train.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "959781": "This explains what I am witnessing, thanks.",
    "959979": "In case you ever find that CPU is becoming a bottleneck (say due to lots of preprocessing), you may explore DALI library by Nvidia, which aims to perform data loading, augmentation, preprocessing all in GPU. I haven't tried it though. https://github.com/NVIDIA/DALI"
  },
  "source": "meta"
}