{
  "id": 202292,
  "title": "What's yours GPU Usage?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202292",
  "author_name": "",
  "post_date": "2020-12-09T09:24:19.295742400Z",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I have been tinkering with one or two models in this competition, using GPU. However, the usage is very low. This hasn't happened much with me before, therefore I would really love some guidance from the community.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F43785c916832780f856f06c7d55dd709%2FScreenshot%20(73).png?generation=1607505987560700&amp;alt=media\" alt=\"\"></p>\n<p>I have been working using Keras ImageDatagenerator on 256 X 256 resized images. It's taking about 8-9 min/epoch, which I think is really slow by GPU standards.</p>\n<p>Any advice on optimizing this performance?</p>",
  "messages": [
    {
      "id": "1106992",
      "postDate": "12/09/2020 09:24:19",
      "content": "<p>I have been tinkering with one or two models in this competition, using GPU. However, the usage is very low. This hasn't happened much with me before, therefore I would really love some guidance from the community.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F43785c916832780f856f06c7d55dd709%2FScreenshot%20(73).png?generation=1607505987560700&amp;alt=media\" alt=\"\"></p>\n<p>I have been working using Keras ImageDatagenerator on 256 X 256 resized images. It's taking about 8-9 min/epoch, which I think is really slow by GPU standards.</p>\n<p>Any advice on optimizing this performance?</p>",
      "rawMarkdown": "I have been tinkering with one or two models in this competition, using GPU. However, the usage is very low. This hasn't happened much with me before, therefore I would really love some guidance from the community.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F43785c916832780f856f06c7d55dd709%2FScreenshot%20(73).png?generation=1607505987560700&alt=media)\n\nI have been working using Keras ImageDatagenerator on 256 X 256 resized images. It's taking about 8-9 min/epoch, which I think is really slow by GPU standards.\n\nAny advice on optimizing this performance?",
      "votes": null
    },
    {
      "id": "1107073",
      "postDate": "12/09/2020 11:15:04",
      "content": "<p>Good question <a href=\"https://www.kaggle.com/fireheart7\" target=\"_blank\">@fireheart7</a> I'm with the same problem when running the runtime in GPU, It shows that the CPU use is high but most part of the time the GPU shows no use. This is strange, but I don't really know if it's just a bug in the system monitor or if it's really running on CPU (because the epochs are not so slow based on my experience running deep learning models on a CPU). Just looking for the system monitor looks like there's a competition for resources, where the notebook doesn't keep the GPU exclusively. </p>",
      "rawMarkdown": "Good question @fireheart7 I'm with the same problem when running the runtime in GPU, It shows that the CPU use is high but most part of the time the GPU shows no use. This is strange, but I don't really know if it's just a bug in the system monitor or if it's really running on CPU (because the epochs are not so slow based on my experience running deep learning models on a CPU). Just looking for the system monitor looks like there's a competition for resources, where the notebook doesn't keep the GPU exclusively.",
      "votes": null
    },
    {
      "id": "1107089",
      "postDate": "12/09/2020 11:35:47",
      "content": "<p>Thank you for the input <a href=\"https://www.kaggle.com/alvarole\" target=\"_blank\">@alvarole</a> . </p>\n<p>I myself have seen the GPU bursting past the 50% mark on several occasions during an epoch, however almost half the time it seems inactive. I believe this might be due to ImageDataGenerator itself. Maybe the supply of images in batches might be causing some context switching, rendering GPU inactive?</p>",
      "rawMarkdown": "Thank you for the input @alvarole . \n\nI myself have seen the GPU bursting past the 50% mark on several occasions during an epoch, however almost half the time it seems inactive. I believe this might be due to ImageDataGenerator itself. Maybe the supply of images in batches might be causing some context switching, rendering GPU inactive?",
      "votes": null
    },
    {
      "id": "1107117",
      "postDate": "12/09/2020 12:06:20",
      "content": "<p><a href=\"https://www.kaggle.com/fireheart7\" target=\"_blank\">@fireheart7</a> definitely the ImageDataGenerator isn't the best way to feed the model, but I've used on some other occasions (in Google Colab and also in my machine) and never see behavior like that. I'm using a batch size of 64 images with the 380x380 resolution, I think the GPU would have some work to train the model, and show inactivity in the system monitor is really strange for me. I'll modify the data ingestion to use the TF data API and let you know if have some change! </p>",
      "rawMarkdown": "fireheart7 definitely the ImageDataGenerator isn't the best way to feed the model, but I've used on some other occasions (in Google Colab and also in my machine) and never see behavior like that. I'm using a batch size of 64 images with the 380x380 resolution, I think the GPU would have some work to train the model, and show inactivity in the system monitor is really strange for me. I'll modify the data ingestion to use the TF data API and let you know if have some change!",
      "votes": null
    },
    {
      "id": "1107123",
      "postDate": "12/09/2020 12:12:33",
      "content": "<p>Yes, even I am using batch size of 64 with 256 X 256 images.</p>\n<p>Thank you so much!! Looking forward to those insights from TF data API pipeline.</p>",
      "rawMarkdown": "Yes, even I am using batch size of 64 with 256 X 256 images.\n\nThank you so much!! Looking forward to those insights from TF data API pipeline.",
      "votes": null
    },
    {
      "id": "1107127",
      "postDate": "12/09/2020 12:15:21",
      "content": "<p>I will be soon sharing a kernel where I will demonstrate how to move a Keras Image generator to GPU's rather than CPU. Will update here once done! Thanks!</p>",
      "rawMarkdown": "I will be soon sharing a kernel where I will demonstrate how to move a Keras Image generator to GPU's rather than CPU. Will update here once done! Thanks!",
      "votes": null
    },
    {
      "id": "1107136",
      "postDate": "12/09/2020 12:24:09",
      "content": "<p>That will be much helpful!! Thank you for this!! </p>",
      "rawMarkdown": "That will be much helpful!! Thank you for this!!",
      "votes": null
    },
    {
      "id": "1109003",
      "postDate": "12/11/2020 08:35:30",
      "content": "<p><strong>So, I implemented the tf data API pipeline on GPU, using TFRecords. Unlike Keras Image Data Generator, this was lightning fast!! Around 56 seconds per epoch and 90% + consistent GPU usage</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F77df974d258f093f98476825601ae83d%2FScreenshot%20(76).png?generation=1607675710582751&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "**So, I implemented the tf data API pipeline on GPU, using TFRecords. Unlike Keras Image Data Generator, this was lightning fast!! Around 56 seconds per epoch and 90% + consistent GPU usage**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F77df974d258f093f98476825601ae83d%2FScreenshot%20(76).png?generation=1607675710582751&alt=media)",
      "votes": null
    },
    {
      "id": "1109011",
      "postDate": "12/11/2020 08:42:16",
      "content": "<p>That's good news! Congratulations on your findings. :) </p>",
      "rawMarkdown": "That's good news! Congratulations on your findings. :)",
      "votes": null
    },
    {
      "id": "1109015",
      "postDate": "12/11/2020 08:47:00",
      "content": "<p>Thank you for your previous insights and assistance!! </p>",
      "rawMarkdown": "Thank you for your previous insights and assistance!!",
      "votes": null
    },
    {
      "id": "1109152",
      "postDate": "12/11/2020 11:23:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/fireheart7\" target=\"_blank\">@fireheart7</a>, the bottleneck in this case is probably due to the augmentations that you applied to your dataset. AFAIK the augmentations take place in the CPU - unless you explicitly specify the contrary. I train on my local PC and, since I want to reserve the full power to the training, I do the transformations on the CPU - which is always running high. Image transformations are computational demanding.</p>\n<p>Since we have 2CPU x 1 huge GPU in these notebooks, the CPU just can't keep the pace.</p>\n<p>One thing that you can do is to explore how to build your own Dataloader and transforms and to make them directly interact with the GPU (probably, sending the tensors to the GPU, doing the transformations and then get them back to CPU feed it to the model).</p>\n<p>Unfortunately, now I'm in PyTorch (you can take a look over <a href=\"https://www.kaggle.com/mawanda/load-entire-dataset-with-pytorch-api\" target=\"_blank\">there</a>, but some time ago I create a <a href=\"https://github.com/mawanda-jun/NoLabels/blob/master/Dataset/data_generator.py\" target=\"_blank\">custom dataloader</a> in TF 2.0.</p>\n<p>In these examples I do not take advantage of transformations in GPU, but you can understand the structure and it'd be easy to improve the CPU-GPU passing.</p>",
      "rawMarkdown": "Hi @fireheart7, the bottleneck in this case is probably due to the augmentations that you applied to your dataset. AFAIK the augmentations take place in the CPU - unless you explicitly specify the contrary. I train on my local PC and, since I want to reserve the full power to the training, I do the transformations on the CPU - which is always running high. Image transformations are computational demanding.\n\nSince we have 2CPU x 1 huge GPU in these notebooks, the CPU just can't keep the pace.\n\nOne thing that you can do is to explore how to build your own Dataloader and transforms and to make them directly interact with the GPU (probably, sending the tensors to the GPU, doing the transformations and then get them back to CPU feed it to the model).\n\nUnfortunately, now I'm in PyTorch (you can take a look over [there](https://www.kaggle.com/mawanda/load-entire-dataset-with-pytorch-api), but some time ago I create a [custom dataloader](https://github.com/mawanda-jun/NoLabels/blob/master/Dataset/data_generator.py) in TF 2.0.\n\nIn these examples I do not take advantage of transformations in GPU, but you can understand the structure and it'd be easy to improve the CPU-GPU passing.",
      "votes": null
    },
    {
      "id": "1113390",
      "postDate": "12/15/2020 12:12:07",
      "content": "<p>I was finally able to put something in my <a href=\"https://www.kaggle.com/harveenchadha/effnetb3-tf-data-gpu-aug-5x-speedup-baseline\" target=\"_blank\">kernel</a> here. Will continue to find other techniques like mixed precision training and will see if I get a boost!</p>",
      "rawMarkdown": "I was finally able to put something in my [kernel](https://www.kaggle.com/harveenchadha/effnetb3-tf-data-gpu-aug-5x-speedup-baseline) here. Will continue to find other techniques like mixed precision training and will see if I get a boost!",
      "votes": null
    },
    {
      "id": "1113508",
      "postDate": "12/15/2020 14:11:44",
      "content": "<p>Oh, nice! Thank you for sharing this. Much helpful!! </p>",
      "rawMarkdown": "Oh, nice! Thank you for sharing this. Much helpful!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1107073,
      "author_name": "alvarole",
      "author_url": "",
      "post_date": "12/09/2020 11:15:04",
      "content": "<p>Good question <a href=\"https://www.kaggle.com/fireheart7\" target=\"_blank\">@fireheart7</a> I'm with the same problem when running the runtime in GPU, It shows that the CPU use is high but most part of the time the GPU shows no use. This is strange, but I don't really know if it's just a bug in the system monitor or if it's really running on CPU (because the epochs are not so slow based on my experience running deep learning models on a CPU). Just looking for the system monitor looks like there's a competition for resources, where the notebook doesn't keep the GPU exclusively. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1107089,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "12/09/2020 11:35:47",
          "content": "<p>Thank you for the input <a href=\"https://www.kaggle.com/alvarole\" target=\"_blank\">@alvarole</a> . </p>\n<p>I myself have seen the GPU bursting past the 50% mark on several occasions during an epoch, however almost half the time it seems inactive. I believe this might be due to ImageDataGenerator itself. Maybe the supply of images in batches might be causing some context switching, rendering GPU inactive?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107117,
          "author_name": "alvarole",
          "author_url": "",
          "post_date": "12/09/2020 12:06:20",
          "content": "<p><a href=\"https://www.kaggle.com/fireheart7\" target=\"_blank\">@fireheart7</a> definitely the ImageDataGenerator isn't the best way to feed the model, but I've used on some other occasions (in Google Colab and also in my machine) and never see behavior like that. I'm using a batch size of 64 images with the 380x380 resolution, I think the GPU would have some work to train the model, and show inactivity in the system monitor is really strange for me. I'll modify the data ingestion to use the TF data API and let you know if have some change! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107123,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "12/09/2020 12:12:33",
          "content": "<p>Yes, even I am using batch size of 64 with 256 X 256 images.</p>\n<p>Thank you so much!! Looking forward to those insights from TF data API pipeline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109003,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "12/11/2020 08:35:30",
          "content": "<p><strong>So, I implemented the tf data API pipeline on GPU, using TFRecords. Unlike Keras Image Data Generator, this was lightning fast!! Around 56 seconds per epoch and 90% + consistent GPU usage</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F77df974d258f093f98476825601ae83d%2FScreenshot%20(76).png?generation=1607675710582751&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109011,
          "author_name": "alvarole",
          "author_url": "",
          "post_date": "12/11/2020 08:42:16",
          "content": "<p>That's good news! Congratulations on your findings. :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109015,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "12/11/2020 08:47:00",
          "content": "<p>Thank you for your previous insights and assistance!! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1107127,
      "author_name": "harveenchadha",
      "author_url": "",
      "post_date": "12/09/2020 12:15:21",
      "content": "<p>I will be soon sharing a kernel where I will demonstrate how to move a Keras Image generator to GPU's rather than CPU. Will update here once done! Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107136,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "12/09/2020 12:24:09",
          "content": "<p>That will be much helpful!! Thank you for this!! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1113390,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "12/15/2020 12:12:07",
          "content": "<p>I was finally able to put something in my <a href=\"https://www.kaggle.com/harveenchadha/effnetb3-tf-data-gpu-aug-5x-speedup-baseline\" target=\"_blank\">kernel</a> here. Will continue to find other techniques like mixed precision training and will see if I get a boost!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1113508,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "12/15/2020 14:11:44",
          "content": "<p>Oh, nice! Thank you for sharing this. Much helpful!! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109152,
      "author_name": "mawanda",
      "author_url": "",
      "post_date": "12/11/2020 11:23:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/fireheart7\" target=\"_blank\">@fireheart7</a>, the bottleneck in this case is probably due to the augmentations that you applied to your dataset. AFAIK the augmentations take place in the CPU - unless you explicitly specify the contrary. I train on my local PC and, since I want to reserve the full power to the training, I do the transformations on the CPU - which is always running high. Image transformations are computational demanding.</p>\n<p>Since we have 2CPU x 1 huge GPU in these notebooks, the CPU just can't keep the pace.</p>\n<p>One thing that you can do is to explore how to build your own Dataloader and transforms and to make them directly interact with the GPU (probably, sending the tensors to the GPU, doing the transformations and then get them back to CPU feed it to the model).</p>\n<p>Unfortunately, now I'm in PyTorch (you can take a look over <a href=\"https://www.kaggle.com/mawanda/load-entire-dataset-with-pytorch-api\" target=\"_blank\">there</a>, but some time ago I create a <a href=\"https://github.com/mawanda-jun/NoLabels/blob/master/Dataset/data_generator.py\" target=\"_blank\">custom dataloader</a> in TF 2.0.</p>\n<p>In these examples I do not take advantage of transformations in GPU, but you can understand the structure and it'd be easy to improve the CPU-GPU passing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1106992": "I have been tinkering with one or two models in this competition, using GPU. However, the usage is very low. This hasn't happened much with me before, therefore I would really love some guidance from the community.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F43785c916832780f856f06c7d55dd709%2FScreenshot%20(73).png?generation=1607505987560700&alt=media)\n\nI have been working using Keras ImageDatagenerator on 256 X 256 resized images. It's taking about 8-9 min/epoch, which I think is really slow by GPU standards.\n\nAny advice on optimizing this performance?",
    "1107073": "Good question @fireheart7 I'm with the same problem when running the runtime in GPU, It shows that the CPU use is high but most part of the time the GPU shows no use. This is strange, but I don't really know if it's just a bug in the system monitor or if it's really running on CPU (because the epochs are not so slow based on my experience running deep learning models on a CPU). Just looking for the system monitor looks like there's a competition for resources, where the notebook doesn't keep the GPU exclusively.",
    "1107089": "Thank you for the input @alvarole . \n\nI myself have seen the GPU bursting past the 50% mark on several occasions during an epoch, however almost half the time it seems inactive. I believe this might be due to ImageDataGenerator itself. Maybe the supply of images in batches might be causing some context switching, rendering GPU inactive?",
    "1107117": "fireheart7 definitely the ImageDataGenerator isn't the best way to feed the model, but I've used on some other occasions (in Google Colab and also in my machine) and never see behavior like that. I'm using a batch size of 64 images with the 380x380 resolution, I think the GPU would have some work to train the model, and show inactivity in the system monitor is really strange for me. I'll modify the data ingestion to use the TF data API and let you know if have some change!",
    "1107123": "Yes, even I am using batch size of 64 with 256 X 256 images.\n\nThank you so much!! Looking forward to those insights from TF data API pipeline.",
    "1107127": "I will be soon sharing a kernel where I will demonstrate how to move a Keras Image generator to GPU's rather than CPU. Will update here once done! Thanks!",
    "1107136": "That will be much helpful!! Thank you for this!!",
    "1109003": "**So, I implemented the tf data API pipeline on GPU, using TFRecords. Unlike Keras Image Data Generator, this was lightning fast!! Around 56 seconds per epoch and 90% + consistent GPU usage**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2599569%2F77df974d258f093f98476825601ae83d%2FScreenshot%20(76).png?generation=1607675710582751&alt=media)",
    "1109011": "That's good news! Congratulations on your findings. :)",
    "1109015": "Thank you for your previous insights and assistance!!",
    "1109152": "Hi @fireheart7, the bottleneck in this case is probably due to the augmentations that you applied to your dataset. AFAIK the augmentations take place in the CPU - unless you explicitly specify the contrary. I train on my local PC and, since I want to reserve the full power to the training, I do the transformations on the CPU - which is always running high. Image transformations are computational demanding.\n\nSince we have 2CPU x 1 huge GPU in these notebooks, the CPU just can't keep the pace.\n\nOne thing that you can do is to explore how to build your own Dataloader and transforms and to make them directly interact with the GPU (probably, sending the tensors to the GPU, doing the transformations and then get them back to CPU feed it to the model).\n\nUnfortunately, now I'm in PyTorch (you can take a look over [there](https://www.kaggle.com/mawanda/load-entire-dataset-with-pytorch-api), but some time ago I create a [custom dataloader](https://github.com/mawanda-jun/NoLabels/blob/master/Dataset/data_generator.py) in TF 2.0.\n\nIn these examples I do not take advantage of transformations in GPU, but you can understand the structure and it'd be easy to improve the CPU-GPU passing.",
    "1113390": "I was finally able to put something in my [kernel](https://www.kaggle.com/harveenchadha/effnetb3-tf-data-gpu-aug-5x-speedup-baseline) here. Will continue to find other techniques like mixed precision training and will see if I get a boost!",
    "1113508": "Oh, nice! Thank you for sharing this. Much helpful!!"
  },
  "source": "meta"
}