{
  "id": 313979,
  "title": "Need tips for better GPU utilization",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/313979",
  "author_name": "",
  "post_date": "2022-03-20T06:55:57.280679400Z",
  "votes": 6,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I was training a pytorch model on the competition data, but when training I found that the pytorch model is not utilizing GPU completely [highest 12-17%]. And because of that its taking ~25min for each epoch and img size is 1x256x256.</p>\n<p>\n<img src=\"https://i.imgur.com/dCmC9Mw.png\">\n</p>\n<p>I tried different things to fix that, eg:</p>\n<ul>\n<li>increasing <code>batch size</code></li>\n<li>adjusting <code>learning rate</code></li>\n<li>adding <code>num_workers = 4</code> in the data loader</li>\n<li>adding <code>pin_memory = True</code> in the data loader</li>\n<li><code>Shuffle= Flase</code> in the data loader</li>\n</ul>\n<p>I also checked the training loop, both model and the data are in cuda. I don't know what I'm missing here.</p>\n<p>I would be great if anyone can suggest some ways to increase the GPU utilization. Thank you.</p>\n<p>W&amp;B logs: <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline?workspace=user-somusan\" target=\"_blank\">https://wandb.ai/somusan/PogChamp2%20Baseline?workspace=user-somusan</a></p>\n<p>Below is the dataset class and dataloader that I'm using,</p>\n<pre><code>################## DATASET CLASS ##################\nclass POGdata(Dataset):\n\n    def __init__(self,\n                 df,\n                 data_dir,\n                 transform = None):\n        super(POGdata, self).__init__()\n        self.data_dir = data_dir\n        self.df = df\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, index):\n\n        path = self.df[\"filename\"].iloc[index]\n        path = os.path.join(self.data_dir, path)\n\n        label = self.df[\"genre_id\"].iloc[index]\n        mono_audio = self.load_audio(path)\n        mono_audio = mono_audio.unsqueeze(dim=0)\n        return mono_audio, label\n\n\n    def load_audio(self, path):\n        audio, _ = torchaudio.load(path)\n        if self.transform != None:\n            for aug in self.transform:\n                audio = aug(audio)\n        return audio[0,:]\n\n################## TRANSFORMS ##################\nSAMPLE_RATE = 22050\n\naugm = [\n    MelSpectrogram(sample_rate=SAMPLE_RATE,\n                    n_fft=1024,\n                    hop_length=512,\n                    n_mels=256),\n    AmplitudeToDB(),\n    Resize((256, 256))\n]\n\n########## DEFINING TRAIN AND VALIDATION DATA LOADER #########\n\ntrain_ds = POGdata(train_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\nvalid_ds = POGdata(valid_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\n\ntrain_dl = DataLoader(train_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\nvalid_dl = DataLoader(valid_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\n</code></pre>",
  "messages": [
    {
      "id": "1729519",
      "postDate": "03/20/2022 06:55:57",
      "content": "<p>I was training a pytorch model on the competition data, but when training I found that the pytorch model is not utilizing GPU completely [highest 12-17%]. And because of that its taking ~25min for each epoch and img size is 1x256x256.</p>\n<p>\n<img src=\"https://i.imgur.com/dCmC9Mw.png\">\n</p>\n<p>I tried different things to fix that, eg:</p>\n<ul>\n<li>increasing <code>batch size</code></li>\n<li>adjusting <code>learning rate</code></li>\n<li>adding <code>num_workers = 4</code> in the data loader</li>\n<li>adding <code>pin_memory = True</code> in the data loader</li>\n<li><code>Shuffle= Flase</code> in the data loader</li>\n</ul>\n<p>I also checked the training loop, both model and the data are in cuda. I don't know what I'm missing here.</p>\n<p>I would be great if anyone can suggest some ways to increase the GPU utilization. Thank you.</p>\n<p>W&amp;B logs: <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline?workspace=user-somusan\" target=\"_blank\">https://wandb.ai/somusan/PogChamp2%20Baseline?workspace=user-somusan</a></p>\n<p>Below is the dataset class and dataloader that I'm using,</p>\n<pre><code>################## DATASET CLASS ##################\nclass POGdata(Dataset):\n\n    def __init__(self,\n                 df,\n                 data_dir,\n                 transform = None):\n        super(POGdata, self).__init__()\n        self.data_dir = data_dir\n        self.df = df\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, index):\n\n        path = self.df[\"filename\"].iloc[index]\n        path = os.path.join(self.data_dir, path)\n\n        label = self.df[\"genre_id\"].iloc[index]\n        mono_audio = self.load_audio(path)\n        mono_audio = mono_audio.unsqueeze(dim=0)\n        return mono_audio, label\n\n\n    def load_audio(self, path):\n        audio, _ = torchaudio.load(path)\n        if self.transform != None:\n            for aug in self.transform:\n                audio = aug(audio)\n        return audio[0,:]\n\n################## TRANSFORMS ##################\nSAMPLE_RATE = 22050\n\naugm = [\n    MelSpectrogram(sample_rate=SAMPLE_RATE,\n                    n_fft=1024,\n                    hop_length=512,\n                    n_mels=256),\n    AmplitudeToDB(),\n    Resize((256, 256))\n]\n\n########## DEFINING TRAIN AND VALIDATION DATA LOADER #########\n\ntrain_ds = POGdata(train_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\nvalid_ds = POGdata(valid_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\n\ntrain_dl = DataLoader(train_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\nvalid_dl = DataLoader(valid_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\n</code></pre>",
      "rawMarkdown": "I was training a pytorch model on the competition data, but when training I found that the pytorch model is not utilizing GPU completely [highest 12-17%]. And because of that its taking ~25min for each epoch and img size is 1x256x256.\n\n<p align=\"center\">\n<img width=\"600\" src=\"https://i.imgur.com/dCmC9Mw.png\">\n</p>\n\nI tried different things to fix that, eg:\n- increasing `batch size`\n- adjusting `learning rate`\n- adding `num_workers = 4` in the data loader\n- adding `pin_memory = True` in the data loader\n- `Shuffle= Flase` in the data loader\n\nI also checked the training loop, both model and the data are in cuda. I don't know what I'm missing here.\n\nI would be great if anyone can suggest some ways to increase the GPU utilization. Thank you.\n\nW&B logs: https://wandb.ai/somusan/PogChamp2%20Baseline?workspace=user-somusan\n\nBelow is the dataset class and dataloader that I'm using,\n```python\n################## DATASET CLASS ##################\nclass POGdata(Dataset):\n    \n    def __init__(self,\n                 df,\n                 data_dir,\n                 transform = None):\n        super(POGdata, self).__init__()\n        self.data_dir = data_dir\n        self.df = df\n        self.transform = transform\n        \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, index):\n\n        path = self.df[\"filename\"].iloc[index]\n        path = os.path.join(self.data_dir, path)\n\n        label = self.df[\"genre_id\"].iloc[index]\n        mono_audio = self.load_audio(path)\n        mono_audio = mono_audio.unsqueeze(dim=0)\n        return mono_audio, label\n    \n    \n    def load_audio(self, path):\n        audio, _ = torchaudio.load(path)\n        if self.transform != None:\n            for aug in self.transform:\n                audio = aug(audio)\n        return audio[0,:]\n\n################## TRANSFORMS ##################\nSAMPLE_RATE = 22050\n\naugm = [\n    MelSpectrogram(sample_rate=SAMPLE_RATE,\n                    n_fft=1024,\n                    hop_length=512,\n                    n_mels=256),\n    AmplitudeToDB(),\n    Resize((256, 256))\n]\n\n########## DEFINING TRAIN AND VALIDATION DATA LOADER #########\n\ntrain_ds = POGdata(train_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\nvalid_ds = POGdata(valid_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\n\ntrain_dl = DataLoader(train_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\nvalid_dl = DataLoader(valid_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\n```",
      "votes": null
    },
    {
      "id": "1729889",
      "postDate": "03/20/2022 16:56:40",
      "content": "<p>I am not a pytorch user but noob answer.</p>\n<p>Do you move you model and input to cuda using .to('cuda') ?</p>",
      "rawMarkdown": "I am not a pytorch user but noob answer.\n\nDo you move you model and input to cuda using .to('cuda') ?",
      "votes": null
    },
    {
      "id": "1729949",
      "postDate": "03/20/2022 18:09:48",
      "content": "<p>Maybe its the pin_memory thingy? I never use pin memory = True. Maybe try switching it off?</p>",
      "rawMarkdown": "Maybe its the pin_memory thingy? I never use pin memory = True. Maybe try switching it off?",
      "votes": null
    },
    {
      "id": "1730004",
      "postDate": "03/20/2022 19:50:59",
      "content": "<p>thank you for answering, Ok I will try that.<br>\nI was actually following an article by William Falcon</p>\n<ul>\n<li><a href=\"https://towardsdatascience.com/7-tips-for-squeezing-maximum-performance-from-pytorch-ca4a40951259\" target=\"_blank\">7 Tips To Maximize PyTorch Performance</a> <br>\nthere he mentioned keeping the pin_memory=True.</li>\n</ul>\n<p>BTW, can I ask you how much time per epoch its taking for your model?</p>",
      "rawMarkdown": "thank you for answering, Ok I will try that.\nI was actually following an article by William Falcon\n- [7 Tips To Maximize PyTorch Performance](https://towardsdatascience.com/7-tips-for-squeezing-maximum-performance-from-pytorch-ca4a40951259) \nthere he mentioned keeping the pin_memory=True.\n\nBTW, can I ask you how much time per epoch its taking for your model?",
      "votes": null
    },
    {
      "id": "1730006",
      "postDate": "03/20/2022 19:54:30",
      "content": "<p>thank you for answering, <br>\nYeah I made sure that the model and input are in cuda. Also checked using <code>.is_cuda()</code> if that is in cuda or not.</p>",
      "rawMarkdown": "thank you for answering, \nYeah I made sure that the model and input are in cuda. Also checked using `.is_cuda()` if that is in cuda or not.",
      "votes": null
    },
    {
      "id": "1730080",
      "postDate": "03/20/2022 22:42:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>, I experimented with the <code>pin_memory</code> and <code>num_workers</code>. These are the results,</p>\n<pre><code>Exp 1:\npin_memory = False    ---|___&gt; 32 min/epoch\nnum_workers = 4       ---|   \n\nExp:2\npin_memory = True   ---|___&gt; 27 min/epoch\nnum_workers = 2     ---|   \n\nExp:3\npin_memory = False    ---|___&gt; 26 min/epoch\nnum_workers = 2       ---|   \n</code></pre>\n<p>Exp 1 --&gt; <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/36z1f61a?workspace=user-somusan\" target=\"_blank\">W&amp;B link</a><br>\nExp 2--&gt; <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/35a4odb4?workspace=user-somusan\" target=\"_blank\">W&amp;B link</a><br>\nExp 3 --&gt; <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/5hz1rfi6?workspace=user-somusan\" target=\"_blank\">W&amp;B link</a></p>\n<p>I was using num_workers=4 because it was given <code>num_worker = 4 * num_GPU\n</code> but when I used <code>num_workers=4</code> it was giving me this warning,</p>\n<blockquote>\n  <p>UserWarning: This DataLoader will create 4 worker processes in total. Our suggested max number of worker in current system is 2, which is smaller than what this DataLoader is going to create. Please be aware that excessive worker creation might get DataLoader running slow or even freeze, lower the worker number to avoid potential slowness/freeze if necessary.</p>\n</blockquote>\n<p>That's why I changed it to 2. Seems like not using <code>pin_memory</code> is comparatively good [as you mentioned]. But still, there was not a big boost in GPU usage. Im using a very small CNN model architecture, do you think that is a problem here? </p>",
      "rawMarkdown": "Hi @pheadrus, I experimented with the `pin_memory` and `num_workers`. These are the results,\n\n```\nExp 1:\npin_memory = False    ---|___> 32 min/epoch\nnum_workers = 4       ---|   \n\nExp:2\npin_memory = True   ---|___> 27 min/epoch\nnum_workers = 2     ---|   \n\nExp:3\npin_memory = False    ---|___> 26 min/epoch\nnum_workers = 2       ---|   \n```\nExp 1 --> [W&B link](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/36z1f61a?workspace=user-somusan)\nExp 2--> [W&B link](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/35a4odb4?workspace=user-somusan)\nExp 3 --> [W&B link](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/5hz1rfi6?workspace=user-somusan)\n\nI was using num_workers=4 because it was given `num_worker = 4 * num_GPU\n` but when I used `num_workers=4` it was giving me this warning,\n> UserWarning: This DataLoader will create 4 worker processes in total. Our suggested max number of worker in current system is 2, which is smaller than what this DataLoader is going to create. Please be aware that excessive worker creation might get DataLoader running slow or even freeze, lower the worker number to avoid potential slowness/freeze if necessary.\n\nThat's why I changed it to 2. Seems like not using `pin_memory` is comparatively good [as you mentioned]. But still, there was not a big boost in GPU usage. Im using a very small CNN model architecture, do you think that is a problem here?",
      "votes": null
    },
    {
      "id": "1730235",
      "postDate": "03/21/2022 03:42:22",
      "content": "<p>Maybe you should try saving raw spec images on disk and then load from there? Typically its the I/O that takes most time if your net is small. Also try reducing image size to see how things change. </p>",
      "rawMarkdown": "Maybe you should try saving raw spec images on disk and then load from there? Typically its the I/O that takes most time if your net is small. Also try reducing image size to see how things change.",
      "votes": null
    },
    {
      "id": "1730257",
      "postDate": "03/21/2022 04:14:40",
      "content": "<p>I tried that previously with <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> 's spectrogram data, it was taking less time ~10min, but it was not using GPU to its full limits,<br>\nYou can see the logs <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/1tub37h5/system?workspace=user-somusan\" target=\"_blank\">here</a><br>\nIn the data loader I was using <code>pin_memory=True,num_workers  = 4</code> maybe I need to change that. I will also try using small image size.</p>",
      "rawMarkdown": "I tried that previously with @harveenchadha 's spectrogram data, it was taking less time ~10min, but it was not using GPU to its full limits,\nYou can see the logs [here](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/1tub37h5/system?workspace=user-somusan)\nIn the data loader I was using `pin_memory=True,num_workers  = 4` maybe I need to change that. I will also try using small image size.",
      "votes": null
    },
    {
      "id": "1730709",
      "postDate": "03/21/2022 14:45:45",
      "content": "<table>\n<thead>\n<tr>\n<th>_</th>\n<th>_</th>\n</tr>\n</thead>\n<tbody>\n</tbody>\n</table>\n<h2><img src=\"https://i.imgur.com/3jWSz9C.png\">  |  <img src=\"https://i.imgur.com/CMzB0gA.png\"></h2>\n<p>I changed the small model with <code>resnet50</code> and now it seems like its utilizing the GPU way better. These small peaks near 20 are actually near 100% in the kaggle GPU usage metric. And I believe It will reduce even more if I use saved raw spec images directly in the pipeline.<br>\n<img src=\"https://i.imgur.com/0qB0564.png\"><br>\nHad to reduce the batch size bcoz of OOM, and including training and validation its taking a total of 40 min  [32min/8min] per epoch [its an 80/20 split]. Im not sure if anyone is having the same time per epoch. please let me know if anyone is getting the same ETA. Thanks, <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> and <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> for helping.<br>\nWandB: <a href=\"https://wandb.ai/somusan/pogchamp2/runs/3j7yisba/system?workspace=user-somusan\" target=\"_blank\">https://wandb.ai/somusan/pogchamp2/runs/3j7yisba/system?workspace=user-somusan</a></p>",
      "rawMarkdown": "_    |     _\n:-------------------------:|:-------------------------:\n<img width=\"120\" src=\"https://i.imgur.com/3jWSz9C.png\">  |  <img width=\"120\" src=\"https://i.imgur.com/CMzB0gA.png\">\n\n\n---\n\nI changed the small model with `resnet50` and now it seems like its utilizing the GPU way better. These small peaks near 20 are actually near 100% in the kaggle GPU usage metric. And I believe It will reduce even more if I use saved raw spec images directly in the pipeline.\n<img width=\"600\" src=\"https://i.imgur.com/0qB0564.png\">\nHad to reduce the batch size bcoz of OOM, and including training and validation its taking a total of 40 min  [32min/8min] per epoch [its an 80/20 split]. Im not sure if anyone is having the same time per epoch. please let me know if anyone is getting the same ETA. Thanks, @pheadrus and @harveenchadha for helping.\nWandB: https://wandb.ai/somusan/pogchamp2/runs/3j7yisba/system?workspace=user-somusan",
      "votes": null
    },
    {
      "id": "1730829",
      "postDate": "03/21/2022 16:37:45",
      "content": "<p>Great stuff <a href=\"https://www.kaggle.com/soumya9977\" target=\"_blank\">@soumya9977</a> . You can also do gradient accumulation if you BS is too low (eg 8-16 ish).</p>",
      "rawMarkdown": "Great stuff @soumya9977 . You can also do gradient accumulation if you BS is too low (eg 8-16 ish).",
      "votes": null
    },
    {
      "id": "1730876",
      "postDate": "03/21/2022 17:47:28",
      "content": "<p>I'm sorry for this silly question, but what do you mean by <code>gradient accumulation if your BS is too low</code>? is it like accumulating gradients from all small mini-batches and updating the params after whole epoch?</p>",
      "rawMarkdown": "I'm sorry for this silly question, but what do you mean by `gradient accumulation if your BS is too low`? is it like accumulating gradients from all small mini-batches and updating the params after whole epoch?",
      "votes": null
    },
    {
      "id": "1731628",
      "postDate": "03/22/2022 14:49:57",
      "content": "<p>You can check this <a href=\"https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html\" target=\"_blank\">blog</a> to learn about gradient accumulation </p>",
      "rawMarkdown": "You can check this [blog](https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html) to learn about gradient accumulation",
      "votes": null
    },
    {
      "id": "1731707",
      "postDate": "03/22/2022 16:07:15",
      "content": "<p>thank you so much for sharing <a href=\"https://www.kaggle.com/kpriyanshu256\" target=\"_blank\">@kpriyanshu256</a>, I will definitely try that. seems like the core idea is almost the same.</p>",
      "rawMarkdown": "thank you so much for sharing @kpriyanshu256, I will definitely try that. seems like the core idea is almost the same.",
      "votes": null
    },
    {
      "id": "1731989",
      "postDate": "03/22/2022 23:01:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/soumya9977\" target=\"_blank\">@soumya9977</a>,</p>\n<p>accessing a Pandas Dateframe in <strong>getitem</strong> is slow. I always convert it into a numpy or TorchTensor in <strong>init</strong>. You can also do the path transformation before.</p>",
      "rawMarkdown": "Hi @soumya9977,\n\naccessing a Pandas Dateframe in __getitem__ is slow. I always convert it into a numpy or TorchTensor in __init__. You can also do the path transformation before.",
      "votes": null
    },
    {
      "id": "1732156",
      "postDate": "03/23/2022 04:40:30",
      "content": "<p>That is a nice point, I generally use pandas df in getitem but from now on I will try to convert it into a numpy or TorchTensor in init and then use that in getitem. Thank you so muchhh.</p>",
      "rawMarkdown": "That is a nice point, I generally use pandas df in getitem but from now on I will try to convert it into a numpy or TorchTensor in init and then use that in getitem. Thank you so muchhh.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1729889,
      "author_name": "harveenchadha",
      "author_url": "",
      "post_date": "03/20/2022 16:56:40",
      "content": "<p>I am not a pytorch user but noob answer.</p>\n<p>Do you move you model and input to cuda using .to('cuda') ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1730006,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/20/2022 19:54:30",
          "content": "<p>thank you for answering, <br>\nYeah I made sure that the model and input are in cuda. Also checked using <code>.is_cuda()</code> if that is in cuda or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1729949,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "03/20/2022 18:09:48",
      "content": "<p>Maybe its the pin_memory thingy? I never use pin memory = True. Maybe try switching it off?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1730004,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/20/2022 19:50:59",
          "content": "<p>thank you for answering, Ok I will try that.<br>\nI was actually following an article by William Falcon</p>\n<ul>\n<li><a href=\"https://towardsdatascience.com/7-tips-for-squeezing-maximum-performance-from-pytorch-ca4a40951259\" target=\"_blank\">7 Tips To Maximize PyTorch Performance</a> <br>\nthere he mentioned keeping the pin_memory=True.</li>\n</ul>\n<p>BTW, can I ask you how much time per epoch its taking for your model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730080,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/20/2022 22:42:23",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>, I experimented with the <code>pin_memory</code> and <code>num_workers</code>. These are the results,</p>\n<pre><code>Exp 1:\npin_memory = False    ---|___&gt; 32 min/epoch\nnum_workers = 4       ---|   \n\nExp:2\npin_memory = True   ---|___&gt; 27 min/epoch\nnum_workers = 2     ---|   \n\nExp:3\npin_memory = False    ---|___&gt; 26 min/epoch\nnum_workers = 2       ---|   \n</code></pre>\n<p>Exp 1 --&gt; <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/36z1f61a?workspace=user-somusan\" target=\"_blank\">W&amp;B link</a><br>\nExp 2--&gt; <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/35a4odb4?workspace=user-somusan\" target=\"_blank\">W&amp;B link</a><br>\nExp 3 --&gt; <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/5hz1rfi6?workspace=user-somusan\" target=\"_blank\">W&amp;B link</a></p>\n<p>I was using num_workers=4 because it was given <code>num_worker = 4 * num_GPU\n</code> but when I used <code>num_workers=4</code> it was giving me this warning,</p>\n<blockquote>\n  <p>UserWarning: This DataLoader will create 4 worker processes in total. Our suggested max number of worker in current system is 2, which is smaller than what this DataLoader is going to create. Please be aware that excessive worker creation might get DataLoader running slow or even freeze, lower the worker number to avoid potential slowness/freeze if necessary.</p>\n</blockquote>\n<p>That's why I changed it to 2. Seems like not using <code>pin_memory</code> is comparatively good [as you mentioned]. But still, there was not a big boost in GPU usage. Im using a very small CNN model architecture, do you think that is a problem here? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730235,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/21/2022 03:42:22",
          "content": "<p>Maybe you should try saving raw spec images on disk and then load from there? Typically its the I/O that takes most time if your net is small. Also try reducing image size to see how things change. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730257,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/21/2022 04:14:40",
          "content": "<p>I tried that previously with <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> 's spectrogram data, it was taking less time ~10min, but it was not using GPU to its full limits,<br>\nYou can see the logs <a href=\"https://wandb.ai/somusan/PogChamp2%20Baseline/runs/1tub37h5/system?workspace=user-somusan\" target=\"_blank\">here</a><br>\nIn the data loader I was using <code>pin_memory=True,num_workers  = 4</code> maybe I need to change that. I will also try using small image size.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730709,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/21/2022 14:45:45",
          "content": "<table>\n<thead>\n<tr>\n<th>_</th>\n<th>_</th>\n</tr>\n</thead>\n<tbody>\n</tbody>\n</table>\n<h2><img src=\"https://i.imgur.com/3jWSz9C.png\">  |  <img src=\"https://i.imgur.com/CMzB0gA.png\"></h2>\n<p>I changed the small model with <code>resnet50</code> and now it seems like its utilizing the GPU way better. These small peaks near 20 are actually near 100% in the kaggle GPU usage metric. And I believe It will reduce even more if I use saved raw spec images directly in the pipeline.<br>\n<img src=\"https://i.imgur.com/0qB0564.png\"><br>\nHad to reduce the batch size bcoz of OOM, and including training and validation its taking a total of 40 min  [32min/8min] per epoch [its an 80/20 split]. Im not sure if anyone is having the same time per epoch. please let me know if anyone is getting the same ETA. Thanks, <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> and <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> for helping.<br>\nWandB: <a href=\"https://wandb.ai/somusan/pogchamp2/runs/3j7yisba/system?workspace=user-somusan\" target=\"_blank\">https://wandb.ai/somusan/pogchamp2/runs/3j7yisba/system?workspace=user-somusan</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730829,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/21/2022 16:37:45",
          "content": "<p>Great stuff <a href=\"https://www.kaggle.com/soumya9977\" target=\"_blank\">@soumya9977</a> . You can also do gradient accumulation if you BS is too low (eg 8-16 ish).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730876,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/21/2022 17:47:28",
          "content": "<p>I'm sorry for this silly question, but what do you mean by <code>gradient accumulation if your BS is too low</code>? is it like accumulating gradients from all small mini-batches and updating the params after whole epoch?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1731628,
          "author_name": "kpriyanshu256",
          "author_url": "",
          "post_date": "03/22/2022 14:49:57",
          "content": "<p>You can check this <a href=\"https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html\" target=\"_blank\">blog</a> to learn about gradient accumulation </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1731707,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/22/2022 16:07:15",
          "content": "<p>thank you so much for sharing <a href=\"https://www.kaggle.com/kpriyanshu256\" target=\"_blank\">@kpriyanshu256</a>, I will definitely try that. seems like the core idea is almost the same.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1731989,
      "author_name": "joatom",
      "author_url": "",
      "post_date": "03/22/2022 23:01:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/soumya9977\" target=\"_blank\">@soumya9977</a>,</p>\n<p>accessing a Pandas Dateframe in <strong>getitem</strong> is slow. I always convert it into a numpy or TorchTensor in <strong>init</strong>. You can also do the path transformation before.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1732156,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "03/23/2022 04:40:30",
          "content": "<p>That is a nice point, I generally use pandas df in getitem but from now on I will try to convert it into a numpy or TorchTensor in init and then use that in getitem. Thank you so muchhh.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1729519": "I was training a pytorch model on the competition data, but when training I found that the pytorch model is not utilizing GPU completely [highest 12-17%]. And because of that its taking ~25min for each epoch and img size is 1x256x256.\n\n<p align=\"center\">\n<img width=\"600\" src=\"https://i.imgur.com/dCmC9Mw.png\">\n</p>\n\nI tried different things to fix that, eg:\n- increasing `batch size`\n- adjusting `learning rate`\n- adding `num_workers = 4` in the data loader\n- adding `pin_memory = True` in the data loader\n- `Shuffle= Flase` in the data loader\n\nI also checked the training loop, both model and the data are in cuda. I don't know what I'm missing here.\n\nI would be great if anyone can suggest some ways to increase the GPU utilization. Thank you.\n\nW&B logs: https://wandb.ai/somusan/PogChamp2%20Baseline?workspace=user-somusan\n\nBelow is the dataset class and dataloader that I'm using,\n```python\n################## DATASET CLASS ##################\nclass POGdata(Dataset):\n    \n    def __init__(self,\n                 df,\n                 data_dir,\n                 transform = None):\n        super(POGdata, self).__init__()\n        self.data_dir = data_dir\n        self.df = df\n        self.transform = transform\n        \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, index):\n\n        path = self.df[\"filename\"].iloc[index]\n        path = os.path.join(self.data_dir, path)\n\n        label = self.df[\"genre_id\"].iloc[index]\n        mono_audio = self.load_audio(path)\n        mono_audio = mono_audio.unsqueeze(dim=0)\n        return mono_audio, label\n    \n    \n    def load_audio(self, path):\n        audio, _ = torchaudio.load(path)\n        if self.transform != None:\n            for aug in self.transform:\n                audio = aug(audio)\n        return audio[0,:]\n\n################## TRANSFORMS ##################\nSAMPLE_RATE = 22050\n\naugm = [\n    MelSpectrogram(sample_rate=SAMPLE_RATE,\n                    n_fft=1024,\n                    hop_length=512,\n                    n_mels=256),\n    AmplitudeToDB(),\n    Resize((256, 256))\n]\n\n########## DEFINING TRAIN AND VALIDATION DATA LOADER #########\n\ntrain_ds = POGdata(train_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\nvalid_ds = POGdata(valid_df, \"../input/kaggle-pog-series-s01e02/train\", transform = augm)\n\ntrain_dl = DataLoader(train_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\nvalid_dl = DataLoader(valid_ds, batch_size=512, shuffle=False,pin_memory=True,num_workers  = 4)\n```",
    "1729889": "I am not a pytorch user but noob answer.\n\nDo you move you model and input to cuda using .to('cuda') ?",
    "1729949": "Maybe its the pin_memory thingy? I never use pin memory = True. Maybe try switching it off?",
    "1730004": "thank you for answering, Ok I will try that.\nI was actually following an article by William Falcon\n- [7 Tips To Maximize PyTorch Performance](https://towardsdatascience.com/7-tips-for-squeezing-maximum-performance-from-pytorch-ca4a40951259) \nthere he mentioned keeping the pin_memory=True.\n\nBTW, can I ask you how much time per epoch its taking for your model?",
    "1730006": "thank you for answering, \nYeah I made sure that the model and input are in cuda. Also checked using `.is_cuda()` if that is in cuda or not.",
    "1730080": "Hi @pheadrus, I experimented with the `pin_memory` and `num_workers`. These are the results,\n\n```\nExp 1:\npin_memory = False    ---|___> 32 min/epoch\nnum_workers = 4       ---|   \n\nExp:2\npin_memory = True   ---|___> 27 min/epoch\nnum_workers = 2     ---|   \n\nExp:3\npin_memory = False    ---|___> 26 min/epoch\nnum_workers = 2       ---|   \n```\nExp 1 --> [W&B link](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/36z1f61a?workspace=user-somusan)\nExp 2--> [W&B link](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/35a4odb4?workspace=user-somusan)\nExp 3 --> [W&B link](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/5hz1rfi6?workspace=user-somusan)\n\nI was using num_workers=4 because it was given `num_worker = 4 * num_GPU\n` but when I used `num_workers=4` it was giving me this warning,\n> UserWarning: This DataLoader will create 4 worker processes in total. Our suggested max number of worker in current system is 2, which is smaller than what this DataLoader is going to create. Please be aware that excessive worker creation might get DataLoader running slow or even freeze, lower the worker number to avoid potential slowness/freeze if necessary.\n\nThat's why I changed it to 2. Seems like not using `pin_memory` is comparatively good [as you mentioned]. But still, there was not a big boost in GPU usage. Im using a very small CNN model architecture, do you think that is a problem here?",
    "1730235": "Maybe you should try saving raw spec images on disk and then load from there? Typically its the I/O that takes most time if your net is small. Also try reducing image size to see how things change.",
    "1730257": "I tried that previously with @harveenchadha 's spectrogram data, it was taking less time ~10min, but it was not using GPU to its full limits,\nYou can see the logs [here](https://wandb.ai/somusan/PogChamp2%20Baseline/runs/1tub37h5/system?workspace=user-somusan)\nIn the data loader I was using `pin_memory=True,num_workers  = 4` maybe I need to change that. I will also try using small image size.",
    "1730709": "_    |     _\n:-------------------------:|:-------------------------:\n<img width=\"120\" src=\"https://i.imgur.com/3jWSz9C.png\">  |  <img width=\"120\" src=\"https://i.imgur.com/CMzB0gA.png\">\n\n\n---\n\nI changed the small model with `resnet50` and now it seems like its utilizing the GPU way better. These small peaks near 20 are actually near 100% in the kaggle GPU usage metric. And I believe It will reduce even more if I use saved raw spec images directly in the pipeline.\n<img width=\"600\" src=\"https://i.imgur.com/0qB0564.png\">\nHad to reduce the batch size bcoz of OOM, and including training and validation its taking a total of 40 min  [32min/8min] per epoch [its an 80/20 split]. Im not sure if anyone is having the same time per epoch. please let me know if anyone is getting the same ETA. Thanks, @pheadrus and @harveenchadha for helping.\nWandB: https://wandb.ai/somusan/pogchamp2/runs/3j7yisba/system?workspace=user-somusan",
    "1730829": "Great stuff @soumya9977 . You can also do gradient accumulation if you BS is too low (eg 8-16 ish).",
    "1730876": "I'm sorry for this silly question, but what do you mean by `gradient accumulation if your BS is too low`? is it like accumulating gradients from all small mini-batches and updating the params after whole epoch?",
    "1731628": "You can check this [blog](https://kozodoi.me/python/deep%20learning/pytorch/tutorial/2021/02/19/gradient-accumulation.html) to learn about gradient accumulation",
    "1731707": "thank you so much for sharing @kpriyanshu256, I will definitely try that. seems like the core idea is almost the same.",
    "1731989": "Hi @soumya9977,\n\naccessing a Pandas Dateframe in __getitem__ is slow. I always convert it into a numpy or TorchTensor in __init__. You can also do the path transformation before.",
    "1732156": "That is a nice point, I generally use pandas df in getitem but from now on I will try to convert it into a numpy or TorchTensor in init and then use that in getitem. Thank you so muchhh."
  },
  "source": "meta"
}