{
  "id": 438746,
  "title": "Question about GPU memory allocation",
  "url": "/competitions/bengaliai-speech/discussion/438746",
  "author_name": "",
  "post_date": "2023-09-12T13:09:44.665766800Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have trained at batch_size = 4, and number of epoch set to 15. If all the input was padded and presumably has a similar size, why GPU memory allocated increases by time?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F11ce23811f2f788197011ada0e16b1cb%2FScreenshot%202023-09-12%20at%209.08.59%20PM.png?generation=1694524177028073&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2434691",
      "postDate": "09/12/2023 13:09:44",
      "content": "<p>I have trained at batch_size = 4, and number of epoch set to 15. If all the input was padded and presumably has a similar size, why GPU memory allocated increases by time?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F11ce23811f2f788197011ada0e16b1cb%2FScreenshot%202023-09-12%20at%209.08.59%20PM.png?generation=1694524177028073&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I have trained at batch_size = 4, and number of epoch set to 15. If all the input was padded and presumably has a similar size, why GPU memory allocated increases by time?![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F11ce23811f2f788197011ada0e16b1cb%2FScreenshot%202023-09-12%20at%209.08.59%20PM.png?generation=1694524177028073&alt=media)",
      "votes": null
    },
    {
      "id": "2434693",
      "postDate": "09/12/2023 13:11:10",
      "content": "<p>Also, how to explain the plateau? My GPU is 24 GiB btw</p>",
      "rawMarkdown": "Also, how to explain the plateau? My GPU is 24 GiB btw",
      "votes": null
    },
    {
      "id": "2448536",
      "postDate": "09/20/2023 16:43:33",
      "content": "<p>Not sure, but I've been getting some out of memory errors on my end as well (I'm just using the P100 here on kaggle for any training at the moment). What's your dataset or dataloader look like? My suspicion is that there is some pernicious memroy leak happening with torch, but I'm not nearly skilled enough to understand completely.</p>\n<p>Maybe you could try </p>\n<pre><code>torch.cuda.empty_cache() \ngc.collect()\n</code></pre>\n<p>after every batch. That's what I've seen recommended online, but I'm still getting those errors even when using that.</p>\n<p>For context, I'm limiting myself to a batch size of only 2, and I can train for around 3000 steps before the out of memory error happens.</p>",
      "rawMarkdown": "Not sure, but I've been getting some out of memory errors on my end as well (I'm just using the P100 here on kaggle for any training at the moment). What's your dataset or dataloader look like? My suspicion is that there is some pernicious memroy leak happening with torch, but I'm not nearly skilled enough to understand completely.\n\nMaybe you could try \n```python\ntorch.cuda.empty_cache() \ngc.collect()\n```\nafter every batch. That's what I've seen recommended online, but I'm still getting those errors even when using that.\n\nFor context, I'm limiting myself to a batch size of only 2, and I can train for around 3000 steps before the out of memory error happens.",
      "votes": null
    },
    {
      "id": "2448965",
      "postDate": "09/20/2023 22:33:54",
      "content": "<p>What I have found is that after using huggingface trainer in <a href=\"https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training\" target=\"_blank\">this notebook</a>, the nemory would increased throughout the training session. I guess it’s because it was optimized to save memory in comparison to manual implementation </p>",
      "rawMarkdown": "What I have found is that after using huggingface trainer in [this notebook](https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training), the nemory would increased throughout the training session. I guess it’s because it was optimized to save memory in comparison to manual implementation",
      "votes": null
    },
    {
      "id": "2449003",
      "postDate": "09/20/2023 23:42:12",
      "content": "<p>I don't know your collator implementation, but most likely you are padding inputs to the maximum length of each batch. If the maximum length of each batch increase as time goes, the memory usage also goes up </p>",
      "rawMarkdown": "I don't know your collator implementation, but most likely you are padding inputs to the maximum length of each batch. If the maximum length of each batch increase as time goes, the memory usage also goes up",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2434693,
      "author_name": "renyiwei",
      "author_url": "",
      "post_date": "09/12/2023 13:11:10",
      "content": "<p>Also, how to explain the plateau? My GPU is 24 GiB btw</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2448536,
      "author_name": "msthil",
      "author_url": "",
      "post_date": "09/20/2023 16:43:33",
      "content": "<p>Not sure, but I've been getting some out of memory errors on my end as well (I'm just using the P100 here on kaggle for any training at the moment). What's your dataset or dataloader look like? My suspicion is that there is some pernicious memroy leak happening with torch, but I'm not nearly skilled enough to understand completely.</p>\n<p>Maybe you could try </p>\n<pre><code>torch.cuda.empty_cache() \ngc.collect()\n</code></pre>\n<p>after every batch. That's what I've seen recommended online, but I'm still getting those errors even when using that.</p>\n<p>For context, I'm limiting myself to a batch size of only 2, and I can train for around 3000 steps before the out of memory error happens.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2448965,
          "author_name": "renyiwei",
          "author_url": "",
          "post_date": "09/20/2023 22:33:54",
          "content": "<p>What I have found is that after using huggingface trainer in <a href=\"https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training\" target=\"_blank\">this notebook</a>, the nemory would increased throughout the training session. I guess it’s because it was optimized to save memory in comparison to manual implementation </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2449003,
      "author_name": "ludditep",
      "author_url": "",
      "post_date": "09/20/2023 23:42:12",
      "content": "<p>I don't know your collator implementation, but most likely you are padding inputs to the maximum length of each batch. If the maximum length of each batch increase as time goes, the memory usage also goes up </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2434691": "I have trained at batch_size = 4, and number of epoch set to 15. If all the input was padded and presumably has a similar size, why GPU memory allocated increases by time?![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F11ce23811f2f788197011ada0e16b1cb%2FScreenshot%202023-09-12%20at%209.08.59%20PM.png?generation=1694524177028073&alt=media)",
    "2434693": "Also, how to explain the plateau? My GPU is 24 GiB btw",
    "2448536": "Not sure, but I've been getting some out of memory errors on my end as well (I'm just using the P100 here on kaggle for any training at the moment). What's your dataset or dataloader look like? My suspicion is that there is some pernicious memroy leak happening with torch, but I'm not nearly skilled enough to understand completely.\n\nMaybe you could try \n```python\ntorch.cuda.empty_cache() \ngc.collect()\n```\nafter every batch. That's what I've seen recommended online, but I'm still getting those errors even when using that.\n\nFor context, I'm limiting myself to a batch size of only 2, and I can train for around 3000 steps before the out of memory error happens.",
    "2448965": "What I have found is that after using huggingface trainer in [this notebook](https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training), the nemory would increased throughout the training session. I guess it’s because it was optimized to save memory in comparison to manual implementation",
    "2449003": "I don't know your collator implementation, but most likely you are padding inputs to the maximum length of each batch. If the maximum length of each batch increase as time goes, the memory usage also goes up"
  },
  "source": "meta"
}