{
  "id": 500877,
  "title": "low memory pytorch training ... remove dataloader and use in memory data + direct index",
  "url": "/competitions/leash-BELKA/discussion/500877",
  "author_name": "",
  "post_date": "2024-05-07T08:55:52.217841500Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>for detail, please refer to<br>\n<a href=\"https://github.com/pytorch/pytorch/issues/20433\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/20433</a><br>\n<a href=\"https://discuss.pytorch.org/t/using-dataloader-when-all-data-is-stored-in-memory/129074\" target=\"_blank\">https://discuss.pytorch.org/t/using-dataloader-when-all-data-is-stored-in-memory/129074</a></p>\n<p>assume that you read all data and store in array</p>\n<pre><code>SMILES,SMILES_MASK,ALLFP,BIND = load\n(about GB)\n</code></pre>\n<p>if you use the nomal dataloader, you will get this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2ba2c8a3d4617db39cb81ebced2425b7%2FSelection_071.png?generation=1715072061898663&amp;alt=media\"></p>\n<p>this is due to multiporcessing lib of python.<br>\n(memory is increasing when dataloader has num worker &gt;0. It put get item function of dataset to queue during training for each epoch. it is reset only after the dataloader ends its iteration)</p>\n<p>note that memory usage of GPU is only 6.5 GB for the example shown above</p>\n<hr>\n<p>if anyone knows how to clear pytorch dataloader queue, let me know!</p>",
  "messages": [
    {
      "id": "2798505",
      "postDate": "05/07/2024 08:55:52",
      "content": "<p>for detail, please refer to<br>\n<a href=\"https://github.com/pytorch/pytorch/issues/20433\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/20433</a><br>\n<a href=\"https://discuss.pytorch.org/t/using-dataloader-when-all-data-is-stored-in-memory/129074\" target=\"_blank\">https://discuss.pytorch.org/t/using-dataloader-when-all-data-is-stored-in-memory/129074</a></p>\n<p>assume that you read all data and store in array</p>\n<pre><code>SMILES,SMILES_MASK,ALLFP,BIND = load\n(about GB)\n</code></pre>\n<p>if you use the nomal dataloader, you will get this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2ba2c8a3d4617db39cb81ebced2425b7%2FSelection_071.png?generation=1715072061898663&amp;alt=media\"></p>\n<p>this is due to multiporcessing lib of python.<br>\n(memory is increasing when dataloader has num worker &gt;0. It put get item function of dataset to queue during training for each epoch. it is reset only after the dataloader ends its iteration)</p>\n<p>note that memory usage of GPU is only 6.5 GB for the example shown above</p>\n<hr>\n<p>if anyone knows how to clear pytorch dataloader queue, let me know!</p>",
      "rawMarkdown": "for detail, please refer to\nhttps://github.com/pytorch/pytorch/issues/20433\nhttps://discuss.pytorch.org/t/using-dataloader-when-all-data-is-stored-in-memory/129074\n\nassume that you read all data and store in array\n```\nSMILES,SMILES_MASK,ALLFP,BIND = load_data(mode='train')\n(about 90GB)\n```\n\nif you use the nomal dataloader, you will get this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2ba2c8a3d4617db39cb81ebced2425b7%2FSelection_071.png?generation=1715072061898663&alt=media)\n\nthis is due to multiporcessing lib of python.\n(memory is increasing when dataloader has num worker >0. It put get item function of dataset to queue during training for each epoch. it is reset only after the dataloader ends its iteration)\n\nnote that memory usage of GPU is only 6.5 GB for the example shown above\n\n---\n\nif anyone knows how to clear pytorch dataloader queue, let me know!",
      "votes": null
    },
    {
      "id": "2798508",
      "postDate": "05/07/2024 08:57:45",
      "content": "<p>solution:<br>\njust use direct indexing</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2deabc8ad0c8333104309b95d6feef98%2FSelection_073.png?generation=1715072224684708&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F01e90f17c2677a3f22a66c0800e62e47%2FSelection_075.png?generation=1715072236024984&amp;alt=media\"></p>\n<p>results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F780bac96ac4d5bcc8febbd864f255754%2FSelection_072.png?generation=1715072262793541&amp;alt=media\"></p>",
      "rawMarkdown": "solution:\njust use direct indexing\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2deabc8ad0c8333104309b95d6feef98%2FSelection_073.png?generation=1715072224684708&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F01e90f17c2677a3f22a66c0800e62e47%2FSelection_075.png?generation=1715072236024984&alt=media)\n\nresults:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F780bac96ac4d5bcc8febbd864f255754%2FSelection_072.png?generation=1715072262793541&alt=media)",
      "votes": null
    },
    {
      "id": "2808271",
      "postDate": "05/12/2024 05:44:22",
      "content": "<p>I have the same issue on LEAP competition and it looks like there is no clear solution.</p>",
      "rawMarkdown": "I have the same issue on LEAP competition and it looks like there is no clear solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2798508,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/07/2024 08:57:45",
      "content": "<p>solution:<br>\njust use direct indexing</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2deabc8ad0c8333104309b95d6feef98%2FSelection_073.png?generation=1715072224684708&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F01e90f17c2677a3f22a66c0800e62e47%2FSelection_075.png?generation=1715072236024984&amp;alt=media\"></p>\n<p>results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F780bac96ac4d5bcc8febbd864f255754%2FSelection_072.png?generation=1715072262793541&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2808271,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "05/12/2024 05:44:22",
      "content": "<p>I have the same issue on LEAP competition and it looks like there is no clear solution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2798505": "for detail, please refer to\nhttps://github.com/pytorch/pytorch/issues/20433\nhttps://discuss.pytorch.org/t/using-dataloader-when-all-data-is-stored-in-memory/129074\n\nassume that you read all data and store in array\n```\nSMILES,SMILES_MASK,ALLFP,BIND = load_data(mode='train')\n(about 90GB)\n```\n\nif you use the nomal dataloader, you will get this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2ba2c8a3d4617db39cb81ebced2425b7%2FSelection_071.png?generation=1715072061898663&alt=media)\n\nthis is due to multiporcessing lib of python.\n(memory is increasing when dataloader has num worker >0. It put get item function of dataset to queue during training for each epoch. it is reset only after the dataloader ends its iteration)\n\nnote that memory usage of GPU is only 6.5 GB for the example shown above\n\n---\n\nif anyone knows how to clear pytorch dataloader queue, let me know!",
    "2798508": "solution:\njust use direct indexing\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2deabc8ad0c8333104309b95d6feef98%2FSelection_073.png?generation=1715072224684708&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F01e90f17c2677a3f22a66c0800e62e47%2FSelection_075.png?generation=1715072236024984&alt=media)\n\nresults:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F780bac96ac4d5bcc8febbd864f255754%2FSelection_072.png?generation=1715072262793541&alt=media)",
    "2808271": "I have the same issue on LEAP competition and it looks like there is no clear solution."
  },
  "source": "meta"
}