{
  "id": 358557,
  "title": "Low memory data loading with Pytorch Dataset class",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/358557",
  "author_name": "ehekatlact",
  "post_date": "2022-10-08T13:08:44.796000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Pytorch has a class called Dataset that reads data with low memory.<br>\nIf you have participated in image-related competitions, you may have seen this function before.</p>\n<p>Unfortunately, due to my poor English, I could not find an example of this applied to csv data.<br>\nTherefore, the following is a notebook I implemented while reading the manual.</p>\n<p><a href=\"https://www.kaggle.com/code/ehekatlact/tps2210-pytorch-lightning-with-iterabledataset\" target=\"_blank\">https://www.kaggle.com/code/ehekatlact/tps2210-pytorch-lightning-with-iterabledataset</a></p>\n<p>By taking the time to do the appropriate preprocessing, it apparently reads all 10 csv records in about 1 minute. This is fast enough. Because with mini-batch learning, it only takes a couple of epochs to finish learning.</p>\n<ul>\n<li>For the Beginner, I hope this helps!</li>\n<li>For the more seasoned, if you know of a better way, please let me know! I am very interested in this.</li>\n</ul>\n<p>PS.</p>\n<p>The Normal Dataset class was not using cache memory well and IO was very slow. I am hoping that using IterableDataset for sequential access will make it faster. I will go check the source code when I have time.</p>",
  "messages": [
    {
      "id": 1978048,
      "postDate": "2022-10-08T13:08:44.797Z",
      "content": "<p>Pytorch has a class called Dataset that reads data with low memory.<br>\nIf you have participated in image-related competitions, you may have seen this function before.</p>\n<p>Unfortunately, due to my poor English, I could not find an example of this applied to csv data.<br>\nTherefore, the following is a notebook I implemented while reading the manual.</p>\n<p><a href=\"https://www.kaggle.com/code/ehekatlact/tps2210-pytorch-lightning-with-iterabledataset\" target=\"_blank\">https://www.kaggle.com/code/ehekatlact/tps2210-pytorch-lightning-with-iterabledataset</a></p>\n<p>By taking the time to do the appropriate preprocessing, it apparently reads all 10 csv records in about 1 minute. This is fast enough. Because with mini-batch learning, it only takes a couple of epochs to finish learning.</p>\n<ul>\n<li>For the Beginner, I hope this helps!</li>\n<li>For the more seasoned, if you know of a better way, please let me know! I am very interested in this.</li>\n</ul>\n<p>PS.</p>\n<p>The Normal Dataset class was not using cache memory well and IO was very slow. I am hoping that using IterableDataset for sequential access will make it faster. I will go check the source code when I have time.</p>",
      "rawMarkdown": "Pytorch has a class called Dataset that reads data with low memory.\nIf you have participated in image-related competitions, you may have seen this function before.\n\nUnfortunately, due to my poor English, I could not find an example of this applied to csv data.\nTherefore, the following is a notebook I implemented while reading the manual.\n\nhttps://www.kaggle.com/code/ehekatlact/tps2210-pytorch-lightning-with-iterabledataset\n\nBy taking the time to do the appropriate preprocessing, it apparently reads all 10 csv records in about 1 minute. This is fast enough. Because with mini-batch learning, it only takes a couple of epochs to finish learning.\n\n* For the Beginner, I hope this helps!\n* For the more seasoned, if you know of a better way, please let me know! I am very interested in this.\n\nPS.\n\nThe Normal Dataset class was not using cache memory well and IO was very slow. I am hoping that using IterableDataset for sequential access will make it faster. I will go check the source code when I have time.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1978048": "Pytorch has a class called Dataset that reads data with low memory.\nIf you have participated in image-related competitions, you may have seen this function before.\n\nUnfortunately, due to my poor English, I could not find an example of this applied to csv data.\nTherefore, the following is a notebook I implemented while reading the manual.\n\nhttps://www.kaggle.com/code/ehekatlact/tps2210-pytorch-lightning-with-iterabledataset\n\nBy taking the time to do the appropriate preprocessing, it apparently reads all 10 csv records in about 1 minute. This is fast enough. Because with mini-batch learning, it only takes a couple of epochs to finish learning.\n\n* For the Beginner, I hope this helps!\n* For the more seasoned, if you know of a better way, please let me know! I am very interested in this.\n\nPS.\n\nThe Normal Dataset class was not using cache memory well and IO was very slow. I am hoping that using IterableDataset for sequential access will make it faster. I will go check the source code when I have time."
  }
}