{
  "id": 394596,
  "title": "Chunk based data loading with caching",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/394596",
  "author_name": "Iafoss",
  "post_date": "2023-03-14T05:17:36.301000",
  "votes": 31,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Dealing with large chunked data is one of the challenges of this competition, which may impose an additional barrier for people willing to join. Since the competition was not very active recently, I decided to share quite a simple way of dealing with the provided data without the time-consuming conversion of data into other formats or trying to load everything into the RAM, which is quite challenging, especially at Kaggle.</p>\n<p>The chunked data format is not very friendly to Pytorch, and some additional tricks are needed. My approach consists of two things: (1) A <strong>dataloader caching the considered chunks</strong> and (2) <strong>Random Chunked Sampler</strong>, which randomly selects a chunk and then goes through all ids in the chunk before selecting a new one. With this sampler, the caching pipeline spends time reading the corresponding chunk only at the first request, and the following requests are finished quickly with reading data from RAM. Meanwhile, when the cache is full, the earliest record is removed, which limits RAM usage. At inference and test, when reading is sequential, no extra modifications are needed to the data sampler. It would be great if the Pytorch team implemented native support of chunked data.</p>\n<p>The notebook is available -&gt; <a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": 2180791,
      "postDate": "2023-03-14T05:17:36.303Z",
      "content": "<p>Dealing with large chunked data is one of the challenges of this competition, which may impose an additional barrier for people willing to join. Since the competition was not very active recently, I decided to share quite a simple way of dealing with the provided data without the time-consuming conversion of data into other formats or trying to load everything into the RAM, which is quite challenging, especially at Kaggle.</p>\n<p>The chunked data format is not very friendly to Pytorch, and some additional tricks are needed. My approach consists of two things: (1) A <strong>dataloader caching the considered chunks</strong> and (2) <strong>Random Chunked Sampler</strong>, which randomly selects a chunk and then goes through all ids in the chunk before selecting a new one. With this sampler, the caching pipeline spends time reading the corresponding chunk only at the first request, and the following requests are finished quickly with reading data from RAM. Meanwhile, when the cache is full, the earliest record is removed, which limits RAM usage. At inference and test, when reading is sequential, no extra modifications are needed to the data sampler. It would be great if the Pytorch team implemented native support of chunked data.</p>\n<p>The notebook is available -&gt; <a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Dealing with large chunked data is one of the challenges of this competition, which may impose an additional barrier for people willing to join. Since the competition was not very active recently, I decided to share quite a simple way of dealing with the provided data without the time-consuming conversion of data into other formats or trying to load everything into the RAM, which is quite challenging, especially at Kaggle.\n\nThe chunked data format is not very friendly to Pytorch, and some additional tricks are needed. My approach consists of two things: (1) A **dataloader caching the considered chunks** and (2) **Random Chunked Sampler**, which randomly selects a chunk and then goes through all ids in the chunk before selecting a new one. With this sampler, the caching pipeline spends time reading the corresponding chunk only at the first request, and the following requests are finished quickly with reading data from RAM. Meanwhile, when the cache is full, the earliest record is removed, which limits RAM usage. At inference and test, when reading is sequential, no extra modifications are needed to the data sampler. It would be great if the Pytorch team implemented native support of chunked data.\n\nThe notebook is available -> [here](https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching).",
      "votes": 31
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2180791": "Dealing with large chunked data is one of the challenges of this competition, which may impose an additional barrier for people willing to join. Since the competition was not very active recently, I decided to share quite a simple way of dealing with the provided data without the time-consuming conversion of data into other formats or trying to load everything into the RAM, which is quite challenging, especially at Kaggle.\n\nThe chunked data format is not very friendly to Pytorch, and some additional tricks are needed. My approach consists of two things: (1) A **dataloader caching the considered chunks** and (2) **Random Chunked Sampler**, which randomly selects a chunk and then goes through all ids in the chunk before selecting a new one. With this sampler, the caching pipeline spends time reading the corresponding chunk only at the first request, and the following requests are finished quickly with reading data from RAM. Meanwhile, when the cache is full, the earliest record is removed, which limits RAM usage. At inference and test, when reading is sequential, no extra modifications are needed to the data sampler. It would be great if the Pytorch team implemented native support of chunked data.\n\nThe notebook is available -> [here](https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching)."
  }
}