{
  "id": 198106,
  "title": "The dataloader - .zarr cache issue",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/198106",
  "author_name": "",
  "post_date": "2020-11-19T20:12:46.205200900Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>It is a bit late for this post, but maybe it will be useful anyway.</p>\n<h2>Problem</h2>\n<p>I noticed a few days ago that something weird is going on with the data loading (train_full). The speed varies a lot. I looked into it, and I think I have an answer (and a solution). At least it seems stable now.</p>\n<p>The <code>l5kit</code> and the <code>.zarr</code> package use an internal caching mechanism. Zarr handles the data in chunks, and it stores these chunks in an in-memory cache. Of course, the <code>train_full</code> is not fit 100% in memory, so (I assume) it keeps those chunks in memory that have higher hit rates.</p>\n<p>The problem occurs when we use the default <code>RandomSampler</code> (shuffle=True). The hit rate is low because of the size of the dataset. The default size of the cache is 1Gb (per pytorch data worker!). It fills up quickly, and from that point, it has to load/replace the cached chunks very often. (I haven't checked the actual numbers)</p>\n<h2>Solution</h2>\n<p>One solution if the sampler asks for samples that are already cached. I divided the <code>train_full</code> into blocks (10M samples) and force the sampler to return samples from this block only. It moves on to the next block if the first one has no more unused data. <br>\n(The order of the blocks and the samples in a block are randomized.)</p>\n<p>It will overfit a bit, as you can see in the image below. (At the spikes is where the loader/sampler switches block) This is because of the similar images in a block.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F130d13ea2da052096b7176aa879155c4%2Fsampler1.png?generation=1605815584463283&amp;alt=media\" alt=\"\"></p>\n<p>The sampler will use only 1/4 of the samples in a block in every epoch to fix this. For example, in epoch 1: block0 = [0, 4, 8, …], epoch 2, block0 = [1, 5, 9, …]. In one epoch it will use all of the blocks (1/4 of the train_full dataset)<br>\n(Samples and the order of the blocks are randomized)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fcde023f7684a8d984fa84d9b2810ca70%2Fsampler2.png?generation=1605815689825469&amp;alt=media\" alt=\"\"></p>\n<h2>Code</h2>\n<pre><code>class LyftZarrCacheFixedSampler(Sampler):\n\n    def __init__(self, datasource: LyftGPUDataset, chunk_size=1000000):\n        super(LyftZarrCacheFixedSampler, self).__init__(datasource)\n\n        self.chunk_size = chunk_size\n        self.datasource = datasource\n        self.datasource.agents_indices.sort(kind='stable')\n\n        self.epoch = 0\n        self.n_chunks = len(self.datasource.agents_indices) // chunk_size\n        self.n_last_chunk = len(self.datasource.agents_indices) % chunk_size\n\n        if self.n_last_chunk &gt; 0:\n            self.n_chunks += 1\n\n    def __len__(self) -&gt; int:\n        return len(self.datasource.agents_indices) // 4\n\n    def __iter__(self):\n        indices = np.array([x for x in range(len(self.datasource.agents_indices))])\n\n        res_idx = []\n\n        for chunk in range(self.n_chunks):\n            from_idx = self.chunk_size * chunk\n            to_idx = min([self.chunk_size * (chunk + 1), len(indices)])\n\n            x = indices[from_idx:to_idx]\n            x = x[self.epoch::4]\n\n            np.random.shuffle(x)\n            res_idx.append(x)\n\n        self.epoch += 1\n        if self.epoch == 4:\n            self.epoch = 0\n\n        random.shuffle(res_idx)\n        indices = np.hstack(res_idx)\n\n        return iter(indices)\n\n\n...\ndataset = AgentDataset(...)\nsampler = LyftZarrCacheFixedSampler(dataset, chunk_size=...)\ndataloader = DataLoader(dataset, sampler=sampler, ...)\n....\n</code></pre>\n<h2>Benchmark</h2>\n<p>More details, benchmark (train.zarr) <a href=\"https://www.kaggle.com/pestipeti/lyft-zarr-cache-data-sampler-benchmark\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "1084180",
      "postDate": "11/19/2020 20:12:46",
      "content": "<p>It is a bit late for this post, but maybe it will be useful anyway.</p>\n<h2>Problem</h2>\n<p>I noticed a few days ago that something weird is going on with the data loading (train_full). The speed varies a lot. I looked into it, and I think I have an answer (and a solution). At least it seems stable now.</p>\n<p>The <code>l5kit</code> and the <code>.zarr</code> package use an internal caching mechanism. Zarr handles the data in chunks, and it stores these chunks in an in-memory cache. Of course, the <code>train_full</code> is not fit 100% in memory, so (I assume) it keeps those chunks in memory that have higher hit rates.</p>\n<p>The problem occurs when we use the default <code>RandomSampler</code> (shuffle=True). The hit rate is low because of the size of the dataset. The default size of the cache is 1Gb (per pytorch data worker!). It fills up quickly, and from that point, it has to load/replace the cached chunks very often. (I haven't checked the actual numbers)</p>\n<h2>Solution</h2>\n<p>One solution if the sampler asks for samples that are already cached. I divided the <code>train_full</code> into blocks (10M samples) and force the sampler to return samples from this block only. It moves on to the next block if the first one has no more unused data. <br>\n(The order of the blocks and the samples in a block are randomized.)</p>\n<p>It will overfit a bit, as you can see in the image below. (At the spikes is where the loader/sampler switches block) This is because of the similar images in a block.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F130d13ea2da052096b7176aa879155c4%2Fsampler1.png?generation=1605815584463283&amp;alt=media\" alt=\"\"></p>\n<p>The sampler will use only 1/4 of the samples in a block in every epoch to fix this. For example, in epoch 1: block0 = [0, 4, 8, …], epoch 2, block0 = [1, 5, 9, …]. In one epoch it will use all of the blocks (1/4 of the train_full dataset)<br>\n(Samples and the order of the blocks are randomized)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fcde023f7684a8d984fa84d9b2810ca70%2Fsampler2.png?generation=1605815689825469&amp;alt=media\" alt=\"\"></p>\n<h2>Code</h2>\n<pre><code>class LyftZarrCacheFixedSampler(Sampler):\n\n    def __init__(self, datasource: LyftGPUDataset, chunk_size=1000000):\n        super(LyftZarrCacheFixedSampler, self).__init__(datasource)\n\n        self.chunk_size = chunk_size\n        self.datasource = datasource\n        self.datasource.agents_indices.sort(kind='stable')\n\n        self.epoch = 0\n        self.n_chunks = len(self.datasource.agents_indices) // chunk_size\n        self.n_last_chunk = len(self.datasource.agents_indices) % chunk_size\n\n        if self.n_last_chunk &gt; 0:\n            self.n_chunks += 1\n\n    def __len__(self) -&gt; int:\n        return len(self.datasource.agents_indices) // 4\n\n    def __iter__(self):\n        indices = np.array([x for x in range(len(self.datasource.agents_indices))])\n\n        res_idx = []\n\n        for chunk in range(self.n_chunks):\n            from_idx = self.chunk_size * chunk\n            to_idx = min([self.chunk_size * (chunk + 1), len(indices)])\n\n            x = indices[from_idx:to_idx]\n            x = x[self.epoch::4]\n\n            np.random.shuffle(x)\n            res_idx.append(x)\n\n        self.epoch += 1\n        if self.epoch == 4:\n            self.epoch = 0\n\n        random.shuffle(res_idx)\n        indices = np.hstack(res_idx)\n\n        return iter(indices)\n\n\n...\ndataset = AgentDataset(...)\nsampler = LyftZarrCacheFixedSampler(dataset, chunk_size=...)\ndataloader = DataLoader(dataset, sampler=sampler, ...)\n....\n</code></pre>\n<h2>Benchmark</h2>\n<p>More details, benchmark (train.zarr) <a href=\"https://www.kaggle.com/pestipeti/lyft-zarr-cache-data-sampler-benchmark\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "It is a bit late for this post, but maybe it will be useful anyway.\n\n## Problem\n\nI noticed a few days ago that something weird is going on with the data loading (train_full). The speed varies a lot. I looked into it, and I think I have an answer (and a solution). At least it seems stable now.\n\nThe `l5kit` and the `.zarr` package use an internal caching mechanism. Zarr handles the data in chunks, and it stores these chunks in an in-memory cache. Of course, the `train_full` is not fit 100% in memory, so (I assume) it keeps those chunks in memory that have higher hit rates.\n\nThe problem occurs when we use the default `RandomSampler` (shuffle=True). The hit rate is low because of the size of the dataset. The default size of the cache is 1Gb (per pytorch data worker!). It fills up quickly, and from that point, it has to load/replace the cached chunks very often. (I haven't checked the actual numbers)\n\n## Solution\n\nOne solution if the sampler asks for samples that are already cached. I divided the `train_full` into blocks (10M samples) and force the sampler to return samples from this block only. It moves on to the next block if the first one has no more unused data. \n(The order of the blocks and the samples in a block are randomized.)\n\nIt will overfit a bit, as you can see in the image below. (At the spikes is where the loader/sampler switches block) This is because of the similar images in a block.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F130d13ea2da052096b7176aa879155c4%2Fsampler1.png?generation=1605815584463283&alt=media)\n\nThe sampler will use only 1/4 of the samples in a block in every epoch to fix this. For example, in epoch 1: block0 = [0, 4, 8, ...], epoch 2, block0 = [1, 5, 9, ...]. In one epoch it will use all of the blocks (1/4 of the train_full dataset)\n(Samples and the order of the blocks are randomized)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fcde023f7684a8d984fa84d9b2810ca70%2Fsampler2.png?generation=1605815689825469&alt=media)\n\n## Code\n```\nclass LyftZarrCacheFixedSampler(Sampler):\n\n    def __init__(self, datasource: LyftGPUDataset, chunk_size=1000000):\n        super(LyftZarrCacheFixedSampler, self).__init__(datasource)\n\n        self.chunk_size = chunk_size\n        self.datasource = datasource\n        self.datasource.agents_indices.sort(kind='stable')\n\n        self.epoch = 0\n        self.n_chunks = len(self.datasource.agents_indices) // chunk_size\n        self.n_last_chunk = len(self.datasource.agents_indices) % chunk_size\n\n        if self.n_last_chunk > 0:\n            self.n_chunks += 1\n\n    def __len__(self) -> int:\n        return len(self.datasource.agents_indices) // 4\n\n    def __iter__(self):\n        indices = np.array([x for x in range(len(self.datasource.agents_indices))])\n\n        res_idx = []\n\n        for chunk in range(self.n_chunks):\n            from_idx = self.chunk_size * chunk\n            to_idx = min([self.chunk_size * (chunk + 1), len(indices)])\n\n            x = indices[from_idx:to_idx]\n            x = x[self.epoch::4]\n\n            np.random.shuffle(x)\n            res_idx.append(x)\n\n        self.epoch += 1\n        if self.epoch == 4:\n            self.epoch = 0\n\n        random.shuffle(res_idx)\n        indices = np.hstack(res_idx)\n\n        return iter(indices)\n\n\n...\ndataset = AgentDataset(...)\nsampler = LyftZarrCacheFixedSampler(dataset, chunk_size=...)\ndataloader = DataLoader(dataset, sampler=sampler, ...)\n....\n\n```\n\n## Benchmark\nMore details, benchmark (train.zarr) [here](https://www.kaggle.com/pestipeti/lyft-zarr-cache-data-sampler-benchmark)",
      "votes": null
    },
    {
      "id": "1084341",
      "postDate": "11/20/2020 00:24:55",
      "content": "<p>What about old style ramdisk solution <code>sudo mount -t tmpfs -o rw,size=90G tmpfs /mnt/ramdisk</code> :)</p>",
      "rawMarkdown": "What about old style ramdisk solution `sudo mount -t tmpfs -o rw,size=90G tmpfs /mnt/ramdisk` :)",
      "votes": null
    },
    {
      "id": "1087154",
      "postDate": "11/22/2020 12:13:33",
      "content": "<p>Hi thanks for sharing! When I implemented the code I got the error Sampler is not defined?</p>",
      "rawMarkdown": "Hi thanks for sharing! When I implemented the code I got the error Sampler is not defined?",
      "votes": null
    },
    {
      "id": "1087186",
      "postDate": "11/22/2020 12:59:56",
      "content": "<p>you should import <code>Sampler</code> from torch:</p>\n<pre><code>from torch.utils.data.sampler import Sampler\n</code></pre>",
      "rawMarkdown": "you should import `Sampler` from torch:\n\n```\nfrom torch.utils.data.sampler import Sampler\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1084341,
      "author_name": "ggrizzly",
      "author_url": "",
      "post_date": "11/20/2020 00:24:55",
      "content": "<p>What about old style ramdisk solution <code>sudo mount -t tmpfs -o rw,size=90G tmpfs /mnt/ramdisk</code> :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1087154,
      "author_name": "huanvo",
      "author_url": "",
      "post_date": "11/22/2020 12:13:33",
      "content": "<p>Hi thanks for sharing! When I implemented the code I got the error Sampler is not defined?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1087186,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/22/2020 12:59:56",
          "content": "<p>you should import <code>Sampler</code> from torch:</p>\n<pre><code>from torch.utils.data.sampler import Sampler\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1084180": "It is a bit late for this post, but maybe it will be useful anyway.\n\n## Problem\n\nI noticed a few days ago that something weird is going on with the data loading (train_full). The speed varies a lot. I looked into it, and I think I have an answer (and a solution). At least it seems stable now.\n\nThe `l5kit` and the `.zarr` package use an internal caching mechanism. Zarr handles the data in chunks, and it stores these chunks in an in-memory cache. Of course, the `train_full` is not fit 100% in memory, so (I assume) it keeps those chunks in memory that have higher hit rates.\n\nThe problem occurs when we use the default `RandomSampler` (shuffle=True). The hit rate is low because of the size of the dataset. The default size of the cache is 1Gb (per pytorch data worker!). It fills up quickly, and from that point, it has to load/replace the cached chunks very often. (I haven't checked the actual numbers)\n\n## Solution\n\nOne solution if the sampler asks for samples that are already cached. I divided the `train_full` into blocks (10M samples) and force the sampler to return samples from this block only. It moves on to the next block if the first one has no more unused data. \n(The order of the blocks and the samples in a block are randomized.)\n\nIt will overfit a bit, as you can see in the image below. (At the spikes is where the loader/sampler switches block) This is because of the similar images in a block.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2F130d13ea2da052096b7176aa879155c4%2Fsampler1.png?generation=1605815584463283&alt=media)\n\nThe sampler will use only 1/4 of the samples in a block in every epoch to fix this. For example, in epoch 1: block0 = [0, 4, 8, ...], epoch 2, block0 = [1, 5, 9, ...]. In one epoch it will use all of the blocks (1/4 of the train_full dataset)\n(Samples and the order of the blocks are randomized)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fcde023f7684a8d984fa84d9b2810ca70%2Fsampler2.png?generation=1605815689825469&alt=media)\n\n## Code\n```\nclass LyftZarrCacheFixedSampler(Sampler):\n\n    def __init__(self, datasource: LyftGPUDataset, chunk_size=1000000):\n        super(LyftZarrCacheFixedSampler, self).__init__(datasource)\n\n        self.chunk_size = chunk_size\n        self.datasource = datasource\n        self.datasource.agents_indices.sort(kind='stable')\n\n        self.epoch = 0\n        self.n_chunks = len(self.datasource.agents_indices) // chunk_size\n        self.n_last_chunk = len(self.datasource.agents_indices) % chunk_size\n\n        if self.n_last_chunk > 0:\n            self.n_chunks += 1\n\n    def __len__(self) -> int:\n        return len(self.datasource.agents_indices) // 4\n\n    def __iter__(self):\n        indices = np.array([x for x in range(len(self.datasource.agents_indices))])\n\n        res_idx = []\n\n        for chunk in range(self.n_chunks):\n            from_idx = self.chunk_size * chunk\n            to_idx = min([self.chunk_size * (chunk + 1), len(indices)])\n\n            x = indices[from_idx:to_idx]\n            x = x[self.epoch::4]\n\n            np.random.shuffle(x)\n            res_idx.append(x)\n\n        self.epoch += 1\n        if self.epoch == 4:\n            self.epoch = 0\n\n        random.shuffle(res_idx)\n        indices = np.hstack(res_idx)\n\n        return iter(indices)\n\n\n...\ndataset = AgentDataset(...)\nsampler = LyftZarrCacheFixedSampler(dataset, chunk_size=...)\ndataloader = DataLoader(dataset, sampler=sampler, ...)\n....\n\n```\n\n## Benchmark\nMore details, benchmark (train.zarr) [here](https://www.kaggle.com/pestipeti/lyft-zarr-cache-data-sampler-benchmark)",
    "1084341": "What about old style ramdisk solution `sudo mount -t tmpfs -o rw,size=90G tmpfs /mnt/ramdisk` :)",
    "1087154": "Hi thanks for sharing! When I implemented the code I got the error Sampler is not defined?",
    "1087186": "you should import `Sampler` from torch:\n\n```\nfrom torch.utils.data.sampler import Sampler\n```"
  },
  "source": "meta"
}