{
  "id": 182428,
  "title": "Running into a memory leak with the Agent Dataset",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/182428",
  "author_name": "ryches",
  "post_date": "2020-09-12T18:00:02.507000",
  "votes": 12,
  "comment_count": 29,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/ryches/memory-leak\" target=\"_blank\">https://www.kaggle.com/ryches/memory-leak</a></p>\n<p>I don't know if others have run into this but I have been fighting against a memory leak. I set up a minimal example to just iterate through the data and do nothing else. It isnt storing anything or accumulating gradients or anything, but memory usage seems to continue to drift overtime until eventually it kills the kernel. Have others run into this? Am I just doing something wrong?</p>",
  "messages": [
    {
      "id": 1008037,
      "postDate": "2020-09-12T18:00:02.507Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryches/memory-leak\" target=\"_blank\">https://www.kaggle.com/ryches/memory-leak</a></p>\n<p>I don't know if others have run into this but I have been fighting against a memory leak. I set up a minimal example to just iterate through the data and do nothing else. It isnt storing anything or accumulating gradients or anything, but memory usage seems to continue to drift overtime until eventually it kills the kernel. Have others run into this? Am I just doing something wrong?</p>",
      "rawMarkdown": "https://www.kaggle.com/ryches/memory-leak\n\nI don't know if others have run into this but I have been fighting against a memory leak. I set up a minimal example to just iterate through the data and do nothing else. It isnt storing anything or accumulating gradients or anything, but memory usage seems to continue to drift overtime until eventually it kills the kernel. Have others run into this? Am I just doing something wrong?",
      "votes": 12
    },
    {
      "id": 1056316,
      "postDate": "2020-10-21T15:31:51.600Z",
      "content": "<p>Did anyone find the source of the memory leak? This keeps popping up from time to time!</p>",
      "rawMarkdown": "Did anyone find the source of the memory leak? This keeps popping up from time to time!",
      "votes": 3
    },
    {
      "id": 1008104,
      "postDate": "2020-09-12T18:49:23.907Z",
      "content": "<p>The issue is not with your code, I had the same problem. After a few hours, my opened chrome tabs started to crash.. (I have 128Gb of RAM)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fa7f6a78395137962153a5e3a29449edf%2Fmemory_leak.png?generation=1599936382798391&amp;alt=media\" alt=\"\"></p>\n<p>I am using the latest code (l5kit repo's master branch) and it seems fixed now.</p>",
      "rawMarkdown": "The issue is not with your code, I had the same problem. After a few hours, my opened chrome tabs started to crash.. (I have 128Gb of RAM)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fa7f6a78395137962153a5e3a29449edf%2Fmemory_leak.png?generation=1599936382798391&alt=media)\n\nI am using the latest code (l5kit repo's master branch) and it seems fixed now.",
      "votes": 4,
      "replies": [
        {
          "id": 1008482,
          "postDate": "2020-09-13T06:40:06.110Z",
          "content": "<p>Hmm I just updated and updated my pytorch to 1.6 and still seem to be getting it</p>",
          "rawMarkdown": "Hmm I just updated and updated my pytorch to 1.6 and still seem to be getting it"
        },
        {
          "id": 1009007,
          "postDate": "2020-09-13T15:24:34.970Z",
          "content": "<p>also, I've been using the utility script workaround you configured. Do you know which version that is?</p>",
          "rawMarkdown": "also, I've been using the utility script workaround you configured. Do you know which version that is?"
        },
        {
          "id": 1009017,
          "postDate": "2020-09-13T15:37:54.717Z",
          "content": "<p>No idea. I think it is the latest released version (1.0.6). Try this:</p>\n<pre><code>import l5kit\nprint(l5kit.__version__)\n</code></pre>",
          "rawMarkdown": "No idea. I think it is the latest released version (1.0.6). Try this:\n\n```\nimport l5kit\nprint(l5kit.__version__)\n```"
        }
      ]
    },
    {
      "id": 1082898,
      "postDate": "2020-11-18T11:00:35.913Z",
      "content": "<p>Are there any solutions for this? We are still experiencing this problem with the l5kit 1.1.0. It uses all my available physical memory (~11GB) in like 20 seconds and stay at 150-200MB free all the time.</p>",
      "rawMarkdown": "Are there any solutions for this? We are still experiencing this problem with the l5kit 1.1.0. It uses all my available physical memory (~11GB) in like 20 seconds and stay at 150-200MB free all the time.",
      "votes": 1,
      "replies": [
        {
          "id": 1082944,
          "postDate": "2020-11-18T12:14:46.570Z",
          "content": "<p>Have you disabled caching? That worked for me.</p>",
          "rawMarkdown": "Have you disabled caching? That worked for me."
        },
        {
          "id": 1083094,
          "postDate": "2020-11-18T15:38:10.380Z",
          "content": "<p>I was never able to make this totally go away and ended up just engineering around this by creating a script that I just run repeatedly from terminal. It's a janky solution but I had to move on from probing dependency versions. </p>",
          "rawMarkdown": "I was never able to make this totally go away and ended up just engineering around this by creating a script that I just run repeatedly from terminal. It's a janky solution but I had to move on from probing dependency versions. "
        },
        {
          "id": 1083194,
          "postDate": "2020-11-18T17:41:04.053Z",
          "content": "<p>If you're using mixed precision, you can simply cast image to fp16 inside rasterizer. This is what helped me. This is pretty crude solution, but I was able to stabilize RAM usage.</p>",
          "rawMarkdown": "If you're using mixed precision, you can simply cast image to fp16 inside rasterizer. This is what helped me. This is pretty crude solution, but I was able to stabilize RAM usage."
        }
      ]
    },
    {
      "id": 1011941,
      "postDate": "2020-09-15T19:13:35.937Z",
      "content": "<p>I ran a data loader loop in a script (not within a notebook). It certainly eats up memory and CPU like crazy but seemed to stabilize. 4 workers stabilized at 6.4GB virtual, 3GB physical. It started at 4/2GB in first epoch and jumped to the stable value in epoch 2/3 and then stayed steady. </p>\n<p>Bumping up the workers to 16, memory usage was 9GB virt, btw 3-4GB phys after epoch 3 with larger jumps per epoch. I killed it but suspect there is a steady state, just larger. There is likely a relationship between num workers and the steady state memory use. </p>\n<p>I feel the design as a while isn't optimal, each worker process zipping through the zarrs independently without co-ordination, but not sure how else you'd get something simple working without threading and tighter synchronization (needing C++).</p>",
      "rawMarkdown": "I ran a data loader loop in a script (not within a notebook). It certainly eats up memory and CPU like crazy but seemed to stabilize. 4 workers stabilized at 6.4GB virtual, 3GB physical. It started at 4/2GB in first epoch and jumped to the stable value in epoch 2/3 and then stayed steady. \n\nBumping up the workers to 16, memory usage was 9GB virt, btw 3-4GB phys after epoch 3 with larger jumps per epoch. I killed it but suspect there is a steady state, just larger. There is likely a relationship between num workers and the steady state memory use. \n\nI feel the design as a while isn't optimal, each worker process zipping through the zarrs independently without co-ordination, but not sure how else you'd get something simple working without threading and tighter synchronization (needing C++).",
      "votes": 1,
      "replies": [
        {
          "id": 1011956,
          "postDate": "2020-09-15T19:35:20.073Z",
          "content": "<p>Hmm, I am wondering if it is a package version issue. I believe I am on newest build but I will have to check on that. </p>",
          "rawMarkdown": "Hmm, I am wondering if it is a package version issue. I believe I am on newest build but I will have to check on that. "
        },
        {
          "id": 1011958,
          "postDate": "2020-09-15T19:39:41.780Z",
          "content": "<p>Was what you ran close to what I ran on the kaggle kernels or do you think there was maybe a code difference there that is worth investigating?</p>",
          "rawMarkdown": "Was what you ran close to what I ran on the kaggle kernels or do you think there was maybe a code difference there that is worth investigating?"
        },
        {
          "id": 1011964,
          "postDate": "2020-09-15T19:42:58.953Z",
          "content": "<p>Basically the same, just with the normal tqdm instead of notebook one. Latest l5kit.</p>",
          "rawMarkdown": "Basically the same, just with the normal tqdm instead of notebook one. Latest l5kit."
        },
        {
          "id": 1011982,
          "postDate": "2020-09-15T20:03:41.107Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1008813,
      "postDate": "2020-09-13T12:14:49.333Z",
      "content": "<p>I've been investigated this but It's quite difficult to pin this down to a specific component. It's not clear if this is:</p>\n<ul>\n<li>a l5kit issue;</li>\n<li>a torch issue;</li>\n<li>a combination of both.</li>\n</ul>\n<p>We use caching in the zarr, which may conflict with torch loaders (it's not clear whether the cache is entirely copied between workers).</p>\n<p>Something you can try is disabling cache for the zarr dataset by using:</p>\n<p><code>ChunkedDataset(\"path\").open(cached=False)</code></p>\n<p>I'll take a deep look into this next week, as it's affecting multiple users</p>",
      "rawMarkdown": "I've been investigated this but It's quite difficult to pin this down to a specific component. It's not clear if this is:\n- a l5kit issue;\n- a torch issue;\n- a combination of both.\n\nWe use caching in the zarr, which may conflict with torch loaders (it's not clear whether the cache is entirely copied between workers).\n\nSomething you can try is disabling cache for the zarr dataset by using:\n\n`ChunkedDataset(\"path\").open(cached=False)`\n\nI'll take a deep look into this next week, as it's affecting multiple users",
      "votes": 1,
      "replies": [
        {
          "id": 1009010,
          "postDate": "2020-09-13T15:26:00.930Z",
          "content": "<p>Yes, it is a bit unclear where it is coming from exactly. I just tried with caching turned off and it still seems to accumulate. Across 30 minutes it gained an extra 3gb. </p>",
          "rawMarkdown": "Yes, it is a bit unclear where it is coming from exactly. I just tried with caching turned off and it still seems to accumulate. Across 30 minutes it gained an extra 3gb. ",
          "votes": 1
        },
        {
          "id": 1009294,
          "postDate": "2020-09-13T20:25:47.433Z",
          "content": "<p>Thanks for reporting the issue in L5kit GitHub page <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>, I'll report back there what I'll find!</p>",
          "rawMarkdown": "Thanks for reporting the issue in L5kit GitHub page @ryches, I'll report back there what I'll find!"
        }
      ]
    },
    {
      "id": 1010197,
      "postDate": "2020-09-14T15:15:22.440Z",
      "content": "<p>I may have found a work around. Change the max number of steps to 100 and correspondingly increase the number of epochs to make up for the same number of training examples. This dumps the memory after the first 100 iterations. This along with a batch size of 16 seems to work on the Kaggle kernel with GPU. However, the bottleneck is the rasterization, so it maybe better to just use CPU for this project. </p>",
      "rawMarkdown": "I may have found a work around. Change the max number of steps to 100 and correspondingly increase the number of epochs to make up for the same number of training examples. This dumps the memory after the first 100 iterations. This along with a batch size of 16 seems to work on the Kaggle kernel with GPU. However, the bottleneck is the rasterization, so it maybe better to just use CPU for this project. ",
      "replies": [
        {
          "id": 1010425,
          "postDate": "2020-09-14T18:23:21.657Z",
          "content": "<p>hmm so maybe it only starts accumulating after taking a certain number of steps? Is that what this is implying?</p>",
          "rawMarkdown": "hmm so maybe it only starts accumulating after taking a certain number of steps? Is that what this is implying?"
        },
        {
          "id": 1010429,
          "postDate": "2020-09-14T18:28:59.933Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> For me, the leak was there when I only ran the loader/iterator. I mean without any deep learning (no move to cuda, no  forward/backward, no optimizer, etc). I am not 100% sure why is fixed now, but I'll try to reproduce, maybe I'll find something useful.</p>",
          "rawMarkdown": "@ryches For me, the leak was there when I only ran the loader/iterator. I mean without any deep learning (no move to cuda, no  forward/backward, no optimizer, etc). I am not 100% sure why is fixed now, but I'll try to reproduce, maybe I'll find something useful."
        },
        {
          "id": 1010448,
          "postDate": "2020-09-14T18:44:02.890Z",
          "content": "<p>Here is what I did. Put in a dummy generator that gets reassigned after the inner loop. It ran for over an hour but finally crashed because of memory issue. Previously it would not even run for 10 mins.</p>\n<p>`for i in range(epochs):</p>\n<pre><code> def gen(dataloader):\n\n    for batch in dataloader:\n\n        yield batch\n    train_gen = gen(train_dataloader)\n\n    progress_bar = tqdm(range(cfg[\"train_params\"][\"max_num_steps\"]))\n   for _ in progress_bar:\n       try:\n          data = next(train_gen)\n    except StopIteration:\n        tr_it = iter(train_dataloader)\n        data = next(train_gen)`\n</code></pre>",
          "rawMarkdown": "Here is what I did. Put in a dummy generator that gets reassigned after the inner loop. It ran for over an hour but finally crashed because of memory issue. Previously it would not even run for 10 mins.\n\n`for i in range(epochs):\n\n     def gen(dataloader):\n\n        for batch in dataloader:\n\n            yield batch\n        train_gen = gen(train_dataloader)\n\n        progress_bar = tqdm(range(cfg[\"train_params\"][\"max_num_steps\"]))\n       for _ in progress_bar:\n           try:\n              data = next(train_gen)\n        except StopIteration:\n            tr_it = iter(train_dataloader)\n            data = next(train_gen)`\n"
        },
        {
          "id": 1010559,
          "postDate": "2020-09-14T21:28:12.603Z",
          "content": "<p>Yeah I was able to narrow my issue down to just the dataloader as well. I was originally trying with my training and validation and thinking maybe my validation was accidentally accumulating gradients over time so I turned that off. Then I turned off training after to see if just data loading alone was causing it and that seems to be the case. </p>",
          "rawMarkdown": "Yeah I was able to narrow my issue down to just the dataloader as well. I was originally trying with my training and validation and thinking maybe my validation was accidentally accumulating gradients over time so I turned that off. Then I turned off training after to see if just data loading alone was causing it and that seems to be the case. "
        },
        {
          "id": 1011413,
          "postDate": "2020-09-15T12:54:59.277Z",
          "content": "<blockquote>\n  <p>maybe it only starts accumulating after taking a certain number of steps? Is that what this is implying?</p>\n</blockquote>\n<p>PyTorch dataloader workers are created in the iterator. So restarting iteration will kill all the previous workers and release any leaked memory.<br>\nRestarting will also clear the LRU caches in the zarr dataset which otherwise seem liable to cause issues just through correct operation. With the default config of 16 workers each with a maximum cache of 1Gb that's 16Gb just for caching. Can't see any checks around available memory.</p>",
          "rawMarkdown": "> maybe it only starts accumulating after taking a certain number of steps? Is that what this is implying?\n\nPyTorch dataloader workers are created in the iterator. So restarting iteration will kill all the previous workers and release any leaked memory.\nRestarting will also clear the LRU caches in the zarr dataset which otherwise seem liable to cause issues just through correct operation. With the default config of 16 workers each with a maximum cache of 1Gb that's 16Gb just for caching. Can't see any checks around available memory.",
          "votes": 3
        },
        {
          "id": 1014181,
          "postDate": "2020-09-17T08:58:12.557Z",
          "content": "<p>You can completely disable cache or reduce the memory footprint by tuning the parameters of the <a href=\"https://github.com/lyft/l5kit/blob/8f75068a7b594f424c5b09ffe8ae5b0375c1c3ba/l5kit/l5kit/data/zarr_dataset.py#L126\" target=\"_blank\"><code>open()</code> method in the <code>ChunkedDataset</code></a></p>",
          "rawMarkdown": "You can completely disable cache or reduce the memory footprint by tuning the parameters of the [`open()` method in the `ChunkedDataset`](https://github.com/lyft/l5kit/blob/8f75068a7b594f424c5b09ffe8ae5b0375c1c3ba/l5kit/l5kit/data/zarr_dataset.py#L126)",
          "votes": 2
        },
        {
          "id": 1018062,
          "postDate": "2020-09-19T12:27:08.417Z",
          "content": "<p>For me, a disabling caching solved the problem. Memory usage tends to fluctuate more now, but it doesn't accumulate and the process is no longer killed by the OS. <br>\nI didn't notice any measureable performance drop due to disabling caching. </p>",
          "rawMarkdown": "For me, a disabling caching solved the problem. Memory usage tends to fluctuate more now, but it doesn't accumulate and the process is no longer killed by the OS. \nI didn't notice any measureable performance drop due to disabling caching. ",
          "votes": 4
        },
        {
          "id": 1018440,
          "postDate": "2020-09-19T17:23:14.923Z",
          "content": "<p>I also didnt see any appreciable change in performance. I haven't inspected closely, but it seems like caching in this case isn't super useful. Maybe default should be set to False instead of True</p>",
          "rawMarkdown": "I also didnt see any appreciable change in performance. I haven't inspected closely, but it seems like caching in this case isn't super useful. Maybe default should be set to False instead of True"
        },
        {
          "id": 1020620,
          "postDate": "2020-09-21T10:00:12.197Z",
          "content": "<p>Yes, the issue is that zarr caches the raw chunk and not the decompressed one (no idea why that's not an option). That means it still needs to run blosc on the chunk even if it's cached</p>",
          "rawMarkdown": "Yes, the issue is that zarr caches the raw chunk and not the decompressed one (no idea why that's not an option). That means it still needs to run blosc on the chunk even if it's cached"
        }
      ]
    },
    {
      "id": 1008277,
      "postDate": "2020-09-13T00:21:03.053Z",
      "content": "<p>I face the same issue when <code>num_workers &gt; 0</code> in the dataloader. I think it has to do with using dicts for loading the data. See here: <a href=\"url\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/13246#issuecomment-445770039</a></p>\n<p>I haven't tried it with the latest l5kit update yet though.</p>",
      "rawMarkdown": "I face the same issue when `num_workers > 0` in the dataloader. I think it has to do with using dicts for loading the data. See here: [https://github.com/pytorch/pytorch/issues/13246#issuecomment-445770039](url)\n\nI haven't tried it with the latest l5kit update yet though."
    },
    {
      "id": 1008039,
      "postDate": "2020-09-12T18:00:24.023Z",
      "content": "<p>This is for the most part just code pulled from the example notebook from l5kit that was provided to us</p>",
      "rawMarkdown": "This is for the most part just code pulled from the example notebook from l5kit that was provided to us"
    }
  ],
  "comments": [
    {
      "id": 1056316,
      "author_name": "A_Elsheikh",
      "author_url": "",
      "post_date": "2020-10-21T15:31:51.600000",
      "content": "<p>Did anyone find the source of the memory leak? This keeps popping up from time to time!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1008104,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2020-09-12T18:49:23.907000",
      "content": "<p>The issue is not with your code, I had the same problem. After a few hours, my opened chrome tabs started to crash.. (I have 128Gb of RAM)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fa7f6a78395137962153a5e3a29449edf%2Fmemory_leak.png?generation=1599936382798391&amp;alt=media\" alt=\"\"></p>\n<p>I am using the latest code (l5kit repo's master branch) and it seems fixed now.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1008482,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-13T06:40:06.110000",
          "content": "<p>Hmm I just updated and updated my pytorch to 1.6 and still seem to be getting it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1009007,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-13T15:24:34.970000",
          "content": "<p>also, I've been using the utility script workaround you configured. Do you know which version that is?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1009017,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-13T15:37:54.717000",
          "content": "<p>No idea. I think it is the latest released version (1.0.6). Try this:</p>\n<pre><code>import l5kit\nprint(l5kit.__version__)\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1082898,
      "author_name": "Anil Ozturk",
      "author_url": "",
      "post_date": "2020-11-18T11:00:35.913000",
      "content": "<p>Are there any solutions for this? We are still experiencing this problem with the l5kit 1.1.0. It uses all my available physical memory (~11GB) in like 20 seconds and stay at 150-200MB free all the time.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1082944,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-11-18T12:14:46.570000",
          "content": "<p>Have you disabled caching? That worked for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1083094,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-11-18T15:38:10.380000",
          "content": "<p>I was never able to make this totally go away and ended up just engineering around this by creating a script that I just run repeatedly from terminal. It's a janky solution but I had to move on from probing dependency versions. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1083194,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-11-18T17:41:04.053000",
          "content": "<p>If you're using mixed precision, you can simply cast image to fp16 inside rasterizer. This is what helped me. This is pretty crude solution, but I was able to stabilize RAM usage.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1011941,
      "author_name": "RossWightman",
      "author_url": "",
      "post_date": "2020-09-15T19:13:35.937000",
      "content": "<p>I ran a data loader loop in a script (not within a notebook). It certainly eats up memory and CPU like crazy but seemed to stabilize. 4 workers stabilized at 6.4GB virtual, 3GB physical. It started at 4/2GB in first epoch and jumped to the stable value in epoch 2/3 and then stayed steady. </p>\n<p>Bumping up the workers to 16, memory usage was 9GB virt, btw 3-4GB phys after epoch 3 with larger jumps per epoch. I killed it but suspect there is a steady state, just larger. There is likely a relationship between num workers and the steady state memory use. </p>\n<p>I feel the design as a while isn't optimal, each worker process zipping through the zarrs independently without co-ordination, but not sure how else you'd get something simple working without threading and tighter synchronization (needing C++).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1011956,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-15T19:35:20.073000",
          "content": "<p>Hmm, I am wondering if it is a package version issue. I believe I am on newest build but I will have to check on that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1011958,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-15T19:39:41.780000",
          "content": "<p>Was what you ran close to what I ran on the kaggle kernels or do you think there was maybe a code difference there that is worth investigating?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1011964,
          "author_name": "RossWightman",
          "author_url": "",
          "post_date": "2020-09-15T19:42:58.953000",
          "content": "<p>Basically the same, just with the normal tqdm instead of notebook one. Latest l5kit.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1011982,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-15T20:03:41.107000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1008813,
      "author_name": "Luca Bergamini",
      "author_url": "",
      "post_date": "2020-09-13T12:14:49.333000",
      "content": "<p>I've been investigated this but It's quite difficult to pin this down to a specific component. It's not clear if this is:</p>\n<ul>\n<li>a l5kit issue;</li>\n<li>a torch issue;</li>\n<li>a combination of both.</li>\n</ul>\n<p>We use caching in the zarr, which may conflict with torch loaders (it's not clear whether the cache is entirely copied between workers).</p>\n<p>Something you can try is disabling cache for the zarr dataset by using:</p>\n<p><code>ChunkedDataset(\"path\").open(cached=False)</code></p>\n<p>I'll take a deep look into this next week, as it's affecting multiple users</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1009010,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-13T15:26:00.930000",
          "content": "<p>Yes, it is a bit unclear where it is coming from exactly. I just tried with caching turned off and it still seems to accumulate. Across 30 minutes it gained an extra 3gb. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1009294,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-13T20:25:47.433000",
          "content": "<p>Thanks for reporting the issue in L5kit GitHub page <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>, I'll report back there what I'll find!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1010197,
      "author_name": "Srinivas Gopal Krishna",
      "author_url": "",
      "post_date": "2020-09-14T15:15:22.440000",
      "content": "<p>I may have found a work around. Change the max number of steps to 100 and correspondingly increase the number of epochs to make up for the same number of training examples. This dumps the memory after the first 100 iterations. This along with a batch size of 16 seems to work on the Kaggle kernel with GPU. However, the bottleneck is the rasterization, so it maybe better to just use CPU for this project. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1010425,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-14T18:23:21.657000",
          "content": "<p>hmm so maybe it only starts accumulating after taking a certain number of steps? Is that what this is implying?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1010429,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-14T18:28:59.933000",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> For me, the leak was there when I only ran the loader/iterator. I mean without any deep learning (no move to cuda, no  forward/backward, no optimizer, etc). I am not 100% sure why is fixed now, but I'll try to reproduce, maybe I'll find something useful.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1010448,
          "author_name": "Srinivas Gopal Krishna",
          "author_url": "",
          "post_date": "2020-09-14T18:44:02.890000",
          "content": "<p>Here is what I did. Put in a dummy generator that gets reassigned after the inner loop. It ran for over an hour but finally crashed because of memory issue. Previously it would not even run for 10 mins.</p>\n<p>`for i in range(epochs):</p>\n<pre><code> def gen(dataloader):\n\n    for batch in dataloader:\n\n        yield batch\n    train_gen = gen(train_dataloader)\n\n    progress_bar = tqdm(range(cfg[\"train_params\"][\"max_num_steps\"]))\n   for _ in progress_bar:\n       try:\n          data = next(train_gen)\n    except StopIteration:\n        tr_it = iter(train_dataloader)\n        data = next(train_gen)`\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1010559,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-14T21:28:12.603000",
          "content": "<p>Yeah I was able to narrow my issue down to just the dataloader as well. I was originally trying with my training and validation and thinking maybe my validation was accidentally accumulating gradients over time so I turned that off. Then I turned off training after to see if just data loading alone was causing it and that seems to be the case. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1011413,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2020-09-15T12:54:59.277000",
          "content": "<blockquote>\n  <p>maybe it only starts accumulating after taking a certain number of steps? Is that what this is implying?</p>\n</blockquote>\n<p>PyTorch dataloader workers are created in the iterator. So restarting iteration will kill all the previous workers and release any leaked memory.<br>\nRestarting will also clear the LRU caches in the zarr dataset which otherwise seem liable to cause issues just through correct operation. With the default config of 16 workers each with a maximum cache of 1Gb that's 16Gb just for caching. Can't see any checks around available memory.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1014181,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-17T08:58:12.557000",
          "content": "<p>You can completely disable cache or reduce the memory footprint by tuning the parameters of the <a href=\"https://github.com/lyft/l5kit/blob/8f75068a7b594f424c5b09ffe8ae5b0375c1c3ba/l5kit/l5kit/data/zarr_dataset.py#L126\" target=\"_blank\"><code>open()</code> method in the <code>ChunkedDataset</code></a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1018062,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-19T12:27:08.417000",
          "content": "<p>For me, a disabling caching solved the problem. Memory usage tends to fluctuate more now, but it doesn't accumulate and the process is no longer killed by the OS. <br>\nI didn't notice any measureable performance drop due to disabling caching. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1018440,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-09-19T17:23:14.923000",
          "content": "<p>I also didnt see any appreciable change in performance. I haven't inspected closely, but it seems like caching in this case isn't super useful. Maybe default should be set to False instead of True</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1020620,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-21T10:00:12.197000",
          "content": "<p>Yes, the issue is that zarr caches the raw chunk and not the decompressed one (no idea why that's not an option). That means it still needs to run blosc on the chunk even if it's cached</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1008277,
      "author_name": "Khushal B",
      "author_url": "",
      "post_date": "2020-09-13T00:21:03.053000",
      "content": "<p>I face the same issue when <code>num_workers &gt; 0</code> in the dataloader. I think it has to do with using dicts for loading the data. See here: <a href=\"url\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/13246#issuecomment-445770039</a></p>\n<p>I haven't tried it with the latest l5kit update yet though.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1008039,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-09-12T18:00:24.023000",
      "content": "<p>This is for the most part just code pulled from the example notebook from l5kit that was provided to us</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1008037": "https://www.kaggle.com/ryches/memory-leak\n\nI don't know if others have run into this but I have been fighting against a memory leak. I set up a minimal example to just iterate through the data and do nothing else. It isnt storing anything or accumulating gradients or anything, but memory usage seems to continue to drift overtime until eventually it kills the kernel. Have others run into this? Am I just doing something wrong?",
    "1056316": "Did anyone find the source of the memory leak? This keeps popping up from time to time!",
    "1008104": "The issue is not with your code, I had the same problem. After a few hours, my opened chrome tabs started to crash.. (I have 128Gb of RAM)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F864684%2Fa7f6a78395137962153a5e3a29449edf%2Fmemory_leak.png?generation=1599936382798391&alt=media)\n\nI am using the latest code (l5kit repo's master branch) and it seems fixed now.",
    "1082898": "Are there any solutions for this? We are still experiencing this problem with the l5kit 1.1.0. It uses all my available physical memory (~11GB) in like 20 seconds and stay at 150-200MB free all the time.",
    "1011941": "I ran a data loader loop in a script (not within a notebook). It certainly eats up memory and CPU like crazy but seemed to stabilize. 4 workers stabilized at 6.4GB virtual, 3GB physical. It started at 4/2GB in first epoch and jumped to the stable value in epoch 2/3 and then stayed steady. \n\nBumping up the workers to 16, memory usage was 9GB virt, btw 3-4GB phys after epoch 3 with larger jumps per epoch. I killed it but suspect there is a steady state, just larger. There is likely a relationship between num workers and the steady state memory use. \n\nI feel the design as a while isn't optimal, each worker process zipping through the zarrs independently without co-ordination, but not sure how else you'd get something simple working without threading and tighter synchronization (needing C++).",
    "1008813": "I've been investigated this but It's quite difficult to pin this down to a specific component. It's not clear if this is:\n- a l5kit issue;\n- a torch issue;\n- a combination of both.\n\nWe use caching in the zarr, which may conflict with torch loaders (it's not clear whether the cache is entirely copied between workers).\n\nSomething you can try is disabling cache for the zarr dataset by using:\n\n`ChunkedDataset(\"path\").open(cached=False)`\n\nI'll take a deep look into this next week, as it's affecting multiple users",
    "1010197": "I may have found a work around. Change the max number of steps to 100 and correspondingly increase the number of epochs to make up for the same number of training examples. This dumps the memory after the first 100 iterations. This along with a batch size of 16 seems to work on the Kaggle kernel with GPU. However, the bottleneck is the rasterization, so it maybe better to just use CPU for this project. ",
    "1008277": "I face the same issue when `num_workers > 0` in the dataloader. I think it has to do with using dicts for loading the data. See here: [https://github.com/pytorch/pytorch/issues/13246#issuecomment-445770039](url)\n\nI haven't tried it with the latest l5kit update yet though.",
    "1008039": "This is for the most part just code pulled from the example notebook from l5kit that was provided to us"
  }
}