{
  "id": 195259,
  "title": "Fix L5kit memory leak on Windows 10",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/195259",
  "author_name": "",
  "post_date": "2020-11-04T10:07:05.250072100Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My notebook run fine on Kaggle kernel but get into serious memory leak on Windows. I was wondering if anyone had that issue?</p>\n<p>The notebook I run is this <a href=\"https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline\" target=\"_blank\">https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline</a>. I run it for 12000 batches with <code>batch_size = 16</code>. Due to the l5kit and pytorch multiprocess issue on Windows, I can only run it with <code>num_workers=0</code>. I have also set the <code>.open(cached=False)</code> for the <code>ChunkedDataset</code> so zarr is not caching the data.</p>\n<p>On Kaggle machine, it can easily run through 12000 batches with just 5GB of memory. But on my Windows 10, it uses all the 32GB of RAM at about 11000 batches. When I look into the usage, looks like the memory was used by nonpool memory. And, seems to be caused by the massive amount of handles (1,000,000+) created by python.exe. Do people have this issue? How do you avoid this?</p>\n<p>I am using 2080 TI, AMD 3900X, 32GB RAM, Window 10 64bits, pytorch 1.6.0, CUDA 11.1, python 3.7.5.</p>\n<p>Edit:<br>\nRemove the line <code>os.environ[\"BLOSC_NOLOCK\"] = \"1\"</code> in the code<br>\n<a href=\"https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\" target=\"_blank\">https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26</a><br>\nFix the problem on Windows.</p>",
  "messages": [
    {
      "id": "1069315",
      "postDate": "11/04/2020 10:07:05",
      "content": "<p>My notebook run fine on Kaggle kernel but get into serious memory leak on Windows. I was wondering if anyone had that issue?</p>\n<p>The notebook I run is this <a href=\"https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline\" target=\"_blank\">https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline</a>. I run it for 12000 batches with <code>batch_size = 16</code>. Due to the l5kit and pytorch multiprocess issue on Windows, I can only run it with <code>num_workers=0</code>. I have also set the <code>.open(cached=False)</code> for the <code>ChunkedDataset</code> so zarr is not caching the data.</p>\n<p>On Kaggle machine, it can easily run through 12000 batches with just 5GB of memory. But on my Windows 10, it uses all the 32GB of RAM at about 11000 batches. When I look into the usage, looks like the memory was used by nonpool memory. And, seems to be caused by the massive amount of handles (1,000,000+) created by python.exe. Do people have this issue? How do you avoid this?</p>\n<p>I am using 2080 TI, AMD 3900X, 32GB RAM, Window 10 64bits, pytorch 1.6.0, CUDA 11.1, python 3.7.5.</p>\n<p>Edit:<br>\nRemove the line <code>os.environ[\"BLOSC_NOLOCK\"] = \"1\"</code> in the code<br>\n<a href=\"https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\" target=\"_blank\">https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26</a><br>\nFix the problem on Windows.</p>",
      "rawMarkdown": "My notebook run fine on Kaggle kernel but get into serious memory leak on Windows. I was wondering if anyone had that issue?\n\nThe notebook I run is this [https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline](https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline). I run it for 12000 batches with `batch_size = 16`. Due to the l5kit and pytorch multiprocess issue on Windows, I can only run it with `num_workers=0`. I have also set the `.open(cached=False)` for the `ChunkedDataset` so zarr is not caching the data.\n\nOn Kaggle machine, it can easily run through 12000 batches with just 5GB of memory. But on my Windows 10, it uses all the 32GB of RAM at about 11000 batches. When I look into the usage, looks like the memory was used by nonpool memory. And, seems to be caused by the massive amount of handles (1,000,000+) created by python.exe. Do people have this issue? How do you avoid this?\n\nI am using 2080 TI, AMD 3900X, 32GB RAM, Window 10 64bits, pytorch 1.6.0, CUDA 11.1, python 3.7.5.\n\nEdit:\nRemove the line `os.environ[\"BLOSC_NOLOCK\"] = \"1\"` in the code\nhttps://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\nFix the problem on Windows.",
      "votes": null
    },
    {
      "id": "1069549",
      "postDate": "11/04/2020 15:34:44",
      "content": "<p>I have run into this same issue myself with almost the same setup. I simply run the bare bones notebook <a href=\"https://www.kaggle.com/ryches/memory-leak\" target=\"_blank\">here</a> and get a memory leak no matter what I change. I have played around with pytorch and python versions but all seem to have the leak. If anyone knows a fix that would be greatly appreciated!</p>",
      "rawMarkdown": "I have run into this same issue myself with almost the same setup. I simply run the bare bones notebook [here](https://www.kaggle.com/ryches/memory-leak) and get a memory leak no matter what I change. I have played around with pytorch and python versions but all seem to have the leak. If anyone knows a fix that would be greatly appreciated!",
      "votes": null
    },
    {
      "id": "1069578",
      "postDate": "11/04/2020 16:11:17",
      "content": "<p>You can try to install dual boot setup and then install ubuntu and also install nvidia cudatoolkit according to the gpu in your system.You can easily run and can adjust numworkers according to the <strong>os.cpu_count()</strong>.For RAM usage usage you can use <strong>top</strong> command in your ubuntu terminal and <strong>watch -n0.1  nvidia-smi</strong> for your GPU usage as well as for its utilization</p>",
      "rawMarkdown": "You can try to install dual boot setup and then install ubuntu and also install nvidia cudatoolkit according to the gpu in your system.You can easily run and can adjust numworkers according to the **os.cpu_count()**.For RAM usage usage you can use **top** command in your ubuntu terminal and **watch -n0.1  nvidia-smi** for your GPU usage as well as for its utilization",
      "votes": null
    },
    {
      "id": "1069827",
      "postDate": "11/05/2020 01:20:25",
      "content": "<p>I try to iterate through the AgentDataset without pytorch DataLoader, the memory leak remains.</p>\n<pre><code>train_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open(cached=False)\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\nfor i in range(4000):\n        data = [train_dataset[k + i * batch_size] for k in range(batch_size)]\n</code></pre>\n<p>So the issue is from l5kit.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F1d9e81e4fda4cb216115ba8cd73e6707%2F20201104%20l5kit%20memory%20leak%20in%20api.png?generation=1604539018730841&amp;alt=media\" alt=\"l5kit memory leak as a function of time\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F16fb971257d81e7c9c36bb355033e188%2F20201104%20l5kit%20memory%20leak%20in%20api%202.png?generation=1604539028411364&amp;alt=media\" alt=\"l5kit memory leak as a function of iteration\"></p>",
      "rawMarkdown": "I try to iterate through the AgentDataset without pytorch DataLoader, the memory leak remains.\n```\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open(cached=False)\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\nfor i in range(4000):\n        data = [train_dataset[k + i * batch_size] for k in range(batch_size)]\n```\nSo the issue is from l5kit.\n\n![l5kit memory leak as a function of time](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F1d9e81e4fda4cb216115ba8cd73e6707%2F20201104%20l5kit%20memory%20leak%20in%20api.png?generation=1604539018730841&alt=media)\n\n![l5kit memory leak as a function of iteration](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F16fb971257d81e7c9c36bb355033e188%2F20201104%20l5kit%20memory%20leak%20in%20api%202.png?generation=1604539028411364&alt=media)",
      "votes": null
    },
    {
      "id": "1071785",
      "postDate": "11/07/2020 12:06:37",
      "content": "<p>After debugging for like 2 days, I think I figured out! <br>\nIt is the random line<br>\n<code>os.environ[\"BLOSC_NOLOCK\"] = \"1\"</code><br>\nin the code<br>\n<a href=\"https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\" target=\"_blank\">https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26</a><br>\nThis terrible line will modify your system variable secretly whenever you import any module from <code>l5kit.dataset</code>.<br>\nOnce you remove this line, on Windows you will no longer see the exploding number of handles created by python.</p>",
      "rawMarkdown": "After debugging for like 2 days, I think I figured out! \nIt is the random line\n`os.environ[\"BLOSC_NOLOCK\"] = \"1\"`\nin the code\nhttps://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\nThis terrible line will modify your system variable secretly whenever you import any module from `l5kit.dataset`.\nOnce you remove this line, on Windows you will no longer see the exploding number of handles created by python.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1069549,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "11/04/2020 15:34:44",
      "content": "<p>I have run into this same issue myself with almost the same setup. I simply run the bare bones notebook <a href=\"https://www.kaggle.com/ryches/memory-leak\" target=\"_blank\">here</a> and get a memory leak no matter what I change. I have played around with pytorch and python versions but all seem to have the leak. If anyone knows a fix that would be greatly appreciated!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1069578,
      "author_name": "",
      "author_url": "",
      "post_date": "11/04/2020 16:11:17",
      "content": "<p>You can try to install dual boot setup and then install ubuntu and also install nvidia cudatoolkit according to the gpu in your system.You can easily run and can adjust numworkers according to the <strong>os.cpu_count()</strong>.For RAM usage usage you can use <strong>top</strong> command in your ubuntu terminal and <strong>watch -n0.1  nvidia-smi</strong> for your GPU usage as well as for its utilization</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1069827,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/05/2020 01:20:25",
      "content": "<p>I try to iterate through the AgentDataset without pytorch DataLoader, the memory leak remains.</p>\n<pre><code>train_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open(cached=False)\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\nfor i in range(4000):\n        data = [train_dataset[k + i * batch_size] for k in range(batch_size)]\n</code></pre>\n<p>So the issue is from l5kit.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F1d9e81e4fda4cb216115ba8cd73e6707%2F20201104%20l5kit%20memory%20leak%20in%20api.png?generation=1604539018730841&amp;alt=media\" alt=\"l5kit memory leak as a function of time\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F16fb971257d81e7c9c36bb355033e188%2F20201104%20l5kit%20memory%20leak%20in%20api%202.png?generation=1604539028411364&amp;alt=media\" alt=\"l5kit memory leak as a function of iteration\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1071785,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/07/2020 12:06:37",
      "content": "<p>After debugging for like 2 days, I think I figured out! <br>\nIt is the random line<br>\n<code>os.environ[\"BLOSC_NOLOCK\"] = \"1\"</code><br>\nin the code<br>\n<a href=\"https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\" target=\"_blank\">https://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26</a><br>\nThis terrible line will modify your system variable secretly whenever you import any module from <code>l5kit.dataset</code>.<br>\nOnce you remove this line, on Windows you will no longer see the exploding number of handles created by python.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1069315": "My notebook run fine on Kaggle kernel but get into serious memory leak on Windows. I was wondering if anyone had that issue?\n\nThe notebook I run is this [https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline](https://www.kaggle.com/louis925/lyft-complete-train-and-prediction-pipeline). I run it for 12000 batches with `batch_size = 16`. Due to the l5kit and pytorch multiprocess issue on Windows, I can only run it with `num_workers=0`. I have also set the `.open(cached=False)` for the `ChunkedDataset` so zarr is not caching the data.\n\nOn Kaggle machine, it can easily run through 12000 batches with just 5GB of memory. But on my Windows 10, it uses all the 32GB of RAM at about 11000 batches. When I look into the usage, looks like the memory was used by nonpool memory. And, seems to be caused by the massive amount of handles (1,000,000+) created by python.exe. Do people have this issue? How do you avoid this?\n\nI am using 2080 TI, AMD 3900X, 32GB RAM, Window 10 64bits, pytorch 1.6.0, CUDA 11.1, python 3.7.5.\n\nEdit:\nRemove the line `os.environ[\"BLOSC_NOLOCK\"] = \"1\"` in the code\nhttps://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\nFix the problem on Windows.",
    "1069549": "I have run into this same issue myself with almost the same setup. I simply run the bare bones notebook [here](https://www.kaggle.com/ryches/memory-leak) and get a memory leak no matter what I change. I have played around with pytorch and python versions but all seem to have the leak. If anyone knows a fix that would be greatly appreciated!",
    "1069578": "You can try to install dual boot setup and then install ubuntu and also install nvidia cudatoolkit according to the gpu in your system.You can easily run and can adjust numworkers according to the **os.cpu_count()**.For RAM usage usage you can use **top** command in your ubuntu terminal and **watch -n0.1  nvidia-smi** for your GPU usage as well as for its utilization",
    "1069827": "I try to iterate through the AgentDataset without pytorch DataLoader, the memory leak remains.\n```\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open(cached=False)\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\nfor i in range(4000):\n        data = [train_dataset[k + i * batch_size] for k in range(batch_size)]\n```\nSo the issue is from l5kit.\n\n![l5kit memory leak as a function of time](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F1d9e81e4fda4cb216115ba8cd73e6707%2F20201104%20l5kit%20memory%20leak%20in%20api.png?generation=1604539018730841&alt=media)\n\n![l5kit memory leak as a function of iteration](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1010129%2F16fb971257d81e7c9c36bb355033e188%2F20201104%20l5kit%20memory%20leak%20in%20api%202.png?generation=1604539028411364&alt=media)",
    "1071785": "After debugging for like 2 days, I think I figured out! \nIt is the random line\n`os.environ[\"BLOSC_NOLOCK\"] = \"1\"`\nin the code\nhttps://github.com/lyft/l5kit/blob/0c5d1c5123ead9f5ce25bea6042f1062667a9fc0/l5kit/l5kit/dataset/select_agents.py#L26\nThis terrible line will modify your system variable secretly whenever you import any module from `l5kit.dataset`.\nOnce you remove this line, on Windows you will no longer see the exploding number of handles created by python."
  },
  "source": "meta"
}