{
  "id": 195936,
  "title": "How to solve pytorch DataLoader parallelization issue on Windows",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/195936",
  "author_name": "",
  "post_date": "2020-11-08T11:14:24.527640Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Currently, when you run pytorch <code>DataLoader</code> with l5kit's <code>AgentDataset</code> and <code>num_workers &gt; 0</code> on Windows, you will encounter the error like</p>\n<pre><code>Can't pickle google.protobuf.pyext._message.RepeatedCompositeContainer objects\n</code></pre>\n<p>when you train the model. </p>\n<p>I found a workaround by making a Dataset wrapper which will only load the actual dataset and rasterizer after the wrapper Dataset has been loaded to the workers.<br>\nPut this into a <code>mydataset.py</code> in the same folder of your model jupyter notebook:</p>\n<pre><code>class MyTrainDataset:\n    def __init__(self, cfg, dm):\n        self.cfg = cfg\n        self.dm = dm\n        self.has_init = False\n    def initialize(self, worker_id):\n        print('initialize called with worker_id', worker_id)\n        from l5kit.data import ChunkedDataset\n        from l5kit.dataset import AgentDataset #, EgoDataset\n        from l5kit.rasterization import build_rasterizer\n        rasterizer = build_rasterizer(self.cfg, self.dm)\n        train_cfg = self.cfg[\"train_data_loader\"]\n        train_zarr = ChunkedDataset(self.dm.require(train_cfg[\"key\"])).open(cached=False)  # try to turn off cache\n        self.dataset = AgentDataset(self.cfg, train_zarr, rasterizer)\n        self.has_init = True\n    def reset(self):\n        self.dataset = None\n        self.has_init = False\n    def __len__(self):\n        # note you have to figure out the actual length beforehand since once the rasterizer and/or AgentDataset been constructed, you cannot pickle it anymore! So we can't compute the size from the real dataset. However, DataLoader require the len to determine the sampling.\n        return 22496709\n    def __getitem__(self, index):\n        return self.dataset[index]    \n\nfrom torch.utils.data import get_worker_info\ndef my_dataset_worker_init_func(worker_id):\n    worker_info = get_worker_info()\n    dataset = worker_info.dataset\n    dataset.initialize(worker_id)\n</code></pre>\n<p>Then in your jupyter notebook, you do</p>\n<pre><code>from mydataset import MyTrainDataset, my_dataset_worker_init_func\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_dataset = MyTrainDataset(cfg, dm)\ntrain_dataloader = DataLoader(\n    train_dataset,\n    shuffle=train_cfg[\"shuffle\"], \n    batch_size=16,\n    num_workers=2,\n    persistent_workers=True,\n    worker_init_fn=my_dataset_worker_init_func,\n)\ntr_it = iter(train_dataloader)\n</code></pre>\n<p>The downside is that each worker process will try to load the entire pytorch library, which somehow take like 4GB of commit memory. During training, each worker can take upto 7GB of commit memory. Since I only got 32GB of ram, with the virtual memory turn on, I can only run 4 workers. Nevertheless, now I got 4 times faster training.</p>\n<p>Reference: example 2 of <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.IterableDataset\" target=\"_blank\">https://pytorch.org/docs/stable/data.html#torch.utils.data.IterableDataset</a></p>\n<p>Edit: Note that to close the workers and release the memory, you can run<br>\n<code>tr_it._shutdown_workers()</code></p>",
  "messages": [
    {
      "id": "1072521",
      "postDate": "11/08/2020 11:14:24",
      "content": "<p>Currently, when you run pytorch <code>DataLoader</code> with l5kit's <code>AgentDataset</code> and <code>num_workers &gt; 0</code> on Windows, you will encounter the error like</p>\n<pre><code>Can't pickle google.protobuf.pyext._message.RepeatedCompositeContainer objects\n</code></pre>\n<p>when you train the model. </p>\n<p>I found a workaround by making a Dataset wrapper which will only load the actual dataset and rasterizer after the wrapper Dataset has been loaded to the workers.<br>\nPut this into a <code>mydataset.py</code> in the same folder of your model jupyter notebook:</p>\n<pre><code>class MyTrainDataset:\n    def __init__(self, cfg, dm):\n        self.cfg = cfg\n        self.dm = dm\n        self.has_init = False\n    def initialize(self, worker_id):\n        print('initialize called with worker_id', worker_id)\n        from l5kit.data import ChunkedDataset\n        from l5kit.dataset import AgentDataset #, EgoDataset\n        from l5kit.rasterization import build_rasterizer\n        rasterizer = build_rasterizer(self.cfg, self.dm)\n        train_cfg = self.cfg[\"train_data_loader\"]\n        train_zarr = ChunkedDataset(self.dm.require(train_cfg[\"key\"])).open(cached=False)  # try to turn off cache\n        self.dataset = AgentDataset(self.cfg, train_zarr, rasterizer)\n        self.has_init = True\n    def reset(self):\n        self.dataset = None\n        self.has_init = False\n    def __len__(self):\n        # note you have to figure out the actual length beforehand since once the rasterizer and/or AgentDataset been constructed, you cannot pickle it anymore! So we can't compute the size from the real dataset. However, DataLoader require the len to determine the sampling.\n        return 22496709\n    def __getitem__(self, index):\n        return self.dataset[index]    \n\nfrom torch.utils.data import get_worker_info\ndef my_dataset_worker_init_func(worker_id):\n    worker_info = get_worker_info()\n    dataset = worker_info.dataset\n    dataset.initialize(worker_id)\n</code></pre>\n<p>Then in your jupyter notebook, you do</p>\n<pre><code>from mydataset import MyTrainDataset, my_dataset_worker_init_func\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_dataset = MyTrainDataset(cfg, dm)\ntrain_dataloader = DataLoader(\n    train_dataset,\n    shuffle=train_cfg[\"shuffle\"], \n    batch_size=16,\n    num_workers=2,\n    persistent_workers=True,\n    worker_init_fn=my_dataset_worker_init_func,\n)\ntr_it = iter(train_dataloader)\n</code></pre>\n<p>The downside is that each worker process will try to load the entire pytorch library, which somehow take like 4GB of commit memory. During training, each worker can take upto 7GB of commit memory. Since I only got 32GB of ram, with the virtual memory turn on, I can only run 4 workers. Nevertheless, now I got 4 times faster training.</p>\n<p>Reference: example 2 of <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.IterableDataset\" target=\"_blank\">https://pytorch.org/docs/stable/data.html#torch.utils.data.IterableDataset</a></p>\n<p>Edit: Note that to close the workers and release the memory, you can run<br>\n<code>tr_it._shutdown_workers()</code></p>",
      "rawMarkdown": "Currently, when you run pytorch `DataLoader` with l5kit's `AgentDataset` and `num_workers > 0` on Windows, you will encounter the error like\n```\nCan't pickle google.protobuf.pyext._message.RepeatedCompositeContainer objects\n```\nwhen you train the model. \n\nI found a workaround by making a Dataset wrapper which will only load the actual dataset and rasterizer after the wrapper Dataset has been loaded to the workers.\nPut this into a `mydataset.py` in the same folder of your model jupyter notebook:\n```python\nclass MyTrainDataset:\n    def __init__(self, cfg, dm):\n        self.cfg = cfg\n        self.dm = dm\n        self.has_init = False\n    def initialize(self, worker_id):\n        print('initialize called with worker_id', worker_id)\n        from l5kit.data import ChunkedDataset\n        from l5kit.dataset import AgentDataset #, EgoDataset\n        from l5kit.rasterization import build_rasterizer\n        rasterizer = build_rasterizer(self.cfg, self.dm)\n        train_cfg = self.cfg[\"train_data_loader\"]\n        train_zarr = ChunkedDataset(self.dm.require(train_cfg[\"key\"])).open(cached=False)  # try to turn off cache\n        self.dataset = AgentDataset(self.cfg, train_zarr, rasterizer)\n        self.has_init = True\n    def reset(self):\n        self.dataset = None\n        self.has_init = False\n    def __len__(self):\n        # note you have to figure out the actual length beforehand since once the rasterizer and/or AgentDataset been constructed, you cannot pickle it anymore! So we can't compute the size from the real dataset. However, DataLoader require the len to determine the sampling.\n        return 22496709\n    def __getitem__(self, index):\n        return self.dataset[index]    \n\nfrom torch.utils.data import get_worker_info\ndef my_dataset_worker_init_func(worker_id):\n    worker_info = get_worker_info()\n    dataset = worker_info.dataset\n    dataset.initialize(worker_id)\n```\nThen in your jupyter notebook, you do\n```\nfrom mydataset import MyTrainDataset, my_dataset_worker_init_func\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_dataset = MyTrainDataset(cfg, dm)\ntrain_dataloader = DataLoader(\n    train_dataset,\n    shuffle=train_cfg[\"shuffle\"], \n    batch_size=16,\n    num_workers=2,\n    persistent_workers=True,\n    worker_init_fn=my_dataset_worker_init_func,\n)\ntr_it = iter(train_dataloader)\n```\nThe downside is that each worker process will try to load the entire pytorch library, which somehow take like 4GB of commit memory. During training, each worker can take upto 7GB of commit memory. Since I only got 32GB of ram, with the virtual memory turn on, I can only run 4 workers. Nevertheless, now I got 4 times faster training.\n\nReference: example 2 of https://pytorch.org/docs/stable/data.html#torch.utils.data.IterableDataset\n\nEdit: Note that to close the workers and release the memory, you can run\n```tr_it._shutdown_workers()```",
      "votes": null
    },
    {
      "id": "1077124",
      "postDate": "11/13/2020 09:01:12",
      "content": "<p>The <code>prefetch_factor</code> actually help a lot. I increase that to <code>prefetch_factor=8</code> with <code>num_workers=4</code> help a lot! </p>",
      "rawMarkdown": "The `prefetch_factor` actually help a lot. I increase that to `prefetch_factor=8` with `num_workers=4` help a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1077124,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/13/2020 09:01:12",
      "content": "<p>The <code>prefetch_factor</code> actually help a lot. I increase that to <code>prefetch_factor=8</code> with <code>num_workers=4</code> help a lot! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1072521": "Currently, when you run pytorch `DataLoader` with l5kit's `AgentDataset` and `num_workers > 0` on Windows, you will encounter the error like\n```\nCan't pickle google.protobuf.pyext._message.RepeatedCompositeContainer objects\n```\nwhen you train the model. \n\nI found a workaround by making a Dataset wrapper which will only load the actual dataset and rasterizer after the wrapper Dataset has been loaded to the workers.\nPut this into a `mydataset.py` in the same folder of your model jupyter notebook:\n```python\nclass MyTrainDataset:\n    def __init__(self, cfg, dm):\n        self.cfg = cfg\n        self.dm = dm\n        self.has_init = False\n    def initialize(self, worker_id):\n        print('initialize called with worker_id', worker_id)\n        from l5kit.data import ChunkedDataset\n        from l5kit.dataset import AgentDataset #, EgoDataset\n        from l5kit.rasterization import build_rasterizer\n        rasterizer = build_rasterizer(self.cfg, self.dm)\n        train_cfg = self.cfg[\"train_data_loader\"]\n        train_zarr = ChunkedDataset(self.dm.require(train_cfg[\"key\"])).open(cached=False)  # try to turn off cache\n        self.dataset = AgentDataset(self.cfg, train_zarr, rasterizer)\n        self.has_init = True\n    def reset(self):\n        self.dataset = None\n        self.has_init = False\n    def __len__(self):\n        # note you have to figure out the actual length beforehand since once the rasterizer and/or AgentDataset been constructed, you cannot pickle it anymore! So we can't compute the size from the real dataset. However, DataLoader require the len to determine the sampling.\n        return 22496709\n    def __getitem__(self, index):\n        return self.dataset[index]    \n\nfrom torch.utils.data import get_worker_info\ndef my_dataset_worker_init_func(worker_id):\n    worker_info = get_worker_info()\n    dataset = worker_info.dataset\n    dataset.initialize(worker_id)\n```\nThen in your jupyter notebook, you do\n```\nfrom mydataset import MyTrainDataset, my_dataset_worker_init_func\ntrain_cfg = cfg[\"train_data_loader\"]\ntrain_dataset = MyTrainDataset(cfg, dm)\ntrain_dataloader = DataLoader(\n    train_dataset,\n    shuffle=train_cfg[\"shuffle\"], \n    batch_size=16,\n    num_workers=2,\n    persistent_workers=True,\n    worker_init_fn=my_dataset_worker_init_func,\n)\ntr_it = iter(train_dataloader)\n```\nThe downside is that each worker process will try to load the entire pytorch library, which somehow take like 4GB of commit memory. During training, each worker can take upto 7GB of commit memory. Since I only got 32GB of ram, with the virtual memory turn on, I can only run 4 workers. Nevertheless, now I got 4 times faster training.\n\nReference: example 2 of https://pytorch.org/docs/stable/data.html#torch.utils.data.IterableDataset\n\nEdit: Note that to close the workers and release the memory, you can run\n```tr_it._shutdown_workers()```",
    "1077124": "The `prefetch_factor` actually help a lot. I increase that to `prefetch_factor=8` with `num_workers=4` help a lot!"
  },
  "source": "meta"
}