{
  "id": 195303,
  "title": "How does num_workers work?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/195303",
  "author_name": "",
  "post_date": "2020-11-04T15:01:36.999547500Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>I'm still a bit confused about how pytorch parameter num_workers work.</p>\n<p>From what I learned, num_workers 0 only uses the main process to load data. num_workers &gt; 0 uses multiple processes to load data. Therefore, it prepares the data used for later batches and loads them into RAM so that later we don't have to prepare data. Is my understanding correct?</p>\n<p>I've been heard that CPU for data preparing is a key bottleneck for this competition, and it seems that this is the case given that 1x2080Ti gives a similar speed as 1xQuadro8000 with a larger batch size. Therefore, we tried to play with num_workers value. However, num_workers values of 4 and 32 do not really result in very different speed, which contradicts with my understanding of num_workers.</p>\n<p>Is this a normal thing? Is there some issue with my understanding of num_workers?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1069526",
      "postDate": "11/04/2020 15:01:37",
      "content": "<p>Hi all!</p>\n<p>I'm still a bit confused about how pytorch parameter num_workers work.</p>\n<p>From what I learned, num_workers 0 only uses the main process to load data. num_workers &gt; 0 uses multiple processes to load data. Therefore, it prepares the data used for later batches and loads them into RAM so that later we don't have to prepare data. Is my understanding correct?</p>\n<p>I've been heard that CPU for data preparing is a key bottleneck for this competition, and it seems that this is the case given that 1x2080Ti gives a similar speed as 1xQuadro8000 with a larger batch size. Therefore, we tried to play with num_workers value. However, num_workers values of 4 and 32 do not really result in very different speed, which contradicts with my understanding of num_workers.</p>\n<p>Is this a normal thing? Is there some issue with my understanding of num_workers?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi all!\n\nI'm still a bit confused about how pytorch parameter num_workers work.\n\nFrom what I learned, num_workers 0 only uses the main process to load data. num_workers > 0 uses multiple processes to load data. Therefore, it prepares the data used for later batches and loads them into RAM so that later we don't have to prepare data. Is my understanding correct?\n\nI've been heard that CPU for data preparing is a key bottleneck for this competition, and it seems that this is the case given that 1x2080Ti gives a similar speed as 1xQuadro8000 with a larger batch size. Therefore, we tried to play with num_workers value. However, num_workers values of 4 and 32 do not really result in very different speed, which contradicts with my understanding of num_workers.\n\nIs this a normal thing? Is there some issue with my understanding of num_workers?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1069686",
      "postDate": "11/04/2020 19:02:20",
      "content": "<p>\". However, num_workers values of 4 and 32 do not really result in very different speed\"</p>\n<p>i think retrieval of one dataset[i] already uses multi-process. that may be the reason why increase num_workers doesn't seems to help. ideally setting  num_workers  &gt; batch size should help.</p>\n<p>pytorch seems to have a \"bug\". if one process get stuck and do not return, the dataloader will wait infinitely long (but there is a timeout argument in dataloader to stop the python process for this)</p>\n<p>caching each dataset[i] record to file (or memory) seems to be only viable option for this. if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10</p>",
      "rawMarkdown": "\". However, num_workers values of 4 and 32 do not really result in very different speed\"\n\ni think retrieval of one dataset[i] already uses multi-process. that may be the reason why increase num\\_workers doesn't seems to help. ideally setting  num\\_workers  > batch size should help.\n\npytorch seems to have a \"bug\". if one process get stuck and do not return, the dataloader will wait infinitely long (but there is a timeout argument in dataloader to stop the python process for this)\n\ncaching each dataset[i] record to file (or memory) seems to be only viable option for this. if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10",
      "votes": null
    },
    {
      "id": "1069769",
      "postDate": "11/04/2020 22:32:29",
      "content": "<p><code>if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10.</code>@hengck23 Good idea for better using GPU or TPU. Could you share total disk size needed for train.zarr?</p>",
      "rawMarkdown": "`if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10. `@hengck23 Good idea for better using GPU or TPU. Could you share total disk size needed for train.zarr?",
      "votes": null
    },
    {
      "id": "1069884",
      "postDate": "11/05/2020 03:38:57",
      "content": "<p>You can try comment out the training code inside the loop and view the CPU usage to find the bottleneck of the dataloader, if 32 worker still could not consume all the CPU power, you may consider the Disk IO are too slow. </p>",
      "rawMarkdown": "You can try comment out the training code inside the loop and view the CPU usage to find the bottleneck of the dataloader, if 32 worker still could not consume all the CPU power, you may consider the Disk IO are too slow.",
      "votes": null
    },
    {
      "id": "1071404",
      "postDate": "11/06/2020 20:47:35",
      "content": "<p>I can confirm the same on kaggle, the speed does not change for more than 4 workers for me. </p>",
      "rawMarkdown": "I can confirm the same on kaggle, the speed does not change for more than 4 workers for me.",
      "votes": null
    },
    {
      "id": "1071474",
      "postDate": "11/06/2020 23:42:23",
      "content": "<p>My 2080Ti never use more than 1% through out the training.<br>\nIt can also be the DISK IO. What kind of disk speed de you have? Do you see the CPU 100% utilized? Or, is it heavy more on the disk?</p>",
      "rawMarkdown": "My 2080Ti never use more than 1% through out the training.\nIt can also be the DISK IO. What kind of disk speed de you have? Do you see the CPU 100% utilized? Or, is it heavy more on the disk?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1069686,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/04/2020 19:02:20",
      "content": "<p>\". However, num_workers values of 4 and 32 do not really result in very different speed\"</p>\n<p>i think retrieval of one dataset[i] already uses multi-process. that may be the reason why increase num_workers doesn't seems to help. ideally setting  num_workers  &gt; batch size should help.</p>\n<p>pytorch seems to have a \"bug\". if one process get stuck and do not return, the dataloader will wait infinitely long (but there is a timeout argument in dataloader to stop the python process for this)</p>\n<p>caching each dataset[i] record to file (or memory) seems to be only viable option for this. if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10</p>",
      "votes": null,
      "replies": [
        {
          "id": 1069769,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "11/04/2020 22:32:29",
          "content": "<p><code>if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10.</code>@hengck23 Good idea for better using GPU or TPU. Could you share total disk size needed for train.zarr?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1069884,
      "author_name": "steamedsheep",
      "author_url": "",
      "post_date": "11/05/2020 03:38:57",
      "content": "<p>You can try comment out the training code inside the loop and view the CPU usage to find the bottleneck of the dataloader, if 32 worker still could not consume all the CPU power, you may consider the Disk IO are too slow. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1071404,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "11/06/2020 20:47:35",
      "content": "<p>I can confirm the same on kaggle, the speed does not change for more than 4 workers for me. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1071474,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/06/2020 23:42:23",
      "content": "<p>My 2080Ti never use more than 1% through out the training.<br>\nIt can also be the DISK IO. What kind of disk speed de you have? Do you see the CPU 100% utilized? Or, is it heavy more on the disk?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1069526": "Hi all!\n\nI'm still a bit confused about how pytorch parameter num_workers work.\n\nFrom what I learned, num_workers 0 only uses the main process to load data. num_workers > 0 uses multiple processes to load data. Therefore, it prepares the data used for later batches and loads them into RAM so that later we don't have to prepare data. Is my understanding correct?\n\nI've been heard that CPU for data preparing is a key bottleneck for this competition, and it seems that this is the case given that 1x2080Ti gives a similar speed as 1xQuadro8000 with a larger batch size. Therefore, we tried to play with num_workers value. However, num_workers values of 4 and 32 do not really result in very different speed, which contradicts with my understanding of num_workers.\n\nIs this a normal thing? Is there some issue with my understanding of num_workers?\n\nThanks!",
    "1069686": "\". However, num_workers values of 4 and 32 do not really result in very different speed\"\n\ni think retrieval of one dataset[i] already uses multi-process. that may be the reason why increase num\\_workers doesn't seems to help. ideally setting  num\\_workers  > batch size should help.\n\npytorch seems to have a \"bug\". if one process get stuck and do not return, the dataloader will wait infinitely long (but there is a timeout argument in dataloader to stop the python process for this)\n\ncaching each dataset[i] record to file (or memory) seems to be only viable option for this. if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10",
    "1069769": "`if you save image as np.uint8, this results in 1.3 mb per record for 224x224 for historyframe=10. `@hengck23 Good idea for better using GPU or TPU. Could you share total disk size needed for train.zarr?",
    "1069884": "You can try comment out the training code inside the loop and view the CPU usage to find the bottleneck of the dataloader, if 32 worker still could not consume all the CPU power, you may consider the Disk IO are too slow.",
    "1071404": "I can confirm the same on kaggle, the speed does not change for more than 4 workers for me.",
    "1071474": "My 2080Ti never use more than 1% through out the training.\nIt can also be the DISK IO. What kind of disk speed de you have? Do you see the CPU 100% utilized? Or, is it heavy more on the disk?"
  },
  "source": "meta"
}