{
  "id": 291773,
  "title": "What is \"This training is CPU bounding\" meaning ?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/291773",
  "author_name": "",
  "post_date": "2021-12-01T03:49:19.349149900Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In detectron2 <br>\nIt seems GPU is not busy <br>\nsome people said </p>\n<p>\"this is due to data generator cannot sample enough to make GPU Busy \"<br>\n\"Training is CPU bounding\"<br>\n\"TPU will not make improvement\"</p>\n<p>what is  this meaning?<br>\nAny idea for approaches to fix ?<br>\nHow to make full use of GPU?</p>",
  "messages": [
    {
      "id": "1601136",
      "postDate": "12/01/2021 03:49:19",
      "content": "<p>In detectron2 <br>\nIt seems GPU is not busy <br>\nsome people said </p>\n<p>\"this is due to data generator cannot sample enough to make GPU Busy \"<br>\n\"Training is CPU bounding\"<br>\n\"TPU will not make improvement\"</p>\n<p>what is  this meaning?<br>\nAny idea for approaches to fix ?<br>\nHow to make full use of GPU?</p>",
      "rawMarkdown": "In detectron2 \nIt seems GPU is not busy \nsome people said \n\n\"this is due to data generator cannot sample enough to make GPU Busy \"\n\"Training is CPU bounding\"\n\"TPU will not make improvement\"\n\nwhat is  this meaning?\nAny idea for approaches to fix ?\nHow to make full use of GPU?",
      "votes": null
    },
    {
      "id": "1601403",
      "postDate": "12/01/2021 08:48:45",
      "content": "<p>I guess using more workers will help.</p>",
      "rawMarkdown": "I guess using more workers will help.",
      "votes": null
    },
    {
      "id": "1601418",
      "postDate": "12/01/2021 09:03:20",
      "content": "<p>CPU and GPU work together passing data from one to another. Usually CPU does loading/preprocessing/augmenting input data, converting it all to tensors so it can be copied to GPU memory. Then GPU passes it through the model, computes losses and updates weights. </p>\n<p>This process repeats for each batch of the data. Ideally things happen in parallel -While GPU processes one batch the CPU already prepares the next. However if the dataloading takes longer than the training step the GPU has to wait idle for the next batch to be ready. That’s what we call CPU bound and changing to a faster GPU won’t affect the total training time.</p>\n<p>How to fix it? If you are using a custom machine you can play with your setup, use more CPU cores, faster processor, faster disk, more memory etc. if that’s not possible (for example on colab/kaggle) the only thing left is to make the preprocessing code faster. I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.</p>",
      "rawMarkdown": "CPU and GPU work together passing data from one to another. Usually CPU does loading/preprocessing/augmenting input data, converting it all to tensors so it can be copied to GPU memory. Then GPU passes it through the model, computes losses and updates weights. \n\nThis process repeats for each batch of the data. Ideally things happen in parallel -While GPU processes one batch the CPU already prepares the next. However if the dataloading takes longer than the training step the GPU has to wait idle for the next batch to be ready. That’s what we call CPU bound and changing to a faster GPU won’t affect the total training time.\n\nHow to fix it? If you are using a custom machine you can play with your setup, use more CPU cores, faster processor, faster disk, more memory etc. if that’s not possible (for example on colab/kaggle) the only thing left is to make the preprocessing code faster. I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.",
      "votes": null
    },
    {
      "id": "1603407",
      "postDate": "12/02/2021 13:53:15",
      "content": "<p>pipeline bottleneck</p>",
      "rawMarkdown": "pipeline bottleneck",
      "votes": null
    },
    {
      "id": "1608261",
      "postDate": "12/06/2021 10:22:57",
      "content": "<p>Thank you for your very delicate reply. </p>\n<p><strong>&gt; I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.</strong></p>\n<p>I faced this problem and knew your mentioned solution, but I can not find the code. Can you share this notebook that has the above solution? Thank you in advance!</p>",
      "rawMarkdown": "Thank you for your very delicate reply. \n\n**> I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.**\n\nI faced this problem and knew your mentioned solution, but I can not find the code. Can you share this notebook that has the above solution? Thank you in advance!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1601403,
      "author_name": "plugin1689",
      "author_url": "",
      "post_date": "12/01/2021 08:48:45",
      "content": "<p>I guess using more workers will help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1601418,
      "author_name": "slawekbiel",
      "author_url": "",
      "post_date": "12/01/2021 09:03:20",
      "content": "<p>CPU and GPU work together passing data from one to another. Usually CPU does loading/preprocessing/augmenting input data, converting it all to tensors so it can be copied to GPU memory. Then GPU passes it through the model, computes losses and updates weights. </p>\n<p>This process repeats for each batch of the data. Ideally things happen in parallel -While GPU processes one batch the CPU already prepares the next. However if the dataloading takes longer than the training step the GPU has to wait idle for the next batch to be ready. That’s what we call CPU bound and changing to a faster GPU won’t affect the total training time.</p>\n<p>How to fix it? If you are using a custom machine you can play with your setup, use more CPU cores, faster processor, faster disk, more memory etc. if that’s not possible (for example on colab/kaggle) the only thing left is to make the preprocessing code faster. I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1608261,
          "author_name": "aengusng",
          "author_url": "",
          "post_date": "12/06/2021 10:22:57",
          "content": "<p>Thank you for your very delicate reply. </p>\n<p><strong>&gt; I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.</strong></p>\n<p>I faced this problem and knew your mentioned solution, but I can not find the code. Can you share this notebook that has the above solution? Thank you in advance!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1603407,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "12/02/2021 13:53:15",
      "content": "<p>pipeline bottleneck</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1601136": "In detectron2 \nIt seems GPU is not busy \nsome people said \n\n\"this is due to data generator cannot sample enough to make GPU Busy \"\n\"Training is CPU bounding\"\n\"TPU will not make improvement\"\n\nwhat is  this meaning?\nAny idea for approaches to fix ?\nHow to make full use of GPU?",
    "1601403": "I guess using more workers will help.",
    "1601418": "CPU and GPU work together passing data from one to another. Usually CPU does loading/preprocessing/augmenting input data, converting it all to tensors so it can be copied to GPU memory. Then GPU passes it through the model, computes losses and updates weights. \n\nThis process repeats for each batch of the data. Ideally things happen in parallel -While GPU processes one batch the CPU already prepares the next. However if the dataloading takes longer than the training step the GPU has to wait idle for the next batch to be ready. That’s what we call CPU bound and changing to a faster GPU won’t affect the total training time.\n\nHow to fix it? If you are using a custom machine you can play with your setup, use more CPU cores, faster processor, faster disk, more memory etc. if that’s not possible (for example on colab/kaggle) the only thing left is to make the preprocessing code faster. I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.",
    "1603407": "pipeline bottleneck",
    "1608261": "Thank you for your very delicate reply. \n\n**> I’ve seen a notebook here of caching images in memory to speed up loading so that’s a possible avenue.**\n\nI faced this problem and knew your mentioned solution, but I can not find the code. Can you share this notebook that has the above solution? Thank you in advance!"
  },
  "source": "meta"
}