{
  "id": 241915,
  "title": "How to get accumulated gradients without lowering the resolution using PyTorch?",
  "url": "/competitions/siim-covid19-detection/discussion/241915",
  "author_name": "",
  "post_date": "2021-05-26T15:35:46.277047600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Lowering the resolution will give worse results as we might miss some fine details of lungs. What I'm hoping is to break one high resolution image into patches and feed them to the network sequentially to get accumulated gradients.</p>\n<p>Let me explain a bit more. Let's use a 3000x3000 example image:</p>\n<ol>\n<li>break the image into 10 patches.</li>\n<li>pass each patch into the model and save gradients</li>\n<li>Now, take all 10 patches gradients and do one backward pass.</li>\n</ol>\n<p>can anyone please tell me how we can do this using PyTorch?<br>\nThere is something called accumulated gradients. I'm not sure if that is related to this method?</p>",
  "messages": [
    {
      "id": "1324068",
      "postDate": "05/26/2021 15:35:46",
      "content": "<p>Lowering the resolution will give worse results as we might miss some fine details of lungs. What I'm hoping is to break one high resolution image into patches and feed them to the network sequentially to get accumulated gradients.</p>\n<p>Let me explain a bit more. Let's use a 3000x3000 example image:</p>\n<ol>\n<li>break the image into 10 patches.</li>\n<li>pass each patch into the model and save gradients</li>\n<li>Now, take all 10 patches gradients and do one backward pass.</li>\n</ol>\n<p>can anyone please tell me how we can do this using PyTorch?<br>\nThere is something called accumulated gradients. I'm not sure if that is related to this method?</p>",
      "rawMarkdown": "Lowering the resolution will give worse results as we might miss some fine details of lungs. What I'm hoping is to break one high resolution image into patches and feed them to the network sequentially to get accumulated gradients.\n\nLet me explain a bit more. Let's use a 3000x3000 example image:\n1. break the image into 10 patches.\n2. pass each patch into the model and save gradients\n3. Now, take all 10 patches gradients and do one backward pass.\n\ncan anyone please tell me how we can do this using PyTorch?\nThere is something called accumulated gradients. I'm not sure if that is related to this method?",
      "votes": null
    },
    {
      "id": "1324409",
      "postDate": "05/26/2021 22:03:39",
      "content": "<p>Um what you're describing sounds like just using smaller images (patches) and a larger batch size… unless you want to use so many examples that you can't fit them all on your GPU. Then yes, you can use gradient accumulation, which is very simple: just call optimizer.step() every n forward passes.</p>",
      "rawMarkdown": "Um what you're describing sounds like just using smaller images (patches) and a larger batch size... unless you want to use so many examples that you can't fit them all on your GPU. Then yes, you can use gradient accumulation, which is very simple: just call optimizer.step() every n forward passes.",
      "votes": null
    },
    {
      "id": "1324413",
      "postDate": "05/26/2021 22:10:32",
      "content": "<p>As far as I know accumulated gradients works different way.<br>\nTo save the memory instead using batch size 16 you can use batch size 8, or batch size 4, or 2, or 1.<br>\nThe problem is with smaller batch size the gradients are different.<br>\nTo fix it you do following: use small batch, then another small batch, then another one, etc… and don't execute loss.backward and optimizer.step after each batch but accumulate gradients and do it after multiple batches.<br>\nSo for example you can use batch_size=4 4 times and it will be like using batch_size = 16</p>",
      "rawMarkdown": "As far as I know accumulated gradients works different way.\nTo save the memory instead using batch size 16 you can use batch size 8, or batch size 4, or 2, or 1.\nThe problem is with smaller batch size the gradients are different.\nTo fix it you do following: use small batch, then another small batch, then another one, etc... and don't execute loss.backward and optimizer.step after each batch but accumulate gradients and do it after multiple batches.\nSo for example you can use batch_size=4 4 times and it will be like using batch_size = 16",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1324409,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "05/26/2021 22:03:39",
      "content": "<p>Um what you're describing sounds like just using smaller images (patches) and a larger batch size… unless you want to use so many examples that you can't fit them all on your GPU. Then yes, you can use gradient accumulation, which is very simple: just call optimizer.step() every n forward passes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1324413,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "05/26/2021 22:10:32",
      "content": "<p>As far as I know accumulated gradients works different way.<br>\nTo save the memory instead using batch size 16 you can use batch size 8, or batch size 4, or 2, or 1.<br>\nThe problem is with smaller batch size the gradients are different.<br>\nTo fix it you do following: use small batch, then another small batch, then another one, etc… and don't execute loss.backward and optimizer.step after each batch but accumulate gradients and do it after multiple batches.<br>\nSo for example you can use batch_size=4 4 times and it will be like using batch_size = 16</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1324068": "Lowering the resolution will give worse results as we might miss some fine details of lungs. What I'm hoping is to break one high resolution image into patches and feed them to the network sequentially to get accumulated gradients.\n\nLet me explain a bit more. Let's use a 3000x3000 example image:\n1. break the image into 10 patches.\n2. pass each patch into the model and save gradients\n3. Now, take all 10 patches gradients and do one backward pass.\n\ncan anyone please tell me how we can do this using PyTorch?\nThere is something called accumulated gradients. I'm not sure if that is related to this method?",
    "1324409": "Um what you're describing sounds like just using smaller images (patches) and a larger batch size... unless you want to use so many examples that you can't fit them all on your GPU. Then yes, you can use gradient accumulation, which is very simple: just call optimizer.step() every n forward passes.",
    "1324413": "As far as I know accumulated gradients works different way.\nTo save the memory instead using batch size 16 you can use batch size 8, or batch size 4, or 2, or 1.\nThe problem is with smaller batch size the gradients are different.\nTo fix it you do following: use small batch, then another small batch, then another one, etc... and don't execute loss.backward and optimizer.step after each batch but accumulate gradients and do it after multiple batches.\nSo for example you can use batch_size=4 4 times and it will be like using batch_size = 16"
  },
  "source": "meta"
}