{
  "id": 348618,
  "title": "Training on large images/patches tips?",
  "url": "/competitions/hubmap-organ-segmentation/discussion/348618",
  "author_name": "",
  "post_date": "2022-08-29T11:04:43.617491500Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have a bunch of pipelines on small patches (512, 768) that give decent baseline solutions (0.75+). But all of them are about tiling original image into small pieces.</p>\n<p>When I try to apply exact same solutions into larger image tiles or whole image rescale (1024, 1536 etc), a lot of problems arises.</p>\n<p>Images hardly fits on GPU memory, and the maximum batch_size is 1-2 in such cases. But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2 (and after that can even decrease back to 0.02).<br>\nThe same situation happens for UNet-based and transformer-based solutions.</p>\n<p>Of course, something is wrong here.</p>\n<p>I've tried playing with (virtual) batch_size and learning rate, but without any success. I did not notice any improvement from here, it's still not converging, and there's no intuition about what values these parameters should have.</p>\n<p>Can you suggest what might be going wrong? </p>\n<p>If the model is overfitting, then for some reason the model (i.e. UNet) cannot remember even the examples from the training set. Should I add more augmentations or soften them?</p>\n<p>Maybe there are some good articles/recommendations on training models on very large images?</p>",
  "messages": [
    {
      "id": "1918148",
      "postDate": "08/29/2022 11:04:43",
      "content": "<p>I have a bunch of pipelines on small patches (512, 768) that give decent baseline solutions (0.75+). But all of them are about tiling original image into small pieces.</p>\n<p>When I try to apply exact same solutions into larger image tiles or whole image rescale (1024, 1536 etc), a lot of problems arises.</p>\n<p>Images hardly fits on GPU memory, and the maximum batch_size is 1-2 in such cases. But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2 (and after that can even decrease back to 0.02).<br>\nThe same situation happens for UNet-based and transformer-based solutions.</p>\n<p>Of course, something is wrong here.</p>\n<p>I've tried playing with (virtual) batch_size and learning rate, but without any success. I did not notice any improvement from here, it's still not converging, and there's no intuition about what values these parameters should have.</p>\n<p>Can you suggest what might be going wrong? </p>\n<p>If the model is overfitting, then for some reason the model (i.e. UNet) cannot remember even the examples from the training set. Should I add more augmentations or soften them?</p>\n<p>Maybe there are some good articles/recommendations on training models on very large images?</p>",
      "rawMarkdown": "I have a bunch of pipelines on small patches (512, 768) that give decent baseline solutions (0.75+). But all of them are about tiling original image into small pieces.\n\nWhen I try to apply exact same solutions into larger image tiles or whole image rescale (1024, 1536 etc), a lot of problems arises.\n\nImages hardly fits on GPU memory, and the maximum batch_size is 1-2 in such cases. But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2 (and after that can even decrease back to 0.02).\nThe same situation happens for UNet-based and transformer-based solutions.\n\nOf course, something is wrong here.\n\nI've tried playing with (virtual) batch_size and learning rate, but without any success. I did not notice any improvement from here, it's still not converging, and there's no intuition about what values these parameters should have.\n\nCan you suggest what might be going wrong? \n\nIf the model is overfitting, then for some reason the model (i.e. UNet) cannot remember even the examples from the training set. Should I add more augmentations or soften them?\n\nMaybe there are some good articles/recommendations on training models on very large images?",
      "votes": null
    },
    {
      "id": "1918177",
      "postDate": "08/29/2022 11:19:49",
      "content": "<p>Some other discussions about large images<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941#1907369\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941#1907369</a> </p>",
      "rawMarkdown": "Some other discussions about large images\nhttps://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941#1907369",
      "votes": null
    },
    {
      "id": "1918286",
      "postDate": "08/29/2022 13:11:17",
      "content": "<blockquote>\n  <blockquote>\n    <p>But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2</p>\n  </blockquote>\n</blockquote>\n<p>each model have a \"receptive field\", it determines the the \"size of the region\" a model can see.</p>\n<p>when you enlarge the image, for the \"same region size\"  the model actually see \"less\".  </p>\n<p>assume the receptive field for an image is NxN pixels, which is 10% x 10% of the image <br>\nfor an 2x enlarged image, the receptive field is still NxN pixels, but now looking at only 5%x5% of the enlarged image.</p>\n<p>when the model \"see less\", there is too little information to made prediction.</p>\n<p>large image needs deeper model (for larger receptive field ).<br>\nif you check the image net papers, etc you will see larger models uses larger images to get better results.</p>\n<hr>\n<p>note:</p>\n<ul>\n<li>it doesn't make sense if you can infinitely get better performance by just increasing image size for a fixed model. if so, we need not design new model … just increase image size.</li>\n</ul>\n<hr>\n<blockquote>\n  <blockquote>\n    <p>Images hardly fits on GPU memory, and the maximum batch_size is 1-2 i</p>\n  </blockquote>\n</blockquote>\n<p>there is a min batch size requirement to train model. too large or too small batch size doesn't work</p>",
      "rawMarkdown": ">>But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2\n\neach model have a \"receptive field\", it determines the the \"size of the region\" a model can see.\n\nwhen you enlarge the image, for the \"same region size\"  the model actually see \"less\".  \n\nassume the receptive field for an image is NxN pixels, which is 10% x 10% of the image \nfor an 2x enlarged image, the receptive field is still NxN pixels, but now looking at only 5%x5% of the enlarged image.\n\nwhen the model \"see less\", there is too little information to made prediction.\n\nlarge image needs deeper model (for larger receptive field ).\nif you check the image net papers, etc you will see larger models uses larger images to get better results.\n\n\n---\n\nnote:\n- it doesn't make sense if you can infinitely get better performance by just increasing image size for a fixed model. if so, we need not design new model ... just increase image size.\n\n\n----\n>>Images hardly fits on GPU memory, and the maximum batch_size is 1-2 i\n\nthere is a min batch size requirement to train model. too large or too small batch size doesn't work",
      "votes": null
    },
    {
      "id": "1918369",
      "postDate": "08/29/2022 14:42:49",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for your detailed explanation</p>\n<p>Receptive field definitely stays the same</p>\n<p>Working setup:<br>\nEffUNet<br>\nOriginal image 3000 x 3000 &gt; scale 0.5 &gt; tiling with size 768x768<br>\nInput for training: patches 768 x 768, spatial size of pixel is 0.8 um<br>\nReceptive field for every output mask pixel is N x N</p>\n<p>For example, I'm replacing it with<br>\nEffUNet <br>\nOriginal image 3000 x 3000 &gt; scale 0.5 (no tiling)<br>\nInput for training: images 1500 x 1500, spatial size of pixel is still 0.8 um<br>\nReceptive field for every output mask pixel is still N x N (because we didn't make any changes to the model)</p>\n<p>So the model sees the same spatial region as before at every mask pixel</p>\n<p>What worries me is that although we did not benefit from using a larger size, we should not have drastically ruined something</p>",
      "rawMarkdown": "hengck23 Thank you for your detailed explanation\n\nReceptive field definitely stays the same\n\nWorking setup:\nEffUNet\nOriginal image 3000 x 3000 > scale 0.5 > tiling with size 768x768\nInput for training: patches 768 x 768, spatial size of pixel is 0.8 um\nReceptive field for every output mask pixel is N x N\n\nFor example, I'm replacing it with\nEffUNet \nOriginal image 3000 x 3000 > scale 0.5 (no tiling)\nInput for training: images 1500 x 1500, spatial size of pixel is still 0.8 um\nReceptive field for every output mask pixel is still N x N (because we didn't make any changes to the model)\n\nSo the model sees the same spatial region as before at every mask pixel\n\nWhat worries me is that although we did not benefit from using a larger size, we should not have drastically ruined something",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1918177,
      "author_name": "celidos",
      "author_url": "",
      "post_date": "08/29/2022 11:19:49",
      "content": "<p>Some other discussions about large images<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941#1907369\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941#1907369</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1918286,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/29/2022 13:11:17",
      "content": "<blockquote>\n  <blockquote>\n    <p>But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2</p>\n  </blockquote>\n</blockquote>\n<p>each model have a \"receptive field\", it determines the the \"size of the region\" a model can see.</p>\n<p>when you enlarge the image, for the \"same region size\"  the model actually see \"less\".  </p>\n<p>assume the receptive field for an image is NxN pixels, which is 10% x 10% of the image <br>\nfor an 2x enlarged image, the receptive field is still NxN pixels, but now looking at only 5%x5% of the enlarged image.</p>\n<p>when the model \"see less\", there is too little information to made prediction.</p>\n<p>large image needs deeper model (for larger receptive field ).<br>\nif you check the image net papers, etc you will see larger models uses larger images to get better results.</p>\n<hr>\n<p>note:</p>\n<ul>\n<li>it doesn't make sense if you can infinitely get better performance by just increasing image size for a fixed model. if so, we need not design new model … just increase image size.</li>\n</ul>\n<hr>\n<blockquote>\n  <blockquote>\n    <p>Images hardly fits on GPU memory, and the maximum batch_size is 1-2 i</p>\n  </blockquote>\n</blockquote>\n<p>there is a min batch size requirement to train model. too large or too small batch size doesn't work</p>",
      "votes": null,
      "replies": [
        {
          "id": 1918369,
          "author_name": "celidos",
          "author_url": "",
          "post_date": "08/29/2022 14:42:49",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for your detailed explanation</p>\n<p>Receptive field definitely stays the same</p>\n<p>Working setup:<br>\nEffUNet<br>\nOriginal image 3000 x 3000 &gt; scale 0.5 &gt; tiling with size 768x768<br>\nInput for training: patches 768 x 768, spatial size of pixel is 0.8 um<br>\nReceptive field for every output mask pixel is N x N</p>\n<p>For example, I'm replacing it with<br>\nEffUNet <br>\nOriginal image 3000 x 3000 &gt; scale 0.5 (no tiling)<br>\nInput for training: images 1500 x 1500, spatial size of pixel is still 0.8 um<br>\nReceptive field for every output mask pixel is still N x N (because we didn't make any changes to the model)</p>\n<p>So the model sees the same spatial region as before at every mask pixel</p>\n<p>What worries me is that although we did not benefit from using a larger size, we should not have drastically ruined something</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1918148": "I have a bunch of pipelines on small patches (512, 768) that give decent baseline solutions (0.75+). But all of them are about tiling original image into small pieces.\n\nWhen I try to apply exact same solutions into larger image tiles or whole image rescale (1024, 1536 etc), a lot of problems arises.\n\nImages hardly fits on GPU memory, and the maximum batch_size is 1-2 in such cases. But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2 (and after that can even decrease back to 0.02).\nThe same situation happens for UNet-based and transformer-based solutions.\n\nOf course, something is wrong here.\n\nI've tried playing with (virtual) batch_size and learning rate, but without any success. I did not notice any improvement from here, it's still not converging, and there's no intuition about what values these parameters should have.\n\nCan you suggest what might be going wrong? \n\nIf the model is overfitting, then for some reason the model (i.e. UNet) cannot remember even the examples from the training set. Should I add more augmentations or soften them?\n\nMaybe there are some good articles/recommendations on training models on very large images?",
    "1918177": "Some other discussions about large images\nhttps://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941#1907369",
    "1918286": ">>But what's worse, the solution doesn't converge at all, and CV dice score stuck at 0.1 - 0.2\n\neach model have a \"receptive field\", it determines the the \"size of the region\" a model can see.\n\nwhen you enlarge the image, for the \"same region size\"  the model actually see \"less\".  \n\nassume the receptive field for an image is NxN pixels, which is 10% x 10% of the image \nfor an 2x enlarged image, the receptive field is still NxN pixels, but now looking at only 5%x5% of the enlarged image.\n\nwhen the model \"see less\", there is too little information to made prediction.\n\nlarge image needs deeper model (for larger receptive field ).\nif you check the image net papers, etc you will see larger models uses larger images to get better results.\n\n\n---\n\nnote:\n- it doesn't make sense if you can infinitely get better performance by just increasing image size for a fixed model. if so, we need not design new model ... just increase image size.\n\n\n----\n>>Images hardly fits on GPU memory, and the maximum batch_size is 1-2 i\n\nthere is a min batch size requirement to train model. too large or too small batch size doesn't work",
    "1918369": "hengck23 Thank you for your detailed explanation\n\nReceptive field definitely stays the same\n\nWorking setup:\nEffUNet\nOriginal image 3000 x 3000 > scale 0.5 > tiling with size 768x768\nInput for training: patches 768 x 768, spatial size of pixel is 0.8 um\nReceptive field for every output mask pixel is N x N\n\nFor example, I'm replacing it with\nEffUNet \nOriginal image 3000 x 3000 > scale 0.5 (no tiling)\nInput for training: images 1500 x 1500, spatial size of pixel is still 0.8 um\nReceptive field for every output mask pixel is still N x N (because we didn't make any changes to the model)\n\nSo the model sees the same spatial region as before at every mask pixel\n\nWhat worries me is that although we did not benefit from using a larger size, we should not have drastically ruined something"
  },
  "source": "meta"
}