{
  "id": 333614,
  "title": "Decompose image to tiles?",
  "url": "/competitions/hubmap-organ-segmentation/discussion/333614",
  "author_name": "",
  "post_date": "2022-06-27T11:39:53.247800700Z",
  "votes": 23,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Has anyone experimented and observed the benefits of using detailed patches?</p>\n<p><a href=\"https://postimg.cc/jCDNvLs4\" target=\"_blank\"><img src=\"https://i.postimg.cc/KcNNGMww/ftu-tiles.png\" alt=\"ftu-tiles.png\"></a></p>\n<p>so the training batch may look like:<br>\n<a href=\"https://postimg.cc/mch2R1pd\" target=\"_blank\"><img src=\"https://i.postimg.cc/7LnC3g3y/ftu-tiles-batch.png\" alt=\"ftu-tiles-batch.png\"></a></p>\n<p>see related training notebook: <a href=\"https://www.kaggle.com/jirkaborovec/ftus-segm-baseline-flash-unet-on-tiled-images\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/ftus-segm-baseline-flash-unet-on-tiled-images</a></p>",
  "messages": [
    {
      "id": "1835021",
      "postDate": "06/27/2022 11:39:53",
      "content": "<p>Has anyone experimented and observed the benefits of using detailed patches?</p>\n<p><a href=\"https://postimg.cc/jCDNvLs4\" target=\"_blank\"><img src=\"https://i.postimg.cc/KcNNGMww/ftu-tiles.png\" alt=\"ftu-tiles.png\"></a></p>\n<p>so the training batch may look like:<br>\n<a href=\"https://postimg.cc/mch2R1pd\" target=\"_blank\"><img src=\"https://i.postimg.cc/7LnC3g3y/ftu-tiles-batch.png\" alt=\"ftu-tiles-batch.png\"></a></p>\n<p>see related training notebook: <a href=\"https://www.kaggle.com/jirkaborovec/ftus-segm-baseline-flash-unet-on-tiled-images\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/ftus-segm-baseline-flash-unet-on-tiled-images</a></p>",
      "rawMarkdown": "Has anyone experimented and observed the benefits of using detailed patches?\n\n[![ftu-tiles.png](https://i.postimg.cc/KcNNGMww/ftu-tiles.png)](https://postimg.cc/jCDNvLs4)\n\nso the training batch may look like:\n[![ftu-tiles-batch.png](https://i.postimg.cc/7LnC3g3y/ftu-tiles-batch.png)](https://postimg.cc/mch2R1pd)\n\nsee related training notebook: https://www.kaggle.com/jirkaborovec/ftus-segm-baseline-flash-unet-on-tiled-images",
      "votes": null
    },
    {
      "id": "1835529",
      "postDate": "06/27/2022 19:55:28",
      "content": "<p>I know patching is used extensively in 3D segmentation. I plan on using it extensively in this competition.</p>",
      "rawMarkdown": "I know patching is used extensively in 3D segmentation. I plan on using it extensively in this competition.",
      "votes": null
    },
    {
      "id": "1835766",
      "postDate": "06/28/2022 04:07:23",
      "content": "<p>This is what i did in the previous HuBMAP and it worked well. We don't need to permanently create the tiles and save to disk. We can create a dataloader which randomly crops tiles during training. And keep picking a random crop until we get one with a certain percentage of mask (before inputting the crop from dataloader into training model). I wrote about it <a href=\"https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238308\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "This is what i did in the previous HuBMAP and it worked well. We don't need to permanently create the tiles and save to disk. We can create a dataloader which randomly crops tiles during training. And keep picking a random crop until we get one with a certain percentage of mask (before inputting the crop from dataloader into training model). I wrote about it [here][1]\n\n[1]: https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238308",
      "votes": null
    },
    {
      "id": "1838450",
      "postDate": "06/30/2022 14:32:13",
      "content": "<p>Hey Chris, did you check how many tries it usually needed to find <code>random crop with a certain percentage of mask</code>? Shouldn't it create some kind of bottleneck? I tried a similar method once with 3D data in <strong>BraTS</strong> dataset. It was quite slow. Did you do anything special to overcome this bottleneck?</p>",
      "rawMarkdown": "Hey Chris, did you check how many tries it usually needed to find `random crop with a certain percentage of mask`? Shouldn't it create some kind of bottleneck? I tried a similar method once with 3D data in **BraTS** dataset. It was quite slow. Did you do anything special to overcome this bottleneck?",
      "votes": null
    },
    {
      "id": "1838494",
      "postDate": "06/30/2022 15:27:08",
      "content": "<p>I guess if you have a reasonable batch size even empty images (without mask) shall not be a problem…</p>",
      "rawMarkdown": "I guess if you have a reasonable batch size even empty images (without mask) shall not be a problem...",
      "votes": null
    },
    {
      "id": "1838899",
      "postDate": "07/01/2022 01:08:05",
      "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> the trick to speed it up (and avoid bottlenecks) is to read images from disk and decode RLE masks before using the dataloader (and put both into memory). Also we can set a limit of 10 tries:</p>\n<pre><code># PUT EVERYTHING INTO MEMORY RAM\nimages = read_images() #load images from disk into memory\nmasks = decode_RLE() #0's with 1's where mask is\nLIMIT = 10\n\n# DATA LOADER BELOW\nfor j in range(LIMIT):\n    k = pick_image() #returns a number from 0 to len(images)-1\n    a,b,c,d = pick_random_crop(masks[k]) #returns random bounds\n    mask = masks[k][a:b,c:d]\n    if np.sum( mask )&gt;0: break\nimage = images[k][a:b,c:d, :]\n</code></pre>",
      "rawMarkdown": "awsaf49 the trick to speed it up (and avoid bottlenecks) is to read images from disk and decode RLE masks before using the dataloader (and put both into memory). Also we can set a limit of 10 tries:\n\n    # PUT EVERYTHING INTO MEMORY RAM\n    images = read_images() #load images from disk into memory\n    masks = decode_RLE() #0's with 1's where mask is\n    LIMIT = 10\n\n    # DATA LOADER BELOW\n    for j in range(LIMIT):\n        k = pick_image() #returns a number from 0 to len(images)-1\n        a,b,c,d = pick_random_crop(masks[k]) #returns random bounds\n        mask = masks[k][a:b,c:d]\n        if np.sum( mask )>0: break\n    image = images[k][a:b,c:d, :]",
      "votes": null
    },
    {
      "id": "1840710",
      "postDate": "07/02/2022 13:59:15",
      "content": "<p>Hi Chris, thanks for your response. I'm new to large image segmentation so I was curious wouldn't patching these images affect the accuracy of the model, since the segmented foreground FTUs in some organ images (unlike the kidney ones) do not entirely fit even in a large enough patch, say 512, so will this \"not fitting in\" affect the model training in some way?</p>\n<p>Also what if we just follow the usual way and resize the images to a fixed size and then upsample the segmentation map using the nearest interpolation or something similar?</p>",
      "rawMarkdown": "Hi Chris, thanks for your response. I'm new to large image segmentation so I was curious wouldn't patching these images affect the accuracy of the model, since the segmented foreground FTUs in some organ images (unlike the kidney ones) do not entirely fit even in a large enough patch, say 512, so will this \"not fitting in\" affect the model training in some way?\n\nAlso what if we just follow the usual way and resize the images to a fixed size and then upsample the segmentation map using the nearest interpolation or something similar?",
      "votes": null
    },
    {
      "id": "1840758",
      "postDate": "07/02/2022 14:35:25",
      "content": "<p>Deciding crop patch size and crop resize are hyperparameters for us to determine. If your GPU has enough VRAM, you can use patches of size <code>768x768</code> or <code>1024x1024</code> or <code>1536x1536</code>. And you can also resize which effectively gives larger patches for example check this out:</p>\n<pre><code># CROP 1024x1024\nmask = masks[k][a:a+1024, c:c+1024]\nimage = images[k][a:a+1024, c:c+1024, :]\n# RESIZE 512x512\nmask = mask[::2, ::2]\nimage = image[::2, ::2, :]\n</code></pre>\n<p>The code above uses patches of size 1024x1024 then resizes them to patches of size 512x512. In previous HuBMAP, many teams used this trick. They also resized to 256x256 (after cropping 1024x1024) with</p>\n<pre><code># RESIZE 256x256\nmask = mask[::4, ::4]\nimage = image[::4, ::4, :]\n</code></pre>",
      "rawMarkdown": "Deciding crop patch size and crop resize are hyperparameters for us to determine. If your GPU has enough VRAM, you can use patches of size `768x768` or `1024x1024` or `1536x1536`. And you can also resize which effectively gives larger patches for example check this out:\n\n    # CROP 1024x1024\n    mask = masks[k][a:a+1024, c:c+1024]\n    image = images[k][a:a+1024, c:c+1024, :]\n    # RESIZE 512x512\n    mask = mask[::2, ::2]\n    image = image[::2, ::2, :]\n\nThe code above uses patches of size 1024x1024 then resizes them to patches of size 512x512. In previous HuBMAP, many teams used this trick. They also resized to 256x256 (after cropping 1024x1024) with\n\n    # RESIZE 256x256\n    mask = mask[::4, ::4]\n    image = image[::4, ::4, :]",
      "votes": null
    },
    {
      "id": "1842833",
      "postDate": "07/04/2022 10:28:22",
      "content": "<p>2 more ideas is to create a dataloader with a batch consisting of such decomposed segments beloning to 1 image. The other idea is to train several models depending on what exactly segment it is (central, left upper, right bottom, etc)</p>",
      "rawMarkdown": "2 more ideas is to create a dataloader with a batch consisting of such decomposed segments beloning to 1 image. The other idea is to train several models depending on what exactly segment it is (central, left upper, right bottom, etc)",
      "votes": null
    },
    {
      "id": "1884344",
      "postDate": "08/04/2022 11:44:07",
      "content": "<p>If you work with patched data, how is your inference pipeline implemented? My intuitive idea is:<br>\nJust create patches deterministically from upper left to lower right and assemble individual masks by using post-processing.<br>\nHere, however, I am bothered by the fact that marginal areas, where a segmentation goes over two patches, can be inconsistent. One possibility would be to perform the inference with overlapping patches and then merge the results. Do you have any ideas, insights or approaches regarding the inference?</p>",
      "rawMarkdown": "If you work with patched data, how is your inference pipeline implemented? My intuitive idea is:\nJust create patches deterministically from upper left to lower right and assemble individual masks by using post-processing.\nHere, however, I am bothered by the fact that marginal areas, where a segmentation goes over two patches, can be inconsistent. One possibility would be to perform the inference with overlapping patches and then merge the results. Do you have any ideas, insights or approaches regarding the inference?",
      "votes": null
    },
    {
      "id": "1884430",
      "postDate": "08/04/2022 12:26:36",
      "content": "<p>Very good point, that any overlap or transition from one tile to another is not patched, and I am assuming a simple/naive approach… just putting it all together, as you can see in the following kernel with inference at the very end:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/jirkaborovec/ftus-segm-lightning-flash-tiled-inference\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/ftus-segm-lightning-flash-tiled-inference</a></p>\n</blockquote>",
      "rawMarkdown": "Very good point, that any overlap or transition from one tile to another is not patched, and I am assuming a simple/naive approach... just putting it all together, as you can see in the following kernel with inference at the very end:\n\n>https://www.kaggle.com/jirkaborovec/ftus-segm-lightning-flash-tiled-inference",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1835529,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "06/27/2022 19:55:28",
      "content": "<p>I know patching is used extensively in 3D segmentation. I plan on using it extensively in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1835766,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/28/2022 04:07:23",
      "content": "<p>This is what i did in the previous HuBMAP and it worked well. We don't need to permanently create the tiles and save to disk. We can create a dataloader which randomly crops tiles during training. And keep picking a random crop until we get one with a certain percentage of mask (before inputting the crop from dataloader into training model). I wrote about it <a href=\"https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238308\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1838450,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "06/30/2022 14:32:13",
          "content": "<p>Hey Chris, did you check how many tries it usually needed to find <code>random crop with a certain percentage of mask</code>? Shouldn't it create some kind of bottleneck? I tried a similar method once with 3D data in <strong>BraTS</strong> dataset. It was quite slow. Did you do anything special to overcome this bottleneck?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1838494,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "06/30/2022 15:27:08",
          "content": "<p>I guess if you have a reasonable batch size even empty images (without mask) shall not be a problem…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1838899,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/01/2022 01:08:05",
          "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> the trick to speed it up (and avoid bottlenecks) is to read images from disk and decode RLE masks before using the dataloader (and put both into memory). Also we can set a limit of 10 tries:</p>\n<pre><code># PUT EVERYTHING INTO MEMORY RAM\nimages = read_images() #load images from disk into memory\nmasks = decode_RLE() #0's with 1's where mask is\nLIMIT = 10\n\n# DATA LOADER BELOW\nfor j in range(LIMIT):\n    k = pick_image() #returns a number from 0 to len(images)-1\n    a,b,c,d = pick_random_crop(masks[k]) #returns random bounds\n    mask = masks[k][a:b,c:d]\n    if np.sum( mask )&gt;0: break\nimage = images[k][a:b,c:d, :]\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1840710,
          "author_name": "adnanpen",
          "author_url": "",
          "post_date": "07/02/2022 13:59:15",
          "content": "<p>Hi Chris, thanks for your response. I'm new to large image segmentation so I was curious wouldn't patching these images affect the accuracy of the model, since the segmented foreground FTUs in some organ images (unlike the kidney ones) do not entirely fit even in a large enough patch, say 512, so will this \"not fitting in\" affect the model training in some way?</p>\n<p>Also what if we just follow the usual way and resize the images to a fixed size and then upsample the segmentation map using the nearest interpolation or something similar?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1840758,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/02/2022 14:35:25",
          "content": "<p>Deciding crop patch size and crop resize are hyperparameters for us to determine. If your GPU has enough VRAM, you can use patches of size <code>768x768</code> or <code>1024x1024</code> or <code>1536x1536</code>. And you can also resize which effectively gives larger patches for example check this out:</p>\n<pre><code># CROP 1024x1024\nmask = masks[k][a:a+1024, c:c+1024]\nimage = images[k][a:a+1024, c:c+1024, :]\n# RESIZE 512x512\nmask = mask[::2, ::2]\nimage = image[::2, ::2, :]\n</code></pre>\n<p>The code above uses patches of size 1024x1024 then resizes them to patches of size 512x512. In previous HuBMAP, many teams used this trick. They also resized to 256x256 (after cropping 1024x1024) with</p>\n<pre><code># RESIZE 256x256\nmask = mask[::4, ::4]\nimage = image[::4, ::4, :]\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1842833,
      "author_name": "starlighter",
      "author_url": "",
      "post_date": "07/04/2022 10:28:22",
      "content": "<p>2 more ideas is to create a dataloader with a batch consisting of such decomposed segments beloning to 1 image. The other idea is to train several models depending on what exactly segment it is (central, left upper, right bottom, etc)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1884344,
      "author_name": "fabianhorst",
      "author_url": "",
      "post_date": "08/04/2022 11:44:07",
      "content": "<p>If you work with patched data, how is your inference pipeline implemented? My intuitive idea is:<br>\nJust create patches deterministically from upper left to lower right and assemble individual masks by using post-processing.<br>\nHere, however, I am bothered by the fact that marginal areas, where a segmentation goes over two patches, can be inconsistent. One possibility would be to perform the inference with overlapping patches and then merge the results. Do you have any ideas, insights or approaches regarding the inference?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1884430,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "08/04/2022 12:26:36",
          "content": "<p>Very good point, that any overlap or transition from one tile to another is not patched, and I am assuming a simple/naive approach… just putting it all together, as you can see in the following kernel with inference at the very end:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/jirkaborovec/ftus-segm-lightning-flash-tiled-inference\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/ftus-segm-lightning-flash-tiled-inference</a></p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1835021": "Has anyone experimented and observed the benefits of using detailed patches?\n\n[![ftu-tiles.png](https://i.postimg.cc/KcNNGMww/ftu-tiles.png)](https://postimg.cc/jCDNvLs4)\n\nso the training batch may look like:\n[![ftu-tiles-batch.png](https://i.postimg.cc/7LnC3g3y/ftu-tiles-batch.png)](https://postimg.cc/mch2R1pd)\n\nsee related training notebook: https://www.kaggle.com/jirkaborovec/ftus-segm-baseline-flash-unet-on-tiled-images",
    "1835529": "I know patching is used extensively in 3D segmentation. I plan on using it extensively in this competition.",
    "1835766": "This is what i did in the previous HuBMAP and it worked well. We don't need to permanently create the tiles and save to disk. We can create a dataloader which randomly crops tiles during training. And keep picking a random crop until we get one with a certain percentage of mask (before inputting the crop from dataloader into training model). I wrote about it [here][1]\n\n[1]: https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238308",
    "1838450": "Hey Chris, did you check how many tries it usually needed to find `random crop with a certain percentage of mask`? Shouldn't it create some kind of bottleneck? I tried a similar method once with 3D data in **BraTS** dataset. It was quite slow. Did you do anything special to overcome this bottleneck?",
    "1838494": "I guess if you have a reasonable batch size even empty images (without mask) shall not be a problem...",
    "1838899": "awsaf49 the trick to speed it up (and avoid bottlenecks) is to read images from disk and decode RLE masks before using the dataloader (and put both into memory). Also we can set a limit of 10 tries:\n\n    # PUT EVERYTHING INTO MEMORY RAM\n    images = read_images() #load images from disk into memory\n    masks = decode_RLE() #0's with 1's where mask is\n    LIMIT = 10\n\n    # DATA LOADER BELOW\n    for j in range(LIMIT):\n        k = pick_image() #returns a number from 0 to len(images)-1\n        a,b,c,d = pick_random_crop(masks[k]) #returns random bounds\n        mask = masks[k][a:b,c:d]\n        if np.sum( mask )>0: break\n    image = images[k][a:b,c:d, :]",
    "1840710": "Hi Chris, thanks for your response. I'm new to large image segmentation so I was curious wouldn't patching these images affect the accuracy of the model, since the segmented foreground FTUs in some organ images (unlike the kidney ones) do not entirely fit even in a large enough patch, say 512, so will this \"not fitting in\" affect the model training in some way?\n\nAlso what if we just follow the usual way and resize the images to a fixed size and then upsample the segmentation map using the nearest interpolation or something similar?",
    "1840758": "Deciding crop patch size and crop resize are hyperparameters for us to determine. If your GPU has enough VRAM, you can use patches of size `768x768` or `1024x1024` or `1536x1536`. And you can also resize which effectively gives larger patches for example check this out:\n\n    # CROP 1024x1024\n    mask = masks[k][a:a+1024, c:c+1024]\n    image = images[k][a:a+1024, c:c+1024, :]\n    # RESIZE 512x512\n    mask = mask[::2, ::2]\n    image = image[::2, ::2, :]\n\nThe code above uses patches of size 1024x1024 then resizes them to patches of size 512x512. In previous HuBMAP, many teams used this trick. They also resized to 256x256 (after cropping 1024x1024) with\n\n    # RESIZE 256x256\n    mask = mask[::4, ::4]\n    image = image[::4, ::4, :]",
    "1842833": "2 more ideas is to create a dataloader with a batch consisting of such decomposed segments beloning to 1 image. The other idea is to train several models depending on what exactly segment it is (central, left upper, right bottom, etc)",
    "1884344": "If you work with patched data, how is your inference pipeline implemented? My intuitive idea is:\nJust create patches deterministically from upper left to lower right and assemble individual masks by using post-processing.\nHere, however, I am bothered by the fact that marginal areas, where a segmentation goes over two patches, can be inconsistent. One possibility would be to perform the inference with overlapping patches and then merge the results. Do you have any ideas, insights or approaches regarding the inference?",
    "1884430": "Very good point, that any overlap or transition from one tile to another is not patched, and I am assuming a simple/naive approach... just putting it all together, as you can see in the following kernel with inference at the very end:\n\n>https://www.kaggle.com/jirkaborovec/ftus-segm-lightning-flash-tiled-inference"
  },
  "source": "meta"
}