{
  "id": 475117,
  "title": "13th place solution: 4-panel solo model",
  "url": "/competitions/blood-vessel-segmentation/writeups/menno-13th-place-solution-4-panel-solo-model",
  "author_name": "",
  "post_date": "2024-02-10T12:45:34.490Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First I would like to thank the host for this wonderful competition, I really enjoyed participating in it! I would also like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for his valuable and interesting comments, I learned a lot from this! As my solution was inspired by the <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\" target=\"_blank\">winning solution</a> of the contrails competition 6 months ago, I want to thank <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> as well! </p>\n<p><strong>4-panel image</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5960015%2F1cc75a5e44c828cbb2b05af1a2d91558%2FSenNetHOA23_4Panel_example.png?generation=1707287151880105&amp;alt=media\"></p>\n<p>Inspired by <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> winning solution, the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512px by 512px images). Therefore, I was able to keep 2.5D dimensionality in a 2D image. The idea was that the model would learn this relationship between slices and therefore would predict a more continuous segmentation.<br>\nThe 256px patches were made with the <a href=\"https://github.com/Mr-TalhaIlyas/EMPatches\" target=\"_blank\">EMPatches</a> library. Which made stitching the separate 256px patches from each 4-panel image a lot easier.<br>\nThe images were normalized based on the percentile of the whole kidney. All values lower than 0.5 were clamped to 0.5, to normalize the background better:</p>\n<pre><code>lo, hi = .percentile(kidney_volume.numpy(), (, ))\ndef preprocess_image(, lo, hi):\n     = .to(torch.float32)\n     = ( - lo) / (hi - lo)\n     = torch.clamp(, =)\n     \n</code></pre>\n<p><strong>Augmentations</strong><br>\nI used simple augmentations for training: </p>\n<pre><code>train_transform = A.Compose([\n    A.RandomRotate90(=1),\n    A.HorizontalFlip(=0.5),\n    A.VerticalFlip(=0.5),\n    A.RandomBrightness(=1),\n    A.OneOf(\n        [\n            A.Blur(=3, =1),\n            A.MotionBlur(=3, =1),\n        ],\n        =0.9,\n    ),\n])\n</code></pre>\n<p>What did not work for me were the augmentations based on scaling of the image.</p>\n<p><strong>Submission</strong><br>\nFor the submission part: each patch in the 4-panel image was rotated on each own. Where after, the mean was taken of the rotated patches. The mean was taken as well of all the patches in the separate 4-panel images. These patches were merged with the ‘max’ setting. This was performed for the xy, xz, yz rotations of the whole kidney volume.</p>\n<p><strong>Model</strong><br>\nThe model that was trained on these 4-panel images was a Unet maxvit_tiny_tf_512 using segmentation models pytorch (SMP). The model was trained on 3 whole kidney volume rotations with 0.4 overlap in the patches (~490.000 different images). The model was trained for 9 epochs with a 1e-4 lr and then another 6 epochs with CosineAnnealingLR to 1e-6.</p>\n<ul>\n<li>Link to Kaggle submission notebook: <a href=\"https://www.kaggle.com/code/menno1111/sennet-hoa23-2-5d-4-panel-submission/\" target=\"_blank\">SenNet-HOA23 | 2.5D 4-panel | submission</a></li>\n<li>Link to the training and validation notebooks: <a href=\"https://github.com/Menno-Meijer/SenNet_VasculatureSegmentation_Competition\" target=\"_blank\">GitHub</a></li>\n</ul>",
  "messages": [
    {
      "id": "2640902",
      "postDate": "02/07/2024 06:31:08",
      "content": "<p>First I would like to thank the host for this wonderful competition, I really enjoyed participating in it! I would also like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for his valuable and interesting comments, I learned a lot from this! As my solution was inspired by the <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\" target=\"_blank\">winning solution</a> of the contrails competition 6 months ago, I want to thank <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> as well! </p>\n<p><strong>4-panel image</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5960015%2F1cc75a5e44c828cbb2b05af1a2d91558%2FSenNetHOA23_4Panel_example.png?generation=1707287151880105&amp;alt=media\"></p>\n<p>Inspired by <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> winning solution, the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512px by 512px images). Therefore, I was able to keep 2.5D dimensionality in a 2D image. The idea was that the model would learn this relationship between slices and therefore would predict a more continuous segmentation.<br>\nThe 256px patches were made with the <a href=\"https://github.com/Mr-TalhaIlyas/EMPatches\" target=\"_blank\">EMPatches</a> library. Which made stitching the separate 256px patches from each 4-panel image a lot easier.<br>\nThe images were normalized based on the percentile of the whole kidney. All values lower than 0.5 were clamped to 0.5, to normalize the background better:</p>\n<pre><code>lo, hi = .percentile(kidney_volume.numpy(), (, ))\ndef preprocess_image(, lo, hi):\n     = .to(torch.float32)\n     = ( - lo) / (hi - lo)\n     = torch.clamp(, =)\n     \n</code></pre>\n<p><strong>Augmentations</strong><br>\nI used simple augmentations for training: </p>\n<pre><code>train_transform = A.Compose([\n    A.RandomRotate90(=1),\n    A.HorizontalFlip(=0.5),\n    A.VerticalFlip(=0.5),\n    A.RandomBrightness(=1),\n    A.OneOf(\n        [\n            A.Blur(=3, =1),\n            A.MotionBlur(=3, =1),\n        ],\n        =0.9,\n    ),\n])\n</code></pre>\n<p>What did not work for me were the augmentations based on scaling of the image.</p>\n<p><strong>Submission</strong><br>\nFor the submission part: each patch in the 4-panel image was rotated on each own. Where after, the mean was taken of the rotated patches. The mean was taken as well of all the patches in the separate 4-panel images. These patches were merged with the ‘max’ setting. This was performed for the xy, xz, yz rotations of the whole kidney volume.</p>\n<p><strong>Model</strong><br>\nThe model that was trained on these 4-panel images was a Unet maxvit_tiny_tf_512 using segmentation models pytorch (SMP). The model was trained on 3 whole kidney volume rotations with 0.4 overlap in the patches (~490.000 different images). The model was trained for 9 epochs with a 1e-4 lr and then another 6 epochs with CosineAnnealingLR to 1e-6.</p>\n<ul>\n<li>Link to Kaggle submission notebook: <a href=\"https://www.kaggle.com/code/menno1111/sennet-hoa23-2-5d-4-panel-submission/\" target=\"_blank\">SenNet-HOA23 | 2.5D 4-panel | submission</a></li>\n<li>Link to the training and validation notebooks: <a href=\"https://github.com/Menno-Meijer/SenNet_VasculatureSegmentation_Competition\" target=\"_blank\">GitHub</a></li>\n</ul>",
      "rawMarkdown": "First I would like to thank the host for this wonderful competition, I really enjoyed participating in it! I would also like to thank @hengck23 for his valuable and interesting comments, I learned a lot from this! As my solution was inspired by the [winning solution](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618) of the contrails competition 6 months ago, I want to thank @junkoda as well! \n\n**4-panel image**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5960015%2F1cc75a5e44c828cbb2b05af1a2d91558%2FSenNetHOA23_4Panel_example.png?generation=1707287151880105&alt=media)\n\nInspired by @junkoda winning solution, the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512px by 512px images). Therefore, I was able to keep 2.5D dimensionality in a 2D image. The idea was that the model would learn this relationship between slices and therefore would predict a more continuous segmentation.\nThe 256px patches were made with the [EMPatches](https://github.com/Mr-TalhaIlyas/EMPatches) library. Which made stitching the separate 256px patches from each 4-panel image a lot easier.\nThe images were normalized based on the percentile of the whole kidney. All values lower than 0.5 were clamped to 0.5, to normalize the background better:\n```\nlo, hi = np.percentile(kidney_volume.numpy(), (2, 98))\ndef preprocess_image(image, lo, hi):\n    image = image.to(torch.float32)\n    image = (image - lo) / (hi - lo)\n    image = torch.clamp(image, min=0.5)\n    return image\n```\n\n**Augmentations**\nI used simple augmentations for training: \n```\ntrain_transform = A.Compose([\n    A.RandomRotate90(p=1),\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomBrightness(p=1),\n    A.OneOf(\n        [\n            A.Blur(blur_limit=3, p=1),\n            A.MotionBlur(blur_limit=3, p=1),\n        ],\n        p=0.9,\n    ),\n])\n```\nWhat did not work for me were the augmentations based on scaling of the image.\n\n**Submission**\nFor the submission part: each patch in the 4-panel image was rotated on each own. Where after, the mean was taken of the rotated patches. The mean was taken as well of all the patches in the separate 4-panel images. These patches were merged with the ‘max’ setting. This was performed for the xy, xz, yz rotations of the whole kidney volume.\n\n**Model**\nThe model that was trained on these 4-panel images was a Unet maxvit_tiny_tf_512 using segmentation models pytorch (SMP). The model was trained on 3 whole kidney volume rotations with 0.4 overlap in the patches (~490.000 different images). The model was trained for 9 epochs with a 1e-4 lr and then another 6 epochs with CosineAnnealingLR to 1e-6.\n\n* Link to Kaggle submission notebook: [SenNet-HOA23 | 2.5D 4-panel | submission](https://www.kaggle.com/code/menno1111/sennet-hoa23-2-5d-4-panel-submission/)\n* Link to the training and validation notebooks: [GitHub](https://github.com/Menno-Meijer/SenNet_VasculatureSegmentation_Competition)",
      "votes": null
    },
    {
      "id": "2641428",
      "postDate": "02/07/2024 13:33:22",
      "content": "<p>Thanks for sharing.</p>\n<ol>\n<li>have you compared the difference in cv and lb with and without using this 4 panel image?</li>\n2. \n</ol>\n<blockquote>\n  <p>the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512 by 512 images). </p>\n</blockquote>\n<p>just to confirm so the model input is 512x512, and during inference, you probably do sliding window of 256x256, make the panel and feed the 512 into the model?</p>\n<ol>\n3. \n</ol>\n<blockquote>\n  <p>These patches were merged with the ‘max’ setting. <br>\n  Why did you choose <code>max</code>? e.g. i felt <code>mean</code> would be <em>safer</em>?</p>\n</blockquote>\n<p>4.For inference do you do overlaps? and how long did it take you for lb to score?<br>\n5.Any thoughts on the difference between have 4 consecutive slices as channel vs a panel in one image?</p>",
      "rawMarkdown": "Thanks for sharing.\n\n1. have you compared the difference in cv and lb with and without using this 4 panel image?\n2. \n>the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512 by 512 images). \n\njust to confirm so the model input is 512x512, and during inference, you probably do sliding window of 256x256, make the panel and feed the 512 into the model?\n3. \n>These patches were merged with the ‘max’ setting. \nWhy did you choose `max`? e.g. i felt `mean` would be _safer_?\n\n4.For inference do you do overlaps? and how long did it take you for lb to score?\n5.Any thoughts on the difference between have 4 consecutive slices as channel vs a panel in one image?",
      "votes": null
    },
    {
      "id": "2642007",
      "postDate": "02/07/2024 20:27:32",
      "content": "<ol>\n<li>Yes, both cv and lb scores without the 4 panel (256px patches separate) were worse. That is why I pursuit this 4-panel image model further.</li>\n<li>Yes that is correct! Another nice thing to add to that: most sliding windows are seen by the model 4 times, once in each quarter of the 4-panel image. If that makes sense.</li>\n<li>I think this is a setting that I should have tested further, but due to time and other priorities I did not. Setting the merge option on mean or max did not change my cv or lb score too much. My cv score was slightly worse with max and the lb slightly improved with max. Why I eventually went for the max setting was the idea that it would better predict blood vessels cut off on the edge of the patch. If the model would be unable to predict the blood vessels that were cut off by the edge and in an overlapping patch it would be able to predict the segmentation of these blood vessels, taking the mean of these predictions would lower the probability. With the risk that these blood vessels would be excluded after thresholding. But taking with the max you would keep the highest probabilities.</li>\n<li>Yes for the interference I put the overlap on 0.1. I did not see and difference by putting it on 0.2. However this was when I still had the setting to merge the patches by mean. Unfortunately, I ran out of time and submissions to test it with max hahah:) It took around 7 hours to score on the lb.</li>\n<li>I did not try using the slices as channels, mostly because of the comments saying that a normal 2D model outperformed such 2.5D model. Therefore, I cannot say if such a model would have outperformed this 4-panel model.</li>\n</ol>\n<p>Hope this answers your questions!</p>",
      "rawMarkdown": "1. Yes, both cv and lb scores without the 4 panel (256px patches separate) were worse. That is why I pursuit this 4-panel image model further.\n2. Yes that is correct! Another nice thing to add to that: most sliding windows are seen by the model 4 times, once in each quarter of the 4-panel image. If that makes sense.\n3. I think this is a setting that I should have tested further, but due to time and other priorities I did not. Setting the merge option on mean or max did not change my cv or lb score too much. My cv score was slightly worse with max and the lb slightly improved with max. Why I eventually went for the max setting was the idea that it would better predict blood vessels cut off on the edge of the patch. If the model would be unable to predict the blood vessels that were cut off by the edge and in an overlapping patch it would be able to predict the segmentation of these blood vessels, taking the mean of these predictions would lower the probability. With the risk that these blood vessels would be excluded after thresholding. But taking with the max you would keep the highest probabilities.\n4. Yes for the interference I put the overlap on 0.1. I did not see and difference by putting it on 0.2. However this was when I still had the setting to merge the patches by mean. Unfortunately, I ran out of time and submissions to test it with max hahah:) It took around 7 hours to score on the lb.\n5. I did not try using the slices as channels, mostly because of the comments saying that a normal 2D model outperformed such 2.5D model. Therefore, I cannot say if such a model would have outperformed this 4-panel model.\n\nHope this answers your questions!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2641428,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "02/07/2024 13:33:22",
      "content": "<p>Thanks for sharing.</p>\n<ol>\n<li>have you compared the difference in cv and lb with and without using this 4 panel image?</li>\n2. \n</ol>\n<blockquote>\n  <p>the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512 by 512 images). </p>\n</blockquote>\n<p>just to confirm so the model input is 512x512, and during inference, you probably do sliding window of 256x256, make the panel and feed the 512 into the model?</p>\n<ol>\n3. \n</ol>\n<blockquote>\n  <p>These patches were merged with the ‘max’ setting. <br>\n  Why did you choose <code>max</code>? e.g. i felt <code>mean</code> would be <em>safer</em>?</p>\n</blockquote>\n<p>4.For inference do you do overlaps? and how long did it take you for lb to score?<br>\n5.Any thoughts on the difference between have 4 consecutive slices as channel vs a panel in one image?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2642007,
          "author_name": "menno1111",
          "author_url": "",
          "post_date": "02/07/2024 20:27:32",
          "content": "<ol>\n<li>Yes, both cv and lb scores without the 4 panel (256px patches separate) were worse. That is why I pursuit this 4-panel image model further.</li>\n<li>Yes that is correct! Another nice thing to add to that: most sliding windows are seen by the model 4 times, once in each quarter of the 4-panel image. If that makes sense.</li>\n<li>I think this is a setting that I should have tested further, but due to time and other priorities I did not. Setting the merge option on mean or max did not change my cv or lb score too much. My cv score was slightly worse with max and the lb slightly improved with max. Why I eventually went for the max setting was the idea that it would better predict blood vessels cut off on the edge of the patch. If the model would be unable to predict the blood vessels that were cut off by the edge and in an overlapping patch it would be able to predict the segmentation of these blood vessels, taking the mean of these predictions would lower the probability. With the risk that these blood vessels would be excluded after thresholding. But taking with the max you would keep the highest probabilities.</li>\n<li>Yes for the interference I put the overlap on 0.1. I did not see and difference by putting it on 0.2. However this was when I still had the setting to merge the patches by mean. Unfortunately, I ran out of time and submissions to test it with max hahah:) It took around 7 hours to score on the lb.</li>\n<li>I did not try using the slices as channels, mostly because of the comments saying that a normal 2D model outperformed such 2.5D model. Therefore, I cannot say if such a model would have outperformed this 4-panel model.</li>\n</ol>\n<p>Hope this answers your questions!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2640902": "First I would like to thank the host for this wonderful competition, I really enjoyed participating in it! I would also like to thank @hengck23 for his valuable and interesting comments, I learned a lot from this! As my solution was inspired by the [winning solution](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618) of the contrails competition 6 months ago, I want to thank @junkoda as well! \n\n**4-panel image**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5960015%2F1cc75a5e44c828cbb2b05af1a2d91558%2FSenNetHOA23_4Panel_example.png?generation=1707287151880105&alt=media)\n\nInspired by @junkoda winning solution, the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512px by 512px images). Therefore, I was able to keep 2.5D dimensionality in a 2D image. The idea was that the model would learn this relationship between slices and therefore would predict a more continuous segmentation.\nThe 256px patches were made with the [EMPatches](https://github.com/Mr-TalhaIlyas/EMPatches) library. Which made stitching the separate 256px patches from each 4-panel image a lot easier.\nThe images were normalized based on the percentile of the whole kidney. All values lower than 0.5 were clamped to 0.5, to normalize the background better:\n```\nlo, hi = np.percentile(kidney_volume.numpy(), (2, 98))\ndef preprocess_image(image, lo, hi):\n    image = image.to(torch.float32)\n    image = (image - lo) / (hi - lo)\n    image = torch.clamp(image, min=0.5)\n    return image\n```\n\n**Augmentations**\nI used simple augmentations for training: \n```\ntrain_transform = A.Compose([\n    A.RandomRotate90(p=1),\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomBrightness(p=1),\n    A.OneOf(\n        [\n            A.Blur(blur_limit=3, p=1),\n            A.MotionBlur(blur_limit=3, p=1),\n        ],\n        p=0.9,\n    ),\n])\n```\nWhat did not work for me were the augmentations based on scaling of the image.\n\n**Submission**\nFor the submission part: each patch in the 4-panel image was rotated on each own. Where after, the mean was taken of the rotated patches. The mean was taken as well of all the patches in the separate 4-panel images. These patches were merged with the ‘max’ setting. This was performed for the xy, xz, yz rotations of the whole kidney volume.\n\n**Model**\nThe model that was trained on these 4-panel images was a Unet maxvit_tiny_tf_512 using segmentation models pytorch (SMP). The model was trained on 3 whole kidney volume rotations with 0.4 overlap in the patches (~490.000 different images). The model was trained for 9 epochs with a 1e-4 lr and then another 6 epochs with CosineAnnealingLR to 1e-6.\n\n* Link to Kaggle submission notebook: [SenNet-HOA23 | 2.5D 4-panel | submission](https://www.kaggle.com/code/menno1111/sennet-hoa23-2-5d-4-panel-submission/)\n* Link to the training and validation notebooks: [GitHub](https://github.com/Menno-Meijer/SenNet_VasculatureSegmentation_Competition)",
    "2641428": "Thanks for sharing.\n\n1. have you compared the difference in cv and lb with and without using this 4 panel image?\n2. \n>the model was trained on 4-panel images consisting of 256px sized patches of consecutive slices (creating 512 by 512 images). \n\njust to confirm so the model input is 512x512, and during inference, you probably do sliding window of 256x256, make the panel and feed the 512 into the model?\n3. \n>These patches were merged with the ‘max’ setting. \nWhy did you choose `max`? e.g. i felt `mean` would be _safer_?\n\n4.For inference do you do overlaps? and how long did it take you for lb to score?\n5.Any thoughts on the difference between have 4 consecutive slices as channel vs a panel in one image?",
    "2642007": "1. Yes, both cv and lb scores without the 4 panel (256px patches separate) were worse. That is why I pursuit this 4-panel image model further.\n2. Yes that is correct! Another nice thing to add to that: most sliding windows are seen by the model 4 times, once in each quarter of the 4-panel image. If that makes sense.\n3. I think this is a setting that I should have tested further, but due to time and other priorities I did not. Setting the merge option on mean or max did not change my cv or lb score too much. My cv score was slightly worse with max and the lb slightly improved with max. Why I eventually went for the max setting was the idea that it would better predict blood vessels cut off on the edge of the patch. If the model would be unable to predict the blood vessels that were cut off by the edge and in an overlapping patch it would be able to predict the segmentation of these blood vessels, taking the mean of these predictions would lower the probability. With the risk that these blood vessels would be excluded after thresholding. But taking with the max you would keep the highest probabilities.\n4. Yes for the interference I put the overlap on 0.1. I did not see and difference by putting it on 0.2. However this was when I still had the setting to merge the patches by mean. Unfortunately, I ran out of time and submissions to test it with max hahah:) It took around 7 hours to score on the lb.\n5. I did not try using the slices as channels, mostly because of the comments saying that a normal 2D model outperformed such 2.5D model. Therefore, I cannot say if such a model would have outperformed this 4-panel model.\n\nHope this answers your questions!"
  },
  "source": "meta"
}