{
  "id": 478109,
  "title": "41st place solution",
  "url": "/competitions/blood-vessel-segmentation/discussion/478109",
  "author_name": "Jow",
  "post_date": "2024-02-19T07:36:48.792000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Firstly, we would like to express our gratitude to Kaggle and the organizers for hosting this exceptional competition. Through participating in this contest, we have gained a deeper understanding of the challenges and methodologies involved in medical image recognition.</p>\n<h2>Introduction</h2>\n<p>We submitted separate solutions within our team.</p>\n<ul>\n<li>I submitted an ensemble model of se_resnext101_32x4d and Vision Transformer (mit_b2), which achieved a score of <strong>0.834</strong> on the public leaderboard. The private leaderboard score was <strong>0.586</strong>.</li>\n<li><a href=\"https://www.kaggle.com/ryosukesaito\" target=\"_blank\">@ryosukesaito</a> submitted an ensemble model of EfficientNet and SE-ResNeXt which achieved a score of <strong>0.857</strong> on the public leaderboard. The private leaderboard score was <strong>0.519</strong>.</li>\n<li>The high public leaderboard score achieved by <a href=\"https://www.kaggle.com/ryosukesaito\" target=\"_blank\">@ryosukesaito</a>’s submission might have been a contributing factor to our ability to submit my somewhat ambitious notebook, possibly leading to our winning a silver medal.</li>\n</ul>\n<h2>My (@jooott) Solution</h2>\n<h3>Overview</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F580d1b944e38bb1924f6ec8480ac2b65%2F5.PNG?generation=1708327636242492&amp;alt=media\"></p>\n<h3>Key points</h3>\n<p>I struggled significantly with stabilizing the training process.</p>\n<ul>\n<li>To address this, I used Accumulate Grad Batches to effectively increase the batch size to 128, which stabilized the training.</li>\n<li>A major factor in the significant improvement in score was the application of stronger data augmentation. The data augmentation strategy was inspired by <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417496\" target=\"_blank\">the 1st place solution of the Vesuvius Challenge - Ink Detection</a>.</li>\n<li>I also think that scaling up the training images from 512px to 1024px contributed to the increase in score.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fe3e49e0014af22209252039314919b6d%2F10.PNG?generation=1708327710646310&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F01cf2045817c40092ded7e721c4ff52e%2F6.PNG?generation=1708327729992246&amp;alt=media\"></p>\n<pre><code>train_transform = A.Compose(\n            [\n                A.RandomScale(\n                    scale_limit=(1.0, 1.20),\n                    =cv2.INTER_CUBIC,\n                    =0.1,\n                ),\n                A.RandomResizedCrop(\n                    image_size,\n                    image_size,\n                    scale=(0.8, 1.0),\n                    =1\n                ),\n                A.RandomBrightnessContrast(=0.75),\n                A.ShiftScaleRotate(=0.75),\n                A.OneOf([\n                        A.GaussNoise(var_limit=[10, 50]),\n                        A.GaussianBlur(),\n                        A.MotionBlur(),\n                        ], =0.4),\n                A.CoarseDropout(\n                    =1, =int(image_size * 0.1),\n                    =int(image_size * 0.1),\n                    =0, =0.5),\n                A.CLAHE(=0.2),\n                A.GridDistortion(=5, =0.3, =0.05),\n                ToTensorV2(=),\n            ]\n        )\n</code></pre>\n<h2>Muku's (@ryosukesaito) solution</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fa10c8d4428a2dd305f8f94b3bddbb6c7%2FUntitled%20(4).png?generation=1708327848313831&amp;alt=media\"></p>\n<h3>key points</h3>\n<ul>\n<li><p>In my architecture, Detection/Segmentation of kidney region is performed before predicting blood vessel area.</p>\n<ul>\n<li>Detection contributed to inference speedup (especially in the yz/zx direction), since it is possible to skip vessel segmentation in frames where kidney is not detected, and to reduce image size by cropping.</li>\n<li>Segmentation masks were used to reduce FP outside the kidney.</li>\n<li>For both annotations, I used LangSAM <a href=\"https://github.com/luca-medeiros/lang-segment-anything\" target=\"_blank\">(luca-medeiros/lang-segment-anything: SAM with text prompt</a>). This allowed me to prepare annotation data with a few manual adjustments.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F3f044b330cb47140058deac70a7db57c%2FUntitled%20(5).png?generation=1708327892920221&amp;alt=media\"></li>\n<li>I use YOLOv8n for Detection and EfficientNet-B0 for Segmentation.</li></ul></li>\n<li><p>Various pre/post processing improved LB/PB scores slightly, but steadily.</p>\n<ul>\n<li><p>In the yz/zx axis image, blood vessels at the edge may be cut off. Since the inference accuracy was poor in this area, I improved the inference accuracy by pseudo-closing the vessels with mirror-padding before inference.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fb30520c1fb82b034bf913908cf65487d%2FUntitled%20(6).png?generation=1708327958129459&amp;alt=media\"></p></li>\n<li><p>After binarization of the results, defects may occur in the vascular prediction region as shown below. For this reason, morphological closing and fillPoly processing were added as post-processing steps.<br>\nThese contributed to a slight score improvement in CV/LB/PB.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F84de68a4219ff31a9f3623c4e83fc8f2%2FUntitled%20(7).png?generation=1708328014528398&amp;alt=media\"></p></li></ul></li>\n<li><p>In my experiments, ideas that contribute to generalization ability (strong augmentation, pseudo labeling, etc…) could not adopted as final submits, because they resulted in a decrease in CV/LB…<br>\nHowever, I regret that I should not have been too aware of the unstable CV/LB, as the sample was not large enough for this competition.</p></li>\n</ul>",
  "messages": [
    {
      "id": 2658452,
      "postDate": "2024-02-19T07:36:48.793Z",
      "content": "<p>Firstly, we would like to express our gratitude to Kaggle and the organizers for hosting this exceptional competition. Through participating in this contest, we have gained a deeper understanding of the challenges and methodologies involved in medical image recognition.</p>\n<h2>Introduction</h2>\n<p>We submitted separate solutions within our team.</p>\n<ul>\n<li>I submitted an ensemble model of se_resnext101_32x4d and Vision Transformer (mit_b2), which achieved a score of <strong>0.834</strong> on the public leaderboard. The private leaderboard score was <strong>0.586</strong>.</li>\n<li><a href=\"https://www.kaggle.com/ryosukesaito\" target=\"_blank\">@ryosukesaito</a> submitted an ensemble model of EfficientNet and SE-ResNeXt which achieved a score of <strong>0.857</strong> on the public leaderboard. The private leaderboard score was <strong>0.519</strong>.</li>\n<li>The high public leaderboard score achieved by <a href=\"https://www.kaggle.com/ryosukesaito\" target=\"_blank\">@ryosukesaito</a>’s submission might have been a contributing factor to our ability to submit my somewhat ambitious notebook, possibly leading to our winning a silver medal.</li>\n</ul>\n<h2>My (@jooott) Solution</h2>\n<h3>Overview</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F580d1b944e38bb1924f6ec8480ac2b65%2F5.PNG?generation=1708327636242492&amp;alt=media\"></p>\n<h3>Key points</h3>\n<p>I struggled significantly with stabilizing the training process.</p>\n<ul>\n<li>To address this, I used Accumulate Grad Batches to effectively increase the batch size to 128, which stabilized the training.</li>\n<li>A major factor in the significant improvement in score was the application of stronger data augmentation. The data augmentation strategy was inspired by <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417496\" target=\"_blank\">the 1st place solution of the Vesuvius Challenge - Ink Detection</a>.</li>\n<li>I also think that scaling up the training images from 512px to 1024px contributed to the increase in score.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fe3e49e0014af22209252039314919b6d%2F10.PNG?generation=1708327710646310&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F01cf2045817c40092ded7e721c4ff52e%2F6.PNG?generation=1708327729992246&amp;alt=media\"></p>\n<pre><code>train_transform = A.Compose(\n            [\n                A.RandomScale(\n                    scale_limit=(1.0, 1.20),\n                    =cv2.INTER_CUBIC,\n                    =0.1,\n                ),\n                A.RandomResizedCrop(\n                    image_size,\n                    image_size,\n                    scale=(0.8, 1.0),\n                    =1\n                ),\n                A.RandomBrightnessContrast(=0.75),\n                A.ShiftScaleRotate(=0.75),\n                A.OneOf([\n                        A.GaussNoise(var_limit=[10, 50]),\n                        A.GaussianBlur(),\n                        A.MotionBlur(),\n                        ], =0.4),\n                A.CoarseDropout(\n                    =1, =int(image_size * 0.1),\n                    =int(image_size * 0.1),\n                    =0, =0.5),\n                A.CLAHE(=0.2),\n                A.GridDistortion(=5, =0.3, =0.05),\n                ToTensorV2(=),\n            ]\n        )\n</code></pre>\n<h2>Muku's (@ryosukesaito) solution</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fa10c8d4428a2dd305f8f94b3bddbb6c7%2FUntitled%20(4).png?generation=1708327848313831&amp;alt=media\"></p>\n<h3>key points</h3>\n<ul>\n<li><p>In my architecture, Detection/Segmentation of kidney region is performed before predicting blood vessel area.</p>\n<ul>\n<li>Detection contributed to inference speedup (especially in the yz/zx direction), since it is possible to skip vessel segmentation in frames where kidney is not detected, and to reduce image size by cropping.</li>\n<li>Segmentation masks were used to reduce FP outside the kidney.</li>\n<li>For both annotations, I used LangSAM <a href=\"https://github.com/luca-medeiros/lang-segment-anything\" target=\"_blank\">(luca-medeiros/lang-segment-anything: SAM with text prompt</a>). This allowed me to prepare annotation data with a few manual adjustments.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F3f044b330cb47140058deac70a7db57c%2FUntitled%20(5).png?generation=1708327892920221&amp;alt=media\"></li>\n<li>I use YOLOv8n for Detection and EfficientNet-B0 for Segmentation.</li></ul></li>\n<li><p>Various pre/post processing improved LB/PB scores slightly, but steadily.</p>\n<ul>\n<li><p>In the yz/zx axis image, blood vessels at the edge may be cut off. Since the inference accuracy was poor in this area, I improved the inference accuracy by pseudo-closing the vessels with mirror-padding before inference.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fb30520c1fb82b034bf913908cf65487d%2FUntitled%20(6).png?generation=1708327958129459&amp;alt=media\"></p></li>\n<li><p>After binarization of the results, defects may occur in the vascular prediction region as shown below. For this reason, morphological closing and fillPoly processing were added as post-processing steps.<br>\nThese contributed to a slight score improvement in CV/LB/PB.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F84de68a4219ff31a9f3623c4e83fc8f2%2FUntitled%20(7).png?generation=1708328014528398&amp;alt=media\"></p></li></ul></li>\n<li><p>In my experiments, ideas that contribute to generalization ability (strong augmentation, pseudo labeling, etc…) could not adopted as final submits, because they resulted in a decrease in CV/LB…<br>\nHowever, I regret that I should not have been too aware of the unstable CV/LB, as the sample was not large enough for this competition.</p></li>\n</ul>",
      "rawMarkdown": "Firstly, we would like to express our gratitude to Kaggle and the organizers for hosting this exceptional competition. Through participating in this contest, we have gained a deeper understanding of the challenges and methodologies involved in medical image recognition.\n\n## Introduction\nWe submitted separate solutions within our team.\n- I submitted an ensemble model of se_resnext101_32x4d and Vision Transformer (mit_b2), which achieved a score of **0.834** on the public leaderboard. The private leaderboard score was **0.586**.\n- @ryosukesaito submitted an ensemble model of EfficientNet and SE-ResNeXt which achieved a score of **0.857** on the public leaderboard. The private leaderboard score was **0.519**.\n- The high public leaderboard score achieved by @ryosukesaito’s submission might have been a contributing factor to our ability to submit my somewhat ambitious notebook, possibly leading to our winning a silver medal.\n\n## My (@jooott) Solution\n\n### Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F580d1b944e38bb1924f6ec8480ac2b65%2F5.PNG?generation=1708327636242492&alt=media)\n\n### Key points\n\nI struggled significantly with stabilizing the training process.\n- To address this, I used Accumulate Grad Batches to effectively increase the batch size to 128, which stabilized the training.\n- A major factor in the significant improvement in score was the application of stronger data augmentation. The data augmentation strategy was inspired by [the 1st place solution of the Vesuvius Challenge - Ink Detection](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417496).\n- I also think that scaling up the training images from 512px to 1024px contributed to the increase in score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fe3e49e0014af22209252039314919b6d%2F10.PNG?generation=1708327710646310&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F01cf2045817c40092ded7e721c4ff52e%2F6.PNG?generation=1708327729992246&alt=media)\n\n```\ntrain_transform = A.Compose(\n            [\n                A.RandomScale(\n                    scale_limit=(1.0, 1.20),\n                    interpolation=cv2.INTER_CUBIC,\n                    p=0.1,\n                ),\n                A.RandomResizedCrop(\n                    image_size,\n                    image_size,\n                    scale=(0.8, 1.0),\n                    p=1\n                ),\n                A.RandomBrightnessContrast(p=0.75),\n                A.ShiftScaleRotate(p=0.75),\n                A.OneOf([\n                        A.GaussNoise(var_limit=[10, 50]),\n                        A.GaussianBlur(),\n                        A.MotionBlur(),\n                        ], p=0.4),\n                A.CoarseDropout(\n                    max_holes=1, max_width=int(image_size * 0.1),\n                    max_height=int(image_size * 0.1),\n                    mask_fill_value=0, p=0.5),\n                A.CLAHE(p=0.2),\n                A.GridDistortion(num_steps=5, distort_limit=0.3, p=0.05),\n                ToTensorV2(transpose_mask=True),\n            ]\n        )\n```\n\n## Muku's (@ryosukesaito) solution\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fa10c8d4428a2dd305f8f94b3bddbb6c7%2FUntitled%20(4).png?generation=1708327848313831&alt=media)\n\n### key points\n- In my architecture, Detection/Segmentation of kidney region is performed before predicting blood vessel area.\n    - Detection contributed to inference speedup (especially in the yz/zx direction), since it is possible to skip vessel segmentation in frames where kidney is not detected, and to reduce image size by cropping.\n    - Segmentation masks were used to reduce FP outside the kidney.\n    - For both annotations, I used LangSAM [(luca-medeiros/lang-segment-anything: SAM with text prompt](https://github.com/luca-medeiros/lang-segment-anything)). This allowed me to prepare annotation data with a few manual adjustments.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F3f044b330cb47140058deac70a7db57c%2FUntitled%20(5).png?generation=1708327892920221&alt=media)\n    - I use YOLOv8n for Detection and EfficientNet-B0 for Segmentation.\n\n- Various pre/post processing improved LB/PB scores slightly, but steadily.\n    - In the yz/zx axis image, blood vessels at the edge may be cut off. Since the inference accuracy was poor in this area, I improved the inference accuracy by pseudo-closing the vessels with mirror-padding before inference.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fb30520c1fb82b034bf913908cf65487d%2FUntitled%20(6).png?generation=1708327958129459&alt=media)\n\n    - After binarization of the results, defects may occur in the vascular prediction region as shown below. For this reason, morphological closing and fillPoly processing were added as post-processing steps.\nThese contributed to a slight score improvement in CV/LB/PB.\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F84de68a4219ff31a9f3623c4e83fc8f2%2FUntitled%20(7).png?generation=1708328014528398&alt=media)\n\n- In my experiments, ideas that contribute to generalization ability (strong augmentation, pseudo labeling, etc…) could not adopted as final submits, because they resulted in a decrease in CV/LB…\nHowever, I regret that I should not have been too aware of the unstable CV/LB, as the sample was not large enough for this competition.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2658452": "Firstly, we would like to express our gratitude to Kaggle and the organizers for hosting this exceptional competition. Through participating in this contest, we have gained a deeper understanding of the challenges and methodologies involved in medical image recognition.\n\n## Introduction\nWe submitted separate solutions within our team.\n- I submitted an ensemble model of se_resnext101_32x4d and Vision Transformer (mit_b2), which achieved a score of **0.834** on the public leaderboard. The private leaderboard score was **0.586**.\n- @ryosukesaito submitted an ensemble model of EfficientNet and SE-ResNeXt which achieved a score of **0.857** on the public leaderboard. The private leaderboard score was **0.519**.\n- The high public leaderboard score achieved by @ryosukesaito’s submission might have been a contributing factor to our ability to submit my somewhat ambitious notebook, possibly leading to our winning a silver medal.\n\n## My (@jooott) Solution\n\n### Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F580d1b944e38bb1924f6ec8480ac2b65%2F5.PNG?generation=1708327636242492&alt=media)\n\n### Key points\n\nI struggled significantly with stabilizing the training process.\n- To address this, I used Accumulate Grad Batches to effectively increase the batch size to 128, which stabilized the training.\n- A major factor in the significant improvement in score was the application of stronger data augmentation. The data augmentation strategy was inspired by [the 1st place solution of the Vesuvius Challenge - Ink Detection](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417496).\n- I also think that scaling up the training images from 512px to 1024px contributed to the increase in score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fe3e49e0014af22209252039314919b6d%2F10.PNG?generation=1708327710646310&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F01cf2045817c40092ded7e721c4ff52e%2F6.PNG?generation=1708327729992246&alt=media)\n\n```\ntrain_transform = A.Compose(\n            [\n                A.RandomScale(\n                    scale_limit=(1.0, 1.20),\n                    interpolation=cv2.INTER_CUBIC,\n                    p=0.1,\n                ),\n                A.RandomResizedCrop(\n                    image_size,\n                    image_size,\n                    scale=(0.8, 1.0),\n                    p=1\n                ),\n                A.RandomBrightnessContrast(p=0.75),\n                A.ShiftScaleRotate(p=0.75),\n                A.OneOf([\n                        A.GaussNoise(var_limit=[10, 50]),\n                        A.GaussianBlur(),\n                        A.MotionBlur(),\n                        ], p=0.4),\n                A.CoarseDropout(\n                    max_holes=1, max_width=int(image_size * 0.1),\n                    max_height=int(image_size * 0.1),\n                    mask_fill_value=0, p=0.5),\n                A.CLAHE(p=0.2),\n                A.GridDistortion(num_steps=5, distort_limit=0.3, p=0.05),\n                ToTensorV2(transpose_mask=True),\n            ]\n        )\n```\n\n## Muku's (@ryosukesaito) solution\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fa10c8d4428a2dd305f8f94b3bddbb6c7%2FUntitled%20(4).png?generation=1708327848313831&alt=media)\n\n### key points\n- In my architecture, Detection/Segmentation of kidney region is performed before predicting blood vessel area.\n    - Detection contributed to inference speedup (especially in the yz/zx direction), since it is possible to skip vessel segmentation in frames where kidney is not detected, and to reduce image size by cropping.\n    - Segmentation masks were used to reduce FP outside the kidney.\n    - For both annotations, I used LangSAM [(luca-medeiros/lang-segment-anything: SAM with text prompt](https://github.com/luca-medeiros/lang-segment-anything)). This allowed me to prepare annotation data with a few manual adjustments.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F3f044b330cb47140058deac70a7db57c%2FUntitled%20(5).png?generation=1708327892920221&alt=media)\n    - I use YOLOv8n for Detection and EfficientNet-B0 for Segmentation.\n\n- Various pre/post processing improved LB/PB scores slightly, but steadily.\n    - In the yz/zx axis image, blood vessels at the edge may be cut off. Since the inference accuracy was poor in this area, I improved the inference accuracy by pseudo-closing the vessels with mirror-padding before inference.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fb30520c1fb82b034bf913908cf65487d%2FUntitled%20(6).png?generation=1708327958129459&alt=media)\n\n    - After binarization of the results, defects may occur in the vascular prediction region as shown below. For this reason, morphological closing and fillPoly processing were added as post-processing steps.\nThese contributed to a slight score improvement in CV/LB/PB.\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F84de68a4219ff31a9f3623c4e83fc8f2%2FUntitled%20(7).png?generation=1708328014528398&alt=media)\n\n- In my experiments, ideas that contribute to generalization ability (strong augmentation, pseudo labeling, etc…) could not adopted as final submits, because they resulted in a decrease in CV/LB…\nHowever, I regret that I should not have been too aware of the unstable CV/LB, as the sample was not large enough for this competition."
  }
}