{
  "id": 475964,
  "title": "7th solution",
  "url": "/competitions/blood-vessel-segmentation/discussion/475964",
  "author_name": "sakaku",
  "post_date": "2024-02-10T15:51:31.473000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First and foremost, I'd like to extend my gratitude to Kaggle and the competition organizers for creating such a compelling event. Despite joining the contest relatively late, I was able to quickly get up to speed thanks to insightful discussions by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , the informative videos by <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> , and the vibrant exchanges within the community.</p>\n<h2>Overview</h2>\n<p>I employed a hybrid approach for the model, utilizing 2.5D images as input. The architecture combines a 2D Unet framework with 3D convolutional layers. From my research and the community's insights, it seemed that a full 3D model might offer superior results compared to 2D models. However, due to the high computational costs associated with 3D models, I choose a blend of 3D convolutions within a 2D Unet structure, striking a balance between efficiency and performance.</p>\n<h2>Data Preparation</h2>\n<h3>Generating 3D Rotational Slices</h3>\n<p>I augmented the dataset with 3D rotation. The process begins by assembling the images into a 3D volume, followed by rotating two axes and extracting slices along the remaining axis. The rotation angles used are as follows:</p>\n<pre><code>rotation_angles = [\n    [, ], [, -], [-, ], [-, -],\n    [, ], [, -], [-, ], [-, -],\n    [, ], [, -], [-, ], [-, -]\n]\n</code></pre>\n<p>Post-rotation, some slices exhibited increased areas of black background. To maintain data quality, I retained only those slices where the target segmentation was present and the black background constituted less than 50% of the slice area. </p>\n<h4>Rotate data sample</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2d8dc94653829769d2cc750903ab67de%2Fk1_z_pseudo_rotset7_1257.jpg?generation=1707579928485450&amp;alt=media\"></p>\n<h3>Pseudo Labeling</h3>\n<p>The pseudo labeling process involved:</p>\n<ol>\n<li>Generating additional slices for kidney1 and kidney3 using the aforementioned technique.</li>\n<li>Training a model with the augmented dataset.</li>\n<li>Applying the model to pseudo label kidney2, followed by generating extra slices for it in a similar manner.</li>\n</ol>\n<h2>Model</h2>\n<h3>Architecture</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2152324997361e303fb5fdfed3811b7d%2FScreenshot%202024-02-10%20at%2023.52.55.png?generation=1707579666652474&amp;alt=media\"></p>\n<h3>Training and Inference Details</h3>\n<h4>Training</h4>\n<p>The training setup was as follows:</p>\n<ol>\n<li>The model takes a 3-channel 2.5D image as input and outputs a 3-channel prediction.</li>\n<li>Normalize the input base on the std of each kidney</li>\n<li>I used a combination of loss functions: BCEWithLogitsLoss, DiceLoss, and FocalLoss. The loss for each of the three channels was calculated separately, with the middle channel assigned a higher weight of 0.9.</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: CosineAnnealingLR</li>\n<li>Images were cropped to a size of 1024x1024 for processing.</li>\n</ol>\n<h4>Inference</h4>\n<p>For inference:</p>\n<ol>\n<li>A single model was used for predictions.</li>\n<li>The model operated at the original image resolution.</li>\n<li>Similar to training, the input comprised 3-channel 2.5D images, with the output being a 3-channel prediction, primarily focusing on the middle channel.</li>\n<li>Normalize the input base on the std of each kidney</li>\n<li>Predictions were made along the x, y, and z axes, and the results were averaged.</li>\n<li>Test Time Augmentation (TTA) included horizontal flipping.</li>\n<li>Post-processing steps involved applying thresholds of 0.2, followed by a 3D closing operation.</li>\n</ol>",
  "messages": [
    {
      "id": 2646004,
      "postDate": "2024-02-10T15:51:31.473Z",
      "content": "<p>First and foremost, I'd like to extend my gratitude to Kaggle and the competition organizers for creating such a compelling event. Despite joining the contest relatively late, I was able to quickly get up to speed thanks to insightful discussions by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , the informative videos by <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> , and the vibrant exchanges within the community.</p>\n<h2>Overview</h2>\n<p>I employed a hybrid approach for the model, utilizing 2.5D images as input. The architecture combines a 2D Unet framework with 3D convolutional layers. From my research and the community's insights, it seemed that a full 3D model might offer superior results compared to 2D models. However, due to the high computational costs associated with 3D models, I choose a blend of 3D convolutions within a 2D Unet structure, striking a balance between efficiency and performance.</p>\n<h2>Data Preparation</h2>\n<h3>Generating 3D Rotational Slices</h3>\n<p>I augmented the dataset with 3D rotation. The process begins by assembling the images into a 3D volume, followed by rotating two axes and extracting slices along the remaining axis. The rotation angles used are as follows:</p>\n<pre><code>rotation_angles = [\n    [, ], [, -], [-, ], [-, -],\n    [, ], [, -], [-, ], [-, -],\n    [, ], [, -], [-, ], [-, -]\n]\n</code></pre>\n<p>Post-rotation, some slices exhibited increased areas of black background. To maintain data quality, I retained only those slices where the target segmentation was present and the black background constituted less than 50% of the slice area. </p>\n<h4>Rotate data sample</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2d8dc94653829769d2cc750903ab67de%2Fk1_z_pseudo_rotset7_1257.jpg?generation=1707579928485450&amp;alt=media\"></p>\n<h3>Pseudo Labeling</h3>\n<p>The pseudo labeling process involved:</p>\n<ol>\n<li>Generating additional slices for kidney1 and kidney3 using the aforementioned technique.</li>\n<li>Training a model with the augmented dataset.</li>\n<li>Applying the model to pseudo label kidney2, followed by generating extra slices for it in a similar manner.</li>\n</ol>\n<h2>Model</h2>\n<h3>Architecture</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2152324997361e303fb5fdfed3811b7d%2FScreenshot%202024-02-10%20at%2023.52.55.png?generation=1707579666652474&amp;alt=media\"></p>\n<h3>Training and Inference Details</h3>\n<h4>Training</h4>\n<p>The training setup was as follows:</p>\n<ol>\n<li>The model takes a 3-channel 2.5D image as input and outputs a 3-channel prediction.</li>\n<li>Normalize the input base on the std of each kidney</li>\n<li>I used a combination of loss functions: BCEWithLogitsLoss, DiceLoss, and FocalLoss. The loss for each of the three channels was calculated separately, with the middle channel assigned a higher weight of 0.9.</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: CosineAnnealingLR</li>\n<li>Images were cropped to a size of 1024x1024 for processing.</li>\n</ol>\n<h4>Inference</h4>\n<p>For inference:</p>\n<ol>\n<li>A single model was used for predictions.</li>\n<li>The model operated at the original image resolution.</li>\n<li>Similar to training, the input comprised 3-channel 2.5D images, with the output being a 3-channel prediction, primarily focusing on the middle channel.</li>\n<li>Normalize the input base on the std of each kidney</li>\n<li>Predictions were made along the x, y, and z axes, and the results were averaged.</li>\n<li>Test Time Augmentation (TTA) included horizontal flipping.</li>\n<li>Post-processing steps involved applying thresholds of 0.2, followed by a 3D closing operation.</li>\n</ol>",
      "rawMarkdown": "First and foremost, I'd like to extend my gratitude to Kaggle and the competition organizers for creating such a compelling event. Despite joining the contest relatively late, I was able to quickly get up to speed thanks to insightful discussions by @hengck23 , the informative videos by @yoyobar , and the vibrant exchanges within the community.\n\n## Overview\n\nI employed a hybrid approach for the model, utilizing 2.5D images as input. The architecture combines a 2D Unet framework with 3D convolutional layers. From my research and the community's insights, it seemed that a full 3D model might offer superior results compared to 2D models. However, due to the high computational costs associated with 3D models, I choose a blend of 3D convolutions within a 2D Unet structure, striking a balance between efficiency and performance.\n\n## Data Preparation\n\n### Generating 3D Rotational Slices\n\nI augmented the dataset with 3D rotation. The process begins by assembling the images into a 3D volume, followed by rotating two axes and extracting slices along the remaining axis. The rotation angles used are as follows:\n\n```python\nrotation_angles = [\n    [10, 10], [10, -10], [-10, 10], [-10, -10],\n    [30, 30], [30, -30], [-30, 30], [-30, -30],\n    [45, 45], [45, -45], [-45, 45], [-45, -45]\n]\n```\n\nPost-rotation, some slices exhibited increased areas of black background. To maintain data quality, I retained only those slices where the target segmentation was present and the black background constituted less than 50% of the slice area. \n#### Rotate data sample\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2d8dc94653829769d2cc750903ab67de%2Fk1_z_pseudo_rotset7_1257.jpg?generation=1707579928485450&alt=media)\n### Pseudo Labeling\n\nThe pseudo labeling process involved:\n1. Generating additional slices for kidney1 and kidney3 using the aforementioned technique.\n2. Training a model with the augmented dataset.\n3. Applying the model to pseudo label kidney2, followed by generating extra slices for it in a similar manner.\n\n## Model\n\n### Architecture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2152324997361e303fb5fdfed3811b7d%2FScreenshot%202024-02-10%20at%2023.52.55.png?generation=1707579666652474&alt=media)\n### Training and Inference Details\n\n#### Training\n\nThe training setup was as follows:\n1. The model takes a 3-channel 2.5D image as input and outputs a 3-channel prediction.\n2. Normalize the input base on the std of each kidney\n2. I used a combination of loss functions: BCEWithLogitsLoss, DiceLoss, and FocalLoss. The loss for each of the three channels was calculated separately, with the middle channel assigned a higher weight of 0.9.\n3. Optimizer: AdamW\n4. Scheduler: CosineAnnealingLR\n5. Images were cropped to a size of 1024x1024 for processing.\n\n#### Inference\n\nFor inference:\n1. A single model was used for predictions.\n2. The model operated at the original image resolution.\n3. Similar to training, the input comprised 3-channel 2.5D images, with the output being a 3-channel prediction, primarily focusing on the middle channel.\n2. Normalize the input base on the std of each kidney\n4. Predictions were made along the x, y, and z axes, and the results were averaged.\n5. Test Time Augmentation (TTA) included horizontal flipping.\n6. Post-processing steps involved applying thresholds of 0.2, followed by a 3D closing operation.\n",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2646004": "First and foremost, I'd like to extend my gratitude to Kaggle and the competition organizers for creating such a compelling event. Despite joining the contest relatively late, I was able to quickly get up to speed thanks to insightful discussions by @hengck23 , the informative videos by @yoyobar , and the vibrant exchanges within the community.\n\n## Overview\n\nI employed a hybrid approach for the model, utilizing 2.5D images as input. The architecture combines a 2D Unet framework with 3D convolutional layers. From my research and the community's insights, it seemed that a full 3D model might offer superior results compared to 2D models. However, due to the high computational costs associated with 3D models, I choose a blend of 3D convolutions within a 2D Unet structure, striking a balance between efficiency and performance.\n\n## Data Preparation\n\n### Generating 3D Rotational Slices\n\nI augmented the dataset with 3D rotation. The process begins by assembling the images into a 3D volume, followed by rotating two axes and extracting slices along the remaining axis. The rotation angles used are as follows:\n\n```python\nrotation_angles = [\n    [10, 10], [10, -10], [-10, 10], [-10, -10],\n    [30, 30], [30, -30], [-30, 30], [-30, -30],\n    [45, 45], [45, -45], [-45, 45], [-45, -45]\n]\n```\n\nPost-rotation, some slices exhibited increased areas of black background. To maintain data quality, I retained only those slices where the target segmentation was present and the black background constituted less than 50% of the slice area. \n#### Rotate data sample\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2d8dc94653829769d2cc750903ab67de%2Fk1_z_pseudo_rotset7_1257.jpg?generation=1707579928485450&alt=media)\n### Pseudo Labeling\n\nThe pseudo labeling process involved:\n1. Generating additional slices for kidney1 and kidney3 using the aforementioned technique.\n2. Training a model with the augmented dataset.\n3. Applying the model to pseudo label kidney2, followed by generating extra slices for it in a similar manner.\n\n## Model\n\n### Architecture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369671%2F2152324997361e303fb5fdfed3811b7d%2FScreenshot%202024-02-10%20at%2023.52.55.png?generation=1707579666652474&alt=media)\n### Training and Inference Details\n\n#### Training\n\nThe training setup was as follows:\n1. The model takes a 3-channel 2.5D image as input and outputs a 3-channel prediction.\n2. Normalize the input base on the std of each kidney\n2. I used a combination of loss functions: BCEWithLogitsLoss, DiceLoss, and FocalLoss. The loss for each of the three channels was calculated separately, with the middle channel assigned a higher weight of 0.9.\n3. Optimizer: AdamW\n4. Scheduler: CosineAnnealingLR\n5. Images were cropped to a size of 1024x1024 for processing.\n\n#### Inference\n\nFor inference:\n1. A single model was used for predictions.\n2. The model operated at the original image resolution.\n3. Similar to training, the input comprised 3-channel 2.5D images, with the output being a 3-channel prediction, primarily focusing on the middle channel.\n2. Normalize the input base on the std of each kidney\n4. Predictions were made along the x, y, and z axes, and the results were averaged.\n5. Test Time Augmentation (TTA) included horizontal flipping.\n6. Post-processing steps involved applying thresholds of 0.2, followed by a 3D closing operation.\n"
  }
}