{
  "id": 417383,
  "title": "8th Place Solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/writeups/luck-is-all-you-need-8th-place-solution",
  "author_name": "",
  "post_date": "2023-06-28T02:30:09.237Z",
  "votes": 20,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Team <a href=\"https://www.kaggle.com/renman\" target=\"_blank\">@renman</a> and <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> </p>\n<h2>Acknowledgement:</h2>\n<p>First of all we want to thank Kaggle for organizing such a great competition. The host has been very helpful and responsive and made this competition such a great experience for us!</p>\n<h2>1. Summary of solution:</h2>\n<ul>\n<li>Ensemble of 5 Unet models with 3D seresnet101, 3D Resnet34, and Segformer-b3 as backbones</li>\n<li>Post-processing methods such as masking out edge pixels for inference and applying rot90 and \"channels\" TTA</li>\n<li>Our final two submissions including one with a fixed threshold (0.55) and the other with a dynamic threshold based on pixel percentile (top 3% pixels)</li>\n</ul>\n<h2>2. Dataset:</h2>\n<h3>Preprocessing:</h3>\n<ul>\n<li>Set maximum pixel value to 0.78 (i.e. if pixel value &gt; 0.78, set it to 0.78)</li>\n</ul>\n<h3>CV Strategies:</h3>\n<ul>\n<li>We tried two different CV strategies:<ul>\n<li>3-fold CV: Split into fragment 1,2,3</li>\n<li>4-fold CV: split fragment 2 into top and bottom half (“2a” and “2b”)</li></ul></li>\n</ul>\n<h3>Data Augmentation:</h3>\n<ul>\n<li>Image Resizing and Sampling:<ul>\n<li>We augment the original dataset with three different resolutions categories (1, 0.75x, 0.5x)</li></ul></li>\n<li>Compression in z dimension (rate=+-0.2)</li>\n<li>3d Rotation (rate=+-5°)</li>\n<li>Channel Dropout</li>\n<li>Rotate and Flip</li>\n<li>Albumentation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F461631%2F6f55b04646dc7cfdea4163f03aa4fa91%2FScreen%20Shot%202023-06-15%20at%208.43.22%20PM.png?generation=1686834309651934&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>3. Models:</h2>\n<ul>\n<li><h3>3D SEResnet101:</h3>\n<ul>\n<li>Change to the architecture:<ul>\n<li>Add Squeeze and Excitation Layer and Dilated Convolution to the encoder</li>\n<li>Use BasicBlock instead of Bottleneck</li></ul></li>\n<li>No pretrained model is used. Model is trained from scratch with a 2 stage process:<ul>\n<li>The first stage is to train the model with image size of 96 which is easier to train</li>\n<li>The second stage is to continue training the model with bigger image size of 192 and 256</li></ul></li>\n<li>Fixed stride of 112</li>\n<li>Randomly pick 20 channels between 15 and 40 as input to the model:<ul>\n<li>At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (i.e. \"channel\" TTA)</li></ul></li>\n<li>BCE + Dice Loss</li>\n<li>Exclude areas outside binary mask<br>\n<br></li></ul></li>\n<li><h3>Segformer-b3:</h3>\n<ul>\n<li>Use Segformer-b3 with 3 channels input. Model pretrained on imagenet</li>\n<li>First method:<ul>\n<li>Select channels 25-36 (11 channels), split them into groups of [25-27, 28-30, 31-33, 34-36] and feed them into the model</li>\n<li>Merge the feature maps in the channel dimension with \"Attention Pooling\":<ul>\n<li>Refer to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s excellent elaborate in this thread: <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972</a></li></ul></li></ul></li>\n<li>Second Method:<ul>\n<li>Select channels 29-34 (6 channels). Split them into two groups of [29-31, 32-34]</li>\n<li>Feed each group into the same model and average the logit outputs</li></ul></li>\n<li>Image size 224</li>\n<li>BCE Loss</li>\n<li>Exclude areas outside binary mask<br>\n<br></li></ul></li>\n<li><h3>3D Resnet34:</h3>\n<ul>\n<li>Pick channel 22-34 (18 channels) as input to the model</li>\n<li>Image size 224 and 192</li>\n<li>BCE + Dice Loss</li>\n<li>Exclude areas outside binary mask</li></ul></li>\n</ul>\n<h2>4. Inference:</h2>\n<ul>\n<li>Fp16</li>\n<li>Ignore areas outside binary mask</li>\n<li>Rotation TTA (90 degree rotation)</li>\n<li>Channels TTA:<ul>\n<li>At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (3D SEResnet101 only)</li></ul></li>\n<li>Mask out edge pixels:<ul>\n<li>Only generate predictions in the “middle part” of the output</li></ul></li>\n<li>Thresholding:<ul>\n<li>One submission we use a fixed Threshold of 0.55</li>\n<li>The other submission we use a dynamic threshold strategy based on pixel percentile (top 3% pixels), i.e. sort the pixels and pick a threshold value such that top 3 percentile of pixels are selected</li></ul></li>\n</ul>\n<h2>5. Final Ensemble:</h2>\n<ul>\n<li>The final ensemble consists of a weighted average of the 5 models below:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Models</th>\n<th>Image size</th>\n<th>Kfold</th>\n<th>Output size</th>\n<th>Stride</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3D Seresnet101</td>\n<td>256</td>\n<td>Single Fold validated on fragment \"2a\"</td>\n<td>192</td>\n<td>192 // 4</td>\n<td>0.76</td>\n<td>0.61</td>\n</tr>\n<tr>\n<td>Segformer-b3 with attention pooling</td>\n<td>224</td>\n<td>4 folds</td>\n<td>160</td>\n<td>160 // 2</td>\n<td>0.7</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>Segformer-b3 with averaged logit</td>\n<td>224</td>\n<td>3 folds</td>\n<td>176</td>\n<td>176 // 2</td>\n<td>0.66</td>\n<td>0.56</td>\n</tr>\n<tr>\n<td>3D Resnet34</td>\n<td>224</td>\n<td>4 folds</td>\n<td>192</td>\n<td>192 // 2</td>\n<td>0.66</td>\n<td>0.58</td>\n</tr>\n<tr>\n<td>3D Resnet34</td>\n<td>192</td>\n<td>3 folds</td>\n<td>160</td>\n<td>160 // 2</td>\n<td>-</td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Our final 2 submissions including one with a fixed threshold (public LB: 0.79, private LB: 0.65) and one with a pixel percentile threshold (public LB: 0.79, private LB: 0.63)</li>\n</ul>\n<h2>6. Methods tried but did not work:</h2>\n<ul>\n<li>Tried to add Convolutional Block Attention Module (CBAM) for the decoder of 3D Resnet but CV dropped<ul>\n<li>Paper to CBAM: <a href=\"https://arxiv.org/abs/1807.06521\" target=\"_blank\">https://arxiv.org/abs/1807.06521</a></li></ul></li>\n<li>Use IR image as additional segmentation head but it did not lead to any improvement in CV</li>\n<li>Mixup and label noise for data augmentation</li>\n</ul>\n<p>Update: The code and detailed instructions for replicating our models can be found in the following publicly available github repo:</p>\n<p><a href=\"https://github.com/flyyufelix/vesuvius_challenge_8th_place_solution\" target=\"_blank\">https://github.com/flyyufelix/vesuvius_challenge_8th_place_solution</a></p>",
  "messages": [
    {
      "id": "2303799",
      "postDate": "06/15/2023 13:32:51",
      "content": "<p>Team <a href=\"https://www.kaggle.com/renman\" target=\"_blank\">@renman</a> and <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> </p>\n<h2>Acknowledgement:</h2>\n<p>First of all we want to thank Kaggle for organizing such a great competition. The host has been very helpful and responsive and made this competition such a great experience for us!</p>\n<h2>1. Summary of solution:</h2>\n<ul>\n<li>Ensemble of 5 Unet models with 3D seresnet101, 3D Resnet34, and Segformer-b3 as backbones</li>\n<li>Post-processing methods such as masking out edge pixels for inference and applying rot90 and \"channels\" TTA</li>\n<li>Our final two submissions including one with a fixed threshold (0.55) and the other with a dynamic threshold based on pixel percentile (top 3% pixels)</li>\n</ul>\n<h2>2. Dataset:</h2>\n<h3>Preprocessing:</h3>\n<ul>\n<li>Set maximum pixel value to 0.78 (i.e. if pixel value &gt; 0.78, set it to 0.78)</li>\n</ul>\n<h3>CV Strategies:</h3>\n<ul>\n<li>We tried two different CV strategies:<ul>\n<li>3-fold CV: Split into fragment 1,2,3</li>\n<li>4-fold CV: split fragment 2 into top and bottom half (“2a” and “2b”)</li></ul></li>\n</ul>\n<h3>Data Augmentation:</h3>\n<ul>\n<li>Image Resizing and Sampling:<ul>\n<li>We augment the original dataset with three different resolutions categories (1, 0.75x, 0.5x)</li></ul></li>\n<li>Compression in z dimension (rate=+-0.2)</li>\n<li>3d Rotation (rate=+-5°)</li>\n<li>Channel Dropout</li>\n<li>Rotate and Flip</li>\n<li>Albumentation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F461631%2F6f55b04646dc7cfdea4163f03aa4fa91%2FScreen%20Shot%202023-06-15%20at%208.43.22%20PM.png?generation=1686834309651934&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>3. Models:</h2>\n<ul>\n<li><h3>3D SEResnet101:</h3>\n<ul>\n<li>Change to the architecture:<ul>\n<li>Add Squeeze and Excitation Layer and Dilated Convolution to the encoder</li>\n<li>Use BasicBlock instead of Bottleneck</li></ul></li>\n<li>No pretrained model is used. Model is trained from scratch with a 2 stage process:<ul>\n<li>The first stage is to train the model with image size of 96 which is easier to train</li>\n<li>The second stage is to continue training the model with bigger image size of 192 and 256</li></ul></li>\n<li>Fixed stride of 112</li>\n<li>Randomly pick 20 channels between 15 and 40 as input to the model:<ul>\n<li>At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (i.e. \"channel\" TTA)</li></ul></li>\n<li>BCE + Dice Loss</li>\n<li>Exclude areas outside binary mask<br>\n<br></li></ul></li>\n<li><h3>Segformer-b3:</h3>\n<ul>\n<li>Use Segformer-b3 with 3 channels input. Model pretrained on imagenet</li>\n<li>First method:<ul>\n<li>Select channels 25-36 (11 channels), split them into groups of [25-27, 28-30, 31-33, 34-36] and feed them into the model</li>\n<li>Merge the feature maps in the channel dimension with \"Attention Pooling\":<ul>\n<li>Refer to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s excellent elaborate in this thread: <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972</a></li></ul></li></ul></li>\n<li>Second Method:<ul>\n<li>Select channels 29-34 (6 channels). Split them into two groups of [29-31, 32-34]</li>\n<li>Feed each group into the same model and average the logit outputs</li></ul></li>\n<li>Image size 224</li>\n<li>BCE Loss</li>\n<li>Exclude areas outside binary mask<br>\n<br></li></ul></li>\n<li><h3>3D Resnet34:</h3>\n<ul>\n<li>Pick channel 22-34 (18 channels) as input to the model</li>\n<li>Image size 224 and 192</li>\n<li>BCE + Dice Loss</li>\n<li>Exclude areas outside binary mask</li></ul></li>\n</ul>\n<h2>4. Inference:</h2>\n<ul>\n<li>Fp16</li>\n<li>Ignore areas outside binary mask</li>\n<li>Rotation TTA (90 degree rotation)</li>\n<li>Channels TTA:<ul>\n<li>At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (3D SEResnet101 only)</li></ul></li>\n<li>Mask out edge pixels:<ul>\n<li>Only generate predictions in the “middle part” of the output</li></ul></li>\n<li>Thresholding:<ul>\n<li>One submission we use a fixed Threshold of 0.55</li>\n<li>The other submission we use a dynamic threshold strategy based on pixel percentile (top 3% pixels), i.e. sort the pixels and pick a threshold value such that top 3 percentile of pixels are selected</li></ul></li>\n</ul>\n<h2>5. Final Ensemble:</h2>\n<ul>\n<li>The final ensemble consists of a weighted average of the 5 models below:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Models</th>\n<th>Image size</th>\n<th>Kfold</th>\n<th>Output size</th>\n<th>Stride</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3D Seresnet101</td>\n<td>256</td>\n<td>Single Fold validated on fragment \"2a\"</td>\n<td>192</td>\n<td>192 // 4</td>\n<td>0.76</td>\n<td>0.61</td>\n</tr>\n<tr>\n<td>Segformer-b3 with attention pooling</td>\n<td>224</td>\n<td>4 folds</td>\n<td>160</td>\n<td>160 // 2</td>\n<td>0.7</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>Segformer-b3 with averaged logit</td>\n<td>224</td>\n<td>3 folds</td>\n<td>176</td>\n<td>176 // 2</td>\n<td>0.66</td>\n<td>0.56</td>\n</tr>\n<tr>\n<td>3D Resnet34</td>\n<td>224</td>\n<td>4 folds</td>\n<td>192</td>\n<td>192 // 2</td>\n<td>0.66</td>\n<td>0.58</td>\n</tr>\n<tr>\n<td>3D Resnet34</td>\n<td>192</td>\n<td>3 folds</td>\n<td>160</td>\n<td>160 // 2</td>\n<td>-</td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Our final 2 submissions including one with a fixed threshold (public LB: 0.79, private LB: 0.65) and one with a pixel percentile threshold (public LB: 0.79, private LB: 0.63)</li>\n</ul>\n<h2>6. Methods tried but did not work:</h2>\n<ul>\n<li>Tried to add Convolutional Block Attention Module (CBAM) for the decoder of 3D Resnet but CV dropped<ul>\n<li>Paper to CBAM: <a href=\"https://arxiv.org/abs/1807.06521\" target=\"_blank\">https://arxiv.org/abs/1807.06521</a></li></ul></li>\n<li>Use IR image as additional segmentation head but it did not lead to any improvement in CV</li>\n<li>Mixup and label noise for data augmentation</li>\n</ul>\n<p>Update: The code and detailed instructions for replicating our models can be found in the following publicly available github repo:</p>\n<p><a href=\"https://github.com/flyyufelix/vesuvius_challenge_8th_place_solution\" target=\"_blank\">https://github.com/flyyufelix/vesuvius_challenge_8th_place_solution</a></p>",
      "rawMarkdown": "Team @renman and @yoyobar \n\n## Acknowledgement:\n\nFirst of all we want to thank Kaggle for organizing such a great competition. The host has been very helpful and responsive and made this competition such a great experience for us!\n \n##1. Summary of solution:\n\n- Ensemble of 5 Unet models with 3D seresnet101, 3D Resnet34, and Segformer-b3 as backbones\n- Post-processing methods such as masking out edge pixels for inference and applying rot90 and \"channels\" TTA\n- Our final two submissions including one with a fixed threshold (0.55) and the other with a dynamic threshold based on pixel percentile (top 3% pixels)\n\n##2. Dataset:\n\n###Preprocessing:\n\n- Set maximum pixel value to 0.78 (i.e. if pixel value > 0.78, set it to 0.78)\n \n###CV Strategies:\n\n- We tried two different CV strategies:\n    - 3-fold CV: Split into fragment 1,2,3\n    - 4-fold CV: split fragment 2 into top and bottom half (“2a” and “2b”)\n\n###Data Augmentation:\n\n- Image Resizing and Sampling:\n    - We augment the original dataset with three different resolutions categories (1, 0.75x, 0.5x)\n- Compression in z dimension (rate=+-0.2)\n- 3d Rotation (rate=+-5°)\n- Channel Dropout\n- Rotate and Flip\n- Albumentation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F461631%2F6f55b04646dc7cfdea4163f03aa4fa91%2FScreen%20Shot%202023-06-15%20at%208.43.22%20PM.png?generation=1686834309651934&alt=media)\n\n##3. Models:\n\n- ###3D SEResnet101:\n\n    - Change to the architecture:\n        - Add Squeeze and Excitation Layer and Dilated Convolution to the encoder\n        - Use BasicBlock instead of Bottleneck\n    - No pretrained model is used. Model is trained from scratch with a 2 stage process:\n        - The first stage is to train the model with image size of 96 which is easier to train\n        - The second stage is to continue training the model with bigger image size of 192 and 256\n    - Fixed stride of 112\n    - Randomly pick 20 channels between 15 and 40 as input to the model:\n        - At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (i.e. \"channel\" TTA)\n    - BCE + Dice Loss\n    - Exclude areas outside binary mask\n<br />\n\n- ###Segformer-b3:\n\n    - Use Segformer-b3 with 3 channels input. Model pretrained on imagenet\n    - First method:\n        - Select channels 25-36 (11 channels), split them into groups of [25-27, 28-30, 31-33, 34-36] and feed them into the model\n        - Merge the feature maps in the channel dimension with \"Attention Pooling\":\n            - Refer to @hengck23's excellent elaborate in this thread: https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\n    - Second Method:\n        - Select channels 29-34 (6 channels). Split them into two groups of [29-31, 32-34]\n        - Feed each group into the same model and average the logit outputs\n    - Image size 224\n    - BCE Loss\n    - Exclude areas outside binary mask\n<br />\n\n- ###3D Resnet34:\n\n    - Pick channel 22-34 (18 channels) as input to the model\n    - Image size 224 and 192\n    - BCE + Dice Loss\n    - Exclude areas outside binary mask\n\n##4. Inference:\n\n- Fp16\n- Ignore areas outside binary mask\n- Rotation TTA (90 degree rotation)\n- Channels TTA:\n    - At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (3D SEResnet101 only)\n- Mask out edge pixels:\n    - Only generate predictions in the “middle part” of the output\n- Thresholding:\n    - One submission we use a fixed Threshold of 0.55\n    - The other submission we use a dynamic threshold strategy based on pixel percentile (top 3% pixels), i.e. sort the pixels and pick a threshold value such that top 3 percentile of pixels are selected\n\n##5. Final Ensemble:\n\n- The final ensemble consists of a weighted average of the 5 models below:\n\n| Models | Image size | Kfold | Output size | Stride | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- | --- |\n| 3D Seresnet101                       | 256        | Single Fold validated on fragment \"2a\"  | 192         | 192 // 4 | 0.76      | 0.61     |\n| Segformer-b3 with attention pooling  | 224        | 4 folds                                 | 160         | 160 // 2 | 0.7       | 0.6     |\n| Segformer-b3 with averaged logit     | 224        | 3 folds                                 | 176         | 176 // 2 | 0.66      | 0.56   |    \n| 3D Resnet34                          | 224        | 4 folds                                 | 192         | 192 // 2 | 0.66      | 0.58       |\n| 3D Resnet34                          | 192        | 3 folds                                 | 160         | 160 // 2 | -         | -          |\n\n- Our final 2 submissions including one with a fixed threshold (public LB: 0.79, private LB: 0.65) and one with a pixel percentile threshold (public LB: 0.79, private LB: 0.63)\n\n##6. Methods tried but did not work:\n- Tried to add Convolutional Block Attention Module (CBAM) for the decoder of 3D Resnet but CV dropped\n    - Paper to CBAM: https://arxiv.org/abs/1807.06521\n- Use IR image as additional segmentation head but it did not lead to any improvement in CV\n- Mixup and label noise for data augmentation\n\nUpdate: The code and detailed instructions for replicating our models can be found in the following publicly available github repo:\n\nhttps://github.com/flyyufelix/vesuvius_challenge_8th_place_solution",
      "votes": null
    },
    {
      "id": "2310186",
      "postDate": "06/20/2023 07:20:42",
      "content": "<p>Thank you for sharing the different models you used.  I found it interesting the different ways you used the z-channels in the three different models.  I think more experimentation with approaches would help my understanding of what information is contained in that dimension.</p>",
      "rawMarkdown": "Thank you for sharing the different models you used.  I found it interesting the different ways you used the z-channels in the three different models.  I think more experimentation with approaches would help my understanding of what information is contained in that dimension.",
      "votes": null
    },
    {
      "id": "2319217",
      "postDate": "06/27/2023 01:46:21",
      "content": "<p>May I ask how your segformer is implemented and is there any available source code？</p>",
      "rawMarkdown": "May I ask how your segformer is implemented and is there any available source code？",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2310186,
      "author_name": "socated",
      "author_url": "",
      "post_date": "06/20/2023 07:20:42",
      "content": "<p>Thank you for sharing the different models you used.  I found it interesting the different ways you used the z-channels in the three different models.  I think more experimentation with approaches would help my understanding of what information is contained in that dimension.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2319217,
      "author_name": "wangxuc",
      "author_url": "",
      "post_date": "06/27/2023 01:46:21",
      "content": "<p>May I ask how your segformer is implemented and is there any available source code？</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2303799": "Team @renman and @yoyobar \n\n## Acknowledgement:\n\nFirst of all we want to thank Kaggle for organizing such a great competition. The host has been very helpful and responsive and made this competition such a great experience for us!\n \n##1. Summary of solution:\n\n- Ensemble of 5 Unet models with 3D seresnet101, 3D Resnet34, and Segformer-b3 as backbones\n- Post-processing methods such as masking out edge pixels for inference and applying rot90 and \"channels\" TTA\n- Our final two submissions including one with a fixed threshold (0.55) and the other with a dynamic threshold based on pixel percentile (top 3% pixels)\n\n##2. Dataset:\n\n###Preprocessing:\n\n- Set maximum pixel value to 0.78 (i.e. if pixel value > 0.78, set it to 0.78)\n \n###CV Strategies:\n\n- We tried two different CV strategies:\n    - 3-fold CV: Split into fragment 1,2,3\n    - 4-fold CV: split fragment 2 into top and bottom half (“2a” and “2b”)\n\n###Data Augmentation:\n\n- Image Resizing and Sampling:\n    - We augment the original dataset with three different resolutions categories (1, 0.75x, 0.5x)\n- Compression in z dimension (rate=+-0.2)\n- 3d Rotation (rate=+-5°)\n- Channel Dropout\n- Rotate and Flip\n- Albumentation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F461631%2F6f55b04646dc7cfdea4163f03aa4fa91%2FScreen%20Shot%202023-06-15%20at%208.43.22%20PM.png?generation=1686834309651934&alt=media)\n\n##3. Models:\n\n- ###3D SEResnet101:\n\n    - Change to the architecture:\n        - Add Squeeze and Excitation Layer and Dilated Convolution to the encoder\n        - Use BasicBlock instead of Bottleneck\n    - No pretrained model is used. Model is trained from scratch with a 2 stage process:\n        - The first stage is to train the model with image size of 96 which is easier to train\n        - The second stage is to continue training the model with bigger image size of 192 and 256\n    - Fixed stride of 112\n    - Randomly pick 20 channels between 15 and 40 as input to the model:\n        - At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (i.e. \"channel\" TTA)\n    - BCE + Dice Loss\n    - Exclude areas outside binary mask\n<br />\n\n- ###Segformer-b3:\n\n    - Use Segformer-b3 with 3 channels input. Model pretrained on imagenet\n    - First method:\n        - Select channels 25-36 (11 channels), split them into groups of [25-27, 28-30, 31-33, 34-36] and feed them into the model\n        - Merge the feature maps in the channel dimension with \"Attention Pooling\":\n            - Refer to @hengck23's excellent elaborate in this thread: https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/407972\n    - Second Method:\n        - Select channels 29-34 (6 channels). Split them into two groups of [29-31, 32-34]\n        - Feed each group into the same model and average the logit outputs\n    - Image size 224\n    - BCE Loss\n    - Exclude areas outside binary mask\n<br />\n\n- ###3D Resnet34:\n\n    - Pick channel 22-34 (18 channels) as input to the model\n    - Image size 224 and 192\n    - BCE + Dice Loss\n    - Exclude areas outside binary mask\n\n##4. Inference:\n\n- Fp16\n- Ignore areas outside binary mask\n- Rotation TTA (90 degree rotation)\n- Channels TTA:\n    - At inference time, slide through channels 15 to 40 with a moving window of 20 and stride of 5 and take average (3D SEResnet101 only)\n- Mask out edge pixels:\n    - Only generate predictions in the “middle part” of the output\n- Thresholding:\n    - One submission we use a fixed Threshold of 0.55\n    - The other submission we use a dynamic threshold strategy based on pixel percentile (top 3% pixels), i.e. sort the pixels and pick a threshold value such that top 3 percentile of pixels are selected\n\n##5. Final Ensemble:\n\n- The final ensemble consists of a weighted average of the 5 models below:\n\n| Models | Image size | Kfold | Output size | Stride | Public LB | Private LB |\n| --- | --- | --- | --- | --- | --- | --- |\n| 3D Seresnet101                       | 256        | Single Fold validated on fragment \"2a\"  | 192         | 192 // 4 | 0.76      | 0.61     |\n| Segformer-b3 with attention pooling  | 224        | 4 folds                                 | 160         | 160 // 2 | 0.7       | 0.6     |\n| Segformer-b3 with averaged logit     | 224        | 3 folds                                 | 176         | 176 // 2 | 0.66      | 0.56   |    \n| 3D Resnet34                          | 224        | 4 folds                                 | 192         | 192 // 2 | 0.66      | 0.58       |\n| 3D Resnet34                          | 192        | 3 folds                                 | 160         | 160 // 2 | -         | -          |\n\n- Our final 2 submissions including one with a fixed threshold (public LB: 0.79, private LB: 0.65) and one with a pixel percentile threshold (public LB: 0.79, private LB: 0.63)\n\n##6. Methods tried but did not work:\n- Tried to add Convolutional Block Attention Module (CBAM) for the decoder of 3D Resnet but CV dropped\n    - Paper to CBAM: https://arxiv.org/abs/1807.06521\n- Use IR image as additional segmentation head but it did not lead to any improvement in CV\n- Mixup and label noise for data augmentation\n\nUpdate: The code and detailed instructions for replicating our models can be found in the following publicly available github repo:\n\nhttps://github.com/flyyufelix/vesuvius_challenge_8th_place_solution",
    "2310186": "Thank you for sharing the different models you used.  I found it interesting the different ways you used the z-channels in the three different models.  I think more experimentation with approaches would help my understanding of what information is contained in that dimension.",
    "2319217": "May I ask how your segformer is implemented and is there any available source code？"
  },
  "source": "meta"
}