{
  "id": 417563,
  "title": "73th Place Solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417563",
  "author_name": "WangXuC",
  "post_date": "2023-06-16T08:32:13.046000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Acknowledgement:</h2>\n<p>Firstly, I would like to express my gratitude to the organizers of the competition and Kaggle. Next, I want to thank <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a>, I started the competition based on his/her open-source NoteBooks, thank you!🎉. This competition is my first and I am glad that my ranking is in the top 6%, which is a milestone for me. This achievement cannot be separated from my enthusiastic senior brothers, and I am grateful for their valuable suggestions!</p>\n<h2>1. Overview:👀</h2>\n<ul>\n<li>Ensemble of 8 Unet models with mit, pvt as backbones</li>\n<li>TTA in four directions(0°, 90°, 180°, 270°) in the inference stage</li>\n</ul>\n<h2>2. DataFlow:👇</h2>\n<ul>\n<li>Select channels 28-34 (11 channels), split them into groups of [28-30, 30-32, 32-34] and concatenate them together in batch_ Size dimension</li>\n<li>Split fragment2 into two sub fragmentA and B, thus we have four parts of the data, we can achieve 4fold cross validation</li>\n</ul>\n<h2>3. Data Augmentation:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2F80b2b3e3d5d6a89f77c70a14eb5a75d3%2F1686900312738.jpg?generation=1686900353043727&amp;alt=media\" alt=\"\"></p>\n<h2>4. Model Selection</h2>\n<p>I initially tried using Segformer, but the results of this model were not good at that time, so I did not use Segformer in subsequent experiments.Now it seems like I did something wrong.<br>\nI have tried many different encoder models, such as resnet, resnext, seresnext, efficientnet, convnext, mit, and pvt, but <strong>only mit and pvt have shown superior performance</strong>, so all subsequent experiments are based on these two models.Meanwhile, it should be noted that overly complex and deep networks may lead to overfitting of the model on training data, so I chose <strong>mit-b2 and pvt-v2-b2</strong> as the encoder models.</p>\n<h2>5. Improvement of Network Structure</h2>\n<ul>\n<li>All shape feature maps output by the encoder will go through an additional SK attention module</li>\n<li>Then perform the following operations on the feature map that has undergone SK attention, as referenced by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> .<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fdf703992b4d39cfa8c5758910e4ddceb%2F1686902361015.png?generation=1686902393516010&amp;alt=media\" alt=\"\"><br>\nThe structure of self.weight1 is as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fccf94645bbad0c6de308f39056f3293a%2F1686902566772.png?generation=1686902593485222&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>6. Other Training Details</h2>\n<ul>\n<li>Image Size: 224 * 224</li>\n<li>Stride: 112</li>\n<li>Loss Function: SoftBCELoss</li>\n<li>Optimizer: AdamW</li>\n<li>Initial LR: 0.0001</li>\n<li>Scheduler: GradualWarmupSchedulerV2 with CosineAnnealingLR</li>\n<li>Batch_size: 32</li>\n<li>Epoch: 30</li>\n</ul>\n<h2>7. Inference</h2>\n<ul>\n<li>Ensemble of 8 Unet models with mit, pvt as backbones</li>\n<li>TTA in four directions(0°, 90°, 180°, 270°) in the inference stage</li>\n<li>Mit &amp; Pvt ensemble 👉Public LB: 0.721 Private LB: 0.586</li>\n<li>Single Pvt 👉Public LB: 0.726 Private LB: 0.591</li>\n</ul>\n<h2>8. Methods tried but did not work</h2>\n<ul>\n<li>DiceLoss❌</li>\n<li>DeNoise module❌</li>\n<li>Image size larger than 224 * 224❌</li>\n<li>Adding attention only to the output of the encoder without adding modules, referenced by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , is not effective❌</li>\n<li>The ensemble of mit and pvt seems to have less good results than the individual pvt❌</li>\n</ul>",
  "messages": [
    {
      "id": 2304785,
      "postDate": "2023-06-16T08:32:13.047Z",
      "content": "<h2>Acknowledgement:</h2>\n<p>Firstly, I would like to express my gratitude to the organizers of the competition and Kaggle. Next, I want to thank <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a>, I started the competition based on his/her open-source NoteBooks, thank you!🎉. This competition is my first and I am glad that my ranking is in the top 6%, which is a milestone for me. This achievement cannot be separated from my enthusiastic senior brothers, and I am grateful for their valuable suggestions!</p>\n<h2>1. Overview:👀</h2>\n<ul>\n<li>Ensemble of 8 Unet models with mit, pvt as backbones</li>\n<li>TTA in four directions(0°, 90°, 180°, 270°) in the inference stage</li>\n</ul>\n<h2>2. DataFlow:👇</h2>\n<ul>\n<li>Select channels 28-34 (11 channels), split them into groups of [28-30, 30-32, 32-34] and concatenate them together in batch_ Size dimension</li>\n<li>Split fragment2 into two sub fragmentA and B, thus we have four parts of the data, we can achieve 4fold cross validation</li>\n</ul>\n<h2>3. Data Augmentation:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2F80b2b3e3d5d6a89f77c70a14eb5a75d3%2F1686900312738.jpg?generation=1686900353043727&amp;alt=media\" alt=\"\"></p>\n<h2>4. Model Selection</h2>\n<p>I initially tried using Segformer, but the results of this model were not good at that time, so I did not use Segformer in subsequent experiments.Now it seems like I did something wrong.<br>\nI have tried many different encoder models, such as resnet, resnext, seresnext, efficientnet, convnext, mit, and pvt, but <strong>only mit and pvt have shown superior performance</strong>, so all subsequent experiments are based on these two models.Meanwhile, it should be noted that overly complex and deep networks may lead to overfitting of the model on training data, so I chose <strong>mit-b2 and pvt-v2-b2</strong> as the encoder models.</p>\n<h2>5. Improvement of Network Structure</h2>\n<ul>\n<li>All shape feature maps output by the encoder will go through an additional SK attention module</li>\n<li>Then perform the following operations on the feature map that has undergone SK attention, as referenced by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> .<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fdf703992b4d39cfa8c5758910e4ddceb%2F1686902361015.png?generation=1686902393516010&amp;alt=media\" alt=\"\"><br>\nThe structure of self.weight1 is as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fccf94645bbad0c6de308f39056f3293a%2F1686902566772.png?generation=1686902593485222&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>6. Other Training Details</h2>\n<ul>\n<li>Image Size: 224 * 224</li>\n<li>Stride: 112</li>\n<li>Loss Function: SoftBCELoss</li>\n<li>Optimizer: AdamW</li>\n<li>Initial LR: 0.0001</li>\n<li>Scheduler: GradualWarmupSchedulerV2 with CosineAnnealingLR</li>\n<li>Batch_size: 32</li>\n<li>Epoch: 30</li>\n</ul>\n<h2>7. Inference</h2>\n<ul>\n<li>Ensemble of 8 Unet models with mit, pvt as backbones</li>\n<li>TTA in four directions(0°, 90°, 180°, 270°) in the inference stage</li>\n<li>Mit &amp; Pvt ensemble 👉Public LB: 0.721 Private LB: 0.586</li>\n<li>Single Pvt 👉Public LB: 0.726 Private LB: 0.591</li>\n</ul>\n<h2>8. Methods tried but did not work</h2>\n<ul>\n<li>DiceLoss❌</li>\n<li>DeNoise module❌</li>\n<li>Image size larger than 224 * 224❌</li>\n<li>Adding attention only to the output of the encoder without adding modules, referenced by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , is not effective❌</li>\n<li>The ensemble of mit and pvt seems to have less good results than the individual pvt❌</li>\n</ul>",
      "rawMarkdown": "## Acknowledgement:\nFirstly, I would like to express my gratitude to the organizers of the competition and Kaggle. Next, I want to thank @yoyobar, I started the competition based on his/her open-source NoteBooks, thank you!🎉. This competition is my first and I am glad that my ranking is in the top 6%, which is a milestone for me. This achievement cannot be separated from my enthusiastic senior brothers, and I am grateful for their valuable suggestions!\n## 1. Overview:👀\n- Ensemble of 8 Unet models with mit, pvt as backbones\n- TTA in four directions(0°, 90°, 180°, 270°) in the inference stage\n## 2. DataFlow:👇\n- Select channels 28-34 (11 channels), split them into groups of [28-30, 30-32, 32-34] and concatenate them together in batch_ Size dimension\n- Split fragment2 into two sub fragmentA and B, thus we have four parts of the data, we can achieve 4fold cross validation\n## 3. Data Augmentation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2F80b2b3e3d5d6a89f77c70a14eb5a75d3%2F1686900312738.jpg?generation=1686900353043727&alt=media)\n## 4. Model Selection\nI initially tried using Segformer, but the results of this model were not good at that time, so I did not use Segformer in subsequent experiments.Now it seems like I did something wrong.\nI have tried many different encoder models, such as resnet, resnext, seresnext, efficientnet, convnext, mit, and pvt, but **only mit and pvt have shown superior performance**, so all subsequent experiments are based on these two models.Meanwhile, it should be noted that overly complex and deep networks may lead to overfitting of the model on training data, so I chose **mit-b2 and pvt-v2-b2** as the encoder models.\n## 5. Improvement of Network Structure\n- All shape feature maps output by the encoder will go through an additional SK attention module\n- Then perform the following operations on the feature map that has undergone SK attention, as referenced by @hengck23 .\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fdf703992b4d39cfa8c5758910e4ddceb%2F1686902361015.png?generation=1686902393516010&alt=media)\nThe structure of self.weight1 is as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fccf94645bbad0c6de308f39056f3293a%2F1686902566772.png?generation=1686902593485222&alt=media)\n## 6. Other Training Details\n- Image Size: 224 * 224\n- Stride: 112\n- Loss Function: SoftBCELoss\n- Optimizer: AdamW\n- Initial LR: 0.0001\n- Scheduler: GradualWarmupSchedulerV2 with CosineAnnealingLR\n- Batch_size: 32\n- Epoch: 30\n## 7. Inference\n- Ensemble of 8 Unet models with mit, pvt as backbones\n- TTA in four directions(0°, 90°, 180°, 270°) in the inference stage\n- Mit & Pvt ensemble 👉Public LB: 0.721 Private LB: 0.586\n- Single Pvt 👉Public LB: 0.726 Private LB: 0.591\n## 8. Methods tried but did not work\n- DiceLoss❌\n- DeNoise module❌\n- Image size larger than 224 * 224❌\n- Adding attention only to the output of the encoder without adding modules, referenced by @hengck23 , is not effective❌\n- The ensemble of mit and pvt seems to have less good results than the individual pvt❌",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2304785": "## Acknowledgement:\nFirstly, I would like to express my gratitude to the organizers of the competition and Kaggle. Next, I want to thank @yoyobar, I started the competition based on his/her open-source NoteBooks, thank you!🎉. This competition is my first and I am glad that my ranking is in the top 6%, which is a milestone for me. This achievement cannot be separated from my enthusiastic senior brothers, and I am grateful for their valuable suggestions!\n## 1. Overview:👀\n- Ensemble of 8 Unet models with mit, pvt as backbones\n- TTA in four directions(0°, 90°, 180°, 270°) in the inference stage\n## 2. DataFlow:👇\n- Select channels 28-34 (11 channels), split them into groups of [28-30, 30-32, 32-34] and concatenate them together in batch_ Size dimension\n- Split fragment2 into two sub fragmentA and B, thus we have four parts of the data, we can achieve 4fold cross validation\n## 3. Data Augmentation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2F80b2b3e3d5d6a89f77c70a14eb5a75d3%2F1686900312738.jpg?generation=1686900353043727&alt=media)\n## 4. Model Selection\nI initially tried using Segformer, but the results of this model were not good at that time, so I did not use Segformer in subsequent experiments.Now it seems like I did something wrong.\nI have tried many different encoder models, such as resnet, resnext, seresnext, efficientnet, convnext, mit, and pvt, but **only mit and pvt have shown superior performance**, so all subsequent experiments are based on these two models.Meanwhile, it should be noted that overly complex and deep networks may lead to overfitting of the model on training data, so I chose **mit-b2 and pvt-v2-b2** as the encoder models.\n## 5. Improvement of Network Structure\n- All shape feature maps output by the encoder will go through an additional SK attention module\n- Then perform the following operations on the feature map that has undergone SK attention, as referenced by @hengck23 .\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fdf703992b4d39cfa8c5758910e4ddceb%2F1686902361015.png?generation=1686902393516010&alt=media)\nThe structure of self.weight1 is as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11574777%2Fccf94645bbad0c6de308f39056f3293a%2F1686902566772.png?generation=1686902593485222&alt=media)\n## 6. Other Training Details\n- Image Size: 224 * 224\n- Stride: 112\n- Loss Function: SoftBCELoss\n- Optimizer: AdamW\n- Initial LR: 0.0001\n- Scheduler: GradualWarmupSchedulerV2 with CosineAnnealingLR\n- Batch_size: 32\n- Epoch: 30\n## 7. Inference\n- Ensemble of 8 Unet models with mit, pvt as backbones\n- TTA in four directions(0°, 90°, 180°, 270°) in the inference stage\n- Mit & Pvt ensemble 👉Public LB: 0.721 Private LB: 0.586\n- Single Pvt 👉Public LB: 0.726 Private LB: 0.591\n## 8. Methods tried but did not work\n- DiceLoss❌\n- DeNoise module❌\n- Image size larger than 224 * 224❌\n- Adding attention only to the output of the encoder without adding modules, referenced by @hengck23 , is not effective❌\n- The ensemble of mit and pvt seems to have less good results than the individual pvt❌"
  }
}