{
  "id": 430242,
  "title": "3rd Place Solution - How to properly utilise noisy annotations",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/writeups/nischay-3rd-place-solution-how-to-properly-utilise",
  "author_name": "",
  "post_date": "2023-08-09T03:18:35.130Z",
  "votes": 61,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi everyone, sorry for taking a bit longer to publish the complete solution. Thanks to the Kaggle team and HubMap team for hosting the competition. In this post I’m going to explain my winning solution in detail. Again I am really happy that I became one of the youngest GrandMaster after this competition. This competition also taught how important it is to keep your updated and trying out the recent researches made in the field. </p>\n<p>Also, we have already made our inference notebook along with model weights public: you may visit that with this <a href=\"https://www.kaggle.com/code/nischaydnk/cv-wala-mega-ensemble-hubmap-2023\" target=\"_blank\">notebook link</a></p>\n<p>All codes related to training or preprocessing data codes are also made public (In Progress): <a href=\"https://github.com/Nischaydnk/HubMap-2023-3rd-Place-Solution\" target=\"_blank\">https://github.com/Nischaydnk/HubMap-2023-3rd-Place-Solution</a> </p>\n<p>You may find most of the configs I used in all_configs folder. I will clean up the repo in upcoming days.</p>\n<p>Here you can refer to the coco annotations used for training the model:  <a href=\"https://www.kaggle.com/datasets/nischaydnk/hubmap-coco-datasets\" target=\"_blank\">dataset link</a></p>\n<h2>Overview</h2>\n<p><strong>Winning solution consists of :</strong></p>\n<p><strong>5</strong> MMdet based models with different architectures.<br>\n<strong>2x ViT-Adapter-L</strong> (<a href=\"https://github.com/czczup/ViT-Adapter/tree/main/detection\" target=\"_blank\">https://github.com/czczup/ViT-Adapter/tree/main/detection</a>)<br>\n<strong>1x CBNetV2 Base</strong> (<a href=\"https://github.com/VDIGPKU/CBNetV2\" target=\"_blank\">https://github.com/VDIGPKU/CBNetV2</a>)<br>\n<strong>1x Detectors ResNeXt-101-32x4d</strong> (<a href=\"https://github.com/joe-siyuan-qiao/DetectoRS\" target=\"_blank\">https://github.com/joe-siyuan-qiao/DetectoRS</a>)<br>\n<strong>1x Detectors Resnet 50</strong> </p>\n<p>I also had few Vit Adapter based single models which could have placed me on 1st/2nd ranks but I didn't select. No regrets :))</p>\n<h2>Image Data Used</h2>\n<p>I only used competition for training models. <em>No external image data was used.</em> </p>\n<h2>How to use dataset 2??</h2>\n<p>Making the best use of dataset 2 was one of the key things to figure out in the competition. For me multi stage approach turned to be giving the highest boost. Basically during stage 1, A coco pretrained model will be loaded and pretrained on all the WSIs present in noisy annotations (dataset 2) for less epochs (~10), using really high learning rate (0.02+), with a cosine scheduler with minimum lr around (0.01), light augmentations.  </p>\n<p>In stage 2, we will load the pretained model from stage 1 and fine-tune it on dataset 1 with higher number of epochs (15-25), heavy augmentations, higher image resolution (for some models), slightly lower starting learning rate and minimum LR till 1e-7. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fdfbae6a6f634c755d9ea39ca2daa9751%2FScreenshot%202023-08-09%20at%203.36.59%20AM.png?generation=1691532480119818&amp;alt=media\" alt=\"\"></p>\n<p><em>I have used Pseudo labels using dataset 3 in training few models of final ensemble solution, although I didn't find any boost using them in the leaderboard scores, I will still talk about it as they were used in final solution</em></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Facf9ee081989a328c9f150458b4fce95%2FScreenshot%202023-08-09%20at%204.01.34%20AM.png?generation=1691533922010932&amp;alt=media\" alt=\"\"></p>\n<p>Using multistage approach gave around consistent 2-3% boost on validation and 4-6% improvement in leaderboard scores which is quite huge. The gap between cv &amp; lb scores was quite obvious as models were now learning on WSI 3 &amp; 4. Also, I never had to worry about using dilation or not as my later stage was just fine-tuned on dataset 1 (noise free annotations), so dilation doesn't help if applied directly on the masks.</p>\n<h1>Models Summary</h1>\n<h2>Vit Adapter: These models were published in recent ICLR 2023 &amp; turned out to be highest scoring architectures.</h2>\n<ul>\n<li>Pretrained coco weights were used. </li>\n<li>1400 x 1400 Image size (dataset fold-1 with pseudo threshold 0.5) &amp; (full data dataset 1 with pseudo threshold 0.6)</li>\n<li>Loss fnc: Mask Head loss multiplied by 2x in decoder.</li>\n<li>1200 x 1200 Image size used in stage 1.</li>\n<li>Cosine Scheduler with warmup were used.</li>\n<li>SGD optimizer for fold 1 model &amp; AdamW for full data model</li>\n<li>Higher Image Size + Multi Scale Inference (1600x1600, 1400x1400)</li>\n</ul>\n<p><strong>Best Public Leaderboard single model: 0.600</strong><br>\n<strong>Best Private Leaderboard single model: 0.589</strong></p>\n<h2>CBNetV2: Another popular set of architectures based on Swin transformers.</h2>\n<ul>\n<li>Pretrained coco weights were used. </li>\n<li>1600 x 1600 Image size (dataset 1 fold-5 without Pseudo)</li>\n<li>1400 x 1400 Image size used in stage 1.</li>\n<li>Cosine Scheduler with warmup were used.</li>\n<li>Higher Image Size during Inference (2048x2048)</li>\n<li>SGD optimizer </li>\n</ul>\n<p><strong>Best Public Leaderboard single model: 0.567</strong></p>\n<h2>Detectors HTC based models:  CNN based encoders for more diversity</h2>\n<ul>\n<li>Pretrained coco weights were used. </li>\n<li>2048 x 2048 image size (Resnet50 fold 1 w/ pseudo threshold 0.5 , Resnext101d without pseudo)</li>\n<li>Loss fnc: Mask Head loss 4x for Resnext101, 1x for Resnet50 </li>\n<li>Cosine Scheduler with warmup were used.</li>\n<li>SGD optimizer </li>\n</ul>\n<p><strong>Public Leaderboard single model: 0.573 ( resnext 101) , 0.558 (resnet50)</strong></p>\n<h2><strong>Techniques which provided consistent boost:</strong></h2>\n<ol>\n<li>Multi Stage Training</li>\n<li>Flip based Test Time Augmentation</li>\n<li>Higher weights to Mask head in HTC based models</li>\n<li>SGD optimizer </li>\n<li>Weighted Box Fusion for Ensemble</li>\n<li>Post Processing</li>\n</ol>\n<h2>Post Processing &amp; Ensemble</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F0fb5299755c5bfa0d6da43263a8be223%2FScreenshot%202023-08-09%20at%204.38.49%20AM.png?generation=1691536182102653&amp;alt=media\" alt=\"\"></p>\n<p>As mentioned earlier, I used WBF to do ensemble. To increase the diversity, I kept NMS for TTA and WBF for ensemble. Also, using both CNN / Transformer based encoders helped in increasing higher diversity and hence more impactful ensemble. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F764ea7995470a76b07549107c4b531a2%2FScreenshot%202023-08-09%20at%204.26.47%20AM.png?generation=1691536206638981&amp;alt=media\" alt=\"\"></p>\n<p>After the ensemble, I think some of my mask predictions got a little distorted. Therefore, I applied erosion followed by single iteration of dilation. This Post-processing gave me a decent amount of boost in both cross validation score as well as on leaderboard (+ 0.005)</p>\n<h2>Light Augmentations</h2>\n<pre><code>dict(\n    =,\n    direction=[, ],\n    =0.5),\ndict(\n    =,\n    policies=[[{\n        : ,\n        : 0.4,\n        : 0\n    }], [{\n        : ,\n        : 0.4,\n        : 5\n    }],\n              [{\n                  : ,\n                  : 1.0,\n                  : 6\n              }, {\n                  : \n              }]]),\ndict(\n    =,\n    transforms=[\n        dict(\n            =,\n            =0.0625,\n            =0.15,\n            =15,\n            =0.4)\n    ],\n    =dict(\n        =,\n        =,\n        label_fields=[],\n        =0.0,\n        =)\n</code></pre>\n<h2>Heavy Augmentations</h2>\n<pre><code>dict(\n                =,\n                direction=[, ],\n                =0.5),\n            dict(\n                =,\n                policies=[[{\n                    : ,\n                    : 0.4,\n                    : 0\n                }], [{\n                    : ,\n                    : 0.4,\n                    : 5\n                }],\n                          [{\n                              : ,\n                              : 0.6,\n                              : 10\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 3\n                          }],\n                          [{\n                              : ,\n                              : 0.6,\n                              : 10\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 5\n                          }],\n                          [{\n                              : ,\n                              : 32,\n                              : (0.5, 1.5),\n                              : 15\n                          }],\n                          [{\n                              : ,\n                              : (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              : 0.2\n                          }],\n                          [{\n                              :\n                              ,\n                              : (3, 8),\n                              : [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              : ,\n                              : 0.6\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 3\n                          }],\n                          [{\n                              : ,\n                              : 32,\n                              : (0.5, 1.5),\n                              : 18\n                          }],\n                          [{\n                              : ,\n                              : (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              : 0.3\n                          }],\n                          [{\n                              :\n                              ,\n                              : (5, 10),\n                              : [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              : ,\n                              : 0.6,\n                              : 4\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 6\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 10\n                          }],\n                          [{\n                              : ,\n                              : 1.0,\n                              : 6\n                          }, {\n                              : \n                          }]]),\n            dict(\n                =,\n                transforms=[\n                    dict(\n                        =,\n                        =0.0625,\n                        =0.15,\n                        =15,\n                        =0.5),\n                    dict(=, =0.5),\n                    dict(\n                        =,\n                        transforms=[\n                            dict(\n                                =,\n                                =120,\n                                =6.0,\n                                =3.5999999999999996,\n                                =1),\n                            dict(=, =1),\n                            dict(\n                                =,\n                                =2,\n                                =0.5,\n                                =1)\n                        ],\n                        =0.3)\n                ],\n                =dict(\n                    =,\n                    =,\n                    label_fields=[],\n                    =0.0,\n                    =) \n</code></pre>\n<p>Thank you all, I've tried my best to cover most part of my solution. Again, I am super happy to win the solo gold, feel free to reach out in case you find difficulty understanding any part of it.</p>",
  "messages": [
    {
      "id": "2381000",
      "postDate": "08/08/2023 23:23:54",
      "content": "<p>Hi everyone, sorry for taking a bit longer to publish the complete solution. Thanks to the Kaggle team and HubMap team for hosting the competition. In this post I’m going to explain my winning solution in detail. Again I am really happy that I became one of the youngest GrandMaster after this competition. This competition also taught how important it is to keep your updated and trying out the recent researches made in the field. </p>\n<p>Also, we have already made our inference notebook along with model weights public: you may visit that with this <a href=\"https://www.kaggle.com/code/nischaydnk/cv-wala-mega-ensemble-hubmap-2023\" target=\"_blank\">notebook link</a></p>\n<p>All codes related to training or preprocessing data codes are also made public (In Progress): <a href=\"https://github.com/Nischaydnk/HubMap-2023-3rd-Place-Solution\" target=\"_blank\">https://github.com/Nischaydnk/HubMap-2023-3rd-Place-Solution</a> </p>\n<p>You may find most of the configs I used in all_configs folder. I will clean up the repo in upcoming days.</p>\n<p>Here you can refer to the coco annotations used for training the model:  <a href=\"https://www.kaggle.com/datasets/nischaydnk/hubmap-coco-datasets\" target=\"_blank\">dataset link</a></p>\n<h2>Overview</h2>\n<p><strong>Winning solution consists of :</strong></p>\n<p><strong>5</strong> MMdet based models with different architectures.<br>\n<strong>2x ViT-Adapter-L</strong> (<a href=\"https://github.com/czczup/ViT-Adapter/tree/main/detection\" target=\"_blank\">https://github.com/czczup/ViT-Adapter/tree/main/detection</a>)<br>\n<strong>1x CBNetV2 Base</strong> (<a href=\"https://github.com/VDIGPKU/CBNetV2\" target=\"_blank\">https://github.com/VDIGPKU/CBNetV2</a>)<br>\n<strong>1x Detectors ResNeXt-101-32x4d</strong> (<a href=\"https://github.com/joe-siyuan-qiao/DetectoRS\" target=\"_blank\">https://github.com/joe-siyuan-qiao/DetectoRS</a>)<br>\n<strong>1x Detectors Resnet 50</strong> </p>\n<p>I also had few Vit Adapter based single models which could have placed me on 1st/2nd ranks but I didn't select. No regrets :))</p>\n<h2>Image Data Used</h2>\n<p>I only used competition for training models. <em>No external image data was used.</em> </p>\n<h2>How to use dataset 2??</h2>\n<p>Making the best use of dataset 2 was one of the key things to figure out in the competition. For me multi stage approach turned to be giving the highest boost. Basically during stage 1, A coco pretrained model will be loaded and pretrained on all the WSIs present in noisy annotations (dataset 2) for less epochs (~10), using really high learning rate (0.02+), with a cosine scheduler with minimum lr around (0.01), light augmentations.  </p>\n<p>In stage 2, we will load the pretained model from stage 1 and fine-tune it on dataset 1 with higher number of epochs (15-25), heavy augmentations, higher image resolution (for some models), slightly lower starting learning rate and minimum LR till 1e-7. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fdfbae6a6f634c755d9ea39ca2daa9751%2FScreenshot%202023-08-09%20at%203.36.59%20AM.png?generation=1691532480119818&amp;alt=media\" alt=\"\"></p>\n<p><em>I have used Pseudo labels using dataset 3 in training few models of final ensemble solution, although I didn't find any boost using them in the leaderboard scores, I will still talk about it as they were used in final solution</em></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Facf9ee081989a328c9f150458b4fce95%2FScreenshot%202023-08-09%20at%204.01.34%20AM.png?generation=1691533922010932&amp;alt=media\" alt=\"\"></p>\n<p>Using multistage approach gave around consistent 2-3% boost on validation and 4-6% improvement in leaderboard scores which is quite huge. The gap between cv &amp; lb scores was quite obvious as models were now learning on WSI 3 &amp; 4. Also, I never had to worry about using dilation or not as my later stage was just fine-tuned on dataset 1 (noise free annotations), so dilation doesn't help if applied directly on the masks.</p>\n<h1>Models Summary</h1>\n<h2>Vit Adapter: These models were published in recent ICLR 2023 &amp; turned out to be highest scoring architectures.</h2>\n<ul>\n<li>Pretrained coco weights were used. </li>\n<li>1400 x 1400 Image size (dataset fold-1 with pseudo threshold 0.5) &amp; (full data dataset 1 with pseudo threshold 0.6)</li>\n<li>Loss fnc: Mask Head loss multiplied by 2x in decoder.</li>\n<li>1200 x 1200 Image size used in stage 1.</li>\n<li>Cosine Scheduler with warmup were used.</li>\n<li>SGD optimizer for fold 1 model &amp; AdamW for full data model</li>\n<li>Higher Image Size + Multi Scale Inference (1600x1600, 1400x1400)</li>\n</ul>\n<p><strong>Best Public Leaderboard single model: 0.600</strong><br>\n<strong>Best Private Leaderboard single model: 0.589</strong></p>\n<h2>CBNetV2: Another popular set of architectures based on Swin transformers.</h2>\n<ul>\n<li>Pretrained coco weights were used. </li>\n<li>1600 x 1600 Image size (dataset 1 fold-5 without Pseudo)</li>\n<li>1400 x 1400 Image size used in stage 1.</li>\n<li>Cosine Scheduler with warmup were used.</li>\n<li>Higher Image Size during Inference (2048x2048)</li>\n<li>SGD optimizer </li>\n</ul>\n<p><strong>Best Public Leaderboard single model: 0.567</strong></p>\n<h2>Detectors HTC based models:  CNN based encoders for more diversity</h2>\n<ul>\n<li>Pretrained coco weights were used. </li>\n<li>2048 x 2048 image size (Resnet50 fold 1 w/ pseudo threshold 0.5 , Resnext101d without pseudo)</li>\n<li>Loss fnc: Mask Head loss 4x for Resnext101, 1x for Resnet50 </li>\n<li>Cosine Scheduler with warmup were used.</li>\n<li>SGD optimizer </li>\n</ul>\n<p><strong>Public Leaderboard single model: 0.573 ( resnext 101) , 0.558 (resnet50)</strong></p>\n<h2><strong>Techniques which provided consistent boost:</strong></h2>\n<ol>\n<li>Multi Stage Training</li>\n<li>Flip based Test Time Augmentation</li>\n<li>Higher weights to Mask head in HTC based models</li>\n<li>SGD optimizer </li>\n<li>Weighted Box Fusion for Ensemble</li>\n<li>Post Processing</li>\n</ol>\n<h2>Post Processing &amp; Ensemble</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F0fb5299755c5bfa0d6da43263a8be223%2FScreenshot%202023-08-09%20at%204.38.49%20AM.png?generation=1691536182102653&amp;alt=media\" alt=\"\"></p>\n<p>As mentioned earlier, I used WBF to do ensemble. To increase the diversity, I kept NMS for TTA and WBF for ensemble. Also, using both CNN / Transformer based encoders helped in increasing higher diversity and hence more impactful ensemble. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F764ea7995470a76b07549107c4b531a2%2FScreenshot%202023-08-09%20at%204.26.47%20AM.png?generation=1691536206638981&amp;alt=media\" alt=\"\"></p>\n<p>After the ensemble, I think some of my mask predictions got a little distorted. Therefore, I applied erosion followed by single iteration of dilation. This Post-processing gave me a decent amount of boost in both cross validation score as well as on leaderboard (+ 0.005)</p>\n<h2>Light Augmentations</h2>\n<pre><code>dict(\n    =,\n    direction=[, ],\n    =0.5),\ndict(\n    =,\n    policies=[[{\n        : ,\n        : 0.4,\n        : 0\n    }], [{\n        : ,\n        : 0.4,\n        : 5\n    }],\n              [{\n                  : ,\n                  : 1.0,\n                  : 6\n              }, {\n                  : \n              }]]),\ndict(\n    =,\n    transforms=[\n        dict(\n            =,\n            =0.0625,\n            =0.15,\n            =15,\n            =0.4)\n    ],\n    =dict(\n        =,\n        =,\n        label_fields=[],\n        =0.0,\n        =)\n</code></pre>\n<h2>Heavy Augmentations</h2>\n<pre><code>dict(\n                =,\n                direction=[, ],\n                =0.5),\n            dict(\n                =,\n                policies=[[{\n                    : ,\n                    : 0.4,\n                    : 0\n                }], [{\n                    : ,\n                    : 0.4,\n                    : 5\n                }],\n                          [{\n                              : ,\n                              : 0.6,\n                              : 10\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 3\n                          }],\n                          [{\n                              : ,\n                              : 0.6,\n                              : 10\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 5\n                          }],\n                          [{\n                              : ,\n                              : 32,\n                              : (0.5, 1.5),\n                              : 15\n                          }],\n                          [{\n                              : ,\n                              : (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              : 0.2\n                          }],\n                          [{\n                              :\n                              ,\n                              : (3, 8),\n                              : [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              : ,\n                              : 0.6\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 3\n                          }],\n                          [{\n                              : ,\n                              : 32,\n                              : (0.5, 1.5),\n                              : 18\n                          }],\n                          [{\n                              : ,\n                              : (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              : 0.3\n                          }],\n                          [{\n                              :\n                              ,\n                              : (5, 10),\n                              : [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              : ,\n                              : 0.6,\n                              : 4\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 6\n                          }, {\n                              : ,\n                              : 0.6,\n                              : 10\n                          }],\n                          [{\n                              : ,\n                              : 1.0,\n                              : 6\n                          }, {\n                              : \n                          }]]),\n            dict(\n                =,\n                transforms=[\n                    dict(\n                        =,\n                        =0.0625,\n                        =0.15,\n                        =15,\n                        =0.5),\n                    dict(=, =0.5),\n                    dict(\n                        =,\n                        transforms=[\n                            dict(\n                                =,\n                                =120,\n                                =6.0,\n                                =3.5999999999999996,\n                                =1),\n                            dict(=, =1),\n                            dict(\n                                =,\n                                =2,\n                                =0.5,\n                                =1)\n                        ],\n                        =0.3)\n                ],\n                =dict(\n                    =,\n                    =,\n                    label_fields=[],\n                    =0.0,\n                    =) \n</code></pre>\n<p>Thank you all, I've tried my best to cover most part of my solution. Again, I am super happy to win the solo gold, feel free to reach out in case you find difficulty understanding any part of it.</p>",
      "rawMarkdown": "Hi everyone, sorry for taking a bit longer to publish the complete solution. Thanks to the Kaggle team and HubMap team for hosting the competition. In this post I’m going to explain my winning solution in detail. Again I am really happy that I became one of the youngest GrandMaster after this competition. This competition also taught how important it is to keep your updated and trying out the recent researches made in the field. \n\nAlso, we have already made our inference notebook along with model weights public: you may visit that with this [notebook link](https://www.kaggle.com/code/nischaydnk/cv-wala-mega-ensemble-hubmap-2023)\n \nAll codes related to training or preprocessing data codes are also made public (In Progress): https://github.com/Nischaydnk/HubMap-2023-3rd-Place-Solution \n\nYou may find most of the configs I used in all_configs folder. I will clean up the repo in upcoming days.\n\nHere you can refer to the coco annotations used for training the model:  [dataset link](https://www.kaggle.com/datasets/nischaydnk/hubmap-coco-datasets)\n\n## Overview\n\n**Winning solution consists of :**\n\n**5** MMdet based models with different architectures.\n**2x ViT-Adapter-L** (https://github.com/czczup/ViT-Adapter/tree/main/detection)\n**1x CBNetV2 Base** (https://github.com/VDIGPKU/CBNetV2)\n**1x Detectors ResNeXt-101-32x4d** (https://github.com/joe-siyuan-qiao/DetectoRS)\n**1x Detectors Resnet 50** \n\nI also had few Vit Adapter based single models which could have placed me on 1st/2nd ranks but I didn't select. No regrets :))\n\n## Image Data Used\n\nI only used competition for training models. *No external image data was used.* \n\n## How to use dataset 2??\n\nMaking the best use of dataset 2 was one of the key things to figure out in the competition. For me multi stage approach turned to be giving the highest boost. Basically during stage 1, A coco pretrained model will be loaded and pretrained on all the WSIs present in noisy annotations (dataset 2) for less epochs (~10), using really high learning rate (0.02+), with a cosine scheduler with minimum lr around (0.01), light augmentations.  \n\nIn stage 2, we will load the pretained model from stage 1 and fine-tune it on dataset 1 with higher number of epochs (15-25), heavy augmentations, higher image resolution (for some models), slightly lower starting learning rate and minimum LR till 1e-7. \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fdfbae6a6f634c755d9ea39ca2daa9751%2FScreenshot%202023-08-09%20at%203.36.59%20AM.png?generation=1691532480119818&alt=media)\n\n*I have used Pseudo labels using dataset 3 in training few models of final ensemble solution, although I didn't find any boost using them in the leaderboard scores, I will still talk about it as they were used in final solution*\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Facf9ee081989a328c9f150458b4fce95%2FScreenshot%202023-08-09%20at%204.01.34%20AM.png?generation=1691533922010932&alt=media)\n\n\nUsing multistage approach gave around consistent 2-3% boost on validation and 4-6% improvement in leaderboard scores which is quite huge. The gap between cv & lb scores was quite obvious as models were now learning on WSI 3 & 4. Also, I never had to worry about using dilation or not as my later stage was just fine-tuned on dataset 1 (noise free annotations), so dilation doesn't help if applied directly on the masks.\n\n# Models Summary\n\n## Vit Adapter: These models were published in recent ICLR 2023 & turned out to be highest scoring architectures. \n- Pretrained coco weights were used. \n- 1400 x 1400 Image size (dataset fold-1 with pseudo threshold 0.5) & (full data dataset 1 with pseudo threshold 0.6)\n- Loss fnc: Mask Head loss multiplied by 2x in decoder.\n- 1200 x 1200 Image size used in stage 1.\n- Cosine Scheduler with warmup were used.\n- SGD optimizer for fold 1 model & AdamW for full data model\n- Higher Image Size + Multi Scale Inference (1600x1600, 1400x1400)\n\n**Best Public Leaderboard single model: 0.600**\n**Best Private Leaderboard single model: 0.589**\n\n\n##CBNetV2: Another popular set of architectures based on Swin transformers. \n- Pretrained coco weights were used. \n- 1600 x 1600 Image size (dataset 1 fold-5 without Pseudo)\n- 1400 x 1400 Image size used in stage 1.\n- Cosine Scheduler with warmup were used.\n- Higher Image Size during Inference (2048x2048)\n- SGD optimizer \n\n**Best Public Leaderboard single model: 0.567**\n\n## Detectors HTC based models:  CNN based encoders for more diversity\n- Pretrained coco weights were used. \n- 2048 x 2048 image size (Resnet50 fold 1 w/ pseudo threshold 0.5 , Resnext101d without pseudo)\n- Loss fnc: Mask Head loss 4x for Resnext101, 1x for Resnet50 \n- Cosine Scheduler with warmup were used.\n- SGD optimizer \n\n**Public Leaderboard single model: 0.573 ( resnext 101) , 0.558 (resnet50)**\n\n\n##**Techniques which provided consistent boost:**\n\n  1. Multi Stage Training\n  2. Flip based Test Time Augmentation\n  3. Higher weights to Mask head in HTC based models\n  4. SGD optimizer \n  5. Weighted Box Fusion for Ensemble\n  6. Post Processing\n\n\n## Post Processing & Ensemble\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F0fb5299755c5bfa0d6da43263a8be223%2FScreenshot%202023-08-09%20at%204.38.49%20AM.png?generation=1691536182102653&alt=media)\n\nAs mentioned earlier, I used WBF to do ensemble. To increase the diversity, I kept NMS for TTA and WBF for ensemble. Also, using both CNN / Transformer based encoders helped in increasing higher diversity and hence more impactful ensemble. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F764ea7995470a76b07549107c4b531a2%2FScreenshot%202023-08-09%20at%204.26.47%20AM.png?generation=1691536206638981&alt=media)\n\nAfter the ensemble, I think some of my mask predictions got a little distorted. Therefore, I applied erosion followed by single iteration of dilation. This Post-processing gave me a decent amount of boost in both cross validation score as well as on leaderboard (+ 0.005)\n \n\n##Light Augmentations\n\n```\ndict(\n    type='RandomFlip',\n    direction=['horizontal', 'vertical'],\n    flip_ratio=0.5),\ndict(\n    type='AutoAugment',\n    policies=[[{\n        'type': 'Shear',\n        'prob': 0.4,\n        'level': 0\n    }], [{\n        'type': 'Translate',\n        'prob': 0.4,\n        'level': 5\n    }],\n              [{\n                  'type': 'ColorTransform',\n                  'prob': 1.0,\n                  'level': 6\n              }, {\n                  'type': 'EqualizeTransform'\n              }]]),\ndict(\n    type='Albu',\n    transforms=[\n        dict(\n            type='ShiftScaleRotate',\n            shift_limit=0.0625,\n            scale_limit=0.15,\n            rotate_limit=15,\n            p=0.4)\n    ],\n    bbox_params=dict(\n        type='BboxParams',\n        format='pascal_voc',\n        label_fields=['gt_labels'],\n        min_visibility=0.0,\n        filter_lost_elements=True)\n```\n\n\n## Heavy Augmentations\n\n```\ndict(\n                type='RandomFlip',\n                direction=['horizontal', 'vertical'],\n                flip_ratio=0.5),\n            dict(\n                type='AutoAugment',\n                policies=[[{\n                    'type': 'Shear',\n                    'prob': 0.4,\n                    'level': 0\n                }], [{\n                    'type': 'Translate',\n                    'prob': 0.4,\n                    'level': 5\n                }],\n                          [{\n                              'type': 'ColorTransform',\n                              'prob': 0.6,\n                              'level': 10\n                          }, {\n                              'type': 'BrightnessTransform',\n                              'prob': 0.6,\n                              'level': 3\n                          }],\n                          [{\n                              'type': 'ColorTransform',\n                              'prob': 0.6,\n                              'level': 10\n                          }, {\n                              'type': 'ContrastTransform',\n                              'prob': 0.6,\n                              'level': 5\n                          }],\n                          [{\n                              'type': 'PhotoMetricDistortion',\n                              'brightness_delta': 32,\n                              'contrast_range': (0.5, 1.5),\n                              'hue_delta': 15\n                          }],\n                          [{\n                              'type': 'MinIoURandomCrop',\n                              'min_ious': (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              'min_crop_size': 0.2\n                          }],\n                          [{\n                              'type':\n                              'CutOut',\n                              'n_holes': (3, 8),\n                              'cutout_shape': [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              'type': 'EqualizeTransform',\n                              'prob': 0.6\n                          }, {\n                              'type': 'BrightnessTransform',\n                              'prob': 0.6,\n                              'level': 3\n                          }],\n                          [{\n                              'type': 'PhotoMetricDistortion',\n                              'brightness_delta': 32,\n                              'contrast_range': (0.5, 1.5),\n                              'hue_delta': 18\n                          }],\n                          [{\n                              'type': 'MinIoURandomCrop',\n                              'min_ious': (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              'min_crop_size': 0.3\n                          }],\n                          [{\n                              'type':\n                              'CutOut',\n                              'n_holes': (5, 10),\n                              'cutout_shape': [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              'type': 'BrightnessTransform',\n                              'prob': 0.6,\n                              'level': 4\n                          }, {\n                              'type': 'ContrastTransform',\n                              'prob': 0.6,\n                              'level': 6\n                          }, {\n                              'type': 'Rotate',\n                              'prob': 0.6,\n                              'level': 10\n                          }],\n                          [{\n                              'type': 'ColorTransform',\n                              'prob': 1.0,\n                              'level': 6\n                          }, {\n                              'type': 'EqualizeTransform'\n                          }]]),\n            dict(\n                type='Albu',\n                transforms=[\n                    dict(\n                        type='ShiftScaleRotate',\n                        shift_limit=0.0625,\n                        scale_limit=0.15,\n                        rotate_limit=15,\n                        p=0.5),\n                    dict(type='RandomRotate90', p=0.5),\n                    dict(\n                        type='OneOf',\n                        transforms=[\n                            dict(\n                                type='ElasticTransform',\n                                alpha=120,\n                                sigma=6.0,\n                                alpha_affine=3.5999999999999996,\n                                p=1),\n                            dict(type='GridDistortion', p=1),\n                            dict(\n                                type='OpticalDistortion',\n                                distort_limit=2,\n                                shift_limit=0.5,\n                                p=1)\n                        ],\n                        p=0.3)\n                ],\n                bbox_params=dict(\n                    type='BboxParams',\n                    format='pascal_voc',\n                    label_fields=['gt_labels'],\n                    min_visibility=0.0,\n                    filter_lost_elements=True) \n```\n\n\nThank you all, I've tried my best to cover most part of my solution. Again, I am super happy to win the solo gold, feel free to reach out in case you find difficulty understanding any part of it.",
      "votes": null
    },
    {
      "id": "2381165",
      "postDate": "08/09/2023 03:13:22",
      "content": "<blockquote>\n  <p>This competition also taught how important it is to keep your updated and trying out the recent researches made in the field.</p>\n</blockquote>\n<p>So will you finally start reading LLM Papers as well? 😁</p>",
      "rawMarkdown": "> This competition also taught how important it is to keep your updated and trying out the recent researches made in the field.\n\nSo will you finally start reading LLM Papers as well? 😁",
      "votes": null
    },
    {
      "id": "2381263",
      "postDate": "08/09/2023 05:31:22",
      "content": "<p>Congratulations on achieving GM Norm and also getting top spot.<br>\nThanks for sharing the interesting info about your approach. Good illustrations for understanding the techniques applied in your notebook. </p>",
      "rawMarkdown": "Congratulations on achieving GM Norm and also getting top spot.\nThanks for sharing the interesting info about your approach. Good illustrations for understanding the techniques applied in your notebook.",
      "votes": null
    },
    {
      "id": "2381448",
      "postDate": "08/09/2023 07:42:26",
      "content": "<p>What an awesome piece of paper <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>, you published for the community. Thanks for showing your great work to the community.👍</p>",
      "rawMarkdown": "What an awesome piece of paper @nischaydnk, you published for the community. Thanks for showing your great work to the community.👍",
      "votes": null
    },
    {
      "id": "2382234",
      "postDate": "08/09/2023 16:50:14",
      "content": "<p>Solution worthy of the GM title :) Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Solution worthy of the GM title :) Congrats and thanks for sharing!",
      "votes": null
    },
    {
      "id": "2383370",
      "postDate": "08/10/2023 10:22:19",
      "content": "<p>Hello, Congrats for the 3rd place and the Competition GM!<br>\nI want to know more detail about how you split the train - valid set in the \"Multi Stage Approach\" and \"Multi Stage Pseudo Label Approach\". Like how to you split so that i can guarantee that there will be no data leakage and you can trust that split </p>\n<p>Thank you</p>",
      "rawMarkdown": "Hello, Congrats for the 3rd place and the Competition GM!\nI want to know more detail about how you split the train - valid set in the \"Multi Stage Approach\" and \"Multi Stage Pseudo Label Approach\". Like how to you split so that i can guarantee that there will be no data leakage and you can trust that split \n\nThank you",
      "votes": null
    },
    {
      "id": "2385862",
      "postDate": "08/11/2023 15:41:41",
      "content": "<p>Thanks for reminding about split, I will update the solution. I simply used kfold on ds1, so basically in stage 1 , training data is kept ds2 (complete) , validation split is single fold from ds1. In stage 2, 4 folds from ds1 in training keeping the same validation fold.</p>",
      "rawMarkdown": "Thanks for reminding about split, I will update the solution. I simply used kfold on ds1, so basically in stage 1 , training data is kept ds2 (complete) , validation split is single fold from ds1. In stage 2, 4 folds from ds1 in training keeping the same validation fold.",
      "votes": null
    },
    {
      "id": "2386167",
      "postDate": "08/11/2023 18:45:45",
      "content": "<p>Thank you 😀</p>",
      "rawMarkdown": "Thank you 😀",
      "votes": null
    },
    {
      "id": "2386809",
      "postDate": "08/12/2023 06:54:09",
      "content": "<p>Congratulations and thank you for the report <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> !<br>\nCould you elaborate on how you generated masks? Have you used a single mask head with WBF fused boxes?</p>",
      "rawMarkdown": "Congratulations and thank you for the report @nischaydnk !\nCould you elaborate on how you generated masks? Have you used a single mask head with WBF fused boxes?",
      "votes": null
    },
    {
      "id": "2387711",
      "postDate": "08/12/2023 22:35:11",
      "content": "<p>Your work has merit and seems very well supported… I have experience in hardware and in the future I hope to implement a medical application, and I will be inspired by your work… Kind regards and congrats!</p>",
      "rawMarkdown": "Your work has merit and seems very well supported... I have experience in hardware and in the future I hope to implement a medical application, and I will be inspired by your work... Kind regards and congrats!",
      "votes": null
    },
    {
      "id": "2389303",
      "postDate": "08/14/2023 00:49:03",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "2389440",
      "postDate": "08/14/2023 04:03:12",
      "content": "<p>Thank you for the explanatory report. Learned a few steps and will try using those. Thank you and also many Congratulations ! <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "rawMarkdown": "Thank you for the explanatory report. Learned a few steps and will try using those. Thank you and also many Congratulations ! @nischaydnk",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2381165,
      "author_name": "init27",
      "author_url": "",
      "post_date": "08/09/2023 03:13:22",
      "content": "<blockquote>\n  <p>This competition also taught how important it is to keep your updated and trying out the recent researches made in the field.</p>\n</blockquote>\n<p>So will you finally start reading LLM Papers as well? 😁</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2381263,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "08/09/2023 05:31:22",
      "content": "<p>Congratulations on achieving GM Norm and also getting top spot.<br>\nThanks for sharing the interesting info about your approach. Good illustrations for understanding the techniques applied in your notebook. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2381448,
      "author_name": "tariqbashir",
      "author_url": "",
      "post_date": "08/09/2023 07:42:26",
      "content": "<p>What an awesome piece of paper <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a>, you published for the community. Thanks for showing your great work to the community.👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2382234,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "08/09/2023 16:50:14",
      "content": "<p>Solution worthy of the GM title :) Congrats and thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2386167,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "08/11/2023 18:45:45",
          "content": "<p>Thank you 😀</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2383370,
      "author_name": "researchbntz",
      "author_url": "",
      "post_date": "08/10/2023 10:22:19",
      "content": "<p>Hello, Congrats for the 3rd place and the Competition GM!<br>\nI want to know more detail about how you split the train - valid set in the \"Multi Stage Approach\" and \"Multi Stage Pseudo Label Approach\". Like how to you split so that i can guarantee that there will be no data leakage and you can trust that split </p>\n<p>Thank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 2385862,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "08/11/2023 15:41:41",
          "content": "<p>Thanks for reminding about split, I will update the solution. I simply used kfold on ds1, so basically in stage 1 , training data is kept ds2 (complete) , validation split is single fold from ds1. In stage 2, 4 folds from ds1 in training keeping the same validation fold.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2389303,
              "author_name": "researchbntz",
              "author_url": "",
              "post_date": "08/14/2023 00:49:03",
              "content": "<p>Thank you!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2386809,
      "author_name": "vslaykovsky",
      "author_url": "",
      "post_date": "08/12/2023 06:54:09",
      "content": "<p>Congratulations and thank you for the report <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> !<br>\nCould you elaborate on how you generated masks? Have you used a single mask head with WBF fused boxes?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2387711,
      "author_name": "guillermoperezg",
      "author_url": "",
      "post_date": "08/12/2023 22:35:11",
      "content": "<p>Your work has merit and seems very well supported… I have experience in hardware and in the future I hope to implement a medical application, and I will be inspired by your work… Kind regards and congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2389440,
      "author_name": "sanidhyajadaun",
      "author_url": "",
      "post_date": "08/14/2023 04:03:12",
      "content": "<p>Thank you for the explanatory report. Learned a few steps and will try using those. Thank you and also many Congratulations ! <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2381000": "Hi everyone, sorry for taking a bit longer to publish the complete solution. Thanks to the Kaggle team and HubMap team for hosting the competition. In this post I’m going to explain my winning solution in detail. Again I am really happy that I became one of the youngest GrandMaster after this competition. This competition also taught how important it is to keep your updated and trying out the recent researches made in the field. \n\nAlso, we have already made our inference notebook along with model weights public: you may visit that with this [notebook link](https://www.kaggle.com/code/nischaydnk/cv-wala-mega-ensemble-hubmap-2023)\n \nAll codes related to training or preprocessing data codes are also made public (In Progress): https://github.com/Nischaydnk/HubMap-2023-3rd-Place-Solution \n\nYou may find most of the configs I used in all_configs folder. I will clean up the repo in upcoming days.\n\nHere you can refer to the coco annotations used for training the model:  [dataset link](https://www.kaggle.com/datasets/nischaydnk/hubmap-coco-datasets)\n\n## Overview\n\n**Winning solution consists of :**\n\n**5** MMdet based models with different architectures.\n**2x ViT-Adapter-L** (https://github.com/czczup/ViT-Adapter/tree/main/detection)\n**1x CBNetV2 Base** (https://github.com/VDIGPKU/CBNetV2)\n**1x Detectors ResNeXt-101-32x4d** (https://github.com/joe-siyuan-qiao/DetectoRS)\n**1x Detectors Resnet 50** \n\nI also had few Vit Adapter based single models which could have placed me on 1st/2nd ranks but I didn't select. No regrets :))\n\n## Image Data Used\n\nI only used competition for training models. *No external image data was used.* \n\n## How to use dataset 2??\n\nMaking the best use of dataset 2 was one of the key things to figure out in the competition. For me multi stage approach turned to be giving the highest boost. Basically during stage 1, A coco pretrained model will be loaded and pretrained on all the WSIs present in noisy annotations (dataset 2) for less epochs (~10), using really high learning rate (0.02+), with a cosine scheduler with minimum lr around (0.01), light augmentations.  \n\nIn stage 2, we will load the pretained model from stage 1 and fine-tune it on dataset 1 with higher number of epochs (15-25), heavy augmentations, higher image resolution (for some models), slightly lower starting learning rate and minimum LR till 1e-7. \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fdfbae6a6f634c755d9ea39ca2daa9751%2FScreenshot%202023-08-09%20at%203.36.59%20AM.png?generation=1691532480119818&alt=media)\n\n*I have used Pseudo labels using dataset 3 in training few models of final ensemble solution, although I didn't find any boost using them in the leaderboard scores, I will still talk about it as they were used in final solution*\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Facf9ee081989a328c9f150458b4fce95%2FScreenshot%202023-08-09%20at%204.01.34%20AM.png?generation=1691533922010932&alt=media)\n\n\nUsing multistage approach gave around consistent 2-3% boost on validation and 4-6% improvement in leaderboard scores which is quite huge. The gap between cv & lb scores was quite obvious as models were now learning on WSI 3 & 4. Also, I never had to worry about using dilation or not as my later stage was just fine-tuned on dataset 1 (noise free annotations), so dilation doesn't help if applied directly on the masks.\n\n# Models Summary\n\n## Vit Adapter: These models were published in recent ICLR 2023 & turned out to be highest scoring architectures. \n- Pretrained coco weights were used. \n- 1400 x 1400 Image size (dataset fold-1 with pseudo threshold 0.5) & (full data dataset 1 with pseudo threshold 0.6)\n- Loss fnc: Mask Head loss multiplied by 2x in decoder.\n- 1200 x 1200 Image size used in stage 1.\n- Cosine Scheduler with warmup were used.\n- SGD optimizer for fold 1 model & AdamW for full data model\n- Higher Image Size + Multi Scale Inference (1600x1600, 1400x1400)\n\n**Best Public Leaderboard single model: 0.600**\n**Best Private Leaderboard single model: 0.589**\n\n\n##CBNetV2: Another popular set of architectures based on Swin transformers. \n- Pretrained coco weights were used. \n- 1600 x 1600 Image size (dataset 1 fold-5 without Pseudo)\n- 1400 x 1400 Image size used in stage 1.\n- Cosine Scheduler with warmup were used.\n- Higher Image Size during Inference (2048x2048)\n- SGD optimizer \n\n**Best Public Leaderboard single model: 0.567**\n\n## Detectors HTC based models:  CNN based encoders for more diversity\n- Pretrained coco weights were used. \n- 2048 x 2048 image size (Resnet50 fold 1 w/ pseudo threshold 0.5 , Resnext101d without pseudo)\n- Loss fnc: Mask Head loss 4x for Resnext101, 1x for Resnet50 \n- Cosine Scheduler with warmup were used.\n- SGD optimizer \n\n**Public Leaderboard single model: 0.573 ( resnext 101) , 0.558 (resnet50)**\n\n\n##**Techniques which provided consistent boost:**\n\n  1. Multi Stage Training\n  2. Flip based Test Time Augmentation\n  3. Higher weights to Mask head in HTC based models\n  4. SGD optimizer \n  5. Weighted Box Fusion for Ensemble\n  6. Post Processing\n\n\n## Post Processing & Ensemble\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F0fb5299755c5bfa0d6da43263a8be223%2FScreenshot%202023-08-09%20at%204.38.49%20AM.png?generation=1691536182102653&alt=media)\n\nAs mentioned earlier, I used WBF to do ensemble. To increase the diversity, I kept NMS for TTA and WBF for ensemble. Also, using both CNN / Transformer based encoders helped in increasing higher diversity and hence more impactful ensemble. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F764ea7995470a76b07549107c4b531a2%2FScreenshot%202023-08-09%20at%204.26.47%20AM.png?generation=1691536206638981&alt=media)\n\nAfter the ensemble, I think some of my mask predictions got a little distorted. Therefore, I applied erosion followed by single iteration of dilation. This Post-processing gave me a decent amount of boost in both cross validation score as well as on leaderboard (+ 0.005)\n \n\n##Light Augmentations\n\n```\ndict(\n    type='RandomFlip',\n    direction=['horizontal', 'vertical'],\n    flip_ratio=0.5),\ndict(\n    type='AutoAugment',\n    policies=[[{\n        'type': 'Shear',\n        'prob': 0.4,\n        'level': 0\n    }], [{\n        'type': 'Translate',\n        'prob': 0.4,\n        'level': 5\n    }],\n              [{\n                  'type': 'ColorTransform',\n                  'prob': 1.0,\n                  'level': 6\n              }, {\n                  'type': 'EqualizeTransform'\n              }]]),\ndict(\n    type='Albu',\n    transforms=[\n        dict(\n            type='ShiftScaleRotate',\n            shift_limit=0.0625,\n            scale_limit=0.15,\n            rotate_limit=15,\n            p=0.4)\n    ],\n    bbox_params=dict(\n        type='BboxParams',\n        format='pascal_voc',\n        label_fields=['gt_labels'],\n        min_visibility=0.0,\n        filter_lost_elements=True)\n```\n\n\n## Heavy Augmentations\n\n```\ndict(\n                type='RandomFlip',\n                direction=['horizontal', 'vertical'],\n                flip_ratio=0.5),\n            dict(\n                type='AutoAugment',\n                policies=[[{\n                    'type': 'Shear',\n                    'prob': 0.4,\n                    'level': 0\n                }], [{\n                    'type': 'Translate',\n                    'prob': 0.4,\n                    'level': 5\n                }],\n                          [{\n                              'type': 'ColorTransform',\n                              'prob': 0.6,\n                              'level': 10\n                          }, {\n                              'type': 'BrightnessTransform',\n                              'prob': 0.6,\n                              'level': 3\n                          }],\n                          [{\n                              'type': 'ColorTransform',\n                              'prob': 0.6,\n                              'level': 10\n                          }, {\n                              'type': 'ContrastTransform',\n                              'prob': 0.6,\n                              'level': 5\n                          }],\n                          [{\n                              'type': 'PhotoMetricDistortion',\n                              'brightness_delta': 32,\n                              'contrast_range': (0.5, 1.5),\n                              'hue_delta': 15\n                          }],\n                          [{\n                              'type': 'MinIoURandomCrop',\n                              'min_ious': (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              'min_crop_size': 0.2\n                          }],\n                          [{\n                              'type':\n                              'CutOut',\n                              'n_holes': (3, 8),\n                              'cutout_shape': [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              'type': 'EqualizeTransform',\n                              'prob': 0.6\n                          }, {\n                              'type': 'BrightnessTransform',\n                              'prob': 0.6,\n                              'level': 3\n                          }],\n                          [{\n                              'type': 'PhotoMetricDistortion',\n                              'brightness_delta': 32,\n                              'contrast_range': (0.5, 1.5),\n                              'hue_delta': 18\n                          }],\n                          [{\n                              'type': 'MinIoURandomCrop',\n                              'min_ious': (0.4, 0.5, 0.6, 0.7, 0.8, 0.9),\n                              'min_crop_size': 0.3\n                          }],\n                          [{\n                              'type':\n                              'CutOut',\n                              'n_holes': (5, 10),\n                              'cutout_shape': [(4, 4), (4, 8), (8, 4), (8, 8),\n                                               (16, 32), (32, 16), (32, 32),\n                                               (32, 48), (48, 32), (48, 48)]\n                          }],\n                          [{\n                              'type': 'BrightnessTransform',\n                              'prob': 0.6,\n                              'level': 4\n                          }, {\n                              'type': 'ContrastTransform',\n                              'prob': 0.6,\n                              'level': 6\n                          }, {\n                              'type': 'Rotate',\n                              'prob': 0.6,\n                              'level': 10\n                          }],\n                          [{\n                              'type': 'ColorTransform',\n                              'prob': 1.0,\n                              'level': 6\n                          }, {\n                              'type': 'EqualizeTransform'\n                          }]]),\n            dict(\n                type='Albu',\n                transforms=[\n                    dict(\n                        type='ShiftScaleRotate',\n                        shift_limit=0.0625,\n                        scale_limit=0.15,\n                        rotate_limit=15,\n                        p=0.5),\n                    dict(type='RandomRotate90', p=0.5),\n                    dict(\n                        type='OneOf',\n                        transforms=[\n                            dict(\n                                type='ElasticTransform',\n                                alpha=120,\n                                sigma=6.0,\n                                alpha_affine=3.5999999999999996,\n                                p=1),\n                            dict(type='GridDistortion', p=1),\n                            dict(\n                                type='OpticalDistortion',\n                                distort_limit=2,\n                                shift_limit=0.5,\n                                p=1)\n                        ],\n                        p=0.3)\n                ],\n                bbox_params=dict(\n                    type='BboxParams',\n                    format='pascal_voc',\n                    label_fields=['gt_labels'],\n                    min_visibility=0.0,\n                    filter_lost_elements=True) \n```\n\n\nThank you all, I've tried my best to cover most part of my solution. Again, I am super happy to win the solo gold, feel free to reach out in case you find difficulty understanding any part of it.",
    "2381165": "> This competition also taught how important it is to keep your updated and trying out the recent researches made in the field.\n\nSo will you finally start reading LLM Papers as well? 😁",
    "2381263": "Congratulations on achieving GM Norm and also getting top spot.\nThanks for sharing the interesting info about your approach. Good illustrations for understanding the techniques applied in your notebook.",
    "2381448": "What an awesome piece of paper @nischaydnk, you published for the community. Thanks for showing your great work to the community.👍",
    "2382234": "Solution worthy of the GM title :) Congrats and thanks for sharing!",
    "2383370": "Hello, Congrats for the 3rd place and the Competition GM!\nI want to know more detail about how you split the train - valid set in the \"Multi Stage Approach\" and \"Multi Stage Pseudo Label Approach\". Like how to you split so that i can guarantee that there will be no data leakage and you can trust that split \n\nThank you",
    "2385862": "Thanks for reminding about split, I will update the solution. I simply used kfold on ds1, so basically in stage 1 , training data is kept ds2 (complete) , validation split is single fold from ds1. In stage 2, 4 folds from ds1 in training keeping the same validation fold.",
    "2386167": "Thank you 😀",
    "2386809": "Congratulations and thank you for the report @nischaydnk !\nCould you elaborate on how you generated masks? Have you used a single mask head with WBF fused boxes?",
    "2387711": "Your work has merit and seems very well supported... I have experience in hardware and in the future I hope to implement a medical application, and I will be inspired by your work... Kind regards and congrats!",
    "2389303": "Thank you!",
    "2389440": "Thank you for the explanatory report. Learned a few steps and will try using those. Thank you and also many Congratulations ! @nischaydnk"
  },
  "source": "meta"
}