{
  "id": 429240,
  "title": "2nd place solution",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/429240",
  "author_name": "qdv206",
  "post_date": "2023-08-04T15:33:49.694000",
  "votes": 28,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Many thanks to Kaggle and HuBMAP who organized another very interesting competition. This is a joint writeup with my long-time teammate <a href=\"https://www.kaggle.com/phamvanlinh143\" target=\"_blank\">@phamvanlinh143</a> , who did most of the hard work in this one.</p>\n<p><strong>Summary</strong></p>\n<p>This competition is an especially tricky one with several problems that needed to be addressed:</p>\n<ul>\n<li>Public and private test set are structured differently, with private set coming from an unseen WSI. Public test set size is also quite small. With these two reasons combined we have a completely unreliable LB.</li>\n<li>The strange effect of dilation on public LB result.</li>\n</ul>\n<p>Fortunately, we came up with a strategy that we believed can deal with each of these issues accordingly:</p>\n<ul>\n<li>Build a reliable CV and trust it.</li>\n<li>Train a large and diverse ensemble for stability.</li>\n<li>Submit the ensemble with dilate and without dilate for our two submissions.</li>\n</ul>\n<p><strong>Cross-validation</strong></p>\n<p>To build a trustworthy CV, it must follow as closely as possible to how private test set is created. That means:</p>\n<ul>\n<li>Validation must be done on dataset 1 labels.</li>\n<li>Train and validation set must not contain the same WSI.</li>\n</ul>\n<p>A simple solution for this is to use Dataset1-WSI1 as one validation fold and Dataset1-WSI2 as the other, but this would leave out a lot of training samples, especially if we want to train using only Dataset 1. </p>\n<p>What we did is using metadata to split the Dataset1 tiles from each WSI into left and right side, resulting in 4 folds:</p>\n<ul>\n<li>Dataset1 – WSI1 - Left</li>\n<li>Dataset1 – WSI1 - Right</li>\n<li>Dataset1 – WSI2 - Left</li>\n<li>Dataset1 – WSI2 – Right</li>\n</ul>\n<p>Next, for each tile in training data we use <code>staintools</code> to generate 9 according tiles transformed into the style of the 9 additional WSIs in Dataset 3. In training, when the sampled tile come from the same WSI of the validation set, one of the 9 generated tiles are sampled instead. While this method doesn’t completely remove the characteristics of the original WSI, we felt that it is a good enough compromise to what we wanted to achieve.</p>\n<p>An example of an original tile and its 9 variations:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F0e621e649f5c50cba0d6c9d51a15889a%2Ftest2.jpg?generation=1691163031076471&amp;alt=media\" alt=\"“”\"></p>\n<p>Splitting the dataset in this manner also allowed us to train models using tile concatenation with minimal leak, as we only had to handle the middle tile column that separated left and right side of the WSI.</p>\n<p><strong>Model training</strong></p>\n<p><strong>Input data:</strong>  we used one of these two data types:</p>\n<ul>\n<li>original tiles</li>\n<li>padded tiles similar to what <a href=\"https://www.kaggle.com/hengck\" target=\"_blank\">@hengck</a> <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2316842\" target=\"_blank\">proposed</a>. We padded 128 pixels around the original tile using available neighboring tiles. Instance labels are modified accordingly. At inference time we predicted on padded region and then center crop.</li>\n</ul>\n<p>An example of original tile (left) and padded tile (right):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F07696cca441ed581263843a0a919752d%2Ftest.jpg?generation=1691162170897200&amp;alt=media\" alt=\"“”\"></p>\n<p><strong>Augmentation:</strong> we applied strong augmentation:</p>\n<ul>\n<li>stain augmentation: p=1.0 for tiles from same WSI as validation set, p=0.5 otherwise</li>\n<li>RandomRotate90, RandomFlip, ElasticTransform, ShiftScaleRotate, RandomBrightnessContrast, HueSaturationValue, ImageCompression, GaussNoise, GaussianBlur…</li>\n<li>AutoAugment similar to DETR <a href=\"https://github.com/open-mmlab/mmdetection/blob/master/configs/detr/detr_r50_8x2_150e_coco.py\" target=\"_blank\">training config</a>.</li>\n</ul>\n<p><strong>Training details:</strong> we trained the models in 2 stages, using only blood vessel class. </p>\n<ul>\n<li>Stage 1: train for 30 epochs using 3 folds dataset 1 + all dataset 2</li>\n<li>Stage 2: finetune for 15 epochs using only 3 folds dataset 1</li>\n</ul>\n<p>Then we picked 5 checkpoints from finetune stage and do SWA to provide the final model.</p>\n<p><strong>Models:</strong> Cascade Mask-RCNN models with <code>swin-t</code>, <code>coat-small</code>, <code>convnext-t</code> and <code>convnext-s</code> backbones. We used mmdet 2.x to train our models.</p>\n<p>We experimented with different combination of backbones and input data types (original or padded). We selected the models with best CV and trained a version using full data. The final ensemble consists of both fold models and full data models.</p>\n<p><strong>Postprocessing:</strong> The masks that met following criteria are filtered:</p>\n<ul>\n<li>Glomeruli filter: remove masks with more than 60% area inside glomeruli regions.</li>\n<li>Confidence filter: remove masks with confidence &lt;0.01.</li>\n<li>Ensemble filter by pixel: count for each pixels the number of blood vessel prediction by the ensemble (how many models predicted positive for that pixel). A threshold is then calculated from the resulting list of pixel counts with quantile q=0.05 (ignore zero value). Masks of which all pixels had count smaller than this threshold are removed.</li>\n<li>Ensemble filter by instance: perform nms with iou_thresh=0.65. For each selected mask we save the number of masks that satisfied iou_thresh with it. Similar to pixel counts filter, we remove masks with small number of overlapping masks using quantile q=0.075.</li>\n<li>Small mask filter: remove masks with fewer than 64 pixels.</li>\n</ul>\n<p><strong>To dilate or not to dilate</strong></p>\n<p>As others have reported, we observed significant change of score on Public LB with and without dilation. As adding dilation didn’t work at all on our cross validation, we knew we cannot trust it. However there exists also the probability that private set would present the same label patterns as public set, as the host have given confirmation that they were verified using the same procedure. Fortunately, we had two submissions, so this is where we decided to put it to good use.</p>\n<p>Our final ensemble scored: </p>\n<ul>\n<li>With dilation: 0.551 public, 0.526 private.</li>\n<li>Without dilation: 0.465 public, 0.588 private.</li>\n</ul>\n<p>Our late submissions showed that submitting any single backbone in our ensemble without dilation would have landed us in the private gold zone anyway while scoring 0.4x on public LB. Even though we chose correctly by trusting our CV, this public LB behavior is still a mystery to us. Hopefully we can shed some light into it by reading the solutions of other top teams.</p>\n<p>Thank you very much for reading and let us know if you have any questions.</p>\n<p>Edit: </p>\n<ul>\n<li>training code: <a href=\"https://github.com/phamvanlinh143/HubMap_2023_2nd_Place_Solution\" target=\"_blank\">https://github.com/phamvanlinh143/HubMap_2023_2nd_Place_Solution</a></li>\n<li>inference notebook: <a href=\"https://www.kaggle.com/code/phamvanlinh143/hubmap-2nd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/phamvanlinh143/hubmap-2nd-place-inference</a></li>\n</ul>",
  "messages": [
    {
      "id": 2374025,
      "postDate": "2023-08-04T15:33:49.693Z",
      "content": "<p>Many thanks to Kaggle and HuBMAP who organized another very interesting competition. This is a joint writeup with my long-time teammate <a href=\"https://www.kaggle.com/phamvanlinh143\" target=\"_blank\">@phamvanlinh143</a> , who did most of the hard work in this one.</p>\n<p><strong>Summary</strong></p>\n<p>This competition is an especially tricky one with several problems that needed to be addressed:</p>\n<ul>\n<li>Public and private test set are structured differently, with private set coming from an unseen WSI. Public test set size is also quite small. With these two reasons combined we have a completely unreliable LB.</li>\n<li>The strange effect of dilation on public LB result.</li>\n</ul>\n<p>Fortunately, we came up with a strategy that we believed can deal with each of these issues accordingly:</p>\n<ul>\n<li>Build a reliable CV and trust it.</li>\n<li>Train a large and diverse ensemble for stability.</li>\n<li>Submit the ensemble with dilate and without dilate for our two submissions.</li>\n</ul>\n<p><strong>Cross-validation</strong></p>\n<p>To build a trustworthy CV, it must follow as closely as possible to how private test set is created. That means:</p>\n<ul>\n<li>Validation must be done on dataset 1 labels.</li>\n<li>Train and validation set must not contain the same WSI.</li>\n</ul>\n<p>A simple solution for this is to use Dataset1-WSI1 as one validation fold and Dataset1-WSI2 as the other, but this would leave out a lot of training samples, especially if we want to train using only Dataset 1. </p>\n<p>What we did is using metadata to split the Dataset1 tiles from each WSI into left and right side, resulting in 4 folds:</p>\n<ul>\n<li>Dataset1 – WSI1 - Left</li>\n<li>Dataset1 – WSI1 - Right</li>\n<li>Dataset1 – WSI2 - Left</li>\n<li>Dataset1 – WSI2 – Right</li>\n</ul>\n<p>Next, for each tile in training data we use <code>staintools</code> to generate 9 according tiles transformed into the style of the 9 additional WSIs in Dataset 3. In training, when the sampled tile come from the same WSI of the validation set, one of the 9 generated tiles are sampled instead. While this method doesn’t completely remove the characteristics of the original WSI, we felt that it is a good enough compromise to what we wanted to achieve.</p>\n<p>An example of an original tile and its 9 variations:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F0e621e649f5c50cba0d6c9d51a15889a%2Ftest2.jpg?generation=1691163031076471&amp;alt=media\" alt=\"“”\"></p>\n<p>Splitting the dataset in this manner also allowed us to train models using tile concatenation with minimal leak, as we only had to handle the middle tile column that separated left and right side of the WSI.</p>\n<p><strong>Model training</strong></p>\n<p><strong>Input data:</strong>  we used one of these two data types:</p>\n<ul>\n<li>original tiles</li>\n<li>padded tiles similar to what <a href=\"https://www.kaggle.com/hengck\" target=\"_blank\">@hengck</a> <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2316842\" target=\"_blank\">proposed</a>. We padded 128 pixels around the original tile using available neighboring tiles. Instance labels are modified accordingly. At inference time we predicted on padded region and then center crop.</li>\n</ul>\n<p>An example of original tile (left) and padded tile (right):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F07696cca441ed581263843a0a919752d%2Ftest.jpg?generation=1691162170897200&amp;alt=media\" alt=\"“”\"></p>\n<p><strong>Augmentation:</strong> we applied strong augmentation:</p>\n<ul>\n<li>stain augmentation: p=1.0 for tiles from same WSI as validation set, p=0.5 otherwise</li>\n<li>RandomRotate90, RandomFlip, ElasticTransform, ShiftScaleRotate, RandomBrightnessContrast, HueSaturationValue, ImageCompression, GaussNoise, GaussianBlur…</li>\n<li>AutoAugment similar to DETR <a href=\"https://github.com/open-mmlab/mmdetection/blob/master/configs/detr/detr_r50_8x2_150e_coco.py\" target=\"_blank\">training config</a>.</li>\n</ul>\n<p><strong>Training details:</strong> we trained the models in 2 stages, using only blood vessel class. </p>\n<ul>\n<li>Stage 1: train for 30 epochs using 3 folds dataset 1 + all dataset 2</li>\n<li>Stage 2: finetune for 15 epochs using only 3 folds dataset 1</li>\n</ul>\n<p>Then we picked 5 checkpoints from finetune stage and do SWA to provide the final model.</p>\n<p><strong>Models:</strong> Cascade Mask-RCNN models with <code>swin-t</code>, <code>coat-small</code>, <code>convnext-t</code> and <code>convnext-s</code> backbones. We used mmdet 2.x to train our models.</p>\n<p>We experimented with different combination of backbones and input data types (original or padded). We selected the models with best CV and trained a version using full data. The final ensemble consists of both fold models and full data models.</p>\n<p><strong>Postprocessing:</strong> The masks that met following criteria are filtered:</p>\n<ul>\n<li>Glomeruli filter: remove masks with more than 60% area inside glomeruli regions.</li>\n<li>Confidence filter: remove masks with confidence &lt;0.01.</li>\n<li>Ensemble filter by pixel: count for each pixels the number of blood vessel prediction by the ensemble (how many models predicted positive for that pixel). A threshold is then calculated from the resulting list of pixel counts with quantile q=0.05 (ignore zero value). Masks of which all pixels had count smaller than this threshold are removed.</li>\n<li>Ensemble filter by instance: perform nms with iou_thresh=0.65. For each selected mask we save the number of masks that satisfied iou_thresh with it. Similar to pixel counts filter, we remove masks with small number of overlapping masks using quantile q=0.075.</li>\n<li>Small mask filter: remove masks with fewer than 64 pixels.</li>\n</ul>\n<p><strong>To dilate or not to dilate</strong></p>\n<p>As others have reported, we observed significant change of score on Public LB with and without dilation. As adding dilation didn’t work at all on our cross validation, we knew we cannot trust it. However there exists also the probability that private set would present the same label patterns as public set, as the host have given confirmation that they were verified using the same procedure. Fortunately, we had two submissions, so this is where we decided to put it to good use.</p>\n<p>Our final ensemble scored: </p>\n<ul>\n<li>With dilation: 0.551 public, 0.526 private.</li>\n<li>Without dilation: 0.465 public, 0.588 private.</li>\n</ul>\n<p>Our late submissions showed that submitting any single backbone in our ensemble without dilation would have landed us in the private gold zone anyway while scoring 0.4x on public LB. Even though we chose correctly by trusting our CV, this public LB behavior is still a mystery to us. Hopefully we can shed some light into it by reading the solutions of other top teams.</p>\n<p>Thank you very much for reading and let us know if you have any questions.</p>\n<p>Edit: </p>\n<ul>\n<li>training code: <a href=\"https://github.com/phamvanlinh143/HubMap_2023_2nd_Place_Solution\" target=\"_blank\">https://github.com/phamvanlinh143/HubMap_2023_2nd_Place_Solution</a></li>\n<li>inference notebook: <a href=\"https://www.kaggle.com/code/phamvanlinh143/hubmap-2nd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/phamvanlinh143/hubmap-2nd-place-inference</a></li>\n</ul>",
      "rawMarkdown": "Many thanks to Kaggle and HuBMAP who organized another very interesting competition. This is a joint writeup with my long-time teammate @phamvanlinh143 , who did most of the hard work in this one.\n\n**Summary**\n\nThis competition is an especially tricky one with several problems that needed to be addressed:\n-\tPublic and private test set are structured differently, with private set coming from an unseen WSI. Public test set size is also quite small. With these two reasons combined we have a completely unreliable LB.\n-\tThe strange effect of dilation on public LB result.\n\nFortunately, we came up with a strategy that we believed can deal with each of these issues accordingly:\n-\tBuild a reliable CV and trust it.\n-\tTrain a large and diverse ensemble for stability.\n-\tSubmit the ensemble with dilate and without dilate for our two submissions.\n\n**Cross-validation**\n\nTo build a trustworthy CV, it must follow as closely as possible to how private test set is created. That means:\n-\tValidation must be done on dataset 1 labels.\n-\tTrain and validation set must not contain the same WSI.\n\nA simple solution for this is to use Dataset1-WSI1 as one validation fold and Dataset1-WSI2 as the other, but this would leave out a lot of training samples, especially if we want to train using only Dataset 1. \n\nWhat we did is using metadata to split the Dataset1 tiles from each WSI into left and right side, resulting in 4 folds:\n-\tDataset1 – WSI1 - Left\n-\tDataset1 – WSI1 - Right\n-\tDataset1 – WSI2 - Left\n-\tDataset1 – WSI2 – Right\n\nNext, for each tile in training data we use `staintools` to generate 9 according tiles transformed into the style of the 9 additional WSIs in Dataset 3. In training, when the sampled tile come from the same WSI of the validation set, one of the 9 generated tiles are sampled instead. While this method doesn’t completely remove the characteristics of the original WSI, we felt that it is a good enough compromise to what we wanted to achieve.\n\nAn example of an original tile and its 9 variations:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F0e621e649f5c50cba0d6c9d51a15889a%2Ftest2.jpg?generation=1691163031076471&alt=media\" alt= “” width=\"640\" height=\"256\">\n\nSplitting the dataset in this manner also allowed us to train models using tile concatenation with minimal leak, as we only had to handle the middle tile column that separated left and right side of the WSI.\n\n**Model training**\n\n**Input data:**  we used one of these two data types:\n-\toriginal tiles\n-\tpadded tiles similar to what @hengck [proposed](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2316842). We padded 128 pixels around the original tile using available neighboring tiles. Instance labels are modified accordingly. At inference time we predicted on padded region and then center crop.\n\nAn example of original tile (left) and padded tile (right):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F07696cca441ed581263843a0a919752d%2Ftest.jpg?generation=1691162170897200&alt=media\" alt= “” width=\"544\" height=\"256\">\n\n**Augmentation:** we applied strong augmentation:\n-\tstain augmentation: p=1.0 for tiles from same WSI as validation set, p=0.5 otherwise\n-\tRandomRotate90, RandomFlip, ElasticTransform, ShiftScaleRotate, RandomBrightnessContrast, HueSaturationValue, ImageCompression, GaussNoise, GaussianBlur…\n-\tAutoAugment similar to DETR [training config](https://github.com/open-mmlab/mmdetection/blob/master/configs/detr/detr_r50_8x2_150e_coco.py).\n\n**Training details:** we trained the models in 2 stages, using only blood vessel class. \n-\tStage 1: train for 30 epochs using 3 folds dataset 1 + all dataset 2\n-\tStage 2: finetune for 15 epochs using only 3 folds dataset 1\n\nThen we picked 5 checkpoints from finetune stage and do SWA to provide the final model.\n\n**Models:** Cascade Mask-RCNN models with `swin-t`, `coat-small`, `convnext-t` and `convnext-s` backbones. We used mmdet 2.x to train our models.\n\nWe experimented with different combination of backbones and input data types (original or padded). We selected the models with best CV and trained a version using full data. The final ensemble consists of both fold models and full data models.\n\n**Postprocessing:** The masks that met following criteria are filtered:\n-\tGlomeruli filter: remove masks with more than 60% area inside glomeruli regions.\n-\tConfidence filter: remove masks with confidence <0.01.\n-\tEnsemble filter by pixel: count for each pixels the number of blood vessel prediction by the ensemble (how many models predicted positive for that pixel). A threshold is then calculated from the resulting list of pixel counts with quantile q=0.05 (ignore zero value). Masks of which all pixels had count smaller than this threshold are removed.\n-\tEnsemble filter by instance: perform nms with iou_thresh=0.65. For each selected mask we save the number of masks that satisfied iou_thresh with it. Similar to pixel counts filter, we remove masks with small number of overlapping masks using quantile q=0.075.\n-\tSmall mask filter: remove masks with fewer than 64 pixels.\n\n**To dilate or not to dilate**\n\nAs others have reported, we observed significant change of score on Public LB with and without dilation. As adding dilation didn’t work at all on our cross validation, we knew we cannot trust it. However there exists also the probability that private set would present the same label patterns as public set, as the host have given confirmation that they were verified using the same procedure. Fortunately, we had two submissions, so this is where we decided to put it to good use.\n\nOur final ensemble scored: \n-\tWith dilation: 0.551 public, 0.526 private.\n-\tWithout dilation: 0.465 public, 0.588 private.\n\nOur late submissions showed that submitting any single backbone in our ensemble without dilation would have landed us in the private gold zone anyway while scoring 0.4x on public LB. Even though we chose correctly by trusting our CV, this public LB behavior is still a mystery to us. Hopefully we can shed some light into it by reading the solutions of other top teams.\n\nThank you very much for reading and let us know if you have any questions.\n\nEdit: \n- training code: https://github.com/phamvanlinh143/HubMap_2023_2nd_Place_Solution\n- inference notebook: https://www.kaggle.com/code/phamvanlinh143/hubmap-2nd-place-inference\n\n\n",
      "votes": 28
    },
    {
      "id": 2376090,
      "postDate": "2023-08-06T07:00:20.240Z",
      "content": "<p><a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> Thanks for the excellent solution sharing.</p>",
      "rawMarkdown": "@qdv206 Thanks for the excellent solution sharing.",
      "votes": 2
    },
    {
      "id": 2375371,
      "postDate": "2023-08-05T15:23:11.560Z",
      "content": "<p>Great achievement and really thanks for sharing your solution and insights!</p>",
      "rawMarkdown": "Great achievement and really thanks for sharing your solution and insights!",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2376090,
      "author_name": "An Dang",
      "author_url": "",
      "post_date": "2023-08-06T07:00:20.240000",
      "content": "<p><a href=\"https://www.kaggle.com/qdv206\" target=\"_blank\">@qdv206</a> Thanks for the excellent solution sharing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2375371,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-05T15:23:11.560000",
      "content": "<p>Great achievement and really thanks for sharing your solution and insights!</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2374025": "Many thanks to Kaggle and HuBMAP who organized another very interesting competition. This is a joint writeup with my long-time teammate @phamvanlinh143 , who did most of the hard work in this one.\n\n**Summary**\n\nThis competition is an especially tricky one with several problems that needed to be addressed:\n-\tPublic and private test set are structured differently, with private set coming from an unseen WSI. Public test set size is also quite small. With these two reasons combined we have a completely unreliable LB.\n-\tThe strange effect of dilation on public LB result.\n\nFortunately, we came up with a strategy that we believed can deal with each of these issues accordingly:\n-\tBuild a reliable CV and trust it.\n-\tTrain a large and diverse ensemble for stability.\n-\tSubmit the ensemble with dilate and without dilate for our two submissions.\n\n**Cross-validation**\n\nTo build a trustworthy CV, it must follow as closely as possible to how private test set is created. That means:\n-\tValidation must be done on dataset 1 labels.\n-\tTrain and validation set must not contain the same WSI.\n\nA simple solution for this is to use Dataset1-WSI1 as one validation fold and Dataset1-WSI2 as the other, but this would leave out a lot of training samples, especially if we want to train using only Dataset 1. \n\nWhat we did is using metadata to split the Dataset1 tiles from each WSI into left and right side, resulting in 4 folds:\n-\tDataset1 – WSI1 - Left\n-\tDataset1 – WSI1 - Right\n-\tDataset1 – WSI2 - Left\n-\tDataset1 – WSI2 – Right\n\nNext, for each tile in training data we use `staintools` to generate 9 according tiles transformed into the style of the 9 additional WSIs in Dataset 3. In training, when the sampled tile come from the same WSI of the validation set, one of the 9 generated tiles are sampled instead. While this method doesn’t completely remove the characteristics of the original WSI, we felt that it is a good enough compromise to what we wanted to achieve.\n\nAn example of an original tile and its 9 variations:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F0e621e649f5c50cba0d6c9d51a15889a%2Ftest2.jpg?generation=1691163031076471&alt=media\" alt= “” width=\"640\" height=\"256\">\n\nSplitting the dataset in this manner also allowed us to train models using tile concatenation with minimal leak, as we only had to handle the middle tile column that separated left and right side of the WSI.\n\n**Model training**\n\n**Input data:**  we used one of these two data types:\n-\toriginal tiles\n-\tpadded tiles similar to what @hengck [proposed](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2316842). We padded 128 pixels around the original tile using available neighboring tiles. Instance labels are modified accordingly. At inference time we predicted on padded region and then center crop.\n\nAn example of original tile (left) and padded tile (right):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9250575%2F07696cca441ed581263843a0a919752d%2Ftest.jpg?generation=1691162170897200&alt=media\" alt= “” width=\"544\" height=\"256\">\n\n**Augmentation:** we applied strong augmentation:\n-\tstain augmentation: p=1.0 for tiles from same WSI as validation set, p=0.5 otherwise\n-\tRandomRotate90, RandomFlip, ElasticTransform, ShiftScaleRotate, RandomBrightnessContrast, HueSaturationValue, ImageCompression, GaussNoise, GaussianBlur…\n-\tAutoAugment similar to DETR [training config](https://github.com/open-mmlab/mmdetection/blob/master/configs/detr/detr_r50_8x2_150e_coco.py).\n\n**Training details:** we trained the models in 2 stages, using only blood vessel class. \n-\tStage 1: train for 30 epochs using 3 folds dataset 1 + all dataset 2\n-\tStage 2: finetune for 15 epochs using only 3 folds dataset 1\n\nThen we picked 5 checkpoints from finetune stage and do SWA to provide the final model.\n\n**Models:** Cascade Mask-RCNN models with `swin-t`, `coat-small`, `convnext-t` and `convnext-s` backbones. We used mmdet 2.x to train our models.\n\nWe experimented with different combination of backbones and input data types (original or padded). We selected the models with best CV and trained a version using full data. The final ensemble consists of both fold models and full data models.\n\n**Postprocessing:** The masks that met following criteria are filtered:\n-\tGlomeruli filter: remove masks with more than 60% area inside glomeruli regions.\n-\tConfidence filter: remove masks with confidence <0.01.\n-\tEnsemble filter by pixel: count for each pixels the number of blood vessel prediction by the ensemble (how many models predicted positive for that pixel). A threshold is then calculated from the resulting list of pixel counts with quantile q=0.05 (ignore zero value). Masks of which all pixels had count smaller than this threshold are removed.\n-\tEnsemble filter by instance: perform nms with iou_thresh=0.65. For each selected mask we save the number of masks that satisfied iou_thresh with it. Similar to pixel counts filter, we remove masks with small number of overlapping masks using quantile q=0.075.\n-\tSmall mask filter: remove masks with fewer than 64 pixels.\n\n**To dilate or not to dilate**\n\nAs others have reported, we observed significant change of score on Public LB with and without dilation. As adding dilation didn’t work at all on our cross validation, we knew we cannot trust it. However there exists also the probability that private set would present the same label patterns as public set, as the host have given confirmation that they were verified using the same procedure. Fortunately, we had two submissions, so this is where we decided to put it to good use.\n\nOur final ensemble scored: \n-\tWith dilation: 0.551 public, 0.526 private.\n-\tWithout dilation: 0.465 public, 0.588 private.\n\nOur late submissions showed that submitting any single backbone in our ensemble without dilation would have landed us in the private gold zone anyway while scoring 0.4x on public LB. Even though we chose correctly by trusting our CV, this public LB behavior is still a mystery to us. Hopefully we can shed some light into it by reading the solutions of other top teams.\n\nThank you very much for reading and let us know if you have any questions.\n\nEdit: \n- training code: https://github.com/phamvanlinh143/HubMap_2023_2nd_Place_Solution\n- inference notebook: https://www.kaggle.com/code/phamvanlinh143/hubmap-2nd-place-inference\n\n\n",
    "2376090": "@qdv206 Thanks for the excellent solution sharing.",
    "2375371": "Great achievement and really thanks for sharing your solution and insights!"
  }
}