{
  "id": 428327,
  "title": "19th Place Solution: Repeated Pseudo Labeling",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/writeups/moritake04-19th-place-solution-repeated-pseudo-lab",
  "author_name": "",
  "post_date": "2023-08-06T10:49:39.480Z",
  "votes": 22,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First, I would like to thank the competition hosts for organizing this competition and also the participants.</p>\n<p>I was disappointed to be kicked out of the gold zone after the shakedown, but the competition was a great learning experience for instance segmentation.</p>\n<h2>Overview</h2>\n<p>My final submission was a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny).</p>\n<p>However, the training process had multiple stages. Figure 1 outlines the training and submission pipeline.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F210e2b0ce449c03a2c77e93ae50c41df%2Fhubmap_train-sumarry.drawio%20(3).png?generation=1690860001136772&amp;alt=media\" alt=\"\"><br>\nFigure 1: Training and Submission Pipeline</p>\n<ul>\n<li>1st Training Phase:<br>\nCascade Mask R-CNN was trained using Dataset 1 and 2. Three different backbones (ConvNext-tiny, ResNext-101, Resnet-50) were used for Cascade Mask R-CNN.<br>\nUsing the three models, we performed pseudo labeling on dataset 3, extracting only labels with a confidence greater than 0.5. We also performed TTA and NMS during pseudo　labeling.</li>\n<li>2nd Training Phase:<br>\nWe used Dataset 1 and 2, and the pseudo labeled dataset 3 created in the 1st phase to train the model whose backbone is convnext. In this phase, the model trained in the 1st phase was fine tuned.</li>\n<li>3rd Training Phase:<br>\nFine tuning the model trained in the 2nd phase using only Dataset 1. This was done to bring the output of the model closer to the annotation quality of dataset 1.<br>\nUsing the model created here, we performed pseudo labeling on dataset 2 and 3, extracting only labels with a confidence greater than or equal to 0.95. We also performed TTA and NMS during pseudo labeling.</li>\n<li>Final Training Phase:<br>\nDataset 1, and pseudo labeled dataset 2 and 3 created in the 3rd phase were used to train the model whose backbone is convnext. Here, the model trained in the 2nd phase was used for fine tuning.</li>\n<li>Submission:<br>\nThe model developed in the Final Training Phase was used to infer the test dataset; it is a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny), but we have performed TTA and NMS.</li>\n</ul>\n<h2>My Approach</h2>\n<p>I proceeded to train the model, focusing on the quality of the annotations in each dataset. The following is a quote from the description of the datasets in this competition.</p>\n<blockquote>\n  <p>The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.</p>\n  <p>All of the test set tiles are from Dataset 1.</p>\n  <p>Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.</p>\n  <p>The training data includes Dataset 2 tiles from the&nbsp;<em>public</em>&nbsp;test WSI, but&nbsp;<em>not</em>&nbsp;from the&nbsp;<em>private</em>&nbsp;test WSI.</p>\n  <p>We also include, as Dataset 3, tiles extracted from an additional nine WSIs. These tiles have not been annotated. You may wish to apply semi- or self-supervised learning techniques on this data to support your predictions.</p>\n</blockquote>\n<p>This means,</p>\n<ul>\n<li>Dataset 1 → accurate annotations and the same annotation quality as the test dataset</li>\n<li>Dataset 2 → not so accurate annotations</li>\n<li>Dataset 3 → no annotations</li>\n</ul>\n<p>Therefore, I decided to trust Dataset 1 and tried to improve the quality of the annotations output by the model to be closer to Dataset 1. I also tried to make good use of the remaining Dataset 2 and 3. Based on this policy, I proceeded with my experiments.</p>\n<ul>\n<li>Model<ul>\n<li>convnext-tiny mainly used.</li>\n<li>resnet, resnext were also used, but finally not used because they did not give good results when used with pseudo label.<ul>\n<li>resnet and resnext were used for the first pseudo labeling.</li></ul></li>\n<li>mmdetection 3.x was used.</li></ul></li>\n<li>Annotations Type<ul>\n<li>Only blood_vessel was used.</li>\n<li>unsure, glomerulus not used at all.</li></ul></li>\n<li>Data Augmentation<ul>\n<li>The images were randomly augmented to multiple scales during training as shown below.<ul>\n<li>scales=[(640, 640), (768, 768), (896, 896), (1024, 1024), (1152, 1152), (1280, 1280), (1408, 1408), (1536, 1536)]</li></ul></li></ul></li>\n<li>Test Time Augmentation (TTA)<ul>\n<li>The ensemble of output results from multiple scales was used.<ul>\n<li>scales=[(1024, 1024), (1280, 1280), (1536, 1536)]</li></ul></li></ul></li>\n<li>Post-Processing<ul>\n<li>Non-Maximum Suppression (NMS) was used during TTA.</li>\n<li>No dilation.</li></ul></li>\n<li>Pseudo Labeling<ul>\n<li>Pseudo labeling was performed in two parts as shown in Figure 1. TTA was performed for three different image sizes, scales=[(1024, 1024), (1280, 1280), (1536, 1536)], and NMS was used for the ensemble.</li></ul></li>\n<li>Cross Validation (CV) Strategy<ul>\n<li>CV was set to be similar to private LB.<ul>\n<li>Source WSI == 1 &amp; dataset == 1 → valid</li>\n<li>Source WSI == 1 &amp; dataset != 1 → It was not used for cv.</li>\n<li>Others → train</li></ul></li>\n<li>At the time of submission, all data was used for training.</li>\n<li>CV was not always correlated with Public LB. However, it became more correlated as the experiment progressed to the latter part of the experiment.</li></ul></li>\n</ul>\n<h2>What Worked</h2>\n<ul>\n<li>Pseudo Labeling</li>\n<li>Trust the quality of the annotations in Dataset 1</li>\n<li>TTA</li>\n<li>NMS</li>\n<li>Scale Up Images</li>\n</ul>\n<h2>What Didn’t Worked</h2>\n<ul>\n<li>Models with large parameter sizes<ul>\n<li>mask2former, convnext-small</li>\n<li>Maybe because of lack of parameter tuning…?</li></ul></li>\n<li>Use unsure and glomerulus for training</li>\n<li>Dilation<ul>\n<li>LB was effective when raw dataset 2 was used for training data</li>\n<li>LB was decreased when using only dataset 1 as training data</li></ul></li>\n<li>Flip Data Augmentation</li>\n<li>Weighted Boxes Fusion (WBF) (my implementation may have been suspect…)</li>\n<li>Smooth inference around edges (my implementation may have been suspect…)<ul>\n<li>Reference: <a href=\"https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238013\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238013</a></li></ul></li>\n<li>Ensemble with other models</li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li>GitHub (My training code. I'm sorry but it's not well maintained) -&gt; <a href=\"https://github.com/moritake04/hubmap-2023\" target=\"_blank\">https://github.com/moritake04/hubmap-2023</a></li>\n<li>Submitted notebook -&gt; <a href=\"https://www.kaggle.com/code/moritake04/private19th-final-sub\" target=\"_blank\">https://www.kaggle.com/code/moritake04/private19th-final-sub</a></li>\n</ul>",
  "messages": [
    {
      "id": "2368160",
      "postDate": "08/01/2023 03:30:17",
      "content": "<p>First, I would like to thank the competition hosts for organizing this competition and also the participants.</p>\n<p>I was disappointed to be kicked out of the gold zone after the shakedown, but the competition was a great learning experience for instance segmentation.</p>\n<h2>Overview</h2>\n<p>My final submission was a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny).</p>\n<p>However, the training process had multiple stages. Figure 1 outlines the training and submission pipeline.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F210e2b0ce449c03a2c77e93ae50c41df%2Fhubmap_train-sumarry.drawio%20(3).png?generation=1690860001136772&amp;alt=media\" alt=\"\"><br>\nFigure 1: Training and Submission Pipeline</p>\n<ul>\n<li>1st Training Phase:<br>\nCascade Mask R-CNN was trained using Dataset 1 and 2. Three different backbones (ConvNext-tiny, ResNext-101, Resnet-50) were used for Cascade Mask R-CNN.<br>\nUsing the three models, we performed pseudo labeling on dataset 3, extracting only labels with a confidence greater than 0.5. We also performed TTA and NMS during pseudo　labeling.</li>\n<li>2nd Training Phase:<br>\nWe used Dataset 1 and 2, and the pseudo labeled dataset 3 created in the 1st phase to train the model whose backbone is convnext. In this phase, the model trained in the 1st phase was fine tuned.</li>\n<li>3rd Training Phase:<br>\nFine tuning the model trained in the 2nd phase using only Dataset 1. This was done to bring the output of the model closer to the annotation quality of dataset 1.<br>\nUsing the model created here, we performed pseudo labeling on dataset 2 and 3, extracting only labels with a confidence greater than or equal to 0.95. We also performed TTA and NMS during pseudo labeling.</li>\n<li>Final Training Phase:<br>\nDataset 1, and pseudo labeled dataset 2 and 3 created in the 3rd phase were used to train the model whose backbone is convnext. Here, the model trained in the 2nd phase was used for fine tuning.</li>\n<li>Submission:<br>\nThe model developed in the Final Training Phase was used to infer the test dataset; it is a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny), but we have performed TTA and NMS.</li>\n</ul>\n<h2>My Approach</h2>\n<p>I proceeded to train the model, focusing on the quality of the annotations in each dataset. The following is a quote from the description of the datasets in this competition.</p>\n<blockquote>\n  <p>The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.</p>\n  <p>All of the test set tiles are from Dataset 1.</p>\n  <p>Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.</p>\n  <p>The training data includes Dataset 2 tiles from the&nbsp;<em>public</em>&nbsp;test WSI, but&nbsp;<em>not</em>&nbsp;from the&nbsp;<em>private</em>&nbsp;test WSI.</p>\n  <p>We also include, as Dataset 3, tiles extracted from an additional nine WSIs. These tiles have not been annotated. You may wish to apply semi- or self-supervised learning techniques on this data to support your predictions.</p>\n</blockquote>\n<p>This means,</p>\n<ul>\n<li>Dataset 1 → accurate annotations and the same annotation quality as the test dataset</li>\n<li>Dataset 2 → not so accurate annotations</li>\n<li>Dataset 3 → no annotations</li>\n</ul>\n<p>Therefore, I decided to trust Dataset 1 and tried to improve the quality of the annotations output by the model to be closer to Dataset 1. I also tried to make good use of the remaining Dataset 2 and 3. Based on this policy, I proceeded with my experiments.</p>\n<ul>\n<li>Model<ul>\n<li>convnext-tiny mainly used.</li>\n<li>resnet, resnext were also used, but finally not used because they did not give good results when used with pseudo label.<ul>\n<li>resnet and resnext were used for the first pseudo labeling.</li></ul></li>\n<li>mmdetection 3.x was used.</li></ul></li>\n<li>Annotations Type<ul>\n<li>Only blood_vessel was used.</li>\n<li>unsure, glomerulus not used at all.</li></ul></li>\n<li>Data Augmentation<ul>\n<li>The images were randomly augmented to multiple scales during training as shown below.<ul>\n<li>scales=[(640, 640), (768, 768), (896, 896), (1024, 1024), (1152, 1152), (1280, 1280), (1408, 1408), (1536, 1536)]</li></ul></li></ul></li>\n<li>Test Time Augmentation (TTA)<ul>\n<li>The ensemble of output results from multiple scales was used.<ul>\n<li>scales=[(1024, 1024), (1280, 1280), (1536, 1536)]</li></ul></li></ul></li>\n<li>Post-Processing<ul>\n<li>Non-Maximum Suppression (NMS) was used during TTA.</li>\n<li>No dilation.</li></ul></li>\n<li>Pseudo Labeling<ul>\n<li>Pseudo labeling was performed in two parts as shown in Figure 1. TTA was performed for three different image sizes, scales=[(1024, 1024), (1280, 1280), (1536, 1536)], and NMS was used for the ensemble.</li></ul></li>\n<li>Cross Validation (CV) Strategy<ul>\n<li>CV was set to be similar to private LB.<ul>\n<li>Source WSI == 1 &amp; dataset == 1 → valid</li>\n<li>Source WSI == 1 &amp; dataset != 1 → It was not used for cv.</li>\n<li>Others → train</li></ul></li>\n<li>At the time of submission, all data was used for training.</li>\n<li>CV was not always correlated with Public LB. However, it became more correlated as the experiment progressed to the latter part of the experiment.</li></ul></li>\n</ul>\n<h2>What Worked</h2>\n<ul>\n<li>Pseudo Labeling</li>\n<li>Trust the quality of the annotations in Dataset 1</li>\n<li>TTA</li>\n<li>NMS</li>\n<li>Scale Up Images</li>\n</ul>\n<h2>What Didn’t Worked</h2>\n<ul>\n<li>Models with large parameter sizes<ul>\n<li>mask2former, convnext-small</li>\n<li>Maybe because of lack of parameter tuning…?</li></ul></li>\n<li>Use unsure and glomerulus for training</li>\n<li>Dilation<ul>\n<li>LB was effective when raw dataset 2 was used for training data</li>\n<li>LB was decreased when using only dataset 1 as training data</li></ul></li>\n<li>Flip Data Augmentation</li>\n<li>Weighted Boxes Fusion (WBF) (my implementation may have been suspect…)</li>\n<li>Smooth inference around edges (my implementation may have been suspect…)<ul>\n<li>Reference: <a href=\"https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238013\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238013</a></li></ul></li>\n<li>Ensemble with other models</li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li>GitHub (My training code. I'm sorry but it's not well maintained) -&gt; <a href=\"https://github.com/moritake04/hubmap-2023\" target=\"_blank\">https://github.com/moritake04/hubmap-2023</a></li>\n<li>Submitted notebook -&gt; <a href=\"https://www.kaggle.com/code/moritake04/private19th-final-sub\" target=\"_blank\">https://www.kaggle.com/code/moritake04/private19th-final-sub</a></li>\n</ul>",
      "rawMarkdown": "First, I would like to thank the competition hosts for organizing this competition and also the participants.\n\nI was disappointed to be kicked out of the gold zone after the shakedown, but the competition was a great learning experience for instance segmentation.\n\n## Overview\n\nMy final submission was a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny).\n\nHowever, the training process had multiple stages. Figure 1 outlines the training and submission pipeline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F210e2b0ce449c03a2c77e93ae50c41df%2Fhubmap_train-sumarry.drawio%20(3).png?generation=1690860001136772&alt=media)\nFigure 1: Training and Submission Pipeline\n\n- 1st Training Phase:\nCascade Mask R-CNN was trained using Dataset 1 and 2. Three different backbones (ConvNext-tiny, ResNext-101, Resnet-50) were used for Cascade Mask R-CNN.\nUsing the three models, we performed pseudo labeling on dataset 3, extracting only labels with a confidence greater than 0.5. We also performed TTA and NMS during pseudo　labeling.\n- 2nd Training Phase:\nWe used Dataset 1 and 2, and the pseudo labeled dataset 3 created in the 1st phase to train the model whose backbone is convnext. In this phase, the model trained in the 1st phase was fine tuned.\n- 3rd Training Phase:\nFine tuning the model trained in the 2nd phase using only Dataset 1. This was done to bring the output of the model closer to the annotation quality of dataset 1.\nUsing the model created here, we performed pseudo labeling on dataset 2 and 3, extracting only labels with a confidence greater than or equal to 0.95. We also performed TTA and NMS during pseudo labeling.\n- Final Training Phase:\nDataset 1, and pseudo labeled dataset 2 and 3 created in the 3rd phase were used to train the model whose backbone is convnext. Here, the model trained in the 2nd phase was used for fine tuning.\n- Submission:\nThe model developed in the Final Training Phase was used to infer the test dataset; it is a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny), but we have performed TTA and NMS.\n\n## My Approach\n\nI proceeded to train the model, focusing on the quality of the annotations in each dataset. The following is a quote from the description of the datasets in this competition.\n\n> The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\n> \n\n> All of the test set tiles are from Dataset 1.\n> \n\n> Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.\n> \n\n> The training data includes Dataset 2 tiles from the *public* test WSI, but *not* from the *private* test WSI.\n> \n\n> We also include, as Dataset 3, tiles extracted from an additional nine WSIs. These tiles have not been annotated. You may wish to apply semi- or self-supervised learning techniques on this data to support your predictions.\n>\n\nThis means,\n\n- Dataset 1 → accurate annotations and the same annotation quality as the test dataset\n- Dataset 2 → not so accurate annotations\n- Dataset 3 → no annotations\n\nTherefore, I decided to trust Dataset 1 and tried to improve the quality of the annotations output by the model to be closer to Dataset 1. I also tried to make good use of the remaining Dataset 2 and 3. Based on this policy, I proceeded with my experiments.\n\n- Model\n    - convnext-tiny mainly used.\n    - resnet, resnext were also used, but finally not used because they did not give good results when used with pseudo label.\n        - resnet and resnext were used for the first pseudo labeling.\n    - mmdetection 3.x was used.\n- Annotations Type\n    - Only blood_vessel was used.\n    - unsure, glomerulus not used at all.\n- Data Augmentation\n    - The images were randomly augmented to multiple scales during training as shown below.\n        - scales=[(640, 640), (768, 768), (896, 896), (1024, 1024), (1152, 1152), (1280, 1280), (1408, 1408), (1536, 1536)]\n- Test Time Augmentation (TTA)\n    - The ensemble of output results from multiple scales was used.\n        - scales=[(1024, 1024), (1280, 1280), (1536, 1536)]\n- Post-Processing\n    - Non-Maximum Suppression (NMS) was used during TTA.\n    - No dilation.\n- Pseudo Labeling\n    - Pseudo labeling was performed in two parts as shown in Figure 1. TTA was performed for three different image sizes, scales=[(1024, 1024), (1280, 1280), (1536, 1536)], and NMS was used for the ensemble.\n- Cross Validation (CV) Strategy\n    - CV was set to be similar to private LB.\n        - Source WSI == 1 & dataset == 1 → valid\n        - Source WSI == 1 & dataset != 1 → It was not used for cv.\n        - Others → train\n    - At the time of submission, all data was used for training.\n    - CV was not always correlated with Public LB. However, it became more correlated as the experiment progressed to the latter part of the experiment.\n\n## What Worked\n\n- Pseudo Labeling\n- Trust the quality of the annotations in Dataset 1\n- TTA\n- NMS\n- Scale Up Images\n\n## What Didn’t Worked\n\n- Models with large parameter sizes\n    - mask2former, convnext-small\n    - Maybe because of lack of parameter tuning…?\n- Use unsure and glomerulus for training\n- Dilation\n    - LB was effective when raw dataset 2 was used for training data\n    - LB was decreased when using only dataset 1 as training data\n- Flip Data Augmentation\n- Weighted Boxes Fusion (WBF) (my implementation may have been suspect…)\n- Smooth inference around edges (my implementation may have been suspect…)\n    - Reference: https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238013\n- Ensemble with other models\n\n## Code\n- GitHub (My training code. I'm sorry but it's not well maintained) -> https://github.com/moritake04/hubmap-2023\n- Submitted notebook -> https://www.kaggle.com/code/moritake04/private19th-final-sub",
      "votes": null
    },
    {
      "id": "2369423",
      "postDate": "08/01/2023 18:13:45",
      "content": "<p>Thanks for sharing how went about tackling this.  Your diagram and explanation makes it easier for a newbie such as myself to follow.</p>",
      "rawMarkdown": "Thanks for sharing how went about tackling this.  Your diagram and explanation makes it easier for a newbie such as myself to follow.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2369423,
      "author_name": "ahmadxy",
      "author_url": "",
      "post_date": "08/01/2023 18:13:45",
      "content": "<p>Thanks for sharing how went about tackling this.  Your diagram and explanation makes it easier for a newbie such as myself to follow.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2368160": "First, I would like to thank the competition hosts for organizing this competition and also the participants.\n\nI was disappointed to be kicked out of the gold zone after the shakedown, but the competition was a great learning experience for instance segmentation.\n\n## Overview\n\nMy final submission was a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny).\n\nHowever, the training process had multiple stages. Figure 1 outlines the training and submission pipeline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F210e2b0ce449c03a2c77e93ae50c41df%2Fhubmap_train-sumarry.drawio%20(3).png?generation=1690860001136772&alt=media)\nFigure 1: Training and Submission Pipeline\n\n- 1st Training Phase:\nCascade Mask R-CNN was trained using Dataset 1 and 2. Three different backbones (ConvNext-tiny, ResNext-101, Resnet-50) were used for Cascade Mask R-CNN.\nUsing the three models, we performed pseudo labeling on dataset 3, extracting only labels with a confidence greater than 0.5. We also performed TTA and NMS during pseudo　labeling.\n- 2nd Training Phase:\nWe used Dataset 1 and 2, and the pseudo labeled dataset 3 created in the 1st phase to train the model whose backbone is convnext. In this phase, the model trained in the 1st phase was fine tuned.\n- 3rd Training Phase:\nFine tuning the model trained in the 2nd phase using only Dataset 1. This was done to bring the output of the model closer to the annotation quality of dataset 1.\nUsing the model created here, we performed pseudo labeling on dataset 2 and 3, extracting only labels with a confidence greater than or equal to 0.95. We also performed TTA and NMS during pseudo labeling.\n- Final Training Phase:\nDataset 1, and pseudo labeled dataset 2 and 3 created in the 3rd phase were used to train the model whose backbone is convnext. Here, the model trained in the 2nd phase was used for fine tuning.\n- Submission:\nThe model developed in the Final Training Phase was used to infer the test dataset; it is a single model of Cascade Mask R-CNN (Backbone: ConvNeXt-tiny), but we have performed TTA and NMS.\n\n## My Approach\n\nI proceeded to train the model, focusing on the quality of the annotations in each dataset. The following is a quote from the description of the datasets in this competition.\n\n> The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\n> \n\n> All of the test set tiles are from Dataset 1.\n> \n\n> Two of the WSIs make up the training set, two WSIs make up the public test set, and one WSI makes up the private test set.\n> \n\n> The training data includes Dataset 2 tiles from the *public* test WSI, but *not* from the *private* test WSI.\n> \n\n> We also include, as Dataset 3, tiles extracted from an additional nine WSIs. These tiles have not been annotated. You may wish to apply semi- or self-supervised learning techniques on this data to support your predictions.\n>\n\nThis means,\n\n- Dataset 1 → accurate annotations and the same annotation quality as the test dataset\n- Dataset 2 → not so accurate annotations\n- Dataset 3 → no annotations\n\nTherefore, I decided to trust Dataset 1 and tried to improve the quality of the annotations output by the model to be closer to Dataset 1. I also tried to make good use of the remaining Dataset 2 and 3. Based on this policy, I proceeded with my experiments.\n\n- Model\n    - convnext-tiny mainly used.\n    - resnet, resnext were also used, but finally not used because they did not give good results when used with pseudo label.\n        - resnet and resnext were used for the first pseudo labeling.\n    - mmdetection 3.x was used.\n- Annotations Type\n    - Only blood_vessel was used.\n    - unsure, glomerulus not used at all.\n- Data Augmentation\n    - The images were randomly augmented to multiple scales during training as shown below.\n        - scales=[(640, 640), (768, 768), (896, 896), (1024, 1024), (1152, 1152), (1280, 1280), (1408, 1408), (1536, 1536)]\n- Test Time Augmentation (TTA)\n    - The ensemble of output results from multiple scales was used.\n        - scales=[(1024, 1024), (1280, 1280), (1536, 1536)]\n- Post-Processing\n    - Non-Maximum Suppression (NMS) was used during TTA.\n    - No dilation.\n- Pseudo Labeling\n    - Pseudo labeling was performed in two parts as shown in Figure 1. TTA was performed for three different image sizes, scales=[(1024, 1024), (1280, 1280), (1536, 1536)], and NMS was used for the ensemble.\n- Cross Validation (CV) Strategy\n    - CV was set to be similar to private LB.\n        - Source WSI == 1 & dataset == 1 → valid\n        - Source WSI == 1 & dataset != 1 → It was not used for cv.\n        - Others → train\n    - At the time of submission, all data was used for training.\n    - CV was not always correlated with Public LB. However, it became more correlated as the experiment progressed to the latter part of the experiment.\n\n## What Worked\n\n- Pseudo Labeling\n- Trust the quality of the annotations in Dataset 1\n- TTA\n- NMS\n- Scale Up Images\n\n## What Didn’t Worked\n\n- Models with large parameter sizes\n    - mask2former, convnext-small\n    - Maybe because of lack of parameter tuning…?\n- Use unsure and glomerulus for training\n- Dilation\n    - LB was effective when raw dataset 2 was used for training data\n    - LB was decreased when using only dataset 1 as training data\n- Flip Data Augmentation\n- Weighted Boxes Fusion (WBF) (my implementation may have been suspect…)\n- Smooth inference around edges (my implementation may have been suspect…)\n    - Reference: https://www.kaggle.com/competitions/hubmap-kidney-segmentation/discussion/238013\n- Ensemble with other models\n\n## Code\n- GitHub (My training code. I'm sorry but it's not well maintained) -> https://github.com/moritake04/hubmap-2023\n- Submitted notebook -> https://www.kaggle.com/code/moritake04/private19th-final-sub",
    "2369423": "Thanks for sharing how went about tackling this.  Your diagram and explanation makes it easier for a newbie such as myself to follow."
  },
  "source": "meta"
}