{
  "id": 428994,
  "title": "4th Place Solution [SDSRV.AI] GoN",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/writeups/sdsrv-ai-gon-4th-place-solution-sdsrv-ai-gon",
  "author_name": "",
  "post_date": "2023-08-03T17:22:50.043Z",
  "votes": 29,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to the organizers for this compelling competition. We present the 4th place solution from the GoN team at SDSRV.AI. It emphasizes the importance of dataset, validation, ensemble and post processing method. Special thanks to <a href=\"https://www.kaggle.com/ducmanvo\" target=\"_blank\">@ducmanvo</a>, an amazing teammate.</p>\n<h1>Summary</h1>\n<p>Our solution employs a combination of 4 instance segmentation models comprising CascadeRCNN (ResNeXt, Regnet), MaskRCNN (swint), HybridTaskCascade (Re2Net), and 1 object detection model, Yolov6m. Each model is trained with a dataset consisting of 1549 images (ds1+ds2), with 84 images from dataset1 utilized for validation. During the inference process, the Weighted Boxes Fusion (WBF) technique is employed to ensemble RPN boxes and ROI boxes, while the masks generated by the ensemble of models utilize the mean average approach. For post-processing, we filter out small instances and refine instance score using mask scores.</p>\n<h1>Cross-Validation and Preprocessing</h1>\n<p><strong>Train-test split</strong><br>\n    - Training: 1549 images (1211 images from ds2 + 338 images from ds1)<br>\n     - Validation: 84 images from ds1<br>\n     - During the initial stages of approaching the problem (in the final month of the competition), we initially utilized K-fold cross-validation. However, upon recognizing significant differences in distribution, label assignment methodologies between dataset1 and dataset2, as well as variations in label assignments among Whole Slide Images (WSI), we decided to split the data into two train-validation sets. The validation data was exclusively taken from dataset1 (20% of dataset1). Another reason for this data split was the scarcity of dataset1 images in the training set, which was a concern given that the private test data was exclusively sourced from dataset1 (with only 422 images available).<br>\n      - To determine the dissimilarities in the label distribution between dataset1 and dataset2, as well as variations among different Whole Slide Images (WSI), we trained a base model, MaskRCNN R50, on the training data from dataset1 (fold1) and validated it on dataset1, dataset2 (fold1). Our observations revealed that the model's performance scored considerably higher on dataset1 compared to dataset2. Similar procedures were applied when comparing performance across various Whole Slide Images (WSI).<br>\n<strong>Preprocessing</strong>: remove duplicate annotations<br>\n     - Approximately 4.5% of training data was duplicated. We directly removed these duplicate labels to ensure data integrity and improve model performance.</p>\n<h1>Training &amp; valid with blood_vessel &amp; unsure only</h1>\n<p>During training and validation, we specifically focused on the \"blood_vessel\" and \"unsure\" classes. This decision was influenced by the significant size difference between glomerulus and the other two classes, with glomerulus being much larger. Surprisingly, when training with all three classes, the model's performance on glomerulus remained exceptionally high. Additionally, the competition organizers informed us that during the testing phase, we could utilize glomerulus labels to eliminate false positives. Consequently, we concluded that training with additional glomerulus data was unnecessary. After removing glomerulus labels from the training dataset, we observed an improvement in the performance score for the \"blood_vessel\" class. Therefore, we made the decision to exclude glomerulus labels from all subsequent experiments.</p>\n<h1>Models</h1>\n<p><strong>2 stages model</strong>: MaskRCNN, CascadeRCNN, HybridTaskCascade</p>\n<ul>\n<li>When training with the base model (MaskRCNN-R50), we noticed that the model struggled to converge using the default mmdet configurations. Consequently, in subsequent experiments, we reduced the utilization of augmentation methods, retaining only multiscale training. Additionally, we adopted a larger backbone (ResNeXt-101) and replaced SGD with AdamW optimizer. To further enhance model diversity during ensemble, we incorporated CascadeRCNN and HybridTaskCascade into our approach. These modifications aimed to improve convergence and overall performance of the models.</li>\n</ul>\n<p><strong>Yolov6m</strong>: Two-stage models like MaskRCNN, CascadeRCNN, and HybridTaskCascade exhibit good Mean Average Precision (MAP) at a certain intersection-over-union (IoU) threshold range [0.5:0.95]. However, their overall MAP [0.5:0.95] scores may not reach high values. To address this limitation, YOLO series models offer a solution. Specifically, YOLOv6, with its default training configurations, achieves a significantly higher Mean Average Recall (MAR) of 55.1, compared to the 44.x obtained by the two-stage models. This improvement in MAR highlights the effectiveness of YOLO series models in overcoming the mentioned drawback.<br>\n<strong>Detail training config</strong></p>\n<ul>\n<li>Optimizer: AdamW, warmup 3 epoch</li>\n<li>Select best model use MAP[0.5:0.95] blood_vessel class</li>\n<li>Data augmentation: Multi scale training: [(512x512), (640,640), (768, 768), (896, 896), (1024, 1024)]</li>\n</ul>\n<h1>Training: 2 stages</h1>\n<ul>\n<li>Stage 1: training all 1549 images, 2 classes</li>\n<li>Stage2: training 338 images ds1, 2 classes</li>\n</ul>\n<p>Focusing on optimizing the validation set with dataset1 exclusively, fine-tuning the model for a few epochs using dataset1 in the training set improved performance on the validation data.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Fdd8e2019007298da32e43ffe92e1fe32%2Fsingle_model.png?generation=1691082591623984&amp;alt=media\">\n</div>\n<h1>Ensemble &amp; Post processing</h1>\n<p><strong>Ensemble</strong></p>\n<ul>\n<li>The ensemble process is as depicted in the following diagram; We utilize Weighted Boxes Fusion (WBF) When ensembling bounding boxes, as it produces superior results compared to NMS, SoftNMS, or NMW.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2F5a2d33de596e95e08f227e514cc49350%2Fensemble.png?generation=1691081285365346&amp;alt=media\" alt=\"Ensemble diagram\"><ul>\n<li>For WBF, equal weights are assigned to each model during the ensembling process.</li></ul></li>\n</ul>\n<p><strong>PostProcessing</strong></p>\n<ul>\n<li>Remove instances with an area &lt; 80 pixels.</li>\n<li>Refine instance scores using the formula: <code>score_instance = score_bbox * mean(mask[mask &gt; 0.5])</code></li>\n</ul>\n<p>Some detail experiments:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Ffabaef83a462c8b1d1f54d8ad11389ea%2Ffinal%20experiments.png?generation=1691081498124640&amp;alt=media\" alt=\"\"><br>\n<em>We observed a strong correlation between the local validation and the public leaderboard (LB) scores. Therefore, we only made submissions when there was a significant improvement in the validation score.</em></p>\n<h1>Conclusion</h1>\n<ul>\n<li>Several techniques in our approach significantly influenced the private score: using only dataset1 for validation and fine-tuning, employing light augmentation, multi-stage ensemble, utilizing YOLOv6 for higher MAR, refining prediction scores, and, undoubtedly, placing strong trust in the validation score - no dilated!!!</li>\n<li>Adopting larger models and training with larger image sizes can potentially enhance model quality. However, due to resource limitations, we were unable to implement this strategy.</li>\n</ul>\n<p><strong>Keywords</strong>: Instance segmentation, ensemble instance segmentation, re-ranking instance, refine instance score.</p>",
  "messages": [
    {
      "id": "2372420",
      "postDate": "08/03/2023 16:57:36",
      "content": "<p>Thanks to the organizers for this compelling competition. We present the 4th place solution from the GoN team at SDSRV.AI. It emphasizes the importance of dataset, validation, ensemble and post processing method. Special thanks to <a href=\"https://www.kaggle.com/ducmanvo\" target=\"_blank\">@ducmanvo</a>, an amazing teammate.</p>\n<h1>Summary</h1>\n<p>Our solution employs a combination of 4 instance segmentation models comprising CascadeRCNN (ResNeXt, Regnet), MaskRCNN (swint), HybridTaskCascade (Re2Net), and 1 object detection model, Yolov6m. Each model is trained with a dataset consisting of 1549 images (ds1+ds2), with 84 images from dataset1 utilized for validation. During the inference process, the Weighted Boxes Fusion (WBF) technique is employed to ensemble RPN boxes and ROI boxes, while the masks generated by the ensemble of models utilize the mean average approach. For post-processing, we filter out small instances and refine instance score using mask scores.</p>\n<h1>Cross-Validation and Preprocessing</h1>\n<p><strong>Train-test split</strong><br>\n    - Training: 1549 images (1211 images from ds2 + 338 images from ds1)<br>\n     - Validation: 84 images from ds1<br>\n     - During the initial stages of approaching the problem (in the final month of the competition), we initially utilized K-fold cross-validation. However, upon recognizing significant differences in distribution, label assignment methodologies between dataset1 and dataset2, as well as variations in label assignments among Whole Slide Images (WSI), we decided to split the data into two train-validation sets. The validation data was exclusively taken from dataset1 (20% of dataset1). Another reason for this data split was the scarcity of dataset1 images in the training set, which was a concern given that the private test data was exclusively sourced from dataset1 (with only 422 images available).<br>\n      - To determine the dissimilarities in the label distribution between dataset1 and dataset2, as well as variations among different Whole Slide Images (WSI), we trained a base model, MaskRCNN R50, on the training data from dataset1 (fold1) and validated it on dataset1, dataset2 (fold1). Our observations revealed that the model's performance scored considerably higher on dataset1 compared to dataset2. Similar procedures were applied when comparing performance across various Whole Slide Images (WSI).<br>\n<strong>Preprocessing</strong>: remove duplicate annotations<br>\n     - Approximately 4.5% of training data was duplicated. We directly removed these duplicate labels to ensure data integrity and improve model performance.</p>\n<h1>Training &amp; valid with blood_vessel &amp; unsure only</h1>\n<p>During training and validation, we specifically focused on the \"blood_vessel\" and \"unsure\" classes. This decision was influenced by the significant size difference between glomerulus and the other two classes, with glomerulus being much larger. Surprisingly, when training with all three classes, the model's performance on glomerulus remained exceptionally high. Additionally, the competition organizers informed us that during the testing phase, we could utilize glomerulus labels to eliminate false positives. Consequently, we concluded that training with additional glomerulus data was unnecessary. After removing glomerulus labels from the training dataset, we observed an improvement in the performance score for the \"blood_vessel\" class. Therefore, we made the decision to exclude glomerulus labels from all subsequent experiments.</p>\n<h1>Models</h1>\n<p><strong>2 stages model</strong>: MaskRCNN, CascadeRCNN, HybridTaskCascade</p>\n<ul>\n<li>When training with the base model (MaskRCNN-R50), we noticed that the model struggled to converge using the default mmdet configurations. Consequently, in subsequent experiments, we reduced the utilization of augmentation methods, retaining only multiscale training. Additionally, we adopted a larger backbone (ResNeXt-101) and replaced SGD with AdamW optimizer. To further enhance model diversity during ensemble, we incorporated CascadeRCNN and HybridTaskCascade into our approach. These modifications aimed to improve convergence and overall performance of the models.</li>\n</ul>\n<p><strong>Yolov6m</strong>: Two-stage models like MaskRCNN, CascadeRCNN, and HybridTaskCascade exhibit good Mean Average Precision (MAP) at a certain intersection-over-union (IoU) threshold range [0.5:0.95]. However, their overall MAP [0.5:0.95] scores may not reach high values. To address this limitation, YOLO series models offer a solution. Specifically, YOLOv6, with its default training configurations, achieves a significantly higher Mean Average Recall (MAR) of 55.1, compared to the 44.x obtained by the two-stage models. This improvement in MAR highlights the effectiveness of YOLO series models in overcoming the mentioned drawback.<br>\n<strong>Detail training config</strong></p>\n<ul>\n<li>Optimizer: AdamW, warmup 3 epoch</li>\n<li>Select best model use MAP[0.5:0.95] blood_vessel class</li>\n<li>Data augmentation: Multi scale training: [(512x512), (640,640), (768, 768), (896, 896), (1024, 1024)]</li>\n</ul>\n<h1>Training: 2 stages</h1>\n<ul>\n<li>Stage 1: training all 1549 images, 2 classes</li>\n<li>Stage2: training 338 images ds1, 2 classes</li>\n</ul>\n<p>Focusing on optimizing the validation set with dataset1 exclusively, fine-tuning the model for a few epochs using dataset1 in the training set improved performance on the validation data.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Fdd8e2019007298da32e43ffe92e1fe32%2Fsingle_model.png?generation=1691082591623984&amp;alt=media\">\n</div>\n<h1>Ensemble &amp; Post processing</h1>\n<p><strong>Ensemble</strong></p>\n<ul>\n<li>The ensemble process is as depicted in the following diagram; We utilize Weighted Boxes Fusion (WBF) When ensembling bounding boxes, as it produces superior results compared to NMS, SoftNMS, or NMW.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2F5a2d33de596e95e08f227e514cc49350%2Fensemble.png?generation=1691081285365346&amp;alt=media\" alt=\"Ensemble diagram\"><ul>\n<li>For WBF, equal weights are assigned to each model during the ensembling process.</li></ul></li>\n</ul>\n<p><strong>PostProcessing</strong></p>\n<ul>\n<li>Remove instances with an area &lt; 80 pixels.</li>\n<li>Refine instance scores using the formula: <code>score_instance = score_bbox * mean(mask[mask &gt; 0.5])</code></li>\n</ul>\n<p>Some detail experiments:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Ffabaef83a462c8b1d1f54d8ad11389ea%2Ffinal%20experiments.png?generation=1691081498124640&amp;alt=media\" alt=\"\"><br>\n<em>We observed a strong correlation between the local validation and the public leaderboard (LB) scores. Therefore, we only made submissions when there was a significant improvement in the validation score.</em></p>\n<h1>Conclusion</h1>\n<ul>\n<li>Several techniques in our approach significantly influenced the private score: using only dataset1 for validation and fine-tuning, employing light augmentation, multi-stage ensemble, utilizing YOLOv6 for higher MAR, refining prediction scores, and, undoubtedly, placing strong trust in the validation score - no dilated!!!</li>\n<li>Adopting larger models and training with larger image sizes can potentially enhance model quality. However, due to resource limitations, we were unable to implement this strategy.</li>\n</ul>\n<p><strong>Keywords</strong>: Instance segmentation, ensemble instance segmentation, re-ranking instance, refine instance score.</p>",
      "rawMarkdown": "Thanks to the organizers for this compelling competition. We present the 4th place solution from the GoN team at SDSRV.AI. It emphasizes the importance of dataset, validation, ensemble and post processing method. Special thanks to @ducmanvo, an amazing teammate.\n\n# Summary\n\nOur solution employs a combination of 4 instance segmentation models comprising CascadeRCNN (ResNeXt, Regnet), MaskRCNN (swint), HybridTaskCascade (Re2Net), and 1 object detection model, Yolov6m. Each model is trained with a dataset consisting of 1549 images (ds1+ds2), with 84 images from dataset1 utilized for validation. During the inference process, the Weighted Boxes Fusion (WBF) technique is employed to ensemble RPN boxes and ROI boxes, while the masks generated by the ensemble of models utilize the mean average approach. For post-processing, we filter out small instances and refine instance score using mask scores.\n\n# Cross-Validation and Preprocessing\n**Train-test split**\n    - Training: 1549 images (1211 images from ds2 + 338 images from ds1)\n     - Validation: 84 images from ds1\n     - During the initial stages of approaching the problem (in the final month of the competition), we initially utilized K-fold cross-validation. However, upon recognizing significant differences in distribution, label assignment methodologies between dataset1 and dataset2, as well as variations in label assignments among Whole Slide Images (WSI), we decided to split the data into two train-validation sets. The validation data was exclusively taken from dataset1 (20% of dataset1). Another reason for this data split was the scarcity of dataset1 images in the training set, which was a concern given that the private test data was exclusively sourced from dataset1 (with only 422 images available).\n      - To determine the dissimilarities in the label distribution between dataset1 and dataset2, as well as variations among different Whole Slide Images (WSI), we trained a base model, MaskRCNN R50, on the training data from dataset1 (fold1) and validated it on dataset1, dataset2 (fold1). Our observations revealed that the model's performance scored considerably higher on dataset1 compared to dataset2. Similar procedures were applied when comparing performance across various Whole Slide Images (WSI).\n**Preprocessing**: remove duplicate annotations\n     - Approximately 4.5% of training data was duplicated. We directly removed these duplicate labels to ensure data integrity and improve model performance.\n\n# Training & valid with blood_vessel & unsure only\nDuring training and validation, we specifically focused on the \"blood_vessel\" and \"unsure\" classes. This decision was influenced by the significant size difference between glomerulus and the other two classes, with glomerulus being much larger. Surprisingly, when training with all three classes, the model's performance on glomerulus remained exceptionally high. Additionally, the competition organizers informed us that during the testing phase, we could utilize glomerulus labels to eliminate false positives. Consequently, we concluded that training with additional glomerulus data was unnecessary. After removing glomerulus labels from the training dataset, we observed an improvement in the performance score for the \"blood_vessel\" class. Therefore, we made the decision to exclude glomerulus labels from all subsequent experiments.\n\n# Models\n**2 stages model**: MaskRCNN, CascadeRCNN, HybridTaskCascade\n- When training with the base model (MaskRCNN-R50), we noticed that the model struggled to converge using the default mmdet configurations. Consequently, in subsequent experiments, we reduced the utilization of augmentation methods, retaining only multiscale training. Additionally, we adopted a larger backbone (ResNeXt-101) and replaced SGD with AdamW optimizer. To further enhance model diversity during ensemble, we incorporated CascadeRCNN and HybridTaskCascade into our approach. These modifications aimed to improve convergence and overall performance of the models.\n\n**Yolov6m**: Two-stage models like MaskRCNN, CascadeRCNN, and HybridTaskCascade exhibit good Mean Average Precision (MAP) at a certain intersection-over-union (IoU) threshold range [0.5:0.95]. However, their overall MAP [0.5:0.95] scores may not reach high values. To address this limitation, YOLO series models offer a solution. Specifically, YOLOv6, with its default training configurations, achieves a significantly higher Mean Average Recall (MAR) of 55.1, compared to the 44.x obtained by the two-stage models. This improvement in MAR highlights the effectiveness of YOLO series models in overcoming the mentioned drawback.\n**Detail training config**\n   - Optimizer: AdamW, warmup 3 epoch\n   - Select best model use MAP[0.5:0.95] blood_vessel class\n   - Data augmentation: Multi scale training: [(512x512), (640,640), (768, 768), (896, 896), (1024, 1024)]\n\n# Training: 2 stages\n- Stage 1: training all 1549 images, 2 classes\n- Stage2: training 338 images ds1, 2 classes\n\nFocusing on optimizing the validation set with dataset1 exclusively, fine-tuning the model for a few epochs using dataset1 in the training set improved performance on the validation data.\n\n<div style=\"display: flex; justify-content: center;\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Fdd8e2019007298da32e43ffe92e1fe32%2Fsingle_model.png?generation=1691082591623984&alt=media\" width=\"430\" height=\"213\">\n</div>\n\n# Ensemble & Post processing\n**Ensemble**\n   - The ensemble process is as depicted in the following diagram; We utilize Weighted Boxes Fusion (WBF) When ensembling bounding boxes, as it produces superior results compared to NMS, SoftNMS, or NMW.\n![Ensemble diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2F5a2d33de596e95e08f227e514cc49350%2Fensemble.png?generation=1691081285365346&alt=media)\n  - For WBF, equal weights are assigned to each model during the ensembling process.\n\n**PostProcessing**\n  - Remove instances with an area < 80 pixels.\n  - Refine instance scores using the formula: `score_instance = score_bbox * mean(mask[mask > 0.5])`\n\nSome detail experiments:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Ffabaef83a462c8b1d1f54d8ad11389ea%2Ffinal%20experiments.png?generation=1691081498124640&alt=media)\n*We observed a strong correlation between the local validation and the public leaderboard (LB) scores. Therefore, we only made submissions when there was a significant improvement in the validation score.*\n\n# Conclusion\n- Several techniques in our approach significantly influenced the private score: using only dataset1 for validation and fine-tuning, employing light augmentation, multi-stage ensemble, utilizing YOLOv6 for higher MAR, refining prediction scores, and, undoubtedly, placing strong trust in the validation score - no dilated!!!\n- Adopting larger models and training with larger image sizes can potentially enhance model quality. However, due to resource limitations, we were unable to implement this strategy.\n\n\n**Keywords**: Instance segmentation, ensemble instance segmentation, re-ranking instance, refine instance score.",
      "votes": null
    },
    {
      "id": "2372780",
      "postDate": "08/04/2023 00:30:48",
      "content": "<p>Congratulations!</p>\n<p>In your Ensemble Diagram, you WBF the RPN head output of Detection (HTC, cascade, mask r cnn). Then, it was used as an input for ROI align + Box head. How is this possible in mmdetection?? Is there a notebook or github I can refer to?</p>",
      "rawMarkdown": "Congratulations!\n\nIn your Ensemble Diagram, you WBF the RPN head output of Detection (HTC, cascade, mask r cnn). Then, it was used as an input for ROI align + Box head. How is this possible in mmdetection?? Is there a notebook or github I can refer to?",
      "votes": null
    },
    {
      "id": "2372805",
      "postDate": "08/04/2023 01:17:57",
      "content": "<p>I think there is currently no public code that does that. You need to customize your inference flow. My training and inference code will be prepared and made public soon!</p>",
      "rawMarkdown": "I think there is currently no public code that does that. You need to customize your inference flow. My training and inference code will be prepared and made public soon!",
      "votes": null
    },
    {
      "id": "2372806",
      "postDate": "08/04/2023 01:19:41",
      "content": "<p>Looking forward to it thank you!!</p>",
      "rawMarkdown": "Looking forward to it thank you!!",
      "votes": null
    },
    {
      "id": "2372994",
      "postDate": "08/04/2023 05:12:47",
      "content": "<p>Great job, Looking forward to your training and inference code.</p>",
      "rawMarkdown": "Great job, Looking forward to your training and inference code.",
      "votes": null
    },
    {
      "id": "2373307",
      "postDate": "08/04/2023 07:56:26",
      "content": "<p><a href=\"https://www.kaggle.com/damtrongtuyen\" target=\"_blank\">@damtrongtuyen</a>  The reason why you use 2 step wbf. In your architecture, I check that you only use yolo like a detection model for ensemble wbf after RPN head, the yolo model is't any effect in segmentation mask, right? </p>",
      "rawMarkdown": "damtrongtuyen  The reason why you use 2 step wbf. In your architecture, I check that you only use yolo like a detection model for ensemble wbf after RPN head, the yolo model is't any effect in segmentation mask, right?",
      "votes": null
    },
    {
      "id": "2373359",
      "postDate": "08/04/2023 08:34:06",
      "content": "<p>The bbox output of rpn head and box head are different. Yes, we only use mask head from 2 stage models.</p>",
      "rawMarkdown": "The bbox output of rpn head and box head are different. Yes, we only use mask head from 2 stage models.",
      "votes": null
    },
    {
      "id": "2373396",
      "postDate": "08/04/2023 08:52:45",
      "content": "<p>But how did you ensemble RPN bbox with yolo bbox. I think that yolo bbox will be like box head bbox. They have difference size, shape, conf score,…</p>",
      "rawMarkdown": "But how did you ensemble RPN bbox with yolo bbox. I think that yolo bbox will be like box head bbox. They have difference size, shape, conf score,...",
      "votes": null
    },
    {
      "id": "2373423",
      "postDate": "08/04/2023 09:07:11",
      "content": "<p>Great finding! I drew the wrong arrow endpoint. It should be the ensemble bounding box after the boxhead, instead of the rpnhead! Thanks</p>",
      "rawMarkdown": "Great finding! I drew the wrong arrow endpoint. It should be the ensemble bounding box after the boxhead, instead of the rpnhead! Thanks",
      "votes": null
    },
    {
      "id": "2378417",
      "postDate": "08/07/2023 16:23:37",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/damtrongtuyen\" target=\"_blank\">@damtrongtuyen</a> and ChatGPT for this great report :)</p>",
      "rawMarkdown": "Thanks @damtrongtuyen and ChatGPT for this great report :)",
      "votes": null
    },
    {
      "id": "2380200",
      "postDate": "08/08/2023 13:17:13",
      "content": "<p>Code training <a href=\"https://www.kaggle.com/code/damtrongtuyen/training-hubmap-yolov6m\" target=\"_blank\">yolov6m</a> and <a href=\"https://www.kaggle.com/damtrongtuyen/training-hubmap-2-stages-model\" target=\"_blank\">2 stages model </a>: mmdet maskrcnn, cascade, htc</p>",
      "rawMarkdown": "Code training [yolov6m](https://www.kaggle.com/code/damtrongtuyen/training-hubmap-yolov6m) and [2 stages model ](https://www.kaggle.com/damtrongtuyen/training-hubmap-2-stages-model): mmdet maskrcnn, cascade, htc",
      "votes": null
    },
    {
      "id": "2381125",
      "postDate": "08/09/2023 02:29:04",
      "content": "<p><a href=\"https://github.com/amirassov/kaggle-imaterialist/blob/f1ae37100801203500d20119b9de7e19b0d89a1c/mmdetection/mmdet/models/detectors/ensemble_htc.py\" target=\"_blank\">Inference code for ensemble both rpn head and roi head(boxhead)</a></p>",
      "rawMarkdown": "[Inference code for ensemble both rpn head and roi head(boxhead)] (https://github.com/amirassov/kaggle-imaterialist/blob/f1ae37100801203500d20119b9de7e19b0d89a1c/mmdetection/mmdet/models/detectors/ensemble_htc.py)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2372780,
      "author_name": "devchopin",
      "author_url": "",
      "post_date": "08/04/2023 00:30:48",
      "content": "<p>Congratulations!</p>\n<p>In your Ensemble Diagram, you WBF the RPN head output of Detection (HTC, cascade, mask r cnn). Then, it was used as an input for ROI align + Box head. How is this possible in mmdetection?? Is there a notebook or github I can refer to?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2372805,
          "author_name": "damtrongtuyen",
          "author_url": "",
          "post_date": "08/04/2023 01:17:57",
          "content": "<p>I think there is currently no public code that does that. You need to customize your inference flow. My training and inference code will be prepared and made public soon!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2372806,
              "author_name": "devchopin",
              "author_url": "",
              "post_date": "08/04/2023 01:19:41",
              "content": "<p>Looking forward to it thank you!!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2372994,
      "author_name": "allenwpr",
      "author_url": "",
      "post_date": "08/04/2023 05:12:47",
      "content": "<p>Great job, Looking forward to your training and inference code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2373307,
      "author_name": "hidngnguyna",
      "author_url": "",
      "post_date": "08/04/2023 07:56:26",
      "content": "<p><a href=\"https://www.kaggle.com/damtrongtuyen\" target=\"_blank\">@damtrongtuyen</a>  The reason why you use 2 step wbf. In your architecture, I check that you only use yolo like a detection model for ensemble wbf after RPN head, the yolo model is't any effect in segmentation mask, right? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2373359,
          "author_name": "damtrongtuyen",
          "author_url": "",
          "post_date": "08/04/2023 08:34:06",
          "content": "<p>The bbox output of rpn head and box head are different. Yes, we only use mask head from 2 stage models.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2373396,
              "author_name": "hidngnguyna",
              "author_url": "",
              "post_date": "08/04/2023 08:52:45",
              "content": "<p>But how did you ensemble RPN bbox with yolo bbox. I think that yolo bbox will be like box head bbox. They have difference size, shape, conf score,…</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2373423,
                  "author_name": "damtrongtuyen",
                  "author_url": "",
                  "post_date": "08/04/2023 09:07:11",
                  "content": "<p>Great finding! I drew the wrong arrow endpoint. It should be the ensemble bounding box after the boxhead, instead of the rpnhead! Thanks</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2378417,
      "author_name": "vslaykovsky",
      "author_url": "",
      "post_date": "08/07/2023 16:23:37",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/damtrongtuyen\" target=\"_blank\">@damtrongtuyen</a> and ChatGPT for this great report :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2380200,
      "author_name": "damtrongtuyen",
      "author_url": "",
      "post_date": "08/08/2023 13:17:13",
      "content": "<p>Code training <a href=\"https://www.kaggle.com/code/damtrongtuyen/training-hubmap-yolov6m\" target=\"_blank\">yolov6m</a> and <a href=\"https://www.kaggle.com/damtrongtuyen/training-hubmap-2-stages-model\" target=\"_blank\">2 stages model </a>: mmdet maskrcnn, cascade, htc</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2381125,
      "author_name": "damtrongtuyen",
      "author_url": "",
      "post_date": "08/09/2023 02:29:04",
      "content": "<p><a href=\"https://github.com/amirassov/kaggle-imaterialist/blob/f1ae37100801203500d20119b9de7e19b0d89a1c/mmdetection/mmdet/models/detectors/ensemble_htc.py\" target=\"_blank\">Inference code for ensemble both rpn head and roi head(boxhead)</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2372420": "Thanks to the organizers for this compelling competition. We present the 4th place solution from the GoN team at SDSRV.AI. It emphasizes the importance of dataset, validation, ensemble and post processing method. Special thanks to @ducmanvo, an amazing teammate.\n\n# Summary\n\nOur solution employs a combination of 4 instance segmentation models comprising CascadeRCNN (ResNeXt, Regnet), MaskRCNN (swint), HybridTaskCascade (Re2Net), and 1 object detection model, Yolov6m. Each model is trained with a dataset consisting of 1549 images (ds1+ds2), with 84 images from dataset1 utilized for validation. During the inference process, the Weighted Boxes Fusion (WBF) technique is employed to ensemble RPN boxes and ROI boxes, while the masks generated by the ensemble of models utilize the mean average approach. For post-processing, we filter out small instances and refine instance score using mask scores.\n\n# Cross-Validation and Preprocessing\n**Train-test split**\n    - Training: 1549 images (1211 images from ds2 + 338 images from ds1)\n     - Validation: 84 images from ds1\n     - During the initial stages of approaching the problem (in the final month of the competition), we initially utilized K-fold cross-validation. However, upon recognizing significant differences in distribution, label assignment methodologies between dataset1 and dataset2, as well as variations in label assignments among Whole Slide Images (WSI), we decided to split the data into two train-validation sets. The validation data was exclusively taken from dataset1 (20% of dataset1). Another reason for this data split was the scarcity of dataset1 images in the training set, which was a concern given that the private test data was exclusively sourced from dataset1 (with only 422 images available).\n      - To determine the dissimilarities in the label distribution between dataset1 and dataset2, as well as variations among different Whole Slide Images (WSI), we trained a base model, MaskRCNN R50, on the training data from dataset1 (fold1) and validated it on dataset1, dataset2 (fold1). Our observations revealed that the model's performance scored considerably higher on dataset1 compared to dataset2. Similar procedures were applied when comparing performance across various Whole Slide Images (WSI).\n**Preprocessing**: remove duplicate annotations\n     - Approximately 4.5% of training data was duplicated. We directly removed these duplicate labels to ensure data integrity and improve model performance.\n\n# Training & valid with blood_vessel & unsure only\nDuring training and validation, we specifically focused on the \"blood_vessel\" and \"unsure\" classes. This decision was influenced by the significant size difference between glomerulus and the other two classes, with glomerulus being much larger. Surprisingly, when training with all three classes, the model's performance on glomerulus remained exceptionally high. Additionally, the competition organizers informed us that during the testing phase, we could utilize glomerulus labels to eliminate false positives. Consequently, we concluded that training with additional glomerulus data was unnecessary. After removing glomerulus labels from the training dataset, we observed an improvement in the performance score for the \"blood_vessel\" class. Therefore, we made the decision to exclude glomerulus labels from all subsequent experiments.\n\n# Models\n**2 stages model**: MaskRCNN, CascadeRCNN, HybridTaskCascade\n- When training with the base model (MaskRCNN-R50), we noticed that the model struggled to converge using the default mmdet configurations. Consequently, in subsequent experiments, we reduced the utilization of augmentation methods, retaining only multiscale training. Additionally, we adopted a larger backbone (ResNeXt-101) and replaced SGD with AdamW optimizer. To further enhance model diversity during ensemble, we incorporated CascadeRCNN and HybridTaskCascade into our approach. These modifications aimed to improve convergence and overall performance of the models.\n\n**Yolov6m**: Two-stage models like MaskRCNN, CascadeRCNN, and HybridTaskCascade exhibit good Mean Average Precision (MAP) at a certain intersection-over-union (IoU) threshold range [0.5:0.95]. However, their overall MAP [0.5:0.95] scores may not reach high values. To address this limitation, YOLO series models offer a solution. Specifically, YOLOv6, with its default training configurations, achieves a significantly higher Mean Average Recall (MAR) of 55.1, compared to the 44.x obtained by the two-stage models. This improvement in MAR highlights the effectiveness of YOLO series models in overcoming the mentioned drawback.\n**Detail training config**\n   - Optimizer: AdamW, warmup 3 epoch\n   - Select best model use MAP[0.5:0.95] blood_vessel class\n   - Data augmentation: Multi scale training: [(512x512), (640,640), (768, 768), (896, 896), (1024, 1024)]\n\n# Training: 2 stages\n- Stage 1: training all 1549 images, 2 classes\n- Stage2: training 338 images ds1, 2 classes\n\nFocusing on optimizing the validation set with dataset1 exclusively, fine-tuning the model for a few epochs using dataset1 in the training set improved performance on the validation data.\n\n<div style=\"display: flex; justify-content: center;\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Fdd8e2019007298da32e43ffe92e1fe32%2Fsingle_model.png?generation=1691082591623984&alt=media\" width=\"430\" height=\"213\">\n</div>\n\n# Ensemble & Post processing\n**Ensemble**\n   - The ensemble process is as depicted in the following diagram; We utilize Weighted Boxes Fusion (WBF) When ensembling bounding boxes, as it produces superior results compared to NMS, SoftNMS, or NMW.\n![Ensemble diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2F5a2d33de596e95e08f227e514cc49350%2Fensemble.png?generation=1691081285365346&alt=media)\n  - For WBF, equal weights are assigned to each model during the ensembling process.\n\n**PostProcessing**\n  - Remove instances with an area < 80 pixels.\n  - Refine instance scores using the formula: `score_instance = score_bbox * mean(mask[mask > 0.5])`\n\nSome detail experiments:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6439446%2Ffabaef83a462c8b1d1f54d8ad11389ea%2Ffinal%20experiments.png?generation=1691081498124640&alt=media)\n*We observed a strong correlation between the local validation and the public leaderboard (LB) scores. Therefore, we only made submissions when there was a significant improvement in the validation score.*\n\n# Conclusion\n- Several techniques in our approach significantly influenced the private score: using only dataset1 for validation and fine-tuning, employing light augmentation, multi-stage ensemble, utilizing YOLOv6 for higher MAR, refining prediction scores, and, undoubtedly, placing strong trust in the validation score - no dilated!!!\n- Adopting larger models and training with larger image sizes can potentially enhance model quality. However, due to resource limitations, we were unable to implement this strategy.\n\n\n**Keywords**: Instance segmentation, ensemble instance segmentation, re-ranking instance, refine instance score.",
    "2372780": "Congratulations!\n\nIn your Ensemble Diagram, you WBF the RPN head output of Detection (HTC, cascade, mask r cnn). Then, it was used as an input for ROI align + Box head. How is this possible in mmdetection?? Is there a notebook or github I can refer to?",
    "2372805": "I think there is currently no public code that does that. You need to customize your inference flow. My training and inference code will be prepared and made public soon!",
    "2372806": "Looking forward to it thank you!!",
    "2372994": "Great job, Looking forward to your training and inference code.",
    "2373307": "damtrongtuyen  The reason why you use 2 step wbf. In your architecture, I check that you only use yolo like a detection model for ensemble wbf after RPN head, the yolo model is't any effect in segmentation mask, right?",
    "2373359": "The bbox output of rpn head and box head are different. Yes, we only use mask head from 2 stage models.",
    "2373396": "But how did you ensemble RPN bbox with yolo bbox. I think that yolo bbox will be like box head bbox. They have difference size, shape, conf score,...",
    "2373423": "Great finding! I drew the wrong arrow endpoint. It should be the ensemble bounding box after the boxhead, instead of the rpnhead! Thanks",
    "2378417": "Thanks @damtrongtuyen and ChatGPT for this great report :)",
    "2380200": "Code training [yolov6m](https://www.kaggle.com/code/damtrongtuyen/training-hubmap-yolov6m) and [2 stages model ](https://www.kaggle.com/damtrongtuyen/training-hubmap-2-stages-model): mmdet maskrcnn, cascade, htc",
    "2381125": "[Inference code for ensemble both rpn head and roi head(boxhead)] (https://github.com/amirassov/kaggle-imaterialist/blob/f1ae37100801203500d20119b9de7e19b0d89a1c/mmdetection/mmdet/models/detectors/ensemble_htc.py)"
  },
  "source": "meta"
}