{
  "id": 429409,
  "title": "(Best private 0.602) 297th Place Solution for the HuBMAP - Hacking the Human Vasculature Competition. How I lost the gold medal.",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/429409",
  "author_name": "Quan Vu",
  "post_date": "2023-08-05T13:10:04.047000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This is the first time I solve a segmentation task.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>My final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&amp;alt=media\" alt=\"\"></p>\n<h1>Details of the submission</h1>\n<h2>Cross validation experiments</h2>\n<p>Maskdino swin-L, image size 1440:</p>\n<ul>\n<li>Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60</li>\n<li>Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66</li>\n<li>Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42</li>\n<li>Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32<br>\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.</li>\n</ul>\n<h2>Experiments on dataset 1</h2>\n<ul>\n<li>Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.</li>\n<li>Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.</li>\n<li>Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.</li>\n<li>Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.</li>\n<li>Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.</li>\n</ul>\n<h2>Important of image size</h2>\n<p>Dataset have a lot of small mask. I found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -&gt; 1440</p>\n<h2>My best private is 0.597</h2>\n<p>Public notebook: <a href=\"https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\" target=\"_blank\">https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597</a></p>\n<h2>My final submission</h2>\n<p>Final submission 1 (Version 8), mask threshold &gt; 0.6, using dialate: Private 0.334, Public: 0.551<br>\nFinal submission 2 (Version 1), mask threshold &gt; 0.03, not using dialate: Private 0.394, Public: 0.539<br>\nLate submission: modify from <strong>Final submission 1 (version 8)</strong> not using dilate and I achieve Private leaderboard 0.602<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&amp;alt=media\" alt=\"\"></p>\n<h2>Code optimization</h2>\n<ul>\n<li>To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.</li>\n<li>Modify maskdino source code to add EMA.</li>\n<li>Modify WBF to work with mask.</li>\n<li>Fix bug to use maskdino on kaggle notebook.</li>\n<li>Fix bug to train Internimage with cuda 12.0.</li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.</li>\n<li>Using SAM.</li>\n</ul>\n<h2>Conclusion</h2>\n<ul>\n<li>I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate =&gt; Size of masks not change between two submission.</li>\n<li>IOU is very sensitive with small mask. I experiment with dilate and not dilate, IOU change 40% with bbox area &lt; 1000.</li>\n<li>Using only dataset 1 good for both public and private.</li>\n<li>When train with dataset 2, dilate good for public but not good for private.</li>\n</ul>\n<h1>Sources</h1>\n<h2>Two my submission notebook:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version</a></li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/not-dialate-version-7\" target=\"_blank\">https://www.kaggle.com/code/quan0095/not-dialate-version-7</a></li>\n</ul>\n<h2>Late submission notebook (Private 0.602):</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303</a></li>\n</ul>\n<h2>Maskdino inference:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\" target=\"_blank\">https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100</a></li>\n</ul>\n<h2>Cascade maskrcnn internimage:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\" target=\"_blank\">https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference</a></li>\n</ul>",
  "messages": [
    {
      "id": 2375164,
      "postDate": "2023-08-05T13:10:04.047Z",
      "content": "<p>This is the first time I solve a segmentation task.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>My final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&amp;alt=media\" alt=\"\"></p>\n<h1>Details of the submission</h1>\n<h2>Cross validation experiments</h2>\n<p>Maskdino swin-L, image size 1440:</p>\n<ul>\n<li>Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60</li>\n<li>Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66</li>\n<li>Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42</li>\n<li>Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32<br>\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.</li>\n</ul>\n<h2>Experiments on dataset 1</h2>\n<ul>\n<li>Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.</li>\n<li>Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.</li>\n<li>Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.</li>\n<li>Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.</li>\n<li>Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.</li>\n</ul>\n<h2>Important of image size</h2>\n<p>Dataset have a lot of small mask. I found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -&gt; 1440</p>\n<h2>My best private is 0.597</h2>\n<p>Public notebook: <a href=\"https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\" target=\"_blank\">https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597</a></p>\n<h2>My final submission</h2>\n<p>Final submission 1 (Version 8), mask threshold &gt; 0.6, using dialate: Private 0.334, Public: 0.551<br>\nFinal submission 2 (Version 1), mask threshold &gt; 0.03, not using dialate: Private 0.394, Public: 0.539<br>\nLate submission: modify from <strong>Final submission 1 (version 8)</strong> not using dilate and I achieve Private leaderboard 0.602<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&amp;alt=media\" alt=\"\"></p>\n<h2>Code optimization</h2>\n<ul>\n<li>To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.</li>\n<li>Modify maskdino source code to add EMA.</li>\n<li>Modify WBF to work with mask.</li>\n<li>Fix bug to use maskdino on kaggle notebook.</li>\n<li>Fix bug to train Internimage with cuda 12.0.</li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.</li>\n<li>Using SAM.</li>\n</ul>\n<h2>Conclusion</h2>\n<ul>\n<li>I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate =&gt; Size of masks not change between two submission.</li>\n<li>IOU is very sensitive with small mask. I experiment with dilate and not dilate, IOU change 40% with bbox area &lt; 1000.</li>\n<li>Using only dataset 1 good for both public and private.</li>\n<li>When train with dataset 2, dilate good for public but not good for private.</li>\n</ul>\n<h1>Sources</h1>\n<h2>Two my submission notebook:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version</a></li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/not-dialate-version-7\" target=\"_blank\">https://www.kaggle.com/code/quan0095/not-dialate-version-7</a></li>\n</ul>\n<h2>Late submission notebook (Private 0.602):</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303</a></li>\n</ul>\n<h2>Maskdino inference:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\" target=\"_blank\">https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100</a></li>\n</ul>\n<h2>Cascade maskrcnn internimage:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\" target=\"_blank\">https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference</a></li>\n</ul>",
      "rawMarkdown": "This is the first time I solve a segmentation task.\n# Context \n- Business context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\n- Data context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n# Overview of the approach\nMy final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&alt=media)\n# Details of the submission\n## Cross validation experiments\nMaskdino swin-L, image size 1440:\n- Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60\n- Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66\n- Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42\n- Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.\n## Experiments on dataset 1\n- Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.\n- Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.\n- Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.\n- Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.\n- Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.\n## Important of image size\nDataset have a lot of small mask. I found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -> 1440\n## My best private is 0.597\nPublic notebook: https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\n## My final submission \nFinal submission 1 (Version 8), mask threshold > 0.6, using dialate: Private 0.334, Public: 0.551\nFinal submission 2 (Version 1), mask threshold > 0.03, not using dialate: Private 0.394, Public: 0.539\nLate submission: modify from **Final submission 1 (version 8)** not using dilate and I achieve Private leaderboard 0.602\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&alt=media)\n## Code optimization\n- To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.\n- Modify maskdino source code to add EMA.\n- Modify WBF to work with mask.\n- Fix bug to use maskdino on kaggle notebook.\n- Fix bug to train Internimage with cuda 12.0.\n## What didn't work\n- Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.\n- Using SAM.\n## Conclusion\n- I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate => Size of masks not change between two submission.\n- IOU is very sensitive with small mask. I experiment with dilate and not dilate, IOU change 40% with bbox area < 1000.\n- Using only dataset 1 good for both public and private.\n- When train with dataset 2, dilate good for public but not good for private.\n# Sources\n##Two my submission notebook: \n- https://www.kaggle.com/code/quan0095/dialate-version\n- https://www.kaggle.com/code/quan0095/not-dialate-version-7\n##Late submission notebook (Private 0.602):\n- https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\n##Maskdino inference:\n- https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\n##Cascade maskrcnn internimage:\n- https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2375164": "This is the first time I solve a segmentation task.\n# Context \n- Business context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\n- Data context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n# Overview of the approach\nMy final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&alt=media)\n# Details of the submission\n## Cross validation experiments\nMaskdino swin-L, image size 1440:\n- Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60\n- Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66\n- Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42\n- Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.\n## Experiments on dataset 1\n- Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.\n- Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.\n- Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.\n- Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.\n- Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.\n## Important of image size\nDataset have a lot of small mask. I found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -> 1440\n## My best private is 0.597\nPublic notebook: https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\n## My final submission \nFinal submission 1 (Version 8), mask threshold > 0.6, using dialate: Private 0.334, Public: 0.551\nFinal submission 2 (Version 1), mask threshold > 0.03, not using dialate: Private 0.394, Public: 0.539\nLate submission: modify from **Final submission 1 (version 8)** not using dilate and I achieve Private leaderboard 0.602\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&alt=media)\n## Code optimization\n- To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.\n- Modify maskdino source code to add EMA.\n- Modify WBF to work with mask.\n- Fix bug to use maskdino on kaggle notebook.\n- Fix bug to train Internimage with cuda 12.0.\n## What didn't work\n- Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.\n- Using SAM.\n## Conclusion\n- I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate => Size of masks not change between two submission.\n- IOU is very sensitive with small mask. I experiment with dilate and not dilate, IOU change 40% with bbox area < 1000.\n- Using only dataset 1 good for both public and private.\n- When train with dataset 2, dilate good for public but not good for private.\n# Sources\n##Two my submission notebook: \n- https://www.kaggle.com/code/quan0095/dialate-version\n- https://www.kaggle.com/code/quan0095/not-dialate-version-7\n##Late submission notebook (Private 0.602):\n- https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\n##Maskdino inference:\n- https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\n##Cascade maskrcnn internimage:\n- https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference"
  }
}