{
  "id": 428315,
  "title": "Best private 0.597. Key is dataset 1?",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/428315",
  "author_name": "Quan Vu",
  "post_date": "2023-08-01T02:11:10.288000",
  "votes": 13,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This is the first time I solve a segmentation task.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>My final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&amp;alt=media\" alt=\"\"></p>\n<h1>Details of the submission</h1>\n<h2>Cross validation experiments</h2>\n<p>Maskdino swin-L, image size 1440:</p>\n<ul>\n<li>Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60</li>\n<li>Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66</li>\n<li>Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42</li>\n<li>Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32<br>\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.</li>\n</ul>\n<h2>Experiments on dataset 1</h2>\n<ul>\n<li>Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.</li>\n<li>Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.</li>\n<li>Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.</li>\n<li>Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.</li>\n<li>Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.</li>\n</ul>\n<h2>Important of image size</h2>\n<p>I found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -&gt; 1440</p>\n<h2>My best private is 0.597</h2>\n<p>Public notebook: <a href=\"https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\" target=\"_blank\">https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597</a></p>\n<h2>My final submission</h2>\n<p>Final submission 1 (Version 8), mask threshold &gt; 0.6, using dialate: Private 0.334, Public: 0.551<br>\nFinal submission 2 (Version 1), mask threshold &gt; 0.03, not using dialate: Private 0.394, Public: 0.539<br>\nLate submission: modify from final submission 1 (version 8) not using dilate and I achieve Private leaderboard 0.602<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&amp;alt=media\" alt=\"\"></p>\n<h2>Code optimization</h2>\n<ul>\n<li>To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.</li>\n<li>Modify maskdino source code to add EMA.</li>\n<li>Modify WBF to work with mask.</li>\n<li>Fix bug to use maskdino on kaggle notebook.</li>\n<li>Fix bug to train Internimage with cuda 12.0.</li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.</li>\n<li>Using SAM.</li>\n</ul>\n<h2>Conclusion</h2>\n<ul>\n<li>I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate =&gt; Size of masks not change between two submission.</li>\n<li>Using only dataset 1 good for both public and private.</li>\n<li>When train with dataset 2, dilate good for public but not good for private.</li>\n</ul>\n<h1>Sources</h1>\n<p>Two my submission notebook: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version</a></li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/not-dialate-version-7\" target=\"_blank\">https://www.kaggle.com/code/quan0095/not-dialate-version-7</a><br>\nLate submission notebook (Private 0.602):</li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303</a><br>\nMaskdino inference:</li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\" target=\"_blank\">https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100</a><br>\nCascade maskrcnn internimage:</li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\" target=\"_blank\">https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference</a></li>\n</ul>",
  "messages": [
    {
      "id": 2368078,
      "postDate": "2023-08-01T02:11:10.290Z",
      "content": "<p>This is the first time I solve a segmentation task.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>My final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&amp;alt=media\" alt=\"\"></p>\n<h1>Details of the submission</h1>\n<h2>Cross validation experiments</h2>\n<p>Maskdino swin-L, image size 1440:</p>\n<ul>\n<li>Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60</li>\n<li>Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66</li>\n<li>Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42</li>\n<li>Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32<br>\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.</li>\n</ul>\n<h2>Experiments on dataset 1</h2>\n<ul>\n<li>Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.</li>\n<li>Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.</li>\n<li>Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.</li>\n<li>Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.</li>\n<li>Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.</li>\n</ul>\n<h2>Important of image size</h2>\n<p>I found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -&gt; 1440</p>\n<h2>My best private is 0.597</h2>\n<p>Public notebook: <a href=\"https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\" target=\"_blank\">https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597</a></p>\n<h2>My final submission</h2>\n<p>Final submission 1 (Version 8), mask threshold &gt; 0.6, using dialate: Private 0.334, Public: 0.551<br>\nFinal submission 2 (Version 1), mask threshold &gt; 0.03, not using dialate: Private 0.394, Public: 0.539<br>\nLate submission: modify from final submission 1 (version 8) not using dilate and I achieve Private leaderboard 0.602<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&amp;alt=media\" alt=\"\"></p>\n<h2>Code optimization</h2>\n<ul>\n<li>To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.</li>\n<li>Modify maskdino source code to add EMA.</li>\n<li>Modify WBF to work with mask.</li>\n<li>Fix bug to use maskdino on kaggle notebook.</li>\n<li>Fix bug to train Internimage with cuda 12.0.</li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.</li>\n<li>Using SAM.</li>\n</ul>\n<h2>Conclusion</h2>\n<ul>\n<li>I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate =&gt; Size of masks not change between two submission.</li>\n<li>Using only dataset 1 good for both public and private.</li>\n<li>When train with dataset 2, dilate good for public but not good for private.</li>\n</ul>\n<h1>Sources</h1>\n<p>Two my submission notebook: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version</a></li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/not-dialate-version-7\" target=\"_blank\">https://www.kaggle.com/code/quan0095/not-dialate-version-7</a><br>\nLate submission notebook (Private 0.602):</li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\" target=\"_blank\">https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303</a><br>\nMaskdino inference:</li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\" target=\"_blank\">https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100</a><br>\nCascade maskrcnn internimage:</li>\n<li><a href=\"https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\" target=\"_blank\">https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference</a></li>\n</ul>",
      "rawMarkdown": "This is the first time I solve a segmentation task.\n# Context \n- Business context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\n- Data context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n# Overview of the approach\nMy final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&alt=media)\n# Details of the submission\n## Cross validation experiments\nMaskdino swin-L, image size 1440:\n- Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60\n- Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66\n- Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42\n- Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.\n## Experiments on dataset 1\n- Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.\n- Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.\n- Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.\n- Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.\n- Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.\n## Important of image size\nI found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -> 1440\n## My best private is 0.597\nPublic notebook: https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\n## My final submission \nFinal submission 1 (Version 8), mask threshold > 0.6, using dialate: Private 0.334, Public: 0.551\nFinal submission 2 (Version 1), mask threshold > 0.03, not using dialate: Private 0.394, Public: 0.539\nLate submission: modify from final submission 1 (version 8) not using dilate and I achieve Private leaderboard 0.602\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&alt=media)\n## Code optimization\n- To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.\n- Modify maskdino source code to add EMA.\n- Modify WBF to work with mask.\n- Fix bug to use maskdino on kaggle notebook.\n- Fix bug to train Internimage with cuda 12.0.\n## What didn't work\n- Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.\n- Using SAM.\n## Conclusion\n- I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate => Size of masks not change between two submission.\n- Using only dataset 1 good for both public and private.\n- When train with dataset 2, dilate good for public but not good for private.\n# Sources\nTwo my submission notebook: \n- https://www.kaggle.com/code/quan0095/dialate-version\n- https://www.kaggle.com/code/quan0095/not-dialate-version-7\nLate submission notebook (Private 0.602):\n- https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\nMaskdino inference:\n- https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\nCascade maskrcnn internimage:\n- https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\n",
      "votes": 13
    },
    {
      "id": 2368335,
      "postDate": "2023-08-01T05:58:34.300Z",
      "content": "<p>thanks for sharing!! I can't use maskdino in kaggle notebook.<br>\nIf possible, could you publish one of the inference notebooks using maskdino?</p>",
      "rawMarkdown": "thanks for sharing!! I can't use maskdino in kaggle notebook.\nIf possible, could you publish one of the inference notebooks using maskdino?",
      "votes": 1,
      "replies": [
        {
          "id": 2368391,
          "postDate": "2023-08-01T06:31:35.457Z",
          "content": "<p>I have published inference notebooks using maskdino with swin-L dataset 1. Notebook is running.<br>\n<a href=\"https://www.kaggle.com/quan0095/inference-of-maskdino-p100\" target=\"_blank\">https://www.kaggle.com/quan0095/inference-of-maskdino-p100</a></p>",
          "rawMarkdown": "I have published inference notebooks using maskdino with swin-L dataset 1. Notebook is running.\nhttps://www.kaggle.com/quan0095/inference-of-maskdino-p100",
          "votes": 1
        }
      ]
    },
    {
      "id": 2368098,
      "postDate": "2023-08-01T02:30:12.860Z",
      "content": "<p>can you share how did you train with Internimage for instance segmentation? I tried but it wasn't working</p>",
      "rawMarkdown": "can you share how did you train with Internimage for instance segmentation? I tried but it wasn't working",
      "replies": [
        {
          "id": 2368105,
          "postDate": "2023-08-01T02:34:47.907Z",
          "content": "<p>I modify config <a href=\"https://github.com/OpenGVLab/InternImage/blob/master/detection/configs/coco/cascade_internimage_l_fpn_3x_coco.py\" target=\"_blank\">https://github.com/OpenGVLab/InternImage/blob/master/detection/configs/coco/cascade_internimage_l_fpn_3x_coco.py</a></p>",
          "rawMarkdown": "I modify config https://github.com/OpenGVLab/InternImage/blob/master/detection/configs/coco/cascade_internimage_l_fpn_3x_coco.py"
        },
        {
          "id": 2368449,
          "postDate": "2023-08-01T07:11:14.923Z",
          "content": "<p>I train Internimage on cuda driver 12 and I can't train with ddp. I reinstall pycocotool and modify file train.py.<br>\nMy inference notebook with train config: <a href=\"https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\" target=\"_blank\">https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference</a></p>",
          "rawMarkdown": "I train Internimage on cuda driver 12 and I can't train with ddp. I reinstall pycocotool and modify file train.py.\nMy inference notebook with train config: https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference"
        }
      ]
    },
    {
      "id": 2368147,
      "postDate": "2023-08-01T03:09:12.377Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2368335,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2023-08-01T05:58:34.300000",
      "content": "<p>thanks for sharing!! I can't use maskdino in kaggle notebook.<br>\nIf possible, could you publish one of the inference notebooks using maskdino?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2368391,
          "author_name": "Quan Vu",
          "author_url": "",
          "post_date": "2023-08-01T06:31:35.457000",
          "content": "<p>I have published inference notebooks using maskdino with swin-L dataset 1. Notebook is running.<br>\n<a href=\"https://www.kaggle.com/quan0095/inference-of-maskdino-p100\" target=\"_blank\">https://www.kaggle.com/quan0095/inference-of-maskdino-p100</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2368098,
      "author_name": "Mr.CheckerChichChich",
      "author_url": "",
      "post_date": "2023-08-01T02:30:12.860000",
      "content": "<p>can you share how did you train with Internimage for instance segmentation? I tried but it wasn't working</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2368105,
          "author_name": "Quan Vu",
          "author_url": "",
          "post_date": "2023-08-01T02:34:47.907000",
          "content": "<p>I modify config <a href=\"https://github.com/OpenGVLab/InternImage/blob/master/detection/configs/coco/cascade_internimage_l_fpn_3x_coco.py\" target=\"_blank\">https://github.com/OpenGVLab/InternImage/blob/master/detection/configs/coco/cascade_internimage_l_fpn_3x_coco.py</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2368449,
          "author_name": "Quan Vu",
          "author_url": "",
          "post_date": "2023-08-01T07:11:14.923000",
          "content": "<p>I train Internimage on cuda driver 12 and I can't train with ddp. I reinstall pycocotool and modify file train.py.<br>\nMy inference notebook with train config: <a href=\"https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\" target=\"_blank\">https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2368147,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-01T03:09:12.377000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2368078": "This is the first time I solve a segmentation task.\n# Context \n- Business context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview\n- Data context: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n# Overview of the approach\nMy final model was a combination of 13 single models. 4 model maskdino swin-L train on different data, 4 model yolov7-seg, 1 model maskdino resnet 50, 4 model Cascade maskrcnn Internimage-L. For validation, I split to 4 folds, 3 wsi for training, the other for testing. Inference with multiple size, Maskdino swin-L use 3 image size [512, 1024, 1440], yolov7 use 2 image size [512, 1024], maskdino resnet50 use 2 image size [1024, 1440], cascade maskrcnn internimage use 1 image size [1440]. Using Weight Mask Fusion for ensemble. No use external data or pseudo label.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F0df50177802f0231fdb3a944e6808cd4%2FScreenshot%20from%202023-08-05%2019-33-28.png?generation=1691238892230812&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F576a5d3446d03ac0951a64c898d75ec4%2Fhubmap.drawio.png?generation=1691237415800097&alt=media)\n# Details of the submission\n## Cross validation experiments\nMaskdino swin-L, image size 1440:\n- Fold 0 (val on wsi1): bbox mAP@0.5-0.95=34.60\n- Fold 1 (val on wsi2): bbox mAP@0.5-0.95=14.66\n- Fold 2 (val on wsi3): bbox mAP@0.5-0.95=40.42\n- Fold 3 (val on wsi4): bbox mAP@0.5-0.95=28.32\nResult on fold 0 and fold 1 are very strange. This make me confuse. Should choose dataset 1 or all dataset for training? These results are the reason why I don't choose only dataset 1 for final submission.\n## Experiments on dataset 1\n- Maskdino swin-L, image size 1408, train only all dataset1. One model without dilate: Public 0.51, Private 0.543.\n- Yolov7-seg, image size 1024, train only all dataset1. One model without dilate: Public 0.3, Private 0.446.\n- Internimage-L, image size 1440, train only all dataset1. One model without dilate: Public 0.453, Private 0.502.\n- Maskdino resnet50, image size 1024, train only all dataset1. One model without dilate: Public 0.46, Private 0.464.\n- Internimage-L, image size [(1024, 1024), (1440, 1440), (1536, 1536)], train only all wsi2. One model without dilate: Public 0.439, Private 0.469.\n## Important of image size\nI found increase image size, boost both CV and Public leader board. However, CV not increase when image size 1024 -> 1440\n## My best private is 0.597\nPublic notebook: https://www.kaggle.com/code/quan0095/maskdino-yolov7-internimage-private-0-597\n## My final submission \nFinal submission 1 (Version 8), mask threshold > 0.6, using dialate: Private 0.334, Public: 0.551\nFinal submission 2 (Version 1), mask threshold > 0.03, not using dialate: Private 0.394, Public: 0.539\nLate submission: modify from final submission 1 (version 8) not using dilate and I achieve Private leaderboard 0.602\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7aa3482c9ca8b4a2830348cf3468d942%2FScreenshot%20from%202023-08-01%2022-41-23.png?generation=1690904619445688&alt=media)\n## Code optimization\n- To ensemble many models, I crop mask using bbox from 512x512 original mask to reduce CPU RAM usage. Using multi thread to reduce time run notebook.\n- Modify maskdino source code to add EMA.\n- Modify WBF to work with mask.\n- Fix bug to use maskdino on kaggle notebook.\n- Fix bug to train Internimage with cuda 12.0.\n## What didn't work\n- Self-supervised learning using dinov1 with dataset3 and hubmap kidney dataset.\n- Using SAM.\n## Conclusion\n- I fail gold medal, because I have a mistake: I submit two submission with dilate and not dilate, but I have changed mask threshold when not dilate => Size of masks not change between two submission.\n- Using only dataset 1 good for both public and private.\n- When train with dataset 2, dilate good for public but not good for private.\n# Sources\nTwo my submission notebook: \n- https://www.kaggle.com/code/quan0095/dialate-version\n- https://www.kaggle.com/code/quan0095/not-dialate-version-7\nLate submission notebook (Private 0.602):\n- https://www.kaggle.com/code/quan0095/dialate-version?scriptVersionId=138564303\nMaskdino inference:\n- https://www.kaggle.com/code/quan0095/inference-of-maskdino-p100\nCascade maskrcnn internimage:\n- https://www.kaggle.com/code/quan0095/internimage-mmdet-2-26-inference\n",
    "2368335": "thanks for sharing!! I can't use maskdino in kaggle notebook.\nIf possible, could you publish one of the inference notebooks using maskdino?",
    "2368098": "can you share how did you train with Internimage for instance segmentation? I tried but it wasn't working",
    "2368147": ""
  }
}