{
  "id": 428319,
  "title": "12th place solution",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/writeups/devchopin-12th-place-solution",
  "author_name": "",
  "post_date": "2023-08-01T03:14:55.327Z",
  "votes": 18,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I am so happy to have won my first gold medal and even a solo medal in this competition. Also thanks to the  OpenMMLab, SenseTime and The Chinese University of Hong Kong for making a really great library mmdetection.</p>\n<p><strong>Summary</strong><br>\nI used the pseudo labels of Dataset3 to train the Cascade Mask RCNN + Convnext v2 large.</p>\n<p><strong>Model</strong></p>\n<ul>\n<li>Detector : Cascade Mask RCNN </li>\n<li>Backbone : Convnext v2 Large</li>\n<li>Loss : FocalLoss</li>\n</ul>\n<p><strong>Augmentation</strong></p>\n<ul>\n<li>Albumentations<ul>\n<li>Distortion worked very well for me.</li></ul></li>\n</ul>\n<pre><code>            (\n                =,\n                transforms=[\n                    (=, p=),\n                    (=, p=),\n                    (=, p=),\n                            ], p=),\n</code></pre>\n<ul>\n<li>mmdet augmentation<ul>\n<li>AutoAugment, MixUp, Mosaic, RandomErasing</li></ul></li>\n</ul>\n<p><strong>Training process</strong></p>\n<p>First, weighted segment fusion was performed on the output values ​​from the 4 models made of 4 types of data sets to create pseudo labels for Dataset3.</p>\n<p>I did WSF by referring to the guide below -&gt; <a href=\"https://www.kaggle.com/code/mistag/sartorius-tta-with-weighted-segments-fusion\" target=\"_blank\">https://www.kaggle.com/code/mistag/sartorius-tta-with-weighted-segments-fusion</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F74ad73b605892f8de8d3bc401dae6d0a%2Fmake%20pseudo%20labels.png?generation=1690855159125845&amp;alt=media\" alt=\"\"></p>\n<p>second. The model was pre-trained using the pseudo label of Dataset 3 and the human label of Dataset 2.</p>\n<p>Finally, the model was fine-tuned using only Dataset1 data and submitted.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F49ef424a4d828f38c7e44d374a2021c2%2Ffinetune.png?generation=1690855729462161&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2368093",
      "postDate": "08/01/2023 02:23:32",
      "content": "<p>I am so happy to have won my first gold medal and even a solo medal in this competition. Also thanks to the  OpenMMLab, SenseTime and The Chinese University of Hong Kong for making a really great library mmdetection.</p>\n<p><strong>Summary</strong><br>\nI used the pseudo labels of Dataset3 to train the Cascade Mask RCNN + Convnext v2 large.</p>\n<p><strong>Model</strong></p>\n<ul>\n<li>Detector : Cascade Mask RCNN </li>\n<li>Backbone : Convnext v2 Large</li>\n<li>Loss : FocalLoss</li>\n</ul>\n<p><strong>Augmentation</strong></p>\n<ul>\n<li>Albumentations<ul>\n<li>Distortion worked very well for me.</li></ul></li>\n</ul>\n<pre><code>            (\n                =,\n                transforms=[\n                    (=, p=),\n                    (=, p=),\n                    (=, p=),\n                            ], p=),\n</code></pre>\n<ul>\n<li>mmdet augmentation<ul>\n<li>AutoAugment, MixUp, Mosaic, RandomErasing</li></ul></li>\n</ul>\n<p><strong>Training process</strong></p>\n<p>First, weighted segment fusion was performed on the output values ​​from the 4 models made of 4 types of data sets to create pseudo labels for Dataset3.</p>\n<p>I did WSF by referring to the guide below -&gt; <a href=\"https://www.kaggle.com/code/mistag/sartorius-tta-with-weighted-segments-fusion\" target=\"_blank\">https://www.kaggle.com/code/mistag/sartorius-tta-with-weighted-segments-fusion</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F74ad73b605892f8de8d3bc401dae6d0a%2Fmake%20pseudo%20labels.png?generation=1690855159125845&amp;alt=media\" alt=\"\"></p>\n<p>second. The model was pre-trained using the pseudo label of Dataset 3 and the human label of Dataset 2.</p>\n<p>Finally, the model was fine-tuned using only Dataset1 data and submitted.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F49ef424a4d828f38c7e44d374a2021c2%2Ffinetune.png?generation=1690855729462161&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I am so happy to have won my first gold medal and even a solo medal in this competition. Also thanks to the  OpenMMLab, SenseTime and The Chinese University of Hong Kong for making a really great library mmdetection.\n\n**Summary**\nI used the pseudo labels of Dataset3 to train the Cascade Mask RCNN + Convnext v2 large.\n\n**Model**\n- Detector : Cascade Mask RCNN \n- Backbone : Convnext v2 Large\n- Loss : FocalLoss\n\n**Augmentation**\n- Albumentations\n   - Distortion worked very well for me.\n```python\n            dict(\n                type='OneOf',\n                transforms=[\n                    dict(type='OpticalDistortion', p=0.3),\n                    dict(type='GridDistortion', p=0.3),\n                    dict(type='ElasticTransform', p=0.1),\n                            ], p=0.5),\n```\n\n- mmdet augmentation\n   - AutoAugment, MixUp, Mosaic, RandomErasing\n\n**Training process**\n\nFirst, weighted segment fusion was performed on the output values ​​from the 4 models made of 4 types of data sets to create pseudo labels for Dataset3.\n\nI did WSF by referring to the guide below -> https://www.kaggle.com/code/mistag/sartorius-tta-with-weighted-segments-fusion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F74ad73b605892f8de8d3bc401dae6d0a%2Fmake%20pseudo%20labels.png?generation=1690855159125845&alt=media)\n\nsecond. The model was pre-trained using the pseudo label of Dataset 3 and the human label of Dataset 2.\n\nFinally, the model was fine-tuned using only Dataset1 data and submitted.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F49ef424a4d828f38c7e44d374a2021c2%2Ffinetune.png?generation=1690855729462161&alt=media)",
      "votes": null
    },
    {
      "id": "2368137",
      "postDate": "08/01/2023 02:59:34",
      "content": "<p>Hi, what is the concept of pseudo label here?</p>\n<p>Thanks in advance.</p>",
      "rawMarkdown": "Hi, what is the concept of pseudo label here?\n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "2368145",
      "postDate": "08/01/2023 03:08:12",
      "content": "<p>Hi What is the meaning of the concept here?</p>\n<p>I thought that there would be differences for each WSI, so I created 4 folds for each WSI. The predicted values ​​of the four models thus created were WSFed and used as pseudo labels. It is.</p>",
      "rawMarkdown": "Hi What is the meaning of the concept here?\n\nI thought that there would be differences for each WSI, so I created 4 folds for each WSI. The predicted values ​​of the four models thus created were WSFed and used as pseudo labels. It is.",
      "votes": null
    },
    {
      "id": "2368149",
      "postDate": "08/01/2023 03:11:43",
      "content": "<p>Congrats on solo gold, but I want to correct you that the mmdetection is made by OpenMMLab which is created by SenseTime and The Chinese University of Hong Kong😌</p>",
      "rawMarkdown": "Congrats on solo gold, but I want to correct you that the mmdetection is made by OpenMMLab which is created by SenseTime and The Chinese University of Hong Kong😌",
      "votes": null
    },
    {
      "id": "2368151",
      "postDate": "08/01/2023 03:14:04",
      "content": "<p>Congrats for the solo gold medal!<br>\nI have 2 questions:</p>\n<ol>\n<li>What is your configs in the final fine tuning step ? Do you still keep the same config from the previous step or use something else ?</li>\n<li>In the final training pre-train model and fine tuning step, what set do you use for valid ?</li>\n</ol>",
      "rawMarkdown": "Congrats for the solo gold medal!\nI have 2 questions:\n1. What is your configs in the final fine tuning step ? Do you still keep the same config from the previous step or use something else ?\n2. In the final training pre-train model and fine tuning step, what set do you use for valid ?",
      "votes": null
    },
    {
      "id": "2368153",
      "postDate": "08/01/2023 03:15:37",
      "content": "<p>Sorry My mistake, Thanks for the advice.</p>",
      "rawMarkdown": "Sorry My mistake, Thanks for the advice.",
      "votes": null
    },
    {
      "id": "2368155",
      "postDate": "08/01/2023 03:21:34",
      "content": "<ol>\n<li>Yes, In the pre-training step, I select the best epoch model pth and resume fine-tuning in the same configs</li>\n<li>First, find the epoch with the best valid loss using 20% ​​of Dataset1 as validatoion. After that, fine-tuning was performed on all Dataset1 data under the same conditions.</li>\n</ol>",
      "rawMarkdown": "1. Yes, In the pre-training step, I select the best epoch model pth and resume fine-tuning in the same configs\n2. First, find the epoch with the best valid loss using 20% ​​of Dataset1 as validatoion. After that, fine-tuning was performed on all Dataset1 data under the same conditions.",
      "votes": null
    },
    {
      "id": "2368177",
      "postDate": "08/01/2023 03:49:50",
      "content": "<ol>\n<li>By \"best epoch model pth\" u mean best model of training pre-train model (dataset 2 of wsi 1,2 and dataset 3) right ?</li>\n<li>Do you freeze any layer in the fine tuning step ?</li>\n</ol>",
      "rawMarkdown": "1. By \"best epoch model pth\" u mean best model of training pre-train model (dataset 2 of wsi 1,2 and dataset 3) right ?\n2. Do you freeze any layer in the fine tuning step ?",
      "votes": null
    },
    {
      "id": "2368183",
      "postDate": "08/01/2023 03:54:30",
      "content": "<ol>\n<li>yes, \"best epoch model pth\" with Train (dataset 2 of wsi 1,2 and dataset 3) valid (only dataset 1)</li>\n<li>no i didn't freeze layer</li>\n</ol>\n<p>Sorry for not being able to explain clearly.</p>",
      "rawMarkdown": "1. yes, \"best epoch model pth\" with Train (dataset 2 of wsi 1,2 and dataset 3) valid (only dataset 1)\n2. no i didn't freeze layer\n\nSorry for not being able to explain clearly.",
      "votes": null
    },
    {
      "id": "2368189",
      "postDate": "08/01/2023 03:59:52",
      "content": "<p>Thank you so much for your response!<br>\nThis is very similar with my first approach for this competition. But i don't use this because i think this is data leakage. What do you think about this ?</p>",
      "rawMarkdown": "Thank you so much for your response!\nThis is very similar with my first approach for this competition. But i don't use this because i think this is data leakage. What do you think about this ?",
      "votes": null
    },
    {
      "id": "2368220",
      "postDate": "08/01/2023 04:32:00",
      "content": "<p>Yes, I thought it was difficult to trust the validation score because Dataset1 was used together when pseudo labeling was performed.</p>\n<p>Therefore, this idea was not applied right away, but was applied after finding the best hyperparameters and epochs through CV until just before the end of the competition.</p>",
      "rawMarkdown": "Yes, I thought it was difficult to trust the validation score because Dataset1 was used together when pseudo labeling was performed.\n\nTherefore, this idea was not applied right away, but was applied after finding the best hyperparameters and epochs through CV until just before the end of the competition.",
      "votes": null
    },
    {
      "id": "2368250",
      "postDate": "08/01/2023 04:49:11",
      "content": "<p>Congrats on the medal <a href=\"https://www.kaggle.com/devchopin\" target=\"_blank\">@devchopin</a> !</p>\n<p>Even I am a university student (just completed my first year). I have not taken any formal courses yet but I try to read up on articles and am currently reading hands on ML with sci-kit learn and TensorFlow. How exactly do you go about preparing for competitions and how much time do you spend?</p>",
      "rawMarkdown": "Congrats on the medal @devchopin !\n\nEven I am a university student (just completed my first year). I have not taken any formal courses yet but I try to read up on articles and am currently reading hands on ML with sci-kit learn and TensorFlow. How exactly do you go about preparing for competitions and how much time do you spend?",
      "votes": null
    },
    {
      "id": "2368272",
      "postDate": "08/01/2023 05:04:48",
      "content": "<p>Since I am also a beginner, I am cautious about giving advice, but in my case, participating in various competitions was helpful. I also read a lot of META research papers.</p>",
      "rawMarkdown": "Since I am also a beginner, I am cautious about giving advice, but in my case, participating in various competitions was helpful. I also read a lot of META research papers.",
      "votes": null
    },
    {
      "id": "2368398",
      "postDate": "08/01/2023 06:37:56",
      "content": "<p>Thank you </p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "2368764",
      "postDate": "08/01/2023 11:15:19",
      "content": "<p>Congrads on your solo gold! I think this pseudo labeling method is insightful. I want to know a little detail about it. When you fine tuning the pretrained model, what validation data you use for it?(you say \"under the same conditions\", you mean you use 20% dataset1 to validate for fine tuning stage too?)</p>",
      "rawMarkdown": "Congrads on your solo gold! I think this pseudo labeling method is insightful. I want to know a little detail about it. When you fine tuning the pretrained model, what validation data you use for it?(you say \"under the same conditions\", you mean you use 20% dataset1 to validate for fine tuning stage too?)",
      "votes": null
    },
    {
      "id": "2369631",
      "postDate": "08/01/2023 22:12:22",
      "content": "<p>The pseudo labels were blood vessel, glomerulus and unsure?.</p>\n<p>Thanks in advance.</p>",
      "rawMarkdown": "The pseudo labels were blood vessel, glomerulus and unsure?.\n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "2369673",
      "postDate": "08/02/2023 00:02:08",
      "content": "<p>Only blood vessel, Thank you</p>",
      "rawMarkdown": "Only blood vessel, Thank you",
      "votes": null
    },
    {
      "id": "2369676",
      "postDate": "08/02/2023 00:11:19",
      "content": "<p>Hello, thank you for your kind words.</p>\n<p>There are a total of 4 steps.</p>\n<ol>\n<li><p>Create a pre-trained model using Dataset2 + Dataset3 data. (24 epoch was the best result for me here.)</p></li>\n<li><p>Train Dataset1 data from 24 epoch(from pre-train model). Here, it was verified using train (80%) and validation (20%). (This step is to find out which epoch is best.)</p></li>\n<li><p>I found that training up to 28 epochs in step 2 is the best mAP60.</p></li>\n<li><p>Finally, submit the model trained on 28 epochs of 100% of Dataset1 data.</p></li>\n</ol>\n<p>If there is something I don't understand, I will answer again. thank you</p>",
      "rawMarkdown": "Hello, thank you for your kind words.\n\nThere are a total of 4 steps.\n\n1. Create a pre-trained model using Dataset2 + Dataset3 data. (24 epoch was the best result for me here.)\n\n2. Train Dataset1 data from 24 epoch(from pre-train model). Here, it was verified using train (80%) and validation (20%). (This step is to find out which epoch is best.)\n\n3. I found that training up to 28 epochs in step 2 is the best mAP60.\n\n4. Finally, submit the model trained on 28 epochs of 100% of Dataset1 data.\n\nIf there is something I don't understand, I will answer again. thank you",
      "votes": null
    },
    {
      "id": "2369760",
      "postDate": "08/02/2023 02:30:15",
      "content": "<p>I understand, thank you</p>",
      "rawMarkdown": "I understand, thank you",
      "votes": null
    },
    {
      "id": "2369765",
      "postDate": "08/02/2023 02:40:18",
      "content": "<p>Congratulations on your solo gold! </p>\n<p>I have a question for you. I assume that the model used for pseudo-labeling contains a fair amount of low-score false positives. After applying the WSF, did you pseudo-label all the predicted masks directly, or did you apply any post-processing to the masks before pseudo-labeling them? (e.g., excluding small instances, filtering by probability values)</p>",
      "rawMarkdown": "Congratulations on your solo gold! \n\nI have a question for you. I assume that the model used for pseudo-labeling contains a fair amount of low-score false positives. After applying the WSF, did you pseudo-label all the predicted masks directly, or did you apply any post-processing to the masks before pseudo-labeling them? (e.g., excluding small instances, filtering by probability values)",
      "votes": null
    },
    {
      "id": "2369769",
      "postDate": "08/02/2023 02:46:14",
      "content": "<p>there is two step</p>\n<ol>\n<li>After WSF, only masks with a confidence score of 0.5 or higher were used.</li>\n<li>And I filtered the mask of area smaller than 100.</li>\n</ol>",
      "rawMarkdown": "there is two step\n\n1. After WSF, only masks with a confidence score of 0.5 or higher were used.\n2. And I filtered the mask of area smaller than 100.",
      "votes": null
    },
    {
      "id": "2369805",
      "postDate": "08/02/2023 03:28:27",
      "content": "<p>Thank you for your response!<br>\n I understand your explanation.</p>",
      "rawMarkdown": "Thank you for your response!\n I understand your explanation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2368137,
      "author_name": "pablolarrosa",
      "author_url": "",
      "post_date": "08/01/2023 02:59:34",
      "content": "<p>Hi, what is the concept of pseudo label here?</p>\n<p>Thanks in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2368145,
          "author_name": "devchopin",
          "author_url": "",
          "post_date": "08/01/2023 03:08:12",
          "content": "<p>Hi What is the meaning of the concept here?</p>\n<p>I thought that there would be differences for each WSI, so I created 4 folds for each WSI. The predicted values ​​of the four models thus created were WSFed and used as pseudo labels. It is.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2369631,
              "author_name": "pablolarrosa",
              "author_url": "",
              "post_date": "08/01/2023 22:12:22",
              "content": "<p>The pseudo labels were blood vessel, glomerulus and unsure?.</p>\n<p>Thanks in advance.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2369673,
                  "author_name": "devchopin",
                  "author_url": "",
                  "post_date": "08/02/2023 00:02:08",
                  "content": "<p>Only blood vessel, Thank you</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2369765,
              "author_name": "yuyagi",
              "author_url": "",
              "post_date": "08/02/2023 02:40:18",
              "content": "<p>Congratulations on your solo gold! </p>\n<p>I have a question for you. I assume that the model used for pseudo-labeling contains a fair amount of low-score false positives. After applying the WSF, did you pseudo-label all the predicted masks directly, or did you apply any post-processing to the masks before pseudo-labeling them? (e.g., excluding small instances, filtering by probability values)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2369769,
                  "author_name": "devchopin",
                  "author_url": "",
                  "post_date": "08/02/2023 02:46:14",
                  "content": "<p>there is two step</p>\n<ol>\n<li>After WSF, only masks with a confidence score of 0.5 or higher were used.</li>\n<li>And I filtered the mask of area smaller than 100.</li>\n</ol>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2369805,
                      "author_name": "yuyagi",
                      "author_url": "",
                      "post_date": "08/02/2023 03:28:27",
                      "content": "<p>Thank you for your response!<br>\n I understand your explanation.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2368149,
      "author_name": "traptinblur",
      "author_url": "",
      "post_date": "08/01/2023 03:11:43",
      "content": "<p>Congrats on solo gold, but I want to correct you that the mmdetection is made by OpenMMLab which is created by SenseTime and The Chinese University of Hong Kong😌</p>",
      "votes": null,
      "replies": [
        {
          "id": 2368153,
          "author_name": "devchopin",
          "author_url": "",
          "post_date": "08/01/2023 03:15:37",
          "content": "<p>Sorry My mistake, Thanks for the advice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2368151,
      "author_name": "researchbntz",
      "author_url": "",
      "post_date": "08/01/2023 03:14:04",
      "content": "<p>Congrats for the solo gold medal!<br>\nI have 2 questions:</p>\n<ol>\n<li>What is your configs in the final fine tuning step ? Do you still keep the same config from the previous step or use something else ?</li>\n<li>In the final training pre-train model and fine tuning step, what set do you use for valid ?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2368155,
          "author_name": "devchopin",
          "author_url": "",
          "post_date": "08/01/2023 03:21:34",
          "content": "<ol>\n<li>Yes, In the pre-training step, I select the best epoch model pth and resume fine-tuning in the same configs</li>\n<li>First, find the epoch with the best valid loss using 20% ​​of Dataset1 as validatoion. After that, fine-tuning was performed on all Dataset1 data under the same conditions.</li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2368177,
              "author_name": "researchbntz",
              "author_url": "",
              "post_date": "08/01/2023 03:49:50",
              "content": "<ol>\n<li>By \"best epoch model pth\" u mean best model of training pre-train model (dataset 2 of wsi 1,2 and dataset 3) right ?</li>\n<li>Do you freeze any layer in the fine tuning step ?</li>\n</ol>",
              "votes": null,
              "replies": [
                {
                  "id": 2368183,
                  "author_name": "devchopin",
                  "author_url": "",
                  "post_date": "08/01/2023 03:54:30",
                  "content": "<ol>\n<li>yes, \"best epoch model pth\" with Train (dataset 2 of wsi 1,2 and dataset 3) valid (only dataset 1)</li>\n<li>no i didn't freeze layer</li>\n</ol>\n<p>Sorry for not being able to explain clearly.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2368189,
                      "author_name": "researchbntz",
                      "author_url": "",
                      "post_date": "08/01/2023 03:59:52",
                      "content": "<p>Thank you so much for your response!<br>\nThis is very similar with my first approach for this competition. But i don't use this because i think this is data leakage. What do you think about this ?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2368220,
                          "author_name": "devchopin",
                          "author_url": "",
                          "post_date": "08/01/2023 04:32:00",
                          "content": "<p>Yes, I thought it was difficult to trust the validation score because Dataset1 was used together when pseudo labeling was performed.</p>\n<p>Therefore, this idea was not applied right away, but was applied after finding the best hyperparameters and epochs through CV until just before the end of the competition.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2368398,
                              "author_name": "researchbntz",
                              "author_url": "",
                              "post_date": "08/01/2023 06:37:56",
                              "content": "<p>Thank you </p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2368764,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "08/01/2023 11:15:19",
              "content": "<p>Congrads on your solo gold! I think this pseudo labeling method is insightful. I want to know a little detail about it. When you fine tuning the pretrained model, what validation data you use for it?(you say \"under the same conditions\", you mean you use 20% dataset1 to validate for fine tuning stage too?)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2369676,
                  "author_name": "devchopin",
                  "author_url": "",
                  "post_date": "08/02/2023 00:11:19",
                  "content": "<p>Hello, thank you for your kind words.</p>\n<p>There are a total of 4 steps.</p>\n<ol>\n<li><p>Create a pre-trained model using Dataset2 + Dataset3 data. (24 epoch was the best result for me here.)</p></li>\n<li><p>Train Dataset1 data from 24 epoch(from pre-train model). Here, it was verified using train (80%) and validation (20%). (This step is to find out which epoch is best.)</p></li>\n<li><p>I found that training up to 28 epochs in step 2 is the best mAP60.</p></li>\n<li><p>Finally, submit the model trained on 28 epochs of 100% of Dataset1 data.</p></li>\n</ol>\n<p>If there is something I don't understand, I will answer again. thank you</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2369760,
                      "author_name": "m1dsolo",
                      "author_url": "",
                      "post_date": "08/02/2023 02:30:15",
                      "content": "<p>I understand, thank you</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2368250,
      "author_name": "arnavmodi",
      "author_url": "",
      "post_date": "08/01/2023 04:49:11",
      "content": "<p>Congrats on the medal <a href=\"https://www.kaggle.com/devchopin\" target=\"_blank\">@devchopin</a> !</p>\n<p>Even I am a university student (just completed my first year). I have not taken any formal courses yet but I try to read up on articles and am currently reading hands on ML with sci-kit learn and TensorFlow. How exactly do you go about preparing for competitions and how much time do you spend?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2368272,
          "author_name": "devchopin",
          "author_url": "",
          "post_date": "08/01/2023 05:04:48",
          "content": "<p>Since I am also a beginner, I am cautious about giving advice, but in my case, participating in various competitions was helpful. I also read a lot of META research papers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2368093": "I am so happy to have won my first gold medal and even a solo medal in this competition. Also thanks to the  OpenMMLab, SenseTime and The Chinese University of Hong Kong for making a really great library mmdetection.\n\n**Summary**\nI used the pseudo labels of Dataset3 to train the Cascade Mask RCNN + Convnext v2 large.\n\n**Model**\n- Detector : Cascade Mask RCNN \n- Backbone : Convnext v2 Large\n- Loss : FocalLoss\n\n**Augmentation**\n- Albumentations\n   - Distortion worked very well for me.\n```python\n            dict(\n                type='OneOf',\n                transforms=[\n                    dict(type='OpticalDistortion', p=0.3),\n                    dict(type='GridDistortion', p=0.3),\n                    dict(type='ElasticTransform', p=0.1),\n                            ], p=0.5),\n```\n\n- mmdet augmentation\n   - AutoAugment, MixUp, Mosaic, RandomErasing\n\n**Training process**\n\nFirst, weighted segment fusion was performed on the output values ​​from the 4 models made of 4 types of data sets to create pseudo labels for Dataset3.\n\nI did WSF by referring to the guide below -> https://www.kaggle.com/code/mistag/sartorius-tta-with-weighted-segments-fusion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F74ad73b605892f8de8d3bc401dae6d0a%2Fmake%20pseudo%20labels.png?generation=1690855159125845&alt=media)\n\nsecond. The model was pre-trained using the pseudo label of Dataset 3 and the human label of Dataset 2.\n\nFinally, the model was fine-tuned using only Dataset1 data and submitted.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9334354%2F49ef424a4d828f38c7e44d374a2021c2%2Ffinetune.png?generation=1690855729462161&alt=media)",
    "2368137": "Hi, what is the concept of pseudo label here?\n\nThanks in advance.",
    "2368145": "Hi What is the meaning of the concept here?\n\nI thought that there would be differences for each WSI, so I created 4 folds for each WSI. The predicted values ​​of the four models thus created were WSFed and used as pseudo labels. It is.",
    "2368149": "Congrats on solo gold, but I want to correct you that the mmdetection is made by OpenMMLab which is created by SenseTime and The Chinese University of Hong Kong😌",
    "2368151": "Congrats for the solo gold medal!\nI have 2 questions:\n1. What is your configs in the final fine tuning step ? Do you still keep the same config from the previous step or use something else ?\n2. In the final training pre-train model and fine tuning step, what set do you use for valid ?",
    "2368153": "Sorry My mistake, Thanks for the advice.",
    "2368155": "1. Yes, In the pre-training step, I select the best epoch model pth and resume fine-tuning in the same configs\n2. First, find the epoch with the best valid loss using 20% ​​of Dataset1 as validatoion. After that, fine-tuning was performed on all Dataset1 data under the same conditions.",
    "2368177": "1. By \"best epoch model pth\" u mean best model of training pre-train model (dataset 2 of wsi 1,2 and dataset 3) right ?\n2. Do you freeze any layer in the fine tuning step ?",
    "2368183": "1. yes, \"best epoch model pth\" with Train (dataset 2 of wsi 1,2 and dataset 3) valid (only dataset 1)\n2. no i didn't freeze layer\n\nSorry for not being able to explain clearly.",
    "2368189": "Thank you so much for your response!\nThis is very similar with my first approach for this competition. But i don't use this because i think this is data leakage. What do you think about this ?",
    "2368220": "Yes, I thought it was difficult to trust the validation score because Dataset1 was used together when pseudo labeling was performed.\n\nTherefore, this idea was not applied right away, but was applied after finding the best hyperparameters and epochs through CV until just before the end of the competition.",
    "2368250": "Congrats on the medal @devchopin !\n\nEven I am a university student (just completed my first year). I have not taken any formal courses yet but I try to read up on articles and am currently reading hands on ML with sci-kit learn and TensorFlow. How exactly do you go about preparing for competitions and how much time do you spend?",
    "2368272": "Since I am also a beginner, I am cautious about giving advice, but in my case, participating in various competitions was helpful. I also read a lot of META research papers.",
    "2368398": "Thank you",
    "2368764": "Congrads on your solo gold! I think this pseudo labeling method is insightful. I want to know a little detail about it. When you fine tuning the pretrained model, what validation data you use for it?(you say \"under the same conditions\", you mean you use 20% dataset1 to validate for fine tuning stage too?)",
    "2369631": "The pseudo labels were blood vessel, glomerulus and unsure?.\n\nThanks in advance.",
    "2369673": "Only blood vessel, Thank you",
    "2369676": "Hello, thank you for your kind words.\n\nThere are a total of 4 steps.\n\n1. Create a pre-trained model using Dataset2 + Dataset3 data. (24 epoch was the best result for me here.)\n\n2. Train Dataset1 data from 24 epoch(from pre-train model). Here, it was verified using train (80%) and validation (20%). (This step is to find out which epoch is best.)\n\n3. I found that training up to 28 epochs in step 2 is the best mAP60.\n\n4. Finally, submit the model trained on 28 epochs of 100% of Dataset1 data.\n\nIf there is something I don't understand, I will answer again. thank you",
    "2369760": "I understand, thank you",
    "2369765": "Congratulations on your solo gold! \n\nI have a question for you. I assume that the model used for pseudo-labeling contains a fair amount of low-score false positives. After applying the WSF, did you pseudo-label all the predicted masks directly, or did you apply any post-processing to the masks before pseudo-labeling them? (e.g., excluding small instances, filtering by probability values)",
    "2369769": "there is two step\n\n1. After WSF, only masks with a confidence score of 0.5 or higher were used.\n2. And I filtered the mask of area smaller than 100.",
    "2369805": "Thank you for your response!\n I understand your explanation."
  },
  "source": "meta"
}