{
  "id": 430062,
  "title": "14th Place solution: Yolov8",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/writeups/vts-ai-14th-place-solution-yolov8",
  "author_name": "",
  "post_date": "2023-08-08T06:36:30.440Z",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congratuations to all the winner and thank the organizers for hosting such an interesting competition!</p>\n<ul>\n<li>My solution is a simple ensemble of Yolov8l and Yolov8m trained on image size 640 and 1024</li>\n<li>All of my submission using Dilation with Kernel size =  3 and Iters = 1 because on my local CV, dilation don't hurt (and sometimes improve) <strong>dataset1</strong> but decrease mAP on <strong>dataset2</strong> significantly.</li>\n</ul>\n<p>Things worked for me:</p>\n<ul>\n<li>Yolov8l-seg and Yolov8m-seg</li>\n<li>Copy-paste and Mixup augmentation (0.5 probability)</li>\n<li>Balance sampling between dataset1, dataset2 and unlabeled data (pseudo label) </li>\n<li>Pseudo-labeling on Unlabeled data</li>\n<li>Larger image size (Increase CV but not LB/PB 😭)</li>\n<li>Pseudo labeling on Test data (i.e: Retrain during submission)</li>\n<li>Ensemble using Weighted mask fusion from 8th place Sartorius competition <a href=\"https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation/discussion/297998\" target=\"_blank\">https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation/discussion/297998</a></li>\n</ul>\n<p>Things not worked for me:</p>\n<ul>\n<li>Pseudo label more than one round</li>\n<li>Ensemble on different fold</li>\n</ul>\n<h1>1. Validation strategy</h1>\n<ul>\n<li>I split the data into train/valid by WSI (4 folds), during the competition, I mostly use WSI 1 and 2 as validation data and models were trained on WSI 1, 3, 4 because WSI1,2 also contains <code>Dataset1</code></li>\n<li>I belive because of the split, my solution can survive when Private leaderboard released. My Validation scores on WSI 2 are nearly correlated to PB/LB</li>\n</ul>\n<h1>2. Modeling</h1>\n<ul>\n<li>Step 1: Train 8 models Yolov8l and Yolov8m with image size of 640 on 4 folds for 100 epochs</li>\n<li>Step 2: Use WMF (Weighted mask fusion) to generate pseudo label on unlabeled data</li>\n<li>Step 3: Retrain models on Labeled data + Pseudo data for 70 epochs, Image size are 640 and 1024</li>\n<li>Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.</li>\n</ul>\n<p>All models are trained on the following hyper parameters, I only show the param that I change compared to default values:</p>\n<table>\n<thead>\n<tr>\n<th>Param</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mixup</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>copy-paste</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>flipud</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>degrees</td>\n<td>45</td>\n</tr>\n<tr>\n<td>lr0</td>\n<td>0.001</td>\n</tr>\n</tbody>\n</table>\n<p>I also modified the dataloader of Yolov8 to sampling uniformly <code>dataset1</code>, <code>dataset2</code> and <code>pseudo data</code>.</p>\n<h1>3. The effect of pseudo labeling</h1>\n<ul>\n<li>I used pseudo labeling two times, one on local with unlabeled data (<strong>Stage1</strong>) and one during submission on hidden test data (<strong>Stage2</strong>) as mentioned in <strong>2.</strong></li>\n<li>Pseudo labeling <strong>stage1</strong> improve my mask mAP50 on WSI 1 and 2 about ~ 0.04</li>\n<li><strong>Stage2</strong> Pseudo labeling improved both my LB/PB about 0.005, I think if base models are better, it could improve more.</li>\n</ul>\n<p>Compare of Stage2 pseudo labeling</p>\n<table>\n<thead>\n<tr>\n<th>Sub</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ENS1 + Stage2 pseudo</td>\n<td>0.537</td>\n<td>0.505</td>\n</tr>\n<tr>\n<td>ENS1</td>\n<td>0.532</td>\n<td>0.500</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Choice of submission</h1>\n<ul>\n<li>I selected my two last submission on two set of my best CV models on different fold, all models are trained with pseudo labeling <strong>Stage1</strong>. Called ENS1 and ENS2.</li>\n<li>My best private score is 0.544 (LB 0.551) which is ensemble of 8 models <code>Fold 2</code></li>\n</ul>\n<p>Submission:</p>\n<table>\n<thead>\n<tr>\n<th>Sub</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ENS1 + Stage2 pseudo</td>\n<td>0.537</td>\n<td>0.505</td>\n</tr>\n<tr>\n<td>ENS2 + Stage2 pseudo</td>\n<td>0.549</td>\n<td>0.532</td>\n</tr>\n</tbody>\n</table>\n<h1>5. Why shake-up (Just my guess)</h1>\n<ul>\n<li>Clearly because the test data is a completly different WSI, this is why my validation strategy is to split the data by WSI</li>\n<li>I think the private test data is small and very sensitive (only one WSI), I think It's not large enough to choose the best model for real-world data (I understand that the annotation process is very costly). I observed my local vaildation score is oscillated during training too.</li>\n<li>Many people reported that dilation also cause the shake-up, may be my dilation iters is only 1, so the effect is minimized</li>\n</ul>",
  "messages": [
    {
      "id": "2379386",
      "postDate": "08/08/2023 06:25:04",
      "content": "<p>Congratuations to all the winner and thank the organizers for hosting such an interesting competition!</p>\n<ul>\n<li>My solution is a simple ensemble of Yolov8l and Yolov8m trained on image size 640 and 1024</li>\n<li>All of my submission using Dilation with Kernel size =  3 and Iters = 1 because on my local CV, dilation don't hurt (and sometimes improve) <strong>dataset1</strong> but decrease mAP on <strong>dataset2</strong> significantly.</li>\n</ul>\n<p>Things worked for me:</p>\n<ul>\n<li>Yolov8l-seg and Yolov8m-seg</li>\n<li>Copy-paste and Mixup augmentation (0.5 probability)</li>\n<li>Balance sampling between dataset1, dataset2 and unlabeled data (pseudo label) </li>\n<li>Pseudo-labeling on Unlabeled data</li>\n<li>Larger image size (Increase CV but not LB/PB 😭)</li>\n<li>Pseudo labeling on Test data (i.e: Retrain during submission)</li>\n<li>Ensemble using Weighted mask fusion from 8th place Sartorius competition <a href=\"https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation/discussion/297998\" target=\"_blank\">https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation/discussion/297998</a></li>\n</ul>\n<p>Things not worked for me:</p>\n<ul>\n<li>Pseudo label more than one round</li>\n<li>Ensemble on different fold</li>\n</ul>\n<h1>1. Validation strategy</h1>\n<ul>\n<li>I split the data into train/valid by WSI (4 folds), during the competition, I mostly use WSI 1 and 2 as validation data and models were trained on WSI 1, 3, 4 because WSI1,2 also contains <code>Dataset1</code></li>\n<li>I belive because of the split, my solution can survive when Private leaderboard released. My Validation scores on WSI 2 are nearly correlated to PB/LB</li>\n</ul>\n<h1>2. Modeling</h1>\n<ul>\n<li>Step 1: Train 8 models Yolov8l and Yolov8m with image size of 640 on 4 folds for 100 epochs</li>\n<li>Step 2: Use WMF (Weighted mask fusion) to generate pseudo label on unlabeled data</li>\n<li>Step 3: Retrain models on Labeled data + Pseudo data for 70 epochs, Image size are 640 and 1024</li>\n<li>Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.</li>\n</ul>\n<p>All models are trained on the following hyper parameters, I only show the param that I change compared to default values:</p>\n<table>\n<thead>\n<tr>\n<th>Param</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mixup</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>copy-paste</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>flipud</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>degrees</td>\n<td>45</td>\n</tr>\n<tr>\n<td>lr0</td>\n<td>0.001</td>\n</tr>\n</tbody>\n</table>\n<p>I also modified the dataloader of Yolov8 to sampling uniformly <code>dataset1</code>, <code>dataset2</code> and <code>pseudo data</code>.</p>\n<h1>3. The effect of pseudo labeling</h1>\n<ul>\n<li>I used pseudo labeling two times, one on local with unlabeled data (<strong>Stage1</strong>) and one during submission on hidden test data (<strong>Stage2</strong>) as mentioned in <strong>2.</strong></li>\n<li>Pseudo labeling <strong>stage1</strong> improve my mask mAP50 on WSI 1 and 2 about ~ 0.04</li>\n<li><strong>Stage2</strong> Pseudo labeling improved both my LB/PB about 0.005, I think if base models are better, it could improve more.</li>\n</ul>\n<p>Compare of Stage2 pseudo labeling</p>\n<table>\n<thead>\n<tr>\n<th>Sub</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ENS1 + Stage2 pseudo</td>\n<td>0.537</td>\n<td>0.505</td>\n</tr>\n<tr>\n<td>ENS1</td>\n<td>0.532</td>\n<td>0.500</td>\n</tr>\n</tbody>\n</table>\n<h1>4. Choice of submission</h1>\n<ul>\n<li>I selected my two last submission on two set of my best CV models on different fold, all models are trained with pseudo labeling <strong>Stage1</strong>. Called ENS1 and ENS2.</li>\n<li>My best private score is 0.544 (LB 0.551) which is ensemble of 8 models <code>Fold 2</code></li>\n</ul>\n<p>Submission:</p>\n<table>\n<thead>\n<tr>\n<th>Sub</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ENS1 + Stage2 pseudo</td>\n<td>0.537</td>\n<td>0.505</td>\n</tr>\n<tr>\n<td>ENS2 + Stage2 pseudo</td>\n<td>0.549</td>\n<td>0.532</td>\n</tr>\n</tbody>\n</table>\n<h1>5. Why shake-up (Just my guess)</h1>\n<ul>\n<li>Clearly because the test data is a completly different WSI, this is why my validation strategy is to split the data by WSI</li>\n<li>I think the private test data is small and very sensitive (only one WSI), I think It's not large enough to choose the best model for real-world data (I understand that the annotation process is very costly). I observed my local vaildation score is oscillated during training too.</li>\n<li>Many people reported that dilation also cause the shake-up, may be my dilation iters is only 1, so the effect is minimized</li>\n</ul>",
      "rawMarkdown": "Congratuations to all the winner and thank the organizers for hosting such an interesting competition!\n\n\n- My solution is a simple ensemble of Yolov8l and Yolov8m trained on image size 640 and 1024\n- All of my submission using Dilation with Kernel size =  3 and Iters = 1 because on my local CV, dilation don't hurt (and sometimes improve) **dataset1** but decrease mAP on **dataset2** significantly.\n\n\nThings worked for me:\n- Yolov8l-seg and Yolov8m-seg\n- Copy-paste and Mixup augmentation (0.5 probability)\n- Balance sampling between dataset1, dataset2 and unlabeled data (pseudo label) \n- Pseudo-labeling on Unlabeled data\n- Larger image size (Increase CV but not LB/PB 😭)\n- Pseudo labeling on Test data (i.e: Retrain during submission)\n- Ensemble using Weighted mask fusion from 8th place Sartorius competition https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation/discussion/297998\n\nThings not worked for me:\n- Pseudo label more than one round\n- Ensemble on different fold\n\n\n# 1. Validation strategy\n\n- I split the data into train/valid by WSI (4 folds), during the competition, I mostly use WSI 1 and 2 as validation data and models were trained on WSI 1, 3, 4 because WSI1,2 also contains `Dataset1 `\n- I belive because of the split, my solution can survive when Private leaderboard released. My Validation scores on WSI 2 are nearly correlated to PB/LB\n\n\n#2. Modeling\n- Step 1: Train 8 models Yolov8l and Yolov8m with image size of 640 on 4 folds for 100 epochs\n- Step 2: Use WMF (Weighted mask fusion) to generate pseudo label on unlabeled data\n- Step 3: Retrain models on Labeled data + Pseudo data for 70 epochs, Image size are 640 and 1024\n- Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.\n\nAll models are trained on the following hyper parameters, I only show the param that I change compared to default values:\n\n| Param | Value |\n| --- | --- |\n| mixup | 0.5 |\n| copy-paste | 0.5 |\n| flipud | 0.5 |\n| degrees | 45 |\n| lr0 | 0.001 |\n\nI also modified the dataloader of Yolov8 to sampling uniformly `dataset1`, `dataset2` and `pseudo data`.\n\n#3. The effect of pseudo labeling\n- I used pseudo labeling two times, one on local with unlabeled data (**Stage1**) and one during submission on hidden test data (**Stage2**) as mentioned in **2.**\n- Pseudo labeling **stage1** improve my mask mAP50 on WSI 1 and 2 about ~ 0.04\n- **Stage2** Pseudo labeling improved both my LB/PB about 0.005, I think if base models are better, it could improve more.\n\nCompare of Stage2 pseudo labeling\n\n| Sub | LB | PB | \n| --- | --- | --- |\n| ENS1 + Stage2 pseudo | 0.537  | 0.505  \n| ENS1|  0.532 | 0.500  | \n\n\n#4. Choice of submission\n- I selected my two last submission on two set of my best CV models on different fold, all models are trained with pseudo labeling **Stage1**. Called ENS1 and ENS2.\n- My best private score is 0.544 (LB 0.551) which is ensemble of 8 models `Fold 2`\n\nSubmission:\n\n| Sub | LB | PB | \n| --- | --- | --- |\n| ENS1 + Stage2 pseudo | 0.537  | 0.505  \n| ENS2 + Stage2 pseudo | 0.549  | 0.532 \n\n\n#5. Why shake-up (Just my guess)\n- Clearly because the test data is a completly different WSI, this is why my validation strategy is to split the data by WSI\n- I think the private test data is small and very sensitive (only one WSI), I think It's not large enough to choose the best model for real-world data (I understand that the annotation process is very costly). I observed my local vaildation score is oscillated during training too.\n- Many people reported that dilation also cause the shake-up, may be my dilation iters is only 1, so the effect is minimized",
      "votes": null
    },
    {
      "id": "2379406",
      "postDate": "08/08/2023 06:33:16",
      "content": "<p>what a great solution. Thanks for sharing</p>",
      "rawMarkdown": "what a great solution. Thanks for sharing",
      "votes": null
    },
    {
      "id": "2379408",
      "postDate": "08/08/2023 06:34:37",
      "content": "<p>Cảm ơn bro</p>",
      "rawMarkdown": "Cảm ơn bro",
      "votes": null
    },
    {
      "id": "2379425",
      "postDate": "08/08/2023 06:46:01",
      "content": "<p>Greate works, thanks for sharing!!</p>",
      "rawMarkdown": "Greate works, thanks for sharing!!",
      "votes": null
    },
    {
      "id": "2383411",
      "postDate": "08/10/2023 10:38:51",
      "content": "<p>\"Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.\" Wow this part blow my mind. A truely unique solution. But how do you know this newly train model is good enough ? How do you evaluate it ?</p>\n<p>Also i have a question about your pseudo labeling method. From my understanding you have 4 folds (for each WSI) and for each fold you ensemble prediction of YoloV8M + YoloV8L on dataset 3. Finally add pseudo label dataset 3 on the train set and perform training and evaluate with the specific fold. Am i right ?</p>\n<p>Thank you and chúc mừng</p>",
      "rawMarkdown": "\"Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.\" Wow this part blow my mind. A truely unique solution. But how do you know this newly train model is good enough ? How do you evaluate it ?\n\nAlso i have a question about your pseudo labeling method. From my understanding you have 4 folds (for each WSI) and for each fold you ensemble prediction of YoloV8M + YoloV8L on dataset 3. Finally add pseudo label dataset 3 on the train set and perform training and evaluate with the specific fold. Am i right ?\n\nThank you and chúc mừng",
      "votes": null
    },
    {
      "id": "2384486",
      "postDate": "08/11/2023 01:37:52",
      "content": "<p>Thank you</p>\n<ul>\n<li><p>For hidden pseudo label part, I could not evaluate the exact performance  of new model. I just trust it, because I use the best hyper-parameters and number of epochs I found during K-fold training. And in my experiments, pseudo label on external data improved my results, so I believe it would as good when training on hidden dataset. Beside of that, training on hidden data is very hard to evaluate the performance because we have no access to the validation mAP during submission.</p></li>\n<li><p>Yes, I trained 4 models Yolov8M and 4 models Yolov8L on WSI dataset1,2 the run inference on dataset 3 using 8 models. Then threshold confidence scores and add them to training set, retrain and evaluate again on 4 folds as I did without pseudo label.</p></li>\n</ul>",
      "rawMarkdown": "Thank you\n\n- For hidden pseudo label part, I could not evaluate the exact performance  of new model. I just trust it, because I use the best hyper-parameters and number of epochs I found during K-fold training. And in my experiments, pseudo label on external data improved my results, so I believe it would as good when training on hidden dataset. Beside of that, training on hidden data is very hard to evaluate the performance because we have no access to the validation mAP during submission.\n\n- Yes, I trained 4 models Yolov8M and 4 models Yolov8L on WSI dataset1,2 the run inference on dataset 3 using 8 models. Then threshold confidence scores and add them to training set, retrain and evaluate again on 4 folds as I did without pseudo label.",
      "votes": null
    },
    {
      "id": "2384746",
      "postDate": "08/11/2023 04:17:40",
      "content": "<p>Thank you<br>\nThe hidden pseudo label part was truely a bold move</p>",
      "rawMarkdown": "Thank you\nThe hidden pseudo label part was truely a bold move",
      "votes": null
    },
    {
      "id": "2524230",
      "postDate": "11/14/2023 04:55:26",
      "content": "<p>Thank you for sharing, and thank you for explaining also what didn't work as well, it is always interesting learning from other's people experience! Pseudo-labelling seem tricky indeed but it seem to have improved slightly your score.</p>",
      "rawMarkdown": "Thank you for sharing, and thank you for explaining also what didn't work as well, it is always interesting learning from other's people experience! Pseudo-labelling seem tricky indeed but it seem to have improved slightly your score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2379406,
      "author_name": "syaoranclone",
      "author_url": "",
      "post_date": "08/08/2023 06:33:16",
      "content": "<p>what a great solution. Thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 2379408,
          "author_name": "ptran1203",
          "author_url": "",
          "post_date": "08/08/2023 06:34:37",
          "content": "<p>Cảm ơn bro</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2379425,
      "author_name": "cngminhtrn99",
      "author_url": "",
      "post_date": "08/08/2023 06:46:01",
      "content": "<p>Greate works, thanks for sharing!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2383411,
      "author_name": "researchbntz",
      "author_url": "",
      "post_date": "08/10/2023 10:38:51",
      "content": "<p>\"Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.\" Wow this part blow my mind. A truely unique solution. But how do you know this newly train model is good enough ? How do you evaluate it ?</p>\n<p>Also i have a question about your pseudo labeling method. From my understanding you have 4 folds (for each WSI) and for each fold you ensemble prediction of YoloV8M + YoloV8L on dataset 3. Finally add pseudo label dataset 3 on the train set and perform training and evaluate with the specific fold. Am i right ?</p>\n<p>Thank you and chúc mừng</p>",
      "votes": null,
      "replies": [
        {
          "id": 2384486,
          "author_name": "ptran1203",
          "author_url": "",
          "post_date": "08/11/2023 01:37:52",
          "content": "<p>Thank you</p>\n<ul>\n<li><p>For hidden pseudo label part, I could not evaluate the exact performance  of new model. I just trust it, because I use the best hyper-parameters and number of epochs I found during K-fold training. And in my experiments, pseudo label on external data improved my results, so I believe it would as good when training on hidden dataset. Beside of that, training on hidden data is very hard to evaluate the performance because we have no access to the validation mAP during submission.</p></li>\n<li><p>Yes, I trained 4 models Yolov8M and 4 models Yolov8L on WSI dataset1,2 the run inference on dataset 3 using 8 models. Then threshold confidence scores and add them to training set, retrain and evaluate again on 4 folds as I did without pseudo label.</p></li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2384746,
              "author_name": "researchbntz",
              "author_url": "",
              "post_date": "08/11/2023 04:17:40",
              "content": "<p>Thank you<br>\nThe hidden pseudo label part was truely a bold move</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2524230,
      "author_name": "awillame",
      "author_url": "",
      "post_date": "11/14/2023 04:55:26",
      "content": "<p>Thank you for sharing, and thank you for explaining also what didn't work as well, it is always interesting learning from other's people experience! Pseudo-labelling seem tricky indeed but it seem to have improved slightly your score.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2379386": "Congratuations to all the winner and thank the organizers for hosting such an interesting competition!\n\n\n- My solution is a simple ensemble of Yolov8l and Yolov8m trained on image size 640 and 1024\n- All of my submission using Dilation with Kernel size =  3 and Iters = 1 because on my local CV, dilation don't hurt (and sometimes improve) **dataset1** but decrease mAP on **dataset2** significantly.\n\n\nThings worked for me:\n- Yolov8l-seg and Yolov8m-seg\n- Copy-paste and Mixup augmentation (0.5 probability)\n- Balance sampling between dataset1, dataset2 and unlabeled data (pseudo label) \n- Pseudo-labeling on Unlabeled data\n- Larger image size (Increase CV but not LB/PB 😭)\n- Pseudo labeling on Test data (i.e: Retrain during submission)\n- Ensemble using Weighted mask fusion from 8th place Sartorius competition https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation/discussion/297998\n\nThings not worked for me:\n- Pseudo label more than one round\n- Ensemble on different fold\n\n\n# 1. Validation strategy\n\n- I split the data into train/valid by WSI (4 folds), during the competition, I mostly use WSI 1 and 2 as validation data and models were trained on WSI 1, 3, 4 because WSI1,2 also contains `Dataset1 `\n- I belive because of the split, my solution can survive when Private leaderboard released. My Validation scores on WSI 2 are nearly correlated to PB/LB\n\n\n#2. Modeling\n- Step 1: Train 8 models Yolov8l and Yolov8m with image size of 640 on 4 folds for 100 epochs\n- Step 2: Use WMF (Weighted mask fusion) to generate pseudo label on unlabeled data\n- Step 3: Retrain models on Labeled data + Pseudo data for 70 epochs, Image size are 640 and 1024\n- Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.\n\nAll models are trained on the following hyper parameters, I only show the param that I change compared to default values:\n\n| Param | Value |\n| --- | --- |\n| mixup | 0.5 |\n| copy-paste | 0.5 |\n| flipud | 0.5 |\n| degrees | 45 |\n| lr0 | 0.001 |\n\nI also modified the dataloader of Yolov8 to sampling uniformly `dataset1`, `dataset2` and `pseudo data`.\n\n#3. The effect of pseudo labeling\n- I used pseudo labeling two times, one on local with unlabeled data (**Stage1**) and one during submission on hidden test data (**Stage2**) as mentioned in **2.**\n- Pseudo labeling **stage1** improve my mask mAP50 on WSI 1 and 2 about ~ 0.04\n- **Stage2** Pseudo labeling improved both my LB/PB about 0.005, I think if base models are better, it could improve more.\n\nCompare of Stage2 pseudo labeling\n\n| Sub | LB | PB | \n| --- | --- | --- |\n| ENS1 + Stage2 pseudo | 0.537  | 0.505  \n| ENS1|  0.532 | 0.500  | \n\n\n#4. Choice of submission\n- I selected my two last submission on two set of my best CV models on different fold, all models are trained with pseudo labeling **Stage1**. Called ENS1 and ENS2.\n- My best private score is 0.544 (LB 0.551) which is ensemble of 8 models `Fold 2`\n\nSubmission:\n\n| Sub | LB | PB | \n| --- | --- | --- |\n| ENS1 + Stage2 pseudo | 0.537  | 0.505  \n| ENS2 + Stage2 pseudo | 0.549  | 0.532 \n\n\n#5. Why shake-up (Just my guess)\n- Clearly because the test data is a completly different WSI, this is why my validation strategy is to split the data by WSI\n- I think the private test data is small and very sensitive (only one WSI), I think It's not large enough to choose the best model for real-world data (I understand that the annotation process is very costly). I observed my local vaildation score is oscillated during training too.\n- Many people reported that dilation also cause the shake-up, may be my dilation iters is only 1, so the effect is minimized",
    "2379406": "what a great solution. Thanks for sharing",
    "2379408": "Cảm ơn bro",
    "2379425": "Greate works, thanks for sharing!!",
    "2383411": "\"Step 4 (During submission) I use my ensemble models to predict hidden test data and train new model Yolov8l x 640 on Labeled data + Private pseudo data, then append that model to ensemble and run inference again.\" Wow this part blow my mind. A truely unique solution. But how do you know this newly train model is good enough ? How do you evaluate it ?\n\nAlso i have a question about your pseudo labeling method. From my understanding you have 4 folds (for each WSI) and for each fold you ensemble prediction of YoloV8M + YoloV8L on dataset 3. Finally add pseudo label dataset 3 on the train set and perform training and evaluate with the specific fold. Am i right ?\n\nThank you and chúc mừng",
    "2384486": "Thank you\n\n- For hidden pseudo label part, I could not evaluate the exact performance  of new model. I just trust it, because I use the best hyper-parameters and number of epochs I found during K-fold training. And in my experiments, pseudo label on external data improved my results, so I believe it would as good when training on hidden dataset. Beside of that, training on hidden data is very hard to evaluate the performance because we have no access to the validation mAP during submission.\n\n- Yes, I trained 4 models Yolov8M and 4 models Yolov8L on WSI dataset1,2 the run inference on dataset 3 using 8 models. Then threshold confidence scores and add them to training set, retrain and evaluate again on 4 folds as I did without pseudo label.",
    "2384746": "Thank you\nThe hidden pseudo label part was truely a bold move",
    "2524230": "Thank you for sharing, and thank you for explaining also what didn't work as well, it is always interesting learning from other's people experience! Pseudo-labelling seem tricky indeed but it seem to have improved slightly your score."
  },
  "source": "meta"
}