{
  "id": 295539,
  "title": "Semi-supervised Experiments",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/295539",
  "author_name": "",
  "post_date": "2021-12-16T12:16:08.769788100Z",
  "votes": 11,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I haven't seen any discussion about semi-supervised learning so I'm starting it.</p>\n<ul>\n<li>I predicted pseudo labels with my best 5 fold model's blend (0.32 LB score) and verified them visually. I didn't saw any major problems but there were some false negatives.</li>\n<li>I tried training with train images + all semi supervised images and model overfits. Validation loss and mAP doesn't stop improving. Validation mAP became 0.45 after 20k iterations.</li>\n<li>I thought shsy5y and astro pseudo labels could be bad so I trained with train images + cort semi supervised images and model overfits again. Even though I only added cort semi supervised images, validation mAP scores of astro and shsy5y don't stop improving. Validation mAP became 0.40 after 20k iterations.</li>\n</ul>\n<p>I couldn't figure out the source of crazy leakage. Copying and pasting objects from annotations (<a href=\"https://arxiv.org/pdf/2012.07177.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.07177.pdf</a>) might help dealing with this overfitting issue. Does anyone able to improve their score using semi-supervised images with pseudo labels?</p>",
  "messages": [
    {
      "id": "1620081",
      "postDate": "12/16/2021 12:16:08",
      "content": "<p>I haven't seen any discussion about semi-supervised learning so I'm starting it.</p>\n<ul>\n<li>I predicted pseudo labels with my best 5 fold model's blend (0.32 LB score) and verified them visually. I didn't saw any major problems but there were some false negatives.</li>\n<li>I tried training with train images + all semi supervised images and model overfits. Validation loss and mAP doesn't stop improving. Validation mAP became 0.45 after 20k iterations.</li>\n<li>I thought shsy5y and astro pseudo labels could be bad so I trained with train images + cort semi supervised images and model overfits again. Even though I only added cort semi supervised images, validation mAP scores of astro and shsy5y don't stop improving. Validation mAP became 0.40 after 20k iterations.</li>\n</ul>\n<p>I couldn't figure out the source of crazy leakage. Copying and pasting objects from annotations (<a href=\"https://arxiv.org/pdf/2012.07177.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.07177.pdf</a>) might help dealing with this overfitting issue. Does anyone able to improve their score using semi-supervised images with pseudo labels?</p>",
      "rawMarkdown": "I haven't seen any discussion about semi-supervised learning so I'm starting it.\n\n* I predicted pseudo labels with my best 5 fold model's blend (0.32 LB score) and verified them visually. I didn't saw any major problems but there were some false negatives.\n*  I tried training with train images + all semi supervised images and model overfits. Validation loss and mAP doesn't stop improving. Validation mAP became 0.45 after 20k iterations.\n* I thought shsy5y and astro pseudo labels could be bad so I trained with train images + cort semi supervised images and model overfits again. Even though I only added cort semi supervised images, validation mAP scores of astro and shsy5y don't stop improving. Validation mAP became 0.40 after 20k iterations.\n\nI couldn't figure out the source of crazy leakage. Copying and pasting objects from annotations (https://arxiv.org/pdf/2012.07177.pdf) might help dealing with this overfitting issue. Does anyone able to improve their score using semi-supervised images with pseudo labels?",
      "votes": null
    },
    {
      "id": "1620098",
      "postDate": "12/16/2021 12:42:41",
      "content": "<p>Sorry if I misunderstood, but if you use a 5 fold model ensemble to produce labels, doesn't that leak the information from the validation set into your pseudo labels? Or do you use a validation set without including samples from these folds?</p>",
      "rawMarkdown": "Sorry if I misunderstood, but if you use a 5 fold model ensemble to produce labels, doesn't that leak the information from the validation set into your pseudo labels? Or do you use a validation set without including samples from these folds?",
      "votes": null
    },
    {
      "id": "1620113",
      "postDate": "12/16/2021 12:52:51",
      "content": "<p>No, you didn't misunderstand. You actually solved my leakage problem. I should have noticed that earlier…</p>\n<p>Besides my leakage, did you able to get any boost from semi-supervised pseudo labels?</p>",
      "rawMarkdown": "No, you didn't misunderstand. You actually solved my leakage problem. I should have noticed that earlier...\n\nBesides my leakage, did you able to get any boost from semi-supervised pseudo labels?",
      "votes": null
    },
    {
      "id": "1620119",
      "postDate": "12/16/2021 13:00:21",
      "content": "<p>Not sure, I did try to train a model similar to your approach, although I've not seen any significant improvement. To be fair, I've not done any semi-supervised method earlier, so I'm not entirely sure if I've done it correctly. 😑</p>",
      "rawMarkdown": "Not sure, I did try to train a model similar to your approach, although I've not seen any significant improvement. To be fair, I've not done any semi-supervised method earlier, so I'm not entirely sure if I've done it correctly. 😑",
      "votes": null
    },
    {
      "id": "1620138",
      "postDate": "12/16/2021 13:20:11",
      "content": "<p>as <a href=\"https://www.kaggle.com/woprime\" target=\"_blank\">@woprime</a> pointed out that is the sort of leakage you have, try to use your best single fold model to generate pseudos and validate on same valid fold data<br>\nFrom few experiments I tried so far I see CV improv ~0.004 but not reflect to LB (marginal increase or similar score) - maybe we need to move to proper semi-supervised methods </p>",
      "rawMarkdown": "as @woprime pointed out that is the sort of leakage you have, try to use your best single fold model to generate pseudos and validate on same valid fold data\nFrom few experiments I tried so far I see CV improv ~0.004 but not reflect to LB (marginal increase or similar score) - maybe we need to move to proper semi-supervised methods",
      "votes": null
    },
    {
      "id": "1620558",
      "postDate": "12/16/2021 22:02:23",
      "content": "<p>Haha i was making the same <a href=\"https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/294434\" target=\"_blank\">mistake</a> earlier this week on the Petfinder competition. Glad to see I'm not alone :)<br>\nUPD: using SSL Data worked for me and gave close to ~0.015 cv increase and ~0.005 lb ^. I suspect using 2 or 3 fold ensemble might give ~0.004-5 more boost(along with more robustness) since that's what happened with my normal pipeline </p>",
      "rawMarkdown": "Haha i was making the same [mistake](https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/294434) earlier this week on the Petfinder competition. Glad to see I'm not alone :)\nUPD: using SSL Data worked for me and gave close to ~0.015 cv increase and ~0.005 lb ^. I suspect using 2 or 3 fold ensemble might give ~0.004-5 more boost(along with more robustness) since that's what happened with my normal pipeline",
      "votes": null
    },
    {
      "id": "1620617",
      "postDate": "12/17/2021 01:14:37",
      "content": "<p>hey has copy paste augmentation increased accuracy over your normal pipeline?</p>",
      "rawMarkdown": "hey has copy paste augmentation increased accuracy over your normal pipeline?",
      "votes": null
    },
    {
      "id": "1620655",
      "postDate": "12/17/2021 02:40:49",
      "content": "<p>I got a similar result, I'm worried that the improvement in CV is also due to some leakage, i.e. I gave information to the pseudo label by <em>choosing</em> the parameters fine-tuned on the validation set.</p>",
      "rawMarkdown": "I got a similar result, I'm worried that the improvement in CV is also due to some leakage, i.e. I gave information to the pseudo label by *choosing* the parameters fine-tuned on the validation set.",
      "votes": null
    },
    {
      "id": "1620712",
      "postDate": "12/17/2021 04:51:00",
      "content": "<p>I haven't tried it yet but they got huge boost with that method in their experiments.</p>",
      "rawMarkdown": "I haven't tried it yet but they got huge boost with that method in their experiments.",
      "votes": null
    },
    {
      "id": "1626135",
      "postDate": "12/22/2021 13:50:52",
      "content": "<p>Wait… So SSL actually worked? And that's a huge boost in both CV and LB!</p>",
      "rawMarkdown": "Wait... So SSL actually worked? And that's a huge boost in both CV and LB!",
      "votes": null
    },
    {
      "id": "1626744",
      "postDate": "12/23/2021 07:11:04",
      "content": "<p>I have some updates regarding this topic.</p>\n<p>My baselines models' 5 fold OOF score is 0.305829 and LB score is 0.321. As <a href=\"https://www.kaggle.com/woprime\" target=\"_blank\">@woprime</a> pointed my leakage is caused by predicting semi supervised images with 5 models and blending their predictions. I run my inference code on them again and saved every models' predictions separately.</p>\n<p>I trained my new models with same folds + individual semi supervised predictions. For example baseline model 1 is trained with fold 2, 3, 4, 5 and validated on fold 1. Baseline model 1 is used for predicting semi supervised images and predictions are saved as a coco dataset. Later, new model is trained with fold 2, 3, 4, 5, semi supervised predictions of baseline model 1 and validated on fold 1. I did this for only fold 1 so far.</p>\n<p>With that training setup, my model reached 0.38222 validation mAP in 10 epochs and it scored 0.322 on LB when I blend this model with previous 5 baseline models. I continue training for 50 epochs and val mAP became 0.531260, and it scored 0.323 on LB when it is blended with previous 5 baseline models.</p>\n<p>The leakage still exists and I can't find the source of it, but it doesn't overfit like my previous experiments. I don't know what to do anymore :D</p>",
      "rawMarkdown": "I have some updates regarding this topic.\n\nMy baselines models' 5 fold OOF score is 0.305829 and LB score is 0.321. As @woprime pointed my leakage is caused by predicting semi supervised images with 5 models and blending their predictions. I run my inference code on them again and saved every models' predictions separately.\n\nI trained my new models with same folds + individual semi supervised predictions. For example baseline model 1 is trained with fold 2, 3, 4, 5 and validated on fold 1. Baseline model 1 is used for predicting semi supervised images and predictions are saved as a coco dataset. Later, new model is trained with fold 2, 3, 4, 5, semi supervised predictions of baseline model 1 and validated on fold 1. I did this for only fold 1 so far.\n\nWith that training setup, my model reached 0.38222 validation mAP in 10 epochs and it scored 0.322 on LB when I blend this model with previous 5 baseline models. I continue training for 50 epochs and val mAP became 0.531260, and it scored 0.323 on LB when it is blended with previous 5 baseline models.\n\nThe leakage still exists and I can't find the source of it, but it doesn't overfit like my previous experiments. I don't know what to do anymore :D",
      "votes": null
    },
    {
      "id": "1626773",
      "postDate": "12/23/2021 07:54:17",
      "content": "<p>The leakage is still exist, when you mention</p>\n<blockquote>\n  <p>my model reached 0.38222 validation mAP in 10 epochs</p>\n</blockquote>\n<p>Since your setup is leak free in theory: use fold 1's model predict simi data, train with merged pseudo label and the fold 2, 3, 4, 5; validate with fold 1. 0.382 is an unusual score, it seems your simi train set contains the information of validation label. <br>\nMaybe your code has some bug that load the fold1 to your train coco dataset? I suggest u check the overlap between train and validate coco json first.</p>",
      "rawMarkdown": "The leakage is still exist, when you mention\n> my model reached 0.38222 validation mAP in 10 epochs\n\nSince your setup is leak free in theory: use fold 1's model predict simi data, train with merged pseudo label and the fold 2, 3, 4, 5; validate with fold 1. 0.382 is an unusual score, it seems your simi train set contains the information of validation label. \nMaybe your code has some bug that load the fold1 to your train coco dataset? I suggest u check the overlap between train and validate coco json first.",
      "votes": null
    },
    {
      "id": "1626800",
      "postDate": "12/23/2021 08:18:03",
      "content": "<p>I checked my code and json files. There is no overlap. As you said, the setup is leak free in theory but there is one thing that I'm suspecting. NMS IoU thresholds, score thresholds and area thresholds are tuned based on validation set mAP. Maybe, I shouldn't even remove overlaps while predicting semi supervised images?</p>",
      "rawMarkdown": "I checked my code and json files. There is no overlap. As you said, the setup is leak free in theory but there is one thing that I'm suspecting. NMS IoU thresholds, score thresholds and area thresholds are tuned based on validation set mAP. Maybe, I shouldn't even remove overlaps while predicting semi supervised images?",
      "votes": null
    },
    {
      "id": "1626833",
      "postDate": "12/23/2021 08:35:59",
      "content": "<p>I don't believe that the process of the simi label could cause such serious leakage. <code>0.531260</code> is the score that model could not reach without learning the label of the validation set. Since your validation score without simi pseudo label training is ok it should not cause by your validation setup, I feel you still feed your model with validation set info.<br>\nHowever I can't put forward any other hyposis, wish other guys can solve your leakage.</p>",
      "rawMarkdown": "I don't believe that the process of the simi label could cause such serious leakage. `0.531260` is the score that model could not reach without learning the label of the validation set. Since your validation score without simi pseudo label training is ok it should not cause by your validation setup, I feel you still feed your model with validation set info.\nHowever I can't put forward any other hyposis, wish other guys can solve your leakage.",
      "votes": null
    },
    {
      "id": "1627211",
      "postDate": "12/23/2021 16:16:48",
      "content": "<p>Do you know whether the leakage occurs for all 3 categories or only for specific ones?</p>\n<p>One thing which might be relevant is that in the Sartorius dataset, images belong to a \"sample\" according to the <code>sample_id</code> column in the <code>train.csv</code> file. As I understand it, a sample is a probe of cells which are imaged multiple times over time (see also the <code>plate_time</code> column). When looking at different images with the same <code>sample_id</code>, images of the cort category are very similar. For the other categories, that's less the case. I believe that if you have images from the same sample spread over the training and validation sets, this might also introduce some leakage. By adding semi supervised data to the training set, you might be adding more images which belong to samples which are also present in the validation set. If that's the case, I think you would notice it mainly for the cort category. This is a general observation and, if relevant, the <code>sample_id</code> should probably even be considered for constructing the initial folds.</p>",
      "rawMarkdown": "Do you know whether the leakage occurs for all 3 categories or only for specific ones?\n\nOne thing which might be relevant is that in the Sartorius dataset, images belong to a \"sample\" according to the `sample_id` column in the `train.csv` file. As I understand it, a sample is a probe of cells which are imaged multiple times over time (see also the `plate_time` column). When looking at different images with the same `sample_id`, images of the cort category are very similar. For the other categories, that's less the case. I believe that if you have images from the same sample spread over the training and validation sets, this might also introduce some leakage. By adding semi supervised data to the training set, you might be adding more images which belong to samples which are also present in the validation set. If that's the case, I think you would notice it mainly for the cort category. This is a general observation and, if relevant, the `sample_id` should probably even be considered for constructing the initial folds.",
      "votes": null
    },
    {
      "id": "1632084",
      "postDate": "12/29/2021 06:42:55",
      "content": "<p>Final updates:</p>\n<p>I finally found the source of leakage. I was using path of validation set instead of semi supervised so it was training on training + validation set. </p>\n<p>After fixing the leakage, I improved my models OOF score from 0.3058 to 0.308, but LB score decreased from 0.323 to 0.316. I noticed astro's mAP increased from 0.2 to 0.22, but cort's mAP decreased from 0.395 to 0.389. I guess the models weren't strong enough for semi supervised after all.</p>",
      "rawMarkdown": "Final updates:\n\nI finally found the source of leakage. I was using path of validation set instead of semi supervised so it was training on training + validation set. \n\nAfter fixing the leakage, I improved my models OOF score from 0.3058 to 0.308, but LB score decreased from 0.323 to 0.316. I noticed astro's mAP increased from 0.2 to 0.22, but cort's mAP decreased from 0.395 to 0.389. I guess the models weren't strong enough for semi supervised after all.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1620098,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "12/16/2021 12:42:41",
      "content": "<p>Sorry if I misunderstood, but if you use a 5 fold model ensemble to produce labels, doesn't that leak the information from the validation set into your pseudo labels? Or do you use a validation set without including samples from these folds?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1620113,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/16/2021 12:52:51",
          "content": "<p>No, you didn't misunderstand. You actually solved my leakage problem. I should have noticed that earlier…</p>\n<p>Besides my leakage, did you able to get any boost from semi-supervised pseudo labels?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1620119,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "12/16/2021 13:00:21",
          "content": "<p>Not sure, I did try to train a model similar to your approach, although I've not seen any significant improvement. To be fair, I've not done any semi-supervised method earlier, so I'm not entirely sure if I've done it correctly. 😑</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1620558,
          "author_name": "ferlockx",
          "author_url": "",
          "post_date": "12/16/2021 22:02:23",
          "content": "<p>Haha i was making the same <a href=\"https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/294434\" target=\"_blank\">mistake</a> earlier this week on the Petfinder competition. Glad to see I'm not alone :)<br>\nUPD: using SSL Data worked for me and gave close to ~0.015 cv increase and ~0.005 lb ^. I suspect using 2 or 3 fold ensemble might give ~0.004-5 more boost(along with more robustness) since that's what happened with my normal pipeline </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626135,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "12/22/2021 13:50:52",
          "content": "<p>Wait… So SSL actually worked? And that's a huge boost in both CV and LB!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1620138,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "12/16/2021 13:20:11",
      "content": "<p>as <a href=\"https://www.kaggle.com/woprime\" target=\"_blank\">@woprime</a> pointed out that is the sort of leakage you have, try to use your best single fold model to generate pseudos and validate on same valid fold data<br>\nFrom few experiments I tried so far I see CV improv ~0.004 but not reflect to LB (marginal increase or similar score) - maybe we need to move to proper semi-supervised methods </p>",
      "votes": null,
      "replies": [
        {
          "id": 1620655,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "12/17/2021 02:40:49",
          "content": "<p>I got a similar result, I'm worried that the improvement in CV is also due to some leakage, i.e. I gave information to the pseudo label by <em>choosing</em> the parameters fine-tuned on the validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1620617,
      "author_name": "ferlockx",
      "author_url": "",
      "post_date": "12/17/2021 01:14:37",
      "content": "<p>hey has copy paste augmentation increased accuracy over your normal pipeline?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1620712,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/17/2021 04:51:00",
          "content": "<p>I haven't tried it yet but they got huge boost with that method in their experiments.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1626744,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "12/23/2021 07:11:04",
      "content": "<p>I have some updates regarding this topic.</p>\n<p>My baselines models' 5 fold OOF score is 0.305829 and LB score is 0.321. As <a href=\"https://www.kaggle.com/woprime\" target=\"_blank\">@woprime</a> pointed my leakage is caused by predicting semi supervised images with 5 models and blending their predictions. I run my inference code on them again and saved every models' predictions separately.</p>\n<p>I trained my new models with same folds + individual semi supervised predictions. For example baseline model 1 is trained with fold 2, 3, 4, 5 and validated on fold 1. Baseline model 1 is used for predicting semi supervised images and predictions are saved as a coco dataset. Later, new model is trained with fold 2, 3, 4, 5, semi supervised predictions of baseline model 1 and validated on fold 1. I did this for only fold 1 so far.</p>\n<p>With that training setup, my model reached 0.38222 validation mAP in 10 epochs and it scored 0.322 on LB when I blend this model with previous 5 baseline models. I continue training for 50 epochs and val mAP became 0.531260, and it scored 0.323 on LB when it is blended with previous 5 baseline models.</p>\n<p>The leakage still exists and I can't find the source of it, but it doesn't overfit like my previous experiments. I don't know what to do anymore :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 1626773,
          "author_name": "steamedsheep",
          "author_url": "",
          "post_date": "12/23/2021 07:54:17",
          "content": "<p>The leakage is still exist, when you mention</p>\n<blockquote>\n  <p>my model reached 0.38222 validation mAP in 10 epochs</p>\n</blockquote>\n<p>Since your setup is leak free in theory: use fold 1's model predict simi data, train with merged pseudo label and the fold 2, 3, 4, 5; validate with fold 1. 0.382 is an unusual score, it seems your simi train set contains the information of validation label. <br>\nMaybe your code has some bug that load the fold1 to your train coco dataset? I suggest u check the overlap between train and validate coco json first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626800,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/23/2021 08:18:03",
          "content": "<p>I checked my code and json files. There is no overlap. As you said, the setup is leak free in theory but there is one thing that I'm suspecting. NMS IoU thresholds, score thresholds and area thresholds are tuned based on validation set mAP. Maybe, I shouldn't even remove overlaps while predicting semi supervised images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1626833,
          "author_name": "steamedsheep",
          "author_url": "",
          "post_date": "12/23/2021 08:35:59",
          "content": "<p>I don't believe that the process of the simi label could cause such serious leakage. <code>0.531260</code> is the score that model could not reach without learning the label of the validation set. Since your validation score without simi pseudo label training is ok it should not cause by your validation setup, I feel you still feed your model with validation set info.<br>\nHowever I can't put forward any other hyposis, wish other guys can solve your leakage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1627211,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "12/23/2021 16:16:48",
          "content": "<p>Do you know whether the leakage occurs for all 3 categories or only for specific ones?</p>\n<p>One thing which might be relevant is that in the Sartorius dataset, images belong to a \"sample\" according to the <code>sample_id</code> column in the <code>train.csv</code> file. As I understand it, a sample is a probe of cells which are imaged multiple times over time (see also the <code>plate_time</code> column). When looking at different images with the same <code>sample_id</code>, images of the cort category are very similar. For the other categories, that's less the case. I believe that if you have images from the same sample spread over the training and validation sets, this might also introduce some leakage. By adding semi supervised data to the training set, you might be adding more images which belong to samples which are also present in the validation set. If that's the case, I think you would notice it mainly for the cort category. This is a general observation and, if relevant, the <code>sample_id</code> should probably even be considered for constructing the initial folds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1632084,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "12/29/2021 06:42:55",
      "content": "<p>Final updates:</p>\n<p>I finally found the source of leakage. I was using path of validation set instead of semi supervised so it was training on training + validation set. </p>\n<p>After fixing the leakage, I improved my models OOF score from 0.3058 to 0.308, but LB score decreased from 0.323 to 0.316. I noticed astro's mAP increased from 0.2 to 0.22, but cort's mAP decreased from 0.395 to 0.389. I guess the models weren't strong enough for semi supervised after all.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1620081": "I haven't seen any discussion about semi-supervised learning so I'm starting it.\n\n* I predicted pseudo labels with my best 5 fold model's blend (0.32 LB score) and verified them visually. I didn't saw any major problems but there were some false negatives.\n*  I tried training with train images + all semi supervised images and model overfits. Validation loss and mAP doesn't stop improving. Validation mAP became 0.45 after 20k iterations.\n* I thought shsy5y and astro pseudo labels could be bad so I trained with train images + cort semi supervised images and model overfits again. Even though I only added cort semi supervised images, validation mAP scores of astro and shsy5y don't stop improving. Validation mAP became 0.40 after 20k iterations.\n\nI couldn't figure out the source of crazy leakage. Copying and pasting objects from annotations (https://arxiv.org/pdf/2012.07177.pdf) might help dealing with this overfitting issue. Does anyone able to improve their score using semi-supervised images with pseudo labels?",
    "1620098": "Sorry if I misunderstood, but if you use a 5 fold model ensemble to produce labels, doesn't that leak the information from the validation set into your pseudo labels? Or do you use a validation set without including samples from these folds?",
    "1620113": "No, you didn't misunderstand. You actually solved my leakage problem. I should have noticed that earlier...\n\nBesides my leakage, did you able to get any boost from semi-supervised pseudo labels?",
    "1620119": "Not sure, I did try to train a model similar to your approach, although I've not seen any significant improvement. To be fair, I've not done any semi-supervised method earlier, so I'm not entirely sure if I've done it correctly. 😑",
    "1620138": "as @woprime pointed out that is the sort of leakage you have, try to use your best single fold model to generate pseudos and validate on same valid fold data\nFrom few experiments I tried so far I see CV improv ~0.004 but not reflect to LB (marginal increase or similar score) - maybe we need to move to proper semi-supervised methods",
    "1620558": "Haha i was making the same [mistake](https://www.kaggle.com/c/petfinder-pawpularity-score/discussion/294434) earlier this week on the Petfinder competition. Glad to see I'm not alone :)\nUPD: using SSL Data worked for me and gave close to ~0.015 cv increase and ~0.005 lb ^. I suspect using 2 or 3 fold ensemble might give ~0.004-5 more boost(along with more robustness) since that's what happened with my normal pipeline",
    "1620617": "hey has copy paste augmentation increased accuracy over your normal pipeline?",
    "1620655": "I got a similar result, I'm worried that the improvement in CV is also due to some leakage, i.e. I gave information to the pseudo label by *choosing* the parameters fine-tuned on the validation set.",
    "1620712": "I haven't tried it yet but they got huge boost with that method in their experiments.",
    "1626135": "Wait... So SSL actually worked? And that's a huge boost in both CV and LB!",
    "1626744": "I have some updates regarding this topic.\n\nMy baselines models' 5 fold OOF score is 0.305829 and LB score is 0.321. As @woprime pointed my leakage is caused by predicting semi supervised images with 5 models and blending their predictions. I run my inference code on them again and saved every models' predictions separately.\n\nI trained my new models with same folds + individual semi supervised predictions. For example baseline model 1 is trained with fold 2, 3, 4, 5 and validated on fold 1. Baseline model 1 is used for predicting semi supervised images and predictions are saved as a coco dataset. Later, new model is trained with fold 2, 3, 4, 5, semi supervised predictions of baseline model 1 and validated on fold 1. I did this for only fold 1 so far.\n\nWith that training setup, my model reached 0.38222 validation mAP in 10 epochs and it scored 0.322 on LB when I blend this model with previous 5 baseline models. I continue training for 50 epochs and val mAP became 0.531260, and it scored 0.323 on LB when it is blended with previous 5 baseline models.\n\nThe leakage still exists and I can't find the source of it, but it doesn't overfit like my previous experiments. I don't know what to do anymore :D",
    "1626773": "The leakage is still exist, when you mention\n> my model reached 0.38222 validation mAP in 10 epochs\n\nSince your setup is leak free in theory: use fold 1's model predict simi data, train with merged pseudo label and the fold 2, 3, 4, 5; validate with fold 1. 0.382 is an unusual score, it seems your simi train set contains the information of validation label. \nMaybe your code has some bug that load the fold1 to your train coco dataset? I suggest u check the overlap between train and validate coco json first.",
    "1626800": "I checked my code and json files. There is no overlap. As you said, the setup is leak free in theory but there is one thing that I'm suspecting. NMS IoU thresholds, score thresholds and area thresholds are tuned based on validation set mAP. Maybe, I shouldn't even remove overlaps while predicting semi supervised images?",
    "1626833": "I don't believe that the process of the simi label could cause such serious leakage. `0.531260` is the score that model could not reach without learning the label of the validation set. Since your validation score without simi pseudo label training is ok it should not cause by your validation setup, I feel you still feed your model with validation set info.\nHowever I can't put forward any other hyposis, wish other guys can solve your leakage.",
    "1627211": "Do you know whether the leakage occurs for all 3 categories or only for specific ones?\n\nOne thing which might be relevant is that in the Sartorius dataset, images belong to a \"sample\" according to the `sample_id` column in the `train.csv` file. As I understand it, a sample is a probe of cells which are imaged multiple times over time (see also the `plate_time` column). When looking at different images with the same `sample_id`, images of the cort category are very similar. For the other categories, that's less the case. I believe that if you have images from the same sample spread over the training and validation sets, this might also introduce some leakage. By adding semi supervised data to the training set, you might be adding more images which belong to samples which are also present in the validation set. If that's the case, I think you would notice it mainly for the cort category. This is a general observation and, if relevant, the `sample_id` should probably even be considered for constructing the initial folds.",
    "1632084": "Final updates:\n\nI finally found the source of leakage. I was using path of validation set instead of semi supervised so it was training on training + validation set. \n\nAfter fixing the leakage, I improved my models OOF score from 0.3058 to 0.308, but LB score decreased from 0.323 to 0.316. I noticed astro's mAP increased from 0.2 to 0.22, but cort's mAP decreased from 0.395 to 0.389. I guess the models weren't strong enough for semi supervised after all."
  },
  "source": "meta"
}