{
  "id": 237999,
  "title": "FYI, Hand Label and Pseudo Label Did Not Help and Did Not Hurt",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/237999",
  "author_name": "",
  "post_date": "2021-05-11T00:19:44.855904900Z",
  "votes": 22,
  "comment_count": 21,
  "views": 0,
  "content": "<p>For the record, training with hand labeled and/or pseudo labeled public test <strong>did not help</strong> and <strong>did not hurt</strong> private test score. </p>\n<p>I had one of the best public test labels for image d48 (as I was in third place public LB with 0.944). I obtained it through an iterative process of pseudo labeling, hand labeling, and guided ensemble techniques to maximize public LB score. (In this comp but not all comps, these techniques were allowed by Kaggle).</p>\n<p>In the image below, you can see that my models trained <strong>with</strong> my pseudo/hand label d48 and <strong>without</strong> my d48 <strong>score the same</strong> private LB score. (The two public scores of 919 and 920 do not use pseudo labels/hand label, the others use pseudo labels/hand label).</p>\n<p>In the end, this competition came down to maximizing CV score, not exploiting public LB or hand labeling. (FYI, using pseudo labels did not change my CV score. FYI, my best private LB was without pseudo labels, but using pseudo labels achieved nearly the same private score).</p>\n<p><img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/pseudo.png\" alt=\"image\"></p>",
  "messages": [
    {
      "id": "1301091",
      "postDate": "05/11/2021 00:19:44",
      "content": "<p>For the record, training with hand labeled and/or pseudo labeled public test <strong>did not help</strong> and <strong>did not hurt</strong> private test score. </p>\n<p>I had one of the best public test labels for image d48 (as I was in third place public LB with 0.944). I obtained it through an iterative process of pseudo labeling, hand labeling, and guided ensemble techniques to maximize public LB score. (In this comp but not all comps, these techniques were allowed by Kaggle).</p>\n<p>In the image below, you can see that my models trained <strong>with</strong> my pseudo/hand label d48 and <strong>without</strong> my d48 <strong>score the same</strong> private LB score. (The two public scores of 919 and 920 do not use pseudo labels/hand label, the others use pseudo labels/hand label).</p>\n<p>In the end, this competition came down to maximizing CV score, not exploiting public LB or hand labeling. (FYI, using pseudo labels did not change my CV score. FYI, my best private LB was without pseudo labels, but using pseudo labels achieved nearly the same private score).</p>\n<p><img src=\"https://www.ccom.ucsd.edu/~cdeotte/Kaggle/pseudo.png\" alt=\"image\"></p>",
      "rawMarkdown": "For the record, training with hand labeled and/or pseudo labeled public test **did not help** and **did not hurt** private test score. \n\nI had one of the best public test labels for image d48 (as I was in third place public LB with 0.944). I obtained it through an iterative process of pseudo labeling, hand labeling, and guided ensemble techniques to maximize public LB score. (In this comp but not all comps, these techniques were allowed by Kaggle).\n\nIn the image below, you can see that my models trained **with** my pseudo/hand label d48 and **without** my d48 **score the same** private LB score. (The two public scores of 919 and 920 do not use pseudo labels/hand label, the others use pseudo labels/hand label).\n\nIn the end, this competition came down to maximizing CV score, not exploiting public LB or hand labeling. (FYI, using pseudo labels did not change my CV score. FYI, my best private LB was without pseudo labels, but using pseudo labels achieved nearly the same private score).\n\n![image](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/pseudo.png)",
      "votes": null
    },
    {
      "id": "1301100",
      "postDate": "05/11/2021 00:24:44",
      "content": "<p>Same here. Our best results (5th place on private) is an old kernel submitted 2 months ago. d48 looks to be a trap. Really weird to have such outlier.</p>",
      "rawMarkdown": "Same here. Our best results (5th place on private) is an old kernel submitted 2 months ago. d48 looks to be a trap. Really weird to have such outlier.",
      "votes": null
    },
    {
      "id": "1301103",
      "postDate": "05/11/2021 00:26:11",
      "content": "<p>Did you hand label sclerotic / Fibrous Crescent glomerulis ? <br>\nI think there were none labeled in the private test set, hence using labels that had a high score for<code>d488xxx</code> could hurt your models.</p>",
      "rawMarkdown": "Did you hand label sclerotic / Fibrous Crescent glomerulis ? \nI think there were none labeled in the private test set, hence using labels that had a high score for`d488xxx` could hurt your models.",
      "votes": null
    },
    {
      "id": "1301107",
      "postDate": "05/11/2021 00:28:23",
      "content": "<p>It is almost the same here, although pseudo label didn’t harm my score much(0.001~0.002), but also not helpful for the private score.</p>",
      "rawMarkdown": "It is almost the same here, although pseudo label didn’t harm my score much(0.001~0.002), but also not helpful for the private score.",
      "votes": null
    },
    {
      "id": "1301110",
      "postDate": "05/11/2021 00:30:13",
      "content": "<p>We were suspecting this and decided not to use labels of d4 for training. I would love to hear about your solution anyway.</p>",
      "rawMarkdown": "We were suspecting this and decided not to use labels of d4 for training. I would love to hear about your solution anyway.",
      "votes": null
    },
    {
      "id": "1301111",
      "postDate": "05/11/2021 00:30:25",
      "content": "<p>I also used d488xxx hand label data, but my score between pesudo labeling dataset and normal dataset don’t have much differences.</p>",
      "rawMarkdown": "I also used d488xxx hand label data, but my score between pesudo labeling dataset and normal dataset don’t have much differences.",
      "votes": null
    },
    {
      "id": "1301115",
      "postDate": "05/11/2021 00:32:29",
      "content": "<p>Yes, and it wasn't just the strange glomeruli in d48 that were labeled positive. I found other outliers. In public test image <code>d488c759a</code> and image <code>57512b7f1</code>, when a glomeruli was in a dark shadowy area on the edge of the image then it wasn't labeled. Removing masks from these dark areas increased public LB another +0.003 or so.</p>",
      "rawMarkdown": "Yes, and it wasn't just the strange glomeruli in d48 that were labeled positive. I found other outliers. In public test image `d488c759a` and image `57512b7f1`, when a glomeruli was in a dark shadowy area on the edge of the image then it wasn't labeled. Removing masks from these dark areas increased public LB another +0.003 or so.",
      "votes": null
    },
    {
      "id": "1301119",
      "postDate": "05/11/2021 00:34:05",
      "content": "<p>Yes we did. I will publish our handlabels tomorrow so we can see how d48 differs from all other images. Maybe we've an image with Fibrous Crescent glomerulis in private dataset but not annotated.</p>",
      "rawMarkdown": "Yes we did. I will publish our handlabels tomorrow so we can see how d48 differs from all other images. Maybe we've an image with Fibrous Crescent glomerulis in private dataset but not annotated.",
      "votes": null
    },
    {
      "id": "1301123",
      "postDate": "05/11/2021 00:36:37",
      "content": "<p>my model was neutral to the pseudo labels. interesting…</p>",
      "rawMarkdown": "my model was neutral to the pseudo labels. interesting...",
      "votes": null
    },
    {
      "id": "1301127",
      "postDate": "05/11/2021 00:38:20",
      "content": "<p>That is what my guessing too :(<br>\nI think private test set is totally difference to public test especially on ground truth. We found many sclerotic cells/fibrous crescent are included in ground truth from probing not only d488 but in other public images as well.</p>\n<p>I think they changed private ground truth from the re-run (1-2 days ago?) so that only non-scelotic cells are included (just my guess).</p>\n<p>Our best private model score only 0.916 on lb and from a month ago. </p>",
      "rawMarkdown": "That is what my guessing too :(\nI think private test set is totally difference to public test especially on ground truth. We found many sclerotic cells/fibrous crescent are included in ground truth from probing not only d488 but in other public images as well.\n\nI think they changed private ground truth from the re-run (1-2 days ago?) so that only non-scelotic cells are included (just my guess).\n\nOur best private model score only 0.916 on lb and from a month ago.",
      "votes": null
    },
    {
      "id": "1301128",
      "postDate": "05/11/2021 00:38:30",
      "content": "<p>I did more than hand label sclerotic / Fibrous Crescent glomerulis. There were many more anomalies in the 5 public test images. By repeatedly training pseudo labeled models and inspecting the test prediction results, i was able to find many outliers in public test. I manually corrected all these outliers, maximized public LB and used the result to train my final models.</p>\n<p>But as you see from my private public scores above, using my pseudo labels did not hurt me nor help me.</p>\n<p>I trained models using 2x 3x 4x scale reduction. My best models ensembled only 2x and 3x scale reduction. Apparently ensembling with 4x scale reduction models slightly decreased my private score.</p>",
      "rawMarkdown": "I did more than hand label sclerotic / Fibrous Crescent glomerulis. There were many more anomalies in the 5 public test images. By repeatedly training pseudo labeled models and inspecting the test prediction results, i was able to find many outliers in public test. I manually corrected all these outliers, maximized public LB and used the result to train my final models.\n\nBut as you see from my private public scores above, using my pseudo labels did not hurt me nor help me.\n\nI trained models using 2x 3x 4x scale reduction. My best models ensembled only 2x and 3x scale reduction. Apparently ensembling with 4x scale reduction models slightly decreased my private score.",
      "votes": null
    },
    {
      "id": "1301130",
      "postDate": "05/11/2021 00:39:28",
      "content": "<p>Here are the per-image scores after our handlabels:</p>\n<ul>\n<li>3589adb90 = 0.191 </li>\n<li>2ec3f1bb9 = 0.191</li>\n<li>aa05346ff = 0.189++</li>\n<li>57512b7f1 = 0.190+</li>\n<li>d488c759a = 0.187++</li>\n</ul>",
      "rawMarkdown": "Here are the per-image scores after our handlabels:\n- 3589adb90 = 0.191 \n- 2ec3f1bb9 = 0.191\n- aa05346ff = 0.189++\n- 57512b7f1 = 0.190+\n- d488c759a = 0.187++",
      "votes": null
    },
    {
      "id": "1301131",
      "postDate": "05/11/2021 00:39:38",
      "content": "<p>Mine were neutral too. Perhaps i wasn't clear above.</p>\n<p>My point is that they \"did not help\". They also \"did not hurt\" my models.</p>",
      "rawMarkdown": "Mine were neutral too. Perhaps i wasn't clear above.\n\nMy point is that they \"did not help\". They also \"did not hurt\" my models.",
      "votes": null
    },
    {
      "id": "1301154",
      "postDate": "05/11/2021 00:54:23",
      "content": "<p>it's totally a lottery, the private score will ascend who gusses there's hard sample like d4, hard to say what's the meaning of this compete, just have a fun with competor.😦</p>",
      "rawMarkdown": "it's totally a lottery, the private score will ascend who gusses there's hard sample like d4, hard to say what's the meaning of this compete, just have a fun with competor.😦",
      "votes": null
    },
    {
      "id": "1301158",
      "postDate": "05/11/2021 00:57:34",
      "content": "<p>Since they are dangerous. Its effect might be counteracted. Personally the wrong choice is more fatal than the lost one.(e.g. 94/100 - (94/101 or 93/99)). Therefore, I used mean and weighted mean which could get a +0.004 (from 0.945 to 0.949) in my submissions. </p>",
      "rawMarkdown": "Since they are dangerous. Its effect might be counteracted. Personally the wrong choice is more fatal than the lost one.(e.g. 94/100 - (94/101 or 93/99)). Therefore, I used mean and weighted mean which could get a +0.004 (from 0.945 to 0.949) in my submissions.",
      "votes": null
    },
    {
      "id": "1301163",
      "postDate": "05/11/2021 01:00:49",
      "content": "<p>Interesting, can you explain this more? I'm looking forward to reading your writeup.</p>",
      "rawMarkdown": "Interesting, can you explain this more? I'm looking forward to reading your writeup.",
      "votes": null
    },
    {
      "id": "1301164",
      "postDate": "05/11/2021 01:01:37",
      "content": "<p>Thanx for sharing.  I dont know if, when one train a model to learn all datasets and an outlier, it become  weaker, in this kind of competition. If the task is  image segmentation, it become more powerful, not the opposite. </p>",
      "rawMarkdown": "Thanx for sharing.  I dont know if, when one train a model to learn all datasets and an outlier, it become  weaker, in this kind of competition. If the task is  image segmentation, it become more powerful, not the opposite.",
      "votes": null
    },
    {
      "id": "1301167",
      "postDate": "05/11/2021 01:05:23",
      "content": "<p>It depends on the private dataset. There were labels in the public test dataset that were not present in the train dataset. Therefore your model cannot learn these. If the private dataset had similar labels then using hand labeled and/or pseudo labeled public test would help.</p>\n<p>Since it did not help (nor hurt) private test score, this means that the private test dataset is similar to the train images and does not contain the outliers present in the public test dataset.</p>",
      "rawMarkdown": "It depends on the private dataset. There were labels in the public test dataset that were not present in the train dataset. Therefore your model cannot learn these. If the private dataset had similar labels then using hand labeled and/or pseudo labeled public test would help.\n\nSince it did not help (nor hurt) private test score, this means that the private test dataset is similar to the train images and does not contain the outliers present in the public test dataset.",
      "votes": null
    },
    {
      "id": "1301190",
      "postDate": "05/11/2021 01:25:31",
      "content": "<p>Actually, it's a sudden idea, which comes from my handing label process.  To be honest, my handing label ability seems not so good which always hurts me. 😂<br>\nBut I found when I add some wrong labels, LB decreased a lot; when I deleted some true labels, LB decreased slowly.<br>\nI recognized a simple mathematic mechanism.👀<br>\nDice score is represented as Intersection divided by union. Assume my current score is 0.94(94/100). If I chose 1% wrong labels, score will decreased to 0.93(94/101). but if I lost 1% true labels, score will deceased to 0.939(93/99).<br>\nI strongly believe the usage of labels in d488, I dont want to drop them in one of my final choices. but lowering its effect is important since it is very dangerous. Thus, using weighted mean between the models which didn't contain any of those labels to increase their effect. After that, using mean between the former results with models containing d488. It could decrease the file size from 5.6MB to 5.34MB. Since I thought smaller size is more safe as mentioned before. LB is lower (from 0.937 to 0.928) but Private LB is much higher (from 0.942 to 0.949).  <br>\nMy formal models pefrom not so good with LB 0.945 .</p>\n<p><strong>Those are the main differences between models using d488 with models not using it.</strong></p>",
      "rawMarkdown": "Actually, it's a sudden idea, which comes from my handing label process.  To be honest, my handing label ability seems not so good which always hurts me. 😂\nBut I found when I add some wrong labels, LB decreased a lot; when I deleted some true labels, LB decreased slowly.\nI recognized a simple mathematic mechanism.👀\nDice score is represented as Intersection divided by union. Assume my current score is 0.94(94/100). If I chose 1% wrong labels, score will decreased to 0.93(94/101). but if I lost 1% true labels, score will deceased to 0.939(93/99).\nI strongly believe the usage of labels in d488, I dont want to drop them in one of my final choices. but lowering its effect is important since it is very dangerous. Thus, using weighted mean between the models which didn't contain any of those labels to increase their effect. After that, using mean between the former results with models containing d488. It could decrease the file size from 5.6MB to 5.34MB. Since I thought smaller size is more safe as mentioned before. LB is lower (from 0.937 to 0.928) but Private LB is much higher (from 0.942 to 0.949).  \nMy formal models pefrom not so good with LB 0.945 .\n \n**Those are the main differences between models using d488 with models not using it.**",
      "votes": null
    },
    {
      "id": "1301205",
      "postDate": "05/11/2021 01:35:00",
      "content": "<p>I have much higher scoring submissions that I didn’t select for final submission. Those were not trained with pseudo labels. In particular one that scored 0.945 on private and 0.919 on the public. I think the private test must have dark sclerotic glomeruli that aren’t labeled.</p>",
      "rawMarkdown": "I have much higher scoring submissions that I didn’t select for final submission. Those were not trained with pseudo labels. In particular one that scored 0.945 on private and 0.919 on the public. I think the private test must have dark sclerotic glomeruli that aren’t labeled.",
      "votes": null
    },
    {
      "id": "1301266",
      "postDate": "05/11/2021 02:09:30",
      "content": "<p>you were crystal clear. my bad :)</p>",
      "rawMarkdown": "you were crystal clear. my bad :)",
      "votes": null
    },
    {
      "id": "1301812",
      "postDate": "05/11/2021 08:54:15",
      "content": "<p>Any handlabels prize? 😄<br>\n<a href=\"https://www.kaggle.com/mpware/handlabels-prize\" target=\"_blank\">https://www.kaggle.com/mpware/handlabels-prize</a></p>\n<p>Here are some examples of the infamous d48. And some reasons why we decided to use it in one of our submissions.</p>",
      "rawMarkdown": "Any handlabels prize? 😄\nhttps://www.kaggle.com/mpware/handlabels-prize\n\nHere are some examples of the infamous d48. And some reasons why we decided to use it in one of our submissions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1301100,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "05/11/2021 00:24:44",
      "content": "<p>Same here. Our best results (5th place on private) is an old kernel submitted 2 months ago. d48 looks to be a trap. Really weird to have such outlier.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301110,
          "author_name": "theudas",
          "author_url": "",
          "post_date": "05/11/2021 00:30:13",
          "content": "<p>We were suspecting this and decided not to use labels of d4 for training. I would love to hear about your solution anyway.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301115,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/11/2021 00:32:29",
          "content": "<p>Yes, and it wasn't just the strange glomeruli in d48 that were labeled positive. I found other outliers. In public test image <code>d488c759a</code> and image <code>57512b7f1</code>, when a glomeruli was in a dark shadowy area on the edge of the image then it wasn't labeled. Removing masks from these dark areas increased public LB another +0.003 or so.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301130,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "05/11/2021 00:39:28",
          "content": "<p>Here are the per-image scores after our handlabels:</p>\n<ul>\n<li>3589adb90 = 0.191 </li>\n<li>2ec3f1bb9 = 0.191</li>\n<li>aa05346ff = 0.189++</li>\n<li>57512b7f1 = 0.190+</li>\n<li>d488c759a = 0.187++</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301164,
          "author_name": "rpsantosakaggle",
          "author_url": "",
          "post_date": "05/11/2021 01:01:37",
          "content": "<p>Thanx for sharing.  I dont know if, when one train a model to learn all datasets and an outlier, it become  weaker, in this kind of competition. If the task is  image segmentation, it become more powerful, not the opposite. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301167,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/11/2021 01:05:23",
          "content": "<p>It depends on the private dataset. There were labels in the public test dataset that were not present in the train dataset. Therefore your model cannot learn these. If the private dataset had similar labels then using hand labeled and/or pseudo labeled public test would help.</p>\n<p>Since it did not help (nor hurt) private test score, this means that the private test dataset is similar to the train images and does not contain the outliers present in the public test dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301103,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "05/11/2021 00:26:11",
      "content": "<p>Did you hand label sclerotic / Fibrous Crescent glomerulis ? <br>\nI think there were none labeled in the private test set, hence using labels that had a high score for<code>d488xxx</code> could hurt your models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301111,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "05/11/2021 00:30:25",
          "content": "<p>I also used d488xxx hand label data, but my score between pesudo labeling dataset and normal dataset don’t have much differences.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301119,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "05/11/2021 00:34:05",
          "content": "<p>Yes we did. I will publish our handlabels tomorrow so we can see how d48 differs from all other images. Maybe we've an image with Fibrous Crescent glomerulis in private dataset but not annotated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301127,
          "author_name": "tom88jerry",
          "author_url": "",
          "post_date": "05/11/2021 00:38:20",
          "content": "<p>That is what my guessing too :(<br>\nI think private test set is totally difference to public test especially on ground truth. We found many sclerotic cells/fibrous crescent are included in ground truth from probing not only d488 but in other public images as well.</p>\n<p>I think they changed private ground truth from the re-run (1-2 days ago?) so that only non-scelotic cells are included (just my guess).</p>\n<p>Our best private model score only 0.916 on lb and from a month ago. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301128,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/11/2021 00:38:30",
          "content": "<p>I did more than hand label sclerotic / Fibrous Crescent glomerulis. There were many more anomalies in the 5 public test images. By repeatedly training pseudo labeled models and inspecting the test prediction results, i was able to find many outliers in public test. I manually corrected all these outliers, maximized public LB and used the result to train my final models.</p>\n<p>But as you see from my private public scores above, using my pseudo labels did not hurt me nor help me.</p>\n<p>I trained models using 2x 3x 4x scale reduction. My best models ensembled only 2x and 3x scale reduction. Apparently ensembling with 4x scale reduction models slightly decreased my private score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301154,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "05/11/2021 00:54:23",
          "content": "<p>it's totally a lottery, the private score will ascend who gusses there's hard sample like d4, hard to say what's the meaning of this compete, just have a fun with competor.😦</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301107,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "05/11/2021 00:28:23",
      "content": "<p>It is almost the same here, although pseudo label didn’t harm my score much(0.001~0.002), but also not helpful for the private score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1301123,
      "author_name": "andrasferenczi",
      "author_url": "",
      "post_date": "05/11/2021 00:36:37",
      "content": "<p>my model was neutral to the pseudo labels. interesting…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301131,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/11/2021 00:39:38",
          "content": "<p>Mine were neutral too. Perhaps i wasn't clear above.</p>\n<p>My point is that they \"did not help\". They also \"did not hurt\" my models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301266,
          "author_name": "andrasferenczi",
          "author_url": "",
          "post_date": "05/11/2021 02:09:30",
          "content": "<p>you were crystal clear. my bad :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301158,
      "author_name": "southsakura",
      "author_url": "",
      "post_date": "05/11/2021 00:57:34",
      "content": "<p>Since they are dangerous. Its effect might be counteracted. Personally the wrong choice is more fatal than the lost one.(e.g. 94/100 - (94/101 or 93/99)). Therefore, I used mean and weighted mean which could get a +0.004 (from 0.945 to 0.949) in my submissions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1301163,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/11/2021 01:00:49",
          "content": "<p>Interesting, can you explain this more? I'm looking forward to reading your writeup.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301190,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "05/11/2021 01:25:31",
          "content": "<p>Actually, it's a sudden idea, which comes from my handing label process.  To be honest, my handing label ability seems not so good which always hurts me. 😂<br>\nBut I found when I add some wrong labels, LB decreased a lot; when I deleted some true labels, LB decreased slowly.<br>\nI recognized a simple mathematic mechanism.👀<br>\nDice score is represented as Intersection divided by union. Assume my current score is 0.94(94/100). If I chose 1% wrong labels, score will decreased to 0.93(94/101). but if I lost 1% true labels, score will deceased to 0.939(93/99).<br>\nI strongly believe the usage of labels in d488, I dont want to drop them in one of my final choices. but lowering its effect is important since it is very dangerous. Thus, using weighted mean between the models which didn't contain any of those labels to increase their effect. After that, using mean between the former results with models containing d488. It could decrease the file size from 5.6MB to 5.34MB. Since I thought smaller size is more safe as mentioned before. LB is lower (from 0.937 to 0.928) but Private LB is much higher (from 0.942 to 0.949).  <br>\nMy formal models pefrom not so good with LB 0.945 .</p>\n<p><strong>Those are the main differences between models using d488 with models not using it.</strong></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301205,
      "author_name": "erikdali",
      "author_url": "",
      "post_date": "05/11/2021 01:35:00",
      "content": "<p>I have much higher scoring submissions that I didn’t select for final submission. Those were not trained with pseudo labels. In particular one that scored 0.945 on private and 0.919 on the public. I think the private test must have dark sclerotic glomeruli that aren’t labeled.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1301812,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "05/11/2021 08:54:15",
      "content": "<p>Any handlabels prize? 😄<br>\n<a href=\"https://www.kaggle.com/mpware/handlabels-prize\" target=\"_blank\">https://www.kaggle.com/mpware/handlabels-prize</a></p>\n<p>Here are some examples of the infamous d48. And some reasons why we decided to use it in one of our submissions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1301091": "For the record, training with hand labeled and/or pseudo labeled public test **did not help** and **did not hurt** private test score. \n\nI had one of the best public test labels for image d48 (as I was in third place public LB with 0.944). I obtained it through an iterative process of pseudo labeling, hand labeling, and guided ensemble techniques to maximize public LB score. (In this comp but not all comps, these techniques were allowed by Kaggle).\n\nIn the image below, you can see that my models trained **with** my pseudo/hand label d48 and **without** my d48 **score the same** private LB score. (The two public scores of 919 and 920 do not use pseudo labels/hand label, the others use pseudo labels/hand label).\n\nIn the end, this competition came down to maximizing CV score, not exploiting public LB or hand labeling. (FYI, using pseudo labels did not change my CV score. FYI, my best private LB was without pseudo labels, but using pseudo labels achieved nearly the same private score).\n\n![image](https://www.ccom.ucsd.edu/~cdeotte/Kaggle/pseudo.png)",
    "1301100": "Same here. Our best results (5th place on private) is an old kernel submitted 2 months ago. d48 looks to be a trap. Really weird to have such outlier.",
    "1301103": "Did you hand label sclerotic / Fibrous Crescent glomerulis ? \nI think there were none labeled in the private test set, hence using labels that had a high score for`d488xxx` could hurt your models.",
    "1301107": "It is almost the same here, although pseudo label didn’t harm my score much(0.001~0.002), but also not helpful for the private score.",
    "1301110": "We were suspecting this and decided not to use labels of d4 for training. I would love to hear about your solution anyway.",
    "1301111": "I also used d488xxx hand label data, but my score between pesudo labeling dataset and normal dataset don’t have much differences.",
    "1301115": "Yes, and it wasn't just the strange glomeruli in d48 that were labeled positive. I found other outliers. In public test image `d488c759a` and image `57512b7f1`, when a glomeruli was in a dark shadowy area on the edge of the image then it wasn't labeled. Removing masks from these dark areas increased public LB another +0.003 or so.",
    "1301119": "Yes we did. I will publish our handlabels tomorrow so we can see how d48 differs from all other images. Maybe we've an image with Fibrous Crescent glomerulis in private dataset but not annotated.",
    "1301123": "my model was neutral to the pseudo labels. interesting...",
    "1301127": "That is what my guessing too :(\nI think private test set is totally difference to public test especially on ground truth. We found many sclerotic cells/fibrous crescent are included in ground truth from probing not only d488 but in other public images as well.\n\nI think they changed private ground truth from the re-run (1-2 days ago?) so that only non-scelotic cells are included (just my guess).\n\nOur best private model score only 0.916 on lb and from a month ago.",
    "1301128": "I did more than hand label sclerotic / Fibrous Crescent glomerulis. There were many more anomalies in the 5 public test images. By repeatedly training pseudo labeled models and inspecting the test prediction results, i was able to find many outliers in public test. I manually corrected all these outliers, maximized public LB and used the result to train my final models.\n\nBut as you see from my private public scores above, using my pseudo labels did not hurt me nor help me.\n\nI trained models using 2x 3x 4x scale reduction. My best models ensembled only 2x and 3x scale reduction. Apparently ensembling with 4x scale reduction models slightly decreased my private score.",
    "1301130": "Here are the per-image scores after our handlabels:\n- 3589adb90 = 0.191 \n- 2ec3f1bb9 = 0.191\n- aa05346ff = 0.189++\n- 57512b7f1 = 0.190+\n- d488c759a = 0.187++",
    "1301131": "Mine were neutral too. Perhaps i wasn't clear above.\n\nMy point is that they \"did not help\". They also \"did not hurt\" my models.",
    "1301154": "it's totally a lottery, the private score will ascend who gusses there's hard sample like d4, hard to say what's the meaning of this compete, just have a fun with competor.😦",
    "1301158": "Since they are dangerous. Its effect might be counteracted. Personally the wrong choice is more fatal than the lost one.(e.g. 94/100 - (94/101 or 93/99)). Therefore, I used mean and weighted mean which could get a +0.004 (from 0.945 to 0.949) in my submissions.",
    "1301163": "Interesting, can you explain this more? I'm looking forward to reading your writeup.",
    "1301164": "Thanx for sharing.  I dont know if, when one train a model to learn all datasets and an outlier, it become  weaker, in this kind of competition. If the task is  image segmentation, it become more powerful, not the opposite.",
    "1301167": "It depends on the private dataset. There were labels in the public test dataset that were not present in the train dataset. Therefore your model cannot learn these. If the private dataset had similar labels then using hand labeled and/or pseudo labeled public test would help.\n\nSince it did not help (nor hurt) private test score, this means that the private test dataset is similar to the train images and does not contain the outliers present in the public test dataset.",
    "1301190": "Actually, it's a sudden idea, which comes from my handing label process.  To be honest, my handing label ability seems not so good which always hurts me. 😂\nBut I found when I add some wrong labels, LB decreased a lot; when I deleted some true labels, LB decreased slowly.\nI recognized a simple mathematic mechanism.👀\nDice score is represented as Intersection divided by union. Assume my current score is 0.94(94/100). If I chose 1% wrong labels, score will decreased to 0.93(94/101). but if I lost 1% true labels, score will deceased to 0.939(93/99).\nI strongly believe the usage of labels in d488, I dont want to drop them in one of my final choices. but lowering its effect is important since it is very dangerous. Thus, using weighted mean between the models which didn't contain any of those labels to increase their effect. After that, using mean between the former results with models containing d488. It could decrease the file size from 5.6MB to 5.34MB. Since I thought smaller size is more safe as mentioned before. LB is lower (from 0.937 to 0.928) but Private LB is much higher (from 0.942 to 0.949).  \nMy formal models pefrom not so good with LB 0.945 .\n \n**Those are the main differences between models using d488 with models not using it.**",
    "1301205": "I have much higher scoring submissions that I didn’t select for final submission. Those were not trained with pseudo labels. In particular one that scored 0.945 on private and 0.919 on the public. I think the private test must have dark sclerotic glomeruli that aren’t labeled.",
    "1301266": "you were crystal clear. my bad :)",
    "1301812": "Any handlabels prize? 😄\nhttps://www.kaggle.com/mpware/handlabels-prize\n\nHere are some examples of the infamous d48. And some reasons why we decided to use it in one of our submissions."
  },
  "source": "meta"
}