{
  "id": 341130,
  "title": "Does anyone able to get any boost from external data?",
  "url": "/competitions/hubmap-organ-segmentation/discussion/341130",
  "author_name": "",
  "post_date": "2022-08-01T11:35:39.340085900Z",
  "votes": 12,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I did this experiment on H&amp;E stained spleen images. I learned about functional tissue units of spleen from HPA dictionary and tried to annotate some images by myself.</p>\n<p><a href=\"https://www.proteinatlas.org/learn/dictionary/normal/spleen\" target=\"_blank\">https://www.proteinatlas.org/learn/dictionary/normal/spleen</a></p>\n<p>I annotated the test image, the image from HPA dictionary and 3 images from <a href=\"https://histology.medicine.umich.edu/resources/lymphatic-system#iv-spleen\" target=\"_blank\">histology.medicine.umich.edu</a>. I added newly annotated images to training set. My validation and lb score decreased after using them. I thought okay I probably messed up the annotations. I tried using HuBMAP Colonic Crypt in my training set this time. There are 7 PAS stained colon images in that dataset. My validation and lb score decreased again.</p>\n<p>What's your experience with using external data?</p>",
  "messages": [
    {
      "id": "1879962",
      "postDate": "08/01/2022 11:35:39",
      "content": "<p>I did this experiment on H&amp;E stained spleen images. I learned about functional tissue units of spleen from HPA dictionary and tried to annotate some images by myself.</p>\n<p><a href=\"https://www.proteinatlas.org/learn/dictionary/normal/spleen\" target=\"_blank\">https://www.proteinatlas.org/learn/dictionary/normal/spleen</a></p>\n<p>I annotated the test image, the image from HPA dictionary and 3 images from <a href=\"https://histology.medicine.umich.edu/resources/lymphatic-system#iv-spleen\" target=\"_blank\">histology.medicine.umich.edu</a>. I added newly annotated images to training set. My validation and lb score decreased after using them. I thought okay I probably messed up the annotations. I tried using HuBMAP Colonic Crypt in my training set this time. There are 7 PAS stained colon images in that dataset. My validation and lb score decreased again.</p>\n<p>What's your experience with using external data?</p>",
      "rawMarkdown": "I did this experiment on H&E stained spleen images. I learned about functional tissue units of spleen from HPA dictionary and tried to annotate some images by myself.\n\nhttps://www.proteinatlas.org/learn/dictionary/normal/spleen\n\nI annotated the test image, the image from HPA dictionary and 3 images from [histology.medicine.umich.edu](https://histology.medicine.umich.edu/resources/lymphatic-system#iv-spleen). I added newly annotated images to training set. My validation and lb score decreased after using them. I thought okay I probably messed up the annotations. I tried using HuBMAP Colonic Crypt in my training set this time. There are 7 PAS stained colon images in that dataset. My validation and lb score decreased again.\n\nWhat's your experience with using external data?",
      "votes": null
    },
    {
      "id": "1880012",
      "postDate": "08/01/2022 12:09:59",
      "content": "<p>the way to do is:<br>\n1) remove akggle annotation from N spleen training images<br>\n2) replace with hand-annotated the N images</p>\n<p>using the original kaggle annotation as ground truth, measure the dice loss of your annotation<br>\nthis is the quality of your annotation</p>\n<p>(in fact if you hand annotate the same objects T times on different days, you will find your annotation is different)</p>\n<p>train 2 model with (1) and (2), and maybe (1)+(2). This measures the how your label quality (label noise) affects dice score</p>\n<p>you have to think of way to use (2) to get better or smiliar results than (1).<br>\nif there is no way to improve results using your hand annotation, you don't have to try hand-label external data at all. </p>\n<p>instead, you should use pesudo label, or self-supervised</p>\n<hr>\n<p>on a side note, find old segmentation competitions with post submission and test image data that you can download.<br>\nyou will find that even if you submit hand annotation to server:</p>\n<p>model results &gt; hand annotation results<br>\n(for hit objects)</p>\n<p>hand annotation is only better when your model completely missed the target object.<br>\nmodel dice is always better than hand annotation for hit object</p>\n<p>that is why we sometimes don't worry about people playing cheat.</p>\n<hr>\n<p>you can treat yourself as a model. measure your error and ensemble (weighted by error) your results into the rest of the machine models.<br>\ni.e. your labels is only pesudo label and not true labels.</p>",
      "rawMarkdown": "the way to do is:\n1) remove akggle annotation from N spleen training images\n2) replace with hand-annotated the N images\n\nusing the original kaggle annotation as ground truth, measure the dice loss of your annotation\nthis is the quality of your annotation\n\n(in fact if you hand annotate the same objects T times on different days, you will find your annotation is different)\n\ntrain 2 model with (1) and (2), and maybe (1)+(2). This measures the how your label quality (label noise) affects dice score\n\nyou have to think of way to use (2) to get better or smiliar results than (1).\nif there is no way to improve results using your hand annotation, you don't have to try hand-label external data at all. \n\n\ninstead, you should use pesudo label, or self-supervised\n\n---\n\non a side note, find old segmentation competitions with post submission and test image data that you can download.\nyou will find that even if you submit hand annotation to server:\n\nmodel results > hand annotation results\n(for hit objects)\n\nhand annotation is only better when your model completely missed the target object.\nmodel dice is always better than hand annotation for hit object\n\nthat is why we sometimes don't worry about people playing cheat.\n\n---\n\nyou can treat yourself as a model. measure your error and ensemble (weighted by error) your results into the rest of the machine models.\ni.e. your labels is only pesudo label and not true labels.",
      "votes": null
    },
    {
      "id": "1888823",
      "postDate": "08/07/2022 19:44:13",
      "content": "<p>Thanks for doing this experiment and sharing your findings! I think your experience is not unique, I also found that adding external data can actually decrease performance initially.</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Thanks for doing this experiment and sharing your findings! I think your experience is not unique, I also found that adding external data can actually decrease performance initially.\n\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "1892272",
      "postDate": "08/10/2022 02:14:33",
      "content": "<p>another possibility is to use auxiliary loss.<br>\ne.g. we cannot segment the image accuracy manually, but we can accuracy annotate the center.<br>\ni called this aux label.</p>\n<p>then we can train the network with kaggle and pesudo center label with auxilay loss.</p>\n<p>other  aux label includes low resolution mask</p>\n<hr>\n<p>a even better solution</p>\n<ol>\n<li>hand label external data + kaggle data  (can be full mask or aux label)</li>\n<li>for kaggle data, you have only kaggle label</li>\n</ol>\n<p>train a model with joint loss (kaggle,external data + your label) and (kaggle data + kaggle label)<br>\nlet's the newtok figure out the relationship between your label and kaggle label itself.</p>\n<hr>\n<p>yet another possibility is to create three class human label:<br>\nbackground (white), unknown(blue), ftu(yellow)</p>\n<p><img src=\"https://i.ibb.co/LNXTsjg/images.png\" alt=\"https://i.ibb.co/LNXTsjg/images.png\"></p>\n<p>for the unknown region, loss in not computed and back propagated.</p>\n<p>to set the unknown region, you can measure the difference of you hand annotation and kaggle true annotation</p>",
      "rawMarkdown": "another possibility is to use auxiliary loss.\ne.g. we cannot segment the image accuracy manually, but we can accuracy annotate the center.\ni called this aux label.\n\nthen we can train the network with kaggle and pesudo center label with auxilay loss.\n\nother  aux label includes low resolution mask\n\n\n---\na even better solution\n\n1. hand label external data + kaggle data  (can be full mask or aux label)\n2. for kaggle data, you have only kaggle label\n\ntrain a model with joint loss (kaggle,external data + your label) and (kaggle data + kaggle label)\nlet's the newtok figure out the relationship between your label and kaggle label itself.\n\n---\n\n\nyet another possibility is to create three class human label:\nbackground (white), unknown(blue), ftu(yellow)\n\n![https://i.ibb.co/LNXTsjg/images.png](https://i.ibb.co/LNXTsjg/images.png)\n\nfor the unknown region, loss in not computed and back propagated.\n\nto set the unknown region, you can measure the difference of you hand annotation and kaggle true annotation",
      "votes": null
    },
    {
      "id": "1892273",
      "postDate": "08/10/2022 02:15:35",
      "content": "<p>HuBMAP Colonic Crypt  did not decrease my LB and CV score.<br>\nBut it shows not improvement in LB (not enough decimal place revealed)</p>",
      "rawMarkdown": "HuBMAP Colonic Crypt  did not decrease my LB and CV score.\nBut it shows not improvement in LB (not enough decimal place revealed)",
      "votes": null
    },
    {
      "id": "1892279",
      "postDate": "08/10/2022 02:18:51",
      "content": "<p>the object of external data is to improve hidden Hubmap LB score. you should just concentrated on Hubmap LB score first when you make submission.</p>",
      "rawMarkdown": "the object of external data is to improve hidden Hubmap LB score. you should just concentrated on Hubmap LB score first when you make submission.",
      "votes": null
    },
    {
      "id": "1892285",
      "postDate": "08/10/2022 02:33:09",
      "content": "<p>yet another ideal to fix poor pesudo label from hand annotation</p>\n<p><img src=\"https://i.ibb.co/d6djCvX/Selection-093.png\" alt=\"https://i.ibb.co/d6djCvX/Selection-093.png\"><br>\n<a href=\"https://paperswithcode.com/task/interactive-segmentation/\" target=\"_blank\">https://paperswithcode.com/task/interactive-segmentation/</a></p>\n<p>in particular<br>\n<a href=\"https://github.com/navidstuv/NuClick\" target=\"_blank\">https://github.com/navidstuv/NuClick</a><br>\n<img src=\"https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif\" alt=\"https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif\"></p>",
      "rawMarkdown": "yet another ideal to fix poor pesudo label from hand annotation\n\n![https://i.ibb.co/d6djCvX/Selection-093.png](https://i.ibb.co/d6djCvX/Selection-093.png)\nhttps://paperswithcode.com/task/interactive-segmentation/\n\nin particular\nhttps://github.com/navidstuv/NuClick\n![https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif](https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif)",
      "votes": null
    },
    {
      "id": "1892318",
      "postDate": "08/10/2022 03:26:56",
      "content": "<p>From my results, HE/PAS stains pseudo-labeling boosted by LB score, with nearly identical CV score. Kidney and colon images come from previous hubmap competition. Spleen and prostate images come from GTEX portal.</p>\n<p>Now, I can get 0.59 hubmap score.</p>\n<p>In further experiments, how to balance between HPA-trained models and HPA/External-trained models is also important to get a good LB score.</p>",
      "rawMarkdown": "From my results, HE/PAS stains pseudo-labeling boosted by LB score, with nearly identical CV score. Kidney and colon images come from previous hubmap competition. Spleen and prostate images come from GTEX portal.\n\nNow, I can get 0.59 hubmap score.\n\nIn further experiments, how to balance between HPA-trained models and HPA/External-trained models is also important to get a good LB score.",
      "votes": null
    },
    {
      "id": "1892554",
      "postDate": "08/10/2022 06:48:16",
      "content": "<p>there is another source for prostate images:<br>\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment\" target=\"_blank\">https://www.kaggle.com/c/prostate-cancer-grade-assessment</a></p>\n<p>but you need to select non cancer tissue</p>",
      "rawMarkdown": "there is another source for prostate images:\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment\n\nbut you need to select non cancer tissue",
      "votes": null
    },
    {
      "id": "1892557",
      "postDate": "08/10/2022 06:50:16",
      "content": "<p>\"In further experiments, how to balance …  important to get a good LB score.\"</p>\n<p>last year hubmap competition, this is quite some shakeup.  we need to prevent the same for here.</p>",
      "rawMarkdown": "\"In further experiments, how to balance ...  important to get a good LB score.\"\n\nlast year hubmap competition, this is quite some shakeup.  we need to prevent the same for here.",
      "votes": null
    },
    {
      "id": "1894850",
      "postDate": "08/11/2022 18:42:31",
      "content": "<p>\"In further experiments, how to balance … important to get a good LB score.\"</p>\n<p>one possible training dynamics is as follows. you can verify it using your learning process.</p>\n<p><a href=\"https://ibb.co/KcPSWHw\"><img src=\"https://i.ibb.co/CpcDQF2/Selection-065.png\" alt=\"Selection-065\"></a><br>\n<a href=\"https://ibb.co/0m3PTyM\"><img src=\"https://i.ibb.co/djynF42/Selection-073.png\" alt=\"Selection-073\"></a></p>",
      "rawMarkdown": "\"In further experiments, how to balance … important to get a good LB score.\"\n\none possible training dynamics is as follows. you can verify it using your learning process.\n\n\n<a href=\"https://ibb.co/KcPSWHw\"><img src=\"https://i.ibb.co/CpcDQF2/Selection-065.png\" alt=\"Selection-065\" border=\"0\"></a>\n<a href=\"https://ibb.co/0m3PTyM\"><img src=\"https://i.ibb.co/djynF42/Selection-073.png\" alt=\"Selection-073\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "1896727",
      "postDate": "08/13/2022 05:15:07",
      "content": "<p>related:<br>\n\"Adaptive Early-Learning Correction for Segmentation from Noisy Annotations\" - cvpr 2022</p>\n<p>\"we propose to update the annotations corresponding to different categories at different<br>\ntimes by detecting when early learning has occurred and memorization is about to begin using the training performance<br>\nof the model\"</p>\n<p>update the annotations  = pesudo label<br>\ndetecting when early learning has occurred = best point of generalisation<br>\n memorization  = overfitting</p>\n<hr>\n<p><a href=\"https://openaccess.thecvf.com/content/CVPR2022/papers/Liu_Adaptive_Early-Learning_Correction_for_Segmentation_From_Noisy_Annotations_CVPR_2022_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/CVPR2022/papers/Liu_Adaptive_Early-Learning_Correction_for_Segmentation_From_Noisy_Annotations_CVPR_2022_paper.pdf</a></p>",
      "rawMarkdown": "related:\n\"Adaptive Early-Learning Correction for Segmentation from Noisy Annotations\" - cvpr 2022\n\n\"we propose to update the annotations corresponding to different categories at different\ntimes by detecting when early learning has occurred and memorization is about to begin using the training performance\nof the model\"\n\nupdate the annotations  = pesudo label\ndetecting when early learning has occurred = best point of generalisation\n memorization  = overfitting\n\n---\n\nhttps://openaccess.thecvf.com/content/CVPR2022/papers/Liu_Adaptive_Early-Learning_Correction_for_Segmentation_From_Noisy_Annotations_CVPR_2022_paper.pdf",
      "votes": null
    },
    {
      "id": "1899507",
      "postDate": "08/15/2022 10:00:41",
      "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> <br>\n\"pseudo-labeling boosted by LB score\"</p>\n<p>do you use previous label for kidney/colon or do you relabel using pseudo-label.</p>\n<p>do you have results or each organ?<br>\ne.g. using pseudo for spleen only, LB score improve by xxx.</p>\n<p>thanks!</p>",
      "rawMarkdown": "carnozhao \n\"pseudo-labeling boosted by LB score\"\n\ndo you use previous label for kidney/colon or do you relabel using pseudo-label.\n\ndo you have results or each organ?\ne.g. using pseudo for spleen only, LB score improve by xxx.\n\nthanks!",
      "votes": null
    },
    {
      "id": "1899526",
      "postDate": "08/15/2022 10:24:52",
      "content": "<p>\"do you use previous label for kidney/colon or do you relabel using pseudo-label.\"</p>\n<p>yes</p>\n<p>\"do you have results or each organ?\"</p>\n<p>nope</p>",
      "rawMarkdown": "\"do you use previous label for kidney/colon or do you relabel using pseudo-label.\"\n\nyes\n\n\"do you have results or each organ?\"\n\nnope",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1880012,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/01/2022 12:09:59",
      "content": "<p>the way to do is:<br>\n1) remove akggle annotation from N spleen training images<br>\n2) replace with hand-annotated the N images</p>\n<p>using the original kaggle annotation as ground truth, measure the dice loss of your annotation<br>\nthis is the quality of your annotation</p>\n<p>(in fact if you hand annotate the same objects T times on different days, you will find your annotation is different)</p>\n<p>train 2 model with (1) and (2), and maybe (1)+(2). This measures the how your label quality (label noise) affects dice score</p>\n<p>you have to think of way to use (2) to get better or smiliar results than (1).<br>\nif there is no way to improve results using your hand annotation, you don't have to try hand-label external data at all. </p>\n<p>instead, you should use pesudo label, or self-supervised</p>\n<hr>\n<p>on a side note, find old segmentation competitions with post submission and test image data that you can download.<br>\nyou will find that even if you submit hand annotation to server:</p>\n<p>model results &gt; hand annotation results<br>\n(for hit objects)</p>\n<p>hand annotation is only better when your model completely missed the target object.<br>\nmodel dice is always better than hand annotation for hit object</p>\n<p>that is why we sometimes don't worry about people playing cheat.</p>\n<hr>\n<p>you can treat yourself as a model. measure your error and ensemble (weighted by error) your results into the rest of the machine models.<br>\ni.e. your labels is only pesudo label and not true labels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1888823,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "08/07/2022 19:44:13",
      "content": "<p>Thanks for doing this experiment and sharing your findings! I think your experience is not unique, I also found that adding external data can actually decrease performance initially.</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1892279,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/10/2022 02:18:51",
          "content": "<p>the object of external data is to improve hidden Hubmap LB score. you should just concentrated on Hubmap LB score first when you make submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1892272,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/10/2022 02:14:33",
      "content": "<p>another possibility is to use auxiliary loss.<br>\ne.g. we cannot segment the image accuracy manually, but we can accuracy annotate the center.<br>\ni called this aux label.</p>\n<p>then we can train the network with kaggle and pesudo center label with auxilay loss.</p>\n<p>other  aux label includes low resolution mask</p>\n<hr>\n<p>a even better solution</p>\n<ol>\n<li>hand label external data + kaggle data  (can be full mask or aux label)</li>\n<li>for kaggle data, you have only kaggle label</li>\n</ol>\n<p>train a model with joint loss (kaggle,external data + your label) and (kaggle data + kaggle label)<br>\nlet's the newtok figure out the relationship between your label and kaggle label itself.</p>\n<hr>\n<p>yet another possibility is to create three class human label:<br>\nbackground (white), unknown(blue), ftu(yellow)</p>\n<p><img src=\"https://i.ibb.co/LNXTsjg/images.png\" alt=\"https://i.ibb.co/LNXTsjg/images.png\"></p>\n<p>for the unknown region, loss in not computed and back propagated.</p>\n<p>to set the unknown region, you can measure the difference of you hand annotation and kaggle true annotation</p>",
      "votes": null,
      "replies": [
        {
          "id": 1892285,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/10/2022 02:33:09",
          "content": "<p>yet another ideal to fix poor pesudo label from hand annotation</p>\n<p><img src=\"https://i.ibb.co/d6djCvX/Selection-093.png\" alt=\"https://i.ibb.co/d6djCvX/Selection-093.png\"><br>\n<a href=\"https://paperswithcode.com/task/interactive-segmentation/\" target=\"_blank\">https://paperswithcode.com/task/interactive-segmentation/</a></p>\n<p>in particular<br>\n<a href=\"https://github.com/navidstuv/NuClick\" target=\"_blank\">https://github.com/navidstuv/NuClick</a><br>\n<img src=\"https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif\" alt=\"https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1892273,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/10/2022 02:15:35",
      "content": "<p>HuBMAP Colonic Crypt  did not decrease my LB and CV score.<br>\nBut it shows not improvement in LB (not enough decimal place revealed)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1892318,
      "author_name": "carnozhao",
      "author_url": "",
      "post_date": "08/10/2022 03:26:56",
      "content": "<p>From my results, HE/PAS stains pseudo-labeling boosted by LB score, with nearly identical CV score. Kidney and colon images come from previous hubmap competition. Spleen and prostate images come from GTEX portal.</p>\n<p>Now, I can get 0.59 hubmap score.</p>\n<p>In further experiments, how to balance between HPA-trained models and HPA/External-trained models is also important to get a good LB score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1892554,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/10/2022 06:48:16",
          "content": "<p>there is another source for prostate images:<br>\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment\" target=\"_blank\">https://www.kaggle.com/c/prostate-cancer-grade-assessment</a></p>\n<p>but you need to select non cancer tissue</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1892557,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/10/2022 06:50:16",
          "content": "<p>\"In further experiments, how to balance …  important to get a good LB score.\"</p>\n<p>last year hubmap competition, this is quite some shakeup.  we need to prevent the same for here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1894850,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/11/2022 18:42:31",
          "content": "<p>\"In further experiments, how to balance … important to get a good LB score.\"</p>\n<p>one possible training dynamics is as follows. you can verify it using your learning process.</p>\n<p><a href=\"https://ibb.co/KcPSWHw\"><img src=\"https://i.ibb.co/CpcDQF2/Selection-065.png\" alt=\"Selection-065\"></a><br>\n<a href=\"https://ibb.co/0m3PTyM\"><img src=\"https://i.ibb.co/djynF42/Selection-073.png\" alt=\"Selection-073\"></a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1896727,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/13/2022 05:15:07",
          "content": "<p>related:<br>\n\"Adaptive Early-Learning Correction for Segmentation from Noisy Annotations\" - cvpr 2022</p>\n<p>\"we propose to update the annotations corresponding to different categories at different<br>\ntimes by detecting when early learning has occurred and memorization is about to begin using the training performance<br>\nof the model\"</p>\n<p>update the annotations  = pesudo label<br>\ndetecting when early learning has occurred = best point of generalisation<br>\n memorization  = overfitting</p>\n<hr>\n<p><a href=\"https://openaccess.thecvf.com/content/CVPR2022/papers/Liu_Adaptive_Early-Learning_Correction_for_Segmentation_From_Noisy_Annotations_CVPR_2022_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/CVPR2022/papers/Liu_Adaptive_Early-Learning_Correction_for_Segmentation_From_Noisy_Annotations_CVPR_2022_paper.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1899507,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/15/2022 10:00:41",
          "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> <br>\n\"pseudo-labeling boosted by LB score\"</p>\n<p>do you use previous label for kidney/colon or do you relabel using pseudo-label.</p>\n<p>do you have results or each organ?<br>\ne.g. using pseudo for spleen only, LB score improve by xxx.</p>\n<p>thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1899526,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "08/15/2022 10:24:52",
          "content": "<p>\"do you use previous label for kidney/colon or do you relabel using pseudo-label.\"</p>\n<p>yes</p>\n<p>\"do you have results or each organ?\"</p>\n<p>nope</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1879962": "I did this experiment on H&E stained spleen images. I learned about functional tissue units of spleen from HPA dictionary and tried to annotate some images by myself.\n\nhttps://www.proteinatlas.org/learn/dictionary/normal/spleen\n\nI annotated the test image, the image from HPA dictionary and 3 images from [histology.medicine.umich.edu](https://histology.medicine.umich.edu/resources/lymphatic-system#iv-spleen). I added newly annotated images to training set. My validation and lb score decreased after using them. I thought okay I probably messed up the annotations. I tried using HuBMAP Colonic Crypt in my training set this time. There are 7 PAS stained colon images in that dataset. My validation and lb score decreased again.\n\nWhat's your experience with using external data?",
    "1880012": "the way to do is:\n1) remove akggle annotation from N spleen training images\n2) replace with hand-annotated the N images\n\nusing the original kaggle annotation as ground truth, measure the dice loss of your annotation\nthis is the quality of your annotation\n\n(in fact if you hand annotate the same objects T times on different days, you will find your annotation is different)\n\ntrain 2 model with (1) and (2), and maybe (1)+(2). This measures the how your label quality (label noise) affects dice score\n\nyou have to think of way to use (2) to get better or smiliar results than (1).\nif there is no way to improve results using your hand annotation, you don't have to try hand-label external data at all. \n\n\ninstead, you should use pesudo label, or self-supervised\n\n---\n\non a side note, find old segmentation competitions with post submission and test image data that you can download.\nyou will find that even if you submit hand annotation to server:\n\nmodel results > hand annotation results\n(for hit objects)\n\nhand annotation is only better when your model completely missed the target object.\nmodel dice is always better than hand annotation for hit object\n\nthat is why we sometimes don't worry about people playing cheat.\n\n---\n\nyou can treat yourself as a model. measure your error and ensemble (weighted by error) your results into the rest of the machine models.\ni.e. your labels is only pesudo label and not true labels.",
    "1888823": "Thanks for doing this experiment and sharing your findings! I think your experience is not unique, I also found that adding external data can actually decrease performance initially.\n\n\nThe Devastator.",
    "1892272": "another possibility is to use auxiliary loss.\ne.g. we cannot segment the image accuracy manually, but we can accuracy annotate the center.\ni called this aux label.\n\nthen we can train the network with kaggle and pesudo center label with auxilay loss.\n\nother  aux label includes low resolution mask\n\n\n---\na even better solution\n\n1. hand label external data + kaggle data  (can be full mask or aux label)\n2. for kaggle data, you have only kaggle label\n\ntrain a model with joint loss (kaggle,external data + your label) and (kaggle data + kaggle label)\nlet's the newtok figure out the relationship between your label and kaggle label itself.\n\n---\n\n\nyet another possibility is to create three class human label:\nbackground (white), unknown(blue), ftu(yellow)\n\n![https://i.ibb.co/LNXTsjg/images.png](https://i.ibb.co/LNXTsjg/images.png)\n\nfor the unknown region, loss in not computed and back propagated.\n\nto set the unknown region, you can measure the difference of you hand annotation and kaggle true annotation",
    "1892273": "HuBMAP Colonic Crypt  did not decrease my LB and CV score.\nBut it shows not improvement in LB (not enough decimal place revealed)",
    "1892279": "the object of external data is to improve hidden Hubmap LB score. you should just concentrated on Hubmap LB score first when you make submission.",
    "1892285": "yet another ideal to fix poor pesudo label from hand annotation\n\n![https://i.ibb.co/d6djCvX/Selection-093.png](https://i.ibb.co/d6djCvX/Selection-093.png)\nhttps://paperswithcode.com/task/interactive-segmentation/\n\nin particular\nhttps://github.com/navidstuv/NuClick\n![https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif](https://raw.githubusercontent.com/navidstuv/NuClick/master/gifs/22.gif)",
    "1892318": "From my results, HE/PAS stains pseudo-labeling boosted by LB score, with nearly identical CV score. Kidney and colon images come from previous hubmap competition. Spleen and prostate images come from GTEX portal.\n\nNow, I can get 0.59 hubmap score.\n\nIn further experiments, how to balance between HPA-trained models and HPA/External-trained models is also important to get a good LB score.",
    "1892554": "there is another source for prostate images:\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment\n\nbut you need to select non cancer tissue",
    "1892557": "\"In further experiments, how to balance ...  important to get a good LB score.\"\n\nlast year hubmap competition, this is quite some shakeup.  we need to prevent the same for here.",
    "1894850": "\"In further experiments, how to balance … important to get a good LB score.\"\n\none possible training dynamics is as follows. you can verify it using your learning process.\n\n\n<a href=\"https://ibb.co/KcPSWHw\"><img src=\"https://i.ibb.co/CpcDQF2/Selection-065.png\" alt=\"Selection-065\" border=\"0\"></a>\n<a href=\"https://ibb.co/0m3PTyM\"><img src=\"https://i.ibb.co/djynF42/Selection-073.png\" alt=\"Selection-073\" border=\"0\"></a>",
    "1896727": "related:\n\"Adaptive Early-Learning Correction for Segmentation from Noisy Annotations\" - cvpr 2022\n\n\"we propose to update the annotations corresponding to different categories at different\ntimes by detecting when early learning has occurred and memorization is about to begin using the training performance\nof the model\"\n\nupdate the annotations  = pesudo label\ndetecting when early learning has occurred = best point of generalisation\n memorization  = overfitting\n\n---\n\nhttps://openaccess.thecvf.com/content/CVPR2022/papers/Liu_Adaptive_Early-Learning_Correction_for_Segmentation_From_Noisy_Annotations_CVPR_2022_paper.pdf",
    "1899507": "carnozhao \n\"pseudo-labeling boosted by LB score\"\n\ndo you use previous label for kidney/colon or do you relabel using pseudo-label.\n\ndo you have results or each organ?\ne.g. using pseudo for spleen only, LB score improve by xxx.\n\nthanks!",
    "1899526": "\"do you use previous label for kidney/colon or do you relabel using pseudo-label.\"\n\nyes\n\n\"do you have results or each organ?\"\n\nnope"
  },
  "source": "meta"
}