{
  "id": 344589,
  "title": "How to use unlabelled external data?",
  "url": "/competitions/hubmap-organ-segmentation/discussion/344589",
  "author_name": "",
  "post_date": "2022-08-15T19:34:25.646289700Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I've never done this before, but it would be really cool if someone can tell me how exactly to use unlabelled data in this competition? Thanks</p>",
  "messages": [
    {
      "id": "1900213",
      "postDate": "08/15/2022 19:34:25",
      "content": "<p>I've never done this before, but it would be really cool if someone can tell me how exactly to use unlabelled data in this competition? Thanks</p>",
      "rawMarkdown": "I've never done this before, but it would be really cool if someone can tell me how exactly to use unlabelled data in this competition? Thanks",
      "votes": null
    },
    {
      "id": "1900238",
      "postDate": "08/15/2022 20:18:01",
      "content": "<p>There are many ways, probably the best way is that you label yourself with the software…</p>\n<p>Also, there is a lazy way to do this… You train your model with original dataset provided by the competition, then you take the best model that you have trained and make predictions on unlabeled data… But you have to still check manually if the model did a good job and correct his predictions to be 100% sure that his job is ok. If you are an advanced user, you can augment unlabeled images, and make several predictions on the same image, and where predictions overlap there is probably a valid mask (but since there is not so much time left in competition I would not recommend for you to do that).</p>\n<p>You can also make the masking threshold bigger if you get some strange artifacts and get only the most certain results…</p>\n<pre><code>mask = prediction[prediction &gt; 0.90]\n</code></pre>\n<p>I wish you all the best.</p>",
      "rawMarkdown": "There are many ways, probably the best way is that you label yourself with the software...\n\nAlso, there is a lazy way to do this... You train your model with original dataset provided by the competition, then you take the best model that you have trained and make predictions on unlabeled data... But you have to still check manually if the model did a good job and correct his predictions to be 100% sure that his job is ok. If you are an advanced user, you can augment unlabeled images, and make several predictions on the same image, and where predictions overlap there is probably a valid mask (but since there is not so much time left in competition I would not recommend for you to do that).\n\nYou can also make the masking threshold bigger if you get some strange artifacts and get only the most certain results...\n\n```\nmask = prediction[prediction > 0.90]\n```\n\nI wish you all the best.",
      "votes": null
    },
    {
      "id": "1900781",
      "postDate": "08/16/2022 08:46:09",
      "content": "<p>\"still check manually if the model did a good job and correct his predictions\"</p>\n<p>the systematically of doing is:</p>\n<ul>\n<li>create pesudo labels (3 class): foreground background, don't care.</li>\n<li>you have to group them into clusters (using some criteria) : C1, C2, C3 …. CN</li>\n</ul>\n<ol>\n<li>set train set = kaggle train set</li>\n<li>train a base model and make a submission.</li>\n<li>select Ci and add to training set. update model and submit.  <br>\ndid it improve? if yes  train set = train set + Ci</li>\n<li>select a new cluster and repeat</li>\n</ol>\n<p>as an example:</p>\n<p>cluster one:    <br>\nforeground : FTU greater than 30 area and not at the edge of tissue and predict &gt;0.90<br>\nbackground : predict &lt;0.30<br>\ndon't care: other FTU, and predict between 0.30 to 0.90</p>\n<p>cluster two<br>\n…</p>",
      "rawMarkdown": "\"still check manually if the model did a good job and correct his predictions\"\n\nthe systematically of doing is:\n\n- create pesudo labels (3 class): foreground background, don't care.\n- you have to group them into clusters (using some criteria) : C1, C2, C3 .... CN\n\n1. set train set = kaggle train set\n2. train a base model and make a submission.\n3. select Ci and add to training set. update model and submit.  \ndid it improve? if yes  train set = train set + Ci\n4. select a new cluster and repeat\n\nas an example:\n\ncluster one:    \nforeground : FTU greater than 30 area and not at the edge of tissue and predict >0.90\nbackground : predict <0.30\ndon't care: other FTU, and predict between 0.30 to 0.90\n\ncluster two\n...",
      "votes": null
    },
    {
      "id": "1900813",
      "postDate": "08/16/2022 09:17:16",
      "content": "<p>a much better way:</p>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook</a><br>\n<a href=\"https://www.kaggle.com/code/hengck23/playground-for-meta-pseudo-label\" target=\"_blank\">https://www.kaggle.com/code/hengck23/playground-for-meta-pseudo-label</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/322832\" target=\"_blank\">https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/322832</a></p>",
      "rawMarkdown": "a much better way:\n\nhttps://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\nhttps://www.kaggle.com/code/hengck23/playground-for-meta-pseudo-label\n\nhttps://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/322832",
      "votes": null
    },
    {
      "id": "1900822",
      "postDate": "08/16/2022 09:22:24",
      "content": "<p>many thanks for writing back! sure ill give it a try</p>",
      "rawMarkdown": "many thanks for writing back! sure ill give it a try",
      "votes": null
    },
    {
      "id": "1900902",
      "postDate": "08/16/2022 10:50:39",
      "content": "<p>you can make a table like this:</p>\n<p><img src=\"https://i.ibb.co/kJnq5SF/Selection-065.png\" alt=\"https://i.ibb.co/kJnq5SF/Selection-065.png\">.</p>\n<p>this only investigate the quality of pesudo labels.<br>\nit answers question like:</p>\n<ul>\n<li><p>how many % of truly labelled pixels (true +ve and true -ve) are required?</p></li>\n<li><p>if there is pixel label noise, how results are affected. what is the max % of label before method of pesudo collapse</p></li>\n<li><p>how accurate is my heurstics or model in predicting truly labelled pixels? </p></li>\n</ul>\n<hr>\n<p>after that you can investigate the quality of unlabelled images, e.g. you unlabelled set only contains 10%, 20%, 30% 50% of your validation set. finally you use download unlabelled HPA images from the website.<br>\nbut matching the performance of the using the unloaded images, you can even estimate :<br>\ne.g.<br>\n1000 download images = 10% validation set.<br>\n10000 download images = 20% validation set.<br>\netc</p>\n<p>finally you think about domain shift.</p>",
      "rawMarkdown": "you can make a table like this:\n\n![https://i.ibb.co/kJnq5SF/Selection-065.png](https://i.ibb.co/kJnq5SF/Selection-065.png).\n\nthis only investigate the quality of pesudo labels.\nit answers question like:\n- how many % of truly labelled pixels (true +ve and true -ve) are required?\n- if there is pixel label noise, how results are affected. what is the max % of label before method of pesudo collapse\n\n- how accurate is my heurstics or model in predicting truly labelled pixels? \n\n\n---\n\n\nafter that you can investigate the quality of unlabelled images, e.g. you unlabelled set only contains 10%, 20%, 30% 50% of your validation set. finally you use download unlabelled HPA images from the website.\nbut matching the performance of the using the unloaded images, you can even estimate :\ne.g.\n1000 download images = 10% validation set.\n10000 download images = 20% validation set.\netc\n\n\nfinally you think about domain shift.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1900238,
      "author_name": "urosjarc",
      "author_url": "",
      "post_date": "08/15/2022 20:18:01",
      "content": "<p>There are many ways, probably the best way is that you label yourself with the software…</p>\n<p>Also, there is a lazy way to do this… You train your model with original dataset provided by the competition, then you take the best model that you have trained and make predictions on unlabeled data… But you have to still check manually if the model did a good job and correct his predictions to be 100% sure that his job is ok. If you are an advanced user, you can augment unlabeled images, and make several predictions on the same image, and where predictions overlap there is probably a valid mask (but since there is not so much time left in competition I would not recommend for you to do that).</p>\n<p>You can also make the masking threshold bigger if you get some strange artifacts and get only the most certain results…</p>\n<pre><code>mask = prediction[prediction &gt; 0.90]\n</code></pre>\n<p>I wish you all the best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1900781,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/16/2022 08:46:09",
          "content": "<p>\"still check manually if the model did a good job and correct his predictions\"</p>\n<p>the systematically of doing is:</p>\n<ul>\n<li>create pesudo labels (3 class): foreground background, don't care.</li>\n<li>you have to group them into clusters (using some criteria) : C1, C2, C3 …. CN</li>\n</ul>\n<ol>\n<li>set train set = kaggle train set</li>\n<li>train a base model and make a submission.</li>\n<li>select Ci and add to training set. update model and submit.  <br>\ndid it improve? if yes  train set = train set + Ci</li>\n<li>select a new cluster and repeat</li>\n</ol>\n<p>as an example:</p>\n<p>cluster one:    <br>\nforeground : FTU greater than 30 area and not at the edge of tissue and predict &gt;0.90<br>\nbackground : predict &lt;0.30<br>\ndon't care: other FTU, and predict between 0.30 to 0.90</p>\n<p>cluster two<br>\n…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900813,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/16/2022 09:17:16",
          "content": "<p>a much better way:</p>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook</a><br>\n<a href=\"https://www.kaggle.com/code/hengck23/playground-for-meta-pseudo-label\" target=\"_blank\">https://www.kaggle.com/code/hengck23/playground-for-meta-pseudo-label</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/322832\" target=\"_blank\">https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/322832</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900822,
          "author_name": "sparkyjunior",
          "author_url": "",
          "post_date": "08/16/2022 09:22:24",
          "content": "<p>many thanks for writing back! sure ill give it a try</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900902,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/16/2022 10:50:39",
          "content": "<p>you can make a table like this:</p>\n<p><img src=\"https://i.ibb.co/kJnq5SF/Selection-065.png\" alt=\"https://i.ibb.co/kJnq5SF/Selection-065.png\">.</p>\n<p>this only investigate the quality of pesudo labels.<br>\nit answers question like:</p>\n<ul>\n<li><p>how many % of truly labelled pixels (true +ve and true -ve) are required?</p></li>\n<li><p>if there is pixel label noise, how results are affected. what is the max % of label before method of pesudo collapse</p></li>\n<li><p>how accurate is my heurstics or model in predicting truly labelled pixels? </p></li>\n</ul>\n<hr>\n<p>after that you can investigate the quality of unlabelled images, e.g. you unlabelled set only contains 10%, 20%, 30% 50% of your validation set. finally you use download unlabelled HPA images from the website.<br>\nbut matching the performance of the using the unloaded images, you can even estimate :<br>\ne.g.<br>\n1000 download images = 10% validation set.<br>\n10000 download images = 20% validation set.<br>\netc</p>\n<p>finally you think about domain shift.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1900213": "I've never done this before, but it would be really cool if someone can tell me how exactly to use unlabelled data in this competition? Thanks",
    "1900238": "There are many ways, probably the best way is that you label yourself with the software...\n\nAlso, there is a lazy way to do this... You train your model with original dataset provided by the competition, then you take the best model that you have trained and make predictions on unlabeled data... But you have to still check manually if the model did a good job and correct his predictions to be 100% sure that his job is ok. If you are an advanced user, you can augment unlabeled images, and make several predictions on the same image, and where predictions overlap there is probably a valid mask (but since there is not so much time left in competition I would not recommend for you to do that).\n\nYou can also make the masking threshold bigger if you get some strange artifacts and get only the most certain results...\n\n```\nmask = prediction[prediction > 0.90]\n```\n\nI wish you all the best.",
    "1900781": "\"still check manually if the model did a good job and correct his predictions\"\n\nthe systematically of doing is:\n\n- create pesudo labels (3 class): foreground background, don't care.\n- you have to group them into clusters (using some criteria) : C1, C2, C3 .... CN\n\n1. set train set = kaggle train set\n2. train a base model and make a submission.\n3. select Ci and add to training set. update model and submit.  \ndid it improve? if yes  train set = train set + Ci\n4. select a new cluster and repeat\n\nas an example:\n\ncluster one:    \nforeground : FTU greater than 30 area and not at the edge of tissue and predict >0.90\nbackground : predict <0.30\ndon't care: other FTU, and predict between 0.30 to 0.90\n\ncluster two\n...",
    "1900813": "a much better way:\n\nhttps://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\nhttps://www.kaggle.com/code/hengck23/playground-for-meta-pseudo-label\n\nhttps://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/322832",
    "1900822": "many thanks for writing back! sure ill give it a try",
    "1900902": "you can make a table like this:\n\n![https://i.ibb.co/kJnq5SF/Selection-065.png](https://i.ibb.co/kJnq5SF/Selection-065.png).\n\nthis only investigate the quality of pesudo labels.\nit answers question like:\n- how many % of truly labelled pixels (true +ve and true -ve) are required?\n- if there is pixel label noise, how results are affected. what is the max % of label before method of pesudo collapse\n\n- how accurate is my heurstics or model in predicting truly labelled pixels? \n\n\n---\n\n\nafter that you can investigate the quality of unlabelled images, e.g. you unlabelled set only contains 10%, 20%, 30% 50% of your validation set. finally you use download unlabelled HPA images from the website.\nbut matching the performance of the using the unloaded images, you can even estimate :\ne.g.\n1000 download images = 10% validation set.\n10000 download images = 20% validation set.\netc\n\n\nfinally you think about domain shift."
  },
  "source": "meta"
}