{
  "id": 301493,
  "title": "How  to use unlabeled data?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/301493",
  "author_name": "",
  "post_date": "2022-01-18T02:19:21.015542200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>There are so many unlabeled data, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments</p>",
  "messages": [
    {
      "id": "1653944",
      "postDate": "01/18/2022 02:19:21",
      "content": "<p>There are so many unlabeled data, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments</p>",
      "rawMarkdown": "There are so many unlabeled data, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments",
      "votes": null
    },
    {
      "id": "1654059",
      "postDate": "01/18/2022 05:33:50",
      "content": "<p>Use a model training on the labelled data to label the new data. Then train again on all the data. This is commonly referred to as pseudo labeling or semi-supervised learning. See this discussion <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/162462\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/162462</a> and notebook <a href=\"https://www.kaggle.com/nvnnghia/yolov5-pseudo-labeling\" target=\"_blank\">https://www.kaggle.com/nvnnghia/yolov5-pseudo-labeling</a> for a good intro</p>",
      "rawMarkdown": "Use a model training on the labelled data to label the new data. Then train again on all the data. This is commonly referred to as pseudo labeling or semi-supervised learning. See this discussion https://www.kaggle.com/c/global-wheat-detection/discussion/162462 and notebook https://www.kaggle.com/nvnnghia/yolov5-pseudo-labeling for a good intro",
      "votes": null
    },
    {
      "id": "1654175",
      "postDate": "01/18/2022 08:34:53",
      "content": "<p><code>There are **so many unlabeled data**, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments</code></p>\n<p>Are you sure?</p>\n<p>Quality of dataset is not TOP. We can see 5 groups of problem:</p>\n<ul>\n<li>not labeled starfish - we can see it on some video frames (so detector see it better then GT - it will be FP in score process)</li>\n<li>labeled but not starfish</li>\n<li>bbox incorrectly located (shifted)</li>\n<li>labeled starfish but could be misleading for model (\"quality\" of starfish is low eg. merges with background or is similar to … stone or other object) </li>\n<li>a few starfish on not labelled frames</li>\n</ul>\n<p>A. you can fix this … but if training videos have the same \"problems / labeling biases) … you can decrease score :) … This is competition not business project where you can fix labelling on whole dataset.  You can leave it (not best business solution but in competition ok).<br>\nB. you can look for unannotated frames and add to your training dataset but …. I am afraid it does not change anything (there little variation in the remaining cages - no very new cases for our model).</p>",
      "rawMarkdown": "`There are **so many unlabeled data**, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments`\n\nAre you sure?\n\nQuality of dataset is not TOP. We can see 5 groups of problem:\n- not labeled starfish - we can see it on some video frames (so detector see it better then GT - it will be FP in score process)\n- labeled but not starfish\n- bbox incorrectly located (shifted)\n- labeled starfish but could be misleading for model (\"quality\" of starfish is low eg. merges with background or is similar to ... stone or other object) \n- a few starfish on not labelled frames\n\nA. you can fix this ... but if training videos have the same \"problems / labeling biases) ... you can decrease score :) ... This is competition not business project where you can fix labelling on whole dataset.  You can leave it (not best business solution but in competition ok).\nB. you can look for unannotated frames and add to your training dataset but .... I am afraid it does not change anything (there little variation in the remaining cages - no very new cases for our model).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1654059,
      "author_name": "maxvandijck",
      "author_url": "",
      "post_date": "01/18/2022 05:33:50",
      "content": "<p>Use a model training on the labelled data to label the new data. Then train again on all the data. This is commonly referred to as pseudo labeling or semi-supervised learning. See this discussion <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/162462\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/162462</a> and notebook <a href=\"https://www.kaggle.com/nvnnghia/yolov5-pseudo-labeling\" target=\"_blank\">https://www.kaggle.com/nvnnghia/yolov5-pseudo-labeling</a> for a good intro</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1654175,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "01/18/2022 08:34:53",
      "content": "<p><code>There are **so many unlabeled data**, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments</code></p>\n<p>Are you sure?</p>\n<p>Quality of dataset is not TOP. We can see 5 groups of problem:</p>\n<ul>\n<li>not labeled starfish - we can see it on some video frames (so detector see it better then GT - it will be FP in score process)</li>\n<li>labeled but not starfish</li>\n<li>bbox incorrectly located (shifted)</li>\n<li>labeled starfish but could be misleading for model (\"quality\" of starfish is low eg. merges with background or is similar to … stone or other object) </li>\n<li>a few starfish on not labelled frames</li>\n</ul>\n<p>A. you can fix this … but if training videos have the same \"problems / labeling biases) … you can decrease score :) … This is competition not business project where you can fix labelling on whole dataset.  You can leave it (not best business solution but in competition ok).<br>\nB. you can look for unannotated frames and add to your training dataset but …. I am afraid it does not change anything (there little variation in the remaining cages - no very new cases for our model).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1653944": "There are so many unlabeled data, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments",
    "1654059": "Use a model training on the labelled data to label the new data. Then train again on all the data. This is commonly referred to as pseudo labeling or semi-supervised learning. See this discussion https://www.kaggle.com/c/global-wheat-detection/discussion/162462 and notebook https://www.kaggle.com/nvnnghia/yolov5-pseudo-labeling for a good intro",
    "1654175": "`There are **so many unlabeled data**, some guys said use these data. But i have no idea how to use it. If you have some ideas, please make some comments`\n\nAre you sure?\n\nQuality of dataset is not TOP. We can see 5 groups of problem:\n- not labeled starfish - we can see it on some video frames (so detector see it better then GT - it will be FP in score process)\n- labeled but not starfish\n- bbox incorrectly located (shifted)\n- labeled starfish but could be misleading for model (\"quality\" of starfish is low eg. merges with background or is similar to ... stone or other object) \n- a few starfish on not labelled frames\n\nA. you can fix this ... but if training videos have the same \"problems / labeling biases) ... you can decrease score :) ... This is competition not business project where you can fix labelling on whole dataset.  You can leave it (not best business solution but in competition ok).\nB. you can look for unannotated frames and add to your training dataset but .... I am afraid it does not change anything (there little variation in the remaining cages - no very new cases for our model)."
  },
  "source": "meta"
}