{
  "id": 202694,
  "title": "Does this competition work?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/202694",
  "author_name": "",
  "post_date": "2020-12-11T11:54:34.490312700Z",
  "votes": 57,
  "comment_count": 11,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI'm wondering if I should continue this competition since the test data is exposed.<br>\nWill the host gain any useful insight in this case?(Does the host want to know how to find the leak?)</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816\" target=\"_blank\">why your cv is not achieving 0.9+</a></li>\n<li><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201248#1101458\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201248#1101458</a><ul>\n<li>private test data may also be exposed</li>\n<li>In addition, the annotation of this task can be partially annotated by non-experts (i.e. anyone)</li>\n<li>Thus everyone can get a private test annotation</li></ul></li>\n</ul>\n<h2>What I want to host</h2>\n<ul>\n<li>New test dataset(participants cannot access)<ul>\n<li>I know this is very difficult… :(</li>\n<li>However, this competition will not work on the test dataset that can be accessed</li></ul></li>\n<li>Do not publish public/private test data<ul>\n<li>This competition is a <code>Code Competition</code></li>\n<li>In <code>Code Competition</code>, participants get a lot of pain for the host</li>\n<li>I think one of the advantages for the participants is that no one can see the test data (public and private)<ul>\n<li>This makes for a more fair competition</li></ul></li></ul></li>\n<li>Clarifying the process of data generation<ul>\n<li>Especially for medical data, the generation process can be important (e.g. <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment\" target=\"_blank\">PANDA</a>)</li></ul></li>\n</ul>\n<p>I know it's a tough time for host right now, but we look forward to host's answers.</p>",
  "messages": [
    {
      "id": "1109184",
      "postDate": "12/11/2020 11:54:34",
      "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI'm wondering if I should continue this competition since the test data is exposed.<br>\nWill the host gain any useful insight in this case?(Does the host want to know how to find the leak?)</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816\" target=\"_blank\">why your cv is not achieving 0.9+</a></li>\n<li><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201248#1101458\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201248#1101458</a><ul>\n<li>private test data may also be exposed</li>\n<li>In addition, the annotation of this task can be partially annotated by non-experts (i.e. anyone)</li>\n<li>Thus everyone can get a private test annotation</li></ul></li>\n</ul>\n<h2>What I want to host</h2>\n<ul>\n<li>New test dataset(participants cannot access)<ul>\n<li>I know this is very difficult… :(</li>\n<li>However, this competition will not work on the test dataset that can be accessed</li></ul></li>\n<li>Do not publish public/private test data<ul>\n<li>This competition is a <code>Code Competition</code></li>\n<li>In <code>Code Competition</code>, participants get a lot of pain for the host</li>\n<li>I think one of the advantages for the participants is that no one can see the test data (public and private)<ul>\n<li>This makes for a more fair competition</li></ul></li></ul></li>\n<li>Clarifying the process of data generation<ul>\n<li>Especially for medical data, the generation process can be important (e.g. <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment\" target=\"_blank\">PANDA</a>)</li></ul></li>\n</ul>\n<p>I know it's a tough time for host right now, but we look forward to host's answers.</p>",
      "rawMarkdown": "philculliton \nI'm wondering if I should continue this competition since the test data is exposed.\nWill the host gain any useful insight in this case?(Does the host want to know how to find the leak?)\n\n- [why your cv is not achieving 0.9+](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816)\n- https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201248#1101458\n    - private test data may also be exposed\n    - In addition, the annotation of this task can be partially annotated by non-experts (i.e. anyone)\n    - Thus everyone can get a private test annotation\n\n## What I want to host\n\n- New test dataset(participants cannot access)\n    - I know this is very difficult... :(\n    - However, this competition will not work on the test dataset that can be accessed\n- Do not publish public/private test data\n    - This competition is a `Code Competition`\n    - In `Code Competition`, participants get a lot of pain for the host\n    - I think one of the advantages for the participants is that no one can see the test data (public and private)\n        - This makes for a more fair competition\n- Clarifying the process of data generation\n   - Especially for medical data, the generation process can be important (e.g. [PANDA](https://www.kaggle.com/c/prostate-cancer-grade-assessment))\n\n\nI know it's a tough time for host right now, but we look forward to host's answers.",
      "votes": null
    },
    {
      "id": "1109545",
      "postDate": "12/11/2020 20:04:46",
      "content": "<p>Hi all,</p>\n<p>Please note that we (the Kaggle team) have been in contact with the HuBMAP team to best understand next steps, interpretation of the rules (e.g. labeling external data), etc. For now, as the private leaderboard data is still unknown (even if some of the images may be available online, unlabeled), there is no way to probe the leaderboard or identify the ground truth for those labels, and we believe that continuing to train models and make submissions is still beneficial to you.</p>\n<p>Expect an official response from the host team shortly regarding next steps. Current discussions have included the <em>possibility</em> of new/additional data (which would include a private leaderboard re-run), updated labels, and/or an updated deadline.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi all,\n\nPlease note that we (the Kaggle team) have been in contact with the HuBMAP team to best understand next steps, interpretation of the rules (e.g. labeling external data), etc. For now, as the private leaderboard data is still unknown (even if some of the images may be available online, unlabeled), there is no way to probe the leaderboard or identify the ground truth for those labels, and we believe that continuing to train models and make submissions is still beneficial to you.\n\nExpect an official response from the host team shortly regarding next steps. Current discussions have included the *possibility* of new/additional data (which would include a private leaderboard re-run), updated labels, and/or an updated deadline.\n\nThanks!",
      "votes": null
    },
    {
      "id": "1109771",
      "postDate": "12/12/2020 03:36:14",
      "content": "<p>Thanks for the sincere answer!</p>\n<p>I'll continue to participate in the competition with the hope that additional data may be added.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Thanks for the sincere answer!\n\nI'll continue to participate in the competition with the hope that additional data may be added.\n\nThanks!",
      "votes": null
    },
    {
      "id": "1113250",
      "postDate": "12/15/2020 10:21:16",
      "content": "<p>hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> ,any new official updates on this topic.</p>",
      "rawMarkdown": "hi @addisonhoward ,any new official updates on this topic.",
      "votes": null
    },
    {
      "id": "1117933",
      "postDate": "12/18/2020 15:43:20",
      "content": "<p>Dear Teams, <br>\nThank you for addressing the dataset situation. We are working in it actively and will provide a detailed update soon.</p>\n<p>Please do continue to bring your creative ideas to this competition!</p>\n<p>In addition, we would like to clarify that hand-labeling of the data is disqualifying. However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.</p>\n<p>Thank you all for your great work on getting us all the best algorithms for this challenging FTU detection task.</p>\n<p>Sincerely,<br>\nKaty Borner </p>",
      "rawMarkdown": "Dear Teams, \nThank you for addressing the dataset situation. We are working in it actively and will provide a detailed update soon.\n\nPlease do continue to bring your creative ideas to this competition!\n\nIn addition, we would like to clarify that hand-labeling of the data is disqualifying. However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.\n\nThank you all for your great work on getting us all the best algorithms for this challenging FTU detection task.\n\nSincerely,\nKaty Borner",
      "votes": null
    },
    {
      "id": "1118033",
      "postDate": "12/18/2020 17:10:54",
      "content": "<p>Is manually fixing the train dataset annotations allowed?</p>",
      "rawMarkdown": "Is manually fixing the train dataset annotations allowed?",
      "votes": null
    },
    {
      "id": "1118064",
      "postDate": "12/18/2020 17:32:12",
      "content": "<p>As stated above, hand-labeling is not allowed. However, we are aware of the issue and are working to fix it as soon as possible. We will ensure the highest quality training data is available. Bear with us!</p>",
      "rawMarkdown": "As stated above, hand-labeling is not allowed. However, we are aware of the issue and are working to fix it as soon as possible. We will ensure the highest quality training data is available. Bear with us!",
      "votes": null
    },
    {
      "id": "1118212",
      "postDate": "12/18/2020 20:52:53",
      "content": "<p>I am very much interested to know why hand-labeling/correcting examples from training dataset would not be allowed. I hope you can share some insight. If not now, then after the competition. Thx</p>\n<p>Hand labeling/relabeling/correcting training examples is usually perfectly fine and actually encouraged.  Any kaggler willing to point out to some cases where this rule is appropriate and explain why?</p>",
      "rawMarkdown": "I am very much interested to know why hand-labeling/correcting examples from training dataset would not be allowed. I hope you can share some insight. If not now, then after the competition. Thx\n\nHand labeling/relabeling/correcting training examples is usually perfectly fine and actually encouraged.  Any kaggler willing to point out to some cases where this rule is appropriate and explain why?",
      "votes": null
    },
    {
      "id": "1118456",
      "postDate": "12/19/2020 04:55:44",
      "content": "<blockquote>\n  <p>However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.</p>\n</blockquote>\n<p>Is it also allowed to apply pseudo-labeling to the training data to modify the labels?</p>",
      "rawMarkdown": "> However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.\n\nIs it also allowed to apply pseudo-labeling to the training data to modify the labels?",
      "votes": null
    },
    {
      "id": "1118483",
      "postDate": "12/19/2020 05:21:07",
      "content": "<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200070\" target=\"_blank\">As discussed here,</a><br>\nI am disappointed that we are not even allowed to correct annotations that are definitely wrong.  I think finding and ruling out these mistakes is part of data science.<br>\nThe host's announcement can be interpreted to mean that if the newly provided data set contains such mistakes, they cannot be corrected.</p>",
      "rawMarkdown": "[As discussed here,](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200070)\nI am disappointed that we are not even allowed to correct annotations that are definitely wrong.  I think finding and ruling out these mistakes is part of data science.\nThe host's announcement can be interpreted to mean that if the newly provided data set contains such mistakes, they cannot be corrected.",
      "votes": null
    },
    {
      "id": "1118885",
      "postDate": "12/19/2020 13:49:50",
      "content": "<p>to correct label shift without hand labeling:</p>\n<ul>\n<li>train model using augmentation</li>\n<li>after training model, do annotation using a written code as:</li>\n</ul>\n<pre><code>corrected_annotation = truth + shift_best, and\nshift_best =  min loss (predicted, truth+shift) for shift in {-8,-7, ...7,8}\n</code></pre>\n<p>Hence correcting shift needs no human intervention </p>\n<p>--<br>\nfor some wrong annotations, there are not many. i don't think they will affect results.<br>\nhowever, if you still want to remove them, they can be filtered away from the model predicted score, e.g.</p>\n<ul>\n<li>train model using augmentation</li>\n<li>then select top annotations, e.g.  95% of the highest prediction score.</li>\n</ul>\n<p>(e.g. <a href=\"http://boqinggong.info/papers/wacv18.pdf\" target=\"_blank\">http://boqinggong.info/papers/wacv18.pdf</a> : A Semi-Supervised Two-Stage Approach to Learning from Noisy Labels, as in  \"apply pseudo-labeling to the training data to modify the labels\" as mentioned by <a href=\"https://www.kaggle.com/sinpcw\" target=\"_blank\">@sinpcw</a>)  </p>\n<hr>\n<p>bottom line, treat this as a slightly noisy label problem. use machine learning to solve this \"noisy label \" issue.</p>\n<p>(there are applications where you \"cannot see the label\" but know that there is some noise.)</p>",
      "rawMarkdown": "to correct label shift without hand labeling:\n- train model using augmentation\n- after training model, do annotation using a written code as:\n\n```\n\ncorrected_annotation = truth + shift_best, and\nshift_best =  min loss (predicted, truth+shift) for shift in {-8,-7, ...7,8}\n\n```\nHence correcting shift needs no human intervention \n\n\n--\nfor some wrong annotations, there are not many. i don't think they will affect results.\nhowever, if you still want to remove them, they can be filtered away from the model predicted score, e.g.\n- train model using augmentation\n- then select top annotations, e.g.  95% of the highest prediction score.\n\n(e.g. http://boqinggong.info/papers/wacv18.pdf : A Semi-Supervised Two-Stage Approach to Learning from Noisy Labels, as in  \"apply pseudo-labeling to the training data to modify the labels\" as mentioned by @sinpcw)  \n\n\n\n---\n\nbottom line, treat this as a slightly noisy label problem. use machine learning to solve this \"noisy label \" issue.\n\n(there are applications where you \"cannot see the label\" but know that there is some noise.)",
      "votes": null
    },
    {
      "id": "1207729",
      "postDate": "02/18/2021 02:12:54",
      "content": "<p>Hello,is this game have a hidden test set?</p>",
      "rawMarkdown": "Hello,is this game have a hidden test set?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1109545,
      "author_name": "addisonhoward",
      "author_url": "",
      "post_date": "12/11/2020 20:04:46",
      "content": "<p>Hi all,</p>\n<p>Please note that we (the Kaggle team) have been in contact with the HuBMAP team to best understand next steps, interpretation of the rules (e.g. labeling external data), etc. For now, as the private leaderboard data is still unknown (even if some of the images may be available online, unlabeled), there is no way to probe the leaderboard or identify the ground truth for those labels, and we believe that continuing to train models and make submissions is still beneficial to you.</p>\n<p>Expect an official response from the host team shortly regarding next steps. Current discussions have included the <em>possibility</em> of new/additional data (which would include a private leaderboard re-run), updated labels, and/or an updated deadline.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109771,
          "author_name": "yukkyo",
          "author_url": "",
          "post_date": "12/12/2020 03:36:14",
          "content": "<p>Thanks for the sincere answer!</p>\n<p>I'll continue to participate in the competition with the hope that additional data may be added.</p>\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1113250,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "12/15/2020 10:21:16",
          "content": "<p>hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> ,any new official updates on this topic.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1117933,
      "author_name": "katyborner",
      "author_url": "",
      "post_date": "12/18/2020 15:43:20",
      "content": "<p>Dear Teams, <br>\nThank you for addressing the dataset situation. We are working in it actively and will provide a detailed update soon.</p>\n<p>Please do continue to bring your creative ideas to this competition!</p>\n<p>In addition, we would like to clarify that hand-labeling of the data is disqualifying. However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.</p>\n<p>Thank you all for your great work on getting us all the best algorithms for this challenging FTU detection task.</p>\n<p>Sincerely,<br>\nKaty Borner </p>",
      "votes": null,
      "replies": [
        {
          "id": 1118033,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "12/18/2020 17:10:54",
          "content": "<p>Is manually fixing the train dataset annotations allowed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118064,
          "author_name": "leahscherschel",
          "author_url": "",
          "post_date": "12/18/2020 17:32:12",
          "content": "<p>As stated above, hand-labeling is not allowed. However, we are aware of the issue and are working to fix it as soon as possible. We will ensure the highest quality training data is available. Bear with us!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118212,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "12/18/2020 20:52:53",
          "content": "<p>I am very much interested to know why hand-labeling/correcting examples from training dataset would not be allowed. I hope you can share some insight. If not now, then after the competition. Thx</p>\n<p>Hand labeling/relabeling/correcting training examples is usually perfectly fine and actually encouraged.  Any kaggler willing to point out to some cases where this rule is appropriate and explain why?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118456,
          "author_name": "sinpcw",
          "author_url": "",
          "post_date": "12/19/2020 04:55:44",
          "content": "<blockquote>\n  <p>However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.</p>\n</blockquote>\n<p>Is it also allowed to apply pseudo-labeling to the training data to modify the labels?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118885,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/19/2020 13:49:50",
          "content": "<p>to correct label shift without hand labeling:</p>\n<ul>\n<li>train model using augmentation</li>\n<li>after training model, do annotation using a written code as:</li>\n</ul>\n<pre><code>corrected_annotation = truth + shift_best, and\nshift_best =  min loss (predicted, truth+shift) for shift in {-8,-7, ...7,8}\n</code></pre>\n<p>Hence correcting shift needs no human intervention </p>\n<p>--<br>\nfor some wrong annotations, there are not many. i don't think they will affect results.<br>\nhowever, if you still want to remove them, they can be filtered away from the model predicted score, e.g.</p>\n<ul>\n<li>train model using augmentation</li>\n<li>then select top annotations, e.g.  95% of the highest prediction score.</li>\n</ul>\n<p>(e.g. <a href=\"http://boqinggong.info/papers/wacv18.pdf\" target=\"_blank\">http://boqinggong.info/papers/wacv18.pdf</a> : A Semi-Supervised Two-Stage Approach to Learning from Noisy Labels, as in  \"apply pseudo-labeling to the training data to modify the labels\" as mentioned by <a href=\"https://www.kaggle.com/sinpcw\" target=\"_blank\">@sinpcw</a>)  </p>\n<hr>\n<p>bottom line, treat this as a slightly noisy label problem. use machine learning to solve this \"noisy label \" issue.</p>\n<p>(there are applications where you \"cannot see the label\" but know that there is some noise.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1118483,
      "author_name": "kashiwagi",
      "author_url": "",
      "post_date": "12/19/2020 05:21:07",
      "content": "<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200070\" target=\"_blank\">As discussed here,</a><br>\nI am disappointed that we are not even allowed to correct annotations that are definitely wrong.  I think finding and ruling out these mistakes is part of data science.<br>\nThe host's announcement can be interpreted to mean that if the newly provided data set contains such mistakes, they cannot be corrected.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207729,
      "author_name": "zekunn",
      "author_url": "",
      "post_date": "02/18/2021 02:12:54",
      "content": "<p>Hello,is this game have a hidden test set?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1109184": "philculliton \nI'm wondering if I should continue this competition since the test data is exposed.\nWill the host gain any useful insight in this case?(Does the host want to know how to find the leak?)\n\n- [why your cv is not achieving 0.9+](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816)\n- https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201248#1101458\n    - private test data may also be exposed\n    - In addition, the annotation of this task can be partially annotated by non-experts (i.e. anyone)\n    - Thus everyone can get a private test annotation\n\n## What I want to host\n\n- New test dataset(participants cannot access)\n    - I know this is very difficult... :(\n    - However, this competition will not work on the test dataset that can be accessed\n- Do not publish public/private test data\n    - This competition is a `Code Competition`\n    - In `Code Competition`, participants get a lot of pain for the host\n    - I think one of the advantages for the participants is that no one can see the test data (public and private)\n        - This makes for a more fair competition\n- Clarifying the process of data generation\n   - Especially for medical data, the generation process can be important (e.g. [PANDA](https://www.kaggle.com/c/prostate-cancer-grade-assessment))\n\n\nI know it's a tough time for host right now, but we look forward to host's answers.",
    "1109545": "Hi all,\n\nPlease note that we (the Kaggle team) have been in contact with the HuBMAP team to best understand next steps, interpretation of the rules (e.g. labeling external data), etc. For now, as the private leaderboard data is still unknown (even if some of the images may be available online, unlabeled), there is no way to probe the leaderboard or identify the ground truth for those labels, and we believe that continuing to train models and make submissions is still beneficial to you.\n\nExpect an official response from the host team shortly regarding next steps. Current discussions have included the *possibility* of new/additional data (which would include a private leaderboard re-run), updated labels, and/or an updated deadline.\n\nThanks!",
    "1109771": "Thanks for the sincere answer!\n\nI'll continue to participate in the competition with the hope that additional data may be added.\n\nThanks!",
    "1113250": "hi @addisonhoward ,any new official updates on this topic.",
    "1117933": "Dear Teams, \nThank you for addressing the dataset situation. We are working in it actively and will provide a detailed update soon.\n\nPlease do continue to bring your creative ideas to this competition!\n\nIn addition, we would like to clarify that hand-labeling of the data is disqualifying. However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.\n\nThank you all for your great work on getting us all the best algorithms for this challenging FTU detection task.\n\nSincerely,\nKaty Borner",
    "1118033": "Is manually fixing the train dataset annotations allowed?",
    "1118064": "As stated above, hand-labeling is not allowed. However, we are aware of the issue and are working to fix it as soon as possible. We will ensure the highest quality training data is available. Bear with us!",
    "1118212": "I am very much interested to know why hand-labeling/correcting examples from training dataset would not be allowed. I hope you can share some insight. If not now, then after the competition. Thx\n\nHand labeling/relabeling/correcting training examples is usually perfectly fine and actually encouraged.  Any kaggler willing to point out to some cases where this rule is appropriate and explain why?",
    "1118456": "> However, pseudo-labeling (e.g., using a model to label outside data not in the provided dataset) is allowed.\n\nIs it also allowed to apply pseudo-labeling to the training data to modify the labels?",
    "1118483": "[As discussed here,](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/200070)\nI am disappointed that we are not even allowed to correct annotations that are definitely wrong.  I think finding and ruling out these mistakes is part of data science.\nThe host's announcement can be interpreted to mean that if the newly provided data set contains such mistakes, they cannot be corrected.",
    "1118885": "to correct label shift without hand labeling:\n- train model using augmentation\n- after training model, do annotation using a written code as:\n\n```\n\ncorrected_annotation = truth + shift_best, and\nshift_best =  min loss (predicted, truth+shift) for shift in {-8,-7, ...7,8}\n\n```\nHence correcting shift needs no human intervention \n\n\n--\nfor some wrong annotations, there are not many. i don't think they will affect results.\nhowever, if you still want to remove them, they can be filtered away from the model predicted score, e.g.\n- train model using augmentation\n- then select top annotations, e.g.  95% of the highest prediction score.\n\n(e.g. http://boqinggong.info/papers/wacv18.pdf : A Semi-Supervised Two-Stage Approach to Learning from Noisy Labels, as in  \"apply pseudo-labeling to the training data to modify the labels\" as mentioned by @sinpcw)  \n\n\n\n---\n\nbottom line, treat this as a slightly noisy label problem. use machine learning to solve this \"noisy label \" issue.\n\n(there are applications where you \"cannot see the label\" but know that there is some noise.)",
    "1207729": "Hello,is this game have a hidden test set?"
  },
  "source": "meta"
}