{
  "id": 243804,
  "title": "Regarding Hand Labeling of Test Set Data ",
  "url": "/competitions/siim-covid19-detection/discussion/243804",
  "author_name": "",
  "post_date": "2021-06-04T03:08:16.102145600Z",
  "votes": -8,
  "comment_count": 22,
  "views": 0,
  "content": "<p>The test set includes images from RICORD and BIMCV, which are publicly available.<br>\nThis competition was organized with the intent to advance research in and accuracy <br>\nof COVID-19 detection. The hosts will disqualify if abuse of hand-labeled images <br>\nthat compose the test set is discovered, but we hope that given the benefit to <br>\nbroader humanity that such a solution could provide, participants will act <br>\nhonorably nevertheless.</p>\n<hr>\n<p>Update to above post (6/9/21):</p>\n<p>I realize now that the original post may seem confusing. After discussing internally, hopefully this will help clarify things:</p>\n<p>Both RICORD and BIMCV are public datasets, which also are used in this competition.<br>\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.</p>\n<p>These annotations done by our group are not available to the public and are used in this competition.<br>\nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes. Those annotations are publicly available. There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.</p>\n<p>Also BIMCV has a public dataset that includes segmentations (for a small portion of the data).</p>\n<p>Because the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.</p>\n<p>However, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets. Hope this makes sense and sorry if the above post caused confusion.</p>",
  "messages": [
    {
      "id": "1335131",
      "postDate": "06/04/2021 03:08:16",
      "content": "<p>The test set includes images from RICORD and BIMCV, which are publicly available.<br>\nThis competition was organized with the intent to advance research in and accuracy <br>\nof COVID-19 detection. The hosts will disqualify if abuse of hand-labeled images <br>\nthat compose the test set is discovered, but we hope that given the benefit to <br>\nbroader humanity that such a solution could provide, participants will act <br>\nhonorably nevertheless.</p>\n<hr>\n<p>Update to above post (6/9/21):</p>\n<p>I realize now that the original post may seem confusing. After discussing internally, hopefully this will help clarify things:</p>\n<p>Both RICORD and BIMCV are public datasets, which also are used in this competition.<br>\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.</p>\n<p>These annotations done by our group are not available to the public and are used in this competition.<br>\nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes. Those annotations are publicly available. There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.</p>\n<p>Also BIMCV has a public dataset that includes segmentations (for a small portion of the data).</p>\n<p>Because the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.</p>\n<p>However, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets. Hope this makes sense and sorry if the above post caused confusion.</p>",
      "rawMarkdown": "The test set includes images from RICORD and BIMCV, which are publicly available.\nThis competition was organized with the intent to advance research in and accuracy \nof COVID-19 detection. The hosts will disqualify if abuse of hand-labeled images \nthat compose the test set is discovered, but we hope that given the benefit to \nbroader humanity that such a solution could provide, participants will act \nhonorably nevertheless.\n\n_________________________________________\n\nUpdate to above post (6/9/21):\n\nI realize now that the original post may seem confusing. After discussing internally, hopefully this will help clarify things:\n\nBoth RICORD and BIMCV are public datasets, which also are used in this competition.\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.\n\nThese annotations done by our group are not available to the public and are used in this competition.\nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes. Those annotations are publicly available. There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.\n\nAlso BIMCV has a public dataset that includes segmentations (for a small portion of the data).\n\nBecause the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.\n\nHowever, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets. Hope this makes sense and sorry if the above post caused confusion.",
      "votes": null
    },
    {
      "id": "1335150",
      "postDate": "06/04/2021 03:20:36",
      "content": "<p>Hand-label is prohibited, I understand.<br>\nDoes \"the test set\" which you mentioned has publicly available images include private test set?<br>\nOr only public test set?</p>",
      "rawMarkdown": "Hand-label is prohibited, I understand.\nDoes \"the test set\" which you mentioned has publicly available images include private test set?\nOr only public test set?",
      "votes": null
    },
    {
      "id": "1335173",
      "postDate": "06/04/2021 03:50:47",
      "content": "<p>Publicly available images are present in the public and private test sets. However, the annotations that we performed on the images (typical, indeterminate, atypical, negative for pneumonia, as well as bounding box information) is private and not publicly available. </p>",
      "rawMarkdown": "Publicly available images are present in the public and private test sets. However, the annotations that we performed on the images (typical, indeterminate, atypical, negative for pneumonia, as well as bounding box information) is private and not publicly available.",
      "votes": null
    },
    {
      "id": "1335175",
      "postDate": "06/04/2021 03:53:33",
      "content": "<p>Thanks for the clarification…<br>\nI asked above because this available kaggle dataset is RICORD and has appearance annotations. <br>\n<a href=\"https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests\" target=\"_blank\">https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests</a></p>",
      "rawMarkdown": "Thanks for the clarification...\nI asked above because this available kaggle dataset is RICORD and has appearance annotations. \nhttps://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests",
      "votes": null
    },
    {
      "id": "1335216",
      "postDate": "06/04/2021 05:01:07",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> <br>\nThe competition test data partially contain images from these publicly available datasets? Or is the competition public/private test data built only from images of these publicly available datasets?</p>",
      "rawMarkdown": "paras42 \nThe competition test data partially contain images from these publicly available datasets? Or is the competition public/private test data built only from images of these publicly available datasets?",
      "votes": null
    },
    {
      "id": "1335276",
      "postDate": "06/04/2021 06:02:01",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> </p>\n<pre><code>The hosts will disqualify if abuse of hand-labeled images that compose the test set is discovered\n</code></pre>\n<p>Does this mean that hand labeling of external data is prohibited?<br>\nIt is also possible that a winner may have hand-labeled even if he or she falsely reports that he or she did not hand-label. For example, if the winner is unlikely to be able to reproduce the solution, I think you should disqualify him/her, is this possible?</p>",
      "rawMarkdown": "paras42 \n```\nThe hosts will disqualify if abuse of hand-labeled images that compose the test set is discovered\n```\nDoes this mean that hand labeling of external data is prohibited?\nIt is also possible that a winner may have hand-labeled even if he or she falsely reports that he or she did not hand-label. For example, if the winner is unlikely to be able to reproduce the solution, I think you should disqualify him/her, is this possible?",
      "votes": null
    },
    {
      "id": "1335287",
      "postDate": "06/04/2021 06:09:50",
      "content": "<p>As far as I know, we've to publish the labels for external data if we do hand-labeling.</p>",
      "rawMarkdown": "As far as I know, we've to publish the labels for external data if we do hand-labeling.",
      "votes": null
    },
    {
      "id": "1335475",
      "postDate": "06/04/2021 08:40:34",
      "content": "<p>Does this mean, we can use this dataset or any other with their given annotations but hand labeling on them is not allowed? Or their given annotations fall into a hand-labeling case too? <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> </p>",
      "rawMarkdown": "Does this mean, we can use this dataset or any other with their given annotations but hand labeling on them is not allowed? Or their given annotations fall into a hand-labeling case too? @paras42",
      "votes": null
    },
    {
      "id": "1335478",
      "postDate": "06/04/2021 08:42:15",
      "content": "<p>Are you saying that hand labeling is allowed if the data is made public? I don't think that is the host's opinion. I will wait for the host's reply.</p>",
      "rawMarkdown": "Are you saying that hand labeling is allowed if the data is made public? I don't think that is the host's opinion. I will wait for the host's reply.",
      "votes": null
    },
    {
      "id": "1336221",
      "postDate": "06/04/2021 18:17:08",
      "content": "<p>Can you define \"abuse\"?  What is acceptable and what isn't?</p>\n<p>Whatever you decide is fine, but ambiguity is not fine.  </p>",
      "rawMarkdown": "Can you define \"abuse\"?  What is acceptable and what isn't?\n\nWhatever you decide is fine, but ambiguity is not fine.",
      "votes": null
    },
    {
      "id": "1336266",
      "postDate": "06/04/2021 18:44:26",
      "content": "<pre><code>Are you saying that hand labeling is allowed if the data is made public?\n</code></pre>\n<p>I think this the basic rule for any external dataset. But if we're talking about <strong>private</strong> or <strong>public</strong> test then I think it's <strong>not allowed</strong>.</p>",
      "rawMarkdown": "```\nAre you saying that hand labeling is allowed if the data is made public?\n```\n I think this the basic rule for any external dataset. But if we're talking about **private** or **public** test then I think it's **not allowed**.",
      "votes": null
    },
    {
      "id": "1336443",
      "postDate": "06/04/2021 23:04:34",
      "content": "<blockquote>\n  <p>Does this mean that hand labeling of external data is prohibited?</p>\n</blockquote>\n<p>Hand labeling of test data is prohibited usually.  I have not checked yet for this competition though.</p>\n<p>If true, and if you hand label data that includes test data without knowing it then you risk removal.  This is tricky, to say the least.</p>",
      "rawMarkdown": "> Does this mean that hand labeling of external data is prohibited?\n\nHand labeling of test data is prohibited usually.  I have not checked yet for this competition though.\n\nIf true, and if you hand label data that includes test data without knowing it then you risk removal.  This is tricky, to say the least.",
      "votes": null
    },
    {
      "id": "1336468",
      "postDate": "06/05/2021 00:24:42",
      "content": "<p>I agree with you. I think the host should draw a clear line.</p>\n<p>1, Prohibit hand labeling on all external data.<br>\n2, Prohibit hand labeling on RICORD and BIMCV.<br>\n3, Hand labeling for external is not prohibited, but if hand labeling is applied to the test set as a result, it will be disqualified.<br>\n4, Other</p>\n<p>I hope it is 1.</p>",
      "rawMarkdown": "I agree with you. I think the host should draw a clear line.\n\n1, Prohibit hand labeling on all external data.\n2, Prohibit hand labeling on RICORD and BIMCV.\n3, Hand labeling for external is not prohibited, but if hand labeling is applied to the test set as a result, it will be disqualified.\n4, Other\n\nI hope it is 1.",
      "votes": null
    },
    {
      "id": "1336509",
      "postDate": "06/05/2021 02:00:18",
      "content": "<p>I think <strong>RICORD</strong> dataset has been already posted in the discussion with annotations. Maybe the annotations don't match. But I'm not sure if someone hand-label a small portion of the data, he'll be caught. It's disappointing to see loopholes in such an interesting competition.</p>",
      "rawMarkdown": "I think **RICORD** dataset has been already posted in the discussion with annotations. Maybe the annotations don't match. But I'm not sure if someone hand-label a small portion of the data, he'll be caught. It's disappointing to see loopholes in such an interesting competition.",
      "votes": null
    },
    {
      "id": "1336568",
      "postDate": "06/05/2021 03:45:01",
      "content": "<p>The RICORD dataset posted only has image-level labels. It does not have bounding box information. </p>",
      "rawMarkdown": "The RICORD dataset posted only has image-level labels. It does not have bounding box information.",
      "votes": null
    },
    {
      "id": "1341447",
      "postDate": "06/08/2021 16:40:26",
      "content": "<p>I realize now that the original post may seem confusing.  After discussing internally, hopefully this will help clarify things:</p>\n<p>Both RICORD and BIMCV are public datasets, which also are used in this competition.<br>\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.</p>\n<p>These annotations done by our group are not available to the public and are used in this competition. <br>\nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes.  Those annotations are publicly available.  There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.</p>\n<p>Also BIMCV has a public dataset that includes segmentations (for a small portion of the data).  </p>\n<p>Because the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.  </p>\n<p>However, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets.  Hope this makes sense and sorry if the above post caused confusion.     </p>",
      "rawMarkdown": "I realize now that the original post may seem confusing.  After discussing internally, hopefully this will help clarify things:\n\nBoth RICORD and BIMCV are public datasets, which also are used in this competition.\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.\n\nThese annotations done by our group are not available to the public and are used in this competition. \nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes.  Those annotations are publicly available.  There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.\n\nAlso BIMCV has a public dataset that includes segmentations (for a small portion of the data).  \n\nBecause the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.  \n\nHowever, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets.  Hope this makes sense and sorry if the above post caused confusion.",
      "votes": null
    },
    {
      "id": "1341487",
      "postDate": "06/08/2021 17:12:10",
      "content": "<p>Thanks, this is clear.  Maybe you can also update the original post with this content?</p>",
      "rawMarkdown": "Thanks, this is clear.  Maybe you can also update the original post with this content?",
      "votes": null
    },
    {
      "id": "1341776",
      "postDate": "06/09/2021 01:54:00",
      "content": "<p>Hello, is lb only caculated by the test we can see in the test file?</p>",
      "rawMarkdown": "Hello, is lb only caculated by the test we can see in the test file?",
      "votes": null
    },
    {
      "id": "1343043",
      "postDate": "06/10/2021 01:45:12",
      "content": "<p>That's a good idea and I just updated the original post above.</p>",
      "rawMarkdown": "That's a good idea and I just updated the original post above.",
      "votes": null
    },
    {
      "id": "1347457",
      "postDate": "06/13/2021 08:18:04",
      "content": "<p>I couldn't understand how it's fair to use the test dataset while building models? <br>\nBecause then my model test metrics are biased, I will never know how the model performs to unseen data, which was the competition's motto.<br>\nCan someone explain this?</p>",
      "rawMarkdown": "I couldn't understand how it's fair to use the test dataset while building models? \nBecause then my model test metrics are biased, I will never know how the model performs to unseen data, which was the competition's motto.\nCan someone explain this?",
      "votes": null
    },
    {
      "id": "1347555",
      "postDate": "06/13/2021 10:12:52",
      "content": "<p><a href=\"https://www.kaggle.com/saikalyan9981\" target=\"_blank\">@saikalyan9981</a> You are mixing fairness and effectiveness.</p>\n<p>You are right that if you used ground truth then you cannot know how your model will perform on new data.  That's about effectiveness.</p>\n<p>But if you could use test ground truth when training your model then your model will perform better in test set than if you don't use test ground truth.  That's what is unfair to those who don't use test set at all.</p>\n<p>I don't know why you got downvoted, all questions are good questions, upvoting you!</p>",
      "rawMarkdown": "saikalyan9981 You are mixing fairness and effectiveness.\n\nYou are right that if you used ground truth then you cannot know how your model will perform on new data.  That's about effectiveness.\n\nBut if you could use test ground truth when training your model then your model will perform better in test set than if you don't use test ground truth.  That's what is unfair to those who don't use test set at all.\n\nI don't know why you got downvoted, all questions are good questions, upvoting you!",
      "votes": null
    },
    {
      "id": "1347581",
      "postDate": "06/13/2021 10:36:59",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thank you for being considerate and explaining fairness and effectiveness. <br>\nI have two more questions. Can you please shed your thoughts on them?</p>\n<ol>\n<li>Can the model created using a test set be valuable and trusted by a radiologist for new real-world data? Which was the ideal motto of the competition.</li>\n<li>Can organisers ban using those two datasets (BIMCV and RICORD) or create a new testset instead?</li>\n</ol>",
      "rawMarkdown": "cpmpml Thank you for being considerate and explaining fairness and effectiveness. \nI have two more questions. Can you please shed your thoughts on them?\n1. Can the model created using a test set be valuable and trusted by a radiologist for new real-world data? Which was the ideal motto of the competition.\n2. Can organisers ban using those two datasets (BIMCV and RICORD) or create a new testset instead?",
      "votes": null
    },
    {
      "id": "1348999",
      "postDate": "06/14/2021 12:54:56",
      "content": "<p>I am afraid I can't answer your questions.  I am not a radiologist an I am not part of the organizer's team.  But I think the host answered your second question already, just read the original post again.</p>",
      "rawMarkdown": "I am afraid I can't answer your questions.  I am not a radiologist an I am not part of the organizer's team.  But I think the host answered your second question already, just read the original post again.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335150,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "06/04/2021 03:20:36",
      "content": "<p>Hand-label is prohibited, I understand.<br>\nDoes \"the test set\" which you mentioned has publicly available images include private test set?<br>\nOr only public test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335173,
          "author_name": "paras42",
          "author_url": "",
          "post_date": "06/04/2021 03:50:47",
          "content": "<p>Publicly available images are present in the public and private test sets. However, the annotations that we performed on the images (typical, indeterminate, atypical, negative for pneumonia, as well as bounding box information) is private and not publicly available. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335175,
          "author_name": "drtausamaru",
          "author_url": "",
          "post_date": "06/04/2021 03:53:33",
          "content": "<p>Thanks for the clarification…<br>\nI asked above because this available kaggle dataset is RICORD and has appearance annotations. <br>\n<a href=\"https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests\" target=\"_blank\">https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335475,
          "author_name": "mohammadzunaed",
          "author_url": "",
          "post_date": "06/04/2021 08:40:34",
          "content": "<p>Does this mean, we can use this dataset or any other with their given annotations but hand labeling on them is not allowed? Or their given annotations fall into a hand-labeling case too? <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1341447,
          "author_name": "paras42",
          "author_url": "",
          "post_date": "06/08/2021 16:40:26",
          "content": "<p>I realize now that the original post may seem confusing.  After discussing internally, hopefully this will help clarify things:</p>\n<p>Both RICORD and BIMCV are public datasets, which also are used in this competition.<br>\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.</p>\n<p>These annotations done by our group are not available to the public and are used in this competition. <br>\nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes.  Those annotations are publicly available.  There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.</p>\n<p>Also BIMCV has a public dataset that includes segmentations (for a small portion of the data).  </p>\n<p>Because the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.  </p>\n<p>However, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets.  Hope this makes sense and sorry if the above post caused confusion.     </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1341487,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/08/2021 17:12:10",
          "content": "<p>Thanks, this is clear.  Maybe you can also update the original post with this content?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1343043,
          "author_name": "paras42",
          "author_url": "",
          "post_date": "06/10/2021 01:45:12",
          "content": "<p>That's a good idea and I just updated the original post above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1347457,
          "author_name": "saikalyan9981",
          "author_url": "",
          "post_date": "06/13/2021 08:18:04",
          "content": "<p>I couldn't understand how it's fair to use the test dataset while building models? <br>\nBecause then my model test metrics are biased, I will never know how the model performs to unseen data, which was the competition's motto.<br>\nCan someone explain this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1347555,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/13/2021 10:12:52",
          "content": "<p><a href=\"https://www.kaggle.com/saikalyan9981\" target=\"_blank\">@saikalyan9981</a> You are mixing fairness and effectiveness.</p>\n<p>You are right that if you used ground truth then you cannot know how your model will perform on new data.  That's about effectiveness.</p>\n<p>But if you could use test ground truth when training your model then your model will perform better in test set than if you don't use test ground truth.  That's what is unfair to those who don't use test set at all.</p>\n<p>I don't know why you got downvoted, all questions are good questions, upvoting you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1347581,
          "author_name": "saikalyan9981",
          "author_url": "",
          "post_date": "06/13/2021 10:36:59",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Thank you for being considerate and explaining fairness and effectiveness. <br>\nI have two more questions. Can you please shed your thoughts on them?</p>\n<ol>\n<li>Can the model created using a test set be valuable and trusted by a radiologist for new real-world data? Which was the ideal motto of the competition.</li>\n<li>Can organisers ban using those two datasets (BIMCV and RICORD) or create a new testset instead?</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1348999,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/14/2021 12:54:56",
          "content": "<p>I am afraid I can't answer your questions.  I am not a radiologist an I am not part of the organizer's team.  But I think the host answered your second question already, just read the original post again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335216,
      "author_name": "inoueu1",
      "author_url": "",
      "post_date": "06/04/2021 05:01:07",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> <br>\nThe competition test data partially contain images from these publicly available datasets? Or is the competition public/private test data built only from images of these publicly available datasets?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335276,
      "author_name": "yujiariyasu",
      "author_url": "",
      "post_date": "06/04/2021 06:02:01",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> </p>\n<pre><code>The hosts will disqualify if abuse of hand-labeled images that compose the test set is discovered\n</code></pre>\n<p>Does this mean that hand labeling of external data is prohibited?<br>\nIt is also possible that a winner may have hand-labeled even if he or she falsely reports that he or she did not hand-label. For example, if the winner is unlikely to be able to reproduce the solution, I think you should disqualify him/her, is this possible?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335287,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "06/04/2021 06:09:50",
          "content": "<p>As far as I know, we've to publish the labels for external data if we do hand-labeling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335478,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/04/2021 08:42:15",
          "content": "<p>Are you saying that hand labeling is allowed if the data is made public? I don't think that is the host's opinion. I will wait for the host's reply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336266,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "06/04/2021 18:44:26",
          "content": "<pre><code>Are you saying that hand labeling is allowed if the data is made public?\n</code></pre>\n<p>I think this the basic rule for any external dataset. But if we're talking about <strong>private</strong> or <strong>public</strong> test then I think it's <strong>not allowed</strong>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336443,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2021 23:04:34",
          "content": "<blockquote>\n  <p>Does this mean that hand labeling of external data is prohibited?</p>\n</blockquote>\n<p>Hand labeling of test data is prohibited usually.  I have not checked yet for this competition though.</p>\n<p>If true, and if you hand label data that includes test data without knowing it then you risk removal.  This is tricky, to say the least.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336468,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "06/05/2021 00:24:42",
          "content": "<p>I agree with you. I think the host should draw a clear line.</p>\n<p>1, Prohibit hand labeling on all external data.<br>\n2, Prohibit hand labeling on RICORD and BIMCV.<br>\n3, Hand labeling for external is not prohibited, but if hand labeling is applied to the test set as a result, it will be disqualified.<br>\n4, Other</p>\n<p>I hope it is 1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336509,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "06/05/2021 02:00:18",
          "content": "<p>I think <strong>RICORD</strong> dataset has been already posted in the discussion with annotations. Maybe the annotations don't match. But I'm not sure if someone hand-label a small portion of the data, he'll be caught. It's disappointing to see loopholes in such an interesting competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336568,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "06/05/2021 03:45:01",
          "content": "<p>The RICORD dataset posted only has image-level labels. It does not have bounding box information. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336221,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2021 18:17:08",
      "content": "<p>Can you define \"abuse\"?  What is acceptable and what isn't?</p>\n<p>Whatever you decide is fine, but ambiguity is not fine.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1341776,
      "author_name": "zekunn",
      "author_url": "",
      "post_date": "06/09/2021 01:54:00",
      "content": "<p>Hello, is lb only caculated by the test we can see in the test file?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335131": "The test set includes images from RICORD and BIMCV, which are publicly available.\nThis competition was organized with the intent to advance research in and accuracy \nof COVID-19 detection. The hosts will disqualify if abuse of hand-labeled images \nthat compose the test set is discovered, but we hope that given the benefit to \nbroader humanity that such a solution could provide, participants will act \nhonorably nevertheless.\n\n_________________________________________\n\nUpdate to above post (6/9/21):\n\nI realize now that the original post may seem confusing. After discussing internally, hopefully this will help clarify things:\n\nBoth RICORD and BIMCV are public datasets, which also are used in this competition.\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.\n\nThese annotations done by our group are not available to the public and are used in this competition.\nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes. Those annotations are publicly available. There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.\n\nAlso BIMCV has a public dataset that includes segmentations (for a small portion of the data).\n\nBecause the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.\n\nHowever, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets. Hope this makes sense and sorry if the above post caused confusion.",
    "1335150": "Hand-label is prohibited, I understand.\nDoes \"the test set\" which you mentioned has publicly available images include private test set?\nOr only public test set?",
    "1335173": "Publicly available images are present in the public and private test sets. However, the annotations that we performed on the images (typical, indeterminate, atypical, negative for pneumonia, as well as bounding box information) is private and not publicly available.",
    "1335175": "Thanks for the clarification...\nI asked above because this available kaggle dataset is RICORD and has appearance annotations. \nhttps://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests",
    "1335216": "paras42 \nThe competition test data partially contain images from these publicly available datasets? Or is the competition public/private test data built only from images of these publicly available datasets?",
    "1335276": "paras42 \n```\nThe hosts will disqualify if abuse of hand-labeled images that compose the test set is discovered\n```\nDoes this mean that hand labeling of external data is prohibited?\nIt is also possible that a winner may have hand-labeled even if he or she falsely reports that he or she did not hand-label. For example, if the winner is unlikely to be able to reproduce the solution, I think you should disqualify him/her, is this possible?",
    "1335287": "As far as I know, we've to publish the labels for external data if we do hand-labeling.",
    "1335475": "Does this mean, we can use this dataset or any other with their given annotations but hand labeling on them is not allowed? Or their given annotations fall into a hand-labeling case too? @paras42",
    "1335478": "Are you saying that hand labeling is allowed if the data is made public? I don't think that is the host's opinion. I will wait for the host's reply.",
    "1336221": "Can you define \"abuse\"?  What is acceptable and what isn't?\n\nWhatever you decide is fine, but ambiguity is not fine.",
    "1336266": "```\nAre you saying that hand labeling is allowed if the data is made public?\n```\n I think this the basic rule for any external dataset. But if we're talking about **private** or **public** test then I think it's **not allowed**.",
    "1336443": "> Does this mean that hand labeling of external data is prohibited?\n\nHand labeling of test data is prohibited usually.  I have not checked yet for this competition though.\n\nIf true, and if you hand label data that includes test data without knowing it then you risk removal.  This is tricky, to say the least.",
    "1336468": "I agree with you. I think the host should draw a clear line.\n\n1, Prohibit hand labeling on all external data.\n2, Prohibit hand labeling on RICORD and BIMCV.\n3, Hand labeling for external is not prohibited, but if hand labeling is applied to the test set as a result, it will be disqualified.\n4, Other\n\nI hope it is 1.",
    "1336509": "I think **RICORD** dataset has been already posted in the discussion with annotations. Maybe the annotations don't match. But I'm not sure if someone hand-label a small portion of the data, he'll be caught. It's disappointing to see loopholes in such an interesting competition.",
    "1336568": "The RICORD dataset posted only has image-level labels. It does not have bounding box information.",
    "1341447": "I realize now that the original post may seem confusing.  After discussing internally, hopefully this will help clarify things:\n\nBoth RICORD and BIMCV are public datasets, which also are used in this competition.\nA group of radiologists (22 all-together) from multiple countries annotated the RICORD and BIMCV data with specific annotations at the exam level (typical, indeterminate, atypical pattern of COVID-19, or negative for pneumonia) and also with bounding boxes in cases of pulmonary opacities.\n\nThese annotations done by our group are not available to the public and are used in this competition. \nHowever, another group of radiologists at the RSNA annotated the RICORD dataset using the same methodology for the exam level (e.g. typical, indeterminate, atypical, or negative for pneumonia), but without bounding boxes.  Those annotations are publicly available.  There may be a few cases where our annotators graded the images differently compared to the radiologists that annotated RICORD, due to differences in interpretation. However, some of the grading will be similar or the same.\n\nAlso BIMCV has a public dataset that includes segmentations (for a small portion of the data).  \n\nBecause the goal in the competition is for Kaggler's to develop the best possible model, and with access to the same data for fairness-sake, it is acceptable for the contestants to use the RICORD data with their annotations and also to use BIMCV data with their annotations.  \n\nHowever, what we don’t want is people getting radiologists to relabel the entire BIMCV and RICORD data themselves and training on those the labels, and that is what was meant by the original post regarding hand-labeling the BIMCV and RICORD datasets.  Hope this makes sense and sorry if the above post caused confusion.",
    "1341487": "Thanks, this is clear.  Maybe you can also update the original post with this content?",
    "1341776": "Hello, is lb only caculated by the test we can see in the test file?",
    "1343043": "That's a good idea and I just updated the original post above.",
    "1347457": "I couldn't understand how it's fair to use the test dataset while building models? \nBecause then my model test metrics are biased, I will never know how the model performs to unseen data, which was the competition's motto.\nCan someone explain this?",
    "1347555": "saikalyan9981 You are mixing fairness and effectiveness.\n\nYou are right that if you used ground truth then you cannot know how your model will perform on new data.  That's about effectiveness.\n\nBut if you could use test ground truth when training your model then your model will perform better in test set than if you don't use test ground truth.  That's what is unfair to those who don't use test set at all.\n\nI don't know why you got downvoted, all questions are good questions, upvoting you!",
    "1347581": "cpmpml Thank you for being considerate and explaining fairness and effectiveness. \nI have two more questions. Can you please shed your thoughts on them?\n1. Can the model created using a test set be valuable and trusted by a radiologist for new real-world data? Which was the ideal motto of the competition.\n2. Can organisers ban using those two datasets (BIMCV and RICORD) or create a new testset instead?",
    "1348999": "I am afraid I can't answer your questions.  I am not a radiologist an I am not part of the organizer's team.  But I think the host answered your second question already, just read the original post again."
  },
  "source": "meta"
}