{
  "id": 246597,
  "title": "Recommendations for handling duplicates on the train dataset",
  "url": "/competitions/siim-covid19-detection/discussion/246597",
  "author_name": "ParasLakhani",
  "post_date": "2021-06-16T04:37:32.980000",
  "votes": 72,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Our annotation team updated the labels for the test datasets (private and public) but the train labels remain the same.</p>\n<p>The updated test labels correct the problem regarding duplicates or extra images on some of the studies, which previously did not have bounding box information.  Bounding box information is now provided for duplicates or similar images that are part of the same study for the public and private test datasets.</p>\n<p>However, regarding the train labels, due to time constraints, those labels were kept the same and not updated. As such, in situations where there are 2 or more images belonging to a study, we recommend only using the labels for the image with the bounding boxes and disregard the other duplicate/similar images. These other duplicate/similar images were likely not looked at by the annotators as we were unaware of them during the initial annotation process.  </p>",
  "messages": [
    {
      "id": 1351107,
      "postDate": "2021-06-16T04:37:32.980Z",
      "content": "<p>Our annotation team updated the labels for the test datasets (private and public) but the train labels remain the same.</p>\n<p>The updated test labels correct the problem regarding duplicates or extra images on some of the studies, which previously did not have bounding box information.  Bounding box information is now provided for duplicates or similar images that are part of the same study for the public and private test datasets.</p>\n<p>However, regarding the train labels, due to time constraints, those labels were kept the same and not updated. As such, in situations where there are 2 or more images belonging to a study, we recommend only using the labels for the image with the bounding boxes and disregard the other duplicate/similar images. These other duplicate/similar images were likely not looked at by the annotators as we were unaware of them during the initial annotation process.  </p>",
      "rawMarkdown": "Our annotation team updated the labels for the test datasets (private and public) but the train labels remain the same.\n\nThe updated test labels correct the problem regarding duplicates or extra images on some of the studies, which previously did not have bounding box information.  Bounding box information is now provided for duplicates or similar images that are part of the same study for the public and private test datasets.\n\nHowever, regarding the train labels, due to time constraints, those labels were kept the same and not updated. As such, in situations where there are 2 or more images belonging to a study, we recommend only using the labels for the image with the bounding boxes and disregard the other duplicate/similar images. These other duplicate/similar images were likely not looked at by the annotators as we were unaware of them during the initial annotation process.  ",
      "votes": 72
    },
    {
      "id": 1351173,
      "postDate": "2021-06-16T05:42:03.843Z",
      "content": "<p>Thanks for the post.<br>\nYou mean…<br>\nFor train data, we have to use only BBoxed images if they are duplicates, but we can still use non-duplicated images without BBox as no-BBox control ones.<br>\nFor test data, true BBox info is corrected now, but the image itself is the same, so we don't have to download it again to make PNGs because it is just a labeling repairment (not images but true labels corrected).<br>\nIs my understanding correct?</p>",
      "rawMarkdown": "Thanks for the post.\nYou mean...\nFor train data, we have to use only BBoxed images if they are duplicates, but we can still use non-duplicated images without BBox as no-BBox control ones.\nFor test data, true BBox info is corrected now, but the image itself is the same, so we don't have to download it again to make PNGs because it is just a labeling repairment (not images but true labels corrected).\nIs my understanding correct?",
      "votes": 3,
      "replies": [
        {
          "id": 1351235,
          "postDate": "2021-06-16T06:56:53.047Z",
          "content": "<p>it is up to you how to decide duplicate images and bbox label in train.<br>\nthe host suggest to use duplicate image with bbox label (but i think he has missed the fact that there may be some train duplicates with different bbox label? … i haven't checked the images and labels in detail yet)</p>\n<p>you can treat the problem now as \"having training dataset with some label noise or uncertainty\".<br>\nit is up to you how to handle this noise.</p>\n<p>but i am worried that the test and train images are labelled \"differently\" since there are no details on \"updated the labels for the test datasets\" (e.g. did the annotator changed?, etc). will there be domain difference?</p>\n<p>but nevertheless, kaggler should take note of any difference of scores in local cross-validation and the public leaderboard to account for domain change or label noise</p>",
          "rawMarkdown": "it is up to you how to decide duplicate images and bbox label in train.\nthe host suggest to use duplicate image with bbox label (but i think he has missed the fact that there may be some train duplicates with different bbox label? ... i haven't checked the images and labels in detail yet)\n\nyou can treat the problem now as \"having training dataset with some label noise or uncertainty\".\nit is up to you how to handle this noise.\n\n\nbut i am worried that the test and train images are labelled \"differently\" since there are no details on \"updated the labels for the test datasets\" (e.g. did the annotator changed?, etc). will there be domain difference?\n\nbut nevertheless, kaggler should take note of any difference of scores in local cross-validation and the public leaderboard to account for domain change or label noise",
          "votes": 8
        },
        {
          "id": 1353916,
          "postDate": "2021-06-17T09:41:46.877Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p><strong>\"We can consider the confidence score to be 1 always for every image at study-level because if you look at the train data, you will see that each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'.\"</strong></p>\n<p>is this true any more  at study level?</p>",
          "rawMarkdown": "@hengck23 \n\n**\"We can consider the confidence score to be 1 always for every image at study-level because if you look at the train data, you will see that each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'.\"**\n\nis this true any more  at study level?",
          "votes": 1
        },
        {
          "id": 1360361,
          "postDate": "2021-06-22T04:01:02.297Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>There seems to be only one image with with different bounding boxes marked. This particular image is part of two different studies and is labeled as <code>atypical appearance</code> in one study and <code>typical appearance</code> in the other study.</p>\n<p>Notebook with complete list of duplicate images here: <a href=\"https://www.kaggle.com/kwk100/siim-covid-19-duplicate-training-images\" target=\"_blank\">siim-covid-19-duplicate-training-images</a></p>\n<p><img src=\"https://i.ibb.co/fFSwFMc/duplicate.png\" alt=\"\"></p>",
          "rawMarkdown": "@hengck23 \n\nThere seems to be only one image with with different bounding boxes marked. This particular image is part of two different studies and is labeled as `atypical appearance` in one study and `typical appearance` in the other study.\n\nNotebook with complete list of duplicate images here: [siim-covid-19-duplicate-training-images](https://www.kaggle.com/kwk100/siim-covid-19-duplicate-training-images)\n\n![](https://i.ibb.co/fFSwFMc/duplicate.png)",
          "votes": 14
        },
        {
          "id": 1388949,
          "postDate": "2021-07-15T10:36:31.387Z",
          "content": "<p>Is this still the case or the dataset has been fixed ?</p>",
          "rawMarkdown": "Is this still the case or the dataset has been fixed ?"
        }
      ]
    },
    {
      "id": 1459996,
      "postDate": "2021-08-08T16:35:08.873Z",
      "content": "<p>Commen by Russian doesn`t count</p>",
      "rawMarkdown": "Commen by Russian doesn`t count"
    },
    {
      "id": 1447013,
      "postDate": "2021-08-04T13:48:03.930Z",
      "content": "<p>Hi I am new Here</p>",
      "rawMarkdown": "Hi I am new Here"
    },
    {
      "id": 1404690,
      "postDate": "2021-07-30T07:04:16.827Z",
      "content": "<p>i am worried that the test and train images are labelled \"differently\" since there are no details on \"updated the labels for the test datasets\" (e.g. did the annotator changed?, etc). will there be domain difference?</p>",
      "rawMarkdown": "i am worried that the test and train images are labelled \"differently\" since there are no details on \"updated the labels for the test datasets\" (e.g. did the annotator changed?, etc). will there be domain difference?"
    },
    {
      "id": 1381047,
      "postDate": "2021-07-08T15:14:07.210Z",
      "content": "<p>Does it mean that study level is also incorrect? Or are all the labels placed correctly at the study level?</p>",
      "rawMarkdown": "Does it mean that study level is also incorrect? Or are all the labels placed correctly at the study level?"
    },
    {
      "id": 1355219,
      "postDate": "2021-06-18T06:52:26.080Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1464849,
      "postDate": "2021-08-10T18:27:58.620Z",
      "content": "<p>Thanks for the post</p>",
      "rawMarkdown": "Thanks for the post"
    },
    {
      "id": 1393003,
      "postDate": "2021-07-19T09:48:13.550Z",
      "content": "<p>thank you for your generous</p>",
      "rawMarkdown": "thank you for your generous"
    }
  ],
  "comments": [
    {
      "id": 1351173,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-06-16T05:42:03.843000",
      "content": "<p>Thanks for the post.<br>\nYou mean…<br>\nFor train data, we have to use only BBoxed images if they are duplicates, but we can still use non-duplicated images without BBox as no-BBox control ones.<br>\nFor test data, true BBox info is corrected now, but the image itself is the same, so we don't have to download it again to make PNGs because it is just a labeling repairment (not images but true labels corrected).<br>\nIs my understanding correct?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1351235,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-06-16T06:56:53.047000",
          "content": "<p>it is up to you how to decide duplicate images and bbox label in train.<br>\nthe host suggest to use duplicate image with bbox label (but i think he has missed the fact that there may be some train duplicates with different bbox label? … i haven't checked the images and labels in detail yet)</p>\n<p>you can treat the problem now as \"having training dataset with some label noise or uncertainty\".<br>\nit is up to you how to handle this noise.</p>\n<p>but i am worried that the test and train images are labelled \"differently\" since there are no details on \"updated the labels for the test datasets\" (e.g. did the annotator changed?, etc). will there be domain difference?</p>\n<p>but nevertheless, kaggler should take note of any difference of scores in local cross-validation and the public leaderboard to account for domain change or label noise</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1353916,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-06-17T09:41:46.877000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p><strong>\"We can consider the confidence score to be 1 always for every image at study-level because if you look at the train data, you will see that each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'.\"</strong></p>\n<p>is this true any more  at study level?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1360361,
          "author_name": "KWK",
          "author_url": "",
          "post_date": "2021-06-22T04:01:02.297000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>There seems to be only one image with with different bounding boxes marked. This particular image is part of two different studies and is labeled as <code>atypical appearance</code> in one study and <code>typical appearance</code> in the other study.</p>\n<p>Notebook with complete list of duplicate images here: <a href=\"https://www.kaggle.com/kwk100/siim-covid-19-duplicate-training-images\" target=\"_blank\">siim-covid-19-duplicate-training-images</a></p>\n<p><img src=\"https://i.ibb.co/fFSwFMc/duplicate.png\" alt=\"\"></p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1388949,
          "author_name": "Akshay Pratap Singh",
          "author_url": "",
          "post_date": "2021-07-15T10:36:31.387000",
          "content": "<p>Is this still the case or the dataset has been fixed ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1459996,
      "author_name": "Ulymov Igor",
      "author_url": "",
      "post_date": "2021-08-08T16:35:08.873000",
      "content": "<p>Commen by Russian doesn`t count</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1447013,
      "author_name": "tanmay walke",
      "author_url": "",
      "post_date": "2021-08-04T13:48:03.930000",
      "content": "<p>Hi I am new Here</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1404690,
      "author_name": "Manjula Srihari",
      "author_url": "",
      "post_date": "2021-07-30T07:04:16.827000",
      "content": "<p>i am worried that the test and train images are labelled \"differently\" since there are no details on \"updated the labels for the test datasets\" (e.g. did the annotator changed?, etc). will there be domain difference?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1381047,
      "author_name": "Aleksandr Emchinov",
      "author_url": "",
      "post_date": "2021-07-08T15:14:07.210000",
      "content": "<p>Does it mean that study level is also incorrect? Or are all the labels placed correctly at the study level?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1355219,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-18T06:52:26.080000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1464849,
      "author_name": "Anand Krishnan",
      "author_url": "",
      "post_date": "2021-08-10T18:27:58.620000",
      "content": "<p>Thanks for the post</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1393003,
      "author_name": "jinpaigiegie",
      "author_url": "",
      "post_date": "2021-07-19T09:48:13.550000",
      "content": "<p>thank you for your generous</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1351107": "Our annotation team updated the labels for the test datasets (private and public) but the train labels remain the same.\n\nThe updated test labels correct the problem regarding duplicates or extra images on some of the studies, which previously did not have bounding box information.  Bounding box information is now provided for duplicates or similar images that are part of the same study for the public and private test datasets.\n\nHowever, regarding the train labels, due to time constraints, those labels were kept the same and not updated. As such, in situations where there are 2 or more images belonging to a study, we recommend only using the labels for the image with the bounding boxes and disregard the other duplicate/similar images. These other duplicate/similar images were likely not looked at by the annotators as we were unaware of them during the initial annotation process.  ",
    "1351173": "Thanks for the post.\nYou mean...\nFor train data, we have to use only BBoxed images if they are duplicates, but we can still use non-duplicated images without BBox as no-BBox control ones.\nFor test data, true BBox info is corrected now, but the image itself is the same, so we don't have to download it again to make PNGs because it is just a labeling repairment (not images but true labels corrected).\nIs my understanding correct?",
    "1459996": "Commen by Russian doesn`t count",
    "1447013": "Hi I am new Here",
    "1404690": "i am worried that the test and train images are labelled \"differently\" since there are no details on \"updated the labels for the test datasets\" (e.g. did the annotator changed?, etc). will there be domain difference?",
    "1381047": "Does it mean that study level is also incorrect? Or are all the labels placed correctly at the study level?",
    "1355219": "",
    "1464849": "Thanks for the post",
    "1393003": "thank you for your generous"
  }
}