{
  "id": 329298,
  "title": "Train Data Labels Questions",
  "url": "/competitions/unifesp-x-ray-body-part-classifier/discussion/329298",
  "author_name": "",
  "post_date": "2022-06-06T02:24:27.395152200Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Dear All,</p>\n<p>I have the following questiosn regarding the train subset labels:</p>\n<ol>\n<li><p>I would like more elaboration on the annotation methodology used for the train dataset. For example, <code>SOPInstanceUID</code> <code>1.2.826.0.1.3680043.8.498.10065930002825553435161793347987832017</code> is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?</p></li>\n<li><p>While digging deep into the training data, I believe there is am image that is not correctly labelled. The <code>SOPInstanceUID</code> of the image is: <code>1.2.826.0.1.3680043.8.498.11911833370040822194823098700931537024</code> which is labelled as Wrist (21) however the image is an ankle (11).</p></li>\n</ol>\n<p>Thanks a lot</p>\n<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a>  <a href=\"https://www.kaggle.com/eduardofarina\" target=\"_blank\">@eduardofarina</a> </p>",
  "messages": [
    {
      "id": "1812551",
      "postDate": "06/06/2022 02:24:27",
      "content": "<p>Dear All,</p>\n<p>I have the following questiosn regarding the train subset labels:</p>\n<ol>\n<li><p>I would like more elaboration on the annotation methodology used for the train dataset. For example, <code>SOPInstanceUID</code> <code>1.2.826.0.1.3680043.8.498.10065930002825553435161793347987832017</code> is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?</p></li>\n<li><p>While digging deep into the training data, I believe there is am image that is not correctly labelled. The <code>SOPInstanceUID</code> of the image is: <code>1.2.826.0.1.3680043.8.498.11911833370040822194823098700931537024</code> which is labelled as Wrist (21) however the image is an ankle (11).</p></li>\n</ol>\n<p>Thanks a lot</p>\n<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a>  <a href=\"https://www.kaggle.com/eduardofarina\" target=\"_blank\">@eduardofarina</a> </p>",
      "rawMarkdown": "Dear All,\n\nI have the following questiosn regarding the train subset labels:\n\n1. I would like more elaboration on the annotation methodology used for the train dataset. For example, `SOPInstanceUID` `1.2.826.0.1.3680043.8.498.10065930002825553435161793347987832017` is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?\n\n2. While digging deep into the training data, I believe there is am image that is not correctly labelled. The `SOPInstanceUID` of the image is: `1.2.826.0.1.3680043.8.498.11911833370040822194823098700931537024` which is labelled as Wrist (21) however the image is an ankle (11).\n\n\nThanks a lot\n\n@felipekitamura  @eduardofarina",
      "votes": null
    },
    {
      "id": "1815936",
      "postDate": "06/09/2022 16:36:12",
      "content": "<p>I agree with you about why something \"is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?\" - this is very subjective in my opinion.</p>",
      "rawMarkdown": "I agree with you about why something \"is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?\" - this is very subjective in my opinion.",
      "votes": null
    },
    {
      "id": "1815938",
      "postDate": "06/09/2022 16:36:52",
      "content": "<p>Errors in DICOM tags: addressed in the Background as a reason for this competition.</p>",
      "rawMarkdown": "Errors in DICOM tags: addressed in the Background as a reason for this competition.",
      "votes": null
    },
    {
      "id": "1816020",
      "postDate": "06/09/2022 18:06:59",
      "content": "<p>Thank you. Can you please elaborate more on this ? I don't think I fully understand. </p>",
      "rawMarkdown": "Thank you. Can you please elaborate more on this ? I don't think I fully understand.",
      "votes": null
    },
    {
      "id": "1832295",
      "postDate": "06/24/2022 22:54:03",
      "content": "<p>Hi,</p>\n<p>Thanks for your question.</p>\n<p>We built a real-world data competition. Hosting it can be challenging, and errors can occur when extracting data and labeling it with multiple source annotators. </p>\n<p>The organization has been thinking about what would be the most suitable option for the completion, and as the data was randomly split into training and test set, the labeling process was the same for both datasets.</p>\n<p>We've decided it is fairer in this competition stage not to change any label, neither in the training nor in the test set, because it could cause a huge shake-up, and be a disadvantage for the competitors that joined since the launching and have struggled all the way here fine-tuning its models.</p>\n<p>We know every decision is hard, and changing the labels or not changing them would affect participants. however, we've decided to keep the dataset and the labels the same way they are now.</p>\n<p>We apologize for any inconvenience and we hope everyone will keep motivated to address a real-world dataset to solve a real-world problem</p>\n<p>Thanks,</p>\n<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> <a href=\"https://www.kaggle.com/eduardofarina\" target=\"_blank\">@eduardofarina</a> </p>",
      "rawMarkdown": "Hi,\n\nThanks for your question.\n\nWe built a real-world data competition. Hosting it can be challenging, and errors can occur when extracting data and labeling it with multiple source annotators. \n\nThe organization has been thinking about what would be the most suitable option for the completion, and as the data was randomly split into training and test set, the labeling process was the same for both datasets.\n\nWe've decided it is fairer in this competition stage not to change any label, neither in the training nor in the test set, because it could cause a huge shake-up, and be a disadvantage for the competitors that joined since the launching and have struggled all the way here fine-tuning its models.\n\nWe know every decision is hard, and changing the labels or not changing them would affect participants. however, we've decided to keep the dataset and the labels the same way they are now.\n\nWe apologize for any inconvenience and we hope everyone will keep motivated to address a real-world dataset to solve a real-world problem\n\nThanks,\n\n@felipekitamura @eduardofarina",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1815936,
      "author_name": "alexanderyyy",
      "author_url": "",
      "post_date": "06/09/2022 16:36:12",
      "content": "<p>I agree with you about why something \"is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?\" - this is very subjective in my opinion.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1815938,
      "author_name": "alexanderyyy",
      "author_url": "",
      "post_date": "06/09/2022 16:36:52",
      "content": "<p>Errors in DICOM tags: addressed in the Background as a reason for this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1816020,
          "author_name": "elbanan",
          "author_url": "",
          "post_date": "06/09/2022 18:06:59",
          "content": "<p>Thank you. Can you please elaborate more on this ? I don't think I fully understand. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1832295,
      "author_name": "eduardofarina",
      "author_url": "",
      "post_date": "06/24/2022 22:54:03",
      "content": "<p>Hi,</p>\n<p>Thanks for your question.</p>\n<p>We built a real-world data competition. Hosting it can be challenging, and errors can occur when extracting data and labeling it with multiple source annotators. </p>\n<p>The organization has been thinking about what would be the most suitable option for the completion, and as the data was randomly split into training and test set, the labeling process was the same for both datasets.</p>\n<p>We've decided it is fairer in this competition stage not to change any label, neither in the training nor in the test set, because it could cause a huge shake-up, and be a disadvantage for the competitors that joined since the launching and have struggled all the way here fine-tuning its models.</p>\n<p>We know every decision is hard, and changing the labels or not changing them would affect participants. however, we've decided to keep the dataset and the labels the same way they are now.</p>\n<p>We apologize for any inconvenience and we hope everyone will keep motivated to address a real-world dataset to solve a real-world problem</p>\n<p>Thanks,</p>\n<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> <a href=\"https://www.kaggle.com/eduardofarina\" target=\"_blank\">@eduardofarina</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1812551": "Dear All,\n\nI have the following questiosn regarding the train subset labels:\n\n1. I would like more elaboration on the annotation methodology used for the train dataset. For example, `SOPInstanceUID` `1.2.826.0.1.3680043.8.498.10065930002825553435161793347987832017` is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?\n\n2. While digging deep into the training data, I believe there is am image that is not correctly labelled. The `SOPInstanceUID` of the image is: `1.2.826.0.1.3680043.8.498.11911833370040822194823098700931537024` which is labelled as Wrist (21) however the image is an ankle (11).\n\n\nThanks a lot\n\n@felipekitamura  @eduardofarina",
    "1815936": "I agree with you about why something \"is labelled as thoracic and lumbar spine. Why isn't it labelled as chest for example ?\" - this is very subjective in my opinion.",
    "1815938": "Errors in DICOM tags: addressed in the Background as a reason for this competition.",
    "1816020": "Thank you. Can you please elaborate more on this ? I don't think I fully understand.",
    "1832295": "Hi,\n\nThanks for your question.\n\nWe built a real-world data competition. Hosting it can be challenging, and errors can occur when extracting data and labeling it with multiple source annotators. \n\nThe organization has been thinking about what would be the most suitable option for the completion, and as the data was randomly split into training and test set, the labeling process was the same for both datasets.\n\nWe've decided it is fairer in this competition stage not to change any label, neither in the training nor in the test set, because it could cause a huge shake-up, and be a disadvantage for the competitors that joined since the launching and have struggled all the way here fine-tuning its models.\n\nWe know every decision is hard, and changing the labels or not changing them would affect participants. however, we've decided to keep the dataset and the labels the same way they are now.\n\nWe apologize for any inconvenience and we hope everyone will keep motivated to address a real-world dataset to solve a real-world problem\n\nThanks,\n\n@felipekitamura @eduardofarina"
  },
  "source": "meta"
}