{
  "id": 106000,
  "title": "Overlapping Labels in Train Data?",
  "url": "/competitions/understanding_cloud_organization/discussion/106000",
  "author_name": "",
  "post_date": "2019-08-27T16:31:12.383675Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey all,</p>\n\n<p>So I am having a poke around at the data while I get up and running, and it looks to me like some  (at least one) of the images have overlapping labels. Intuitively, this seems like this should not be the case, as each cloud/pixel can only belong to a single cloud-type. </p>\n\n<p>I have included an example below on what I am seeing, the image that I am looking at is \"0011165.jpg\" for anyone that wants to play around at home.</p>\n\n<p>If someone can enlighten me on where my logical shortfalls are, be that in reading and understanding the images, or in understanding the labels, that would be fantastic!</p>\n\n<p><img src=\"https://puu.sh/EadmW/09658b92dc.png\" alt=\"My labels\"></p>",
  "messages": [
    {
      "id": "609361",
      "postDate": "08/27/2019 16:31:12",
      "content": "<p>Hey all,</p>\n\n<p>So I am having a poke around at the data while I get up and running, and it looks to me like some  (at least one) of the images have overlapping labels. Intuitively, this seems like this should not be the case, as each cloud/pixel can only belong to a single cloud-type. </p>\n\n<p>I have included an example below on what I am seeing, the image that I am looking at is \"0011165.jpg\" for anyone that wants to play around at home.</p>\n\n<p>If someone can enlighten me on where my logical shortfalls are, be that in reading and understanding the images, or in understanding the labels, that would be fantastic!</p>\n\n<p><img src=\"https://puu.sh/EadmW/09658b92dc.png\" alt=\"My labels\"></p>",
      "rawMarkdown": "Hey all,\n\nSo I am having a poke around at the data while I get up and running, and it looks to me like some  (at least one) of the images have overlapping labels. Intuitively, this seems like this should not be the case, as each cloud/pixel can only belong to a single cloud-type. \n\nI have included an example below on what I am seeing, the image that I am looking at is \"0011165.jpg\" for anyone that wants to play around at home.\n\nIf someone can enlighten me on where my logical shortfalls are, be that in reading and understanding the images, or in understanding the labels, that would be fantastic!\n\n![My labels](https://puu.sh/EadmW/09658b92dc.png)",
      "votes": null
    },
    {
      "id": "609845",
      "postDate": "08/28/2019 06:51:27",
      "content": "<p>Hi,\nthanks for pointing this out since it probably requires some clarification. In traditional image segmentation each pixel can only belong to one class, e.g. it's either a car or a person. In other words, there is a definite ground truth. </p>\n\n<p>In this competition, this is not the case. The classes are subjective. One person's flower can be another persons fish. The labelers were given rough guidelines on what each class looks like but here we have no ground truth. </p>\n\n<p>Each image was labeled by several people (2-4), so the labels can overlap. In addition, there was no restriction that the labels from a single labeler cannot overlap. To create the masks for this competition, we simply used the union of all labels for each class. So naturally there will be some overlap.</p>\n\n<p>This is actually one of the interesting aspects of this competition. How will the algorithms deal with such uncertainty.</p>\n\n<p>I hope this clarifies it. Let me know if you have any further questions.</p>",
      "rawMarkdown": "Hi,\nthanks for pointing this out since it probably requires some clarification. In traditional image segmentation each pixel can only belong to one class, e.g. it's either a car or a person. In other words, there is a definite ground truth. \n\nIn this competition, this is not the case. The classes are subjective. One person's flower can be another persons fish. The labelers were given rough guidelines on what each class looks like but here we have no ground truth. \n\nEach image was labeled by several people (2-4), so the labels can overlap. In addition, there was no restriction that the labels from a single labeler cannot overlap. To create the masks for this competition, we simply used the union of all labels for each class. So naturally there will be some overlap.\n\nThis is actually one of the interesting aspects of this competition. How will the algorithms deal with such uncertainty.\n\nI hope this clarifies it. Let me know if you have any further questions.",
      "votes": null
    },
    {
      "id": "610048",
      "postDate": "08/28/2019 11:06:36",
      "content": "<p>Thanks for the response, what you are saying makes a fair amount of sense. </p>\n\n<p>Following up, are the test images labelled in a similar manner? I.e., some parts of test images are treated as multiple types (e.g., flower and fish?). This would have implications on how the DICE coefficient for each image is calculated...</p>",
      "rawMarkdown": "Thanks for the response, what you are saying makes a fair amount of sense. \n\nFollowing up, are the test images labelled in a similar manner? I.e., some parts of test images are treated as multiple types (e.g., flower and fish?). This would have implications on how the DICE coefficient for each image is calculated...",
      "votes": null
    },
    {
      "id": "610289",
      "postDate": "08/28/2019 16:11:08",
      "content": "<p>From what I can tell, the split between TEST and TRAIN is random.</p>\n\n<p>Note that about 3/5 of images have overlapping labels, and indeed 1/6 of labelled pixels are tagged with more than one label. There are even some areas with all 4 labels e.g. see 0cb6144.jpg</p>\n\n<p>Dealing with noisy labels will be a key part of the challenge. Even images with masks from a single class may be in disagreement amongst labellers, but we are not provided with this information. It would have been useful if we were provided with labels and labeller_id's instead of just a union of labels, but its not Christmas.</p>",
      "rawMarkdown": "From what I can tell, the split between TEST and TRAIN is random.\n\nNote that about 3/5 of images have overlapping labels, and indeed 1/6 of labelled pixels are tagged with more than one label. There are even some areas with all 4 labels e.g. see 0cb6144.jpg\n\nDealing with noisy labels will be a key part of the challenge. Even images with masks from a single class may be in disagreement amongst labellers, but we are not provided with this information. It would have been useful if we were provided with labels and labeller_id's instead of just a union of labels, but its not Christmas.",
      "votes": null
    },
    {
      "id": "610437",
      "postDate": "08/28/2019 19:47:43",
      "content": "<p>Yes, the test set is just a random split, so there will be overlapping labels in the test set. </p>\n\n<p>We were thinking about including information about the labeler id but decided against it. The reason is that we, as scientists, do not actually care about the individual labelers. We want the model to make predictions that represent the mean agreement between humans. So while this information might be helpful for getting a good score in the competition, it is not something that we would want to use in practice.</p>",
      "rawMarkdown": "Yes, the test set is just a random split, so there will be overlapping labels in the test set. \n\nWe were thinking about including information about the labeler id but decided against it. The reason is that we, as scientists, do not actually care about the individual labelers. We want the model to make predictions that represent the mean agreement between humans. So while this information might be helpful for getting a good score in the competition, it is not something that we would want to use in practice.",
      "votes": null
    },
    {
      "id": "614849",
      "postDate": "09/01/2019 06:43:29",
      "content": "<p>just a small question is it necessary to predict mask for every test image or it can be resulted as nan like in train dataset\nthankx</p>",
      "rawMarkdown": "just a small question is it necessary to predict mask for every test image or it can be resulted as nan like in train dataset\nthankx",
      "votes": null
    },
    {
      "id": "616938",
      "postDate": "09/03/2019 15:55:51",
      "content": "<p>The train and test sets are random splits. This means that the train can also contain images that have no masks.</p>",
      "rawMarkdown": "The train and test sets are random splits. This means that the train can also contain images that have no masks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 609845,
      "author_name": "raspstephan",
      "author_url": "",
      "post_date": "08/28/2019 06:51:27",
      "content": "<p>Hi,\nthanks for pointing this out since it probably requires some clarification. In traditional image segmentation each pixel can only belong to one class, e.g. it's either a car or a person. In other words, there is a definite ground truth. </p>\n\n<p>In this competition, this is not the case. The classes are subjective. One person's flower can be another persons fish. The labelers were given rough guidelines on what each class looks like but here we have no ground truth. </p>\n\n<p>Each image was labeled by several people (2-4), so the labels can overlap. In addition, there was no restriction that the labels from a single labeler cannot overlap. To create the masks for this competition, we simply used the union of all labels for each class. So naturally there will be some overlap.</p>\n\n<p>This is actually one of the interesting aspects of this competition. How will the algorithms deal with such uncertainty.</p>\n\n<p>I hope this clarifies it. Let me know if you have any further questions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 610048,
          "author_name": "adrianball",
          "author_url": "",
          "post_date": "08/28/2019 11:06:36",
          "content": "<p>Thanks for the response, what you are saying makes a fair amount of sense. </p>\n\n<p>Following up, are the test images labelled in a similar manner? I.e., some parts of test images are treated as multiple types (e.g., flower and fish?). This would have implications on how the DICE coefficient for each image is calculated...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 610289,
          "author_name": "robga",
          "author_url": "",
          "post_date": "08/28/2019 16:11:08",
          "content": "<p>From what I can tell, the split between TEST and TRAIN is random.</p>\n\n<p>Note that about 3/5 of images have overlapping labels, and indeed 1/6 of labelled pixels are tagged with more than one label. There are even some areas with all 4 labels e.g. see 0cb6144.jpg</p>\n\n<p>Dealing with noisy labels will be a key part of the challenge. Even images with masks from a single class may be in disagreement amongst labellers, but we are not provided with this information. It would have been useful if we were provided with labels and labeller_id's instead of just a union of labels, but its not Christmas.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 610437,
          "author_name": "raspstephan",
          "author_url": "",
          "post_date": "08/28/2019 19:47:43",
          "content": "<p>Yes, the test set is just a random split, so there will be overlapping labels in the test set. </p>\n\n<p>We were thinking about including information about the labeler id but decided against it. The reason is that we, as scientists, do not actually care about the individual labelers. We want the model to make predictions that represent the mean agreement between humans. So while this information might be helpful for getting a good score in the competition, it is not something that we would want to use in practice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 614849,
      "author_name": "pranshu29",
      "author_url": "",
      "post_date": "09/01/2019 06:43:29",
      "content": "<p>just a small question is it necessary to predict mask for every test image or it can be resulted as nan like in train dataset\nthankx</p>",
      "votes": null,
      "replies": [
        {
          "id": 616938,
          "author_name": "raspstephan",
          "author_url": "",
          "post_date": "09/03/2019 15:55:51",
          "content": "<p>The train and test sets are random splits. This means that the train can also contain images that have no masks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "609361": "Hey all,\n\nSo I am having a poke around at the data while I get up and running, and it looks to me like some  (at least one) of the images have overlapping labels. Intuitively, this seems like this should not be the case, as each cloud/pixel can only belong to a single cloud-type. \n\nI have included an example below on what I am seeing, the image that I am looking at is \"0011165.jpg\" for anyone that wants to play around at home.\n\nIf someone can enlighten me on where my logical shortfalls are, be that in reading and understanding the images, or in understanding the labels, that would be fantastic!\n\n![My labels](https://puu.sh/EadmW/09658b92dc.png)",
    "609845": "Hi,\nthanks for pointing this out since it probably requires some clarification. In traditional image segmentation each pixel can only belong to one class, e.g. it's either a car or a person. In other words, there is a definite ground truth. \n\nIn this competition, this is not the case. The classes are subjective. One person's flower can be another persons fish. The labelers were given rough guidelines on what each class looks like but here we have no ground truth. \n\nEach image was labeled by several people (2-4), so the labels can overlap. In addition, there was no restriction that the labels from a single labeler cannot overlap. To create the masks for this competition, we simply used the union of all labels for each class. So naturally there will be some overlap.\n\nThis is actually one of the interesting aspects of this competition. How will the algorithms deal with such uncertainty.\n\nI hope this clarifies it. Let me know if you have any further questions.",
    "610048": "Thanks for the response, what you are saying makes a fair amount of sense. \n\nFollowing up, are the test images labelled in a similar manner? I.e., some parts of test images are treated as multiple types (e.g., flower and fish?). This would have implications on how the DICE coefficient for each image is calculated...",
    "610289": "From what I can tell, the split between TEST and TRAIN is random.\n\nNote that about 3/5 of images have overlapping labels, and indeed 1/6 of labelled pixels are tagged with more than one label. There are even some areas with all 4 labels e.g. see 0cb6144.jpg\n\nDealing with noisy labels will be a key part of the challenge. Even images with masks from a single class may be in disagreement amongst labellers, but we are not provided with this information. It would have been useful if we were provided with labels and labeller_id's instead of just a union of labels, but its not Christmas.",
    "610437": "Yes, the test set is just a random split, so there will be overlapping labels in the test set. \n\nWe were thinking about including information about the labeler id but decided against it. The reason is that we, as scientists, do not actually care about the individual labelers. We want the model to make predictions that represent the mean agreement between humans. So while this information might be helpful for getting a good score in the competition, it is not something that we would want to use in practice.",
    "614849": "just a small question is it necessary to predict mask for every test image or it can be resulted as nan like in train dataset\nthankx",
    "616938": "The train and test sets are random splits. This means that the train can also contain images that have no masks."
  },
  "source": "meta"
}