{
  "id": 280250,
  "title": "Overlaps - ambiguous or not?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/280250",
  "author_name": "",
  "post_date": "2021-10-20T21:51:27.863547200Z",
  "votes": 47,
  "comment_count": 13,
  "views": 0,
  "content": "<p>The competition states:</p>\n<blockquote>\n  <p>Note: while predictions are not allowed to overlap, the training labels are provided in full (with overlapping portions included). This is to ensure that models are provided the full data for each object. Removing overlap in predictions is a task for the competitor.</p>\n</blockquote>\n<p><img src=\"https://i.postimg.cc/nLvMfSy7/ambiguous.jpg\" alt=\"overlap\"></p>\n<p>This means we will have <strong>possibly</strong> more instances assigned to the <strong>same pixel</strong> if we use an instance segmentation network… but the objects are transparent, this means the overlapping areas are equivalent (at least theoretically - or do i miss something here?)</p>\n<p>Do the contours cross each other in a way one can say - hey, this cell it OVER the other? If so, the pixels of the overlapping areas could equally belong to the instances == not ambiguous.</p>\n<p>Is there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision… </p>\n<p>… anyway, how many cases are there of overlapping areas at all… if like 0.1% the cases do we care… what if 3%?</p>\n<p>;)</p>",
  "messages": [
    {
      "id": "1551793",
      "postDate": "10/20/2021 21:51:27",
      "content": "<p>The competition states:</p>\n<blockquote>\n  <p>Note: while predictions are not allowed to overlap, the training labels are provided in full (with overlapping portions included). This is to ensure that models are provided the full data for each object. Removing overlap in predictions is a task for the competitor.</p>\n</blockquote>\n<p><img src=\"https://i.postimg.cc/nLvMfSy7/ambiguous.jpg\" alt=\"overlap\"></p>\n<p>This means we will have <strong>possibly</strong> more instances assigned to the <strong>same pixel</strong> if we use an instance segmentation network… but the objects are transparent, this means the overlapping areas are equivalent (at least theoretically - or do i miss something here?)</p>\n<p>Do the contours cross each other in a way one can say - hey, this cell it OVER the other? If so, the pixels of the overlapping areas could equally belong to the instances == not ambiguous.</p>\n<p>Is there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision… </p>\n<p>… anyway, how many cases are there of overlapping areas at all… if like 0.1% the cases do we care… what if 3%?</p>\n<p>;)</p>",
      "rawMarkdown": "The competition states:\n\n> Note: while predictions are not allowed to overlap, the training labels are provided in full (with overlapping portions included). This is to ensure that models are provided the full data for each object. Removing overlap in predictions is a task for the competitor.\n\n![overlap](https://i.postimg.cc/nLvMfSy7/ambiguous.jpg)\n\nThis means we will have **possibly** more instances assigned to the **same pixel** if we use an instance segmentation network... but the objects are transparent, this means the overlapping areas are equivalent (at least theoretically - or do i miss something here?)\n\nDo the contours cross each other in a way one can say - hey, this cell it OVER the other? If so, the pixels of the overlapping areas could equally belong to the instances == not ambiguous.\n\nIs there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision... \n\n... anyway, how many cases are there of overlapping areas at all... if like 0.1% the cases do we care... what if 3%?\n\n ;)",
      "votes": null
    },
    {
      "id": "1553646",
      "postDate": "10/22/2021 11:23:18",
      "content": "<blockquote>\n  <p>Is there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision…</p>\n</blockquote>\n<p>Yes, this is a very important problem. We can simply remove those overlaps, but we can not guarantee the correctness of the removal unless we know how the pathologists handle such cases.</p>",
      "rawMarkdown": "> Is there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision…\n\nYes, this is a very important problem. We can simply remove those overlaps, but we can not guarantee the correctness of the removal unless we know how the pathologists handle such cases.",
      "votes": null
    },
    {
      "id": "1555133",
      "postDate": "10/23/2021 15:58:35",
      "content": "<p><a href=\"https://www.kaggle.com/stupidios\" target=\"_blank\">@stupidios</a> exactly!<br>\nI made some analysis of the overlaps <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "stupidios exactly!\nI made some analysis of the overlaps [here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084).",
      "votes": null
    },
    {
      "id": "1571081",
      "postDate": "11/04/2021 15:29:47",
      "content": "<p>Hi there dr Konya, just wanted to contribute a bit and state that the IoU overlap in the evaluation data is low. The evaluation metric supported by Kaggle either had only one IoU threshold (but with overlap support) or multiple IoUs thresholds but no overlap.</p>\n<p>The amount of overlap in the evaluation dataset is very low (lower than the training dataset) and we therefor went with no-overlap but multiple IoUs evaluation. When we internally work with this data we prefer the <a href=\"https://cocodataset.org/#detection-eval\" target=\"_blank\">coco-mAP </a> metric since it supports both overlaps of objects and evaluation of multiple IoUs. Meaning that our biologists will be able to keep track of overlapping cells in their experiments.</p>\n<p>Kindly,<br>\nChristoffer</p>",
      "rawMarkdown": "Hi there dr Konya, just wanted to contribute a bit and state that the IoU overlap in the evaluation data is low. The evaluation metric supported by Kaggle either had only one IoU threshold (but with overlap support) or multiple IoUs thresholds but no overlap.\n\nThe amount of overlap in the evaluation dataset is very low (lower than the training dataset) and we therefor went with no-overlap but multiple IoUs evaluation. When we internally work with this data we prefer the [coco-mAP ](https://cocodataset.org/#detection-eval) metric since it supports both overlaps of objects and evaluation of multiple IoUs. Meaning that our biologists will be able to keep track of overlapping cells in their experiments.\n\nKindly,\nChristoffer",
      "votes": null
    },
    {
      "id": "1575582",
      "postDate": "11/08/2021 13:39:04",
      "content": "<p><a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> How are the overlapping regions in the test set reduced to a single category? Randomly?</p>",
      "rawMarkdown": "christoffersartorius How are the overlapping regions in the test set reduced to a single category? Randomly?",
      "votes": null
    },
    {
      "id": "1579568",
      "postDate": "11/12/2021 02:16:45",
      "content": "<blockquote>\n  <p>how many cases are there of overlapping areas at all…</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287029\" target=\"_blank\">Here</a> is the answer of your question. Three images contain very high fraction of overlap.</p>",
      "rawMarkdown": "> how many cases are there of overlapping areas at all…\n\n@sandorkonya  [Here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287029) is the answer of your question. Three images contain very high fraction of overlap.",
      "votes": null
    },
    {
      "id": "1593454",
      "postDate": "11/23/2021 23:42:09",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> for the useful and clear information. Since the LB ground truth masks were originally overlapping, can you let us know how the overlapping pixels were assigned to instances in the competition ground truth ? I can think of 4 possibilities:</p>\n<ol>\n<li>Randomly</li>\n<li>Overlapping pixels assigned to the largest overlapping mask</li>\n<li>Overlapping pixels assigned to the smallest overlapping mask</li>\n<li>By the first index appearance when looping trough a list of instances</li>\n</ol>\n<p>We could probably probe the public LB test set with these hypothesis but letting us know officially would be more fair and transparent for everyone since the ground truth between train and LB test is by definition different.</p>",
      "rawMarkdown": "Thanks a lot @christoffersartorius for the useful and clear information. Since the LB ground truth masks were originally overlapping, can you let us know how the overlapping pixels were assigned to instances in the competition ground truth ? I can think of 4 possibilities:\n1.  Randomly\n2. Overlapping pixels assigned to the largest overlapping mask\n3. Overlapping pixels assigned to the smallest overlapping mask\n4. By the first index appearance when looping trough a list of instances\n\nWe could probably probe the public LB test set with these hypothesis but letting us know officially would be more fair and transparent for everyone since the ground truth between train and LB test is by definition different.",
      "votes": null
    },
    {
      "id": "1608570",
      "postDate": "12/06/2021 13:04:40",
      "content": "<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a>, Hi there! Great questions. The Kaggle team did the conversion so I am not sure what method they used. But i definitely agree that transparency on this would be a benefit for everyone involved, including us as host.</p>\n<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>, do you the possibility to find the response to these questions?</p>",
      "rawMarkdown": "alexandrecc, Hi there! Great questions. The Kaggle team did the conversion so I am not sure what method they used. But i definitely agree that transparency on this would be a benefit for everyone involved, including us as host.\n\n@addisonhoward, do you the possibility to find the response to these questions?",
      "votes": null
    },
    {
      "id": "1608807",
      "postDate": "12/06/2021 17:11:25",
      "content": "<p>Hi there - The correct answer is Option 4 - we assigned them in index order</p>",
      "rawMarkdown": "Hi there - The correct answer is Option 4 - we assigned them in index order",
      "votes": null
    },
    {
      "id": "1614838",
      "postDate": "12/11/2021 14:09:51",
      "content": "<p><a href=\"https://www.kaggle.com/ChristofferSartorius\" target=\"_blank\">@ChristofferSartorius</a> I am curious why the evaluation set has less overlap than the training set. Is the evaluation set is annotated differently (e.g., by experts) or ambiguously-overlapped cells are merged into one cell?</p>",
      "rawMarkdown": "ChristofferSartorius I am curious why the evaluation set has less overlap than the training set. Is the evaluation set is annotated differently (e.g., by experts) or ambiguously-overlapped cells are merged into one cell?",
      "votes": null
    },
    {
      "id": "1621769",
      "postDate": "12/18/2021 03:28:14",
      "content": "<p>But an important question is in what order are these instances listed. Confidence? Randomness?</p>",
      "rawMarkdown": "But an important question is in what order are these instances listed. Confidence? Randomness?",
      "votes": null
    },
    {
      "id": "1621797",
      "postDate": "12/18/2021 04:16:12",
      "content": "<p>What do you mean? Ground truth labels doesn't have confidence. On the other note, NOICE ranking😉.</p>",
      "rawMarkdown": "What do you mean? Ground truth labels doesn't have confidence. On the other note, NOICE ranking😉.",
      "votes": null
    },
    {
      "id": "1621835",
      "postDate": "12/18/2021 05:55:18",
      "content": "<p>The order of the predicted label of the instances.</p>",
      "rawMarkdown": "The order of the predicted label of the instances.",
      "votes": null
    },
    {
      "id": "1622430",
      "postDate": "12/18/2021 18:16:27",
      "content": "<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a> Good question. I was sorting masks by their sizes (tried both descending and ascending order) before removing overlaps but it didn't improve my score. I guess <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>'s answer explains why it didn't work. I don't think we can do much about this since models' predictions orders are random just like annotations' orders.</p>",
      "rawMarkdown": "alexandrecc Good question. I was sorting masks by their sizes (tried both descending and ascending order) before removing overlaps but it didn't improve my score. I guess @addisonhoward's answer explains why it didn't work. I don't think we can do much about this since models' predictions orders are random just like annotations' orders.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1553646,
      "author_name": "stupidios",
      "author_url": "",
      "post_date": "10/22/2021 11:23:18",
      "content": "<blockquote>\n  <p>Is there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision…</p>\n</blockquote>\n<p>Yes, this is a very important problem. We can simply remove those overlaps, but we can not guarantee the correctness of the removal unless we know how the pathologists handle such cases.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1555133,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "10/23/2021 15:58:35",
          "content": "<p><a href=\"https://www.kaggle.com/stupidios\" target=\"_blank\">@stupidios</a> exactly!<br>\nI made some analysis of the overlaps <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1571081,
      "author_name": "christoffersartorius",
      "author_url": "",
      "post_date": "11/04/2021 15:29:47",
      "content": "<p>Hi there dr Konya, just wanted to contribute a bit and state that the IoU overlap in the evaluation data is low. The evaluation metric supported by Kaggle either had only one IoU threshold (but with overlap support) or multiple IoUs thresholds but no overlap.</p>\n<p>The amount of overlap in the evaluation dataset is very low (lower than the training dataset) and we therefor went with no-overlap but multiple IoUs evaluation. When we internally work with this data we prefer the <a href=\"https://cocodataset.org/#detection-eval\" target=\"_blank\">coco-mAP </a> metric since it supports both overlaps of objects and evaluation of multiple IoUs. Meaning that our biologists will be able to keep track of overlapping cells in their experiments.</p>\n<p>Kindly,<br>\nChristoffer</p>",
      "votes": null,
      "replies": [
        {
          "id": 1575582,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "11/08/2021 13:39:04",
          "content": "<p><a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> How are the overlapping regions in the test set reduced to a single category? Randomly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1593454,
          "author_name": "alexandrecc",
          "author_url": "",
          "post_date": "11/23/2021 23:42:09",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/christoffersartorius\" target=\"_blank\">@christoffersartorius</a> for the useful and clear information. Since the LB ground truth masks were originally overlapping, can you let us know how the overlapping pixels were assigned to instances in the competition ground truth ? I can think of 4 possibilities:</p>\n<ol>\n<li>Randomly</li>\n<li>Overlapping pixels assigned to the largest overlapping mask</li>\n<li>Overlapping pixels assigned to the smallest overlapping mask</li>\n<li>By the first index appearance when looping trough a list of instances</li>\n</ol>\n<p>We could probably probe the public LB test set with these hypothesis but letting us know officially would be more fair and transparent for everyone since the ground truth between train and LB test is by definition different.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1608570,
          "author_name": "christoffersartorius",
          "author_url": "",
          "post_date": "12/06/2021 13:04:40",
          "content": "<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a>, Hi there! Great questions. The Kaggle team did the conversion so I am not sure what method they used. But i definitely agree that transparency on this would be a benefit for everyone involved, including us as host.</p>\n<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>, do you the possibility to find the response to these questions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1608807,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "12/06/2021 17:11:25",
          "content": "<p>Hi there - The correct answer is Option 4 - we assigned them in index order</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1614838,
          "author_name": "thangvubka",
          "author_url": "",
          "post_date": "12/11/2021 14:09:51",
          "content": "<p><a href=\"https://www.kaggle.com/ChristofferSartorius\" target=\"_blank\">@ChristofferSartorius</a> I am curious why the evaluation set has less overlap than the training set. Is the evaluation set is annotated differently (e.g., by experts) or ambiguously-overlapped cells are merged into one cell?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1621769,
          "author_name": "deeeeeeeplearning",
          "author_url": "",
          "post_date": "12/18/2021 03:28:14",
          "content": "<p>But an important question is in what order are these instances listed. Confidence? Randomness?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1621797,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "12/18/2021 04:16:12",
          "content": "<p>What do you mean? Ground truth labels doesn't have confidence. On the other note, NOICE ranking😉.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1621835,
          "author_name": "deeeeeeeplearning",
          "author_url": "",
          "post_date": "12/18/2021 05:55:18",
          "content": "<p>The order of the predicted label of the instances.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1622430,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/18/2021 18:16:27",
          "content": "<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a> Good question. I was sorting masks by their sizes (tried both descending and ascending order) before removing overlaps but it didn't improve my score. I guess <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a>'s answer explains why it didn't work. I don't think we can do much about this since models' predictions orders are random just like annotations' orders.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1579568,
      "author_name": "tolgadincer",
      "author_url": "",
      "post_date": "11/12/2021 02:16:45",
      "content": "<blockquote>\n  <p>how many cases are there of overlapping areas at all…</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287029\" target=\"_blank\">Here</a> is the answer of your question. Three images contain very high fraction of overlap.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1551793": "The competition states:\n\n> Note: while predictions are not allowed to overlap, the training labels are provided in full (with overlapping portions included). This is to ensure that models are provided the full data for each object. Removing overlap in predictions is a task for the competitor.\n\n![overlap](https://i.postimg.cc/nLvMfSy7/ambiguous.jpg)\n\nThis means we will have **possibly** more instances assigned to the **same pixel** if we use an instance segmentation network... but the objects are transparent, this means the overlapping areas are equivalent (at least theoretically - or do i miss something here?)\n\nDo the contours cross each other in a way one can say - hey, this cell it OVER the other? If so, the pixels of the overlapping areas could equally belong to the instances == not ambiguous.\n\nIs there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision... \n\n... anyway, how many cases are there of overlapping areas at all... if like 0.1% the cases do we care... what if 3%?\n\n ;)",
    "1553646": "> Is there any standard how the pathologists handle such cases, could it be then the annotator's own decision and thus subject to a very subjective decision…\n\nYes, this is a very important problem. We can simply remove those overlaps, but we can not guarantee the correctness of the removal unless we know how the pathologists handle such cases.",
    "1555133": "stupidios exactly!\nI made some analysis of the overlaps [here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281084).",
    "1571081": "Hi there dr Konya, just wanted to contribute a bit and state that the IoU overlap in the evaluation data is low. The evaluation metric supported by Kaggle either had only one IoU threshold (but with overlap support) or multiple IoUs thresholds but no overlap.\n\nThe amount of overlap in the evaluation dataset is very low (lower than the training dataset) and we therefor went with no-overlap but multiple IoUs evaluation. When we internally work with this data we prefer the [coco-mAP ](https://cocodataset.org/#detection-eval) metric since it supports both overlaps of objects and evaluation of multiple IoUs. Meaning that our biologists will be able to keep track of overlapping cells in their experiments.\n\nKindly,\nChristoffer",
    "1575582": "christoffersartorius How are the overlapping regions in the test set reduced to a single category? Randomly?",
    "1579568": "> how many cases are there of overlapping areas at all…\n\n@sandorkonya  [Here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287029) is the answer of your question. Three images contain very high fraction of overlap.",
    "1593454": "Thanks a lot @christoffersartorius for the useful and clear information. Since the LB ground truth masks were originally overlapping, can you let us know how the overlapping pixels were assigned to instances in the competition ground truth ? I can think of 4 possibilities:\n1.  Randomly\n2. Overlapping pixels assigned to the largest overlapping mask\n3. Overlapping pixels assigned to the smallest overlapping mask\n4. By the first index appearance when looping trough a list of instances\n\nWe could probably probe the public LB test set with these hypothesis but letting us know officially would be more fair and transparent for everyone since the ground truth between train and LB test is by definition different.",
    "1608570": "alexandrecc, Hi there! Great questions. The Kaggle team did the conversion so I am not sure what method they used. But i definitely agree that transparency on this would be a benefit for everyone involved, including us as host.\n\n@addisonhoward, do you the possibility to find the response to these questions?",
    "1608807": "Hi there - The correct answer is Option 4 - we assigned them in index order",
    "1614838": "ChristofferSartorius I am curious why the evaluation set has less overlap than the training set. Is the evaluation set is annotated differently (e.g., by experts) or ambiguously-overlapped cells are merged into one cell?",
    "1621769": "But an important question is in what order are these instances listed. Confidence? Randomness?",
    "1621797": "What do you mean? Ground truth labels doesn't have confidence. On the other note, NOICE ranking😉.",
    "1621835": "The order of the predicted label of the instances.",
    "1622430": "alexandrecc Good question. I was sorting masks by their sizes (tried both descending and ascending order) before removing overlaps but it didn't improve my score. I guess @addisonhoward's answer explains why it didn't work. I don't think we can do much about this since models' predictions orders are random just like annotations' orders."
  },
  "source": "meta"
}