{
  "id": 279200,
  "title": "Is this a One-class Semantic Segmentation Competition?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/279200",
  "author_name": "",
  "post_date": "2021-10-17T04:13:29.182922500Z",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am confused that the problem is not the Instance Segmentation. For some reasons, I think the problem is just the One-class Semantic Segmentation:</p>\n<ul>\n<li>Some public kernels use UNet-like models to solve the problem</li>\n<li>The evaluation metric is mean average precision, but it doesn’t care about the prediction of cell type (the same is true for submission structure)</li>\n</ul>\n<blockquote>\n  <p>A true positive is counted when a single predicted object matches a ground truth object with an IoU above the threshold. A false positive indicates a predicted object had no associated ground truth object. A false negative indicates a ground truth object had no associated predicted object. The average precision of a single image is then calculated as the mean of the above precision values at each IoU threshold</p>\n</blockquote>\n<p>The only difference is that the evaluation metric is mean average precision instead of dice coefficient.<br>\nAm I correct? Can anyone explain or correct me<br>\nThanks in advance</p>",
  "messages": [
    {
      "id": "1547298",
      "postDate": "10/17/2021 04:13:29",
      "content": "<p>I am confused that the problem is not the Instance Segmentation. For some reasons, I think the problem is just the One-class Semantic Segmentation:</p>\n<ul>\n<li>Some public kernels use UNet-like models to solve the problem</li>\n<li>The evaluation metric is mean average precision, but it doesn’t care about the prediction of cell type (the same is true for submission structure)</li>\n</ul>\n<blockquote>\n  <p>A true positive is counted when a single predicted object matches a ground truth object with an IoU above the threshold. A false positive indicates a predicted object had no associated ground truth object. A false negative indicates a ground truth object had no associated predicted object. The average precision of a single image is then calculated as the mean of the above precision values at each IoU threshold</p>\n</blockquote>\n<p>The only difference is that the evaluation metric is mean average precision instead of dice coefficient.<br>\nAm I correct? Can anyone explain or correct me<br>\nThanks in advance</p>",
      "rawMarkdown": "I am confused that the problem is not the Instance Segmentation. For some reasons, I think the problem is just the One-class Semantic Segmentation:\n- Some public kernels use UNet-like models to solve the problem\n- The evaluation metric is mean average precision, but it doesn’t care about the prediction of cell type (the same is true for submission structure)\n> A true positive is counted when a single predicted object matches a ground truth object with an IoU above the threshold. A false positive indicates a predicted object had no associated ground truth object. A false negative indicates a ground truth object had no associated predicted object. The average precision of a single image is then calculated as the mean of the above precision values at each IoU threshold\n\nThe only difference is that the evaluation metric is mean average precision instead of dice coefficient.\nAm I correct? Can anyone explain or correct me\nThanks in advance",
      "votes": null
    },
    {
      "id": "1547366",
      "postDate": "10/17/2021 06:33:28",
      "content": "<p>There are three types of cells represented in the images, but only one type per image. Thus the cell type is not included in the submission. It is probably beneficial to stratify the train/test sets on cell type though.</p>",
      "rawMarkdown": "There are three types of cells represented in the images, but only one type per image. Thus the cell type is not included in the submission. It is probably beneficial to stratify the train/test sets on cell type though.",
      "votes": null
    },
    {
      "id": "1547376",
      "postDate": "10/17/2021 06:48:05",
      "content": "<p>I got it. Thank you <a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a> </p>",
      "rawMarkdown": "I got it. Thank you @mistag",
      "votes": null
    },
    {
      "id": "1548064",
      "postDate": "10/17/2021 23:29:32",
      "content": "<p>No.</p>\n<p>This is an instance segmentation problem.</p>\n<ul>\n<li><a href=\"https://www.linkedin.com/pulse/quick-understanding-instance-segmentation-vs-semantic-rohan-chikorde/\" target=\"_blank\"><strong>See here</strong></a> for a nice post on it on Linkedin</li>\n<li>See below for a visual representation of the differences between Object Detection, Semantic Segmentation (all instances are masked with the same pixel value), and Instance Segmentation (all instances are masked with their own respective value)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/tDbJmgh/1-J1qp-G6-TUYsq43wh-Aqezjyg.png\" alt=\"hi\"></p>\n<hr>\n<p><strong>Think of instance segmentation as Object Detection… except instead of identifying the bounding boxes for all instances, you are required to report the segmentation masks for all instances</strong> </p>",
      "rawMarkdown": "No.\n\nThis is an instance segmentation problem.\n* [**See here**](https://www.linkedin.com/pulse/quick-understanding-instance-segmentation-vs-semantic-rohan-chikorde/) for a nice post on it on Linkedin\n* See below for a visual representation of the differences between Object Detection, Semantic Segmentation (all instances are masked with the same pixel value), and Instance Segmentation (all instances are masked with their own respective value)\n\n![hi](https://i.ibb.co/tDbJmgh/1-J1qp-G6-TUYsq43wh-Aqezjyg.png)\n\n---\n\n**Think of instance segmentation as Object Detection... except instead of identifying the bounding boxes for all instances, you are required to report the segmentation masks for all instances**",
      "votes": null
    },
    {
      "id": "1548126",
      "postDate": "10/18/2021 02:41:00",
      "content": "<p>But there is only one type of cell in an image, and the score is calculated without taking cell type into account. So, we can treat the problem similar to semantic segmentation</p>",
      "rawMarkdown": "But there is only one type of cell in an image, and the score is calculated without taking cell type into account. So, we can treat the problem similar to semantic segmentation",
      "votes": null
    },
    {
      "id": "1548150",
      "postDate": "10/18/2021 03:14:35",
      "content": "<p>No. Please see below for the case with a single class.</p>\n<p><img src=\"https://i.ibb.co/hcjjnMf/0-Qe-Os5-Rv-Xlkb-Dk-LOy.png\" alt=\"\"></p>",
      "rawMarkdown": "No. Please see below for the case with a single class.\n\n![](https://i.ibb.co/hcjjnMf/0-Qe-Os5-Rv-Xlkb-Dk-LOy.png)",
      "votes": null
    },
    {
      "id": "1548157",
      "postDate": "10/18/2021 03:30:46",
      "content": "<p>Yeah, you're right. Thank you <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> <br>\nCan UNet solve this problem in some way?</p>",
      "rawMarkdown": "Yeah, you're right. Thank you @dschettler8845 \nCan UNet solve this problem in some way?",
      "votes": null
    },
    {
      "id": "1548927",
      "postDate": "10/18/2021 16:22:34",
      "content": "<p>UNET is a semantic segmenter.</p>\n<p>I will answer your question by saying what my upcoming baseline solution attempt will be.</p>\n<hr>\n<p><strong>I plan on leveraging EfficientDet to output BOTH bounding boxes and a semantic segmentation map</strong></p>\n<p>This would be equivalent to trying to use the two images on the left (in the above graphic) to generate the image on the right. There are shortcomings to this method, but it's something I want to try.</p>\n<hr>\n<p>In other words. To use UNET, I believe you would have to also build an object detector and use the two of them in concert. Alternatively… you could hope that the semantic segmentation map has a lot of gaps (non-touching segments), because then you could simply call all non-touching segments instances.</p>",
      "rawMarkdown": "UNET is a semantic segmenter.\n\nI will answer your question by saying what my upcoming baseline solution attempt will be.\n\n---\n\n**I plan on leveraging EfficientDet to output BOTH bounding boxes and a semantic segmentation map**\n\nThis would be equivalent to trying to use the two images on the left (in the above graphic) to generate the image on the right. There are shortcomings to this method, but it's something I want to try.\n\n---\n\nIn other words. To use UNET, I believe you would have to also build an object detector and use the two of them in concert. Alternatively... you could hope that the semantic segmentation map has a lot of gaps (non-touching segments), because then you could simply call all non-touching segments instances.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1547366,
      "author_name": "mistag",
      "author_url": "",
      "post_date": "10/17/2021 06:33:28",
      "content": "<p>There are three types of cells represented in the images, but only one type per image. Thus the cell type is not included in the submission. It is probably beneficial to stratify the train/test sets on cell type though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1547376,
          "author_name": "lhkhiem28",
          "author_url": "",
          "post_date": "10/17/2021 06:48:05",
          "content": "<p>I got it. Thank you <a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1548064,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "10/17/2021 23:29:32",
      "content": "<p>No.</p>\n<p>This is an instance segmentation problem.</p>\n<ul>\n<li><a href=\"https://www.linkedin.com/pulse/quick-understanding-instance-segmentation-vs-semantic-rohan-chikorde/\" target=\"_blank\"><strong>See here</strong></a> for a nice post on it on Linkedin</li>\n<li>See below for a visual representation of the differences between Object Detection, Semantic Segmentation (all instances are masked with the same pixel value), and Instance Segmentation (all instances are masked with their own respective value)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/tDbJmgh/1-J1qp-G6-TUYsq43wh-Aqezjyg.png\" alt=\"hi\"></p>\n<hr>\n<p><strong>Think of instance segmentation as Object Detection… except instead of identifying the bounding boxes for all instances, you are required to report the segmentation masks for all instances</strong> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1548126,
          "author_name": "lhkhiem28",
          "author_url": "",
          "post_date": "10/18/2021 02:41:00",
          "content": "<p>But there is only one type of cell in an image, and the score is calculated without taking cell type into account. So, we can treat the problem similar to semantic segmentation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1548150,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "10/18/2021 03:14:35",
          "content": "<p>No. Please see below for the case with a single class.</p>\n<p><img src=\"https://i.ibb.co/hcjjnMf/0-Qe-Os5-Rv-Xlkb-Dk-LOy.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1548157,
          "author_name": "lhkhiem28",
          "author_url": "",
          "post_date": "10/18/2021 03:30:46",
          "content": "<p>Yeah, you're right. Thank you <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> <br>\nCan UNet solve this problem in some way?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1548927,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "10/18/2021 16:22:34",
          "content": "<p>UNET is a semantic segmenter.</p>\n<p>I will answer your question by saying what my upcoming baseline solution attempt will be.</p>\n<hr>\n<p><strong>I plan on leveraging EfficientDet to output BOTH bounding boxes and a semantic segmentation map</strong></p>\n<p>This would be equivalent to trying to use the two images on the left (in the above graphic) to generate the image on the right. There are shortcomings to this method, but it's something I want to try.</p>\n<hr>\n<p>In other words. To use UNET, I believe you would have to also build an object detector and use the two of them in concert. Alternatively… you could hope that the semantic segmentation map has a lot of gaps (non-touching segments), because then you could simply call all non-touching segments instances.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1547298": "I am confused that the problem is not the Instance Segmentation. For some reasons, I think the problem is just the One-class Semantic Segmentation:\n- Some public kernels use UNet-like models to solve the problem\n- The evaluation metric is mean average precision, but it doesn’t care about the prediction of cell type (the same is true for submission structure)\n> A true positive is counted when a single predicted object matches a ground truth object with an IoU above the threshold. A false positive indicates a predicted object had no associated ground truth object. A false negative indicates a ground truth object had no associated predicted object. The average precision of a single image is then calculated as the mean of the above precision values at each IoU threshold\n\nThe only difference is that the evaluation metric is mean average precision instead of dice coefficient.\nAm I correct? Can anyone explain or correct me\nThanks in advance",
    "1547366": "There are three types of cells represented in the images, but only one type per image. Thus the cell type is not included in the submission. It is probably beneficial to stratify the train/test sets on cell type though.",
    "1547376": "I got it. Thank you @mistag",
    "1548064": "No.\n\nThis is an instance segmentation problem.\n* [**See here**](https://www.linkedin.com/pulse/quick-understanding-instance-segmentation-vs-semantic-rohan-chikorde/) for a nice post on it on Linkedin\n* See below for a visual representation of the differences between Object Detection, Semantic Segmentation (all instances are masked with the same pixel value), and Instance Segmentation (all instances are masked with their own respective value)\n\n![hi](https://i.ibb.co/tDbJmgh/1-J1qp-G6-TUYsq43wh-Aqezjyg.png)\n\n---\n\n**Think of instance segmentation as Object Detection... except instead of identifying the bounding boxes for all instances, you are required to report the segmentation masks for all instances**",
    "1548126": "But there is only one type of cell in an image, and the score is calculated without taking cell type into account. So, we can treat the problem similar to semantic segmentation",
    "1548150": "No. Please see below for the case with a single class.\n\n![](https://i.ibb.co/hcjjnMf/0-Qe-Os5-Rv-Xlkb-Dk-LOy.png)",
    "1548157": "Yeah, you're right. Thank you @dschettler8845 \nCan UNet solve this problem in some way?",
    "1548927": "UNET is a semantic segmenter.\n\nI will answer your question by saying what my upcoming baseline solution attempt will be.\n\n---\n\n**I plan on leveraging EfficientDet to output BOTH bounding boxes and a semantic segmentation map**\n\nThis would be equivalent to trying to use the two images on the left (in the above graphic) to generate the image on the right. There are shortcomings to this method, but it's something I want to try.\n\n---\n\nIn other words. To use UNET, I believe you would have to also build an object detector and use the two of them in concert. Alternatively... you could hope that the semantic segmentation map has a lot of gaps (non-touching segments), because then you could simply call all non-touching segments instances."
  },
  "source": "meta"
}