{
  "id": 214659,
  "title": "How to do local evaluation?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/214659",
  "author_name": "Chenglu",
  "post_date": "2021-01-27T10:16:34.511000",
  "votes": 9,
  "comment_count": 14,
  "views": 0,
  "content": "<p>As the images are \"weak labeled\", how can we do local evaluation without the label of the mask?</p>",
  "messages": [
    {
      "id": 1172248,
      "postDate": "2021-01-27T10:16:34.510Z",
      "content": "<p>As the images are \"weak labeled\", how can we do local evaluation without the label of the mask?</p>",
      "rawMarkdown": "As the images are \"weak labeled\", how can we do local evaluation without the label of the mask?",
      "votes": 9
    },
    {
      "id": 1172998,
      "postDate": "2021-01-27T16:39:22.063Z",
      "content": "<p>Semantic segmentation with only weak classification labels — that's what makes this competition sexy 😄</p>",
      "rawMarkdown": "Semantic segmentation with only weak classification labels &mdash; that's what makes this competition sexy 😄",
      "votes": 7,
      "replies": [
        {
          "id": 1173425,
          "postDate": "2021-01-27T21:37:30.287Z",
          "content": "<p>Isn’t the competition for instance segmentation not semantic segmentation?</p>",
          "rawMarkdown": "Isn’t the competition for instance segmentation not semantic segmentation?",
          "votes": 2
        },
        {
          "id": 1173621,
          "postDate": "2021-01-28T02:30:57.307Z",
          "content": "<p>Soooo sexy, never met a competition like this 🤒</p>",
          "rawMarkdown": "Soooo sexy, never met a competition like this 🤒"
        }
      ]
    },
    {
      "id": 1172968,
      "postDate": "2021-01-27T16:22:07.190Z",
      "content": "<p>I had the same question <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>. At first glance it looks like a multi-label classification challenge but when I saw this with the evaluation tab:</p>\n<blockquote>\n  <p>The OpenImages version of the metric is described in detail here. See also this tutorial on running the evaluation in Python, with the only difference being the use of F1 rather than average precision.</p>\n  <p>Segmentation is calculated using IoU with a threshold of 0.6</p>\n</blockquote>\n<p>so it leads me to believe its label + masks that are used for the evaluation but there are not ground truth / bounding boxes masks provided.</p>\n<p>I could be mistaken though and looking to much into it.</p>\n<p><strong>UPDATE</strong>:<br>\nif I actually click into the link that describes the metric in detail it leads here - <a href=\"https://storage.googleapis.com/openimages/web/evaluation.html#instance_segmentation_eval\" target=\"_blank\">https://storage.googleapis.com/openimages/web/evaluation.html#instance_segmentation_eval</a></p>\n<p>and there is a description that include ground truth masks:</p>\n<blockquote>\n  <p>For each image:</p>\n  <p>First, all detection masks are matched to all existing ground-truth masks. Two masks are considered as a potential match if mask-to-mask IoU &gt; 0.5. Matched detections are considered True Positives.<br>\n  The remaining unmatched detections are now matched with ground-truth boxes that do not have a corresponding mask and that are not marked as group-of. Two such boxes are considered a potential match if box-to-box IoU &gt; 0.5. Matched detections are ignored. These are probably true detections, but no ground-truth mask is available to evaluate in detail.<br>\n  The remaining unmatched detections are finally matched to ground-truth group-of boxes. Two such boxes are considered a potential match if box-to-box intersection over area &gt; 0.5 (like in the object detection case, multiple detections can match the same group-of box). Matched detections are ignored. The remaining non-matched detections do not match any mask nor box, and are thus considered False Positives.</p>\n  <p>After this three-stage matching all detections have been tagged as True Positive, False Positive or to be ignored, and precision/recall values can be computed to generate per-class AP values.</p>\n</blockquote>\n<p>and they state the only difference is at the end here instead of using AP we use F1. I think another difference would also be that mask-to-mask IoU &gt; 0.6.</p>",
      "rawMarkdown": "I had the same question @snaker. At first glance it looks like a multi-label classification challenge but when I saw this with the evaluation tab:\n\n> The OpenImages version of the metric is described in detail here. See also this tutorial on running the evaluation in Python, with the only difference being the use of F1 rather than average precision.\n\n> Segmentation is calculated using IoU with a threshold of 0.6\n\nso it leads me to believe its label + masks that are used for the evaluation but there are not ground truth / bounding boxes masks provided.\n\nI could be mistaken though and looking to much into it.\n\n**UPDATE**:\nif I actually click into the link that describes the metric in detail it leads here - https://storage.googleapis.com/openimages/web/evaluation.html#instance_segmentation_eval\n\nand there is a description that include ground truth masks:\n\n> For each image:\n\n> First, all detection masks are matched to all existing ground-truth masks. Two masks are considered as a potential match if mask-to-mask IoU > 0.5. Matched detections are considered True Positives.\nThe remaining unmatched detections are now matched with ground-truth boxes that do not have a corresponding mask and that are not marked as group-of. Two such boxes are considered a potential match if box-to-box IoU > 0.5. Matched detections are ignored. These are probably true detections, but no ground-truth mask is available to evaluate in detail.\nThe remaining unmatched detections are finally matched to ground-truth group-of boxes. Two such boxes are considered a potential match if box-to-box intersection over area > 0.5 (like in the object detection case, multiple detections can match the same group-of box). Matched detections are ignored. The remaining non-matched detections do not match any mask nor box, and are thus considered False Positives.\n\n> After this three-stage matching all detections have been tagged as True Positive, False Positive or to be ignored, and precision/recall values can be computed to generate per-class AP values.\n\nand they state the only difference is at the end here instead of using AP we use F1. I think another difference would also be that mask-to-mask IoU > 0.6.",
      "votes": 1,
      "replies": [
        {
          "id": 1173074,
          "postDate": "2021-01-27T17:08:28.027Z",
          "content": "<p>This is a weakly-labeled multi-labeled classification challenge. And you need to do 2 tasks here: identify the cell (segmentation) and classify it. The metrics works like this: For each image, all predicted masks are matched with ground-truth masks (this is not provided, but you can easily get this by any type of segmentation model, an example here <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a>.  . Then mAP is calculated. If the model predict a class 0 wrongly for cell 1 in image A, this is counted as 1 False Positive for class 0. If the model predicts a class on a region of the image that is not a cell, this is also counted as 1 False Positive for that class. </p>",
          "rawMarkdown": "This is a weakly-labeled multi-labeled classification challenge. And you need to do 2 tasks here: identify the cell (segmentation) and classify it. The metrics works like this: For each image, all predicted masks are matched with ground-truth masks (this is not provided, but you can easily get this by any type of segmentation model, an example here https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg.  ~~Then F1 is calculated for each class on these cells. Then macro-F1 is calculated for all classes. So ideally you want to maximize the macro-F1 of all cells in the images with defined patterns (high confidence score for a class)~~. Then mAP is calculated. If the model predict a class 0 wrongly for cell 1 in image A, this is counted as 1 False Positive for class 0. If the model predicts a class on a region of the image that is not a cell, this is also counted as 1 False Positive for that class. ",
          "votes": 6
        },
        {
          "id": 1173229,
          "postDate": "2021-01-27T18:34:45.110Z",
          "content": "<p>Thank you for the explanation <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>. The link is leading to a 404 error though:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291298%2Fb8afbe951cc83c92dfd405af7f112317%2FScreen%20Shot%202021-01-27%20at%2010.34.19%20AM.png?generation=1611772475581218&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you for the explanation @lnhtrang. The link is leading to a 404 error though:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291298%2Fb8afbe951cc83c92dfd405af7f112317%2FScreen%20Shot%202021-01-27%20at%2010.34.19%20AM.png?generation=1611772475581218&alt=media)\n"
        },
        {
          "id": 1173339,
          "postDate": "2021-01-27T20:03:38.433Z",
          "content": "<p>Updated! typo ')' in the link :)</p>",
          "rawMarkdown": "Updated! typo ')' in the link :)",
          "votes": 1
        },
        {
          "id": 1173676,
          "postDate": "2021-01-28T03:51:08.453Z",
          "content": "<p>Thanks for the great explaination. Just one question</p>\n<blockquote>\n  <p>\"you need to do 2 tasks here: identify the cell (segmentation) and classify it\"</p>\n</blockquote>\n<p>Actually we are classifying every pixel of the image right?</p>",
          "rawMarkdown": "Thanks for the great explaination. Just one question\n\n> \"you need to do 2 tasks here: identify the cell (segmentation) and classify it\"\n\nActually we are classifying every pixel of the image right?"
        },
        {
          "id": 1174316,
          "postDate": "2021-01-28T12:01:34.783Z",
          "content": "<p>You can certainly view the problem as such!</p>",
          "rawMarkdown": "You can certainly view the problem as such!"
        },
        {
          "id": 1174380,
          "postDate": "2021-01-28T12:29:15.797Z",
          "content": "<p>It depends on your approach to solve the problem. What we are asking you to submit is labels per cell (not per pixel).</p>",
          "rawMarkdown": "It depends on your approach to solve the problem. What we are asking you to submit is labels per cell (not per pixel)."
        },
        {
          "id": 1179923,
          "postDate": "2021-02-01T01:27:54.683Z",
          "content": "<p>Does this mean that every single cell should just belongs to ony one cell type? (so the pixel type of one cell should all be the same type)</p>",
          "rawMarkdown": "Does this mean that every single cell should just belongs to ony one cell type? (so the pixel type of one cell should all be the same type)"
        },
        {
          "id": 1180264,
          "postDate": "2021-02-01T07:35:57.830Z",
          "content": "<p>In the challenge dataset there are 17 different cell types. All cells in one image are always the same cell type. The task is perform 'protein organelle classification label' per cell, which may differ between the cells in one image.</p>",
          "rawMarkdown": "In the challenge dataset there are 17 different cell types. All cells in one image are always the same cell type. The task is perform 'protein organelle classification label' per cell, which may differ between the cells in one image."
        },
        {
          "id": 1180659,
          "postDate": "2021-02-01T11:59:10.303Z",
          "content": "<p>Hey.</p>\n<p>We will be using instance segmentation to perform segmentation (pixel wise prediction) to identify masks for all cells in an image.</p>\n<p>You will then have to predict, for a given instance/cell, what the labels are (which organelles is the protein of interest found in).</p>",
          "rawMarkdown": "Hey.\n\nWe will be using instance segmentation to perform segmentation (pixel wise prediction) to identify masks for all cells in an image.\n\nYou will then have to predict, for a given instance/cell, what the labels are (which organelles is the protein of interest found in)."
        },
        {
          "id": 1220281,
          "postDate": "2021-02-27T20:00:15.677Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1172998,
      "author_name": "Chan Kha Vu",
      "author_url": "",
      "post_date": "2021-01-27T16:39:22.063000",
      "content": "<p>Semantic segmentation with only weak classification labels — that's what makes this competition sexy 😄</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1173425,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-01-27T21:37:30.287000",
          "content": "<p>Isn’t the competition for instance segmentation not semantic segmentation?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1173621,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2021-01-28T02:30:57.307000",
          "content": "<p>Soooo sexy, never met a competition like this 🤒</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1172968,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2021-01-27T16:22:07.190000",
      "content": "<p>I had the same question <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>. At first glance it looks like a multi-label classification challenge but when I saw this with the evaluation tab:</p>\n<blockquote>\n  <p>The OpenImages version of the metric is described in detail here. See also this tutorial on running the evaluation in Python, with the only difference being the use of F1 rather than average precision.</p>\n  <p>Segmentation is calculated using IoU with a threshold of 0.6</p>\n</blockquote>\n<p>so it leads me to believe its label + masks that are used for the evaluation but there are not ground truth / bounding boxes masks provided.</p>\n<p>I could be mistaken though and looking to much into it.</p>\n<p><strong>UPDATE</strong>:<br>\nif I actually click into the link that describes the metric in detail it leads here - <a href=\"https://storage.googleapis.com/openimages/web/evaluation.html#instance_segmentation_eval\" target=\"_blank\">https://storage.googleapis.com/openimages/web/evaluation.html#instance_segmentation_eval</a></p>\n<p>and there is a description that include ground truth masks:</p>\n<blockquote>\n  <p>For each image:</p>\n  <p>First, all detection masks are matched to all existing ground-truth masks. Two masks are considered as a potential match if mask-to-mask IoU &gt; 0.5. Matched detections are considered True Positives.<br>\n  The remaining unmatched detections are now matched with ground-truth boxes that do not have a corresponding mask and that are not marked as group-of. Two such boxes are considered a potential match if box-to-box IoU &gt; 0.5. Matched detections are ignored. These are probably true detections, but no ground-truth mask is available to evaluate in detail.<br>\n  The remaining unmatched detections are finally matched to ground-truth group-of boxes. Two such boxes are considered a potential match if box-to-box intersection over area &gt; 0.5 (like in the object detection case, multiple detections can match the same group-of box). Matched detections are ignored. The remaining non-matched detections do not match any mask nor box, and are thus considered False Positives.</p>\n  <p>After this three-stage matching all detections have been tagged as True Positive, False Positive or to be ignored, and precision/recall values can be computed to generate per-class AP values.</p>\n</blockquote>\n<p>and they state the only difference is at the end here instead of using AP we use F1. I think another difference would also be that mask-to-mask IoU &gt; 0.6.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1173074,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-27T17:08:28.027000",
          "content": "<p>This is a weakly-labeled multi-labeled classification challenge. And you need to do 2 tasks here: identify the cell (segmentation) and classify it. The metrics works like this: For each image, all predicted masks are matched with ground-truth masks (this is not provided, but you can easily get this by any type of segmentation model, an example here <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a>.  . Then mAP is calculated. If the model predict a class 0 wrongly for cell 1 in image A, this is counted as 1 False Positive for class 0. If the model predicts a class on a region of the image that is not a cell, this is also counted as 1 False Positive for that class. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1173229,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2021-01-27T18:34:45.110000",
          "content": "<p>Thank you for the explanation <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>. The link is leading to a 404 error though:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291298%2Fb8afbe951cc83c92dfd405af7f112317%2FScreen%20Shot%202021-01-27%20at%2010.34.19%20AM.png?generation=1611772475581218&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1173339,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-27T20:03:38.433000",
          "content": "<p>Updated! typo ')' in the link :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1173676,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2021-01-28T03:51:08.453000",
          "content": "<p>Thanks for the great explaination. Just one question</p>\n<blockquote>\n  <p>\"you need to do 2 tasks here: identify the cell (segmentation) and classify it\"</p>\n</blockquote>\n<p>Actually we are classifying every pixel of the image right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1174316,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-28T12:01:34.783000",
          "content": "<p>You can certainly view the problem as such!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1174380,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-28T12:29:15.797000",
          "content": "<p>It depends on your approach to solve the problem. What we are asking you to submit is labels per cell (not per pixel).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1179923,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2021-02-01T01:27:54.683000",
          "content": "<p>Does this mean that every single cell should just belongs to ony one cell type? (so the pixel type of one cell should all be the same type)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1180264,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-02-01T07:35:57.830000",
          "content": "<p>In the challenge dataset there are 17 different cell types. All cells in one image are always the same cell type. The task is perform 'protein organelle classification label' per cell, which may differ between the cells in one image.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1180659,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-01T11:59:10.303000",
          "content": "<p>Hey.</p>\n<p>We will be using instance segmentation to perform segmentation (pixel wise prediction) to identify masks for all cells in an image.</p>\n<p>You will then have to predict, for a given instance/cell, what the labels are (which organelles is the protein of interest found in).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220281,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-27T20:00:15.677000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1172248": "As the images are \"weak labeled\", how can we do local evaluation without the label of the mask?",
    "1172998": "Semantic segmentation with only weak classification labels &mdash; that's what makes this competition sexy 😄",
    "1172968": "I had the same question @snaker. At first glance it looks like a multi-label classification challenge but when I saw this with the evaluation tab:\n\n> The OpenImages version of the metric is described in detail here. See also this tutorial on running the evaluation in Python, with the only difference being the use of F1 rather than average precision.\n\n> Segmentation is calculated using IoU with a threshold of 0.6\n\nso it leads me to believe its label + masks that are used for the evaluation but there are not ground truth / bounding boxes masks provided.\n\nI could be mistaken though and looking to much into it.\n\n**UPDATE**:\nif I actually click into the link that describes the metric in detail it leads here - https://storage.googleapis.com/openimages/web/evaluation.html#instance_segmentation_eval\n\nand there is a description that include ground truth masks:\n\n> For each image:\n\n> First, all detection masks are matched to all existing ground-truth masks. Two masks are considered as a potential match if mask-to-mask IoU > 0.5. Matched detections are considered True Positives.\nThe remaining unmatched detections are now matched with ground-truth boxes that do not have a corresponding mask and that are not marked as group-of. Two such boxes are considered a potential match if box-to-box IoU > 0.5. Matched detections are ignored. These are probably true detections, but no ground-truth mask is available to evaluate in detail.\nThe remaining unmatched detections are finally matched to ground-truth group-of boxes. Two such boxes are considered a potential match if box-to-box intersection over area > 0.5 (like in the object detection case, multiple detections can match the same group-of box). Matched detections are ignored. The remaining non-matched detections do not match any mask nor box, and are thus considered False Positives.\n\n> After this three-stage matching all detections have been tagged as True Positive, False Positive or to be ignored, and precision/recall values can be computed to generate per-class AP values.\n\nand they state the only difference is at the end here instead of using AP we use F1. I think another difference would also be that mask-to-mask IoU > 0.6."
  }
}