{
  "id": 287163,
  "title": "Understanding COCO metrics",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/287163",
  "author_name": "",
  "post_date": "2021-11-12T13:22:00.933901400Z",
  "votes": 50,
  "comment_count": 3,
  "views": 0,
  "content": "<p>If you lately read some paper on segmentation, trained a model with detectron, or looked at <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285927\" target=\"_blank\">LIVECell results</a>. You've seen numbers like these:</p>\n<pre><code>Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.212\nAverage Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=2000 ] = 0.491\nAverage Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=2000 ] = 0.150\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.201 \nAverage Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.207\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.114\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.252\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=500 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.287\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.293\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.184\n</code></pre>\n<p>If like me you haven't worked on image segmentation before it all can be confusing. So I'm going to explain what each of these things mean and put it in the context of this competition.</p>\n<h3>COCO</h3>\n<p>It's the most common dataset and benchmark for image object detection, segmentation, and captioning. Most of the tools and papers on image segmentation report scores on COCO. Other datasets like LIVECell. Also adopt the same (or similar) metrics. You can see the official evaluation description here: <a href=\"https://cocodataset.org/#detection-eval\" target=\"_blank\">https://cocodataset.org/#detection-eval</a></p>\n<h3>IoU</h3>\n<p>Intersection over Union - this is a measure of overlap between two objects, in our case two bitmasks. You take the area of the common part (intersection) and divide by the total area (union)</p>\n<h3>Average Precision</h3>\n<p>As you know Precision simply tells us what ratio of our predictions are correct. When we want to apply it to an instance segmentation problem there are two questions we need to answer:</p>\n<ul>\n<li>What constitutes a correct prediction - The targets and predictions are masks with different shapes, that rarely match exactly. That's why we use IoU. We consider a prediction correct if the IoU between it and a target is greater or equal than some threshold. That's <code>AP@[Iou=0.50]</code> in our table - precision at the IoU threshold 0.5. <code>AP @[ IoU=0.50:0.95]</code> means calculating it a 10 different thresholds and averaging the results.</li>\n<li>How many predictions to consider - A model can output hundreds or even thousands of predictions, each with a different confidence score. Usually adding more predictions will hurt our precision, but improve recall. The way to handle it is: Rather than calculation Precision on a fixed set predictions we calculate the area under the Precision-Recall curve. <a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">Here</a> I found a blog post which explains it in detail and with illustrations. This way adding low confidence predictions won't hurt the score, but won't improve it either. <strong>Note that this is different from this competition metric where it's on the competitor to decide which predictions to include and it has a large effect on the score received</strong></li>\n</ul>\n<h3>Average Recall</h3>\n<p>Like in the case of precision we calculate it over different IoU thresholds, but we don't do any area under the curve calculation, instead we take a fixed number of predictions with the highest score. That's the <code>maxDets</code> in the table. I'm using the values from the <a href=\"https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py\" target=\"_blank\">LIVECell version</a> of the evaluation function, showing recall at 100, 500 and 2000 predictions. Which makes more sense in the case of our images with many cells, versus what COCO uses for pictures of real life objects 1, 10, and 100.</p>\n<h3>Scales</h3>\n<p>Targets are divided into three classes (small, medium, large) based on the area of their bitmask. Again I'm using the values from the LIVECell repository better suited for cells. Small - area under 18 ** 2, medium - between 18 ** 2 and 31 ** 2, large - all above that. </p>\n<h3>Putting it all together</h3>\n<p>Now we can understand the results of my experiment I've pasted at the top. The first row shows the Average Precision  over the 10 different IoU thresholds is 0.212 - this is the most similar to the competition metric and probably the number you care about the most. We can see that AP it's much higher if we only consider only the 0.5 IoU but then falls sharply at 0.75. This supports <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281205\" target=\"_blank\">the discussion</a> that high IoU thresholds are hard to reach on this dataset.</p>\n<p>As far as sizes go my model works the best on the medium targets and much worse on the large - that's something worth further investigation.</p>\n<p>Looking at recall we can see there is a benefit in taking 500 rather than 100 predictions but no extra gain from taking 2000. Also all in all it's only able to find less than 30% of the targets, but that's averaged over IoU threshold, might be interesting to look at the 0.50 alone.</p>",
  "messages": [
    {
      "id": "1580096",
      "postDate": "11/12/2021 13:22:00",
      "content": "<p>If you lately read some paper on segmentation, trained a model with detectron, or looked at <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285927\" target=\"_blank\">LIVECell results</a>. You've seen numbers like these:</p>\n<pre><code>Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.212\nAverage Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=2000 ] = 0.491\nAverage Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=2000 ] = 0.150\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.201 \nAverage Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.207\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.114\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.252\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=500 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.287\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.293\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.184\n</code></pre>\n<p>If like me you haven't worked on image segmentation before it all can be confusing. So I'm going to explain what each of these things mean and put it in the context of this competition.</p>\n<h3>COCO</h3>\n<p>It's the most common dataset and benchmark for image object detection, segmentation, and captioning. Most of the tools and papers on image segmentation report scores on COCO. Other datasets like LIVECell. Also adopt the same (or similar) metrics. You can see the official evaluation description here: <a href=\"https://cocodataset.org/#detection-eval\" target=\"_blank\">https://cocodataset.org/#detection-eval</a></p>\n<h3>IoU</h3>\n<p>Intersection over Union - this is a measure of overlap between two objects, in our case two bitmasks. You take the area of the common part (intersection) and divide by the total area (union)</p>\n<h3>Average Precision</h3>\n<p>As you know Precision simply tells us what ratio of our predictions are correct. When we want to apply it to an instance segmentation problem there are two questions we need to answer:</p>\n<ul>\n<li>What constitutes a correct prediction - The targets and predictions are masks with different shapes, that rarely match exactly. That's why we use IoU. We consider a prediction correct if the IoU between it and a target is greater or equal than some threshold. That's <code>AP@[Iou=0.50]</code> in our table - precision at the IoU threshold 0.5. <code>AP @[ IoU=0.50:0.95]</code> means calculating it a 10 different thresholds and averaging the results.</li>\n<li>How many predictions to consider - A model can output hundreds or even thousands of predictions, each with a different confidence score. Usually adding more predictions will hurt our precision, but improve recall. The way to handle it is: Rather than calculation Precision on a fixed set predictions we calculate the area under the Precision-Recall curve. <a href=\"https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173\" target=\"_blank\">Here</a> I found a blog post which explains it in detail and with illustrations. This way adding low confidence predictions won't hurt the score, but won't improve it either. <strong>Note that this is different from this competition metric where it's on the competitor to decide which predictions to include and it has a large effect on the score received</strong></li>\n</ul>\n<h3>Average Recall</h3>\n<p>Like in the case of precision we calculate it over different IoU thresholds, but we don't do any area under the curve calculation, instead we take a fixed number of predictions with the highest score. That's the <code>maxDets</code> in the table. I'm using the values from the <a href=\"https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py\" target=\"_blank\">LIVECell version</a> of the evaluation function, showing recall at 100, 500 and 2000 predictions. Which makes more sense in the case of our images with many cells, versus what COCO uses for pictures of real life objects 1, 10, and 100.</p>\n<h3>Scales</h3>\n<p>Targets are divided into three classes (small, medium, large) based on the area of their bitmask. Again I'm using the values from the LIVECell repository better suited for cells. Small - area under 18 ** 2, medium - between 18 ** 2 and 31 ** 2, large - all above that. </p>\n<h3>Putting it all together</h3>\n<p>Now we can understand the results of my experiment I've pasted at the top. The first row shows the Average Precision  over the 10 different IoU thresholds is 0.212 - this is the most similar to the competition metric and probably the number you care about the most. We can see that AP it's much higher if we only consider only the 0.5 IoU but then falls sharply at 0.75. This supports <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281205\" target=\"_blank\">the discussion</a> that high IoU thresholds are hard to reach on this dataset.</p>\n<p>As far as sizes go my model works the best on the medium targets and much worse on the large - that's something worth further investigation.</p>\n<p>Looking at recall we can see there is a benefit in taking 500 rather than 100 predictions but no extra gain from taking 2000. Also all in all it's only able to find less than 30% of the targets, but that's averaged over IoU threshold, might be interesting to look at the 0.50 alone.</p>",
      "rawMarkdown": "If you lately read some paper on segmentation, trained a model with detectron, or looked at [LIVECell results](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285927). You've seen numbers like these:\n```\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.212\nAverage Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=2000 ] = 0.491\nAverage Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=2000 ] = 0.150\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.201 \nAverage Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.207\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.114\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.252\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=500 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.287\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.293\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.184\n```\nIf like me you haven't worked on image segmentation before it all can be confusing. So I'm going to explain what each of these things mean and put it in the context of this competition.\n### COCO\nIt's the most common dataset and benchmark for image object detection, segmentation, and captioning. Most of the tools and papers on image segmentation report scores on COCO. Other datasets like LIVECell. Also adopt the same (or similar) metrics. You can see the official evaluation description here: https://cocodataset.org/#detection-eval\n### IoU\nIntersection over Union - this is a measure of overlap between two objects, in our case two bitmasks. You take the area of the common part (intersection) and divide by the total area (union)\n### Average Precision \nAs you know Precision simply tells us what ratio of our predictions are correct. When we want to apply it to an instance segmentation problem there are two questions we need to answer:\n- What constitutes a correct prediction - The targets and predictions are masks with different shapes, that rarely match exactly. That's why we use IoU. We consider a prediction correct if the IoU between it and a target is greater or equal than some threshold. That's `AP@[Iou=0.50]` in our table - precision at the IoU threshold 0.5. `AP @[ IoU=0.50:0.95]` means calculating it a 10 different thresholds and averaging the results.\n- How many predictions to consider - A model can output hundreds or even thousands of predictions, each with a different confidence score. Usually adding more predictions will hurt our precision, but improve recall. The way to handle it is: Rather than calculation Precision on a fixed set predictions we calculate the area under the Precision-Recall curve. [Here](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173) I found a blog post which explains it in detail and with illustrations. This way adding low confidence predictions won't hurt the score, but won't improve it either. **Note that this is different from this competition metric where it's on the competitor to decide which predictions to include and it has a large effect on the score received**\n### Average Recall\nLike in the case of precision we calculate it over different IoU thresholds, but we don't do any area under the curve calculation, instead we take a fixed number of predictions with the highest score. That's the `maxDets` in the table. I'm using the values from the [LIVECell version](https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py) of the evaluation function, showing recall at 100, 500 and 2000 predictions. Which makes more sense in the case of our images with many cells, versus what COCO uses for pictures of real life objects 1, 10, and 100.\n### Scales\nTargets are divided into three classes (small, medium, large) based on the area of their bitmask. Again I'm using the values from the LIVECell repository better suited for cells. Small - area under 18 ** 2, medium - between 18 ** 2 and 31 ** 2, large - all above that. \n### Putting it all together\nNow we can understand the results of my experiment I've pasted at the top. The first row shows the Average Precision  over the 10 different IoU thresholds is 0.212 - this is the most similar to the competition metric and probably the number you care about the most. We can see that AP it's much higher if we only consider only the 0.5 IoU but then falls sharply at 0.75. This supports [the discussion](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281205) that high IoU thresholds are hard to reach on this dataset.\n\nAs far as sizes go my model works the best on the medium targets and much worse on the large - that's something worth further investigation.\n\nLooking at recall we can see there is a benefit in taking 500 rather than 100 predictions but no extra gain from taking 2000. Also all in all it's only able to find less than 30% of the targets, but that's averaged over IoU threshold, might be interesting to look at the 0.50 alone.",
      "votes": null
    },
    {
      "id": "1580102",
      "postDate": "11/12/2021 13:27:05",
      "content": "<p>And this is how I generate those numbers in detectron2 training:</p>\n<pre><code>@classmethod\ndef build_evaluator(cls, cfg, dataset_name, output_folder=None):\n    return COCOEvaluator(dataset_name, cfg, False, output_folder)\n</code></pre>",
      "rawMarkdown": "And this is how I generate those numbers in detectron2 training:\n```\n@classmethod\ndef build_evaluator(cls, cfg, dataset_name, output_folder=None):\n    return COCOEvaluator(dataset_name, cfg, False, output_folder)\n```",
      "votes": null
    },
    {
      "id": "1652083",
      "postDate": "01/16/2022 10:22:03",
      "content": "<p>What do you think is the reason for the lower score of the large segmentation targets?</p>",
      "rawMarkdown": "What do you think is the reason for the lower score of the large segmentation targets?",
      "votes": null
    },
    {
      "id": "1652087",
      "postDate": "01/16/2022 10:24:41",
      "content": "<p>I think: the big categories are mostly SHSY5Y or Astro, which are inherently more difficult than cort. Do you have any additional explanation for this?</p>",
      "rawMarkdown": "I think: the big categories are mostly SHSY5Y or Astro, which are inherently more difficult than cort. Do you have any additional explanation for this?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1580102,
      "author_name": "slawekbiel",
      "author_url": "",
      "post_date": "11/12/2021 13:27:05",
      "content": "<p>And this is how I generate those numbers in detectron2 training:</p>\n<pre><code>@classmethod\ndef build_evaluator(cls, cfg, dataset_name, output_folder=None):\n    return COCOEvaluator(dataset_name, cfg, False, output_folder)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1652083,
      "author_name": "zaopolearning",
      "author_url": "",
      "post_date": "01/16/2022 10:22:03",
      "content": "<p>What do you think is the reason for the lower score of the large segmentation targets?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1652087,
      "author_name": "zaopolearning",
      "author_url": "",
      "post_date": "01/16/2022 10:24:41",
      "content": "<p>I think: the big categories are mostly SHSY5Y or Astro, which are inherently more difficult than cort. Do you have any additional explanation for this?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1580096": "If you lately read some paper on segmentation, trained a model with detectron, or looked at [LIVECell results](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/285927). You've seen numbers like these:\n```\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.212\nAverage Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=2000 ] = 0.491\nAverage Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=2000 ] = 0.150\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.201 \nAverage Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.207\nAverage Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.114\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.252\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=500 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=2000 ] = 0.294\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=2000 ] = 0.287\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=2000 ] = 0.293\nAverage Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=2000 ] = 0.184\n```\nIf like me you haven't worked on image segmentation before it all can be confusing. So I'm going to explain what each of these things mean and put it in the context of this competition.\n### COCO\nIt's the most common dataset and benchmark for image object detection, segmentation, and captioning. Most of the tools and papers on image segmentation report scores on COCO. Other datasets like LIVECell. Also adopt the same (or similar) metrics. You can see the official evaluation description here: https://cocodataset.org/#detection-eval\n### IoU\nIntersection over Union - this is a measure of overlap between two objects, in our case two bitmasks. You take the area of the common part (intersection) and divide by the total area (union)\n### Average Precision \nAs you know Precision simply tells us what ratio of our predictions are correct. When we want to apply it to an instance segmentation problem there are two questions we need to answer:\n- What constitutes a correct prediction - The targets and predictions are masks with different shapes, that rarely match exactly. That's why we use IoU. We consider a prediction correct if the IoU between it and a target is greater or equal than some threshold. That's `AP@[Iou=0.50]` in our table - precision at the IoU threshold 0.5. `AP @[ IoU=0.50:0.95]` means calculating it a 10 different thresholds and averaging the results.\n- How many predictions to consider - A model can output hundreds or even thousands of predictions, each with a different confidence score. Usually adding more predictions will hurt our precision, but improve recall. The way to handle it is: Rather than calculation Precision on a fixed set predictions we calculate the area under the Precision-Recall curve. [Here](https://jonathan-hui.medium.com/map-mean-average-precision-for-object-detection-45c121a31173) I found a blog post which explains it in detail and with illustrations. This way adding low confidence predictions won't hurt the score, but won't improve it either. **Note that this is different from this competition metric where it's on the competitor to decide which predictions to include and it has a large effect on the score received**\n### Average Recall\nLike in the case of precision we calculate it over different IoU thresholds, but we don't do any area under the curve calculation, instead we take a fixed number of predictions with the highest score. That's the `maxDets` in the table. I'm using the values from the [LIVECell version](https://github.com/sartorius-research/LIVECell/blob/main/code/coco_evaluation.py) of the evaluation function, showing recall at 100, 500 and 2000 predictions. Which makes more sense in the case of our images with many cells, versus what COCO uses for pictures of real life objects 1, 10, and 100.\n### Scales\nTargets are divided into three classes (small, medium, large) based on the area of their bitmask. Again I'm using the values from the LIVECell repository better suited for cells. Small - area under 18 ** 2, medium - between 18 ** 2 and 31 ** 2, large - all above that. \n### Putting it all together\nNow we can understand the results of my experiment I've pasted at the top. The first row shows the Average Precision  over the 10 different IoU thresholds is 0.212 - this is the most similar to the competition metric and probably the number you care about the most. We can see that AP it's much higher if we only consider only the 0.5 IoU but then falls sharply at 0.75. This supports [the discussion](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/281205) that high IoU thresholds are hard to reach on this dataset.\n\nAs far as sizes go my model works the best on the medium targets and much worse on the large - that's something worth further investigation.\n\nLooking at recall we can see there is a benefit in taking 500 rather than 100 predictions but no extra gain from taking 2000. Also all in all it's only able to find less than 30% of the targets, but that's averaged over IoU threshold, might be interesting to look at the 0.50 alone.",
    "1580102": "And this is how I generate those numbers in detectron2 training:\n```\n@classmethod\ndef build_evaluator(cls, cfg, dataset_name, output_folder=None):\n    return COCOEvaluator(dataset_name, cfg, False, output_folder)\n```",
    "1652083": "What do you think is the reason for the lower score of the large segmentation targets?",
    "1652087": "I think: the big categories are mostly SHSY5Y or Astro, which are inherently more difficult than cort. Do you have any additional explanation for this?"
  },
  "source": "meta"
}