{
  "id": 94245,
  "title": "Help to calculate mAP",
  "url": "/competitions/imaterialist-fashion-2019-FGVC6/discussion/94245",
  "author_name": "",
  "post_date": "2019-06-03T12:32:51.756835100Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi All,</p>\n\n<p>I am trying to calculate the score of my validation set however I am not entirely sure how to calculate the mAP.</p>\n\n<p>To start from the beginning, I am not even sure that my definition of how to calculate the positives and negatives are correct:</p>\n\n<h3>TP:</h3>\n\n<ul>\n<li>IoU of predicted mask to gt mask is greater than threshold</li>\n</ul>\n\n<h3>FP:</h3>\n\n<ul>\n<li>Evaluation page clearly states this is when a predicted class has no associated ground truth. Ie predict class A but the image does not have class A in it.</li>\n<li>But, is it also when IoU of predicted mask to gt mask is LESS than threshold?</li>\n</ul>\n\n<h3>FN:</h3>\n\n<ul>\n<li>Again evaluation page is pretty clear that a ground truth object has no predicted object. Ie image has A in it but we don't even predict A.</li>\n<li>But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?</li>\n</ul>\n\n<p>Even if I am clear on how to classify positives and negatives, I am still unclear on how to calculate precision from them. On the evaluation page,  it seems to suggest for each ClassId and ImageId pair, we calculate the precision according to the given equation. However this does not seem right to me, since in most cases there is only one or two of each ClassId for an ImageId. That doesn't seem like enough to calculate a precision from. In most cases, that fraction is going to be 0 or 1.</p>\n\n<blockquote>\n  <p>ImageId ClassID pairs with exactly  1 entries: 173160\n  ImageId ClassID pairs with exactly  2 entries: 62687\n  ImageId ClassID pairs with exactly  3 entries: 4242\n  ImageId ClassID pairs with exactly  4 entries: 2882\n  ImageId ClassID pairs with exactly  5 entries: 921\n  ImageId ClassID pairs with more than  5 entries: 650\n  ImageId ClassID pairs with more than 10 entries: 125\n  ImageId ClassID pairs with more than 15 entries: 62\n  ImageId ClassID pairs with more than 20 entries: 31\n  ImageId ClassID pairs with more than 25 entries: 18\n  ImageId ClassID pairs with more than 30 entries: 11\n  ImageId ClassID pairs with more than 35 entries: 10\n  ImageId ClassID pairs with more than 40 entries: 6\n  ImageId ClassID pairs with more than 45 entries: 4\n  ImageId ClassID pairs with more than 50 entries: 2\n  ImageId ClassID pairs with more than 55 entries: 1</p>\n</blockquote>\n\n<p>I have seen examples elsewhere in which a precision at each threshold is calculated from the sum of all TP, FP and FN for all images and class ids.\nI have also seen examples of where precision is calculated for each image by the sum of TP, FP, FN for all class ids for that image, then final precision is average of all images\nI have also seen it calculated as the precision for each class by summing TP, FP, FN for all images, then final precision the average of all classes. \nI am really not sure which one is correct here - maybe the evaluation page does mean exactly what it says?</p>\n\n<p>And last - but certainly not least - I am not sure how to calculate IoU in cases when there is more than once ClassId in an image. How do I know which gt mask to compare my predicted mask to? Surely I can't calculate compared to all, since a perfect model would then get false positives. Do I compare the predicted mask to all possible gt masks and just take the maximum (since predicted masks cannot overlap)?</p>\n\n<p>I know I have just asked a lot of questions, but any help would be much appreciated. </p>\n\n<p>Cheers,\nChris</p>",
  "messages": [
    {
      "id": "542080",
      "postDate": "06/03/2019 12:32:51",
      "content": "<p>Hi All,</p>\n\n<p>I am trying to calculate the score of my validation set however I am not entirely sure how to calculate the mAP.</p>\n\n<p>To start from the beginning, I am not even sure that my definition of how to calculate the positives and negatives are correct:</p>\n\n<h3>TP:</h3>\n\n<ul>\n<li>IoU of predicted mask to gt mask is greater than threshold</li>\n</ul>\n\n<h3>FP:</h3>\n\n<ul>\n<li>Evaluation page clearly states this is when a predicted class has no associated ground truth. Ie predict class A but the image does not have class A in it.</li>\n<li>But, is it also when IoU of predicted mask to gt mask is LESS than threshold?</li>\n</ul>\n\n<h3>FN:</h3>\n\n<ul>\n<li>Again evaluation page is pretty clear that a ground truth object has no predicted object. Ie image has A in it but we don't even predict A.</li>\n<li>But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?</li>\n</ul>\n\n<p>Even if I am clear on how to classify positives and negatives, I am still unclear on how to calculate precision from them. On the evaluation page,  it seems to suggest for each ClassId and ImageId pair, we calculate the precision according to the given equation. However this does not seem right to me, since in most cases there is only one or two of each ClassId for an ImageId. That doesn't seem like enough to calculate a precision from. In most cases, that fraction is going to be 0 or 1.</p>\n\n<blockquote>\n  <p>ImageId ClassID pairs with exactly  1 entries: 173160\n  ImageId ClassID pairs with exactly  2 entries: 62687\n  ImageId ClassID pairs with exactly  3 entries: 4242\n  ImageId ClassID pairs with exactly  4 entries: 2882\n  ImageId ClassID pairs with exactly  5 entries: 921\n  ImageId ClassID pairs with more than  5 entries: 650\n  ImageId ClassID pairs with more than 10 entries: 125\n  ImageId ClassID pairs with more than 15 entries: 62\n  ImageId ClassID pairs with more than 20 entries: 31\n  ImageId ClassID pairs with more than 25 entries: 18\n  ImageId ClassID pairs with more than 30 entries: 11\n  ImageId ClassID pairs with more than 35 entries: 10\n  ImageId ClassID pairs with more than 40 entries: 6\n  ImageId ClassID pairs with more than 45 entries: 4\n  ImageId ClassID pairs with more than 50 entries: 2\n  ImageId ClassID pairs with more than 55 entries: 1</p>\n</blockquote>\n\n<p>I have seen examples elsewhere in which a precision at each threshold is calculated from the sum of all TP, FP and FN for all images and class ids.\nI have also seen examples of where precision is calculated for each image by the sum of TP, FP, FN for all class ids for that image, then final precision is average of all images\nI have also seen it calculated as the precision for each class by summing TP, FP, FN for all images, then final precision the average of all classes. \nI am really not sure which one is correct here - maybe the evaluation page does mean exactly what it says?</p>\n\n<p>And last - but certainly not least - I am not sure how to calculate IoU in cases when there is more than once ClassId in an image. How do I know which gt mask to compare my predicted mask to? Surely I can't calculate compared to all, since a perfect model would then get false positives. Do I compare the predicted mask to all possible gt masks and just take the maximum (since predicted masks cannot overlap)?</p>\n\n<p>I know I have just asked a lot of questions, but any help would be much appreciated. </p>\n\n<p>Cheers,\nChris</p>",
      "rawMarkdown": "Hi All,\n\nI am trying to calculate the score of my validation set however I am not entirely sure how to calculate the mAP.\n\nTo start from the beginning, I am not even sure that my definition of how to calculate the positives and negatives are correct:\n### TP: \n- IoU of predicted mask to gt mask is greater than threshold\n\n### FP: \n- Evaluation page clearly states this is when a predicted class has no associated ground truth. Ie predict class A but the image does not have class A in it.\n- But, is it also when IoU of predicted mask to gt mask is LESS than threshold?\n\n### FN:\n- Again evaluation page is pretty clear that a ground truth object has no predicted object. Ie image has A in it but we don't even predict A.\n- But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?\n\nEven if I am clear on how to classify positives and negatives, I am still unclear on how to calculate precision from them. On the evaluation page,  it seems to suggest for each ClassId and ImageId pair, we calculate the precision according to the given equation. However this does not seem right to me, since in most cases there is only one or two of each ClassId for an ImageId. That doesn't seem like enough to calculate a precision from. In most cases, that fraction is going to be 0 or 1.\n\n&gt; ImageId ClassID pairs with exactly  1 entries: 173160\nImageId ClassID pairs with exactly  2 entries: 62687\nImageId ClassID pairs with exactly  3 entries: 4242\nImageId ClassID pairs with exactly  4 entries: 2882\nImageId ClassID pairs with exactly  5 entries: 921\nImageId ClassID pairs with more than  5 entries: 650\nImageId ClassID pairs with more than 10 entries: 125\nImageId ClassID pairs with more than 15 entries: 62\nImageId ClassID pairs with more than 20 entries: 31\nImageId ClassID pairs with more than 25 entries: 18\nImageId ClassID pairs with more than 30 entries: 11\nImageId ClassID pairs with more than 35 entries: 10\nImageId ClassID pairs with more than 40 entries: 6\nImageId ClassID pairs with more than 45 entries: 4\nImageId ClassID pairs with more than 50 entries: 2\nImageId ClassID pairs with more than 55 entries: 1\n\nI have seen examples elsewhere in which a precision at each threshold is calculated from the sum of all TP, FP and FN for all images and class ids.\nI have also seen examples of where precision is calculated for each image by the sum of TP, FP, FN for all class ids for that image, then final precision is average of all images\nI have also seen it calculated as the precision for each class by summing TP, FP, FN for all images, then final precision the average of all classes. \nI am really not sure which one is correct here - maybe the evaluation page does mean exactly what it says?\n\nAnd last - but certainly not least - I am not sure how to calculate IoU in cases when there is more than once ClassId in an image. How do I know which gt mask to compare my predicted mask to? Surely I can't calculate compared to all, since a perfect model would then get false positives. Do I compare the predicted mask to all possible gt masks and just take the maximum (since predicted masks cannot overlap)?\n\nI know I have just asked a lot of questions, but any help would be much appreciated. \n\nCheers,\nChris",
      "votes": null
    },
    {
      "id": "542108",
      "postDate": "06/03/2019 13:20:38",
      "content": "<p>The evaluation page says</p>\n\n<blockquote>\n  <p>The average precision of a single ClassId and a single image is then calculated as the mean of the above precision values at each IoU threshold</p>\n</blockquote>\n\n<p>So, my take is that\n- first calculate precision of every instance\n- then take average for each ClassId. Example, if there are 2 shoes, then get a mean of that\n- then get a mean of all ClassIds in that image\n- then do averaging of all images</p>",
      "rawMarkdown": "The evaluation page says\n&gt; The average precision of a single ClassId and a single image is then calculated as the mean of the above precision values at each IoU threshold\n\nSo, my take is that\n- first calculate precision of every instance\n- then take average for each ClassId. Example, if there are 2 shoes, then get a mean of that\n- then get a mean of all ClassIds in that image\n- then do averaging of all images",
      "votes": null
    },
    {
      "id": "542315",
      "postDate": "06/03/2019 17:40:22",
      "content": "<blockquote>\n  <p>But, is it also when IoU of predicted mask to gt mask is LESS than threshold?</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>Do I compare the predicted mask to all possible gt masks</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>and just take the maximum (since predicted masks cannot overlap)?</p>\n</blockquote>\n\n<p>This is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.</p>\n\n<p>Regarding averaging, you need to take one big average overall all class ids, all images, all iou thresholds, and exclude items where f1 is not defined (this means that for each image you loop over classes that are predicted or true).</p>",
      "rawMarkdown": "&gt; But, is it also when IoU of predicted mask to gt mask is LESS than threshold?\n\nyes\n\n&gt; But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?\n\nyes\n\n&gt; Do I compare the predicted mask to all possible gt masks\n\nyes\n\n&gt; and just take the maximum (since predicted masks cannot overlap)?\n\nThis is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.\n\nRegarding averaging, you need to take one big average overall all class ids, all images, all iou thresholds, and exclude items where f1 is not defined (this means that for each image you loop over classes that are predicted or true).",
      "votes": null
    },
    {
      "id": "542329",
      "postDate": "06/03/2019 17:56:52",
      "content": "<p>Also it may help that COCO and most other detection/instance segmentation competitions use the same metric with only minor differences, so for example if you understand how COCO is evaluated, you'll also understand how this competition is evaluated. I'm not using COCO code myself but it's descriptions helped me.</p>",
      "rawMarkdown": "Also it may help that COCO and most other detection/instance segmentation competitions use the same metric with only minor differences, so for example if you understand how COCO is evaluated, you'll also understand how this competition is evaluated. I'm not using COCO code myself but it's descriptions helped me.",
      "votes": null
    },
    {
      "id": "543014",
      "postDate": "06/04/2019 09:16:35",
      "content": "<p>Thanks that makes things clearer,</p>\n\n<p>But what do you mean by:\n&gt; This is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.</p>\n\n<p>You compare the prediction to all ground truths, for each of those if IoU is greater than threshold then TP, less than threshold FP and intersection is 0 then FN? Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN? But in practive, there aren't that many of those therefore it is not important?</p>\n\n<p>I am still not sure what you mean for averaging. Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?\nOr you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)</p>",
      "rawMarkdown": "Thanks that makes things clearer,\n\nBut what do you mean by:\n&gt; This is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.\n\nYou compare the prediction to all ground truths, for each of those if IoU is greater than threshold then TP, less than threshold FP and intersection is 0 then FN? Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN? But in practive, there aren't that many of those therefore it is not important?\n\nI am still not sure what you mean for averaging. Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?\nOr you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)",
      "votes": null
    },
    {
      "id": "543020",
      "postDate": "06/04/2019 09:23:31",
      "content": "<blockquote>\n  <p>Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>Or you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)</p>\n</blockquote>\n\n<p>no, as far as I understand we need to calculate F1 first and then average</p>\n\n<blockquote>\n  <p>Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN?</p>\n</blockquote>\n\n<p>I was thinking of a case where we have an instance which has IoU &gt; 0.5 with several GT instances (actually now I wrote it and I'm not sure if it's even possible, since GT masks for one class can't overlap). To be honest at the moment I won't be able to formulate that exactly :) This is more like a place in my code I'm not 100% sure about. </p>",
      "rawMarkdown": "&gt; Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?\n\nyes\n\n&gt; Or you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)\n\nno, as far as I understand we need to calculate F1 first and then average\n\n&gt; Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN?\n\nI was thinking of a case where we have an instance which has IoU &gt; 0.5 with several GT instances (actually now I wrote it and I'm not sure if it's even possible, since GT masks for one class can't overlap). To be honest at the moment I won't be able to formulate that exactly :) This is more like a place in my code I'm not 100% sure about.",
      "votes": null
    },
    {
      "id": "543274",
      "postDate": "06/04/2019 12:34:37",
      "content": "<p>Great thanks,  I should be able to put together a metric that will at least give me an indication of performance now :)</p>",
      "rawMarkdown": "Great thanks,  I should be able to put together a metric that will at least give me an indication of performance now :)",
      "votes": null
    },
    {
      "id": "546630",
      "postDate": "06/06/2019 18:52:46",
      "content": "<p>I've tried implementing the evaluation codes. If you'd like, check it and give me feedbacks;)\n<a href=\"https://www.kaggle.com/kyazuki/calculate-evaluation-score\">https://www.kaggle.com/kyazuki/calculate-evaluation-score</a></p>",
      "rawMarkdown": "I've tried implementing the evaluation codes. If you'd like, check it and give me feedbacks;)\nhttps://www.kaggle.com/kyazuki/calculate-evaluation-score",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542108,
      "author_name": "soonyau",
      "author_url": "",
      "post_date": "06/03/2019 13:20:38",
      "content": "<p>The evaluation page says</p>\n\n<blockquote>\n  <p>The average precision of a single ClassId and a single image is then calculated as the mean of the above precision values at each IoU threshold</p>\n</blockquote>\n\n<p>So, my take is that\n- first calculate precision of every instance\n- then take average for each ClassId. Example, if there are 2 shoes, then get a mean of that\n- then get a mean of all ClassIds in that image\n- then do averaging of all images</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542315,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "06/03/2019 17:40:22",
      "content": "<blockquote>\n  <p>But, is it also when IoU of predicted mask to gt mask is LESS than threshold?</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>Do I compare the predicted mask to all possible gt masks</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>and just take the maximum (since predicted masks cannot overlap)?</p>\n</blockquote>\n\n<p>This is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.</p>\n\n<p>Regarding averaging, you need to take one big average overall all class ids, all images, all iou thresholds, and exclude items where f1 is not defined (this means that for each image you loop over classes that are predicted or true).</p>",
      "votes": null,
      "replies": [
        {
          "id": 542329,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/03/2019 17:56:52",
          "content": "<p>Also it may help that COCO and most other detection/instance segmentation competitions use the same metric with only minor differences, so for example if you understand how COCO is evaluated, you'll also understand how this competition is evaluated. I'm not using COCO code myself but it's descriptions helped me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543014,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "06/04/2019 09:16:35",
          "content": "<p>Thanks that makes things clearer,</p>\n\n<p>But what do you mean by:\n&gt; This is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.</p>\n\n<p>You compare the prediction to all ground truths, for each of those if IoU is greater than threshold then TP, less than threshold FP and intersection is 0 then FN? Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN? But in practive, there aren't that many of those therefore it is not important?</p>\n\n<p>I am still not sure what you mean for averaging. Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?\nOr you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543020,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/04/2019 09:23:31",
          "content": "<blockquote>\n  <p>Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?</p>\n</blockquote>\n\n<p>yes</p>\n\n<blockquote>\n  <p>Or you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)</p>\n</blockquote>\n\n<p>no, as far as I understand we need to calculate F1 first and then average</p>\n\n<blockquote>\n  <p>Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN?</p>\n</blockquote>\n\n<p>I was thinking of a case where we have an instance which has IoU &gt; 0.5 with several GT instances (actually now I wrote it and I'm not sure if it's even possible, since GT masks for one class can't overlap). To be honest at the moment I won't be able to formulate that exactly :) This is more like a place in my code I'm not 100% sure about. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543274,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "06/04/2019 12:34:37",
          "content": "<p>Great thanks,  I should be able to put together a metric that will at least give me an indication of performance now :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546630,
      "author_name": "kyazuki",
      "author_url": "",
      "post_date": "06/06/2019 18:52:46",
      "content": "<p>I've tried implementing the evaluation codes. If you'd like, check it and give me feedbacks;)\n<a href=\"https://www.kaggle.com/kyazuki/calculate-evaluation-score\">https://www.kaggle.com/kyazuki/calculate-evaluation-score</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542080": "Hi All,\n\nI am trying to calculate the score of my validation set however I am not entirely sure how to calculate the mAP.\n\nTo start from the beginning, I am not even sure that my definition of how to calculate the positives and negatives are correct:\n### TP: \n- IoU of predicted mask to gt mask is greater than threshold\n\n### FP: \n- Evaluation page clearly states this is when a predicted class has no associated ground truth. Ie predict class A but the image does not have class A in it.\n- But, is it also when IoU of predicted mask to gt mask is LESS than threshold?\n\n### FN:\n- Again evaluation page is pretty clear that a ground truth object has no predicted object. Ie image has A in it but we don't even predict A.\n- But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?\n\nEven if I am clear on how to classify positives and negatives, I am still unclear on how to calculate precision from them. On the evaluation page,  it seems to suggest for each ClassId and ImageId pair, we calculate the precision according to the given equation. However this does not seem right to me, since in most cases there is only one or two of each ClassId for an ImageId. That doesn't seem like enough to calculate a precision from. In most cases, that fraction is going to be 0 or 1.\n\n&gt; ImageId ClassID pairs with exactly  1 entries: 173160\nImageId ClassID pairs with exactly  2 entries: 62687\nImageId ClassID pairs with exactly  3 entries: 4242\nImageId ClassID pairs with exactly  4 entries: 2882\nImageId ClassID pairs with exactly  5 entries: 921\nImageId ClassID pairs with more than  5 entries: 650\nImageId ClassID pairs with more than 10 entries: 125\nImageId ClassID pairs with more than 15 entries: 62\nImageId ClassID pairs with more than 20 entries: 31\nImageId ClassID pairs with more than 25 entries: 18\nImageId ClassID pairs with more than 30 entries: 11\nImageId ClassID pairs with more than 35 entries: 10\nImageId ClassID pairs with more than 40 entries: 6\nImageId ClassID pairs with more than 45 entries: 4\nImageId ClassID pairs with more than 50 entries: 2\nImageId ClassID pairs with more than 55 entries: 1\n\nI have seen examples elsewhere in which a precision at each threshold is calculated from the sum of all TP, FP and FN for all images and class ids.\nI have also seen examples of where precision is calculated for each image by the sum of TP, FP, FN for all class ids for that image, then final precision is average of all images\nI have also seen it calculated as the precision for each class by summing TP, FP, FN for all images, then final precision the average of all classes. \nI am really not sure which one is correct here - maybe the evaluation page does mean exactly what it says?\n\nAnd last - but certainly not least - I am not sure how to calculate IoU in cases when there is more than once ClassId in an image. How do I know which gt mask to compare my predicted mask to? Surely I can't calculate compared to all, since a perfect model would then get false positives. Do I compare the predicted mask to all possible gt masks and just take the maximum (since predicted masks cannot overlap)?\n\nI know I have just asked a lot of questions, but any help would be much appreciated. \n\nCheers,\nChris",
    "542108": "The evaluation page says\n&gt; The average precision of a single ClassId and a single image is then calculated as the mean of the above precision values at each IoU threshold\n\nSo, my take is that\n- first calculate precision of every instance\n- then take average for each ClassId. Example, if there are 2 shoes, then get a mean of that\n- then get a mean of all ClassIds in that image\n- then do averaging of all images",
    "542315": "&gt; But, is it also when IoU of predicted mask to gt mask is LESS than threshold?\n\nyes\n\n&gt; But, is this also when intersection of masks is zero? ie we predict a class that is in the image, however it is not associated at all?\n\nyes\n\n&gt; Do I compare the predicted mask to all possible gt masks\n\nyes\n\n&gt; and just take the maximum (since predicted masks cannot overlap)?\n\nThis is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.\n\nRegarding averaging, you need to take one big average overall all class ids, all images, all iou thresholds, and exclude items where f1 is not defined (this means that for each image you loop over classes that are predicted or true).",
    "542329": "Also it may help that COCO and most other detection/instance segmentation competitions use the same metric with only minor differences, so for example if you understand how COCO is evaluated, you'll also understand how this competition is evaluated. I'm not using COCO code myself but it's descriptions helped me.",
    "543014": "Thanks that makes things clearer,\n\nBut what do you mean by:\n&gt; This is a corner-case I'm not sure of myself (I take all matches), but in practice this seems not important.\n\nYou compare the prediction to all ground truths, for each of those if IoU is greater than threshold then TP, less than threshold FP and intersection is 0 then FN? Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN? But in practive, there aren't that many of those therefore it is not important?\n\nI am still not sure what you mean for averaging. Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?\nOr you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)",
    "543020": "&gt; Do you mean for each class in each image you calculate one average precision according to the evaluation page, then take a big average of all of those precisions?\n\nyes\n\n&gt; Or you sum up all TP, FP, TN for all image and class ids for a given IoU threshold, then take one big weighted average (ie Calculate 10 different average precisions, weight them according to threshold and average)\n\nno, as far as I understand we need to calculate F1 first and then average\n\n&gt; Would that mean that in the case where there are 50 instances of a class in an image, and you have a perfect predictor that predicts exactly one instance, then you have one TP and 49 FN?\n\nI was thinking of a case where we have an instance which has IoU &gt; 0.5 with several GT instances (actually now I wrote it and I'm not sure if it's even possible, since GT masks for one class can't overlap). To be honest at the moment I won't be able to formulate that exactly :) This is more like a place in my code I'm not 100% sure about.",
    "543274": "Great thanks,  I should be able to put together a metric that will at least give me an indication of performance now :)",
    "546630": "I've tried implementing the evaluation codes. If you'd like, check it and give me feedbacks;)\nhttps://www.kaggle.com/kyazuki/calculate-evaluation-score"
  },
  "source": "meta"
}