{
  "id": 241275,
  "title": "How Competition Metric is Calculated Internally?",
  "url": "/competitions/siim-covid19-detection/discussion/241275",
  "author_name": "Kerem Turgutlu",
  "post_date": "2021-05-23T21:39:01.292000",
  "votes": 31,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I wonder how the competition metric is calculated for a submission with both study and image level predictions using PASCAL-VOC map@0.5 metric. </p>\n<p>I would like to replicate this process locally for validating my models but I am not sure if I understand the calculation 100%. I assume a tool similar to pycocotools is used for the calculations.</p>\n<p><strong>Image Level</strong></p>\n<p><strong>Do we consider each image level as a binary classification e.g. opacity vs none ?</strong></p>\n<p>Normally AP is calculated per class and then an average among different classes is taken for the final mAP. In this case for image-level predictions we are expected to predicted either with a class label of <strong>opacity</strong> or <strong>none</strong>.  Does it mean there are only 2 classes (or 1 class depending on how you look at it) for image-level predictions?</p>\n<p><strong>Study Level</strong></p>\n<p><strong>How to correctly calculate study level mAP?</strong></p>\n<p>We know that study level predictions are converted into a bounding box target of 0 0 1 1 to make AP calculation possible similar to other past competitions such <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection\" target=\"_blank\">vinbigdata</a>.  For each study we can make the following predictions:  <code>['negative', 'typical', 'indeterminate', 'atypical']</code>. This makes an additional of 4 classes on top of what we already have from image level classes (opacity and none).</p>\n<p>Taking all of these into account does it mean mAP is calculated by taking an average over the classes <code>none, opacity, negative, typical, indeterminate, atypical</code>. If a tool like pycocotools is used then I guess this would be the case. </p>\n<p>I am currently trying to follow this great kernel <a href=\"https://www.kaggle.com/its7171/map-understanding-with-code-and-its-tips\" target=\"_blank\">https://www.kaggle.com/its7171/map-understanding-with-code-and-its-tips</a> for my local validation. I tried using <a href=\"https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\" target=\"_blank\">https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes</a> on a sample study level predictions however mAP comes out 0:</p>\n<pre><code>from map_boxes import mean_average_precision_for_boxes\n#ann[:5], det[:5]\n(array([['059fb5105884_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0]],\n       dtype=object),\n array([['059fb5105884_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0]],\n       dtype=object))\n\nmean_ap, average_precisions = mean_average_precision_for_boxes(ann[:5], det[:5])\n\nOutput:\nNumber of files in annotations: 5\nNumber of files in predictions: 5\nUnique classes: 2\nDetections length: 5\nAnnotations length: 5\nindeterminate                  | 0.000000 |       3\nnegative                       | 0.000000 |       2\nmAP: 0.000000\n</code></pre>",
  "messages": [
    {
      "id": 1320210,
      "postDate": "2021-05-23T21:39:01.293Z",
      "content": "<p>I wonder how the competition metric is calculated for a submission with both study and image level predictions using PASCAL-VOC map@0.5 metric. </p>\n<p>I would like to replicate this process locally for validating my models but I am not sure if I understand the calculation 100%. I assume a tool similar to pycocotools is used for the calculations.</p>\n<p><strong>Image Level</strong></p>\n<p><strong>Do we consider each image level as a binary classification e.g. opacity vs none ?</strong></p>\n<p>Normally AP is calculated per class and then an average among different classes is taken for the final mAP. In this case for image-level predictions we are expected to predicted either with a class label of <strong>opacity</strong> or <strong>none</strong>.  Does it mean there are only 2 classes (or 1 class depending on how you look at it) for image-level predictions?</p>\n<p><strong>Study Level</strong></p>\n<p><strong>How to correctly calculate study level mAP?</strong></p>\n<p>We know that study level predictions are converted into a bounding box target of 0 0 1 1 to make AP calculation possible similar to other past competitions such <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection\" target=\"_blank\">vinbigdata</a>.  For each study we can make the following predictions:  <code>['negative', 'typical', 'indeterminate', 'atypical']</code>. This makes an additional of 4 classes on top of what we already have from image level classes (opacity and none).</p>\n<p>Taking all of these into account does it mean mAP is calculated by taking an average over the classes <code>none, opacity, negative, typical, indeterminate, atypical</code>. If a tool like pycocotools is used then I guess this would be the case. </p>\n<p>I am currently trying to follow this great kernel <a href=\"https://www.kaggle.com/its7171/map-understanding-with-code-and-its-tips\" target=\"_blank\">https://www.kaggle.com/its7171/map-understanding-with-code-and-its-tips</a> for my local validation. I tried using <a href=\"https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes\" target=\"_blank\">https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes</a> on a sample study level predictions however mAP comes out 0:</p>\n<pre><code>from map_boxes import mean_average_precision_for_boxes\n#ann[:5], det[:5]\n(array([['059fb5105884_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0]],\n       dtype=object),\n array([['059fb5105884_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0]],\n       dtype=object))\n\nmean_ap, average_precisions = mean_average_precision_for_boxes(ann[:5], det[:5])\n\nOutput:\nNumber of files in annotations: 5\nNumber of files in predictions: 5\nUnique classes: 2\nDetections length: 5\nAnnotations length: 5\nindeterminate                  | 0.000000 |       3\nnegative                       | 0.000000 |       2\nmAP: 0.000000\n</code></pre>",
      "rawMarkdown": "I wonder how the competition metric is calculated for a submission with both study and image level predictions using PASCAL-VOC map@0.5 metric. \n\nI would like to replicate this process locally for validating my models but I am not sure if I understand the calculation 100%. I assume a tool similar to pycocotools is used for the calculations.\n\n**Image Level**\n\n**Do we consider each image level as a binary classification e.g. opacity vs none ?**\n\nNormally AP is calculated per class and then an average among different classes is taken for the final mAP. In this case for image-level predictions we are expected to predicted either with a class label of **opacity** or **none**.  Does it mean there are only 2 classes (or 1 class depending on how you look at it) for image-level predictions?\n\n**Study Level**\n\n**How to correctly calculate study level mAP?**\n\nWe know that study level predictions are converted into a bounding box target of 0 0 1 1 to make AP calculation possible similar to other past competitions such [vinbigdata](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection).  For each study we can make the following predictions:  `['negative', 'typical', 'indeterminate', 'atypical']`. This makes an additional of 4 classes on top of what we already have from image level classes (opacity and none).\n\n\nTaking all of these into account does it mean mAP is calculated by taking an average over the classes `none, opacity, negative, typical, indeterminate, atypical`. If a tool like pycocotools is used then I guess this would be the case. \n\n\nI am currently trying to follow this great kernel https://www.kaggle.com/its7171/map-understanding-with-code-and-its-tips for my local validation. I tried using https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes on a sample study level predictions however mAP comes out 0:\n\n```python\nfrom map_boxes import mean_average_precision_for_boxes\n#ann[:5], det[:5]\n(array([['059fb5105884_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0]],\n       dtype=object),\n array([['059fb5105884_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0]],\n       dtype=object))\n\nmean_ap, average_precisions = mean_average_precision_for_boxes(ann[:5], det[:5])\n\nOutput:\nNumber of files in annotations: 5\nNumber of files in predictions: 5\nUnique classes: 2\nDetections length: 5\nAnnotations length: 5\nindeterminate                  | 0.000000 |       3\nnegative                       | 0.000000 |       2\nmAP: 0.000000\n```\n\n\n\n\n\n\n",
      "votes": 31
    },
    {
      "id": 1330995,
      "postDate": "2021-06-01T08:02:59.823Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/novice03\" target=\"_blank\">@novice03</a>  - we are still patiently waiting. It is rather strange that we received no explanation regarding the metric calculation.<br>\nYour mentioned notebook predicted first study level labels and at the end (you can see that by downloading the file and scrolling to the end of it) image level predictions were made. Althought image level predictions were all \"none 1 0 0 1 1\"</p>",
      "rawMarkdown": "Hi @novice03  - we are still patiently waiting. It is rather strange that we received no explanation regarding the metric calculation.\nYour mentioned notebook predicted first study level labels and at the end (you can see that by downloading the file and scrolling to the end of it) image level predictions were made. Althought image level predictions were all \"none 1 0 0 1 1\"",
      "votes": 1,
      "replies": [
        {
          "id": 1331240,
          "postDate": "2021-06-01T11:03:21.040Z",
          "content": "<p>Oh, I see. Thank you for your reply.</p>",
          "rawMarkdown": "Oh, I see. Thank you for your reply.\n"
        }
      ]
    },
    {
      "id": 1351792,
      "postDate": "2021-06-16T15:56:47.947Z",
      "content": "<p>The evaluation metric is as described on the <a href=\"https://www.kaggle.com/c/siim-covid19-detection/overview/evaluation\" target=\"_blank\">Evaluation Page</a>. The best resource will be the link provided there to the documentation on PASCAL VOC 2010/12. We won't be providing any further code examples regarding this, but others are welcome to develop and share what they've created.</p>",
      "rawMarkdown": "The evaluation metric is as described on the [Evaluation Page](https://www.kaggle.com/c/siim-covid19-detection/overview/evaluation). The best resource will be the link provided there to the documentation on PASCAL VOC 2010/12. We won't be providing any further code examples regarding this, but others are welcome to develop and share what they've created."
    },
    {
      "id": 1381879,
      "postDate": "2021-07-09T09:59:47.657Z",
      "content": "<p>Hi, the evaluation metrics is calculating mAP at IoU &gt; 0.5, instead of loU = 0.5.</p>",
      "rawMarkdown": "Hi, the evaluation metrics is calculating mAP at IoU > 0.5, instead of loU = 0.5."
    },
    {
      "id": 1351771,
      "postDate": "2021-06-16T15:31:20.807Z",
      "content": "<p>Is it clear how the metric is calculated?<br>\nIs there some good code example with correct evaluation metric?</p>",
      "rawMarkdown": "Is it clear how the metric is calculated?\nIs there some good code example with correct evaluation metric?"
    },
    {
      "id": 1329962,
      "postDate": "2021-05-31T14:01:45.800Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a>, did you get an answer for this? I just joined the competition, and I can't find anything that answers this question. I don't understand how this (<a href=\"https://www.kaggle.com/h053473666/siim-cov19-efnb7-infer-study\" target=\"_blank\">https://www.kaggle.com/h053473666/siim-cov19-efnb7-infer-study</a>) notebook did inference for study level only and still got a score. </p>",
      "rawMarkdown": "Hi @keremt, did you get an answer for this? I just joined the competition, and I can't find anything that answers this question. I don't understand how this (https://www.kaggle.com/h053473666/siim-cov19-efnb7-infer-study) notebook did inference for study level only and still got a score. "
    },
    {
      "id": 1373346,
      "postDate": "2021-07-02T12:22:56.823Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a> :) The issue is that you pass the string of <strong>|ImgId, Label, Score (optional), Xmin, Ymin, Xmax, Ymax|</strong> to <em>mean_average_precision_for_boxes</em>, but it requires **|ImgId, Label, Score (optional), Xmin, Xmax, Ymin, Ymax| ** format. It works fine.</p>",
      "rawMarkdown": "Hey @keremt :) The issue is that you pass the string of **|ImgId, Label, Score (optional), Xmin, Ymin, Xmax, Ymax|** to *mean_average_precision_for_boxes*, but it requires **|ImgId, Label, Score (optional), Xmin, Xmax, Ymin, Ymax| ** format. It works fine.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1330995,
      "author_name": "kulverstukas",
      "author_url": "",
      "post_date": "2021-06-01T08:02:59.823000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/novice03\" target=\"_blank\">@novice03</a>  - we are still patiently waiting. It is rather strange that we received no explanation regarding the metric calculation.<br>\nYour mentioned notebook predicted first study level labels and at the end (you can see that by downloading the file and scrolling to the end of it) image level predictions were made. Althought image level predictions were all \"none 1 0 0 1 1\"</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1331240,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "2021-06-01T11:03:21.040000",
          "content": "<p>Oh, I see. Thank you for your reply.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1351792,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2021-06-16T15:56:47.947000",
      "content": "<p>The evaluation metric is as described on the <a href=\"https://www.kaggle.com/c/siim-covid19-detection/overview/evaluation\" target=\"_blank\">Evaluation Page</a>. The best resource will be the link provided there to the documentation on PASCAL VOC 2010/12. We won't be providing any further code examples regarding this, but others are welcome to develop and share what they've created.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1381879,
      "author_name": "Yue Sun",
      "author_url": "",
      "post_date": "2021-07-09T09:59:47.657000",
      "content": "<p>Hi, the evaluation metrics is calculating mAP at IoU &gt; 0.5, instead of loU = 0.5.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1351771,
      "author_name": "Aleksandr Emchinov",
      "author_url": "",
      "post_date": "2021-06-16T15:31:20.807000",
      "content": "<p>Is it clear how the metric is calculated?<br>\nIs there some good code example with correct evaluation metric?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1329962,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "2021-05-31T14:01:45.800000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a>, did you get an answer for this? I just joined the competition, and I can't find anything that answers this question. I don't understand how this (<a href=\"https://www.kaggle.com/h053473666/siim-cov19-efnb7-infer-study\" target=\"_blank\">https://www.kaggle.com/h053473666/siim-cov19-efnb7-infer-study</a>) notebook did inference for study level only and still got a score. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1373346,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-02T12:22:56.823000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/keremt\" target=\"_blank\">@keremt</a> :) The issue is that you pass the string of <strong>|ImgId, Label, Score (optional), Xmin, Ymin, Xmax, Ymax|</strong> to <em>mean_average_precision_for_boxes</em>, but it requires **|ImgId, Label, Score (optional), Xmin, Xmax, Ymin, Ymax| ** format. It works fine.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1320210": "I wonder how the competition metric is calculated for a submission with both study and image level predictions using PASCAL-VOC map@0.5 metric. \n\nI would like to replicate this process locally for validating my models but I am not sure if I understand the calculation 100%. I assume a tool similar to pycocotools is used for the calculations.\n\n**Image Level**\n\n**Do we consider each image level as a binary classification e.g. opacity vs none ?**\n\nNormally AP is calculated per class and then an average among different classes is taken for the final mAP. In this case for image-level predictions we are expected to predicted either with a class label of **opacity** or **none**.  Does it mean there are only 2 classes (or 1 class depending on how you look at it) for image-level predictions?\n\n**Study Level**\n\n**How to correctly calculate study level mAP?**\n\nWe know that study level predictions are converted into a bounding box target of 0 0 1 1 to make AP calculation possible similar to other past competitions such [vinbigdata](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection).  For each study we can make the following predictions:  `['negative', 'typical', 'indeterminate', 'atypical']`. This makes an additional of 4 classes on top of what we already have from image level classes (opacity and none).\n\n\nTaking all of these into account does it mean mAP is calculated by taking an average over the classes `none, opacity, negative, typical, indeterminate, atypical`. If a tool like pycocotools is used then I guess this would be the case. \n\n\nI am currently trying to follow this great kernel https://www.kaggle.com/its7171/map-understanding-with-code-and-its-tips for my local validation. I tried using https://github.com/ZFTurbo/Mean-Average-Precision-for-Boxes on a sample study level predictions however mAP comes out 0:\n\n```python\nfrom map_boxes import mean_average_precision_for_boxes\n#ann[:5], det[:5]\n(array([['059fb5105884_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'indeterminate', 0.0, 0.0, 1.0, 1.0]],\n       dtype=object),\n array([['059fb5105884_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['066b12d875eb_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['069d6c49a1d1_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['06e5fcfcd2f8_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0],\n        ['070cbb72137d_study', 'negative', 1.0, 0.0, 0.0, 1.0, 1.0]],\n       dtype=object))\n\nmean_ap, average_precisions = mean_average_precision_for_boxes(ann[:5], det[:5])\n\nOutput:\nNumber of files in annotations: 5\nNumber of files in predictions: 5\nUnique classes: 2\nDetections length: 5\nAnnotations length: 5\nindeterminate                  | 0.000000 |       3\nnegative                       | 0.000000 |       2\nmAP: 0.000000\n```\n\n\n\n\n\n\n",
    "1330995": "Hi @novice03  - we are still patiently waiting. It is rather strange that we received no explanation regarding the metric calculation.\nYour mentioned notebook predicted first study level labels and at the end (you can see that by downloading the file and scrolling to the end of it) image level predictions were made. Althought image level predictions were all \"none 1 0 0 1 1\"",
    "1351792": "The evaluation metric is as described on the [Evaluation Page](https://www.kaggle.com/c/siim-covid19-detection/overview/evaluation). The best resource will be the link provided there to the documentation on PASCAL VOC 2010/12. We won't be providing any further code examples regarding this, but others are welcome to develop and share what they've created.",
    "1381879": "Hi, the evaluation metrics is calculating mAP at IoU > 0.5, instead of loU = 0.5.",
    "1351771": "Is it clear how the metric is calculated?\nIs there some good code example with correct evaluation metric?",
    "1329962": "Hi @keremt, did you get an answer for this? I just joined the competition, and I can't find anything that answers this question. I don't understand how this (https://www.kaggle.com/h053473666/siim-cov19-efnb7-infer-study) notebook did inference for study level only and still got a score. ",
    "1373346": "Hey @keremt :) The issue is that you pass the string of **|ImgId, Label, Score (optional), Xmin, Ymin, Xmax, Ymax|** to *mean_average_precision_for_boxes*, but it requires **|ImgId, Label, Score (optional), Xmin, Xmax, Ymin, Ymax| ** format. It works fine."
  }
}