{
  "id": 116661,
  "title": "leaderboard's mAP calculation wrong?",
  "url": "/competitions/pku-autonomous-driving/discussion/116661",
  "author_name": "Bo",
  "post_date": "2019-11-10T15:31:30.442000",
  "votes": 64,
  "comment_count": 39,
  "views": 0,
  "content": "<p>In coco style mAP metric, for the same model, the score should be a monotone non-increasing function of the probability threshold. In other words, the lower the threshold, the higher the score should be. That was the case in the <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection\">open images competitions</a></p>\n\n<p>But apparently this is not the case in this competition. For the same <a href=\"https://www.kaggle.com/hocop1/centernet-baseline?scriptVersionId=22630284\">model</a> (best public notebook), the score is:\n- 0.110 with thres=0.75 (i.e. only output predicted car instances with prob&gt;=0.75, which is a subset of below two submissions)\n- 0.093 with thres=0.5\n- 0.074 with thres=0.25</p>\n\n<p>Does anyone notice the same thing?\n@philculliton and @addisonhoward, is it possible to check this, or even better, post the evaluation code? Thank you.</p>",
  "messages": [
    {
      "id": 669866,
      "postDate": "2019-11-10T15:31:30.443Z",
      "content": "<p>In coco style mAP metric, for the same model, the score should be a monotone non-increasing function of the probability threshold. In other words, the lower the threshold, the higher the score should be. That was the case in the <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection\">open images competitions</a></p>\n\n<p>But apparently this is not the case in this competition. For the same <a href=\"https://www.kaggle.com/hocop1/centernet-baseline?scriptVersionId=22630284\">model</a> (best public notebook), the score is:\n- 0.110 with thres=0.75 (i.e. only output predicted car instances with prob&gt;=0.75, which is a subset of below two submissions)\n- 0.093 with thres=0.5\n- 0.074 with thres=0.25</p>\n\n<p>Does anyone notice the same thing?\n@philculliton and @addisonhoward, is it possible to check this, or even better, post the evaluation code? Thank you.</p>",
      "rawMarkdown": "In coco style mAP metric, for the same model, the score should be a monotone non-increasing function of the probability threshold. In other words, the lower the threshold, the higher the score should be. That was the case in the [open images competitions](https://www.kaggle.com/c/open-images-2019-object-detection)\n\nBut apparently this is not the case in this competition. For the same [model](https://www.kaggle.com/hocop1/centernet-baseline?scriptVersionId=22630284) (best public notebook), the score is:\n- 0.110 with thres=0.75 (i.e. only output predicted car instances with prob&gt;=0.75, which is a subset of below two submissions)\n- 0.093 with thres=0.5\n- 0.074 with thres=0.25\n\nDoes anyone notice the same thing?\n@philculliton and @addisonhoward, is it possible to check this, or even better, post the evaluation code? Thank you.",
      "votes": 63
    },
    {
      "id": 681067,
      "postDate": "2019-11-25T16:18:27.703Z",
      "content": "<p>Hi all - thanks again for your feedback. Sorry for the delay! I've been working with the host and other data scientists to make sure we didn't have bugs in the metric re: the submissions presented in this thread, and also working through next steps. We will very likely be adding FN to the calculation and rescoring this week.</p>",
      "rawMarkdown": "Hi all - thanks again for your feedback. Sorry for the delay! I've been working with the host and other data scientists to make sure we didn't have bugs in the metric re: the submissions presented in this thread, and also working through next steps. We will very likely be adding FN to the calculation and rescoring this week.",
      "votes": 27,
      "replies": [
        {
          "id": 681094,
          "postDate": "2019-11-25T16:48:42.507Z",
          "content": "<p>Thanks for the update!</p>",
          "rawMarkdown": "Thanks for the update!",
          "votes": 2
        },
        {
          "id": 681138,
          "postDate": "2019-11-25T18:04:24.570Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Thanks a lot for managing this and thanks for giving us an update on this matter. Much appreciated!</p>",
          "rawMarkdown": "@philculliton Thanks a lot for managing this and thanks for giving us an update on this matter. Much appreciated!\n",
          "votes": 2
        },
        {
          "id": 681257,
          "postDate": "2019-11-25T22:59:04.353Z",
          "content": "<p>Great to hear that! Sure will switch gears after reflected :)</p>",
          "rawMarkdown": "Great to hear that! Sure will switch gears after reflected :)"
        },
        {
          "id": 681461,
          "postDate": "2019-11-26T06:25:12.070Z",
          "content": "<p>Thank you for attention to the problem! Nevertheless, I think that just rescoring and changing the metrics is not enough. Would you mind to change the Evaluation page as well and describe the metic more thoroughly, because without understanding the work of the metric, we will not be able to properly validate our models. </p>",
          "rawMarkdown": "Thank you for attention to the problem! Nevertheless, I think that just rescoring and changing the metrics is not enough. Would you mind to change the Evaluation page as well and describe the metic more thoroughly, because without understanding the work of the metric, we will not be able to properly validate our models. ",
          "votes": 9
        },
        {
          "id": 681559,
          "postDate": "2019-11-26T08:37:56.227Z",
          "content": "<p>Yes, you are right, we need the evaluation code for the validation. </p>",
          "rawMarkdown": "Yes, you are right, we need the evaluation code for the validation. ",
          "votes": 5
        },
        {
          "id": 682042,
          "postDate": "2019-11-26T20:52:15.070Z",
          "content": "<p>I don't think you need to give us the code (though maybe pseudocode would be nice), but a very thorough description of how it works with an equation would be great.</p>\n\n<p>Assuming this change in the evaluation metric impacts the leaderboard significantly, would it be normal to ask for an extension by a few weeks for this competition? A lot of people probably spent time chasing the leaderboard results without any regard for actually having a better model, though maybe thats their fault 🤷‍♂️</p>",
          "rawMarkdown": "I don't think you need to give us the code (though maybe pseudocode would be nice), but a very thorough description of how it works with an equation would be great.\n\nAssuming this change in the evaluation metric impacts the leaderboard significantly, would it be normal to ask for an extension by a few weeks for this competition? A lot of people probably spent time chasing the leaderboard results without any regard for actually having a better model, though maybe thats their fault 🤷‍♂️",
          "votes": 2
        },
        {
          "id": 684551,
          "postDate": "2019-11-30T02:16:03.117Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi philculliton , This week is drawing to a close, will the score be recalculated within this week, and then the new score calculation method will be announced, is that right?</p>",
          "rawMarkdown": "@philculliton Hi philculliton , This week is drawing to a close, will the score be recalculated within this week, and then the new score calculation method will be announced, is that right?",
          "votes": 7
        },
        {
          "id": 686265,
          "postDate": "2019-12-03T01:28:16.603Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi Phil, could you kindly give us some update about the evaluation metric rescoring? Your plan was to update last week....thanks</p>",
          "rawMarkdown": "@philculliton Hi Phil, could you kindly give us some update about the evaluation metric rescoring? Your plan was to update last week....thanks",
          "votes": 4
        },
        {
          "id": 686351,
          "postDate": "2019-12-03T03:54:13.610Z",
          "content": "<p>Hi - thanks everybody! Sorry for the delay - we and the host needed to do some additional due diligence before going ahead. I need to coordinate some resources on the Kaggle side, and will rescore / update the discussion forums tomorrow morning.</p>\n\n<p>Re: additional detail on the metric, I'll add some. Thanks for the feedback!</p>",
          "rawMarkdown": "Hi - thanks everybody! Sorry for the delay - we and the host needed to do some additional due diligence before going ahead. I need to coordinate some resources on the Kaggle side, and will rescore / update the discussion forums tomorrow morning.\n\nRe: additional detail on the metric, I'll add some. Thanks for the feedback!",
          "votes": 5
        },
        {
          "id": 694739,
          "postDate": "2019-12-14T02:51:33.410Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi Phil,</p>\n\n<p>Could you give us more detail on the metric as you wrote <code>additional detail on the metric, I'll add some.</code>?</p>\n\n<p>Thank you in advance.</p>",
          "rawMarkdown": "@philculliton Hi Phil,\n\nCould you give us more detail on the metric as you wrote `additional detail on the metric, I'll add some.`?\n\nThank you in advance.",
          "votes": 3
        }
      ]
    },
    {
      "id": 671936,
      "postDate": "2019-11-13T10:45:41.073Z",
      "content": "<p>Hi,  <a href=\"/philculliton\">@philculliton</a>, i also think that metric's implementation is right. But the idea to calculate TP/FP is very strange. What about FN? Because without FN such behavior of metric is reasonable. Now metric evaluate only mean quality of predictions, no matter how much cars we predict.   So, for 10 good predicted car metric is higher then for 20000 not so accurate predicted cars. I suggest to talk with organizers and change metric, because now this competition  looks like: \"Choose small subset of cars (maybe one car, in private part of test) which easiest for prediction and predict coordinates of these cars (this car) as good as possible \".  I think organizers wont get any profit from such competition.  Correct me if i wrong about something </p>",
      "rawMarkdown": "Hi,  @philculliton, i also think that metric's implementation is right. But the idea to calculate TP/FP is very strange. What about FN? Because without FN such behavior of metric is reasonable. Now metric evaluate only mean quality of predictions, no matter how much cars we predict.   So, for 10 good predicted car metric is higher then for 20000 not so accurate predicted cars. I suggest to talk with organizers and change metric, because now this competition  looks like: \"Choose small subset of cars (maybe one car, in private part of test) which easiest for prediction and predict coordinates of these cars (this car) as good as possible \".  I think organizers wont get any profit from such competition.  Correct me if i wrong about something ",
      "votes": 22,
      "replies": [
        {
          "id": 672006,
          "postDate": "2019-11-13T12:51:44.457Z",
          "content": "<p>Looks like it's the case for now</p>",
          "rawMarkdown": "Looks like it's the case for now",
          "votes": 1
        },
        {
          "id": 673344,
          "postDate": "2019-11-14T21:34:49.160Z",
          "content": "<p>I think you are correct. The metric is correct but just meaningless for self-driving</p>",
          "rawMarkdown": "I think you are correct. The metric is correct but just meaningless for self-driving",
          "votes": 1
        },
        {
          "id": 674364,
          "postDate": "2019-11-16T10:37:41.620Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> could you please look into this. This seems to be the case here and can not be the aim for the competition.</p>",
          "rawMarkdown": "@philculliton could you please look into this. This seems to be the case here and can not be the aim for the competition.",
          "votes": 7
        }
      ]
    },
    {
      "id": 670626,
      "postDate": "2019-11-11T17:01:04.500Z",
      "content": "<p>I have to say that it is almost impossible to participate in a competition without an adequate understanding of the metric. But the proposed description brings up more questions than answers, even if the leaderboard is calculated correctly.\nPlease, provide the code.</p>",
      "rawMarkdown": "I have to say that it is almost impossible to participate in a competition without an adequate understanding of the metric. But the proposed description brings up more questions than answers, even if the leaderboard is calculated correctly.\nPlease, provide the code.\n",
      "votes": 12
    },
    {
      "id": 670297,
      "postDate": "2019-11-11T09:38:27.833Z",
      "content": "<p>I agree, the LB score is almost definitely wrong. My score is from a submission that only predicts ~15 cars in 2021 images...</p>",
      "rawMarkdown": "I agree, the LB score is almost definitely wrong. My score is from a submission that only predicts ~15 cars in 2021 images...",
      "votes": 12,
      "replies": [
        {
          "id": 671651,
          "postDate": "2019-11-13T01:54:23.520Z",
          "content": "<p>Hey Branden, are you saying you only provide in total ~15 car for all the 2021 images or 15*2021 cars? I do find the evaluation rather confusing in this challenge.</p>",
          "rawMarkdown": "Hey Branden, are you saying you only provide in total ~15 car for all the 2021 images or 15*2021 cars? I do find the evaluation rather confusing in this challenge."
        },
        {
          "id": 671707,
          "postDate": "2019-11-13T04:19:00.743Z",
          "content": "<p>Yes, I'd also like clarity re: what you mean here, <a href=\"/brandenkmurray\">@brandenkmurray</a>. Thanks!</p>",
          "rawMarkdown": "Yes, I'd also like clarity re: what you mean here, @brandenkmurray. Thanks!"
        },
        {
          "id": 671751,
          "postDate": "2019-11-13T05:43:20.663Z",
          "content": "<p>I am sure that he means ~15 cars for all 2021 images. I am from 6th place and my prediction consists only 200 cars (in total) in 2021 images.</p>",
          "rawMarkdown": "I am sure that he means ~15 cars for all 2021 images. I am from 6th place and my prediction consists only 200 cars (in total) in 2021 images.",
          "votes": 3
        },
        {
          "id": 671906,
          "postDate": "2019-11-13T09:50:45.960Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> ~15 car for all the 2021 images. I predicted nothing for almost every image.</p>",
          "rawMarkdown": "@philculliton ~15 car for all the 2021 images. I predicted nothing for almost every image.",
          "votes": 5
        },
        {
          "id": 672548,
          "postDate": "2019-11-14T01:35:02.273Z",
          "content": "<p>Then it's likely that the evaluation metric does not taken False Negative into account which is obviously wrong....</p>",
          "rawMarkdown": "Then it's likely that the evaluation metric does not taken False Negative into account which is obviously wrong...."
        },
        {
          "id": 672851,
          "postDate": "2019-11-14T08:22:49.017Z",
          "content": "<p>sounds like manual labeling one of easy test data, submit and you win this competition lol</p>\n\n<p>hope that this will be fixed ASAP</p>",
          "rawMarkdown": "sounds like manual labeling one of easy test data, submit and you win this competition lol\n\nhope that this will be fixed ASAP",
          "votes": 2
        },
        {
          "id": 672857,
          "postDate": "2019-11-14T08:30:35.610Z",
          "content": "<p>Hhhhh... ... You are a genius!</p>",
          "rawMarkdown": "Hhhhh... ... You are a genius!"
        },
        {
          "id": 675357,
          "postDate": "2019-11-18T01:03:56.487Z",
          "content": "<p>OMG!!!!\nI'm now having my Public LB 0.122 score by predicting 8000+ cars!!!!!\nHow competition hosts think about this !?</p>",
          "rawMarkdown": "OMG!!!!\nI'm now having my Public LB 0.122 score by predicting 8000+ cars!!!!!\nHow competition hosts think about this !?"
        }
      ]
    },
    {
      "id": 669885,
      "postDate": "2019-11-10T15:57:10.440Z",
      "content": "<p>Just posted about this too. My original threshold was &gt;=99 which was a score of .178. Next threshold was &gt;=99.5 which was a score of .199. I actually noticed this at the beginning of the competition and I've always been confused...</p>",
      "rawMarkdown": "Just posted about this too. My original threshold was &gt;=99 which was a score of .178. Next threshold was &gt;=99.5 which was a score of .199. I actually noticed this at the beginning of the competition and I've always been confused...",
      "votes": 7
    },
    {
      "id": 671000,
      "postDate": "2019-11-12T06:10:37.907Z",
      "content": "<p>I really hope the hosts are going to do something about that. Please publish the LB calculations, it would help to pinpoint the problem.</p>",
      "rawMarkdown": "I really hope the hosts are going to do something about that. Please publish the LB calculations, it would help to pinpoint the problem.",
      "votes": 5
    },
    {
      "id": 671706,
      "postDate": "2019-11-13T04:18:18.317Z",
      "content": "<p>Hi all! Thanks for your feedback and questions. I'm looking into this. This metric is not a classic COCO-style metric, so behaviors may differ from what you're expecting. I'll try and nail down where the dissonance is, and make sure it's not a bug in the metric itself.</p>\n\n<p><a href=\"/boliu0\">@boliu0</a> - what threshold are you changing? By <code>prob</code> do you mean your confidence measurement?</p>",
      "rawMarkdown": "Hi all! Thanks for your feedback and questions. I'm looking into this. This metric is not a classic COCO-style metric, so behaviors may differ from what you're expecting. I'll try and nail down where the dissonance is, and make sure it's not a bug in the metric itself.\n\n@boliu0 - what threshold are you changing? By `prob` do you mean your confidence measurement?",
      "votes": 4,
      "replies": [
        {
          "id": 671721,
          "postDate": "2019-11-13T04:42:36.157Z",
          "content": "<p>Hi Phil, yes by <code>prob</code> I mean the confidence score of each predicted car instance.</p>\n\n<p>Consider two submissions:\nA: thres=0.5. It has all the predicted car instances with prob&gt;0.5\nB: thres=0.25. It has all the predicted car instances with prob&gt;0.25. So it's a superset of all the above car instances, plus more. So submission B's score should be &gt;= submission A's score. This is because at each recall level (0.01, 0.02, ... , 0.99, 1.00), B's precision &gt;= A's precision.</p>\n\n<p>Do you agree with above statements? They should hold true for any mAP metric.</p>\n\n<p>Regardless, I strongly encourage you publish the evaluation code, because of the uniqueness (3-D object detection with a custom metric) of this competition. For example, many of us also have doubts in the distance thresholds as discussed <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/115630#latest-665163\">here</a>. </p>\n\n<p>The evaluation code can give all participants a crystal clear understanding of the metric without any ambiguity. If it has bug, crowd-sourcing the bug-finding can save you a lot of time and effort.</p>\n\n<p>Thanks.</p>",
          "rawMarkdown": "Hi Phil, yes by `prob` I mean the confidence score of each predicted car instance.\n\nConsider two submissions:\nA: thres=0.5. It has all the predicted car instances with prob&gt;0.5\nB: thres=0.25. It has all the predicted car instances with prob&gt;0.25. So it's a superset of all the above car instances, plus more. So submission B's score should be &gt;= submission A's score. This is because at each recall level (0.01, 0.02, ... , 0.99, 1.00), B's precision &gt;= A's precision.\n\nDo you agree with above statements? They should hold true for any mAP metric.\n\nRegardless, I strongly encourage you publish the evaluation code, because of the uniqueness (3-D object detection with a custom metric) of this competition. For example, many of us also have doubts in the distance thresholds as discussed [here](https://www.kaggle.com/c/pku-autonomous-driving/discussion/115630#latest-665163). \n\nThe evaluation code can give all participants a crystal clear understanding of the metric without any ambiguity. If it has bug, crowd-sourcing the bug-finding can save you a lot of time and effort.\n\nThanks.\n",
          "votes": 7
        },
        {
          "id": 671784,
          "postDate": "2019-11-13T06:33:27.913Z",
          "content": "<p>Hey Bo, I think you have a misunderstanding of mAP\n<a href=\"https://medium.com/&lt;a href=\">@jonathan</a>_hui/map-mean-average-precision-for-object-detection-45c121a31173\"&gt;https://medium.com/<a href=\"/jonathan\">@jonathan</a>_hui/map-mean-average-precision-for-object-detection-45c121a31173\nAs in the picture below if there are 5 apples (5 true positives).</p>\n\n<p></p>\n\n<p>Hence, it's possible your set A can have higher mAP than set B.</p>",
          "rawMarkdown": "Hey Bo, I think you have a misunderstanding of mAP\nhttps://medium.com/@jonathan_hui/map-mean-average-precision-for-object-detection-45c121a31173\nAs in the picture below if there are 5 apples (5 true positives).\n\n![mAP](https://miro.medium.com/max/3336/1*9ordwhXD68cKCGzuJaH2Rg.png)\n\nHence, it's possible your set A can have higher mAP than set B.",
          "votes": -1
        },
        {
          "id": 671786,
          "postDate": "2019-11-13T06:34:39.130Z",
          "content": "<p>But still, I think the challenge's evaluation code has bug...</p>",
          "rawMarkdown": "But still, I think the challenge's evaluation code has bug..."
        },
        {
          "id": 672170,
          "postDate": "2019-11-13T15:25:06.280Z",
          "content": "<p>Hi stevenwudi, I don't think I follow your example.</p>\n\n<p>In your example, if submission A has the first 5 predictions and submission B has first 7 predictions.\nThen A's AP is: \n<code>(Precision@0.1 recall + Precision@0.2 recall + ... + Precision@1.0 recall)/10 = (1+1+1+1+0+0+0+0+0+0)/10=0.4</code>\nbecause A didn't reach 0.5 recall</p>\n\n<p>B's AP is <code>(1+1+1+1+0.57+0.57+0.57+0.57+0+0)/10=0.628</code></p>\n\n<p>Note that Average Precision is not the same as Precision. It looks like LB's metric is Precision. At the extreme, if you only predict 1 instance in the whole test set and it is TP, then you get 1.0 score, which is clearly wrong.</p>",
          "rawMarkdown": "Hi stevenwudi, I don't think I follow your example.\n\nIn your example, if submission A has the first 5 predictions and submission B has first 7 predictions.\nThen A's AP is: \n`(Precision@0.1 recall + Precision@0.2 recall + ... + Precision@1.0 recall)/10 = (1+1+1+1+0+0+0+0+0+0)/10=0.4`\nbecause A didn't reach 0.5 recall\n\nB's AP is `(1+1+1+1+0.57+0.57+0.57+0.57+0+0)/10=0.628`\n\n\nNote that Average Precision is not the same as Precision. It looks like LB's metric is Precision. At the extreme, if you only predict 1 instance in the whole test set and it is TP, then you get 1.0 score, which is clearly wrong.",
          "votes": 3
        },
        {
          "id": 672573,
          "postDate": "2019-11-14T02:01:51.597Z",
          "content": "<p>Hi Bo, I have a question about your explanation <code>This is because at each recall level (0.01, 0.02, … , 0.99, 1.00), B's precision &gt;= A's precision.</code>\nFor the recall level that A and B both can reach, their precision will always be the same. It is the recall level that differs the score. Correct me if i am wrong about something</p>",
          "rawMarkdown": "Hi Bo, I have a question about your explanation `This is because at each recall level (0.01, 0.02, … , 0.99, 1.00), B's precision &gt;= A's precision.`\nFor the recall level that A and B both can reach, their precision will always be the same. It is the recall level that differs the score. Correct me if i am wrong about something"
        },
        {
          "id": 672630,
          "postDate": "2019-11-14T03:06:28.950Z",
          "content": "<p>hi WuYang, yes I think that's right. When A cannot reach a certain recall level but B can, then B's precision at this recall level &gt; 0, and A's precision at this recall level = 0.</p>",
          "rawMarkdown": "hi WuYang, yes I think that's right. When A cannot reach a certain recall level but B can, then B's precision at this recall level &gt; 0, and A's precision at this recall level = 0."
        },
        {
          "id": 676629,
          "postDate": "2019-11-19T11:23:58.127Z",
          "content": "<p>I believe you should look into this metric of evaluating mAP as it includes TP, FP and FN: <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/overview/evaluation\">mAP eval</a></p>",
          "rawMarkdown": "I believe you should look into this metric of evaluating mAP as it includes TP, FP and FN: [mAP eval](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/overview/evaluation)"
        },
        {
          "id": 676951,
          "postDate": "2019-11-19T16:40:13.460Z",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , is there an update?</p>",
          "rawMarkdown": "Hi @philculliton , is there an update?",
          "votes": 12
        },
        {
          "id": 680743,
          "postDate": "2019-11-25T06:48:49.310Z",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> ,\nIt would be great to hear an update, since a lot of kagglers would like to participate in a competition with fair metrics. Thanks!</p>",
          "rawMarkdown": "Hi @philculliton ,\nIt would be great to hear an update, since a lot of kagglers would like to participate in a competition with fair metrics. Thanks!",
          "votes": 3
        }
      ]
    },
    {
      "id": 671026,
      "postDate": "2019-11-12T06:55:31.003Z",
      "content": "<p>I have written this question in the \"Welcome\" topic. Hope hosts will answer us soon. Btw I did a small research of this topic. First of all we have to notice that our leaderboard is build by 180 images and this score wouldn't  correlate with private. Probably they gives a low weight for false negative cars and at least one correctly found car raise your score into the sky.</p>",
      "rawMarkdown": "I have written this question in the \"Welcome\" topic. Hope hosts will answer us soon. Btw I did a small research of this topic. First of all we have to notice that our leaderboard is build by 180 images and this score wouldn't  correlate with private. Probably they gives a low weight for false negative cars and at least one correctly found car raise your score into the sky."
    },
    {
      "id": 674492,
      "postDate": "2019-11-16T15:19:25.763Z",
      "rawMarkdown": "",
      "votes": 12,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 681067,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2019-11-25T16:18:27.703000",
      "content": "<p>Hi all - thanks again for your feedback. Sorry for the delay! I've been working with the host and other data scientists to make sure we didn't have bugs in the metric re: the submissions presented in this thread, and also working through next steps. We will very likely be adding FN to the calculation and rescoring this week.</p>",
      "votes": 27,
      "replies": [
        {
          "id": 681094,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2019-11-25T16:48:42.507000",
          "content": "<p>Thanks for the update!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 681138,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-11-25T18:04:24.570000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Thanks a lot for managing this and thanks for giving us an update on this matter. Much appreciated!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 681257,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2019-11-25T22:59:04.353000",
          "content": "<p>Great to hear that! Sure will switch gears after reflected :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 681461,
          "author_name": "Anton-br",
          "author_url": "",
          "post_date": "2019-11-26T06:25:12.070000",
          "content": "<p>Thank you for attention to the problem! Nevertheless, I think that just rescoring and changing the metrics is not enough. Would you mind to change the Evaluation page as well and describe the metic more thoroughly, because without understanding the work of the metric, we will not be able to properly validate our models. </p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 681559,
          "author_name": "OkkBand",
          "author_url": "",
          "post_date": "2019-11-26T08:37:56.227000",
          "content": "<p>Yes, you are right, we need the evaluation code for the validation. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 682042,
          "author_name": "shatz",
          "author_url": "",
          "post_date": "2019-11-26T20:52:15.070000",
          "content": "<p>I don't think you need to give us the code (though maybe pseudocode would be nice), but a very thorough description of how it works with an equation would be great.</p>\n\n<p>Assuming this change in the evaluation metric impacts the leaderboard significantly, would it be normal to ask for an extension by a few weeks for this competition? A lot of people probably spent time chasing the leaderboard results without any regard for actually having a better model, though maybe thats their fault 🤷‍♂️</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 684551,
          "author_name": "He",
          "author_url": "",
          "post_date": "2019-11-30T02:16:03.117000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi philculliton , This week is drawing to a close, will the score be recalculated within this week, and then the new score calculation method will be announced, is that right?</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 686265,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2019-12-03T01:28:16.603000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi Phil, could you kindly give us some update about the evaluation metric rescoring? Your plan was to update last week....thanks</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 686351,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-12-03T03:54:13.610000",
          "content": "<p>Hi - thanks everybody! Sorry for the delay - we and the host needed to do some additional due diligence before going ahead. I need to coordinate some resources on the Kaggle side, and will rescore / update the discussion forums tomorrow morning.</p>\n\n<p>Re: additional detail on the metric, I'll add some. Thanks for the feedback!</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 694739,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2019-12-14T02:51:33.410000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Hi Phil,</p>\n\n<p>Could you give us more detail on the metric as you wrote <code>additional detail on the metric, I'll add some.</code>?</p>\n\n<p>Thank you in advance.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 671936,
      "author_name": "Evgenii Kononenko",
      "author_url": "",
      "post_date": "2019-11-13T10:45:41.073000",
      "content": "<p>Hi,  <a href=\"/philculliton\">@philculliton</a>, i also think that metric's implementation is right. But the idea to calculate TP/FP is very strange. What about FN? Because without FN such behavior of metric is reasonable. Now metric evaluate only mean quality of predictions, no matter how much cars we predict.   So, for 10 good predicted car metric is higher then for 20000 not so accurate predicted cars. I suggest to talk with organizers and change metric, because now this competition  looks like: \"Choose small subset of cars (maybe one car, in private part of test) which easiest for prediction and predict coordinates of these cars (this car) as good as possible \".  I think organizers wont get any profit from such competition.  Correct me if i wrong about something </p>",
      "votes": 22,
      "replies": [
        {
          "id": 672006,
          "author_name": "YALICKJ",
          "author_url": "",
          "post_date": "2019-11-13T12:51:44.457000",
          "content": "<p>Looks like it's the case for now</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673344,
          "author_name": "Chuyao Shen",
          "author_url": "",
          "post_date": "2019-11-14T21:34:49.160000",
          "content": "<p>I think you are correct. The metric is correct but just meaningless for self-driving</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674364,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-11-16T10:37:41.620000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> could you please look into this. This seems to be the case here and can not be the aim for the competition.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 670626,
      "author_name": "Michael Diskin",
      "author_url": "",
      "post_date": "2019-11-11T17:01:04.500000",
      "content": "<p>I have to say that it is almost impossible to participate in a competition without an adequate understanding of the metric. But the proposed description brings up more questions than answers, even if the leaderboard is calculated correctly.\nPlease, provide the code.</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 670297,
      "author_name": "Branden Murray",
      "author_url": "",
      "post_date": "2019-11-11T09:38:27.833000",
      "content": "<p>I agree, the LB score is almost definitely wrong. My score is from a submission that only predicts ~15 cars in 2021 images...</p>",
      "votes": 12,
      "replies": [
        {
          "id": 671651,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2019-11-13T01:54:23.520000",
          "content": "<p>Hey Branden, are you saying you only provide in total ~15 car for all the 2021 images or 15*2021 cars? I do find the evaluation rather confusing in this challenge.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 671707,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-11-13T04:19:00.743000",
          "content": "<p>Yes, I'd also like clarity re: what you mean here, <a href=\"/brandenkmurray\">@brandenkmurray</a>. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 671751,
          "author_name": "Anton-br",
          "author_url": "",
          "post_date": "2019-11-13T05:43:20.663000",
          "content": "<p>I am sure that he means ~15 cars for all 2021 images. I am from 6th place and my prediction consists only 200 cars (in total) in 2021 images.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 671906,
          "author_name": "Branden Murray",
          "author_url": "",
          "post_date": "2019-11-13T09:50:45.960000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> ~15 car for all the 2021 images. I predicted nothing for almost every image.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 672548,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2019-11-14T01:35:02.273000",
          "content": "<p>Then it's likely that the evaluation metric does not taken False Negative into account which is obviously wrong....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672851,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2019-11-14T08:22:49.017000",
          "content": "<p>sounds like manual labeling one of easy test data, submit and you win this competition lol</p>\n\n<p>hope that this will be fixed ASAP</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 672857,
          "author_name": "Xianzhong",
          "author_url": "",
          "post_date": "2019-11-14T08:30:35.610000",
          "content": "<p>Hhhhh... ... You are a genius!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 675357,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2019-11-18T01:03:56.487000",
          "content": "<p>OMG!!!!\nI'm now having my Public LB 0.122 score by predicting 8000+ cars!!!!!\nHow competition hosts think about this !?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 669885,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2019-11-10T15:57:10.440000",
      "content": "<p>Just posted about this too. My original threshold was &gt;=99 which was a score of .178. Next threshold was &gt;=99.5 which was a score of .199. I actually noticed this at the beginning of the competition and I've always been confused...</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 671000,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2019-11-12T06:10:37.907000",
      "content": "<p>I really hope the hosts are going to do something about that. Please publish the LB calculations, it would help to pinpoint the problem.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 671706,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2019-11-13T04:18:18.317000",
      "content": "<p>Hi all! Thanks for your feedback and questions. I'm looking into this. This metric is not a classic COCO-style metric, so behaviors may differ from what you're expecting. I'll try and nail down where the dissonance is, and make sure it's not a bug in the metric itself.</p>\n\n<p><a href=\"/boliu0\">@boliu0</a> - what threshold are you changing? By <code>prob</code> do you mean your confidence measurement?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 671721,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2019-11-13T04:42:36.157000",
          "content": "<p>Hi Phil, yes by <code>prob</code> I mean the confidence score of each predicted car instance.</p>\n\n<p>Consider two submissions:\nA: thres=0.5. It has all the predicted car instances with prob&gt;0.5\nB: thres=0.25. It has all the predicted car instances with prob&gt;0.25. So it's a superset of all the above car instances, plus more. So submission B's score should be &gt;= submission A's score. This is because at each recall level (0.01, 0.02, ... , 0.99, 1.00), B's precision &gt;= A's precision.</p>\n\n<p>Do you agree with above statements? They should hold true for any mAP metric.</p>\n\n<p>Regardless, I strongly encourage you publish the evaluation code, because of the uniqueness (3-D object detection with a custom metric) of this competition. For example, many of us also have doubts in the distance thresholds as discussed <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/115630#latest-665163\">here</a>. </p>\n\n<p>The evaluation code can give all participants a crystal clear understanding of the metric without any ambiguity. If it has bug, crowd-sourcing the bug-finding can save you a lot of time and effort.</p>\n\n<p>Thanks.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 671784,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2019-11-13T06:33:27.913000",
          "content": "<p>Hey Bo, I think you have a misunderstanding of mAP\n<a href=\"https://medium.com/&lt;a href=\">@jonathan</a>_hui/map-mean-average-precision-for-object-detection-45c121a31173\"&gt;https://medium.com/<a href=\"/jonathan\">@jonathan</a>_hui/map-mean-average-precision-for-object-detection-45c121a31173\nAs in the picture below if there are 5 apples (5 true positives).</p>\n\n<p></p>\n\n<p>Hence, it's possible your set A can have higher mAP than set B.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 671786,
          "author_name": "stevenwudi",
          "author_url": "",
          "post_date": "2019-11-13T06:34:39.130000",
          "content": "<p>But still, I think the challenge's evaluation code has bug...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672170,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2019-11-13T15:25:06.280000",
          "content": "<p>Hi stevenwudi, I don't think I follow your example.</p>\n\n<p>In your example, if submission A has the first 5 predictions and submission B has first 7 predictions.\nThen A's AP is: \n<code>(Precision@0.1 recall + Precision@0.2 recall + ... + Precision@1.0 recall)/10 = (1+1+1+1+0+0+0+0+0+0)/10=0.4</code>\nbecause A didn't reach 0.5 recall</p>\n\n<p>B's AP is <code>(1+1+1+1+0.57+0.57+0.57+0.57+0+0)/10=0.628</code></p>\n\n<p>Note that Average Precision is not the same as Precision. It looks like LB's metric is Precision. At the extreme, if you only predict 1 instance in the whole test set and it is TP, then you get 1.0 score, which is clearly wrong.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 672573,
          "author_name": "WuYang",
          "author_url": "",
          "post_date": "2019-11-14T02:01:51.597000",
          "content": "<p>Hi Bo, I have a question about your explanation <code>This is because at each recall level (0.01, 0.02, … , 0.99, 1.00), B's precision &gt;= A's precision.</code>\nFor the recall level that A and B both can reach, their precision will always be the same. It is the recall level that differs the score. Correct me if i am wrong about something</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672630,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2019-11-14T03:06:28.950000",
          "content": "<p>hi WuYang, yes I think that's right. When A cannot reach a certain recall level but B can, then B's precision at this recall level &gt; 0, and A's precision at this recall level = 0.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676629,
          "author_name": "Muccici",
          "author_url": "",
          "post_date": "2019-11-19T11:23:58.127000",
          "content": "<p>I believe you should look into this metric of evaluating mAP as it includes TP, FP and FN: <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/overview/evaluation\">mAP eval</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676951,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2019-11-19T16:40:13.460000",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , is there an update?</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 680743,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2019-11-25T06:48:49.310000",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> ,\nIt would be great to hear an update, since a lot of kagglers would like to participate in a competition with fair metrics. Thanks!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 671026,
      "author_name": "Anton-br",
      "author_url": "",
      "post_date": "2019-11-12T06:55:31.003000",
      "content": "<p>I have written this question in the \"Welcome\" topic. Hope hosts will answer us soon. Btw I did a small research of this topic. First of all we have to notice that our leaderboard is build by 180 images and this score wouldn't  correlate with private. Probably they gives a low weight for false negative cars and at least one correctly found car raise your score into the sky.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 674492,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-16T15:19:25.763000",
      "content": "",
      "votes": 12,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "669866": "In coco style mAP metric, for the same model, the score should be a monotone non-increasing function of the probability threshold. In other words, the lower the threshold, the higher the score should be. That was the case in the [open images competitions](https://www.kaggle.com/c/open-images-2019-object-detection)\n\nBut apparently this is not the case in this competition. For the same [model](https://www.kaggle.com/hocop1/centernet-baseline?scriptVersionId=22630284) (best public notebook), the score is:\n- 0.110 with thres=0.75 (i.e. only output predicted car instances with prob&gt;=0.75, which is a subset of below two submissions)\n- 0.093 with thres=0.5\n- 0.074 with thres=0.25\n\nDoes anyone notice the same thing?\n@philculliton and @addisonhoward, is it possible to check this, or even better, post the evaluation code? Thank you.",
    "681067": "Hi all - thanks again for your feedback. Sorry for the delay! I've been working with the host and other data scientists to make sure we didn't have bugs in the metric re: the submissions presented in this thread, and also working through next steps. We will very likely be adding FN to the calculation and rescoring this week.",
    "671936": "Hi,  @philculliton, i also think that metric's implementation is right. But the idea to calculate TP/FP is very strange. What about FN? Because without FN such behavior of metric is reasonable. Now metric evaluate only mean quality of predictions, no matter how much cars we predict.   So, for 10 good predicted car metric is higher then for 20000 not so accurate predicted cars. I suggest to talk with organizers and change metric, because now this competition  looks like: \"Choose small subset of cars (maybe one car, in private part of test) which easiest for prediction and predict coordinates of these cars (this car) as good as possible \".  I think organizers wont get any profit from such competition.  Correct me if i wrong about something ",
    "670626": "I have to say that it is almost impossible to participate in a competition without an adequate understanding of the metric. But the proposed description brings up more questions than answers, even if the leaderboard is calculated correctly.\nPlease, provide the code.\n",
    "670297": "I agree, the LB score is almost definitely wrong. My score is from a submission that only predicts ~15 cars in 2021 images...",
    "669885": "Just posted about this too. My original threshold was &gt;=99 which was a score of .178. Next threshold was &gt;=99.5 which was a score of .199. I actually noticed this at the beginning of the competition and I've always been confused...",
    "671000": "I really hope the hosts are going to do something about that. Please publish the LB calculations, it would help to pinpoint the problem.",
    "671706": "Hi all! Thanks for your feedback and questions. I'm looking into this. This metric is not a classic COCO-style metric, so behaviors may differ from what you're expecting. I'll try and nail down where the dissonance is, and make sure it's not a bug in the metric itself.\n\n@boliu0 - what threshold are you changing? By `prob` do you mean your confidence measurement?",
    "671026": "I have written this question in the \"Welcome\" topic. Hope hosts will answer us soon. Btw I did a small research of this topic. First of all we have to notice that our leaderboard is build by 180 images and this score wouldn't  correlate with private. Probably they gives a low weight for false negative cars and at least one correctly found car raise your score into the sky.",
    "674492": ""
  }
}