{
  "id": 64860,
  "title": "How to get official scoring metric?",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64860",
  "author_name": "",
  "post_date": "2018-09-03T11:03:49.266256900Z",
  "votes": 48,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hi @philculliton !</p>\n\n<p>Can you pls provide example of scoring method? Because now it is unclear how to use confidence score data.</p>",
  "messages": [
    {
      "id": "380739",
      "postDate": "09/03/2018 11:03:49",
      "content": "<p>Hi @philculliton !</p>\n\n<p>Can you pls provide example of scoring method? Because now it is unclear how to use confidence score data.</p>",
      "rawMarkdown": "Hi @philculliton !\n\nCan you pls provide example of scoring method? Because now it is unclear how to use confidence score data.",
      "votes": null
    },
    {
      "id": "380755",
      "postDate": "09/03/2018 12:03:00",
      "content": "<p>Hello, Konstantin. Are you want to get theoretical example of metric computation, or you need programm implementation?</p>",
      "rawMarkdown": "Hello, Konstantin. Are you want to get theoretical example of metric computation, or you need programm implementation?",
      "votes": null
    },
    {
      "id": "380761",
      "postDate": "09/03/2018 12:18:13",
      "content": "<p>I want to get <strong>official</strong> programm implementation.</p>",
      "rawMarkdown": "I want to get **official** programm implementation.",
      "votes": null
    },
    {
      "id": "381195",
      "postDate": "09/04/2018 08:54:24",
      "content": "<p>I want it too!!!</p>",
      "rawMarkdown": "I want it too!!!",
      "votes": null
    },
    {
      "id": "381208",
      "postDate": "09/04/2018 09:25:28",
      "content": "<p>+</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "381485",
      "postDate": "09/04/2018 17:33:42",
      "content": "<p>Thanks for your interest.  I have updated the <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge#evaluation\">Evaluation</a> page with more information.  Confidence is a general-purpose tool to resolve edge cases in submission box ordering - it is there primarily as a failsafe if any extreme edge cases arise.  In this particular data set I have not seen <code>confidence</code> change scores.</p>\n\n<p>I would recommend that you calculate a <code>confidence</code> score to ensure that, if any edge cases do arise, your submission is evaluated in order of how sure you are of your submission boxes.</p>",
      "rawMarkdown": "Thanks for your interest.  I have updated the [Evaluation][1] page with more information.  Confidence is a general-purpose tool to resolve edge cases in submission box ordering - it is there primarily as a failsafe if any extreme edge cases arise.  In this particular data set I have not seen `confidence` change scores.\n\nI would recommend that you calculate a `confidence` score to ensure that, if any edge cases do arise, your submission is evaluated in order of how sure you are of your submission boxes.\n\n  [1]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge#evaluation",
      "votes": null
    },
    {
      "id": "381803",
      "postDate": "09/05/2018 08:19:45",
      "content": "<p>Thank you for clarification.</p>",
      "rawMarkdown": "Thank you for clarification.",
      "votes": null
    },
    {
      "id": "382187",
      "postDate": "09/05/2018 20:50:01",
      "content": "<p>Hi, Phil. For negative cases, does it mean I can predict anything without receiving any penalty (since TP will always be 0)?</p>",
      "rawMarkdown": "Hi, Phil. For negative cases, does it mean I can predict anything without receiving any penalty (since TP will always be 0)?",
      "votes": null
    },
    {
      "id": "382196",
      "postDate": "09/05/2018 21:08:40",
      "content": "<p>Hey xinario,</p>\n\n<p>No - if the sample is a true negative with <strong>no</strong> predictions, then it is not counted in the mean.  If <em>any</em> predictions are made, the whole sample counts as a 0 and is counted in the mean.</p>",
      "rawMarkdown": "Hey xinario,\n\nNo - if the sample is a true negative with **no** predictions, then it is not counted in the mean.  If *any* predictions are made, the whole sample counts as a 0 and is counted in the mean.",
      "votes": null
    },
    {
      "id": "382197",
      "postDate": "09/05/2018 21:09:24",
      "content": "<p>Hi <a href=\"/xinario\">@xinario</a>. On the Evaluation page, they state:</p>\n\n<blockquote>\n  <p>Important note: if there are no ground truth objects at all for a\n  given image, ANY number of predictions (false positives) will result\n  in the image receiving a score of zero, and being included in the mean\n  average precision.</p>\n</blockquote>\n\n<p>So in the basics of numerator/denominator, if you are correct on a negative case, you add 0 to the numerator and denominator; it's not included in result--you don't gain anything, but at least you don't lose anything. If you are incorrect in the negative case you add 0 to the numerator and 1 to the denominator as if you guessed entirely the wrong box. So the negatives can only hurt, they can't help. But an important part of getting a good score is surely avoiding those False Positives.</p>\n\n<p>Hope that helps.</p>",
      "rawMarkdown": "Hi @xinario. On the Evaluation page, they state:\n\n&gt; Important note: if there are no ground truth objects at all for a\n&gt; given image, ANY number of predictions (false positives) will result\n&gt; in the image receiving a score of zero, and being included in the mean\n&gt; average precision.\n\nSo in the basics of numerator/denominator, if you are correct on a negative case, you add 0 to the numerator and denominator; it's not included in result--you don't gain anything, but at least you don't lose anything. If you are incorrect in the negative case you add 0 to the numerator and 1 to the denominator as if you guessed entirely the wrong box. So the negatives can only hurt, they can't help. But an important part of getting a good score is surely avoiding those False Positives.\n\nHope that helps.",
      "votes": null
    },
    {
      "id": "382199",
      "postDate": "09/05/2018 21:19:57",
      "content": "<p>So that means for a single negative sample, it actually doesn't matter how many false positive predictions (&gt;0) I got. </p>",
      "rawMarkdown": "So that means for a single negative sample, it actually doesn't matter how many false positive predictions (&gt;0) I got.",
      "votes": null
    },
    {
      "id": "382201",
      "postDate": "09/05/2018 21:24:09",
      "content": "<p>Yeah, thanks for the classification. What I got so far is that we have to avoid predicting bounding boxes on negative samples. But once you do, it actually does matter if it's 1 or 2 or even more. The final score will be the same.</p>",
      "rawMarkdown": "Yeah, thanks for the classification. What I got so far is that we have to avoid predicting bounding boxes on negative samples. But once you do, it actually does matter if it's 1 or 2 or even more. The final score will be the same.",
      "votes": null
    },
    {
      "id": "382833",
      "postDate": "09/07/2018 07:37:17",
      "content": "<p>What score is used for single patient when there are no real boxes and we didn't predict any boxes? 1.0?</p>\n\n<p>I made an empty submission and it gets score = 0.0. It can be in 2 cases only: </p>\n\n<ul>\n<li>There are no empty cases in test set</li>\n<li>Scroing metric doesn't use empty boxes in test set.</li>\n</ul>",
      "rawMarkdown": "What score is used for single patient when there are no real boxes and we didn't predict any boxes? 1.0?\n\nI made an empty submission and it gets score = 0.0. It can be in 2 cases only: \n\n - There are no empty cases in test set\n - Scroing metric doesn't use empty boxes in test set.",
      "votes": null
    },
    {
      "id": "382843",
      "postDate": "09/07/2018 07:58:04",
      "content": "<p>On this case this patient don't take into account in competition metric. I mean\n we have three step of metric calculation:</p>\n\n<ol>\n<li><p>Founding amount of patient, who had no real boxes and had no predicted boxes. Remove these patient from our test set. Let's denote amount of remained patient by n</p></li>\n<li><p>For each other case we calculate average precision. So, we have n values: AP_1, AP_2, ..., AP_n</p></li>\n<li>Final score is (AP_1 + AP_2 +  AP_n) / n</li>\n</ol>",
      "rawMarkdown": "On this case this patient don't take into account in competition metric. I mean\n we have three step of metric calculation:\n\n 1. Founding amount of patient, who had no real boxes and had no predicted boxes. Remove these patient from our test set. Let's denote amount of remained patient by n\n\n 2. For each other case we calculate average precision. So, we have n values: AP_1, AP_2, ..., AP_n\n 3. Final score is (AP_1 + AP_2 +  AP_n) / n",
      "votes": null
    },
    {
      "id": "382847",
      "postDate": "09/07/2018 08:07:35",
      "content": "<p>It mean's that in your submition you have relatively small value n, and that's good for your final score cause you get little denominator. But you have AP_1 = AP_2 = ... = AP_n = 0, so that's bad for your metric, cause you get zero in numerator.</p>",
      "rawMarkdown": "It mean's that in your submition you have relatively small value n, and that's good for your final score cause you get little denominator. But you have AP_1 = AP_2 = ... = AP_n = 0, so that's bad for your metric, cause you get zero in numerator.",
      "votes": null
    },
    {
      "id": "382849",
      "postDate": "09/07/2018 08:09:18",
      "content": "<p>Thanks for explanation!</p>",
      "rawMarkdown": "Thanks for explanation!",
      "votes": null
    },
    {
      "id": "383687",
      "postDate": "09/09/2018 11:03:47",
      "content": "<p>I wonder what happens with the scoring if I have two predicted bounding boxes overlapping with a single ground-true object. An example (purple is grounf truth, light blue are predictions):</p>\n\n<p><a href=\"https://imgur.com/a/dy0tZbC\">example of overlapping predictions</a></p>\n\n<p>If I go with an algorithm:</p>\n\n<pre><code>for each prediction sorted by confidence descending:\n    find the best matching ground truth.\n    calculate IoU and note  TP, FP, FN for every threshold\n    remove 'processed' ground truth box from the list\n</code></pre>\n\n<p>Then I start from the lower predicted bbox, which has IoU &lt; 0.4 so I get 0 precision, even though my next bbox could have a lower confidence but higher IoU.</p>\n\n<p>Is this the algorithm used when scoring? In this case the 'confidence' is really important and it does not depend on an 'edge case' in the dataset but more on the 'edge case' in predictions.</p>",
      "rawMarkdown": "I wonder what happens with the scoring if I have two predicted bounding boxes overlapping with a single ground-true object. An example (purple is grounf truth, light blue are predictions):\n\n[example of overlapping predictions][1]\n\nIf I go with an algorithm:\n\n    for each prediction sorted by confidence descending:\n        find the best matching ground truth.\n        calculate IoU and note  TP, FP, FN for every threshold\n        remove 'processed' ground truth box from the list\n\nThen I start from the lower predicted bbox, which has IoU &lt; 0.4 so I get 0 precision, even though my next bbox could have a lower confidence but higher IoU.\n\nIs this the algorithm used when scoring? In this case the 'confidence' is really important and it does not depend on an 'edge case' in the dataset but more on the 'edge case' in predictions.\n\n  [1]: https://imgur.com/a/dy0tZbC",
      "votes": null
    },
    {
      "id": "383718",
      "postDate": "09/09/2018 12:38:31",
      "content": "<p>Good example. Yes, in this case confidence is important. But if seems, that situation like this, when we have to prediction boxes on one lung, have small probability. </p>",
      "rawMarkdown": "Good example. Yes, in this case confidence is important. But if seems, that situation like this, when we have to prediction boxes on one lung, have small probability.",
      "votes": null
    },
    {
      "id": "385141",
      "postDate": "09/10/2018 12:10:20",
      "content": "<p>Tomasz,</p>\n\n<p>IoU is also taken into account.  The predicted box with the highest IoU would be taken as the match regardless of confidence order.</p>\n\n<p>All ground truth boxes are scored against all predicted boxes.  No boxes are removed from contention.</p>",
      "rawMarkdown": "Tomasz,\n\nIoU is also taken into account.  The predicted box with the highest IoU would be taken as the match regardless of confidence order.\n\nAll ground truth boxes are scored against all predicted boxes.  No boxes are removed from contention.",
      "votes": null
    },
    {
      "id": "385146",
      "postDate": "09/10/2018 12:21:29",
      "content": "<p>I am confused. So, if i have many boxes that cover the gt to some extent, only the one with the highest iou will be scored? All the rest will be ignored?\nIf you can't publish the code, can you at least post official psaudo code? </p>",
      "rawMarkdown": "I am confused. So, if i have many boxes that cover the gt to some extent, only the one with the highest iou will be scored? All the rest will be ignored?\nIf you can't publish the code, can you at least post official psaudo code?",
      "votes": null
    },
    {
      "id": "385149",
      "postDate": "09/10/2018 12:27:53",
      "content": "<p>The one with the highest IoU will be scored.  The rest will be marked as false positives.</p>\n\n<p>I will consider creating and posting pseudo code, thanks for the suggestion.</p>",
      "rawMarkdown": "The one with the highest IoU will be scored.  The rest will be marked as false positives.\n\nI will consider creating and posting pseudo code, thanks for the suggestion.",
      "votes": null
    },
    {
      "id": "385326",
      "postDate": "09/10/2018 19:31:39",
      "content": "<p>Well the 'unofficial' but commonly used version is this: <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a></p>\n\n<p>and it goes with the following consequences - for any given threshold look at the predicted bbox in the order of confidence and if IoU is higher than the threshold - consider that a hit. That is mostly ok.\nBut then - this bbox cannot be used to match any other gt boxes.</p>\n\n<p>@Phil - It would be great if we can get a pseudo code. \nIt would be perfect if we can get a real implementation :)</p>",
      "rawMarkdown": "Well the 'unofficial' but commonly used version is this: https://www.kaggle.com/chenyc15/mean-average-precision-metric\n\nand it goes with the following consequences - for any given threshold look at the predicted bbox in the order of confidence and if IoU is higher than the threshold - consider that a hit. That is mostly ok.\nBut then - this bbox cannot be used to match any other gt boxes.\n\n\n@Phil - It would be great if we can get a pseudo code. \nIt would be perfect if we can get a real implementation :)",
      "votes": null
    },
    {
      "id": "386851",
      "postDate": "09/13/2018 18:33:06",
      "content": "<p>@Phil pseudo code would be nice for clarity</p>",
      "rawMarkdown": "Phil pseudo code would be nice for clarity",
      "votes": null
    },
    {
      "id": "387375",
      "postDate": "09/14/2018 20:03:59",
      "content": "<p>++</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "387384",
      "postDate": "09/14/2018 20:14:07",
      "content": "<p>@ Phil, why does it matter which box has the highest IoU? Isn't it enough that a predicted box has an IoU above the threshold with AT LEAST one ground truth box? Can you please post a flow chart or pseudo code? </p>",
      "rawMarkdown": "Phil, why does it matter which box has the highest IoU? Isn't it enough that a predicted box has an IoU above the threshold with AT LEAST one ground truth box? Can you please post a flow chart or pseudo code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 380755,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "09/03/2018 12:03:00",
      "content": "<p>Hello, Konstantin. Are you want to get theoretical example of metric computation, or you need programm implementation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 380761,
          "author_name": "maksimovka",
          "author_url": "",
          "post_date": "09/03/2018 12:18:13",
          "content": "<p>I want to get <strong>official</strong> programm implementation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 381195,
          "author_name": "aantonova",
          "author_url": "",
          "post_date": "09/04/2018 08:54:24",
          "content": "<p>I want it too!!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 381208,
          "author_name": "msalnikov",
          "author_url": "",
          "post_date": "09/04/2018 09:25:28",
          "content": "<p>+</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 387375,
          "author_name": "giuliasavorgnan",
          "author_url": "",
          "post_date": "09/14/2018 20:03:59",
          "content": "<p>++</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 381485,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "09/04/2018 17:33:42",
      "content": "<p>Thanks for your interest.  I have updated the <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge#evaluation\">Evaluation</a> page with more information.  Confidence is a general-purpose tool to resolve edge cases in submission box ordering - it is there primarily as a failsafe if any extreme edge cases arise.  In this particular data set I have not seen <code>confidence</code> change scores.</p>\n\n<p>I would recommend that you calculate a <code>confidence</code> score to ensure that, if any edge cases do arise, your submission is evaluated in order of how sure you are of your submission boxes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 381803,
          "author_name": "maksimovka",
          "author_url": "",
          "post_date": "09/05/2018 08:19:45",
          "content": "<p>Thank you for clarification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382187,
          "author_name": "xinario",
          "author_url": "",
          "post_date": "09/05/2018 20:50:01",
          "content": "<p>Hi, Phil. For negative cases, does it mean I can predict anything without receiving any penalty (since TP will always be 0)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382196,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "09/05/2018 21:08:40",
          "content": "<p>Hey xinario,</p>\n\n<p>No - if the sample is a true negative with <strong>no</strong> predictions, then it is not counted in the mean.  If <em>any</em> predictions are made, the whole sample counts as a 0 and is counted in the mean.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382197,
          "author_name": "mlandry",
          "author_url": "",
          "post_date": "09/05/2018 21:09:24",
          "content": "<p>Hi <a href=\"/xinario\">@xinario</a>. On the Evaluation page, they state:</p>\n\n<blockquote>\n  <p>Important note: if there are no ground truth objects at all for a\n  given image, ANY number of predictions (false positives) will result\n  in the image receiving a score of zero, and being included in the mean\n  average precision.</p>\n</blockquote>\n\n<p>So in the basics of numerator/denominator, if you are correct on a negative case, you add 0 to the numerator and denominator; it's not included in result--you don't gain anything, but at least you don't lose anything. If you are incorrect in the negative case you add 0 to the numerator and 1 to the denominator as if you guessed entirely the wrong box. So the negatives can only hurt, they can't help. But an important part of getting a good score is surely avoiding those False Positives.</p>\n\n<p>Hope that helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382199,
          "author_name": "xinario",
          "author_url": "",
          "post_date": "09/05/2018 21:19:57",
          "content": "<p>So that means for a single negative sample, it actually doesn't matter how many false positive predictions (&gt;0) I got. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382201,
          "author_name": "xinario",
          "author_url": "",
          "post_date": "09/05/2018 21:24:09",
          "content": "<p>Yeah, thanks for the classification. What I got so far is that we have to avoid predicting bounding boxes on negative samples. But once you do, it actually does matter if it's 1 or 2 or even more. The final score will be the same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383687,
          "author_name": "kretes",
          "author_url": "",
          "post_date": "09/09/2018 11:03:47",
          "content": "<p>I wonder what happens with the scoring if I have two predicted bounding boxes overlapping with a single ground-true object. An example (purple is grounf truth, light blue are predictions):</p>\n\n<p><a href=\"https://imgur.com/a/dy0tZbC\">example of overlapping predictions</a></p>\n\n<p>If I go with an algorithm:</p>\n\n<pre><code>for each prediction sorted by confidence descending:\n    find the best matching ground truth.\n    calculate IoU and note  TP, FP, FN for every threshold\n    remove 'processed' ground truth box from the list\n</code></pre>\n\n<p>Then I start from the lower predicted bbox, which has IoU &lt; 0.4 so I get 0 precision, even though my next bbox could have a lower confidence but higher IoU.</p>\n\n<p>Is this the algorithm used when scoring? In this case the 'confidence' is really important and it does not depend on an 'edge case' in the dataset but more on the 'edge case' in predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383718,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/09/2018 12:38:31",
          "content": "<p>Good example. Yes, in this case confidence is important. But if seems, that situation like this, when we have to prediction boxes on one lung, have small probability. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385141,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "09/10/2018 12:10:20",
          "content": "<p>Tomasz,</p>\n\n<p>IoU is also taken into account.  The predicted box with the highest IoU would be taken as the match regardless of confidence order.</p>\n\n<p>All ground truth boxes are scored against all predicted boxes.  No boxes are removed from contention.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385146,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/10/2018 12:21:29",
          "content": "<p>I am confused. So, if i have many boxes that cover the gt to some extent, only the one with the highest iou will be scored? All the rest will be ignored?\nIf you can't publish the code, can you at least post official psaudo code? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385149,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "09/10/2018 12:27:53",
          "content": "<p>The one with the highest IoU will be scored.  The rest will be marked as false positives.</p>\n\n<p>I will consider creating and posting pseudo code, thanks for the suggestion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385326,
          "author_name": "kretes",
          "author_url": "",
          "post_date": "09/10/2018 19:31:39",
          "content": "<p>Well the 'unofficial' but commonly used version is this: <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a></p>\n\n<p>and it goes with the following consequences - for any given threshold look at the predicted bbox in the order of confidence and if IoU is higher than the threshold - consider that a hit. That is mostly ok.\nBut then - this bbox cannot be used to match any other gt boxes.</p>\n\n<p>@Phil - It would be great if we can get a pseudo code. \nIt would be perfect if we can get a real implementation :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 386851,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "09/13/2018 18:33:06",
          "content": "<p>@Phil pseudo code would be nice for clarity</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 387384,
          "author_name": "giuliasavorgnan",
          "author_url": "",
          "post_date": "09/14/2018 20:14:07",
          "content": "<p>@ Phil, why does it matter which box has the highest IoU? Isn't it enough that a predicted box has an IoU above the threshold with AT LEAST one ground truth box? Can you please post a flow chart or pseudo code? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 382833,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "09/07/2018 07:37:17",
      "content": "<p>What score is used for single patient when there are no real boxes and we didn't predict any boxes? 1.0?</p>\n\n<p>I made an empty submission and it gets score = 0.0. It can be in 2 cases only: </p>\n\n<ul>\n<li>There are no empty cases in test set</li>\n<li>Scroing metric doesn't use empty boxes in test set.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 382843,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/07/2018 07:58:04",
          "content": "<p>On this case this patient don't take into account in competition metric. I mean\n we have three step of metric calculation:</p>\n\n<ol>\n<li><p>Founding amount of patient, who had no real boxes and had no predicted boxes. Remove these patient from our test set. Let's denote amount of remained patient by n</p></li>\n<li><p>For each other case we calculate average precision. So, we have n values: AP_1, AP_2, ..., AP_n</p></li>\n<li>Final score is (AP_1 + AP_2 +  AP_n) / n</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382847,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/07/2018 08:07:35",
          "content": "<p>It mean's that in your submition you have relatively small value n, and that's good for your final score cause you get little denominator. But you have AP_1 = AP_2 = ... = AP_n = 0, so that's bad for your metric, cause you get zero in numerator.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 382849,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "09/07/2018 08:09:18",
          "content": "<p>Thanks for explanation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "380739": "Hi @philculliton !\n\nCan you pls provide example of scoring method? Because now it is unclear how to use confidence score data.",
    "380755": "Hello, Konstantin. Are you want to get theoretical example of metric computation, or you need programm implementation?",
    "380761": "I want to get **official** programm implementation.",
    "381195": "I want it too!!!",
    "381208": "",
    "381485": "Thanks for your interest.  I have updated the [Evaluation][1] page with more information.  Confidence is a general-purpose tool to resolve edge cases in submission box ordering - it is there primarily as a failsafe if any extreme edge cases arise.  In this particular data set I have not seen `confidence` change scores.\n\nI would recommend that you calculate a `confidence` score to ensure that, if any edge cases do arise, your submission is evaluated in order of how sure you are of your submission boxes.\n\n  [1]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge#evaluation",
    "381803": "Thank you for clarification.",
    "382187": "Hi, Phil. For negative cases, does it mean I can predict anything without receiving any penalty (since TP will always be 0)?",
    "382196": "Hey xinario,\n\nNo - if the sample is a true negative with **no** predictions, then it is not counted in the mean.  If *any* predictions are made, the whole sample counts as a 0 and is counted in the mean.",
    "382197": "Hi @xinario. On the Evaluation page, they state:\n\n&gt; Important note: if there are no ground truth objects at all for a\n&gt; given image, ANY number of predictions (false positives) will result\n&gt; in the image receiving a score of zero, and being included in the mean\n&gt; average precision.\n\nSo in the basics of numerator/denominator, if you are correct on a negative case, you add 0 to the numerator and denominator; it's not included in result--you don't gain anything, but at least you don't lose anything. If you are incorrect in the negative case you add 0 to the numerator and 1 to the denominator as if you guessed entirely the wrong box. So the negatives can only hurt, they can't help. But an important part of getting a good score is surely avoiding those False Positives.\n\nHope that helps.",
    "382199": "So that means for a single negative sample, it actually doesn't matter how many false positive predictions (&gt;0) I got.",
    "382201": "Yeah, thanks for the classification. What I got so far is that we have to avoid predicting bounding boxes on negative samples. But once you do, it actually does matter if it's 1 or 2 or even more. The final score will be the same.",
    "382833": "What score is used for single patient when there are no real boxes and we didn't predict any boxes? 1.0?\n\nI made an empty submission and it gets score = 0.0. It can be in 2 cases only: \n\n - There are no empty cases in test set\n - Scroing metric doesn't use empty boxes in test set.",
    "382843": "On this case this patient don't take into account in competition metric. I mean\n we have three step of metric calculation:\n\n 1. Founding amount of patient, who had no real boxes and had no predicted boxes. Remove these patient from our test set. Let's denote amount of remained patient by n\n\n 2. For each other case we calculate average precision. So, we have n values: AP_1, AP_2, ..., AP_n\n 3. Final score is (AP_1 + AP_2 +  AP_n) / n",
    "382847": "It mean's that in your submition you have relatively small value n, and that's good for your final score cause you get little denominator. But you have AP_1 = AP_2 = ... = AP_n = 0, so that's bad for your metric, cause you get zero in numerator.",
    "382849": "Thanks for explanation!",
    "383687": "I wonder what happens with the scoring if I have two predicted bounding boxes overlapping with a single ground-true object. An example (purple is grounf truth, light blue are predictions):\n\n[example of overlapping predictions][1]\n\nIf I go with an algorithm:\n\n    for each prediction sorted by confidence descending:\n        find the best matching ground truth.\n        calculate IoU and note  TP, FP, FN for every threshold\n        remove 'processed' ground truth box from the list\n\nThen I start from the lower predicted bbox, which has IoU &lt; 0.4 so I get 0 precision, even though my next bbox could have a lower confidence but higher IoU.\n\nIs this the algorithm used when scoring? In this case the 'confidence' is really important and it does not depend on an 'edge case' in the dataset but more on the 'edge case' in predictions.\n\n  [1]: https://imgur.com/a/dy0tZbC",
    "383718": "Good example. Yes, in this case confidence is important. But if seems, that situation like this, when we have to prediction boxes on one lung, have small probability.",
    "385141": "Tomasz,\n\nIoU is also taken into account.  The predicted box with the highest IoU would be taken as the match regardless of confidence order.\n\nAll ground truth boxes are scored against all predicted boxes.  No boxes are removed from contention.",
    "385146": "I am confused. So, if i have many boxes that cover the gt to some extent, only the one with the highest iou will be scored? All the rest will be ignored?\nIf you can't publish the code, can you at least post official psaudo code?",
    "385149": "The one with the highest IoU will be scored.  The rest will be marked as false positives.\n\nI will consider creating and posting pseudo code, thanks for the suggestion.",
    "385326": "Well the 'unofficial' but commonly used version is this: https://www.kaggle.com/chenyc15/mean-average-precision-metric\n\nand it goes with the following consequences - for any given threshold look at the predicted bbox in the order of confidence and if IoU is higher than the threshold - consider that a hit. That is mostly ok.\nBut then - this bbox cannot be used to match any other gt boxes.\n\n\n@Phil - It would be great if we can get a pseudo code. \nIt would be perfect if we can get a real implementation :)",
    "386851": "Phil pseudo code would be nice for clarity",
    "387375": "",
    "387384": "Phil, why does it matter which box has the highest IoU? Isn't it enough that a predicted box has an IoU above the threshold with AT LEAST one ground truth box? Can you please post a flow chart or pseudo code?"
  },
  "source": "meta"
}