{
  "id": 62365,
  "title": "Metric clarification",
  "url": "/competitions/airbus-ship-detection/discussion/62365",
  "author_name": "Evgeny Nizhibitsky",
  "post_date": "2018-07-31T17:50:42.461000",
  "votes": 13,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Is it true that in this contest unlike the salt one each ship is a different object and thus TP(t)/FP(t)/FN(t) are referring to ship instances and not pixels? Or what's the purpose of the IoU introduction otherwise? Both in this and salt competitions there is no clarification on the \"objects\" from Evaluation page, so the question arises the second time.</p>\n\n<p>P.S. +1 for publishing metrics code can be useful and self-claryfying like it happens in some hidden place starting from \"Driven\" and ending with \"Data\".</p>\n\n<p>/cc @inversion</p>",
  "messages": [
    {
      "id": 364533,
      "postDate": "2018-07-31T17:50:42.460Z",
      "content": "<p>Is it true that in this contest unlike the salt one each ship is a different object and thus TP(t)/FP(t)/FN(t) are referring to ship instances and not pixels? Or what's the purpose of the IoU introduction otherwise? Both in this and salt competitions there is no clarification on the \"objects\" from Evaluation page, so the question arises the second time.</p>\n\n<p>P.S. +1 for publishing metrics code can be useful and self-claryfying like it happens in some hidden place starting from \"Driven\" and ending with \"Data\".</p>\n\n<p>/cc @inversion</p>",
      "rawMarkdown": "Is it true that in this contest unlike the salt one each ship is a different object and thus TP(t)/FP(t)/FN(t) are referring to ship instances and not pixels? Or what's the purpose of the IoU introduction otherwise? Both in this and salt competitions there is no clarification on the \"objects\" from Evaluation page, so the question arises the second time.\n\nP.S. +1 for publishing metrics code can be useful and self-claryfying like it happens in some hidden place starting from \"Driven\" and ending with \"Data\".\n\n/cc @inversion",
      "votes": 12
    },
    {
      "id": 365038,
      "postDate": "2018-08-01T19:07:48.470Z",
      "content": "<p>From the description it seems that the submission file is asking for the pixel mask for each ship that you predict in an image.  I take it that if there is more than one ship in an image the submission file will have multiple lines for the same ImageID.    At the top of the description they state  \"The IoU of a proposed set of object pixels and a set of true object pixels\"  I think that the IoU has to be at the pixel level.   That's what it seems to read.</p>",
      "rawMarkdown": "From the description it seems that the submission file is asking for the pixel mask for each ship that you predict in an image.  I take it that if there is more than one ship in an image the submission file will have multiple lines for the same ImageID.    At the top of the description they state  \"The IoU of a proposed set of object pixels and a set of true object pixels\"  I think that the IoU has to be at the pixel level.   That's what it seems to read.",
      "votes": 1
    },
    {
      "id": 366681,
      "postDate": "2018-08-06T07:42:41.267Z",
      "content": "<p>My interpretation is indeed also that for every image they look at the number of lines in submission file and every line is a instance of a ship. Then they compare every submitted ship instance to every target ship instance (all per image) and determine the total score.</p>\n\n<p>So I guess the code could look something like this (pseudo code):</p>\n\n<pre><code>def score(pred_masks, target_masks):\n    f2_sum=0\n    thresholds = range(0.5, 1, 0.05)\n\n    for threshold in thresholds:\n        false_neg = len(target_masks)\n        false_pos = len(pred_masks)\n        true_pos = 0\n\n        for p in pred_masks:\n            for t in target_masks:\n                score = calculate_IoU(t,p)\n                if score &amp;gt; threshold:\n                    true_pos += 1\n                    false_pos -= 1 \n                    false_neg -= 1\n\n        f2_sum += calculate_f2(false_neg, false_pos, true_pos)          \n\n    return f2_sum/len(thresholds)\n</code></pre>\n\n<p>It is a might be a bit simplified and for sure not the most compute efficient, but hopefully it clarifies the logic.</p>",
      "rawMarkdown": "My interpretation is indeed also that for every image they look at the number of lines in submission file and every line is a instance of a ship. Then they compare every submitted ship instance to every target ship instance (all per image) and determine the total score.\n\nSo I guess the code could look something like this (pseudo code):\n\n\n\n    def score(pred_masks, target_masks):\n        f2_sum=0\n        thresholds = range(0.5, 1, 0.05)\n\n        for threshold in thresholds:\n            false_neg = len(target_masks)\n            false_pos = len(pred_masks)\n            true_pos = 0\n\n            for p in pred_masks:\n                for t in target_masks:\n                    score = calculate_IoU(t,p)\n                    if score &gt; threshold:\n                        true_pos += 1\n                        false_pos -= 1 \n                        false_neg -= 1\n\n            f2_sum += calculate_f2(false_neg, false_pos, true_pos)\t\t\t\n\n        return f2_sum/len(thresholds)\n\n\nIt is a might be a bit simplified and for sure not the most compute efficient, but hopefully it clarifies the logic.",
      "replies": [
        {
          "id": 366965,
          "postDate": "2018-08-06T20:00:47.857Z",
          "content": "<p>Actually the thresholds start at 0.5, and masks don't overlap, so at any given threshold every ground truth ship mask can only \"hit\" with one submitted ship mask.</p>",
          "rawMarkdown": "Actually the thresholds start at 0.5, and masks don't overlap, so at any given threshold every ground truth ship mask can only \"hit\" with one submitted ship mask."
        },
        {
          "id": 367087,
          "postDate": "2018-08-07T04:06:13.073Z",
          "content": "<p>Ok this discussion has confused me further. Let us assume for a moment that you get the predicted ship paired up correctly with the ground truth ship.  Then does it matter what the threshold is set to?  If there is only one predicted mask per ship and only one GT mask per ship, the IoU (and the F2 score) will be the same in all cases.  </p>",
          "rawMarkdown": "Ok this discussion has confused me further. Let us assume for a moment that you get the predicted ship paired up correctly with the ground truth ship.  Then does it matter what the threshold is set to?  If there is only one predicted mask per ship and only one GT mask per ship, the IoU (and the F2 score) will be the same in all cases.  "
        },
        {
          "id": 367208,
          "postDate": "2018-08-07T10:13:37.227Z",
          "content": "<p>Wether a predicted ship is paired correctly to a gt ship or not depends on the threshold. For example if your predicted ship has IoU=0.83 with gt ship, it will be counted as True Positive for all thresholds in [0.5, ... , 0.8], but for thresholds in [0.85, 0.9, 0.95] you will get one False Positive (for your predicted ship that does not match to any gt ship) and one False Negative (for the gt ship that does not have any predicted ship paired to it). Hope it helps</p>",
          "rawMarkdown": "Wether a predicted ship is paired correctly to a gt ship or not depends on the threshold. For example if your predicted ship has IoU=0.83 with gt ship, it will be counted as True Positive for all thresholds in [0.5, ... , 0.8], but for thresholds in [0.85, 0.9, 0.95] you will get one False Positive (for your predicted ship that does not match to any gt ship) and one False Negative (for the gt ship that does not have any predicted ship paired to it). Hope it helps",
          "votes": 1
        }
      ]
    },
    {
      "id": 366179,
      "postDate": "2018-08-04T08:00:16.110Z",
      "content": "<p>The submission page says that submission lines must be sorted.\nIf it's the way to match ships from one image (submission vs GT), than it could lead to problems, as errors in ship masks can lead to different ordering in submission.</p>",
      "rawMarkdown": "The submission page says that submission lines must be sorted.\nIf it's the way to match ships from one image (submission vs GT), than it could lead to problems, as errors in ship masks can lead to different ordering in submission.",
      "replies": [
        {
          "id": 366197,
          "postDate": "2018-08-04T10:02:51.787Z",
          "content": "<p>I think they meant the rle pairs needed to be sorted, but the boat masks themselves can be listed in any order in the submision.csv.</p>",
          "rawMarkdown": "I think they meant the rle pairs needed to be sorted, but the boat masks themselves can be listed in any order in the submision.csv."
        },
        {
          "id": 366630,
          "postDate": "2018-08-06T04:57:31.047Z",
          "content": "<p>Then how are they supposed to match the ship masks for same image for calculating IoU? If you mismatch them, intersection will be zero everywhere.</p>",
          "rawMarkdown": "Then how are they supposed to match the ship masks for same image for calculating IoU? If you mismatch them, intersection will be zero everywhere.",
          "votes": 1
        },
        {
          "id": 367085,
          "postDate": "2018-08-07T03:53:46.287Z",
          "content": "<p>Ok, right. Now I understand. This is a good question.</p>",
          "rawMarkdown": "Ok, right. Now I understand. This is a good question."
        },
        {
          "id": 371716,
          "postDate": "2018-08-17T13:58:19.423Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 371516,
      "postDate": "2018-08-17T01:30:09.247Z",
      "content": "<p>Can one of the organizers comment here?  There still seems to be confusion about how TP/FN/FP are computed on the Kaggle backend when there are multiple ships in an image.</p>\n\n<ul>\n<li>How are predictions and ground-truths matched to each other when there are multiple ground-truths and/or multiple predictions for the same image?</li>\n<li>Can one prediction 'hit' multiple ground-truths?</li>\n</ul>",
      "rawMarkdown": "Can one of the organizers comment here?  There still seems to be confusion about how TP/FN/FP are computed on the Kaggle backend when there are multiple ships in an image.\n\n- How are predictions and ground-truths matched to each other when there are multiple ground-truths and/or multiple predictions for the same image?\n- Can one prediction 'hit' multiple ground-truths?\n",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 371919,
          "postDate": "2018-08-17T20:23:12.740Z",
          "content": "<p>The final scoring is a little different but the evaluation section is very similar to: <a href=\"https://www.kaggle.com/c/data-science-bowl-2018#evaluation\">https://www.kaggle.com/c/data-science-bowl-2018#evaluation</a></p>\n\n<p>Someone kindly wrote up what I think is a pretty good description of the evaluation metric for that competition here:</p>\n\n<p><a href=\"https://www.kaggle.com/stkbailey/step-by-step-explanation-of-scoring-metric\">https://www.kaggle.com/stkbailey/step-by-step-explanation-of-scoring-metric</a></p>",
          "rawMarkdown": "The final scoring is a little different but the evaluation section is very similar to: https://www.kaggle.com/c/data-science-bowl-2018#evaluation\n\nSomeone kindly wrote up what I think is a pretty good description of the evaluation metric for that competition here:\n\nhttps://www.kaggle.com/stkbailey/step-by-step-explanation-of-scoring-metric"
        },
        {
          "id": 373167,
          "postDate": "2018-08-21T01:41:45.517Z",
          "content": "<p>Thanks for the link Paul, the description is well-written, but nonetheless it's a user's interpretation of the evaluation. Without validation by one of the moderators it's not clear that this interpretation matches what Kaggle computes on the backend.</p>",
          "rawMarkdown": "Thanks for the link Paul, the description is well-written, but nonetheless it's a user's interpretation of the evaluation. Without validation by one of the moderators it's not clear that this interpretation matches what Kaggle computes on the backend.",
          "votes": 3,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 365038,
      "author_name": "Gerry Traicoff",
      "author_url": "",
      "post_date": "2018-08-01T19:07:48.470000",
      "content": "<p>From the description it seems that the submission file is asking for the pixel mask for each ship that you predict in an image.  I take it that if there is more than one ship in an image the submission file will have multiple lines for the same ImageID.    At the top of the description they state  \"The IoU of a proposed set of object pixels and a set of true object pixels\"  I think that the IoU has to be at the pixel level.   That's what it seems to read.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 366681,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2018-08-06T07:42:41.267000",
      "content": "<p>My interpretation is indeed also that for every image they look at the number of lines in submission file and every line is a instance of a ship. Then they compare every submitted ship instance to every target ship instance (all per image) and determine the total score.</p>\n\n<p>So I guess the code could look something like this (pseudo code):</p>\n\n<pre><code>def score(pred_masks, target_masks):\n    f2_sum=0\n    thresholds = range(0.5, 1, 0.05)\n\n    for threshold in thresholds:\n        false_neg = len(target_masks)\n        false_pos = len(pred_masks)\n        true_pos = 0\n\n        for p in pred_masks:\n            for t in target_masks:\n                score = calculate_IoU(t,p)\n                if score &amp;gt; threshold:\n                    true_pos += 1\n                    false_pos -= 1 \n                    false_neg -= 1\n\n        f2_sum += calculate_f2(false_neg, false_pos, true_pos)          \n\n    return f2_sum/len(thresholds)\n</code></pre>\n\n<p>It is a might be a bit simplified and for sure not the most compute efficient, but hopefully it clarifies the logic.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 366965,
          "author_name": "Mark Ayzenshtadt",
          "author_url": "",
          "post_date": "2018-08-06T20:00:47.857000",
          "content": "<p>Actually the thresholds start at 0.5, and masks don't overlap, so at any given threshold every ground truth ship mask can only \"hit\" with one submitted ship mask.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367087,
          "author_name": "WillieMaddox",
          "author_url": "",
          "post_date": "2018-08-07T04:06:13.073000",
          "content": "<p>Ok this discussion has confused me further. Let us assume for a moment that you get the predicted ship paired up correctly with the ground truth ship.  Then does it matter what the threshold is set to?  If there is only one predicted mask per ship and only one GT mask per ship, the IoU (and the F2 score) will be the same in all cases.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367208,
          "author_name": "dzem90",
          "author_url": "",
          "post_date": "2018-08-07T10:13:37.227000",
          "content": "<p>Wether a predicted ship is paired correctly to a gt ship or not depends on the threshold. For example if your predicted ship has IoU=0.83 with gt ship, it will be counted as True Positive for all thresholds in [0.5, ... , 0.8], but for thresholds in [0.85, 0.9, 0.95] you will get one False Positive (for your predicted ship that does not match to any gt ship) and one False Negative (for the gt ship that does not have any predicted ship paired to it). Hope it helps</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 366179,
      "author_name": "Mark Ayzenshtadt",
      "author_url": "",
      "post_date": "2018-08-04T08:00:16.110000",
      "content": "<p>The submission page says that submission lines must be sorted.\nIf it's the way to match ships from one image (submission vs GT), than it could lead to problems, as errors in ship masks can lead to different ordering in submission.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 366197,
          "author_name": "WillieMaddox",
          "author_url": "",
          "post_date": "2018-08-04T10:02:51.787000",
          "content": "<p>I think they meant the rle pairs needed to be sorted, but the boat masks themselves can be listed in any order in the submision.csv.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366630,
          "author_name": "Mark Ayzenshtadt",
          "author_url": "",
          "post_date": "2018-08-06T04:57:31.047000",
          "content": "<p>Then how are they supposed to match the ship masks for same image for calculating IoU? If you mismatch them, intersection will be zero everywhere.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367085,
          "author_name": "WillieMaddox",
          "author_url": "",
          "post_date": "2018-08-07T03:53:46.287000",
          "content": "<p>Ok, right. Now I understand. This is a good question.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 371716,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-17T13:58:19.423000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 371516,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-17T01:30:09.247000",
      "content": "<p>Can one of the organizers comment here?  There still seems to be confusion about how TP/FN/FP are computed on the Kaggle backend when there are multiple ships in an image.</p>\n\n<ul>\n<li>How are predictions and ground-truths matched to each other when there are multiple ground-truths and/or multiple predictions for the same image?</li>\n<li>Can one prediction 'hit' multiple ground-truths?</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 371919,
          "author_name": "Paul Johnson",
          "author_url": "",
          "post_date": "2018-08-17T20:23:12.740000",
          "content": "<p>The final scoring is a little different but the evaluation section is very similar to: <a href=\"https://www.kaggle.com/c/data-science-bowl-2018#evaluation\">https://www.kaggle.com/c/data-science-bowl-2018#evaluation</a></p>\n\n<p>Someone kindly wrote up what I think is a pretty good description of the evaluation metric for that competition here:</p>\n\n<p><a href=\"https://www.kaggle.com/stkbailey/step-by-step-explanation-of-scoring-metric\">https://www.kaggle.com/stkbailey/step-by-step-explanation-of-scoring-metric</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 373167,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-21T01:41:45.517000",
          "content": "<p>Thanks for the link Paul, the description is well-written, but nonetheless it's a user's interpretation of the evaluation. Without validation by one of the moderators it's not clear that this interpretation matches what Kaggle computes on the backend.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "364533": "Is it true that in this contest unlike the salt one each ship is a different object and thus TP(t)/FP(t)/FN(t) are referring to ship instances and not pixels? Or what's the purpose of the IoU introduction otherwise? Both in this and salt competitions there is no clarification on the \"objects\" from Evaluation page, so the question arises the second time.\n\nP.S. +1 for publishing metrics code can be useful and self-claryfying like it happens in some hidden place starting from \"Driven\" and ending with \"Data\".\n\n/cc @inversion",
    "365038": "From the description it seems that the submission file is asking for the pixel mask for each ship that you predict in an image.  I take it that if there is more than one ship in an image the submission file will have multiple lines for the same ImageID.    At the top of the description they state  \"The IoU of a proposed set of object pixels and a set of true object pixels\"  I think that the IoU has to be at the pixel level.   That's what it seems to read.",
    "366681": "My interpretation is indeed also that for every image they look at the number of lines in submission file and every line is a instance of a ship. Then they compare every submitted ship instance to every target ship instance (all per image) and determine the total score.\n\nSo I guess the code could look something like this (pseudo code):\n\n\n\n    def score(pred_masks, target_masks):\n        f2_sum=0\n        thresholds = range(0.5, 1, 0.05)\n\n        for threshold in thresholds:\n            false_neg = len(target_masks)\n            false_pos = len(pred_masks)\n            true_pos = 0\n\n            for p in pred_masks:\n                for t in target_masks:\n                    score = calculate_IoU(t,p)\n                    if score &gt; threshold:\n                        true_pos += 1\n                        false_pos -= 1 \n                        false_neg -= 1\n\n            f2_sum += calculate_f2(false_neg, false_pos, true_pos)\t\t\t\n\n        return f2_sum/len(thresholds)\n\n\nIt is a might be a bit simplified and for sure not the most compute efficient, but hopefully it clarifies the logic.",
    "366179": "The submission page says that submission lines must be sorted.\nIf it's the way to match ships from one image (submission vs GT), than it could lead to problems, as errors in ship masks can lead to different ordering in submission.",
    "371516": "Can one of the organizers comment here?  There still seems to be confusion about how TP/FN/FP are computed on the Kaggle backend when there are multiple ships in an image.\n\n- How are predictions and ground-truths matched to each other when there are multiple ground-truths and/or multiple predictions for the same image?\n- Can one prediction 'hit' multiple ground-truths?\n"
  }
}