{
  "id": 115444,
  "title": "Confused about the metrics",
  "url": "/competitions/pku-autonomous-driving/discussion/115444",
  "author_name": "",
  "post_date": "2019-11-02T21:46:54.264814800Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello! \nIt's the first time I participate in a competition. I hope the question isn't too blunt, but I would like to double check with you if my understanding of the <strong>metrics</strong> is correct. Here it comes:</p>\n\n<p>Let's say we have an image X, where we detect three cars and have an estimate about their <em>pose</em> (yaw, pitch, roll, x, y, z). Let's call them A, B and C.\nAt the same time, there exists a solution to that image, which says, for example, we have cars A', B' and C' and their poses. Now:</p>\n\n<p>1.</p>\n\n<blockquote>\n  <p>We then take the resulting distances between all pairs of objects</p>\n</blockquote>\n\n<p>I understand we take 9 pairs then: A-A', A-B', A-C', B-A', B-B', B-C', C-A', C-B', C-C', then,\nwe calculate the distances (in accordance with the definition in C#), we have:\nd(A, A'), d(A, B'), ..., d(C, C') that are numbers, correct?</p>\n\n<p>2.</p>\n\n<blockquote>\n  <p>and determine which predicted objects are closest to solution objects, </p>\n</blockquote>\n\n<p>OK, so having 9 numbers (d( . , . )), we can sort them in ascending manner, and since the solution for that image has three cars (A', B', C') then we can pick the top three predicted objects with the lowest distances, correct?</p>\n\n<p>3.</p>\n\n<blockquote>\n  <p>and apply thresholds for both translation and rotation.</p>\n</blockquote>\n\n<p>My understanding is that there are two separate multi-level thresholds, so depending on how good we are, we can get different contribution for the final score from each object, from perfect to none.</p>\n\n<p>4.</p>\n\n<blockquote>\n  <p>Confidence scores are used to sort submission objects. </p>\n</blockquote>\n\n<p>Why would we need that? I know that <em>confidence</em> value is one of the outputs in the submission file, but didn't we actually sort the distances in step 1.? Or do you think they refer to something different here?</p>\n\n<p>5.</p>\n\n<blockquote>\n  <p>Units for rotation are radians; translation is meters.</p>\n</blockquote>\n\n<p>OK. Good.</p>\n\n<p>6.</p>\n\n<blockquote>\n  <p>If both of the distances between prediction and solution (as calculated above) are less than the threshold, then that prediction object is counted as a true positive for that threshold. If not the predicted object is counted as a false positive for that threshold.</p>\n</blockquote>\n\n<p>Yes, that is what I understood in point 3.</p>\n\n<p>7.</p>\n\n<blockquote>\n  <p>Finally, mAP is calculated using these TP/FP determinations across all thresholds.</p>\n</blockquote>\n\n<p>Yes. We need to go across the whole dataset.</p>\n\n<p>Summary:\nPlease, let me know if my general understanding is correct.\nAlso, I would appreciate if you can clarify point 4. that discusses the confidence sorting. It is very unclear to me.</p>\n\n<p>Thank you!\nPerhaps it is a simple question, but if someone benefits from this, the time wasn't wasted ;)</p>",
  "messages": [
    {
      "id": "663924",
      "postDate": "11/02/2019 21:46:54",
      "content": "<p>Hello! \nIt's the first time I participate in a competition. I hope the question isn't too blunt, but I would like to double check with you if my understanding of the <strong>metrics</strong> is correct. Here it comes:</p>\n\n<p>Let's say we have an image X, where we detect three cars and have an estimate about their <em>pose</em> (yaw, pitch, roll, x, y, z). Let's call them A, B and C.\nAt the same time, there exists a solution to that image, which says, for example, we have cars A', B' and C' and their poses. Now:</p>\n\n<p>1.</p>\n\n<blockquote>\n  <p>We then take the resulting distances between all pairs of objects</p>\n</blockquote>\n\n<p>I understand we take 9 pairs then: A-A', A-B', A-C', B-A', B-B', B-C', C-A', C-B', C-C', then,\nwe calculate the distances (in accordance with the definition in C#), we have:\nd(A, A'), d(A, B'), ..., d(C, C') that are numbers, correct?</p>\n\n<p>2.</p>\n\n<blockquote>\n  <p>and determine which predicted objects are closest to solution objects, </p>\n</blockquote>\n\n<p>OK, so having 9 numbers (d( . , . )), we can sort them in ascending manner, and since the solution for that image has three cars (A', B', C') then we can pick the top three predicted objects with the lowest distances, correct?</p>\n\n<p>3.</p>\n\n<blockquote>\n  <p>and apply thresholds for both translation and rotation.</p>\n</blockquote>\n\n<p>My understanding is that there are two separate multi-level thresholds, so depending on how good we are, we can get different contribution for the final score from each object, from perfect to none.</p>\n\n<p>4.</p>\n\n<blockquote>\n  <p>Confidence scores are used to sort submission objects. </p>\n</blockquote>\n\n<p>Why would we need that? I know that <em>confidence</em> value is one of the outputs in the submission file, but didn't we actually sort the distances in step 1.? Or do you think they refer to something different here?</p>\n\n<p>5.</p>\n\n<blockquote>\n  <p>Units for rotation are radians; translation is meters.</p>\n</blockquote>\n\n<p>OK. Good.</p>\n\n<p>6.</p>\n\n<blockquote>\n  <p>If both of the distances between prediction and solution (as calculated above) are less than the threshold, then that prediction object is counted as a true positive for that threshold. If not the predicted object is counted as a false positive for that threshold.</p>\n</blockquote>\n\n<p>Yes, that is what I understood in point 3.</p>\n\n<p>7.</p>\n\n<blockquote>\n  <p>Finally, mAP is calculated using these TP/FP determinations across all thresholds.</p>\n</blockquote>\n\n<p>Yes. We need to go across the whole dataset.</p>\n\n<p>Summary:\nPlease, let me know if my general understanding is correct.\nAlso, I would appreciate if you can clarify point 4. that discusses the confidence sorting. It is very unclear to me.</p>\n\n<p>Thank you!\nPerhaps it is a simple question, but if someone benefits from this, the time wasn't wasted ;)</p>",
      "rawMarkdown": "Hello! \nIt's the first time I participate in a competition. I hope the question isn't too blunt, but I would like to double check with you if my understanding of the **metrics** is correct. Here it comes:\n\nLet's say we have an image X, where we detect three cars and have an estimate about their *pose* (yaw, pitch, roll, x, y, z). Let's call them A, B and C.\nAt the same time, there exists a solution to that image, which says, for example, we have cars A', B' and C' and their poses. Now:\n\n1.\n&gt; We then take the resulting distances between all pairs of objects\n\nI understand we take 9 pairs then: A-A', A-B', A-C', B-A', B-B', B-C', C-A', C-B', C-C', then,\nwe calculate the distances (in accordance with the definition in C#), we have:\nd(A, A'), d(A, B'), ..., d(C, C') that are numbers, correct?\n\n2.\n&gt; and determine which predicted objects are closest to solution objects, \n\nOK, so having 9 numbers (d( . , . )), we can sort them in ascending manner, and since the solution for that image has three cars (A', B', C') then we can pick the top three predicted objects with the lowest distances, correct?\n\n3.\n&gt; and apply thresholds for both translation and rotation.\n\nMy understanding is that there are two separate multi-level thresholds, so depending on how good we are, we can get different contribution for the final score from each object, from perfect to none.\n\n4.\n&gt; Confidence scores are used to sort submission objects. \n\nWhy would we need that? I know that *confidence* value is one of the outputs in the submission file, but didn't we actually sort the distances in step 1.? Or do you think they refer to something different here?\n\n5.\n&gt; Units for rotation are radians; translation is meters.\n\nOK. Good.\n\n6.\n&gt; If both of the distances between prediction and solution (as calculated above) are less than the threshold, then that prediction object is counted as a true positive for that threshold. If not the predicted object is counted as a false positive for that threshold.\n\nYes, that is what I understood in point 3.\n\n7.\n&gt; Finally, mAP is calculated using these TP/FP determinations across all thresholds.\n\nYes. We need to go across the whole dataset.\n\nSummary:\nPlease, let me know if my general understanding is correct.\nAlso, I would appreciate if you can clarify point 4. that discusses the confidence sorting. It is very unclear to me.\n\nThank you!\nPerhaps it is a simple question, but if someone benefits from this, the time wasn't wasted ;)",
      "votes": null
    },
    {
      "id": "663930",
      "postDate": "11/02/2019 22:13:55",
      "content": "<p>My interpretation of the evaluation metrics I put into code here: <a href=\"https://www.kaggle.com/greatgamedota/cv-util-functions\">CV Util Functions</a> <br>\nNot saying my way is correct, need more feed back</p>",
      "rawMarkdown": "My interpretation of the evaluation metrics I put into code here: [CV Util Functions](https://www.kaggle.com/greatgamedota/cv-util-functions)  \nNot saying my way is correct, need more feed back",
      "votes": null
    },
    {
      "id": "663934",
      "postDate": "11/02/2019 22:24:36",
      "content": "<p>Thank you! I will take a look.\nPS. Have you accounted for this <em>confidence</em> they use for sorting? Also, the metrics enforces normalizing of  quaternions. I don't know if scipy does that by default. Anyway, thank you for sharing. I will test it on Monday and let you know ;)</p>",
      "rawMarkdown": "Thank you! I will take a look.\nPS. Have you accounted for this *confidence* they use for sorting? Also, the metrics enforces normalizing of  quaternions. I don't know if scipy does that by default. Anyway, thank you for sharing. I will test it on Monday and let you know ;)",
      "votes": null
    },
    {
      "id": "663935",
      "postDate": "11/02/2019 22:30:18",
      "content": "<p>I just think confidence is worthless, just sort what car to calc accuracy for next. <br>\nI assume skipy normalizes the quaternions because you only give euler angles which don't have a magnitude.</p>",
      "rawMarkdown": "I just think confidence is worthless, just sort what car to calc accuracy for next.  \nI assume skipy normalizes the quaternions because you only give euler angles which don't have a magnitude.",
      "votes": null
    },
    {
      "id": "663938",
      "postDate": "11/02/2019 22:42:59",
      "content": "<p>Right... they say \"Confidence scores are used to sort submission objects\", but it would be kind of stupid to use a <em>subjective</em> (aka a participant's own \"opinion\" on how confident he/she is) to determine which car is which. So I think you're correct. It is used for sorting answers, but not really grading them. \nIn the end: \"We then take the resulting distances between all pairs of objects and determine which predicted objects are closest to solution objects\". It makes total sense to use the distance...  I think.</p>",
      "rawMarkdown": "Right... they say \"Confidence scores are used to sort submission objects\", but it would be kind of stupid to use a *subjective* (aka a participant's own \"opinion\" on how confident he/she is) to determine which car is which. So I think you're correct. It is used for sorting answers, but not really grading them. \nIn the end: \"We then take the resulting distances between all pairs of objects and determine which predicted objects are closest to solution objects\". It makes total sense to use the distance...  I think.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 663930,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "11/02/2019 22:13:55",
      "content": "<p>My interpretation of the evaluation metrics I put into code here: <a href=\"https://www.kaggle.com/greatgamedota/cv-util-functions\">CV Util Functions</a> <br>\nNot saying my way is correct, need more feed back</p>",
      "votes": null,
      "replies": [
        {
          "id": 663934,
          "author_name": "olegzero13",
          "author_url": "",
          "post_date": "11/02/2019 22:24:36",
          "content": "<p>Thank you! I will take a look.\nPS. Have you accounted for this <em>confidence</em> they use for sorting? Also, the metrics enforces normalizing of  quaternions. I don't know if scipy does that by default. Anyway, thank you for sharing. I will test it on Monday and let you know ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663935,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "11/02/2019 22:30:18",
          "content": "<p>I just think confidence is worthless, just sort what car to calc accuracy for next. <br>\nI assume skipy normalizes the quaternions because you only give euler angles which don't have a magnitude.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663938,
          "author_name": "olegzero13",
          "author_url": "",
          "post_date": "11/02/2019 22:42:59",
          "content": "<p>Right... they say \"Confidence scores are used to sort submission objects\", but it would be kind of stupid to use a <em>subjective</em> (aka a participant's own \"opinion\" on how confident he/she is) to determine which car is which. So I think you're correct. It is used for sorting answers, but not really grading them. \nIn the end: \"We then take the resulting distances between all pairs of objects and determine which predicted objects are closest to solution objects\". It makes total sense to use the distance...  I think.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "663924": "Hello! \nIt's the first time I participate in a competition. I hope the question isn't too blunt, but I would like to double check with you if my understanding of the **metrics** is correct. Here it comes:\n\nLet's say we have an image X, where we detect three cars and have an estimate about their *pose* (yaw, pitch, roll, x, y, z). Let's call them A, B and C.\nAt the same time, there exists a solution to that image, which says, for example, we have cars A', B' and C' and their poses. Now:\n\n1.\n&gt; We then take the resulting distances between all pairs of objects\n\nI understand we take 9 pairs then: A-A', A-B', A-C', B-A', B-B', B-C', C-A', C-B', C-C', then,\nwe calculate the distances (in accordance with the definition in C#), we have:\nd(A, A'), d(A, B'), ..., d(C, C') that are numbers, correct?\n\n2.\n&gt; and determine which predicted objects are closest to solution objects, \n\nOK, so having 9 numbers (d( . , . )), we can sort them in ascending manner, and since the solution for that image has three cars (A', B', C') then we can pick the top three predicted objects with the lowest distances, correct?\n\n3.\n&gt; and apply thresholds for both translation and rotation.\n\nMy understanding is that there are two separate multi-level thresholds, so depending on how good we are, we can get different contribution for the final score from each object, from perfect to none.\n\n4.\n&gt; Confidence scores are used to sort submission objects. \n\nWhy would we need that? I know that *confidence* value is one of the outputs in the submission file, but didn't we actually sort the distances in step 1.? Or do you think they refer to something different here?\n\n5.\n&gt; Units for rotation are radians; translation is meters.\n\nOK. Good.\n\n6.\n&gt; If both of the distances between prediction and solution (as calculated above) are less than the threshold, then that prediction object is counted as a true positive for that threshold. If not the predicted object is counted as a false positive for that threshold.\n\nYes, that is what I understood in point 3.\n\n7.\n&gt; Finally, mAP is calculated using these TP/FP determinations across all thresholds.\n\nYes. We need to go across the whole dataset.\n\nSummary:\nPlease, let me know if my general understanding is correct.\nAlso, I would appreciate if you can clarify point 4. that discusses the confidence sorting. It is very unclear to me.\n\nThank you!\nPerhaps it is a simple question, but if someone benefits from this, the time wasn't wasted ;)",
    "663930": "My interpretation of the evaluation metrics I put into code here: [CV Util Functions](https://www.kaggle.com/greatgamedota/cv-util-functions)  \nNot saying my way is correct, need more feed back",
    "663934": "Thank you! I will take a look.\nPS. Have you accounted for this *confidence* they use for sorting? Also, the metrics enforces normalizing of  quaternions. I don't know if scipy does that by default. Anyway, thank you for sharing. I will test it on Monday and let you know ;)",
    "663935": "I just think confidence is worthless, just sort what car to calc accuracy for next.  \nI assume skipy normalizes the quaternions because you only give euler angles which don't have a magnitude.",
    "663938": "Right... they say \"Confidence scores are used to sort submission objects\", but it would be kind of stupid to use a *subjective* (aka a participant's own \"opinion\" on how confident he/she is) to determine which car is which. So I think you're correct. It is used for sorting answers, but not really grading them. \nIn the end: \"We then take the resulting distances between all pairs of objects and determine which predicted objects are closest to solution objects\". It makes total sense to use the distance...  I think."
  },
  "source": "meta"
}