{
  "id": 119775,
  "title": "Can someone explain the evaluation metric?",
  "url": "/competitions/pku-autonomous-driving/discussion/119775",
  "author_name": "Rafael Natan",
  "post_date": "2019-12-01T11:06:51.646000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am aware that the evaluation metric currently does not take into account false negatives when less cars are predicted than in the image and that that will be fixed.</p>\n\n<p>However, I still don't understand how the evaluation workflow goes.</p>\n\n<p>Can someone please walk me through an example of the current evaluation, please? Or at least point me to a good resource that explains this?</p>\n\n<p>As I currently understand it (in order): \n1) Translational distances are calculated between all car pairs (between prediction and ground truth labels).\n2) Each prediction is assigned to the closest ground ground truth (solution) car.\n3) All car predictions associated with a given ground truth (solution) car are sorted by confidence, among all other predictions associated with that same given ground truth (solution) car.\n4) For each threshold, all predicted cars are labeled True positive if they are within threshold and False positive if they are beyond the thresholds (translational and orientation).\n5) For each threshold, precision and recall are calculated for all predictions, where the predicted cars are like \"retrieved documents\" (labeled false or true positive at step 4), and ground truth (solution) cars are like \"relevant documents\".\n6) Then the precision and recall values of each threshold are used to calculate Average Precision, for the image.\n7) And then the mean of all images' average precision is taken to give the MAP of the whole submission.</p>\n\n<ul>\n<li>I know I'm wrong somewhere because the sorting based on the confidence (step 3).</li>\n</ul>\n\n<p>Any help is appreciated.\nThank you.</p>",
  "messages": [
    {
      "id": 685305,
      "postDate": "2019-12-01T11:06:51.647Z",
      "content": "<p>I am aware that the evaluation metric currently does not take into account false negatives when less cars are predicted than in the image and that that will be fixed.</p>\n\n<p>However, I still don't understand how the evaluation workflow goes.</p>\n\n<p>Can someone please walk me through an example of the current evaluation, please? Or at least point me to a good resource that explains this?</p>\n\n<p>As I currently understand it (in order): \n1) Translational distances are calculated between all car pairs (between prediction and ground truth labels).\n2) Each prediction is assigned to the closest ground ground truth (solution) car.\n3) All car predictions associated with a given ground truth (solution) car are sorted by confidence, among all other predictions associated with that same given ground truth (solution) car.\n4) For each threshold, all predicted cars are labeled True positive if they are within threshold and False positive if they are beyond the thresholds (translational and orientation).\n5) For each threshold, precision and recall are calculated for all predictions, where the predicted cars are like \"retrieved documents\" (labeled false or true positive at step 4), and ground truth (solution) cars are like \"relevant documents\".\n6) Then the precision and recall values of each threshold are used to calculate Average Precision, for the image.\n7) And then the mean of all images' average precision is taken to give the MAP of the whole submission.</p>\n\n<ul>\n<li>I know I'm wrong somewhere because the sorting based on the confidence (step 3).</li>\n</ul>\n\n<p>Any help is appreciated.\nThank you.</p>",
      "rawMarkdown": "I am aware that the evaluation metric currently does not take into account false negatives when less cars are predicted than in the image and that that will be fixed.\n\nHowever, I still don't understand how the evaluation workflow goes.\n\nCan someone please walk me through an example of the current evaluation, please? Or at least point me to a good resource that explains this?\n\nAs I currently understand it (in order): \n1) Translational distances are calculated between all car pairs (between prediction and ground truth labels).\n2) Each prediction is assigned to the closest ground ground truth (solution) car.\n3) All car predictions associated with a given ground truth (solution) car are sorted by confidence, among all other predictions associated with that same given ground truth (solution) car.\n4) For each threshold, all predicted cars are labeled True positive if they are within threshold and False positive if they are beyond the thresholds (translational and orientation).\n5) For each threshold, precision and recall are calculated for all predictions, where the predicted cars are like \"retrieved documents\" (labeled false or true positive at step 4), and ground truth (solution) cars are like \"relevant documents\".\n6) Then the precision and recall values of each threshold are used to calculate Average Precision, for the image.\n7) And then the mean of all images' average precision is taken to give the MAP of the whole submission.\n\n* I know I'm wrong somewhere because the sorting based on the confidence (step 3).\n\nAny help is appreciated.\nThank you.",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "685305": "I am aware that the evaluation metric currently does not take into account false negatives when less cars are predicted than in the image and that that will be fixed.\n\nHowever, I still don't understand how the evaluation workflow goes.\n\nCan someone please walk me through an example of the current evaluation, please? Or at least point me to a good resource that explains this?\n\nAs I currently understand it (in order): \n1) Translational distances are calculated between all car pairs (between prediction and ground truth labels).\n2) Each prediction is assigned to the closest ground ground truth (solution) car.\n3) All car predictions associated with a given ground truth (solution) car are sorted by confidence, among all other predictions associated with that same given ground truth (solution) car.\n4) For each threshold, all predicted cars are labeled True positive if they are within threshold and False positive if they are beyond the thresholds (translational and orientation).\n5) For each threshold, precision and recall are calculated for all predictions, where the predicted cars are like \"retrieved documents\" (labeled false or true positive at step 4), and ground truth (solution) cars are like \"relevant documents\".\n6) Then the precision and recall values of each threshold are used to calculate Average Precision, for the image.\n7) And then the mean of all images' average precision is taken to give the MAP of the whole submission.\n\n* I know I'm wrong somewhere because the sorting based on the confidence (step 3).\n\nAny help is appreciated.\nThank you."
  }
}