{
  "id": 109161,
  "title": "Expectation for final LB score",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/109161",
  "author_name": "",
  "post_date": "2019-09-17T03:42:40.480395800Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>We can see on KITTI benchmark, the current top evaluation score is around 0.88 for easy and 0.76  for hard, considering car IoU 0.7 and pedestrian/cyclist IoU 0.5. </p>\n\n<p>What's the expectation here for the LB score in 2 months?</p>",
  "messages": [
    {
      "id": "628266",
      "postDate": "09/17/2019 03:42:40",
      "content": "<p>We can see on KITTI benchmark, the current top evaluation score is around 0.88 for easy and 0.76  for hard, considering car IoU 0.7 and pedestrian/cyclist IoU 0.5. </p>\n\n<p>What's the expectation here for the LB score in 2 months?</p>",
      "rawMarkdown": "We can see on KITTI benchmark, the current top evaluation score is around 0.88 for easy and 0.76  for hard, considering car IoU 0.7 and pedestrian/cyclist IoU 0.5. \n\nWhat's the expectation here for the LB score in 2 months?",
      "votes": null
    },
    {
      "id": "641065",
      "postDate": "10/04/2019 12:19:42",
      "content": "<p>KITTI/nuScenes challenge benchmark use distance thresholds for mAP calculation, but this challenge uses IoU thresholds (that too from 0.5 to 0.95) which makes this even more challenging.</p>\n\n<p>Just for comparison, here are my stats for my model's predictions evaluated on both the metrics:</p>\n\n<p>nuScenes' mAP (from <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/eval/detection/evaluate.py\">here</a>):\n<code>\ncar: 0.9187143408655809,\npedestrian: 0.25615435803934755,\nother_vehicle: 0.7742371803950503,\nbus: 0.35341638450263824,\nmotorcycle: 0.20038938394575653,\ntruck: 0.549069458817806,\nbicycle: 0.3405492518182246\n</code>\nlyft's mAP (from <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\">here</a>):\n<code>\ncar: 0.26481168914758046\npedestrian: 0.01484931969844282\nother_vehicle: 0.19598733527638573\nbus: 0.08244517085140785\nmotorcycle: 0.00706896551724138\ntruck: 0.11083860627446822\nbicycle: 0.026110721574923158\n</code>\n[don't have stats for animal and emergency due to my bad validation set, it didn't have any objects from these categories. lol]</p>\n\n<p>As you can see, there's huge difference in both the mAPs. \nWe also have imbalance in class distribution, here's category vs total object count for the lyft dataset:</p>\n\n<p>category, n= total object count \n<code>\nanimal    n=  186\nbicycle    n=20928\nbus     n= 8729\ncar    n=534911\nemergency_vehicle   n= 132\nmotorcycle      n=  818\nother_vehicle   n=33376\npedestrian  n=24935\ntruck   n=14164\n</code>\nAs the lyft's mAP is mean of APs of all the classes, animal and emergency are gonna pull down the overall score. \nSo, It's gonna be tough to score high scores but not impossible :)</p>",
      "rawMarkdown": "KITTI/nuScenes challenge benchmark use distance thresholds for mAP calculation, but this challenge uses IoU thresholds (that too from 0.5 to 0.95) which makes this even more challenging.\n\nJust for comparison, here are my stats for my model's predictions evaluated on both the metrics:\n\nnuScenes' mAP (from [here](https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/eval/detection/evaluate.py)):\n```\ncar: 0.9187143408655809,\npedestrian: 0.25615435803934755,\nother_vehicle: 0.7742371803950503,\nbus: 0.35341638450263824,\nmotorcycle: 0.20038938394575653,\ntruck: 0.549069458817806,\nbicycle: 0.3405492518182246\n```\nlyft's mAP (from [here](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py)):\n```\ncar: 0.26481168914758046\npedestrian: 0.01484931969844282\nother_vehicle: 0.19598733527638573\nbus: 0.08244517085140785\nmotorcycle: 0.00706896551724138\ntruck: 0.11083860627446822\nbicycle: 0.026110721574923158\n```\n[don't have stats for animal and emergency due to my bad validation set, it didn't have any objects from these categories. lol]\n\nAs you can see, there's huge difference in both the mAPs. \nWe also have imbalance in class distribution, here's category vs total object count for the lyft dataset:\n\ncategory, n= total object count \n```\nanimal    n=  186\nbicycle    n=20928\nbus     n= 8729\ncar    n=534911\nemergency_vehicle   n= 132\nmotorcycle      n=  818\nother_vehicle   n=33376\npedestrian  n=24935\ntruck   n=14164\n```\nAs the lyft's mAP is mean of APs of all the classes, animal and emergency are gonna pull down the overall score. \nSo, It's gonna be tough to score high scores but not impossible :)",
      "votes": null
    },
    {
      "id": "641136",
      "postDate": "10/04/2019 13:03:31",
      "content": "<p>I think you are mistaken here, the scoring here is not averaging between classes.</p>\n\n<p>&gt; At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object <strong>to all ground truth objects</strong></p>\n\n<p>After that, you are right, it will be an average of all thresholds\n&gt; The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold</p>\n\n<p>and</p>\n\n<blockquote>\n  <p>Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset</p>\n</blockquote>\n\n<p>So, even a model only predicting cars should do OK.</p>",
      "rawMarkdown": "I think you are mistaken here, the scoring here is not averaging between classes.\n\n&gt; At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object **to all ground truth objects**\n\nAfter that, you are right, it will be an average of all thresholds\n&gt; The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold\n\nand\n\n&gt; Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset\n\nSo, even a model only predicting cars should do OK.",
      "votes": null
    },
    {
      "id": "641145",
      "postDate": "10/04/2019 13:08:57",
      "content": "<p>I am at <code>Average per sample mAP (IoU= 0.5) =  0.0940</code>\nnot even looking at the harder IoUs, yet ...</p>\n\n<p>I believe we can end somewhere in the 0.2 area. This challenge has some really hard to predict cases and a metric that includes IoU thresholds of e.g. 0.95 is somewhat insane ;) </p>",
      "rawMarkdown": "I am at `Average per sample mAP (IoU= 0.5) =  0.0940`\nnot even looking at the harder IoUs, yet ...\n\n\nI believe we can end somewhere in the 0.2 area. This challenge has some really hard to predict cases and a metric that includes IoU thresholds of e.g. 0.95 is somewhat insane ;)",
      "votes": null
    },
    {
      "id": "641147",
      "postDate": "10/04/2019 13:11:24",
      "content": "<p>Hey, actually the evaluation page of this competition is a little shady. </p>\n\n<p>&gt;At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object to all ground truth objects</p>\n\n<p>Before doing this, they group predictions by classes (<a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L361\">here</a>) .Once they have class wise average precisions, then they take the mean (<a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L411\">here</a>)</p>",
      "rawMarkdown": "Hey, actually the evaluation page of this competition is a little shady. \n\n&gt;At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object to all ground truth objects\n\nBefore doing this, they group predictions by classes ([here](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L361)) .Once they have class wise average precisions, then they take the mean ([here](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L411))",
      "votes": null
    },
    {
      "id": "641382",
      "postDate": "10/04/2019 16:30:20",
      "content": "<p>I have seen that commit. But I am not convinced that this will be the metric, that they use here for evaluation. \nSome clarification would be greatly appreciated.</p>",
      "rawMarkdown": "I have seen that commit. But I am not convinced that this will be the metric, that they use here for evaluation. \nSome clarification would be greatly appreciated.",
      "votes": null
    },
    {
      "id": "641389",
      "postDate": "10/04/2019 16:38:21",
      "content": "<p>see <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175\">this</a></p>",
      "rawMarkdown": "see [this](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175)",
      "votes": null
    },
    {
      "id": "641403",
      "postDate": "10/04/2019 16:54:42",
      "content": "<p>Right <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> , <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175\">this</a> really sounds like it's doing a per group average, which in my opinion doesn't make sense in such a highly skewed set. Or is is somehow weighted and I have missed it?</p>",
      "rawMarkdown": "Right @rishabhiitbhu , [this](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175) really sounds like it's doing a per group average, which in my opinion doesn't make sense in such a highly skewed set. Or is is somehow weighted and I have missed it?",
      "votes": null
    },
    {
      "id": "641454",
      "postDate": "10/04/2019 17:45:28",
      "content": "<p>&gt; which in my opinion doesn't make sense in such a highly skewed set</p>\n\n<p>Exactly.</p>\n\n<p>(your hyperlink is giving 404)</p>\n\n<p>There's no mention about weighted average as such, even I think this metric is not apt if not weighted by class or averaged across samples.</p>",
      "rawMarkdown": "&gt; which in my opinion doesn't make sense in such a highly skewed set\n\nExactly.\n\n(your hyperlink is giving 404)\n\nThere's no mention about weighted average as such, even I think this metric is not apt if not weighted by class or averaged across samples.",
      "votes": null
    },
    {
      "id": "641475",
      "postDate": "10/04/2019 18:11:47",
      "content": "<p>fixed the typo ;)\nI might give it some submission tries as soon as i have a good model to evaluate the metric used here</p>",
      "rawMarkdown": "fixed the typo ;)\nI might give it some submission tries as soon as i have a good model to evaluate the metric used here",
      "votes": null
    },
    {
      "id": "641971",
      "postDate": "10/05/2019 11:45:32",
      "content": "<p>Hi, I can confirm through submission results, this is definitely an unweighted average across all 9 classes</p>",
      "rawMarkdown": "Hi, I can confirm through submission results, this is definitely an unweighted average across all 9 classes",
      "votes": null
    },
    {
      "id": "642004",
      "postDate": "10/05/2019 12:26:57",
      "content": "<p>thanks <a href=\"/thomasgilles\">@thomasgilles</a>, even I can confirm that.\nMy mAP stats (mAP calculated using lyft's sdk evaluation script):\nLocal train mAP: 0.12, val mAP: 0.102 \nLB: 0.093</p>",
      "rawMarkdown": "thanks @thomasgilles, even I can confirm that.\nMy mAP stats (mAP calculated using lyft's sdk evaluation script):\nLocal train mAP: 0.12, val mAP: 0.102 \nLB: 0.093",
      "votes": null
    },
    {
      "id": "642075",
      "postDate": "10/05/2019 14:11:41",
      "content": "<p>Thanks a lot <a href=\"/thomasgilles\">@thomasgilles</a> and <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> \nThat overvalues those small groups such as animals and emergency vehicles by a lot!</p>",
      "rawMarkdown": "Thanks a lot @thomasgilles and @rishabhiitbhu \nThat overvalues those small groups such as animals and emergency vehicles by a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 641065,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "10/04/2019 12:19:42",
      "content": "<p>KITTI/nuScenes challenge benchmark use distance thresholds for mAP calculation, but this challenge uses IoU thresholds (that too from 0.5 to 0.95) which makes this even more challenging.</p>\n\n<p>Just for comparison, here are my stats for my model's predictions evaluated on both the metrics:</p>\n\n<p>nuScenes' mAP (from <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/eval/detection/evaluate.py\">here</a>):\n<code>\ncar: 0.9187143408655809,\npedestrian: 0.25615435803934755,\nother_vehicle: 0.7742371803950503,\nbus: 0.35341638450263824,\nmotorcycle: 0.20038938394575653,\ntruck: 0.549069458817806,\nbicycle: 0.3405492518182246\n</code>\nlyft's mAP (from <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py\">here</a>):\n<code>\ncar: 0.26481168914758046\npedestrian: 0.01484931969844282\nother_vehicle: 0.19598733527638573\nbus: 0.08244517085140785\nmotorcycle: 0.00706896551724138\ntruck: 0.11083860627446822\nbicycle: 0.026110721574923158\n</code>\n[don't have stats for animal and emergency due to my bad validation set, it didn't have any objects from these categories. lol]</p>\n\n<p>As you can see, there's huge difference in both the mAPs. \nWe also have imbalance in class distribution, here's category vs total object count for the lyft dataset:</p>\n\n<p>category, n= total object count \n<code>\nanimal    n=  186\nbicycle    n=20928\nbus     n= 8729\ncar    n=534911\nemergency_vehicle   n= 132\nmotorcycle      n=  818\nother_vehicle   n=33376\npedestrian  n=24935\ntruck   n=14164\n</code>\nAs the lyft's mAP is mean of APs of all the classes, animal and emergency are gonna pull down the overall score. \nSo, It's gonna be tough to score high scores but not impossible :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 641136,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/04/2019 13:03:31",
          "content": "<p>I think you are mistaken here, the scoring here is not averaging between classes.</p>\n\n<p>&gt; At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object <strong>to all ground truth objects</strong></p>\n\n<p>After that, you are right, it will be an average of all thresholds\n&gt; The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold</p>\n\n<p>and</p>\n\n<blockquote>\n  <p>Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset</p>\n</blockquote>\n\n<p>So, even a model only predicting cars should do OK.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 641147,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "10/04/2019 13:11:24",
          "content": "<p>Hey, actually the evaluation page of this competition is a little shady. </p>\n\n<p>&gt;At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object to all ground truth objects</p>\n\n<p>Before doing this, they group predictions by classes (<a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L361\">here</a>) .Once they have class wise average precisions, then they take the mean (<a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L411\">here</a>)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 641382,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/04/2019 16:30:20",
          "content": "<p>I have seen that commit. But I am not convinced that this will be the metric, that they use here for evaluation. \nSome clarification would be greatly appreciated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 641389,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "10/04/2019 16:38:21",
          "content": "<p>see <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175\">this</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 641403,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/04/2019 16:54:42",
          "content": "<p>Right <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> , <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175\">this</a> really sounds like it's doing a per group average, which in my opinion doesn't make sense in such a highly skewed set. Or is is somehow weighted and I have missed it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 641454,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "10/04/2019 17:45:28",
          "content": "<p>&gt; which in my opinion doesn't make sense in such a highly skewed set</p>\n\n<p>Exactly.</p>\n\n<p>(your hyperlink is giving 404)</p>\n\n<p>There's no mention about weighted average as such, even I think this metric is not apt if not weighted by class or averaged across samples.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 641475,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/04/2019 18:11:47",
          "content": "<p>fixed the typo ;)\nI might give it some submission tries as soon as i have a good model to evaluate the metric used here</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 641971,
          "author_name": "thomasgilles",
          "author_url": "",
          "post_date": "10/05/2019 11:45:32",
          "content": "<p>Hi, I can confirm through submission results, this is definitely an unweighted average across all 9 classes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 642004,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "10/05/2019 12:26:57",
          "content": "<p>thanks <a href=\"/thomasgilles\">@thomasgilles</a>, even I can confirm that.\nMy mAP stats (mAP calculated using lyft's sdk evaluation script):\nLocal train mAP: 0.12, val mAP: 0.102 \nLB: 0.093</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 642075,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "10/05/2019 14:11:41",
          "content": "<p>Thanks a lot <a href=\"/thomasgilles\">@thomasgilles</a> and <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> \nThat overvalues those small groups such as animals and emergency vehicles by a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 641145,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "10/04/2019 13:08:57",
      "content": "<p>I am at <code>Average per sample mAP (IoU= 0.5) =  0.0940</code>\nnot even looking at the harder IoUs, yet ...</p>\n\n<p>I believe we can end somewhere in the 0.2 area. This challenge has some really hard to predict cases and a metric that includes IoU thresholds of e.g. 0.95 is somewhat insane ;) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "628266": "We can see on KITTI benchmark, the current top evaluation score is around 0.88 for easy and 0.76  for hard, considering car IoU 0.7 and pedestrian/cyclist IoU 0.5. \n\nWhat's the expectation here for the LB score in 2 months?",
    "641065": "KITTI/nuScenes challenge benchmark use distance thresholds for mAP calculation, but this challenge uses IoU thresholds (that too from 0.5 to 0.95) which makes this even more challenging.\n\nJust for comparison, here are my stats for my model's predictions evaluated on both the metrics:\n\nnuScenes' mAP (from [here](https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/eval/detection/evaluate.py)):\n```\ncar: 0.9187143408655809,\npedestrian: 0.25615435803934755,\nother_vehicle: 0.7742371803950503,\nbus: 0.35341638450263824,\nmotorcycle: 0.20038938394575653,\ntruck: 0.549069458817806,\nbicycle: 0.3405492518182246\n```\nlyft's mAP (from [here](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py)):\n```\ncar: 0.26481168914758046\npedestrian: 0.01484931969844282\nother_vehicle: 0.19598733527638573\nbus: 0.08244517085140785\nmotorcycle: 0.00706896551724138\ntruck: 0.11083860627446822\nbicycle: 0.026110721574923158\n```\n[don't have stats for animal and emergency due to my bad validation set, it didn't have any objects from these categories. lol]\n\nAs you can see, there's huge difference in both the mAPs. \nWe also have imbalance in class distribution, here's category vs total object count for the lyft dataset:\n\ncategory, n= total object count \n```\nanimal    n=  186\nbicycle    n=20928\nbus     n= 8729\ncar    n=534911\nemergency_vehicle   n= 132\nmotorcycle      n=  818\nother_vehicle   n=33376\npedestrian  n=24935\ntruck   n=14164\n```\nAs the lyft's mAP is mean of APs of all the classes, animal and emergency are gonna pull down the overall score. \nSo, It's gonna be tough to score high scores but not impossible :)",
    "641136": "I think you are mistaken here, the scoring here is not averaging between classes.\n\n&gt; At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object **to all ground truth objects**\n\nAfter that, you are right, it will be an average of all thresholds\n&gt; The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold\n\nand\n\n&gt; Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset\n\nSo, even a model only predicting cars should do OK.",
    "641145": "I am at `Average per sample mAP (IoU= 0.5) =  0.0940`\nnot even looking at the harder IoUs, yet ...\n\n\nI believe we can end somewhere in the 0.2 area. This challenge has some really hard to predict cases and a metric that includes IoU thresholds of e.g. 0.95 is somewhat insane ;)",
    "641147": "Hey, actually the evaluation page of this competition is a little shady. \n\n&gt;At each threshold value t, a precision value is calculated based on the number of true positives (TP), false negatives (FN), and false positives (FP) resulting from comparing the predicted object to all ground truth objects\n\nBefore doing this, they group predictions by classes ([here](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L361)) .Once they have class wise average precisions, then they take the mean ([here](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/eval/detection/mAP_evaluation.py#L411))",
    "641382": "I have seen that commit. But I am not convinced that this will be the metric, that they use here for evaluation. \nSome clarification would be greatly appreciated.",
    "641389": "see [this](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175)",
    "641403": "Right @rishabhiitbhu , [this](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/110207#634175) really sounds like it's doing a per group average, which in my opinion doesn't make sense in such a highly skewed set. Or is is somehow weighted and I have missed it?",
    "641454": "&gt; which in my opinion doesn't make sense in such a highly skewed set\n\nExactly.\n\n(your hyperlink is giving 404)\n\nThere's no mention about weighted average as such, even I think this metric is not apt if not weighted by class or averaged across samples.",
    "641475": "fixed the typo ;)\nI might give it some submission tries as soon as i have a good model to evaluate the metric used here",
    "641971": "Hi, I can confirm through submission results, this is definitely an unweighted average across all 9 classes",
    "642004": "thanks @thomasgilles, even I can confirm that.\nMy mAP stats (mAP calculated using lyft's sdk evaluation script):\nLocal train mAP: 0.12, val mAP: 0.102 \nLB: 0.093",
    "642075": "Thanks a lot @thomasgilles and @rishabhiitbhu \nThat overvalues those small groups such as animals and emergency vehicles by a lot!"
  },
  "source": "meta"
}