{
  "id": 116434,
  "title": "Small classes problem",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/116434",
  "author_name": "",
  "post_date": "2019-11-09T03:34:19.848558200Z",
  "votes": 11,
  "comment_count": 16,
  "views": 0,
  "content": "<p>As you know, there are several classes with small amount of samples.\nCategory stats\nemergency_vehicle  n=  132, width= 2.45±0.43, len= 6.52±1.44, height= 2.39±0.59\nanimal                      n=  186, width= 0.36±0.12, len= 0.73±0.19, height= 0.51±0.16\nmotorcycle               n=  818, width= 0.96±0.20, len= 2.35±0.22, height= 1.59±0.16\nbus                            n= 8729, width= 2.96±0.24, len=12.34±3.41, height= 3.44±0.31</p>\n\n<p>I have very low validation metrics (close to zeros) on that classes with my lidar-based NN. And I guess that it gives a huge impact on overall metric score due to averaging mAP over all categories. </p>\n\n<p>Do you have any proven ideas, what to do with small categories? \nI have several ideas, but don't have enough time to check them all :D\nSome thoughts:\n1. Try to use coco pretrained object detection models to detect classes with only camera images. Maybe this kernel can help with that (<a href=\"https://www.kaggle.com/rustemiskuzhin/animal-evaluation\">https://www.kaggle.com/rustemiskuzhin/animal-evaluation</a>)\n2. Try to oversample small categories during training.\n3. Increase loss weight for that categories.\n4. Separate lidar-based NN for small classes problem.\n5. Ur ideas :)</p>\n\n<p>Could u share ur insights or results please?</p>",
  "messages": [
    {
      "id": "668899",
      "postDate": "11/09/2019 03:34:19",
      "content": "<p>As you know, there are several classes with small amount of samples.\nCategory stats\nemergency_vehicle  n=  132, width= 2.45±0.43, len= 6.52±1.44, height= 2.39±0.59\nanimal                      n=  186, width= 0.36±0.12, len= 0.73±0.19, height= 0.51±0.16\nmotorcycle               n=  818, width= 0.96±0.20, len= 2.35±0.22, height= 1.59±0.16\nbus                            n= 8729, width= 2.96±0.24, len=12.34±3.41, height= 3.44±0.31</p>\n\n<p>I have very low validation metrics (close to zeros) on that classes with my lidar-based NN. And I guess that it gives a huge impact on overall metric score due to averaging mAP over all categories. </p>\n\n<p>Do you have any proven ideas, what to do with small categories? \nI have several ideas, but don't have enough time to check them all :D\nSome thoughts:\n1. Try to use coco pretrained object detection models to detect classes with only camera images. Maybe this kernel can help with that (<a href=\"https://www.kaggle.com/rustemiskuzhin/animal-evaluation\">https://www.kaggle.com/rustemiskuzhin/animal-evaluation</a>)\n2. Try to oversample small categories during training.\n3. Increase loss weight for that categories.\n4. Separate lidar-based NN for small classes problem.\n5. Ur ideas :)</p>\n\n<p>Could u share ur insights or results please?</p>",
      "rawMarkdown": "As you know, there are several classes with small amount of samples.\nCategory stats\nemergency_vehicle  n=  132, width= 2.45±0.43, len= 6.52±1.44, height= 2.39±0.59\nanimal                      n=  186, width= 0.36±0.12, len= 0.73±0.19, height= 0.51±0.16\nmotorcycle               n=  818, width= 0.96±0.20, len= 2.35±0.22, height= 1.59±0.16\nbus                            n= 8729, width= 2.96±0.24, len=12.34±3.41, height= 3.44±0.31\n\nI have very low validation metrics (close to zeros) on that classes with my lidar-based NN. And I guess that it gives a huge impact on overall metric score due to averaging mAP over all categories. \n\nDo you have any proven ideas, what to do with small categories? \nI have several ideas, but don't have enough time to check them all :D\nSome thoughts:\n1. Try to use coco pretrained object detection models to detect classes with only camera images. Maybe this kernel can help with that (https://www.kaggle.com/rustemiskuzhin/animal-evaluation)\n2. Try to oversample small categories during training.\n3. Increase loss weight for that categories.\n4. Separate lidar-based NN for small classes problem.\n5. Ur ideas :)\n\nCould u share ur insights or results please?",
      "votes": null
    },
    {
      "id": "669168",
      "postDate": "11/09/2019 15:55:58",
      "content": "<p>The metric is not averaged with class. You do not need to take too much care of it. classes with more samples will have more importance proportional to its occurrence.  I have tried to balance them but observe no improvement.</p>",
      "rawMarkdown": "The metric is not averaged with class. You do not need to take too much care of it. classes with more samples will have more importance proportional to its occurrence.  I have tried to balance them but observe no improvement.",
      "votes": null
    },
    {
      "id": "669371",
      "postDate": "11/10/2019 00:22:53",
      "content": "<p>Are you sure about this? According to the evaluation script, it is averaged by the class. I also think it is more resonable to be weighted by class number </p>",
      "rawMarkdown": "Are you sure about this? According to the evaluation script, it is averaged by the class. I also think it is more resonable to be weighted by class number",
      "votes": null
    },
    {
      "id": "669392",
      "postDate": "11/10/2019 01:32:07",
      "content": "<p>I trained two models with/without balanced classes. For two models, both with/without balanced classes give the same score. So either my class balance does not work nor the metric is not averaged over classes. It would be more sure if the organizer can give further explanations. According to the evaluation part, the metric is averaged over every frames. It will be problematic to average over class for each frame. Let’s assume we have a frame with 0 ground truth object. If we give a detection with 0 object too, we get a score of 1.0(if I am right for empty ground truth frames), otherwise we get score 0.0. Say the animal is absent for most frames. So if we do not detect any object at all but submit an empty result, for most frames we got score 1.0 for animal class (and also 1.0 for other classes absent from that frame). If we average over classes, we are still at around 0.1 from each frame. The overall score should thus be above 0.1 just for an empty submission!</p>",
      "rawMarkdown": "I trained two models with/without balanced classes. For two models, both with/without balanced classes give the same score. So either my class balance does not work nor the metric is not averaged over classes. It would be more sure if the organizer can give further explanations. According to the evaluation part, the metric is averaged over every frames. It will be problematic to average over class for each frame. Let’s assume we have a frame with 0 ground truth object. If we give a detection with 0 object too, we get a score of 1.0(if I am right for empty ground truth frames), otherwise we get score 0.0. Say the animal is absent for most frames. So if we do not detect any object at all but submit an empty result, for most frames we got score 1.0 for animal class (and also 1.0 for other classes absent from that frame). If we average over classes, we are still at around 0.1 from each frame. The overall score should thus be above 0.1 just for an empty submission!",
      "votes": null
    },
    {
      "id": "669627",
      "postDate": "11/10/2019 08:55:19",
      "content": "<p>So Could you inform me what is the ap of small objects of your model in evaluation. I mean, to achieve about 0.21 mAP which is very high if averaged by class, the ap of small object should not close to zero. </p>",
      "rawMarkdown": "So Could you inform me what is the ap of small objects of your model in evaluation. I mean, to achieve about 0.21 mAP which is very high if averaged by class, the ap of small object should not close to zero.",
      "votes": null
    },
    {
      "id": "669656",
      "postDate": "11/10/2019 09:52:22",
      "content": "<p>I don't know. The dataset is so large so that I did not do any validation. But I guess it should be near to 0 too.  </p>",
      "rawMarkdown": "I don't know. The dataset is so large so that I did not do any validation. But I guess it should be near to 0 too.",
      "votes": null
    },
    {
      "id": "669679",
      "postDate": "11/10/2019 10:16:35",
      "content": "<p>I tried to correct baseline model for animal class with 2D image and PCA <a href=\"https://www.kaggle.com/rustemiskuzhin/animal-evaluation\">https://www.kaggle.com/rustemiskuzhin/animal-evaluation</a>. </p>\n\n<p>I validated model for 10000 samples (predicted 3d boxes): \n<a href=\"https://www.kaggle.com/rustemiskuzhin/voxelnet-train-evaluation\">https://www.kaggle.com/rustemiskuzhin/voxelnet-train-evaluation</a></p>\n\n<p>For animal class and IoU = 0.01 I have AP(animal) = 0.66 but for IoU = 0.5 I have AP(animal) = 0.001. \nI think AP is so small because height intersection is low and estimation of yaw  from PCA  is not good for this class.</p>",
      "rawMarkdown": "I tried to correct baseline model for animal class with 2D image and PCA https://www.kaggle.com/rustemiskuzhin/animal-evaluation. \n\nI validated model for 10000 samples (predicted 3d boxes): \nhttps://www.kaggle.com/rustemiskuzhin/voxelnet-train-evaluation\n\nFor animal class and IoU = 0.01 I have AP(animal) = 0.66 but for IoU = 0.5 I have AP(animal) = 0.001. \nI think AP is so small because height intersection is low and estimation of yaw  from PCA  is not good for this class.",
      "votes": null
    },
    {
      "id": "670668",
      "postDate": "11/11/2019 17:49:45",
      "content": "<ol>\n<li>The evaluate metrics is definitely based on class-average, which you could refer to the evaluation code. 😄 </li>\n<li>Your example on animal I think is not correct. Let's say you only have small number(G) of animal ground truth, and you submit(P) number of predictions. <strong>*<em>They will calculate the True Positive(TP) over all frames</em>*</strong> based on different IOU threshold, then AP@IOU = TP / (G + P - TP). Your example say you submit 0 prediction, which means TP is AP@IOU = 0 / (G + 0 - 0) = 0.</li>\n</ol>",
      "rawMarkdown": "1. The evaluate metrics is definitely based on class-average, which you could refer to the evaluation code. 😄 \n2. Your example on animal I think is not correct. Let's say you only have small number(G) of animal ground truth, and you submit(P) number of predictions. ****They will calculate the True Positive(TP) over all frames**** based on different IOU threshold, then AP@IOU = TP / (G + P - TP). Your example say you submit 0 prediction, which means TP is AP@IOU = 0 / (G + 0 - 0) = 0.",
      "votes": null
    },
    {
      "id": "670856",
      "postDate": "11/12/2019 00:36:52",
      "content": "<p>According to the organizer: Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset.</p>",
      "rawMarkdown": "According to the organizer: Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset.",
      "votes": null
    },
    {
      "id": "670858",
      "postDate": "11/12/2019 00:43:20",
      "content": "<p>And also: The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold. \n It seems the ap is calculated frame by frame.</p>",
      "rawMarkdown": "And also: The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold. \n It seems the ap is calculated frame by frame.",
      "votes": null
    },
    {
      "id": "670880",
      "postDate": "11/12/2019 01:43:07",
      "content": "<p><a href=\"/hanxiaodeng\">@hanxiaodeng</a> , for clarity, I actually agree with your calculation and it is the way how the eval script does. I am just not sure the LB is computed exactly with that code. Also, the script only compute mAP for each class. It does not show how to average them to get the final score: simple average or weighted average? </p>\n\n<p>I tried the work \" zhu etc. Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection \" to balance the classes but got no improvement: two version gave exactly the same score.</p>\n\n<p>Of course, it is quite possible my class-balance scheme is incorrect or it does not work on Lyft dataset. All the judgement is based my intuition.</p>",
      "rawMarkdown": "hanxiaodeng , for clarity, I actually agree with your calculation and it is the way how the eval script does. I am just not sure the LB is computed exactly with that code. Also, the script only compute mAP for each class. It does not show how to average them to get the final score: simple average or weighted average? \n\nI tried the work \" zhu etc. Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection \" to balance the classes but got no improvement: two version gave exactly the same score.\n\nOf course, it is quite possible my class-balance scheme is incorrect or it does not work on Lyft dataset. All the judgement is based my intuition.",
      "votes": null
    },
    {
      "id": "670891",
      "postDate": "11/12/2019 02:02:15",
      "content": "<p><a href=\"/nywenjing\">@nywenjing</a> Thanks for reply. </p>\n\n<ol>\n<li><p><strong>Simple average</strong>. I try to find the competition host for his reply on some discussion thread that they confirm that they just do simple average.</p></li>\n<li><p>I also tried \"Classes-balanced Grouping and Sampling for Point Cloud 3D Object Detection\" <strong>which doesn't work either</strong>. That only means this strategy doesn't work for this dataset, I don't see any strong relationship between evaluation metrics and the result.</p></li>\n</ol>",
      "rawMarkdown": "nywenjing Thanks for reply. \n\n1. **Simple average**. I try to find the competition host for his reply on some discussion thread that they confirm that they just do simple average.\n\n2. I also tried \"Classes-balanced Grouping and Sampling for Point Cloud 3D Object Detection\" **which doesn't work either**. That only means this strategy doesn't work for this dataset, I don't see any strong relationship between evaluation metrics and the result.",
      "votes": null
    },
    {
      "id": "671937",
      "postDate": "11/13/2019 10:47:17",
      "content": "<p>I was disappointed too. I put a lot of time in finding most emergency-vehicles (police-car, ambulance, fire-truck) and got about 500 with about 10% being false positives. So i am pretty sure, that i found about &gt;50% of the emergency-vehicles, but it doesn't reflect in the score at all. I was thinking that it was due to lack of emergency-vehicle in the public leaderboard data, but now it seems the low-count classes don't have impact at all. I loved the competition but it would have been better if the host would have been clearer about the metric regarding the low-count classes. I still don't understand why the score for animals and emergency-vehicle don't seem to affect the score at all. Or maybe C# evaluation on kaggle is different to the evaluation in the devkit.</p>\n\n<p>Besides this i loved the competition. Congratulations to you guys!</p>",
      "rawMarkdown": "I was disappointed too. I put a lot of time in finding most emergency-vehicles (police-car, ambulance, fire-truck) and got about 500 with about 10% being false positives. So i am pretty sure, that i found about &gt;50% of the emergency-vehicles, but it doesn't reflect in the score at all. I was thinking that it was due to lack of emergency-vehicle in the public leaderboard data, but now it seems the low-count classes don't have impact at all. I loved the competition but it would have been better if the host would have been clearer about the metric regarding the low-count classes. I still don't understand why the score for animals and emergency-vehicle don't seem to affect the score at all. Or maybe C# evaluation on kaggle is different to the evaluation in the devkit.\n\nBesides this i loved the competition. Congratulations to you guys!",
      "votes": null
    },
    {
      "id": "673846",
      "postDate": "11/15/2019 14:52:57",
      "content": "<p>Let me also add, that this metric class-avg-per-frame instead of frame-avg-per-class is bad for safety as it will tend to ignore or put low importance on motorcycles, emergeny-vehicles and animals for the sake of the high-count classes :-/</p>",
      "rawMarkdown": "Let me also add, that this metric class-avg-per-frame instead of frame-avg-per-class is bad for safety as it will tend to ignore or put low importance on motorcycles, emergeny-vehicles and animals for the sake of the high-count classes :-/",
      "votes": null
    },
    {
      "id": "678713",
      "postDate": "11/21/2019 18:45:33",
      "content": "<p>nice hint</p>",
      "rawMarkdown": "nice hint",
      "votes": null
    },
    {
      "id": "679749",
      "postDate": "11/23/2019 10:14:58",
      "content": "<p>Separate lidar-based NN for small classes problem - it is bad ways. In my experience, when you have class with small amount of samples, necessary pick simple model. Try use logistic regression for such classes.</p>",
      "rawMarkdown": "Separate lidar-based NN for small classes problem - it is bad ways. In my experience, when you have class with small amount of samples, necessary pick simple model. Try use logistic regression for such classes.",
      "votes": null
    },
    {
      "id": "680451",
      "postDate": "11/24/2019 17:13:28",
      "content": "<p>nobody tried focal loss? <a href=\"https://arxiv.org/pdf/1708.02002.pdf\">https://arxiv.org/pdf/1708.02002.pdf</a> </p>",
      "rawMarkdown": "nobody tried focal loss? https://arxiv.org/pdf/1708.02002.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 669168,
      "author_name": "nywenjing",
      "author_url": "",
      "post_date": "11/09/2019 15:55:58",
      "content": "<p>The metric is not averaged with class. You do not need to take too much care of it. classes with more samples will have more importance proportional to its occurrence.  I have tried to balance them but observe no improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 669371,
          "author_name": "jeremydzx",
          "author_url": "",
          "post_date": "11/10/2019 00:22:53",
          "content": "<p>Are you sure about this? According to the evaluation script, it is averaged by the class. I also think it is more resonable to be weighted by class number </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 669392,
          "author_name": "nywenjing",
          "author_url": "",
          "post_date": "11/10/2019 01:32:07",
          "content": "<p>I trained two models with/without balanced classes. For two models, both with/without balanced classes give the same score. So either my class balance does not work nor the metric is not averaged over classes. It would be more sure if the organizer can give further explanations. According to the evaluation part, the metric is averaged over every frames. It will be problematic to average over class for each frame. Let’s assume we have a frame with 0 ground truth object. If we give a detection with 0 object too, we get a score of 1.0(if I am right for empty ground truth frames), otherwise we get score 0.0. Say the animal is absent for most frames. So if we do not detect any object at all but submit an empty result, for most frames we got score 1.0 for animal class (and also 1.0 for other classes absent from that frame). If we average over classes, we are still at around 0.1 from each frame. The overall score should thus be above 0.1 just for an empty submission!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 669627,
          "author_name": "jeremydzx",
          "author_url": "",
          "post_date": "11/10/2019 08:55:19",
          "content": "<p>So Could you inform me what is the ap of small objects of your model in evaluation. I mean, to achieve about 0.21 mAP which is very high if averaged by class, the ap of small object should not close to zero. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 669656,
          "author_name": "nywenjing",
          "author_url": "",
          "post_date": "11/10/2019 09:52:22",
          "content": "<p>I don't know. The dataset is so large so that I did not do any validation. But I guess it should be near to 0 too.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670668,
          "author_name": "hanxiaodeng",
          "author_url": "",
          "post_date": "11/11/2019 17:49:45",
          "content": "<ol>\n<li>The evaluate metrics is definitely based on class-average, which you could refer to the evaluation code. 😄 </li>\n<li>Your example on animal I think is not correct. Let's say you only have small number(G) of animal ground truth, and you submit(P) number of predictions. <strong>*<em>They will calculate the True Positive(TP) over all frames</em>*</strong> based on different IOU threshold, then AP@IOU = TP / (G + P - TP). Your example say you submit 0 prediction, which means TP is AP@IOU = 0 / (G + 0 - 0) = 0.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670856,
          "author_name": "nywenjing",
          "author_url": "",
          "post_date": "11/12/2019 00:36:52",
          "content": "<p>According to the organizer: Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670858,
          "author_name": "nywenjing",
          "author_url": "",
          "post_date": "11/12/2019 00:43:20",
          "content": "<p>And also: The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold. \n It seems the ap is calculated frame by frame.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670880,
          "author_name": "nywenjing",
          "author_url": "",
          "post_date": "11/12/2019 01:43:07",
          "content": "<p><a href=\"/hanxiaodeng\">@hanxiaodeng</a> , for clarity, I actually agree with your calculation and it is the way how the eval script does. I am just not sure the LB is computed exactly with that code. Also, the script only compute mAP for each class. It does not show how to average them to get the final score: simple average or weighted average? </p>\n\n<p>I tried the work \" zhu etc. Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection \" to balance the classes but got no improvement: two version gave exactly the same score.</p>\n\n<p>Of course, it is quite possible my class-balance scheme is incorrect or it does not work on Lyft dataset. All the judgement is based my intuition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670891,
          "author_name": "hanxiaodeng",
          "author_url": "",
          "post_date": "11/12/2019 02:02:15",
          "content": "<p><a href=\"/nywenjing\">@nywenjing</a> Thanks for reply. </p>\n\n<ol>\n<li><p><strong>Simple average</strong>. I try to find the competition host for his reply on some discussion thread that they confirm that they just do simple average.</p></li>\n<li><p>I also tried \"Classes-balanced Grouping and Sampling for Point Cloud 3D Object Detection\" <strong>which doesn't work either</strong>. That only means this strategy doesn't work for this dataset, I don't see any strong relationship between evaluation metrics and the result.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 671937,
          "author_name": "marekwyborski",
          "author_url": "",
          "post_date": "11/13/2019 10:47:17",
          "content": "<p>I was disappointed too. I put a lot of time in finding most emergency-vehicles (police-car, ambulance, fire-truck) and got about 500 with about 10% being false positives. So i am pretty sure, that i found about &gt;50% of the emergency-vehicles, but it doesn't reflect in the score at all. I was thinking that it was due to lack of emergency-vehicle in the public leaderboard data, but now it seems the low-count classes don't have impact at all. I loved the competition but it would have been better if the host would have been clearer about the metric regarding the low-count classes. I still don't understand why the score for animals and emergency-vehicle don't seem to affect the score at all. Or maybe C# evaluation on kaggle is different to the evaluation in the devkit.</p>\n\n<p>Besides this i loved the competition. Congratulations to you guys!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 669679,
      "author_name": "rustemiskuzhin",
      "author_url": "",
      "post_date": "11/10/2019 10:16:35",
      "content": "<p>I tried to correct baseline model for animal class with 2D image and PCA <a href=\"https://www.kaggle.com/rustemiskuzhin/animal-evaluation\">https://www.kaggle.com/rustemiskuzhin/animal-evaluation</a>. </p>\n\n<p>I validated model for 10000 samples (predicted 3d boxes): \n<a href=\"https://www.kaggle.com/rustemiskuzhin/voxelnet-train-evaluation\">https://www.kaggle.com/rustemiskuzhin/voxelnet-train-evaluation</a></p>\n\n<p>For animal class and IoU = 0.01 I have AP(animal) = 0.66 but for IoU = 0.5 I have AP(animal) = 0.001. \nI think AP is so small because height intersection is low and estimation of yaw  from PCA  is not good for this class.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 673846,
      "author_name": "marekwyborski",
      "author_url": "",
      "post_date": "11/15/2019 14:52:57",
      "content": "<p>Let me also add, that this metric class-avg-per-frame instead of frame-avg-per-class is bad for safety as it will tend to ignore or put low importance on motorcycles, emergeny-vehicles and animals for the sake of the high-count classes :-/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 678713,
      "author_name": "wickjon",
      "author_url": "",
      "post_date": "11/21/2019 18:45:33",
      "content": "<p>nice hint</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 679749,
      "author_name": "petrchuikov",
      "author_url": "",
      "post_date": "11/23/2019 10:14:58",
      "content": "<p>Separate lidar-based NN for small classes problem - it is bad ways. In my experience, when you have class with small amount of samples, necessary pick simple model. Try use logistic regression for such classes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 680451,
      "author_name": "oarphme",
      "author_url": "",
      "post_date": "11/24/2019 17:13:28",
      "content": "<p>nobody tried focal loss? <a href=\"https://arxiv.org/pdf/1708.02002.pdf\">https://arxiv.org/pdf/1708.02002.pdf</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "668899": "As you know, there are several classes with small amount of samples.\nCategory stats\nemergency_vehicle  n=  132, width= 2.45±0.43, len= 6.52±1.44, height= 2.39±0.59\nanimal                      n=  186, width= 0.36±0.12, len= 0.73±0.19, height= 0.51±0.16\nmotorcycle               n=  818, width= 0.96±0.20, len= 2.35±0.22, height= 1.59±0.16\nbus                            n= 8729, width= 2.96±0.24, len=12.34±3.41, height= 3.44±0.31\n\nI have very low validation metrics (close to zeros) on that classes with my lidar-based NN. And I guess that it gives a huge impact on overall metric score due to averaging mAP over all categories. \n\nDo you have any proven ideas, what to do with small categories? \nI have several ideas, but don't have enough time to check them all :D\nSome thoughts:\n1. Try to use coco pretrained object detection models to detect classes with only camera images. Maybe this kernel can help with that (https://www.kaggle.com/rustemiskuzhin/animal-evaluation)\n2. Try to oversample small categories during training.\n3. Increase loss weight for that categories.\n4. Separate lidar-based NN for small classes problem.\n5. Ur ideas :)\n\nCould u share ur insights or results please?",
    "669168": "The metric is not averaged with class. You do not need to take too much care of it. classes with more samples will have more importance proportional to its occurrence.  I have tried to balance them but observe no improvement.",
    "669371": "Are you sure about this? According to the evaluation script, it is averaged by the class. I also think it is more resonable to be weighted by class number",
    "669392": "I trained two models with/without balanced classes. For two models, both with/without balanced classes give the same score. So either my class balance does not work nor the metric is not averaged over classes. It would be more sure if the organizer can give further explanations. According to the evaluation part, the metric is averaged over every frames. It will be problematic to average over class for each frame. Let’s assume we have a frame with 0 ground truth object. If we give a detection with 0 object too, we get a score of 1.0(if I am right for empty ground truth frames), otherwise we get score 0.0. Say the animal is absent for most frames. So if we do not detect any object at all but submit an empty result, for most frames we got score 1.0 for animal class (and also 1.0 for other classes absent from that frame). If we average over classes, we are still at around 0.1 from each frame. The overall score should thus be above 0.1 just for an empty submission!",
    "669627": "So Could you inform me what is the ap of small objects of your model in evaluation. I mean, to achieve about 0.21 mAP which is very high if averaged by class, the ap of small object should not close to zero.",
    "669656": "I don't know. The dataset is so large so that I did not do any validation. But I guess it should be near to 0 too.",
    "669679": "I tried to correct baseline model for animal class with 2D image and PCA https://www.kaggle.com/rustemiskuzhin/animal-evaluation. \n\nI validated model for 10000 samples (predicted 3d boxes): \nhttps://www.kaggle.com/rustemiskuzhin/voxelnet-train-evaluation\n\nFor animal class and IoU = 0.01 I have AP(animal) = 0.66 but for IoU = 0.5 I have AP(animal) = 0.001. \nI think AP is so small because height intersection is low and estimation of yaw  from PCA  is not good for this class.",
    "670668": "1. The evaluate metrics is definitely based on class-average, which you could refer to the evaluation code. 😄 \n2. Your example on animal I think is not correct. Let's say you only have small number(G) of animal ground truth, and you submit(P) number of predictions. ****They will calculate the True Positive(TP) over all frames**** based on different IOU threshold, then AP@IOU = TP / (G + P - TP). Your example say you submit 0 prediction, which means TP is AP@IOU = 0 / (G + 0 - 0) = 0.",
    "670856": "According to the organizer: Lastly, the score returned by the competition metric is the mean taken over the individual average precisions of each image in the test dataset.",
    "670858": "And also: The average precision of a single image is calculated as the mean of the above precision values at each IoU threshold. \n It seems the ap is calculated frame by frame.",
    "670880": "hanxiaodeng , for clarity, I actually agree with your calculation and it is the way how the eval script does. I am just not sure the LB is computed exactly with that code. Also, the script only compute mAP for each class. It does not show how to average them to get the final score: simple average or weighted average? \n\nI tried the work \" zhu etc. Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection \" to balance the classes but got no improvement: two version gave exactly the same score.\n\nOf course, it is quite possible my class-balance scheme is incorrect or it does not work on Lyft dataset. All the judgement is based my intuition.",
    "670891": "nywenjing Thanks for reply. \n\n1. **Simple average**. I try to find the competition host for his reply on some discussion thread that they confirm that they just do simple average.\n\n2. I also tried \"Classes-balanced Grouping and Sampling for Point Cloud 3D Object Detection\" **which doesn't work either**. That only means this strategy doesn't work for this dataset, I don't see any strong relationship between evaluation metrics and the result.",
    "671937": "I was disappointed too. I put a lot of time in finding most emergency-vehicles (police-car, ambulance, fire-truck) and got about 500 with about 10% being false positives. So i am pretty sure, that i found about &gt;50% of the emergency-vehicles, but it doesn't reflect in the score at all. I was thinking that it was due to lack of emergency-vehicle in the public leaderboard data, but now it seems the low-count classes don't have impact at all. I loved the competition but it would have been better if the host would have been clearer about the metric regarding the low-count classes. I still don't understand why the score for animals and emergency-vehicle don't seem to affect the score at all. Or maybe C# evaluation on kaggle is different to the evaluation in the devkit.\n\nBesides this i loved the competition. Congratulations to you guys!",
    "673846": "Let me also add, that this metric class-avg-per-frame instead of frame-avg-per-class is bad for safety as it will tend to ignore or put low importance on motorcycles, emergeny-vehicles and animals for the sake of the high-count classes :-/",
    "678713": "nice hint",
    "679749": "Separate lidar-based NN for small classes problem - it is bad ways. In my experience, when you have class with small amount of samples, necessary pick simple model. Try use logistic regression for such classes.",
    "680451": "nobody tried focal loss? https://arxiv.org/pdf/1708.02002.pdf"
  },
  "source": "meta"
}