{
  "id": 110969,
  "title": "Missing annotations?",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/110969",
  "author_name": "",
  "post_date": "2019-10-02T16:42:15.090838500Z",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Scene token: 71dfb15d2f88bf2aab2c5d4800c0d10a76c279b9fda98720781a406cbacc583b</p>\n\n<p>Birds eye view of one of the samples in this scene with ground truth annotations:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F9e968f72cdd20fa85d5022dbab46d9dc%2Fgt.png?generation=1570033982309217&amp;alt=media\" alt=\"\"></p>\n\n<p>My model's predictions for the sample:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F42002c59a595801bf13ddcdfc452d17f%2Fpred.png?generation=1570034331122527&amp;alt=media\" alt=\"\"></p>\n\n<p>Notice the difference, don't you think we have many missing ground truth annotations here?</p>\n\n<p>Here are some cam views: (screenshots from <code>render_scene</code>):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2Fb781279019379f683364fdc569a6b306%2Fpic-selected-191002-2140-16.png?generation=1570034416454010&amp;alt=media\" alt=\"\"></p>\n\n<p>the cars on the other lane aren't annotated, but why?</p>\n\n<p>CC: <a href=\"/iglovikov\">@iglovikov</a> </p>",
  "messages": [
    {
      "id": "638985",
      "postDate": "10/02/2019 16:42:15",
      "content": "<p>Scene token: 71dfb15d2f88bf2aab2c5d4800c0d10a76c279b9fda98720781a406cbacc583b</p>\n\n<p>Birds eye view of one of the samples in this scene with ground truth annotations:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F9e968f72cdd20fa85d5022dbab46d9dc%2Fgt.png?generation=1570033982309217&amp;alt=media\" alt=\"\"></p>\n\n<p>My model's predictions for the sample:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F42002c59a595801bf13ddcdfc452d17f%2Fpred.png?generation=1570034331122527&amp;alt=media\" alt=\"\"></p>\n\n<p>Notice the difference, don't you think we have many missing ground truth annotations here?</p>\n\n<p>Here are some cam views: (screenshots from <code>render_scene</code>):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2Fb781279019379f683364fdc569a6b306%2Fpic-selected-191002-2140-16.png?generation=1570034416454010&amp;alt=media\" alt=\"\"></p>\n\n<p>the cars on the other lane aren't annotated, but why?</p>\n\n<p>CC: <a href=\"/iglovikov\">@iglovikov</a> </p>",
      "rawMarkdown": "Scene token: 71dfb15d2f88bf2aab2c5d4800c0d10a76c279b9fda98720781a406cbacc583b\n\nBirds eye view of one of the samples in this scene with ground truth annotations:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F9e968f72cdd20fa85d5022dbab46d9dc%2Fgt.png?generation=1570033982309217&amp;alt=media)\n\nMy model's predictions for the sample:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F42002c59a595801bf13ddcdfc452d17f%2Fpred.png?generation=1570034331122527&amp;alt=media)\n\nNotice the difference, don't you think we have many missing ground truth annotations here?\n\nHere are some cam views: (screenshots from `render_scene`):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2Fb781279019379f683364fdc569a6b306%2Fpic-selected-191002-2140-16.png?generation=1570034416454010&amp;alt=media)\n\nthe cars on the other lane aren't annotated, but why?\n\nCC: @iglovikov",
      "votes": null
    },
    {
      "id": "639886",
      "postDate": "10/03/2019 17:10:12",
      "content": "<p>I have read a blog from Nvidia a while ago, where they said parked cars in parking lots should not be annotated. But for nuscenes this is not the case.</p>\n\n<p>Have a look at their annotation instructions here:\n<a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/instructions.md\">https://github.com/nutonomy/nuscenes-devkit/blob/master/instructions.md</a></p>\n\n<p>So, yes, I think there a missing some annotations. Could we have some clarification from the host?</p>",
      "rawMarkdown": "I have read a blog from Nvidia a while ago, where they said parked cars in parking lots should not be annotated. But for nuscenes this is not the case.\n\nHave a look at their annotation instructions here:\nhttps://github.com/nutonomy/nuscenes-devkit/blob/master/instructions.md\n\nSo, yes, I think there a missing some annotations. Could we have some clarification from the host?",
      "votes": null
    },
    {
      "id": "639973",
      "postDate": "10/03/2019 18:23:51",
      "content": "<p>Yeah, it would have been good if we had a detailed paper/description about the lyft dataset as we have for the nuscenes. We know that lyft dataset follows format of nuscenes, but there's no clue about the methodology of data annotation, sensor synchronization, sensor calibration and many other things.</p>",
      "rawMarkdown": "Yeah, it would have been good if we had a detailed paper/description about the lyft dataset as we have for the nuscenes. We know that lyft dataset follows format of nuscenes, but there's no clue about the methodology of data annotation, sensor synchronization, sensor calibration and many other things.",
      "votes": null
    },
    {
      "id": "640530",
      "postDate": "10/04/2019 04:48:30",
      "content": "<p>It looks like we will need another model to train on false positives.</p>",
      "rawMarkdown": "It looks like we will need another model to train on false positives.",
      "votes": null
    },
    {
      "id": "641923",
      "postDate": "10/05/2019 10:15:22",
      "content": "<p>There's also this non-parked car, that is an active traffic participant, that is missing a label and does not appear to be beyond lidar range:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3779103%2Fbece70672306ba447eec4e0884e9e2b5%2FScreen%20Shot%202019-10-02%20at%2017.13.40.png?generation=1570269992785155&amp;alt=media\" alt=\"\"></p>\n\n<p>This is a pretty typical data error that you'd see with any data provider.  It's not really Scale's fault because they do not do quality assurance for <em>every</em> single label.  Everybody does review, but there are probably at least 5% of scenes that have at least one error.</p>\n\n<p>I think the organizer can fix this a few different ways:\n 1) Have Scale or somebody update the test set to fix labels\n 2) Remove the scene or part of the scene from the test set entirely for today\n 3) Use the map to introduce polygons where predictions are ignored.  This would be similar to what Waymo does in their dataset.  It's also not uncommon in industry to simply ignore certain polygons.  </p>\n\n<p>My guess is #2 is easiest and perhaps #1 is possible.  Not sure what Kaggle's position is on updates to the test set.  It probably only matters if the leaderboard starts to show algorithms with error rate approaching human error on the dataset.</p>\n\n<p>I think if there are truly a lot of scenes, like order of 1,000 or so, then we have to adopt polygons to ignore things like parking lots and humans inside buildings and stuff.  But for a test set this small, maybe just remove the parking lot cases for now.</p>",
      "rawMarkdown": "There's also this non-parked car, that is an active traffic participant, that is missing a label and does not appear to be beyond lidar range:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3779103%2Fbece70672306ba447eec4e0884e9e2b5%2FScreen%20Shot%202019-10-02%20at%2017.13.40.png?generation=1570269992785155&amp;alt=media)\n\nThis is a pretty typical data error that you'd see with any data provider.  It's not really Scale's fault because they do not do quality assurance for *every* single label.  Everybody does review, but there are probably at least 5% of scenes that have at least one error.\n\nI think the organizer can fix this a few different ways:\n 1) Have Scale or somebody update the test set to fix labels\n 2) Remove the scene or part of the scene from the test set entirely for today\n 3) Use the map to introduce polygons where predictions are ignored.  This would be similar to what Waymo does in their dataset.  It's also not uncommon in industry to simply ignore certain polygons.  \n\nMy guess is #2 is easiest and perhaps #1 is possible.  Not sure what Kaggle's position is on updates to the test set.  It probably only matters if the leaderboard starts to show algorithms with error rate approaching human error on the dataset.\n\nI think if there are truly a lot of scenes, like order of 1,000 or so, then we have to adopt polygons to ignore things like parking lots and humans inside buildings and stuff.  But for a test set this small, maybe just remove the parking lot cases for now.",
      "votes": null
    },
    {
      "id": "642003",
      "postDate": "10/05/2019 12:25:46",
      "content": "<p>Unfortunately, we have to train the model which reproduces even missing annotation 😧</p>",
      "rawMarkdown": "Unfortunately, we have to train the model which reproduces even missing annotation 😧",
      "votes": null
    },
    {
      "id": "652283",
      "postDate": "10/18/2019 16:20:15",
      "content": "<p>Any news on updated annotations? Or do we have to deal with the missing annotations? I attached some more examples of missing annotations i found during training (the green boxes depict predictions here not GT). The problem i think with those cars not annotated is, that they are really obvious positives to the model, while other cars are really not as obvious but are still annotated.</p>",
      "rawMarkdown": "Any news on updated annotations? Or do we have to deal with the missing annotations? I attached some more examples of missing annotations i found during training (the green boxes depict predictions here not GT). The problem i think with those cars not annotated is, that they are really obvious positives to the model, while other cars are really not as obvious but are still annotated.",
      "votes": null
    },
    {
      "id": "652308",
      "postDate": "10/18/2019 16:50:53",
      "content": "<p>Sadly, there's no reply from the mods. I've no clue what to do with this.</p>",
      "rawMarkdown": "Sadly, there's no reply from the mods. I've no clue what to do with this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 639886,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "10/03/2019 17:10:12",
      "content": "<p>I have read a blog from Nvidia a while ago, where they said parked cars in parking lots should not be annotated. But for nuscenes this is not the case.</p>\n\n<p>Have a look at their annotation instructions here:\n<a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/instructions.md\">https://github.com/nutonomy/nuscenes-devkit/blob/master/instructions.md</a></p>\n\n<p>So, yes, I think there a missing some annotations. Could we have some clarification from the host?</p>",
      "votes": null,
      "replies": [
        {
          "id": 639973,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "10/03/2019 18:23:51",
          "content": "<p>Yeah, it would have been good if we had a detailed paper/description about the lyft dataset as we have for the nuscenes. We know that lyft dataset follows format of nuscenes, but there's no clue about the methodology of data annotation, sensor synchronization, sensor calibration and many other things.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 640530,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "10/04/2019 04:48:30",
      "content": "<p>It looks like we will need another model to train on false positives.</p>",
      "votes": null,
      "replies": [
        {
          "id": 642003,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "10/05/2019 12:25:46",
          "content": "<p>Unfortunately, we have to train the model which reproduces even missing annotation 😧</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 641923,
      "author_name": "oarphme",
      "author_url": "",
      "post_date": "10/05/2019 10:15:22",
      "content": "<p>There's also this non-parked car, that is an active traffic participant, that is missing a label and does not appear to be beyond lidar range:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3779103%2Fbece70672306ba447eec4e0884e9e2b5%2FScreen%20Shot%202019-10-02%20at%2017.13.40.png?generation=1570269992785155&amp;alt=media\" alt=\"\"></p>\n\n<p>This is a pretty typical data error that you'd see with any data provider.  It's not really Scale's fault because they do not do quality assurance for <em>every</em> single label.  Everybody does review, but there are probably at least 5% of scenes that have at least one error.</p>\n\n<p>I think the organizer can fix this a few different ways:\n 1) Have Scale or somebody update the test set to fix labels\n 2) Remove the scene or part of the scene from the test set entirely for today\n 3) Use the map to introduce polygons where predictions are ignored.  This would be similar to what Waymo does in their dataset.  It's also not uncommon in industry to simply ignore certain polygons.  </p>\n\n<p>My guess is #2 is easiest and perhaps #1 is possible.  Not sure what Kaggle's position is on updates to the test set.  It probably only matters if the leaderboard starts to show algorithms with error rate approaching human error on the dataset.</p>\n\n<p>I think if there are truly a lot of scenes, like order of 1,000 or so, then we have to adopt polygons to ignore things like parking lots and humans inside buildings and stuff.  But for a test set this small, maybe just remove the parking lot cases for now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 652283,
      "author_name": "tobiasfshr",
      "author_url": "",
      "post_date": "10/18/2019 16:20:15",
      "content": "<p>Any news on updated annotations? Or do we have to deal with the missing annotations? I attached some more examples of missing annotations i found during training (the green boxes depict predictions here not GT). The problem i think with those cars not annotated is, that they are really obvious positives to the model, while other cars are really not as obvious but are still annotated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 652308,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "10/18/2019 16:50:53",
          "content": "<p>Sadly, there's no reply from the mods. I've no clue what to do with this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "638985": "Scene token: 71dfb15d2f88bf2aab2c5d4800c0d10a76c279b9fda98720781a406cbacc583b\n\nBirds eye view of one of the samples in this scene with ground truth annotations:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F9e968f72cdd20fa85d5022dbab46d9dc%2Fgt.png?generation=1570033982309217&amp;alt=media)\n\nMy model's predictions for the sample:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2F42002c59a595801bf13ddcdfc452d17f%2Fpred.png?generation=1570034331122527&amp;alt=media)\n\nNotice the difference, don't you think we have many missing ground truth annotations here?\n\nHere are some cam views: (screenshots from `render_scene`):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F761152%2Fb781279019379f683364fdc569a6b306%2Fpic-selected-191002-2140-16.png?generation=1570034416454010&amp;alt=media)\n\nthe cars on the other lane aren't annotated, but why?\n\nCC: @iglovikov",
    "639886": "I have read a blog from Nvidia a while ago, where they said parked cars in parking lots should not be annotated. But for nuscenes this is not the case.\n\nHave a look at their annotation instructions here:\nhttps://github.com/nutonomy/nuscenes-devkit/blob/master/instructions.md\n\nSo, yes, I think there a missing some annotations. Could we have some clarification from the host?",
    "639973": "Yeah, it would have been good if we had a detailed paper/description about the lyft dataset as we have for the nuscenes. We know that lyft dataset follows format of nuscenes, but there's no clue about the methodology of data annotation, sensor synchronization, sensor calibration and many other things.",
    "640530": "It looks like we will need another model to train on false positives.",
    "641923": "There's also this non-parked car, that is an active traffic participant, that is missing a label and does not appear to be beyond lidar range:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3779103%2Fbece70672306ba447eec4e0884e9e2b5%2FScreen%20Shot%202019-10-02%20at%2017.13.40.png?generation=1570269992785155&amp;alt=media)\n\nThis is a pretty typical data error that you'd see with any data provider.  It's not really Scale's fault because they do not do quality assurance for *every* single label.  Everybody does review, but there are probably at least 5% of scenes that have at least one error.\n\nI think the organizer can fix this a few different ways:\n 1) Have Scale or somebody update the test set to fix labels\n 2) Remove the scene or part of the scene from the test set entirely for today\n 3) Use the map to introduce polygons where predictions are ignored.  This would be similar to what Waymo does in their dataset.  It's also not uncommon in industry to simply ignore certain polygons.  \n\nMy guess is #2 is easiest and perhaps #1 is possible.  Not sure what Kaggle's position is on updates to the test set.  It probably only matters if the leaderboard starts to show algorithms with error rate approaching human error on the dataset.\n\nI think if there are truly a lot of scenes, like order of 1,000 or so, then we have to adopt polygons to ignore things like parking lots and humans inside buildings and stuff.  But for a test set this small, maybe just remove the parking lot cases for now.",
    "642003": "Unfortunately, we have to train the model which reproduces even missing annotation 😧",
    "652283": "Any news on updated annotations? Or do we have to deal with the missing annotations? I attached some more examples of missing annotations i found during training (the green boxes depict predictions here not GT). The problem i think with those cars not annotated is, that they are really obvious positives to the model, while other cars are really not as obvious but are still annotated.",
    "652308": "Sadly, there's no reply from the mods. I've no clue what to do with this."
  },
  "source": "meta"
}