{
  "id": 108758,
  "title": "Lyft Dataset SDK",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/108758",
  "author_name": "Vladimir Iglovikov",
  "post_date": "2019-09-13T17:56:13.571000",
  "votes": 73,
  "comment_count": 19,
  "views": 0,
  "content": "<p>In addition to the data, we released  SDK that contains helper and visualization functions. <a href=\"https://github.com/lyft/nuscenes-devkit\">https://github.com/lyft/nuscenes-devkit</a></p>\n\n<p>The SDK is based on the Devkit by the Nuscence team.</p>\n\n<p>I am thrilled to see that the tutorial for the SDK is already ported to Kaggle Kernels. We were planning to do it today, but the community is very fast. :)</p>\n\n<p>I want to ask a favor from you:</p>\n\n<ol>\n<li>If you find bugs/typos/something is unclear/lack of documentation or some other issues.</li>\n<li>If you want to see some functionality that is not there.</li>\n</ol>\n\n<p>Could you please create an issue at <a href=\"https://github.com/lyft/nuscenes-devkit/issues\">https://github.com/lyft/nuscenes-devkit/issues</a> or share it in this thread?</p>\n\n<p>We will try to address them as soon as possible.</p>\n\n<p>P.S. In my free time I develop image augmentation library <a href=\"https://github.com/albu/albumentations\">Albumentations</a>. The impact of the community is tremendous there. I hope something similar can be done to improve the Lyft Dataset SDK.</p>",
  "messages": [
    {
      "id": 626033,
      "postDate": "2019-09-13T17:56:13.570Z",
      "content": "<p>In addition to the data, we released  SDK that contains helper and visualization functions. <a href=\"https://github.com/lyft/nuscenes-devkit\">https://github.com/lyft/nuscenes-devkit</a></p>\n\n<p>The SDK is based on the Devkit by the Nuscence team.</p>\n\n<p>I am thrilled to see that the tutorial for the SDK is already ported to Kaggle Kernels. We were planning to do it today, but the community is very fast. :)</p>\n\n<p>I want to ask a favor from you:</p>\n\n<ol>\n<li>If you find bugs/typos/something is unclear/lack of documentation or some other issues.</li>\n<li>If you want to see some functionality that is not there.</li>\n</ol>\n\n<p>Could you please create an issue at <a href=\"https://github.com/lyft/nuscenes-devkit/issues\">https://github.com/lyft/nuscenes-devkit/issues</a> or share it in this thread?</p>\n\n<p>We will try to address them as soon as possible.</p>\n\n<p>P.S. In my free time I develop image augmentation library <a href=\"https://github.com/albu/albumentations\">Albumentations</a>. The impact of the community is tremendous there. I hope something similar can be done to improve the Lyft Dataset SDK.</p>",
      "rawMarkdown": "In addition to the data, we released  SDK that contains helper and visualization functions. https://github.com/lyft/nuscenes-devkit\n\n\nThe SDK is based on the Devkit by the Nuscence team.\n\nI am thrilled to see that the tutorial for the SDK is already ported to Kaggle Kernels. We were planning to do it today, but the community is very fast. :)\n\n\nI want to ask a favor from you:\n\n1. If you find bugs/typos/something is unclear/lack of documentation or some other issues.\n2. If you want to see some functionality that is not there.\n\nCould you please create an issue at https://github.com/lyft/nuscenes-devkit/issues or share it in this thread?\n\nWe will try to address them as soon as possible.\n\nP.S. In my free time I develop image augmentation library [Albumentations](https://github.com/albu/albumentations). The impact of the community is tremendous there. I hope something similar can be done to improve the Lyft Dataset SDK.",
      "votes": 73
    },
    {
      "id": 626702,
      "postDate": "2019-09-14T17:40:58.207Z",
      "content": "<p>1) functionality - it seems that we can switch to LyftDataset().get simply by token, without specifying the type of table - tokens are still unique.\n2) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F0c847f2eb29f78fdf664aad0c8ffa191%2FScreenshot_1.png?generation=1568482593348897&amp;alt=media\" alt=\"\">\nseems like error in bounding box. Maybe timestamp bug? Did i miss something? token 5eeaa3e4ba898b43996288c6895e8cc38ab3ba6cf7f2ae68112d8b0b936f81e5\n3) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F3420a5e569ce515c2fd3982a981c0290%2FScreenshot_2.png?generation=1568482700186538&amp;alt=media\" alt=\"\">\nseems like bug in visualization. I'm briefly dive into code for render and didn't see sensor translation corrections, but maybe miss in code of SDK. But bug (?) is here - non critical, in visualization. Token 9805befc0348d1f922c4a859d1751573e1a9689d6b34aa50c2416e6386eeee84</p>",
      "rawMarkdown": "1) functionality - it seems that we can switch to LyftDataset().get simply by token, without specifying the type of table - tokens are still unique.\n2) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F0c847f2eb29f78fdf664aad0c8ffa191%2FScreenshot_1.png?generation=1568482593348897&amp;alt=media)\nseems like error in bounding box. Maybe timestamp bug? Did i miss something? token 5eeaa3e4ba898b43996288c6895e8cc38ab3ba6cf7f2ae68112d8b0b936f81e5\n3) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F3420a5e569ce515c2fd3982a981c0290%2FScreenshot_2.png?generation=1568482700186538&amp;alt=media)\nseems like bug in visualization. I'm briefly dive into code for render and didn't see sensor translation corrections, but maybe miss in code of SDK. But bug (?) is here - non critical, in visualization. Token 9805befc0348d1f922c4a859d1751573e1a9689d6b34aa50c2416e6386eeee84",
      "votes": 3
    },
    {
      "id": 631617,
      "postDate": "2019-09-22T10:51:09.443Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2539019%2Fad73b218aedc5fe8182264bcc4a9f79d%2FQQ20190922184729.png?generation=1569149433888030&amp;alt=media\" alt=\"\">\ntwo problems.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2539019%2Fad73b218aedc5fe8182264bcc4a9f79d%2FQQ20190922184729.png?generation=1569149433888030&amp;alt=media)\ntwo problems.",
      "votes": 4,
      "replies": [
        {
          "id": 631878,
          "postDate": "2019-09-22T21:34:21.750Z",
          "content": "<p>Found the same second problem.</p>",
          "rawMarkdown": "Found the same second problem.",
          "votes": 2
        }
      ]
    },
    {
      "id": 664459,
      "postDate": "2019-11-03T17:42:38.963Z",
      "content": "<p>There was a bug pointed in <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/115477\">https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/115477</a> related to the fact that in Kitti, Lyft, Nuscences Lidar is mounted in three different ways.</p>\n\n<p>Pull request with the fix was merged.</p>",
      "rawMarkdown": "There was a bug pointed in https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/115477 related to the fact that in Kitti, Lyft, Nuscences Lidar is mounted in three different ways.\n\nPull request with the fix was merged.",
      "votes": 1
    },
    {
      "id": 662815,
      "postDate": "2019-11-01T03:08:28.930Z",
      "content": "<p>A few questions on the data model:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3747125%2F405864be2c1d6ab6fa79b76f2af4d6e4%2Flyft%20object%20model-2.png?generation=1572707818413006&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Calibrated sensor:</strong>\n1. What do the values inside camera_intrinsic represent? Do they have to be factored into transformations between sensor and car frame of reference</p>\n\n<p><strong>Sample data:</strong>\n1.  What is the significance of \"Is key frame\" attribute? </p>\n\n<p><strong>Sample annotations:</strong>\n1. The annotations appear to reference only lidar data (size, translation and rotation). I follow how humans could use the images as secondary input to annotate the lidar data. But are the images ever directly annotated ? If not, how can we make use of the images as direct training input?\n2. Does num_lidar_points carry any significance? The values appear to be all -1</p>\n\n<p><strong>Attribute</strong>\nDo the object attributes have any impact on the training objectives for this completion (which is limited to identifying objects in a scene with probability)? Maybe identify dust clouds and obstacles?</p>",
      "rawMarkdown": "A few questions on the data model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3747125%2F405864be2c1d6ab6fa79b76f2af4d6e4%2Flyft%20object%20model-2.png?generation=1572707818413006&amp;alt=media)\n\n\n\n**Calibrated sensor:**\n1. What do the values inside camera_intrinsic represent? Do they have to be factored into transformations between sensor and car frame of reference\n\n**Sample data:**\n1.  What is the significance of \"Is key frame\" attribute? \n\n**Sample annotations:**\n1. The annotations appear to reference only lidar data (size, translation and rotation). I follow how humans could use the images as secondary input to annotate the lidar data. But are the images ever directly annotated ? If not, how can we make use of the images as direct training input?\n2. Does num_lidar_points carry any significance? The values appear to be all -1\n\n**Attribute**\nDo the object attributes have any impact on the training objectives for this completion (which is limited to identifying objects in a scene with probability)? Maybe identify dust clouds and obstacles?",
      "votes": 1,
      "replies": [
        {
          "id": 663015,
          "postDate": "2019-11-01T10:47:43.990Z",
          "content": "<p>calibration\ncamera_intrinsic I believe is straight-up the camera matrix K, and included in that record are the extrinsics R|T.  You can project lidar / cuboid points into the camera frame using this calibration.  See the nuscenes or lyft code for examples.</p>\n\n<p>keyframes\nis_key_frame doesn’t carry much meaning in the Lyft dataset, because everything is sampled at 5Hz I think.  In nuscenes, labels are 2Hz, lidar is 20Hz, and cameras are 12Hz. Key Frames are samples for which the sensor timestamps are very well-aligned.  Depending on the speed, the disagreement from temporal differences can be a tenth or a quarter of a meter in the world frame.   Nuscenes by default interpolates cuboid labels using the labels that ‘straddle’ a target timestamp to help correct for this error.  For Lyft, you probably don’t need to worry so much about this because it appears Lyft wanted to only sample and distribute keyframes.</p>\n\n<p>lidar labels\nScale.ai ‘s tool was used to label the Lyft data (according to a scale engineer who went way out of his way on hacker news to make that fact prominent).  I believe Scale’s tool lets the user see both the point clouds <em>and</em> camera images.  Moreover, I believe the labeler creates a cuboid for a track and then just scrubs over the scene and moves the track to where the lidar points appear to be.  </p>\n\n<p>So yeah from the camera’s perspective, the label accuracy is a function of the quality of the camera-lidar calibration (which doesn’t look too bad in general) as well as the time offset between camera and lidar (which I think I found to be always 100ms or less?)</p>\n\n<p>All that said, the contest (and — in the real world — the car’s planning stack) cars most about predictions in the lidar frame, so any camera-based model should probably need to adapt to any camera-lidar disagreement.</p>\n\n<p>I think num lidar points might just be the number of points that fall within the cuboid.  Probably only supported by nuscenes and not Lyft.  In nuscenes, they only have a single velodyne 64, so it’s common for them to only get a few points on cars / pedestrians, and so this stat is interesting during evaluation.  (Also in real life— if you only get a couple of points on a pedestrian, you have a very real chance of screwing up and hitting them.  That’s why most companies use safety drivers).</p>\n\n<p>attributes\nI believe this competition is only about category, not attribute?\nIn nuscenes, you have to look at the attributes to see if a bicycle / motorcycle has a person riding it or not.  I think in Lyft I have not seen them use the attributes or not a bike without a rider.  In Waymo Open, they sadly only label bikes that have riders, leaving parked bikes completely unlabeled.  (The Waymo dataset is pretty sad in terms of research potential and was recorded in 2017; Lyft is from 2019).  While for this competition I don’t think the rider state matters, in the real world it can make a big difference.  For example, one might have some dumb motion prediction system that says bikes with riders will move in the lane while bikes without riders will remain static, so you’d want to be able to detect static bikes and/or evaluate you bike prediction as a function of rider vs no rider.</p>",
          "rawMarkdown": "calibration\ncamera_intrinsic I believe is straight-up the camera matrix K, and included in that record are the extrinsics R|T.  You can project lidar / cuboid points into the camera frame using this calibration.  See the nuscenes or lyft code for examples.\n\nkeyframes\nis_key_frame doesn’t carry much meaning in the Lyft dataset, because everything is sampled at 5Hz I think.  In nuscenes, labels are 2Hz, lidar is 20Hz, and cameras are 12Hz. Key Frames are samples for which the sensor timestamps are very well-aligned.  Depending on the speed, the disagreement from temporal differences can be a tenth or a quarter of a meter in the world frame.   Nuscenes by default interpolates cuboid labels using the labels that ‘straddle’ a target timestamp to help correct for this error.  For Lyft, you probably don’t need to worry so much about this because it appears Lyft wanted to only sample and distribute keyframes.\n\nlidar labels\nScale.ai ‘s tool was used to label the Lyft data (according to a scale engineer who went way out of his way on hacker news to make that fact prominent).  I believe Scale’s tool lets the user see both the point clouds *and* camera images.  Moreover, I believe the labeler creates a cuboid for a track and then just scrubs over the scene and moves the track to where the lidar points appear to be.  \n\nSo yeah from the camera’s perspective, the label accuracy is a function of the quality of the camera-lidar calibration (which doesn’t look too bad in general) as well as the time offset between camera and lidar (which I think I found to be always 100ms or less?)\n\nAll that said, the contest (and — in the real world — the car’s planning stack) cars most about predictions in the lidar frame, so any camera-based model should probably need to adapt to any camera-lidar disagreement.\n\nI think num lidar points might just be the number of points that fall within the cuboid.  Probably only supported by nuscenes and not Lyft.  In nuscenes, they only have a single velodyne 64, so it’s common for them to only get a few points on cars / pedestrians, and so this stat is interesting during evaluation.  (Also in real life— if you only get a couple of points on a pedestrian, you have a very real chance of screwing up and hitting them.  That’s why most companies use safety drivers).\n\nattributes\nI believe this competition is only about category, not attribute?\nIn nuscenes, you have to look at the attributes to see if a bicycle / motorcycle has a person riding it or not.  I think in Lyft I have not seen them use the attributes or not a bike without a rider.  In Waymo Open, they sadly only label bikes that have riders, leaving parked bikes completely unlabeled.  (The Waymo dataset is pretty sad in terms of research potential and was recorded in 2017; Lyft is from 2019).  While for this competition I don’t think the rider state matters, in the real world it can make a big difference.  For example, one might have some dumb motion prediction system that says bikes with riders will move in the lane while bikes without riders will remain static, so you’d want to be able to detect static bikes and/or evaluate you bike prediction as a function of rider vs no rider.\n\n\n",
          "votes": 1
        },
        {
          "id": 663063,
          "postDate": "2019-11-01T12:08:19.523Z",
          "content": "<p>Great explanation. I do not fully understand the points you make about lidar labels,  but my take aways are  a) lidar trumps camera b) any learning from camera data, would have already been acquired from the lidar data and c)  as a newbie, it would be simpler to just focus on the lidar data for now.</p>",
          "rawMarkdown": "Great explanation. I do not fully understand the points you make about lidar labels,  but my take aways are  a) lidar trumps camera b) any learning from camera data, would have already been acquired from the lidar data and c)  as a newbie, it would be simpler to just focus on the lidar data for now."
        },
        {
          "id": 663482,
          "postDate": "2019-11-02T06:11:56.113Z",
          "content": "<p>eep!  well starting with a single sensor modality isn't a bad idea (combining the two modalities is an open area of research).  since this contest asks for 3D pose, you're likely going to estimate that better with lidar (modulo pi -- vision is pretty important for determining front vs back).  plus there are some shared notebooks showing how to handle the lidar data.</p>\n\n<p>To give a little more color on the lidar labels: if you look at the render_sample() function, it will render labels in both lidar and camera images.  The tool used to label the Lyft images let users draw boxes in 3d (lidar space) and likely also allowed them to immediately see their box as projected into the cameras.  This is important for things with a very small number of returns.  So in other words: the cuboids should be pretty good to use as camera labels, but they likely won't be as accurate as, say, the image bounding boxes found in Waymo Open, BDD100k, MSCOCO, etc, where those labels are explicitly just for images.  </p>",
          "rawMarkdown": "eep!  well starting with a single sensor modality isn't a bad idea (combining the two modalities is an open area of research).  since this contest asks for 3D pose, you're likely going to estimate that better with lidar (modulo pi -- vision is pretty important for determining front vs back).  plus there are some shared notebooks showing how to handle the lidar data.\n\nTo give a little more color on the lidar labels: if you look at the render_sample() function, it will render labels in both lidar and camera images.  The tool used to label the Lyft images let users draw boxes in 3d (lidar space) and likely also allowed them to immediately see their box as projected into the cameras.  This is important for things with a very small number of returns.  So in other words: the cuboids should be pretty good to use as camera labels, but they likely won't be as accurate as, say, the image bounding boxes found in Waymo Open, BDD100k, MSCOCO, etc, where those labels are explicitly just for images.  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 638660,
      "postDate": "2019-10-02T08:24:08.487Z",
      "content": "<p>Based upon the output of <code>list_scenes()</code>, it looks like the Kaggle train split consists of the same exact 180 scenes as previously released in the Level 5 Dataset earlier this year: <a href=\"https://level5.lyft.com/dataset/\">https://level5.lyft.com/dataset/</a>    Can we get any confirmation?   Just curious because then any work on the original release is transferrable to the Kaggle venue.</p>\n\n<p>And the Kaggle test split appears to be an entirely new data release of 218 scenes (but from the same cars and a slightly narrower time range; the train split has data from May 2019).  </p>",
      "rawMarkdown": "Based upon the output of `list_scenes()`, it looks like the Kaggle train split consists of the same exact 180 scenes as previously released in the Level 5 Dataset earlier this year: https://level5.lyft.com/dataset/    Can we get any confirmation?   Just curious because then any work on the original release is transferrable to the Kaggle venue.\n\nAnd the Kaggle test split appears to be an entirely new data release of 218 scenes (but from the same cars and a slightly narrower time range; the train split has data from May 2019).  ",
      "votes": 1,
      "replies": [
        {
          "id": 638843,
          "postDate": "2019-10-02T13:34:58.487Z",
          "content": "<p>This is correct. Train on the released dataset and Kaggle are the same.</p>\n\n<p>The test set is new.</p>",
          "rawMarkdown": "This is correct. Train on the released dataset and Kaggle are the same.\n\nThe test set is new.",
          "votes": 1
        }
      ]
    },
    {
      "id": 631415,
      "postDate": "2019-09-22T01:27:28.680Z",
      "content": "<p>Many thanks to you and the Lyft Level 5 team for putting this competition together.</p>\n\n<p>Are you able to tell us a bit about the hardware used for training models at Lyft Level 5? Not necessarily the models that will be deployed on the vehicle but for day to day experimentation and development on similar sized datasets as this competition. </p>",
      "rawMarkdown": "Many thanks to you and the Lyft Level 5 team for putting this competition together.\n\nAre you able to tell us a bit about the hardware used for training models at Lyft Level 5? Not necessarily the models that will be deployed on the vehicle but for day to day experimentation and development on similar sized datasets as this competition. ",
      "replies": [
        {
          "id": 631459,
          "postDate": "2019-09-22T04:10:22.180Z",
          "content": "<p>Thank you for the question.</p>\n\n<p>I would love to tell you the details, but I do not think that I am not allowed to do this.</p>\n\n<p>I would ask internally, and maybe we will write a blog post covering this topic.</p>",
          "rawMarkdown": "Thank you for the question.\n\nI would love to tell you the details, but I do not think that I am not allowed to do this.\n\nI would ask internally, and maybe we will write a blog post covering this topic.\n",
          "votes": 4
        },
        {
          "id": 631723,
          "postDate": "2019-09-22T14:49:18.587Z",
          "content": "<p>No worries, it would be great to read a blog post about it in the future! </p>\n\n<p>I'm mainly trying to determine if my hardware is anywhere near capable enough to train a model on this dataset and I'm also generally curious/interested in the hardware being used at Lyft Level 5.</p>\n\n<p>I have a GTX 1080 (with 8gb ram), 6 core i5 cpu and 32gb ram, which I think/hope should be enough but guessing it will be quite slow. I'm waiting for a new SSD to arrive today so haven't downloaded the dataset and tried to train a model yet, so I will find out soon enough if it is up to the task!</p>",
          "rawMarkdown": "No worries, it would be great to read a blog post about it in the future! \n\nI'm mainly trying to determine if my hardware is anywhere near capable enough to train a model on this dataset and I'm also generally curious/interested in the hardware being used at Lyft Level 5.\n\nI have a GTX 1080 (with 8gb ram), 6 core i5 cpu and 32gb ram, which I think/hope should be enough but guessing it will be quite slow. I'm waiting for a new SSD to arrive today so haven't downloaded the dataset and tried to train a model yet, so I will find out soon enough if it is up to the task!",
          "votes": 2
        },
        {
          "id": 631883,
          "postDate": "2019-09-22T21:46:14.027Z",
          "content": "<p>I'm playing around with PointRCNN (one of the SOTA arch.), I'm able to train it with batch size 2 on RTX 2070 (8GB GPU RAM), though it heavily depends on the training configuration. This is without apex, I think I'll be able to use higher batch size with it.</p>",
          "rawMarkdown": "I'm playing around with PointRCNN (one of the SOTA arch.), I'm able to train it with batch size 2 on RTX 2070 (8GB GPU RAM), though it heavily depends on the training configuration. This is without apex, I think I'll be able to use higher batch size with it.",
          "votes": 3
        },
        {
          "id": 632553,
          "postDate": "2019-09-23T18:29:46.710Z",
          "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> thank you, that gives me some hope 😄  The lidar data in this competition is also 3d point cloud? </p>",
          "rawMarkdown": "@rishabhiitbhu thank you, that gives me some hope 😄  The lidar data in this competition is also 3d point cloud? ",
          "votes": 2
        },
        {
          "id": 632584,
          "postDate": "2019-09-23T19:04:52.527Z",
          "content": "<p>Yup it's 3D, about ~ 70k points 🤓 </p>",
          "rawMarkdown": "Yup it's 3D, about ~ 70k points 🤓 ",
          "votes": 1
        },
        {
          "id": 632766,
          "postDate": "2019-09-24T03:15:47.490Z",
          "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> I played around with the reference model on my machine this evening and was able to train with batch size of 16! (~ 7/8 GB memory) 😀</p>",
          "rawMarkdown": "@rishabhiitbhu I played around with the reference model on my machine this evening and was able to train with batch size of 16! (~ 7/8 GB memory) 😀",
          "votes": 1
        }
      ]
    },
    {
      "id": 660958,
      "postDate": "2019-10-29T20:01:50.110Z",
      "content": "<p>Hi,</p>\n\n<p>I tried to render 3d sample using:\n<code>lyft_dataset.render_sample_3d_interactive(lyft_dataset.sample[0][\"token\"])</code>\nI get a lot of text printed. I was hoping to get some kind of 3D interactive visualization. Is this expected or do I need to include something to see the visualization?</p>",
      "rawMarkdown": "Hi,\n\nI tried to render 3d sample using:\n`lyft_dataset.render_sample_3d_interactive(lyft_dataset.sample[0][\"token\"])`\nI get a lot of text printed. I was hoping to get some kind of 3D interactive visualization. Is this expected or do I need to include something to see the visualization?"
    },
    {
      "id": 626097,
      "postDate": "2019-09-13T20:01:46.250Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 626702,
      "author_name": "Igor Kotenkov",
      "author_url": "",
      "post_date": "2019-09-14T17:40:58.207000",
      "content": "<p>1) functionality - it seems that we can switch to LyftDataset().get simply by token, without specifying the type of table - tokens are still unique.\n2) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F0c847f2eb29f78fdf664aad0c8ffa191%2FScreenshot_1.png?generation=1568482593348897&amp;alt=media\" alt=\"\">\nseems like error in bounding box. Maybe timestamp bug? Did i miss something? token 5eeaa3e4ba898b43996288c6895e8cc38ab3ba6cf7f2ae68112d8b0b936f81e5\n3) <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F3420a5e569ce515c2fd3982a981c0290%2FScreenshot_2.png?generation=1568482700186538&amp;alt=media\" alt=\"\">\nseems like bug in visualization. I'm briefly dive into code for render and didn't see sensor translation corrections, but maybe miss in code of SDK. But bug (?) is here - non critical, in visualization. Token 9805befc0348d1f922c4a859d1751573e1a9689d6b34aa50c2416e6386eeee84</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 631617,
      "author_name": "gakki",
      "author_url": "",
      "post_date": "2019-09-22T10:51:09.443000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2539019%2Fad73b218aedc5fe8182264bcc4a9f79d%2FQQ20190922184729.png?generation=1569149433888030&amp;alt=media\" alt=\"\">\ntwo problems.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 631878,
          "author_name": "Hanxiao Deng",
          "author_url": "",
          "post_date": "2019-09-22T21:34:21.750000",
          "content": "<p>Found the same second problem.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 664459,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2019-11-03T17:42:38.963000",
      "content": "<p>There was a bug pointed in <a href=\"https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/115477\">https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/115477</a> related to the fact that in Kitti, Lyft, Nuscences Lidar is mounted in three different ways.</p>\n\n<p>Pull request with the fix was merged.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 662815,
      "author_name": "Rajaram G",
      "author_url": "",
      "post_date": "2019-11-01T03:08:28.930000",
      "content": "<p>A few questions on the data model:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3747125%2F405864be2c1d6ab6fa79b76f2af4d6e4%2Flyft%20object%20model-2.png?generation=1572707818413006&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Calibrated sensor:</strong>\n1. What do the values inside camera_intrinsic represent? Do they have to be factored into transformations between sensor and car frame of reference</p>\n\n<p><strong>Sample data:</strong>\n1.  What is the significance of \"Is key frame\" attribute? </p>\n\n<p><strong>Sample annotations:</strong>\n1. The annotations appear to reference only lidar data (size, translation and rotation). I follow how humans could use the images as secondary input to annotate the lidar data. But are the images ever directly annotated ? If not, how can we make use of the images as direct training input?\n2. Does num_lidar_points carry any significance? The values appear to be all -1</p>\n\n<p><strong>Attribute</strong>\nDo the object attributes have any impact on the training objectives for this completion (which is limited to identifying objects in a scene with probability)? Maybe identify dust clouds and obstacles?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 663015,
          "author_name": "oarph",
          "author_url": "",
          "post_date": "2019-11-01T10:47:43.990000",
          "content": "<p>calibration\ncamera_intrinsic I believe is straight-up the camera matrix K, and included in that record are the extrinsics R|T.  You can project lidar / cuboid points into the camera frame using this calibration.  See the nuscenes or lyft code for examples.</p>\n\n<p>keyframes\nis_key_frame doesn’t carry much meaning in the Lyft dataset, because everything is sampled at 5Hz I think.  In nuscenes, labels are 2Hz, lidar is 20Hz, and cameras are 12Hz. Key Frames are samples for which the sensor timestamps are very well-aligned.  Depending on the speed, the disagreement from temporal differences can be a tenth or a quarter of a meter in the world frame.   Nuscenes by default interpolates cuboid labels using the labels that ‘straddle’ a target timestamp to help correct for this error.  For Lyft, you probably don’t need to worry so much about this because it appears Lyft wanted to only sample and distribute keyframes.</p>\n\n<p>lidar labels\nScale.ai ‘s tool was used to label the Lyft data (according to a scale engineer who went way out of his way on hacker news to make that fact prominent).  I believe Scale’s tool lets the user see both the point clouds <em>and</em> camera images.  Moreover, I believe the labeler creates a cuboid for a track and then just scrubs over the scene and moves the track to where the lidar points appear to be.  </p>\n\n<p>So yeah from the camera’s perspective, the label accuracy is a function of the quality of the camera-lidar calibration (which doesn’t look too bad in general) as well as the time offset between camera and lidar (which I think I found to be always 100ms or less?)</p>\n\n<p>All that said, the contest (and — in the real world — the car’s planning stack) cars most about predictions in the lidar frame, so any camera-based model should probably need to adapt to any camera-lidar disagreement.</p>\n\n<p>I think num lidar points might just be the number of points that fall within the cuboid.  Probably only supported by nuscenes and not Lyft.  In nuscenes, they only have a single velodyne 64, so it’s common for them to only get a few points on cars / pedestrians, and so this stat is interesting during evaluation.  (Also in real life— if you only get a couple of points on a pedestrian, you have a very real chance of screwing up and hitting them.  That’s why most companies use safety drivers).</p>\n\n<p>attributes\nI believe this competition is only about category, not attribute?\nIn nuscenes, you have to look at the attributes to see if a bicycle / motorcycle has a person riding it or not.  I think in Lyft I have not seen them use the attributes or not a bike without a rider.  In Waymo Open, they sadly only label bikes that have riders, leaving parked bikes completely unlabeled.  (The Waymo dataset is pretty sad in terms of research potential and was recorded in 2017; Lyft is from 2019).  While for this competition I don’t think the rider state matters, in the real world it can make a big difference.  For example, one might have some dumb motion prediction system that says bikes with riders will move in the lane while bikes without riders will remain static, so you’d want to be able to detect static bikes and/or evaluate you bike prediction as a function of rider vs no rider.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 663063,
          "author_name": "Rajaram G",
          "author_url": "",
          "post_date": "2019-11-01T12:08:19.523000",
          "content": "<p>Great explanation. I do not fully understand the points you make about lidar labels,  but my take aways are  a) lidar trumps camera b) any learning from camera data, would have already been acquired from the lidar data and c)  as a newbie, it would be simpler to just focus on the lidar data for now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 663482,
          "author_name": "oarph",
          "author_url": "",
          "post_date": "2019-11-02T06:11:56.113000",
          "content": "<p>eep!  well starting with a single sensor modality isn't a bad idea (combining the two modalities is an open area of research).  since this contest asks for 3D pose, you're likely going to estimate that better with lidar (modulo pi -- vision is pretty important for determining front vs back).  plus there are some shared notebooks showing how to handle the lidar data.</p>\n\n<p>To give a little more color on the lidar labels: if you look at the render_sample() function, it will render labels in both lidar and camera images.  The tool used to label the Lyft images let users draw boxes in 3d (lidar space) and likely also allowed them to immediately see their box as projected into the cameras.  This is important for things with a very small number of returns.  So in other words: the cuboids should be pretty good to use as camera labels, but they likely won't be as accurate as, say, the image bounding boxes found in Waymo Open, BDD100k, MSCOCO, etc, where those labels are explicitly just for images.  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 638660,
      "author_name": "oarph",
      "author_url": "",
      "post_date": "2019-10-02T08:24:08.487000",
      "content": "<p>Based upon the output of <code>list_scenes()</code>, it looks like the Kaggle train split consists of the same exact 180 scenes as previously released in the Level 5 Dataset earlier this year: <a href=\"https://level5.lyft.com/dataset/\">https://level5.lyft.com/dataset/</a>    Can we get any confirmation?   Just curious because then any work on the original release is transferrable to the Kaggle venue.</p>\n\n<p>And the Kaggle test split appears to be an entirely new data release of 218 scenes (but from the same cars and a slightly narrower time range; the train split has data from May 2019).  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 638843,
          "author_name": "Vladimir Iglovikov",
          "author_url": "",
          "post_date": "2019-10-02T13:34:58.487000",
          "content": "<p>This is correct. Train on the released dataset and Kaggle are the same.</p>\n\n<p>The test set is new.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 631415,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2019-09-22T01:27:28.680000",
      "content": "<p>Many thanks to you and the Lyft Level 5 team for putting this competition together.</p>\n\n<p>Are you able to tell us a bit about the hardware used for training models at Lyft Level 5? Not necessarily the models that will be deployed on the vehicle but for day to day experimentation and development on similar sized datasets as this competition. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 631459,
          "author_name": "Vladimir Iglovikov",
          "author_url": "",
          "post_date": "2019-09-22T04:10:22.180000",
          "content": "<p>Thank you for the question.</p>\n\n<p>I would love to tell you the details, but I do not think that I am not allowed to do this.</p>\n\n<p>I would ask internally, and maybe we will write a blog post covering this topic.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 631723,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-09-22T14:49:18.587000",
          "content": "<p>No worries, it would be great to read a blog post about it in the future! </p>\n\n<p>I'm mainly trying to determine if my hardware is anywhere near capable enough to train a model on this dataset and I'm also generally curious/interested in the hardware being used at Lyft Level 5.</p>\n\n<p>I have a GTX 1080 (with 8gb ram), 6 core i5 cpu and 32gb ram, which I think/hope should be enough but guessing it will be quite slow. I'm waiting for a new SSD to arrive today so haven't downloaded the dataset and tried to train a model yet, so I will find out soon enough if it is up to the task!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 631883,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-09-22T21:46:14.027000",
          "content": "<p>I'm playing around with PointRCNN (one of the SOTA arch.), I'm able to train it with batch size 2 on RTX 2070 (8GB GPU RAM), though it heavily depends on the training configuration. This is without apex, I think I'll be able to use higher batch size with it.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 632553,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-09-23T18:29:46.710000",
          "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> thank you, that gives me some hope 😄  The lidar data in this competition is also 3d point cloud? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 632584,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-09-23T19:04:52.527000",
          "content": "<p>Yup it's 3D, about ~ 70k points 🤓 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 632766,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-09-24T03:15:47.490000",
          "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> I played around with the reference model on my machine this evening and was able to train with batch size of 16! (~ 7/8 GB memory) 😀</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 660958,
      "author_name": "AshutoshNirala",
      "author_url": "",
      "post_date": "2019-10-29T20:01:50.110000",
      "content": "<p>Hi,</p>\n\n<p>I tried to render 3d sample using:\n<code>lyft_dataset.render_sample_3d_interactive(lyft_dataset.sample[0][\"token\"])</code>\nI get a lot of text printed. I was hoping to get some kind of 3D interactive visualization. Is this expected or do I need to include something to see the visualization?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 626097,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-13T20:01:46.250000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "626033": "In addition to the data, we released  SDK that contains helper and visualization functions. https://github.com/lyft/nuscenes-devkit\n\n\nThe SDK is based on the Devkit by the Nuscence team.\n\nI am thrilled to see that the tutorial for the SDK is already ported to Kaggle Kernels. We were planning to do it today, but the community is very fast. :)\n\n\nI want to ask a favor from you:\n\n1. If you find bugs/typos/something is unclear/lack of documentation or some other issues.\n2. If you want to see some functionality that is not there.\n\nCould you please create an issue at https://github.com/lyft/nuscenes-devkit/issues or share it in this thread?\n\nWe will try to address them as soon as possible.\n\nP.S. In my free time I develop image augmentation library [Albumentations](https://github.com/albu/albumentations). The impact of the community is tremendous there. I hope something similar can be done to improve the Lyft Dataset SDK.",
    "626702": "1) functionality - it seems that we can switch to LyftDataset().get simply by token, without specifying the type of table - tokens are still unique.\n2) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F0c847f2eb29f78fdf664aad0c8ffa191%2FScreenshot_1.png?generation=1568482593348897&amp;alt=media)\nseems like error in bounding box. Maybe timestamp bug? Did i miss something? token 5eeaa3e4ba898b43996288c6895e8cc38ab3ba6cf7f2ae68112d8b0b936f81e5\n3) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F3420a5e569ce515c2fd3982a981c0290%2FScreenshot_2.png?generation=1568482700186538&amp;alt=media)\nseems like bug in visualization. I'm briefly dive into code for render and didn't see sensor translation corrections, but maybe miss in code of SDK. But bug (?) is here - non critical, in visualization. Token 9805befc0348d1f922c4a859d1751573e1a9689d6b34aa50c2416e6386eeee84",
    "631617": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2539019%2Fad73b218aedc5fe8182264bcc4a9f79d%2FQQ20190922184729.png?generation=1569149433888030&amp;alt=media)\ntwo problems.",
    "664459": "There was a bug pointed in https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/discussion/115477 related to the fact that in Kitti, Lyft, Nuscences Lidar is mounted in three different ways.\n\nPull request with the fix was merged.",
    "662815": "A few questions on the data model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3747125%2F405864be2c1d6ab6fa79b76f2af4d6e4%2Flyft%20object%20model-2.png?generation=1572707818413006&amp;alt=media)\n\n\n\n**Calibrated sensor:**\n1. What do the values inside camera_intrinsic represent? Do they have to be factored into transformations between sensor and car frame of reference\n\n**Sample data:**\n1.  What is the significance of \"Is key frame\" attribute? \n\n**Sample annotations:**\n1. The annotations appear to reference only lidar data (size, translation and rotation). I follow how humans could use the images as secondary input to annotate the lidar data. But are the images ever directly annotated ? If not, how can we make use of the images as direct training input?\n2. Does num_lidar_points carry any significance? The values appear to be all -1\n\n**Attribute**\nDo the object attributes have any impact on the training objectives for this completion (which is limited to identifying objects in a scene with probability)? Maybe identify dust clouds and obstacles?",
    "638660": "Based upon the output of `list_scenes()`, it looks like the Kaggle train split consists of the same exact 180 scenes as previously released in the Level 5 Dataset earlier this year: https://level5.lyft.com/dataset/    Can we get any confirmation?   Just curious because then any work on the original release is transferrable to the Kaggle venue.\n\nAnd the Kaggle test split appears to be an entirely new data release of 218 scenes (but from the same cars and a slightly narrower time range; the train split has data from May 2019).  ",
    "631415": "Many thanks to you and the Lyft Level 5 team for putting this competition together.\n\nAre you able to tell us a bit about the hardware used for training models at Lyft Level 5? Not necessarily the models that will be deployed on the vehicle but for day to day experimentation and development on similar sized datasets as this competition. ",
    "660958": "Hi,\n\nI tried to render 3d sample using:\n`lyft_dataset.render_sample_3d_interactive(lyft_dataset.sample[0][\"token\"])`\nI get a lot of text printed. I was hoping to get some kind of 3D interactive visualization. Is this expected or do I need to include something to see the visualization?",
    "626097": ""
  }
}