{
  "id": 110086,
  "title": "Regarding is_key_frame and sampling frequency",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/110086",
  "author_name": "Rishabh Agrahari",
  "post_date": "2019-09-24T15:57:33.471000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The <code>sample_data</code> <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/schema.md#sample_data\">description</a> says</p>\n\n<blockquote>\n  <p>For sample_data with is_key_frame=True, the time-stamps should be very close to the sample it points to. For non key-frames the sample_data points to the sample that follows closest in time.</p>\n</blockquote>\n\n<p>My questions:</p>\n\n<ol>\n<li>What's the significance of having this extra information of <code>is_key_frame</code>? How can we use it? </li>\n<li>Regarding sensor synchronization in the dataset, what's the sampling frequency of LIDAR and camera? Is the sensor synchronization methodology and annotation methodology same as that of <a href=\"https://arxiv.org/abs/1903.11027\">NuScenes</a>?</li>\n</ol>\n\n<p>Here's what nuscenes paper says:\n```\n<em>Sensor synchronization</em>: In order to achieve good crossmodality data alignment between the lidar and the cameras, the exposure of a camera is triggered when the top lidar sweeps across the center of the camera’s FOV. The timestamp of the image is the exposure trigger time; and the timestamp of the lidar scan is the time when the full rotation of the current lidar frame is achieved. Given that the camera’s exposure time is nearly instantaneous, this method generally yields good data alignment.</p>\n\n<p><em>Data annotation</em>: Having selected the scenes, we sample keyframes (image, lidar, radar) at 2Hz. ...</p>\n\n<p>The cameras run at 12Hz while the lidar runs at 20Hz. The 12 camera exposures are spread as evenly as possible across the 20 lidar scans, so not all lidar scans have a corresponding camera frame.\n```</p>",
  "messages": [
    {
      "id": 633253,
      "postDate": "2019-09-24T15:57:33.473Z",
      "content": "<p>The <code>sample_data</code> <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/schema.md#sample_data\">description</a> says</p>\n\n<blockquote>\n  <p>For sample_data with is_key_frame=True, the time-stamps should be very close to the sample it points to. For non key-frames the sample_data points to the sample that follows closest in time.</p>\n</blockquote>\n\n<p>My questions:</p>\n\n<ol>\n<li>What's the significance of having this extra information of <code>is_key_frame</code>? How can we use it? </li>\n<li>Regarding sensor synchronization in the dataset, what's the sampling frequency of LIDAR and camera? Is the sensor synchronization methodology and annotation methodology same as that of <a href=\"https://arxiv.org/abs/1903.11027\">NuScenes</a>?</li>\n</ol>\n\n<p>Here's what nuscenes paper says:\n```\n<em>Sensor synchronization</em>: In order to achieve good crossmodality data alignment between the lidar and the cameras, the exposure of a camera is triggered when the top lidar sweeps across the center of the camera’s FOV. The timestamp of the image is the exposure trigger time; and the timestamp of the lidar scan is the time when the full rotation of the current lidar frame is achieved. Given that the camera’s exposure time is nearly instantaneous, this method generally yields good data alignment.</p>\n\n<p><em>Data annotation</em>: Having selected the scenes, we sample keyframes (image, lidar, radar) at 2Hz. ...</p>\n\n<p>The cameras run at 12Hz while the lidar runs at 20Hz. The 12 camera exposures are spread as evenly as possible across the 20 lidar scans, so not all lidar scans have a corresponding camera frame.\n```</p>",
      "rawMarkdown": "The `sample_data` [description](https://github.com/nutonomy/nuscenes-devkit/blob/master/schema.md#sample_data) says\n&gt;For sample\\_data with is\\_key\\_frame=True, the time-stamps should be very close to the sample it points to. For non key-frames the sample\\_data points to the sample that follows closest in time.\n\nMy questions:\n\n1. What's the significance of having this extra information of `is_key_frame`? How can we use it? \n2. Regarding sensor synchronization in the dataset, what's the sampling frequency of LIDAR and camera? Is the sensor synchronization methodology and annotation methodology same as that of [NuScenes](https://arxiv.org/abs/1903.11027)?\n\nHere's what nuscenes paper says:\n```\n*Sensor synchronization*: In order to achieve good crossmodality data alignment between the lidar and the cameras, the exposure of a camera is triggered when the top lidar sweeps across the center of the camera’s FOV. The timestamp of the image is the exposure trigger time; and the timestamp of the lidar scan is the time when the full rotation of the current lidar frame is achieved. Given that the camera’s exposure time is nearly instantaneous, this method generally yields good data alignment.\n\n*Data annotation*: Having selected the scenes, we sample keyframes (image, lidar, radar) at 2Hz. ...\n\nThe cameras run at 12Hz while the lidar runs at 20Hz. The 12 camera exposures are spread as evenly as possible across the 20 lidar scans, so not all lidar scans have a corresponding camera frame.\n```\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "633253": "The `sample_data` [description](https://github.com/nutonomy/nuscenes-devkit/blob/master/schema.md#sample_data) says\n&gt;For sample\\_data with is\\_key\\_frame=True, the time-stamps should be very close to the sample it points to. For non key-frames the sample\\_data points to the sample that follows closest in time.\n\nMy questions:\n\n1. What's the significance of having this extra information of `is_key_frame`? How can we use it? \n2. Regarding sensor synchronization in the dataset, what's the sampling frequency of LIDAR and camera? Is the sensor synchronization methodology and annotation methodology same as that of [NuScenes](https://arxiv.org/abs/1903.11027)?\n\nHere's what nuscenes paper says:\n```\n*Sensor synchronization*: In order to achieve good crossmodality data alignment between the lidar and the cameras, the exposure of a camera is triggered when the top lidar sweeps across the center of the camera’s FOV. The timestamp of the image is the exposure trigger time; and the timestamp of the lidar scan is the time when the full rotation of the current lidar frame is achieved. Given that the camera’s exposure time is nearly instantaneous, this method generally yields good data alignment.\n\n*Data annotation*: Having selected the scenes, we sample keyframes (image, lidar, radar) at 2Hz. ...\n\nThe cameras run at 12Hz while the lidar runs at 20Hz. The 12 camera exposures are spread as evenly as possible across the 20 lidar scans, so not all lidar scans have a corresponding camera frame.\n```\n"
  }
}