{
  "id": 115511,
  "title": "bad cloud?",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/115511",
  "author_name": "",
  "post_date": "2019-11-03T09:30:37.724805900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>has anybody else see some issues with the cloud in file <code>train_lidar/host-a011_lidar1_1233090652702363606.bin</code> ?  the array needs to have a number of elements divisble by 5 for the python code to work (else an error <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/utils/data_classes.py#L286\">https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/utils/data_classes.py#L286</a> )</p>\n\n<p>In the kaggle zip file, I get 265728 values:\n```</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>arr = np.fromfile('/lyft_kaggle_root/train_lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n      arr.shape\n      (265728,)\n      ```</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>But from the Lyft Level 5 download site ( <a href=\"https://level5.lyft.com/dataset/\">https://level5.lyft.com/dataset/</a> -- which also must have this cloud, because it's part of the training set and not the kaggle-only test set), I see 341800 values:\n```</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>arr = np.fromfile('/lyft_level_5_root/train/lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n      arr.shape\n      (341800,)\n      ```</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>It looks like the cloud file might have gotten truncated when added to the Kaggle tarball?</p>\n\n<p>I ran a job over the 30,744 lidar files in the kaggle training set, and this is the only file that causes an error.  </p>\n\n<p>I ran a job over the 27,468 lidar files in the kaggle <em>test</em> set, and everything looks ok (or at least divisible by 5 ...)...</p>",
  "messages": [
    {
      "id": "664178",
      "postDate": "11/03/2019 09:30:37",
      "content": "<p>has anybody else see some issues with the cloud in file <code>train_lidar/host-a011_lidar1_1233090652702363606.bin</code> ?  the array needs to have a number of elements divisble by 5 for the python code to work (else an error <a href=\"https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/utils/data_classes.py#L286\">https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/utils/data_classes.py#L286</a> )</p>\n\n<p>In the kaggle zip file, I get 265728 values:\n```</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>arr = np.fromfile('/lyft_kaggle_root/train_lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n      arr.shape\n      (265728,)\n      ```</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>But from the Lyft Level 5 download site ( <a href=\"https://level5.lyft.com/dataset/\">https://level5.lyft.com/dataset/</a> -- which also must have this cloud, because it's part of the training set and not the kaggle-only test set), I see 341800 values:\n```</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>arr = np.fromfile('/lyft_level_5_root/train/lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n      arr.shape\n      (341800,)\n      ```</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>It looks like the cloud file might have gotten truncated when added to the Kaggle tarball?</p>\n\n<p>I ran a job over the 30,744 lidar files in the kaggle training set, and this is the only file that causes an error.  </p>\n\n<p>I ran a job over the 27,468 lidar files in the kaggle <em>test</em> set, and everything looks ok (or at least divisible by 5 ...)...</p>",
      "rawMarkdown": "has anybody else see some issues with the cloud in file `train_lidar/host-a011_lidar1_1233090652702363606.bin` ?  the array needs to have a number of elements divisble by 5 for the python code to work (else an error https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/utils/data_classes.py#L286 )\n\nIn the kaggle zip file, I get 265728 values:\n```\n&gt;&gt;&gt; arr = np.fromfile('/lyft_kaggle_root/train_lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n&gt;&gt;&gt; arr.shape\n(265728,)\n```\n\nBut from the Lyft Level 5 download site ( https://level5.lyft.com/dataset/ -- which also must have this cloud, because it's part of the training set and not the kaggle-only test set), I see 341800 values:\n```\n&gt;&gt;&gt; arr = np.fromfile('/lyft_level_5_root/train/lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n&gt;&gt;&gt; arr.shape\n(341800,)\n```\n\nIt looks like the cloud file might have gotten truncated when added to the Kaggle tarball?\n\n\nI ran a job over the 30,744 lidar files in the kaggle training set, and this is the only file that causes an error.  \n\nI ran a job over the 27,468 lidar files in the kaggle *test* set, and everything looks ok (or at least divisible by 5 ...)...",
      "votes": null
    },
    {
      "id": "664201",
      "postDate": "11/03/2019 10:17:19",
      "content": "<p>this issue has already been discussed in one of the threads.</p>",
      "rawMarkdown": "this issue has already been discussed in one of the threads.",
      "votes": null
    },
    {
      "id": "664698",
      "postDate": "11/04/2019 04:42:47",
      "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a>  thank you!  derp could you link to that thread? or might you just know if the result was that the kaggle data was updated / fixed?  I admit I appear to have an old copy of the zipfile.  </p>",
      "rawMarkdown": "rishabhiitbhu  thank you!  derp could you link to that thread? or might you just know if the result was that the kaggle data was updated / fixed?  I admit I appear to have an old copy of the zipfile.",
      "votes": null
    },
    {
      "id": "664725",
      "postDate": "11/04/2019 06:08:09",
      "content": "<p>Here's the fixed version:</p>",
      "rawMarkdown": "Here's the fixed version:",
      "votes": null
    },
    {
      "id": "664827",
      "postDate": "11/04/2019 09:29:08",
      "content": "<p>oh thank you! yes I have that though from <a href=\"https://level5.lyft.com/dataset/\">https://level5.lyft.com/dataset/</a>  I was just curious if they fixed the kaggle download perhaps</p>",
      "rawMarkdown": "oh thank you! yes I have that though from https://level5.lyft.com/dataset/  I was just curious if they fixed the kaggle download perhaps",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 664201,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "11/03/2019 10:17:19",
      "content": "<p>this issue has already been discussed in one of the threads.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 664698,
      "author_name": "oarphme",
      "author_url": "",
      "post_date": "11/04/2019 04:42:47",
      "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a>  thank you!  derp could you link to that thread? or might you just know if the result was that the kaggle data was updated / fixed?  I admit I appear to have an old copy of the zipfile.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 664725,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "11/04/2019 06:08:09",
          "content": "<p>Here's the fixed version:</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 664827,
          "author_name": "oarphme",
          "author_url": "",
          "post_date": "11/04/2019 09:29:08",
          "content": "<p>oh thank you! yes I have that though from <a href=\"https://level5.lyft.com/dataset/\">https://level5.lyft.com/dataset/</a>  I was just curious if they fixed the kaggle download perhaps</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "664178": "has anybody else see some issues with the cloud in file `train_lidar/host-a011_lidar1_1233090652702363606.bin` ?  the array needs to have a number of elements divisble by 5 for the python code to work (else an error https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/utils/data_classes.py#L286 )\n\nIn the kaggle zip file, I get 265728 values:\n```\n&gt;&gt;&gt; arr = np.fromfile('/lyft_kaggle_root/train_lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n&gt;&gt;&gt; arr.shape\n(265728,)\n```\n\nBut from the Lyft Level 5 download site ( https://level5.lyft.com/dataset/ -- which also must have this cloud, because it's part of the training set and not the kaggle-only test set), I see 341800 values:\n```\n&gt;&gt;&gt; arr = np.fromfile('/lyft_level_5_root/train/lidar/host-a011_lidar1_1233090652702363606.bin', dtype=np.float32)\n&gt;&gt;&gt; arr.shape\n(341800,)\n```\n\nIt looks like the cloud file might have gotten truncated when added to the Kaggle tarball?\n\n\nI ran a job over the 30,744 lidar files in the kaggle training set, and this is the only file that causes an error.  \n\nI ran a job over the 27,468 lidar files in the kaggle *test* set, and everything looks ok (or at least divisible by 5 ...)...",
    "664201": "this issue has already been discussed in one of the threads.",
    "664698": "rishabhiitbhu  thank you!  derp could you link to that thread? or might you just know if the result was that the kaggle data was updated / fixed?  I admit I appear to have an old copy of the zipfile.",
    "664725": "Here's the fixed version:",
    "664827": "oh thank you! yes I have that though from https://level5.lyft.com/dataset/  I was just curious if they fixed the kaggle download perhaps"
  },
  "source": "meta"
}