{
  "id": 110931,
  "title": "split the dataset",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/110931",
  "author_name": "",
  "post_date": "2019-10-02T10:22:56.694306400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello, I'm new here and I'm wondering whether we could split the dataset with offical  SDK into smaller ones, still we'would be okay to load them? Cause the dataset is really big and too heavy for us to debug and test our new thoughts. Thanks in advance!</p>",
  "messages": [
    {
      "id": "638709",
      "postDate": "10/02/2019 10:22:56",
      "content": "<p>Hello, I'm new here and I'm wondering whether we could split the dataset with offical  SDK into smaller ones, still we'would be okay to load them? Cause the dataset is really big and too heavy for us to debug and test our new thoughts. Thanks in advance!</p>",
      "rawMarkdown": "Hello, I'm new here and I'm wondering whether we could split the dataset with offical  SDK into smaller ones, still we'would be okay to load them? Cause the dataset is really big and too heavy for us to debug and test our new thoughts. Thanks in advance!",
      "votes": null
    },
    {
      "id": "638724",
      "postDate": "10/02/2019 10:58:49",
      "content": "<p>You can manually select and split scenes into train/val sets, then use samples only from those splitted scenes. \nSomething similar in nuscenes <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/utils/splits.py\">here</a> </p>",
      "rawMarkdown": "You can manually select and split scenes into train/val sets, then use samples only from those splitted scenes. \nSomething similar in nuscenes [here](https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/utils/splits.py)",
      "votes": null
    },
    {
      "id": "638742",
      "postDate": "10/02/2019 11:28:28",
      "content": "<p>Wow! Thank you for your reply but to my understanding, the train/Val set is still part of the original dataset, that's to say, we still have to deal with the whole dataset, which is too big to store in a laptop, for example. What I really want to do is split out a mini dataset just as you mentioned, like v0.1-mini of nuscenes dataset. Do you have any idea?</p>",
      "rawMarkdown": "Wow! Thank you for your reply but to my understanding, the train/Val set is still part of the original dataset, that's to say, we still have to deal with the whole dataset, which is too big to store in a laptop, for example. What I really want to do is split out a mini dataset just as you mentioned, like v0.1-mini of nuscenes dataset. Do you have any idea?",
      "votes": null
    },
    {
      "id": "638746",
      "postDate": "10/02/2019 11:35:11",
      "content": "<p>Oh okay, so you don't have enough disk space. As of now there's no 'mini' version of lyft yet afaik.</p>",
      "rawMarkdown": "Oh okay, so you don't have enough disk space. As of now there's no 'mini' version of lyft yet afaik.",
      "votes": null
    },
    {
      "id": "638750",
      "postDate": "10/02/2019 11:38:38",
      "content": "<p>Yeah, that's the point</p>",
      "rawMarkdown": "Yeah, that's the point",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 638724,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "10/02/2019 10:58:49",
      "content": "<p>You can manually select and split scenes into train/val sets, then use samples only from those splitted scenes. \nSomething similar in nuscenes <a href=\"https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/utils/splits.py\">here</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 638742,
          "author_name": "alfredocheng",
          "author_url": "",
          "post_date": "10/02/2019 11:28:28",
          "content": "<p>Wow! Thank you for your reply but to my understanding, the train/Val set is still part of the original dataset, that's to say, we still have to deal with the whole dataset, which is too big to store in a laptop, for example. What I really want to do is split out a mini dataset just as you mentioned, like v0.1-mini of nuscenes dataset. Do you have any idea?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 638746,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "10/02/2019 11:35:11",
          "content": "<p>Oh okay, so you don't have enough disk space. As of now there's no 'mini' version of lyft yet afaik.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 638750,
          "author_name": "alfredocheng",
          "author_url": "",
          "post_date": "10/02/2019 11:38:38",
          "content": "<p>Yeah, that's the point</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "638709": "Hello, I'm new here and I'm wondering whether we could split the dataset with offical  SDK into smaller ones, still we'would be okay to load them? Cause the dataset is really big and too heavy for us to debug and test our new thoughts. Thanks in advance!",
    "638724": "You can manually select and split scenes into train/val sets, then use samples only from those splitted scenes. \nSomething similar in nuscenes [here](https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/utils/splits.py)",
    "638742": "Wow! Thank you for your reply but to my understanding, the train/Val set is still part of the original dataset, that's to say, we still have to deal with the whole dataset, which is too big to store in a laptop, for example. What I really want to do is split out a mini dataset just as you mentioned, like v0.1-mini of nuscenes dataset. Do you have any idea?",
    "638746": "Oh okay, so you don't have enough disk space. As of now there's no 'mini' version of lyft yet afaik.",
    "638750": "Yeah, that's the point"
  },
  "source": "meta"
}