{
  "id": 24218,
  "title": "Smaller dataset",
  "url": "/competitions/outbrain-click-prediction/discussion/24218",
  "author_name": "",
  "post_date": "2016-10-09T13:30:54.657Z",
  "votes": null,
  "comment_count": 2,
  "views": 304,
  "content": "<p>Hi,\nI am a hobbyist and participate in competitions to hone my skills. </p>\n\n<p>I need help in getting smaller datasets as I have a PC with only 50 GB free space .  I also want to get randomized records . \nIs there something I can do ?</p>\n\n<p>Help will be highly appreciated </p>",
  "messages": [
    {
      "id": "138543",
      "postDate": "10/09/2016 13:30:54",
      "content": "<p>Hi,\nI am a hobbyist and participate in competitions to hone my skills. </p>\n\n<p>I need help in getting smaller datasets as I have a PC with only 50 GB free space .  I also want to get randomized records . \nIs there something I can do ?</p>\n\n<p>Help will be highly appreciated </p>",
      "rawMarkdown": "Hi,\r\nI am a hobbyist and participate in competitions to hone my skills. \r\n\r\nI need help in getting smaller datasets as I have a PC with only 50 GB free space .  I also want to get randomized records . \r\nIs there something I can do ?\r\n\r\nHelp will be highly appreciated",
      "votes": null
    },
    {
      "id": "138544",
      "postDate": "10/09/2016 13:32:42",
      "content": "<p>First thing to try would be to read in chunks (code should be available in the kernels - either here or in the Bosch contest, which also has a rather large dataset). In particular the pandas package in Python, and <code>read_csv</code> therein - look for the <code>chunksize</code> parameter.</p>",
      "rawMarkdown": "First thing to try would be to read in chunks (code should be available in the kernels - either here or in the Bosch contest, which also has a rather large dataset). In particular the pandas package in Python, and `read_csv` therein - look for the `chunksize` parameter.",
      "votes": null
    },
    {
      "id": "138546",
      "postDate": "10/09/2016 13:47:58",
      "content": "<p>If you use python you can take a look at zipfile library.  It allows you to read files in zip files without having to unzip first. It may take a while to go through the file this way - on my machine it took half an hour just to go through the page view file line by line.</p>",
      "rawMarkdown": "If you use python you can take a look at zipfile library.  It allows you to read files in zip files without having to unzip first. It may take a while to go through the file this way - on my machine it took half an hour just to go through the page view file line by line.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 138544,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "10/09/2016 13:32:42",
      "content": "<p>First thing to try would be to read in chunks (code should be available in the kernels - either here or in the Bosch contest, which also has a rather large dataset). In particular the pandas package in Python, and <code>read_csv</code> therein - look for the <code>chunksize</code> parameter.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138546,
      "author_name": "sangxia",
      "author_url": "",
      "post_date": "10/09/2016 13:47:58",
      "content": "<p>If you use python you can take a look at zipfile library.  It allows you to read files in zip files without having to unzip first. It may take a while to go through the file this way - on my machine it took half an hour just to go through the page view file line by line.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "138543": "Hi,\r\nI am a hobbyist and participate in competitions to hone my skills. \r\n\r\nI need help in getting smaller datasets as I have a PC with only 50 GB free space .  I also want to get randomized records . \r\nIs there something I can do ?\r\n\r\nHelp will be highly appreciated",
    "138544": "First thing to try would be to read in chunks (code should be available in the kernels - either here or in the Bosch contest, which also has a rather large dataset). In particular the pandas package in Python, and `read_csv` therein - look for the `chunksize` parameter.",
    "138546": "If you use python you can take a look at zipfile library.  It allows you to read files in zip files without having to unzip first. It may take a while to go through the file this way - on my machine it took half an hour just to go through the page view file line by line."
  },
  "source": "meta"
}