{
  "id": 20758,
  "title": "How do you handle the nearly 4gb train data",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20758",
  "author_name": "",
  "post_date": "2016-05-06T06:42:25.687Z",
  "votes": 1,
  "comment_count": 7,
  "views": 2081,
  "content": "<p>I see the train data size is nearly 4gb, it even takes a lot of time to read train.csv using pandas, how do model the train data? do you use any sampling methods to extract part of the data or use any distributed platforms like hadoop, spark?</p>",
  "messages": [
    {
      "id": "118927",
      "postDate": "05/06/2016 06:42:25",
      "content": "<p>I see the train data size is nearly 4gb, it even takes a lot of time to read train.csv using pandas, how do model the train data? do you use any sampling methods to extract part of the data or use any distributed platforms like hadoop, spark?</p>",
      "rawMarkdown": "I see the train data size is nearly 4gb, it even takes a lot of time to read train.csv using pandas, how do model the train data? do you use any sampling methods to extract part of the data or use any distributed platforms like hadoop, spark?",
      "votes": null
    },
    {
      "id": "118965",
      "postDate": "05/06/2016 12:23:31",
      "content": "<ol>\n<li>Data Reduction</li>\n<li>Online Learning</li>\n<li>Cloud computing on Amazon</li>\n<li>Hope your birthday is coming up and buy a bigger computer ;)</li>\n</ol>",
      "rawMarkdown": "1. Data Reduction\r\n2. Online Learning\r\n3. Cloud computing on Amazon\r\n4. Hope your birthday is coming up and buy a bigger computer ;)",
      "votes": null
    },
    {
      "id": "118968",
      "postDate": "05/06/2016 12:57:39",
      "content": "<p>Take a look at <a href=\"https://www.kaggle.com/dvasyukova/expedia-hotel-recommendations/predict-hotel-type-with-pandas\">this script</a> which is basically suggestion #1.  In this case your training data can be aggregated.  For your feature exploration, take a sampling as described in the tutorial <a href=\"https://www.dataquest.io/blog/kaggle-tutorial/\">here</a></p>\n\n<p>Both links are courtesy of forum dwellers here.</p>",
      "rawMarkdown": "Take a look at [this script][1] which is basically suggestion #1.  In this case your training data can be aggregated.  For your feature exploration, take a sampling as described in the tutorial [here][2]\r\n\r\nBoth links are courtesy of forum dwellers here.\r\n\r\n\r\n  [1]: https://www.kaggle.com/dvasyukova/expedia-hotel-recommendations/predict-hotel-type-with-pandas\r\n  [2]: https://www.dataquest.io/blog/kaggle-tutorial/",
      "votes": null
    },
    {
      "id": "119099",
      "postDate": "05/07/2016 07:16:04",
      "content": "<p>I went ahead with pandas load_csv on  my MacAir with 4GB memory. Took about 7 mins to load the csv file. Then I ran a downsample of the data set . Works like a charm !! </p>",
      "rawMarkdown": "I went ahead with pandas load_csv on  my MacAir with 4GB memory. Took about 7 mins to load the csv file. Then I ran a downsample of the data set . Works like a charm !!",
      "votes": null
    },
    {
      "id": "119115",
      "postDate": "05/07/2016 11:50:29",
      "content": "<p>There is another thread in the forum covering that topic. From what I see in the scripts and in my own experience it seems like the way to go is online learning (or parsing the files line by line). Not sure if anyone used Machine Learning successfully yet.</p>\n\n<p>3 is only an option if money is not an issue, I think...</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "There is another thread in the forum covering that topic. From what I see in the scripts and in my own experience it seems like the way to go is online learning (or parsing the files line by line). Not sure if anyone used Machine Learning successfully yet.\r\n\r\n3 is only an option if money is not an issue, I think...\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "119298",
      "postDate": "05/08/2016 23:50:12",
      "content": "<p>I'm using GraphLab (which claims to be more efficient) and AWS. Using spot pricing I can get an 8 core, 30GB memory server for about 6c/hr. Since I only use it a few hours per day its cost effective.  </p>",
      "rawMarkdown": "I'm using GraphLab (which claims to be more efficient) and AWS. Using spot pricing I can get an 8 core, 30GB memory server for about 6c/hr. Since I only use it a few hours per day its cost effective.",
      "votes": null
    },
    {
      "id": "120200",
      "postDate": "05/16/2016 09:57:59",
      "content": "<p>I use a super slow machine and graphlab - data reading works brilliantly well. \nMy machine is an old, rickety pentium with 4GB RAM, all is well. \nexp_train = graphlab.SFrame(&quot;d:\\datasets\\kaggle-expedia\\train.csv&quot;)\n(SFrame being disk-backed, it works quite well). Do explore, gets all the data loading and model training parts out of the way easily. </p>",
      "rawMarkdown": "I use a super slow machine and graphlab - data reading works brilliantly well. \r\nMy machine is an old, rickety pentium with 4GB RAM, all is well. \r\nexp_train = graphlab.SFrame(\"d:\\\\datasets\\\\kaggle-expedia\\\\train.csv\")\r\n(SFrame being disk-backed, it works quite well). Do explore, gets all the data loading and model training parts out of the way easily.",
      "votes": null
    },
    {
      "id": "120205",
      "postDate": "05/16/2016 11:29:24",
      "content": "<p>I'm using Python with blaze + dask + bcolz for working on the full data set.</p>\n\n<p>Things can go a bit slow from time to time, but avoids running out of memory.</p>",
      "rawMarkdown": "I'm using Python with blaze + dask + bcolz for working on the full data set.\r\n\r\nThings can go a bit slow from time to time, but avoids running out of memory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118965,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "05/06/2016 12:23:31",
      "content": "<ol>\n<li>Data Reduction</li>\n<li>Online Learning</li>\n<li>Cloud computing on Amazon</li>\n<li>Hope your birthday is coming up and buy a bigger computer ;)</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118968,
      "author_name": "rdslater",
      "author_url": "",
      "post_date": "05/06/2016 12:57:39",
      "content": "<p>Take a look at <a href=\"https://www.kaggle.com/dvasyukova/expedia-hotel-recommendations/predict-hotel-type-with-pandas\">this script</a> which is basically suggestion #1.  In this case your training data can be aggregated.  For your feature exploration, take a sampling as described in the tutorial <a href=\"https://www.dataquest.io/blog/kaggle-tutorial/\">here</a></p>\n\n<p>Both links are courtesy of forum dwellers here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119099,
      "author_name": "karandewan",
      "author_url": "",
      "post_date": "05/07/2016 07:16:04",
      "content": "<p>I went ahead with pandas load_csv on  my MacAir with 4GB memory. Took about 7 mins to load the csv file. Then I ran a downsample of the data set . Works like a charm !! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119115,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/07/2016 11:50:29",
      "content": "<p>There is another thread in the forum covering that topic. From what I see in the scripts and in my own experience it seems like the way to go is online learning (or parsing the files line by line). Not sure if anyone used Machine Learning successfully yet.</p>\n\n<p>3 is only an option if money is not an issue, I think...</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119298,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "05/08/2016 23:50:12",
      "content": "<p>I'm using GraphLab (which claims to be more efficient) and AWS. Using spot pricing I can get an 8 core, 30GB memory server for about 6c/hr. Since I only use it a few hours per day its cost effective.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120200,
      "author_name": "krish240574",
      "author_url": "",
      "post_date": "05/16/2016 09:57:59",
      "content": "<p>I use a super slow machine and graphlab - data reading works brilliantly well. \nMy machine is an old, rickety pentium with 4GB RAM, all is well. \nexp_train = graphlab.SFrame(&quot;d:\\datasets\\kaggle-expedia\\train.csv&quot;)\n(SFrame being disk-backed, it works quite well). Do explore, gets all the data loading and model training parts out of the way easily. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120205,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "05/16/2016 11:29:24",
      "content": "<p>I'm using Python with blaze + dask + bcolz for working on the full data set.</p>\n\n<p>Things can go a bit slow from time to time, but avoids running out of memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118927": "I see the train data size is nearly 4gb, it even takes a lot of time to read train.csv using pandas, how do model the train data? do you use any sampling methods to extract part of the data or use any distributed platforms like hadoop, spark?",
    "118965": "1. Data Reduction\r\n2. Online Learning\r\n3. Cloud computing on Amazon\r\n4. Hope your birthday is coming up and buy a bigger computer ;)",
    "118968": "Take a look at [this script][1] which is basically suggestion #1.  In this case your training data can be aggregated.  For your feature exploration, take a sampling as described in the tutorial [here][2]\r\n\r\nBoth links are courtesy of forum dwellers here.\r\n\r\n\r\n  [1]: https://www.kaggle.com/dvasyukova/expedia-hotel-recommendations/predict-hotel-type-with-pandas\r\n  [2]: https://www.dataquest.io/blog/kaggle-tutorial/",
    "119099": "I went ahead with pandas load_csv on  my MacAir with 4GB memory. Took about 7 mins to load the csv file. Then I ran a downsample of the data set . Works like a charm !!",
    "119115": "There is another thread in the forum covering that topic. From what I see in the scripts and in my own experience it seems like the way to go is online learning (or parsing the files line by line). Not sure if anyone used Machine Learning successfully yet.\r\n\r\n3 is only an option if money is not an issue, I think...\r\n\r\nGerhard",
    "119298": "I'm using GraphLab (which claims to be more efficient) and AWS. Using spot pricing I can get an 8 core, 30GB memory server for about 6c/hr. Since I only use it a few hours per day its cost effective.",
    "120200": "I use a super slow machine and graphlab - data reading works brilliantly well. \r\nMy machine is an old, rickety pentium with 4GB RAM, all is well. \r\nexp_train = graphlab.SFrame(\"d:\\\\datasets\\\\kaggle-expedia\\\\train.csv\")\r\n(SFrame being disk-backed, it works quite well). Do explore, gets all the data loading and model training parts out of the way easily.",
    "120205": "I'm using Python with blaze + dask + bcolz for working on the full data set.\r\n\r\nThings can go a bit slow from time to time, but avoids running out of memory."
  },
  "source": "meta"
}