{
  "id": 21151,
  "title": "Newbie - Question on loading data & others",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21151",
  "author_name": "",
  "post_date": "2016-05-23T10:12:45.410Z",
  "votes": null,
  "comment_count": 1,
  "views": 395,
  "content": "<p>Hi \nPlease absolve my technical in-depth as this is my first time here in Kaggle and first competition. Very new here and I have these queries</p>\n\n<p>I have a 32 bit laptop with 4GB RAM. So I am unable to load the 37 million rows of training data set. So,</p>\n\n<ol>\n<li>I had loaded 3 million records, then filtered for a random 40,000 user IDs to get 1.2 million records and then a further filtering on dates to get to a 881 K records on training. </li>\n</ol>\n\n<p>Is this sufficient for training as I am unable to run with more records than this. </p>\n\n<p>Thanks\nSri</p>",
  "messages": [
    {
      "id": "121055",
      "postDate": "05/23/2016 10:12:45",
      "content": "<p>Hi \nPlease absolve my technical in-depth as this is my first time here in Kaggle and first competition. Very new here and I have these queries</p>\n\n<p>I have a 32 bit laptop with 4GB RAM. So I am unable to load the 37 million rows of training data set. So,</p>\n\n<ol>\n<li>I had loaded 3 million records, then filtered for a random 40,000 user IDs to get 1.2 million records and then a further filtering on dates to get to a 881 K records on training. </li>\n</ol>\n\n<p>Is this sufficient for training as I am unable to run with more records than this. </p>\n\n<p>Thanks\nSri</p>",
      "rawMarkdown": "Hi \r\nPlease absolve my technical in-depth as this is my first time here in Kaggle and first competition. Very new here and I have these queries\r\n\r\nI have a 32 bit laptop with 4GB RAM. So I am unable to load the 37 million rows of training data set. So,\r\n\r\n1. I had loaded 3 million records, then filtered for a random 40,000 user IDs to get 1.2 million records and then a further filtering on dates to get to a 881 K records on training. \r\n\r\nIs this sufficient for training as I am unable to run with more records than this. \r\n\r\nThanks\r\nSri",
      "votes": null
    },
    {
      "id": "121402",
      "postDate": "05/26/2016 04:52:23",
      "content": "<p>Hey Sri,</p>\n\n<p>Welcome to Kaggle.  It's pretty fun here, and there are loads of opportunities to learn cool stuff, so keep at it even if you feel in over your head a little at the outset.</p>\n\n<p>I suspect it will be tough to get a model expressive enough with so few records...  more data is always better, and you're trying to model 100 hotel clusters in 50,000 destinations, which gives you 5M combinations you want to consider, and that's before you add any context about the user or the booking itself.  Only considering 900k records seems a bit light for the task.</p>\n\n<p>BUT...</p>\n\n<p>Who's to say what's &quot;enough&quot;.  I will bet big bucks a top 10 finisher is going to have run their approach with &lt; 4GB or RAM and make it look simple.  Maybe one approach is to not try to load all the records at once, but read them one at a time...  and increment summary tables as you read them.  Not sure that's a great idea, but it's an idea that lets you leverage more records.  Also spend time looking around the forums in previous competitions, and look for solutions that used small amounts of RAM to find approaches you can adopt for this one.  That's a great way to learn, and to uncover ideas that have been proven in the past.</p>\n\n<p>Good luck!\nkevin</p>",
      "rawMarkdown": "Hey Sri,\r\n\r\nWelcome to Kaggle.  It's pretty fun here, and there are loads of opportunities to learn cool stuff, so keep at it even if you feel in over your head a little at the outset.\r\n\r\nI suspect it will be tough to get a model expressive enough with so few records...  more data is always better, and you're trying to model 100 hotel clusters in 50,000 destinations, which gives you 5M combinations you want to consider, and that's before you add any context about the user or the booking itself.  Only considering 900k records seems a bit light for the task.\r\n\r\nBUT...\r\n\r\nWho's to say what's \"enough\".  I will bet big bucks a top 10 finisher is going to have run their approach with < 4GB or RAM and make it look simple.  Maybe one approach is to not try to load all the records at once, but read them one at a time...  and increment summary tables as you read them.  Not sure that's a great idea, but it's an idea that lets you leverage more records.  Also spend time looking around the forums in previous competitions, and look for solutions that used small amounts of RAM to find approaches you can adopt for this one.  That's a great way to learn, and to uncover ideas that have been proven in the past.\r\n\r\nGood luck!\r\nkevin",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 121402,
      "author_name": "kevinhinson",
      "author_url": "",
      "post_date": "05/26/2016 04:52:23",
      "content": "<p>Hey Sri,</p>\n\n<p>Welcome to Kaggle.  It's pretty fun here, and there are loads of opportunities to learn cool stuff, so keep at it even if you feel in over your head a little at the outset.</p>\n\n<p>I suspect it will be tough to get a model expressive enough with so few records...  more data is always better, and you're trying to model 100 hotel clusters in 50,000 destinations, which gives you 5M combinations you want to consider, and that's before you add any context about the user or the booking itself.  Only considering 900k records seems a bit light for the task.</p>\n\n<p>BUT...</p>\n\n<p>Who's to say what's &quot;enough&quot;.  I will bet big bucks a top 10 finisher is going to have run their approach with &lt; 4GB or RAM and make it look simple.  Maybe one approach is to not try to load all the records at once, but read them one at a time...  and increment summary tables as you read them.  Not sure that's a great idea, but it's an idea that lets you leverage more records.  Also spend time looking around the forums in previous competitions, and look for solutions that used small amounts of RAM to find approaches you can adopt for this one.  That's a great way to learn, and to uncover ideas that have been proven in the past.</p>\n\n<p>Good luck!\nkevin</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121055": "Hi \r\nPlease absolve my technical in-depth as this is my first time here in Kaggle and first competition. Very new here and I have these queries\r\n\r\nI have a 32 bit laptop with 4GB RAM. So I am unable to load the 37 million rows of training data set. So,\r\n\r\n1. I had loaded 3 million records, then filtered for a random 40,000 user IDs to get 1.2 million records and then a further filtering on dates to get to a 881 K records on training. \r\n\r\nIs this sufficient for training as I am unable to run with more records than this. \r\n\r\nThanks\r\nSri",
    "121402": "Hey Sri,\r\n\r\nWelcome to Kaggle.  It's pretty fun here, and there are loads of opportunities to learn cool stuff, so keep at it even if you feel in over your head a little at the outset.\r\n\r\nI suspect it will be tough to get a model expressive enough with so few records...  more data is always better, and you're trying to model 100 hotel clusters in 50,000 destinations, which gives you 5M combinations you want to consider, and that's before you add any context about the user or the booking itself.  Only considering 900k records seems a bit light for the task.\r\n\r\nBUT...\r\n\r\nWho's to say what's \"enough\".  I will bet big bucks a top 10 finisher is going to have run their approach with < 4GB or RAM and make it look simple.  Maybe one approach is to not try to load all the records at once, but read them one at a time...  and increment summary tables as you read them.  Not sure that's a great idea, but it's an idea that lets you leverage more records.  Also spend time looking around the forums in previous competitions, and look for solutions that used small amounts of RAM to find approaches you can adopt for this one.  That's a great way to learn, and to uncover ideas that have been proven in the past.\r\n\r\nGood luck!\r\nkevin"
  },
  "source": "meta"
}