{
  "id": 21153,
  "title": "How to work with this dataset on 4Gb Ram",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21153",
  "author_name": "",
  "post_date": "2016-05-23T11:54:32.967Z",
  "votes": null,
  "comment_count": 1,
  "views": 674,
  "content": "",
  "messages": [
    {
      "id": "121064",
      "postDate": "05/23/2016 11:54:32",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "122761",
      "postDate": "06/07/2016 01:17:59",
      "content": "<p>Sorry this is a bit belated, but I had similar problems on a 8GB laptop. I used a SQL database, and you can too. SQLite might be easiest to get started with, or PostgreSQL if you're looking to get a little fancier.</p>\n\n<p>Define a table with something like the following:</p>\n\n<p>create table train (\n    date_time DATETIME,\n    site_name VARCHAR(8),\n    posa_continent VARCHAR(8),\n    user_location_country VARCHAR(8),\n    user_location_region VARCHAR(8),\n    user_location_city VARCHAR(8),\n    orig_destination_distance REAL,\n    user_id VARCHAR(8),\n    is_mobile VARCHAR(8),\n    is_package VARCHAR(8),\n    channel VARCHAR(8),\n    srch_ci DATE,\n    srch_co DATE,\n    srch_adults_cnt REAL,\n    srch_children_cnt REAL,\n    srch_rm_cnt REAL,\n    srch_destination_id VARCHAR(8),\n    srch_destination_type_id VARCHAR(8),\n    is_booking VARCHAR(8),\n    cnt REAL,\n    hotel_continent VARCHAR(8),\n    hotel_country VARCHAR(8),\n    hotel_market VARCHAR(8),\n    hotel_cluster VARCHAR(8)\n);</p>\n\n<p>Then load in the data. In SQLite, you'd strip off the first line by hand, and then use .mode csv and the .import command to load the CSV file into the database.</p>\n\n<p>From there you can compute conditional probabilities for clusters, given other features, and save those to other tables. When you have a CP table and want to use it to predict something, you can use a SQL join against the 'test' table.</p>\n\n<p>SQLite will do a lot of the work on-disk, which means it doesn't require a lot of memory. But beware it's going to be very, very slow.</p>\n\n<p>The R dplyr:: functions can work against SQLite and PostgresSQL, and are a handy tool if you don't want to hand-write SQL. </p>\n\n<p>Hope this helps.</p>",
      "rawMarkdown": "Sorry this is a bit belated, but I had similar problems on a 8GB laptop. I used a SQL database, and you can too. SQLite might be easiest to get started with, or PostgreSQL if you're looking to get a little fancier.\r\n\r\nDefine a table with something like the following:\r\n\r\ncreate table train (\r\n    date_time DATETIME,\r\n    site_name VARCHAR(8),\r\n    posa_continent VARCHAR(8),\r\n    user_location_country VARCHAR(8),\r\n    user_location_region VARCHAR(8),\r\n    user_location_city VARCHAR(8),\r\n    orig_destination_distance REAL,\r\n    user_id VARCHAR(8),\r\n    is_mobile VARCHAR(8),\r\n    is_package VARCHAR(8),\r\n    channel VARCHAR(8),\r\n    srch_ci DATE,\r\n    srch_co DATE,\r\n    srch_adults_cnt REAL,\r\n    srch_children_cnt REAL,\r\n    srch_rm_cnt REAL,\r\n    srch_destination_id VARCHAR(8),\r\n    srch_destination_type_id VARCHAR(8),\r\n    is_booking VARCHAR(8),\r\n    cnt REAL,\r\n    hotel_continent VARCHAR(8),\r\n    hotel_country VARCHAR(8),\r\n    hotel_market VARCHAR(8),\r\n    hotel_cluster VARCHAR(8)\r\n);\r\n\r\nThen load in the data. In SQLite, you'd strip off the first line by hand, and then use .mode csv and the .import command to load the CSV file into the database.\r\n\r\nFrom there you can compute conditional probabilities for clusters, given other features, and save those to other tables. When you have a CP table and want to use it to predict something, you can use a SQL join against the 'test' table.\r\n\r\nSQLite will do a lot of the work on-disk, which means it doesn't require a lot of memory. But beware it's going to be very, very slow.\r\n\r\nThe R dplyr:: functions can work against SQLite and PostgresSQL, and are a handy tool if you don't want to hand-write SQL. \r\n\r\nHope this helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122761,
      "author_name": "rareitmeyer",
      "author_url": "",
      "post_date": "06/07/2016 01:17:59",
      "content": "<p>Sorry this is a bit belated, but I had similar problems on a 8GB laptop. I used a SQL database, and you can too. SQLite might be easiest to get started with, or PostgreSQL if you're looking to get a little fancier.</p>\n\n<p>Define a table with something like the following:</p>\n\n<p>create table train (\n    date_time DATETIME,\n    site_name VARCHAR(8),\n    posa_continent VARCHAR(8),\n    user_location_country VARCHAR(8),\n    user_location_region VARCHAR(8),\n    user_location_city VARCHAR(8),\n    orig_destination_distance REAL,\n    user_id VARCHAR(8),\n    is_mobile VARCHAR(8),\n    is_package VARCHAR(8),\n    channel VARCHAR(8),\n    srch_ci DATE,\n    srch_co DATE,\n    srch_adults_cnt REAL,\n    srch_children_cnt REAL,\n    srch_rm_cnt REAL,\n    srch_destination_id VARCHAR(8),\n    srch_destination_type_id VARCHAR(8),\n    is_booking VARCHAR(8),\n    cnt REAL,\n    hotel_continent VARCHAR(8),\n    hotel_country VARCHAR(8),\n    hotel_market VARCHAR(8),\n    hotel_cluster VARCHAR(8)\n);</p>\n\n<p>Then load in the data. In SQLite, you'd strip off the first line by hand, and then use .mode csv and the .import command to load the CSV file into the database.</p>\n\n<p>From there you can compute conditional probabilities for clusters, given other features, and save those to other tables. When you have a CP table and want to use it to predict something, you can use a SQL join against the 'test' table.</p>\n\n<p>SQLite will do a lot of the work on-disk, which means it doesn't require a lot of memory. But beware it's going to be very, very slow.</p>\n\n<p>The R dplyr:: functions can work against SQLite and PostgresSQL, and are a handy tool if you don't want to hand-write SQL. </p>\n\n<p>Hope this helps.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121064": "",
    "122761": "Sorry this is a bit belated, but I had similar problems on a 8GB laptop. I used a SQL database, and you can too. SQLite might be easiest to get started with, or PostgreSQL if you're looking to get a little fancier.\r\n\r\nDefine a table with something like the following:\r\n\r\ncreate table train (\r\n    date_time DATETIME,\r\n    site_name VARCHAR(8),\r\n    posa_continent VARCHAR(8),\r\n    user_location_country VARCHAR(8),\r\n    user_location_region VARCHAR(8),\r\n    user_location_city VARCHAR(8),\r\n    orig_destination_distance REAL,\r\n    user_id VARCHAR(8),\r\n    is_mobile VARCHAR(8),\r\n    is_package VARCHAR(8),\r\n    channel VARCHAR(8),\r\n    srch_ci DATE,\r\n    srch_co DATE,\r\n    srch_adults_cnt REAL,\r\n    srch_children_cnt REAL,\r\n    srch_rm_cnt REAL,\r\n    srch_destination_id VARCHAR(8),\r\n    srch_destination_type_id VARCHAR(8),\r\n    is_booking VARCHAR(8),\r\n    cnt REAL,\r\n    hotel_continent VARCHAR(8),\r\n    hotel_country VARCHAR(8),\r\n    hotel_market VARCHAR(8),\r\n    hotel_cluster VARCHAR(8)\r\n);\r\n\r\nThen load in the data. In SQLite, you'd strip off the first line by hand, and then use .mode csv and the .import command to load the CSV file into the database.\r\n\r\nFrom there you can compute conditional probabilities for clusters, given other features, and save those to other tables. When you have a CP table and want to use it to predict something, you can use a SQL join against the 'test' table.\r\n\r\nSQLite will do a lot of the work on-disk, which means it doesn't require a lot of memory. But beware it's going to be very, very slow.\r\n\r\nThe R dplyr:: functions can work against SQLite and PostgresSQL, and are a handy tool if you don't want to hand-write SQL. \r\n\r\nHope this helps."
  },
  "source": "meta"
}