{
  "id": 20306,
  "title": "How to load the extracted file into the R???",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20306",
  "author_name": "",
  "post_date": "2016-04-21T09:20:12.800Z",
  "votes": null,
  "comment_count": 12,
  "views": 1600,
  "content": "<p>I have been trying to extract the given data file and load it into R . but its not responding </p>",
  "messages": [
    {
      "id": "115998",
      "postDate": "04/21/2016 09:20:12",
      "content": "<p>I have been trying to extract the given data file and load it into R . but its not responding </p>",
      "rawMarkdown": "I have been trying to extract the given data file and load it into R . but its not responding",
      "votes": null
    },
    {
      "id": "115999",
      "postDate": "04/21/2016 09:27:39",
      "content": "<p>me too facing the same issue,  Its not evening opening in Excel file. </p>",
      "rawMarkdown": "me too facing the same issue,  Its not evening opening in Excel file.",
      "votes": null
    },
    {
      "id": "116003",
      "postDate": "04/21/2016 10:55:12",
      "content": "<p>The file has over 40 million observations. </p>\n\n<p>It is far too large to be loaded in Excel, and the default read.csv function in R might take some time. </p>\n\n<p>You can try with the read_csv one from the readr package, which is much faster. It might mess up the type inference for some columns though, so you'd need to tell it the correct types.</p>",
      "rawMarkdown": "The file has over 40 million observations. \r\n\r\nIt is far too large to be loaded in Excel, and the default read.csv function in R might take some time. \r\n\r\nYou can try with the read_csv one from the readr package, which is much faster. It might mess up the type inference for some columns though, so you'd need to tell it the correct types.",
      "votes": null
    },
    {
      "id": "116004",
      "postDate": "04/21/2016 11:20:46",
      "content": "<p>Yesterday I tried to load train into R using a 32GB machine. When memory consumption was around 22GB R got stuck as far as I can tell.</p>\n\n<p>I will thus only use samples of train in R.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Yesterday I tried to load train into R using a 32GB machine. When memory consumption was around 22GB R got stuck as far as I can tell.\r\n\r\nI will thus only use samples of train in R.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "116033",
      "postDate": "04/21/2016 16:16:20",
      "content": "<p>Use fread() from data.table</p>",
      "rawMarkdown": "Use fread() from data.table",
      "votes": null
    },
    {
      "id": "116051",
      "postDate": "04/21/2016 17:38:14",
      "content": "<p>As mentioned, use <strong>fread</strong>:</p>\n\n<pre><code>library(data.table)\ntrain &lt;- fread(&quot;train.csv&quot;, header = T, stringsAsFactors = F)\n</code></pre>\n\n<p>After that you may save resulting data.table to disk with:</p>\n\n<pre><code>save(train,file = &quot;train_saved.dat&quot;)\n</code></pre>\n\n<p>This will allow you to read it fast later with:</p>\n\n<pre><code>load(&quot;train_saved.dat&quot;)\n</code></pre>",
      "rawMarkdown": "As mentioned, use **fread**:\r\n\r\n    library(data.table)\r\n    train <- fread(\"train.csv\", header = T, stringsAsFactors = F)\r\n\r\nAfter that you may save resulting data.table to disk with:\r\n\r\n    save(train,file = \"train_saved.dat\")\r\n\r\nThis will allow you to read it fast later with:\r\n\r\n    load(\"train_saved.dat\")",
      "votes": null
    },
    {
      "id": "116410",
      "postDate": "04/24/2016 01:10:10",
      "content": "<p>Even if I loaded the data into r, but I still cannot run any model. Do not how to deal with this.</p>",
      "rawMarkdown": "Even if I loaded the data into r, but I still cannot run any model. Do not how to deal with this.",
      "votes": null
    },
    {
      "id": "116443",
      "postDate": "04/24/2016 09:34:53",
      "content": "<p>The problem is with the size of the data. It is too big to fit into RAM, so R is not able to work on it in any normal laptop / desktop machine. </p>\n\n<p>In cases such as these, a good way is to sample the data right away. Instead of pulling the full data, pull a reasonably sized random sample of observations. Nobody builds a model on 38MM observations - so this should be just fine. Also, you could sample only bookings (and not clicks) since in the end the test data is booking only (and if I am understanding right, the leaderboard data is also booking only).</p>",
      "rawMarkdown": "The problem is with the size of the data. It is too big to fit into RAM, so R is not able to work on it in any normal laptop / desktop machine. \r\n\r\nIn cases such as these, a good way is to sample the data right away. Instead of pulling the full data, pull a reasonably sized random sample of observations. Nobody builds a model on 38MM observations - so this should be just fine. Also, you could sample only bookings (and not clicks) since in the end the test data is booking only (and if I am understanding right, the leaderboard data is also booking only).",
      "votes": null
    },
    {
      "id": "118986",
      "postDate": "05/06/2016 14:59:47",
      "content": "<p>Another newbie here: I downloaded the files, all of which have the .csv.gz extension.  How do I unzip this into csv using R?  On another note, I tried downloading using 7-zip, but the output of the submission file after extraction doesn't look right (I get the ID# plus &quot;99 1&quot; for all row IDs).</p>",
      "rawMarkdown": "Another newbie here: I downloaded the files, all of which have the .csv.gz extension.  How do I unzip this into csv using R?  On another note, I tried downloading using 7-zip, but the output of the submission file after extraction doesn't look right (I get the ID# plus \"99 1\" for all row IDs).",
      "votes": null
    },
    {
      "id": "119076",
      "postDate": "05/07/2016 01:45:03",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "119300",
      "postDate": "05/09/2016 00:26:35",
      "content": "<p>I succesfully have used readr or fread for big files with 16GB RAM. If you use read.csv, use option stringsAsFactors=F, but either you specify with previously with colClasses that some dates are present, or you convert to dates later. After reading, I use package SOAR for storing files in the HDD. You can use the files later, when needed. After storing, or when needed, I use gc() to free space. Maybe it's good to split columns, for example, date-time related in a separate file to do some transformations, for example, year, weekday, etc. Other option is to split rows, but I think it is not needed.</p>",
      "rawMarkdown": "I succesfully have used readr or fread for big files with 16GB RAM. If you use read.csv, use option stringsAsFactors=F, but either you specify with previously with colClasses that some dates are present, or you convert to dates later. After reading, I use package SOAR for storing files in the HDD. You can use the files later, when needed. After storing, or when needed, I use gc() to free space. Maybe it's good to split columns, for example, date-time related in a separate file to do some transformations, for example, year, weekday, etc. Other option is to split rows, but I think it is not needed.",
      "votes": null
    },
    {
      "id": "119308",
      "postDate": "05/09/2016 04:31:32",
      "content": "<p>Install cygwin, then run:</p>\n\n<p>awk -F&quot;,&quot; '$8%100==37 {print}' train.csv &gt; train.37.csv</p>\n\n<p>to create a 1/100th sample of data where mod(user_id, 100)=37.\nRepeat M times to get a M/100th sample of data, or of course, use\na divisor N less than 100 to get a 1/Nth sample of the data.</p>",
      "rawMarkdown": "Install cygwin, then run:\r\n\r\nawk -F\",\" '$8%100==37 {print}' train.csv > train.37.csv\r\n\r\nto create a 1/100th sample of data where mod(user_id, 100)=37.\r\nRepeat M times to get a M/100th sample of data, or of course, use\r\na divisor N less than 100 to get a 1/Nth sample of the data.",
      "votes": null
    },
    {
      "id": "119362",
      "postDate": "05/09/2016 14:42:18",
      "content": "<p>Look into the ff package (and ffbase) - this has been very helpful to me.</p>",
      "rawMarkdown": "Look into the ff package (and ffbase) - this has been very helpful to me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115999,
      "author_name": "vishalrshukla",
      "author_url": "",
      "post_date": "04/21/2016 09:27:39",
      "content": "<p>me too facing the same issue,  Its not evening opening in Excel file. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116003,
      "author_name": "gpistre",
      "author_url": "",
      "post_date": "04/21/2016 10:55:12",
      "content": "<p>The file has over 40 million observations. </p>\n\n<p>It is far too large to be loaded in Excel, and the default read.csv function in R might take some time. </p>\n\n<p>You can try with the read_csv one from the readr package, which is much faster. It might mess up the type inference for some columns though, so you'd need to tell it the correct types.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116004,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/21/2016 11:20:46",
      "content": "<p>Yesterday I tried to load train into R using a 32GB machine. When memory consumption was around 22GB R got stuck as far as I can tell.</p>\n\n<p>I will thus only use samples of train in R.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116033,
      "author_name": "mpjdem",
      "author_url": "",
      "post_date": "04/21/2016 16:16:20",
      "content": "<p>Use fread() from data.table</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116051,
      "author_name": "dimetrix",
      "author_url": "",
      "post_date": "04/21/2016 17:38:14",
      "content": "<p>As mentioned, use <strong>fread</strong>:</p>\n\n<pre><code>library(data.table)\ntrain &lt;- fread(&quot;train.csv&quot;, header = T, stringsAsFactors = F)\n</code></pre>\n\n<p>After that you may save resulting data.table to disk with:</p>\n\n<pre><code>save(train,file = &quot;train_saved.dat&quot;)\n</code></pre>\n\n<p>This will allow you to read it fast later with:</p>\n\n<pre><code>load(&quot;train_saved.dat&quot;)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116410,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "04/24/2016 01:10:10",
      "content": "<p>Even if I loaded the data into r, but I still cannot run any model. Do not how to deal with this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116443,
      "author_name": "malamute",
      "author_url": "",
      "post_date": "04/24/2016 09:34:53",
      "content": "<p>The problem is with the size of the data. It is too big to fit into RAM, so R is not able to work on it in any normal laptop / desktop machine. </p>\n\n<p>In cases such as these, a good way is to sample the data right away. Instead of pulling the full data, pull a reasonably sized random sample of observations. Nobody builds a model on 38MM observations - so this should be just fine. Also, you could sample only bookings (and not clicks) since in the end the test data is booking only (and if I am understanding right, the leaderboard data is also booking only).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118986,
      "author_name": "margiehertneck",
      "author_url": "",
      "post_date": "05/06/2016 14:59:47",
      "content": "<p>Another newbie here: I downloaded the files, all of which have the .csv.gz extension.  How do I unzip this into csv using R?  On another note, I tried downloading using 7-zip, but the output of the submission file after extraction doesn't look right (I get the ID# plus &quot;99 1&quot; for all row IDs).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119076,
      "author_name": "",
      "author_url": "",
      "post_date": "05/07/2016 01:45:03",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 119300,
      "author_name": "gmobaz",
      "author_url": "",
      "post_date": "05/09/2016 00:26:35",
      "content": "<p>I succesfully have used readr or fread for big files with 16GB RAM. If you use read.csv, use option stringsAsFactors=F, but either you specify with previously with colClasses that some dates are present, or you convert to dates later. After reading, I use package SOAR for storing files in the HDD. You can use the files later, when needed. After storing, or when needed, I use gc() to free space. Maybe it's good to split columns, for example, date-time related in a separate file to do some transformations, for example, year, weekday, etc. Other option is to split rows, but I think it is not needed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119308,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "05/09/2016 04:31:32",
      "content": "<p>Install cygwin, then run:</p>\n\n<p>awk -F&quot;,&quot; '$8%100==37 {print}' train.csv &gt; train.37.csv</p>\n\n<p>to create a 1/100th sample of data where mod(user_id, 100)=37.\nRepeat M times to get a M/100th sample of data, or of course, use\na divisor N less than 100 to get a 1/Nth sample of the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119362,
      "author_name": "margiehertneck",
      "author_url": "",
      "post_date": "05/09/2016 14:42:18",
      "content": "<p>Look into the ff package (and ffbase) - this has been very helpful to me.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115998": "I have been trying to extract the given data file and load it into R . but its not responding",
    "115999": "me too facing the same issue,  Its not evening opening in Excel file.",
    "116003": "The file has over 40 million observations. \r\n\r\nIt is far too large to be loaded in Excel, and the default read.csv function in R might take some time. \r\n\r\nYou can try with the read_csv one from the readr package, which is much faster. It might mess up the type inference for some columns though, so you'd need to tell it the correct types.",
    "116004": "Yesterday I tried to load train into R using a 32GB machine. When memory consumption was around 22GB R got stuck as far as I can tell.\r\n\r\nI will thus only use samples of train in R.\r\n\r\nGerhard",
    "116033": "Use fread() from data.table",
    "116051": "As mentioned, use **fread**:\r\n\r\n    library(data.table)\r\n    train <- fread(\"train.csv\", header = T, stringsAsFactors = F)\r\n\r\nAfter that you may save resulting data.table to disk with:\r\n\r\n    save(train,file = \"train_saved.dat\")\r\n\r\nThis will allow you to read it fast later with:\r\n\r\n    load(\"train_saved.dat\")",
    "116410": "Even if I loaded the data into r, but I still cannot run any model. Do not how to deal with this.",
    "116443": "The problem is with the size of the data. It is too big to fit into RAM, so R is not able to work on it in any normal laptop / desktop machine. \r\n\r\nIn cases such as these, a good way is to sample the data right away. Instead of pulling the full data, pull a reasonably sized random sample of observations. Nobody builds a model on 38MM observations - so this should be just fine. Also, you could sample only bookings (and not clicks) since in the end the test data is booking only (and if I am understanding right, the leaderboard data is also booking only).",
    "118986": "Another newbie here: I downloaded the files, all of which have the .csv.gz extension.  How do I unzip this into csv using R?  On another note, I tried downloading using 7-zip, but the output of the submission file after extraction doesn't look right (I get the ID# plus \"99 1\" for all row IDs).",
    "119076": "",
    "119300": "I succesfully have used readr or fread for big files with 16GB RAM. If you use read.csv, use option stringsAsFactors=F, but either you specify with previously with colClasses that some dates are present, or you convert to dates later. After reading, I use package SOAR for storing files in the HDD. You can use the files later, when needed. After storing, or when needed, I use gc() to free space. Maybe it's good to split columns, for example, date-time related in a separate file to do some transformations, for example, year, weekday, etc. Other option is to split rows, but I think it is not needed.",
    "119308": "Install cygwin, then run:\r\n\r\nawk -F\",\" '$8%100==37 {print}' train.csv > train.37.csv\r\n\r\nto create a 1/100th sample of data where mod(user_id, 100)=37.\r\nRepeat M times to get a M/100th sample of data, or of course, use\r\na divisor N less than 100 to get a 1/Nth sample of the data.",
    "119362": "Look into the ff package (and ffbase) - this has been very helpful to me."
  },
  "source": "meta"
}