{
  "id": 20779,
  "title": "Please confirm correct gz extraction?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20779",
  "author_name": "",
  "post_date": "2016-05-06T23:18:31.537Z",
  "votes": null,
  "comment_count": 3,
  "views": 388,
  "content": "<p>Apologies for the newbie question, but I'm not sure I'm extracting the csv.gz files correctly.  I've used both winzip, 7-zip and the gzfile and fread functions within R to extract the sample_submission.csv.gz file as a test...and it looks like this:</p>\n\n<p>id,hotel_cluster\n0, 99 1\n1, 99 1\n3, 99 1\netc...</p>\n\n<p>According to the contest info, it should look something like this:</p>\n\n<p>id,hotel_cluster\n0,99 3 1 75 20\n1,2 50 30 23 9\netc...</p>\n\n<p>I don't want to spend hours incorrectly downloading the test/train/destination data; could someone please help me with some R code or direction on how to do it outside of programming?  (I only know R at this point).</p>\n\n<p>Thanks much!</p>",
  "messages": [
    {
      "id": "119057",
      "postDate": "05/06/2016 23:18:31",
      "content": "<p>Apologies for the newbie question, but I'm not sure I'm extracting the csv.gz files correctly.  I've used both winzip, 7-zip and the gzfile and fread functions within R to extract the sample_submission.csv.gz file as a test...and it looks like this:</p>\n\n<p>id,hotel_cluster\n0, 99 1\n1, 99 1\n3, 99 1\netc...</p>\n\n<p>According to the contest info, it should look something like this:</p>\n\n<p>id,hotel_cluster\n0,99 3 1 75 20\n1,2 50 30 23 9\netc...</p>\n\n<p>I don't want to spend hours incorrectly downloading the test/train/destination data; could someone please help me with some R code or direction on how to do it outside of programming?  (I only know R at this point).</p>\n\n<p>Thanks much!</p>",
      "rawMarkdown": "Apologies for the newbie question, but I'm not sure I'm extracting the csv.gz files correctly.  I've used both winzip, 7-zip and the gzfile and fread functions within R to extract the sample_submission.csv.gz file as a test...and it looks like this:\r\n\r\nid,hotel_cluster\r\n0, 99 1\r\n1, 99 1\r\n3, 99 1\r\netc...\r\n\r\nAccording to the contest info, it should look something like this:\r\n\r\nid,hotel_cluster\r\n0,99 3 1 75 20\r\n1,2 50 30 23 9\r\netc...\r\n\r\nI don't want to spend hours incorrectly downloading the test/train/destination data; could someone please help me with some R code or direction on how to do it outside of programming?  (I only know R at this point).\r\n\r\nThanks much!",
      "votes": null
    },
    {
      "id": "119060",
      "postDate": "05/06/2016 23:39:38",
      "content": "<p>hi, my file is the same as yours. It should be correct.</p>",
      "rawMarkdown": "hi, my file is the same as yours. It should be correct.",
      "votes": null
    },
    {
      "id": "119063",
      "postDate": "05/06/2016 23:54:26",
      "content": "<p>Don't worry, the sample submission file is correct that way. If you look at the lower benchmark on the leaderboard, it says:</p>\n\n<blockquote>\n  <p>Sample Submission Benchmark   0.01744          Benchmark Info always\n  predicting 99 1</p>\n</blockquote>\n\n<p>What's shown on the Evaluation page is just an example of what it could look like with five predictions.\nBtw, R seems to be able to handle .csv.gz files just like unzipped .csv's. I at least tried it with:</p>\n\n<blockquote>\n  <p>submission &lt;- read.csv(&quot;sample_submission.csv.gz&quot;)</p>\n  \n  <p>submission &lt;- read.table(&quot;sample_submission.csv.gz&quot;)</p>\n  \n  <p>library(readr)</p>\n  \n  <p>submission &lt;- read_csv(&quot;sample_submission.csv.gz&quot;)</p>\n</blockquote>",
      "rawMarkdown": "Don't worry, the sample submission file is correct that way. If you look at the lower benchmark on the leaderboard, it says:\r\n\r\n> Sample Submission Benchmark \t0.01744 \t\t Benchmark Info always\r\n> predicting 99 1\r\n\r\nWhat's shown on the Evaluation page is just an example of what it could look like with five predictions.\r\nBtw, R seems to be able to handle .csv.gz files just like unzipped .csv's. I at least tried it with:\r\n\r\n> submission <- read.csv(\"sample_submission.csv.gz\")\r\n> \r\n> submission <- read.table(\"sample_submission.csv.gz\")\r\n> \r\n> library(readr)\r\n> \r\n> submission <- read_csv(\"sample_submission.csv.gz\")",
      "votes": null
    },
    {
      "id": "119066",
      "postDate": "05/07/2016 00:11:46",
      "content": "<p>Excellent, thanks for checking this!</p>",
      "rawMarkdown": "Excellent, thanks for checking this!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119060,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "05/06/2016 23:39:38",
      "content": "<p>hi, my file is the same as yours. It should be correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119063,
      "author_name": "hnowak",
      "author_url": "",
      "post_date": "05/06/2016 23:54:26",
      "content": "<p>Don't worry, the sample submission file is correct that way. If you look at the lower benchmark on the leaderboard, it says:</p>\n\n<blockquote>\n  <p>Sample Submission Benchmark   0.01744          Benchmark Info always\n  predicting 99 1</p>\n</blockquote>\n\n<p>What's shown on the Evaluation page is just an example of what it could look like with five predictions.\nBtw, R seems to be able to handle .csv.gz files just like unzipped .csv's. I at least tried it with:</p>\n\n<blockquote>\n  <p>submission &lt;- read.csv(&quot;sample_submission.csv.gz&quot;)</p>\n  \n  <p>submission &lt;- read.table(&quot;sample_submission.csv.gz&quot;)</p>\n  \n  <p>library(readr)</p>\n  \n  <p>submission &lt;- read_csv(&quot;sample_submission.csv.gz&quot;)</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119066,
      "author_name": "margiehertneck",
      "author_url": "",
      "post_date": "05/07/2016 00:11:46",
      "content": "<p>Excellent, thanks for checking this!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119057": "Apologies for the newbie question, but I'm not sure I'm extracting the csv.gz files correctly.  I've used both winzip, 7-zip and the gzfile and fread functions within R to extract the sample_submission.csv.gz file as a test...and it looks like this:\r\n\r\nid,hotel_cluster\r\n0, 99 1\r\n1, 99 1\r\n3, 99 1\r\netc...\r\n\r\nAccording to the contest info, it should look something like this:\r\n\r\nid,hotel_cluster\r\n0,99 3 1 75 20\r\n1,2 50 30 23 9\r\netc...\r\n\r\nI don't want to spend hours incorrectly downloading the test/train/destination data; could someone please help me with some R code or direction on how to do it outside of programming?  (I only know R at this point).\r\n\r\nThanks much!",
    "119060": "hi, my file is the same as yours. It should be correct.",
    "119063": "Don't worry, the sample submission file is correct that way. If you look at the lower benchmark on the leaderboard, it says:\r\n\r\n> Sample Submission Benchmark \t0.01744 \t\t Benchmark Info always\r\n> predicting 99 1\r\n\r\nWhat's shown on the Evaluation page is just an example of what it could look like with five predictions.\r\nBtw, R seems to be able to handle .csv.gz files just like unzipped .csv's. I at least tried it with:\r\n\r\n> submission <- read.csv(\"sample_submission.csv.gz\")\r\n> \r\n> submission <- read.table(\"sample_submission.csv.gz\")\r\n> \r\n> library(readr)\r\n> \r\n> submission <- read_csv(\"sample_submission.csv.gz\")",
    "119066": "Excellent, thanks for checking this!"
  },
  "source": "meta"
}