{
  "id": 20684,
  "title": "A tutorial to get you started",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20684",
  "author_name": "",
  "post_date": "2016-05-03T20:24:43.177Z",
  "votes": 112,
  "comment_count": 40,
  "views": 8757,
  "content": "<p>I wrote a tutorial on how to get into the top 15 on the leaderboard in the competition (as of now):  <a href=\"https://www.dataquest.io/blog/kaggle-tutorial/\">https://www.dataquest.io/blog/kaggle-tutorial/</a> . Hope you enjoy it!</p>",
  "messages": [
    {
      "id": "118471",
      "postDate": "05/03/2016 20:24:43",
      "content": "<p>I wrote a tutorial on how to get into the top 15 on the leaderboard in the competition (as of now):  <a href=\"https://www.dataquest.io/blog/kaggle-tutorial/\">https://www.dataquest.io/blog/kaggle-tutorial/</a> . Hope you enjoy it!</p>",
      "rawMarkdown": "I wrote a tutorial on how to get into the top 15 on the leaderboard in the competition (as of now):  https://www.dataquest.io/blog/kaggle-tutorial/ . Hope you enjoy it!",
      "votes": null
    },
    {
      "id": "118474",
      "postDate": "05/03/2016 20:57:43",
      "content": "<p>Thanks!  I wish you had used a better language [R :) ] but I will enjoy working through your examples and reasoning. Thanks for teaching us!</p>",
      "rawMarkdown": "Thanks!  I wish you had used a better language [R :) ] but I will enjoy working through your examples and reasoning. Thanks for teaching us!",
      "votes": null
    },
    {
      "id": "118503",
      "postDate": "05/04/2016 01:42:59",
      "content": "<p>Very detailed and well explained.. I really enjoyed reading through it and definitely learned more than one thing or two from the material. Thanks for the tutorial!</p>",
      "rawMarkdown": "Very detailed and well explained.. I really enjoyed reading through it and definitely learned more than one thing or two from the material. Thanks for the tutorial!",
      "votes": null
    },
    {
      "id": "118544",
      "postDate": "05/04/2016 07:45:09",
      "content": "<p>Very nice tutorial. Thank you for sharing your knowledge.</p>",
      "rawMarkdown": "Very nice tutorial. Thank you for sharing your knowledge.",
      "votes": null
    },
    {
      "id": "118557",
      "postDate": "05/04/2016 09:03:25",
      "content": "<p>I wish there was a post like this in every competition :)</p>",
      "rawMarkdown": "I wish there was a post like this in every competition :)",
      "votes": null
    },
    {
      "id": "118564",
      "postDate": "05/04/2016 09:35:47",
      "content": "<p>@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? </p>",
      "rawMarkdown": "Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb?",
      "votes": null
    },
    {
      "id": "118573",
      "postDate": "05/04/2016 10:16:42",
      "content": "<p>@Vik Thank you very much</p>\n\n<p>With a very similar (and in times more granular) approach in perl I made 0.47x. Will surely look into your solution.</p>\n\n<p>Can anyone tell how much memory is needed for this script?</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Vik Thank you very much\r\n\r\nWith a very similar (and in times more granular) approach in perl I made 0.47x. Will surely look into your solution.\r\n\r\nCan anyone tell how much memory is needed for this script?\r\n\r\nThanks\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "118611",
      "postDate": "05/04/2016 13:51:39",
      "content": "<p>One of the easier to follow and understand tutorials that I have read, awesome job!</p>",
      "rawMarkdown": "One of the easier to follow and understand tutorials that I have read, awesome job!",
      "votes": null
    },
    {
      "id": "118623",
      "postDate": "05/04/2016 14:40:31",
      "content": "<p>[quote=Konrad Banachewicz;118564]</p>\n\n<p>@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? </p>\n\n<p>[/quote]</p>\n\n<p>Thanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.</p>",
      "rawMarkdown": "[quote=Konrad Banachewicz;118564]\r\n\r\n@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? \r\n\r\n[/quote]\r\n\r\nThanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.",
      "votes": null
    },
    {
      "id": "118626",
      "postDate": "05/04/2016 14:46:16",
      "content": "<p>[quote=Vik Paruchuri;118623]</p>\n\n<p>[quote=Konrad Banachewicz;118564]</p>\n\n<p>@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? </p>\n\n<p>[/quote]</p>\n\n<p>Thanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you Vik! Very helpful!</p>\n\n<p>I'm doing a much more careless validation but on 670k data. My local score is about 0.01 higher than lb.</p>",
      "rawMarkdown": "[quote=Vik Paruchuri;118623]\r\n\r\n[quote=Konrad Banachewicz;118564]\r\n\r\n@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? \r\n\r\n[/quote]\r\n\r\nThanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.\r\n\r\n[/quote]\r\n\r\nThank you Vik! Very helpful!\r\n\r\nI'm doing a much more careless validation but on 670k data. My local score is about 0.01 higher than lb.",
      "votes": null
    },
    {
      "id": "118628",
      "postDate": "05/04/2016 14:47:08",
      "content": "<p>@Vik: that's interesting - and potentially another part of the challenge. I suspect (didnt have time to verify it yet) it might have sth to do with the subsampling, i.e. the outcome is quite sensitive to <em>which</em> 10000 are sampled.</p>",
      "rawMarkdown": "Vik: that's interesting - and potentially another part of the challenge. I suspect (didnt have time to verify it yet) it might have sth to do with the subsampling, i.e. the outcome is quite sensitive to *which* 10000 are sampled.",
      "votes": null
    },
    {
      "id": "118636",
      "postDate": "05/04/2016 15:04:32",
      "content": "<p>@Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). </p>\n\n<p>The only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).</p>",
      "rawMarkdown": "Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). \r\n\r\nThe only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).",
      "votes": null
    },
    {
      "id": "118637",
      "postDate": "05/04/2016 15:06:58",
      "content": "<p>@rcarson , but using the same approach or you actually used some ML?</p>\n\n<p>How would you suggest splitting the set into train/test? \nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?</p>",
      "rawMarkdown": "rcarson , but using the same approach or you actually used some ML?\r\n\r\nHow would you suggest splitting the set into train/test? \r\nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?",
      "votes": null
    },
    {
      "id": "118638",
      "postDate": "05/04/2016 15:14:17",
      "content": "<p>[quote=Marcin P&#281;kalski;118637]</p>\n\n<p>@rcarson , but using the same approach or you actually used some ML?</p>\n\n<p>How would you suggest splitting the set into train/test? \nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?</p>\n\n<p>[/quote]</p>\n\n<p>We're simply using the last 670k data (very arbitrary choice) in train for validation. But Vik's way (by dates) makes more sense, you just need more data in validation to reduce local-lb difference.</p>\n\n<p>Our ml is still a little worse than our rule-based model but we are getting there.  </p>",
      "rawMarkdown": "[quote=Marcin Pękalski;118637]\r\n\r\n@rcarson , but using the same approach or you actually used some ML?\r\n\r\nHow would you suggest splitting the set into train/test? \r\nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?\r\n\r\n[/quote]\r\n\r\nWe're simply using the last 670k data (very arbitrary choice) in train for validation. But Vik's way (by dates) makes more sense, you just need more data in validation to reduce local-lb difference.\r\n\r\nOur ml is still a little worse than our rule-based model but we are getting there.",
      "votes": null
    },
    {
      "id": "118644",
      "postDate": "05/04/2016 15:32:49",
      "content": "<p>[quote=dune_dweller;118636]</p>\n\n<p>...If logistic regression doesn't work here the reason lies somewhere else :).</p>\n\n<p>[/quote]</p>\n\n<p>Agreed. I guess the reason is that there are too many unique values for each feature and their interactions. To accommodate all the features, the logistic regression will have to store and update a huge number of parameters - which means a memory problem. If we try to hash down the number of features, random hash collisions would be unavoidable, and hence the model would be compromised.</p>",
      "rawMarkdown": "[quote=dune_dweller;118636]\r\n\r\n...If logistic regression doesn't work here the reason lies somewhere else :).\r\n\r\n[/quote]\r\n\r\nAgreed. I guess the reason is that there are too many unique values for each feature and their interactions. To accommodate all the features, the logistic regression will have to store and update a huge number of parameters - which means a memory problem. If we try to hash down the number of features, random hash collisions would be unavoidable, and hence the model would be compromised.",
      "votes": null
    },
    {
      "id": "118646",
      "postDate": "05/04/2016 15:41:07",
      "content": "<p>Hi,</p>\n\n<p>so. After work I ran the first lines in python and RAM used max was nearly 30GB. Dropping to ~21GB after load finished.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nso. After work I ran the first lines in python and RAM used max was nearly 30GB. Dropping to ~21GB after load finished.\r\n\r\nCheers\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "118649",
      "postDate": "05/04/2016 15:46:35",
      "content": "<p>@MightyBird, try changing data types from 64bit to 32bit. That big precision is not needed, and the RAM requirements should go down a bit. Just remember to use garbage collector afterwards, to force the cleaning.</p>",
      "rawMarkdown": "MightyBird, try changing data types from 64bit to 32bit. That big precision is not needed, and the RAM requirements should go down a bit. Just remember to use garbage collector afterwards, to force the cleaning.",
      "votes": null
    },
    {
      "id": "118661",
      "postDate": "05/04/2016 16:44:52",
      "content": "<p>@Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard? </p>",
      "rawMarkdown": "Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard?",
      "votes": null
    },
    {
      "id": "118672",
      "postDate": "05/04/2016 17:20:02",
      "content": "<p>Thank you Vik! Definitely very useful for a beginner Kaggler like me :)</p>",
      "rawMarkdown": "Thank you Vik! Definitely very useful for a beginner Kaggler like me :)",
      "votes": null
    },
    {
      "id": "118693",
      "postDate": "05/04/2016 19:31:30",
      "content": "<p>A new benchmark :)</p>",
      "rawMarkdown": "A new benchmark :)",
      "votes": null
    },
    {
      "id": "118696",
      "postDate": "05/04/2016 19:45:47",
      "content": "<p>Thanks for the great tutorial. \nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard. </p>",
      "rawMarkdown": "Thanks for the great tutorial. \r\nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard.",
      "votes": null
    },
    {
      "id": "118697",
      "postDate": "05/04/2016 19:50:21",
      "content": "<p>[quote=Mohamed Ali Jamaoui;118696]</p>\n\n<p>Thanks for the great tutorial. \nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard. </p>\n\n<p>[/quote]</p>\n\n<p>The data section states that the test set only consists of booking events.</p>",
      "rawMarkdown": "[quote=Mohamed Ali Jamaoui;118696]\r\n\r\nThanks for the great tutorial. \r\nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard. \r\n\r\n[/quote]\r\n\r\nThe data section states that the test set only consists of booking events.",
      "votes": null
    },
    {
      "id": "118712",
      "postDate": "05/04/2016 21:38:38",
      "content": "<p>Nice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.</p>",
      "rawMarkdown": "Nice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.",
      "votes": null
    },
    {
      "id": "118729",
      "postDate": "05/05/2016 00:57:36",
      "content": "<p>Sorry for the double post, but I did want to send an alert to everyone following the thread.  </p>\n\n<p>For me the killer code was this</p>\n\n<pre><code>for i in range(test.shape[0]):\n     exact_matches.append(generate_exact_matches(test.iloc[i], match_cols))\n</code></pre>\n\n<p>after about an hour+ I was only at 100k entries.</p>\n\n<p>I think a panda merge (pd.merge)--basically an SQL left join for those of us in database world, might be more efficient here for those of us with lesser hardware.  I'll try and figure it out in a bit--its been a day and a half unrelated to kaggle, but I do want to contribute and make the script a bit more accessible.</p>",
      "rawMarkdown": "Sorry for the double post, but I did want to send an alert to everyone following the thread.  \r\n\r\nFor me the killer code was this\r\n\r\n    for i in range(test.shape[0]):\r\n         exact_matches.append(generate_exact_matches(test.iloc[i], match_cols))\r\n\r\nafter about an hour+ I was only at 100k entries.\r\n\r\nI think a panda merge (pd.merge)--basically an SQL left join for those of us in database world, might be more efficient here for those of us with lesser hardware.  I'll try and figure it out in a bit--its been a day and a half unrelated to kaggle, but I do want to contribute and make the script a bit more accessible.",
      "votes": null
    },
    {
      "id": "118742",
      "postDate": "05/05/2016 04:03:49",
      "content": "<p>Great script Vik and very well written article!</p>\n\n<p>One critique - you should probably join the destinations dataset to the training data (or a sample of the training data) <em>before</em> you run PCA instead of running PCA on the destinations dataset in isolation.  The distribution of destinations within the samples could have an effect on the variance within the features.</p>",
      "rawMarkdown": "Great script Vik and very well written article!\r\n\r\nOne critique - you should probably join the destinations dataset to the training data (or a sample of the training data) *before* you run PCA instead of running PCA on the destinations dataset in isolation.  The distribution of destinations within the samples could have an effect on the variance within the features.",
      "votes": null
    },
    {
      "id": "118753",
      "postDate": "05/05/2016 06:17:39",
      "content": "<p>If you really take a look at the data in there, many of the rows have duplicate responses for one of the values. They just look like fills with one value from what I saw in the first few rows. Even doing a transform with Rtsne noted &quot;duplicates&quot;. To be succinct - I highly doubt that each destination was reviewed 150 times. In that highly unlikely situation, I also doubt that many got a same sentiment value (derived from solely from text analysis, if not a weighted average of rating and sentiment analysis) several times.</p>\n\n<p>The unique number of values for each row will differ if you look closely. I fear that doing a linear (or otherwise) transform, like PCA, on these values may not be the best choice of action without prior cleaning.</p>\n\n<p>Since I'm unsure of whether the mean is being used as the duplicate value for each row, it may be good to use that as the fill, instead of the value that's being displayed in each row.</p>\n\n<p>Excellent script though.</p>",
      "rawMarkdown": "If you really take a look at the data in there, many of the rows have duplicate responses for one of the values. They just look like fills with one value from what I saw in the first few rows. Even doing a transform with Rtsne noted \"duplicates\". To be succinct - I highly doubt that each destination was reviewed 150 times. In that highly unlikely situation, I also doubt that many got a same sentiment value (derived from solely from text analysis, if not a weighted average of rating and sentiment analysis) several times.\r\n\r\nThe unique number of values for each row will differ if you look closely. I fear that doing a linear (or otherwise) transform, like PCA, on these values may not be the best choice of action without prior cleaning.\r\n\r\nSince I'm unsure of whether the mean is being used as the duplicate value for each row, it may be good to use that as the fill, instead of the value that's being displayed in each row.\r\n\r\nExcellent script though.",
      "votes": null
    },
    {
      "id": "118844",
      "postDate": "05/05/2016 16:55:24",
      "content": "<p>[quote=Bob, Lord of Cats;118729]\nNice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook. 12 hours and it still was cranking away on the final step.\n[/quote]</p>\n\n<p>[quote=Bob, Lord of Cats;118729]\nafter about an hour+ I was only at 100k entries.\n[/quote]</p>\n\n<p>He actually mentions in the post:</p>\n\n<p>[quote]\nGiven the amount of memory on your system, it may or may not be feasible to read all the data in. If it isn&#8217;t, you should consider creating a machine on EC2 or DigitalOcean to process the data with.[/quote]</p>\n\n<p>and then links to this post <a href=\"https://www.dataquest.io/blog/data-science-ec2-quickstart/\">https://www.dataquest.io/blog/data-science-ec2-quickstart/</a></p>\n\n<p>Obviously there's a cost involved here, but if it helps you iterate faster and doesn't lock up your machine, $30 doesn't seem like that much to me.</p>",
      "rawMarkdown": "[quote=Bob, Lord of Cats;118729]\r\nNice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook. 12 hours and it still was cranking away on the final step.\r\n[/quote]\r\n\r\n[quote=Bob, Lord of Cats;118729]\r\nafter about an hour+ I was only at 100k entries.\r\n[/quote]\r\n\r\nHe actually mentions in the post:\r\n\r\n[quote]\r\nGiven the amount of memory on your system, it may or may not be feasible to read all the data in. If it isn’t, you should consider creating a machine on EC2 or DigitalOcean to process the data with.[/quote]\r\n\r\nand then links to this post https://www.dataquest.io/blog/data-science-ec2-quickstart/\r\n\r\nObviously there's a cost involved here, but if it helps you iterate faster and doesn't lock up your machine, $30 doesn't seem like that much to me.",
      "votes": null
    },
    {
      "id": "118845",
      "postDate": "05/05/2016 17:02:49",
      "content": "<p>@Bob, Lord of Cats \nTry to run this <a href=\"https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial\">version</a>. The <em>killer code</em> you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)</p>",
      "rawMarkdown": "Bob, Lord of Cats \r\nTry to run this [version][1]. The *killer code* you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)\r\n\r\n\r\n  [1]: https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial",
      "votes": null
    },
    {
      "id": "118847",
      "postDate": "05/05/2016 17:07:20",
      "content": "<p>[quote=Bob, Lord of Cats;118712]</p>\n\n<p>Nice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.</p>\n\n<p>[/quote]</p>\n\n<p>It took about an hour to run on my machine (16gb RAM, though).  Maybe a library or system package version difference?  Also, the ML code will be impractical to run on the whole training / test set, so that should be left out.</p>",
      "rawMarkdown": "[quote=Bob, Lord of Cats;118712]\r\n\r\nNice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.\r\n\r\n[/quote]\r\n\r\nIt took about an hour to run on my machine (16gb RAM, though).  Maybe a library or system package version difference?  Also, the ML code will be impractical to run on the whole training / test set, so that should be left out.",
      "votes": null
    },
    {
      "id": "118848",
      "postDate": "05/05/2016 17:09:52",
      "content": "<p>[quote=dune_dweller;118636]</p>\n\n<p>@Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). </p>\n\n<p>The only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, I should have explained this better -- it comes down more to not being able to linearly separate the clusters with the given features because the boundaries are so fuzzy.</p>",
      "rawMarkdown": "[quote=dune_dweller;118636]\r\n\r\n@Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). \r\n\r\nThe only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).\r\n\r\n[/quote]\r\n\r\nYeah, I should have explained this better -- it comes down more to not being able to linearly separate the clusters with the given features because the boundaries are so fuzzy.",
      "votes": null
    },
    {
      "id": "118849",
      "postDate": "05/05/2016 17:10:39",
      "content": "<p>[quote=MelvinDunn;118661]</p>\n\n<p>@Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard? </p>\n\n<p>[/quote]</p>\n\n<p>Please do! :)</p>",
      "rawMarkdown": "[quote=MelvinDunn;118661]\r\n\r\n@Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard? \r\n\r\n[/quote]\r\n\r\nPlease do! :)",
      "votes": null
    },
    {
      "id": "118851",
      "postDate": "05/05/2016 17:29:37",
      "content": "<p>@Vik: Excellent tutorial! I also confirm that 100 x binary classification does not work here, I created 100 Online Logistic models using feature hashing (with the entire training dataset) and the resulting precision was the same (about 0.05) , in addition the AUC/LogLoss metrics for each classifier were about 0.68/0.29.\nWell done!</p>",
      "rawMarkdown": "Vik: Excellent tutorial! I also confirm that 100 x binary classification does not work here, I created 100 Online Logistic models using feature hashing (with the entire training dataset) and the resulting precision was the same (about 0.05) , in addition the AUC/LogLoss metrics for each classifier were about 0.68/0.29.\r\nWell done!",
      "votes": null
    },
    {
      "id": "118855",
      "postDate": "05/05/2016 18:23:18",
      "content": "<p>@Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?</p>",
      "rawMarkdown": "Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?",
      "votes": null
    },
    {
      "id": "118856",
      "postDate": "05/05/2016 18:24:00",
      "content": "<p>[quote=Manel;118845]</p>\n\n<p>@Bob, Lord of Cats \nTry to run this <a href=\"https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial\">version</a>. The <em>killer code</em> you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)</p>\n\n<p>[/quote]\nI'll give it a shot after work.  My idea wasn't to complain but improve because I think there is a more efficient way to do it (I actually have been modifying Dune Dweller's code as I couldn't quite get this one to the point I wanted where I could get my join/merge to work).  My apologies as in my studies in R/Python/ML it has been pounded into me LOOPS BAD/VECTORIZE YOUR CODE.  Plus in the grand scheme of research, grabbing several users ideas, combining them into my own and then have all of you tell me what a genius I am. (I am sort of slow on the combine part, but still... ;) )</p>\n\n<p>update:  I dropped out of ipython into a straight shell--the improvement for me is significant--running at about 100K samples / 45 sec.  Much quicker.  Methinks there is an ID:10T error in my notebook config.  R simply blows away my ipython notebook for speed. </p>\n\n<p>23:09 for the &quot;Leak&quot; section.  I knew the old gal had it in her!</p>",
      "rawMarkdown": "[quote=Manel;118845]\r\n\r\n@Bob, Lord of Cats \r\nTry to run this [version][1]. The *killer code* you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)\r\n\r\n\r\n  [1]: https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial\r\n\r\n[/quote]\r\nI'll give it a shot after work.  My idea wasn't to complain but improve because I think there is a more efficient way to do it (I actually have been modifying Dune Dweller's code as I couldn't quite get this one to the point I wanted where I could get my join/merge to work).  My apologies as in my studies in R/Python/ML it has been pounded into me LOOPS BAD/VECTORIZE YOUR CODE.  Plus in the grand scheme of research, grabbing several users ideas, combining them into my own and then have all of you tell me what a genius I am. (I am sort of slow on the combine part, but still... ;) )\r\n\r\nupdate:  I dropped out of ipython into a straight shell--the improvement for me is significant--running at about 100K samples / 45 sec.  Much quicker.  Methinks there is an ID:10T error in my notebook config.  R simply blows away my ipython notebook for speed. \r\n\r\n23:09 for the \"Leak\" section.  I knew the old gal had it in her!",
      "votes": null
    },
    {
      "id": "118858",
      "postDate": "05/05/2016 18:44:12",
      "content": "<p>[quote=Marcin P&#281;kalski;118855]</p>\n\n<p>@Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?</p>\n\n<p>[/quote]\nHello Marcin, I used basically the raw data excluding the dates, I created a few features (year, month, day of week, total days of stay)  using the dates. Regarding passes, just passed the training dataset 2 times.\nNo fine tuning here, but I guess not so much chance for this to work.</p>",
      "rawMarkdown": "[quote=Marcin Pękalski;118855]\r\n\r\n@Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?\r\n\r\n[/quote]\r\nHello Marcin, I used basically the raw data excluding the dates, I created a few features (year, month, day of week, total days of stay)  using the dates. Regarding passes, just passed the training dataset 2 times.\r\nNo fine tuning here, but I guess not so much chance for this to work.",
      "votes": null
    },
    {
      "id": "118903",
      "postDate": "05/06/2016 01:28:07",
      "content": "<p>I owe @Vik an apology--the program wasn't the problem--apparently my Jupyter Notebook/IPython shell and chrome do not play well at all (10X or more performance hit).  Running from a prompt gives excellent performance (under an hour as advertised).</p>",
      "rawMarkdown": "I owe @Vik an apology--the program wasn't the problem--apparently my Jupyter Notebook/IPython shell and chrome do not play well at all (10X or more performance hit).  Running from a prompt gives excellent performance (under an hour as advertised).",
      "votes": null
    },
    {
      "id": "120832",
      "postDate": "05/20/2016 20:14:04",
      "content": "<p>Glad you got it sorted in the end Bob!</p>",
      "rawMarkdown": "Glad you got it sorted in the end Bob!",
      "votes": null
    },
    {
      "id": "160460",
      "postDate": "02/07/2017 19:37:13",
      "content": "<p>@Vik many thanks! Following a lot of your tips/tutorials is helping a lot. </p>",
      "rawMarkdown": "Vik many thanks! Following a lot of your tips/tutorials is helping a lot.",
      "votes": null
    },
    {
      "id": "160461",
      "postDate": "02/07/2017 19:37:16",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "344911",
      "postDate": "06/18/2018 22:40:18",
      "content": "<p>Thanks for the link @Vik </p>",
      "rawMarkdown": "Thanks for the link @Vik",
      "votes": null
    },
    {
      "id": "3080956",
      "postDate": "12/26/2024 02:14:38",
      "content": "<p>thanks for the tutorial</p>",
      "rawMarkdown": "thanks for the tutorial",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3080956,
      "author_name": "anowersha",
      "author_url": "",
      "post_date": "12/26/2024 02:14:38",
      "content": "<p>thanks for the tutorial</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118474,
      "author_name": "piinformatics",
      "author_url": "",
      "post_date": "05/03/2016 20:57:43",
      "content": "<p>Thanks!  I wish you had used a better language [R :) ] but I will enjoy working through your examples and reasoning. Thanks for teaching us!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118503,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "05/04/2016 01:42:59",
      "content": "<p>Very detailed and well explained.. I really enjoyed reading through it and definitely learned more than one thing or two from the material. Thanks for the tutorial!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118544,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/04/2016 07:45:09",
      "content": "<p>Very nice tutorial. Thank you for sharing your knowledge.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118557,
      "author_name": "vykhand",
      "author_url": "",
      "post_date": "05/04/2016 09:03:25",
      "content": "<p>I wish there was a post like this in every competition :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118564,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/04/2016 09:35:47",
      "content": "<p>@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118573,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/04/2016 10:16:42",
      "content": "<p>@Vik Thank you very much</p>\n\n<p>With a very similar (and in times more granular) approach in perl I made 0.47x. Will surely look into your solution.</p>\n\n<p>Can anyone tell how much memory is needed for this script?</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118611,
      "author_name": "nitini",
      "author_url": "",
      "post_date": "05/04/2016 13:51:39",
      "content": "<p>One of the easier to follow and understand tutorials that I have read, awesome job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118623,
      "author_name": "vikpar",
      "author_url": "",
      "post_date": "05/04/2016 14:40:31",
      "content": "<p>[quote=Konrad Banachewicz;118564]</p>\n\n<p>@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? </p>\n\n<p>[/quote]</p>\n\n<p>Thanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118626,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "05/04/2016 14:46:16",
      "content": "<p>[quote=Vik Paruchuri;118623]</p>\n\n<p>[quote=Konrad Banachewicz;118564]</p>\n\n<p>@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? </p>\n\n<p>[/quote]</p>\n\n<p>Thanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you Vik! Very helpful!</p>\n\n<p>I'm doing a much more careless validation but on 670k data. My local score is about 0.01 higher than lb.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118628,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/04/2016 14:47:08",
      "content": "<p>@Vik: that's interesting - and potentially another part of the challenge. I suspect (didnt have time to verify it yet) it might have sth to do with the subsampling, i.e. the outcome is quite sensitive to <em>which</em> 10000 are sampled.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118636,
      "author_name": "dvasyukova",
      "author_url": "",
      "post_date": "05/04/2016 15:04:32",
      "content": "<p>@Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). </p>\n\n<p>The only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118637,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "05/04/2016 15:06:58",
      "content": "<p>@rcarson , but using the same approach or you actually used some ML?</p>\n\n<p>How would you suggest splitting the set into train/test? \nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118638,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "05/04/2016 15:14:17",
      "content": "<p>[quote=Marcin P&#281;kalski;118637]</p>\n\n<p>@rcarson , but using the same approach or you actually used some ML?</p>\n\n<p>How would you suggest splitting the set into train/test? \nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?</p>\n\n<p>[/quote]</p>\n\n<p>We're simply using the last 670k data (very arbitrary choice) in train for validation. But Vik's way (by dates) makes more sense, you just need more data in validation to reduce local-lb difference.</p>\n\n<p>Our ml is still a little worse than our rule-based model but we are getting there.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118644,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "05/04/2016 15:32:49",
      "content": "<p>[quote=dune_dweller;118636]</p>\n\n<p>...If logistic regression doesn't work here the reason lies somewhere else :).</p>\n\n<p>[/quote]</p>\n\n<p>Agreed. I guess the reason is that there are too many unique values for each feature and their interactions. To accommodate all the features, the logistic regression will have to store and update a huge number of parameters - which means a memory problem. If we try to hash down the number of features, random hash collisions would be unavoidable, and hence the model would be compromised.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118646,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/04/2016 15:41:07",
      "content": "<p>Hi,</p>\n\n<p>so. After work I ran the first lines in python and RAM used max was nearly 30GB. Dropping to ~21GB after load finished.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118649,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "05/04/2016 15:46:35",
      "content": "<p>@MightyBird, try changing data types from 64bit to 32bit. That big precision is not needed, and the RAM requirements should go down a bit. Just remember to use garbage collector afterwards, to force the cleaning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118661,
      "author_name": "",
      "author_url": "",
      "post_date": "05/04/2016 16:44:52",
      "content": "<p>@Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118672,
      "author_name": "dannychan",
      "author_url": "",
      "post_date": "05/04/2016 17:20:02",
      "content": "<p>Thank you Vik! Definitely very useful for a beginner Kaggler like me :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118693,
      "author_name": "yangyang",
      "author_url": "",
      "post_date": "05/04/2016 19:31:30",
      "content": "<p>A new benchmark :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118696,
      "author_name": "mohamedali",
      "author_url": "",
      "post_date": "05/04/2016 19:45:47",
      "content": "<p>Thanks for the great tutorial. \nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118697,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "05/04/2016 19:50:21",
      "content": "<p>[quote=Mohamed Ali Jamaoui;118696]</p>\n\n<p>Thanks for the great tutorial. \nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard. </p>\n\n<p>[/quote]</p>\n\n<p>The data section states that the test set only consists of booking events.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118712,
      "author_name": "rdslater",
      "author_url": "",
      "post_date": "05/04/2016 21:38:38",
      "content": "<p>Nice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118729,
      "author_name": "rdslater",
      "author_url": "",
      "post_date": "05/05/2016 00:57:36",
      "content": "<p>Sorry for the double post, but I did want to send an alert to everyone following the thread.  </p>\n\n<p>For me the killer code was this</p>\n\n<pre><code>for i in range(test.shape[0]):\n     exact_matches.append(generate_exact_matches(test.iloc[i], match_cols))\n</code></pre>\n\n<p>after about an hour+ I was only at 100k entries.</p>\n\n<p>I think a panda merge (pd.merge)--basically an SQL left join for those of us in database world, might be more efficient here for those of us with lesser hardware.  I'll try and figure it out in a bit--its been a day and a half unrelated to kaggle, but I do want to contribute and make the script a bit more accessible.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118742,
      "author_name": "ben519",
      "author_url": "",
      "post_date": "05/05/2016 04:03:49",
      "content": "<p>Great script Vik and very well written article!</p>\n\n<p>One critique - you should probably join the destinations dataset to the training data (or a sample of the training data) <em>before</em> you run PCA instead of running PCA on the destinations dataset in isolation.  The distribution of destinations within the samples could have an effect on the variance within the features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118753,
      "author_name": "",
      "author_url": "",
      "post_date": "05/05/2016 06:17:39",
      "content": "<p>If you really take a look at the data in there, many of the rows have duplicate responses for one of the values. They just look like fills with one value from what I saw in the first few rows. Even doing a transform with Rtsne noted &quot;duplicates&quot;. To be succinct - I highly doubt that each destination was reviewed 150 times. In that highly unlikely situation, I also doubt that many got a same sentiment value (derived from solely from text analysis, if not a weighted average of rating and sentiment analysis) several times.</p>\n\n<p>The unique number of values for each row will differ if you look closely. I fear that doing a linear (or otherwise) transform, like PCA, on these values may not be the best choice of action without prior cleaning.</p>\n\n<p>Since I'm unsure of whether the mean is being used as the duplicate value for each row, it may be good to use that as the fill, instead of the value that's being displayed in each row.</p>\n\n<p>Excellent script though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118844,
      "author_name": "jaypeedevlin",
      "author_url": "",
      "post_date": "05/05/2016 16:55:24",
      "content": "<p>[quote=Bob, Lord of Cats;118729]\nNice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook. 12 hours and it still was cranking away on the final step.\n[/quote]</p>\n\n<p>[quote=Bob, Lord of Cats;118729]\nafter about an hour+ I was only at 100k entries.\n[/quote]</p>\n\n<p>He actually mentions in the post:</p>\n\n<p>[quote]\nGiven the amount of memory on your system, it may or may not be feasible to read all the data in. If it isn&#8217;t, you should consider creating a machine on EC2 or DigitalOcean to process the data with.[/quote]</p>\n\n<p>and then links to this post <a href=\"https://www.dataquest.io/blog/data-science-ec2-quickstart/\">https://www.dataquest.io/blog/data-science-ec2-quickstart/</a></p>\n\n<p>Obviously there's a cost involved here, but if it helps you iterate faster and doesn't lock up your machine, $30 doesn't seem like that much to me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118845,
      "author_name": "manels",
      "author_url": "",
      "post_date": "05/05/2016 17:02:49",
      "content": "<p>@Bob, Lord of Cats \nTry to run this <a href=\"https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial\">version</a>. The <em>killer code</em> you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118847,
      "author_name": "vikpar",
      "author_url": "",
      "post_date": "05/05/2016 17:07:20",
      "content": "<p>[quote=Bob, Lord of Cats;118712]</p>\n\n<p>Nice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.</p>\n\n<p>[/quote]</p>\n\n<p>It took about an hour to run on my machine (16gb RAM, though).  Maybe a library or system package version difference?  Also, the ML code will be impractical to run on the whole training / test set, so that should be left out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118848,
      "author_name": "vikpar",
      "author_url": "",
      "post_date": "05/05/2016 17:09:52",
      "content": "<p>[quote=dune_dweller;118636]</p>\n\n<p>@Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). </p>\n\n<p>The only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).</p>\n\n<p>[/quote]</p>\n\n<p>Yeah, I should have explained this better -- it comes down more to not being able to linearly separate the clusters with the given features because the boundaries are so fuzzy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118849,
      "author_name": "vikpar",
      "author_url": "",
      "post_date": "05/05/2016 17:10:39",
      "content": "<p>[quote=MelvinDunn;118661]</p>\n\n<p>@Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard? </p>\n\n<p>[/quote]</p>\n\n<p>Please do! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118851,
      "author_name": "dimitrislev",
      "author_url": "",
      "post_date": "05/05/2016 17:29:37",
      "content": "<p>@Vik: Excellent tutorial! I also confirm that 100 x binary classification does not work here, I created 100 Online Logistic models using feature hashing (with the entire training dataset) and the resulting precision was the same (about 0.05) , in addition the AUC/LogLoss metrics for each classifier were about 0.68/0.29.\nWell done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118855,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "05/05/2016 18:23:18",
      "content": "<p>@Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118856,
      "author_name": "rdslater",
      "author_url": "",
      "post_date": "05/05/2016 18:24:00",
      "content": "<p>[quote=Manel;118845]</p>\n\n<p>@Bob, Lord of Cats \nTry to run this <a href=\"https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial\">version</a>. The <em>killer code</em> you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)</p>\n\n<p>[/quote]\nI'll give it a shot after work.  My idea wasn't to complain but improve because I think there is a more efficient way to do it (I actually have been modifying Dune Dweller's code as I couldn't quite get this one to the point I wanted where I could get my join/merge to work).  My apologies as in my studies in R/Python/ML it has been pounded into me LOOPS BAD/VECTORIZE YOUR CODE.  Plus in the grand scheme of research, grabbing several users ideas, combining them into my own and then have all of you tell me what a genius I am. (I am sort of slow on the combine part, but still... ;) )</p>\n\n<p>update:  I dropped out of ipython into a straight shell--the improvement for me is significant--running at about 100K samples / 45 sec.  Much quicker.  Methinks there is an ID:10T error in my notebook config.  R simply blows away my ipython notebook for speed. </p>\n\n<p>23:09 for the &quot;Leak&quot; section.  I knew the old gal had it in her!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118858,
      "author_name": "dimitrislev",
      "author_url": "",
      "post_date": "05/05/2016 18:44:12",
      "content": "<p>[quote=Marcin P&#281;kalski;118855]</p>\n\n<p>@Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?</p>\n\n<p>[/quote]\nHello Marcin, I used basically the raw data excluding the dates, I created a few features (year, month, day of week, total days of stay)  using the dates. Regarding passes, just passed the training dataset 2 times.\nNo fine tuning here, but I guess not so much chance for this to work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118903,
      "author_name": "rdslater",
      "author_url": "",
      "post_date": "05/06/2016 01:28:07",
      "content": "<p>I owe @Vik an apology--the program wasn't the problem--apparently my Jupyter Notebook/IPython shell and chrome do not play well at all (10X or more performance hit).  Running from a prompt gives excellent performance (under an hour as advertised).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120832,
      "author_name": "jaypeedevlin",
      "author_url": "",
      "post_date": "05/20/2016 20:14:04",
      "content": "<p>Glad you got it sorted in the end Bob!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 160460,
      "author_name": "rayjohnsoncomedy",
      "author_url": "",
      "post_date": "02/07/2017 19:37:13",
      "content": "<p>@Vik many thanks! Following a lot of your tips/tutorials is helping a lot. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 160461,
      "author_name": "rayjohnsoncomedy",
      "author_url": "",
      "post_date": "02/07/2017 19:37:16",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 344911,
      "author_name": "funkegoodvibe",
      "author_url": "",
      "post_date": "06/18/2018 22:40:18",
      "content": "<p>Thanks for the link @Vik </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118471": "I wrote a tutorial on how to get into the top 15 on the leaderboard in the competition (as of now):  https://www.dataquest.io/blog/kaggle-tutorial/ . Hope you enjoy it!",
    "118474": "Thanks!  I wish you had used a better language [R :) ] but I will enjoy working through your examples and reasoning. Thanks for teaching us!",
    "118503": "Very detailed and well explained.. I really enjoyed reading through it and definitely learned more than one thing or two from the material. Thanks for the tutorial!",
    "118544": "Very nice tutorial. Thank you for sharing your knowledge.",
    "118557": "I wish there was a post like this in every competition :)",
    "118564": "Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb?",
    "118573": "Vik Thank you very much\r\n\r\nWith a very similar (and in times more granular) approach in perl I made 0.47x. Will surely look into your solution.\r\n\r\nCan anyone tell how much memory is needed for this script?\r\n\r\nThanks\r\n\r\nGerhard",
    "118611": "One of the easier to follow and understand tutorials that I have read, awesome job!",
    "118623": "[quote=Konrad Banachewicz;118564]\r\n\r\n@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? \r\n\r\n[/quote]\r\n\r\nThanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.",
    "118626": "[quote=Vik Paruchuri;118623]\r\n\r\n[quote=Konrad Banachewicz;118564]\r\n\r\n@Vik: fantastic tutorial, well done. One thing that got me curious: the validation score shown in the script is ~ 0.28 and the snapshot from the leaderboard says 0.48. Is there really such a large discrepancy between validation and lb? \r\n\r\n[/quote]\r\n\r\nThanks!  Yup, in this case, there was.  The validation set was much smaller that the testing set due to sampling.  I outlined a few reasons in the post for why the validation accuracy might be lower than leaderboard accuracy.\r\n\r\n[/quote]\r\n\r\nThank you Vik! Very helpful!\r\n\r\nI'm doing a much more careless validation but on 670k data. My local score is about 0.01 higher than lb.",
    "118628": "Vik: that's interesting - and potentially another part of the challenge. I suspect (didnt have time to verify it yet) it might have sth to do with the subsampling, i.e. the outcome is quite sensitive to *which* 10000 are sampled.",
    "118636": "Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). \r\n\r\nThe only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).",
    "118637": "rcarson , but using the same approach or you actually used some ML?\r\n\r\nHow would you suggest splitting the set into train/test? \r\nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?",
    "118638": "[quote=Marcin Pękalski;118637]\r\n\r\n@rcarson , but using the same approach or you actually used some ML?\r\n\r\nHow would you suggest splitting the set into train/test? \r\nSeems like the most logical would be simply by dates, to have similar issues like with real data (new srch_destination_id in test set), but maybe not?\r\n\r\n[/quote]\r\n\r\nWe're simply using the last 670k data (very arbitrary choice) in train for validation. But Vik's way (by dates) makes more sense, you just need more data in validation to reduce local-lb difference.\r\n\r\nOur ml is still a little worse than our rule-based model but we are getting there.",
    "118644": "[quote=dune_dweller;118636]\r\n\r\n...If logistic regression doesn't work here the reason lies somewhere else :).\r\n\r\n[/quote]\r\n\r\nAgreed. I guess the reason is that there are too many unique values for each feature and their interactions. To accommodate all the features, the logistic regression will have to store and update a huge number of parameters - which means a memory problem. If we try to hash down the number of features, random hash collisions would be unavoidable, and hence the model would be compromised.",
    "118646": "Hi,\r\n\r\nso. After work I ran the first lines in python and RAM used max was nearly 30GB. Dropping to ~21GB after load finished.\r\n\r\nCheers\r\n\r\nGerhard",
    "118649": "MightyBird, try changing data types from 64bit to 32bit. That big precision is not needed, and the RAM requirements should go down a bit. Just remember to use garbage collector afterwards, to force the cleaning.",
    "118661": "Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard?",
    "118672": "Thank you Vik! Definitely very useful for a beginner Kaggler like me :)",
    "118693": "A new benchmark :)",
    "118696": "Thanks for the great tutorial. \r\nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard.",
    "118697": "[quote=Mohamed Ali Jamaoui;118696]\r\n\r\nThanks for the great tutorial. \r\nI wonder, though, if removing the click events is a good strategy as we are not sure if the private test set is also only composed of booking events (has no click events). If it's the case, your solution can overfit to the public leaderboard. \r\n\r\n[/quote]\r\n\r\nThe data section states that the test set only consists of booking events.",
    "118712": "Nice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.",
    "118729": "Sorry for the double post, but I did want to send an alert to everyone following the thread.  \r\n\r\nFor me the killer code was this\r\n\r\n    for i in range(test.shape[0]):\r\n         exact_matches.append(generate_exact_matches(test.iloc[i], match_cols))\r\n\r\nafter about an hour+ I was only at 100k entries.\r\n\r\nI think a panda merge (pd.merge)--basically an SQL left join for those of us in database world, might be more efficient here for those of us with lesser hardware.  I'll try and figure it out in a bit--its been a day and a half unrelated to kaggle, but I do want to contribute and make the script a bit more accessible.",
    "118742": "Great script Vik and very well written article!\r\n\r\nOne critique - you should probably join the destinations dataset to the training data (or a sample of the training data) *before* you run PCA instead of running PCA on the destinations dataset in isolation.  The distribution of destinations within the samples could have an effect on the variance within the features.",
    "118753": "If you really take a look at the data in there, many of the rows have duplicate responses for one of the values. They just look like fills with one value from what I saw in the first few rows. Even doing a transform with Rtsne noted \"duplicates\". To be succinct - I highly doubt that each destination was reviewed 150 times. In that highly unlikely situation, I also doubt that many got a same sentiment value (derived from solely from text analysis, if not a weighted average of rating and sentiment analysis) several times.\r\n\r\nThe unique number of values for each row will differ if you look closely. I fear that doing a linear (or otherwise) transform, like PCA, on these values may not be the best choice of action without prior cleaning.\r\n\r\nSince I'm unsure of whether the mean is being used as the duplicate value for each row, it may be good to use that as the fill, instead of the value that's being displayed in each row.\r\n\r\nExcellent script though.",
    "118844": "[quote=Bob, Lord of Cats;118729]\r\nNice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook. 12 hours and it still was cranking away on the final step.\r\n[/quote]\r\n\r\n[quote=Bob, Lord of Cats;118729]\r\nafter about an hour+ I was only at 100k entries.\r\n[/quote]\r\n\r\nHe actually mentions in the post:\r\n\r\n[quote]\r\nGiven the amount of memory on your system, it may or may not be feasible to read all the data in. If it isn’t, you should consider creating a machine on EC2 or DigitalOcean to process the data with.[/quote]\r\n\r\nand then links to this post https://www.dataquest.io/blog/data-science-ec2-quickstart/\r\n\r\nObviously there's a cost involved here, but if it helps you iterate faster and doesn't lock up your machine, $30 doesn't seem like that much to me.",
    "118845": "Bob, Lord of Cats \r\nTry to run this [version][1]. The *killer code* you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)\r\n\r\n\r\n  [1]: https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial",
    "118847": "[quote=Bob, Lord of Cats;118712]\r\n\r\nNice tutorial but it did crush my machine (e8400+8gb ram) running in an ipython notebook.  12 hours and it still was cranking away on the final step.\r\n\r\n[/quote]\r\n\r\nIt took about an hour to run on my machine (16gb RAM, though).  Maybe a library or system package version difference?  Also, the ML code will be impractical to run on the whole training / test set, so that should be left out.",
    "118848": "[quote=dune_dweller;118636]\r\n\r\n@Vik, nice tutorial! I liked your method for subsampling data, and visiting the Expedia site was a nice touch :). \r\n\r\nThe only thing that seemed off was the comment about how logistic regression is unlikely to work because  no features correlate with hotel cluster. Logistic regression is a classification method, so it shouldn't care at all about the numeric values of class labels. If logistic regression doesn't work here the reason lies somewhere else :).\r\n\r\n[/quote]\r\n\r\nYeah, I should have explained this better -- it comes down more to not being able to linearly separate the clusters with the given features because the boundaries are so fuzzy.",
    "118849": "[quote=MelvinDunn;118661]\r\n\r\n@Vik, you mind if I put your tutorial in the scripts (you'll be cited) and add a bit of chaos to the Leaderboard? \r\n\r\n[/quote]\r\n\r\nPlease do! :)",
    "118851": "Vik: Excellent tutorial! I also confirm that 100 x binary classification does not work here, I created 100 Online Logistic models using feature hashing (with the entire training dataset) and the resulting precision was the same (about 0.05) , in addition the AUC/LogLoss metrics for each classifier were about 0.68/0.29.\r\nWell done!",
    "118855": "Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?",
    "118856": "[quote=Manel;118845]\r\n\r\n@Bob, Lord of Cats \r\nTry to run this [version][1]. The *killer code* you said runs in 0:24:40.167909h on my computer (i7 +8gb ram)\r\n\r\n\r\n  [1]: https://www.kaggle.com/manels/expedia-hotel-recommendations/dataquest-tutorial\r\n\r\n[/quote]\r\nI'll give it a shot after work.  My idea wasn't to complain but improve because I think there is a more efficient way to do it (I actually have been modifying Dune Dweller's code as I couldn't quite get this one to the point I wanted where I could get my join/merge to work).  My apologies as in my studies in R/Python/ML it has been pounded into me LOOPS BAD/VECTORIZE YOUR CODE.  Plus in the grand scheme of research, grabbing several users ideas, combining them into my own and then have all of you tell me what a genius I am. (I am sort of slow on the combine part, but still... ;) )\r\n\r\nupdate:  I dropped out of ipython into a straight shell--the improvement for me is significant--running at about 100K samples / 45 sec.  Much quicker.  Methinks there is an ID:10T error in my notebook config.  R simply blows away my ipython notebook for speed. \r\n\r\n23:09 for the \"Leak\" section.  I knew the old gal had it in her!",
    "118858": "[quote=Marcin Pękalski;118855]\r\n\r\n@Dimitris, but what new features have you created? Or you just run it on the raw data? How many passes you made through the dataset?\r\n\r\n[/quote]\r\nHello Marcin, I used basically the raw data excluding the dates, I created a few features (year, month, day of week, total days of stay)  using the dates. Regarding passes, just passed the training dataset 2 times.\r\nNo fine tuning here, but I guess not so much chance for this to work.",
    "118903": "I owe @Vik an apology--the program wasn't the problem--apparently my Jupyter Notebook/IPython shell and chrome do not play well at all (10X or more performance hit).  Running from a prompt gives excellent performance (under an hour as advertised).",
    "120832": "Glad you got it sorted in the end Bob!",
    "160460": "Vik many thanks! Following a lot of your tips/tutorials is helping a lot.",
    "160461": "",
    "344911": "Thanks for the link @Vik",
    "3080956": "thanks for the tutorial"
  },
  "source": "meta"
}