{
  "id": 20345,
  "title": "Data leak",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20345",
  "author_name": "Adam",
  "post_date": "2016-04-22T20:44:16.257000",
  "votes": 19,
  "comment_count": 48,
  "views": 19918,
  "content": "<p>Thank you all for your patience for waiting for the clarification on the data leak.  We wanted to analyze the full extent before responding.</p>\n\n<p>I confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. Thanks to Arnaud who was the first to point this out. Good catch!</p>\n\n<p>Our estimate is this directly affects approx. 1/3 rd of the holdout data. </p>\n\n<p>The contest will continue without any changes.  For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).</p>\n\n<p>I hope you will still enjoy the rest of the competition.</p>\n\n<p>Adam</p>",
  "messages": [
    {
      "id": 116277,
      "postDate": "2016-04-22T20:44:16.257Z",
      "content": "<p>Thank you all for your patience for waiting for the clarification on the data leak.  We wanted to analyze the full extent before responding.</p>\n\n<p>I confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. Thanks to Arnaud who was the first to point this out. Good catch!</p>\n\n<p>Our estimate is this directly affects approx. 1/3 rd of the holdout data. </p>\n\n<p>The contest will continue without any changes.  For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).</p>\n\n<p>I hope you will still enjoy the rest of the competition.</p>\n\n<p>Adam</p>",
      "rawMarkdown": "Thank you all for your patience for waiting for the clarification on the data leak.  We wanted to analyze the full extent before responding.\r\n\r\nI confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. Thanks to Arnaud who was the first to point this out. Good catch!\r\n\r\nOur estimate is this directly affects approx. 1/3 rd of the holdout data. \r\n\r\nThe contest will continue without any changes.  For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).\r\n\r\nI hope you will still enjoy the rest of the competition.\r\n\r\nAdam\r\n",
      "votes": 19
    },
    {
      "id": 117278,
      "postDate": "2016-04-28T10:31:56.693Z",
      "content": "<p>[quote=Wendy Kan;116282]</p>\n\n<p>@Sergey, </p>\n\n<p>They won't be excluded.</p>\n\n<p>[/quote]</p>\n\n<p>So the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\nThe winner will be the best puzzle solver.</p>\n\n<p>Is more leaderboard efficient invest time in solve the puzzle.</p>\n\n<p>Could you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?</p>",
      "rawMarkdown": "[quote=Wendy Kan;116282]\r\n\r\n@Sergey, \r\n\r\nThey won't be excluded.\r\n\r\n[/quote]\r\n\r\nSo the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\r\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\r\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\r\nThe winner will be the best puzzle solver.\r\n\r\nIs more leaderboard efficient invest time in solve the puzzle.\r\n\r\nCould you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?\r\n\r\n\r\n\r\n",
      "votes": 10
    },
    {
      "id": 116279,
      "postDate": "2016-04-22T21:02:44.277Z",
      "content": "<p>Can someone explain this to me?</p>\n\n<p>How do I use this leak to remain competitive with the other submissions?</p>",
      "rawMarkdown": "Can someone explain this to me?\r\n\r\nHow do I use this leak to remain competitive with the other submissions?\r\n\r\n",
      "votes": 9
    },
    {
      "id": 121250,
      "postDate": "2016-05-25T06:26:05.897Z",
      "content": "<p>Again... :(</p>\n\n<p>many people working and there is a data leak in the data set...</p>\n\n<p>kaggle needs to improve its data evaluation process before the competitions... maybe hiring companies from the same area of the competition to evaluate the data set... maybe hiring new data scientists from this comunity...</p>\n\n<p>The question is: we cannot spend our time in wrong data sets. </p>\n\n<p>I got bad times when I spent 2 months in the Winton Challenge and Kaggle changed the data set. And I can tell you: that was a very easy mistake to detect (considering that people know a little about stock markets). My friend and I detected the mistake almost at the same time. Few days after to start... Kaggle identified the mistake almost 2 months later... They did not identify the mistake because they were trying... they identify because the score of the first one told they (a quick improvement)... For this reason I think Kaggle needs to evaluate the possibility to validate data before competitions...</p>\n\n<p><strong>It is very serious. We are talking about our time!!!</strong></p>\n\n<p>This is just my opinion... </p>\n\n<p>and I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...</p>",
      "rawMarkdown": "Again... :(\r\n\r\nmany people working and there is a data leak in the data set...\r\n\r\nkaggle needs to improve its data evaluation process before the competitions... maybe hiring companies from the same area of the competition to evaluate the data set... maybe hiring new data scientists from this comunity...\r\n\r\nThe question is: we cannot spend our time in wrong data sets. \r\n\r\nI got bad times when I spent 2 months in the Winton Challenge and Kaggle changed the data set. And I can tell you: that was a very easy mistake to detect (considering that people know a little about stock markets). My friend and I detected the mistake almost at the same time. Few days after to start... Kaggle identified the mistake almost 2 months later... They did not identify the mistake because they were trying... they identify because the score of the first one told they (a quick improvement)... For this reason I think Kaggle needs to evaluate the possibility to validate data before competitions...\r\n\r\n**It is very serious. We are talking about our time!!!**\r\n\r\nThis is just my opinion... \r\n\r\nand I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...\r\n",
      "votes": 7
    },
    {
      "id": 116597,
      "postDate": "2016-04-25T11:07:02.323Z",
      "content": "<p>[quote=Sangxia;116573]</p>\n\n<p>[quote=Manish Maheshwari;116568]</p>\n\n<p>@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...</p>\n\n<p>[/quote]</p>\n\n<p>Not sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?</p>\n\n<p>[/quote]</p>\n\n<p>Correct, cluster assignments can change. For example, in winter a hotel is cheap and not-popular and will be assigned to cluster A, and in summer the same hotel will become expensive and popular and will be assigned to cluster B. In addition to yearly seasonality, we can observe weekly seasonality as well (hotels are cheaper and not popular in weekdays, and more expensive and popular during weekends).</p>",
      "rawMarkdown": "[quote=Sangxia;116573]\r\n\r\n[quote=Manish Maheshwari;116568]\r\n\r\n@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...\r\n\r\n[/quote]\r\n\r\nNot sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?\r\n\r\n[/quote]\r\n\r\nCorrect, cluster assignments can change. For example, in winter a hotel is cheap and not-popular and will be assigned to cluster A, and in summer the same hotel will become expensive and popular and will be assigned to cluster B. In addition to yearly seasonality, we can observe weekly seasonality as well (hotels are cheaper and not popular in weekdays, and more expensive and popular during weekends).\r\n\r\n",
      "votes": 5
    },
    {
      "id": 116387,
      "postDate": "2016-04-23T20:37:05.683Z",
      "content": "<p>[quote=Adam;116277]\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\n[/quote]</p>\n\n<p>Hi Adam,</p>\n\n<p>Are hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?</p>\n\n<p>Thanks,\nkevin</p>",
      "rawMarkdown": "[quote=Adam;116277]\r\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\r\n[/quote]\r\n\r\nHi Adam,\r\n\r\nAre hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?\r\n\r\nThanks,\r\nkevin",
      "votes": 6
    },
    {
      "id": 118348,
      "postDate": "2016-05-03T11:32:18.233Z",
      "content": "<p>I am curious how did Arnaud originally found the data leak? and also if test data was leaked into train data then should't the complete rows match?</p>",
      "rawMarkdown": "I am curious how did Arnaud originally found the data leak? and also if test data was leaked into train data then should't the complete rows match?",
      "votes": 3
    },
    {
      "id": 121279,
      "postDate": "2016-05-25T10:03:03.143Z",
      "content": "<p>@Humberto Brand&#227;o, I agree with you.</p>\n\n<p>As we can see one or two admins for a competition is not enough or they are underpaid. Maybe there should be a small group of confidential Kagglers earning extra money and obtaining a prestigious rank for finding a leakage in the data set.  The competition would start two weeks after their work. Then our work will make sense.</p>",
      "rawMarkdown": "@Humberto Brandão, I agree with you.\r\n\r\nAs we can see one or two admins for a competition is not enough or they are underpaid. Maybe there should be a small group of confidential Kagglers earning extra money and obtaining a prestigious rank for finding a leakage in the data set.  The competition would start two weeks after their work. Then our work will make sense.\r\n",
      "votes": 4
    },
    {
      "id": 116594,
      "postDate": "2016-04-25T10:57:05.943Z",
      "content": "<p>[quote=vtKMH;116387]</p>\n\n<p>[quote=Adam;116277]\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\n[/quote]</p>\n\n<p>Hi Adam,</p>\n\n<p>Are hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?</p>\n\n<p>Thanks,\nkevin</p>\n\n<p>[/quote]</p>\n\n<p>Clustering is based on static hotel attributes and dynamic hotels attributes. Two most important dynamic variables are historical (log) price and historical (log) # of transactions. If possible, both of them are taken from the same season as check-in date but previous year.</p>",
      "rawMarkdown": "[quote=vtKMH;116387]\r\n\r\n[quote=Adam;116277]\r\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\r\n[/quote]\r\n\r\nHi Adam,\r\n\r\nAre hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?\r\n\r\nThanks,\r\nkevin\r\n\r\n[/quote]\r\n\r\nClustering is based on static hotel attributes and dynamic hotels attributes. Two most important dynamic variables are historical (log) price and historical (log) # of transactions. If possible, both of them are taken from the same season as check-in date but previous year.\r\n\r\n\r\n",
      "votes": 4
    },
    {
      "id": 117292,
      "postDate": "2016-04-28T11:52:57.947Z",
      "content": "<p>[quote=narsil;117291]</p>\n\n<p>[quote=Jos&#233; A. Guerrero;117278]</p>\n\n<p>[quote=Wendy Kan;116282]</p>\n\n<p>@Sergey, </p>\n\n<p>They won't be excluded.</p>\n\n<p>[/quote]</p>\n\n<p>So the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\nThe winner will be the best puzzle solver.</p>\n\n<p>Is more leaderboard efficient invest time in solve the puzzle.</p>\n\n<p>Could you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?</p>\n\n<p>[/quote]</p>\n\n<p>I guess it is too late now to exclude them. Many kagglers (including myself) spent a lot of time and effort to analyze the &quot;puzzles&quot; as you call them, and it would be unfair to trash those efforts now.</p>\n\n<p>At the same time, I agree with you in principle. However, imo, you have only one moment to set such issues: when you publish the first ruling. You cannot change the rules while the game is on. This would in turn give an unfair advantage to those who did not spend any time on &quot;puzzles&quot;.</p>\n\n<p>[/quote]</p>\n\n<p>Last time the decision was delete the rows for the scoring process, after two weeks of competition...</p>\n\n<p><a href=\"https://www.kaggle.com/c/cervical-cancer-screening/forums/t/18029/is-screener-and-patient-age/102621#post102621\">https://www.kaggle.com/c/cervical-cancer-screening/forums/t/18029/is-screener-and-patient-age/102621#post102621</a></p>\n\n<p>Don't worry. I like puzzles too!</p>",
      "rawMarkdown": "[quote=narsil;117291]\r\n\r\n[quote=José A. Guerrero;117278]\r\n\r\n[quote=Wendy Kan;116282]\r\n\r\n@Sergey, \r\n\r\nThey won't be excluded.\r\n\r\n[/quote]\r\n\r\nSo the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\r\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\r\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\r\nThe winner will be the best puzzle solver.\r\n\r\nIs more leaderboard efficient invest time in solve the puzzle.\r\n\r\nCould you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?\r\n\r\n\r\n[/quote]\r\n\r\nI guess it is too late now to exclude them. Many kagglers (including myself) spent a lot of time and effort to analyze the \"puzzles\" as you call them, and it would be unfair to trash those efforts now.\r\n\r\nAt the same time, I agree with you in principle. However, imo, you have only one moment to set such issues: when you publish the first ruling. You cannot change the rules while the game is on. This would in turn give an unfair advantage to those who did not spend any time on \"puzzles\".\r\n\r\n[/quote]\r\n\r\n\r\nLast time the decision was delete the rows for the scoring process, after two weeks of competition...\r\n\r\nhttps://www.kaggle.com/c/cervical-cancer-screening/forums/t/18029/is-screener-and-patient-age/102621#post102621\r\n\r\nDon't worry. I like puzzles too!\r\n \r\n",
      "votes": 2
    },
    {
      "id": 119222,
      "postDate": "2016-05-08T07:51:31.437Z",
      "content": "<p>[quote=Adam;116277]\nI confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. \nOur estimate is this directly affects approx. 1/3 rd of the holdout data. \n[/quote]\nIs it  allowed to utilize this leak by any way?\nThis potentially might affect up to 2/3 of the holdout data. </p>",
      "rawMarkdown": "[quote=Adam;116277]\r\nI confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. \r\nOur estimate is this directly affects approx. 1/3 rd of the holdout data. \r\n[/quote]\r\nIs it  allowed to utilize this leak by any way?\r\nThis potentially might affect up to 2/3 of the holdout data. ",
      "votes": 1
    },
    {
      "id": 117090,
      "postDate": "2016-04-27T11:16:47.587Z",
      "content": "<p>Unfortunately, I cannot provide you with this information. One reason is that in many cases boundaries between clusters are quite fuzzy.</p>",
      "rawMarkdown": "Unfortunately, I cannot provide you with this information. One reason is that in many cases boundaries between clusters are quite fuzzy.\r\n",
      "votes": 1
    },
    {
      "id": 116739,
      "postDate": "2016-04-25T20:46:58.537Z",
      "content": "<p>One more clarification: the hotels assigned to a cluster change, but what defines a cluster does not. </p>",
      "rawMarkdown": "One more clarification: the hotels assigned to a cluster change, but what defines a cluster does not. ",
      "votes": 1
    },
    {
      "id": 116573,
      "postDate": "2016-04-25T08:29:52.100Z",
      "content": "<p>[quote=Manish Maheshwari;116568]</p>\n\n<p>@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...</p>\n\n<p>[/quote]</p>\n\n<p>Not sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?</p>",
      "rawMarkdown": "[quote=Manish Maheshwari;116568]\r\n\r\n@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...\r\n\r\n[/quote]\r\n\r\nNot sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?",
      "votes": 1
    },
    {
      "id": 116282,
      "postDate": "2016-04-22T21:11:59.553Z",
      "content": "<p>@Sergey, </p>\n\n<p>They won't be excluded.</p>",
      "rawMarkdown": "@Sergey, \r\n\r\nThey won't be excluded.",
      "votes": 1
    },
    {
      "id": 116281,
      "postDate": "2016-04-22T21:08:20.357Z",
      "content": "<p>Thanks. Are you going to continue using effected rows in score calculations or they will be excluded?</p>",
      "rawMarkdown": "Thanks. Are you going to continue using effected rows in score calculations or they will be excluded?\r\n",
      "votes": 1
    },
    {
      "id": 117359,
      "postDate": "2016-04-28T18:01:05.633Z",
      "content": "<p>Navot, </p>\n\n<p>Data leakage is explained at wiki- <a href=\"https://www.kaggle.com/wiki/Leakage\">https://www.kaggle.com/wiki/Leakage</a></p>\n\n<p>[quote=Navot Silberstein;117305]</p>\n\n<p>Hi kagglers!\nI must admit I find the explanation confusing.\ncan someone please refer to a more thorough explanation of what exactly <strong>is</strong> the data leak?\n10x a lot</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Navot, \r\n\r\nData leakage is explained at wiki- https://www.kaggle.com/wiki/Leakage\r\n\r\n[quote=Navot Silberstein;117305]\r\n\r\nHi kagglers!\r\nI must admit I find the explanation confusing.\r\ncan someone please refer to a more thorough explanation of what exactly **is** the data leak?\r\n10x a lot\r\n\r\n[/quote]\r\n",
      "votes": 2
    },
    {
      "id": 117305,
      "postDate": "2016-04-28T13:02:38.143Z",
      "content": "<p>Hi kagglers!\nI must admit I find the explanation confusing.\ncan someone please refer to a more thorough explanation of what exactly <strong>is</strong> the data leak?\n10x a lot</p>",
      "rawMarkdown": "Hi kagglers!\r\nI must admit I find the explanation confusing.\r\ncan someone please refer to a more thorough explanation of what exactly **is** the data leak?\r\n10x a lot",
      "votes": 2
    },
    {
      "id": 116596,
      "postDate": "2016-04-25T10:59:40.490Z",
      "content": "<p>[quote=Marcin P&#281;kalski;116549]</p>\n\n<p>I meant that within one search session for a user you may have varying hotel_clusters withing the same srch_destination_id, and you will not know if the change is caused by a hotel changing cluster or if it is user checking a different hotel. </p>\n\n<p>[/quote]</p>\n\n<p>For the same check-in dates and the same srch_destination_id, it usually means that a user is checking a different hotel. </p>",
      "rawMarkdown": "[quote=Marcin Pękalski;116549]\r\n\r\nI meant that within one search session for a user you may have varying hotel_clusters withing the same srch_destination_id, and you will not know if the change is caused by a hotel changing cluster or if it is user checking a different hotel. \r\n\r\n[/quote]\r\n\r\nFor the same check-in dates and the same srch_destination_id, it usually means that a user is checking a different hotel. ",
      "votes": 2
    },
    {
      "id": 121303,
      "postDate": "2016-05-25T13:20:26.773Z",
      "rawMarkdown": "",
      "votes": -2
    },
    {
      "id": 121847,
      "postDate": "2016-05-30T09:38:36.777Z",
      "content": "<p>How do we 'utilize' the data leak?? What I am able to understand is we may have to use a benchmark for 33% of the data..and use ML algo on the rest of it!! But m not sure how to do this..Can someone help?!</p>",
      "rawMarkdown": "How do we 'utilize' the data leak?? What I am able to understand is we may have to use a benchmark for 33% of the data..and use ML algo on the rest of it!! But m not sure how to do this..Can someone help?!"
    },
    {
      "id": 121764,
      "postDate": "2016-05-29T15:02:50.963Z",
      "content": "<p>[quote=TomM;120118]</p>\n\n<p>I read Arnaud's description of the leak and now it makes sense. Should have read more deeply into the forums before posting...</p>\n\n<p>[/quote]</p>\n\n<p>@TomM, can you tell me where can I find Arnaud's description? Is this ( <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20282/a-leak-in-the-data?limit=all\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20282/a-leak-in-the-data?limit=all</a> ) what you mean?</p>",
      "rawMarkdown": "[quote=TomM;120118]\r\n\r\nI read Arnaud's description of the leak and now it makes sense. Should have read more deeply into the forums before posting...\r\n\r\n[/quote]\r\n\r\n@TomM, can you tell me where can I find Arnaud's description? Is this ( https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20282/a-leak-in-the-data?limit=all ) what you mean?"
    },
    {
      "id": 121398,
      "postDate": "2016-05-26T04:04:43.307Z",
      "content": "<p>Can I know how many data points from the test set (the &quot;test.csv&quot; file given, which contains ~2,000,000 data points) is affected? I just want to make sure that what I am doing is correct. Thanks.</p>",
      "rawMarkdown": "Can I know how many data points from the test set (the \"test.csv\" file given, which contains ~2,000,000 data points) is affected? I just want to make sure that what I am doing is correct. Thanks."
    },
    {
      "id": 121325,
      "postDate": "2016-05-25T16:53:30.423Z",
      "content": "<p>[quote=Humberto Brand&#227;o;121250]</p>\n\n<p>and I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...</p>\n\n<p>[/quote]</p>\n\n<p>What it was? Do you have a link?</p>",
      "rawMarkdown": "[quote=Humberto Brandão;121250]\r\n\r\nand I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...\r\n\r\n[/quote]\r\n\r\nWhat it was? Do you have a link?\r\n"
    },
    {
      "id": 121321,
      "postDate": "2016-05-25T16:23:27.130Z",
      "content": "<p>I actually think the admins are spending their time removing the rows that would benefit from the data leak from the private LB test set&#8230;just to see all those scores generated from the scripts dropping to the end of the final LB:)</p>",
      "rawMarkdown": "I actually think the admins are spending their time removing the rows that would benefit from the data leak from the private LB test set…just to see all those scores generated from the scripts dropping to the end of the final LB:)"
    },
    {
      "id": 120839,
      "postDate": "2016-05-20T21:22:37.633Z",
      "content": "<p>Realistically, I think if you had 97% efficiency on the leak portion of the data, your LB score would go way down.  The added score for having a correct HC in positions #3..#5 for leak-identified test rows, as thin as those positions may be, outweighs what can be accomplished by the other modeling.  My instinct is that around 90% efficiency for the leak data is about the max, balancing the desire to score higher on the LB.  FWIW, I no longer work on leak, I focus only on the remaining data, for which scoring 0.29 to 0.33 should result in a LB score of &gt;= 0.50000.</p>",
      "rawMarkdown": "Realistically, I think if you had 97% efficiency on the leak portion of the data, your LB score would go way down.  The added score for having a correct HC in positions #3..#5 for leak-identified test rows, as thin as those positions may be, outweighs what can be accomplished by the other modeling.  My instinct is that around 90% efficiency for the leak data is about the max, balancing the desire to score higher on the LB.  FWIW, I no longer work on leak, I focus only on the remaining data, for which scoring 0.29 to 0.33 should result in a LB score of >= 0.50000."
    },
    {
      "id": 120836,
      "postDate": "2016-05-20T20:22:52.893Z",
      "content": "<p>[quote=Piyush Jaiswal;120237]</p>\n\n<p>@Jos&#233; A. Guerrero</p>\n\n<p>What exactly do you mean by this 'puzzle solving' phrase? Could you please elaborate? </p>\n\n<p>[/quote]</p>\n\n<p>Just improve the leakage part of test dataset.</p>",
      "rawMarkdown": "[quote=Piyush Jaiswal;120237]\r\n\r\n@José A. Guerrero\r\n\r\nWhat exactly do you mean by this 'puzzle solving' phrase? Could you please elaborate? \r\n\r\n[/quote]\r\n\r\nJust improve the leakage part of test dataset."
    },
    {
      "id": 120820,
      "postDate": "2016-05-20T18:24:39.180Z",
      "content": "<p>Leakage is an unpleasant side effect. Probably, the awards will go to those competitors who will concentrate on improving leak efficiency from 0.87 to 0.97 (completely useless for the organizers) instead of improving the efficiency of main model from 0.25 to 0.30 (work that makes sense). It would be nice, if the organizers think about a consolation prize for the best model working on the test set without leakage.</p>",
      "rawMarkdown": "Leakage is an unpleasant side effect. Probably, the awards will go to those competitors who will concentrate on improving leak efficiency from 0.87 to 0.97 (completely useless for the organizers) instead of improving the efficiency of main model from 0.25 to 0.30 (work that makes sense). It would be nice, if the organizers think about a consolation prize for the best model working on the test set without leakage."
    },
    {
      "id": 120812,
      "postDate": "2016-05-20T17:37:52.063Z",
      "content": "<p>I frame my leak solution in terms of leak efficiency - if my leak solution put the correct HC in position 1 for every test row in my submission, I'd score it 100%.  Using Grzegorz' number as an example, 0.3373 would be the max score for a leak only submit; if an actual submission scored 0.2950, then 0.2950 / 0.3373 = 87.459% efficiency.  There are ways to get leak efficiency to 88%+.  In so doing, the percentage of test rows flagged as leak may go down.  If this the case, the golden question is: for the test rows removed, that say you match in position #5 for a score of 1/5 point, can you do better with your models, on average?  In the extreme case, say you could find 10,000 test rows flagged as leak and you know you are scoring 0% efficiency for them - then of course, you could only do better by omitting them from the leak portion of your submission and trying to score for them using other modeling.  The puzzle-solving aspect to me is a balance of raising leak efficiency toward 90%, while at the same time, raising score on the LB - the magic is to get both numbers to rise at the same time.</p>",
      "rawMarkdown": "I frame my leak solution in terms of leak efficiency - if my leak solution put the correct HC in position 1 for every test row in my submission, I'd score it 100%.  Using Grzegorz' number as an example, 0.3373 would be the max score for a leak only submit; if an actual submission scored 0.2950, then 0.2950 / 0.3373 = 87.459% efficiency.  There are ways to get leak efficiency to 88%+.  In so doing, the percentage of test rows flagged as leak may go down.  If this the case, the golden question is: for the test rows removed, that say you match in position #5 for a score of 1/5 point, can you do better with your models, on average?  In the extreme case, say you could find 10,000 test rows flagged as leak and you know you are scoring 0% efficiency for them - then of course, you could only do better by omitting them from the leak portion of your submission and trying to score for them using other modeling.  The puzzle-solving aspect to me is a balance of raising leak efficiency toward 90%, while at the same time, raising score on the LB - the magic is to get both numbers to rise at the same time."
    },
    {
      "id": 120811,
      "postDate": "2016-05-20T17:12:53.173Z",
      "content": "<p>@zyazzy, Leakage is not proportional to the number of orig_destination_distance available in test data set. It is proportional to the number of only those pairs of (user_location_city,orig_destination_distance) in test set which exist in train set. Only in that case you may assume (you shouldn't be sure) that hotel_cluster in test set is equal to that in train set.\nI have added a piece of code to my <a href=\"https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/run/242886\">Simple Validation</a> script (version 20) which calculates percentage of leakage. You may check the code (lines 95, 144, 237) and the result (33.73%)</p>",
      "rawMarkdown": "@zyazzy, Leakage is not proportional to the number of orig_destination_distance available in test data set. It is proportional to the number of only those pairs of (user_location_city,orig_destination_distance) in test set which exist in train set. Only in that case you may assume (you shouldn't be sure) that hotel_cluster in test set is equal to that in train set.\r\nI have added a piece of code to my [Simple Validation][1] script (version 20) which calculates percentage of leakage. You may check the code (lines 95, 144, 237) and the result (33.73%)\r\n\r\n\r\n  [1]: https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/run/242886"
    },
    {
      "id": 120807,
      "postDate": "2016-05-20T16:29:44.237Z",
      "content": "<p>To cite Adam's words: </p>\n\n<blockquote>\n  <p>Our estimate is this directly affects approx. 1/3 rd of the holdout data.</p>\n</blockquote>\n\n<p>He said 1/3 of holdout data, so ~66% of the data leaked in the test data is in the 1/3 of the test set used to calculate the public LB. The remaining 33% of the data leaked is still in the holdout data.</p>",
      "rawMarkdown": "To cite Adam's words: \r\n\r\n> Our estimate is this directly affects approx. 1/3 rd of the holdout data.\r\n\r\nHe said 1/3 of holdout data, so ~66% of the data leaked in the test data is in the 1/3 of the test set used to calculate the public LB. The remaining 33% of the data leaked is still in the holdout data.\r\n\r\n"
    },
    {
      "id": 120806,
      "postDate": "2016-05-20T16:25:02.630Z",
      "content": "<p>Why is everyone saying the data leak affects 33% of the test data? </p>\n\n<p>In my test set, orig_destination_distance is NA for 33% of the data, so data leak should affect 66% of the test data because the leaked parameter - orig_destination_distance - is available.</p>\n\n<p><code>test['orig_destination_distance'].isnull().sum()/len(test)</code></p>\n\n<pre><code>0.3351976056099038\n</code></pre>",
      "rawMarkdown": "Why is everyone saying the data leak affects 33% of the test data? \r\n\r\nIn my test set, orig_destination_distance is NA for 33% of the data, so data leak should affect 66% of the test data because the leaked parameter - orig_destination_distance - is available.\r\n\r\n`test['orig_destination_distance'].isnull().sum()/len(test)`\r\n\r\n    0.3351976056099038"
    },
    {
      "id": 120593,
      "postDate": "2016-05-19T11:12:37.233Z",
      "content": "<p>[quote=felix000;120584]</p>\n\n<p>I'm still trying to understand what is going on with the leak. The test set contains user booking events for 2015, correct? And we want to list up to 5 hotel clusters which they might like to book based on the session data. </p>\n\n<p>However, we know the history of bookings and can match their session data (user_location_city, orig_destination_distance) to other bookings and session activity from 2013 and 2014.</p>\n\n<p>However, the test set is 2015 and training is pre-2015. So is it a leak because the way the target has been defined by the administrators is also based on this same data from 2013 and 2014?</p>\n\n<p>[/quote]</p>\n\n<p>If I get it right, I think that the data leak is related to the fact that one of the fields (distance between origin and destination) is based on the target hotel that the user picked. since this is calculated with high precision it created a situation where user location fields + distance from destination creates a unique(ish) mapping to hotel cluster.</p>",
      "rawMarkdown": "[quote=felix000;120584]\r\n\r\nI'm still trying to understand what is going on with the leak. The test set contains user booking events for 2015, correct? And we want to list up to 5 hotel clusters which they might like to book based on the session data. \r\n\r\nHowever, we know the history of bookings and can match their session data (user_location_city, orig_destination_distance) to other bookings and session activity from 2013 and 2014.\r\n\r\nHowever, the test set is 2015 and training is pre-2015. So is it a leak because the way the target has been defined by the administrators is also based on this same data from 2013 and 2014?\r\n\r\n\r\n[/quote]\r\n\r\nIf I get it right, I think that the data leak is related to the fact that one of the fields (distance between origin and destination) is based on the target hotel that the user picked. since this is calculated with high precision it created a situation where user location fields + distance from destination creates a unique(ish) mapping to hotel cluster.\r\n"
    },
    {
      "id": 120592,
      "postDate": "2016-05-19T11:07:03.527Z",
      "content": "<p>[quote=felix000;120584]\nI'm still trying to understand what is going on with the leak...</p>\n\n<p>So is it a leak because the way the target has been defined by the administrators is also based on this same data from 2013 and 2014?\n[/quote]</p>\n\n<p>@felix000, Not exactly the same data. The same combination of two, three specific features is enough to change the problem from type (1) to (2):</p>\n\n<p>1) A customer has entered a shop. What will he buy,  if he bought a car last year and a shirt two years ago?</p>\n\n<p>2) A customer has entered a shop. What will he buy,  if other customers bought <strong>here</strong> shoes last year?</p>\n\n<p>EDIT: It is hard to write something new if the answer is in your question.</p>",
      "rawMarkdown": "[quote=felix000;120584]\r\nI'm still trying to understand what is going on with the leak...\r\n\r\n So is it a leak because the way the target has been defined by the administrators is also based on this same data from 2013 and 2014?\r\n[/quote]\r\n\r\n@felix000, Not exactly the same data. The same combination of two, three specific features is enough to change the problem from type (1) to (2):\r\n\r\n1) A customer has entered a shop. What will he buy,  if he bought a car last year and a shirt two years ago?\r\n\r\n2) A customer has entered a shop. What will he buy,  if other customers bought **here** shoes last year?\r\n\r\nEDIT: It is hard to write something new if the answer is in your question.\r\n"
    },
    {
      "id": 120316,
      "postDate": "2016-05-17T11:40:35.383Z",
      "content": "<p>I am not able to find the last 3 columns in the test and train csv files. i.e. is_booking, cnt and hotel_cluster. Am i missing something?</p>",
      "rawMarkdown": "I am not able to find the last 3 columns in the test and train csv files. i.e. is_booking, cnt and hotel_cluster. Am i missing something?"
    },
    {
      "id": 120237,
      "postDate": "2016-05-16T16:08:34.410Z",
      "content": "<p>@Jos&#233; A. Guerrero</p>\n\n<p>What exactly do you mean by this 'puzzle solving' phrase? Could you please elaborate? </p>",
      "rawMarkdown": "@José A. Guerrero\r\n\r\nWhat exactly do you mean by this 'puzzle solving' phrase? Could you please elaborate? "
    },
    {
      "id": 120128,
      "postDate": "2016-05-15T18:42:36.460Z",
      "content": "<p>[quote=TomM;120115]\nIn my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?</p>\n\n<p>[/quote]\n@TomM, It is not needed to know the user_id to obtain data leakage. Common user_location_city is enough, because orig_destination_distance is not the distance from the user's house to the destination hotel, but (probably) the distance from the Post Office no. 1 in the user's city to the destination hotel.</p>",
      "rawMarkdown": "[quote=TomM;120115]\r\nIn my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?\r\n\r\n[/quote]\r\n@TomM, It is not needed to know the user_id to obtain data leakage. Common user_location_city is enough, because orig_destination_distance is not the distance from the user's house to the destination hotel, but (probably) the distance from the Post Office no. 1 in the user's city to the destination hotel."
    },
    {
      "id": 120118,
      "postDate": "2016-05-15T16:38:47.737Z",
      "content": "<p>I read Arnaud's description of the leak and now it makes sense. Should have read more deeply into the forums before posting...</p>\n\n<p>[quote=TomM;120115]</p>\n\n<p>I'm a bit confused by this. Is this really data leakage? To me this just seems like a property of the data. For example, if I'm from Palm Springs, CA (where a lot of wealthy people live) and I'm looking for a hotel in New York City, it's far more likely that I'm going to pick and expensive hotel. That's not data leakage, that's a model...</p>\n\n<p>In my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I read Arnaud's description of the leak and now it makes sense. Should have read more deeply into the forums before posting...\r\n\r\n\r\n[quote=TomM;120115]\r\n\r\nI'm a bit confused by this. Is this really data leakage? To me this just seems like a property of the data. For example, if I'm from Palm Springs, CA (where a lot of wealthy people live) and I'm looking for a hotel in New York City, it's far more likely that I'm going to pick and expensive hotel. That's not data leakage, that's a model...\r\n\r\nIn my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 120115,
      "postDate": "2016-05-15T16:20:39.470Z",
      "content": "<p>I'm a bit confused by this. Is this really data leakage? To me this just seems like a property of the data. For example, if I'm from Palm Springs, CA (where a lot of wealthy people live) and I'm looking for a hotel in New York City, it's far more likely that I'm going to pick and expensive hotel. That's not data leakage, that's a model...</p>\n\n<p>In my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?</p>",
      "rawMarkdown": "I'm a bit confused by this. Is this really data leakage? To me this just seems like a property of the data. For example, if I'm from Palm Springs, CA (where a lot of wealthy people live) and I'm looking for a hotel in New York City, it's far more likely that I'm going to pick and expensive hotel. That's not data leakage, that's a model...\r\n\r\nIn my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?"
    },
    {
      "id": 119680,
      "postDate": "2016-05-12T07:39:05.543Z",
      "content": "<p>Missing values in logs do happen. It's likely a system error.</p>",
      "rawMarkdown": "Missing values in logs do happen. It's likely a system error."
    },
    {
      "id": 119675,
      "postDate": "2016-05-12T07:07:51.550Z",
      "content": "<p>Hi Adam, \nThere are also missing data in the fields 'srch_ci' and 'srch_co'. What is the reason? How will it affect if those records are not considered?</p>\n\n<p>In[29]: trdf[['srch_ci','srch_co']].isnull().any()</p>\n\n<p>Out[27]: </p>\n\n<p>srch_ci    True</p>\n\n<p>srch_co    True</p>",
      "rawMarkdown": "Hi Adam, \r\nThere are also missing data in the fields 'srch_ci' and 'srch_co'. What is the reason? How will it affect if those records are not considered?\r\n\r\nIn[29]: trdf[['srch_ci','srch_co']].isnull().any()\r\n\r\nOut[27]: \r\n\r\nsrch_ci    True\r\n\r\nsrch_co    True"
    },
    {
      "id": 117291,
      "postDate": "2016-04-28T11:30:06.460Z",
      "content": "<p>[quote=Jos&#233; A. Guerrero;117278]</p>\n\n<p>[quote=Wendy Kan;116282]</p>\n\n<p>@Sergey, </p>\n\n<p>They won't be excluded.</p>\n\n<p>[/quote]</p>\n\n<p>So the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\nThe winner will be the best puzzle solver.</p>\n\n<p>Is more leaderboard efficient invest time in solve the puzzle.</p>\n\n<p>Could you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?</p>\n\n<p>[/quote]</p>\n\n<p>I guess it is too late now to exclude them. Many kagglers (including myself) spent a lot of time and effort to analyze the &quot;puzzles&quot; as you call them, and it would be unfair to trash those efforts now.</p>\n\n<p>At the same time, I agree with you in principle. However, imo, you have only one moment to set such issues: when you publish the first ruling. You cannot change the rules while the game is on. This would in turn give an unfair advantage to those who did not spend any time on &quot;puzzles&quot;.</p>",
      "rawMarkdown": "[quote=José A. Guerrero;117278]\r\n\r\n[quote=Wendy Kan;116282]\r\n\r\n@Sergey, \r\n\r\nThey won't be excluded.\r\n\r\n[/quote]\r\n\r\nSo the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\r\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\r\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\r\nThe winner will be the best puzzle solver.\r\n\r\nIs more leaderboard efficient invest time in solve the puzzle.\r\n\r\nCould you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?\r\n\r\n\r\n[/quote]\r\n\r\nI guess it is too late now to exclude them. Many kagglers (including myself) spent a lot of time and effort to analyze the \"puzzles\" as you call them, and it would be unfair to trash those efforts now.\r\n\r\nAt the same time, I agree with you in principle. However, imo, you have only one moment to set such issues: when you publish the first ruling. You cannot change the rules while the game is on. This would in turn give an unfair advantage to those who did not spend any time on \"puzzles\"."
    },
    {
      "id": 117061,
      "postDate": "2016-04-27T07:31:26.323Z",
      "content": "<p>@ Adam - can expedia give us brief descriptions of hotel clusters? I think this will help us add a lot of value for expedia.</p>\n\n<p>For example, when we know that for example cluster 1 is a premium group of 5-star hotels, we can detect traits in user behaviors which lead to booking of premium hotels. Knowing definition of cluster helps a lot in detecting traits in customer behavior.</p>",
      "rawMarkdown": "@ Adam - can expedia give us brief descriptions of hotel clusters? I think this will help us add a lot of value for expedia.\r\n\r\nFor example, when we know that for example cluster 1 is a premium group of 5-star hotels, we can detect traits in user behaviors which lead to booking of premium hotels. Knowing definition of cluster helps a lot in detecting traits in customer behavior."
    },
    {
      "id": 116661,
      "postDate": "2016-04-25T15:43:54.297Z",
      "content": "<p>[quote=Adam;116594]\nIf possible, both of them are taken from the same season as check-in date but previous year.\n[/quote]</p>\n\n<p>Great!  thanks for the clarification Adam!</p>",
      "rawMarkdown": "[quote=Adam;116594]\r\nIf possible, both of them are taken from the same season as check-in date but previous year.\r\n[/quote]\r\n\r\nGreat!  thanks for the clarification Adam!"
    },
    {
      "id": 116568,
      "postDate": "2016-04-25T07:53:32.693Z",
      "content": "<p>@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...</p>",
      "rawMarkdown": "@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change..."
    },
    {
      "id": 116549,
      "postDate": "2016-04-25T04:48:32.433Z",
      "content": "<p>I meant that within one search session for a user you may have varying hotel_clusters withing the same srch_destination_id, and you will not know if the change is caused by a hotel changing cluster or if it is user checking a different hotel. </p>",
      "rawMarkdown": "I meant that within one search session for a user you may have varying hotel_clusters withing the same srch_destination_id, and you will not know if the change is caused by a hotel changing cluster or if it is user checking a different hotel. "
    },
    {
      "id": 116546,
      "postDate": "2016-04-25T04:36:54.193Z",
      "content": "<p>[quote=Marcin P&#281;kalski;116389]</p>\n\n<p>Hotel cluster can vary even within user's session or searches, without booking.</p>\n\n<p>[/quote]</p>\n\n<p>What!!! @Adam - Can you pls clarify this</p>",
      "rawMarkdown": "[quote=Marcin Pękalski;116389]\r\n\r\nHotel cluster can vary even within user's session or searches, without booking.\r\n\r\n[/quote]\r\n\r\nWhat!!! @Adam - Can you pls clarify this"
    },
    {
      "id": 116389,
      "postDate": "2016-04-23T20:58:26.563Z",
      "content": "<p>Hotel cluster can vary even within user's session or searches, without booking.</p>",
      "rawMarkdown": "Hotel cluster can vary even within user's session or searches, without booking."
    },
    {
      "id": 120584,
      "postDate": "2016-05-19T09:17:17.170Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 117278,
      "author_name": "José A. Guerrero",
      "author_url": "",
      "post_date": "2016-04-28T10:31:56.693000",
      "content": "<p>[quote=Wendy Kan;116282]</p>\n\n<p>@Sergey, </p>\n\n<p>They won't be excluded.</p>\n\n<p>[/quote]</p>\n\n<p>So the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\nThe winner will be the best puzzle solver.</p>\n\n<p>Is more leaderboard efficient invest time in solve the puzzle.</p>\n\n<p>Could you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 116279,
      "author_name": "Kieran",
      "author_url": "",
      "post_date": "2016-04-22T21:02:44.277000",
      "content": "<p>Can someone explain this to me?</p>\n\n<p>How do I use this leak to remain competitive with the other submissions?</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 121250,
      "author_name": "Humberto Brandão, Ph.D.",
      "author_url": "",
      "post_date": "2016-05-25T06:26:05.897000",
      "content": "<p>Again... :(</p>\n\n<p>many people working and there is a data leak in the data set...</p>\n\n<p>kaggle needs to improve its data evaluation process before the competitions... maybe hiring companies from the same area of the competition to evaluate the data set... maybe hiring new data scientists from this comunity...</p>\n\n<p>The question is: we cannot spend our time in wrong data sets. </p>\n\n<p>I got bad times when I spent 2 months in the Winton Challenge and Kaggle changed the data set. And I can tell you: that was a very easy mistake to detect (considering that people know a little about stock markets). My friend and I detected the mistake almost at the same time. Few days after to start... Kaggle identified the mistake almost 2 months later... They did not identify the mistake because they were trying... they identify because the score of the first one told they (a quick improvement)... For this reason I think Kaggle needs to evaluate the possibility to validate data before competitions...</p>\n\n<p><strong>It is very serious. We are talking about our time!!!</strong></p>\n\n<p>This is just my opinion... </p>\n\n<p>and I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 116597,
      "author_name": "Adam",
      "author_url": "",
      "post_date": "2016-04-25T11:07:02.323000",
      "content": "<p>[quote=Sangxia;116573]</p>\n\n<p>[quote=Manish Maheshwari;116568]</p>\n\n<p>@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...</p>\n\n<p>[/quote]</p>\n\n<p>Not sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?</p>\n\n<p>[/quote]</p>\n\n<p>Correct, cluster assignments can change. For example, in winter a hotel is cheap and not-popular and will be assigned to cluster A, and in summer the same hotel will become expensive and popular and will be assigned to cluster B. In addition to yearly seasonality, we can observe weekly seasonality as well (hotels are cheaper and not popular in weekdays, and more expensive and popular during weekends).</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 116387,
      "author_name": "vtKMH",
      "author_url": "",
      "post_date": "2016-04-23T20:37:05.683000",
      "content": "<p>[quote=Adam;116277]\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\n[/quote]</p>\n\n<p>Hi Adam,</p>\n\n<p>Are hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?</p>\n\n<p>Thanks,\nkevin</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 118348,
      "author_name": "Sameh Faidi",
      "author_url": "",
      "post_date": "2016-05-03T11:32:18.233000",
      "content": "<p>I am curious how did Arnaud originally found the data leak? and also if test data was leaked into train data then should't the complete rows match?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 121279,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2016-05-25T10:03:03.143000",
      "content": "<p>@Humberto Brand&#227;o, I agree with you.</p>\n\n<p>As we can see one or two admins for a competition is not enough or they are underpaid. Maybe there should be a small group of confidential Kagglers earning extra money and obtaining a prestigious rank for finding a leakage in the data set.  The competition would start two weeks after their work. Then our work will make sense.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 116594,
      "author_name": "Adam",
      "author_url": "",
      "post_date": "2016-04-25T10:57:05.943000",
      "content": "<p>[quote=vtKMH;116387]</p>\n\n<p>[quote=Adam;116277]\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\n[/quote]</p>\n\n<p>Hi Adam,</p>\n\n<p>Are hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?</p>\n\n<p>Thanks,\nkevin</p>\n\n<p>[/quote]</p>\n\n<p>Clustering is based on static hotel attributes and dynamic hotels attributes. Two most important dynamic variables are historical (log) price and historical (log) # of transactions. If possible, both of them are taken from the same season as check-in date but previous year.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 117292,
      "author_name": "José A. Guerrero",
      "author_url": "",
      "post_date": "2016-04-28T11:52:57.947000",
      "content": "<p>[quote=narsil;117291]</p>\n\n<p>[quote=Jos&#233; A. Guerrero;117278]</p>\n\n<p>[quote=Wendy Kan;116282]</p>\n\n<p>@Sergey, </p>\n\n<p>They won't be excluded.</p>\n\n<p>[/quote]</p>\n\n<p>So the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\nThe winner will be the best puzzle solver.</p>\n\n<p>Is more leaderboard efficient invest time in solve the puzzle.</p>\n\n<p>Could you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?</p>\n\n<p>[/quote]</p>\n\n<p>I guess it is too late now to exclude them. Many kagglers (including myself) spent a lot of time and effort to analyze the &quot;puzzles&quot; as you call them, and it would be unfair to trash those efforts now.</p>\n\n<p>At the same time, I agree with you in principle. However, imo, you have only one moment to set such issues: when you publish the first ruling. You cannot change the rules while the game is on. This would in turn give an unfair advantage to those who did not spend any time on &quot;puzzles&quot;.</p>\n\n<p>[/quote]</p>\n\n<p>Last time the decision was delete the rows for the scoring process, after two weeks of competition...</p>\n\n<p><a href=\"https://www.kaggle.com/c/cervical-cancer-screening/forums/t/18029/is-screener-and-patient-age/102621#post102621\">https://www.kaggle.com/c/cervical-cancer-screening/forums/t/18029/is-screener-and-patient-age/102621#post102621</a></p>\n\n<p>Don't worry. I like puzzles too!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 119222,
      "author_name": "Victor",
      "author_url": "",
      "post_date": "2016-05-08T07:51:31.437000",
      "content": "<p>[quote=Adam;116277]\nI confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. \nOur estimate is this directly affects approx. 1/3 rd of the holdout data. \n[/quote]\nIs it  allowed to utilize this leak by any way?\nThis potentially might affect up to 2/3 of the holdout data. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 117090,
      "author_name": "Adam",
      "author_url": "",
      "post_date": "2016-04-27T11:16:47.587000",
      "content": "<p>Unfortunately, I cannot provide you with this information. One reason is that in many cases boundaries between clusters are quite fuzzy.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116739,
      "author_name": "MPC24",
      "author_url": "",
      "post_date": "2016-04-25T20:46:58.537000",
      "content": "<p>One more clarification: the hotels assigned to a cluster change, but what defines a cluster does not. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116573,
      "author_name": "Sangxia",
      "author_url": "",
      "post_date": "2016-04-25T08:29:52.100000",
      "content": "<p>[quote=Manish Maheshwari;116568]</p>\n\n<p>@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...</p>\n\n<p>[/quote]</p>\n\n<p>Not sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116282,
      "author_name": "Wendy Kan",
      "author_url": "",
      "post_date": "2016-04-22T21:11:59.553000",
      "content": "<p>@Sergey, </p>\n\n<p>They won't be excluded.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 116281,
      "author_name": "Sergey Yurgenson",
      "author_url": "",
      "post_date": "2016-04-22T21:08:20.357000",
      "content": "<p>Thanks. Are you going to continue using effected rows in score calculations or they will be excluded?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 117359,
      "author_name": "Mrityunjay",
      "author_url": "",
      "post_date": "2016-04-28T18:01:05.633000",
      "content": "<p>Navot, </p>\n\n<p>Data leakage is explained at wiki- <a href=\"https://www.kaggle.com/wiki/Leakage\">https://www.kaggle.com/wiki/Leakage</a></p>\n\n<p>[quote=Navot Silberstein;117305]</p>\n\n<p>Hi kagglers!\nI must admit I find the explanation confusing.\ncan someone please refer to a more thorough explanation of what exactly <strong>is</strong> the data leak?\n10x a lot</p>\n\n<p>[/quote]</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 117305,
      "author_name": "Navot Silberstein",
      "author_url": "",
      "post_date": "2016-04-28T13:02:38.143000",
      "content": "<p>Hi kagglers!\nI must admit I find the explanation confusing.\ncan someone please refer to a more thorough explanation of what exactly <strong>is</strong> the data leak?\n10x a lot</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 116596,
      "author_name": "Adam",
      "author_url": "",
      "post_date": "2016-04-25T10:59:40.490000",
      "content": "<p>[quote=Marcin P&#281;kalski;116549]</p>\n\n<p>I meant that within one search session for a user you may have varying hotel_clusters withing the same srch_destination_id, and you will not know if the change is caused by a hotel changing cluster or if it is user checking a different hotel. </p>\n\n<p>[/quote]</p>\n\n<p>For the same check-in dates and the same srch_destination_id, it usually means that a user is checking a different hotel. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 121303,
      "author_name": "Subhajit Mandal",
      "author_url": "",
      "post_date": "2016-05-25T13:20:26.773000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 121847,
      "author_name": "Abhinav Sharma ",
      "author_url": "",
      "post_date": "2016-05-30T09:38:36.777000",
      "content": "<p>How do we 'utilize' the data leak?? What I am able to understand is we may have to use a benchmark for 33% of the data..and use ML algo on the rest of it!! But m not sure how to do this..Can someone help?!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121764,
      "author_name": "Hye Jin Jang",
      "author_url": "",
      "post_date": "2016-05-29T15:02:50.963000",
      "content": "<p>[quote=TomM;120118]</p>\n\n<p>I read Arnaud's description of the leak and now it makes sense. Should have read more deeply into the forums before posting...</p>\n\n<p>[/quote]</p>\n\n<p>@TomM, can you tell me where can I find Arnaud's description? Is this ( <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20282/a-leak-in-the-data?limit=all\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20282/a-leak-in-the-data?limit=all</a> ) what you mean?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121398,
      "author_name": "YongXien",
      "author_url": "",
      "post_date": "2016-05-26T04:04:43.307000",
      "content": "<p>Can I know how many data points from the test set (the &quot;test.csv&quot; file given, which contains ~2,000,000 data points) is affected? I just want to make sure that what I am doing is correct. Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121325,
      "author_name": "==>(AL)<==",
      "author_url": "",
      "post_date": "2016-05-25T16:53:30.423000",
      "content": "<p>[quote=Humberto Brand&#227;o;121250]</p>\n\n<p>and I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...</p>\n\n<p>[/quote]</p>\n\n<p>What it was? Do you have a link?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121321,
      "author_name": "Diego",
      "author_url": "",
      "post_date": "2016-05-25T16:23:27.130000",
      "content": "<p>I actually think the admins are spending their time removing the rows that would benefit from the data leak from the private LB test set&#8230;just to see all those scores generated from the scripts dropping to the end of the final LB:)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120839,
      "author_name": "eipiplus1",
      "author_url": "",
      "post_date": "2016-05-20T21:22:37.633000",
      "content": "<p>Realistically, I think if you had 97% efficiency on the leak portion of the data, your LB score would go way down.  The added score for having a correct HC in positions #3..#5 for leak-identified test rows, as thin as those positions may be, outweighs what can be accomplished by the other modeling.  My instinct is that around 90% efficiency for the leak data is about the max, balancing the desire to score higher on the LB.  FWIW, I no longer work on leak, I focus only on the remaining data, for which scoring 0.29 to 0.33 should result in a LB score of &gt;= 0.50000.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120836,
      "author_name": "José A. Guerrero",
      "author_url": "",
      "post_date": "2016-05-20T20:22:52.893000",
      "content": "<p>[quote=Piyush Jaiswal;120237]</p>\n\n<p>@Jos&#233; A. Guerrero</p>\n\n<p>What exactly do you mean by this 'puzzle solving' phrase? Could you please elaborate? </p>\n\n<p>[/quote]</p>\n\n<p>Just improve the leakage part of test dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120820,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2016-05-20T18:24:39.180000",
      "content": "<p>Leakage is an unpleasant side effect. Probably, the awards will go to those competitors who will concentrate on improving leak efficiency from 0.87 to 0.97 (completely useless for the organizers) instead of improving the efficiency of main model from 0.25 to 0.30 (work that makes sense). It would be nice, if the organizers think about a consolation prize for the best model working on the test set without leakage.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120812,
      "author_name": "eipiplus1",
      "author_url": "",
      "post_date": "2016-05-20T17:37:52.063000",
      "content": "<p>I frame my leak solution in terms of leak efficiency - if my leak solution put the correct HC in position 1 for every test row in my submission, I'd score it 100%.  Using Grzegorz' number as an example, 0.3373 would be the max score for a leak only submit; if an actual submission scored 0.2950, then 0.2950 / 0.3373 = 87.459% efficiency.  There are ways to get leak efficiency to 88%+.  In so doing, the percentage of test rows flagged as leak may go down.  If this the case, the golden question is: for the test rows removed, that say you match in position #5 for a score of 1/5 point, can you do better with your models, on average?  In the extreme case, say you could find 10,000 test rows flagged as leak and you know you are scoring 0% efficiency for them - then of course, you could only do better by omitting them from the leak portion of your submission and trying to score for them using other modeling.  The puzzle-solving aspect to me is a balance of raising leak efficiency toward 90%, while at the same time, raising score on the LB - the magic is to get both numbers to rise at the same time.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120811,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2016-05-20T17:12:53.173000",
      "content": "<p>@zyazzy, Leakage is not proportional to the number of orig_destination_distance available in test data set. It is proportional to the number of only those pairs of (user_location_city,orig_destination_distance) in test set which exist in train set. Only in that case you may assume (you shouldn't be sure) that hotel_cluster in test set is equal to that in train set.\nI have added a piece of code to my <a href=\"https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/run/242886\">Simple Validation</a> script (version 20) which calculates percentage of leakage. You may check the code (lines 95, 144, 237) and the result (33.73%)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120807,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-20T16:29:44.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120806,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-20T16:25:02.630000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120593,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-19T11:12:37.233000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120592,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-19T11:07:03.527000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120316,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-17T11:40:35.383000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120237,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-16T16:08:34.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120128,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-15T18:42:36.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120118,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-15T16:38:47.737000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120115,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-15T16:20:39.470000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119680,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-12T07:39:05.543000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119675,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-12T07:07:51.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 117291,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-28T11:30:06.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 117061,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-27T07:31:26.323000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116661,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-25T15:43:54.297000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116568,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-25T07:53:32.693000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116549,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-25T04:48:32.433000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116546,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-25T04:36:54.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 116389,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-23T20:58:26.563000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120584,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-19T09:17:17.170000",
      "content": "",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116277": "Thank you all for your patience for waiting for the clarification on the data leak.  We wanted to analyze the full extent before responding.\r\n\r\nI confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. Thanks to Arnaud who was the first to point this out. Good catch!\r\n\r\nOur estimate is this directly affects approx. 1/3 rd of the holdout data. \r\n\r\nThe contest will continue without any changes.  For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).\r\n\r\nI hope you will still enjoy the rest of the competition.\r\n\r\nAdam\r\n",
    "117278": "[quote=Wendy Kan;116282]\r\n\r\n@Sergey, \r\n\r\nThey won't be excluded.\r\n\r\n[/quote]\r\n\r\nSo the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\r\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\r\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\r\nThe winner will be the best puzzle solver.\r\n\r\nIs more leaderboard efficient invest time in solve the puzzle.\r\n\r\nCould you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?\r\n\r\n\r\n\r\n",
    "116279": "Can someone explain this to me?\r\n\r\nHow do I use this leak to remain competitive with the other submissions?\r\n\r\n",
    "121250": "Again... :(\r\n\r\nmany people working and there is a data leak in the data set...\r\n\r\nkaggle needs to improve its data evaluation process before the competitions... maybe hiring companies from the same area of the competition to evaluate the data set... maybe hiring new data scientists from this comunity...\r\n\r\nThe question is: we cannot spend our time in wrong data sets. \r\n\r\nI got bad times when I spent 2 months in the Winton Challenge and Kaggle changed the data set. And I can tell you: that was a very easy mistake to detect (considering that people know a little about stock markets). My friend and I detected the mistake almost at the same time. Few days after to start... Kaggle identified the mistake almost 2 months later... They did not identify the mistake because they were trying... they identify because the score of the first one told they (a quick improvement)... For this reason I think Kaggle needs to evaluate the possibility to validate data before competitions...\r\n\r\n**It is very serious. We are talking about our time!!!**\r\n\r\nThis is just my opinion... \r\n\r\nand I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...\r\n",
    "116597": "[quote=Sangxia;116573]\r\n\r\n[quote=Manish Maheshwari;116568]\r\n\r\n@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...\r\n\r\n[/quote]\r\n\r\nNot sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?\r\n\r\n[/quote]\r\n\r\nCorrect, cluster assignments can change. For example, in winter a hotel is cheap and not-popular and will be assigned to cluster A, and in summer the same hotel will become expensive and popular and will be assigned to cluster B. In addition to yearly seasonality, we can observe weekly seasonality as well (hotels are cheaper and not popular in weekdays, and more expensive and popular during weekends).\r\n\r\n",
    "116387": "[quote=Adam;116277]\r\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\r\n[/quote]\r\n\r\nHi Adam,\r\n\r\nAre hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?\r\n\r\nThanks,\r\nkevin",
    "118348": "I am curious how did Arnaud originally found the data leak? and also if test data was leaked into train data then should't the complete rows match?",
    "121279": "@Humberto Brandão, I agree with you.\r\n\r\nAs we can see one or two admins for a competition is not enough or they are underpaid. Maybe there should be a small group of confidential Kagglers earning extra money and obtaining a prestigious rank for finding a leakage in the data set.  The competition would start two weeks after their work. Then our work will make sense.\r\n",
    "116594": "[quote=vtKMH;116387]\r\n\r\n[quote=Adam;116277]\r\n... because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics)\r\n[/quote]\r\n\r\nHi Adam,\r\n\r\nAre hotel clusters assigned based on the booking date or based on the check-in date.  i.e. if Hotel X is in cluster 1 in the winter and cluster 2 in the summer, and in January I book a stay in that hotel for the following July...  will that record show up as being cluster 1 or 2?\r\n\r\nThanks,\r\nkevin\r\n\r\n[/quote]\r\n\r\nClustering is based on static hotel attributes and dynamic hotels attributes. Two most important dynamic variables are historical (log) price and historical (log) # of transactions. If possible, both of them are taken from the same season as check-in date but previous year.\r\n\r\n\r\n",
    "117292": "[quote=narsil;117291]\r\n\r\n[quote=José A. Guerrero;117278]\r\n\r\n[quote=Wendy Kan;116282]\r\n\r\n@Sergey, \r\n\r\nThey won't be excluded.\r\n\r\n[/quote]\r\n\r\nSo the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\r\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\r\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\r\nThe winner will be the best puzzle solver.\r\n\r\nIs more leaderboard efficient invest time in solve the puzzle.\r\n\r\nCould you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?\r\n\r\n\r\n[/quote]\r\n\r\nI guess it is too late now to exclude them. Many kagglers (including myself) spent a lot of time and effort to analyze the \"puzzles\" as you call them, and it would be unfair to trash those efforts now.\r\n\r\nAt the same time, I agree with you in principle. However, imo, you have only one moment to set such issues: when you publish the first ruling. You cannot change the rules while the game is on. This would in turn give an unfair advantage to those who did not spend any time on \"puzzles\".\r\n\r\n[/quote]\r\n\r\n\r\nLast time the decision was delete the rows for the scoring process, after two weeks of competition...\r\n\r\nhttps://www.kaggle.com/c/cervical-cancer-screening/forums/t/18029/is-screener-and-patient-age/102621#post102621\r\n\r\nDon't worry. I like puzzles too!\r\n \r\n",
    "119222": "[quote=Adam;116277]\r\nI confirm there is a leak in the data and it has to do with the orig_destination_distance attribute. \r\nOur estimate is this directly affects approx. 1/3 rd of the holdout data. \r\n[/quote]\r\nIs it  allowed to utilize this leak by any way?\r\nThis potentially might affect up to 2/3 of the holdout data. ",
    "117090": "Unfortunately, I cannot provide you with this information. One reason is that in many cases boundaries between clusters are quite fuzzy.\r\n",
    "116739": "One more clarification: the hotels assigned to a cluster change, but what defines a cluster does not. ",
    "116573": "[quote=Manish Maheshwari;116568]\r\n\r\n@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...\r\n\r\n[/quote]\r\n\r\nNot sure I fully understand. Could you elaborate? Does this mean that the hotel cluster tagging in test and train may be different because they are from different years?",
    "116282": "@Sergey, \r\n\r\nThey won't be excluded.",
    "116281": "Thanks. Are you going to continue using effected rows in score calculations or they will be excluded?\r\n",
    "117359": "Navot, \r\n\r\nData leakage is explained at wiki- https://www.kaggle.com/wiki/Leakage\r\n\r\n[quote=Navot Silberstein;117305]\r\n\r\nHi kagglers!\r\nI must admit I find the explanation confusing.\r\ncan someone please refer to a more thorough explanation of what exactly **is** the data leak?\r\n10x a lot\r\n\r\n[/quote]\r\n",
    "117305": "Hi kagglers!\r\nI must admit I find the explanation confusing.\r\ncan someone please refer to a more thorough explanation of what exactly **is** the data leak?\r\n10x a lot",
    "116596": "[quote=Marcin Pękalski;116549]\r\n\r\nI meant that within one search session for a user you may have varying hotel_clusters withing the same srch_destination_id, and you will not know if the change is caused by a hotel changing cluster or if it is user checking a different hotel. \r\n\r\n[/quote]\r\n\r\nFor the same check-in dates and the same srch_destination_id, it usually means that a user is checking a different hotel. ",
    "121303": "",
    "121847": "How do we 'utilize' the data leak?? What I am able to understand is we may have to use a benchmark for 33% of the data..and use ML algo on the rest of it!! But m not sure how to do this..Can someone help?!",
    "121764": "[quote=TomM;120118]\r\n\r\nI read Arnaud's description of the leak and now it makes sense. Should have read more deeply into the forums before posting...\r\n\r\n[/quote]\r\n\r\n@TomM, can you tell me where can I find Arnaud's description? Is this ( https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20282/a-leak-in-the-data?limit=all ) what you mean?",
    "121398": "Can I know how many data points from the test set (the \"test.csv\" file given, which contains ~2,000,000 data points) is affected? I just want to make sure that what I am doing is correct. Thanks.",
    "121325": "[quote=Humberto Brandão;121250]\r\n\r\nand I saw today that people identified another big mistake in the Draper data set... again!!! in another data set...\r\n\r\n[/quote]\r\n\r\nWhat it was? Do you have a link?\r\n",
    "121321": "I actually think the admins are spending their time removing the rows that would benefit from the data leak from the private LB test set…just to see all those scores generated from the scripts dropping to the end of the final LB:)",
    "120839": "Realistically, I think if you had 97% efficiency on the leak portion of the data, your LB score would go way down.  The added score for having a correct HC in positions #3..#5 for leak-identified test rows, as thin as those positions may be, outweighs what can be accomplished by the other modeling.  My instinct is that around 90% efficiency for the leak data is about the max, balancing the desire to score higher on the LB.  FWIW, I no longer work on leak, I focus only on the remaining data, for which scoring 0.29 to 0.33 should result in a LB score of >= 0.50000.",
    "120836": "[quote=Piyush Jaiswal;120237]\r\n\r\n@José A. Guerrero\r\n\r\nWhat exactly do you mean by this 'puzzle solving' phrase? Could you please elaborate? \r\n\r\n[/quote]\r\n\r\nJust improve the leakage part of test dataset.",
    "120820": "Leakage is an unpleasant side effect. Probably, the awards will go to those competitors who will concentrate on improving leak efficiency from 0.87 to 0.97 (completely useless for the organizers) instead of improving the efficiency of main model from 0.25 to 0.30 (work that makes sense). It would be nice, if the organizers think about a consolation prize for the best model working on the test set without leakage.",
    "120812": "I frame my leak solution in terms of leak efficiency - if my leak solution put the correct HC in position 1 for every test row in my submission, I'd score it 100%.  Using Grzegorz' number as an example, 0.3373 would be the max score for a leak only submit; if an actual submission scored 0.2950, then 0.2950 / 0.3373 = 87.459% efficiency.  There are ways to get leak efficiency to 88%+.  In so doing, the percentage of test rows flagged as leak may go down.  If this the case, the golden question is: for the test rows removed, that say you match in position #5 for a score of 1/5 point, can you do better with your models, on average?  In the extreme case, say you could find 10,000 test rows flagged as leak and you know you are scoring 0% efficiency for them - then of course, you could only do better by omitting them from the leak portion of your submission and trying to score for them using other modeling.  The puzzle-solving aspect to me is a balance of raising leak efficiency toward 90%, while at the same time, raising score on the LB - the magic is to get both numbers to rise at the same time.",
    "120811": "@zyazzy, Leakage is not proportional to the number of orig_destination_distance available in test data set. It is proportional to the number of only those pairs of (user_location_city,orig_destination_distance) in test set which exist in train set. Only in that case you may assume (you shouldn't be sure) that hotel_cluster in test set is equal to that in train set.\r\nI have added a piece of code to my [Simple Validation][1] script (version 20) which calculates percentage of leakage. You may check the code (lines 95, 144, 237) and the result (33.73%)\r\n\r\n\r\n  [1]: https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/run/242886",
    "120807": "To cite Adam's words: \r\n\r\n> Our estimate is this directly affects approx. 1/3 rd of the holdout data.\r\n\r\nHe said 1/3 of holdout data, so ~66% of the data leaked in the test data is in the 1/3 of the test set used to calculate the public LB. The remaining 33% of the data leaked is still in the holdout data.\r\n\r\n",
    "120806": "Why is everyone saying the data leak affects 33% of the test data? \r\n\r\nIn my test set, orig_destination_distance is NA for 33% of the data, so data leak should affect 66% of the test data because the leaked parameter - orig_destination_distance - is available.\r\n\r\n`test['orig_destination_distance'].isnull().sum()/len(test)`\r\n\r\n    0.3351976056099038",
    "120593": "[quote=felix000;120584]\r\n\r\nI'm still trying to understand what is going on with the leak. The test set contains user booking events for 2015, correct? And we want to list up to 5 hotel clusters which they might like to book based on the session data. \r\n\r\nHowever, we know the history of bookings and can match their session data (user_location_city, orig_destination_distance) to other bookings and session activity from 2013 and 2014.\r\n\r\nHowever, the test set is 2015 and training is pre-2015. So is it a leak because the way the target has been defined by the administrators is also based on this same data from 2013 and 2014?\r\n\r\n\r\n[/quote]\r\n\r\nIf I get it right, I think that the data leak is related to the fact that one of the fields (distance between origin and destination) is based on the target hotel that the user picked. since this is calculated with high precision it created a situation where user location fields + distance from destination creates a unique(ish) mapping to hotel cluster.\r\n",
    "120592": "[quote=felix000;120584]\r\nI'm still trying to understand what is going on with the leak...\r\n\r\n So is it a leak because the way the target has been defined by the administrators is also based on this same data from 2013 and 2014?\r\n[/quote]\r\n\r\n@felix000, Not exactly the same data. The same combination of two, three specific features is enough to change the problem from type (1) to (2):\r\n\r\n1) A customer has entered a shop. What will he buy,  if he bought a car last year and a shirt two years ago?\r\n\r\n2) A customer has entered a shop. What will he buy,  if other customers bought **here** shoes last year?\r\n\r\nEDIT: It is hard to write something new if the answer is in your question.\r\n",
    "120316": "I am not able to find the last 3 columns in the test and train csv files. i.e. is_booking, cnt and hotel_cluster. Am i missing something?",
    "120237": "@José A. Guerrero\r\n\r\nWhat exactly do you mean by this 'puzzle solving' phrase? Could you please elaborate? ",
    "120128": "[quote=TomM;120115]\r\nIn my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?\r\n\r\n[/quote]\r\n@TomM, It is not needed to know the user_id to obtain data leakage. Common user_location_city is enough, because orig_destination_distance is not the distance from the user's house to the destination hotel, but (probably) the distance from the Post Office no. 1 in the user's city to the destination hotel.",
    "120118": "I read Arnaud's description of the leak and now it makes sense. Should have read more deeply into the forums before posting...\r\n\r\n\r\n[quote=TomM;120115]\r\n\r\nI'm a bit confused by this. Is this really data leakage? To me this just seems like a property of the data. For example, if I'm from Palm Springs, CA (where a lot of wealthy people live) and I'm looking for a hotel in New York City, it's far more likely that I'm going to pick and expensive hotel. That's not data leakage, that's a model...\r\n\r\nIn my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?\r\n\r\n[/quote]\r\n",
    "120115": "I'm a bit confused by this. Is this really data leakage? To me this just seems like a property of the data. For example, if I'm from Palm Springs, CA (where a lot of wealthy people live) and I'm looking for a hotel in New York City, it's far more likely that I'm going to pick and expensive hotel. That's not data leakage, that's a model...\r\n\r\nIn my opinion, for it to be data leakage, there would need to be overlap between users in the test/train data sets or the same row in both data sets. Is that the actual case?",
    "119680": "Missing values in logs do happen. It's likely a system error.",
    "119675": "Hi Adam, \r\nThere are also missing data in the fields 'srch_ci' and 'srch_co'. What is the reason? How will it affect if those records are not considered?\r\n\r\nIn[29]: trdf[['srch_ci','srch_co']].isnull().any()\r\n\r\nOut[27]: \r\n\r\nsrch_ci    True\r\n\r\nsrch_co    True",
    "117291": "[quote=José A. Guerrero;117278]\r\n\r\n[quote=Wendy Kan;116282]\r\n\r\n@Sergey, \r\n\r\nThey won't be excluded.\r\n\r\n[/quote]\r\n\r\nSo the competition score is a mix, 33 % of a puzzle scoring and the rest of a true model.\r\nI don't understand the sense doesn't delete this rows in the evaluation (as in previous competition with similar problem).\r\nThink in a submission with a good performace in the valid model and a moderate puzzle solution performace  and, in other hand, a tight puzzle solver reaching nearly perfect performance in the 33% and average in valid model.\r\nThe winner will be the best puzzle solver.\r\n\r\nIs more leaderboard efficient invest time in solve the puzzle.\r\n\r\nCould you do an internal test and compute the scores without affected rows and check the robustness of the current leaderboard?\r\n\r\n\r\n[/quote]\r\n\r\nI guess it is too late now to exclude them. Many kagglers (including myself) spent a lot of time and effort to analyze the \"puzzles\" as you call them, and it would be unfair to trash those efforts now.\r\n\r\nAt the same time, I agree with you in principle. However, imo, you have only one moment to set such issues: when you publish the first ruling. You cannot change the rules while the game is on. This would in turn give an unfair advantage to those who did not spend any time on \"puzzles\".",
    "117061": "@ Adam - can expedia give us brief descriptions of hotel clusters? I think this will help us add a lot of value for expedia.\r\n\r\nFor example, when we know that for example cluster 1 is a premium group of 5-star hotels, we can detect traits in user behaviors which lead to booking of premium hotels. Knowing definition of cluster helps a lot in detecting traits in customer behavior.",
    "116661": "[quote=Adam;116594]\r\nIf possible, both of them are taken from the same season as check-in date but previous year.\r\n[/quote]\r\n\r\nGreat!  thanks for the clarification Adam!",
    "116568": "@Marcin - Thats ok. But the hotel to cluster Tagging will change only once per Seasonal period. In one session, the user can search the same hotel multiple times. But the cluster tagging wont change...",
    "116549": "I meant that within one search session for a user you may have varying hotel_clusters withing the same srch_destination_id, and you will not know if the change is caused by a hotel changing cluster or if it is user checking a different hotel. ",
    "116546": "[quote=Marcin Pękalski;116389]\r\n\r\nHotel cluster can vary even within user's session or searches, without booking.\r\n\r\n[/quote]\r\n\r\nWhat!!! @Adam - Can you pls clarify this",
    "116389": "Hotel cluster can vary even within user's session or searches, without booking.",
    "120584": ""
  }
}