{
  "id": 20672,
  "title": "A question for adam about the data leak?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20672",
  "author_name": "theonlyone",
  "post_date": "2016-05-03T13:31:04.950000",
  "votes": 1,
  "comment_count": 0,
  "views": 418,
  "content": "<p>Adam wrote - &quot;The contest will continue without any changes. For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).&quot;</p>\n\n<p>I have read kaggle's post about data leak.\n1. I am a bit confused about the meaning of the leak.\nWhat is said, is that the: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance columns are highly predictive of the hotel_cluster in the test data (1/3 of the rows)?</p>\n\n<ol start=\"2\">\n<li>How is that caused? I mean we have logs from past events,\nWhat is the scenario in which they could predict the future?\nI would love a clarification.</li>\n</ol>",
  "messages": [
    {
      "id": 118382,
      "postDate": "2016-05-03T13:31:04.950Z",
      "content": "<p>Adam wrote - &quot;The contest will continue without any changes. For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).&quot;</p>\n\n<p>I have read kaggle's post about data leak.\n1. I am a bit confused about the meaning of the leak.\nWhat is said, is that the: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance columns are highly predictive of the hotel_cluster in the test data (1/3 of the rows)?</p>\n\n<ol start=\"2\">\n<li>How is that caused? I mean we have logs from past events,\nWhat is the scenario in which they could predict the future?\nI would love a clarification.</li>\n</ol>",
      "rawMarkdown": "Adam wrote - \"The contest will continue without any changes. For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).\"\r\n\r\nI have read kaggle's post about data leak.\r\n1. I am a bit confused about the meaning of the leak.\r\nWhat is said, is that the: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance columns are highly predictive of the hotel_cluster in the test data (1/3 of the rows)?\r\n\r\n2. How is that caused? I mean we have logs from past events,\r\nWhat is the scenario in which they could predict the future?\r\nI would love a clarification.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "118382": "Adam wrote - \"The contest will continue without any changes. For clarity, we are confirming you can find hotel_clusters for the affected rows by matching rows from the train dataset based on the following columns: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance. However, this will not be 100% accurate because hotels can change cluster assignments (hotels popularity and price have seasonal characteristics).\"\r\n\r\nI have read kaggle's post about data leak.\r\n1. I am a bit confused about the meaning of the leak.\r\nWhat is said, is that the: user_location_country, user_location_region, user_location_city, hotel_market and orig_destination_distance columns are highly predictive of the hotel_cluster in the test data (1/3 of the rows)?\r\n\r\n2. How is that caused? I mean we have logs from past events,\r\nWhat is the scenario in which they could predict the future?\r\nI would love a clarification."
  }
}