{
  "id": 20730,
  "title": "Data Leak Explanation",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20730",
  "author_name": "",
  "post_date": "2016-05-05T11:54:00.670Z",
  "votes": null,
  "comment_count": 4,
  "views": 1088,
  "content": "<p>Hello All,</p>\n\n<p>I have to say that so far, I have looked at many sites and videos to understand what is Data Leak but sadly, I have not found a good explanation that is simple and intuitive.</p>\n\n<p>Even the explanation on <a href=\"https://www.kaggle.com/wiki/Leakage\">Kaggle wiki</a> is confusing.</p>\n\n<p>It would be really great if someone can spend sometime to explain what is a Data Leak.</p>\n\n<p>Best Regards</p>\n\n<p>Chintan</p>",
  "messages": [
    {
      "id": "118800",
      "postDate": "05/05/2016 11:54:00",
      "content": "<p>Hello All,</p>\n\n<p>I have to say that so far, I have looked at many sites and videos to understand what is Data Leak but sadly, I have not found a good explanation that is simple and intuitive.</p>\n\n<p>Even the explanation on <a href=\"https://www.kaggle.com/wiki/Leakage\">Kaggle wiki</a> is confusing.</p>\n\n<p>It would be really great if someone can spend sometime to explain what is a Data Leak.</p>\n\n<p>Best Regards</p>\n\n<p>Chintan</p>",
      "rawMarkdown": "Hello All,\r\n\r\nI have to say that so far, I have looked at many sites and videos to understand what is Data Leak but sadly, I have not found a good explanation that is simple and intuitive.\r\n\r\nEven the explanation on [Kaggle wiki][1] is confusing.\r\n\r\nIt would be really great if someone can spend sometime to explain what is a Data Leak.\r\n\r\nBest Regards\r\n\r\nChintan\r\n\r\n\r\n  [1]: https://www.kaggle.com/wiki/Leakage",
      "votes": null
    },
    {
      "id": "118802",
      "postDate": "05/05/2016 12:13:09",
      "content": "<p>Data Leak is a problem in the construction of the training set for a ML algorithm.</p>\n\n<p>The most simple case would be to include the target variable (the class) as a input feature - a magic feature. If you do that, then most ML algorithms will build models based only on that feature, which would give a very high performance in CV, but the model itself would be useless (since you won't have that information for your test set).</p>\n\n<p>In the real world, this is pretty bad, and the symptom usually is a very high score in training (e.g. AUC almost 1.0), but with a very poor performance over the test set.</p>\n\n<p>Now, regarding this competition, the 'Data Leak' consists on the fact that the distances to the hotel from a given location are incredibly precise and unique. In some sense (assuming that hotels never change their clusters - which some do), this is the same as stating directly which hotel was booked, and, ultimately, the hotel cluster.</p>\n\n<p>And taking advantage of this is as easy as just reading the distance and the location (user_location_country, city and region) and searching in your training set for someone from the same location and the same distance. If you found someone, then it is very likely that he went to the same hotel - remember, distances are very precise - and thus you can just copy the hotel cluster.</p>\n\n<p>This won't work if the distance is missing from either the training or the train set. And there's a chance that the hotel changed it's cluster. But aside from that, it is a good way of figuring out the hotel cluster for some % of the test set samples. </p>",
      "rawMarkdown": "Data Leak is a problem in the construction of the training set for a ML algorithm.\r\n\r\nThe most simple case would be to include the target variable (the class) as a input feature - a magic feature. If you do that, then most ML algorithms will build models based only on that feature, which would give a very high performance in CV, but the model itself would be useless (since you won't have that information for your test set).\r\n\r\nIn the real world, this is pretty bad, and the symptom usually is a very high score in training (e.g. AUC almost 1.0), but with a very poor performance over the test set.\r\n\r\nNow, regarding this competition, the 'Data Leak' consists on the fact that the distances to the hotel from a given location are incredibly precise and unique. In some sense (assuming that hotels never change their clusters - which some do), this is the same as stating directly which hotel was booked, and, ultimately, the hotel cluster.\r\n\r\nAnd taking advantage of this is as easy as just reading the distance and the location (user_location_country, city and region) and searching in your training set for someone from the same location and the same distance. If you found someone, then it is very likely that he went to the same hotel - remember, distances are very precise - and thus you can just copy the hotel cluster.\r\n\r\nThis won't work if the distance is missing from either the training or the train set. And there's a chance that the hotel changed it's cluster. But aside from that, it is a good way of figuring out the hotel cluster for some % of the test set samples.",
      "votes": null
    },
    {
      "id": "118813",
      "postDate": "05/05/2016 14:00:51",
      "content": "<p>@CarrDelling This has been very helpful. </p>\n\n<p>In that case, the model should not use the distance in the training?</p>",
      "rawMarkdown": "CarrDelling This has been very helpful. \r\n\r\nIn that case, the model should not use the distance in the training?",
      "votes": null
    },
    {
      "id": "118815",
      "postDate": "05/05/2016 14:30:01",
      "content": "<p>Not necessarily; this case (this competition) is different because we are given the true distances for about 1/3 of the test set (if I am not mistaken), so that could somewhat help the models (although it is much easier to just match the exact distance numbers outside of your modelling process).</p>\n\n<p>But only in this case. And you have to promise that you won't do the same in a real world problem ;)</p>",
      "rawMarkdown": "Not necessarily; this case (this competition) is different because we are given the true distances for about 1/3 of the test set (if I am not mistaken), so that could somewhat help the models (although it is much easier to just match the exact distance numbers outside of your modelling process).\r\n\r\nBut only in this case. And you have to promise that you won't do the same in a real world problem ;)",
      "votes": null
    },
    {
      "id": "118817",
      "postDate": "05/05/2016 14:34:21",
      "content": "<p>I am just getting started with Kaggle and Data Science so lets hope I don't do it in practice :)</p>",
      "rawMarkdown": "I am just getting started with Kaggle and Data Science so lets hope I don't do it in practice :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118802,
      "author_name": "carrdelling",
      "author_url": "",
      "post_date": "05/05/2016 12:13:09",
      "content": "<p>Data Leak is a problem in the construction of the training set for a ML algorithm.</p>\n\n<p>The most simple case would be to include the target variable (the class) as a input feature - a magic feature. If you do that, then most ML algorithms will build models based only on that feature, which would give a very high performance in CV, but the model itself would be useless (since you won't have that information for your test set).</p>\n\n<p>In the real world, this is pretty bad, and the symptom usually is a very high score in training (e.g. AUC almost 1.0), but with a very poor performance over the test set.</p>\n\n<p>Now, regarding this competition, the 'Data Leak' consists on the fact that the distances to the hotel from a given location are incredibly precise and unique. In some sense (assuming that hotels never change their clusters - which some do), this is the same as stating directly which hotel was booked, and, ultimately, the hotel cluster.</p>\n\n<p>And taking advantage of this is as easy as just reading the distance and the location (user_location_country, city and region) and searching in your training set for someone from the same location and the same distance. If you found someone, then it is very likely that he went to the same hotel - remember, distances are very precise - and thus you can just copy the hotel cluster.</p>\n\n<p>This won't work if the distance is missing from either the training or the train set. And there's a chance that the hotel changed it's cluster. But aside from that, it is a good way of figuring out the hotel cluster for some % of the test set samples. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118813,
      "author_name": "thecpshah",
      "author_url": "",
      "post_date": "05/05/2016 14:00:51",
      "content": "<p>@CarrDelling This has been very helpful. </p>\n\n<p>In that case, the model should not use the distance in the training?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118815,
      "author_name": "carrdelling",
      "author_url": "",
      "post_date": "05/05/2016 14:30:01",
      "content": "<p>Not necessarily; this case (this competition) is different because we are given the true distances for about 1/3 of the test set (if I am not mistaken), so that could somewhat help the models (although it is much easier to just match the exact distance numbers outside of your modelling process).</p>\n\n<p>But only in this case. And you have to promise that you won't do the same in a real world problem ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118817,
      "author_name": "thecpshah",
      "author_url": "",
      "post_date": "05/05/2016 14:34:21",
      "content": "<p>I am just getting started with Kaggle and Data Science so lets hope I don't do it in practice :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118800": "Hello All,\r\n\r\nI have to say that so far, I have looked at many sites and videos to understand what is Data Leak but sadly, I have not found a good explanation that is simple and intuitive.\r\n\r\nEven the explanation on [Kaggle wiki][1] is confusing.\r\n\r\nIt would be really great if someone can spend sometime to explain what is a Data Leak.\r\n\r\nBest Regards\r\n\r\nChintan\r\n\r\n\r\n  [1]: https://www.kaggle.com/wiki/Leakage",
    "118802": "Data Leak is a problem in the construction of the training set for a ML algorithm.\r\n\r\nThe most simple case would be to include the target variable (the class) as a input feature - a magic feature. If you do that, then most ML algorithms will build models based only on that feature, which would give a very high performance in CV, but the model itself would be useless (since you won't have that information for your test set).\r\n\r\nIn the real world, this is pretty bad, and the symptom usually is a very high score in training (e.g. AUC almost 1.0), but with a very poor performance over the test set.\r\n\r\nNow, regarding this competition, the 'Data Leak' consists on the fact that the distances to the hotel from a given location are incredibly precise and unique. In some sense (assuming that hotels never change their clusters - which some do), this is the same as stating directly which hotel was booked, and, ultimately, the hotel cluster.\r\n\r\nAnd taking advantage of this is as easy as just reading the distance and the location (user_location_country, city and region) and searching in your training set for someone from the same location and the same distance. If you found someone, then it is very likely that he went to the same hotel - remember, distances are very precise - and thus you can just copy the hotel cluster.\r\n\r\nThis won't work if the distance is missing from either the training or the train set. And there's a chance that the hotel changed it's cluster. But aside from that, it is a good way of figuring out the hotel cluster for some % of the test set samples.",
    "118813": "CarrDelling This has been very helpful. \r\n\r\nIn that case, the model should not use the distance in the training?",
    "118815": "Not necessarily; this case (this competition) is different because we are given the true distances for about 1/3 of the test set (if I am not mistaken), so that could somewhat help the models (although it is much easier to just match the exact distance numbers outside of your modelling process).\r\n\r\nBut only in this case. And you have to promise that you won't do the same in a real world problem ;)",
    "118817": "I am just getting started with Kaggle and Data Science so lets hope I don't do it in practice :)"
  },
  "source": "meta"
}