{
  "id": 508750,
  "title": "Data Leakage?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508750",
  "author_name": "",
  "post_date": "2024-05-30T20:30:22.057339200Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>We've seen a lot of discussion about the metric hacking. However it doesn't seem to be the only problem. Given the lack of a time series API, it seems that the design is also leaky: we can use future data to predict previous data. Any idea why this hasn't be discussed that much ? or is it because everyone has the same leak ? </p>\n<p>Edit: I've been trying to push a panel data dedicated cv object into sklearn (<a href=\"https://github.com/scikit-learn/scikit-learn/issues/28873\" target=\"_blank\">https://github.com/scikit-learn/scikit-learn/issues/28873</a>) feel free to chime in.  </p>",
  "messages": [
    {
      "id": "2846035",
      "postDate": "05/30/2024 20:30:22",
      "content": "<p>We've seen a lot of discussion about the metric hacking. However it doesn't seem to be the only problem. Given the lack of a time series API, it seems that the design is also leaky: we can use future data to predict previous data. Any idea why this hasn't be discussed that much ? or is it because everyone has the same leak ? </p>\n<p>Edit: I've been trying to push a panel data dedicated cv object into sklearn (<a href=\"https://github.com/scikit-learn/scikit-learn/issues/28873\" target=\"_blank\">https://github.com/scikit-learn/scikit-learn/issues/28873</a>) feel free to chime in.  </p>",
      "rawMarkdown": "We've seen a lot of discussion about the metric hacking. However it doesn't seem to be the only problem. Given the lack of a time series API, it seems that the design is also leaky: we can use future data to predict previous data. Any idea why this hasn't be discussed that much ? or is it because everyone has the same leak ? \n\nEdit: I've been trying to push a panel data dedicated cv object into sklearn (https://github.com/scikit-learn/scikit-learn/issues/28873) feel free to chime in.",
      "votes": null
    },
    {
      "id": "2846068",
      "postDate": "05/30/2024 20:58:43",
      "content": "<p>That is an excellent point and it falls in the same category as adversarial validation. The issue with that is that in order to prevent that, the evaluation should be done in sequential steps with more and more months of data to replicate the actual real-life scenario. It would be a major complication for scoring and the added benefit for kagglers and for the possible new findings in this competition weren't that high to be evaluated as worthy of persuading. </p>\n<p>The original idea was to keep the scoring as possible. </p>",
      "rawMarkdown": "That is an excellent point and it falls in the same category as adversarial validation. The issue with that is that in order to prevent that, the evaluation should be done in sequential steps with more and more months of data to replicate the actual real-life scenario. It would be a major complication for scoring and the added benefit for kagglers and for the possible new findings in this competition weren't that high to be evaluated as worthy of persuading. \n\nThe original idea was to keep the scoring as possible.",
      "votes": null
    },
    {
      "id": "2853098",
      "postDate": "06/03/2024 15:38:47",
      "content": "<p>As I have started figuring this myself, building a competition is quite difficult, and I understand the idea to keep it simple. <br>\nHowever, I am bit concerned the drop in auc is not due to model instability, but to data leakage. (i.e. better performance on most recent data is due to having more future data, and the drop of performance occurs along time we have less and less access to their future)</p>",
      "rawMarkdown": "As I have started figuring this myself, building a competition is quite difficult, and I understand the idea to keep it simple. \nHowever, I am bit concerned the drop in auc is not due to model instability, but to data leakage. (i.e. better performance on most recent data is due to having more future data, and the drop of performance occurs along time we have less and less access to their future)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2846068,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "05/30/2024 20:58:43",
      "content": "<p>That is an excellent point and it falls in the same category as adversarial validation. The issue with that is that in order to prevent that, the evaluation should be done in sequential steps with more and more months of data to replicate the actual real-life scenario. It would be a major complication for scoring and the added benefit for kagglers and for the possible new findings in this competition weren't that high to be evaluated as worthy of persuading. </p>\n<p>The original idea was to keep the scoring as possible. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2853098,
          "author_name": "lucasmorin",
          "author_url": "",
          "post_date": "06/03/2024 15:38:47",
          "content": "<p>As I have started figuring this myself, building a competition is quite difficult, and I understand the idea to keep it simple. <br>\nHowever, I am bit concerned the drop in auc is not due to model instability, but to data leakage. (i.e. better performance on most recent data is due to having more future data, and the drop of performance occurs along time we have less and less access to their future)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2846035": "We've seen a lot of discussion about the metric hacking. However it doesn't seem to be the only problem. Given the lack of a time series API, it seems that the design is also leaky: we can use future data to predict previous data. Any idea why this hasn't be discussed that much ? or is it because everyone has the same leak ? \n\nEdit: I've been trying to push a panel data dedicated cv object into sklearn (https://github.com/scikit-learn/scikit-learn/issues/28873) feel free to chime in.",
    "2846068": "That is an excellent point and it falls in the same category as adversarial validation. The issue with that is that in order to prevent that, the evaluation should be done in sequential steps with more and more months of data to replicate the actual real-life scenario. It would be a major complication for scoring and the added benefit for kagglers and for the possible new findings in this competition weren't that high to be evaluated as worthy of persuading. \n\nThe original idea was to keep the scoring as possible.",
    "2853098": "As I have started figuring this myself, building a competition is quite difficult, and I understand the idea to keep it simple. \nHowever, I am bit concerned the drop in auc is not due to model instability, but to data leakage. (i.e. better performance on most recent data is due to having more future data, and the drop of performance occurs along time we have less and less access to their future)"
  },
  "source": "meta"
}