{
  "id": 477768,
  "title": "Negative correlation from CV to leaderboard",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/477768",
  "author_name": "David Cano Rosillo",
  "post_date": "2024-02-17T18:59:18.158000",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone! We've achieved a cross-validation performance of 0.8384 AUC and 0.6569 Gini stability metric. However, our leaderboard results are surprisingly poor. While we anticipate a drop in performance, our training scores surpass those of competitors ranked higher on the leaderboard.</p>\n<p>We acknowledge the risk of overfitting to the weeks in the training set, potentially resulting in a significant penalty on stability. Interestingly, competitors with similar cross-validation setups are outperforming us in submissions despite having lower AUC and Gini scores during training.</p>\n<p>We've noticed discussions where participants express concerns about the complexity of the metrics leading to increased randomness between the training and test sets.</p>\n<p>Our primary concern is the observed negative correlation between cross-validation and leaderboard performance, which we find puzzling.</p>\n<p>Any insights or similar experiences from your end would be greatly appreciated. Also what do you guys think is a reasonable score without metric hacking?</p>",
  "messages": [
    {
      "id": 2656495,
      "postDate": "2024-02-17T18:59:18.160Z",
      "content": "<p>Hello everyone! We've achieved a cross-validation performance of 0.8384 AUC and 0.6569 Gini stability metric. However, our leaderboard results are surprisingly poor. While we anticipate a drop in performance, our training scores surpass those of competitors ranked higher on the leaderboard.</p>\n<p>We acknowledge the risk of overfitting to the weeks in the training set, potentially resulting in a significant penalty on stability. Interestingly, competitors with similar cross-validation setups are outperforming us in submissions despite having lower AUC and Gini scores during training.</p>\n<p>We've noticed discussions where participants express concerns about the complexity of the metrics leading to increased randomness between the training and test sets.</p>\n<p>Our primary concern is the observed negative correlation between cross-validation and leaderboard performance, which we find puzzling.</p>\n<p>Any insights or similar experiences from your end would be greatly appreciated. Also what do you guys think is a reasonable score without metric hacking?</p>",
      "rawMarkdown": "Hello everyone! We've achieved a cross-validation performance of 0.8384 AUC and 0.6569 Gini stability metric. However, our leaderboard results are surprisingly poor. While we anticipate a drop in performance, our training scores surpass those of competitors ranked higher on the leaderboard.\n\nWe acknowledge the risk of overfitting to the weeks in the training set, potentially resulting in a significant penalty on stability. Interestingly, competitors with similar cross-validation setups are outperforming us in submissions despite having lower AUC and Gini scores during training.\n\nWe've noticed discussions where participants express concerns about the complexity of the metrics leading to increased randomness between the training and test sets.\n\nOur primary concern is the observed negative correlation between cross-validation and leaderboard performance, which we find puzzling.\n\nAny insights or similar experiences from your end would be greatly appreciated. Also what do you guys think is a reasonable score without metric hacking?",
      "votes": 5
    },
    {
      "id": 2656522,
      "postDate": "2024-02-17T19:19:22.160Z",
      "content": "<p>There is some amount of disconnect between the problem the produced models are meant to address (open up credit opporortunities), the evaluation metric and some of the data that is included in the train/test for modelling. Does that mean you should assume it's some weirdness causing what you're observing - no, but I think it's fair to suspect it may be.</p>\n<p>Re: performance, for a traditional metric like AUC for these kind of models I'd expect to see AUCs in the range of 0.60 to 0.80 from parametric models w/FICO information included. 0.50 to 0.70 seems like a fair range for non-parametric models without FICO (looks like about 5% of the data has something that behaves/looks like a transformed/obfuscated FICO score)</p>",
      "rawMarkdown": "There is some amount of disconnect between the problem the produced models are meant to address (open up credit opporortunities), the evaluation metric and some of the data that is included in the train/test for modelling. Does that mean you should assume it's some weirdness causing what you're observing - no, but I think it's fair to suspect it may be.\n\nRe: performance, for a traditional metric like AUC for these kind of models I'd expect to see AUCs in the range of 0.60 to 0.80 from parametric models w/FICO information included. 0.50 to 0.70 seems like a fair range for non-parametric models without FICO (looks like about 5% of the data has something that behaves/looks like a transformed/obfuscated FICO score)",
      "votes": 3
    },
    {
      "id": 2656700,
      "postDate": "2024-02-18T00:34:54.917Z",
      "content": "<p>The data that are presented in this competition are meant to represent what we normally have to deal with when developing risk models. They are not row data, but definitely not synthetic data. Naturally, you can expect the data to be changing in time (distribution of features). Metric will punish unstable solutions on test sample. You should think about better ways how to incorporate stability into the development just cross-validation. </p>",
      "rawMarkdown": "The data that are presented in this competition are meant to represent what we normally have to deal with when developing risk models. They are not row data, but definitely not synthetic data. Naturally, you can expect the data to be changing in time (distribution of features). Metric will punish unstable solutions on test sample. You should think about better ways how to incorporate stability into the development just cross-validation. ",
      "votes": 2,
      "replies": [
        {
          "id": 2656741,
          "postDate": "2024-02-18T01:37:46.803Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2656522,
      "author_name": "Sean McManus",
      "author_url": "",
      "post_date": "2024-02-17T19:19:22.160000",
      "content": "<p>There is some amount of disconnect between the problem the produced models are meant to address (open up credit opporortunities), the evaluation metric and some of the data that is included in the train/test for modelling. Does that mean you should assume it's some weirdness causing what you're observing - no, but I think it's fair to suspect it may be.</p>\n<p>Re: performance, for a traditional metric like AUC for these kind of models I'd expect to see AUCs in the range of 0.60 to 0.80 from parametric models w/FICO information included. 0.50 to 0.70 seems like a fair range for non-parametric models without FICO (looks like about 5% of the data has something that behaves/looks like a transformed/obfuscated FICO score)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2656700,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-18T00:34:54.917000",
      "content": "<p>The data that are presented in this competition are meant to represent what we normally have to deal with when developing risk models. They are not row data, but definitely not synthetic data. Naturally, you can expect the data to be changing in time (distribution of features). Metric will punish unstable solutions on test sample. You should think about better ways how to incorporate stability into the development just cross-validation. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2656741,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-18T01:37:46.803000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2656495": "Hello everyone! We've achieved a cross-validation performance of 0.8384 AUC and 0.6569 Gini stability metric. However, our leaderboard results are surprisingly poor. While we anticipate a drop in performance, our training scores surpass those of competitors ranked higher on the leaderboard.\n\nWe acknowledge the risk of overfitting to the weeks in the training set, potentially resulting in a significant penalty on stability. Interestingly, competitors with similar cross-validation setups are outperforming us in submissions despite having lower AUC and Gini scores during training.\n\nWe've noticed discussions where participants express concerns about the complexity of the metrics leading to increased randomness between the training and test sets.\n\nOur primary concern is the observed negative correlation between cross-validation and leaderboard performance, which we find puzzling.\n\nAny insights or similar experiences from your end would be greatly appreciated. Also what do you guys think is a reasonable score without metric hacking?",
    "2656522": "There is some amount of disconnect between the problem the produced models are meant to address (open up credit opporortunities), the evaluation metric and some of the data that is included in the train/test for modelling. Does that mean you should assume it's some weirdness causing what you're observing - no, but I think it's fair to suspect it may be.\n\nRe: performance, for a traditional metric like AUC for these kind of models I'd expect to see AUCs in the range of 0.60 to 0.80 from parametric models w/FICO information included. 0.50 to 0.70 seems like a fair range for non-parametric models without FICO (looks like about 5% of the data has something that behaves/looks like a transformed/obfuscated FICO score)",
    "2656700": "The data that are presented in this competition are meant to represent what we normally have to deal with when developing risk models. They are not row data, but definitely not synthetic data. Naturally, you can expect the data to be changing in time (distribution of features). Metric will punish unstable solutions on test sample. You should think about better ways how to incorporate stability into the development just cross-validation. "
  }
}