{
  "id": 483125,
  "title": "What the Slope Term in the Metric Really Does",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/483125",
  "author_name": "",
  "post_date": "2024-03-11T02:50:04.952150800Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The hosts say the slope and standard deviation terms make a “stability metric,” but what do they actually accomplish? Well, the effect of the slope term is simple enough:</p>\n<p>Let's recall that the mean elevation of a line segment with endpoint elevations of <code>y1</code> and <code>y2</code> is <code>(y1+y2)/2</code>, and that the elevation of a point displaced by distance <code>dx</code> along a line from some initial point is <code>m dx + y0</code> where <code>m</code> is the slope and <code>y0</code> the initial elevation. Then when we have a metric of the form <code>mean(y) + 88*slope(y)</code>, we effectively have <code>(y1+y2)/2 + 88*m</code> which is equivalent to <code>(y1 + (y2 + 176*m))/2</code>. So we can see in the equivalent equation that the effect of the slope is to simulate extending the test period by 176 weeks or about 3 years by assuming a linear trend. </p>\n<p>This of course begs the question, if we want the effect of a longer test period, why not just use a longer test period? After all, the test period we have already includes past data, and extrapolation is hackable and prone to finding false trends particularly when we have the risk of features being dropped from the test set. The host's answer is that they have to use this metric because “it approximates well their business decisions.“ But is this really a good reason? </p>\n<p>Cargo cultism is the naive copying of practices that are successful in one context to other contexts where they may not fit without regard for the purpose of the original practice. It is the fallacy that what works in one situation must work everywhere all the time. And it is all too common in large companies. </p>\n<p>Extrapolating future performance to decide whether to put a model into production is necessary because it is that model's future performance that will determine profitability. Does that need for extrapolation also apply to this contest? I think not. It would be pretty silly to immediately put the winning models from this competition into production considering that they've been trained on transformed data and, with the new changes, they don't even have the means to construct potentially useful time-based aggregation features. More likely the models will need to be retrained with the full feature set and then extrapolation can be performed.</p>\n<p>As much as extrapolation is necessary in production, it is entirely undesirable in competition scoring. It essentially requires placing negative weight on the scores of early rounds which both creates the possibility of hacks and effectively reduces the sample size of the test set making the scoring much more noisy. It replaces real test data with guessing, but competition ranks should not be guesses. This is an area where I think Kaggle should enforce some standards. To allow noisy and hackable metrics like this in ranked contests degrades the value of everyone's accomplishments on the platform.</p>\n<p>“But without the slope term, how can we test stability?” you might ask. Well as I've shown, the slope term only tests stability in the sense that it simulates a longer test set. Simply using a longer test set would do a better job of testing stability because doing so replaces guessing with real data. Slope is also a pretty poor measure of stability because it only takes into account linear performance decline and not random performance drift as a real stability metric should. The hosts might expect the standard deviation term to handle this, but that is likely to be more influenced by week to week performance variability than any tendency towards long-term drift. A longer test set with a real stability metric that is punishing to sustained periods of poor performance would do much more to test stability than the metric we have now, and could do so without the risk of hackability.</p>",
  "messages": [
    {
      "id": "2691066",
      "postDate": "03/11/2024 02:50:04",
      "content": "<p>The hosts say the slope and standard deviation terms make a “stability metric,” but what do they actually accomplish? Well, the effect of the slope term is simple enough:</p>\n<p>Let's recall that the mean elevation of a line segment with endpoint elevations of <code>y1</code> and <code>y2</code> is <code>(y1+y2)/2</code>, and that the elevation of a point displaced by distance <code>dx</code> along a line from some initial point is <code>m dx + y0</code> where <code>m</code> is the slope and <code>y0</code> the initial elevation. Then when we have a metric of the form <code>mean(y) + 88*slope(y)</code>, we effectively have <code>(y1+y2)/2 + 88*m</code> which is equivalent to <code>(y1 + (y2 + 176*m))/2</code>. So we can see in the equivalent equation that the effect of the slope is to simulate extending the test period by 176 weeks or about 3 years by assuming a linear trend. </p>\n<p>This of course begs the question, if we want the effect of a longer test period, why not just use a longer test period? After all, the test period we have already includes past data, and extrapolation is hackable and prone to finding false trends particularly when we have the risk of features being dropped from the test set. The host's answer is that they have to use this metric because “it approximates well their business decisions.“ But is this really a good reason? </p>\n<p>Cargo cultism is the naive copying of practices that are successful in one context to other contexts where they may not fit without regard for the purpose of the original practice. It is the fallacy that what works in one situation must work everywhere all the time. And it is all too common in large companies. </p>\n<p>Extrapolating future performance to decide whether to put a model into production is necessary because it is that model's future performance that will determine profitability. Does that need for extrapolation also apply to this contest? I think not. It would be pretty silly to immediately put the winning models from this competition into production considering that they've been trained on transformed data and, with the new changes, they don't even have the means to construct potentially useful time-based aggregation features. More likely the models will need to be retrained with the full feature set and then extrapolation can be performed.</p>\n<p>As much as extrapolation is necessary in production, it is entirely undesirable in competition scoring. It essentially requires placing negative weight on the scores of early rounds which both creates the possibility of hacks and effectively reduces the sample size of the test set making the scoring much more noisy. It replaces real test data with guessing, but competition ranks should not be guesses. This is an area where I think Kaggle should enforce some standards. To allow noisy and hackable metrics like this in ranked contests degrades the value of everyone's accomplishments on the platform.</p>\n<p>“But without the slope term, how can we test stability?” you might ask. Well as I've shown, the slope term only tests stability in the sense that it simulates a longer test set. Simply using a longer test set would do a better job of testing stability because doing so replaces guessing with real data. Slope is also a pretty poor measure of stability because it only takes into account linear performance decline and not random performance drift as a real stability metric should. The hosts might expect the standard deviation term to handle this, but that is likely to be more influenced by week to week performance variability than any tendency towards long-term drift. A longer test set with a real stability metric that is punishing to sustained periods of poor performance would do much more to test stability than the metric we have now, and could do so without the risk of hackability.</p>",
      "rawMarkdown": "The hosts say the slope and standard deviation terms make a “stability metric,” but what do they actually accomplish? Well, the effect of the slope term is simple enough:\n\nLet's recall that the mean elevation of a line segment with endpoint elevations of `y1` and `y2` is `(y1+y2)/2`, and that the elevation of a point displaced by distance `dx` along a line from some initial point is `m dx + y0` where `m` is the slope and `y0` the initial elevation. Then when we have a metric of the form `mean(y) + 88*slope(y)`, we effectively have `(y1+y2)/2 + 88*m` which is equivalent to `(y1 + (y2 + 176*m))/2`. So we can see in the equivalent equation that the effect of the slope is to simulate extending the test period by 176 weeks or about 3 years by assuming a linear trend. \n\nThis of course begs the question, if we want the effect of a longer test period, why not just use a longer test period? After all, the test period we have already includes past data, and extrapolation is hackable and prone to finding false trends particularly when we have the risk of features being dropped from the test set. The host's answer is that they have to use this metric because “it approximates well their business decisions.“ But is this really a good reason? \n\nCargo cultism is the naive copying of practices that are successful in one context to other contexts where they may not fit without regard for the purpose of the original practice. It is the fallacy that what works in one situation must work everywhere all the time. And it is all too common in large companies. \n\nExtrapolating future performance to decide whether to put a model into production is necessary because it is that model's future performance that will determine profitability. Does that need for extrapolation also apply to this contest? I think not. It would be pretty silly to immediately put the winning models from this competition into production considering that they've been trained on transformed data and, with the new changes, they don't even have the means to construct potentially useful time-based aggregation features. More likely the models will need to be retrained with the full feature set and then extrapolation can be performed.\n\nAs much as extrapolation is necessary in production, it is entirely undesirable in competition scoring. It essentially requires placing negative weight on the scores of early rounds which both creates the possibility of hacks and effectively reduces the sample size of the test set making the scoring much more noisy. It replaces real test data with guessing, but competition ranks should not be guesses. This is an area where I think Kaggle should enforce some standards. To allow noisy and hackable metrics like this in ranked contests degrades the value of everyone's accomplishments on the platform.\n\n“But without the slope term, how can we test stability?” you might ask. Well as I've shown, the slope term only tests stability in the sense that it simulates a longer test set. Simply using a longer test set would do a better job of testing stability because doing so replaces guessing with real data. Slope is also a pretty poor measure of stability because it only takes into account linear performance decline and not random performance drift as a real stability metric should. The hosts might expect the standard deviation term to handle this, but that is likely to be more influenced by week to week performance variability than any tendency towards long-term drift. A longer test set with a real stability metric that is punishing to sustained periods of poor performance would do much more to test stability than the metric we have now, and could do so without the risk of hackability.",
      "votes": null
    },
    {
      "id": "2693134",
      "postDate": "03/12/2024 08:35:02",
      "content": "<p>I fully support your point, <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a>. I had a <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2650795\" target=\"_blank\">similar discussion</a> on extrapolation vs more test data with organizers about a month ago. From what I understand, it was decided to keep the existing metric despite the possible issues.</p>",
      "rawMarkdown": "I fully support your point, @jacobyjaeger. I had a [similar discussion](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2650795) on extrapolation vs more test data with organizers about a month ago. From what I understand, it was decided to keep the existing metric despite the possible issues.",
      "votes": null
    },
    {
      "id": "2693352",
      "postDate": "03/12/2024 11:30:40",
      "content": "<p>My only hope is that, should this competition prove to be a disaster, Kaggle will consider adopting some requirements for the metrics it allows.</p>",
      "rawMarkdown": "My only hope is that, should this competition prove to be a disaster, Kaggle will consider adopting some requirements for the metrics it allows.",
      "votes": null
    },
    {
      "id": "2693509",
      "postDate": "03/12/2024 13:23:54",
      "content": "<p>Hi,</p>\n<p>Calculating naively the dates, I don't think there was room to extend the testing dataset that much.</p>\n<pre><code>df = (\n    pl.scan_parquet(os.path.join(BASE_PATH, ))\n    .group_by(, maintain_order=)\n    .agg(\n        pl.col()\n        .()\n        .alias()\n    )\n    .collect()\n)\n\nsns.lineplot(\n    df,\n    x=,\n    y=\n)\n\nplt.ylabel()\nplt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ffad58bddd9d6f2675a28dadc2ea5638a%2FScreenshot%202024-03-12%20at%202.09.03PM.png?generation=1710249002193888&amp;alt=media\"></p>\n<pre><code>train_ends = (\n    pl.scan_parquet(os.path.join(BASE_PATH, ))\n    .select(\n        pl.col()\n        .last()\n    )\n    .collect()\n    .item()\n)\n\nweeks = (df[].() *  / df[].mean())\ntest_ends = (date.fromisoformat(train_ends) + timedelta(weeks=weeks)).strftime()\ntrain_to_date = ((date.today() - date.fromisoformat(train_ends)).days / )\n\n()\n()\n()\n</code></pre>\n<p>Test start after: 2020-10-05.<br>\nTest ends around: 2022-05-09.</p>\n<p>Weeks since last training date: 179 weeks.</p>\n<p>Maybe they could have extended the testing dataset by another year, so 135 weeks, still far from their goal.</p>",
      "rawMarkdown": "Hi,\n\nCalculating naively the dates, I don't think there was room to extend the testing dataset that much.\n\n```python\ndf = (\n    pl.scan_parquet(os.path.join(BASE_PATH, 'train_base.parquet'))\n    .group_by('WEEK_NUM', maintain_order=True)\n    .agg(\n        pl.col('case_id')\n        .len()\n        .alias('num_cases')\n    )\n    .collect()\n)\n\nsns.lineplot(\n    df,\n    x=\"WEEK_NUM\",\n    y=\"num_cases\"\n)\n\nplt.ylabel(f\"Amount of cases, mean: {df['num_cases'].mean():.2f}\")\nplt.show()\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ffad58bddd9d6f2675a28dadc2ea5638a%2FScreenshot%202024-03-12%20at%202.09.03PM.png?generation=1710249002193888&alt=media)\n\n```python\ntrain_ends = (\n    pl.scan_parquet(os.path.join(BASE_PATH, 'train_base.parquet'))\n    .select(\n        pl.col('date_decision')\n        .last()\n    )\n    .collect()\n    .item()\n)\n\nweeks = round(df['num_cases'].sum() * 0.9 / df['num_cases'].mean())\ntest_ends = (date.fromisoformat(train_ends) + timedelta(weeks=weeks)).strftime('%Y-%m-%d')\ntrain_to_date = round((date.today() - date.fromisoformat(train_ends)).days / 7)\n\nprint(f\"Test start after: {train_ends}.\")\nprint(f\"Test ends around: {test_ends}.\")\nprint(f\"\\nWeeks since last training date: {train_to_date} weeks.\")\n```\n\nTest start after: 2020-10-05.\nTest ends around: 2022-05-09.\n\nWeeks since last training date: 179 weeks.\n\nMaybe they could have extended the testing dataset by another year, so 135 weeks, still far from their goal.",
      "votes": null
    },
    {
      "id": "2694093",
      "postDate": "03/12/2024 20:30:20",
      "content": "<p>Well, it may not be a disaster already, but the thing is that hundreds of submissions have been invalidated and now score differently when resubmitted is definitely not a good sign.</p>",
      "rawMarkdown": "Well, it may not be a disaster already, but the thing is that hundreds of submissions have been invalidated and now score differently when resubmitted is definitely not a good sign.",
      "votes": null
    },
    {
      "id": "2694209",
      "postDate": "03/12/2024 23:18:09",
      "content": "<p>Let's not hold Kaggle to be infallible when it comes to metrics either.  Some of the more standard metrics can be grossly inappropriate as well in some contexts (usually involving logging the actual and predicted values), and I've seen them used countless times.  They may not always ruin a competition, but they would certainly make the winning models useless.</p>\n<p>I think the real lesson that can be taken from this competition is that determining the metric in a data science project is something that should be approached with utmost care and rigor, and it's often a significant part of the challenge.  If your metric doesn't have some first principles connection to what you're trying to accomplish, you're playing with fire.</p>",
      "rawMarkdown": "Let's not hold Kaggle to be infallible when it comes to metrics either.  Some of the more standard metrics can be grossly inappropriate as well in some contexts (usually involving logging the actual and predicted values), and I've seen them used countless times.  They may not always ruin a competition, but they would certainly make the winning models useless.\n\nI think the real lesson that can be taken from this competition is that determining the metric in a data science project is something that should be approached with utmost care and rigor, and it's often a significant part of the challenge.  If your metric doesn't have some first principles connection to what you're trying to accomplish, you're playing with fire.",
      "votes": null
    },
    {
      "id": "2694993",
      "postDate": "03/13/2024 11:46:17",
      "content": "<p>Probably true, but I don't know that they don't have longer data that they cropped for relevance or dataset size considerations. </p>\n<p>I would point out though that if they used a real stability metric that heavily punished periods of poor performance, models with declining performance would suffer much sooner and less additional data would be required to punish them.</p>\n<p>From the examples they showed, it looks like they're most concerned with the projected performance shortly after the end of the test set. I'm sure this is important when considering whether to put a model into production, but it is a bizarre thing to emphasize in a competition. </p>\n<p>I doubt they have a certain test duration requirement in mind that they're trying to simulate. Rather I think they're just naively copying their production criteria that says whatever test period they have they have to project a bit beyond it. </p>\n<p>For the effective center of the extrapolated test set to be beyond the end of real test set, they have to extrapolate to a bit more than twice the real test set's length. I think that's why they're extrapolating so far here. They're not thinking about the bigger picture, or what testing stability really means, they're just copying what they do in production.</p>",
      "rawMarkdown": "Probably true, but I don't know that they don't have longer data that they cropped for relevance or dataset size considerations. \n\nI would point out though that if they used a real stability metric that heavily punished periods of poor performance, models with declining performance would suffer much sooner and less additional data would be required to punish them.\n\nFrom the examples they showed, it looks like they're most concerned with the projected performance shortly after the end of the test set. I'm sure this is important when considering whether to put a model into production, but it is a bizarre thing to emphasize in a competition. \n\nI doubt they have a certain test duration requirement in mind that they're trying to simulate. Rather I think they're just naively copying their production criteria that says whatever test period they have they have to project a bit beyond it. \n\nFor the effective center of the extrapolated test set to be beyond the end of real test set, they have to extrapolate to a bit more than twice the real test set's length. I think that's why they're extrapolating so far here. They're not thinking about the bigger picture, or what testing stability really means, they're just copying what they do in production.",
      "votes": null
    },
    {
      "id": "2694999",
      "postDate": "03/13/2024 11:53:03",
      "content": "<p>Certainly Kaggle also often makes mistakes. I'm not proposing that we rely on their scrutiny more, but that they adopt some objective disqualifying criteria for transparently bad metrics. Namely, I'd like to see them require that overall scores are monotonic wrt sample scores. If all metrics followed this rule, they would not be hackable in the way that this one is. </p>",
      "rawMarkdown": "Certainly Kaggle also often makes mistakes. I'm not proposing that we rely on their scrutiny more, but that they adopt some objective disqualifying criteria for transparently bad metrics. Namely, I'd like to see them require that overall scores are monotonic wrt sample scores. If all metrics followed this rule, they would not be hackable in the way that this one is.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2693134,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "03/12/2024 08:35:02",
      "content": "<p>I fully support your point, <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a>. I had a <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2650795\" target=\"_blank\">similar discussion</a> on extrapolation vs more test data with organizers about a month ago. From what I understand, it was decided to keep the existing metric despite the possible issues.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2693352,
          "author_name": "jacobyjaeger",
          "author_url": "",
          "post_date": "03/12/2024 11:30:40",
          "content": "<p>My only hope is that, should this competition prove to be a disaster, Kaggle will consider adopting some requirements for the metrics it allows.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2694093,
              "author_name": "kononenko",
              "author_url": "",
              "post_date": "03/12/2024 20:30:20",
              "content": "<p>Well, it may not be a disaster already, but the thing is that hundreds of submissions have been invalidated and now score differently when resubmitted is definitely not a good sign.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2694209,
              "author_name": "dmitriyguller",
              "author_url": "",
              "post_date": "03/12/2024 23:18:09",
              "content": "<p>Let's not hold Kaggle to be infallible when it comes to metrics either.  Some of the more standard metrics can be grossly inappropriate as well in some contexts (usually involving logging the actual and predicted values), and I've seen them used countless times.  They may not always ruin a competition, but they would certainly make the winning models useless.</p>\n<p>I think the real lesson that can be taken from this competition is that determining the metric in a data science project is something that should be approached with utmost care and rigor, and it's often a significant part of the challenge.  If your metric doesn't have some first principles connection to what you're trying to accomplish, you're playing with fire.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2694999,
                  "author_name": "jacobyjaeger",
                  "author_url": "",
                  "post_date": "03/13/2024 11:53:03",
                  "content": "<p>Certainly Kaggle also often makes mistakes. I'm not proposing that we rely on their scrutiny more, but that they adopt some objective disqualifying criteria for transparently bad metrics. Namely, I'd like to see them require that overall scores are monotonic wrt sample scores. If all metrics followed this rule, they would not be hackable in the way that this one is. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2693509,
      "author_name": "enriquezaf",
      "author_url": "",
      "post_date": "03/12/2024 13:23:54",
      "content": "<p>Hi,</p>\n<p>Calculating naively the dates, I don't think there was room to extend the testing dataset that much.</p>\n<pre><code>df = (\n    pl.scan_parquet(os.path.join(BASE_PATH, ))\n    .group_by(, maintain_order=)\n    .agg(\n        pl.col()\n        .()\n        .alias()\n    )\n    .collect()\n)\n\nsns.lineplot(\n    df,\n    x=,\n    y=\n)\n\nplt.ylabel()\nplt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ffad58bddd9d6f2675a28dadc2ea5638a%2FScreenshot%202024-03-12%20at%202.09.03PM.png?generation=1710249002193888&amp;alt=media\"></p>\n<pre><code>train_ends = (\n    pl.scan_parquet(os.path.join(BASE_PATH, ))\n    .select(\n        pl.col()\n        .last()\n    )\n    .collect()\n    .item()\n)\n\nweeks = (df[].() *  / df[].mean())\ntest_ends = (date.fromisoformat(train_ends) + timedelta(weeks=weeks)).strftime()\ntrain_to_date = ((date.today() - date.fromisoformat(train_ends)).days / )\n\n()\n()\n()\n</code></pre>\n<p>Test start after: 2020-10-05.<br>\nTest ends around: 2022-05-09.</p>\n<p>Weeks since last training date: 179 weeks.</p>\n<p>Maybe they could have extended the testing dataset by another year, so 135 weeks, still far from their goal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2694993,
          "author_name": "jacobyjaeger",
          "author_url": "",
          "post_date": "03/13/2024 11:46:17",
          "content": "<p>Probably true, but I don't know that they don't have longer data that they cropped for relevance or dataset size considerations. </p>\n<p>I would point out though that if they used a real stability metric that heavily punished periods of poor performance, models with declining performance would suffer much sooner and less additional data would be required to punish them.</p>\n<p>From the examples they showed, it looks like they're most concerned with the projected performance shortly after the end of the test set. I'm sure this is important when considering whether to put a model into production, but it is a bizarre thing to emphasize in a competition. </p>\n<p>I doubt they have a certain test duration requirement in mind that they're trying to simulate. Rather I think they're just naively copying their production criteria that says whatever test period they have they have to project a bit beyond it. </p>\n<p>For the effective center of the extrapolated test set to be beyond the end of real test set, they have to extrapolate to a bit more than twice the real test set's length. I think that's why they're extrapolating so far here. They're not thinking about the bigger picture, or what testing stability really means, they're just copying what they do in production.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2691066": "The hosts say the slope and standard deviation terms make a “stability metric,” but what do they actually accomplish? Well, the effect of the slope term is simple enough:\n\nLet's recall that the mean elevation of a line segment with endpoint elevations of `y1` and `y2` is `(y1+y2)/2`, and that the elevation of a point displaced by distance `dx` along a line from some initial point is `m dx + y0` where `m` is the slope and `y0` the initial elevation. Then when we have a metric of the form `mean(y) + 88*slope(y)`, we effectively have `(y1+y2)/2 + 88*m` which is equivalent to `(y1 + (y2 + 176*m))/2`. So we can see in the equivalent equation that the effect of the slope is to simulate extending the test period by 176 weeks or about 3 years by assuming a linear trend. \n\nThis of course begs the question, if we want the effect of a longer test period, why not just use a longer test period? After all, the test period we have already includes past data, and extrapolation is hackable and prone to finding false trends particularly when we have the risk of features being dropped from the test set. The host's answer is that they have to use this metric because “it approximates well their business decisions.“ But is this really a good reason? \n\nCargo cultism is the naive copying of practices that are successful in one context to other contexts where they may not fit without regard for the purpose of the original practice. It is the fallacy that what works in one situation must work everywhere all the time. And it is all too common in large companies. \n\nExtrapolating future performance to decide whether to put a model into production is necessary because it is that model's future performance that will determine profitability. Does that need for extrapolation also apply to this contest? I think not. It would be pretty silly to immediately put the winning models from this competition into production considering that they've been trained on transformed data and, with the new changes, they don't even have the means to construct potentially useful time-based aggregation features. More likely the models will need to be retrained with the full feature set and then extrapolation can be performed.\n\nAs much as extrapolation is necessary in production, it is entirely undesirable in competition scoring. It essentially requires placing negative weight on the scores of early rounds which both creates the possibility of hacks and effectively reduces the sample size of the test set making the scoring much more noisy. It replaces real test data with guessing, but competition ranks should not be guesses. This is an area where I think Kaggle should enforce some standards. To allow noisy and hackable metrics like this in ranked contests degrades the value of everyone's accomplishments on the platform.\n\n“But without the slope term, how can we test stability?” you might ask. Well as I've shown, the slope term only tests stability in the sense that it simulates a longer test set. Simply using a longer test set would do a better job of testing stability because doing so replaces guessing with real data. Slope is also a pretty poor measure of stability because it only takes into account linear performance decline and not random performance drift as a real stability metric should. The hosts might expect the standard deviation term to handle this, but that is likely to be more influenced by week to week performance variability than any tendency towards long-term drift. A longer test set with a real stability metric that is punishing to sustained periods of poor performance would do much more to test stability than the metric we have now, and could do so without the risk of hackability.",
    "2693134": "I fully support your point, @jacobyjaeger. I had a [similar discussion](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449#2650795) on extrapolation vs more test data with organizers about a month ago. From what I understand, it was decided to keep the existing metric despite the possible issues.",
    "2693352": "My only hope is that, should this competition prove to be a disaster, Kaggle will consider adopting some requirements for the metrics it allows.",
    "2693509": "Hi,\n\nCalculating naively the dates, I don't think there was room to extend the testing dataset that much.\n\n```python\ndf = (\n    pl.scan_parquet(os.path.join(BASE_PATH, 'train_base.parquet'))\n    .group_by('WEEK_NUM', maintain_order=True)\n    .agg(\n        pl.col('case_id')\n        .len()\n        .alias('num_cases')\n    )\n    .collect()\n)\n\nsns.lineplot(\n    df,\n    x=\"WEEK_NUM\",\n    y=\"num_cases\"\n)\n\nplt.ylabel(f\"Amount of cases, mean: {df['num_cases'].mean():.2f}\")\nplt.show()\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ffad58bddd9d6f2675a28dadc2ea5638a%2FScreenshot%202024-03-12%20at%202.09.03PM.png?generation=1710249002193888&alt=media)\n\n```python\ntrain_ends = (\n    pl.scan_parquet(os.path.join(BASE_PATH, 'train_base.parquet'))\n    .select(\n        pl.col('date_decision')\n        .last()\n    )\n    .collect()\n    .item()\n)\n\nweeks = round(df['num_cases'].sum() * 0.9 / df['num_cases'].mean())\ntest_ends = (date.fromisoformat(train_ends) + timedelta(weeks=weeks)).strftime('%Y-%m-%d')\ntrain_to_date = round((date.today() - date.fromisoformat(train_ends)).days / 7)\n\nprint(f\"Test start after: {train_ends}.\")\nprint(f\"Test ends around: {test_ends}.\")\nprint(f\"\\nWeeks since last training date: {train_to_date} weeks.\")\n```\n\nTest start after: 2020-10-05.\nTest ends around: 2022-05-09.\n\nWeeks since last training date: 179 weeks.\n\nMaybe they could have extended the testing dataset by another year, so 135 weeks, still far from their goal.",
    "2694093": "Well, it may not be a disaster already, but the thing is that hundreds of submissions have been invalidated and now score differently when resubmitted is definitely not a good sign.",
    "2694209": "Let's not hold Kaggle to be infallible when it comes to metrics either.  Some of the more standard metrics can be grossly inappropriate as well in some contexts (usually involving logging the actual and predicted values), and I've seen them used countless times.  They may not always ruin a competition, but they would certainly make the winning models useless.\n\nI think the real lesson that can be taken from this competition is that determining the metric in a data science project is something that should be approached with utmost care and rigor, and it's often a significant part of the challenge.  If your metric doesn't have some first principles connection to what you're trying to accomplish, you're playing with fire.",
    "2694993": "Probably true, but I don't know that they don't have longer data that they cropped for relevance or dataset size considerations. \n\nI would point out though that if they used a real stability metric that heavily punished periods of poor performance, models with declining performance would suffer much sooner and less additional data would be required to punish them.\n\nFrom the examples they showed, it looks like they're most concerned with the projected performance shortly after the end of the test set. I'm sure this is important when considering whether to put a model into production, but it is a bizarre thing to emphasize in a competition. \n\nI doubt they have a certain test duration requirement in mind that they're trying to simulate. Rather I think they're just naively copying their production criteria that says whatever test period they have they have to project a bit beyond it. \n\nFor the effective center of the extrapolated test set to be beyond the end of real test set, they have to extrapolate to a bit more than twice the real test set's length. I think that's why they're extrapolating so far here. They're not thinking about the bigger picture, or what testing stability really means, they're just copying what they do in production.",
    "2694999": "Certainly Kaggle also often makes mistakes. I'm not proposing that we rely on their scrutiny more, but that they adopt some objective disqualifying criteria for transparently bad metrics. Namely, I'd like to see them require that overall scores are monotonic wrt sample scores. If all metrics followed this rule, they would not be hackable in the way that this one is."
  },
  "source": "meta"
}