{
  "id": 501800,
  "title": "Stability metric issue: probabilistic approach",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501800",
  "author_name": "",
  "post_date": "2024-05-10T19:45:15.304259200Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'd like to share some additional thoughts on why the current stability metric is not a reliable way to estimate the stability of a model. I know it's too late to change anything, but I believe it may still be of interest to understand the underlying issues. As kaggle discussions do not currently support inline formulas, the detailed explanation can be found in the <a href=\"https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook\" target=\"_blank\"><strong>notebook</strong></a>. And here is a short summary:</p>\n<ul>\n<li><strong>The penalty term in the metric is not a reliable way to estimate how model performance gets worse over time.</strong> It assumes a predictable, deterministic trend, which can be extrapolated into the future, but the trend in ginis can be stochastic = unpredictable and still greatly penalize a model.</li>\n<li><strong>Negative slope can appear without a deterministic trend.</strong> A simple example is a Brownian motion - it has a stochastic trend and can produce a negative slope due to random shocks, but we can't expect the slope to persist in the future. In this case the penalty is not justified. </li>\n<li><strong>It's possible to distinguish between a stochastic and deterministic trend, but the current metric fails to do so.</strong> The host wants to avoid a deterioration in performance outside the evaluation sample, which they believe will occur if they observe a downward slope during the evaluation sample, but that is not true in case of a stochastic trend, which is still penalized by the current metric.</li>\n<li><strong>A better approach is to check if the average rate of change in ginis is negative.</strong> Penalties should only be applied if this rate is statistically significant, which would ensure we account for real trends and not random changes.</li>\n</ul>",
  "messages": [
    {
      "id": "2805904",
      "postDate": "05/10/2024 19:45:15",
      "content": "<p>I'd like to share some additional thoughts on why the current stability metric is not a reliable way to estimate the stability of a model. I know it's too late to change anything, but I believe it may still be of interest to understand the underlying issues. As kaggle discussions do not currently support inline formulas, the detailed explanation can be found in the <a href=\"https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook\" target=\"_blank\"><strong>notebook</strong></a>. And here is a short summary:</p>\n<ul>\n<li><strong>The penalty term in the metric is not a reliable way to estimate how model performance gets worse over time.</strong> It assumes a predictable, deterministic trend, which can be extrapolated into the future, but the trend in ginis can be stochastic = unpredictable and still greatly penalize a model.</li>\n<li><strong>Negative slope can appear without a deterministic trend.</strong> A simple example is a Brownian motion - it has a stochastic trend and can produce a negative slope due to random shocks, but we can't expect the slope to persist in the future. In this case the penalty is not justified. </li>\n<li><strong>It's possible to distinguish between a stochastic and deterministic trend, but the current metric fails to do so.</strong> The host wants to avoid a deterioration in performance outside the evaluation sample, which they believe will occur if they observe a downward slope during the evaluation sample, but that is not true in case of a stochastic trend, which is still penalized by the current metric.</li>\n<li><strong>A better approach is to check if the average rate of change in ginis is negative.</strong> Penalties should only be applied if this rate is statistically significant, which would ensure we account for real trends and not random changes.</li>\n</ul>",
      "rawMarkdown": "I'd like to share some additional thoughts on why the current stability metric is not a reliable way to estimate the stability of a model. I know it's too late to change anything, but I believe it may still be of interest to understand the underlying issues. As kaggle discussions do not currently support inline formulas, the detailed explanation can be found in the [**notebook**](https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook). And here is a short summary:\n\n- **The penalty term in the metric is not a reliable way to estimate how model performance gets worse over time.** It assumes a predictable, deterministic trend, which can be extrapolated into the future, but the trend in ginis can be stochastic = unpredictable and still greatly penalize a model.\n- **Negative slope can appear without a deterministic trend.** A simple example is a Brownian motion - it has a stochastic trend and can produce a negative slope due to random shocks, but we can't expect the slope to persist in the future. In this case the penalty is not justified. \n- **It's possible to distinguish between a stochastic and deterministic trend, but the current metric fails to do so.** The host wants to avoid a deterioration in performance outside the evaluation sample, which they believe will occur if they observe a downward slope during the evaluation sample, but that is not true in case of a stochastic trend, which is still penalized by the current metric.\n- **A better approach is to check if the average rate of change in ginis is negative.** Penalties should only be applied if this rate is statistically significant, which would ensure we account for real trends and not random changes.",
      "votes": null
    },
    {
      "id": "2805981",
      "postDate": "05/10/2024 20:08:38",
      "content": "<p>100% agree. There are in my opinion multiple much better ways to assess stability (like the approach you mention above for instance). And I can think of at least three other ways as well!  </p>",
      "rawMarkdown": "100% agree. There are in my opinion multiple much better ways to assess stability (like the approach you mention above for instance). And I can think of at least three other ways as well!",
      "votes": null
    },
    {
      "id": "2806475",
      "postDate": "05/11/2024 05:54:16",
      "content": "<p>100% agree.</p>",
      "rawMarkdown": "100% agree.",
      "votes": null
    },
    {
      "id": "2807396",
      "postDate": "05/11/2024 17:03:48",
      "content": "<p>I absolutely agree with you, especially the third point. What is the point of penalty, if a consistent negative slope is not observed? From what I see, irrespective of positive or negative slope, any amount of std, still penalizes the stability (though smaller penalties for positive slopes).</p>",
      "rawMarkdown": "I absolutely agree with you, especially the third point. What is the point of penalty, if a consistent negative slope is not observed? From what I see, irrespective of positive or negative slope, any amount of std, still penalizes the stability (though smaller penalties for positive slopes).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2805981,
      "author_name": "ern711",
      "author_url": "",
      "post_date": "05/10/2024 20:08:38",
      "content": "<p>100% agree. There are in my opinion multiple much better ways to assess stability (like the approach you mention above for instance). And I can think of at least three other ways as well!  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2806475,
          "author_name": "alexxanderlarko",
          "author_url": "",
          "post_date": "05/11/2024 05:54:16",
          "content": "<p>100% agree.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2807396,
      "author_name": "varuniraothumsi",
      "author_url": "",
      "post_date": "05/11/2024 17:03:48",
      "content": "<p>I absolutely agree with you, especially the third point. What is the point of penalty, if a consistent negative slope is not observed? From what I see, irrespective of positive or negative slope, any amount of std, still penalizes the stability (though smaller penalties for positive slopes).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2805904": "I'd like to share some additional thoughts on why the current stability metric is not a reliable way to estimate the stability of a model. I know it's too late to change anything, but I believe it may still be of interest to understand the underlying issues. As kaggle discussions do not currently support inline formulas, the detailed explanation can be found in the [**notebook**](https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach/notebook). And here is a short summary:\n\n- **The penalty term in the metric is not a reliable way to estimate how model performance gets worse over time.** It assumes a predictable, deterministic trend, which can be extrapolated into the future, but the trend in ginis can be stochastic = unpredictable and still greatly penalize a model.\n- **Negative slope can appear without a deterministic trend.** A simple example is a Brownian motion - it has a stochastic trend and can produce a negative slope due to random shocks, but we can't expect the slope to persist in the future. In this case the penalty is not justified. \n- **It's possible to distinguish between a stochastic and deterministic trend, but the current metric fails to do so.** The host wants to avoid a deterioration in performance outside the evaluation sample, which they believe will occur if they observe a downward slope during the evaluation sample, but that is not true in case of a stochastic trend, which is still penalized by the current metric.\n- **A better approach is to check if the average rate of change in ginis is negative.** Penalties should only be applied if this rate is statistically significant, which would ensure we account for real trends and not random changes.",
    "2805981": "100% agree. There are in my opinion multiple much better ways to assess stability (like the approach you mention above for instance). And I can think of at least three other ways as well!",
    "2806475": "100% agree.",
    "2807396": "I absolutely agree with you, especially the third point. What is the point of penalty, if a consistent negative slope is not observed? From what I see, irrespective of positive or negative slope, any amount of std, still penalizes the stability (though smaller penalties for positive slopes)."
  },
  "source": "meta"
}