{
  "id": 501708,
  "title": "A much better metric for stability - For organizers Internal use",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501708",
  "author_name": "Simo Elm",
  "post_date": "2024-05-10T12:13:37.475000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I have joined this competition quite few days ago and it has been fun so far. I haven't done any feature engineering yet but been experimenting with the stability metric.</p>\n<p>The stability issue is not new by any means. If anyone has ever worked in designing automated trading strategies or worked in a quant hf he'd quickly understand the return-volatility issue.</p>\n<p>Usually, HFs wants their strategy to have higher returns while minimizing volatility and drowdowns (deminishing returns over some period of time).</p>\n<p>The metric used for those cases is the Sharpe Ratio. In this case a better (similar) stability metric should be something like:</p>\n<p>**Stability = Mean_week(AUC)/std_week(AUC) ** or some kind of variations to it.</p>\n<p>The problem with the current metric is the regression factor 'a'. Everyone understands that it is unrealistic to have a positive a, but a slightly negative 'a' is highly penalized.  **This is statstically uncorrect. ** </p>\n<p>A correct correction to the metric (in case organizers wants to keep those specific penalizations would be to run a <strong>hypothesis t-test to test 'a=0' against 'a&lt;0' and then use p-value as the trigger for penalization</strong>. With the current setup, a model might have an excellent AUC but a very small negative 'a' (which is desirable) would be ranked worse than a model with less AUC and a slightly positive a.</p>\n<p>Simo</p>",
  "messages": [
    {
      "id": 2805150,
      "postDate": "2024-05-10T12:13:37.477Z",
      "content": "<p>Hello all,</p>\n<p>I have joined this competition quite few days ago and it has been fun so far. I haven't done any feature engineering yet but been experimenting with the stability metric.</p>\n<p>The stability issue is not new by any means. If anyone has ever worked in designing automated trading strategies or worked in a quant hf he'd quickly understand the return-volatility issue.</p>\n<p>Usually, HFs wants their strategy to have higher returns while minimizing volatility and drowdowns (deminishing returns over some period of time).</p>\n<p>The metric used for those cases is the Sharpe Ratio. In this case a better (similar) stability metric should be something like:</p>\n<p>**Stability = Mean_week(AUC)/std_week(AUC) ** or some kind of variations to it.</p>\n<p>The problem with the current metric is the regression factor 'a'. Everyone understands that it is unrealistic to have a positive a, but a slightly negative 'a' is highly penalized.  **This is statstically uncorrect. ** </p>\n<p>A correct correction to the metric (in case organizers wants to keep those specific penalizations would be to run a <strong>hypothesis t-test to test 'a=0' against 'a&lt;0' and then use p-value as the trigger for penalization</strong>. With the current setup, a model might have an excellent AUC but a very small negative 'a' (which is desirable) would be ranked worse than a model with less AUC and a slightly positive a.</p>\n<p>Simo</p>",
      "rawMarkdown": "Hello all,\n\nI have joined this competition quite few days ago and it has been fun so far. I haven't done any feature engineering yet but been experimenting with the stability metric.\n\nThe stability issue is not new by any means. If anyone has ever worked in designing automated trading strategies or worked in a quant hf he'd quickly understand the return-volatility issue.\n\nUsually, HFs wants their strategy to have higher returns while minimizing volatility and drowdowns (deminishing returns over some period of time).\n\nThe metric used for those cases is the Sharpe Ratio. In this case a better (similar) stability metric should be something like:\n\n**Stability = Mean_week(AUC)/std_week(AUC) ** or some kind of variations to it.\n\nThe problem with the current metric is the regression factor 'a'. Everyone understands that it is unrealistic to have a positive a, but a slightly negative 'a' is highly penalized.  **This is statstically uncorrect. ** \n\nA correct correction to the metric (in case organizers wants to keep those specific penalizations would be to run a **hypothesis t-test to test 'a=0' against 'a<0' and then use p-value as the trigger for penalization**. With the current setup, a model might have an excellent AUC but a very small negative 'a' (which is desirable) would be ranked worse than a model with less AUC and a slightly positive a.\n\nSimo\n\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2805150": "Hello all,\n\nI have joined this competition quite few days ago and it has been fun so far. I haven't done any feature engineering yet but been experimenting with the stability metric.\n\nThe stability issue is not new by any means. If anyone has ever worked in designing automated trading strategies or worked in a quant hf he'd quickly understand the return-volatility issue.\n\nUsually, HFs wants their strategy to have higher returns while minimizing volatility and drowdowns (deminishing returns over some period of time).\n\nThe metric used for those cases is the Sharpe Ratio. In this case a better (similar) stability metric should be something like:\n\n**Stability = Mean_week(AUC)/std_week(AUC) ** or some kind of variations to it.\n\nThe problem with the current metric is the regression factor 'a'. Everyone understands that it is unrealistic to have a positive a, but a slightly negative 'a' is highly penalized.  **This is statstically uncorrect. ** \n\nA correct correction to the metric (in case organizers wants to keep those specific penalizations would be to run a **hypothesis t-test to test 'a=0' against 'a<0' and then use p-value as the trigger for penalization**. With the current setup, a model might have an excellent AUC but a very small negative 'a' (which is desirable) would be ranked worse than a model with less AUC and a slightly positive a.\n\nSimo\n\n"
  }
}