{
  "id": 499993,
  "title": "Competition Metric Behavior Case",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/499993",
  "author_name": "",
  "post_date": "2024-05-03T19:40:21.321801500Z",
  "votes": 20,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We have several discussions explaining the competition metric behavior, I decided to do some experiments and was surprised how big the penalties are. Perhaps this example will clarify the metric for someone.</p>\n<p>Assume that we have <code>WEEK_NUM = [1, 2, 3, 4]</code> in the test period (shift <code>WEEK_NUM+100</code> doesn't change metrics).</p>\n<p><strong>CASE 1:</strong><br>\nCorresponding <code>ginis = [0.80, 0.80, 0.80, 0.80]</code>.</p>\n<p><code>stability metric = 0.80</code></p>\n<p><strong>CASE 2:</strong><br>\nCorresponding <code>ginis = [1.00, 0.99, 0.99, 0.99]</code>.</p>\n<p><code>stability metric = 0.73</code></p>\n<h2>Reasoning</h2>\n<p>It shows that <strong>CASE 2</strong> has much worse metric than <strong>CASE 1</strong> (for some people it should be vice versa, and I would prefer almost perfect model from <strong>CASE 2</strong>) because the term <code>88.0 * falling_rate</code> penalizes heavily, and we have a situation where sacrificing performance for first weeks might improve the final score a lot.</p>\n<p>If you want to play with the competition metric or check my calculations, you can use <a href=\"https://www.kaggle.com/code/vitalykudelya/home-credit-metric-playground\" target=\"_blank\">Home Credit: Metric Playground</a>.</p>",
  "messages": [
    {
      "id": "2791678",
      "postDate": "05/03/2024 19:40:21",
      "content": "<p>We have several discussions explaining the competition metric behavior, I decided to do some experiments and was surprised how big the penalties are. Perhaps this example will clarify the metric for someone.</p>\n<p>Assume that we have <code>WEEK_NUM = [1, 2, 3, 4]</code> in the test period (shift <code>WEEK_NUM+100</code> doesn't change metrics).</p>\n<p><strong>CASE 1:</strong><br>\nCorresponding <code>ginis = [0.80, 0.80, 0.80, 0.80]</code>.</p>\n<p><code>stability metric = 0.80</code></p>\n<p><strong>CASE 2:</strong><br>\nCorresponding <code>ginis = [1.00, 0.99, 0.99, 0.99]</code>.</p>\n<p><code>stability metric = 0.73</code></p>\n<h2>Reasoning</h2>\n<p>It shows that <strong>CASE 2</strong> has much worse metric than <strong>CASE 1</strong> (for some people it should be vice versa, and I would prefer almost perfect model from <strong>CASE 2</strong>) because the term <code>88.0 * falling_rate</code> penalizes heavily, and we have a situation where sacrificing performance for first weeks might improve the final score a lot.</p>\n<p>If you want to play with the competition metric or check my calculations, you can use <a href=\"https://www.kaggle.com/code/vitalykudelya/home-credit-metric-playground\" target=\"_blank\">Home Credit: Metric Playground</a>.</p>",
      "rawMarkdown": "We have several discussions explaining the competition metric behavior, I decided to do some experiments and was surprised how big the penalties are. Perhaps this example will clarify the metric for someone.\n\nAssume that we have `WEEK_NUM = [1, 2, 3, 4]` in the test period (shift `WEEK_NUM+100` doesn't change metrics).\n\n**CASE 1:**\nCorresponding `ginis = [0.80, 0.80, 0.80, 0.80]`.\n\n`stability metric = 0.80`\n\n**CASE 2:**\nCorresponding `ginis = [1.00, 0.99, 0.99, 0.99]`.\n\n`stability metric = 0.73`\n\n## Reasoning\nIt shows that **CASE 2** has much worse metric than **CASE 1** (for some people it should be vice versa, and I would prefer almost perfect model from **CASE 2**) because the term `88.0 * falling_rate` penalizes heavily, and we have a situation where sacrificing performance for first weeks might improve the final score a lot.\n\nIf you want to play with the competition metric or check my calculations, you can use [Home Credit: Metric Playground](https://www.kaggle.com/code/vitalykudelya/home-credit-metric-playground).",
      "votes": null
    },
    {
      "id": "2791715",
      "postDate": "05/03/2024 20:09:01",
      "content": "<p>Yes, this <code>88</code> coefficient is probably calculated from some behind-the-scene data (overfitting?) that may have very different distribution from the future ones. When I see constants like this in metrics, I feel like it won’t help models to generalize well…</p>",
      "rawMarkdown": "Yes, this `88` coefficient is probably calculated from some behind-the-scene data (overfitting?) that may have very different distribution from the future ones. When I see constants like this in metrics, I feel like it won’t help models to generalize well…",
      "votes": null
    },
    {
      "id": "2791883",
      "postDate": "05/03/2024 22:20:01",
      "content": "<p>It's better to test on a larger number of observations, corresponding to the expected number of observations in the test set. In your case, the score is decreasing at a rate of -0.003 per week, which extrapolates to -0.33 over 100 weeks.</p>\n<p>If we simulate 100 weeks instead of 4, with the same decrease of 0.01, the penalty is much smaller. The decrease in the score over 100 weeks should be much more pronounced to incur a significant penalty. </p>\n<pre><code>ginis1 = np.ones()*\nginis2 = np.ones()\nginis2[:] = \n\n ginis  [\n    ginis1, ginis2\n]:\n    metric = calculate_stability_metric(ginis)\n\nScore: , mean , slope , std \nScore: , mean , slope -, std \n</code></pre>",
      "rawMarkdown": "It's better to test on a larger number of observations, corresponding to the expected number of observations in the test set. In your case, the score is decreasing at a rate of -0.003 per week, which extrapolates to -0.33 over 100 weeks.\n\nIf we simulate 100 weeks instead of 4, with the same decrease of 0.01, the penalty is much smaller. The decrease in the score over 100 weeks should be much more pronounced to incur a significant penalty. \n\n```python\nginis1 = np.ones(100)*0.88\nginis2 = np.ones(100)\nginis2[25:] = 0.99\n\nfor ginis in [\n    ginis1, ginis2\n]:\n    metric = calculate_stability_metric(ginis)\n\nScore: 0.8800, mean 0.8800, slope 0.0000, std 0.0000\nScore: 0.9812, mean 0.9925, slope -0.0099, std 0.0014\n```",
      "votes": null
    },
    {
      "id": "2791970",
      "postDate": "05/04/2024 01:22:50",
      "content": "<p>I understand the goal of using the stability metric by the host. </p>\n<p>However, I think it would be better to calculate the AUC for each week separately, then calculate the weighted average (assigning appropriate weights to each week).</p>",
      "rawMarkdown": "I understand the goal of using the stability metric by the host. \n\nHowever, I think it would be better to calculate the AUC for each week separately, then calculate the weighted average (assigning appropriate weights to each week).",
      "votes": null
    },
    {
      "id": "2796965",
      "postDate": "05/06/2024 13:39:14",
      "content": "<p>I completely agree… I've used constants in the past in prior competitions to help increase the LB, only to have the models perform poorly when extrapolating over the entire (hidden) test set. You must get really lucky for such constants to work in future samples. </p>",
      "rawMarkdown": "I completely agree... I've used constants in the past in prior competitions to help increase the LB, only to have the models perform poorly when extrapolating over the entire (hidden) test set. You must get really lucky for such constants to work in future samples.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2791715,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "05/03/2024 20:09:01",
      "content": "<p>Yes, this <code>88</code> coefficient is probably calculated from some behind-the-scene data (overfitting?) that may have very different distribution from the future ones. When I see constants like this in metrics, I feel like it won’t help models to generalize well…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2796965,
          "author_name": "eduardastefanescu",
          "author_url": "",
          "post_date": "05/06/2024 13:39:14",
          "content": "<p>I completely agree… I've used constants in the past in prior competitions to help increase the LB, only to have the models perform poorly when extrapolating over the entire (hidden) test set. You must get really lucky for such constants to work in future samples. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2791883,
      "author_name": "eivolkova",
      "author_url": "",
      "post_date": "05/03/2024 22:20:01",
      "content": "<p>It's better to test on a larger number of observations, corresponding to the expected number of observations in the test set. In your case, the score is decreasing at a rate of -0.003 per week, which extrapolates to -0.33 over 100 weeks.</p>\n<p>If we simulate 100 weeks instead of 4, with the same decrease of 0.01, the penalty is much smaller. The decrease in the score over 100 weeks should be much more pronounced to incur a significant penalty. </p>\n<pre><code>ginis1 = np.ones()*\nginis2 = np.ones()\nginis2[:] = \n\n ginis  [\n    ginis1, ginis2\n]:\n    metric = calculate_stability_metric(ginis)\n\nScore: , mean , slope , std \nScore: , mean , slope -, std \n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2791970,
      "author_name": "bruceqdu",
      "author_url": "",
      "post_date": "05/04/2024 01:22:50",
      "content": "<p>I understand the goal of using the stability metric by the host. </p>\n<p>However, I think it would be better to calculate the AUC for each week separately, then calculate the weighted average (assigning appropriate weights to each week).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2791678": "We have several discussions explaining the competition metric behavior, I decided to do some experiments and was surprised how big the penalties are. Perhaps this example will clarify the metric for someone.\n\nAssume that we have `WEEK_NUM = [1, 2, 3, 4]` in the test period (shift `WEEK_NUM+100` doesn't change metrics).\n\n**CASE 1:**\nCorresponding `ginis = [0.80, 0.80, 0.80, 0.80]`.\n\n`stability metric = 0.80`\n\n**CASE 2:**\nCorresponding `ginis = [1.00, 0.99, 0.99, 0.99]`.\n\n`stability metric = 0.73`\n\n## Reasoning\nIt shows that **CASE 2** has much worse metric than **CASE 1** (for some people it should be vice versa, and I would prefer almost perfect model from **CASE 2**) because the term `88.0 * falling_rate` penalizes heavily, and we have a situation where sacrificing performance for first weeks might improve the final score a lot.\n\nIf you want to play with the competition metric or check my calculations, you can use [Home Credit: Metric Playground](https://www.kaggle.com/code/vitalykudelya/home-credit-metric-playground).",
    "2791715": "Yes, this `88` coefficient is probably calculated from some behind-the-scene data (overfitting?) that may have very different distribution from the future ones. When I see constants like this in metrics, I feel like it won’t help models to generalize well…",
    "2791883": "It's better to test on a larger number of observations, corresponding to the expected number of observations in the test set. In your case, the score is decreasing at a rate of -0.003 per week, which extrapolates to -0.33 over 100 weeks.\n\nIf we simulate 100 weeks instead of 4, with the same decrease of 0.01, the penalty is much smaller. The decrease in the score over 100 weeks should be much more pronounced to incur a significant penalty. \n\n```python\nginis1 = np.ones(100)*0.88\nginis2 = np.ones(100)\nginis2[25:] = 0.99\n\nfor ginis in [\n    ginis1, ginis2\n]:\n    metric = calculate_stability_metric(ginis)\n\nScore: 0.8800, mean 0.8800, slope 0.0000, std 0.0000\nScore: 0.9812, mean 0.9925, slope -0.0099, std 0.0014\n```",
    "2791970": "I understand the goal of using the stability metric by the host. \n\nHowever, I think it would be better to calculate the AUC for each week separately, then calculate the weighted average (assigning appropriate weights to each week).",
    "2796965": "I completely agree... I've used constants in the past in prior competitions to help increase the LB, only to have the models perform poorly when extrapolating over the entire (hidden) test set. You must get really lucky for such constants to work in future samples."
  },
  "source": "meta"
}