{
  "id": 478699,
  "title": "Proposed metrics",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/478699",
  "author_name": "Daniel Herman",
  "post_date": "2024-02-21T20:35:12.215000",
  "votes": 39,
  "comment_count": 63,
  "views": 0,
  "content": "<p>Here you can ask me to score any metric you would like to suggest as new stability metric. The score itself does not guarantee that we will use the metric in this competition. The score of a metric is just an approximation of how we rate the metrics, perhaps during the following dates I will improve this method. Think of the score as a percentage of how much it is aligned with an internal business view. </p>\n<h2>Rules</h2>\n<ol>\n<li>If there are any hyperparameters I will optimize them with given boundaries ([-inf, inf] is not a valid interval). </li>\n<li>I won't score a metric with more than three hyperparameters that are not set beforehand. </li>\n<li>The metric should be relatively easy to explain and should have a connection to the current stability metric, in other words, we would like to see a formula that is easy to read (interpretability) and it is not connected directly to the competition data (generality). </li>\n<li>Please send me a simple notebook or easy-to-read code, ideally with sensible variable naming. I won't be implementing metric if you send me just an equation. I will implement it from an equation only if I have spare time, so no guarantee. </li>\n</ol>\n<p>Score range is 0-1, where higher is better. </p>\n<h3>Current stability metric</h3>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n     avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n</code></pre>\n<p>Score: 0.99</p>\n<h3>Metric mean + rolling min/max diff penalty by <a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a></h3>\n<pre><code> ():\n    penalty = \n    reward = \n\n     i  (rolling_per1 - , (gini_scores)):\n        penalty = (\n            penalty\n            + ((gini_scores[(i - rolling_per1 + ) : i + ]) - gini_scores[i]) ** \n        )\n\n     i  (rolling_per2 - , (gini_scores)):\n        reward = (\n            reward\n            + (gini_scores[i] - (gini_scores[(i - rolling_per2 + ) : i + ])) ** \n        )\n\n     (\n        np.mean(gini_scores)\n        - c1 * (np.sqrt(penalty) / ((gini_scores) - rolling_per1 + ))\n        + c2 * (np.sqrt(reward) / ((gini_scores) - rolling_per2 + ))\n    )\n</code></pre>\n<p>Score 0.86</p>\n<h3>Metric mean + rolling max diff penalty by <a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a></h3>\n<pre><code> ():\n    penalty = \n\n     i  (rolling_per-,(gini_scores)):\n        penalty = penalty + ((gini_scores[(i-rolling_per+):i+])-gini_scores[i])**\n\n     np.mean(gini_scores) - c*(np.sqrt(penalty)/((gini_scores)-rolling_per+)) \n</code></pre>\n<p>Score 0.85</p>\n<h3>Metric mean log MA by <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a></h3>\n<pre><code> ():\n    x = np.cumsum(gini_in_time, axis=)\n    x = np.concatenate([*x[:], x], )\n    scores = -np.mean(-np.log(np.maximum((x[ma_len:] - x[:-ma_len])/ma_len, ))**exponent)\n     scores \n</code></pre>\n<p>Score: 0.83</p>\n<h3>Metric with weights by <a href=\"https://www.kaggle.com/viktorian\" target=\"_blank\">@viktorian</a></h3>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/(x)\n     w_avg_gini + w_gstd * gini_std\n</code></pre>\n<h3>Metric mean std by <a href=\"https://www.kaggle.com/seifachour12\" target=\"_blank\">@seifachour12</a></h3>\n<pre><code> ():\n     (np.(beta*(/beta*np.mean(gini_in_time) + (gamma-np.std(gini_in_time)))-))**(/alpha) \n</code></pre>\n<p>Score: 0.76</p>\n<h3>Metric mean std by <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a></h3>\n<pre><code> ():\n     np.mean(gini_in_time) - c*np.std(gini_in_time)\n</code></pre>\n<p>Score: 0.74</p>\n<h3>Metric rolling gini max by <a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a></h3>\n<pre><code> ():\n    max_gini = gini_in_time[]\n    cost = \n     i  (, (gini_in_time)):\n        max_gini = (max_gini, gini_in_time[i])\n        cost += (, max_gini - gini_in_time[i])**\n     np.mean(gini_in_time) + c/((gini_in_time)-)*cost   \n</code></pre>\n<p>Score: 0.67</p>\n<h2>Metric stability with averaging by <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a></h2>\n<pre><code> ():\n    w_fallingrate /= f + \n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far = a\n     s  (,f):\n        start_index = (x) // f * (s)\n        end_index = (x) // f * (s+)\n        x_second_fifth = x[start_index:end_index]\n        y_second_fifth = gini_in_time[start_index:end_index]\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        far += (,a1)\n\n     avg_gini + w_fallingrate * (far) + w_resstd * res_std \n</code></pre>\n<h2>Metric mean gini</h2>\n<pre><code> ():\n     np.mean(gini_in_time)  \n</code></pre>\n<p>Score: 0.59</p>\n<h3>Metric sharpe-like by <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a></h3>\n<pre><code> ():\n     np.mean(gini_in_time) / np.std(gini_in_time)\n</code></pre>\n<p>Score 0.52</p>\n<h2>Stability with sliding window by <a href=\"https://www.kaggle.com/naotokai\" target=\"_blank\">@naotokai</a></h2>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    \n    x_past = np.arange(-win, )\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n\n    \n    a1_lst = []\n     start_index  (, (x)-win+):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, )\n        a1_lst.append((,a1))       \n    a1_avg = np.mean(a1_lst)\n\n     avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n</code></pre>\n<p>Score: 0.49</p>\n<p>I think the target is at least 0.90 for us to consider this metric as a substitution for the current stability metric. Any effort will be appreciated, but again this is outside the scope of the competition LB and definitely not mandatory, completely up to you. </p>",
  "messages": [
    {
      "id": 2662336,
      "postDate": "2024-02-21T20:35:12.217Z",
      "content": "<p>Here you can ask me to score any metric you would like to suggest as new stability metric. The score itself does not guarantee that we will use the metric in this competition. The score of a metric is just an approximation of how we rate the metrics, perhaps during the following dates I will improve this method. Think of the score as a percentage of how much it is aligned with an internal business view. </p>\n<h2>Rules</h2>\n<ol>\n<li>If there are any hyperparameters I will optimize them with given boundaries ([-inf, inf] is not a valid interval). </li>\n<li>I won't score a metric with more than three hyperparameters that are not set beforehand. </li>\n<li>The metric should be relatively easy to explain and should have a connection to the current stability metric, in other words, we would like to see a formula that is easy to read (interpretability) and it is not connected directly to the competition data (generality). </li>\n<li>Please send me a simple notebook or easy-to-read code, ideally with sensible variable naming. I won't be implementing metric if you send me just an equation. I will implement it from an equation only if I have spare time, so no guarantee. </li>\n</ol>\n<p>Score range is 0-1, where higher is better. </p>\n<h3>Current stability metric</h3>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n     avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n</code></pre>\n<p>Score: 0.99</p>\n<h3>Metric mean + rolling min/max diff penalty by <a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a></h3>\n<pre><code> ():\n    penalty = \n    reward = \n\n     i  (rolling_per1 - , (gini_scores)):\n        penalty = (\n            penalty\n            + ((gini_scores[(i - rolling_per1 + ) : i + ]) - gini_scores[i]) ** \n        )\n\n     i  (rolling_per2 - , (gini_scores)):\n        reward = (\n            reward\n            + (gini_scores[i] - (gini_scores[(i - rolling_per2 + ) : i + ])) ** \n        )\n\n     (\n        np.mean(gini_scores)\n        - c1 * (np.sqrt(penalty) / ((gini_scores) - rolling_per1 + ))\n        + c2 * (np.sqrt(reward) / ((gini_scores) - rolling_per2 + ))\n    )\n</code></pre>\n<p>Score 0.86</p>\n<h3>Metric mean + rolling max diff penalty by <a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a></h3>\n<pre><code> ():\n    penalty = \n\n     i  (rolling_per-,(gini_scores)):\n        penalty = penalty + ((gini_scores[(i-rolling_per+):i+])-gini_scores[i])**\n\n     np.mean(gini_scores) - c*(np.sqrt(penalty)/((gini_scores)-rolling_per+)) \n</code></pre>\n<p>Score 0.85</p>\n<h3>Metric mean log MA by <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a></h3>\n<pre><code> ():\n    x = np.cumsum(gini_in_time, axis=)\n    x = np.concatenate([*x[:], x], )\n    scores = -np.mean(-np.log(np.maximum((x[ma_len:] - x[:-ma_len])/ma_len, ))**exponent)\n     scores \n</code></pre>\n<p>Score: 0.83</p>\n<h3>Metric with weights by <a href=\"https://www.kaggle.com/viktorian\" target=\"_blank\">@viktorian</a></h3>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/(x)\n     w_avg_gini + w_gstd * gini_std\n</code></pre>\n<h3>Metric mean std by <a href=\"https://www.kaggle.com/seifachour12\" target=\"_blank\">@seifachour12</a></h3>\n<pre><code> ():\n     (np.(beta*(/beta*np.mean(gini_in_time) + (gamma-np.std(gini_in_time)))-))**(/alpha) \n</code></pre>\n<p>Score: 0.76</p>\n<h3>Metric mean std by <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a></h3>\n<pre><code> ():\n     np.mean(gini_in_time) - c*np.std(gini_in_time)\n</code></pre>\n<p>Score: 0.74</p>\n<h3>Metric rolling gini max by <a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a></h3>\n<pre><code> ():\n    max_gini = gini_in_time[]\n    cost = \n     i  (, (gini_in_time)):\n        max_gini = (max_gini, gini_in_time[i])\n        cost += (, max_gini - gini_in_time[i])**\n     np.mean(gini_in_time) + c/((gini_in_time)-)*cost   \n</code></pre>\n<p>Score: 0.67</p>\n<h2>Metric stability with averaging by <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a></h2>\n<pre><code> ():\n    w_fallingrate /= f + \n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far = a\n     s  (,f):\n        start_index = (x) // f * (s)\n        end_index = (x) // f * (s+)\n        x_second_fifth = x[start_index:end_index]\n        y_second_fifth = gini_in_time[start_index:end_index]\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        far += (,a1)\n\n     avg_gini + w_fallingrate * (far) + w_resstd * res_std \n</code></pre>\n<h2>Metric mean gini</h2>\n<pre><code> ():\n     np.mean(gini_in_time)  \n</code></pre>\n<p>Score: 0.59</p>\n<h3>Metric sharpe-like by <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a></h3>\n<pre><code> ():\n     np.mean(gini_in_time) / np.std(gini_in_time)\n</code></pre>\n<p>Score 0.52</p>\n<h2>Stability with sliding window by <a href=\"https://www.kaggle.com/naotokai\" target=\"_blank\">@naotokai</a></h2>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    \n    x_past = np.arange(-win, )\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n\n    \n    a1_lst = []\n     start_index  (, (x)-win+):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, )\n        a1_lst.append((,a1))       \n    a1_avg = np.mean(a1_lst)\n\n     avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n</code></pre>\n<p>Score: 0.49</p>\n<p>I think the target is at least 0.90 for us to consider this metric as a substitution for the current stability metric. Any effort will be appreciated, but again this is outside the scope of the competition LB and definitely not mandatory, completely up to you. </p>",
      "rawMarkdown": "Here you can ask me to score any metric you would like to suggest as new stability metric. The score itself does not guarantee that we will use the metric in this competition. The score of a metric is just an approximation of how we rate the metrics, perhaps during the following dates I will improve this method. Think of the score as a percentage of how much it is aligned with an internal business view. \n\n## Rules\n1. If there are any hyperparameters I will optimize them with given boundaries ([-inf, inf] is not a valid interval). \n2. I won't score a metric with more than three hyperparameters that are not set beforehand. \n3. The metric should be relatively easy to explain and should have a connection to the current stability metric, in other words, we would like to see a formula that is easy to read (interpretability) and it is not connected directly to the competition data (generality). \n4. Please send me a simple notebook or easy-to-read code, ideally with sensible variable naming. I won't be implementing metric if you send me just an equation. I will implement it from an equation only if I have spare time, so no guarantee. \n\nScore range is 0-1, where higher is better. \n\n### Current stability metric\n```python\ndef gini_stability(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5):\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    return avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n```\nScore: 0.99\n\n### Metric mean + rolling min/max diff penalty by @davutpolat\n```python\ndef stability_score_2(gini_scores, rolling_per1=94, rolling_per2=94, c1=3.1, c2=3.1):\n    penalty = 0\n    reward = 0\n\n    for i in range(rolling_per1 - 1, len(gini_scores)):\n        penalty = (\n            penalty\n            + (max(gini_scores[(i - rolling_per1 + 1) : i + 1]) - gini_scores[i]) ** 2\n        )\n\n    for i in range(rolling_per2 - 1, len(gini_scores)):\n        reward = (\n            reward\n            + (gini_scores[i] - min(gini_scores[(i - rolling_per2 + 1) : i + 1])) ** 2\n        )\n\n    return (\n        np.mean(gini_scores)\n        - c1 * (np.sqrt(penalty) / (len(gini_scores) - rolling_per1 + 1))\n        + c2 * (np.sqrt(reward) / (len(gini_scores) - rolling_per2 + 1))\n    )\n\n```\nScore 0.86\n\n### Metric mean + rolling max diff penalty by @davutpolat\n```python\ndef metric_davutpolat(gini_scores, rolling_per=92, c=1.9):\n    penalty = 0\n\n    for i in range(rolling_per-1,len(gini_scores)):\n        penalty = penalty + (max(gini_scores[(i-rolling_per+1):i+1])-gini_scores[i])**2\n\n    return np.mean(gini_scores) - c*(np.sqrt(penalty)/(len(gini_scores)-rolling_per+1)) \n```\nScore 0.85\n\n### Metric mean log MA by @jacobyjaeger \n```python\ndef metric_jacobyjaeger(gini_in_time, exponent=69, ma_len=4):\n    x = np.cumsum(gini_in_time, axis=0)\n    x = np.concatenate([0*x[:1], x], 0)\n    scores = -np.mean(-np.log(np.maximum((x[ma_len:] - x[:-ma_len])/ma_len, 1e-5))**exponent)\n    return scores \n```\nScore: 0.83\n\n### Metric with weights by @viktorian\n```python\ndef gini_stability_avg(gini_in_time, w_gstd=-0.5):\n    x = np.arange(len(gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/sum(x)\n    return w_avg_gini + w_gstd * gini_std\n```\n\n### Metric mean std by @seifachour12 \n```python\ndef metric_seifachour12(gini_in_time, alpha=1028, beta=1.97, gamma=0.49):\n    return (np.abs(beta*(1/beta*np.mean(gini_in_time) + (gamma-np.std(gini_in_time)))-1))**(1/alpha) \n```\nScore: 0.76\n\n### Metric mean std by @kononenko \n```python\ndef metric_kononenko(gini_in_time, c=0.90):\n    return np.mean(gini_in_time) - c*np.std(gini_in_time)\n```\nScore: 0.74\n\n### Metric rolling gini max by @davutpolat \n```python\ndef metric_davutpolat(gini_in_time, c=1.82):\n    max_gini = gini_in_time[0]\n    cost = 0\n    for i in range(1, len(gini_in_time)):\n        max_gini = max(max_gini, gini_in_time[i])\n        cost += max(0, max_gini - gini_in_time[i])**2\n    return np.mean(gini_in_time) + c/(len(gini_in_time)-1)*cost   \n```\nScore: 0.67\n\n## Metric stability with averaging by @at7459 \n```python\ndef gini_at7459(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5, f=8):\n    w_fallingrate /= f + 1\n\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far = a\n    for s in range(1,f):\n        start_index = len(x) // f * (s)\n        end_index = len(x) // f * (s+1)\n        x_second_fifth = x[start_index:end_index]\n        y_second_fifth = gini_in_time[start_index:end_index]\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, 1)\n        far += min(0,a1)\n\n    return avg_gini + w_fallingrate * (far) + w_resstd * res_std \n```\n\n## Metric mean gini\n```python\ndef metric_mean(gini_in_time):\n    return np.mean(gini_in_time)  \n```\nScore: 0.59\n\n### Metric sharpe-like by @lucasmorin\n```python\ndef metric_sharpe(gini_in_time):\n    return np.mean(gini_in_time) / np.std(gini_in_time)\n```\nScore 0.52\n\n## Stability with sliding window by @naotokai\n```python\ndef gini_stability_slide_window(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5, win=10):\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    # The countermeasures for hacking in the first segment\n    x_past = np.arange(1-win, 0)\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n\n    # linear regression in each segment\n    a1_lst = []\n    for start_index in range(0, len(x)-win+1):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, 1)\n        a1_lst.append(min(0,a1))       \n    a1_avg = np.mean(a1_lst)\n\n    return avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n```\nScore: 0.49\n\nI think the target is at least 0.90 for us to consider this metric as a substitution for the current stability metric. Any effort will be appreciated, but again this is outside the scope of the competition LB and definitely not mandatory, completely up to you. ",
      "votes": 38
    },
    {
      "id": 2662884,
      "postDate": "2024-02-22T07:22:46.280Z",
      "content": "<p>Under what kind of evaluation does the current metric get a perfect score? Are you just taking the correlation of scores under each metric to the scores under the original?</p>\n<p>I do not believe the current metric deserves a perfect score under any reasonable evaluation. It has numerous problems even putting aside the hackablity. </p>\n<p>First of all taking statistics over raw Gini scores is unlikely to be ideal. Gini scores, being bounded by zero and one, cannot take on extreme values even for rounds that see extremely poor performance. This leads to a tendency to under-emphasize bad rounds which is really the opposite of what you want if you value stability. -log(gini) would be a better starting point.</p>\n<p>Now, you might expect the standard deviation term to capture the risk, but it really can do only a mediocre job of that. Again, when you use gini scores, no deviation can be greater than one away from the mean so exceptionally bad rounds still can only have limited impact. Standard deviation, being only one moment of a distribution, cannot tell us anything about the other moments such as skew and kurtosis, but these moments should also be important to a good risk measure as they have an impact on the frequency of very bad rounds. It's better to punish bad rounds, or periods of bad rounds, directly with a metric that becomes exponentially worse as the round does. Unlike standard deviation, this does not punish good round and hence isn't hackable.</p>\n<p>And then the regression slope. Raw gini scores are again a bad choice here because a bounded quantity cannot maintain a stable slope. arctanh(2*gini-1) would be much nicer for fitting lines to since linear behavior is actually possible in that domain. As it is we can expect scores mostly near 1 and scores mostly near 0 to see much smaller slopes than scores mostly near .5 due to the bounded nature of gini scores. </p>\n<p>While the slope can tell us if scores have trended downwards, it cannot necessarily tell us if scores have a tendency to drift over time. For example if scores happen to follow a 'U' shaped pattern over the test period, this results in zero slope but it does demonstrate a pattern of sustained poor scores, a kind of risk we should take account of. My moving average based metric captures this kind of risk and without the hackability.  </p>\n<p>Using the slope is a bit like trying to extrapolate the performance on future rounds by negatively weighting early rounds. But if any rounds have negative weight, our incentive is of course to make those rounds as bad as possible. And even if you make that hard to do, extrapolation is inherently very sensitive to noise and will make the competition much more random and the possibility that some features may disappear in the future makes the competition even more of a gamble. </p>",
      "rawMarkdown": "Under what kind of evaluation does the current metric get a perfect score? Are you just taking the correlation of scores under each metric to the scores under the original?\n\nI do not believe the current metric deserves a perfect score under any reasonable evaluation. It has numerous problems even putting aside the hackablity. \n\nFirst of all taking statistics over raw Gini scores is unlikely to be ideal. Gini scores, being bounded by zero and one, cannot take on extreme values even for rounds that see extremely poor performance. This leads to a tendency to under-emphasize bad rounds which is really the opposite of what you want if you value stability. -log(gini) would be a better starting point.\n\nNow, you might expect the standard deviation term to capture the risk, but it really can do only a mediocre job of that. Again, when you use gini scores, no deviation can be greater than one away from the mean so exceptionally bad rounds still can only have limited impact. Standard deviation, being only one moment of a distribution, cannot tell us anything about the other moments such as skew and kurtosis, but these moments should also be important to a good risk measure as they have an impact on the frequency of very bad rounds. It's better to punish bad rounds, or periods of bad rounds, directly with a metric that becomes exponentially worse as the round does. Unlike standard deviation, this does not punish good round and hence isn't hackable.\n\nAnd then the regression slope. Raw gini scores are again a bad choice here because a bounded quantity cannot maintain a stable slope. arctanh(2*gini-1) would be much nicer for fitting lines to since linear behavior is actually possible in that domain. As it is we can expect scores mostly near 1 and scores mostly near 0 to see much smaller slopes than scores mostly near .5 due to the bounded nature of gini scores. \n\nWhile the slope can tell us if scores have trended downwards, it cannot necessarily tell us if scores have a tendency to drift over time. For example if scores happen to follow a 'U' shaped pattern over the test period, this results in zero slope but it does demonstrate a pattern of sustained poor scores, a kind of risk we should take account of. My moving average based metric captures this kind of risk and without the hackability.  \n  \nUsing the slope is a bit like trying to extrapolate the performance on future rounds by negatively weighting early rounds. But if any rounds have negative weight, our incentive is of course to make those rounds as bad as possible. And even if you make that hard to do, extrapolation is inherently very sensitive to noise and will make the competition much more random and the possibility that some features may disappear in the future makes the competition even more of a gamble. \n",
      "votes": 7,
      "replies": [
        {
          "id": 2662908,
          "postDate": "2024-02-22T08:01:46.337Z",
          "content": "<p>It's a very interesting discussion. I find your arguments very sound, however there seem to be other factors at play. The key is:</p>\n<blockquote>\n  <p>The score of a metric is just an approximation of how we rate the metrics.</p>\n</blockquote>\n<p>I may be wrong, but it appears to me, and as suggested by <a href=\"https://www.kaggle.com/dmitriyguller\" target=\"_blank\">@dmitriyguller</a>, that the original metric itself is modelled after how business experts perceive the performance. If that's the case, mathematical properties don't matter as much as the experts' perception of performance. In a funny way, it's like an argument with the boss, is it the right thing to do at all? A true moral dilemma here.. 55</p>",
          "rawMarkdown": "It's a very interesting discussion. I find your arguments very sound, however there seem to be other factors at play. The key is:\n\n>The score of a metric is just an approximation of how we rate the metrics.\n\nI may be wrong, but it appears to me, and as suggested by @dmitriyguller, that the original metric itself is modelled after how business experts perceive the performance. If that's the case, mathematical properties don't matter as much as the experts' perception of performance. In a funny way, it's like an argument with the boss, is it the right thing to do at all? A true moral dilemma here.. 55",
          "votes": 3,
          "replies": [
            {
              "id": 2663097,
              "postDate": "2024-02-22T09:54:30.590Z",
              "content": "<p><a href=\"https://www.kaggle.com/lohmaa\" target=\"_blank\">@lohmaa</a> Exactly, the metric does a good job at describing how the risk managers internally rate models in production.</p>",
              "rawMarkdown": "@lohmaa Exactly, the metric does a good job at describing how the risk managers internally rate models in production.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2663094,
          "postDate": "2024-02-22T09:53:00.703Z",
          "content": "<p>Thanks for the extensive comment.</p>\n<blockquote>\n  <p>And then the regression slope. Raw gini scores are again a bad choice here because a bounded quantity cannot maintain a stable slope. arctanh(2*gini-1) would be much nicer for fitting lines to since linear behavior is actually possible in that domain. As it is we can expect scores mostly near 1 and scores mostly near 0 to see much smaller slopes than scores mostly near .5 due to the bounded nature of gini scores.</p>\n</blockquote>\n<p>I like this point. I examined over two hundred models in production within Home Credit before the start of this competition and I was aiming initially for exponential decay. To my surprise, the linear decline was most suitable from a statistical point of view. The reason for that is the models are usually very noisy on weekly ginis. However, yeah, good point. </p>\n<blockquote>\n  <p>While the slope can tell us if scores have trended downwards, it cannot necessarily tell us if scores have a tendency to drift over time. For example if scores happen to follow a 'U' shaped pattern over the test period, this results in zero slope but it does demonstrate a pattern of sustained poor scores, a kind of risk we should take account of. My moving average based metric captures this kind of risk and without the hackability.</p>\n  <p>Using the slope is a bit like trying to extrapolate the performance on future rounds by negatively weighting early rounds. But if any rounds have negative weight, our incentive is of course to make those rounds as bad as possible. And even if you make that hard to do, extrapolation is inherently very sensitive to noise and will make the competition much more random and the possibility that some features may disappear in the future makes the competition even more of a gamble.</p>\n</blockquote>\n<p>I like your proposed metric, the only downside now is that it can not take into the account the overall trend of the weekly gini. This is important point and we are unlikely to ignore it even in this competition. I encourage you to think about ways how to incorporate the trends into your metric. </p>",
          "rawMarkdown": "Thanks for the extensive comment.\n\n>And then the regression slope. Raw gini scores are again a bad choice here because a bounded quantity cannot maintain a stable slope. arctanh(2*gini-1) would be much nicer for fitting lines to since linear behavior is actually possible in that domain. As it is we can expect scores mostly near 1 and scores mostly near 0 to see much smaller slopes than scores mostly near .5 due to the bounded nature of gini scores.\n\nI like this point. I examined over two hundred models in production within Home Credit before the start of this competition and I was aiming initially for exponential decay. To my surprise, the linear decline was most suitable from a statistical point of view. The reason for that is the models are usually very noisy on weekly ginis. However, yeah, good point. \n\n>While the slope can tell us if scores have trended downwards, it cannot necessarily tell us if scores have a tendency to drift over time. For example if scores happen to follow a 'U' shaped pattern over the test period, this results in zero slope but it does demonstrate a pattern of sustained poor scores, a kind of risk we should take account of. My moving average based metric captures this kind of risk and without the hackability.\n\n>Using the slope is a bit like trying to extrapolate the performance on future rounds by negatively weighting early rounds. But if any rounds have negative weight, our incentive is of course to make those rounds as bad as possible. And even if you make that hard to do, extrapolation is inherently very sensitive to noise and will make the competition much more random and the possibility that some features may disappear in the future makes the competition even more of a gamble.\n\nI like your proposed metric, the only downside now is that it can not take into the account the overall trend of the weekly gini. This is important point and we are unlikely to ignore it even in this competition. I encourage you to think about ways how to incorporate the trends into your metric. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2665125,
      "postDate": "2024-02-23T12:34:40.737Z",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> </p>\n<p>From reading the discussion on all pinned threads, it's not clear what the real problem this thread is trying to solve.<br>\nIf, as you say <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2663049\" target=\"_blank\">here</a>, [0.3, 0.3, 0.3] is preferable to [0.6, 0.5, 0.4] then artificially reducing gini becomes merely a form of <strong>legitimate hedging or smoothing</strong> to compensate for overconfidence/overfitting the model in \"better\" weeks, and in my mind no more of a \"hack\" than say artificially clipping predictions to [1e-10, 1 - 1e-10] in a log-loss competition.</p>\n<p>Of course I may be missing something and would love to learn where I went wrong.</p>",
      "rawMarkdown": "@jetakow @tomasjeline2 \n\nFrom reading the discussion on all pinned threads, it's not clear what the real problem this thread is trying to solve.\nIf, as you say [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2663049), [0.3, 0.3, 0.3] is preferable to [0.6, 0.5, 0.4] then artificially reducing gini becomes merely a form of **legitimate hedging or smoothing** to compensate for overconfidence/overfitting the model in \"better\" weeks, and in my mind no more of a \"hack\" than say artificially clipping predictions to [1e-10, 1 - 1e-10] in a log-loss competition.\n\nOf course I may be missing something and would love to learn where I went wrong.",
      "votes": 5
    },
    {
      "id": 2663125,
      "postDate": "2024-02-22T10:21:41.317Z",
      "content": "<p>Firstly, I want to express gratitude for launching this competition, which addresses a crucial topic: the stability of models in time series analysis.</p>\n<p>As you pointed out, the previous metric exhibited an undesirable trait where penalizing the model's early predictions actually increased the score, contradicting intuitive expectations. This issue primarily stems from the necessity of ensuring the stability of the area under the curve (AUC) over time, attempted through the min(0, a) component.</p>\n<p>It is paramount that we achieve stability as time progresses. However, what we truly require is stability across all time points, not solely in future predictions. Hence, the metric I propose is rooted in the notion of maximizing both the mean of the Gini coefficient and its standard deviation over time. By doing so, we aim to ensure that the Gini coefficient remains sufficiently high and stable across all time points.</p>\n<p>` Below is the proposed metric:</p>\n<pre><code> mean_std_gini(gini_in_time, alpha=.):\n     (np.abs(*(.*np.mean(gini_in_time) + (.-np.std(gini_in_time)))-))**alpha`\n</code></pre>\n<p>If gini_in_time is perfect (all ones) the metric returns 1 and if it is the worst (all zeros) the metric is 0.<br>\nHopefully this metric allows us to achieve the competition's goal.</p>",
      "rawMarkdown": "Firstly, I want to express gratitude for launching this competition, which addresses a crucial topic: the stability of models in time series analysis.\n\nAs you pointed out, the previous metric exhibited an undesirable trait where penalizing the model's early predictions actually increased the score, contradicting intuitive expectations. This issue primarily stems from the necessity of ensuring the stability of the area under the curve (AUC) over time, attempted through the min(0, a) component.\n\nIt is paramount that we achieve stability as time progresses. However, what we truly require is stability across all time points, not solely in future predictions. Hence, the metric I propose is rooted in the notion of maximizing both the mean of the Gini coefficient and its standard deviation over time. By doing so, we aim to ensure that the Gini coefficient remains sufficiently high and stable across all time points.\n\n` Below is the proposed metric:\n\n    def mean_std_gini(gini_in_time, alpha=0.01):\n        return (np.abs(2*(0.5*np.mean(gini_in_time) + (0.5-np.std(gini_in_time)))-1))**alpha`\n\nIf gini_in_time is perfect (all ones) the metric returns 1 and if it is the worst (all zeros) the metric is 0.\nHopefully this metric allows us to achieve the competition's goal.",
      "votes": 3,
      "replies": [
        {
          "id": 2663156,
          "postDate": "2024-02-22T10:52:01.363Z",
          "content": "<p>Thanks, I edited it slightly and optimized. Your original metric gives 0.74 score.</p>",
          "rawMarkdown": "Thanks, I edited it slightly and optimized. Your original metric gives 0.74 score.",
          "replies": [
            {
              "id": 2663184,
              "postDate": "2024-02-22T11:03:44.357Z",
              "content": "<p>Thank you for the feedback I suggest that you take alpha = 0.0001 as a starting point before optimizing</p>",
              "rawMarkdown": "Thank you for the feedback I suggest that you take alpha = 0.0001 as a starting point before optimizing"
            },
            {
              "id": 2663205,
              "postDate": "2024-02-22T11:12:49.230Z",
              "content": "<p>I thank you for the idea :) any help is appreciated here very much.</p>\n<pre><code> ():\n     (np.(beta*(/beta*np.mean(gini_in_time) + (gamma-np.std(gini_in_time)))-))**(/alpha) \n</code></pre>\n<p>gives</p>\n<pre><code>, beta, gamma, score = (, ., ., .) \n</code></pre>",
              "rawMarkdown": "I thank you for the idea :) any help is appreciated here very much.\n\n```python\ndef metric_mean_std2(gini_in_time, alpha=0.01, beta=2.0, gamma=0.5):\n    return (np.abs(beta*(1/beta*np.mean(gini_in_time) + (gamma-np.std(gini_in_time)))-1))**(1/alpha) \n```\ngives\n```\nalpha, beta, gamma, score = (1028, 1.97, 0.49, 0.76) \n```"
            },
            {
              "id": 2663242,
              "postDate": "2024-02-22T11:47:08.377Z",
              "content": "<p>No problem glad to provide you with any help :) <br>\nIf you want to change it to 1/alpha the starting alpha should be 10000 but I recommand to keep it alpha not 1/alpha and in general as much the power tend to zero the score will tend to one you can try alpha= 10**-10 (but using **alpha)<br>\nBest of luck! </p>",
              "rawMarkdown": "No problem glad to provide you with any help :) \nIf you want to change it to 1/alpha the starting alpha should be 10000 but I recommand to keep it alpha not 1/alpha and in general as much the power tend to zero the score will tend to one you can try alpha= 10**-10 (but using **alpha)\nBest of luck! "
            }
          ]
        }
      ]
    },
    {
      "id": 2663477,
      "postDate": "2024-02-22T14:04:29.207Z",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> is it possible that you can provide some gini scores (at least 10 gini series with size of  number of weeks in test set )  with their preference order, so we can test our ideas and apply on them to see if they holds HC concerns?  it would be more practical for everyone to see how  their suggestions perform :) </p>",
      "rawMarkdown": "@jetakow is it possible that you can provide some gini scores (at least 10 gini series with size of  number of weeks in test set )  with their preference order, so we can test our ideas and apply on them to see if they holds HC concerns?  it would be more practical for everyone to see how  their suggestions perform :) ",
      "votes": 1,
      "replies": [
        {
          "id": 2663491,
          "postDate": "2024-02-22T14:16:19.437Z",
          "content": "<p>Yes, I am working on something at the moment that could be published. Tomorrow I will be taking a day off, so I am not sure how quickly it will be available here on kaggle. </p>",
          "rawMarkdown": "Yes, I am working on something at the moment that could be published. Tomorrow I will be taking a day off, so I am not sure how quickly it will be available here on kaggle. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2662707,
      "postDate": "2024-02-22T04:56:53.837Z",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> it is not what i suggested.</p>\n<p>$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1} \\sqrt{\\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{rollingMax} - GiniScore_i)\\right)^2}<br>\n$$</p>\n<p>GiniScore_rollingMax : is the highest Gini score observed up to week i <br>\nGiniScore_i : is the Gini score for week i </p>\n<p>Can you update the return term with this : </p>\n<p>return np.mean(gini_in_time) - c/(len(gini_in_time)-1)*np.sqrt(cost)</p>\n<p>it is totally different what i suggested. Can you evaluate it again please?</p>",
      "rawMarkdown": "@jetakow it is not what i suggested.\n\n$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1} \\sqrt{\\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{rollingMax} - GiniScore_i)\\right)^2}\n$$\n\nGiniScore_rollingMax : is the highest Gini score observed up to week i \nGiniScore_i : is the Gini score for week i \n\n\nCan you update the return term with this : \n\n\n return np.mean(gini_in_time) - c/(len(gini_in_time)-1)*np.sqrt(cost)\n\n\n\nit is totally different what i suggested. Can you evaluate it again please?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2662984,
          "postDate": "2024-02-22T09:06:48.673Z",
          "content": "<p>Apologies, the optimized c is 1.82 and the score is 0.67.</p>",
          "rawMarkdown": "Apologies, the optimized c is 1.82 and the score is 0.67.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2662586,
      "postDate": "2024-02-22T02:26:01.610Z",
      "content": "<p>I created a notebook to run some tests.<br>\n<a href=\"https://www.kaggle.com/carloshuertas/homecreditmetrictest\" target=\"_blank\">https://www.kaggle.com/carloshuertas/homecreditmetrictest</a></p>\n<p>I still cant understand the objective. </p>\n<blockquote>\n  <p>A model-X that outperforms week-over-week to model-Y over a long (100 weeks) time-horizon still is \"worse\"</p>\n</blockquote>\n<p>While in real-life, even with a slightly faster degradation this would lead to better decisions, as it outperforms the weaker model, every-single-time. I would understand that if eventually model-Y outperforms, Ok, but not the case in the experiment, yet, still the better model scores worse.</p>\n<p>Am I missing something? The math is not clicking for me, is this based on any peer-reviewed material? It doesnt feel ready.</p>",
      "rawMarkdown": "I created a notebook to run some tests.\nhttps://www.kaggle.com/carloshuertas/homecreditmetrictest\n\nI still cant understand the objective. \n\n>A model-X that outperforms week-over-week to model-Y over a long (100 weeks) time-horizon still is \"worse\"\n\n While in real-life, even with a slightly faster degradation this would lead to better decisions, as it outperforms the weaker model, every-single-time. I would understand that if eventually model-Y outperforms, Ok, but not the case in the experiment, yet, still the better model scores worse.\n\nAm I missing something? The math is not clicking for me, is this based on any peer-reviewed material? It doesnt feel ready.",
      "votes": 1,
      "replies": [
        {
          "id": 2662693,
          "postDate": "2024-02-22T04:42:12Z",
          "content": "<p>It's not clicking with me either.  I think one problem could be with the way that the organizers came up with the metric.  From what I understand, they presented various cases of Model A vs Model B, and asked the expert to express their preferences.  Then they formulated a metric that would fit to those preferences.</p>\n<p>Let's assume for a moment that the experts are perfectly infallible and have a perfectly calibrated sense of how costly the model instability is.  There is still a problem with that approach if you didn't go to the next step and try to be adversarial with the metric.  Once you come up with a metric, can you generate cases specifically to get situations where the metric leads to a clearly wrong conclusion?  Those cases might not come up in a business setting where common sense would put a stop to all nonsensical abuse, but a Kaggle competition where the value of the metric is the only thing that counts is a whole other ballgame.  To me a case where one model is strictly better than another for all weeks, and clearly has a better or equal asymptotic behavior, should never ever score worse.</p>",
          "rawMarkdown": "It's not clicking with me either.  I think one problem could be with the way that the organizers came up with the metric.  From what I understand, they presented various cases of Model A vs Model B, and asked the expert to express their preferences.  Then they formulated a metric that would fit to those preferences.\n\nLet's assume for a moment that the experts are perfectly infallible and have a perfectly calibrated sense of how costly the model instability is.  There is still a problem with that approach if you didn't go to the next step and try to be adversarial with the metric.  Once you come up with a metric, can you generate cases specifically to get situations where the metric leads to a clearly wrong conclusion?  Those cases might not come up in a business setting where common sense would put a stop to all nonsensical abuse, but a Kaggle competition where the value of the metric is the only thing that counts is a whole other ballgame.  To me a case where one model is strictly better than another for all weeks, and clearly has a better or equal asymptotic behavior, should never ever score worse.",
          "votes": 5,
          "replies": [
            {
              "id": 2662838,
              "postDate": "2024-02-22T06:39:47.053Z",
              "content": "<p>I imagine the \"expert modelling\" approach could indeed take place and while it's not a wrong approach it may lead to significant perceptual distortions. In particular, there may be something like perceptual extrapolation, which is causing the disconnect between the business expectations and how the metric looks to kagglers. An example:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5871958%2F6867388250ec3d3fc3b6f06834fe9ec2%2FFigure_1.png?generation=1708583025156692&amp;alt=media\"></p>\n<p>When experts look at the chart they may decide the blue model performs better because it's \"flatter\" and clearly it looks like it's net gini scores are soon going to be better than the orange one. So this is reflected in the original metric and the score is correct according to the expectations. Apparently many kagglers would argue the orange one is better. Personally I like the liberal approach -- the metric has been given and is all we should be concerned to optimize. All hacks allowed. In the worst case it will be a lesson for the experts… ,)</p>",
              "rawMarkdown": "I imagine the \"expert modelling\" approach could indeed take place and while it's not a wrong approach it may lead to significant perceptual distortions. In particular, there may be something like perceptual extrapolation, which is causing the disconnect between the business expectations and how the metric looks to kagglers. An example:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5871958%2F6867388250ec3d3fc3b6f06834fe9ec2%2FFigure_1.png?generation=1708583025156692&alt=media)\n\nWhen experts look at the chart they may decide the blue model performs better because it's \"flatter\" and clearly it looks like it's net gini scores are soon going to be better than the orange one. So this is reflected in the original metric and the score is correct according to the expectations. Apparently many kagglers would argue the orange one is better. Personally I like the liberal approach -- the metric has been given and is all we should be concerned to optimize. All hacks allowed. In the worst case it will be a lesson for the experts... ,)",
              "votes": 7
            },
            {
              "id": 2662902,
              "postDate": "2024-02-22T07:48:23.573Z",
              "content": "<p>i think the main disconnect between the metric and what people think about when they say \"better models\" is this asymptotic behavior you refer to. the homecredit people, with their metric, basically say that the trend is a linear slope and this linear slope will continue forever, so the best model will always be one that has a non-negative slope.</p>\n<p>in most situations, however, there should be a flattening of this slope as time goes on. usually that would mean that, even when the better model has a sharper drop off, it will still consistently outperform the model with the flatter slope because the slope of the performance will flatten over time and not just keep dropping off linearly.</p>\n<p>however one of the many pitfalls here is that the hosts have said that data resources can become unreliable. what this means is that this asymptotic behavior you would usually expect, of the flattening of the performance, is not necessarily true here.</p>\n<p>there are actually many points to think about here, and the conclusion ive personally come to is that (without deep insights into their business) their metric, with the weights theyve chosen, because they never have a reason to artificially make a model worse, is probably perfectly fine for their business overall, and it's just unfortunate for them that it turns out that in a competition setting it is exploitable.</p>",
              "rawMarkdown": "i think the main disconnect between the metric and what people think about when they say \"better models\" is this asymptotic behavior you refer to. the homecredit people, with their metric, basically say that the trend is a linear slope and this linear slope will continue forever, so the best model will always be one that has a non-negative slope.\n\nin most situations, however, there should be a flattening of this slope as time goes on. usually that would mean that, even when the better model has a sharper drop off, it will still consistently outperform the model with the flatter slope because the slope of the performance will flatten over time and not just keep dropping off linearly.\n\nhowever one of the many pitfalls here is that the hosts have said that data resources can become unreliable. what this means is that this asymptotic behavior you would usually expect, of the flattening of the performance, is not necessarily true here.\n\nthere are actually many points to think about here, and the conclusion ive personally come to is that (without deep insights into their business) their metric, with the weights theyve chosen, because they never have a reason to artificially make a model worse, is probably perfectly fine for their business overall, and it's just unfortunate for them that it turns out that in a competition setting it is exploitable.",
              "votes": 7
            },
            {
              "id": 2675116,
              "postDate": "2024-02-29T18:04:32.190Z",
              "content": "<p>I think you are right. Overall it depends on the asymptotic behavior (wether the blue line cross the orange line in the end). From my experience on longer time frame I didn't encouter that much drop and \"crossing\". </p>\n<p>I don't entirely know what to think here… A model loosing 10 AUC points in the course of 6 months is indeed quite bad. But I don't entirely know what is at fault here (blattant overfitting ? volatililty of risk drivers ?).</p>\n<p>Maybe a solution would be to use a good baseline, typically test wether its above and stay above a 'fico' score (or a linear model recalibrated each week). </p>",
              "rawMarkdown": "I think you are right. Overall it depends on the asymptotic behavior (wether the blue line cross the orange line in the end). From my experience on longer time frame I didn't encouter that much drop and \"crossing\". \n\nI don't entirely know what to think here... A model loosing 10 AUC points in the course of 6 months is indeed quite bad. But I don't entirely know what is at fault here (blattant overfitting ? volatililty of risk drivers ?).\n\nMaybe a solution would be to use a good baseline, typically test wether its above and stay above a 'fico' score (or a linear model recalibrated each week). "
            }
          ]
        }
      ]
    },
    {
      "id": 2662456,
      "postDate": "2024-02-21T22:25:25.053Z",
      "content": "<p>What about just a simple </p>\n<pre><code>def metric_mean_std(gini_in_time, k=):\n     .(gini_in_time) - k*.(gini_in_time)\n</code></pre>\n<p>with whatever <code>k</code> you deem necessary to ensure proper stability.</p>",
      "rawMarkdown": "What about just a simple \n\n```\ndef metric_mean_std(gini_in_time, k=0.5):\n    return np.mean(gini_in_time) - k*np.std(gini_in_time)\n```\n\nwith whatever `k` you deem necessary to ensure proper stability.",
      "votes": 1,
      "replies": [
        {
          "id": 2662464,
          "postDate": "2024-02-21T22:42:00.797Z",
          "content": "<p>That gives me score 0.64.</p>",
          "rawMarkdown": "That gives me score 0.64.",
          "votes": 2,
          "replies": [
            {
              "id": 2662704,
              "postDate": "2024-02-22T04:55:41.960Z",
              "content": "<p>How this score is calculated and for what <code>k</code>? My feeling is that the score itself can’t be used to rank metric robustness, as we know the current #1 metric has other serious issues. In my opinion, being simple and closer to standard could result in a better generalization and a lower hackability.</p>\n<p>Btw, in the post you say the score is <code>0.74</code>, not <code>0.64</code>.</p>",
              "rawMarkdown": "How this score is calculated and for what `k`? My feeling is that the score itself can’t be used to rank metric robustness, as we know the current #1 metric has other serious issues. In my opinion, being simple and closer to standard could result in a better generalization and a lower hackability.\n\nBtw, in the post you say the score is `0.74`, not `0.64`.",
              "votes": 1
            },
            {
              "id": 2662956,
              "postDate": "2024-02-22T08:47:14.527Z",
              "content": "<p>I can't say exactly how the score is calculated I am afraid, at least not during the competition. Think of them as checks and you get a percentage of how many checks you pass. Those checks were given by in-house risk managers. Yes, indeed it can't be used to score the metric robustness. If anyone happens to find a metric that is around 0.90 I think we would definitely consider it as a suitable replacement for the current candidate. </p>\n<p>There is 0.74, because I optimized the k to 0.9.</p>",
              "rawMarkdown": "I can't say exactly how the score is calculated I am afraid, at least not during the competition. Think of them as checks and you get a percentage of how many checks you pass. Those checks were given by in-house risk managers. Yes, indeed it can't be used to score the metric robustness. If anyone happens to find a metric that is around 0.90 I think we would definitely consider it as a suitable replacement for the current candidate. \n\nThere is 0.74, because I optimized the k to 0.9."
            }
          ]
        }
      ]
    },
    {
      "id": 2673245,
      "postDate": "2024-02-28T15:29:00.520Z",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> could you test this metric function?</p>\n<p>c and rolling_per should be optimized.<br>\nconstraints: <br>\nc &gt;  0<br>\n2 &lt;= rolling_per &lt;= len(gini_scores)</p>\n<pre><code> stability_score(gini_scores,rolling_per=,c=.):\n     = \n\n     i in range(rolling_per-,len(gini_scores)):\n         = penalty + (max(gini_scores[(i-rolling_per+):i+])-gini_scores[i])**\n\n     np.mean(gini_scores) - c*(np.sqrt(penalty)/(len(gini_scores)-rolling_per+))\n</code></pre>",
      "rawMarkdown": "@jetakow could you test this metric function?\n\nc and rolling_per should be optimized.\nconstraints: \nc >  0\n2 <= rolling_per <= len(gini_scores)\n\n\n```\ndef stability_score(gini_scores,rolling_per=5,c=1.0):\n    penalty = 0\n    \n    for i in range(rolling_per-1,len(gini_scores)):\n        penalty = penalty + (max(gini_scores[(i-rolling_per+1):i+1])-gini_scores[i])**2\n\n    return np.mean(gini_scores) - c*(np.sqrt(penalty)/(len(gini_scores)-rolling_per+1))\n```\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 2673405,
          "postDate": "2024-02-28T16:54:27.743Z",
          "content": "<p>Without optimization it is 0.61</p>",
          "rawMarkdown": "Without optimization it is 0.61",
          "replies": [
            {
              "id": 2673642,
              "postDate": "2024-02-28T19:40:18.490Z",
              "content": "<p>could you optimize the period and c ?</p>",
              "rawMarkdown": "could you optimize the period and c ?"
            }
          ]
        },
        {
          "id": 2673460,
          "postDate": "2024-02-28T17:40:49.983Z",
          "content": "<p>I like it! It gives good score, could be still simple enough to explain and I don't see how it could be hacked. </p>",
          "rawMarkdown": "I like it! It gives good score, could be still simple enough to explain and I don't see how it could be hacked. ",
          "replies": [
            {
              "id": 2673590,
              "postDate": "2024-02-28T19:05:35.057Z",
              "content": "<p>Can the exponent in the formula for penalty possibly optimized to give even better score?</p>",
              "rawMarkdown": "Can the exponent in the formula for penalty possibly optimized to give even better score?"
            },
            {
              "id": 2673895,
              "postDate": "2024-02-28T23:06:24.727Z",
              "content": "<p>I could do it, but somehow I think the L2 penalty will be close to optimal </p>",
              "rawMarkdown": "I could do it, but somehow I think the L2 penalty will be close to optimal "
            },
            {
              "id": 2674630,
              "postDate": "2024-02-29T13:07:58.763Z",
              "content": "<p>Would like to see how L1 performs (more robust wrt outlier)</p>",
              "rawMarkdown": "Would like to see how L1 performs (more robust wrt outlier)"
            },
            {
              "id": 2676286,
              "postDate": "2024-03-01T11:50:47.617Z",
              "content": "<p>L1 gives the same score, just with different params.</p>",
              "rawMarkdown": "L1 gives the same score, just with different params."
            }
          ]
        },
        {
          "id": 2674248,
          "postDate": "2024-02-29T07:37:01.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>  can you also test this one? I feel like this one is also unhackable and it should give good score.<br>\nthe idea is that metric focuses on stability over small segments of timeline. when  someone worsen the gini performance over some time segment, metric will be ruined due to mean(gini)+reward part. </p>\n<p>c1,c2, rolling_per1, and rolling_per2 should be optimized.<br>\nconstraints:<br>\nc1 &gt; 0, c2&gt;0</p>\n<p>2 &lt;= rolling_per1 &lt;= len(gini_scores)<br>\n2 &lt;= rolling_per2 &lt;= len(gini_scores)</p>\n<p>for simplicity you can try <br>\nc1=c2 and rolling_per1=rolling_per2</p>\n<pre><code> stability_score_2(gini_scores,rolling_per1=,rolling_per2=,c1=.,c2=.):\n     = \n      = \n\n     i in range(rolling_per1-,len(gini_scores)):\n         = penalty + (max(gini_scores[(i-rolling_per1+):i+])-gini_scores[i])**\n\n     i in range(rolling_per2-,len(gini_scores)):\n         = reward + (gini_scores[i]-min(gini_scores[(i-rolling_per2+):i+]))**\n\n     np.mean(gini_scores) - c1*(np.sqrt(penalty)/(len(gini_scores)-rolling_per1+)) + c2*(np.sqrt(reward)/(len(gini_scores)-rolling_per2+))\n</code></pre>\n<p>i am aware of that the code is a little dirty, after optimizing the params it can be simplified. </p>",
          "rawMarkdown": "@jetakow  can you also test this one? I feel like this one is also unhackable and it should give good score.\nthe idea is that metric focuses on stability over small segments of timeline. when  someone worsen the gini performance over some time segment, metric will be ruined due to mean(gini)+reward part. \n\nc1,c2, rolling_per1, and rolling_per2 should be optimized.\nconstraints:\nc1 > 0, c2>0\n\n2 <= rolling_per1 <= len(gini_scores)\n2 <= rolling_per2 <= len(gini_scores)\n\nfor simplicity you can try \nc1=c2 and rolling_per1=rolling_per2\n\n```\ndef stability_score_2(gini_scores,rolling_per1=5,rolling_per2=5,c1=1.0,c2=1.0):\n    penalty = 0\n    reward  = 0\n    \n    for i in range(rolling_per1-1,len(gini_scores)):\n        penalty = penalty + (max(gini_scores[(i-rolling_per1+1):i+1])-gini_scores[i])**2\n        \n    for i in range(rolling_per2-1,len(gini_scores)):\n        reward = reward + (gini_scores[i]-min(gini_scores[(i-rolling_per2+1):i+1]))**2\n\n    return np.mean(gini_scores) - c1*(np.sqrt(penalty)/(len(gini_scores)-rolling_per1+1)) + c2*(np.sqrt(reward)/(len(gini_scores)-rolling_per2+1))\n\n```\n\ni am aware of that the code is a little dirty, after optimizing the params it can be simplified. ",
          "votes": 1,
          "replies": [
            {
              "id": 2674643,
              "postDate": "2024-02-29T13:12:32.380Z",
              "content": "<p>Seems like that gives a score of 0.86 </p>",
              "rawMarkdown": "Seems like that gives a score of 0.86 "
            },
            {
              "id": 2674645,
              "postDate": "2024-02-29T13:15:07.163Z",
              "content": "<p>did you optimize parameters ? and how was the score for previous one ? </p>",
              "rawMarkdown": "did you optimize parameters ? and how was the score for previous one ? ",
              "votes": 1
            },
            {
              "id": 2674870,
              "postDate": "2024-02-29T15:15:11.300Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nwhat is the len(gini_scores) in  your test code? </p>\n<p>if it is near 94, then rolling_per1_2=94 means nothing.  the formula omits from 0 to rolling_per1_2 for calculating <br>\nreward/penalty part. it focuses on indice range  between rolling_per1_2 and -1(end)</p>",
              "rawMarkdown": "@jetakow \nwhat is the len(gini_scores) in  your test code? \n\nif it is near 94, then rolling_per1_2=94 means nothing.  the formula omits from 0 to rolling_per1_2 for calculating \nreward/penalty part. it focuses on indice range  between rolling_per1_2 and -1(end)"
            },
            {
              "id": 2676285,
              "postDate": "2024-03-01T11:50:26.420Z",
              "content": "<p><a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a> I understand and it is close to that number. Yes, I optimized the params.</p>",
              "rawMarkdown": "@davutpolat I understand and it is close to that number. Yes, I optimized the params."
            }
          ]
        }
      ]
    },
    {
      "id": 2673048,
      "postDate": "2024-02-28T13:07:06.440Z",
      "content": "<p>what about some 'sharpe' ?</p>\n<pre><code> ():\n     np.mean(gini_in_time) / np.std(gini_in_time)\n</code></pre>",
      "rawMarkdown": "what about some 'sharpe' ?\n\n```python\ndef metric_sharpe(gini_in_time):\n    return np.mean(gini_in_time) / np.std(gini_in_time)\n```",
      "votes": 2,
      "replies": [
        {
          "id": 2673119,
          "postDate": "2024-02-28T13:46:05.593Z",
          "content": "<p>gini series1 : [0.7, 0.6, 0.7, 0.6 , 0.7, 0.6 ]<br>\nsharpe1  = 13</p>\n<p>gini series2 : [0.099, 0.1, 0.099, 0.1 , 0.099, 0.1 ]<br>\nsharpe2 : 199</p>\n<p>which one would you prefer  ? </p>",
          "rawMarkdown": "gini series1 : [0.7, 0.6, 0.7, 0.6 , 0.7, 0.6 ]\nsharpe1  = 13\n\n gini series2 : [0.099, 0.1, 0.099, 0.1 , 0.099, 0.1 ]\nsharpe2 : 199\n\nwhich one would you prefer  ? ",
          "replies": [
            {
              "id": 2673152,
              "postDate": "2024-02-28T14:08:29.357Z",
              "content": "<p>I know that weird use case. But is it actually possible to reduce std that way ?</p>",
              "rawMarkdown": "I know that weird use case. But is it actually possible to reduce std that way ?"
            },
            {
              "id": 2673223,
              "postDate": "2024-02-28T15:11:45.823Z",
              "content": "<p>it also does not handle gini performance decrease in time. By stability, it is meant that the model should have   high mean gini score and non decreasing performance in time  as much as possible.  Basic Sharpe  ratio does not have such time variant term.  </p>",
              "rawMarkdown": "it also does not handle gini performance decrease in time. By stability, it is meant that the model should have   high mean gini score and non decreasing performance in time  as much as possible.  Basic Sharpe  ratio does not have such time variant term.  "
            },
            {
              "id": 2673503,
              "postDate": "2024-02-28T18:28:06.033Z",
              "content": "<p>It penalizes variation, both upward and downward. Not sure it is a good idea to introduce an asymmetric sign. </p>",
              "rawMarkdown": "It penalizes variation, both upward and downward. Not sure it is a good idea to introduce an asymmetric sign. "
            }
          ]
        },
        {
          "id": 2674579,
          "postDate": "2024-02-29T12:28:02.600Z",
          "content": "<p>That gives 0.52</p>",
          "rawMarkdown": "That gives 0.52"
        }
      ]
    },
    {
      "id": 2667785,
      "postDate": "2024-02-25T10:55:40.013Z",
      "content": "<p>I think if the sum and standard deviation of scores and are the same, the model would consider the following order preferable in a business context:</p>\n<ol>\n<li>Normal distribution, like <code>[0.61, 0.59, 0.63, 0.59, 0.58]</code></li>\n<li>Monotonically increasing or monotonically decreasing, like <code>[0.62, 0.61, 0.60, 0.59, 0.58]</code></li>\n<li>Domain shift, like <code>[0.7, 0.7, 0.7, 0.7, 0.2]</code></li>\n</ol>\n<p>I have created a notebook to check it with some test cases.<br>\n<a href=\"https://www.kaggle.com/code/kurupical/testcase-of-new-metric-for-business-applications/notebook?scriptVersionId=164230696\" target=\"_blank\">https://www.kaggle.com/code/kurupical/testcase-of-new-metric-for-business-applications/notebook?scriptVersionId=164230696</a></p>",
      "rawMarkdown": "I think if the sum and standard deviation of scores and are the same, the model would consider the following order preferable in a business context:\n1. Normal distribution, like ``[0.61, 0.59, 0.63, 0.59, 0.58]``\n2. Monotonically increasing or monotonically decreasing, like ``[0.62, 0.61, 0.60, 0.59, 0.58]``\n3. Domain shift, like `` [0.7, 0.7, 0.7, 0.7, 0.2]``\n\nI have created a notebook to check it with some test cases.\nhttps://www.kaggle.com/code/kurupical/testcase-of-new-metric-for-business-applications/notebook?scriptVersionId=164230696\n",
      "votes": 2,
      "replies": [
        {
          "id": 2667864,
          "postDate": "2024-02-25T11:58:33.247Z",
          "content": "<p>Thanks for the work. Kudos for the useful notebook. I think what should be also done is to start with declining performance (setting <code>a</code> and <code>std(residues)</code>) and then sample it let's say 10**7 times and see how each hack can improve the score in q=0.99,0.999 compared to stability metric. This could prove the robustness for some of them. </p>",
          "rawMarkdown": "Thanks for the work. Kudos for the useful notebook. I think what should be also done is to start with declining performance (setting `a` and `std(residues)`) and then sample it let's say 10**7 times and see how each hack can improve the score in q=0.99,0.999 compared to stability metric. This could prove the robustness for some of them. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2806853,
      "postDate": "2024-05-11T11:03:41.557Z",
      "content": "<p>Great job and nostalgic…</p>",
      "rawMarkdown": "Great job and nostalgic..."
    },
    {
      "id": 2682520,
      "postDate": "2024-03-05T11:24:09.943Z",
      "content": "<p>This is so great, I'm moved to express my feelings</p>",
      "rawMarkdown": "This is so great, I'm moved to express my feelings"
    },
    {
      "id": 2667232,
      "postDate": "2024-02-25T01:23:47.033Z",
      "content": "<p>I don't know much about this competition, including this field. If there are any mistakes, please feel free to point them out</p>\n<p>I think we can multiply the gini coefficient for each week by a different weight, with the weight increasing later on. For example, if we have four weeks of gini coefficient, we can multiply them by 1/10, 2/10, 3/10, 4/10 respectively,10=1+2+3+4.</p>",
      "rawMarkdown": "I don't know much about this competition, including this field. If there are any mistakes, please feel free to point them out\n\n\nI think we can multiply the gini coefficient for each week by a different weight, with the weight increasing later on. For example, if we have four weeks of gini coefficient, we can multiply them by 1/10, 2/10, 3/10, 4/10 respectively,10=1+2+3+4."
    },
    {
      "id": 2665972,
      "postDate": "2024-02-24T03:13:59.633Z",
      "content": "<p>I have a question: If the week_num field in the test data is removed, how will gini_in_time be calculated?</p>",
      "rawMarkdown": "I have a question: If the week_num field in the test data is removed, how will gini_in_time be calculated?",
      "replies": [
        {
          "id": 2666310,
          "postDate": "2024-02-24T09:57:58.397Z",
          "content": "<p>It will be only removed in the data that are available to kagglers when running the submission. The final decision is still not made.</p>",
          "rawMarkdown": "It will be only removed in the data that are available to kagglers when running the submission. The final decision is still not made.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2665549,
      "postDate": "2024-02-23T18:02:11.073Z",
      "content": "<p>Based on the idea of calculating the slope in various parts by <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>, I have created the following metric:</p>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    \n    x_past = np.arange(-win, )\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n\n    \n    a1_lst = []\n     start_index  (, (x)-win+):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, )\n        a1_lst.append((,a1))       \n    a1_avg = np.mean(a1_lst)\n\n     avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n</code></pre>\n<p>The points are as follows:</p>\n<ul>\n<li>setting the hop length to 1 instead of matching it to the frame length in order to avoid hacking</li>\n<li>taking measures in order to avoid hacking in the first segment</li>\n</ul>\n<p>As far as I have tried, I could not increase the score by hacking. <br>\nHere is the experimental notebook:<br>\n<a href=\"https://www.kaggle.com/code/naotokai/gini-stability-slide-window\" target=\"_blank\">https://www.kaggle.com/code/naotokai/gini-stability-slide-window</a></p>",
      "rawMarkdown": "Based on the idea of calculating the slope in various parts by @at7459, I have created the following metric:\n```python\ndef gini_stability_slide_window(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5, win=10):\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    \n    # The countermeasures for hacking in the first segment\n    x_past = np.arange(1-win, 0)\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n    \n    # linear regression in each segment\n    a1_lst = []\n    for start_index in range(0, len(x)-win+1):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, 1)\n        a1_lst.append(min(0,a1))       \n    a1_avg = np.mean(a1_lst)\n\n    return avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n```\nThe points are as follows:\n- setting the hop length to 1 instead of matching it to the frame length in order to avoid hacking\n- taking measures in order to avoid hacking in the first segment\n\nAs far as I have tried, I could not increase the score by hacking. \nHere is the experimental notebook:\nhttps://www.kaggle.com/code/naotokai/gini-stability-slide-window",
      "replies": [
        {
          "id": 2671488,
          "postDate": "2024-02-27T14:33:16.277Z",
          "content": "<p>That gives score of 0.49</p>",
          "rawMarkdown": "That gives score of 0.49"
        }
      ]
    },
    {
      "id": 2665413,
      "postDate": "2024-02-23T16:39:32.910Z",
      "content": "<p>Wouldn't taking the weighted average over time take care of the hacking problem? While not a perfect metric  in terms of capturing sharp drops in Gini on occasional weeks or skewness of the distribution, it would capture the over-time model deterioration by giving higher weights to further away data points and prioritises the absolute Gini value which prevents the hack. Something like below in simplest version, but how the weights are allocated can be modified.</p>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/(x)\n     w_avg_gini + w_gstd * gini_std\n</code></pre>",
      "rawMarkdown": "Wouldn't taking the weighted average over time take care of the hacking problem? While not a perfect metric  in terms of capturing sharp drops in Gini on occasional weeks or skewness of the distribution, it would capture the over-time model deterioration by giving higher weights to further away data points and prioritises the absolute Gini value which prevents the hack. Something like below in simplest version, but how the weights are allocated can be modified.\n\n```python\ndef gini_stability_avg(gini_in_time, w_gstd=-0.5):\n    x = np.arange(len(gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/sum(x)\n    return w_avg_gini + w_gstd * gini_std\n```",
      "replies": [
        {
          "id": 2674583,
          "postDate": "2024-02-29T12:33:44.020Z",
          "content": "<p>That gives the score 0.77, but I have to say it is a bit misaligned with the intention of the initial metric. At least with such a steep weights. </p>",
          "rawMarkdown": "That gives the score 0.77, but I have to say it is a bit misaligned with the intention of the initial metric. At least with such a steep weights. "
        }
      ]
    },
    {
      "id": 2663212,
      "postDate": "2024-02-22T11:21:30.100Z",
      "content": "<p>In the current stability metric<br>\nStability = mean(gini) + 88<em>min(0,a) -0.5</em>std(residuals)</p>\n<p>I suggest to change the first term to be:<br>\nStability = <strong>mean(gini1,gini2,…giniN)</strong> + 88 * min(0,a) - 0.5 * std(residuals)</p>\n<p>Where N is a parameter, the mean is for first N weeks rather than all weeks. The quality of predictions of the rest of the weeks will be taken care of by the other two terms.</p>",
      "rawMarkdown": "In the current stability metric\nStability = mean(gini) + 88*min(0,a) -0.5*std(residuals)\n\nI suggest to change the first term to be:\nStability = **mean(gini1,gini2,…giniN)** + 88 * min(0,a) - 0.5 * std(residuals)\n\nWhere N is a parameter, the mean is for first N weeks rather than all weeks. The quality of predictions of the rest of the weeks will be taken care of by the other two terms.\n"
    },
    {
      "id": 2663005,
      "postDate": "2024-02-22T09:19:36.890Z",
      "content": "<p>I like the idea, it may be an interesting approach to calculate slopes over various parts of <code>gini_in_time</code> to avoid hacks. <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> </p>",
      "rawMarkdown": "I like the idea, it may be an interesting approach to calculate slopes over various parts of `gini_in_time` to avoid hacks. @at7459 ",
      "replies": [
        {
          "id": 2663083,
          "postDate": "2024-02-22T09:48:37.370Z",
          "content": "<p>the formula i posted was nonsense, but yeah that was the idea i dont know if it would fix anything though 😅 i gotta go for a little while now too</p>",
          "rawMarkdown": "the formula i posted was nonsense, but yeah that was the idea i dont know if it would fix anything though 😅 i gotta go for a little while now too"
        },
        {
          "id": 2663147,
          "postDate": "2024-02-22T10:46:34.700Z",
          "content": "<p>something like this maybe could work (at least i couldnt hack it in my attempts right now and it still has a big emphasis on the stability) i really gotta go now though</p>\n<pre><code>f=\n ():\n    gini_in_time = base.loc[:, [, , ]]\\\n        .sort_values()\\\n        .groupby()[[, ]]\\\n        .apply( x: *roc_auc_score(x[], x[])-).tolist()\n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far=a\n     s  (,f):\n        start_index = (x) // f * (s)\n        end_index = (x) // f * (s+)\n        x_second_fifth = x[start_index:end_index]\n        y_second_fifth = gini_in_time[start_index:end_index]\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        far+=(,a1)\n\n\n     avg_gini + w_fallingrate * (far) + w_resstd * res_std \n</code></pre>",
          "rawMarkdown": "something like this maybe could work (at least i couldnt hack it in my attempts right now and it still has a big emphasis on the stability) i really gotta go now though\n\n```python\nf=8\ndef gini_stability(base, w_fallingrate=88.0/(f+1), w_resstd=-0.5):\n    gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far=a\n    for s in range(1,f):\n        start_index = len(x) // f * (s)\n        end_index = len(x) // f * (s+1)\n        x_second_fifth = x[start_index:end_index]\n        y_second_fifth = gini_in_time[start_index:end_index]\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, 1)\n        far+=min(0,a1)\n    \n\n    return avg_gini + w_fallingrate * (far) + w_resstd * res_std \n```\n",
          "votes": 1,
          "replies": [
            {
              "id": 2663196,
              "postDate": "2024-02-22T11:06:41Z",
              "content": "<p>That gives a score of 0.62. I will try to think of a ways how to improve it. </p>",
              "rawMarkdown": "That gives a score of 0.62. I will try to think of a ways how to improve it. "
            },
            {
              "id": 2665131,
              "postDate": "2024-02-23T12:44:13.473Z",
              "content": "<p>hmm yeah, there is a tradeoff between how much it values stability versus how hackable it is - for the competition probably f=7 is the lowest you can go,</p>\n<p>also, what this is doing is a bit different than the original metric, because it is not only concerned about the overall trend but also the trend of smaller non-overlapping time periods. this is, of course, what makes it not hackable because you can only hack one period (or with this implementation at most two periods) at a time and the weight of the period is small enough that you hurt your overall score while doing it.</p>\n<p>what i noticed in your simulated data, in the other thread, is that it is very swingy in these smaller time periods contrary to my testing is using models trained on the actual training data, which is much more closer aligned with the overall trend over the entire time period…</p>\n<p>i think this averaging might work well for real data because these non-overlapping time periods are more or less the same as the overall time period in terms of the size of stability. in your simulated data, however, there are periods where it suddenly jumps all over the place giving a disproportional amount of weight to this time period for the overall calculation.</p>\n<p>(also, i made a hasty mistake: far=a in the metric is supposed to be far=min(0,a) but it doesnt really matter)</p>",
              "rawMarkdown": "hmm yeah, there is a tradeoff between how much it values stability versus how hackable it is - for the competition probably f=7 is the lowest you can go,\n\nalso, what this is doing is a bit different than the original metric, because it is not only concerned about the overall trend but also the trend of smaller non-overlapping time periods. this is, of course, what makes it not hackable because you can only hack one period (or with this implementation at most two periods) at a time and the weight of the period is small enough that you hurt your overall score while doing it.\n\nwhat i noticed in your simulated data, in the other thread, is that it is very swingy in these smaller time periods contrary to my testing is using models trained on the actual training data, which is much more closer aligned with the overall trend over the entire time period...\n\ni think this averaging might work well for real data because these non-overlapping time periods are more or less the same as the overall time period in terms of the size of stability. in your simulated data, however, there are periods where it suddenly jumps all over the place giving a disproportional amount of weight to this time period for the overall calculation.\n\n(also, i made a hasty mistake: far=a in the metric is supposed to be far=min(0,a) but it doesnt really matter)"
            },
            {
              "id": 2665161,
              "postDate": "2024-02-23T13:23:29.680Z",
              "content": "<p>Please note that it can also be hacked by reducing the gini score at the beginning of each segment:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4416902%2Fa50b6a14215711ebeca4e1cdd03fc998%2Foutput.png?generation=1708694385542872&amp;alt=media\"></p>\n<p>It would be better to set the hop length to 1 instead of matching it to the frame length, to avoid hacking. For example, like this:</p>\n<pre><code> ():\n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    a1_lst = []\n     start_index  (, (x)-win):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        a1_lst.append((,a1))\n    a = np.mean(a1_lst)\n\n     avg_gini + w_fallingrate * a + w_resstd * res_std \n</code></pre>\n<p>Here is the experiment notebook.<br>\n<a href=\"https://www.kaggle.com/naotokai/metic-experiments\" target=\"_blank\">https://www.kaggle.com/naotokai/metic-experiments</a></p>",
              "rawMarkdown": "Please note that it can also be hacked by reducing the gini score at the beginning of each segment:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4416902%2Fa50b6a14215711ebeca4e1cdd03fc998%2Foutput.png?generation=1708694385542872&alt=media)\n\nIt would be better to set the hop length to 1 instead of matching it to the frame length, to avoid hacking. For example, like this:\n```python\ndef gini_at7459_v2(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5, win=6):\n    \n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    \n    a1_lst = []\n    for start_index in range(0, len(x)-win):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, 1)\n        a1_lst.append(min(0,a1))\n    a = np.mean(a1_lst)\n    \n    return avg_gini + w_fallingrate * a + w_resstd * res_std \n```\n\nHere is the experiment notebook.\nhttps://www.kaggle.com/naotokai/metic-experiments",
              "votes": 1
            },
            {
              "id": 2665292,
              "postDate": "2024-02-23T15:20:35.977Z",
              "content": "<p>i just noticed myself that with some good accuracy i can still hack things by like 0.02 for realistic scenarios. i just saw ur comment, and honestly, i am not even gonna bother anymore for now and just give up, i dont think it's possible to fully fix.</p>\n<pre><code>f=\n ():\n    gini_in_time = base.loc[:, [, , ]]\\\n        .sort_values()\\\n        .groupby()[[, ]]\\\n        .apply( x: *roc_auc_score(x[], x[])-).tolist()\n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far=\n\n     s  (,f):\n        start_index = (x) // f * (s)\n        end_index = (x) // f * (s+)\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        far += (,a1)\n     avg_gini + w_fallingrate * (far) + w_resstd * res_std \n</code></pre>\n<p>this is the best i could come up with in the end for realistic scenarios, but it's still possible to hack with good accuracy.</p>",
              "rawMarkdown": "i just noticed myself that with some good accuracy i can still hack things by like 0.02 for realistic scenarios. i just saw ur comment, and honestly, i am not even gonna bother anymore for now and just give up, i dont think it's possible to fully fix.\n```python\nf=14\ndef gini_stability(base, w_fallingrate=88/(6*f), w_resstd=-0.5):\n    gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far=0\n\n    for s in range(0,f):\n        start_index = len(x) // f * (s)\n        end_index = len(x) // f * (s+1)\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, 1)\n        far += min(0,a1)\n    return avg_gini + w_fallingrate * (far) + w_resstd * res_std \n```\n\nthis is the best i could come up with in the end for realistic scenarios, but it's still possible to hack with good accuracy.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2664871,
      "postDate": "2024-02-23T08:58:29.150Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2662979,
      "postDate": "2024-02-22T09:00:47.777Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2662884,
      "author_name": "Jacoby Jaeger",
      "author_url": "",
      "post_date": "2024-02-22T07:22:46.280000",
      "content": "<p>Under what kind of evaluation does the current metric get a perfect score? Are you just taking the correlation of scores under each metric to the scores under the original?</p>\n<p>I do not believe the current metric deserves a perfect score under any reasonable evaluation. It has numerous problems even putting aside the hackablity. </p>\n<p>First of all taking statistics over raw Gini scores is unlikely to be ideal. Gini scores, being bounded by zero and one, cannot take on extreme values even for rounds that see extremely poor performance. This leads to a tendency to under-emphasize bad rounds which is really the opposite of what you want if you value stability. -log(gini) would be a better starting point.</p>\n<p>Now, you might expect the standard deviation term to capture the risk, but it really can do only a mediocre job of that. Again, when you use gini scores, no deviation can be greater than one away from the mean so exceptionally bad rounds still can only have limited impact. Standard deviation, being only one moment of a distribution, cannot tell us anything about the other moments such as skew and kurtosis, but these moments should also be important to a good risk measure as they have an impact on the frequency of very bad rounds. It's better to punish bad rounds, or periods of bad rounds, directly with a metric that becomes exponentially worse as the round does. Unlike standard deviation, this does not punish good round and hence isn't hackable.</p>\n<p>And then the regression slope. Raw gini scores are again a bad choice here because a bounded quantity cannot maintain a stable slope. arctanh(2*gini-1) would be much nicer for fitting lines to since linear behavior is actually possible in that domain. As it is we can expect scores mostly near 1 and scores mostly near 0 to see much smaller slopes than scores mostly near .5 due to the bounded nature of gini scores. </p>\n<p>While the slope can tell us if scores have trended downwards, it cannot necessarily tell us if scores have a tendency to drift over time. For example if scores happen to follow a 'U' shaped pattern over the test period, this results in zero slope but it does demonstrate a pattern of sustained poor scores, a kind of risk we should take account of. My moving average based metric captures this kind of risk and without the hackability.  </p>\n<p>Using the slope is a bit like trying to extrapolate the performance on future rounds by negatively weighting early rounds. But if any rounds have negative weight, our incentive is of course to make those rounds as bad as possible. And even if you make that hard to do, extrapolation is inherently very sensitive to noise and will make the competition much more random and the possibility that some features may disappear in the future makes the competition even more of a gamble. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 2662908,
          "author_name": "loh-maa",
          "author_url": "",
          "post_date": "2024-02-22T08:01:46.337000",
          "content": "<p>It's a very interesting discussion. I find your arguments very sound, however there seem to be other factors at play. The key is:</p>\n<blockquote>\n  <p>The score of a metric is just an approximation of how we rate the metrics.</p>\n</blockquote>\n<p>I may be wrong, but it appears to me, and as suggested by <a href=\"https://www.kaggle.com/dmitriyguller\" target=\"_blank\">@dmitriyguller</a>, that the original metric itself is modelled after how business experts perceive the performance. If that's the case, mathematical properties don't matter as much as the experts' perception of performance. In a funny way, it's like an argument with the boss, is it the right thing to do at all? A true moral dilemma here.. 55</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2663097,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T09:54:30.590000",
              "content": "<p><a href=\"https://www.kaggle.com/lohmaa\" target=\"_blank\">@lohmaa</a> Exactly, the metric does a good job at describing how the risk managers internally rate models in production.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2663094,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-22T09:53:00.703000",
          "content": "<p>Thanks for the extensive comment.</p>\n<blockquote>\n  <p>And then the regression slope. Raw gini scores are again a bad choice here because a bounded quantity cannot maintain a stable slope. arctanh(2*gini-1) would be much nicer for fitting lines to since linear behavior is actually possible in that domain. As it is we can expect scores mostly near 1 and scores mostly near 0 to see much smaller slopes than scores mostly near .5 due to the bounded nature of gini scores.</p>\n</blockquote>\n<p>I like this point. I examined over two hundred models in production within Home Credit before the start of this competition and I was aiming initially for exponential decay. To my surprise, the linear decline was most suitable from a statistical point of view. The reason for that is the models are usually very noisy on weekly ginis. However, yeah, good point. </p>\n<blockquote>\n  <p>While the slope can tell us if scores have trended downwards, it cannot necessarily tell us if scores have a tendency to drift over time. For example if scores happen to follow a 'U' shaped pattern over the test period, this results in zero slope but it does demonstrate a pattern of sustained poor scores, a kind of risk we should take account of. My moving average based metric captures this kind of risk and without the hackability.</p>\n  <p>Using the slope is a bit like trying to extrapolate the performance on future rounds by negatively weighting early rounds. But if any rounds have negative weight, our incentive is of course to make those rounds as bad as possible. And even if you make that hard to do, extrapolation is inherently very sensitive to noise and will make the competition much more random and the possibility that some features may disappear in the future makes the competition even more of a gamble.</p>\n</blockquote>\n<p>I like your proposed metric, the only downside now is that it can not take into the account the overall trend of the weekly gini. This is important point and we are unlikely to ignore it even in this competition. I encourage you to think about ways how to incorporate the trends into your metric. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2665125,
      "author_name": "Nigel A. R. Henry",
      "author_url": "",
      "post_date": "2024-02-23T12:34:40.737000",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> </p>\n<p>From reading the discussion on all pinned threads, it's not clear what the real problem this thread is trying to solve.<br>\nIf, as you say <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2663049\" target=\"_blank\">here</a>, [0.3, 0.3, 0.3] is preferable to [0.6, 0.5, 0.4] then artificially reducing gini becomes merely a form of <strong>legitimate hedging or smoothing</strong> to compensate for overconfidence/overfitting the model in \"better\" weeks, and in my mind no more of a \"hack\" than say artificially clipping predictions to [1e-10, 1 - 1e-10] in a log-loss competition.</p>\n<p>Of course I may be missing something and would love to learn where I went wrong.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2663125,
      "author_name": "Seif Eddine Achour",
      "author_url": "",
      "post_date": "2024-02-22T10:21:41.317000",
      "content": "<p>Firstly, I want to express gratitude for launching this competition, which addresses a crucial topic: the stability of models in time series analysis.</p>\n<p>As you pointed out, the previous metric exhibited an undesirable trait where penalizing the model's early predictions actually increased the score, contradicting intuitive expectations. This issue primarily stems from the necessity of ensuring the stability of the area under the curve (AUC) over time, attempted through the min(0, a) component.</p>\n<p>It is paramount that we achieve stability as time progresses. However, what we truly require is stability across all time points, not solely in future predictions. Hence, the metric I propose is rooted in the notion of maximizing both the mean of the Gini coefficient and its standard deviation over time. By doing so, we aim to ensure that the Gini coefficient remains sufficiently high and stable across all time points.</p>\n<p>` Below is the proposed metric:</p>\n<pre><code> mean_std_gini(gini_in_time, alpha=.):\n     (np.abs(*(.*np.mean(gini_in_time) + (.-np.std(gini_in_time)))-))**alpha`\n</code></pre>\n<p>If gini_in_time is perfect (all ones) the metric returns 1 and if it is the worst (all zeros) the metric is 0.<br>\nHopefully this metric allows us to achieve the competition's goal.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2663156,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-22T10:52:01.363000",
          "content": "<p>Thanks, I edited it slightly and optimized. Your original metric gives 0.74 score.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2663184,
              "author_name": "Seif Eddine Achour",
              "author_url": "",
              "post_date": "2024-02-22T11:03:44.357000",
              "content": "<p>Thank you for the feedback I suggest that you take alpha = 0.0001 as a starting point before optimizing</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2663205,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T11:12:49.230000",
              "content": "<p>I thank you for the idea :) any help is appreciated here very much.</p>\n<pre><code> ():\n     (np.(beta*(/beta*np.mean(gini_in_time) + (gamma-np.std(gini_in_time)))-))**(/alpha) \n</code></pre>\n<p>gives</p>\n<pre><code>, beta, gamma, score = (, ., ., .) \n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2663242,
              "author_name": "Seif Eddine Achour",
              "author_url": "",
              "post_date": "2024-02-22T11:47:08.377000",
              "content": "<p>No problem glad to provide you with any help :) <br>\nIf you want to change it to 1/alpha the starting alpha should be 10000 but I recommand to keep it alpha not 1/alpha and in general as much the power tend to zero the score will tend to one you can try alpha= 10**-10 (but using **alpha)<br>\nBest of luck! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2663477,
      "author_name": "Davut Polat",
      "author_url": "",
      "post_date": "2024-02-22T14:04:29.207000",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> is it possible that you can provide some gini scores (at least 10 gini series with size of  number of weeks in test set )  with their preference order, so we can test our ideas and apply on them to see if they holds HC concerns?  it would be more practical for everyone to see how  their suggestions perform :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2663491,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-22T14:16:19.437000",
          "content": "<p>Yes, I am working on something at the moment that could be published. Tomorrow I will be taking a day off, so I am not sure how quickly it will be available here on kaggle. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2662707,
      "author_name": "Davut Polat",
      "author_url": "",
      "post_date": "2024-02-22T04:56:53.837000",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> it is not what i suggested.</p>\n<p>$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1} \\sqrt{\\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{rollingMax} - GiniScore_i)\\right)^2}<br>\n$$</p>\n<p>GiniScore_rollingMax : is the highest Gini score observed up to week i <br>\nGiniScore_i : is the Gini score for week i </p>\n<p>Can you update the return term with this : </p>\n<p>return np.mean(gini_in_time) - c/(len(gini_in_time)-1)*np.sqrt(cost)</p>\n<p>it is totally different what i suggested. Can you evaluate it again please?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2662984,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-22T09:06:48.673000",
          "content": "<p>Apologies, the optimized c is 1.82 and the score is 0.67.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2662586,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2024-02-22T02:26:01.610000",
      "content": "<p>I created a notebook to run some tests.<br>\n<a href=\"https://www.kaggle.com/carloshuertas/homecreditmetrictest\" target=\"_blank\">https://www.kaggle.com/carloshuertas/homecreditmetrictest</a></p>\n<p>I still cant understand the objective. </p>\n<blockquote>\n  <p>A model-X that outperforms week-over-week to model-Y over a long (100 weeks) time-horizon still is \"worse\"</p>\n</blockquote>\n<p>While in real-life, even with a slightly faster degradation this would lead to better decisions, as it outperforms the weaker model, every-single-time. I would understand that if eventually model-Y outperforms, Ok, but not the case in the experiment, yet, still the better model scores worse.</p>\n<p>Am I missing something? The math is not clicking for me, is this based on any peer-reviewed material? It doesnt feel ready.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2662693,
          "author_name": "Dmitriy Guller",
          "author_url": "",
          "post_date": "2024-02-22T04:42:12",
          "content": "<p>It's not clicking with me either.  I think one problem could be with the way that the organizers came up with the metric.  From what I understand, they presented various cases of Model A vs Model B, and asked the expert to express their preferences.  Then they formulated a metric that would fit to those preferences.</p>\n<p>Let's assume for a moment that the experts are perfectly infallible and have a perfectly calibrated sense of how costly the model instability is.  There is still a problem with that approach if you didn't go to the next step and try to be adversarial with the metric.  Once you come up with a metric, can you generate cases specifically to get situations where the metric leads to a clearly wrong conclusion?  Those cases might not come up in a business setting where common sense would put a stop to all nonsensical abuse, but a Kaggle competition where the value of the metric is the only thing that counts is a whole other ballgame.  To me a case where one model is strictly better than another for all weeks, and clearly has a better or equal asymptotic behavior, should never ever score worse.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2662838,
              "author_name": "loh-maa",
              "author_url": "",
              "post_date": "2024-02-22T06:39:47.053000",
              "content": "<p>I imagine the \"expert modelling\" approach could indeed take place and while it's not a wrong approach it may lead to significant perceptual distortions. In particular, there may be something like perceptual extrapolation, which is causing the disconnect between the business expectations and how the metric looks to kagglers. An example:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5871958%2F6867388250ec3d3fc3b6f06834fe9ec2%2FFigure_1.png?generation=1708583025156692&amp;alt=media\"></p>\n<p>When experts look at the chart they may decide the blue model performs better because it's \"flatter\" and clearly it looks like it's net gini scores are soon going to be better than the orange one. So this is reflected in the original metric and the score is correct according to the expectations. Apparently many kagglers would argue the orange one is better. Personally I like the liberal approach -- the metric has been given and is all we should be concerned to optimize. All hacks allowed. In the worst case it will be a lesson for the experts… ,)</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2662902,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-02-22T07:48:23.573000",
              "content": "<p>i think the main disconnect between the metric and what people think about when they say \"better models\" is this asymptotic behavior you refer to. the homecredit people, with their metric, basically say that the trend is a linear slope and this linear slope will continue forever, so the best model will always be one that has a non-negative slope.</p>\n<p>in most situations, however, there should be a flattening of this slope as time goes on. usually that would mean that, even when the better model has a sharper drop off, it will still consistently outperform the model with the flatter slope because the slope of the performance will flatten over time and not just keep dropping off linearly.</p>\n<p>however one of the many pitfalls here is that the hosts have said that data resources can become unreliable. what this means is that this asymptotic behavior you would usually expect, of the flattening of the performance, is not necessarily true here.</p>\n<p>there are actually many points to think about here, and the conclusion ive personally come to is that (without deep insights into their business) their metric, with the weights theyve chosen, because they never have a reason to artificially make a model worse, is probably perfectly fine for their business overall, and it's just unfortunate for them that it turns out that in a competition setting it is exploitable.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2675116,
              "author_name": "Lucas Morin",
              "author_url": "",
              "post_date": "2024-02-29T18:04:32.190000",
              "content": "<p>I think you are right. Overall it depends on the asymptotic behavior (wether the blue line cross the orange line in the end). From my experience on longer time frame I didn't encouter that much drop and \"crossing\". </p>\n<p>I don't entirely know what to think here… A model loosing 10 AUC points in the course of 6 months is indeed quite bad. But I don't entirely know what is at fault here (blattant overfitting ? volatililty of risk drivers ?).</p>\n<p>Maybe a solution would be to use a good baseline, typically test wether its above and stay above a 'fico' score (or a linear model recalibrated each week). </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2662456,
      "author_name": "Oleksiy Kononenko",
      "author_url": "",
      "post_date": "2024-02-21T22:25:25.053000",
      "content": "<p>What about just a simple </p>\n<pre><code>def metric_mean_std(gini_in_time, k=):\n     .(gini_in_time) - k*.(gini_in_time)\n</code></pre>\n<p>with whatever <code>k</code> you deem necessary to ensure proper stability.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2662464,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-21T22:42:00.797000",
          "content": "<p>That gives me score 0.64.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2662704,
              "author_name": "Oleksiy Kononenko",
              "author_url": "",
              "post_date": "2024-02-22T04:55:41.960000",
              "content": "<p>How this score is calculated and for what <code>k</code>? My feeling is that the score itself can’t be used to rank metric robustness, as we know the current #1 metric has other serious issues. In my opinion, being simple and closer to standard could result in a better generalization and a lower hackability.</p>\n<p>Btw, in the post you say the score is <code>0.74</code>, not <code>0.64</code>.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2662956,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T08:47:14.527000",
              "content": "<p>I can't say exactly how the score is calculated I am afraid, at least not during the competition. Think of them as checks and you get a percentage of how many checks you pass. Those checks were given by in-house risk managers. Yes, indeed it can't be used to score the metric robustness. If anyone happens to find a metric that is around 0.90 I think we would definitely consider it as a suitable replacement for the current candidate. </p>\n<p>There is 0.74, because I optimized the k to 0.9.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2673245,
      "author_name": "Davut Polat",
      "author_url": "",
      "post_date": "2024-02-28T15:29:00.520000",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> could you test this metric function?</p>\n<p>c and rolling_per should be optimized.<br>\nconstraints: <br>\nc &gt;  0<br>\n2 &lt;= rolling_per &lt;= len(gini_scores)</p>\n<pre><code> stability_score(gini_scores,rolling_per=,c=.):\n     = \n\n     i in range(rolling_per-,len(gini_scores)):\n         = penalty + (max(gini_scores[(i-rolling_per+):i+])-gini_scores[i])**\n\n     np.mean(gini_scores) - c*(np.sqrt(penalty)/(len(gini_scores)-rolling_per+))\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2673405,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-28T16:54:27.743000",
          "content": "<p>Without optimization it is 0.61</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2673642,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-28T19:40:18.490000",
              "content": "<p>could you optimize the period and c ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2673460,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-28T17:40:49.983000",
          "content": "<p>I like it! It gives good score, could be still simple enough to explain and I don't see how it could be hacked. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2673590,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-28T19:05:35.057000",
              "content": "<p>Can the exponent in the formula for penalty possibly optimized to give even better score?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2673895,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-28T23:06:24.727000",
              "content": "<p>I could do it, but somehow I think the L2 penalty will be close to optimal </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2674630,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-29T13:07:58.763000",
              "content": "<p>Would like to see how L1 performs (more robust wrt outlier)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2676286,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-01T11:50:47.617000",
              "content": "<p>L1 gives the same score, just with different params.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2674248,
          "author_name": "Davut Polat",
          "author_url": "",
          "post_date": "2024-02-29T07:37:01.793000",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>  can you also test this one? I feel like this one is also unhackable and it should give good score.<br>\nthe idea is that metric focuses on stability over small segments of timeline. when  someone worsen the gini performance over some time segment, metric will be ruined due to mean(gini)+reward part. </p>\n<p>c1,c2, rolling_per1, and rolling_per2 should be optimized.<br>\nconstraints:<br>\nc1 &gt; 0, c2&gt;0</p>\n<p>2 &lt;= rolling_per1 &lt;= len(gini_scores)<br>\n2 &lt;= rolling_per2 &lt;= len(gini_scores)</p>\n<p>for simplicity you can try <br>\nc1=c2 and rolling_per1=rolling_per2</p>\n<pre><code> stability_score_2(gini_scores,rolling_per1=,rolling_per2=,c1=.,c2=.):\n     = \n      = \n\n     i in range(rolling_per1-,len(gini_scores)):\n         = penalty + (max(gini_scores[(i-rolling_per1+):i+])-gini_scores[i])**\n\n     i in range(rolling_per2-,len(gini_scores)):\n         = reward + (gini_scores[i]-min(gini_scores[(i-rolling_per2+):i+]))**\n\n     np.mean(gini_scores) - c1*(np.sqrt(penalty)/(len(gini_scores)-rolling_per1+)) + c2*(np.sqrt(reward)/(len(gini_scores)-rolling_per2+))\n</code></pre>\n<p>i am aware of that the code is a little dirty, after optimizing the params it can be simplified. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2674643,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-29T13:12:32.380000",
              "content": "<p>Seems like that gives a score of 0.86 </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2674645,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-29T13:15:07.163000",
              "content": "<p>did you optimize parameters ? and how was the score for previous one ? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2674870,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-29T15:15:11.300000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nwhat is the len(gini_scores) in  your test code? </p>\n<p>if it is near 94, then rolling_per1_2=94 means nothing.  the formula omits from 0 to rolling_per1_2 for calculating <br>\nreward/penalty part. it focuses on indice range  between rolling_per1_2 and -1(end)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2676285,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-01T11:50:26.420000",
              "content": "<p><a href=\"https://www.kaggle.com/davutpolat\" target=\"_blank\">@davutpolat</a> I understand and it is close to that number. Yes, I optimized the params.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2673048,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2024-02-28T13:07:06.440000",
      "content": "<p>what about some 'sharpe' ?</p>\n<pre><code> ():\n     np.mean(gini_in_time) / np.std(gini_in_time)\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2673119,
          "author_name": "Davut Polat",
          "author_url": "",
          "post_date": "2024-02-28T13:46:05.593000",
          "content": "<p>gini series1 : [0.7, 0.6, 0.7, 0.6 , 0.7, 0.6 ]<br>\nsharpe1  = 13</p>\n<p>gini series2 : [0.099, 0.1, 0.099, 0.1 , 0.099, 0.1 ]<br>\nsharpe2 : 199</p>\n<p>which one would you prefer  ? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2673152,
              "author_name": "Lucas Morin",
              "author_url": "",
              "post_date": "2024-02-28T14:08:29.357000",
              "content": "<p>I know that weird use case. But is it actually possible to reduce std that way ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2673223,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-28T15:11:45.823000",
              "content": "<p>it also does not handle gini performance decrease in time. By stability, it is meant that the model should have   high mean gini score and non decreasing performance in time  as much as possible.  Basic Sharpe  ratio does not have such time variant term.  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2673503,
              "author_name": "Lucas Morin",
              "author_url": "",
              "post_date": "2024-02-28T18:28:06.033000",
              "content": "<p>It penalizes variation, both upward and downward. Not sure it is a good idea to introduce an asymmetric sign. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2674579,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-29T12:28:02.600000",
          "content": "<p>That gives 0.52</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2667785,
      "author_name": "kurupical",
      "author_url": "",
      "post_date": "2024-02-25T10:55:40.013000",
      "content": "<p>I think if the sum and standard deviation of scores and are the same, the model would consider the following order preferable in a business context:</p>\n<ol>\n<li>Normal distribution, like <code>[0.61, 0.59, 0.63, 0.59, 0.58]</code></li>\n<li>Monotonically increasing or monotonically decreasing, like <code>[0.62, 0.61, 0.60, 0.59, 0.58]</code></li>\n<li>Domain shift, like <code>[0.7, 0.7, 0.7, 0.7, 0.2]</code></li>\n</ol>\n<p>I have created a notebook to check it with some test cases.<br>\n<a href=\"https://www.kaggle.com/code/kurupical/testcase-of-new-metric-for-business-applications/notebook?scriptVersionId=164230696\" target=\"_blank\">https://www.kaggle.com/code/kurupical/testcase-of-new-metric-for-business-applications/notebook?scriptVersionId=164230696</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2667864,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-25T11:58:33.247000",
          "content": "<p>Thanks for the work. Kudos for the useful notebook. I think what should be also done is to start with declining performance (setting <code>a</code> and <code>std(residues)</code>) and then sample it let's say 10**7 times and see how each hack can improve the score in q=0.99,0.999 compared to stability metric. This could prove the robustness for some of them. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2806853,
      "author_name": "mh",
      "author_url": "",
      "post_date": "2024-05-11T11:03:41.557000",
      "content": "<p>Great job and nostalgic…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2682520,
      "author_name": "hanjinxu2000",
      "author_url": "",
      "post_date": "2024-03-05T11:24:09.943000",
      "content": "<p>This is so great, I'm moved to express my feelings</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2667232,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-02-25T01:23:47.033000",
      "content": "<p>I don't know much about this competition, including this field. If there are any mistakes, please feel free to point them out</p>\n<p>I think we can multiply the gini coefficient for each week by a different weight, with the weight increasing later on. For example, if we have four weeks of gini coefficient, we can multiply them by 1/10, 2/10, 3/10, 4/10 respectively,10=1+2+3+4.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2665972,
      "author_name": "暗黑AGI",
      "author_url": "",
      "post_date": "2024-02-24T03:13:59.633000",
      "content": "<p>I have a question: If the week_num field in the test data is removed, how will gini_in_time be calculated?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2666310,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-24T09:57:58.397000",
          "content": "<p>It will be only removed in the data that are available to kagglers when running the submission. The final decision is still not made.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2665549,
      "author_name": "N. Kai",
      "author_url": "",
      "post_date": "2024-02-23T18:02:11.073000",
      "content": "<p>Based on the idea of calculating the slope in various parts by <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>, I have created the following metric:</p>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    \n    x_past = np.arange(-win, )\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n\n    \n    a1_lst = []\n     start_index  (, (x)-win+):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, )\n        a1_lst.append((,a1))       \n    a1_avg = np.mean(a1_lst)\n\n     avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n</code></pre>\n<p>The points are as follows:</p>\n<ul>\n<li>setting the hop length to 1 instead of matching it to the frame length in order to avoid hacking</li>\n<li>taking measures in order to avoid hacking in the first segment</li>\n</ul>\n<p>As far as I have tried, I could not increase the score by hacking. <br>\nHere is the experimental notebook:<br>\n<a href=\"https://www.kaggle.com/code/naotokai/gini-stability-slide-window\" target=\"_blank\">https://www.kaggle.com/code/naotokai/gini-stability-slide-window</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2671488,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-27T14:33:16.277000",
          "content": "<p>That gives score of 0.49</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2665413,
      "author_name": "Viktoria",
      "author_url": "",
      "post_date": "2024-02-23T16:39:32.910000",
      "content": "<p>Wouldn't taking the weighted average over time take care of the hacking problem? While not a perfect metric  in terms of capturing sharp drops in Gini on occasional weeks or skewness of the distribution, it would capture the over-time model deterioration by giving higher weights to further away data points and prioritises the absolute Gini value which prevents the hack. Something like below in simplest version, but how the weights are allocated can be modified.</p>\n<pre><code> ():\n    x = np.arange((gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/(x)\n     w_avg_gini + w_gstd * gini_std\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2674583,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-29T12:33:44.020000",
          "content": "<p>That gives the score 0.77, but I have to say it is a bit misaligned with the intention of the initial metric. At least with such a steep weights. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2663212,
      "author_name": "ARMADA",
      "author_url": "",
      "post_date": "2024-02-22T11:21:30.100000",
      "content": "<p>In the current stability metric<br>\nStability = mean(gini) + 88<em>min(0,a) -0.5</em>std(residuals)</p>\n<p>I suggest to change the first term to be:<br>\nStability = <strong>mean(gini1,gini2,…giniN)</strong> + 88 * min(0,a) - 0.5 * std(residuals)</p>\n<p>Where N is a parameter, the mean is for first N weeks rather than all weeks. The quality of predictions of the rest of the weeks will be taken care of by the other two terms.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2663005,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-22T09:19:36.890000",
      "content": "<p>I like the idea, it may be an interesting approach to calculate slopes over various parts of <code>gini_in_time</code> to avoid hacks. <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2663083,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-22T09:48:37.370000",
          "content": "<p>the formula i posted was nonsense, but yeah that was the idea i dont know if it would fix anything though 😅 i gotta go for a little while now too</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2663147,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-22T10:46:34.700000",
          "content": "<p>something like this maybe could work (at least i couldnt hack it in my attempts right now and it still has a big emphasis on the stability) i really gotta go now though</p>\n<pre><code>f=\n ():\n    gini_in_time = base.loc[:, [, , ]]\\\n        .sort_values()\\\n        .groupby()[[, ]]\\\n        .apply( x: *roc_auc_score(x[], x[])-).tolist()\n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far=a\n     s  (,f):\n        start_index = (x) // f * (s)\n        end_index = (x) // f * (s+)\n        x_second_fifth = x[start_index:end_index]\n        y_second_fifth = gini_in_time[start_index:end_index]\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        far+=(,a1)\n\n\n     avg_gini + w_fallingrate * (far) + w_resstd * res_std \n</code></pre>",
          "votes": 1,
          "replies": [
            {
              "id": 2663196,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T11:06:41",
              "content": "<p>That gives a score of 0.62. I will try to think of a ways how to improve it. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2665131,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-02-23T12:44:13.473000",
              "content": "<p>hmm yeah, there is a tradeoff between how much it values stability versus how hackable it is - for the competition probably f=7 is the lowest you can go,</p>\n<p>also, what this is doing is a bit different than the original metric, because it is not only concerned about the overall trend but also the trend of smaller non-overlapping time periods. this is, of course, what makes it not hackable because you can only hack one period (or with this implementation at most two periods) at a time and the weight of the period is small enough that you hurt your overall score while doing it.</p>\n<p>what i noticed in your simulated data, in the other thread, is that it is very swingy in these smaller time periods contrary to my testing is using models trained on the actual training data, which is much more closer aligned with the overall trend over the entire time period…</p>\n<p>i think this averaging might work well for real data because these non-overlapping time periods are more or less the same as the overall time period in terms of the size of stability. in your simulated data, however, there are periods where it suddenly jumps all over the place giving a disproportional amount of weight to this time period for the overall calculation.</p>\n<p>(also, i made a hasty mistake: far=a in the metric is supposed to be far=min(0,a) but it doesnt really matter)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2665161,
              "author_name": "N. Kai",
              "author_url": "",
              "post_date": "2024-02-23T13:23:29.680000",
              "content": "<p>Please note that it can also be hacked by reducing the gini score at the beginning of each segment:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4416902%2Fa50b6a14215711ebeca4e1cdd03fc998%2Foutput.png?generation=1708694385542872&amp;alt=media\"></p>\n<p>It would be better to set the hop length to 1 instead of matching it to the frame length, to avoid hacking. For example, like this:</p>\n<pre><code> ():\n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    a1_lst = []\n     start_index  (, (x)-win):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        a1_lst.append((,a1))\n    a = np.mean(a1_lst)\n\n     avg_gini + w_fallingrate * a + w_resstd * res_std \n</code></pre>\n<p>Here is the experiment notebook.<br>\n<a href=\"https://www.kaggle.com/naotokai/metic-experiments\" target=\"_blank\">https://www.kaggle.com/naotokai/metic-experiments</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2665292,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-02-23T15:20:35.977000",
              "content": "<p>i just noticed myself that with some good accuracy i can still hack things by like 0.02 for realistic scenarios. i just saw ur comment, and honestly, i am not even gonna bother anymore for now and just give up, i dont think it's possible to fully fix.</p>\n<pre><code>f=\n ():\n    gini_in_time = base.loc[:, [, , ]]\\\n        .sort_values()\\\n        .groupby()[[, ]]\\\n        .apply( x: *roc_auc_score(x[], x[])-).tolist()\n\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far=\n\n     s  (,f):\n        start_index = (x) // f * (s)\n        end_index = (x) // f * (s+)\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, )\n        far += (,a1)\n     avg_gini + w_fallingrate * (far) + w_resstd * res_std \n</code></pre>\n<p>this is the best i could come up with in the end for realistic scenarios, but it's still possible to hack with good accuracy.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2664871,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-23T08:58:29.150000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2662979,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-22T09:00:47.777000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2662336": "Here you can ask me to score any metric you would like to suggest as new stability metric. The score itself does not guarantee that we will use the metric in this competition. The score of a metric is just an approximation of how we rate the metrics, perhaps during the following dates I will improve this method. Think of the score as a percentage of how much it is aligned with an internal business view. \n\n## Rules\n1. If there are any hyperparameters I will optimize them with given boundaries ([-inf, inf] is not a valid interval). \n2. I won't score a metric with more than three hyperparameters that are not set beforehand. \n3. The metric should be relatively easy to explain and should have a connection to the current stability metric, in other words, we would like to see a formula that is easy to read (interpretability) and it is not connected directly to the competition data (generality). \n4. Please send me a simple notebook or easy-to-read code, ideally with sensible variable naming. I won't be implementing metric if you send me just an equation. I will implement it from an equation only if I have spare time, so no guarantee. \n\nScore range is 0-1, where higher is better. \n\n### Current stability metric\n```python\ndef gini_stability(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5):\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    return avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n```\nScore: 0.99\n\n### Metric mean + rolling min/max diff penalty by @davutpolat\n```python\ndef stability_score_2(gini_scores, rolling_per1=94, rolling_per2=94, c1=3.1, c2=3.1):\n    penalty = 0\n    reward = 0\n\n    for i in range(rolling_per1 - 1, len(gini_scores)):\n        penalty = (\n            penalty\n            + (max(gini_scores[(i - rolling_per1 + 1) : i + 1]) - gini_scores[i]) ** 2\n        )\n\n    for i in range(rolling_per2 - 1, len(gini_scores)):\n        reward = (\n            reward\n            + (gini_scores[i] - min(gini_scores[(i - rolling_per2 + 1) : i + 1])) ** 2\n        )\n\n    return (\n        np.mean(gini_scores)\n        - c1 * (np.sqrt(penalty) / (len(gini_scores) - rolling_per1 + 1))\n        + c2 * (np.sqrt(reward) / (len(gini_scores) - rolling_per2 + 1))\n    )\n\n```\nScore 0.86\n\n### Metric mean + rolling max diff penalty by @davutpolat\n```python\ndef metric_davutpolat(gini_scores, rolling_per=92, c=1.9):\n    penalty = 0\n\n    for i in range(rolling_per-1,len(gini_scores)):\n        penalty = penalty + (max(gini_scores[(i-rolling_per+1):i+1])-gini_scores[i])**2\n\n    return np.mean(gini_scores) - c*(np.sqrt(penalty)/(len(gini_scores)-rolling_per+1)) \n```\nScore 0.85\n\n### Metric mean log MA by @jacobyjaeger \n```python\ndef metric_jacobyjaeger(gini_in_time, exponent=69, ma_len=4):\n    x = np.cumsum(gini_in_time, axis=0)\n    x = np.concatenate([0*x[:1], x], 0)\n    scores = -np.mean(-np.log(np.maximum((x[ma_len:] - x[:-ma_len])/ma_len, 1e-5))**exponent)\n    return scores \n```\nScore: 0.83\n\n### Metric with weights by @viktorian\n```python\ndef gini_stability_avg(gini_in_time, w_gstd=-0.5):\n    x = np.arange(len(gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/sum(x)\n    return w_avg_gini + w_gstd * gini_std\n```\n\n### Metric mean std by @seifachour12 \n```python\ndef metric_seifachour12(gini_in_time, alpha=1028, beta=1.97, gamma=0.49):\n    return (np.abs(beta*(1/beta*np.mean(gini_in_time) + (gamma-np.std(gini_in_time)))-1))**(1/alpha) \n```\nScore: 0.76\n\n### Metric mean std by @kononenko \n```python\ndef metric_kononenko(gini_in_time, c=0.90):\n    return np.mean(gini_in_time) - c*np.std(gini_in_time)\n```\nScore: 0.74\n\n### Metric rolling gini max by @davutpolat \n```python\ndef metric_davutpolat(gini_in_time, c=1.82):\n    max_gini = gini_in_time[0]\n    cost = 0\n    for i in range(1, len(gini_in_time)):\n        max_gini = max(max_gini, gini_in_time[i])\n        cost += max(0, max_gini - gini_in_time[i])**2\n    return np.mean(gini_in_time) + c/(len(gini_in_time)-1)*cost   \n```\nScore: 0.67\n\n## Metric stability with averaging by @at7459 \n```python\ndef gini_at7459(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5, f=8):\n    w_fallingrate /= f + 1\n\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    far = a\n    for s in range(1,f):\n        start_index = len(x) // f * (s)\n        end_index = len(x) // f * (s+1)\n        x_second_fifth = x[start_index:end_index]\n        y_second_fifth = gini_in_time[start_index:end_index]\n        x1 = x[start_index:end_index]\n        y1 = gini_in_time[start_index:end_index]\n        a1, b1 = np.polyfit(x1, y1, 1)\n        far += min(0,a1)\n\n    return avg_gini + w_fallingrate * (far) + w_resstd * res_std \n```\n\n## Metric mean gini\n```python\ndef metric_mean(gini_in_time):\n    return np.mean(gini_in_time)  \n```\nScore: 0.59\n\n### Metric sharpe-like by @lucasmorin\n```python\ndef metric_sharpe(gini_in_time):\n    return np.mean(gini_in_time) / np.std(gini_in_time)\n```\nScore 0.52\n\n## Stability with sliding window by @naotokai\n```python\ndef gini_stability_slide_window(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5, win=10):\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n\n    # The countermeasures for hacking in the first segment\n    x_past = np.arange(1-win, 0)\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n\n    # linear regression in each segment\n    a1_lst = []\n    for start_index in range(0, len(x)-win+1):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, 1)\n        a1_lst.append(min(0,a1))       \n    a1_avg = np.mean(a1_lst)\n\n    return avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n```\nScore: 0.49\n\nI think the target is at least 0.90 for us to consider this metric as a substitution for the current stability metric. Any effort will be appreciated, but again this is outside the scope of the competition LB and definitely not mandatory, completely up to you. ",
    "2662884": "Under what kind of evaluation does the current metric get a perfect score? Are you just taking the correlation of scores under each metric to the scores under the original?\n\nI do not believe the current metric deserves a perfect score under any reasonable evaluation. It has numerous problems even putting aside the hackablity. \n\nFirst of all taking statistics over raw Gini scores is unlikely to be ideal. Gini scores, being bounded by zero and one, cannot take on extreme values even for rounds that see extremely poor performance. This leads to a tendency to under-emphasize bad rounds which is really the opposite of what you want if you value stability. -log(gini) would be a better starting point.\n\nNow, you might expect the standard deviation term to capture the risk, but it really can do only a mediocre job of that. Again, when you use gini scores, no deviation can be greater than one away from the mean so exceptionally bad rounds still can only have limited impact. Standard deviation, being only one moment of a distribution, cannot tell us anything about the other moments such as skew and kurtosis, but these moments should also be important to a good risk measure as they have an impact on the frequency of very bad rounds. It's better to punish bad rounds, or periods of bad rounds, directly with a metric that becomes exponentially worse as the round does. Unlike standard deviation, this does not punish good round and hence isn't hackable.\n\nAnd then the regression slope. Raw gini scores are again a bad choice here because a bounded quantity cannot maintain a stable slope. arctanh(2*gini-1) would be much nicer for fitting lines to since linear behavior is actually possible in that domain. As it is we can expect scores mostly near 1 and scores mostly near 0 to see much smaller slopes than scores mostly near .5 due to the bounded nature of gini scores. \n\nWhile the slope can tell us if scores have trended downwards, it cannot necessarily tell us if scores have a tendency to drift over time. For example if scores happen to follow a 'U' shaped pattern over the test period, this results in zero slope but it does demonstrate a pattern of sustained poor scores, a kind of risk we should take account of. My moving average based metric captures this kind of risk and without the hackability.  \n  \nUsing the slope is a bit like trying to extrapolate the performance on future rounds by negatively weighting early rounds. But if any rounds have negative weight, our incentive is of course to make those rounds as bad as possible. And even if you make that hard to do, extrapolation is inherently very sensitive to noise and will make the competition much more random and the possibility that some features may disappear in the future makes the competition even more of a gamble. \n",
    "2665125": "@jetakow @tomasjeline2 \n\nFrom reading the discussion on all pinned threads, it's not clear what the real problem this thread is trying to solve.\nIf, as you say [here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2663049), [0.3, 0.3, 0.3] is preferable to [0.6, 0.5, 0.4] then artificially reducing gini becomes merely a form of **legitimate hedging or smoothing** to compensate for overconfidence/overfitting the model in \"better\" weeks, and in my mind no more of a \"hack\" than say artificially clipping predictions to [1e-10, 1 - 1e-10] in a log-loss competition.\n\nOf course I may be missing something and would love to learn where I went wrong.",
    "2663125": "Firstly, I want to express gratitude for launching this competition, which addresses a crucial topic: the stability of models in time series analysis.\n\nAs you pointed out, the previous metric exhibited an undesirable trait where penalizing the model's early predictions actually increased the score, contradicting intuitive expectations. This issue primarily stems from the necessity of ensuring the stability of the area under the curve (AUC) over time, attempted through the min(0, a) component.\n\nIt is paramount that we achieve stability as time progresses. However, what we truly require is stability across all time points, not solely in future predictions. Hence, the metric I propose is rooted in the notion of maximizing both the mean of the Gini coefficient and its standard deviation over time. By doing so, we aim to ensure that the Gini coefficient remains sufficiently high and stable across all time points.\n\n` Below is the proposed metric:\n\n    def mean_std_gini(gini_in_time, alpha=0.01):\n        return (np.abs(2*(0.5*np.mean(gini_in_time) + (0.5-np.std(gini_in_time)))-1))**alpha`\n\nIf gini_in_time is perfect (all ones) the metric returns 1 and if it is the worst (all zeros) the metric is 0.\nHopefully this metric allows us to achieve the competition's goal.",
    "2663477": "@jetakow is it possible that you can provide some gini scores (at least 10 gini series with size of  number of weeks in test set )  with their preference order, so we can test our ideas and apply on them to see if they holds HC concerns?  it would be more practical for everyone to see how  their suggestions perform :) ",
    "2662707": "@jetakow it is not what i suggested.\n\n$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1} \\sqrt{\\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{rollingMax} - GiniScore_i)\\right)^2}\n$$\n\nGiniScore_rollingMax : is the highest Gini score observed up to week i \nGiniScore_i : is the Gini score for week i \n\n\nCan you update the return term with this : \n\n\n return np.mean(gini_in_time) - c/(len(gini_in_time)-1)*np.sqrt(cost)\n\n\n\nit is totally different what i suggested. Can you evaluate it again please?\n",
    "2662586": "I created a notebook to run some tests.\nhttps://www.kaggle.com/carloshuertas/homecreditmetrictest\n\nI still cant understand the objective. \n\n>A model-X that outperforms week-over-week to model-Y over a long (100 weeks) time-horizon still is \"worse\"\n\n While in real-life, even with a slightly faster degradation this would lead to better decisions, as it outperforms the weaker model, every-single-time. I would understand that if eventually model-Y outperforms, Ok, but not the case in the experiment, yet, still the better model scores worse.\n\nAm I missing something? The math is not clicking for me, is this based on any peer-reviewed material? It doesnt feel ready.",
    "2662456": "What about just a simple \n\n```\ndef metric_mean_std(gini_in_time, k=0.5):\n    return np.mean(gini_in_time) - k*np.std(gini_in_time)\n```\n\nwith whatever `k` you deem necessary to ensure proper stability.",
    "2673245": "@jetakow could you test this metric function?\n\nc and rolling_per should be optimized.\nconstraints: \nc >  0\n2 <= rolling_per <= len(gini_scores)\n\n\n```\ndef stability_score(gini_scores,rolling_per=5,c=1.0):\n    penalty = 0\n    \n    for i in range(rolling_per-1,len(gini_scores)):\n        penalty = penalty + (max(gini_scores[(i-rolling_per+1):i+1])-gini_scores[i])**2\n\n    return np.mean(gini_scores) - c*(np.sqrt(penalty)/(len(gini_scores)-rolling_per+1))\n```\n\n",
    "2673048": "what about some 'sharpe' ?\n\n```python\ndef metric_sharpe(gini_in_time):\n    return np.mean(gini_in_time) / np.std(gini_in_time)\n```",
    "2667785": "I think if the sum and standard deviation of scores and are the same, the model would consider the following order preferable in a business context:\n1. Normal distribution, like ``[0.61, 0.59, 0.63, 0.59, 0.58]``\n2. Monotonically increasing or monotonically decreasing, like ``[0.62, 0.61, 0.60, 0.59, 0.58]``\n3. Domain shift, like `` [0.7, 0.7, 0.7, 0.7, 0.2]``\n\nI have created a notebook to check it with some test cases.\nhttps://www.kaggle.com/code/kurupical/testcase-of-new-metric-for-business-applications/notebook?scriptVersionId=164230696\n",
    "2806853": "Great job and nostalgic...",
    "2682520": "This is so great, I'm moved to express my feelings",
    "2667232": "I don't know much about this competition, including this field. If there are any mistakes, please feel free to point them out\n\n\nI think we can multiply the gini coefficient for each week by a different weight, with the weight increasing later on. For example, if we have four weeks of gini coefficient, we can multiply them by 1/10, 2/10, 3/10, 4/10 respectively,10=1+2+3+4.",
    "2665972": "I have a question: If the week_num field in the test data is removed, how will gini_in_time be calculated?",
    "2665549": "Based on the idea of calculating the slope in various parts by @at7459, I have created the following metric:\n```python\ndef gini_stability_slide_window(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5, win=10):\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    \n    # The countermeasures for hacking in the first segment\n    x_past = np.arange(1-win, 0)\n    x = np.hstack([x_past, x])\n    y = np.hstack([np.ones_like(x_past), y])\n    \n    # linear regression in each segment\n    a1_lst = []\n    for start_index in range(0, len(x)-win+1):\n        end_index = start_index + win\n        x1 = x[start_index:end_index]\n        y1 = y[start_index:end_index]\n        a1, _ = np.polyfit(x1, y1, 1)\n        a1_lst.append(min(0,a1))       \n    a1_avg = np.mean(a1_lst)\n\n    return avg_gini + w_fallingrate * a1_avg + w_resstd * res_std \n```\nThe points are as follows:\n- setting the hop length to 1 instead of matching it to the frame length in order to avoid hacking\n- taking measures in order to avoid hacking in the first segment\n\nAs far as I have tried, I could not increase the score by hacking. \nHere is the experimental notebook:\nhttps://www.kaggle.com/code/naotokai/gini-stability-slide-window",
    "2665413": "Wouldn't taking the weighted average over time take care of the hacking problem? While not a perfect metric  in terms of capturing sharp drops in Gini on occasional weeks or skewness of the distribution, it would capture the over-time model deterioration by giving higher weights to further away data points and prioritises the absolute Gini value which prevents the hack. Something like below in simplest version, but how the weights are allocated can be modified.\n\n```python\ndef gini_stability_avg(gini_in_time, w_gstd=-0.5):\n    x = np.arange(len(gini_in_time))\n    gini_std = np.std(gini_in_time)\n    w_avg_gini = np.matmul(gini_in_time, x)/sum(x)\n    return w_avg_gini + w_gstd * gini_std\n```",
    "2663212": "In the current stability metric\nStability = mean(gini) + 88*min(0,a) -0.5*std(residuals)\n\nI suggest to change the first term to be:\nStability = **mean(gini1,gini2,…giniN)** + 88 * min(0,a) - 0.5 * std(residuals)\n\nWhere N is a parameter, the mean is for first N weeks rather than all weeks. The quality of predictions of the rest of the weeks will be taken care of by the other two terms.\n",
    "2663005": "I like the idea, it may be an interesting approach to calculate slopes over various parts of `gini_in_time` to avoid hacks. @at7459 ",
    "2664871": "",
    "2662979": ""
  }
}