{
  "id": 501577,
  "title": "Implement competition metric directly into the model",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501577",
  "author_name": "",
  "post_date": "2024-05-09T19:30:29.422881300Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Based on <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/500868\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/500868</a></p>\n<p>Slightly modified as I don't like lambda functions, so it should work quicker </p>\n<p>Documentation says:<br>\nCustom eval function expects a callable with following signatures: func(y_true, y_pred), func(y_true, y_pred, weight) or func(y_true, y_pred, weight, group) and returns (eval_name, eval_result, is_higher_better) or list of (eval_name, eval_result, is_higher_better)</p>\n<pre><code> ():\n    \n    gini_in_time = []\n    weeks_to_score = X_val[].reset_index(drop=)\n\n     week  weeks_to_score.unique():\n        week_idx = weeks_to_score.eq(week)\n        gini = np.array( * roc_auc_score(y_true[week_idx], y_pred[week_idx]) - )\n        gini_in_time.append(gini)\n\n    w_fallingrate = \n    w_resstd = -\n    x = np.arange((gini_in_time))\n    y = np.array(gini_in_time)\n    a, b = np.polyfit(x, y, )\n    y_hat = a * x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(y)\n    stability_score = avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n    is_higher_better = \n\n     , stability_score, is_higher_better\n\nparams = {: }\nmodel = LGBMClassifier(**params)\nmodel.fit(X_train, y_train,\n         eval_set=[(X_val, y_val)],\n         eval_metric=stability_metric,\n         callbacks=[early_stopping()],\n        )\n</code></pre>",
  "messages": [
    {
      "id": "2804089",
      "postDate": "05/09/2024 19:30:29",
      "content": "<p>Based on <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/500868\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/500868</a></p>\n<p>Slightly modified as I don't like lambda functions, so it should work quicker </p>\n<p>Documentation says:<br>\nCustom eval function expects a callable with following signatures: func(y_true, y_pred), func(y_true, y_pred, weight) or func(y_true, y_pred, weight, group) and returns (eval_name, eval_result, is_higher_better) or list of (eval_name, eval_result, is_higher_better)</p>\n<pre><code> ():\n    \n    gini_in_time = []\n    weeks_to_score = X_val[].reset_index(drop=)\n\n     week  weeks_to_score.unique():\n        week_idx = weeks_to_score.eq(week)\n        gini = np.array( * roc_auc_score(y_true[week_idx], y_pred[week_idx]) - )\n        gini_in_time.append(gini)\n\n    w_fallingrate = \n    w_resstd = -\n    x = np.arange((gini_in_time))\n    y = np.array(gini_in_time)\n    a, b = np.polyfit(x, y, )\n    y_hat = a * x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(y)\n    stability_score = avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n    is_higher_better = \n\n     , stability_score, is_higher_better\n\nparams = {: }\nmodel = LGBMClassifier(**params)\nmodel.fit(X_train, y_train,\n         eval_set=[(X_val, y_val)],\n         eval_metric=stability_metric,\n         callbacks=[early_stopping()],\n        )\n</code></pre>",
      "rawMarkdown": "Based on https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/500868\n\nSlightly modified as I don't like lambda functions, so it should work quicker \n\nDocumentation says:\nCustom eval function expects a callable with following signatures: func(y_true, y_pred), func(y_true, y_pred, weight) or func(y_true, y_pred, weight, group) and returns (eval_name, eval_result, is_higher_better) or list of (eval_name, eval_result, is_higher_better)\n\n```python\ndef stability_metric(y_true, y_pred):\n    \"\"\"\n    Competition metric for model optimization during training\n    \"\"\"\n    gini_in_time = []\n    weeks_to_score = X_val['WEEK_NUM'].reset_index(drop=True)\n    \n    for week in weeks_to_score.unique():\n        week_idx = weeks_to_score.eq(week)\n        gini = np.array(2 * roc_auc_score(y_true[week_idx], y_pred[week_idx]) - 1)\n        gini_in_time.append(gini)\n\n    w_fallingrate = 88.0\n    w_resstd = -0.5\n    x = np.arange(len(gini_in_time))\n    y = np.array(gini_in_time)\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a * x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(y)\n    stability_score = avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n    is_higher_better = True\n\n    return 'stability_score', stability_score, is_higher_better\n\nparams = {'metric': 'None'}\nmodel = LGBMClassifier(**params)\nmodel.fit(X_train, y_train,\n         eval_set=[(X_val, y_val)],\n         eval_metric=stability_metric,\n         callbacks=[early_stopping(50)],\n        )\n```",
      "votes": null
    },
    {
      "id": "2804622",
      "postDate": "05/10/2024 06:07:29",
      "content": "<p>could you plz also write a catboost version metric ,i tried ,but the index was unmatched.i dont know how to fix it</p>",
      "rawMarkdown": "could you plz also write a catboost version metric ,i tried ,but the index was unmatched.i dont know how to fix it",
      "votes": null
    },
    {
      "id": "2804644",
      "postDate": "05/10/2024 06:20:55",
      "content": "<p>CatBoost doesn't support even simple AUC if you train on GPU, so I doubt it will work with a custom metric, but on CPU it doesn't make much sens</p>",
      "rawMarkdown": "CatBoost doesn't support even simple AUC if you train on GPU, so I doubt it will work with a custom metric, but on CPU it doesn't make much sens",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2804622,
      "author_name": "fanhuaiyuan",
      "author_url": "",
      "post_date": "05/10/2024 06:07:29",
      "content": "<p>could you plz also write a catboost version metric ,i tried ,but the index was unmatched.i dont know how to fix it</p>",
      "votes": null,
      "replies": [
        {
          "id": 2804644,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "05/10/2024 06:20:55",
          "content": "<p>CatBoost doesn't support even simple AUC if you train on GPU, so I doubt it will work with a custom metric, but on CPU it doesn't make much sens</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2804089": "Based on https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/500868\n\nSlightly modified as I don't like lambda functions, so it should work quicker \n\nDocumentation says:\nCustom eval function expects a callable with following signatures: func(y_true, y_pred), func(y_true, y_pred, weight) or func(y_true, y_pred, weight, group) and returns (eval_name, eval_result, is_higher_better) or list of (eval_name, eval_result, is_higher_better)\n\n```python\ndef stability_metric(y_true, y_pred):\n    \"\"\"\n    Competition metric for model optimization during training\n    \"\"\"\n    gini_in_time = []\n    weeks_to_score = X_val['WEEK_NUM'].reset_index(drop=True)\n    \n    for week in weeks_to_score.unique():\n        week_idx = weeks_to_score.eq(week)\n        gini = np.array(2 * roc_auc_score(y_true[week_idx], y_pred[week_idx]) - 1)\n        gini_in_time.append(gini)\n\n    w_fallingrate = 88.0\n    w_resstd = -0.5\n    x = np.arange(len(gini_in_time))\n    y = np.array(gini_in_time)\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a * x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(y)\n    stability_score = avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n    is_higher_better = True\n\n    return 'stability_score', stability_score, is_higher_better\n\nparams = {'metric': 'None'}\nmodel = LGBMClassifier(**params)\nmodel.fit(X_train, y_train,\n         eval_set=[(X_val, y_val)],\n         eval_metric=stability_metric,\n         callbacks=[early_stopping(50)],\n        )\n```",
    "2804622": "could you plz also write a catboost version metric ,i tried ,but the index was unmatched.i dont know how to fix it",
    "2804644": "CatBoost doesn't support even simple AUC if you train on GPU, so I doubt it will work with a custom metric, but on CPU it doesn't make much sens"
  },
  "source": "meta"
}