{
  "id": 500868,
  "title": "gini_stability_custom_metric used in lightgbm👀",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/500868",
  "author_name": "RogerOcean",
  "post_date": "2024-05-07T08:17:27.248000",
  "votes": 46,
  "comment_count": 10,
  "views": 0,
  "content": "<p>hi: <br>\nRecently, I have been working on feature generation and feature selection. I realized that when using feature selection methods like REF and boruta, it is necessary to use competition metrics instead of just the AUC metric in lightgbm. Here are some evaluation functions that can be directly used in lgbm. I hope you find them useful.🥳</p>\n<pre><code> ():\n\n   \n\n   w_fallingrate = \n   w_resstd = -\n\n   base = pd.DataFrame()\n   base[] = week\n   base[] = y_true.get_label()\n   base[] = y_pred\n   gini_in_time = base.loc[:, [, , ]]\\\n       .sort_values()\\\n       .groupby()[[, ]]\\\n       .apply( x: *roc_auc_score(x[], x[])-).tolist()\n\n   x = np.arange((gini_in_time))\n   y = gini_in_time\n   a, b = np.polyfit(x, y, )\n   y_hat = a*x + b\n   residuals = y - y_hat\n   res_std = np.std(residuals)\n   avg_gini = np.mean(gini_in_time)\n\n   final_score = avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n\n    , final_score, \n\ntrain_data = lgb.Dataset(df_train.iloc[idx_train][num_cols], label=y.iloc[idx_train], categorical_feature=cat_cols)\nvalid_data = lgb.Dataset(df_train.iloc[idx_valid][num_cols], label=y.iloc[idx_valid], categorical_feature=cat_cols)\nparams = {\n   : ,\n   : ,\n   : , \n   : ,\n   : ,\n   : \n}\n\nmodel = lgb.train(params, train_data, valid_sets=[valid_data], num_boost_round=, \n                 feval= y_true, y_pred: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid][]))\n</code></pre>",
  "messages": [
    {
      "id": 2798431,
      "postDate": "2024-05-07T08:17:27.250Z",
      "content": "<p>hi: <br>\nRecently, I have been working on feature generation and feature selection. I realized that when using feature selection methods like REF and boruta, it is necessary to use competition metrics instead of just the AUC metric in lightgbm. Here are some evaluation functions that can be directly used in lgbm. I hope you find them useful.🥳</p>\n<pre><code> ():\n\n   \n\n   w_fallingrate = \n   w_resstd = -\n\n   base = pd.DataFrame()\n   base[] = week\n   base[] = y_true.get_label()\n   base[] = y_pred\n   gini_in_time = base.loc[:, [, , ]]\\\n       .sort_values()\\\n       .groupby()[[, ]]\\\n       .apply( x: *roc_auc_score(x[], x[])-).tolist()\n\n   x = np.arange((gini_in_time))\n   y = gini_in_time\n   a, b = np.polyfit(x, y, )\n   y_hat = a*x + b\n   residuals = y - y_hat\n   res_std = np.std(residuals)\n   avg_gini = np.mean(gini_in_time)\n\n   final_score = avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n\n    , final_score, \n\ntrain_data = lgb.Dataset(df_train.iloc[idx_train][num_cols], label=y.iloc[idx_train], categorical_feature=cat_cols)\nvalid_data = lgb.Dataset(df_train.iloc[idx_valid][num_cols], label=y.iloc[idx_valid], categorical_feature=cat_cols)\nparams = {\n   : ,\n   : ,\n   : , \n   : ,\n   : ,\n   : \n}\n\nmodel = lgb.train(params, train_data, valid_sets=[valid_data], num_boost_round=, \n                 feval= y_true, y_pred: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid][]))\n</code></pre>",
      "rawMarkdown": "\nhi: \n\nRecently, I have been working on feature generation and feature selection. I realized that when using feature selection methods like REF and boruta, it is necessary to use competition metrics instead of just the AUC metric in lightgbm. Here are some evaluation functions that can be directly used in lgbm. I hope you find them useful.🥳\n\n\n```python\n\n\ndef gini_stability_custom_metric(y_pred: np.array, y_true: lgb.Dataset, week: np.array):\n    \n    '''\n    :param y_pred:\n    :param y_true:\n    :param week: \n    :return eval_name: str\n    :return eval_result: float\n    :return is_higher_better: bool\n    '''\n    \n    w_fallingrate = 88.0\n    w_resstd = -0.5\n    \n    base = pd.DataFrame()\n    base['WEEK_NUM'] = week\n    base['target'] = y_true.get_label()\n    base['score'] = y_pred\n\n    gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    \n    final_score = avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n    \n    return 'gini_stability', final_score, True\n\n \ntrain_data = lgb.Dataset(df_train.iloc[idx_train][num_cols], label=y.iloc[idx_train], categorical_feature=cat_cols)\nvalid_data = lgb.Dataset(df_train.iloc[idx_valid][num_cols], label=y.iloc[idx_valid], categorical_feature=cat_cols)\n\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    'objective': 'binary',\n    'metric': 'custom', \n    'num_leaves': 31,\n    'learning_rate': 0.05,\n    'feature_fraction': 0.9\n}\n\n# During model evaluation, it is done in the order of the input data.\nmodel = lgb.train(params, train_data, valid_sets=[valid_data], num_boost_round=100, \n                  feval=lambda y_true, y_pred: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid]['WEEK_NUM']))\n\n```\n\n\n\n\n\n\n\n",
      "votes": 46
    },
    {
      "id": 2799145,
      "postDate": "2024-05-07T15:42:02.890Z",
      "content": "<p>Nice idea! But shouldn't it be:</p>\n<p><code>feval=lambda y_pred, y_true: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid]['WEEK_NUM']))</code></p>\n<p>according to this <a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/engine.html#train\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/engine.html#train</a>?</p>",
      "rawMarkdown": "Nice idea! But shouldn't it be:\n\n `feval=lambda y_pred, y_true: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid]['WEEK_NUM']))`\n\naccording to this https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/engine.html#train?",
      "votes": 1,
      "replies": [
        {
          "id": 2802351,
          "postDate": "2024-05-09T02:25:59.450Z",
          "content": "<p>Thus, how can we apply this which is suitable?</p>",
          "rawMarkdown": "Thus, how can we apply this which is suitable?"
        }
      ]
    },
    {
      "id": 2803600,
      "postDate": "2024-05-09T15:19:20.623Z",
      "content": "<p>Nice idea!</p>",
      "rawMarkdown": "Nice idea!\n"
    },
    {
      "id": 2798490,
      "postDate": "2024-05-07T08:49:52.370Z",
      "content": "<p>Nice one! Does it give you any direct change of score on LB?</p>",
      "rawMarkdown": "Nice one! Does it give you any direct change of score on LB?",
      "replies": [
        {
          "id": 2798771,
          "postDate": "2024-05-07T12:10:15.690Z",
          "content": "<p>I have some similar experiments on gini_stability_custom_metric but it didn't perform better at LB… <br>\nHas anyone successfully increased their score using this method? Can you share the changes in cv and lb?</p>",
          "rawMarkdown": "I have some similar experiments on gini_stability_custom_metric but it didn't perform better at LB... \nHas anyone successfully increased their score using this method? Can you share the changes in cv and lb?",
          "votes": 1,
          "isDeleted": true,
          "replies": [
            {
              "id": 2799283,
              "postDate": "2024-05-07T17:17:01.320Z",
              "content": "<p>I checked the components of the stabiliy score on my test set like a month ago. If I recall correctly, the slope of the regression line wasn't negative, meaning I didn't see any deterioration. It is of course possible that this was due to the specific time split I used, but if no such deterioration can be found I doubt that optimizing against this will work. (and I kind of doubt it will work even if one can see a deterioration over time)</p>\n<p>Besides, I believe there are other potentially more promising ways of making it stable over time. With that said, nice work with the function <a href=\"https://www.kaggle.com/roger92\" target=\"_blank\">@roger92</a> </p>\n<p>Nontheless (regardless of potential metric hacks) I find the competition interesting and a bit different.</p>",
              "rawMarkdown": "I checked the components of the stabiliy score on my test set like a month ago. If I recall correctly, the slope of the regression line wasn't negative, meaning I didn't see any deterioration. It is of course possible that this was due to the specific time split I used, but if no such deterioration can be found I doubt that optimizing against this will work. (and I kind of doubt it will work even if one can see a deterioration over time)\n\nBesides, I believe there are other potentially more promising ways of making it stable over time. With that said, nice work with the function @roger92 \n\nNontheless (regardless of potential metric hacks) I find the competition interesting and a bit different.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2801235,
      "postDate": "2024-05-08T14:52:15.463Z",
      "content": "<p>Thanks for your sharing</p>",
      "rawMarkdown": "Thanks for your sharing"
    },
    {
      "id": 2800196,
      "postDate": "2024-05-08T05:53:25.857Z",
      "content": "<p>Thanks for this nice post</p>",
      "rawMarkdown": "Thanks for this nice post"
    },
    {
      "id": 2799252,
      "postDate": "2024-05-07T16:53:17.187Z",
      "content": "<p>Great work , Thanks for sharing..</p>",
      "rawMarkdown": "Great work , Thanks for sharing.."
    },
    {
      "id": 2798712,
      "postDate": "2024-05-07T11:30:00.143Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 2799145,
      "author_name": "Bassel Abdelmassih",
      "author_url": "",
      "post_date": "2024-05-07T15:42:02.890000",
      "content": "<p>Nice idea! But shouldn't it be:</p>\n<p><code>feval=lambda y_pred, y_true: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid]['WEEK_NUM']))</code></p>\n<p>according to this <a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/engine.html#train\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/engine.html#train</a>?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2802351,
          "author_name": "Pham Hoang Le Nguyen",
          "author_url": "",
          "post_date": "2024-05-09T02:25:59.450000",
          "content": "<p>Thus, how can we apply this which is suitable?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2803600,
      "author_name": "Sheema Zain",
      "author_url": "",
      "post_date": "2024-05-09T15:19:20.623000",
      "content": "<p>Nice idea!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2798490,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-05-07T08:49:52.370000",
      "content": "<p>Nice one! Does it give you any direct change of score on LB?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2798771,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-07T12:10:15.690000",
          "content": "<p>I have some similar experiments on gini_stability_custom_metric but it didn't perform better at LB… <br>\nHas anyone successfully increased their score using this method? Can you share the changes in cv and lb?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2799283,
              "author_name": "Ern711",
              "author_url": "",
              "post_date": "2024-05-07T17:17:01.320000",
              "content": "<p>I checked the components of the stabiliy score on my test set like a month ago. If I recall correctly, the slope of the regression line wasn't negative, meaning I didn't see any deterioration. It is of course possible that this was due to the specific time split I used, but if no such deterioration can be found I doubt that optimizing against this will work. (and I kind of doubt it will work even if one can see a deterioration over time)</p>\n<p>Besides, I believe there are other potentially more promising ways of making it stable over time. With that said, nice work with the function <a href=\"https://www.kaggle.com/roger92\" target=\"_blank\">@roger92</a> </p>\n<p>Nontheless (regardless of potential metric hacks) I find the competition interesting and a bit different.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2801235,
      "author_name": "Dark yang boy",
      "author_url": "",
      "post_date": "2024-05-08T14:52:15.463000",
      "content": "<p>Thanks for your sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2800196,
      "author_name": "Sabrina Ho",
      "author_url": "",
      "post_date": "2024-05-08T05:53:25.857000",
      "content": "<p>Thanks for this nice post</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2799252,
      "author_name": "Anil Kumar Reddy",
      "author_url": "",
      "post_date": "2024-05-07T16:53:17.187000",
      "content": "<p>Great work , Thanks for sharing..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2798712,
      "author_name": "Sergio Henrique",
      "author_url": "",
      "post_date": "2024-05-07T11:30:00.143000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2798431": "\nhi: \n\nRecently, I have been working on feature generation and feature selection. I realized that when using feature selection methods like REF and boruta, it is necessary to use competition metrics instead of just the AUC metric in lightgbm. Here are some evaluation functions that can be directly used in lgbm. I hope you find them useful.🥳\n\n\n```python\n\n\ndef gini_stability_custom_metric(y_pred: np.array, y_true: lgb.Dataset, week: np.array):\n    \n    '''\n    :param y_pred:\n    :param y_true:\n    :param week: \n    :return eval_name: str\n    :return eval_result: float\n    :return is_higher_better: bool\n    '''\n    \n    w_fallingrate = 88.0\n    w_resstd = -0.5\n    \n    base = pd.DataFrame()\n    base['WEEK_NUM'] = week\n    base['target'] = y_true.get_label()\n    base['score'] = y_pred\n\n    gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    \n    final_score = avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n    \n    return 'gini_stability', final_score, True\n\n \ntrain_data = lgb.Dataset(df_train.iloc[idx_train][num_cols], label=y.iloc[idx_train], categorical_feature=cat_cols)\nvalid_data = lgb.Dataset(df_train.iloc[idx_valid][num_cols], label=y.iloc[idx_valid], categorical_feature=cat_cols)\n\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    'objective': 'binary',\n    'metric': 'custom', \n    'num_leaves': 31,\n    'learning_rate': 0.05,\n    'feature_fraction': 0.9\n}\n\n# During model evaluation, it is done in the order of the input data.\nmodel = lgb.train(params, train_data, valid_sets=[valid_data], num_boost_round=100, \n                  feval=lambda y_true, y_pred: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid]['WEEK_NUM']))\n\n```\n\n\n\n\n\n\n\n",
    "2799145": "Nice idea! But shouldn't it be:\n\n `feval=lambda y_pred, y_true: gini_stability_custom_metric(y_pred, y_true, df_train.iloc[idx_valid]['WEEK_NUM']))`\n\naccording to this https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/engine.html#train?",
    "2803600": "Nice idea!\n",
    "2798490": "Nice one! Does it give you any direct change of score on LB?",
    "2801235": "Thanks for your sharing",
    "2800196": "Thanks for this nice post",
    "2799252": "Great work , Thanks for sharing..",
    "2798712": "Thanks for sharing!"
  }
}