{
  "id": 473800,
  "title": "LGBM Competition Metric",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/473800",
  "author_name": "werus23",
  "post_date": "2024-02-06T04:58:53.866000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I think this is the correct implementation of the competition metric for LGBM validation purposes</p>\n<pre><code> sklearn\n scipy  stats\n\n ():\n    index = week.argsort()\n    true = true[index]\n    pred = pred[index]\n\n    uniq = np.unique(week[index], return_index=)\n    week_unique, week_index = uniq[],uniq[][:]\n    grouped_true = np.split(true, week_index)\n    grouped_pred = np.split(pred, week_index)\n\n    ginis = np.zeros((week_unique))\n     i, (true,pred)  ((grouped_true,grouped_pred)):\n        gini = sklearn.metrics.roc_auc_score(true, pred)*-\n        ginis[i] = gini\n\n    slope, intercept, _, _, _ = stats.linregress(week_unique,ginis)\n    residuals = ginis - (slope*week_unique + intercept)\n\n     np.mean(ginis) +  * (,slope) -  * np.std(residuals)\n\n ():   \n     ():\n        labels = train_data.get_label()\n        weeks = train_data.indexes\n        y_pred = preds.reshape(-, )\n\n        score = hc_metric(labels, \n                        y_pred, \n                        weeks\n                       )\n\n         , score, \n     func\n</code></pre>\n<p>you need to set the indexes of the lgb dataset as the week number</p>\n<pre><code> = lgb.Dataset(X_train, y_train)\n = lgb.Dataset(X_valid, y_valid)\n = week_train\n = week_valid\n</code></pre>",
  "messages": [
    {
      "id": 2638113,
      "postDate": "2024-02-06T04:58:53.867Z",
      "content": "<p>I think this is the correct implementation of the competition metric for LGBM validation purposes</p>\n<pre><code> sklearn\n scipy  stats\n\n ():\n    index = week.argsort()\n    true = true[index]\n    pred = pred[index]\n\n    uniq = np.unique(week[index], return_index=)\n    week_unique, week_index = uniq[],uniq[][:]\n    grouped_true = np.split(true, week_index)\n    grouped_pred = np.split(pred, week_index)\n\n    ginis = np.zeros((week_unique))\n     i, (true,pred)  ((grouped_true,grouped_pred)):\n        gini = sklearn.metrics.roc_auc_score(true, pred)*-\n        ginis[i] = gini\n\n    slope, intercept, _, _, _ = stats.linregress(week_unique,ginis)\n    residuals = ginis - (slope*week_unique + intercept)\n\n     np.mean(ginis) +  * (,slope) -  * np.std(residuals)\n\n ():   \n     ():\n        labels = train_data.get_label()\n        weeks = train_data.indexes\n        y_pred = preds.reshape(-, )\n\n        score = hc_metric(labels, \n                        y_pred, \n                        weeks\n                       )\n\n         , score, \n     func\n</code></pre>\n<p>you need to set the indexes of the lgb dataset as the week number</p>\n<pre><code> = lgb.Dataset(X_train, y_train)\n = lgb.Dataset(X_valid, y_valid)\n = week_train\n = week_valid\n</code></pre>",
      "rawMarkdown": "I think this is the correct implementation of the competition metric for LGBM validation purposes\n\n```python\nimport sklearn\nfrom scipy import stats\n\ndef hc_metric(true, pred, week):\n    index = week.argsort()\n    true = true[index]\n    pred = pred[index]\n\n    uniq = np.unique(week[index], return_index=True)\n    week_unique, week_index = uniq[0],uniq[1][1:]\n    grouped_true = np.split(true, week_index)\n    grouped_pred = np.split(pred, week_index)\n\n    ginis = np.zeros(len(week_unique))\n    for i, (true,pred) in enumerate(zip(grouped_true,grouped_pred)):\n        gini = sklearn.metrics.roc_auc_score(true, pred)*2-1\n        ginis[i] = gini\n        \n    slope, intercept, _, _, _ = stats.linregress(week_unique,ginis)\n    residuals = ginis - (slope*week_unique + intercept)\n    \n    return np.mean(ginis) + 88.0 * min(0,slope) - 0.5 * np.std(residuals)\n\ndef custom_lgb_metric():   \n    def func(preds, train_data):\n        labels = train_data.get_label()\n        weeks = train_data.indexes\n        y_pred = preds.reshape(-1, 1)\n\n        score = hc_metric(labels, \n                        y_pred, \n                        weeks\n                       )\n       \n        return 'hm_metric', score, True\n    return func\n```\n\nyou need to set the indexes of the lgb dataset as the week number\n```\ndtrain = lgb.Dataset(X_train, y_train)\ndval = lgb.Dataset(X_valid, y_valid)\ndtrain.indexes = week_train\ndval.indexes = week_valid\n```",
      "votes": 4
    },
    {
      "id": 2638266,
      "postDate": "2024-02-06T07:17:20.130Z",
      "content": "<p>You can see the implementation directly in the notebook I prepared for the competition <a href=\"https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\" target=\"_blank\">https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook</a></p>",
      "rawMarkdown": "You can see the implementation directly in the notebook I prepared for the competition https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook",
      "votes": 3,
      "replies": [
        {
          "id": 2645138,
          "postDate": "2024-02-10T01:43:48.787Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2662800,
      "postDate": "2024-02-22T06:04:31.500Z",
      "content": "<p>I want to know what the effect will be if only the lightgbm model is used in this competition.</p>",
      "rawMarkdown": "I want to know what the effect will be if only the lightgbm model is used in this competition."
    },
    {
      "id": 2645140,
      "postDate": "2024-02-10T01:44:02.530Z",
      "content": "<p>Outstanding work!</p>",
      "rawMarkdown": "Outstanding work!"
    }
  ],
  "comments": [
    {
      "id": 2638266,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-06T07:17:20.130000",
      "content": "<p>You can see the implementation directly in the notebook I prepared for the competition <a href=\"https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\" target=\"_blank\">https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2645138,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-10T01:43:48.787000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2662800,
      "author_name": "Ziyi Dong",
      "author_url": "",
      "post_date": "2024-02-22T06:04:31.500000",
      "content": "<p>I want to know what the effect will be if only the lightgbm model is used in this competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2645140,
      "author_name": "Linsida",
      "author_url": "",
      "post_date": "2024-02-10T01:44:02.530000",
      "content": "<p>Outstanding work!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2638113": "I think this is the correct implementation of the competition metric for LGBM validation purposes\n\n```python\nimport sklearn\nfrom scipy import stats\n\ndef hc_metric(true, pred, week):\n    index = week.argsort()\n    true = true[index]\n    pred = pred[index]\n\n    uniq = np.unique(week[index], return_index=True)\n    week_unique, week_index = uniq[0],uniq[1][1:]\n    grouped_true = np.split(true, week_index)\n    grouped_pred = np.split(pred, week_index)\n\n    ginis = np.zeros(len(week_unique))\n    for i, (true,pred) in enumerate(zip(grouped_true,grouped_pred)):\n        gini = sklearn.metrics.roc_auc_score(true, pred)*2-1\n        ginis[i] = gini\n        \n    slope, intercept, _, _, _ = stats.linregress(week_unique,ginis)\n    residuals = ginis - (slope*week_unique + intercept)\n    \n    return np.mean(ginis) + 88.0 * min(0,slope) - 0.5 * np.std(residuals)\n\ndef custom_lgb_metric():   \n    def func(preds, train_data):\n        labels = train_data.get_label()\n        weeks = train_data.indexes\n        y_pred = preds.reshape(-1, 1)\n\n        score = hc_metric(labels, \n                        y_pred, \n                        weeks\n                       )\n       \n        return 'hm_metric', score, True\n    return func\n```\n\nyou need to set the indexes of the lgb dataset as the week number\n```\ndtrain = lgb.Dataset(X_train, y_train)\ndval = lgb.Dataset(X_valid, y_valid)\ndtrain.indexes = week_train\ndval.indexes = week_valid\n```",
    "2638266": "You can see the implementation directly in the notebook I prepared for the competition https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook",
    "2662800": "I want to know what the effect will be if only the lightgbm model is used in this competition.",
    "2645140": "Outstanding work!"
  }
}