{
  "id": 344038,
  "title": "Custom LGBM Obj: Weighted LogLoss Function",
  "url": "/competitions/amex-default-prediction/discussion/344038",
  "author_name": "",
  "post_date": "2022-08-13T17:16:03.810726900Z",
  "votes": 23,
  "comment_count": 17,
  "views": 0,
  "content": "<p>This competition has a metric where 50% depends on the percentage of positive cases that are in the top 4%.</p>\n<p>The idea of the proposed loss function is <strong>to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions</strong>. The intention is that the training of the LGB model will put more effort to improve the prediction of these instances.</p>\n<p>In this case, an increasing exponential curve like the one in the figure is used to calculate the weights. Thus, the weights increase as the distance of the positive instances to the 4% zone increases. The challenge is <strong>to determine the best values for this curve (MULT_NO4PERC and MAX_WEIGHTS)</strong>.</p>\n<pre><code>MULT_NO4PERC = 5.0\nMAX_WEIGHTS = 2.0\nlen_preds = 458913//5  # OOF preds\nplt.plot(1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1.0,0,len_preds)))\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F276462%2F62df0ffeb3c271ea493dc07c3803ce64%2Fweigths_curve.png?generation=1660408151561657&amp;alt=media\" alt=\"\"></p>\n<p>The weights of the negative cases are set to 1.0, and the <strong>positive cases that are within the 4% zone are also set to 1.0.</strong></p>\n<p>This loss function can be easily adapted to other algorithms such as: XGB, CATBoost, ANNs, RRNN, SVR, …</p>\n<p>Following this idea, other powerful custom loss functions can be created.</p>\n<p>This is the proposed <em>Weighted LogLoss Function</em>:</p>\n<pre><code>def weighted_logloss(preds, dtrain):\n   global MULT_NO4PERC, MAX_WEIGHTS\n   eps = 1e-16\n   labels = dtrain.get_label()\n   preds = 1.0 / (1.0 + np.exp(-preds))\n\n   # top 4 perc\n   labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n   pos_ord = labels_mat[:, 2].argsort()[::-1]\n   labels_mat = labels_mat[pos_ord]\n   weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n   top4 = np.cumsum(weights_4perc) &lt;= int(0.04 * np.sum(weights_4perc))\n   top4 = top4[labels_mat[:, 0].argsort()]\n\n   weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n   # Set to one weights of positive labels in top 4perc\n   weights[top4 &amp; (labels==1.0)] = 1.0 \n   # Set to one weights of negative labels\n   weights[(labels==0.0)] = 1.0 \n\n   grad = (preds - labels) * weights\n   hess = np.maximum(preds * (1.0 - preds) * weights , eps)\n   return grad, hess\n\n\n# How to include the custom objective loss function   \nmodel = lgb.train(..., \n      fobj = weighted_logloss)\n        )\n</code></pre>\n<p><strong>One basic notebook</strong> with the weighted loss custom function can be found <a href=\"https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function\" target=\"_blank\">here.</a></p>",
  "messages": [
    {
      "id": "1897354",
      "postDate": "08/13/2022 17:16:03",
      "content": "<p>This competition has a metric where 50% depends on the percentage of positive cases that are in the top 4%.</p>\n<p>The idea of the proposed loss function is <strong>to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions</strong>. The intention is that the training of the LGB model will put more effort to improve the prediction of these instances.</p>\n<p>In this case, an increasing exponential curve like the one in the figure is used to calculate the weights. Thus, the weights increase as the distance of the positive instances to the 4% zone increases. The challenge is <strong>to determine the best values for this curve (MULT_NO4PERC and MAX_WEIGHTS)</strong>.</p>\n<pre><code>MULT_NO4PERC = 5.0\nMAX_WEIGHTS = 2.0\nlen_preds = 458913//5  # OOF preds\nplt.plot(1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1.0,0,len_preds)))\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F276462%2F62df0ffeb3c271ea493dc07c3803ce64%2Fweigths_curve.png?generation=1660408151561657&amp;alt=media\" alt=\"\"></p>\n<p>The weights of the negative cases are set to 1.0, and the <strong>positive cases that are within the 4% zone are also set to 1.0.</strong></p>\n<p>This loss function can be easily adapted to other algorithms such as: XGB, CATBoost, ANNs, RRNN, SVR, …</p>\n<p>Following this idea, other powerful custom loss functions can be created.</p>\n<p>This is the proposed <em>Weighted LogLoss Function</em>:</p>\n<pre><code>def weighted_logloss(preds, dtrain):\n   global MULT_NO4PERC, MAX_WEIGHTS\n   eps = 1e-16\n   labels = dtrain.get_label()\n   preds = 1.0 / (1.0 + np.exp(-preds))\n\n   # top 4 perc\n   labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n   pos_ord = labels_mat[:, 2].argsort()[::-1]\n   labels_mat = labels_mat[pos_ord]\n   weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n   top4 = np.cumsum(weights_4perc) &lt;= int(0.04 * np.sum(weights_4perc))\n   top4 = top4[labels_mat[:, 0].argsort()]\n\n   weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n   # Set to one weights of positive labels in top 4perc\n   weights[top4 &amp; (labels==1.0)] = 1.0 \n   # Set to one weights of negative labels\n   weights[(labels==0.0)] = 1.0 \n\n   grad = (preds - labels) * weights\n   hess = np.maximum(preds * (1.0 - preds) * weights , eps)\n   return grad, hess\n\n\n# How to include the custom objective loss function   \nmodel = lgb.train(..., \n      fobj = weighted_logloss)\n        )\n</code></pre>\n<p><strong>One basic notebook</strong> with the weighted loss custom function can be found <a href=\"https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function\" target=\"_blank\">here.</a></p>",
      "rawMarkdown": "This competition has a metric where 50% depends on the percentage of positive cases that are in the top 4%.\n\nThe idea of the proposed loss function is **to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions**. The intention is that the training of the LGB model will put more effort to improve the prediction of these instances.\n\nIn this case, an increasing exponential curve like the one in the figure is used to calculate the weights. Thus, the weights increase as the distance of the positive instances to the 4% zone increases. The challenge is **to determine the best values for this curve (MULT_NO4PERC and MAX_WEIGHTS)**.\n\n    MULT_NO4PERC = 5.0\n    MAX_WEIGHTS = 2.0\n    len_preds = 458913//5  # OOF preds\n    plt.plot(1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1.0,0,len_preds)))\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F276462%2F62df0ffeb3c271ea493dc07c3803ce64%2Fweigths_curve.png?generation=1660408151561657&alt=media)\n\nThe weights of the negative cases are set to 1.0, and the **positive cases that are within the 4% zone are also set to 1.0.**\n\nThis loss function can be easily adapted to other algorithms such as: XGB, CATBoost, ANNs, RRNN, SVR, ...\n\nFollowing this idea, other powerful custom loss functions can be created.\n\nThis is the proposed *Weighted LogLoss Function*:\n\n    def weighted_logloss(preds, dtrain):\n       global MULT_NO4PERC, MAX_WEIGHTS\n       eps = 1e-16\n       labels = dtrain.get_label()\n       preds = 1.0 / (1.0 + np.exp(-preds))\n    \n       # top 4 perc\n       labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n       pos_ord = labels_mat[:, 2].argsort()[::-1]\n       labels_mat = labels_mat[pos_ord]\n       weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n       top4 = np.cumsum(weights_4perc) <= int(0.04 * np.sum(weights_4perc))\n       top4 = top4[labels_mat[:, 0].argsort()]\n\n       weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n       # Set to one weights of positive labels in top 4perc\n       weights[top4 & (labels==1.0)] = 1.0 \n       # Set to one weights of negative labels\n       weights[(labels==0.0)] = 1.0 \n\n       grad = (preds - labels) * weights\n       hess = np.maximum(preds * (1.0 - preds) * weights , eps)\n       return grad, hess\n\n\n    # How to include the custom objective loss function   \n    model = lgb.train(..., \n          fobj = weighted_logloss)\n            )\n\n\n**One basic notebook** with the weighted loss custom function can be found [here.](https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function)",
      "votes": null
    },
    {
      "id": "1897442",
      "postDate": "08/13/2022 18:13:11",
      "content": "<p>thanks for your topic</p>",
      "rawMarkdown": "thanks for your topic",
      "votes": null
    },
    {
      "id": "1898659",
      "postDate": "08/14/2022 17:55:11",
      "content": "<p>Thank you for sharing this. I brought up <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268\" target=\"_blank\"><strong>before</strong></a> the issue of a more appropriate objective function, so I was eager to try your implementation.</p>\n<p>I tried <code>MAX_WEIGHTS = 2.0</code> fixed with <code>MULT_NO4PERC one of [0.5, 5.0, 10.0]</code>, and with two different sets of LGB parameters. In both cases <code>MULT_NO4PERC = 5.0</code> worked the best, followed by <code>MULT_NO4PERC = 10.0</code>, but the improvement was only on the 4th decimal place. In 4 out of 5 folds it got to early stopping faster compared to the binary objective function. In my view that's the biggest value of your objective function, as it will likely allow faster hyperparameter search.</p>\n<p>What could be the problem for some users is that predictions are not calibrated. I suspect they can be submitted as such, or one can rank-order them and scale. Still, the latter solution is not compatible with ensembling, so the predictions should probably be calibrated properly.</p>\n<p>It is not a comprehensive evaluation by any means, but hopefully it is still useful.</p>",
      "rawMarkdown": "Thank you for sharing this. I brought up [**before**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268) the issue of a more appropriate objective function, so I was eager to try your implementation.\n\nI tried `MAX_WEIGHTS = 2.0` fixed with `MULT_NO4PERC one of [0.5, 5.0, 10.0]`, and with two different sets of LGB parameters. In both cases `MULT_NO4PERC = 5.0` worked the best, followed by `MULT_NO4PERC = 10.0`, but the improvement was only on the 4th decimal place. In 4 out of 5 folds it got to early stopping faster compared to the binary objective function. In my view that's the biggest value of your objective function, as it will likely allow faster hyperparameter search.\n\nWhat could be the problem for some users is that predictions are not calibrated. I suspect they can be submitted as such, or one can rank-order them and scale. Still, the latter solution is not compatible with ensembling, so the predictions should probably be calibrated properly.\n\nIt is not a comprehensive evaluation by any means, but hopefully it is still useful.",
      "votes": null
    },
    {
      "id": "1898761",
      "postDate": "08/14/2022 19:33:13",
      "content": "<p>Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights.</p>\n<p>You could simplify things by setting <code>weights = exp(-k * raw_pred_delta)</code>, where k is a tunable constant and <code>raw_pred_delta</code> is the difference between the raw predictions and their mean. This lets you easily calculate the hessian using the chain rule and the fact that e^x is its own derivative:</p>\n<pre><code>k = MY_TUNABLE_CONSTANT\npreds = preds - preds.mean()  # This is \"raw_pred_delta\"\nweights = np.exp(-1 * k * preds)  \nweights[label == 1] = 1\nprobs = 1.0 / (1.0 + np.exp(-preds))\ngrad = (probs - labels) * weights\nhess = probs * (1.0 - probs) * weights\nhess[label == 0] = hess[label  == 0] - k * grad[label == 0]  # The second term is the second half from the product and chain rules.\n</code></pre>",
      "rawMarkdown": "Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights.\n\nYou could simplify things by setting `weights = exp(-k * raw_pred_delta)`, where k is a tunable constant and `raw_pred_delta` is the difference between the raw predictions and their mean. This lets you easily calculate the hessian using the chain rule and the fact that e^x is its own derivative:\n\n```\nk = MY_TUNABLE_CONSTANT\npreds = preds - preds.mean()  # This is \"raw_pred_delta\"\nweights = np.exp(-1 * k * preds)  \nweights[label == 1] = 1\nprobs = 1.0 / (1.0 + np.exp(-preds))\ngrad = (probs - labels) * weights\nhess = probs * (1.0 - probs) * weights\nhess[label == 0] = hess[label  == 0] - k * grad[label == 0]  # The second term is the second half from the product and chain rules.\n```",
      "votes": null
    },
    {
      "id": "1899560",
      "postDate": "08/15/2022 10:58:33",
      "content": "<p>I tried something similar, comments on <a href=\"https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function/\" target=\"_blank\">the notebook</a></p>",
      "rawMarkdown": "I tried something similar, comments on [the notebook](https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function/)",
      "votes": null
    },
    {
      "id": "1899598",
      "postDate": "08/15/2022 11:39:44",
      "content": "<blockquote>\n  <p>Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights. </p>\n</blockquote>\n<p>OK, thanks David. You are right. I oversimplified the problem. As you say, I have to apply the chain rule to calculate the partial derivatives. I must reformulate the problem. Thanks.</p>\n<p>However, in your approach you are not using weights based on ranking. Is it working for you?</p>",
      "rawMarkdown": "> Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights. \n\nOK, thanks David. You are right. I oversimplified the problem. As you say, I have to apply the chain rule to calculate the partial derivatives. I must reformulate the problem. Thanks.\n\nHowever, in your approach you are not using weights based on ranking. Is it working for you?",
      "votes": null
    },
    {
      "id": "1899607",
      "postDate": "08/15/2022 11:46:10",
      "content": "<p>Thanks for your comment. I think the custom loss must be reformulated…</p>",
      "rawMarkdown": "Thanks for your comment. I think the custom loss must be reformulated...",
      "votes": null
    },
    {
      "id": "1899628",
      "postDate": "08/15/2022 12:00:35",
      "content": "<p>ok, thanks for your evaluation. I think you are right but I am not giving up hope of finding a loss function that will produce a slight improvement. Thanks</p>",
      "rawMarkdown": "ok, thanks for your evaluation. I think you are right but I am not giving up hope of finding a loss function that will produce a slight improvement. Thanks",
      "votes": null
    },
    {
      "id": "1899777",
      "postDate": "08/15/2022 13:46:40",
      "content": "<p>Very good idea! Thanks for sharing!</p>",
      "rawMarkdown": "Very good idea! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1900168",
      "postDate": "08/15/2022 18:48:18",
      "content": "<p>Unfortunately, once you move to using explicit ranking, the actual loss function becomes non-parametrizable. I have not found any clear improvements from my many efforts. I tried lots of things:</p>\n<ul>\n<li>Elevating items in the top percentiles by a fixed amount, leading to a discontinuity at the breakpoint.</li>\n<li>Emulating gain from 4-percent_capture by adding a normal distribution curve centered at the 96th percentile with a sigma = 3 percentile's worth of rank.</li>\n<li>Calculating actual gradient and hessian empirically.</li>\n</ul>\n<p>I'm going to try a version of the simpler approach I shared with you and see if it makes a difference.</p>",
      "rawMarkdown": "Unfortunately, once you move to using explicit ranking, the actual loss function becomes non-parametrizable. I have not found any clear improvements from my many efforts. I tried lots of things:\n- Elevating items in the top percentiles by a fixed amount, leading to a discontinuity at the breakpoint.\n- Emulating gain from 4-percent_capture by adding a normal distribution curve centered at the 96th percentile with a sigma = 3 percentile's worth of rank.\n- Calculating actual gradient and hessian empirically.\n\nI'm going to try a version of the simpler approach I shared with you and see if it makes a difference.",
      "votes": null
    },
    {
      "id": "1900773",
      "postDate": "08/16/2022 08:41:23",
      "content": "<p>change loss function doesn't work, I even rebuild a super simple loss function and the final result doesn't change</p>",
      "rawMarkdown": "change loss function doesn't work, I even rebuild a super simple loss function and the final result doesn't change",
      "votes": null
    },
    {
      "id": "1901105",
      "postDate": "08/16/2022 13:16:33",
      "content": "<p>I see the potential improvement in this. But did you check it's performance in practice? Does it improve the score in comarison to not using it?</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "I see the potential improvement in this. But did you check it's performance in practice? Does it improve the score in comarison to not using it?\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "1901460",
      "postDate": "08/16/2022 17:03:12",
      "content": "<p>I have already answered this question on a limited set of 2 examples (see earlier in this thread), and can now extend it to about half a dozen examples. This objective function doesn't converge to a better score, and that also goes for its <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578\" target=\"_blank\"><strong>updated version</strong></a>. Five-fold CV AmEx scores are essentially the same: 0.797648 with binary and 0.797577 with this function. Both score 0.799 on the LB but the former has better score when sorted.</p>\n<p>I think the main advantage is better speed during the hyperparameter search. Full DART run with as single set of parameters and binary objective function (5-fold) takes about 10 hours, compared to 8 hours with weighted loss objective function. So I am using it exclusively for that purpose.</p>",
      "rawMarkdown": "I have already answered this question on a limited set of 2 examples (see earlier in this thread), and can now extend it to about half a dozen examples. This objective function doesn't converge to a better score, and that also goes for its [**updated version**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578). Five-fold CV AmEx scores are essentially the same: 0.797648 with binary and 0.797577 with this function. Both score 0.799 on the LB but the former has better score when sorted.\n\nI think the main advantage is better speed during the hyperparameter search. Full DART run with as single set of parameters and binary objective function (5-fold) takes about 10 hours, compared to 8 hours with weighted loss objective function. So I am using it exclusively for that purpose.",
      "votes": null
    },
    {
      "id": "1903133",
      "postDate": "08/17/2022 06:15:58",
      "content": "<p><a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> Thanks for sharing! So this did you find the improvement you were aiming for with this function? <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578</a></p>",
      "rawMarkdown": "jpison Thanks for sharing! So this did you find the improvement you were aiming for with this function? https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578",
      "votes": null
    },
    {
      "id": "1903172",
      "postDate": "08/17/2022 06:58:48",
      "content": "<p>Thank you for sharing this.</p>",
      "rawMarkdown": "Thank you for sharing this.",
      "votes": null
    },
    {
      "id": "1907022",
      "postDate": "08/20/2022 12:00:09",
      "content": "<p>thanks for your help</p>",
      "rawMarkdown": "thanks for your help",
      "votes": null
    },
    {
      "id": "1910010",
      "postDate": "08/23/2022 05:49:58",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1911287",
      "postDate": "08/24/2022 02:21:05",
      "content": "<p>Be careful in changing your optimization objective. When you capture 1% improvement of metric A, you may obtain a 5% decrease of metric B.</p>",
      "rawMarkdown": "Be careful in changing your optimization objective. When you capture 1% improvement of metric A, you may obtain a 5% decrease of metric B.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1897442,
      "author_name": "hermanshow",
      "author_url": "",
      "post_date": "08/13/2022 18:13:11",
      "content": "<p>thanks for your topic</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1898659,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "08/14/2022 17:55:11",
      "content": "<p>Thank you for sharing this. I brought up <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268\" target=\"_blank\"><strong>before</strong></a> the issue of a more appropriate objective function, so I was eager to try your implementation.</p>\n<p>I tried <code>MAX_WEIGHTS = 2.0</code> fixed with <code>MULT_NO4PERC one of [0.5, 5.0, 10.0]</code>, and with two different sets of LGB parameters. In both cases <code>MULT_NO4PERC = 5.0</code> worked the best, followed by <code>MULT_NO4PERC = 10.0</code>, but the improvement was only on the 4th decimal place. In 4 out of 5 folds it got to early stopping faster compared to the binary objective function. In my view that's the biggest value of your objective function, as it will likely allow faster hyperparameter search.</p>\n<p>What could be the problem for some users is that predictions are not calibrated. I suspect they can be submitted as such, or one can rank-order them and scale. Still, the latter solution is not compatible with ensembling, so the predictions should probably be calibrated properly.</p>\n<p>It is not a comprehensive evaluation by any means, but hopefully it is still useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1899607,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "08/15/2022 11:46:10",
          "content": "<p>Thanks for your comment. I think the custom loss must be reformulated…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1898761,
      "author_name": "davidirudel",
      "author_url": "",
      "post_date": "08/14/2022 19:33:13",
      "content": "<p>Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights.</p>\n<p>You could simplify things by setting <code>weights = exp(-k * raw_pred_delta)</code>, where k is a tunable constant and <code>raw_pred_delta</code> is the difference between the raw predictions and their mean. This lets you easily calculate the hessian using the chain rule and the fact that e^x is its own derivative:</p>\n<pre><code>k = MY_TUNABLE_CONSTANT\npreds = preds - preds.mean()  # This is \"raw_pred_delta\"\nweights = np.exp(-1 * k * preds)  \nweights[label == 1] = 1\nprobs = 1.0 / (1.0 + np.exp(-preds))\ngrad = (probs - labels) * weights\nhess = probs * (1.0 - probs) * weights\nhess[label == 0] = hess[label  == 0] - k * grad[label == 0]  # The second term is the second half from the product and chain rules.\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1899598,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "08/15/2022 11:39:44",
          "content": "<blockquote>\n  <p>Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights. </p>\n</blockquote>\n<p>OK, thanks David. You are right. I oversimplified the problem. As you say, I have to apply the chain rule to calculate the partial derivatives. I must reformulate the problem. Thanks.</p>\n<p>However, in your approach you are not using weights based on ranking. Is it working for you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900168,
          "author_name": "davidirudel",
          "author_url": "",
          "post_date": "08/15/2022 18:48:18",
          "content": "<p>Unfortunately, once you move to using explicit ranking, the actual loss function becomes non-parametrizable. I have not found any clear improvements from my many efforts. I tried lots of things:</p>\n<ul>\n<li>Elevating items in the top percentiles by a fixed amount, leading to a discontinuity at the breakpoint.</li>\n<li>Emulating gain from 4-percent_capture by adding a normal distribution curve centered at the 96th percentile with a sigma = 3 percentile's worth of rank.</li>\n<li>Calculating actual gradient and hessian empirically.</li>\n</ul>\n<p>I'm going to try a version of the simpler approach I shared with you and see if it makes a difference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1899560,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "08/15/2022 10:58:33",
      "content": "<p>I tried something similar, comments on <a href=\"https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function/\" target=\"_blank\">the notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1899628,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "08/15/2022 12:00:35",
          "content": "<p>ok, thanks for your evaluation. I think you are right but I am not giving up hope of finding a loss function that will produce a slight improvement. Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1903133,
          "author_name": "kaito510",
          "author_url": "",
          "post_date": "08/17/2022 06:15:58",
          "content": "<p><a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> Thanks for sharing! So this did you find the improvement you were aiming for with this function? <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1899777,
      "author_name": "leewook",
      "author_url": "",
      "post_date": "08/15/2022 13:46:40",
      "content": "<p>Very good idea! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1900773,
      "author_name": "wenjunzhang323",
      "author_url": "",
      "post_date": "08/16/2022 08:41:23",
      "content": "<p>change loss function doesn't work, I even rebuild a super simple loss function and the final result doesn't change</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1901105,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "08/16/2022 13:16:33",
      "content": "<p>I see the potential improvement in this. But did you check it's performance in practice? Does it improve the score in comarison to not using it?</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1901460,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "08/16/2022 17:03:12",
          "content": "<p>I have already answered this question on a limited set of 2 examples (see earlier in this thread), and can now extend it to about half a dozen examples. This objective function doesn't converge to a better score, and that also goes for its <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578\" target=\"_blank\"><strong>updated version</strong></a>. Five-fold CV AmEx scores are essentially the same: 0.797648 with binary and 0.797577 with this function. Both score 0.799 on the LB but the former has better score when sorted.</p>\n<p>I think the main advantage is better speed during the hyperparameter search. Full DART run with as single set of parameters and binary objective function (5-fold) takes about 10 hours, compared to 8 hours with weighted loss objective function. So I am using it exclusively for that purpose.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1903172,
      "author_name": "",
      "author_url": "",
      "post_date": "08/17/2022 06:58:48",
      "content": "<p>Thank you for sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1907022,
      "author_name": "siyuhuang2022",
      "author_url": "",
      "post_date": "08/20/2022 12:00:09",
      "content": "<p>thanks for your help</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1910010,
      "author_name": "shencaocao",
      "author_url": "",
      "post_date": "08/23/2022 05:49:58",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1911287,
      "author_name": "voix97",
      "author_url": "",
      "post_date": "08/24/2022 02:21:05",
      "content": "<p>Be careful in changing your optimization objective. When you capture 1% improvement of metric A, you may obtain a 5% decrease of metric B.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1897354": "This competition has a metric where 50% depends on the percentage of positive cases that are in the top 4%.\n\nThe idea of the proposed loss function is **to penalize both the Gradient and the Hessian of the positive cases that are further away from the top positions**. The intention is that the training of the LGB model will put more effort to improve the prediction of these instances.\n\nIn this case, an increasing exponential curve like the one in the figure is used to calculate the weights. Thus, the weights increase as the distance of the positive instances to the 4% zone increases. The challenge is **to determine the best values for this curve (MULT_NO4PERC and MAX_WEIGHTS)**.\n\n    MULT_NO4PERC = 5.0\n    MAX_WEIGHTS = 2.0\n    len_preds = 458913//5  # OOF preds\n    plt.plot(1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1.0,0,len_preds)))\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F276462%2F62df0ffeb3c271ea493dc07c3803ce64%2Fweigths_curve.png?generation=1660408151561657&alt=media)\n\nThe weights of the negative cases are set to 1.0, and the **positive cases that are within the 4% zone are also set to 1.0.**\n\nThis loss function can be easily adapted to other algorithms such as: XGB, CATBoost, ANNs, RRNN, SVR, ...\n\nFollowing this idea, other powerful custom loss functions can be created.\n\nThis is the proposed *Weighted LogLoss Function*:\n\n    def weighted_logloss(preds, dtrain):\n       global MULT_NO4PERC, MAX_WEIGHTS\n       eps = 1e-16\n       labels = dtrain.get_label()\n       preds = 1.0 / (1.0 + np.exp(-preds))\n    \n       # top 4 perc\n       labels_mat = np.transpose(np.array([np.arange(len(labels)), labels, preds]))\n       pos_ord = labels_mat[:, 2].argsort()[::-1]\n       labels_mat = labels_mat[pos_ord]\n       weights_4perc = np.where(labels_mat[:,1]==0, 20, 1)\n       top4 = np.cumsum(weights_4perc) <= int(0.04 * np.sum(weights_4perc))\n       top4 = top4[labels_mat[:, 0].argsort()]\n\n       weights = 1+np.exp(-MULT_NO4PERC*np.linspace(MAX_WEIGHTS-1,0,len(top4)))[labels_mat[:, 0].argsort()]\n       # Set to one weights of positive labels in top 4perc\n       weights[top4 & (labels==1.0)] = 1.0 \n       # Set to one weights of negative labels\n       weights[(labels==0.0)] = 1.0 \n\n       grad = (preds - labels) * weights\n       hess = np.maximum(preds * (1.0 - preds) * weights , eps)\n       return grad, hess\n\n\n    # How to include the custom objective loss function   \n    model = lgb.train(..., \n          fobj = weighted_logloss)\n            )\n\n\n**One basic notebook** with the weighted loss custom function can be found [here.](https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function)",
    "1897442": "thanks for your topic",
    "1898659": "Thank you for sharing this. I brought up [**before**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268) the issue of a more appropriate objective function, so I was eager to try your implementation.\n\nI tried `MAX_WEIGHTS = 2.0` fixed with `MULT_NO4PERC one of [0.5, 5.0, 10.0]`, and with two different sets of LGB parameters. In both cases `MULT_NO4PERC = 5.0` worked the best, followed by `MULT_NO4PERC = 10.0`, but the improvement was only on the 4th decimal place. In 4 out of 5 folds it got to early stopping faster compared to the binary objective function. In my view that's the biggest value of your objective function, as it will likely allow faster hyperparameter search.\n\nWhat could be the problem for some users is that predictions are not calibrated. I suspect they can be submitted as such, or one can rank-order them and scale. Still, the latter solution is not compatible with ensembling, so the predictions should probably be calibrated properly.\n\nIt is not a comprehensive evaluation by any means, but hopefully it is still useful.",
    "1898761": "Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights.\n\nYou could simplify things by setting `weights = exp(-k * raw_pred_delta)`, where k is a tunable constant and `raw_pred_delta` is the difference between the raw predictions and their mean. This lets you easily calculate the hessian using the chain rule and the fact that e^x is its own derivative:\n\n```\nk = MY_TUNABLE_CONSTANT\npreds = preds - preds.mean()  # This is \"raw_pred_delta\"\nweights = np.exp(-1 * k * preds)  \nweights[label == 1] = 1\nprobs = 1.0 / (1.0 + np.exp(-preds))\ngrad = (probs - labels) * weights\nhess = probs * (1.0 - probs) * weights\nhess[label == 0] = hess[label  == 0] - k * grad[label == 0]  # The second term is the second half from the product and chain rules.\n```",
    "1899560": "I tried something similar, comments on [the notebook](https://www.kaggle.com/code/jpison/custom-lgbm-obj-weighted-logloss-function/)",
    "1899598": "> Note that the Hessian is incorrect. Since weight depends on y, which itself depends on preds, you have to use the product rule to calculate the hessian. You cannot just multiply the standard hessian by the weights. \n\nOK, thanks David. You are right. I oversimplified the problem. As you say, I have to apply the chain rule to calculate the partial derivatives. I must reformulate the problem. Thanks.\n\nHowever, in your approach you are not using weights based on ranking. Is it working for you?",
    "1899607": "Thanks for your comment. I think the custom loss must be reformulated...",
    "1899628": "ok, thanks for your evaluation. I think you are right but I am not giving up hope of finding a loss function that will produce a slight improvement. Thanks",
    "1899777": "Very good idea! Thanks for sharing!",
    "1900168": "Unfortunately, once you move to using explicit ranking, the actual loss function becomes non-parametrizable. I have not found any clear improvements from my many efforts. I tried lots of things:\n- Elevating items in the top percentiles by a fixed amount, leading to a discontinuity at the breakpoint.\n- Emulating gain from 4-percent_capture by adding a normal distribution curve centered at the 96th percentile with a sigma = 3 percentile's worth of rank.\n- Calculating actual gradient and hessian empirically.\n\nI'm going to try a version of the simpler approach I shared with you and see if it makes a difference.",
    "1900773": "change loss function doesn't work, I even rebuild a super simple loss function and the final result doesn't change",
    "1901105": "I see the potential improvement in this. But did you check it's performance in practice? Does it improve the score in comarison to not using it?\n\nThe Devastator.",
    "1901460": "I have already answered this question on a limited set of 2 examples (see earlier in this thread), and can now extend it to about half a dozen examples. This objective function doesn't converge to a better score, and that also goes for its [**updated version**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578). Five-fold CV AmEx scores are essentially the same: 0.797648 with binary and 0.797577 with this function. Both score 0.799 on the LB but the former has better score when sorted.\n\nI think the main advantage is better speed during the hyperparameter search. Full DART run with as single set of parameters and binary objective function (5-fold) takes about 10 hours, compared to 8 hours with weighted loss objective function. So I am using it exclusively for that purpose.",
    "1903133": "jpison Thanks for sharing! So this did you find the improvement you were aiming for with this function? https://www.kaggle.com/competitions/amex-default-prediction/discussion/344578",
    "1903172": "Thank you for sharing this.",
    "1907022": "thanks for your help",
    "1910010": "Thanks for sharing!",
    "1911287": "Be careful in changing your optimization objective. When you capture 1% improvement of metric A, you may obtain a 5% decrease of metric B."
  },
  "source": "meta"
}