{
  "id": 71328,
  "title": "How LightGBM and XGBoost calculate Hessian?",
  "url": "/competitions/PLAsTiCC-2018/discussion/71328",
  "author_name": "",
  "post_date": "2018-11-12T16:26:08.475342100Z",
  "votes": 6,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I  was trying to implement a customer objective function for the competition. As far as I understand, both LightGBM and XGBoost requires the diagonal part of the Hessian matrix, which is \n\nwhere k is the row index, and i is the column index, and l_k is the loss from the k-th row. In the log loss, this term seems to be \n\nwithout considering weights. But when I check the implementation of LightGBM and XGBoost, they calculate the hessian part using <code>2.0*p*(1.0-p)</code>.</p>\n\n<p>So there is a factor of <strong>2</strong> here. I must be missing something here. Could anyone kindly indicate where am I wrong? Thanks a lot!</p>",
  "messages": [
    {
      "id": "419855",
      "postDate": "11/12/2018 16:26:08",
      "content": "<p>I  was trying to implement a customer objective function for the competition. As far as I understand, both LightGBM and XGBoost requires the diagonal part of the Hessian matrix, which is \n\nwhere k is the row index, and i is the column index, and l_k is the loss from the k-th row. In the log loss, this term seems to be \n\nwithout considering weights. But when I check the implementation of LightGBM and XGBoost, they calculate the hessian part using <code>2.0*p*(1.0-p)</code>.</p>\n\n<p>So there is a factor of <strong>2</strong> here. I must be missing something here. Could anyone kindly indicate where am I wrong? Thanks a lot!</p>",
      "rawMarkdown": "I  was trying to implement a customer objective function for the competition. As far as I understand, both LightGBM and XGBoost requires the diagonal part of the Hessian matrix, which is \n<a href=\"https://www.codecogs.com/eqnedit.php?latex=\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}\" target=\"_blank\"><img src=\"https://latex.codecogs.com/gif.latex?\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}\"></a>\nwhere k is the row index, and i is the column index, and l_k is the loss from the k-th row. In the log loss, this term seems to be \n<a href=\"https://www.codecogs.com/eqnedit.php?latex=\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}=p_{k,i}(1-p_{k,i}),\" target=\"_blank\"><img src=\"https://latex.codecogs.com/gif.latex?\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}=p_{k,i}(1-p_{k,i}),\"></a>\nwithout considering weights. But when I check the implementation of LightGBM and XGBoost, they calculate the hessian part using `2.0*p*(1.0-p)`.\n\nSo there is a factor of **2** here. I must be missing something here. Could anyone kindly indicate where am I wrong? Thanks a lot!",
      "votes": null
    },
    {
      "id": "433892",
      "postDate": "12/05/2018 16:09:34",
      "content": "<p>Hi lucaskg, did you find out how to  implement Hessian matrix in LightGBM, if you can guide me how to calculate and implement Hessian matrix for current loss function in LightGBM. \nThanks!!</p>",
      "rawMarkdown": "Hi lucaskg, did you find out how to  implement Hessian matrix in LightGBM, if you can guide me how to calculate and implement Hessian matrix for current loss function in LightGBM. \nThanks!!",
      "votes": null
    },
    {
      "id": "434024",
      "postDate": "12/05/2018 19:41:41",
      "content": "<p>seems like theres a kernel already implemented that? <a href=\"https://www.kaggle.com/mithrillion/know-your-objective\">https://www.kaggle.com/mithrillion/know-your-objective</a></p>",
      "rawMarkdown": "seems like theres a kernel already implemented that? https://www.kaggle.com/mithrillion/know-your-objective",
      "votes": null
    },
    {
      "id": "434032",
      "postDate": "12/05/2018 19:56:48",
      "content": "<p>This kernel sets the hessian to 1. However you can calculate the hessian yourself by taking the derivative of the gradient again. LGB and XGB only need the diagonal of the hessian, not the full hessian matrix.</p>",
      "rawMarkdown": "This kernel sets the hessian to 1. However you can calculate the hessian yourself by taking the derivative of the gradient again. LGB and XGB only need the diagonal of the hessian, not the full hessian matrix.",
      "votes": null
    },
    {
      "id": "434165",
      "postDate": "12/06/2018 01:45:11",
      "content": "<p>Did you ever find a reference for this? Multiplying my hessian by a factor of 2 improved my CV by a small amount</p>",
      "rawMarkdown": "Did you ever find a reference for this? Multiplying my hessian by a factor of 2 improved my CV by a small amount",
      "votes": null
    },
    {
      "id": "434303",
      "postDate": "12/06/2018 07:16:03",
      "content": "<p>Hi Jonny, does calculating Hessian Analytically is improving the CV score ?</p>",
      "rawMarkdown": "Hi Jonny, does calculating Hessian Analytically is improving the CV score ?",
      "votes": null
    },
    {
      "id": "434327",
      "postDate": "12/06/2018 07:50:03",
      "content": "<p>Maybe this link is the answer. \n<a href=\"https://github.com/Microsoft/LightGBM/issues/1052\">https://github.com/Microsoft/LightGBM/issues/1052</a></p>\n\n<p>We used correct hessian, but CV score does'nt improve much... .  We may be mistaking something, but We don't yet understand it...</p>",
      "rawMarkdown": "Maybe this link is the answer. \nhttps://github.com/Microsoft/LightGBM/issues/1052\n\nWe used correct hessian, but CV score does'nt improve much... .  We may be mistaking something, but We don't yet understand it...",
      "votes": null
    },
    {
      "id": "434425",
      "postDate": "12/06/2018 11:23:54",
      "content": "<p>Yes, I found an improvement in both CV and LB score. I have improved my model a lot since then, I don't know if the difference would be bigger or smaller now.</p>",
      "rawMarkdown": "Yes, I found an improvement in both CV and LB score. I have improved my model a lot since then, I don't know if the difference would be bigger or smaller now.",
      "votes": null
    },
    {
      "id": "434431",
      "postDate": "12/06/2018 11:37:30",
      "content": "<p>Keep in mind parameters like min child weight and the L1/L2 regularization values need to be tuned again after changing the hessian calculation, which could increase your improvement a little</p>",
      "rawMarkdown": "Keep in mind parameters like min child weight and the L1/L2 regularization values need to be tuned again after changing the hessian calculation, which could increase your improvement a little",
      "votes": null
    },
    {
      "id": "434452",
      "postDate": "12/06/2018 12:29:03",
      "content": "<p>Thanks a lot! We tried parameter tuning again , and CV score improved. Thanks for the helpful advice!</p>",
      "rawMarkdown": "Thanks a lot! We tried parameter tuning again , and CV score improved. Thanks for the helpful advice!",
      "votes": null
    },
    {
      "id": "434509",
      "postDate": "12/06/2018 14:20:49",
      "content": "<p>Hi Jonny, did you finally got it working with hard-coded Hessian? I tried autograd as well as manually implemented Hessian, but it never worked properly as expected. In the kernel I said that min_sum_hessian_in_leaf affects how hessians will change your fitted trees, but I still frequently got nonsensical results even after disabling it, so the problem is not really solved for me. I'd be really interested to know if you have successfully implemented the full objective function because I cannot really figure out what went wrong for me.</p>\n\n<p>Also, the way the XGB paper describes how they used Hessian seem to suggest that scale of Hessian should not matter, but why is 2x hessian giving better results?</p>",
      "rawMarkdown": "Hi Jonny, did you finally got it working with hard-coded Hessian? I tried autograd as well as manually implemented Hessian, but it never worked properly as expected. In the kernel I said that min_sum_hessian_in_leaf affects how hessians will change your fitted trees, but I still frequently got nonsensical results even after disabling it, so the problem is not really solved for me. I'd be really interested to know if you have successfully implemented the full objective function because I cannot really figure out what went wrong for me.\n\nAlso, the way the XGB paper describes how they used Hessian seem to suggest that scale of Hessian should not matter, but why is 2x hessian giving better results?",
      "votes": null
    },
    {
      "id": "434514",
      "postDate": "12/06/2018 14:28:20",
      "content": "<p>i used your kernal's objective function</p>\n\n<p><a href=\"https://www.kaggle.com/mithrillion/know-your-objective\">https://www.kaggle.com/mithrillion/know-your-objective</a></p>\n\n<p>but i am lost here, it would be great if you people can help me with objective function !!</p>",
      "rawMarkdown": "i used your kernal's objective function\n\nhttps://www.kaggle.com/mithrillion/know-your-objective\n\nbut i am lost here, it would be great if you people can help me with objective function !!",
      "votes": null
    },
    {
      "id": "434527",
      "postDate": "12/06/2018 14:39:40",
      "content": "<p>You have to be a bit careful with the objective function in my kernel. It works fairly well in basic stratified folds, but the normalisation factors for different classes are actually not constant. If you actually want a more stable objective function, you might have to pass in the class counts in the training set to the objective function instead of calculating the counts from the set being evaluated. I think it's actually more convenient if you only implement the metric function and use the weight trick in some other kernels together with multi_log_loss to achieve the same goal. If you want more control over the objective function, the objective in my kernel with the modification I mentioned above should do the trick. Unfortunately I still don't have a proper Hessian...</p>",
      "rawMarkdown": "You have to be a bit careful with the objective function in my kernel. It works fairly well in basic stratified folds, but the normalisation factors for different classes are actually not constant. If you actually want a more stable objective function, you might have to pass in the class counts in the training set to the objective function instead of calculating the counts from the set being evaluated. I think it's actually more convenient if you only implement the metric function and use the weight trick in some other kernels together with multi\\_log\\_loss to achieve the same goal. If you want more control over the objective function, the objective in my kernel with the modification I mentioned above should do the trick. Unfortunately I still don't have a proper Hessian...",
      "votes": null
    },
    {
      "id": "434541",
      "postDate": "12/06/2018 14:58:50",
      "content": "<p>Using your kernel as a guide I was able to manually implement the gradient and hessian calculations. For a long time I was stuck because my training loss didn't decrease at all, so I thought I did something wrong. Then I realized L1 regularization of 0.1 was too high relative to the  new size of the hessian, so I set it to zero and my model improved. I think it is because I correctly implemented the objective, but maybe it's because L1 regularization was hurting.</p>\n\n<p>I will try to make a small kernel about it when I get bored of feature engineering, maybe Saturday.</p>",
      "rawMarkdown": "Using your kernel as a guide I was able to manually implement the gradient and hessian calculations. For a long time I was stuck because my training loss didn't decrease at all, so I thought I did something wrong. Then I realized L1 regularization of 0.1 was too high relative to the  new size of the hessian, so I set it to zero and my model improved. I think it is because I correctly implemented the objective, but maybe it's because L1 regularization was hurting.\n\nI will try to make a small kernel about it when I get bored of feature engineering, maybe Saturday.",
      "votes": null
    },
    {
      "id": "434598",
      "postDate": "12/06/2018 16:57:58",
      "content": "<p>I'm still new to xgb/lgbm, this is probably a silly question. Why do people still try to implement the customized loss function (which requires output a gradient and hessian), when you can just pass in the desired class weights to the 'class_weight' parameter in the (sklearn api) LGBMClassifier with 'multiclass' as the objective? </p>",
      "rawMarkdown": "I'm still new to xgb/lgbm, this is probably a silly question. Why do people still try to implement the customized loss function (which requires output a gradient and hessian), when you can just pass in the desired class weights to the 'class_weight' parameter in the (sklearn api) LGBMClassifier with 'multiclass' as the objective?",
      "votes": null
    },
    {
      "id": "434612",
      "postDate": "12/06/2018 17:15:29",
      "content": "<p>Principled reason: I don't think the objective function is exactly equivalent to the multiclass objective with sample weights (maybe I am wrong), and we should get a better score if we can directly minimize what we are evaluated by.</p>\n\n<p>Honest reason: I will try anything if there is some hint on the discussion forum or in kernels that it improved someone's score, even if I don't really understand why it helps :)</p>",
      "rawMarkdown": "Principled reason: I don't think the objective function is exactly equivalent to the multiclass objective with sample weights (maybe I am wrong), and we should get a better score if we can directly minimize what we are evaluated by.\n\nHonest reason: I will try anything if there is some hint on the discussion forum or in kernels that it improved someone's score, even if I don't really understand why it helps :)",
      "votes": null
    },
    {
      "id": "434623",
      "postDate": "12/06/2018 17:28:58",
      "content": "<p>I guess there are some benefits to a custom objective that you cannot get from using the weights, like changing the cutoff point of probability clipping, using the logsumexp trick to stabilise the calculation, etc. Nothing major, but could be helpful.</p>",
      "rawMarkdown": "I guess there are some benefits to a custom objective that you cannot get from using the weights, like changing the cutoff point of probability clipping, using the logsumexp trick to stabilise the calculation, etc. Nothing major, but could be helpful.",
      "votes": null
    },
    {
      "id": "434853",
      "postDate": "12/07/2018 03:52:50",
      "content": "<p>A followup question, do you know how does lightgbm implement the multiclass log_loss? I did a quick search online and looked at its code on github but seems like didn't find a clear place. </p>",
      "rawMarkdown": "A followup question, do you know how does lightgbm implement the multiclass log_loss? I did a quick search online and looked at its code on github but seems like didn't find a clear place.",
      "votes": null
    },
    {
      "id": "434952",
      "postDate": "12/07/2018 07:54:04",
      "content": "<p>The code is in c++ but the relevant part is fairly easy to understand:\n<code>\n        Common::Softmax(&amp;rec);\n        for (int k = 0; k &lt; num_class_; ++k) {\n          auto p = rec[k];\n          size_t idx = static_cast&lt;size_t&gt;(num_data_) * k + i;\n          if (label_int_[i] == k) {\n            gradients[idx] = static_cast&lt;score_t&gt;((p - 1.0f) * weights_[i]);\n          } else {\n            gradients[idx] = static_cast&lt;score_t&gt;((p) * weights_[i]);\n          }\n          hessians[idx] = static_cast&lt;score_t&gt;((2.0f * p * (1.0f - p))* weights_[i]);\n        }\n</code></p>",
      "rawMarkdown": "The code is in c++ but the relevant part is fairly easy to understand:\n```\n        Common::Softmax(&amp;rec);\n        for (int k = 0; k &lt; num_class_; ++k) {\n          auto p = rec[k];\n          size_t idx = static_cast",
      "votes": null
    },
    {
      "id": "437347",
      "postDate": "12/11/2018 18:46:56",
      "content": "<p>Thanks for all the replies, I have yet another late follower-up question. So I am trying some of the public kernels thats using sample weights (instead of implementing a customized loss function), the weights they gave is calculated as such:</p>\n\n<pre><code>w = y.value_counts()\nsample_weights = {i : np.sum(w) / w[i] for i in w.index}\n</code></pre>\n\n<p>why there is a <em>np.sum(w)</em>, shouldn't it just be <em>1/w[i]</em> according to the evaluation formula provided? </p>\n\n<pre><code> ... where N is the number of objects in the class set, ...\n</code></pre>\n\n<p>But if I just pass in <em>/w[i]</em> as the sample weights to lgbm fit, the loss (evaluated by Oliver's script) is a lot higher than the one with <em>np.sum(w)</em>, does anyone have an intuition of why is this?</p>",
      "rawMarkdown": "Thanks for all the replies, I have yet another late follower-up question. So I am trying some of the public kernels thats using sample weights (instead of implementing a customized loss function), the weights they gave is calculated as such:\n\n    w = y.value_counts()\n    sample_weights = {i : np.sum(w) / w[i] for i in w.index}\n\nwhy there is a *np.sum(w)*, shouldn't it just be *1/w[i]* according to the evaluation formula provided? \n\n     ... where N is the number of objects in the class set, ...\n\nBut if I just pass in */w[i]* as the sample weights to lgbm fit, the loss (evaluated by Oliver's script) is a lot higher than the one with *np.sum(w)*, does anyone have an intuition of why is this?",
      "votes": null
    },
    {
      "id": "437473",
      "postDate": "12/12/2018 01:36:18",
      "content": "<p>I am missing something here, my model seems it is not learning when using my custom objective function</p>",
      "rawMarkdown": "I am missing something here, my model seems it is not learning when using my custom objective function",
      "votes": null
    },
    {
      "id": "437627",
      "postDate": "12/12/2018 08:14:22",
      "content": "<p>Happens the same to me. It performs worse when using the custom objective; I'm missing something as well...</p>",
      "rawMarkdown": "Happens the same to me. It performs worse when using the custom objective; I'm missing something as well...",
      "votes": null
    },
    {
      "id": "437689",
      "postDate": "12/12/2018 10:10:30",
      "content": "<p>There are many tricky issues when implementing custom objective. Here is a list of roadblocks that took away a few hours of my life:\n1. 'F(ortran)' order vs 'C' order regarding in input and output of objective functions\n2. loss function scale (mean vs sum and how they affect learning rate)\n3. Hessian scale and <code>min_sum_hessian_in_leaf</code>\n4. Per-batch <code>N</code> normalisation factor vs static <code>N</code> based on the training set (should only really matter when doing minibatch or unstratified folds)</p>",
      "rawMarkdown": "There are many tricky issues when implementing custom objective. Here is a list of roadblocks that took away a few hours of my life:\n1. 'F(ortran)' order vs 'C' order regarding in input and output of objective functions\n2. loss function scale (mean vs sum and how they affect learning rate)\n3. Hessian scale and `min_sum_hessian_in_leaf`\n4. Per-batch `N` normalisation factor vs static `N` based on the training set (should only really matter when doing minibatch or unstratified folds)",
      "votes": null
    },
    {
      "id": "438003",
      "postDate": "12/12/2018 23:30:38",
      "content": "<p>i manually calculated gradient and hessian. the loss is decreasing when hessian is set to np.ones but when i implement the second derivative equation the loss is at 0 value always</p>",
      "rawMarkdown": "i manually calculated gradient and hessian. the loss is decreasing when hessian is set to np.ones but when i implement the second derivative equation the loss is at 0 value always",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 433892,
      "author_name": "mks2192",
      "author_url": "",
      "post_date": "12/05/2018 16:09:34",
      "content": "<p>Hi lucaskg, did you find out how to  implement Hessian matrix in LightGBM, if you can guide me how to calculate and implement Hessian matrix for current loss function in LightGBM. \nThanks!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 434024,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "12/05/2018 19:41:41",
          "content": "<p>seems like theres a kernel already implemented that? <a href=\"https://www.kaggle.com/mithrillion/know-your-objective\">https://www.kaggle.com/mithrillion/know-your-objective</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434032,
          "author_name": "jlomond",
          "author_url": "",
          "post_date": "12/05/2018 19:56:48",
          "content": "<p>This kernel sets the hessian to 1. However you can calculate the hessian yourself by taking the derivative of the gradient again. LGB and XGB only need the diagonal of the hessian, not the full hessian matrix.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434303,
          "author_name": "mks2192",
          "author_url": "",
          "post_date": "12/06/2018 07:16:03",
          "content": "<p>Hi Jonny, does calculating Hessian Analytically is improving the CV score ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434425,
          "author_name": "jlomond",
          "author_url": "",
          "post_date": "12/06/2018 11:23:54",
          "content": "<p>Yes, I found an improvement in both CV and LB score. I have improved my model a lot since then, I don't know if the difference would be bigger or smaller now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434509,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/06/2018 14:20:49",
          "content": "<p>Hi Jonny, did you finally got it working with hard-coded Hessian? I tried autograd as well as manually implemented Hessian, but it never worked properly as expected. In the kernel I said that min_sum_hessian_in_leaf affects how hessians will change your fitted trees, but I still frequently got nonsensical results even after disabling it, so the problem is not really solved for me. I'd be really interested to know if you have successfully implemented the full objective function because I cannot really figure out what went wrong for me.</p>\n\n<p>Also, the way the XGB paper describes how they used Hessian seem to suggest that scale of Hessian should not matter, but why is 2x hessian giving better results?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434514,
          "author_name": "mks2192",
          "author_url": "",
          "post_date": "12/06/2018 14:28:20",
          "content": "<p>i used your kernal's objective function</p>\n\n<p><a href=\"https://www.kaggle.com/mithrillion/know-your-objective\">https://www.kaggle.com/mithrillion/know-your-objective</a></p>\n\n<p>but i am lost here, it would be great if you people can help me with objective function !!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434527,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/06/2018 14:39:40",
          "content": "<p>You have to be a bit careful with the objective function in my kernel. It works fairly well in basic stratified folds, but the normalisation factors for different classes are actually not constant. If you actually want a more stable objective function, you might have to pass in the class counts in the training set to the objective function instead of calculating the counts from the set being evaluated. I think it's actually more convenient if you only implement the metric function and use the weight trick in some other kernels together with multi_log_loss to achieve the same goal. If you want more control over the objective function, the objective in my kernel with the modification I mentioned above should do the trick. Unfortunately I still don't have a proper Hessian...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434541,
          "author_name": "jlomond",
          "author_url": "",
          "post_date": "12/06/2018 14:58:50",
          "content": "<p>Using your kernel as a guide I was able to manually implement the gradient and hessian calculations. For a long time I was stuck because my training loss didn't decrease at all, so I thought I did something wrong. Then I realized L1 regularization of 0.1 was too high relative to the  new size of the hessian, so I set it to zero and my model improved. I think it is because I correctly implemented the objective, but maybe it's because L1 regularization was hurting.</p>\n\n<p>I will try to make a small kernel about it when I get bored of feature engineering, maybe Saturday.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434598,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "12/06/2018 16:57:58",
          "content": "<p>I'm still new to xgb/lgbm, this is probably a silly question. Why do people still try to implement the customized loss function (which requires output a gradient and hessian), when you can just pass in the desired class weights to the 'class_weight' parameter in the (sklearn api) LGBMClassifier with 'multiclass' as the objective? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434612,
          "author_name": "jlomond",
          "author_url": "",
          "post_date": "12/06/2018 17:15:29",
          "content": "<p>Principled reason: I don't think the objective function is exactly equivalent to the multiclass objective with sample weights (maybe I am wrong), and we should get a better score if we can directly minimize what we are evaluated by.</p>\n\n<p>Honest reason: I will try anything if there is some hint on the discussion forum or in kernels that it improved someone's score, even if I don't really understand why it helps :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434623,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/06/2018 17:28:58",
          "content": "<p>I guess there are some benefits to a custom objective that you cannot get from using the weights, like changing the cutoff point of probability clipping, using the logsumexp trick to stabilise the calculation, etc. Nothing major, but could be helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434853,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "12/07/2018 03:52:50",
          "content": "<p>A followup question, do you know how does lightgbm implement the multiclass log_loss? I did a quick search online and looked at its code on github but seems like didn't find a clear place. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434952,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/07/2018 07:54:04",
          "content": "<p>The code is in c++ but the relevant part is fairly easy to understand:\n<code>\n        Common::Softmax(&amp;rec);\n        for (int k = 0; k &lt; num_class_; ++k) {\n          auto p = rec[k];\n          size_t idx = static_cast&lt;size_t&gt;(num_data_) * k + i;\n          if (label_int_[i] == k) {\n            gradients[idx] = static_cast&lt;score_t&gt;((p - 1.0f) * weights_[i]);\n          } else {\n            gradients[idx] = static_cast&lt;score_t&gt;((p) * weights_[i]);\n          }\n          hessians[idx] = static_cast&lt;score_t&gt;((2.0f * p * (1.0f - p))* weights_[i]);\n        }\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437347,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "12/11/2018 18:46:56",
          "content": "<p>Thanks for all the replies, I have yet another late follower-up question. So I am trying some of the public kernels thats using sample weights (instead of implementing a customized loss function), the weights they gave is calculated as such:</p>\n\n<pre><code>w = y.value_counts()\nsample_weights = {i : np.sum(w) / w[i] for i in w.index}\n</code></pre>\n\n<p>why there is a <em>np.sum(w)</em>, shouldn't it just be <em>1/w[i]</em> according to the evaluation formula provided? </p>\n\n<pre><code> ... where N is the number of objects in the class set, ...\n</code></pre>\n\n<p>But if I just pass in <em>/w[i]</em> as the sample weights to lgbm fit, the loss (evaluated by Oliver's script) is a lot higher than the one with <em>np.sum(w)</em>, does anyone have an intuition of why is this?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434165,
      "author_name": "jlomond",
      "author_url": "",
      "post_date": "12/06/2018 01:45:11",
      "content": "<p>Did you ever find a reference for this? Multiplying my hessian by a factor of 2 improved my CV by a small amount</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 434327,
      "author_name": "ykaneko1992",
      "author_url": "",
      "post_date": "12/06/2018 07:50:03",
      "content": "<p>Maybe this link is the answer. \n<a href=\"https://github.com/Microsoft/LightGBM/issues/1052\">https://github.com/Microsoft/LightGBM/issues/1052</a></p>\n\n<p>We used correct hessian, but CV score does'nt improve much... .  We may be mistaking something, but We don't yet understand it...</p>",
      "votes": null,
      "replies": [
        {
          "id": 434431,
          "author_name": "jlomond",
          "author_url": "",
          "post_date": "12/06/2018 11:37:30",
          "content": "<p>Keep in mind parameters like min child weight and the L1/L2 regularization values need to be tuned again after changing the hessian calculation, which could increase your improvement a little</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434452,
          "author_name": "ykaneko1992",
          "author_url": "",
          "post_date": "12/06/2018 12:29:03",
          "content": "<p>Thanks a lot! We tried parameter tuning again , and CV score improved. Thanks for the helpful advice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437473,
      "author_name": "niclasdoce",
      "author_url": "",
      "post_date": "12/12/2018 01:36:18",
      "content": "<p>I am missing something here, my model seems it is not learning when using my custom objective function</p>",
      "votes": null,
      "replies": [
        {
          "id": 437627,
          "author_name": "andreusancho",
          "author_url": "",
          "post_date": "12/12/2018 08:14:22",
          "content": "<p>Happens the same to me. It performs worse when using the custom objective; I'm missing something as well...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437689,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/12/2018 10:10:30",
          "content": "<p>There are many tricky issues when implementing custom objective. Here is a list of roadblocks that took away a few hours of my life:\n1. 'F(ortran)' order vs 'C' order regarding in input and output of objective functions\n2. loss function scale (mean vs sum and how they affect learning rate)\n3. Hessian scale and <code>min_sum_hessian_in_leaf</code>\n4. Per-batch <code>N</code> normalisation factor vs static <code>N</code> based on the training set (should only really matter when doing minibatch or unstratified folds)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438003,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "12/12/2018 23:30:38",
          "content": "<p>i manually calculated gradient and hessian. the loss is decreasing when hessian is set to np.ones but when i implement the second derivative equation the loss is at 0 value always</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "419855": "I  was trying to implement a customer objective function for the competition. As far as I understand, both LightGBM and XGBoost requires the diagonal part of the Hessian matrix, which is \n<a href=\"https://www.codecogs.com/eqnedit.php?latex=\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}\" target=\"_blank\"><img src=\"https://latex.codecogs.com/gif.latex?\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}\"></a>\nwhere k is the row index, and i is the column index, and l_k is the loss from the k-th row. In the log loss, this term seems to be \n<a href=\"https://www.codecogs.com/eqnedit.php?latex=\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}=p_{k,i}(1-p_{k,i}),\" target=\"_blank\"><img src=\"https://latex.codecogs.com/gif.latex?\\frac{\\partial^2&amp;space;l_{k}}{\\partial&amp;space;x_{k,i}^2}=p_{k,i}(1-p_{k,i}),\"></a>\nwithout considering weights. But when I check the implementation of LightGBM and XGBoost, they calculate the hessian part using `2.0*p*(1.0-p)`.\n\nSo there is a factor of **2** here. I must be missing something here. Could anyone kindly indicate where am I wrong? Thanks a lot!",
    "433892": "Hi lucaskg, did you find out how to  implement Hessian matrix in LightGBM, if you can guide me how to calculate and implement Hessian matrix for current loss function in LightGBM. \nThanks!!",
    "434024": "seems like theres a kernel already implemented that? https://www.kaggle.com/mithrillion/know-your-objective",
    "434032": "This kernel sets the hessian to 1. However you can calculate the hessian yourself by taking the derivative of the gradient again. LGB and XGB only need the diagonal of the hessian, not the full hessian matrix.",
    "434165": "Did you ever find a reference for this? Multiplying my hessian by a factor of 2 improved my CV by a small amount",
    "434303": "Hi Jonny, does calculating Hessian Analytically is improving the CV score ?",
    "434327": "Maybe this link is the answer. \nhttps://github.com/Microsoft/LightGBM/issues/1052\n\nWe used correct hessian, but CV score does'nt improve much... .  We may be mistaking something, but We don't yet understand it...",
    "434425": "Yes, I found an improvement in both CV and LB score. I have improved my model a lot since then, I don't know if the difference would be bigger or smaller now.",
    "434431": "Keep in mind parameters like min child weight and the L1/L2 regularization values need to be tuned again after changing the hessian calculation, which could increase your improvement a little",
    "434452": "Thanks a lot! We tried parameter tuning again , and CV score improved. Thanks for the helpful advice!",
    "434509": "Hi Jonny, did you finally got it working with hard-coded Hessian? I tried autograd as well as manually implemented Hessian, but it never worked properly as expected. In the kernel I said that min_sum_hessian_in_leaf affects how hessians will change your fitted trees, but I still frequently got nonsensical results even after disabling it, so the problem is not really solved for me. I'd be really interested to know if you have successfully implemented the full objective function because I cannot really figure out what went wrong for me.\n\nAlso, the way the XGB paper describes how they used Hessian seem to suggest that scale of Hessian should not matter, but why is 2x hessian giving better results?",
    "434514": "i used your kernal's objective function\n\nhttps://www.kaggle.com/mithrillion/know-your-objective\n\nbut i am lost here, it would be great if you people can help me with objective function !!",
    "434527": "You have to be a bit careful with the objective function in my kernel. It works fairly well in basic stratified folds, but the normalisation factors for different classes are actually not constant. If you actually want a more stable objective function, you might have to pass in the class counts in the training set to the objective function instead of calculating the counts from the set being evaluated. I think it's actually more convenient if you only implement the metric function and use the weight trick in some other kernels together with multi\\_log\\_loss to achieve the same goal. If you want more control over the objective function, the objective in my kernel with the modification I mentioned above should do the trick. Unfortunately I still don't have a proper Hessian...",
    "434541": "Using your kernel as a guide I was able to manually implement the gradient and hessian calculations. For a long time I was stuck because my training loss didn't decrease at all, so I thought I did something wrong. Then I realized L1 regularization of 0.1 was too high relative to the  new size of the hessian, so I set it to zero and my model improved. I think it is because I correctly implemented the objective, but maybe it's because L1 regularization was hurting.\n\nI will try to make a small kernel about it when I get bored of feature engineering, maybe Saturday.",
    "434598": "I'm still new to xgb/lgbm, this is probably a silly question. Why do people still try to implement the customized loss function (which requires output a gradient and hessian), when you can just pass in the desired class weights to the 'class_weight' parameter in the (sklearn api) LGBMClassifier with 'multiclass' as the objective?",
    "434612": "Principled reason: I don't think the objective function is exactly equivalent to the multiclass objective with sample weights (maybe I am wrong), and we should get a better score if we can directly minimize what we are evaluated by.\n\nHonest reason: I will try anything if there is some hint on the discussion forum or in kernels that it improved someone's score, even if I don't really understand why it helps :)",
    "434623": "I guess there are some benefits to a custom objective that you cannot get from using the weights, like changing the cutoff point of probability clipping, using the logsumexp trick to stabilise the calculation, etc. Nothing major, but could be helpful.",
    "434853": "A followup question, do you know how does lightgbm implement the multiclass log_loss? I did a quick search online and looked at its code on github but seems like didn't find a clear place.",
    "434952": "The code is in c++ but the relevant part is fairly easy to understand:\n```\n        Common::Softmax(&amp;rec);\n        for (int k = 0; k &lt; num_class_; ++k) {\n          auto p = rec[k];\n          size_t idx = static_cast",
    "437347": "Thanks for all the replies, I have yet another late follower-up question. So I am trying some of the public kernels thats using sample weights (instead of implementing a customized loss function), the weights they gave is calculated as such:\n\n    w = y.value_counts()\n    sample_weights = {i : np.sum(w) / w[i] for i in w.index}\n\nwhy there is a *np.sum(w)*, shouldn't it just be *1/w[i]* according to the evaluation formula provided? \n\n     ... where N is the number of objects in the class set, ...\n\nBut if I just pass in */w[i]* as the sample weights to lgbm fit, the loss (evaluated by Oliver's script) is a lot higher than the one with *np.sum(w)*, does anyone have an intuition of why is this?",
    "437473": "I am missing something here, my model seems it is not learning when using my custom objective function",
    "437627": "Happens the same to me. It performs worse when using the custom objective; I'm missing something as well...",
    "437689": "There are many tricky issues when implementing custom objective. Here is a list of roadblocks that took away a few hours of my life:\n1. 'F(ortran)' order vs 'C' order regarding in input and output of objective functions\n2. loss function scale (mean vs sum and how they affect learning rate)\n3. Hessian scale and `min_sum_hessian_in_leaf`\n4. Per-batch `N` normalisation factor vs static `N` based on the training set (should only really matter when doing minibatch or unstratified folds)",
    "438003": "i manually calculated gradient and hessian. the loss is decreasing when hessian is set to np.ones but when i implement the second derivative equation the loss is at 0 value always"
  },
  "source": "meta"
}