{
  "id": 181505,
  "title": "The challenge of directly using the competition loss function (when Hessian is needed)",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/181505",
  "author_name": "Björn",
  "post_date": "2020-09-09T06:23:09.752000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>It clearly would be desirable to directly optimize for the competition loss function and to output both a point estimate and a confidence from a single model. With a different metric this is nicely done e.g. in <a href=\"https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric\" target=\"_blank\">this notebook</a>, but it would be nice to replace the Gaussian loss in that notebook. That loss is quadratic and cares more about extreme outliers than the competition loss. The modified Laplace loss function is essentially<br>\n$$ -\\sqrt{2} \\times \\text{MAE} / \\sigma - \\log(\\sqrt{2} \\times \\sigma) = -\\sqrt{2} \\times |\\text{residual}| / \\sigma - \\log(\\sqrt{2} \\times \\sigma),$$ <br>\nif you ignore the clipping (since residuals are usually &lt;&gt;70, I speculate that that's reasonable to do). </p>\n<p>However, as we know MAE is a problem for methods that require the diagonal entries of the Hessian (like LightGBM or xgboost) and this is still a problem with the modified version above, because of the constant Hessian diagonal entry you get when you differentiate by residual twice.</p>\n<p>So, why not replace MAE with pseudo-Huber loss with \\(\\delta=1\\), log-cosh-loss or fair-loss? From what I've tried, I've run into two issues:</p>\n<ol>\n<li>numerical issues at large residuals (either producing NaNs or at least Hessians becoming constant up to machine precision so that often all predictions ended up being just 0) OR</li>\n<li>I changed the scale to litres instead of mL (i.e. divide FVC values by 1000), in which case I got models to fit and got sort of sensible predictions. However, then I was mostly in the region of residuals&lt;1, where these loss functions behave more quadratically (i.e. less like MAE and more like MSE), which was what I wanted to get away from in the first place.</li>\n</ol>\n<p>Of course, then you wonder how LightGBM and xgboost implement MAE, pseuo-Huber and quantile loss. It seems like that requires a bit more than just defining a loss function with gradients and diagnoal entries of the Hessian (see e.g. <a href=\"https://stackoverflow.com/questions/55793947/implement-custom-huber-loss-in-lightgbm\" target=\"_blank\">the detailed answer to this question</a>, as well as <a href=\"http://jmarkhou.com/lgbqr/\" target=\"_blank\">this blog post</a>).</p>\n<p>Am I overlooking anything obvious here for making this work / ignoring known solutions to the issues I ran into? Or are there similar models that allow for custom loss functions, but are not/less affected by the Hessian (even if it then takes longer to fit them)?</p>",
  "messages": [
    {
      "id": 1003632,
      "postDate": "2020-09-09T06:23:09.753Z",
      "content": "<p>It clearly would be desirable to directly optimize for the competition loss function and to output both a point estimate and a confidence from a single model. With a different metric this is nicely done e.g. in <a href=\"https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric\" target=\"_blank\">this notebook</a>, but it would be nice to replace the Gaussian loss in that notebook. That loss is quadratic and cares more about extreme outliers than the competition loss. The modified Laplace loss function is essentially<br>\n$$ -\\sqrt{2} \\times \\text{MAE} / \\sigma - \\log(\\sqrt{2} \\times \\sigma) = -\\sqrt{2} \\times |\\text{residual}| / \\sigma - \\log(\\sqrt{2} \\times \\sigma),$$ <br>\nif you ignore the clipping (since residuals are usually &lt;&gt;70, I speculate that that's reasonable to do). </p>\n<p>However, as we know MAE is a problem for methods that require the diagonal entries of the Hessian (like LightGBM or xgboost) and this is still a problem with the modified version above, because of the constant Hessian diagonal entry you get when you differentiate by residual twice.</p>\n<p>So, why not replace MAE with pseudo-Huber loss with \\(\\delta=1\\), log-cosh-loss or fair-loss? From what I've tried, I've run into two issues:</p>\n<ol>\n<li>numerical issues at large residuals (either producing NaNs or at least Hessians becoming constant up to machine precision so that often all predictions ended up being just 0) OR</li>\n<li>I changed the scale to litres instead of mL (i.e. divide FVC values by 1000), in which case I got models to fit and got sort of sensible predictions. However, then I was mostly in the region of residuals&lt;1, where these loss functions behave more quadratically (i.e. less like MAE and more like MSE), which was what I wanted to get away from in the first place.</li>\n</ol>\n<p>Of course, then you wonder how LightGBM and xgboost implement MAE, pseuo-Huber and quantile loss. It seems like that requires a bit more than just defining a loss function with gradients and diagnoal entries of the Hessian (see e.g. <a href=\"https://stackoverflow.com/questions/55793947/implement-custom-huber-loss-in-lightgbm\" target=\"_blank\">the detailed answer to this question</a>, as well as <a href=\"http://jmarkhou.com/lgbqr/\" target=\"_blank\">this blog post</a>).</p>\n<p>Am I overlooking anything obvious here for making this work / ignoring known solutions to the issues I ran into? Or are there similar models that allow for custom loss functions, but are not/less affected by the Hessian (even if it then takes longer to fit them)?</p>",
      "rawMarkdown": "It clearly would be desirable to directly optimize for the competition loss function and to output both a point estimate and a confidence from a single model. With a different metric this is nicely done e.g. in [this notebook](https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric), but it would be nice to replace the Gaussian loss in that notebook. That loss is quadratic and cares more about extreme outliers than the competition loss. The modified Laplace loss function is essentially\n$$ -\\sqrt{2} \\times \\text{MAE} / \\sigma - \\log(\\sqrt{2} \\times \\sigma) = -\\sqrt{2} \\times |\\text{residual}| / \\sigma - \\log(\\sqrt{2} \\times \\sigma),$$ \nif you ignore the clipping (since residuals are usually <<1000 and estimated confidences are usually >>70, I speculate that that's reasonable to do). \n\nHowever, as we know MAE is a problem for methods that require the diagonal entries of the Hessian (like LightGBM or xgboost) and this is still a problem with the modified version above, because of the constant Hessian diagonal entry you get when you differentiate by residual twice.\n\nSo, why not replace MAE with pseudo-Huber loss with \\\\(\\delta=1\\\\), log-cosh-loss or fair-loss? From what I've tried, I've run into two issues:\n1. numerical issues at large residuals (either producing NaNs or at least Hessians becoming constant up to machine precision so that often all predictions ended up being just 0) OR\n2. I changed the scale to litres instead of mL (i.e. divide FVC values by 1000), in which case I got models to fit and got sort of sensible predictions. However, then I was mostly in the region of residuals<1, where these loss functions behave more quadratically (i.e. less like MAE and more like MSE), which was what I wanted to get away from in the first place.\n\nOf course, then you wonder how LightGBM and xgboost implement MAE, pseuo-Huber and quantile loss. It seems like that requires a bit more than just defining a loss function with gradients and diagnoal entries of the Hessian (see e.g. [the detailed answer to this question](https://stackoverflow.com/questions/55793947/implement-custom-huber-loss-in-lightgbm), as well as [this blog post](http://jmarkhou.com/lgbqr/)).\n\nAm I overlooking anything obvious here for making this work / ignoring known solutions to the issues I ran into? Or are there similar models that allow for custom loss functions, but are not/less affected by the Hessian (even if it then takes longer to fit them)?",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1003632": "It clearly would be desirable to directly optimize for the competition loss function and to output both a point estimate and a confidence from a single model. With a different metric this is nicely done e.g. in [this notebook](https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric), but it would be nice to replace the Gaussian loss in that notebook. That loss is quadratic and cares more about extreme outliers than the competition loss. The modified Laplace loss function is essentially\n$$ -\\sqrt{2} \\times \\text{MAE} / \\sigma - \\log(\\sqrt{2} \\times \\sigma) = -\\sqrt{2} \\times |\\text{residual}| / \\sigma - \\log(\\sqrt{2} \\times \\sigma),$$ \nif you ignore the clipping (since residuals are usually <<1000 and estimated confidences are usually >>70, I speculate that that's reasonable to do). \n\nHowever, as we know MAE is a problem for methods that require the diagonal entries of the Hessian (like LightGBM or xgboost) and this is still a problem with the modified version above, because of the constant Hessian diagonal entry you get when you differentiate by residual twice.\n\nSo, why not replace MAE with pseudo-Huber loss with \\\\(\\delta=1\\\\), log-cosh-loss or fair-loss? From what I've tried, I've run into two issues:\n1. numerical issues at large residuals (either producing NaNs or at least Hessians becoming constant up to machine precision so that often all predictions ended up being just 0) OR\n2. I changed the scale to litres instead of mL (i.e. divide FVC values by 1000), in which case I got models to fit and got sort of sensible predictions. However, then I was mostly in the region of residuals<1, where these loss functions behave more quadratically (i.e. less like MAE and more like MSE), which was what I wanted to get away from in the first place.\n\nOf course, then you wonder how LightGBM and xgboost implement MAE, pseuo-Huber and quantile loss. It seems like that requires a bit more than just defining a loss function with gradients and diagnoal entries of the Hessian (see e.g. [the detailed answer to this question](https://stackoverflow.com/questions/55793947/implement-custom-huber-loss-in-lightgbm), as well as [this blog post](http://jmarkhou.com/lgbqr/)).\n\nAm I overlooking anything obvious here for making this work / ignoring known solutions to the issues I ran into? Or are there similar models that allow for custom loss functions, but are not/less affected by the Hessian (even if it then takes longer to fit them)?"
  }
}