{
  "id": 85198,
  "title": "Cat Boost Efficiency",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85198",
  "author_name": "",
  "post_date": "2019-03-22T08:22:44.275252400Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Results are a bit wow since I was certain that I have failed due to being late with submitting the revised CatBoost solution.\nMy model was based on forking Jun Koda's approach with a slight modification of the percentile parameters. \nI have tried various approaches, including adding partial autocorrelation features and features, representing correlations between percentiles, but they didn't improve the score (public LB based one).\nThen I have started looking at classifier tuning and was surprised that despite experiment with Bayesian Optimisation I was not able to come close to the Cat Boost efficiency.\nI've done Cat Boost rounds and they have resulted in a huge score gain for the best params (like 740 vs baseline 706 for CV) but due to a mistake in committed one code did not compile into a submission before the deadline).</p>\n\n<p>Now the Question Number One is: why was Cat Boost so robust in this competitions, looking like it tops even Deep Learning?\nMy preliminary conclusion is that it's thanks to CBs unique handling of a lesser known issue of Gradient Boosting Classifiers:\n<strong>Prediction Shift.</strong>\nComplicated Math aside it basically happens due to inadvertent data leakage that occurs when we use the gradient to learn \"where to go\" in the next iteration of the procedure. Baseline gradient usage in both lgbm and xgboost(?) classifiers results in the shift of the conditional distribution used in the optimisation from the test to train one which introduces data leakage.\nWhile developing CatBoost Yandex team has looked into this problem and came up with a solution, that's called <a href=\"https://arxiv.org/pdf/1706.09516.pdf\">ordered boosting</a> that involves the permutation of the training samples when calculating the residuals for the gradient estimation.</p>\n\n<p>In the case of VSB powerline discharge dataset we have a large degree of non-uniformity within datasets, that's introduced both by the presence of different phases and various noisy signals, that are disparate in nature. All of this combined makes the normally mild Prediction Leakage issue much more severe and the gains from CatBoost more pronounced.</p>",
  "messages": [
    {
      "id": "496457",
      "postDate": "03/22/2019 08:22:44",
      "content": "<p>Results are a bit wow since I was certain that I have failed due to being late with submitting the revised CatBoost solution.\nMy model was based on forking Jun Koda's approach with a slight modification of the percentile parameters. \nI have tried various approaches, including adding partial autocorrelation features and features, representing correlations between percentiles, but they didn't improve the score (public LB based one).\nThen I have started looking at classifier tuning and was surprised that despite experiment with Bayesian Optimisation I was not able to come close to the Cat Boost efficiency.\nI've done Cat Boost rounds and they have resulted in a huge score gain for the best params (like 740 vs baseline 706 for CV) but due to a mistake in committed one code did not compile into a submission before the deadline).</p>\n\n<p>Now the Question Number One is: why was Cat Boost so robust in this competitions, looking like it tops even Deep Learning?\nMy preliminary conclusion is that it's thanks to CBs unique handling of a lesser known issue of Gradient Boosting Classifiers:\n<strong>Prediction Shift.</strong>\nComplicated Math aside it basically happens due to inadvertent data leakage that occurs when we use the gradient to learn \"where to go\" in the next iteration of the procedure. Baseline gradient usage in both lgbm and xgboost(?) classifiers results in the shift of the conditional distribution used in the optimisation from the test to train one which introduces data leakage.\nWhile developing CatBoost Yandex team has looked into this problem and came up with a solution, that's called <a href=\"https://arxiv.org/pdf/1706.09516.pdf\">ordered boosting</a> that involves the permutation of the training samples when calculating the residuals for the gradient estimation.</p>\n\n<p>In the case of VSB powerline discharge dataset we have a large degree of non-uniformity within datasets, that's introduced both by the presence of different phases and various noisy signals, that are disparate in nature. All of this combined makes the normally mild Prediction Leakage issue much more severe and the gains from CatBoost more pronounced.</p>",
      "rawMarkdown": "Results are a bit wow since I was certain that I have failed due to being late with submitting the revised CatBoost solution.\nMy model was based on forking Jun Koda's approach with a slight modification of the percentile parameters. \nI have tried various approaches, including adding partial autocorrelation features and features, representing correlations between percentiles, but they didn't improve the score (public LB based one).\nThen I have started looking at classifier tuning and was surprised that despite experiment with Bayesian Optimisation I was not able to come close to the Cat Boost efficiency.\nI've done Cat Boost rounds and they have resulted in a huge score gain for the best params (like 740 vs baseline 706 for CV) but due to a mistake in committed one code did not compile into a submission before the deadline).\n\nNow the Question Number One is: why was Cat Boost so robust in this competitions, looking like it tops even Deep Learning?\nMy preliminary conclusion is that it's thanks to CBs unique handling of a lesser known issue of Gradient Boosting Classifiers:\n**Prediction Shift.**\nComplicated Math aside it basically happens due to inadvertent data leakage that occurs when we use the gradient to learn \"where to go\" in the next iteration of the procedure. Baseline gradient usage in both lgbm and xgboost(?) classifiers results in the shift of the conditional distribution used in the optimisation from the test to train one which introduces data leakage.\nWhile developing CatBoost Yandex team has looked into this problem and came up with a solution, that's called [ordered boosting](https://arxiv.org/pdf/1706.09516.pdf) that involves the permutation of the training samples when calculating the residuals for the gradient estimation.\n\nIn the case of VSB powerline discharge dataset we have a large degree of non-uniformity within datasets, that's introduced both by the presence of different phases and various noisy signals, that are disparate in nature. All of this combined makes the normally mild Prediction Leakage issue much more severe and the gains from CatBoost more pronounced.",
      "votes": null
    },
    {
      "id": "496557",
      "postDate": "03/22/2019 10:34:13",
      "content": "<p>Thanks for the interesting hypothesis George.   Do you have any specific head-to-head examples to share?  Ideal I guess would be local cross validation, public, and private scores for catboost, xgboost, and lightgbm using the exact same data and each one tuned reasonably well.     I did a few runs locally using 5 reps x 5 different CV folds and my local cv ordering is xgboost (~0.74) &gt; lightgbm (~0.72) &gt; catboost (~0.70), but I haven't yet checked lb scores and looking back at Jun Koda's kernel I may have set my catboost learning rate too high (0.02).     Side question:  given the prediction shift difference, is it generally better to use smaller learning rates in catboost?</p>",
      "rawMarkdown": "Thanks for the interesting hypothesis George.   Do you have any specific head-to-head examples to share?  Ideal I guess would be local cross validation, public, and private scores for catboost, xgboost, and lightgbm using the exact same data and each one tuned reasonably well.     I did a few runs locally using 5 reps x 5 different CV folds and my local cv ordering is xgboost (~0.74) &gt; lightgbm (~0.72) &gt; catboost (~0.70), but I haven't yet checked lb scores and looking back at Jun Koda's kernel I may have set my catboost learning rate too high (0.02).     Side question:  given the prediction shift difference, is it generally better to use smaller learning rates in catboost?",
      "votes": null
    },
    {
      "id": "496571",
      "postDate": "03/22/2019 10:55:08",
      "content": "<p>Congrats <a href=\"/george1986\">@george1986</a>, I agree that catboost faired well with this data.</p>\n\n<p>My best model on hindsight is a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">CatBoost model with private LB=0.655</a> but I did not choose it as one of my final submissions.</p>",
      "rawMarkdown": "Congrats @george1986, I agree that catboost faired well with this data.\n\nMy best model on hindsight is a [CatBoost model with private LB=0.655](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) but I did not choose it as one of my final submissions.",
      "votes": null
    },
    {
      "id": "496574",
      "postDate": "03/22/2019 10:57:10",
      "content": "<p>Congrats <a href=\"/sasrdw\">@sasrdw</a>, check <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">my post for the full numbers</a> as well as a link to my catboost kernel.</p>",
      "rawMarkdown": "Congrats @sasrdw, check [my post for the full numbers](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) as well as a link to my catboost kernel.",
      "votes": null
    },
    {
      "id": "496577",
      "postDate": "03/22/2019 11:02:00",
      "content": "<p>Thanks YaGana.  If you get the chance, please try the same kernel but swap in xgboost and/or lightgbm, do a little local tuning, then see what happens with LB scores.</p>",
      "rawMarkdown": "Thanks YaGana.  If you get the chance, please try the same kernel but swap in xgboost and/or lightgbm, do a little local tuning, then see what happens with LB scores.",
      "votes": null
    },
    {
      "id": "496590",
      "postDate": "03/22/2019 11:24:45",
      "content": "<p><a href=\"/sasrdw\">@sasrdw</a> , I did a well tuned LGBM with the same features during the competition and the numbers are not good i.e. public LB = 0.558, private LB = 0.566</p>",
      "rawMarkdown": "sasrdw , I did a well tuned LGBM with the same features during the competition and the numbers are not good i.e. public LB = 0.558, private LB = 0.566",
      "votes": null
    },
    {
      "id": "496609",
      "postDate": "03/22/2019 11:41:47",
      "content": "<p><a href=\"/sasrdw\">@sasrdw</a> - Hi Russ! I only base this on how my lgbm attempts (less than 0.69 CV on forked Jun Joda's kernel no matter the parameter settings) vs catboost ones that consistently scored 0.7-0.72 on CV.\nAs for the learning rate from my experience you could have set it higher but additionally had to tune other parameters. One working combination I have found had it even as high as 0.27!</p>",
      "rawMarkdown": "sasrdw - Hi Russ! I only base this on how my lgbm attempts (less than 0.69 CV on forked Jun Joda's kernel no matter the parameter settings) vs catboost ones that consistently scored 0.7-0.72 on CV.\nAs for the learning rate from my experience you could have set it higher but additionally had to tune other parameters. One working combination I have found had it even as high as 0.27!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496557,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "03/22/2019 10:34:13",
      "content": "<p>Thanks for the interesting hypothesis George.   Do you have any specific head-to-head examples to share?  Ideal I guess would be local cross validation, public, and private scores for catboost, xgboost, and lightgbm using the exact same data and each one tuned reasonably well.     I did a few runs locally using 5 reps x 5 different CV folds and my local cv ordering is xgboost (~0.74) &gt; lightgbm (~0.72) &gt; catboost (~0.70), but I haven't yet checked lb scores and looking back at Jun Koda's kernel I may have set my catboost learning rate too high (0.02).     Side question:  given the prediction shift difference, is it generally better to use smaller learning rates in catboost?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496574,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "03/22/2019 10:57:10",
          "content": "<p>Congrats <a href=\"/sasrdw\">@sasrdw</a>, check <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">my post for the full numbers</a> as well as a link to my catboost kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496577,
          "author_name": "sasrdw",
          "author_url": "",
          "post_date": "03/22/2019 11:02:00",
          "content": "<p>Thanks YaGana.  If you get the chance, please try the same kernel but swap in xgboost and/or lightgbm, do a little local tuning, then see what happens with LB scores.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496590,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "03/22/2019 11:24:45",
          "content": "<p><a href=\"/sasrdw\">@sasrdw</a> , I did a well tuned LGBM with the same features during the competition and the numbers are not good i.e. public LB = 0.558, private LB = 0.566</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496609,
          "author_name": "george1986",
          "author_url": "",
          "post_date": "03/22/2019 11:41:47",
          "content": "<p><a href=\"/sasrdw\">@sasrdw</a> - Hi Russ! I only base this on how my lgbm attempts (less than 0.69 CV on forked Jun Joda's kernel no matter the parameter settings) vs catboost ones that consistently scored 0.7-0.72 on CV.\nAs for the learning rate from my experience you could have set it higher but additionally had to tune other parameters. One working combination I have found had it even as high as 0.27!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496571,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/22/2019 10:55:08",
      "content": "<p>Congrats <a href=\"/george1986\">@george1986</a>, I agree that catboost faired well with this data.</p>\n\n<p>My best model on hindsight is a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">CatBoost model with private LB=0.655</a> but I did not choose it as one of my final submissions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496457": "Results are a bit wow since I was certain that I have failed due to being late with submitting the revised CatBoost solution.\nMy model was based on forking Jun Koda's approach with a slight modification of the percentile parameters. \nI have tried various approaches, including adding partial autocorrelation features and features, representing correlations between percentiles, but they didn't improve the score (public LB based one).\nThen I have started looking at classifier tuning and was surprised that despite experiment with Bayesian Optimisation I was not able to come close to the Cat Boost efficiency.\nI've done Cat Boost rounds and they have resulted in a huge score gain for the best params (like 740 vs baseline 706 for CV) but due to a mistake in committed one code did not compile into a submission before the deadline).\n\nNow the Question Number One is: why was Cat Boost so robust in this competitions, looking like it tops even Deep Learning?\nMy preliminary conclusion is that it's thanks to CBs unique handling of a lesser known issue of Gradient Boosting Classifiers:\n**Prediction Shift.**\nComplicated Math aside it basically happens due to inadvertent data leakage that occurs when we use the gradient to learn \"where to go\" in the next iteration of the procedure. Baseline gradient usage in both lgbm and xgboost(?) classifiers results in the shift of the conditional distribution used in the optimisation from the test to train one which introduces data leakage.\nWhile developing CatBoost Yandex team has looked into this problem and came up with a solution, that's called [ordered boosting](https://arxiv.org/pdf/1706.09516.pdf) that involves the permutation of the training samples when calculating the residuals for the gradient estimation.\n\nIn the case of VSB powerline discharge dataset we have a large degree of non-uniformity within datasets, that's introduced both by the presence of different phases and various noisy signals, that are disparate in nature. All of this combined makes the normally mild Prediction Leakage issue much more severe and the gains from CatBoost more pronounced.",
    "496557": "Thanks for the interesting hypothesis George.   Do you have any specific head-to-head examples to share?  Ideal I guess would be local cross validation, public, and private scores for catboost, xgboost, and lightgbm using the exact same data and each one tuned reasonably well.     I did a few runs locally using 5 reps x 5 different CV folds and my local cv ordering is xgboost (~0.74) &gt; lightgbm (~0.72) &gt; catboost (~0.70), but I haven't yet checked lb scores and looking back at Jun Koda's kernel I may have set my catboost learning rate too high (0.02).     Side question:  given the prediction shift difference, is it generally better to use smaller learning rates in catboost?",
    "496571": "Congrats @george1986, I agree that catboost faired well with this data.\n\nMy best model on hindsight is a [CatBoost model with private LB=0.655](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) but I did not choose it as one of my final submissions.",
    "496574": "Congrats @sasrdw, check [my post for the full numbers](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) as well as a link to my catboost kernel.",
    "496577": "Thanks YaGana.  If you get the chance, please try the same kernel but swap in xgboost and/or lightgbm, do a little local tuning, then see what happens with LB scores.",
    "496590": "sasrdw , I did a well tuned LGBM with the same features during the competition and the numbers are not good i.e. public LB = 0.558, private LB = 0.566",
    "496609": "sasrdw - Hi Russ! I only base this on how my lgbm attempts (less than 0.69 CV on forked Jun Joda's kernel no matter the parameter settings) vs catboost ones that consistently scored 0.7-0.72 on CV.\nAs for the learning rate from my experience you could have set it higher but additionally had to tune other parameters. One working combination I have found had it even as high as 0.27!"
  },
  "source": "meta"
}