{
  "id": 337160,
  "title": "Tree correlation vs highly correlated features",
  "url": "/competitions/amex-default-prediction/discussion/337160",
  "author_name": "",
  "post_date": "2022-07-14T18:34:27.776437900Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>There's a known truth that highly correlated features can cause problems. I'm focusing on tree booster models, so let's simply say they can cause problems for these models. </p>\n<p>But why? I have a theory (okay, I have a hypothesis. You know what I mean). See, the issue isn't the correlation, the issue is the sampling rate of the model causing highly correlated trees. </p>\n<p>It's a great time to bring it up, because there's great examples of both extremes in this competition, and other ongoing discussions that relate to the topic. </p>\n<p>To simplify, let's pretend that a highly related feature pair - in the examples, engineered feature pair - is close enough to identical to cause problems (obviously 90% correlated will cause less problems than 100%). </p>\n<p>If colsample_tree = 0.88 like in the top public XGB model, then this duplicate feature has a whopping 98.6% chance of being considered by each tree. </p>\n<p>However, if you include the rounded version of EVERY float feature for '_last', and drop col sampling way down to 0.2 like in the top model, it's self evident you have tons of highly correlated features, but only a 36% chance a tree will pick up on the duplicated feature!</p>\n<p>A discussion topic recently brought up including both last-mean and last-first. If you assume that some percentage of base columns have highly correlated first and mean values, then it's another example of duplication!</p>\n<p>Trying both at once of two similar features will probably hurt your results if sampling is &gt;.7, but at least <em>might</em> help if sampling &lt; .5. </p>\n<p>I think the reason rounding is effective for the top model is mainly to increase the chance the highly important \"last\" columns are selected by the tree, compared with other features. Slight diversity between the duplicate versions is just a bonus. </p>\n<p>XGB: high sampling rate: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a></p>\n<p>LGBM: low sampling rate: <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></p>",
  "messages": [
    {
      "id": "1855628",
      "postDate": "07/14/2022 18:34:27",
      "content": "<p>There's a known truth that highly correlated features can cause problems. I'm focusing on tree booster models, so let's simply say they can cause problems for these models. </p>\n<p>But why? I have a theory (okay, I have a hypothesis. You know what I mean). See, the issue isn't the correlation, the issue is the sampling rate of the model causing highly correlated trees. </p>\n<p>It's a great time to bring it up, because there's great examples of both extremes in this competition, and other ongoing discussions that relate to the topic. </p>\n<p>To simplify, let's pretend that a highly related feature pair - in the examples, engineered feature pair - is close enough to identical to cause problems (obviously 90% correlated will cause less problems than 100%). </p>\n<p>If colsample_tree = 0.88 like in the top public XGB model, then this duplicate feature has a whopping 98.6% chance of being considered by each tree. </p>\n<p>However, if you include the rounded version of EVERY float feature for '_last', and drop col sampling way down to 0.2 like in the top model, it's self evident you have tons of highly correlated features, but only a 36% chance a tree will pick up on the duplicated feature!</p>\n<p>A discussion topic recently brought up including both last-mean and last-first. If you assume that some percentage of base columns have highly correlated first and mean values, then it's another example of duplication!</p>\n<p>Trying both at once of two similar features will probably hurt your results if sampling is &gt;.7, but at least <em>might</em> help if sampling &lt; .5. </p>\n<p>I think the reason rounding is effective for the top model is mainly to increase the chance the highly important \"last\" columns are selected by the tree, compared with other features. Slight diversity between the duplicate versions is just a bonus. </p>\n<p>XGB: high sampling rate: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a></p>\n<p>LGBM: low sampling rate: <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></p>",
      "rawMarkdown": "There's a known truth that highly correlated features can cause problems. I'm focusing on tree booster models, so let's simply say they can cause problems for these models. \n\nBut why? I have a theory (okay, I have a hypothesis. You know what I mean). See, the issue isn't the correlation, the issue is the sampling rate of the model causing highly correlated trees. \n\nIt's a great time to bring it up, because there's great examples of both extremes in this competition, and other ongoing discussions that relate to the topic. \n\nTo simplify, let's pretend that a highly related feature pair - in the examples, engineered feature pair - is close enough to identical to cause problems (obviously 90% correlated will cause less problems than 100%). \n\nIf colsample_tree = 0.88 like in the top public XGB model, then this duplicate feature has a whopping 98.6% chance of being considered by each tree. \n\nHowever, if you include the rounded version of EVERY float feature for '_last', and drop col sampling way down to 0.2 like in the top model, it's self evident you have tons of highly correlated features, but only a 36% chance a tree will pick up on the duplicated feature!\n\nA discussion topic recently brought up including both last-mean and last-first. If you assume that some percentage of base columns have highly correlated first and mean values, then it's another example of duplication!\n\nTrying both at once of two similar features will probably hurt your results if sampling is >.7, but at least *might* help if sampling < .5. \n\nI think the reason rounding is effective for the top model is mainly to increase the chance the highly important \"last\" columns are selected by the tree, compared with other features. Slight diversity between the duplicate versions is just a bonus. \n\nXGB: high sampling rate: https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\n\nLGBM: low sampling rate: https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977",
      "votes": null
    },
    {
      "id": "1855637",
      "postDate": "07/14/2022 18:37:40",
      "content": "<p>Good explanation! Keep it up</p>",
      "rawMarkdown": "Good explanation! Keep it up",
      "votes": null
    },
    {
      "id": "1855651",
      "postDate": "07/14/2022 19:02:05",
      "content": "<p><a href=\"https://medium.com/data-design/ensembles-of-tree-based-models-why-correlated-features-do-not-trip-them-and-why-na-matters-7658f4752e1b\" target=\"_blank\">Here</a> is an article by the great <a href=\"https://www.kaggle.com/laurae2\" target=\"_blank\">laruae</a> that seems to show that highly correlated features are no problem for ensembled tree methods such as xgboost or lightgbm.</p>",
      "rawMarkdown": "[Here](https://medium.com/data-design/ensembles-of-tree-based-models-why-correlated-features-do-not-trip-them-and-why-na-matters-7658f4752e1b) is an article by the great [laruae](https://www.kaggle.com/laurae2) that seems to show that highly correlated features are no problem for ensembled tree methods such as xgboost or lightgbm.",
      "votes": null
    },
    {
      "id": "1855763",
      "postDate": "07/14/2022 22:09:56",
      "content": "<p>Generally speaking, tree-based methods have no problem with correlated features. I think you are thinking about linear models, where indeed it is highly advisable to remove correlated features. For tree-based methods, they simply take the best split in whichever feature it may be, and they have no problem jumping between features in their decision-making process.</p>",
      "rawMarkdown": "Generally speaking, tree-based methods have no problem with correlated features. I think you are thinking about linear models, where indeed it is highly advisable to remove correlated features. For tree-based methods, they simply take the best split in whichever feature it may be, and they have no problem jumping between features in their decision-making process.",
      "votes": null
    },
    {
      "id": "1856239",
      "postDate": "07/15/2022 07:55:57",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>I agree with <a href=\"https://www.kaggle.com/chrisrichardmiles\" target=\"_blank\">@chrisrichardmiles</a>  and <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> in that with modern estimators highly correlated features is not a major problem. However it is worth mentioning that correlated features have the potential to play havoc with feature importance calculations (for example see <a href=\"https://www.kaggle.com/code/carlmcbrideellis/variance-inflation-factor-vif-and-explainability\" target=\"_blank\">\"Explainability, collinearity and the variance inflation factor (VIF)\"</a>).</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @roberthatch \n\nI agree with @chrisrichardmiles  and @tilii7 in that with modern estimators highly correlated features is not a major problem. However it is worth mentioning that correlated features have the potential to play havoc with feature importance calculations (for example see [\"Explainability, collinearity and the variance inflation factor (VIF)\"](https://www.kaggle.com/code/carlmcbrideellis/variance-inflation-factor-vif-and-explainability)).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1857089",
      "postDate": "07/15/2022 20:24:43",
      "content": "<p>Came here and was about to share this article, well I was slow 😬</p>",
      "rawMarkdown": "Came here and was about to share this article, well I was slow 😬",
      "votes": null
    },
    {
      "id": "1857123",
      "postDate": "07/15/2022 21:23:10",
      "content": "<p>Clearly I used the wrong analogy and comparison.</p>\n<p>My primary point is that I would expect adding rounding features to the XGB 88% sampling model to hurt the model. (aka to require adjusting the sampling percentage by a huge amount) and perhaps even vice versa, removing the rounding on the dart model might benefit from increasing the sampling by a lot.</p>\n<p>Am I wrong there, too? I didn't read the article, let me know if it touches on my main thought?</p>",
      "rawMarkdown": "Clearly I used the wrong analogy and comparison.\n\nMy primary point is that I would expect adding rounding features to the XGB 88% sampling model to hurt the model. (aka to require adjusting the sampling percentage by a huge amount) and perhaps even vice versa, removing the rounding on the dart model might benefit from increasing the sampling by a lot.\n\nAm I wrong there, too? I didn't read the article, let me know if it touches on my main thought?",
      "votes": null
    },
    {
      "id": "1857149",
      "postDate": "07/15/2022 22:12:56",
      "content": "<blockquote>\n  <p>Am I wrong there, too? I didn't read the article, let me know if it touches on my main thought?</p>\n</blockquote>\n<p>The article is not long, and I would imagine you'd want to read carefully something that addresses your point of view from this thread. A simple exercise in that article (column duplication) shows that XGB scores are unaffected. Intuitively, I would have thought the same even before reading the article, though there is always a possibility that the score would be affected if the number of correlated features is substantial, say 10% or so. By this I mean 10% of \"child\" features that are all correlated to a single \"parent\" feature. If we had 5% of total features, each of which had a feature doppelganger, I don't think that would matter. Even though that would again amount to a total of 10% correlated features, for the purpose of decision-making I think it potentially matters only when many columns are all correlated to each other, rather than when many columns have one other correlated column apiece.</p>",
      "rawMarkdown": "> Am I wrong there, too? I didn't read the article, let me know if it touches on my main thought?\n\nThe article is not long, and I would imagine you'd want to read carefully something that addresses your point of view from this thread. A simple exercise in that article (column duplication) shows that XGB scores are unaffected. Intuitively, I would have thought the same even before reading the article, though there is always a possibility that the score would be affected if the number of correlated features is substantial, say 10% or so. By this I mean 10% of \"child\" features that are all correlated to a single \"parent\" feature. If we had 5% of total features, each of which had a feature doppelganger, I don't think that would matter. Even though that would again amount to a total of 10% correlated features, for the purpose of decision-making I think it potentially matters only when many columns are all correlated to each other, rather than when many columns have one other correlated column apiece.",
      "votes": null
    },
    {
      "id": "1857158",
      "postDate": "07/15/2022 22:51:42",
      "content": "<p>I mean, XGBoost is highly robust to everything.</p>\n<p>\"Tree-based models have an innate feature of being robust to correlated features. \"</p>\n<p>To me, taking 20% of the features and doubling them is no where near the same as what's addressed in the article. I would expect XGBoost to be \"robust\" in that scenario… but that doesn't mean it's helpful. Is XGBoost robust against bad colsample values? Or do they cause a problem? In my very limited experience, I'd assume both statements are true. </p>\n<p>So to restate my view: adding (entire sets of) duplicate and near duplicate features will impact the 'true' sampling rate. This can be good or bad. That's all I'm trying to say. </p>\n<p>I didn't see anything contradictory in that article. </p>",
      "rawMarkdown": "I mean, XGBoost is highly robust to everything.\n\n\"Tree-based models have an innate feature of being robust to correlated features. \"\n\nTo me, taking 20% of the features and doubling them is no where near the same as what's addressed in the article. I would expect XGBoost to be \"robust\" in that scenario... but that doesn't mean it's helpful. Is XGBoost robust against bad colsample values? Or do they cause a problem? In my very limited experience, I'd assume both statements are true. \n\nSo to restate my view: adding (entire sets of) duplicate and near duplicate features will impact the 'true' sampling rate. This can be good or bad. That's all I'm trying to say. \n\nI didn't see anything contradictory in that article.",
      "votes": null
    },
    {
      "id": "1857161",
      "postDate": "07/15/2022 22:57:53",
      "content": "<p>More specifically, the duplicate column experiment in the article isn't setting colsample under 1!</p>\n<p>So yes, if it was guaranteed before, still guaranteed after. 1 - (1-1)^1 == 1 - (1-1)^2. Doesn't matter that you squared a term. </p>",
      "rawMarkdown": "More specifically, the duplicate column experiment in the article isn't setting colsample under 1!\n\nSo yes, if it was guaranteed before, still guaranteed after. 1 - (1-1)^1 == 1 - (1-1)^2. Doesn't matter that you squared a term.",
      "votes": null
    },
    {
      "id": "1857179",
      "postDate": "07/15/2022 23:26:00",
      "content": "<p>I have clearly said that the exercise in that article was simple, and wasn't trying to prove anything by it. You asked a question about the article, I gave you a summary.</p>\n<p>But since you are bringing up column sampling - gradient boosters can always pick one of the two correlated columns and be fine with it. Now, it is possible that the result would be slightly better with both correlated columns included, but that is why we do hyperparameter search to determine the best <code>colsample</code> value. A substandard colsample value is not necessarily that way because it eliminated a correlated column, but possibly because it eliminated another more useful column. In other words, correlated features would definitely not be my biggest worry when it comes to column sampling.</p>\n<p>Correlated columns by definition contain very similar information, so it is unlikely that excluding one of them will have a profound effect on the final result. I think that removing any other non-correlated column has greater potential of affecting the model, so that would be of greater concern for me. That of course brings us back to hyperparameter optimization where these things are found in a systematic way. </p>",
      "rawMarkdown": "I have clearly said that the exercise in that article was simple, and wasn't trying to prove anything by it. You asked a question about the article, I gave you a summary.\n\nBut since you are bringing up column sampling - gradient boosters can always pick one of the two correlated columns and be fine with it. Now, it is possible that the result would be slightly better with both correlated columns included, but that is why we do hyperparameter search to determine the best `colsample` value. A substandard colsample value is not necessarily that way because it eliminated a correlated column, but possibly because it eliminated another more useful column. In other words, correlated features would definitely not be my biggest worry when it comes to column sampling.\n\nCorrelated columns by definition contain very similar information, so it is unlikely that excluding one of them will have a profound effect on the final result. I think that removing any other non-correlated column has greater potential of affecting the model, so that would be of greater concern for me. That of course brings us back to hyperparameter optimization where these things are found in a systematic way.",
      "votes": null
    },
    {
      "id": "1857192",
      "postDate": "07/15/2022 23:57:23",
      "content": "<p>thanks for your explanation</p>",
      "rawMarkdown": "thanks for your explanation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1855637,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/14/2022 18:37:40",
      "content": "<p>Good explanation! Keep it up</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1855651,
      "author_name": "chrisrichardmiles",
      "author_url": "",
      "post_date": "07/14/2022 19:02:05",
      "content": "<p><a href=\"https://medium.com/data-design/ensembles-of-tree-based-models-why-correlated-features-do-not-trip-them-and-why-na-matters-7658f4752e1b\" target=\"_blank\">Here</a> is an article by the great <a href=\"https://www.kaggle.com/laurae2\" target=\"_blank\">laruae</a> that seems to show that highly correlated features are no problem for ensembled tree methods such as xgboost or lightgbm.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1857089,
          "author_name": "nitishraj",
          "author_url": "",
          "post_date": "07/15/2022 20:24:43",
          "content": "<p>Came here and was about to share this article, well I was slow 😬</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1857123,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/15/2022 21:23:10",
          "content": "<p>Clearly I used the wrong analogy and comparison.</p>\n<p>My primary point is that I would expect adding rounding features to the XGB 88% sampling model to hurt the model. (aka to require adjusting the sampling percentage by a huge amount) and perhaps even vice versa, removing the rounding on the dart model might benefit from increasing the sampling by a lot.</p>\n<p>Am I wrong there, too? I didn't read the article, let me know if it touches on my main thought?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1857149,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/15/2022 22:12:56",
          "content": "<blockquote>\n  <p>Am I wrong there, too? I didn't read the article, let me know if it touches on my main thought?</p>\n</blockquote>\n<p>The article is not long, and I would imagine you'd want to read carefully something that addresses your point of view from this thread. A simple exercise in that article (column duplication) shows that XGB scores are unaffected. Intuitively, I would have thought the same even before reading the article, though there is always a possibility that the score would be affected if the number of correlated features is substantial, say 10% or so. By this I mean 10% of \"child\" features that are all correlated to a single \"parent\" feature. If we had 5% of total features, each of which had a feature doppelganger, I don't think that would matter. Even though that would again amount to a total of 10% correlated features, for the purpose of decision-making I think it potentially matters only when many columns are all correlated to each other, rather than when many columns have one other correlated column apiece.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1857158,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/15/2022 22:51:42",
          "content": "<p>I mean, XGBoost is highly robust to everything.</p>\n<p>\"Tree-based models have an innate feature of being robust to correlated features. \"</p>\n<p>To me, taking 20% of the features and doubling them is no where near the same as what's addressed in the article. I would expect XGBoost to be \"robust\" in that scenario… but that doesn't mean it's helpful. Is XGBoost robust against bad colsample values? Or do they cause a problem? In my very limited experience, I'd assume both statements are true. </p>\n<p>So to restate my view: adding (entire sets of) duplicate and near duplicate features will impact the 'true' sampling rate. This can be good or bad. That's all I'm trying to say. </p>\n<p>I didn't see anything contradictory in that article. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1857161,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/15/2022 22:57:53",
          "content": "<p>More specifically, the duplicate column experiment in the article isn't setting colsample under 1!</p>\n<p>So yes, if it was guaranteed before, still guaranteed after. 1 - (1-1)^1 == 1 - (1-1)^2. Doesn't matter that you squared a term. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1857179,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/15/2022 23:26:00",
          "content": "<p>I have clearly said that the exercise in that article was simple, and wasn't trying to prove anything by it. You asked a question about the article, I gave you a summary.</p>\n<p>But since you are bringing up column sampling - gradient boosters can always pick one of the two correlated columns and be fine with it. Now, it is possible that the result would be slightly better with both correlated columns included, but that is why we do hyperparameter search to determine the best <code>colsample</code> value. A substandard colsample value is not necessarily that way because it eliminated a correlated column, but possibly because it eliminated another more useful column. In other words, correlated features would definitely not be my biggest worry when it comes to column sampling.</p>\n<p>Correlated columns by definition contain very similar information, so it is unlikely that excluding one of them will have a profound effect on the final result. I think that removing any other non-correlated column has greater potential of affecting the model, so that would be of greater concern for me. That of course brings us back to hyperparameter optimization where these things are found in a systematic way. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1855763,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "07/14/2022 22:09:56",
      "content": "<p>Generally speaking, tree-based methods have no problem with correlated features. I think you are thinking about linear models, where indeed it is highly advisable to remove correlated features. For tree-based methods, they simply take the best split in whichever feature it may be, and they have no problem jumping between features in their decision-making process.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1856239,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "07/15/2022 07:55:57",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>\n<p>I agree with <a href=\"https://www.kaggle.com/chrisrichardmiles\" target=\"_blank\">@chrisrichardmiles</a>  and <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> in that with modern estimators highly correlated features is not a major problem. However it is worth mentioning that correlated features have the potential to play havoc with feature importance calculations (for example see <a href=\"https://www.kaggle.com/code/carlmcbrideellis/variance-inflation-factor-vif-and-explainability\" target=\"_blank\">\"Explainability, collinearity and the variance inflation factor (VIF)\"</a>).</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1857192,
      "author_name": "deepkun1995",
      "author_url": "",
      "post_date": "07/15/2022 23:57:23",
      "content": "<p>thanks for your explanation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1855628": "There's a known truth that highly correlated features can cause problems. I'm focusing on tree booster models, so let's simply say they can cause problems for these models. \n\nBut why? I have a theory (okay, I have a hypothesis. You know what I mean). See, the issue isn't the correlation, the issue is the sampling rate of the model causing highly correlated trees. \n\nIt's a great time to bring it up, because there's great examples of both extremes in this competition, and other ongoing discussions that relate to the topic. \n\nTo simplify, let's pretend that a highly related feature pair - in the examples, engineered feature pair - is close enough to identical to cause problems (obviously 90% correlated will cause less problems than 100%). \n\nIf colsample_tree = 0.88 like in the top public XGB model, then this duplicate feature has a whopping 98.6% chance of being considered by each tree. \n\nHowever, if you include the rounded version of EVERY float feature for '_last', and drop col sampling way down to 0.2 like in the top model, it's self evident you have tons of highly correlated features, but only a 36% chance a tree will pick up on the duplicated feature!\n\nA discussion topic recently brought up including both last-mean and last-first. If you assume that some percentage of base columns have highly correlated first and mean values, then it's another example of duplication!\n\nTrying both at once of two similar features will probably hurt your results if sampling is >.7, but at least *might* help if sampling < .5. \n\nI think the reason rounding is effective for the top model is mainly to increase the chance the highly important \"last\" columns are selected by the tree, compared with other features. Slight diversity between the duplicate versions is just a bonus. \n\nXGB: high sampling rate: https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\n\nLGBM: low sampling rate: https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977",
    "1855637": "Good explanation! Keep it up",
    "1855651": "[Here](https://medium.com/data-design/ensembles-of-tree-based-models-why-correlated-features-do-not-trip-them-and-why-na-matters-7658f4752e1b) is an article by the great [laruae](https://www.kaggle.com/laurae2) that seems to show that highly correlated features are no problem for ensembled tree methods such as xgboost or lightgbm.",
    "1855763": "Generally speaking, tree-based methods have no problem with correlated features. I think you are thinking about linear models, where indeed it is highly advisable to remove correlated features. For tree-based methods, they simply take the best split in whichever feature it may be, and they have no problem jumping between features in their decision-making process.",
    "1856239": "Dear @roberthatch \n\nI agree with @chrisrichardmiles  and @tilii7 in that with modern estimators highly correlated features is not a major problem. However it is worth mentioning that correlated features have the potential to play havoc with feature importance calculations (for example see [\"Explainability, collinearity and the variance inflation factor (VIF)\"](https://www.kaggle.com/code/carlmcbrideellis/variance-inflation-factor-vif-and-explainability)).\n\nAll the best,\ncarl",
    "1857089": "Came here and was about to share this article, well I was slow 😬",
    "1857123": "Clearly I used the wrong analogy and comparison.\n\nMy primary point is that I would expect adding rounding features to the XGB 88% sampling model to hurt the model. (aka to require adjusting the sampling percentage by a huge amount) and perhaps even vice versa, removing the rounding on the dart model might benefit from increasing the sampling by a lot.\n\nAm I wrong there, too? I didn't read the article, let me know if it touches on my main thought?",
    "1857149": "> Am I wrong there, too? I didn't read the article, let me know if it touches on my main thought?\n\nThe article is not long, and I would imagine you'd want to read carefully something that addresses your point of view from this thread. A simple exercise in that article (column duplication) shows that XGB scores are unaffected. Intuitively, I would have thought the same even before reading the article, though there is always a possibility that the score would be affected if the number of correlated features is substantial, say 10% or so. By this I mean 10% of \"child\" features that are all correlated to a single \"parent\" feature. If we had 5% of total features, each of which had a feature doppelganger, I don't think that would matter. Even though that would again amount to a total of 10% correlated features, for the purpose of decision-making I think it potentially matters only when many columns are all correlated to each other, rather than when many columns have one other correlated column apiece.",
    "1857158": "I mean, XGBoost is highly robust to everything.\n\n\"Tree-based models have an innate feature of being robust to correlated features. \"\n\nTo me, taking 20% of the features and doubling them is no where near the same as what's addressed in the article. I would expect XGBoost to be \"robust\" in that scenario... but that doesn't mean it's helpful. Is XGBoost robust against bad colsample values? Or do they cause a problem? In my very limited experience, I'd assume both statements are true. \n\nSo to restate my view: adding (entire sets of) duplicate and near duplicate features will impact the 'true' sampling rate. This can be good or bad. That's all I'm trying to say. \n\nI didn't see anything contradictory in that article.",
    "1857161": "More specifically, the duplicate column experiment in the article isn't setting colsample under 1!\n\nSo yes, if it was guaranteed before, still guaranteed after. 1 - (1-1)^1 == 1 - (1-1)^2. Doesn't matter that you squared a term.",
    "1857179": "I have clearly said that the exercise in that article was simple, and wasn't trying to prove anything by it. You asked a question about the article, I gave you a summary.\n\nBut since you are bringing up column sampling - gradient boosters can always pick one of the two correlated columns and be fine with it. Now, it is possible that the result would be slightly better with both correlated columns included, but that is why we do hyperparameter search to determine the best `colsample` value. A substandard colsample value is not necessarily that way because it eliminated a correlated column, but possibly because it eliminated another more useful column. In other words, correlated features would definitely not be my biggest worry when it comes to column sampling.\n\nCorrelated columns by definition contain very similar information, so it is unlikely that excluding one of them will have a profound effect on the final result. I think that removing any other non-correlated column has greater potential of affecting the model, so that would be of greater concern for me. That of course brings us back to hyperparameter optimization where these things are found in a systematic way.",
    "1857192": "thanks for your explanation"
  },
  "source": "meta"
}