{
  "id": 399957,
  "title": "why doesn't adding new features help improve tree-based models?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/399957",
  "author_name": "",
  "post_date": "2023-04-06T08:59:48.247882600Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As is well known, tree-based models such as xgb and gbdt have feature selection capabilities themselves. If new features are added but do not have a positive effect on label prediction, these features would be filtered out. That is, adding features will at least not make the model worse.</p>\n<p>But based on what I have seen, the addition of new features often leads to poorer performance of the model. What is the reason for this?</p>",
  "messages": [
    {
      "id": "2211782",
      "postDate": "04/06/2023 08:59:48",
      "content": "<p>As is well known, tree-based models such as xgb and gbdt have feature selection capabilities themselves. If new features are added but do not have a positive effect on label prediction, these features would be filtered out. That is, adding features will at least not make the model worse.</p>\n<p>But based on what I have seen, the addition of new features often leads to poorer performance of the model. What is the reason for this?</p>",
      "rawMarkdown": "As is well known, tree-based models such as xgb and gbdt have feature selection capabilities themselves. If new features are added but do not have a positive effect on label prediction, these features would be filtered out. That is, adding features will at least not make the model worse.\n\nBut based on what I have seen, the addition of new features often leads to poorer performance of the model. What is the reason for this?",
      "votes": null
    },
    {
      "id": "2211820",
      "postDate": "04/06/2023 09:31:07",
      "content": "<p>Gradient boost methods  have good feature selection capabilities (way better than neural networks or linear methods).<br>\nBut It does not mean that they cannot overfit the train set !</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fca37b5e39a948f249de3e05c2c279351%2F1677508581309.svg?generation=1680773404303881&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Gradient boost methods  have good feature selection capabilities (way better than neural networks or linear methods).\nBut It does not mean that they cannot overfit the train set !\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fca37b5e39a948f249de3e05c2c279351%2F1677508581309.svg?generation=1680773404303881&alt=media)",
      "votes": null
    },
    {
      "id": "2211835",
      "postDate": "04/06/2023 09:49:00",
      "content": "<p>In that case, I guess I should lower the complexity of my model?</p>",
      "rawMarkdown": "In that case, I guess I should lower the complexity of my model?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2211820,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "04/06/2023 09:31:07",
      "content": "<p>Gradient boost methods  have good feature selection capabilities (way better than neural networks or linear methods).<br>\nBut It does not mean that they cannot overfit the train set !</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fca37b5e39a948f249de3e05c2c279351%2F1677508581309.svg?generation=1680773404303881&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2211835,
          "author_name": "",
          "author_url": "",
          "post_date": "04/06/2023 09:49:00",
          "content": "<p>In that case, I guess I should lower the complexity of my model?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2211782": "As is well known, tree-based models such as xgb and gbdt have feature selection capabilities themselves. If new features are added but do not have a positive effect on label prediction, these features would be filtered out. That is, adding features will at least not make the model worse.\n\nBut based on what I have seen, the addition of new features often leads to poorer performance of the model. What is the reason for this?",
    "2211820": "Gradient boost methods  have good feature selection capabilities (way better than neural networks or linear methods).\nBut It does not mean that they cannot overfit the train set !\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Fca37b5e39a948f249de3e05c2c279351%2F1677508581309.svg?generation=1680773404303881&alt=media)",
    "2211835": "In that case, I guess I should lower the complexity of my model?"
  },
  "source": "meta"
}