{
  "id": 389231,
  "title": "How to do Feature Selection??",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/389231",
  "author_name": "",
  "post_date": "2023-02-21T05:58:58.872817600Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hii,<br>\nA lot of time it happens that when I add a bunch of features, the overall f1 score decreases, then I add another bunch of features and then also the score decreases. But if I add both of them together, the score increases.<br>\nThis confuses me a lot. Why does this happen?<br>\nAnd so how should I go about adding features?<br>\nThanks.</p>",
  "messages": [
    {
      "id": "2152986",
      "postDate": "02/21/2023 05:58:58",
      "content": "<p>Hii,<br>\nA lot of time it happens that when I add a bunch of features, the overall f1 score decreases, then I add another bunch of features and then also the score decreases. But if I add both of them together, the score increases.<br>\nThis confuses me a lot. Why does this happen?<br>\nAnd so how should I go about adding features?<br>\nThanks.</p>",
      "rawMarkdown": "Hii,\nA lot of time it happens that when I add a bunch of features, the overall f1 score decreases, then I add another bunch of features and then also the score decreases. But if I add both of them together, the score increases.\nThis confuses me a lot. Why does this happen?\nAnd so how should I go about adding features?\nThanks.",
      "votes": null
    },
    {
      "id": "2154563",
      "postDate": "02/22/2023 05:34:30",
      "content": "<p>Maybe the model managed to find useful patterns when all the features get added. To know which features were the useful ones try to look into the features importance. Then look into the top ones. You will get good insights about the useful features.</p>",
      "rawMarkdown": "Maybe the model managed to find useful patterns when all the features get added. To know which features were the useful ones try to look into the features importance. Then look into the top ones. You will get good insights about the useful features.",
      "votes": null
    },
    {
      "id": "2156183",
      "postDate": "02/23/2023 06:09:27",
      "content": "<p>If you are using feature sampling, changing the position of columns can also update your CV.  So depending on how big the difference is, your CV changes may be due to randomness.</p>",
      "rawMarkdown": "If you are using feature sampling, changing the position of columns can also update your CV.  So depending on how big the difference is, your CV changes may be due to randomness.",
      "votes": null
    },
    {
      "id": "2156248",
      "postDate": "02/23/2023 07:35:48",
      "content": "<p>Yeah. But I read somewhere that the feature importance approach may not work every time. Sometimes it can lead us to remove some important features, so it's better to try other approaches like backward or forward elimination.</p>",
      "rawMarkdown": "Yeah. But I read somewhere that the feature importance approach may not work every time. Sometimes it can lead us to remove some important features, so it's better to try other approaches like backward or forward elimination.",
      "votes": null
    },
    {
      "id": "2156249",
      "postDate": "02/23/2023 07:39:09",
      "content": "<p>Yes. But if the features are useful, shouldn't they just increase the score?</p>",
      "rawMarkdown": "Yes. But if the features are useful, shouldn't they just increase the score?",
      "votes": null
    },
    {
      "id": "2160416",
      "postDate": "02/26/2023 16:51:43",
      "content": "<p>Yes, they should normally, but I've seen cases where, given the same set of features, I got <strong>slightly</strong> lower CV scores when the column order changes. However note that this can hold true only for small drops; big drops are due to something else.</p>",
      "rawMarkdown": "Yes, they should normally, but I've seen cases where, given the same set of features, I got **slightly** lower CV scores when the column order changes. However note that this can hold true only for small drops; big drops are due to something else.",
      "votes": null
    },
    {
      "id": "2160426",
      "postDate": "02/26/2023 17:00:36",
      "content": "<p>GBT models are very robust to unimportant features. Especially if using regularization. Therefore, the simplest explanation for your results is:</p>\n<p>(Most likely): both feature sets added gave minimal improvement to the model, the improvement, if any, is less than the model variance if rerandomizing the seed. </p>\n<p>(Less likely): there's significant improvement only if including certain feature(s) in A with certain feature(s) in B due to non-linear important relationship. </p>\n<p>For efficiency prize, feature selection could be very important. And you can probably (?) just leave off all these features unless the was major improvement. If there was big improvement, you can try removing one at a time, it won't be perfect but efficiency prize may not need perfection. </p>\n<p>For normal prize: my way is just use all features, and make sure to tune your regularization parameters to prevent overfitting due to too many features compared to number of train labels. As far as I can tell, that is better practice with GBT models than using feature selection. </p>",
      "rawMarkdown": "GBT models are very robust to unimportant features. Especially if using regularization. Therefore, the simplest explanation for your results is:\n\n(Most likely): both feature sets added gave minimal improvement to the model, the improvement, if any, is less than the model variance if rerandomizing the seed. \n\n(Less likely): there's significant improvement only if including certain feature(s) in A with certain feature(s) in B due to non-linear important relationship. \n\nFor efficiency prize, feature selection could be very important. And you can probably (?) just leave off all these features unless the was major improvement. If there was big improvement, you can try removing one at a time, it won't be perfect but efficiency prize may not need perfection. \n\nFor normal prize: my way is just use all features, and make sure to tune your regularization parameters to prevent overfitting due to too many features compared to number of train labels. As far as I can tell, that is better practice with GBT models than using feature selection.",
      "votes": null
    },
    {
      "id": "2168725",
      "postDate": "03/04/2023 13:23:50",
      "content": "<p>Thank you for the reply. So should I tune the alpha parameter of xgboost for regularization? And how often should I change it as I add more and more features?<br>\nThanks.</p>",
      "rawMarkdown": "Thank you for the reply. So should I tune the alpha parameter of xgboost for regularization? And how often should I change it as I add more and more features?\nThanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2154563,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "02/22/2023 05:34:30",
      "content": "<p>Maybe the model managed to find useful patterns when all the features get added. To know which features were the useful ones try to look into the features importance. Then look into the top ones. You will get good insights about the useful features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2156248,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/23/2023 07:35:48",
          "content": "<p>Yeah. But I read somewhere that the feature importance approach may not work every time. Sometimes it can lead us to remove some important features, so it's better to try other approaches like backward or forward elimination.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2156183,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "02/23/2023 06:09:27",
      "content": "<p>If you are using feature sampling, changing the position of columns can also update your CV.  So depending on how big the difference is, your CV changes may be due to randomness.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2156249,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/23/2023 07:39:09",
          "content": "<p>Yes. But if the features are useful, shouldn't they just increase the score?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2160416,
              "author_name": "hoangnguyen719",
              "author_url": "",
              "post_date": "02/26/2023 16:51:43",
              "content": "<p>Yes, they should normally, but I've seen cases where, given the same set of features, I got <strong>slightly</strong> lower CV scores when the column order changes. However note that this can hold true only for small drops; big drops are due to something else.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2160426,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "02/26/2023 17:00:36",
      "content": "<p>GBT models are very robust to unimportant features. Especially if using regularization. Therefore, the simplest explanation for your results is:</p>\n<p>(Most likely): both feature sets added gave minimal improvement to the model, the improvement, if any, is less than the model variance if rerandomizing the seed. </p>\n<p>(Less likely): there's significant improvement only if including certain feature(s) in A with certain feature(s) in B due to non-linear important relationship. </p>\n<p>For efficiency prize, feature selection could be very important. And you can probably (?) just leave off all these features unless the was major improvement. If there was big improvement, you can try removing one at a time, it won't be perfect but efficiency prize may not need perfection. </p>\n<p>For normal prize: my way is just use all features, and make sure to tune your regularization parameters to prevent overfitting due to too many features compared to number of train labels. As far as I can tell, that is better practice with GBT models than using feature selection. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2168725,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "03/04/2023 13:23:50",
          "content": "<p>Thank you for the reply. So should I tune the alpha parameter of xgboost for regularization? And how often should I change it as I add more and more features?<br>\nThanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2152986": "Hii,\nA lot of time it happens that when I add a bunch of features, the overall f1 score decreases, then I add another bunch of features and then also the score decreases. But if I add both of them together, the score increases.\nThis confuses me a lot. Why does this happen?\nAnd so how should I go about adding features?\nThanks.",
    "2154563": "Maybe the model managed to find useful patterns when all the features get added. To know which features were the useful ones try to look into the features importance. Then look into the top ones. You will get good insights about the useful features.",
    "2156183": "If you are using feature sampling, changing the position of columns can also update your CV.  So depending on how big the difference is, your CV changes may be due to randomness.",
    "2156248": "Yeah. But I read somewhere that the feature importance approach may not work every time. Sometimes it can lead us to remove some important features, so it's better to try other approaches like backward or forward elimination.",
    "2156249": "Yes. But if the features are useful, shouldn't they just increase the score?",
    "2160416": "Yes, they should normally, but I've seen cases where, given the same set of features, I got **slightly** lower CV scores when the column order changes. However note that this can hold true only for small drops; big drops are due to something else.",
    "2160426": "GBT models are very robust to unimportant features. Especially if using regularization. Therefore, the simplest explanation for your results is:\n\n(Most likely): both feature sets added gave minimal improvement to the model, the improvement, if any, is less than the model variance if rerandomizing the seed. \n\n(Less likely): there's significant improvement only if including certain feature(s) in A with certain feature(s) in B due to non-linear important relationship. \n\nFor efficiency prize, feature selection could be very important. And you can probably (?) just leave off all these features unless the was major improvement. If there was big improvement, you can try removing one at a time, it won't be perfect but efficiency prize may not need perfection. \n\nFor normal prize: my way is just use all features, and make sure to tune your regularization parameters to prevent overfitting due to too many features compared to number of train labels. As far as I can tell, that is better practice with GBT models than using feature selection.",
    "2168725": "Thank you for the reply. So should I tune the alpha parameter of xgboost for regularization? And how often should I change it as I add more and more features?\nThanks."
  },
  "source": "meta"
}