{
  "id": 401278,
  "title": "How many features should we use?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/401278",
  "author_name": "",
  "post_date": "2023-04-12T15:05:24.625907400Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm now using 2000~10000 features.<br>\nAre there any harms of using too much features?<br>\nHow can I choose features?</p>\n<p>Feature importance contains the information from label data and selecting features based on it boosts cv score but decrease lb score, if I understand the post below.<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682</a></p>",
  "messages": [
    {
      "id": "2219434",
      "postDate": "04/12/2023 15:05:24",
      "content": "<p>I'm now using 2000~10000 features.<br>\nAre there any harms of using too much features?<br>\nHow can I choose features?</p>\n<p>Feature importance contains the information from label data and selecting features based on it boosts cv score but decrease lb score, if I understand the post below.<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682</a></p>",
      "rawMarkdown": "I'm now using 2000~10000 features.\nAre there any harms of using too much features?\nHow can I choose features?\n\nFeature importance contains the information from label data and selecting features based on it boosts cv score but decrease lb score, if I understand the post below.\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682",
      "votes": null
    },
    {
      "id": "2219852",
      "postDate": "04/13/2023 00:13:10",
      "content": "<p>That's a great question!</p>\n<p>Generally using too many features can potentially make your model inefficient or even overfit your data.</p>\n<p>Here is my suggestion: </p>\n<p>You can use mutual information to determine which features are the most useful, then run your model based on that and see if there is a difference.</p>\n<p>This is a good link to explain Mutual Information if needed: <a href=\"https://people.cs.umass.edu/~elm/Teaching/Docs/mutInf.pdf\" target=\"_blank\">Mutual Information</a></p>\n<p>Also, I suggest taking the 'Feature Engineering' course in the learn section if you want a deeper understanding on the importance of features. </p>",
      "rawMarkdown": "That's a great question!\n\nGenerally using too many features can potentially make your model inefficient or even overfit your data.\n\nHere is my suggestion: \n\nYou can use mutual information to determine which features are the most useful, then run your model based on that and see if there is a difference.\n\nThis is a good link to explain Mutual Information if needed: [Mutual Information](https://people.cs.umass.edu/~elm/Teaching/Docs/mutInf.pdf)\n\nAlso, I suggest taking the 'Feature Engineering' course in the learn section if you want a deeper understanding on the importance of features.",
      "votes": null
    },
    {
      "id": "2219891",
      "postDate": "04/13/2023 01:20:12",
      "content": "<p>Thank you for sharing!<br>\nI will study about mutual information.</p>",
      "rawMarkdown": "Thank you for sharing!\nI will study about mutual information.",
      "votes": null
    },
    {
      "id": "2219980",
      "postDate": "04/13/2023 03:50:35",
      "content": "<p>Too many features will suffer from the curse of dimensionality. Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin. </p>",
      "rawMarkdown": "Too many features will suffer from the curse of dimensionality. Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin.",
      "votes": null
    },
    {
      "id": "2220227",
      "postDate": "04/13/2023 07:54:52",
      "content": "<p>Too many features will cause overfitting on the training data. I suggest some feature engineering and dimensionality reduction like PCA. Did you one hot encode all of the categorical data to cause this many dimensions?</p>",
      "rawMarkdown": "Too many features will cause overfitting on the training data. I suggest some feature engineering and dimensionality reduction like PCA. Did you one hot encode all of the categorical data to cause this many dimensions?",
      "votes": null
    },
    {
      "id": "2223443",
      "postDate": "04/16/2023 09:12:17",
      "content": "<blockquote>\n  <p>Too many features will cause overfitting on the training data. I suggest some feature engineering <br>\n  and dimensionality reduction like PCA.</p>\n</blockquote>\n<p>Thanks for your advice. I will study about PCA for dimentionality reduction.</p>\n<blockquote>\n  <p>Did you one hot encode all of the categorical data to cause this many dimensions?</p>\n</blockquote>\n<p>No, I didn't. I only have numerical features.</p>",
      "rawMarkdown": ">Too many features will cause overfitting on the training data. I suggest some feature engineering \nand dimensionality reduction like PCA.\n\nThanks for your advice. I will study about PCA for dimentionality reduction.\n\n>Did you one hot encode all of the categorical data to cause this many dimensions?\n\nNo, I didn't. I only have numerical features.",
      "votes": null
    },
    {
      "id": "2223444",
      "postDate": "04/16/2023 09:14:10",
      "content": "<blockquote>\n  <p>Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin.</p>\n</blockquote>\n<p>So the number of features that can be used is relative to the size of training set. Thank you!</p>",
      "rawMarkdown": ">Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin.\n\nSo the number of features that can be used is relative to the size of training set. Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2219852,
      "author_name": "adriandiazny",
      "author_url": "",
      "post_date": "04/13/2023 00:13:10",
      "content": "<p>That's a great question!</p>\n<p>Generally using too many features can potentially make your model inefficient or even overfit your data.</p>\n<p>Here is my suggestion: </p>\n<p>You can use mutual information to determine which features are the most useful, then run your model based on that and see if there is a difference.</p>\n<p>This is a good link to explain Mutual Information if needed: <a href=\"https://people.cs.umass.edu/~elm/Teaching/Docs/mutInf.pdf\" target=\"_blank\">Mutual Information</a></p>\n<p>Also, I suggest taking the 'Feature Engineering' course in the learn section if you want a deeper understanding on the importance of features. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2219891,
          "author_name": "nynyny67",
          "author_url": "",
          "post_date": "04/13/2023 01:20:12",
          "content": "<p>Thank you for sharing!<br>\nI will study about mutual information.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2219980,
      "author_name": "ataraxian",
      "author_url": "",
      "post_date": "04/13/2023 03:50:35",
      "content": "<p>Too many features will suffer from the curse of dimensionality. Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2223444,
          "author_name": "nynyny67",
          "author_url": "",
          "post_date": "04/16/2023 09:14:10",
          "content": "<blockquote>\n  <p>Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin.</p>\n</blockquote>\n<p>So the number of features that can be used is relative to the size of training set. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2220227,
      "author_name": "haydenlabrie",
      "author_url": "",
      "post_date": "04/13/2023 07:54:52",
      "content": "<p>Too many features will cause overfitting on the training data. I suggest some feature engineering and dimensionality reduction like PCA. Did you one hot encode all of the categorical data to cause this many dimensions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2223443,
          "author_name": "nynyny67",
          "author_url": "",
          "post_date": "04/16/2023 09:12:17",
          "content": "<blockquote>\n  <p>Too many features will cause overfitting on the training data. I suggest some feature engineering <br>\n  and dimensionality reduction like PCA.</p>\n</blockquote>\n<p>Thanks for your advice. I will study about PCA for dimentionality reduction.</p>\n<blockquote>\n  <p>Did you one hot encode all of the categorical data to cause this many dimensions?</p>\n</blockquote>\n<p>No, I didn't. I only have numerical features.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2219434": "I'm now using 2000~10000 features.\nAre there any harms of using too much features?\nHow can I choose features?\n\nFeature importance contains the information from label data and selecting features based on it boosts cv score but decrease lb score, if I understand the post below.\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682",
    "2219852": "That's a great question!\n\nGenerally using too many features can potentially make your model inefficient or even overfit your data.\n\nHere is my suggestion: \n\nYou can use mutual information to determine which features are the most useful, then run your model based on that and see if there is a difference.\n\nThis is a good link to explain Mutual Information if needed: [Mutual Information](https://people.cs.umass.edu/~elm/Teaching/Docs/mutInf.pdf)\n\nAlso, I suggest taking the 'Feature Engineering' course in the learn section if you want a deeper understanding on the importance of features.",
    "2219891": "Thank you for sharing!\nI will study about mutual information.",
    "2219980": "Too many features will suffer from the curse of dimensionality. Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin.",
    "2220227": "Too many features will cause overfitting on the training data. I suggest some feature engineering and dimensionality reduction like PCA. Did you one hot encode all of the categorical data to cause this many dimensions?",
    "2223443": ">Too many features will cause overfitting on the training data. I suggest some feature engineering \nand dimensionality reduction like PCA.\n\nThanks for your advice. I will study about PCA for dimentionality reduction.\n\n>Did you one hot encode all of the categorical data to cause this many dimensions?\n\nNo, I didn't. I only have numerical features.",
    "2223444": ">Considering that we are using a rather small training set (around 25,000), 10,000 features are most likely too much above the safe margin.\n\nSo the number of features that can be used is relative to the size of training set. Thank you!"
  },
  "source": "meta"
}