{
  "id": 386526,
  "title": "The correct way to approach features importance?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/386526",
  "author_name": "",
  "post_date": "2023-02-13T14:10:41.069490700Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello, I was just wondering, as we are building a model for each question, is it a good idea to just have a global features importance (by average the importance of all the models in the 5 folds), or it is better to check the importance of each question individually (or maybe group the questions to 3 or 4 groups and find the importance for each one?)?<br>\nI think the implementation is not a problem, so my question regards which approach is the most reasonable. This may help further in doing a better features selection.</p>",
  "messages": [
    {
      "id": "2142422",
      "postDate": "02/13/2023 14:10:41",
      "content": "<p>Hello, I was just wondering, as we are building a model for each question, is it a good idea to just have a global features importance (by average the importance of all the models in the 5 folds), or it is better to check the importance of each question individually (or maybe group the questions to 3 or 4 groups and find the importance for each one?)?<br>\nI think the implementation is not a problem, so my question regards which approach is the most reasonable. This may help further in doing a better features selection.</p>",
      "rawMarkdown": "Hello, I was just wondering, as we are building a model for each question, is it a good idea to just have a global features importance (by average the importance of all the models in the 5 folds), or it is better to check the importance of each question individually (or maybe group the questions to 3 or 4 groups and find the importance for each one?)?\nI think the implementation is not a problem, so my question regards which approach is the most reasonable. This may help further in doing a better features selection.",
      "votes": null
    },
    {
      "id": "2142474",
      "postDate": "02/13/2023 14:56:24",
      "content": "<p>The answer depends on your goals. There are different reasons to use feature importance. For me, I like to analyze feature importance as a form of EDA (as opposed to solely improving CV LB). By seeing which features are important for which questions, it helps me explore the data better. This allows me to generate better new features. For example, from the hundreds of rows of events in the train data per user, what rows are most important to predict if user gets questions 5 correct or not? Then i ask myself why is this? Then i explore the game and discover new insights.</p>",
      "rawMarkdown": "The answer depends on your goals. There are different reasons to use feature importance. For me, I like to analyze feature importance as a form of EDA (as opposed to solely improving CV LB). By seeing which features are important for which questions, it helps me explore the data better. This allows me to generate better new features. For example, from the hundreds of rows of events in the train data per user, what rows are most important to predict if user gets questions 5 correct or not? Then i ask myself why is this? Then i explore the game and discover new insights.",
      "votes": null
    },
    {
      "id": "2142903",
      "postDate": "02/13/2023 21:52:27",
      "content": "<p>And global avg and global max(!) sounds more important, logically, if trying to use features importance for the purposes of dimensionality reduction and CV LB.</p>\n<p>It's worth noting the there's only about 11,000 unique users in the train data set, so although it's not a tiny dataset, neither is it super large. The curse of dimensionality can definitely make an impact. Chris said he's already made over 1k features…</p>",
      "rawMarkdown": "And global avg and global max(!) sounds more important, logically, if trying to use features importance for the purposes of dimensionality reduction and CV LB.\n\nIt's worth noting the there's only about 11,000 unique users in the train data set, so although it's not a tiny dataset, neither is it super large. The curse of dimensionality can definitely make an impact. Chris said he's already made over 1k features...",
      "votes": null
    },
    {
      "id": "2144109",
      "postDate": "02/14/2023 18:31:31",
      "content": "<p>I wonder how much the o features importance  change depending on the questions…</p>",
      "rawMarkdown": "I wonder how much the o features importance  change depending on the questions...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2142474,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/13/2023 14:56:24",
      "content": "<p>The answer depends on your goals. There are different reasons to use feature importance. For me, I like to analyze feature importance as a form of EDA (as opposed to solely improving CV LB). By seeing which features are important for which questions, it helps me explore the data better. This allows me to generate better new features. For example, from the hundreds of rows of events in the train data per user, what rows are most important to predict if user gets questions 5 correct or not? Then i ask myself why is this? Then i explore the game and discover new insights.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2142903,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "02/13/2023 21:52:27",
          "content": "<p>And global avg and global max(!) sounds more important, logically, if trying to use features importance for the purposes of dimensionality reduction and CV LB.</p>\n<p>It's worth noting the there's only about 11,000 unique users in the train data set, so although it's not a tiny dataset, neither is it super large. The curse of dimensionality can definitely make an impact. Chris said he's already made over 1k features…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2144109,
      "author_name": "mikhaildonskoy",
      "author_url": "",
      "post_date": "02/14/2023 18:31:31",
      "content": "<p>I wonder how much the o features importance  change depending on the questions…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2142422": "Hello, I was just wondering, as we are building a model for each question, is it a good idea to just have a global features importance (by average the importance of all the models in the 5 folds), or it is better to check the importance of each question individually (or maybe group the questions to 3 or 4 groups and find the importance for each one?)?\nI think the implementation is not a problem, so my question regards which approach is the most reasonable. This may help further in doing a better features selection.",
    "2142474": "The answer depends on your goals. There are different reasons to use feature importance. For me, I like to analyze feature importance as a form of EDA (as opposed to solely improving CV LB). By seeing which features are important for which questions, it helps me explore the data better. This allows me to generate better new features. For example, from the hundreds of rows of events in the train data per user, what rows are most important to predict if user gets questions 5 correct or not? Then i ask myself why is this? Then i explore the game and discover new insights.",
    "2142903": "And global avg and global max(!) sounds more important, logically, if trying to use features importance for the purposes of dimensionality reduction and CV LB.\n\nIt's worth noting the there's only about 11,000 unique users in the train data set, so although it's not a tiny dataset, neither is it super large. The curse of dimensionality can definitely make an impact. Chris said he's already made over 1k features...",
    "2144109": "I wonder how much the o features importance  change depending on the questions..."
  },
  "source": "meta"
}