{
  "id": 580617,
  "title": "Key Insights from Feature Selection & Analysis",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580617",
  "author_name": "AC",
  "post_date": "2025-05-25T12:31:11.629000",
  "votes": 72,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Hey, there! </p>\n<p>I’ve been digging deep into the features and wanted to share some insights that might help others. The dataset is large, anonymized, and noisy — so I focused a lot on feature redundancy, signal strength, and interpretability. <br>\nYou can access the entire analysis in this notebook : <a href=\"https://www.kaggle.com/code/ahsuna123/anonymized-features-importance-selection-eda\" target=\"_blank\">https://www.kaggle.com/code/ahsuna123/anonymized-features-importance-selection-eda</a> </p>\n<p>Here’s what stood out:</p>\n<p><strong>1. A Lot of Redundant Features :</strong></p>\n<p>I found over 20 pairs of features with extremely high correlation (Pearson &gt; 0.98). Dropping one from each pair would help reduce dimensionality without losing information. </p>\n<p><strong>2. Weak Linear Correlation with the Target :</strong></p>\n<p>Most features individually have very weak correlation with the target. The best ones were only around ~0.07 — features like X21, X20, X28, and X863. This tells me that the signal isn’t in individual features but likely hidden in non-linear patterns or interactions — a good case for tree-based models or neural nets.</p>\n<p><strong>3. PCA Shows a Lot of Noise :</strong></p>\n<p>I ran Incremental PCA to understand the variance spread. Turns out:</p>\n<p>~90% variance is captured by just 23 components</p>\n<p>~95% by 36 components [Note: I only took top 50 features for computational ease]</p>\n<p><strong>4. Used 3 Different Feature Selection Techniques :</strong></p>\n<p>To get a more robust handle on useful features, I compared:</p>\n<ul>\n<li><p>Univariate F-test → picked out X21, X28, etc. (linear signal)</p></li>\n<li><p>Random Forest Importance →  X198, X95, X179, etc. (non-linear)</p></li>\n<li><p>RFE with Lasso → sparse feature selection that picked overlapping but also unique ones.</p></li>\n</ul>\n<p>Interestingly, 11 features consistently came up across all 3. I’m treating those as reliable “core signals.”</p>\n<p><strong>5. Clusters in the Data :</strong></p>\n<p>I ran KMeans on the dataset (k=5). Some clusters were heavily imbalanced in size — one had 231k+ samples! Early look at label distribution suggests that some clusters may be easier to classify than others. Planning to try cluster-based ensembling next.</p>\n<p>Note: You can tweak the code for your analysis ( I mostly showed top 20 features, but you can always expand).</p>\n<p>Hope, it's helpful! :))</p>",
  "messages": [
    {
      "id": 3209240,
      "postDate": "2025-05-25T12:31:11.630Z",
      "content": "<p>Hey, there! </p>\n<p>I’ve been digging deep into the features and wanted to share some insights that might help others. The dataset is large, anonymized, and noisy — so I focused a lot on feature redundancy, signal strength, and interpretability. <br>\nYou can access the entire analysis in this notebook : <a href=\"https://www.kaggle.com/code/ahsuna123/anonymized-features-importance-selection-eda\" target=\"_blank\">https://www.kaggle.com/code/ahsuna123/anonymized-features-importance-selection-eda</a> </p>\n<p>Here’s what stood out:</p>\n<p><strong>1. A Lot of Redundant Features :</strong></p>\n<p>I found over 20 pairs of features with extremely high correlation (Pearson &gt; 0.98). Dropping one from each pair would help reduce dimensionality without losing information. </p>\n<p><strong>2. Weak Linear Correlation with the Target :</strong></p>\n<p>Most features individually have very weak correlation with the target. The best ones were only around ~0.07 — features like X21, X20, X28, and X863. This tells me that the signal isn’t in individual features but likely hidden in non-linear patterns or interactions — a good case for tree-based models or neural nets.</p>\n<p><strong>3. PCA Shows a Lot of Noise :</strong></p>\n<p>I ran Incremental PCA to understand the variance spread. Turns out:</p>\n<p>~90% variance is captured by just 23 components</p>\n<p>~95% by 36 components [Note: I only took top 50 features for computational ease]</p>\n<p><strong>4. Used 3 Different Feature Selection Techniques :</strong></p>\n<p>To get a more robust handle on useful features, I compared:</p>\n<ul>\n<li><p>Univariate F-test → picked out X21, X28, etc. (linear signal)</p></li>\n<li><p>Random Forest Importance →  X198, X95, X179, etc. (non-linear)</p></li>\n<li><p>RFE with Lasso → sparse feature selection that picked overlapping but also unique ones.</p></li>\n</ul>\n<p>Interestingly, 11 features consistently came up across all 3. I’m treating those as reliable “core signals.”</p>\n<p><strong>5. Clusters in the Data :</strong></p>\n<p>I ran KMeans on the dataset (k=5). Some clusters were heavily imbalanced in size — one had 231k+ samples! Early look at label distribution suggests that some clusters may be easier to classify than others. Planning to try cluster-based ensembling next.</p>\n<p>Note: You can tweak the code for your analysis ( I mostly showed top 20 features, but you can always expand).</p>\n<p>Hope, it's helpful! :))</p>",
      "rawMarkdown": "Hey, there! \n\nI’ve been digging deep into the features and wanted to share some insights that might help others. The dataset is large, anonymized, and noisy — so I focused a lot on feature redundancy, signal strength, and interpretability. \nYou can access the entire analysis in this notebook : https://www.kaggle.com/code/ahsuna123/anonymized-features-importance-selection-eda \n\nHere’s what stood out:\n\n**1. A Lot of Redundant Features :**\n\nI found over 20 pairs of features with extremely high correlation (Pearson > 0.98). Dropping one from each pair would help reduce dimensionality without losing information. \n\n**2. Weak Linear Correlation with the Target :**\n\nMost features individually have very weak correlation with the target. The best ones were only around ~0.07 — features like X21, X20, X28, and X863. This tells me that the signal isn’t in individual features but likely hidden in non-linear patterns or interactions — a good case for tree-based models or neural nets.\n\n**3. PCA Shows a Lot of Noise :**\n\nI ran Incremental PCA to understand the variance spread. Turns out:\n\n~90% variance is captured by just 23 components\n\n~95% by 36 components [Note: I only took top 50 features for computational ease]\n\n**4. Used 3 Different Feature Selection Techniques :**\n\nTo get a more robust handle on useful features, I compared:\n\n- Univariate F-test → picked out X21, X28, etc. (linear signal)\n\n- Random Forest Importance →  X198, X95, X179, etc. (non-linear)\n\n- RFE with Lasso → sparse feature selection that picked overlapping but also unique ones.\n\nInterestingly, 11 features consistently came up across all 3. I’m treating those as reliable “core signals.”\n\n**5. Clusters in the Data :**\n\nI ran KMeans on the dataset (k=5). Some clusters were heavily imbalanced in size — one had 231k+ samples! Early look at label distribution suggests that some clusters may be easier to classify than others. Planning to try cluster-based ensembling next.\n\nNote: You can tweak the code for your analysis ( I mostly showed top 20 features, but you can always expand).\n\nHope, it's helpful! :))",
      "votes": 72
    },
    {
      "id": 3214616,
      "postDate": "2025-05-31T21:26:34.260Z",
      "content": "<p>Correct me if I'm wrong, but a boosted tree based algorithm would be better at non-linear feature selection than a basic random forest.</p>",
      "rawMarkdown": "Correct me if I'm wrong, but a boosted tree based algorithm would be better at non-linear feature selection than a basic random forest.",
      "votes": 7,
      "replies": [
        {
          "id": 3215475,
          "postDate": "2025-06-02T09:43:51.053Z",
          "content": "<p>That's absolutely right. This is only a baseline analysis! :))</p>",
          "rawMarkdown": "That's absolutely right. This is only a baseline analysis! :))",
          "votes": 4
        },
        {
          "id": 3235279,
          "postDate": "2025-06-29T02:27:55.637Z",
          "content": "<p>While normally true, in this competition there is a risk of assigning meaning to noise as the boosted portion of a tree algorithm attempts to find signal from noise to reduce errors.</p>",
          "rawMarkdown": "While normally true, in this competition there is a risk of assigning meaning to noise as the boosted portion of a tree algorithm attempts to find signal from noise to reduce errors.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3209408,
      "postDate": "2025-05-25T18:19:30.423Z",
      "content": "<p>Heyy this was really helpful. It would be really helpful if u could provide some sources where i could learn these feature selection techniques in a detailed and systematic manner. Thank you so much.</p>",
      "rawMarkdown": "Heyy this was really helpful. It would be really helpful if u could provide some sources where i could learn these feature selection techniques in a detailed and systematic manner. Thank you so much.",
      "votes": 3,
      "replies": [
        {
          "id": 3209536,
          "postDate": "2025-05-26T02:43:49.887Z",
          "content": "<table>\n<thead>\n<tr>\n<th>Method Type</th>\n<th>Examples &amp; Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Filter Methods</strong></td>\n<td>Based on statistical tests (independent of models)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Variance Threshold (remove low variance features)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Correlation matrix (remove highly correlated features)</td>\n</tr>\n<tr>\n<td><strong>Wrapper Methods</strong></td>\n<td>Evaluate model performance with subsets of features</td>\n</tr>\n<tr>\n<td></td>\n<td>- Recursive Feature Elimination (RFE)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Sequential Feature Selector (SFS)</td>\n</tr>\n<tr>\n<td><strong>Embedded Methods</strong></td>\n<td>Feature selection is part of the model training</td>\n</tr>\n<tr>\n<td></td>\n<td>- Lasso (L1 Regularization)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Tree-based feature importance (Random Forest, XGBoost)</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "| Method Type          | Examples & Description                                   |\n| -------------------- | -------------------------------------------------------- |\n| **Filter Methods**   | Based on statistical tests (independent of models)       |\n|                      | - Variance Threshold (remove low variance features)      |\n|                      | - Correlation matrix (remove highly correlated features) |\n| **Wrapper Methods**  | Evaluate model performance with subsets of features      |\n|                      | - Recursive Feature Elimination (RFE)                    |\n|                      | - Sequential Feature Selector (SFS)                      |\n| **Embedded Methods** | Feature selection is part of the model training          |\n|                      | - Lasso (L1 Regularization)                              |\n|                      | - Tree-based feature importance (Random Forest, XGBoost) |\n",
          "votes": 8
        },
        {
          "id": 3209537,
          "postDate": "2025-05-26T02:44:28.310Z",
          "content": "<p>These are some of the feature selection method. I hope it will help you.</p>",
          "rawMarkdown": "These are some of the feature selection method. I hope it will help you.",
          "votes": 1
        },
        {
          "id": 3209648,
          "postDate": "2025-05-26T06:08:48.933Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/santopaulm\" target=\"_blank\">@santopaulm</a> <br>\nThank you for your comment. For me, this notebook has proven to be very useful, personally one of my favorite notebooks on Kaggle by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <a href=\"https://www.kaggle.com/code/ambrosm/mcts-eda-which-makes-sense\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/mcts-eda-which-makes-sense</a> <br>\nI got to learn alot about feature importance and selection from this. I'd suggest more than diving deep into the theory look for notebooks from playground competitions and read them. That'd make you understand the concept and it's implication in real time. </p>",
          "rawMarkdown": "Hey @santopaulm \nThank you for your comment. For me, this notebook has proven to be very useful, personally one of my favorite notebooks on Kaggle by @ambrosm https://www.kaggle.com/code/ambrosm/mcts-eda-which-makes-sense \nI got to learn alot about feature importance and selection from this. I'd suggest more than diving deep into the theory look for notebooks from playground competitions and read them. That'd make you understand the concept and it's implication in real time. ",
          "votes": 3,
          "replies": [
            {
              "id": 3209776,
              "postDate": "2025-05-26T09:46:23.320Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3209780,
          "postDate": "2025-05-26T09:48:51.907Z",
          "content": "<p>CatBoost has an inbuilt mechanism for feature selection (RFE) based on SHAP values; see the tutorial <a href=\"https://github.com/catboost/tutorials/blob/master/feature_selection/select_features_tutorial.ipynb\" target=\"_blank\">here</a>. If you are interested in how it works on real Comp data, I have a couple of public notebooks: <a href=\"https://www.kaggle.com/code/yekenot/feature-elimination-by-catboost\" target=\"_blank\"><em>Feature Elimination by CatBoost</em></a> (trading data, drop num feats from 124 to 100 with score improvement) and <a href=\"https://www.kaggle.com/code/yekenot/catboost-as-feature-selector\" target=\"_blank\"><em>CatBoost as Feature Selector</em></a> (banking data, drop num feats from 918 to 618 with score improvement).</p>",
          "rawMarkdown": "CatBoost has an inbuilt mechanism for feature selection (RFE) based on SHAP values; see the tutorial [here](https://github.com/catboost/tutorials/blob/master/feature_selection/select_features_tutorial.ipynb). If you are interested in how it works on real Comp data, I have a couple of public notebooks: [*Feature Elimination by CatBoost*](https://www.kaggle.com/code/yekenot/feature-elimination-by-catboost) (trading data, drop num feats from 124 to 100 with score improvement) and [*CatBoost as Feature Selector*](https://www.kaggle.com/code/yekenot/catboost-as-feature-selector) (banking data, drop num feats from 918 to 618 with score improvement).",
          "votes": 7
        }
      ]
    },
    {
      "id": 3216458,
      "postDate": "2025-06-03T16:00:20.990Z",
      "content": "<p>Thanks for sharing this unique and timely dataset! Crypto market behavior is complex and volatile, so having access to high-quality, granular data like this is a great resource for testing advanced time series models and market prediction strategies. Looking forward to seeing what the community builds with this!</p>",
      "rawMarkdown": "Thanks for sharing this unique and timely dataset! Crypto market behavior is complex and volatile, so having access to high-quality, granular data like this is a great resource for testing advanced time series models and market prediction strategies. Looking forward to seeing what the community builds with this!",
      "votes": 2,
      "replies": [
        {
          "id": 3216509,
          "postDate": "2025-06-03T17:08:45.563Z",
          "content": "<p>Appreciate your words :)</p>",
          "rawMarkdown": "Appreciate your words :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 3218949,
      "postDate": "2025-06-07T00:58:25.757Z",
      "content": "<p>I have seen a lot of different feature Selection &amp; Analysis in the discussion threads across competitions. This is definitely top-notch, very clear and the codes are concise. Great work.</p>",
      "rawMarkdown": "I have seen a lot of different feature Selection & Analysis in the discussion threads across competitions. This is definitely top-notch, very clear and the codes are concise. Great work.",
      "votes": 1,
      "replies": [
        {
          "id": 3218992,
          "postDate": "2025-06-07T03:22:07.030Z",
          "content": "<p>Thanks Andy, appreciate your words! :))</p>",
          "rawMarkdown": "Thanks Andy, appreciate your words! :))",
          "votes": 1
        }
      ]
    },
    {
      "id": 3213688,
      "postDate": "2025-05-30T10:29:05.340Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a>, the above info way really useful. Though, I gone through the dataset and had some observation. Data contains some low variance feature and some high variance features, but major features lay in a specific region. so, working on that sweet spot would get some meaningful results.</p>",
      "rawMarkdown": "Hey @ahsuna123, the above info way really useful. Though, I gone through the dataset and had some observation. Data contains some low variance feature and some high variance features, but major features lay in a specific region. so, working on that sweet spot would get some meaningful results.",
      "votes": 1,
      "replies": [
        {
          "id": 3213936,
          "postDate": "2025-05-30T17:21:32.197Z",
          "content": "<p>I certainly agree! </p>",
          "rawMarkdown": "I certainly agree! "
        }
      ]
    },
    {
      "id": 3210181,
      "postDate": "2025-05-26T20:50:17.900Z",
      "content": "<p>Really helpful. Appreciate it.</p>",
      "rawMarkdown": "Really helpful. Appreciate it.",
      "votes": 1
    },
    {
      "id": 3209614,
      "postDate": "2025-05-26T04:57:44.460Z",
      "content": "<p>this was helpful, wish to see this kind of insights more in future </p>",
      "rawMarkdown": "this was helpful, wish to see this kind of insights more in future ",
      "votes": 1,
      "replies": [
        {
          "id": 3209653,
          "postDate": "2025-05-26T06:11:40.713Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/Sahil\" target=\"_blank\">@Sahil</a><br>\nThanks for the comment. Glad you found it helpful! :))</p>",
          "rawMarkdown": "Hey @Sahil\nThanks for the comment. Glad you found it helpful! :))"
        }
      ]
    },
    {
      "id": 3223646,
      "postDate": "2025-06-13T16:25:16.590Z",
      "content": "<p>i found several problems when i transfrom features . this post really good for me . thank you so much &lt;3</p>",
      "rawMarkdown": "i found several problems when i transfrom features . this post really good for me . thank you so much <3",
      "votes": 2
    },
    {
      "id": 3220378,
      "postDate": "2025-06-09T08:08:59.983Z",
      "content": "<p>feature selection is the key for this competition. Thanks for sharing your insights</p>",
      "rawMarkdown": "feature selection is the key for this competition. Thanks for sharing your insights",
      "votes": 2,
      "replies": [
        {
          "id": 3220383,
          "postDate": "2025-06-09T08:09:48.100Z",
          "content": "<p>Exactly!! That plays a key role in this one! </p>",
          "rawMarkdown": "Exactly!! That plays a key role in this one! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3215023,
      "postDate": "2025-06-01T13:24:53.213Z",
      "content": "<p>what made you choose Incremental PCA over standard or kernel PCA?  Was it mainly a size thing, or did it give better results too?</p>",
      "rawMarkdown": "what made you choose Incremental PCA over standard or kernel PCA?  Was it mainly a size thing, or did it give better results too?",
      "votes": 2,
      "replies": [
        {
          "id": 3215476,
          "postDate": "2025-06-02T09:45:46.277Z",
          "content": "<p>Hey, Aniket! <a href=\"https://www.kaggle.com/aniketpotabatti\" target=\"_blank\">@aniketpotabatti</a> <br>\nSolely to avoid OOM issues.</p>",
          "rawMarkdown": "Hey, Aniket! @aniketpotabatti \nSolely to avoid OOM issues.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3210110,
      "postDate": "2025-05-26T18:27:13.620Z",
      "content": "<p>Helpful guideline, very appreciate! </p>",
      "rawMarkdown": "Helpful guideline, very appreciate! ",
      "votes": 2
    },
    {
      "id": 3209414,
      "postDate": "2025-05-25T18:40:11.733Z",
      "content": "<p>Did feature removal actually boost your cv score? In another post the auther shows results in which removing features outside of unduplicated/non constant features reduces the CV scores</p>",
      "rawMarkdown": "Did feature removal actually boost your cv score? In another post the auther shows results in which removing features outside of unduplicated/non constant features reduces the CV scores",
      "votes": 2,
      "replies": [
        {
          "id": 3209533,
          "postDate": "2025-05-26T02:39:50.217Z",
          "content": "<p>Well, it's not always the case, it depends on the problem type. Sometimes removing features make model simple and hence boost CV score and sometimes it can over fitting issues if not removed. So, removing features is relative and not absolute and depends on the condition.</p>",
          "rawMarkdown": "Well, it's not always the case, it depends on the problem type. Sometimes removing features make model simple and hence boost CV score and sometimes it can over fitting issues if not removed. So, removing features is relative and not absolute and depends on the condition.",
          "votes": 3
        },
        {
          "id": 3209649,
          "postDate": "2025-05-26T06:10:55.033Z",
          "content": "<p>Hey, Max! <br>\nI've not tried removing any features atm. I'd update as soon as I experiment with that. What <a href=\"https://www.kaggle.com/sharmajicoder\" target=\"_blank\">@sharmajicoder</a> said is right though, it's relative.</p>",
          "rawMarkdown": "Hey, Max! \nI've not tried removing any features atm. I'd update as soon as I experiment with that. What @sharmajicoder said is right though, it's relative.",
          "votes": 1
        },
        {
          "id": 3213694,
          "postDate": "2025-05-30T10:38:47.223Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/maxuhl98\" target=\"_blank\">@maxuhl98</a>, features may contain information that is helpful for prediction, but not always. It's important to statistically evaluate how much each feature actually contributes to the CV scores.</p>",
          "rawMarkdown": "Hey @maxuhl98, features may contain information that is helpful for prediction, but not always. It's important to statistically evaluate how much each feature actually contributes to the CV scores.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3223178,
      "postDate": "2025-06-13T01:12:17.697Z",
      "content": "<p>This is helpful, thank you</p>",
      "rawMarkdown": "This is helpful, thank you",
      "votes": 1
    },
    {
      "id": 3213526,
      "postDate": "2025-05-30T06:05:07.620Z",
      "content": "<p>Nice analysis. Thanks for sharing. </p>",
      "rawMarkdown": "Nice analysis. Thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 3210473,
      "postDate": "2025-05-27T08:16:17.527Z",
      "content": "<p>Thank you for your analysis!</p>",
      "rawMarkdown": "Thank you for your analysis!",
      "votes": 1
    },
    {
      "id": 3209405,
      "postDate": "2025-05-25T18:17:23.073Z",
      "content": "<p>Thanks for your insights 🤗🤗</p>",
      "rawMarkdown": "Thanks for your insights 🤗🤗",
      "votes": 1
    },
    {
      "id": 3216856,
      "postDate": "2025-06-04T07:17:59.483Z",
      "content": "<p>Thanks for sharing </p>",
      "rawMarkdown": "Thanks for sharing ",
      "votes": 2
    },
    {
      "id": 3215016,
      "postDate": "2025-06-01T12:57:21.607Z",
      "content": "<p>thanks! Great observation! </p>",
      "rawMarkdown": "thanks! Great observation! ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3214616,
      "author_name": "Migrant Worker Data Hub",
      "author_url": "",
      "post_date": "2025-05-31T21:26:34.260000",
      "content": "<p>Correct me if I'm wrong, but a boosted tree based algorithm would be better at non-linear feature selection than a basic random forest.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3215475,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-06-02T09:43:51.053000",
          "content": "<p>That's absolutely right. This is only a baseline analysis! :))</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3235279,
          "author_name": "Jasmine Z. Li",
          "author_url": "",
          "post_date": "2025-06-29T02:27:55.637000",
          "content": "<p>While normally true, in this competition there is a risk of assigning meaning to noise as the boosted portion of a tree algorithm attempts to find signal from noise to reduce errors.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3209408,
      "author_name": "Santo Paul M",
      "author_url": "",
      "post_date": "2025-05-25T18:19:30.423000",
      "content": "<p>Heyy this was really helpful. It would be really helpful if u could provide some sources where i could learn these feature selection techniques in a detailed and systematic manner. Thank you so much.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3209536,
          "author_name": "Stable Space",
          "author_url": "",
          "post_date": "2025-05-26T02:43:49.887000",
          "content": "<table>\n<thead>\n<tr>\n<th>Method Type</th>\n<th>Examples &amp; Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Filter Methods</strong></td>\n<td>Based on statistical tests (independent of models)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Variance Threshold (remove low variance features)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Correlation matrix (remove highly correlated features)</td>\n</tr>\n<tr>\n<td><strong>Wrapper Methods</strong></td>\n<td>Evaluate model performance with subsets of features</td>\n</tr>\n<tr>\n<td></td>\n<td>- Recursive Feature Elimination (RFE)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Sequential Feature Selector (SFS)</td>\n</tr>\n<tr>\n<td><strong>Embedded Methods</strong></td>\n<td>Feature selection is part of the model training</td>\n</tr>\n<tr>\n<td></td>\n<td>- Lasso (L1 Regularization)</td>\n</tr>\n<tr>\n<td></td>\n<td>- Tree-based feature importance (Random Forest, XGBoost)</td>\n</tr>\n</tbody>\n</table>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 3209537,
          "author_name": "Stable Space",
          "author_url": "",
          "post_date": "2025-05-26T02:44:28.310000",
          "content": "<p>These are some of the feature selection method. I hope it will help you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3209648,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-05-26T06:08:48.933000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/santopaulm\" target=\"_blank\">@santopaulm</a> <br>\nThank you for your comment. For me, this notebook has proven to be very useful, personally one of my favorite notebooks on Kaggle by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <a href=\"https://www.kaggle.com/code/ambrosm/mcts-eda-which-makes-sense\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/mcts-eda-which-makes-sense</a> <br>\nI got to learn alot about feature importance and selection from this. I'd suggest more than diving deep into the theory look for notebooks from playground competitions and read them. That'd make you understand the concept and it's implication in real time. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 3209776,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-26T09:46:23.320000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3209780,
          "author_name": "Vladimir Demidov",
          "author_url": "",
          "post_date": "2025-05-26T09:48:51.907000",
          "content": "<p>CatBoost has an inbuilt mechanism for feature selection (RFE) based on SHAP values; see the tutorial <a href=\"https://github.com/catboost/tutorials/blob/master/feature_selection/select_features_tutorial.ipynb\" target=\"_blank\">here</a>. If you are interested in how it works on real Comp data, I have a couple of public notebooks: <a href=\"https://www.kaggle.com/code/yekenot/feature-elimination-by-catboost\" target=\"_blank\"><em>Feature Elimination by CatBoost</em></a> (trading data, drop num feats from 124 to 100 with score improvement) and <a href=\"https://www.kaggle.com/code/yekenot/catboost-as-feature-selector\" target=\"_blank\"><em>CatBoost as Feature Selector</em></a> (banking data, drop num feats from 918 to 618 with score improvement).</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 3216458,
      "author_name": "Hem Ajit Patel",
      "author_url": "",
      "post_date": "2025-06-03T16:00:20.990000",
      "content": "<p>Thanks for sharing this unique and timely dataset! Crypto market behavior is complex and volatile, so having access to high-quality, granular data like this is a great resource for testing advanced time series models and market prediction strategies. Looking forward to seeing what the community builds with this!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3216509,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-06-03T17:08:45.563000",
          "content": "<p>Appreciate your words :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3218949,
      "author_name": "Andy Wang",
      "author_url": "",
      "post_date": "2025-06-07T00:58:25.757000",
      "content": "<p>I have seen a lot of different feature Selection &amp; Analysis in the discussion threads across competitions. This is definitely top-notch, very clear and the codes are concise. Great work.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3218992,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-06-07T03:22:07.030000",
          "content": "<p>Thanks Andy, appreciate your words! :))</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3213688,
      "author_name": "Vishal Painjane",
      "author_url": "",
      "post_date": "2025-05-30T10:29:05.340000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a>, the above info way really useful. Though, I gone through the dataset and had some observation. Data contains some low variance feature and some high variance features, but major features lay in a specific region. so, working on that sweet spot would get some meaningful results.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3213936,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-05-30T17:21:32.197000",
          "content": "<p>I certainly agree! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3210181,
      "author_name": "Ram Binay Gupta",
      "author_url": "",
      "post_date": "2025-05-26T20:50:17.900000",
      "content": "<p>Really helpful. Appreciate it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3209614,
      "author_name": "Sahil Islam007",
      "author_url": "",
      "post_date": "2025-05-26T04:57:44.460000",
      "content": "<p>this was helpful, wish to see this kind of insights more in future </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3209653,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-05-26T06:11:40.713000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/Sahil\" target=\"_blank\">@Sahil</a><br>\nThanks for the comment. Glad you found it helpful! :))</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3223646,
      "author_name": "Shoeb Ahmad Shamim",
      "author_url": "",
      "post_date": "2025-06-13T16:25:16.590000",
      "content": "<p>i found several problems when i transfrom features . this post really good for me . thank you so much &lt;3</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3220378,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "2025-06-09T08:08:59.983000",
      "content": "<p>feature selection is the key for this competition. Thanks for sharing your insights</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3220383,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-06-09T08:09:48.100000",
          "content": "<p>Exactly!! That plays a key role in this one! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3215023,
      "author_name": "Aniket Potabatti",
      "author_url": "",
      "post_date": "2025-06-01T13:24:53.213000",
      "content": "<p>what made you choose Incremental PCA over standard or kernel PCA?  Was it mainly a size thing, or did it give better results too?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3215476,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-06-02T09:45:46.277000",
          "content": "<p>Hey, Aniket! <a href=\"https://www.kaggle.com/aniketpotabatti\" target=\"_blank\">@aniketpotabatti</a> <br>\nSolely to avoid OOM issues.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3210110,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-26T18:27:13.620000",
      "content": "<p>Helpful guideline, very appreciate! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3209414,
      "author_name": "MaxUhl98",
      "author_url": "",
      "post_date": "2025-05-25T18:40:11.733000",
      "content": "<p>Did feature removal actually boost your cv score? In another post the auther shows results in which removing features outside of unduplicated/non constant features reduces the CV scores</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3209533,
          "author_name": "Stable Space",
          "author_url": "",
          "post_date": "2025-05-26T02:39:50.217000",
          "content": "<p>Well, it's not always the case, it depends on the problem type. Sometimes removing features make model simple and hence boost CV score and sometimes it can over fitting issues if not removed. So, removing features is relative and not absolute and depends on the condition.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3209649,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2025-05-26T06:10:55.033000",
          "content": "<p>Hey, Max! <br>\nI've not tried removing any features atm. I'd update as soon as I experiment with that. What <a href=\"https://www.kaggle.com/sharmajicoder\" target=\"_blank\">@sharmajicoder</a> said is right though, it's relative.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3213694,
          "author_name": "Vishal Painjane",
          "author_url": "",
          "post_date": "2025-05-30T10:38:47.223000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/maxuhl98\" target=\"_blank\">@maxuhl98</a>, features may contain information that is helpful for prediction, but not always. It's important to statistically evaluate how much each feature actually contributes to the CV scores.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3223178,
      "author_name": "mindy_nyc",
      "author_url": "",
      "post_date": "2025-06-13T01:12:17.697000",
      "content": "<p>This is helpful, thank you</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3213526,
      "author_name": "Samith Chimminiyan",
      "author_url": "",
      "post_date": "2025-05-30T06:05:07.620000",
      "content": "<p>Nice analysis. Thanks for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3210473,
      "author_name": "PeizheLi03",
      "author_url": "",
      "post_date": "2025-05-27T08:16:17.527000",
      "content": "<p>Thank you for your analysis!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3209405,
      "author_name": "Stable Space",
      "author_url": "",
      "post_date": "2025-05-25T18:17:23.073000",
      "content": "<p>Thanks for your insights 🤗🤗</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3216856,
      "author_name": "MANAN PATHAK",
      "author_url": "",
      "post_date": "2025-06-04T07:17:59.483000",
      "content": "<p>Thanks for sharing </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3215016,
      "author_name": "Jubayer Hasan",
      "author_url": "",
      "post_date": "2025-06-01T12:57:21.607000",
      "content": "<p>thanks! Great observation! </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3209240": "Hey, there! \n\nI’ve been digging deep into the features and wanted to share some insights that might help others. The dataset is large, anonymized, and noisy — so I focused a lot on feature redundancy, signal strength, and interpretability. \nYou can access the entire analysis in this notebook : https://www.kaggle.com/code/ahsuna123/anonymized-features-importance-selection-eda \n\nHere’s what stood out:\n\n**1. A Lot of Redundant Features :**\n\nI found over 20 pairs of features with extremely high correlation (Pearson > 0.98). Dropping one from each pair would help reduce dimensionality without losing information. \n\n**2. Weak Linear Correlation with the Target :**\n\nMost features individually have very weak correlation with the target. The best ones were only around ~0.07 — features like X21, X20, X28, and X863. This tells me that the signal isn’t in individual features but likely hidden in non-linear patterns or interactions — a good case for tree-based models or neural nets.\n\n**3. PCA Shows a Lot of Noise :**\n\nI ran Incremental PCA to understand the variance spread. Turns out:\n\n~90% variance is captured by just 23 components\n\n~95% by 36 components [Note: I only took top 50 features for computational ease]\n\n**4. Used 3 Different Feature Selection Techniques :**\n\nTo get a more robust handle on useful features, I compared:\n\n- Univariate F-test → picked out X21, X28, etc. (linear signal)\n\n- Random Forest Importance →  X198, X95, X179, etc. (non-linear)\n\n- RFE with Lasso → sparse feature selection that picked overlapping but also unique ones.\n\nInterestingly, 11 features consistently came up across all 3. I’m treating those as reliable “core signals.”\n\n**5. Clusters in the Data :**\n\nI ran KMeans on the dataset (k=5). Some clusters were heavily imbalanced in size — one had 231k+ samples! Early look at label distribution suggests that some clusters may be easier to classify than others. Planning to try cluster-based ensembling next.\n\nNote: You can tweak the code for your analysis ( I mostly showed top 20 features, but you can always expand).\n\nHope, it's helpful! :))",
    "3214616": "Correct me if I'm wrong, but a boosted tree based algorithm would be better at non-linear feature selection than a basic random forest.",
    "3209408": "Heyy this was really helpful. It would be really helpful if u could provide some sources where i could learn these feature selection techniques in a detailed and systematic manner. Thank you so much.",
    "3216458": "Thanks for sharing this unique and timely dataset! Crypto market behavior is complex and volatile, so having access to high-quality, granular data like this is a great resource for testing advanced time series models and market prediction strategies. Looking forward to seeing what the community builds with this!",
    "3218949": "I have seen a lot of different feature Selection & Analysis in the discussion threads across competitions. This is definitely top-notch, very clear and the codes are concise. Great work.",
    "3213688": "Hey @ahsuna123, the above info way really useful. Though, I gone through the dataset and had some observation. Data contains some low variance feature and some high variance features, but major features lay in a specific region. so, working on that sweet spot would get some meaningful results.",
    "3210181": "Really helpful. Appreciate it.",
    "3209614": "this was helpful, wish to see this kind of insights more in future ",
    "3223646": "i found several problems when i transfrom features . this post really good for me . thank you so much <3",
    "3220378": "feature selection is the key for this competition. Thanks for sharing your insights",
    "3215023": "what made you choose Incremental PCA over standard or kernel PCA?  Was it mainly a size thing, or did it give better results too?",
    "3210110": "Helpful guideline, very appreciate! ",
    "3209414": "Did feature removal actually boost your cv score? In another post the auther shows results in which removing features outside of unduplicated/non constant features reduces the CV scores",
    "3223178": "This is helpful, thank you",
    "3213526": "Nice analysis. Thanks for sharing. ",
    "3210473": "Thank you for your analysis!",
    "3209405": "Thanks for your insights 🤗🤗",
    "3216856": "Thanks for sharing ",
    "3215016": "thanks! Great observation! "
  }
}