{
  "id": 580600,
  "title": "Which features are most useful",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580600",
  "author_name": "Mahdi Ravaghi",
  "post_date": "2025-05-25T10:27:18.504000",
  "votes": 21,
  "comment_count": 7,
  "views": 0,
  "content": "<p>There are many features in this dataset, which makes feature engineering quite challenging, especially for the anonymous features. Instead of adding new features, I decided to run an experiment to see whether reducing the number of features could lead to any improvements. I removed the obvious choices, i.e. duplicate features and features with only one unique value. Then, I used the code below to extract the mutual information of each remaining feature:</p>\n<pre><code> sklearn.feature_selection  mutual_info_regression\n\nmutual_info = mutual_info_regression(X, y, random_state=)\n\nmutual_info = pd.Series(mutual_info)\nmutual_info.index = X.columns\nmutual_info = pd.DataFrame(mutual_info.sort_values(ascending=), columns=[])\n(mutual_info[mutual_info[] == ].index.values.tolist())\n</code></pre>\n<p>This resulted in the following features being identified as the top 20 most important.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16179447%2F56faa2d13eef949b853f053d387d78c5%2FSkjermbilde%202025-05-25%20101157.png?generation=1748168367008010&amp;alt=media\" alt=\"\"></p>\n<p>There were also some features with a mutual information score of 0. These are:</p>\n<pre><code>[\n    , , , , , , , , , , \n    , , , , , , , , , ,\n    , , , , , , , , , , \n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , , \n    , , , , , , , , , , \n    , , , , , , , , , , \n]\n</code></pre>\n<p>I retrained the models in my <a href=\"https://www.kaggle.com/code/ravaghi/drw-crypto-market-prediction-ensemble\" target=\"_blank\">public notebook</a> with these low-importance features removed, and observed that the CV scores dropped for all but one model. The table below shows the CV scores before and after feature reduction.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Model</strong></th>\n<th><strong>Before</strong></th>\n<th><strong>After</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGBoost</td>\n<td>0.106846</td>\n<td>0.100844</td>\n</tr>\n<tr>\n<td>LightGBM (goss)</td>\n<td>0.109056</td>\n<td>0.111739</td>\n</tr>\n<tr>\n<td>LightGBM (gbdt)</td>\n<td>0.105010</td>\n<td>0.104289</td>\n</tr>\n<tr>\n<td>Ridge (ensemble)</td>\n<td>0.122975</td>\n<td>0.119562</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 3209164,
      "postDate": "2025-05-25T10:27:18.503Z",
      "content": "<p>There are many features in this dataset, which makes feature engineering quite challenging, especially for the anonymous features. Instead of adding new features, I decided to run an experiment to see whether reducing the number of features could lead to any improvements. I removed the obvious choices, i.e. duplicate features and features with only one unique value. Then, I used the code below to extract the mutual information of each remaining feature:</p>\n<pre><code> sklearn.feature_selection  mutual_info_regression\n\nmutual_info = mutual_info_regression(X, y, random_state=)\n\nmutual_info = pd.Series(mutual_info)\nmutual_info.index = X.columns\nmutual_info = pd.DataFrame(mutual_info.sort_values(ascending=), columns=[])\n(mutual_info[mutual_info[] == ].index.values.tolist())\n</code></pre>\n<p>This resulted in the following features being identified as the top 20 most important.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16179447%2F56faa2d13eef949b853f053d387d78c5%2FSkjermbilde%202025-05-25%20101157.png?generation=1748168367008010&amp;alt=media\" alt=\"\"></p>\n<p>There were also some features with a mutual information score of 0. These are:</p>\n<pre><code>[\n    , , , , , , , , , , \n    , , , , , , , , , ,\n    , , , , , , , , , , \n    , , , , , , , , , ,\n    , , , , , , , , , ,\n    , , , , , , , , , , \n    , , , , , , , , , , \n    , , , , , , , , , , \n]\n</code></pre>\n<p>I retrained the models in my <a href=\"https://www.kaggle.com/code/ravaghi/drw-crypto-market-prediction-ensemble\" target=\"_blank\">public notebook</a> with these low-importance features removed, and observed that the CV scores dropped for all but one model. The table below shows the CV scores before and after feature reduction.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Model</strong></th>\n<th><strong>Before</strong></th>\n<th><strong>After</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGBoost</td>\n<td>0.106846</td>\n<td>0.100844</td>\n</tr>\n<tr>\n<td>LightGBM (goss)</td>\n<td>0.109056</td>\n<td>0.111739</td>\n</tr>\n<tr>\n<td>LightGBM (gbdt)</td>\n<td>0.105010</td>\n<td>0.104289</td>\n</tr>\n<tr>\n<td>Ridge (ensemble)</td>\n<td>0.122975</td>\n<td>0.119562</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "There are many features in this dataset, which makes feature engineering quite challenging, especially for the anonymous features. Instead of adding new features, I decided to run an experiment to see whether reducing the number of features could lead to any improvements. I removed the obvious choices, i.e. duplicate features and features with only one unique value. Then, I used the code below to extract the mutual information of each remaining feature:\n\n```python\nfrom sklearn.feature_selection import mutual_info_regression\n\nmutual_info = mutual_info_regression(X, y, random_state=42)\n\nmutual_info = pd.Series(mutual_info)\nmutual_info.index = X.columns\nmutual_info = pd.DataFrame(mutual_info.sort_values(ascending=False), columns=['Mutual Information'])\nprint(mutual_info[mutual_info['Mutual Information'] == 0].index.values.tolist())\n```\n\nThis resulted in the following features being identified as the top 20 most important.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16179447%2F56faa2d13eef949b853f053d387d78c5%2FSkjermbilde%202025-05-25%20101157.png?generation=1748168367008010&alt=media)\n\nThere were also some features with a mutual information score of 0. These are:\n\n```python\n[\n    'X100', 'X101', 'X102', 'X114', 'X143', 'X237', 'X240', 'X242', 'X244', 'X246', \n    'X266', 'X306', 'X308', 'X309', 'X327', 'X348', 'X392', 'X394', 'X481', 'X484',\n    'X486', 'X492', 'X501', 'X508', 'X515', 'X519', 'X521', 'X522', 'X528', 'X529', \n    'X533', 'X534', 'X535', 'X536', 'X541', 'X542', 'X543', 'X544', 'X550', 'X563',\n    'X569', 'X570', 'X571', 'X576', 'X577', 'X578', 'X582', 'X585', 'X590', 'X591',\n    'X593', 'X597', 'X61', 'X615', 'X616', 'X618', 'X62', 'X621', 'X622', 'X624', \n    'X625', 'X627', 'X628', 'X630', 'X633', 'X642', 'X648', 'X670', 'X68', 'X739', \n    'X74', 'X740', 'X749', 'X767', 'X775', 'X80', 'X807', 'X811', 'X815', 'X848', 'X850'\n]\n```\n\nI retrained the models in my [public notebook](https://www.kaggle.com/code/ravaghi/drw-crypto-market-prediction-ensemble) with these low-importance features removed, and observed that the CV scores dropped for all but one model. The table below shows the CV scores before and after feature reduction.\n|     **Model**    | **Before** | **After** |\n|:----------------:|:----------:|:---------:|\n|      XGBoost     |  0.106846  |  0.100844 |\n|  LightGBM (goss) |  0.109056  |  0.111739 |\n|  LightGBM (gbdt) |  0.105010  |  0.104289 |\n| Ridge (ensemble) |  0.122975  |  0.119562 |",
      "votes": 20
    },
    {
      "id": 3226015,
      "postDate": "2025-06-17T06:36:15.853Z",
      "content": "<p>I think it's better to use SelectFromModel</p>",
      "rawMarkdown": "I think it's better to use SelectFromModel",
      "votes": 1
    },
    {
      "id": 3209202,
      "postDate": "2025-05-25T11:10:46.297Z",
      "content": "<p><a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a> I just feel that masking the time stamps in the test set is not boding well for the problem. This is not a simple regression problem as we are well aware of the temporal nature of the dataset, but using this method may not yield the best results on an anonymized time-stamp test set!</p>\n<p>Also, no time series model can create a 13-month forward prediction without further work also. I will think deeply and continue my efforts here!</p>",
      "rawMarkdown": "@ravaghi I just feel that masking the time stamps in the test set is not boding well for the problem. This is not a simple regression problem as we are well aware of the temporal nature of the dataset, but using this method may not yield the best results on an anonymized time-stamp test set!\n\nAlso, no time series model can create a 13-month forward prediction without further work also. I will think deeply and continue my efforts here!",
      "votes": 2,
      "replies": [
        {
          "id": 3209247,
          "postDate": "2025-05-25T12:46:50.623Z",
          "content": "<p>I completely agree. This should have been a time series forecasting competition.</p>",
          "rawMarkdown": "I completely agree. This should have been a time series forecasting competition.",
          "votes": 1,
          "replies": [
            {
              "id": 3239250,
              "postDate": "2025-07-02T16:03:41.153Z",
              "content": "<p>This is like training a model to predict tomorrow's weather but then asking it to predict weather for \"some random day\" without knowing when that day is! </p>\n<p>On the other hand this is a <strong>clever strategy by the hosts</strong> - it seems they are more eager to find out fundamental patterns than just temporal, trend patterns, which may not generalize for different market regimes.</p>",
              "rawMarkdown": "This is like training a model to predict tomorrow's weather but then asking it to predict weather for \"some random day\" without knowing when that day is! \n\nOn the other hand this is a **clever strategy by the hosts** - it seems they are more eager to find out fundamental patterns than just temporal, trend patterns, which may not generalize for different market regimes."
            }
          ]
        }
      ]
    },
    {
      "id": 3209201,
      "postDate": "2025-05-25T11:08:00.577Z",
      "content": "<p>mutual info seems to have very low calculation speed </p>",
      "rawMarkdown": "mutual info seems to have very low calculation speed ",
      "votes": -2
    },
    {
      "id": 3209372,
      "postDate": "2025-05-25T17:05:17.393Z",
      "content": "<p>Whenever I work with Datasets, I used to select features which are highly related to the target feature and in most of the cases i just decide the feature importance by intuition which obviously is not a good way. I often struggle with this step of feature selection. Could you help me tackle this issue?</p>",
      "rawMarkdown": "Whenever I work with Datasets, I used to select features which are highly related to the target feature and in most of the cases i just decide the feature importance by intuition which obviously is not a good way. I often struggle with this step of feature selection. Could you help me tackle this issue?",
      "replies": [
        {
          "id": 3209442,
          "postDate": "2025-05-25T20:22:36.367Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3226015,
      "author_name": "Rohïth",
      "author_url": "",
      "post_date": "2025-06-17T06:36:15.853000",
      "content": "<p>I think it's better to use SelectFromModel</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3209202,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2025-05-25T11:10:46.297000",
      "content": "<p><a href=\"https://www.kaggle.com/ravaghi\" target=\"_blank\">@ravaghi</a> I just feel that masking the time stamps in the test set is not boding well for the problem. This is not a simple regression problem as we are well aware of the temporal nature of the dataset, but using this method may not yield the best results on an anonymized time-stamp test set!</p>\n<p>Also, no time series model can create a 13-month forward prediction without further work also. I will think deeply and continue my efforts here!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3209247,
          "author_name": "Mahdi Ravaghi",
          "author_url": "",
          "post_date": "2025-05-25T12:46:50.623000",
          "content": "<p>I completely agree. This should have been a time series forecasting competition.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3239250,
              "author_name": "BJBA99",
              "author_url": "",
              "post_date": "2025-07-02T16:03:41.153000",
              "content": "<p>This is like training a model to predict tomorrow's weather but then asking it to predict weather for \"some random day\" without knowing when that day is! </p>\n<p>On the other hand this is a <strong>clever strategy by the hosts</strong> - it seems they are more eager to find out fundamental patterns than just temporal, trend patterns, which may not generalize for different market regimes.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3209201,
      "author_name": "Wenxiao Sun",
      "author_url": "",
      "post_date": "2025-05-25T11:08:00.577000",
      "content": "<p>mutual info seems to have very low calculation speed </p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 3209372,
      "author_name": "Stable Space",
      "author_url": "",
      "post_date": "2025-05-25T17:05:17.393000",
      "content": "<p>Whenever I work with Datasets, I used to select features which are highly related to the target feature and in most of the cases i just decide the feature importance by intuition which obviously is not a good way. I often struggle with this step of feature selection. Could you help me tackle this issue?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3209442,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-25T20:22:36.367000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3209164": "There are many features in this dataset, which makes feature engineering quite challenging, especially for the anonymous features. Instead of adding new features, I decided to run an experiment to see whether reducing the number of features could lead to any improvements. I removed the obvious choices, i.e. duplicate features and features with only one unique value. Then, I used the code below to extract the mutual information of each remaining feature:\n\n```python\nfrom sklearn.feature_selection import mutual_info_regression\n\nmutual_info = mutual_info_regression(X, y, random_state=42)\n\nmutual_info = pd.Series(mutual_info)\nmutual_info.index = X.columns\nmutual_info = pd.DataFrame(mutual_info.sort_values(ascending=False), columns=['Mutual Information'])\nprint(mutual_info[mutual_info['Mutual Information'] == 0].index.values.tolist())\n```\n\nThis resulted in the following features being identified as the top 20 most important.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16179447%2F56faa2d13eef949b853f053d387d78c5%2FSkjermbilde%202025-05-25%20101157.png?generation=1748168367008010&alt=media)\n\nThere were also some features with a mutual information score of 0. These are:\n\n```python\n[\n    'X100', 'X101', 'X102', 'X114', 'X143', 'X237', 'X240', 'X242', 'X244', 'X246', \n    'X266', 'X306', 'X308', 'X309', 'X327', 'X348', 'X392', 'X394', 'X481', 'X484',\n    'X486', 'X492', 'X501', 'X508', 'X515', 'X519', 'X521', 'X522', 'X528', 'X529', \n    'X533', 'X534', 'X535', 'X536', 'X541', 'X542', 'X543', 'X544', 'X550', 'X563',\n    'X569', 'X570', 'X571', 'X576', 'X577', 'X578', 'X582', 'X585', 'X590', 'X591',\n    'X593', 'X597', 'X61', 'X615', 'X616', 'X618', 'X62', 'X621', 'X622', 'X624', \n    'X625', 'X627', 'X628', 'X630', 'X633', 'X642', 'X648', 'X670', 'X68', 'X739', \n    'X74', 'X740', 'X749', 'X767', 'X775', 'X80', 'X807', 'X811', 'X815', 'X848', 'X850'\n]\n```\n\nI retrained the models in my [public notebook](https://www.kaggle.com/code/ravaghi/drw-crypto-market-prediction-ensemble) with these low-importance features removed, and observed that the CV scores dropped for all but one model. The table below shows the CV scores before and after feature reduction.\n|     **Model**    | **Before** | **After** |\n|:----------------:|:----------:|:---------:|\n|      XGBoost     |  0.106846  |  0.100844 |\n|  LightGBM (goss) |  0.109056  |  0.111739 |\n|  LightGBM (gbdt) |  0.105010  |  0.104289 |\n| Ridge (ensemble) |  0.122975  |  0.119562 |",
    "3226015": "I think it's better to use SelectFromModel",
    "3209202": "@ravaghi I just feel that masking the time stamps in the test set is not boding well for the problem. This is not a simple regression problem as we are well aware of the temporal nature of the dataset, but using this method may not yield the best results on an anonymized time-stamp test set!\n\nAlso, no time series model can create a 13-month forward prediction without further work also. I will think deeply and continue my efforts here!",
    "3209201": "mutual info seems to have very low calculation speed ",
    "3209372": "Whenever I work with Datasets, I used to select features which are highly related to the target feature and in most of the cases i just decide the feature importance by intuition which obviously is not a good way. I often struggle with this step of feature selection. Could you help me tackle this issue?"
  }
}