{
  "id": 580291,
  "title": "Redundant columns (infinity and zeroes).",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580291",
  "author_name": "",
  "post_date": "2025-05-23T15:27:53.512464400Z",
  "votes": 26,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've just done some EDA with the following results:</p>\n<p>Features X697..X717 have only -infinity values.<br>\nFeatures X864, X867 and X869..X872 are always 0.<br>\nTherefore, those columns can be removed.</p>\n<p>More details see here: <a href=\"https://www.kaggle.com/code/docxian/drw-crypto-prediction-explore-data\" target=\"_blank\">https://www.kaggle.com/code/docxian/drw-crypto-prediction-explore-data</a> </p>",
  "messages": [
    {
      "id": "3208025",
      "postDate": "05/23/2025 15:27:53",
      "content": "<p>I've just done some EDA with the following results:</p>\n<p>Features X697..X717 have only -infinity values.<br>\nFeatures X864, X867 and X869..X872 are always 0.<br>\nTherefore, those columns can be removed.</p>\n<p>More details see here: <a href=\"https://www.kaggle.com/code/docxian/drw-crypto-prediction-explore-data\" target=\"_blank\">https://www.kaggle.com/code/docxian/drw-crypto-prediction-explore-data</a> </p>",
      "rawMarkdown": "I've just done some EDA with the following results:\n\nFeatures X697..X717 have only -infinity values.\nFeatures X864, X867 and X869..X872 are always 0.\nTherefore, those columns can be removed.\n\nMore details see here: https://www.kaggle.com/code/docxian/drw-crypto-prediction-explore-data",
      "votes": null
    },
    {
      "id": "3208460",
      "postDate": "05/24/2025 07:20:29",
      "content": "<p>Thanks for sharing.<br>\nIn addition to the columns you found, I checked for duplicated columns in train. The dict below lists columns containing the exact same information. Might be worth checking for fuzzy ones too. </p>\n<p><code>{'X62': ['X104', 'X146'], 'X68': ['X110', 'X152'], 'X74': ['X116', 'X158'], 'X80': ['X122', 'X164'], 'X86': ['X128', 'X170'], 'X92': ['X134', 'X176'], 'X98': ['X140', 'X182'], 'X309': ['X351', 'X393'], 'X315': ['X357', 'X399'], 'X321': ['X363', 'X405'], 'X327': ['X369', 'X411'], 'X333': ['X375', 'X417'], 'X339': ['X381', 'X423'], 'X345': ['X387', 'X429'], 'X697': ['X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705', 'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714', 'X715', 'X716', 'X717'], 'X864': ['X867', 'X869', 'X870', 'X871', 'X872']}</code></p>",
      "rawMarkdown": "Thanks for sharing.\nIn addition to the columns you found, I checked for duplicated columns in train. The dict below lists columns containing the exact same information. Might be worth checking for fuzzy ones too. \n\n`{'X62': ['X104', 'X146'], 'X68': ['X110', 'X152'], 'X74': ['X116', 'X158'], 'X80': ['X122', 'X164'], 'X86': ['X128', 'X170'], 'X92': ['X134', 'X176'], 'X98': ['X140', 'X182'], 'X309': ['X351', 'X393'], 'X315': ['X357', 'X399'], 'X321': ['X363', 'X405'], 'X327': ['X369', 'X411'], 'X333': ['X375', 'X417'], 'X339': ['X381', 'X423'], 'X345': ['X387', 'X429'], 'X697': ['X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705', 'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714', 'X715', 'X716', 'X717'], 'X864': ['X867', 'X869', 'X870', 'X871', 'X872']}`",
      "votes": null
    },
    {
      "id": "3208776",
      "postDate": "05/24/2025 16:58:15",
      "content": "<p>so helpful! Could you share your code?</p>",
      "rawMarkdown": "so helpful! Could you share your code?",
      "votes": null
    },
    {
      "id": "3209276",
      "postDate": "05/25/2025 13:46:18",
      "content": "<p>Good catch, <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a>! Inspired by your post I added a calculation of all Pearson correlations for the X-features to the notebook mentioned above. After removing the 0 and infinity columns there still remain 42 [ordered] pairs of features with 100% correlation (this includes the duplicates you found and a few more). In addition, we find further 71 pairs with Pearson correlation &gt; 0.99.</p>",
      "rawMarkdown": "Good catch, @lennarthaupts! Inspired by your post I added a calculation of all Pearson correlations for the X-features to the notebook mentioned above. After removing the 0 and infinity columns there still remain 42 [ordered] pairs of features with 100% correlation (this includes the duplicates you found and a few more). In addition, we find further 71 pairs with Pearson correlation > 0.99.",
      "votes": null
    },
    {
      "id": "3209294",
      "postDate": "05/25/2025 14:27:51",
      "content": "<p>It's a little strange that after removing features with 0.99 &lt; corr &lt; 0.1, my performance becomes worse.</p>",
      "rawMarkdown": "It's a little strange that after removing features with 0.99 < corr < 0.1, my performance becomes worse.",
      "votes": null
    },
    {
      "id": "3209381",
      "postDate": "05/25/2025 17:30:11",
      "content": "<p>Thanks for the info.</p>",
      "rawMarkdown": "Thanks for the info.",
      "votes": null
    },
    {
      "id": "3209625",
      "postDate": "05/26/2025 05:20:44",
      "content": "<p>Great analysis <a href=\"https://www.kaggle.com/ChrisX\" target=\"_blank\">@ChrisX</a> and <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a>! </p>\n<p><a href=\"https://www.kaggle.com/YouruiWang\" target=\"_blank\">@YouruiWang</a> - that's actually expected behavior! High correlation ≠ redundancy in time series. Here's why:</p>\n<p><strong>Correlated features can still add value:</strong></p>\n<ul>\n<li>Different noise patterns</li>\n<li>Temporal lag relationships  </li>\n<li>Non-linear interactions</li>\n</ul>\n<p><strong>Better approach for crypto data:</strong></p>\n<ul>\n<li>Use <strong>variance inflation factor (VIF)</strong> instead of correlation</li>\n<li>Try <strong>PCA/ICA</strong> to capture orthogonal signals</li>\n<li>Keep features if they improve ensemble diversity</li>\n</ul>\n<p>The infinity/zero columns are definitely safe to remove though! 🚀</p>",
      "rawMarkdown": "Great analysis @ChrisX and @lennarthaupts! \n\n@YouruiWang - that's actually expected behavior! High correlation ≠ redundancy in time series. Here's why:\n\n**Correlated features can still add value:**\n- Different noise patterns\n- Temporal lag relationships  \n- Non-linear interactions\n\n**Better approach for crypto data:**\n- Use **variance inflation factor (VIF)** instead of correlation\n- Try **PCA/ICA** to capture orthogonal signals\n- Keep features if they improve ensemble diversity\n\nThe infinity/zero columns are definitely safe to remove though! 🚀",
      "votes": null
    },
    {
      "id": "3210220",
      "postDate": "05/26/2025 22:36:33",
      "content": "<pre><code> itertools  combinations\n\n ():\n    equal_pairs = []\n     c1, c2  combinations(df.columns, ):\n        \n         df[c1].equals(df[c2]):\n            (c1, c2)\n            equal_pairs.append((c1, c2))\n     equal_pairs\n\npairs = find_equal_column_pairs(df)\nto_drop = (((*pairs))[])\n</code></pre>",
      "rawMarkdown": "```python\nfrom itertools import combinations\n\ndef find_equal_column_pairs(df: pd.DataFrame):\n    equal_pairs = []\n    for c1, c2 in combinations(df.columns, 2):\n        # .equals checks every element (and NaNs in the same positions)\n        if df[c1].equals(df[c2]):\n            print(c1, c2)\n            equal_pairs.append((c1, c2))\n    return equal_pairs\n\npairs = find_equal_column_pairs(df)\nto_drop = set(list(zip(*pairs))[1])\n```",
      "votes": null
    },
    {
      "id": "3210690",
      "postDate": "05/27/2025 15:41:30",
      "content": "<p>cols = [\"X864\", \"X867\", \"X869\", \"X870\", \"X871\", \"X872\"] </p>\n<p>do have all zero values in the train set, but they seem to have regular values in the test set as far as i can tell. or they are just noise there. still it makes me question if simply dropping them is the best solution?</p>",
      "rawMarkdown": "cols = [\"X864\", \"X867\", \"X869\", \"X870\", \"X871\", \"X872\"] \n\ndo have all zero values in the train set, but they seem to have regular values in the test set as far as i can tell. or they are just noise there. still it makes me question if simply dropping them is the best solution?",
      "votes": null
    },
    {
      "id": "3210716",
      "postDate": "05/27/2025 16:04:08",
      "content": "<p>Yes for sure. No matter tree-based model or nn they both learn nothing from constant.</p>",
      "rawMarkdown": "Yes for sure. No matter tree-based model or nn they both learn nothing from constant.",
      "votes": null
    },
    {
      "id": "3211605",
      "postDate": "05/28/2025 15:44:11",
      "content": "<p>that is good to know , i was thinking about this, </p>",
      "rawMarkdown": "that is good to know , i was thinking about this,",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3208460,
      "author_name": "lennarthaupts",
      "author_url": "",
      "post_date": "05/24/2025 07:20:29",
      "content": "<p>Thanks for sharing.<br>\nIn addition to the columns you found, I checked for duplicated columns in train. The dict below lists columns containing the exact same information. Might be worth checking for fuzzy ones too. </p>\n<p><code>{'X62': ['X104', 'X146'], 'X68': ['X110', 'X152'], 'X74': ['X116', 'X158'], 'X80': ['X122', 'X164'], 'X86': ['X128', 'X170'], 'X92': ['X134', 'X176'], 'X98': ['X140', 'X182'], 'X309': ['X351', 'X393'], 'X315': ['X357', 'X399'], 'X321': ['X363', 'X405'], 'X327': ['X369', 'X411'], 'X333': ['X375', 'X417'], 'X339': ['X381', 'X423'], 'X345': ['X387', 'X429'], 'X697': ['X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705', 'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714', 'X715', 'X716', 'X717'], 'X864': ['X867', 'X869', 'X870', 'X871', 'X872']}</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 3208776,
          "author_name": "andrewguanzc",
          "author_url": "",
          "post_date": "05/24/2025 16:58:15",
          "content": "<p>so helpful! Could you share your code?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3210220,
              "author_name": "gtoking",
              "author_url": "",
              "post_date": "05/26/2025 22:36:33",
              "content": "<pre><code> itertools  combinations\n\n ():\n    equal_pairs = []\n     c1, c2  combinations(df.columns, ):\n        \n         df[c1].equals(df[c2]):\n            (c1, c2)\n            equal_pairs.append((c1, c2))\n     equal_pairs\n\npairs = find_equal_column_pairs(df)\nto_drop = (((*pairs))[])\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3209276,
          "author_name": "docxian",
          "author_url": "",
          "post_date": "05/25/2025 13:46:18",
          "content": "<p>Good catch, <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a>! Inspired by your post I added a calculation of all Pearson correlations for the X-features to the notebook mentioned above. After removing the 0 and infinity columns there still remain 42 [ordered] pairs of features with 100% correlation (this includes the duplicates you found and a few more). In addition, we find further 71 pairs with Pearson correlation &gt; 0.99.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3209294,
              "author_name": "yw2735",
              "author_url": "",
              "post_date": "05/25/2025 14:27:51",
              "content": "<p>It's a little strange that after removing features with 0.99 &lt; corr &lt; 0.1, my performance becomes worse.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3209381,
      "author_name": "sharmajicoder",
      "author_url": "",
      "post_date": "05/25/2025 17:30:11",
      "content": "<p>Thanks for the info.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3209625,
      "author_name": "harshithvaddiparthy",
      "author_url": "",
      "post_date": "05/26/2025 05:20:44",
      "content": "<p>Great analysis <a href=\"https://www.kaggle.com/ChrisX\" target=\"_blank\">@ChrisX</a> and <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a>! </p>\n<p><a href=\"https://www.kaggle.com/YouruiWang\" target=\"_blank\">@YouruiWang</a> - that's actually expected behavior! High correlation ≠ redundancy in time series. Here's why:</p>\n<p><strong>Correlated features can still add value:</strong></p>\n<ul>\n<li>Different noise patterns</li>\n<li>Temporal lag relationships  </li>\n<li>Non-linear interactions</li>\n</ul>\n<p><strong>Better approach for crypto data:</strong></p>\n<ul>\n<li>Use <strong>variance inflation factor (VIF)</strong> instead of correlation</li>\n<li>Try <strong>PCA/ICA</strong> to capture orthogonal signals</li>\n<li>Keep features if they improve ensemble diversity</li>\n</ul>\n<p>The infinity/zero columns are definitely safe to remove though! 🚀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3210690,
      "author_name": "wolfstefan",
      "author_url": "",
      "post_date": "05/27/2025 15:41:30",
      "content": "<p>cols = [\"X864\", \"X867\", \"X869\", \"X870\", \"X871\", \"X872\"] </p>\n<p>do have all zero values in the train set, but they seem to have regular values in the test set as far as i can tell. or they are just noise there. still it makes me question if simply dropping them is the best solution?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3210716,
          "author_name": "yw2735",
          "author_url": "",
          "post_date": "05/27/2025 16:04:08",
          "content": "<p>Yes for sure. No matter tree-based model or nn they both learn nothing from constant.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3211605,
      "author_name": "hamzabinbutt",
      "author_url": "",
      "post_date": "05/28/2025 15:44:11",
      "content": "<p>that is good to know , i was thinking about this, </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3208025": "I've just done some EDA with the following results:\n\nFeatures X697..X717 have only -infinity values.\nFeatures X864, X867 and X869..X872 are always 0.\nTherefore, those columns can be removed.\n\nMore details see here: https://www.kaggle.com/code/docxian/drw-crypto-prediction-explore-data",
    "3208460": "Thanks for sharing.\nIn addition to the columns you found, I checked for duplicated columns in train. The dict below lists columns containing the exact same information. Might be worth checking for fuzzy ones too. \n\n`{'X62': ['X104', 'X146'], 'X68': ['X110', 'X152'], 'X74': ['X116', 'X158'], 'X80': ['X122', 'X164'], 'X86': ['X128', 'X170'], 'X92': ['X134', 'X176'], 'X98': ['X140', 'X182'], 'X309': ['X351', 'X393'], 'X315': ['X357', 'X399'], 'X321': ['X363', 'X405'], 'X327': ['X369', 'X411'], 'X333': ['X375', 'X417'], 'X339': ['X381', 'X423'], 'X345': ['X387', 'X429'], 'X697': ['X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705', 'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714', 'X715', 'X716', 'X717'], 'X864': ['X867', 'X869', 'X870', 'X871', 'X872']}`",
    "3208776": "so helpful! Could you share your code?",
    "3209276": "Good catch, @lennarthaupts! Inspired by your post I added a calculation of all Pearson correlations for the X-features to the notebook mentioned above. After removing the 0 and infinity columns there still remain 42 [ordered] pairs of features with 100% correlation (this includes the duplicates you found and a few more). In addition, we find further 71 pairs with Pearson correlation > 0.99.",
    "3209294": "It's a little strange that after removing features with 0.99 < corr < 0.1, my performance becomes worse.",
    "3209381": "Thanks for the info.",
    "3209625": "Great analysis @ChrisX and @lennarthaupts! \n\n@YouruiWang - that's actually expected behavior! High correlation ≠ redundancy in time series. Here's why:\n\n**Correlated features can still add value:**\n- Different noise patterns\n- Temporal lag relationships  \n- Non-linear interactions\n\n**Better approach for crypto data:**\n- Use **variance inflation factor (VIF)** instead of correlation\n- Try **PCA/ICA** to capture orthogonal signals\n- Keep features if they improve ensemble diversity\n\nThe infinity/zero columns are definitely safe to remove though! 🚀",
    "3210220": "```python\nfrom itertools import combinations\n\ndef find_equal_column_pairs(df: pd.DataFrame):\n    equal_pairs = []\n    for c1, c2 in combinations(df.columns, 2):\n        # .equals checks every element (and NaNs in the same positions)\n        if df[c1].equals(df[c2]):\n            print(c1, c2)\n            equal_pairs.append((c1, c2))\n    return equal_pairs\n\npairs = find_equal_column_pairs(df)\nto_drop = set(list(zip(*pairs))[1])\n```",
    "3210690": "cols = [\"X864\", \"X867\", \"X869\", \"X870\", \"X871\", \"X872\"] \n\ndo have all zero values in the train set, but they seem to have regular values in the test set as far as i can tell. or they are just noise there. still it makes me question if simply dropping them is the best solution?",
    "3210716": "Yes for sure. No matter tree-based model or nn they both learn nothing from constant.",
    "3211605": "that is good to know , i was thinking about this,"
  },
  "source": "meta"
}