{
  "id": 580704,
  "title": "Highly Correlated Features",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580704",
  "author_name": "",
  "post_date": "2025-05-26T04:31:21.897898400Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<pre><code># high correlation columns\nextreme_high_corr_features = dict()\nextreme_high_corr_features_list = []\n\nX_columns = [   in train_df.columns   not in MD_COLUMNS]\n\n i,  in tqdm(enumerate(X_columns), total=(X_columns)):\n      in extreme_high_corr_features_list:\n        \n    high_corr_wth_cf = []\n     cf2 in X_columns[i+:]:\n        corr_XX = pearsonr(train_df[], train_df[cf2]).statistic\n         corr_XX &gt;= :\n            high_corr_wth_cf.(cf2)\n            extreme_high_corr_features_list.(cf2)\n     (high_corr_wth_cf) &gt; :\n        extreme_high_corr_features[] = high_corr_wth_cf\n        ()\n</code></pre>\n<p>For your convenience:</p>\n<pre><code>extreme_high_corr_features =  {: [], : [], : [], : [], : [], : [, ], : [, ], : [, ], : [, ], : [, , , ], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [, ], : [, ], : [, ], : [, ], : [, , ], : [], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : []}\n</code></pre>",
  "messages": [
    {
      "id": "3209601",
      "postDate": "05/26/2025 04:31:21",
      "content": "<pre><code># high correlation columns\nextreme_high_corr_features = dict()\nextreme_high_corr_features_list = []\n\nX_columns = [   in train_df.columns   not in MD_COLUMNS]\n\n i,  in tqdm(enumerate(X_columns), total=(X_columns)):\n      in extreme_high_corr_features_list:\n        \n    high_corr_wth_cf = []\n     cf2 in X_columns[i+:]:\n        corr_XX = pearsonr(train_df[], train_df[cf2]).statistic\n         corr_XX &gt;= :\n            high_corr_wth_cf.(cf2)\n            extreme_high_corr_features_list.(cf2)\n     (high_corr_wth_cf) &gt; :\n        extreme_high_corr_features[] = high_corr_wth_cf\n        ()\n</code></pre>\n<p>For your convenience:</p>\n<pre><code>extreme_high_corr_features =  {: [], : [], : [], : [], : [], : [, ], : [, ], : [, ], : [, ], : [, , , ], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [, ], : [, ], : [, ], : [, ], : [, , ], : [], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [, ], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : [], : []}\n</code></pre>",
      "rawMarkdown": "```\n# high correlation columns\nextreme_high_corr_features = dict()\nextreme_high_corr_features_list = []\n\nX_columns = [c for c in train_df.columns if c not in MD_COLUMNS]\n\nfor i, cf in tqdm(enumerate(X_columns), total=len(X_columns)):\n    if cf in extreme_high_corr_features_list:\n        continue\n    high_corr_wth_cf = []\n    for cf2 in X_columns[i+1:]:\n        corr_XX = pearsonr(train_df[cf], train_df[cf2]).statistic\n        if corr_XX >= 0.98:\n            high_corr_wth_cf.append(cf2)\n            extreme_high_corr_features_list.append(cf2)\n    if len(high_corr_wth_cf) > 0:\n        extreme_high_corr_features[cf] = high_corr_wth_cf\n        print(f\"i={i} ==> {extreme_high_corr_features}\")\n```\n\nFor your convenience:\n\n```\nextreme_high_corr_features =  {'X1': ['X8'], 'X6': ['X7'], 'X14': ['X15'], 'X33': ['X35'], 'X34': ['X36'], 'X39': ['X41', 'X43'], 'X40': ['X42', 'X44'], 'X45': ['X47', 'X49'], 'X46': ['X48', 'X50'], 'X51': ['X52', 'X53', 'X54', 'X55'], 'X62': ['X104', 'X146'], 'X68': ['X110', 'X152'], 'X74': ['X116', 'X158'], 'X80': ['X122', 'X164'], 'X86': ['X128', 'X170'], 'X92': ['X134', 'X176'], 'X98': ['X140', 'X182'], 'X185': ['X188'], 'X234': ['X241'], 'X235': ['X242'], 'X236': ['X243'], 'X237': ['X244'], 'X238': ['X245'], 'X239': ['X246'], 'X240': ['X247'], 'X253': ['X254'], 'X262': ['X263'], 'X280': ['X282'], 'X281': ['X283'], 'X286': ['X288', 'X290'], 'X287': ['X289', 'X291'], 'X292': ['X294', 'X296'], 'X293': ['X295', 'X297'], 'X298': ['X299', 'X300', 'X302'], 'X301': ['X303'], 'X309': ['X351', 'X393'], 'X315': ['X357', 'X399'], 'X321': ['X363', 'X405'], 'X327': ['X369', 'X411'], 'X333': ['X375', 'X417'], 'X339': ['X381', 'X423'], 'X345': ['X387', 'X429'], 'X431': ['X434'], 'X432': ['X435'], 'X481': ['X488'], 'X482': ['X489'], 'X483': ['X490'], 'X484': ['X491'], 'X485': ['X492'], 'X486': ['X493'], 'X487': ['X494'], 'X613': ['X619'], 'X616': ['X622'], 'X625': ['X631'], 'X628': ['X634'], 'X637': ['X643'], 'X640': ['X646'], 'X649': ['X655'], 'X652': ['X658'], 'X661': ['X667'], 'X663': ['X669'], 'X664': ['X670'], 'X666': ['X672'], 'X673': ['X679'], 'X675': ['X681'], 'X676': ['X682'], 'X678': ['X684'], 'X685': ['X691'], 'X688': ['X694'], 'X690': ['X696'], 'X788': ['X789'], 'X792': ['X793'], 'X796': ['X797'], 'X800': ['X801'], 'X804': ['X805'], 'X816': ['X817'], 'X820': ['X821'], 'X824': ['X825'], 'X884': ['X885'], 'X886': ['X887']}\n```",
      "votes": null
    },
    {
      "id": "3209608",
      "postDate": "05/26/2025 04:52:53",
      "content": "<p>Your code is designed to identify highly correlated feature pairs (Pearson correlation coefficient ≥ 0.98) within a DataFrame.</p>",
      "rawMarkdown": "Your code is designed to identify highly correlated feature pairs (Pearson correlation coefficient ≥ 0.98) within a DataFrame.",
      "votes": null
    },
    {
      "id": "3209622",
      "postDate": "05/26/2025 05:15:48",
      "content": "<p>exactly, for training datasets</p>",
      "rawMarkdown": "exactly, for training datasets",
      "votes": null
    },
    {
      "id": "3209631",
      "postDate": "05/26/2025 05:36:10",
      "content": "<p>But, sadly, looks like rm high corr Xs makes the model performance worse 🤣</p>",
      "rawMarkdown": "But, sadly, looks like rm high corr Xs makes the model performance worse 🤣",
      "votes": null
    },
    {
      "id": "3210026",
      "postDate": "05/26/2025 16:16:02",
      "content": "<p><a href=\"https://www.kaggle.com/henrysun\" target=\"_blank\">@henrysun</a>  hey ! 0.99 , too strict I guess ! also , have you considering negative correlation ?</p>",
      "rawMarkdown": "henrysun  hey ! 0.99 , too strict I guess ! also , have you considering negative correlation ?",
      "votes": null
    },
    {
      "id": "3210369",
      "postDate": "05/27/2025 05:02:12",
      "content": "<p>Can confirm haha!</p>",
      "rawMarkdown": "Can confirm haha!",
      "votes": null
    },
    {
      "id": "3210678",
      "postDate": "05/27/2025 15:01:46",
      "content": "<p>haha yes, we can add .abs() over pearson</p>",
      "rawMarkdown": "haha yes, we can add .abs() over pearson",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3209608,
      "author_name": "sumit08",
      "author_url": "",
      "post_date": "05/26/2025 04:52:53",
      "content": "<p>Your code is designed to identify highly correlated feature pairs (Pearson correlation coefficient ≥ 0.98) within a DataFrame.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3209622,
          "author_name": "henrysun",
          "author_url": "",
          "post_date": "05/26/2025 05:15:48",
          "content": "<p>exactly, for training datasets</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3209631,
      "author_name": "henrysun",
      "author_url": "",
      "post_date": "05/26/2025 05:36:10",
      "content": "<p>But, sadly, looks like rm high corr Xs makes the model performance worse 🤣</p>",
      "votes": null,
      "replies": [
        {
          "id": 3210369,
          "author_name": "jclm43",
          "author_url": "",
          "post_date": "05/27/2025 05:02:12",
          "content": "<p>Can confirm haha!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3210026,
      "author_name": "ayushkhaire",
      "author_url": "",
      "post_date": "05/26/2025 16:16:02",
      "content": "<p><a href=\"https://www.kaggle.com/henrysun\" target=\"_blank\">@henrysun</a>  hey ! 0.99 , too strict I guess ! also , have you considering negative correlation ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3210678,
          "author_name": "henrysun",
          "author_url": "",
          "post_date": "05/27/2025 15:01:46",
          "content": "<p>haha yes, we can add .abs() over pearson</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3209601": "```\n# high correlation columns\nextreme_high_corr_features = dict()\nextreme_high_corr_features_list = []\n\nX_columns = [c for c in train_df.columns if c not in MD_COLUMNS]\n\nfor i, cf in tqdm(enumerate(X_columns), total=len(X_columns)):\n    if cf in extreme_high_corr_features_list:\n        continue\n    high_corr_wth_cf = []\n    for cf2 in X_columns[i+1:]:\n        corr_XX = pearsonr(train_df[cf], train_df[cf2]).statistic\n        if corr_XX >= 0.98:\n            high_corr_wth_cf.append(cf2)\n            extreme_high_corr_features_list.append(cf2)\n    if len(high_corr_wth_cf) > 0:\n        extreme_high_corr_features[cf] = high_corr_wth_cf\n        print(f\"i={i} ==> {extreme_high_corr_features}\")\n```\n\nFor your convenience:\n\n```\nextreme_high_corr_features =  {'X1': ['X8'], 'X6': ['X7'], 'X14': ['X15'], 'X33': ['X35'], 'X34': ['X36'], 'X39': ['X41', 'X43'], 'X40': ['X42', 'X44'], 'X45': ['X47', 'X49'], 'X46': ['X48', 'X50'], 'X51': ['X52', 'X53', 'X54', 'X55'], 'X62': ['X104', 'X146'], 'X68': ['X110', 'X152'], 'X74': ['X116', 'X158'], 'X80': ['X122', 'X164'], 'X86': ['X128', 'X170'], 'X92': ['X134', 'X176'], 'X98': ['X140', 'X182'], 'X185': ['X188'], 'X234': ['X241'], 'X235': ['X242'], 'X236': ['X243'], 'X237': ['X244'], 'X238': ['X245'], 'X239': ['X246'], 'X240': ['X247'], 'X253': ['X254'], 'X262': ['X263'], 'X280': ['X282'], 'X281': ['X283'], 'X286': ['X288', 'X290'], 'X287': ['X289', 'X291'], 'X292': ['X294', 'X296'], 'X293': ['X295', 'X297'], 'X298': ['X299', 'X300', 'X302'], 'X301': ['X303'], 'X309': ['X351', 'X393'], 'X315': ['X357', 'X399'], 'X321': ['X363', 'X405'], 'X327': ['X369', 'X411'], 'X333': ['X375', 'X417'], 'X339': ['X381', 'X423'], 'X345': ['X387', 'X429'], 'X431': ['X434'], 'X432': ['X435'], 'X481': ['X488'], 'X482': ['X489'], 'X483': ['X490'], 'X484': ['X491'], 'X485': ['X492'], 'X486': ['X493'], 'X487': ['X494'], 'X613': ['X619'], 'X616': ['X622'], 'X625': ['X631'], 'X628': ['X634'], 'X637': ['X643'], 'X640': ['X646'], 'X649': ['X655'], 'X652': ['X658'], 'X661': ['X667'], 'X663': ['X669'], 'X664': ['X670'], 'X666': ['X672'], 'X673': ['X679'], 'X675': ['X681'], 'X676': ['X682'], 'X678': ['X684'], 'X685': ['X691'], 'X688': ['X694'], 'X690': ['X696'], 'X788': ['X789'], 'X792': ['X793'], 'X796': ['X797'], 'X800': ['X801'], 'X804': ['X805'], 'X816': ['X817'], 'X820': ['X821'], 'X824': ['X825'], 'X884': ['X885'], 'X886': ['X887']}\n```",
    "3209608": "Your code is designed to identify highly correlated feature pairs (Pearson correlation coefficient ≥ 0.98) within a DataFrame.",
    "3209622": "exactly, for training datasets",
    "3209631": "But, sadly, looks like rm high corr Xs makes the model performance worse 🤣",
    "3210026": "henrysun  hey ! 0.99 , too strict I guess ! also , have you considering negative correlation ?",
    "3210369": "Can confirm haha!",
    "3210678": "haha yes, we can add .abs() over pearson"
  },
  "source": "meta"
}