{
  "id": 338569,
  "title": "Are features correlated and does it matter?",
  "url": "/competitions/amex-default-prediction/discussion/338569",
  "author_name": "",
  "post_date": "2022-07-21T01:44:27.310423Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> posted an interesting <a href=\"https://www.kaggle.com/code/raddar/redundant-features-amex/\" target=\"_blank\"><strong>notebook</strong></a> about correlated features - check it out!</p>\n<p>I don't think this means that any feature can be removed because they are not identical, and for tree-based modeling co-linearity is generally not a problem. There is always a chance that last bit of difference between the two features ends up being important in classifying a subset of data points better. Co-linearity would be a big problem for linear methods, so that's something to keep in mind if going in that direction.</p>\n<p>In my experience there are many more correlated features than just those shown by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. This is a complete heatmap of all feature correlations, though I should preface this by saying that the features are modified and this may not look the same in your hands.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3ad3aa9508fe1745ab2ff4870f04a53a%2Fdf_train_correlations.png?generation=1658367208920732&amp;alt=media\" alt=\"\"></p>\n<p>This image will be too small to view properly in Kaggle webpage, so you should probably right-hand click on it and <code>Save the image as ...</code> to your computer, where you can zoom into it. It is fairly high resolution and should be legible.</p>\n<p>To make it somewhat easier to see, here are only the features for which correlations &gt; 0.8, and even that is a big group.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3b84849b1922cacadf148d49d9601277%2Fdf_train_correlations_subset.png?generation=1658367377457420&amp;alt=media\" alt=\"\"></p>\n<p>You will notice that in general features from the same category (Bs, Ds, Rs and Ss) tend to be more correlated within the group than outside of group.</p>\n<p>I can't yet reveal how I modified the features to get these plots, so you'll have to take my word or simply ignore this post.</p>\n<p>PS In the big heatmap not all feature names are shown, as they would overlap and make it illegible. Their order is the same as in original train file, so hopefully that helps.</p>",
  "messages": [
    {
      "id": "1864361",
      "postDate": "07/21/2022 01:44:27",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> posted an interesting <a href=\"https://www.kaggle.com/code/raddar/redundant-features-amex/\" target=\"_blank\"><strong>notebook</strong></a> about correlated features - check it out!</p>\n<p>I don't think this means that any feature can be removed because they are not identical, and for tree-based modeling co-linearity is generally not a problem. There is always a chance that last bit of difference between the two features ends up being important in classifying a subset of data points better. Co-linearity would be a big problem for linear methods, so that's something to keep in mind if going in that direction.</p>\n<p>In my experience there are many more correlated features than just those shown by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. This is a complete heatmap of all feature correlations, though I should preface this by saying that the features are modified and this may not look the same in your hands.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3ad3aa9508fe1745ab2ff4870f04a53a%2Fdf_train_correlations.png?generation=1658367208920732&amp;alt=media\" alt=\"\"></p>\n<p>This image will be too small to view properly in Kaggle webpage, so you should probably right-hand click on it and <code>Save the image as ...</code> to your computer, where you can zoom into it. It is fairly high resolution and should be legible.</p>\n<p>To make it somewhat easier to see, here are only the features for which correlations &gt; 0.8, and even that is a big group.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3b84849b1922cacadf148d49d9601277%2Fdf_train_correlations_subset.png?generation=1658367377457420&amp;alt=media\" alt=\"\"></p>\n<p>You will notice that in general features from the same category (Bs, Ds, Rs and Ss) tend to be more correlated within the group than outside of group.</p>\n<p>I can't yet reveal how I modified the features to get these plots, so you'll have to take my word or simply ignore this post.</p>\n<p>PS In the big heatmap not all feature names are shown, as they would overlap and make it illegible. Their order is the same as in original train file, so hopefully that helps.</p>",
      "rawMarkdown": "raddar posted an interesting [**notebook**](https://www.kaggle.com/code/raddar/redundant-features-amex/) about correlated features - check it out!\n\nI don't think this means that any feature can be removed because they are not identical, and for tree-based modeling co-linearity is generally not a problem. There is always a chance that last bit of difference between the two features ends up being important in classifying a subset of data points better. Co-linearity would be a big problem for linear methods, so that's something to keep in mind if going in that direction.\n\nIn my experience there are many more correlated features than just those shown by @raddar. This is a complete heatmap of all feature correlations, though I should preface this by saying that the features are modified and this may not look the same in your hands.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3ad3aa9508fe1745ab2ff4870f04a53a%2Fdf_train_correlations.png?generation=1658367208920732&alt=media)\n\nThis image will be too small to view properly in Kaggle webpage, so you should probably right-hand click on it and `Save the image as ...` to your computer, where you can zoom into it. It is fairly high resolution and should be legible.\n\nTo make it somewhat easier to see, here are only the features for which correlations > 0.8, and even that is a big group.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3b84849b1922cacadf148d49d9601277%2Fdf_train_correlations_subset.png?generation=1658367377457420&alt=media)\n\nYou will notice that in general features from the same category (Bs, Ds, Rs and Ss) tend to be more correlated within the group than outside of group.\n\nI can't yet reveal how I modified the features to get these plots, so you'll have to take my word or simply ignore this post.\n\nPS In the big heatmap not all feature names are shown, as they would overlap and make it illegible. Their order is the same as in original train file, so hopefully that helps.",
      "votes": null
    },
    {
      "id": "1865085",
      "postDate": "07/21/2022 14:08:42",
      "content": "<p>If you are using decision trees, you can certainly drop (one of each pair of) the <a href=\"https://www.kaggle.com/code/raddar/redundant-features-amex\" target=\"_blank\">redundant features</a> that <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> found, as their ordering is the same and there's no additional information in the lower cardinality features. These won't necessarily come up as 100% correlated, as that is only a measure of their linear relationship, and not of their conditional entropy. You want to look up <a href=\"https://en.wikipedia.org/wiki/Mutual_information\" target=\"_blank\">mutual information</a>, which is easy to calculate for low cardinality categorical features, much less obvious for sampled continuous features! There's an <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.mutual_info_regression.html#:~:text=Mutual%20information%20(MI)%20%5B1,higher%20values%20mean%20higher%20dependency.\" target=\"_blank\">sklearn implementation</a> for the latter, but I've not used it.</p>",
      "rawMarkdown": "If you are using decision trees, you can certainly drop (one of each pair of) the [redundant features](https://www.kaggle.com/code/raddar/redundant-features-amex) that @raddar found, as their ordering is the same and there's no additional information in the lower cardinality features. These won't necessarily come up as 100% correlated, as that is only a measure of their linear relationship, and not of their conditional entropy. You want to look up [mutual information](https://en.wikipedia.org/wiki/Mutual_information), which is easy to calculate for low cardinality categorical features, much less obvious for sampled continuous features! There's an [sklearn implementation](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.mutual_info_regression.html#:~:text=Mutual%20information%20(MI)%20%5B1,higher%20values%20mean%20higher%20dependency.) for the latter, but I've not used it.",
      "votes": null
    },
    {
      "id": "1865416",
      "postDate": "07/21/2022 19:53:11",
      "content": "<p>I didn't put my statements into proper context. Of course there are many features that could be removed and the effect would be negligible (only on 3rd or 4th decimal place), and likely the training much faster. In a competition setting such as this one, 3rd or 4th decimal places actually make a 1000+ ranking difference. That is why I said that keeping some features is worth it even if it helps rank better only a few hundred out of 9+ million samples.</p>",
      "rawMarkdown": "I didn't put my statements into proper context. Of course there are many features that could be removed and the effect would be negligible (only on 3rd or 4th decimal place), and likely the training much faster. In a competition setting such as this one, 3rd or 4th decimal places actually make a 1000+ ranking difference. That is why I said that keeping some features is worth it even if it helps rank better only a few hundred out of 9+ million samples.",
      "votes": null
    },
    {
      "id": "1868405",
      "postDate": "07/24/2022 00:35:11",
      "content": "<p>Thanks for sharing your thoughts on this. I completely agree that co-linearity is not necessarily a problem for tree-based models, and in fact, I think it can actually be helpful!</p>",
      "rawMarkdown": "Thanks for sharing your thoughts on this. I completely agree that co-linearity is not necessarily a problem for tree-based models, and in fact, I think it can actually be helpful!",
      "votes": null
    },
    {
      "id": "1879532",
      "postDate": "08/01/2022 05:45:18",
      "content": "<p>Hello ! In my notebook <a href=\"https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex\" target=\"_blank\">eda carlos amex</a> I calculated Pearson's (linear) and Spearman's (non-linear) correlation for all features (and made some nice plots of intra-group correlations), but only for the last month of data (due to temporal correlations). This isn't a complete guide for correlation, but, in the different versions of it, I think there's a lot of information about it.</p>\n<p>The notebook isn't well organized though - it's the work of a novice.</p>\n<p>I've made it because I tried to build all linear models I could first (before diving into tree-based ones), but the scores I achieved were miserable, so I abandoned the idea.</p>\n<p>I still don't know , too, what to do with those long lists of correlations. In my notebook I also link some discussions in this forum about correlated features.</p>",
      "rawMarkdown": "Hello ! In my notebook [eda carlos amex] (https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex) I calculated Pearson's (linear) and Spearman's (non-linear) correlation for all features (and made some nice plots of intra-group correlations), but only for the last month of data (due to temporal correlations). This isn't a complete guide for correlation, but, in the different versions of it, I think there's a lot of information about it.\n\nThe notebook isn't well organized though - it's the work of a novice.\n\nI've made it because I tried to build all linear models I could first (before diving into tree-based ones), but the scores I achieved were miserable, so I abandoned the idea.\n\nI still don't know , too, what to do with those long lists of correlations. In my notebook I also link some discussions in this forum about correlated features.",
      "votes": null
    },
    {
      "id": "1881294",
      "postDate": "08/02/2022 12:32:53",
      "content": "<p>Thanks for sharing your experience, as I am running my models on Kaggle, with limited RAM, there are some difficult decisions to take on selecting features and your thoughts are very helpful. </p>",
      "rawMarkdown": "Thanks for sharing your experience, as I am running my models on Kaggle, with limited RAM, there are some difficult decisions to take on selecting features and your thoughts are very helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1865085,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "07/21/2022 14:08:42",
      "content": "<p>If you are using decision trees, you can certainly drop (one of each pair of) the <a href=\"https://www.kaggle.com/code/raddar/redundant-features-amex\" target=\"_blank\">redundant features</a> that <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> found, as their ordering is the same and there's no additional information in the lower cardinality features. These won't necessarily come up as 100% correlated, as that is only a measure of their linear relationship, and not of their conditional entropy. You want to look up <a href=\"https://en.wikipedia.org/wiki/Mutual_information\" target=\"_blank\">mutual information</a>, which is easy to calculate for low cardinality categorical features, much less obvious for sampled continuous features! There's an <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.mutual_info_regression.html#:~:text=Mutual%20information%20(MI)%20%5B1,higher%20values%20mean%20higher%20dependency.\" target=\"_blank\">sklearn implementation</a> for the latter, but I've not used it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1865416,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/21/2022 19:53:11",
          "content": "<p>I didn't put my statements into proper context. Of course there are many features that could be removed and the effect would be negligible (only on 3rd or 4th decimal place), and likely the training much faster. In a competition setting such as this one, 3rd or 4th decimal places actually make a 1000+ ranking difference. That is why I said that keeping some features is worth it even if it helps rank better only a few hundred out of 9+ million samples.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1868405,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/24/2022 00:35:11",
      "content": "<p>Thanks for sharing your thoughts on this. I completely agree that co-linearity is not necessarily a problem for tree-based models, and in fact, I think it can actually be helpful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1879532,
      "author_name": "carlosasdesouza",
      "author_url": "",
      "post_date": "08/01/2022 05:45:18",
      "content": "<p>Hello ! In my notebook <a href=\"https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex\" target=\"_blank\">eda carlos amex</a> I calculated Pearson's (linear) and Spearman's (non-linear) correlation for all features (and made some nice plots of intra-group correlations), but only for the last month of data (due to temporal correlations). This isn't a complete guide for correlation, but, in the different versions of it, I think there's a lot of information about it.</p>\n<p>The notebook isn't well organized though - it's the work of a novice.</p>\n<p>I've made it because I tried to build all linear models I could first (before diving into tree-based ones), but the scores I achieved were miserable, so I abandoned the idea.</p>\n<p>I still don't know , too, what to do with those long lists of correlations. In my notebook I also link some discussions in this forum about correlated features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1881294,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "08/02/2022 12:32:53",
      "content": "<p>Thanks for sharing your experience, as I am running my models on Kaggle, with limited RAM, there are some difficult decisions to take on selecting features and your thoughts are very helpful. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1864361": "raddar posted an interesting [**notebook**](https://www.kaggle.com/code/raddar/redundant-features-amex/) about correlated features - check it out!\n\nI don't think this means that any feature can be removed because they are not identical, and for tree-based modeling co-linearity is generally not a problem. There is always a chance that last bit of difference between the two features ends up being important in classifying a subset of data points better. Co-linearity would be a big problem for linear methods, so that's something to keep in mind if going in that direction.\n\nIn my experience there are many more correlated features than just those shown by @raddar. This is a complete heatmap of all feature correlations, though I should preface this by saying that the features are modified and this may not look the same in your hands.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3ad3aa9508fe1745ab2ff4870f04a53a%2Fdf_train_correlations.png?generation=1658367208920732&alt=media)\n\nThis image will be too small to view properly in Kaggle webpage, so you should probably right-hand click on it and `Save the image as ...` to your computer, where you can zoom into it. It is fairly high resolution and should be legible.\n\nTo make it somewhat easier to see, here are only the features for which correlations > 0.8, and even that is a big group.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F3b84849b1922cacadf148d49d9601277%2Fdf_train_correlations_subset.png?generation=1658367377457420&alt=media)\n\nYou will notice that in general features from the same category (Bs, Ds, Rs and Ss) tend to be more correlated within the group than outside of group.\n\nI can't yet reveal how I modified the features to get these plots, so you'll have to take my word or simply ignore this post.\n\nPS In the big heatmap not all feature names are shown, as they would overlap and make it illegible. Their order is the same as in original train file, so hopefully that helps.",
    "1865085": "If you are using decision trees, you can certainly drop (one of each pair of) the [redundant features](https://www.kaggle.com/code/raddar/redundant-features-amex) that @raddar found, as their ordering is the same and there's no additional information in the lower cardinality features. These won't necessarily come up as 100% correlated, as that is only a measure of their linear relationship, and not of their conditional entropy. You want to look up [mutual information](https://en.wikipedia.org/wiki/Mutual_information), which is easy to calculate for low cardinality categorical features, much less obvious for sampled continuous features! There's an [sklearn implementation](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.mutual_info_regression.html#:~:text=Mutual%20information%20(MI)%20%5B1,higher%20values%20mean%20higher%20dependency.) for the latter, but I've not used it.",
    "1865416": "I didn't put my statements into proper context. Of course there are many features that could be removed and the effect would be negligible (only on 3rd or 4th decimal place), and likely the training much faster. In a competition setting such as this one, 3rd or 4th decimal places actually make a 1000+ ranking difference. That is why I said that keeping some features is worth it even if it helps rank better only a few hundred out of 9+ million samples.",
    "1868405": "Thanks for sharing your thoughts on this. I completely agree that co-linearity is not necessarily a problem for tree-based models, and in fact, I think it can actually be helpful!",
    "1879532": "Hello ! In my notebook [eda carlos amex] (https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex) I calculated Pearson's (linear) and Spearman's (non-linear) correlation for all features (and made some nice plots of intra-group correlations), but only for the last month of data (due to temporal correlations). This isn't a complete guide for correlation, but, in the different versions of it, I think there's a lot of information about it.\n\nThe notebook isn't well organized though - it's the work of a novice.\n\nI've made it because I tried to build all linear models I could first (before diving into tree-based ones), but the scores I achieved were miserable, so I abandoned the idea.\n\nI still don't know , too, what to do with those long lists of correlations. In my notebook I also link some discussions in this forum about correlated features.",
    "1881294": "Thanks for sharing your experience, as I am running my models on Kaggle, with limited RAM, there are some difficult decisions to take on selecting features and your thoughts are very helpful."
  },
  "source": "meta"
}