{
  "id": 334603,
  "title": "Correlation KOs",
  "url": "/competitions/amex-default-prediction/discussion/334603",
  "author_name": "",
  "post_date": "2022-07-02T07:23:25.405373500Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So I wrote this simple function to reduce my feature space (for numerical features). It identifies the features to knockout based on high correlation with other variables in the dataset</p>\n<pre><code>def correlation_KO(dataset, threshold):\n     col_corr = set()  # Set of all the names of correlated columns\n     corr_matrix = dataset.corr()\n     for i in range(len(corr_matrix.columns)):\n         for j in range(i):\n             if abs(corr_matrix.iloc[i, j]) &gt;= threshold: # we are interested in absolute coeff value\n                 colname = corr_matrix.columns[i]  # getting the name of column\n                 col_corr.add(colname)\n     return col_corr\n</code></pre>\n<p>But the XGB actually became weaker when I removed the features having a &gt;= 0.95 correlation! </p>\n<p>I guess even if features are highly correlated, it will use them interchangeably and extract what value it can where they differ. </p>\n<p>Is that a fair assessment? Would this approach be more useful in a Neural Network model?</p>\n<p>May be obvious to many out there but would love to get your thoughts!</p>",
  "messages": [
    {
      "id": "1840350",
      "postDate": "07/02/2022 07:23:25",
      "content": "<p>So I wrote this simple function to reduce my feature space (for numerical features). It identifies the features to knockout based on high correlation with other variables in the dataset</p>\n<pre><code>def correlation_KO(dataset, threshold):\n     col_corr = set()  # Set of all the names of correlated columns\n     corr_matrix = dataset.corr()\n     for i in range(len(corr_matrix.columns)):\n         for j in range(i):\n             if abs(corr_matrix.iloc[i, j]) &gt;= threshold: # we are interested in absolute coeff value\n                 colname = corr_matrix.columns[i]  # getting the name of column\n                 col_corr.add(colname)\n     return col_corr\n</code></pre>\n<p>But the XGB actually became weaker when I removed the features having a &gt;= 0.95 correlation! </p>\n<p>I guess even if features are highly correlated, it will use them interchangeably and extract what value it can where they differ. </p>\n<p>Is that a fair assessment? Would this approach be more useful in a Neural Network model?</p>\n<p>May be obvious to many out there but would love to get your thoughts!</p>",
      "rawMarkdown": "So I wrote this simple function to reduce my feature space (for numerical features). It identifies the features to knockout based on high correlation with other variables in the dataset\n\n```\ndef correlation_KO(dataset, threshold):\n     col_corr = set()  # Set of all the names of correlated columns\n     corr_matrix = dataset.corr()\n     for i in range(len(corr_matrix.columns)):\n         for j in range(i):\n             if abs(corr_matrix.iloc[i, j]) >= threshold: # we are interested in absolute coeff value\n                 colname = corr_matrix.columns[i]  # getting the name of column\n                 col_corr.add(colname)\n     return col_corr\n```\n\nBut the XGB actually became weaker when I removed the features having a >= 0.95 correlation! \n\nI guess even if features are highly correlated, it will use them interchangeably and extract what value it can where they differ. \n\nIs that a fair assessment? Would this approach be more useful in a Neural Network model?\n\nMay be obvious to many out there but would love to get your thoughts!",
      "votes": null
    },
    {
      "id": "1840426",
      "postDate": "07/02/2022 08:28:34",
      "content": "<p>With your function you are computing the pearson correlation which is a measurement of the <strong>linear</strong> relationship between two variables. <br>\nThat means that your variables are correlated for following scenarios: X increases Y increases (correlation &gt;0), X increases Y decreases (correlation &lt;0), X increases Y steady (correlation close to 0)</p>\n<p>Due to the linearity, the algorithm wont be able to find different non-linear relations e.g. X increases when Y is close to 0, which is also a correlation.<br>\nOn top of that reducing the dimensions does not necessarily lead to an increase of performance of the model it rather leads to an decrease of training time. Removing features exchanges variance for more bias.</p>\n<p>You should definetly look into different dimensionality reduction techniques which are more complex e.g. denoising autoencoders or feature importance algorithms e.g. permutation feature importance which even suggest features to be removed due to their negative impact on the model, read more about it here: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131</a></p>",
      "rawMarkdown": "With your function you are computing the pearson correlation which is a measurement of the **linear** relationship between two variables. \nThat means that your variables are correlated for following scenarios: X increases Y increases (correlation >0), X increases Y decreases (correlation <0), X increases Y steady (correlation close to 0)\n\nDue to the linearity, the algorithm wont be able to find different non-linear relations e.g. X increases when Y is close to 0, which is also a correlation.\nOn top of that reducing the dimensions does not necessarily lead to an increase of performance of the model it rather leads to an decrease of training time. Removing features exchanges variance for more bias.\n\nYou should definetly look into different dimensionality reduction techniques which are more complex e.g. denoising autoencoders or feature importance algorithms e.g. permutation feature importance which even suggest features to be removed due to their negative impact on the model, read more about it here: https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131",
      "votes": null
    },
    {
      "id": "1840932",
      "postDate": "07/02/2022 16:30:30",
      "content": "<p>Check to make sure your not eliminating both sides of the correlation - I made that mistake - pretty easy to do.</p>\n<p>A, B = correlation = 0.999<br>\nB, A = correlation = 0.999</p>\n<p>I was eliminating  A due to the first row, and B due to the 2nd.   If you eliminate both than you threw out the baby with the dirty water.</p>",
      "rawMarkdown": "Check to make sure your not eliminating both sides of the correlation - I made that mistake - pretty easy to do.\n\nA, B = correlation = 0.999\nB, A = correlation = 0.999\n\nI was eliminating  A due to the first row, and B due to the 2nd.   If you eliminate both than you threw out the baby with the dirty water.",
      "votes": null
    },
    {
      "id": "1847028",
      "postDate": "07/07/2022 15:01:41",
      "content": "<p>You're right it would be bad to eliminate both correlated variables. But I believe this is handled for in the function</p>\n<pre><code>for i in range(len(corr_matrix.columns)):\n         for j in range(i):\n</code></pre>\n<p>Basically you only traverse till the diagonal element of the correlation matrix, so we're looking for correlation values only one-way</p>",
      "rawMarkdown": "You're right it would be bad to eliminate both correlated variables. But I believe this is handled for in the function\n\n```\nfor i in range(len(corr_matrix.columns)):\n         for j in range(i):\n```\n\nBasically you only traverse till the diagonal element of the correlation matrix, so we're looking for correlation values only one-way",
      "votes": null
    },
    {
      "id": "1847415",
      "postDate": "07/07/2022 23:44:08",
      "content": "<p>Next issue<br>\nA,B = 0.999</p>\n<p>You drop A.  But B has only 10 rows of data and 990 rows of nan.  A has 1000 rows and no nan's.  The correlation matrix fits 10 vs 10 - you drop A when you should drop B.   I think in this data set for training there is not a single row that has no nan's - its full of them.</p>",
      "rawMarkdown": "Next issue\nA,B = 0.999\n\nYou drop A.  But B has only 10 rows of data and 990 rows of nan.  A has 1000 rows and no nan's.  The correlation matrix fits 10 vs 10 - you drop A when you should drop B.   I think in this data set for training there is not a single row that has no nan's - its full of them.",
      "votes": null
    },
    {
      "id": "1847729",
      "postDate": "07/08/2022 05:32:53",
      "content": "<p>Interesting! I did not consider how the correlation function would handle nulls. That opens up its own can of worms with how to handle nulls in a nuanced way so as not to introduce noise to the variable's distribution</p>\n<p>So it probably would be best to limit the corr calculation to variables which have &lt;10% of rows null or something like that. To avoid the scenario you have mentioned</p>",
      "rawMarkdown": "Interesting! I did not consider how the correlation function would handle nulls. That opens up its own can of worms with how to handle nulls in a nuanced way so as not to introduce noise to the variable's distribution\n\nSo it probably would be best to limit the corr calculation to variables which have <10% of rows null or something like that. To avoid the scenario you have mentioned",
      "votes": null
    },
    {
      "id": "1851361",
      "postDate": "07/11/2022 07:30:07",
      "content": "<p>checkout this <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe\" target=\"_blank\">notebook</a>. You may want to do the same for categorical features.</p>",
      "rawMarkdown": "checkout this [notebook](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe). You may want to do the same for categorical features.",
      "votes": null
    },
    {
      "id": "1851788",
      "postDate": "07/11/2022 14:40:13",
      "content": "<p>Thanks for sharing this! That is very helpful</p>",
      "rawMarkdown": "Thanks for sharing this! That is very helpful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1840426,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "07/02/2022 08:28:34",
      "content": "<p>With your function you are computing the pearson correlation which is a measurement of the <strong>linear</strong> relationship between two variables. <br>\nThat means that your variables are correlated for following scenarios: X increases Y increases (correlation &gt;0), X increases Y decreases (correlation &lt;0), X increases Y steady (correlation close to 0)</p>\n<p>Due to the linearity, the algorithm wont be able to find different non-linear relations e.g. X increases when Y is close to 0, which is also a correlation.<br>\nOn top of that reducing the dimensions does not necessarily lead to an increase of performance of the model it rather leads to an decrease of training time. Removing features exchanges variance for more bias.</p>\n<p>You should definetly look into different dimensionality reduction techniques which are more complex e.g. denoising autoencoders or feature importance algorithms e.g. permutation feature importance which even suggest features to be removed due to their negative impact on the model, read more about it here: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1840932,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/02/2022 16:30:30",
      "content": "<p>Check to make sure your not eliminating both sides of the correlation - I made that mistake - pretty easy to do.</p>\n<p>A, B = correlation = 0.999<br>\nB, A = correlation = 0.999</p>\n<p>I was eliminating  A due to the first row, and B due to the 2nd.   If you eliminate both than you threw out the baby with the dirty water.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1847028,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/07/2022 15:01:41",
          "content": "<p>You're right it would be bad to eliminate both correlated variables. But I believe this is handled for in the function</p>\n<pre><code>for i in range(len(corr_matrix.columns)):\n         for j in range(i):\n</code></pre>\n<p>Basically you only traverse till the diagonal element of the correlation matrix, so we're looking for correlation values only one-way</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1847415,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/07/2022 23:44:08",
          "content": "<p>Next issue<br>\nA,B = 0.999</p>\n<p>You drop A.  But B has only 10 rows of data and 990 rows of nan.  A has 1000 rows and no nan's.  The correlation matrix fits 10 vs 10 - you drop A when you should drop B.   I think in this data set for training there is not a single row that has no nan's - its full of them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1847729,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/08/2022 05:32:53",
          "content": "<p>Interesting! I did not consider how the correlation function would handle nulls. That opens up its own can of worms with how to handle nulls in a nuanced way so as not to introduce noise to the variable's distribution</p>\n<p>So it probably would be best to limit the corr calculation to variables which have &lt;10% of rows null or something like that. To avoid the scenario you have mentioned</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1851361,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/11/2022 07:30:07",
      "content": "<p>checkout this <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe\" target=\"_blank\">notebook</a>. You may want to do the same for categorical features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1851788,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/11/2022 14:40:13",
          "content": "<p>Thanks for sharing this! That is very helpful</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1840350": "So I wrote this simple function to reduce my feature space (for numerical features). It identifies the features to knockout based on high correlation with other variables in the dataset\n\n```\ndef correlation_KO(dataset, threshold):\n     col_corr = set()  # Set of all the names of correlated columns\n     corr_matrix = dataset.corr()\n     for i in range(len(corr_matrix.columns)):\n         for j in range(i):\n             if abs(corr_matrix.iloc[i, j]) >= threshold: # we are interested in absolute coeff value\n                 colname = corr_matrix.columns[i]  # getting the name of column\n                 col_corr.add(colname)\n     return col_corr\n```\n\nBut the XGB actually became weaker when I removed the features having a >= 0.95 correlation! \n\nI guess even if features are highly correlated, it will use them interchangeably and extract what value it can where they differ. \n\nIs that a fair assessment? Would this approach be more useful in a Neural Network model?\n\nMay be obvious to many out there but would love to get your thoughts!",
    "1840426": "With your function you are computing the pearson correlation which is a measurement of the **linear** relationship between two variables. \nThat means that your variables are correlated for following scenarios: X increases Y increases (correlation >0), X increases Y decreases (correlation <0), X increases Y steady (correlation close to 0)\n\nDue to the linearity, the algorithm wont be able to find different non-linear relations e.g. X increases when Y is close to 0, which is also a correlation.\nOn top of that reducing the dimensions does not necessarily lead to an increase of performance of the model it rather leads to an decrease of training time. Removing features exchanges variance for more bias.\n\nYou should definetly look into different dimensionality reduction techniques which are more complex e.g. denoising autoencoders or feature importance algorithms e.g. permutation feature importance which even suggest features to be removed due to their negative impact on the model, read more about it here: https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131",
    "1840932": "Check to make sure your not eliminating both sides of the correlation - I made that mistake - pretty easy to do.\n\nA, B = correlation = 0.999\nB, A = correlation = 0.999\n\nI was eliminating  A due to the first row, and B due to the 2nd.   If you eliminate both than you threw out the baby with the dirty water.",
    "1847028": "You're right it would be bad to eliminate both correlated variables. But I believe this is handled for in the function\n\n```\nfor i in range(len(corr_matrix.columns)):\n         for j in range(i):\n```\n\nBasically you only traverse till the diagonal element of the correlation matrix, so we're looking for correlation values only one-way",
    "1847415": "Next issue\nA,B = 0.999\n\nYou drop A.  But B has only 10 rows of data and 990 rows of nan.  A has 1000 rows and no nan's.  The correlation matrix fits 10 vs 10 - you drop A when you should drop B.   I think in this data set for training there is not a single row that has no nan's - its full of them.",
    "1847729": "Interesting! I did not consider how the correlation function would handle nulls. That opens up its own can of worms with how to handle nulls in a nuanced way so as not to introduce noise to the variable's distribution\n\nSo it probably would be best to limit the corr calculation to variables which have <10% of rows null or something like that. To avoid the scenario you have mentioned",
    "1851361": "checkout this [notebook](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe). You may want to do the same for categorical features.",
    "1851788": "Thanks for sharing this! That is very helpful"
  },
  "source": "meta"
}