{
  "id": 332461,
  "title": "Pearson's and Spearman's correlation for high-cardinality features",
  "url": "/competitions/amex-default-prediction/discussion/332461",
  "author_name": "",
  "post_date": "2022-06-21T19:50:55.435334700Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Based on  <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">RADDAR's dataset</a> , I've made a study on linear and nonlinear correlation between high-cardinality (\"continuous\") features in the following <a href=\"https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex\" target=\"_blank\">notebook</a> . In version 10 Pearson's correlation is calculated, in version 11 Spearman's correlation is calculated and in version 14 some regression plots between the most positively correlated features (both linear and nonlinear) are shown. I didn't make a thoroughly feature engineering analysis yet, but I think this may be useful for both deanonymization and Feature Engineering.</p>\n<p>The methodology and the notebooks I used are described in my notebook above.</p>\n<p>In future versions the most negative correlated features will be plotted , together with an analysis of low-cardinality features.</p>",
  "messages": [
    {
      "id": "1828544",
      "postDate": "06/21/2022 19:50:55",
      "content": "<p>Based on  <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">RADDAR's dataset</a> , I've made a study on linear and nonlinear correlation between high-cardinality (\"continuous\") features in the following <a href=\"https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex\" target=\"_blank\">notebook</a> . In version 10 Pearson's correlation is calculated, in version 11 Spearman's correlation is calculated and in version 14 some regression plots between the most positively correlated features (both linear and nonlinear) are shown. I didn't make a thoroughly feature engineering analysis yet, but I think this may be useful for both deanonymization and Feature Engineering.</p>\n<p>The methodology and the notebooks I used are described in my notebook above.</p>\n<p>In future versions the most negative correlated features will be plotted , together with an analysis of low-cardinality features.</p>",
      "rawMarkdown": "Based on  [RADDAR's dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) , I've made a study on linear and nonlinear correlation between high-cardinality (\"continuous\") features in the following [notebook](https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex) . In version 10 Pearson's correlation is calculated, in version 11 Spearman's correlation is calculated and in version 14 some regression plots between the most positively correlated features (both linear and nonlinear) are shown. I didn't make a thoroughly feature engineering analysis yet, but I think this may be useful for both deanonymization and Feature Engineering.\n\nThe methodology and the notebooks I used are described in my notebook above.\n\nIn future versions the most negative correlated features will be plotted , together with an analysis of low-cardinality features.",
      "votes": null
    },
    {
      "id": "1838947",
      "postDate": "07/01/2022 02:48:32",
      "content": "<p>Great job! You may also want to check this <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe\" target=\"_blank\">notebook</a>. I used dython library to handle association and correlation between num-num, num-cat, cat-cat, num-target, and cat-target.</p>",
      "rawMarkdown": "Great job! You may also want to check this [notebook](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe). I used dython library to handle association and correlation between num-num, num-cat, cat-cat, num-target, and cat-target.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1838947,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/01/2022 02:48:32",
      "content": "<p>Great job! You may also want to check this <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe\" target=\"_blank\">notebook</a>. I used dython library to handle association and correlation between num-num, num-cat, cat-cat, num-target, and cat-target.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1828544": "Based on  [RADDAR's dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) , I've made a study on linear and nonlinear correlation between high-cardinality (\"continuous\") features in the following [notebook](https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex) . In version 10 Pearson's correlation is calculated, in version 11 Spearman's correlation is calculated and in version 14 some regression plots between the most positively correlated features (both linear and nonlinear) are shown. I didn't make a thoroughly feature engineering analysis yet, but I think this may be useful for both deanonymization and Feature Engineering.\n\nThe methodology and the notebooks I used are described in my notebook above.\n\nIn future versions the most negative correlated features will be plotted , together with an analysis of low-cardinality features.",
    "1838947": "Great job! You may also want to check this [notebook](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe). I used dython library to handle association and correlation between num-num, num-cat, cat-cat, num-target, and cat-target."
  },
  "source": "meta"
}