{
  "id": 342967,
  "title": "Correlation analysis matters.",
  "url": "/competitions/amex-default-prediction/discussion/342967",
  "author_name": "Cccccccc",
  "post_date": "2022-08-09T13:24:25.443000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Maybe a lot of participants like me are confused by large amounts of features when using some other extensive features created by such as lag or stacking etc.</p>\n<p>As far as I am concerned,large amounts of features will face at least two aspects of problems:<br>\n(1) Long running time and Large ram consuming.<br>\n(2) Features redundancy and Local models overfitting.</p>\n<p>I am pretty sure that this kind of problems are related to many domains,but correlation analysis truly works.<br>\nVerified by various ways,it is glad for me to find that when reducing features from about 2500 to 2000 by discarding some features which share high similarity with other features,the score promotes in lb while cv score is almost the same.</p>\n<p>EDA:<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/girishkumarsahu/american-express-default-prediction-eda</a></p>\n<p>There are many terrific EDA on Kaggle community,they guide me how to make a comprehensive data analysis.Let's focus on the correlation analysis.Deep color remind there are some high correlation features.Please recording them.After collecting the candidate features,we can set a threshold to control the partition of discarding features.</p>\n<p>PS:Do this work before extending original features.</p>",
  "messages": [
    {
      "id": 1891525,
      "postDate": "2022-08-09T13:24:25.443Z",
      "content": "<p>Maybe a lot of participants like me are confused by large amounts of features when using some other extensive features created by such as lag or stacking etc.</p>\n<p>As far as I am concerned,large amounts of features will face at least two aspects of problems:<br>\n(1) Long running time and Large ram consuming.<br>\n(2) Features redundancy and Local models overfitting.</p>\n<p>I am pretty sure that this kind of problems are related to many domains,but correlation analysis truly works.<br>\nVerified by various ways,it is glad for me to find that when reducing features from about 2500 to 2000 by discarding some features which share high similarity with other features,the score promotes in lb while cv score is almost the same.</p>\n<p>EDA:<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/girishkumarsahu/american-express-default-prediction-eda</a></p>\n<p>There are many terrific EDA on Kaggle community,they guide me how to make a comprehensive data analysis.Let's focus on the correlation analysis.Deep color remind there are some high correlation features.Please recording them.After collecting the candidate features,we can set a threshold to control the partition of discarding features.</p>\n<p>PS:Do this work before extending original features.</p>",
      "rawMarkdown": "Maybe a lot of participants like me are confused by large amounts of features when using some other extensive features created by such as lag or stacking etc.\n\nAs far as I am concerned,large amounts of features will face at least two aspects of problems:\n(1) Long running time and Large ram consuming.\n(2) Features redundancy and Local models overfitting.\n\nI am pretty sure that this kind of problems are related to many domains,but correlation analysis truly works.\nVerified by various ways,it is glad for me to find that when reducing features from about 2500 to 2000 by discarding some features which share high similarity with other features,the score promotes in lb while cv score is almost the same.\n\nEDA:[https://www.kaggle.com/code/girishkumarsahu/american-express-default-prediction-eda](url)\n\nThere are many terrific EDA on Kaggle community,they guide me how to make a comprehensive data analysis.Let's focus on the correlation analysis.Deep color remind there are some high correlation features.Please recording them.After collecting the candidate features,we can set a threshold to control the partition of discarding features.\n\nPS:Do this work before extending original features.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1891525": "Maybe a lot of participants like me are confused by large amounts of features when using some other extensive features created by such as lag or stacking etc.\n\nAs far as I am concerned,large amounts of features will face at least two aspects of problems:\n(1) Long running time and Large ram consuming.\n(2) Features redundancy and Local models overfitting.\n\nI am pretty sure that this kind of problems are related to many domains,but correlation analysis truly works.\nVerified by various ways,it is glad for me to find that when reducing features from about 2500 to 2000 by discarding some features which share high similarity with other features,the score promotes in lb while cv score is almost the same.\n\nEDA:[https://www.kaggle.com/code/girishkumarsahu/american-express-default-prediction-eda](url)\n\nThere are many terrific EDA on Kaggle community,they guide me how to make a comprehensive data analysis.Let's focus on the correlation analysis.Deep color remind there are some high correlation features.Please recording them.After collecting the candidate features,we can set a threshold to control the partition of discarding features.\n\nPS:Do this work before extending original features."
  }
}