{
  "id": 328885,
  "title": "Predictive Features",
  "url": "/competitions/amex-default-prediction/discussion/328885",
  "author_name": "1110Ra",
  "post_date": "2022-06-03T12:54:27.443000",
  "votes": 17,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I would like to share my findings on predictive features for this competition. In this <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe\" target=\"_blank\">notebook</a>, I studied the association and correlation between categorical features and numerical features within their own groups and also with the target variable. The result may be used to create a preference for predictive features or remove weak features to reduce computation. The <a href=\"https://en.wikipedia.org/wiki/Uncertainty_coefficient\" target=\"_blank\">Uncertainty coefficient</a> and <a href=\"https://en.wikipedia.org/wiki/Pearson_correlation_coefficient\" target=\"_blank\">Pearson correlation</a> is deployed using Dython library to  measure association and correlation in the train dataset.<br>\nThese rankings are aligned with Mr. Durgaprasad's <a href=\"https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model#Feature-Selection\" target=\"_blank\">work </a> on information value (IV) and may be used as an additional criteria to select predictive features.  Feel free to check the complete list on <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe/data\" target=\"_blank\">here</a>. </p>\n<p>Highest correlation with target for categorical features:</p>\n<table>\n<thead>\n<tr>\n<th>CAT_Feature</th>\n<th>Correlation with Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B_38</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>B_30</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>D_120</td>\n<td>0.24</td>\n</tr>\n<tr>\n<td>D_64</td>\n<td>0.20</td>\n</tr>\n<tr>\n<td>D_68</td>\n<td>0.19</td>\n</tr>\n</tbody>\n</table>\n<p>Highest correlation with target for numerical features:</p>\n<table>\n<thead>\n<tr>\n<th>NUM_Feature</th>\n<th>Correlation with Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>P_2</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>D_48</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>B_18</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>B_9</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>D_55</td>\n<td>0.53</td>\n</tr>\n</tbody>\n</table>\n<p>Sincerely,<br>\n111oRa</p>",
  "messages": [
    {
      "id": 1810317,
      "postDate": "2022-06-03T12:54:27.443Z",
      "content": "<p>Hi,</p>\n<p>I would like to share my findings on predictive features for this competition. In this <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe\" target=\"_blank\">notebook</a>, I studied the association and correlation between categorical features and numerical features within their own groups and also with the target variable. The result may be used to create a preference for predictive features or remove weak features to reduce computation. The <a href=\"https://en.wikipedia.org/wiki/Uncertainty_coefficient\" target=\"_blank\">Uncertainty coefficient</a> and <a href=\"https://en.wikipedia.org/wiki/Pearson_correlation_coefficient\" target=\"_blank\">Pearson correlation</a> is deployed using Dython library to  measure association and correlation in the train dataset.<br>\nThese rankings are aligned with Mr. Durgaprasad's <a href=\"https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model#Feature-Selection\" target=\"_blank\">work </a> on information value (IV) and may be used as an additional criteria to select predictive features.  Feel free to check the complete list on <a href=\"https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe/data\" target=\"_blank\">here</a>. </p>\n<p>Highest correlation with target for categorical features:</p>\n<table>\n<thead>\n<tr>\n<th>CAT_Feature</th>\n<th>Correlation with Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B_38</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>B_30</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>D_120</td>\n<td>0.24</td>\n</tr>\n<tr>\n<td>D_64</td>\n<td>0.20</td>\n</tr>\n<tr>\n<td>D_68</td>\n<td>0.19</td>\n</tr>\n</tbody>\n</table>\n<p>Highest correlation with target for numerical features:</p>\n<table>\n<thead>\n<tr>\n<th>NUM_Feature</th>\n<th>Correlation with Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>P_2</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>D_48</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>B_18</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>B_9</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>D_55</td>\n<td>0.53</td>\n</tr>\n</tbody>\n</table>\n<p>Sincerely,<br>\n111oRa</p>",
      "rawMarkdown": "Hi,\n\nI would like to share my findings on predictive features for this competition. In this [notebook](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe), I studied the association and correlation between categorical features and numerical features within their own groups and also with the target variable. The result may be used to create a preference for predictive features or remove weak features to reduce computation. The [Uncertainty coefficient](https://en.wikipedia.org/wiki/Uncertainty_coefficient) and [Pearson correlation](https://en.wikipedia.org/wiki/Pearson_correlation_coefficient) is deployed using Dython library to  measure association and correlation in the train dataset.\nThese rankings are aligned with Mr. Durgaprasad's [work ](https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model#Feature-Selection) on information value (IV) and may be used as an additional criteria to select predictive features.  Feel free to check the complete list on [here](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe/data). \n\n\nHighest correlation with target for categorical features:\n\n| CAT_Feature| Correlation with Target\n| ---| ---|\n| B_38  |  0.53|\n| B_30 |  0.40|\n| D_120   |  0.24|\n| D_64  |  0.20|\n| D_68   |  0.19|\n\nHighest correlation with target for numerical features:\n\n| NUM_Feature| Correlation with Target\n| ---| ---|\n| P_2 |  0.67|\n| D_48 |  0.60|\n| B_18    |  0.55|\n| B_9 |  0.54|\n| D_55    |  0.53|\n\nSincerely,\n111oRa",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1810317": "Hi,\n\nI would like to share my findings on predictive features for this competition. In this [notebook](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe), I studied the association and correlation between categorical features and numerical features within their own groups and also with the target variable. The result may be used to create a preference for predictive features or remove weak features to reduce computation. The [Uncertainty coefficient](https://en.wikipedia.org/wiki/Uncertainty_coefficient) and [Pearson correlation](https://en.wikipedia.org/wiki/Pearson_correlation_coefficient) is deployed using Dython library to  measure association and correlation in the train dataset.\nThese rankings are aligned with Mr. Durgaprasad's [work ](https://www.kaggle.com/code/gopidurgaprasad/amex-credit-score-model#Feature-Selection) on information value (IV) and may be used as an additional criteria to select predictive features.  Feel free to check the complete list on [here](https://www.kaggle.com/code/mohammadrahmati/associations-and-correlations-useful-for-fe/data). \n\n\nHighest correlation with target for categorical features:\n\n| CAT_Feature| Correlation with Target\n| ---| ---|\n| B_38  |  0.53|\n| B_30 |  0.40|\n| D_120   |  0.24|\n| D_64  |  0.20|\n| D_68   |  0.19|\n\nHighest correlation with target for numerical features:\n\n| NUM_Feature| Correlation with Target\n| ---| ---|\n| P_2 |  0.67|\n| D_48 |  0.60|\n| B_18    |  0.55|\n| B_9 |  0.54|\n| D_55    |  0.53|\n\nSincerely,\n111oRa"
  }
}