{
  "id": 330931,
  "title": "How To Select Features?",
  "url": "/competitions/amex-default-prediction/discussion/330931",
  "author_name": "zakopuro",
  "post_date": "2022-06-15T01:27:15.849000",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I found differences between Train and Test data in some features.</p>\n<h3>Categorical Features</h3>\n<p>As shown in this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327161#1801568\" target=\"_blank\">discussion</a>, some of the Categorical Features contain values in the Train data that are not present in the Test data.</p>\n<h3>Adversarial Validation</h3>\n<p>I have created two Notebooks.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/zakopur0/adversarial-validation-train-vs-test\" target=\"_blank\">Adversarial Validation [Train vs Test]</a></li>\n<li><a href=\"https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public\" target=\"_blank\">Adversarial Validation [Private vs Public]</a></li>\n</ul>\n<p>As these results show, the Train and Test data contain several features with different distributions.<br>\nThere are also different features with different distributions in Public and Private.</p>\n<p>How would you handle these features?<br>\nIs it no problem because there is a correlation between CV and LB?</p>",
  "messages": [
    {
      "id": 1820771,
      "postDate": "2022-06-15T01:27:15.850Z",
      "content": "<p>I found differences between Train and Test data in some features.</p>\n<h3>Categorical Features</h3>\n<p>As shown in this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327161#1801568\" target=\"_blank\">discussion</a>, some of the Categorical Features contain values in the Train data that are not present in the Test data.</p>\n<h3>Adversarial Validation</h3>\n<p>I have created two Notebooks.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/zakopur0/adversarial-validation-train-vs-test\" target=\"_blank\">Adversarial Validation [Train vs Test]</a></li>\n<li><a href=\"https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public\" target=\"_blank\">Adversarial Validation [Private vs Public]</a></li>\n</ul>\n<p>As these results show, the Train and Test data contain several features with different distributions.<br>\nThere are also different features with different distributions in Public and Private.</p>\n<p>How would you handle these features?<br>\nIs it no problem because there is a correlation between CV and LB?</p>",
      "rawMarkdown": "I found differences between Train and Test data in some features.\n\n### Categorical Features\nAs shown in this [discussion](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327161#1801568), some of the Categorical Features contain values in the Train data that are not present in the Test data.\n\n### Adversarial Validation\nI have created two Notebooks.\n- [Adversarial Validation [Train vs Test]](https://www.kaggle.com/code/zakopur0/adversarial-validation-train-vs-test)\n- [Adversarial Validation [Private vs Public]](https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public)\n\nAs these results show, the Train and Test data contain several features with different distributions.\nThere are also different features with different distributions in Public and Private.\n\n\nHow would you handle these features?\nIs it no problem because there is a correlation between CV and LB?",
      "votes": 29
    },
    {
      "id": 1820826,
      "postDate": "2022-06-15T03:20:40.633Z",
      "content": "<p>olivier's  null importance is all you need</p>",
      "rawMarkdown": "olivier's  null importance is all you need",
      "votes": 8,
      "replies": [
        {
          "id": 1820874,
          "postDate": "2022-06-15T04:31:58.153Z",
          "content": "<p>Thank you for comment!<br>\nI will check this Notebook :)<br>\n<a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances</a></p>",
          "rawMarkdown": "Thank you for comment!\nI will check this Notebook :)\nhttps://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances",
          "votes": 3
        }
      ]
    },
    {
      "id": 1821762,
      "postDate": "2022-06-15T20:30:34.197Z",
      "content": "<h5>Wait! before you remove the columns, try this</h5>\n<p>Good catch, but you might not want to remove all columns with unique categories only presented in the test. <br>\nThere are other solutions to consider!</p>\n<p><strong>Removing the column</strong></p>\n<ul>\n<li>When you remove the column, the model doesn't learn anything about the missing category in the train data. Since the model doesn't have any input about this category, its predictions will be equal to the global mean/median/mode. This is the most basic example of regularization. </li>\n</ul>\n<p><strong>Removing only the categories that are problematic</strong></p>\n<ul>\n<li>Maybe you have a bunch of categories that you don't know what to do with. What do you do then?</li>\n<li>What you should do is keep the categories that are present in both the train and test data, and remove the ones that are problematic. </li>\n</ul>\n<pre><code>only_in_train = set(train[col].unique()) - set(test[col].unique())\nonly_in_test = set(test[col].unique()) - set(train[col].unique())\n\n# set those values to np.nan\ntrain.loc[train[col].isin(only_in_train), col] = np.nan\ntest.loc[test[col].isin(only_in_test), col] = np.nan\n</code></pre>\n<p><strong>Frequency Encoding</strong></p>\n<ul>\n<li>Frequency encoding is the same as label encoding, but instead of each category getting its own integer, you use the count or frequency of each category as its representation. </li>\n<li>The values for each category are determined by the train data, but you can use the same calculation for the unknown category the test data as well.</li>\n<li>This will remove the problems with missing categories from the test data. If a new category appears in the test data, it will be assigned a frequency value rather than the category itself. This way the model will simply see that \"this is some rare category [never seen before]\" and if there were some <strong>other</strong> different rare values in the train the model will know \"what to do with rate categories\" even it had never seed the new category.</li>\n<li>be careful doing this if you have high cardinality. This means that you have a lot of unique values in your categorical column. In this case, you will only have a few examples for each category, and this will not be a very good form of representation.</li>\n</ul>\n<pre><code>def frequency_encoding(train, test, col):\n    train_dict = dict(train[col].value_counts())\n    test_dict = dict(test[col].value_counts())\n    train[col+'_count'] = train[col].map(train_dict)\n    test[col+'_count'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = frequency_encoding(train, test, 'category')\n</code></pre>\n<p><strong>Target Encoding</strong></p>\n<ul>\n<li>You can build on the same concept and take it one step further. </li>\n<li>Instead of representing each category with the count or frequency of the category, you can represent each category with the target variable's mean value. </li>\n<li>This way, if there are several categories that are present in the train data, but missing in the test data, the model will know that the rare category that is only in the test data belongs to some other, similar category that is in the train data, and will make a prediction using the target variable's mean value for the other category.</li>\n<li>This is also great for encoding categorical variables that have too many categories, because instead of each category getting its own value, each category will get the value of the target mean for this category.</li>\n</ul>\n<pre><code>def target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n</code></pre>\n<p><strong>Target Encoding with Smoothing</strong></p>\n<ul>\n<li>However, this solution is not without its caveats. </li>\n<li>When the target mean is calculated it is susceptible to noise, and if you get a very large target mean or very small target mean, the model will <strong>overfit</strong> to this value and make erroneous predictions on the test data.</li>\n<li>The best way to fix this is to use a mean of the previous values and the new value, with the new value getting more weight. You can use different weights depending on how much data you have, and depending on what you think the target mean should be getting more weight.</li>\n<li>This is also called \"smoothing\" and you can use different formulas to do it. A good one is to add the value alpha to both the numerator and denominator, and then divide. </li>\n<li>How much you should add to the numerator and denominator depends on how much data you have. Adding the same amount to both will make your new mean closer to the new value, and you should do this if you think your new values are better at representing the real target mean. </li>\n<li>However, if you think your old values are better at representing the target mean, you should add more to the denominator and keep the numerator at the same value, or add very little to the numerator. </li>\n<li>To keep the mean at the same place you should add alpha to the numerator and (alpha * number of observations) to the denominator.</li>\n</ul>\n<pre><code>def target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n</code></pre>\n<p><strong>Target Encoding with K-Fold</strong></p>\n<ul>\n<li>Instead of encoding the category with the average value of the target, you can do target encoding using K-Fold to help the stability of the calculated value. </li>\n<li>This will reduce the impact of noise, but will not fully remove it. </li>\n<li>If the dataset is big enough, this is usually not a problem, and this method is a great way to do target encoding with categories that have too many unique values.</li>\n</ul>\n<pre><code>from sklearn.model_selection import KFold\n\ndef target_encoding_kfold(train, test, col, target_col, n_folds=5, alpha=5):\n    train[col+'_target_enc'] = np.nan\n    test[col+'_target_enc'] = np.nan\n    kf = KFold(n_splits=n_folds, shuffle=True, random_state=1)\n    target_mean = train[target_col].mean()\n    for tr_idx, val_idx in kf.split(train):\n        X_tr, X_val = train.iloc[tr_idx], train.iloc[val_idx]\n        train.loc[train.index[val_idx], col+'_target_enc'] = X_val[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        test[col+'_target_enc'] = test[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        train[col+'_target_enc'].fillna(target_mean, inplace=True)\n        test[col+'_target_enc'].fillna(target_mean, inplace=True)\n        train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n        test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding_kfold(train, test, 'category', 'target')\n</code></pre>",
      "rawMarkdown": "##### Wait! before you remove the columns, try this\n\nGood catch, but you might not want to remove all columns with unique categories only presented in the test. \nThere are other solutions to consider!\n\n\n**Removing the column**\n- When you remove the column, the model doesn't learn anything about the missing category in the train data. Since the model doesn't have any input about this category, its predictions will be equal to the global mean/median/mode. This is the most basic example of regularization. \n\n\n**Removing only the categories that are problematic**\n- Maybe you have a bunch of categories that you don't know what to do with. What do you do then?\n- What you should do is keep the categories that are present in both the train and test data, and remove the ones that are problematic. \n\n```python\nonly_in_train = set(train[col].unique()) - set(test[col].unique())\nonly_in_test = set(test[col].unique()) - set(train[col].unique())\n\n# set those values to np.nan\ntrain.loc[train[col].isin(only_in_train), col] = np.nan\ntest.loc[test[col].isin(only_in_test), col] = np.nan\n```\n\n**Frequency Encoding**\n-  Frequency encoding is the same as label encoding, but instead of each category getting its own integer, you use the count or frequency of each category as its representation. \n-  The values for each category are determined by the train data, but you can use the same calculation for the unknown category the test data as well.\n- This will remove the problems with missing categories from the test data. If a new category appears in the test data, it will be assigned a frequency value rather than the category itself. This way the model will simply see that \"this is some rare category [never seen before]\" and if there were some **other** different rare values in the train the model will know \"what to do with rate categories\" even it had never seed the new category.\n- be careful doing this if you have high cardinality. This means that you have a lot of unique values in your categorical column. In this case, you will only have a few examples for each category, and this will not be a very good form of representation.\n\n```python\ndef frequency_encoding(train, test, col):\n    train_dict = dict(train[col].value_counts())\n    test_dict = dict(test[col].value_counts())\n    train[col+'_count'] = train[col].map(train_dict)\n    test[col+'_count'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = frequency_encoding(train, test, 'category')\n```\n\n**Target Encoding**\n- You can build on the same concept and take it one step further. \n- Instead of representing each category with the count or frequency of the category, you can represent each category with the target variable's mean value. \n- This way, if there are several categories that are present in the train data, but missing in the test data, the model will know that the rare category that is only in the test data belongs to some other, similar category that is in the train data, and will make a prediction using the target variable's mean value for the other category.\n- This is also great for encoding categorical variables that have too many categories, because instead of each category getting its own value, each category will get the value of the target mean for this category.\n\n```python\ndef target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n```\n\n**Target Encoding with Smoothing**\n- However, this solution is not without its caveats. \n- When the target mean is calculated it is susceptible to noise, and if you get a very large target mean or very small target mean, the model will **overfit** to this value and make erroneous predictions on the test data.\n- The best way to fix this is to use a mean of the previous values and the new value, with the new value getting more weight. You can use different weights depending on how much data you have, and depending on what you think the target mean should be getting more weight.\n- This is also called \"smoothing\" and you can use different formulas to do it. A good one is to add the value alpha to both the numerator and denominator, and then divide. \n- How much you should add to the numerator and denominator depends on how much data you have. Adding the same amount to both will make your new mean closer to the new value, and you should do this if you think your new values are better at representing the real target mean. \n- However, if you think your old values are better at representing the target mean, you should add more to the denominator and keep the numerator at the same value, or add very little to the numerator. \n- To keep the mean at the same place you should add alpha to the numerator and (alpha * number of observations) to the denominator.\n\n```python\ndef target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n```\n\n**Target Encoding with K-Fold**\n- Instead of encoding the category with the average value of the target, you can do target encoding using K-Fold to help the stability of the calculated value. \n- This will reduce the impact of noise, but will not fully remove it. \n- If the dataset is big enough, this is usually not a problem, and this method is a great way to do target encoding with categories that have too many unique values.\n\n```python\nfrom sklearn.model_selection import KFold\n\ndef target_encoding_kfold(train, test, col, target_col, n_folds=5, alpha=5):\n    train[col+'_target_enc'] = np.nan\n    test[col+'_target_enc'] = np.nan\n    kf = KFold(n_splits=n_folds, shuffle=True, random_state=1)\n    target_mean = train[target_col].mean()\n    for tr_idx, val_idx in kf.split(train):\n        X_tr, X_val = train.iloc[tr_idx], train.iloc[val_idx]\n        train.loc[train.index[val_idx], col+'_target_enc'] = X_val[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        test[col+'_target_enc'] = test[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        train[col+'_target_enc'].fillna(target_mean, inplace=True)\n        test[col+'_target_enc'].fillna(target_mean, inplace=True)\n        train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n        test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding_kfold(train, test, 'category', 'target')\n```\n\n",
      "votes": 5
    },
    {
      "id": 1824071,
      "postDate": "2022-06-18T00:40:48.130Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1824123,
          "postDate": "2022-06-18T02:02:18.060Z",
          "content": "<p><a href=\"https://www.kaggle.com/yiwangnz\" target=\"_blank\">@yiwangnz</a> <br>\nThank you comment!</p>\n<p>We can see how the scores are affected when these features are removed.<br>\nCV - public LB is correlated, but I think analysis is needed to prevent private LB swing.</p>",
          "rawMarkdown": "@yiwangnz \nThank you comment!\n\nWe can see how the scores are affected when these features are removed.\nCV - public LB is correlated, but I think analysis is needed to prevent private LB swing."
        }
      ]
    },
    {
      "id": 1861989,
      "postDate": "2022-07-19T11:28:03.797Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!"
    }
  ],
  "comments": [
    {
      "id": 1820826,
      "author_name": "Young for you",
      "author_url": "",
      "post_date": "2022-06-15T03:20:40.633000",
      "content": "<p>olivier's  null importance is all you need</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1820874,
          "author_name": "zakopuro",
          "author_url": "",
          "post_date": "2022-06-15T04:31:58.153000",
          "content": "<p>Thank you for comment!<br>\nI will check this Notebook :)<br>\n<a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1821762,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-06-15T20:30:34.197000",
      "content": "<h5>Wait! before you remove the columns, try this</h5>\n<p>Good catch, but you might not want to remove all columns with unique categories only presented in the test. <br>\nThere are other solutions to consider!</p>\n<p><strong>Removing the column</strong></p>\n<ul>\n<li>When you remove the column, the model doesn't learn anything about the missing category in the train data. Since the model doesn't have any input about this category, its predictions will be equal to the global mean/median/mode. This is the most basic example of regularization. </li>\n</ul>\n<p><strong>Removing only the categories that are problematic</strong></p>\n<ul>\n<li>Maybe you have a bunch of categories that you don't know what to do with. What do you do then?</li>\n<li>What you should do is keep the categories that are present in both the train and test data, and remove the ones that are problematic. </li>\n</ul>\n<pre><code>only_in_train = set(train[col].unique()) - set(test[col].unique())\nonly_in_test = set(test[col].unique()) - set(train[col].unique())\n\n# set those values to np.nan\ntrain.loc[train[col].isin(only_in_train), col] = np.nan\ntest.loc[test[col].isin(only_in_test), col] = np.nan\n</code></pre>\n<p><strong>Frequency Encoding</strong></p>\n<ul>\n<li>Frequency encoding is the same as label encoding, but instead of each category getting its own integer, you use the count or frequency of each category as its representation. </li>\n<li>The values for each category are determined by the train data, but you can use the same calculation for the unknown category the test data as well.</li>\n<li>This will remove the problems with missing categories from the test data. If a new category appears in the test data, it will be assigned a frequency value rather than the category itself. This way the model will simply see that \"this is some rare category [never seen before]\" and if there were some <strong>other</strong> different rare values in the train the model will know \"what to do with rate categories\" even it had never seed the new category.</li>\n<li>be careful doing this if you have high cardinality. This means that you have a lot of unique values in your categorical column. In this case, you will only have a few examples for each category, and this will not be a very good form of representation.</li>\n</ul>\n<pre><code>def frequency_encoding(train, test, col):\n    train_dict = dict(train[col].value_counts())\n    test_dict = dict(test[col].value_counts())\n    train[col+'_count'] = train[col].map(train_dict)\n    test[col+'_count'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = frequency_encoding(train, test, 'category')\n</code></pre>\n<p><strong>Target Encoding</strong></p>\n<ul>\n<li>You can build on the same concept and take it one step further. </li>\n<li>Instead of representing each category with the count or frequency of the category, you can represent each category with the target variable's mean value. </li>\n<li>This way, if there are several categories that are present in the train data, but missing in the test data, the model will know that the rare category that is only in the test data belongs to some other, similar category that is in the train data, and will make a prediction using the target variable's mean value for the other category.</li>\n<li>This is also great for encoding categorical variables that have too many categories, because instead of each category getting its own value, each category will get the value of the target mean for this category.</li>\n</ul>\n<pre><code>def target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n</code></pre>\n<p><strong>Target Encoding with Smoothing</strong></p>\n<ul>\n<li>However, this solution is not without its caveats. </li>\n<li>When the target mean is calculated it is susceptible to noise, and if you get a very large target mean or very small target mean, the model will <strong>overfit</strong> to this value and make erroneous predictions on the test data.</li>\n<li>The best way to fix this is to use a mean of the previous values and the new value, with the new value getting more weight. You can use different weights depending on how much data you have, and depending on what you think the target mean should be getting more weight.</li>\n<li>This is also called \"smoothing\" and you can use different formulas to do it. A good one is to add the value alpha to both the numerator and denominator, and then divide. </li>\n<li>How much you should add to the numerator and denominator depends on how much data you have. Adding the same amount to both will make your new mean closer to the new value, and you should do this if you think your new values are better at representing the real target mean. </li>\n<li>However, if you think your old values are better at representing the target mean, you should add more to the denominator and keep the numerator at the same value, or add very little to the numerator. </li>\n<li>To keep the mean at the same place you should add alpha to the numerator and (alpha * number of observations) to the denominator.</li>\n</ul>\n<pre><code>def target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n</code></pre>\n<p><strong>Target Encoding with K-Fold</strong></p>\n<ul>\n<li>Instead of encoding the category with the average value of the target, you can do target encoding using K-Fold to help the stability of the calculated value. </li>\n<li>This will reduce the impact of noise, but will not fully remove it. </li>\n<li>If the dataset is big enough, this is usually not a problem, and this method is a great way to do target encoding with categories that have too many unique values.</li>\n</ul>\n<pre><code>from sklearn.model_selection import KFold\n\ndef target_encoding_kfold(train, test, col, target_col, n_folds=5, alpha=5):\n    train[col+'_target_enc'] = np.nan\n    test[col+'_target_enc'] = np.nan\n    kf = KFold(n_splits=n_folds, shuffle=True, random_state=1)\n    target_mean = train[target_col].mean()\n    for tr_idx, val_idx in kf.split(train):\n        X_tr, X_val = train.iloc[tr_idx], train.iloc[val_idx]\n        train.loc[train.index[val_idx], col+'_target_enc'] = X_val[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        test[col+'_target_enc'] = test[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        train[col+'_target_enc'].fillna(target_mean, inplace=True)\n        test[col+'_target_enc'].fillna(target_mean, inplace=True)\n        train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n        test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding_kfold(train, test, 'category', 'target')\n</code></pre>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1824071,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-18T00:40:48.130000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1824123,
          "author_name": "zakopuro",
          "author_url": "",
          "post_date": "2022-06-18T02:02:18.060000",
          "content": "<p><a href=\"https://www.kaggle.com/yiwangnz\" target=\"_blank\">@yiwangnz</a> <br>\nThank you comment!</p>\n<p>We can see how the scores are affected when these features are removed.<br>\nCV - public LB is correlated, but I think analysis is needed to prevent private LB swing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1861989,
      "author_name": "Rajneesh Singhatiya",
      "author_url": "",
      "post_date": "2022-07-19T11:28:03.797000",
      "content": "<p>Thanks a lot!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1820771": "I found differences between Train and Test data in some features.\n\n### Categorical Features\nAs shown in this [discussion](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327161#1801568), some of the Categorical Features contain values in the Train data that are not present in the Test data.\n\n### Adversarial Validation\nI have created two Notebooks.\n- [Adversarial Validation [Train vs Test]](https://www.kaggle.com/code/zakopur0/adversarial-validation-train-vs-test)\n- [Adversarial Validation [Private vs Public]](https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public)\n\nAs these results show, the Train and Test data contain several features with different distributions.\nThere are also different features with different distributions in Public and Private.\n\n\nHow would you handle these features?\nIs it no problem because there is a correlation between CV and LB?",
    "1820826": "olivier's  null importance is all you need",
    "1821762": "##### Wait! before you remove the columns, try this\n\nGood catch, but you might not want to remove all columns with unique categories only presented in the test. \nThere are other solutions to consider!\n\n\n**Removing the column**\n- When you remove the column, the model doesn't learn anything about the missing category in the train data. Since the model doesn't have any input about this category, its predictions will be equal to the global mean/median/mode. This is the most basic example of regularization. \n\n\n**Removing only the categories that are problematic**\n- Maybe you have a bunch of categories that you don't know what to do with. What do you do then?\n- What you should do is keep the categories that are present in both the train and test data, and remove the ones that are problematic. \n\n```python\nonly_in_train = set(train[col].unique()) - set(test[col].unique())\nonly_in_test = set(test[col].unique()) - set(train[col].unique())\n\n# set those values to np.nan\ntrain.loc[train[col].isin(only_in_train), col] = np.nan\ntest.loc[test[col].isin(only_in_test), col] = np.nan\n```\n\n**Frequency Encoding**\n-  Frequency encoding is the same as label encoding, but instead of each category getting its own integer, you use the count or frequency of each category as its representation. \n-  The values for each category are determined by the train data, but you can use the same calculation for the unknown category the test data as well.\n- This will remove the problems with missing categories from the test data. If a new category appears in the test data, it will be assigned a frequency value rather than the category itself. This way the model will simply see that \"this is some rare category [never seen before]\" and if there were some **other** different rare values in the train the model will know \"what to do with rate categories\" even it had never seed the new category.\n- be careful doing this if you have high cardinality. This means that you have a lot of unique values in your categorical column. In this case, you will only have a few examples for each category, and this will not be a very good form of representation.\n\n```python\ndef frequency_encoding(train, test, col):\n    train_dict = dict(train[col].value_counts())\n    test_dict = dict(test[col].value_counts())\n    train[col+'_count'] = train[col].map(train_dict)\n    test[col+'_count'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = frequency_encoding(train, test, 'category')\n```\n\n**Target Encoding**\n- You can build on the same concept and take it one step further. \n- Instead of representing each category with the count or frequency of the category, you can represent each category with the target variable's mean value. \n- This way, if there are several categories that are present in the train data, but missing in the test data, the model will know that the rare category that is only in the test data belongs to some other, similar category that is in the train data, and will make a prediction using the target variable's mean value for the other category.\n- This is also great for encoding categorical variables that have too many categories, because instead of each category getting its own value, each category will get the value of the target mean for this category.\n\n```python\ndef target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n```\n\n**Target Encoding with Smoothing**\n- However, this solution is not without its caveats. \n- When the target mean is calculated it is susceptible to noise, and if you get a very large target mean or very small target mean, the model will **overfit** to this value and make erroneous predictions on the test data.\n- The best way to fix this is to use a mean of the previous values and the new value, with the new value getting more weight. You can use different weights depending on how much data you have, and depending on what you think the target mean should be getting more weight.\n- This is also called \"smoothing\" and you can use different formulas to do it. A good one is to add the value alpha to both the numerator and denominator, and then divide. \n- How much you should add to the numerator and denominator depends on how much data you have. Adding the same amount to both will make your new mean closer to the new value, and you should do this if you think your new values are better at representing the real target mean. \n- However, if you think your old values are better at representing the target mean, you should add more to the denominator and keep the numerator at the same value, or add very little to the numerator. \n- To keep the mean at the same place you should add alpha to the numerator and (alpha * number of observations) to the denominator.\n\n```python\ndef target_encoding(train, test, col, target_col, alpha=5):\n    train_dict = dict(train.groupby(col)[target_col].mean())\n    test_dict = dict(test.groupby(col)[target_col].mean())\n    train[col+'_target_enc'] = train[col].map(train_dict)\n    test[col+'_target_enc'] = test[col].map(test_dict)\n    train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding(train, test, 'category', 'target')\n```\n\n**Target Encoding with K-Fold**\n- Instead of encoding the category with the average value of the target, you can do target encoding using K-Fold to help the stability of the calculated value. \n- This will reduce the impact of noise, but will not fully remove it. \n- If the dataset is big enough, this is usually not a problem, and this method is a great way to do target encoding with categories that have too many unique values.\n\n```python\nfrom sklearn.model_selection import KFold\n\ndef target_encoding_kfold(train, test, col, target_col, n_folds=5, alpha=5):\n    train[col+'_target_enc'] = np.nan\n    test[col+'_target_enc'] = np.nan\n    kf = KFold(n_splits=n_folds, shuffle=True, random_state=1)\n    target_mean = train[target_col].mean()\n    for tr_idx, val_idx in kf.split(train):\n        X_tr, X_val = train.iloc[tr_idx], train.iloc[val_idx]\n        train.loc[train.index[val_idx], col+'_target_enc'] = X_val[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        test[col+'_target_enc'] = test[col].map(dict(X_tr.groupby(col)[target_col].mean()))\n        train[col+'_target_enc'].fillna(target_mean, inplace=True)\n        test[col+'_target_enc'].fillna(target_mean, inplace=True)\n        train[col+'_target_enc'] = train[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n        test[col+'_target_enc'] = test[col+'_target_enc'] * alpha + train[target_col].mean() * (1 - alpha)\n    return train, test\n\ntrain, test = target_encoding_kfold(train, test, 'category', 'target')\n```\n\n",
    "1824071": "",
    "1861989": "Thanks a lot!"
  }
}