{
  "id": 340714,
  "title": "Imputing values on training and testing sets does not work",
  "url": "/competitions/amex-default-prediction/discussion/340714",
  "author_name": "",
  "post_date": "2022-07-30T15:49:13.891083Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I am trying to impute all of the missing values on both training and testing sets. I created a function which drops the columns if the amount of missing values goes above the threshold, or imputes them if the amount is below the threshold. The code runs fine with no errors, but when I tested the sum of null values it says there are still missing values.</p>\n<p>Here's the code showing the amount of missing values before imputing:</p>\n<p><code>null_vals = all_data.isnull().sum(axis=0).compute()</code><br>\n<code>null_vals.sort_values(ascending=False).head()</code></p>\n<p>The result:</p>\n<blockquote>\n  <p>D_87     16880376<br>\n  D_88     16877944<br>\n  D_108    16798500<br>\n  D_110    16747585<br>\n  D_111    16747585<br>\n  dtype: int64</p>\n</blockquote>\n<p>This is my imputation code:</p>\n<p>`def impute_datasets(removal_thresh=.3):<br>\n    threshold = int(len(all_data) * removal_thresh)<br>\n    cols_before = train_ftr.shape[1]<br>\n    dropping_params = {'axis': 1, 'inplace': True}<br>\n    cat_cols_imputed = 0<br>\n    num_cols_imputed = 0</p>\n<pre><code>for col, val in null_vals.items():\n    if val == 0:\n        continue\n    else:\n        if val &gt; threshold:\n            train_ftr.drop(col, **dropping_params)\n            test_ftr.drop(col, **dropping_params)\n        else:\n            if col in cat_cols:\n                train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mode()[0])\n                test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)\n                cat_cols_imputed += 1\n            else:\n                train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mean())\n                test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mean())\n                num_cols_imputed += 1\n\ncols_after = train_ftr.shape[1]\ncols_removed = cols_before - cols_after\ntotal_cols_imputed = cat_cols_imputed + num_cols_imputed\nprint('{} columns removed with {:.0f}% threshold'.format(cols_removed, removal_thresh * 100))\nprint('{} columns imputed with {:.0f}% threshold ({} categorical, {} numerical)'.format(total_cols_imputed, \n                                                                                        removal_thresh * 100, \n                                                                                        cat_cols_imputed, \n                                                                                        num_cols_imputed))\nprint('There are now {} columns for both training an testing datasets'.format(cols_after))\n</code></pre>\n<p>`</p>\n<p>And these are are results of missing values after imputation:</p>\n<blockquote>\n  <p>customer_ID         0<br>\n  S_2                 0<br>\n  P_2             45985<br>\n  D_39                0<br>\n  B_1                 0<br>\n                  …  <br>\n  D_140           40632<br>\n  D_141          101548<br>\n  D_143          101548<br>\n  D_144           40727<br>\n  D_145          101548<br>\n  Length: 158, dtype: int64</p>\n</blockquote>\n<p>I have tried many various things like changing the loops and functions, but no matter what I do, nothing works. Can anybody lend a hand here?</p>",
  "messages": [
    {
      "id": "1877397",
      "postDate": "07/30/2022 15:49:13",
      "content": "<p>I am trying to impute all of the missing values on both training and testing sets. I created a function which drops the columns if the amount of missing values goes above the threshold, or imputes them if the amount is below the threshold. The code runs fine with no errors, but when I tested the sum of null values it says there are still missing values.</p>\n<p>Here's the code showing the amount of missing values before imputing:</p>\n<p><code>null_vals = all_data.isnull().sum(axis=0).compute()</code><br>\n<code>null_vals.sort_values(ascending=False).head()</code></p>\n<p>The result:</p>\n<blockquote>\n  <p>D_87     16880376<br>\n  D_88     16877944<br>\n  D_108    16798500<br>\n  D_110    16747585<br>\n  D_111    16747585<br>\n  dtype: int64</p>\n</blockquote>\n<p>This is my imputation code:</p>\n<p>`def impute_datasets(removal_thresh=.3):<br>\n    threshold = int(len(all_data) * removal_thresh)<br>\n    cols_before = train_ftr.shape[1]<br>\n    dropping_params = {'axis': 1, 'inplace': True}<br>\n    cat_cols_imputed = 0<br>\n    num_cols_imputed = 0</p>\n<pre><code>for col, val in null_vals.items():\n    if val == 0:\n        continue\n    else:\n        if val &gt; threshold:\n            train_ftr.drop(col, **dropping_params)\n            test_ftr.drop(col, **dropping_params)\n        else:\n            if col in cat_cols:\n                train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mode()[0])\n                test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)\n                cat_cols_imputed += 1\n            else:\n                train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mean())\n                test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mean())\n                num_cols_imputed += 1\n\ncols_after = train_ftr.shape[1]\ncols_removed = cols_before - cols_after\ntotal_cols_imputed = cat_cols_imputed + num_cols_imputed\nprint('{} columns removed with {:.0f}% threshold'.format(cols_removed, removal_thresh * 100))\nprint('{} columns imputed with {:.0f}% threshold ({} categorical, {} numerical)'.format(total_cols_imputed, \n                                                                                        removal_thresh * 100, \n                                                                                        cat_cols_imputed, \n                                                                                        num_cols_imputed))\nprint('There are now {} columns for both training an testing datasets'.format(cols_after))\n</code></pre>\n<p>`</p>\n<p>And these are are results of missing values after imputation:</p>\n<blockquote>\n  <p>customer_ID         0<br>\n  S_2                 0<br>\n  P_2             45985<br>\n  D_39                0<br>\n  B_1                 0<br>\n                  …  <br>\n  D_140           40632<br>\n  D_141          101548<br>\n  D_143          101548<br>\n  D_144           40727<br>\n  D_145          101548<br>\n  Length: 158, dtype: int64</p>\n</blockquote>\n<p>I have tried many various things like changing the loops and functions, but no matter what I do, nothing works. Can anybody lend a hand here?</p>",
      "rawMarkdown": "I am trying to impute all of the missing values on both training and testing sets. I created a function which drops the columns if the amount of missing values goes above the threshold, or imputes them if the amount is below the threshold. The code runs fine with no errors, but when I tested the sum of null values it says there are still missing values.\n\nHere's the code showing the amount of missing values before imputing:\n\n`null_vals = all_data.isnull().sum(axis=0).compute()`\n`null_vals.sort_values(ascending=False).head()`\n\nThe result:\n\n> D_87     16880376\nD_88     16877944\nD_108    16798500\nD_110    16747585\nD_111    16747585\ndtype: int64\n\nThis is my imputation code:\n\n`def impute_datasets(removal_thresh=.3):\n    threshold = int(len(all_data) * removal_thresh)\n    cols_before = train_ftr.shape[1]\n    dropping_params = {'axis': 1, 'inplace': True}\n    cat_cols_imputed = 0\n    num_cols_imputed = 0\n    \n    for col, val in null_vals.items():\n        if val == 0:\n            continue\n        else:\n            if val > threshold:\n                train_ftr.drop(col, **dropping_params)\n                test_ftr.drop(col, **dropping_params)\n            else:\n                if col in cat_cols:\n                    train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mode()[0])\n                    test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)\n                    cat_cols_imputed += 1\n                else:\n                    train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mean())\n                    test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mean())\n                    num_cols_imputed += 1\n\n    cols_after = train_ftr.shape[1]\n    cols_removed = cols_before - cols_after\n    total_cols_imputed = cat_cols_imputed + num_cols_imputed\n    print('{} columns removed with {:.0f}% threshold'.format(cols_removed, removal_thresh * 100))\n    print('{} columns imputed with {:.0f}% threshold ({} categorical, {} numerical)'.format(total_cols_imputed, \n                                                                                            removal_thresh * 100, \n                                                                                            cat_cols_imputed, \n                                                                                            num_cols_imputed))\n    print('There are now {} columns for both training an testing datasets'.format(cols_after))\n\n`\n\nAnd these are are results of missing values after imputation:\n> customer_ID         0\nS_2                 0\nP_2             45985\nD_39                0\nB_1                 0\n                ...  \nD_140           40632\nD_141          101548\nD_143          101548\nD_144           40727\nD_145          101548\nLength: 158, dtype: int64\n\nI have tried many various things like changing the loops and functions, but no matter what I do, nothing works. Can anybody lend a hand here?",
      "votes": null
    },
    {
      "id": "1877408",
      "postDate": "07/30/2022 16:03:17",
      "content": "<p>Your forgot add <code>inplace=True</code> here:</p>\n<pre><code>if val &gt; threshold:\n            train_ftr.drop(col, **dropping_params)\n            test_ftr.drop(col, **dropping_params)\n</code></pre>",
      "rawMarkdown": "Your forgot add `inplace=True` here:\n\n```\nif val > threshold:\n            train_ftr.drop(col, **dropping_params)\n            test_ftr.drop(col, **dropping_params)\n```",
      "votes": null
    },
    {
      "id": "1877413",
      "postDate": "07/30/2022 16:05:43",
      "content": "<p>dropping_params already has the inplace parameter. I coded it as <code>dropping_params = {'axis': 1, 'inplace': True}</code>.</p>",
      "rawMarkdown": "dropping_params already has the inplace parameter. I coded it as `dropping_params = {'axis': 1, 'inplace': True}`.",
      "votes": null
    },
    {
      "id": "1877457",
      "postDate": "07/30/2022 16:57:51",
      "content": "<p>When you say 'nothing works', what do you mean? No score improvement? Have you considered they are null because a customer did not report any delinquency related items? Hence it is vital info and should not be dropped?</p>",
      "rawMarkdown": "When you say 'nothing works', what do you mean? No score improvement? Have you considered they are null because a customer did not report any delinquency related items? Hence it is vital info and should not be dropped?",
      "votes": null
    },
    {
      "id": "1877473",
      "postDate": "07/30/2022 17:08:48",
      "content": "<p>Did you know that there are two types of missing values - missing at random and systemic; Imputation only may/or may not work for missing at random type. This competition does not have (or very little) missing at random NA values.</p>\n<p>That's the main reason nothing works ;)</p>",
      "rawMarkdown": "Did you know that there are two types of missing values - missing at random and systemic; Imputation only may/or may not work for missing at random type. This competition does not have (or very little) missing at random NA values.\n\nThat's the main reason nothing works ;)",
      "votes": null
    },
    {
      "id": "1877487",
      "postDate": "07/30/2022 17:29:08",
      "content": "<p>test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True) - this is wrong,  use instead: <br>\ntest_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=False) <br>\nor<br>\ntest_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)</p>",
      "rawMarkdown": "test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True) - this is wrong,  use instead: \ntest_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=False) \nor\ntest_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)",
      "votes": null
    },
    {
      "id": "1877500",
      "postDate": "07/30/2022 17:47:14",
      "content": "<p>Explore more ways of handling missing data - <a href=\"https://www.kaggle.com/discussions/getting-started/339017\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/339017</a></p>",
      "rawMarkdown": "Explore more ways of handling missing data - https://www.kaggle.com/discussions/getting-started/339017",
      "votes": null
    },
    {
      "id": "1877506",
      "postDate": "07/30/2022 17:58:43",
      "content": "<p>That was a mistake, I will fix that</p>",
      "rawMarkdown": "That was a mistake, I will fix that",
      "votes": null
    },
    {
      "id": "1877508",
      "postDate": "07/30/2022 18:00:02",
      "content": "<p>I meant that there are still missing values in the dataset after imputing them. Hence the results I get when I calculate the sum of missing values for each column after imputing them.</p>",
      "rawMarkdown": "I meant that there are still missing values in the dataset after imputing them. Hence the results I get when I calculate the sum of missing values for each column after imputing them.",
      "votes": null
    },
    {
      "id": "1877892",
      "postDate": "07/31/2022 04:21:26",
      "content": "<p>When the missing values are missing because of an underlying algorithms (which is not random): Their actual existence has information. <br>\nSo imputing them removes this information..</p>",
      "rawMarkdown": "When the missing values are missing because of an underlying algorithms (which is not random): Their actual existence has information. \nSo imputing them removes this information..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1877408,
      "author_name": "alexbruskin",
      "author_url": "",
      "post_date": "07/30/2022 16:03:17",
      "content": "<p>Your forgot add <code>inplace=True</code> here:</p>\n<pre><code>if val &gt; threshold:\n            train_ftr.drop(col, **dropping_params)\n            test_ftr.drop(col, **dropping_params)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1877413,
          "author_name": "une510",
          "author_url": "",
          "post_date": "07/30/2022 16:05:43",
          "content": "<p>dropping_params already has the inplace parameter. I coded it as <code>dropping_params = {'axis': 1, 'inplace': True}</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1877457,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "07/30/2022 16:57:51",
      "content": "<p>When you say 'nothing works', what do you mean? No score improvement? Have you considered they are null because a customer did not report any delinquency related items? Hence it is vital info and should not be dropped?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1877508,
          "author_name": "une510",
          "author_url": "",
          "post_date": "07/30/2022 18:00:02",
          "content": "<p>I meant that there are still missing values in the dataset after imputing them. Hence the results I get when I calculate the sum of missing values for each column after imputing them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1877473,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "07/30/2022 17:08:48",
      "content": "<p>Did you know that there are two types of missing values - missing at random and systemic; Imputation only may/or may not work for missing at random type. This competition does not have (or very little) missing at random NA values.</p>\n<p>That's the main reason nothing works ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1877487,
      "author_name": "andrejvetrov",
      "author_url": "",
      "post_date": "07/30/2022 17:29:08",
      "content": "<p>test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True) - this is wrong,  use instead: <br>\ntest_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=False) <br>\nor<br>\ntest_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1877506,
          "author_name": "une510",
          "author_url": "",
          "post_date": "07/30/2022 17:58:43",
          "content": "<p>That was a mistake, I will fix that</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1877500,
      "author_name": "hskhawaja",
      "author_url": "",
      "post_date": "07/30/2022 17:47:14",
      "content": "<p>Explore more ways of handling missing data - <a href=\"https://www.kaggle.com/discussions/getting-started/339017\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/339017</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1877892,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/31/2022 04:21:26",
      "content": "<p>When the missing values are missing because of an underlying algorithms (which is not random): Their actual existence has information. <br>\nSo imputing them removes this information..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1877397": "I am trying to impute all of the missing values on both training and testing sets. I created a function which drops the columns if the amount of missing values goes above the threshold, or imputes them if the amount is below the threshold. The code runs fine with no errors, but when I tested the sum of null values it says there are still missing values.\n\nHere's the code showing the amount of missing values before imputing:\n\n`null_vals = all_data.isnull().sum(axis=0).compute()`\n`null_vals.sort_values(ascending=False).head()`\n\nThe result:\n\n> D_87     16880376\nD_88     16877944\nD_108    16798500\nD_110    16747585\nD_111    16747585\ndtype: int64\n\nThis is my imputation code:\n\n`def impute_datasets(removal_thresh=.3):\n    threshold = int(len(all_data) * removal_thresh)\n    cols_before = train_ftr.shape[1]\n    dropping_params = {'axis': 1, 'inplace': True}\n    cat_cols_imputed = 0\n    num_cols_imputed = 0\n    \n    for col, val in null_vals.items():\n        if val == 0:\n            continue\n        else:\n            if val > threshold:\n                train_ftr.drop(col, **dropping_params)\n                test_ftr.drop(col, **dropping_params)\n            else:\n                if col in cat_cols:\n                    train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mode()[0])\n                    test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)\n                    cat_cols_imputed += 1\n                else:\n                    train_ftr[col] = train_ftr[col].fillna(train_ftr[col].mean())\n                    test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mean())\n                    num_cols_imputed += 1\n\n    cols_after = train_ftr.shape[1]\n    cols_removed = cols_before - cols_after\n    total_cols_imputed = cat_cols_imputed + num_cols_imputed\n    print('{} columns removed with {:.0f}% threshold'.format(cols_removed, removal_thresh * 100))\n    print('{} columns imputed with {:.0f}% threshold ({} categorical, {} numerical)'.format(total_cols_imputed, \n                                                                                            removal_thresh * 100, \n                                                                                            cat_cols_imputed, \n                                                                                            num_cols_imputed))\n    print('There are now {} columns for both training an testing datasets'.format(cols_after))\n\n`\n\nAnd these are are results of missing values after imputation:\n> customer_ID         0\nS_2                 0\nP_2             45985\nD_39                0\nB_1                 0\n                ...  \nD_140           40632\nD_141          101548\nD_143          101548\nD_144           40727\nD_145          101548\nLength: 158, dtype: int64\n\nI have tried many various things like changing the loops and functions, but no matter what I do, nothing works. Can anybody lend a hand here?",
    "1877408": "Your forgot add `inplace=True` here:\n\n```\nif val > threshold:\n            train_ftr.drop(col, **dropping_params)\n            test_ftr.drop(col, **dropping_params)\n```",
    "1877413": "dropping_params already has the inplace parameter. I coded it as `dropping_params = {'axis': 1, 'inplace': True}`.",
    "1877457": "When you say 'nothing works', what do you mean? No score improvement? Have you considered they are null because a customer did not report any delinquency related items? Hence it is vital info and should not be dropped?",
    "1877473": "Did you know that there are two types of missing values - missing at random and systemic; Imputation only may/or may not work for missing at random type. This competition does not have (or very little) missing at random NA values.\n\nThat's the main reason nothing works ;)",
    "1877487": "test_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True) - this is wrong,  use instead: \ntest_ftr[col] = test_ftr[col].fillna(test_ftr[col].mode()[0], inplace=False) \nor\ntest_ftr[col].fillna(test_ftr[col].mode()[0], inplace=True)",
    "1877500": "Explore more ways of handling missing data - https://www.kaggle.com/discussions/getting-started/339017",
    "1877506": "That was a mistake, I will fix that",
    "1877508": "I meant that there are still missing values in the dataset after imputing them. Hence the results I get when I calculate the sum of missing values for each column after imputing them.",
    "1877892": "When the missing values are missing because of an underlying algorithms (which is not random): Their actual existence has information. \nSo imputing them removes this information.."
  },
  "source": "meta"
}