{
  "id": 331007,
  "title": "AMEX : Missing Value or Outlier Treatment ? ",
  "url": "/competitions/amex-default-prediction/discussion/331007",
  "author_name": "",
  "post_date": "2022-06-15T10:10:44.794789800Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>How to go about <strong>missing value treatment</strong> in the AMEX data or outlier treatment?<br>\nsome pointers:</p>\n<ul>\n<li>There are **122 **columns with missing values.</li>\n<li>Some columns(<strong>R_7,R_14</strong>) has <strong>only 1 rows missing</strong> but removing them would be a trouble because those are missing for a single month in between, so removal can be dangerous, or we can remove the customer altogether - not sure what to do?</li>\n<li>Some columns(<strong>D_49, D_132, D_106, B_29,R_9,D_134,D_135,D_136,D_137,D_138,B_42,D_73,B_39,D_111,D_110,D_108,D_88,D_87</strong>) have <strong>more than 90% of the total records are missing</strong> - how to deal with them?</li>\n<li>similarly how to deal with missing values and outliers in this case.</li>\n</ul>\n<p>So in general how to deal with the missing values and outliers in this case wherein we don't possess the actual meaning of the columns, suggest some notebook or discussion/links or ideas, if any.</p>",
  "messages": [
    {
      "id": "1821198",
      "postDate": "06/15/2022 10:10:44",
      "content": "<p>How to go about <strong>missing value treatment</strong> in the AMEX data or outlier treatment?<br>\nsome pointers:</p>\n<ul>\n<li>There are **122 **columns with missing values.</li>\n<li>Some columns(<strong>R_7,R_14</strong>) has <strong>only 1 rows missing</strong> but removing them would be a trouble because those are missing for a single month in between, so removal can be dangerous, or we can remove the customer altogether - not sure what to do?</li>\n<li>Some columns(<strong>D_49, D_132, D_106, B_29,R_9,D_134,D_135,D_136,D_137,D_138,B_42,D_73,B_39,D_111,D_110,D_108,D_88,D_87</strong>) have <strong>more than 90% of the total records are missing</strong> - how to deal with them?</li>\n<li>similarly how to deal with missing values and outliers in this case.</li>\n</ul>\n<p>So in general how to deal with the missing values and outliers in this case wherein we don't possess the actual meaning of the columns, suggest some notebook or discussion/links or ideas, if any.</p>",
      "rawMarkdown": "How to go about **missing value treatment** in the AMEX data or outlier treatment?\nsome pointers:\n- There are **122 **columns with missing values.\n- Some columns(**R_7,R_14**) has **only 1 rows missing** but removing them would be a trouble because those are missing for a single month in between, so removal can be dangerous, or we can remove the customer altogether - not sure what to do?\n- Some columns(**D_49, D_132, D_106, B_29,R_9,D_134,D_135,D_136,D_137,D_138,B_42,D_73,B_39,D_111,D_110,D_108,D_88,D_87**) have **more than 90% of the total records are missing** - how to deal with them?\n- similarly how to deal with missing values and outliers in this case.\n\nSo in general how to deal with the missing values and outliers in this case wherein we don't possess the actual meaning of the columns, suggest some notebook or discussion/links or ideas, if any.",
      "votes": null
    },
    {
      "id": "1821256",
      "postDate": "06/15/2022 11:22:19",
      "content": "<p>It is important to know that there are 2 types of missing values: 1) missing at random and 2) missing by design.</p>\n<p>I am pretty sure 2) is dominant for this dataset. That means its should be treated like any other number (for example assign to -1)</p>",
      "rawMarkdown": "It is important to know that there are 2 types of missing values: 1) missing at random and 2) missing by design.\n\nI am pretty sure 2) is dominant for this dataset. That means its should be treated like any other number (for example assign to -1)",
      "votes": null
    },
    {
      "id": "1821294",
      "postDate": "06/15/2022 12:01:05",
      "content": "<p>The right treatment depends on the algorithm you're going to use. LightGBM or XGBoost, for instance, handle missing values and outliers automatically. With neural networks you'll get good results by simply filling all missing values with 0.</p>",
      "rawMarkdown": "The right treatment depends on the algorithm you're going to use. LightGBM or XGBoost, for instance, handle missing values and outliers automatically. With neural networks you'll get good results by simply filling all missing values with 0.",
      "votes": null
    },
    {
      "id": "1821359",
      "postDate": "06/15/2022 13:23:41",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> , I have also found these which may be useful :- </p>\n<ul>\n<li>remove columns with high missing values like more than 90% of data missing - but only for classical or traditional algorithm</li>\n<li>mean/median/mode would work - matter of trial &amp; error.</li>\n<li>going for model based imputation - again matter of trial and also depends on the fraction of improvements.</li>\n<li>remove the customers with more number of month values are missing</li>\n</ul>\n<p>Is there a way to identify random missing or missing intentionally - any analysis which may help?</p>",
      "rawMarkdown": "Thanks @raddar , I have also found these which may be useful :- \n- remove columns with high missing values like more than 90% of data missing - but only for classical or traditional algorithm\n- mean/median/mode would work - matter of trial & error.\n- going for model based imputation - again matter of trial and also depends on the fraction of improvements.\n- remove the customers with more number of month values are missing\n\nIs there a way to identify random missing or missing intentionally - any analysis which may help?",
      "votes": null
    },
    {
      "id": "1821365",
      "postDate": "06/15/2022 13:25:44",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <br>\nWhat if we apply classical algorithm like logistic regression(may be slow), any suggestion on those would be helpful for the larger group.</p>",
      "rawMarkdown": "Thanks @ambrosm \nWhat if we apply classical algorithm like logistic regression(may be slow), any suggestion on those would be helpful for the larger group.",
      "votes": null
    },
    {
      "id": "1821432",
      "postDate": "06/15/2022 14:35:18",
      "content": "<p>If you want to apply logistic regression, I think that a nice approach will be: Optimal Binning + WoE + Logistic Regression. Optbinning handles missing values. Also I've read in Naeem Siddiqqi's book \"Intelligent Credit Scoring\" that WoE method allows you not to worry about outliers</p>",
      "rawMarkdown": "If you want to apply logistic regression, I think that a nice approach will be: Optimal Binning + WoE + Logistic Regression. Optbinning handles missing values. Also I've read in Naeem Siddiqqi's book \"Intelligent Credit Scoring\" that WoE method allows you not to worry about outliers",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1821256,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/15/2022 11:22:19",
      "content": "<p>It is important to know that there are 2 types of missing values: 1) missing at random and 2) missing by design.</p>\n<p>I am pretty sure 2) is dominant for this dataset. That means its should be treated like any other number (for example assign to -1)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1821359,
          "author_name": "abhianalytic",
          "author_url": "",
          "post_date": "06/15/2022 13:23:41",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> , I have also found these which may be useful :- </p>\n<ul>\n<li>remove columns with high missing values like more than 90% of data missing - but only for classical or traditional algorithm</li>\n<li>mean/median/mode would work - matter of trial &amp; error.</li>\n<li>going for model based imputation - again matter of trial and also depends on the fraction of improvements.</li>\n<li>remove the customers with more number of month values are missing</li>\n</ul>\n<p>Is there a way to identify random missing or missing intentionally - any analysis which may help?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1821294,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "06/15/2022 12:01:05",
      "content": "<p>The right treatment depends on the algorithm you're going to use. LightGBM or XGBoost, for instance, handle missing values and outliers automatically. With neural networks you'll get good results by simply filling all missing values with 0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1821365,
          "author_name": "abhianalytic",
          "author_url": "",
          "post_date": "06/15/2022 13:25:44",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <br>\nWhat if we apply classical algorithm like logistic regression(may be slow), any suggestion on those would be helpful for the larger group.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1821432,
          "author_name": "denisplaj",
          "author_url": "",
          "post_date": "06/15/2022 14:35:18",
          "content": "<p>If you want to apply logistic regression, I think that a nice approach will be: Optimal Binning + WoE + Logistic Regression. Optbinning handles missing values. Also I've read in Naeem Siddiqqi's book \"Intelligent Credit Scoring\" that WoE method allows you not to worry about outliers</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1821198": "How to go about **missing value treatment** in the AMEX data or outlier treatment?\nsome pointers:\n- There are **122 **columns with missing values.\n- Some columns(**R_7,R_14**) has **only 1 rows missing** but removing them would be a trouble because those are missing for a single month in between, so removal can be dangerous, or we can remove the customer altogether - not sure what to do?\n- Some columns(**D_49, D_132, D_106, B_29,R_9,D_134,D_135,D_136,D_137,D_138,B_42,D_73,B_39,D_111,D_110,D_108,D_88,D_87**) have **more than 90% of the total records are missing** - how to deal with them?\n- similarly how to deal with missing values and outliers in this case.\n\nSo in general how to deal with the missing values and outliers in this case wherein we don't possess the actual meaning of the columns, suggest some notebook or discussion/links or ideas, if any.",
    "1821256": "It is important to know that there are 2 types of missing values: 1) missing at random and 2) missing by design.\n\nI am pretty sure 2) is dominant for this dataset. That means its should be treated like any other number (for example assign to -1)",
    "1821294": "The right treatment depends on the algorithm you're going to use. LightGBM or XGBoost, for instance, handle missing values and outliers automatically. With neural networks you'll get good results by simply filling all missing values with 0.",
    "1821359": "Thanks @raddar , I have also found these which may be useful :- \n- remove columns with high missing values like more than 90% of data missing - but only for classical or traditional algorithm\n- mean/median/mode would work - matter of trial & error.\n- going for model based imputation - again matter of trial and also depends on the fraction of improvements.\n- remove the customers with more number of month values are missing\n\nIs there a way to identify random missing or missing intentionally - any analysis which may help?",
    "1821365": "Thanks @ambrosm \nWhat if we apply classical algorithm like logistic regression(may be slow), any suggestion on those would be helpful for the larger group.",
    "1821432": "If you want to apply logistic regression, I think that a nice approach will be: Optimal Binning + WoE + Logistic Regression. Optbinning handles missing values. Also I've read in Naeem Siddiqqi's book \"Intelligent Credit Scoring\" that WoE method allows you not to worry about outliers"
  },
  "source": "meta"
}