{
  "id": 483387,
  "title": "Handling missing data",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/483387",
  "author_name": "",
  "post_date": "2024-03-12T06:49:24.289574400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>How do you all handle the large amount of null values? I've only joined files on depth=0 but still have like 100 features with a null rate &gt;70%. Do you go through these one by one and see what to replace them with (0 doesn't make sense as a replacement for many of them), or do you just use models that can handle missing data out of the box instead?</p>",
  "messages": [
    {
      "id": "2693014",
      "postDate": "03/12/2024 06:49:24",
      "content": "<p>How do you all handle the large amount of null values? I've only joined files on depth=0 but still have like 100 features with a null rate &gt;70%. Do you go through these one by one and see what to replace them with (0 doesn't make sense as a replacement for many of them), or do you just use models that can handle missing data out of the box instead?</p>",
      "rawMarkdown": "How do you all handle the large amount of null values? I've only joined files on depth=0 but still have like 100 features with a null rate >70%. Do you go through these one by one and see what to replace them with (0 doesn't make sense as a replacement for many of them), or do you just use models that can handle missing data out of the box instead?",
      "votes": null
    },
    {
      "id": "2693143",
      "postDate": "03/12/2024 08:43:20",
      "content": "<p>For now, I am not doing anything for these features and am using boosted tree models for the same. I think dropping features with too many nulls is a recourse, but for a little later <a href=\"https://www.kaggle.com/alexanderholmberg\" target=\"_blank\">@alexanderholmberg</a> </p>",
      "rawMarkdown": "For now, I am not doing anything for these features and am using boosted tree models for the same. I think dropping features with too many nulls is a recourse, but for a little later @alexanderholmberg",
      "votes": null
    },
    {
      "id": "2693537",
      "postDate": "03/12/2024 13:50:50",
      "content": "<p>We generally don't see the practice of dropping columns with a high null ratio as good for modelling, but it definitely can be a good start. You can also in the first approximation treat the features univariate with binning to encode them with mean target encoding <a href=\"https://gnpalencia.org/optbinning/tutorials/tutorial_continuous.html\" target=\"_blank\">https://gnpalencia.org/optbinning/tutorials/tutorial_continuous.html</a> but that is still an approximation.</p>",
      "rawMarkdown": "We generally don't see the practice of dropping columns with a high null ratio as good for modelling, but it definitely can be a good start. You can also in the first approximation treat the features univariate with binning to encode them with mean target encoding https://gnpalencia.org/optbinning/tutorials/tutorial_continuous.html but that is still an approximation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2693143,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "03/12/2024 08:43:20",
      "content": "<p>For now, I am not doing anything for these features and am using boosted tree models for the same. I think dropping features with too many nulls is a recourse, but for a little later <a href=\"https://www.kaggle.com/alexanderholmberg\" target=\"_blank\">@alexanderholmberg</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2693537,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "03/12/2024 13:50:50",
      "content": "<p>We generally don't see the practice of dropping columns with a high null ratio as good for modelling, but it definitely can be a good start. You can also in the first approximation treat the features univariate with binning to encode them with mean target encoding <a href=\"https://gnpalencia.org/optbinning/tutorials/tutorial_continuous.html\" target=\"_blank\">https://gnpalencia.org/optbinning/tutorials/tutorial_continuous.html</a> but that is still an approximation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2693014": "How do you all handle the large amount of null values? I've only joined files on depth=0 but still have like 100 features with a null rate >70%. Do you go through these one by one and see what to replace them with (0 doesn't make sense as a replacement for many of them), or do you just use models that can handle missing data out of the box instead?",
    "2693143": "For now, I am not doing anything for these features and am using boosted tree models for the same. I think dropping features with too many nulls is a recourse, but for a little later @alexanderholmberg",
    "2693537": "We generally don't see the practice of dropping columns with a high null ratio as good for modelling, but it definitely can be a good start. You can also in the first approximation treat the features univariate with binning to encode them with mean target encoding https://gnpalencia.org/optbinning/tutorials/tutorial_continuous.html but that is still an approximation."
  },
  "source": "meta"
}