{
  "id": 548020,
  "title": "Missing values",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/548020",
  "author_name": "",
  "post_date": "2024-11-24T16:55:24.539040500Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all, when I am doing EDA of training dataset, I noticed that missing values are prominent (just as many posts in the discussion point out). When I look at the percentage of missing values in each features (excluding PCIAT), the distribution shows that most of the features have 20% of values missing, and more than 10 features have 50% of values missing. In this case, should the features with more than half of the values missing be discarded?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17330114%2Fc838a475937fb4bce7d308db098c8f02%2Fmissing_value_distribution.png?generation=1732467348401457&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3054443",
      "postDate": "11/24/2024 16:55:24",
      "content": "<p>Hi all, when I am doing EDA of training dataset, I noticed that missing values are prominent (just as many posts in the discussion point out). When I look at the percentage of missing values in each features (excluding PCIAT), the distribution shows that most of the features have 20% of values missing, and more than 10 features have 50% of values missing. In this case, should the features with more than half of the values missing be discarded?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17330114%2Fc838a475937fb4bce7d308db098c8f02%2Fmissing_value_distribution.png?generation=1732467348401457&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi all, when I am doing EDA of training dataset, I noticed that missing values are prominent (just as many posts in the discussion point out). When I look at the percentage of missing values in each features (excluding PCIAT), the distribution shows that most of the features have 20% of values missing, and more than 10 features have 50% of values missing. In this case, should the features with more than half of the values missing be discarded?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17330114%2Fc838a475937fb4bce7d308db098c8f02%2Fmissing_value_distribution.png?generation=1732467348401457&alt=media)",
      "votes": null
    },
    {
      "id": "3054513",
      "postDate": "11/24/2024 18:48:44",
      "content": "<p>try to find the correlation of these features to decide </p>",
      "rawMarkdown": "try to find the correlation of these features to decide",
      "votes": null
    },
    {
      "id": "3054619",
      "postDate": "11/24/2024 22:48:11",
      "content": "<p>In general practice a feature with more than 70% missing values only then a complete feature may be discarded. </p>\n<p>But it also depends on importance of each feature as well<br>\nCheck if these features are strongly correlated with the target variable (using correlation or feature importance score). If they are higly correlative it would be worth imputing the values rather than discarding them.</p>",
      "rawMarkdown": "In general practice a feature with more than 70% missing values only then a complete feature may be discarded. \n\nBut it also depends on importance of each feature as well\nCheck if these features are strongly correlated with the target variable (using correlation or feature importance score). If they are higly correlative it would be worth imputing the values rather than discarding them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3054513,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "11/24/2024 18:48:44",
      "content": "<p>try to find the correlation of these features to decide </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3054619,
      "author_name": "adilnadeem98",
      "author_url": "",
      "post_date": "11/24/2024 22:48:11",
      "content": "<p>In general practice a feature with more than 70% missing values only then a complete feature may be discarded. </p>\n<p>But it also depends on importance of each feature as well<br>\nCheck if these features are strongly correlated with the target variable (using correlation or feature importance score). If they are higly correlative it would be worth imputing the values rather than discarding them.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3054443": "Hi all, when I am doing EDA of training dataset, I noticed that missing values are prominent (just as many posts in the discussion point out). When I look at the percentage of missing values in each features (excluding PCIAT), the distribution shows that most of the features have 20% of values missing, and more than 10 features have 50% of values missing. In this case, should the features with more than half of the values missing be discarded?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17330114%2Fc838a475937fb4bce7d308db098c8f02%2Fmissing_value_distribution.png?generation=1732467348401457&alt=media)",
    "3054513": "try to find the correlation of these features to decide",
    "3054619": "In general practice a feature with more than 70% missing values only then a complete feature may be discarded. \n\nBut it also depends on importance of each feature as well\nCheck if these features are strongly correlated with the target variable (using correlation or feature importance score). If they are higly correlative it would be worth imputing the values rather than discarding them."
  },
  "source": "meta"
}