{
  "id": 535226,
  "title": "Missing values in training dataset",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535226",
  "author_name": "",
  "post_date": "2024-09-20T21:08:16.271855700Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I have noticed that there are a many features in the train.csv file. The simple way to deal with these would be to fill these NaN values using some imputation technique. However, I wanted to know if there is a better way of doing this, possibly using other correlated features that are present?</p>",
  "messages": [
    {
      "id": "2994383",
      "postDate": "09/20/2024 21:08:16",
      "content": "<p>Hi all,</p>\n<p>I have noticed that there are a many features in the train.csv file. The simple way to deal with these would be to fill these NaN values using some imputation technique. However, I wanted to know if there is a better way of doing this, possibly using other correlated features that are present?</p>",
      "rawMarkdown": "Hi all,\n\nI have noticed that there are a many features in the train.csv file. The simple way to deal with these would be to fill these NaN values using some imputation technique. However, I wanted to know if there is a better way of doing this, possibly using other correlated features that are present?",
      "votes": null
    },
    {
      "id": "3002291",
      "postDate": "09/29/2024 21:32:40",
      "content": "<p>Hey Rafid,<br>\nI think it's also can  Regression Imputation and Predictive to imput the missing value,<br>\n For continuous variables, you can build a regression model using other features that are highly correlated with the missing values. The model can predict the missing values based on the relationships between these features. For categorical variables, you can use classification models (e.g., Decision Trees, Random Forests) trained on non-missing data to predict the missing categories.</p>",
      "rawMarkdown": "Hey Rafid,\nI think it's also can  Regression Imputation and Predictive to imput the missing value,\n For continuous variables, you can build a regression model using other features that are highly correlated with the missing values. The model can predict the missing values based on the relationships between these features. For categorical variables, you can use classification models (e.g., Decision Trees, Random Forests) trained on non-missing data to predict the missing categories.",
      "votes": null
    },
    {
      "id": "3002868",
      "postDate": "09/30/2024 13:28:29",
      "content": "<p>Thanks for the suggestion. So far, I have tried KNN imputation with some success. I will try the regression method as well.</p>",
      "rawMarkdown": "Thanks for the suggestion. So far, I have tried KNN imputation with some success. I will try the regression method as well.",
      "votes": null
    },
    {
      "id": "3005526",
      "postDate": "10/03/2024 03:45:15",
      "content": "<p>When the amount of missing data is too large, any gap-filling method will have little success</p>",
      "rawMarkdown": "When the amount of missing data is too large, any gap-filling method will have little success",
      "votes": null
    },
    {
      "id": "3005856",
      "postDate": "10/03/2024 12:29:46",
      "content": "<p>What do you suggest we do in that case?</p>",
      "rawMarkdown": "What do you suggest we do in that case?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3002291,
      "author_name": "lanacai",
      "author_url": "",
      "post_date": "09/29/2024 21:32:40",
      "content": "<p>Hey Rafid,<br>\nI think it's also can  Regression Imputation and Predictive to imput the missing value,<br>\n For continuous variables, you can build a regression model using other features that are highly correlated with the missing values. The model can predict the missing values based on the relationships between these features. For categorical variables, you can use classification models (e.g., Decision Trees, Random Forests) trained on non-missing data to predict the missing categories.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3002868,
          "author_name": "rafidmahbub",
          "author_url": "",
          "post_date": "09/30/2024 13:28:29",
          "content": "<p>Thanks for the suggestion. So far, I have tried KNN imputation with some success. I will try the regression method as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3005526,
      "author_name": "leozwt",
      "author_url": "",
      "post_date": "10/03/2024 03:45:15",
      "content": "<p>When the amount of missing data is too large, any gap-filling method will have little success</p>",
      "votes": null,
      "replies": [
        {
          "id": 3005856,
          "author_name": "rafidmahbub",
          "author_url": "",
          "post_date": "10/03/2024 12:29:46",
          "content": "<p>What do you suggest we do in that case?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2994383": "Hi all,\n\nI have noticed that there are a many features in the train.csv file. The simple way to deal with these would be to fill these NaN values using some imputation technique. However, I wanted to know if there is a better way of doing this, possibly using other correlated features that are present?",
    "3002291": "Hey Rafid,\nI think it's also can  Regression Imputation and Predictive to imput the missing value,\n For continuous variables, you can build a regression model using other features that are highly correlated with the missing values. The model can predict the missing values based on the relationships between these features. For categorical variables, you can use classification models (e.g., Decision Trees, Random Forests) trained on non-missing data to predict the missing categories.",
    "3002868": "Thanks for the suggestion. So far, I have tried KNN imputation with some success. I will try the regression method as well.",
    "3005526": "When the amount of missing data is too large, any gap-filling method will have little success",
    "3005856": "What do you suggest we do in that case?"
  },
  "source": "meta"
}