{
  "id": 551233,
  "title": "Handling NaN Target Values in a Noisy Dataset",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551233",
  "author_name": "",
  "post_date": "2024-12-12T03:14:01.402760600Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello, I hope you have learned some good things :) <br>\nI have a few things I'm curious about. When the data is already noisy and has a lot of randomness, wouldn't filling 1,224 rows with NaN target values using various methods make the dataset even more random and noisy? I don't think we need to do this. Maybe instead of predicting all the NaN data, we could add one-fold of it to the data. What do you think?</p>",
  "messages": [
    {
      "id": "3069907",
      "postDate": "12/12/2024 03:14:01",
      "content": "<p>Hello, I hope you have learned some good things :) <br>\nI have a few things I'm curious about. When the data is already noisy and has a lot of randomness, wouldn't filling 1,224 rows with NaN target values using various methods make the dataset even more random and noisy? I don't think we need to do this. Maybe instead of predicting all the NaN data, we could add one-fold of it to the data. What do you think?</p>",
      "rawMarkdown": "Hello, I hope you have learned some good things :) \nI have a few things I'm curious about. When the data is already noisy and has a lot of randomness, wouldn't filling 1,224 rows with NaN target values using various methods make the dataset even more random and noisy? I don't think we need to do this. Maybe instead of predicting all the NaN data, we could add one-fold of it to the data. What do you think?",
      "votes": null
    },
    {
      "id": "3070038",
      "postDate": "12/12/2024 07:02:18",
      "content": "<p>I agree with everything you said. I also tried to just impute single features. Sometimes it improved CV, but it seemed by chance.<br>\nFor example, I imputed the PAQ Scores and it improved my CV a bit. But after checking the feature importance, PAQ Score went down, but PAQ Season (Unknown) went up. So it took the imputed values and used the season as indicator that these values are imputed.<br>\nI didnt know what to make of it, since my CV only changed a bit and I had to tune the model again with the risk of just including more noise like you said. So I removed it.</p>",
      "rawMarkdown": "I agree with everything you said. I also tried to just impute single features. Sometimes it improved CV, but it seemed by chance.\nFor example, I imputed the PAQ Scores and it improved my CV a bit. But after checking the feature importance, PAQ Score went down, but PAQ Season (Unknown) went up. So it took the imputed values and used the season as indicator that these values are imputed.\nI didnt know what to make of it, since my CV only changed a bit and I had to tune the model again with the risk of just including more noise like you said. So I removed it.",
      "votes": null
    },
    {
      "id": "3070812",
      "postDate": "12/13/2024 04:19:01",
      "content": "<p><a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a>  It's important to consider how the LB scores are affected. If they are significantly impacted, it might be worth looking into those features. Thank you 🎉.</p>",
      "rawMarkdown": "mariusheuser  It's important to consider how the LB scores are affected. If they are significantly impacted, it might be worth looking into those features. Thank you 🎉.",
      "votes": null
    },
    {
      "id": "3075366",
      "postDate": "12/18/2024 18:40:37",
      "content": "<p>I agree with you that filling all the NaN values could make the dataset even noisier. However, I also think that techniques like random forests can do a better job of handling the imputation. I’ve actually tried it before, but it took way too long to run, so it’s not without its drawbacks.</p>",
      "rawMarkdown": "I agree with you that filling all the NaN values could make the dataset even noisier. However, I also think that techniques like random forests can do a better job of handling the imputation. I’ve actually tried it before, but it took way too long to run, so it’s not without its drawbacks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3070038,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "12/12/2024 07:02:18",
      "content": "<p>I agree with everything you said. I also tried to just impute single features. Sometimes it improved CV, but it seemed by chance.<br>\nFor example, I imputed the PAQ Scores and it improved my CV a bit. But after checking the feature importance, PAQ Score went down, but PAQ Season (Unknown) went up. So it took the imputed values and used the season as indicator that these values are imputed.<br>\nI didnt know what to make of it, since my CV only changed a bit and I had to tune the model again with the risk of just including more noise like you said. So I removed it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3070812,
          "author_name": "trcnveli",
          "author_url": "",
          "post_date": "12/13/2024 04:19:01",
          "content": "<p><a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a>  It's important to consider how the LB scores are affected. If they are significantly impacted, it might be worth looking into those features. Thank you 🎉.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3075366,
      "author_name": "viktoriamelkumyan",
      "author_url": "",
      "post_date": "12/18/2024 18:40:37",
      "content": "<p>I agree with you that filling all the NaN values could make the dataset even noisier. However, I also think that techniques like random forests can do a better job of handling the imputation. I’ve actually tried it before, but it took way too long to run, so it’s not without its drawbacks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3069907": "Hello, I hope you have learned some good things :) \nI have a few things I'm curious about. When the data is already noisy and has a lot of randomness, wouldn't filling 1,224 rows with NaN target values using various methods make the dataset even more random and noisy? I don't think we need to do this. Maybe instead of predicting all the NaN data, we could add one-fold of it to the data. What do you think?",
    "3070038": "I agree with everything you said. I also tried to just impute single features. Sometimes it improved CV, but it seemed by chance.\nFor example, I imputed the PAQ Scores and it improved my CV a bit. But after checking the feature importance, PAQ Score went down, but PAQ Season (Unknown) went up. So it took the imputed values and used the season as indicator that these values are imputed.\nI didnt know what to make of it, since my CV only changed a bit and I had to tune the model again with the risk of just including more noise like you said. So I removed it.",
    "3070812": "mariusheuser  It's important to consider how the LB scores are affected. If they are significantly impacted, it might be worth looking into those features. Thank you 🎉.",
    "3075366": "I agree with you that filling all the NaN values could make the dataset even noisier. However, I also think that techniques like random forests can do a better job of handling the imputation. I’ve actually tried it before, but it took way too long to run, so it’s not without its drawbacks."
  },
  "source": "meta"
}