{
  "id": 477248,
  "title": "Data cleaning",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477248",
  "author_name": "",
  "post_date": "2024-02-15T10:35:20.139023100Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In EEG and Spectrogram data, there are numerous NaN values, how should these values be handled? Should the corresponding rows be dropped or mean/median/KNN Imputation is a good idea. This is a tricky question to answer, since we are dealing with signal frequencies, where minor changes may lead to bigger changes in the output.</p>",
  "messages": [
    {
      "id": "2653280",
      "postDate": "02/15/2024 10:35:20",
      "content": "<p>In EEG and Spectrogram data, there are numerous NaN values, how should these values be handled? Should the corresponding rows be dropped or mean/median/KNN Imputation is a good idea. This is a tricky question to answer, since we are dealing with signal frequencies, where minor changes may lead to bigger changes in the output.</p>",
      "rawMarkdown": "In EEG and Spectrogram data, there are numerous NaN values, how should these values be handled? Should the corresponding rows be dropped or mean/median/KNN Imputation is a good idea. This is a tricky question to answer, since we are dealing with signal frequencies, where minor changes may lead to bigger changes in the output.",
      "votes": null
    },
    {
      "id": "2660877",
      "postDate": "02/20/2024 21:24:52",
      "content": "<p>Hi - I dropped NaN rows with a good improvement in CV/LB scores.  Many discussions note the difficulty or (impossibility) of accurately imputing NaNs.  It'd be great to hear if you successfully imputed NaNs.  </p>",
      "rawMarkdown": "Hi - I dropped NaN rows with a good improvement in CV/LB scores.  Many discussions note the difficulty or (impossibility) of accurately imputing NaNs.  It'd be great to hear if you successfully imputed NaNs.",
      "votes": null
    },
    {
      "id": "2660883",
      "postDate": "02/20/2024 21:32:30",
      "content": "<p>You will have to experiment with this and see what gives you a better CV/LB score. This is indeed a tricky question to answer because everyone will have to experiment with it and see what works. </p>",
      "rawMarkdown": "You will have to experiment with this and see what gives you a better CV/LB score. This is indeed a tricky question to answer because everyone will have to experiment with it and see what works.",
      "votes": null
    },
    {
      "id": "2662034",
      "postDate": "02/21/2024 16:58:19",
      "content": "<p>When you dropped them, did you need to do additional preprocessing to fill in the missing pieces in the spectrogram? </p>",
      "rawMarkdown": "When you dropped them, did you need to do additional preprocessing to fill in the missing pieces in the spectrogram?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2660877,
      "author_name": "matucker",
      "author_url": "",
      "post_date": "02/20/2024 21:24:52",
      "content": "<p>Hi - I dropped NaN rows with a good improvement in CV/LB scores.  Many discussions note the difficulty or (impossibility) of accurately imputing NaNs.  It'd be great to hear if you successfully imputed NaNs.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2662034,
          "author_name": "zoestankow",
          "author_url": "",
          "post_date": "02/21/2024 16:58:19",
          "content": "<p>When you dropped them, did you need to do additional preprocessing to fill in the missing pieces in the spectrogram? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2660883,
      "author_name": "feidawei",
      "author_url": "",
      "post_date": "02/20/2024 21:32:30",
      "content": "<p>You will have to experiment with this and see what gives you a better CV/LB score. This is indeed a tricky question to answer because everyone will have to experiment with it and see what works. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2653280": "In EEG and Spectrogram data, there are numerous NaN values, how should these values be handled? Should the corresponding rows be dropped or mean/median/KNN Imputation is a good idea. This is a tricky question to answer, since we are dealing with signal frequencies, where minor changes may lead to bigger changes in the output.",
    "2660877": "Hi - I dropped NaN rows with a good improvement in CV/LB scores.  Many discussions note the difficulty or (impossibility) of accurately imputing NaNs.  It'd be great to hear if you successfully imputed NaNs.",
    "2660883": "You will have to experiment with this and see what gives you a better CV/LB score. This is indeed a tricky question to answer because everyone will have to experiment with it and see what works.",
    "2662034": "When you dropped them, did you need to do additional preprocessing to fill in the missing pieces in the spectrogram?"
  },
  "source": "meta"
}