{
  "id": 471445,
  "title": "bad sample of eeg_id=1457334423(spec_id =1659812292)",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/471445",
  "author_name": "",
  "post_date": "2024-01-28T11:12:19.010788500Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am creating my own spectrograms, while accidentally I found there is one bad sample in the training data, with all values of 0.16, and the corresponding spectrograms being all nans. That affects the results seriously. So I think that  1457334423 should be excluded from traning model. </p>",
  "messages": [
    {
      "id": "2623769",
      "postDate": "01/28/2024 11:12:19",
      "content": "<p>I am creating my own spectrograms, while accidentally I found there is one bad sample in the training data, with all values of 0.16, and the corresponding spectrograms being all nans. That affects the results seriously. So I think that  1457334423 should be excluded from traning model. </p>",
      "rawMarkdown": "I am creating my own spectrograms, while accidentally I found there is one bad sample in the training data, with all values of 0.16, and the corresponding spectrograms being all nans. That affects the results seriously. So I think that  1457334423 should be excluded from traning model.",
      "votes": null
    },
    {
      "id": "2623872",
      "postDate": "01/28/2024 12:11:30",
      "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> i found 853 sub-sequences can be excluded since no 'nan' values in the private dataset <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044\" target=\"_blank\">here</a> </p>",
      "rawMarkdown": "niuweikun i found 853 sub-sequences can be excluded since no 'nan' values in the private dataset [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044)",
      "votes": null
    },
    {
      "id": "2624057",
      "postDate": "01/28/2024 14:47:44",
      "content": "<p>wow great work.  So just remove all that samples with at least one nan seems reasonable in this way</p>",
      "rawMarkdown": "wow great work.  So just remove all that samples with at least one nan seems reasonable in this way",
      "votes": null
    },
    {
      "id": "2624088",
      "postDate": "01/28/2024 15:15:07",
      "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> we can use it as soft labels too, at present i am training with removing these samples. since only 853 samples ~ 0.8% data to be removed as my analysis</p>",
      "rawMarkdown": "niuweikun we can use it as soft labels too, at present i am training with removing these samples. since only 853 samples ~ 0.8% data to be removed as my analysis",
      "votes": null
    },
    {
      "id": "2633202",
      "postDate": "02/02/2024 21:32:27",
      "content": "<p>This is very helpful!  I wondered about the effect of those NaN entries.  Considering several options: 1) if TRAIN.csv has an eeg_id that points to the deleted spectrogram, delete that eeg_id/sprectrogram row from the train dataset generated from the TRAIN.csv.  2) handle the key not found exception &amp; use the original datasets.  Thoughts?</p>",
      "rawMarkdown": "This is very helpful!  I wondered about the effect of those NaN entries.  Considering several options: 1) if TRAIN.csv has an eeg_id that points to the deleted spectrogram, delete that eeg_id/sprectrogram row from the train dataset generated from the TRAIN.csv.  2) handle the key not found exception & use the original datasets.  Thoughts?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2623872,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/28/2024 12:11:30",
      "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> i found 853 sub-sequences can be excluded since no 'nan' values in the private dataset <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044\" target=\"_blank\">here</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2624057,
          "author_name": "niuweikun",
          "author_url": "",
          "post_date": "01/28/2024 14:47:44",
          "content": "<p>wow great work.  So just remove all that samples with at least one nan seems reasonable in this way</p>",
          "votes": null,
          "replies": [
            {
              "id": 2624088,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "01/28/2024 15:15:07",
              "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> we can use it as soft labels too, at present i am training with removing these samples. since only 853 samples ~ 0.8% data to be removed as my analysis</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2633202,
      "author_name": "matucker",
      "author_url": "",
      "post_date": "02/02/2024 21:32:27",
      "content": "<p>This is very helpful!  I wondered about the effect of those NaN entries.  Considering several options: 1) if TRAIN.csv has an eeg_id that points to the deleted spectrogram, delete that eeg_id/sprectrogram row from the train dataset generated from the TRAIN.csv.  2) handle the key not found exception &amp; use the original datasets.  Thoughts?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2623769": "I am creating my own spectrograms, while accidentally I found there is one bad sample in the training data, with all values of 0.16, and the corresponding spectrograms being all nans. That affects the results seriously. So I think that  1457334423 should be excluded from traning model.",
    "2623872": "niuweikun i found 853 sub-sequences can be excluded since no 'nan' values in the private dataset [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044)",
    "2624057": "wow great work.  So just remove all that samples with at least one nan seems reasonable in this way",
    "2624088": "niuweikun we can use it as soft labels too, at present i am training with removing these samples. since only 853 samples ~ 0.8% data to be removed as my analysis",
    "2633202": "This is very helpful!  I wondered about the effect of those NaN entries.  Considering several options: 1) if TRAIN.csv has an eeg_id that points to the deleted spectrogram, delete that eeg_id/sprectrogram row from the train dataset generated from the TRAIN.csv.  2) handle the key not found exception & use the original datasets.  Thoughts?"
  },
  "source": "meta"
}