{
  "id": 473039,
  "title": "📊 A quick remainder about potential label imbalance",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/473039",
  "author_name": "",
  "post_date": "2024-02-03T07:57:40.647603300Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Just a friendly reminder for those following the open-source public notebooks: when you group by eeg_id and then select only one eeg or spec per id, the resulting dataset may end up with imbalanced labels despite the original data being relatively balanced.</p>\n<p>This implies that some adjustments might be necessary at the initial data processing stage. We should Keep this in mind as we develop models!</p>",
  "messages": [
    {
      "id": "2633677",
      "postDate": "02/03/2024 07:57:40",
      "content": "<p>Just a friendly reminder for those following the open-source public notebooks: when you group by eeg_id and then select only one eeg or spec per id, the resulting dataset may end up with imbalanced labels despite the original data being relatively balanced.</p>\n<p>This implies that some adjustments might be necessary at the initial data processing stage. We should Keep this in mind as we develop models!</p>",
      "rawMarkdown": "Just a friendly reminder for those following the open-source public notebooks: when you group by eeg_id and then select only one eeg or spec per id, the resulting dataset may end up with imbalanced labels despite the original data being relatively balanced.\n\nThis implies that some adjustments might be necessary at the initial data processing stage. We should Keep this in mind as we develop models!",
      "votes": null
    },
    {
      "id": "2633898",
      "postDate": "02/03/2024 11:03:37",
      "content": "<p>Thanks for sharing. I believe most people only see the raw train data is balanced.</p>",
      "rawMarkdown": "Thanks for sharing. I believe most people only see the raw train data is balanced.",
      "votes": null
    },
    {
      "id": "2634077",
      "postDate": "02/03/2024 13:35:42",
      "content": "<p>Interesting, I was doing the separation using grouping by spectrograms and taking the first eeg, what would be the most appropriate way to separate the data in your opinion?</p>",
      "rawMarkdown": "Interesting, I was doing the separation using grouping by spectrograms and taking the first eeg, what would be the most appropriate way to separate the data in your opinion?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2633898,
      "author_name": "wannacui",
      "author_url": "",
      "post_date": "02/03/2024 11:03:37",
      "content": "<p>Thanks for sharing. I believe most people only see the raw train data is balanced.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2634077,
      "author_name": "rafaelzimmermann1",
      "author_url": "",
      "post_date": "02/03/2024 13:35:42",
      "content": "<p>Interesting, I was doing the separation using grouping by spectrograms and taking the first eeg, what would be the most appropriate way to separate the data in your opinion?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2633677": "Just a friendly reminder for those following the open-source public notebooks: when you group by eeg_id and then select only one eeg or spec per id, the resulting dataset may end up with imbalanced labels despite the original data being relatively balanced.\n\nThis implies that some adjustments might be necessary at the initial data processing stage. We should Keep this in mind as we develop models!",
    "2633898": "Thanks for sharing. I believe most people only see the raw train data is balanced.",
    "2634077": "Interesting, I was doing the separation using grouping by spectrograms and taking the first eeg, what would be the most appropriate way to separate the data in your opinion?"
  },
  "source": "meta"
}