{
  "id": 475730,
  "title": "Why don't orgs provide raw EEGs?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/475730",
  "author_name": "",
  "post_date": "2024-02-09T15:43:22.735927900Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I just joined the competition, and I'm confused about why we are given 50-second-long EEGs and 10-minute-long spectrograms. Have I correctly understood that spectrograms somehow cover a wider time range compared to raw EEGs?</p>\n<p>Why won't orgs provide us with the complete raw EEGs [10 mins], to allow us to reconstruct the spectrograms ourselves and possibly extract more information from the rawest form of the signal? The current data suggests that to extract the entire signal, one would need to use both spectrograms and raw EEGs, which is cumbersome.</p>",
  "messages": [
    {
      "id": "2644636",
      "postDate": "02/09/2024 15:43:22",
      "content": "<p>I just joined the competition, and I'm confused about why we are given 50-second-long EEGs and 10-minute-long spectrograms. Have I correctly understood that spectrograms somehow cover a wider time range compared to raw EEGs?</p>\n<p>Why won't orgs provide us with the complete raw EEGs [10 mins], to allow us to reconstruct the spectrograms ourselves and possibly extract more information from the rawest form of the signal? The current data suggests that to extract the entire signal, one would need to use both spectrograms and raw EEGs, which is cumbersome.</p>",
      "rawMarkdown": "I just joined the competition, and I'm confused about why we are given 50-second-long EEGs and 10-minute-long spectrograms. Have I correctly understood that spectrograms somehow cover a wider time range compared to raw EEGs?\n\nWhy won't orgs provide us with the complete raw EEGs [10 mins], to allow us to reconstruct the spectrograms ourselves and possibly extract more information from the rawest form of the signal? The current data suggests that to extract the entire signal, one would need to use both spectrograms and raw EEGs, which is cumbersome.",
      "votes": null
    },
    {
      "id": "2644654",
      "postDate": "02/09/2024 15:53:10",
      "content": "<p>Note both EEG and Spectrograms already include extra data. We are classifying the middle 10 seconds of each. Therefore we received 40 extra EEG seconds (which is 8000 extra rows since EEG sampled at 1/200 second). And we received 590 extra Spectrogram seconds (which is 295 extra rows since spectrograms are sampled at 2 second).</p>\n<p>Also note that most train EEG and train Spectrogram are contained inside a longer parquet file. So in many cases we have even more EEG and more Spectrogram (for train but not test).</p>\n<p>More information about competition data in my discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Note both EEG and Spectrograms already include extra data. We are classifying the middle 10 seconds of each. Therefore we received 40 extra EEG seconds (which is 8000 extra rows since EEG sampled at 1/200 second). And we received 590 extra Spectrogram seconds (which is 295 extra rows since spectrograms are sampled at 2 second).\n\nAlso note that most train EEG and train Spectrogram are contained inside a longer parquet file. So in many cases we have even more EEG and more Spectrogram (for train but not test).\n\nMore information about competition data in my discussion [here][1]\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010",
      "votes": null
    },
    {
      "id": "2644672",
      "postDate": "02/09/2024 16:03:49",
      "content": "<p>Sure thing, however seeing that the scores of 1d model doesn't match the score of 2d model which uses wider time context suggests that some signal exists here, also from the reports of people we can see that ensembling these two models boosts the score a lot, therefore I conclude that it not enough to use just 2d [hyperparams of the specs are not optimal?, too much information is lost during conversion to spectrograms?], or just use 1d [time context is too short? signal exists outside 10 or 50 seconds interval?], and it's necessary to create models from both modalities. </p>\n<p>Raw EEGs would've been much better, in competitions like this [e.g. BirdClef] we often don't even have time borders of the label, but still manage to develope classification algorithms that can find signal even if it is very sparse.</p>\n<p>Maybe I'm wrong, I'll try to look into the data more :D</p>",
      "rawMarkdown": "Sure thing, however seeing that the scores of 1d model doesn't match the score of 2d model which uses wider time context suggests that some signal exists here, also from the reports of people we can see that ensembling these two models boosts the score a lot, therefore I conclude that it not enough to use just 2d [hyperparams of the specs are not optimal?, too much information is lost during conversion to spectrograms?], or just use 1d [time context is too short? signal exists outside 10 or 50 seconds interval?], and it's necessary to create models from both modalities. \n\nRaw EEGs would've been much better, in competitions like this [e.g. BirdClef] we often don't even have time borders of the label, but still manage to develope classification algorithms that can find signal even if it is very sparse.\n\nMaybe I'm wrong, I'll try to look into the data more :D",
      "votes": null
    },
    {
      "id": "2644793",
      "postDate": "02/09/2024 17:22:21",
      "content": "<p>The data setup matches the data available to the expert annotators who labelled the dataset.</p>",
      "rawMarkdown": "The data setup matches the data available to the expert annotators who labelled the dataset.",
      "votes": null
    },
    {
      "id": "2645062",
      "postDate": "02/09/2024 22:38:09",
      "content": "<p>With the raw EEG data, we would have data in the order of terabytes</p>",
      "rawMarkdown": "With the raw EEG data, we would have data in the order of terabytes",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2644654,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/09/2024 15:53:10",
      "content": "<p>Note both EEG and Spectrograms already include extra data. We are classifying the middle 10 seconds of each. Therefore we received 40 extra EEG seconds (which is 8000 extra rows since EEG sampled at 1/200 second). And we received 590 extra Spectrogram seconds (which is 295 extra rows since spectrograms are sampled at 2 second).</p>\n<p>Also note that most train EEG and train Spectrogram are contained inside a longer parquet file. So in many cases we have even more EEG and more Spectrogram (for train but not test).</p>\n<p>More information about competition data in my discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2644672,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "02/09/2024 16:03:49",
          "content": "<p>Sure thing, however seeing that the scores of 1d model doesn't match the score of 2d model which uses wider time context suggests that some signal exists here, also from the reports of people we can see that ensembling these two models boosts the score a lot, therefore I conclude that it not enough to use just 2d [hyperparams of the specs are not optimal?, too much information is lost during conversion to spectrograms?], or just use 1d [time context is too short? signal exists outside 10 or 50 seconds interval?], and it's necessary to create models from both modalities. </p>\n<p>Raw EEGs would've been much better, in competitions like this [e.g. BirdClef] we often don't even have time borders of the label, but still manage to develope classification algorithms that can find signal even if it is very sparse.</p>\n<p>Maybe I'm wrong, I'll try to look into the data more :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2644793,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "02/09/2024 17:22:21",
      "content": "<p>The data setup matches the data available to the expert annotators who labelled the dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2645062,
      "author_name": "fabienpv",
      "author_url": "",
      "post_date": "02/09/2024 22:38:09",
      "content": "<p>With the raw EEG data, we would have data in the order of terabytes</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2644636": "I just joined the competition, and I'm confused about why we are given 50-second-long EEGs and 10-minute-long spectrograms. Have I correctly understood that spectrograms somehow cover a wider time range compared to raw EEGs?\n\nWhy won't orgs provide us with the complete raw EEGs [10 mins], to allow us to reconstruct the spectrograms ourselves and possibly extract more information from the rawest form of the signal? The current data suggests that to extract the entire signal, one would need to use both spectrograms and raw EEGs, which is cumbersome.",
    "2644654": "Note both EEG and Spectrograms already include extra data. We are classifying the middle 10 seconds of each. Therefore we received 40 extra EEG seconds (which is 8000 extra rows since EEG sampled at 1/200 second). And we received 590 extra Spectrogram seconds (which is 295 extra rows since spectrograms are sampled at 2 second).\n\nAlso note that most train EEG and train Spectrogram are contained inside a longer parquet file. So in many cases we have even more EEG and more Spectrogram (for train but not test).\n\nMore information about competition data in my discussion [here][1]\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010",
    "2644672": "Sure thing, however seeing that the scores of 1d model doesn't match the score of 2d model which uses wider time context suggests that some signal exists here, also from the reports of people we can see that ensembling these two models boosts the score a lot, therefore I conclude that it not enough to use just 2d [hyperparams of the specs are not optimal?, too much information is lost during conversion to spectrograms?], or just use 1d [time context is too short? signal exists outside 10 or 50 seconds interval?], and it's necessary to create models from both modalities. \n\nRaw EEGs would've been much better, in competitions like this [e.g. BirdClef] we often don't even have time borders of the label, but still manage to develope classification algorithms that can find signal even if it is very sparse.\n\nMaybe I'm wrong, I'll try to look into the data more :D",
    "2644793": "The data setup matches the data available to the expert annotators who labelled the dataset.",
    "2645062": "With the raw EEG data, we would have data in the order of terabytes"
  },
  "source": "meta"
}