{
  "id": 467321,
  "title": "Typo in data description: understanding spectrogram and EEG matching",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467321",
  "author_name": "",
  "post_date": "2024-01-12T03:26:21.587882300Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm struggling to understand how the spectrogram and EEG samples are related. The relevant description in the \"Data\" tab seems to have typos:</p>\n<blockquote>\n  <p>Files<br>\n  train.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.</p>\n</blockquote>\n<p>What does this mean, especially for the spectrogram part? I am a native English speaker, but I cannot parse this. </p>\n<p>To take a concrete example, let's look at the 3rd row from train.csv:</p>\n<pre><code>\n[]: \n                              \n                                   \n                  .\n                          \n                           \n          .\n                            \n                               \n                       Seizure\n                                 \n                                     \n                                     \n                                    \n                                    \n                                   \n</code></pre>\n<p>When I open up the EEG, I suspect I should keep all rows from 18 x 200 to (18 + 50) x 200 (perhaps excluding the last row). Is that correct?</p>\n<p>When I open up the spectrogram file, should I take all rows where time &gt;= 18 and time &lt; 68 (i.e. the same 50 seconds as the EEG)? </p>\n<p>Thanks for any help! </p>",
  "messages": [
    {
      "id": "2597934",
      "postDate": "01/12/2024 03:26:21",
      "content": "<p>I'm struggling to understand how the spectrogram and EEG samples are related. The relevant description in the \"Data\" tab seems to have typos:</p>\n<blockquote>\n  <p>Files<br>\n  train.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.</p>\n</blockquote>\n<p>What does this mean, especially for the spectrogram part? I am a native English speaker, but I cannot parse this. </p>\n<p>To take a concrete example, let's look at the 3rd row from train.csv:</p>\n<pre><code>\n[]: \n                              \n                                   \n                  .\n                          \n                           \n          .\n                            \n                               \n                       Seizure\n                                 \n                                     \n                                     \n                                    \n                                    \n                                   \n</code></pre>\n<p>When I open up the EEG, I suspect I should keep all rows from 18 x 200 to (18 + 50) x 200 (perhaps excluding the last row). Is that correct?</p>\n<p>When I open up the spectrogram file, should I take all rows where time &gt;= 18 and time &lt; 68 (i.e. the same 50 seconds as the EEG)? </p>\n<p>Thanks for any help! </p>",
      "rawMarkdown": "I'm struggling to understand how the spectrogram and EEG samples are related. The relevant description in the \"Data\" tab seems to have typos:\n\n>Files\ntrain.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.\n\nWhat does this mean, especially for the spectrogram part? I am a native English speaker, but I cannot parse this. \n\nTo take a concrete example, let's look at the 3rd row from train.csv:\n```\nIn [33]: mdf.iloc[3, :]\nOut[33]: \neeg_id                              1628180742\neeg_sub_id                                   3\neeg_label_offset_seconds                  18.0\nspectrogram_id                          353733\nspectrogram_sub_id                           3\nspectrogram_label_offset_seconds          18.0\nlabel_id                            2718991173\npatient_id                               42516\nexpert_consensus                       Seizure\nseizure_vote                                 3\nlpd_vote                                     0\ngpd_vote                                     0\nlrda_vote                                    0\ngrda_vote                                    0\nother_vote                                   0\n```\n\nWhen I open up the EEG, I suspect I should keep all rows from 18 x 200 to (18 + 50) x 200 (perhaps excluding the last row). Is that correct?\n\nWhen I open up the spectrogram file, should I take all rows where time >= 18 and time < 68 (i.e. the same 50 seconds as the EEG)? \n\nThanks for any help!",
      "votes": null
    },
    {
      "id": "2597958",
      "postDate": "01/12/2024 04:05:38",
      "content": "<p>The spectrograms are 10 minutes long. So you want <code>time &gt;= 18</code> and <code>time &lt; 18+600</code>. This will be 300 rows because the spectrogram parquets count by 2s. This is the same size as the test spectrograms which are all 300 rows.</p>",
      "rawMarkdown": "The spectrograms are 10 minutes long. So you want `time >= 18` and `time < 18+600`. This will be 300 rows because the spectrogram parquets count by 2s. This is the same size as the test spectrograms which are all 300 rows.",
      "votes": null
    },
    {
      "id": "2597989",
      "postDate": "01/12/2024 04:57:38",
      "content": "<p>Wow, that's really surprising to me. Why would a 50 second EEG portion and its corresponding 10 minute spectrogram even get the same label? They're looking at two really different periods of time!</p>\n<p>The EEG time series overlap for two adjacent sub-samples is like 80-90%, but the spectrogram overlap for the same adjacent subsamples is going to be ~99%. Why bother subsampling at all?</p>\n<p>Relatedly, my impression is that neurologists primarily look at the spectrograms, not the raw EEG time series. Does anyone know if that's the case for the annotators here as well?</p>",
      "rawMarkdown": "Wow, that's really surprising to me. Why would a 50 second EEG portion and its corresponding 10 minute spectrogram even get the same label? They're looking at two really different periods of time!\n\nThe EEG time series overlap for two adjacent sub-samples is like 80-90%, but the spectrogram overlap for the same adjacent subsamples is going to be ~99%. Why bother subsampling at all?\n\nRelatedly, my impression is that neurologists primarily look at the spectrograms, not the raw EEG time series. Does anyone know if that's the case for the annotators here as well?",
      "votes": null
    },
    {
      "id": "2597994",
      "postDate": "01/12/2024 05:00:57",
      "content": "<p><a href=\"https://www.kaggle.com/keithcrum\" target=\"_blank\">@keithcrum</a> </p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window <strong>centered at the same time and labeled the central 10 seconds</strong>. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated</p>\n</blockquote>\n<p>EEG sub sequence focus part : central 10 secs<br>\nSpectorgram sub sequence focus part : centrail 10secs of EEG sub sequence's frequency around</p>",
      "rawMarkdown": "keithcrum \n> The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window **centered at the same time and labeled the central 10 seconds**. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated\n\nEEG sub sequence focus part : central 10 secs\nSpectorgram sub sequence focus part : centrail 10secs of EEG sub sequence's frequency around",
      "votes": null
    },
    {
      "id": "2599536",
      "postDate": "01/12/2024 23:10:39",
      "content": "<p>Can anyone please clarify? </p>\n<p>What does 'centered at the same time and labeled the central 10 seconds' imply? Does it mean that the midpoint of the 50-second EEG and the 10-minute spectrogram coincide in time? Example, if I examine the EEG data from 11 to 60 seconds, should the corresponding spectrogram be centered around the 35th second, within <em>5mins&lt;35thSecond&lt;5mins</em>? Also, if the minimum threshold time of the spectrogram window calculated by this method falls below zero, should it be adjusted to start at zero instead?</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> <a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a>  - I see you already performed some analysis, so marking you for attention.</p>",
      "rawMarkdown": "Can anyone please clarify? \n\nWhat does 'centered at the same time and labeled the central 10 seconds' imply? Does it mean that the midpoint of the 50-second EEG and the 10-minute spectrogram coincide in time? Example, if I examine the EEG data from 11 to 60 seconds, should the corresponding spectrogram be centered around the 35th second, within *5mins<35thSecond<5mins*? Also, if the minimum threshold time of the spectrogram window calculated by this method falls below zero, should it be adjusted to start at zero instead?\n\n@cdeotte @seshurajup @markwijkhuizen @yousof9  - I see you already performed some analysis, so marking you for attention.",
      "votes": null
    },
    {
      "id": "2600486",
      "postDate": "01/13/2024 17:11:04",
      "content": "<p>For each patient, we have a very long EEG record, but we are interested only in certain 'events', each of which correspond to a row in train dataset. For each event we do the following.<br>\n1) We take the time interval starting 25 seconds before the event and ending 25 seconds later. We sample it with 200 Hz sampling frequency and present it \"as is\" in eeg folder.<br>\n2) We take the time interval starting 5 minutes before the event and ending 5 minutes later. For this interval we calculate a spectrogram. Each sample in spectrogram can be considered as a power spectrum of a 2-second long slice of a signal. The <code>time</code> column denotes the center of each slice, that's why it starts with '1' and not '0'.</p>\n<p>So, we have two independent descriptions of each event. The short-term EEG signal with high sampling frequency makes it possible to study the fine details of events. The long-term spectrograms can be used to observe, what happened before and after the event.</p>\n<p>The time between two events can be small. So it is possible to save the disk space, if we provide the whole data and only mark the offsets of events. The spectrogram window sould never fall below zero.</p>\n<p>I also have prepared a couple of functions, which can be useful for understanding the data as well as looking at the records:<br>\n<a href=\"https://www.kaggle.com/kdmitrie/hms-simple-functions-to-explore-the-records\" target=\"_blank\">https://www.kaggle.com/kdmitrie/hms-simple-functions-to-explore-the-records</a></p>",
      "rawMarkdown": "For each patient, we have a very long EEG record, but we are interested only in certain 'events', each of which correspond to a row in train dataset. For each event we do the following.\n1) We take the time interval starting 25 seconds before the event and ending 25 seconds later. We sample it with 200 Hz sampling frequency and present it \"as is\" in eeg folder.\n2) We take the time interval starting 5 minutes before the event and ending 5 minutes later. For this interval we calculate a spectrogram. Each sample in spectrogram can be considered as a power spectrum of a 2-second long slice of a signal. The `time` column denotes the center of each slice, that's why it starts with '1' and not '0'.\n\nSo, we have two independent descriptions of each event. The short-term EEG signal with high sampling frequency makes it possible to study the fine details of events. The long-term spectrograms can be used to observe, what happened before and after the event.\n\nThe time between two events can be small. So it is possible to save the disk space, if we provide the whole data and only mark the offsets of events. The spectrogram window sould never fall below zero.\n\nI also have prepared a couple of functions, which can be useful for understanding the data as well as looking at the records:\nhttps://www.kaggle.com/kdmitrie/hms-simple-functions-to-explore-the-records",
      "votes": null
    },
    {
      "id": "2601008",
      "postDate": "01/14/2024 04:07:06",
      "content": "<p><a href=\"https://www.kaggle.com/kdmitrie\" target=\"_blank\">@kdmitrie</a> Wonderful! Thanks a lot for the functions. This provides a clear understanding of how the data is derived.</p>\n<p>To ensure a clear understanding of how the data is derived, let’s consider an example below with the EEG segment. For instance, suppose the 4th EEG segment (sub_id = 3) is recorded between 1:00:00 PM and 1:00:50 PM. In relation to this, the corresponding 4th spectrogram data is generated from the EEG recorded during a broader timeframe, specifically from 12:55:25 PM to 1:05:25 PM. Am I getting this right?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745307%2Fc110094877c9f11a7f1805ff267a9075%2FScreenshot%202024-01-13%20at%2010.51.59PM.png?generation=1705204329670409&amp;alt=media\"></p>",
      "rawMarkdown": "kdmitrie Wonderful! Thanks a lot for the functions. This provides a clear understanding of how the data is derived.\n\nTo ensure a clear understanding of how the data is derived, let’s consider an example below with the EEG segment. For instance, suppose the 4th EEG segment (sub_id = 3) is recorded between 1:00:00 PM and 1:00:50 PM. In relation to this, the corresponding 4th spectrogram data is generated from the EEG recorded during a broader timeframe, specifically from 12:55:25 PM to 1:05:25 PM. Am I getting this right?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745307%2Fc110094877c9f11a7f1805ff267a9075%2FScreenshot%202024-01-13%20at%2010.51.59PM.png?generation=1705204329670409&alt=media)",
      "votes": null
    },
    {
      "id": "2601207",
      "postDate": "01/14/2024 08:39:23",
      "content": "<p>Yes, I think, you are right.</p>",
      "rawMarkdown": "Yes, I think, you are right.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2597958,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/12/2024 04:05:38",
      "content": "<p>The spectrograms are 10 minutes long. So you want <code>time &gt;= 18</code> and <code>time &lt; 18+600</code>. This will be 300 rows because the spectrogram parquets count by 2s. This is the same size as the test spectrograms which are all 300 rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2597989,
          "author_name": "keithcrum",
          "author_url": "",
          "post_date": "01/12/2024 04:57:38",
          "content": "<p>Wow, that's really surprising to me. Why would a 50 second EEG portion and its corresponding 10 minute spectrogram even get the same label? They're looking at two really different periods of time!</p>\n<p>The EEG time series overlap for two adjacent sub-samples is like 80-90%, but the spectrogram overlap for the same adjacent subsamples is going to be ~99%. Why bother subsampling at all?</p>\n<p>Relatedly, my impression is that neurologists primarily look at the spectrograms, not the raw EEG time series. Does anyone know if that's the case for the annotators here as well?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2597994,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "01/12/2024 05:00:57",
              "content": "<p><a href=\"https://www.kaggle.com/keithcrum\" target=\"_blank\">@keithcrum</a> </p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window <strong>centered at the same time and labeled the central 10 seconds</strong>. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated</p>\n</blockquote>\n<p>EEG sub sequence focus part : central 10 secs<br>\nSpectorgram sub sequence focus part : centrail 10secs of EEG sub sequence's frequency around</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2599536,
      "author_name": "narareddy",
      "author_url": "",
      "post_date": "01/12/2024 23:10:39",
      "content": "<p>Can anyone please clarify? </p>\n<p>What does 'centered at the same time and labeled the central 10 seconds' imply? Does it mean that the midpoint of the 50-second EEG and the 10-minute spectrogram coincide in time? Example, if I examine the EEG data from 11 to 60 seconds, should the corresponding spectrogram be centered around the 35th second, within <em>5mins&lt;35thSecond&lt;5mins</em>? Also, if the minimum threshold time of the spectrogram window calculated by this method falls below zero, should it be adjusted to start at zero instead?</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> <a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">@yousof9</a>  - I see you already performed some analysis, so marking you for attention.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2600486,
          "author_name": "kdmitrie",
          "author_url": "",
          "post_date": "01/13/2024 17:11:04",
          "content": "<p>For each patient, we have a very long EEG record, but we are interested only in certain 'events', each of which correspond to a row in train dataset. For each event we do the following.<br>\n1) We take the time interval starting 25 seconds before the event and ending 25 seconds later. We sample it with 200 Hz sampling frequency and present it \"as is\" in eeg folder.<br>\n2) We take the time interval starting 5 minutes before the event and ending 5 minutes later. For this interval we calculate a spectrogram. Each sample in spectrogram can be considered as a power spectrum of a 2-second long slice of a signal. The <code>time</code> column denotes the center of each slice, that's why it starts with '1' and not '0'.</p>\n<p>So, we have two independent descriptions of each event. The short-term EEG signal with high sampling frequency makes it possible to study the fine details of events. The long-term spectrograms can be used to observe, what happened before and after the event.</p>\n<p>The time between two events can be small. So it is possible to save the disk space, if we provide the whole data and only mark the offsets of events. The spectrogram window sould never fall below zero.</p>\n<p>I also have prepared a couple of functions, which can be useful for understanding the data as well as looking at the records:<br>\n<a href=\"https://www.kaggle.com/kdmitrie/hms-simple-functions-to-explore-the-records\" target=\"_blank\">https://www.kaggle.com/kdmitrie/hms-simple-functions-to-explore-the-records</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2601008,
              "author_name": "narareddy",
              "author_url": "",
              "post_date": "01/14/2024 04:07:06",
              "content": "<p><a href=\"https://www.kaggle.com/kdmitrie\" target=\"_blank\">@kdmitrie</a> Wonderful! Thanks a lot for the functions. This provides a clear understanding of how the data is derived.</p>\n<p>To ensure a clear understanding of how the data is derived, let’s consider an example below with the EEG segment. For instance, suppose the 4th EEG segment (sub_id = 3) is recorded between 1:00:00 PM and 1:00:50 PM. In relation to this, the corresponding 4th spectrogram data is generated from the EEG recorded during a broader timeframe, specifically from 12:55:25 PM to 1:05:25 PM. Am I getting this right?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745307%2Fc110094877c9f11a7f1805ff267a9075%2FScreenshot%202024-01-13%20at%2010.51.59PM.png?generation=1705204329670409&amp;alt=media\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2601207,
                  "author_name": "kdmitrie",
                  "author_url": "",
                  "post_date": "01/14/2024 08:39:23",
                  "content": "<p>Yes, I think, you are right.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2597934": "I'm struggling to understand how the spectrogram and EEG samples are related. The relevant description in the \"Data\" tab seems to have typos:\n\n>Files\ntrain.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated.\n\nWhat does this mean, especially for the spectrogram part? I am a native English speaker, but I cannot parse this. \n\nTo take a concrete example, let's look at the 3rd row from train.csv:\n```\nIn [33]: mdf.iloc[3, :]\nOut[33]: \neeg_id                              1628180742\neeg_sub_id                                   3\neeg_label_offset_seconds                  18.0\nspectrogram_id                          353733\nspectrogram_sub_id                           3\nspectrogram_label_offset_seconds          18.0\nlabel_id                            2718991173\npatient_id                               42516\nexpert_consensus                       Seizure\nseizure_vote                                 3\nlpd_vote                                     0\ngpd_vote                                     0\nlrda_vote                                    0\ngrda_vote                                    0\nother_vote                                   0\n```\n\nWhen I open up the EEG, I suspect I should keep all rows from 18 x 200 to (18 + 50) x 200 (perhaps excluding the last row). Is that correct?\n\nWhen I open up the spectrogram file, should I take all rows where time >= 18 and time < 68 (i.e. the same 50 seconds as the EEG)? \n\nThanks for any help!",
    "2597958": "The spectrograms are 10 minutes long. So you want `time >= 18` and `time < 18+600`. This will be 300 rows because the spectrogram parquets count by 2s. This is the same size as the test spectrograms which are all 300 rows.",
    "2597989": "Wow, that's really surprising to me. Why would a 50 second EEG portion and its corresponding 10 minute spectrogram even get the same label? They're looking at two really different periods of time!\n\nThe EEG time series overlap for two adjacent sub-samples is like 80-90%, but the spectrogram overlap for the same adjacent subsamples is going to be ~99%. Why bother subsampling at all?\n\nRelatedly, my impression is that neurologists primarily look at the spectrograms, not the raw EEG time series. Does anyone know if that's the case for the annotators here as well?",
    "2597994": "keithcrum \n> The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window **centered at the same time and labeled the central 10 seconds**. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated\n\nEEG sub sequence focus part : central 10 secs\nSpectorgram sub sequence focus part : centrail 10secs of EEG sub sequence's frequency around",
    "2599536": "Can anyone please clarify? \n\nWhat does 'centered at the same time and labeled the central 10 seconds' imply? Does it mean that the midpoint of the 50-second EEG and the 10-minute spectrogram coincide in time? Example, if I examine the EEG data from 11 to 60 seconds, should the corresponding spectrogram be centered around the 35th second, within *5mins<35thSecond<5mins*? Also, if the minimum threshold time of the spectrogram window calculated by this method falls below zero, should it be adjusted to start at zero instead?\n\n@cdeotte @seshurajup @markwijkhuizen @yousof9  - I see you already performed some analysis, so marking you for attention.",
    "2600486": "For each patient, we have a very long EEG record, but we are interested only in certain 'events', each of which correspond to a row in train dataset. For each event we do the following.\n1) We take the time interval starting 25 seconds before the event and ending 25 seconds later. We sample it with 200 Hz sampling frequency and present it \"as is\" in eeg folder.\n2) We take the time interval starting 5 minutes before the event and ending 5 minutes later. For this interval we calculate a spectrogram. Each sample in spectrogram can be considered as a power spectrum of a 2-second long slice of a signal. The `time` column denotes the center of each slice, that's why it starts with '1' and not '0'.\n\nSo, we have two independent descriptions of each event. The short-term EEG signal with high sampling frequency makes it possible to study the fine details of events. The long-term spectrograms can be used to observe, what happened before and after the event.\n\nThe time between two events can be small. So it is possible to save the disk space, if we provide the whole data and only mark the offsets of events. The spectrogram window sould never fall below zero.\n\nI also have prepared a couple of functions, which can be useful for understanding the data as well as looking at the records:\nhttps://www.kaggle.com/kdmitrie/hms-simple-functions-to-explore-the-records",
    "2601008": "kdmitrie Wonderful! Thanks a lot for the functions. This provides a clear understanding of how the data is derived.\n\nTo ensure a clear understanding of how the data is derived, let’s consider an example below with the EEG segment. For instance, suppose the 4th EEG segment (sub_id = 3) is recorded between 1:00:00 PM and 1:00:50 PM. In relation to this, the corresponding 4th spectrogram data is generated from the EEG recorded during a broader timeframe, specifically from 12:55:25 PM to 1:05:25 PM. Am I getting this right?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745307%2Fc110094877c9f11a7f1805ff267a9075%2FScreenshot%202024-01-13%20at%2010.51.59PM.png?generation=1705204329670409&alt=media)",
    "2601207": "Yes, I think, you are right."
  },
  "source": "meta"
}