{
  "id": 466759,
  "title": "Spectrogram Data Acquisition Rate",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/466759",
  "author_name": "",
  "post_date": "2024-01-09T21:48:44.054226700Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>From exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length.</p>\n<p>In the Data page, information about the spectrograms refers to \"spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\", and later \"Spectrograms assembled using exactly 10 minutes of EEG data\".</p>\n<p>From exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length. The sample lengths range from 300 to 9116 samples, with most spectrograms being close to 300 samples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369640%2Ff98484839623a737f36f1579821d08bd%2FScreenshot%202024-01-09%20214620.png?generation=1704836849249121&amp;alt=media\" alt=\"\"></p>\n<p>Should these spectrograms be the same length? If we need to extract the 10 minute window, can we have the information about the sampling frequency so we know how many samples this is? Is it ok to assume that 300 samples corresponds to 10 minutes (aka 0.5 Hz sample rate)?</p>",
  "messages": [
    {
      "id": "2594434",
      "postDate": "01/09/2024 21:48:44",
      "content": "<p>From exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length.</p>\n<p>In the Data page, information about the spectrograms refers to \"spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\", and later \"Spectrograms assembled using exactly 10 minutes of EEG data\".</p>\n<p>From exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length. The sample lengths range from 300 to 9116 samples, with most spectrograms being close to 300 samples.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369640%2Ff98484839623a737f36f1579821d08bd%2FScreenshot%202024-01-09%20214620.png?generation=1704836849249121&amp;alt=media\" alt=\"\"></p>\n<p>Should these spectrograms be the same length? If we need to extract the 10 minute window, can we have the information about the sampling frequency so we know how many samples this is? Is it ok to assume that 300 samples corresponds to 10 minutes (aka 0.5 Hz sample rate)?</p>",
      "rawMarkdown": "From exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length.\n\nIn the Data page, information about the spectrograms refers to \"spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\", and later \"Spectrograms assembled using exactly 10 minutes of EEG data\".\n\nFrom exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length. The sample lengths range from 300 to 9116 samples, with most spectrograms being close to 300 samples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369640%2Ff98484839623a737f36f1579821d08bd%2FScreenshot%202024-01-09%20214620.png?generation=1704836849249121&alt=media)\n\nShould these spectrograms be the same length? If we need to extract the 10 minute window, can we have the information about the sampling frequency so we know how many samples this is? Is it ok to assume that 300 samples corresponds to 10 minutes (aka 0.5 Hz sample rate)?",
      "votes": null
    },
    {
      "id": "2595713",
      "postDate": "01/10/2024 15:44:11",
      "content": "<p>I am not 100% sure because the train.csv file is empty at the moment but I believe \"spectogram_label_offset_seconds\" can be used here.  Higher frequency data acquisition will mean more samples and that column lets you sort of calculate how far along the 10 minute window the sample was taken (because I guess we cannot assume a constant frequency throughout).  </p>",
      "rawMarkdown": "I am not 100% sure because the train.csv file is empty at the moment but I believe \"spectogram_label_offset_seconds\" can be used here.  Higher frequency data acquisition will mean more samples and that column lets you sort of calculate how far along the 10 minute window the sample was taken (because I guess we cannot assume a constant frequency throughout).",
      "votes": null
    },
    {
      "id": "2597908",
      "postDate": "01/12/2024 02:41:56",
      "content": "<p><a href=\"https://www.kaggle.com/connorjd\" target=\"_blank\">@connorjd</a> train.csv file empty issue is resolved - <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002</a></p>",
      "rawMarkdown": "connorjd train.csv file empty issue is resolved - [https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002)",
      "votes": null
    },
    {
      "id": "2597911",
      "postDate": "01/12/2024 02:48:12",
      "content": "<p>For each row in <code>train.csv</code>, you need to extract the corresponding 300 consecutive rows from the corresponding spectrogram parquet. Here is the code to get each spectrogram for each row of <code>train.csv</code>:</p>\n<pre><code>train = pd.read_csv()\n = train.iloc[]\nspectrogram = pd.read_parquet(+str(.spectrogram_id)+)\nstart = (.spectrogram_label_offset_seconds)\n %==:  += \nend =  + \nspectrogram = spectrogram.loc[(spectrogram.time&gt;=)&amp;(spectrogram.time&lt;=)]\n</code></pre>",
      "rawMarkdown": "For each row in `train.csv`, you need to extract the corresponding 300 consecutive rows from the corresponding spectrogram parquet. Here is the code to get each spectrogram for each row of `train.csv`:\n\n    train = pd.read_csv('train.csv')\n    row = train.iloc[ROW]\n    spectrogram = pd.read_parquet(PATH+str(row.spectrogram_id)+'.parquet')\n    start = int(row.spectrogram_label_offset_seconds)\n    if start%2==0: start += 1\n    end = start + 598\n    spectrogram = spectrogram.loc[(spectrogram.time>=start)&(spectrogram.time<=end)]",
      "votes": null
    },
    {
      "id": "2600411",
      "postDate": "01/13/2024 15:41:51",
      "content": "<p>Hi! Can you explain, what are magical \"+1 if start is even\" and 598? Am I right that it means that 300 rows is equal to 50 seconds of eeg (or 50 * 200 = 10k samples in the EEG dataframe) so the sample rate of each spectrogram is 6 samples per second (so 50 * 6 = 300 rows you mentioned). </p>\n<p>I got same numbers when checked on the test data (50 seconds, 10k samples of eeg, 300 rows of spectrograms), but on the train data it seems to be different. For example if I select first pair (eeg_id, spectrogram_id), which is (1628180742, 353733), then eeg have 18000 samples, which corresponds to the 90 seconds, and spectrogram have only 320 rows, or about 53 seconds (if the spectrogram sample rate is also 6). So either the spectrogram is much shorter than eeg in this example, or it have different sample rate (3.55 in this case). In first case it means that not all the 50sec subsamples have corresponding spectrogram and the second means that the resolution of the spectrograms is extremely low.</p>\n<p>Can't understand how to align specs with eegs, maybe you can help with it?</p>",
      "rawMarkdown": "Hi! Can you explain, what are magical \"+1 if start is even\" and 598? Am I right that it means that 300 rows is equal to 50 seconds of eeg (or 50 * 200 = 10k samples in the EEG dataframe) so the sample rate of each spectrogram is 6 samples per second (so 50 * 6 = 300 rows you mentioned). \n\nI got same numbers when checked on the test data (50 seconds, 10k samples of eeg, 300 rows of spectrograms), but on the train data it seems to be different. For example if I select first pair (eeg_id, spectrogram_id), which is (1628180742, 353733), then eeg have 18000 samples, which corresponds to the 90 seconds, and spectrogram have only 320 rows, or about 53 seconds (if the spectrogram sample rate is also 6). So either the spectrogram is much shorter than eeg in this example, or it have different sample rate (3.55 in this case). In first case it means that not all the 50sec subsamples have corresponding spectrogram and the second means that the resolution of the spectrograms is extremely low.\n \nCan't understand how to align specs with eegs, maybe you can help with it?",
      "votes": null
    },
    {
      "id": "2600472",
      "postDate": "01/13/2024 16:43:01",
      "content": "<p>As far, as I understood,<br>\n1) Each spectrogram is 10 minutes long, i.e., 600 sec. However, according to the <code>time</code> column, the sampling rate is 0.5 Hz. What does it mean? For example, we can take the raw 10 minutes long record, split it into 300 slices with the length of 2 sec each, and perform Fourier transform. Then we cut off everything above 20 Hz, and finally we get the spectrogram. </p>\n<p>2) Time column correspond to the center of each slice. For example, <code>time</code>==1 means the slice starting at 0 sec and ending at 2 sec. If we want the spectrogram with, for example, <code>spectrogram_label_offset_seconds</code>==6, then the first slice of it corresponds to time interval [6; 8], and its center is <code>time</code>==7. The last sample corresponds to time interval [606; 608], and its center is <code>time</code>==607. </p>\n<p>So, we are interesting in (offset + 1)&lt;= time &lt;= (offset + 599). This is equal to int(offset/2)&lt;= row_index &lt;= int(offset/2) + 300</p>\n<p>3) There are several records with odd spectrogram_label_offset_seconds, which correspond to 9 patients:</p>\n<pre><code>df_train()\n&gt;&gt;&gt; ()\n</code></pre>\n<p>I don't know, how to handle this correctly. But, perhaps, ±1 sec doesn't matter, since we are looking at least at 10 sec interval, and we can still use the same formula.</p>\n<p>Here is a simple function to extract all needed data:</p>\n<pre><code>TRAIN = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'\nSG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_spectrograms/%i.parquet'\nEEG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_eegs/%i.parquet'\nEEG_FS = \nEEG_LEN = EEG_FS\nSG_LEN = \n\ndef get:\n    item = df_train.iloc\n    eeg_start = (EEG_FSitem.eeg_label_offset_seconds)\n    eeg = pd.read).iloc.reset\n    sg_start = (item.spectrogram_label_offset_seconds)\n    sg = pd.read).iloc.reset\n    return sg, eeg, item.iloc\n</code></pre>",
      "rawMarkdown": "As far, as I understood,\n1) Each spectrogram is 10 minutes long, i.e., 600 sec. However, according to the `time` column, the sampling rate is 0.5 Hz. What does it mean? For example, we can take the raw 10 minutes long record, split it into 300 slices with the length of 2 sec each, and perform Fourier transform. Then we cut off everything above 20 Hz, and finally we get the spectrogram. \n\n2) Time column correspond to the center of each slice. For example, `time`==1 means the slice starting at 0 sec and ending at 2 sec. If we want the spectrogram with, for example, `spectrogram_label_offset_seconds`==6, then the first slice of it corresponds to time interval [6; 8], and its center is `time`==7. The last sample corresponds to time interval [606; 608], and its center is `time`==607. \n\nSo, we are interesting in (offset + 1)<= time <= (offset + 599). This is equal to int(offset/2)<= row_index <= int(offset/2) + 300\n\n3) There are several records with odd spectrogram_label_offset_seconds, which correspond to 9 patients:\n```\ndf_train[df_train.spectrogram_label_offset_seconds % 2 == 1].patient_id.unique()\n>>> array([56450, 64704, 35627, 29441, 12251, 17408, 36297, 21260, 17089])\n\n```\nI don't know, how to handle this correctly. But, perhaps, ±1 sec doesn't matter, since we are looking at least at 10 sec interval, and we can still use the same formula.\n\nHere is a simple function to extract all needed data:\n```\nTRAIN = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'\nSG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_spectrograms/%i.parquet'\nEEG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_eegs/%i.parquet'\nEEG_FS = 200\nEEG_LEN = EEG_FS * 50\nSG_LEN = 300\n\ndef get_train_item(item_id):\n    item = df_train.iloc[item_id]\n    eeg_start = int(EEG_FS * item.eeg_label_offset_seconds)\n    eeg = pd.read_parquet(EEG % ('train', item.eeg_id)).iloc[eeg_start:eeg_start + EEG_LEN].reset_index()\n    sg_start = int(item.spectrogram_label_offset_seconds / 2)\n    sg = pd.read_parquet(SG % ('train', item.spectrogram_id)).iloc[sg_start:sg_start + SG_LEN, 1:].reset_index()\n    return sg, eeg, item.iloc[-6:]\n```",
      "votes": null
    },
    {
      "id": "2600487",
      "postDate": "01/13/2024 17:11:20",
      "content": "<p>Does anyone know what signal is used to create the spectrograms? The eeg has 20 electrodes. How are these signals combined to create 4 signals to make the 4 spectrograms LL RL LP RP?</p>",
      "rawMarkdown": "Does anyone know what signal is used to create the spectrograms? The eeg has 20 electrodes. How are these signals combined to create 4 signals to make the 4 spectrograms LL RL LP RP?",
      "votes": null
    },
    {
      "id": "2600499",
      "postDate": "01/13/2024 17:23:23",
      "content": "<p>I suppose, this is some kind of a linear combination of electrodes. It's hard to know exactly, but I found such reference with an image:<br>\n<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6776236/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6776236/</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Ff790d9fba9d01f39b5c688819cc90f64%2Ftileshop.jpg?generation=1705166531784497&amp;alt=media\"></p>",
      "rawMarkdown": "I suppose, this is some kind of a linear combination of electrodes. It's hard to know exactly, but I found such reference with an image:\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6776236/\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Ff790d9fba9d01f39b5c688819cc90f64%2Ftileshop.jpg?generation=1705166531784497&alt=media)",
      "votes": null
    },
    {
      "id": "2601461",
      "postDate": "01/14/2024 12:41:49",
      "content": "<p>Thanks for the example. <br>\nLooks like your assumption is correct. For example LL spectrogram -&gt; SomeFuntion(Fp1, F7, T3, T5, O1)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F384f80c6972217b03108598b072fbccf%2F.png?generation=1705237135204580&amp;alt=media\" alt=\"Example\"></p>\n<p>I'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the eeg has 90 seconds. It's supposed to be the same lenght. </p>\n<p>For example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz</p>",
      "rawMarkdown": "Thanks for the example. \nLooks like your assumption is correct. For example LL spectrogram -> SomeFuntion(Fp1, F7, T3, T5, O1)\n\n![Example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F384f80c6972217b03108598b072fbccf%2F.png?generation=1705237135204580&alt=media)\n\nI'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the eeg has 90 seconds. It's supposed to be the same lenght. \n\nFor example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz",
      "votes": null
    },
    {
      "id": "2601634",
      "postDate": "01/14/2024 15:15:18",
      "content": "<p>I think, I have covered this below: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/466759#2600472\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/466759#2600472</a></p>\n<p>The reason is we have a long record of a signal, from which we select only a part. This approach intended to save the space on the disk or in memory.</p>",
      "rawMarkdown": "I think, I have covered this below: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/466759#2600472\n\nThe reason is we have a long record of a signal, from which we select only a part. This approach intended to save the space on the disk or in memory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2595713,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "01/10/2024 15:44:11",
      "content": "<p>I am not 100% sure because the train.csv file is empty at the moment but I believe \"spectogram_label_offset_seconds\" can be used here.  Higher frequency data acquisition will mean more samples and that column lets you sort of calculate how far along the 10 minute window the sample was taken (because I guess we cannot assume a constant frequency throughout).  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2597908,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/12/2024 02:41:56",
          "content": "<p><a href=\"https://www.kaggle.com/connorjd\" target=\"_blank\">@connorjd</a> train.csv file empty issue is resolved - <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2597911,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/12/2024 02:48:12",
      "content": "<p>For each row in <code>train.csv</code>, you need to extract the corresponding 300 consecutive rows from the corresponding spectrogram parquet. Here is the code to get each spectrogram for each row of <code>train.csv</code>:</p>\n<pre><code>train = pd.read_csv()\n = train.iloc[]\nspectrogram = pd.read_parquet(+str(.spectrogram_id)+)\nstart = (.spectrogram_label_offset_seconds)\n %==:  += \nend =  + \nspectrogram = spectrogram.loc[(spectrogram.time&gt;=)&amp;(spectrogram.time&lt;=)]\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2600411,
          "author_name": "kst179",
          "author_url": "",
          "post_date": "01/13/2024 15:41:51",
          "content": "<p>Hi! Can you explain, what are magical \"+1 if start is even\" and 598? Am I right that it means that 300 rows is equal to 50 seconds of eeg (or 50 * 200 = 10k samples in the EEG dataframe) so the sample rate of each spectrogram is 6 samples per second (so 50 * 6 = 300 rows you mentioned). </p>\n<p>I got same numbers when checked on the test data (50 seconds, 10k samples of eeg, 300 rows of spectrograms), but on the train data it seems to be different. For example if I select first pair (eeg_id, spectrogram_id), which is (1628180742, 353733), then eeg have 18000 samples, which corresponds to the 90 seconds, and spectrogram have only 320 rows, or about 53 seconds (if the spectrogram sample rate is also 6). So either the spectrogram is much shorter than eeg in this example, or it have different sample rate (3.55 in this case). In first case it means that not all the 50sec subsamples have corresponding spectrogram and the second means that the resolution of the spectrograms is extremely low.</p>\n<p>Can't understand how to align specs with eegs, maybe you can help with it?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2600472,
              "author_name": "kdmitrie",
              "author_url": "",
              "post_date": "01/13/2024 16:43:01",
              "content": "<p>As far, as I understood,<br>\n1) Each spectrogram is 10 minutes long, i.e., 600 sec. However, according to the <code>time</code> column, the sampling rate is 0.5 Hz. What does it mean? For example, we can take the raw 10 minutes long record, split it into 300 slices with the length of 2 sec each, and perform Fourier transform. Then we cut off everything above 20 Hz, and finally we get the spectrogram. </p>\n<p>2) Time column correspond to the center of each slice. For example, <code>time</code>==1 means the slice starting at 0 sec and ending at 2 sec. If we want the spectrogram with, for example, <code>spectrogram_label_offset_seconds</code>==6, then the first slice of it corresponds to time interval [6; 8], and its center is <code>time</code>==7. The last sample corresponds to time interval [606; 608], and its center is <code>time</code>==607. </p>\n<p>So, we are interesting in (offset + 1)&lt;= time &lt;= (offset + 599). This is equal to int(offset/2)&lt;= row_index &lt;= int(offset/2) + 300</p>\n<p>3) There are several records with odd spectrogram_label_offset_seconds, which correspond to 9 patients:</p>\n<pre><code>df_train()\n&gt;&gt;&gt; ()\n</code></pre>\n<p>I don't know, how to handle this correctly. But, perhaps, ±1 sec doesn't matter, since we are looking at least at 10 sec interval, and we can still use the same formula.</p>\n<p>Here is a simple function to extract all needed data:</p>\n<pre><code>TRAIN = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'\nSG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_spectrograms/%i.parquet'\nEEG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_eegs/%i.parquet'\nEEG_FS = \nEEG_LEN = EEG_FS\nSG_LEN = \n\ndef get:\n    item = df_train.iloc\n    eeg_start = (EEG_FSitem.eeg_label_offset_seconds)\n    eeg = pd.read).iloc.reset\n    sg_start = (item.spectrogram_label_offset_seconds)\n    sg = pd.read).iloc.reset\n    return sg, eeg, item.iloc\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2600487,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/13/2024 17:11:20",
      "content": "<p>Does anyone know what signal is used to create the spectrograms? The eeg has 20 electrodes. How are these signals combined to create 4 signals to make the 4 spectrograms LL RL LP RP?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2600499,
          "author_name": "kdmitrie",
          "author_url": "",
          "post_date": "01/13/2024 17:23:23",
          "content": "<p>I suppose, this is some kind of a linear combination of electrodes. It's hard to know exactly, but I found such reference with an image:<br>\n<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6776236/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6776236/</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Ff790d9fba9d01f39b5c688819cc90f64%2Ftileshop.jpg?generation=1705166531784497&amp;alt=media\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2601461,
              "author_name": "evgeny000",
              "author_url": "",
              "post_date": "01/14/2024 12:41:49",
              "content": "<p>Thanks for the example. <br>\nLooks like your assumption is correct. For example LL spectrogram -&gt; SomeFuntion(Fp1, F7, T3, T5, O1)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F384f80c6972217b03108598b072fbccf%2F.png?generation=1705237135204580&amp;alt=media\" alt=\"Example\"></p>\n<p>I'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the eeg has 90 seconds. It's supposed to be the same lenght. </p>\n<p>For example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2601634,
                  "author_name": "kdmitrie",
                  "author_url": "",
                  "post_date": "01/14/2024 15:15:18",
                  "content": "<p>I think, I have covered this below: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/466759#2600472\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/466759#2600472</a></p>\n<p>The reason is we have a long record of a signal, from which we select only a part. This approach intended to save the space on the disk or in memory.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2594434": "From exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length.\n\nIn the Data page, information about the spectrograms refers to \"spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\", and later \"Spectrograms assembled using exactly 10 minutes of EEG data\".\n\nFrom exploring the spectrogram data, I noticed that each of the spectrograms is a slightly different length. The sample lengths range from 300 to 9116 samples, with most spectrograms being close to 300 samples.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2369640%2Ff98484839623a737f36f1579821d08bd%2FScreenshot%202024-01-09%20214620.png?generation=1704836849249121&alt=media)\n\nShould these spectrograms be the same length? If we need to extract the 10 minute window, can we have the information about the sampling frequency so we know how many samples this is? Is it ok to assume that 300 samples corresponds to 10 minutes (aka 0.5 Hz sample rate)?",
    "2595713": "I am not 100% sure because the train.csv file is empty at the moment but I believe \"spectogram_label_offset_seconds\" can be used here.  Higher frequency data acquisition will mean more samples and that column lets you sort of calculate how far along the 10 minute window the sample was taken (because I guess we cannot assume a constant frequency throughout).",
    "2597908": "connorjd train.csv file empty issue is resolved - [https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467002)",
    "2597911": "For each row in `train.csv`, you need to extract the corresponding 300 consecutive rows from the corresponding spectrogram parquet. Here is the code to get each spectrogram for each row of `train.csv`:\n\n    train = pd.read_csv('train.csv')\n    row = train.iloc[ROW]\n    spectrogram = pd.read_parquet(PATH+str(row.spectrogram_id)+'.parquet')\n    start = int(row.spectrogram_label_offset_seconds)\n    if start%2==0: start += 1\n    end = start + 598\n    spectrogram = spectrogram.loc[(spectrogram.time>=start)&(spectrogram.time<=end)]",
    "2600411": "Hi! Can you explain, what are magical \"+1 if start is even\" and 598? Am I right that it means that 300 rows is equal to 50 seconds of eeg (or 50 * 200 = 10k samples in the EEG dataframe) so the sample rate of each spectrogram is 6 samples per second (so 50 * 6 = 300 rows you mentioned). \n\nI got same numbers when checked on the test data (50 seconds, 10k samples of eeg, 300 rows of spectrograms), but on the train data it seems to be different. For example if I select first pair (eeg_id, spectrogram_id), which is (1628180742, 353733), then eeg have 18000 samples, which corresponds to the 90 seconds, and spectrogram have only 320 rows, or about 53 seconds (if the spectrogram sample rate is also 6). So either the spectrogram is much shorter than eeg in this example, or it have different sample rate (3.55 in this case). In first case it means that not all the 50sec subsamples have corresponding spectrogram and the second means that the resolution of the spectrograms is extremely low.\n \nCan't understand how to align specs with eegs, maybe you can help with it?",
    "2600472": "As far, as I understood,\n1) Each spectrogram is 10 minutes long, i.e., 600 sec. However, according to the `time` column, the sampling rate is 0.5 Hz. What does it mean? For example, we can take the raw 10 minutes long record, split it into 300 slices with the length of 2 sec each, and perform Fourier transform. Then we cut off everything above 20 Hz, and finally we get the spectrogram. \n\n2) Time column correspond to the center of each slice. For example, `time`==1 means the slice starting at 0 sec and ending at 2 sec. If we want the spectrogram with, for example, `spectrogram_label_offset_seconds`==6, then the first slice of it corresponds to time interval [6; 8], and its center is `time`==7. The last sample corresponds to time interval [606; 608], and its center is `time`==607. \n\nSo, we are interesting in (offset + 1)<= time <= (offset + 599). This is equal to int(offset/2)<= row_index <= int(offset/2) + 300\n\n3) There are several records with odd spectrogram_label_offset_seconds, which correspond to 9 patients:\n```\ndf_train[df_train.spectrogram_label_offset_seconds % 2 == 1].patient_id.unique()\n>>> array([56450, 64704, 35627, 29441, 12251, 17408, 36297, 21260, 17089])\n\n```\nI don't know, how to handle this correctly. But, perhaps, ±1 sec doesn't matter, since we are looking at least at 10 sec interval, and we can still use the same formula.\n\nHere is a simple function to extract all needed data:\n```\nTRAIN = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'\nSG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_spectrograms/%i.parquet'\nEEG = '/kaggle/input/hms-harmful-brain-activity-classification/%s_eegs/%i.parquet'\nEEG_FS = 200\nEEG_LEN = EEG_FS * 50\nSG_LEN = 300\n\ndef get_train_item(item_id):\n    item = df_train.iloc[item_id]\n    eeg_start = int(EEG_FS * item.eeg_label_offset_seconds)\n    eeg = pd.read_parquet(EEG % ('train', item.eeg_id)).iloc[eeg_start:eeg_start + EEG_LEN].reset_index()\n    sg_start = int(item.spectrogram_label_offset_seconds / 2)\n    sg = pd.read_parquet(SG % ('train', item.spectrogram_id)).iloc[sg_start:sg_start + SG_LEN, 1:].reset_index()\n    return sg, eeg, item.iloc[-6:]\n```",
    "2600487": "Does anyone know what signal is used to create the spectrograms? The eeg has 20 electrodes. How are these signals combined to create 4 signals to make the 4 spectrograms LL RL LP RP?",
    "2600499": "I suppose, this is some kind of a linear combination of electrodes. It's hard to know exactly, but I found such reference with an image:\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6776236/\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Ff790d9fba9d01f39b5c688819cc90f64%2Ftileshop.jpg?generation=1705166531784497&alt=media)",
    "2601461": "Thanks for the example. \nLooks like your assumption is correct. For example LL spectrogram -> SomeFuntion(Fp1, F7, T3, T5, O1)\n\n![Example](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F384f80c6972217b03108598b072fbccf%2F.png?generation=1705237135204580&alt=media)\n\nI'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the eeg has 90 seconds. It's supposed to be the same lenght. \n\nFor example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz",
    "2601634": "I think, I have covered this below: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/466759#2600472\n\nThe reason is we have a long record of a signal, from which we select only a part. This approach intended to save the space on the disk or in memory."
  },
  "source": "meta"
}