{
  "id": 475599,
  "title": "how to locate the correct spectrogram sample using the offset.",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/475599",
  "author_name": "",
  "post_date": "2024-02-09T05:49:16.387457100Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have just started on the competition, and i run into this problem, where I m not sure how the spectrogram offset and the EEG offset translate to the exact location in the dataframe. as far as I know the spectrogram are of the shape 300<em>400 . and eeg should be of the shape 20</em>10000. but the offset are often larger than these numbers. so my question is, how is the offset calculated and how do we use them? thanks.</p>",
  "messages": [
    {
      "id": "2643810",
      "postDate": "02/09/2024 05:49:16",
      "content": "<p>I have just started on the competition, and i run into this problem, where I m not sure how the spectrogram offset and the EEG offset translate to the exact location in the dataframe. as far as I know the spectrogram are of the shape 300<em>400 . and eeg should be of the shape 20</em>10000. but the offset are often larger than these numbers. so my question is, how is the offset calculated and how do we use them? thanks.</p>",
      "rawMarkdown": "I have just started on the competition, and i run into this problem, where I m not sure how the spectrogram offset and the EEG offset translate to the exact location in the dataframe. as far as I know the spectrogram are of the shape 300*400 . and eeg should be of the shape 20*10000. but the offset are often larger than these numbers. so my question is, how is the offset calculated and how do we use them? thanks.",
      "votes": null
    },
    {
      "id": "2644383",
      "postDate": "02/09/2024 12:41:29",
      "content": "<p>Here is code. Each row in EEG file is 1/200 of second. Each row in Spectrogram file is 2 seconds. Therefore we use offsets in train.csv like code below:</p>\n<pre><code> = \n = \n = \n\n = pd.read_csv()\n = train.iloc[GET_ROW]\n\n = pd.read_parquet(f)\n = int( row.eeg_label_set_seconds )\n = eeg.iloc[eeg_set*:(eeg_set+)*]\n\n = pd.read_parquet(f)\n = int( row.spectrogram_label_set_seconds )\n = spectrogram.loc[(spectrogram.time&gt;=spec_set)\n                 &amp;(spectrogram.time&lt;spec_set+)]\n</code></pre>",
      "rawMarkdown": "Here is code. Each row in EEG file is 1/200 of second. Each row in Spectrogram file is 2 seconds. Therefore we use offsets in train.csv like code below:\n\n    GET_ROW = 0\n    EEG_PATH = 'train_eegs/'\n    SPEC_PATH = 'train_spectrograms/'\n\n    train = pd.read_csv('train.csv')\n    row = train.iloc[GET_ROW]\n\n    eeg = pd.read_parquet(f'{EEG_PATH}{row.eeg_id}.parquet')\n    eeg_offset = int( row.eeg_label_offset_seconds )\n    eeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\n\n    spectrogram = pd.read_parquet(f'{SPEC_PATH}{row.spectrogram_id}.parquet')\n    spec_offset = int( row.spectrogram_label_offset_seconds )\n    spectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                     &(spectrogram.time<spec_offset+600)]",
      "votes": null
    },
    {
      "id": "2645196",
      "postDate": "02/10/2024 03:52:26",
      "content": "<p>thanks chris! very helpful as always!</p>",
      "rawMarkdown": "thanks chris! very helpful as always!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2644383,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/09/2024 12:41:29",
      "content": "<p>Here is code. Each row in EEG file is 1/200 of second. Each row in Spectrogram file is 2 seconds. Therefore we use offsets in train.csv like code below:</p>\n<pre><code> = \n = \n = \n\n = pd.read_csv()\n = train.iloc[GET_ROW]\n\n = pd.read_parquet(f)\n = int( row.eeg_label_set_seconds )\n = eeg.iloc[eeg_set*:(eeg_set+)*]\n\n = pd.read_parquet(f)\n = int( row.spectrogram_label_set_seconds )\n = spectrogram.loc[(spectrogram.time&gt;=spec_set)\n                 &amp;(spectrogram.time&lt;spec_set+)]\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2645196,
          "author_name": "danlxr",
          "author_url": "",
          "post_date": "02/10/2024 03:52:26",
          "content": "<p>thanks chris! very helpful as always!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2643810": "I have just started on the competition, and i run into this problem, where I m not sure how the spectrogram offset and the EEG offset translate to the exact location in the dataframe. as far as I know the spectrogram are of the shape 300*400 . and eeg should be of the shape 20*10000. but the offset are often larger than these numbers. so my question is, how is the offset calculated and how do we use them? thanks.",
    "2644383": "Here is code. Each row in EEG file is 1/200 of second. Each row in Spectrogram file is 2 seconds. Therefore we use offsets in train.csv like code below:\n\n    GET_ROW = 0\n    EEG_PATH = 'train_eegs/'\n    SPEC_PATH = 'train_spectrograms/'\n\n    train = pd.read_csv('train.csv')\n    row = train.iloc[GET_ROW]\n\n    eeg = pd.read_parquet(f'{EEG_PATH}{row.eeg_id}.parquet')\n    eeg_offset = int( row.eeg_label_offset_seconds )\n    eeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\n\n    spectrogram = pd.read_parquet(f'{SPEC_PATH}{row.spectrogram_id}.parquet')\n    spec_offset = int( row.spectrogram_label_offset_seconds )\n    spectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                     &(spectrogram.time<spec_offset+600)]",
    "2645196": "thanks chris! very helpful as always!"
  },
  "source": "meta"
}