{
  "id": 467940,
  "title": "Offset of EEGs and Spectrograms",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467940",
  "author_name": "",
  "post_date": "2024-01-14T18:15:42.918431300Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone.</p>\n<p>I am still in the first stage of exploring the data. I am a bit confused about the offset columns eeg_label_offset_seconds and spectrogram_label_offset_seconds.<br>\nHere are my questions and considerations.</p>\n<ol>\n<li><p>offset is the time from where we start to track the signal. If offset is 6 we start tracking EEG at 6s and track it for the next 50 seconds. For Spectrogram we start tracking it at 6 second and track it for 10 minutes. We than use this data for our models. Do I understand the first part correctly?</p></li>\n<li><p>EEG do not have a time column. So how much samples do we take and where do we start. We see the frequency is 200 Hz so this means 200 rows per second. So if we start with 6 second we should start with row 1200. The last row for 50 second of EEG signal should then be at 10000+1200=11200. Is my logic correct?</p></li>\n<li><p>Then if we go to spectrographs we have a time columns so we should just filter by it from 6 seconds 10min 6 seconds (606 seconds). Do you agree with it?</p></li>\n<li><p>We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?</p></li>\n</ol>\n<p>I will be thankful for any constructive comment under this topic</p>",
  "messages": [
    {
      "id": "2601857",
      "postDate": "01/14/2024 18:15:42",
      "content": "<p>Hi everyone.</p>\n<p>I am still in the first stage of exploring the data. I am a bit confused about the offset columns eeg_label_offset_seconds and spectrogram_label_offset_seconds.<br>\nHere are my questions and considerations.</p>\n<ol>\n<li><p>offset is the time from where we start to track the signal. If offset is 6 we start tracking EEG at 6s and track it for the next 50 seconds. For Spectrogram we start tracking it at 6 second and track it for 10 minutes. We than use this data for our models. Do I understand the first part correctly?</p></li>\n<li><p>EEG do not have a time column. So how much samples do we take and where do we start. We see the frequency is 200 Hz so this means 200 rows per second. So if we start with 6 second we should start with row 1200. The last row for 50 second of EEG signal should then be at 10000+1200=11200. Is my logic correct?</p></li>\n<li><p>Then if we go to spectrographs we have a time columns so we should just filter by it from 6 seconds 10min 6 seconds (606 seconds). Do you agree with it?</p></li>\n<li><p>We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?</p></li>\n</ol>\n<p>I will be thankful for any constructive comment under this topic</p>",
      "rawMarkdown": "Hi everyone.\n\nI am still in the first stage of exploring the data. I am a bit confused about the offset columns eeg_label_offset_seconds and spectrogram_label_offset_seconds.\nHere are my questions and considerations.\n\n1. offset is the time from where we start to track the signal. If offset is 6 we start tracking EEG at 6s and track it for the next 50 seconds. For Spectrogram we start tracking it at 6 second and track it for 10 minutes. We than use this data for our models. Do I understand the first part correctly?\n\n2. EEG do not have a time column. So how much samples do we take and where do we start. We see the frequency is 200 Hz so this means 200 rows per second. So if we start with 6 second we should start with row 1200. The last row for 50 second of EEG signal should then be at 10000+1200=11200. Is my logic correct?\n\n3. Then if we go to spectrographs we have a time columns so we should just filter by it from 6 seconds 10min 6 seconds (606 seconds). Do you agree with it?\n\n4. We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?\n\nI will be thankful for any constructive comment under this topic",
      "votes": null
    },
    {
      "id": "2602512",
      "postDate": "01/15/2024 06:54:24",
      "content": "<blockquote>\n  <p>We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?</p>\n</blockquote>\n<p>different <code>eeg_label_offset_seconds</code> implies different number of subsamples(50s) taken and is present in same .parquet file present in the data</p>",
      "rawMarkdown": ">We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?\n\ndifferent `eeg_label_offset_seconds` implies different number of subsamples(50s) taken and is present in same .parquet file present in the data",
      "votes": null
    },
    {
      "id": "2602543",
      "postDate": "01/15/2024 07:17:27",
      "content": "<p><a href=\"https://www.kaggle.com/klemenvrhovec\" target=\"_blank\">@klemenvrhovec</a> check this  <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">Discussion</a>, it answer all your questions</p>",
      "rawMarkdown": "klemenvrhovec check this  [Discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010), it answer all your questions",
      "votes": null
    },
    {
      "id": "2606511",
      "postDate": "01/17/2024 17:31:17",
      "content": "<p>Thank you 🙂</p>",
      "rawMarkdown": "Thank you 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2602512,
      "author_name": "shashankshuklacodedl",
      "author_url": "",
      "post_date": "01/15/2024 06:54:24",
      "content": "<blockquote>\n  <p>We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?</p>\n</blockquote>\n<p>different <code>eeg_label_offset_seconds</code> implies different number of subsamples(50s) taken and is present in same .parquet file present in the data</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2602543,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/15/2024 07:17:27",
      "content": "<p><a href=\"https://www.kaggle.com/klemenvrhovec\" target=\"_blank\">@klemenvrhovec</a> check this  <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">Discussion</a>, it answer all your questions</p>",
      "votes": null,
      "replies": [
        {
          "id": 2606511,
          "author_name": "klemenvrhovec",
          "author_url": "",
          "post_date": "01/17/2024 17:31:17",
          "content": "<p>Thank you 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2601857": "Hi everyone.\n\nI am still in the first stage of exploring the data. I am a bit confused about the offset columns eeg_label_offset_seconds and spectrogram_label_offset_seconds.\nHere are my questions and considerations.\n\n1. offset is the time from where we start to track the signal. If offset is 6 we start tracking EEG at 6s and track it for the next 50 seconds. For Spectrogram we start tracking it at 6 second and track it for 10 minutes. We than use this data for our models. Do I understand the first part correctly?\n\n2. EEG do not have a time column. So how much samples do we take and where do we start. We see the frequency is 200 Hz so this means 200 rows per second. So if we start with 6 second we should start with row 1200. The last row for 50 second of EEG signal should then be at 10000+1200=11200. Is my logic correct?\n\n3. Then if we go to spectrographs we have a time columns so we should just filter by it from 6 seconds 10min 6 seconds (606 seconds). Do you agree with it?\n\n4. We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?\n\nI will be thankful for any constructive comment under this topic",
    "2602512": ">We have many rows in train.csv file that have the same eeg_id and spectrogram_id with different eeg_label_offset_seconds. Is this just so we cover more different cases of the EEGs? Or is there something else behind it?\n\ndifferent `eeg_label_offset_seconds` implies different number of subsamples(50s) taken and is present in same .parquet file present in the data",
    "2602543": "klemenvrhovec check this  [Discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010), it answer all your questions",
    "2606511": "Thank you 🙂"
  },
  "source": "meta"
}