{
  "id": 468010,
  "title": "Understanding Competition Data and EfficientNetB2 Starter - LB 0.43 🎉",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/468010",
  "author_name": "Chris Deotte",
  "post_date": "2024-01-15T02:17:55.586000",
  "votes": 603,
  "comment_count": 165,
  "views": 0,
  "content": "<p>The data in this competition is confusing. <strong>Question:</strong> We are given <code>train.csv</code> with 106,800 rows but there are only 17089 unique <code>eeg_ids</code>, 11138 unique <code>spectrogram_ids</code>, and 1950 unique <code>patients</code>. What is going on?? <strong>Answer:</strong> Each <code>row</code> of train is a <code>window of time</code> from one specific <code>patient</code>. And the corresponding <code>eeg</code> and <code>spectrogram</code> can be found in the corresponding Kaggle <code>parquet</code> files.</p>\n<h1>Data Explained</h1>\n<p>Since each <code>row</code> of <code>train.csv</code> is a specific <code>window in time</code> (for a specific <code>patient_id</code>), each row has a specific middle timestamp in seconds. For example, maybe <code>row 235</code> has center timestamp <code>T = May 3 2023 19:30:06</code> exactly in the middle of both its <code>EEG time window</code> and <code>Spectrogram time window</code>. (Note that the dataframe does <strong>not</strong> give us the middle timestamp).</p>\n<p>The <code>EEG time window</code> is length 50 seconds and the <code>Spectrogram time window</code> is length 600 seconds. And both have the <strong>same center timestamp</strong>. In this competition, we are asked to predict the event occurring in the middle 10 seconds of both these time windows:</p>\n<ul>\n<li>Center is <code>T</code></li>\n<li>EEG is <code>[T-25:T+25]</code></li>\n<li>Spectrogram is <code>[T-300:T+300]</code></li>\n<li>We predict event in <code>[T-5:T+5]</code></li>\n</ul>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain8.png\"></p>\n<h1>EEG Parquet Files</h1>\n<p>The <code>EEG parquet files</code> are longer than 50 seconds. One <code>EEG parquet file</code> has multiple rows of <code>time windows</code> inside it. Similarily, the <code>Spectrogram parquet files</code> are longer than 600 seconds. One <code>Spectrogram parquet file</code> has multiple rows of <code>time windows</code> inside it.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain9.png\"></p>\n<h1>Spectrogram Parquet Files</h1>\n<p>There are fewer <code>Spectrogram parquet files</code> than <code>EEG parquet files</code>. This is because two rows may have the same <code>Spectrogram parquet file</code> but different <code>EEG parquet files</code>.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain11.png\"></p>\n<h1>Code to Retrieve EEG and Spectrogram</h1>\n<p>For a specific row of <code>train.csv</code>, here is code to retrieve the corresponding EEG and Spectrogram. Note that <code>train.csv</code> does <strong>not</strong> give us the middle timestamp. Instead it gives us the beginning timestamp for each timewindow. The two beginnings are determined from <code>eeg_label_offset_seconds</code> and <code>spectrogram_label_offset_seconds</code>. These are offsets from the start of the parquet file which tells us where the <code>eeg</code> and <code>spectrogram</code> time windows begin respectively.</p>\n<pre><code> = \n = \n = \n\n = pd.read_csv()\n = train.iloc[GET_ROW]\n\n = pd.read_parquet(f)\n = int( row.eeg_label_set_seconds )\n = eeg.iloc[eeg_set*:(eeg_set+)*]\n\n = pd.read_parquet(f)\n = int( row.spectrogram_label_set_seconds )\n = spectrogram.loc[(spectrogram.time&gt;=spec_set)\n                     &amp;(spectrogram.time&lt;spec_set+)]\n</code></pre>\n<h1>EfficientNetB2 Starter - CV 0.59 - LB 0.43 - (Deep Learning)</h1>\n<p>When training our models we have (at least) 4 choices of weighting train samples. For example the 4th bullet point says that our dataloader will give an equal chance of outputting a sample for each <code>patient</code>. Whereas bullet point 2 says that our dataloader will give an equal chance of outputting a sample for each <code>eeg id</code>.</p>\n<ul>\n<li>106,800 unique rows</li>\n<li>17,089 unique eeg ids</li>\n<li>11,138 unique spectrogram ids</li>\n<li>1950 unique patient ids</li>\n</ul>\n<p>From LB probing (discussed <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">here</a>), we achieve the best LB score using bullet point 2 (i.e. eeg ids). I published a TensorFlow EfficientNetB2 starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a> that trains with bullet point 2 and achieves CV 0.73 and LB 0.57. (UPDATE: <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> converted to PyTorch <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092\" target=\"_blank\">here</a>)</p>\n<h1>CatBoost Starter - CV 0.74 - LB 0.60 - (Machine Learning)</h1>\n<p>I published a CatBoost starter <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">here</a> which achieves CV 0.74 and LB 0.60 (using bullet point 2 above).</p>\n<h1>WaveNet Starter - CV 0.81 - LB 0.52 - (Deep Learning)</h1>\n<p>I published a TensorFlow WaveNet starter <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\" target=\"_blank\">here</a> which achieves CV 0.81 and LB 0.52 (using bullet point 2 above). And it only uses EEG data (and not spectrograms). (UPDATE: <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> converted to PyTorch <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477610\" target=\"_blank\">here</a>)</p>\n<h1>Spectrogram versus EEG</h1>\n<p>My EffNet and CatBoost starter notebooks currently only use information from Kaggle's spectrograms. We can improve CV and LB score by incorporating information from EEGs. <strong>UPDATE</strong>: I published a starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">here</a> to convert EEG data into spectrograms (and we can incorporate these EEG spectrograms into my EffNet and CatBoost starters). My WaveNet starter above trains using the raw EEG waveforms. Consider combining my WaveNet and EffNet into a single model which accepts input of both EEG waveforms and Kaggle spectrograms! <strong>UPDATE</strong> Recent versions of EfficientNet starter and CatBoost starter now use both EEG spectrograms and Kaggle spectrograms.</p>\n<h1>Kaggle Dataset</h1>\n<p>Our dataloader needs to read 11138 spectrogram files. The most efficient way to design our dataloader is to first read all 11138 spectrogram files into memory once. Then our dataload uses RAM during epoch training instead of continually reading the parquets from disk.</p>\n<p>I created a Kaggle <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">dataset</a> with one file that contains all 11138 spectrograms. First we load this one file (i.e. Python dictionary of parquets converted into NumPy arrays) into memory. (Reading one file is faster than reading 11k files). Then our dataloader speeds through training!</p>\n<p>UPDATE: Additionally, i have datasets <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eegs\" target=\"_blank\">here</a> containing EEG spectrograms and EEG raw data respectively.</p>\n<h1>Enjoy - Have Fun!</h1>",
  "messages": [
    {
      "id": 2602232,
      "postDate": "2024-01-15T02:17:55.587Z",
      "content": "<p>The data in this competition is confusing. <strong>Question:</strong> We are given <code>train.csv</code> with 106,800 rows but there are only 17089 unique <code>eeg_ids</code>, 11138 unique <code>spectrogram_ids</code>, and 1950 unique <code>patients</code>. What is going on?? <strong>Answer:</strong> Each <code>row</code> of train is a <code>window of time</code> from one specific <code>patient</code>. And the corresponding <code>eeg</code> and <code>spectrogram</code> can be found in the corresponding Kaggle <code>parquet</code> files.</p>\n<h1>Data Explained</h1>\n<p>Since each <code>row</code> of <code>train.csv</code> is a specific <code>window in time</code> (for a specific <code>patient_id</code>), each row has a specific middle timestamp in seconds. For example, maybe <code>row 235</code> has center timestamp <code>T = May 3 2023 19:30:06</code> exactly in the middle of both its <code>EEG time window</code> and <code>Spectrogram time window</code>. (Note that the dataframe does <strong>not</strong> give us the middle timestamp).</p>\n<p>The <code>EEG time window</code> is length 50 seconds and the <code>Spectrogram time window</code> is length 600 seconds. And both have the <strong>same center timestamp</strong>. In this competition, we are asked to predict the event occurring in the middle 10 seconds of both these time windows:</p>\n<ul>\n<li>Center is <code>T</code></li>\n<li>EEG is <code>[T-25:T+25]</code></li>\n<li>Spectrogram is <code>[T-300:T+300]</code></li>\n<li>We predict event in <code>[T-5:T+5]</code></li>\n</ul>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain8.png\"></p>\n<h1>EEG Parquet Files</h1>\n<p>The <code>EEG parquet files</code> are longer than 50 seconds. One <code>EEG parquet file</code> has multiple rows of <code>time windows</code> inside it. Similarily, the <code>Spectrogram parquet files</code> are longer than 600 seconds. One <code>Spectrogram parquet file</code> has multiple rows of <code>time windows</code> inside it.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain9.png\"></p>\n<h1>Spectrogram Parquet Files</h1>\n<p>There are fewer <code>Spectrogram parquet files</code> than <code>EEG parquet files</code>. This is because two rows may have the same <code>Spectrogram parquet file</code> but different <code>EEG parquet files</code>.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain11.png\"></p>\n<h1>Code to Retrieve EEG and Spectrogram</h1>\n<p>For a specific row of <code>train.csv</code>, here is code to retrieve the corresponding EEG and Spectrogram. Note that <code>train.csv</code> does <strong>not</strong> give us the middle timestamp. Instead it gives us the beginning timestamp for each timewindow. The two beginnings are determined from <code>eeg_label_offset_seconds</code> and <code>spectrogram_label_offset_seconds</code>. These are offsets from the start of the parquet file which tells us where the <code>eeg</code> and <code>spectrogram</code> time windows begin respectively.</p>\n<pre><code> = \n = \n = \n\n = pd.read_csv()\n = train.iloc[GET_ROW]\n\n = pd.read_parquet(f)\n = int( row.eeg_label_set_seconds )\n = eeg.iloc[eeg_set*:(eeg_set+)*]\n\n = pd.read_parquet(f)\n = int( row.spectrogram_label_set_seconds )\n = spectrogram.loc[(spectrogram.time&gt;=spec_set)\n                     &amp;(spectrogram.time&lt;spec_set+)]\n</code></pre>\n<h1>EfficientNetB2 Starter - CV 0.59 - LB 0.43 - (Deep Learning)</h1>\n<p>When training our models we have (at least) 4 choices of weighting train samples. For example the 4th bullet point says that our dataloader will give an equal chance of outputting a sample for each <code>patient</code>. Whereas bullet point 2 says that our dataloader will give an equal chance of outputting a sample for each <code>eeg id</code>.</p>\n<ul>\n<li>106,800 unique rows</li>\n<li>17,089 unique eeg ids</li>\n<li>11,138 unique spectrogram ids</li>\n<li>1950 unique patient ids</li>\n</ul>\n<p>From LB probing (discussed <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">here</a>), we achieve the best LB score using bullet point 2 (i.e. eeg ids). I published a TensorFlow EfficientNetB2 starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a> that trains with bullet point 2 and achieves CV 0.73 and LB 0.57. (UPDATE: <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> converted to PyTorch <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092\" target=\"_blank\">here</a>)</p>\n<h1>CatBoost Starter - CV 0.74 - LB 0.60 - (Machine Learning)</h1>\n<p>I published a CatBoost starter <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">here</a> which achieves CV 0.74 and LB 0.60 (using bullet point 2 above).</p>\n<h1>WaveNet Starter - CV 0.81 - LB 0.52 - (Deep Learning)</h1>\n<p>I published a TensorFlow WaveNet starter <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\" target=\"_blank\">here</a> which achieves CV 0.81 and LB 0.52 (using bullet point 2 above). And it only uses EEG data (and not spectrograms). (UPDATE: <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> converted to PyTorch <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477610\" target=\"_blank\">here</a>)</p>\n<h1>Spectrogram versus EEG</h1>\n<p>My EffNet and CatBoost starter notebooks currently only use information from Kaggle's spectrograms. We can improve CV and LB score by incorporating information from EEGs. <strong>UPDATE</strong>: I published a starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">here</a> to convert EEG data into spectrograms (and we can incorporate these EEG spectrograms into my EffNet and CatBoost starters). My WaveNet starter above trains using the raw EEG waveforms. Consider combining my WaveNet and EffNet into a single model which accepts input of both EEG waveforms and Kaggle spectrograms! <strong>UPDATE</strong> Recent versions of EfficientNet starter and CatBoost starter now use both EEG spectrograms and Kaggle spectrograms.</p>\n<h1>Kaggle Dataset</h1>\n<p>Our dataloader needs to read 11138 spectrogram files. The most efficient way to design our dataloader is to first read all 11138 spectrogram files into memory once. Then our dataload uses RAM during epoch training instead of continually reading the parquets from disk.</p>\n<p>I created a Kaggle <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">dataset</a> with one file that contains all 11138 spectrograms. First we load this one file (i.e. Python dictionary of parquets converted into NumPy arrays) into memory. (Reading one file is faster than reading 11k files). Then our dataloader speeds through training!</p>\n<p>UPDATE: Additionally, i have datasets <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eegs\" target=\"_blank\">here</a> containing EEG spectrograms and EEG raw data respectively.</p>\n<h1>Enjoy - Have Fun!</h1>",
      "rawMarkdown": "The data in this competition is confusing. **Question:** We are given `train.csv` with 106,800 rows but there are only 17089 unique `eeg_ids`, 11138 unique `spectrogram_ids`, and 1950 unique `patients`. What is going on?? **Answer:** Each `row` of train is a `window of time` from one specific `patient`. And the corresponding `eeg` and `spectrogram` can be found in the corresponding Kaggle `parquet` files.\n\n# Data Explained\nSince each `row` of `train.csv` is a specific `window in time` (for a specific `patient_id`), each row has a specific middle timestamp in seconds. For example, maybe `row 235` has center timestamp `T = May 3 2023 19:30:06` exactly in the middle of both its `EEG time window` and `Spectrogram time window`. (Note that the dataframe does **not** give us the middle timestamp).\n\nThe `EEG time window` is length 50 seconds and the `Spectrogram time window` is length 600 seconds. And both have the **same center timestamp**. In this competition, we are asked to predict the event occurring in the middle 10 seconds of both these time windows:\n* Center is `T`\n* EEG is `[T-25:T+25]`\n* Spectrogram is `[T-300:T+300]`\n* We predict event in `[T-5:T+5]`\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain8.png)\n\n# EEG Parquet Files\nThe `EEG parquet files` are longer than 50 seconds. One `EEG parquet file` has multiple rows of `time windows` inside it. Similarily, the `Spectrogram parquet files` are longer than 600 seconds. One `Spectrogram parquet file` has multiple rows of `time windows` inside it.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain9.png)\n\n# Spectrogram Parquet Files\nThere are fewer `Spectrogram parquet files` than `EEG parquet files`. This is because two rows may have the same `Spectrogram parquet file` but different `EEG parquet files`.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain11.png)\n\n# Code to Retrieve EEG and Spectrogram\nFor a specific row of `train.csv`, here is code to retrieve the corresponding EEG and Spectrogram. Note that `train.csv` does **not** give us the middle timestamp. Instead it gives us the beginning timestamp for each timewindow. The two beginnings are determined from `eeg_label_offset_seconds` and `spectrogram_label_offset_seconds`. These are offsets from the start of the parquet file which tells us where the `eeg` and `spectrogram` time windows begin respectively.\n\n    GET_ROW = 0\n    EEG_PATH = 'train_eegs/'\n    SPEC_PATH = 'train_spectrograms/'\n\n    train = pd.read_csv('train.csv')\n    row = train.iloc[GET_ROW]\n\n    eeg = pd.read_parquet(f'{EEG_PATH}{row.eeg_id}.parquet')\n    eeg_offset = int( row.eeg_label_offset_seconds )\n    eeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\n\n    spectrogram = pd.read_parquet(f'{SPEC_PATH}{row.spectrogram_id}.parquet')\n    spec_offset = int( row.spectrogram_label_offset_seconds )\n    spectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                         &(spectrogram.time<spec_offset+600)]\n\n# EfficientNetB2 Starter - CV 0.59 - LB 0.43 - (Deep Learning)\nWhen training our models we have (at least) 4 choices of weighting train samples. For example the 4th bullet point says that our dataloader will give an equal chance of outputting a sample for each `patient`. Whereas bullet point 2 says that our dataloader will give an equal chance of outputting a sample for each `eeg id`.\n* 106,800 unique rows\n* 17,089 unique eeg ids\n* 11,138 unique spectrogram ids\n* 1950 unique patient ids\n\nFrom LB probing (discussed [here][1]), we achieve the best LB score using bullet point 2 (i.e. eeg ids). I published a TensorFlow EfficientNetB2 starter notebook [here][2] that trains with bullet point 2 and achieves CV 0.73 and LB 0.57. (UPDATE: @alejopaullier converted to PyTorch [here][10])\n\n# CatBoost Starter - CV 0.74 - LB 0.60 - (Machine Learning)\n\nI published a CatBoost starter [here][4] which achieves CV 0.74 and LB 0.60 (using bullet point 2 above).\n\n# WaveNet Starter - CV 0.81 - LB 0.52 - (Deep Learning)\n\nI published a TensorFlow WaveNet starter [here][5] which achieves CV 0.81 and LB 0.52 (using bullet point 2 above). And it only uses EEG data (and not spectrograms). (UPDATE: @alejopaullier converted to PyTorch [here][9])\n\n# Spectrogram versus EEG\nMy EffNet and CatBoost starter notebooks currently only use information from Kaggle's spectrograms. We can improve CV and LB score by incorporating information from EEGs. **UPDATE**: I published a starter notebook [here][6] to convert EEG data into spectrograms (and we can incorporate these EEG spectrograms into my EffNet and CatBoost starters). My WaveNet starter above trains using the raw EEG waveforms. Consider combining my WaveNet and EffNet into a single model which accepts input of both EEG waveforms and Kaggle spectrograms! **UPDATE** Recent versions of EfficientNet starter and CatBoost starter now use both EEG spectrograms and Kaggle spectrograms.\n\n# Kaggle Dataset\nOur dataloader needs to read 11138 spectrogram files. The most efficient way to design our dataloader is to first read all 11138 spectrogram files into memory once. Then our dataload uses RAM during epoch training instead of continually reading the parquets from disk.\n\nI created a Kaggle [dataset][3] with one file that contains all 11138 spectrograms. First we load this one file (i.e. Python dictionary of parquets converted into NumPy arrays) into memory. (Reading one file is faster than reading 11k files). Then our dataloader speeds through training!\n\nUPDATE: Additionally, i have datasets [here][7] and [here][8] containing EEG spectrograms and EEG raw data respectively.\n\n# Enjoy - Have Fun!\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\n[2]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[3]: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n[4]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[5]: https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\n[6]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n[7]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\n[8]: https://www.kaggle.com/datasets/cdeotte/brain-eegs\n[9]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477610\n[10]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092",
      "votes": 596
    },
    {
      "id": 2602237,
      "postDate": "2024-01-15T02:24:35.027Z",
      "content": "<p>Chris, the Gandalf of Kaggle.</p>",
      "rawMarkdown": "Chris, the Gandalf of Kaggle.",
      "votes": 20
    },
    {
      "id": 2619843,
      "postDate": "2024-01-25T18:42:27.200Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> are these offsets right? EEG offsets being measured with the \"zero\" placed on the begging of the EEG not from the start of the spectogram.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F698fcf40886403a5bc93d19a58eee8e4%2FScreen%20Shot%202024-01-25%20at%2015.37.22.png?generation=1706207919227935&amp;alt=media\"></p>",
      "rawMarkdown": "@cdeotte are these offsets right? EEG offsets being measured with the \"zero\" placed on the begging of the EEG not from the start of the spectogram.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F698fcf40886403a5bc93d19a58eee8e4%2FScreen%20Shot%202024-01-25%20at%2015.37.22.png?generation=1706207919227935&alt=media)",
      "votes": 8,
      "replies": [
        {
          "id": 2620008,
          "postDate": "2024-01-25T20:34:55.213Z",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> that is correct. All offsets are relative to the specific parquet on disk.</p>\n<p>I like your diagram, I updated my post with new diagrams to make it clearer.</p>",
          "rawMarkdown": "Yes @alejopaullier that is correct. All offsets are relative to the specific parquet on disk.\n\nI like your diagram, I updated my post with new diagrams to make it clearer.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2602866,
      "postDate": "2024-01-15T11:49:27.293Z",
      "content": "<p>Thanks for the explanation. I understand your scheme for non-overlapping segments. I'm sorry, but I still don't understand the overlapping segments. Is my below diagram correctly? </p>\n<p>For example EEG_ID 1628180742</p>\n<ol>\n<li><strong>eeg_label_offset_seconds</strong> should be used for both spectrogram and EEG. </li>\n<li>Marking of labels starts from the center of the EEG_ID, then if the offset is 40 seconds we have no EEG data for this section ? 🥴</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F0c1f5de7da95ba7d40d27c0e99ea958b%2F3.drawio%20(2).png?generation=1705319245246205&amp;alt=media\"></p>",
      "rawMarkdown": "Thanks for the explanation. I understand your scheme for non-overlapping segments. I'm sorry, but I still don't understand the overlapping segments. Is my below diagram correctly? \n\nFor example EEG_ID 1628180742\n\n1. **eeg_label_offset_seconds** should be used for both spectrogram and EEG. \n2. Marking of labels starts from the center of the EEG_ID, then if the offset is 40 seconds we have no EEG data for this section ? 🥴\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F0c1f5de7da95ba7d40d27c0e99ea958b%2F3.drawio%20(2).png?generation=1705319245246205&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 2602876,
          "postDate": "2024-01-15T12:02:17.503Z",
          "content": "<p>Hi. The variable <code>eeg_label_offset_seconds</code> is the beginning, not the middle. (We are not given the middle). The file <code>train_eegs/1628180742.parquet</code> is 90 seconds long. (It is 18,000 rows and each row is 1/200 seconds). Therefore when <code>eeg_label_offset_seconds = 40</code>, we use the 50 seconds between time 40 and time 90. These are rows <code>40*200</code> thru <code>90*200</code> of <code>1628180742.parquet</code>.</p>\n<p>And the file <code>train_spectrograms/1628180742.parquet</code> is 10 minutes and 40 seconds long. (It is 320 rows and each row is 2 seconds). Note to retrieve spectrograms, we <strong>do not</strong> use <code>eeg_label_offset_seconds</code>. Instead we use <code>spectrogram_label_offset_seconds</code>. For some rows in <code>train.csv</code> the <code>eeg_label_offset_seconds != spectrogram_label_offset_seconds</code>. In your example, they are equal.</p>",
          "rawMarkdown": "Hi. The variable `eeg_label_offset_seconds` is the beginning, not the middle. (We are not given the middle). The file `train_eegs/1628180742.parquet` is 90 seconds long. (It is 18,000 rows and each row is 1/200 seconds). Therefore when `eeg_label_offset_seconds = 40`, we use the 50 seconds between time 40 and time 90. These are rows `40*200` thru `90*200` of `1628180742.parquet`.\n\nAnd the file `train_spectrograms/1628180742.parquet` is 10 minutes and 40 seconds long. (It is 320 rows and each row is 2 seconds). Note to retrieve spectrograms, we **do not** use `eeg_label_offset_seconds`. Instead we use `spectrogram_label_offset_seconds`. For some rows in `train.csv` the `eeg_label_offset_seconds != spectrogram_label_offset_seconds`. In your example, they are equal.",
          "votes": 12
        },
        {
          "id": 2602878,
          "postDate": "2024-01-15T12:03:14.740Z",
          "content": "<p>I updated my discussion post with code to extract the eeg and spectrogram for a specific row of <code>train.csv</code>.</p>",
          "rawMarkdown": "I updated my discussion post with code to extract the eeg and spectrogram for a specific row of `train.csv`.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2608513,
      "postDate": "2024-01-18T20:57:48.227Z",
      "content": "<p>Did you figured out all this alone? My first thought was that eeg_sub_id and spectogram_sub_id were foreign key to the corresponding files. <br>\nThe eeg_offset*200 is also a tricky part.</p>\n<p>The competition description should be updated with this explanation. Much more clear now.</p>",
      "rawMarkdown": "Did you figured out all this alone? My first thought was that eeg_sub_id and spectogram_sub_id were foreign key to the corresponding files. \nThe eeg_offset*200 is also a tricky part.\n\nThe competition description should be updated with this explanation. Much more clear now.",
      "votes": 6
    },
    {
      "id": 2602297,
      "postDate": "2024-01-15T03:51:53.703Z",
      "content": "<p>I thought I understood, but now I understand, thank you! =)</p>",
      "rawMarkdown": "I thought I understood, but now I understand, thank you! =)",
      "votes": 6
    },
    {
      "id": 2617670,
      "postDate": "2024-01-24T11:40:14.957Z",
      "content": "<p>Thank you for your kind explanation. <br>\nI have a question. Why is there a difference in time between the provided spectrograms and the EEG data?<br>\n50 sec vs 600 sec</p>",
      "rawMarkdown": "Thank you for your kind explanation. \nI have a question. Why is there a difference in time between the provided spectrograms and the EEG data?\n50 sec vs 600 sec",
      "votes": 3,
      "replies": [
        {
          "id": 2617757,
          "postDate": "2024-01-24T12:50:35.583Z",
          "content": "<p>When doctors evaluate patients, they use the spectrograms to see the \"big picture\" of how the patient were feeling for 10 minute window. And they use the EEG waveforms to see the \"zoom in picture\" of what what happening close to the possible event. (Note the event we are classifying is the middle 10 second window).</p>",
          "rawMarkdown": "When doctors evaluate patients, they use the spectrograms to see the \"big picture\" of how the patient were feeling for 10 minute window. And they use the EEG waveforms to see the \"zoom in picture\" of what what happening close to the possible event. (Note the event we are classifying is the middle 10 second window).",
          "votes": 11,
          "replies": [
            {
              "id": 2618940,
              "postDate": "2024-01-25T05:32:07.160Z",
              "content": "<p>Thank you!<br>\n I understand that the time of the EEG is different from that of the spectrogram because of the doctor's diagnostic method.</p>",
              "rawMarkdown": "Thank you!\n I understand that the time of the EEG is different from that of the spectrogram because of the doctor's diagnostic method.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2613345,
      "postDate": "2024-01-22T03:02:19.187Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Amazing discussion thanks for sharing .  Can you please explain why do we multiply by  200 while slicing the eeg_offset , Whereas while slicing spectogram we dont use the same formula . Thanks</p>\n<pre><code>eeg = eeg.iloc[eeg_offset*:(eeg_offset+)*]\nspectrogram = spectrogram.loc[(spectrogram.time&gt;=spec_offset)\n                     &amp;(spectrogram.time&lt;spec_offset+)]\n</code></pre>\n<p>Can you please elaborate on this 2 parts please </p>",
      "rawMarkdown": "@cdeotte  Amazing discussion thanks for sharing .  Can you please explain why do we multiply by  200 while slicing the eeg_offset , Whereas while slicing spectogram we dont use the same formula . Thanks\n```python\neeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\nspectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                     &(spectrogram.time<spec_offset+600)]\n```\nCan you please elaborate on this 2 parts please \n",
      "votes": 3,
      "replies": [
        {
          "id": 2614177,
          "postDate": "2024-01-22T13:33:50.893Z",
          "content": "<p>Hi. Each row of the eeg parquet is <code>1/200</code> of a second. And each row of the spectrogram parquet is <code>2</code> seconds.</p>",
          "rawMarkdown": "Hi. Each row of the eeg parquet is `1/200` of a second. And each row of the spectrogram parquet is `2` seconds.",
          "votes": 5,
          "replies": [
            {
              "id": 2614321,
              "postDate": "2024-01-22T14:49:32.403Z",
              "content": "<p>Thanks alot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for reply I also noticed the data description which says eegs are sampled at the rate of 200 samples per second . </p>",
              "rawMarkdown": "Thanks alot @cdeotte for reply I also noticed the data description which says eegs are sampled at the rate of 200 samples per second . "
            },
            {
              "id": 2624000,
              "postDate": "2024-01-28T14:26:29.170Z",
              "content": "<blockquote>\n  <p>And each row of the spectrogram parquet is 2 seconds.<br>\n  Hey Chris, how did you conclude this?</p>\n</blockquote>",
              "rawMarkdown": "> And each row of the spectrogram parquet is 2 seconds.\nHey Chris, how did you conclude this?",
              "votes": 3
            },
            {
              "id": 2626189,
              "postDate": "2024-01-29T20:10:52.900Z",
              "content": "<p>In the spectrogram parquet files there is a column named <code>time</code>. And each subsequent row increases by 2. (i.e. first row is 1, second row is 3, third row is 5, etc)</p>",
              "rawMarkdown": "In the spectrogram parquet files there is a column named `time`. And each subsequent row increases by 2. (i.e. first row is 1, second row is 3, third row is 5, etc)",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2604067,
      "postDate": "2024-01-16T08:34:25.933Z",
      "content": "<p>Chris, congratulations on your innovative approach to optimizing data loading for the Kaggle competition, particularly your method of loading all 11,138 spectrogram files into memory at once. This strategy is not only efficient but also a great learning point for the community in handling large datasets.</p>\n<p>I have a question: In your experiments with EfficientNetB2 and CatBoost models, how do you address the potential issue of overfitting, especially given the high dimensionality of the EEG and spectrogram data and the relatively smaller number of unique patient IDs?</p>",
      "rawMarkdown": "Chris, congratulations on your innovative approach to optimizing data loading for the Kaggle competition, particularly your method of loading all 11,138 spectrogram files into memory at once. This strategy is not only efficient but also a great learning point for the community in handling large datasets.\n\nI have a question: In your experiments with EfficientNetB2 and CatBoost models, how do you address the potential issue of overfitting, especially given the high dimensionality of the EEG and spectrogram data and the relatively smaller number of unique patient IDs?",
      "votes": 3,
      "replies": [
        {
          "id": 2604341,
          "postDate": "2024-01-16T11:25:49.307Z",
          "content": "<p>Hi. CatBoost has internal tricks to prevent overfitting. With EfficientNet, i am only training for 3 epochs (i.e. few epochs). Furthermore, we can improve EfficientNet using data augmentation. And we can add more regularization to both models.</p>",
          "rawMarkdown": "Hi. CatBoost has internal tricks to prevent overfitting. With EfficientNet, i am only training for 3 epochs (i.e. few epochs). Furthermore, we can improve EfficientNet using data augmentation. And we can add more regularization to both models.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2603434,
      "postDate": "2024-01-15T19:13:35.947Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for the explanation. I have one question. I see in your code that you aggregated the target across eeg_id. My understanding is that the fine-grained target for each row instead of the whole egg_id is more time approprate (maybe the waves show seizure in some time period but not others). Any reason why you treat the target this way? Thank you,</p>\n<pre><code>tmp = df()()\n t  TARGETS:\n    train = tmp\n\ny_data = train\ny_data = y_data / y_data(axis=,keepdims=True)\ntrain = y_data\n</code></pre>",
      "rawMarkdown": "@cdeotte Thank you for the explanation. I have one question. I see in your code that you aggregated the target across eeg_id. My understanding is that the fine-grained target for each row instead of the whole egg_id is more time approprate (maybe the waves show seizure in some time period but not others). Any reason why you treat the target this way? Thank you,\n\n```\ntmp = df.groupby('eeg_id')[TARGETS].agg('sum')\nfor t in TARGETS:\n    train[t] = tmp[t].values\n    \ny_data = train[TARGETS].values\ny_data = y_data / y_data.sum(axis=1,keepdims=True)\ntrain[TARGETS] = y_data\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2603456,
          "postDate": "2024-01-15T19:25:51.893Z",
          "content": "<p>We can run <code>train.groupby('eeg_id').expert_consensus.agg('nunique').value_counts()</code> and get</p>\n<pre><code>   \n     \n      \n      \n       \n</code></pre>\n<p>Therefore if we naively just take 1 target to represent the entire <code>eeg_id</code> we have expected value of being correct = 97.6%. Because <code>(16306 + 333 + 31 + 5)/17089 = 0.976</code>. So yes, we could be more precise, but this simple technique works very well and is easy.</p>",
          "rawMarkdown": "We can run `train.groupby('eeg_id').expert_consensus.agg('nunique').value_counts()` and get\n\n    1    16306\n    2      666\n    3       94\n    4       22\n    5        1\n\nTherefore if we naively just take 1 target to represent the entire `eeg_id` we have expected value of being correct = 97.6%. Because `(16306 + 333 + 31 + 5)/17089 = 0.976`. So yes, we could be more precise, but this simple technique works very well and is easy.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2603127,
      "postDate": "2024-01-15T15:32:11.020Z",
      "content": "<p>Thanks for the explanation, now it much cleaner for me, how to align the data. However I still do not understand, why the predictions are made by 50 seconds of the raw eeg and 10 minutes of spectrograms, because it means that spectrogram have much less resolution that eeg? If we predicting for only 10 seconds, then why should look at the whole 10 minutes interval, isn't 50 seconds enough?</p>",
      "rawMarkdown": "Thanks for the explanation, now it much cleaner for me, how to align the data. However I still do not understand, why the predictions are made by 50 seconds of the raw eeg and 10 minutes of spectrograms, because it means that spectrogram have much less resolution that eeg? If we predicting for only 10 seconds, then why should look at the whole 10 minutes interval, isn't 50 seconds enough?",
      "votes": 3,
      "replies": [
        {
          "id": 2603143,
          "postDate": "2024-01-15T15:47:22.113Z",
          "content": "<p>The larger window provides context. Imagine if i asked you to predict what i will eat for lunch today. I can give you the 30 minute window of me sitting at my kitchen table. I can also give you the previous 24 hours of me eating and the future 24 hours.  Then you can use information from the larger window to help predict what I do in the smaller 30 minute window.</p>",
          "rawMarkdown": "The larger window provides context. Imagine if i asked you to predict what i will eat for lunch today. I can give you the 30 minute window of me sitting at my kitchen table. I can also give you the previous 24 hours of me eating and the future 24 hours.  Then you can use information from the larger window to help predict what I do in the smaller 30 minute window.",
          "votes": 14,
          "replies": [
            {
              "id": 2603760,
              "postDate": "2024-01-16T03:07:38.793Z",
              "content": "<p>Super great example.</p>",
              "rawMarkdown": "Super great example.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2618452,
      "postDate": "2024-01-24T18:28:52.277Z",
      "content": "<p>The relationship between <code>spectogram_ids</code> and <code>eeg_id</code> is 1:M? For example, 1 spectogram may have multiple eegs but 1 eeg is always \"inside\" 1 spectogram</p>",
      "rawMarkdown": "The relationship between `spectogram_ids` and `eeg_id` is 1:M? For example, 1 spectogram may have multiple eegs but 1 eeg is always \"inside\" 1 spectogram",
      "votes": 4,
      "replies": [
        {
          "id": 2618469,
          "postDate": "2024-01-24T18:36:06.647Z",
          "content": "<p>Yes. We can confirm this by doing <code>groupby('eeg_id').spectrogram_id.agg('nunique').max()</code> and vice versa (with <code>train.csv</code>)</p>",
          "rawMarkdown": "Yes. We can confirm this by doing `groupby('eeg_id').spectrogram_id.agg('nunique').max()` and vice versa (with `train.csv`)",
          "votes": 5
        }
      ]
    },
    {
      "id": 2608407,
      "postDate": "2024-01-18T18:54:33.187Z",
      "content": "<p>What do the sub_ids do in the dataset? </p>",
      "rawMarkdown": "What do the sub_ids do in the dataset? ",
      "votes": 4,
      "replies": [
        {
          "id": 2608504,
          "postDate": "2024-01-18T20:38:23.163Z",
          "content": "<p>Ignore them. I think they are just identified unique rows of train.csv (Furthermore they do not exist in test.csv)</p>",
          "rawMarkdown": "Ignore them. I think they are just identified unique rows of train.csv (Furthermore they do not exist in test.csv)",
          "votes": 5,
          "replies": [
            {
              "id": 2613246,
              "postDate": "2024-01-21T23:55:18.610Z",
              "content": "<p>ok. thanks </p>",
              "rawMarkdown": "ok. thanks "
            }
          ]
        }
      ]
    },
    {
      "id": 2606539,
      "postDate": "2024-01-17T17:59:34.393Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>  -- I was trying to understand both of your baseline . My understanding is Chris you are taking  approach 2 and Tawara has taken approach 3 . Am I correct ?</p>\n<ol>\n<li>106,800 unique rows</li>\n<li>17,089 unique eeg ids</li>\n<li>11,138 unique spectrogram ids</li>\n<li>1950 unique patient ids</li>\n</ol>\n<p>I also see that Tawara images are 400<em>300 which looks like having all the 4 series together vertically stacked and yours is 128</em>256*4 the four series are placed in channel ? How is your 256 values coming?</p>",
      "rawMarkdown": "Hello @cdeotte @ttahara  -- I was trying to understand both of your baseline . My understanding is Chris you are taking  approach 2 and Tawara has taken approach 3 . Am I correct ?\n1. 106,800 unique rows\n2. 17,089 unique eeg ids\n3. 11,138 unique spectrogram ids\n4. 1950 unique patient ids\n\nI also see that Tawara images are 400*300 which looks like having all the 4 series together vertically stacked and yours is 128*256*4 the four series are placed in channel ? How is your 256 values coming?",
      "votes": 4,
      "replies": [
        {
          "id": 2606561,
          "postDate": "2024-01-17T18:08:48.187Z",
          "content": "<p>Yes, I do 2 and Tawara does 3. My 256 is because I take the middle 512 seconds (instead of full 600 seconds). We both do 400x300.</p>\n<p>My data loader provides 128x256x4. There is padding on top 14 and bottom 14, so my height is 100. My width is 300 cropped to 256. Inside my model I reshape it to 512x256 which is basically same as 400x300</p>",
          "rawMarkdown": "Yes, I do 2 and Tawara does 3. My 256 is because I take the middle 512 seconds (instead of full 600 seconds). We both do 400x300.\n\nMy data loader provides 128x256x4. There is padding on top 14 and bottom 14, so my height is 100. My width is 300 cropped to 256. Inside my model I reshape it to 512x256 which is basically same as 400x300",
          "votes": 11
        },
        {
          "id": 2606961,
          "postDate": "2024-01-18T01:21:15.457Z",
          "content": "<p>Yes, you are right. I use the <strong>first</strong> 600 seconds of each unique spectrogram id.</p>",
          "rawMarkdown": "Yes, you are right. I use the **first** 600 seconds of each unique spectrogram id.",
          "votes": 8
        }
      ]
    },
    {
      "id": 2604227,
      "postDate": "2024-01-16T10:29:07.413Z",
      "content": "<p>Does time matters here?,<br>\nSo solutions like transformers or RNN, will it be effective ? </p>",
      "rawMarkdown": "Does time matters here?,\nSo solutions like transformers or RNN, will it be effective ? ",
      "votes": 4,
      "replies": [
        {
          "id": 2604336,
          "postDate": "2024-01-16T11:24:08.080Z",
          "content": "<p>Time matters for 2 reasons:</p>\n<ul>\n<li>we need time to determine what frequencies are in the signal. (i.e. without time, there is no such thing as 10Hz).</li>\n<li>in 3% of cases, the <code>expert_consensus</code> changes over time</li>\n</ul>\n<p>Most people are currently ignoring the second bullet point and just assigning a single target to a single <code>eeg_id</code>. Later in the competition, if we use the 3% to train we can boost our CV score and LB score slightly.</p>",
          "rawMarkdown": "Time matters for 2 reasons:\n* we need time to determine what frequencies are in the signal. (i.e. without time, there is no such thing as 10Hz).\n* in 3% of cases, the `expert_consensus` changes over time\n\nMost people are currently ignoring the second bullet point and just assigning a single target to a single `eeg_id`. Later in the competition, if we use the 3% to train we can boost our CV score and LB score slightly.",
          "votes": 5
        }
      ]
    },
    {
      "id": 2602420,
      "postDate": "2024-01-15T05:53:18.670Z",
      "content": "<p>Chris, the god of kaggle || thank you for giving such a great code</p>",
      "rawMarkdown": "Chris, the god of kaggle || thank you for giving such a great code",
      "votes": 4
    },
    {
      "id": 2602321,
      "postDate": "2024-01-15T04:27:27.297Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> you are on fire in this comp 🔥.  Btw, loved your explanation about the dataset here. Also, glad to see DL beating ML 😅</p>",
      "rawMarkdown": "@cdeotte you are on fire in this comp 🔥.  Btw, loved your explanation about the dataset here. Also, glad to see DL beating ML 😅\n",
      "votes": 4,
      "replies": [
        {
          "id": 2602824,
          "postDate": "2024-01-15T11:09:00.347Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> . This is fun competition with lots to explore and learn. I see that you are competition host. Thanks for hosting.</p>",
          "rawMarkdown": "Thanks @awsaf49 . This is fun competition with lots to explore and learn. I see that you are competition host. Thanks for hosting.",
          "votes": 2,
          "replies": [
            {
              "id": 2602973,
              "postDate": "2024-01-15T13:38:58.917Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdoette\" target=\"_blank\">@cdoette</a>, I wish I was hosting this amazing competition but sadly I'm not; I am working in this competition on behalf of the Keras Team to share starter notebooks.</p>",
              "rawMarkdown": "@cdoette, I wish I was hosting this amazing competition but sadly I'm not; I am working in this competition on behalf of the Keras Team to share starter notebooks.",
              "votes": 5
            }
          ]
        },
        {
          "id": 2602959,
          "postDate": "2024-01-15T13:20:45.740Z",
          "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> Thanks for hosting this competition, many new topics explored recently (majorly eeg). </p>\n<p><strong>is it good to pin this topic for everyone to understand data better</strong> ? ( we have 5-6 discussions about same topic )</p>",
          "rawMarkdown": "@awsaf49 Thanks for hosting this competition, many new topics explored recently (majorly eeg). \n\n**is it good to pin this topic for everyone to understand data better** ? ( we have 5-6 discussions about same topic )",
          "votes": 1,
          "replies": [
            {
              "id": 2602968,
              "postDate": "2024-01-15T13:34:58.517Z",
              "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> I'm not the competition host per se; I am working in this competition on behalf of the Keras Team to share starter notebooks. You may need to contact the Kaggle Team. Sorry for the confusion.</p>",
              "rawMarkdown": "@seshurajup I'm not the competition host per se; I am working in this competition on behalf of the Keras Team to share starter notebooks. You may need to contact the Kaggle Team. Sorry for the confusion.",
              "votes": 2
            },
            {
              "id": 2602970,
              "postDate": "2024-01-15T13:36:56.043Z",
              "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> Sorry it is my mistake of miss understanding from the tag. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> is it good to keep this topic as pin</p>",
              "rawMarkdown": "@awsaf49 Sorry it is my mistake of miss understanding from the tag. @sohier is it good to keep this topic as pin"
            }
          ]
        }
      ]
    },
    {
      "id": 2745038,
      "postDate": "2024-04-10T10:20:31.303Z",
      "content": "<p>Chris, congratulations on your gold medal in this competition! I would also like to express my deep gratitude for sharing your various knowledge and insights. This was the first competition I attended in earnest. I learned so much from your notebooks and discussions. Thank you!</p>",
      "rawMarkdown": "Chris, congratulations on your gold medal in this competition! I would also like to express my deep gratitude for sharing your various knowledge and insights. This was the first competition I attended in earnest. I learned so much from your notebooks and discussions. Thank you!",
      "votes": 1
    },
    {
      "id": 2744051,
      "postDate": "2024-04-09T18:20:41.433Z",
      "content": "<p>Congratulations and thank you again Chris for your invaluable guides!!</p>",
      "rawMarkdown": "Congratulations and thank you again Chris for your invaluable guides!!",
      "votes": 1
    },
    {
      "id": 2743337,
      "postDate": "2024-04-09T11:46:25.300Z",
      "content": "<p>Congratulations on your gold medal, Chris. I always appreciate your information sharing.  I was wondering why you used GroupKFold in the starter notebook? I'm sure I've seen this before but I can't find it, so if you don't mind, I'd appreciate a refresher. I used StratifiedGroupKFold.</p>",
      "rawMarkdown": "Congratulations on your gold medal, Chris. I always appreciate your information sharing.  I was wondering why you used GroupKFold in the starter notebook? I'm sure I've seen this before but I can't find it, so if you don't mind, I'd appreciate a refresher. I used StratifiedGroupKFold.",
      "votes": 1,
      "replies": [
        {
          "id": 2743573,
          "postDate": "2024-04-09T14:22:57.047Z",
          "content": "<p>Thanks. Congratulations on your strong Silver finish.</p>\n<p>The purpose of CV is to make validation folds as similar to test data as possible. Then using this CV will optimize a model to perform its best on test data.</p>\n<p>In the beginning of this competition, we did not know if test data has the same mean of 6 targets as train data. Therefore I choose to allow my validation folds to have more variation in their 6 target means. This way, the optimized model can perform well even if the test data has a different mean of 6 targets than train.</p>",
          "rawMarkdown": "Thanks. Congratulations on your strong Silver finish.\n\nThe purpose of CV is to make validation folds as similar to test data as possible. Then using this CV will optimize a model to perform its best on test data.\n\nIn the beginning of this competition, we did not know if test data has the same mean of 6 targets as train data. Therefore I choose to allow my validation folds to have more variation in their 6 target means. This way, the optimized model can perform well even if the test data has a different mean of 6 targets than train.",
          "votes": 1,
          "replies": [
            {
              "id": 2744407,
              "postDate": "2024-04-09T22:21:32.517Z",
              "content": "<p>Thank you so much for your detailed answer!</p>",
              "rawMarkdown": "Thank you so much for your detailed answer!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2741085,
      "postDate": "2024-04-08T06:10:16.413Z",
      "content": "<p>This is a very odd way of storing EEG data, I tried parsing this dataset but eventually gave up: <a href=\"https://bionichaos.com/SeizureFuz\" target=\"_blank\">https://bionichaos.com/SeizureFuz</a></p>",
      "rawMarkdown": "This is a very odd way of storing EEG data, I tried parsing this dataset but eventually gave up: https://bionichaos.com/SeizureFuz",
      "votes": 1
    },
    {
      "id": 2736518,
      "postDate": "2024-04-05T08:42:59.267Z",
      "content": "<p>Very Helpful!!</p>",
      "rawMarkdown": "Very Helpful!!",
      "votes": 1
    },
    {
      "id": 2725003,
      "postDate": "2024-03-31T09:22:53.123Z",
      "content": "<p>Great explanation!</p>",
      "rawMarkdown": "Great explanation!",
      "votes": 1
    },
    {
      "id": 2714150,
      "postDate": "2024-03-24T18:01:20.250Z",
      "content": "<p>Great explanation…</p>",
      "rawMarkdown": "Great explanation...",
      "votes": 1
    },
    {
      "id": 2701005,
      "postDate": "2024-03-16T19:54:29.850Z",
      "content": "<p>Very helpful！</p>",
      "rawMarkdown": "Very helpful！",
      "votes": 1
    },
    {
      "id": 2695983,
      "postDate": "2024-03-14T01:31:55.627Z",
      "content": "<p>Great explanation of the dataset！</p>",
      "rawMarkdown": "Great explanation of the dataset！",
      "votes": 1,
      "replies": [
        {
          "id": 2695995,
          "postDate": "2024-03-14T01:57:12.477Z",
          "content": "<p>Thanks Y.S !</p>",
          "rawMarkdown": "Thanks Y.S !"
        }
      ]
    },
    {
      "id": 2684979,
      "postDate": "2024-03-06T22:13:59.137Z",
      "content": "<p>Thankyou so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your detailed and dedicated explanation. I read all the comments, but still have 1 doubt unanswered. If we can create create the 50seconds spectrogram data using the EEG data, then why are we given both EEG and corresponding Spectrogram data? Does it have to be our choice on which data to train our model..?</p>",
      "rawMarkdown": "Thankyou so much @cdeotte for your detailed and dedicated explanation. I read all the comments, but still have 1 doubt unanswered. If we can create create the 50seconds spectrogram data using the EEG data, then why are we given both EEG and corresponding Spectrogram data? Does it have to be our choice on which data to train our model..?",
      "votes": 1,
      "replies": [
        {
          "id": 2685019,
          "postDate": "2024-03-06T23:28:00.613Z",
          "content": "<p>The given spectrograms are 10 minute low resolution windows whereas the given EEG waveforms are 50 second high resolution windows. One is like a zoom out view and one is like the zoom in view. We cannot create each from the other because of the difference in time window and difference in resolution. Both have something unique to contribute to our solution.</p>\n<p>(The rows of given spectrogram parquets are 2 seconds each whereas the rows of given EEG waveforms are 1/200 seconds each. That is why i say low-res and hi-res).</p>",
          "rawMarkdown": "The given spectrograms are 10 minute low resolution windows whereas the given EEG waveforms are 50 second high resolution windows. One is like a zoom out view and one is like the zoom in view. We cannot create each from the other because of the difference in time window and difference in resolution. Both have something unique to contribute to our solution.\n\n(The rows of given spectrogram parquets are 2 seconds each whereas the rows of given EEG waveforms are 1/200 seconds each. That is why i say low-res and hi-res).",
          "votes": 1
        }
      ]
    },
    {
      "id": 2660610,
      "postDate": "2024-02-20T17:44:24.243Z",
      "content": "<p>why do different EEGs point to the same spectrogram? I mean the spectrograms are created by doing FFT over overlapping windows, wouldn't that have at some point overlapping windows between two different EEGs( different samples essentially), though I'm pretty sure if that's the case its not at all a big deal… just trying to understand how they came up with it, generally you have EEG and you have a spectrogram for that EEG alone… Also if the reason is faster loading because you have multiple spectrograms we could do better (as <a href=\"https://www.kaggle.com/chris\" target=\"_blank\">@chris</a> did in the custom spectrogram creator at the end)…just spit balling</p>",
      "rawMarkdown": "why do different EEGs point to the same spectrogram? I mean the spectrograms are created by doing FFT over overlapping windows, wouldn't that have at some point overlapping windows between two different EEGs( different samples essentially), though I'm pretty sure if that's the case its not at all a big deal... just trying to understand how they came up with it, generally you have EEG and you have a spectrogram for that EEG alone... Also if the reason is faster loading because you have multiple spectrograms we could do better (as @chris did in the custom spectrogram creator at the end)...just spit balling",
      "votes": 1,
      "replies": [
        {
          "id": 2660833,
          "postDate": "2024-02-20T20:35:17.627Z",
          "content": "<p>EEG-1 and EEG-2 may be eegs for the same patient-A for time 12:00p and time 3:00p respectively. Then Spectrogram-1 is the entire day for patient-A from 11a thru 6p. So both EEG-1 and EEG-2 point to the same Spectrogram-1 file. And EEG-1 uses one offset and EEG-2 uses a different offset.</p>",
          "rawMarkdown": "EEG-1 and EEG-2 may be eegs for the same patient-A for time 12:00p and time 3:00p respectively. Then Spectrogram-1 is the entire day for patient-A from 11a thru 6p. So both EEG-1 and EEG-2 point to the same Spectrogram-1 file. And EEG-1 uses one offset and EEG-2 uses a different offset.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2655389,
      "postDate": "2024-02-16T20:27:27.040Z",
      "content": "<p>Hi! I am a bit confused if a single EEG parquet file contains information only from a single EEG_ID or if there could be multiple EEG_IDs associated with it. The latter seems unlikely to me but I cannot find any concrete answer on this either. Thank you in advance! :)</p>",
      "rawMarkdown": "Hi! I am a bit confused if a single EEG parquet file contains information only from a single EEG_ID or if there could be multiple EEG_IDs associated with it. The latter seems unlikely to me but I cannot find any concrete answer on this either. Thank you in advance! :)",
      "votes": 1,
      "replies": [
        {
          "id": 2655395,
          "postDate": "2024-02-16T20:30:41.560Z",
          "content": "<p>There is information about a single eeg_id. But can contain multiple subsamples with sub_eeg_id refering to a specific windows starting on eeg_offset.</p>",
          "rawMarkdown": "There is information about a single eeg_id. But can contain multiple subsamples with sub_eeg_id refering to a specific windows starting on eeg_offset.",
          "votes": 1
        },
        {
          "id": 2655429,
          "postDate": "2024-02-16T21:02:30.597Z",
          "content": "<p>Each EEG parquet file only contains information about one <code>eeg_id</code>. Then in <code>train.csv</code> file, many rows have the same <code>eeg_id</code> but access a different time window within the associated EEG parquet file.</p>",
          "rawMarkdown": "Each EEG parquet file only contains information about one `eeg_id`. Then in `train.csv` file, many rows have the same `eeg_id` but access a different time window within the associated EEG parquet file.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2654291,
      "postDate": "2024-02-16T03:38:33.970Z",
      "content": "<p>How are you saying the prediction should be done for 10 seconds. The test EEG has a size of 50 seconds and unique ID. And the  submission also requires predicting each eeg_id with 50 seconds of test EEG. I don't see anywhere that the prediction was done for the 10 seconds. (t-5 to t+5, where t is the middle time stamp). Can you explain that? </p>",
      "rawMarkdown": "How are you saying the prediction should be done for 10 seconds. The test EEG has a size of 50 seconds and unique ID. And the  submission also requires predicting each eeg_id with 50 seconds of test EEG. I don't see anywhere that the prediction was done for the 10 seconds. (t-5 to t+5, where t is the middle time stamp). Can you explain that? ",
      "votes": 1,
      "replies": [
        {
          "id": 2654544,
          "postDate": "2024-02-16T07:58:05.697Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data</a><br>\n\"Files<br>\ntrain.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and <strong>labeled the central 10 seconds.</strong> Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\"</p>",
          "rawMarkdown": "https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\n\"Files\ntrain.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and **labeled the central 10 seconds.** Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\"",
          "votes": 2
        },
        {
          "id": 2654868,
          "postDate": "2024-02-16T13:43:21.030Z",
          "content": "<p>We input the 50 second EEG and the 10 minute Spectrogram into our model. The model trains with the 6 target columns. The 6 target columns describe the central 10 seconds from both inputs. Therefore the model will learn during training that the most important region is the central 10 seconds. Hence the model learns to focus on the middle 10 seconds itself (and uses the extra time data when needed).</p>",
          "rawMarkdown": "We input the 50 second EEG and the 10 minute Spectrogram into our model. The model trains with the 6 target columns. The 6 target columns describe the central 10 seconds from both inputs. Therefore the model will learn during training that the most important region is the central 10 seconds. Hence the model learns to focus on the middle 10 seconds itself (and uses the extra time data when needed).",
          "votes": 3,
          "replies": [
            {
              "id": 2679457,
              "postDate": "2024-03-03T14:47:08.663Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for all your useful contributions! Somewhat related question. I notice that the submission must contain a prediction for each eeg_id. In the train set, each eeg_id is often repeated (multiple rows per eeg_id). Is the structure of the test set going to be the same? Because this means that the model needs to make one prediction for multiple rows. Or will the test set have each eeg_id appearing only once? Im a bit confused about this. </p>",
              "rawMarkdown": "@cdeotte Thank you for all your useful contributions! Somewhat related question. I notice that the submission must contain a prediction for each eeg_id. In the train set, each eeg_id is often repeated (multiple rows per eeg_id). Is the structure of the test set going to be the same? Because this means that the model needs to make one prediction for multiple rows. Or will the test set have each eeg_id appearing only once? Im a bit confused about this. ",
              "votes": 1
            },
            {
              "id": 2679607,
              "postDate": "2024-03-03T16:22:48.827Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/ketanjaltare\" target=\"_blank\">@ketanjaltare</a> the test set does not repeat <code>eeg_id</code>. In the test set each <code>egg_id</code> only appears once. This is stated in this competition's data page:</p>\n<blockquote>\n  <p>Metadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.</p>\n</blockquote>\n<p>i.e. we don't need <code>eeg_label_offset_seconds</code> and we don't need <code>spectogram_label_offset_seconds</code> because each test eeq parquet is unique and only contains 10,000 rows (and each test spectrogram parquet is only 300 rows)</p>",
              "rawMarkdown": "Hi @ketanjaltare the test set does not repeat `eeg_id`. In the test set each `egg_id` only appears once. This is stated in this competition's data page:\n\n>Metadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.\n\ni.e. we don't need `eeg_label_offset_seconds` and we don't need `spectogram_label_offset_seconds` because each test eeq parquet is unique and only contains 10,000 rows (and each test spectrogram parquet is only 300 rows)"
            },
            {
              "id": 2680016,
              "postDate": "2024-03-03T22:42:09.123Z",
              "content": "<p>Great! Thanks a lot. </p>",
              "rawMarkdown": "Great! Thanks a lot. "
            },
            {
              "id": 2713267,
              "postDate": "2024-03-24T04:42:55.707Z",
              "content": "<blockquote>\n  <p>Therefore the model will learn during training that the most important region is the central 10 seconds. Hence the model learns to focus on the middle 10 seconds itself (and uses the extra time data when needed).</p>\n</blockquote>\n<p>I dont get it. model doesnt really get the fact that votes are based on middle 10 seconds. we just know it. can u explain it more?</p>",
              "rawMarkdown": ">  Therefore the model will learn during training that the most important region is the central 10 seconds. Hence the model learns to focus on the middle 10 seconds itself (and uses the extra time data when needed).\n\nI dont get it. model doesnt really get the fact that votes are based on middle 10 seconds. we just know it. can u explain it more?\n\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2652235,
      "postDate": "2024-02-14T16:29:25.353Z",
      "content": "<p>Is it true that all data contained in the EEG parquet files can be constructed from the spectogram parquet files? If so why are they both given, and why not only the spectrogram files?</p>",
      "rawMarkdown": "Is it true that all data contained in the EEG parquet files can be constructed from the spectogram parquet files? If so why are they both given, and why not only the spectrogram files?",
      "votes": 1,
      "replies": [
        {
          "id": 2652271,
          "postDate": "2024-02-14T16:48:34.773Z",
          "content": "<p>No. The provided Spectrogram parquets have sampling rate of 0.5 measurements per second. Whereas the EEG parquets have sampling rate of 200 measurements per second. Additionally, I don't think we can convert spectrograms back to waveforms. (I think some information like phase is lost).</p>",
          "rawMarkdown": "No. The provided Spectrogram parquets have sampling rate of 0.5 measurements per second. Whereas the EEG parquets have sampling rate of 200 measurements per second. Additionally, I don't think we can convert spectrograms back to waveforms. (I think some information like phase is lost).",
          "votes": 2,
          "replies": [
            {
              "id": 2683628,
              "postDate": "2024-03-06T05:58:26.197Z",
              "content": "<p>Amazing content in this discussion page. My question is, can the given spectograms be created from the given EEGs?</p>",
              "rawMarkdown": "Amazing content in this discussion page. My question is, can the given spectograms be created from the given EEGs?",
              "votes": 1
            },
            {
              "id": 2683844,
              "postDate": "2024-03-06T08:53:33.810Z",
              "content": "<p>The given spectrograms span a 10 minute window while the given EEGs span a 50 second window. Therefore the middle 50 seconds of the given spectrograms can be created but not the additional 550 seconds.</p>",
              "rawMarkdown": "The given spectrograms span a 10 minute window while the given EEGs span a 50 second window. Therefore the middle 50 seconds of the given spectrograms can be created but not the additional 550 seconds.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2650274,
      "postDate": "2024-02-13T11:09:48.847Z",
      "content": "<p>Hi I am new to the competition and I am here to learn new things as fresher. But still I am in great fear of where to start. Is there any one could help me.</p>",
      "rawMarkdown": "Hi I am new to the competition and I am here to learn new things as fresher. But still I am in great fear of where to start. Is there any one could help me.",
      "votes": 1,
      "replies": [
        {
          "id": 2650575,
          "postDate": "2024-02-13T14:40:15.493Z",
          "content": "<p>Welcome <a href=\"https://www.kaggle.com/jahnavikurugundla\" target=\"_blank\">@jahnavikurugundla</a> . I suggest beginning with a starter notebook (either mine or someone else's). Then try to improve the CV score and LB score by changing things. Note that all 3 of my stater notebooks can achieve much better CV score and LB score.</p>",
          "rawMarkdown": "Welcome @jahnavikurugundla . I suggest beginning with a starter notebook (either mine or someone else's). Then try to improve the CV score and LB score by changing things. Note that all 3 of my stater notebooks can achieve much better CV score and LB score.",
          "votes": 6,
          "replies": [
            {
              "id": 2650757,
              "postDate": "2024-02-13T17:29:08.863Z",
              "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> sir. It means a lot to me. Thank you sir</p>",
              "rawMarkdown": "Thank you so much @cdeotte sir. It means a lot to me. Thank you sir",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2650052,
      "postDate": "2024-02-13T07:53:52.080Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for the clear explanation. <br>\nI have a question related to inference. As I am aware, there is a one-to-many mapping between the kaggle-spectrogram and the EEG-signals in train.csv, in the test.csv we don't have the time_offset for kaggle-spectrogram then how do we select the window of kaggle-spectrogram to EEG-signals or there is no one-to-many mapping in test.csv?</p>",
      "rawMarkdown": "Hi @cdeotte, thanks for the clear explanation. \nI have a question related to inference. As I am aware, there is a one-to-many mapping between the kaggle-spectrogram and the EEG-signals in train.csv, in the test.csv we don't have the time_offset for kaggle-spectrogram then how do we select the window of kaggle-spectrogram to EEG-signals or there is no one-to-many mapping in test.csv?",
      "votes": 1,
      "replies": [
        {
          "id": 2650215,
          "postDate": "2024-02-13T10:13:19.503Z",
          "content": "<p>Test are single samples. With a single label for middle timestamp.</p>\n<p>EDIT: To clarify, offset in 0.</p>",
          "rawMarkdown": "Test are single samples. With a single label for middle timestamp.\n\nEDIT: To clarify, offset in 0.",
          "votes": 1,
          "replies": [
            {
              "id": 2650225,
              "postDate": "2024-02-13T10:21:10.200Z",
              "content": "<p>Thanks !!!</p>",
              "rawMarkdown": "Thanks !!!"
            }
          ]
        }
      ]
    },
    {
      "id": 2647292,
      "postDate": "2024-02-11T13:20:03.817Z",
      "content": "<p>Impressive work. I express my gratitude to you. 🙏</p>",
      "rawMarkdown": "Impressive work. I express my gratitude to you. 🙏",
      "votes": 1
    },
    {
      "id": 2635357,
      "postDate": "2024-02-04T10:52:07.923Z",
      "content": "<p>Great work sir</p>",
      "rawMarkdown": "Great work sir",
      "votes": 1
    },
    {
      "id": 2629759,
      "postDate": "2024-02-01T02:31:34.380Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for all the EDA work you did. I'm very confused about the way the offset seconds are labeled in both the spectrograms and eeg. Consider the attached table which I made by printing out all the rows in which <code>eeg_id</code> == 3638862953. If we are supposed to index from the offset label to 50 seconds after for the eeg, then that would produce a lot of overlap. For example, do I start at the 42 second offset mark and index up to the 94 second offset mark? If so, I would have to overlap with <code>eeg_sub_id</code> 2,3,4, and 5. There would be even more overlap if I were to index with 10 minutes on the spectrogram, because the <code>spectrogram_label_offset_seconds</code> in this example are the same as the <code>eeg_offset_label_seconds</code> I'm just not sure how to index this effectively? How am I supposed to deduce which spectrograms/eegs the neurologists actually reviewed? Thank you anyone who might be able to clarify. </p>\n<p>Update: I noticed is that this eeg has 10000 rows, which means it is exactly 50 seconds long. I guess this means we might want to err on the lowest offset label in each train.csv file? This is because if you start any later in this the 3638862953 eeg, you won't be able to get the full 50 seconds.</p>",
      "rawMarkdown": "Thank you @cdeotte for all the EDA work you did. I'm very confused about the way the offset seconds are labeled in both the spectrograms and eeg. Consider the attached table which I made by printing out all the rows in which `eeg_id` == 3638862953. If we are supposed to index from the offset label to 50 seconds after for the eeg, then that would produce a lot of overlap. For example, do I start at the 42 second offset mark and index up to the 94 second offset mark? If so, I would have to overlap with `eeg_sub_id` 2,3,4, and 5. There would be even more overlap if I were to index with 10 minutes on the spectrogram, because the `spectrogram_label_offset_seconds` in this example are the same as the `eeg_offset_label_seconds` I'm just not sure how to index this effectively? How am I supposed to deduce which spectrograms/eegs the neurologists actually reviewed? Thank you anyone who might be able to clarify. \n\nUpdate: I noticed is that this eeg has 10000 rows, which means it is exactly 50 seconds long. I guess this means we might want to err on the lowest offset label in each train.csv file? This is because if you start any later in this the 3638862953 eeg, you won't be able to get the full 50 seconds.",
      "votes": 1
    },
    {
      "id": 2627292,
      "postDate": "2024-01-30T15:54:45.417Z",
      "content": "<p>Thanks so much for laying out these details clearly and visually!<br>\nMy inner pedant (\"one who unduly emphasizes minutiae\") wants to note that there is no \"middle 10 seconds\" of data in the spectrograms: The middle of the 300 spectra is between the 150th and 151st, so there is actually no (or two) middle 10 seconds 😯  There are middle 8 seconds and middle 12 seconds sections, though 😃<br>\nLess pedantly, unlike other cases where we had to determine the event center, here it is given. But I wonder how accurate/meaningful it is? If it is just an approximate location (is it <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> ?) then the data analysis method should be insensitive to few-seconds (or more?) time shifts; and time shifts could be used as a data augmentation. Doing image analysis on the whole spectrogram is probably insensitive to time shifts if shifts are effectively already in the data.</p>",
      "rawMarkdown": "Thanks so much for laying out these details clearly and visually!\nMy inner pedant (\"one who unduly emphasizes minutiae\") wants to note that there is no \"middle 10 seconds\" of data in the spectrograms: The middle of the 300 spectra is between the 150th and 151st, so there is actually no (or two) middle 10 seconds 😯  There are middle 8 seconds and middle 12 seconds sections, though 😃\nLess pedantly, unlike other cases where we had to determine the event center, here it is given. But I wonder how accurate/meaningful it is? If it is just an approximate location (is it @awsaf49 ?) then the data analysis method should be insensitive to few-seconds (or more?) time shifts; and time shifts could be used as a data augmentation. Doing image analysis on the whole spectrogram is probably insensitive to time shifts if shifts are effectively already in the data.",
      "votes": 1
    },
    {
      "id": 2626779,
      "postDate": "2024-01-30T09:01:52.163Z",
      "content": "<p>Thank you for your explanation, that makes me clearer.</p>",
      "rawMarkdown": "Thank you for your explanation, that makes me clearer.",
      "votes": 1
    },
    {
      "id": 2624988,
      "postDate": "2024-01-29T06:10:52.623Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for making the competition a lot clearer.<br>\nI had a doubt, shouldn't the diagram be like this in <code>Spectrogram parquet files</code> section<br>\ni.e. the <code>eeg_label_offset_seconds</code> should be <code>spectrogram_label_offset_seconds</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4359632%2F559a7782fd26aba6d1569a58d0500ba8%2Fdownload.png?generation=1706508510916891&amp;alt=media\"></p>",
      "rawMarkdown": "Hi @cdeotte, thanks for making the competition a lot clearer.\nI had a doubt, shouldn't the diagram be like this in `Spectrogram parquet files` section\ni.e. the `eeg_label_offset_seconds` should be `spectrogram_label_offset_seconds`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4359632%2F559a7782fd26aba6d1569a58d0500ba8%2Fdownload.png?generation=1706508510916891&alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 2625239,
          "postDate": "2024-01-29T09:36:16.437Z",
          "content": "<p><a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> Yes, you are correct. Thank you for finding this typo. I updated my discussion post.</p>",
          "rawMarkdown": "@atharvaingle Yes, you are correct. Thank you for finding this typo. I updated my discussion post.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2624946,
      "postDate": "2024-01-29T05:12:54.563Z",
      "content": "<p>Very informative, thanks for the explanation! </p>",
      "rawMarkdown": "Very informative, thanks for the explanation! ",
      "votes": 1
    },
    {
      "id": 2624901,
      "postDate": "2024-01-29T04:30:10.583Z",
      "content": "<p>Thanks a lot, Great explanation of the dataset.</p>",
      "rawMarkdown": "Thanks a lot, Great explanation of the dataset.",
      "votes": 1
    },
    {
      "id": 2696311,
      "postDate": "2024-03-14T08:27:06.313Z",
      "content": "<p>Thank you for sharing the good analysis:)</p>\n<p>I have a question. Then, wouldn’t it be right to create a spectrogram for each offset and learn it?</p>\n<p>A label for 50 seconds is given based on the offset, and creating a spectrogram of more detailed parts seems to be a more accurate learning method.</p>\n<pre><code>meta_data = row[]\noffset_sec = (meta_data[])\neeg = original_eeg.iloc[offset_sec*:(offset_sec+)*].reset_index(drop=)\nmiddle = ((eeg)//)\neeg = eeg.iloc[middle-(*):middle+(*)] \n</code></pre>\n<p>Wouldn’t it be right to structure the dataset like this?</p>",
      "rawMarkdown": "Thank you for sharing the good analysis:)\n\nI have a question. Then, wouldn’t it be right to create a spectrogram for each offset and learn it?\n\nA label for 50 seconds is given based on the offset, and creating a spectrogram of more detailed parts seems to be a more accurate learning method.\n\n```python\nmeta_data = row[1]\noffset_sec = int(meta_data['eeg_label_offset_seconds'])\neeg = original_eeg.iloc[offset_sec*200:(offset_sec+50)*200].reset_index(drop=True)\nmiddle = int(len(eeg)//2)\neeg = eeg.iloc[middle-(5*200):middle+(5*200)] # 2000\n ```\n\nWouldn’t it be right to structure the dataset like this?",
      "votes": 2,
      "replies": [
        {
          "id": 2696415,
          "postDate": "2024-03-14T10:00:26.403Z",
          "content": "<p>We can try this. We must experiment with preprocessing data in different ways to see what achieves the best CV score and LB score.</p>",
          "rawMarkdown": "We can try this. We must experiment with preprocessing data in different ways to see what achieves the best CV score and LB score.",
          "votes": 1,
          "replies": [
            {
              "id": 2696775,
              "postDate": "2024-03-14T14:33:54.533Z",
              "content": "<p>you're right. That is the fun and flower of data analysis. I will try an experiment and leave the results. <br>\nThank you for reply ! 👍</p>",
              "rawMarkdown": "you're right. That is the fun and flower of data analysis. I will try an experiment and leave the results. \nThank you for reply ! 👍",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2622218,
      "postDate": "2024-01-27T10:46:35.157Z",
      "content": "<p>Many thanks for the great description. <br>\nIt was very helpful to get started with the competition!</p>",
      "rawMarkdown": "Many thanks for the great description. \nIt was very helpful to get started with the competition!",
      "votes": 1
    },
    {
      "id": 2620648,
      "postDate": "2024-01-26T09:06:07.147Z",
      "content": "<p>One point i couldn't understand İf EGG signals measured 50 seconds for each patients. Why spectrograms shows 10 minute. I mean to create 10 minute spectrogram shouldn't we use 10 minute of EEG signals? So these are 2 different measurements?</p>",
      "rawMarkdown": "One point i couldn't understand İf EGG signals measured 50 seconds for each patients. Why spectrograms shows 10 minute. I mean to create 10 minute spectrogram shouldn't we use 10 minute of EEG signals? So these are 2 different measurements?",
      "votes": 1,
      "replies": [
        {
          "id": 2620666,
          "postDate": "2024-01-26T09:24:16.390Z",
          "content": "<p>The middle 50 seconds is the same for each. If we take the middle of Kaggle spectrogram which is <code>[T-25:T+25]</code> then this is the same as our EEG spectrogram. The Kaggle provided spectrogram has more information. It has </p>\n<ul>\n<li>beginning 4 minutes 5 seconds</li>\n<li>middle 50 seconds</li>\n<li>ending 4 minutes 5 seconds</li>\n</ul>\n<p>The first and last bullet point is information that Kaggle spectrogram includes which EEG spectrogram does not include.</p>",
          "rawMarkdown": "The middle 50 seconds is the same for each. If we take the middle of Kaggle spectrogram which is `[T-25:T+25]` then this is the same as our EEG spectrogram. The Kaggle provided spectrogram has more information. It has \n* beginning 4 minutes 5 seconds\n* middle 50 seconds\n* ending 4 minutes 5 seconds\n\nThe first and last bullet point is information that Kaggle spectrogram includes which EEG spectrogram does not include.",
          "votes": 8,
          "replies": [
            {
              "id": 2651564,
              "postDate": "2024-02-14T08:41:45.383Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . Thank you for this explanation. I wanted to clear something. The sentence \"The first and last bullet point is information that Kaggle spectrogram includes which <strong>EEG spectrogram</strong> does not include.\" shouldn't be like this \"The first and last bullet point is information that Kaggle spectrogram includes which <strong>EEG waveform</strong> does not include.\" ? Because kaggle didn't provided EEG spectrograms and you made them in one of your notebook. Also it would be nice if you could also break down your EEG spectrogram in the same manner. </p>",
              "rawMarkdown": "Hi @cdeotte . Thank you for this explanation. I wanted to clear something. The sentence \"The first and last bullet point is information that Kaggle spectrogram includes which **EEG spectrogram** does not include.\" shouldn't be like this \"The first and last bullet point is information that Kaggle spectrogram includes which **EEG waveform** does not include.\" ? Because kaggle didn't provided EEG spectrograms and you made them in one of your notebook. Also it would be nice if you could also break down your EEG spectrogram in the same manner. "
            }
          ]
        }
      ]
    },
    {
      "id": 2620394,
      "postDate": "2024-01-26T05:35:37.050Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Sorry if it's a silly question. Why does the input shape for EfficientNet have to be (128, 256, x) when the base model is loaded without the top?</p>",
      "rawMarkdown": "@cdeotte Sorry if it's a silly question. Why does the input shape for EfficientNet have to be (128, 256, x) when the base model is loaded without the top?",
      "votes": 1,
      "replies": [
        {
          "id": 2620661,
          "postDate": "2024-01-26T09:21:56.140Z",
          "content": "<p>The input shape to EfficientNet is <strong>not</strong> (128,256,x). The input shape is <code>None</code> as seen here <code>base_model = efn.EfficientNetB0(include_top=False, weights=None, input_shape=None)</code>. This means that we can input any size into EfficientNet. Remember that EfficientNet is fully convolutional. That means we can train with any size and infer with any size including a different size. For example we can train with <code>(32,32)</code> and then later infer with <code>(256,256)</code>. We can also change from batch to batch. Or epoch to epoch. This is advantage of fully convolutional model.</p>\n<p>In my starter code, the input to the model is (128,256,8). Then we reshape it into (512,512,3) and input (512,512) into EfficientNet. The model is EfficientNet plus the layers which reshape input, plus the head which performs global average pooling and the output layer of softmax with 6 units.</p>",
          "rawMarkdown": "The input shape to EfficientNet is **not** (128,256,x). The input shape is `None` as seen here `base_model = efn.EfficientNetB0(include_top=False, weights=None, input_shape=None)`. This means that we can input any size into EfficientNet. Remember that EfficientNet is fully convolutional. That means we can train with any size and infer with any size including a different size. For example we can train with `(32,32)` and then later infer with `(256,256)`. We can also change from batch to batch. Or epoch to epoch. This is advantage of fully convolutional model.\n\nIn my starter code, the input to the model is (128,256,8). Then we reshape it into (512,512,3) and input (512,512) into EfficientNet. The model is EfficientNet plus the layers which reshape input, plus the head which performs global average pooling and the output layer of softmax with 6 units.",
          "votes": 4,
          "replies": [
            {
              "id": 2621745,
              "postDate": "2024-01-27T00:15:49.300Z",
              "content": "<p>Really appreciate your reply! Sorry for my bad English. Why would you set your input shape as (128,256,8) and later go through the reshaping and cropping painstakingly but not set it as (300,100,4) which is compatible with the spectrogram dataset?</p>",
              "rawMarkdown": "Really appreciate your reply! Sorry for my bad English. Why would you set your input shape as (128,256,8) and later go through the reshaping and cropping painstakingly but not set it as (300,100,4) which is compatible with the spectrogram dataset?"
            },
            {
              "id": 2621795,
              "postDate": "2024-01-27T02:49:41.930Z",
              "content": "<p>CNN like EfficientNet require that their image input dimensions be multiples of 32. So we can't have height 300 for example, we need either 288 or 320. Likewise we cannot have 100 but rather 96 or 128.</p>\n<p>My dataloader, inserts Kaggle's 300x100 into 128x256 (after taking transpose =&gt; 100x300 i.e. frequency x time). Also my EfficientNet starter uses spectrograms made by me. And i make my spectrograms to have dimensions 128x256 (frequency x time)</p>\n<p>Hence I put all 8 spectrograms (4 of Kaggle, 4 of mine) into an input of 128x256x8. Then we can experiment how to reshape it. For example, perhaps my starter notebook is not best. Maybe we want to stack the spectrograms into 512x256x2 instead of 512x512x1 so that the model can compare pairs of spectrograms. By making the input a generic 128x256x8, we are free to reshape however we want to achieve best CV score and LB score.</p>",
              "rawMarkdown": "CNN like EfficientNet require that their image input dimensions be multiples of 32. So we can't have height 300 for example, we need either 288 or 320. Likewise we cannot have 100 but rather 96 or 128.\n\nMy dataloader, inserts Kaggle's 300x100 into 128x256 (after taking transpose => 100x300 i.e. frequency x time). Also my EfficientNet starter uses spectrograms made by me. And i make my spectrograms to have dimensions 128x256 (frequency x time)\n\nHence I put all 8 spectrograms (4 of Kaggle, 4 of mine) into an input of 128x256x8. Then we can experiment how to reshape it. For example, perhaps my starter notebook is not best. Maybe we want to stack the spectrograms into 512x256x2 instead of 512x512x1 so that the model can compare pairs of spectrograms. By making the input a generic 128x256x8, we are free to reshape however we want to achieve best CV score and LB score.",
              "votes": 8
            },
            {
              "id": 2621839,
              "postDate": "2024-01-27T04:35:57.160Z",
              "content": "<p>Thank you so much Chris for the detailed explanation! So glad just learned a critical point about EfficientNet.</p>",
              "rawMarkdown": "Thank you so much Chris for the detailed explanation! So glad just learned a critical point about EfficientNet.",
              "votes": 1
            },
            {
              "id": 2623083,
              "postDate": "2024-01-28T00:17:45.477Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in your example (alternative) size will be 512x256x2 instead of 512x512x3? Is it not expected a 3 channels image?</p>",
              "rawMarkdown": "@cdeotte in your example (alternative) size will be 512x256x2 instead of 512x512x3? Is it not expected a 3 channels image?\n",
              "votes": 1
            },
            {
              "id": 2624037,
              "postDate": "2024-01-28T14:37:45.020Z",
              "content": "<p>Yes <a href=\"https://www.kaggle.com/alexandervega\" target=\"_blank\">@alexandervega</a> for TensorFlow EffNet we need 3 channels. Note that when using timm and PyTorch, we can have any number of channels by setting parameter <code>in_chans</code>: </p>\n<pre><code>timm.create_model(\n        =model_name, =pretrained,\n        =num_classes, =in_channels)\n</code></pre>\n<p>In TF, after making 512x256x2, we can just add a third channel of all zeros <code>512x256x2 =&gt; 512x256x3</code>. Or we can repeat one of the two channels. Or we can take the difference of the first two channels as the third channel. Here is pseudo code:</p>\n<pre><code> = img_old[:,:,] - img_old[:,:,]\nimg_new = Concatenate(axis=)([img_old,])\n</code></pre>",
              "rawMarkdown": "Yes @alexandervega for TensorFlow EffNet we need 3 channels. Note that when using timm and PyTorch, we can have any number of channels by setting parameter `in_chans`: \n\n    timm.create_model(\n            model_name=model_name, pretrained=pretrained,\n            num_classes=num_classes, in_chans=in_channels)\n\nIn TF, after making 512x256x2, we can just add a third channel of all zeros `512x256x2 => 512x256x3`. Or we can repeat one of the two channels. Or we can take the difference of the first two channels as the third channel. Here is pseudo code:\n\n    new_channel = img_old[:,:,0] - img_old[:,:,1]\n    img_new = Concatenate(axis=-1)([img_old,new_channel])",
              "votes": 4
            },
            {
              "id": 2624547,
              "postDate": "2024-01-28T19:58:25.623Z",
              "content": "<p>ok, thanks!</p>",
              "rawMarkdown": "ok, thanks!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2617816,
      "postDate": "2024-01-24T13:22:18.400Z",
      "content": "<p>Impressive work. It is inspiring.</p>",
      "rawMarkdown": "Impressive work. It is inspiring.",
      "votes": 1
    },
    {
      "id": 2688344,
      "postDate": "2024-03-09T07:06:53.667Z",
      "content": "<p>Hi, I have two inquiries:</p>\n<ol>\n<li><p>Could you explain the integration of both EEG data and spectrogram into the processes of model training?</p></li>\n<li><p>Is it imperative to apply filtering, denoising, or normalization techniques to the provided readings?</p></li>\n</ol>",
      "rawMarkdown": "Hi, I have two inquiries:\n\n1. Could you explain the integration of both EEG data and spectrogram into the processes of model training?\n\n2. Is it imperative to apply filtering, denoising, or normalization techniques to the provided readings?",
      "votes": 2,
      "replies": [
        {
          "id": 2688345,
          "postDate": "2024-03-09T07:08:37.427Z",
          "content": "<p>Hi. Check out my EfficientNet starter which uses both EEG and Spectrogram and normalization technique. And check out my WaveNet starter which uses filtering, denoising, and normalization. </p>",
          "rawMarkdown": "Hi. Check out my EfficientNet starter which uses both EEG and Spectrogram and normalization technique. And check out my WaveNet starter which uses filtering, denoising, and normalization. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2613300,
      "postDate": "2024-01-22T02:04:03.023Z",
      "content": "<p>Thank you for the great read. Easy and well explained</p>",
      "rawMarkdown": "Thank you for the great read. Easy and well explained",
      "votes": 1
    },
    {
      "id": 2612941,
      "postDate": "2024-01-21T18:07:35.110Z",
      "content": "<p>thank you for the explanation and the example codes</p>",
      "rawMarkdown": "thank you for the explanation and the example codes",
      "votes": 1
    },
    {
      "id": 2612519,
      "postDate": "2024-01-21T12:50:39.987Z",
      "content": "<p>Thanks for the insight. Was a good read. Great explanation !!</p>",
      "rawMarkdown": "Thanks for the insight. Was a good read. Great explanation !!\n\n",
      "votes": 1
    },
    {
      "id": 2612064,
      "postDate": "2024-01-21T06:56:09.403Z",
      "content": "<p>Thanks for the work, its quite useful and enriching.</p>",
      "rawMarkdown": "Thanks for the work, its quite useful and enriching.",
      "votes": 1
    },
    {
      "id": 2611074,
      "postDate": "2024-01-20T14:49:10.133Z",
      "content": "<p>Thank you for your interesting information</p>",
      "rawMarkdown": "Thank you for your interesting information",
      "votes": 1
    },
    {
      "id": 2607462,
      "postDate": "2024-01-18T08:44:09.723Z",
      "content": "<p>Great！but why 10 sec here？</p>",
      "rawMarkdown": "Great！but why 10 sec here？",
      "votes": 1,
      "replies": [
        {
          "id": 2607640,
          "postDate": "2024-01-18T10:52:27.507Z",
          "content": "<p>This is said in competition description <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\" target=\"_blank\">here</a> in section data, subsection files train.csv</p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.</p>\n</blockquote>",
          "rawMarkdown": "This is said in competition description [here][1] in section data, subsection files train.csv\n>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data",
          "votes": 4,
          "replies": [
            {
              "id": 2610634,
              "postDate": "2024-01-20T08:41:48.447Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2610681,
              "postDate": "2024-01-20T09:19:42.020Z",
              "content": "<p>I believe it provides context. Imagine we want to know if someone has a cold. When we look at them and see their symptoms (i.e. runny nose etc) we are observing the \"10 seconds\", i.e the recent small window of the moment. Then we ask the person, how were you feeling this morning? How were you feeling yesterday.</p>\n<p>We use information from our observation of current 10 seconds, plus information from the past 24 hours to make a decision whether the person has a cold. So our decision is based on multiple time windows of information.</p>",
              "rawMarkdown": "I believe it provides context. Imagine we want to know if someone has a cold. When we look at them and see their symptoms (i.e. runny nose etc) we are observing the \"10 seconds\", i.e the recent small window of the moment. Then we ask the person, how were you feeling this morning? How were you feeling yesterday.\n\nWe use information from our observation of current 10 seconds, plus information from the past 24 hours to make a decision whether the person has a cold. So our decision is based on multiple time windows of information.",
              "votes": 3
            },
            {
              "id": 2611098,
              "postDate": "2024-01-20T14:58:59.813Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2611105,
              "postDate": "2024-01-20T15:01:33.243Z",
              "content": "<p>What do you mean \"without 10 seconds\". The middle 10 seconds of the EEG waveform is the \"central 10 seconds\" and the middle 10 seconds of the Spectrogram is the \"central 10 seconds\". So we do have them.</p>\n<p>If we want we can ignore all data except these two 10 seconds windows and build a model using only this limited data (if the extra data is bothering us).</p>",
              "rawMarkdown": "What do you mean \"without 10 seconds\". The middle 10 seconds of the EEG waveform is the \"central 10 seconds\" and the middle 10 seconds of the Spectrogram is the \"central 10 seconds\". So we do have them.\n\nIf we want we can ignore all data except these two 10 seconds windows and build a model using only this limited data (if the extra data is bothering us)."
            },
            {
              "id": 2612126,
              "postDate": "2024-01-21T07:58:51.690Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2604952,
      "postDate": "2024-01-16T19:02:27.053Z",
      "content": "<p>Thanks for the insight. Was a good read.</p>",
      "rawMarkdown": "Thanks for the insight. Was a good read.",
      "votes": 1
    },
    {
      "id": 2604052,
      "postDate": "2024-01-16T08:26:39.893Z",
      "content": "<p>hello sir you are great</p>",
      "rawMarkdown": "hello sir you are great",
      "votes": 1
    },
    {
      "id": 2602334,
      "postDate": "2024-01-15T04:39:38.270Z",
      "content": "<p>Thanks for your work. I now got to understand how data is given.</p>",
      "rawMarkdown": "Thanks for your work. I now got to understand how data is given.",
      "votes": 1
    },
    {
      "id": 2602247,
      "postDate": "2024-01-15T02:56:04.057Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> till now everyone imagine it, after seeing your well designed flow. All questions are answered mainly</p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.</p>\n</blockquote>\n<p>Thanks for sharing</p>",
      "rawMarkdown": "@cdeotte till now everyone imagine it, after seeing your well designed flow. All questions are answered mainly\n> The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\n\nThanks for sharing\n",
      "votes": 1
    },
    {
      "id": 2641848,
      "postDate": "2024-02-07T17:50:49.113Z",
      "content": "<p>Hi and thanks for the clear explanation. Just to check if I understood it correctly. <br>\nAt<br>\neeg = eeg.iloc[eeg_offset<em>200:(eeg_offset+50)</em>200]<br>\nthe 200 factor is because the EEG have been measured at 200 Hz so there is 200 rows per second measured?</p>\n<p>EDIT: My bad. Just answered question a couple comments up.</p>",
      "rawMarkdown": "Hi and thanks for the clear explanation. Just to check if I understood it correctly. \nAt\neeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\nthe 200 factor is because the EEG have been measured at 200 Hz so there is 200 rows per second measured?\n\nEDIT: My bad. Just answered question a couple comments up.",
      "votes": 2,
      "replies": [
        {
          "id": 2641861,
          "postDate": "2024-02-07T17:59:52.147Z",
          "content": "<p>Yes, your understanding is correct</p>",
          "rawMarkdown": "Yes, your understanding is correct",
          "votes": 1
        }
      ]
    },
    {
      "id": 2632774,
      "postDate": "2024-02-02T16:10:58.070Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, the link for the wavenet starter is taking me to the catboost starter</p>",
      "rawMarkdown": "Hi @cdeotte, the link for the wavenet starter is taking me to the catboost starter",
      "votes": 2,
      "replies": [
        {
          "id": 2632809,
          "postDate": "2024-02-02T16:33:09.373Z",
          "content": "<p>Thanks for letting me know. I fixed the link</p>",
          "rawMarkdown": "Thanks for letting me know. I fixed the link",
          "votes": 1
        }
      ]
    },
    {
      "id": 2625831,
      "postDate": "2024-01-29T16:42:40.297Z",
      "content": "<p>Thanks a lot Chris. this will help me hit the ground running!</p>",
      "rawMarkdown": "Thanks a lot Chris. this will help me hit the ground running!",
      "votes": 2
    },
    {
      "id": 2613649,
      "postDate": "2024-01-22T07:46:43.700Z",
      "content": "<p>Can anyone please tell me why the sum of votes for each rows are different?</p>",
      "rawMarkdown": "Can anyone please tell me why the sum of votes for each rows are different?",
      "votes": 2,
      "replies": [
        {
          "id": 2617760,
          "postDate": "2024-01-24T12:51:26.607Z",
          "content": "<p>I'm not sure exactly. This is the data that Kaggle provides. It seems that some data was evaluated by more doctors and some data was evaluated by less doctors.</p>",
          "rawMarkdown": "I'm not sure exactly. This is the data that Kaggle provides. It seems that some data was evaluated by more doctors and some data was evaluated by less doctors.",
          "votes": 4
        },
        {
          "id": 2634697,
          "postDate": "2024-02-03T23:59:25.720Z",
          "content": "<p>From the dataset description, <br>\n<code>Note that the test samples had between 3 and 20 annotators.</code><br>\nSeems that we just have different annotators for different samples.</p>",
          "rawMarkdown": " From the dataset description, \n\n```Note that the test samples had between 3 and 20 annotators.```\n\nSeems that we just have different annotators for different samples.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2607374,
      "postDate": "2024-01-18T07:39:15.370Z",
      "content": "<p>Thanks for the knowledge! Really Helpful!</p>",
      "rawMarkdown": "Thanks for the knowledge! Really Helpful!",
      "votes": 2
    },
    {
      "id": 2605063,
      "postDate": "2024-01-16T20:05:15.833Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nwhat is the relationship between the EEG and the spectrogram?  plz explain to me on dataset point of view</p>",
      "rawMarkdown": "Hi @cdeotte \nwhat is the relationship between the EEG and the spectrogram?  plz explain to me on dataset point of view",
      "votes": 2,
      "replies": [
        {
          "id": 2605069,
          "postDate": "2024-01-16T20:13:21.090Z",
          "content": "<p>A spectrogram is a visual representation of the EEG signals.</p>",
          "rawMarkdown": "A spectrogram is a visual representation of the EEG signals.",
          "votes": 3,
          "replies": [
            {
              "id": 2605161,
              "postDate": "2024-01-16T22:19:13.153Z",
              "content": "<p>In our dataset, for each eeg_id we have only one spectrogram_id is there but for every spectrogram_id, we have more than 1 eeg_id present.  any suggestion?</p>",
              "rawMarkdown": "In our dataset, for each eeg_id we have only one spectrogram_id is there but for every spectrogram_id, we have more than 1 eeg_id present.  any suggestion?",
              "votes": 1
            },
            {
              "id": 2608520,
              "postDate": "2024-01-18T21:07:40.270Z",
              "content": "<p>Decide whether you want to create 1 train sample per 1 unique eeg_id. Or create 1 train sample per 1 unique spectrogram id. Then make all your train samples. Then train your model.</p>",
              "rawMarkdown": "Decide whether you want to create 1 train sample per 1 unique eeg_id. Or create 1 train sample per 1 unique spectrogram id. Then make all your train samples. Then train your model.",
              "votes": 5
            },
            {
              "id": 2608815,
              "postDate": "2024-01-19T05:47:37.210Z",
              "content": "<p>So, In this way, we train our model.<br>\nIs there any way, we can form training data, that includes both eeg and spectrogram data? . <br>\nthank you for answering my question <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. </p>",
              "rawMarkdown": "So, In this way, we train our model.\nIs there any way, we can form training data, that includes both eeg and spectrogram data? . \nthank you for answering my question @cdeotte. ",
              "votes": -1
            },
            {
              "id": 2609079,
              "postDate": "2024-01-19T08:39:09.413Z",
              "content": "<p>Yes, there are many ways to combine them. This is one of the main challenge of this competition. How we combine them depends on our model. When using GBT (i.e ML models), we can just engineer features from both. When using DL models, we must be more creative how we combine them. </p>\n<p>Two ideas for DL are</p>\n<ul>\n<li>twin tower</li>\n<li>interaction</li>\n</ul>\n<p>In twin tower, our DL model has two separate \"towers\", where each \"tower\" extracts features from either data 1 or data 2. Then we add hidden layers (with optional self attention and/or CNN, RNN) to mix the two. In bullet point two, we start interaction between the two data in the beginning layers using cross features and/or attention.</p>",
              "rawMarkdown": "Yes, there are many ways to combine them. This is one of the main challenge of this competition. How we combine them depends on our model. When using GBT (i.e ML models), we can just engineer features from both. When using DL models, we must be more creative how we combine them. \n\nTwo ideas for DL are\n* twin tower\n* interaction\n\nIn twin tower, our DL model has two separate \"towers\", where each \"tower\" extracts features from either data 1 or data 2. Then we add hidden layers (with optional self attention and/or CNN, RNN) to mix the two. In bullet point two, we start interaction between the two data in the beginning layers using cross features and/or attention.",
              "votes": 6
            },
            {
              "id": 2609618,
              "postDate": "2024-01-19T15:52:28.993Z",
              "content": "<p>Thanks, Chris.  Have you implemented this DL type of approach?  <br>\nif yes Please share the notebook. I love to go through it once.</p>",
              "rawMarkdown": "Thanks, Chris.  Have you implemented this DL type of approach?  \nif yes Please share the notebook. I love to go through it once.",
              "votes": -1
            }
          ]
        }
      ]
    },
    {
      "id": 2602764,
      "postDate": "2024-01-15T10:19:48.847Z",
      "content": "<p>Thank you for clear explanation!<br>\nSo when calculating Center of spectrogram window(T), we need to add 'spectrogram_label_offset_seconds' to 300sec.<br>\nAnd also, predicting labels using [T-5, T+5].<br>\nIs it correct?</p>",
      "rawMarkdown": "Thank you for clear explanation!\nSo when calculating Center of spectrogram window(T), we need to add 'spectrogram_label_offset_seconds' to 300sec.\nAnd also, predicting labels using [T-5, T+5].\nIs it correct?",
      "votes": 2,
      "replies": [
        {
          "id": 2602807,
          "postDate": "2024-01-15T10:58:24.693Z",
          "content": "<p>This is a great question. We do not use <code>Center</code>, we use <code>Offset</code> which is provided in the <code>train.csv</code>. Here is code to get the EEG and Spectrogram for a given row:</p>\n<pre><code> = \n = \n = \n\n = pd.read_csv()\n = train.iloc[GET_ROW]\n\n = pd.read_parquet(f)\n = int( row.eeg_label_set_seconds )\n = eeg.iloc[eeg_set*:(eeg_set+)*]\n\n = pd.read_parquet(f)\n = int( row.spectrogram_label_set_seconds )\n = spectrogram.loc[(spectrogram.time&gt;=spec_set)\n                         &amp;(spectrogram.time&lt;spec_set+)]\n</code></pre>\n<p>I added this code to the original discussion post because others will want to know too. Thanks!</p>",
          "rawMarkdown": "This is a great question. We do not use `Center`, we use `Offset` which is provided in the `train.csv`. Here is code to get the EEG and Spectrogram for a given row:\n\n    GET_ROW = 0\n    EEG_PATH = 'train_eegs/'\n    SPEC_PATH = 'train_spectrograms/'\n\n    train = pd.read_csv('train.csv')\n    row = train.iloc[GET_ROW]\n\n    eeg = pd.read_parquet(f'{EEG_PATH}{row.eeg_id}.parquet')\n    eeg_offset = int( row.eeg_label_offset_seconds )\n    eeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\n\n    spectrogram = pd.read_parquet(f'{SPEC_PATH}{row.spectrogram_id}.parquet')\n    spec_offset = int( row.spectrogram_label_offset_seconds )\n    spectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                             &(spectrogram.time<spec_offset+600)]\n\nI added this code to the original discussion post because others will want to know too. Thanks!",
          "votes": 5,
          "replies": [
            {
              "id": 2602932,
              "postDate": "2024-01-15T12:59:23.360Z",
              "content": "<p>Thank you Update!<br>\neeg_offset<em>200:(eeg_offset+50)</em>200</p>\n<p>so, sampling frequency of eeg is 200hz??</p>",
              "rawMarkdown": "Thank you Update!\neeg_offset*200:(eeg_offset+50)*200\n\nso, sampling frequency of eeg is 200hz??",
              "votes": 1
            },
            {
              "id": 2602947,
              "postDate": "2024-01-15T13:11:36.480Z",
              "content": "<p>so, sampling frequency of eeg is 200hz??<br>\n→I got it!! This was described on the following page.<br>\n　Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Your post is always helpful.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data</a></p>",
              "rawMarkdown": "so, sampling frequency of eeg is 200hz??\n→I got it!! This was described on the following page.\n　Thank you @cdeotte. Your post is always helpful.\n[https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data](url)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2602445,
      "postDate": "2024-01-15T06:10:28.003Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> can you explain this part of your code in data generator <br>\n<code>\n           for k in range(4):\n                # EXTRACT 300 ROWS OF SPECTROGRAM\n                img = self.specs[row.spec_id][r:r+300,k*100:(k+1)*100].T</code><br>\n<strong>Edit</strong> : okay got it.<br>\nWhy only first 300 rows is being extracted only?</p>",
      "rawMarkdown": "@cdeotte can you explain this part of your code in data generator \n`\n           for k in range(4):\n                # EXTRACT 300 ROWS OF SPECTROGRAM\n                img = self.specs[row.spec_id][r:r+300,k*100:(k+1)*100].T`\n**Edit** : okay got it.\nWhy only first 300 rows is being extracted only?",
      "votes": 2,
      "replies": [
        {
          "id": 2602816,
          "postDate": "2024-01-15T11:03:46.047Z",
          "content": "<p>We extract 300 rows because each row is 2 seconds and we want 600 seconds which is 10 minutes.</p>",
          "rawMarkdown": "We extract 300 rows because each row is 2 seconds and we want 600 seconds which is 10 minutes.",
          "votes": 3,
          "replies": [
            {
              "id": 2602888,
              "postDate": "2024-01-15T12:10:46.663Z",
              "content": "<p>But we can sample 300 rows among number of rows. Is there any specific reason for choosing first 300 rows. And what about different sub samples.</p>",
              "rawMarkdown": "But we can sample 300 rows among number of rows. Is there any specific reason for choosing first 300 rows. And what about different sub samples.",
              "votes": 2
            },
            {
              "id": 2602891,
              "postDate": "2024-01-15T12:14:25.740Z",
              "content": "<p>We don't take the first 300 rows. We take rows starting from the variable <code>r</code> and the variable <code>r</code> is randomly chosen with <code>r = np.random.randint(row['min'], row['max']+1)//2</code>. This range of <code>r</code> is the entire spectrogram dataframe. All rows have a chance of being selected.</p>",
              "rawMarkdown": "We don't take the first 300 rows. We take rows starting from the variable `r` and the variable `r` is randomly chosen with `r = np.random.randint(row['min'], row['max']+1)//2`. This range of `r` is the entire spectrogram dataframe. All rows have a chance of being selected.",
              "votes": 2
            },
            {
              "id": 2602899,
              "postDate": "2024-01-15T12:23:36.587Z",
              "content": "<p>Okay, so for a given patient id you took only one spectrogram_id?</p>",
              "rawMarkdown": "Okay, so for a given patient id you took only one spectrogram_id?",
              "votes": 2
            },
            {
              "id": 2602903,
              "postDate": "2024-01-15T12:26:50.993Z",
              "content": "<p>No, for each <code>eeg_id</code>, i take only one spectrogram random crop. So I perform KFold on the <code>17,089</code> unique eeg ids. (This is more train data than only using 1950 unique patient id)</p>",
              "rawMarkdown": "No, for each `eeg_id`, i take only one spectrogram random crop. So I perform KFold on the `17,089` unique eeg ids. (This is more train data than only using 1950 unique patient id)",
              "votes": 3
            },
            {
              "id": 2602917,
              "postDate": "2024-01-15T12:37:15.447Z",
              "content": "<p>Sorry, but can you clarify what is meant by multiple crop in your notebook. And what you refer here random crop </p>",
              "rawMarkdown": "Sorry, but can you clarify what is meant by multiple crop in your notebook. And what you refer here random crop ",
              "votes": 2
            },
            {
              "id": 2602922,
              "postDate": "2024-01-15T12:45:38.443Z",
              "content": "<p>We train with a batch of 32 <code>eeg_id</code>. When the data loader create the first sample (out of 32), it uses the <code>eeg_id</code> to determine the corresponding <code>spectrogram_id</code>. We then load the full length <code>spectrogram</code>. The full spectrogram might be 40 minutes. We take a random 10 minute crop somewhere inside those 40 minutes. We train the model with this 10 min crop.</p>\n<p>During epoch 2, we will see the same <code>eeg_id</code> again. This time we take a different random 10 minute crop inside the full 40 minutes. So each epoch, our model sees a different 10 minute crop from the full length <code>spectrogram id</code> (which is associated with the <code>eeg_id</code> of the train sample).</p>",
              "rawMarkdown": "We train with a batch of 32 `eeg_id`. When the data loader create the first sample (out of 32), it uses the `eeg_id` to determine the corresponding `spectrogram_id`. We then load the full length `spectrogram`. The full spectrogram might be 40 minutes. We take a random 10 minute crop somewhere inside those 40 minutes. We train the model with this 10 min crop.\n\nDuring epoch 2, we will see the same `eeg_id` again. This time we take a different random 10 minute crop inside the full 40 minutes. So each epoch, our model sees a different 10 minute crop from the full length `spectrogram id` (which is associated with the `eeg_id` of the train sample).",
              "votes": 5
            },
            {
              "id": 2602941,
              "postDate": "2024-01-15T13:06:27.897Z",
              "content": "<p>Thanks for your explanation 😊.</p>",
              "rawMarkdown": "Thanks for your explanation 😊.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2602404,
      "postDate": "2024-01-15T05:36:32.280Z",
      "content": "<p>We may need a GPT modeled on Chris's work for sure. Thank you</p>",
      "rawMarkdown": "We may need a GPT modeled on Chris's work for sure. Thank you",
      "votes": 2
    },
    {
      "id": 2602332,
      "postDate": "2024-01-15T04:38:03.330Z",
      "content": "<p>Yet another great share! Super interested to see the direction of this competition, thank you for leading the way! I may join and try to add to your pile here!</p>",
      "rawMarkdown": "Yet another great share! Super interested to see the direction of this competition, thank you for leading the way! I may join and try to add to your pile here!",
      "votes": 2,
      "replies": [
        {
          "id": 2602818,
          "postDate": "2024-01-15T11:05:07.160Z",
          "content": "<p>Hi Cody. Good to see you here. I'm looking forward to your helpful discussions and notebooks.</p>",
          "rawMarkdown": "Hi Cody. Good to see you here. I'm looking forward to your helpful discussions and notebooks.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2619322,
      "postDate": "2024-01-25T12:04:41.250Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2626078,
          "postDate": "2024-01-29T19:14:16.140Z",
          "content": "<p>You have a great heart…</p>",
          "rawMarkdown": "You have a great heart...",
          "votes": 1
        }
      ]
    },
    {
      "id": 2610296,
      "postDate": "2024-01-20T04:02:03.177Z",
      "content": "<p>I really appreciate how insightful your thorough explanations have been! Thank you for taking the time to sort through the dataset's complex. I'd like to know if you worked mainly with EfficientNetB2, CatBoost, and WaveNet, or have you tried other models as well?</p>",
      "rawMarkdown": "I really appreciate how insightful your thorough explanations have been! Thank you for taking the time to sort through the dataset's complex. I'd like to know if you worked mainly with EfficientNetB2, CatBoost, and WaveNet, or have you tried other models as well?",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2726259,
      "postDate": "2024-04-01T04:54:03.887Z",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!",
      "votes": 1
    },
    {
      "id": 2715462,
      "postDate": "2024-03-25T14:08:48.423Z",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!\n",
      "votes": 1
    },
    {
      "id": 2704151,
      "postDate": "2024-03-18T15:42:43.513Z",
      "content": "<p>Thanks, it helps a lot!</p>",
      "rawMarkdown": "Thanks, it helps a lot!",
      "votes": 1
    },
    {
      "id": 2697836,
      "postDate": "2024-03-15T05:52:28.433Z",
      "content": "<p>thanks,learned a lot</p>",
      "rawMarkdown": "thanks,learned a lot",
      "votes": 1
    },
    {
      "id": 2670087,
      "postDate": "2024-02-26T16:24:54.187Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": 1
    },
    {
      "id": 2644590,
      "postDate": "2024-02-09T15:28:26.127Z",
      "content": "<p>Thanks for the explanation</p>",
      "rawMarkdown": "Thanks for the explanation",
      "votes": 1
    },
    {
      "id": 2636206,
      "postDate": "2024-02-04T22:16:54.167Z",
      "content": "<p>Thanks for sharing. Great post. </p>",
      "rawMarkdown": "Thanks for sharing. Great post. ",
      "votes": 1
    },
    {
      "id": 2627106,
      "postDate": "2024-01-30T13:10:09.553Z",
      "content": "<p>Thank you for your great work!</p>",
      "rawMarkdown": "Thank you for your great work!",
      "votes": 1
    },
    {
      "id": 2610497,
      "postDate": "2024-01-20T06:54:35.577Z",
      "content": "<p>I understand, thank you! </p>",
      "rawMarkdown": "I understand, thank you! ",
      "votes": 1
    },
    {
      "id": 2608444,
      "postDate": "2024-01-18T19:25:50.213Z",
      "content": "<p>insightful !! thanks</p>",
      "rawMarkdown": "insightful !! thanks",
      "votes": 1
    },
    {
      "id": 2602840,
      "postDate": "2024-01-15T11:22:14.050Z",
      "content": "<p>Great explanation, thank you!</p>",
      "rawMarkdown": "Great explanation, thank you!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2602237,
      "author_name": "Yan Teixeira",
      "author_url": "",
      "post_date": "2024-01-15T02:24:35.027000",
      "content": "<p>Chris, the Gandalf of Kaggle.</p>",
      "votes": 20,
      "replies": []
    },
    {
      "id": 2619843,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2024-01-25T18:42:27.200000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> are these offsets right? EEG offsets being measured with the \"zero\" placed on the begging of the EEG not from the start of the spectogram.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F698fcf40886403a5bc93d19a58eee8e4%2FScreen%20Shot%202024-01-25%20at%2015.37.22.png?generation=1706207919227935&amp;alt=media\"></p>",
      "votes": 8,
      "replies": [
        {
          "id": 2620008,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-25T20:34:55.213000",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> that is correct. All offsets are relative to the specific parquet on disk.</p>\n<p>I like your diagram, I updated my post with new diagrams to make it clearer.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2602866,
      "author_name": "eugene",
      "author_url": "",
      "post_date": "2024-01-15T11:49:27.293000",
      "content": "<p>Thanks for the explanation. I understand your scheme for non-overlapping segments. I'm sorry, but I still don't understand the overlapping segments. Is my below diagram correctly? </p>\n<p>For example EEG_ID 1628180742</p>\n<ol>\n<li><strong>eeg_label_offset_seconds</strong> should be used for both spectrogram and EEG. </li>\n<li>Marking of labels starts from the center of the EEG_ID, then if the offset is 40 seconds we have no EEG data for this section ? 🥴</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F0c1f5de7da95ba7d40d27c0e99ea958b%2F3.drawio%20(2).png?generation=1705319245246205&amp;alt=media\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2602876,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-15T12:02:17.503000",
          "content": "<p>Hi. The variable <code>eeg_label_offset_seconds</code> is the beginning, not the middle. (We are not given the middle). The file <code>train_eegs/1628180742.parquet</code> is 90 seconds long. (It is 18,000 rows and each row is 1/200 seconds). Therefore when <code>eeg_label_offset_seconds = 40</code>, we use the 50 seconds between time 40 and time 90. These are rows <code>40*200</code> thru <code>90*200</code> of <code>1628180742.parquet</code>.</p>\n<p>And the file <code>train_spectrograms/1628180742.parquet</code> is 10 minutes and 40 seconds long. (It is 320 rows and each row is 2 seconds). Note to retrieve spectrograms, we <strong>do not</strong> use <code>eeg_label_offset_seconds</code>. Instead we use <code>spectrogram_label_offset_seconds</code>. For some rows in <code>train.csv</code> the <code>eeg_label_offset_seconds != spectrogram_label_offset_seconds</code>. In your example, they are equal.</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 2602878,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-15T12:03:14.740000",
          "content": "<p>I updated my discussion post with code to extract the eeg and spectrogram for a specific row of <code>train.csv</code>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2608513,
      "author_name": "Sergio Henrique",
      "author_url": "",
      "post_date": "2024-01-18T20:57:48.227000",
      "content": "<p>Did you figured out all this alone? My first thought was that eeg_sub_id and spectogram_sub_id were foreign key to the corresponding files. <br>\nThe eeg_offset*200 is also a tricky part.</p>\n<p>The competition description should be updated with this explanation. Much more clear now.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2602297,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2024-01-15T03:51:53.703000",
      "content": "<p>I thought I understood, but now I understand, thank you! =)</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2617670,
      "author_name": "Occccn",
      "author_url": "",
      "post_date": "2024-01-24T11:40:14.957000",
      "content": "<p>Thank you for your kind explanation. <br>\nI have a question. Why is there a difference in time between the provided spectrograms and the EEG data?<br>\n50 sec vs 600 sec</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2617757,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-24T12:50:35.583000",
          "content": "<p>When doctors evaluate patients, they use the spectrograms to see the \"big picture\" of how the patient were feeling for 10 minute window. And they use the EEG waveforms to see the \"zoom in picture\" of what what happening close to the possible event. (Note the event we are classifying is the middle 10 second window).</p>",
          "votes": 11,
          "replies": [
            {
              "id": 2618940,
              "author_name": "Occccn",
              "author_url": "",
              "post_date": "2024-01-25T05:32:07.160000",
              "content": "<p>Thank you!<br>\n I understand that the time of the EEG is different from that of the spectrogram because of the doctor's diagnostic method.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2613345,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2024-01-22T03:02:19.187000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Amazing discussion thanks for sharing .  Can you please explain why do we multiply by  200 while slicing the eeg_offset , Whereas while slicing spectogram we dont use the same formula . Thanks</p>\n<pre><code>eeg = eeg.iloc[eeg_offset*:(eeg_offset+)*]\nspectrogram = spectrogram.loc[(spectrogram.time&gt;=spec_offset)\n                     &amp;(spectrogram.time&lt;spec_offset+)]\n</code></pre>\n<p>Can you please elaborate on this 2 parts please </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2614177,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-22T13:33:50.893000",
          "content": "<p>Hi. Each row of the eeg parquet is <code>1/200</code> of a second. And each row of the spectrogram parquet is <code>2</code> seconds.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2614321,
              "author_name": "Athar Sayed",
              "author_url": "",
              "post_date": "2024-01-22T14:49:32.403000",
              "content": "<p>Thanks alot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for reply I also noticed the data description which says eegs are sampled at the rate of 200 samples per second . </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2624000,
              "author_name": "Akshat Singhal",
              "author_url": "",
              "post_date": "2024-01-28T14:26:29.170000",
              "content": "<blockquote>\n  <p>And each row of the spectrogram parquet is 2 seconds.<br>\n  Hey Chris, how did you conclude this?</p>\n</blockquote>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2626189,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-29T20:10:52.900000",
              "content": "<p>In the spectrogram parquet files there is a column named <code>time</code>. And each subsequent row increases by 2. (i.e. first row is 1, second row is 3, third row is 5, etc)</p>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2604067,
      "author_name": "Nishant Singhal",
      "author_url": "",
      "post_date": "2024-01-16T08:34:25.933000",
      "content": "<p>Chris, congratulations on your innovative approach to optimizing data loading for the Kaggle competition, particularly your method of loading all 11,138 spectrogram files into memory at once. This strategy is not only efficient but also a great learning point for the community in handling large datasets.</p>\n<p>I have a question: In your experiments with EfficientNetB2 and CatBoost models, how do you address the potential issue of overfitting, especially given the high dimensionality of the EEG and spectrogram data and the relatively smaller number of unique patient IDs?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2604341,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-16T11:25:49.307000",
          "content": "<p>Hi. CatBoost has internal tricks to prevent overfitting. With EfficientNet, i am only training for 3 epochs (i.e. few epochs). Furthermore, we can improve EfficientNet using data augmentation. And we can add more regularization to both models.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2603434,
      "author_name": "Zhenlan",
      "author_url": "",
      "post_date": "2024-01-15T19:13:35.947000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for the explanation. I have one question. I see in your code that you aggregated the target across eeg_id. My understanding is that the fine-grained target for each row instead of the whole egg_id is more time approprate (maybe the waves show seizure in some time period but not others). Any reason why you treat the target this way? Thank you,</p>\n<pre><code>tmp = df()()\n t  TARGETS:\n    train = tmp\n\ny_data = train\ny_data = y_data / y_data(axis=,keepdims=True)\ntrain = y_data\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2603456,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-15T19:25:51.893000",
          "content": "<p>We can run <code>train.groupby('eeg_id').expert_consensus.agg('nunique').value_counts()</code> and get</p>\n<pre><code>   \n     \n      \n      \n       \n</code></pre>\n<p>Therefore if we naively just take 1 target to represent the entire <code>eeg_id</code> we have expected value of being correct = 97.6%. Because <code>(16306 + 333 + 31 + 5)/17089 = 0.976</code>. So yes, we could be more precise, but this simple technique works very well and is easy.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2603127,
      "author_name": "Konstantin Kozlovtsev",
      "author_url": "",
      "post_date": "2024-01-15T15:32:11.020000",
      "content": "<p>Thanks for the explanation, now it much cleaner for me, how to align the data. However I still do not understand, why the predictions are made by 50 seconds of the raw eeg and 10 minutes of spectrograms, because it means that spectrogram have much less resolution that eeg? If we predicting for only 10 seconds, then why should look at the whole 10 minutes interval, isn't 50 seconds enough?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2603143,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-15T15:47:22.113000",
          "content": "<p>The larger window provides context. Imagine if i asked you to predict what i will eat for lunch today. I can give you the 30 minute window of me sitting at my kitchen table. I can also give you the previous 24 hours of me eating and the future 24 hours.  Then you can use information from the larger window to help predict what I do in the smaller 30 minute window.</p>",
          "votes": 14,
          "replies": [
            {
              "id": 2603760,
              "author_name": "Haru",
              "author_url": "",
              "post_date": "2024-01-16T03:07:38.793000",
              "content": "<p>Super great example.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2618452,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2024-01-24T18:28:52.277000",
      "content": "<p>The relationship between <code>spectogram_ids</code> and <code>eeg_id</code> is 1:M? For example, 1 spectogram may have multiple eegs but 1 eeg is always \"inside\" 1 spectogram</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2618469,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-24T18:36:06.647000",
          "content": "<p>Yes. We can confirm this by doing <code>groupby('eeg_id').spectrogram_id.agg('nunique').max()</code> and vice versa (with <code>train.csv</code>)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2608407,
      "author_name": "Shaofeng Kang",
      "author_url": "",
      "post_date": "2024-01-18T18:54:33.187000",
      "content": "<p>What do the sub_ids do in the dataset? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2608504,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-18T20:38:23.163000",
          "content": "<p>Ignore them. I think they are just identified unique rows of train.csv (Furthermore they do not exist in test.csv)</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2613246,
              "author_name": "Shaofeng Kang",
              "author_url": "",
              "post_date": "2024-01-21T23:55:18.610000",
              "content": "<p>ok. thanks </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2606539,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2024-01-17T17:59:34.393000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>  -- I was trying to understand both of your baseline . My understanding is Chris you are taking  approach 2 and Tawara has taken approach 3 . Am I correct ?</p>\n<ol>\n<li>106,800 unique rows</li>\n<li>17,089 unique eeg ids</li>\n<li>11,138 unique spectrogram ids</li>\n<li>1950 unique patient ids</li>\n</ol>\n<p>I also see that Tawara images are 400<em>300 which looks like having all the 4 series together vertically stacked and yours is 128</em>256*4 the four series are placed in channel ? How is your 256 values coming?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2606561,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-17T18:08:48.187000",
          "content": "<p>Yes, I do 2 and Tawara does 3. My 256 is because I take the middle 512 seconds (instead of full 600 seconds). We both do 400x300.</p>\n<p>My data loader provides 128x256x4. There is padding on top 14 and bottom 14, so my height is 100. My width is 300 cropped to 256. Inside my model I reshape it to 512x256 which is basically same as 400x300</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 2606961,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2024-01-18T01:21:15.457000",
          "content": "<p>Yes, you are right. I use the <strong>first</strong> 600 seconds of each unique spectrogram id.</p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 2604227,
      "author_name": "FarisML",
      "author_url": "",
      "post_date": "2024-01-16T10:29:07.413000",
      "content": "<p>Does time matters here?,<br>\nSo solutions like transformers or RNN, will it be effective ? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2604336,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-16T11:24:08.080000",
          "content": "<p>Time matters for 2 reasons:</p>\n<ul>\n<li>we need time to determine what frequencies are in the signal. (i.e. without time, there is no such thing as 10Hz).</li>\n<li>in 3% of cases, the <code>expert_consensus</code> changes over time</li>\n</ul>\n<p>Most people are currently ignoring the second bullet point and just assigning a single target to a single <code>eeg_id</code>. Later in the competition, if we use the 3% to train we can boost our CV score and LB score slightly.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2602420,
      "author_name": "Tanishq dublish",
      "author_url": "",
      "post_date": "2024-01-15T05:53:18.670000",
      "content": "<p>Chris, the god of kaggle || thank you for giving such a great code</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2602321,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2024-01-15T04:27:27.297000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> you are on fire in this comp 🔥.  Btw, loved your explanation about the dataset here. Also, glad to see DL beating ML 😅</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2602824,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-15T11:09:00.347000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> . This is fun competition with lots to explore and learn. I see that you are competition host. Thanks for hosting.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2602973,
              "author_name": "Awsaf",
              "author_url": "",
              "post_date": "2024-01-15T13:38:58.917000",
              "content": "<p><a href=\"https://www.kaggle.com/cdoette\" target=\"_blank\">@cdoette</a>, I wish I was hosting this amazing competition but sadly I'm not; I am working in this competition on behalf of the Keras Team to share starter notebooks.</p>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 2602959,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-15T13:20:45.740000",
          "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> Thanks for hosting this competition, many new topics explored recently (majorly eeg). </p>\n<p><strong>is it good to pin this topic for everyone to understand data better</strong> ? ( we have 5-6 discussions about same topic )</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2602968,
              "author_name": "Awsaf",
              "author_url": "",
              "post_date": "2024-01-15T13:34:58.517000",
              "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> I'm not the competition host per se; I am working in this competition on behalf of the Keras Team to share starter notebooks. You may need to contact the Kaggle Team. Sorry for the confusion.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2602970,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-15T13:36:56.043000",
              "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> Sorry it is my mistake of miss understanding from the tag. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> is it good to keep this topic as pin</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2745038,
      "author_name": "TkgglO",
      "author_url": "",
      "post_date": "2024-04-10T10:20:31.303000",
      "content": "<p>Chris, congratulations on your gold medal in this competition! I would also like to express my deep gratitude for sharing your various knowledge and insights. This was the first competition I attended in earnest. I learned so much from your notebooks and discussions. Thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2744051,
      "author_name": "Shi Luo",
      "author_url": "",
      "post_date": "2024-04-09T18:20:41.433000",
      "content": "<p>Congratulations and thank you again Chris for your invaluable guides!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2743337,
      "author_name": "HB",
      "author_url": "",
      "post_date": "2024-04-09T11:46:25.300000",
      "content": "<p>Congratulations on your gold medal, Chris. I always appreciate your information sharing.  I was wondering why you used GroupKFold in the starter notebook? I'm sure I've seen this before but I can't find it, so if you don't mind, I'd appreciate a refresher. I used StratifiedGroupKFold.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2743573,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T14:22:57.047000",
          "content": "<p>Thanks. Congratulations on your strong Silver finish.</p>\n<p>The purpose of CV is to make validation folds as similar to test data as possible. Then using this CV will optimize a model to perform its best on test data.</p>\n<p>In the beginning of this competition, we did not know if test data has the same mean of 6 targets as train data. Therefore I choose to allow my validation folds to have more variation in their 6 target means. This way, the optimized model can perform well even if the test data has a different mean of 6 targets than train.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2744407,
              "author_name": "HB",
              "author_url": "",
              "post_date": "2024-04-09T22:21:32.517000",
              "content": "<p>Thank you so much for your detailed answer!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2741085,
      "author_name": "BioniChaos",
      "author_url": "",
      "post_date": "2024-04-08T06:10:16.413000",
      "content": "<p>This is a very odd way of storing EEG data, I tried parsing this dataset but eventually gave up: <a href=\"https://bionichaos.com/SeizureFuz\" target=\"_blank\">https://bionichaos.com/SeizureFuz</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2736518,
      "author_name": "nyl0522",
      "author_url": "",
      "post_date": "2024-04-05T08:42:59.267000",
      "content": "<p>Very Helpful!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2725003,
      "author_name": "Charlie Doan",
      "author_url": "",
      "post_date": "2024-03-31T09:22:53.123000",
      "content": "<p>Great explanation!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2714150,
      "author_name": "NGK",
      "author_url": "",
      "post_date": "2024-03-24T18:01:20.250000",
      "content": "<p>Great explanation…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2701005,
      "author_name": "qianmulin",
      "author_url": "",
      "post_date": "2024-03-16T19:54:29.850000",
      "content": "<p>Very helpful！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2695983,
      "author_name": "Y.S",
      "author_url": "",
      "post_date": "2024-03-14T01:31:55.627000",
      "content": "<p>Great explanation of the dataset！</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2695995,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-03-14T01:57:12.477000",
          "content": "<p>Thanks Y.S !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2684979,
      "author_name": "Megha",
      "author_url": "",
      "post_date": "2024-03-06T22:13:59.137000",
      "content": "<p>Thankyou so much <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for your detailed and dedicated explanation. I read all the comments, but still have 1 doubt unanswered. If we can create create the 50seconds spectrogram data using the EEG data, then why are we given both EEG and corresponding Spectrogram data? Does it have to be our choice on which data to train our model..?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2685019,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-03-06T23:28:00.613000",
          "content": "<p>The given spectrograms are 10 minute low resolution windows whereas the given EEG waveforms are 50 second high resolution windows. One is like a zoom out view and one is like the zoom in view. We cannot create each from the other because of the difference in time window and difference in resolution. Both have something unique to contribute to our solution.</p>\n<p>(The rows of given spectrogram parquets are 2 seconds each whereas the rows of given EEG waveforms are 1/200 seconds each. That is why i say low-res and hi-res).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2660610,
      "author_name": "Rauf",
      "author_url": "",
      "post_date": "2024-02-20T17:44:24.243000",
      "content": "<p>why do different EEGs point to the same spectrogram? I mean the spectrograms are created by doing FFT over overlapping windows, wouldn't that have at some point overlapping windows between two different EEGs( different samples essentially), though I'm pretty sure if that's the case its not at all a big deal… just trying to understand how they came up with it, generally you have EEG and you have a spectrogram for that EEG alone… Also if the reason is faster loading because you have multiple spectrograms we could do better (as <a href=\"https://www.kaggle.com/chris\" target=\"_blank\">@chris</a> did in the custom spectrogram creator at the end)…just spit balling</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2660833,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-02-20T20:35:17.627000",
          "content": "<p>EEG-1 and EEG-2 may be eegs for the same patient-A for time 12:00p and time 3:00p respectively. Then Spectrogram-1 is the entire day for patient-A from 11a thru 6p. So both EEG-1 and EEG-2 point to the same Spectrogram-1 file. And EEG-1 uses one offset and EEG-2 uses a different offset.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2655389,
      "author_name": "Tsvetomira Krikoryan",
      "author_url": "",
      "post_date": "2024-02-16T20:27:27.040000",
      "content": "<p>Hi! I am a bit confused if a single EEG parquet file contains information only from a single EEG_ID or if there could be multiple EEG_IDs associated with it. The latter seems unlikely to me but I cannot find any concrete answer on this either. Thank you in advance! :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2655395,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2024-02-16T20:30:41.560000",
          "content": "<p>There is information about a single eeg_id. But can contain multiple subsamples with sub_eeg_id refering to a specific windows starting on eeg_offset.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2655429,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-02-16T21:02:30.597000",
          "content": "<p>Each EEG parquet file only contains information about one <code>eeg_id</code>. Then in <code>train.csv</code> file, many rows have the same <code>eeg_id</code> but access a different time window within the associated EEG parquet file.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2654291,
      "author_name": "Bhoj Raj Thapa",
      "author_url": "",
      "post_date": "2024-02-16T03:38:33.970000",
      "content": "<p>How are you saying the prediction should be done for 10 seconds. The test EEG has a size of 50 seconds and unique ID. And the  submission also requires predicting each eeg_id with 50 seconds of test EEG. I don't see anywhere that the prediction was done for the 10 seconds. (t-5 to t+5, where t is the middle time stamp). Can you explain that? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2654544,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2024-02-16T07:58:05.697000",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data</a><br>\n\"Files<br>\ntrain.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and <strong>labeled the central 10 seconds.</strong> Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\"</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2654868,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-02-16T13:43:21.030000",
          "content": "<p>We input the 50 second EEG and the 10 minute Spectrogram into our model. The model trains with the 6 target columns. The 6 target columns describe the central 10 seconds from both inputs. Therefore the model will learn during training that the most important region is the central 10 seconds. Hence the model learns to focus on the middle 10 seconds itself (and uses the extra time data when needed).</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2679457,
              "author_name": "Ketan Jaltare",
              "author_url": "",
              "post_date": "2024-03-03T14:47:08.663000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for all your useful contributions! Somewhat related question. I notice that the submission must contain a prediction for each eeg_id. In the train set, each eeg_id is often repeated (multiple rows per eeg_id). Is the structure of the test set going to be the same? Because this means that the model needs to make one prediction for multiple rows. Or will the test set have each eeg_id appearing only once? Im a bit confused about this. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2679607,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-03-03T16:22:48.827000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/ketanjaltare\" target=\"_blank\">@ketanjaltare</a> the test set does not repeat <code>eeg_id</code>. In the test set each <code>egg_id</code> only appears once. This is stated in this competition's data page:</p>\n<blockquote>\n  <p>Metadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.</p>\n</blockquote>\n<p>i.e. we don't need <code>eeg_label_offset_seconds</code> and we don't need <code>spectogram_label_offset_seconds</code> because each test eeq parquet is unique and only contains 10,000 rows (and each test spectrogram parquet is only 300 rows)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2680016,
              "author_name": "Ketan Jaltare",
              "author_url": "",
              "post_date": "2024-03-03T22:42:09.123000",
              "content": "<p>Great! Thanks a lot. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2713267,
              "author_name": "Simon Beck",
              "author_url": "",
              "post_date": "2024-03-24T04:42:55.707000",
              "content": "<blockquote>\n  <p>Therefore the model will learn during training that the most important region is the central 10 seconds. Hence the model learns to focus on the middle 10 seconds itself (and uses the extra time data when needed).</p>\n</blockquote>\n<p>I dont get it. model doesnt really get the fact that votes are based on middle 10 seconds. we just know it. can u explain it more?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2652235,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-14T16:29:25.353000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2652271,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-14T16:48:34.773000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 2683628,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-06T05:58:26.197000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2683844,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-06T08:53:33.810000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2650274,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-13T11:09:48.847000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2650575,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-13T14:40:15.493000",
          "content": "",
          "votes": 6,
          "replies": [
            {
              "id": 2650757,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-13T17:29:08.863000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2650052,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-13T07:53:52.080000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2650215,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-13T10:13:19.503000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2650225,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-13T10:21:10.200000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2647292,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-11T13:20:03.817000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2635357,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-04T10:52:07.923000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2629759,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-01T02:31:34.380000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2627292,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-30T15:54:45.417000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2626779,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-30T09:01:52.163000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2624988,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-29T06:10:52.623000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2625239,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-29T09:36:16.437000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2624946,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-29T05:12:54.563000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2624901,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-29T04:30:10.583000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2696311,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-14T08:27:06.313000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2696415,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-14T10:00:26.403000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2696775,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-14T14:33:54.533000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2622218,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-27T10:46:35.157000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2620648,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-26T09:06:07.147000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2620666,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-26T09:24:16.390000",
          "content": "",
          "votes": 8,
          "replies": [
            {
              "id": 2651564,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-14T08:41:45.383000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2620394,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-26T05:35:37.050000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2620661,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-26T09:21:56.140000",
          "content": "",
          "votes": 4,
          "replies": [
            {
              "id": 2621745,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-27T00:15:49.300000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2621795,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-27T02:49:41.930000",
              "content": "",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2621839,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-27T04:35:57.160000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2623083,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-28T00:17:45.477000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2624037,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-28T14:37:45.020000",
              "content": "",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2624547,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-28T19:58:25.623000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2617816,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-24T13:22:18.400000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2688344,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-09T07:06:53.667000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2688345,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-09T07:08:37.427000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2613300,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-22T02:04:03.023000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2612941,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-21T18:07:35.110000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2612519,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-21T12:50:39.987000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2612064,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-21T06:56:09.403000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2611074,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-20T14:49:10.133000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2607462,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-18T08:44:09.723000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2607640,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-18T10:52:27.507000",
          "content": "",
          "votes": 4,
          "replies": [
            {
              "id": 2610634,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-20T08:41:48.447000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2610681,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-20T09:19:42.020000",
              "content": "",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2611098,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-20T14:58:59.813000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2611105,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-20T15:01:33.243000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2612126,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-21T07:58:51.690000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2604952,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-16T19:02:27.053000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2604052,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-16T08:26:39.893000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2602334,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T04:39:38.270000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2602247,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T02:56:04.057000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2641848,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-07T17:50:49.113000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2641861,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-07T17:59:52.147000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2632774,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-02T16:10:58.070000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2632809,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-02T16:33:09.373000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2625831,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-29T16:42:40.297000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2613649,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-22T07:46:43.700000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2617760,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-24T12:51:26.607000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2634697,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-03T23:59:25.720000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2607374,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-18T07:39:15.370000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2605063,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-16T20:05:15.833000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2605069,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-16T20:13:21.090000",
          "content": "",
          "votes": 3,
          "replies": [
            {
              "id": 2605161,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-16T22:19:13.153000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2608520,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-18T21:07:40.270000",
              "content": "",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2608815,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-19T05:47:37.210000",
              "content": "",
              "votes": -1,
              "replies": []
            },
            {
              "id": 2609079,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-19T08:39:09.413000",
              "content": "",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2609618,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-19T15:52:28.993000",
              "content": "",
              "votes": -1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2602764,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T10:19:48.847000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2602807,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-15T10:58:24.693000",
          "content": "",
          "votes": 5,
          "replies": [
            {
              "id": 2602932,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T12:59:23.360000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2602947,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T13:11:36.480000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2602445,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T06:10:28.003000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2602816,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-15T11:03:46.047000",
          "content": "",
          "votes": 3,
          "replies": [
            {
              "id": 2602888,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T12:10:46.663000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2602891,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T12:14:25.740000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2602899,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T12:23:36.587000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2602903,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T12:26:50.993000",
              "content": "",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2602917,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T12:37:15.447000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2602922,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T12:45:38.443000",
              "content": "",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2602941,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-15T13:06:27.897000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2602404,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T05:36:32.280000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2602332,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T04:38:03.330000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2602818,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-15T11:05:07.160000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2619322,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-25T12:04:41.250000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2626078,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-29T19:14:16.140000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2610296,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-20T04:02:03.177000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2726259,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-01T04:54:03.887000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2715462,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-25T14:08:48.423000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2704151,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-18T15:42:43.513000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2697836,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-15T05:52:28.433000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2670087,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-26T16:24:54.187000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2644590,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-09T15:28:26.127000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2636206,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-04T22:16:54.167000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2627106,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-30T13:10:09.553000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2610497,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-20T06:54:35.577000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2608444,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-18T19:25:50.213000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2602840,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T11:22:14.050000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2602232": "The data in this competition is confusing. **Question:** We are given `train.csv` with 106,800 rows but there are only 17089 unique `eeg_ids`, 11138 unique `spectrogram_ids`, and 1950 unique `patients`. What is going on?? **Answer:** Each `row` of train is a `window of time` from one specific `patient`. And the corresponding `eeg` and `spectrogram` can be found in the corresponding Kaggle `parquet` files.\n\n# Data Explained\nSince each `row` of `train.csv` is a specific `window in time` (for a specific `patient_id`), each row has a specific middle timestamp in seconds. For example, maybe `row 235` has center timestamp `T = May 3 2023 19:30:06` exactly in the middle of both its `EEG time window` and `Spectrogram time window`. (Note that the dataframe does **not** give us the middle timestamp).\n\nThe `EEG time window` is length 50 seconds and the `Spectrogram time window` is length 600 seconds. And both have the **same center timestamp**. In this competition, we are asked to predict the event occurring in the middle 10 seconds of both these time windows:\n* Center is `T`\n* EEG is `[T-25:T+25]`\n* Spectrogram is `[T-300:T+300]`\n* We predict event in `[T-5:T+5]`\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain8.png)\n\n# EEG Parquet Files\nThe `EEG parquet files` are longer than 50 seconds. One `EEG parquet file` has multiple rows of `time windows` inside it. Similarily, the `Spectrogram parquet files` are longer than 600 seconds. One `Spectrogram parquet file` has multiple rows of `time windows` inside it.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain9.png)\n\n# Spectrogram Parquet Files\nThere are fewer `Spectrogram parquet files` than `EEG parquet files`. This is because two rows may have the same `Spectrogram parquet file` but different `EEG parquet files`.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/brain11.png)\n\n# Code to Retrieve EEG and Spectrogram\nFor a specific row of `train.csv`, here is code to retrieve the corresponding EEG and Spectrogram. Note that `train.csv` does **not** give us the middle timestamp. Instead it gives us the beginning timestamp for each timewindow. The two beginnings are determined from `eeg_label_offset_seconds` and `spectrogram_label_offset_seconds`. These are offsets from the start of the parquet file which tells us where the `eeg` and `spectrogram` time windows begin respectively.\n\n    GET_ROW = 0\n    EEG_PATH = 'train_eegs/'\n    SPEC_PATH = 'train_spectrograms/'\n\n    train = pd.read_csv('train.csv')\n    row = train.iloc[GET_ROW]\n\n    eeg = pd.read_parquet(f'{EEG_PATH}{row.eeg_id}.parquet')\n    eeg_offset = int( row.eeg_label_offset_seconds )\n    eeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\n\n    spectrogram = pd.read_parquet(f'{SPEC_PATH}{row.spectrogram_id}.parquet')\n    spec_offset = int( row.spectrogram_label_offset_seconds )\n    spectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                         &(spectrogram.time<spec_offset+600)]\n\n# EfficientNetB2 Starter - CV 0.59 - LB 0.43 - (Deep Learning)\nWhen training our models we have (at least) 4 choices of weighting train samples. For example the 4th bullet point says that our dataloader will give an equal chance of outputting a sample for each `patient`. Whereas bullet point 2 says that our dataloader will give an equal chance of outputting a sample for each `eeg id`.\n* 106,800 unique rows\n* 17,089 unique eeg ids\n* 11,138 unique spectrogram ids\n* 1950 unique patient ids\n\nFrom LB probing (discussed [here][1]), we achieve the best LB score using bullet point 2 (i.e. eeg ids). I published a TensorFlow EfficientNetB2 starter notebook [here][2] that trains with bullet point 2 and achieves CV 0.73 and LB 0.57. (UPDATE: @alejopaullier converted to PyTorch [here][10])\n\n# CatBoost Starter - CV 0.74 - LB 0.60 - (Machine Learning)\n\nI published a CatBoost starter [here][4] which achieves CV 0.74 and LB 0.60 (using bullet point 2 above).\n\n# WaveNet Starter - CV 0.81 - LB 0.52 - (Deep Learning)\n\nI published a TensorFlow WaveNet starter [here][5] which achieves CV 0.81 and LB 0.52 (using bullet point 2 above). And it only uses EEG data (and not spectrograms). (UPDATE: @alejopaullier converted to PyTorch [here][9])\n\n# Spectrogram versus EEG\nMy EffNet and CatBoost starter notebooks currently only use information from Kaggle's spectrograms. We can improve CV and LB score by incorporating information from EEGs. **UPDATE**: I published a starter notebook [here][6] to convert EEG data into spectrograms (and we can incorporate these EEG spectrograms into my EffNet and CatBoost starters). My WaveNet starter above trains using the raw EEG waveforms. Consider combining my WaveNet and EffNet into a single model which accepts input of both EEG waveforms and Kaggle spectrograms! **UPDATE** Recent versions of EfficientNet starter and CatBoost starter now use both EEG spectrograms and Kaggle spectrograms.\n\n# Kaggle Dataset\nOur dataloader needs to read 11138 spectrogram files. The most efficient way to design our dataloader is to first read all 11138 spectrogram files into memory once. Then our dataload uses RAM during epoch training instead of continually reading the parquets from disk.\n\nI created a Kaggle [dataset][3] with one file that contains all 11138 spectrograms. First we load this one file (i.e. Python dictionary of parquets converted into NumPy arrays) into memory. (Reading one file is faster than reading 11k files). Then our dataloader speeds through training!\n\nUPDATE: Additionally, i have datasets [here][7] and [here][8] containing EEG spectrograms and EEG raw data respectively.\n\n# Enjoy - Have Fun!\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\n[2]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[3]: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n[4]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[5]: https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\n[6]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n[7]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\n[8]: https://www.kaggle.com/datasets/cdeotte/brain-eegs\n[9]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477610\n[10]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092",
    "2602237": "Chris, the Gandalf of Kaggle.",
    "2619843": "@cdeotte are these offsets right? EEG offsets being measured with the \"zero\" placed on the begging of the EEG not from the start of the spectogram.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F698fcf40886403a5bc93d19a58eee8e4%2FScreen%20Shot%202024-01-25%20at%2015.37.22.png?generation=1706207919227935&alt=media)",
    "2602866": "Thanks for the explanation. I understand your scheme for non-overlapping segments. I'm sorry, but I still don't understand the overlapping segments. Is my below diagram correctly? \n\nFor example EEG_ID 1628180742\n\n1. **eeg_label_offset_seconds** should be used for both spectrogram and EEG. \n2. Marking of labels starts from the center of the EEG_ID, then if the offset is 40 seconds we have no EEG data for this section ? 🥴\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2094652%2F0c1f5de7da95ba7d40d27c0e99ea958b%2F3.drawio%20(2).png?generation=1705319245246205&alt=media)",
    "2608513": "Did you figured out all this alone? My first thought was that eeg_sub_id and spectogram_sub_id were foreign key to the corresponding files. \nThe eeg_offset*200 is also a tricky part.\n\nThe competition description should be updated with this explanation. Much more clear now.",
    "2602297": "I thought I understood, but now I understand, thank you! =)",
    "2617670": "Thank you for your kind explanation. \nI have a question. Why is there a difference in time between the provided spectrograms and the EEG data?\n50 sec vs 600 sec",
    "2613345": "@cdeotte  Amazing discussion thanks for sharing .  Can you please explain why do we multiply by  200 while slicing the eeg_offset , Whereas while slicing spectogram we dont use the same formula . Thanks\n```python\neeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\nspectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                     &(spectrogram.time<spec_offset+600)]\n```\nCan you please elaborate on this 2 parts please \n",
    "2604067": "Chris, congratulations on your innovative approach to optimizing data loading for the Kaggle competition, particularly your method of loading all 11,138 spectrogram files into memory at once. This strategy is not only efficient but also a great learning point for the community in handling large datasets.\n\nI have a question: In your experiments with EfficientNetB2 and CatBoost models, how do you address the potential issue of overfitting, especially given the high dimensionality of the EEG and spectrogram data and the relatively smaller number of unique patient IDs?",
    "2603434": "@cdeotte Thank you for the explanation. I have one question. I see in your code that you aggregated the target across eeg_id. My understanding is that the fine-grained target for each row instead of the whole egg_id is more time approprate (maybe the waves show seizure in some time period but not others). Any reason why you treat the target this way? Thank you,\n\n```\ntmp = df.groupby('eeg_id')[TARGETS].agg('sum')\nfor t in TARGETS:\n    train[t] = tmp[t].values\n    \ny_data = train[TARGETS].values\ny_data = y_data / y_data.sum(axis=1,keepdims=True)\ntrain[TARGETS] = y_data\n```",
    "2603127": "Thanks for the explanation, now it much cleaner for me, how to align the data. However I still do not understand, why the predictions are made by 50 seconds of the raw eeg and 10 minutes of spectrograms, because it means that spectrogram have much less resolution that eeg? If we predicting for only 10 seconds, then why should look at the whole 10 minutes interval, isn't 50 seconds enough?",
    "2618452": "The relationship between `spectogram_ids` and `eeg_id` is 1:M? For example, 1 spectogram may have multiple eegs but 1 eeg is always \"inside\" 1 spectogram",
    "2608407": "What do the sub_ids do in the dataset? ",
    "2606539": "Hello @cdeotte @ttahara  -- I was trying to understand both of your baseline . My understanding is Chris you are taking  approach 2 and Tawara has taken approach 3 . Am I correct ?\n1. 106,800 unique rows\n2. 17,089 unique eeg ids\n3. 11,138 unique spectrogram ids\n4. 1950 unique patient ids\n\nI also see that Tawara images are 400*300 which looks like having all the 4 series together vertically stacked and yours is 128*256*4 the four series are placed in channel ? How is your 256 values coming?",
    "2604227": "Does time matters here?,\nSo solutions like transformers or RNN, will it be effective ? ",
    "2602420": "Chris, the god of kaggle || thank you for giving such a great code",
    "2602321": "@cdeotte you are on fire in this comp 🔥.  Btw, loved your explanation about the dataset here. Also, glad to see DL beating ML 😅\n",
    "2745038": "Chris, congratulations on your gold medal in this competition! I would also like to express my deep gratitude for sharing your various knowledge and insights. This was the first competition I attended in earnest. I learned so much from your notebooks and discussions. Thank you!",
    "2744051": "Congratulations and thank you again Chris for your invaluable guides!!",
    "2743337": "Congratulations on your gold medal, Chris. I always appreciate your information sharing.  I was wondering why you used GroupKFold in the starter notebook? I'm sure I've seen this before but I can't find it, so if you don't mind, I'd appreciate a refresher. I used StratifiedGroupKFold.",
    "2741085": "This is a very odd way of storing EEG data, I tried parsing this dataset but eventually gave up: https://bionichaos.com/SeizureFuz",
    "2736518": "Very Helpful!!",
    "2725003": "Great explanation!",
    "2714150": "Great explanation...",
    "2701005": "Very helpful！",
    "2695983": "Great explanation of the dataset！",
    "2684979": "Thankyou so much @cdeotte for your detailed and dedicated explanation. I read all the comments, but still have 1 doubt unanswered. If we can create create the 50seconds spectrogram data using the EEG data, then why are we given both EEG and corresponding Spectrogram data? Does it have to be our choice on which data to train our model..?",
    "2660610": "why do different EEGs point to the same spectrogram? I mean the spectrograms are created by doing FFT over overlapping windows, wouldn't that have at some point overlapping windows between two different EEGs( different samples essentially), though I'm pretty sure if that's the case its not at all a big deal... just trying to understand how they came up with it, generally you have EEG and you have a spectrogram for that EEG alone... Also if the reason is faster loading because you have multiple spectrograms we could do better (as @chris did in the custom spectrogram creator at the end)...just spit balling",
    "2655389": "Hi! I am a bit confused if a single EEG parquet file contains information only from a single EEG_ID or if there could be multiple EEG_IDs associated with it. The latter seems unlikely to me but I cannot find any concrete answer on this either. Thank you in advance! :)",
    "2654291": "How are you saying the prediction should be done for 10 seconds. The test EEG has a size of 50 seconds and unique ID. And the  submission also requires predicting each eeg_id with 50 seconds of test EEG. I don't see anywhere that the prediction was done for the 10 seconds. (t-5 to t+5, where t is the middle time stamp). Can you explain that? ",
    "2652235": "Is it true that all data contained in the EEG parquet files can be constructed from the spectogram parquet files? If so why are they both given, and why not only the spectrogram files?",
    "2650274": "Hi I am new to the competition and I am here to learn new things as fresher. But still I am in great fear of where to start. Is there any one could help me.",
    "2650052": "Hi @cdeotte, thanks for the clear explanation. \nI have a question related to inference. As I am aware, there is a one-to-many mapping between the kaggle-spectrogram and the EEG-signals in train.csv, in the test.csv we don't have the time_offset for kaggle-spectrogram then how do we select the window of kaggle-spectrogram to EEG-signals or there is no one-to-many mapping in test.csv?",
    "2647292": "Impressive work. I express my gratitude to you. 🙏",
    "2635357": "Great work sir",
    "2629759": "Thank you @cdeotte for all the EDA work you did. I'm very confused about the way the offset seconds are labeled in both the spectrograms and eeg. Consider the attached table which I made by printing out all the rows in which `eeg_id` == 3638862953. If we are supposed to index from the offset label to 50 seconds after for the eeg, then that would produce a lot of overlap. For example, do I start at the 42 second offset mark and index up to the 94 second offset mark? If so, I would have to overlap with `eeg_sub_id` 2,3,4, and 5. There would be even more overlap if I were to index with 10 minutes on the spectrogram, because the `spectrogram_label_offset_seconds` in this example are the same as the `eeg_offset_label_seconds` I'm just not sure how to index this effectively? How am I supposed to deduce which spectrograms/eegs the neurologists actually reviewed? Thank you anyone who might be able to clarify. \n\nUpdate: I noticed is that this eeg has 10000 rows, which means it is exactly 50 seconds long. I guess this means we might want to err on the lowest offset label in each train.csv file? This is because if you start any later in this the 3638862953 eeg, you won't be able to get the full 50 seconds.",
    "2627292": "Thanks so much for laying out these details clearly and visually!\nMy inner pedant (\"one who unduly emphasizes minutiae\") wants to note that there is no \"middle 10 seconds\" of data in the spectrograms: The middle of the 300 spectra is between the 150th and 151st, so there is actually no (or two) middle 10 seconds 😯  There are middle 8 seconds and middle 12 seconds sections, though 😃\nLess pedantly, unlike other cases where we had to determine the event center, here it is given. But I wonder how accurate/meaningful it is? If it is just an approximate location (is it @awsaf49 ?) then the data analysis method should be insensitive to few-seconds (or more?) time shifts; and time shifts could be used as a data augmentation. Doing image analysis on the whole spectrogram is probably insensitive to time shifts if shifts are effectively already in the data.",
    "2626779": "Thank you for your explanation, that makes me clearer.",
    "2624988": "Hi @cdeotte, thanks for making the competition a lot clearer.\nI had a doubt, shouldn't the diagram be like this in `Spectrogram parquet files` section\ni.e. the `eeg_label_offset_seconds` should be `spectrogram_label_offset_seconds`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4359632%2F559a7782fd26aba6d1569a58d0500ba8%2Fdownload.png?generation=1706508510916891&alt=media)\n",
    "2624946": "Very informative, thanks for the explanation! ",
    "2624901": "Thanks a lot, Great explanation of the dataset.",
    "2696311": "Thank you for sharing the good analysis:)\n\nI have a question. Then, wouldn’t it be right to create a spectrogram for each offset and learn it?\n\nA label for 50 seconds is given based on the offset, and creating a spectrogram of more detailed parts seems to be a more accurate learning method.\n\n```python\nmeta_data = row[1]\noffset_sec = int(meta_data['eeg_label_offset_seconds'])\neeg = original_eeg.iloc[offset_sec*200:(offset_sec+50)*200].reset_index(drop=True)\nmiddle = int(len(eeg)//2)\neeg = eeg.iloc[middle-(5*200):middle+(5*200)] # 2000\n ```\n\nWouldn’t it be right to structure the dataset like this?",
    "2622218": "Many thanks for the great description. \nIt was very helpful to get started with the competition!",
    "2620648": "One point i couldn't understand İf EGG signals measured 50 seconds for each patients. Why spectrograms shows 10 minute. I mean to create 10 minute spectrogram shouldn't we use 10 minute of EEG signals? So these are 2 different measurements?",
    "2620394": "@cdeotte Sorry if it's a silly question. Why does the input shape for EfficientNet have to be (128, 256, x) when the base model is loaded without the top?",
    "2617816": "Impressive work. It is inspiring.",
    "2688344": "Hi, I have two inquiries:\n\n1. Could you explain the integration of both EEG data and spectrogram into the processes of model training?\n\n2. Is it imperative to apply filtering, denoising, or normalization techniques to the provided readings?",
    "2613300": "Thank you for the great read. Easy and well explained",
    "2612941": "thank you for the explanation and the example codes",
    "2612519": "Thanks for the insight. Was a good read. Great explanation !!\n\n",
    "2612064": "Thanks for the work, its quite useful and enriching.",
    "2611074": "Thank you for your interesting information",
    "2607462": "Great！but why 10 sec here？",
    "2604952": "Thanks for the insight. Was a good read.",
    "2604052": "hello sir you are great",
    "2602334": "Thanks for your work. I now got to understand how data is given.",
    "2602247": "@cdeotte till now everyone imagine it, after seeing your well designed flow. All questions are answered mainly\n> The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\n\nThanks for sharing\n",
    "2641848": "Hi and thanks for the clear explanation. Just to check if I understood it correctly. \nAt\neeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\nthe 200 factor is because the EEG have been measured at 200 Hz so there is 200 rows per second measured?\n\nEDIT: My bad. Just answered question a couple comments up.",
    "2632774": "Hi @cdeotte, the link for the wavenet starter is taking me to the catboost starter",
    "2625831": "Thanks a lot Chris. this will help me hit the ground running!",
    "2613649": "Can anyone please tell me why the sum of votes for each rows are different?",
    "2607374": "Thanks for the knowledge! Really Helpful!",
    "2605063": "Hi @cdeotte \nwhat is the relationship between the EEG and the spectrogram?  plz explain to me on dataset point of view",
    "2602764": "Thank you for clear explanation!\nSo when calculating Center of spectrogram window(T), we need to add 'spectrogram_label_offset_seconds' to 300sec.\nAnd also, predicting labels using [T-5, T+5].\nIs it correct?",
    "2602445": "@cdeotte can you explain this part of your code in data generator \n`\n           for k in range(4):\n                # EXTRACT 300 ROWS OF SPECTROGRAM\n                img = self.specs[row.spec_id][r:r+300,k*100:(k+1)*100].T`\n**Edit** : okay got it.\nWhy only first 300 rows is being extracted only?",
    "2602404": "We may need a GPT modeled on Chris's work for sure. Thank you",
    "2602332": "Yet another great share! Super interested to see the direction of this competition, thank you for leading the way! I may join and try to add to your pile here!",
    "2619322": "",
    "2610296": "I really appreciate how insightful your thorough explanations have been! Thank you for taking the time to sort through the dataset's complex. I'd like to know if you worked mainly with EfficientNetB2, CatBoost, and WaveNet, or have you tried other models as well?",
    "2726259": "thanks a lot!",
    "2715462": "thanks a lot!\n",
    "2704151": "Thanks, it helps a lot!",
    "2697836": "thanks,learned a lot",
    "2670087": "Thanks a lot!",
    "2644590": "Thanks for the explanation",
    "2636206": "Thanks for sharing. Great post. ",
    "2627106": "Thank you for your great work!",
    "2610497": "I understand, thank you! ",
    "2608444": "insightful !! thanks",
    "2602840": "Great explanation, thank you!"
  }
}