{
  "id": 467576,
  "title": "UPDATED - CatBoost Starter Notebook and Kaggle Dataset - LB 0.60",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467576",
  "author_name": "",
  "post_date": "2024-01-13T04:46:27.010870200Z",
  "votes": 132,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hi everyone. I published a CatBoost starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-8\" target=\"_blank\">here</a>. It trains a model using simple features from spectrograms (and does not use features from eeg parquets). It achieves CV 0.74 and LB 0.60! 🎉</p>\n<p>I also published a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">here</a> with all 11138 spectrograms in one single file. This makes reading all the spectrograms 10x faster to facilitate faster experimentation! And I converted all the 17089 raw eeg waveforms into spectrograms and uploaded them to a dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a>. These new spectrograms boost the CV score and LB score of all models that are only using Kaggle spectrograms! 🎉</p>\n<h1>How CatBoost Works</h1>\n<p>In this competition, the <code>train.csv</code> file has 17,089 unique <code>eeg_id</code>. The reason that <code>train.csv</code> has 106,800 rows is because we get multiple time windows into these 17,089 unique eeq. So we can think of the 100k rows as data augmentation random crops into 17k train samples.</p>\n<p>The CatBoost starter notebook only uses 17,089 train samples to train the 5-Fold model. For each unique <code>eeg_id</code>, we use multiclass target with one class equal 1 and the other 5 classes equal 0 as indicated from <code>train.csv</code>. Then for each unique <code>eeg_id</code>, we create features from the middle 10 minutes of the associated spectrogram.</p>\n<p>Our features are simple. A 10 minute spectrogram has 300 readings (taking every 2 seconds). Readings are taken for 100 frequencies from 4 quadrants of the brain. We take the average over time of each of these 400 time series. This produces 400 features to be used with each <code>eeg_id</code>.</p>\n<h2>UPDATE 1:</h2>\n<p>Version 2 uses <code>means</code> and <code>mins</code>. And Version 2 uses both 10 minute time window and 20 second time window.</p>\n<h2>UPDATE 2:</h2>\n<p>Version 3 adds features from EEG waveforms and boosts CV from CV 0.82 =&gt; CV 0.74. And boosts LB 0.67 =&gt; LB 0.60. Hooray!</p>\n<p>We use these features to train our CatBoost model. There are many ways to improve CV score and LB score. From the spectrograms, we can create lots of more features. We can also create 10 (or whatever) samples for each <code>eeg_id</code> where each sample uses a different 10 minute spectrogram window (i.e. this would be like random crop data augmentation). This would create 10x more data. Also we can create features from the eeg parquets.</p>",
  "messages": [
    {
      "id": "2599691",
      "postDate": "01/13/2024 04:46:27",
      "content": "<p>Hi everyone. I published a CatBoost starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-8\" target=\"_blank\">here</a>. It trains a model using simple features from spectrograms (and does not use features from eeg parquets). It achieves CV 0.74 and LB 0.60! 🎉</p>\n<p>I also published a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">here</a> with all 11138 spectrograms in one single file. This makes reading all the spectrograms 10x faster to facilitate faster experimentation! And I converted all the 17089 raw eeg waveforms into spectrograms and uploaded them to a dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a>. These new spectrograms boost the CV score and LB score of all models that are only using Kaggle spectrograms! 🎉</p>\n<h1>How CatBoost Works</h1>\n<p>In this competition, the <code>train.csv</code> file has 17,089 unique <code>eeg_id</code>. The reason that <code>train.csv</code> has 106,800 rows is because we get multiple time windows into these 17,089 unique eeq. So we can think of the 100k rows as data augmentation random crops into 17k train samples.</p>\n<p>The CatBoost starter notebook only uses 17,089 train samples to train the 5-Fold model. For each unique <code>eeg_id</code>, we use multiclass target with one class equal 1 and the other 5 classes equal 0 as indicated from <code>train.csv</code>. Then for each unique <code>eeg_id</code>, we create features from the middle 10 minutes of the associated spectrogram.</p>\n<p>Our features are simple. A 10 minute spectrogram has 300 readings (taking every 2 seconds). Readings are taken for 100 frequencies from 4 quadrants of the brain. We take the average over time of each of these 400 time series. This produces 400 features to be used with each <code>eeg_id</code>.</p>\n<h2>UPDATE 1:</h2>\n<p>Version 2 uses <code>means</code> and <code>mins</code>. And Version 2 uses both 10 minute time window and 20 second time window.</p>\n<h2>UPDATE 2:</h2>\n<p>Version 3 adds features from EEG waveforms and boosts CV from CV 0.82 =&gt; CV 0.74. And boosts LB 0.67 =&gt; LB 0.60. Hooray!</p>\n<p>We use these features to train our CatBoost model. There are many ways to improve CV score and LB score. From the spectrograms, we can create lots of more features. We can also create 10 (or whatever) samples for each <code>eeg_id</code> where each sample uses a different 10 minute spectrogram window (i.e. this would be like random crop data augmentation). This would create 10x more data. Also we can create features from the eeg parquets.</p>",
      "rawMarkdown": "Hi everyone. I published a CatBoost starter notebook [here][1]. It trains a model using simple features from spectrograms (and does not use features from eeg parquets). It achieves CV 0.74 and LB 0.60! 🎉\n\nI also published a Kaggle dataset [here][2] with all 11138 spectrograms in one single file. This makes reading all the spectrograms 10x faster to facilitate faster experimentation! And I converted all the 17089 raw eeg waveforms into spectrograms and uploaded them to a dataset [here][3]. These new spectrograms boost the CV score and LB score of all models that are only using Kaggle spectrograms! 🎉\n\n# How CatBoost Works\nIn this competition, the `train.csv` file has 17,089 unique `eeg_id`. The reason that `train.csv` has 106,800 rows is because we get multiple time windows into these 17,089 unique eeq. So we can think of the 100k rows as data augmentation random crops into 17k train samples.\n\nThe CatBoost starter notebook only uses 17,089 train samples to train the 5-Fold model. For each unique `eeg_id`, we use multiclass target with one class equal 1 and the other 5 classes equal 0 as indicated from `train.csv`. Then for each unique `eeg_id`, we create features from the middle 10 minutes of the associated spectrogram.\n\nOur features are simple. A 10 minute spectrogram has 300 readings (taking every 2 seconds). Readings are taken for 100 frequencies from 4 quadrants of the brain. We take the average over time of each of these 400 time series. This produces 400 features to be used with each `eeg_id`.\n\n## UPDATE 1: \nVersion 2 uses `means` and `mins`. And Version 2 uses both 10 minute time window and 20 second time window.\n\n## UPDATE 2:\nVersion 3 adds features from EEG waveforms and boosts CV from CV 0.82 => CV 0.74. And boosts LB 0.67 => LB 0.60. Hooray!\n\nWe use these features to train our CatBoost model. There are many ways to improve CV score and LB score. From the spectrograms, we can create lots of more features. We can also create 10 (or whatever) samples for each `eeg_id` where each sample uses a different 10 minute spectrogram window (i.e. this would be like random crop data augmentation). This would create 10x more data. Also we can create features from the eeg parquets.\n\n[1]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-8\n[2]: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n[3]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms",
      "votes": null
    },
    {
      "id": "2599845",
      "postDate": "01/13/2024 07:06:11",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?</p>",
      "rawMarkdown": "cdeotte Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?",
      "votes": null
    },
    {
      "id": "2599910",
      "postDate": "01/13/2024 07:41:38",
      "content": "<p>It's great to see you again in this competition😀</p>",
      "rawMarkdown": "It's great to see you again in this competition😀",
      "votes": null
    },
    {
      "id": "2600139",
      "postDate": "01/13/2024 12:02:47",
      "content": "<p>Since Chris has started sharing from very beginning of this competition this is going to be exciting filled with lot of learnings :) </p>",
      "rawMarkdown": "Since Chris has started sharing from very beginning of this competition this is going to be exciting filled with lot of learnings :)",
      "votes": null
    },
    {
      "id": "2600491",
      "postDate": "01/13/2024 17:15:23",
      "content": "<p>I have the same question</p>",
      "rawMarkdown": "I have the same question",
      "votes": null
    },
    {
      "id": "2600727",
      "postDate": "01/13/2024 19:53:53",
      "content": "<p>Generally I use Stratified Group KFold in the following two situations:</p>\n<ul>\n<li>one target class is very rare and we want to make sure to include this target class in both train and valid (for each K fold split)</li>\n<li>the test data has the same proportion of target classes as train data</li>\n</ul>\n<p>In this competition, we are not sure if the test data has the same proportions (of 6 target classes) as train data. Therefore if we use Group KFold (instead of Stratified Group KFold) then we are evaluating how well our model can perform when the test data may have slightly different proportions as train data.</p>\n<p>If we use Stratified Group KFold, then we are optimizing a model to make predictions on test data that has the same proportion of target classes as train data. In conclusion, they will both work well, but perhaps Group KFold will produce a model that generalizes slightly better to an unknown test proportion.</p>\n<p>(Note we can probe the public test target class proportions, but we cannot know the private test target class proportions).</p>",
      "rawMarkdown": "Generally I use Stratified Group KFold in the following two situations:\n* one target class is very rare and we want to make sure to include this target class in both train and valid (for each K fold split)\n* the test data has the same proportion of target classes as train data\n\nIn this competition, we are not sure if the test data has the same proportions (of 6 target classes) as train data. Therefore if we use Group KFold (instead of Stratified Group KFold) then we are evaluating how well our model can perform when the test data may have slightly different proportions as train data.\n\nIf we use Stratified Group KFold, then we are optimizing a model to make predictions on test data that has the same proportion of target classes as train data. In conclusion, they will both work well, but perhaps Group KFold will produce a model that generalizes slightly better to an unknown test proportion.\n\n(Note we can probe the public test target class proportions, but we cannot know the private test target class proportions).",
      "votes": null
    },
    {
      "id": "2600888",
      "postDate": "01/14/2024 01:43:26",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the detailed explanation. is if-else approach for probe the public test target class proportions?</p>\n<p><em>Note</em>: <strong>No major difference in local CV with StratifiedGroupKFold vs GroupKFold.</strong>. I will go with your CV approach as it correlate with LB except gap between Local CV and LB is [0.2 - 0.15] as per your notebook (14Jan).</p>",
      "rawMarkdown": "cdeotte Thanks for the detailed explanation. is if-else approach for probe the public test target class proportions?\n\n*Note*: **No major difference in local CV with StratifiedGroupKFold vs GroupKFold.**. I will go with your CV approach as it correlate with LB except gap between Local CV and LB is [0.2 - 0.15] as per your notebook (14Jan).",
      "votes": null
    },
    {
      "id": "2600966",
      "postDate": "01/14/2024 03:28:49",
      "content": "<p>Out of curiosity has anyone tried building a pytorch variant of this dataloader and training an lstm?</p>",
      "rawMarkdown": "Out of curiosity has anyone tried building a pytorch variant of this dataloader and training an lstm?",
      "votes": null
    },
    {
      "id": "2600997",
      "postDate": "01/14/2024 03:56:27",
      "content": "<p>I have tried EfficientNet on spectrograms. And i have tried RNN, WaveNet, and Transformer on eegs. So far I cannot beat my CatBoost yet. But i will keep trying.</p>",
      "rawMarkdown": "I have tried EfficientNet on spectrograms. And i have tried RNN, WaveNet, and Transformer on eegs. So far I cannot beat my CatBoost yet. But i will keep trying.",
      "votes": null
    },
    {
      "id": "2601786",
      "postDate": "01/14/2024 17:06:58",
      "content": "<p>Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?</p>",
      "rawMarkdown": "Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?",
      "votes": null
    },
    {
      "id": "2601849",
      "postDate": "01/14/2024 18:06:52",
      "content": "<p>They both work well. For my full answer see my other comment below <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467576#2600727\" target=\"_blank\">link</a></p>",
      "rawMarkdown": "They both work well. For my full answer see my other comment below [link][1]\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467576#2600727",
      "votes": null
    },
    {
      "id": "2601914",
      "postDate": "01/14/2024 19:16:17",
      "content": "<p>Hi Chris, Thank you so much for your response! This is really fascinating. I was wondering if I could take a look at your RNN code, namely the dataloader? Just getting my feet wet with ML and would love to learn from it!</p>",
      "rawMarkdown": "Hi Chris, Thank you so much for your response! This is really fascinating. I was wondering if I could take a look at your RNN code, namely the dataloader? Just getting my feet wet with ML and would love to learn from it!",
      "votes": null
    },
    {
      "id": "2601971",
      "postDate": "01/14/2024 20:06:23",
      "content": "<p>I will publish it soon when I begin working with eeg again. </p>\n<p>Right now i am working with spectrogram. I got my EfficientNet on spectrogram code to work. It now performs better than my CatBoost on spectrogram. </p>",
      "rawMarkdown": "I will publish it soon when I begin working with eeg again. \n\nRight now i am working with spectrogram. I got my EfficientNet on spectrogram code to work. It now performs better than my CatBoost on spectrogram.",
      "votes": null
    },
    {
      "id": "2601978",
      "postDate": "01/14/2024 20:08:39",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> since spectrograms are generated without noise. do we need to do preprocess for spectrogram?</p>",
      "rawMarkdown": "cdeotte since spectrograms are generated without noise. do we need to do preprocess for spectrogram?",
      "votes": null
    },
    {
      "id": "2601994",
      "postDate": "01/14/2024 20:16:49",
      "content": "<p>So far the following preprocess works best with my EfficientNet model. (I learned some of this from Tawara's great starter notebooks)</p>\n<ul>\n<li>arrange the 4 spectrograms flat as 400x300x1 (as opposed to 100x300x4 or 200x600x2 etc etc)</li>\n<li>log transform (i.e. <code>img = np.clip(img,np.exp(-4),np.exp(8)); img = np.log(img)</code>)</li>\n<li>standardize per image (with minus mean divide std, then each img has mean=0 )</li>\n<li>fill nan with 0</li>\n<li>repeat 1 channel to make 3 channel monotone image 400x300x1 =&gt; 400x300x3</li>\n</ul>",
      "rawMarkdown": "So far the following preprocess works best with my EfficientNet model. (I learned some of this from Tawara's great starter notebooks)\n* arrange the 4 spectrograms flat as 400x300x1 (as opposed to 100x300x4 or 200x600x2 etc etc)\n* log transform (i.e. `img = np.clip(img,np.exp(-4),np.exp(8)); img = np.log(img)`)\n* standardize per image (with minus mean divide std, then each img has mean=0 )\n* fill nan with 0\n* repeat 1 channel to make 3 channel monotone image 400x300x1 => 400x300x3",
      "votes": null
    },
    {
      "id": "2601998",
      "postDate": "01/14/2024 20:19:20",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for detailed approach <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, i will spend time on now spectrogram sequences.</p>",
      "rawMarkdown": "cdeotte Thanks for detailed approach @cdeotte, i will spend time on now spectrogram sequences.",
      "votes": null
    },
    {
      "id": "2602005",
      "postDate": "01/14/2024 20:28:57",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I expected eeg sequences are better choice then spectrogram, but training model to learn patterns failing. Is it same for you also about eeg sequence experiments.</p>",
      "rawMarkdown": "cdeotte I expected eeg sequences are better choice then spectrogram, but training model to learn patterns failing. Is it same for you also about eeg sequence experiments.",
      "votes": null
    },
    {
      "id": "2602010",
      "postDate": "01/14/2024 20:30:25",
      "content": "<p><a href=\"https://www.kaggle.com/tanishqdublish\" target=\"_blank\">@tanishqdublish</a> Both public LB top notebooks as 15Jan used GroupKFold &amp; StratifiedGroupKFold -&gt; got same 0.67 score with different approach :)<br>\n<a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">Notebook - CatBoost Starter - [LB 0.67]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  =&gt; <strong>GroupKFold</strong><br>\n<a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training\" target=\"_blank\">Notebook - HMS-HBAC: ResNet34d Baseline [Training]</a> by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> =&gt; <strong>StratifiedGroupKFold</strong></p>",
      "rawMarkdown": "tanishqdublish Both public LB top notebooks as 15Jan used GroupKFold & StratifiedGroupKFold -> got same 0.67 score with different approach :)\n[Notebook - CatBoost Starter - [LB 0.67]](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67) by @cdeotte  => **GroupKFold**\n[Notebook - HMS-HBAC: ResNet34d Baseline [Training]](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training) by @ttahara => **StratifiedGroupKFold**",
      "votes": null
    },
    {
      "id": "2602212",
      "postDate": "01/15/2024 01:34:31",
      "content": "<p>So…it would seem this would be yet another competition where <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is gonna destroy the LB. Why do I have flashbacks lol</p>",
      "rawMarkdown": "So...it would seem this would be yet another competition where @cdeotte is gonna destroy the LB. Why do I have flashbacks lol",
      "votes": null
    },
    {
      "id": "2602289",
      "postDate": "01/15/2024 03:46:30",
      "content": "<p>Yes. My current <code>LB = 0.46</code> only uses spectrogram. All my attempts to use EEG have not improved my CV score. I even created the 4 signals for LL, LP, RR, RP from the 16 pairs. And inputted these signals into WaveNet which should extract frequencies like spectrogram but it did not do better than using train means.</p>",
      "rawMarkdown": "Yes. My current `LB = 0.46` only uses spectrogram. All my attempts to use EEG have not improved my CV score. I even created the 4 signals for LL, LP, RR, RP from the 16 pairs. And inputted these signals into WaveNet which should extract frequencies like spectrogram but it did not do better than using train means.",
      "votes": null
    },
    {
      "id": "2602308",
      "postDate": "01/15/2024 04:08:09",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for conforming about EEG sequences</p>",
      "rawMarkdown": "cdeotte Thanks for conforming about EEG sequences",
      "votes": null
    },
    {
      "id": "2602364",
      "postDate": "01/15/2024 05:13:45",
      "content": "<p>Sir i don't think , such parameters  exists. could to please elaborate the following term….</p>",
      "rawMarkdown": "Sir i don't think , such parameters  exists. could to please elaborate the following term....",
      "votes": null
    },
    {
      "id": "2605710",
      "postDate": "01/17/2024 07:37:44",
      "content": "<p>Hi，thx for your codes, which really helps me alot. Some lines that I can not quite understand:</p>\n<pre><code>     = int( (row['min'] + row['max'])// ) \n    \n     = np.nanmean(spectrograms[row.spec_id][r:r+,:],axis=)\n    [k,:] = x\n     = np.nanmin(spectrograms[row.spec_id][r:r+,:],axis=)\n    [k,:] = x\n\n    \n     = np.nanmean(spectrograms[row.spec_id][r+:r+,:],axis=)\n    [k,:] = x\n     = np.nanmin(spectrograms[row.spec_id][r+:r+,:],axis=)\n    [k,:] = x\n</code></pre>\n<p>So how come  r = int( (row['min'] + row['max'])//4 ) and r+145? <br>\nlooking forward your reply.</p>",
      "rawMarkdown": "Hi，thx for your codes, which really helps me alot. Some lines that I can not quite understand:\n\n        r = int( (row['min'] + row['max'])//4 ) \n        # 10 MINUTE WINDOW FEATURES (MEANS and MINS)\n        x = np.nanmean(spectrograms[row.spec_id][r:r+300,:],axis=0)\n        data[k,:400] = x\n        x = np.nanmin(spectrograms[row.spec_id][r:r+300,:],axis=0)\n        data[k,400:800] = x\n        \n        # 20 SECOND WINDOW FEATURES (MEANS and MINS)\n        x = np.nanmean(spectrograms[row.spec_id][r+145:r+155,:],axis=0)\n        data[k,800:1200] = x\n        x = np.nanmin(spectrograms[row.spec_id][r+145:r+155,:],axis=0)\n        data[k,1200:1600] = x\n\nSo how come  r = int( (row['min'] + row['max'])//4 ) and r+145? \nlooking forward your reply.",
      "votes": null
    },
    {
      "id": "2605715",
      "postDate": "01/17/2024 07:44:48",
      "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> in train.pqt, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> keep maintaining min and max sub-sequence offsets</p>\n<blockquote>\n  <p>r = int( (row['min'] + row['max'])//4 )  =&gt; <strong>mid point is (min+max)/2 but we have array where each unit is 2secs so r =  ((min+max)/2)/2</strong></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>r+145:r+155 =&gt; <strong>10 units which is 20secs starts from offset \"r\" (mid sub-sequence)</strong></p>\n</blockquote>",
      "rawMarkdown": "niuweikun in train.pqt, @cdeotte keep maintaining min and max sub-sequence offsets\n\n\n> r = int( (row['min'] + row['max'])//4 )  => **mid point is (min+max)/2 but we have array where each unit is 2secs so r =  ((min+max)/2)/2**\n\n---\n>r+145:r+155 => **10 units which is 20secs starts from offset \"r\" (mid sub-sequence)**",
      "votes": null
    },
    {
      "id": "2605727",
      "postDate": "01/17/2024 07:55:20",
      "content": "<p>wow great！Thank you so much😀❤️</p>",
      "rawMarkdown": "wow great！Thank you so much😀❤️",
      "votes": null
    },
    {
      "id": "2605756",
      "postDate": "01/17/2024 08:11:24",
      "content": "<p>so 145 in r+145 is randomly selected? </p>",
      "rawMarkdown": "so 145 in r+145 is randomly selected?",
      "votes": null
    },
    {
      "id": "2605767",
      "postDate": "01/17/2024 08:22:29",
      "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> spectrogram is 10mins i.e 600secs so 300 units, midpoint is 150 so 145:155 is 20secs</p>\n<p>Go through Discussion - <strong>Understanding Competition Data and EfficientNetB2 Starter - LB 0.57</strong></p>",
      "rawMarkdown": "niuweikun spectrogram is 10mins i.e 600secs so 300 units, midpoint is 150 so 145:155 is 20secs\n\nGo through Discussion - **Understanding Competition Data and EfficientNetB2 Starter - LB 0.57**",
      "votes": null
    },
    {
      "id": "2605781",
      "postDate": "01/17/2024 08:35:26",
      "content": "<p>really cool. I am so lucky to meet u here♥️</p>",
      "rawMarkdown": "really cool. I am so lucky to meet u here♥️",
      "votes": null
    },
    {
      "id": "2616799",
      "postDate": "01/23/2024 20:37:35",
      "content": "<p><strong>UPDATE</strong> I added EEG waveform features and boosted CV score and LB score!</p>",
      "rawMarkdown": "**UPDATE** I added EEG waveform features and boosted CV score and LB score!",
      "votes": null
    },
    {
      "id": "2620620",
      "postDate": "01/26/2024 08:41:44",
      "content": "<p>I'm new in this competition, so <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> if I'm right you first found this approach which outperformed all the rest before settling on the effnet and wavenet approach? Very nice notebook btw, currently trying to improve your catboost notebook.</p>",
      "rawMarkdown": "I'm new in this competition, so @cdeotte if I'm right you first found this approach which outperformed all the rest before settling on the effnet and wavenet approach? Very nice notebook btw, currently trying to improve your catboost notebook.",
      "votes": null
    },
    {
      "id": "2620668",
      "postDate": "01/26/2024 09:25:40",
      "content": "<p>Yes. I created CatBoost first. It was quicker to create, train, and infer. Also by playing around with CatBoost features and looking at feature importance, we begin to see what is important. This helps us create a better architecture for our EfficientNet model.</p>",
      "rawMarkdown": "Yes. I created CatBoost first. It was quicker to create, train, and infer. Also by playing around with CatBoost features and looking at feature importance, we begin to see what is important. This helps us create a better architecture for our EfficientNet model.",
      "votes": null
    },
    {
      "id": "2621064",
      "postDate": "01/26/2024 14:59:25",
      "content": "<p>Okay nice, I'll have a look in your other notebooks to and see how you incorporated the most important features from your catboost starter.</p>",
      "rawMarkdown": "Okay nice, I'll have a look in your other notebooks to and see how you incorporated the most important features from your catboost starter.",
      "votes": null
    },
    {
      "id": "2627238",
      "postDate": "01/30/2024 14:53:23",
      "content": "<p>Greetings. <br>\nI am new to competitions as a whole and wonder, how exactly have you \"zipped\" spectogram parquets?<br>\nAs I understood you \"jsoned\" them to specific format with numpy arrays as values. Can I read about this \".npy\" anywhere?<br>\nAnd are there any ideas about eeg.parquets reduction?</p>",
      "rawMarkdown": "Greetings. \nI am new to competitions as a whole and wonder, how exactly have you \"zipped\" spectogram parquets?\nAs I understood you \"jsoned\" them to specific format with numpy arrays as values. Can I read about this \".npy\" anywhere?\nAnd are there any ideas about eeg.parquets reduction?",
      "votes": null
    },
    {
      "id": "2627247",
      "postDate": "01/30/2024 15:07:31",
      "content": "<p>Correct. I have not compressed them. I have just converted each parquet into a NumPy array. Then I put all the NumPy arrays into a single Python dictionary. Finally i saved with <code>np.save()</code>. If we want to reduce the file size on disk, we can compress with NumPy using</p>\n<pre><code>np.savez\n</code></pre>\n<p>Also note that all my starter notebooks only use the middle 300 rows of the Kaggle spectrogram parquets (which is middle 10 minutes). So we can also just save the middle 300 rows instead of all the rows of data to reduce disk size more. For my EEG spectrogram dataset, I also use <code>np.save()</code>. We can do similar things to compress them.</p>",
      "rawMarkdown": "Correct. I have not compressed them. I have just converted each parquet into a NumPy array. Then I put all the NumPy arrays into a single Python dictionary. Finally i saved with `np.save()`. If we want to reduce the file size on disk, we can compress with NumPy using\n\n    np.savez_compressed('compressed_array.npz', arr=arr)\n\nAlso note that all my starter notebooks only use the middle 300 rows of the Kaggle spectrogram parquets (which is middle 10 minutes). So we can also just save the middle 300 rows instead of all the rows of data to reduce disk size more. For my EEG spectrogram dataset, I also use `np.save()`. We can do similar things to compress them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2599845,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/13/2024 07:06:11",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2600491,
          "author_name": "yantxx",
          "author_url": "",
          "post_date": "01/13/2024 17:15:23",
          "content": "<p>I have the same question</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2600727,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/13/2024 19:53:53",
          "content": "<p>Generally I use Stratified Group KFold in the following two situations:</p>\n<ul>\n<li>one target class is very rare and we want to make sure to include this target class in both train and valid (for each K fold split)</li>\n<li>the test data has the same proportion of target classes as train data</li>\n</ul>\n<p>In this competition, we are not sure if the test data has the same proportions (of 6 target classes) as train data. Therefore if we use Group KFold (instead of Stratified Group KFold) then we are evaluating how well our model can perform when the test data may have slightly different proportions as train data.</p>\n<p>If we use Stratified Group KFold, then we are optimizing a model to make predictions on test data that has the same proportion of target classes as train data. In conclusion, they will both work well, but perhaps Group KFold will produce a model that generalizes slightly better to an unknown test proportion.</p>\n<p>(Note we can probe the public test target class proportions, but we cannot know the private test target class proportions).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2600888,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "01/14/2024 01:43:26",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the detailed explanation. is if-else approach for probe the public test target class proportions?</p>\n<p><em>Note</em>: <strong>No major difference in local CV with StratifiedGroupKFold vs GroupKFold.</strong>. I will go with your CV approach as it correlate with LB except gap between Local CV and LB is [0.2 - 0.15] as per your notebook (14Jan).</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2599910,
      "author_name": "gentlezdh",
      "author_url": "",
      "post_date": "01/13/2024 07:41:38",
      "content": "<p>It's great to see you again in this competition😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2600139,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "01/13/2024 12:02:47",
      "content": "<p>Since Chris has started sharing from very beginning of this competition this is going to be exciting filled with lot of learnings :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2600966,
      "author_name": "krithikramesh",
      "author_url": "",
      "post_date": "01/14/2024 03:28:49",
      "content": "<p>Out of curiosity has anyone tried building a pytorch variant of this dataloader and training an lstm?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2600997,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/14/2024 03:56:27",
          "content": "<p>I have tried EfficientNet on spectrograms. And i have tried RNN, WaveNet, and Transformer on eegs. So far I cannot beat my CatBoost yet. But i will keep trying.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2601914,
              "author_name": "krithikramesh",
              "author_url": "",
              "post_date": "01/14/2024 19:16:17",
              "content": "<p>Hi Chris, Thank you so much for your response! This is really fascinating. I was wondering if I could take a look at your RNN code, namely the dataloader? Just getting my feet wet with ML and would love to learn from it!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2601971,
                  "author_name": "cdeotte",
                  "author_url": "",
                  "post_date": "01/14/2024 20:06:23",
                  "content": "<p>I will publish it soon when I begin working with eeg again. </p>\n<p>Right now i am working with spectrogram. I got my EfficientNet on spectrogram code to work. It now performs better than my CatBoost on spectrogram. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2601978,
                      "author_name": "seshurajup",
                      "author_url": "",
                      "post_date": "01/14/2024 20:08:39",
                      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> since spectrograms are generated without noise. do we need to do preprocess for spectrogram?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2601994,
                          "author_name": "cdeotte",
                          "author_url": "",
                          "post_date": "01/14/2024 20:16:49",
                          "content": "<p>So far the following preprocess works best with my EfficientNet model. (I learned some of this from Tawara's great starter notebooks)</p>\n<ul>\n<li>arrange the 4 spectrograms flat as 400x300x1 (as opposed to 100x300x4 or 200x600x2 etc etc)</li>\n<li>log transform (i.e. <code>img = np.clip(img,np.exp(-4),np.exp(8)); img = np.log(img)</code>)</li>\n<li>standardize per image (with minus mean divide std, then each img has mean=0 )</li>\n<li>fill nan with 0</li>\n<li>repeat 1 channel to make 3 channel monotone image 400x300x1 =&gt; 400x300x3</li>\n</ul>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2601998,
                              "author_name": "seshurajup",
                              "author_url": "",
                              "post_date": "01/14/2024 20:19:20",
                              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for detailed approach <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, i will spend time on now spectrogram sequences.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2602005,
                                  "author_name": "seshurajup",
                                  "author_url": "",
                                  "post_date": "01/14/2024 20:28:57",
                                  "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I expected eeg sequences are better choice then spectrogram, but training model to learn patterns failing. Is it same for you also about eeg sequence experiments.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2602289,
                                      "author_name": "cdeotte",
                                      "author_url": "",
                                      "post_date": "01/15/2024 03:46:30",
                                      "content": "<p>Yes. My current <code>LB = 0.46</code> only uses spectrogram. All my attempts to use EEG have not improved my CV score. I even created the 4 signals for LL, LP, RR, RP from the 16 pairs. And inputted these signals into WaveNet which should extract frequencies like spectrogram but it did not do better than using train means.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2602308,
                                          "author_name": "seshurajup",
                                          "author_url": "",
                                          "post_date": "01/15/2024 04:08:09",
                                          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for conforming about EEG sequences</p>",
                                          "votes": null,
                                          "replies": []
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2601786,
      "author_name": "tanishqdublish",
      "author_url": "",
      "post_date": "01/14/2024 17:06:58",
      "content": "<p>Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2601849,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/14/2024 18:06:52",
          "content": "<p>They both work well. For my full answer see my other comment below <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467576#2600727\" target=\"_blank\">link</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2602010,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/14/2024 20:30:25",
          "content": "<p><a href=\"https://www.kaggle.com/tanishqdublish\" target=\"_blank\">@tanishqdublish</a> Both public LB top notebooks as 15Jan used GroupKFold &amp; StratifiedGroupKFold -&gt; got same 0.67 score with different approach :)<br>\n<a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">Notebook - CatBoost Starter - [LB 0.67]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  =&gt; <strong>GroupKFold</strong><br>\n<a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training\" target=\"_blank\">Notebook - HMS-HBAC: ResNet34d Baseline [Training]</a> by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> =&gt; <strong>StratifiedGroupKFold</strong></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2602212,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "01/15/2024 01:34:31",
      "content": "<p>So…it would seem this would be yet another competition where <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is gonna destroy the LB. Why do I have flashbacks lol</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2602364,
      "author_name": "devitachi",
      "author_url": "",
      "post_date": "01/15/2024 05:13:45",
      "content": "<p>Sir i don't think , such parameters  exists. could to please elaborate the following term….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2605710,
      "author_name": "niuweikun",
      "author_url": "",
      "post_date": "01/17/2024 07:37:44",
      "content": "<p>Hi，thx for your codes, which really helps me alot. Some lines that I can not quite understand:</p>\n<pre><code>     = int( (row['min'] + row['max'])// ) \n    \n     = np.nanmean(spectrograms[row.spec_id][r:r+,:],axis=)\n    [k,:] = x\n     = np.nanmin(spectrograms[row.spec_id][r:r+,:],axis=)\n    [k,:] = x\n\n    \n     = np.nanmean(spectrograms[row.spec_id][r+:r+,:],axis=)\n    [k,:] = x\n     = np.nanmin(spectrograms[row.spec_id][r+:r+,:],axis=)\n    [k,:] = x\n</code></pre>\n<p>So how come  r = int( (row['min'] + row['max'])//4 ) and r+145? <br>\nlooking forward your reply.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2605715,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/17/2024 07:44:48",
          "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> in train.pqt, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> keep maintaining min and max sub-sequence offsets</p>\n<blockquote>\n  <p>r = int( (row['min'] + row['max'])//4 )  =&gt; <strong>mid point is (min+max)/2 but we have array where each unit is 2secs so r =  ((min+max)/2)/2</strong></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>r+145:r+155 =&gt; <strong>10 units which is 20secs starts from offset \"r\" (mid sub-sequence)</strong></p>\n</blockquote>",
          "votes": null,
          "replies": [
            {
              "id": 2605727,
              "author_name": "niuweikun",
              "author_url": "",
              "post_date": "01/17/2024 07:55:20",
              "content": "<p>wow great！Thank you so much😀❤️</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2605756,
              "author_name": "niuweikun",
              "author_url": "",
              "post_date": "01/17/2024 08:11:24",
              "content": "<p>so 145 in r+145 is randomly selected? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2605767,
                  "author_name": "seshurajup",
                  "author_url": "",
                  "post_date": "01/17/2024 08:22:29",
                  "content": "<p><a href=\"https://www.kaggle.com/niuweikun\" target=\"_blank\">@niuweikun</a> spectrogram is 10mins i.e 600secs so 300 units, midpoint is 150 so 145:155 is 20secs</p>\n<p>Go through Discussion - <strong>Understanding Competition Data and EfficientNetB2 Starter - LB 0.57</strong></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2605781,
                      "author_name": "niuweikun",
                      "author_url": "",
                      "post_date": "01/17/2024 08:35:26",
                      "content": "<p>really cool. I am so lucky to meet u here♥️</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2616799,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/23/2024 20:37:35",
      "content": "<p><strong>UPDATE</strong> I added EEG waveform features and boosted CV score and LB score!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2620620,
      "author_name": "stefanoclss",
      "author_url": "",
      "post_date": "01/26/2024 08:41:44",
      "content": "<p>I'm new in this competition, so <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> if I'm right you first found this approach which outperformed all the rest before settling on the effnet and wavenet approach? Very nice notebook btw, currently trying to improve your catboost notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2620668,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/26/2024 09:25:40",
          "content": "<p>Yes. I created CatBoost first. It was quicker to create, train, and infer. Also by playing around with CatBoost features and looking at feature importance, we begin to see what is important. This helps us create a better architecture for our EfficientNet model.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2621064,
              "author_name": "stefanoclss",
              "author_url": "",
              "post_date": "01/26/2024 14:59:25",
              "content": "<p>Okay nice, I'll have a look in your other notebooks to and see how you incorporated the most important features from your catboost starter.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2627238,
      "author_name": "heimdahl",
      "author_url": "",
      "post_date": "01/30/2024 14:53:23",
      "content": "<p>Greetings. <br>\nI am new to competitions as a whole and wonder, how exactly have you \"zipped\" spectogram parquets?<br>\nAs I understood you \"jsoned\" them to specific format with numpy arrays as values. Can I read about this \".npy\" anywhere?<br>\nAnd are there any ideas about eeg.parquets reduction?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2627247,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/30/2024 15:07:31",
          "content": "<p>Correct. I have not compressed them. I have just converted each parquet into a NumPy array. Then I put all the NumPy arrays into a single Python dictionary. Finally i saved with <code>np.save()</code>. If we want to reduce the file size on disk, we can compress with NumPy using</p>\n<pre><code>np.savez\n</code></pre>\n<p>Also note that all my starter notebooks only use the middle 300 rows of the Kaggle spectrogram parquets (which is middle 10 minutes). So we can also just save the middle 300 rows instead of all the rows of data to reduce disk size more. For my EEG spectrogram dataset, I also use <code>np.save()</code>. We can do similar things to compress them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2599691": "Hi everyone. I published a CatBoost starter notebook [here][1]. It trains a model using simple features from spectrograms (and does not use features from eeg parquets). It achieves CV 0.74 and LB 0.60! 🎉\n\nI also published a Kaggle dataset [here][2] with all 11138 spectrograms in one single file. This makes reading all the spectrograms 10x faster to facilitate faster experimentation! And I converted all the 17089 raw eeg waveforms into spectrograms and uploaded them to a dataset [here][3]. These new spectrograms boost the CV score and LB score of all models that are only using Kaggle spectrograms! 🎉\n\n# How CatBoost Works\nIn this competition, the `train.csv` file has 17,089 unique `eeg_id`. The reason that `train.csv` has 106,800 rows is because we get multiple time windows into these 17,089 unique eeq. So we can think of the 100k rows as data augmentation random crops into 17k train samples.\n\nThe CatBoost starter notebook only uses 17,089 train samples to train the 5-Fold model. For each unique `eeg_id`, we use multiclass target with one class equal 1 and the other 5 classes equal 0 as indicated from `train.csv`. Then for each unique `eeg_id`, we create features from the middle 10 minutes of the associated spectrogram.\n\nOur features are simple. A 10 minute spectrogram has 300 readings (taking every 2 seconds). Readings are taken for 100 frequencies from 4 quadrants of the brain. We take the average over time of each of these 400 time series. This produces 400 features to be used with each `eeg_id`.\n\n## UPDATE 1: \nVersion 2 uses `means` and `mins`. And Version 2 uses both 10 minute time window and 20 second time window.\n\n## UPDATE 2:\nVersion 3 adds features from EEG waveforms and boosts CV from CV 0.82 => CV 0.74. And boosts LB 0.67 => LB 0.60. Hooray!\n\nWe use these features to train our CatBoost model. There are many ways to improve CV score and LB score. From the spectrograms, we can create lots of more features. We can also create 10 (or whatever) samples for each `eeg_id` where each sample uses a different 10 minute spectrogram window (i.e. this would be like random crop data augmentation). This would create 10x more data. Also we can create features from the eeg parquets.\n\n[1]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-8\n[2]: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n[3]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms",
    "2599845": "cdeotte Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?",
    "2599910": "It's great to see you again in this competition😀",
    "2600139": "Since Chris has started sharing from very beginning of this competition this is going to be exciting filled with lot of learnings :)",
    "2600491": "I have the same question",
    "2600727": "Generally I use Stratified Group KFold in the following two situations:\n* one target class is very rare and we want to make sure to include this target class in both train and valid (for each K fold split)\n* the test data has the same proportion of target classes as train data\n\nIn this competition, we are not sure if the test data has the same proportions (of 6 target classes) as train data. Therefore if we use Group KFold (instead of Stratified Group KFold) then we are evaluating how well our model can perform when the test data may have slightly different proportions as train data.\n\nIf we use Stratified Group KFold, then we are optimizing a model to make predictions on test data that has the same proportion of target classes as train data. In conclusion, they will both work well, but perhaps Group KFold will produce a model that generalizes slightly better to an unknown test proportion.\n\n(Note we can probe the public test target class proportions, but we cannot know the private test target class proportions).",
    "2600888": "cdeotte Thanks for the detailed explanation. is if-else approach for probe the public test target class proportions?\n\n*Note*: **No major difference in local CV with StratifiedGroupKFold vs GroupKFold.**. I will go with your CV approach as it correlate with LB except gap between Local CV and LB is [0.2 - 0.15] as per your notebook (14Jan).",
    "2600966": "Out of curiosity has anyone tried building a pytorch variant of this dataloader and training an lstm?",
    "2600997": "I have tried EfficientNet on spectrograms. And i have tried RNN, WaveNet, and Transformer on eegs. So far I cannot beat my CatBoost yet. But i will keep trying.",
    "2601786": "Thanks for sharing. is any reason for GroupKFold instead StratifiedGroupKFold ?",
    "2601849": "They both work well. For my full answer see my other comment below [link][1]\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467576#2600727",
    "2601914": "Hi Chris, Thank you so much for your response! This is really fascinating. I was wondering if I could take a look at your RNN code, namely the dataloader? Just getting my feet wet with ML and would love to learn from it!",
    "2601971": "I will publish it soon when I begin working with eeg again. \n\nRight now i am working with spectrogram. I got my EfficientNet on spectrogram code to work. It now performs better than my CatBoost on spectrogram.",
    "2601978": "cdeotte since spectrograms are generated without noise. do we need to do preprocess for spectrogram?",
    "2601994": "So far the following preprocess works best with my EfficientNet model. (I learned some of this from Tawara's great starter notebooks)\n* arrange the 4 spectrograms flat as 400x300x1 (as opposed to 100x300x4 or 200x600x2 etc etc)\n* log transform (i.e. `img = np.clip(img,np.exp(-4),np.exp(8)); img = np.log(img)`)\n* standardize per image (with minus mean divide std, then each img has mean=0 )\n* fill nan with 0\n* repeat 1 channel to make 3 channel monotone image 400x300x1 => 400x300x3",
    "2601998": "cdeotte Thanks for detailed approach @cdeotte, i will spend time on now spectrogram sequences.",
    "2602005": "cdeotte I expected eeg sequences are better choice then spectrogram, but training model to learn patterns failing. Is it same for you also about eeg sequence experiments.",
    "2602010": "tanishqdublish Both public LB top notebooks as 15Jan used GroupKFold & StratifiedGroupKFold -> got same 0.67 score with different approach :)\n[Notebook - CatBoost Starter - [LB 0.67]](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67) by @cdeotte  => **GroupKFold**\n[Notebook - HMS-HBAC: ResNet34d Baseline [Training]](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training) by @ttahara => **StratifiedGroupKFold**",
    "2602212": "So...it would seem this would be yet another competition where @cdeotte is gonna destroy the LB. Why do I have flashbacks lol",
    "2602289": "Yes. My current `LB = 0.46` only uses spectrogram. All my attempts to use EEG have not improved my CV score. I even created the 4 signals for LL, LP, RR, RP from the 16 pairs. And inputted these signals into WaveNet which should extract frequencies like spectrogram but it did not do better than using train means.",
    "2602308": "cdeotte Thanks for conforming about EEG sequences",
    "2602364": "Sir i don't think , such parameters  exists. could to please elaborate the following term....",
    "2605710": "Hi，thx for your codes, which really helps me alot. Some lines that I can not quite understand:\n\n        r = int( (row['min'] + row['max'])//4 ) \n        # 10 MINUTE WINDOW FEATURES (MEANS and MINS)\n        x = np.nanmean(spectrograms[row.spec_id][r:r+300,:],axis=0)\n        data[k,:400] = x\n        x = np.nanmin(spectrograms[row.spec_id][r:r+300,:],axis=0)\n        data[k,400:800] = x\n        \n        # 20 SECOND WINDOW FEATURES (MEANS and MINS)\n        x = np.nanmean(spectrograms[row.spec_id][r+145:r+155,:],axis=0)\n        data[k,800:1200] = x\n        x = np.nanmin(spectrograms[row.spec_id][r+145:r+155,:],axis=0)\n        data[k,1200:1600] = x\n\nSo how come  r = int( (row['min'] + row['max'])//4 ) and r+145? \nlooking forward your reply.",
    "2605715": "niuweikun in train.pqt, @cdeotte keep maintaining min and max sub-sequence offsets\n\n\n> r = int( (row['min'] + row['max'])//4 )  => **mid point is (min+max)/2 but we have array where each unit is 2secs so r =  ((min+max)/2)/2**\n\n---\n>r+145:r+155 => **10 units which is 20secs starts from offset \"r\" (mid sub-sequence)**",
    "2605727": "wow great！Thank you so much😀❤️",
    "2605756": "so 145 in r+145 is randomly selected?",
    "2605767": "niuweikun spectrogram is 10mins i.e 600secs so 300 units, midpoint is 150 so 145:155 is 20secs\n\nGo through Discussion - **Understanding Competition Data and EfficientNetB2 Starter - LB 0.57**",
    "2605781": "really cool. I am so lucky to meet u here♥️",
    "2616799": "**UPDATE** I added EEG waveform features and boosted CV score and LB score!",
    "2620620": "I'm new in this competition, so @cdeotte if I'm right you first found this approach which outperformed all the rest before settling on the effnet and wavenet approach? Very nice notebook btw, currently trying to improve your catboost notebook.",
    "2620668": "Yes. I created CatBoost first. It was quicker to create, train, and infer. Also by playing around with CatBoost features and looking at feature importance, we begin to see what is important. This helps us create a better architecture for our EfficientNet model.",
    "2621064": "Okay nice, I'll have a look in your other notebooks to and see how you incorporated the most important features from your catboost starter.",
    "2627238": "Greetings. \nI am new to competitions as a whole and wonder, how exactly have you \"zipped\" spectogram parquets?\nAs I understood you \"jsoned\" them to specific format with numpy arrays as values. Can I read about this \".npy\" anywhere?\nAnd are there any ideas about eeg.parquets reduction?",
    "2627247": "Correct. I have not compressed them. I have just converted each parquet into a NumPy array. Then I put all the NumPy arrays into a single Python dictionary. Finally i saved with `np.save()`. If we want to reduce the file size on disk, we can compress with NumPy using\n\n    np.savez_compressed('compressed_array.npz', arr=arr)\n\nAlso note that all my starter notebooks only use the middle 300 rows of the Kaggle spectrogram parquets (which is middle 10 minutes). So we can also just save the middle 300 rows instead of all the rows of data to reduce disk size more. For my EEG spectrogram dataset, I also use `np.save()`. We can do similar things to compress them."
  },
  "source": "meta"
}