{
  "id": 492985,
  "title": "70th Place Solution: 5 Model Ensemble Combining 1D-ResNet, Wavenet, Chrononet, and Two Variations of Efficientnet-B0",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492985",
  "author_name": "Jan Brederecke",
  "post_date": "2024-04-11T17:07:37.418000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>First of all, thank you to Kaggle and the hosts for the great competition and everyone for all the great discussions and notebooks!</h1>\n<h2>Team : Brain Freeze</h2>\n<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> <br>\n<a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> <br>\n<a href=\"https://www.kaggle.com/medali1992\" target=\"_blank\">@medali1992</a> <br>\n<a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> <br>\n<a href=\"https://www.kaggle.com/janbrederecke\" target=\"_blank\">@janbrederecke</a> </p>\n<h2>Summary</h2>\n<p>We achieved 70th place by a combination of various 1D models that used the raw EEGs and 2D-CNNs that used spectrograms as the input. The final ensemble achieved public LB: 0.289 / private LB: 0.345 and was our best submission after all.</p>\n<h2>EDA</h2>\n<p>Especially <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> did lots of EDA in the beginning where he tried to extract useful features for tree-based machine learning algorithms. While doing so he found out about the two different distributions and kindly shared this here: <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact\" target=\"_blank\">https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact</a></p>\n<h3>Two parallel approaches</h3>\n<p>As we were unsure if the hidden test set would really be from the distribution of data that was labeled by &gt;=10 physicians, we decided to create stage 1 and stage 2 submissions.</p>\n<h3>Data and CV</h3>\n<p>While individually starting out with the 17k rows approach, it was decided to use the 20k rows approach posted by <a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@seanbearden</a>. Even though we repeatedly discussed using more or all the data available we did make it work as most of us worked on Kaggle GPUs and the limitations made it too inefficient. <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> found out that binning the samples by number of voters using 5 bins resulted in a very even distribution which we used for stratified group k-fold cross-validation using 5 folds yielding a 0.03 LB boost in our experiments for all models.</p>\n<pre><code>bin_edges =   \nnum_bins = (bin_edges) -   \ntrain = pd(train, bins=bin_edges, labels=False)  \npatient_groups = train()()  \nstratified_kfold = (n_splits=CFG.n_fold)  \ntrain_indices =   \nvalid_indices =   \n fold, (train_idx, valid_idx)  (stratified_kfold(X=train, y=train, groups=patient_groups)):  \ntrain_indices(train_idx)  \nvalid_indices(valid_idx)  \ntrain = fold\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2F1211f98a73ee3130de65b2f236eb3205%2FPicture%201.png?generation=1712854341400842&amp;alt=media\"></p>\n<h2>Training strategy and the one big mistake</h2>\n<p>We trained some of the models as different people had different (local) resources and workflows (established individually before the team merge) but generally speaking we trained for most epochs on stage 1 data and then trained a few epochs on stage 2 data only. Likely our one big mistake was to validate the second stage on the same data as in stage 1 which might have led to suboptimal model selection at that stage.</p>\n<h2>Models used</h2>\n<p>For a long time, we tried to make the raw input 1D approach work using variations of 1D-ResNets, Wavenet (<a href=\"https://arxiv.org/abs/1609.03499\" target=\"_blank\">https://arxiv.org/abs/1609.03499</a>), and in the end Chrononet (<a href=\"https://arxiv.org/abs/1802.00308\" target=\"_blank\">https://arxiv.org/abs/1802.00308</a>) as well. We then added different flavors of 2D CNNs that used spectrograms as their input, dual input models were tried but we were not successful enough to keep them in our final ensemble (Now having read about more successful solutions we see what we did wrong…).</p>\n<h3>1D-ResNet (only used in stage1)</h3>\n<p><strong>Input</strong>: 50s raw data downsampled to 50 Hz, preprocessed with 60 Hz notch-filter and 0.5 / 50 Hz Low-High-Cut</p>\n<p>The 1D-ResNet shared by <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> was very useful and we modified it slightly to have more planes and ResBlocks. This model was used in our stage 1 ensemble but not in stage 2 (it is complicated). We used downsampled data with 50 Hz. This model achieved LB .37 when trained for stage 1 on the 17k rows data selection. We never really transferred this to our final data / CV scheme as we found the next model to be a little bit better. We used Adan for LR scheduling in all 1D-ResNet variations.</p>\n<h3>1D-Dual-Input-ResNet with Attention (otherwise known as the Turbomodel)</h3>\n<p><strong>Input</strong>: 50s raw data downsampled to 25 / 50 Hz, 0.5 / 50 Hz Low-High-Cut</p>\n<p>We used a dual input version of the given 1D-Resnet that was made of a large and a small instance of the ResNet-1D. The large one was fed the 50Hz input data and the small one was fed 25Hz data. Each output was given a separate scaled dot product attention mechanism before the individual results were concatenated. This model achieved LB .33 in stage 2 when retrained for 10 epochs on stage 2 data only.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2Fc68ab9af340d695ebbe2402c37cc30c6%2FPicture%202.png?generation=1712854373761825&amp;alt=media\"></p>\n<h3>Wavenet-GRU</h3>\n<p>We used the pytorch wavenet shared by <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and added to it a bidirectional GRU layer followed by a SeqPool that behaves like a mean pooling layer used for transformer.</p>\n<p>We used the Quad/Penta montage with an overall 15/17 features and each chain is evaluated separately by the model.</p>\n<p>Notes on structure of wavenet gru with legs corresponding to parts of the brain LP, RP,FP,BP,CP</p>\n<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220</a>. The CP did not work for us in stage 2 though the Penta model was trained for the same.</p>\n<p><strong>Preprocessing</strong></p>\n<ul>\n<li><p>Apply notch filter (60Hz) followed by forward and backward digital filter</p></li>\n<li><p>Apply bandpass filter (0.7Hz-20Hz) followed by forward and backward digital filter</p></li>\n</ul>\n<p><strong>Training</strong></p>\n<ul>\n<li><p>Train PentaTail :  <a href=\"https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail/edit\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail</a></p></li>\n<li><p>Train Quad : <a href=\"https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267</a></p></li>\n<li><p>Private STAGE 2 Quad and Penta: 0.40</p></li>\n<li><p>Infer Stage 2 :<a href=\"https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167</a></p></li>\n</ul>\n<pre><code>  (nn.Module):  \n      ():  \n        ().__init__()  \n        .dense = nn.Linear(emb_dim, )  \n        .softmax = nn.Softmax(dim=-)  \n\n      ():  \n        bs, seq_len, emb_dim = x.shape  \n        identity = x  \n        x = .dense(x)  \n        x = x.permute(, , )  \n        x = .softmax(x)  \n        x = x @ identity  \n        x = x.reshape(x.shape[], -)  \n         x\n</code></pre>\n<p><strong>Feature code</strong></p>\n<pre><code>def extract_features(self, x):\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x6 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x7 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    #x8 = self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x9 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z1 = torch.mean(torch.stack([x1, x2, x3, x4, x6, x7, x9]), dim=)\n\n\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x5 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x6 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x7 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x9 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z2 = torch.mean(torch.stack([x1, x2, x3, x4, x5, x6, x7, x9]), dim=)\n\n\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z3 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=)\n\n\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z4 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=)\n\n\n    y = torch.cat([z1, z2, z3, z4], dim=)\n</code></pre>\n<h3>Chrononet</h3>\n<p>“ChronoNet is formed by stacking multiple 1D convolution layers followed by deep gated recurrent unit (GRU) layers where each 1D convolution layer uses multiple filters of exponentially varying lengths and the stacked GRU layers are densely connected in a feed-forward manner.” - <a href=\"https://arxiv.org/abs/1802.00308\" target=\"_blank\">https://arxiv.org/abs/1802.00308</a>.</p>\n<h4>ChronoNet Infer :</h4>\n<p><a href=\"https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323\" target=\"_blank\">https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323</a></p>\n<p>Feature code similar to wavenet-gru but the chrononet blocks look like below</p>\n<pre><code>  (nn.Module):\n\n      ():\n\n    ().__init__()\n\n    self.inception_block1=InceptionBlock(channel)        \n\n    self.inception_block2=InceptionBlock() \n\n    self.inception_block3=InceptionBlock() \n\n    self.gru1 = nn.GRU(input_size =  , hidden_size =  , batch_first =  , bidirectional=) \n\n    self.gru2 = nn.GRU(input_size =  *,  hidden_size =  , batch_first =  , bidirectional=) \n\n    self.gru3 = nn.GRU(input_size =  *, hidden_size =  , batch_first =  , bidirectional=) \n\n    self.gru4 = nn.GRU(input_size =  *, hidden_size =  , batch_first =  , bidirectional=) \n\n    self.relu = nn.ReLU() \n\n    self.gru_linear=nn.Linear(in_features =  , out_features =  ) \n\n    self.flatten = nn.Flatten() \n\n    self.seqpool = SeqPool()\n\n    self.fc1 = nn.Linear(,) \n\n\n\n     (): \n\n    x = x.permute(, , )\n\n    x=self.inception_block1(x) \n\n    x=self.inception_block2(x) \n\n    x=self.inception_block3(x) \n\n    x=x.permute(,,) \n\n    gru_out1,_=self.gru1(x) \n\n    gru_out2,_=self.gru2(gru_out1) \n\n    gru_out=torch.cat((gru_out1, gru_out2), dim =  ) \n\n    gru_out3,_=self.gru3(gru_out) \n\n    gru_out = torch.cat((gru_out1, gru_out2, gru_out3), dim =  ) \n\n    gru_out4,_=self.gru4(gru_out) \n\n    seqpool =  self.seqpool(gru_out4)\n\n     seqpool \n</code></pre>\n<h3>Efficientnet</h3>\n<p>**Public Model shared by <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> **</p>\n<p>After this model was shared shortly before the deadline we decided to tweak it a little bit and just incorporate it as this was a great solution. Thanks <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>!</p>\n<p><strong>Pytorch Zimmermann model</strong></p>\n<p>For the CNN model we used <code>tf_efficientnet_b0.ns_jft_in1k</code>backbone from timm followed by GEM pooling. What helped improved the model performance is the usage of hflip+XYMasking and using a probability which dictates if mixup should be applied or not, and when it is applied no other augmentations are applied.</p>\n<p>Using this trick + Zimmermann specs gave a boost around 0.1 (@alejopaullier trick).</p>\n<h3>Ensemble</h3>\n<p>We tried different procedures for the ensembles and went with a mixture of OOF-predictions and Nelder-Mead optimization and manual weights for the weighting as some of our models did not have OOF-predictions (pipeline from before teammerge and time was running out).</p>\n<p>Our final ensemble was weighted as follows:</p>\n<p>(Dual-input 1D-ResNet + Wavenet-GRU + Chrononet) / 3 * 0.39 +</p>\n<p>Efficientnet-B0_melspecs * 0.30 +</p>\n<p>Efficientnet-B0_zimmermann_specs * 0.31</p>\n<h2>Little “tricks” and augmentations we used</h2>\n<p>We used a binary auxiliary target that discriminated those with more than 9 voters from those with less and it helped especially in stage 1.</p>\n<p>The 1D-ResNet-based models were trained using heavy augmentation which can be found in the dataset class in the submission. The augmentations include:</p>\n<ul>\n<li><p>Horizontal flip</p></li>\n<li><p>Slightly moving the single channels against each other randomly</p></li>\n<li><p>Masking parts of the signal, especially outside the middle 10 s</p></li>\n<li><p>Dropout of single channels</p></li>\n<li><p>Random increase/decrease of wave amplitude</p></li>\n<li><p>Gaussian noise</p></li>\n</ul>\n<p>Cutmix and Mixup were also used heavily for their training which was done mostly using OneCycleLR-Scheduler in the end. Preprocessing for the 1D-Resnets was simple using the eight double banana features Chris proposed as all experiments we did with other features yielded worse results for the 1D-ResNet-based models.</p>\n<h3>Lessons learned</h3>\n<p>Maxing out the batch size and tuning the learning rate to match it will help you to be able to smoothly run more experiments. Using OneCycleLR scheduler as a starter is also not a bad idea. Looking at all the top solutions, using the raw input and turning it in whichever way to 2D input would have been amazing but maybe we did not have the trust that this would be a good idea as we definitely thought about similar things - so: If you think it, try it!</p>",
  "messages": [
    {
      "id": 2747089,
      "postDate": "2024-04-11T17:07:37.417Z",
      "content": "<h1>First of all, thank you to Kaggle and the hosts for the great competition and everyone for all the great discussions and notebooks!</h1>\n<h2>Team : Brain Freeze</h2>\n<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> <br>\n<a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> <br>\n<a href=\"https://www.kaggle.com/medali1992\" target=\"_blank\">@medali1992</a> <br>\n<a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> <br>\n<a href=\"https://www.kaggle.com/janbrederecke\" target=\"_blank\">@janbrederecke</a> </p>\n<h2>Summary</h2>\n<p>We achieved 70th place by a combination of various 1D models that used the raw EEGs and 2D-CNNs that used spectrograms as the input. The final ensemble achieved public LB: 0.289 / private LB: 0.345 and was our best submission after all.</p>\n<h2>EDA</h2>\n<p>Especially <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> did lots of EDA in the beginning where he tried to extract useful features for tree-based machine learning algorithms. While doing so he found out about the two different distributions and kindly shared this here: <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact\" target=\"_blank\">https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact</a></p>\n<h3>Two parallel approaches</h3>\n<p>As we were unsure if the hidden test set would really be from the distribution of data that was labeled by &gt;=10 physicians, we decided to create stage 1 and stage 2 submissions.</p>\n<h3>Data and CV</h3>\n<p>While individually starting out with the 17k rows approach, it was decided to use the 20k rows approach posted by <a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@seanbearden</a>. Even though we repeatedly discussed using more or all the data available we did make it work as most of us worked on Kaggle GPUs and the limitations made it too inefficient. <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> found out that binning the samples by number of voters using 5 bins resulted in a very even distribution which we used for stratified group k-fold cross-validation using 5 folds yielding a 0.03 LB boost in our experiments for all models.</p>\n<pre><code>bin_edges =   \nnum_bins = (bin_edges) -   \ntrain = pd(train, bins=bin_edges, labels=False)  \npatient_groups = train()()  \nstratified_kfold = (n_splits=CFG.n_fold)  \ntrain_indices =   \nvalid_indices =   \n fold, (train_idx, valid_idx)  (stratified_kfold(X=train, y=train, groups=patient_groups)):  \ntrain_indices(train_idx)  \nvalid_indices(valid_idx)  \ntrain = fold\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2F1211f98a73ee3130de65b2f236eb3205%2FPicture%201.png?generation=1712854341400842&amp;alt=media\"></p>\n<h2>Training strategy and the one big mistake</h2>\n<p>We trained some of the models as different people had different (local) resources and workflows (established individually before the team merge) but generally speaking we trained for most epochs on stage 1 data and then trained a few epochs on stage 2 data only. Likely our one big mistake was to validate the second stage on the same data as in stage 1 which might have led to suboptimal model selection at that stage.</p>\n<h2>Models used</h2>\n<p>For a long time, we tried to make the raw input 1D approach work using variations of 1D-ResNets, Wavenet (<a href=\"https://arxiv.org/abs/1609.03499\" target=\"_blank\">https://arxiv.org/abs/1609.03499</a>), and in the end Chrononet (<a href=\"https://arxiv.org/abs/1802.00308\" target=\"_blank\">https://arxiv.org/abs/1802.00308</a>) as well. We then added different flavors of 2D CNNs that used spectrograms as their input, dual input models were tried but we were not successful enough to keep them in our final ensemble (Now having read about more successful solutions we see what we did wrong…).</p>\n<h3>1D-ResNet (only used in stage1)</h3>\n<p><strong>Input</strong>: 50s raw data downsampled to 50 Hz, preprocessed with 60 Hz notch-filter and 0.5 / 50 Hz Low-High-Cut</p>\n<p>The 1D-ResNet shared by <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> was very useful and we modified it slightly to have more planes and ResBlocks. This model was used in our stage 1 ensemble but not in stage 2 (it is complicated). We used downsampled data with 50 Hz. This model achieved LB .37 when trained for stage 1 on the 17k rows data selection. We never really transferred this to our final data / CV scheme as we found the next model to be a little bit better. We used Adan for LR scheduling in all 1D-ResNet variations.</p>\n<h3>1D-Dual-Input-ResNet with Attention (otherwise known as the Turbomodel)</h3>\n<p><strong>Input</strong>: 50s raw data downsampled to 25 / 50 Hz, 0.5 / 50 Hz Low-High-Cut</p>\n<p>We used a dual input version of the given 1D-Resnet that was made of a large and a small instance of the ResNet-1D. The large one was fed the 50Hz input data and the small one was fed 25Hz data. Each output was given a separate scaled dot product attention mechanism before the individual results were concatenated. This model achieved LB .33 in stage 2 when retrained for 10 epochs on stage 2 data only.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2Fc68ab9af340d695ebbe2402c37cc30c6%2FPicture%202.png?generation=1712854373761825&amp;alt=media\"></p>\n<h3>Wavenet-GRU</h3>\n<p>We used the pytorch wavenet shared by <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and added to it a bidirectional GRU layer followed by a SeqPool that behaves like a mean pooling layer used for transformer.</p>\n<p>We used the Quad/Penta montage with an overall 15/17 features and each chain is evaluated separately by the model.</p>\n<p>Notes on structure of wavenet gru with legs corresponding to parts of the brain LP, RP,FP,BP,CP</p>\n<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220</a>. The CP did not work for us in stage 2 though the Penta model was trained for the same.</p>\n<p><strong>Preprocessing</strong></p>\n<ul>\n<li><p>Apply notch filter (60Hz) followed by forward and backward digital filter</p></li>\n<li><p>Apply bandpass filter (0.7Hz-20Hz) followed by forward and backward digital filter</p></li>\n</ul>\n<p><strong>Training</strong></p>\n<ul>\n<li><p>Train PentaTail :  <a href=\"https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail/edit\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail</a></p></li>\n<li><p>Train Quad : <a href=\"https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267</a></p></li>\n<li><p>Private STAGE 2 Quad and Penta: 0.40</p></li>\n<li><p>Infer Stage 2 :<a href=\"https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167</a></p></li>\n</ul>\n<pre><code>  (nn.Module):  \n      ():  \n        ().__init__()  \n        .dense = nn.Linear(emb_dim, )  \n        .softmax = nn.Softmax(dim=-)  \n\n      ():  \n        bs, seq_len, emb_dim = x.shape  \n        identity = x  \n        x = .dense(x)  \n        x = x.permute(, , )  \n        x = .softmax(x)  \n        x = x @ identity  \n        x = x.reshape(x.shape[], -)  \n         x\n</code></pre>\n<p><strong>Feature code</strong></p>\n<pre><code>def extract_features(self, x):\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x6 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x7 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    #x8 = self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x9 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z1 = torch.mean(torch.stack([x1, x2, x3, x4, x6, x7, x9]), dim=)\n\n\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x5 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x6 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x7 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x9 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z2 = torch.mean(torch.stack([x1, x2, x3, x4, x5, x6, x7, x9]), dim=)\n\n\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z3 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=)\n\n\n\n    #  part\n\n    x1 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x2 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x3 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    x4 =  self.model(x[:, :, self.index_dict[(, )]:self.index_dict[(, )]+])\n\n    z4 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=)\n\n\n    y = torch.cat([z1, z2, z3, z4], dim=)\n</code></pre>\n<h3>Chrononet</h3>\n<p>“ChronoNet is formed by stacking multiple 1D convolution layers followed by deep gated recurrent unit (GRU) layers where each 1D convolution layer uses multiple filters of exponentially varying lengths and the stacked GRU layers are densely connected in a feed-forward manner.” - <a href=\"https://arxiv.org/abs/1802.00308\" target=\"_blank\">https://arxiv.org/abs/1802.00308</a>.</p>\n<h4>ChronoNet Infer :</h4>\n<p><a href=\"https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323\" target=\"_blank\">https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323</a></p>\n<p>Feature code similar to wavenet-gru but the chrononet blocks look like below</p>\n<pre><code>  (nn.Module):\n\n      ():\n\n    ().__init__()\n\n    self.inception_block1=InceptionBlock(channel)        \n\n    self.inception_block2=InceptionBlock() \n\n    self.inception_block3=InceptionBlock() \n\n    self.gru1 = nn.GRU(input_size =  , hidden_size =  , batch_first =  , bidirectional=) \n\n    self.gru2 = nn.GRU(input_size =  *,  hidden_size =  , batch_first =  , bidirectional=) \n\n    self.gru3 = nn.GRU(input_size =  *, hidden_size =  , batch_first =  , bidirectional=) \n\n    self.gru4 = nn.GRU(input_size =  *, hidden_size =  , batch_first =  , bidirectional=) \n\n    self.relu = nn.ReLU() \n\n    self.gru_linear=nn.Linear(in_features =  , out_features =  ) \n\n    self.flatten = nn.Flatten() \n\n    self.seqpool = SeqPool()\n\n    self.fc1 = nn.Linear(,) \n\n\n\n     (): \n\n    x = x.permute(, , )\n\n    x=self.inception_block1(x) \n\n    x=self.inception_block2(x) \n\n    x=self.inception_block3(x) \n\n    x=x.permute(,,) \n\n    gru_out1,_=self.gru1(x) \n\n    gru_out2,_=self.gru2(gru_out1) \n\n    gru_out=torch.cat((gru_out1, gru_out2), dim =  ) \n\n    gru_out3,_=self.gru3(gru_out) \n\n    gru_out = torch.cat((gru_out1, gru_out2, gru_out3), dim =  ) \n\n    gru_out4,_=self.gru4(gru_out) \n\n    seqpool =  self.seqpool(gru_out4)\n\n     seqpool \n</code></pre>\n<h3>Efficientnet</h3>\n<p>**Public Model shared by <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> **</p>\n<p>After this model was shared shortly before the deadline we decided to tweak it a little bit and just incorporate it as this was a great solution. Thanks <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>!</p>\n<p><strong>Pytorch Zimmermann model</strong></p>\n<p>For the CNN model we used <code>tf_efficientnet_b0.ns_jft_in1k</code>backbone from timm followed by GEM pooling. What helped improved the model performance is the usage of hflip+XYMasking and using a probability which dictates if mixup should be applied or not, and when it is applied no other augmentations are applied.</p>\n<p>Using this trick + Zimmermann specs gave a boost around 0.1 (@alejopaullier trick).</p>\n<h3>Ensemble</h3>\n<p>We tried different procedures for the ensembles and went with a mixture of OOF-predictions and Nelder-Mead optimization and manual weights for the weighting as some of our models did not have OOF-predictions (pipeline from before teammerge and time was running out).</p>\n<p>Our final ensemble was weighted as follows:</p>\n<p>(Dual-input 1D-ResNet + Wavenet-GRU + Chrononet) / 3 * 0.39 +</p>\n<p>Efficientnet-B0_melspecs * 0.30 +</p>\n<p>Efficientnet-B0_zimmermann_specs * 0.31</p>\n<h2>Little “tricks” and augmentations we used</h2>\n<p>We used a binary auxiliary target that discriminated those with more than 9 voters from those with less and it helped especially in stage 1.</p>\n<p>The 1D-ResNet-based models were trained using heavy augmentation which can be found in the dataset class in the submission. The augmentations include:</p>\n<ul>\n<li><p>Horizontal flip</p></li>\n<li><p>Slightly moving the single channels against each other randomly</p></li>\n<li><p>Masking parts of the signal, especially outside the middle 10 s</p></li>\n<li><p>Dropout of single channels</p></li>\n<li><p>Random increase/decrease of wave amplitude</p></li>\n<li><p>Gaussian noise</p></li>\n</ul>\n<p>Cutmix and Mixup were also used heavily for their training which was done mostly using OneCycleLR-Scheduler in the end. Preprocessing for the 1D-Resnets was simple using the eight double banana features Chris proposed as all experiments we did with other features yielded worse results for the 1D-ResNet-based models.</p>\n<h3>Lessons learned</h3>\n<p>Maxing out the batch size and tuning the learning rate to match it will help you to be able to smoothly run more experiments. Using OneCycleLR scheduler as a starter is also not a bad idea. Looking at all the top solutions, using the raw input and turning it in whichever way to 2D input would have been amazing but maybe we did not have the trust that this would be a good idea as we definitely thought about similar things - so: If you think it, try it!</p>",
      "rawMarkdown": "# First of all, thank you to Kaggle and the hosts for the great competition and everyone for all the great discussions and notebooks!\n\n## Team : Brain Freeze\n@alejopaullier \n@gauravbrills \n@medali1992 \n@pcjimmmy \n@janbrederecke \n\n## Summary\n\nWe achieved 70th place by a combination of various 1D models that used the raw EEGs and 2D-CNNs that used spectrograms as the input. The final ensemble achieved public LB: 0.289 / private LB: 0.345 and was our best submission after all.\n\n  \n\n## EDA\n\nEspecially [@pcjimmmy](https://www.kaggle.com/pcjimmmy) did lots of EDA in the beginning where he tried to extract useful features for tree-based machine learning algorithms. While doing so he found out about the two different distributions and kindly shared this here: [https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact)\n\n  \n\n### Two parallel approaches\n\nAs we were unsure if the hidden test set would really be from the distribution of data that was labeled by >=10 physicians, we decided to create stage 1 and stage 2 submissions.\n\n  \n\n### Data and CV\n\nWhile individually starting out with the 17k rows approach, it was decided to use the 20k rows approach posted by @seanbearden. Even though we repeatedly discussed using more or all the data available we did make it work as most of us worked on Kaggle GPUs and the limitations made it too inefficient. @gauravbrills found out that binning the samples by number of voters using 5 bins resulted in a very even distribution which we used for stratified group k-fold cross-validation using 5 folds yielding a 0.03 LB boost in our experiments for all models.\n\n```\nbin_edges = [0, 5, 10, 15, 20, np.inf]  \nnum_bins = len(bin_edges) - 1  \ntrain['evaluator_bin'] = pd.cut(train['total_evaluators'], bins=bin_edges, labels=False)  \npatient_groups = train.groupby('patient_id').ngroup()  \nstratified_kfold = StratifiedGroupKFold(n_splits=CFG.n_fold)  \ntrain_indices = []  \nvalid_indices = []  \nfor fold, (train_idx, valid_idx) in enumerate(stratified_kfold.split(X=train, y=train['evaluator_bin'], groups=patient_groups)):  \ntrain_indices.append(train_idx)  \nvalid_indices.append(valid_idx)  \ntrain.loc[valid_idx, \"fold\"] = fold\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2F1211f98a73ee3130de65b2f236eb3205%2FPicture%201.png?generation=1712854341400842&alt=media)\n\n## Training strategy and the one big mistake\n\nWe trained some of the models as different people had different (local) resources and workflows (established individually before the team merge) but generally speaking we trained for most epochs on stage 1 data and then trained a few epochs on stage 2 data only. Likely our one big mistake was to validate the second stage on the same data as in stage 1 which might have led to suboptimal model selection at that stage.\n\n## Models used\n\nFor a long time, we tried to make the raw input 1D approach work using variations of 1D-ResNets, Wavenet ([https://arxiv.org/abs/1609.03499](https://arxiv.org/abs/1609.03499)), and in the end Chrononet ([https://arxiv.org/abs/1802.00308](https://arxiv.org/abs/1802.00308)) as well. We then added different flavors of 2D CNNs that used spectrograms as their input, dual input models were tried but we were not successful enough to keep them in our final ensemble (Now having read about more successful solutions we see what we did wrong…).\n\n  \n\n### 1D-ResNet (only used in stage1)\n\n**Input**: 50s raw data downsampled to 50 Hz, preprocessed with 60 Hz notch-filter and 0.5 / 50 Hz Low-High-Cut\n\nThe 1D-ResNet shared by @nischaydnk was very useful and we modified it slightly to have more planes and ResBlocks. This model was used in our stage 1 ensemble but not in stage 2 (it is complicated). We used downsampled data with 50 Hz. This model achieved LB .37 when trained for stage 1 on the 17k rows data selection. We never really transferred this to our final data / CV scheme as we found the next model to be a little bit better. We used Adan for LR scheduling in all 1D-ResNet variations.\n\n  \n\n### 1D-Dual-Input-ResNet with Attention (otherwise known as the Turbomodel)\n\n**Input**: 50s raw data downsampled to 25 / 50 Hz, 0.5 / 50 Hz Low-High-Cut\n\nWe used a dual input version of the given 1D-Resnet that was made of a large and a small instance of the ResNet-1D. The large one was fed the 50Hz input data and the small one was fed 25Hz data. Each output was given a separate scaled dot product attention mechanism before the individual results were concatenated. This model achieved LB .33 in stage 2 when retrained for 10 epochs on stage 2 data only.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2Fc68ab9af340d695ebbe2402c37cc30c6%2FPicture%202.png?generation=1712854373761825&alt=media)\n\n\n\n### Wavenet-GRU\n\nWe used the pytorch wavenet shared by @alejopaullier @cdeotte and added to it a bidirectional GRU layer followed by a SeqPool that behaves like a mean pooling layer used for transformer.\n\nWe used the Quad/Penta montage with an overall 15/17 features and each chain is evaluated separately by the model.\n\nNotes on structure of wavenet gru with legs corresponding to parts of the brain LP, RP,FP,BP,CP\n\n[https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220). The CP did not work for us in stage 2 though the Penta model was trained for the same.\n\n**Preprocessing**\n\n-   Apply notch filter (60Hz) followed by forward and backward digital filter\n    \n-   Apply bandpass filter (0.7Hz-20Hz) followed by forward and backward digital filter\n    \n\n**Training**\n\n-   Train PentaTail :  [https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail](https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail/edit)\n    \n-   Train Quad : [https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267](https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267)\n    \n-   Private STAGE 2 Quad and Penta: 0.40\n    \n-   Infer Stage 2 :[https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167](https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167)\n\n```\nclass  SeqPool(nn.Module):  \n\tdef  __init__(self, emb_dim=192):  \n\t\tsuper().__init__()  \n\t\tself.dense = nn.Linear(emb_dim, 1)  \n\t\tself.softmax = nn.Softmax(dim=-1)  \n  \n\tdef  forward(self, x):  \n\t\tbs, seq_len, emb_dim = x.shape  \n\t\tidentity = x  \n\t\tx = self.dense(x)  \n\t\tx = x.permute(0, 2, 1)  \n\t\tx = self.softmax(x)  \n\t\tx = x @ identity  \n\t\tx = x.reshape(x.shape[0], -1)  \n\t\treturn x\n```\n**Feature code**\n```\ndef extract_features(self, x):\n\n\t# Left part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('Fp1', 'F7')]:self.index_dict[('Fp1', 'F7')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('F7', 'T3')]:self.index_dict[('F7', 'T3')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('T3', 'T5')]:self.index_dict[('T3', 'T5')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('T5', 'O1')]:self.index_dict[('T5', 'O1')]+1])\n\n\tx6 =  self.model(x[:, :, self.index_dict[('Fp1', 'F3')]:self.index_dict[('Fp1', 'F3')]+1])\n\n\tx7 =  self.model(x[:, :, self.index_dict[('F3', 'C3')]:self.index_dict[('F3', 'C3')]+1])\n\n\t#x8 = self.model(x[:, :, self.index_dict[('C3', 'P3')]:self.index_dict[('C3', 'P3')]+1])\n\n\tx9 =  self.model(x[:, :, self.index_dict[('P3', 'O1')]:self.index_dict[('P3', 'O1')]+1])\n\n\tz1 = torch.mean(torch.stack([x1, x2, x3, x4, x6, x7, x9]), dim=0)\n\n  \n\n\t# Right part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('Fp2', 'F8')]:self.index_dict[('Fp2', 'F8')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('F8', 'T4')]:self.index_dict[('F8', 'T4')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('T4', 'T6')]:self.index_dict[('T4', 'T6')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('T6', 'O2')]:self.index_dict[('T6', 'O2')]+1])\n\n\tx5 =  self.model(x[:, :, self.index_dict[('C4', 'T4')]:self.index_dict[('C4', 'T4')]+1])\n\n\tx6 =  self.model(x[:, :, self.index_dict[('Fp2', 'F4')]:self.index_dict[('Fp2', 'F4')]+1])\n\n\tx7 =  self.model(x[:, :, self.index_dict[('F4', 'C4')]:self.index_dict[('F4', 'C4')]+1])\n\n\tx9 =  self.model(x[:, :, self.index_dict[('P4', 'O2')]:self.index_dict[('P4', 'O2')]+1])\n\n\tz2 = torch.mean(torch.stack([x1, x2, x3, x4, x5, x6, x7, x9]), dim=0)\n\n  \n\n\t# Front part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('Fp1', 'F7')]:self.index_dict[('Fp1', 'F7')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('Fp2', 'F8')]:self.index_dict[('Fp2', 'F8')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('Fp1', 'F3')]:self.index_dict[('Fp1', 'F3')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('Fp2', 'F4')]:self.index_dict[('Fp2', 'F4')]+1])\n\n\tz3 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=0)\n\n  \n\n\t# Back part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('T5', 'O1')]:self.index_dict[('T5', 'O1')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('T6', 'O2')]:self.index_dict[('T6', 'O2')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('P3', 'O1')]:self.index_dict[('P3', 'O1')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('P4', 'O2')]:self.index_dict[('P4', 'O2')]+1])\n\n\tz4 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=0)\n\n  \n\ty = torch.cat([z1, z2, z3, z4], dim=1)\n```\n\n\n### Chrononet\n\n“ChronoNet is formed by stacking multiple 1D convolution layers followed by deep gated recurrent unit (GRU) layers where each 1D convolution layer uses multiple filters of exponentially varying lengths and the stacked GRU layers are densely connected in a feed-forward manner.” - [https://arxiv.org/abs/1802.00308](https://arxiv.org/abs/1802.00308).\n\n  \n#### ChronoNet Infer : \n[https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323](https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323)\n\nFeature code similar to wavenet-gru but the chrononet blocks look like below\n\n```\nclass  ChronoNet(nn.Module):\n\n\tdef  __init__(self, channel):\n\n\tsuper().__init__()\n\n\tself.inception_block1=InceptionBlock(channel) \t\t # 1st Inception Block\n\n\tself.inception_block2=InceptionBlock(96) # 2nd Inception Block\n\n\tself.inception_block3=InceptionBlock(96) # 3rd Inception Block\n\n\tself.gru1 = nn.GRU(input_size =  96, hidden_size =  32, batch_first =  True, bidirectional=True) # 1st GRU layer\n\n\tself.gru2 = nn.GRU(input_size =  32*2, \thidden_size =  32, batch_first =  True, bidirectional=True) # 2nd GRU layer\n\n\tself.gru3 = nn.GRU(input_size =  64*2, hidden_size =  32, batch_first =  True, bidirectional=True) # 3rd GRU layer\n\n\tself.gru4 = nn.GRU(input_size =  96*2, hidden_size =  32, batch_first =  True, bidirectional=True) # 4th GRU layer\n\n\tself.relu = nn.ReLU() # ReLU Activation Function\n\n\tself.gru_linear=nn.Linear(in_features =  1250, out_features =  1) # Linear Layer for the 4th GRU\n\n\tself.flatten = nn.Flatten() # Flattening Layer\n\n\tself.seqpool = SeqPool(64)\n\n\tself.fc1 = nn.Linear(32,6) # Fully Connected Layer / Output Layer.\n\n  \n\n\tdef forward(self,x): # Defining the feed forward function\n\n\tx = x.permute(0, 2, 1)\n\n\tx=self.inception_block1(x) # Fed to Inception Block 1\n\n\tx=self.inception_block2(x) # Fed to Inception Block 2\n\n\tx=self.inception_block3(x) # Fed to Inception Block 3\n\n\tx=x.permute(0,2,1) # Permuted for GRU layers\n\n\tgru_out1,_=self.gru1(x) # Fed into GRU layer 1\n\n\tgru_out2,_=self.gru2(gru_out1) # Fed into GRU layer 2\n\n\tgru_out=torch.cat((gru_out1, gru_out2), dim =  2) # Concatenated, defining the skip connection\n\n\tgru_out3,_=self.gru3(gru_out) # Fed into GRU layer 3\n\n\tgru_out = torch.cat((gru_out1, gru_out2, gru_out3), dim =  2) #C Concatenated, defining the next 2 skip connections\n\n\tgru_out4,_=self.gru4(gru_out) # Fed into the 4th GRU Layer\n\n\tseqpool =  self.seqpool(gru_out4)\n\n\treturn seqpool # Output\n```\n\n\n### Efficientnet\n\n**Public Model shared by @rafaelzimmermann1 **\n\nAfter this model was shared shortly before the deadline we decided to tweak it a little bit and just incorporate it as this was a great solution. Thanks @rafaelzimmermann1!\n\n  \n\n**Pytorch Zimmermann model**\n\nFor the CNN model we used `tf_efficientnet_b0.ns_jft_in1k`backbone from timm followed by GEM pooling. What helped improved the model performance is the usage of hflip+XYMasking and using a probability which dictates if mixup should be applied or not, and when it is applied no other augmentations are applied.\n\nUsing this trick + Zimmermann specs gave a boost around 0.1 (@alejopaullier trick).\n\n  \n\n### Ensemble\n\nWe tried different procedures for the ensembles and went with a mixture of OOF-predictions and Nelder-Mead optimization and manual weights for the weighting as some of our models did not have OOF-predictions (pipeline from before teammerge and time was running out).\n\nOur final ensemble was weighted as follows:\n\n  \n\n(Dual-input 1D-ResNet + Wavenet-GRU + Chrononet) / 3 * 0.39 +\n\nEfficientnet-B0_melspecs * 0.30 +\n\nEfficientnet-B0_zimmermann_specs * 0.31\n\n  \n\n## Little “tricks” and augmentations we used\n\nWe used a binary auxiliary target that discriminated those with more than 9 voters from those with less and it helped especially in stage 1.\n\nThe 1D-ResNet-based models were trained using heavy augmentation which can be found in the dataset class in the submission. The augmentations include:\n\n-   Horizontal flip\n    \n-   Slightly moving the single channels against each other randomly\n    \n-   Masking parts of the signal, especially outside the middle 10 s\n    \n-   Dropout of single channels\n    \n-   Random increase/decrease of wave amplitude\n    \n-   Gaussian noise\n    \n\n  \n\nCutmix and Mixup were also used heavily for their training which was done mostly using OneCycleLR-Scheduler in the end. Preprocessing for the 1D-Resnets was simple using the eight double banana features Chris proposed as all experiments we did with other features yielded worse results for the 1D-ResNet-based models.\n\n  \n\n### Lessons learned\n\nMaxing out the batch size and tuning the learning rate to match it will help you to be able to smoothly run more experiments. Using OneCycleLR scheduler as a starter is also not a bad idea. Looking at all the top solutions, using the raw input and turning it in whichever way to 2D input would have been amazing but maybe we did not have the trust that this would be a good idea as we definitely thought about similar things - so: If you think it, try it!",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2747089": "# First of all, thank you to Kaggle and the hosts for the great competition and everyone for all the great discussions and notebooks!\n\n## Team : Brain Freeze\n@alejopaullier \n@gauravbrills \n@medali1992 \n@pcjimmmy \n@janbrederecke \n\n## Summary\n\nWe achieved 70th place by a combination of various 1D models that used the raw EEGs and 2D-CNNs that used spectrograms as the input. The final ensemble achieved public LB: 0.289 / private LB: 0.345 and was our best submission after all.\n\n  \n\n## EDA\n\nEspecially [@pcjimmmy](https://www.kaggle.com/pcjimmmy) did lots of EDA in the beginning where he tried to extract useful features for tree-based machine learning algorithms. While doing so he found out about the two different distributions and kindly shared this here: [https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda#Evaluator-group-size-impact)\n\n  \n\n### Two parallel approaches\n\nAs we were unsure if the hidden test set would really be from the distribution of data that was labeled by >=10 physicians, we decided to create stage 1 and stage 2 submissions.\n\n  \n\n### Data and CV\n\nWhile individually starting out with the 17k rows approach, it was decided to use the 20k rows approach posted by @seanbearden. Even though we repeatedly discussed using more or all the data available we did make it work as most of us worked on Kaggle GPUs and the limitations made it too inefficient. @gauravbrills found out that binning the samples by number of voters using 5 bins resulted in a very even distribution which we used for stratified group k-fold cross-validation using 5 folds yielding a 0.03 LB boost in our experiments for all models.\n\n```\nbin_edges = [0, 5, 10, 15, 20, np.inf]  \nnum_bins = len(bin_edges) - 1  \ntrain['evaluator_bin'] = pd.cut(train['total_evaluators'], bins=bin_edges, labels=False)  \npatient_groups = train.groupby('patient_id').ngroup()  \nstratified_kfold = StratifiedGroupKFold(n_splits=CFG.n_fold)  \ntrain_indices = []  \nvalid_indices = []  \nfor fold, (train_idx, valid_idx) in enumerate(stratified_kfold.split(X=train, y=train['evaluator_bin'], groups=patient_groups)):  \ntrain_indices.append(train_idx)  \nvalid_indices.append(valid_idx)  \ntrain.loc[valid_idx, \"fold\"] = fold\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2F1211f98a73ee3130de65b2f236eb3205%2FPicture%201.png?generation=1712854341400842&alt=media)\n\n## Training strategy and the one big mistake\n\nWe trained some of the models as different people had different (local) resources and workflows (established individually before the team merge) but generally speaking we trained for most epochs on stage 1 data and then trained a few epochs on stage 2 data only. Likely our one big mistake was to validate the second stage on the same data as in stage 1 which might have led to suboptimal model selection at that stage.\n\n## Models used\n\nFor a long time, we tried to make the raw input 1D approach work using variations of 1D-ResNets, Wavenet ([https://arxiv.org/abs/1609.03499](https://arxiv.org/abs/1609.03499)), and in the end Chrononet ([https://arxiv.org/abs/1802.00308](https://arxiv.org/abs/1802.00308)) as well. We then added different flavors of 2D CNNs that used spectrograms as their input, dual input models were tried but we were not successful enough to keep them in our final ensemble (Now having read about more successful solutions we see what we did wrong…).\n\n  \n\n### 1D-ResNet (only used in stage1)\n\n**Input**: 50s raw data downsampled to 50 Hz, preprocessed with 60 Hz notch-filter and 0.5 / 50 Hz Low-High-Cut\n\nThe 1D-ResNet shared by @nischaydnk was very useful and we modified it slightly to have more planes and ResBlocks. This model was used in our stage 1 ensemble but not in stage 2 (it is complicated). We used downsampled data with 50 Hz. This model achieved LB .37 when trained for stage 1 on the 17k rows data selection. We never really transferred this to our final data / CV scheme as we found the next model to be a little bit better. We used Adan for LR scheduling in all 1D-ResNet variations.\n\n  \n\n### 1D-Dual-Input-ResNet with Attention (otherwise known as the Turbomodel)\n\n**Input**: 50s raw data downsampled to 25 / 50 Hz, 0.5 / 50 Hz Low-High-Cut\n\nWe used a dual input version of the given 1D-Resnet that was made of a large and a small instance of the ResNet-1D. The large one was fed the 50Hz input data and the small one was fed 25Hz data. Each output was given a separate scaled dot product attention mechanism before the individual results were concatenated. This model achieved LB .33 in stage 2 when retrained for 10 epochs on stage 2 data only.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9189772%2Fc68ab9af340d695ebbe2402c37cc30c6%2FPicture%202.png?generation=1712854373761825&alt=media)\n\n\n\n### Wavenet-GRU\n\nWe used the pytorch wavenet shared by @alejopaullier @cdeotte and added to it a bidirectional GRU layer followed by a SeqPool that behaves like a mean pooling layer used for transformer.\n\nWe used the Quad/Penta montage with an overall 15/17 features and each chain is evaluated separately by the model.\n\nNotes on structure of wavenet gru with legs corresponding to parts of the brain LP, RP,FP,BP,CP\n\n[https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492220). The CP did not work for us in stage 2 though the Penta model was trained for the same.\n\n**Preprocessing**\n\n-   Apply notch filter (60Hz) followed by forward and backward digital filter\n    \n-   Apply bandpass filter (0.7Hz-20Hz) followed by forward and backward digital filter\n    \n\n**Training**\n\n-   Train PentaTail :  [https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail](https://www.kaggle.com/code/gauravbrills/hms-wave-gru-train-penta-tail/edit)\n    \n-   Train Quad : [https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267](https://www.kaggle.com/code/gauravbrills/hms-wavenet-gru-train-v3-diff-feats?scriptVersionId=168621267)\n    \n-   Private STAGE 2 Quad and Penta: 0.40\n    \n-   Infer Stage 2 :[https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167](https://www.kaggle.com/code/gauravbrills/eeg-wave-gru-strat-folds?scriptVersionId=168858167)\n\n```\nclass  SeqPool(nn.Module):  \n\tdef  __init__(self, emb_dim=192):  \n\t\tsuper().__init__()  \n\t\tself.dense = nn.Linear(emb_dim, 1)  \n\t\tself.softmax = nn.Softmax(dim=-1)  \n  \n\tdef  forward(self, x):  \n\t\tbs, seq_len, emb_dim = x.shape  \n\t\tidentity = x  \n\t\tx = self.dense(x)  \n\t\tx = x.permute(0, 2, 1)  \n\t\tx = self.softmax(x)  \n\t\tx = x @ identity  \n\t\tx = x.reshape(x.shape[0], -1)  \n\t\treturn x\n```\n**Feature code**\n```\ndef extract_features(self, x):\n\n\t# Left part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('Fp1', 'F7')]:self.index_dict[('Fp1', 'F7')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('F7', 'T3')]:self.index_dict[('F7', 'T3')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('T3', 'T5')]:self.index_dict[('T3', 'T5')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('T5', 'O1')]:self.index_dict[('T5', 'O1')]+1])\n\n\tx6 =  self.model(x[:, :, self.index_dict[('Fp1', 'F3')]:self.index_dict[('Fp1', 'F3')]+1])\n\n\tx7 =  self.model(x[:, :, self.index_dict[('F3', 'C3')]:self.index_dict[('F3', 'C3')]+1])\n\n\t#x8 = self.model(x[:, :, self.index_dict[('C3', 'P3')]:self.index_dict[('C3', 'P3')]+1])\n\n\tx9 =  self.model(x[:, :, self.index_dict[('P3', 'O1')]:self.index_dict[('P3', 'O1')]+1])\n\n\tz1 = torch.mean(torch.stack([x1, x2, x3, x4, x6, x7, x9]), dim=0)\n\n  \n\n\t# Right part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('Fp2', 'F8')]:self.index_dict[('Fp2', 'F8')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('F8', 'T4')]:self.index_dict[('F8', 'T4')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('T4', 'T6')]:self.index_dict[('T4', 'T6')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('T6', 'O2')]:self.index_dict[('T6', 'O2')]+1])\n\n\tx5 =  self.model(x[:, :, self.index_dict[('C4', 'T4')]:self.index_dict[('C4', 'T4')]+1])\n\n\tx6 =  self.model(x[:, :, self.index_dict[('Fp2', 'F4')]:self.index_dict[('Fp2', 'F4')]+1])\n\n\tx7 =  self.model(x[:, :, self.index_dict[('F4', 'C4')]:self.index_dict[('F4', 'C4')]+1])\n\n\tx9 =  self.model(x[:, :, self.index_dict[('P4', 'O2')]:self.index_dict[('P4', 'O2')]+1])\n\n\tz2 = torch.mean(torch.stack([x1, x2, x3, x4, x5, x6, x7, x9]), dim=0)\n\n  \n\n\t# Front part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('Fp1', 'F7')]:self.index_dict[('Fp1', 'F7')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('Fp2', 'F8')]:self.index_dict[('Fp2', 'F8')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('Fp1', 'F3')]:self.index_dict[('Fp1', 'F3')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('Fp2', 'F4')]:self.index_dict[('Fp2', 'F4')]+1])\n\n\tz3 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=0)\n\n  \n\n\t# Back part\n\n\tx1 =  self.model(x[:, :, self.index_dict[('T5', 'O1')]:self.index_dict[('T5', 'O1')]+1])\n\n\tx2 =  self.model(x[:, :, self.index_dict[('T6', 'O2')]:self.index_dict[('T6', 'O2')]+1])\n\n\tx3 =  self.model(x[:, :, self.index_dict[('P3', 'O1')]:self.index_dict[('P3', 'O1')]+1])\n\n\tx4 =  self.model(x[:, :, self.index_dict[('P4', 'O2')]:self.index_dict[('P4', 'O2')]+1])\n\n\tz4 = torch.mean(torch.stack([x1, x2, x3, x4]), dim=0)\n\n  \n\ty = torch.cat([z1, z2, z3, z4], dim=1)\n```\n\n\n### Chrononet\n\n“ChronoNet is formed by stacking multiple 1D convolution layers followed by deep gated recurrent unit (GRU) layers where each 1D convolution layer uses multiple filters of exponentially varying lengths and the stacked GRU layers are densely connected in a feed-forward manner.” - [https://arxiv.org/abs/1802.00308](https://arxiv.org/abs/1802.00308).\n\n  \n#### ChronoNet Infer : \n[https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323](https://www.kaggle.com/code/medali1992/eeg-chrononet-strat-quad-inference?scriptVersionId=170923323)\n\nFeature code similar to wavenet-gru but the chrononet blocks look like below\n\n```\nclass  ChronoNet(nn.Module):\n\n\tdef  __init__(self, channel):\n\n\tsuper().__init__()\n\n\tself.inception_block1=InceptionBlock(channel) \t\t # 1st Inception Block\n\n\tself.inception_block2=InceptionBlock(96) # 2nd Inception Block\n\n\tself.inception_block3=InceptionBlock(96) # 3rd Inception Block\n\n\tself.gru1 = nn.GRU(input_size =  96, hidden_size =  32, batch_first =  True, bidirectional=True) # 1st GRU layer\n\n\tself.gru2 = nn.GRU(input_size =  32*2, \thidden_size =  32, batch_first =  True, bidirectional=True) # 2nd GRU layer\n\n\tself.gru3 = nn.GRU(input_size =  64*2, hidden_size =  32, batch_first =  True, bidirectional=True) # 3rd GRU layer\n\n\tself.gru4 = nn.GRU(input_size =  96*2, hidden_size =  32, batch_first =  True, bidirectional=True) # 4th GRU layer\n\n\tself.relu = nn.ReLU() # ReLU Activation Function\n\n\tself.gru_linear=nn.Linear(in_features =  1250, out_features =  1) # Linear Layer for the 4th GRU\n\n\tself.flatten = nn.Flatten() # Flattening Layer\n\n\tself.seqpool = SeqPool(64)\n\n\tself.fc1 = nn.Linear(32,6) # Fully Connected Layer / Output Layer.\n\n  \n\n\tdef forward(self,x): # Defining the feed forward function\n\n\tx = x.permute(0, 2, 1)\n\n\tx=self.inception_block1(x) # Fed to Inception Block 1\n\n\tx=self.inception_block2(x) # Fed to Inception Block 2\n\n\tx=self.inception_block3(x) # Fed to Inception Block 3\n\n\tx=x.permute(0,2,1) # Permuted for GRU layers\n\n\tgru_out1,_=self.gru1(x) # Fed into GRU layer 1\n\n\tgru_out2,_=self.gru2(gru_out1) # Fed into GRU layer 2\n\n\tgru_out=torch.cat((gru_out1, gru_out2), dim =  2) # Concatenated, defining the skip connection\n\n\tgru_out3,_=self.gru3(gru_out) # Fed into GRU layer 3\n\n\tgru_out = torch.cat((gru_out1, gru_out2, gru_out3), dim =  2) #C Concatenated, defining the next 2 skip connections\n\n\tgru_out4,_=self.gru4(gru_out) # Fed into the 4th GRU Layer\n\n\tseqpool =  self.seqpool(gru_out4)\n\n\treturn seqpool # Output\n```\n\n\n### Efficientnet\n\n**Public Model shared by @rafaelzimmermann1 **\n\nAfter this model was shared shortly before the deadline we decided to tweak it a little bit and just incorporate it as this was a great solution. Thanks @rafaelzimmermann1!\n\n  \n\n**Pytorch Zimmermann model**\n\nFor the CNN model we used `tf_efficientnet_b0.ns_jft_in1k`backbone from timm followed by GEM pooling. What helped improved the model performance is the usage of hflip+XYMasking and using a probability which dictates if mixup should be applied or not, and when it is applied no other augmentations are applied.\n\nUsing this trick + Zimmermann specs gave a boost around 0.1 (@alejopaullier trick).\n\n  \n\n### Ensemble\n\nWe tried different procedures for the ensembles and went with a mixture of OOF-predictions and Nelder-Mead optimization and manual weights for the weighting as some of our models did not have OOF-predictions (pipeline from before teammerge and time was running out).\n\nOur final ensemble was weighted as follows:\n\n  \n\n(Dual-input 1D-ResNet + Wavenet-GRU + Chrononet) / 3 * 0.39 +\n\nEfficientnet-B0_melspecs * 0.30 +\n\nEfficientnet-B0_zimmermann_specs * 0.31\n\n  \n\n## Little “tricks” and augmentations we used\n\nWe used a binary auxiliary target that discriminated those with more than 9 voters from those with less and it helped especially in stage 1.\n\nThe 1D-ResNet-based models were trained using heavy augmentation which can be found in the dataset class in the submission. The augmentations include:\n\n-   Horizontal flip\n    \n-   Slightly moving the single channels against each other randomly\n    \n-   Masking parts of the signal, especially outside the middle 10 s\n    \n-   Dropout of single channels\n    \n-   Random increase/decrease of wave amplitude\n    \n-   Gaussian noise\n    \n\n  \n\nCutmix and Mixup were also used heavily for their training which was done mostly using OneCycleLR-Scheduler in the end. Preprocessing for the 1D-Resnets was simple using the eight double banana features Chris proposed as all experiments we did with other features yielded worse results for the 1D-ResNet-based models.\n\n  \n\n### Lessons learned\n\nMaxing out the batch size and tuning the learning rate to match it will help you to be able to smoothly run more experiments. Using OneCycleLR scheduler as a starter is also not a bad idea. Looking at all the top solutions, using the raw input and turning it in whichever way to 2D input would have been amazing but maybe we did not have the trust that this would be a good idea as we definitely thought about similar things - so: If you think it, try it!"
  }
}