{
  "id": 492603,
  "title": "Team BNR: 12th Place Solution, Exploring Various Custom Architectures ",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/bnr-team-bnr-12th-place-solution-exploring-various",
  "author_name": "",
  "post_date": "2024-04-10T08:15:53.093Z",
  "votes": 43,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Firstly, Thank you Kaggle &amp; HBS team for hosting such an interesting competition &amp; my teammates <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> <a href=\"https://www.kaggle.com/romainhardy\" target=\"_blank\">@romainhardy</a> , It was the collective efforts of all of us which led to gold position. Congrats <a href=\"https://www.kaggle.com/romainhardy\" target=\"_blank\">@romainhardy</a> on becoming Master after multiple close to solo gold performances.   </p>\n<p>Our solution mainly comprises <strong>five different modalities</strong>, I will try to cover all of them in detail. Each model was trained using two stage custom pipeline (w/ pseudo+noisy labels). It’s always fun to work with signals based data since you can try out variety of architectures 1D,2D and even 3D networks :))</p>\n<h2>1. Preprocessing/ data split:</h2>\n<p>Our split was common throughout the competition, we used 5 fold GroupStratifiedKFold based on patient_id and labels. </p>\n<p><strong>Our preprocessing pipeline is divided into several parts:</strong></p>\n<ul>\n<li>1d models: we used the calculated differences of the Banana montage -&gt; other montages like the Transverse or Average montage did not work for us. Also, we could not use the mean EEG signal series and the ECG data in a meaningful way</li>\n<li>2d models: We had two different approaches here. For Romain's spectrogram models, we chose the 4 public montage. Also, outputting a Spectrogram for each banana pair gave us a good boost (0.03 on lb/cv).</li>\n<li>Bandpass filter with a lower bound between 0.5 and 1 and an upper bound between 20 and 35.</li>\n<li>We already identified the training split in the first month after the start of the competition, i.e. we trained and validated the whole time on stage 2 (&gt;=10 votes). Later, we used the noisy labels in the pre-training phase. </li>\n</ul>\n<h2>2. Augmentations / TTA</h2>\n<p>Augmentations gave us almost <strong>0.1 boost</strong> on both CV/LB. Some of common augmentations which we used across all models were:<br>\nMixUp<br>\nCutmix<br>\nRandom Crop (20/30/40 secs)<br>\nRandom Flip (channel level)<br>\nRandom Sign Flip (channel level)<br>\nRandom Masking (channel level)</p>\n<p><strong>We got additional 0.01 boost using offset level TTA. Instead of generating predictions on middle N samples, we changed the starting point of sequence with (+-) offsets, ranging from -2400 to + 2400</strong></p>\n<h2>3. Architectures</h2>\n<h3>1D Learnable Frontend w/ EffcientNet:</h3>\n<ul>\n<li>Using 1D convolutions as Feature extractor</li>\n<li>Starting layers similar to the EEGNet architecture (link) I released </li>\n<li>Use efficientnet as backbone</li>\n<li>Best single model score: 0.246 CV | LB: 0.25</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F90b83b641cb87f089958a1d9abc70dd3%2FScreenshot%202024-04-10%20at%201.23.15%20PM.png?generation=1712735828273679&amp;alt=media\"></p>\n<h3>2D Learnable Frontend w/ EffcientNet:</h3>\n<ul>\n<li>Using 2D convolutions as Feature extractor, turns out to be most powerful model for us.</li>\n<li>Use multiple parallel conv2D blocks and create 3 channels input</li>\n<li>Use efficientnet B0 / v2s as backbone</li>\n<li>Best single model score: <strong>0.235</strong> CV | LB: <strong>0.24x</strong></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F2af8b3a4b4c12f3d08c0f1b8cae2ab8d%2FScreenshot%202024-04-10%20at%201.27.40%20PM.png?generation=1712735875191718&amp;alt=media\"></p>\n<h3>1D Resnet based EEGNet with Transformer blocks:</h3>\n<ul>\n<li>Similar to the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\" target=\"_blank\">EegNet baseline</a> I shared early in competition</li>\n<li>After 4 resnet blocks, once the time dimension is shorter, we applied transformer based residual networks on the output.</li>\n<li>Use Linear Attention Transformer with GRU </li>\n<li>Best single model score: 0.269 CV | LB: 0.26</li>\n</ul>\n<h3>CWT w/ EffcientNet:</h3>\n<ul>\n<li>Using trainable Continuous Wavelet Transform as feature extractor</li>\n<li>Use efficientnet based models as backbone</li>\n<li>Best single model score: 0.26 CV | LB: 0.26</li>\n</ul>\n<h3>Spectrogram based models</h3>\n<ul>\n<li>As already described, we used two different versions of Spectrograms. The most important backbones were different versions of EfficientNet as well as SwinTransformer.</li>\n<li>The public lb score for the Spectrograms was around 0.29/0.30. However, we reached this plateau very early on in the competition and were no longer able to improve performance significantly. Perhaps our bottleneck was that we did not increase the image size and had too little information in the image. Adding Kaggle Spectrograms, other resolutions or other time slices did not lead to an improvement in model performance.</li>\n<li>We tried to use the raw EEG signals as 3 channel RGB images, but again the performance was not promising. We also tried splitting the time dimension, but this did not improve the model either.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F8043d2b02aded0d1653a32fe0f2dc1f9%2FScreenshot%202024-04-10%20at%201.29.22%20PM.png?generation=1712735978734534&amp;alt=media\"></p>\n<h2>4. Ensembling &amp; Pseudo Labelling</h2>\n<p>Our final ensemble included these models, and scored around <strong>0.215 CV and 0.23 LB.</strong> Our winning models consists of full data trained models on 3 different seeds and ensemble weights were kept same as 5 fold scores.</p>\n<ul>\n<li>We tried different stacking techniques and also tried to use ANNs. However, the performance was not good. In the end we used <strong>hill climbing</strong>:</li>\n<li>We wanted to avoid overfitting too much on the training data</li>\n<li>We could quickly identify diversity in our models without investing a lot of time</li>\n<li>Every improvement in the hill climbing cv led to an improvement in the leaderboard</li>\n<li>We were able to derive the weights for the models trained on the whole data from the 5-fold scores</li>\n<li>Due to the specific metric, the ensemble methods were limited. Hill climbing was a simple and efficient method to continuously optimize our ensemble.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fed70cf14d416e800b59baaf73200f89d%2FScreenshot%202024-04-10%20at%201.36.31%20PM.png?generation=1712736406689577&amp;alt=media\"></p>\n<p>For pseudo labels, steps were as follows.</p>\n<ul>\n<li>Taking \"current\"-best-ensemble, pseudo label noisy data</li>\n<li>Training in two stages: apply pseudo labels by a chance of 50% instead of noisy labels</li>\n</ul>\n<p>Also shoutout to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for his immense sharing in the competition, it also encouraged me to publish multiple baselines in the competition. And my teammates, despite the time zone differences we used to connect over calls and made great progress together.</p>",
  "messages": [
    {
      "id": "2744928",
      "postDate": "04/10/2024 08:07:06",
      "content": "<p>Firstly, Thank you Kaggle &amp; HBS team for hosting such an interesting competition &amp; my teammates <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> <a href=\"https://www.kaggle.com/romainhardy\" target=\"_blank\">@romainhardy</a> , It was the collective efforts of all of us which led to gold position. Congrats <a href=\"https://www.kaggle.com/romainhardy\" target=\"_blank\">@romainhardy</a> on becoming Master after multiple close to solo gold performances.   </p>\n<p>Our solution mainly comprises <strong>five different modalities</strong>, I will try to cover all of them in detail. Each model was trained using two stage custom pipeline (w/ pseudo+noisy labels). It’s always fun to work with signals based data since you can try out variety of architectures 1D,2D and even 3D networks :))</p>\n<h2>1. Preprocessing/ data split:</h2>\n<p>Our split was common throughout the competition, we used 5 fold GroupStratifiedKFold based on patient_id and labels. </p>\n<p><strong>Our preprocessing pipeline is divided into several parts:</strong></p>\n<ul>\n<li>1d models: we used the calculated differences of the Banana montage -&gt; other montages like the Transverse or Average montage did not work for us. Also, we could not use the mean EEG signal series and the ECG data in a meaningful way</li>\n<li>2d models: We had two different approaches here. For Romain's spectrogram models, we chose the 4 public montage. Also, outputting a Spectrogram for each banana pair gave us a good boost (0.03 on lb/cv).</li>\n<li>Bandpass filter with a lower bound between 0.5 and 1 and an upper bound between 20 and 35.</li>\n<li>We already identified the training split in the first month after the start of the competition, i.e. we trained and validated the whole time on stage 2 (&gt;=10 votes). Later, we used the noisy labels in the pre-training phase. </li>\n</ul>\n<h2>2. Augmentations / TTA</h2>\n<p>Augmentations gave us almost <strong>0.1 boost</strong> on both CV/LB. Some of common augmentations which we used across all models were:<br>\nMixUp<br>\nCutmix<br>\nRandom Crop (20/30/40 secs)<br>\nRandom Flip (channel level)<br>\nRandom Sign Flip (channel level)<br>\nRandom Masking (channel level)</p>\n<p><strong>We got additional 0.01 boost using offset level TTA. Instead of generating predictions on middle N samples, we changed the starting point of sequence with (+-) offsets, ranging from -2400 to + 2400</strong></p>\n<h2>3. Architectures</h2>\n<h3>1D Learnable Frontend w/ EffcientNet:</h3>\n<ul>\n<li>Using 1D convolutions as Feature extractor</li>\n<li>Starting layers similar to the EEGNet architecture (link) I released </li>\n<li>Use efficientnet as backbone</li>\n<li>Best single model score: 0.246 CV | LB: 0.25</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F90b83b641cb87f089958a1d9abc70dd3%2FScreenshot%202024-04-10%20at%201.23.15%20PM.png?generation=1712735828273679&amp;alt=media\"></p>\n<h3>2D Learnable Frontend w/ EffcientNet:</h3>\n<ul>\n<li>Using 2D convolutions as Feature extractor, turns out to be most powerful model for us.</li>\n<li>Use multiple parallel conv2D blocks and create 3 channels input</li>\n<li>Use efficientnet B0 / v2s as backbone</li>\n<li>Best single model score: <strong>0.235</strong> CV | LB: <strong>0.24x</strong></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F2af8b3a4b4c12f3d08c0f1b8cae2ab8d%2FScreenshot%202024-04-10%20at%201.27.40%20PM.png?generation=1712735875191718&amp;alt=media\"></p>\n<h3>1D Resnet based EEGNet with Transformer blocks:</h3>\n<ul>\n<li>Similar to the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\" target=\"_blank\">EegNet baseline</a> I shared early in competition</li>\n<li>After 4 resnet blocks, once the time dimension is shorter, we applied transformer based residual networks on the output.</li>\n<li>Use Linear Attention Transformer with GRU </li>\n<li>Best single model score: 0.269 CV | LB: 0.26</li>\n</ul>\n<h3>CWT w/ EffcientNet:</h3>\n<ul>\n<li>Using trainable Continuous Wavelet Transform as feature extractor</li>\n<li>Use efficientnet based models as backbone</li>\n<li>Best single model score: 0.26 CV | LB: 0.26</li>\n</ul>\n<h3>Spectrogram based models</h3>\n<ul>\n<li>As already described, we used two different versions of Spectrograms. The most important backbones were different versions of EfficientNet as well as SwinTransformer.</li>\n<li>The public lb score for the Spectrograms was around 0.29/0.30. However, we reached this plateau very early on in the competition and were no longer able to improve performance significantly. Perhaps our bottleneck was that we did not increase the image size and had too little information in the image. Adding Kaggle Spectrograms, other resolutions or other time slices did not lead to an improvement in model performance.</li>\n<li>We tried to use the raw EEG signals as 3 channel RGB images, but again the performance was not promising. We also tried splitting the time dimension, but this did not improve the model either.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F8043d2b02aded0d1653a32fe0f2dc1f9%2FScreenshot%202024-04-10%20at%201.29.22%20PM.png?generation=1712735978734534&amp;alt=media\"></p>\n<h2>4. Ensembling &amp; Pseudo Labelling</h2>\n<p>Our final ensemble included these models, and scored around <strong>0.215 CV and 0.23 LB.</strong> Our winning models consists of full data trained models on 3 different seeds and ensemble weights were kept same as 5 fold scores.</p>\n<ul>\n<li>We tried different stacking techniques and also tried to use ANNs. However, the performance was not good. In the end we used <strong>hill climbing</strong>:</li>\n<li>We wanted to avoid overfitting too much on the training data</li>\n<li>We could quickly identify diversity in our models without investing a lot of time</li>\n<li>Every improvement in the hill climbing cv led to an improvement in the leaderboard</li>\n<li>We were able to derive the weights for the models trained on the whole data from the 5-fold scores</li>\n<li>Due to the specific metric, the ensemble methods were limited. Hill climbing was a simple and efficient method to continuously optimize our ensemble.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fed70cf14d416e800b59baaf73200f89d%2FScreenshot%202024-04-10%20at%201.36.31%20PM.png?generation=1712736406689577&amp;alt=media\"></p>\n<p>For pseudo labels, steps were as follows.</p>\n<ul>\n<li>Taking \"current\"-best-ensemble, pseudo label noisy data</li>\n<li>Training in two stages: apply pseudo labels by a chance of 50% instead of noisy labels</li>\n</ul>\n<p>Also shoutout to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for his immense sharing in the competition, it also encouraged me to publish multiple baselines in the competition. And my teammates, despite the time zone differences we used to connect over calls and made great progress together.</p>",
      "rawMarkdown": "Firstly, Thank you Kaggle & HBS team for hosting such an interesting competition & my teammates @benbla @romainhardy , It was the collective efforts of all of us which led to gold position. Congrats @romainhardy on becoming Master after multiple close to solo gold performances.   \n\nOur solution mainly comprises **five different modalities**, I will try to cover all of them in detail. Each model was trained using two stage custom pipeline (w/ pseudo+noisy labels). It’s always fun to work with signals based data since you can try out variety of architectures 1D,2D and even 3D networks :))\n\n\n## 1. Preprocessing/ data split:\n\nOur split was common throughout the competition, we used 5 fold GroupStratifiedKFold based on patient_id and labels. \n\n**Our preprocessing pipeline is divided into several parts:**\n- 1d models: we used the calculated differences of the Banana montage -> other montages like the Transverse or Average montage did not work for us. Also, we could not use the mean EEG signal series and the ECG data in a meaningful way\n- 2d models: We had two different approaches here. For Romain's spectrogram models, we chose the 4 public montage. Also, outputting a Spectrogram for each banana pair gave us a good boost (0.03 on lb/cv).\n- Bandpass filter with a lower bound between 0.5 and 1 and an upper bound between 20 and 35.\n- We already identified the training split in the first month after the start of the competition, i.e. we trained and validated the whole time on stage 2 (>=10 votes). Later, we used the noisy labels in the pre-training phase. \n\n\n## 2. Augmentations / TTA\nAugmentations gave us almost **0.1 boost** on both CV/LB. Some of common augmentations which we used across all models were:\nMixUp\nCutmix\nRandom Crop (20/30/40 secs)\nRandom Flip (channel level)\nRandom Sign Flip (channel level)\nRandom Masking (channel level)\n\n**We got additional 0.01 boost using offset level TTA. Instead of generating predictions on middle N samples, we changed the starting point of sequence with (+-) offsets, ranging from -2400 to + 2400**\n\n## 3. Architectures\n\n### 1D Learnable Frontend w/ EffcientNet:\n\n- Using 1D convolutions as Feature extractor\n- Starting layers similar to the EEGNet architecture (link) I released \n- Use efficientnet as backbone\n- Best single model score: 0.246 CV | LB: 0.25\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F90b83b641cb87f089958a1d9abc70dd3%2FScreenshot%202024-04-10%20at%201.23.15%20PM.png?generation=1712735828273679&alt=media)\n\n### 2D Learnable Frontend w/ EffcientNet:\n\n- Using 2D convolutions as Feature extractor, turns out to be most powerful model for us.\n- Use multiple parallel conv2D blocks and create 3 channels input\n- Use efficientnet B0 / v2s as backbone\n- Best single model score: **0.235** CV | LB: **0.24x**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F2af8b3a4b4c12f3d08c0f1b8cae2ab8d%2FScreenshot%202024-04-10%20at%201.27.40%20PM.png?generation=1712735875191718&alt=media)\n\n### 1D Resnet based EEGNet with Transformer blocks:\n\n- Similar to the [EegNet baseline](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666) I shared early in competition\n- After 4 resnet blocks, once the time dimension is shorter, we applied transformer based residual networks on the output.\n- Use Linear Attention Transformer with GRU \n- Best single model score: 0.269 CV | LB: 0.26\n\n\n### CWT w/ EffcientNet:\n\n- Using trainable Continuous Wavelet Transform as feature extractor\n- Use efficientnet based models as backbone\n- Best single model score: 0.26 CV | LB: 0.26\n\n### Spectrogram based models\n\n- As already described, we used two different versions of Spectrograms. The most important backbones were different versions of EfficientNet as well as SwinTransformer.\n- The public lb score for the Spectrograms was around 0.29/0.30. However, we reached this plateau very early on in the competition and were no longer able to improve performance significantly. Perhaps our bottleneck was that we did not increase the image size and had too little information in the image. Adding Kaggle Spectrograms, other resolutions or other time slices did not lead to an improvement in model performance.\n- We tried to use the raw EEG signals as 3 channel RGB images, but again the performance was not promising. We also tried splitting the time dimension, but this did not improve the model either.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F8043d2b02aded0d1653a32fe0f2dc1f9%2FScreenshot%202024-04-10%20at%201.29.22%20PM.png?generation=1712735978734534&alt=media)\n\n\n## 4. Ensembling & Pseudo Labelling\n\nOur final ensemble included these models, and scored around **0.215 CV and 0.23 LB.** Our winning models consists of full data trained models on 3 different seeds and ensemble weights were kept same as 5 fold scores.\n\n- We tried different stacking techniques and also tried to use ANNs. However, the performance was not good. In the end we used **hill climbing**:\n- We wanted to avoid overfitting too much on the training data\n- We could quickly identify diversity in our models without investing a lot of time\n- Every improvement in the hill climbing cv led to an improvement in the leaderboard\n- We were able to derive the weights for the models trained on the whole data from the 5-fold scores\n- Due to the specific metric, the ensemble methods were limited. Hill climbing was a simple and efficient method to continuously optimize our ensemble.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fed70cf14d416e800b59baaf73200f89d%2FScreenshot%202024-04-10%20at%201.36.31%20PM.png?generation=1712736406689577&alt=media)\n\nFor pseudo labels, steps were as follows.\n- Taking \"current\"-best-ensemble, pseudo label noisy data\n- Training in two stages: apply pseudo labels by a chance of 50% instead of noisy labels\n\nAlso shoutout to @cdeotte for his immense sharing in the competition, it also encouraged me to publish multiple baselines in the competition. And my teammates, despite the time zone differences we used to connect over calls and made great progress together.",
      "votes": null
    },
    {
      "id": "2748271",
      "postDate": "04/12/2024 10:51:12",
      "content": "<p>Congratulations on achieving 12th place in this competition. Thanks for sharing details of your solution with colorful diagrams. </p>",
      "rawMarkdown": "Congratulations on achieving 12th place in this competition. Thanks for sharing details of your solution with colorful diagrams.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2748271,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "04/12/2024 10:51:12",
      "content": "<p>Congratulations on achieving 12th place in this competition. Thanks for sharing details of your solution with colorful diagrams. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2744928": "Firstly, Thank you Kaggle & HBS team for hosting such an interesting competition & my teammates @benbla @romainhardy , It was the collective efforts of all of us which led to gold position. Congrats @romainhardy on becoming Master after multiple close to solo gold performances.   \n\nOur solution mainly comprises **five different modalities**, I will try to cover all of them in detail. Each model was trained using two stage custom pipeline (w/ pseudo+noisy labels). It’s always fun to work with signals based data since you can try out variety of architectures 1D,2D and even 3D networks :))\n\n\n## 1. Preprocessing/ data split:\n\nOur split was common throughout the competition, we used 5 fold GroupStratifiedKFold based on patient_id and labels. \n\n**Our preprocessing pipeline is divided into several parts:**\n- 1d models: we used the calculated differences of the Banana montage -> other montages like the Transverse or Average montage did not work for us. Also, we could not use the mean EEG signal series and the ECG data in a meaningful way\n- 2d models: We had two different approaches here. For Romain's spectrogram models, we chose the 4 public montage. Also, outputting a Spectrogram for each banana pair gave us a good boost (0.03 on lb/cv).\n- Bandpass filter with a lower bound between 0.5 and 1 and an upper bound between 20 and 35.\n- We already identified the training split in the first month after the start of the competition, i.e. we trained and validated the whole time on stage 2 (>=10 votes). Later, we used the noisy labels in the pre-training phase. \n\n\n## 2. Augmentations / TTA\nAugmentations gave us almost **0.1 boost** on both CV/LB. Some of common augmentations which we used across all models were:\nMixUp\nCutmix\nRandom Crop (20/30/40 secs)\nRandom Flip (channel level)\nRandom Sign Flip (channel level)\nRandom Masking (channel level)\n\n**We got additional 0.01 boost using offset level TTA. Instead of generating predictions on middle N samples, we changed the starting point of sequence with (+-) offsets, ranging from -2400 to + 2400**\n\n## 3. Architectures\n\n### 1D Learnable Frontend w/ EffcientNet:\n\n- Using 1D convolutions as Feature extractor\n- Starting layers similar to the EEGNet architecture (link) I released \n- Use efficientnet as backbone\n- Best single model score: 0.246 CV | LB: 0.25\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F90b83b641cb87f089958a1d9abc70dd3%2FScreenshot%202024-04-10%20at%201.23.15%20PM.png?generation=1712735828273679&alt=media)\n\n### 2D Learnable Frontend w/ EffcientNet:\n\n- Using 2D convolutions as Feature extractor, turns out to be most powerful model for us.\n- Use multiple parallel conv2D blocks and create 3 channels input\n- Use efficientnet B0 / v2s as backbone\n- Best single model score: **0.235** CV | LB: **0.24x**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F2af8b3a4b4c12f3d08c0f1b8cae2ab8d%2FScreenshot%202024-04-10%20at%201.27.40%20PM.png?generation=1712735875191718&alt=media)\n\n### 1D Resnet based EEGNet with Transformer blocks:\n\n- Similar to the [EegNet baseline](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666) I shared early in competition\n- After 4 resnet blocks, once the time dimension is shorter, we applied transformer based residual networks on the output.\n- Use Linear Attention Transformer with GRU \n- Best single model score: 0.269 CV | LB: 0.26\n\n\n### CWT w/ EffcientNet:\n\n- Using trainable Continuous Wavelet Transform as feature extractor\n- Use efficientnet based models as backbone\n- Best single model score: 0.26 CV | LB: 0.26\n\n### Spectrogram based models\n\n- As already described, we used two different versions of Spectrograms. The most important backbones were different versions of EfficientNet as well as SwinTransformer.\n- The public lb score for the Spectrograms was around 0.29/0.30. However, we reached this plateau very early on in the competition and were no longer able to improve performance significantly. Perhaps our bottleneck was that we did not increase the image size and had too little information in the image. Adding Kaggle Spectrograms, other resolutions or other time slices did not lead to an improvement in model performance.\n- We tried to use the raw EEG signals as 3 channel RGB images, but again the performance was not promising. We also tried splitting the time dimension, but this did not improve the model either.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F8043d2b02aded0d1653a32fe0f2dc1f9%2FScreenshot%202024-04-10%20at%201.29.22%20PM.png?generation=1712735978734534&alt=media)\n\n\n## 4. Ensembling & Pseudo Labelling\n\nOur final ensemble included these models, and scored around **0.215 CV and 0.23 LB.** Our winning models consists of full data trained models on 3 different seeds and ensemble weights were kept same as 5 fold scores.\n\n- We tried different stacking techniques and also tried to use ANNs. However, the performance was not good. In the end we used **hill climbing**:\n- We wanted to avoid overfitting too much on the training data\n- We could quickly identify diversity in our models without investing a lot of time\n- Every improvement in the hill climbing cv led to an improvement in the leaderboard\n- We were able to derive the weights for the models trained on the whole data from the 5-fold scores\n- Due to the specific metric, the ensemble methods were limited. Hill climbing was a simple and efficient method to continuously optimize our ensemble.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fed70cf14d416e800b59baaf73200f89d%2FScreenshot%202024-04-10%20at%201.36.31%20PM.png?generation=1712736406689577&alt=media)\n\nFor pseudo labels, steps were as follows.\n- Taking \"current\"-best-ensemble, pseudo label noisy data\n- Training in two stages: apply pseudo labels by a chance of 50% instead of noisy labels\n\nAlso shoutout to @cdeotte for his immense sharing in the competition, it also encouraged me to publish multiple baselines in the competition. And my teammates, despite the time zone differences we used to connect over calls and made great progress together.",
    "2748271": "Congratulations on achieving 12th place in this competition. Thanks for sharing details of your solution with colorful diagrams."
  },
  "source": "meta"
}