{
  "id": 492301,
  "title": "11th Place Solution",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/kansai-kaggler-11th-place-solution",
  "author_name": "",
  "post_date": "2024-04-09T07:20:42.934268900Z",
  "votes": 47,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>11th place solution for the Harmful Brain Activity Classification Competition</h1>\n<p>Team : kansai-kaggler ( <a href=\"https://www.kaggle.com/tomohiroh\" target=\"_blank\">@tomohiroh</a> <a href=\"https://www.kaggle.com/ryushisa\" target=\"_blank\">@ryushisa</a> <a href=\"https://www.kaggle.com/shunyaiida\" target=\"_blank\">@shunyaiida</a> <a href=\"https://www.kaggle.com/ekichi\" target=\"_blank\">@ekichi</a> <a href=\"https://www.kaggle.com/t88take\" target=\"_blank\">@t88take</a>)  </p>\n<p>First of all, I would like to express my sincere gratitude to the Kaggle staff and all the hosts of the competition for organizing such an amazing competition. It was a very enjoyable competition where we could consider various approaches to EEG data. Furthermore, We are very happy that three people from our team have become Kaggle Competitions Master.</p>\n<p>We would like to share our solution.   </p>\n<h1>Overview</h1>\n<p>Our team merged with the P-SHA and TR teams during the competition. We ensemble the models developed by each team. The single models were trained in two stages: the first stage used all EEG IDs, and the second stage used data with a vote count of 8 or more. These single models were then ensemble using stacking.</p>\n<h1>Details of the submission</h1>\n<h1>P-SHA Team Single Model</h1>\n<h2>iida part</h2>\n<p>＜Model architecture＞<br>\n・Various spectrogram transformations (STFT, MFCC, LFCC, RMS) were applied to the Raw EEG data.<br>\n・After spectrogram transformation, the data, including Kaggle spectrograms, were fed into a Timm Backbone or GRU for training.<br>\n・For the Raw EEG data, training was also conducted after passing through a 1D CNN, followed by a Timm Backbone.  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3484256%2Fb4ac74e0b7f9b5b129d2fb8fc9711caa%2Fimage.png?generation=1712646730283771&amp;alt=media\">  </p>\n<p>＜1D CNN＞<br>\nI made some improvements to the 1D CNN model provided by Nischay Dhankhar. I preserved the sequence length at 10000 for all convolutions by using various kernel sizes (kernels=[1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,35,63,127]) and setting padding=(kernel_size-1)//2. The outputs were concatenated in the channel direction. Then, the data was passed through a 1D CNN, and trained using a Timm backbone with the two-dimensional shape of (features from various kernel sizes, time).</p>\n<p>＜Calculation of train loss＞<br>\nThe individual learning outcomes of various spectrograms were also taken into account in the train loss calculation.</p>\n<pre><code># train loss\n amp.autocast(cfg.enable_amp):\n     cfg.mixup_p &gt; rand:\n         y,h1,h2,h3,h4,h5,h6 = model(xs)\n         loss = mixup\n         loss_h1 = mixup\n         loss_h2 = mixup\n         loss_h3 = mixup\n         loss_h4 = mixup\n         loss_h5 = mixup\n         loss_h6 = mixup\n    :\n         y,h1,h2,h3,h4,h5,h6 = model(x)\n         loss = loss\n         loss_h1 = loss\n         loss_h2 = loss\n         loss_h3 = loss\n         loss_h4 = loss\n         loss_h5 = loss\n         loss_h6 = loss\n    loss = loss+(loss_h1+loss_h2+loss_h3+loss_h4+loss_h5+loss_h6)\n</code></pre>\n<p>＜Augmentation＞<br>\n・mixup(p=0.5, alpha=1)<br>\n・HorizontalFlip<br>\n・MultiplicativeNoise<br>\n・CoarseDropout<br>\n・XYMasking</p>\n<p>＜timm backbone＞<br>\n・tf_efficientnetv2_s.in21k<br>\n・tf_efficientnet_b3.in1k</p>\n<p>＜Other tips＞<br>\nSimilar to eikichi part, two-stage learning was effective.</p>\n<h2>eikichi part</h2>\n<h2>Model:</h2>\n<p>I constructed 5 models and combined them with fully connected layers.</p>\n<ul>\n<li>model1: timm model <ul>\n<li>backbone: tf_efficientnetv2_s.in21k</li>\n<li>input: kaggle_spec</li></ul></li>\n<li>model2: timm model <ul>\n<li>backbone: tf_efficientnetv2_s.in21k</li>\n<li>input: spec generated from eeg</li></ul></li>\n<li>model3: EEG1D model<ul>\n<li>Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features</li>\n<li>input: EEG at the front of the brain (Fp1-F7, F7-T3, Fp1-F3, F3-C3, Fp2-F8, F8-T4, Fp2-F4, F4-C4)</li></ul></li>\n<li>model4: EEG1D model<ul>\n<li>Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features</li>\n<li>input: EEG at the back of the brain (T3-T5, T5-O1, C3-P3, P3-O1, T4-T6, T6-O2, C4-P4, P4-O2)</li></ul></li>\n<li>model5: EEGMegaNet model</li>\n</ul>\n<h2>Augmentation</h2>\n<ul>\n<li>mixup(p=0.5, alpha=0.4)</li>\n<li>HorizontalFlip(p=0.5)</li>\n</ul>\n<h2>Tips</h2>\n<ul>\n<li>2-stage training <ul>\n<li>1st stage: train: all, valid: extract only middle data from all</li>\n<li>2nd stage: train: total_vote ≥ 8, valid: extract only middle data from total_vote ≥ 8 &amp; total_vote ≤ 20</li></ul></li>\n<li>I did not use butter_lowpass_filter because the CV was better without it.</li>\n</ul>\n<h1>TR Team Single Model</h1>\n<ul>\n<li>feature extractor (wave to image)</li>\n</ul>\n<p>CWT or MelSpectrogam</p>\n<pre><code># CWT \nCWT(fmin  // \n\n# MelSpec \nnn.Sequential(\n    T.MelSpectrogram(\n        sample_rate \n        n_fft \n        win_length \n        hop_length  // \n        f_min \n        f_max \n        n_mels \n    ),\n    T.AmplitudeToDB(top_db \n)\n</code></pre>\n<ul>\n<li>timm backbone</li>\n</ul>\n<p>maxvit_tiny(cwt or melspec), swinv2_base (cwt), convnextv2_base (cwt)<br>\nmaxvit was so good.</p>\n<ul>\n<li>augmentation</li>\n</ul>\n<p>TimeMaking, left-right-flip(like below)</p>\n<pre><code>\n self.training:\n     random.random() &lt; .:\n         = x.clone()\n       \n        [:, :, :, :] = _x[:, :, :, :]\n        [:, :, :, :] = _x[:, :, :, :]\n       \n        [:, :, :, :] = _x[:, :, :, :]\n        [:, :, :, :] = _x[:, :, :, :]\n</code></pre>\n<ul>\n<li>other tips</li>\n</ul>\n<p>We trained models with 20 epochs. (first 10 epochs, we used all votes data. after 10 epochs, swithced  votes &gt;= 8 data)<br>\nWe didn’t use kaggle-spectrogram. Using only eeg-data was better.</p>\n<h1>Stacking Model (T88 part)</h1>\n<p>2dcnn stacking model was trained using single model oof data. Only data with a total vote of 8 or more were used for training and evaluation.   <br>\nThe stacking model was very effective, achieving a final score of (public : 0.23 / private : 0.28). The simple average for the same single model configuration was (public : 0.24 / Private : 0.29), showing the superiority of stacking.</p>\n<h2>Input Single Models</h2>\n<p>9 models were used.</p>\n<ul>\n<li>iida062 (Public : 0.27 / Private : 0.33)</li>\n<li>iida139 (Public : 0.25 / Private : 0.31)</li>\n<li>iida164 (Public : 0.25 / Private : 0.31)</li>\n<li>iida169 (Public : 0.25 / Private : 0.32)</li>\n<li>eikichi038 (Public : 0.28 / Private : 0.34)</li>\n<li>tomo089 (Public : 0.26 / Private : 0.31)</li>\n<li>tomo092 (Public : 0.27 / Private : 0.31)</li>\n<li>tomo094 (Public : 0.28 / Private : 0.34)</li>\n<li>tomo095 (Public : 0.29 / Private : 0.34 )</li>\n</ul>\n<h2>Model Architecture</h2>\n<p>We adopted the 2D-CNN stacking model, which was the best after trying several NN-based stacking models.    </p>\n<ul>\n<li>The format of the input data</li>\n</ul>\n<p>(Channel, Height, Width) = (1, Class, Model) = (1, 6, 9)  </p>\n<ul>\n<li>architecture</li>\n</ul>\n<pre><code> :\n    def :\n        super.\n        self.conv2d1 = nn., padding=, stride=)\n        self.conv2d2 = nn., padding=, stride=)\n        self.fc1 = nn.\n        self.dropout1 = nn.\n        self.fc2 = nn.\n\n    def forward(self, x):\n        x = relu(self.conv2d1(x))\n        x = relu(self.conv2d2(x))\n        x = torch.flatten(x, start_dim=)\n        x = relu(self.fc1(x))\n        x = self.dropout1(x)\n        x = self.fc2(x)\n        return x\n</code></pre>\n<h2>Model Optimization</h2>\n<p>The combination of single model as input and hyper parameters such as dropout were optimized using optuna. To reduce overfitting, the average of 3 cvs with different seed was used as the optimization metric.</p>\n<h2>Training</h2>\n<ul>\n<li>lr = 5e-4</li>\n<li>epoch = 20</li>\n<li>batch_size = 16</li>\n<li>The data used for the final sub is trained using all data without cutting folds　　</li>\n<li>5 SeedAvg  </li>\n</ul>\n<h1>others</h1>\n<ul>\n<li>The distribution of the private dataset was a matter of interest. To clarify this, we probed based on the similarity of the distribution of predictions and determined that the entire test data was similar to the train data for vote8 and above. Based on this, CV strategy and sub-selections were determined.</li>\n</ul>",
  "messages": [
    {
      "id": "2743025",
      "postDate": "04/09/2024 07:20:42",
      "content": "<h1>11th place solution for the Harmful Brain Activity Classification Competition</h1>\n<p>Team : kansai-kaggler ( <a href=\"https://www.kaggle.com/tomohiroh\" target=\"_blank\">@tomohiroh</a> <a href=\"https://www.kaggle.com/ryushisa\" target=\"_blank\">@ryushisa</a> <a href=\"https://www.kaggle.com/shunyaiida\" target=\"_blank\">@shunyaiida</a> <a href=\"https://www.kaggle.com/ekichi\" target=\"_blank\">@ekichi</a> <a href=\"https://www.kaggle.com/t88take\" target=\"_blank\">@t88take</a>)  </p>\n<p>First of all, I would like to express my sincere gratitude to the Kaggle staff and all the hosts of the competition for organizing such an amazing competition. It was a very enjoyable competition where we could consider various approaches to EEG data. Furthermore, We are very happy that three people from our team have become Kaggle Competitions Master.</p>\n<p>We would like to share our solution.   </p>\n<h1>Overview</h1>\n<p>Our team merged with the P-SHA and TR teams during the competition. We ensemble the models developed by each team. The single models were trained in two stages: the first stage used all EEG IDs, and the second stage used data with a vote count of 8 or more. These single models were then ensemble using stacking.</p>\n<h1>Details of the submission</h1>\n<h1>P-SHA Team Single Model</h1>\n<h2>iida part</h2>\n<p>＜Model architecture＞<br>\n・Various spectrogram transformations (STFT, MFCC, LFCC, RMS) were applied to the Raw EEG data.<br>\n・After spectrogram transformation, the data, including Kaggle spectrograms, were fed into a Timm Backbone or GRU for training.<br>\n・For the Raw EEG data, training was also conducted after passing through a 1D CNN, followed by a Timm Backbone.  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3484256%2Fb4ac74e0b7f9b5b129d2fb8fc9711caa%2Fimage.png?generation=1712646730283771&amp;alt=media\">  </p>\n<p>＜1D CNN＞<br>\nI made some improvements to the 1D CNN model provided by Nischay Dhankhar. I preserved the sequence length at 10000 for all convolutions by using various kernel sizes (kernels=[1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,35,63,127]) and setting padding=(kernel_size-1)//2. The outputs were concatenated in the channel direction. Then, the data was passed through a 1D CNN, and trained using a Timm backbone with the two-dimensional shape of (features from various kernel sizes, time).</p>\n<p>＜Calculation of train loss＞<br>\nThe individual learning outcomes of various spectrograms were also taken into account in the train loss calculation.</p>\n<pre><code># train loss\n amp.autocast(cfg.enable_amp):\n     cfg.mixup_p &gt; rand:\n         y,h1,h2,h3,h4,h5,h6 = model(xs)\n         loss = mixup\n         loss_h1 = mixup\n         loss_h2 = mixup\n         loss_h3 = mixup\n         loss_h4 = mixup\n         loss_h5 = mixup\n         loss_h6 = mixup\n    :\n         y,h1,h2,h3,h4,h5,h6 = model(x)\n         loss = loss\n         loss_h1 = loss\n         loss_h2 = loss\n         loss_h3 = loss\n         loss_h4 = loss\n         loss_h5 = loss\n         loss_h6 = loss\n    loss = loss+(loss_h1+loss_h2+loss_h3+loss_h4+loss_h5+loss_h6)\n</code></pre>\n<p>＜Augmentation＞<br>\n・mixup(p=0.5, alpha=1)<br>\n・HorizontalFlip<br>\n・MultiplicativeNoise<br>\n・CoarseDropout<br>\n・XYMasking</p>\n<p>＜timm backbone＞<br>\n・tf_efficientnetv2_s.in21k<br>\n・tf_efficientnet_b3.in1k</p>\n<p>＜Other tips＞<br>\nSimilar to eikichi part, two-stage learning was effective.</p>\n<h2>eikichi part</h2>\n<h2>Model:</h2>\n<p>I constructed 5 models and combined them with fully connected layers.</p>\n<ul>\n<li>model1: timm model <ul>\n<li>backbone: tf_efficientnetv2_s.in21k</li>\n<li>input: kaggle_spec</li></ul></li>\n<li>model2: timm model <ul>\n<li>backbone: tf_efficientnetv2_s.in21k</li>\n<li>input: spec generated from eeg</li></ul></li>\n<li>model3: EEG1D model<ul>\n<li>Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features</li>\n<li>input: EEG at the front of the brain (Fp1-F7, F7-T3, Fp1-F3, F3-C3, Fp2-F8, F8-T4, Fp2-F4, F4-C4)</li></ul></li>\n<li>model4: EEG1D model<ul>\n<li>Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features</li>\n<li>input: EEG at the back of the brain (T3-T5, T5-O1, C3-P3, P3-O1, T4-T6, T6-O2, C4-P4, P4-O2)</li></ul></li>\n<li>model5: EEGMegaNet model</li>\n</ul>\n<h2>Augmentation</h2>\n<ul>\n<li>mixup(p=0.5, alpha=0.4)</li>\n<li>HorizontalFlip(p=0.5)</li>\n</ul>\n<h2>Tips</h2>\n<ul>\n<li>2-stage training <ul>\n<li>1st stage: train: all, valid: extract only middle data from all</li>\n<li>2nd stage: train: total_vote ≥ 8, valid: extract only middle data from total_vote ≥ 8 &amp; total_vote ≤ 20</li></ul></li>\n<li>I did not use butter_lowpass_filter because the CV was better without it.</li>\n</ul>\n<h1>TR Team Single Model</h1>\n<ul>\n<li>feature extractor (wave to image)</li>\n</ul>\n<p>CWT or MelSpectrogam</p>\n<pre><code># CWT \nCWT(fmin  // \n\n# MelSpec \nnn.Sequential(\n    T.MelSpectrogram(\n        sample_rate \n        n_fft \n        win_length \n        hop_length  // \n        f_min \n        f_max \n        n_mels \n    ),\n    T.AmplitudeToDB(top_db \n)\n</code></pre>\n<ul>\n<li>timm backbone</li>\n</ul>\n<p>maxvit_tiny(cwt or melspec), swinv2_base (cwt), convnextv2_base (cwt)<br>\nmaxvit was so good.</p>\n<ul>\n<li>augmentation</li>\n</ul>\n<p>TimeMaking, left-right-flip(like below)</p>\n<pre><code>\n self.training:\n     random.random() &lt; .:\n         = x.clone()\n       \n        [:, :, :, :] = _x[:, :, :, :]\n        [:, :, :, :] = _x[:, :, :, :]\n       \n        [:, :, :, :] = _x[:, :, :, :]\n        [:, :, :, :] = _x[:, :, :, :]\n</code></pre>\n<ul>\n<li>other tips</li>\n</ul>\n<p>We trained models with 20 epochs. (first 10 epochs, we used all votes data. after 10 epochs, swithced  votes &gt;= 8 data)<br>\nWe didn’t use kaggle-spectrogram. Using only eeg-data was better.</p>\n<h1>Stacking Model (T88 part)</h1>\n<p>2dcnn stacking model was trained using single model oof data. Only data with a total vote of 8 or more were used for training and evaluation.   <br>\nThe stacking model was very effective, achieving a final score of (public : 0.23 / private : 0.28). The simple average for the same single model configuration was (public : 0.24 / Private : 0.29), showing the superiority of stacking.</p>\n<h2>Input Single Models</h2>\n<p>9 models were used.</p>\n<ul>\n<li>iida062 (Public : 0.27 / Private : 0.33)</li>\n<li>iida139 (Public : 0.25 / Private : 0.31)</li>\n<li>iida164 (Public : 0.25 / Private : 0.31)</li>\n<li>iida169 (Public : 0.25 / Private : 0.32)</li>\n<li>eikichi038 (Public : 0.28 / Private : 0.34)</li>\n<li>tomo089 (Public : 0.26 / Private : 0.31)</li>\n<li>tomo092 (Public : 0.27 / Private : 0.31)</li>\n<li>tomo094 (Public : 0.28 / Private : 0.34)</li>\n<li>tomo095 (Public : 0.29 / Private : 0.34 )</li>\n</ul>\n<h2>Model Architecture</h2>\n<p>We adopted the 2D-CNN stacking model, which was the best after trying several NN-based stacking models.    </p>\n<ul>\n<li>The format of the input data</li>\n</ul>\n<p>(Channel, Height, Width) = (1, Class, Model) = (1, 6, 9)  </p>\n<ul>\n<li>architecture</li>\n</ul>\n<pre><code> :\n    def :\n        super.\n        self.conv2d1 = nn., padding=, stride=)\n        self.conv2d2 = nn., padding=, stride=)\n        self.fc1 = nn.\n        self.dropout1 = nn.\n        self.fc2 = nn.\n\n    def forward(self, x):\n        x = relu(self.conv2d1(x))\n        x = relu(self.conv2d2(x))\n        x = torch.flatten(x, start_dim=)\n        x = relu(self.fc1(x))\n        x = self.dropout1(x)\n        x = self.fc2(x)\n        return x\n</code></pre>\n<h2>Model Optimization</h2>\n<p>The combination of single model as input and hyper parameters such as dropout were optimized using optuna. To reduce overfitting, the average of 3 cvs with different seed was used as the optimization metric.</p>\n<h2>Training</h2>\n<ul>\n<li>lr = 5e-4</li>\n<li>epoch = 20</li>\n<li>batch_size = 16</li>\n<li>The data used for the final sub is trained using all data without cutting folds　　</li>\n<li>5 SeedAvg  </li>\n</ul>\n<h1>others</h1>\n<ul>\n<li>The distribution of the private dataset was a matter of interest. To clarify this, we probed based on the similarity of the distribution of predictions and determined that the entire test data was similar to the train data for vote8 and above. Based on this, CV strategy and sub-selections were determined.</li>\n</ul>",
      "rawMarkdown": "# 11th place solution for the Harmful Brain Activity Classification Competition  \nTeam : kansai-kaggler ( @tomohiroh @ryushisa @shunyaiida @ekichi @t88take)  \n\nFirst of all, I would like to express my sincere gratitude to the Kaggle staff and all the hosts of the competition for organizing such an amazing competition. It was a very enjoyable competition where we could consider various approaches to EEG data. Furthermore, We are very happy that three people from our team have become Kaggle Competitions Master.\n\nWe would like to share our solution.   \n\n\n# Overview\n\nOur team merged with the P-SHA and TR teams during the competition. We ensemble the models developed by each team. The single models were trained in two stages: the first stage used all EEG IDs, and the second stage used data with a vote count of 8 or more. These single models were then ensemble using stacking.\n\n# Details of the submission \n# P-SHA Team Single Model\n\n## iida part\n\n＜Model architecture＞\n・Various spectrogram transformations (STFT, MFCC, LFCC, RMS) were applied to the Raw EEG data.\n・After spectrogram transformation, the data, including Kaggle spectrograms, were fed into a Timm Backbone or GRU for training.\n・For the Raw EEG data, training was also conducted after passing through a 1D CNN, followed by a Timm Backbone.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3484256%2Fb4ac74e0b7f9b5b129d2fb8fc9711caa%2Fimage.png?generation=1712646730283771&alt=media)  \n\n＜1D CNN＞\nI made some improvements to the 1D CNN model provided by Nischay Dhankhar. I preserved the sequence length at 10000 for all convolutions by using various kernel sizes (kernels=[1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,35,63,127]) and setting padding=(kernel_size-1)//2. The outputs were concatenated in the channel direction. Then, the data was passed through a 1D CNN, and trained using a Timm backbone with the two-dimensional shape of (features from various kernel sizes, time).\n\n＜Calculation of train loss＞\nThe individual learning outcomes of various spectrograms were also taken into account in the train loss calculation.\n\n```\n# train loss\nwith amp.autocast(cfg.enable_amp):\n    if cfg.mixup_p > rand:\n         y,h1,h2,h3,h4,h5,h6 = model(xs)\n         loss = mixup_loss(loss_func, y, y_a, y_b, lam)\n         loss_h1 = mixup_loss(loss_func, h1, y_a, y_b, lam)\n         loss_h2 = mixup_loss(loss_func, h2, y_a, y_b, lam)\n         loss_h3 = mixup_loss(loss_func, h3, y_a, y_b, lam)\n         loss_h4 = mixup_loss(loss_func, h4, y_a, y_b, lam)\n         loss_h5 = mixup_loss(loss_func, h5, y_a, y_b, lam)\n         loss_h6 = mixup_loss(loss_func, h6, y_a, y_b, lam)\n    else:\n         y,h1,h2,h3,h4,h5,h6 = model(x)\n         loss = loss_func(y, t)\n         loss_h1 = loss_func(h1, t)\n         loss_h2 = loss_func(h2, t)\n         loss_h3 = loss_func(h3, t)\n         loss_h4 = loss_func(h4, t)\n         loss_h5 = loss_func(h5, t)\n         loss_h6 = loss_func(h6, t)\n    loss = loss+(loss_h1+loss_h2+loss_h3+loss_h4+loss_h5+loss_h6) / 6.\n```\n＜Augmentation＞\n・mixup(p=0.5, alpha=1)\n・HorizontalFlip\n・MultiplicativeNoise\n・CoarseDropout\n・XYMasking\n\n＜timm backbone＞\n・tf_efficientnetv2_s.in21k\n・tf_efficientnet_b3.in1k\n\n＜Other tips＞\nSimilar to eikichi part, two-stage learning was effective.\n\n## eikichi part\n\n## Model:\n I constructed 5 models and combined them with fully connected layers.\n\n* model1: timm model \n    * backbone: tf_efficientnetv2_s.in21k\n    * input: kaggle_spec\n* model2: timm model \n    * backbone: tf_efficientnetv2_s.in21k\n    * input: spec generated from eeg\n* model3: EEG1D model\n    * Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features\n    * input: EEG at the front of the brain (Fp1-F7, F7-T3, Fp1-F3, F3-C3, Fp2-F8, F8-T4, Fp2-F4, F4-C4)\n* model4: EEG1D model\n    * Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features\n    * input: EEG at the back of the brain (T3-T5, T5-O1, C3-P3, P3-O1, T4-T6, T6-O2, C4-P4, P4-O2)\n* model5: EEGMegaNet model\n\n## Augmentation\n\n* mixup(p=0.5, alpha=0.4)\n* HorizontalFlip(p=0.5)\n\n## Tips\n\n* 2-stage training \n    * 1st stage: train: all, valid: extract only middle data from all\n    * 2nd stage: train: total_vote ≥ 8, valid: extract only middle data from total_vote ≥ 8 & total_vote ≤ 20\n* I did not use butter_lowpass_filter because the CV was better without it.\n\n# TR Team Single Model\n\n* feature extractor (wave to image)\n\nCWT or MelSpectrogam\n```\n# CWT parameters\nCWT(fmin = 0.5, fmax = 25, hop_length = 200 * 50 // self.size)\n\n# MelSpec parameters\nnn.Sequential(\n    T.MelSpectrogram(\n        sample_rate = 200,\n        n_fft = 1024,\n        win_length = 128,\n        hop_length = (200 * 50) // self.size,\n        f_min = 0,\n        f_max = 20,\n        n_mels = 128,\n    ),\n    T.AmplitudeToDB(top_db = 80),\n)\n```\n\n\n* timm backbone\n\nmaxvit_tiny(cwt or melspec), swinv2_base (cwt), convnextv2_base (cwt)\nmaxvit was so good.\n\n\n* augmentation\n\nTimeMaking, left-right-flip(like below)\n```\n# left-right-flip\nif self.training:\n    if random.random() < 0.50:\n        _x = x.clone()\n       # LL <=> RL\n        x[:, 0:4, :, :] = _x[:, 4:8, :, :]\n        x[:, 4:8, :, :] = _x[:, 0:4, :, :]\n       # LP <=> RP\n        x[:, 8:12, :, :] = _x[:, 12:16, :, :]\n        x[:, 12:16, :, :] = _x[:, 8:12, :, :]\n```\n\n* other tips\n\nWe trained models with 20 epochs. (first 10 epochs, we used all votes data. after 10 epochs, swithced  votes >= 8 data)\nWe didn’t use kaggle-spectrogram. Using only eeg-data was better.\n\n# Stacking Model (T88 part)\n\n2dcnn stacking model was trained using single model oof data. Only data with a total vote of 8 or more were used for training and evaluation.   \nThe stacking model was very effective, achieving a final score of (public : 0.23 / private : 0.28). The simple average for the same single model configuration was (public : 0.24 / Private : 0.29), showing the superiority of stacking.\n\n## Input Single Models\n\n9 models were used.\n\n* iida062 (Public : 0.27 / Private : 0.33)\n* iida139 (Public : 0.25 / Private : 0.31)\n* iida164 (Public : 0.25 / Private : 0.31)\n* iida169 (Public : 0.25 / Private : 0.32)\n* eikichi038 (Public : 0.28 / Private : 0.34)\n* tomo089 (Public : 0.26 / Private : 0.31)\n* tomo092 (Public : 0.27 / Private : 0.31)\n* tomo094 (Public : 0.28 / Private : 0.34)\n* tomo095 (Public : 0.29 / Private : 0.34 )\n\n## Model Architecture\n\nWe adopted the 2D-CNN stacking model, which was the best after trying several NN-based stacking models.    \n\n* The format of the input data\n\n(Channel, Height, Width) = (1, Class, Model) = (1, 6, 9)  \n\n\n* architecture\n```\nclass HMSStacking2DCNN(nn.Module):\n    def __init__(self, num_classes):\n        super().__init__()\n        self.conv2d1 = nn.Conv2d(1, 8, kernel_size=(1,3), padding=0, stride=1)\n        self.conv2d2 = nn.Conv2d(8, 16, kernel_size=(1,3), padding=0, stride=1)\n        self.fc1 = nn.Linear(480, 480)\n        self.dropout1 = nn.Dropout(p=cfg.dropout)\n        self.fc2 = nn.Linear(480, num_classes)\n        \n    def forward(self, x):\n        x = F.relu(self.conv2d1(x))\n        x = F.relu(self.conv2d2(x))\n        x = torch.flatten(x, start_dim=1)\n        x = F.relu(self.fc1(x))\n        x = self.dropout1(x)\n        x = self.fc2(x)\n        return x\n```\n\n## Model Optimization\n\nThe combination of single model as input and hyper parameters such as dropout were optimized using optuna. To reduce overfitting, the average of 3 cvs with different seed was used as the optimization metric.\n\n\n## Training\n\n* lr = 5e-4\n* epoch = 20\n* batch_size = 16\n* The data used for the final sub is trained using all data without cutting folds　　\n* 5 SeedAvg  \n\n# others\n\n* The distribution of the private dataset was a matter of interest. To clarify this, we probed based on the similarity of the distribution of predictions and determined that the entire test data was similar to the train data for vote8 and above. Based on this, CV strategy and sub-selections were determined.",
      "votes": null
    },
    {
      "id": "2743063",
      "postDate": "04/09/2024 07:50:04",
      "content": "<p>Congratulations for the gold <a href=\"https://www.kaggle.com/t88take\" target=\"_blank\">@t88take</a> and team! Great solution.</p>",
      "rawMarkdown": "Congratulations for the gold @t88take and team! Great solution.",
      "votes": null
    },
    {
      "id": "2743132",
      "postDate": "04/09/2024 08:39:03",
      "content": "<p>Congratulations. It would be interesting to see your work in more detail.</p>",
      "rawMarkdown": "Congratulations. It would be interesting to see your work in more detail.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2743063,
      "author_name": "chaudharypriyanshu",
      "author_url": "",
      "post_date": "04/09/2024 07:50:04",
      "content": "<p>Congratulations for the gold <a href=\"https://www.kaggle.com/t88take\" target=\"_blank\">@t88take</a> and team! Great solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2743132,
      "author_name": "zaakciiru",
      "author_url": "",
      "post_date": "04/09/2024 08:39:03",
      "content": "<p>Congratulations. It would be interesting to see your work in more detail.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2743025": "# 11th place solution for the Harmful Brain Activity Classification Competition  \nTeam : kansai-kaggler ( @tomohiroh @ryushisa @shunyaiida @ekichi @t88take)  \n\nFirst of all, I would like to express my sincere gratitude to the Kaggle staff and all the hosts of the competition for organizing such an amazing competition. It was a very enjoyable competition where we could consider various approaches to EEG data. Furthermore, We are very happy that three people from our team have become Kaggle Competitions Master.\n\nWe would like to share our solution.   \n\n\n# Overview\n\nOur team merged with the P-SHA and TR teams during the competition. We ensemble the models developed by each team. The single models were trained in two stages: the first stage used all EEG IDs, and the second stage used data with a vote count of 8 or more. These single models were then ensemble using stacking.\n\n# Details of the submission \n# P-SHA Team Single Model\n\n## iida part\n\n＜Model architecture＞\n・Various spectrogram transformations (STFT, MFCC, LFCC, RMS) were applied to the Raw EEG data.\n・After spectrogram transformation, the data, including Kaggle spectrograms, were fed into a Timm Backbone or GRU for training.\n・For the Raw EEG data, training was also conducted after passing through a 1D CNN, followed by a Timm Backbone.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3484256%2Fb4ac74e0b7f9b5b129d2fb8fc9711caa%2Fimage.png?generation=1712646730283771&alt=media)  \n\n＜1D CNN＞\nI made some improvements to the 1D CNN model provided by Nischay Dhankhar. I preserved the sequence length at 10000 for all convolutions by using various kernel sizes (kernels=[1,3,5,7,9,11,13,15,17,19,21,23,25,27,29,35,63,127]) and setting padding=(kernel_size-1)//2. The outputs were concatenated in the channel direction. Then, the data was passed through a 1D CNN, and trained using a Timm backbone with the two-dimensional shape of (features from various kernel sizes, time).\n\n＜Calculation of train loss＞\nThe individual learning outcomes of various spectrograms were also taken into account in the train loss calculation.\n\n```\n# train loss\nwith amp.autocast(cfg.enable_amp):\n    if cfg.mixup_p > rand:\n         y,h1,h2,h3,h4,h5,h6 = model(xs)\n         loss = mixup_loss(loss_func, y, y_a, y_b, lam)\n         loss_h1 = mixup_loss(loss_func, h1, y_a, y_b, lam)\n         loss_h2 = mixup_loss(loss_func, h2, y_a, y_b, lam)\n         loss_h3 = mixup_loss(loss_func, h3, y_a, y_b, lam)\n         loss_h4 = mixup_loss(loss_func, h4, y_a, y_b, lam)\n         loss_h5 = mixup_loss(loss_func, h5, y_a, y_b, lam)\n         loss_h6 = mixup_loss(loss_func, h6, y_a, y_b, lam)\n    else:\n         y,h1,h2,h3,h4,h5,h6 = model(x)\n         loss = loss_func(y, t)\n         loss_h1 = loss_func(h1, t)\n         loss_h2 = loss_func(h2, t)\n         loss_h3 = loss_func(h3, t)\n         loss_h4 = loss_func(h4, t)\n         loss_h5 = loss_func(h5, t)\n         loss_h6 = loss_func(h6, t)\n    loss = loss+(loss_h1+loss_h2+loss_h3+loss_h4+loss_h5+loss_h6) / 6.\n```\n＜Augmentation＞\n・mixup(p=0.5, alpha=1)\n・HorizontalFlip\n・MultiplicativeNoise\n・CoarseDropout\n・XYMasking\n\n＜timm backbone＞\n・tf_efficientnetv2_s.in21k\n・tf_efficientnet_b3.in1k\n\n＜Other tips＞\nSimilar to eikichi part, two-stage learning was effective.\n\n## eikichi part\n\n## Model:\n I constructed 5 models and combined them with fully connected layers.\n\n* model1: timm model \n    * backbone: tf_efficientnetv2_s.in21k\n    * input: kaggle_spec\n* model2: timm model \n    * backbone: tf_efficientnetv2_s.in21k\n    * input: spec generated from eeg\n* model3: EEG1D model\n    * Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features\n    * input: EEG at the front of the brain (Fp1-F7, F7-T3, Fp1-F3, F3-C3, Fp2-F8, F8-T4, Fp2-F4, F4-C4)\n* model4: EEG1D model\n    * Expanded the kernels from [3,7,5,9] to [3,7,5,9,15,17,19,21] to put more emphasis on long-term features\n    * input: EEG at the back of the brain (T3-T5, T5-O1, C3-P3, P3-O1, T4-T6, T6-O2, C4-P4, P4-O2)\n* model5: EEGMegaNet model\n\n## Augmentation\n\n* mixup(p=0.5, alpha=0.4)\n* HorizontalFlip(p=0.5)\n\n## Tips\n\n* 2-stage training \n    * 1st stage: train: all, valid: extract only middle data from all\n    * 2nd stage: train: total_vote ≥ 8, valid: extract only middle data from total_vote ≥ 8 & total_vote ≤ 20\n* I did not use butter_lowpass_filter because the CV was better without it.\n\n# TR Team Single Model\n\n* feature extractor (wave to image)\n\nCWT or MelSpectrogam\n```\n# CWT parameters\nCWT(fmin = 0.5, fmax = 25, hop_length = 200 * 50 // self.size)\n\n# MelSpec parameters\nnn.Sequential(\n    T.MelSpectrogram(\n        sample_rate = 200,\n        n_fft = 1024,\n        win_length = 128,\n        hop_length = (200 * 50) // self.size,\n        f_min = 0,\n        f_max = 20,\n        n_mels = 128,\n    ),\n    T.AmplitudeToDB(top_db = 80),\n)\n```\n\n\n* timm backbone\n\nmaxvit_tiny(cwt or melspec), swinv2_base (cwt), convnextv2_base (cwt)\nmaxvit was so good.\n\n\n* augmentation\n\nTimeMaking, left-right-flip(like below)\n```\n# left-right-flip\nif self.training:\n    if random.random() < 0.50:\n        _x = x.clone()\n       # LL <=> RL\n        x[:, 0:4, :, :] = _x[:, 4:8, :, :]\n        x[:, 4:8, :, :] = _x[:, 0:4, :, :]\n       # LP <=> RP\n        x[:, 8:12, :, :] = _x[:, 12:16, :, :]\n        x[:, 12:16, :, :] = _x[:, 8:12, :, :]\n```\n\n* other tips\n\nWe trained models with 20 epochs. (first 10 epochs, we used all votes data. after 10 epochs, swithced  votes >= 8 data)\nWe didn’t use kaggle-spectrogram. Using only eeg-data was better.\n\n# Stacking Model (T88 part)\n\n2dcnn stacking model was trained using single model oof data. Only data with a total vote of 8 or more were used for training and evaluation.   \nThe stacking model was very effective, achieving a final score of (public : 0.23 / private : 0.28). The simple average for the same single model configuration was (public : 0.24 / Private : 0.29), showing the superiority of stacking.\n\n## Input Single Models\n\n9 models were used.\n\n* iida062 (Public : 0.27 / Private : 0.33)\n* iida139 (Public : 0.25 / Private : 0.31)\n* iida164 (Public : 0.25 / Private : 0.31)\n* iida169 (Public : 0.25 / Private : 0.32)\n* eikichi038 (Public : 0.28 / Private : 0.34)\n* tomo089 (Public : 0.26 / Private : 0.31)\n* tomo092 (Public : 0.27 / Private : 0.31)\n* tomo094 (Public : 0.28 / Private : 0.34)\n* tomo095 (Public : 0.29 / Private : 0.34 )\n\n## Model Architecture\n\nWe adopted the 2D-CNN stacking model, which was the best after trying several NN-based stacking models.    \n\n* The format of the input data\n\n(Channel, Height, Width) = (1, Class, Model) = (1, 6, 9)  \n\n\n* architecture\n```\nclass HMSStacking2DCNN(nn.Module):\n    def __init__(self, num_classes):\n        super().__init__()\n        self.conv2d1 = nn.Conv2d(1, 8, kernel_size=(1,3), padding=0, stride=1)\n        self.conv2d2 = nn.Conv2d(8, 16, kernel_size=(1,3), padding=0, stride=1)\n        self.fc1 = nn.Linear(480, 480)\n        self.dropout1 = nn.Dropout(p=cfg.dropout)\n        self.fc2 = nn.Linear(480, num_classes)\n        \n    def forward(self, x):\n        x = F.relu(self.conv2d1(x))\n        x = F.relu(self.conv2d2(x))\n        x = torch.flatten(x, start_dim=1)\n        x = F.relu(self.fc1(x))\n        x = self.dropout1(x)\n        x = self.fc2(x)\n        return x\n```\n\n## Model Optimization\n\nThe combination of single model as input and hyper parameters such as dropout were optimized using optuna. To reduce overfitting, the average of 3 cvs with different seed was used as the optimization metric.\n\n\n## Training\n\n* lr = 5e-4\n* epoch = 20\n* batch_size = 16\n* The data used for the final sub is trained using all data without cutting folds　　\n* 5 SeedAvg  \n\n# others\n\n* The distribution of the private dataset was a matter of interest. To clarify this, we probed based on the similarity of the distribution of predictions and determined that the entire test data was similar to the train data for vote8 and above. Based on this, CV strategy and sub-selections were determined.",
    "2743063": "Congratulations for the gold @t88take and team! Great solution.",
    "2743132": "Congratulations. It would be interesting to see your work in more detail."
  },
  "source": "meta"
}