{
  "id": 492193,
  "title": "15th Place Solution",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/r-b-15th-place-solution",
  "author_name": "",
  "post_date": "2024-04-09T15:47:47.370Z",
  "votes": 60,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to Harvard Medical School and Kaggle for hosting this fun competition. The competition was very rewarding for us, as there was so much to learn!</p>\n<p>TLDR; <a href=\"https://www.kaggle.com/ihebch\" target=\"_blank\">@ihebch</a> and I built a multimodal model based on three feature extractors. These are an improved 1D-wavenet, a trainable STFT, and an image model with attention pooling for the competition spectrograms. We used <code>hgnetv2_b4.ssld_stage2_ft_in1k</code> and <code>tf_efficientnetv2_s.in21k_ft_in1k</code> backbones with different architecture configurations to diversify our ensemble. For more details, keep reading!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F0bdfe4bb5759dfd538b2e473c9cac962%2Fhms-GeneralArch.jpg?generation=1712621808361384&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<h2>Fast Experimentation</h2>\n<p>One of the most important parts of our pipeline was fast experimentation. This allowed us to try creative and crazy ideas quickly. The key points for us were to use smaller inputs, use smaller backbones, and create spectrograms on GPU using nnAudio. In the last few weeks of the competition, we scaled up from 8x to 24x node differences, increased the size of the model backbones, and concatenated the three feature vectors into one multimodal model.</p>\n<p><strong>Model 1: Improved 1D Wavenet</strong></p>\n<p>Our best single architecture is a wavenet model with some modifications. We add Mish as a third activation function and we use residual connections from the input convolution rather than within each layer. These changes allowed us to use deeper wavenets but we found no significant improvements going beyond 7 blocks. Finally, we added a downsample convolution after the last wave block to significantly reduce vRam usage.</p>\n<p>Ours</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F60bcbfa7f8bcde74c05b35c4b81eb392%2Fhms-WaveBlock_Ours.jpg?generation=1712621883006223&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 95%\"></p>\n<p>Original</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F151172231e6f88767c2bcd19a5b86e3b%2Fhms-WaveBlock_Original.jpg?generation=1712621909762554&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 95%\"></p>\n<p>Each input sequence is passed through a wavenet with the improved blocks, and the outputs are stacked into a single-channel image. All but two of the output channels are stacked together, with the last two output channels being appended to the end of the image. We did this to create an area where features from all 24 node differences are nearby. The stacked outputs are then passed through a timm backbone to get a feature vector.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F049d909627d392a2b69bc10430ee048e%2Fhms-WaveNet_stack.jpg?generation=1712621986486094&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p><strong>Model 2: STFT</strong></p>\n<p>We generated STFTs on GPU using nnAudio. This is very fast, and each training run for this part of the model took 3-10 mins (depending on the STFT parameters). After generating the STFT, we restack the frequency bins into 2 columns. Finally, the STFTs for each input are stacked together and passed through an image model to get a feature vector.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F5c61eda5775994d937aca74efb840e19%2Fhms-STFTStack.jpg?generation=1712622013956043&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p>We are still unsure why the restacking method worked, but we wanted to share some more thoughts on this. We thought that maybe by reducing the height of the STFT, we were bringing information from different nodes closer together. We tested this by changing the order in which we stacked the STFTs, and saw significant changes to our CV. This indicated to us that the model was learning inter-STFT features and that the order we stacked mattered. </p>\n<p>Based on these findings, we believed that we were missing some sort of attention mechanism between inputs. We started experimenting with this during the last few days of the competition but unfortunately ran out of time.</p>\n<p><strong>Model 3: Attention Pooling</strong></p>\n<p>The final feature extractor we used was a timm backbone with attention pooling for the 10-min spectrograms. We found that adding the differences between the spectrograms as new images helped, and applying heavy dropout to the edges of images (0-95%) during training also helped.</p>\n<h2>Other Strategies</h2>\n<p><strong>Augmentations</strong></p>\n<p>The most important augmentations for us were vertical brain flipping, horizontal brain flipping, and a variant of the CropCat[1] augmentation. Other augmentations we used were brain-size scaling and random noise. We used Brain Flipping during inference as TTA.</p>\n<p>CropCat</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb6df8385512f3fa34cc45942a7cd3977%2Fhms-CropCat.jpg?generation=1712622065747848&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p><strong>Training Data</strong></p>\n<p>Like other competitors, we found that training on data with &gt;= 10 votes resulted in better LB scores. We used GKF on this subset to create a correlated CV/LB. To create our training dataset, we grouped by <code>eeg_id</code>, selected the rows with the max number of votes, and then removed overlapping segments. We then sampled one data point per eeg_id per epoch. During training, we used weighted kldivergence to emphasize data points with more votes. We also sampled noisy student labels with a probability of 15% to increase model diversity. </p>\n<h2>Final Note</h2>\n<p>Our solution would not have been possible without the informative notebooks and sharing of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Also, we mimicked the code structure of <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>'s solution from the ASL Fingerspelling Competition <a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\" target=\"_blank\">here</a>. Thank you both for sharing. </p>\n<p>Our training code can be found <a href=\"https://github.com/brendanartley/HMS-Competition\" target=\"_blank\">here</a>. Happy Kaggling!</p>\n<h2>Frameworks</h2>\n<ul>\n<li><a href=\"https://lightning.ai/docs/pytorch/stable/\" target=\"_blank\">Pytorch Lightning</a> (training)</li>\n<li><a href=\"https://wandb.ai/site\" target=\"_blank\">Weights + Biases</a> (logging)</li>\n<li><a href=\"https://huggingface.co/timm\" target=\"_blank\">Timm</a> (backbones)</li>\n<li><a href=\"https://github.com/KinWaiCheuk/nnAudio\" target=\"_blank\">nnAudio</a> (STFTs)</li>\n</ul>\n<h1>Sources</h1>\n<p>[1] CropCat: Data Augmentation for Smoothing the Feature Distribution of EEG Signals. <a href=\"https://arxiv.org/abs/2212.06413\" target=\"_blank\">https://arxiv.org/abs/2212.06413</a></p>",
  "messages": [
    {
      "id": "2742505",
      "postDate": "04/09/2024 00:22:23",
      "content": "<p>Thanks to Harvard Medical School and Kaggle for hosting this fun competition. The competition was very rewarding for us, as there was so much to learn!</p>\n<p>TLDR; <a href=\"https://www.kaggle.com/ihebch\" target=\"_blank\">@ihebch</a> and I built a multimodal model based on three feature extractors. These are an improved 1D-wavenet, a trainable STFT, and an image model with attention pooling for the competition spectrograms. We used <code>hgnetv2_b4.ssld_stage2_ft_in1k</code> and <code>tf_efficientnetv2_s.in21k_ft_in1k</code> backbones with different architecture configurations to diversify our ensemble. For more details, keep reading!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F0bdfe4bb5759dfd538b2e473c9cac962%2Fhms-GeneralArch.jpg?generation=1712621808361384&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<h2>Fast Experimentation</h2>\n<p>One of the most important parts of our pipeline was fast experimentation. This allowed us to try creative and crazy ideas quickly. The key points for us were to use smaller inputs, use smaller backbones, and create spectrograms on GPU using nnAudio. In the last few weeks of the competition, we scaled up from 8x to 24x node differences, increased the size of the model backbones, and concatenated the three feature vectors into one multimodal model.</p>\n<p><strong>Model 1: Improved 1D Wavenet</strong></p>\n<p>Our best single architecture is a wavenet model with some modifications. We add Mish as a third activation function and we use residual connections from the input convolution rather than within each layer. These changes allowed us to use deeper wavenets but we found no significant improvements going beyond 7 blocks. Finally, we added a downsample convolution after the last wave block to significantly reduce vRam usage.</p>\n<p>Ours</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F60bcbfa7f8bcde74c05b35c4b81eb392%2Fhms-WaveBlock_Ours.jpg?generation=1712621883006223&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 95%\"></p>\n<p>Original</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F151172231e6f88767c2bcd19a5b86e3b%2Fhms-WaveBlock_Original.jpg?generation=1712621909762554&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 95%\"></p>\n<p>Each input sequence is passed through a wavenet with the improved blocks, and the outputs are stacked into a single-channel image. All but two of the output channels are stacked together, with the last two output channels being appended to the end of the image. We did this to create an area where features from all 24 node differences are nearby. The stacked outputs are then passed through a timm backbone to get a feature vector.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F049d909627d392a2b69bc10430ee048e%2Fhms-WaveNet_stack.jpg?generation=1712621986486094&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p><strong>Model 2: STFT</strong></p>\n<p>We generated STFTs on GPU using nnAudio. This is very fast, and each training run for this part of the model took 3-10 mins (depending on the STFT parameters). After generating the STFT, we restack the frequency bins into 2 columns. Finally, the STFTs for each input are stacked together and passed through an image model to get a feature vector.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F5c61eda5775994d937aca74efb840e19%2Fhms-STFTStack.jpg?generation=1712622013956043&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p>We are still unsure why the restacking method worked, but we wanted to share some more thoughts on this. We thought that maybe by reducing the height of the STFT, we were bringing information from different nodes closer together. We tested this by changing the order in which we stacked the STFTs, and saw significant changes to our CV. This indicated to us that the model was learning inter-STFT features and that the order we stacked mattered. </p>\n<p>Based on these findings, we believed that we were missing some sort of attention mechanism between inputs. We started experimenting with this during the last few days of the competition but unfortunately ran out of time.</p>\n<p><strong>Model 3: Attention Pooling</strong></p>\n<p>The final feature extractor we used was a timm backbone with attention pooling for the 10-min spectrograms. We found that adding the differences between the spectrograms as new images helped, and applying heavy dropout to the edges of images (0-95%) during training also helped.</p>\n<h2>Other Strategies</h2>\n<p><strong>Augmentations</strong></p>\n<p>The most important augmentations for us were vertical brain flipping, horizontal brain flipping, and a variant of the CropCat[1] augmentation. Other augmentations we used were brain-size scaling and random noise. We used Brain Flipping during inference as TTA.</p>\n<p>CropCat</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb6df8385512f3fa34cc45942a7cd3977%2Fhms-CropCat.jpg?generation=1712622065747848&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 75%\"></p>\n<p><strong>Training Data</strong></p>\n<p>Like other competitors, we found that training on data with &gt;= 10 votes resulted in better LB scores. We used GKF on this subset to create a correlated CV/LB. To create our training dataset, we grouped by <code>eeg_id</code>, selected the rows with the max number of votes, and then removed overlapping segments. We then sampled one data point per eeg_id per epoch. During training, we used weighted kldivergence to emphasize data points with more votes. We also sampled noisy student labels with a probability of 15% to increase model diversity. </p>\n<h2>Final Note</h2>\n<p>Our solution would not have been possible without the informative notebooks and sharing of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Also, we mimicked the code structure of <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>'s solution from the ASL Fingerspelling Competition <a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution\" target=\"_blank\">here</a>. Thank you both for sharing. </p>\n<p>Our training code can be found <a href=\"https://github.com/brendanartley/HMS-Competition\" target=\"_blank\">here</a>. Happy Kaggling!</p>\n<h2>Frameworks</h2>\n<ul>\n<li><a href=\"https://lightning.ai/docs/pytorch/stable/\" target=\"_blank\">Pytorch Lightning</a> (training)</li>\n<li><a href=\"https://wandb.ai/site\" target=\"_blank\">Weights + Biases</a> (logging)</li>\n<li><a href=\"https://huggingface.co/timm\" target=\"_blank\">Timm</a> (backbones)</li>\n<li><a href=\"https://github.com/KinWaiCheuk/nnAudio\" target=\"_blank\">nnAudio</a> (STFTs)</li>\n</ul>\n<h1>Sources</h1>\n<p>[1] CropCat: Data Augmentation for Smoothing the Feature Distribution of EEG Signals. <a href=\"https://arxiv.org/abs/2212.06413\" target=\"_blank\">https://arxiv.org/abs/2212.06413</a></p>",
      "rawMarkdown": "Thanks to Harvard Medical School and Kaggle for hosting this fun competition. The competition was very rewarding for us, as there was so much to learn!\n\nTLDR; @ihebch and I built a multimodal model based on three feature extractors. These are an improved 1D-wavenet, a trainable STFT, and an image model with attention pooling for the competition spectrograms. We used `hgnetv2_b4.ssld_stage2_ft_in1k` and `tf_efficientnetv2_s.in21k_ft_in1k` backbones with different architecture configurations to diversify our ensemble. For more details, keep reading!\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F0bdfe4bb5759dfd538b2e473c9cac962%2Fhms-GeneralArch.jpg?generation=1712621808361384&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n## Fast Experimentation\n\nOne of the most important parts of our pipeline was fast experimentation. This allowed us to try creative and crazy ideas quickly. The key points for us were to use smaller inputs, use smaller backbones, and create spectrograms on GPU using nnAudio. In the last few weeks of the competition, we scaled up from 8x to 24x node differences, increased the size of the model backbones, and concatenated the three feature vectors into one multimodal model.\n\n**Model 1: Improved 1D Wavenet**\n\nOur best single architecture is a wavenet model with some modifications. We add Mish as a third activation function and we use residual connections from the input convolution rather than within each layer. These changes allowed us to use deeper wavenets but we found no significant improvements going beyond 7 blocks. Finally, we added a downsample convolution after the last wave block to significantly reduce vRam usage.\n\nOurs\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F60bcbfa7f8bcde74c05b35c4b81eb392%2Fhms-WaveBlock_Ours.jpg?generation=1712621883006223&alt=media\" alt=\"Cropper\" style=\"max-width: 95%;\">\n\nOriginal\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F151172231e6f88767c2bcd19a5b86e3b%2Fhms-WaveBlock_Original.jpg?generation=1712621909762554&alt=media\" alt=\"Cropper\" style=\"max-width: 95%;\">\n\nEach input sequence is passed through a wavenet with the improved blocks, and the outputs are stacked into a single-channel image. All but two of the output channels are stacked together, with the last two output channels being appended to the end of the image. We did this to create an area where features from all 24 node differences are nearby. The stacked outputs are then passed through a timm backbone to get a feature vector.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F049d909627d392a2b69bc10430ee048e%2Fhms-WaveNet_stack.jpg?generation=1712621986486094&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n**Model 2: STFT**\n\nWe generated STFTs on GPU using nnAudio. This is very fast, and each training run for this part of the model took 3-10 mins (depending on the STFT parameters). After generating the STFT, we restack the frequency bins into 2 columns. Finally, the STFTs for each input are stacked together and passed through an image model to get a feature vector.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F5c61eda5775994d937aca74efb840e19%2Fhms-STFTStack.jpg?generation=1712622013956043&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\nWe are still unsure why the restacking method worked, but we wanted to share some more thoughts on this. We thought that maybe by reducing the height of the STFT, we were bringing information from different nodes closer together. We tested this by changing the order in which we stacked the STFTs, and saw significant changes to our CV. This indicated to us that the model was learning inter-STFT features and that the order we stacked mattered. \n\nBased on these findings, we believed that we were missing some sort of attention mechanism between inputs. We started experimenting with this during the last few days of the competition but unfortunately ran out of time.\n\n**Model 3: Attention Pooling**\n\nThe final feature extractor we used was a timm backbone with attention pooling for the 10-min spectrograms. We found that adding the differences between the spectrograms as new images helped, and applying heavy dropout to the edges of images (0-95%) during training also helped.\n\n## Other Strategies\n\n**Augmentations**\n\nThe most important augmentations for us were vertical brain flipping, horizontal brain flipping, and a variant of the CropCat[1] augmentation. Other augmentations we used were brain-size scaling and random noise. We used Brain Flipping during inference as TTA.\n\nCropCat\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb6df8385512f3fa34cc45942a7cd3977%2Fhms-CropCat.jpg?generation=1712622065747848&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n**Training Data**\n\nLike other competitors, we found that training on data with >= 10 votes resulted in better LB scores. We used GKF on this subset to create a correlated CV/LB. To create our training dataset, we grouped by `eeg_id`, selected the rows with the max number of votes, and then removed overlapping segments. We then sampled one data point per eeg_id per epoch. During training, we used weighted kldivergence to emphasize data points with more votes. We also sampled noisy student labels with a probability of 15% to increase model diversity. \n\n## Final Note\n\nOur solution would not have been possible without the informative notebooks and sharing of @cdeotte. Also, we mimicked the code structure of @christofhenkel's solution from the ASL Fingerspelling Competition [here](https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution). Thank you both for sharing. \n\nOur training code can be found [here](https://github.com/brendanartley/HMS-Competition). Happy Kaggling!\n\n## Frameworks\n\n- [Pytorch Lightning](https://lightning.ai/docs/pytorch/stable/) (training)\n- [Weights + Biases](https://wandb.ai/site) (logging)\n- [Timm](https://huggingface.co/timm) (backbones)\n- [nnAudio](https://github.com/KinWaiCheuk/nnAudio) (STFTs)\n\n# Sources\n\n[1] CropCat: Data Augmentation for Smoothing the Feature Distribution of EEG Signals. https://arxiv.org/abs/2212.06413",
      "votes": null
    },
    {
      "id": "2742534",
      "postDate": "04/09/2024 00:41:25",
      "content": "<p>Thanks for Sharing !</p>",
      "rawMarkdown": "Thanks for Sharing !",
      "votes": null
    },
    {
      "id": "2742550",
      "postDate": "04/09/2024 00:50:26",
      "content": "<p>Great work! Congrats to you both on your new tiers!!!! </p>",
      "rawMarkdown": "Great work! Congrats to you both on your new tiers!!!!",
      "votes": null
    },
    {
      "id": "2742610",
      "postDate": "04/09/2024 01:43:28",
      "content": "<p>What a solid solution! I adopted attention pooling from <a href=\"https://github.com/openai/CLIP/blob/main/clip/model.py#L58\" target=\"_blank\">https://github.com/openai/CLIP/blob/main/clip/model.py#L58</a>, however it didn't work well. May I ask which attention pooling you used <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> </p>",
      "rawMarkdown": "What a solid solution! I adopted attention pooling from https://github.com/openai/CLIP/blob/main/clip/model.py#L58, however it didn't work well. May I ask which attention pooling you used @brendanartley",
      "votes": null
    },
    {
      "id": "2742615",
      "postDate": "04/09/2024 01:49:22",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Nice detail explanation and great achievement and is possible to share the training code.</p>",
      "rawMarkdown": "brendanartley Nice detail explanation and great achievement and is possible to share the training code.",
      "votes": null
    },
    {
      "id": "2742619",
      "postDate": "04/09/2024 01:55:26",
      "content": "<p>Of course! </p>\n<p>Code uploaded <a href=\"https://github.com/brendanartley/HMS-Competition\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Of course! ~~We plan to have our code cleaned and uploaded tomorrow.~~\n\nCode uploaded [here](https://github.com/brendanartley/HMS-Competition).",
      "votes": null
    },
    {
      "id": "2742621",
      "postDate": "04/09/2024 01:58:17",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/huyduong7101\" target=\"_blank\">@huyduong7101</a>. We used attention pooling in the same way as <a href=\"https://github.com/brendanartley/UBCO-Competition/blob/a6bc71bfc057be979866cfffc58b2855bcb7870f/ubco_stage2/mil_model/model.py#L5\" target=\"_blank\">this model</a>.</p>",
      "rawMarkdown": "Thanks @huyduong7101. We used attention pooling in the same way as [this model](https://github.com/brendanartley/UBCO-Competition/blob/a6bc71bfc057be979866cfffc58b2855bcb7870f/ubco_stage2/mil_model/model.py#L5).",
      "votes": null
    },
    {
      "id": "2742790",
      "postDate": "04/09/2024 04:57:09",
      "content": "<p>Thanks for this detailed post about your solution.<br>\nMay I ask for more information about this kind of method? </p>\n<blockquote>\n  <p>We used GKF on this subset to create a correlated CV/LB.</p>\n</blockquote>",
      "rawMarkdown": "Thanks for this detailed post about your solution.\nMay I ask for more information about this kind of method? \n>We used GKF on this subset to create a correlated CV/LB.",
      "votes": null
    },
    {
      "id": "2743496",
      "postDate": "04/09/2024 13:37:25",
      "content": "<p>Sure. When we calculate our CV locally, we used <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html\" target=\"_blank\">StratifiedGroupKFold</a>. This meant that each out of fold training set had a similar distribution, and that the <code>eeg_id</code>s were \"unseen\". The reason we say this is correlated, is because when we improved our CV score, we also improved when submitting to the LB. </p>\n<p>On the other hand, when we trained + evaluated with data that had total_votes &lt; 10, improvements to our CV did not always lead to improvements on the LB. In this case the CV/LB was not correlated.</p>",
      "rawMarkdown": "Sure. When we calculate our CV locally, we used [StratifiedGroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html). This meant that each out of fold training set had a similar distribution, and that the `eeg_id`s were \"unseen\". The reason we say this is correlated, is because when we improved our CV score, we also improved when submitting to the LB. \n\nOn the other hand, when we trained + evaluated with data that had total_votes < 10, improvements to our CV did not always lead to improvements on the LB. In this case the CV/LB was not correlated.",
      "votes": null
    },
    {
      "id": "2746177",
      "postDate": "04/11/2024 05:08:49",
      "content": "<p>Thanks for detailed explanation, it will be very helpful. Congrats to both on your new tiers!!!</p>",
      "rawMarkdown": "Thanks for detailed explanation, it will be very helpful. Congrats to both on your new tiers!!!",
      "votes": null
    },
    {
      "id": "2752363",
      "postDate": "04/14/2024 23:46:46",
      "content": "<p>Hi! Are you sure you used 10min spectrograms instead of 10 sec? Looks like a typo</p>",
      "rawMarkdown": "Hi! Are you sure you used 10min spectrograms instead of 10 sec? Looks like a typo",
      "votes": null
    },
    {
      "id": "2752365",
      "postDate": "04/14/2024 23:52:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/asimandia\" target=\"_blank\">@asimandia</a>, the 10-min spectrogram refers to the spectrograms provided by the competition hosts. See <code>spectrogram_id</code> and <code>spectrogram_sub_id</code> on the data page <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\" target=\"_blank\">here</a>. Hope this helps 🙂</p>",
      "rawMarkdown": "Hi @asimandia, the 10-min spectrogram refers to the spectrograms provided by the competition hosts. See `spectrogram_id` and `spectrogram_sub_id` on the data page [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data). Hope this helps 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2742534,
      "author_name": "yitounian",
      "author_url": "",
      "post_date": "04/09/2024 00:41:25",
      "content": "<p>Thanks for Sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2742550,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "04/09/2024 00:50:26",
      "content": "<p>Great work! Congrats to you both on your new tiers!!!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2742610,
      "author_name": "huyduong7101",
      "author_url": "",
      "post_date": "04/09/2024 01:43:28",
      "content": "<p>What a solid solution! I adopted attention pooling from <a href=\"https://github.com/openai/CLIP/blob/main/clip/model.py#L58\" target=\"_blank\">https://github.com/openai/CLIP/blob/main/clip/model.py#L58</a>, however it didn't work well. May I ask which attention pooling you used <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2742621,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "04/09/2024 01:58:17",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/huyduong7101\" target=\"_blank\">@huyduong7101</a>. We used attention pooling in the same way as <a href=\"https://github.com/brendanartley/UBCO-Competition/blob/a6bc71bfc057be979866cfffc58b2855bcb7870f/ubco_stage2/mil_model/model.py#L5\" target=\"_blank\">this model</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2742615,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "04/09/2024 01:49:22",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Nice detail explanation and great achievement and is possible to share the training code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2742619,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "04/09/2024 01:55:26",
          "content": "<p>Of course! </p>\n<p>Code uploaded <a href=\"https://github.com/brendanartley/HMS-Competition\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2742790,
      "author_name": "lmhongkhnh",
      "author_url": "",
      "post_date": "04/09/2024 04:57:09",
      "content": "<p>Thanks for this detailed post about your solution.<br>\nMay I ask for more information about this kind of method? </p>\n<blockquote>\n  <p>We used GKF on this subset to create a correlated CV/LB.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2743496,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "04/09/2024 13:37:25",
          "content": "<p>Sure. When we calculate our CV locally, we used <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html\" target=\"_blank\">StratifiedGroupKFold</a>. This meant that each out of fold training set had a similar distribution, and that the <code>eeg_id</code>s were \"unseen\". The reason we say this is correlated, is because when we improved our CV score, we also improved when submitting to the LB. </p>\n<p>On the other hand, when we trained + evaluated with data that had total_votes &lt; 10, improvements to our CV did not always lead to improvements on the LB. In this case the CV/LB was not correlated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2746177,
      "author_name": "aadityaporwal",
      "author_url": "",
      "post_date": "04/11/2024 05:08:49",
      "content": "<p>Thanks for detailed explanation, it will be very helpful. Congrats to both on your new tiers!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2752363,
      "author_name": "asimandia",
      "author_url": "",
      "post_date": "04/14/2024 23:46:46",
      "content": "<p>Hi! Are you sure you used 10min spectrograms instead of 10 sec? Looks like a typo</p>",
      "votes": null,
      "replies": [
        {
          "id": 2752365,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "04/14/2024 23:52:33",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/asimandia\" target=\"_blank\">@asimandia</a>, the 10-min spectrogram refers to the spectrograms provided by the competition hosts. See <code>spectrogram_id</code> and <code>spectrogram_sub_id</code> on the data page <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\" target=\"_blank\">here</a>. Hope this helps 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2742505": "Thanks to Harvard Medical School and Kaggle for hosting this fun competition. The competition was very rewarding for us, as there was so much to learn!\n\nTLDR; @ihebch and I built a multimodal model based on three feature extractors. These are an improved 1D-wavenet, a trainable STFT, and an image model with attention pooling for the competition spectrograms. We used `hgnetv2_b4.ssld_stage2_ft_in1k` and `tf_efficientnetv2_s.in21k_ft_in1k` backbones with different architecture configurations to diversify our ensemble. For more details, keep reading!\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F0bdfe4bb5759dfd538b2e473c9cac962%2Fhms-GeneralArch.jpg?generation=1712621808361384&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n## Fast Experimentation\n\nOne of the most important parts of our pipeline was fast experimentation. This allowed us to try creative and crazy ideas quickly. The key points for us were to use smaller inputs, use smaller backbones, and create spectrograms on GPU using nnAudio. In the last few weeks of the competition, we scaled up from 8x to 24x node differences, increased the size of the model backbones, and concatenated the three feature vectors into one multimodal model.\n\n**Model 1: Improved 1D Wavenet**\n\nOur best single architecture is a wavenet model with some modifications. We add Mish as a third activation function and we use residual connections from the input convolution rather than within each layer. These changes allowed us to use deeper wavenets but we found no significant improvements going beyond 7 blocks. Finally, we added a downsample convolution after the last wave block to significantly reduce vRam usage.\n\nOurs\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F60bcbfa7f8bcde74c05b35c4b81eb392%2Fhms-WaveBlock_Ours.jpg?generation=1712621883006223&alt=media\" alt=\"Cropper\" style=\"max-width: 95%;\">\n\nOriginal\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F151172231e6f88767c2bcd19a5b86e3b%2Fhms-WaveBlock_Original.jpg?generation=1712621909762554&alt=media\" alt=\"Cropper\" style=\"max-width: 95%;\">\n\nEach input sequence is passed through a wavenet with the improved blocks, and the outputs are stacked into a single-channel image. All but two of the output channels are stacked together, with the last two output channels being appended to the end of the image. We did this to create an area where features from all 24 node differences are nearby. The stacked outputs are then passed through a timm backbone to get a feature vector.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F049d909627d392a2b69bc10430ee048e%2Fhms-WaveNet_stack.jpg?generation=1712621986486094&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n**Model 2: STFT**\n\nWe generated STFTs on GPU using nnAudio. This is very fast, and each training run for this part of the model took 3-10 mins (depending on the STFT parameters). After generating the STFT, we restack the frequency bins into 2 columns. Finally, the STFTs for each input are stacked together and passed through an image model to get a feature vector.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F5c61eda5775994d937aca74efb840e19%2Fhms-STFTStack.jpg?generation=1712622013956043&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\nWe are still unsure why the restacking method worked, but we wanted to share some more thoughts on this. We thought that maybe by reducing the height of the STFT, we were bringing information from different nodes closer together. We tested this by changing the order in which we stacked the STFTs, and saw significant changes to our CV. This indicated to us that the model was learning inter-STFT features and that the order we stacked mattered. \n\nBased on these findings, we believed that we were missing some sort of attention mechanism between inputs. We started experimenting with this during the last few days of the competition but unfortunately ran out of time.\n\n**Model 3: Attention Pooling**\n\nThe final feature extractor we used was a timm backbone with attention pooling for the 10-min spectrograms. We found that adding the differences between the spectrograms as new images helped, and applying heavy dropout to the edges of images (0-95%) during training also helped.\n\n## Other Strategies\n\n**Augmentations**\n\nThe most important augmentations for us were vertical brain flipping, horizontal brain flipping, and a variant of the CropCat[1] augmentation. Other augmentations we used were brain-size scaling and random noise. We used Brain Flipping during inference as TTA.\n\nCropCat\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb6df8385512f3fa34cc45942a7cd3977%2Fhms-CropCat.jpg?generation=1712622065747848&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n**Training Data**\n\nLike other competitors, we found that training on data with >= 10 votes resulted in better LB scores. We used GKF on this subset to create a correlated CV/LB. To create our training dataset, we grouped by `eeg_id`, selected the rows with the max number of votes, and then removed overlapping segments. We then sampled one data point per eeg_id per epoch. During training, we used weighted kldivergence to emphasize data points with more votes. We also sampled noisy student labels with a probability of 15% to increase model diversity. \n\n## Final Note\n\nOur solution would not have been possible without the informative notebooks and sharing of @cdeotte. Also, we mimicked the code structure of @christofhenkel's solution from the ASL Fingerspelling Competition [here](https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution). Thank you both for sharing. \n\nOur training code can be found [here](https://github.com/brendanartley/HMS-Competition). Happy Kaggling!\n\n## Frameworks\n\n- [Pytorch Lightning](https://lightning.ai/docs/pytorch/stable/) (training)\n- [Weights + Biases](https://wandb.ai/site) (logging)\n- [Timm](https://huggingface.co/timm) (backbones)\n- [nnAudio](https://github.com/KinWaiCheuk/nnAudio) (STFTs)\n\n# Sources\n\n[1] CropCat: Data Augmentation for Smoothing the Feature Distribution of EEG Signals. https://arxiv.org/abs/2212.06413",
    "2742534": "Thanks for Sharing !",
    "2742550": "Great work! Congrats to you both on your new tiers!!!!",
    "2742610": "What a solid solution! I adopted attention pooling from https://github.com/openai/CLIP/blob/main/clip/model.py#L58, however it didn't work well. May I ask which attention pooling you used @brendanartley",
    "2742615": "brendanartley Nice detail explanation and great achievement and is possible to share the training code.",
    "2742619": "Of course! ~~We plan to have our code cleaned and uploaded tomorrow.~~\n\nCode uploaded [here](https://github.com/brendanartley/HMS-Competition).",
    "2742621": "Thanks @huyduong7101. We used attention pooling in the same way as [this model](https://github.com/brendanartley/UBCO-Competition/blob/a6bc71bfc057be979866cfffc58b2855bcb7870f/ubco_stage2/mil_model/model.py#L5).",
    "2742790": "Thanks for this detailed post about your solution.\nMay I ask for more information about this kind of method? \n>We used GKF on this subset to create a correlated CV/LB.",
    "2743496": "Sure. When we calculate our CV locally, we used [StratifiedGroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html). This meant that each out of fold training set had a similar distribution, and that the `eeg_id`s were \"unseen\". The reason we say this is correlated, is because when we improved our CV score, we also improved when submitting to the LB. \n\nOn the other hand, when we trained + evaluated with data that had total_votes < 10, improvements to our CV did not always lead to improvements on the LB. In this case the CV/LB was not correlated.",
    "2746177": "Thanks for detailed explanation, it will be very helpful. Congrats to both on your new tiers!!!",
    "2752363": "Hi! Are you sure you used 10min spectrograms instead of 10 sec? Looks like a typo",
    "2752365": "Hi @asimandia, the 10-min spectrogram refers to the spectrograms provided by the competition hosts. See `spectrogram_id` and `spectrogram_sub_id` on the data page [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data). Hope this helps 🙂"
  },
  "source": "meta"
}