{
  "id": 492471,
  "title": "3rd place solution",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/nvidia-dd-3rd-place-solution",
  "author_name": "",
  "post_date": "2024-04-30T21:34:15.113Z",
  "votes": 133,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about EEG data and how to create strong models for it. Special thanks to <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> who carried most of the workload in the last weeks of the competition, when I was busy with training my first real life neural network. </p>\n<h2>TLDR</h2>\n<p>Our solution is an ensemble of multiple models from two diverse modeling approaches. The first approach is to apply pretrained 2d-CNN architectures to the MelSpectrogram transformation of the data. The second approach uses 1D-Convolutions to encode the raw eeg data before modeling with Squeezeformer blocks. <br>\nKey ingredients for our solution were a robust cross validation based on data selection and creative augmentations. </p>\n<h2>Cross validation and Data filtering</h2>\n<p>Having a solid cross validation was certainly key to doing well in this competition! So we spend a lot of time figuring out a good validation scheme. In general we used a 4fold validation setup, split by patient id. One key ingredient was filtering the val dataset for those rows only having more than 9 votes. The <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf\" target=\"_blank\">SPaRCNet paper</a> most likely explained how the data was created.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F3690bc48fab4ed74631ebd292ad0e421%2FScreenshot%202024-04-09%20at%2018.39.18.png?generation=1712680785774576&amp;alt=media\"></p>\n<p>Not only was the test data filtered by having more than 9 votes, but also the data has been augmented by using the same labels but shifting the eeg data. This means that the given 106k training rows are highly redundant and can/ should be filtered. </p>\n<blockquote>\n  <p>We expanded the high- and low-quality sets of EEG segments by adding additional segments belonging to the<br>\n  same stationary period of the EEG</p>\n</blockquote>\n<p>This basically means the authors applied a shift-like augmentation to the raw eeg data to create more data and used same labels for the newly created data. This explains why there were so many rows for the same EEG_id having the same label. We believe that separating the original data from the augmented data makes the data cleaner and improved our cross validation. So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of <strong>only 6350 rows</strong> which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows. At the end we had a quite good cv/ lb correlation</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fccecc2d899a329f4d796640a966d5f53%2FScreenshot%202024-04-09%20at%2018.47.16.png?generation=1712681252795117&amp;alt=media\"></p>\n<h2>Data sources</h2>\n<p>For 2D models, we had pretrained weights and we observed high data quality was important and we got no benefit from using pseudo labels and data with &lt;8 votes in the label. All models were trained with high quality data. </p>\n<p>Our 1D models were trained from scratch so it helped to also use the low count data. See below in the training procedure how low vote data was used in training. Pseudo labelling was not helpful for us - we tried a lot of things here.  </p>\n<p>In the end we did not use the 10 minute spectrograms. They may have helped slightly in the 2d model in some experiments but not in any major way. </p>\n<h2>Data preprocessing</h2>\n<p>No preprocessing of the data to disk was made, which really sped up experiments and allowed flexibility in approaches. We used torchaudio‘s GPU implementation to create the MelSpectrogram on the fly and normalized the signal. </p>\n<p>The double banana montage, explained by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> was used throughout our solution. As mentioned by others, stacking 16 signals next to one another had limitations in that the each node would interact more strongly in the model with it’s neighbours. We experimented with ordering and found this order worked well, which essentially looks at the left and right side of the brain for each node together, <br>\nFp1&gt;F7 Fp2&gt;F8 F7&gt;T3 F8&gt;T4 T3&gt;T5 T4&gt;T6 T5&gt;O1 T6&gt;O2 Fp1&gt;F3 Fp2&gt;F4 F3&gt;C3 F4&gt;C4 C3&gt;P3 C4&gt;P4 P3&gt;O1 P4&gt;O2</p>\n<p>Scipy implementation of butter bandpass filter was used with an order of 2, and lowpass of between 0 and 1.5 Hz on different models, and a highpass of between 20 and 30 Hz. No notchfilter was used. </p>\n<p>We observed it was important not to normalize the data batch-wise or sample-wise. We saw this in <a href=\"https://www.kaggle.com/medali1992\" target=\"_blank\">@medali1992</a> Resnet 1D GRU implementation -  Not sure if this is where the idea originally came from. Instead, after the butter filter, we used the logic implemented <code>x = x.clip(-1024, 1024) / 32.</code>. </p>\n<h2>Augmentations</h2>\n<p>For 2D models, we had pretrained weights and we observed high data quality was more important - not all augmentations helped. Our 1D models were trained from scratch so heavy augmentations were used. In both models, augmentation was highly effective. </p>\n<p>The following augmentations were used by boths models,</p>\n<ul>\n<li>Overwrite with zeros between 1 to 8 melspec nodes, from the total 16, in 50% of samples.</li>\n<li>Randomly choose a different narrower butter bandpass range in between 1 and 8 melspec nodes, from the total 16, in 20% of samples.</li>\n<li>In 50% of samples, randomly shift the 50s window around the center point by up to 20 seconds. </li>\n</ul>\n<p>In the 1D models only, </p>\n<ul>\n<li>In 50% of samples, left-right flip the signal time wise. </li>\n<li>In 50% of samples, switch the sides of the brain left-right.   </li>\n</ul>\n<p>In the 2D models only, </p>\n<ul>\n<li>In 50% of samples for the zoomed center, using a static 50s window, randomly shift the 10 second center point within the window by up to 5 seconds. </li>\n</ul>\n<h2>Models</h2>\n<h3>Melspec + 2D-CNN backbones</h3>\n<p>Like most other competitors we used pretrained 2D CNN backbones which we fed with melspec transformation of the data. However, we made several improvements to publicly known approaches. The first major boost comes from not combining the data into regions of the EEG electrodes (LL, LP, RP, RL) but creating 16 individual mel spectrograms for the double banana montage. We then concatenated the 16 images into one large image. In the last week we found another big improvement, by “zooming” onto the center 10sec. Zooming was applied with different window and hop lengths to increase diversity.  Mel-Spectrogram transformation are done as part of the model directly on GPU using torchaudios Mel-Spectrogram. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F2088a92e1f8bf11dbff9735ce5c9be94%2FScreenshot%202024-04-09%20at%2018.53.19.png?generation=1712681633116409&amp;alt=media\"></p>\n<h3>1D-CNN + Squeezeformer blocks</h3>\n<p>Our 1d CNN was heavily inspired by the work of Yuri Sun <a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> , and particularly his implementation of a lightweight CNN for seizure detection (<a href=\"https://github.com/ThreePoundUniverse/2023JNE/blob/main/The_proposed_model.py\" target=\"_blank\">code and paper here</a>). As shown on the architecture the model focuses on grouped convolutions timewise and then channel wise. The first two blocks focus time-wise, keeping the channel information independent with grouped convolutions. After pooling the timewise information a channel wise convolution is made to interact across the nodes on the brain, and then another timewise convoltion. Technically we use conv2d, but the effect is like conv1d, only we can convolve over all channels in parallel. The original implementation used a CBAM attention block, we swapped this out and instead used 3 layers of squeezeformer. In addition we added pointwise convolutions before and after the depthwise convolutions which helped stabilite and learn better within the net. <br>\nAs we train from scratch on relatively small data, keeping parameter count low was important. During ASL competition we had spent a lot of time on that, so leveraged our squeezeformer implementation from <a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution/tree/main\" target=\"_blank\">there</a>. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc57c62c4a5e83a98435a20b780b6860%2FScreenshot%202024-04-09%20at%2018.55.46.png?generation=1712681758605758&amp;alt=media\"></p>\n<h2>Training procedure</h2>\n<h3>Melspec + 2d CNN backbones</h3>\n<p>Training parameters used were generally 12 to 16 epochs, cosine learning rate decay with initial lr 0.0012 and batch size 32. Drop path was quite effective and set at 0.2. The main difference in the training was the granularity of the window lengths in the magnified 10 second window, different models were trained with different granularity and a mixnet_l and mixnet_xl backbones. </p>\n<h3>1D CNN + SqueezeFormer blocks</h3>\n<p>Training parameters used were ~32 epochs, cosine learning rate decay with initial lr 0.001 and batch size 64. When we used low vote samples, we increased batch size to 256. We used a hidden dimension of 128 and dropout (attention, ff, head, conv) were all 0.1, going deeper on hidden dimension or more layers did not help. <br>\nWe trained 3 different version of the 1d model with respect to data. </p>\n<ul>\n<li>Only 8+ vote samples.</li>\n<li>3-8 vote count samples, and 8+ vote count samples. In each epoch the 3-8 votes dataset was subsampled; and a different loss was used for each set. We decayed the weight of the 3-8 batch over time, so it started with 0.8 weight and decayed with linear or cosine schedule to ~0.2 weight. This decay weighted loss was added to the loss of the 8+ vote dataset. </li>\n<li>Finally, we added 1-2 votes as a separate decayed weight loss, and trained it in the same manner together with the 3-8 vote set and the 8+ set. The 1-2 vote set was very noisy and thus started with a weight on 0.5 and finished with 0. </li>\n</ul>\n<p>For diversity, we trained some 1d models with butter bandpass order 0 and 1. They were not as strong but added some diversity in the blend. </p>\n<h2>Ensembling + Postprocessing</h2>\n<p>Up until the last day we had been using mean blend of weights, with a higher weighting on 2d models. 15 hours before the end, with the help of GPT4, we created a simple neural net to learn the best blend weight on the 23 base models. 5 of the models were 1D CNN with different training procedures. 18 of the models were Mixnet 2d CNNs, with different backbones, and 10 second magnification windows. <br>\nOn CV the learned blend weights lifted the score but affect on public LB was not so large. <br>\nEarlier, we had been thinking that the distribution of the predictions were different to the distribution seen in the &gt;8 votes portion of the training dataset. So about 10 hours from the end we added a bias term for each of the 6 classes to the nnet. This bias would be added to the logits. CV difference was small, but when the Public LB score came through it jumped us from 9th to 4th position which was exciting. The nnet is quite simple, see below. On private score, over the final days, we had been in the 0.27 LB prize zone with most of our blends even with mean weights. </p>\n<pre><code> (nn.Module):\n     ():\n        (Net, ).__init__()\n        .fc = nn.Linear(n_models, , bias = False) \n        .fc_c = torch.nn.Parameter(torch.zeros()[None,,None])\n     ():\n         . fc(x) + .fc_c\n</code></pre>\n<h2>Supplemental Data</h2>\n<p>No supplemental data was used</p>\n<h2>Ablation study (roughly)</h2>\n<ul>\n<li>Ordering of the 16 nodes signals  -0.02 </li>\n<li>Including low votes samples in 1D model  -0.01</li>\n<li>Augmentations -0.03 or more</li>\n<li>Magnify annotated window -0.006; and more by blending different views</li>\n<li>Remove sample wise and batch wise normalisation -0.02</li>\n<li>FC layer to learn blend weights -0.003</li>\n<li>Postprocessing with bias term -0.001</li>\n</ul>\n<p>Best single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30<br>\nBest single 2D model - 4 folds, 2 seed per fold - CV/Public/Private 0.232 / 0.24 / 0.29<br>\nFinal blend 4 folds, 2 seed, 23 model per fold - CV/Public/Private 0.207 / 0.22 / 0.27</p>\n<h2>Used tools/ repos</h2>\n<p>Pytorch<br>\nTimm, huggingface, albumentations.<br>\n<a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> used a 4090 GPU instance rented from runpod.io<br>\nNeptune.ai was our MLOps stack to track, compare and share models. It was heavily used and allowed easy 4 fold <strong>grouped</strong> view of models and mean of the val scores for all folds - which made tracking a lot easier. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F7ef0101b2bb47fdcc1b99688221b47b9%2FScreenshot%202024-04-09%20at%2019.02.09.png?generation=1712682149182806&amp;alt=media\"></p>\n<p>Thanks for reading. Questions are welcome. </p>\n<p>Edit:<br>\nInference kernel: <a href=\"https://www.kaggle.com/code/darraghdog/3rd-place-solution\" target=\"_blank\">https://www.kaggle.com/code/darraghdog/3rd-place-solution</a> <br>\ngithub: <a href=\"https://github.com/darraghdog/kaggle-hms-3rd-place-solution\" target=\"_blank\">https://github.com/darraghdog/kaggle-hms-3rd-place-solution</a></p>",
  "messages": [
    {
      "id": "2743949",
      "postDate": "04/09/2024 17:09:55",
      "content": "<p>Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about EEG data and how to create strong models for it. Special thanks to <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> who carried most of the workload in the last weeks of the competition, when I was busy with training my first real life neural network. </p>\n<h2>TLDR</h2>\n<p>Our solution is an ensemble of multiple models from two diverse modeling approaches. The first approach is to apply pretrained 2d-CNN architectures to the MelSpectrogram transformation of the data. The second approach uses 1D-Convolutions to encode the raw eeg data before modeling with Squeezeformer blocks. <br>\nKey ingredients for our solution were a robust cross validation based on data selection and creative augmentations. </p>\n<h2>Cross validation and Data filtering</h2>\n<p>Having a solid cross validation was certainly key to doing well in this competition! So we spend a lot of time figuring out a good validation scheme. In general we used a 4fold validation setup, split by patient id. One key ingredient was filtering the val dataset for those rows only having more than 9 votes. The <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf\" target=\"_blank\">SPaRCNet paper</a> most likely explained how the data was created.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F3690bc48fab4ed74631ebd292ad0e421%2FScreenshot%202024-04-09%20at%2018.39.18.png?generation=1712680785774576&amp;alt=media\"></p>\n<p>Not only was the test data filtered by having more than 9 votes, but also the data has been augmented by using the same labels but shifting the eeg data. This means that the given 106k training rows are highly redundant and can/ should be filtered. </p>\n<blockquote>\n  <p>We expanded the high- and low-quality sets of EEG segments by adding additional segments belonging to the<br>\n  same stationary period of the EEG</p>\n</blockquote>\n<p>This basically means the authors applied a shift-like augmentation to the raw eeg data to create more data and used same labels for the newly created data. This explains why there were so many rows for the same EEG_id having the same label. We believe that separating the original data from the augmented data makes the data cleaner and improved our cross validation. So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of <strong>only 6350 rows</strong> which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows. At the end we had a quite good cv/ lb correlation</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fccecc2d899a329f4d796640a966d5f53%2FScreenshot%202024-04-09%20at%2018.47.16.png?generation=1712681252795117&amp;alt=media\"></p>\n<h2>Data sources</h2>\n<p>For 2D models, we had pretrained weights and we observed high data quality was important and we got no benefit from using pseudo labels and data with &lt;8 votes in the label. All models were trained with high quality data. </p>\n<p>Our 1D models were trained from scratch so it helped to also use the low count data. See below in the training procedure how low vote data was used in training. Pseudo labelling was not helpful for us - we tried a lot of things here.  </p>\n<p>In the end we did not use the 10 minute spectrograms. They may have helped slightly in the 2d model in some experiments but not in any major way. </p>\n<h2>Data preprocessing</h2>\n<p>No preprocessing of the data to disk was made, which really sped up experiments and allowed flexibility in approaches. We used torchaudio‘s GPU implementation to create the MelSpectrogram on the fly and normalized the signal. </p>\n<p>The double banana montage, explained by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> was used throughout our solution. As mentioned by others, stacking 16 signals next to one another had limitations in that the each node would interact more strongly in the model with it’s neighbours. We experimented with ordering and found this order worked well, which essentially looks at the left and right side of the brain for each node together, <br>\nFp1&gt;F7 Fp2&gt;F8 F7&gt;T3 F8&gt;T4 T3&gt;T5 T4&gt;T6 T5&gt;O1 T6&gt;O2 Fp1&gt;F3 Fp2&gt;F4 F3&gt;C3 F4&gt;C4 C3&gt;P3 C4&gt;P4 P3&gt;O1 P4&gt;O2</p>\n<p>Scipy implementation of butter bandpass filter was used with an order of 2, and lowpass of between 0 and 1.5 Hz on different models, and a highpass of between 20 and 30 Hz. No notchfilter was used. </p>\n<p>We observed it was important not to normalize the data batch-wise or sample-wise. We saw this in <a href=\"https://www.kaggle.com/medali1992\" target=\"_blank\">@medali1992</a> Resnet 1D GRU implementation -  Not sure if this is where the idea originally came from. Instead, after the butter filter, we used the logic implemented <code>x = x.clip(-1024, 1024) / 32.</code>. </p>\n<h2>Augmentations</h2>\n<p>For 2D models, we had pretrained weights and we observed high data quality was more important - not all augmentations helped. Our 1D models were trained from scratch so heavy augmentations were used. In both models, augmentation was highly effective. </p>\n<p>The following augmentations were used by boths models,</p>\n<ul>\n<li>Overwrite with zeros between 1 to 8 melspec nodes, from the total 16, in 50% of samples.</li>\n<li>Randomly choose a different narrower butter bandpass range in between 1 and 8 melspec nodes, from the total 16, in 20% of samples.</li>\n<li>In 50% of samples, randomly shift the 50s window around the center point by up to 20 seconds. </li>\n</ul>\n<p>In the 1D models only, </p>\n<ul>\n<li>In 50% of samples, left-right flip the signal time wise. </li>\n<li>In 50% of samples, switch the sides of the brain left-right.   </li>\n</ul>\n<p>In the 2D models only, </p>\n<ul>\n<li>In 50% of samples for the zoomed center, using a static 50s window, randomly shift the 10 second center point within the window by up to 5 seconds. </li>\n</ul>\n<h2>Models</h2>\n<h3>Melspec + 2D-CNN backbones</h3>\n<p>Like most other competitors we used pretrained 2D CNN backbones which we fed with melspec transformation of the data. However, we made several improvements to publicly known approaches. The first major boost comes from not combining the data into regions of the EEG electrodes (LL, LP, RP, RL) but creating 16 individual mel spectrograms for the double banana montage. We then concatenated the 16 images into one large image. In the last week we found another big improvement, by “zooming” onto the center 10sec. Zooming was applied with different window and hop lengths to increase diversity.  Mel-Spectrogram transformation are done as part of the model directly on GPU using torchaudios Mel-Spectrogram. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F2088a92e1f8bf11dbff9735ce5c9be94%2FScreenshot%202024-04-09%20at%2018.53.19.png?generation=1712681633116409&amp;alt=media\"></p>\n<h3>1D-CNN + Squeezeformer blocks</h3>\n<p>Our 1d CNN was heavily inspired by the work of Yuri Sun <a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> , and particularly his implementation of a lightweight CNN for seizure detection (<a href=\"https://github.com/ThreePoundUniverse/2023JNE/blob/main/The_proposed_model.py\" target=\"_blank\">code and paper here</a>). As shown on the architecture the model focuses on grouped convolutions timewise and then channel wise. The first two blocks focus time-wise, keeping the channel information independent with grouped convolutions. After pooling the timewise information a channel wise convolution is made to interact across the nodes on the brain, and then another timewise convoltion. Technically we use conv2d, but the effect is like conv1d, only we can convolve over all channels in parallel. The original implementation used a CBAM attention block, we swapped this out and instead used 3 layers of squeezeformer. In addition we added pointwise convolutions before and after the depthwise convolutions which helped stabilite and learn better within the net. <br>\nAs we train from scratch on relatively small data, keeping parameter count low was important. During ASL competition we had spent a lot of time on that, so leveraged our squeezeformer implementation from <a href=\"https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution/tree/main\" target=\"_blank\">there</a>. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc57c62c4a5e83a98435a20b780b6860%2FScreenshot%202024-04-09%20at%2018.55.46.png?generation=1712681758605758&amp;alt=media\"></p>\n<h2>Training procedure</h2>\n<h3>Melspec + 2d CNN backbones</h3>\n<p>Training parameters used were generally 12 to 16 epochs, cosine learning rate decay with initial lr 0.0012 and batch size 32. Drop path was quite effective and set at 0.2. The main difference in the training was the granularity of the window lengths in the magnified 10 second window, different models were trained with different granularity and a mixnet_l and mixnet_xl backbones. </p>\n<h3>1D CNN + SqueezeFormer blocks</h3>\n<p>Training parameters used were ~32 epochs, cosine learning rate decay with initial lr 0.001 and batch size 64. When we used low vote samples, we increased batch size to 256. We used a hidden dimension of 128 and dropout (attention, ff, head, conv) were all 0.1, going deeper on hidden dimension or more layers did not help. <br>\nWe trained 3 different version of the 1d model with respect to data. </p>\n<ul>\n<li>Only 8+ vote samples.</li>\n<li>3-8 vote count samples, and 8+ vote count samples. In each epoch the 3-8 votes dataset was subsampled; and a different loss was used for each set. We decayed the weight of the 3-8 batch over time, so it started with 0.8 weight and decayed with linear or cosine schedule to ~0.2 weight. This decay weighted loss was added to the loss of the 8+ vote dataset. </li>\n<li>Finally, we added 1-2 votes as a separate decayed weight loss, and trained it in the same manner together with the 3-8 vote set and the 8+ set. The 1-2 vote set was very noisy and thus started with a weight on 0.5 and finished with 0. </li>\n</ul>\n<p>For diversity, we trained some 1d models with butter bandpass order 0 and 1. They were not as strong but added some diversity in the blend. </p>\n<h2>Ensembling + Postprocessing</h2>\n<p>Up until the last day we had been using mean blend of weights, with a higher weighting on 2d models. 15 hours before the end, with the help of GPT4, we created a simple neural net to learn the best blend weight on the 23 base models. 5 of the models were 1D CNN with different training procedures. 18 of the models were Mixnet 2d CNNs, with different backbones, and 10 second magnification windows. <br>\nOn CV the learned blend weights lifted the score but affect on public LB was not so large. <br>\nEarlier, we had been thinking that the distribution of the predictions were different to the distribution seen in the &gt;8 votes portion of the training dataset. So about 10 hours from the end we added a bias term for each of the 6 classes to the nnet. This bias would be added to the logits. CV difference was small, but when the Public LB score came through it jumped us from 9th to 4th position which was exciting. The nnet is quite simple, see below. On private score, over the final days, we had been in the 0.27 LB prize zone with most of our blends even with mean weights. </p>\n<pre><code> (nn.Module):\n     ():\n        (Net, ).__init__()\n        .fc = nn.Linear(n_models, , bias = False) \n        .fc_c = torch.nn.Parameter(torch.zeros()[None,,None])\n     ():\n         . fc(x) + .fc_c\n</code></pre>\n<h2>Supplemental Data</h2>\n<p>No supplemental data was used</p>\n<h2>Ablation study (roughly)</h2>\n<ul>\n<li>Ordering of the 16 nodes signals  -0.02 </li>\n<li>Including low votes samples in 1D model  -0.01</li>\n<li>Augmentations -0.03 or more</li>\n<li>Magnify annotated window -0.006; and more by blending different views</li>\n<li>Remove sample wise and batch wise normalisation -0.02</li>\n<li>FC layer to learn blend weights -0.003</li>\n<li>Postprocessing with bias term -0.001</li>\n</ul>\n<p>Best single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30<br>\nBest single 2D model - 4 folds, 2 seed per fold - CV/Public/Private 0.232 / 0.24 / 0.29<br>\nFinal blend 4 folds, 2 seed, 23 model per fold - CV/Public/Private 0.207 / 0.22 / 0.27</p>\n<h2>Used tools/ repos</h2>\n<p>Pytorch<br>\nTimm, huggingface, albumentations.<br>\n<a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> used a 4090 GPU instance rented from runpod.io<br>\nNeptune.ai was our MLOps stack to track, compare and share models. It was heavily used and allowed easy 4 fold <strong>grouped</strong> view of models and mean of the val scores for all folds - which made tracking a lot easier. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F7ef0101b2bb47fdcc1b99688221b47b9%2FScreenshot%202024-04-09%20at%2019.02.09.png?generation=1712682149182806&amp;alt=media\"></p>\n<p>Thanks for reading. Questions are welcome. </p>\n<p>Edit:<br>\nInference kernel: <a href=\"https://www.kaggle.com/code/darraghdog/3rd-place-solution\" target=\"_blank\">https://www.kaggle.com/code/darraghdog/3rd-place-solution</a> <br>\ngithub: <a href=\"https://github.com/darraghdog/kaggle-hms-3rd-place-solution\" target=\"_blank\">https://github.com/darraghdog/kaggle-hms-3rd-place-solution</a></p>",
      "rawMarkdown": "Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about EEG data and how to create strong models for it. Special thanks to @darraghdog who carried most of the workload in the last weeks of the competition, when I was busy with training my first real life neural network. \n\n## TLDR\n\nOur solution is an ensemble of multiple models from two diverse modeling approaches. The first approach is to apply pretrained 2d-CNN architectures to the MelSpectrogram transformation of the data. The second approach uses 1D-Convolutions to encode the raw eeg data before modeling with Squeezeformer blocks. \nKey ingredients for our solution were a robust cross validation based on data selection and creative augmentations. \n\n## Cross validation and Data filtering\n\nHaving a solid cross validation was certainly key to doing well in this competition! So we spend a lot of time figuring out a good validation scheme. In general we used a 4fold validation setup, split by patient id. One key ingredient was filtering the val dataset for those rows only having more than 9 votes. The [SPaRCNet paper](https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf) most likely explained how the data was created.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F3690bc48fab4ed74631ebd292ad0e421%2FScreenshot%202024-04-09%20at%2018.39.18.png?generation=1712680785774576&alt=media)\n\nNot only was the test data filtered by having more than 9 votes, but also the data has been augmented by using the same labels but shifting the eeg data. This means that the given 106k training rows are highly redundant and can/ should be filtered. \n\n> We expanded the high- and low-quality sets of EEG segments by adding additional segments belonging to the\nsame stationary period of the EEG\n\nThis basically means the authors applied a shift-like augmentation to the raw eeg data to create more data and used same labels for the newly created data. This explains why there were so many rows for the same EEG_id having the same label. We believe that separating the original data from the augmented data makes the data cleaner and improved our cross validation. So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of **only 6350 rows** which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows. At the end we had a quite good cv/ lb correlation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fccecc2d899a329f4d796640a966d5f53%2FScreenshot%202024-04-09%20at%2018.47.16.png?generation=1712681252795117&alt=media)\n\n## Data sources\nFor 2D models, we had pretrained weights and we observed high data quality was important and we got no benefit from using pseudo labels and data with <8 votes in the label. All models were trained with high quality data. \n\nOur 1D models were trained from scratch so it helped to also use the low count data. See below in the training procedure how low vote data was used in training. Pseudo labelling was not helpful for us - we tried a lot of things here.  \n\nIn the end we did not use the 10 minute spectrograms. They may have helped slightly in the 2d model in some experiments but not in any major way. \n\n## Data preprocessing\nNo preprocessing of the data to disk was made, which really sped up experiments and allowed flexibility in approaches. We used torchaudio‘s GPU implementation to create the MelSpectrogram on the fly and normalized the signal. \n\nThe double banana montage, explained by @cdeotte was used throughout our solution. As mentioned by others, stacking 16 signals next to one another had limitations in that the each node would interact more strongly in the model with it’s neighbours. We experimented with ordering and found this order worked well, which essentially looks at the left and right side of the brain for each node together, \nFp1>F7 Fp2>F8 F7>T3 F8>T4 T3>T5 T4>T6 T5>O1 T6>O2 Fp1>F3 Fp2>F4 F3>C3 F4>C4 C3>P3 C4>P4 P3>O1 P4>O2\n\nScipy implementation of butter bandpass filter was used with an order of 2, and lowpass of between 0 and 1.5 Hz on different models, and a highpass of between 20 and 30 Hz. No notchfilter was used. \n \nWe observed it was important not to normalize the data batch-wise or sample-wise. We saw this in @medali1992 Resnet 1D GRU implementation -  Not sure if this is where the idea originally came from. Instead, after the butter filter, we used the logic implemented `x = x.clip(-1024, 1024) / 32.`. \n\n## Augmentations\nFor 2D models, we had pretrained weights and we observed high data quality was more important - not all augmentations helped. Our 1D models were trained from scratch so heavy augmentations were used. In both models, augmentation was highly effective. \n\nThe following augmentations were used by boths models,\n- Overwrite with zeros between 1 to 8 melspec nodes, from the total 16, in 50% of samples.\n- Randomly choose a different narrower butter bandpass range in between 1 and 8 melspec nodes, from the total 16, in 20% of samples.\n- In 50% of samples, randomly shift the 50s window around the center point by up to 20 seconds. \n\nIn the 1D models only, \n\n- In 50% of samples, left-right flip the signal time wise. \n- In 50% of samples, switch the sides of the brain left-right.   \n\nIn the 2D models only, \n\n- In 50% of samples for the zoomed center, using a static 50s window, randomly shift the 10 second center point within the window by up to 5 seconds. \n\n## Models\n\n### Melspec + 2D-CNN backbones\n\nLike most other competitors we used pretrained 2D CNN backbones which we fed with melspec transformation of the data. However, we made several improvements to publicly known approaches. The first major boost comes from not combining the data into regions of the EEG electrodes (LL, LP, RP, RL) but creating 16 individual mel spectrograms for the double banana montage. We then concatenated the 16 images into one large image. In the last week we found another big improvement, by “zooming” onto the center 10sec. Zooming was applied with different window and hop lengths to increase diversity.  Mel-Spectrogram transformation are done as part of the model directly on GPU using torchaudios Mel-Spectrogram. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F2088a92e1f8bf11dbff9735ce5c9be94%2FScreenshot%202024-04-09%20at%2018.53.19.png?generation=1712681633116409&alt=media)\n\n### 1D-CNN + Squeezeformer blocks\n\nOur 1d CNN was heavily inspired by the work of Yuri Sun @sunyuri , and particularly his implementation of a lightweight CNN for seizure detection ([code and paper here](https://github.com/ThreePoundUniverse/2023JNE/blob/main/The_proposed_model.py)). As shown on the architecture the model focuses on grouped convolutions timewise and then channel wise. The first two blocks focus time-wise, keeping the channel information independent with grouped convolutions. After pooling the timewise information a channel wise convolution is made to interact across the nodes on the brain, and then another timewise convoltion. Technically we use conv2d, but the effect is like conv1d, only we can convolve over all channels in parallel. The original implementation used a CBAM attention block, we swapped this out and instead used 3 layers of squeezeformer. In addition we added pointwise convolutions before and after the depthwise convolutions which helped stabilite and learn better within the net. \nAs we train from scratch on relatively small data, keeping parameter count low was important. During ASL competition we had spent a lot of time on that, so leveraged our squeezeformer implementation from [there](https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution/tree/main). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc57c62c4a5e83a98435a20b780b6860%2FScreenshot%202024-04-09%20at%2018.55.46.png?generation=1712681758605758&alt=media)\n\n## Training procedure\n\n### Melspec + 2d CNN backbones\n\nTraining parameters used were generally 12 to 16 epochs, cosine learning rate decay with initial lr 0.0012 and batch size 32. Drop path was quite effective and set at 0.2. The main difference in the training was the granularity of the window lengths in the magnified 10 second window, different models were trained with different granularity and a mixnet_l and mixnet_xl backbones. \n\n### 1D CNN + SqueezeFormer blocks\n\nTraining parameters used were ~32 epochs, cosine learning rate decay with initial lr 0.001 and batch size 64. When we used low vote samples, we increased batch size to 256. We used a hidden dimension of 128 and dropout (attention, ff, head, conv) were all 0.1, going deeper on hidden dimension or more layers did not help. \nWe trained 3 different version of the 1d model with respect to data. \n- Only 8+ vote samples.\n- 3-8 vote count samples, and 8+ vote count samples. In each epoch the 3-8 votes dataset was subsampled; and a different loss was used for each set. We decayed the weight of the 3-8 batch over time, so it started with 0.8 weight and decayed with linear or cosine schedule to ~0.2 weight. This decay weighted loss was added to the loss of the 8+ vote dataset. \n- Finally, we added 1-2 votes as a separate decayed weight loss, and trained it in the same manner together with the 3-8 vote set and the 8+ set. The 1-2 vote set was very noisy and thus started with a weight on 0.5 and finished with 0. \n\nFor diversity, we trained some 1d models with butter bandpass order 0 and 1. They were not as strong but added some diversity in the blend. \n\n## Ensembling + Postprocessing\nUp until the last day we had been using mean blend of weights, with a higher weighting on 2d models. 15 hours before the end, with the help of GPT4, we created a simple neural net to learn the best blend weight on the 23 base models. 5 of the models were 1D CNN with different training procedures. 18 of the models were Mixnet 2d CNNs, with different backbones, and 10 second magnification windows. \nOn CV the learned blend weights lifted the score but affect on public LB was not so large. \nEarlier, we had been thinking that the distribution of the predictions were different to the distribution seen in the >8 votes portion of the training dataset. So about 10 hours from the end we added a bias term for each of the 6 classes to the nnet. This bias would be added to the logits. CV difference was small, but when the Public LB score came through it jumped us from 9th to 4th position which was exciting. The nnet is quite simple, see below. On private score, over the final days, we had been in the 0.27 LB prize zone with most of our blends even with mean weights. \n\n```\nclass Net(nn.Module):\n    def __init__(self, n_models = 23):\n        super(Net, self).__init__()\n        self.fc = nn.Linear(n_models, 1, bias = False) \n        self.fc_c = torch.nn.Parameter(torch.zeros(6)[None,:,None])\n    def forward(self, x):\n        return self. fc(x) + self.fc_c\n```  \n\n## Supplemental Data\nNo supplemental data was used\n\n## Ablation study (roughly)\n\n- Ordering of the 16 nodes signals  -0.02 \n- Including low votes samples in 1D model  -0.01\n- Augmentations -0.03 or more\n- Magnify annotated window -0.006; and more by blending different views\n- Remove sample wise and batch wise normalisation -0.02\n- FC layer to learn blend weights -0.003\n- Postprocessing with bias term -0.001\n\n\nBest single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30\nBest single 2D model - 4 folds, 2 seed per fold - CV/Public/Private 0.232 / 0.24 / 0.29\nFinal blend 4 folds, 2 seed, 23 model per fold - CV/Public/Private 0.207 / 0.22 / 0.27\n\n## Used tools/ repos\nPytorch\nTimm, huggingface, albumentations.\n@darraghdog used a 4090 GPU instance rented from runpod.io\nNeptune.ai was our MLOps stack to track, compare and share models. It was heavily used and allowed easy 4 fold **grouped** view of models and mean of the val scores for all folds - which made tracking a lot easier. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F7ef0101b2bb47fdcc1b99688221b47b9%2FScreenshot%202024-04-09%20at%2019.02.09.png?generation=1712682149182806&alt=media)\n\nThanks for reading. Questions are welcome. \n\nEdit:\nInference kernel: https://www.kaggle.com/code/darraghdog/3rd-place-solution \ngithub: https://github.com/darraghdog/kaggle-hms-3rd-place-solution",
      "votes": null
    },
    {
      "id": "2743973",
      "postDate": "04/09/2024 17:24:49",
      "content": "<p>Congrats! Thanks for sharing. </p>",
      "rawMarkdown": "Congrats! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2743979",
      "postDate": "04/09/2024 17:27:04",
      "content": "<p>Great solution (as always). I'll reread in more details, but first reaction is that you explain why one hing I tried did not work: I tried to sample eeg raw data from other locations than the middle and did not see any significant change.</p>",
      "rawMarkdown": "Great solution (as always). I'll reread in more details, but first reaction is that you explain why one hing I tried did not work: I tried to sample eeg raw data from other locations than the middle and did not see any significant change.",
      "votes": null
    },
    {
      "id": "2744026",
      "postDate": "04/09/2024 18:00:28",
      "content": "<p><code>I was busy with training my first real life neural network.</code><br>\nCongratulations!</p>",
      "rawMarkdown": "`I was busy with training my first real life neural network.`\nCongratulations!",
      "votes": null
    },
    {
      "id": "2744036",
      "postDate": "04/09/2024 18:06:34",
      "content": "<p>congratulations!<br>\nNice work and nice solution writing<br>\nThanks for your sharing</p>",
      "rawMarkdown": "congratulations!\nNice work and nice solution writing\nThanks for your sharing",
      "votes": null
    },
    {
      "id": "2744060",
      "postDate": "04/09/2024 18:25:44",
      "content": "<p>thank you for the explanation post. Amazing work!</p>",
      "rawMarkdown": "thank you for the explanation post. Amazing work!",
      "votes": null
    },
    {
      "id": "2744065",
      "postDate": "04/09/2024 18:27:44",
      "content": "<p>Congrats! Thank you very much for sharing. Amazing work!</p>",
      "rawMarkdown": "Congrats! Thank you very much for sharing. Amazing work!",
      "votes": null
    },
    {
      "id": "2744247",
      "postDate": "04/09/2024 20:06:25",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on achieveing 3rd place.</p>",
      "rawMarkdown": "Congrats @christofhenkel on achieveing 3rd place.",
      "votes": null
    },
    {
      "id": "2744248",
      "postDate": "04/09/2024 20:06:30",
      "content": "<p>Thank you very much for sharing this magnificent work,<br>\nis there any chance for that final code released ?<br>\nregards</p>",
      "rawMarkdown": "Thank you very much for sharing this magnificent work,\nis there any chance for that final code released ?\nregards",
      "votes": null
    },
    {
      "id": "2744312",
      "postDate": "04/09/2024 20:33:31",
      "content": "<p>Thanks for sharing! <br>\nRegarding cv, we thought based on the data section that votes in test rounding from 3 to 20 so our validation was stratifiedgroupkfold stratifying on votes 3-20 and grouping by patients. We optimized this a lot and reached cv 0.45 using 1 stage only and trusted this too much and stopped working on 2 stages entirely. Really sad that all our work went into dust at the end due to wrong cv. <br>\nOne interesting thing we found is how high is the impact of correct normalization. We dropped the normalization-per-image and adopted global normalization with some modifications. That boosted our cv by around 0.1 and public lb by 0.05.</p>",
      "rawMarkdown": "Thanks for sharing! \nRegarding cv, we thought based on the data section that votes in test rounding from 3 to 20 so our validation was stratifiedgroupkfold stratifying on votes 3-20 and grouping by patients. We optimized this a lot and reached cv 0.45 using 1 stage only and trusted this too much and stopped working on 2 stages entirely. Really sad that all our work went into dust at the end due to wrong cv. \nOne interesting thing we found is how high is the impact of correct normalization. We dropped the normalization-per-image and adopted global normalization with some modifications. That boosted our cv by around 0.1 and public lb by 0.05.",
      "votes": null
    },
    {
      "id": "2744327",
      "postDate": "04/09/2024 20:42:25",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> and yes, we had similar experiences on normalisation. Also, the effort you spent will payoff in different ways with experience gained 💪🏻, I always find that the best learning comes when things do not go as well as expected </p>",
      "rawMarkdown": "Thanks @mohammad2012191 and yes, we had similar experiences on normalisation. Also, the effort you spent will payoff in different ways with experience gained 💪🏻, I always find that the best learning comes when things do not go as well as expected",
      "votes": null
    },
    {
      "id": "2744329",
      "postDate": "04/09/2024 20:44:01",
      "content": "<p>A repo with source code will be added in the thread 🧵 in the couple of weeks</p>",
      "rawMarkdown": "A repo with source code will be added in the thread 🧵 in the couple of weeks",
      "votes": null
    },
    {
      "id": "2744395",
      "postDate": "04/09/2024 21:52:52",
      "content": "<p>Lol it took me a sec to get what this meant 😂 </p>",
      "rawMarkdown": "Lol it took me a sec to get what this meant 😂",
      "votes": null
    },
    {
      "id": "2744424",
      "postDate": "04/09/2024 22:48:27",
      "content": "<p>Wow congratulations this is very helpful </p>",
      "rawMarkdown": "Wow congratulations this is very helpful",
      "votes": null
    },
    {
      "id": "2744455",
      "postDate": "04/09/2024 23:26:49",
      "content": "<p>Congratulations Dieter and Darragh</p>\n<blockquote>\n  <p>Best single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30</p>\n</blockquote>\n<p>Wow. That's a very powerful 1D model. Great job!</p>",
      "rawMarkdown": "Congratulations Dieter and Darragh\n>Best single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30\n\nWow. That's a very powerful 1D model. Great job!",
      "votes": null
    },
    {
      "id": "2744526",
      "postDate": "04/10/2024 01:01:08",
      "content": "<p>Congrats! These single models are powerful. Excellent!</p>",
      "rawMarkdown": "Congrats! These single models are powerful. Excellent!",
      "votes": null
    },
    {
      "id": "2744675",
      "postDate": "04/10/2024 04:20:40",
      "content": "<p>Congratulations on 3rd place in this competition. Thanks for sharing the details of your solutions with colorful diagrams. </p>",
      "rawMarkdown": "Congratulations on 3rd place in this competition. Thanks for sharing the details of your solutions with colorful diagrams.",
      "votes": null
    },
    {
      "id": "2744918",
      "postDate": "04/10/2024 07:58:38",
      "content": "<p>Have fun with the little one</p>",
      "rawMarkdown": "Have fun with the little one",
      "votes": null
    },
    {
      "id": "2745291",
      "postDate": "04/10/2024 14:27:48",
      "content": "<p>Congratulations on 3rd place in this competition. It's very nice to sharing with us !</p>",
      "rawMarkdown": "Congratulations on 3rd place in this competition. It's very nice to sharing with us !",
      "votes": null
    },
    {
      "id": "2745364",
      "postDate": "04/10/2024 15:16:42",
      "content": "<p>Congratulations!! 🎉🎉, Love to see Squeezeformer again…</p>",
      "rawMarkdown": "Congratulations!! 🎉🎉, Love to see Squeezeformer again...",
      "votes": null
    },
    {
      "id": "2745791",
      "postDate": "04/10/2024 20:08:38",
      "content": "<p>Congratulations! I am so surprised that you understood that <code>the authors applied a shift-like augmentation to the rEEGeeg data</code> and identified 6350 original data. While the data section of the competition site says</p>\n<p><code>Many of these samples overlapped and have been consolidated.</code></p>\n<p>And I thought that they gathered the 106,800 data (I was honestly surprised that they would collect them experimentally) as the row numbers and merged them.</p>\n<p>Along with the high-quality data being used for the test, I wonder why the completion host would not mention it.</p>",
      "rawMarkdown": "Congratulations! I am so surprised that you understood that `the authors applied a shift-like augmentation to the rEEGeeg data` and identified 6350 original data. While the data section of the competition site says\n\n`Many of these samples overlapped and have been consolidated. `\n\nAnd I thought that they gathered the 106,800 data (I was honestly surprised that they would collect them experimentally) as the row numbers and merged them.\n\nAlong with the high-quality data being used for the test, I wonder why the completion host would not mention it.",
      "votes": null
    },
    {
      "id": "2746009",
      "postDate": "04/11/2024 01:40:44",
      "content": "<p>Congratulations. I especially liked idea of zooming center region of EEG.</p>\n<p>And our team also noticed the trick of processing test data, but only my team mate <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> successfully squeezed gain from that. I should have dig deeper on that.</p>\n<p>Thanks for sharing exciting report.</p>",
      "rawMarkdown": "Congratulations. I especially liked idea of zooming center region of EEG.\n\nAnd our team also noticed the trick of processing test data, but only my team mate @yujiariyasu successfully squeezed gain from that. I should have dig deeper on that.\n\nThanks for sharing exciting report.",
      "votes": null
    },
    {
      "id": "2746279",
      "postDate": "04/11/2024 06:20:51",
      "content": "<p>Congratulations! I have learned a lot.</p>",
      "rawMarkdown": "Congratulations! I have learned a lot.",
      "votes": null
    },
    {
      "id": "2746401",
      "postDate": "04/11/2024 07:43:26",
      "content": "<p>When extra precision was released we saw the difference of the postprocessing bias term was CV/Public/Private 0.001 / 0.004 / 0.003 </p>",
      "rawMarkdown": "When extra precision was released we saw the difference of the postprocessing bias term was CV/Public/Private 0.001 / 0.004 / 0.003",
      "votes": null
    },
    {
      "id": "2746403",
      "postDate": "04/11/2024 07:44:41",
      "content": "<blockquote>\n  <p>So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of only 6350 rows which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows.</p>\n</blockquote>\n<p>Anyway, could you share us more details of reverse engineering <em>true</em> labels?</p>\n<p>Our approach is simple:</p>\n<ol>\n<li>grouping nearby labels into sub-groups which has same 6-dim vote vectors</li>\n<li>take middle point of each sub-groups</li>\n<li>filter only middle point</li>\n</ol>\n<p>Do you use more sophisticated ways?</p>",
      "rawMarkdown": "> So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of only 6350 rows which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows.\n\nAnyway, could you share us more details of reverse engineering *true* labels?\n\nOur approach is simple:\n1. grouping nearby labels into sub-groups which has same 6-dim vote vectors\n2. take middle point of each sub-groups\n3. filter only middle point\n\nDo you use more sophisticated ways?",
      "votes": null
    },
    {
      "id": "2746407",
      "postDate": "04/11/2024 07:45:47",
      "content": "<p>The CBAM work from you &amp; your team really helped a lot. It is a clever way to approach the problem as 1D model. </p>",
      "rawMarkdown": "The CBAM work from you & your team really helped a lot. It is a clever way to approach the problem as 1D model.",
      "votes": null
    },
    {
      "id": "2746447",
      "postDate": "04/11/2024 08:10:01",
      "content": "<p>Here is the main block.</p>\n<pre><code>  eeg_id with same votes\ntrain = train + label_columns](str)(, axis=)\n\n each unique eeg  combination calculate  overlaping time frames of same label and take the one row that has the highest overlap to all \nrows = \n eeg_id_l  (train()):\n    df0 = train==eeg_id_l](drop=True)()\n    offsets = df0(int)\n    x = np(offsets()+)\n     o  offsets:\n        x += \n    best_idx = np(()  o  offsets])\n    rows += ]\n\ndf = pd(rows)\n (although few) eeg have multiple labels, just take the first one - this last part is not essential\ndf2 = df(subset=)()\n</code></pre>",
      "rawMarkdown": "Here is the main block.\n```\n#group all eeg_id with same votes\ntrain['eeg_id_l'] = train[['eeg_id'] + label_columns].astype(str).agg('_'.join, axis=1)\n\n#for each unique eeg label combination calculate all overlaping time frames of same label and take the one row that has the highest overlap to all \nrows = []\nfor eeg_id_l in tqdm(train['eeg_id_l'].unique()):\n    df0 = train[train['eeg_id_l']==eeg_id_l].reset_index(drop=True).copy()\n    offsets = df0['spectrogram_label_offset_seconds'].astype(int).values\n    x = np.zeros(offsets.max()+600)\n    for o in offsets:\n        x[o:o+600] += 1\n    best_idx = np.argmax([x[o:o+600].sum() for o in offsets])\n    rows += [df0.iloc[best_idx]]\n\ndf = pd.DataFrame(rows)\n#some (although few) eeg have multiple labels, just take the first one - this last part is not essential\ndf2 = df.drop_duplicates(subset='eeg_id').copy()\n```",
      "votes": null
    },
    {
      "id": "2746463",
      "postDate": "04/11/2024 08:24:00",
      "content": "<p>Most of the time take middle point of each sub-groups is good enough. But only if there is only one \"true\" point for the eeg_id. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fe5db07f3b4bf72f51f41c9b80dd27ac4%2FScreenshot%202024-04-11%20at%2010.11.20.png?generation=1712823564444253&amp;alt=media\"><br>\nThis would be the resulting <code>x</code> from the block above for a more complex example. As you can see there are 2 local maxima. Those are most likely the offsets of the true datapoints. The plot also shows why in this example middle point is not the optimal way for filtering</p>",
      "rawMarkdown": "Most of the time take middle point of each sub-groups is good enough. But only if there is only one \"true\" point for the eeg_id. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fe5db07f3b4bf72f51f41c9b80dd27ac4%2FScreenshot%202024-04-11%20at%2010.11.20.png?generation=1712823564444253&alt=media)\n\nThis would be the resulting `x` from the block above for a more complex example. As you can see there are 2 local maxima. Those are most likely the offsets of the true datapoints. The plot also shows why in this example middle point is not the optimal way for filtering",
      "votes": null
    },
    {
      "id": "2746561",
      "postDate": "04/11/2024 10:18:02",
      "content": "<p>Thanks. It all make sense.</p>",
      "rawMarkdown": "Thanks. It all make sense.",
      "votes": null
    },
    {
      "id": "2746710",
      "postDate": "04/11/2024 12:16:26",
      "content": "<p>btw, our pp got public: 0.0045 / private: 0.005 gain</p>",
      "rawMarkdown": "btw, our pp got public: 0.0045 / private: 0.005 gain",
      "votes": null
    },
    {
      "id": "2746714",
      "postDate": "04/11/2024 12:18:47",
      "content": "<p>Thanks for sharing exciting report.</p>",
      "rawMarkdown": "Thanks for sharing exciting report.",
      "votes": null
    },
    {
      "id": "2746930",
      "postDate": "04/11/2024 14:44:32",
      "content": "<p>that is a good jump!</p>",
      "rawMarkdown": "that is a good jump!",
      "votes": null
    },
    {
      "id": "2748296",
      "postDate": "04/12/2024 11:07:32",
      "content": "<p>Congrats! This is actually very similar to our solution. I tried melspectrograms but I couldn't make them work. They were way worse compared to cusignal STFT spectrograms. Our best models were MaxViT tiny and ConvNeXt base. I think you can get first place using those models.</p>",
      "rawMarkdown": "Congrats! This is actually very similar to our solution. I tried melspectrograms but I couldn't make them work. They were way worse compared to cusignal STFT spectrograms. Our best models were MaxViT tiny and ConvNeXt base. I think you can get first place using those models.",
      "votes": null
    },
    {
      "id": "2748429",
      "postDate": "04/12/2024 12:44:33",
      "content": "<p>Congratulations on the 3rd place and thanks for sharing!</p>\n<p>Very interesting approach with the split. I spent most of the competition using all data, because I could see no reason for using only samples with votes &gt;=10, even if they worked well in the LB. In the last weeks I finally took time to read some papers, starting with those on Sparcnet. It was an eye opener. It finally made sense to use only the high quality samples. I followed an approach slightly different from yours: </p>\n<ol>\n<li>pick only the eeg_ids whose maximum number of votes &gt;=10</li>\n<li>take all consecutive samples with the same share of votes (and more than 10 votes) and pick the middle of those samples (10m, 50s, 20s, and 10s). I wanted to use the whole range and randomize, but given the kaggle limits on memory and GPU hours, I ended up going with just the middle. Samples with less than 10 votes were ignored.</li>\n</ol>\n<p>I ended up with 6366 samples. I used all for training, applying weights to account for eeg_ids with a very large number of samples, which didn't exist in the test data, and the fact that the share of \"Other\" in the test sample was lower, according to the paper. For validation I picked only one sample per eeg_id: the one with the most votes or the longest in case of a tie.</p>\n<p>What puzzled me is that adding low quality samples with pseudo labels to the training data, I was able to significantly improve the CV (0.23 for single model), but only get a small LB improvement. Moreover, after a certain point, the CV kept going down but the LB would start degrading just slightly. The public notebooks I looked at had a leak (they sampled high quality and low quality separately, thus mixing eegs from the same patient in train and test), but I couldn't find any leak in my case. I still don't know why CV improved so much, but LB didn't. Maybe this approach introduced a leak because of labeler bias. Any thoughts from your experience using pseudo labels for 1D models?</p>",
      "rawMarkdown": "Congratulations on the 3rd place and thanks for sharing!\n\nVery interesting approach with the split. I spent most of the competition using all data, because I could see no reason for using only samples with votes >=10, even if they worked well in the LB. In the last weeks I finally took time to read some papers, starting with those on Sparcnet. It was an eye opener. It finally made sense to use only the high quality samples. I followed an approach slightly different from yours: \n1. pick only the eeg_ids whose maximum number of votes >=10\n2. take all consecutive samples with the same share of votes (and more than 10 votes) and pick the middle of those samples (10m, 50s, 20s, and 10s). I wanted to use the whole range and randomize, but given the kaggle limits on memory and GPU hours, I ended up going with just the middle. Samples with less than 10 votes were ignored.\n\nI ended up with 6366 samples. I used all for training, applying weights to account for eeg_ids with a very large number of samples, which didn't exist in the test data, and the fact that the share of \"Other\" in the test sample was lower, according to the paper. For validation I picked only one sample per eeg_id: the one with the most votes or the longest in case of a tie.\n\nWhat puzzled me is that adding low quality samples with pseudo labels to the training data, I was able to significantly improve the CV (0.23 for single model), but only get a small LB improvement. Moreover, after a certain point, the CV kept going down but the LB would start degrading just slightly. The public notebooks I looked at had a leak (they sampled high quality and low quality separately, thus mixing eegs from the same patient in train and test), but I couldn't find any leak in my case. I still don't know why CV improved so much, but LB didn't. Maybe this approach introduced a leak because of labeler bias. Any thoughts from your experience using pseudo labels for 1D models?",
      "votes": null
    },
    {
      "id": "2749092",
      "postDate": "04/12/2024 20:30:31",
      "content": "<p><a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> thanks for your comments; re: the pseudo, we had a very similar experience - very promising CV score, but then on LB the pseudo trained model did worse than the base model which generated the pseudo. I just checked private LB scores and they were also the same. <br>\nI do not see any convincing reason; one possible idea I thought is that diversity among folds is higher with non-pseudo training, so when averaged the score is better - as opposed to with pseudo, where we may get similar predictions per fold which do not blend as well even though each individual checkpoint is better. I'm not sure if this is the reason though. Perhaps as you say it could have been a leak because of labeler bias. When a lot of effort was spent on it and score didnt improve, we switched to something else - there were a lot of other things to try.<br>\ncongratulations on your result in the competition.</p>",
      "rawMarkdown": "vialactea thanks for your comments; re: the pseudo, we had a very similar experience - very promising CV score, but then on LB the pseudo trained model did worse than the base model which generated the pseudo. I just checked private LB scores and they were also the same. \nI do not see any convincing reason; one possible idea I thought is that diversity among folds is higher with non-pseudo training, so when averaged the score is better - as opposed to with pseudo, where we may get similar predictions per fold which do not blend as well even though each individual checkpoint is better. I'm not sure if this is the reason though. Perhaps as you say it could have been a leak because of labeler bias. When a lot of effort was spent on it and score didnt improve, we switched to something else - there were a lot of other things to try.\ncongratulations on your result in the competition.",
      "votes": null
    },
    {
      "id": "2749106",
      "postDate": "04/12/2024 20:47:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> When you use pseudo are you careful to make 5 sets of pseudo labels? one for each fold?</p>\n<p>For example we use <code>fold 1</code> to pseudo label the <code>vote&lt;10</code> on <code>fold 1</code> train data only (i.e don't use patients  outside for fold 1) and save that as <code>pseudo set 1</code>. Then we use <code>fold 2</code> to pseudo label the <code>vote&lt;10</code> on <code>fold 2</code> and save that  as <code>pseudo set 2</code> etc.</p>\n<p>Then in stage 2 training, we add the <code>pseudo set 1</code> to <code>fold 1</code> during training. Doing this prevents leaks and prevents optimistic CV score.</p>",
      "rawMarkdown": "Hi @vialactea When you use pseudo are you careful to make 5 sets of pseudo labels? one for each fold?\n\nFor example we use `fold 1` to pseudo label the `vote<10` on `fold 1` train data only (i.e don't use patients  outside for fold 1) and save that as `pseudo set 1`. Then we use `fold 2` to pseudo label the `vote<10` on `fold 2` and save that  as `pseudo set 2` etc.\n\nThen in stage 2 training, we add the `pseudo set 1` to `fold 1` during training. Doing this prevents leaks and prevents optimistic CV score.",
      "votes": null
    },
    {
      "id": "2749205",
      "postDate": "04/12/2024 22:47:25",
      "content": "<p>I experienced the same experience (improve CV drastically and deteriorated LB) with pseudo label. I pseudo label to OOF: first, I trained fold0 train with all data and make pseudo label to fold0 valid set (vote&lt;10). I did this because I believed pseudo label to OOF set is safe since it only leaks features, not labels. </p>\n<p>However, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s approach is safest because there are no leaks of information about validation set including features. Then I wonder what is the mechanism of leaks in pseudo labeling OOF set.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Do you have in mind some particular case of leak when pseudo labeling OOF set?</p>",
      "rawMarkdown": "I experienced the same experience (improve CV drastically and deteriorated LB) with pseudo label. I pseudo label to OOF: first, I trained fold0 train with all data and make pseudo label to fold0 valid set (vote<10). I did this because I believed pseudo label to OOF set is safe since it only leaks features, not labels. \n\nHowever, @cdeotte's approach is safest because there are no leaks of information about validation set including features. Then I wonder what is the mechanism of leaks in pseudo labeling OOF set.\n\n@cdeotte Do you have in mind some particular case of leak when pseudo labeling OOF set?",
      "votes": null
    },
    {
      "id": "2749223",
      "postDate": "04/12/2024 23:13:03",
      "content": "<p>Yes. Each fold has subset of <code>patient_id</code>. Therefore we cannot pseudo label all data. We must use <code>fold 1</code> to only pseudo label the <code>patient_id</code> in fold 1. Otherwise there is a leak.</p>",
      "rawMarkdown": "Yes. Each fold has subset of `patient_id`. Therefore we cannot pseudo label all data. We must use `fold 1` to only pseudo label the `patient_id` in fold 1. Otherwise there is a leak.",
      "votes": null
    },
    {
      "id": "2749553",
      "postDate": "04/13/2024 05:45:43",
      "content": "<p>Congratulations Dieter and Darragh!!!!!!!!</p>",
      "rawMarkdown": "Congratulations Dieter and Darragh!!!!!!!!",
      "votes": null
    },
    {
      "id": "2751855",
      "postDate": "04/14/2024 15:01:38",
      "content": "<p>Seems so 😏 will know next time. We had an idea in the final weeks that our 2d model was weaker but there were so many things to try. Congrats on your great work!</p>",
      "rawMarkdown": "Seems so 😏 will know next time. We had an idea in the final weeks that our 2d model was weaker but there were so many things to try. Congrats on your great work!",
      "votes": null
    },
    {
      "id": "2761071",
      "postDate": "04/19/2024 16:48:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeote\" target=\"_blank\">@cdeote</a>. Yes, I made sure that the pseudo labels for a stage 2 fold came from models trained without any of the patients in that fold. Moreover, all stage 1 and stage 2 models used the exact same 5 folds split of patients, with the caveat that some patients didn't have high quality hard labels and thus were not used in stage 1 training.</p>\n<p>I couldn´t understand how come the stage 2 score for a fold would improve significantly when using pseudo labels produced from models that had been trained without seeing any of the patients in that fold, but that improvement would vanish in the LB. I expected ensembling to be affected due to a potential loss of diversity, as <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> mentioned, but LB of single models didn't improve either.</p>",
      "rawMarkdown": "Hi @cdeote. Yes, I made sure that the pseudo labels for a stage 2 fold came from models trained without any of the patients in that fold. Moreover, all stage 1 and stage 2 models used the exact same 5 folds split of patients, with the caveat that some patients didn't have high quality hard labels and thus were not used in stage 1 training.\n\nI couldn´t understand how come the stage 2 score for a fold would improve significantly when using pseudo labels produced from models that had been trained without seeing any of the patients in that fold, but that improvement would vanish in the LB. I expected ensembling to be affected due to a potential loss of diversity, as @darraghdog mentioned, but LB of single models didn't improve either.",
      "votes": null
    },
    {
      "id": "2764109",
      "postDate": "04/20/2024 23:00:07",
      "content": "<p>I know, right.  I even read the original paper and saw that chart, but discounted it in thinking \"they surely have given us original data and would tell us if anything was augmentation.\"</p>",
      "rawMarkdown": "I know, right.  I even read the original paper and saw that chart, but discounted it in thinking \"they surely have given us original data and would tell us if anything was augmentation.\"",
      "votes": null
    },
    {
      "id": "2799058",
      "postDate": "05/07/2024 15:17:56",
      "content": "<p>Congratulations on 3rd place! I'm sorry for the late question.</p>\n<p>Currently, I'm tracing your released code on github. And there's one thing I can't understand. I notice that you use <strong>masking</strong> in each <code>SqueezeformerBlock</code>. I categorize these masking operations into 3 types,</p>\n<ol>\n<li><code>mask_pad</code> for <code>LlamaAttention</code>.</li>\n<li><code>mask</code> for convolution module.</li>\n<li><code>mask_flat</code> for those commented with <code>Skip/Unskip pad</code>.</li>\n</ol>\n<p>Could you provide some hints about the intuition behind these design choices? Thanks a lot!</p>",
      "rawMarkdown": "Congratulations on 3rd place! I'm sorry for the late question.\n\nCurrently, I'm tracing your released code on github. And there's one thing I can't understand. I notice that you use **masking** in each `SqueezeformerBlock`. I categorize these masking operations into 3 types,\n1. `mask_pad` for `LlamaAttention`.\n2. `mask` for convolution module.\n3. `mask_flat` for those commented with `Skip/Unskip pad`.\n\nCould you provide some hints about the intuition behind these design choices? Thanks a lot!",
      "votes": null
    },
    {
      "id": "2807608",
      "postDate": "05/11/2024 18:28:11",
      "content": "<p>The masking is not used as our inputs have constant length. The mask was all 1, so just a dummy. We leveraged from a previous competition and <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">discussed</a> how it is used there. </p>",
      "rawMarkdown": "The masking is not used as our inputs have constant length. The mask was all 1, so just a dummy. We leveraged from a previous competition and [discussed](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485) how it is used there.",
      "votes": null
    },
    {
      "id": "2807992",
      "postDate": "05/12/2024 03:25:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>,</p>\n<p>Thanks for your explanation. I'll check how it's used in the fingerspelling competition!</p>",
      "rawMarkdown": "Hi @darraghdog,\n\nThanks for your explanation. I'll check how it's used in the fingerspelling competition!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2743973,
      "author_name": "haohantsao",
      "author_url": "",
      "post_date": "04/09/2024 17:24:49",
      "content": "<p>Congrats! Thanks for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2743979,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/09/2024 17:27:04",
      "content": "<p>Great solution (as always). I'll reread in more details, but first reaction is that you explain why one hing I tried did not work: I tried to sample eeg raw data from other locations than the middle and did not see any significant change.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744026,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "04/09/2024 18:00:28",
      "content": "<p><code>I was busy with training my first real life neural network.</code><br>\nCongratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2744395,
          "author_name": "kylekylekyle",
          "author_url": "",
          "post_date": "04/09/2024 21:52:52",
          "content": "<p>Lol it took me a sec to get what this meant 😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2744918,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "04/10/2024 07:58:38",
          "content": "<p>Have fun with the little one</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2744036,
      "author_name": "horikitasaku",
      "author_url": "",
      "post_date": "04/09/2024 18:06:34",
      "content": "<p>congratulations!<br>\nNice work and nice solution writing<br>\nThanks for your sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744060,
      "author_name": "djtrainwreckx",
      "author_url": "",
      "post_date": "04/09/2024 18:25:44",
      "content": "<p>thank you for the explanation post. Amazing work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744065,
      "author_name": "zlemglsmklkaya",
      "author_url": "",
      "post_date": "04/09/2024 18:27:44",
      "content": "<p>Congrats! Thank you very much for sharing. Amazing work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744247,
      "author_name": "adnanalaref",
      "author_url": "",
      "post_date": "04/09/2024 20:06:25",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on achieveing 3rd place.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744248,
      "author_name": "letemoin",
      "author_url": "",
      "post_date": "04/09/2024 20:06:30",
      "content": "<p>Thank you very much for sharing this magnificent work,<br>\nis there any chance for that final code released ?<br>\nregards</p>",
      "votes": null,
      "replies": [
        {
          "id": 2744329,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "04/09/2024 20:44:01",
          "content": "<p>A repo with source code will be added in the thread 🧵 in the couple of weeks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2744312,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "04/09/2024 20:33:31",
      "content": "<p>Thanks for sharing! <br>\nRegarding cv, we thought based on the data section that votes in test rounding from 3 to 20 so our validation was stratifiedgroupkfold stratifying on votes 3-20 and grouping by patients. We optimized this a lot and reached cv 0.45 using 1 stage only and trusted this too much and stopped working on 2 stages entirely. Really sad that all our work went into dust at the end due to wrong cv. <br>\nOne interesting thing we found is how high is the impact of correct normalization. We dropped the normalization-per-image and adopted global normalization with some modifications. That boosted our cv by around 0.1 and public lb by 0.05.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2744327,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "04/09/2024 20:42:25",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> and yes, we had similar experiences on normalisation. Also, the effort you spent will payoff in different ways with experience gained 💪🏻, I always find that the best learning comes when things do not go as well as expected </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2744424,
      "author_name": "rashidrk",
      "author_url": "",
      "post_date": "04/09/2024 22:48:27",
      "content": "<p>Wow congratulations this is very helpful </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744455,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/09/2024 23:26:49",
      "content": "<p>Congratulations Dieter and Darragh</p>\n<blockquote>\n  <p>Best single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30</p>\n</blockquote>\n<p>Wow. That's a very powerful 1D model. Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744526,
      "author_name": "zijiangyang1116",
      "author_url": "",
      "post_date": "04/10/2024 01:01:08",
      "content": "<p>Congrats! These single models are powerful. Excellent!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744675,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "04/10/2024 04:20:40",
      "content": "<p>Congratulations on 3rd place in this competition. Thanks for sharing the details of your solutions with colorful diagrams. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2745291,
      "author_name": "zephyrus1",
      "author_url": "",
      "post_date": "04/10/2024 14:27:48",
      "content": "<p>Congratulations on 3rd place in this competition. It's very nice to sharing with us !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2745364,
      "author_name": "gowrishankarp",
      "author_url": "",
      "post_date": "04/10/2024 15:16:42",
      "content": "<p>Congratulations!! 🎉🎉, Love to see Squeezeformer again…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2745791,
      "author_name": "makio323",
      "author_url": "",
      "post_date": "04/10/2024 20:08:38",
      "content": "<p>Congratulations! I am so surprised that you understood that <code>the authors applied a shift-like augmentation to the rEEGeeg data</code> and identified 6350 original data. While the data section of the competition site says</p>\n<p><code>Many of these samples overlapped and have been consolidated.</code></p>\n<p>And I thought that they gathered the 106,800 data (I was honestly surprised that they would collect them experimentally) as the row numbers and merged them.</p>\n<p>Along with the high-quality data being used for the test, I wonder why the completion host would not mention it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2764109,
          "author_name": "idithhaber",
          "author_url": "",
          "post_date": "04/20/2024 23:00:07",
          "content": "<p>I know, right.  I even read the original paper and saw that chart, but discounted it in thinking \"they surely have given us original data and would tell us if anything was augmentation.\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2746009,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "04/11/2024 01:40:44",
      "content": "<p>Congratulations. I especially liked idea of zooming center region of EEG.</p>\n<p>And our team also noticed the trick of processing test data, but only my team mate <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> successfully squeezed gain from that. I should have dig deeper on that.</p>\n<p>Thanks for sharing exciting report.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2746401,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "04/11/2024 07:43:26",
          "content": "<p>When extra precision was released we saw the difference of the postprocessing bias term was CV/Public/Private 0.001 / 0.004 / 0.003 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2746403,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "04/11/2024 07:44:41",
          "content": "<blockquote>\n  <p>So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of only 6350 rows which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows.</p>\n</blockquote>\n<p>Anyway, could you share us more details of reverse engineering <em>true</em> labels?</p>\n<p>Our approach is simple:</p>\n<ol>\n<li>grouping nearby labels into sub-groups which has same 6-dim vote vectors</li>\n<li>take middle point of each sub-groups</li>\n<li>filter only middle point</li>\n</ol>\n<p>Do you use more sophisticated ways?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2746447,
              "author_name": "darraghdog",
              "author_url": "",
              "post_date": "04/11/2024 08:10:01",
              "content": "<p>Here is the main block.</p>\n<pre><code>  eeg_id with same votes\ntrain = train + label_columns](str)(, axis=)\n\n each unique eeg  combination calculate  overlaping time frames of same label and take the one row that has the highest overlap to all \nrows = \n eeg_id_l  (train()):\n    df0 = train==eeg_id_l](drop=True)()\n    offsets = df0(int)\n    x = np(offsets()+)\n     o  offsets:\n        x += \n    best_idx = np(()  o  offsets])\n    rows += ]\n\ndf = pd(rows)\n (although few) eeg have multiple labels, just take the first one - this last part is not essential\ndf2 = df(subset=)()\n</code></pre>",
              "votes": null,
              "replies": [
                {
                  "id": 2748429,
                  "author_name": "vialactea",
                  "author_url": "",
                  "post_date": "04/12/2024 12:44:33",
                  "content": "<p>Congratulations on the 3rd place and thanks for sharing!</p>\n<p>Very interesting approach with the split. I spent most of the competition using all data, because I could see no reason for using only samples with votes &gt;=10, even if they worked well in the LB. In the last weeks I finally took time to read some papers, starting with those on Sparcnet. It was an eye opener. It finally made sense to use only the high quality samples. I followed an approach slightly different from yours: </p>\n<ol>\n<li>pick only the eeg_ids whose maximum number of votes &gt;=10</li>\n<li>take all consecutive samples with the same share of votes (and more than 10 votes) and pick the middle of those samples (10m, 50s, 20s, and 10s). I wanted to use the whole range and randomize, but given the kaggle limits on memory and GPU hours, I ended up going with just the middle. Samples with less than 10 votes were ignored.</li>\n</ol>\n<p>I ended up with 6366 samples. I used all for training, applying weights to account for eeg_ids with a very large number of samples, which didn't exist in the test data, and the fact that the share of \"Other\" in the test sample was lower, according to the paper. For validation I picked only one sample per eeg_id: the one with the most votes or the longest in case of a tie.</p>\n<p>What puzzled me is that adding low quality samples with pseudo labels to the training data, I was able to significantly improve the CV (0.23 for single model), but only get a small LB improvement. Moreover, after a certain point, the CV kept going down but the LB would start degrading just slightly. The public notebooks I looked at had a leak (they sampled high quality and low quality separately, thus mixing eegs from the same patient in train and test), but I couldn't find any leak in my case. I still don't know why CV improved so much, but LB didn't. Maybe this approach introduced a leak because of labeler bias. Any thoughts from your experience using pseudo labels for 1D models?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2749092,
                      "author_name": "darraghdog",
                      "author_url": "",
                      "post_date": "04/12/2024 20:30:31",
                      "content": "<p><a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> thanks for your comments; re: the pseudo, we had a very similar experience - very promising CV score, but then on LB the pseudo trained model did worse than the base model which generated the pseudo. I just checked private LB scores and they were also the same. <br>\nI do not see any convincing reason; one possible idea I thought is that diversity among folds is higher with non-pseudo training, so when averaged the score is better - as opposed to with pseudo, where we may get similar predictions per fold which do not blend as well even though each individual checkpoint is better. I'm not sure if this is the reason though. Perhaps as you say it could have been a leak because of labeler bias. When a lot of effort was spent on it and score didnt improve, we switched to something else - there were a lot of other things to try.<br>\ncongratulations on your result in the competition.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2749106,
                      "author_name": "cdeotte",
                      "author_url": "",
                      "post_date": "04/12/2024 20:47:51",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/vialactea\" target=\"_blank\">@vialactea</a> When you use pseudo are you careful to make 5 sets of pseudo labels? one for each fold?</p>\n<p>For example we use <code>fold 1</code> to pseudo label the <code>vote&lt;10</code> on <code>fold 1</code> train data only (i.e don't use patients  outside for fold 1) and save that as <code>pseudo set 1</code>. Then we use <code>fold 2</code> to pseudo label the <code>vote&lt;10</code> on <code>fold 2</code> and save that  as <code>pseudo set 2</code> etc.</p>\n<p>Then in stage 2 training, we add the <code>pseudo set 1</code> to <code>fold 1</code> during training. Doing this prevents leaks and prevents optimistic CV score.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2749205,
                          "author_name": "tatamikenn",
                          "author_url": "",
                          "post_date": "04/12/2024 22:47:25",
                          "content": "<p>I experienced the same experience (improve CV drastically and deteriorated LB) with pseudo label. I pseudo label to OOF: first, I trained fold0 train with all data and make pseudo label to fold0 valid set (vote&lt;10). I did this because I believed pseudo label to OOF set is safe since it only leaks features, not labels. </p>\n<p>However, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s approach is safest because there are no leaks of information about validation set including features. Then I wonder what is the mechanism of leaks in pseudo labeling OOF set.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Do you have in mind some particular case of leak when pseudo labeling OOF set?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2749223,
                              "author_name": "cdeotte",
                              "author_url": "",
                              "post_date": "04/12/2024 23:13:03",
                              "content": "<p>Yes. Each fold has subset of <code>patient_id</code>. Therefore we cannot pseudo label all data. We must use <code>fold 1</code> to only pseudo label the <code>patient_id</code> in fold 1. Otherwise there is a leak.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        },
                        {
                          "id": 2761071,
                          "author_name": "vialactea",
                          "author_url": "",
                          "post_date": "04/19/2024 16:48:49",
                          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeote\" target=\"_blank\">@cdeote</a>. Yes, I made sure that the pseudo labels for a stage 2 fold came from models trained without any of the patients in that fold. Moreover, all stage 1 and stage 2 models used the exact same 5 folds split of patients, with the caveat that some patients didn't have high quality hard labels and thus were not used in stage 1 training.</p>\n<p>I couldn´t understand how come the stage 2 score for a fold would improve significantly when using pseudo labels produced from models that had been trained without seeing any of the patients in that fold, but that improvement would vanish in the LB. I expected ensembling to be affected due to a potential loss of diversity, as <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> mentioned, but LB of single models didn't improve either.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2746463,
              "author_name": "christofhenkel",
              "author_url": "",
              "post_date": "04/11/2024 08:24:00",
              "content": "<p>Most of the time take middle point of each sub-groups is good enough. But only if there is only one \"true\" point for the eeg_id. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fe5db07f3b4bf72f51f41c9b80dd27ac4%2FScreenshot%202024-04-11%20at%2010.11.20.png?generation=1712823564444253&amp;alt=media\"><br>\nThis would be the resulting <code>x</code> from the block above for a more complex example. As you can see there are 2 local maxima. Those are most likely the offsets of the true datapoints. The plot also shows why in this example middle point is not the optimal way for filtering</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2746561,
                  "author_name": "tatamikenn",
                  "author_url": "",
                  "post_date": "04/11/2024 10:18:02",
                  "content": "<p>Thanks. It all make sense.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2746710,
                      "author_name": "yujiariyasu",
                      "author_url": "",
                      "post_date": "04/11/2024 12:16:26",
                      "content": "<p>btw, our pp got public: 0.0045 / private: 0.005 gain</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2746930,
                          "author_name": "darraghdog",
                          "author_url": "",
                          "post_date": "04/11/2024 14:44:32",
                          "content": "<p>that is a good jump!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2746279,
      "author_name": "sunyuri",
      "author_url": "",
      "post_date": "04/11/2024 06:20:51",
      "content": "<p>Congratulations! I have learned a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2746407,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "04/11/2024 07:45:47",
          "content": "<p>The CBAM work from you &amp; your team really helped a lot. It is a clever way to approach the problem as 1D model. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2746714,
      "author_name": "ganeshtalwar",
      "author_url": "",
      "post_date": "04/11/2024 12:18:47",
      "content": "<p>Thanks for sharing exciting report.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2748296,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "04/12/2024 11:07:32",
      "content": "<p>Congrats! This is actually very similar to our solution. I tried melspectrograms but I couldn't make them work. They were way worse compared to cusignal STFT spectrograms. Our best models were MaxViT tiny and ConvNeXt base. I think you can get first place using those models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2751855,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "04/14/2024 15:01:38",
          "content": "<p>Seems so 😏 will know next time. We had an idea in the final weeks that our 2d model was weaker but there were so many things to try. Congrats on your great work!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2749553,
      "author_name": "aadityaporwal",
      "author_url": "",
      "post_date": "04/13/2024 05:45:43",
      "content": "<p>Congratulations Dieter and Darragh!!!!!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2799058,
      "author_name": "abaojiang",
      "author_url": "",
      "post_date": "05/07/2024 15:17:56",
      "content": "<p>Congratulations on 3rd place! I'm sorry for the late question.</p>\n<p>Currently, I'm tracing your released code on github. And there's one thing I can't understand. I notice that you use <strong>masking</strong> in each <code>SqueezeformerBlock</code>. I categorize these masking operations into 3 types,</p>\n<ol>\n<li><code>mask_pad</code> for <code>LlamaAttention</code>.</li>\n<li><code>mask</code> for convolution module.</li>\n<li><code>mask_flat</code> for those commented with <code>Skip/Unskip pad</code>.</li>\n</ol>\n<p>Could you provide some hints about the intuition behind these design choices? Thanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2807608,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "05/11/2024 18:28:11",
          "content": "<p>The masking is not used as our inputs have constant length. The mask was all 1, so just a dummy. We leveraged from a previous competition and <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">discussed</a> how it is used there. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2807992,
              "author_name": "abaojiang",
              "author_url": "",
              "post_date": "05/12/2024 03:25:53",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>,</p>\n<p>Thanks for your explanation. I'll check how it's used in the fingerspelling competition!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2743949": "Thanks to kaggle and everyone involved for hosting such an interesting competition. We learned a lot about EEG data and how to create strong models for it. Special thanks to @darraghdog who carried most of the workload in the last weeks of the competition, when I was busy with training my first real life neural network. \n\n## TLDR\n\nOur solution is an ensemble of multiple models from two diverse modeling approaches. The first approach is to apply pretrained 2d-CNN architectures to the MelSpectrogram transformation of the data. The second approach uses 1D-Convolutions to encode the raw eeg data before modeling with Squeezeformer blocks. \nKey ingredients for our solution were a robust cross validation based on data selection and creative augmentations. \n\n## Cross validation and Data filtering\n\nHaving a solid cross validation was certainly key to doing well in this competition! So we spend a lot of time figuring out a good validation scheme. In general we used a 4fold validation setup, split by patient id. One key ingredient was filtering the val dataset for those rows only having more than 9 votes. The [SPaRCNet paper](https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf) most likely explained how the data was created.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F3690bc48fab4ed74631ebd292ad0e421%2FScreenshot%202024-04-09%20at%2018.39.18.png?generation=1712680785774576&alt=media)\n\nNot only was the test data filtered by having more than 9 votes, but also the data has been augmented by using the same labels but shifting the eeg data. This means that the given 106k training rows are highly redundant and can/ should be filtered. \n\n> We expanded the high- and low-quality sets of EEG segments by adding additional segments belonging to the\nsame stationary period of the EEG\n\nThis basically means the authors applied a shift-like augmentation to the raw eeg data to create more data and used same labels for the newly created data. This explains why there were so many rows for the same EEG_id having the same label. We believe that separating the original data from the augmented data makes the data cleaner and improved our cross validation. So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of **only 6350 rows** which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows. At the end we had a quite good cv/ lb correlation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fccecc2d899a329f4d796640a966d5f53%2FScreenshot%202024-04-09%20at%2018.47.16.png?generation=1712681252795117&alt=media)\n\n## Data sources\nFor 2D models, we had pretrained weights and we observed high data quality was important and we got no benefit from using pseudo labels and data with <8 votes in the label. All models were trained with high quality data. \n\nOur 1D models were trained from scratch so it helped to also use the low count data. See below in the training procedure how low vote data was used in training. Pseudo labelling was not helpful for us - we tried a lot of things here.  \n\nIn the end we did not use the 10 minute spectrograms. They may have helped slightly in the 2d model in some experiments but not in any major way. \n\n## Data preprocessing\nNo preprocessing of the data to disk was made, which really sped up experiments and allowed flexibility in approaches. We used torchaudio‘s GPU implementation to create the MelSpectrogram on the fly and normalized the signal. \n\nThe double banana montage, explained by @cdeotte was used throughout our solution. As mentioned by others, stacking 16 signals next to one another had limitations in that the each node would interact more strongly in the model with it’s neighbours. We experimented with ordering and found this order worked well, which essentially looks at the left and right side of the brain for each node together, \nFp1>F7 Fp2>F8 F7>T3 F8>T4 T3>T5 T4>T6 T5>O1 T6>O2 Fp1>F3 Fp2>F4 F3>C3 F4>C4 C3>P3 C4>P4 P3>O1 P4>O2\n\nScipy implementation of butter bandpass filter was used with an order of 2, and lowpass of between 0 and 1.5 Hz on different models, and a highpass of between 20 and 30 Hz. No notchfilter was used. \n \nWe observed it was important not to normalize the data batch-wise or sample-wise. We saw this in @medali1992 Resnet 1D GRU implementation -  Not sure if this is where the idea originally came from. Instead, after the butter filter, we used the logic implemented `x = x.clip(-1024, 1024) / 32.`. \n\n## Augmentations\nFor 2D models, we had pretrained weights and we observed high data quality was more important - not all augmentations helped. Our 1D models were trained from scratch so heavy augmentations were used. In both models, augmentation was highly effective. \n\nThe following augmentations were used by boths models,\n- Overwrite with zeros between 1 to 8 melspec nodes, from the total 16, in 50% of samples.\n- Randomly choose a different narrower butter bandpass range in between 1 and 8 melspec nodes, from the total 16, in 20% of samples.\n- In 50% of samples, randomly shift the 50s window around the center point by up to 20 seconds. \n\nIn the 1D models only, \n\n- In 50% of samples, left-right flip the signal time wise. \n- In 50% of samples, switch the sides of the brain left-right.   \n\nIn the 2D models only, \n\n- In 50% of samples for the zoomed center, using a static 50s window, randomly shift the 10 second center point within the window by up to 5 seconds. \n\n## Models\n\n### Melspec + 2D-CNN backbones\n\nLike most other competitors we used pretrained 2D CNN backbones which we fed with melspec transformation of the data. However, we made several improvements to publicly known approaches. The first major boost comes from not combining the data into regions of the EEG electrodes (LL, LP, RP, RL) but creating 16 individual mel spectrograms for the double banana montage. We then concatenated the 16 images into one large image. In the last week we found another big improvement, by “zooming” onto the center 10sec. Zooming was applied with different window and hop lengths to increase diversity.  Mel-Spectrogram transformation are done as part of the model directly on GPU using torchaudios Mel-Spectrogram. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F2088a92e1f8bf11dbff9735ce5c9be94%2FScreenshot%202024-04-09%20at%2018.53.19.png?generation=1712681633116409&alt=media)\n\n### 1D-CNN + Squeezeformer blocks\n\nOur 1d CNN was heavily inspired by the work of Yuri Sun @sunyuri , and particularly his implementation of a lightweight CNN for seizure detection ([code and paper here](https://github.com/ThreePoundUniverse/2023JNE/blob/main/The_proposed_model.py)). As shown on the architecture the model focuses on grouped convolutions timewise and then channel wise. The first two blocks focus time-wise, keeping the channel information independent with grouped convolutions. After pooling the timewise information a channel wise convolution is made to interact across the nodes on the brain, and then another timewise convoltion. Technically we use conv2d, but the effect is like conv1d, only we can convolve over all channels in parallel. The original implementation used a CBAM attention block, we swapped this out and instead used 3 layers of squeezeformer. In addition we added pointwise convolutions before and after the depthwise convolutions which helped stabilite and learn better within the net. \nAs we train from scratch on relatively small data, keeping parameter count low was important. During ASL competition we had spent a lot of time on that, so leveraged our squeezeformer implementation from [there](https://github.com/ChristofHenkel/kaggle-asl-fingerspelling-1st-place-solution/tree/main). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fdc57c62c4a5e83a98435a20b780b6860%2FScreenshot%202024-04-09%20at%2018.55.46.png?generation=1712681758605758&alt=media)\n\n## Training procedure\n\n### Melspec + 2d CNN backbones\n\nTraining parameters used were generally 12 to 16 epochs, cosine learning rate decay with initial lr 0.0012 and batch size 32. Drop path was quite effective and set at 0.2. The main difference in the training was the granularity of the window lengths in the magnified 10 second window, different models were trained with different granularity and a mixnet_l and mixnet_xl backbones. \n\n### 1D CNN + SqueezeFormer blocks\n\nTraining parameters used were ~32 epochs, cosine learning rate decay with initial lr 0.001 and batch size 64. When we used low vote samples, we increased batch size to 256. We used a hidden dimension of 128 and dropout (attention, ff, head, conv) were all 0.1, going deeper on hidden dimension or more layers did not help. \nWe trained 3 different version of the 1d model with respect to data. \n- Only 8+ vote samples.\n- 3-8 vote count samples, and 8+ vote count samples. In each epoch the 3-8 votes dataset was subsampled; and a different loss was used for each set. We decayed the weight of the 3-8 batch over time, so it started with 0.8 weight and decayed with linear or cosine schedule to ~0.2 weight. This decay weighted loss was added to the loss of the 8+ vote dataset. \n- Finally, we added 1-2 votes as a separate decayed weight loss, and trained it in the same manner together with the 3-8 vote set and the 8+ set. The 1-2 vote set was very noisy and thus started with a weight on 0.5 and finished with 0. \n\nFor diversity, we trained some 1d models with butter bandpass order 0 and 1. They were not as strong but added some diversity in the blend. \n\n## Ensembling + Postprocessing\nUp until the last day we had been using mean blend of weights, with a higher weighting on 2d models. 15 hours before the end, with the help of GPT4, we created a simple neural net to learn the best blend weight on the 23 base models. 5 of the models were 1D CNN with different training procedures. 18 of the models were Mixnet 2d CNNs, with different backbones, and 10 second magnification windows. \nOn CV the learned blend weights lifted the score but affect on public LB was not so large. \nEarlier, we had been thinking that the distribution of the predictions were different to the distribution seen in the >8 votes portion of the training dataset. So about 10 hours from the end we added a bias term for each of the 6 classes to the nnet. This bias would be added to the logits. CV difference was small, but when the Public LB score came through it jumped us from 9th to 4th position which was exciting. The nnet is quite simple, see below. On private score, over the final days, we had been in the 0.27 LB prize zone with most of our blends even with mean weights. \n\n```\nclass Net(nn.Module):\n    def __init__(self, n_models = 23):\n        super(Net, self).__init__()\n        self.fc = nn.Linear(n_models, 1, bias = False) \n        self.fc_c = torch.nn.Parameter(torch.zeros(6)[None,:,None])\n    def forward(self, x):\n        return self. fc(x) + self.fc_c\n```  \n\n## Supplemental Data\nNo supplemental data was used\n\n## Ablation study (roughly)\n\n- Ordering of the 16 nodes signals  -0.02 \n- Including low votes samples in 1D model  -0.01\n- Augmentations -0.03 or more\n- Magnify annotated window -0.006; and more by blending different views\n- Remove sample wise and batch wise normalisation -0.02\n- FC layer to learn blend weights -0.003\n- Postprocessing with bias term -0.001\n\n\nBest single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30\nBest single 2D model - 4 folds, 2 seed per fold - CV/Public/Private 0.232 / 0.24 / 0.29\nFinal blend 4 folds, 2 seed, 23 model per fold - CV/Public/Private 0.207 / 0.22 / 0.27\n\n## Used tools/ repos\nPytorch\nTimm, huggingface, albumentations.\n@darraghdog used a 4090 GPU instance rented from runpod.io\nNeptune.ai was our MLOps stack to track, compare and share models. It was heavily used and allowed easy 4 fold **grouped** view of models and mean of the val scores for all folds - which made tracking a lot easier. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F7ef0101b2bb47fdcc1b99688221b47b9%2FScreenshot%202024-04-09%20at%2019.02.09.png?generation=1712682149182806&alt=media)\n\nThanks for reading. Questions are welcome. \n\nEdit:\nInference kernel: https://www.kaggle.com/code/darraghdog/3rd-place-solution \ngithub: https://github.com/darraghdog/kaggle-hms-3rd-place-solution",
    "2743973": "Congrats! Thanks for sharing.",
    "2743979": "Great solution (as always). I'll reread in more details, but first reaction is that you explain why one hing I tried did not work: I tried to sample eeg raw data from other locations than the middle and did not see any significant change.",
    "2744026": "`I was busy with training my first real life neural network.`\nCongratulations!",
    "2744036": "congratulations!\nNice work and nice solution writing\nThanks for your sharing",
    "2744060": "thank you for the explanation post. Amazing work!",
    "2744065": "Congrats! Thank you very much for sharing. Amazing work!",
    "2744247": "Congrats @christofhenkel on achieveing 3rd place.",
    "2744248": "Thank you very much for sharing this magnificent work,\nis there any chance for that final code released ?\nregards",
    "2744312": "Thanks for sharing! \nRegarding cv, we thought based on the data section that votes in test rounding from 3 to 20 so our validation was stratifiedgroupkfold stratifying on votes 3-20 and grouping by patients. We optimized this a lot and reached cv 0.45 using 1 stage only and trusted this too much and stopped working on 2 stages entirely. Really sad that all our work went into dust at the end due to wrong cv. \nOne interesting thing we found is how high is the impact of correct normalization. We dropped the normalization-per-image and adopted global normalization with some modifications. That boosted our cv by around 0.1 and public lb by 0.05.",
    "2744327": "Thanks @mohammad2012191 and yes, we had similar experiences on normalisation. Also, the effort you spent will payoff in different ways with experience gained 💪🏻, I always find that the best learning comes when things do not go as well as expected",
    "2744329": "A repo with source code will be added in the thread 🧵 in the couple of weeks",
    "2744395": "Lol it took me a sec to get what this meant 😂",
    "2744424": "Wow congratulations this is very helpful",
    "2744455": "Congratulations Dieter and Darragh\n>Best single 1D model - 4 folds, 2 seed per fold - CV/Public/Private 0.257 / 0.25 / 0.30\n\nWow. That's a very powerful 1D model. Great job!",
    "2744526": "Congrats! These single models are powerful. Excellent!",
    "2744675": "Congratulations on 3rd place in this competition. Thanks for sharing the details of your solutions with colorful diagrams.",
    "2744918": "Have fun with the little one",
    "2745291": "Congratulations on 3rd place in this competition. It's very nice to sharing with us !",
    "2745364": "Congratulations!! 🎉🎉, Love to see Squeezeformer again...",
    "2745791": "Congratulations! I am so surprised that you understood that `the authors applied a shift-like augmentation to the rEEGeeg data` and identified 6350 original data. While the data section of the competition site says\n\n`Many of these samples overlapped and have been consolidated. `\n\nAnd I thought that they gathered the 106,800 data (I was honestly surprised that they would collect them experimentally) as the row numbers and merged them.\n\nAlong with the high-quality data being used for the test, I wonder why the completion host would not mention it.",
    "2746009": "Congratulations. I especially liked idea of zooming center region of EEG.\n\nAnd our team also noticed the trick of processing test data, but only my team mate @yujiariyasu successfully squeezed gain from that. I should have dig deeper on that.\n\nThanks for sharing exciting report.",
    "2746279": "Congratulations! I have learned a lot.",
    "2746401": "When extra precision was released we saw the difference of the postprocessing bias term was CV/Public/Private 0.001 / 0.004 / 0.003",
    "2746403": "> So we put some effort into reverse-engineering this process and finding the original “true” data point belonging to a given label. So we ended up with a filtered dataset of only 6350 rows which we used not only for validation using our 4fold scheme but also mainly used that for training, and discarded the other 100k rows.\n\nAnyway, could you share us more details of reverse engineering *true* labels?\n\nOur approach is simple:\n1. grouping nearby labels into sub-groups which has same 6-dim vote vectors\n2. take middle point of each sub-groups\n3. filter only middle point\n\nDo you use more sophisticated ways?",
    "2746407": "The CBAM work from you & your team really helped a lot. It is a clever way to approach the problem as 1D model.",
    "2746447": "Here is the main block.\n```\n#group all eeg_id with same votes\ntrain['eeg_id_l'] = train[['eeg_id'] + label_columns].astype(str).agg('_'.join, axis=1)\n\n#for each unique eeg label combination calculate all overlaping time frames of same label and take the one row that has the highest overlap to all \nrows = []\nfor eeg_id_l in tqdm(train['eeg_id_l'].unique()):\n    df0 = train[train['eeg_id_l']==eeg_id_l].reset_index(drop=True).copy()\n    offsets = df0['spectrogram_label_offset_seconds'].astype(int).values\n    x = np.zeros(offsets.max()+600)\n    for o in offsets:\n        x[o:o+600] += 1\n    best_idx = np.argmax([x[o:o+600].sum() for o in offsets])\n    rows += [df0.iloc[best_idx]]\n\ndf = pd.DataFrame(rows)\n#some (although few) eeg have multiple labels, just take the first one - this last part is not essential\ndf2 = df.drop_duplicates(subset='eeg_id').copy()\n```",
    "2746463": "Most of the time take middle point of each sub-groups is good enough. But only if there is only one \"true\" point for the eeg_id. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fe5db07f3b4bf72f51f41c9b80dd27ac4%2FScreenshot%202024-04-11%20at%2010.11.20.png?generation=1712823564444253&alt=media)\n\nThis would be the resulting `x` from the block above for a more complex example. As you can see there are 2 local maxima. Those are most likely the offsets of the true datapoints. The plot also shows why in this example middle point is not the optimal way for filtering",
    "2746561": "Thanks. It all make sense.",
    "2746710": "btw, our pp got public: 0.0045 / private: 0.005 gain",
    "2746714": "Thanks for sharing exciting report.",
    "2746930": "that is a good jump!",
    "2748296": "Congrats! This is actually very similar to our solution. I tried melspectrograms but I couldn't make them work. They were way worse compared to cusignal STFT spectrograms. Our best models were MaxViT tiny and ConvNeXt base. I think you can get first place using those models.",
    "2748429": "Congratulations on the 3rd place and thanks for sharing!\n\nVery interesting approach with the split. I spent most of the competition using all data, because I could see no reason for using only samples with votes >=10, even if they worked well in the LB. In the last weeks I finally took time to read some papers, starting with those on Sparcnet. It was an eye opener. It finally made sense to use only the high quality samples. I followed an approach slightly different from yours: \n1. pick only the eeg_ids whose maximum number of votes >=10\n2. take all consecutive samples with the same share of votes (and more than 10 votes) and pick the middle of those samples (10m, 50s, 20s, and 10s). I wanted to use the whole range and randomize, but given the kaggle limits on memory and GPU hours, I ended up going with just the middle. Samples with less than 10 votes were ignored.\n\nI ended up with 6366 samples. I used all for training, applying weights to account for eeg_ids with a very large number of samples, which didn't exist in the test data, and the fact that the share of \"Other\" in the test sample was lower, according to the paper. For validation I picked only one sample per eeg_id: the one with the most votes or the longest in case of a tie.\n\nWhat puzzled me is that adding low quality samples with pseudo labels to the training data, I was able to significantly improve the CV (0.23 for single model), but only get a small LB improvement. Moreover, after a certain point, the CV kept going down but the LB would start degrading just slightly. The public notebooks I looked at had a leak (they sampled high quality and low quality separately, thus mixing eegs from the same patient in train and test), but I couldn't find any leak in my case. I still don't know why CV improved so much, but LB didn't. Maybe this approach introduced a leak because of labeler bias. Any thoughts from your experience using pseudo labels for 1D models?",
    "2749092": "vialactea thanks for your comments; re: the pseudo, we had a very similar experience - very promising CV score, but then on LB the pseudo trained model did worse than the base model which generated the pseudo. I just checked private LB scores and they were also the same. \nI do not see any convincing reason; one possible idea I thought is that diversity among folds is higher with non-pseudo training, so when averaged the score is better - as opposed to with pseudo, where we may get similar predictions per fold which do not blend as well even though each individual checkpoint is better. I'm not sure if this is the reason though. Perhaps as you say it could have been a leak because of labeler bias. When a lot of effort was spent on it and score didnt improve, we switched to something else - there were a lot of other things to try.\ncongratulations on your result in the competition.",
    "2749106": "Hi @vialactea When you use pseudo are you careful to make 5 sets of pseudo labels? one for each fold?\n\nFor example we use `fold 1` to pseudo label the `vote<10` on `fold 1` train data only (i.e don't use patients  outside for fold 1) and save that as `pseudo set 1`. Then we use `fold 2` to pseudo label the `vote<10` on `fold 2` and save that  as `pseudo set 2` etc.\n\nThen in stage 2 training, we add the `pseudo set 1` to `fold 1` during training. Doing this prevents leaks and prevents optimistic CV score.",
    "2749205": "I experienced the same experience (improve CV drastically and deteriorated LB) with pseudo label. I pseudo label to OOF: first, I trained fold0 train with all data and make pseudo label to fold0 valid set (vote<10). I did this because I believed pseudo label to OOF set is safe since it only leaks features, not labels. \n\nHowever, @cdeotte's approach is safest because there are no leaks of information about validation set including features. Then I wonder what is the mechanism of leaks in pseudo labeling OOF set.\n\n@cdeotte Do you have in mind some particular case of leak when pseudo labeling OOF set?",
    "2749223": "Yes. Each fold has subset of `patient_id`. Therefore we cannot pseudo label all data. We must use `fold 1` to only pseudo label the `patient_id` in fold 1. Otherwise there is a leak.",
    "2749553": "Congratulations Dieter and Darragh!!!!!!!!",
    "2751855": "Seems so 😏 will know next time. We had an idea in the final weeks that our 2d model was weaker but there were so many things to try. Congrats on your great work!",
    "2761071": "Hi @cdeote. Yes, I made sure that the pseudo labels for a stage 2 fold came from models trained without any of the patients in that fold. Moreover, all stage 1 and stage 2 models used the exact same 5 folds split of patients, with the caveat that some patients didn't have high quality hard labels and thus were not used in stage 1 training.\n\nI couldn´t understand how come the stage 2 score for a fold would improve significantly when using pseudo labels produced from models that had been trained without seeing any of the patients in that fold, but that improvement would vanish in the LB. I expected ensembling to be affected due to a potential loss of diversity, as @darraghdog mentioned, but LB of single models didn't improve either.",
    "2764109": "I know, right.  I even read the original paper and saw that chart, but discounted it in thinking \"they surely have given us original data and would tell us if anything was augmentation.\"",
    "2799058": "Congratulations on 3rd place! I'm sorry for the late question.\n\nCurrently, I'm tracing your released code on github. And there's one thing I can't understand. I notice that you use **masking** in each `SqueezeformerBlock`. I categorize these masking operations into 3 types,\n1. `mask_pad` for `LlamaAttention`.\n2. `mask` for convolution module.\n3. `mask_flat` for those commented with `Skip/Unskip pad`.\n\nCould you provide some hints about the intuition behind these design choices? Thanks a lot!",
    "2807608": "The masking is not used as our inputs have constant length. The mask was all 1, so just a dummy. We leveraged from a previous competition and [discussed](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485) how it is used there.",
    "2807992": "Hi @darraghdog,\n\nThanks for your explanation. I'll check how it's used in the fingerspelling competition!"
  },
  "source": "meta"
}