{
  "id": 492560,
  "title": "1st place solution, team Sony",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492560",
  "author_name": "suguuuuu",
  "post_date": "2024-04-10T03:43:11.723000",
  "votes": 114,
  "comment_count": 29,
  "views": 0,
  "content": "<p>I would like to thank the organizers and the community of the competition!<br>\nI'm grateful to the members who joined the team.</p>\n<h1>Code ( updated on 4/23/2024)</h1>\n<ul>\n<li>train : <a href=\"https://www.kaggle.com/code/sugupoko/1st-place-all-train-code/notebook\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/1st-place-all-train-code/notebook</a></li>\n<li>inferenece: <a href=\"https://www.kaggle.com/code/sugupoko/1st-place-hms-inference-code\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/1st-place-hms-inference-code</a></li>\n</ul>\n<h1>Summary</h1>\n<p>Our prediction pipeline is this.<br>\nEach approaches are will added in the comments!</p>\n<ul>\n<li>yamash's part   : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744629\" target=\"_blank\">Link</a></li>\n<li>suguuuuu's part: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628\" target=\"_blank\">Link</a></li>\n<li>kfuji's part         : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744643\" target=\"_blank\">Link</a></li>\n<li>Muku's part       : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744634\" target=\"_blank\">Link</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fdb96007aaf4d5c3278bcda3b17c1bb14%2Fkaggle_overview.jpg?generation=1712721482613740&amp;alt=media\"></p>\n<h1>Team Validation Strategy</h1>\n<p>The verification method was the same for the entire team.</p>\n<ul>\n<li>Validation using Chris's method<ul>\n<li>link : <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52</a></li>\n<li>Normalize by summing labels with eeg_id, crop the middle 10,000 frames of EEG.</li></ul></li>\n<li>Folding strategy : Group K-fold with a 5-fold split.</li>\n<li>vote &gt;=10</li>\n</ul>",
  "messages": [
    {
      "id": 2744626,
      "postDate": "2024-04-10T03:43:11.723Z",
      "content": "<p>I would like to thank the organizers and the community of the competition!<br>\nI'm grateful to the members who joined the team.</p>\n<h1>Code ( updated on 4/23/2024)</h1>\n<ul>\n<li>train : <a href=\"https://www.kaggle.com/code/sugupoko/1st-place-all-train-code/notebook\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/1st-place-all-train-code/notebook</a></li>\n<li>inferenece: <a href=\"https://www.kaggle.com/code/sugupoko/1st-place-hms-inference-code\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/1st-place-hms-inference-code</a></li>\n</ul>\n<h1>Summary</h1>\n<p>Our prediction pipeline is this.<br>\nEach approaches are will added in the comments!</p>\n<ul>\n<li>yamash's part   : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744629\" target=\"_blank\">Link</a></li>\n<li>suguuuuu's part: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628\" target=\"_blank\">Link</a></li>\n<li>kfuji's part         : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744643\" target=\"_blank\">Link</a></li>\n<li>Muku's part       : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744634\" target=\"_blank\">Link</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fdb96007aaf4d5c3278bcda3b17c1bb14%2Fkaggle_overview.jpg?generation=1712721482613740&amp;alt=media\"></p>\n<h1>Team Validation Strategy</h1>\n<p>The verification method was the same for the entire team.</p>\n<ul>\n<li>Validation using Chris's method<ul>\n<li>link : <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52</a></li>\n<li>Normalize by summing labels with eeg_id, crop the middle 10,000 frames of EEG.</li></ul></li>\n<li>Folding strategy : Group K-fold with a 5-fold split.</li>\n<li>vote &gt;=10</li>\n</ul>",
      "rawMarkdown": "I would like to thank the organizers and the community of the competition!\nI'm grateful to the members who joined the team.\n\n# Code ( updated on 4/23/2024)\n- train : https://www.kaggle.com/code/sugupoko/1st-place-all-train-code/notebook\n- inferenece: https://www.kaggle.com/code/sugupoko/1st-place-hms-inference-code\n \n# Summary\nOur prediction pipeline is this.\nEach approaches are will added in the comments!\n - yamash's part   : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744629)\n - suguuuuu's part: [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628)\n - kfuji's part         : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744643)\n - Muku's part       : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744634)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fdb96007aaf4d5c3278bcda3b17c1bb14%2Fkaggle_overview.jpg?generation=1712721482613740&alt=media)\n\n# Team Validation Strategy\nThe verification method was the same for the entire team.\n - Validation using Chris's method\n  - link : https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\n   - Normalize by summing labels with eeg_id, crop the middle 10,000 frames of EEG.\n - Folding strategy : Group K-fold with a 5-fold split.\n - vote >=10\n",
      "votes": 114
    },
    {
      "id": 2744629,
      "postDate": "2024-04-10T03:48:49.807Z",
      "content": "<p>Thank you Kaggle &amp; hosts for organizing an interesting competition.<br>\nThanks also to my teammates who took part with me. It was fun to participate in close discussions.</p>\n<h2>yamash’s part</h2>\n<h3>Model</h3>\n<p>I lined up the raw EEG signal as it was and created a single 2D image, which was then processed with a 2D CNN model.<br>\nIn more detail, 18 signals of bipolar montage were cropped at three different intervals (2000, 5000 and 10000 samples),  resized and concatenated into a single image.<br>\nI trained several models based on this method. Depending on the model, I changed the range of the bandpass filter to create a 2D image.</p>\n<p>This is a very simple method, but the CV was around 0.24.</p>\n<p>Following 5 models are used for final sub:</p>\n<ul>\n<li>4 x convnext atto models with different seed and bandpass filter (CV: 0.2452, 0.2385, 0.2457, 0.2351)</li>\n<li>1 x inception next tiny (CV: 0.2309)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5568744%2Fb0284055fbb92fab6ac8f4a71bdcc343%2Fhms_figure_yamash.png?generation=1712720787699483&amp;alt=media\"></p>\n<h3>Ensemble</h3>\n<p>I used non-negative linear regression for ensemble.<br>\nEach target variable was estimated by non-negative linear regression from the values of the same target variable predicted by each model.</p>\n<p>At first, our ensemble method is the average of all models, or a single weight was given to each model. We decided on using non-negative linear regression  when we realized that the correlation between CV and public lb was maintained even when more overfit to training data.</p>\n<h3>Replacing softmax with entmax</h3>\n<p>The output of softmax is not zero for every target variable, but many training data labels have some zero probability classes. <br>\nTo deal with this difference, I tried to replace softmax with sparsemax[1] and entmax[2]. <br>\nThese functions can output sparser results compared to softmax. <br>\nFinally, we decided to use entmax with a small alpha parameter (~1.03) for all our single models, which improved both public and private LB score around 0.004.<br>\nIn practice, small alpha didn’t output sparse results, but only of making the values a little more crisp.</p>\n<p>I found this almost the end of this competition, so I couldn’t have enough experiments of using entmax and sparsemax in the training stage.</p>\n<p>[1] <a href=\"https://arxiv.org/abs/1602.02068\" target=\"_blank\">https://arxiv.org/abs/1602.02068</a><br>\n[2] <a href=\"https://arxiv.org/abs/1905.05702\" target=\"_blank\">https://arxiv.org/abs/1905.05702</a></p>",
      "rawMarkdown": "Thank you Kaggle & hosts for organizing an interesting competition.\nThanks also to my teammates who took part with me. It was fun to participate in close discussions.\n\n## yamash’s part\n\n\n### Model\n\nI lined up the raw EEG signal as it was and created a single 2D image, which was then processed with a 2D CNN model.\nIn more detail, 18 signals of bipolar montage were cropped at three different intervals (2000, 5000 and 10000 samples),  resized and concatenated into a single image.\nI trained several models based on this method. Depending on the model, I changed the range of the bandpass filter to create a 2D image.\n\nThis is a very simple method, but the CV was around 0.24.\n\nFollowing 5 models are used for final sub:\n- 4 x convnext atto models with different seed and bandpass filter (CV: 0.2452, 0.2385, 0.2457, 0.2351)\n- 1 x inception next tiny (CV: 0.2309)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5568744%2Fb0284055fbb92fab6ac8f4a71bdcc343%2Fhms_figure_yamash.png?generation=1712720787699483&alt=media)\n\n### Ensemble\n\nI used non-negative linear regression for ensemble.\nEach target variable was estimated by non-negative linear regression from the values of the same target variable predicted by each model.\n\nAt first, our ensemble method is the average of all models, or a single weight was given to each model. We decided on using non-negative linear regression  when we realized that the correlation between CV and public lb was maintained even when more overfit to training data.\n\n### Replacing softmax with entmax\n\nThe output of softmax is not zero for every target variable, but many training data labels have some zero probability classes. \nTo deal with this difference, I tried to replace softmax with sparsemax[1] and entmax[2]. \nThese functions can output sparser results compared to softmax. \nFinally, we decided to use entmax with a small alpha parameter (~1.03) for all our single models, which improved both public and private LB score around 0.004.\nIn practice, small alpha didn’t output sparse results, but only of making the values a little more crisp.\n\nI found this almost the end of this competition, so I couldn’t have enough experiments of using entmax and sparsemax in the training stage.\n\n[1] [https://arxiv.org/abs/1602.02068](https://arxiv.org/abs/1602.02068)\n[2] [https://arxiv.org/abs/1905.05702](https://arxiv.org/abs/1905.05702)\n",
      "votes": 24,
      "replies": [
        {
          "id": 2744905,
          "postDate": "2024-04-10T07:40:40.133Z",
          "content": "<p>I really like the entmax idea 💡 thanks for sharing. For such a small change it has a relatively big effect. Are you aware if it gave similar lift in CV and LB?</p>",
          "rawMarkdown": "I really like the entmax idea 💡 thanks for sharing. For such a small change it has a relatively big effect. Are you aware if it gave similar lift in CV and LB?",
          "votes": 1,
          "replies": [
            {
              "id": 2745003,
              "postDate": "2024-04-10T09:40:55.203Z",
              "content": "<p>Thank you for the comment.<br>\nAs for CV, there were some improvements and some not, depending on each single model. And there was no overall improvement in CV. CV was calculated with labels averaged over the same eeg_id, so it is possible that the effect of entmax was less pronounced than with the labels in the actual test data.<br>\nOn the other hand, for public LB, scores improved consistently regardless of the single model combination.</p>",
              "rawMarkdown": "Thank you for the comment.\nAs for CV, there were some improvements and some not, depending on each single model. And there was no overall improvement in CV. CV was calculated with labels averaged over the same eeg_id, so it is possible that the effect of entmax was less pronounced than with the labels in the actual test data.\nOn the other hand, for public LB, scores improved consistently regardless of the single model combination.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2765977,
          "postDate": "2024-04-21T12:39:38.373Z",
          "content": "<p>how can u use different activation softmax and entmax in train and test? </p>",
          "rawMarkdown": "how can u use different activation softmax and entmax in train and test? "
        }
      ]
    },
    {
      "id": 2744634,
      "postDate": "2024-04-10T03:53:57.120Z",
      "content": "<h1>Muku's part</h1>\n<p>First of all, I would like to thank Harvard Medical School and Kaggle team for organizing a great competition, the participants for sharing useful ideas, and my teammates for fighting alongside me.</p>\n<p><br></p>\n<h2>Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F0a9de44bccbe20701429d7762e229940%2Fhms_architecture.png?generation=1750902445619521&amp;alt=media\" alt=\"\"></p>\n<p><br></p>\n<h2>1. Model</h2>\n<ul>\n<li>Based on the 16-channel anterior-posterior montages generated from raw EEG, two types of image representations are generated and input to timm model.<br>\n<br></li>\n<li><strong>First：Temporal feature maps by 1D temporal convolution</strong><ul>\n<li>This idea was inspired by <a href=\"https://arxiv.org/pdf/1611.08024.pdf\" target=\"_blank\">EEGNet paper</a> and <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476\" target=\"_blank\">G2Net Gravitational Wave Detection | Top 1 solution</a>.<ul>\n<li>I expected to obtain filters for the rhythmic temporal features specific to each symptom, by learning.</li></ul></li>\n<li>The kernel size was set to the same value as the sampling rate (200).</li>\n<li>Stack the generated temporal feature maps by montage channels or feature maps to obtain input images for timm mddel.<ul>\n<li>Stacking each montage channels has better accuracy, and parallel use of both methods improves accuracy.</li></ul></li>\n<li>Optional: Temporal feature maps are directly input to the GRU layer, which improves accuracy, so added as a variation of Ensemble.<br>\n<br></li></ul></li>\n<li><strong>Second：Time-Frequency Transform (superlets / STFT)</strong><ul>\n<li>In the early stages of the study, I used STFT, but it did not appear to be a good representation because of the loss of resolution in the time or frequency.</li>\n<li>Then, I use <strong>superlets</strong> as a better representation method. (CWT was being investigated by suguuuuu and kfuji, so I used this as a variation).<ul>\n<li>Use the following repositories published under the MIT License<br>\n<a href=\"https://github.com/irhum/superlets\" target=\"_blank\">https://github.com/irhum/superlets</a></li>\n<li>settings：<ul>\n<li>min_freq, max_freq = 0.5, 20.0</li>\n<li>base_cycle, min_order, max_order = 1, 1, 16</li>\n<li>Adjusted for better resolution in the time direction</li></ul></li></ul></li>\n<li>The superlet is more compatible with frequency/time expressivity compared to STFT, and the representation is more predictable for labels even for humans.<ul>\n<li>Examples：<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F6e0ad7491adf39bb468b27237d3bd7c7%2Fsuperlets_1.png?generation=1712721056221695&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fa40bf80b7b07f12e04bb3b031dfd00a1%2Fsuperlets_2.png?generation=1712721073292828&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F296c1d4610df6d2f526aae619495e29d%2Fsuperlets_3.png?generation=1712721081681421&amp;alt=media\" alt=\"# \"><br>\n<br></li></ul></li></ul></li>\n<li><strong>timm model</strong><ul>\n<li>Try various timm backbones and add them to Ensemble. \nDiverse ensemble was effective because the KL divergence is sensitive to errors for extreme predictions.<ul>\n<li>Overall, smaller models performed well.</li>\n<li>Best backbone：<strong>swinv2_tiny_window16 (CV: 0.2229</strong>)</li>\n<li>The following also showed good accuracy and contributed to Ensemble<ul>\n<li>swinv2_tiny_window8</li>\n<li>caformer_s18</li>\n<li>gcvit_xtiny, xxtiny</li>\n<li>convnextv2_atto</li>\n<li>maxvit_pico</li>\n<li>inception_next_tiny</li>\n<li>poolformerv2_s12</li>\n<li>nextvit_small</li></ul></li></ul></li></ul></li>\n</ul>\n<p><br></p>\n<h2>2. Training</h2>\n<ul>\n<li>2-stage training (Stage1：&gt; 1 vote, Stage2：&gt; 9 votes)<ul>\n<li>Labels of vote: 1 appeared to be unreliable, so excluded from training.<br>\n(I had thought about applying pseudo labeling, but I didn't have enough time)</li></ul></li>\n<li>Learning rate：[Stage1：1e-3, Stage2：1e-4]</li>\n<li>Loss function：KLDivLoss（with aux loss for each model output）</li>\n<li>scheduler：CosineAnnealingLR</li>\n<li>Optimizer：Adan</li>\n<li>Data sampling：Each unique eeg_id &amp; label（Multiple samples may be generated from the same eeg_id）</li>\n<li>Label smoothing：Add offset of 0.02 to each vote before normalization.<ul>\n<li>By adding this before normalization, a stronger regularization is applied to labels with low vote counts, which are considered to have a relatively low confidence.</li></ul></li>\n<li>Augmentation：<ul>\n<li>±5 sec random time-shift</li>\n<li>Random bandpass filter (Only for waves)</li>\n<li>XYMasking（Only for specs）</li></ul></li>\n<li>butter bandpass filter (Only for waves)：Set highcut to 30~40Hz<ul>\n<li>When high-frequency noise exists, \"Others\" votes tended to increase, so these noises are determined to be important for predicition.</li></ul></li>\n</ul>",
      "rawMarkdown": "# Muku's part\n\nFirst of all, I would like to thank Harvard Medical School and Kaggle team for organizing a great competition, the participants for sharing useful ideas, and my teammates for fighting alongside me.\n\n<br>\n\n## Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F0a9de44bccbe20701429d7762e229940%2Fhms_architecture.png?generation=1750902445619521&alt=media)\n\n<br>\n\t\n## 1. Model\n- Based on the 16-channel anterior-posterior montages generated from raw EEG, two types of image representations are generated and input to timm model.\n<br>\n- **First：Temporal feature maps by 1D temporal convolution**\n    - This idea was inspired by [EEGNet paper](https://arxiv.org/pdf/1611.08024.pdf) and [G2Net Gravitational Wave Detection | Top 1 solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476).\n        - I expected to obtain filters for the rhythmic temporal features specific to each symptom, by learning.\n    - The kernel size was set to the same value as the sampling rate (200).\n    - Stack the generated temporal feature maps by montage channels or feature maps to obtain input images for timm mddel.\n        - Stacking each montage channels has better accuracy, and parallel use of both methods improves accuracy.\n    -  Optional: Temporal feature maps are directly input to the GRU layer, which improves accuracy, so added as a variation of Ensemble.\n<br>\n- **Second：Time-Frequency Transform (superlets / STFT)**\n    - In the early stages of the study, I used STFT, but it did not appear to be a good representation because of the loss of resolution in the time or frequency.\n    - Then, I use **superlets** as a better representation method. (CWT was being investigated by suguuuuu and kfuji, so I used this as a variation).\n        - Use the following repositories published under the MIT License\n[https://github.com/irhum/superlets](https://github.com/irhum/superlets)\n        - settings：\n            - min_freq, max_freq = 0.5, 20.0\n            - base_cycle, min_order, max_order = 1, 1, 16\n            - Adjusted for better resolution in the time direction\n    - The superlet is more compatible with frequency/time expressivity compared to STFT, and the representation is more predictable for labels even for humans.\n        - Examples：\n            ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F6e0ad7491adf39bb468b27237d3bd7c7%2Fsuperlets_1.png?generation=1712721056221695&alt=media)\n            ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fa40bf80b7b07f12e04bb3b031dfd00a1%2Fsuperlets_2.png?generation=1712721073292828&alt=media)\n            ![# ](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F296c1d4610df6d2f526aae619495e29d%2Fsuperlets_3.png?generation=1712721081681421&alt=media)\n<br>\n- **timm model**\n    - Try various timm backbones and add them to Ensemble. \nDiverse ensemble was effective because the KL divergence is sensitive to errors for extreme predictions.\n        - Overall, smaller models performed well.\n        - Best backbone：**swinv2_tiny_window16 (CV: 0.2229**)\n        - The following also showed good accuracy and contributed to Ensemble\n            - swinv2_tiny_window8\n            - caformer_s18\n            - gcvit_xtiny, xxtiny\n            - convnextv2_atto\n            - maxvit_pico\n            - inception_next_tiny\n            - poolformerv2_s12\n            - nextvit_small\n\n<br>\n## 2. Training\n- 2-stage training (Stage1：> 1 vote, Stage2：> 9 votes)\n    - Labels of vote: 1 appeared to be unreliable, so excluded from training.\n(I had thought about applying pseudo labeling, but I didn't have enough time)\n- Learning rate：[Stage1：1e-3, Stage2：1e-4]\n- Loss function：KLDivLoss（with aux loss for each model output）\n- scheduler：CosineAnnealingLR\n- Optimizer：Adan\n- Data sampling：Each unique eeg_id & label（Multiple samples may be generated from the same eeg_id）\n- Label smoothing：Add offset of 0.02 to each vote before normalization.\n    - By adding this before normalization, a stronger regularization is applied to labels with low vote counts, which are considered to have a relatively low confidence.\n- Augmentation：\n    - ±5 sec random time-shift\n    - Random bandpass filter (Only for waves)\n    - XYMasking（Only for specs）\n- butter bandpass filter (Only for waves)：Set highcut to 30~40Hz\n    - When high-frequency noise exists, \"Others\" votes tended to increase, so these noises are determined to be important for predicition.",
      "votes": 22,
      "replies": [
        {
          "id": 2744910,
          "postDate": "2024-04-10T07:53:38.050Z",
          "content": "<p>Very clever idea on the label smoothing.. applying before normalisation to smooth low count labels more</p>",
          "rawMarkdown": "Very clever idea on the label smoothing.. applying before normalisation to smooth low count labels more",
          "votes": 2
        },
        {
          "id": 2753536,
          "postDate": "2024-04-15T14:50:31.197Z",
          "content": "<p>Congrats for 1st place!</p>\n<p>I have one question about 2D images created by superlet.</p>\n<p>It seems you made scalograms by setting the time-axis as about 800. <br>\nBut when resizing it to (256,256), the images are totally different from the original one, which would lead to the loss of information in time direction.</p>\n<p>Did you make any treatment against that problem?</p>\n<p>P.S. If possible, I'd like to know the image's shape you made when creating scalogram (before resizing).</p>",
          "rawMarkdown": "Congrats for 1st place!\n\nI have one question about 2D images created by superlet.\n\nIt seems you made scalograms by setting the time-axis as about 800. \nBut when resizing it to (256,256), the images are totally different from the original one, which would lead to the loss of information in time direction.\n\nDid you make any treatment against that problem?\n\nP.S. If possible, I'd like to know the image's shape you made when creating scalogram (before resizing).",
          "votes": 1,
          "replies": [
            {
              "id": 2755239,
              "postDate": "2024-04-16T12:42:06.467Z",
              "content": "<p>Dear MBOOK, </p>\n<p>Thank you for your comment! And sorry for the late reply.<br>\n</p>\n<blockquote>\n  <p>P.S. If possible, I'd like to know the image's shape you made when creating scalogram (before resizing).</p>\n</blockquote>\n<p>The shape of scalogram before resize is (32, 10000), for each montage channel.</p>\n<ul>\n<li>I set the frequency axis resolution to 32, which is enough for my needs.<br>\nThe reason is that higher resolution increases generation time of scalogram significantly.<br>\n(It is also intended to prevent time-out during inference)</li>\n<li>The freqs input to the function is as follows<br>\n<code>freqs = jnp.linspace(0.5, 20.0, 32)</code><br>\n</li>\n</ul>\n<p>Then, to save storage consumption, I resize it once to (32, 1000) and save it in .npy format.<br>\n<br>\nNext, during the training phase, scalogram 2d image is generated by the following steps:</p>\n<ol>\n<li>Crop the center of each montage channnel image in the range of 256/512/768px. (i.e., shape=(32, 256or512or768))<ul>\n<li>The reason for not using max 1000px is to use 5sec time-shift augmentation.</li></ul></li>\n<li>Resize the image of each montage channnel to (16, 256) (when using 16ch montage; when using 4ch average: (64, 256))<ul>\n<li>You are correct that resize will result in lack of information. \nHowever, comparing the images before and after resize, I judge this effect is not that significant. \n(I may have the impression that information is \"aggregated\")<ul>\n<li>before resize (32, 768)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fd2a7966d9b78289f4def02d4dbcbf284%2Ftmp_1372816239_1.png?generation=1713269431324714&amp;alt=media\"></li>\n<li>after resize (32, 256)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F1cffbda4a7a07e394325ea8135987314%2Ftmp_1372816239_1_resize.png?generation=1713269448710728&amp;alt=media\"></li></ul></li>\n<li>Also, some timm backbones, such as swinv2, which has good accuracy this time, have a limitation of 256px (or 224px) resolution.<br>\nTherefore, I used resize because it was worth accepting the trade-off for lack of information to use these models.</li>\n<li>There are no special tricks here, only application of resize by bilinear.</li></ul></li>\n<li>Stack each montages channel in the frequency direction to obtain single (256, 256) image.<br>\nIn the case of 16 montages, (16x16=256, 256)</li>\n</ol>\n<p>If you have any questions, please feel free to ask again. :)</p>",
              "rawMarkdown": "Dear MBOOK, \n\t\t\nThank you for your comment! And sorry for the late reply.\n<bs>\n\t\t\n>P.S. If possible, I'd like to know the image's shape you made when creating scalogram (before resizing).\n\nThe shape of scalogram before resize is (32, 10000), for each montage channel.\n- I set the frequency axis resolution to 32, which is enough for my needs.\nThe reason is that higher resolution increases generation time of scalogram significantly.\n(It is also intended to prevent time-out during inference)\n- The freqs input to the function is as follows\n  `freqs = jnp.linspace(0.5, 20.0, 32)`\n<bs>\n\nThen, to save storage consumption, I resize it once to (32, 1000) and save it in .npy format.\n<bs>\nNext, during the training phase, scalogram 2d image is generated by the following steps:\n1. Crop the center of each montage channnel image in the range of 256/512/768px. (i.e., shape=(32, 256or512or768))\n    - The reason for not using max 1000px is to use 5sec time-shift augmentation.\n2. Resize the image of each montage channnel to (16, 256) (when using 16ch montage; when using 4ch average: (64, 256))\n    - You are correct that resize will result in lack of information. \nHowever, comparing the images before and after resize, I judge this effect is not that significant. \n(I may have the impression that information is \"aggregated\")\n        - before resize (32, 768)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fd2a7966d9b78289f4def02d4dbcbf284%2Ftmp_1372816239_1.png?generation=1713269431324714&alt=media)\n        - after resize (32, 256)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F1cffbda4a7a07e394325ea8135987314%2Ftmp_1372816239_1_resize.png?generation=1713269448710728&alt=media)\n    - Also, some timm backbones, such as swinv2, which has good accuracy this time, have a limitation of 256px (or 224px) resolution.\nTherefore, I used resize because it was worth accepting the trade-off for lack of information to use these models.\n    - There are no special tricks here, only application of resize by bilinear.\n3. Stack each montages channel in the frequency direction to obtain single (256, 256) image.\nIn the case of 16 montages, (16x16=256, 256)\n\nIf you have any questions, please feel free to ask again. :)",
              "votes": 1
            },
            {
              "id": 2757127,
              "postDate": "2024-04-17T10:55:54.633Z",
              "content": "<p>Thanks for very clear and logical representations!<br>\nI learned a loft of things from you. </p>\n<p>Thanks again, and happy kaggling!</p>",
              "rawMarkdown": "Thanks for very clear and logical representations!\nI learned a loft of things from you. \n\nThanks again, and happy kaggling!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2744628,
      "postDate": "2024-04-10T03:48:04.947Z",
      "content": "<h1>suguuuu part</h1>\n<h2>Overview</h2>\n<ul>\n<li>I first created Model1.<ul>\n<li>Model2 combines the Yamash preprocessing (I only use 50sec). Please refer to the Yamash solution.</li></ul></li>\n</ul>\n<h3>model 1 (4models seed ensemble, one model cv: 0.2399)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F9393eef898d0837f198d56e3727d281e%2Fmodel1.jpg?generation=1712721648829864&amp;alt=media\"></p>\n<h3>model 2 (4models seed ensemble, one model cv: 0.23810)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb477d11e17f58820f36beabfb4406579%2Fmodel2.jpg?generation=1712721635166923&amp;alt=media\"></p>\n<h2>Preprocessing ( How to create a stacked scalogram image )</h2>\n<ul>\n<li>Extract features from rawEEG</li>\n<li>x.clip(-1024.1024)/32</li>\n<li>Crop EEG: crop the middle 50sec(10,000 frames) of EEG.<ul>\n<li>Crop 50 seconds (10,000 frames) performed better than 25 seconds or 10 seconds.</li></ul></li>\n<li>label : Normalize by summing labels with eeg_id</li>\n<li>Continueous Wavelet Transform(CWT)<ul>\n<li>code : <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275356#1529639\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275356#1529639</a><ul>\n<li>Settings: wavelet_width=7, fs=200, lower_freq=0.5, upper_freq=40, n_scales=40, border_crop=1, stride=16.</li>\n<li>Adjust the parameters to approximate a 512x512 size when stacked.<ul>\n<li>Configured for 0.5-40Hz, which outperformed the 0.5-20Hz setting.</li></ul></li>\n<li>input EEG:18 x 10000, output scalogram:18x40x625</li></ul></li></ul></li>\n<li>Stack CWT images vertically.</li>\n<li>Resize to 512 x 512</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>2stage training (5epoch =&gt; 15epoch)<ul>\n<li>Augmentation<ul>\n<li>XYMasking, Mixup</li></ul></li></ul></li>\n<li>optimizer : Adan</li>\n<li>scheduler : CosineAnnealingLR</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>The MaxViT_base is the best for the scalogram image.</li>\n</ul>\n<h2>Do not work for me</h2>\n<ul>\n<li>STFT, Kaggle spec, CQT.</li>\n<li>Train with plotted image.</li>\n<li>Various montage</li>\n</ul>\n<h2>Why did I use CWT?</h2>\n<ul>\n<li>I tried various experiments with STFT, but couldn't achieve good performance. </li>\n<li>So, I chatted to ChatGPT and was able to get the following information:</li>\n</ul>\n<pre><code> capture the    signals  high non-stationarity, its determine the best .\n\n### Short- Fourier  (STFT)\n- **Strengths:**\n  - Relatively easy  implement  widely used.\n  - Offers an intuitive presentation  -frequency information.\n- **Weaknesses:**\n  - Fixed  size creates a trade-    frequency resolution.\n  - Limited ability  capture  features  highly non-stationary signals.\n\n### Wavelet \n- **Strengths:**\n  - Capable  multi-resolution analysis, capturing signal  at different scales.\n  - Excellently captures  features  non-stationary  complex signals.\n  - Suitable  detecting short-duration events, analyzing abrupt changes,  non-linear   signals.\n- **Weaknesses:**\n  - Requires the selection  an appropriate wavelet , which can demand specialized knowledge.\n  - Implementation  interpretation can become complex.\n\n### Superlet \n- **Strengths:**\n  - Provides high -frequency resolution, capturing fine details  the signal.\n  - High capability  distinguish overlapped  short-duration signal components.\n  - Particularly effective  analyzing signals  high non-stationarity, such  complex brain wave patterns.\n- **Weaknesses:**\n  - Relatively ,  potentially limited resources  examples  implementation available.\n  - May incur high computational .\n\n### Best   Analyzing High Non-stationary Signals\n capture the    highly non-stationary signals, **Wavelet **  **Superlet ** are particularly suitable. These methods provide high flexibility  adaptability  temporal  frequency changes  the signal, making them effective  analyzing complex,  signals. Wavelet ,  its versatility   feature extraction capability,  broadly adopted. The Superlet  may be chosen  even higher -frequency resolution needs   analyzing extremely complex signals.\n</code></pre>",
      "rawMarkdown": "# suguuuu part\n## Overview\n - I first created Model1.\n    - Model2 combines the Yamash preprocessing (I only use 50sec). Please refer to the Yamash solution.\n\n### model 1 (4models seed ensemble, one model cv: 0.2399)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F9393eef898d0837f198d56e3727d281e%2Fmodel1.jpg?generation=1712721648829864&alt=media)\n### model 2 (4models seed ensemble, one model cv: 0.23810)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb477d11e17f58820f36beabfb4406579%2Fmodel2.jpg?generation=1712721635166923&alt=media)\n\n## Preprocessing ( How to create a stacked scalogram image )\n - Extract features from rawEEG\n - x.clip(-1024.1024)/32\n - Crop EEG: crop the middle 50sec(10,000 frames) of EEG.\n   - Crop 50 seconds (10,000 frames) performed better than 25 seconds or 10 seconds.\n - label : Normalize by summing labels with eeg_id\n - Continueous Wavelet Transform(CWT)\n   - code : https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275356#1529639\n      - Settings: wavelet_width=7, fs=200, lower_freq=0.5, upper_freq=40, n_scales=40, border_crop=1, stride=16.\n      - Adjust the parameters to approximate a 512x512 size when stacked.\n         - Configured for 0.5-40Hz, which outperformed the 0.5-20Hz setting.\n      - input EEG:18 x 10000, output scalogram:18x40x625\n - Stack CWT images vertically.\n - Resize to 512 x 512\n\n## Training\n - 2stage training (5epoch => 15epoch)\n  - Augmentation\n      - XYMasking, Mixup\n - optimizer : Adan\n - scheduler : CosineAnnealingLR\n\n## Model\n - The MaxViT_base is the best for the scalogram image.\n\n## Do not work for me\n - STFT, Kaggle spec, CQT.\n - Train with plotted image.\n - Various montage\n\n## Why did I use CWT?\n - I tried various experiments with STFT, but couldn't achieve good performance. \n - So, I chatted to ChatGPT and was able to get the following information:\n```\nTo capture the local characteristics of signals with high non-stationarity, it's essential to choose an analysis method that can adapt to the varying nature of the signal. Considering the strengths and weaknesses of Wavelet Transform, Superlet Transform, and Short-Time Fourier Transform (STFT), let's determine the best option.\n\n### Short-Time Fourier Transform (STFT)\n- **Strengths:**\n  - Relatively easy to implement and widely used.\n  - Offers an intuitive presentation of time-frequency information.\n- **Weaknesses:**\n  - Fixed window size creates a trade-off between time and frequency resolution.\n  - Limited ability to capture local features of highly non-stationary signals.\n\n### Wavelet Transform\n- **Strengths:**\n  - Capable of multi-resolution analysis, capturing signal characteristics at different scales.\n  - Excellently captures local features of non-stationary or complex signals.\n  - Suitable for detecting short-duration events, analyzing abrupt changes, and non-linear characteristics within signals.\n- **Weaknesses:**\n  - Requires the selection of an appropriate wavelet function, which can demand specialized knowledge.\n  - Implementation and interpretation can become complex.\n\n### Superlet Transform\n- **Strengths:**\n  - Provides high time-frequency resolution, capturing fine details of the signal.\n  - High capability to distinguish overlapped or short-duration signal components.\n  - Particularly effective for analyzing signals with high non-stationarity, such as complex brain wave patterns.\n- **Weaknesses:**\n  - Relatively new, with potentially limited resources or examples of implementation available.\n  - May incur high computational costs.\n\n### Best Transform for Analyzing High Non-stationary Signals\nTo capture the local characteristics of highly non-stationary signals, **Wavelet Transform** or **Superlet Transform** are particularly suitable. These methods provide high flexibility and adaptability to temporal and frequency changes in the signal, making them effective for analyzing complex, varying signals. Wavelet Transform, with its versatility and local feature extraction capability, is broadly adopted. The Superlet Transform may be chosen for even higher time-frequency resolution needs or when analyzing extremely complex signals.\n```",
      "votes": 16
    },
    {
      "id": 2744643,
      "postDate": "2024-04-10T04:07:32.053Z",
      "content": "<p>First and foremost, I would like to express my gratitude to the competition host who organized this interesting competition, as well as the Kaggle team. Additionally, I am thankful for my teammates who worked together with me until the end.</p>\n<h1>kfuji's part</h1>\n<h2>Preprocessing:</h2>\n<ul>\n<li>Using all raw EEG data excluding EKG, taking differences between specific pairs to create 18 types of signals.</li>\n<li>For each of the 18 types of signals, the continuous wavelet transform (CWT) was applied following the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628\" target=\"_blank\">suguuuu's method</a>. To increase the variation in input data, I used Paul as the mother wavelet instead of Morlet. I used the implementation of <a href=\"https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py\" target=\"_blank\">https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py</a> and <a href=\"https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation\" target=\"_blank\">https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation</a>.</li>\n<li>The scalograms obtained through the CWT will be stacked vertically to form a single 2D image and resized to 512x512 as the final step.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8448924%2F06de133392f73c8ea8e823be09c55c25%2FHMS_CWT_Paul.png?generation=1712722022756620&amp;alt=media\"></li>\n</ul>\n<h2>Model:</h2>\n<ul>\n<li>maxvit_base_tf_512.in21k_ft_in1k</li>\n</ul>\n<h2>Training:</h2>\n<ul>\n<li>During training, one eeg_sub_id was randomly selected from eeg_id for each epoch. The input data consisted of 10,000 frames (50 seconds) starting from the corresponding eeg_label_offset_seconds. Furthermore, the labels assigned to the eeg_sub_id were utilized for training.</li>\n<li>Lr: 1e-3</li>\n<li>Scheduelr: CosineAnnealingLR</li>\n<li>Optimizer: Adan</li>\n<li>Augmentation: XYMasking, mixup</li>\n<li>2stage training<ul>\n<li>Stage1: 5 epochs (all data) </li>\n<li>Stage2: 15 epochs (10 votes or more)</li></ul></li>\n</ul>\n<h2>Variation:</h2>\n<ul>\n<li>The following three models, trained with variations in input frame length and Paul's parameters, were merged into the final team submission.<ul>\n<li>2000frame(10sec), Paul(m=4), CV:0.2475</li>\n<li>5000 frame(25sec), Paul(m=4), CV:0.2309</li>\n<li>5000frame(25sec), Paul(m=16), CV:0.2311</li></ul></li>\n</ul>\n<h2>What didn't work well:</h2>\n<ul>\n<li>Using DOG (Derivative of Gaussian) instead of Paul.</li>\n<li>Longer input frames (performed well individually but did not contribute to the ensemble).</li>\n</ul>",
      "rawMarkdown": "First and foremost, I would like to express my gratitude to the competition host who organized this interesting competition, as well as the Kaggle team. Additionally, I am thankful for my teammates who worked together with me until the end.\n\n# kfuji's part\n## Preprocessing:\n- Using all raw EEG data excluding EKG, taking differences between specific pairs to create 18 types of signals.\n- For each of the 18 types of signals, the continuous wavelet transform (CWT) was applied following the [suguuuu's method](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628). To increase the variation in input data, I used Paul as the mother wavelet instead of Morlet. I used the implementation of [https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py](https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py) and [https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation](https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation).\n- The scalograms obtained through the CWT will be stacked vertically to form a single 2D image and resized to 512x512 as the final step.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8448924%2F06de133392f73c8ea8e823be09c55c25%2FHMS_CWT_Paul.png?generation=1712722022756620&alt=media)\n\n## Model:\n- maxvit_base_tf_512.in21k_ft_in1k\n\n## Training:\n- During training, one eeg_sub_id was randomly selected from eeg_id for each epoch. The input data consisted of 10,000 frames (50 seconds) starting from the corresponding eeg_label_offset_seconds. Furthermore, the labels assigned to the eeg_sub_id were utilized for training.\n- Lr: 1e-3\n- Scheduelr: CosineAnnealingLR\n- Optimizer: Adan\n- Augmentation: XYMasking, mixup\n- 2stage training\n  - Stage1: 5 epochs (all data) \n  -  Stage2: 15 epochs (10 votes or more)\n\n## Variation:\n- The following three models, trained with variations in input frame length and Paul's parameters, were merged into the final team submission.\n  - 2000frame(10sec), Paul(m=4), CV:0.2475\n  - 5000 frame(25sec), Paul(m=4), CV:0.2309\n  - 5000frame(25sec), Paul(m=16), CV:0.2311\n\n## What didn't work well:\n- Using DOG (Derivative of Gaussian) instead of Paul.\n- Longer input frames (performed well individually but did not contribute to the ensemble).",
      "votes": 12,
      "replies": [
        {
          "id": 2755815,
          "postDate": "2024-04-16T17:12:57.117Z",
          "content": "<p>This comment's detailed methodical approach really impresses me. It's clear that cleaning the EEG data and improving the model architecture took a lot of work on the part of Fuji and the team. A thorough understanding of the issue domain is demonstrated by the choice to use wavelet transforms using Paul as the mother wavelet as well as the meticulous selection of training settings and augmentation strategies.</p>",
          "rawMarkdown": "This comment's detailed methodical approach really impresses me. It's clear that cleaning the EEG data and improving the model architecture took a lot of work on the part of Fuji and the team. A thorough understanding of the issue domain is demonstrated by the choice to use wavelet transforms using Paul as the mother wavelet as well as the meticulous selection of training settings and augmentation strategies.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2744890,
      "postDate": "2024-04-10T07:21:34.207Z",
      "content": "<p>Congratulations ! Well I used to think Cris's method was not the best for eval as per eeg id we might have many labels, only same label parts could be seen as aug, but from your wok seems I am wrong…, I will try to change this strategy to see the diff, thanks for sharing, entmax is also cool!</p>",
      "rawMarkdown": "Congratulations ! Well I used to think Cris's method was not the best for eval as per eeg id we might have many labels, only same label parts could be seen as aug, but from your wok seems I am wrong..., I will try to change this strategy to see the diff, thanks for sharing, entmax is also cool!",
      "votes": 6
    },
    {
      "id": 2744757,
      "postDate": "2024-04-10T05:22:54.197Z",
      "content": "<p>Awesome way of reporting, lots of interesting content. Well done for the win!</p>",
      "rawMarkdown": "Awesome way of reporting, lots of interesting content. Well done for the win!",
      "votes": 4
    },
    {
      "id": 2745238,
      "postDate": "2024-04-10T13:40:44.080Z",
      "content": "<p>Congratulations! 🎉 Thanks for sharing your excellent work. In this competition, I didn't make full use of the raw EEG data, but I learned a lot from your detailed analysis.</p>",
      "rawMarkdown": "Congratulations! 🎉 Thanks for sharing your excellent work. In this competition, I didn't make full use of the raw EEG data, but I learned a lot from your detailed analysis.",
      "votes": 1
    },
    {
      "id": 2744666,
      "postDate": "2024-04-10T04:17:41.513Z",
      "content": "<p>Congratulations on first prize in this competition. Thanks for sharing the details of your solution. </p>",
      "rawMarkdown": "Congratulations on first prize in this competition. Thanks for sharing the details of your solution. ",
      "votes": 2
    },
    {
      "id": 2756393,
      "postDate": "2024-04-17T02:38:57.060Z",
      "content": "<p>Congrats! Really makes me wonder how much of raw EEG data is actually noise vs relevant information. It seems like depending on the data format, it can be interpreted as either</p>",
      "rawMarkdown": "Congrats! Really makes me wonder how much of raw EEG data is actually noise vs relevant information. It seems like depending on the data format, it can be interpreted as either"
    },
    {
      "id": 2754400,
      "postDate": "2024-04-16T02:49:50.607Z",
      "content": "<p>Congrats !!! for securing 1st position</p>",
      "rawMarkdown": "Congrats !!! for securing 1st position"
    },
    {
      "id": 2748002,
      "postDate": "2024-04-12T07:40:33.433Z",
      "content": "<p>excellent work!</p>",
      "rawMarkdown": "excellent work!"
    },
    {
      "id": 2747661,
      "postDate": "2024-04-12T02:37:52.977Z",
      "content": "<p>fantastic work. congrats to your team :)</p>",
      "rawMarkdown": "fantastic work. congrats to your team :)"
    },
    {
      "id": 2747097,
      "postDate": "2024-04-11T17:15:44.873Z",
      "content": "<p>Congratulations on winning competition. ⭐</p>",
      "rawMarkdown": "Congratulations on winning competition. ⭐"
    },
    {
      "id": 2746718,
      "postDate": "2024-04-11T12:20:31.267Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations"
    },
    {
      "id": 2746543,
      "postDate": "2024-04-11T09:54:32.500Z",
      "content": "<p>Congratulations on taking first place. 🎉</p>",
      "rawMarkdown": "Congratulations on taking first place. 🎉"
    },
    {
      "id": 2745797,
      "postDate": "2024-04-10T20:20:03.573Z",
      "content": "<p>Congratulations!  Along with the models, I am stunned to find that the basic CV strategy can still lead the 1st place even many top solutions use the high-number voting data for validation along with training.  It is encouraging to stick to the basic CV, which still works for training with the selected data and may hedge a potential shake-up.</p>",
      "rawMarkdown": "Congratulations!  Along with the models, I am stunned to find that the basic CV strategy can still lead the 1st place even many top solutions use the high-number voting data for validation along with training.  It is encouraging to stick to the basic CV, which still works for training with the selected data and may hedge a potential shake-up.",
      "replies": [
        {
          "id": 2745972,
          "postDate": "2024-04-11T00:51:45.013Z",
          "content": "<p>Sorry, we didn't explain it well enough, but we were also using validation data with vote &gt;= 10.</p>",
          "rawMarkdown": "Sorry, we didn't explain it well enough, but we were also using validation data with vote >= 10.",
          "votes": 1,
          "replies": [
            {
              "id": 2747130,
              "postDate": "2024-04-11T17:32:05.060Z",
              "content": "<p>Thank you for the clarification.  It is yet good to know it!</p>",
              "rawMarkdown": "Thank you for the clarification.  It is yet good to know it!"
            }
          ]
        }
      ]
    },
    {
      "id": 2745011,
      "postDate": "2024-04-10T09:49:00.537Z",
      "content": "<p>I have a question， when counting for votes》10 it is per eeg subid or per eeg id？</p>",
      "rawMarkdown": "I have a question， when counting for votes》10 it is per eeg subid or per eeg id？",
      "replies": [
        {
          "id": 2745112,
          "postDate": "2024-04-10T11:41:22.440Z",
          "content": "<p>One data from each eeg_id with more than 10 votes was used for CV calculations. We treated the number of votes of the first eeg_sub_id as for that eeg_id.</p>",
          "rawMarkdown": "One data from each eeg_id with more than 10 votes was used for CV calculations. We treated the number of votes of the first eeg_sub_id as for that eeg_id.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2764048,
      "postDate": "2024-04-20T21:11:59.503Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2744629,
      "author_name": "yamash",
      "author_url": "",
      "post_date": "2024-04-10T03:48:49.807000",
      "content": "<p>Thank you Kaggle &amp; hosts for organizing an interesting competition.<br>\nThanks also to my teammates who took part with me. It was fun to participate in close discussions.</p>\n<h2>yamash’s part</h2>\n<h3>Model</h3>\n<p>I lined up the raw EEG signal as it was and created a single 2D image, which was then processed with a 2D CNN model.<br>\nIn more detail, 18 signals of bipolar montage were cropped at three different intervals (2000, 5000 and 10000 samples),  resized and concatenated into a single image.<br>\nI trained several models based on this method. Depending on the model, I changed the range of the bandpass filter to create a 2D image.</p>\n<p>This is a very simple method, but the CV was around 0.24.</p>\n<p>Following 5 models are used for final sub:</p>\n<ul>\n<li>4 x convnext atto models with different seed and bandpass filter (CV: 0.2452, 0.2385, 0.2457, 0.2351)</li>\n<li>1 x inception next tiny (CV: 0.2309)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5568744%2Fb0284055fbb92fab6ac8f4a71bdcc343%2Fhms_figure_yamash.png?generation=1712720787699483&amp;alt=media\"></p>\n<h3>Ensemble</h3>\n<p>I used non-negative linear regression for ensemble.<br>\nEach target variable was estimated by non-negative linear regression from the values of the same target variable predicted by each model.</p>\n<p>At first, our ensemble method is the average of all models, or a single weight was given to each model. We decided on using non-negative linear regression  when we realized that the correlation between CV and public lb was maintained even when more overfit to training data.</p>\n<h3>Replacing softmax with entmax</h3>\n<p>The output of softmax is not zero for every target variable, but many training data labels have some zero probability classes. <br>\nTo deal with this difference, I tried to replace softmax with sparsemax[1] and entmax[2]. <br>\nThese functions can output sparser results compared to softmax. <br>\nFinally, we decided to use entmax with a small alpha parameter (~1.03) for all our single models, which improved both public and private LB score around 0.004.<br>\nIn practice, small alpha didn’t output sparse results, but only of making the values a little more crisp.</p>\n<p>I found this almost the end of this competition, so I couldn’t have enough experiments of using entmax and sparsemax in the training stage.</p>\n<p>[1] <a href=\"https://arxiv.org/abs/1602.02068\" target=\"_blank\">https://arxiv.org/abs/1602.02068</a><br>\n[2] <a href=\"https://arxiv.org/abs/1905.05702\" target=\"_blank\">https://arxiv.org/abs/1905.05702</a></p>",
      "votes": 24,
      "replies": [
        {
          "id": 2744905,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2024-04-10T07:40:40.133000",
          "content": "<p>I really like the entmax idea 💡 thanks for sharing. For such a small change it has a relatively big effect. Are you aware if it gave similar lift in CV and LB?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2745003,
              "author_name": "yamash",
              "author_url": "",
              "post_date": "2024-04-10T09:40:55.203000",
              "content": "<p>Thank you for the comment.<br>\nAs for CV, there were some improvements and some not, depending on each single model. And there was no overall improvement in CV. CV was calculated with labels averaged over the same eeg_id, so it is possible that the effect of entmax was less pronounced than with the labels in the actual test data.<br>\nOn the other hand, for public LB, scores improved consistently regardless of the single model combination.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2765977,
          "author_name": "Simon Beck",
          "author_url": "",
          "post_date": "2024-04-21T12:39:38.373000",
          "content": "<p>how can u use different activation softmax and entmax in train and test? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2744634,
      "author_name": "Muku",
      "author_url": "",
      "post_date": "2024-04-10T03:53:57.120000",
      "content": "<h1>Muku's part</h1>\n<p>First of all, I would like to thank Harvard Medical School and Kaggle team for organizing a great competition, the participants for sharing useful ideas, and my teammates for fighting alongside me.</p>\n<p><br></p>\n<h2>Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F0a9de44bccbe20701429d7762e229940%2Fhms_architecture.png?generation=1750902445619521&amp;alt=media\" alt=\"\"></p>\n<p><br></p>\n<h2>1. Model</h2>\n<ul>\n<li>Based on the 16-channel anterior-posterior montages generated from raw EEG, two types of image representations are generated and input to timm model.<br>\n<br></li>\n<li><strong>First：Temporal feature maps by 1D temporal convolution</strong><ul>\n<li>This idea was inspired by <a href=\"https://arxiv.org/pdf/1611.08024.pdf\" target=\"_blank\">EEGNet paper</a> and <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476\" target=\"_blank\">G2Net Gravitational Wave Detection | Top 1 solution</a>.<ul>\n<li>I expected to obtain filters for the rhythmic temporal features specific to each symptom, by learning.</li></ul></li>\n<li>The kernel size was set to the same value as the sampling rate (200).</li>\n<li>Stack the generated temporal feature maps by montage channels or feature maps to obtain input images for timm mddel.<ul>\n<li>Stacking each montage channels has better accuracy, and parallel use of both methods improves accuracy.</li></ul></li>\n<li>Optional: Temporal feature maps are directly input to the GRU layer, which improves accuracy, so added as a variation of Ensemble.<br>\n<br></li></ul></li>\n<li><strong>Second：Time-Frequency Transform (superlets / STFT)</strong><ul>\n<li>In the early stages of the study, I used STFT, but it did not appear to be a good representation because of the loss of resolution in the time or frequency.</li>\n<li>Then, I use <strong>superlets</strong> as a better representation method. (CWT was being investigated by suguuuuu and kfuji, so I used this as a variation).<ul>\n<li>Use the following repositories published under the MIT License<br>\n<a href=\"https://github.com/irhum/superlets\" target=\"_blank\">https://github.com/irhum/superlets</a></li>\n<li>settings：<ul>\n<li>min_freq, max_freq = 0.5, 20.0</li>\n<li>base_cycle, min_order, max_order = 1, 1, 16</li>\n<li>Adjusted for better resolution in the time direction</li></ul></li></ul></li>\n<li>The superlet is more compatible with frequency/time expressivity compared to STFT, and the representation is more predictable for labels even for humans.<ul>\n<li>Examples：<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F6e0ad7491adf39bb468b27237d3bd7c7%2Fsuperlets_1.png?generation=1712721056221695&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fa40bf80b7b07f12e04bb3b031dfd00a1%2Fsuperlets_2.png?generation=1712721073292828&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F296c1d4610df6d2f526aae619495e29d%2Fsuperlets_3.png?generation=1712721081681421&amp;alt=media\" alt=\"# \"><br>\n<br></li></ul></li></ul></li>\n<li><strong>timm model</strong><ul>\n<li>Try various timm backbones and add them to Ensemble. \nDiverse ensemble was effective because the KL divergence is sensitive to errors for extreme predictions.<ul>\n<li>Overall, smaller models performed well.</li>\n<li>Best backbone：<strong>swinv2_tiny_window16 (CV: 0.2229</strong>)</li>\n<li>The following also showed good accuracy and contributed to Ensemble<ul>\n<li>swinv2_tiny_window8</li>\n<li>caformer_s18</li>\n<li>gcvit_xtiny, xxtiny</li>\n<li>convnextv2_atto</li>\n<li>maxvit_pico</li>\n<li>inception_next_tiny</li>\n<li>poolformerv2_s12</li>\n<li>nextvit_small</li></ul></li></ul></li></ul></li>\n</ul>\n<p><br></p>\n<h2>2. Training</h2>\n<ul>\n<li>2-stage training (Stage1：&gt; 1 vote, Stage2：&gt; 9 votes)<ul>\n<li>Labels of vote: 1 appeared to be unreliable, so excluded from training.<br>\n(I had thought about applying pseudo labeling, but I didn't have enough time)</li></ul></li>\n<li>Learning rate：[Stage1：1e-3, Stage2：1e-4]</li>\n<li>Loss function：KLDivLoss（with aux loss for each model output）</li>\n<li>scheduler：CosineAnnealingLR</li>\n<li>Optimizer：Adan</li>\n<li>Data sampling：Each unique eeg_id &amp; label（Multiple samples may be generated from the same eeg_id）</li>\n<li>Label smoothing：Add offset of 0.02 to each vote before normalization.<ul>\n<li>By adding this before normalization, a stronger regularization is applied to labels with low vote counts, which are considered to have a relatively low confidence.</li></ul></li>\n<li>Augmentation：<ul>\n<li>±5 sec random time-shift</li>\n<li>Random bandpass filter (Only for waves)</li>\n<li>XYMasking（Only for specs）</li></ul></li>\n<li>butter bandpass filter (Only for waves)：Set highcut to 30~40Hz<ul>\n<li>When high-frequency noise exists, \"Others\" votes tended to increase, so these noises are determined to be important for predicition.</li></ul></li>\n</ul>",
      "votes": 22,
      "replies": [
        {
          "id": 2744910,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2024-04-10T07:53:38.050000",
          "content": "<p>Very clever idea on the label smoothing.. applying before normalisation to smooth low count labels more</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2753536,
          "author_name": "MBOOK",
          "author_url": "",
          "post_date": "2024-04-15T14:50:31.197000",
          "content": "<p>Congrats for 1st place!</p>\n<p>I have one question about 2D images created by superlet.</p>\n<p>It seems you made scalograms by setting the time-axis as about 800. <br>\nBut when resizing it to (256,256), the images are totally different from the original one, which would lead to the loss of information in time direction.</p>\n<p>Did you make any treatment against that problem?</p>\n<p>P.S. If possible, I'd like to know the image's shape you made when creating scalogram (before resizing).</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2755239,
              "author_name": "Muku",
              "author_url": "",
              "post_date": "2024-04-16T12:42:06.467000",
              "content": "<p>Dear MBOOK, </p>\n<p>Thank you for your comment! And sorry for the late reply.<br>\n</p>\n<blockquote>\n  <p>P.S. If possible, I'd like to know the image's shape you made when creating scalogram (before resizing).</p>\n</blockquote>\n<p>The shape of scalogram before resize is (32, 10000), for each montage channel.</p>\n<ul>\n<li>I set the frequency axis resolution to 32, which is enough for my needs.<br>\nThe reason is that higher resolution increases generation time of scalogram significantly.<br>\n(It is also intended to prevent time-out during inference)</li>\n<li>The freqs input to the function is as follows<br>\n<code>freqs = jnp.linspace(0.5, 20.0, 32)</code><br>\n</li>\n</ul>\n<p>Then, to save storage consumption, I resize it once to (32, 1000) and save it in .npy format.<br>\n<br>\nNext, during the training phase, scalogram 2d image is generated by the following steps:</p>\n<ol>\n<li>Crop the center of each montage channnel image in the range of 256/512/768px. (i.e., shape=(32, 256or512or768))<ul>\n<li>The reason for not using max 1000px is to use 5sec time-shift augmentation.</li></ul></li>\n<li>Resize the image of each montage channnel to (16, 256) (when using 16ch montage; when using 4ch average: (64, 256))<ul>\n<li>You are correct that resize will result in lack of information. \nHowever, comparing the images before and after resize, I judge this effect is not that significant. \n(I may have the impression that information is \"aggregated\")<ul>\n<li>before resize (32, 768)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fd2a7966d9b78289f4def02d4dbcbf284%2Ftmp_1372816239_1.png?generation=1713269431324714&amp;alt=media\"></li>\n<li>after resize (32, 256)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F1cffbda4a7a07e394325ea8135987314%2Ftmp_1372816239_1_resize.png?generation=1713269448710728&amp;alt=media\"></li></ul></li>\n<li>Also, some timm backbones, such as swinv2, which has good accuracy this time, have a limitation of 256px (or 224px) resolution.<br>\nTherefore, I used resize because it was worth accepting the trade-off for lack of information to use these models.</li>\n<li>There are no special tricks here, only application of resize by bilinear.</li></ul></li>\n<li>Stack each montages channel in the frequency direction to obtain single (256, 256) image.<br>\nIn the case of 16 montages, (16x16=256, 256)</li>\n</ol>\n<p>If you have any questions, please feel free to ask again. :)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2757127,
              "author_name": "MBOOK",
              "author_url": "",
              "post_date": "2024-04-17T10:55:54.633000",
              "content": "<p>Thanks for very clear and logical representations!<br>\nI learned a loft of things from you. </p>\n<p>Thanks again, and happy kaggling!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2744628,
      "author_name": "suguuuuu",
      "author_url": "",
      "post_date": "2024-04-10T03:48:04.947000",
      "content": "<h1>suguuuu part</h1>\n<h2>Overview</h2>\n<ul>\n<li>I first created Model1.<ul>\n<li>Model2 combines the Yamash preprocessing (I only use 50sec). Please refer to the Yamash solution.</li></ul></li>\n</ul>\n<h3>model 1 (4models seed ensemble, one model cv: 0.2399)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F9393eef898d0837f198d56e3727d281e%2Fmodel1.jpg?generation=1712721648829864&amp;alt=media\"></p>\n<h3>model 2 (4models seed ensemble, one model cv: 0.23810)</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb477d11e17f58820f36beabfb4406579%2Fmodel2.jpg?generation=1712721635166923&amp;alt=media\"></p>\n<h2>Preprocessing ( How to create a stacked scalogram image )</h2>\n<ul>\n<li>Extract features from rawEEG</li>\n<li>x.clip(-1024.1024)/32</li>\n<li>Crop EEG: crop the middle 50sec(10,000 frames) of EEG.<ul>\n<li>Crop 50 seconds (10,000 frames) performed better than 25 seconds or 10 seconds.</li></ul></li>\n<li>label : Normalize by summing labels with eeg_id</li>\n<li>Continueous Wavelet Transform(CWT)<ul>\n<li>code : <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275356#1529639\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275356#1529639</a><ul>\n<li>Settings: wavelet_width=7, fs=200, lower_freq=0.5, upper_freq=40, n_scales=40, border_crop=1, stride=16.</li>\n<li>Adjust the parameters to approximate a 512x512 size when stacked.<ul>\n<li>Configured for 0.5-40Hz, which outperformed the 0.5-20Hz setting.</li></ul></li>\n<li>input EEG:18 x 10000, output scalogram:18x40x625</li></ul></li></ul></li>\n<li>Stack CWT images vertically.</li>\n<li>Resize to 512 x 512</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>2stage training (5epoch =&gt; 15epoch)<ul>\n<li>Augmentation<ul>\n<li>XYMasking, Mixup</li></ul></li></ul></li>\n<li>optimizer : Adan</li>\n<li>scheduler : CosineAnnealingLR</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>The MaxViT_base is the best for the scalogram image.</li>\n</ul>\n<h2>Do not work for me</h2>\n<ul>\n<li>STFT, Kaggle spec, CQT.</li>\n<li>Train with plotted image.</li>\n<li>Various montage</li>\n</ul>\n<h2>Why did I use CWT?</h2>\n<ul>\n<li>I tried various experiments with STFT, but couldn't achieve good performance. </li>\n<li>So, I chatted to ChatGPT and was able to get the following information:</li>\n</ul>\n<pre><code> capture the    signals  high non-stationarity, its determine the best .\n\n### Short- Fourier  (STFT)\n- **Strengths:**\n  - Relatively easy  implement  widely used.\n  - Offers an intuitive presentation  -frequency information.\n- **Weaknesses:**\n  - Fixed  size creates a trade-    frequency resolution.\n  - Limited ability  capture  features  highly non-stationary signals.\n\n### Wavelet \n- **Strengths:**\n  - Capable  multi-resolution analysis, capturing signal  at different scales.\n  - Excellently captures  features  non-stationary  complex signals.\n  - Suitable  detecting short-duration events, analyzing abrupt changes,  non-linear   signals.\n- **Weaknesses:**\n  - Requires the selection  an appropriate wavelet , which can demand specialized knowledge.\n  - Implementation  interpretation can become complex.\n\n### Superlet \n- **Strengths:**\n  - Provides high -frequency resolution, capturing fine details  the signal.\n  - High capability  distinguish overlapped  short-duration signal components.\n  - Particularly effective  analyzing signals  high non-stationarity, such  complex brain wave patterns.\n- **Weaknesses:**\n  - Relatively ,  potentially limited resources  examples  implementation available.\n  - May incur high computational .\n\n### Best   Analyzing High Non-stationary Signals\n capture the    highly non-stationary signals, **Wavelet **  **Superlet ** are particularly suitable. These methods provide high flexibility  adaptability  temporal  frequency changes  the signal, making them effective  analyzing complex,  signals. Wavelet ,  its versatility   feature extraction capability,  broadly adopted. The Superlet  may be chosen  even higher -frequency resolution needs   analyzing extremely complex signals.\n</code></pre>",
      "votes": 16,
      "replies": []
    },
    {
      "id": 2744643,
      "author_name": "kfuji",
      "author_url": "",
      "post_date": "2024-04-10T04:07:32.053000",
      "content": "<p>First and foremost, I would like to express my gratitude to the competition host who organized this interesting competition, as well as the Kaggle team. Additionally, I am thankful for my teammates who worked together with me until the end.</p>\n<h1>kfuji's part</h1>\n<h2>Preprocessing:</h2>\n<ul>\n<li>Using all raw EEG data excluding EKG, taking differences between specific pairs to create 18 types of signals.</li>\n<li>For each of the 18 types of signals, the continuous wavelet transform (CWT) was applied following the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628\" target=\"_blank\">suguuuu's method</a>. To increase the variation in input data, I used Paul as the mother wavelet instead of Morlet. I used the implementation of <a href=\"https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py\" target=\"_blank\">https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py</a> and <a href=\"https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation\" target=\"_blank\">https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation</a>.</li>\n<li>The scalograms obtained through the CWT will be stacked vertically to form a single 2D image and resized to 512x512 as the final step.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8448924%2F06de133392f73c8ea8e823be09c55c25%2FHMS_CWT_Paul.png?generation=1712722022756620&amp;alt=media\"></li>\n</ul>\n<h2>Model:</h2>\n<ul>\n<li>maxvit_base_tf_512.in21k_ft_in1k</li>\n</ul>\n<h2>Training:</h2>\n<ul>\n<li>During training, one eeg_sub_id was randomly selected from eeg_id for each epoch. The input data consisted of 10,000 frames (50 seconds) starting from the corresponding eeg_label_offset_seconds. Furthermore, the labels assigned to the eeg_sub_id were utilized for training.</li>\n<li>Lr: 1e-3</li>\n<li>Scheduelr: CosineAnnealingLR</li>\n<li>Optimizer: Adan</li>\n<li>Augmentation: XYMasking, mixup</li>\n<li>2stage training<ul>\n<li>Stage1: 5 epochs (all data) </li>\n<li>Stage2: 15 epochs (10 votes or more)</li></ul></li>\n</ul>\n<h2>Variation:</h2>\n<ul>\n<li>The following three models, trained with variations in input frame length and Paul's parameters, were merged into the final team submission.<ul>\n<li>2000frame(10sec), Paul(m=4), CV:0.2475</li>\n<li>5000 frame(25sec), Paul(m=4), CV:0.2309</li>\n<li>5000frame(25sec), Paul(m=16), CV:0.2311</li></ul></li>\n</ul>\n<h2>What didn't work well:</h2>\n<ul>\n<li>Using DOG (Derivative of Gaussian) instead of Paul.</li>\n<li>Longer input frames (performed well individually but did not contribute to the ensemble).</li>\n</ul>",
      "votes": 12,
      "replies": [
        {
          "id": 2755815,
          "author_name": "ANTONY ALEXIA",
          "author_url": "",
          "post_date": "2024-04-16T17:12:57.117000",
          "content": "<p>This comment's detailed methodical approach really impresses me. It's clear that cleaning the EEG data and improving the model architecture took a lot of work on the part of Fuji and the team. A thorough understanding of the issue domain is demonstrated by the choice to use wavelet transforms using Paul as the mother wavelet as well as the meticulous selection of training settings and augmentation strategies.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2744890,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2024-04-10T07:21:34.207000",
      "content": "<p>Congratulations ! Well I used to think Cris's method was not the best for eval as per eeg id we might have many labels, only same label parts could be seen as aug, but from your wok seems I am wrong…, I will try to change this strategy to see the diff, thanks for sharing, entmax is also cool!</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2744757,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-04-10T05:22:54.197000",
      "content": "<p>Awesome way of reporting, lots of interesting content. Well done for the win!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2745238,
      "author_name": "smalltrain",
      "author_url": "",
      "post_date": "2024-04-10T13:40:44.080000",
      "content": "<p>Congratulations! 🎉 Thanks for sharing your excellent work. In this competition, I didn't make full use of the raw EEG data, but I learned a lot from your detailed analysis.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2744666,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-04-10T04:17:41.513000",
      "content": "<p>Congratulations on first prize in this competition. Thanks for sharing the details of your solution. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2756393,
      "author_name": "ModeEric",
      "author_url": "",
      "post_date": "2024-04-17T02:38:57.060000",
      "content": "<p>Congrats! Really makes me wonder how much of raw EEG data is actually noise vs relevant information. It seems like depending on the data format, it can be interpreted as either</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2754400,
      "author_name": "Aaditya Porwal",
      "author_url": "",
      "post_date": "2024-04-16T02:49:50.607000",
      "content": "<p>Congrats !!! for securing 1st position</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2748002,
      "author_name": "Enes Koşar",
      "author_url": "",
      "post_date": "2024-04-12T07:40:33.433000",
      "content": "<p>excellent work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2747661,
      "author_name": "Niclo",
      "author_url": "",
      "post_date": "2024-04-12T02:37:52.977000",
      "content": "<p>fantastic work. congrats to your team :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2747097,
      "author_name": "zephyrus1",
      "author_url": "",
      "post_date": "2024-04-11T17:15:44.873000",
      "content": "<p>Congratulations on winning competition. ⭐</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2746718,
      "author_name": "Ganesh Talwar",
      "author_url": "",
      "post_date": "2024-04-11T12:20:31.267000",
      "content": "<p>congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2746543,
      "author_name": "SungyoonKim_KoK",
      "author_url": "",
      "post_date": "2024-04-11T09:54:32.500000",
      "content": "<p>Congratulations on taking first place. 🎉</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2745797,
      "author_name": "makio323",
      "author_url": "",
      "post_date": "2024-04-10T20:20:03.573000",
      "content": "<p>Congratulations!  Along with the models, I am stunned to find that the basic CV strategy can still lead the 1st place even many top solutions use the high-number voting data for validation along with training.  It is encouraging to stick to the basic CV, which still works for training with the selected data and may hedge a potential shake-up.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2745972,
          "author_name": "yamash",
          "author_url": "",
          "post_date": "2024-04-11T00:51:45.013000",
          "content": "<p>Sorry, we didn't explain it well enough, but we were also using validation data with vote &gt;= 10.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2747130,
              "author_name": "makio323",
              "author_url": "",
              "post_date": "2024-04-11T17:32:05.060000",
              "content": "<p>Thank you for the clarification.  It is yet good to know it!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2745011,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2024-04-10T09:49:00.537000",
      "content": "<p>I have a question， when counting for votes》10 it is per eeg subid or per eeg id？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2745112,
          "author_name": "yamash",
          "author_url": "",
          "post_date": "2024-04-10T11:41:22.440000",
          "content": "<p>One data from each eeg_id with more than 10 votes was used for CV calculations. We treated the number of votes of the first eeg_sub_id as for that eeg_id.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2764048,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-20T21:11:59.503000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2744626": "I would like to thank the organizers and the community of the competition!\nI'm grateful to the members who joined the team.\n\n# Code ( updated on 4/23/2024)\n- train : https://www.kaggle.com/code/sugupoko/1st-place-all-train-code/notebook\n- inferenece: https://www.kaggle.com/code/sugupoko/1st-place-hms-inference-code\n \n# Summary\nOur prediction pipeline is this.\nEach approaches are will added in the comments!\n - yamash's part   : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744629)\n - suguuuuu's part: [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628)\n - kfuji's part         : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744643)\n - Muku's part       : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744634)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fdb96007aaf4d5c3278bcda3b17c1bb14%2Fkaggle_overview.jpg?generation=1712721482613740&alt=media)\n\n# Team Validation Strategy\nThe verification method was the same for the entire team.\n - Validation using Chris's method\n  - link : https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\n   - Normalize by summing labels with eeg_id, crop the middle 10,000 frames of EEG.\n - Folding strategy : Group K-fold with a 5-fold split.\n - vote >=10\n",
    "2744629": "Thank you Kaggle & hosts for organizing an interesting competition.\nThanks also to my teammates who took part with me. It was fun to participate in close discussions.\n\n## yamash’s part\n\n\n### Model\n\nI lined up the raw EEG signal as it was and created a single 2D image, which was then processed with a 2D CNN model.\nIn more detail, 18 signals of bipolar montage were cropped at three different intervals (2000, 5000 and 10000 samples),  resized and concatenated into a single image.\nI trained several models based on this method. Depending on the model, I changed the range of the bandpass filter to create a 2D image.\n\nThis is a very simple method, but the CV was around 0.24.\n\nFollowing 5 models are used for final sub:\n- 4 x convnext atto models with different seed and bandpass filter (CV: 0.2452, 0.2385, 0.2457, 0.2351)\n- 1 x inception next tiny (CV: 0.2309)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5568744%2Fb0284055fbb92fab6ac8f4a71bdcc343%2Fhms_figure_yamash.png?generation=1712720787699483&alt=media)\n\n### Ensemble\n\nI used non-negative linear regression for ensemble.\nEach target variable was estimated by non-negative linear regression from the values of the same target variable predicted by each model.\n\nAt first, our ensemble method is the average of all models, or a single weight was given to each model. We decided on using non-negative linear regression  when we realized that the correlation between CV and public lb was maintained even when more overfit to training data.\n\n### Replacing softmax with entmax\n\nThe output of softmax is not zero for every target variable, but many training data labels have some zero probability classes. \nTo deal with this difference, I tried to replace softmax with sparsemax[1] and entmax[2]. \nThese functions can output sparser results compared to softmax. \nFinally, we decided to use entmax with a small alpha parameter (~1.03) for all our single models, which improved both public and private LB score around 0.004.\nIn practice, small alpha didn’t output sparse results, but only of making the values a little more crisp.\n\nI found this almost the end of this competition, so I couldn’t have enough experiments of using entmax and sparsemax in the training stage.\n\n[1] [https://arxiv.org/abs/1602.02068](https://arxiv.org/abs/1602.02068)\n[2] [https://arxiv.org/abs/1905.05702](https://arxiv.org/abs/1905.05702)\n",
    "2744634": "# Muku's part\n\nFirst of all, I would like to thank Harvard Medical School and Kaggle team for organizing a great competition, the participants for sharing useful ideas, and my teammates for fighting alongside me.\n\n<br>\n\n## Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F0a9de44bccbe20701429d7762e229940%2Fhms_architecture.png?generation=1750902445619521&alt=media)\n\n<br>\n\t\n## 1. Model\n- Based on the 16-channel anterior-posterior montages generated from raw EEG, two types of image representations are generated and input to timm model.\n<br>\n- **First：Temporal feature maps by 1D temporal convolution**\n    - This idea was inspired by [EEGNet paper](https://arxiv.org/pdf/1611.08024.pdf) and [G2Net Gravitational Wave Detection | Top 1 solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476).\n        - I expected to obtain filters for the rhythmic temporal features specific to each symptom, by learning.\n    - The kernel size was set to the same value as the sampling rate (200).\n    - Stack the generated temporal feature maps by montage channels or feature maps to obtain input images for timm mddel.\n        - Stacking each montage channels has better accuracy, and parallel use of both methods improves accuracy.\n    -  Optional: Temporal feature maps are directly input to the GRU layer, which improves accuracy, so added as a variation of Ensemble.\n<br>\n- **Second：Time-Frequency Transform (superlets / STFT)**\n    - In the early stages of the study, I used STFT, but it did not appear to be a good representation because of the loss of resolution in the time or frequency.\n    - Then, I use **superlets** as a better representation method. (CWT was being investigated by suguuuuu and kfuji, so I used this as a variation).\n        - Use the following repositories published under the MIT License\n[https://github.com/irhum/superlets](https://github.com/irhum/superlets)\n        - settings：\n            - min_freq, max_freq = 0.5, 20.0\n            - base_cycle, min_order, max_order = 1, 1, 16\n            - Adjusted for better resolution in the time direction\n    - The superlet is more compatible with frequency/time expressivity compared to STFT, and the representation is more predictable for labels even for humans.\n        - Examples：\n            ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F6e0ad7491adf39bb468b27237d3bd7c7%2Fsuperlets_1.png?generation=1712721056221695&alt=media)\n            ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2Fa40bf80b7b07f12e04bb3b031dfd00a1%2Fsuperlets_2.png?generation=1712721073292828&alt=media)\n            ![# ](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671348%2F296c1d4610df6d2f526aae619495e29d%2Fsuperlets_3.png?generation=1712721081681421&alt=media)\n<br>\n- **timm model**\n    - Try various timm backbones and add them to Ensemble. \nDiverse ensemble was effective because the KL divergence is sensitive to errors for extreme predictions.\n        - Overall, smaller models performed well.\n        - Best backbone：**swinv2_tiny_window16 (CV: 0.2229**)\n        - The following also showed good accuracy and contributed to Ensemble\n            - swinv2_tiny_window8\n            - caformer_s18\n            - gcvit_xtiny, xxtiny\n            - convnextv2_atto\n            - maxvit_pico\n            - inception_next_tiny\n            - poolformerv2_s12\n            - nextvit_small\n\n<br>\n## 2. Training\n- 2-stage training (Stage1：> 1 vote, Stage2：> 9 votes)\n    - Labels of vote: 1 appeared to be unreliable, so excluded from training.\n(I had thought about applying pseudo labeling, but I didn't have enough time)\n- Learning rate：[Stage1：1e-3, Stage2：1e-4]\n- Loss function：KLDivLoss（with aux loss for each model output）\n- scheduler：CosineAnnealingLR\n- Optimizer：Adan\n- Data sampling：Each unique eeg_id & label（Multiple samples may be generated from the same eeg_id）\n- Label smoothing：Add offset of 0.02 to each vote before normalization.\n    - By adding this before normalization, a stronger regularization is applied to labels with low vote counts, which are considered to have a relatively low confidence.\n- Augmentation：\n    - ±5 sec random time-shift\n    - Random bandpass filter (Only for waves)\n    - XYMasking（Only for specs）\n- butter bandpass filter (Only for waves)：Set highcut to 30~40Hz\n    - When high-frequency noise exists, \"Others\" votes tended to increase, so these noises are determined to be important for predicition.",
    "2744628": "# suguuuu part\n## Overview\n - I first created Model1.\n    - Model2 combines the Yamash preprocessing (I only use 50sec). Please refer to the Yamash solution.\n\n### model 1 (4models seed ensemble, one model cv: 0.2399)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F9393eef898d0837f198d56e3727d281e%2Fmodel1.jpg?generation=1712721648829864&alt=media)\n### model 2 (4models seed ensemble, one model cv: 0.23810)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb477d11e17f58820f36beabfb4406579%2Fmodel2.jpg?generation=1712721635166923&alt=media)\n\n## Preprocessing ( How to create a stacked scalogram image )\n - Extract features from rawEEG\n - x.clip(-1024.1024)/32\n - Crop EEG: crop the middle 50sec(10,000 frames) of EEG.\n   - Crop 50 seconds (10,000 frames) performed better than 25 seconds or 10 seconds.\n - label : Normalize by summing labels with eeg_id\n - Continueous Wavelet Transform(CWT)\n   - code : https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275356#1529639\n      - Settings: wavelet_width=7, fs=200, lower_freq=0.5, upper_freq=40, n_scales=40, border_crop=1, stride=16.\n      - Adjust the parameters to approximate a 512x512 size when stacked.\n         - Configured for 0.5-40Hz, which outperformed the 0.5-20Hz setting.\n      - input EEG:18 x 10000, output scalogram:18x40x625\n - Stack CWT images vertically.\n - Resize to 512 x 512\n\n## Training\n - 2stage training (5epoch => 15epoch)\n  - Augmentation\n      - XYMasking, Mixup\n - optimizer : Adan\n - scheduler : CosineAnnealingLR\n\n## Model\n - The MaxViT_base is the best for the scalogram image.\n\n## Do not work for me\n - STFT, Kaggle spec, CQT.\n - Train with plotted image.\n - Various montage\n\n## Why did I use CWT?\n - I tried various experiments with STFT, but couldn't achieve good performance. \n - So, I chatted to ChatGPT and was able to get the following information:\n```\nTo capture the local characteristics of signals with high non-stationarity, it's essential to choose an analysis method that can adapt to the varying nature of the signal. Considering the strengths and weaknesses of Wavelet Transform, Superlet Transform, and Short-Time Fourier Transform (STFT), let's determine the best option.\n\n### Short-Time Fourier Transform (STFT)\n- **Strengths:**\n  - Relatively easy to implement and widely used.\n  - Offers an intuitive presentation of time-frequency information.\n- **Weaknesses:**\n  - Fixed window size creates a trade-off between time and frequency resolution.\n  - Limited ability to capture local features of highly non-stationary signals.\n\n### Wavelet Transform\n- **Strengths:**\n  - Capable of multi-resolution analysis, capturing signal characteristics at different scales.\n  - Excellently captures local features of non-stationary or complex signals.\n  - Suitable for detecting short-duration events, analyzing abrupt changes, and non-linear characteristics within signals.\n- **Weaknesses:**\n  - Requires the selection of an appropriate wavelet function, which can demand specialized knowledge.\n  - Implementation and interpretation can become complex.\n\n### Superlet Transform\n- **Strengths:**\n  - Provides high time-frequency resolution, capturing fine details of the signal.\n  - High capability to distinguish overlapped or short-duration signal components.\n  - Particularly effective for analyzing signals with high non-stationarity, such as complex brain wave patterns.\n- **Weaknesses:**\n  - Relatively new, with potentially limited resources or examples of implementation available.\n  - May incur high computational costs.\n\n### Best Transform for Analyzing High Non-stationary Signals\nTo capture the local characteristics of highly non-stationary signals, **Wavelet Transform** or **Superlet Transform** are particularly suitable. These methods provide high flexibility and adaptability to temporal and frequency changes in the signal, making them effective for analyzing complex, varying signals. Wavelet Transform, with its versatility and local feature extraction capability, is broadly adopted. The Superlet Transform may be chosen for even higher time-frequency resolution needs or when analyzing extremely complex signals.\n```",
    "2744643": "First and foremost, I would like to express my gratitude to the competition host who organized this interesting competition, as well as the Kaggle team. Additionally, I am thankful for my teammates who worked together with me until the end.\n\n# kfuji's part\n## Preprocessing:\n- Using all raw EEG data excluding EKG, taking differences between specific pairs to create 18 types of signals.\n- For each of the 18 types of signals, the continuous wavelet transform (CWT) was applied following the [suguuuu's method](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560#2744628). To increase the variation in input data, I used Paul as the mother wavelet instead of Morlet. I used the implementation of [https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py](https://github.com/tomrunia/PyTorchWavelets/blob/master/wavelets_pytorch/wavelets.py) and [https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation](https://www.kaggle.com/code/anjum48/continuous-wavelet-transform-cwt-in-pytorch#PyTorch-implementation).\n- The scalograms obtained through the CWT will be stacked vertically to form a single 2D image and resized to 512x512 as the final step.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8448924%2F06de133392f73c8ea8e823be09c55c25%2FHMS_CWT_Paul.png?generation=1712722022756620&alt=media)\n\n## Model:\n- maxvit_base_tf_512.in21k_ft_in1k\n\n## Training:\n- During training, one eeg_sub_id was randomly selected from eeg_id for each epoch. The input data consisted of 10,000 frames (50 seconds) starting from the corresponding eeg_label_offset_seconds. Furthermore, the labels assigned to the eeg_sub_id were utilized for training.\n- Lr: 1e-3\n- Scheduelr: CosineAnnealingLR\n- Optimizer: Adan\n- Augmentation: XYMasking, mixup\n- 2stage training\n  - Stage1: 5 epochs (all data) \n  -  Stage2: 15 epochs (10 votes or more)\n\n## Variation:\n- The following three models, trained with variations in input frame length and Paul's parameters, were merged into the final team submission.\n  - 2000frame(10sec), Paul(m=4), CV:0.2475\n  - 5000 frame(25sec), Paul(m=4), CV:0.2309\n  - 5000frame(25sec), Paul(m=16), CV:0.2311\n\n## What didn't work well:\n- Using DOG (Derivative of Gaussian) instead of Paul.\n- Longer input frames (performed well individually but did not contribute to the ensemble).",
    "2744890": "Congratulations ! Well I used to think Cris's method was not the best for eval as per eeg id we might have many labels, only same label parts could be seen as aug, but from your wok seems I am wrong..., I will try to change this strategy to see the diff, thanks for sharing, entmax is also cool!",
    "2744757": "Awesome way of reporting, lots of interesting content. Well done for the win!",
    "2745238": "Congratulations! 🎉 Thanks for sharing your excellent work. In this competition, I didn't make full use of the raw EEG data, but I learned a lot from your detailed analysis.",
    "2744666": "Congratulations on first prize in this competition. Thanks for sharing the details of your solution. ",
    "2756393": "Congrats! Really makes me wonder how much of raw EEG data is actually noise vs relevant information. It seems like depending on the data format, it can be interpreted as either",
    "2754400": "Congrats !!! for securing 1st position",
    "2748002": "excellent work!",
    "2747661": "fantastic work. congrats to your team :)",
    "2747097": "Congratulations on winning competition. ⭐",
    "2746718": "congratulations",
    "2746543": "Congratulations on taking first place. 🎉",
    "2745797": "Congratulations!  Along with the models, I am stunned to find that the basic CV strategy can still lead the 1st place even many top solutions use the high-number voting data for validation along with training.  It is encouraging to stick to the basic CV, which still works for training with the selected data and may hedge a potential shake-up.",
    "2745011": "I have a question， when counting for votes》10 it is per eeg subid or per eeg id？",
    "2764048": ""
  }
}