{
  "id": 492652,
  "title": "5th place solution, team KTMUD",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492652",
  "author_name": "kazumax",
  "post_date": "2024-04-10T13:00:11.237000",
  "votes": 40,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Team : KTMUD ( <a href=\"https://www.kaggle.com/kazumax0720\" target=\"_blank\">@kazumax0720</a> <a href=\"https://www.kaggle.com/hiroitakafumi\" target=\"_blank\">@hiroitakafumi</a> <a href=\"https://www.kaggle.com/maxchen303\" target=\"_blank\">@maxchen303</a> <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a> <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>)  </p>\n<p>First of all, Thank you to the organizers and Congrats to all the participants and winners!</p>\n<h2>Summary</h2>\n<p>Our team merged with the UEMU&amp;T.H.&amp;kazumax teams and MaxChen303 and D.Imanishi during the competition.<br>\nWe ensemble the models developed by each team.<br>\nOur prediciction pipeline is this.<br>\nEach approaches are will added in the comments!</p>\n<ul>\n<li>UEMU's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745803\" target=\"_blank\">Link</a></li>\n<li>T.H.'s part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745305\" target=\"_blank\">Link</a></li>\n<li>kazumax's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745192\" target=\"_blank\">Link</a></li>\n<li>MaxChen303's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745233\" target=\"_blank\">Link</a></li>\n<li>D.Imanishi's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745220\" target=\"_blank\">Link</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F93297936355b66b3d6b10b6f70a6016c%2Fprediction_pipeline.png?generation=1712796229936116&amp;alt=media\" alt=\"\"></p>\n<p>We borrowed the format from the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560\" target=\"_blank\">1st place solution</a>.<br>\nThank you, team SONY!</p>",
  "messages": [
    {
      "id": 2745188,
      "postDate": "2024-04-10T13:00:11.237Z",
      "content": "<p>Team : KTMUD ( <a href=\"https://www.kaggle.com/kazumax0720\" target=\"_blank\">@kazumax0720</a> <a href=\"https://www.kaggle.com/hiroitakafumi\" target=\"_blank\">@hiroitakafumi</a> <a href=\"https://www.kaggle.com/maxchen303\" target=\"_blank\">@maxchen303</a> <a href=\"https://www.kaggle.com/asaliquid1011\" target=\"_blank\">@asaliquid1011</a> <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>)  </p>\n<p>First of all, Thank you to the organizers and Congrats to all the participants and winners!</p>\n<h2>Summary</h2>\n<p>Our team merged with the UEMU&amp;T.H.&amp;kazumax teams and MaxChen303 and D.Imanishi during the competition.<br>\nWe ensemble the models developed by each team.<br>\nOur prediciction pipeline is this.<br>\nEach approaches are will added in the comments!</p>\n<ul>\n<li>UEMU's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745803\" target=\"_blank\">Link</a></li>\n<li>T.H.'s part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745305\" target=\"_blank\">Link</a></li>\n<li>kazumax's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745192\" target=\"_blank\">Link</a></li>\n<li>MaxChen303's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745233\" target=\"_blank\">Link</a></li>\n<li>D.Imanishi's part : <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745220\" target=\"_blank\">Link</a></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F93297936355b66b3d6b10b6f70a6016c%2Fprediction_pipeline.png?generation=1712796229936116&amp;alt=media\" alt=\"\"></p>\n<p>We borrowed the format from the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560\" target=\"_blank\">1st place solution</a>.<br>\nThank you, team SONY!</p>",
      "rawMarkdown": "Team : KTMUD ( @kazumax0720 @hiroitakafumi @maxchen303 @asaliquid1011 @dimanishi)  \n\nFirst of all, Thank you to the organizers and Congrats to all the participants and winners!\n\n## Summary\nOur team merged with the UEMU&T.H.&kazumax teams and MaxChen303 and D.Imanishi during the competition.   \nWe ensemble the models developed by each team.  \nOur prediciction pipeline is this.  \nEach approaches are will added in the comments!\n\n* UEMU's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745803)\n* T.H.'s part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745305)\n* kazumax's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745192)\n* MaxChen303's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745233)\n* D.Imanishi's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745220)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F93297936355b66b3d6b10b6f70a6016c%2Fprediction_pipeline.png?generation=1712796229936116&alt=media)\n\nWe borrowed the format from the [1st place solution](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560).  \nThank you, team SONY!",
      "votes": 39
    },
    {
      "id": 2745220,
      "postDate": "2024-04-10T13:18:33.623Z",
      "content": "<h2>D.Imanishi's part</h2>\n<p>Thanks to the organizers and congrats to all the winners! I think it was a very exciting and hard competition.<br>\nAnd I am very happy to have finally achieved my goal of becoming a grandmaster!</p>\n<h3>Approach</h3>\n<p>2D-CNN Models were trained for two types of input data and ensemble them. </p>\n<ul>\n<li>concatenation of waveform graph image and kaggle spectrogram image (public LB:0.25)</li>\n<li>waveform graph only (public LB:0.25)</li>\n</ul>\n<p>The input image of my model is like following. 19 signals are the same as competition's sample figure.<br>\n<img src=\"https://i.postimg.cc/BQHHgFyw/specgraph.png\"><br>\nI had also taken the approach of converting EEG to spectrogram, but did not use it in the end because the score was better without it in the process of team merging.</p>\n<h3>Training</h3>\n<ul>\n<li>2 stage training (1st: all data, 2nd: votes&gt;=10)</li>\n<li>After training the first model, the second model was trained by applying the pseudo label to votes&lt;10 data using OOF prediction. Instead of replacing it with the pseudo label, it was mixed with the original label according to the number of votes (for samples with fewer votes, increase the weight of the pseudo label).</li>\n<li>3 types of CNN backbones were ensembled (EfficientNetV2B0, EfficientNetV2B1, EfficientNetV1B0)</li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>Swap left and right electrode (Fp1&lt;=&gt;Fp2, F3&lt;=&gt;F4, …)</li>\n<li>Positive and negative swap (data=-data)</li>\n<li>Mixup</li>\n<li>Flipping time axis (kaggle spectrogram only)</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>TTA (Same as training augmentations except mixup)</li>\n<li>Logits averaging (not after softmax) for ensemble</li>\n</ul>\n<h3>Code</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/dimanishi/hms-waveform-image-generation\" target=\"_blank\">Waveform image generation</a></li>\n<li><a href=\"https://www.kaggle.com/code/dimanishi/hms-graph-model-training\" target=\"_blank\">Graph model training</a></li>\n<li><a href=\"https://www.kaggle.com/code/dimanishi/hms-graph-model-inference\" target=\"_blank\">Graph model inference</a></li>\n</ul>",
      "rawMarkdown": "## D.Imanishi's part\nThanks to the organizers and congrats to all the winners! I think it was a very exciting and hard competition.\nAnd I am very happy to have finally achieved my goal of becoming a grandmaster!\n\n### Approach\n2D-CNN Models were trained for two types of input data and ensemble them. \n- concatenation of waveform graph image and kaggle spectrogram image (public LB:0.25)\n- waveform graph only (public LB:0.25)\n\nThe input image of my model is like following. 19 signals are the same as competition's sample figure.\n![](https://i.postimg.cc/BQHHgFyw/specgraph.png)\nI had also taken the approach of converting EEG to spectrogram, but did not use it in the end because the score was better without it in the process of team merging.\n\n### Training\n- 2 stage training (1st: all data, 2nd: votes>=10)\n- After training the first model, the second model was trained by applying the pseudo label to votes<10 data using OOF prediction. Instead of replacing it with the pseudo label, it was mixed with the original label according to the number of votes (for samples with fewer votes, increase the weight of the pseudo label).\n- 3 types of CNN backbones were ensembled (EfficientNetV2B0, EfficientNetV2B1, EfficientNetV1B0)\n\n### Augmentation\n- Swap left and right electrode (Fp1<=>Fp2, F3<=>F4, ...)\n- Positive and negative swap (data=-data)\n- Mixup\n- Flipping time axis (kaggle spectrogram only)\n\n### Inference\n- TTA (Same as training augmentations except mixup)\n- Logits averaging (not after softmax) for ensemble\n\n### Code\n- [Waveform image generation](https://www.kaggle.com/code/dimanishi/hms-waveform-image-generation)\n- [Graph model training](https://www.kaggle.com/code/dimanishi/hms-graph-model-training)\n- [Graph model inference](https://www.kaggle.com/code/dimanishi/hms-graph-model-inference)",
      "votes": 18,
      "replies": [
        {
          "id": 2745827,
          "postDate": "2024-04-10T21:26:06.010Z",
          "content": "<p>Congrats on achieving grandmaster <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>, well deserved!</p>",
          "rawMarkdown": "Congrats on achieving grandmaster @dimanishi, well deserved!",
          "votes": 2
        },
        {
          "id": 2746172,
          "postDate": "2024-04-11T05:07:40.513Z",
          "content": "<p>Congrats on achieving grandmaster <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>, you deserve it!</p>",
          "rawMarkdown": "Congrats on achieving grandmaster @dimanishi, you deserve it!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2745233,
      "postDate": "2024-04-10T13:35:32.363Z",
      "content": "<h2>MaxChen303's part</h2>\n<p>This is the first time I have made a team with others on Kaggle. Thanks to my teammates for bringing the great experience.</p>\n<h3>Summary</h3>\n<p>The main idea of my solution is to generate inputs similar to what the annotators see.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fef831958400ea7348dfe6ecebf6c511a%2FScreenshot%20from%202024-04-09%2012-53-16.png?generation=1712681638219917&amp;alt=media\"></p>\n<ul>\n<li>The size of each part is configurable.<ul>\n<li>Final models use size 512x512.</li>\n<li>The width of raw eeg part is 500 for 10s. (Should be enough for sampling 20Hz signals)</li>\n<li>The height of raw eeg part is controlled by the number of repeat. E.g. 16 channels * 8 repeat = 128 height.</li></ul></li>\n<li>Removing the EEG spectrogram part may not affect the performance.</li>\n</ul>\n<p>Due to a mistake in my HIGH vote CV implementation (accidentally included some low votes samples), I think I didn't fully optimize the performance in the last week.</p>\n<ul>\n<li>The best score of my part with ensemble (3 models) is Private 0.295527/Public 0.247586. </li>\n<li>The best single model score: Private=0.303614, Public = 0.250533</li>\n</ul>\n<h3>EEG Preprocessing</h3>\n<p>After checking the provided EEG pattern images many times, I found that the provided raw EEGs look much cleaner. It was much easier to see the pattern in the provided image than my first few raw EEG plots.</p>\n<p>It turns out that simply applying a bandpass filter can make the raw EEGs look much closer to the provided examples.<br>\nI also noticed that the scipy <code>sosfilt</code> is preferred over the commonly used <code>lfilter</code> according to the <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.lfilter.html\" target=\"_blank\">document</a> and the <code>sosfiltfilt</code> can reduce the effect of phase shift. </p>\n<ul>\n<li><a href=\"https://stackoverflow.com/questions/12093594/how-to-implement-band-pass-butterworth-filter-with-scipy-signal-butter/48677312#48677312\" target=\"_blank\">https://stackoverflow.com/questions/12093594/how-to-implement-band-pass-butterworth-filter-with-scipy-signal-butter/48677312#48677312</a></li>\n<li>Thanks to <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> for sharing this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/483839\" target=\"_blank\">discussion</a></li>\n</ul>\n<p>The red boxes in this diagram show the possible phase shifts:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Ff2e4493929bebfd7149582d9e23e8c0b%2FScreenshot%20from%202024-04-09%2012-56-41.png?generation=1712681817517725&amp;alt=media\"></p>\n<h3>Training</h3>\n<ul>\n<li>One stage training with lower weight for LOW vote samples.<ul>\n<li>Monitor CV score for both All and HIGH vote samples under different weight ratios.</li>\n<li>The final weight selected is 0.25 for low vote samples (&lt;=7) for best ALL and HIGH vote CV score.</li>\n<li>The weight is very close to (3.04/14.29 ~= 0.21), which is <code>average # LOW votes / average # HIGH votes</code>.</li></ul></li>\n<li>Exponential Moving Average (EMA) did well in stabilizing the validation loss during training. </li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>Flip Left and Right components</li>\n<li>Horizontal flip for spectrograms</li>\n<li>Time and Frequency mask for the spectrograms</li>\n<li>Time and Channel mask for the raw EEG. (Randomly hide 4-6 channels in the raw EEG)</li>\n<li>Time shift of 10s Raw EEG</li>\n<li>Raw EEG * -1 </li>\n<li>Small random contrast adjustment for the raw eeg.</li>\n</ul>\n<p>Test time augmentation(TTA):</p>\n<ul>\n<li>Flip Left and Right components.</li>\n</ul>\n<h3>Model</h3>\n<p>Pretrained backbones from TIMM with different pooling layers.</p>\n<ul>\n<li>Models used in final submission:<ul>\n<li>maxvit_small_tf_512, </li>\n<li>maxvit_tiny_tf_512,</li>\n<li>convnext_small</li></ul></li>\n<li>Average pooling performs the best in my experiments.</li>\n<li>LayerNorm -&gt; Dropout -&gt; FC head</li>\n</ul>",
      "rawMarkdown": "## MaxChen303's part\nThis is the first time I have made a team with others on Kaggle. Thanks to my teammates for bringing the great experience.\n### Summary\nThe main idea of my solution is to generate inputs similar to what the annotators see.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fef831958400ea7348dfe6ecebf6c511a%2FScreenshot%20from%202024-04-09%2012-53-16.png?generation=1712681638219917&alt=media)\n- The size of each part is configurable.\n    - Final models use size 512x512.\n    - The width of raw eeg part is 500 for 10s. (Should be enough for sampling 20Hz signals)\n    - The height of raw eeg part is controlled by the number of repeat. E.g. 16 channels * 8 repeat = 128 height.\n- Removing the EEG spectrogram part may not affect the performance.\n\nDue to a mistake in my HIGH vote CV implementation (accidentally included some low votes samples), I think I didn't fully optimize the performance in the last week.\n- The best score of my part with ensemble (3 models) is Private 0.295527/Public 0.247586. \n- The best single model score: Private=0.303614, Public = 0.250533\n\n### EEG Preprocessing \nAfter checking the provided EEG pattern images many times, I found that the provided raw EEGs look much cleaner. It was much easier to see the pattern in the provided image than my first few raw EEG plots.\n\nIt turns out that simply applying a bandpass filter can make the raw EEGs look much closer to the provided examples.\nI also noticed that the scipy `sosfilt` is preferred over the commonly used `lfilter` according to the [document](https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.lfilter.html) and the `sosfiltfilt` can reduce the effect of phase shift. \n\n- https://stackoverflow.com/questions/12093594/how-to-implement-band-pass-butterworth-filter-with-scipy-signal-butter/48677312#48677312\n- Thanks to @gunesevitan for sharing this [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/483839)\n\nThe red boxes in this diagram show the possible phase shifts:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Ff2e4493929bebfd7149582d9e23e8c0b%2FScreenshot%20from%202024-04-09%2012-56-41.png?generation=1712681817517725&alt=media)\n\n\n### Training\n- One stage training with lower weight for LOW vote samples.\n    - Monitor CV score for both All and HIGH vote samples under different weight ratios.\n    - The final weight selected is 0.25 for low vote samples (<=7) for best ALL and HIGH vote CV score.\n    - The weight is very close to (3.04/14.29 ~= 0.21), which is `average # LOW votes / average # HIGH votes`.\n- Exponential Moving Average (EMA) did well in stabilizing the validation loss during training. \n\n### Augmentation\n- Flip Left and Right components\n- Horizontal flip for spectrograms\n- Time and Frequency mask for the spectrograms\n- Time and Channel mask for the raw EEG. (Randomly hide 4-6 channels in the raw EEG)\n- Time shift of 10s Raw EEG\n- Raw EEG * -1 \n- Small random contrast adjustment for the raw eeg.\n\nTest time augmentation(TTA):\n- Flip Left and Right components.\n\n### Model\nPretrained backbones from TIMM with different pooling layers.\n- Models used in final submission:\n    - maxvit_small_tf_512, \n    - maxvit_tiny_tf_512,\n    - convnext_small\n- Average pooling performs the best in my experiments.\n- LayerNorm -> Dropout -> FC head\n",
      "votes": 13
    },
    {
      "id": 2745305,
      "postDate": "2024-04-10T14:33:43.407Z",
      "content": "<h2>T.H.'s part</h2>\n<p>UEMU &amp; T.H. &amp; kazumax created models as a mini team.  <br>\nI was in charge of the raw EEG part.  <br>\nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.</p>\n<h3>pipeline</h3>\n<p>I developed a wavenet+2dCNN model using only raw EEG.  <br>\nBelow is an outline of my pipeline.  <br>\nI achieved a public LB = 0.25 (cv = 0.244) with a single model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2449267%2Ff01ca637106a7b304d97a892205fbabc%2Fth_pipeline.png?generation=1712759565877010&amp;alt=media\" alt=\"th_pipeline\"></p>\n<h3>Input</h3>\n<ul>\n<li>Compute a total of 16 signals by calculating 4 neighboring difference signals each for LL, RL, LP, RP.</li>\n<li>20Hz butter lowpass filter.</li>\n<li>Clipping of signal intensity.</li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>Wavenet+2DCNN.</li>\n<li>For Wavenet, I referred to this <a href=\"https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train\" target=\"_blank\">notebook</a>. Thank you for sharing the code!</li>\n<li>Input 16 signals separately into Wavenet and average the output feature maps within the groups of LL, RL, LP, RP.</li>\n<li>Stack the four feature maps vertically to input into 2DCNN.</li>\n<li>For the 2DCNN model, MaxxvitV2-nano was the most accurate. Maxxvit-small and effnetb4 were also used for the ensemble.</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>2-stage training:<ul>\n<li>1st stage: train with all data for 50 epochs (lr=1e-3).</li>\n<li>2nd stage: fine-tune with data where n_votes&gt;=10 for 30 epochs (lr=1e-4).</li></ul></li>\n<li>Optimizer: AdamW.</li>\n<li>Scheduler: Cosine Annealing (+Warmup).</li>\n<li>Augmentations (in order of most effective):<ul>\n<li>ChannelSwap ([LL/LP] &lt;=&gt; [RL/RP]).</li>\n<li>Inversion (inversion of signal strength).</li>\n<li>TimeMask.</li>\n<li>Reverse (time reversal of the signal).</li></ul></li>\n</ul>",
      "rawMarkdown": "## T.H.'s part\nUEMU & T.H. & kazumax created models as a mini team.  \nI was in charge of the raw EEG part.  \nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.\n\n### pipeline\nI developed a wavenet+2dCNN model using only raw EEG.  \nBelow is an outline of my pipeline.  \nI achieved a public LB = 0.25 (cv = 0.244) with a single model.\n\n![th_pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2449267%2Ff01ca637106a7b304d97a892205fbabc%2Fth_pipeline.png?generation=1712759565877010&alt=media)\n\n### Input\n\n- Compute a total of 16 signals by calculating 4 neighboring difference signals each for LL, RL, LP, RP.\n- 20Hz butter lowpass filter.\n- Clipping of signal intensity.\n\n### Model\n\n- Wavenet+2DCNN.\n- For Wavenet, I referred to this [notebook](https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train \"HMS | WaveNet PyTorch [Train]\"). Thank you for sharing the code!\n- Input 16 signals separately into Wavenet and average the output feature maps within the groups of LL, RL, LP, RP.\n- Stack the four feature maps vertically to input into 2DCNN.\n- For the 2DCNN model, MaxxvitV2-nano was the most accurate. Maxxvit-small and effnetb4 were also used for the ensemble.\n\n### Training\n\n- 2-stage training:\n  - 1st stage: train with all data for 50 epochs (lr=1e-3).\n  - 2nd stage: fine-tune with data where n_votes>=10 for 30 epochs (lr=1e-4).\n- Optimizer: AdamW.\n- Scheduler: Cosine Annealing (+Warmup).\n- Augmentations (in order of most effective):\n  - ChannelSwap ([LL/LP] <=> [RL/RP]).\n  - Inversion (inversion of signal strength).\n  - TimeMask.\n  - Reverse (time reversal of the signal).",
      "votes": 12,
      "replies": [
        {
          "id": 2753869,
          "postDate": "2024-04-15T17:26:10.563Z",
          "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/hiroitakafumi\" target=\"_blank\">@hiroitakafumi</a> and congrats on achieving 5th place. Is it possible you share your training and inference code?</p>",
          "rawMarkdown": "thanks for sharing @hiroitakafumi and congrats on achieving 5th place. Is it possible you share your training and inference code?"
        },
        {
          "id": 2755394,
          "postDate": "2024-04-16T13:48:38.570Z",
          "content": "<p>Thanks for sharing.<br>\nWhy did you add a CNN Block between WaveBlock?<br>\nI just want to know the logic to do that.</p>",
          "rawMarkdown": "Thanks for sharing.\nWhy did you add a CNN Block between WaveBlock?\nI just want to know the logic to do that.",
          "replies": [
            {
              "id": 2756287,
              "postDate": "2024-04-17T01:11:40.680Z",
              "content": "<p>I wanted to reduce the width of the Wavenet output by including a CNN with stride=2.<br>\nWithout CNN, the output shape of Wavenet is [64, 10000], but with CNN it becomes [64, 625]. <br>\nThe aim was to improve accuracy by lowering the downsampling rate due to pooling layer.<br>\nPublic LB improved by about 0.01.</p>",
              "rawMarkdown": "I wanted to reduce the width of the Wavenet output by including a CNN with stride=2.\nWithout CNN, the output shape of Wavenet is [64, 10000], but with CNN it becomes [64, 625]. \nThe aim was to improve accuracy by lowering the downsampling rate due to pooling layer.\nPublic LB improved by about 0.01."
            }
          ]
        }
      ]
    },
    {
      "id": 2745192,
      "postDate": "2024-04-10T13:04:19.773Z",
      "content": "<h2>kazumax's part</h2>\n<p>UEMU &amp; T.H. &amp; kazumax created models as a mini team.  <br>\nI was in charge of the spectrogram part.  <br>\nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.</p>\n<h3>Input</h3>\n<ul>\n<li><p>24 spectrograms</p>\n<ul>\n<li>4 KaggleSpec (LL,LR,RP,RL)</li>\n<li>20 EEGSpec  <ul>\n<li>16 specs: Adjacent Difference (Fp1-F7, F7-T3, …) </li>\n<li>4 specs (LL,LP,RP,RL)  </li></ul></li></ul></li>\n<li><p>EEGSpec is a power spectrogram truncated at 20Hz not mel spectrogram.</p>\n<pre><code>spec = librosa.stft(\n    y=x,\n    hop_length=(x) // ,\n    n_fft=,\n)\nspec = np.(spec) ** \n\n\n\n\nspec = spec[:, :] \n</code></pre></li>\n</ul>\n<h3>Model</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2Fdc133f535c8f382e26ed9a77956e3f62%2Fmodel.png?generation=1712754127506152&amp;alt=media\"></p>\n<h3>Training</h3>\n<ul>\n<li>2stage training<ul>\n<li>1st stage: all data</li>\n<li>2nd stage: vote &gt;= 10</li></ul></li>\n<li>Augmentation<ul>\n<li>TimeFreqMask(p=0.5)</li>\n<li>swap left-right specs(p=0.5)<ul>\n<li>LL ↔ RL, LR ↔ RP</li>\n<li>Fp1 - F7 ↔ FP2 - F8, …</li></ul></li></ul></li>\n<li>label smoothing<ul>\n<li>epsilon=0.1 for samples with vote &lt;= 3</li></ul></li>\n</ul>",
      "rawMarkdown": "## kazumax's part\n\nUEMU & T.H. & kazumax created models as a mini team.  \nI was in charge of the spectrogram part.  \nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.\n### Input\n* 24 spectrograms\n    * 4 KaggleSpec (LL,LR,RP,RL)\n    * 20 EEGSpec  \n        * 16 specs: Adjacent Difference (Fp1-F7, F7-T3, …) \n        * 4 specs (LL,LP,RP,RL)  \n\n* EEGSpec is a power spectrogram truncated at 20Hz not mel spectrogram.\n    ```python\n    spec = librosa.stft(\n        y=x,\n        hop_length=len(x) // 256,\n        n_fft=1024,\n    )\n    spec = np.abs(spec) ** 2\n\n    # truncate at 20Hz.\n    # frequency resoluion = sr / n_fft = 200[Hz] / 1024 = 0.1953[Hz]\n    # 0.1953[Hz] * 100 ≃ 20[Hz]\n    spec = spec[:100, :] \n    ```\n\n### Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2Fdc133f535c8f382e26ed9a77956e3f62%2Fmodel.png?generation=1712754127506152&alt=media)\n### Training\n* 2stage training\n    * 1st stage: all data\n    * 2nd stage: vote >= 10\n* Augmentation\n    * TimeFreqMask(p=0.5)\n    * swap left-right specs(p=0.5)\n        * LL ↔ RL, LR ↔ RP\n        * Fp1 - F7 ↔ FP2 - F8, ...\n* label smoothing\n    * epsilon=0.1 for samples with vote <= 3",
      "votes": 12,
      "replies": [
        {
          "id": 2745217,
          "postDate": "2024-04-10T13:17:54.140Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2745803,
      "postDate": "2024-04-10T20:41:25.840Z",
      "content": "<h2>UEMU's part</h2>\n<h3>1.Multi Input Model</h3>\n<ul>\n<li>input = Kaggle_Spec + EEG_Spec + Raw_RRG</li>\n<li>model = large model(2DCNN + 1DCNN + MLP)</li>\n<li>CV(single)    = 0.2436<br>\n(update later)</li>\n</ul>\n<h3>2.UEMU&amp;kazumax&amp;TH Ensemble</h3>\n<h4>2_1.Integrated model for Each Data Type</h4>\n<p>Our team (UEMU&amp;kazumax&amp;TH) developed models, each member focusing on a different data type.<br>\n However, in this competition, it's necessary to integrate the inference values of each data type, similar to what the annotators are doing.<br>\n Therefore, we created a model that integrates the models for each data type through a fully connected (FC) layer to provide a multifaceted perspective.<br>\n We created some models of each data type for ensemble, and we also developed this model for all combinations.</p>\n<ul>\n<li>Feature<ul>\n<li>prediction of Raw EEG Models(T.H. Part)</li>\n<li>prediction of Multi Spec Models(kazumax Part)</li>\n<li>prediction of EEG Spec  Models(UEMU Part)</li>\n<li>prediction of Multi Input Models(UEMU Part)</li></ul></li>\n<li>Model<ul>\n<li>multi-layer fully connected (FC) layer (No activation function is used)</li></ul></li>\n</ul>\n<h4>2_2.PostProcess</h4>\n<p>As mentioned in this discussion, there are some data where the eeg_id is different but the patient_id is the same in the test dataset.<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471890\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471890</a></p>\n<p>We aggregated the prediction values for each patient_id to create new features.<br>\nUsing the original prediction values and these aggregated features, we built a another inference model.<br>\nThis model did not improve the score on its own, but by including this model in the 2_3.Weight Average, we achieved an improvement in CV of 0.002 to 0.003.</p>\n<ul>\n<li>Feature<ul>\n<li>Original prediction<ul>\n<li>Average prediction of Raw EEG Models(T.H. Part)</li>\n<li>Average prediction of Multi Spec Models(kazumax Part)</li>\n<li>Average prediction of Multi Input Models(UEMU Part)</li></ul></li>\n<li>Aggregated feature<ul>\n<li>Average/Max/Min/Std/Delta/Count of original prediction groupby patient_id</li></ul></li></ul></li>\n<li>Model<ul>\n<li>multi-layer fully connected (FC) layer (No activation function is used)</li></ul></li>\n</ul>\n<h4>2_3.Weight Average All Models</h4>\n<p>We took a weighted average using all the models(each single model,2_1.FC model and the 2_2.PostProcess, about 50 models).<br>\nThe weights were optimized based on CV. Although it seemed excessive, the correlation between the LB and CV was good.</p>\n<p>The best scores(UEMU&amp;kazumax&amp;TH's part)</p>\n<ul>\n<li>CV      = 0.204</li>\n<li>public  = 0.2382</li>\n<li>private = 0.2859(≒10th place)</li>\n</ul>\n<h3>3.Team  KTMUD Ensemble</h3>\n<p>Simple weighted average of three teams.<br>\nThe weighted average was taken using the logits before applying the softmax function.<br>\nThe weights were decided based on the public leaderboard (LB), where the public best score was equal to the private best score.</p>\n<ul>\n<li>weights = [D.Imanishi , MaxChen, and UEMU&amp;kazumax&amp;TH] = [0.28, 0.28, 0.44] </li>\n</ul>\n<p>The best scores(Team KTMUD)</p>\n<ul>\n<li>public  = 0.2258</li>\n<li>private = 0.2761</li>\n</ul>",
      "rawMarkdown": "## UEMU's part\n\n### 1.Multi Input Model \n  - input = Kaggle_Spec + EEG_Spec + Raw_RRG\n  - model = large model(2DCNN + 1DCNN + MLP)\n  - CV(single)    = 0.2436\n  (update later)\n\n### 2.UEMU&kazumax&TH Ensemble\n#### 2_1.Integrated model for Each Data Type\n Our team (UEMU&kazumax&TH) developed models, each member focusing on a different data type.\n However, in this competition, it's necessary to integrate the inference values of each data type, similar to what the annotators are doing.\n Therefore, we created a model that integrates the models for each data type through a fully connected (FC) layer to provide a multifaceted perspective.\n We created some models of each data type for ensemble, and we also developed this model for all combinations.\n  - Feature\n    - prediction of Raw EEG Models(T.H. Part)\n    - prediction of Multi Spec Models(kazumax Part)\n    - prediction of EEG Spec  Models(UEMU Part)\n    - prediction of Multi Input Models(UEMU Part)\n  - Model\n    - multi-layer fully connected (FC) layer (No activation function is used)\n \n#### 2_2.PostProcess\nAs mentioned in this discussion, there are some data where the eeg_id is different but the patient_id is the same in the test dataset.\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471890\n\nWe aggregated the prediction values for each patient_id to create new features.\nUsing the original prediction values and these aggregated features, we built a another inference model.\nThis model did not improve the score on its own, but by including this model in the 2_3.Weight Average, we achieved an improvement in CV of 0.002 to 0.003.\n\n  - Feature\n    - Original prediction\n      - Average prediction of Raw EEG Models(T.H. Part)\n      - Average prediction of Multi Spec Models(kazumax Part)\n      - Average prediction of Multi Input Models(UEMU Part)\n    - Aggregated feature\n      - Average/Max/Min/Std/Delta/Count of original prediction groupby patient_id\n  - Model\n    - multi-layer fully connected (FC) layer (No activation function is used)\n\n\n#### 2_3.Weight Average All Models\nWe took a weighted average using all the models(each single model,2_1.FC model and the 2_2.PostProcess, about 50 models).\nThe weights were optimized based on CV. Although it seemed excessive, the correlation between the LB and CV was good.\n\nThe best scores(UEMU&kazumax&TH's part)\n - CV      = 0.204\n - public  = 0.2382\n - private = 0.2859(≒10th place)\n\n### 3.Team  KTMUD Ensemble\nSimple weighted average of three teams.\nThe weighted average was taken using the logits before applying the softmax function.\nThe weights were decided based on the public leaderboard (LB), where the public best score was equal to the private best score.\n - weights = [D.Imanishi , MaxChen, and UEMU&kazumax&TH] = [0.28, 0.28, 0.44] \n\nThe best scores(Team KTMUD)\n - public  = 0.2258\n - private = 0.2761",
      "votes": 9
    },
    {
      "id": 2748281,
      "postDate": "2024-04-12T10:54:10.677Z",
      "content": "<p>Congratulations on achieving 5th place in this competition. Thanks for sharing the insights into your solution. </p>",
      "rawMarkdown": "Congratulations on achieving 5th place in this competition. Thanks for sharing the insights into your solution. "
    }
  ],
  "comments": [
    {
      "id": 2745220,
      "author_name": "D.Imanishi",
      "author_url": "",
      "post_date": "2024-04-10T13:18:33.623000",
      "content": "<h2>D.Imanishi's part</h2>\n<p>Thanks to the organizers and congrats to all the winners! I think it was a very exciting and hard competition.<br>\nAnd I am very happy to have finally achieved my goal of becoming a grandmaster!</p>\n<h3>Approach</h3>\n<p>2D-CNN Models were trained for two types of input data and ensemble them. </p>\n<ul>\n<li>concatenation of waveform graph image and kaggle spectrogram image (public LB:0.25)</li>\n<li>waveform graph only (public LB:0.25)</li>\n</ul>\n<p>The input image of my model is like following. 19 signals are the same as competition's sample figure.<br>\n<img src=\"https://i.postimg.cc/BQHHgFyw/specgraph.png\"><br>\nI had also taken the approach of converting EEG to spectrogram, but did not use it in the end because the score was better without it in the process of team merging.</p>\n<h3>Training</h3>\n<ul>\n<li>2 stage training (1st: all data, 2nd: votes&gt;=10)</li>\n<li>After training the first model, the second model was trained by applying the pseudo label to votes&lt;10 data using OOF prediction. Instead of replacing it with the pseudo label, it was mixed with the original label according to the number of votes (for samples with fewer votes, increase the weight of the pseudo label).</li>\n<li>3 types of CNN backbones were ensembled (EfficientNetV2B0, EfficientNetV2B1, EfficientNetV1B0)</li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>Swap left and right electrode (Fp1&lt;=&gt;Fp2, F3&lt;=&gt;F4, …)</li>\n<li>Positive and negative swap (data=-data)</li>\n<li>Mixup</li>\n<li>Flipping time axis (kaggle spectrogram only)</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>TTA (Same as training augmentations except mixup)</li>\n<li>Logits averaging (not after softmax) for ensemble</li>\n</ul>\n<h3>Code</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/dimanishi/hms-waveform-image-generation\" target=\"_blank\">Waveform image generation</a></li>\n<li><a href=\"https://www.kaggle.com/code/dimanishi/hms-graph-model-training\" target=\"_blank\">Graph model training</a></li>\n<li><a href=\"https://www.kaggle.com/code/dimanishi/hms-graph-model-inference\" target=\"_blank\">Graph model inference</a></li>\n</ul>",
      "votes": 18,
      "replies": [
        {
          "id": 2745827,
          "author_name": "Romain Hardy",
          "author_url": "",
          "post_date": "2024-04-10T21:26:06.010000",
          "content": "<p>Congrats on achieving grandmaster <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>, well deserved!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2746172,
          "author_name": "Aaditya Porwal",
          "author_url": "",
          "post_date": "2024-04-11T05:07:40.513000",
          "content": "<p>Congrats on achieving grandmaster <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>, you deserve it!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2745233,
      "author_name": "MaxChen303",
      "author_url": "",
      "post_date": "2024-04-10T13:35:32.363000",
      "content": "<h2>MaxChen303's part</h2>\n<p>This is the first time I have made a team with others on Kaggle. Thanks to my teammates for bringing the great experience.</p>\n<h3>Summary</h3>\n<p>The main idea of my solution is to generate inputs similar to what the annotators see.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fef831958400ea7348dfe6ecebf6c511a%2FScreenshot%20from%202024-04-09%2012-53-16.png?generation=1712681638219917&amp;alt=media\"></p>\n<ul>\n<li>The size of each part is configurable.<ul>\n<li>Final models use size 512x512.</li>\n<li>The width of raw eeg part is 500 for 10s. (Should be enough for sampling 20Hz signals)</li>\n<li>The height of raw eeg part is controlled by the number of repeat. E.g. 16 channels * 8 repeat = 128 height.</li></ul></li>\n<li>Removing the EEG spectrogram part may not affect the performance.</li>\n</ul>\n<p>Due to a mistake in my HIGH vote CV implementation (accidentally included some low votes samples), I think I didn't fully optimize the performance in the last week.</p>\n<ul>\n<li>The best score of my part with ensemble (3 models) is Private 0.295527/Public 0.247586. </li>\n<li>The best single model score: Private=0.303614, Public = 0.250533</li>\n</ul>\n<h3>EEG Preprocessing</h3>\n<p>After checking the provided EEG pattern images many times, I found that the provided raw EEGs look much cleaner. It was much easier to see the pattern in the provided image than my first few raw EEG plots.</p>\n<p>It turns out that simply applying a bandpass filter can make the raw EEGs look much closer to the provided examples.<br>\nI also noticed that the scipy <code>sosfilt</code> is preferred over the commonly used <code>lfilter</code> according to the <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.lfilter.html\" target=\"_blank\">document</a> and the <code>sosfiltfilt</code> can reduce the effect of phase shift. </p>\n<ul>\n<li><a href=\"https://stackoverflow.com/questions/12093594/how-to-implement-band-pass-butterworth-filter-with-scipy-signal-butter/48677312#48677312\" target=\"_blank\">https://stackoverflow.com/questions/12093594/how-to-implement-band-pass-butterworth-filter-with-scipy-signal-butter/48677312#48677312</a></li>\n<li>Thanks to <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> for sharing this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/483839\" target=\"_blank\">discussion</a></li>\n</ul>\n<p>The red boxes in this diagram show the possible phase shifts:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Ff2e4493929bebfd7149582d9e23e8c0b%2FScreenshot%20from%202024-04-09%2012-56-41.png?generation=1712681817517725&amp;alt=media\"></p>\n<h3>Training</h3>\n<ul>\n<li>One stage training with lower weight for LOW vote samples.<ul>\n<li>Monitor CV score for both All and HIGH vote samples under different weight ratios.</li>\n<li>The final weight selected is 0.25 for low vote samples (&lt;=7) for best ALL and HIGH vote CV score.</li>\n<li>The weight is very close to (3.04/14.29 ~= 0.21), which is <code>average # LOW votes / average # HIGH votes</code>.</li></ul></li>\n<li>Exponential Moving Average (EMA) did well in stabilizing the validation loss during training. </li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>Flip Left and Right components</li>\n<li>Horizontal flip for spectrograms</li>\n<li>Time and Frequency mask for the spectrograms</li>\n<li>Time and Channel mask for the raw EEG. (Randomly hide 4-6 channels in the raw EEG)</li>\n<li>Time shift of 10s Raw EEG</li>\n<li>Raw EEG * -1 </li>\n<li>Small random contrast adjustment for the raw eeg.</li>\n</ul>\n<p>Test time augmentation(TTA):</p>\n<ul>\n<li>Flip Left and Right components.</li>\n</ul>\n<h3>Model</h3>\n<p>Pretrained backbones from TIMM with different pooling layers.</p>\n<ul>\n<li>Models used in final submission:<ul>\n<li>maxvit_small_tf_512, </li>\n<li>maxvit_tiny_tf_512,</li>\n<li>convnext_small</li></ul></li>\n<li>Average pooling performs the best in my experiments.</li>\n<li>LayerNorm -&gt; Dropout -&gt; FC head</li>\n</ul>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 2745305,
      "author_name": "T.H.",
      "author_url": "",
      "post_date": "2024-04-10T14:33:43.407000",
      "content": "<h2>T.H.'s part</h2>\n<p>UEMU &amp; T.H. &amp; kazumax created models as a mini team.  <br>\nI was in charge of the raw EEG part.  <br>\nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.</p>\n<h3>pipeline</h3>\n<p>I developed a wavenet+2dCNN model using only raw EEG.  <br>\nBelow is an outline of my pipeline.  <br>\nI achieved a public LB = 0.25 (cv = 0.244) with a single model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2449267%2Ff01ca637106a7b304d97a892205fbabc%2Fth_pipeline.png?generation=1712759565877010&amp;alt=media\" alt=\"th_pipeline\"></p>\n<h3>Input</h3>\n<ul>\n<li>Compute a total of 16 signals by calculating 4 neighboring difference signals each for LL, RL, LP, RP.</li>\n<li>20Hz butter lowpass filter.</li>\n<li>Clipping of signal intensity.</li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>Wavenet+2DCNN.</li>\n<li>For Wavenet, I referred to this <a href=\"https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train\" target=\"_blank\">notebook</a>. Thank you for sharing the code!</li>\n<li>Input 16 signals separately into Wavenet and average the output feature maps within the groups of LL, RL, LP, RP.</li>\n<li>Stack the four feature maps vertically to input into 2DCNN.</li>\n<li>For the 2DCNN model, MaxxvitV2-nano was the most accurate. Maxxvit-small and effnetb4 were also used for the ensemble.</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>2-stage training:<ul>\n<li>1st stage: train with all data for 50 epochs (lr=1e-3).</li>\n<li>2nd stage: fine-tune with data where n_votes&gt;=10 for 30 epochs (lr=1e-4).</li></ul></li>\n<li>Optimizer: AdamW.</li>\n<li>Scheduler: Cosine Annealing (+Warmup).</li>\n<li>Augmentations (in order of most effective):<ul>\n<li>ChannelSwap ([LL/LP] &lt;=&gt; [RL/RP]).</li>\n<li>Inversion (inversion of signal strength).</li>\n<li>TimeMask.</li>\n<li>Reverse (time reversal of the signal).</li></ul></li>\n</ul>",
      "votes": 12,
      "replies": [
        {
          "id": 2753869,
          "author_name": "TooNovice",
          "author_url": "",
          "post_date": "2024-04-15T17:26:10.563000",
          "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/hiroitakafumi\" target=\"_blank\">@hiroitakafumi</a> and congrats on achieving 5th place. Is it possible you share your training and inference code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2755394,
          "author_name": "MBOOK",
          "author_url": "",
          "post_date": "2024-04-16T13:48:38.570000",
          "content": "<p>Thanks for sharing.<br>\nWhy did you add a CNN Block between WaveBlock?<br>\nI just want to know the logic to do that.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2756287,
              "author_name": "T.H.",
              "author_url": "",
              "post_date": "2024-04-17T01:11:40.680000",
              "content": "<p>I wanted to reduce the width of the Wavenet output by including a CNN with stride=2.<br>\nWithout CNN, the output shape of Wavenet is [64, 10000], but with CNN it becomes [64, 625]. <br>\nThe aim was to improve accuracy by lowering the downsampling rate due to pooling layer.<br>\nPublic LB improved by about 0.01.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2745192,
      "author_name": "kazumax",
      "author_url": "",
      "post_date": "2024-04-10T13:04:19.773000",
      "content": "<h2>kazumax's part</h2>\n<p>UEMU &amp; T.H. &amp; kazumax created models as a mini team.  <br>\nI was in charge of the spectrogram part.  <br>\nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.</p>\n<h3>Input</h3>\n<ul>\n<li><p>24 spectrograms</p>\n<ul>\n<li>4 KaggleSpec (LL,LR,RP,RL)</li>\n<li>20 EEGSpec  <ul>\n<li>16 specs: Adjacent Difference (Fp1-F7, F7-T3, …) </li>\n<li>4 specs (LL,LP,RP,RL)  </li></ul></li></ul></li>\n<li><p>EEGSpec is a power spectrogram truncated at 20Hz not mel spectrogram.</p>\n<pre><code>spec = librosa.stft(\n    y=x,\n    hop_length=(x) // ,\n    n_fft=,\n)\nspec = np.(spec) ** \n\n\n\n\nspec = spec[:, :] \n</code></pre></li>\n</ul>\n<h3>Model</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2Fdc133f535c8f382e26ed9a77956e3f62%2Fmodel.png?generation=1712754127506152&amp;alt=media\"></p>\n<h3>Training</h3>\n<ul>\n<li>2stage training<ul>\n<li>1st stage: all data</li>\n<li>2nd stage: vote &gt;= 10</li></ul></li>\n<li>Augmentation<ul>\n<li>TimeFreqMask(p=0.5)</li>\n<li>swap left-right specs(p=0.5)<ul>\n<li>LL ↔ RL, LR ↔ RP</li>\n<li>Fp1 - F7 ↔ FP2 - F8, …</li></ul></li></ul></li>\n<li>label smoothing<ul>\n<li>epsilon=0.1 for samples with vote &lt;= 3</li></ul></li>\n</ul>",
      "votes": 12,
      "replies": [
        {
          "id": 2745217,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-10T13:17:54.140000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2745803,
      "author_name": "UEMU",
      "author_url": "",
      "post_date": "2024-04-10T20:41:25.840000",
      "content": "<h2>UEMU's part</h2>\n<h3>1.Multi Input Model</h3>\n<ul>\n<li>input = Kaggle_Spec + EEG_Spec + Raw_RRG</li>\n<li>model = large model(2DCNN + 1DCNN + MLP)</li>\n<li>CV(single)    = 0.2436<br>\n(update later)</li>\n</ul>\n<h3>2.UEMU&amp;kazumax&amp;TH Ensemble</h3>\n<h4>2_1.Integrated model for Each Data Type</h4>\n<p>Our team (UEMU&amp;kazumax&amp;TH) developed models, each member focusing on a different data type.<br>\n However, in this competition, it's necessary to integrate the inference values of each data type, similar to what the annotators are doing.<br>\n Therefore, we created a model that integrates the models for each data type through a fully connected (FC) layer to provide a multifaceted perspective.<br>\n We created some models of each data type for ensemble, and we also developed this model for all combinations.</p>\n<ul>\n<li>Feature<ul>\n<li>prediction of Raw EEG Models(T.H. Part)</li>\n<li>prediction of Multi Spec Models(kazumax Part)</li>\n<li>prediction of EEG Spec  Models(UEMU Part)</li>\n<li>prediction of Multi Input Models(UEMU Part)</li></ul></li>\n<li>Model<ul>\n<li>multi-layer fully connected (FC) layer (No activation function is used)</li></ul></li>\n</ul>\n<h4>2_2.PostProcess</h4>\n<p>As mentioned in this discussion, there are some data where the eeg_id is different but the patient_id is the same in the test dataset.<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471890\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471890</a></p>\n<p>We aggregated the prediction values for each patient_id to create new features.<br>\nUsing the original prediction values and these aggregated features, we built a another inference model.<br>\nThis model did not improve the score on its own, but by including this model in the 2_3.Weight Average, we achieved an improvement in CV of 0.002 to 0.003.</p>\n<ul>\n<li>Feature<ul>\n<li>Original prediction<ul>\n<li>Average prediction of Raw EEG Models(T.H. Part)</li>\n<li>Average prediction of Multi Spec Models(kazumax Part)</li>\n<li>Average prediction of Multi Input Models(UEMU Part)</li></ul></li>\n<li>Aggregated feature<ul>\n<li>Average/Max/Min/Std/Delta/Count of original prediction groupby patient_id</li></ul></li></ul></li>\n<li>Model<ul>\n<li>multi-layer fully connected (FC) layer (No activation function is used)</li></ul></li>\n</ul>\n<h4>2_3.Weight Average All Models</h4>\n<p>We took a weighted average using all the models(each single model,2_1.FC model and the 2_2.PostProcess, about 50 models).<br>\nThe weights were optimized based on CV. Although it seemed excessive, the correlation between the LB and CV was good.</p>\n<p>The best scores(UEMU&amp;kazumax&amp;TH's part)</p>\n<ul>\n<li>CV      = 0.204</li>\n<li>public  = 0.2382</li>\n<li>private = 0.2859(≒10th place)</li>\n</ul>\n<h3>3.Team  KTMUD Ensemble</h3>\n<p>Simple weighted average of three teams.<br>\nThe weighted average was taken using the logits before applying the softmax function.<br>\nThe weights were decided based on the public leaderboard (LB), where the public best score was equal to the private best score.</p>\n<ul>\n<li>weights = [D.Imanishi , MaxChen, and UEMU&amp;kazumax&amp;TH] = [0.28, 0.28, 0.44] </li>\n</ul>\n<p>The best scores(Team KTMUD)</p>\n<ul>\n<li>public  = 0.2258</li>\n<li>private = 0.2761</li>\n</ul>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 2748281,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-04-12T10:54:10.677000",
      "content": "<p>Congratulations on achieving 5th place in this competition. Thanks for sharing the insights into your solution. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2745188": "Team : KTMUD ( @kazumax0720 @hiroitakafumi @maxchen303 @asaliquid1011 @dimanishi)  \n\nFirst of all, Thank you to the organizers and Congrats to all the participants and winners!\n\n## Summary\nOur team merged with the UEMU&T.H.&kazumax teams and MaxChen303 and D.Imanishi during the competition.   \nWe ensemble the models developed by each team.  \nOur prediciction pipeline is this.  \nEach approaches are will added in the comments!\n\n* UEMU's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745803)\n* T.H.'s part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745305)\n* kazumax's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745192)\n* MaxChen303's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745233)\n* D.Imanishi's part : [Link](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492652#2745220)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F93297936355b66b3d6b10b6f70a6016c%2Fprediction_pipeline.png?generation=1712796229936116&alt=media)\n\nWe borrowed the format from the [1st place solution](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560).  \nThank you, team SONY!",
    "2745220": "## D.Imanishi's part\nThanks to the organizers and congrats to all the winners! I think it was a very exciting and hard competition.\nAnd I am very happy to have finally achieved my goal of becoming a grandmaster!\n\n### Approach\n2D-CNN Models were trained for two types of input data and ensemble them. \n- concatenation of waveform graph image and kaggle spectrogram image (public LB:0.25)\n- waveform graph only (public LB:0.25)\n\nThe input image of my model is like following. 19 signals are the same as competition's sample figure.\n![](https://i.postimg.cc/BQHHgFyw/specgraph.png)\nI had also taken the approach of converting EEG to spectrogram, but did not use it in the end because the score was better without it in the process of team merging.\n\n### Training\n- 2 stage training (1st: all data, 2nd: votes>=10)\n- After training the first model, the second model was trained by applying the pseudo label to votes<10 data using OOF prediction. Instead of replacing it with the pseudo label, it was mixed with the original label according to the number of votes (for samples with fewer votes, increase the weight of the pseudo label).\n- 3 types of CNN backbones were ensembled (EfficientNetV2B0, EfficientNetV2B1, EfficientNetV1B0)\n\n### Augmentation\n- Swap left and right electrode (Fp1<=>Fp2, F3<=>F4, ...)\n- Positive and negative swap (data=-data)\n- Mixup\n- Flipping time axis (kaggle spectrogram only)\n\n### Inference\n- TTA (Same as training augmentations except mixup)\n- Logits averaging (not after softmax) for ensemble\n\n### Code\n- [Waveform image generation](https://www.kaggle.com/code/dimanishi/hms-waveform-image-generation)\n- [Graph model training](https://www.kaggle.com/code/dimanishi/hms-graph-model-training)\n- [Graph model inference](https://www.kaggle.com/code/dimanishi/hms-graph-model-inference)",
    "2745233": "## MaxChen303's part\nThis is the first time I have made a team with others on Kaggle. Thanks to my teammates for bringing the great experience.\n### Summary\nThe main idea of my solution is to generate inputs similar to what the annotators see.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Fef831958400ea7348dfe6ecebf6c511a%2FScreenshot%20from%202024-04-09%2012-53-16.png?generation=1712681638219917&alt=media)\n- The size of each part is configurable.\n    - Final models use size 512x512.\n    - The width of raw eeg part is 500 for 10s. (Should be enough for sampling 20Hz signals)\n    - The height of raw eeg part is controlled by the number of repeat. E.g. 16 channels * 8 repeat = 128 height.\n- Removing the EEG spectrogram part may not affect the performance.\n\nDue to a mistake in my HIGH vote CV implementation (accidentally included some low votes samples), I think I didn't fully optimize the performance in the last week.\n- The best score of my part with ensemble (3 models) is Private 0.295527/Public 0.247586. \n- The best single model score: Private=0.303614, Public = 0.250533\n\n### EEG Preprocessing \nAfter checking the provided EEG pattern images many times, I found that the provided raw EEGs look much cleaner. It was much easier to see the pattern in the provided image than my first few raw EEG plots.\n\nIt turns out that simply applying a bandpass filter can make the raw EEGs look much closer to the provided examples.\nI also noticed that the scipy `sosfilt` is preferred over the commonly used `lfilter` according to the [document](https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.lfilter.html) and the `sosfiltfilt` can reduce the effect of phase shift. \n\n- https://stackoverflow.com/questions/12093594/how-to-implement-band-pass-butterworth-filter-with-scipy-signal-butter/48677312#48677312\n- Thanks to @gunesevitan for sharing this [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/483839)\n\nThe red boxes in this diagram show the possible phase shifts:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3964695%2Ff2e4493929bebfd7149582d9e23e8c0b%2FScreenshot%20from%202024-04-09%2012-56-41.png?generation=1712681817517725&alt=media)\n\n\n### Training\n- One stage training with lower weight for LOW vote samples.\n    - Monitor CV score for both All and HIGH vote samples under different weight ratios.\n    - The final weight selected is 0.25 for low vote samples (<=7) for best ALL and HIGH vote CV score.\n    - The weight is very close to (3.04/14.29 ~= 0.21), which is `average # LOW votes / average # HIGH votes`.\n- Exponential Moving Average (EMA) did well in stabilizing the validation loss during training. \n\n### Augmentation\n- Flip Left and Right components\n- Horizontal flip for spectrograms\n- Time and Frequency mask for the spectrograms\n- Time and Channel mask for the raw EEG. (Randomly hide 4-6 channels in the raw EEG)\n- Time shift of 10s Raw EEG\n- Raw EEG * -1 \n- Small random contrast adjustment for the raw eeg.\n\nTest time augmentation(TTA):\n- Flip Left and Right components.\n\n### Model\nPretrained backbones from TIMM with different pooling layers.\n- Models used in final submission:\n    - maxvit_small_tf_512, \n    - maxvit_tiny_tf_512,\n    - convnext_small\n- Average pooling performs the best in my experiments.\n- LayerNorm -> Dropout -> FC head\n",
    "2745305": "## T.H.'s part\nUEMU & T.H. & kazumax created models as a mini team.  \nI was in charge of the raw EEG part.  \nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.\n\n### pipeline\nI developed a wavenet+2dCNN model using only raw EEG.  \nBelow is an outline of my pipeline.  \nI achieved a public LB = 0.25 (cv = 0.244) with a single model.\n\n![th_pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2449267%2Ff01ca637106a7b304d97a892205fbabc%2Fth_pipeline.png?generation=1712759565877010&alt=media)\n\n### Input\n\n- Compute a total of 16 signals by calculating 4 neighboring difference signals each for LL, RL, LP, RP.\n- 20Hz butter lowpass filter.\n- Clipping of signal intensity.\n\n### Model\n\n- Wavenet+2DCNN.\n- For Wavenet, I referred to this [notebook](https://www.kaggle.com/code/alejopaullier/hms-wavenet-pytorch-train \"HMS | WaveNet PyTorch [Train]\"). Thank you for sharing the code!\n- Input 16 signals separately into Wavenet and average the output feature maps within the groups of LL, RL, LP, RP.\n- Stack the four feature maps vertically to input into 2DCNN.\n- For the 2DCNN model, MaxxvitV2-nano was the most accurate. Maxxvit-small and effnetb4 were also used for the ensemble.\n\n### Training\n\n- 2-stage training:\n  - 1st stage: train with all data for 50 epochs (lr=1e-3).\n  - 2nd stage: fine-tune with data where n_votes>=10 for 30 epochs (lr=1e-4).\n- Optimizer: AdamW.\n- Scheduler: Cosine Annealing (+Warmup).\n- Augmentations (in order of most effective):\n  - ChannelSwap ([LL/LP] <=> [RL/RP]).\n  - Inversion (inversion of signal strength).\n  - TimeMask.\n  - Reverse (time reversal of the signal).",
    "2745192": "## kazumax's part\n\nUEMU & T.H. & kazumax created models as a mini team.  \nI was in charge of the spectrogram part.  \nCV was only calculated for the middle sample when samples from the same eeg_id were lined up in time order.\n### Input\n* 24 spectrograms\n    * 4 KaggleSpec (LL,LR,RP,RL)\n    * 20 EEGSpec  \n        * 16 specs: Adjacent Difference (Fp1-F7, F7-T3, …) \n        * 4 specs (LL,LP,RP,RL)  \n\n* EEGSpec is a power spectrogram truncated at 20Hz not mel spectrogram.\n    ```python\n    spec = librosa.stft(\n        y=x,\n        hop_length=len(x) // 256,\n        n_fft=1024,\n    )\n    spec = np.abs(spec) ** 2\n\n    # truncate at 20Hz.\n    # frequency resoluion = sr / n_fft = 200[Hz] / 1024 = 0.1953[Hz]\n    # 0.1953[Hz] * 100 ≃ 20[Hz]\n    spec = spec[:100, :] \n    ```\n\n### Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2Fdc133f535c8f382e26ed9a77956e3f62%2Fmodel.png?generation=1712754127506152&alt=media)\n### Training\n* 2stage training\n    * 1st stage: all data\n    * 2nd stage: vote >= 10\n* Augmentation\n    * TimeFreqMask(p=0.5)\n    * swap left-right specs(p=0.5)\n        * LL ↔ RL, LR ↔ RP\n        * Fp1 - F7 ↔ FP2 - F8, ...\n* label smoothing\n    * epsilon=0.1 for samples with vote <= 3",
    "2745803": "## UEMU's part\n\n### 1.Multi Input Model \n  - input = Kaggle_Spec + EEG_Spec + Raw_RRG\n  - model = large model(2DCNN + 1DCNN + MLP)\n  - CV(single)    = 0.2436\n  (update later)\n\n### 2.UEMU&kazumax&TH Ensemble\n#### 2_1.Integrated model for Each Data Type\n Our team (UEMU&kazumax&TH) developed models, each member focusing on a different data type.\n However, in this competition, it's necessary to integrate the inference values of each data type, similar to what the annotators are doing.\n Therefore, we created a model that integrates the models for each data type through a fully connected (FC) layer to provide a multifaceted perspective.\n We created some models of each data type for ensemble, and we also developed this model for all combinations.\n  - Feature\n    - prediction of Raw EEG Models(T.H. Part)\n    - prediction of Multi Spec Models(kazumax Part)\n    - prediction of EEG Spec  Models(UEMU Part)\n    - prediction of Multi Input Models(UEMU Part)\n  - Model\n    - multi-layer fully connected (FC) layer (No activation function is used)\n \n#### 2_2.PostProcess\nAs mentioned in this discussion, there are some data where the eeg_id is different but the patient_id is the same in the test dataset.\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471890\n\nWe aggregated the prediction values for each patient_id to create new features.\nUsing the original prediction values and these aggregated features, we built a another inference model.\nThis model did not improve the score on its own, but by including this model in the 2_3.Weight Average, we achieved an improvement in CV of 0.002 to 0.003.\n\n  - Feature\n    - Original prediction\n      - Average prediction of Raw EEG Models(T.H. Part)\n      - Average prediction of Multi Spec Models(kazumax Part)\n      - Average prediction of Multi Input Models(UEMU Part)\n    - Aggregated feature\n      - Average/Max/Min/Std/Delta/Count of original prediction groupby patient_id\n  - Model\n    - multi-layer fully connected (FC) layer (No activation function is used)\n\n\n#### 2_3.Weight Average All Models\nWe took a weighted average using all the models(each single model,2_1.FC model and the 2_2.PostProcess, about 50 models).\nThe weights were optimized based on CV. Although it seemed excessive, the correlation between the LB and CV was good.\n\nThe best scores(UEMU&kazumax&TH's part)\n - CV      = 0.204\n - public  = 0.2382\n - private = 0.2859(≒10th place)\n\n### 3.Team  KTMUD Ensemble\nSimple weighted average of three teams.\nThe weighted average was taken using the logits before applying the softmax function.\nThe weights were decided based on the public leaderboard (LB), where the public best score was equal to the private best score.\n - weights = [D.Imanishi , MaxChen, and UEMU&kazumax&TH] = [0.28, 0.28, 0.44] \n\nThe best scores(Team KTMUD)\n - public  = 0.2258\n - private = 0.2761",
    "2748281": "Congratulations on achieving 5th place in this competition. Thanks for sharing the insights into your solution. "
  }
}