{
  "id": 492281,
  "title": "7th Place solution",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/jebastin-gunes-7th-place-solution",
  "author_name": "",
  "post_date": "2024-04-09T05:43:30.069060800Z",
  "votes": 51,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the organizers for this fun competition. This was another great team-up experience with <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>\n<p>Our solution consists of 2D raw eeg model and 2D EEG spectrogram models. We did not use Kaggle 10 min spectrograms, as they consistently showed poor CV score in experiments. We used 18 bipolar signal differences for all our models. EKG signal gave no improvement. We used <code>convnext_base</code> and <code>maxvit_tiny_rw_224</code> for all types of data.</p>\n<h3>2D Raw EEG</h3>\n<p>We reshaped each (10000,) signal into (20, 500), using a stride of 20, and then vstack 18 signals to form a 360x500 2D image. We trained 2 types of models, 50s and 50/10s models.</p>\n<pre><code>inputs = inputs.T\n\n center_10_stack:\n    center_10 = inputs[:, :]\n\nimage = []\nn_channels = \n\n i  (n_channels):\n     j  ():\n        image.append(inputs[i, j::])\n\n     center_10_stack:\n         j  ():\n            image.append(center_10[i, j::])\n\nimage = np.stack(image, axis=)\n</code></pre>\n<h3>2D STFT Spectrogram</h3>\n<p>We used cusignal to generate spectrograms for each signal, and vstack them to from a 2D image. The parameters were chosen in such a way that the final image shape was (512, 512). We tried other resolutions, but found this to be optimal one. Models trained with this data gave best CV scores, so we trained 4 types of models - 50s, 50/10s, 50/30/10s, 30/10s.</p>\n<pre><code>spectrogram_50_second = []\n signal_idx  ():\n    frequencies, _, spectrogram = cusignal.spectrogram(\n        eeg_differences[:, signal_idx],\n        fs=,\n        nperseg=,\n        noverlap=,\n        nfft=\n    )\n    frequency_mask = (frequencies &gt;= ) &amp; (frequencies &lt;= )\n    spectrogram = spectrogram[frequency_mask, :]\n    spectrogram = spectrogram.get().astype(np.float32)\n\n    spectrogram_50_second.append(spectrogram)\n\nspectrogram_50_second = np.concatenate(spectrogram_50_second, axis=)\nspectrogram_50_second = np.log1p(spectrogram_50_second)\n</code></pre>\n<h3>2D Multitaper Spectrogram</h3>\n<p>We found out that the Kaggle 10-min spectrograms are multitaper spectrograms, from <a href=\"https://github.com/bdsp-core/Rapid_IIIC_Labeling_GUI/blob/main/Tools/qEEG/mtspecgram_jj.m\" target=\"_blank\">here</a>. We used the python code <a href=\"https://github.com/preraulab/multitaper_toolbox/blob/master/python/multitaper_spectrogram_python.py\" target=\"_blank\">here</a> to generate EEG spectrograms, and then a combination of hstack, vstack and padding to create (500, 372) image. We trained 2 types of models, 50s and 50/10s models.</p>\n<pre><code>eeg_spec = []\n idx  ():\n    spect, stimes, sfreqs = multitaper_spectrogram(\n        eeg_differences[:, idx],\n        fs=,\n        frequency_range=[, ],\n        window_params=[, ],\n        multiprocess=,\n        plot_on=,\n        verbose=\n    )\n    eeg_spec.append(spect)\n\neeg_spec = np.stack(eeg_spec, axis=).astype(np.float32)\neeg_spec = np.log1p(eeg_spec)\n\nimg = []\n eeg_spec_idx  [(, ), (, ), (, ), (, ), (, )]:\n    eeg_spec_temp = eeg_spec[eeg_spec_idx[]:eeg_spec_idx[]]\n\n     eeg_spec_temp.shape[] == :\n        eeg_spec_temp = np.concatenate([eeg_spec_temp, np.zeros_like(eeg_spec_temp)], axis=)\n\n    img.append(np.concatenate([eeg_spec_temp[], eeg_spec_temp[],\n                               eeg_spec_temp[], eeg_spec_temp[]], axis=))\n\nimg = np.concatenate(img, axis=)\n</code></pre>\n<p>Our final submission is a weighted blend, with weights 0.2, 0.65 and 0.15 for the 3 types of data.</p>",
  "messages": [
    {
      "id": "2742850",
      "postDate": "04/09/2024 05:43:30",
      "content": "<p>Thanks to Kaggle and the organizers for this fun competition. This was another great team-up experience with <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>\n<p>Our solution consists of 2D raw eeg model and 2D EEG spectrogram models. We did not use Kaggle 10 min spectrograms, as they consistently showed poor CV score in experiments. We used 18 bipolar signal differences for all our models. EKG signal gave no improvement. We used <code>convnext_base</code> and <code>maxvit_tiny_rw_224</code> for all types of data.</p>\n<h3>2D Raw EEG</h3>\n<p>We reshaped each (10000,) signal into (20, 500), using a stride of 20, and then vstack 18 signals to form a 360x500 2D image. We trained 2 types of models, 50s and 50/10s models.</p>\n<pre><code>inputs = inputs.T\n\n center_10_stack:\n    center_10 = inputs[:, :]\n\nimage = []\nn_channels = \n\n i  (n_channels):\n     j  ():\n        image.append(inputs[i, j::])\n\n     center_10_stack:\n         j  ():\n            image.append(center_10[i, j::])\n\nimage = np.stack(image, axis=)\n</code></pre>\n<h3>2D STFT Spectrogram</h3>\n<p>We used cusignal to generate spectrograms for each signal, and vstack them to from a 2D image. The parameters were chosen in such a way that the final image shape was (512, 512). We tried other resolutions, but found this to be optimal one. Models trained with this data gave best CV scores, so we trained 4 types of models - 50s, 50/10s, 50/30/10s, 30/10s.</p>\n<pre><code>spectrogram_50_second = []\n signal_idx  ():\n    frequencies, _, spectrogram = cusignal.spectrogram(\n        eeg_differences[:, signal_idx],\n        fs=,\n        nperseg=,\n        noverlap=,\n        nfft=\n    )\n    frequency_mask = (frequencies &gt;= ) &amp; (frequencies &lt;= )\n    spectrogram = spectrogram[frequency_mask, :]\n    spectrogram = spectrogram.get().astype(np.float32)\n\n    spectrogram_50_second.append(spectrogram)\n\nspectrogram_50_second = np.concatenate(spectrogram_50_second, axis=)\nspectrogram_50_second = np.log1p(spectrogram_50_second)\n</code></pre>\n<h3>2D Multitaper Spectrogram</h3>\n<p>We found out that the Kaggle 10-min spectrograms are multitaper spectrograms, from <a href=\"https://github.com/bdsp-core/Rapid_IIIC_Labeling_GUI/blob/main/Tools/qEEG/mtspecgram_jj.m\" target=\"_blank\">here</a>. We used the python code <a href=\"https://github.com/preraulab/multitaper_toolbox/blob/master/python/multitaper_spectrogram_python.py\" target=\"_blank\">here</a> to generate EEG spectrograms, and then a combination of hstack, vstack and padding to create (500, 372) image. We trained 2 types of models, 50s and 50/10s models.</p>\n<pre><code>eeg_spec = []\n idx  ():\n    spect, stimes, sfreqs = multitaper_spectrogram(\n        eeg_differences[:, idx],\n        fs=,\n        frequency_range=[, ],\n        window_params=[, ],\n        multiprocess=,\n        plot_on=,\n        verbose=\n    )\n    eeg_spec.append(spect)\n\neeg_spec = np.stack(eeg_spec, axis=).astype(np.float32)\neeg_spec = np.log1p(eeg_spec)\n\nimg = []\n eeg_spec_idx  [(, ), (, ), (, ), (, ), (, )]:\n    eeg_spec_temp = eeg_spec[eeg_spec_idx[]:eeg_spec_idx[]]\n\n     eeg_spec_temp.shape[] == :\n        eeg_spec_temp = np.concatenate([eeg_spec_temp, np.zeros_like(eeg_spec_temp)], axis=)\n\n    img.append(np.concatenate([eeg_spec_temp[], eeg_spec_temp[],\n                               eeg_spec_temp[], eeg_spec_temp[]], axis=))\n\nimg = np.concatenate(img, axis=)\n</code></pre>\n<p>Our final submission is a weighted blend, with weights 0.2, 0.65 and 0.15 for the 3 types of data.</p>",
      "rawMarkdown": "Thanks to Kaggle and the organizers for this fun competition. This was another great team-up experience with @gunesevitan \n\nOur solution consists of 2D raw eeg model and 2D EEG spectrogram models. We did not use Kaggle 10 min spectrograms, as they consistently showed poor CV score in experiments. We used 18 bipolar signal differences for all our models. EKG signal gave no improvement. We used `convnext_base` and `maxvit_tiny_rw_224` for all types of data.\n\n### 2D Raw EEG\nWe reshaped each (10000,) signal into (20, 500), using a stride of 20, and then vstack 18 signals to form a 360x500 2D image. We trained 2 types of models, 50s and 50/10s models.\n\n```python\ninputs = inputs.T\n\nif center_10_stack:\n    center_10 = inputs[:, 4000:6000]\n\nimage = []\nn_channels = 18\n\nfor i in range(n_channels):\n    for j in range(20):\n        image.append(inputs[i, j::20])\n\n    if center_10_stack:\n        for j in range(4):\n            image.append(center_10[i, j::4])\n\nimage = np.stack(image, axis=0)\n```\n\n### 2D STFT Spectrogram\nWe used cusignal to generate spectrograms for each signal, and vstack them to from a 2D image. The parameters were chosen in such a way that the final image shape was (512, 512). We tried other resolutions, but found this to be optimal one. Models trained with this data gave best CV scores, so we trained 4 types of models - 50s, 50/10s, 50/30/10s, 30/10s.\n\n```python\nspectrogram_50_second = []\nfor signal_idx in range(18):\n    frequencies, _, spectrogram = cusignal.spectrogram(\n        eeg_differences[:, signal_idx],\n        fs=200,\n        nperseg=280,\n        noverlap=261,\n        nfft=None\n    )\n    frequency_mask = (frequencies >= 0.5) & (frequencies <= 20)\n    spectrogram = spectrogram[frequency_mask, :]\n    spectrogram = spectrogram.get().astype(np.float32)\n        \n    spectrogram_50_second.append(spectrogram)\n        \nspectrogram_50_second = np.concatenate(spectrogram_50_second, axis=0)\nspectrogram_50_second = np.log1p(spectrogram_50_second)\n```\n\n### 2D Multitaper Spectrogram\nWe found out that the Kaggle 10-min spectrograms are multitaper spectrograms, from [here](https://github.com/bdsp-core/Rapid_IIIC_Labeling_GUI/blob/main/Tools/qEEG/mtspecgram_jj.m). We used the python code [here](https://github.com/preraulab/multitaper_toolbox/blob/master/python/multitaper_spectrogram_python.py) to generate EEG spectrograms, and then a combination of hstack, vstack and padding to create (500, 372) image. We trained 2 types of models, 50s and 50/10s models.\n\n```python\neeg_spec = []\nfor idx in range(18):\n    spect, stimes, sfreqs = multitaper_spectrogram(\n        eeg_differences[:, idx],\n        fs=200,\n        frequency_range=[0.5, 20],\n        window_params=[4, 0.5],\n        multiprocess=True,\n        plot_on=False,\n        verbose=False\n    )\n    eeg_spec.append(spect)\n     \neeg_spec = np.stack(eeg_spec, axis=0).astype(np.float32)\neeg_spec = np.log1p(eeg_spec)\n\nimg = []\nfor eeg_spec_idx in [(0, 4), (4, 8), (8, 12), (12, 16), (16, 18)]:\n    eeg_spec_temp = eeg_spec[eeg_spec_idx[0]:eeg_spec_idx[1]]\n\n    if eeg_spec_temp.shape[0] == 2:\n        eeg_spec_temp = np.concatenate([eeg_spec_temp, np.zeros_like(eeg_spec_temp)], axis=0)\n\n    img.append(np.concatenate([eeg_spec_temp[0], eeg_spec_temp[1],\n                               eeg_spec_temp[2], eeg_spec_temp[3]], axis=1))\n\nimg = np.concatenate(img, axis=0)\n```\n\nOur final submission is a weighted blend, with weights 0.2, 0.65 and 0.15 for the 3 types of data.",
      "votes": null
    },
    {
      "id": "2743072",
      "postDate": "04/09/2024 07:59:38",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> , Keep up the good work, and thanks for these great insights.</p>",
      "rawMarkdown": "Congratulations @samfc10 , Keep up the good work, and thanks for these great insights.",
      "votes": null
    },
    {
      "id": "2743201",
      "postDate": "04/09/2024 09:25:35",
      "content": "<p>Congratulations! <br>\nCould you please explain, what does \"50/10s models\" mean?</p>",
      "rawMarkdown": "Congratulations! \nCould you please explain, what does \"50/10s models\" mean?",
      "votes": null
    },
    {
      "id": "2743207",
      "postDate": "04/09/2024 09:28:27",
      "content": "<p>Great work, congratulations! The solution is bold and different.</p>",
      "rawMarkdown": "Great work, congratulations! The solution is bold and different.",
      "votes": null
    },
    {
      "id": "2743268",
      "postDate": "04/09/2024 10:34:00",
      "content": "<p>vstack 50s data and centre 10s data. We did this because the label is for the centre 10s, and helps the model learn from both the resolutions. It gave good improvement in all our models.</p>",
      "rawMarkdown": "vstack 50s data and centre 10s data. We did this because the label is for the centre 10s, and helps the model learn from both the resolutions. It gave good improvement in all our models.",
      "votes": null
    },
    {
      "id": "2744072",
      "postDate": "04/09/2024 18:30:27",
      "content": "<p>Congrats! Thanks for sharing. This solution is different. Great work!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing. This solution is different. Great work!",
      "votes": null
    },
    {
      "id": "2744692",
      "postDate": "04/10/2024 04:25:38",
      "content": "<p>Congratulations on 20th place in this competition. Thanks for sharing analysis of your code. </p>",
      "rawMarkdown": "Congratulations on 20th place in this competition. Thanks for sharing analysis of your code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2743072,
      "author_name": "pranayakuniyal",
      "author_url": "",
      "post_date": "04/09/2024 07:59:38",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/samfc10\" target=\"_blank\">@samfc10</a> , Keep up the good work, and thanks for these great insights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2743201,
      "author_name": "germanroev",
      "author_url": "",
      "post_date": "04/09/2024 09:25:35",
      "content": "<p>Congratulations! <br>\nCould you please explain, what does \"50/10s models\" mean?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2743268,
          "author_name": "samfc10",
          "author_url": "",
          "post_date": "04/09/2024 10:34:00",
          "content": "<p>vstack 50s data and centre 10s data. We did this because the label is for the centre 10s, and helps the model learn from both the resolutions. It gave good improvement in all our models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2743207,
      "author_name": "nartaa",
      "author_url": "",
      "post_date": "04/09/2024 09:28:27",
      "content": "<p>Great work, congratulations! The solution is bold and different.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744072,
      "author_name": "zlemglsmklkaya",
      "author_url": "",
      "post_date": "04/09/2024 18:30:27",
      "content": "<p>Congrats! Thanks for sharing. This solution is different. Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744692,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "04/10/2024 04:25:38",
      "content": "<p>Congratulations on 20th place in this competition. Thanks for sharing analysis of your code. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742850": "Thanks to Kaggle and the organizers for this fun competition. This was another great team-up experience with @gunesevitan \n\nOur solution consists of 2D raw eeg model and 2D EEG spectrogram models. We did not use Kaggle 10 min spectrograms, as they consistently showed poor CV score in experiments. We used 18 bipolar signal differences for all our models. EKG signal gave no improvement. We used `convnext_base` and `maxvit_tiny_rw_224` for all types of data.\n\n### 2D Raw EEG\nWe reshaped each (10000,) signal into (20, 500), using a stride of 20, and then vstack 18 signals to form a 360x500 2D image. We trained 2 types of models, 50s and 50/10s models.\n\n```python\ninputs = inputs.T\n\nif center_10_stack:\n    center_10 = inputs[:, 4000:6000]\n\nimage = []\nn_channels = 18\n\nfor i in range(n_channels):\n    for j in range(20):\n        image.append(inputs[i, j::20])\n\n    if center_10_stack:\n        for j in range(4):\n            image.append(center_10[i, j::4])\n\nimage = np.stack(image, axis=0)\n```\n\n### 2D STFT Spectrogram\nWe used cusignal to generate spectrograms for each signal, and vstack them to from a 2D image. The parameters were chosen in such a way that the final image shape was (512, 512). We tried other resolutions, but found this to be optimal one. Models trained with this data gave best CV scores, so we trained 4 types of models - 50s, 50/10s, 50/30/10s, 30/10s.\n\n```python\nspectrogram_50_second = []\nfor signal_idx in range(18):\n    frequencies, _, spectrogram = cusignal.spectrogram(\n        eeg_differences[:, signal_idx],\n        fs=200,\n        nperseg=280,\n        noverlap=261,\n        nfft=None\n    )\n    frequency_mask = (frequencies >= 0.5) & (frequencies <= 20)\n    spectrogram = spectrogram[frequency_mask, :]\n    spectrogram = spectrogram.get().astype(np.float32)\n        \n    spectrogram_50_second.append(spectrogram)\n        \nspectrogram_50_second = np.concatenate(spectrogram_50_second, axis=0)\nspectrogram_50_second = np.log1p(spectrogram_50_second)\n```\n\n### 2D Multitaper Spectrogram\nWe found out that the Kaggle 10-min spectrograms are multitaper spectrograms, from [here](https://github.com/bdsp-core/Rapid_IIIC_Labeling_GUI/blob/main/Tools/qEEG/mtspecgram_jj.m). We used the python code [here](https://github.com/preraulab/multitaper_toolbox/blob/master/python/multitaper_spectrogram_python.py) to generate EEG spectrograms, and then a combination of hstack, vstack and padding to create (500, 372) image. We trained 2 types of models, 50s and 50/10s models.\n\n```python\neeg_spec = []\nfor idx in range(18):\n    spect, stimes, sfreqs = multitaper_spectrogram(\n        eeg_differences[:, idx],\n        fs=200,\n        frequency_range=[0.5, 20],\n        window_params=[4, 0.5],\n        multiprocess=True,\n        plot_on=False,\n        verbose=False\n    )\n    eeg_spec.append(spect)\n     \neeg_spec = np.stack(eeg_spec, axis=0).astype(np.float32)\neeg_spec = np.log1p(eeg_spec)\n\nimg = []\nfor eeg_spec_idx in [(0, 4), (4, 8), (8, 12), (12, 16), (16, 18)]:\n    eeg_spec_temp = eeg_spec[eeg_spec_idx[0]:eeg_spec_idx[1]]\n\n    if eeg_spec_temp.shape[0] == 2:\n        eeg_spec_temp = np.concatenate([eeg_spec_temp, np.zeros_like(eeg_spec_temp)], axis=0)\n\n    img.append(np.concatenate([eeg_spec_temp[0], eeg_spec_temp[1],\n                               eeg_spec_temp[2], eeg_spec_temp[3]], axis=1))\n\nimg = np.concatenate(img, axis=0)\n```\n\nOur final submission is a weighted blend, with weights 0.2, 0.65 and 0.15 for the 3 types of data.",
    "2743072": "Congratulations @samfc10 , Keep up the good work, and thanks for these great insights.",
    "2743201": "Congratulations! \nCould you please explain, what does \"50/10s models\" mean?",
    "2743207": "Great work, congratulations! The solution is bold and different.",
    "2743268": "vstack 50s data and centre 10s data. We did this because the label is for the centre 10s, and helps the model learn from both the resolutions. It gave good improvement in all our models.",
    "2744072": "Congrats! Thanks for sharing. This solution is different. Great work!",
    "2744692": "Congratulations on 20th place in this competition. Thanks for sharing analysis of your code."
  },
  "source": "meta"
}