{
  "id": 492472,
  "title": "56th place solution: multimodal, treat montage as 2D image, channel swapping",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/nakano-56th-place-solution-multimodal-treat-montag",
  "author_name": "",
  "post_date": "2024-04-12T12:19:55.830Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This was my first time participating in a Kaggle competition, and I found it to be a highly valuable learning experience. Thanks to Kaggle and the organizers for such a nice competition.</p>\n<h2>Overview</h2>\n<ul>\n<li>Implemented an ensemble of three model types: Spectrogram (Spec), Electroencephalogram (EEG), and Multimodal.</li>\n<li>EEG montages were treated as 2D images with time on the x-axis and channels on the y-axis.</li>\n<li>Channel swapping was an effective data augmentation.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2869454%2Fe596b31c92ab72082fa209b42e3a366a%2Fsolution.png?generation=1712680892467375&amp;alt=media\"></p>\n<h2>Training Data</h2>\n<ul>\n<li>Used train table provided by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>.</li>\n<li>Followed a two-stage training strategy by <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a>. The first stage used all data, while the second stage filtered data to include only instances with a KL loss &gt; 5.5.<br>\n(Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>, and <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a> !)</li>\n</ul>\n<h2>Model</h2>\n<p>Most of my models used EfficientNet as their backbone.</p>\n<h3>Spec Model</h3>\n<ul>\n<li>Employed stacked specs from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's and <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> 's work.</li>\n<li>Found EffnetB0 was most effective. It achieved a second-stage CV of 0.3490, public LB of 0.319433, and private LB of 0.369891.</li>\n</ul>\n<h3>EEG Model</h3>\n<ul>\n<li>Eleven montages were calculated, with time points downsampled from 10,000 to 2,000.</li>\n<li>Prefrontal, Central, and Occipital were calculated by subtracting the right channel from the left one. They helped to identify L and G.</li>\n<li>Stacking each montage into a single image and using EfficientNet improved the CV score by 0.04 over a 1D Resnet1D+GRU model.</li>\n<li>EffnetB7 performed best, reaching a second-stage CV of 0.3508, public LB of 0.322728, and private LB of 0.381983.</li>\n</ul>\n<pre><code>MONTAGE = [\n    [,],[,],  \n    [,],[,],  \n    [,],[,],  \n    [,],[,],  \n    [,], \n    [,], \n    [,], \n]\n</code></pre>\n<h3>Multimodal Model</h3>\n<ul>\n<li>Spec and EEG data were processed through two different backbones and then combined.</li>\n<li>EffnetB0 for both got the best results: second-stage CV of 0.3184, public LB of 0.287670, and private LB of 0.346987.</li>\n<li>Variations with EffnetB0 for Spec and EffnetB7 for EEG, and Resnet50 for Spec and EffnetB0 for EEG were also used for ensemble.</li>\n</ul>\n<h2>Data Augmentation</h2>\n<ul>\n<li>Based on the ACNS guideline, I theorized that the linear symmetrical swapping (right and left, forward and backward) wouldn't impact the target status. This hypothesis appeared to be correct, and channel swapping improved the CV score by more than 0.05 for the EEG model.</li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Weight optimized using SLSQP.</li>\n<li>Reached a second stage CV 0.2768, public LB 0.284148, private LB 0.335807.</li>\n</ul>",
  "messages": [
    {
      "id": "2743950",
      "postDate": "04/09/2024 17:10:14",
      "content": "<p>This was my first time participating in a Kaggle competition, and I found it to be a highly valuable learning experience. Thanks to Kaggle and the organizers for such a nice competition.</p>\n<h2>Overview</h2>\n<ul>\n<li>Implemented an ensemble of three model types: Spectrogram (Spec), Electroencephalogram (EEG), and Multimodal.</li>\n<li>EEG montages were treated as 2D images with time on the x-axis and channels on the y-axis.</li>\n<li>Channel swapping was an effective data augmentation.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2869454%2Fe596b31c92ab72082fa209b42e3a366a%2Fsolution.png?generation=1712680892467375&amp;alt=media\"></p>\n<h2>Training Data</h2>\n<ul>\n<li>Used train table provided by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>.</li>\n<li>Followed a two-stage training strategy by <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a>. The first stage used all data, while the second stage filtered data to include only instances with a KL loss &gt; 5.5.<br>\n(Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>, and <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a> !)</li>\n</ul>\n<h2>Model</h2>\n<p>Most of my models used EfficientNet as their backbone.</p>\n<h3>Spec Model</h3>\n<ul>\n<li>Employed stacked specs from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's and <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> 's work.</li>\n<li>Found EffnetB0 was most effective. It achieved a second-stage CV of 0.3490, public LB of 0.319433, and private LB of 0.369891.</li>\n</ul>\n<h3>EEG Model</h3>\n<ul>\n<li>Eleven montages were calculated, with time points downsampled from 10,000 to 2,000.</li>\n<li>Prefrontal, Central, and Occipital were calculated by subtracting the right channel from the left one. They helped to identify L and G.</li>\n<li>Stacking each montage into a single image and using EfficientNet improved the CV score by 0.04 over a 1D Resnet1D+GRU model.</li>\n<li>EffnetB7 performed best, reaching a second-stage CV of 0.3508, public LB of 0.322728, and private LB of 0.381983.</li>\n</ul>\n<pre><code>MONTAGE = [\n    [,],[,],  \n    [,],[,],  \n    [,],[,],  \n    [,],[,],  \n    [,], \n    [,], \n    [,], \n]\n</code></pre>\n<h3>Multimodal Model</h3>\n<ul>\n<li>Spec and EEG data were processed through two different backbones and then combined.</li>\n<li>EffnetB0 for both got the best results: second-stage CV of 0.3184, public LB of 0.287670, and private LB of 0.346987.</li>\n<li>Variations with EffnetB0 for Spec and EffnetB7 for EEG, and Resnet50 for Spec and EffnetB0 for EEG were also used for ensemble.</li>\n</ul>\n<h2>Data Augmentation</h2>\n<ul>\n<li>Based on the ACNS guideline, I theorized that the linear symmetrical swapping (right and left, forward and backward) wouldn't impact the target status. This hypothesis appeared to be correct, and channel swapping improved the CV score by more than 0.05 for the EEG model.</li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>Weight optimized using SLSQP.</li>\n<li>Reached a second stage CV 0.2768, public LB 0.284148, private LB 0.335807.</li>\n</ul>",
      "rawMarkdown": "This was my first time participating in a Kaggle competition, and I found it to be a highly valuable learning experience. Thanks to Kaggle and the organizers for such a nice competition.\n\n## Overview\n- Implemented an ensemble of three model types: Spectrogram (Spec), Electroencephalogram (EEG), and Multimodal.\n- EEG montages were treated as 2D images with time on the x-axis and channels on the y-axis.\n- Channel swapping was an effective data augmentation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2869454%2Fe596b31c92ab72082fa209b42e3a366a%2Fsolution.png?generation=1712680892467375&alt=media)\n\n## Training Data\n- Used train table provided by @cdeotte and @rafaelzimmermann1.\n- Followed a two-stage training strategy by @zijiangyang1116. The first stage used all data, while the second stage filtered data to include only instances with a KL loss > 5.5.\n(Thanks a lot @cdeotte, @rafaelzimmermann1, and @zijiangyang1116 !)\n\n## Model\n\nMost of my models used EfficientNet as their backbone.\n\n### Spec Model\n- Employed stacked specs from @cdeotte 's and @rafaelzimmermann1 's work.\n- Found EffnetB0 was most effective. It achieved a second-stage CV of 0.3490, public LB of 0.319433, and private LB of 0.369891.\n\n### EEG Model\n- Eleven montages were calculated, with time points downsampled from 10,000 to 2,000.\n- Prefrontal, Central, and Occipital were calculated by subtracting the right channel from the left one. They helped to identify L and G.\n- Stacking each montage into a single image and using EfficientNet improved the CV score by 0.04 over a 1D Resnet1D+GRU model.\n- EffnetB7 performed best, reaching a second-stage CV of 0.3508, public LB of 0.322728, and private LB of 0.381983.\n```python\nMONTAGE = [\n    ['Fp1','T3'],['T3','O1'],  # LL\n    ['Fp1','C3'],['C3','O1'],  # LP\n    ['Fp2','C4'],['C4','O2'],  # RP\n    ['Fp2','T4'],['T4','O2'],  # RL\n    ['Fp1','Fp2'], # Prefrontal\n    ['C3','C4'], # # Central\n    ['O1','O2'], # Occipital\n]\n```\n\n### Multimodal Model\n- Spec and EEG data were processed through two different backbones and then combined.\n- EffnetB0 for both got the best results: second-stage CV of 0.3184, public LB of 0.287670, and private LB of 0.346987.\n- Variations with EffnetB0 for Spec and EffnetB7 for EEG, and Resnet50 for Spec and EffnetB0 for EEG were also used for ensemble.\n\n## Data Augmentation\n- Based on the ACNS guideline, I theorized that the linear symmetrical swapping (right and left, forward and backward) wouldn't impact the target status. This hypothesis appeared to be correct, and channel swapping improved the CV score by more than 0.05 for the EEG model.\n\n## Ensemble\n- Weight optimized using SLSQP.\n- Reached a second stage CV 0.2768, public LB 0.284148, private LB 0.335807.",
      "votes": null
    },
    {
      "id": "2744677",
      "postDate": "04/10/2024 04:21:43",
      "content": "<p>Congratulations on securing the 56th rank in this competition. Thanks for sharing insights into your solution. </p>",
      "rawMarkdown": "Congratulations on securing the 56th rank in this competition. Thanks for sharing insights into your solution.",
      "votes": null
    },
    {
      "id": "2747557",
      "postDate": "04/12/2024 00:53:51",
      "content": "<p>thanks for sharing your great insight🥳</p>",
      "rawMarkdown": "thanks for sharing your great insight🥳",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2744677,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "04/10/2024 04:21:43",
      "content": "<p>Congratulations on securing the 56th rank in this competition. Thanks for sharing insights into your solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2747557,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "04/12/2024 00:53:51",
      "content": "<p>thanks for sharing your great insight🥳</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2743950": "This was my first time participating in a Kaggle competition, and I found it to be a highly valuable learning experience. Thanks to Kaggle and the organizers for such a nice competition.\n\n## Overview\n- Implemented an ensemble of three model types: Spectrogram (Spec), Electroencephalogram (EEG), and Multimodal.\n- EEG montages were treated as 2D images with time on the x-axis and channels on the y-axis.\n- Channel swapping was an effective data augmentation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2869454%2Fe596b31c92ab72082fa209b42e3a366a%2Fsolution.png?generation=1712680892467375&alt=media)\n\n## Training Data\n- Used train table provided by @cdeotte and @rafaelzimmermann1.\n- Followed a two-stage training strategy by @zijiangyang1116. The first stage used all data, while the second stage filtered data to include only instances with a KL loss > 5.5.\n(Thanks a lot @cdeotte, @rafaelzimmermann1, and @zijiangyang1116 !)\n\n## Model\n\nMost of my models used EfficientNet as their backbone.\n\n### Spec Model\n- Employed stacked specs from @cdeotte 's and @rafaelzimmermann1 's work.\n- Found EffnetB0 was most effective. It achieved a second-stage CV of 0.3490, public LB of 0.319433, and private LB of 0.369891.\n\n### EEG Model\n- Eleven montages were calculated, with time points downsampled from 10,000 to 2,000.\n- Prefrontal, Central, and Occipital were calculated by subtracting the right channel from the left one. They helped to identify L and G.\n- Stacking each montage into a single image and using EfficientNet improved the CV score by 0.04 over a 1D Resnet1D+GRU model.\n- EffnetB7 performed best, reaching a second-stage CV of 0.3508, public LB of 0.322728, and private LB of 0.381983.\n```python\nMONTAGE = [\n    ['Fp1','T3'],['T3','O1'],  # LL\n    ['Fp1','C3'],['C3','O1'],  # LP\n    ['Fp2','C4'],['C4','O2'],  # RP\n    ['Fp2','T4'],['T4','O2'],  # RL\n    ['Fp1','Fp2'], # Prefrontal\n    ['C3','C4'], # # Central\n    ['O1','O2'], # Occipital\n]\n```\n\n### Multimodal Model\n- Spec and EEG data were processed through two different backbones and then combined.\n- EffnetB0 for both got the best results: second-stage CV of 0.3184, public LB of 0.287670, and private LB of 0.346987.\n- Variations with EffnetB0 for Spec and EffnetB7 for EEG, and Resnet50 for Spec and EffnetB0 for EEG were also used for ensemble.\n\n## Data Augmentation\n- Based on the ACNS guideline, I theorized that the linear symmetrical swapping (right and left, forward and backward) wouldn't impact the target status. This hypothesis appeared to be correct, and channel swapping improved the CV score by more than 0.05 for the EEG model.\n\n## Ensemble\n- Weight optimized using SLSQP.\n- Reached a second stage CV 0.2768, public LB 0.284148, private LB 0.335807.",
    "2744677": "Congratulations on securing the 56th rank in this competition. Thanks for sharing insights into your solution.",
    "2747557": "thanks for sharing your great insight🥳"
  },
  "source": "meta"
}