{
  "id": 492254,
  "title": "2nd place solution ",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492254",
  "author_name": "coolz",
  "post_date": "2024-04-09T03:01:28.941000",
  "votes": 133,
  "comment_count": 40,
  "views": 0,
  "content": "<p>First of all, I would like to thanks to Kaggle and the organizers. I have learned a lot from this experience.  </p>\n<h2>1. model</h2>\n<p>In early time, feeding a bsx4xHxW input to the 2D-CNN produce worse result than the one-image method.  I start to think why?</p>\n<p>Due to the position-sensitive nature of the labels (LPD, GPD, etc.), We should take care of the channel dimension. 2D-CNN is not well-equipped to capture positional information within the channels. Because there is not pad in channel direction. And that's why we need to concat the spectrum into one image for a vision model other than a bsx16xHxW image( double banana montage).<br>\nSo I decided to use 3d-CNN for spectrum, and 2d-CNN model for raw-eeg signal.</p>\n<p>diagram like:</p>\n<p>total solution like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F9d2ab8354f85f39d092a75fd13922db7%2Fimg.png?generation=1715226166682613&amp;alt=media\" alt=\"solution figure\"></p>\n<ul>\n<li>mne 0.5-20hz means use MNE-tool do filter.</li>\n<li>scipy.signal means use scipy.signal to do filter.</li>\n<li>For the reshape operator and the stft params please refer to the codes below.</li>\n<li>And the number in () means final weights of the ensemble.</li>\n</ul>\n<h3>1.1 x3d-l (spectrum model)</h3>\n<p>After double banana montage, +-1024 clip and 0.5-20hz filter,<br>\nuse stft to extract the spectrum, then feed to a 3d-CNN(x3d-l).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2Fcbfe301faa8741dc262aa334e620e4c9%2Fimg_2.png?generation=1715226202040353&amp;alt=media\"></p>\n<p>The input data is a 16-channels spectrum image.</p>\n<p>X3d-l cv 0.21+, public lb0.25 private lb0.29.</p>\n<p>stft code like below, use it as a nn.Module :</p>\n<pre><code></code></pre>\n<h3>1.2 oneimage</h3>\n<p>It's a 2d vision model, efficientnetb5, public lb 0.26 2490 private lb0.304877.</p>\n<p>For 2d model, I concat the 16-channels spectrum image like many people does.</p>\n<pre><code>         = torch.reshape(image, shape=[n, , -, w])\n         = image[:, :, ...]\n         = image[:, :, ...]\n         = torch.cat([x1, x2], dim=-)\n         = torch.cat([image, image, image], dim=)\n</code></pre>\n<p>I check the submission score, combine these two spectrum model, get 0.28 private lb.</p>\n<h3>1.3 eeg model</h3>\n<p>View eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.<br>\nThe model define as below</p>\n<pre><code>rclass Net(nn.Module):\n     ():\n        (Net, self).__init__()\n        self.model = timm.create_model(, pretrained=, in_chans=)\n        self.pool = nn.AdaptiveAvgPool2d()\n        self.fc = nn.Linear(, out_features=, bias=)\n        self.dropout = nn.Dropout(p=)\n\n     ():\n        feature1 = self.model.forward_features(x)\n         feature1\n\n     ():\n        bs = x.size()\n        reshaped_tensor = x.view(bs, , , )\n        reshaped_and_permuted_tensor = reshaped_tensor.permute(, , , )\n        reshaped_and_permuted_tensor = reshaped_and_permuted_tensor.reshape(bs,  * , )\n        x = torch.unsqueeze(reshaped_and_permuted_tensor, dim=)\n\n        x = torch.cat([x, x, x], dim=)\n        bs = x.size()\n\n        x = self.extract_features(x)\n        x = self.pool(x)\n        x = x.view(bs, -)\n        x = self.dropout(x)\n        x = self.fc(x)\n         x\n</code></pre>\n<p>Then, the eeg 'reshaped' to bsx3x160x1000.&nbsp; With efficinetnetb5, actives public lb 0.230 978, private lb 0.282 873.</p>\n<p>There are 3 models , but slightly different, mne-filter-efficinetnetb5, scipy.signal-filter-efficientnetb5 and mne-filter-hgnetb5. <br>\nAnd the scires are slightly different,also.</p>\n<h3>1.4  doublehead (eeg+spectrum)</h3>\n<p>I use x3d-l to extract the spectrum feature ( with Transform50s only ), efficientnetb5 to extract the raw eeg feature, <br>\nlike the solution figure.<br>\nconcat the last feature. public lb 0.24, private lb 0.29 ? ? not sure</p>\n<h1>2 Preprocess</h1>\n<ul>\n<li>2.1.Double banana montage, eeg as 16x10000</li>\n<li>2.2.Filter with 0.5-20hz</li>\n<li>2.3.Clip by +-1024</li>\n</ul>\n<h1>3. Train</h1>\n<ul>\n<li>3.1 Stage1, 15 epochs, with loss weight voters_num/20, Adamw lr=0.001, cos scheduel</li>\n<li>3.2 Stage2, 5 epoch, loss weight=1, voters_num&gt;=6.   , Adamw lr=0.0001, cos scheduel</li>\n<li>3.3 use eeg_label_offset_seconds. I random choose an offset for each eegid, and the target is average according to eegid for each train iter.</li>\n<li>3.4 augmentation, mirror eeg, flip between left brain data and right brain data.</li>\n<li>3.5 10 folds, then move 1000 samples from val to train, left 709 samples in val set. And use vote_num&gt;=6 to do validation. </li>\n</ul>\n<h1>4. ensemble</h1>\n<p>By combining these models, I think it could get my current score. However my score is 6 model emsemble, but not improvement that much (0.28-&gt;0.27 private lb). <br>\nFinal ensemble including:  </p>\n<ul>\n<li>2 spectrum model (x3d, efficientnetb5), both use mne-filter</li>\n<li>3 raw eeg model( efficientnetb5 with-mne.filter, efficientnetb5 -butter filter, 1 hgnetb5 mne.filter),</li>\n<li>1 eeg-spectrum mix model, butter filter,   </li>\n</ul>\n<p>With weights=[0.1,0.1,0.2,0.2,0.2,0.2].</p>\n<p>ps. with-mne.filter means use mne lib to do filter, butter filter means use scipy.signal, just to add some diversity.</p>\n<h1>Some thought</h1>\n<p>I think raw eeg is more important in this task.  Feeding raw EEG data into a 2D visual model is somewhat similar to how humans observe EEG signals. Observe in time and channel  dimensions！</p>",
  "messages": [
    {
      "id": 2742687,
      "postDate": "2024-04-09T03:01:28.943Z",
      "content": "<p>First of all, I would like to thanks to Kaggle and the organizers. I have learned a lot from this experience.  </p>\n<h2>1. model</h2>\n<p>In early time, feeding a bsx4xHxW input to the 2D-CNN produce worse result than the one-image method.  I start to think why?</p>\n<p>Due to the position-sensitive nature of the labels (LPD, GPD, etc.), We should take care of the channel dimension. 2D-CNN is not well-equipped to capture positional information within the channels. Because there is not pad in channel direction. And that's why we need to concat the spectrum into one image for a vision model other than a bsx16xHxW image( double banana montage).<br>\nSo I decided to use 3d-CNN for spectrum, and 2d-CNN model for raw-eeg signal.</p>\n<p>diagram like:</p>\n<p>total solution like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F9d2ab8354f85f39d092a75fd13922db7%2Fimg.png?generation=1715226166682613&amp;alt=media\" alt=\"solution figure\"></p>\n<ul>\n<li>mne 0.5-20hz means use MNE-tool do filter.</li>\n<li>scipy.signal means use scipy.signal to do filter.</li>\n<li>For the reshape operator and the stft params please refer to the codes below.</li>\n<li>And the number in () means final weights of the ensemble.</li>\n</ul>\n<h3>1.1 x3d-l (spectrum model)</h3>\n<p>After double banana montage, +-1024 clip and 0.5-20hz filter,<br>\nuse stft to extract the spectrum, then feed to a 3d-CNN(x3d-l).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2Fcbfe301faa8741dc262aa334e620e4c9%2Fimg_2.png?generation=1715226202040353&amp;alt=media\"></p>\n<p>The input data is a 16-channels spectrum image.</p>\n<p>X3d-l cv 0.21+, public lb0.25 private lb0.29.</p>\n<p>stft code like below, use it as a nn.Module :</p>\n<pre><code></code></pre>\n<h3>1.2 oneimage</h3>\n<p>It's a 2d vision model, efficientnetb5, public lb 0.26 2490 private lb0.304877.</p>\n<p>For 2d model, I concat the 16-channels spectrum image like many people does.</p>\n<pre><code>         = torch.reshape(image, shape=[n, , -, w])\n         = image[:, :, ...]\n         = image[:, :, ...]\n         = torch.cat([x1, x2], dim=-)\n         = torch.cat([image, image, image], dim=)\n</code></pre>\n<p>I check the submission score, combine these two spectrum model, get 0.28 private lb.</p>\n<h3>1.3 eeg model</h3>\n<p>View eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.<br>\nThe model define as below</p>\n<pre><code>rclass Net(nn.Module):\n     ():\n        (Net, self).__init__()\n        self.model = timm.create_model(, pretrained=, in_chans=)\n        self.pool = nn.AdaptiveAvgPool2d()\n        self.fc = nn.Linear(, out_features=, bias=)\n        self.dropout = nn.Dropout(p=)\n\n     ():\n        feature1 = self.model.forward_features(x)\n         feature1\n\n     ():\n        bs = x.size()\n        reshaped_tensor = x.view(bs, , , )\n        reshaped_and_permuted_tensor = reshaped_tensor.permute(, , , )\n        reshaped_and_permuted_tensor = reshaped_and_permuted_tensor.reshape(bs,  * , )\n        x = torch.unsqueeze(reshaped_and_permuted_tensor, dim=)\n\n        x = torch.cat([x, x, x], dim=)\n        bs = x.size()\n\n        x = self.extract_features(x)\n        x = self.pool(x)\n        x = x.view(bs, -)\n        x = self.dropout(x)\n        x = self.fc(x)\n         x\n</code></pre>\n<p>Then, the eeg 'reshaped' to bsx3x160x1000.&nbsp; With efficinetnetb5, actives public lb 0.230 978, private lb 0.282 873.</p>\n<p>There are 3 models , but slightly different, mne-filter-efficinetnetb5, scipy.signal-filter-efficientnetb5 and mne-filter-hgnetb5. <br>\nAnd the scires are slightly different,also.</p>\n<h3>1.4  doublehead (eeg+spectrum)</h3>\n<p>I use x3d-l to extract the spectrum feature ( with Transform50s only ), efficientnetb5 to extract the raw eeg feature, <br>\nlike the solution figure.<br>\nconcat the last feature. public lb 0.24, private lb 0.29 ? ? not sure</p>\n<h1>2 Preprocess</h1>\n<ul>\n<li>2.1.Double banana montage, eeg as 16x10000</li>\n<li>2.2.Filter with 0.5-20hz</li>\n<li>2.3.Clip by +-1024</li>\n</ul>\n<h1>3. Train</h1>\n<ul>\n<li>3.1 Stage1, 15 epochs, with loss weight voters_num/20, Adamw lr=0.001, cos scheduel</li>\n<li>3.2 Stage2, 5 epoch, loss weight=1, voters_num&gt;=6.   , Adamw lr=0.0001, cos scheduel</li>\n<li>3.3 use eeg_label_offset_seconds. I random choose an offset for each eegid, and the target is average according to eegid for each train iter.</li>\n<li>3.4 augmentation, mirror eeg, flip between left brain data and right brain data.</li>\n<li>3.5 10 folds, then move 1000 samples from val to train, left 709 samples in val set. And use vote_num&gt;=6 to do validation. </li>\n</ul>\n<h1>4. ensemble</h1>\n<p>By combining these models, I think it could get my current score. However my score is 6 model emsemble, but not improvement that much (0.28-&gt;0.27 private lb). <br>\nFinal ensemble including:  </p>\n<ul>\n<li>2 spectrum model (x3d, efficientnetb5), both use mne-filter</li>\n<li>3 raw eeg model( efficientnetb5 with-mne.filter, efficientnetb5 -butter filter, 1 hgnetb5 mne.filter),</li>\n<li>1 eeg-spectrum mix model, butter filter,   </li>\n</ul>\n<p>With weights=[0.1,0.1,0.2,0.2,0.2,0.2].</p>\n<p>ps. with-mne.filter means use mne lib to do filter, butter filter means use scipy.signal, just to add some diversity.</p>\n<h1>Some thought</h1>\n<p>I think raw eeg is more important in this task.  Feeding raw EEG data into a 2D visual model is somewhat similar to how humans observe EEG signals. Observe in time and channel  dimensions！</p>",
      "rawMarkdown": "First of all, I would like to thanks to Kaggle and the organizers. I have learned a lot from this experience.  \n\n\n\n##  1. model\n\nIn early time, feeding a bsx4xHxW input to the 2D-CNN produce worse result than the one-image method.  I start to think why?\n\nDue to the position-sensitive nature of the labels (LPD, GPD, etc.), We should take care of the channel dimension. 2D-CNN is not well-equipped to capture positional information within the channels. Because there is not pad in channel direction. And that's why we need to concat the spectrum into one image for a vision model other than a bsx16xHxW image( double banana montage).\nSo I decided to use 3d-CNN for spectrum, and 2d-CNN model for raw-eeg signal.\n\n\ndiagram like:\n\n\ntotal solution like this:\n![solution figure](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F9d2ab8354f85f39d092a75fd13922db7%2Fimg.png?generation=1715226166682613&alt=media)\n+ mne 0.5-20hz means use MNE-tool do filter.\n+ scipy.signal means use scipy.signal to do filter.\n+ For the reshape operator and the stft params please refer to the codes below.\n+ And the number in () means final weights of the ensemble.\n\n### 1.1 x3d-l (spectrum model)\n\n\nAfter double banana montage, +-1024 clip and 0.5-20hz filter,\nuse stft to extract the spectrum, then feed to a 3d-CNN(x3d-l).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2Fcbfe301faa8741dc262aa334e620e4c9%2Fimg_2.png?generation=1715226202040353&alt=media)\n\nThe input data is a 16-channels spectrum image.\n\nX3d-l cv 0.21+, public lb0.25 private lb0.29.\n\n\n\nstft code like below, use it as a nn.Module :\n\n```\nclass Transform50s(nn.Module):\n    def __init__(self, ):\n        super().__init__()\n        self.wave_transform = torchaudio.transforms.Spectrogram(n_fft=512, win_length=128, hop_length=50, power=1)\n\n    def forward(self, x):\n        image = self.wave_transform(x)\n        image = torch.clip(image, min=0, max=10000) / 1000\n        n, c, h, w = image.size()\n        image = image[:, :, :int(20 / 100 * h + 10), :]\n        return image\n\n\nclass Transform10s(nn.Module):\n    def __init__(self, ):\n        super().__init__()\n        self.wave_transform = torchaudio.transforms.Spectrogram(n_fft=512, win_length=128, hop_length=10, power=1)\n\n    def forward(self, x):\n        image = self.wave_transform(x)\n        image = torch.clip(image, min=0, max=10000) / 1000\n        n, c, h, w = image.size()\n        image = image[:, :, :int(20 / 100 * h + 10), :]\n        return image\n\n\nclass Model(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        model_name = \"x3d_l\"\n        self.net = torch.hub.load('facebookresearch/pytorchvideo',\n                                  model_name, pretrained=True)\n\n        self.net.blocks[5].pool.pool = nn.AdaptiveAvgPool3d(1)\n        # self.net.blocks[5]=nn.Identity()\n        # self.net.avgpool = nn.Identity()\n        self.net.blocks[5].dropout = nn.Identity()\n        self.net.blocks[5].proj = nn.Identity()\n        self.net.blocks[5].activation = nn.Identity()\n        self.net.blocks[5].output_pool = nn.Identity()\n\n    def forward(self, x):\n        x = self.net(x)\n        return x\n\n\nclass Net(nn.Module):\n    def __init__(self, num_classes=1):\n        super().__init__()\n\n        self.preprocess50s = Transform50s()\n        self.preprocess10s = Transform10s()\n\n        self.model = Model()\n\n        self.pool = nn.AdaptiveAvgPool3d(1)\n        self.fc = nn.Linear(2048, 6, bias=True)\n        self.drop = nn.Dropout(0.5)\n\n    def forward(self, eeg):\n        # do preprocess\n        bs = eeg.size(0)\n        eeg_50s = eeg\n        eeg_10s = eeg[:, :, 4000:6000]\n        x_50 = self.preprocess50s(eeg_50s)\n        x_10 = self.preprocess10s(eeg_10s)\n        x = torch.cat([x_10, x_50], dim=1)\n        x = torch.unsqueeze(x, dim=1)\n        x = torch.cat([x, x, x], dim=1)\n        x = self.model(x)\n        # x = self.pool(x)\n        x = x.view(bs, -1)\n        x = self.drop(x)\n        x = self.fc(x)\n        return x\n\n```\n\n### 1.2 oneimage\nIt's a 2d vision model, efficientnetb5, public lb 0.26 2490 private lb0.304877.\n\nFor 2d model, I concat the 16-channels spectrum image like many people does.\n\n```\n        image = torch.reshape(image, shape=[n, 2, -1, w])\n        x1 = image[:, 0:1, ...]\n        x2 = image[:, 1:2, ...]\n        image = torch.cat([x1, x2], dim=-1)\n        image = torch.cat([image, image, image], dim=1)\n```\n\nI check the submission score, combine these two spectrum model, get 0.28 private lb.\n\n\n### 1.3 eeg model\n\nView eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.\nThe model define as below\n```python\nrclass Net(nn.Module):\n    def __init__(self,):\n        super(Net, self).__init__()\n        self.model = timm.create_model('efficientnet_b5', pretrained=True, in_chans=3)\n        self.pool = nn.AdaptiveAvgPool2d(1)\n        self.fc = nn.Linear(2048, out_features=6, bias=True)\n        self.dropout = nn.Dropout(p=0.5)\n\n    def extract_features(self, x):\n        feature1 = self.model.forward_features(x)\n        return feature1\n\n    def forward(self, x):\n        bs = x.size(0)\n        reshaped_tensor = x.view(bs, 16, 1000, 10)\n        reshaped_and_permuted_tensor = reshaped_tensor.permute(0, 1, 3, 2)\n        reshaped_and_permuted_tensor = reshaped_and_permuted_tensor.reshape(bs, 16 * 10, 1000)\n        x = torch.unsqueeze(reshaped_and_permuted_tensor, dim=1)\n\n        x = torch.cat([x, x, x], dim=1)\n        bs = x.size(0)\n\n        x = self.extract_features(x)\n        x = self.pool(x)\n        x = x.view(bs, -1)\n        x = self.dropout(x)\n        x = self.fc(x)\n        return x\n```\nThen, the eeg 'reshaped' to bsx3x160x1000.  With efficinetnetb5, actives public lb 0.230 978, private lb 0.282 873.\n\nThere are 3 models , but slightly different, mne-filter-efficinetnetb5, scipy.signal-filter-efficientnetb5 and mne-filter-hgnetb5. \nAnd the scires are slightly different,also.\n\n### 1.4  doublehead (eeg+spectrum)\n\nI use x3d-l to extract the spectrum feature ( with Transform50s only ), efficientnetb5 to extract the raw eeg feature, \nlike the solution figure.\nconcat the last feature. public lb 0.24, private lb 0.29 ? ? not sure\n\n# 2 Preprocess\n- 2.1.Double banana montage, eeg as 16x10000\n- 2.2.Filter with 0.5-20hz\n- 2.3.Clip by +-1024\n\n# 3. Train\n- 3.1 Stage1, 15 epochs, with loss weight voters_num/20, Adamw lr=0.001, cos scheduel\n- 3.2 Stage2, 5 epoch, loss weight=1, voters_num>=6.   , Adamw lr=0.0001, cos scheduel\n- 3.3 use eeg_label_offset_seconds. I random choose an offset for each eegid, and the target is average according to eegid for each train iter.\n- 3.4 augmentation, mirror eeg, flip between left brain data and right brain data.\n- 3.5 10 folds, then move 1000 samples from val to train, left 709 samples in val set. And use vote_num>=6 to do validation. \n\n# 4. ensemble\n\nBy combining these models, I think it could get my current score. However my score is 6 model emsemble, but not improvement that much (0.28->0.27 private lb). \nFinal ensemble including:  \n- 2 spectrum model (x3d, efficientnetb5), both use mne-filter\n- 3 raw eeg model( efficientnetb5 with-mne.filter, efficientnetb5 -butter filter, 1 hgnetb5 mne.filter),\n- 1 eeg-spectrum mix model, butter filter,   \n\nWith weights=[0.1,0.1,0.2,0.2,0.2,0.2].\n\nps. with-mne.filter means use mne lib to do filter, butter filter means use scipy.signal, just to add some diversity.\n\n# Some thought\nI think raw eeg is more important in this task.  Feeding raw EEG data into a 2D visual model is somewhat similar to how humans observe EEG signals. Observe in time and channel  dimensions！\n\n",
      "votes": 133
    },
    {
      "id": 2743455,
      "postDate": "2024-04-09T13:08:35.670Z",
      "content": "<p><a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a> </p>\n<p>Congratulations on your solo 2nd place. Very impressed with your very unique solution!</p>\n<pre><code> = x.view(bs, , , )\n = reshaped_tensor.permute(, , , )\n = reshaped_and_permuted_tensor.reshape(bs,  * , )\n = torch.unsqueeze(reshaped_and_permuted_tensor, dim=)\n = torch.cat([x, x, x], dim=)\n</code></pre>\n<blockquote>\n  <p>Then, the eeg 'reshaped' to bsx3x80x1000.</p>\n</blockquote>\n<p>Is the size of the input to the model <code>bsx3x80x1000</code> or <code>bsx3x160x1000</code>?<br>\nThe former seems to contradict the shape of x in the pseudo code.</p>",
      "rawMarkdown": "@cooolz \n\nCongratulations on your solo 2nd place. Very impressed with your very unique solution!\n\n```\nreshaped_tensor = x.view(bs, 16, 1000, 10)\nreshaped_and_permuted_tensor = reshaped_tensor.permute(0, 1, 3, 2)\nreshaped_and_permuted_tensor = reshaped_and_permuted_tensor.reshape(bs, 16 * 10, 1000)\nx = torch.unsqueeze(reshaped_and_permuted_tensor, dim=1)\nx = torch.cat([x, x, x], dim=1)\n```\n\n> Then, the eeg 'reshaped' to bsx3x80x1000.\n\nIs the size of the input to the model `bsx3x80x1000` or `bsx3x160x1000`?\nThe former seems to contradict the shape of x in the pseudo code.",
      "votes": 5,
      "replies": [
        {
          "id": 2743487,
          "postDate": "2024-04-09T13:27:56.800Z",
          "content": "<p>It should be bsx3x160x1000, Thanks for pointing out the mistake.</p>",
          "rawMarkdown": "It should be bsx3x160x1000, Thanks for pointing out the mistake.",
          "votes": 5,
          "replies": [
            {
              "id": 2743495,
              "postDate": "2024-04-09T13:36:35.067Z",
              "content": "<p>Thank you for your prompt reply!</p>",
              "rawMarkdown": "Thank you for your prompt reply!",
              "votes": 1
            },
            {
              "id": 2743606,
              "postDate": "2024-04-09T14:43:52.207Z",
              "content": "<p>thanks for the suggestion</p>",
              "rawMarkdown": "thanks for the suggestion\n",
              "votes": 1
            }
          ]
        },
        {
          "id": 2743504,
          "postDate": "2024-04-09T13:45:29.900Z",
          "content": "<p>Very nice solution. Did you see performance differences with 3-channels vs 1-channel images?</p>\n<p><code>x = torch.cat([x, x, x], dim=1)</code></p>",
          "rawMarkdown": "Very nice solution. Did you see performance differences with 3-channels vs 1-channel images?\n\n`x = torch.cat([x, x, x], dim=1)`",
          "votes": 2,
          "replies": [
            {
              "id": 2743647,
              "postDate": "2024-04-09T15:04:26.723Z",
              "content": "<p>I didn't try. I think there should be no difference after some tuning.</p>",
              "rawMarkdown": "I didn't try. I think there should be no difference after some tuning.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2742899,
      "postDate": "2024-04-09T06:15:00.850Z",
      "content": "<p>Great solution, congratulations!</p>\n<blockquote>\n  <p>2D-CNN is not well-equipped to capture positional information within the channels</p>\n</blockquote>\n<p>What do you mean here, can you elaborate?</p>",
      "rawMarkdown": "Great solution, congratulations!\n> 2D-CNN is not well-equipped to capture positional information within the channels\n\nWhat do you mean here, can you elaborate?",
      "votes": 1,
      "replies": [
        {
          "id": 2742917,
          "postDate": "2024-04-09T06:21:42.780Z",
          "content": "<p>In a paper（I forget which one）said, the position in the CNN is obtained through padding, but in general, padding is not done in the channel dimension for vision models. </p>",
          "rawMarkdown": "In a paper（I forget which one）said, the position in the CNN is obtained through padding, but in general, padding is not done in the channel dimension for vision models. ",
          "votes": 2
        },
        {
          "id": 2742921,
          "postDate": "2024-04-09T06:23:17.757Z",
          "content": "<p>《Position, Padding and Predictions:<br>\nA Deeper Look at Position Information in CNNs》     this paper</p>",
          "rawMarkdown": "《Position, Padding and Predictions:\nA Deeper Look at Position Information in CNNs》     this paper",
          "votes": 4,
          "replies": [
            {
              "id": 2742951,
              "postDate": "2024-04-09T06:41:54.477Z",
              "content": "<p>Ah I see, thanks for clarifying and for the paper link. </p>",
              "rawMarkdown": "Ah I see, thanks for clarifying and for the paper link. "
            }
          ]
        }
      ]
    },
    {
      "id": 2742695,
      "postDate": "2024-04-09T03:08:30.883Z",
      "content": "<p>In your opinion, would Vision transformer-based models for example CCT, CvT, and ViT do well on the spectrum stage?</p>",
      "rawMarkdown": "In your opinion, would Vision transformer-based models for example CCT, CvT, and ViT do well on the spectrum stage?",
      "votes": 1,
      "replies": [
        {
          "id": 2742702,
          "postDate": "2024-04-09T03:12:12.973Z",
          "content": "<p>Vit not work well in this task in my experiments. I think there is 2 reason. 1. the patch embeding has no overlap. 2. It's the same with plain vision model, not care too much on the channel direction.</p>",
          "rawMarkdown": "Vit not work well in this task in my experiments. I think there is 2 reason. 1. the patch embeding has no overlap. 2. It's the same with plain vision model, not care too much on the channel direction.",
          "votes": 1
        },
        {
          "id": 2742852,
          "postDate": "2024-04-09T05:45:45.023Z",
          "content": "<p>However, because of the position encoding, vit should work better than the cnn arch. And it seems that  was approved in other teams.</p>",
          "rawMarkdown": "However, because of the position encoding, vit should work better than the cnn arch. And it seems that  was approved in other teams.",
          "votes": 1
        },
        {
          "id": 2742968,
          "postDate": "2024-04-09T06:50:16.687Z",
          "content": "<p>I only using vit in all final submission. My single have CV 0.217, LB 0.22, Private 0.28.</p>",
          "rawMarkdown": "I only using vit in all final submission. My single have CV 0.217, LB 0.22, Private 0.28.",
          "votes": 1,
          "replies": [
            {
              "id": 2743261,
              "postDate": "2024-04-09T10:20:03.643Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2743266,
              "postDate": "2024-04-09T10:25:04.437Z",
              "content": "<p>Vit not work well, i mean there, with separate channel spectrum, not the one-image method. <br>\nAnd congratulation !!!</p>",
              "rawMarkdown": "Vit not work well, i mean there, with separate channel spectrum, not the one-image method. \nAnd congratulation !!!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2744746,
      "postDate": "2024-04-10T05:14:54.910Z",
      "content": "<p>Using 3d cnn is a great idea. Like many I was puzzled by why using channels was not effective. But you moved past it and addressed it with 3d models. Well done, and well deserved 2nd prize!</p>",
      "rawMarkdown": "Using 3d cnn is a great idea. Like many I was puzzled by why using channels was not effective. But you moved past it and addressed it with 3d models. Well done, and well deserved 2nd prize!",
      "votes": 2,
      "replies": [
        {
          "id": 2744852,
          "postDate": "2024-04-10T06:58:01.810Z",
          "content": "<p>Thanks!<br>\nI actually spent a lot of time trying to figure out why using channels was not effective in the 2D vision model. Now i think it may because of the target. For example, lpd, gpd may have same type wave shape in some channels, but in different position( channels ), <br>\nthere position( channels ) means electrode placement.</p>",
          "rawMarkdown": "Thanks!\nI actually spent a lot of time trying to figure out why using channels was not effective in the 2D vision model. Now i think it may because of the target. For example, lpd, gpd may have same type wave shape in some channels, but in different position( channels ), \nthere position( channels ) means electrode placement.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2743871,
      "postDate": "2024-04-09T16:35:22.470Z",
      "content": "<p>Well done <a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>  on solo 2nd. Impressive 💪<br>\nI had not tried stft transform as preprocess; I am interested to see a snippet of the code for preprocessing the signal for input to the model; or understand how it was done. I guess the detail is important. <br>\nWe tried x3d_m and _l as 3d model of spectrograms. They performed well, but in a blend they added no extra information over our 2d CNN. I guess with spectrum as input data it would have performed better in the blend. </p>",
      "rawMarkdown": "Well done @cooolz  on solo 2nd. Impressive 💪\nI had not tried stft transform as preprocess; I am interested to see a snippet of the code for preprocessing the signal for input to the model; or understand how it was done. I guess the detail is important. \nWe tried x3d_m and _l as 3d model of spectrograms. They performed well, but in a blend they added no extra information over our 2d CNN. I guess with spectrum as input data it would have performed better in the blend. ",
      "votes": 2,
      "replies": [
        {
          "id": 2744735,
          "postDate": "2024-04-10T05:06:50.670Z",
          "content": "<p>Thanks. <br>\nI update the stft codes in the post, torchaudio was used. I tried 50s-spectrum  and 10s-spectrum concat to 32 channels, but they not ouperform that much with just 50s-spectrum.</p>",
          "rawMarkdown": "Thanks. \nI update the stft codes in the post, torchaudio was used. I tried 50s-spectrum  and 10s-spectrum concat to 32 channels, but they not ouperform that much with just 50s-spectrum.",
          "votes": 3
        },
        {
          "id": 2744747,
          "postDate": "2024-04-10T05:17:36.797Z",
          "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> mel is really similar to stsft. The main difference is that with mel the frequency scale is logarithmic (or close to it depending on the settings you use) while it is linear for stft. Mel zooms in the low frequency are compared to stft of same size.</p>",
          "rawMarkdown": "@darraghdog mel is really similar to stsft. The main difference is that with mel the frequency scale is logarithmic (or close to it depending on the settings you use) while it is linear for stft. Mel zooms in the low frequency are compared to stft of same size.",
          "votes": 4,
          "replies": [
            {
              "id": 2744798,
              "postDate": "2024-04-10T06:15:48.013Z",
              "content": "<p>Got it, thanks !</p>",
              "rawMarkdown": "Got it, thanks !"
            },
            {
              "id": 2744838,
              "postDate": "2024-04-10T06:41:04.930Z",
              "content": "<p>Yeah, from my experiments, they perform similar, actually stft is a bit better then mel, and interestingly is scipy.spectrogram still work best for me(better then torchaudio.stft).</p>",
              "rawMarkdown": "Yeah, from my experiments, they perform similar, actually stft is a bit better then mel, and interestingly is scipy.spectrogram still work best for me(better then torchaudio.stft).",
              "votes": 1
            },
            {
              "id": 2745399,
              "postDate": "2024-04-10T15:47:47.723Z",
              "content": "<p>ah, interesting - will know for next time to try them out. I guess log vs linear frequency scaling may give some diversity on what is learned.</p>",
              "rawMarkdown": "ah, interesting - will know for next time to try them out. I guess log vs linear frequency scaling may give some diversity on what is learned.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2742774,
      "postDate": "2024-04-09T04:31:44.970Z",
      "content": "<blockquote>\n  <p>View eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.</p>\n</blockquote>\n<p>Wow, brilliant. Great idea!</p>",
      "rawMarkdown": ">View eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.\n\nWow, brilliant. Great idea!",
      "votes": 2
    },
    {
      "id": 2742773,
      "postDate": "2024-04-09T04:27:03.787Z",
      "content": "<p>Thanks for sharing , 3d image model and your method of eeg reshape seems much better then juse simplly resize, thanks for sharing, learned a lot and congratulation on solo 2nd place!</p>",
      "rawMarkdown": "Thanks for sharing , 3d image model and your method of eeg reshape seems much better then juse simplly resize, thanks for sharing, learned a lot and congratulation on solo 2nd place!",
      "votes": 2
    },
    {
      "id": 2886257,
      "postDate": "2024-06-23T14:45:59.083Z",
      "content": "<p>I will learn it in future</p>",
      "rawMarkdown": "I will learn it in future"
    },
    {
      "id": 2825957,
      "postDate": "2024-05-20T16:41:02.520Z",
      "content": "<p>nice solution for me💝</p>",
      "rawMarkdown": "nice solution for me💝"
    },
    {
      "id": 2764034,
      "postDate": "2024-04-20T20:54:56.797Z",
      "content": "<p>Congratulations! Thanks for sharing your work.  Are you using transfer learning or fitting these models from scratch.  As a beginner, can you please help me understand how you compare between model architectures (either with transfer learning or \"fresh\"?  Do you make sure to use the same 2-stage regime or did 2-stage learning come after you have chosen the final model?</p>",
      "rawMarkdown": "\nCongratulations! Thanks for sharing your work.  Are you using transfer learning or fitting these models from scratch.  As a beginner, can you please help me understand how you compare between model architectures (either with transfer learning or \"fresh\"?  Do you make sure to use the same 2-stage regime or did 2-stage learning come after you have chosen the final model?\n"
    },
    {
      "id": 2747447,
      "postDate": "2024-04-11T22:13:58.230Z",
      "content": "<p>Thank you for sharing your comprehensive approach and the detailed insights into your modeling techniques. Your innovative use of both spectral and raw EEG data with a blend of 3D and 2D CNNs provides valuable lessons on tackling position-sensitive label challenges in EEG analysis.</p>",
      "rawMarkdown": "Thank you for sharing your comprehensive approach and the detailed insights into your modeling techniques. Your innovative use of both spectral and raw EEG data with a blend of 3D and 2D CNNs provides valuable lessons on tackling position-sensitive label challenges in EEG analysis."
    },
    {
      "id": 2746452,
      "postDate": "2024-04-11T08:14:37.863Z",
      "content": "<p>Very helpful. 3D-CNN is a model structure that I have never thought of before, which has also inspired my personal research topic.</p>",
      "rawMarkdown": "Very helpful. 3D-CNN is a model structure that I have never thought of before, which has also inspired my personal research topic."
    },
    {
      "id": 2744869,
      "postDate": "2024-04-10T07:10:03.123Z",
      "content": "<p>Congratulations on second place and thanks for sharing this solution!<br>\nWe tried using 3D-CNN on the raw EEG data, but not on the spectrograms, so it is great to see that you tried it on the spectrograms and it worked.</p>",
      "rawMarkdown": "Congratulations on second place and thanks for sharing this solution!\nWe tried using 3D-CNN on the raw EEG data, but not on the spectrograms, so it is great to see that you tried it on the spectrograms and it worked."
    },
    {
      "id": 2744693,
      "postDate": "2024-04-10T04:26:31.350Z",
      "content": "<p>Congratulations on winning the 2nd prize in this competition. </p>",
      "rawMarkdown": "Congratulations on winning the 2nd prize in this competition. "
    },
    {
      "id": 2743571,
      "postDate": "2024-04-09T14:21:41.980Z",
      "content": "<p>Hello, congratulations on winning second place~~<br>\nThis competition is not an easy task for anyone.</p>\n<p>I would like to ask a question about the ensemble, which can shed some light on the technique and direction of thinking. If I have doubts about the results, where should I test or modify?<br>\nTHX!!</p>",
      "rawMarkdown": "Hello, congratulations on winning second place~~\nThis competition is not an easy task for anyone.\n\nI would like to ask a question about the ensemble, which can shed some light on the technique and direction of thinking. If I have doubts about the results, where should I test or modify?\nTHX!!"
    },
    {
      "id": 2743009,
      "postDate": "2024-04-09T07:07:34.993Z",
      "content": "<p>Interesting approach to eeg model. Thanks for sharing and congratulations!</p>",
      "rawMarkdown": "Interesting approach to eeg model. Thanks for sharing and congratulations!"
    },
    {
      "id": 2742967,
      "postDate": "2024-04-09T06:50:13.413Z",
      "content": "<p>An incredible solution! You have no idea how painful it is for me! At the very beginning of the competition, I tested the eeg as an image, but then it seemed to me that it was not working. I also tried to use 3d models, but training did not start on tensorflow (tensorflow clearly had a problem here). Congratulations again!</p>",
      "rawMarkdown": "An incredible solution! You have no idea how painful it is for me! At the very beginning of the competition, I tested the eeg as an image, but then it seemed to me that it was not working. I also tried to use 3d models, but training did not start on tensorflow (tensorflow clearly had a problem here). Congratulations again!"
    },
    {
      "id": 2742896,
      "postDate": "2024-04-09T06:13:08.250Z",
      "content": "<p>Our team never thought of 3d architecture, and I learn a lot from your method. Congratulation !</p>",
      "rawMarkdown": "Our team never thought of 3d architecture, and I learn a lot from your method. Congratulation !"
    },
    {
      "id": 2742844,
      "postDate": "2024-04-09T05:41:32.697Z",
      "content": "<p>Wow, great idea! Thanks for sharing your approach. It will be very helpful in the upcoming competitions.</p>",
      "rawMarkdown": "Wow, great idea! Thanks for sharing your approach. It will be very helpful in the upcoming competitions."
    },
    {
      "id": 2742736,
      "postDate": "2024-04-09T03:49:37.430Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2742719,
      "postDate": "2024-04-09T03:40:53.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2743455,
      "author_name": "yu4u",
      "author_url": "",
      "post_date": "2024-04-09T13:08:35.670000",
      "content": "<p><a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a> </p>\n<p>Congratulations on your solo 2nd place. Very impressed with your very unique solution!</p>\n<pre><code> = x.view(bs, , , )\n = reshaped_tensor.permute(, , , )\n = reshaped_and_permuted_tensor.reshape(bs,  * , )\n = torch.unsqueeze(reshaped_and_permuted_tensor, dim=)\n = torch.cat([x, x, x], dim=)\n</code></pre>\n<blockquote>\n  <p>Then, the eeg 'reshaped' to bsx3x80x1000.</p>\n</blockquote>\n<p>Is the size of the input to the model <code>bsx3x80x1000</code> or <code>bsx3x160x1000</code>?<br>\nThe former seems to contradict the shape of x in the pseudo code.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2743487,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-04-09T13:27:56.800000",
          "content": "<p>It should be bsx3x160x1000, Thanks for pointing out the mistake.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2743495,
              "author_name": "yu4u",
              "author_url": "",
              "post_date": "2024-04-09T13:36:35.067000",
              "content": "<p>Thank you for your prompt reply!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2743606,
              "author_name": "ZACKY_ZAC",
              "author_url": "",
              "post_date": "2024-04-09T14:43:52.207000",
              "content": "<p>thanks for the suggestion</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2743504,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-04-09T13:45:29.900000",
          "content": "<p>Very nice solution. Did you see performance differences with 3-channels vs 1-channel images?</p>\n<p><code>x = torch.cat([x, x, x], dim=1)</code></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2743647,
              "author_name": "coolz",
              "author_url": "",
              "post_date": "2024-04-09T15:04:26.723000",
              "content": "<p>I didn't try. I think there should be no difference after some tuning.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2742899,
      "author_name": "Viktor Cikojevic",
      "author_url": "",
      "post_date": "2024-04-09T06:15:00.850000",
      "content": "<p>Great solution, congratulations!</p>\n<blockquote>\n  <p>2D-CNN is not well-equipped to capture positional information within the channels</p>\n</blockquote>\n<p>What do you mean here, can you elaborate?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2742917,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-04-09T06:21:42.780000",
          "content": "<p>In a paper（I forget which one）said, the position in the CNN is obtained through padding, but in general, padding is not done in the channel dimension for vision models. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2742921,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-04-09T06:23:17.757000",
          "content": "<p>《Position, Padding and Predictions:<br>\nA Deeper Look at Position Information in CNNs》     this paper</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2742951,
              "author_name": "Viktor Cikojevic",
              "author_url": "",
              "post_date": "2024-04-09T06:41:54.477000",
              "content": "<p>Ah I see, thanks for clarifying and for the paper link. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2742695,
      "author_name": "MoZayed",
      "author_url": "",
      "post_date": "2024-04-09T03:08:30.883000",
      "content": "<p>In your opinion, would Vision transformer-based models for example CCT, CvT, and ViT do well on the spectrum stage?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2742702,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-04-09T03:12:12.973000",
          "content": "<p>Vit not work well in this task in my experiments. I think there is 2 reason. 1. the patch embeding has no overlap. 2. It's the same with plain vision model, not care too much on the channel direction.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2742852,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-04-09T05:45:45.023000",
          "content": "<p>However, because of the position encoding, vit should work better than the cnn arch. And it seems that  was approved in other teams.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2742968,
          "author_name": "Quan Vu",
          "author_url": "",
          "post_date": "2024-04-09T06:50:16.687000",
          "content": "<p>I only using vit in all final submission. My single have CV 0.217, LB 0.22, Private 0.28.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2743261,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-09T10:20:03.643000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2743266,
              "author_name": "coolz",
              "author_url": "",
              "post_date": "2024-04-09T10:25:04.437000",
              "content": "<p>Vit not work well, i mean there, with separate channel spectrum, not the one-image method. <br>\nAnd congratulation !!!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2744746,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-04-10T05:14:54.910000",
      "content": "<p>Using 3d cnn is a great idea. Like many I was puzzled by why using channels was not effective. But you moved past it and addressed it with 3d models. Well done, and well deserved 2nd prize!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2744852,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-04-10T06:58:01.810000",
          "content": "<p>Thanks!<br>\nI actually spent a lot of time trying to figure out why using channels was not effective in the 2D vision model. Now i think it may because of the target. For example, lpd, gpd may have same type wave shape in some channels, but in different position( channels ), <br>\nthere position( channels ) means electrode placement.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2743871,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2024-04-09T16:35:22.470000",
      "content": "<p>Well done <a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>  on solo 2nd. Impressive 💪<br>\nI had not tried stft transform as preprocess; I am interested to see a snippet of the code for preprocessing the signal for input to the model; or understand how it was done. I guess the detail is important. <br>\nWe tried x3d_m and _l as 3d model of spectrograms. They performed well, but in a blend they added no extra information over our 2d CNN. I guess with spectrum as input data it would have performed better in the blend. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2744735,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-04-10T05:06:50.670000",
          "content": "<p>Thanks. <br>\nI update the stft codes in the post, torchaudio was used. I tried 50s-spectrum  and 10s-spectrum concat to 32 channels, but they not ouperform that much with just 50s-spectrum.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2744747,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2024-04-10T05:17:36.797000",
          "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> mel is really similar to stsft. The main difference is that with mel the frequency scale is logarithmic (or close to it depending on the settings you use) while it is linear for stft. Mel zooms in the low frequency are compared to stft of same size.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2744798,
              "author_name": "Darragh",
              "author_url": "",
              "post_date": "2024-04-10T06:15:48.013000",
              "content": "<p>Got it, thanks !</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2744838,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2024-04-10T06:41:04.930000",
              "content": "<p>Yeah, from my experiments, they perform similar, actually stft is a bit better then mel, and interestingly is scipy.spectrogram still work best for me(better then torchaudio.stft).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2745399,
              "author_name": "Darragh",
              "author_url": "",
              "post_date": "2024-04-10T15:47:47.723000",
              "content": "<p>ah, interesting - will know for next time to try them out. I guess log vs linear frequency scaling may give some diversity on what is learned.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2742774,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-04-09T04:31:44.970000",
      "content": "<blockquote>\n  <p>View eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.</p>\n</blockquote>\n<p>Wow, brilliant. Great idea!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2742773,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2024-04-09T04:27:03.787000",
      "content": "<p>Thanks for sharing , 3d image model and your method of eeg reshape seems much better then juse simplly resize, thanks for sharing, learned a lot and congratulation on solo 2nd place!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2886257,
      "author_name": "DongLibin",
      "author_url": "",
      "post_date": "2024-06-23T14:45:59.083000",
      "content": "<p>I will learn it in future</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2825957,
      "author_name": "Wachirawit Premthaisong-KMUTT",
      "author_url": "",
      "post_date": "2024-05-20T16:41:02.520000",
      "content": "<p>nice solution for me💝</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2764034,
      "author_name": "Idith Haber",
      "author_url": "",
      "post_date": "2024-04-20T20:54:56.797000",
      "content": "<p>Congratulations! Thanks for sharing your work.  Are you using transfer learning or fitting these models from scratch.  As a beginner, can you please help me understand how you compare between model architectures (either with transfer learning or \"fresh\"?  Do you make sure to use the same 2-stage regime or did 2-stage learning come after you have chosen the final model?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2747447,
      "author_name": "Suhasini Polampelly",
      "author_url": "",
      "post_date": "2024-04-11T22:13:58.230000",
      "content": "<p>Thank you for sharing your comprehensive approach and the detailed insights into your modeling techniques. Your innovative use of both spectral and raw EEG data with a blend of 3D and 2D CNNs provides valuable lessons on tackling position-sensitive label challenges in EEG analysis.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2746452,
      "author_name": "Yuri Sun",
      "author_url": "",
      "post_date": "2024-04-11T08:14:37.863000",
      "content": "<p>Very helpful. 3D-CNN is a model structure that I have never thought of before, which has also inspired my personal research topic.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2744869,
      "author_name": "movieminer",
      "author_url": "",
      "post_date": "2024-04-10T07:10:03.123000",
      "content": "<p>Congratulations on second place and thanks for sharing this solution!<br>\nWe tried using 3D-CNN on the raw EEG data, but not on the spectrograms, so it is great to see that you tried it on the spectrograms and it worked.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2744693,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-04-10T04:26:31.350000",
      "content": "<p>Congratulations on winning the 2nd prize in this competition. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2743571,
      "author_name": "warrenkuo",
      "author_url": "",
      "post_date": "2024-04-09T14:21:41.980000",
      "content": "<p>Hello, congratulations on winning second place~~<br>\nThis competition is not an easy task for anyone.</p>\n<p>I would like to ask a question about the ensemble, which can shed some light on the technique and direction of thinking. If I have doubts about the results, where should I test or modify?<br>\nTHX!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2743009,
      "author_name": "Viji",
      "author_url": "",
      "post_date": "2024-04-09T07:07:34.993000",
      "content": "<p>Interesting approach to eeg model. Thanks for sharing and congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2742967,
      "author_name": "Andrij",
      "author_url": "",
      "post_date": "2024-04-09T06:50:13.413000",
      "content": "<p>An incredible solution! You have no idea how painful it is for me! At the very beginning of the competition, I tested the eeg as an image, but then it seemed to me that it was not working. I also tried to use 3d models, but training did not start on tensorflow (tensorflow clearly had a problem here). Congratulations again!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2742896,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2024-04-09T06:13:08.250000",
      "content": "<p>Our team never thought of 3d architecture, and I learn a lot from your method. Congratulation !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2742844,
      "author_name": "Aaditya Porwal",
      "author_url": "",
      "post_date": "2024-04-09T05:41:32.697000",
      "content": "<p>Wow, great idea! Thanks for sharing your approach. It will be very helpful in the upcoming competitions.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2742736,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-09T03:49:37.430000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2742719,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-09T03:40:53.780000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742687": "First of all, I would like to thanks to Kaggle and the organizers. I have learned a lot from this experience.  \n\n\n\n##  1. model\n\nIn early time, feeding a bsx4xHxW input to the 2D-CNN produce worse result than the one-image method.  I start to think why?\n\nDue to the position-sensitive nature of the labels (LPD, GPD, etc.), We should take care of the channel dimension. 2D-CNN is not well-equipped to capture positional information within the channels. Because there is not pad in channel direction. And that's why we need to concat the spectrum into one image for a vision model other than a bsx16xHxW image( double banana montage).\nSo I decided to use 3d-CNN for spectrum, and 2d-CNN model for raw-eeg signal.\n\n\ndiagram like:\n\n\ntotal solution like this:\n![solution figure](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2F9d2ab8354f85f39d092a75fd13922db7%2Fimg.png?generation=1715226166682613&alt=media)\n+ mne 0.5-20hz means use MNE-tool do filter.\n+ scipy.signal means use scipy.signal to do filter.\n+ For the reshape operator and the stft params please refer to the codes below.\n+ And the number in () means final weights of the ensemble.\n\n### 1.1 x3d-l (spectrum model)\n\n\nAfter double banana montage, +-1024 clip and 0.5-20hz filter,\nuse stft to extract the spectrum, then feed to a 3d-CNN(x3d-l).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1562037%2Fcbfe301faa8741dc262aa334e620e4c9%2Fimg_2.png?generation=1715226202040353&alt=media)\n\nThe input data is a 16-channels spectrum image.\n\nX3d-l cv 0.21+, public lb0.25 private lb0.29.\n\n\n\nstft code like below, use it as a nn.Module :\n\n```\nclass Transform50s(nn.Module):\n    def __init__(self, ):\n        super().__init__()\n        self.wave_transform = torchaudio.transforms.Spectrogram(n_fft=512, win_length=128, hop_length=50, power=1)\n\n    def forward(self, x):\n        image = self.wave_transform(x)\n        image = torch.clip(image, min=0, max=10000) / 1000\n        n, c, h, w = image.size()\n        image = image[:, :, :int(20 / 100 * h + 10), :]\n        return image\n\n\nclass Transform10s(nn.Module):\n    def __init__(self, ):\n        super().__init__()\n        self.wave_transform = torchaudio.transforms.Spectrogram(n_fft=512, win_length=128, hop_length=10, power=1)\n\n    def forward(self, x):\n        image = self.wave_transform(x)\n        image = torch.clip(image, min=0, max=10000) / 1000\n        n, c, h, w = image.size()\n        image = image[:, :, :int(20 / 100 * h + 10), :]\n        return image\n\n\nclass Model(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        model_name = \"x3d_l\"\n        self.net = torch.hub.load('facebookresearch/pytorchvideo',\n                                  model_name, pretrained=True)\n\n        self.net.blocks[5].pool.pool = nn.AdaptiveAvgPool3d(1)\n        # self.net.blocks[5]=nn.Identity()\n        # self.net.avgpool = nn.Identity()\n        self.net.blocks[5].dropout = nn.Identity()\n        self.net.blocks[5].proj = nn.Identity()\n        self.net.blocks[5].activation = nn.Identity()\n        self.net.blocks[5].output_pool = nn.Identity()\n\n    def forward(self, x):\n        x = self.net(x)\n        return x\n\n\nclass Net(nn.Module):\n    def __init__(self, num_classes=1):\n        super().__init__()\n\n        self.preprocess50s = Transform50s()\n        self.preprocess10s = Transform10s()\n\n        self.model = Model()\n\n        self.pool = nn.AdaptiveAvgPool3d(1)\n        self.fc = nn.Linear(2048, 6, bias=True)\n        self.drop = nn.Dropout(0.5)\n\n    def forward(self, eeg):\n        # do preprocess\n        bs = eeg.size(0)\n        eeg_50s = eeg\n        eeg_10s = eeg[:, :, 4000:6000]\n        x_50 = self.preprocess50s(eeg_50s)\n        x_10 = self.preprocess10s(eeg_10s)\n        x = torch.cat([x_10, x_50], dim=1)\n        x = torch.unsqueeze(x, dim=1)\n        x = torch.cat([x, x, x], dim=1)\n        x = self.model(x)\n        # x = self.pool(x)\n        x = x.view(bs, -1)\n        x = self.drop(x)\n        x = self.fc(x)\n        return x\n\n```\n\n### 1.2 oneimage\nIt's a 2d vision model, efficientnetb5, public lb 0.26 2490 private lb0.304877.\n\nFor 2d model, I concat the 16-channels spectrum image like many people does.\n\n```\n        image = torch.reshape(image, shape=[n, 2, -1, w])\n        x1 = image[:, 0:1, ...]\n        x2 = image[:, 1:2, ...]\n        image = torch.cat([x1, x2], dim=-1)\n        image = torch.cat([image, image, image], dim=1)\n```\n\nI check the submission score, combine these two spectrum model, get 0.28 private lb.\n\n\n### 1.3 eeg model\n\nView eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.\nThe model define as below\n```python\nrclass Net(nn.Module):\n    def __init__(self,):\n        super(Net, self).__init__()\n        self.model = timm.create_model('efficientnet_b5', pretrained=True, in_chans=3)\n        self.pool = nn.AdaptiveAvgPool2d(1)\n        self.fc = nn.Linear(2048, out_features=6, bias=True)\n        self.dropout = nn.Dropout(p=0.5)\n\n    def extract_features(self, x):\n        feature1 = self.model.forward_features(x)\n        return feature1\n\n    def forward(self, x):\n        bs = x.size(0)\n        reshaped_tensor = x.view(bs, 16, 1000, 10)\n        reshaped_and_permuted_tensor = reshaped_tensor.permute(0, 1, 3, 2)\n        reshaped_and_permuted_tensor = reshaped_and_permuted_tensor.reshape(bs, 16 * 10, 1000)\n        x = torch.unsqueeze(reshaped_and_permuted_tensor, dim=1)\n\n        x = torch.cat([x, x, x], dim=1)\n        bs = x.size(0)\n\n        x = self.extract_features(x)\n        x = self.pool(x)\n        x = x.view(bs, -1)\n        x = self.dropout(x)\n        x = self.fc(x)\n        return x\n```\nThen, the eeg 'reshaped' to bsx3x160x1000.  With efficinetnetb5, actives public lb 0.230 978, private lb 0.282 873.\n\nThere are 3 models , but slightly different, mne-filter-efficinetnetb5, scipy.signal-filter-efficientnetb5 and mne-filter-hgnetb5. \nAnd the scires are slightly different,also.\n\n### 1.4  doublehead (eeg+spectrum)\n\nI use x3d-l to extract the spectrum feature ( with Transform50s only ), efficientnetb5 to extract the raw eeg feature, \nlike the solution figure.\nconcat the last feature. public lb 0.24, private lb 0.29 ? ? not sure\n\n# 2 Preprocess\n- 2.1.Double banana montage, eeg as 16x10000\n- 2.2.Filter with 0.5-20hz\n- 2.3.Clip by +-1024\n\n# 3. Train\n- 3.1 Stage1, 15 epochs, with loss weight voters_num/20, Adamw lr=0.001, cos scheduel\n- 3.2 Stage2, 5 epoch, loss weight=1, voters_num>=6.   , Adamw lr=0.0001, cos scheduel\n- 3.3 use eeg_label_offset_seconds. I random choose an offset for each eegid, and the target is average according to eegid for each train iter.\n- 3.4 augmentation, mirror eeg, flip between left brain data and right brain data.\n- 3.5 10 folds, then move 1000 samples from val to train, left 709 samples in val set. And use vote_num>=6 to do validation. \n\n# 4. ensemble\n\nBy combining these models, I think it could get my current score. However my score is 6 model emsemble, but not improvement that much (0.28->0.27 private lb). \nFinal ensemble including:  \n- 2 spectrum model (x3d, efficientnetb5), both use mne-filter\n- 3 raw eeg model( efficientnetb5 with-mne.filter, efficientnetb5 -butter filter, 1 hgnetb5 mne.filter),\n- 1 eeg-spectrum mix model, butter filter,   \n\nWith weights=[0.1,0.1,0.2,0.2,0.2,0.2].\n\nps. with-mne.filter means use mne lib to do filter, butter filter means use scipy.signal, just to add some diversity.\n\n# Some thought\nI think raw eeg is more important in this task.  Feeding raw EEG data into a 2D visual model is somewhat similar to how humans observe EEG signals. Observe in time and channel  dimensions！\n\n",
    "2743455": "@cooolz \n\nCongratulations on your solo 2nd place. Very impressed with your very unique solution!\n\n```\nreshaped_tensor = x.view(bs, 16, 1000, 10)\nreshaped_and_permuted_tensor = reshaped_tensor.permute(0, 1, 3, 2)\nreshaped_and_permuted_tensor = reshaped_and_permuted_tensor.reshape(bs, 16 * 10, 1000)\nx = torch.unsqueeze(reshaped_and_permuted_tensor, dim=1)\nx = torch.cat([x, x, x], dim=1)\n```\n\n> Then, the eeg 'reshaped' to bsx3x80x1000.\n\nIs the size of the input to the model `bsx3x80x1000` or `bsx3x160x1000`?\nThe former seems to contradict the shape of x in the pseudo code.",
    "2742899": "Great solution, congratulations!\n> 2D-CNN is not well-equipped to capture positional information within the channels\n\nWhat do you mean here, can you elaborate?",
    "2742695": "In your opinion, would Vision transformer-based models for example CCT, CvT, and ViT do well on the spectrum stage?",
    "2744746": "Using 3d cnn is a great idea. Like many I was puzzled by why using channels was not effective. But you moved past it and addressed it with 3d models. Well done, and well deserved 2nd prize!",
    "2743871": "Well done @cooolz  on solo 2nd. Impressive 💪\nI had not tried stft transform as preprocess; I am interested to see a snippet of the code for preprocessing the signal for input to the model; or understand how it was done. I guess the detail is important. \nWe tried x3d_m and _l as 3d model of spectrograms. They performed well, but in a blend they added no extra information over our 2d CNN. I guess with spectrum as input data it would have performed better in the blend. ",
    "2742774": ">View eeg (bsx16x10000) as an image. So expand dim=1( bsx1x16x10000), but the time dimension is too large. Then I make a reshape.\n\nWow, brilliant. Great idea!",
    "2742773": "Thanks for sharing , 3d image model and your method of eeg reshape seems much better then juse simplly resize, thanks for sharing, learned a lot and congratulation on solo 2nd place!",
    "2886257": "I will learn it in future",
    "2825957": "nice solution for me💝",
    "2764034": "\nCongratulations! Thanks for sharing your work.  Are you using transfer learning or fitting these models from scratch.  As a beginner, can you please help me understand how you compare between model architectures (either with transfer learning or \"fresh\"?  Do you make sure to use the same 2-stage regime or did 2-stage learning come after you have chosen the final model?\n",
    "2747447": "Thank you for sharing your comprehensive approach and the detailed insights into your modeling techniques. Your innovative use of both spectral and raw EEG data with a blend of 3D and 2D CNNs provides valuable lessons on tackling position-sensitive label challenges in EEG analysis.",
    "2746452": "Very helpful. 3D-CNN is a model structure that I have never thought of before, which has also inspired my personal research topic.",
    "2744869": "Congratulations on second place and thanks for sharing this solution!\nWe tried using 3D-CNN on the raw EEG data, but not on the spectrograms, so it is great to see that you tried it on the spectrograms and it worked.",
    "2744693": "Congratulations on winning the 2nd prize in this competition. ",
    "2743571": "Hello, congratulations on winning second place~~\nThis competition is not an easy task for anyone.\n\nI would like to ask a question about the ensemble, which can shed some light on the technique and direction of thinking. If I have doubts about the results, where should I test or modify?\nTHX!!",
    "2743009": "Interesting approach to eeg model. Thanks for sharing and congratulations!",
    "2742967": "An incredible solution! You have no idea how painful it is for me! At the very beginning of the competition, I tested the eeg as an image, but then it seemed to me that it was not working. I also tried to use 3d models, but training did not start on tensorflow (tensorflow clearly had a problem here). Congratulations again!",
    "2742896": "Our team never thought of 3d architecture, and I learn a lot from your method. Congratulation !",
    "2742844": "Wow, great idea! Thanks for sharing your approach. It will be very helpful in the upcoming competitions.",
    "2742736": "",
    "2742719": ""
  }
}