{
  "id": 481700,
  "title": "Channel Attention Hybrid Transformer",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/481700",
  "author_name": "Gunes Evitan",
  "post_date": "2024-03-04T17:25:35.433000",
  "votes": 26,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I found <a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a>'s paper <a href=\"https://ieeexplore.ieee.org/document/9858598\" target=\"_blank\">here</a> implemented in on pytorch but couldn't make it work. I'm not good at transformers so I'm sharing it and maybe someone else can make it work.</p>\n<pre><code> numpy  np\n torch\n torch.nn  nn\n\n heads  ClassificationHead\n\n\n ():\n\n    dim = embed_dim // \n\n    position = np.arange(length)[:, np.newaxis]\n    dim = np.arange(dim)[np.newaxis, :] / dim\n\n    angle =  / ( ** dim)\n    angle = position * angle\n\n    pos_embed = np.concatenate([np.sin(angle), np.cos(angle)], axis=-)\n    pos_embed = torch.from_numpy(pos_embed).()\n\n     pos_embed\n\n\n (nn.Module):\n\n     ():\n\n        (ConvBlock, self).__init__()\n\n        self.conv1 = nn.Conv2d(in_channels=in_channels, out_channels=, kernel_size=(, ), stride=(, ), padding=(, ), padding_mode=, bias=)\n        self.bn1 = nn.BatchNorm2d(num_features=)\n        self.conv2 = nn.Conv2d(in_channels=, out_channels=, kernel_size=(, ), stride=(, ), padding=(, ), padding_mode=, bias=)\n        self.bn2 = nn.BatchNorm2d(num_features=)\n        self.conv3 = nn.Conv2d(in_channels=, out_channels=out_channels, kernel_size=(, ), stride=(, ), padding=(, ), padding_mode=, bias=)\n        self.activation = nn.LeakyReLU(inplace=)\n\n     ():\n\n        x = self.conv1(x)\n        x = self.bn1(x)\n        x = self.activation(x)\n\n        x = self.conv2(x)\n        x = self.bn2(x)\n        x = self.activation(x)\n\n        x = self.conv3(x)\n\n         x\n\n\n (nn.Module):\n\n     ():\n\n        (MLP, self).__init__()\n\n        self.mlp = nn.Sequential(\n            nn.Linear(in_features=embed_dim, out_features=hidden_dim),\n            nn.ReLU(inplace=),\n            nn.Linear(in_features=hidden_dim, out_features=embed_dim),\n        )\n\n     ():\n         self.mlp(x)\n\n\n (nn.Module):\n\n     ():\n        (TransformerBlock, self).__init__()\n\n        self.attention = nn.MultiheadAttention(embed_dim=embed_dim, num_heads=num_heads, batch_first=)\n        self.mlp = MLP(embed_dim=embed_dim, hidden_dim=out_dim)\n        self.ln1 = nn.LayerNorm(embed_dim)\n        self.ln2 = nn.LayerNorm(out_dim)\n\n     ():\n\n        x = self.ln1(x)\n        x = self.attention(x, x, x)[]\n        x = self.ln2(x)\n        x = self.mlp(x)\n\n         x\n\n\n (nn.Module):\n\n     ():\n\n        (HybridTransformer, self).__init__()\n\n        self.conv_block = ConvBlock(in_channels=, out_channels=hidden_size)\n        self.encoder = nn.ModuleList([\n            TransformerBlock(\n                embed_dim=hidden_size,\n                num_heads=num_heads,\n                out_dim=hidden_size,\n            )  _  (num_blocks)\n        ])\n\n        self.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))\n        self.cls_token = nn.Parameter(torch.zeros((, hidden_size)))\n\n        self.pooling_type = pooling_type\n        self.dropout = nn.Dropout(dropout_rate)  dropout_rate &gt;   nn.Identity()\n        self.head = ClassificationHead(input_dimensions=hidden_size, **head_args)\n\n     ():\n\n        \n        x = x.unsqueeze(dim=).permute(, , , )\n        x = self.conv_block(x)\n\n        \n        x = torch.mean(x, dim=).permute(, , )\n\n        \n        x += self.positional_embeddings\n        x = torch.cat([\n            self.cls_token.unsqueeze().repeat(x.size(), , ),\n            x\n        ], )\n\n        \n         block  self.encoder:\n            x = block(x)\n\n         self.pooling_type == :\n            x = torch.mean(x, dim=)\n         self.pooling_type == :\n            x = torch.(x, dim=)[]\n         self.pooling_type == :\n            x = x[:, , :]\n\n        x = self.dropout(x)\n        output = self.head(x)\n\n         output\n</code></pre>",
  "messages": [
    {
      "id": 2681525,
      "postDate": "2024-03-04T17:25:35.433Z",
      "content": "<p>I found <a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a>'s paper <a href=\"https://ieeexplore.ieee.org/document/9858598\" target=\"_blank\">here</a> implemented in on pytorch but couldn't make it work. I'm not good at transformers so I'm sharing it and maybe someone else can make it work.</p>\n<pre><code> numpy  np\n torch\n torch.nn  nn\n\n heads  ClassificationHead\n\n\n ():\n\n    dim = embed_dim // \n\n    position = np.arange(length)[:, np.newaxis]\n    dim = np.arange(dim)[np.newaxis, :] / dim\n\n    angle =  / ( ** dim)\n    angle = position * angle\n\n    pos_embed = np.concatenate([np.sin(angle), np.cos(angle)], axis=-)\n    pos_embed = torch.from_numpy(pos_embed).()\n\n     pos_embed\n\n\n (nn.Module):\n\n     ():\n\n        (ConvBlock, self).__init__()\n\n        self.conv1 = nn.Conv2d(in_channels=in_channels, out_channels=, kernel_size=(, ), stride=(, ), padding=(, ), padding_mode=, bias=)\n        self.bn1 = nn.BatchNorm2d(num_features=)\n        self.conv2 = nn.Conv2d(in_channels=, out_channels=, kernel_size=(, ), stride=(, ), padding=(, ), padding_mode=, bias=)\n        self.bn2 = nn.BatchNorm2d(num_features=)\n        self.conv3 = nn.Conv2d(in_channels=, out_channels=out_channels, kernel_size=(, ), stride=(, ), padding=(, ), padding_mode=, bias=)\n        self.activation = nn.LeakyReLU(inplace=)\n\n     ():\n\n        x = self.conv1(x)\n        x = self.bn1(x)\n        x = self.activation(x)\n\n        x = self.conv2(x)\n        x = self.bn2(x)\n        x = self.activation(x)\n\n        x = self.conv3(x)\n\n         x\n\n\n (nn.Module):\n\n     ():\n\n        (MLP, self).__init__()\n\n        self.mlp = nn.Sequential(\n            nn.Linear(in_features=embed_dim, out_features=hidden_dim),\n            nn.ReLU(inplace=),\n            nn.Linear(in_features=hidden_dim, out_features=embed_dim),\n        )\n\n     ():\n         self.mlp(x)\n\n\n (nn.Module):\n\n     ():\n        (TransformerBlock, self).__init__()\n\n        self.attention = nn.MultiheadAttention(embed_dim=embed_dim, num_heads=num_heads, batch_first=)\n        self.mlp = MLP(embed_dim=embed_dim, hidden_dim=out_dim)\n        self.ln1 = nn.LayerNorm(embed_dim)\n        self.ln2 = nn.LayerNorm(out_dim)\n\n     ():\n\n        x = self.ln1(x)\n        x = self.attention(x, x, x)[]\n        x = self.ln2(x)\n        x = self.mlp(x)\n\n         x\n\n\n (nn.Module):\n\n     ():\n\n        (HybridTransformer, self).__init__()\n\n        self.conv_block = ConvBlock(in_channels=, out_channels=hidden_size)\n        self.encoder = nn.ModuleList([\n            TransformerBlock(\n                embed_dim=hidden_size,\n                num_heads=num_heads,\n                out_dim=hidden_size,\n            )  _  (num_blocks)\n        ])\n\n        self.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))\n        self.cls_token = nn.Parameter(torch.zeros((, hidden_size)))\n\n        self.pooling_type = pooling_type\n        self.dropout = nn.Dropout(dropout_rate)  dropout_rate &gt;   nn.Identity()\n        self.head = ClassificationHead(input_dimensions=hidden_size, **head_args)\n\n     ():\n\n        \n        x = x.unsqueeze(dim=).permute(, , , )\n        x = self.conv_block(x)\n\n        \n        x = torch.mean(x, dim=).permute(, , )\n\n        \n        x += self.positional_embeddings\n        x = torch.cat([\n            self.cls_token.unsqueeze().repeat(x.size(), , ),\n            x\n        ], )\n\n        \n         block  self.encoder:\n            x = block(x)\n\n         self.pooling_type == :\n            x = torch.mean(x, dim=)\n         self.pooling_type == :\n            x = torch.(x, dim=)[]\n         self.pooling_type == :\n            x = x[:, , :]\n\n        x = self.dropout(x)\n        output = self.head(x)\n\n         output\n</code></pre>",
      "rawMarkdown": "I found @sunyuri's paper [here](https://ieeexplore.ieee.org/document/9858598) implemented in on pytorch but couldn't make it work. I'm not good at transformers so I'm sharing it and maybe someone else can make it work.\n\n```\nimport numpy as np\nimport torch\nimport torch.nn as nn\n\nfrom heads import ClassificationHead\n\n\ndef positional_encoding(length, embed_dim):\n\n    dim = embed_dim // 2\n\n    position = np.arange(length)[:, np.newaxis]\n    dim = np.arange(dim)[np.newaxis, :] / dim\n\n    angle = 1 / (10000 ** dim)\n    angle = position * angle\n\n    pos_embed = np.concatenate([np.sin(angle), np.cos(angle)], axis=-1)\n    pos_embed = torch.from_numpy(pos_embed).float()\n\n    return pos_embed\n\n\nclass ConvBlock(nn.Module):\n\n    def __init__(self, in_channels, out_channels):\n\n        super(ConvBlock, self).__init__()\n\n        self.conv1 = nn.Conv2d(in_channels=in_channels, out_channels=32, kernel_size=(4, 1), stride=(2, 1), padding=(1, 0), padding_mode='zeros', bias=True)\n        self.bn1 = nn.BatchNorm2d(num_features=32)\n        self.conv2 = nn.Conv2d(in_channels=32, out_channels=64, kernel_size=(4, 1), stride=(2, 1), padding=(1, 0), padding_mode='zeros', bias=True)\n        self.bn2 = nn.BatchNorm2d(num_features=64)\n        self.conv3 = nn.Conv2d(in_channels=64, out_channels=out_channels, kernel_size=(4, 1), stride=(2, 1), padding=(1, 0), padding_mode='zeros', bias=True)\n        self.activation = nn.LeakyReLU(inplace=False)\n\n    def forward(self, x):\n\n        x = self.conv1(x)\n        x = self.bn1(x)\n        x = self.activation(x)\n\n        x = self.conv2(x)\n        x = self.bn2(x)\n        x = self.activation(x)\n\n        x = self.conv3(x)\n\n        return x\n\n\nclass MLP(nn.Module):\n\n    def __init__(self, embed_dim, hidden_dim):\n\n        super(MLP, self).__init__()\n\n        self.mlp = nn.Sequential(\n            nn.Linear(in_features=embed_dim, out_features=hidden_dim),\n            nn.ReLU(inplace=True),\n            nn.Linear(in_features=hidden_dim, out_features=embed_dim),\n        )\n\n    def forward(self, x):\n        return self.mlp(x)\n\n\nclass TransformerBlock(nn.Module):\n\n    def __init__(self, embed_dim, num_heads, out_dim):\n        super(TransformerBlock, self).__init__()\n\n        self.attention = nn.MultiheadAttention(embed_dim=embed_dim, num_heads=num_heads, batch_first=True)\n        self.mlp = MLP(embed_dim=embed_dim, hidden_dim=out_dim)\n        self.ln1 = nn.LayerNorm(embed_dim)\n        self.ln2 = nn.LayerNorm(out_dim)\n\n    def forward(self, x):\n\n        x = self.ln1(x)\n        x = self.attention(x, x, x)[0]\n        x = self.ln2(x)\n        x = self.mlp(x)\n\n        return x\n\n\nclass HybridTransformer(nn.Module):\n\n    def __init__(self, in_channels, hidden_size, num_heads, num_blocks, pooling_type, dropout_rate, head_args):\n\n        super(HybridTransformer, self).__init__()\n\n        self.conv_block = ConvBlock(in_channels=1, out_channels=hidden_size)\n        self.encoder = nn.ModuleList([\n            TransformerBlock(\n                embed_dim=hidden_size,\n                num_heads=num_heads,\n                out_dim=hidden_size,\n            ) for _ in range(num_blocks)\n        ])\n\n        self.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))\n        self.cls_token = nn.Parameter(torch.zeros((1, hidden_size)))\n\n        self.pooling_type = pooling_type\n        self.dropout = nn.Dropout(dropout_rate) if dropout_rate > 0 else nn.Identity()\n        self.head = ClassificationHead(input_dimensions=hidden_size, **head_args)\n\n    def forward(self, x):\n\n        # Add channel dimension and pass it to conv 2d block\n        x = x.unsqueeze(dim=1).permute(0, 1, 3, 2)\n        x = self.conv_block(x)\n\n        # Average features along time dimension\n        x = torch.mean(x, dim=2).permute(0, 2, 1)\n\n        # Add positional embeddings and concatenate cls token\n        x += self.positional_embeddings\n        x = torch.cat([\n            self.cls_token.unsqueeze(0).repeat(x.size(0), 1, 1),\n            x\n        ], 1)\n\n        # Pass it to transformer encoder\n        for block in self.encoder:\n            x = block(x)\n\n        if self.pooling_type == 'avg':\n            x = torch.mean(x, dim=1)\n        elif self.pooling_type == 'max':\n            x = torch.max(x, dim=1)[0]\n        elif self.pooling_type == 'cls':\n            x = x[:, 0, :]\n\n        x = self.dropout(x)\n        output = self.head(x)\n\n        return output\n```",
      "votes": 26
    },
    {
      "id": 2681972,
      "postDate": "2024-03-05T03:19:07.347Z",
      "content": "<p>I have tried using the attention mechanism in this competition, but the results were not as good as using pre-trained EfficientNet-B0. My initial idea was to treat the four brain regions of LL, RL, LP, and RP as four tokens, which would be a very short \"sentence\".</p>",
      "rawMarkdown": "I have tried using the attention mechanism in this competition, but the results were not as good as using pre-trained EfficientNet-B0. My initial idea was to treat the four brain regions of LL, RL, LP, and RP as four tokens, which would be a very short \"sentence\".",
      "votes": 7,
      "replies": [
        {
          "id": 2682293,
          "postDate": "2024-03-05T08:17:04.803Z",
          "content": "<p>I also conduct a similar experiment, but it's not better than concatenation (like Chris's works)</p>",
          "rawMarkdown": "I also conduct a similar experiment, but it's not better than concatenation (like Chris's works)",
          "votes": 2,
          "replies": [
            {
              "id": 2682313,
              "postDate": "2024-03-05T08:40:42.130Z",
              "content": "<p>This result is confusing for me, perhaps finding a suitable concatenation method is the key. For the original EEG signal, the performance of EEGNet is better than that of the EfficientNet without pre-trained weights, but at the same time, it is lower than that of pre-trained EfficientNet (In my experiments).<br>\nI think we need a pre-trained EEG large model (like this <a href=\"https://openreview.net/pdf?id=QzTpTRVtrP\" target=\"_blank\">paper</a>)</p>",
              "rawMarkdown": "This result is confusing for me, perhaps finding a suitable concatenation method is the key. For the original EEG signal, the performance of EEGNet is better than that of the EfficientNet without pre-trained weights, but at the same time, it is lower than that of pre-trained EfficientNet (In my experiments).\nI think we need a pre-trained EEG large model (like this [paper](https://openreview.net/pdf?id=QzTpTRVtrP))",
              "votes": 10
            },
            {
              "id": 2682436,
              "postDate": "2024-03-05T10:25:28.163Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> , by  <code>pre-trained EfficientNet</code> you mean ImageNet weights?</p>",
              "rawMarkdown": "Hi @sunyuri , by  `pre-trained EfficientNet` you mean ImageNet weights?"
            },
            {
              "id": 2682542,
              "postDate": "2024-03-05T11:41:18.823Z",
              "content": "<p>Yes.             </p>",
              "rawMarkdown": "Yes.             ",
              "votes": 2
            },
            {
              "id": 2698816,
              "postDate": "2024-03-15T16:47:23.800Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2698821,
              "postDate": "2024-03-15T16:49:33.203Z",
              "content": "<p><a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> By \"EEGNet\" do you mean this EEGNet? <a href=\"https://github.com/vlawhern/arl-eegmodels\" target=\"_blank\">https://github.com/vlawhern/arl-eegmodels</a></p>\n<p>I tried implementing a similar model to that EEGNet, and did not get good performance.</p>",
              "rawMarkdown": "@sunyuri By \"EEGNet\" do you mean this EEGNet? https://github.com/vlawhern/arl-eegmodels\n\nI tried implementing a similar model to that EEGNet, and did not get good performance."
            },
            {
              "id": 2700168,
              "postDate": "2024-03-16T11:36:24.063Z",
              "content": "<p>Some parameters need to be adjusted, and the original versions of many models are not suitable.</p>",
              "rawMarkdown": "Some parameters need to be adjusted, and the original versions of many models are not suitable."
            },
            {
              "id": 2700289,
              "postDate": "2024-03-16T13:02:01.190Z",
              "content": "<p><a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> I see, but is that the EEGNet you were referring to in your original comment? (the one in that repo)?</p>",
              "rawMarkdown": "@sunyuri I see, but is that the EEGNet you were referring to in your original comment? (the one in that repo)?"
            },
            {
              "id": 2700461,
              "postDate": "2024-03-16T14:29:05.020Z",
              "content": "<p>The code was written by myself based on the structure of the network in the <a href=\"https://iopscience.iop.org/article/10.1088/1741-2552/aace8c\" target=\"_blank\">paper</a>. (Not that repo)</p>",
              "rawMarkdown": "The code was written by myself based on the structure of the network in the [paper](https://iopscience.iop.org/article/10.1088/1741-2552/aace8c). (Not that repo)"
            }
          ]
        }
      ]
    },
    {
      "id": 2681672,
      "postDate": "2024-03-04T19:32:09.127Z",
      "content": "<p>I can't acces the full publication. But if I undestood correctly for the abstract. The idea is to apply attention to the in our case 19 (or 20 with EKG) eeg channels rigth? So we need to apply positional embedding to each of those channels. Being a sequence 19 (or 20) \"words\" long.</p>\n<p>EDIT: Is working now. But I'm still not confident with making embeddings trainable. Or if they're actually trainable by this implementation. The think is that as far I know, they shouldn't be.</p>\n<p>EDIT2: May be that's the key. Initializes embeddings as a combination of sinusoidal and zeros. And then let them train. Not bad.</p>",
      "rawMarkdown": "I can't acces the full publication. But if I undestood correctly for the abstract. The idea is to apply attention to the in our case 19 (or 20 with EKG) eeg channels rigth? So we need to apply positional embedding to each of those channels. Being a sequence 19 (or 20) \"words\" long.\n\nEDIT: Is working now. But I'm still not confident with making embeddings trainable. Or if they're actually trainable by this implementation. The think is that as far I know, they shouldn't be.\n\nEDIT2: May be that's the key. Initializes embeddings as a combination of sinusoidal and zeros. And then let them train. Not bad.",
      "votes": 2,
      "replies": [
        {
          "id": 2681859,
          "postDate": "2024-03-04T23:27:20.370Z",
          "content": "<p>You can access full paper from here</p>\n<p><a href=\"https://www.researchgate.net/publication/362766298_Continuous_Seizure_Detection_Based_on_Transformer_and_Long-Term_iEEG\" target=\"_blank\">https://www.researchgate.net/publication/362766298_Continuous_Seizure_Detection_Based_on_Transformer_and_Long-Term_iEEG</a></p>",
          "rawMarkdown": "You can access full paper from here\n\nhttps://www.researchgate.net/publication/362766298_Continuous_Seizure_Detection_Based_on_Transformer_and_Long-Term_iEEG",
          "votes": 6
        }
      ]
    },
    {
      "id": 2681773,
      "postDate": "2024-03-04T20:50:50.550Z",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> why did you put the layer normalization before the attention and the MLP? Shouldn't it be after?</p>",
      "rawMarkdown": "@gunesevitan why did you put the layer normalization before the attention and the MLP? Shouldn't it be after?",
      "replies": [
        {
          "id": 2681804,
          "postDate": "2024-03-04T21:35:16.087Z",
          "content": "<p>The last convolution at convolution blocks hasn't been normalized since it applyes mean cross time dimension to it. Normalizes after mean and before attention between channels.</p>",
          "rawMarkdown": "The last convolution at convolution blocks hasn't been normalized since it applyes mean cross time dimension to it. Normalizes after mean and before attention between channels.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2681650,
      "postDate": "2024-03-04T18:55:56.943Z",
      "content": "<p><a href=\"https://www.kaggle.com/sacuscreed/channel-attention-hybrid-transformer/edit\" target=\"_blank\">https://www.kaggle.com/sacuscreed/channel-attention-hybrid-transformer/edit</a><br>\nI've changed Head for a Linear to 6 classes. </p>\n<p>EDIT: Not yet. I guess I need to read the paper but <br>\nself.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))<br>\nBy doing that those embeddings will not be trainable? They shouldn't.</p>",
      "rawMarkdown": "https://www.kaggle.com/sacuscreed/channel-attention-hybrid-transformer/edit\nI've changed Head for a Linear to 6 classes. \n\nEDIT: Not yet. I guess I need to read the paper but \nself.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))\nBy doing that those embeddings will not be trainable? They shouldn't.",
      "replies": [
        {
          "id": 2682004,
          "postDate": "2024-03-05T03:43:02.653Z",
          "content": "<p><a href=\"https://pytorch.org/docs/stable/generated/torch.nn.parameter.Parameter.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.parameter.Parameter.html</a></p>\n<p><code>nn.Parameter</code> has <code>requires_grad=True</code> by default so it's trainable.</p>",
          "rawMarkdown": "https://pytorch.org/docs/stable/generated/torch.nn.parameter.Parameter.html\n\n`nn.Parameter` has `requires_grad=True` by default so it's trainable.",
          "votes": 2,
          "replies": [
            {
              "id": 2682487,
              "postDate": "2024-03-05T10:59:49.150Z",
              "content": "<p>Thanks. So they will be stored and accessible but not necessarily affected by backpropagation. Will they?</p>",
              "rawMarkdown": "Thanks. So they will be stored and accessible but not necessarily affected by backpropagation. Will they?"
            },
            {
              "id": 2682729,
              "postDate": "2024-03-05T14:52:36.547Z",
              "content": "<p>They are affected by backprop and their weights are updated since they are added and concatenated here.</p>\n<pre><code>        \n        x += self.positional_embeddings\n        x = torch.cat([\n            self.cls_token.unsqueeze().repeat(x.size(), , ),\n            x\n        ], )\n</code></pre>",
              "rawMarkdown": "They are affected by backprop and their weights are updated since they are added and concatenated here.\n```python\n        # Add positional embeddings and concatenate cls token\n        x += self.positional_embeddings\n        x = torch.cat([\n            self.cls_token.unsqueeze(0).repeat(x.size(0), 1, 1),\n            x\n        ], 1)\n```",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2681632,
      "postDate": "2024-03-04T18:42:04.257Z",
      "content": "<p>Hi. Diving into it. What's heads?</p>",
      "rawMarkdown": "Hi. Diving into it. What's heads?",
      "replies": [
        {
          "id": 2681647,
          "postDate": "2024-03-04T18:53:26.057Z",
          "content": "<p>It's just one of my modules that has model heads. It's basically a fully connected layer that outputs 6 classes.</p>",
          "rawMarkdown": "It's just one of my modules that has model heads. It's basically a fully connected layer that outputs 6 classes.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2681972,
      "author_name": "Yuri Sun",
      "author_url": "",
      "post_date": "2024-03-05T03:19:07.347000",
      "content": "<p>I have tried using the attention mechanism in this competition, but the results were not as good as using pre-trained EfficientNet-B0. My initial idea was to treat the four brain regions of LL, RL, LP, and RP as four tokens, which would be a very short \"sentence\".</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2682293,
          "author_name": "Dive Deeper",
          "author_url": "",
          "post_date": "2024-03-05T08:17:04.803000",
          "content": "<p>I also conduct a similar experiment, but it's not better than concatenation (like Chris's works)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2682313,
              "author_name": "Yuri Sun",
              "author_url": "",
              "post_date": "2024-03-05T08:40:42.130000",
              "content": "<p>This result is confusing for me, perhaps finding a suitable concatenation method is the key. For the original EEG signal, the performance of EEGNet is better than that of the EfficientNet without pre-trained weights, but at the same time, it is lower than that of pre-trained EfficientNet (In my experiments).<br>\nI think we need a pre-trained EEG large model (like this <a href=\"https://openreview.net/pdf?id=QzTpTRVtrP\" target=\"_blank\">paper</a>)</p>",
              "votes": 10,
              "replies": []
            },
            {
              "id": 2682436,
              "author_name": "Reacher",
              "author_url": "",
              "post_date": "2024-03-05T10:25:28.163000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> , by  <code>pre-trained EfficientNet</code> you mean ImageNet weights?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2682542,
              "author_name": "Yuri Sun",
              "author_url": "",
              "post_date": "2024-03-05T11:41:18.823000",
              "content": "<p>Yes.             </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2698816,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-15T16:47:23.800000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2698821,
              "author_name": "Mandeep",
              "author_url": "",
              "post_date": "2024-03-15T16:49:33.203000",
              "content": "<p><a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> By \"EEGNet\" do you mean this EEGNet? <a href=\"https://github.com/vlawhern/arl-eegmodels\" target=\"_blank\">https://github.com/vlawhern/arl-eegmodels</a></p>\n<p>I tried implementing a similar model to that EEGNet, and did not get good performance.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2700168,
              "author_name": "Yuri Sun",
              "author_url": "",
              "post_date": "2024-03-16T11:36:24.063000",
              "content": "<p>Some parameters need to be adjusted, and the original versions of many models are not suitable.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2700289,
              "author_name": "Mandeep",
              "author_url": "",
              "post_date": "2024-03-16T13:02:01.190000",
              "content": "<p><a href=\"https://www.kaggle.com/sunyuri\" target=\"_blank\">@sunyuri</a> I see, but is that the EEGNet you were referring to in your original comment? (the one in that repo)?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2700461,
              "author_name": "Yuri Sun",
              "author_url": "",
              "post_date": "2024-03-16T14:29:05.020000",
              "content": "<p>The code was written by myself based on the structure of the network in the <a href=\"https://iopscience.iop.org/article/10.1088/1741-2552/aace8c\" target=\"_blank\">paper</a>. (Not that repo)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2681672,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-03-04T19:32:09.127000",
      "content": "<p>I can't acces the full publication. But if I undestood correctly for the abstract. The idea is to apply attention to the in our case 19 (or 20 with EKG) eeg channels rigth? So we need to apply positional embedding to each of those channels. Being a sequence 19 (or 20) \"words\" long.</p>\n<p>EDIT: Is working now. But I'm still not confident with making embeddings trainable. Or if they're actually trainable by this implementation. The think is that as far I know, they shouldn't be.</p>\n<p>EDIT2: May be that's the key. Initializes embeddings as a combination of sinusoidal and zeros. And then let them train. Not bad.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2681859,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-03-04T23:27:20.370000",
          "content": "<p>You can access full paper from here</p>\n<p><a href=\"https://www.researchgate.net/publication/362766298_Continuous_Seizure_Detection_Based_on_Transformer_and_Long-Term_iEEG\" target=\"_blank\">https://www.researchgate.net/publication/362766298_Continuous_Seizure_Detection_Based_on_Transformer_and_Long-Term_iEEG</a></p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 2681773,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2024-03-04T20:50:50.550000",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> why did you put the layer normalization before the attention and the MLP? Shouldn't it be after?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2681804,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2024-03-04T21:35:16.087000",
          "content": "<p>The last convolution at convolution blocks hasn't been normalized since it applyes mean cross time dimension to it. Normalizes after mean and before attention between channels.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2681650,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-03-04T18:55:56.943000",
      "content": "<p><a href=\"https://www.kaggle.com/sacuscreed/channel-attention-hybrid-transformer/edit\" target=\"_blank\">https://www.kaggle.com/sacuscreed/channel-attention-hybrid-transformer/edit</a><br>\nI've changed Head for a Linear to 6 classes. </p>\n<p>EDIT: Not yet. I guess I need to read the paper but <br>\nself.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))<br>\nBy doing that those embeddings will not be trainable? They shouldn't.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2682004,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-03-05T03:43:02.653000",
          "content": "<p><a href=\"https://pytorch.org/docs/stable/generated/torch.nn.parameter.Parameter.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.parameter.Parameter.html</a></p>\n<p><code>nn.Parameter</code> has <code>requires_grad=True</code> by default so it's trainable.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2682487,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-03-05T10:59:49.150000",
              "content": "<p>Thanks. So they will be stored and accessible but not necessarily affected by backpropagation. Will they?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2682729,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-05T14:52:36.547000",
              "content": "<p>They are affected by backprop and their weights are updated since they are added and concatenated here.</p>\n<pre><code>        \n        x += self.positional_embeddings\n        x = torch.cat([\n            self.cls_token.unsqueeze().repeat(x.size(), , ),\n            x\n        ], )\n</code></pre>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2681632,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-03-04T18:42:04.257000",
      "content": "<p>Hi. Diving into it. What's heads?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2681647,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-03-04T18:53:26.057000",
          "content": "<p>It's just one of my modules that has model heads. It's basically a fully connected layer that outputs 6 classes.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2681525": "I found @sunyuri's paper [here](https://ieeexplore.ieee.org/document/9858598) implemented in on pytorch but couldn't make it work. I'm not good at transformers so I'm sharing it and maybe someone else can make it work.\n\n```\nimport numpy as np\nimport torch\nimport torch.nn as nn\n\nfrom heads import ClassificationHead\n\n\ndef positional_encoding(length, embed_dim):\n\n    dim = embed_dim // 2\n\n    position = np.arange(length)[:, np.newaxis]\n    dim = np.arange(dim)[np.newaxis, :] / dim\n\n    angle = 1 / (10000 ** dim)\n    angle = position * angle\n\n    pos_embed = np.concatenate([np.sin(angle), np.cos(angle)], axis=-1)\n    pos_embed = torch.from_numpy(pos_embed).float()\n\n    return pos_embed\n\n\nclass ConvBlock(nn.Module):\n\n    def __init__(self, in_channels, out_channels):\n\n        super(ConvBlock, self).__init__()\n\n        self.conv1 = nn.Conv2d(in_channels=in_channels, out_channels=32, kernel_size=(4, 1), stride=(2, 1), padding=(1, 0), padding_mode='zeros', bias=True)\n        self.bn1 = nn.BatchNorm2d(num_features=32)\n        self.conv2 = nn.Conv2d(in_channels=32, out_channels=64, kernel_size=(4, 1), stride=(2, 1), padding=(1, 0), padding_mode='zeros', bias=True)\n        self.bn2 = nn.BatchNorm2d(num_features=64)\n        self.conv3 = nn.Conv2d(in_channels=64, out_channels=out_channels, kernel_size=(4, 1), stride=(2, 1), padding=(1, 0), padding_mode='zeros', bias=True)\n        self.activation = nn.LeakyReLU(inplace=False)\n\n    def forward(self, x):\n\n        x = self.conv1(x)\n        x = self.bn1(x)\n        x = self.activation(x)\n\n        x = self.conv2(x)\n        x = self.bn2(x)\n        x = self.activation(x)\n\n        x = self.conv3(x)\n\n        return x\n\n\nclass MLP(nn.Module):\n\n    def __init__(self, embed_dim, hidden_dim):\n\n        super(MLP, self).__init__()\n\n        self.mlp = nn.Sequential(\n            nn.Linear(in_features=embed_dim, out_features=hidden_dim),\n            nn.ReLU(inplace=True),\n            nn.Linear(in_features=hidden_dim, out_features=embed_dim),\n        )\n\n    def forward(self, x):\n        return self.mlp(x)\n\n\nclass TransformerBlock(nn.Module):\n\n    def __init__(self, embed_dim, num_heads, out_dim):\n        super(TransformerBlock, self).__init__()\n\n        self.attention = nn.MultiheadAttention(embed_dim=embed_dim, num_heads=num_heads, batch_first=True)\n        self.mlp = MLP(embed_dim=embed_dim, hidden_dim=out_dim)\n        self.ln1 = nn.LayerNorm(embed_dim)\n        self.ln2 = nn.LayerNorm(out_dim)\n\n    def forward(self, x):\n\n        x = self.ln1(x)\n        x = self.attention(x, x, x)[0]\n        x = self.ln2(x)\n        x = self.mlp(x)\n\n        return x\n\n\nclass HybridTransformer(nn.Module):\n\n    def __init__(self, in_channels, hidden_size, num_heads, num_blocks, pooling_type, dropout_rate, head_args):\n\n        super(HybridTransformer, self).__init__()\n\n        self.conv_block = ConvBlock(in_channels=1, out_channels=hidden_size)\n        self.encoder = nn.ModuleList([\n            TransformerBlock(\n                embed_dim=hidden_size,\n                num_heads=num_heads,\n                out_dim=hidden_size,\n            ) for _ in range(num_blocks)\n        ])\n\n        self.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))\n        self.cls_token = nn.Parameter(torch.zeros((1, hidden_size)))\n\n        self.pooling_type = pooling_type\n        self.dropout = nn.Dropout(dropout_rate) if dropout_rate > 0 else nn.Identity()\n        self.head = ClassificationHead(input_dimensions=hidden_size, **head_args)\n\n    def forward(self, x):\n\n        # Add channel dimension and pass it to conv 2d block\n        x = x.unsqueeze(dim=1).permute(0, 1, 3, 2)\n        x = self.conv_block(x)\n\n        # Average features along time dimension\n        x = torch.mean(x, dim=2).permute(0, 2, 1)\n\n        # Add positional embeddings and concatenate cls token\n        x += self.positional_embeddings\n        x = torch.cat([\n            self.cls_token.unsqueeze(0).repeat(x.size(0), 1, 1),\n            x\n        ], 1)\n\n        # Pass it to transformer encoder\n        for block in self.encoder:\n            x = block(x)\n\n        if self.pooling_type == 'avg':\n            x = torch.mean(x, dim=1)\n        elif self.pooling_type == 'max':\n            x = torch.max(x, dim=1)[0]\n        elif self.pooling_type == 'cls':\n            x = x[:, 0, :]\n\n        x = self.dropout(x)\n        output = self.head(x)\n\n        return output\n```",
    "2681972": "I have tried using the attention mechanism in this competition, but the results were not as good as using pre-trained EfficientNet-B0. My initial idea was to treat the four brain regions of LL, RL, LP, and RP as four tokens, which would be a very short \"sentence\".",
    "2681672": "I can't acces the full publication. But if I undestood correctly for the abstract. The idea is to apply attention to the in our case 19 (or 20 with EKG) eeg channels rigth? So we need to apply positional embedding to each of those channels. Being a sequence 19 (or 20) \"words\" long.\n\nEDIT: Is working now. But I'm still not confident with making embeddings trainable. Or if they're actually trainable by this implementation. The think is that as far I know, they shouldn't be.\n\nEDIT2: May be that's the key. Initializes embeddings as a combination of sinusoidal and zeros. And then let them train. Not bad.",
    "2681773": "@gunesevitan why did you put the layer normalization before the attention and the MLP? Shouldn't it be after?",
    "2681650": "https://www.kaggle.com/sacuscreed/channel-attention-hybrid-transformer/edit\nI've changed Head for a Linear to 6 classes. \n\nEDIT: Not yet. I guess I need to read the paper but \nself.positional_embeddings = torch.nn.Parameter(positional_encoding(in_channels, hidden_size))\nBy doing that those embeddings will not be trainable? They shouldn't.",
    "2681632": "Hi. Diving into it. What's heads?"
  }
}