{
  "id": 476000,
  "title": "Getting into MultiModal Approach Raw Signals + Spectrograms CV: 0.607 LB: 0.41 ",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/476000",
  "author_name": "",
  "post_date": "2024-02-10T18:34:55.070428300Z",
  "votes": 87,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi again, sharing another baseline for everyone to play around with. I am able to get good results within couple of experiments. </p>\n<p>In continuation to my previous <strong>1D ResNet based [approach]</strong>(<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666)\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666)</a>, this pipeline basically introduces 2D encoder (<strong>effcientnet-b0</strong>). </p>\n<p>The idea behind the approach is to utilise both Raw Signals as well as Kaggle/EEG signals based Spectrograms within a single architecture &amp; train the entire network together. Also, I do not claim this solution to be superior than methods like separately finetuning 1D &amp; 2D models and then ensemble, but this definitely does have potential to reach much higher scores.</p>\n<p><strong>Training notebook:</strong> <a href=\"https://www.kaggle.com/nischaydnk/training-multimodal-1d-2d-approach-eegs\" target=\"_blank\">link</a> <br>\n<strong>Submission notebook:</strong> <a href=\"https://www.kaggle.com/nischaydnk/multimodal-1d-2d-eeg-approach-submission\" target=\"_blank\">link</a> </p>\n<h2>Architecture:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fc4e39641fa8ec13588b62065ae1e9b58%2FScreenshot%202024-02-10%20at%2010.40.36%20PM.png?generation=1707585942607444&amp;alt=media\"></p>\n<p>Basically, now we send two inputs within the architecture: Raw signals w/ Feature engineering from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> + <em>concatenated Kaggle+EEG spectrograms</em> , torch version shared <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> </p>\n<p>Each input is separately processed by the 1D &amp; 2D CNN based Encoders, followed by fully connected layer to make equi-dimensional features. These features are then normalized &amp; used to calculate contrastive loss. Since we do not have any pariwise labels, we keep all the labels as 1 while calculating loss, this will basically encourage the hidden respresentahions from 1D and 2D models to more closer to one another. </p>\n<p>`        </p>\n<pre><code>     = torch.nn.functional.normalize(embedding_1d, p=, dim=)\n     = torch.nn.functional.normalize(embedding_2d, p=, dim=)\n\n\n     = torch.es(embedding_1d.size()).to(self.device)  \n     = self.contrastive_loss(embedding_1d, embedding_2d, contrastive_target)\n\n     = classification_loss + classification_loss1* + classification_loss2* + contrastive_loss*  \n</code></pre>\n<p>Apart from this, we are also using KL Divergence Loss on multiple outputs generated from linear heads: concatenated features from 1D + 2D encoders, 1D CNN features, 2D CNN features. I just tried this to get higher weightage to task based loss function than contrastive loss &amp; this worked better than just having single linear head for making predictions. <br>\nThis is the weightage I am giving to calculate the final predictions.</p>\n<p>`         </p>\n<pre><code>         = outputs*. + y1*. + y2*.\n\n         = nn.Softmax(dim=)(outputs)`\n</code></pre>\n<h4>**Few more additions in training strategy: **</h4>\n<ul>\n<li>Horizontal Flip Augmentation for Spectrograms</li>\n<li>8 channel Feature Engineering shared by Chris</li>\n<li>Auxiliary Loss<br>\n    <em>total_loss = classification_loss + classification_loss1</em>0.5 + classification_loss2<em>0.5 + contrastive_loss</em>0.5*</li>\n<li>Multiple outputs</li>\n</ul>\n<p>Thanks. Feel free to do more experiments to improve the results or share new ideas related to multimodal approaches in the thread.</p>",
  "messages": [
    {
      "id": "2646248",
      "postDate": "02/10/2024 18:34:55",
      "content": "<p>Hi again, sharing another baseline for everyone to play around with. I am able to get good results within couple of experiments. </p>\n<p>In continuation to my previous <strong>1D ResNet based [approach]</strong>(<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666)\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666)</a>, this pipeline basically introduces 2D encoder (<strong>effcientnet-b0</strong>). </p>\n<p>The idea behind the approach is to utilise both Raw Signals as well as Kaggle/EEG signals based Spectrograms within a single architecture &amp; train the entire network together. Also, I do not claim this solution to be superior than methods like separately finetuning 1D &amp; 2D models and then ensemble, but this definitely does have potential to reach much higher scores.</p>\n<p><strong>Training notebook:</strong> <a href=\"https://www.kaggle.com/nischaydnk/training-multimodal-1d-2d-approach-eegs\" target=\"_blank\">link</a> <br>\n<strong>Submission notebook:</strong> <a href=\"https://www.kaggle.com/nischaydnk/multimodal-1d-2d-eeg-approach-submission\" target=\"_blank\">link</a> </p>\n<h2>Architecture:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fc4e39641fa8ec13588b62065ae1e9b58%2FScreenshot%202024-02-10%20at%2010.40.36%20PM.png?generation=1707585942607444&amp;alt=media\"></p>\n<p>Basically, now we send two inputs within the architecture: Raw signals w/ Feature engineering from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> + <em>concatenated Kaggle+EEG spectrograms</em> , torch version shared <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> </p>\n<p>Each input is separately processed by the 1D &amp; 2D CNN based Encoders, followed by fully connected layer to make equi-dimensional features. These features are then normalized &amp; used to calculate contrastive loss. Since we do not have any pariwise labels, we keep all the labels as 1 while calculating loss, this will basically encourage the hidden respresentahions from 1D and 2D models to more closer to one another. </p>\n<p>`        </p>\n<pre><code>     = torch.nn.functional.normalize(embedding_1d, p=, dim=)\n     = torch.nn.functional.normalize(embedding_2d, p=, dim=)\n\n\n     = torch.es(embedding_1d.size()).to(self.device)  \n     = self.contrastive_loss(embedding_1d, embedding_2d, contrastive_target)\n\n     = classification_loss + classification_loss1* + classification_loss2* + contrastive_loss*  \n</code></pre>\n<p>Apart from this, we are also using KL Divergence Loss on multiple outputs generated from linear heads: concatenated features from 1D + 2D encoders, 1D CNN features, 2D CNN features. I just tried this to get higher weightage to task based loss function than contrastive loss &amp; this worked better than just having single linear head for making predictions. <br>\nThis is the weightage I am giving to calculate the final predictions.</p>\n<p>`         </p>\n<pre><code>         = outputs*. + y1*. + y2*.\n\n         = nn.Softmax(dim=)(outputs)`\n</code></pre>\n<h4>**Few more additions in training strategy: **</h4>\n<ul>\n<li>Horizontal Flip Augmentation for Spectrograms</li>\n<li>8 channel Feature Engineering shared by Chris</li>\n<li>Auxiliary Loss<br>\n    <em>total_loss = classification_loss + classification_loss1</em>0.5 + classification_loss2<em>0.5 + contrastive_loss</em>0.5*</li>\n<li>Multiple outputs</li>\n</ul>\n<p>Thanks. Feel free to do more experiments to improve the results or share new ideas related to multimodal approaches in the thread.</p>",
      "rawMarkdown": "Hi again, sharing another baseline for everyone to play around with. I am able to get good results within couple of experiments. \n\nIn continuation to my previous **1D ResNet based [approach]**(https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666), this pipeline basically introduces 2D encoder (**effcientnet-b0**). \n \nThe idea behind the approach is to utilise both Raw Signals as well as Kaggle/EEG signals based Spectrograms within a single architecture & train the entire network together. Also, I do not claim this solution to be superior than methods like separately finetuning 1D & 2D models and then ensemble, but this definitely does have potential to reach much higher scores.\n\n**Training notebook:** [link](https://www.kaggle.com/nischaydnk/training-multimodal-1d-2d-approach-eegs) \n**Submission notebook:** [link](https://www.kaggle.com/nischaydnk/multimodal-1d-2d-eeg-approach-submission) \n\n\n## Architecture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fc4e39641fa8ec13588b62065ae1e9b58%2FScreenshot%202024-02-10%20at%2010.40.36%20PM.png?generation=1707585942607444&alt=media)\n\nBasically, now we send two inputs within the architecture: Raw signals w/ Feature engineering from @cdeotte + *concatenated Kaggle+EEG spectrograms* , torch version shared [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092) by @alejopaullier \n\nEach input is separately processed by the 1D & 2D CNN based Encoders, followed by fully connected layer to make equi-dimensional features. These features are then normalized & used to calculate contrastive loss. Since we do not have any pariwise labels, we keep all the labels as 1 while calculating loss, this will basically encourage the hidden respresentahions from 1D and 2D models to more closer to one another. \n\n`        \n        \n        embedding_1d = torch.nn.functional.normalize(embedding_1d, p=2, dim=1)\n        embedding_2d = torch.nn.functional.normalize(embedding_2d, p=2, dim=1)\n\n        \n        contrastive_target = torch.ones(embedding_1d.size(0)).to(self.device)  # Assuming all pairs are similar\n        contrastive_loss = self.contrastive_loss(embedding_1d, embedding_2d, contrastive_target)\n\n        total_loss = classification_loss + classification_loss1*0.5 + classification_loss2*0.5 + contrastive_loss*0.5  # Aux losses`\n\nApart from this, we are also using KL Divergence Loss on multiple outputs generated from linear heads: concatenated features from 1D + 2D encoders, 1D CNN features, 2D CNN features. I just tried this to get higher weightage to task based loss function than contrastive loss & this worked better than just having single linear head for making predictions. \nThis is the weightage I am giving to calculate the final predictions.\n\n`         \n\n            outputs = outputs*0.5 + y1*0.25 + y2*0.25\n            \n            outputs = nn.Softmax(dim=1)(outputs)`\n\n\n#### **Few more additions in training strategy: **\n- Horizontal Flip Augmentation for Spectrograms\n- 8 channel Feature Engineering shared by Chris\n- Auxiliary Loss\n        *total_loss = classification_loss + classification_loss1*0.5 + classification_loss2*0.5 + contrastive_loss*0.5*\n- Multiple outputs\n\nThanks. Feel free to do more experiments to improve the results or share new ideas related to multimodal approaches in the thread.",
      "votes": null
    },
    {
      "id": "2648075",
      "postDate": "02/12/2024 04:36:13",
      "content": "<p>Maybe this is a dumb question but I wonder if instead of trying to minimise contrastive loss here, one could get interesting results by trying to maximise contrastive loss, i.e., maybe it would help increase diversity of the information processed.</p>",
      "rawMarkdown": "Maybe this is a dumb question but I wonder if instead of trying to minimise contrastive loss here, one could get interesting results by trying to maximise contrastive loss, i.e., maybe it would help increase diversity of the information processed.",
      "votes": null
    },
    {
      "id": "2649260",
      "postDate": "02/12/2024 17:41:10",
      "content": "<p>Very nice!</p>",
      "rawMarkdown": "Very nice!",
      "votes": null
    },
    {
      "id": "2649527",
      "postDate": "02/12/2024 22:52:41",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> very nice! How do you choose the weights for the auxiliary loss?</p>",
      "rawMarkdown": "Hey @nischaydnk very nice! How do you choose the weights for the auxiliary loss?",
      "votes": null
    },
    {
      "id": "2650386",
      "postDate": "02/13/2024 12:35:26",
      "content": "<p>It could be an option, if we penalize the loss function and use the contrastive loss as a regularization term to increase diversity, that might help with overfitting.</p>",
      "rawMarkdown": "It could be an option, if we penalize the loss function and use the contrastive loss as a regularization term to increase diversity, that might help with overfitting.",
      "votes": null
    },
    {
      "id": "2651200",
      "postDate": "02/14/2024 02:59:56",
      "content": "<p>Was anyone able to run the training notebook on Kaggle? I'm currently facing two issues:</p>\n<ul>\n<li>When I update local paths to what I believe is the correct Kaggle path, they just don't work.</li>\n<li>If the brain-eeg-spectrograms dataset is loaded, the notebook runs out of memory right at the beginning. Since it seems that this dataset is necessary for the training, I find myself unable to continue.</li>\n</ul>",
      "rawMarkdown": "Was anyone able to run the training notebook on Kaggle? I'm currently facing two issues:\n- When I update local paths to what I believe is the correct Kaggle path, they just don't work.\n- If the brain-eeg-spectrograms dataset is loaded, the notebook runs out of memory right at the beginning. Since it seems that this dataset is necessary for the training, I find myself unable to continue.",
      "votes": null
    },
    {
      "id": "2652001",
      "postDate": "02/14/2024 13:59:05",
      "content": "<p>Nice work, I even have similar ideas in my scientific work!</p>",
      "rawMarkdown": "Nice work, I even have similar ideas in my scientific work!",
      "votes": null
    },
    {
      "id": "2662543",
      "postDate": "02/22/2024 01:04:52",
      "content": "<p>hi Have you tried this method and has the effect increased</p>",
      "rawMarkdown": "hi Have you tried this method and has the effect increased",
      "votes": null
    },
    {
      "id": "2667259",
      "postDate": "02/25/2024 02:51:32",
      "content": "<p>Hehe not yet. Been working on other stuff. Would like to see if anyone else has tried though.</p>",
      "rawMarkdown": "Hehe not yet. Been working on other stuff. Would like to see if anyone else has tried though.",
      "votes": null
    },
    {
      "id": "3058956",
      "postDate": "11/30/2024 07:05:34",
      "content": "<p>Very cool, thanks!</p>",
      "rawMarkdown": "Very cool, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2648075,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "02/12/2024 04:36:13",
      "content": "<p>Maybe this is a dumb question but I wonder if instead of trying to minimise contrastive loss here, one could get interesting results by trying to maximise contrastive loss, i.e., maybe it would help increase diversity of the information processed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2650386,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "02/13/2024 12:35:26",
          "content": "<p>It could be an option, if we penalize the loss function and use the contrastive loss as a regularization term to increase diversity, that might help with overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2662543,
          "author_name": "chenboluo",
          "author_url": "",
          "post_date": "02/22/2024 01:04:52",
          "content": "<p>hi Have you tried this method and has the effect increased</p>",
          "votes": null,
          "replies": [
            {
              "id": 2667259,
              "author_name": "caelhasse",
              "author_url": "",
              "post_date": "02/25/2024 02:51:32",
              "content": "<p>Hehe not yet. Been working on other stuff. Would like to see if anyone else has tried though.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2649260,
      "author_name": "",
      "author_url": "",
      "post_date": "02/12/2024 17:41:10",
      "content": "<p>Very nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2649527,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "02/12/2024 22:52:41",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> very nice! How do you choose the weights for the auxiliary loss?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2651200,
      "author_name": "yantxx",
      "author_url": "",
      "post_date": "02/14/2024 02:59:56",
      "content": "<p>Was anyone able to run the training notebook on Kaggle? I'm currently facing two issues:</p>\n<ul>\n<li>When I update local paths to what I believe is the correct Kaggle path, they just don't work.</li>\n<li>If the brain-eeg-spectrograms dataset is loaded, the notebook runs out of memory right at the beginning. Since it seems that this dataset is necessary for the training, I find myself unable to continue.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2652001,
      "author_name": "chenboluo",
      "author_url": "",
      "post_date": "02/14/2024 13:59:05",
      "content": "<p>Nice work, I even have similar ideas in my scientific work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3058956,
      "author_name": "guneevarora",
      "author_url": "",
      "post_date": "11/30/2024 07:05:34",
      "content": "<p>Very cool, thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2646248": "Hi again, sharing another baseline for everyone to play around with. I am able to get good results within couple of experiments. \n\nIn continuation to my previous **1D ResNet based [approach]**(https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666), this pipeline basically introduces 2D encoder (**effcientnet-b0**). \n \nThe idea behind the approach is to utilise both Raw Signals as well as Kaggle/EEG signals based Spectrograms within a single architecture & train the entire network together. Also, I do not claim this solution to be superior than methods like separately finetuning 1D & 2D models and then ensemble, but this definitely does have potential to reach much higher scores.\n\n**Training notebook:** [link](https://www.kaggle.com/nischaydnk/training-multimodal-1d-2d-approach-eegs) \n**Submission notebook:** [link](https://www.kaggle.com/nischaydnk/multimodal-1d-2d-eeg-approach-submission) \n\n\n## Architecture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fc4e39641fa8ec13588b62065ae1e9b58%2FScreenshot%202024-02-10%20at%2010.40.36%20PM.png?generation=1707585942607444&alt=media)\n\nBasically, now we send two inputs within the architecture: Raw signals w/ Feature engineering from @cdeotte + *concatenated Kaggle+EEG spectrograms* , torch version shared [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/472092) by @alejopaullier \n\nEach input is separately processed by the 1D & 2D CNN based Encoders, followed by fully connected layer to make equi-dimensional features. These features are then normalized & used to calculate contrastive loss. Since we do not have any pariwise labels, we keep all the labels as 1 while calculating loss, this will basically encourage the hidden respresentahions from 1D and 2D models to more closer to one another. \n\n`        \n        \n        embedding_1d = torch.nn.functional.normalize(embedding_1d, p=2, dim=1)\n        embedding_2d = torch.nn.functional.normalize(embedding_2d, p=2, dim=1)\n\n        \n        contrastive_target = torch.ones(embedding_1d.size(0)).to(self.device)  # Assuming all pairs are similar\n        contrastive_loss = self.contrastive_loss(embedding_1d, embedding_2d, contrastive_target)\n\n        total_loss = classification_loss + classification_loss1*0.5 + classification_loss2*0.5 + contrastive_loss*0.5  # Aux losses`\n\nApart from this, we are also using KL Divergence Loss on multiple outputs generated from linear heads: concatenated features from 1D + 2D encoders, 1D CNN features, 2D CNN features. I just tried this to get higher weightage to task based loss function than contrastive loss & this worked better than just having single linear head for making predictions. \nThis is the weightage I am giving to calculate the final predictions.\n\n`         \n\n            outputs = outputs*0.5 + y1*0.25 + y2*0.25\n            \n            outputs = nn.Softmax(dim=1)(outputs)`\n\n\n#### **Few more additions in training strategy: **\n- Horizontal Flip Augmentation for Spectrograms\n- 8 channel Feature Engineering shared by Chris\n- Auxiliary Loss\n        *total_loss = classification_loss + classification_loss1*0.5 + classification_loss2*0.5 + contrastive_loss*0.5*\n- Multiple outputs\n\nThanks. Feel free to do more experiments to improve the results or share new ideas related to multimodal approaches in the thread.",
    "2648075": "Maybe this is a dumb question but I wonder if instead of trying to minimise contrastive loss here, one could get interesting results by trying to maximise contrastive loss, i.e., maybe it would help increase diversity of the information processed.",
    "2649260": "Very nice!",
    "2649527": "Hey @nischaydnk very nice! How do you choose the weights for the auxiliary loss?",
    "2650386": "It could be an option, if we penalize the loss function and use the contrastive loss as a regularization term to increase diversity, that might help with overfitting.",
    "2651200": "Was anyone able to run the training notebook on Kaggle? I'm currently facing two issues:\n- When I update local paths to what I believe is the correct Kaggle path, they just don't work.\n- If the brain-eeg-spectrograms dataset is loaded, the notebook runs out of memory right at the beginning. Since it seems that this dataset is necessary for the training, I find myself unable to continue.",
    "2652001": "Nice work, I even have similar ideas in my scientific work!",
    "2662543": "hi Have you tried this method and has the effect increased",
    "2667259": "Hehe not yet. Been working on other stuff. Would like to see if anyone else has tried though.",
    "3058956": "Very cool, thanks!"
  },
  "source": "meta"
}