{
  "id": 471666,
  "title": "1D ResNet Based Architecture Baseline - Achieve 0.4x LB with Raw Signals in 2 minutes",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/471666",
  "author_name": "Nischay Dhankhar",
  "post_date": "2024-01-29T05:36:59.406000",
  "votes": 177,
  "comment_count": 26,
  "views": 0,
  "content": "<p><strong>Just joined the competition and I am already loving it. There are so many ways to tackle such signals dataset problems, and it always hype me up.</strong></p>\n<p>I am releasing a 1D CNN based baseline which scores <strong>0.72CV &amp; 0.48 LB</strong> without using any Spectrograms / Augmentations or 2D CNN Networks &amp; takes only 2 minutes to run submission. The idea is similar to what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has already shared in his Wavenet Notebook, but I worked upon a lot on replicating scores in Pytorch &amp; improving the architecture. </p>\n<h4>Inference Notebook: <a href=\"https://www.kaggle.com/code/nischaydnk/hms-submission-1d-eegnet-pipeline-lightning/notebook?scriptVersionId=160814854\" target=\"_blank\">link</a></h4>\n<h4>Training Notebook: <a href=\"https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs?scriptVersionId=160814948\" target=\"_blank\">link</a></h4>\n<p>The model is slightly based on our 4th Place solution from G2Net competition hosted 2 years back, which includes using Parallel 1D convolutions as feature extractors and then using 1D ResNet based blocks. You can read more about the solution in the <a href=\"https://www.researchgate.net/publication/359051366_GWNET_Detecting_Gravitational_Waves_using_Hierarchical_and_Residual_Learning_based_1D_CNNs\" target=\"_blank\">paper</a> we released. </p>\n<h4>**Current training settings &amp; Methods Used: **</h4>\n<ul>\n<li>Basic Preprocessing: <strong>Raw Signals</strong> --&gt; <strong>LowPass Filter</strong>--&gt;<strong>mu_law_encoding</strong></li>\n<li>Cosine Annealing Scheduler</li>\n<li>Validate every 0.5 epochs</li>\n<li>5 Fold GroupKFold</li>\n<li>KLDivLoss</li>\n</ul>\n<p>I haven't experimented much with the baseline, but I believe even 1D solutions are capable of scoring as high as 2D Spectrograms &amp; Deep CNN based approaches. Maybe with right configuration &amp; adding augmentations, you can improve scores further. </p>\n<h2>Architecture:</h2>\n<h4>It is mainly based on 3 main ideas: <strong>Parallel Convolution Blocks</strong> + <strong>ResNet Like 1D Blocks</strong> + <strong>RNN head</strong></h4>\n<p>Whenever working with any signals data, I mainly use this kind of generic architecture while developing 1D networks. In most cases, this works well for me. I start off with basic 1D CNN based networks and slowly upgrading it with introducing different ideas like adding residuals, etc. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fabbddaa44e5c4220f6bed3d3db688c17%2FScreenshot%202024-01-29%20at%209.50.30%20AM.png?generation=1706505660519033&amp;alt=media\"></p>\n<p><strong>Parallel Convolution Blocks</strong></p>\n<p>Specific to this competition data, I found smaller kernels (3,5,7) to work better in comparision to larger Kernels which we used in G2Net competition. Again, its more upto the experimentation, you can play around with different kernel size and decide to go with which one.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9317ffc90ca7cf9cd990cc87250b66ab%2FScreenshot%202024-01-29%20at%209.49.18%20AM.png?generation=1706503620778537&amp;alt=media\"></p>\n<p><strong>ResNet Like 1D Blocks</strong></p>\n<p>Followed by Parallel 1D CNNs, these features are then passed into Residual networks which basically helps in better feature extraction &amp; downsampling. Again, you can play around with different number of blocks you want to use. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F3c07868f0ac8069fee743203c9903107%2FScreenshot%202024-01-29%20at%209.49.13%20AM.png?generation=1706503747689647&amp;alt=media\"></p>\n<p><strong>I am planning to release more code with improved baseline and on augmentations.</strong></p>",
  "messages": [
    {
      "id": 2624962,
      "postDate": "2024-01-29T05:36:59.407Z",
      "content": "<p><strong>Just joined the competition and I am already loving it. There are so many ways to tackle such signals dataset problems, and it always hype me up.</strong></p>\n<p>I am releasing a 1D CNN based baseline which scores <strong>0.72CV &amp; 0.48 LB</strong> without using any Spectrograms / Augmentations or 2D CNN Networks &amp; takes only 2 minutes to run submission. The idea is similar to what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has already shared in his Wavenet Notebook, but I worked upon a lot on replicating scores in Pytorch &amp; improving the architecture. </p>\n<h4>Inference Notebook: <a href=\"https://www.kaggle.com/code/nischaydnk/hms-submission-1d-eegnet-pipeline-lightning/notebook?scriptVersionId=160814854\" target=\"_blank\">link</a></h4>\n<h4>Training Notebook: <a href=\"https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs?scriptVersionId=160814948\" target=\"_blank\">link</a></h4>\n<p>The model is slightly based on our 4th Place solution from G2Net competition hosted 2 years back, which includes using Parallel 1D convolutions as feature extractors and then using 1D ResNet based blocks. You can read more about the solution in the <a href=\"https://www.researchgate.net/publication/359051366_GWNET_Detecting_Gravitational_Waves_using_Hierarchical_and_Residual_Learning_based_1D_CNNs\" target=\"_blank\">paper</a> we released. </p>\n<h4>**Current training settings &amp; Methods Used: **</h4>\n<ul>\n<li>Basic Preprocessing: <strong>Raw Signals</strong> --&gt; <strong>LowPass Filter</strong>--&gt;<strong>mu_law_encoding</strong></li>\n<li>Cosine Annealing Scheduler</li>\n<li>Validate every 0.5 epochs</li>\n<li>5 Fold GroupKFold</li>\n<li>KLDivLoss</li>\n</ul>\n<p>I haven't experimented much with the baseline, but I believe even 1D solutions are capable of scoring as high as 2D Spectrograms &amp; Deep CNN based approaches. Maybe with right configuration &amp; adding augmentations, you can improve scores further. </p>\n<h2>Architecture:</h2>\n<h4>It is mainly based on 3 main ideas: <strong>Parallel Convolution Blocks</strong> + <strong>ResNet Like 1D Blocks</strong> + <strong>RNN head</strong></h4>\n<p>Whenever working with any signals data, I mainly use this kind of generic architecture while developing 1D networks. In most cases, this works well for me. I start off with basic 1D CNN based networks and slowly upgrading it with introducing different ideas like adding residuals, etc. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fabbddaa44e5c4220f6bed3d3db688c17%2FScreenshot%202024-01-29%20at%209.50.30%20AM.png?generation=1706505660519033&amp;alt=media\"></p>\n<p><strong>Parallel Convolution Blocks</strong></p>\n<p>Specific to this competition data, I found smaller kernels (3,5,7) to work better in comparision to larger Kernels which we used in G2Net competition. Again, its more upto the experimentation, you can play around with different kernel size and decide to go with which one.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9317ffc90ca7cf9cd990cc87250b66ab%2FScreenshot%202024-01-29%20at%209.49.18%20AM.png?generation=1706503620778537&amp;alt=media\"></p>\n<p><strong>ResNet Like 1D Blocks</strong></p>\n<p>Followed by Parallel 1D CNNs, these features are then passed into Residual networks which basically helps in better feature extraction &amp; downsampling. Again, you can play around with different number of blocks you want to use. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F3c07868f0ac8069fee743203c9903107%2FScreenshot%202024-01-29%20at%209.49.13%20AM.png?generation=1706503747689647&amp;alt=media\"></p>\n<p><strong>I am planning to release more code with improved baseline and on augmentations.</strong></p>",
      "rawMarkdown": "**Just joined the competition and I am already loving it. There are so many ways to tackle such signals dataset problems, and it always hype me up.**\n\nI am releasing a 1D CNN based baseline which scores **0.72CV & 0.48 LB** without using any Spectrograms / Augmentations or 2D CNN Networks & takes only 2 minutes to run submission. The idea is similar to what @cdeotte has already shared in his Wavenet Notebook, but I worked upon a lot on replicating scores in Pytorch & improving the architecture. \n\n#### Inference Notebook: [link](https://www.kaggle.com/code/nischaydnk/hms-submission-1d-eegnet-pipeline-lightning/notebook?scriptVersionId=160814854) \n\n#### Training Notebook: [link](https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs?scriptVersionId=160814948)\n\n\nThe model is slightly based on our 4th Place solution from G2Net competition hosted 2 years back, which includes using Parallel 1D convolutions as feature extractors and then using 1D ResNet based blocks. You can read more about the solution in the [paper](https://www.researchgate.net/publication/359051366_GWNET_Detecting_Gravitational_Waves_using_Hierarchical_and_Residual_Learning_based_1D_CNNs) we released. \n \n#### **Current training settings & Methods Used: **\n- Basic Preprocessing: **Raw Signals** --> **LowPass Filter**-->**mu_law_encoding**\n- Cosine Annealing Scheduler\n- Validate every 0.5 epochs\n- 5 Fold GroupKFold\n- KLDivLoss\n\nI haven't experimented much with the baseline, but I believe even 1D solutions are capable of scoring as high as 2D Spectrograms & Deep CNN based approaches. Maybe with right configuration & adding augmentations, you can improve scores further. \n\n## Architecture:\n\n#### It is mainly based on 3 main ideas: **Parallel Convolution Blocks** + **ResNet Like 1D Blocks** + **RNN head**\n\nWhenever working with any signals data, I mainly use this kind of generic architecture while developing 1D networks. In most cases, this works well for me. I start off with basic 1D CNN based networks and slowly upgrading it with introducing different ideas like adding residuals, etc. \n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fabbddaa44e5c4220f6bed3d3db688c17%2FScreenshot%202024-01-29%20at%209.50.30%20AM.png?generation=1706505660519033&alt=media)\n\n\n **Parallel Convolution Blocks**\n\nSpecific to this competition data, I found smaller kernels (3,5,7) to work better in comparision to larger Kernels which we used in G2Net competition. Again, its more upto the experimentation, you can play around with different kernel size and decide to go with which one.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9317ffc90ca7cf9cd990cc87250b66ab%2FScreenshot%202024-01-29%20at%209.49.18%20AM.png?generation=1706503620778537&alt=media)\n\n**ResNet Like 1D Blocks**\n\nFollowed by Parallel 1D CNNs, these features are then passed into Residual networks which basically helps in better feature extraction & downsampling. Again, you can play around with different number of blocks you want to use. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F3c07868f0ac8069fee743203c9903107%2FScreenshot%202024-01-29%20at%209.49.13%20AM.png?generation=1706503747689647&alt=media)\n\n**I am planning to release more code with improved baseline and on augmentations.**",
      "votes": 176
    },
    {
      "id": 2627263,
      "postDate": "2024-01-30T15:29:06.337Z",
      "content": "<p>Thanks for sharing the architecture! I think it's really good you are looking at 1D convolutions per channel. If you watch some of the youtube videos on how experts use spectrograms (e.g. <a href=\"https://www.youtube.com/watch?v=W5aWLOgMKEE\" target=\"_blank\">here</a> ), you'll see that they look at the spectrogram in time, and then the each of the four EEG channels per group (i.e. LT, LP, RP, RT). Sometimes the spectrogram makes it easy and the EEG is hard, sometimes vice-versa, sometimes a greater than 10 second window helps you classify the 10 second window, etc. You'll notice that to classify something as, for example, Lateralized Periodic Discharges (LPDs), that you'll need (1) for the events to occur on one side (e.g. LT and/or LP but not RP and RT) which satisfies the \"Lateral\" label, and (2) for the discharges to be periodic with pauses in between spikes and at least 6 in the 10 second interval. On the other hand, for Generalized Periodic Discharges (GPDs), you still need (2), but (1) is modified such that you see the events happening on both sides. Note here that this competition appears to follow the double banana montage. </p>\n<p>Therefore, it is a great idea to try to replicate the way experts process these signals by incorporating it directly into the model architecture, then letting the model learn the necessary features. This will include both temporal and spatial aspects of the EEG signals, along with associated spectrograms.</p>",
      "rawMarkdown": "Thanks for sharing the architecture! I think it's really good you are looking at 1D convolutions per channel. If you watch some of the youtube videos on how experts use spectrograms (e.g. [here](https://www.youtube.com/watch?v=W5aWLOgMKEE) ), you'll see that they look at the spectrogram in time, and then the each of the four EEG channels per group (i.e. LT, LP, RP, RT). Sometimes the spectrogram makes it easy and the EEG is hard, sometimes vice-versa, sometimes a greater than 10 second window helps you classify the 10 second window, etc. You'll notice that to classify something as, for example, Lateralized Periodic Discharges (LPDs), that you'll need (1) for the events to occur on one side (e.g. LT and/or LP but not RP and RT) which satisfies the \"Lateral\" label, and (2) for the discharges to be periodic with pauses in between spikes and at least 6 in the 10 second interval. On the other hand, for Generalized Periodic Discharges (GPDs), you still need (2), but (1) is modified such that you see the events happening on both sides. Note here that this competition appears to follow the double banana montage. \n\nTherefore, it is a great idea to try to replicate the way experts process these signals by incorporating it directly into the model architecture, then letting the model learn the necessary features. This will include both temporal and spatial aspects of the EEG signals, along with associated spectrograms.",
      "votes": 12
    },
    {
      "id": 2624998,
      "postDate": "2024-01-29T06:24:59.127Z",
      "content": "<p>Very interesting. How long does it takes to train?</p>",
      "rawMarkdown": "Very interesting. How long does it takes to train?",
      "votes": 4,
      "replies": [
        {
          "id": 2625119,
          "postDate": "2024-01-29T07:36:50.947Z",
          "content": "<p>Around 1 hour on fp32 training ( 5 folds) on single A6000 machine, you can set it to fp16 and run. Should be more faster. </p>",
          "rawMarkdown": "Around 1 hour on fp32 training ( 5 folds) on single A6000 machine, you can set it to fp16 and run. Should be more faster. ",
          "votes": 8
        }
      ]
    },
    {
      "id": 2629580,
      "postDate": "2024-01-31T21:46:50.240Z",
      "content": "<p>1D ResNet still a good model.</p>",
      "rawMarkdown": "1D ResNet still a good model.",
      "votes": 1
    },
    {
      "id": 2628555,
      "postDate": "2024-01-31T11:27:12.733Z",
      "content": "<p>I've tried it (though slightly modified), and get CV score approximately 0.8. <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> have you used 20 channels raw eeg instead of 8 channel eegs to test the model. How is it the result? It will be very kind of you to post the 20 channel eegs data as public dataset, since the generation of which requires quite a lot of memory and both Kaggle notebook and local kernel crash during the process.</p>",
      "rawMarkdown": "I've tried it (though slightly modified), and get CV score approximately 0.8. @nischaydnk have you used 20 channels raw eeg instead of 8 channel eegs to test the model. How is it the result? It will be very kind of you to post the 20 channel eegs data as public dataset, since the generation of which requires quite a lot of memory and both Kaggle notebook and local kernel crash during the process.",
      "votes": 1,
      "replies": [
        {
          "id": 2629642,
          "postDate": "2024-01-31T23:38:40.130Z",
          "content": "<p>As I increase the number of epochs to 50, I get 0.758 for 5 Gfold</p>",
          "rawMarkdown": "As I increase the number of epochs to 50, I get 0.758 for 5 Gfold",
          "replies": [
            {
              "id": 2629737,
              "postDate": "2024-02-01T02:04:48.187Z",
              "content": "<p>UPD: CV 0.758, LB 0.51</p>",
              "rawMarkdown": "UPD: CV 0.758, LB 0.51"
            },
            {
              "id": 2631363,
              "postDate": "2024-02-01T18:17:52.780Z",
              "content": "<p>Is this using all electrodes or the 8 aggregated channels?</p>",
              "rawMarkdown": "Is this using all electrodes or the 8 aggregated channels?"
            },
            {
              "id": 2631811,
              "postDate": "2024-02-02T00:31:30.787Z",
              "content": "<p>8 electrodes selected from 20 electrodes. No aggregation</p>",
              "rawMarkdown": "8 electrodes selected from 20 electrodes. No aggregation"
            }
          ]
        }
      ]
    },
    {
      "id": 2627926,
      "postDate": "2024-01-31T01:10:10.933Z",
      "content": "<p>interesting</p>",
      "rawMarkdown": "interesting",
      "votes": 1
    },
    {
      "id": 2627354,
      "postDate": "2024-01-30T16:25:42.497Z",
      "content": "<p>Very interesting. How long does it takes to train?</p>",
      "rawMarkdown": "Very interesting. How long does it takes to train?",
      "votes": 1
    },
    {
      "id": 2626235,
      "postDate": "2024-01-29T20:53:27.930Z",
      "content": "<p>Great model architecture! Thanks for sharing.</p>",
      "rawMarkdown": "Great model architecture! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 2624971,
      "postDate": "2024-01-29T05:56:34.480Z",
      "content": "<p>cool works!</p>",
      "rawMarkdown": "cool works!",
      "votes": 1
    },
    {
      "id": 2641781,
      "postDate": "2024-02-07T17:00:03.217Z",
      "content": "<p>Thanks for sharing, do you think an RNN-based approach for the signal can be beneficial to capture the temporal features?</p>",
      "rawMarkdown": "Thanks for sharing, do you think an RNN-based approach for the signal can be beneficial to capture the temporal features?"
    },
    {
      "id": 2629153,
      "postDate": "2024-01-31T16:49:11.767Z",
      "content": "<p>thanks for the info, Why opt for smaller kernels (3, 5, 7) in your 1D ResNet model? Any specific observations on their effectiveness?</p>",
      "rawMarkdown": "thanks for the info, Why opt for smaller kernels (3, 5, 7) in your 1D ResNet model? Any specific observations on their effectiveness?"
    },
    {
      "id": 2626656,
      "postDate": "2024-01-30T06:45:50.737Z",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> I tried to reproduce your results in tensorflow on kaggle GPU but somehow model is not able to train. There is no utilisation of either gpu or cpu after few seconds. I guess memory requirement is limiting this training.</p>",
      "rawMarkdown": "@nischaydnk I tried to reproduce your results in tensorflow on kaggle GPU but somehow model is not able to train. There is no utilisation of either gpu or cpu after few seconds. I guess memory requirement is limiting this training."
    },
    {
      "id": 2626583,
      "postDate": "2024-01-30T05:53:22.937Z",
      "content": "<p>Thank you for your sharing.  Could you clarify why a lowpass filter is used?</p>",
      "rawMarkdown": "Thank you for your sharing.  Could you clarify why a lowpass filter is used?"
    },
    {
      "id": 2625990,
      "postDate": "2024-01-29T18:15:05.087Z",
      "content": "<p>Informative! Keep sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "rawMarkdown": "Informative! Keep sharing @nischaydnk "
    },
    {
      "id": 2625782,
      "postDate": "2024-01-29T15:39:02.617Z",
      "content": "<p>Thanks for your insight. Could you elaborate on the mu law encoding? Specifically, how you choose your value for mu? In your function definition for <code>mu_law_encoding()</code> the second argument is <code>mu</code> and in <code>quantize_data()</code> it is called <code>classes</code>. The value you choose for this argument in your <code>EEGDataset</code> when <code>quantize_data()</code> is called is 1 but there are 6 classes. What exactly does the value of mu represent?</p>",
      "rawMarkdown": "Thanks for your insight. Could you elaborate on the mu law encoding? Specifically, how you choose your value for mu? In your function definition for `mu_law_encoding()` the second argument is `mu` and in `quantize_data()` it is called `classes`. The value you choose for this argument in your `EEGDataset` when `quantize_data()` is called is 1 but there are 6 classes. What exactly does the value of mu represent?"
    },
    {
      "id": 2625475,
      "postDate": "2024-01-29T12:50:10.487Z",
      "content": "<p>Very interesting. How long does it takes to train?</p>\n<p>Great job! Thanks for sharing.<br>\nI was also able to get LB 0.45 with a simple 1D-CNN network</p>",
      "rawMarkdown": "Very interesting. How long does it takes to train?\n\nGreat job! Thanks for sharing.\nI was also able to get LB 0.45 with a simple 1D-CNN network"
    },
    {
      "id": 2625134,
      "postDate": "2024-01-29T07:51:07.973Z",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Thanks for sharing, like your paper. How Kernel size capture the signal len?</p>",
      "rawMarkdown": "@nischaydnk Thanks for sharing, like your paper. How Kernel size capture the signal len?"
    },
    {
      "id": 2625464,
      "postDate": "2024-01-29T12:43:03.860Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2631340,
      "postDate": "2024-02-01T18:04:23.613Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 2626708,
      "postDate": "2024-01-30T07:56:52.723Z",
      "content": "<p>Great work. Thanks for sharing.</p>",
      "rawMarkdown": "Great work. Thanks for sharing."
    },
    {
      "id": 2626416,
      "postDate": "2024-01-30T02:19:48.047Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 2625391,
      "postDate": "2024-01-29T11:34:48.630Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 2627263,
      "author_name": "ajland",
      "author_url": "",
      "post_date": "2024-01-30T15:29:06.337000",
      "content": "<p>Thanks for sharing the architecture! I think it's really good you are looking at 1D convolutions per channel. If you watch some of the youtube videos on how experts use spectrograms (e.g. <a href=\"https://www.youtube.com/watch?v=W5aWLOgMKEE\" target=\"_blank\">here</a> ), you'll see that they look at the spectrogram in time, and then the each of the four EEG channels per group (i.e. LT, LP, RP, RT). Sometimes the spectrogram makes it easy and the EEG is hard, sometimes vice-versa, sometimes a greater than 10 second window helps you classify the 10 second window, etc. You'll notice that to classify something as, for example, Lateralized Periodic Discharges (LPDs), that you'll need (1) for the events to occur on one side (e.g. LT and/or LP but not RP and RT) which satisfies the \"Lateral\" label, and (2) for the discharges to be periodic with pauses in between spikes and at least 6 in the 10 second interval. On the other hand, for Generalized Periodic Discharges (GPDs), you still need (2), but (1) is modified such that you see the events happening on both sides. Note here that this competition appears to follow the double banana montage. </p>\n<p>Therefore, it is a great idea to try to replicate the way experts process these signals by incorporating it directly into the model architecture, then letting the model learn the necessary features. This will include both temporal and spatial aspects of the EEG signals, along with associated spectrograms.</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 2624998,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-01-29T06:24:59.127000",
      "content": "<p>Very interesting. How long does it takes to train?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2625119,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2024-01-29T07:36:50.947000",
          "content": "<p>Around 1 hour on fp32 training ( 5 folds) on single A6000 machine, you can set it to fp16 and run. Should be more faster. </p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 2629580,
      "author_name": "Sofia García",
      "author_url": "",
      "post_date": "2024-01-31T21:46:50.240000",
      "content": "<p>1D ResNet still a good model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2628555,
      "author_name": "Roy Wei",
      "author_url": "",
      "post_date": "2024-01-31T11:27:12.733000",
      "content": "<p>I've tried it (though slightly modified), and get CV score approximately 0.8. <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> have you used 20 channels raw eeg instead of 8 channel eegs to test the model. How is it the result? It will be very kind of you to post the 20 channel eegs data as public dataset, since the generation of which requires quite a lot of memory and both Kaggle notebook and local kernel crash during the process.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2629642,
          "author_name": "Roy Wei",
          "author_url": "",
          "post_date": "2024-01-31T23:38:40.130000",
          "content": "<p>As I increase the number of epochs to 50, I get 0.758 for 5 Gfold</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2629737,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2024-02-01T02:04:48.187000",
              "content": "<p>UPD: CV 0.758, LB 0.51</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2631363,
              "author_name": "Luis Pinto",
              "author_url": "",
              "post_date": "2024-02-01T18:17:52.780000",
              "content": "<p>Is this using all electrodes or the 8 aggregated channels?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2631811,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2024-02-02T00:31:30.787000",
              "content": "<p>8 electrodes selected from 20 electrodes. No aggregation</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2627926,
      "author_name": "David Monserratu",
      "author_url": "",
      "post_date": "2024-01-31T01:10:10.933000",
      "content": "<p>interesting</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2627354,
      "author_name": "Ali Zain",
      "author_url": "",
      "post_date": "2024-01-30T16:25:42.497000",
      "content": "<p>Very interesting. How long does it takes to train?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2626235,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-01-29T20:53:27.930000",
      "content": "<p>Great model architecture! Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2624971,
      "author_name": "Donghui Zhang",
      "author_url": "",
      "post_date": "2024-01-29T05:56:34.480000",
      "content": "<p>cool works!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2641781,
      "author_name": "goyalshubh",
      "author_url": "",
      "post_date": "2024-02-07T17:00:03.217000",
      "content": "<p>Thanks for sharing, do you think an RNN-based approach for the signal can be beneficial to capture the temporal features?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2629153,
      "author_name": "Devang Giri Goswami",
      "author_url": "",
      "post_date": "2024-01-31T16:49:11.767000",
      "content": "<p>thanks for the info, Why opt for smaller kernels (3, 5, 7) in your 1D ResNet model? Any specific observations on their effectiveness?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2626656,
      "author_name": "Shashank Shukla",
      "author_url": "",
      "post_date": "2024-01-30T06:45:50.737000",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> I tried to reproduce your results in tensorflow on kaggle GPU but somehow model is not able to train. There is no utilisation of either gpu or cpu after few seconds. I guess memory requirement is limiting this training.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2626583,
      "author_name": "Zhuqiang Lu",
      "author_url": "",
      "post_date": "2024-01-30T05:53:22.937000",
      "content": "<p>Thank you for your sharing.  Could you clarify why a lowpass filter is used?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2625990,
      "author_name": "Disha Agarwal",
      "author_url": "",
      "post_date": "2024-01-29T18:15:05.087000",
      "content": "<p>Informative! Keep sharing <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2625782,
      "author_name": "Dawson Huth",
      "author_url": "",
      "post_date": "2024-01-29T15:39:02.617000",
      "content": "<p>Thanks for your insight. Could you elaborate on the mu law encoding? Specifically, how you choose your value for mu? In your function definition for <code>mu_law_encoding()</code> the second argument is <code>mu</code> and in <code>quantize_data()</code> it is called <code>classes</code>. The value you choose for this argument in your <code>EEGDataset</code> when <code>quantize_data()</code> is called is 1 but there are 6 classes. What exactly does the value of mu represent?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2625475,
      "author_name": "Tanishq dublish",
      "author_url": "",
      "post_date": "2024-01-29T12:50:10.487000",
      "content": "<p>Very interesting. How long does it takes to train?</p>\n<p>Great job! Thanks for sharing.<br>\nI was also able to get LB 0.45 with a simple 1D-CNN network</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2625134,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-01-29T07:51:07.973000",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> Thanks for sharing, like your paper. How Kernel size capture the signal len?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2625464,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-29T12:43:03.860000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2631340,
      "author_name": "Sushanth Raj Singh K",
      "author_url": "",
      "post_date": "2024-02-01T18:04:23.613000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2626708,
      "author_name": "AIDLRE001",
      "author_url": "",
      "post_date": "2024-01-30T07:56:52.723000",
      "content": "<p>Great work. Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2626416,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2024-01-30T02:19:48.047000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2625391,
      "author_name": "Giselle Teixeira",
      "author_url": "",
      "post_date": "2024-01-29T11:34:48.630000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2624962": "**Just joined the competition and I am already loving it. There are so many ways to tackle such signals dataset problems, and it always hype me up.**\n\nI am releasing a 1D CNN based baseline which scores **0.72CV & 0.48 LB** without using any Spectrograms / Augmentations or 2D CNN Networks & takes only 2 minutes to run submission. The idea is similar to what @cdeotte has already shared in his Wavenet Notebook, but I worked upon a lot on replicating scores in Pytorch & improving the architecture. \n\n#### Inference Notebook: [link](https://www.kaggle.com/code/nischaydnk/hms-submission-1d-eegnet-pipeline-lightning/notebook?scriptVersionId=160814854) \n\n#### Training Notebook: [link](https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs?scriptVersionId=160814948)\n\n\nThe model is slightly based on our 4th Place solution from G2Net competition hosted 2 years back, which includes using Parallel 1D convolutions as feature extractors and then using 1D ResNet based blocks. You can read more about the solution in the [paper](https://www.researchgate.net/publication/359051366_GWNET_Detecting_Gravitational_Waves_using_Hierarchical_and_Residual_Learning_based_1D_CNNs) we released. \n \n#### **Current training settings & Methods Used: **\n- Basic Preprocessing: **Raw Signals** --> **LowPass Filter**-->**mu_law_encoding**\n- Cosine Annealing Scheduler\n- Validate every 0.5 epochs\n- 5 Fold GroupKFold\n- KLDivLoss\n\nI haven't experimented much with the baseline, but I believe even 1D solutions are capable of scoring as high as 2D Spectrograms & Deep CNN based approaches. Maybe with right configuration & adding augmentations, you can improve scores further. \n\n## Architecture:\n\n#### It is mainly based on 3 main ideas: **Parallel Convolution Blocks** + **ResNet Like 1D Blocks** + **RNN head**\n\nWhenever working with any signals data, I mainly use this kind of generic architecture while developing 1D networks. In most cases, this works well for me. I start off with basic 1D CNN based networks and slowly upgrading it with introducing different ideas like adding residuals, etc. \n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fabbddaa44e5c4220f6bed3d3db688c17%2FScreenshot%202024-01-29%20at%209.50.30%20AM.png?generation=1706505660519033&alt=media)\n\n\n **Parallel Convolution Blocks**\n\nSpecific to this competition data, I found smaller kernels (3,5,7) to work better in comparision to larger Kernels which we used in G2Net competition. Again, its more upto the experimentation, you can play around with different kernel size and decide to go with which one.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9317ffc90ca7cf9cd990cc87250b66ab%2FScreenshot%202024-01-29%20at%209.49.18%20AM.png?generation=1706503620778537&alt=media)\n\n**ResNet Like 1D Blocks**\n\nFollowed by Parallel 1D CNNs, these features are then passed into Residual networks which basically helps in better feature extraction & downsampling. Again, you can play around with different number of blocks you want to use. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F3c07868f0ac8069fee743203c9903107%2FScreenshot%202024-01-29%20at%209.49.13%20AM.png?generation=1706503747689647&alt=media)\n\n**I am planning to release more code with improved baseline and on augmentations.**",
    "2627263": "Thanks for sharing the architecture! I think it's really good you are looking at 1D convolutions per channel. If you watch some of the youtube videos on how experts use spectrograms (e.g. [here](https://www.youtube.com/watch?v=W5aWLOgMKEE) ), you'll see that they look at the spectrogram in time, and then the each of the four EEG channels per group (i.e. LT, LP, RP, RT). Sometimes the spectrogram makes it easy and the EEG is hard, sometimes vice-versa, sometimes a greater than 10 second window helps you classify the 10 second window, etc. You'll notice that to classify something as, for example, Lateralized Periodic Discharges (LPDs), that you'll need (1) for the events to occur on one side (e.g. LT and/or LP but not RP and RT) which satisfies the \"Lateral\" label, and (2) for the discharges to be periodic with pauses in between spikes and at least 6 in the 10 second interval. On the other hand, for Generalized Periodic Discharges (GPDs), you still need (2), but (1) is modified such that you see the events happening on both sides. Note here that this competition appears to follow the double banana montage. \n\nTherefore, it is a great idea to try to replicate the way experts process these signals by incorporating it directly into the model architecture, then letting the model learn the necessary features. This will include both temporal and spatial aspects of the EEG signals, along with associated spectrograms.",
    "2624998": "Very interesting. How long does it takes to train?",
    "2629580": "1D ResNet still a good model.",
    "2628555": "I've tried it (though slightly modified), and get CV score approximately 0.8. @nischaydnk have you used 20 channels raw eeg instead of 8 channel eegs to test the model. How is it the result? It will be very kind of you to post the 20 channel eegs data as public dataset, since the generation of which requires quite a lot of memory and both Kaggle notebook and local kernel crash during the process.",
    "2627926": "interesting",
    "2627354": "Very interesting. How long does it takes to train?",
    "2626235": "Great model architecture! Thanks for sharing.",
    "2624971": "cool works!",
    "2641781": "Thanks for sharing, do you think an RNN-based approach for the signal can be beneficial to capture the temporal features?",
    "2629153": "thanks for the info, Why opt for smaller kernels (3, 5, 7) in your 1D ResNet model? Any specific observations on their effectiveness?",
    "2626656": "@nischaydnk I tried to reproduce your results in tensorflow on kaggle GPU but somehow model is not able to train. There is no utilisation of either gpu or cpu after few seconds. I guess memory requirement is limiting this training.",
    "2626583": "Thank you for your sharing.  Could you clarify why a lowpass filter is used?",
    "2625990": "Informative! Keep sharing @nischaydnk ",
    "2625782": "Thanks for your insight. Could you elaborate on the mu law encoding? Specifically, how you choose your value for mu? In your function definition for `mu_law_encoding()` the second argument is `mu` and in `quantize_data()` it is called `classes`. The value you choose for this argument in your `EEGDataset` when `quantize_data()` is called is 1 but there are 6 classes. What exactly does the value of mu represent?",
    "2625475": "Very interesting. How long does it takes to train?\n\nGreat job! Thanks for sharing.\nI was also able to get LB 0.45 with a simple 1D-CNN network",
    "2625134": "@nischaydnk Thanks for sharing, like your paper. How Kernel size capture the signal len?",
    "2625464": "",
    "2631340": "Thanks for sharing",
    "2626708": "Great work. Thanks for sharing.",
    "2626416": "Thanks for sharing",
    "2625391": "Thanks for sharing"
  }
}