{
  "id": 467877,
  "title": "How To Create Spectrogram From Eeg? SOLVED? Boost CV and LB!",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467877",
  "author_name": "Chris Deotte",
  "post_date": "2024-01-14T12:22:35.796000",
  "votes": 169,
  "comment_count": 50,
  "views": 0,
  "content": "<h1>Original Discussion</h1>\n<p>My original post had title \"How can we create spectrograms from eeg?\". The post asked; Does anyone know which 4 time series signals are used to create the 4 spectrograms? We are given 20 eeg time series signals for each time window. How can we combine these 20 signals to make 4 signals which are used to create spectrograms?</p>\n<p>I realize that we won't have 10 minute window (for the newly create 4 series), so we can't recreate the full spectrograms from 50 second windows, but none-the-less, I'm curious how to combine the 20 signals into 4 signals.</p>\n<h1>New Discussion</h1>\n<p>After asking this question many Kagglers provided valuable information and I was able to create spectrograms from eegs. And I trained a model using only these new spectrograms and it performed well. So I think we are close to producing the spectrograms correctly. If anyone has feedback or suggestions, please comment below</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2024/montage2.png\"></p>\n<h1>Starter Notebook</h1>\n<p>Using the information below, I published a starter notebook to create Spectrograms from EEG <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrograms-from-eeg\" target=\"_blank\">here</a></p>\n<h1>UPDATE</h1>\n<p>Exciting news! I think the Magic Formula has been discovered! Discussion post <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a></p>\n<h1>Boost CV and LB Score</h1>\n<p>I have verified that using only these newly created EEG spectrograms (and not Kaggle's spectrograms), we can achieve a CV and LB score better than many public notebooks (including submitting train means)</p>\n<h1>Enjoy!</h1>",
  "messages": [
    {
      "id": 2601428,
      "postDate": "2024-01-14T12:22:35.797Z",
      "content": "<h1>Original Discussion</h1>\n<p>My original post had title \"How can we create spectrograms from eeg?\". The post asked; Does anyone know which 4 time series signals are used to create the 4 spectrograms? We are given 20 eeg time series signals for each time window. How can we combine these 20 signals to make 4 signals which are used to create spectrograms?</p>\n<p>I realize that we won't have 10 minute window (for the newly create 4 series), so we can't recreate the full spectrograms from 50 second windows, but none-the-less, I'm curious how to combine the 20 signals into 4 signals.</p>\n<h1>New Discussion</h1>\n<p>After asking this question many Kagglers provided valuable information and I was able to create spectrograms from eegs. And I trained a model using only these new spectrograms and it performed well. So I think we are close to producing the spectrograms correctly. If anyone has feedback or suggestions, please comment below</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2024/montage2.png\"></p>\n<h1>Starter Notebook</h1>\n<p>Using the information below, I published a starter notebook to create Spectrograms from EEG <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrograms-from-eeg\" target=\"_blank\">here</a></p>\n<h1>UPDATE</h1>\n<p>Exciting news! I think the Magic Formula has been discovered! Discussion post <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a></p>\n<h1>Boost CV and LB Score</h1>\n<p>I have verified that using only these newly created EEG spectrograms (and not Kaggle's spectrograms), we can achieve a CV and LB score better than many public notebooks (including submitting train means)</p>\n<h1>Enjoy!</h1>",
      "rawMarkdown": "# Original Discussion\nMy original post had title \"How can we create spectrograms from eeg?\". The post asked; Does anyone know which 4 time series signals are used to create the 4 spectrograms? We are given 20 eeg time series signals for each time window. How can we combine these 20 signals to make 4 signals which are used to create spectrograms?\n\nI realize that we won't have 10 minute window (for the newly create 4 series), so we can't recreate the full spectrograms from 50 second windows, but none-the-less, I'm curious how to combine the 20 signals into 4 signals.\n\n# New Discussion\nAfter asking this question many Kagglers provided valuable information and I was able to create spectrograms from eegs. And I trained a model using only these new spectrograms and it performed well. So I think we are close to producing the spectrograms correctly. If anyone has feedback or suggestions, please comment below\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2024/montage2.png)\n\n# Starter Notebook\nUsing the information below, I published a starter notebook to create Spectrograms from EEG [here][1]\n\n# UPDATE\nExciting news! I think the Magic Formula has been discovered! Discussion post [here][2]\n\n# Boost CV and LB Score\nI have verified that using only these newly created EEG spectrograms (and not Kaggle's spectrograms), we can achieve a CV and LB score better than many public notebooks (including submitting train means)\n\n# Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrograms-from-eeg\n[2]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760",
      "votes": 168
    },
    {
      "id": 2601478,
      "postDate": "2024-01-14T12:53:50.303Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> from my understanding</p>\n<p><strong>Note</strong> 16 channels out of 19 used and 1 EKG</p>\n<pre><code>pairing = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : , //  part of LL, LP, RP, RR\n    :   //  part of LL, LP, RP, RR\n}\n</code></pre>\n<p>from  <a href=\"https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features\" target=\"_blank\">Notebook - Eegs Pairing Analysis &amp; Features</a></p>\n<p><strong>LL</strong> -&gt;  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F42abf3341243af65c92c8ba2083b9a79%2FScreenshot%202024-01-14%20at%205.49.27PM.png?generation=1705234855564208&amp;alt=media\"><br>\n<strong>LP</strong> -&gt;<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa725bf3a965b4299d0af07f4937216f5%2FScreenshot%202024-01-14%20at%205.49.38PM.png?generation=1705234875852002&amp;alt=media\"><br>\n<strong>RP</strong> -&gt;<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5869dccab21b6cc3ba7b31e849622d5a%2FScreenshot%202024-01-14%20at%205.50.02PM.png?generation=1705234897860474&amp;alt=media\"><br>\n<strong>RR</strong> -&gt; <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa13eff4393d4fc67fbf3a2bc52294553%2FScreenshot%202024-01-14%20at%205.50.12PM.png?generation=1705234918197292&amp;alt=media\"><br>\nfrom  <a href=\"https://www.kaggle.com/code/seshurajup/eegs-10-20-system\" target=\"_blank\">Notebook - EEGS 10–20 system</a></p>\n<p>This video is awesome to understand basics <a href=\"https://www.youtube.com/watch?v=XMizSSOejg0&amp;ab_channel=JeremyMoeller\" target=\"_blank\">Source Youtube</a></p>\n<p><strong>Ignored Middle part channels</strong></p>",
      "rawMarkdown": "@cdeotte from my understanding\n\n**Note** 16 channels out of 19 used and 1 EKG\n```python\npairing = {\n    \"Fp1\": \"F7\",\n    \"F7\": \"T3\",\n    \"T3\": \"T5\",\n    \"T5\": \"O1\",\n    \"Fp2\": \"F8\",\n    \"F8\": \"T4\",\n    \"T4\": \"T6\",\n    \"T6\": \"O2\",\n    \"Fp1\": \"F3\",\n    \"F3\": \"C3\",\n    \"C3\": \"P3\",\n    \"P3\": \"O1\",\n    \"Fp2\": \"F4\",\n    \"F4\": \"C4\",\n    \"C4\": \"P4\",\n    \"P4\": \"O2\",\n    \"Fz\": \"Cz\", // not part of LL, LP, RP, RR\n    \"Cz\": \"Pz\"  // not part of LL, LP, RP, RR\n}\n```\nfrom  [Notebook - Eegs Pairing Analysis & Features](https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features)\n\n**LL** ->  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F42abf3341243af65c92c8ba2083b9a79%2FScreenshot%202024-01-14%20at%205.49.27PM.png?generation=1705234855564208&alt=media)\n**LP** ->\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa725bf3a965b4299d0af07f4937216f5%2FScreenshot%202024-01-14%20at%205.49.38PM.png?generation=1705234875852002&alt=media)\n**RP** ->\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5869dccab21b6cc3ba7b31e849622d5a%2FScreenshot%202024-01-14%20at%205.50.02PM.png?generation=1705234897860474&alt=media)\n**RR** -> \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa13eff4393d4fc67fbf3a2bc52294553%2FScreenshot%202024-01-14%20at%205.50.12PM.png?generation=1705234918197292&alt=media)\nfrom  [Notebook - EEGS 10–20 system](https://www.kaggle.com/code/seshurajup/eegs-10-20-system)\n\nThis video is awesome to understand basics [Source Youtube](https://www.youtube.com/watch?v=XMizSSOejg0&ab_channel=JeremyMoeller)\n\n**Ignored Middle part channels**",
      "votes": 15,
      "replies": [
        {
          "id": 2601504,
          "postDate": "2024-01-14T13:25:18.537Z",
          "content": "<p>In your diagram above, LL is 4 signals (i.e. 4 differences). We need LP to be 1 signal. Do we just average the 4? (Also LP is 4, RP is 4, RR is 4)</p>",
          "rawMarkdown": "In your diagram above, LL is 4 signals (i.e. 4 differences). We need LP to be 1 signal. Do we just average the 4? (Also LP is 4, RP is 4, RR is 4)",
          "votes": 1,
          "replies": [
            {
              "id": 2601518,
              "postDate": "2024-01-14T13:34:21.800Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> As per my understanding,</p>\n<p>calculated spectrograms of each pair of LL channels. then average of these spectrograms as final LL spectrogram<br>\nFp1 -&gt; F7 =&gt; 1st spectrogram<br>\nF7 -&gt; T7 =&gt; 2nd spectrogram <br>\nT7 -&gt; P7 =&gt; 3rd spectrogram <br>\nP7 -&gt; 01 =&gt; 4th spectrogram<br>\nAverage of all 4 spectrograms =&gt; LL spectrogram</p>\n<p><em>Note</em> correct me if am wrong</p>",
              "rawMarkdown": "@cdeotte As per my understanding,\n\ncalculated spectrograms of each pair of LL channels. then average of these spectrograms as final LL spectrogram\nFp1 -> F7 => 1st spectrogram\nF7 -> T7 => 2nd spectrogram \nT7 -> P7 => 3rd spectrogram \nP7 -> 01 => 4th spectrogram\nAverage of all 4 spectrograms => LL spectrogram\n\n*Note* correct me if am wrong",
              "votes": 3
            },
            {
              "id": 2601523,
              "postDate": "2024-01-14T13:40:05.813Z",
              "content": "<p>Oh, that would work. Did you read that somewhere or are you guessing?</p>\n<p>This means we can take the average of the four groups (LL LP RR RP) of 4 signals to create 4 new signals. Then if we apply WaveNet to these 4 newly created eeg signals, we should be able to achieve similar performance as spectrogram only models because WaveNet extracts frequencies from signals.</p>\n<p>(If you prefer, you can crop the middle 10 (or 20 or 50) seconds from both spectrogram and eeg. And ignore the extra information. However more information will create better models).</p>",
              "rawMarkdown": "Oh, that would work. Did you read that somewhere or are you guessing?\n\nThis means we can take the average of the four groups (LL LP RR RP) of 4 signals to create 4 new signals. Then if we apply WaveNet to these 4 newly created eeg signals, we should be able to achieve similar performance as spectrogram only models because WaveNet extracts frequencies from signals.\n\n(If you prefer, you can crop the middle 10 (or 20 or 50) seconds from both spectrogram and eeg. And ignore the extra information. However more information will create better models).",
              "votes": 3
            },
            {
              "id": 2601620,
              "postDate": "2024-01-14T15:01:02.297Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> From these video from where i assumed</p>\n<p><a href=\"https://www.youtube.com/@eegforanesthesia3954\" target=\"_blank\">Source from this youtube channel</a></p>",
              "rawMarkdown": "@cdeotte From these video from where i assumed\n\n[Source from this youtube channel](https://www.youtube.com/@eegforanesthesia3954)",
              "votes": 2
            },
            {
              "id": 2614562,
              "postDate": "2024-01-22T17:32:09.387Z",
              "content": "<p>Hey Chris, did you get a chance to test this? Due to memory constraints, I was able to only use LL and RR but performance is the same as your base wavenet notebook</p>",
              "rawMarkdown": "Hey Chris, did you get a chance to test this? Due to memory constraints, I was able to only use LL and RR but performance is the same as your base wavenet notebook"
            },
            {
              "id": 2614591,
              "postDate": "2024-01-22T17:55:11.210Z",
              "content": "<p>Yes, I did this and posted a bunch of starter notebook. It works well. Check out my notebooks in the code section.</p>",
              "rawMarkdown": "Yes, I did this and posted a bunch of starter notebook. It works well. Check out my notebooks in the code section.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2632308,
          "postDate": "2024-02-02T09:01:55.660Z",
          "content": "<p>Thanks! The video helped a lot :)</p>",
          "rawMarkdown": "Thanks! The video helped a lot :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2604075,
      "postDate": "2024-01-16T08:36:57.343Z",
      "content": "<p>Congratulations, Chris, on your innovative method of generating EEG-derived spectrograms, leading to improved CV and LB scores in the Kaggle competition. This approach opens up new avenues for data utilization.</p>\n<p>I'm curious to know: How do the characteristics and information content of your EEG-derived spectrograms differ from the original Kaggle-provided ones, and how do these differences contribute to the enhanced model performance?</p>",
      "rawMarkdown": "Congratulations, Chris, on your innovative method of generating EEG-derived spectrograms, leading to improved CV and LB scores in the Kaggle competition. This approach opens up new avenues for data utilization.\n\nI'm curious to know: How do the characteristics and information content of your EEG-derived spectrograms differ from the original Kaggle-provided ones, and how do these differences contribute to the enhanced model performance?",
      "votes": 7,
      "replies": [
        {
          "id": 2604470,
          "postDate": "2024-01-16T12:52:03.553Z",
          "content": "<p>Thanks. Great question Nishant about comparing. I have not compared yet but I have observed that training models with both Kaggle's spectrograms and my EEG spectrograms achieves better CV score and LB score (than Kaggle's alone). Therefore there is definitely new information in my EEG spectrograms compared with Kaggle's spectrograms.</p>\n<p>One difference is that Kaggle's are 10 minutes long whereas mine are only 50 seconds. It would be interesting to compare my 50 seconds with Kaggle's middle 50 seconds and see how they differ. I suspect they will differ because:</p>\n<ul>\n<li>Kaggle may create their spectrograms from signals not available to us</li>\n<li>My formula of <code>LL = ( (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1) )/4.</code> may not be the correct way to create the single LL time series.</li>\n<li>My spectrogram hyperparameter settings differ from Kaggle's</li>\n<li>Kaggle may have used denoise before creating spectrograms</li>\n<li>etc etc</li>\n</ul>",
          "rawMarkdown": "Thanks. Great question Nishant about comparing. I have not compared yet but I have observed that training models with both Kaggle's spectrograms and my EEG spectrograms achieves better CV score and LB score (than Kaggle's alone). Therefore there is definitely new information in my EEG spectrograms compared with Kaggle's spectrograms.\n\nOne difference is that Kaggle's are 10 minutes long whereas mine are only 50 seconds. It would be interesting to compare my 50 seconds with Kaggle's middle 50 seconds and see how they differ. I suspect they will differ because:\n* Kaggle may create their spectrograms from signals not available to us\n* My formula of `LL = ( (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1) )/4.` may not be the correct way to create the single LL time series.\n* My spectrogram hyperparameter settings differ from Kaggle's\n* Kaggle may have used denoise before creating spectrograms\n* etc etc",
          "votes": 12,
          "replies": [
            {
              "id": 2605136,
              "postDate": "2024-01-16T21:38:09.123Z",
              "content": "<p>Isn't the LL just 1/4 * (Fp1 - O1) then, because all the middle terms are eliminated? So all the specs are defined only by two sensors thus LL and LP are the same (and also RL and RP). Or there is some time lag between difference terms? I've checked in your notebook about specs calculation and it looks like the spectrograms and raw signals are really identical.</p>",
              "rawMarkdown": "Isn't the LL just 1/4 * (Fp1 - O1) then, because all the middle terms are eliminated? So all the specs are defined only by two sensors thus LL and LP are the same (and also RL and RP). Or there is some time lag between difference terms? I've checked in your notebook about specs calculation and it looks like the spectrograms and raw signals are really identical.",
              "votes": 3
            },
            {
              "id": 2605142,
              "postDate": "2024-01-16T21:42:43.773Z",
              "content": "<p>Hmm… that's a great point. I overlooked that the formula reduces to that. Thus, that formula doesn't look correct. I would think that the spectrogram would take into account all 5 electrodes Fp1, F7, T3, T5, O1 in the Left Temporal Chain.</p>\n<p>I wonder if i create 5 spectrograms for the 5 differences and then take the average of the 5 spectrograms. I wonder if that would be mathematically different than taking 1 spectrogram of Fp1-O1?</p>",
              "rawMarkdown": "Hmm... that's a great point. I overlooked that the formula reduces to that. Thus, that formula doesn't look correct. I would think that the spectrogram would take into account all 5 electrodes Fp1, F7, T3, T5, O1 in the Left Temporal Chain.\n\nI wonder if i create 5 spectrograms for the 5 differences and then take the average of the 5 spectrograms. I wonder if that would be mathematically different than taking 1 spectrogram of Fp1-O1?",
              "votes": 3
            },
            {
              "id": 2605202,
              "postDate": "2024-01-16T23:33:57.210Z",
              "content": "<p>It should be the same again since STFT is based on the fourier transform (i.e. f(x) = sum a_i * h_i(x) where h_i(x) are basis functions, sin/cos in the terms of the fourier transform) so f(x) + g(x) = sum (a_i + b_i) * h_i(x). However if you do it in the log (dB) space  (actually no, mel spectrogram is just linear transform), then the result can be different. However the simpliest way to use this approach is just make 4 different channels, but I do not think it will bring a big difference scince you will probably train a network so all possible linear calculations can be made in the first layer, so you can just use raw channels and pairwise differences can be (theoretically) trained.</p>",
              "rawMarkdown": "It should be the same again since STFT is based on the fourier transform (i.e. f(x) = sum a_i * h_i(x) where h_i(x) are basis functions, sin/cos in the terms of the fourier transform) so f(x) + g(x) = sum (a_i + b_i) * h_i(x). However if you do it in the log (dB) space ~~and probably take some mel-features~~ (actually no, mel spectrogram is just linear transform), then the result can be different. However the simpliest way to use this approach is just make 4 different channels, but I do not think it will bring a big difference scince you will probably train a network so all possible linear calculations can be made in the first layer, so you can just use raw channels and pairwise differences can be (theoretically) trained.",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2601534,
      "postDate": "2024-01-14T13:53:08.980Z",
      "content": "<p>I'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the EEG has 90 seconds. It's supposed to be the same lenght.</p>\n<p>For example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz</p>",
      "rawMarkdown": "I'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the EEG has 90 seconds. It's supposed to be the same lenght.\n\nFor example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz",
      "votes": 6,
      "replies": [
        {
          "id": 2601590,
          "postDate": "2024-01-14T14:27:57.687Z",
          "content": "<p>The train data is just two different perspectives. The eeg is a \"zoom in\" to what was happening during 50 seconds and the spectrogram is a \"zoom out\" of what was happening during 10 minutes. They are both centered at the same moment of time (i.e. the timestamp in middle of spectrogram is the same timestamp in middle of eeg).</p>\n<p>Our task is to classify what is happening in the middle 10 seconds.</p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.</p>\n</blockquote>",
          "rawMarkdown": "The train data is just two different perspectives. The eeg is a \"zoom in\" to what was happening during 50 seconds and the spectrogram is a \"zoom out\" of what was happening during 10 minutes. They are both centered at the same moment of time (i.e. the timestamp in middle of spectrogram is the same timestamp in middle of eeg).\n\nOur task is to classify what is happening in the middle 10 seconds.\n>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.",
          "votes": 13,
          "replies": [
            {
              "id": 2611416,
              "postDate": "2024-01-20T18:06:56.510Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Thank you for this amazing explanation. Looking forward to learn some more amazing things. 😀</p>",
              "rawMarkdown": "Hi @cdeotte. Thank you for this amazing explanation. Looking forward to learn some more amazing things. 😀",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2612157,
      "postDate": "2024-01-21T08:20:58.450Z",
      "content": "<p>single fold resnet34d</p>\n<p>kaggle provided:val_loss = 0.6532<br>\nspectrogram_from_eeg:val_loss=0.8286</p>\n<pre><code>\n = np(f),,\n = cv2(np(,axis=),(,))\n = self(image=img)\n....\n</code></pre>\n<p>spectrogram_from_eeg+kaggle provided 2ch:val_loss=0.6695</p>\n<pre><code> = np(f)\n = cv2(,(,))\n\neeg_spec = np(f),,\neeg_spec = np(eeg_spec,axis=)\neeg_spec = cv2(eeg_spec,(,))\n\n = np(,axis=-)\n = self(image=img)\n....\n</code></pre>\n<p>I don't seem to be making good use of the new features …</p>",
      "rawMarkdown": "single fold resnet34d\n\nkaggle provided:val_loss = 0.6532\nspectrogram_from_eeg:val_loss=0.8286\n```\n#spectrogram_from_eeg\nimg = np.load(f\"spec_from_eeg/{eeg_id}.npy\")#128,256,4\nimg = cv2.resize(np.concatenate(img,axis=1).T,(512,512))\nimg = self.transform(image=img)[\"image\"]\n....\n```\n\nspectrogram_from_eeg+kaggle provided 2ch:val_loss=0.6695\n```\nimg = np.load(f\"kaggle_spec/{label_id}.npy\")\nimg = cv2.resize(img,(512,512))\n        \neeg_spec = np.load(f\"spec_from_eeg/{eeg_id}.npy\")#128,256,4\neeg_spec = np.concatenate(eeg_spec,axis=1).T\neeg_spec = cv2.resize(eeg_spec,(512,512))\n        \nimg = np.stack([img,eeg_spec],axis=-1)\nimg = self.transform(image=img)[\"image\"]\n....\n```\n\nI don't seem to be making good use of the new features ...",
      "votes": 3,
      "replies": [
        {
          "id": 2612394,
          "postDate": "2024-01-21T11:20:20.113Z",
          "content": "<p>Soon, i will update my CatBoost starter and EffNet starter to use the new features. Afterward, they will become the new top scoring public notebooks 😀</p>",
          "rawMarkdown": "Soon, i will update my CatBoost starter and EffNet starter to use the new features. Afterward, they will become the new top scoring public notebooks 😀",
          "votes": 11
        },
        {
          "id": 2612461,
          "postDate": "2024-01-21T12:19:25.190Z",
          "content": "<p>same issue and wait for Chris the new updated notebook, so appreciate that</p>",
          "rawMarkdown": "same issue and wait for Chris the new updated notebook, so appreciate that",
          "votes": 2,
          "replies": [
            {
              "id": 2613268,
              "postDate": "2024-01-22T00:45:06.753Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> and <a href=\"https://www.kaggle.com/leehann\" target=\"_blank\">@leehann</a> I published examples of how to use my eeg spectrograms in my EfficientNet starter notebook. I also published new eeg spectrograms in my \"how to create spectrogram\" starter notebook. Versions 1-3 of my \"how to create spectrograms\" starter notebook are called \"my old eeg spectrograms\". Version 4 uses a new better formula and is called \"my new eeg spectrograms\". The example code (with its hyperparameters etc) in my EfficientNet starter notebook works for both the new eeg spectrograms and my old eeg spectrograms. Here are the CV scores:</p>\n<ul>\n<li>Using only Kaggle spectrograms achieves 5Fold CV 0.66</li>\n<li>Using only old eeg spectrograms achieves 5Fold CV 0.78</li>\n<li>Using only new eeg spectrograms achieves 5Fold CV 0.63</li>\n</ul>",
              "rawMarkdown": "Hi @abebe9849 and @leehann I published examples of how to use my eeg spectrograms in my EfficientNet starter notebook. I also published new eeg spectrograms in my \"how to create spectrogram\" starter notebook. Versions 1-3 of my \"how to create spectrograms\" starter notebook are called \"my old eeg spectrograms\". Version 4 uses a new better formula and is called \"my new eeg spectrograms\". The example code (with its hyperparameters etc) in my EfficientNet starter notebook works for both the new eeg spectrograms and my old eeg spectrograms. Here are the CV scores:\n* Using only Kaggle spectrograms achieves 5Fold CV 0.66\n* Using only old eeg spectrograms achieves 5Fold CV 0.78\n* Using only new eeg spectrograms achieves 5Fold CV 0.63",
              "votes": 6
            },
            {
              "id": 2613790,
              "postDate": "2024-01-22T09:18:32.697Z",
              "content": "<p>thanks so much agian, respect</p>",
              "rawMarkdown": "thanks so much agian, respect",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2609331,
      "postDate": "2024-01-19T11:51:10.890Z",
      "content": "<p>Hi Chris</p>\n<p>so so great job you had done, i am using the data generated from above method, but the cv and lb get worse, may i ask how to use these data, i must missed something. P.S i just used the genreated data to replace official data</p>",
      "rawMarkdown": "Hi Chris\n\nso so great job you had done, i am using the data generated from above method, but the cv and lb get worse, may i ask how to use these data, i must missed something. P.S i just used the genreated data to replace official data",
      "votes": 4,
      "replies": [
        {
          "id": 2609346,
          "postDate": "2024-01-19T12:09:55.130Z",
          "content": "<p>This data is already standardized to be values between -1 and 1. So  if you use existing public notebooks with this data, remove the log transform and standardization.</p>\n<p>(The Kaggle spectrograms need preprocess of log transform and standardization whereas my EEG spectrograms do not)</p>",
          "rawMarkdown": "This data is already standardized to be values between -1 and 1. So  if you use existing public notebooks with this data, remove the log transform and standardization.\n\n(The Kaggle spectrograms need preprocess of log transform and standardization whereas my EEG spectrograms do not)",
          "votes": 3,
          "replies": [
            {
              "id": 2609375,
              "postDate": "2024-01-19T12:40:26.187Z",
              "content": "<p>ok, thanks so much and let me try it</p>",
              "rawMarkdown": "ok, thanks so much and let me try it"
            },
            {
              "id": 2610484,
              "postDate": "2024-01-20T06:47:31.197Z",
              "content": "<p>Hi Chris<br>\nSorry to bother you again, i tried your data output without any further process, but it still did not work, may i ask if it is convenient that you can get a small example that how to use your spec data, if i did not miss something, all the code were based on official code right now.</p>\n<blockquote>\n  <p>This data is already standardized to be values between -1 and 1. So  if you use existing public notebooks with this data, remove the log transform and standardization.</p>\n  <p>(The Kaggle spectrograms need preprocess of log transform and standardization whereas my EEG spectrograms do not)</p>\n</blockquote>",
              "rawMarkdown": "Hi Chris\nSorry to bother you again, i tried your data output without any further process, but it still did not work, may i ask if it is convenient that you can get a small example that how to use your spec data, if i did not miss something, all the code were based on official code right now.\n\n\n> This data is already standardized to be values between -1 and 1. So  if you use existing public notebooks with this data, remove the log transform and standardization.\n> \n> (The Kaggle spectrograms need preprocess of log transform and standardization whereas my EEG spectrograms do not)\n\n"
            },
            {
              "id": 2610521,
              "postDate": "2024-01-20T07:19:09.450Z",
              "content": "<p>What does \"did not work\" mean? What model are you using and what CV score and LB score are you achieving? There is an example of an MLP in the notebook which achieves CV 1.0 and LB 0.77 (whereas submitting train means achieves CV 1.26 LB 0.97)</p>",
              "rawMarkdown": "What does \"did not work\" mean? What model are you using and what CV score and LB score are you achieving? There is an example of an MLP in the notebook which achieves CV 1.0 and LB 0.77 (whereas submitting train means achieves CV 1.26 LB 0.97)",
              "votes": 2
            },
            {
              "id": 2610558,
              "postDate": "2024-01-20T07:48:04.763Z",
              "content": "<p>the same model and traing hyper, the cv is alway 1.0+ did not drop during the traing process, but the official spec training is normal</p>",
              "rawMarkdown": "the same model and traing hyper, the cv is alway 1.0+ did not drop during the traing process, but the official spec training is normal",
              "votes": 1
            },
            {
              "id": 2611595,
              "postDate": "2024-01-20T20:59:03.133Z",
              "content": "<p><a href=\"https://www.kaggle.com/leehann\" target=\"_blank\">@leehann</a> I'm experiencing the same</p>",
              "rawMarkdown": "@leehann I'm experiencing the same",
              "votes": 1
            },
            {
              "id": 2612432,
              "postDate": "2024-01-21T11:56:52.853Z",
              "content": "<p>Hi HZM and Yan, soon I will update my CatBoost and EffNet starter notebooks to use the EEG spectrograms.</p>",
              "rawMarkdown": "Hi HZM and Yan, soon I will update my CatBoost and EffNet starter notebooks to use the EEG spectrograms.",
              "votes": 1
            },
            {
              "id": 2612520,
              "postDate": "2024-01-21T12:51:41.040Z",
              "content": "<p>thanks so much for your help Chris, you are save</p>",
              "rawMarkdown": "thanks so much for your help Chris, you are save",
              "votes": 1
            },
            {
              "id": 2614233,
              "postDate": "2024-01-22T14:04:21.467Z",
              "content": "<p>Hi Yan, Are u using pytorch? I still had the issue when using updated datasets</p>",
              "rawMarkdown": "Hi Yan, Are u using pytorch? I still had the issue when using updated datasets"
            }
          ]
        }
      ]
    },
    {
      "id": 2603599,
      "postDate": "2024-01-15T23:04:55.577Z",
      "content": "<p>to add to the previous comments, spectrograms were computed with some variant of Fourier transform (most likely Wavelet Convolution, or Complex Morelet Wavelet), and then averaged across electrode channels, as suggested by <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> </p>\n<p>To compute your own time-frequency spectrograms (which is a good idea IMHO), you could follow the User Guide of MNE package <a href=\"https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html\" target=\"_blank\">https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html</a></p>",
      "rawMarkdown": "to add to the previous comments, spectrograms were computed with some variant of Fourier transform (most likely Wavelet Convolution, or Complex Morelet Wavelet), and then averaged across electrode channels, as suggested by @seshurajup \n\nTo compute your own time-frequency spectrograms (which is a good idea IMHO), you could follow the User Guide of MNE package https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html",
      "votes": 4,
      "replies": [
        {
          "id": 2603600,
          "postDate": "2024-01-15T23:12:07.470Z",
          "content": "<p><a href=\"https://www.kaggle.com/iworeushankaonce\" target=\"_blank\">@iworeushankaonce</a> </p>\n<blockquote>\n  <p>to add to the previous comments, spectrograms were computed with some variant of Fourier transform (most likely Wavelet Convolution, or Complex Morelet Wavelet), and then averaged across electrode channels, as suggested by <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> </p>\n  <ul>\n  <li>Its my assumption need to conform from experts in this area</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <p>To compute your own time-frequency spectrograms (which is a good idea IMHO), you could follow docudocumentat of MNE package <a href=\"https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html\" target=\"_blank\">https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html</a></p>\n  <ul>\n  <li>spectrograms are of sequence len 10mins long, where egg sequences not having complete 10mins long sequences.</li>\n  </ul>\n</blockquote>",
          "rawMarkdown": "@iworeushankaonce \n> to add to the previous comments, spectrograms were computed with some variant of Fourier transform (most likely Wavelet Convolution, or Complex Morelet Wavelet), and then averaged across electrode channels, as suggested by @seshurajup \n- Its my assumption need to conform from experts in this area\n\n-------\n> To compute your own time-frequency spectrograms (which is a good idea IMHO), you could follow docudocumentat of MNE package https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html\n- spectrograms are of sequence len 10mins long, where egg sequences not having complete 10mins long sequences.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2612552,
      "postDate": "2024-01-21T13:18:37.167Z",
      "content": "<p>Thank you for your great post again !<br>\nI have a question about below.</p>\n<p>mel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, <br>\n              n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)</p>\n<p>How did you decide the parameter of (hop_length=len(x)//256,  n_fft=1024, n_mels=128, win_length=128) ?<br>\nThank you first.</p>",
      "rawMarkdown": "Thank you for your great post again !\nI have a question about below.\n\nmel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, \n              n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)\n\nHow did you decide the parameter of (hop_length=len(x)//256,  n_fft=1024, n_mels=128, win_length=128) ?\nThank you first.",
      "votes": 1,
      "replies": [
        {
          "id": 2612609,
          "postDate": "2024-01-21T13:49:30.620Z",
          "content": "<p>This is explained in the notebook. <code>hop_length</code> will determine the width of our image. I wanted an image of width 256, so i chose <code>hop_length=len(x)//256</code> where <code>x</code> is the input waveform. The variable <code>n_mels</code> will determine the height of the image. I wanted 128. I picked 256 and 128 because I wanted it to be similar to Kaggle's 300x100 and for image models it is best to have images with dimensions that have multiples of 32.</p>\n<p>The variables <code>n_fft</code> and <code>win_lenth</code> control the sliding window size of the Fourier transform. It is my understanding that a sliding window moves over the original waveform and computes the image in pieces. Hence these variables control the image resolution. I tried a few common ones until the image looked good to my eye. I think <code>win_length</code> is the length of time window and hence affects the image horizontal resolution. And <code>n_fft</code> affects vertical image resolution by deciding the frequency resolution when creating a DFT (discrete fourier transform)</p>",
          "rawMarkdown": "This is explained in the notebook. `hop_length` will determine the width of our image. I wanted an image of width 256, so i chose `hop_length=len(x)//256` where `x` is the input waveform. The variable `n_mels` will determine the height of the image. I wanted 128. I picked 256 and 128 because I wanted it to be similar to Kaggle's 300x100 and for image models it is best to have images with dimensions that have multiples of 32.\n\nThe variables `n_fft` and `win_lenth` control the sliding window size of the Fourier transform. It is my understanding that a sliding window moves over the original waveform and computes the image in pieces. Hence these variables control the image resolution. I tried a few common ones until the image looked good to my eye. I think `win_length` is the length of time window and hence affects the image horizontal resolution. And `n_fft` affects vertical image resolution by deciding the frequency resolution when creating a DFT (discrete fourier transform)",
          "votes": 6,
          "replies": [
            {
              "id": 2613713,
              "postDate": "2024-01-22T08:36:24.410Z",
              "content": "<p>Thank you for quick replay and explanation !</p>",
              "rawMarkdown": "Thank you for quick replay and explanation !",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2609539,
      "postDate": "2024-01-19T14:45:21.553Z",
      "content": "<p>Can we combine this created Spectrograms with  Kaggle's spectrograms,and put them together for training?</p>",
      "rawMarkdown": "Can we combine this created Spectrograms with  Kaggle's spectrograms,and put them together for training?",
      "votes": 1,
      "replies": [
        {
          "id": 2609554,
          "postDate": "2024-01-19T15:09:09.923Z",
          "content": "<p>Yes, we can do whatever we want. The best way to see if it improves CV score and LB score is to try it.</p>",
          "rawMarkdown": "Yes, we can do whatever we want. The best way to see if it improves CV score and LB score is to try it.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2632593,
      "postDate": "2024-02-02T13:40:13.190Z",
      "content": "<p>I'm seeking help interpreting the seizure activity labels in the train.csv file. Some of the seizure events seem shorter than the 50 second test windows.<br>\n10-second durations are used as well including in the discussion and in the example_figures directory. <br>\nThe eeg_label_offset_seconds in train.csv range from 1-6 seconds for some rows with identical eeg_sub_id, i.e. the same eeg signal.</p>\n<p>For example, these rows from train.csv:<br>\nEeg_id,eeg_sub_id, eeg_label_offset_seconds, seizure_vote<br>\n110, 3132287169, 3, Seizure<br>\n111, 3132287169, 4, Seizure<br>\n112, 3132287169, 5, Seizure<br>\nDo these 3 rows indicate 3 separate seizure events in the same eeg ? Taking a 10 seconds window will cover, in these cases, the neighbouring seizure event.</p>\n<p>I'm trying to understand how these 10 second analysis windows with seizure labels correlate to the specific timing of the seizure events themselves. What am I missing about how the 10 second windows relate to the labeled seizure activity start/stop times?</p>\n<p>If anyone has already covered this in the discussions, please point me there! I'm still catching up on all the great info.</p>",
      "rawMarkdown": "I'm seeking help interpreting the seizure activity labels in the train.csv file. Some of the seizure events seem shorter than the 50 second test windows.\n10-second durations are used as well including in the discussion and in the example_figures directory. \nThe eeg_label_offset_seconds in train.csv range from 1-6 seconds for some rows with identical eeg_sub_id, i.e. the same eeg signal.\n\n\nFor example, these rows from train.csv:\nEeg_id,eeg_sub_id, eeg_label_offset_seconds, seizure_vote\n110, 3132287169, 3, Seizure\n111, 3132287169, 4, Seizure\n112, 3132287169, 5, Seizure\nDo these 3 rows indicate 3 separate seizure events in the same eeg ? Taking a 10 seconds window will cover, in these cases, the neighbouring seizure event.\n\nI'm trying to understand how these 10 second analysis windows with seizure labels correlate to the specific timing of the seizure events themselves. What am I missing about how the 10 second windows relate to the labeled seizure activity start/stop times?\n\nIf anyone has already covered this in the discussions, please point me there! I'm still catching up on all the great info.\n",
      "votes": 2,
      "replies": [
        {
          "id": 2632655,
          "postDate": "2024-02-02T14:26:35.410Z",
          "content": "<p>These 3 rows overlap in time. The first row is 3 seconds thru 53 seconds with 10second middle at 23 seconds thru 33 seconds. The second row is 4 seconds thru 54 seconds with 10second middle at 24 seconds thru 34 seconds. Therefore the two 10second middle windows overlap by 9 seconds.</p>",
          "rawMarkdown": "These 3 rows overlap in time. The first row is 3 seconds thru 53 seconds with 10second middle at 23 seconds thru 33 seconds. The second row is 4 seconds thru 54 seconds with 10second middle at 24 seconds thru 34 seconds. Therefore the two 10second middle windows overlap by 9 seconds.",
          "votes": 1,
          "replies": [
            {
              "id": 2636280,
              "postDate": "2024-02-05T00:44:46.450Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2636282,
              "postDate": "2024-02-05T00:49:09.380Z",
              "content": "<p>Hi, Chris, you do a great work for this competition, I m learning your notebook----How To Make Spectrogram from EEG, But I can't understand why you select eeg from 'middle = (len(eeg)-10_000)//2, eeg = eeg.iloc[middle:middle+10_000]'. rather than from the eeg.iloc[0:] or select the eeg.iloc[:the end]? Maybe I miss some information? thanks for you reply.</p>",
              "rawMarkdown": "Hi, Chris, you do a great work for this competition, I m learning your notebook----How To Make Spectrogram from EEG, But I can't understand why you select eeg from 'middle = (len(eeg)-10_000)//2, eeg = eeg.iloc[middle:middle+10_000]'. rather than from the eeg.iloc[0:] or select the eeg.iloc[:the end]? Maybe I miss some information? thanks for you reply.",
              "votes": 1
            },
            {
              "id": 2636835,
              "postDate": "2024-02-05T10:23:23.433Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, the 10-second window and the overlap periods are now clear. My question is, where is the signal period that indicates Seizure?<br>\nIn the paper <a href=\"https://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf\" target=\"_blank\">https://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf</a> Figure 1 and Figure 2, the types of  Seizure are explained.  some are continuous and some are discontinuous. <br>\nMy question may come down to this:  Each row indicates that there is a Seizure in the relevant 10-second period, but the actual location is not given. <br>\nIn the original example I referenced, it may be that eeg_id 110 and 111  actually refer to the same Seizure signal if the duration is short.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F69790%2F155405b9063640a4a3e094401d886e4d%2FScreenshot%202024-02-05%20102103.png?generation=1707128521445316&amp;alt=media\" alt=\"Seizure Period illustration\"> </p>",
              "rawMarkdown": "@cdeotte, the 10-second window and the overlap periods are now clear. My question is, where is the signal period that indicates Seizure?\nIn the paper https://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf Figure 1 and Figure 2, the types of  Seizure are explained.  some are continuous and some are discontinuous. \nMy question may come down to this:  Each row indicates that there is a Seizure in the relevant 10-second period, but the actual location is not given. \nIn the original example I referenced, it may be that eeg_id 110 and 111  actually refer to the same Seizure signal if the duration is short.\n![Seizure Period illustration](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F69790%2F155405b9063640a4a3e094401d886e4d%2FScreenshot%202024-02-05%20102103.png?generation=1707128521445316&alt=media) \n",
              "votes": 1
            },
            {
              "id": 2637073,
              "postDate": "2024-02-05T13:39:15.707Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/wangjianxing206\" target=\"_blank\">@wangjianxing206</a> and <a href=\"https://www.kaggle.com/tzurvaich\" target=\"_blank\">@tzurvaich</a> We can try using first and/or last instead of middle. I picked middle because it was a safe choice. We don't know where the true Seizure is. Kaggle has provided a range of windows. I guess that they provided windows before and after the true Seizure, so I take the middle.</p>\n<p>We can train a model using a different location and evaluate CV score and LB score. Also we can pick more than one window per eeg_id. For example, we can use first, middle, and last. And create 3 spectrograms per eeg_id and train with all of them.</p>",
              "rawMarkdown": "Hi @wangjianxing206 and @tzurvaich We can try using first and/or last instead of middle. I picked middle because it was a safe choice. We don't know where the true Seizure is. Kaggle has provided a range of windows. I guess that they provided windows before and after the true Seizure, so I take the middle.\n\nWe can train a model using a different location and evaluate CV score and LB score. Also we can pick more than one window per eeg_id. For example, we can use first, middle, and last. And create 3 spectrograms per eeg_id and train with all of them."
            },
            {
              "id": 2638775,
              "postDate": "2024-02-06T13:28:41.137Z",
              "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, for clarifying what you look at at the centre. The duration and LOCATION would differ for different types of seizures.  I assume there is no way of getting clarification and additional information from the people who have created the dataset. </p>",
              "rawMarkdown": "Thanks, @cdeotte, for clarifying what you look at at the centre. The duration and LOCATION would differ for different types of seizures.  I assume there is no way of getting clarification and additional information from the people who have created the dataset. ",
              "votes": 1
            },
            {
              "id": 2638969,
              "postDate": "2024-02-06T15:43:50.087Z",
              "content": "<p>Reading again the training data description, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\" target=\"_blank\">it says</a> \"<strong>train.csv</strong> Metadata for the train set. The expert annotators <em>reviewed 50 second long EEG samples</em> plus matched spectrograms covering 10 a minute window <em>centered at the same time</em> and <strong>labeled the central 10 seconds.</strong> Many of these samples overlapped and have been consolidated.\"</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I believe that your understanding is correct. In other words:</p>\n<ul>\n<li>The 50 second long EEG samples are reviewed and labeled by experts.</li>\n<li>Many of these 50 second samples overlap and have been consolidated into a single EEG recording.</li>\n<li><em>eeg_label_offset_seconds</em> provides the time offset in seconds from the beginning of this consolidated EEG to the start of the specific 50 second subsample.</li>\n<li>The experts labeled the central 10 seconds of each 50 second subsample.</li>\n</ul>",
              "rawMarkdown": "Reading again the training data description, [it says](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data) \"**train.csv** Metadata for the train set. The expert annotators *reviewed 50 second long EEG samples* plus matched spectrograms covering 10 a minute window *centered at the same time* and **labeled the central 10 seconds.** Many of these samples overlapped and have been consolidated.\"\n\n@cdeotte, I believe that your understanding is correct. In other words:\n- The 50 second long EEG samples are reviewed and labeled by experts.\n- Many of these 50 second samples overlap and have been consolidated into a single EEG recording.\n- *eeg_label_offset_seconds* provides the time offset in seconds from the beginning of this consolidated EEG to the start of the specific 50 second subsample.\n- The experts labeled the central 10 seconds of each 50 second subsample.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2629909,
      "postDate": "2024-02-01T04:03:22.770Z",
      "content": "<p>Obviously we only have EEG data in 50s snippets, not the 600s we have for spectrograms, so we don't have as much information when trying to recreate the spectrograms from the EEG data.<br>\nKaggle Spectrograms have Freq(Hz) between 0 and 20, why do your spectrograms have with levels between 0 and 120? Is this because our EEG data is sampled at 200Hz while the Kaggle Spectrogram is more like 0.5 Hz?</p>",
      "rawMarkdown": "Obviously we only have EEG data in 50s snippets, not the 600s we have for spectrograms, so we don't have as much information when trying to recreate the spectrograms from the EEG data.\nKaggle Spectrograms have Freq(Hz) between 0 and 20, why do your spectrograms have with levels between 0 and 120? Is this because our EEG data is sampled at 200Hz while the Kaggle Spectrogram is more like 0.5 Hz?",
      "votes": 2,
      "replies": [
        {
          "id": 2629927,
          "postDate": "2024-02-01T04:27:09.423Z",
          "content": "<p>My spectrograms are also between 0 and 20Hz. Each unit on the height dimension is about 1/5 Hz, so 128 height is actually 20 Hz.</p>",
          "rawMarkdown": "My spectrograms are also between 0 and 20Hz. Each unit on the height dimension is about 1/5 Hz, so 128 height is actually 20 Hz.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2689709,
      "postDate": "2024-03-10T04:49:57.160Z",
      "content": "<p>halo all, i am a newbie can someone guide me where and how to start</p>",
      "rawMarkdown": "halo all, i am a newbie can someone guide me where and how to start",
      "isDeleted": true
    },
    {
      "id": 2614957,
      "postDate": "2024-01-22T23:05:53.587Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2601478,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-01-14T12:53:50.303000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> from my understanding</p>\n<p><strong>Note</strong> 16 channels out of 19 used and 1 EKG</p>\n<pre><code>pairing = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : , //  part of LL, LP, RP, RR\n    :   //  part of LL, LP, RP, RR\n}\n</code></pre>\n<p>from  <a href=\"https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features\" target=\"_blank\">Notebook - Eegs Pairing Analysis &amp; Features</a></p>\n<p><strong>LL</strong> -&gt;  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F42abf3341243af65c92c8ba2083b9a79%2FScreenshot%202024-01-14%20at%205.49.27PM.png?generation=1705234855564208&amp;alt=media\"><br>\n<strong>LP</strong> -&gt;<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa725bf3a965b4299d0af07f4937216f5%2FScreenshot%202024-01-14%20at%205.49.38PM.png?generation=1705234875852002&amp;alt=media\"><br>\n<strong>RP</strong> -&gt;<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5869dccab21b6cc3ba7b31e849622d5a%2FScreenshot%202024-01-14%20at%205.50.02PM.png?generation=1705234897860474&amp;alt=media\"><br>\n<strong>RR</strong> -&gt; <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa13eff4393d4fc67fbf3a2bc52294553%2FScreenshot%202024-01-14%20at%205.50.12PM.png?generation=1705234918197292&amp;alt=media\"><br>\nfrom  <a href=\"https://www.kaggle.com/code/seshurajup/eegs-10-20-system\" target=\"_blank\">Notebook - EEGS 10–20 system</a></p>\n<p>This video is awesome to understand basics <a href=\"https://www.youtube.com/watch?v=XMizSSOejg0&amp;ab_channel=JeremyMoeller\" target=\"_blank\">Source Youtube</a></p>\n<p><strong>Ignored Middle part channels</strong></p>",
      "votes": 15,
      "replies": [
        {
          "id": 2601504,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-14T13:25:18.537000",
          "content": "<p>In your diagram above, LL is 4 signals (i.e. 4 differences). We need LP to be 1 signal. Do we just average the 4? (Also LP is 4, RP is 4, RR is 4)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2601518,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-14T13:34:21.800000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> As per my understanding,</p>\n<p>calculated spectrograms of each pair of LL channels. then average of these spectrograms as final LL spectrogram<br>\nFp1 -&gt; F7 =&gt; 1st spectrogram<br>\nF7 -&gt; T7 =&gt; 2nd spectrogram <br>\nT7 -&gt; P7 =&gt; 3rd spectrogram <br>\nP7 -&gt; 01 =&gt; 4th spectrogram<br>\nAverage of all 4 spectrograms =&gt; LL spectrogram</p>\n<p><em>Note</em> correct me if am wrong</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2601523,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-14T13:40:05.813000",
              "content": "<p>Oh, that would work. Did you read that somewhere or are you guessing?</p>\n<p>This means we can take the average of the four groups (LL LP RR RP) of 4 signals to create 4 new signals. Then if we apply WaveNet to these 4 newly created eeg signals, we should be able to achieve similar performance as spectrogram only models because WaveNet extracts frequencies from signals.</p>\n<p>(If you prefer, you can crop the middle 10 (or 20 or 50) seconds from both spectrogram and eeg. And ignore the extra information. However more information will create better models).</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2601620,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-14T15:01:02.297000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> From these video from where i assumed</p>\n<p><a href=\"https://www.youtube.com/@eegforanesthesia3954\" target=\"_blank\">Source from this youtube channel</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2614562,
              "author_name": "Luis Pinto",
              "author_url": "",
              "post_date": "2024-01-22T17:32:09.387000",
              "content": "<p>Hey Chris, did you get a chance to test this? Due to memory constraints, I was able to only use LL and RR but performance is the same as your base wavenet notebook</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2614591,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-22T17:55:11.210000",
              "content": "<p>Yes, I did this and posted a bunch of starter notebook. It works well. Check out my notebooks in the code section.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2632308,
          "author_name": "Raki",
          "author_url": "",
          "post_date": "2024-02-02T09:01:55.660000",
          "content": "<p>Thanks! The video helped a lot :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2604075,
      "author_name": "Nishant Singhal",
      "author_url": "",
      "post_date": "2024-01-16T08:36:57.343000",
      "content": "<p>Congratulations, Chris, on your innovative method of generating EEG-derived spectrograms, leading to improved CV and LB scores in the Kaggle competition. This approach opens up new avenues for data utilization.</p>\n<p>I'm curious to know: How do the characteristics and information content of your EEG-derived spectrograms differ from the original Kaggle-provided ones, and how do these differences contribute to the enhanced model performance?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2604470,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-16T12:52:03.553000",
          "content": "<p>Thanks. Great question Nishant about comparing. I have not compared yet but I have observed that training models with both Kaggle's spectrograms and my EEG spectrograms achieves better CV score and LB score (than Kaggle's alone). Therefore there is definitely new information in my EEG spectrograms compared with Kaggle's spectrograms.</p>\n<p>One difference is that Kaggle's are 10 minutes long whereas mine are only 50 seconds. It would be interesting to compare my 50 seconds with Kaggle's middle 50 seconds and see how they differ. I suspect they will differ because:</p>\n<ul>\n<li>Kaggle may create their spectrograms from signals not available to us</li>\n<li>My formula of <code>LL = ( (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1) )/4.</code> may not be the correct way to create the single LL time series.</li>\n<li>My spectrogram hyperparameter settings differ from Kaggle's</li>\n<li>Kaggle may have used denoise before creating spectrograms</li>\n<li>etc etc</li>\n</ul>",
          "votes": 12,
          "replies": [
            {
              "id": 2605136,
              "author_name": "Konstantin Kozlovtsev",
              "author_url": "",
              "post_date": "2024-01-16T21:38:09.123000",
              "content": "<p>Isn't the LL just 1/4 * (Fp1 - O1) then, because all the middle terms are eliminated? So all the specs are defined only by two sensors thus LL and LP are the same (and also RL and RP). Or there is some time lag between difference terms? I've checked in your notebook about specs calculation and it looks like the spectrograms and raw signals are really identical.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2605142,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-16T21:42:43.773000",
              "content": "<p>Hmm… that's a great point. I overlooked that the formula reduces to that. Thus, that formula doesn't look correct. I would think that the spectrogram would take into account all 5 electrodes Fp1, F7, T3, T5, O1 in the Left Temporal Chain.</p>\n<p>I wonder if i create 5 spectrograms for the 5 differences and then take the average of the 5 spectrograms. I wonder if that would be mathematically different than taking 1 spectrogram of Fp1-O1?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2605202,
              "author_name": "Konstantin Kozlovtsev",
              "author_url": "",
              "post_date": "2024-01-16T23:33:57.210000",
              "content": "<p>It should be the same again since STFT is based on the fourier transform (i.e. f(x) = sum a_i * h_i(x) where h_i(x) are basis functions, sin/cos in the terms of the fourier transform) so f(x) + g(x) = sum (a_i + b_i) * h_i(x). However if you do it in the log (dB) space  (actually no, mel spectrogram is just linear transform), then the result can be different. However the simpliest way to use this approach is just make 4 different channels, but I do not think it will bring a big difference scince you will probably train a network so all possible linear calculations can be made in the first layer, so you can just use raw channels and pairwise differences can be (theoretically) trained.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2601534,
      "author_name": "eugene",
      "author_url": "",
      "post_date": "2024-01-14T13:53:08.980000",
      "content": "<p>I'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the EEG has 90 seconds. It's supposed to be the same lenght.</p>\n<p>For example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2601590,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-14T14:27:57.687000",
          "content": "<p>The train data is just two different perspectives. The eeg is a \"zoom in\" to what was happening during 50 seconds and the spectrogram is a \"zoom out\" of what was happening during 10 minutes. They are both centered at the same moment of time (i.e. the timestamp in middle of spectrogram is the same timestamp in middle of eeg).</p>\n<p>Our task is to classify what is happening in the middle 10 seconds.</p>\n<blockquote>\n  <p>The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds.</p>\n</blockquote>",
          "votes": 13,
          "replies": [
            {
              "id": 2611416,
              "author_name": "JamshaidSohail",
              "author_url": "",
              "post_date": "2024-01-20T18:06:56.510000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Thank you for this amazing explanation. Looking forward to learn some more amazing things. 😀</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2612157,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2024-01-21T08:20:58.450000",
      "content": "<p>single fold resnet34d</p>\n<p>kaggle provided:val_loss = 0.6532<br>\nspectrogram_from_eeg:val_loss=0.8286</p>\n<pre><code>\n = np(f),,\n = cv2(np(,axis=),(,))\n = self(image=img)\n....\n</code></pre>\n<p>spectrogram_from_eeg+kaggle provided 2ch:val_loss=0.6695</p>\n<pre><code> = np(f)\n = cv2(,(,))\n\neeg_spec = np(f),,\neeg_spec = np(eeg_spec,axis=)\neeg_spec = cv2(eeg_spec,(,))\n\n = np(,axis=-)\n = self(image=img)\n....\n</code></pre>\n<p>I don't seem to be making good use of the new features …</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2612394,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-21T11:20:20.113000",
          "content": "<p>Soon, i will update my CatBoost starter and EffNet starter to use the new features. Afterward, they will become the new top scoring public notebooks 😀</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 2612461,
          "author_name": "HZM",
          "author_url": "",
          "post_date": "2024-01-21T12:19:25.190000",
          "content": "<p>same issue and wait for Chris the new updated notebook, so appreciate that</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2613268,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-22T00:45:06.753000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> and <a href=\"https://www.kaggle.com/leehann\" target=\"_blank\">@leehann</a> I published examples of how to use my eeg spectrograms in my EfficientNet starter notebook. I also published new eeg spectrograms in my \"how to create spectrogram\" starter notebook. Versions 1-3 of my \"how to create spectrograms\" starter notebook are called \"my old eeg spectrograms\". Version 4 uses a new better formula and is called \"my new eeg spectrograms\". The example code (with its hyperparameters etc) in my EfficientNet starter notebook works for both the new eeg spectrograms and my old eeg spectrograms. Here are the CV scores:</p>\n<ul>\n<li>Using only Kaggle spectrograms achieves 5Fold CV 0.66</li>\n<li>Using only old eeg spectrograms achieves 5Fold CV 0.78</li>\n<li>Using only new eeg spectrograms achieves 5Fold CV 0.63</li>\n</ul>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2613790,
              "author_name": "HZM",
              "author_url": "",
              "post_date": "2024-01-22T09:18:32.697000",
              "content": "<p>thanks so much agian, respect</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2609331,
      "author_name": "HZM",
      "author_url": "",
      "post_date": "2024-01-19T11:51:10.890000",
      "content": "<p>Hi Chris</p>\n<p>so so great job you had done, i am using the data generated from above method, but the cv and lb get worse, may i ask how to use these data, i must missed something. P.S i just used the genreated data to replace official data</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2609346,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-19T12:09:55.130000",
          "content": "<p>This data is already standardized to be values between -1 and 1. So  if you use existing public notebooks with this data, remove the log transform and standardization.</p>\n<p>(The Kaggle spectrograms need preprocess of log transform and standardization whereas my EEG spectrograms do not)</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2609375,
              "author_name": "HZM",
              "author_url": "",
              "post_date": "2024-01-19T12:40:26.187000",
              "content": "<p>ok, thanks so much and let me try it</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2610484,
              "author_name": "HZM",
              "author_url": "",
              "post_date": "2024-01-20T06:47:31.197000",
              "content": "<p>Hi Chris<br>\nSorry to bother you again, i tried your data output without any further process, but it still did not work, may i ask if it is convenient that you can get a small example that how to use your spec data, if i did not miss something, all the code were based on official code right now.</p>\n<blockquote>\n  <p>This data is already standardized to be values between -1 and 1. So  if you use existing public notebooks with this data, remove the log transform and standardization.</p>\n  <p>(The Kaggle spectrograms need preprocess of log transform and standardization whereas my EEG spectrograms do not)</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2610521,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-20T07:19:09.450000",
              "content": "<p>What does \"did not work\" mean? What model are you using and what CV score and LB score are you achieving? There is an example of an MLP in the notebook which achieves CV 1.0 and LB 0.77 (whereas submitting train means achieves CV 1.26 LB 0.97)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2610558,
              "author_name": "HZM",
              "author_url": "",
              "post_date": "2024-01-20T07:48:04.763000",
              "content": "<p>the same model and traing hyper, the cv is alway 1.0+ did not drop during the traing process, but the official spec training is normal</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2611595,
              "author_name": "Yan Teixeira",
              "author_url": "",
              "post_date": "2024-01-20T20:59:03.133000",
              "content": "<p><a href=\"https://www.kaggle.com/leehann\" target=\"_blank\">@leehann</a> I'm experiencing the same</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2612432,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-21T11:56:52.853000",
              "content": "<p>Hi HZM and Yan, soon I will update my CatBoost and EffNet starter notebooks to use the EEG spectrograms.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2612520,
              "author_name": "HZM",
              "author_url": "",
              "post_date": "2024-01-21T12:51:41.040000",
              "content": "<p>thanks so much for your help Chris, you are save</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2614233,
              "author_name": "HZM",
              "author_url": "",
              "post_date": "2024-01-22T14:04:21.467000",
              "content": "<p>Hi Yan, Are u using pytorch? I still had the issue when using updated datasets</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2603599,
      "author_name": "Arslan Gabdulkhakov",
      "author_url": "",
      "post_date": "2024-01-15T23:04:55.577000",
      "content": "<p>to add to the previous comments, spectrograms were computed with some variant of Fourier transform (most likely Wavelet Convolution, or Complex Morelet Wavelet), and then averaged across electrode channels, as suggested by <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> </p>\n<p>To compute your own time-frequency spectrograms (which is a good idea IMHO), you could follow the User Guide of MNE package <a href=\"https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html\" target=\"_blank\">https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 2603600,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-15T23:12:07.470000",
          "content": "<p><a href=\"https://www.kaggle.com/iworeushankaonce\" target=\"_blank\">@iworeushankaonce</a> </p>\n<blockquote>\n  <p>to add to the previous comments, spectrograms were computed with some variant of Fourier transform (most likely Wavelet Convolution, or Complex Morelet Wavelet), and then averaged across electrode channels, as suggested by <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> </p>\n  <ul>\n  <li>Its my assumption need to conform from experts in this area</li>\n  </ul>\n</blockquote>\n<hr>\n<blockquote>\n  <p>To compute your own time-frequency spectrograms (which is a good idea IMHO), you could follow docudocumentat of MNE package <a href=\"https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html\" target=\"_blank\">https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html</a></p>\n  <ul>\n  <li>spectrograms are of sequence len 10mins long, where egg sequences not having complete 10mins long sequences.</li>\n  </ul>\n</blockquote>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2612552,
      "author_name": "Haru",
      "author_url": "",
      "post_date": "2024-01-21T13:18:37.167000",
      "content": "<p>Thank you for your great post again !<br>\nI have a question about below.</p>\n<p>mel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, <br>\n              n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)</p>\n<p>How did you decide the parameter of (hop_length=len(x)//256,  n_fft=1024, n_mels=128, win_length=128) ?<br>\nThank you first.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2612609,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-21T13:49:30.620000",
          "content": "<p>This is explained in the notebook. <code>hop_length</code> will determine the width of our image. I wanted an image of width 256, so i chose <code>hop_length=len(x)//256</code> where <code>x</code> is the input waveform. The variable <code>n_mels</code> will determine the height of the image. I wanted 128. I picked 256 and 128 because I wanted it to be similar to Kaggle's 300x100 and for image models it is best to have images with dimensions that have multiples of 32.</p>\n<p>The variables <code>n_fft</code> and <code>win_lenth</code> control the sliding window size of the Fourier transform. It is my understanding that a sliding window moves over the original waveform and computes the image in pieces. Hence these variables control the image resolution. I tried a few common ones until the image looked good to my eye. I think <code>win_length</code> is the length of time window and hence affects the image horizontal resolution. And <code>n_fft</code> affects vertical image resolution by deciding the frequency resolution when creating a DFT (discrete fourier transform)</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2613713,
              "author_name": "Haru",
              "author_url": "",
              "post_date": "2024-01-22T08:36:24.410000",
              "content": "<p>Thank you for quick replay and explanation !</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2609539,
      "author_name": "Joser",
      "author_url": "",
      "post_date": "2024-01-19T14:45:21.553000",
      "content": "<p>Can we combine this created Spectrograms with  Kaggle's spectrograms,and put them together for training?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2609554,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-19T15:09:09.923000",
          "content": "<p>Yes, we can do whatever we want. The best way to see if it improves CV score and LB score is to try it.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2632593,
      "author_name": "Zuuk",
      "author_url": "",
      "post_date": "2024-02-02T13:40:13.190000",
      "content": "<p>I'm seeking help interpreting the seizure activity labels in the train.csv file. Some of the seizure events seem shorter than the 50 second test windows.<br>\n10-second durations are used as well including in the discussion and in the example_figures directory. <br>\nThe eeg_label_offset_seconds in train.csv range from 1-6 seconds for some rows with identical eeg_sub_id, i.e. the same eeg signal.</p>\n<p>For example, these rows from train.csv:<br>\nEeg_id,eeg_sub_id, eeg_label_offset_seconds, seizure_vote<br>\n110, 3132287169, 3, Seizure<br>\n111, 3132287169, 4, Seizure<br>\n112, 3132287169, 5, Seizure<br>\nDo these 3 rows indicate 3 separate seizure events in the same eeg ? Taking a 10 seconds window will cover, in these cases, the neighbouring seizure event.</p>\n<p>I'm trying to understand how these 10 second analysis windows with seizure labels correlate to the specific timing of the seizure events themselves. What am I missing about how the 10 second windows relate to the labeled seizure activity start/stop times?</p>\n<p>If anyone has already covered this in the discussions, please point me there! I'm still catching up on all the great info.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2632655,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-02-02T14:26:35.410000",
          "content": "<p>These 3 rows overlap in time. The first row is 3 seconds thru 53 seconds with 10second middle at 23 seconds thru 33 seconds. The second row is 4 seconds thru 54 seconds with 10second middle at 24 seconds thru 34 seconds. Therefore the two 10second middle windows overlap by 9 seconds.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2636280,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-05T00:44:46.450000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2636282,
              "author_name": "wangjianxing206",
              "author_url": "",
              "post_date": "2024-02-05T00:49:09.380000",
              "content": "<p>Hi, Chris, you do a great work for this competition, I m learning your notebook----How To Make Spectrogram from EEG, But I can't understand why you select eeg from 'middle = (len(eeg)-10_000)//2, eeg = eeg.iloc[middle:middle+10_000]'. rather than from the eeg.iloc[0:] or select the eeg.iloc[:the end]? Maybe I miss some information? thanks for you reply.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2636835,
              "author_name": "Zuuk",
              "author_url": "",
              "post_date": "2024-02-05T10:23:23.433000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, the 10-second window and the overlap periods are now clear. My question is, where is the signal period that indicates Seizure?<br>\nIn the paper <a href=\"https://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf\" target=\"_blank\">https://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf</a> Figure 1 and Figure 2, the types of  Seizure are explained.  some are continuous and some are discontinuous. <br>\nMy question may come down to this:  Each row indicates that there is a Seizure in the relevant 10-second period, but the actual location is not given. <br>\nIn the original example I referenced, it may be that eeg_id 110 and 111  actually refer to the same Seizure signal if the duration is short.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F69790%2F155405b9063640a4a3e094401d886e4d%2FScreenshot%202024-02-05%20102103.png?generation=1707128521445316&amp;alt=media\" alt=\"Seizure Period illustration\"> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2637073,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-02-05T13:39:15.707000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/wangjianxing206\" target=\"_blank\">@wangjianxing206</a> and <a href=\"https://www.kaggle.com/tzurvaich\" target=\"_blank\">@tzurvaich</a> We can try using first and/or last instead of middle. I picked middle because it was a safe choice. We don't know where the true Seizure is. Kaggle has provided a range of windows. I guess that they provided windows before and after the true Seizure, so I take the middle.</p>\n<p>We can train a model using a different location and evaluate CV score and LB score. Also we can pick more than one window per eeg_id. For example, we can use first, middle, and last. And create 3 spectrograms per eeg_id and train with all of them.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2638775,
              "author_name": "Zuuk",
              "author_url": "",
              "post_date": "2024-02-06T13:28:41.137000",
              "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, for clarifying what you look at at the centre. The duration and LOCATION would differ for different types of seizures.  I assume there is no way of getting clarification and additional information from the people who have created the dataset. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2638969,
              "author_name": "Zuuk",
              "author_url": "",
              "post_date": "2024-02-06T15:43:50.087000",
              "content": "<p>Reading again the training data description, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/data\" target=\"_blank\">it says</a> \"<strong>train.csv</strong> Metadata for the train set. The expert annotators <em>reviewed 50 second long EEG samples</em> plus matched spectrograms covering 10 a minute window <em>centered at the same time</em> and <strong>labeled the central 10 seconds.</strong> Many of these samples overlapped and have been consolidated.\"</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I believe that your understanding is correct. In other words:</p>\n<ul>\n<li>The 50 second long EEG samples are reviewed and labeled by experts.</li>\n<li>Many of these 50 second samples overlap and have been consolidated into a single EEG recording.</li>\n<li><em>eeg_label_offset_seconds</em> provides the time offset in seconds from the beginning of this consolidated EEG to the start of the specific 50 second subsample.</li>\n<li>The experts labeled the central 10 seconds of each 50 second subsample.</li>\n</ul>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2629909,
      "author_name": "Sarah Hedberg",
      "author_url": "",
      "post_date": "2024-02-01T04:03:22.770000",
      "content": "<p>Obviously we only have EEG data in 50s snippets, not the 600s we have for spectrograms, so we don't have as much information when trying to recreate the spectrograms from the EEG data.<br>\nKaggle Spectrograms have Freq(Hz) between 0 and 20, why do your spectrograms have with levels between 0 and 120? Is this because our EEG data is sampled at 200Hz while the Kaggle Spectrogram is more like 0.5 Hz?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2629927,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-02-01T04:27:09.423000",
          "content": "<p>My spectrograms are also between 0 and 20Hz. Each unit on the height dimension is about 1/5 Hz, so 128 height is actually 20 Hz.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2689709,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-10T04:49:57.160000",
      "content": "<p>halo all, i am a newbie can someone guide me where and how to start</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2614957,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-22T23:05:53.587000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2601428": "# Original Discussion\nMy original post had title \"How can we create spectrograms from eeg?\". The post asked; Does anyone know which 4 time series signals are used to create the 4 spectrograms? We are given 20 eeg time series signals for each time window. How can we combine these 20 signals to make 4 signals which are used to create spectrograms?\n\nI realize that we won't have 10 minute window (for the newly create 4 series), so we can't recreate the full spectrograms from 50 second windows, but none-the-less, I'm curious how to combine the 20 signals into 4 signals.\n\n# New Discussion\nAfter asking this question many Kagglers provided valuable information and I was able to create spectrograms from eegs. And I trained a model using only these new spectrograms and it performed well. So I think we are close to producing the spectrograms correctly. If anyone has feedback or suggestions, please comment below\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2024/montage2.png)\n\n# Starter Notebook\nUsing the information below, I published a starter notebook to create Spectrograms from EEG [here][1]\n\n# UPDATE\nExciting news! I think the Magic Formula has been discovered! Discussion post [here][2]\n\n# Boost CV and LB Score\nI have verified that using only these newly created EEG spectrograms (and not Kaggle's spectrograms), we can achieve a CV and LB score better than many public notebooks (including submitting train means)\n\n# Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrograms-from-eeg\n[2]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760",
    "2601478": "@cdeotte from my understanding\n\n**Note** 16 channels out of 19 used and 1 EKG\n```python\npairing = {\n    \"Fp1\": \"F7\",\n    \"F7\": \"T3\",\n    \"T3\": \"T5\",\n    \"T5\": \"O1\",\n    \"Fp2\": \"F8\",\n    \"F8\": \"T4\",\n    \"T4\": \"T6\",\n    \"T6\": \"O2\",\n    \"Fp1\": \"F3\",\n    \"F3\": \"C3\",\n    \"C3\": \"P3\",\n    \"P3\": \"O1\",\n    \"Fp2\": \"F4\",\n    \"F4\": \"C4\",\n    \"C4\": \"P4\",\n    \"P4\": \"O2\",\n    \"Fz\": \"Cz\", // not part of LL, LP, RP, RR\n    \"Cz\": \"Pz\"  // not part of LL, LP, RP, RR\n}\n```\nfrom  [Notebook - Eegs Pairing Analysis & Features](https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features)\n\n**LL** ->  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F42abf3341243af65c92c8ba2083b9a79%2FScreenshot%202024-01-14%20at%205.49.27PM.png?generation=1705234855564208&alt=media)\n**LP** ->\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa725bf3a965b4299d0af07f4937216f5%2FScreenshot%202024-01-14%20at%205.49.38PM.png?generation=1705234875852002&alt=media)\n**RP** ->\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5869dccab21b6cc3ba7b31e849622d5a%2FScreenshot%202024-01-14%20at%205.50.02PM.png?generation=1705234897860474&alt=media)\n**RR** -> \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa13eff4393d4fc67fbf3a2bc52294553%2FScreenshot%202024-01-14%20at%205.50.12PM.png?generation=1705234918197292&alt=media)\nfrom  [Notebook - EEGS 10–20 system](https://www.kaggle.com/code/seshurajup/eegs-10-20-system)\n\nThis video is awesome to understand basics [Source Youtube](https://www.youtube.com/watch?v=XMizSSOejg0&ab_channel=JeremyMoeller)\n\n**Ignored Middle part channels**",
    "2604075": "Congratulations, Chris, on your innovative method of generating EEG-derived spectrograms, leading to improved CV and LB scores in the Kaggle competition. This approach opens up new avenues for data utilization.\n\nI'm curious to know: How do the characteristics and information content of your EEG-derived spectrograms differ from the original Kaggle-provided ones, and how do these differences contribute to the enhanced model performance?",
    "2601534": "I'm also trying to figure out how the spectrograms were constructed. I can't understand why the spectrogram has 10 minutes of data and the EEG has 90 seconds. It's supposed to be the same lenght.\n\nFor example: eeg_id 1628180742 has 18000 rows it mean there 90 second of data if freq is 200hz",
    "2612157": "single fold resnet34d\n\nkaggle provided:val_loss = 0.6532\nspectrogram_from_eeg:val_loss=0.8286\n```\n#spectrogram_from_eeg\nimg = np.load(f\"spec_from_eeg/{eeg_id}.npy\")#128,256,4\nimg = cv2.resize(np.concatenate(img,axis=1).T,(512,512))\nimg = self.transform(image=img)[\"image\"]\n....\n```\n\nspectrogram_from_eeg+kaggle provided 2ch:val_loss=0.6695\n```\nimg = np.load(f\"kaggle_spec/{label_id}.npy\")\nimg = cv2.resize(img,(512,512))\n        \neeg_spec = np.load(f\"spec_from_eeg/{eeg_id}.npy\")#128,256,4\neeg_spec = np.concatenate(eeg_spec,axis=1).T\neeg_spec = cv2.resize(eeg_spec,(512,512))\n        \nimg = np.stack([img,eeg_spec],axis=-1)\nimg = self.transform(image=img)[\"image\"]\n....\n```\n\nI don't seem to be making good use of the new features ...",
    "2609331": "Hi Chris\n\nso so great job you had done, i am using the data generated from above method, but the cv and lb get worse, may i ask how to use these data, i must missed something. P.S i just used the genreated data to replace official data",
    "2603599": "to add to the previous comments, spectrograms were computed with some variant of Fourier transform (most likely Wavelet Convolution, or Complex Morelet Wavelet), and then averaged across electrode channels, as suggested by @seshurajup \n\nTo compute your own time-frequency spectrograms (which is a good idea IMHO), you could follow the User Guide of MNE package https://mne.tools/stable/auto_tutorials/time-freq/20_sensors_time_frequency.html",
    "2612552": "Thank you for your great post again !\nI have a question about below.\n\nmel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, \n              n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)\n\nHow did you decide the parameter of (hop_length=len(x)//256,  n_fft=1024, n_mels=128, win_length=128) ?\nThank you first.",
    "2609539": "Can we combine this created Spectrograms with  Kaggle's spectrograms,and put them together for training?",
    "2632593": "I'm seeking help interpreting the seizure activity labels in the train.csv file. Some of the seizure events seem shorter than the 50 second test windows.\n10-second durations are used as well including in the discussion and in the example_figures directory. \nThe eeg_label_offset_seconds in train.csv range from 1-6 seconds for some rows with identical eeg_sub_id, i.e. the same eeg signal.\n\n\nFor example, these rows from train.csv:\nEeg_id,eeg_sub_id, eeg_label_offset_seconds, seizure_vote\n110, 3132287169, 3, Seizure\n111, 3132287169, 4, Seizure\n112, 3132287169, 5, Seizure\nDo these 3 rows indicate 3 separate seizure events in the same eeg ? Taking a 10 seconds window will cover, in these cases, the neighbouring seizure event.\n\nI'm trying to understand how these 10 second analysis windows with seizure labels correlate to the specific timing of the seizure events themselves. What am I missing about how the 10 second windows relate to the labeled seizure activity start/stop times?\n\nIf anyone has already covered this in the discussions, please point me there! I'm still catching up on all the great info.\n",
    "2629909": "Obviously we only have EEG data in 50s snippets, not the 600s we have for spectrograms, so we don't have as much information when trying to recreate the spectrograms from the EEG data.\nKaggle Spectrograms have Freq(Hz) between 0 and 20, why do your spectrograms have with levels between 0 and 120? Is this because our EEG data is sampled at 200Hz while the Kaggle Spectrogram is more like 0.5 Hz?",
    "2689709": "halo all, i am a newbie can someone guide me where and how to start",
    "2614957": ""
  }
}