{
  "id": 469760,
  "title": "Magic Formula to Convert EEG to Spectrograms!",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/469760",
  "author_name": "Chris Deotte",
  "post_date": "2024-01-21T23:13:37.717000",
  "votes": 233,
  "comment_count": 37,
  "views": 0,
  "content": "<h1>Magic Formula</h1>\n<p>Exciting news everyone. I think we found the <strong>magic formula</strong> to convert EEG waveforms into Kaggle Spectrograms.</p>\n<pre><code>LL Spec = ( ( - F7) + spec( - T3) + spec( - T5) + spec( - O1) )/4\nLP Spec = ( ( - F3) + spec( - C3) + spec( - P3) + spec( - O1) )/4\nRP Spec = ( ( - F4) + spec( - C4) + spec( - P4) + spec( - O2) )/4\nRL Spec = ( ( - F8) + spec( - T4) + spec( - T6) + spec( - O2) )/4\n</code></pre>\n<p>The trick is that we need to average four spectrograms to produce the LL (i.e. Left Temporal Chain) single spectrogram instead of combining Fp1, F7, T3, T5, O1 into a single waveform and computing a single spectrogram. If we combine waveforms first with <code>W = (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1)</code> then this reduces to <code>W = Fp1 - O1</code> and it doesn't use information from <code>F7 , T3, T5</code>. But the above <strong>magic formula</strong> does use information from all EEG waveforms.</p>\n<h1>History</h1>\n<p>One week ago, we started a discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467877\" target=\"_blank\">here</a> to determine how the Kaggle spectrograms were made. Kaggle users <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> and others have been active and helpful in this conversation. I tried various ways to combine the 19 EEG waveforms into the 4 spectrograms <code>LL, LP, RL, RP</code>. I have finally created spectrograms that achieve a better CV and LB score than the Kaggle spectrograms!</p>\n<h1>Experiments</h1>\n<p>We can review the following experiments in my EfficientNet starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a></p>\n<table>\n<thead>\n<tr>\n<th>Spectrogram</th>\n<th>EffNet 5Fold CV</th>\n<th>EffNet LB</th>\n<th>Notebook version</th>\n<th>Model</th>\n<th>Loss</th>\n<th>Epochs</th>\n<th>Data Augmention</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Kaggle Spectrograms</td>\n<td>0.73</td>\n<td>0.57</td>\n<td>ver 1</td>\n<td>EffNetB2</td>\n<td>CE</td>\n<td>3</td>\n<td>none</td>\n</tr>\n<tr>\n<td>Kaggle Spectrograms</td>\n<td>0.66</td>\n<td>???</td>\n<td>ver 4</td>\n<td>EffNetB0</td>\n<td>KL-Div</td>\n<td>4</td>\n<td>none</td>\n</tr>\n<tr>\n<td>EEG Spectrograms</td>\n<td>0.63</td>\n<td>???</td>\n<td>ver 3</td>\n<td>EffNetB0</td>\n<td>KL-Div</td>\n<td>4</td>\n<td>none</td>\n</tr>\n<tr>\n<td>Both Spectrograms</td>\n<td>0.59</td>\n<td>0.44</td>\n<td>ver 5</td>\n<td>EffNetB0</td>\n<td>KL-Div</td>\n<td>4</td>\n<td>none</td>\n</tr>\n</tbody>\n</table>\n<p>If we compare row 2 and row 3 above, we see that training EfficientNet B0 on only the Magic formula EEG spectrograms achieves a better CV score and LB score than training EfficientNet B0 on only the Kaggle spectrograms.</p>\n<p>Furthermore, training a single model with both spectrograms achieves the best CV score and LB score. I think this confirms that we have a powerful formula for converting 19 EEG waveforms into 4 spectrograms!</p>\n<h1>How To Use EEG Spectrograms</h1>\n<p>Examples of how to use new EEG spectrograms to boost CV score and LB score are published in recent versions of my EfficientNet starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a> and CatBoost starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">here</a>.</p>\n<h1>Kaggle Dataset</h1>\n<p>The new EEG spectrograms using the <strong>Magic Formula</strong> have been uploaded to a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a>. We can attach this Kaggle dataset to our future Kaggle notebooks to boost our CV scores and LB scores! </p>\n<p>If working offline, we can download one 7.5GB Python dictionary file that contains all 17089 <strong>EEG spectrograms</strong> with the command: </p>\n<pre><code>kaggle datasets download -f eeg_specs.npy cdeotte/brain-eeg-spectrograms\n</code></pre>\n<p>And we can download all of Kaggle spectrograms in one 2.5GB Python dictionary that contains all 11138 Kaggle spectrograms with command:</p>\n<pre><code>kaggle datasets download cdeotte/brain-spectrograms\n</code></pre>\n<h1>Future Work</h1>\n<p>Since these new spectrograms help improve our model, this means we can try exploring more formulas and alternative ways to create images from EEG waveforms. For example <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> is exploring using different types of Fourier transform in his notebook <a href=\"https://www.kaggle.com/code/abebe9849/make-cwt-from-eeg-compare-3-types-of-images\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": 2613222,
      "postDate": "2024-01-21T23:13:37.717Z",
      "content": "<h1>Magic Formula</h1>\n<p>Exciting news everyone. I think we found the <strong>magic formula</strong> to convert EEG waveforms into Kaggle Spectrograms.</p>\n<pre><code>LL Spec = ( ( - F7) + spec( - T3) + spec( - T5) + spec( - O1) )/4\nLP Spec = ( ( - F3) + spec( - C3) + spec( - P3) + spec( - O1) )/4\nRP Spec = ( ( - F4) + spec( - C4) + spec( - P4) + spec( - O2) )/4\nRL Spec = ( ( - F8) + spec( - T4) + spec( - T6) + spec( - O2) )/4\n</code></pre>\n<p>The trick is that we need to average four spectrograms to produce the LL (i.e. Left Temporal Chain) single spectrogram instead of combining Fp1, F7, T3, T5, O1 into a single waveform and computing a single spectrogram. If we combine waveforms first with <code>W = (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1)</code> then this reduces to <code>W = Fp1 - O1</code> and it doesn't use information from <code>F7 , T3, T5</code>. But the above <strong>magic formula</strong> does use information from all EEG waveforms.</p>\n<h1>History</h1>\n<p>One week ago, we started a discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467877\" target=\"_blank\">here</a> to determine how the Kaggle spectrograms were made. Kaggle users <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> and others have been active and helpful in this conversation. I tried various ways to combine the 19 EEG waveforms into the 4 spectrograms <code>LL, LP, RL, RP</code>. I have finally created spectrograms that achieve a better CV and LB score than the Kaggle spectrograms!</p>\n<h1>Experiments</h1>\n<p>We can review the following experiments in my EfficientNet starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a></p>\n<table>\n<thead>\n<tr>\n<th>Spectrogram</th>\n<th>EffNet 5Fold CV</th>\n<th>EffNet LB</th>\n<th>Notebook version</th>\n<th>Model</th>\n<th>Loss</th>\n<th>Epochs</th>\n<th>Data Augmention</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Kaggle Spectrograms</td>\n<td>0.73</td>\n<td>0.57</td>\n<td>ver 1</td>\n<td>EffNetB2</td>\n<td>CE</td>\n<td>3</td>\n<td>none</td>\n</tr>\n<tr>\n<td>Kaggle Spectrograms</td>\n<td>0.66</td>\n<td>???</td>\n<td>ver 4</td>\n<td>EffNetB0</td>\n<td>KL-Div</td>\n<td>4</td>\n<td>none</td>\n</tr>\n<tr>\n<td>EEG Spectrograms</td>\n<td>0.63</td>\n<td>???</td>\n<td>ver 3</td>\n<td>EffNetB0</td>\n<td>KL-Div</td>\n<td>4</td>\n<td>none</td>\n</tr>\n<tr>\n<td>Both Spectrograms</td>\n<td>0.59</td>\n<td>0.44</td>\n<td>ver 5</td>\n<td>EffNetB0</td>\n<td>KL-Div</td>\n<td>4</td>\n<td>none</td>\n</tr>\n</tbody>\n</table>\n<p>If we compare row 2 and row 3 above, we see that training EfficientNet B0 on only the Magic formula EEG spectrograms achieves a better CV score and LB score than training EfficientNet B0 on only the Kaggle spectrograms.</p>\n<p>Furthermore, training a single model with both spectrograms achieves the best CV score and LB score. I think this confirms that we have a powerful formula for converting 19 EEG waveforms into 4 spectrograms!</p>\n<h1>How To Use EEG Spectrograms</h1>\n<p>Examples of how to use new EEG spectrograms to boost CV score and LB score are published in recent versions of my EfficientNet starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a> and CatBoost starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">here</a>.</p>\n<h1>Kaggle Dataset</h1>\n<p>The new EEG spectrograms using the <strong>Magic Formula</strong> have been uploaded to a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a>. We can attach this Kaggle dataset to our future Kaggle notebooks to boost our CV scores and LB scores! </p>\n<p>If working offline, we can download one 7.5GB Python dictionary file that contains all 17089 <strong>EEG spectrograms</strong> with the command: </p>\n<pre><code>kaggle datasets download -f eeg_specs.npy cdeotte/brain-eeg-spectrograms\n</code></pre>\n<p>And we can download all of Kaggle spectrograms in one 2.5GB Python dictionary that contains all 11138 Kaggle spectrograms with command:</p>\n<pre><code>kaggle datasets download cdeotte/brain-spectrograms\n</code></pre>\n<h1>Future Work</h1>\n<p>Since these new spectrograms help improve our model, this means we can try exploring more formulas and alternative ways to create images from EEG waveforms. For example <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> is exploring using different types of Fourier transform in his notebook <a href=\"https://www.kaggle.com/code/abebe9849/make-cwt-from-eeg-compare-3-types-of-images\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "# Magic Formula\n\nExciting news everyone. I think we found the **magic formula** to convert EEG waveforms into Kaggle Spectrograms.\n\n    LL Spec = ( spec(Fp1 - F7) + spec(F7 - T3) + spec(T3 - T5) + spec(T5 - O1) )/4\n    LP Spec = ( spec(Fp1 - F3) + spec(F3 - C3) + spec(C3 - P3) + spec(P3 - O1) )/4\n    RP Spec = ( spec(Fp2 - F4) + spec(F4 - C4) + spec(C4 - P4) + spec(P4 - O2) )/4\n    RL Spec = ( spec(Fp2 - F8) + spec(F8 - T4) + spec(T4 - T6) + spec(T6 - O2) )/4\n\nThe trick is that we need to average four spectrograms to produce the LL (i.e. Left Temporal Chain) single spectrogram instead of combining Fp1, F7, T3, T5, O1 into a single waveform and computing a single spectrogram. If we combine waveforms first with `W = (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1)` then this reduces to `W = Fp1 - O1` and it doesn't use information from `F7 , T3, T5`. But the above **magic formula** does use information from all EEG waveforms.\n\n# History\n\nOne week ago, we started a discussion [here][1] to determine how the Kaggle spectrograms were made. Kaggle users @seshurajup and others have been active and helpful in this conversation. I tried various ways to combine the 19 EEG waveforms into the 4 spectrograms `LL, LP, RL, RP`. I have finally created spectrograms that achieve a better CV and LB score than the Kaggle spectrograms!\n\n# Experiments\nWe can review the following experiments in my EfficientNet starter notebook [here][3]\n\n| Spectrogram | EffNet 5Fold CV | EffNet LB | Notebook version | Model | Loss | Epochs | Data Augmention |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| Kaggle Spectrograms | 0.73 | 0.57 | ver 1 | EffNetB2 | CE | 3 | none |\n| Kaggle Spectrograms | 0.66 | ??? | ver 4 | EffNetB0 | KL-Div | 4 | none |\n| EEG Spectrograms | 0.63 | ??? | ver 3 | EffNetB0 | KL-Div | 4 | none |\n| Both Spectrograms | 0.59 | 0.44 | ver 5| EffNetB0 | KL-Div | 4 | none | \n\n\nIf we compare row 2 and row 3 above, we see that training EfficientNet B0 on only the Magic formula EEG spectrograms achieves a better CV score and LB score than training EfficientNet B0 on only the Kaggle spectrograms.\n\nFurthermore, training a single model with both spectrograms achieves the best CV score and LB score. I think this confirms that we have a powerful formula for converting 19 EEG waveforms into 4 spectrograms!\n\n# How To Use EEG Spectrograms\nExamples of how to use new EEG spectrograms to boost CV score and LB score are published in recent versions of my EfficientNet starter notebook [here][3] and CatBoost starter notebook [here][4].\n\n# Kaggle Dataset\nThe new EEG spectrograms using the **Magic Formula** have been uploaded to a Kaggle dataset [here][5]. We can attach this Kaggle dataset to our future Kaggle notebooks to boost our CV scores and LB scores! \n\nIf working offline, we can download one 7.5GB Python dictionary file that contains all 17089 **EEG spectrograms** with the command: \n\n    kaggle datasets download -f eeg_specs.npy cdeotte/brain-eeg-spectrograms\n\nAnd we can download all of Kaggle spectrograms in one 2.5GB Python dictionary that contains all 11138 Kaggle spectrograms with command:\n\n    kaggle datasets download cdeotte/brain-spectrograms\n\n# Future Work\nSince these new spectrograms help improve our model, this means we can try exploring more formulas and alternative ways to create images from EEG waveforms. For example @abebe9849 is exploring using different types of Fourier transform in his notebook [here][2].\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467877\n[2]: https://www.kaggle.com/code/abebe9849/make-cwt-from-eeg-compare-3-types-of-images\n[3]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[4]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[5]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms",
      "votes": 231
    },
    {
      "id": 2632312,
      "postDate": "2024-02-02T09:09:53.267Z",
      "content": "<p>I found a visualization in a video of how the anterior-posterior bipolar montage is created (the names of a few electrode locations are different though).<br>\n<a href=\"https://www.youtube.com/clip/UgkxqTZZRvcuSHO4la1-qGMCpGai_VF6h-QD\" target=\"_blank\">https://www.youtube.com/clip/UgkxqTZZRvcuSHO4la1-qGMCpGai_VF6h-QD</a></p>\n<p>Disclaimer: I'm not 100% certain it's the same as the magic formula.</p>",
      "rawMarkdown": "I found a visualization in a video of how the anterior-posterior bipolar montage is created (the names of a few electrode locations are different though).\nhttps://www.youtube.com/clip/UgkxqTZZRvcuSHO4la1-qGMCpGai_VF6h-QD\n\nDisclaimer: I'm not 100% certain it's the same as the magic formula.",
      "votes": 5
    },
    {
      "id": 2618146,
      "postDate": "2024-01-24T15:31:01.580Z",
      "content": "<p>Spectrum is a linear transform. Sum of the spec should be equal to the spec of the sum which should reduce to spec(Fp1 - O1) similarly. </p>",
      "rawMarkdown": "Spectrum is a linear transform. Sum of the spec should be equal to the spec of the sum which should reduce to spec(Fp1 - O1) similarly. ",
      "votes": 3,
      "replies": [
        {
          "id": 2618159,
          "postDate": "2024-01-24T15:42:02.533Z",
          "content": "<p>I don't think so. The spectrogram ignores the phase of sine waves. If i have two 60Hz sine waves that are 180 degrees out of phase, then if i add up the raw waveforms, the result is zero everywhere. However if i take their 2 spectrograms and add them up the result is <strong>not</strong> zero everywhere.</p>\n<p>(In other words, raw waveforms are positive and negative. Whereas spectrograms are non-negative. Adding spectrograms together does not cancel out the spectrograms).</p>",
          "rawMarkdown": "I don't think so. The spectrogram ignores the phase of sine waves. If i have two 60Hz sine waves that are 180 degrees out of phase, then if i add up the raw waveforms, the result is zero everywhere. However if i take their 2 spectrograms and add them up the result is **not** zero everywhere.\n\n(In other words, raw waveforms are positive and negative. Whereas spectrograms are non-negative. Adding spectrograms together does not cancel out the spectrograms).",
          "votes": 7
        }
      ]
    },
    {
      "id": 2613568,
      "postDate": "2024-01-22T06:19:56.927Z",
      "content": "<p>Good idea, in fact I used it to train the model on a data set you posted for making spectrograms, but the training loss deviated greatly from the validation loss.</p>",
      "rawMarkdown": "Good idea, in fact I used it to train the model on a data set you posted for making spectrograms, but the training loss deviated greatly from the validation loss.",
      "votes": 3,
      "replies": [
        {
          "id": 2615026,
          "postDate": "2024-01-23T00:32:11.163Z",
          "content": "<p>same issue, pytoch did not work as good as tensorflow</p>",
          "rawMarkdown": "same issue, pytoch did not work as good as tensorflow",
          "votes": 2
        }
      ]
    },
    {
      "id": 2613249,
      "postDate": "2024-01-22T00:02:28.667Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for sharing</p>\n<blockquote>\n  <h3>LOG TRANSFORM SPECTROGRAM</h3>\n  <p>img = np.clip(img,np.exp(-4),np.exp(8)) ==&gt; for Spectrogram </p>\n</blockquote>\n<hr>\n<blockquote>\n  <h3>LOG TRANSFORM</h3>\n  <p>width = (mel_spec.shape[1]//32)*32<br>\n  mel_spec_db = librosa.power_to_db(mel_spec, ref=np.max).astype(np.float32)[:,:width] ==&gt; for EGG -&gt; Spectrogram</p>\n</blockquote>\n<p>How to choice transformations? </p>",
      "rawMarkdown": "@cdeotte Thank you for sharing\n\n> ### LOG TRANSFORM SPECTROGRAM\n img = np.clip(img,np.exp(-4),np.exp(8)) ==> for Spectrogram \n\n---\n> ### LOG TRANSFORM\nwidth = (mel_spec.shape[1]//32)*32\nmel_spec_db = librosa.power_to_db(mel_spec, ref=np.max).astype(np.float32)[:,:width] ==> for EGG -> Spectrogram\n\nHow to choice transformations? ",
      "votes": 4,
      "replies": [
        {
          "id": 2614723,
          "postDate": "2024-01-22T19:10:47.443Z",
          "content": "<p>Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.</p>\n<p>The most common transformation is <code>standardization</code> which is <code>new data = (old data - mean) / std</code>. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a <code>log transform</code> to remove the skew. So first we <code>log transform</code> and next we <code>standardize</code>. </p>\n<p>Log transforms can only handle positive numbers (i.e <code>x&gt;0</code>), so we must shift, clip, and/or flip all the data to be <code>data &gt; 0</code> before we perform <code>log transform</code>.</p>",
          "rawMarkdown": "Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.\n\nThe most common transformation is `standardization` which is `new data = (old data - mean) / std`. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a `log transform` to remove the skew. So first we `log transform` and next we `standardize`. \n\nLog transforms can only handle positive numbers (i.e `x>0`), so we must shift, clip, and/or flip all the data to be `data > 0` before we perform `log transform`.",
          "votes": 17
        },
        {
          "id": 2614725,
          "postDate": "2024-01-22T19:13:28.727Z",
          "content": "<p>Another trick is using <code>Gauss Rank Transform</code> (shown <a href=\"https://medium.com/rapids-ai/gauss-rank-transformation-is-100x-faster-with-rapids-and-cupy-7c947e3397da\" target=\"_blank\">here</a>). Or using a <code>Quantile Transform</code> <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.QuantileTransformer.html\" target=\"_blank\">here</a>. Both can transform any distribution into a nearly perfect Gaussian distribution afterward. </p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/dog3.png\"></p>",
          "rawMarkdown": "Another trick is using `Gauss Rank Transform` (shown [here][1]). Or using a `Quantile Transform` [here][2]. Both can transform any distribution into a nearly perfect Gaussian distribution afterward. \n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/dog3.png)\n\n[1]: https://medium.com/rapids-ai/gauss-rank-transformation-is-100x-faster-with-rapids-and-cupy-7c947e3397da\n[2]: https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.QuantileTransformer.html",
          "votes": 13,
          "replies": [
            {
              "id": 2615051,
              "postDate": "2024-01-23T00:46:18.267Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thank you Chris, i will learn and explore. If i understood well, then will make a notebook about it.</p>\n<blockquote>\n  <p>I will try to reverse engineer one of your magic formula <strong>np.clip(img,np.exp(-4),np.exp(8))</strong></p>\n</blockquote>",
              "rawMarkdown": "@cdeotte thank you Chris, i will learn and explore. If i understood well, then will make a notebook about it.\n> I will try to reverse engineer one of your magic formula **np.clip(img,np.exp(-4),np.exp(8))**",
              "votes": 1
            },
            {
              "id": 2615129,
              "postDate": "2024-01-23T01:47:33.703Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> now i can see, what you seen (left skewed distribution). Thank you</p>\n<blockquote>\n  <p><strong>We eliminated outliers using np.clip from e^-4 to e^8 values, is the major improvement of LB 0.10 score by just adding this line in most of top public notebooks</strong> - is this conclusion right?</p>\n</blockquote>\n<p>Before clip</p>\n<pre><code> ~ -\n ~ \n mean ~ . and std ~ .\n</code></pre>\n<p>After clip</p>\n<pre><code> ~ -\n ~  \n mean ~ . and std ~ .\n</code></pre>",
              "rawMarkdown": "@cdeotte now i can see, what you seen (left skewed distribution). Thank you\n> **We eliminated outliers using np.clip from e^-4 to e^8 values, is the major improvement of LB 0.10 score by just adding this line in most of top public notebooks** - is this conclusion right?\n\nBefore clip\n```\nmin ~ -4\nmax ~ 10\nbut mean ~ 0.7 and std ~ 1.7\n```\n\nAfter clip\n```\nmin ~ -4\nmax ~  8\nbut mean ~ 0.7 and std ~ 1.7\n```"
            },
            {
              "id": 2616976,
              "postDate": "2024-01-24T02:12:20.487Z",
              "content": "<p>very helpful thanks!</p>",
              "rawMarkdown": "very helpful thanks!",
              "votes": 1
            },
            {
              "id": 2617043,
              "postDate": "2024-01-24T03:38:36.663Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> i'm able to do analysis only for few samples (so did with 1 sample) <a href=\"https://www.kaggle.com/code/seshurajup/spectrogram-distribution-analysis/notebook\" target=\"_blank\">here</a>, otherwise notebook crashing with memory limit. what is best way to do for complete data same analysis?.</p>",
              "rawMarkdown": "@cdeotte i'm able to do analysis only for few samples (so did with 1 sample) [here](https://www.kaggle.com/code/seshurajup/spectrogram-distribution-analysis/notebook), otherwise notebook crashing with memory limit. what is best way to do for complete data same analysis?.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2670611,
      "postDate": "2024-02-27T01:44:04.243Z",
      "content": "<p>I'm creating dataset of EEGs &gt;20GB so I think I have to use the Kaggle API to upload a dataset. Has anyone successfully done this? I'm getting issues with this code:<br>\n`from kaggle_secrets import UserSecretsClient<br>\nuser_secrets = UserSecretsClient()<br>\nsecret_value_0 = user_secrets.get_secret(\"KAGGLE_KEY\")<br>\nsecret_value_1 = user_secrets.get_secret(\"KAGGLE_USERNAME\")<br>\nos.environ['KAGGLE_USERNAME'] = secret_value_0<br>\nos.environ['KAGGLE_KEY'] = secret_value_1</p>\n<p>import os<br>\nimport json<br>\nos.makedirs('/kaggle/dataset/', exist_ok=True)<br>\n!kaggle datasets init -p '/kaggle/dataset'</p>\n<h1>Change below</h1>\n<p>meta = dict(<br>\n    id=\"maxdli262/hms-cwt\",<br>\n    title=\"My CWTs for HMS competition\",<br>\n    isPrivate=True,<br>\n    licenses=[dict(name=\"CC0-1.0\")],<br>\n)</p>\n<p>with open('/kaggle/dataset/dataset-metadata.json', 'w') as f:<br>\n    json.dump(meta, f)</p>\n<p>!kaggle datasets create -p \"/kaggle/dataset\" --dir-mode zip<br>\n`</p>\n<p>I get this output:</p>\n<blockquote>\n  <p>Starting upload for file EEG_Spectrograms.zip<br>\n  Invalid value for <code>token</code>, must not be <code>None</code></p>\n</blockquote>\n<p>Can't seem to trace what 'token' is. Anyone deal with this before?</p>",
      "rawMarkdown": "I'm creating dataset of EEGs >20GB so I think I have to use the Kaggle API to upload a dataset. Has anyone successfully done this? I'm getting issues with this code:\n`from kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nsecret_value_0 = user_secrets.get_secret(\"KAGGLE_KEY\")\nsecret_value_1 = user_secrets.get_secret(\"KAGGLE_USERNAME\")\nos.environ['KAGGLE_USERNAME'] = secret_value_0\nos.environ['KAGGLE_KEY'] = secret_value_1\n\nimport os\nimport json\nos.makedirs('/kaggle/dataset/', exist_ok=True)\n!kaggle datasets init -p '/kaggle/dataset'\n\n# Change below\nmeta = dict(\n    id=\"maxdli262/hms-cwt\",\n    title=\"My CWTs for HMS competition\",\n    isPrivate=True,\n    licenses=[dict(name=\"CC0-1.0\")],\n)\n\nwith open('/kaggle/dataset/dataset-metadata.json', 'w') as f:\n    json.dump(meta, f)\n\n\n!kaggle datasets create -p \"/kaggle/dataset\" --dir-mode zip\n`\n\nI get this output:\n>Starting upload for file EEG_Spectrograms.zip\nInvalid value for `token`, must not be `None`\n\nCan't seem to trace what 'token' is. Anyone deal with this before?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2670626,
          "postDate": "2024-02-27T02:06:09.980Z",
          "content": "<p>I'm not sure. Another option is to make two notebooks and each notebook can have output less than 20GB (i.e. creates half of the spectrograms). Then you don't need to upload to a Kaggle dataset. Just connect the output of that notebook to your next notebook. </p>",
          "rawMarkdown": "I'm not sure. Another option is to make two notebooks and each notebook can have output less than 20GB (i.e. creates half of the spectrograms). Then you don't need to upload to a Kaggle dataset. Just connect the output of that notebook to your next notebook. "
        }
      ]
    },
    {
      "id": 2655198,
      "postDate": "2024-02-16T18:07:13.297Z",
      "content": "<p>LL Spec = ( spec(Fp1 - F7) + spec(F7 - T3) + spec(T3 - T5) + spec(T5 - O1) )/4</p>\n<p>Sorry, what does spec mean here? </p>",
      "rawMarkdown": "LL Spec = ( spec(Fp1 - F7) + spec(F7 - T3) + spec(T3 - T5) + spec(T5 - O1) )/4\n\nSorry, what does spec mean here? ",
      "votes": 1,
      "replies": [
        {
          "id": 2655203,
          "postDate": "2024-02-16T18:10:59.997Z",
          "content": "<p>Spectrogram</p>",
          "rawMarkdown": "Spectrogram",
          "votes": 2,
          "replies": [
            {
              "id": 2657103,
              "postDate": "2024-02-18T09:57:30.730Z",
              "content": "<p>Ok, but spec(Fp1 - F7), for instance, does that mean spec(Fp1) - spec(F7)? I'm sorry, but I don't know how to interpret this expression.</p>",
              "rawMarkdown": "Ok, but spec(Fp1 - F7), for instance, does that mean spec(Fp1) - spec(F7)? I'm sorry, but I don't know how to interpret this expression."
            },
            {
              "id": 2657107,
              "postDate": "2024-02-18T10:01:54.150Z",
              "content": "<p>It means create a new column (i.e. waveform) from parquet column <code>Fp1</code> minus column <code>F7</code>. Then create a spectrogram from this column (i.e. waveform). Now we have a NumPy array of size <code>(128x256)</code> which is <code>(frequence x time)</code> spectrogram. This is <code>spec(Fp1-F7)</code>.</p>\n<p>Afterward, we make 3 more spectrograms for LL montage. And lastly we average the 4 NumPy array spectrograms to create a single NumPy array spectrogram</p>",
              "rawMarkdown": "It means create a new column (i.e. waveform) from parquet column `Fp1` minus column `F7`. Then create a spectrogram from this column (i.e. waveform). Now we have a NumPy array of size `(128x256)` which is `(frequence x time)` spectrogram. This is `spec(Fp1-F7)`.\n\nAfterward, we make 3 more spectrograms for LL montage. And lastly we average the 4 NumPy array spectrograms to create a single NumPy array spectrogram",
              "votes": 3
            },
            {
              "id": 2659230,
              "postDate": "2024-02-19T17:20:08.843Z",
              "content": "<p>Thank you very much for the answer, now everything is much clearer.</p>",
              "rawMarkdown": "Thank you very much for the answer, now everything is much clearer.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2625005,
      "postDate": "2024-01-29T06:27:25.590Z",
      "content": "<p>Thank you!!…i will try to integrate this new information <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thank you!!...i will try to integrate this new information @cdeotte ",
      "votes": 1
    },
    {
      "id": 2624964,
      "postDate": "2024-01-29T05:37:43.917Z",
      "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nAs a beginner of deep learning, I would like to ask you, for general classification task, what learning rate and epoch would you choose to achieve the best training effect?</p>",
      "rawMarkdown": "Thank you for sharing @cdeotte \nAs a beginner of deep learning, I would like to ask you, for general classification task, what learning rate and epoch would you choose to achieve the best training effect?",
      "votes": 1,
      "replies": [
        {
          "id": 2625245,
          "postDate": "2024-01-29T09:39:47.643Z",
          "content": "<p>Hi. Learning rate and epoch depends on your type of model. For example CatBoost, MLP, WaveNet, EfficientNet, and BERT all have different requirements. For NN, we can try Adam 1e-3, 1e-4, and 1e-5. And try 5, 10, 20, 40, 80 epochs. And we can try various schedulers like Cosine schedule or Step schedule.</p>\n<p>If you name a specific model, i can give more specific recommendations.</p>",
          "rawMarkdown": "Hi. Learning rate and epoch depends on your type of model. For example CatBoost, MLP, WaveNet, EfficientNet, and BERT all have different requirements. For NN, we can try Adam 1e-3, 1e-4, and 1e-5. And try 5, 10, 20, 40, 80 epochs. And we can try various schedulers like Cosine schedule or Step schedule.\n\nIf you name a specific model, i can give more specific recommendations.",
          "votes": 6,
          "replies": [
            {
              "id": 2625304,
              "postDate": "2024-01-29T10:37:28.860Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2625503,
              "postDate": "2024-01-29T13:04:54.077Z",
              "content": "<p>Thank you for your reply, these answers are enough</p>",
              "rawMarkdown": "Thank you for your reply, these answers are enough",
              "votes": 1
            },
            {
              "id": 2650406,
              "postDate": "2024-02-13T12:51:12.713Z",
              "content": "<p>This is very helpful for me as deep learning beginner.<br>\nQuestion about hyper parameter tuning.</p>\n<ol>\n<li>How do you tune hyper parameter in python code ? Do you use optimization library or just grid search in for-loop ?</li>\n<li>In what order do you tune hyper parameter ? There is no general way to tune ?</li>\n<li>How do you decide parameter of scheduler ?</li>\n</ol>\n<p>I apologize for many questions, but I would appreciate it if you could provide some answers.</p>",
              "rawMarkdown": "This is very helpful for me as deep learning beginner.\nQuestion about hyper parameter tuning.\n1. How do you tune hyper parameter in python code ? Do you use optimization library or just grid search in for-loop ?\n2. In what order do you tune hyper parameter ? There is no general way to tune ?\n3. How do you decide parameter of scheduler ?\n\nI apologize for many questions, but I would appreciate it if you could provide some answers.",
              "votes": 1
            },
            {
              "id": 2726172,
              "postDate": "2024-04-01T03:01:39.517Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>  Sorry for the late reply. I posted a discussion about tuning learning rates <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/488083\" target=\"_blank\">here</a></p>",
              "rawMarkdown": "Hi @clearwaterkzk  Sorry for the late reply. I posted a discussion about tuning learning rates [here][1]\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/488083"
            },
            {
              "id": 2726231,
              "postDate": "2024-04-01T04:13:42.730Z",
              "content": "<p>Thank you very much.<br>\nI will definitely buy it if you would release a book on deep learning.😀</p>",
              "rawMarkdown": "Thank you very much.\nI will definitely buy it if you would release a book on deep learning.😀",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2623462,
      "postDate": "2024-01-28T08:39:49.123Z",
      "content": "<p>This is very useful. Thank you for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "This is very useful. Thank you for sharing @cdeotte ",
      "votes": 1
    },
    {
      "id": 2619352,
      "postDate": "2024-01-25T12:14:33.130Z",
      "content": "<p>This is really nice! Thank you for sharing! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "This is really nice! Thank you for sharing! @cdeotte ",
      "votes": 1
    },
    {
      "id": 2617844,
      "postDate": "2024-01-24T13:37:53.190Z",
      "content": "<p>Impressive! Thank you for sharing.<br>\nI was able to reproduce your results with pytorch.</p>",
      "rawMarkdown": "Impressive! Thank you for sharing.\nI was able to reproduce your results with pytorch.",
      "votes": 1
    },
    {
      "id": 2615728,
      "postDate": "2024-01-23T09:29:50.573Z",
      "content": "<p>Thank you for sharing!<br>\nWhat is the difference between ['LL','LP','RP','RR'] in your code and ['LL','LP','RL','RP'] in spectrogram parquet columns?<br>\nIs it just typo?</p>",
      "rawMarkdown": "Thank you for sharing!\nWhat is the difference between ['LL','LP','RP','RR'] in your code and ['LL','LP','RL','RP'] in spectrogram parquet columns?\nIs it just typo?",
      "votes": 1,
      "replies": [
        {
          "id": 2615743,
          "postDate": "2024-01-23T09:36:02.097Z",
          "content": "<p>Yes, just a typo. The RR should be RL</p>",
          "rawMarkdown": "Yes, just a typo. The RR should be RL",
          "votes": 2
        }
      ]
    },
    {
      "id": 2617237,
      "postDate": "2024-01-24T06:11:50.853Z",
      "content": "<p>Thank you for sharing the useful info 🤓</p>",
      "rawMarkdown": "Thank you for sharing the useful info 🤓"
    },
    {
      "id": 2659631,
      "postDate": "2024-02-20T03:38:42.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2619314,
      "postDate": "2024-01-25T12:02:29.453Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2632849,
      "postDate": "2024-02-02T16:55:08.153Z",
      "content": "<p>Thank you for sharing the information ! </p>",
      "rawMarkdown": "Thank you for sharing the information ! ",
      "votes": 1
    },
    {
      "id": 2618768,
      "postDate": "2024-01-25T01:26:18.127Z",
      "content": "<p>Thanks Chris! This is so inspiring!</p>",
      "rawMarkdown": "Thanks Chris! This is so inspiring!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2632312,
      "author_name": "Raki",
      "author_url": "",
      "post_date": "2024-02-02T09:09:53.267000",
      "content": "<p>I found a visualization in a video of how the anterior-posterior bipolar montage is created (the names of a few electrode locations are different though).<br>\n<a href=\"https://www.youtube.com/clip/UgkxqTZZRvcuSHO4la1-qGMCpGai_VF6h-QD\" target=\"_blank\">https://www.youtube.com/clip/UgkxqTZZRvcuSHO4la1-qGMCpGai_VF6h-QD</a></p>\n<p>Disclaimer: I'm not 100% certain it's the same as the magic formula.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2618146,
      "author_name": "Aindriú",
      "author_url": "",
      "post_date": "2024-01-24T15:31:01.580000",
      "content": "<p>Spectrum is a linear transform. Sum of the spec should be equal to the spec of the sum which should reduce to spec(Fp1 - O1) similarly. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2618159,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-24T15:42:02.533000",
          "content": "<p>I don't think so. The spectrogram ignores the phase of sine waves. If i have two 60Hz sine waves that are 180 degrees out of phase, then if i add up the raw waveforms, the result is zero everywhere. However if i take their 2 spectrograms and add them up the result is <strong>not</strong> zero everywhere.</p>\n<p>(In other words, raw waveforms are positive and negative. Whereas spectrograms are non-negative. Adding spectrograms together does not cancel out the spectrograms).</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 2613568,
      "author_name": "Donghui Zhang",
      "author_url": "",
      "post_date": "2024-01-22T06:19:56.927000",
      "content": "<p>Good idea, in fact I used it to train the model on a data set you posted for making spectrograms, but the training loss deviated greatly from the validation loss.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2615026,
          "author_name": "HZM",
          "author_url": "",
          "post_date": "2024-01-23T00:32:11.163000",
          "content": "<p>same issue, pytoch did not work as good as tensorflow</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2613249,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-01-22T00:02:28.667000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for sharing</p>\n<blockquote>\n  <h3>LOG TRANSFORM SPECTROGRAM</h3>\n  <p>img = np.clip(img,np.exp(-4),np.exp(8)) ==&gt; for Spectrogram </p>\n</blockquote>\n<hr>\n<blockquote>\n  <h3>LOG TRANSFORM</h3>\n  <p>width = (mel_spec.shape[1]//32)*32<br>\n  mel_spec_db = librosa.power_to_db(mel_spec, ref=np.max).astype(np.float32)[:,:width] ==&gt; for EGG -&gt; Spectrogram</p>\n</blockquote>\n<p>How to choice transformations? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2614723,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-22T19:10:47.443000",
          "content": "<p>Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.</p>\n<p>The most common transformation is <code>standardization</code> which is <code>new data = (old data - mean) / std</code>. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a <code>log transform</code> to remove the skew. So first we <code>log transform</code> and next we <code>standardize</code>. </p>\n<p>Log transforms can only handle positive numbers (i.e <code>x&gt;0</code>), so we must shift, clip, and/or flip all the data to be <code>data &gt; 0</code> before we perform <code>log transform</code>.</p>",
          "votes": 17,
          "replies": []
        },
        {
          "id": 2614725,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-22T19:13:28.727000",
          "content": "<p>Another trick is using <code>Gauss Rank Transform</code> (shown <a href=\"https://medium.com/rapids-ai/gauss-rank-transformation-is-100x-faster-with-rapids-and-cupy-7c947e3397da\" target=\"_blank\">here</a>). Or using a <code>Quantile Transform</code> <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.QuantileTransformer.html\" target=\"_blank\">here</a>. Both can transform any distribution into a nearly perfect Gaussian distribution afterward. </p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/dog3.png\"></p>",
          "votes": 13,
          "replies": [
            {
              "id": 2615051,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-23T00:46:18.267000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thank you Chris, i will learn and explore. If i understood well, then will make a notebook about it.</p>\n<blockquote>\n  <p>I will try to reverse engineer one of your magic formula <strong>np.clip(img,np.exp(-4),np.exp(8))</strong></p>\n</blockquote>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2615129,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-23T01:47:33.703000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> now i can see, what you seen (left skewed distribution). Thank you</p>\n<blockquote>\n  <p><strong>We eliminated outliers using np.clip from e^-4 to e^8 values, is the major improvement of LB 0.10 score by just adding this line in most of top public notebooks</strong> - is this conclusion right?</p>\n</blockquote>\n<p>Before clip</p>\n<pre><code> ~ -\n ~ \n mean ~ . and std ~ .\n</code></pre>\n<p>After clip</p>\n<pre><code> ~ -\n ~  \n mean ~ . and std ~ .\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2616976,
              "author_name": "Edward Tuck",
              "author_url": "",
              "post_date": "2024-01-24T02:12:20.487000",
              "content": "<p>very helpful thanks!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2617043,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-24T03:38:36.663000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> i'm able to do analysis only for few samples (so did with 1 sample) <a href=\"https://www.kaggle.com/code/seshurajup/spectrogram-distribution-analysis/notebook\" target=\"_blank\">here</a>, otherwise notebook crashing with memory limit. what is best way to do for complete data same analysis?.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2670611,
      "author_name": "Max Li",
      "author_url": "",
      "post_date": "2024-02-27T01:44:04.243000",
      "content": "<p>I'm creating dataset of EEGs &gt;20GB so I think I have to use the Kaggle API to upload a dataset. Has anyone successfully done this? I'm getting issues with this code:<br>\n`from kaggle_secrets import UserSecretsClient<br>\nuser_secrets = UserSecretsClient()<br>\nsecret_value_0 = user_secrets.get_secret(\"KAGGLE_KEY\")<br>\nsecret_value_1 = user_secrets.get_secret(\"KAGGLE_USERNAME\")<br>\nos.environ['KAGGLE_USERNAME'] = secret_value_0<br>\nos.environ['KAGGLE_KEY'] = secret_value_1</p>\n<p>import os<br>\nimport json<br>\nos.makedirs('/kaggle/dataset/', exist_ok=True)<br>\n!kaggle datasets init -p '/kaggle/dataset'</p>\n<h1>Change below</h1>\n<p>meta = dict(<br>\n    id=\"maxdli262/hms-cwt\",<br>\n    title=\"My CWTs for HMS competition\",<br>\n    isPrivate=True,<br>\n    licenses=[dict(name=\"CC0-1.0\")],<br>\n)</p>\n<p>with open('/kaggle/dataset/dataset-metadata.json', 'w') as f:<br>\n    json.dump(meta, f)</p>\n<p>!kaggle datasets create -p \"/kaggle/dataset\" --dir-mode zip<br>\n`</p>\n<p>I get this output:</p>\n<blockquote>\n  <p>Starting upload for file EEG_Spectrograms.zip<br>\n  Invalid value for <code>token</code>, must not be <code>None</code></p>\n</blockquote>\n<p>Can't seem to trace what 'token' is. Anyone deal with this before?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2670626,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-02-27T02:06:09.980000",
          "content": "<p>I'm not sure. Another option is to make two notebooks and each notebook can have output less than 20GB (i.e. creates half of the spectrograms). Then you don't need to upload to a Kaggle dataset. Just connect the output of that notebook to your next notebook. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2655198,
      "author_name": "Adrian Tif",
      "author_url": "",
      "post_date": "2024-02-16T18:07:13.297000",
      "content": "<p>LL Spec = ( spec(Fp1 - F7) + spec(F7 - T3) + spec(T3 - T5) + spec(T5 - O1) )/4</p>\n<p>Sorry, what does spec mean here? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2655203,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-02-16T18:10:59.997000",
          "content": "<p>Spectrogram</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2657103,
              "author_name": "Adrian Tif",
              "author_url": "",
              "post_date": "2024-02-18T09:57:30.730000",
              "content": "<p>Ok, but spec(Fp1 - F7), for instance, does that mean spec(Fp1) - spec(F7)? I'm sorry, but I don't know how to interpret this expression.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2657107,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-02-18T10:01:54.150000",
              "content": "<p>It means create a new column (i.e. waveform) from parquet column <code>Fp1</code> minus column <code>F7</code>. Then create a spectrogram from this column (i.e. waveform). Now we have a NumPy array of size <code>(128x256)</code> which is <code>(frequence x time)</code> spectrogram. This is <code>spec(Fp1-F7)</code>.</p>\n<p>Afterward, we make 3 more spectrograms for LL montage. And lastly we average the 4 NumPy array spectrograms to create a single NumPy array spectrogram</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2659230,
              "author_name": "Adrian Tif",
              "author_url": "",
              "post_date": "2024-02-19T17:20:08.843000",
              "content": "<p>Thank you very much for the answer, now everything is much clearer.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2625005,
      "author_name": "N Sai Harshith Varma",
      "author_url": "",
      "post_date": "2024-01-29T06:27:25.590000",
      "content": "<p>Thank you!!…i will try to integrate this new information <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2624964,
      "author_name": "zhiyue666",
      "author_url": "",
      "post_date": "2024-01-29T05:37:43.917000",
      "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nAs a beginner of deep learning, I would like to ask you, for general classification task, what learning rate and epoch would you choose to achieve the best training effect?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2625245,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-29T09:39:47.643000",
          "content": "<p>Hi. Learning rate and epoch depends on your type of model. For example CatBoost, MLP, WaveNet, EfficientNet, and BERT all have different requirements. For NN, we can try Adam 1e-3, 1e-4, and 1e-5. And try 5, 10, 20, 40, 80 epochs. And we can try various schedulers like Cosine schedule or Step schedule.</p>\n<p>If you name a specific model, i can give more specific recommendations.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2625304,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-29T10:37:28.860000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2625503,
              "author_name": "zhiyue666",
              "author_url": "",
              "post_date": "2024-01-29T13:04:54.077000",
              "content": "<p>Thank you for your reply, these answers are enough</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2650406,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-02-13T12:51:12.713000",
              "content": "<p>This is very helpful for me as deep learning beginner.<br>\nQuestion about hyper parameter tuning.</p>\n<ol>\n<li>How do you tune hyper parameter in python code ? Do you use optimization library or just grid search in for-loop ?</li>\n<li>In what order do you tune hyper parameter ? There is no general way to tune ?</li>\n<li>How do you decide parameter of scheduler ?</li>\n</ol>\n<p>I apologize for many questions, but I would appreciate it if you could provide some answers.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2726172,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-04-01T03:01:39.517000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>  Sorry for the late reply. I posted a discussion about tuning learning rates <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/488083\" target=\"_blank\">here</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2726231,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-01T04:13:42.730000",
              "content": "<p>Thank you very much.<br>\nI will definitely buy it if you would release a book on deep learning.😀</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2623462,
      "author_name": "Virendra Kumar",
      "author_url": "",
      "post_date": "2024-01-28T08:39:49.123000",
      "content": "<p>This is very useful. Thank you for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2619352,
      "author_name": "Daeeon_H",
      "author_url": "",
      "post_date": "2024-01-25T12:14:33.130000",
      "content": "<p>This is really nice! Thank you for sharing! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2617844,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2024-01-24T13:37:53.190000",
      "content": "<p>Impressive! Thank you for sharing.<br>\nI was able to reproduce your results with pytorch.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2615728,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2024-01-23T09:29:50.573000",
      "content": "<p>Thank you for sharing!<br>\nWhat is the difference between ['LL','LP','RP','RR'] in your code and ['LL','LP','RL','RP'] in spectrogram parquet columns?<br>\nIs it just typo?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2615743,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-23T09:36:02.097000",
          "content": "<p>Yes, just a typo. The RR should be RL</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2617237,
      "author_name": "Devang Giri Goswami",
      "author_url": "",
      "post_date": "2024-01-24T06:11:50.853000",
      "content": "<p>Thank you for sharing the useful info 🤓</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2659631,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-20T03:38:42.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2619314,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-25T12:02:29.453000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2632849,
      "author_name": "TheEventHorizons",
      "author_url": "",
      "post_date": "2024-02-02T16:55:08.153000",
      "content": "<p>Thank you for sharing the information ! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2618768,
      "author_name": "Qianyi Zhao",
      "author_url": "",
      "post_date": "2024-01-25T01:26:18.127000",
      "content": "<p>Thanks Chris! This is so inspiring!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2613222": "# Magic Formula\n\nExciting news everyone. I think we found the **magic formula** to convert EEG waveforms into Kaggle Spectrograms.\n\n    LL Spec = ( spec(Fp1 - F7) + spec(F7 - T3) + spec(T3 - T5) + spec(T5 - O1) )/4\n    LP Spec = ( spec(Fp1 - F3) + spec(F3 - C3) + spec(C3 - P3) + spec(P3 - O1) )/4\n    RP Spec = ( spec(Fp2 - F4) + spec(F4 - C4) + spec(C4 - P4) + spec(P4 - O2) )/4\n    RL Spec = ( spec(Fp2 - F8) + spec(F8 - T4) + spec(T4 - T6) + spec(T6 - O2) )/4\n\nThe trick is that we need to average four spectrograms to produce the LL (i.e. Left Temporal Chain) single spectrogram instead of combining Fp1, F7, T3, T5, O1 into a single waveform and computing a single spectrogram. If we combine waveforms first with `W = (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1)` then this reduces to `W = Fp1 - O1` and it doesn't use information from `F7 , T3, T5`. But the above **magic formula** does use information from all EEG waveforms.\n\n# History\n\nOne week ago, we started a discussion [here][1] to determine how the Kaggle spectrograms were made. Kaggle users @seshurajup and others have been active and helpful in this conversation. I tried various ways to combine the 19 EEG waveforms into the 4 spectrograms `LL, LP, RL, RP`. I have finally created spectrograms that achieve a better CV and LB score than the Kaggle spectrograms!\n\n# Experiments\nWe can review the following experiments in my EfficientNet starter notebook [here][3]\n\n| Spectrogram | EffNet 5Fold CV | EffNet LB | Notebook version | Model | Loss | Epochs | Data Augmention |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| Kaggle Spectrograms | 0.73 | 0.57 | ver 1 | EffNetB2 | CE | 3 | none |\n| Kaggle Spectrograms | 0.66 | ??? | ver 4 | EffNetB0 | KL-Div | 4 | none |\n| EEG Spectrograms | 0.63 | ??? | ver 3 | EffNetB0 | KL-Div | 4 | none |\n| Both Spectrograms | 0.59 | 0.44 | ver 5| EffNetB0 | KL-Div | 4 | none | \n\n\nIf we compare row 2 and row 3 above, we see that training EfficientNet B0 on only the Magic formula EEG spectrograms achieves a better CV score and LB score than training EfficientNet B0 on only the Kaggle spectrograms.\n\nFurthermore, training a single model with both spectrograms achieves the best CV score and LB score. I think this confirms that we have a powerful formula for converting 19 EEG waveforms into 4 spectrograms!\n\n# How To Use EEG Spectrograms\nExamples of how to use new EEG spectrograms to boost CV score and LB score are published in recent versions of my EfficientNet starter notebook [here][3] and CatBoost starter notebook [here][4].\n\n# Kaggle Dataset\nThe new EEG spectrograms using the **Magic Formula** have been uploaded to a Kaggle dataset [here][5]. We can attach this Kaggle dataset to our future Kaggle notebooks to boost our CV scores and LB scores! \n\nIf working offline, we can download one 7.5GB Python dictionary file that contains all 17089 **EEG spectrograms** with the command: \n\n    kaggle datasets download -f eeg_specs.npy cdeotte/brain-eeg-spectrograms\n\nAnd we can download all of Kaggle spectrograms in one 2.5GB Python dictionary that contains all 11138 Kaggle spectrograms with command:\n\n    kaggle datasets download cdeotte/brain-spectrograms\n\n# Future Work\nSince these new spectrograms help improve our model, this means we can try exploring more formulas and alternative ways to create images from EEG waveforms. For example @abebe9849 is exploring using different types of Fourier transform in his notebook [here][2].\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467877\n[2]: https://www.kaggle.com/code/abebe9849/make-cwt-from-eeg-compare-3-types-of-images\n[3]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[4]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[5]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms",
    "2632312": "I found a visualization in a video of how the anterior-posterior bipolar montage is created (the names of a few electrode locations are different though).\nhttps://www.youtube.com/clip/UgkxqTZZRvcuSHO4la1-qGMCpGai_VF6h-QD\n\nDisclaimer: I'm not 100% certain it's the same as the magic formula.",
    "2618146": "Spectrum is a linear transform. Sum of the spec should be equal to the spec of the sum which should reduce to spec(Fp1 - O1) similarly. ",
    "2613568": "Good idea, in fact I used it to train the model on a data set you posted for making spectrograms, but the training loss deviated greatly from the validation loss.",
    "2613249": "@cdeotte Thank you for sharing\n\n> ### LOG TRANSFORM SPECTROGRAM\n img = np.clip(img,np.exp(-4),np.exp(8)) ==> for Spectrogram \n\n---\n> ### LOG TRANSFORM\nwidth = (mel_spec.shape[1]//32)*32\nmel_spec_db = librosa.power_to_db(mel_spec, ref=np.max).astype(np.float32)[:,:width] ==> for EGG -> Spectrogram\n\nHow to choice transformations? ",
    "2670611": "I'm creating dataset of EEGs >20GB so I think I have to use the Kaggle API to upload a dataset. Has anyone successfully done this? I'm getting issues with this code:\n`from kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nsecret_value_0 = user_secrets.get_secret(\"KAGGLE_KEY\")\nsecret_value_1 = user_secrets.get_secret(\"KAGGLE_USERNAME\")\nos.environ['KAGGLE_USERNAME'] = secret_value_0\nos.environ['KAGGLE_KEY'] = secret_value_1\n\nimport os\nimport json\nos.makedirs('/kaggle/dataset/', exist_ok=True)\n!kaggle datasets init -p '/kaggle/dataset'\n\n# Change below\nmeta = dict(\n    id=\"maxdli262/hms-cwt\",\n    title=\"My CWTs for HMS competition\",\n    isPrivate=True,\n    licenses=[dict(name=\"CC0-1.0\")],\n)\n\nwith open('/kaggle/dataset/dataset-metadata.json', 'w') as f:\n    json.dump(meta, f)\n\n\n!kaggle datasets create -p \"/kaggle/dataset\" --dir-mode zip\n`\n\nI get this output:\n>Starting upload for file EEG_Spectrograms.zip\nInvalid value for `token`, must not be `None`\n\nCan't seem to trace what 'token' is. Anyone deal with this before?\n",
    "2655198": "LL Spec = ( spec(Fp1 - F7) + spec(F7 - T3) + spec(T3 - T5) + spec(T5 - O1) )/4\n\nSorry, what does spec mean here? ",
    "2625005": "Thank you!!...i will try to integrate this new information @cdeotte ",
    "2624964": "Thank you for sharing @cdeotte \nAs a beginner of deep learning, I would like to ask you, for general classification task, what learning rate and epoch would you choose to achieve the best training effect?",
    "2623462": "This is very useful. Thank you for sharing @cdeotte ",
    "2619352": "This is really nice! Thank you for sharing! @cdeotte ",
    "2617844": "Impressive! Thank you for sharing.\nI was able to reproduce your results with pytorch.",
    "2615728": "Thank you for sharing!\nWhat is the difference between ['LL','LP','RP','RR'] in your code and ['LL','LP','RL','RP'] in spectrogram parquet columns?\nIs it just typo?",
    "2617237": "Thank you for sharing the useful info 🤓",
    "2659631": "",
    "2619314": "",
    "2632849": "Thank you for sharing the information ! ",
    "2618768": "Thanks Chris! This is so inspiring!"
  }
}