{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.12"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30635,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"papermill":{"default_parameters":{},"duration":102.485133,"end_time":"2024-01-16T12:14:21.10455","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-01-16T12:12:38.619417","version":"2.4.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# How To Make Spectrograms from EEG\nIn this notebook, we learn how to make spectrograms from EEG. The EEGs are waveforms and the Spectrograms are images. There is a discussion about this notebook [here][1].\n\nIn version 1-3, we also train a simple model using our EEG spectrograms to confirm that they work well. We observe that a model trained with EEG spectrograms performs better than baseline models using only train means.\n\n# Exciting UPDATE!\nVersion 4 of this notebook uses a different formula to make spectrograms than earlier versions. I trained an EfficientNet model using the old version eeg spectrograms, new version spectrograms, and Kaggle spectrograms. We can see that the new version eeg spectrograms are **powerful**!\n\n| Spectrogram | 5-Fold CV | LB |\n| --- | --- | --- |\n| Kaggle spectrogram | 0.73 | 0.57 |\n| Old EEG formula | 0.84 on fold 1 | ?? |\n| New EEG formula | 0.70 on fold 1| ?? |\n| Use both Kaggle and New | 0.64 | 0.44 |\n\nFrom the results above, we conclude that our new formula is probably similar or better than the true formula used to create the Kaggle spectrograms. Details about the old and new formula are in the next notebook section. \n\n# How To Use EEG Spectrograms\nExamples of how to use new EEG spectrograms to boost CV score and LB score will be (or already are) published in recent versions of my EfficientNet starter notebook [here][2] and CatBoost starter notebook [here][3]\n\n# Kaggle Dataset\nThe new EEG spectrograms from version 4 of this notebook have been uploaded to a Kaggle dataset [here][4]. We can attach this Kaggle dataset to our future notebooks to boost our CV scores and LB scores! Thank you everyone for upvoting my new EEG spectrogram Kaggle dataset!\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467877\n[2]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[3]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[4]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms","metadata":{"papermill":{"duration":0.00726,"end_time":"2024-01-16T12:12:41.9057","exception":false,"start_time":"2024-01-16T12:12:41.89844","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pandas as pd, numpy as np, os\nimport matplotlib.pyplot as plt, gc\n\ntrain = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\nprint('Train shape', train.shape )\ndisplay( train.head() )","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":0.693827,"end_time":"2024-01-16T12:12:42.606147","exception":false,"start_time":"2024-01-16T12:12:41.91232","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-11-01T18:07:39.711029Z","iopub.execute_input":"2024-11-01T18:07:39.711430Z","iopub.status.idle":"2024-11-01T18:07:39.902798Z","shell.execute_reply.started":"2024-11-01T18:07:39.711399Z","shell.execute_reply":"2024-11-01T18:07:39.901612Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# The Bipolar Double Banana Montage\nIn the Kaggle discussion [here][1], we learn what information we need to make spectrograms from eegs. The following website [here][2] is helpful also. To build 1 spectrogram, we need 1 time series signal. Kaggle provides us with 19 eeg time signals, so we must combine them into 4 time signals to make 4 spectrograms.\n\nIn the diagram below, we see which electrode signals are needed to make the `LL, LP, RP, RR` spectrograms. Furthermore Kaggle discussions imply that most likely we create differences between consecutive electrodes and average the differences. For example, we create `LL spectrogram` with the formula: \n\n    LL = ( (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1) )/4.\n    \nI am not positive that this is the correct formula. I also tried the formula below but it produced a worse CV score than the above formula, so perhaps the above is correct. I am confident that we only use these 5 electrodes to create `LL spectrogram`. I'm just a little unsure about the formula:\n\n    LL = ( Fp1 + F7 + T3 + T5 + O1 )/5.\n    \n# Exciting UPDATE!\nI believe the above two formulas are wrong. Many Kagglers have pointed out that the above formula reduces to `LL = ( Fp1 - O1 )/4` which means that it does not use all the EEG signals. The new formula below utilizes all the EEG signals and produces EEG spectrograms that achieve better CV score and LB score than the Kaggle spectrograms. Therefore I think the following formula is the correct one:\n\n    LL Spec = ( spec(Fp1 - F7) + spec(F7 - T3) + spec(T3 - T5) + spec(T5 - O1) )/4.\n    \nSince creating a spectrogram is a non-linear operation, the above formula which computes 4 spectrograms and then takes the average is different than the formula below which computes 1 spectrogam. And the above formula does utilize all EEG signals and cannot be reduced to a shorter formula (like the one below).\n\n    LL Spec = spec( ( (Fp1 - F7) + (F7 - T3) + (T3 - T5) + (T5 - O1) )/4. )\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jan-2024/montage.png)\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467877\n[2]: https://www.learningeeg.com/montages-and-technical-components","metadata":{"papermill":{"duration":0.007102,"end_time":"2024-01-16T12:12:42.620233","exception":false,"start_time":"2024-01-16T12:12:42.613131","status":"completed"},"tags":[]}},{"cell_type":"code","source":"NAMES = ['LL','LP','RP','RR']\n\nFEATS = [['Fp1','F7','T3','T5','O1'],\n         ['Fp1','F3','C3','P3','O1'],\n         ['Fp2','F8','T4','T6','O2'],\n         ['Fp2','F4','C4','P4','O2']]\n\ndirectory_path = 'EEG_Spectrograms/'\nif not os.path.exists(directory_path):\n    os.makedirs(directory_path)","metadata":{"papermill":{"duration":0.017131,"end_time":"2024-01-16T12:12:42.644893","exception":false,"start_time":"2024-01-16T12:12:42.627762","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-11-01T18:07:39.905464Z","iopub.execute_input":"2024-11-01T18:07:39.905946Z","iopub.status.idle":"2024-11-01T18:07:39.913423Z","shell.execute_reply.started":"2024-11-01T18:07:39.905900Z","shell.execute_reply":"2024-11-01T18:07:39.912031Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Optional Signal Denoising with Wavelet transform\nWe can optionally denoise the signal before creating the spectrogram. I'm not sure yet if this creates better or worse spectrograms. We can experiment with this. This code comes from Yusaku5738 notebook [here][1] and was suggested by SeshuRajuP in the comments. We have many parent functions to use for denoising. Yusaku5738 suggests using `wavelet = db8`.\n\n[1]: https://www.kaggle.com/code/yusaku5739/eeg-signal-denosing-using-wavelet-transform","metadata":{"papermill":{"duration":0.006935,"end_time":"2024-01-16T12:12:42.659252","exception":false,"start_time":"2024-01-16T12:12:42.652317","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pywt\nprint(\"The wavelet functions we can use:\")\n# print(pywt.wavelist(kind='continuous'))\nprint(pywt.wavelist(kind='discrete'))\n\nUSE_WAVELET = 'db8' #or \"db8\" or anything below","metadata":{"papermill":{"duration":0.540601,"end_time":"2024-01-16T12:12:43.206759","exception":false,"start_time":"2024-01-16T12:12:42.666158","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-11-01T18:07:39.914935Z","iopub.execute_input":"2024-11-01T18:07:39.915953Z","iopub.status.idle":"2024-11-01T18:07:39.928432Z","shell.execute_reply.started":"2024-11-01T18:07:39.915901Z","shell.execute_reply":"2024-11-01T18:07:39.926990Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# DENOISE FUNCTION\ndef maddest(d, axis=None):\n    return np.mean(np.absolute(d - np.mean(d, axis)), axis)\n\ndef denoise(x, wavelet='haar', level=1):    \n    coeff = pywt.wavedec(x, wavelet, mode=\"per\")\n    sigma = (1/0.6745) * maddest(coeff[-level])\n\n    uthresh = sigma * np.sqrt(2*np.log(len(x)))\n    coeff[1:] = (pywt.threshold(i, value=uthresh, mode='hard') for i in coeff[1:])\n\n    ret=pywt.waverec(coeff, wavelet, mode='per')\n    \n    return ret","metadata":{"papermill":{"duration":0.018392,"end_time":"2024-01-16T12:12:43.23335","exception":false,"start_time":"2024-01-16T12:12:43.214958","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-11-01T18:07:39.930269Z","iopub.execute_input":"2024-11-01T18:07:39.931159Z","iopub.status.idle":"2024-11-01T18:07:39.943115Z","shell.execute_reply.started":"2024-11-01T18:07:39.931099Z","shell.execute_reply":"2024-11-01T18:07:39.942026Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create Spectrograms with Librosa\nWe can use library librosa to create spectrograms. We will save them to disk. For each `eeg_id` we will make 1 spectrogram from the middle 50 seconds. We don't want to use more information than 50 seconds at a time because during test inference, we only have access to 50 seconds of EEG for each test `eeg_id`. We will create spectrograms of `size = 128x256 (freq x time)`.\n\nThe main function is \n\n    mel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, \n              n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)\n              \nLet's explain these variables.\n* `y` is the input time series signal\n* `sr` is the sampling frequency. In this competition EEG is sample 200 times per sec\n* `hop_length` produces image with `width = len(x)/hop_length`\n* `n_fft` controls vertical resolution and quality of spectrogram\n* `n_mels` produces image with `height = n_mels`\n* `fmin` is smallest frequency in our spectrogram\n* `fmax` is largest frequency in our spectrogram\n* `win_length` controls hortizonal resolution and quality of spectrogram","metadata":{"papermill":{"duration":0.007002,"end_time":"2024-01-16T12:12:43.247431","exception":false,"start_time":"2024-01-16T12:12:43.240429","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import librosa\n\ndef spectrogram_from_eeg(parquet_path, display=False):\n    \n    # LOAD MIDDLE 50 SECONDS OF EEG SERIES\n    eeg = pd.read_parquet(parquet_path)\n    middle = (len(eeg)-10_000)//2\n    eeg = eeg.iloc[middle:middle+10_000]\n    \n    # VARIABLE TO HOLD SPECTROGRAM\n    img = np.zeros((128,256,4),dtype='float32')\n    \n    if display: plt.figure(figsize=(10,7))\n    signals = []\n    for k in range(4):\n        COLS = FEATS[k]\n        \n        for kk in range(4):\n        \n            # COMPUTE PAIR DIFFERENCES\n            x = eeg[COLS[kk]].values - eeg[COLS[kk+1]].values\n\n            # FILL NANS\n            m = np.nanmean(x)\n            if np.isnan(x).mean()<1: x = np.nan_to_num(x,nan=m)\n            else: x[:] = 0\n\n            # DENOISE\n            if USE_WAVELET:\n                x = denoise(x, wavelet=USE_WAVELET)\n            signals.append(x)\n\n            # RAW SPECTROGRAM\n            mel_spec = librosa.feature.melspectrogram(y=x, sr=200, hop_length=len(x)//256, \n                  n_fft=1024, n_mels=128, fmin=0, fmax=20, win_length=128)\n\n            # LOG TRANSFORM\n            width = (mel_spec.shape[1]//32)*32\n            mel_spec_db = librosa.power_to_db(mel_spec, ref=np.max).astype(np.float32)[:,:width]\n\n            # STANDARDIZE TO -1 TO 1\n            mel_spec_db = (mel_spec_db+40)/40 \n            img[:,:,k] += mel_spec_db\n                \n        # AVERAGE THE 4 MONTAGE DIFFERENCES\n        img[:,:,k] /= 4.0\n        \n        if display:\n            plt.subplot(2,2,k+1)\n            plt.imshow(img[:,:,k],aspect='auto',origin='lower')\n            plt.title(f'EEG {eeg_id} - Spectrogram {NAMES[k]}')\n            \n    if display: \n        plt.show()\n        plt.figure(figsize=(10,5))\n        offset = 0\n        for k in range(4):\n            if k>0: offset -= signals[3-k].min()\n            plt.plot(range(10_000),signals[k]+offset,label=NAMES[3-k])\n            offset += signals[3-k].max()\n        plt.legend()\n        plt.title(f'EEG {eeg_id} Signals')\n        plt.show()\n        print(); print('#'*25); print()\n        \n    return img","metadata":{"papermill":{"duration":0.032361,"end_time":"2024-01-16T12:12:43.287163","exception":false,"start_time":"2024-01-16T12:12:43.254802","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-11-01T18:07:39.945728Z","iopub.execute_input":"2024-11-01T18:07:39.946496Z","iopub.status.idle":"2024-11-01T18:07:39.963096Z","shell.execute_reply.started":"2024-11-01T18:07:39.946458Z","shell.execute_reply":"2024-11-01T18:07:39.961746Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\nPATH = '/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/'\nDISPLAY = 4\nEEG_IDS = train.eeg_id.unique()\nall_eegs = {}\n\nfor i,eeg_id in enumerate(EEG_IDS):\n    if (i%100==0)&(i!=0): print(i,', ',end='')\n        \n    # CREATE SPECTROGRAM FROM EEG PARQUET\n    img = spectrogram_from_eeg(f'{PATH}{eeg_id}.parquet', i<DISPLAY)\n    \n    # SAVE TO DISK\n    if i==DISPLAY:\n        print(f'Creating and writing {len(EEG_IDS)} spectrograms to disk... ',end='')\n#     np.save(f'{directory_path}{eeg_id}',img)\n    all_eegs[eeg_id] = img\n   \n# SAVE EEG SPECTROGRAM DICTIONARY\nnp.save('eeg_specs_new',all_eegs)","metadata":{"papermill":{"duration":66.745724,"end_time":"2024-01-16T12:13:50.040166","exception":false,"start_time":"2024-01-16T12:12:43.294442","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-11-01T18:07:39.964868Z","iopub.execute_input":"2024-11-01T18:07:39.965341Z","iopub.status.idle":"2024-11-01T18:07:42.866249Z","shell.execute_reply.started":"2024-11-01T18:07:39.965308Z","shell.execute_reply":"2024-11-01T18:07:42.865085Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%cd /kaggle/working\nfrom IPython.display import FileLink\nFileLink('.tgz')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-01T18:07:42.867652Z","iopub.execute_input":"2024-11-01T18:07:42.868077Z","iopub.status.idle":"2024-11-01T18:07:42.878143Z","shell.execute_reply.started":"2024-11-01T18:07:42.868041Z","shell.execute_reply":"2024-11-01T18:07:42.877070Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%cd /kaggle/working\n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-01T18:07:42.879686Z","iopub.execute_input":"2024-11-01T18:07:42.880863Z","iopub.status.idle":"2024-11-01T18:07:43.981023Z","shell.execute_reply.started":"2024-11-01T18:07:42.880812Z","shell.execute_reply":"2024-11-01T18:07:43.979382Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from IPython.display import FileLink\nFileLink(r'eeg_specs.npy')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-01T18:07:43.983296Z","iopub.execute_input":"2024-11-01T18:07:43.983734Z","iopub.status.idle":"2024-11-01T18:07:43.992227Z","shell.execute_reply.started":"2024-11-01T18:07:43.983696Z","shell.execute_reply":"2024-11-01T18:07:43.990925Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Kaggle Dataset\nThe new EEG spectrograms from version 4 of this notebook have been uploaded to a Kaggle dataset [here][4]. We can attach this Kaggle dataset to our future notebooks to boost our CV scores and LB scores! Thank you everyone for upvoting my new EEG spectrogram Kaggle dataset!\n\nExamples of how to use EEG spectrograms to boost CV score and LB score will be (or already are) published in recent versions of my EfficientNet starter notebook [here][2] and CatBoost starter notebook [here][3]\n\nEnjoy! Happy Kaggling!\n\n[2]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[3]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[4]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}