{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction to EEG Spectrogram Generation\n\nIn this notebook, we focus on the generation of spectrograms from EEG data with the MNE library. EEG recordings capture the electrical activity of the brain, offering valuable insights into neurological phenomena such as seizures, cognitive processes, and sleep stages.\n\nOur main goal is to establish a reliable pipeline for transforming raw EEG signals into spectrograms. Spectrograms provide a visual depiction of the frequency distribution of EEG signals over time, enabling the identification of patterns, anomalies, and unique features in the data.\n\nThroughout this notebook, we harness the power of the MNE Python library, renowned for its extensive capabilities in EEG/MEG data analysis. Leveraging MNE, we will explore various montage configurations, define frequency ranges, and utilize wavelet transforms to generate spectrograms.\n\nFor an introductory experience with MNE, you can start with my first notebook available [here](https://www.kaggle.com/code/theeventhorizons/understanding-the-basics-of-mne).","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport tqdm as tqdm\nimport mne\nfrom mne.time_frequency import tfr_morlet\nimport cv2\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\ndata_train_path = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:36:48.339736Z","iopub.execute_input":"2024-03-29T14:36:48.340199Z","iopub.status.idle":"2024-03-29T14:36:48.346783Z","shell.execute_reply.started":"2024-03-29T14:36:48.340139Z","shell.execute_reply":"2024-03-29T14:36:48.345565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read the data","metadata":{}},{"cell_type":"code","source":"# Read the data\ndata = pd.read_csv(data_train_path)\n\n# Copy the Data\ndf = data.copy()\n\n# Observe few lines \nprint(df.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:25.599688Z","iopub.execute_input":"2024-03-29T10:11:25.600257Z","iopub.status.idle":"2024-03-29T10:11:25.920079Z","shell.execute_reply.started":"2024-03-29T10:11:25.600221Z","shell.execute_reply":"2024-03-29T10:11:25.919117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Shape of the data\nprint('The shape of df is:', df.shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:25.923805Z","iopub.execute_input":"2024-03-29T10:11:25.924325Z","iopub.status.idle":"2024-03-29T10:11:25.931823Z","shell.execute_reply.started":"2024-03-29T10:11:25.924280Z","shell.execute_reply":"2024-03-29T10:11:25.930385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Target\nTarget = {'Seizure':0, 'LPD':1, 'GPD':2, 'LRDA':3, 'GRDA':4, 'Other':5}","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:25.934525Z","iopub.execute_input":"2024-03-29T10:11:25.935505Z","iopub.status.idle":"2024-03-29T10:11:25.945068Z","shell.execute_reply.started":"2024-03-29T10:11:25.935468Z","shell.execute_reply":"2024-03-29T10:11:25.944219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Montages Used\n\nIn EEG (Electroencephalography), a montage refers to the configuration of electrodes used to record brain activity. It's crucial for manipulating raw EEG data to enhance signal quality and extract specific information. One common goal of montage design is to reduce noise and artifacts in the EEG signal, improving clarity and accuracy. \n\nMontages come in various types, including bipolar and referential configurations, which can augment the number of spectrograms generated (data generation) for the future machine learning models. Here, we introduce some of these montages, although we'll only utilize the double banana montage in subsequent notebooks.","metadata":{}},{"cell_type":"markdown","source":"- ## The Bipolar montages\n\nBipolar montages are used in electroencephalography (EEG) to detect and record the differences in electrical potential between two adjacent electrodes on the scalp. They are often used to capture local variations in brain activity and to identify sources of abnormal electrical activity in the brain.\n\nIn general, bipolar montages are useful for:\n\n1. Precise localization: They allow for the detection of subtle changes in electrical potential between two specific points on the scalp, which can help accurately localize abnormal brain activity associated with conditions such as epilepsy or other neurological disorders.\n\n2. Noise reduction: By subtracting the signals recorded by two adjacent electrodes, bipolar montages can reduce artifacts and noise that may interfere with EEG analysis, thereby improving the overall quality of the recorded signal.\n\n3. Differentiation of activity sources: By measuring the potential differences between specific pairs of electrodes, bipolar montages enable the differentiation of different sources of brain activity and mapping their spatial distribution on the scalp.\n","metadata":{}},{"cell_type":"markdown","source":"### 1. The Bipolar double banana montage or longitudinal montage\n\nThe \"double banana\" montage, also known as the longitudinal bipolar montage, is a widely used configuration in EEG electrode placement. In this montage, electrodes are arranged in series, and the channels of electrodes placed in series are subtracted from each other. This arrangement offers versatility and is commonly employed in EEG recordings.\n\nThe placement of electrodes and the selection of electrodes in the montage can significantly influence EEG readings. Electrodes are typically positioned according to the 10-20 system, where they are evenly spaced between anatomical landmarks such as the nasion, inion, and preauricular points. Adjusting the distance between electrodes and the number of electrodes used aims to effectively localize the seizure onset zone (SOZ) in EEG recordings.\n\nEEGs play a crucial role not only in diagnosing epilepsy but also in identifying epileptic brain activity patterns and determining the type of epilepsy. Different types of epilepsy exhibit distinct characteristics, including generalized or localized seizures, and may be idiopathic, cryptogenic, or symptomatic. These epileptic activities are associated with various brainwave patterns, such as delta, theta, alpha, and beta waves. Ensuring accuracy in EEG diagnostics involves calibrating montages to reduce noise and artifacts, with evaluations based on band power.","metadata":{}},{"cell_type":"code","source":"NAMES_BANANA = ['LL','LP','RP','RR','BC']\n\nFEATS_BANANA = [['Fp1','F7','T3','T5','O1'],\n                ['Fp1','F3','C3','P3','O1'],\n                ['Fp2','F8','T4','T6','O2'],\n                ['Fp2','F4','C4','P4','O2'],\n                ['Fz','Cz','Pz']]\n\nDIC_BANANA = {}\nfor name, feat in zip(NAMES_BANANA,FEATS_BANANA):\n    DIC_BANANA[name]=feat\n\n\nDIC_BANANA_SPEC = {}\nfor name, feat in zip(NAMES_BANANA, FEATS_BANANA):\n    ch_names = []\n    for kk in range(len(feat)-1):\n        # COMPUTE PAIR DIFFERENCES\n        ch_names.append(f'{feat[kk]}-{feat[kk+1]}')\n    DIC_BANANA_SPEC[name] = ch_names\n\nprint(DIC_BANANA_SPEC)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:25.946542Z","iopub.execute_input":"2024-03-29T10:11:25.947656Z","iopub.status.idle":"2024-03-29T10:11:25.960000Z","shell.execute_reply.started":"2024-03-29T10:11:25.947616Z","shell.execute_reply":"2024-03-29T10:11:25.958966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. The transverse bipolar montage\n\nThe transverse montage, a form of bipolar montage, aims to localize the peak discharge in the left-to-right direction. It involves linking adjacent electrodes in a chain starting from electrode F7. While effective in certain localization tasks, transverse montages have drawbacks. In-phase cancellation may occur when synchronous waveforms lead to false localization. Additionally, bipolar montages, including the transverse type, are prone to the end-of-chain phenomenon, where the last electrode in a chain produces a downward deflection without a phase reversal. To address this, three chains were grouped instead of five.","metadata":{}},{"cell_type":"code","source":"NAMES_TRANSVERSE = ['Tr1','Tr2','Tr3','Tr4','Tr5']\n\nFEATS_TRANSVERSE = [['F7','Fp1','Fp2','F8'],\n                    ['F7','F3','Fz', 'F4','F8'],\n                    ['T3','C3','Cz','C4','T4'],\n                    ['T5','P3','Pz','P4','T6'],\n                    ['T5','O1','O2','T6']]\n\nDIC_TRANSVERSE = {}\nfor name, feat in zip(NAMES_TRANSVERSE,FEATS_TRANSVERSE):\n    DIC_TRANSVERSE[name]=feat\n\n\nDIC_TRANSVERSE_SPEC = {}\nfor name, feat in zip(NAMES_TRANSVERSE, FEATS_TRANSVERSE):\n    ch_names = []\n    for kk in range(len(feat)-1):\n        # COMPUTE PAIR DIFFERENCES\n        ch_names.append(f'{feat[kk]}-{feat[kk+1]}')\n    DIC_TRANSVERSE_SPEC[name] = ch_names\n\nprint(DIC_TRANSVERSE_SPEC)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:25.961321Z","iopub.execute_input":"2024-03-29T10:11:25.962435Z","iopub.status.idle":"2024-03-29T10:11:25.972284Z","shell.execute_reply.started":"2024-03-29T10:11:25.962393Z","shell.execute_reply":"2024-03-29T10:11:25.971348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3. The Crossed Bipolar montage","metadata":{}},{"cell_type":"code","source":"NAMES_CROSSED = ['Cr1','Cr2','Cr3','Cr4','Cr5']\n\nFEATS_CROSSED = [['Fp1','Fp2'],\n                 ['F7','F3','Fz','F4','F8'],\n                 ['T3', 'C3', 'Cz', 'C4', 'T4'],\n                 ['T5', 'P3', 'Pz', 'P4', 'T6'],\n                 ['O1','O2']]\n\nDIC_CROSSED = {}\nfor name, feat in zip(NAMES_CROSSED,FEATS_CROSSED):\n    DIC_CROSSED[name]=feat\n\n\nDIC_CROSSED_SPEC = {}\nfor name, feat in zip(NAMES_CROSSED, FEATS_CROSSED):\n    ch_names = []\n    for kk in range(len(feat)-1):\n        # COMPUTE PAIR DIFFERENCES\n        ch_names.append(f'{feat[kk]}-{feat[kk+1]}')\n    DIC_CROSSED_SPEC[name] = ch_names\n\nprint(DIC_CROSSED_SPEC)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:25.973647Z","iopub.execute_input":"2024-03-29T10:11:25.974708Z","iopub.status.idle":"2024-03-29T10:11:25.984397Z","shell.execute_reply.started":"2024-03-29T10:11:25.974666Z","shell.execute_reply":"2024-03-29T10:11:25.982936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4. Circumferential montage \n\n\nThe circumferential montage is a configuration of electrodes used in electroencephalography (EEG). In this montage, the electrodes are placed around the subject's head, forming a \"circumference\" or ring. The main purpose of this montage is to capture the brain's electrical activity in a global manner by recording electrical signals from different regions of the cerebral cortex.\n\nThis montage provides a panoramic view of brain activity, which can be useful in certain applications such as general monitoring of brain activity or detecting abnormal electrical activity patterns across the entire cortex. However, due to its circular arrangement and limited number of electrodes, the circumferential montage may not be as effective as other more targeted montages for precisely localizing specific sources of brain activity.","metadata":{}},{"cell_type":"code","source":"NAMES_CIRCUMFERENTIAL = ['CE','CI']\n\nFEATS_CIRCUMFERENTIAL = [['Fp1','F7','T3','T5','O1','O2','T6','T4','F8','Fp2'],\n                         ['F3','C3','P3','Pz','P4','C4','F4','Fz']]\n\nDIC_CIRCUMFERENTIAL = {}\nfor name, feat in zip(NAMES_CIRCUMFERENTIAL,FEATS_CIRCUMFERENTIAL):\n    DIC_CIRCUMFERENTIAL[name]=feat\n\n\nDIC_CIRCUMFERENTIAL_SPEC = {}\nfor name, feat in zip(NAMES_CIRCUMFERENTIAL, FEATS_CIRCUMFERENTIAL):\n    ch_names = []\n    for kk in range(len(feat)-1):\n        # COMPUTE PAIR DIFFERENCES\n        ch_names.append(f'{feat[kk]}-{feat[kk+1]}')\n    DIC_CIRCUMFERENTIAL_SPEC[name] = ch_names\n\nprint(DIC_CIRCUMFERENTIAL_SPEC)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:25.986040Z","iopub.execute_input":"2024-03-29T10:11:25.986725Z","iopub.status.idle":"2024-03-29T10:11:26.002080Z","shell.execute_reply.started":"2024-03-29T10:11:25.986683Z","shell.execute_reply":"2024-03-29T10:11:26.000767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- ## Referential montages\n\nReferential montages are used in electroencephalography (EEG) to record and analyze the differences in electrical potential between a reference electrode and other electrodes placed on the scalp. Unlike bipolar montages where each pair of electrodes is used to measure potential difference, referential montages compare each electrode to a single reference electrode.\n\nReferential montages serve several purposes:\n\n- Noise reduction: By using a common reference electrode to compare the signals recorded by other electrodes, referential montages can help attenuate artifacts and noise that may be present in the EEG signal.\n\n- Comparative analysis: Referential montages allow for comparison of the recorded electrical activity at different positions on the scalp relative to a fixed reference electrode. This can be useful for examining spatial variations in brain activity and identifying regions of the brain showing significant changes.\n\n- Study of cerebral asymmetry: By using a reference electrode on one side of the scalp, referential montages can be used to study differences in electrical activity between the left and right cerebral hemispheres, which may be relevant in certain studies of cognition and brain function.\n\n- Flexibility in reference electrode placement: Referential montages offer the option to choose the placement of the reference electrode based on the study's needs or clinician's preferences, allowing for some flexibility in EEG experiment design.\n\nHere we will introduce only the Laplacian montage.\n\n### The Laplacian montage\n\nThe Laplacian montage serves as a valuable tool in EEG analysis, particularly in cases of focal abnormal brain activity. Unlike common and average reference montages, which can suffer from referential contamination due to outliers in electrode potentials, the Laplacian montage minimizes this issue. It achieves this by using only the average electrical potential of the nearest electrodes as a reference. This approach enhances the accuracy of EEG readings, making it especially effective for pinpointing localized abnormalities and identifying the Seizure Onset Zone (SOZ) in epilepsy diagnosis.","metadata":{}},{"cell_type":"code","source":"NAMES_LAPLACIAN = ['LLL', 'LLP', 'LRP', 'LRR']\n\nFEATS_LAPLACIAN = [['Fp1', 'F7', 'C3', 'Fz', 'F3'],\n                   ['Fp2', 'F8', 'C4', 'Fz', 'F4'],\n                   ['O1', 'T5', 'C3', 'Pz', 'P3'],\n                   ['O2', 'T6', 'C4', 'Pz', 'P4']]\n\nDIC_LAPLACIAN = {}\nfor name, feat in zip(NAMES_LAPLACIAN, FEATS_LAPLACIAN):\n    DIC_LAPLACIAN[name] = feat\n\nDIC_LAPLACIAN_SPEC = {}\nfor name, feat in zip(NAMES_LAPLACIAN, FEATS_LAPLACIAN):\n    ch_names = []\n    for kk in range(len(feat) - 4):  # Loop until the last 4 elements of feat\n        # COMPUTE SUM\n        ch_names.append(f'{feat[kk]}+{feat[kk + 1]}+{feat[kk + 2]}+{feat[kk + 3]}+{feat[kk + 4]}')\n    DIC_LAPLACIAN_SPEC[name] = ch_names\n\nprint(DIC_LAPLACIAN_SPEC)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:26.003678Z","iopub.execute_input":"2024-03-29T10:11:26.004517Z","iopub.status.idle":"2024-03-29T10:11:26.016417Z","shell.execute_reply.started":"2024-03-29T10:11:26.004474Z","shell.execute_reply":"2024-03-29T10:11:26.015197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#MONTAGES_BIP = [DIC_BANANA, DIC_TRANSVERSE, DIC_CROSSED, DIC_CIRCUMFERENTIAL]\n\ndirectory_path = 'EEG_Spectrograms/'\nif not os.path.exists(directory_path):\n    os.makedirs(directory_path)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:26.019816Z","iopub.execute_input":"2024-03-29T10:11:26.020490Z","iopub.status.idle":"2024-03-29T10:11:26.026829Z","shell.execute_reply.started":"2024-03-29T10:11:26.020456Z","shell.execute_reply":"2024-03-29T10:11:26.025829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Select EEG\n\nHere we are working with a DataFrame containing evaluations from various evaluators. Our objective is to add additional columns to the DataFrame to analyze the data more effectively. Firstly, we compute the sum of evaluations from six specified columns to create a new column called 'total_evaluators'. Next, we modify the previous code to add an additional column named 'consensus_column' to the DataFrame. This column holds the name of the column with the highest number for each row, indicating the consensus among the evaluators. Finally, we filter the DataFrame to include only rows with a agreement percentage greater than 30% and we create individual DataFrames for each EEG ID, selecting non-overlapping EEG events with a minimum separation of 10 seconds between them.","metadata":{}},{"cell_type":"code","source":"# Adding a new column 'total_evaluators' that sums up the six specified columns\ndf.loc[:,'total_evaluators'] = df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].sum(axis=1)\nprint(df.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:26.028041Z","iopub.execute_input":"2024-03-29T10:11:26.028378Z","iopub.status.idle":"2024-03-29T10:11:26.071635Z","shell.execute_reply.started":"2024-03-29T10:11:26.028351Z","shell.execute_reply":"2024-03-29T10:11:26.070094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Modifying the previous code to add an additional column 'consensus_column' to 'df'\n\n# Finding the column with the largest number for each row and storing the value in 'consensus'\ndf.loc[:,'consensus'] = df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].max(axis=1)\n\n# Identifying the column name that corresponds to the max value for each row\ndf.loc[:,'consensus_column'] = df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].idxmax(axis=1)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:26.073232Z","iopub.execute_input":"2024-03-29T10:11:26.073573Z","iopub.status.idle":"2024-03-29T10:11:26.123136Z","shell.execute_reply.started":"2024-03-29T10:11:26.073545Z","shell.execute_reply":"2024-03-29T10:11:26.121861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create a new column that shows the percentage agreement\ndf.loc[:,'row_agreement'] = df['consensus']/df['total_evaluators']","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:26.124826Z","iopub.execute_input":"2024-03-29T10:11:26.125622Z","iopub.status.idle":"2024-03-29T10:11:26.133646Z","shell.execute_reply.started":"2024-03-29T10:11:26.125547Z","shell.execute_reply":"2024-03-29T10:11:26.132584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Delete rows for which the DataFrame has less than 30% of row_agreement\ndf = df[df['row_agreement']>0.3]\nprint(df)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:26.135029Z","iopub.execute_input":"2024-03-29T10:11:26.135840Z","iopub.status.idle":"2024-03-29T10:11:26.177090Z","shell.execute_reply.started":"2024-03-29T10:11:26.135799Z","shell.execute_reply":"2024-03-29T10:11:26.175874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define a dictionnary of non-overlapping events of ten seconds\ndf_dict={}\n\nfor eeg_id in df['eeg_id'].unique():\n\n    # Create a DataFrame for each 'eeg_id' and store it in the dictionary\n    df_dict[f'df_{eeg_id}'] = df[df['eeg_id'] == eeg_id].copy()\n\n    # Access the DataFrame for the current 'eeg_id'\n    df_eeg = df_dict[f'df_{eeg_id}'].copy()\n\n    # Create the non overlapping events\n    events = df_eeg[['eeg_label_offset_seconds']]\n    non_overlapping_events = pd.DataFrame(columns=events.columns)\n\n    # Create the list of all eeg_label_offset_seconds\n    list_eeg_label_offset_seconds = list(df_eeg['eeg_label_offset_seconds'])\n    #print(list_eeg_label_offset_seconds)\n\n    # Generate a list of eeg_label_offset_seconds with a minimum separation of 10 seconds   \n    non_overlapping_mask=[]\n    current_offset = 0\n    min_distance = 10\n\n    while current_offset <= max(list_eeg_label_offset_seconds):\n        non_overlapping_mask.append(current_offset)\n        next_offset = next((x for x in list_eeg_label_offset_seconds if x >= current_offset + min_distance), None)\n        if next_offset is None:\n            break\n        current_offset = next_offset\n\n    # Create a mask using the non_overlapping_mask\n    mask = df_eeg['eeg_label_offset_seconds'].isin(non_overlapping_mask)\n\n    # Apply the mask to get non-overlapping events\n    df_dict[f'df_{eeg_id}'] = df_eeg[mask]\n","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:11:26.185829Z","iopub.execute_input":"2024-03-29T10:11:26.186190Z","iopub.status.idle":"2024-03-29T10:12:07.677673Z","shell.execute_reply.started":"2024-03-29T10:11:26.186143Z","shell.execute_reply":"2024-03-29T10:12:07.676262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the non overlapping events\nevents = df_eeg[['eeg_label_offset_seconds']]\nnon_overlapping_events = pd.DataFrame(columns=events.columns)\n\n# Create the list of all eeg_label_offset_seconds\nlist_eeg_label_offset_seconds = list(df['eeg_label_offset_seconds'])\n#print(list_eeg_label_offset_seconds)\n\n# Generate a list of eeg_label_offset_seconds with a minimum separation of 10 seconds   \nnon_overlapping_mask=[]\ncurrent_offset = 0\nmin_distance = 10\n\nwhile current_offset <= max(list_eeg_label_offset_seconds):\n    non_overlapping_mask.append(current_offset)\n    next_offset = next((x for x in list_eeg_label_offset_seconds if x >= current_offset + min_distance), None)\n    if next_offset is None:\n        break\n    current_offset = next_offset\n\n# Create a mask using the non_overlapping_mask\nmask = df['eeg_label_offset_seconds'].isin(non_overlapping_mask)\n\ndf = df[mask]\nprint(df)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:38:13.844516Z","iopub.execute_input":"2024-03-29T14:38:13.845019Z","iopub.status.idle":"2024-03-29T14:38:14.308164Z","shell.execute_reply.started":"2024-03-29T14:38:13.844986Z","shell.execute_reply":"2024-03-29T14:38:14.306526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create objects\n\nIn this portion of code, we are working with EEG data using a predefined montage configuration represented by the dictionaries `NAMES_MONTAGE` and `FEATS_MONTAGE`. Firstly, we construct a dictionary called `DIC_MONTAGE`, where each key represents a specific montage name associated with its corresponding list of channels. Next, we define another dictionary named `DIC_MONTAGE_SPEC`, where each montage name is associated with a list of channel pairs representing the computed pair differences.\n\nThe code then proceeds to create EEG objects using the `create_objects` function, which reads EEG data from Parquet files, preprocesses it, computes pair differences, and constructs MNE RawArray and Epochs objects for each montage configuration. The function returns a dictionary containing the epochs for each montage.\n\nFinally, the code plots the first four epochs for all channels using the `plot` method (as an example).\n\nOverall, this code segment demonstrates the preprocessing and visualization steps involved in analyzing EEG data using MNE-Python.","metadata":{}},{"cell_type":"markdown","source":"## Creation of the main function\n\nThis function, `create_objects`, is designed to create and return a dictionary of epochs based on the `eeg_id`, `events`, and the montage specified in the `DIC_MONTAGE`.\n\nHere's a breakdown of what the function does:\n\n1. **Read EEG Data**: It reads the EEG data from the corresponding parquet file based on the given `eeg_id`.\n\n2. **Preprocessing**: It drops the EKG column from the EEG data.\n\n3. **Compute Pair Differences**: For each pair of channels in the specified montage (`DIC_MONTAGE`), it computes the pair differences and loads them into a matrix.\n\n4. **Create MNE Raw Objects**: It creates MNE raw objects for each pair of channels, specifying the channel names, types, and sampling frequency. These raw objects are used to store the EEG data.\n\n5. **Create Epochs**: It creates epochs from the raw objects based on the provided events and event IDs. Each epoch represents a segment of EEG data centered around an event, with a specified time window (`tmin` to `tmax`).\n\n6. **Return**: It returns a dictionary (`dic_epoch`) containing the epochs for each montage specified in `DIC_MONTAGE`.\n\nOverall, this function is responsible for preparing the EEG data, computing pair differences, creating MNE raw objects, and extracting epochs for further analysis and processing.","metadata":{}},{"cell_type":"code","source":"#CHANGER EVENTS EET EVENTIDDS EN EVENT_DICT\n# Define a function which return a dictionnary of epochs based on eeeg_id, events and the montage \ndef create_objects(eeg_id, event_dict, DIC_MONTAGE):\n    \n    # Read the corresponding eeg\n    data = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet').copy()\n    \n    # Drop the EKG column\n    data = data.drop(['EKG'], axis=1)\n    \n    # Compute pair differences and load it in a matrix\n    dic_raw = {}\n    dic_epoch={}\n    \n    for name, COLS in DIC_MONTAGE.items():\n        raw=[] \n        ch_names = []\n        for kk in range(len(COLS)-1):\n            # COMPUTE PAIR DIFFERENCES\n            ch_names.append(f'{COLS[kk]}-{COLS[kk+1]}')\n            x = data[COLS[kk]].values - data[COLS[kk+1]].values\n            # FILL NANS\n            m = np.nanmean(x)\n            if np.isnan(x).mean()<1: x = np.nan_to_num(x,nan=m)\n            else: x[:] = 0\n            # Put it in the matrix\n            raw.append(x)\n        ch_types = ['misc']*(len(ch_names)) \n        info = mne.create_info(ch_names, ch_types=ch_types, sfreq=200,verbose=False)\n        raw_array = np.array(raw)\n        raw_object = mne.io.RawArray(raw_array, info,verbose=False)\n        dic_raw[name] = raw_object\n        dic_epoch[name]= mne.Epochs(raw_object, event_dict[eeg_id]['events'], event_id = event_dict[eeg_id]['event_ids'], tmin=-5, tmax=5, preload=True, baseline=(None, 0),verbose=False)\n    return dic_epoch","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:12:09.210689Z","iopub.execute_input":"2024-03-29T10:12:09.211072Z","iopub.status.idle":"2024-03-29T10:12:09.222088Z","shell.execute_reply.started":"2024-03-29T10:12:09.211043Z","shell.execute_reply":"2024-03-29T10:12:09.220738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Example\n\nThis example shows how we prepare event data for EEG analysis and visualize epochs for the 'LL' montage. It first creates event IDs based on unique activities, then creates a DataFrame with event information. After mapping expert consensus to numerical labels, it defines the sample time point for events. Finally, it converts the DataFrame to an array and plots the first epoch for the 'LL' montage.","metadata":{}},{"cell_type":"code","source":"event_dict = {}\nEEG_IDS = df.eeg_id.unique()\n\nfor eeg_id in EEG_IDS:\n    event_ids = {}\n    # Create event ids for this eeg\n    for activity in df_dict[f'df_{eeg_id}']['expert_consensus'].unique():\n        event_ids[activity] = Target[activity]\n\n    # Create the events\n    events = df_dict[f'df_{eeg_id}'][['eeg_label_offset_seconds', 'expert_consensus']]\n    events.insert(1, 'New', 0)\n\n    # Map 'expert_consensus' values to numerical labels using the 'event_ids' dictionary\n    events.loc[:,'expert_consensus'] = events['expert_consensus'].map(event_ids)\n\n    # Define the sample time point where the event occurs\n    events.loc[:,'eeg_label_offset_seconds'] = (events['eeg_label_offset_seconds']+25)*200\n\n    # Convert the data frame into an array\n    events = events.values.astype(int)\n    \n    # Add the events for this EEG to the event dictionary\n    event_dict[eeg_id] = {'event_ids': event_ids, 'events': events}","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:12:59.494002Z","iopub.execute_input":"2024-03-29T10:12:59.494663Z","iopub.status.idle":"2024-03-29T10:13:35.725459Z","shell.execute_reply.started":"2024-03-29T10:12:59.494620Z","shell.execute_reply":"2024-03-29T10:13:35.723977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the first epochs for all channels.\ncreate_objects(1628180742, event_dict, DIC_BANANA)['LL'][0].plot(n_epochs=4, events=True, picks = 'all', show_scrollbars=False, show_scalebars=False, scalings=300)\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-29T10:13:42.675544Z","iopub.execute_input":"2024-03-29T10:13:42.675980Z","iopub.status.idle":"2024-03-29T10:13:43.107930Z","shell.execute_reply.started":"2024-03-29T10:13:42.675948Z","shell.execute_reply":"2024-03-29T10:13:43.106750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Generating Spectrograms with MNE\n\n\nIn this section, we are creating a function to convert EEG data into spectrograms. We will demonstrate this process with an example and then proceed to save all the spectrograms. \n\nEach spectrogram will be resized to 128x256 pixels and associated with each montage configuration and epoch (it's my wish but I couldn't with MNE time frequency function, if someone knows how to do that it will be nice). These spectrograms will be stored in a dictionary where the keys correspond to the EEG IDs. Each spectrogram will have a shape of (128x256 x number of keys in the montage dictionary x number of epochs).\n","metadata":{}},{"cell_type":"markdown","source":"## Converting EEG into spectrograms\n\n\nThe `spectrogram_from_eeg` function generates spectrograms from EEG data epochs provided in a dictionary format. It takes parameters such as `dic_epoch` (containing EEG epochs), `NAME_MONTAGE` (the name of the montage configuration), `num_epoch` (the index of the epoch to analyze), and `display` (a boolean indicating whether to display the spectrogram). \n\n1. **Frequency and Cycle Definition:**\n   - The function defines the range of frequencies and the number of cycles for the Morlet wavelet transform.\n\n2. **Morlet Wavelet Transform:**\n   - It performs the Morlet wavelet transform on the EEG epoch associated with the event 'seizure' using the `tfr_morlet` function from the MNE library.\n\n3. **Spectrogram Calculation:**\n   - The spectrogram is computed for each EEG channel specified in the `DIC_MONTAGE_SPEC` dictionary.\n   - Each spectrogram is normalized and stored in a list.\n\n4. **Image Resizing:**\n   - The spectrograms are averaged and resized to a lower resolution using OpenCV's `cv2.resize` function.\n\n5. **Visualization:**\n   - The resized spectrogram is displayed using Matplotlib.\n   - A colorbar indicates the power intensity.\n   - The function saves the spectrogram image as 'img.png'.\n\n6. **Return:**\n   - The function returns the resized spectrogram image.","metadata":{}},{"cell_type":"code","source":"def spectrogram_from_eeg(eeg_id, dic_epoch, DIC_MONTAGE_SPEC, display=False):\n    # Define the range of frequencies\n    freq = np.arange(0.1, 20, 0.01)\n\n    # Define the number of cycles\n    n_cycles = freq/2\n\n\n    # Initialize the list to store the spectrograms\n    spectrograms = []\n\n    # Define the height and width of the resized image\n    height = 128\n    width = 256\n\n    num_epoch = df[df['eeg_id']==1628180742].shape[0]\n\n    # VARIABLE TO HOLD SPECTROGRAM\n    img = np.zeros((1990,2001, len(DIC_BANANA.keys()), num_epoch), dtype=np.float32)\n\n\n\n    # Compute the spectrograms for each channel\n    for j in range(num_epoch):\n        index=0\n        for name in DIC_BANANA_SPEC.keys():\n\n            #  Define the Morlet wavelet transform for the epoch num_epoch associated to the event NAME_MONTAGE\n            power_db = tfr_morlet(create_objects(1628180742, event_dict, DIC_BANANA)[name][j], freq, n_cycles = n_cycles, return_itc = False, picks='all', verbose=False)\n\n            for title in range(len(DIC_BANANA_SPEC[name])):\n\n                # Get the data array of the spectrogram\n                spectrogram = power_db.data[title] \n\n                # Normalize the spectrogram\n                spectrogram = (spectrogram + 40) / 40  \n                img[:,:,index,j] += spectrogram\n\n            # AVERAGE THE 4 MONTAGE DIFFERENCES\n            img[:,:,index,j] /= len(DIC_BANANA.keys())\n\n\n\n\n            if display:\n                # Display the resized spectrogram\n                plt.figure()\n                plt.imshow(img[:,:,index,j], aspect='auto', origin='lower')\n                plt.colorbar(label='Power')\n                plt.xlabel('Time')\n                plt.ylabel('Frequency')\n                plt.title(f'EEG {1628180742} {name} {j}')\n                plt.show()\n            index += 1\n    return img","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:28:03.764212Z","iopub.status.idle":"2024-03-29T14:28:03.764629Z","shell.execute_reply.started":"2024-03-29T14:28:03.764436Z","shell.execute_reply":"2024-03-29T14:28:03.764456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Example\n\nThis code generates and displays a spectrogram for EEG data associated with EEG ID `1628180742`. It uses the `create_objects` function to prepare EEG epochs and then calls `spectrogram_from_eeg` with specific parameters, including the montage configuration 'LL' and displaying the spectrogram.","metadata":{}},{"cell_type":"code","source":"spectrogram_from_eeg(1628180742, create_objects(1628180742, event_dict, DIC_BANANA),DIC_BANANA_SPEC, display=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:53:58.843478Z","iopub.execute_input":"2024-03-29T14:53:58.844061Z","iopub.status.idle":"2024-03-29T14:54:56.685869Z","shell.execute_reply.started":"2024-03-29T14:53:58.844023Z","shell.execute_reply":"2024-03-29T14:54:56.684666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\n\n#EEG_IDS = df.eeg_id.unique()\nEEG_IDS = [1628180742, 351917269]\nall_eegs = {}\n\nprint('Starting the creation of the dictionary')\n\n# Utiliser tqdm pour visualiser la progression de la boucle\nfor eeg_id in tqdm(EEG_IDS):\n    print(f'Starting the creation of spectrograms associated to the eeg_id {eeg_id}')    \n\n    # Create the necessary objects to generate the spectrograms\n    create_objects(eeg_id, event_dict, DIC_BANANA)\n\n    img = spectrogram_from_eeg(eeg_id, create_objects(eeg_id, event_dict, DIC_BANANA), DIC_BANANA_SPEC, display=False).astype(np.float32)\n\n    # Save the spectrogram array\n    np.save(f'{directory_path}{eeg_id}', img)\n\n    # Add the array of spectrograms for this EEG to the global dictionary\n    all_eegs[eeg_id] = img\n    print(f'Spectrograms associated to the eeg_id {eeg_id} saved')\n\n# Save the dictionary containing all the spectrograms\nnp.save('eeg_specs', all_eegs)\n\nprint('Dictionary saved')\n","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:44:52.529417Z","iopub.execute_input":"2024-03-29T14:44:52.529876Z","iopub.status.idle":"2024-03-29T14:46:24.577216Z","shell.execute_reply.started":"2024-03-29T14:44:52.529845Z","shell.execute_reply":"2024-03-29T14:46:24.575932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\n\nIt appears that the process I employed to create the objects and spectrograms is time-consuming and not optimal. This slowness can largely be attributed to the significant size of the spectrograms, measuring (1990, 2001), which requires a substantial amount of resources for processing. Exploring alternatives or optimizing our current approach could greatly enhance efficiency and reduce the time required to perform these tasks.\n\nI am open to any suggestions or ideas for improving the process of creating spectrograms or optimizing their processing. Your knowledge and expertise could help us find more effective solutions and optimize the spectrogram creation process with MNE. Feel free to share your ideas, as they could significantly contribute to our goal of improvement.","metadata":{}},{"cell_type":"markdown","source":"# Acknowledgement \n\nA significant portion of this notebook draws inspiration from the work of [@cdeotte](https://www.kaggle.com/code/cdeotte). Their insightful notebooks, particularly [\"How to Make Spectrogram from EEG\"](https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg/notebook) served as valuable references in developing the spectrogram generation process using Librosa. We extend our gratitude to @cdeotte for their contributions to the EEG analysis community on Kaggle.","metadata":{}},{"cell_type":"markdown","source":"# External References\n\n- https://www.youtube.com/watch?time_continue=5&v=l7EbPvIFWYg&embeds_referring_euri=https%3A%2F%2Ftheinformaticists.com%2F&source_ve_path=Mjg2NjY&feature=emb_logo\n- https://theinformaticists.com/2020/08/25/developing-and-testing-new-montage-methods-in-electroencephalography/\n- https://mne.tools/stable/index.html\n","metadata":{}},{"cell_type":"markdown","source":"# Thank you for reading! ","metadata":{}}]}