{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\nimport os.path\nimport tqdm\nimport pandas as pd, numpy as np\nimport pyarrow.parquet as pq\nimport matplotlib.pyplot as plt\nimport scipy\nfrom scipy import signal\nfrom scipy.signal import savgol_filter\nimport tensorflow as tf\nfrom pywt import wavedec\nimport concurrent\nimport pickle,gzip\nimport xgboost as xgb\nimport sklearn\nimport keras\nfrom keras import layers\nfrom keras import ops\nimport glob\nimport matplotlib.pyplot as plt\nimport matplotlib.cm as cm\nfrom sklearn.decomposition import PCA\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:44:54.377069Z","iopub.execute_input":"2024-05-25T15:44:54.377602Z","iopub.status.idle":"2024-05-25T15:44:54.388756Z","shell.execute_reply.started":"2024-05-25T15:44:54.377559Z","shell.execute_reply":"2024-05-25T15:44:54.386986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"raw","source":"","metadata":{}},{"cell_type":"markdown","source":"train.csv Metadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\n\neeg_id - A unique identifier for the entire EEG recording.\n\neeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.\n\neeg_label_offset_seconds - The time between the beginning of the consolidated EEG and this subsample.\n\nspectrogram_id - A unique identifier for the entire EEG recording.\n\nspectrogram_sub_id - An ID for the specific 10 minute subsample this row's labels apply to.\n\nspectogram_label_offset_seconds - The time between the beginning of the consolidated spectrogram and this subsample.\n\nlabel_id - An ID for this set of labels.\n\npatient_id - An ID for the patient who donated the data.\n\nexpert_consensus - The consensus annotator label. Provided for convenience only.\n\n[seizure/lpd/gpd/lrda/grda/other]_vote - The count of annotator votes for a given brain activity class. The full names of the activity classes are as follows: lpd: lateralized periodic discharges, gpd: generalized periodic discharges, lrd: lateralized rhythmic delta activity, and grda: generalized rhythmic delta activity . A detailed explanations of these patterns is available here.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\ntrain","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:45:02.465918Z","iopub.execute_input":"2024-05-25T15:45:02.466341Z","iopub.status.idle":"2024-05-25T15:45:02.666812Z","shell.execute_reply.started":"2024-05-25T15:45:02.466309Z","shell.execute_reply":"2024-05-25T15:45:02.665586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_values = pd.DataFrame()\nfor col in train.columns[:-6]:\n    unique_values.at['unique_count', col] = len(train[col].unique())\nunique_values.astype(int)","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:45:05.288936Z","iopub.execute_input":"2024-05-25T15:45:05.289338Z","iopub.status.idle":"2024-05-25T15:45:05.337008Z","shell.execute_reply.started":"2024-05-25T15:45:05.289307Z","shell.execute_reply":"2024-05-25T15:45:05.335807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"train_eegs/ EEG data from one or more overlapping samples. Use the metadata in train.csv to select specific annotated subsets. The column names are the names of the individual electrode locations for EEG leads, with one exception. The EKG column is for an electrocardiogram lead that records data from the heart. All of the EEG data (for both train and test) was collected at a frequency of 200 samples per second.\n![electrodes](https://www.researchgate.net/profile/Gloria-Mcanulty/publication/254260385/figure/fig1/AS:202868222107648@1425378961756/Standard-EEG-electrode-names-and-positions-Head-in-vertex-view-nose-above-left-ear-to.png)\n","metadata":{}},{"cell_type":"code","source":"# load eeg by id\neeg0_id =2289322082\neeg0 = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg0_id}.parquet')\neeg0","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:45:10.294627Z","iopub.execute_input":"2024-05-25T15:45:10.295026Z","iopub.status.idle":"2024-05-25T15:45:10.339662Z","shell.execute_reply.started":"2024-05-25T15:45:10.294997Z","shell.execute_reply":"2024-05-25T15:45:10.338307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# clop corresponding 50 second long subsample\ni = 2024\n\neeg1_id = train['eeg_id'][i] # == 6.0s\noffset_idx = int(train['eeg_label_offset_seconds'][i]*100) \n\neeg1 = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg1_id}.parquet')\neeg1 = eeg1[offset_idx:offset_idx+5000] #slice 5000 slices = 50 sec.\nprint(f\"offset_sec = {train['eeg_label_offset_seconds'][i]}\")\neeg1","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:45:12.865900Z","iopub.execute_input":"2024-05-25T15:45:12.866800Z","iopub.status.idle":"2024-05-25T15:45:12.916154Z","shell.execute_reply.started":"2024-05-25T15:45:12.866760Z","shell.execute_reply":"2024-05-25T15:45:12.914807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"eeg_np=eeg1.to_numpy()\nx = np.linspace(9,50, 5000)\nfig, axs = plt.subplots(1,1, figsize=(8,4))\naxs.plot(x,eeg_np[:,9][:5000])\naxs.set_title('Cz')\naxs.axis([9, 50, -500, 500])\naxs.set_xlabel('Time [s]')\naxs.set_ylabel('Voltage [μV]')\n\nthreshold = -300\n\naxs.fill_between(x,25,-600,where=(eeg_np[:,0][:5000]< threshold) | (eeg_np[:,0][:5000]> 250) , color='cyan', alpha=0.4, transform=axs.get_xaxis_transform())\n","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:45:16.266570Z","iopub.execute_input":"2024-05-25T15:45:16.267001Z","iopub.status.idle":"2024-05-25T15:45:16.648768Z","shell.execute_reply.started":"2024-05-25T15:45:16.266961Z","shell.execute_reply":"2024-05-25T15:45:16.647351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = np.linspace(9,50, 5000)\nfig, axs = plt.subplots(1,1, figsize=(8,4))\nsos = signal.butter(2, 0.1, 'hp', fs=200, output='sos')\nfiltered = signal.sosfilt(sos, eeg_np[:,0][:5000])\nx_sg=savgol_filter(filtered, 17, 2)\naxs.plot(x, x_sg)\naxs.set_title('Filtered signal (Cz)')\naxs.set_xlabel('Time [s]')\naxs.set_ylabel('Voltage [μV]')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:45:20.081473Z","iopub.execute_input":"2024-05-25T15:45:20.081879Z","iopub.status.idle":"2024-05-25T15:45:20.373915Z","shell.execute_reply.started":"2024-05-25T15:45:20.081848Z","shell.execute_reply":"2024-05-25T15:45:20.372687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Standardize the data\nscaler = StandardScaler()\neeg_data_scaled = scaler.fit_transform(eeg1)\n\n# Apply PCA\npca = PCA(n_components=5)  # Adjust n_components as needed\neeg_data_pca = pca.fit_transform(eeg_data_scaled)\neeg_data_pca.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:45:23.572976Z","iopub.execute_input":"2024-05-25T15:45:23.573358Z","iopub.status.idle":"2024-05-25T15:45:23.631289Z","shell.execute_reply.started":"2024-05-25T15:45:23.573325Z","shell.execute_reply":"2024-05-25T15:45:23.629820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def Filter(x):\n    sos = signal.butter(2, 0.1, 'hp', fs=200, output='sos')\n    filt_1Hz = signal.sosfilt(sos, x)\n    \n    x_filtered=savgol_filter(filt_1Hz, 17, 2)\n    return  x_filtered\n\n\ndef Filtering(M):\n    \n    eeg_np=M.fillna(0).to_numpy()\n    for i in range(19):\n        eeg_np[:,i]=Filter(eeg_np[:,i])\n    return pd.DataFrame(eeg_np)\ndef PCA_EEG(M):\n    scaler = StandardScaler()\n    eeg_data_scaled = scaler.fit_transform(M)\n\n    # Apply PCA\n    pca = PCA(n_components=5)  # Adjust n_components as needed\n    eeg_data_pca = pca.fit_transform(eeg_data_scaled)\n    # Reshape to 1D array\n    eeg_data_pca = eeg_data_pca.flatten()\n\n    # Convert to pandas Series\n    eeg_data_pca = pd.DataFrame(eeg_data_pca)\n\n    return eeg_data_pca","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-05-25T15:45:26.072723Z","iopub.execute_input":"2024-05-25T15:45:26.073126Z","iopub.status.idle":"2024-05-25T15:45:26.083832Z","shell.execute_reply.started":"2024-05-25T15:45:26.073097Z","shell.execute_reply":"2024-05-25T15:45:26.082414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sampling_frequency = 200 \ndata_collection_duration = 50\ntotal_samples = sampling_frequency * data_collection_duration \n\n\nnum_train_data_points = 25000  \ntraining_data_df = pd.DataFrame()\n\nfor i in tqdm.tqdm(range(num_train_data_points)):\n    # Loading EEG data for a specified eeg_id\n    eeg_id = train.loc[i, 'eeg_id']\n    eeg_data = pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/{eeg_id}.parquet')\n    eeg_data=Filtering(eeg_data)\n    # Extracting EEG data for 50 seconds\n    label_offset_time = train.loc[i, 'eeg_label_offset_seconds']  # Offset time for the EEG label\n    label_offset_index = int(sampling_frequency * label_offset_time)  # Calculating offset index\n    # aPPLY pca\n    eeg_data=PCA_EEG(eeg_data.iloc[label_offset_index:label_offset_index + total_samples, :])\n\n    # Adding the extracted data as a row to the training DataFrame\n    training_data_df = pd.concat([training_data_df, eeg_data.transpose()], axis=0)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-25T17:12:29.736185Z","iopub.execute_input":"2024-05-25T17:12:29.737290Z","iopub.status.idle":"2024-05-25T17:42:54.541750Z","shell.execute_reply.started":"2024-05-25T17:12:29.737240Z","shell.execute_reply":"2024-05-25T17:42:54.540509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_data_df['expert_consensus'] = train[:num_train_data_points]['expert_consensus'].values\n\n# Removing rows with missing values\ntraining_data_df = training_data_df.dropna()\ntraining_data_df = training_data_df.reset_index(drop=True)\n\n# Separating data into features and target\ny_train = training_data_df['expert_consensus']\nX_train = training_data_df.drop('expert_consensus', axis=1)\n\n# Displaying the first few rows of the feature  dataset\ny_train.size","metadata":{"execution":{"iopub.status.busy":"2024-05-25T18:24:17.817653Z","iopub.execute_input":"2024-05-25T18:24:17.818058Z","iopub.status.idle":"2024-05-25T18:24:18.268481Z","shell.execute_reply.started":"2024-05-25T18:24:17.818024Z","shell.execute_reply":"2024-05-25T18:24:18.266025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n# Create a DataFrame from y_train\ndf = pd.DataFrame(y_train, columns=['label'])\nplt.figure(figsize=(10, 6))\nsns.countplot(x=y_train, data=df)\nplt.title('Distribution of Labels in y_train')\nplt.xlabel('Label')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-25T16:38:23.785921Z","iopub.execute_input":"2024-05-25T16:38:23.786359Z","iopub.status.idle":"2024-05-25T16:38:24.089931Z","shell.execute_reply.started":"2024-05-25T16:38:23.786316Z","shell.execute_reply":"2024-05-25T16:38:24.088692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the mapping\nlabel_mapping = {'Seizure': 0, 'GPD': 1, 'LPD':2, 'LRDA': 3, 'GRDA': 4,'Other': 5}\n\ny_train = [label_mapping[label] for label in y_train]","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:49:09.066951Z","iopub.execute_input":"2024-05-25T15:49:09.067456Z","iopub.status.idle":"2024-05-25T15:49:09.074460Z","shell.execute_reply.started":"2024-05-25T15:49:09.067397Z","shell.execute_reply":"2024-05-25T15:49:09.072903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nmodel = RandomForestClassifier(max_depth=5, min_samples_split=10, n_estimators=12)\n\nmodel.fit(X_train, y_train)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:57:39.853298Z","iopub.execute_input":"2024-05-25T15:57:39.854092Z","iopub.status.idle":"2024-05-25T15:59:00.499805Z","shell.execute_reply.started":"2024-05-25T15:57:39.854053Z","shell.execute_reply":"2024-05-25T15:59:00.498521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')\n\n# Displaying the first few rows of the DataFrame\ndisplay(df_test.head())","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:52:40.867521Z","iopub.execute_input":"2024-05-25T15:52:40.867981Z","iopub.status.idle":"2024-05-25T15:52:40.893041Z","shell.execute_reply.started":"2024-05-25T15:52:40.867923Z","shell.execute_reply":"2024-05-25T15:52:40.892123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = pd.DataFrame()\n\n# Iterating over each test data point\nfor i in tqdm.tqdm(range(len(df_test))):\n    # Loading EEG data for a specified eeg_id\n    eeg_id = df_test.loc[i, 'eeg_id']\n    eeg_data= pd.read_parquet(f'/kaggle/input/hms-harmful-brain-activity-classification/test_eegs/{eeg_id}.parquet')\n    eeg_data=Filtering(eeg_data)\n    eeg_data=PCA_EEG(eeg_data)\n    \n    # Adding the extracted data as a row to the testing DataFrame\n    X_test = pd.concat([X_test, eeg_data.transpose()], axis=0)\n\nX_test","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:52:44.176349Z","iopub.execute_input":"2024-05-25T15:52:44.176800Z","iopub.status.idle":"2024-05-25T15:52:44.364783Z","shell.execute_reply.started":"2024-05-25T15:52:44.176767Z","shell.execute_reply":"2024-05-25T15:52:44.363560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model.predict_proba(X_test)\n\nTARGETS = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\nsub = pd.DataFrame({'eeg_id': df_test.eeg_id.values})\nsub[TARGETS] = predictions\nsub.to_csv('submission.csv',index=False)\nprint(f'Submissionn shape: {sub.shape}')\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:59:07.125834Z","iopub.execute_input":"2024-05-25T15:59:07.126219Z","iopub.status.idle":"2024-05-25T15:59:08.257837Z","shell.execute_reply.started":"2024-05-25T15:59:07.126188Z","shell.execute_reply":"2024-05-25T15:59:08.256232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-05-25T15:59:17.555571Z","iopub.execute_input":"2024-05-25T15:59:17.555962Z","iopub.status.idle":"2024-05-25T15:59:17.562944Z","shell.execute_reply.started":"2024-05-25T15:59:17.555932Z","shell.execute_reply":"2024-05-25T15:59:17.561534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"train_spectrograms/ Spectrograms assembled EEG data. Use the metadata in train.csv to select specific annotated subsets. The column names indicate the frequency in hertz and the recording regions of the EEG electrodes. The latter are abbreviated as LL = left lateral; RL = right lateral; LP = left parasagittal; RP = right parasagittal.","metadata":{}}]}