{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction to my EDA**\n\nLet's start the kaggle EDA of the HMS harmful brain activity classification.\n\nThe goal: \"the goal of this competition is to detect and classify seizures and other types of harmful brain activity\".\n\nThe measure: \"Submissions are evaluated on the Kullback Liebler divergence between the predicted probability and the observed target\".\n","metadata":{}},{"cell_type":"code","source":"# Import Libraries\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/hms-harmful-brain-activity-classification'):\n#    for filename in filenames:\n        #print(os.path.join(dirname, filename))\n\ndata_train_path = '/kaggle/input/hms-harmful-brain-activity-classification/train.csv'","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:17:31.537942Z","iopub.execute_input":"2024-02-28T10:17:31.538968Z","iopub.status.idle":"2024-02-28T10:17:31.546271Z","shell.execute_reply.started":"2024-02-28T10:17:31.538914Z","shell.execute_reply":"2024-02-28T10:17:31.544814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Form analysis: train.csv**","metadata":{}},{"cell_type":"code","source":"# Read the data\ndata = pd.read_csv(data_train_path)\n\n# Copy the Data\ndf = data.copy()\n\n# Observe few lines \nprint(df.head())\n\n# Shape of the data\nprint('The shape of df is:', df.shape)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:17:31.554035Z","iopub.execute_input":"2024-02-28T10:17:31.554396Z","iopub.status.idle":"2024-02-28T10:17:31.835520Z","shell.execute_reply.started":"2024-02-28T10:17:31.554361Z","shell.execute_reply":"2024-02-28T10:17:31.834359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create columns\nColumn_name = list(df.columns)\n\n# Number of NaN in each column\nnumber_na = df.isna().sum()\n\n# Type of Data and the number\ntypes = df.dtypes\nnumber_types = df.dtypes.value_counts()\nprint(number_types)\n\n# Create a resume table\ndf_resume = pd.DataFrame({'features': Column_name, 'Type': types, 'Number of NaN': number_na})\nprint(df_resume)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:17:31.837184Z","iopub.execute_input":"2024-02-28T10:17:31.837542Z","iopub.status.idle":"2024-02-28T10:17:31.863650Z","shell.execute_reply.started":"2024-02-28T10:17:31.837513Z","shell.execute_reply":"2024-02-28T10:17:31.862523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Background analysis: train.csv**","metadata":{}},{"cell_type":"markdown","source":"# 1. **Target vizualisation**","metadata":{}},{"cell_type":"code","source":"# Define the target columns\ntarget_columns = ['seizure_vote', 'lpd_vote','gpd_vote','lrda_vote','grda_vote','other_vote']\n\n# expert_consensus column is a part of target column since it represent the max of votes\nconsensus_column = ['expert_consensus']\n\n# Display few line of the target columns\nprint(df[target_columns].head())","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:17:31.864990Z","iopub.execute_input":"2024-02-28T10:17:31.865277Z","iopub.status.idle":"2024-02-28T10:17:31.881868Z","shell.execute_reply.started":"2024-02-28T10:17:31.865252Z","shell.execute_reply":"2024-02-28T10:17:31.880735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the proportion of types in the column 'expert_consensus'\ncount_df_consensus = df['expert_consensus'].value_counts().reset_index()\ncount_df_consensus.columns=['expert_consensus','count']\n\n\n# Define colors\ncolors = sns.color_palette(\"viridis\", len(count_df_consensus))\n\n# Display df_consensus in a bar\nplt.figure(figsize=(12,8))\nplt.bar(count_df_consensus['expert_consensus'], count_df_consensus['count'], color=colors, zorder=2)\nplt.xlabel('Expert consensus')\nplt.ylabel('Count')\nplt.title('Distribution of Expert Consensus')\nplt.legend()\nplt.grid(ls='--')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:17:31.883539Z","iopub.execute_input":"2024-02-28T10:17:31.883970Z","iopub.status.idle":"2024-02-28T10:17:32.289921Z","shell.execute_reply.started":"2024-02-28T10:17:31.883937Z","shell.execute_reply":"2024-02-28T10:17:32.288857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. **Significance of Variables**","metadata":{}},{"cell_type":"markdown","source":"* **label_id**: Nothing to say since it's a unique number representing each cases in the dataset","metadata":{}},{"cell_type":"markdown","source":"* **eeg_id, eeg_sub_id, eeg_label_offset_seconds:**","metadata":{}},{"cell_type":"code","source":"# Count the number of unique eeg_id\ndf_count_eeg_id = df['eeg_id'].value_counts().reset_index()\ndf_count_eeg_id.columns = ['eeg_id','Count']\nprint(df_count_eeg_id)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:17:32.293413Z","iopub.execute_input":"2024-02-28T10:17:32.294200Z","iopub.status.idle":"2024-02-28T10:17:32.310302Z","shell.execute_reply.started":"2024-02-28T10:17:32.294160Z","shell.execute_reply":"2024-02-28T10:17:32.309092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display if in a figure\nplt.figure(figsize=(12,8))\nplt.bar(df_count_eeg_id.index, df_count_eeg_id['Count'], color=colors, zorder=2)\nplt.xlabel('eeg_id')\nplt.ylabel('Count')\nplt.title('Distribution of eeg_id')\nplt.legend()\nplt.grid(ls='--')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:17:32.311457Z","iopub.execute_input":"2024-02-28T10:17:32.311790Z","iopub.status.idle":"2024-02-28T10:18:02.373908Z","shell.execute_reply.started":"2024-02-28T10:17:32.311740Z","shell.execute_reply":"2024-02-28T10:18:02.372838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the dataframe corresponding to the eeg_id which appears the most 'eeg_id'=2259539799\ndf_max_eeg_id = df[df['eeg_id']==df_count_eeg_id.loc[df_count_eeg_id['Count'].idxmax(),'eeg_id']]\nprint(df_max_eeg_id)\n\n# Display the different result of expert consensus for this eeg_id\nprint(df_max_eeg_id['expert_consensus'].unique())\n\n# Display the dataframe corresponding to the eeg_id which appears the least 'eeg_id'=98046913 (one of them)\ndf_min_eeg_id = df[df['eeg_id']==df_count_eeg_id.loc[df_count_eeg_id['Count'].idxmin(),'eeg_id']]\nprint(df_min_eeg_id)\n\n# Display the different result of expert consensus for this eeg_id\nprint(df_min_eeg_id['expert_consensus'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:02.375033Z","iopub.execute_input":"2024-02-28T10:18:02.375701Z","iopub.status.idle":"2024-02-28T10:18:02.403191Z","shell.execute_reply.started":"2024-02-28T10:18:02.375667Z","shell.execute_reply":"2024-02-28T10:18:02.402347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the min-max eeg_label_offset_seconds\nprint(df['eeg_label_offset_seconds'].min())\nprint(df['eeg_label_offset_seconds'].max())\n\n# Count the unique spectrogram_ids for each eeg_id\neeg_spectrogram_count = df.groupby('eeg_id')['spectrogram_id'].nunique()\n\n# Filter the eeg_ids that have more than one unique spectrogram_id\neeg_ids_with_multiple_spectrogram = eeg_spectrogram_count[eeg_spectrogram_count > 1]\n\n# Display the result\nprint(eeg_ids_with_multiple_spectrogram)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:02.404061Z","iopub.execute_input":"2024-02-28T10:18:02.404644Z","iopub.status.idle":"2024-02-28T10:18:02.425549Z","shell.execute_reply.started":"2024-02-28T10:18:02.404615Z","shell.execute_reply":"2024-02-28T10:18:02.424518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There are 17089 different eeg_id.\n- Each eeg_id contains eeg_sub_id which can go from 0 to 742 associated to a eeg_label_offset_seconds which goes from 0.0 to 3372.0. Thus the occurency of a eeg_id goes from 1 to 743.\n- A eeg_id is associated to a unique spectrogram_id.","metadata":{}},{"cell_type":"markdown","source":"* **spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds:**","metadata":{}},{"cell_type":"code","source":"# Count the number of unique spectrogram_id\ndf_count_spectrogram_id = df['spectrogram_id'].value_counts().reset_index()\ndf_count_spectrogram_id.columns = ['spectrogram_id','Count']\nprint(df_count_spectrogram_id)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:02.427063Z","iopub.execute_input":"2024-02-28T10:18:02.427504Z","iopub.status.idle":"2024-02-28T10:18:02.440582Z","shell.execute_reply.started":"2024-02-28T10:18:02.427460Z","shell.execute_reply":"2024-02-28T10:18:02.439116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display if in a figure\nplt.figure(figsize=(12,8))\nplt.bar(df_count_spectrogram_id.index, df_count_spectrogram_id['Count'], color=colors, zorder=2)\nplt.xlabel('spectrogram_id')\nplt.ylabel('Count')\nplt.title('Distribution of spectrogram_id')\nplt.legend()\nplt.grid(ls='--')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:02.442096Z","iopub.execute_input":"2024-02-28T10:18:02.442606Z","iopub.status.idle":"2024-02-28T10:18:22.349705Z","shell.execute_reply.started":"2024-02-28T10:18:02.442575Z","shell.execute_reply":"2024-02-28T10:18:22.348525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the dataframe corresponding to the spectrogram_id which appears the most 'spectrogram_id'=2259539799\ndf_max_spectrogram_id = df[df['spectrogram_id']==df_count_spectrogram_id.loc[df_count_spectrogram_id['Count'].idxmax(),'spectrogram_id']]\nprint(df_max_spectrogram_id)\n\n# Display the different result of expert consensus for this spectrogram_id\nprint(df_max_spectrogram_id['expert_consensus'].unique())\n\n# Display the dataframe corresponding to the spectrogram_id which appears the least 'spectrogram_id'=98046913 (one of them)\ndf_min_spectrogram_id = df[df['spectrogram_id']==df_count_spectrogram_id.loc[df_count_spectrogram_id['Count'].idxmin(),'spectrogram_id']]\nprint(df_min_spectrogram_id)\n\n# Display the different result of expert consensus for this spectrogram_id\nprint(df_min_spectrogram_id['expert_consensus'].unique())","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:22.351635Z","iopub.execute_input":"2024-02-28T10:18:22.352098Z","iopub.status.idle":"2024-02-28T10:18:22.373980Z","shell.execute_reply.started":"2024-02-28T10:18:22.352056Z","shell.execute_reply":"2024-02-28T10:18:22.372925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the min-max spectrogram_label_offset_seconds\nprint(df['spectrogram_label_offset_seconds'].min())\nprint(df['spectrogram_label_offset_seconds'].max())","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:22.375556Z","iopub.execute_input":"2024-02-28T10:18:22.376264Z","iopub.status.idle":"2024-02-28T10:18:22.385055Z","shell.execute_reply.started":"2024-02-28T10:18:22.376225Z","shell.execute_reply":"2024-02-28T10:18:22.383700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There are 11138 different spectrogram_id\n- Each spectrogram_id contains spectrogram_sub_id which can go from 0 to 1021 associated to a spectrogram_label_offset_seconds which goes from 0.0 to 17632.0?. Thus the occurency of a spectrogram_id goes from 1 to 1022.\n- A spectrogram_id can contain different eeg_id","metadata":{}},{"cell_type":"markdown","source":"* **patient_id**:","metadata":{}},{"cell_type":"code","source":"# Count the number of unique Patient_id\ndf_count_patient_id = df['patient_id'].value_counts().reset_index()\ndf_count_patient_id.columns = ['patient_id','Count']\nprint(df_count_patient_id)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:22.386666Z","iopub.execute_input":"2024-02-28T10:18:22.387010Z","iopub.status.idle":"2024-02-28T10:18:22.398481Z","shell.execute_reply.started":"2024-02-28T10:18:22.386980Z","shell.execute_reply":"2024-02-28T10:18:22.397435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display if in a figure\nplt.figure(figsize=(12,8))\nplt.bar(df_count_patient_id.index, df_count_patient_id['Count'], color=colors, zorder=2)\nplt.xlabel('patient_id')\nplt.ylabel('Count')\nplt.title('Distribution of patient_id')\nplt.legend()\nplt.grid(ls='--')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:22.403605Z","iopub.execute_input":"2024-02-28T10:18:22.403970Z","iopub.status.idle":"2024-02-28T10:18:26.470062Z","shell.execute_reply.started":"2024-02-28T10:18:22.403941Z","shell.execute_reply":"2024-02-28T10:18:26.468937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the dataframe corresponding to the patient_id which appears the most ('patient_id'=30631)\ndf_max_patient_id = df[df['patient_id']==df_count_patient_id.loc[df_count_patient_id['Count'].idxmax(),'patient_id']]\nprint(df_max_patient_id)\n\n\n# Display the dataframe corresponding to the patient_id which appears the least ('patient_id'=10324)\ndf_min_patient_id = df[df['patient_id']==df_count_patient_id.loc[df_count_patient_id['Count'].idxmin(),'patient_id']]\nprint(df_min_patient_id)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:26.471706Z","iopub.execute_input":"2024-02-28T10:18:26.472193Z","iopub.status.idle":"2024-02-28T10:18:26.491838Z","shell.execute_reply.started":"2024-02-28T10:18:26.472154Z","shell.execute_reply":"2024-02-28T10:18:26.490787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The occurrence is between 1 and 2215 corresponding to different combination of ('eeg_id','eeg_sub_id', 'eeg_label_offset_seconds',spectrogram_id',spectrogram_sub_id',spectrogram_label_offset_seconds).","metadata":{}},{"cell_type":"markdown","source":"# 3. **Relationship Target/Target**","metadata":{}},{"cell_type":"code","source":"# Define a correlation matrix between target columns\ncorrelation_matrix_target = df[target_columns].corr()\n\n# Plot the correlation matrix with clusters\nplt.figure(figsize=(12,8))\nsns.heatmap(correlation_matrix_target, annot=True, cmap='viridis')\nplt.title('Correlation matrix between target columns')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:26.492860Z","iopub.execute_input":"2024-02-28T10:18:26.493168Z","iopub.status.idle":"2024-02-28T10:18:26.894712Z","shell.execute_reply.started":"2024-02-28T10:18:26.493140Z","shell.execute_reply":"2024-02-28T10:18:26.893552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that seizure is correlated with grda_vote (24%), other_vote (21%), lrda_vote (17%), gpd_vote (13%), lpd_vote (13%). lrda_vote and gpd_vote (15%). The rest seems to be \"negligeable\" (-9%)","metadata":{}},{"cell_type":"markdown","source":"# 4. **Relationship Variable/Target**","metadata":{}},{"cell_type":"markdown","source":"* **Patient_id/expert_consensus:**","metadata":{}},{"cell_type":"code","source":"# Display the dataframe corresponding to columns 'patient_id', 'expert_consensus'\ndf_patient_id_expert_consensus = df[['patient_id','expert_consensus']]\nprint(df_patient_id_expert_consensus)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:26.896146Z","iopub.execute_input":"2024-02-28T10:18:26.896530Z","iopub.status.idle":"2024-02-28T10:18:26.908281Z","shell.execute_reply.started":"2024-02-28T10:18:26.896496Z","shell.execute_reply":"2024-02-28T10:18:26.907042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the unique expert_consensus for each patient_id\npatient_consensus_count = df.groupby('patient_id')['expert_consensus'].nunique()\n\n# Filter the eeg_ids that have more than one unique spectrogram_id\npatient_ids_with_multiple_expert_consensus = patient_consensus_count[patient_consensus_count > 1]\n\n# Display the result\nprint(patient_ids_with_multiple_expert_consensus)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:26.909714Z","iopub.execute_input":"2024-02-28T10:18:26.910144Z","iopub.status.idle":"2024-02-28T10:18:26.929521Z","shell.execute_reply.started":"2024-02-28T10:18:26.910108Z","shell.execute_reply":"2024-02-28T10:18:26.928460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the percentage of patients for each number of unique expert_consensus\npercentage_dict = {}\nfor i in range(1, 6):\n    count = (patient_consensus_count == i).sum()\n    percentage = (count / len(patient_consensus_count)) * 100\n    print(percentage)\n    percentage_dict[i] = percentage\n\n# Create a bar plot\nplt.bar(percentage_dict.keys(), percentage_dict.values())\nplt.xlabel('Number of unique expert_consensus')\nplt.ylabel('Percentage of patients')\nplt.title('Percentage of patients with different numbers of unique expert_consensus')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:26.930664Z","iopub.execute_input":"2024-02-28T10:18:26.930978Z","iopub.status.idle":"2024-02-28T10:18:27.135255Z","shell.execute_reply.started":"2024-02-28T10:18:26.930952Z","shell.execute_reply":"2024-02-28T10:18:27.134223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example of a patient_id with diffent expert_consensus\nprint(df[df['patient_id']==105])","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.136502Z","iopub.execute_input":"2024-02-28T10:18:27.137060Z","iopub.status.idle":"2024-02-28T10:18:27.155083Z","shell.execute_reply.started":"2024-02-28T10:18:27.137026Z","shell.execute_reply":"2024-02-28T10:18:27.153991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each patient_id may be associated with multiple expert_consensus records. Specifically, approximately 43% possess a singular expert_consensus, 35% exhibit two, 13% manifest three, 4% showcase four, and 1% display more than four.","metadata":{}},{"cell_type":"markdown","source":"* **eeg_id, eeg_sub_id, eeg_label_offset_seconds/expert_consensus:**","metadata":{}},{"cell_type":"code","source":"# Display the dataframe corresponding to columns 'eeg_id', 'eeg_sub_id', 'eeg_label_offset_seconds', 'expert_consensus'\ndf_eeg_expert_consensus = df[['eeg_id','eeg_sub_id','eeg_label_offset_seconds','expert_consensus']]\nprint(df_eeg_expert_consensus)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.156395Z","iopub.execute_input":"2024-02-28T10:18:27.156706Z","iopub.status.idle":"2024-02-28T10:18:27.170302Z","shell.execute_reply.started":"2024-02-28T10:18:27.156671Z","shell.execute_reply":"2024-02-28T10:18:27.169236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the unique expert_consensus for each eeg_id\neeg_consensus_count = df.groupby(['eeg_id'])['expert_consensus'].nunique()\n\n# Filter the eeg_ids that have more than one unique expert_consensus\neeg_ids_with_multiple_expert_consensus = eeg_consensus_count[eeg_consensus_count > 1]\n\n# Display the result\nprint(eeg_ids_with_multiple_expert_consensus)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.171994Z","iopub.execute_input":"2024-02-28T10:18:27.172316Z","iopub.status.idle":"2024-02-28T10:18:27.197205Z","shell.execute_reply.started":"2024-02-28T10:18:27.172288Z","shell.execute_reply":"2024-02-28T10:18:27.196047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the percentage of eeg for each number of unique expert_consensus\npercentage_dict = {}\nfor i in range(1, 6):\n    count = (eeg_consensus_count == i).sum()\n    percentage = (count / len(eeg_consensus_count)) * 100\n    print(percentage)\n    percentage_dict[i] = percentage\n\n# Create a bar plot\nplt.bar(percentage_dict.keys(), percentage_dict.values())\nplt.xlabel('Number of unique expert_consensus')\nplt.ylabel('Percentage of eeg')\nplt.title('Percentage of eeg with different numbers of unique expert_consensus')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.198696Z","iopub.execute_input":"2024-02-28T10:18:27.199357Z","iopub.status.idle":"2024-02-28T10:18:27.401915Z","shell.execute_reply.started":"2024-02-28T10:18:27.199316Z","shell.execute_reply":"2024-02-28T10:18:27.400826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example of a eeg_id with diffent expert_consensus\nprint(df[df['eeg_id']==21379701])","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.403266Z","iopub.execute_input":"2024-02-28T10:18:27.403572Z","iopub.status.idle":"2024-02-28T10:18:27.416563Z","shell.execute_reply.started":"2024-02-28T10:18:27.403547Z","shell.execute_reply":"2024-02-28T10:18:27.415184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each eeg_id can have different expert consensus based on the eeg_sub_id or eeg_label_offset_seconds we look at. But more that 95% of eeg_id has only one expert_consensus.","metadata":{}},{"cell_type":"markdown","source":"* **spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds/expert_consensus:**","metadata":{}},{"cell_type":"code","source":"# Display the dataframe corresponding to columns 'spectrogram_id', 'spectrogram_sub_id', 'spectrogram_label_offset_seconds', 'expert_consensus'\ndf_spec_expert_consensus = df[['spectrogram_id','spectrogram_sub_id','spectrogram_label_offset_seconds','expert_consensus']]\nprint(df_spec_expert_consensus)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.418057Z","iopub.execute_input":"2024-02-28T10:18:27.418354Z","iopub.status.idle":"2024-02-28T10:18:27.432322Z","shell.execute_reply.started":"2024-02-28T10:18:27.418330Z","shell.execute_reply":"2024-02-28T10:18:27.430989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the unique expert_consensus for each spectrogram\nspec_consensus_count = df.groupby(['spectrogram_id'])['expert_consensus'].nunique()\n\n# Filter the spectrogram_ids that have more than one unique expert_consensus\nspec_ids_with_multiple_expert_consensus = spec_consensus_count[spec_consensus_count > 1]\n\n# Display the result\nprint(spec_ids_with_multiple_expert_consensus)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.433920Z","iopub.execute_input":"2024-02-28T10:18:27.434294Z","iopub.status.idle":"2024-02-28T10:18:27.455895Z","shell.execute_reply.started":"2024-02-28T10:18:27.434265Z","shell.execute_reply":"2024-02-28T10:18:27.454696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the percentage of spectrogram for each number of unique expert_consensus\npercentage_dict = {}\nfor i in range(1, 6):\n    count = (spec_consensus_count == i).sum()\n    percentage = (count / len(spec_consensus_count)) * 100\n    print(percentage)\n    percentage_dict[i] = percentage\n\n# Create a bar plot\nplt.bar(percentage_dict.keys(), percentage_dict.values())\nplt.xlabel('Number of unique expert_consensus')\nplt.ylabel('Percentage of spectrograms')\nplt.title('Percentage of spectrograms with different numbers of unique expert_consensus')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.457478Z","iopub.execute_input":"2024-02-28T10:18:27.458060Z","iopub.status.idle":"2024-02-28T10:18:27.658924Z","shell.execute_reply.started":"2024-02-28T10:18:27.458021Z","shell.execute_reply":"2024-02-28T10:18:27.657838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example of a spectrogram_id with diffent expert_consensus\nprint(df[df['spectrogram_id']==12849827])","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.660370Z","iopub.execute_input":"2024-02-28T10:18:27.660798Z","iopub.status.idle":"2024-02-28T10:18:27.673378Z","shell.execute_reply.started":"2024-02-28T10:18:27.660738Z","shell.execute_reply":"2024-02-28T10:18:27.672198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each spectrogram_id can have different expert consensus based on spectrogram_sub_id, spectrogram_label_offset_seconds but also the eeg_id we look at. More than 90% of spectrogram_id has one expert_consensus.","metadata":{}},{"cell_type":"markdown","source":"# 5. **Relationship Variables/Variables**","metadata":{}},{"cell_type":"markdown","source":"* **eeg_id/spectrogram_id:**","metadata":{}},{"cell_type":"code","source":"# Count the unique spectrogram_ids for each eeg_id\neeg_spectrogram_count = df.groupby('eeg_id')['spectrogram_id'].nunique()\n\n# Filter the eeg_ids that have more than one unique spectrogram_id and display the result\neeg_ids_with_multiple_spectrogram = eeg_spectrogram_count[eeg_spectrogram_count > 1]\nprint(eeg_ids_with_multiple_spectrogram)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.674605Z","iopub.execute_input":"2024-02-28T10:18:27.674924Z","iopub.status.idle":"2024-02-28T10:18:27.691566Z","shell.execute_reply.started":"2024-02-28T10:18:27.674897Z","shell.execute_reply":"2024-02-28T10:18:27.690395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the number of unique spectrogram_id\ndf_count_spectrogram_id = df['spectrogram_id'].value_counts().reset_index()\ndf_count_spectrogram_id.columns = ['spectrogram_id','Count']\n\n# Display the dataframe corresponding to the spectrogram_id which appears the most 'spectrogram_id'=2259539799\ndf_max_spectrogram_id = df[df['spectrogram_id']==df_count_spectrogram_id.loc[df_count_spectrogram_id['Count'].idxmax(),'spectrogram_id']]\nprint(df_max_spectrogram_id)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.692873Z","iopub.execute_input":"2024-02-28T10:18:27.693181Z","iopub.status.idle":"2024-02-28T10:18:27.711297Z","shell.execute_reply.started":"2024-02-28T10:18:27.693156Z","shell.execute_reply":"2024-02-28T10:18:27.710059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A eeg_id is associated to a unique spectrogram_id but a spectrogram_id can contain different eeg_id","metadata":{}},{"cell_type":"markdown","source":"* **eeg_id/patient_id and spectrogram_id/patient_id:**","metadata":{}},{"cell_type":"code","source":"# Count the number of unique Patient_id\ndf_count_patient_id = df['patient_id'].value_counts().reset_index()\ndf_count_patient_id.columns = ['patient_id','Count']\nprint(df_count_patient_id)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.712660Z","iopub.execute_input":"2024-02-28T10:18:27.713005Z","iopub.status.idle":"2024-02-28T10:18:27.724787Z","shell.execute_reply.started":"2024-02-28T10:18:27.712975Z","shell.execute_reply":"2024-02-28T10:18:27.723534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the dataframe corresponding to the patient_id which appears the most ('patient_id'=30631)\ndf_max_patient_id = df[df['patient_id']==df_count_patient_id.loc[df_count_patient_id['Count'].idxmax(),'patient_id']]\nprint(df_max_patient_id)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.726427Z","iopub.execute_input":"2024-02-28T10:18:27.726766Z","iopub.status.idle":"2024-02-28T10:18:27.743054Z","shell.execute_reply.started":"2024-02-28T10:18:27.726720Z","shell.execute_reply":"2024-02-28T10:18:27.741710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" Each patient can have several different eeg_id and spectrogram_id.","metadata":{}},{"cell_type":"markdown","source":"# Checklist\n\n**Initial Form Analysis**:\n\n- Target Variable: (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote)\n- Rows and Columns: (106800, 15)\n- Types of Variables: 12 int64, 2 float64, 1 object\n- Analysis of Missing Variables: No missing value\n\n**Initial Background Analysis**:\n\n- Target Visualization: The target variable is well balanced. Number of expert consensus for each type of activity is between 15000 and 2000\n\n- Significance of Variables:\n    * label_id: Nothing to say since it's a unique number representing each cases in the dataset\n    * eeg_id: There are 17089 different eeg_id. Each eeg_id contains eeg_sub_id which can go from 0 to 742 associated to a eeg_label_offset_seconds which goes from 0.0 to 3372.0. Thus the occurency of a eeg_id goes from 1 to 743. A eeg_id is associated to a unique spectrogram_id.\n    * spectrogram_id: There are 11138 different spectrogram_id Each spectrogram_id contains spectrogram_sub_id which can go from 0 to 1021 associated to a spectrogram_label_offset_seconds which goes from 0.0 to 17632.0?. Thus the occurency of a spectrogram_id goes from 1 to 1022. A spectrogram_id can contain different eeg_id  \n    * patient_id: There are 1950 different patient_id. The occurrence is between 1 and 2215 corresponding to different combination of ('eeg_id','eeg_sub_id', 'eeg_label_offset_seconds', 'spectrogram_id', 'spectrogram_sub_id', 'spectrogram_label_offset_seconds').  \n    \n    \n- Relationship Target/target: it seems that seizure is correlated with grda_vote (24%), other_vote (21%), lrda_vote (17%), gpd_vote (13%), lpd_vote (13%). lrda_vote and gpd_vote (15%). The rest seems to be negligeable (-8%).\n\n- Relationship Variables/Target: \n    * patient_id/expert_consensus: Each patient_id can have more than one expert_consensus corresponding to a particular configuration (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote).\n    * eeg_id, eeg_sub_id, eeg_label_offset_seconds/expert_consensus: Each eeg_id can have different expert consensus based on the eeg_sub_id or eeg_label_offset_seconds we look at.\n    * spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds/expert_consensus: Each spectrogram_id can have different expert consensus based on spectrogram_sub_id, spectrogram_label_offset_seconds \n      but also the eeg_id we look at.\n      \n- Relationship Variables/variables: \n    * eeg_id/spectrogram_id: A eeg_id is associated to a unique spectrogram_id but a spectrogram_id can contain different eeg_id.\n    * eeg_id/patient_id and spectrogram_id/patient_id: Each patient can have several different eeg_id and spectrogram_id.","metadata":{}},{"cell_type":"markdown","source":"# **Analysis of one particular eeg_id from train_eegs.csv**","metadata":{}},{"cell_type":"code","source":"# To locate the correct spectrogram sample using the offset, I used the following code of @cdeotte.\n'''\nGET_ROW = 0\nEEG_PATH = 'train_eegs/'\nSPEC_PATH = 'train_spectrograms/'\n\ntrain = pd.read_csv('train.csv')\nrow = train.iloc[GET_ROW]\n\neeg = pd.read_parquet(f'{EEG_PATH}{row.eeg_id}.parquet')\neeg_offset = int( row.eeg_label_offset_seconds )\neeg = eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]\n\nspectrogram = pd.read_parquet(f'{SPEC_PATH}{row.spectrogram_id}.parquet')\nspec_offset = int( row.spectrogram_label_offset_seconds )\nspectrogram = spectrogram.loc[(spectrogram.time>=spec_offset)\n                 &(spectrogram.time<spec_offset+600)]\n'''","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.744224Z","iopub.execute_input":"2024-02-28T10:18:27.744522Z","iopub.status.idle":"2024-02-28T10:18:27.751722Z","shell.execute_reply.started":"2024-02-28T10:18:27.744496Z","shell.execute_reply":"2024-02-28T10:18:27.750935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_train_eeg_path = '/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/'\ntrain = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.753107Z","iopub.execute_input":"2024-02-28T10:18:27.753630Z","iopub.status.idle":"2024-02-28T10:18:27.912264Z","shell.execute_reply.started":"2024-02-28T10:18:27.753603Z","shell.execute_reply":"2024-02-28T10:18:27.911236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select the first row of the dataset train\nGET_ROW = 0\nrow = train.iloc[GET_ROW]\ndf_eeg = pd.read_parquet(f'{data_train_eeg_path}{row.eeg_id}.parquet')\neeg_offset = int( row.eeg_label_offset_seconds )\ndf_eeg = df_eeg.iloc[eeg_offset*200:(eeg_offset+50)*200]","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:27.913792Z","iopub.execute_input":"2024-02-28T10:18:27.914205Z","iopub.status.idle":"2024-02-28T10:18:28.099023Z","shell.execute_reply.started":"2024-02-28T10:18:27.914168Z","shell.execute_reply":"2024-02-28T10:18:28.098010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the number of offset_seconds for a unique eeg_id\nprint(df.iloc[GET_ROW]['eeg_id'])\nprint(df[df['eeg_id'] == df.iloc[GET_ROW]['eeg_id']]['eeg_label_offset_seconds'].unique())\n\n# Print the unique offset_seconds for a specific row\nprint(df[df.reset_index()['index']==GET_ROW]['eeg_label_offset_seconds'])","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:28.101923Z","iopub.execute_input":"2024-02-28T10:18:28.102345Z","iopub.status.idle":"2024-02-28T10:18:28.121114Z","shell.execute_reply.started":"2024-02-28T10:18:28.102306Z","shell.execute_reply":"2024-02-28T10:18:28.120056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Copy the Data\ndf_eeg = df_eeg.copy()\n\n# Create columns time with offset_seconds\nsample_per_second = 200\ndf_eeg['seconds_with_offset'] = range(df_eeg.shape[0]) \ndf_eeg['seconds_with_offset'] = df_eeg['seconds_with_offset']/sample_per_second +df[df.reset_index()['index']==GET_ROW]['eeg_label_offset_seconds'].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:28.122491Z","iopub.execute_input":"2024-02-28T10:18:28.123437Z","iopub.status.idle":"2024-02-28T10:18:28.138592Z","shell.execute_reply.started":"2024-02-28T10:18:28.123387Z","shell.execute_reply":"2024-02-28T10:18:28.137824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Observe few lines \nprint(df_eeg.head())\n\n# Shape of the data\nprint('The shape of df is:', df_eeg.shape)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:28.141550Z","iopub.execute_input":"2024-02-28T10:18:28.141920Z","iopub.status.idle":"2024-02-28T10:18:28.156015Z","shell.execute_reply.started":"2024-02-28T10:18:28.141891Z","shell.execute_reply":"2024-02-28T10:18:28.154825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create columns\nColumn_name = list(df_eeg.columns)\n\n# Number of NaN in each column\nnumber_na = df_eeg.isna().sum()\n\n# Type of Data and the number\ntypes = df_eeg.dtypes\nnumber_types = df_eeg.dtypes.value_counts()\nprint(number_types)\n\n# Create a resume table\ndf_eeg_resume = pd.DataFrame({'features': Column_name, 'Type': types, 'Number of NaN': number_na})\nprint(df_eeg_resume)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:28.157395Z","iopub.execute_input":"2024-02-28T10:18:28.157842Z","iopub.status.idle":"2024-02-28T10:18:28.173359Z","shell.execute_reply.started":"2024-02-28T10:18:28.157811Z","shell.execute_reply":"2024-02-28T10:18:28.172441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the 3 first rows for the columns Fp1 and observe that there are overlaps\nplt.figure(figsize=(12, 8))\nfor i in range(3):\n    GET_ROW = i\n    row = train.iloc[GET_ROW]\n    df_eeg = pd.read_parquet(f'{data_train_eeg_path}{row.eeg_id}.parquet')\n    eeg_offset = int( row.eeg_label_offset_seconds )\n    df_eeg = df_eeg.iloc[eeg_offset*200:(eeg_offset+50)*200] \n    df_eeg = df_eeg.copy()\n    df_eeg['seconds_with_offset'] = range(df_eeg.shape[0]) \n    df_eeg['seconds_with_offset'] = df_eeg['seconds_with_offset']/sample_per_second +df[df.reset_index()['index']==GET_ROW]['eeg_label_offset_seconds'].unique()\n    color = plt.cm.viridis(i/8.0)\n    plt.plot(df_eeg['seconds_with_offset'], df_eeg['Fp1'], c='black')\n    plt.fill_between(df_eeg['seconds_with_offset'], -250, 0, where=[(x >= df_eeg['seconds_with_offset'].min()) and (x <= df_eeg['seconds_with_offset'].max()) for x in df_eeg['seconds_with_offset']], color=color, alpha=0.7, label=f'row {i}')\n    plt.axvline(df_eeg['seconds_with_offset'].min(), color=color, linestyle='-', linewidth=2)\n    plt.axvline(df_eeg['seconds_with_offset'].max(), color=color, linestyle='-', linewidth=2)\nplt.ylabel('Fp1')\nplt.xlabel('seconds')\nplt.legend()\nplt.grid(ls='--')    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:28.174470Z","iopub.execute_input":"2024-02-28T10:18:28.175424Z","iopub.status.idle":"2024-02-28T10:18:31.046040Z","shell.execute_reply.started":"2024-02-28T10:18:28.175394Z","shell.execute_reply":"2024-02-28T10:18:31.045167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the full eeg_id for the columns Fp1 and the ten seconds that we need to predict and observe that there are overlaps\nplt.figure(figsize=(12, 8))\nfor i in range(len(df[df['eeg_id'] == df.iloc[GET_ROW]['eeg_id']]['eeg_label_offset_seconds'].unique())):\n    GET_ROW = i\n    row = train.iloc[GET_ROW]\n    df_eeg = pd.read_parquet(f'{data_train_eeg_path}{row.eeg_id}.parquet')\n    eeg_offset = int( row.eeg_label_offset_seconds )\n    df_eeg = df_eeg.iloc[eeg_offset*200:(eeg_offset+50)*200] \n    df_eeg = df_eeg.copy()\n    df_eeg['seconds_with_offset'] = range(df_eeg.shape[0]) \n    df_eeg['seconds_with_offset'] = df_eeg['seconds_with_offset']/sample_per_second +df[df.reset_index()['index']==GET_ROW]['eeg_label_offset_seconds'].unique()\n    color = plt.cm.viridis(i/8.0)\n    plt.plot(df_eeg['seconds_with_offset'], df_eeg['Fp1'], c='black')\n    plt.fill_between(df_eeg['seconds_with_offset'], -250, 0, where=[(x >= ((df_eeg['seconds_with_offset'].max()+df_eeg['seconds_with_offset'].min())/2-5)) and (x <= ((df_eeg['seconds_with_offset'].max()+df_eeg['seconds_with_offset'].min())/2+5)) for x in df_eeg['seconds_with_offset']], color=color, alpha=0.7, label=f'row {i}')\n    plt.axvline((df_eeg['seconds_with_offset'].max()+df_eeg['seconds_with_offset'].min())/2-5, color=color, linestyle='-', linewidth=2)\n    plt.axvline((df_eeg['seconds_with_offset'].max()+df_eeg['seconds_with_offset'].min())/2+5, color=color, linestyle='-', linewidth=2)\nplt.ylabel('Fp1')\nplt.legend()\nplt.grid(ls='--')    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:31.052150Z","iopub.execute_input":"2024-02-28T10:18:31.053095Z","iopub.status.idle":"2024-02-28T10:18:43.376714Z","shell.execute_reply.started":"2024-02-28T10:18:31.053060Z","shell.execute_reply":"2024-02-28T10:18:43.375507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display all the columns of a particular eeg_id, eeg_sub_id  and the ten seconds that we need to predict\nplt.figure(figsize=(12, 8))\nfor i, col in enumerate(df_eeg.columns[:-2]):\n    plt.subplot(len(df_eeg.columns),1,i+1)\n    plt.plot(df_eeg['seconds_with_offset'], df_eeg[col], c='black')\n    #plt.tick_params(axis='y', which='both', left=False, right=False, labelleft=False)\n    plt.fill_between(df_eeg['seconds_with_offset'], df_eeg[col].min(), df_eeg[col].max(), where=[(x >= (df_eeg['seconds_with_offset'].min()+df_eeg['seconds_with_offset'].max())/2-5) and (x <= (df_eeg['seconds_with_offset'].min()+df_eeg['seconds_with_offset'].max())/2+5) for x in df_eeg['seconds_with_offset']], color='lightblue', alpha=0.7)\n    plt.axvline((df_eeg['seconds_with_offset'].max()+df_eeg['seconds_with_offset'].min())/2-5, color='lightblue', linestyle='-', linewidth=2)\n    plt.axvline((df_eeg['seconds_with_offset'].max()+df_eeg['seconds_with_offset'].min())/2+5, color='lightblue', linestyle='-', linewidth=2)\n    plt.ylabel(col)\n    plt.grid(ls='--')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:18:43.378203Z","iopub.execute_input":"2024-02-28T10:18:43.378706Z","iopub.status.idle":"2024-02-28T10:19:09.233956Z","shell.execute_reply.started":"2024-02-28T10:18:43.378668Z","shell.execute_reply":"2024-02-28T10:19:09.232766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Relationship variables eeg/variables eeg \n\n# Define the dataframe of electrodes\ndf_electrode = df_eeg.drop(['EKG','seconds_with_offset'], axis=1)\n\n# Plot the correlation matrix between electrodes\nplt.figure(figsize=(14,8))\nsns.heatmap(df_electrode[df_electrode.columns].corr(), annot=True, cmap='viridis')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:09.235191Z","iopub.execute_input":"2024-02-28T10:19:09.235497Z","iopub.status.idle":"2024-02-28T10:19:10.519488Z","shell.execute_reply.started":"2024-02-28T10:19:09.235473Z","shell.execute_reply":"2024-02-28T10:19:10.518331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'm uncertain about the practical applications of this information. Perhaps we can explore how to associate electrodes to create various optimal montages?","metadata":{}},{"cell_type":"markdown","source":"# Checklist\n **Form Analysis**:\n\n- Rows and Columns: (10000, 20) + (0,1) to facilitate our comprehension, corresponding the 'eeg_label_offset_seconds' column. Each combination ('eeg_id', 'eeg_sub_id') corresponds to a 50 second long subsample starting at time 'eeg_label_offset_seconds' where 200 samples were taken each second. \n- Types of Variables: 20 float64\n- Analysis of Missing Variables: No missing value\n\n**Background Analysis**:\n\n- Significance of Variables: Each column represents a measure done by a particular electrode placed on the head of the patient.\n\n- Relationship Variables/variables: \n    * Variables seem to be generally correlated according to the distance in a defined montage\n    * Correlations between variables change according to the label_id\n","metadata":{}},{"cell_type":"markdown","source":"# **Analysis of one particular spectrogram_id from train_spectrograms.csv**","metadata":{}},{"cell_type":"code","source":"train_spec_path = '/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/'","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:10.521060Z","iopub.execute_input":"2024-02-28T10:19:10.521449Z","iopub.status.idle":"2024-02-28T10:19:10.527087Z","shell.execute_reply.started":"2024-02-28T10:19:10.521417Z","shell.execute_reply":"2024-02-28T10:19:10.525994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select the first row of the dataset train\nGET_ROW = 0\nrow = train.iloc[GET_ROW]\ndf_spectrogram = pd.read_parquet(f'{train_spec_path}{row.spectrogram_id}.parquet')\nspec_offset = int( row.spectrogram_label_offset_seconds )\ndf_spectrogram = df_spectrogram.loc[(df_spectrogram.time>=spec_offset)\n                     &(df_spectrogram.time<spec_offset+600)]","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:10.528933Z","iopub.execute_input":"2024-02-28T10:19:10.529628Z","iopub.status.idle":"2024-02-28T10:19:10.589800Z","shell.execute_reply.started":"2024-02-28T10:19:10.529589Z","shell.execute_reply":"2024-02-28T10:19:10.588663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the number of offset_seconds for a unique eeg_id\nprint(df.iloc[GET_ROW]['spectrogram_id'])\nprint(df[df['spectrogram_id'] == df.iloc[GET_ROW]['spectrogram_id']]['spectrogram_label_offset_seconds'].unique())\n\n# Print the unique offset_seconds for a specific row\nprint(df[df.reset_index()['index']==GET_ROW]['spectrogram_label_offset_seconds'])","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:10.591100Z","iopub.execute_input":"2024-02-28T10:19:10.591963Z","iopub.status.idle":"2024-02-28T10:19:10.607727Z","shell.execute_reply.started":"2024-02-28T10:19:10.591932Z","shell.execute_reply":"2024-02-28T10:19:10.606671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Copy the Data\ndf_spectrogram = df_spectrogram.copy()\n\n# Create columns time with offset_seconds\nsample_per_second = 0.5\ndf_spectrogram['seconds_with_offset'] = range(df_spectrogram.shape[0]) \ndf_spectrogram['seconds_with_offset'] = df_spectrogram['seconds_with_offset']/sample_per_second +df[df.reset_index()['index']==GET_ROW]['spectrogram_label_offset_seconds'].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:10.608961Z","iopub.execute_input":"2024-02-28T10:19:10.609284Z","iopub.status.idle":"2024-02-28T10:19:10.630500Z","shell.execute_reply.started":"2024-02-28T10:19:10.609256Z","shell.execute_reply":"2024-02-28T10:19:10.629535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Observe few lines \nprint(df_spectrogram.head())\n\n# Shape of the data\nprint('The shape of df is:', df_spectrogram.shape)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:10.632250Z","iopub.execute_input":"2024-02-28T10:19:10.633099Z","iopub.status.idle":"2024-02-28T10:19:10.654120Z","shell.execute_reply.started":"2024-02-28T10:19:10.633059Z","shell.execute_reply":"2024-02-28T10:19:10.652449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create columns\nColumn_name = list(df_spectrogram.columns)\n\n# Number of NaN in each column\nnumber_na = df_spectrogram.isna().sum()\n\n# Type of Data and the number\ntypes = df_spectrogram.dtypes\nnumber_types = df_spectrogram.dtypes.value_counts()\nprint(number_types)\n\n# Create a resume table\ndf_spectrogram_resume = pd.DataFrame({'features': Column_name, 'Type': types, 'Number of NaN': number_na})\nprint(df_spectrogram_resume)","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:10.655722Z","iopub.execute_input":"2024-02-28T10:19:10.656401Z","iopub.status.idle":"2024-02-28T10:19:10.671106Z","shell.execute_reply.started":"2024-02-28T10:19:10.656359Z","shell.execute_reply":"2024-02-28T10:19:10.669843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the 3 first rows for the columns LL_0.59 and observe that there are overlaps\nplt.figure(figsize=(12, 8))\nfor i in range(3):\n    GET_ROW = i\n    row = train.iloc[GET_ROW]\n    df_spectrogram = pd.read_parquet(f'{train_spec_path}{row.spectrogram_id}.parquet')\n    spec_offset = int( row.spectrogram_label_offset_seconds )\n    df_spectrogram = df_spectrogram.loc[(df_spectrogram.time>=spec_offset)\n                     &(df_spectrogram.time<spec_offset+600)]\n    df_spectrogram = df_spectrogram.copy()\n    df_spectrogram['seconds_with_offset'] = range(df_spectrogram.shape[0]) \n    df_spectrogram['seconds_with_offset'] = df_spectrogram['seconds_with_offset']/sample_per_second +df[df.reset_index()['index']==GET_ROW]['spectrogram_label_offset_seconds'].unique()\n    color = plt.cm.viridis(i/8.0)\n    plt.plot(df_spectrogram['seconds_with_offset'], df_spectrogram['LL_0.59'], c='black')\n    plt.fill_between(df_spectrogram['seconds_with_offset'], 0, 18, where=[(x >= df_spectrogram['seconds_with_offset'].min()) and (x <= df_spectrogram['seconds_with_offset'].max()) for x in df_spectrogram['seconds_with_offset']], color=color, alpha=0.7, label=f'row {i}')\n    plt.axvline(df_spectrogram['seconds_with_offset'].min(), color=color, linestyle='-', linewidth=2)\n    plt.axvline(df_spectrogram['seconds_with_offset'].max(), color=color, linestyle='-', linewidth=2)\nplt.ylabel('Fp1')\nplt.legend()\nplt.grid(ls='--')    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:10.673858Z","iopub.execute_input":"2024-02-28T10:19:10.674517Z","iopub.status.idle":"2024-02-28T10:19:11.198816Z","shell.execute_reply.started":"2024-02-28T10:19:10.674477Z","shell.execute_reply":"2024-02-28T10:19:11.197867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the full spectrogram_id for the columns LL_0.59 and the ten seconds that we need to predict and observe that there are overlaps\nplt.figure(figsize=(12, 8))\nfor i in range(9):\n    GET_ROW = i\n    row = train.iloc[GET_ROW]\n    df_spectrogram = pd.read_parquet(f'{train_spec_path}{row.spectrogram_id}.parquet')\n    spec_offset = int( row.spectrogram_label_offset_seconds )\n    df_spectrogram = df_spectrogram.loc[(df_spectrogram.time>=spec_offset)\n                     &(df_spectrogram.time<spec_offset+600)]\n    df_spectrogram = df_spectrogram.copy()\n    df_spectrogram['seconds_with_offset'] = range(df_spectrogram.shape[0]) \n    df_spectrogram['seconds_with_offset'] = df_spectrogram['seconds_with_offset']/sample_per_second +df[df.reset_index()['index']==GET_ROW]['spectrogram_label_offset_seconds'].unique()\n    color = plt.cm.viridis(i/8.0)\n    plt.plot(df_spectrogram['seconds_with_offset'], df_spectrogram['LL_0.59'], c='black')\n    plt.fill_between(df_spectrogram['seconds_with_offset'], 0, 18, where=[(x >= ((df_spectrogram['seconds_with_offset'].max()+df_spectrogram['seconds_with_offset'].min())/2-5)) and (x <= ((df_spectrogram['seconds_with_offset'].max()+df_spectrogram['seconds_with_offset'].min())/2+5)) for x in df_spectrogram['seconds_with_offset']], color=color, alpha=0.7, label=f'row {i}')\n    plt.axvline((df_spectrogram['seconds_with_offset'].max()+df_spectrogram['seconds_with_offset'].min())/2-5, color=color, linestyle='-', linewidth=2)\n    plt.axvline((df_spectrogram['seconds_with_offset'].max()+df_spectrogram['seconds_with_offset'].min())/2+5, color=color, linestyle='-', linewidth=2)\nplt.ylabel('Fp1')\nplt.legend()\nplt.grid(ls='--')    \nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:11.200176Z","iopub.execute_input":"2024-02-28T10:19:11.201169Z","iopub.status.idle":"2024-02-28T10:19:12.222199Z","shell.execute_reply.started":"2024-02-28T10:19:11.201136Z","shell.execute_reply":"2024-02-28T10:19:12.220773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the columns (LL_0.59, RL_0.59, LP_0.59, RP_0.59) of a particular spectrogram_id, spectrogram__sub_id  and the ten seconds that we need to predict\nplt.figure(figsize=(12, 8))\nfor i, col in enumerate(['LL_0.59', 'RL_0.59', 'LP_0.59', 'RP_0.59']):\n    plt.subplot(len(['LL_0.59', 'RL_0.59', 'LP_0.59', 'RP_0.59']),1,i+1)\n    plt.plot(df_spectrogram['seconds_with_offset'], df_spectrogram[col], c='black')\n    #plt.tick_params(axis='y', which='both', left=False, right=False, labelleft=False)\n    plt.fill_between(df_spectrogram['seconds_with_offset'], df_spectrogram[col].min(), df_spectrogram[col].max(), where=[(x >= (df_spectrogram['seconds_with_offset'].min()+df_spectrogram['seconds_with_offset'].max())/2-5) and (x <= (df_spectrogram['seconds_with_offset'].min()+df_spectrogram['seconds_with_offset'].max())/2+5) for x in df_spectrogram['seconds_with_offset']], color='lightblue', alpha=0.7)\n    plt.axvline((df_spectrogram['seconds_with_offset'].max()+df_spectrogram['seconds_with_offset'].min())/2-5, color='lightblue', linestyle='-', linewidth=2)\n    plt.axvline((df_spectrogram['seconds_with_offset'].max()+df_spectrogram['seconds_with_offset'].min())/2+5, color='lightblue', linestyle='-', linewidth=2)\n    plt.ylabel(col)\n    plt.grid(ls='--')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:12.223807Z","iopub.execute_input":"2024-02-28T10:19:12.224224Z","iopub.status.idle":"2024-02-28T10:19:12.949415Z","shell.execute_reply.started":"2024-02-28T10:19:12.224184Z","shell.execute_reply":"2024-02-28T10:19:12.948124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Relationship variables spectrogram/variables sprectrogram zoom\n\n# Define the dataframe of montage\ndf_montage = df_spectrogram[['LL_0.78', 'RL_0.78', 'LP_0.78', 'RP_0.78']]\n\n# Plot the correlation matrix between electrodes\nsns.heatmap(df_montage[df_montage.columns].corr(), annot=True, cmap='viridis')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:19:12.950987Z","iopub.execute_input":"2024-02-28T10:19:12.951528Z","iopub.status.idle":"2024-02-28T10:19:13.230397Z","shell.execute_reply.started":"2024-02-28T10:19:12.951489Z","shell.execute_reply":"2024-02-28T10:19:13.229176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Checklist\n\n**Form Analysis**:\n\n- Rows and Columns: (300, 401). Each combination ('spectrograms_id', 'spectrograms_sub_id') corresponds to a 600 seconds (ten minutes) long subsample starting at time 'spectrograms_label_offset_seconds' where 1 sample was taken each 2 seconds. \n- Types of Variables: 400 float32, 1 int64\n- Analysis of Missing Variables: No missing value\n\n**Background Analysis**:\n\n- Significance of Variables:\n    * Each column signifies a measurement conducted by a specific group of electrodes (LL, LP, RL, RP) positioned at a distinct location on the patient's head, along with a corresponding time column. Each suffix in the column header corresponds to a specific frequency of brain activity.\n    * Observe that different spectrograms_sub_id with the same id coincide for a large part \n\n- Relationship Variables/variables: \n    * Variables RL and RP with the same suffix number seem to be generally highly correlated (>88%) \n    * Correlations between variables change according to the label_id","metadata":{}},{"cell_type":"markdown","source":"# Patient Variation","metadata":{}},{"cell_type":"markdown","source":"The rest of this EDA was taken from the notebook \"Patient Variation - EDA\" done by @pcjimmmy (link here https://www.kaggle.com/code/pcjimmmy/patient-variation-eda/notebook) that we sum up here. \n\nTo enhance understanding, several new columns have been defined:\n* total_evaluators : number of evaluator for one eeg_id\n* consensus : maximum number of votes between each cerebral activity\n* agreement : ratio between consensus and total_evaluators (consensus/total_evaluators)\n\nRemarks: \n- It appears that the training data may have been aggregated from at least two different studies, with varying numbers of evaluators participating the first study has less than 6 voters whereas the second one has more that 9 voters. (Considering the weight of the data, those with more evaluators should carry more significance, particularly if a strong consensus is observed)\n- Rows with a small number of evaluators often receive perfect ratings (1.0), raising questions about their reliability.\n- Rows with very low agreement on consensus (20% or less) present a significant challenge. The decision on how to handle such problematic cases remains uncertain.\n- The rating difficulty varies, with 'seizure_vote' appearing easier to rate compared to 'gpd_vote,' hinting at potential differences in the studies contributing to the training data.\n    -  Seizure_vote - Patients with this condition either exhibit consistent EEG patterns or their symptoms are easily recognizable and agreeable among evaluators\n    - gpd_vote: Patients with this condition either do not display the issue consistently over time (it comes and goes), or it is challenging for evaluators to identify. The plot suggests that there are no EEG recordings for this type where all evaluators unanimously agreed.\n    - Conclusion: Stuies relying only on the 'first' may miss valuable information. \n- Descriptions of eegs where experts agree are termed \"idealized.\" However, questions arise about the confidence in these idealized descriptions when the number of experts is limited. Similarly, when only a single eeg is available for a patient, even with agreement from many experts, the confidence in labeling it as 'ideal' is questioned.\n- Analyzing the distribution of ratings reveals unexpected results. For Seizure, when the number of evaluators is six or fewer, it is the major label. However, with more than nine evaluators, it becomes the least observed. The small vs. large group of evaluators presents a challenge, with the predominance of 'other' in larger groups potentially indicating measurement error or different study characteristics.\n- The issue of outliers, especially in the Seizure category, prompts consideration for removal. Additionally, the expected agreement for 'lateral' is not observed, challenging assumptions about the reliability of right vs. left eeg distributions.\n","metadata":{}},{"cell_type":"markdown","source":"# **To sum up**","metadata":{}},{"cell_type":"markdown","source":"# 1. train.csv:\n\n**Initial Form Analysis**:\n\n- Target Variable: (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote)\n- Rows and Columns: (106800, 15)\n- Types of Variables: 12 int64, 2 float64, 1 object\n- Analysis of Missing Variables: No missing value\n\n**Initial Background Analysis**:\n\n- Target Visualization: The target variable is well balanced. Number of expert consensus for each type of activity is between 15000 and 2000\n\n- Significance of Variables:\n    * label_id: Nothing to say since it's a unique number representing each cases in the dataset\n    * eeg_id: There are 17089 different eeg_id. Each eeg_id contains eeg_sub_id which can go from 0 to 742 associated to a eeg_label_offset_seconds which goes from 0.0 to 3372.0. Thus the occurency of a eeg_id goes from 1 to 743. A eeg_id is associated to a unique spectrogram_id.\n    * spectrogram_id: There are 11138 different spectrogram_id Each spectrogram_id contains spectrogram_sub_id which can go from 0 to 1021 associated to a spectrogram_label_offset_seconds which goes from 0.0 to 17632.0?. Thus the occurency of a spectrogram_id goes from 1 to 1022. A spectrogram_id can contain different eeg_id  \n    * patient_id: There are 1950 different patient_id. The occurrence is between 1 and 2215 corresponding to different combination of ('eeg_id','eeg_sub_id', 'eeg_label_offset_seconds', 'spectrogram_id', 'spectrogram_sub_id', 'spectrogram_label_offset_seconds').  \n    \n    \n- Relationship Target/target: it seems that seizure is correlated with grda_vote (24%), other_vote (21%), lrda_vote (17%), gpd_vote (13%), lpd_vote (13%). lrda_vote and gpd_vote (15%). The rest seems to be negligeable (-8%).\n\n- Relationship Variables/Target: \n    * patient_id/expert_consensus: Each patient_id can have more than one expert_consensus corresponding to a particular configuration (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote).\n    * eeg_id, eeg_sub_id, eeg_label_offset_seconds/expert_consensus: Each eeg_id can have different expert consensus based on the eeg_sub_id or eeg_label_offset_seconds we look at.\n    * spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds/expert_consensus: Each spectrogram_id can have different expert consensus based on spectrogram_sub_id, spectrogram_label_offset_seconds \n      but also the eeg_id we look at.\n      \n- Relationship Variables/variables: \n    * eeg_id/spectrogram_id: A eeg_id is associated to a unique spectrogram_id but a spectrogram_id can contain different eeg_id.\n    * eeg_id/patient_id and spectrogram_id/patient_id: Each patient can have several different eeg_id and spectrogram_id.\n    \n    ","metadata":{}},{"cell_type":"markdown","source":"**detailled Analysis**\n\n- It appears that the training data may have been aggregated from at least two different studies, with varying numbers of evaluators participating the first study has less than 6 voters whereas the second one has more that 9 voters. (Considering the weight of the data, those with more evaluators should carry more significance, particularly if a strong consensus is observed)\n- Rows with a small number of evaluators often receive perfect ratings (1.0), raising questions about their reliability.\n- Rows with very low agreement on consensus (20% or less) present a significant challenge. The decision on how to handle such problematic cases remains uncertain.\n- The rating difficulty varies, with 'seizure_vote' appearing easier to rate compared to 'gpd_vote,' hinting at potential differences in the studies contributing to the training data.\n    -  Seizure_vote - Patients with this condition either exhibit consistent EEG patterns or their symptoms are easily recognizable and agreeable among evaluators\n    - gpd_vote: Patients with this condition either do not display the issue consistently over time (it comes and goes), or it is challenging for evaluators to identify. The plot suggests that there are no EEG recordings for this type where all evaluators unanimously agreed.\n    - Conclusion: Stuies relying only on the 'first' may miss valuable information. \n- Descriptions of eegs where experts agree are termed \"idealized.\" However, questions arise about the confidence in these idealized descriptions when the number of experts is limited. Similarly, when only a single eeg is available for a patient, even with agreement from many experts, the confidence in labeling it as 'ideal' is questioned.\n- Analyzing the distribution of ratings reveals unexpected results. For Seizure, when the number of evaluators is six or fewer, it is the major label. However, with more than nine evaluators, it becomes the least observed. The small vs. large group of evaluators presents a challenge, with the predominance of 'other' in larger groups potentially indicating measurement error or different study characteristics.\n- The issue of outliers, especially in the Seizure category, prompts consideration for removal. Additionally, the expected agreement for 'lateral' is not observed, challenging assumptions about the reliability of right vs. left eeg distributions.","metadata":{}},{"cell_type":"markdown","source":"# 2. train_eegs (for a particular eeg):\n\n **Form Analysis**:\n\n- Rows and Columns: (10000, 20) + (0,1) to facilitate our comprehension, corresponding the 'eeg_label_offset_seconds' column. Each combination ('eeg_id', 'eeg_sub_id') corresponds to a 50 second long subsample starting at time 'eeg_label_offset_seconds' where 200 samples were taken each second. \n- Types of Variables: 20 float64\n- Analysis of Missing Variables: No missing value\n\n**Background Analysis**:\n\n- Significance of Variables: Each column represents a measure done by a particular electrode placed on the head of the patient.\n\n- Relationship Variables/variables: \n    * Variables seem to be generally correlated according to the distance in a defined montage\n    * Correlations between variables change according to the label_id","metadata":{}},{"cell_type":"markdown","source":"# 3. train_spectrograms (for a particular spectrogram):\n\n**Form Analysis**:\n\n- Rows and Columns: (300, 401). Each combination ('spectrograms_id', 'spectrograms_sub_id') corresponds to a 600 seconds (ten minutes) long subsample starting at time 'spectrograms_label_offset_seconds' where 1 sample was taken each 2 seconds. \n- Types of Variables: 400 float32, 1 int64\n- Analysis of Missing Variables: No missing value\n\n**Background Analysis**:\n\n- Significance of Variables:\n    * Each column signifies a measurement conducted by a specific group of electrodes (LL, LP, RL, RP) positioned at a distinct location on the patient's head, along with a corresponding time column. Each suffix in the column header corresponds to a specific frequency of brain activity.\n    * Observe that different spectrograms_sub_id with the same id coincide for a large part \n\n- Relationship Variables/variables: \n    * Variables RL and RP with the same suffix number seem to be generally highly correlated (>88%) \n    * Correlations between variables change according to the label_id","metadata":{}},{"cell_type":"markdown","source":"# Thank you for reading !","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}