{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7392733,"sourceType":"datasetVersion","datasetId":4297749},{"sourceId":7414022,"sourceType":"datasetVersion","datasetId":4312784},{"sourceId":7526248,"sourceType":"datasetVersion","datasetId":4308295},{"sourceId":7898593,"sourceType":"datasetVersion","datasetId":4638484},{"sourceId":8121200,"sourceType":"datasetVersion","datasetId":4798711},{"sourceId":6125,"sourceType":"modelInstanceVersion","modelInstanceId":4596},{"sourceId":6127,"sourceType":"modelInstanceVersion","modelInstanceId":4598}],"dockerImageVersionId":30683,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n    <img src=\"https://i.imgur.com/1v2C2o7.png\" alt=\"Image Description\">\n</div>\n\n\n<div style=\"text-align:center;\">\n<h1><strong> Harmful Brain Activity Classification</strong></h1>\n    <h2>Harvard Medical School</h2> </div>","metadata":{}},{"cell_type":"markdown","source":"# Objectives and Goal\n* To detect and classify seizures and other types of harmful brain activity in critical ill patients.\n* To develop a model trained on electroencephalography (EEG) signals recorded from critically ill hospital patients.\n* To create a model that can automates EEG analysis will help doctors and brain researchers detect seizures and other types of brain activity that can cause brain damage, so that they can give treatments more quickly and accurately. \n* May also help researchers who are working to develop drugs to treat and prevent seizures.\n","metadata":{}},{"cell_type":"markdown","source":"<h2> Dataset Flowchart </h2>\n<a href=\"https://imgur.com/9pvOTk6\"><img src=\"https://i.imgur.com/9pvOTk6.png\" title=\"source: imgur.com\" /></a>","metadata":{"execution":{"iopub.status.busy":"2024-04-15T01:53:13.422210Z","iopub.execute_input":"2024-04-15T01:53:13.422580Z","iopub.status.idle":"2024-04-15T01:53:13.429019Z","shell.execute_reply.started":"2024-04-15T01:53:13.422551Z","shell.execute_reply":"2024-04-15T01:53:13.427145Z"}}},{"cell_type":"markdown","source":"## Patterns for Classification:\n1. Seizure (SZ)\n2. Generalized periodic discharges (GPD)\n3. Lateralized periodic discharges (LPD)\n4. Lateralized rhythmic delta activity (LRDA)\n5. Generalized rhythmic delta activity (GRDA)\n6. Other\n\n##### Detailed explanation of above patterns:\n* https://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf \n\n\n* In some cases experts completely agree about the correct label. On other cases the experts disagree. They call segments where there are high levels of agreement “idealized” patterns. Cases where ~1/2 of experts give a label as “other” and ~1/2 give one of the remaining five labels, we call “proto patterns”. Cases where experts are approximately split between 2 of the 5 named patterns, we call “edge cases”.","metadata":{"execution":{"iopub.status.busy":"2024-04-15T01:53:01.194450Z","iopub.execute_input":"2024-04-15T01:53:01.195051Z","iopub.status.idle":"2024-04-15T01:53:01.204939Z","shell.execute_reply.started":"2024-04-15T01:53:01.195021Z","shell.execute_reply":"2024-04-15T01:53:01.203625Z"}}},{"cell_type":"markdown","source":"1. High level of agreement                        - Idealized pattern\n2. Half of experts give a label                    - Other\n3. Half of experts give one of the remaining five label      - proto patterns\n4. Experts split between 2 of the 5 named patterns - edge cases","metadata":{}},{"cell_type":"markdown","source":"## Import Libraries","metadata":{}},{"cell_type":"code","source":"import os\nos.environ[\"KERAS_BACKEND\"] = \"jax\"\nimport pandas as pd\nimport numpy as np\nimport keras_cv\nimport keras\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport shutil\nfrom tqdm.notebook import tqdm\nimport multiprocessing\nimport librosa.display\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom sklearn.metrics import classification_report, confusion_matrix, precision_score, recall_score, roc_auc_score\nimport tensorflow as tf\nimport tensorflow_hub as hub\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nstrategy =tf.distribute.MirroredStrategy()\nprint('DEVICE AVALABLE: {}'. format(strategy.num_replicas_in_sync))\n\n# Use mixed precision\ntf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-21T21:07:27.450654Z","iopub.execute_input":"2024-04-21T21:07:27.451506Z","iopub.status.idle":"2024-04-21T21:07:47.188892Z","shell.execute_reply.started":"2024-04-21T21:07:27.451471Z","shell.execute_reply":"2024-04-21T21:07:47.187626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !pip install -q /kaggle/input/kerasv3-lib-ds/keras_cv-0.8.2-py3-none-any.whl --no-deps\n# !pip install -q /kaggle/input/kerasv3-lib-ds/tensorflow-2.15.0.post1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl --no-deps\n# !pip install -q /kaggle/input/kerasv3-lib-ds/keras-3.0.4-py3-none-any.whl --no-deps","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:48.460924Z","iopub.execute_input":"2024-04-15T02:24:48.461967Z","iopub.status.idle":"2024-04-15T02:24:48.466219Z","shell.execute_reply.started":"2024-04-15T02:24:48.461889Z","shell.execute_reply":"2024-04-15T02:24:48.465219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nprint(\"TensorFlow:\", tf.__version__)\n\nimport keras\nprint(\"Keras:\", keras.__version__)\n\nimport keras_cv\nprint(\"KerasCV:\", keras_cv.__version__)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:07:47.190573Z","iopub.execute_input":"2024-04-21T21:07:47.191280Z","iopub.status.idle":"2024-04-21T21:07:47.197096Z","shell.execute_reply.started":"2024-04-21T21:07:47.191252Z","shell.execute_reply":"2024-04-21T21:07:47.196213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.random.set_seed(42)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:48.480806Z","iopub.execute_input":"2024-04-15T02:24:48.481162Z","iopub.status.idle":"2024-04-15T02:24:48.487550Z","shell.execute_reply.started":"2024-04-15T02:24:48.481134Z","shell.execute_reply":"2024-04-15T02:24:48.486659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\nprint(f\"Shape of the train metadata:\", train_df.shape)\ndisplay(train_df.head())","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:36.573289Z","iopub.execute_input":"2024-04-21T21:08:36.573654Z","iopub.status.idle":"2024-04-21T21:08:36.851879Z","shell.execute_reply.started":"2024-04-21T21:08:36.573625Z","shell.execute_reply":"2024-04-21T21:08:36.850344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Unique eeg_id associated with unqiue expert consensus\ntrain_df= train_df.drop_duplicates(subset=['eeg_id', 'expert_consensus'])\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:36.907378Z","iopub.execute_input":"2024-04-21T21:08:36.907862Z","iopub.status.idle":"2024-04-21T21:08:36.938825Z","shell.execute_reply.started":"2024-04-21T21:08:36.907821Z","shell.execute_reply":"2024-04-21T21:08:36.937927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Target columns stored in a variable\ntarget_classes = [col for col in train_df.columns if col.endswith(\"_vote\")]\nprint(list(target_classes))\n\n\nclass_names = ['Seizure', 'LPD', 'GPD', 'LRDA','GRDA', 'Other']\nlabel_dict = dict(zip(class_names, range(len(class_names))))\nlabel_dict\n\n# Give Label to the classes \ntrain_df['target_name'] = train_df.expert_consensus.copy()\ntrain_df['target_labels'] = train_df['expert_consensus'].map(label_dict)\n\ntrain_df['target_labels'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:37.246516Z","iopub.execute_input":"2024-04-21T21:08:37.247430Z","iopub.status.idle":"2024-04-21T21:08:37.268175Z","shell.execute_reply.started":"2024-04-21T21:08:37.247395Z","shell.execute_reply":"2024-04-21T21:08:37.267284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read Train EEG and Spectrogram Data\n<a href=\"https://imgur.com/kf1sLtK\"><img src=\"https://i.imgur.com/kf1sLtK.png\" title=\"source: imgur.com\" /></a>","metadata":{}},{"cell_type":"code","source":"# Paths\nbase_path = \"/kaggle/input/hms-harmful-brain-activity-classification\"\ntrain_eeg_path = os.path.join(base_path, \"train_eegs\")\ntrain_spec_path = os.path.join(base_path, \"train_spectrograms\")\n","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:37.982490Z","iopub.execute_input":"2024-04-21T21:08:37.982903Z","iopub.status.idle":"2024-04-21T21:08:37.987791Z","shell.execute_reply.started":"2024-04-21T21:08:37.982872Z","shell.execute_reply":"2024-04-21T21:08:37.986782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get first file\nspec_file = os.listdir(train_spec_path)\nfile= spec_file[0]\nprint(file)\n\n# DIsplay first file of train spectrograms\ntrain_parquet1 = pd.read_parquet(f'{train_spec_path}/{file}')\ntrain_parquet1.head(2)\n\n# pd.set_option(\"display.max_columns\",None)\n# train_parquet1","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:38.349400Z","iopub.execute_input":"2024-04-21T21:08:38.349797Z","iopub.status.idle":"2024-04-21T21:08:38.793289Z","shell.execute_reply.started":"2024-04-21T21:08:38.349765Z","shell.execute_reply":"2024-04-21T21:08:38.792318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the Spectrogram data","metadata":{}},{"cell_type":"code","source":"## Visualize the single LL0.59 column of 32728416.parquet spectrogram file\nplt.figure(figsize=(10, 5))  \nplt.plot(train_parquet1['time'], train_parquet1['LL_0.59'], label='LL_0.59')\nplt.xlabel('Time')\nplt.ylabel('Frequency')\nplt.title('Left Lateral(0.59) over Time')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:39.066014Z","iopub.execute_input":"2024-04-21T21:08:39.066368Z","iopub.status.idle":"2024-04-21T21:08:39.446730Z","shell.execute_reply.started":"2024-04-21T21:08:39.066340Z","shell.execute_reply":"2024-04-21T21:08:39.445691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Read first EEG file\n\n* EEG Data: The EEG recordings contain electrical activity data captured from electrodes placed on the scalp. ","metadata":{}},{"cell_type":"code","source":"# Get first file\neeg_file = os.listdir(train_eeg_path)\nfile= eeg_file[0]\nprint(file)\n\n# DIsplay first file of train EEG\ntrain_parquet_eeg = pd.read_parquet(f'{train_eeg_path}/{file}')\ntrain_parquet_eeg.head(2)","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:39.657173Z","iopub.execute_input":"2024-04-21T21:08:39.658109Z","iopub.status.idle":"2024-04-21T21:08:39.965871Z","shell.execute_reply.started":"2024-04-21T21:08:39.658073Z","shell.execute_reply":"2024-04-21T21:08:39.964955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the EEG Data","metadata":{}},{"cell_type":"code","source":"## Visualize the single Fp1 electrode of 2208063991.parquet EEG file\nplt.figure(figsize=(10,5))\nplt.plot(train_parquet_eeg[\"Fp1\"], label = \"Fp1\") \nplt.xlabel('Time')\nplt.ylabel('Amplitude')\nplt.title('EEG Data from Fp1 Electrode')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:40.461183Z","iopub.execute_input":"2024-04-21T21:08:40.461567Z","iopub.status.idle":"2024-04-21T21:08:40.791695Z","shell.execute_reply.started":"2024-04-21T21:08:40.461537Z","shell.execute_reply":"2024-04-21T21:08:40.790759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h1><strong> Explanatory Analysis</strong></h1>","metadata":{}},{"cell_type":"markdown","source":"## Class distribution based on experts votes","metadata":{}},{"cell_type":"code","source":"# Calculate class distribution based on votes\nclass_distribution = train_df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].sum()\n\na=[\"#d7658b\", \"#e27c7c\", \"#98d1d1\",  \"#e2e2e2\", \"#e1a692\",  \"#776bcd\"]\n# Plotting the pie chart\nplt.figure(figsize=(8, 8))\nplt.pie(class_distribution, labels=class_distribution.index, autopct='%1.1f%%', colors=a)\nplt.title('Class Distribution Based on Votes')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:46:39.799599Z","iopub.execute_input":"2024-04-21T21:46:39.800296Z","iopub.status.idle":"2024-04-21T21:46:40.012364Z","shell.execute_reply.started":"2024-04-21T21:46:39.800262Z","shell.execute_reply":"2024-04-21T21:46:40.011052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the provided percentages, the class distribution appears moderately imbalanced, with the \"Other\" class being the most dominant. ","metadata":{}},{"cell_type":"code","source":"## Barplot of expert's label based on the majority votes to the labes\nmax_expert_consensus_votes = train_df[\"expert_consensus\"].value_counts()\nplt.figure(figsize=(10, 10))\nsns.barplot(x=max_expert_consensus_votes.index, y=max_expert_consensus_votes.values, palette = 'crest')\nplt.title(\"Counts of Expert Labels Based on Majority Votes\", fontsize=12, fontweight = 'bold')\nplt.xlabel(\" Expert Consensus\", fontsize=10, fontweight = 'bold')\nplt.ylabel( \"Majority Votes\", fontsize=10, fontweight = 'bold')\n\n# Add labels to the bars\nfor index, value in enumerate(max_expert_consensus_votes):\n    plt.text(index, value + 0.5, str(value), ha='center', va='bottom')\nsize=1\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:49.716442Z","iopub.execute_input":"2024-04-15T02:24:49.717031Z","iopub.status.idle":"2024-04-15T02:24:50.076054Z","shell.execute_reply.started":"2024-04-15T02:24:49.716984Z","shell.execute_reply":"2024-04-15T02:24:50.075060Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><strong> 1.Feature Engineering </strong>\n    <br><br>Using Below Information </h1>\n\n1. High level of agreement                        - Idealized pattern\n2. Half of experts give a label                    - Other\n3. Half of experts give remaining five label      - proto patterns\n4. Experts split between 2 of the 5 named patterns - edge cases","metadata":{}},{"cell_type":"code","source":"# function to generate the new column based on the rules to understand the data\ndef generate_pattern_type(df):\n    votes = [df['seizure_vote'], df['lpd_vote'], df['gpd_vote'], df['lrda_vote'], df['grda_vote']]\n    total_votes = sum(votes)\n    max_vote = max(votes)\n    other_vote = df['other_vote']\n    \n    # Rule 1: All experts agreed on one label\n    if  total_votes == max_vote:\n        return 'Idealized'\n    # Rule 2: ~1/2 of experts give a label as \"other\" and ~1/2 give one of the remaining five labels\n             ## used abs() to avoid negative numbers\n    elif abs(other_vote - total_votes / 2) <= 1 and max_vote == total_votes / 2:\n        return 'Proto'\n    # Rule 3: Experts are approximately split between 2 of the 5 named patterns\n    elif max_vote >= total_votes / 2  or len([vote for vote in votes if vote == max_vote]) == 2:\n        return 'Edge '\n    else:\n        return 'Undefined'  # other label for cases not covered by the rules\n\n# Apply the function \ntrain_df['pattern_type'] = train_df.apply(generate_pattern_type, axis=1)\n\n# # Display the DataFrame with the new column\ntrain_df.head(2)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:50.077329Z","iopub.execute_input":"2024-04-15T02:24:50.077696Z","iopub.status.idle":"2024-04-15T02:24:50.692144Z","shell.execute_reply.started":"2024-04-15T02:24:50.077669Z","shell.execute_reply":"2024-04-15T02:24:50.691151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"value = train_df['pattern_type'].value_counts().values\nsns.barplot(x= train_df['pattern_type'].unique(), y=value , palette ='crest')\n\nfor i, value in enumerate(value):\n    plt.text(i, value +0.5, str(value), ha= 'center', va='bottom')","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:50.693401Z","iopub.execute_input":"2024-04-15T02:24:50.693779Z","iopub.status.idle":"2024-04-15T02:24:50.877883Z","shell.execute_reply.started":"2024-04-15T02:24:50.693746Z","shell.execute_reply":"2024-04-15T02:24:50.876941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize Spectrogram Data Based On Idealized Pattern Category","metadata":{}},{"cell_type":"code","source":"#  Get Unique spectrogram_id with the Idealized pattern types\nunique_spectrogram_ids = train_df['spectrogram_id'].unique()\nidealized_pattern_data = train_df[train_df['pattern_type'] == 'Idealized'].reset_index(drop=True)\n\nidealized_pattern_data\n\n# define number of spectrogram to display in each row\nn=3\n\n# initialize an empty dictionary to store the list of spectrogram ids for each idealized pattern category\nspec_dict = {\n    \"seizure_vote\": [],\n    \"lpd_vote\": [],\n    \"gpd_vote\": [],\n    \"lrda_vote\": [],\n    \"grda_vote\": [],\n    \"other_vote\": []\n}\n\n\nfor key in spec_dict.keys():\n    #sort the spectrogrram based on the category and select the top 3\n    idx = idealized_pattern_data[key].sort_values(ascending= False).head(n).index\n    spec_dict[key] = idealized_pattern_data.loc[idx, \"spectrogram_id\"].values\n    \n# spec_dict","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:50.879151Z","iopub.execute_input":"2024-04-15T02:24:50.880035Z","iopub.status.idle":"2024-04-15T02:24:50.903154Z","shell.execute_reply.started":"2024-04-15T02:24:50.879999Z","shell.execute_reply":"2024-04-15T02:24:50.902484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Convert the EEG signals into spectrogram images. You can use libraries like matplotlib or scipy to generate spectrograms from the EEG signals.","metadata":{}},{"cell_type":"code","source":"def plot_spectrogram_by_cateogry(spect_dict):\n    train_spec_path = \"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/\"\n\n    fig, axes = plt.subplots(len(spec_dict), n, figsize= (15, len(spec_dict)*5))\n    plt.subplots_adjust(wspace=0.3, hspace=0.5)\n    \n    for i , (key, values) in enumerate(spec_dict.items()):\n        for j, value in enumerate(values):\n            \n            spec_data = pd.read_parquet(train_spec_path + str(value) + '.parquet')\n            #plot the spectrogram\n            ax = axes[i, j]\n            ax.imshow(np.log(spec_data.T), aspect ='auto')\n            ax.set_title(f\"{key} - ID: {value}\", size=12)\n            ax.set_xlabel(\"Time\", size=10)\n            ax.set_ylabel(\"(Hz)\", size=10)\n            ax.tick_params(axis='both', which='both', labelsize=8)\n            \n    plt.show()\n\nplot_spectrogram_by_cateogry(spec_dict)\n            ","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:50.904137Z","iopub.execute_input":"2024-04-15T02:24:50.904482Z","iopub.status.idle":"2024-04-15T02:24:56.510277Z","shell.execute_reply.started":"2024-04-15T02:24:50.904458Z","shell.execute_reply":"2024-04-15T02:24:56.509158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Adding EEG and Spectrogram Paths to the Train CSV DataFrame\n*  Adding 'eeg_path' and 'spec_path' columns to the train_df can efficiently scale model to handle larger datasets without being limited by memory constraints. This scalability is essential for handling real-world large datasets .\n* 'eeg_path' provides the location of the EEG data files associated with each 'eeg_id'. This allows the model to access the actual data when needed for training or evaluation.\n* 'spec_path' provides the location of the spectrogram data files associated with 'spec_id. This allows the model to access the actual data during training or evaluation.","metadata":{}},{"cell_type":"markdown","source":"## Convert spectrogram .parquet file into the numpy\n\n#### Reasons:\n* Compatibility: Tensorflow and pytorch expect the input data to be in the form of Numpy arrays.\n* Efficient Data Loading: NumPy arrays are stored in a binary format, which allows for faster loading compared to text-based formats like Parquet.Since working with larger dataset as it reduce the overhead associated with data loading and improves overall training efficiency.\n* Ease of Manipulation: NumPy arrays provide a convenient interface for performing various operations and manipulations on the data, such as normalization, reshaping, and augmentation.","metadata":{}},{"cell_type":"code","source":"\n# # # Run this code to delete the working data before generating saving the new npy data to remove the space\n\n# Specify the directory path\ndirectory_path = \"/tmp/dataset/hms-hbac\"  # Replace with the path to your directory\n\n# Check if the directory exists before attempting to delete it\nif os.path.exists(directory_path):\n    shutil.rmtree(directory_path)  # Use shutil.rmtree() to remove a directory and its contents\n    print(f\"Directory '{directory_path}' and its contents have been deleted successfully.\")\nelse:\n    print(f\"Directory '{directory_path}' does not exist.\")","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:56.511728Z","iopub.execute_input":"2024-04-15T02:24:56.512789Z","iopub.status.idle":"2024-04-15T02:24:57.349052Z","shell.execute_reply.started":"2024-04-15T02:24:56.512739Z","shell.execute_reply":"2024-04-15T02:24:57.348084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from concurrent.futures import ProcessPoolExecutor\n# #Define a function to process with single spectrogram file\ndef process_spectrogram(spec_id, base_path,split = 'train'):\n    spec_path =os.path.join(base_path, f\"{split}_spectrograms\", f\"{spec_id}.parquet\")\n    spec_df = pd.read_parquet(spec_path)\n    spec_values = spec_df.fillna(0).values[:, 1:].T.astype(np.float32) # removing the time column\n    spec_dir = os.path.join(\"/tmp/dataset/hms-hbac\", f\"{split}_spectrograms\")\n    os.makedirs(spec_dir, exist_ok=True)\n    \n    np.save(os.path.join(spec_dir,f\"{spec_id}.npy\"),spec_values)\n \n\n\n # Define a function to process spectrograms\ndef process_spec_wrapper(spec_id):\n    process_spectrogram(spec_id, base_path, split='train')\n\n# Get unique spectrogram IDs\nspec_ids = train_df['spectrogram_id'].unique()\n\n# Create a ProcessPoolExecutor with a specified number of processes\nwith ProcessPoolExecutor() as executor:\n    # Submit the process_spec_wrapper function to the executor for each spectrogram ID\n    list(tqdm(executor.map(process_spec_wrapper, spec_ids), total=len(spec_ids)))   ","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:24:57.350505Z","iopub.execute_input":"2024-04-15T02:24:57.351176Z","iopub.status.idle":"2024-04-15T02:28:11.646475Z","shell.execute_reply.started":"2024-04-15T02:24:57.351136Z","shell.execute_reply":"2024-04-15T02:28:11.645181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numpy_spec_path = \"/tmp/dataset/hms-hbac/train_spectrograms/\"\ntrain_df['spec_path_numpy'] = f'{numpy_spec_path}' + train_df['spectrogram_id'].astype(str) + '.npy'\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:28:11.648131Z","iopub.execute_input":"2024-04-15T02:28:11.649216Z","iopub.status.idle":"2024-04-15T02:28:11.669154Z","shell.execute_reply.started":"2024-04-15T02:28:11.649182Z","shell.execute_reply":"2024-04-15T02:28:11.668166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Function for Augmentation, to generate Dataset and Preprocess the Images\n * Decoding an image which involves converting the image from its raw format into a format that can be processed by the machine learning model. ","metadata":{}},{"cell_type":"code","source":"def build_augmenter(dim=[400,300]):\n    augmenters =[\n        keras_cv.layers.RandomCutout(height_factor=(0.5, 0.8), width_factor=(0.2, 0.4)), \n        keras_cv.layers.RandomCutout(height_factor=(0.2, 0.4), width_factor=(0.2, 0.4)), \n        keras_cv.layers.RandomFlip(),\n        keras_cv.layers.MixUp(alpha=2.0)\n    ]\n\n    def augment(img, labels):\n        inputs = {\"images\": img, \"labels\": labels}\n        data =inputs\n        for aug in augmenters:\n            if tf.random.uniform([]) < 0.5:\n                data = aug(data, training=True)\n        return data[\"images\"], data[\"labels\"]\n    \n    return augment\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:28:11.670443Z","iopub.execute_input":"2024-04-15T02:28:11.670829Z","iopub.status.idle":"2024-04-15T02:28:12.372878Z","shell.execute_reply.started":"2024-04-15T02:28:11.670802Z","shell.execute_reply":"2024-04-15T02:28:12.371792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def build_decoder(with_labels=True, target_size=[400,300], dtype=32):\n    def decode_signal(path, offset=None):\n        # Read .npy files and process the signal\n        file_bytes = tf.io.read_file(path)\n        sig = tf.io.decode_raw(file_bytes, tf.float32)\n        sig = sig[1024//dtype:]  # Remove header tag\n        sig = tf.reshape(sig, [400, -1])\n        \n        # Extract labeled subsample from full spectrogram using \"offset\"\n        if offset is not None: \n            offset = offset // 2  # Only odd values are given\n            sig = sig[:, offset:offset+300]\n            \n            # Pad spectrogram to ensure the same input shape of [400, 300]\n            pad_size = tf.math.maximum(0, 300 - tf.shape(sig)[1])\n            sig = tf.pad(sig, [[0, 0], [0, pad_size]])\n            sig = tf.reshape(sig, [400, 300])\n        \n        # Log spectrogram \n        sig = tf.clip_by_value(sig, tf.math.exp(-4.0), tf.math.exp(8.0)) # avoid 0 in log\n        sig = tf.math.log(sig)\n        \n        # Normalize spectrogram\n        sig -= tf.math.reduce_mean(sig)\n        sig /= tf.math.reduce_std(sig) + 1e-6\n        \n        # Mono channel to 3 channels to use \"ImageNet\" weights\n        sig = tf.tile(sig[..., None], [1, 1, 3])\n        return sig\n    \n    def decode_label(label):\n        label = tf.one_hot(label, 6)\n        label = tf.cast(label, tf.float32)\n        label = tf.reshape(label, [6])\n        return label\n    \n    def decode_with_labels(path, offset=None, label=None):\n        sig = decode_signal(path, offset)\n        label = decode_label(label)\n        return (sig, label)\n    \n    return decode_with_labels if with_labels else decode_signal\n\n\ndef create_dataset(paths, offsets=None, labels=None, batch_size=32, cache=True,\n                  decode_fn=None, augment_fn=None,\n                  augment=False, repeat=True, shuffle=1024, \n                  cache_dir=\"\", drop_remainder=False):\n    if cache_dir != \"\" and cache is True:\n        os.makedirs(cache_dir, exist_ok=True)\n    \n    if decode_fn is None:\n        decode_fn = build_decoder(labels is not None)\n    \n    if augment_fn is None:\n        augment_fn = build_augmenter()\n    \n    AUTO = tf.data.experimental.AUTOTUNE\n    slices = (paths, offsets) if labels is None else (paths, offsets, labels)\n    \n    ds = tf.data.Dataset.from_tensor_slices(slices)\n    ds = ds.map(decode_fn, num_parallel_calls=AUTO)\n    ds = ds.cache(cache_dir) if cache else ds\n    ds = ds.repeat() if repeat else ds\n    if shuffle: \n        ds = ds.shuffle(shuffle, seed=42)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n    ds = ds.batch(batch_size, drop_remainder=drop_remainder)\n    ds = ds.map(augment_fn, num_parallel_calls=AUTO) if augment else ds\n    ds = ds.prefetch(AUTO)\n    return ds","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:28:12.374062Z","iopub.execute_input":"2024-04-15T02:28:12.374423Z","iopub.status.idle":"2024-04-15T02:28:12.393679Z","shell.execute_reply.started":"2024-04-15T02:28:12.374395Z","shell.execute_reply":"2024-04-15T02:28:12.392763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Handle Imbalanced Data\n\nUsing **\"StratifiedGroupKFold\"** from **\"sklearn.model_selection\"** for dealing with imbalanced datasets because it prevents one class from being overly represented or underrepresented in any given fold.\n","metadata":{}},{"cell_type":"code","source":"sgkf = StratifiedGroupKFold(n_splits = 5, shuffle = True, random_state = 42)\n# Assign the fold the data\ntrain_df[\"fold\"] =-1\ntrain_df.reset_index(drop=True, inplace=True)\nfor fold, (_, valid_idx) in enumerate(sgkf.split(train_df, y=train_df['target_labels'], groups = train_df['patient_id'])):\n    train_df.loc[valid_idx, \"fold\"] = fold\n    \n#to display  the all columns of fold\npd.set_option('display.max_columns', None)\n\n# fold group by the eed_id and xpert consensus\ncounts = train_df.groupby([\"fold\", \"expert_consensus\"])[[\"eeg_id\"]].count().T\n\n# display the ocunt\ncounts","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:28:12.394975Z","iopub.execute_input":"2024-04-15T02:28:12.395323Z","iopub.status.idle":"2024-04-15T02:28:13.379589Z","shell.execute_reply.started":"2024-04-15T02:28:12.395299Z","shell.execute_reply":"2024-04-15T02:28:13.378523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_df = train_df.groupby('spectrogram_id').head(1).reset_index(drop=True)\ntrain = sample_df[sample_df.fold != 0]\nvalid = sample_df[sample_df.fold == 0]\nprint(f\"# Num Train: {len(train)} | Num Valid: {len(valid)}\")","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:28:13.380894Z","iopub.execute_input":"2024-04-15T02:28:13.381202Z","iopub.status.idle":"2024-04-15T02:28:13.398368Z","shell.execute_reply.started":"2024-04-15T02:28:13.381177Z","shell.execute_reply":"2024-04-15T02:28:13.397220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_paths = train[\"spec_path_numpy\"].values\ntrain_offsets = train[\"spectrogram_label_offset_seconds\"].astype(int)\ntrain_labels = train[\"target_labels\"].values\ntrain_ds = create_dataset(train_paths, offsets=train_offsets, labels=train_labels, batch_size=32, cache=True,\n                  augment=True,repeat=True, shuffle=True)\n# Valid\nvalid_paths = valid[\"spec_path_numpy\"].values\nvalid_offsets = valid[\"spectrogram_label_offset_seconds\"].values.astype(int)\nvalid_labels = valid[\"target_labels\"].values\nvalid_ds = create_dataset(valid_paths, offsets=valid_offsets, labels=valid_labels, batch_size=32, cache=True,\n                   augment=False,repeat=False, shuffle=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:28:13.399686Z","iopub.execute_input":"2024-04-15T02:28:13.400050Z","iopub.status.idle":"2024-04-15T02:28:16.249861Z","shell.execute_reply.started":"2024-04-15T02:28:13.400018Z","shell.execute_reply":"2024-04-15T02:28:16.248970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import math\n\ndef get_lr_callback(batch_size=8, mode='cos', epochs=10, plot=False):\n    lr_start, lr_max, lr_min = 5e-5, 6e-6 * batch_size, 1e-5\n    lr_ramp_ep, lr_sus_ep, lr_decay = 3, 0, 0.75\n\n    def lrfn(epoch):  # Learning rate update function\n        if epoch < lr_ramp_ep: lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n        elif epoch < lr_ramp_ep + lr_sus_ep: lr = lr_max\n        elif mode == 'exp': lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n        elif mode == 'step': lr = lr_max * lr_decay**((epoch - lr_ramp_ep - lr_sus_ep) // 2)\n        elif mode == 'cos':\n            decay_total_epochs, decay_epoch_index = epochs - lr_ramp_ep - lr_sus_ep + 3, epoch - lr_ramp_ep - lr_sus_ep\n            phase = math.pi * decay_epoch_index / decay_total_epochs\n            lr = (lr_max - lr_min) * 0.5 * (1 + math.cos(phase)) + lr_min\n        return lr\n\n    if plot:  # Plot lr curve if plot is True\n        plt.figure(figsize=(10, 5))\n        plt.plot(np.arange(epochs), [lrfn(epoch) for epoch in np.arange(epochs)], marker='o')\n        plt.xlabel('epoch'); plt.ylabel('lr')\n        plt.title('LR Scheduler')\n        plt.show()\n\n    return keras.callbacks.LearningRateScheduler(lrfn, verbose=False)\n\n## Model checkpoint\ncallback = keras.callbacks.ModelCheckpoint('best_model.keras.weights.h5',\n                                             monitor ='val_loss',\n                                             verbose=0,\n                                             save_best_only=True,\n                                             save_weights_only=True,\n                                             mode=\"min\")","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:34:51.263248Z","iopub.execute_input":"2024-04-15T02:34:51.263938Z","iopub.status.idle":"2024-04-15T02:34:51.275205Z","shell.execute_reply.started":"2024-04-15T02:34:51.263906Z","shell.execute_reply":"2024-04-15T02:34:51.274072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr_scheduler = get_lr_callback(64, mode='cos', plot=True)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:34:56.742569Z","iopub.execute_input":"2024-04-15T02:34:56.743255Z","iopub.status.idle":"2024-04-15T02:34:57.039212Z","shell.execute_reply.started":"2024-04-15T02:34:56.743211Z","shell.execute_reply":"2024-04-15T02:34:57.038271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Models Used:\n1. EfficientNetV2-B2\n2. ResNet50\n3. EfficientNetV2-B0\n","metadata":{}},{"cell_type":"markdown","source":"# Model 1\n## efficientnetv2_b2","metadata":{}},{"cell_type":"code","source":"# Build Classifier \nwith strategy.scope():\n    model1 = keras_cv.models.ImageClassifier.from_preset(\n       \"efficientnetv2_b2_imagenet\", num_classes=6\n    )    \n\n    loss = tf.keras.losses.KLDivergence()\n\n    # Compile the model within the scope of the strategy\n\n    model1.compile(optimizer=keras.optimizers.Adam(learning_rate=1e-3),\n                      loss=loss, metrics=['accuracy'])\n\n#display the summary of model\nmodel1.summary()","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:34:59.525839Z","iopub.execute_input":"2024-04-15T02:34:59.526590Z","iopub.status.idle":"2024-04-15T02:35:22.009284Z","shell.execute_reply.started":"2024-04-15T02:34:59.526557Z","shell.execute_reply":"2024-04-15T02:35:22.008305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.random.set_seed(1)\n# train the model\n\nh1 = model1.fit(\n    train_ds, \n    batch_size = 64,\n    epochs=15,\n    validation_data = valid_ds,\n    callbacks = [lr_scheduler,callback],\n    steps_per_epoch=len(train)//64,\n      verbose=1\n)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:35:22.011010Z","iopub.execute_input":"2024-04-15T02:35:22.011327Z","iopub.status.idle":"2024-04-15T02:52:26.615495Z","shell.execute_reply.started":"2024-04-15T02:35:22.011300Z","shell.execute_reply":"2024-04-15T02:52:26.614684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n\n# Plot training & validation loss values\nplt.plot(h1.history['loss'])\nplt.plot(h1.history['val_loss'])\nplt.title('Model1 loss')\nplt.ylabel('Loss')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n\n# Plot training & validation accuracy values\nplt.plot(h1.history['accuracy'])\nplt.plot(h1.history['val_accuracy'])\nplt.title('Model1 accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:43:56.676415Z","iopub.execute_input":"2024-04-15T03:43:56.676737Z","iopub.status.idle":"2024-04-15T03:43:57.230733Z","shell.execute_reply.started":"2024-04-15T03:43:56.676711Z","shell.execute_reply":"2024-04-15T03:43:57.229728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"efficient_pred = model1.predict(valid_ds)\npredicted_prob_efficientnet = np.argmax(efficient_pred, axis=1)\n\ntarget_labels = valid_labels","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:52:58.883511Z","iopub.execute_input":"2024-04-15T02:52:58.883897Z","iopub.status.idle":"2024-04-15T02:53:32.054808Z","shell.execute_reply.started":"2024-04-15T02:52:58.883866Z","shell.execute_reply":"2024-04-15T02:53:32.053788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"efficient_pred","metadata":{"execution":{"iopub.status.busy":"2024-04-15T02:53:32.056849Z","iopub.execute_input":"2024-04-15T02:53:32.057722Z","iopub.status.idle":"2024-04-15T02:53:32.065194Z","shell.execute_reply.started":"2024-04-15T02:53:32.057682Z","shell.execute_reply":"2024-04-15T02:53:32.064170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model-2\n## ResNet50","metadata":{}},{"cell_type":"code","source":"# Import ResNet50 from TensorFlow\nfrom tensorflow.keras.applications import ResNet50\nfrom tensorflow.keras.layers import Input, GlobalAveragePooling2D, Dense, Dropout\nfrom tensorflow.keras.regularizers import l2\n\n\nweight_path = \"/kaggle/input/resner50-weight/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5\"\n\n# Define a function to build ResNet50 model\nwith strategy.scope():\n    def build_resnet50_model(num_classes=6, input_shape=(400, 300, 3)):\n        # Load pre-trained ResNet50 model with imagenet weights\n        base_model = ResNet50(weights=None , include_top=False, input_shape=input_shape)\n        base_model.load_weights(weight_path)\n        # Freeze the base model layers\n        base_model.trainable = True\n        \n        inputs = Input(shape=input_shape)\n        x = base_model(inputs, training=False)\n        \n        # Add classification head\n        x = keras.layers.GlobalAveragePooling2D()(x)\n        # Adding Dropout\n        x = Dropout(0.5)(x)\n        output = keras.layers.Dense(6, activation='softmax',kernel_regularizer=l2(0.01))(x)\n\n        # Create the model\n        model2 = keras.Model(inputs=inputs, outputs=output)\n\n        return model2\n\n    # Build ResNet50 model\n    resnet50_model = build_resnet50_model()\n\n    # Compile the model\n    resnet50_model.compile(optimizer=keras.optimizers.Adam(learning_rate=1e-5),\n                            loss=keras.losses.KLDivergence(),\n                            metrics=['accuracy'])\n\n# Summary of the ResNet50 model\nresnet50_model.summary()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:15:32.114942Z","iopub.execute_input":"2024-04-15T03:15:32.115940Z","iopub.status.idle":"2024-04-15T03:15:33.844623Z","shell.execute_reply.started":"2024-04-15T03:15:32.115900Z","shell.execute_reply":"2024-04-15T03:15:33.843713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nwith strategy.scope():\n# Train the ResNet50 model\n    h2 = resnet50_model.fit(train_ds,\n                             batch_size=64,\n                             epochs=15,\n                             validation_data=valid_ds,\n                             callbacks=[lr_scheduler,callback],\n                             steps_per_epoch=len(train)//64,\n                             verbose=1)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:15:39.779352Z","iopub.execute_input":"2024-04-15T03:15:39.779732Z","iopub.status.idle":"2024-04-15T03:43:56.674444Z","shell.execute_reply.started":"2024-04-15T03:15:39.779703Z","shell.execute_reply":"2024-04-15T03:43:56.673354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import matplotlib.pyplot as plt\n\n# Plot training & validation loss values\nplt.plot(h2.history['loss'])\nplt.plot(h2.history['val_loss'])\nplt.title('Model2 loss')\nplt.ylabel('Loss')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n\n# Plot training & validation accuracy values\nplt.plot(h2.history['accuracy'])\nplt.plot(h2.history['val_accuracy'])\nplt.title('Model2 accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:43:57.232107Z","iopub.execute_input":"2024-04-15T03:43:57.232506Z","iopub.status.idle":"2024-04-15T03:43:57.824226Z","shell.execute_reply.started":"2024-04-15T03:43:57.232470Z","shell.execute_reply":"2024-04-15T03:43:57.823221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model-3\n## efficientnetv2_b0","metadata":{}},{"cell_type":"code","source":"# Build Classifier \nwith strategy.scope():\n    model3 = keras_cv.models.ImageClassifier.from_preset(\n       \"efficientnetv2_b0_imagenet\", num_classes=6\n    )    \n\n    loss = tf.keras.losses.KLDivergence()\n\n    # Compile the model within the scope of the strategy\n\n    model3.compile(optimizer=keras.optimizers.Adam(learning_rate=1e-4),\n                      loss=loss, metrics=['accuracy'])\n\n#display the summary of model\nmodel3.summary()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:43:57.826587Z","iopub.execute_input":"2024-04-15T03:43:57.826937Z","iopub.status.idle":"2024-04-15T03:44:03.246937Z","shell.execute_reply.started":"2024-04-15T03:43:57.826909Z","shell.execute_reply":"2024-04-15T03:44:03.245900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n    h3 = model3.fit(\n        train_ds,\n        epochs=15,\n        batch_size=64,\n        validation_data=valid_ds,\n        callbacks = [lr_scheduler,callback],\n        steps_per_epoch=len(train)//64,\n        verbose=1)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:44:03.248199Z","iopub.execute_input":"2024-04-15T03:44:03.248532Z","iopub.status.idle":"2024-04-15T03:55:38.934830Z","shell.execute_reply.started":"2024-04-15T03:44:03.248504Z","shell.execute_reply":"2024-04-15T03:55:38.933807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import matplotlib.pyplot as plt\n\n# Plot training & validation loss values\nplt.plot(h3.history['loss'])\nplt.plot(h3.history['val_loss'])\nplt.title('Model3 loss')\nplt.ylabel('Loss')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n\n# Plot training & validation accuracy values\nplt.plot(h3.history['accuracy'])\nplt.plot(h3.history['val_accuracy'])\nplt.title('Model3 accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:55:38.936320Z","iopub.execute_input":"2024-04-15T03:55:38.936606Z","iopub.status.idle":"2024-04-15T03:55:39.549321Z","shell.execute_reply.started":"2024-04-15T03:55:38.936582Z","shell.execute_reply":"2024-04-15T03:55:39.548512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predict with all three models\n\npredictions_efficientnet = model1.predict(valid_ds)\npredictions_resnet50 = resnet50_model.predict(valid_ds)\npredictions_model3= model3.predict(valid_ds)\n  \n","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:55:39.550499Z","iopub.execute_input":"2024-04-15T03:55:39.550871Z","iopub.status.idle":"2024-04-15T03:56:27.963795Z","shell.execute_reply.started":"2024-04-15T03:55:39.550843Z","shell.execute_reply":"2024-04-15T03:56:27.962972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ensemble predictions by averaging\nensemble_predictions = ( predictions_efficientnet + predictions_resnet50 + predictions_model3  ) / 3\nfinal_predictions = np.argmax(ensemble_predictions, axis=1)\nfrom sklearn.metrics import accuracy_score\nprint(f'Ensemble model accuracy on validation set: {accuracy_score(valid_labels, final_predictions)}')","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:27.965044Z","iopub.execute_input":"2024-04-15T03:56:27.965363Z","iopub.status.idle":"2024-04-15T03:56:27.972819Z","shell.execute_reply.started":"2024-04-15T03:56:27.965337Z","shell.execute_reply":"2024-04-15T03:56:27.971828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Ensemble Models Classification Report","metadata":{}},{"cell_type":"code","source":"names = ['Seizure', 'LPD', 'GPD', 'LRDA','GRDA', 'Other']\nprint(\"Classification Report:\")\nprint(classification_report(target_labels,final_predictions , target_names=names))\n\n# Calculate confusion matrix\nprint(\"Confusion Matrix:\")\nprint(confusion_matrix(target_labels,final_predictions))\n\n# Calculate precision\nprecision = precision_score(target_labels, final_predictions, average='macro')\nprint(\"\\nPrecision:\", round(precision,4))\n\n# Calculate average precision across all folds\naverage_precision = np.mean(precision)\nprint(\"Average Precision:\", average_precision)\n\n# Calculate recall\nrecall = recall_score(target_labels,final_predictions, average='macro')\nprint(\"\\nRecall:\", round(recall,4))\n\n\n# Calculate average precision across all folds\naverage_recall = np.mean(recall)\nprint(\"Average Precision:\", average_recall)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:27.974222Z","iopub.execute_input":"2024-04-15T03:56:27.974682Z","iopub.status.idle":"2024-04-15T03:56:27.998612Z","shell.execute_reply.started":"2024-04-15T03:56:27.974654Z","shell.execute_reply":"2024-04-15T03:56:27.997630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# # Access the loss history from the history object of each model\nloss_history_model1 = h1.history['loss']\nval_loss_history_model1 = h1.history['val_loss']\n\nloss_history_model2 = h2.history['loss']\nval_loss_history_model2 = h2.history['val_loss']\n\nloss_history_model3 = h3.history['loss']\nval_loss_history_model3 = h3.history['val_loss']\n\n# Plot the loss history\nplt.figure(figsize=(10, 5))\n\n# Plot model 1 loss\nplt.plot(loss_history_model1, label='Model 1 Training Loss', color='blue')\nplt.plot(val_loss_history_model1, label='Model 1 Validation Loss', color='skyblue')\n\n# Plot model 2 loss\nplt.plot(loss_history_model2, label='Model 2 Training Loss', color='red')\nplt.plot(val_loss_history_model2, label='Model 2 Validation Loss', color='salmon')\n\n# Plot model 3 loss\nplt.plot(loss_history_model3, label='Model 3 Training Loss', color='orange')\nplt.plot(val_loss_history_model3, label='Model 3 Validation Loss', color='yellow')\n\n\nplt.title('Training and Validation Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.002088Z","iopub.execute_input":"2024-04-15T03:56:28.002397Z","iopub.status.idle":"2024-04-15T03:56:28.373893Z","shell.execute_reply.started":"2024-04-15T03:56:28.002372Z","shell.execute_reply":"2024-04-15T03:56:28.372924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# # Access the loss history from the history object of each model\nloss_history_model1 = h1.history['accuracy']\nval_loss_history_model1 = h1.history['val_accuracy']\n\nloss_history_model2 = h2.history['accuracy']\nval_loss_history_model2 = h2.history['val_accuracy']\n\nloss_history_model3 = h3.history['accuracy']\nval_loss_history_model3 = h3.history['val_accuracy']\n\n# Plot the loss history\nplt.figure(figsize=(10, 5))\n\n# Plot model 1 loss\nplt.plot(loss_history_model1, label='Model 1 Training Accuracy', color='blue')\nplt.plot(val_loss_history_model1, label='Model 1 Validation Accuracy', color='skyblue')\n\n# Plot model 2 loss\nplt.plot(loss_history_model2, label='Model 2 Training Accuracy', color='red')\nplt.plot(val_loss_history_model2, label='Model 2 Validation Accuracy', color='salmon')\n\n# Plot model 3 loss\nplt.plot(loss_history_model3, label='Model 3 Training Accuracy', color='orange')\nplt.plot(val_loss_history_model3, label='Model 3 Validation Accuracy', color='yellow')\n\n\nplt.title('Training and Validation Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.375356Z","iopub.execute_input":"2024-04-15T03:56:28.375746Z","iopub.status.idle":"2024-04-15T03:56:28.707531Z","shell.execute_reply.started":"2024-04-15T03:56:28.375712Z","shell.execute_reply":"2024-04-15T03:56:28.706525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test Data","metadata":{}},{"cell_type":"code","source":"# Read the test dataset\ntest_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/test.csv\")\ntest_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.709018Z","iopub.execute_input":"2024-04-15T03:56:28.709347Z","iopub.status.idle":"2024-04-15T03:56:28.732748Z","shell.execute_reply.started":"2024-04-15T03:56:28.709321Z","shell.execute_reply":"2024-04-15T03:56:28.731844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Define path of all test files\ntest_spec = \"/kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms\"\ntest_eeg = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs\"","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.733958Z","iopub.execute_input":"2024-04-15T03:56:28.734362Z","iopub.status.idle":"2024-04-15T03:56:28.739107Z","shell.execute_reply.started":"2024-04-15T03:56:28.734330Z","shell.execute_reply":"2024-04-15T03:56:28.738123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Get Unique spectrogram ids of test data\n ### Test data has only one spectrogram id\nspec_ids_test = test_df['spectrogram_id'].unique()\n\n# Apply process function on test data\nfor spec_id in spec_ids_test:\n    process_spectrogram(spec_id, base_path, split = 'test')","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.740429Z","iopub.execute_input":"2024-04-15T03:56:28.740800Z","iopub.status.idle":"2024-04-15T03:56:28.804807Z","shell.execute_reply.started":"2024-04-15T03:56:28.740774Z","shell.execute_reply":"2024-04-15T03:56:28.804068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df['eeg_path'] = test_eeg+'/' +  test_df['eeg_id'].astype(str) + '.parquet'\ntest_df['spec_path'] = test_spec+ '/' + test_df['spectrogram_id'].astype(str) + '.parquet'\n\nnumpy_test_spec= \"/tmp/dataset/hms-hbac/test_spectrograms\"\n# #Replace the .parquet file to the numpy.\ntest_df['spec_path_numpy'] = test_df['spectrogram_id'].apply(lambda x: os.path.join(numpy_test_spec , f\"{x}.npy\"))\n\n#Display the test data \n# test_df","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.806041Z","iopub.execute_input":"2024-04-15T03:56:28.806459Z","iopub.status.idle":"2024-04-15T03:56:28.815606Z","shell.execute_reply.started":"2024-04-15T03:56:28.806427Z","shell.execute_reply":"2024-04-15T03:56:28.814490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Applt the preprocess functions to the test data\ntest_paths = test_df['spec_path_numpy'].values\ntest_ds = create_dataset(test_paths, batch_size=min(64, len(test_df)),\n                         repeat=False, shuffle=False, cache=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.816633Z","iopub.execute_input":"2024-04-15T03:56:28.816897Z","iopub.status.idle":"2024-04-15T03:56:28.868446Z","shell.execute_reply.started":"2024-04-15T03:56:28.816874Z","shell.execute_reply":"2024-04-15T03:56:28.867711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Predict all three models on the test set\npredictions_resnet50_test = resnet50_model.predict(test_ds)\npredictions_model3_test = model3.predict(test_ds)\npredictions_efficientnet_test = model1.predict(test_ds)","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:56:28.869763Z","iopub.execute_input":"2024-04-15T03:56:28.870137Z","iopub.status.idle":"2024-04-15T03:57:05.344894Z","shell.execute_reply.started":"2024-04-15T03:56:28.870103Z","shell.execute_reply":"2024-04-15T03:57:05.343962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Ensemble predictions by averaging\nensemble_predictions_test = (predictions_efficientnet_test + predictions_resnet50_test + predictions_model3_test) / 3\nfinal_predictions_test = np.argmax(ensemble_predictions_test, axis=1)\nensemble_predictions_test","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:57:05.346408Z","iopub.execute_input":"2024-04-15T03:57:05.346709Z","iopub.status.idle":"2024-04-15T03:57:05.354388Z","shell.execute_reply.started":"2024-04-15T03:57:05.346683Z","shell.execute_reply":"2024-04-15T03:57:05.353314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_predictions_test.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:57:05.355708Z","iopub.execute_input":"2024-04-15T03:57:05.356042Z","iopub.status.idle":"2024-04-15T03:57:05.368517Z","shell.execute_reply.started":"2024-04-15T03:57:05.356016Z","shell.execute_reply":"2024-04-15T03:57:05.367552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Prepare the predicted porbabilties for the submission and convert them into the csv file\npred_df = test_df[[\"eeg_id\"]]\npred_df[target_classes] = ensemble_predictions_test.tolist()\npred_df[target_classes] \n\nsubmission= pd.read_csv(f'{base_path}/sample_submission.csv')\ndf =submission[[\"eeg_id\"]]\ndf = df.merge(pred_df, on=\"eeg_id\")\ndf.to_csv(\"submission.csv\", index=False)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-15T03:57:05.369815Z","iopub.execute_input":"2024-04-15T03:57:05.370467Z","iopub.status.idle":"2024-04-15T03:57:05.411768Z","shell.execute_reply.started":"2024-04-15T03:57:05.370441Z","shell.execute_reply":"2024-04-15T03:57:05.410967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# References\n\n\n* [HMS-HBAC KerasCV Starter Notebook](https://www.kaggle.com/code/awsaf49/hms-hbac-kerascv-starter-notebook)\n* [HMS- dataset Analysis Xgboost Classification](https://www.kaggle.com/code/andradaolteanu/hms-dataset-analysis-xgboost-classification)\n* [Keras_CV](https://keras.io/keras_cv/)","metadata":{}}]}