{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30626,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Examining the Training File\n\nThis notebook aims to given an introductory to the metadata in the **train.csv** file for this competition. I will focus the  the analysis on the patients, brain activity classes, the EEG samples, and the spectogram samples. Later on, I will focus on the features within the EEG and spectogram files. \n\n* [Custom Functions](#custom_functions)\n* [Read Data](#read_file)\n* [EDA](#eda)\n    * [Patient Analysis](#patient_analysis)\n    * [Brain Activity Classes](#classes)\n    \n### Summary \n\n**Patient Info**:\n* There are 1950 patients in the training dataset.\n* Most patients have *less than 5 EEG samples* in the training dataset with 2 EEG samples being the most common\n* A few patients have *more than 100* different EEG samples in the training set\n* Similarly, most patients have *less than 5 spectrogram samples* in the training dataset with *2 spectrogram sample being the most common*\n* A few patients have *more than 50* different spectrogram samples in the training set\n* Based on expert consensus, most patients only have *1 or 2 unique brain activity classes* in the training sample\n\n\n**Brain Activity Classes**:\n* Most of the consenus is evenly split among the classes\n* Almost half of the records only have 1 class with a vote\n    \nMore to come...","metadata":{}},{"cell_type":"code","source":"# Analysis packages\nimport pandas as pd \n\n# Visualization packages\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-01-12T19:02:17.063435Z","iopub.execute_input":"2024-01-12T19:02:17.063847Z","iopub.status.idle":"2024-01-12T19:02:18.154249Z","shell.execute_reply.started":"2024-01-12T19:02:17.063813Z","shell.execute_reply":"2024-01-12T19:02:18.153407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom Functions\n<a id=\"custom_functions\" ></a>\n","metadata":{}},{"cell_type":"code","source":"primary_color = '#FF595E'\nsecondary_color = '#1982C4'\nthird_color = '#6A4C93'\nsns.set_style(\n    \"whitegrid\",\n    {\n        \"axes.facecolor\": \"#FFFDD0\",\n        \"figure.facecolor\": \"#FFFDD0\",\n        \"patch.facecolor\":primary_color,\n        \"patch.edgecolor\":'black',\n        \"axes.edgecolor\":'black',\n        \"grid.color\":'black'\n    })","metadata":{"execution":{"iopub.status.busy":"2024-01-12T02:05:28.524114Z","iopub.execute_input":"2024-01-12T02:05:28.524496Z","iopub.status.idle":"2024-01-12T02:05:28.530831Z","shell.execute_reply.started":"2024-01-12T02:05:28.524466Z","shell.execute_reply":"2024-01-12T02:05:28.529473Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def custom_barplot(data, x, y, xlabel, ylabel, title, ax, vertical = True):\n    ax = sns.barplot(\n        data = data,\n        x = x,\n        y = y,\n        color=primary_color,\n        ax=ax\n    )\n    \n    # Add percentage labels above the bars\n    bars = ax.containers[0]\n    if vertical:\n        total = sum(data[y])\n        ax.bar_label(\n            bars, \n            labels = [f'{(x.get_height()/total):.1%}' for x in bars]\n        )\n    else:\n        total = sum(data[x])\n        ax.bar_label(\n            bars, \n            labels = [f'{(x.get_width()/total):.1%}' for x in bars]\n        )\n        \n\n    ax.set_xlabel(xlabel)\n    ax.set_ylabel(ylabel)\n    ax.set_title(title)\n        \n    return ax\n\ndef custom_boxenplot(data, x, xlabel, title, ax):\n    ax = sns.boxenplot(\n        data = data,\n        x = x,\n        ax=ax,\n        color=primary_color\n    )\n        \n\n    ax.set_xlabel(xlabel)\n    ax.set_title(title)\n        \n    return ax","metadata":{"execution":{"iopub.status.busy":"2024-01-12T02:05:28.536494Z","iopub.execute_input":"2024-01-12T02:05:28.537429Z","iopub.status.idle":"2024-01-12T02:05:28.548880Z","shell.execute_reply.started":"2024-01-12T02:05:28.537388Z","shell.execute_reply":"2024-01-12T02:05:28.547936Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read Data\n<a id=\"read_file\" ></a>","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2024-01-12T02:05:28.561351Z","iopub.execute_input":"2024-01-12T02:05:28.562419Z","iopub.status.idle":"2024-01-12T02:05:28.741132Z","shell.execute_reply.started":"2024-01-12T02:05:28.562381Z","shell.execute_reply":"2024-01-12T02:05:28.739625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA\n<a id=\"eda\" ></a>","metadata":{}},{"cell_type":"markdown","source":"## Patient Analysis\n<a id=\"patient_analysis\" ></a>\n\nLet's look at the patients in the dataset and some characteristics.","metadata":{}},{"cell_type":"code","source":"print(f'There are {train_df.patient_id.nunique()} patients in the training dataset.')","metadata":{"execution":{"iopub.status.busy":"2024-01-12T02:05:28.744030Z","iopub.execute_input":"2024-01-12T02:05:28.744773Z","iopub.status.idle":"2024-01-12T02:05:28.752694Z","shell.execute_reply.started":"2024-01-12T02:05:28.744720Z","shell.execute_reply":"2024-01-12T02:05:28.751094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(ncols=2, figsize=(16,6))\n\n# Left plot\nax = custom_boxenplot(\n    data= (train_df\n         .groupby('patient_id')\n         .eeg_id\n         .nunique()\n         .reset_index()\n         .set_axis(['patient_id', 'num_eegs'], axis=1)\n        ),\n    x = 'num_eegs',\n    ax=axs[0],\n    xlabel = 'Number of EEGs per patient',\n    title = 'There are a few patients with more than 150 EEG samples'\n)\nax.grid(axis='x', color='black', linestyle='--', alpha=0.8)\nax.set_axisbelow(True)\n\n# Right plot\nax = custom_barplot(\n    data = (train_df\n             .groupby('patient_id')\n             .eeg_id\n             .nunique()\n             .reset_index()\n             .set_axis(['patient_id', 'num_eegs'], axis=1)\n             .assign(\n                 label = lambda x: x.num_eegs.apply(lambda x: '> 10' if x > 10 else x)\n             )\n            .groupby('label')\n            .num_eegs\n            .count()\n            .reset_index()\n        ),\n    x = 'label',\n    y = 'num_eegs',\n    xlabel = 'Number of EEGs per patient',\n    ylabel = 'Number of patients',\n    title = 'Most patients have less than 5 EEG samples in the training set',\n    ax = axs[1]\n)\nax.grid(axis='y', color='black', linestyle='--', alpha=0.8)\nax.set_axisbelow(True)\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-12T02:05:28.754283Z","iopub.execute_input":"2024-01-12T02:05:28.755781Z","iopub.status.idle":"2024-01-12T02:05:29.465040Z","shell.execute_reply.started":"2024-01-12T02:05:28.755719Z","shell.execute_reply":"2024-01-12T02:05:29.463933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(ncols=2, figsize=(16,6))\n\n# Left plot\nax = custom_boxenplot(\n    data= (train_df\n     .groupby('patient_id')\n     .spectrogram_id\n     .nunique()\n     .reset_index()\n     .set_axis(['patient_id', 'num_spectrograms'], axis=1)\n    ),\n    x='num_spectrograms',\n    xlabel = 'Number of spectrograms per patient',\n    title = 'There are a few patients with more than 50 spectrogram samples',\n    ax=axs[0]\n)\nax.grid(axis='x', color='black', linestyle='--', alpha=0.8)\nax.set_axisbelow(True)\n\n# Right plot\nax = custom_barplot(\n    data = (train_df\n             .groupby('patient_id')\n             .spectrogram_id\n             .nunique()\n             .reset_index()\n             .set_axis(['patient_id', 'num_spectrograms'], axis=1)\n             .assign(\n                 label = lambda x: x.num_spectrograms.apply(lambda x: '> 10' if x > 10 else x)\n             )\n            .groupby('label')\n            .num_spectrograms\n            .count()\n            .reset_index()\n        ),\n    x = 'label',\n    y = 'num_spectrograms',\n    xlabel = 'Number of spectrograms per patient',\n    ylabel = 'Number of patients',\n    title = 'Most patients have less than 5 spectrogram samples in the training set',\n    ax = axs[1]\n)\nax.grid(axis='y', color='black', linestyle='--', alpha=0.8)\nax.set_axisbelow(True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-12T02:05:29.467080Z","iopub.execute_input":"2024-01-12T02:05:29.467497Z","iopub.status.idle":"2024-01-12T02:05:30.242834Z","shell.execute_reply.started":"2024-01-12T02:05:29.467467Z","shell.execute_reply":"2024-01-12T02:05:30.241763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Brain Activity for Patients\n","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nax = custom_barplot(\n    data = (train_df\n         .groupby('patient_id')\n         .expert_consensus\n         .nunique()\n         .reset_index()\n         .set_axis(['patient_id', 'unique_classes'], axis=1)\n         .groupby('unique_classes', as_index=False)\n         .patient_id\n         .count()\n        ),\n    x = 'unique_classes',\n    y = 'patient_id',\n    xlabel = 'Number of unique brain activity classes per patient',\n    ylabel = 'Number of patients',\n    title = 'Most patients only have 1 or 2 unique brain activity classes',\n    ax=ax\n)\nax.grid(axis='y', color='black', linestyle='--', alpha=0.8)\nax.set_axisbelow(True)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-12T02:05:30.244249Z","iopub.execute_input":"2024-01-12T02:05:30.244658Z","iopub.status.idle":"2024-01-12T02:05:30.593281Z","shell.execute_reply.started":"2024-01-12T02:05:30.244619Z","shell.execute_reply":"2024-01-12T02:05:30.591586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Brain Activity Classes\n<a id=\"classes\" ></a>","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\ncustom_barplot(\n    data=train_df.expert_consensus.value_counts().reset_index(),\n    x='count',\n    y='expert_consensus',\n    xlabel = 'Number of subsamples',\n    ylabel = 'Expert consensus',\n    title = 'Expert consensus is roughly evenly split among subsamples',\n    vertical=False,\n    ax=ax\n)\nax.set_xlim(0,23000)\nax.grid(axis='x', color='black', linestyle='--', alpha=0.8)\nax.set_axisbelow(True)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-12T02:05:30.595184Z","iopub.execute_input":"2024-01-12T02:05:30.595645Z","iopub.status.idle":"2024-01-12T02:05:31.004899Z","shell.execute_reply.started":"2024-01-12T02:05:30.595603Z","shell.execute_reply":"2024-01-12T02:05:31.003622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots()\n\ncustom_barplot(\n    data = (pd.concat(\n                [train_df['label_id'],\n                 train_df.filter(like='_vote')\n                ], axis=1)\n             .melt(id_vars='label_id')\n             .query(\"value > 0\")\n             .groupby('label_id')\n             .variable\n             .nunique()\n             .reset_index()\n             .groupby('variable', as_index=False)\n             .count()\n             .set_axis(['unique_votes', 'count'], axis=1)\n            ),\n    x='unique_votes',\n    y='count',\n    xlabel='Unique brain activity classes voted for',\n    ylabel='Subsamples',\n    title='Around 47.8% of the labels are fully-agreed upon by experts',\n    ax = ax\n)\nax.grid(axis='y', color='black', linestyle='--', alpha=0.8)\nax.set_axisbelow(True)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-12T02:06:06.652942Z","iopub.execute_input":"2024-01-12T02:06:06.653348Z","iopub.status.idle":"2024-01-12T02:06:07.658769Z","shell.execute_reply.started":"2024-01-12T02:06:06.653319Z","shell.execute_reply":"2024-01-12T02:06:07.657910Z"},"trusted":true},"execution_count":null,"outputs":[]}]}