{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Objectives :\n\nThis competition aims to improve the detection and classification of seizures and harmful brain activity using EEG signals from critically ill patients. Currently, doctors manually analyze EEG signals, which is time-consuming and prone to errors. The goal is to develop algorithms that can automate this process, helping doctors identify brain issues more quickly and accurately.\n\nThe competition focuses on six patterns of brain activity: \n- **seizure** \n- **generalized periodic discharges (GPD)**\n- **lateralized periodic discharges (LPD)**\n- **lateralized rhythmic delta activity (LRDA)**\n- **generalized rhythmic delta activity (GRDA)**\n\nExperts have annotated EEG segments with labels, but there's disagreement among them in some cases. \n- **\"Idealized\"** patterns have high agreement among experts\n- **\"proto patterns\"** ( one specific pattern and the other half labeled them as \"other\") and **\"edge cases\"** have varying levels of disagreement (about half of the raters labeled them as one specific pattern and the other half labeled them as another specific pattern.","metadata":{}},{"cell_type":"markdown","source":"### Import packages","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport os\nimport random\nimport gc\n\n# pytorch\nimport torch\nimport torch.nn as nn\n\n\n# Graph\nimport matplotlib.pyplot as plt \nimport seaborn as sns\nimport plotly.express as px\nimport plotly.graph_objs as go\nsns.set_theme()\n\n# warnings\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:47:39.221832Z","iopub.execute_input":"2024-02-12T10:47:39.222279Z","iopub.status.idle":"2024-02-12T10:47:39.231064Z","shell.execute_reply.started":"2024-02-12T10:47:39.222247Z","shell.execute_reply":"2024-02-12T10:47:39.229655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Cuda Device","metadata":{}},{"cell_type":"code","source":"# set up device \ndevice =\"cuda\" if torch.cuda.is_available() else \"cpu\"\nprint( device)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:14.297183Z","iopub.execute_input":"2024-02-12T10:35:14.298507Z","iopub.status.idle":"2024-02-12T10:35:14.304818Z","shell.execute_reply.started":"2024-02-12T10:35:14.298455Z","shell.execute_reply":"2024-02-12T10:35:14.303667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Configuration ","metadata":{}},{"cell_type":"code","source":"class CFG:\n    dataset_path=\"/kaggle/input/hms-harmful-brain-activity-classification\"\n    targets=['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\n    seed = 42\n\nCFG = CFG()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:14.306342Z","iopub.execute_input":"2024-02-12T10:35:14.307208Z","iopub.status.idle":"2024-02-12T10:35:14.324139Z","shell.execute_reply.started":"2024-02-12T10:35:14.307170Z","shell.execute_reply":"2024-02-12T10:35:14.322683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed(seed ):\n    '''Sets the seed of the entire notebook so results are the same every time we run.\n    This is for REPRODUCIBILITY.'''\n    np.random.seed(seed)\n    random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    # When running on the CuDNN backend, two further options must be set\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n    # Set a fixed value for the hash seed\n    os.environ['PYTHONHASHSEED'] = str(seed)\nset_seed(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:15.528546Z","iopub.execute_input":"2024-02-12T10:35:15.529180Z","iopub.status.idle":"2024-02-12T10:35:15.541505Z","shell.execute_reply.started":"2024-02-12T10:35:15.529134Z","shell.execute_reply":"2024-02-12T10:35:15.539809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Datasets Description : \n\n1. **train.csv**: This file contains metadata for the training set. It includes information about EEG recordings and spectrograms. Each row corresponds to a 50-second EEG sample, with labels provided by expert annotators. The columns include identifiers for EEG and spectrogram samples, offsets, label IDs, patient IDs, and expert consensus. Additionally, there are columns indicating the count of votes for each brain activity class (seizure, LPD, GPD, LRDA, GRDA, Other).\n\n2. **test.csv**: This file contains metadata for the test set. Unlike the training set, there are no overlapping samples in the test set. Columns include identifiers for EEG and spectrogram samples, as well as patient IDs.\n\n3. **sample_submission.csv**: This file is a template for submitting predictions for the test set. It contains columns for EEG identifiers and target columns for seizure, LPD, GPD, LRDA, GRDA, and Other brain activity classes. Predictions must be probabilities, and each test sample had between 3 and 20 annotators.\n\n4. **train_eegs/**: This directory contains EEG data from one or more overlapping samples for the training set. The metadata in train.csv helps select specific annotated subsets. Column names represent individual electrode locations for EEG leads, with one exception: the EKG column records data from the heart. The EEG data was collected at a frequency of 200 samples per second.\n\n5. **test_eegs/**: This directory contains exactly 50 seconds of EEG data for the test set.\n\n6. **train_spectrograms/**: This directory contains spectrograms assembled from EEG data for the training set. The metadata in train.csv helps select specific annotated subsets. Column names indicate frequency in hertz and recording regions of EEG electrodes (LL = left lateral; RL = right lateral; LP = left parasagittal; RP = right parasagittal).\n\n7. **test_spectrograms/**: This directory contains spectrograms assembled using exactly 10 minutes of EEG data for the test set.\n\n8. **example_figures/**: This directory contains larger copies of example case images used on the overview tab.\n\nOverall, the provided data and metadata facilitate the development and evaluation of algorithms for EEG pattern classification in the context of detecting seizures and harmful brain activity.\n","metadata":{}},{"cell_type":"markdown","source":"### Explore the Data","metadata":{}},{"cell_type":"code","source":"train=pd.read_csv(os.path.join(CFG.dataset_path,\"train.csv\"))\ntrain.head(20)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:16.535488Z","iopub.execute_input":"2024-02-12T10:35:16.535960Z","iopub.status.idle":"2024-02-12T10:35:16.935529Z","shell.execute_reply.started":"2024-02-12T10:35:16.535926Z","shell.execute_reply":"2024-02-12T10:35:16.934099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:16.938323Z","iopub.execute_input":"2024-02-12T10:35:16.938831Z","iopub.status.idle":"2024-02-12T10:35:16.948821Z","shell.execute_reply.started":"2024-02-12T10:35:16.938790Z","shell.execute_reply":"2024-02-12T10:35:16.947291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:16.950884Z","iopub.execute_input":"2024-02-12T10:35:16.951415Z","iopub.status.idle":"2024-02-12T10:35:16.998513Z","shell.execute_reply.started":"2024-02-12T10:35:16.951352Z","shell.execute_reply":"2024-02-12T10:35:16.997380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.describe()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:17.000260Z","iopub.execute_input":"2024-02-12T10:35:17.001763Z","iopub.status.idle":"2024-02-12T10:35:17.093129Z","shell.execute_reply.started":"2024-02-12T10:35:17.001715Z","shell.execute_reply":"2024-02-12T10:35:17.091992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Unique Indentifiers :\n\nConfirm the unique identifiers (eeg_id, eeg_sub_id, spectrogram_id, spectrogram_sub_id, label_id, patient_id) and their roles in identifying specific EEG and spectrogram samples, subsamples, labels, and patients.\n","metadata":{}},{"cell_type":"code","source":"train[\"spectrogram_id\"].nunique()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:17.095701Z","iopub.execute_input":"2024-02-12T10:35:17.097012Z","iopub.status.idle":"2024-02-12T10:35:17.108629Z","shell.execute_reply.started":"2024-02-12T10:35:17.096969Z","shell.execute_reply":"2024-02-12T10:35:17.107307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(os.listdir(CFG.dataset_path+\"/train_spectrograms\"))","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:17.369032Z","iopub.execute_input":"2024-02-12T10:35:17.369497Z","iopub.status.idle":"2024-02-12T10:35:22.080526Z","shell.execute_reply.started":"2024-02-12T10:35:17.369422Z","shell.execute_reply":"2024-02-12T10:35:22.079117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"spectrogram_id\"].nunique()-len(os.listdir(CFG.dataset_path+\"/train_spectrograms\"))","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:22.082835Z","iopub.execute_input":"2024-02-12T10:35:22.083263Z","iopub.status.idle":"2024-02-12T10:35:22.099386Z","shell.execute_reply.started":"2024-02-12T10:35:22.083230Z","shell.execute_reply":"2024-02-12T10:35:22.098038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, regarding the spectrogtam id there are not missing files","metadata":{}},{"cell_type":"code","source":"len(os.listdir(CFG.dataset_path+\"/train_eegs\"))","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:22.101090Z","iopub.execute_input":"2024-02-12T10:35:22.101509Z","iopub.status.idle":"2024-02-12T10:35:25.115709Z","shell.execute_reply.started":"2024-02-12T10:35:22.101469Z","shell.execute_reply":"2024-02-12T10:35:25.114307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"eeg_id\"].nunique()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:25.120011Z","iopub.execute_input":"2024-02-12T10:35:25.120609Z","iopub.status.idle":"2024-02-12T10:35:25.131562Z","shell.execute_reply.started":"2024-02-12T10:35:25.120561Z","shell.execute_reply":"2024-02-12T10:35:25.130097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(os.listdir(CFG.dataset_path+\"/train_eegs\"))-train[\"eeg_id\"].nunique()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:25.135215Z","iopub.execute_input":"2024-02-12T10:35:25.135703Z","iopub.status.idle":"2024-02-12T10:35:25.152915Z","shell.execute_reply.started":"2024-02-12T10:35:25.135668Z","shell.execute_reply":"2024-02-12T10:35:25.151476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Regarding the eeg id , there 211 missing files , and we should knoown where the problem come from","metadata":{}},{"cell_type":"code","source":"train_eggs_files=os.listdir(CFG.dataset_path+\"/train_eegs\")","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:25.154994Z","iopub.execute_input":"2024-02-12T10:35:25.155866Z","iopub.status.idle":"2024-02-12T10:35:25.175216Z","shell.execute_reply.started":"2024-02-12T10:35:25.155814Z","shell.execute_reply":"2024-02-12T10:35:25.174054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"only_file_eeg_id=[int(x.split(\".\")[0]) for x in train_eggs_files]","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:25.177051Z","iopub.execute_input":"2024-02-12T10:35:25.177901Z","iopub.status.idle":"2024-02-12T10:35:25.196209Z","shell.execute_reply.started":"2024-02-12T10:35:25.177854Z","shell.execute_reply":"2024-02-12T10:35:25.194984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_eegs_id=[int(x) for x in list(set(train[\"eeg_id\"])) ]","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:25.697428Z","iopub.execute_input":"2024-02-12T10:35:25.698641Z","iopub.status.idle":"2024-02-12T10:35:25.729071Z","shell.execute_reply.started":"2024-02-12T10:35:25.698568Z","shell.execute_reply":"2024-02-12T10:35:25.727505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\", \".join([str(x) for x in list(set(only_file_eeg_id)-set(train_eegs_id))])","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:25.731682Z","iopub.execute_input":"2024-02-12T10:35:25.732132Z","iopub.status.idle":"2024-02-12T10:35:25.753855Z","shell.execute_reply.started":"2024-02-12T10:35:25.732089Z","shell.execute_reply":"2024-02-12T10:35:25.752255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[train[\"eeg_id\"]==3576774144]","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:25.755060Z","iopub.execute_input":"2024-02-12T10:35:25.755407Z","iopub.status.idle":"2024-02-12T10:35:25.776889Z","shell.execute_reply.started":"2024-02-12T10:35:25.755378Z","shell.execute_reply":"2024-02-12T10:35:25.775209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can infer based on what we did, that ther are files that are'nt in the train.csv file ","metadata":{}},{"cell_type":"markdown","source":"### Expert Consensus Analysis","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,8))\nsns.countplot(train, x=\"expert_consensus\",palette=\"GnBu\")\nplt.title(\"Distribution of expert consensus\")\nplt.xlabel(\"Expert consensus\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:26.571994Z","iopub.execute_input":"2024-02-12T10:35:26.572769Z","iopub.status.idle":"2024-02-12T10:35:27.060626Z","shell.execute_reply.started":"2024-02-12T10:35:26.572730Z","shell.execute_reply":"2024-02-12T10:35:27.059259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As shown in the figure above , there is a smal inbalance between the experts consensus","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,14))\nfor i, column in enumerate(CFG.targets,1):\n    plt.subplot(4, 2, i)\n    plt.subplots_adjust(hspace=0.5)\n    sns.violinplot(data=train, x='expert_consensus', y=column)\n    plt.title(f'Distribution of {column} by Expert Consensus')\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:27.448250Z","iopub.execute_input":"2024-02-12T10:35:27.448679Z","iopub.status.idle":"2024-02-12T10:35:34.191946Z","shell.execute_reply.started":"2024-02-12T10:35:27.448648Z","shell.execute_reply":"2024-02-12T10:35:34.190766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As Shown above, the distribution of expert_consensus for each vote patterns indicates the  moste agreement however we don't have to forget the others consensus  ","metadata":{}},{"cell_type":"markdown","source":"### Label Distribution Analysis","metadata":{}},{"cell_type":"code","source":"train[CFG.targets].sum()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:34.194299Z","iopub.execute_input":"2024-02-12T10:35:34.194709Z","iopub.status.idle":"2024-02-12T10:35:34.207229Z","shell.execute_reply.started":"2024-02-12T10:35:34.194675Z","shell.execute_reply":"2024-02-12T10:35:34.206222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"percentages_label=(train[CFG.targets].sum()*100)/train[CFG.targets].sum().sum()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:34.208888Z","iopub.execute_input":"2024-02-12T10:35:34.209794Z","iopub.status.idle":"2024-02-12T10:35:34.223582Z","shell.execute_reply.started":"2024-02-12T10:35:34.209749Z","shell.execute_reply":"2024-02-12T10:35:34.222249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Frequency\nplt.figure(figsize=(10,8))\nsns.barplot(x=percentages_label.index,y=percentages_label,palette=\"GnBu\")\nplt.xlabel('Target Variables')\nplt.ylabel('Percentage')\nplt.title('Percentage Distribution of Target Variables')\nplt.xticks(rotation=45)  # Rotate x-axis labels if needed for better readability\nplt.tight_layout()  # Adjust layout to prevent clipping of labels\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:34.225078Z","iopub.execute_input":"2024-02-12T10:35:34.225522Z","iopub.status.idle":"2024-02-12T10:35:34.601486Z","shell.execute_reply.started":"2024-02-12T10:35:34.225484Z","shell.execute_reply":"2024-02-12T10:35:34.599937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vote_counts_by_consensus = train.groupby('expert_consensus')[CFG.targets].sum()\n\nplt.figure(figsize=(12, 8))\nvote_counts_by_consensus.plot(kind='bar', stacked=True)\nplt.title('Overall Vote Counts by Expert Consensus')\nplt.xlabel('Expert Consensus')\nplt.ylabel('Total Votes')\nplt.xticks(rotation=45)\nplt.legend(title='Vote Types')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:34.605656Z","iopub.execute_input":"2024-02-12T10:35:34.606196Z","iopub.status.idle":"2024-02-12T10:35:35.122835Z","shell.execute_reply.started":"2024-02-12T10:35:34.606150Z","shell.execute_reply":"2024-02-12T10:35:35.121371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15, 10))\nfor i, column in enumerate(CFG.targets, 1):\n    plt.subplot(4, 2, i)\n    sns.histplot(train[column], kde=False, bins=30)\n    plt.title(column)\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:35.124866Z","iopub.execute_input":"2024-02-12T10:35:35.125423Z","iopub.status.idle":"2024-02-12T10:35:38.363677Z","shell.execute_reply.started":"2024-02-12T10:35:35.125378Z","shell.execute_reply":"2024-02-12T10:35:38.362352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the correlation \n\nplt.figure(figsize=(10,8))\nsns.heatmap(train[CFG.targets].corr(), annot=True,cmap=\"Blues\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:38.365551Z","iopub.execute_input":"2024-02-12T10:35:38.365972Z","iopub.status.idle":"2024-02-12T10:35:38.846387Z","shell.execute_reply.started":"2024-02-12T10:35:38.365937Z","shell.execute_reply":"2024-02-12T10:35:38.844981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Regarding the correlation, there are no insights interpretation.","metadata":{}},{"cell_type":"markdown","source":"### Time Offset Analysis","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,6))\nsns.histplot(train['eeg_label_offset_seconds'], bins='auto',log_scale=(False, True))\nplt.xlabel('Offset Seconds')\nplt.ylabel('Frequency (Log Scale)')\nplt.title('Histogram of Offset Seconds (Log Scale)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:38.848410Z","iopub.execute_input":"2024-02-12T10:35:38.848928Z","iopub.status.idle":"2024-02-12T10:35:41.812965Z","shell.execute_reply.started":"2024-02-12T10:35:38.848891Z","shell.execute_reply":"2024-02-12T10:35:41.811318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,6))\nsns.histplot(train['spectrogram_label_offset_seconds'], bins='auto',log_scale=(False, True))\nplt.xlabel('Offset Seconds')\nplt.ylabel('Frequency (Log Scale)')\nplt.title('Histogram of Offset Seconds (Log Scale)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:41.815071Z","iopub.execute_input":"2024-02-12T10:35:41.816835Z","iopub.status.idle":"2024-02-12T10:35:44.926547Z","shell.execute_reply.started":"2024-02-12T10:35:41.816774Z","shell.execute_reply":"2024-02-12T10:35:44.925524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The majority of EEG and Spectrogramm samples have offset times close to the beginning of the EEG recordings. many EEG samples start almost immediately after the recording begin","metadata":{}},{"cell_type":"markdown","source":"Here's a summary of the key points regarding the training data:\n\n1. **Data Overview**:\n   - The training data consists of 106,800 rows, each representing a labeled EEG subsample.\n   - Each row contains an ID corresponding to both EEG and spectrogram data, which will be used for making predictions.\n\n2. **Target Columns**:\n   - The target columns include `lpd_vote`, `gpd_vote`, `lrda_vote`, `grda_vote`, and `other_vote`.\n   - These columns must contain probabilities and represent different types of brain activity classes.\n\n3. **Subsampling and Offset Information**:\n   - Each row includes `sub_id` and `offset_seconds` for both EEG and spectrogram data.\n   - During inference, 50-second-long EEG subsamples and 10-minute-long spectrogram subsamples will be used to make predictions.\n   - Therefore, target values apply to these subsamples rather than the entire EEG and spectrograms.\n\n4. **Patient ID**:\n   - The `patient_id` column is provided and can be useful for creating a train/validation/test split.\n   - It helps ensure that data from the same patient does not appear in both the training and validation/test sets to prevent data leakage.\n\n5. **Expert Consensus**:\n   - The `expert_consensus` column indicates the level of agreement among expert annotators.\n   - low consensus among expert annotators indicates uncertainty or controversy regarding the classification of certain samples in the dataset, which could potentially affect the performance of predictive models trained on that data. Excluding such samples can help ensure the quality and consistency of the dataset for model training and evaluation.\n6. **Offsets**\n    - Most offsets in the training data are very low: This indicates that the majority of EEG samples have offset times close to the beginning of the EEG recordings. In other words, many EEG samples start almost immediately after the recording begins.\n\nIn summary, the training data contains labeled EEG subsamples, with target columns representing probabilities of different brain activity classes. During inference, subsamples of EEG and spectrogram data are used for prediction, and patient IDs can aid in dataset splitting. The expert consensus column helps identify potentially contentious samples.","metadata":{}},{"cell_type":"markdown","source":"### Explore train EEGs","metadata":{}},{"cell_type":"code","source":"train_eegs_sample=pd.read_parquet(CFG.dataset_path+\"/train_eegs/1000913311.parquet\")","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:36:16.836409Z","iopub.execute_input":"2024-02-12T10:36:16.836856Z","iopub.status.idle":"2024-02-12T10:36:16.853506Z","shell.execute_reply.started":"2024-02-12T10:36:16.836826Z","shell.execute_reply":"2024-02-12T10:36:16.851737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_eegs_sample.head()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:36:29.361865Z","iopub.execute_input":"2024-02-12T10:36:29.362430Z","iopub.status.idle":"2024-02-12T10:36:29.397519Z","shell.execute_reply.started":"2024-02-12T10:36:29.362385Z","shell.execute_reply":"2024-02-12T10:36:29.396523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(os.listdir(CFG.dataset_path+\"/train_eegs\"))","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:35:45.105806Z","iopub.execute_input":"2024-02-12T10:35:45.106308Z","iopub.status.idle":"2024-02-12T10:35:45.123340Z","shell.execute_reply.started":"2024-02-12T10:35:45.106264Z","shell.execute_reply":"2024-02-12T10:35:45.122469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_eegs_sample.index=pd.to_timedelta(train_eegs_sample.index / 200, unit='s')","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:44:34.795579Z","iopub.execute_input":"2024-02-12T10:44:34.796150Z","iopub.status.idle":"2024-02-12T10:44:34.816928Z","shell.execute_reply.started":"2024-02-12T10:44:34.796098Z","shell.execute_reply":"2024-02-12T10:44:34.815655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_eegs_sample.columns)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:46:12.012826Z","iopub.execute_input":"2024-02-12T10:46:12.013278Z","iopub.status.idle":"2024-02-12T10:46:12.021765Z","shell.execute_reply.started":"2024-02-12T10:46:12.013245Z","shell.execute_reply":"2024-02-12T10:46:12.020269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traces = []\nfor column in train_eegs_sample.columns:\n    trace = go.Scatter(\n        x=train_eegs_sample.index,  \n        y=train_eegs_sample[column],\n        mode='lines',\n        name=column  \n    )\n    traces.append(trace)\n\n# Create layout for the plot\nlayout = go.Layout(\n    title='EEG Data Visualization',\n    xaxis=dict(title='Time'),\n    yaxis=dict(title='EEG Signal')\n)\n\n# Create figure and plot\nfig = go.Figure(data=traces, layout=layout)\n\n# Display the plot\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T10:48:17.374257Z","iopub.execute_input":"2024-02-12T10:48:17.374945Z","iopub.status.idle":"2024-02-12T10:48:18.382544Z","shell.execute_reply.started":"2024-02-12T10:48:17.374913Z","shell.execute_reply":"2024-02-12T10:48:18.380210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All the columns which represent the electrodes' parts are between -200 and 200, except for the EKG, which ranges from -4000 to 4000","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}