{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30635,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Overview\n\nThe goal of this competition is to detect and classify seizures and other types of harmful brain activity. You will develop a model trained on electroencephalography (EEG) signals recorded from critically ill hospital patients.\n\n## Dataset Description\n\nThe goal of this competition is to detect and classify seizures and other types of harmful brain activity in electroencephalography (EEG) data. Even experts find this to be a challenging task and often disagree about the correct labels.\n\n\nThis is a code competition. Only a few examples from the test set are available for download. When your submission is scored the test folders will be replaced with versions containing the complete test set.\n\n## Files\n**train.csv** \nMetadata for the train set. The expert annotators reviewed 50 second long EEG samples plus matched spectrograms covering 10 a minute window centered at the same time and labeled the central 10 seconds. Many of these samples overlapped and have been consolidated. train.csv provides the metadata that allows you to extract the original subsets that the raters annotated.\n\n- eeg_id - A unique identifier for the entire EEG recording.\n- eeg_sub_id - An ID for the specific 50 second long subsample this row's labels apply to.\n- eeg_label_offset_seconds - The time between the beginning of the consolidated EEG and this subsample.\n- spectrogram_id - A unique identifier for the entire EEG recording.\n- spectrogram_sub_id - An ID for the specific 10 minute subsample this row's labels apply to.\n- spectogram_label_offset_seconds - The time between the beginning of the consolidated spectrogram and this subsample.\n- label_id - An ID for this set of labels.\n- patient_id - An ID for the patient who donated the data.\n- expert_consensus - The consensus annotator label. Provided for convenience only.\n- [seizure/lpd/gpd/lrda/grda/other]_vote - The count of annotator votes for a given brain activity class. The full names of the activity classes are as follows: lpd: lateralized periodic discharges, gpd: generalized periodic discharges, lrd: lateralized rhythmic delta activity, and grda: generalized rhythmic delta activity . A detailed explanations of these patterns is available here.\n\n**test.csv** \nMetadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.\n\n- eeg_id\n- spectrogram_id\n- patient_id\n\n**sample_submission.csv**\n\n- eeg_id\n- [seizure/lpd/gpd/lrda/grda/other]_vote - The target columns. Your predictions must be probabilities. Note that the test samples had between 3 and 20 annotators.\n\n**train_eegs**/ EEG data from one or more overlapping samples. Use the metadata in **train.csv** to select specific annotated subsets. The column names are the names of the individual electrode locations for EEG leads, with one exception. The EKG column is for an electrocardiogram lead that records data from the heart. All of the EEG data (for both train and test) was collected at a frequency of 200 samples per second.\n\n**test_eegs**/ Exactly 50 seconds of EEG data.\n\n**train_spectrograms**/ Spectrograms assembled EEG data. Use the metadata in **train.csv** to select specific annotated subsets. The column names indicate the frequency in hertz and the recording regions of the EEG electrodes. The latter are abbreviated as LL = left lateral; RL = right lateral; LP = left parasagittal; RP = right parasagittal.\n\n**test_spectrograms**/ Spectrograms assembled using exactly 10 minutes of EEG data.\n\n**example_figures**/ Larger copies of the example case images used on the overview tab.","metadata":{}},{"cell_type":"code","source":"# Libraries for reading and manipulating data\nimport numpy as np\nimport pandas as pd\n\n# Libraries for data visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load sample submission format\nsample_submission = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv')\n\n# Load training and testing metadata\ntrain_df = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\ntest_df = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Path for example EEG figures\nexample_figures_path = \"/kaggle/input/hms-harmful-brain-activity-classification/example_figures\"\n\n# Paths to folders containing EEG data\ntrain_eegs_path = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs\"\ntest_eegs_path = \"/kaggle/input/hms-harmful-brain-activity-classification/test_eegs\"\n\n# Paths to folders containing spectrogram data\ntrain_spectrograms_path = \"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms\"\ntest_spectrograms_path = \"/kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms\"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Preprocessing and Feature Engineering:","metadata":{}},{"cell_type":"markdown","source":"### Loading EEG and Spectrogram Data Function:\n\nDefine functions to load EEG and spectrogram data given a file path.","metadata":{}},{"cell_type":"code","source":"# Function to load EEG data\ndef load_eeg_data(file_path):\n    # Load EEG data from the file path\n    # ...\n\n# Example usage\nsome_eeg_data = load_eeg_data(train_eegs_path + '/some_file_name.csv')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Preprocessing Functions:\n\nDevelop preprocessing functions for both EEG and spectrogram data to handle any data cleaning, transformation, or feature extraction necessary for your model.","metadata":{}},{"cell_type":"code","source":"# Function for preprocessing EEG data\ndef preprocess_eeg_data(eeg_data):\n    # Preprocess the EEG data\n    # ...\n\n# Function for preprocessing Spectrogram data\ndef preprocess_spectrogram_data(spectrogram_data):\n    # Preprocess the Spectrogram data\n    # ...","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis\n\nVisualization and statistical analysis:\n","metadata":{}},{"cell_type":"markdown","source":"# Model Developement\n\nFeature Selection\n\nModel Training\n\nHyperparameter Tuning\n\nModel Evaluation","metadata":{}},{"cell_type":"markdown","source":"## Model Training and Evaluation Functions:","metadata":{}},{"cell_type":"code","source":"# Function to train the model\ndef train_model(X_train, y_train):\n    # Train the model\n    # ...\n\n# Function to evaluate the model\ndef evaluate_model(model, X_test, y_test):\n    # Evaluate the model\n    # ...\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Predictions and Submission Preparation:","metadata":{}},{"cell_type":"code","source":"# Make predictions on the test data\ntest_predictions = model.predict(test_eegs_data)\n\n# Prepare the submission DataFrame\nsubmission_df = pd.DataFrame({\n    'eeg_id': test_df['eeg_id'],\n    'predicted_class': test_predictions\n})","metadata":{},"execution_count":null,"outputs":[]}]}