{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7457433,"sourceType":"competition"}],"dockerImageVersionId":30626,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"\n# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-01-10T20:08:40.696304Z","iopub.execute_input":"2024-01-10T20:08:40.696674Z","iopub.status.idle":"2024-01-10T20:08:40.701843Z","shell.execute_reply.started":"2024-01-10T20:08:40.696645Z","shell.execute_reply":"2024-01-10T20:08:40.700945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Overview\n\nKey points about the competition include:\n\n* **Objective:** The primary goal is to detect and classify seizures as well as other harmful brain activities using machine learning models.\n\n* **Data Source**: The data used for training the model comes from EEG signals recorded from critically ill patients in a hospital. EEG is a technique that measures electrical activity in the brain through electrodes placed on the scalp.\n\n* **Potential Impact**: The competition emphasizes the potential transformative benefits for various medical fields, particularly in neurocritical care, epilepsy treatment, and drug development.\n\n* **Medical Application**: Successful advancements in EEG pattern classification accuracy have the potential to significantly impact neurocritical care. This could lead to quicker and more precise treatments for patients experiencing seizures or other forms of brain damage.\n\n* **Broader Implications**: The competition suggests that improvements in this area could have far-reaching effects, enabling doctors and brain researchers to rapidly detect and respond to abnormal brain activities. This, in turn, may enhance the overall understanding of neurological disorders and contribute to the development of more effective treatment strategies.\n\nIn summary, the competition aims to harness machine learning and EEG data to enhance the accuracy of detecting and classifying seizures and harmful brain activities. The ultimate goal is to bring about positive changes in neurocritical care, epilepsy treatment, and drug development by providing faster and more accurate diagnostic tools for medical professionals.\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# 2. Description\n\n**1. Current Medical Practice with EEG:**\n\nDoctors currently use electroencephalography (EEG) as a diagnostic tool for critically ill patients to detect seizures and other forms of brain activity that could lead to brain damage.\nEEG signals are interpreted by specialized neurologists through a manual analysis process.\n\n**2. Challenges in Manual EEG Analysis:**\n\nManual analysis of EEG recordings is a time-consuming and expensive process.\nIt is prone to errors due to fatigue, and reliability issues may arise among different reviewers, even if they are experts.\n\n**3. Objective of the Competition:**\n\nThe competition aims to automate EEG analysis, which would significantly alleviate the challenges associated with manual analysis.\nThe goal is to develop algorithms that can accurately detect six patterns of interest: seizure (SZ), generalized periodic discharges (GPD), lateralized periodic discharges (LPD), lateralized rhythmic delta activity (LRDA), generalized rhythmic delta activity (GRDA), or \"other.\"\n\n**4. Competition Host - Sunstella Foundation:**\n\nThe Sunstella Foundation was established in 2021 during the COVID pandemic with a focus on supporting minority graduate students in technology.\nThe foundation aims to help these students overcome challenges and celebrate their achievements through workshops, forums, and competitions.\n\n**5. Partners and Collaborators:**\n\nSunstella Foundation collaborates with Persyst, Jazz Pharmaceuticals, and the Clinical Data Animation Center (CDAC).\nThe partners share a common research goal of preserving and enhancing brain health.\n\n**6. Significance of Automated EEG Analysis:**\n\nAutomating EEG analysis is expected to expedite the detection of seizures and other harmful brain activities.\nThis, in turn, will facilitate quicker and more accurate treatments for patients.\nThe developed algorithms may also aid researchers in drug development for the treatment and prevention of seizures.\n\n**7. Patterns of Interest:**\n\nThere are six patterns of interest, each representing different types of brain activity.\nDetailed explanations of these patterns are available for reference.\nAnnotation of EEG Segments:\n\nEEG segments used in the competition have been annotated or classified by a group of experts.\nAgreement levels among experts vary, with some cases having high agreement (\"idealized\" patterns), some having disagreements (~1/2 experts label as \"other\" and ~1/2 label as one of the remaining five patterns - \"proto patterns\"), and others having experts split between two of the five named patterns (\"edge cases\").\nIn summary, the competition seeks to leverage machine learning to automate the analysis of EEG signals, with the ultimate goal of improving the speed and accuracy of detecting various patterns of brain activity, especially seizures, in critically ill patients. This has the potential to revolutionize neurocritical care, epilepsy treatment, and drug development.\n\n\n![](https://storage.googleapis.com/kaggle-media/competitions/Harvard%20Medical%20School/eFig2.png)\n\n**Figure Structure:**\n\n**Rows**:\n\n1. Seizure (SZ)\n2. Lateralized Periodic Discharges (LPDs)\n3. Generalized Periodic Discharges (GPDs)\n4. Lateralized Rhythmic Delta Activity (LRDA)\n5. Generalized Rhythmic Delta Activity (GRDA)\n\n**Columns**:\n\n1. Idealized forms of patterns (A) - Patterns with uniform expert agreement.\n2. Proto or partially formed patterns (B) - About half of raters labeled these as one pattern and the other half labeled as \"Other.\"\n3, Edge cases (C) - About half of raters labeled these as one pattern and half labeled them as another pattern (column C).\n4. More edge cases (D) - Similar to column C.\n\n**Explanation of Examples:**\n\nColumn B (Proto or Partially Formed Patterns):\n\n* B-1: Rhythmic delta activity with some sharp discharges, potentially the tail end of a seizure, causing disagreement between SZ and \"Other.\"\n* B-2: Frontal lateralized sharp transients with reversed polarity, suggesting non-cerebral source, leading to split between LPD and \"Other.\"\n* B-3: Diffused semi-rhythmic delta background with poorly formed low amplitude periodic discharges, a proto-GPD.\n* B-4: Semi-rhythmic delta activity with unstable morphology over the right hemisphere, a proto-LRDA.\n* B-5: Waves of rhythmic delta activity with unstable morphology, a proto-GRDA.\n\nColumns C and D (Edge Cases):\nExamples with features straddling two patterns:\n\n* C-1: LPDs evolving into a seizure, an edge-case.\n* D-1: GPDs on a suppressed background, suggesting a seizure, another edge case.\n* C-2: Split between LPDs and GPDs.\n* D-2: Tied between LPDs and LRDA, sharing features of both.\n* C-3: Split between GPDs and LRDA.\n* D-3: Split between GPDs and GRDA, showing asymmetry in slope.\n* C-4: Split between LRDA and seizure.\n* D-4: Split between LRDA and GRDA, asymmetry in delta wave.\n* C-5: Split between GRDA and seizure.\n* D-5: Split between GRDA and LPDs, showing generalized rhythmic delta activity with features suggestive of LPDs.\n\n**Note:**\n* EEG electrode recording regions are abbreviated as LL (left lateral), RL (right lateral), LP (left parasagittal), and RP (right parasagittal).\n\nThis detailed explanation provides insights into the complexities of EEG pattern classification, showcasing examples that challenge clear categorization and highlight the nuances involved in expert interpretation.\n\n\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# 3. Data Analysis \n\n**1. Files and Folders**\n\n**2. Input data**\n\n    * train.csv\n    * test.csv\n    * XXXXX.parquet\n    \n**3. Output Data**\n\n    *  sample_submission.csv\n    \n","metadata":{}},{"cell_type":"code","source":"# Import\nimport os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:08:40.703230Z","iopub.execute_input":"2024-01-10T20:08:40.703561Z","iopub.status.idle":"2024-01-10T20:08:40.716122Z","shell.execute_reply.started":"2024-01-10T20:08:40.703519Z","shell.execute_reply":"2024-01-10T20:08:40.714788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_folder_info(folder_path):\n    # Initialize lists to store information\n    files_list = []\n    folders_list = []\n    sizes_list = []\n    types_list = []\n\n    # Iterate over each item in the folder\n    for item in os.listdir(folder_path):\n        item_path = os.path.join(folder_path, item)\n\n        # Check if it's a file or a folder\n        if os.path.isfile(item_path):\n            files_list.append(item)\n            sizes_list.append(os.path.getsize(item_path))\n            types_list.append(item.split('.')[-1].lower()) if '.' in item else 'Unknown'\n            folders_list.append('')\n        elif os.path.isdir(item_path):\n            files_list.append('')\n            sizes_list.append('')\n            types_list.append('')\n            folders_list.append(item)\n\n    # Create a DataFrame for tabular representation\n    df = pd.DataFrame({\n        'File': files_list,\n        'Folder': folders_list,\n        'Size (bytes)': sizes_list,\n        'Type': types_list\n    })\n\n    return df\n\n# Provide the path to the folder you want to analyze\nfolder_path = '/kaggle/input/hms-harmful-brain-activity-classification'\n\n# Get information about the folder\nfolder_info = get_folder_info(folder_path)\n\n# Display the information\nprint(folder_info)\n","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:08:40.717692Z","iopub.execute_input":"2024-01-10T20:08:40.718109Z","iopub.status.idle":"2024-01-10T20:08:40.736509Z","shell.execute_reply.started":"2024-01-10T20:08:40.718068Z","shell.execute_reply":"2024-01-10T20:08:40.735099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\nprint(train_df)\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:08:40.738389Z","iopub.execute_input":"2024-01-10T20:08:40.739236Z","iopub.status.idle":"2024-01-10T20:08:40.757977Z","shell.execute_reply.started":"2024-01-10T20:08:40.739197Z","shell.execute_reply":"2024-01-10T20:08:40.756595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/test.csv\")\ntest_df","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:08:40.759215Z","iopub.execute_input":"2024-01-10T20:08:40.759813Z","iopub.status.idle":"2024-01-10T20:08:40.770942Z","shell.execute_reply.started":"2024-01-10T20:08:40.759774Z","shell.execute_reply":"2024-01-10T20:08:40.769043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission_df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv\")\nsample_submission_df\n","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:08:40.773704Z","iopub.execute_input":"2024-01-10T20:08:40.774179Z","iopub.status.idle":"2024-01-10T20:08:40.789157Z","shell.execute_reply.started":"2024-01-10T20:08:40.774141Z","shell.execute_reply":"2024-01-10T20:08:40.787979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pandas as pd\n\ndef get_folder_info(folder_path):\n    # Initialize lists to store information\n    files_list = []\n    folders_list = []\n    sizes_list = []\n    types_list = []\n\n    # Recursive function to explore subfolders\n    def explore_folder(current_path):\n        for item in os.listdir(current_path):\n            item_path = os.path.join(current_path, item)\n\n            if os.path.isfile(item_path):\n                files_list.append(item)\n                sizes_list.append(os.path.getsize(item_path))\n                types_list.append(item.split('.')[-1].lower()) if '.' in item else 'Unknown'\n                folders_list.append(current_path[len(folder_path) + 1:])  # Relative path to the main folder\n            elif os.path.isdir(item_path):\n                explore_folder(item_path)\n\n    # Call the recursive function\n    explore_folder(folder_path)\n\n    # Create a DataFrame for tabular representation\n    df = pd.DataFrame({\n        'File': files_list,\n        'Folder': folders_list,\n        'Size (bytes)': sizes_list,\n        'Type': types_list\n    })\n\n    return df\n\n# Provide the path to the main folder you want to analyze\nfolder_path = '/kaggle/input/hms-harmful-brain-activity-classification'\n\n# Get information about the folder and its subfolders\nfolder_info = get_folder_info(folder_path)\n\n# Display the information\nprint(folder_info)\n","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:08:40.790622Z","iopub.execute_input":"2024-01-10T20:08:40.792108Z","iopub.status.idle":"2024-01-10T20:09:25.170987Z","shell.execute_reply.started":"2024-01-10T20:08:40.790960Z","shell.execute_reply":"2024-01-10T20:09:25.168205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_eeg_0001 = pd.read_parquet(\"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/2208063991.parquet\")\ntrain_eeg_0001","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:09:25.173600Z","iopub.execute_input":"2024-01-10T20:09:25.174199Z","iopub.status.idle":"2024-01-10T20:09:25.209539Z","shell.execute_reply.started":"2024-01-10T20:09:25.174155Z","shell.execute_reply":"2024-01-10T20:09:25.208213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\n\n# Read the parquet file\nfile_path = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/2208063991.parquet\"\ntrain_eeg_0001 = pd.read_parquet(file_path)\n\n# List of EEG channels\neeg_channels = train_eeg_0001.columns[:-1]  # Exclude the last column (EKG)\n\n# Plotting each EEG channel\nfor channel in eeg_channels:\n    plt.figure(figsize=(10, 5))\n    plt.plot(train_eeg_0001[channel], label=channel)\n    plt.title(f'EEG Channel: {channel}')\n    plt.xlabel('Time Steps')\n    plt.ylabel('EEG Signal Value')\n    plt.legend()\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:09:25.211256Z","iopub.execute_input":"2024-01-10T20:09:25.211927Z","iopub.status.idle":"2024-01-10T20:09:30.166262Z","shell.execute_reply.started":"2024-01-10T20:09:25.211890Z","shell.execute_reply":"2024-01-10T20:09:30.165163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_spectrogram_0001 = pd.read_parquet(\"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/1000086677.parquet\")","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:09:30.168104Z","iopub.execute_input":"2024-01-10T20:09:30.168710Z","iopub.status.idle":"2024-01-10T20:09:30.200198Z","shell.execute_reply.started":"2024-01-10T20:09:30.168679Z","shell.execute_reply":"2024-01-10T20:09:30.199231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_spectrogram_0001","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:09:30.201462Z","iopub.execute_input":"2024-01-10T20:09:30.201718Z","iopub.status.idle":"2024-01-10T20:09:30.231326Z","shell.execute_reply.started":"2024-01-10T20:09:30.201693Z","shell.execute_reply":"2024-01-10T20:09:30.229625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# Read the data\ntrain_spectrogram_0001 = pd.read_parquet(\"/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/1000086677.parquet\")\n\n# Display basic information about the DataFrame\nprint(train_spectrogram_0001.info())\n\n# Display summary statistics for numeric columns\nprint(train_spectrogram_0001.describe())\n\n# Count missing values in each column\nmissing_values = train_spectrogram_0001.isnull().sum()\nprint(\"Missing Values:\\n\", missing_values[missing_values > 0])\n\n# Calculate correlation matrix for a subset of columns\nsubset_columns = ['LL_0.59', 'LL_0.78', 'LL_0.98', 'RP_18.16', 'RP_18.36', 'RP_18.55']\ncorrelation_matrix = train_spectrogram_0001[subset_columns].corr()\nprint(\"Correlation Matrix:\\n\", correlation_matrix)\n\n# Plot histograms for a subset of columns\ntrain_spectrogram_0001[subset_columns].hist(bins=20, figsize=(15, 8))\nplt.suptitle(\"Histograms of Selected Columns\")\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-01-10T20:09:30.233764Z","iopub.execute_input":"2024-01-10T20:09:30.234195Z","iopub.status.idle":"2024-01-10T20:09:31.948864Z","shell.execute_reply.started":"2024-01-10T20:09:30.234164Z","shell.execute_reply":"2024-01-10T20:09:31.947629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Reference\n \nhttps://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf\n\nhttps://media.journals.elsevier.com/content/files/clinph-chapter1-5-14083047.pdf\n\nhttp://ulae.org.ua/attachments/article/123/ICU%20EEG%20Terminology%20SB2018_SHORT.pdf","metadata":{}}]}