{"metadata":{"colab":{"provenance":[]},"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30787,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# KAGGLE COMPETITION","metadata":{}},{"cell_type":"markdown","source":"# HMS - Harmful Brain Activity Classification","metadata":{}},{"cell_type":"markdown","source":"### IMPORTING THE NECESSARY LIBRARIES","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"id":"a6684e18","execution":{"iopub.status.busy":"2024-10-19T18:57:10.768203Z","iopub.execute_input":"2024-10-19T18:57:10.768926Z","iopub.status.idle":"2024-10-19T18:57:10.773343Z","shell.execute_reply.started":"2024-10-19T18:57:10.768884Z","shell.execute_reply":"2024-10-19T18:57:10.772418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')","metadata":{"id":"bd783d7e","execution":{"iopub.status.busy":"2024-10-19T18:57:10.775427Z","iopub.execute_input":"2024-10-19T18:57:10.775806Z","iopub.status.idle":"2024-10-19T18:57:10.783676Z","shell.execute_reply.started":"2024-10-19T18:57:10.775763Z","shell.execute_reply":"2024-10-19T18:57:10.782781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LOADING THE DATASET","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")","metadata":{"id":"ab2bfbe1","execution":{"iopub.status.busy":"2024-10-19T18:57:10.784623Z","iopub.execute_input":"2024-10-19T18:57:10.784941Z","iopub.status.idle":"2024-10-19T18:57:10.933218Z","shell.execute_reply.started":"2024-10-19T18:57:10.784893Z","shell.execute_reply":"2024-10-19T18:57:10.932173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"id":"4bb1c2a7","outputId":"29e6dd13-9524-44b8-a874-684a0fbab19c","execution":{"iopub.status.busy":"2024-10-19T18:57:10.935218Z","iopub.execute_input":"2024-10-19T18:57:10.935539Z","iopub.status.idle":"2024-10-19T18:57:10.950872Z","shell.execute_reply.started":"2024-10-19T18:57:10.935506Z","shell.execute_reply":"2024-10-19T18:57:10.949926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### BASIC CHECKS","metadata":{}},{"cell_type":"code","source":"df.tail()","metadata":{"id":"c0c2ac52","outputId":"8d61aefa-a834-4c02-db27-cb90b28ed904","execution":{"iopub.status.busy":"2024-10-19T18:57:10.952052Z","iopub.execute_input":"2024-10-19T18:57:10.952414Z","iopub.status.idle":"2024-10-19T18:57:10.972918Z","shell.execute_reply.started":"2024-10-19T18:57:10.952375Z","shell.execute_reply":"2024-10-19T18:57:10.971985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"id":"a5fb31fa","outputId":"bcb8422f-8f44-4a4a-ac90-344f9c9793d4","execution":{"iopub.status.busy":"2024-10-19T18:57:10.974043Z","iopub.execute_input":"2024-10-19T18:57:10.974358Z","iopub.status.idle":"2024-10-19T18:57:10.983090Z","shell.execute_reply.started":"2024-10-19T18:57:10.974327Z","shell.execute_reply":"2024-10-19T18:57:10.982271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns","metadata":{"id":"3c203298","outputId":"87a8fd0e-f916-4f70-c72c-c8da658fccbe","execution":{"iopub.status.busy":"2024-10-19T18:57:10.984057Z","iopub.execute_input":"2024-10-19T18:57:10.984392Z","iopub.status.idle":"2024-10-19T18:57:10.993675Z","shell.execute_reply.started":"2024-10-19T18:57:10.984351Z","shell.execute_reply":"2024-10-19T18:57:10.992783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## DATA ANALYSIS","metadata":{}},{"cell_type":"markdown","source":"1. **`eeg_id`**: Unique identifier for the entire EEG recording. It represents a specific EEG session or recording from a patient.\n\n2. **`eeg_sub_id`**: An ID for the specific 50-second long subsample to which the row's labels apply. This column helps identify a particular subset or segment within the larger EEG recording.\n\n3. **`eeg_label_offset_seconds`**: The time between the beginning of the consolidated EEG and the subsample. It indicates the time offset for the EEG subsample within the entire EEG recording.\n\n4. **`spectrogram_id`**: Unique identifier for the entire EEG recording, similar to `eeg_id`. It is related to the spectrogram data.\n\n5. **`spectrogram_sub_id`**: An ID for the specific 10-minute subsample to which the row's labels apply. This corresponds to a subset within the larger spectrogram data.\n\n6. **`spectrogram_label_offset_seconds`**: The time between the beginning of the consolidated spectrogram and the subsample. It indicates the time offset for the spectrogram subsample within the entire spectrogram recording.\n\n7. **`label_id`**: An ID for this set of labels. It helps distinguish different sets of labels within the dataset.\n\n8. **`patient_id`**: An ID for the patient who donated the data. It uniquely identifies each patient.\n\n9. **`expert_consensus`**: The consensus annotator label for convenience. This column may provide a summary or agreement among expert annotators regarding the type of brain activity in the given subsample.\n\n10. **`seizure_vote`**, **`lpd_vote`**, **`gpd_vote`**, **`lrda_vote`**, **`grda_vote`**, **`other_vote`**: These columns represent the count of annotator votes for specific brain activity classes. The classes are:\n    - `seizure_vote`: Count of votes for seizure.\n    - `lpd_vote`: Count of votes for lateralized periodic discharges.\n    - `gpd_vote`: Count of votes for generalized periodic discharges.\n    - `lrda_vote`: Count of votes for lateralized rhythmic delta activity.\n    - `grda_vote`: Count of votes for generalized rhythmic delta activity.\n    - `other_vote`: Count of votes for other types of brain activity.\n\n11. **Target Variable**: The target variable in this dataset is the actual brain activity class for each subsample. It could be any of the following:\n    - Seizure (`seizure_vote`): Represents the count of votes for seizure.\n    - Lateralized Periodic Discharges (`lpd_vote`): Represents the count of votes for lateralized periodic discharges.\n    - Generalized Periodic Discharges (`gpd_vote`): Represents the count of votes for generalized periodic discharges.\n    - Lateralized Rhythmic Delta Activity (`lrda_vote`): Represents the count of votes for lateralized rhythmic delta activity.\n    - Generalized Rhythmic Delta Activity (`grda_vote`): Represents the count of votes for generalized rhythmic delta activity.\n    - Other (`other_vote`): Represents the count of votes for other types of brain activity.\n\nThe target variable is the type of brain activity class (seizure, lpd, gpd, lrda, grda, other), and the goal of the competition or analysis is likely to predict or classify the correct brain activity class for each EEG subsample.","metadata":{"id":"24fe4409"}},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"id":"6a4d1266","outputId":"2897b253-e47a-4bdb-e90f-8dbeff3de890","execution":{"iopub.status.busy":"2024-10-19T18:57:10.997121Z","iopub.execute_input":"2024-10-19T18:57:10.997488Z","iopub.status.idle":"2024-10-19T18:57:11.016671Z","shell.execute_reply.started":"2024-10-19T18:57:10.997446Z","shell.execute_reply":"2024-10-19T18:57:11.015851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.nunique()","metadata":{"id":"16931f94","outputId":"45ab9b55-f749-422f-c6c7-dd0028bbf309","execution":{"iopub.status.busy":"2024-10-19T18:57:11.017881Z","iopub.execute_input":"2024-10-19T18:57:11.018171Z","iopub.status.idle":"2024-10-19T18:57:11.047714Z","shell.execute_reply.started":"2024-10-19T18:57:11.018141Z","shell.execute_reply":"2024-10-19T18:57:11.046889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"object_columns = df.select_dtypes(include=['object', 'bool']).columns\nprint(\"Object type columns:\")\nprint(object_columns)\n\nnumerical_columns = df.select_dtypes(include=['int64', 'float64']).columns\nprint(\"\\nNumerical type columns:\")\nprint(numerical_columns)","metadata":{"id":"3dbfbdf0","outputId":"e12bcadd-c5a8-4a1a-daf4-5c8a77c22dea","execution":{"iopub.status.busy":"2024-10-19T18:57:11.048917Z","iopub.execute_input":"2024-10-19T18:57:11.049505Z","iopub.status.idle":"2024-10-19T18:57:11.059754Z","shell.execute_reply.started":"2024-10-19T18:57:11.049460Z","shell.execute_reply":"2024-10-19T18:57:11.058792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def classify_features(df):\n    categorical_features = []\n    non_categorical_features = []\n    discrete_features = []\n    continuous_features = []\n\n    for column in df.columns:\n        if df[column].dtype in ['object', 'bool']:\n            if df[column].nunique() < 15:\n                categorical_features.append(column)\n            else:\n                non_categorical_features.append(column)\n        elif df[column].dtype in ['int64', 'float64']:\n            if df[column].nunique() < 10:\n                discrete_features.append(column)\n            else:\n                continuous_features.append(column)\n\n    return categorical_features, non_categorical_features, discrete_features, continuous_features","metadata":{"id":"8ee6e8ce","execution":{"iopub.status.busy":"2024-10-19T18:57:11.060848Z","iopub.execute_input":"2024-10-19T18:57:11.061196Z","iopub.status.idle":"2024-10-19T18:57:11.069827Z","shell.execute_reply.started":"2024-10-19T18:57:11.061147Z","shell.execute_reply":"2024-10-19T18:57:11.068903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical, non_categorical, discrete, continuous = classify_features(df)","metadata":{"id":"dc6734b4","execution":{"iopub.status.busy":"2024-10-19T18:57:11.071078Z","iopub.execute_input":"2024-10-19T18:57:11.071413Z","iopub.status.idle":"2024-10-19T18:57:11.101115Z","shell.execute_reply.started":"2024-10-19T18:57:11.071373Z","shell.execute_reply":"2024-10-19T18:57:11.100300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Categorical Features:\", categorical)\nprint(\"Non-Categorical Features:\", non_categorical)\nprint(\"Discrete Features:\", discrete)\nprint(\"Continuous Features:\", continuous)","metadata":{"id":"de33f77a","outputId":"a29604e3-9f3a-4342-c9ed-6892e0ae0ecb","execution":{"iopub.status.busy":"2024-10-19T18:57:11.102225Z","iopub.execute_input":"2024-10-19T18:57:11.102496Z","iopub.status.idle":"2024-10-19T18:57:11.107311Z","shell.execute_reply.started":"2024-10-19T18:57:11.102466Z","shell.execute_reply":"2024-10-19T18:57:11.106477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in categorical:\n    print(i, ':')\n    print(df[i].unique())\n    print()","metadata":{"id":"855580ed","outputId":"c0387edf-ad70-430e-eac3-98a6ba0ac49c","execution":{"iopub.status.busy":"2024-10-19T18:57:11.108623Z","iopub.execute_input":"2024-10-19T18:57:11.109010Z","iopub.status.idle":"2024-10-19T18:57:11.129215Z","shell.execute_reply.started":"2024-10-19T18:57:11.108969Z","shell.execute_reply":"2024-10-19T18:57:11.128280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in categorical:\n    print(i, ':')\n    print(df[i].value_counts())\n    print()","metadata":{"id":"69615077","outputId":"869ac4c1-f917-46ed-e036-74fb6c40bfd1","execution":{"iopub.status.busy":"2024-10-19T18:57:11.130532Z","iopub.execute_input":"2024-10-19T18:57:11.130977Z","iopub.status.idle":"2024-10-19T18:57:11.155362Z","shell.execute_reply.started":"2024-10-19T18:57:11.130933Z","shell.execute_reply":"2024-10-19T18:57:11.154559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## EXPLORATORY DATA ANALYSIS","metadata":{}},{"cell_type":"code","source":"for i in categorical:\n    plt.figure(figsize=(15, 6))\n    ax = sns.countplot(x=i, data=df, palette='hls')\n    for p in ax.patches:\n        ax.annotate(f'{p.get_height()}', (p.get_x() + p.get_width() / 2., p.get_height()),\n                    ha='center', va='center', xytext=(0, 10), textcoords='offset points')\n    plt.show()","metadata":{"id":"10dc4fbd","outputId":"31b6fa18-b42a-4448-830f-38ef9eb28a62","execution":{"iopub.status.busy":"2024-10-19T18:57:11.156604Z","iopub.execute_input":"2024-10-19T18:57:11.156943Z","iopub.status.idle":"2024-10-19T18:57:11.474685Z","shell.execute_reply.started":"2024-10-19T18:57:11.156904Z","shell.execute_reply":"2024-10-19T18:57:11.473761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in categorical:\n    plt.figure(figsize=(20,10))\n    plt.pie(df[i].value_counts(), labels=df[i].value_counts().index, autopct='%1.1f%%', textprops={'fontsize': 15,\n                                           'color': 'black',\n                                           'weight': 'bold',\n                                           'family': 'serif' })\n    hfont = {'fontname':'serif', 'weight': 'bold'}\n    plt.title(i, size=20, **hfont)\n    plt.show()","metadata":{"id":"230b9f16","outputId":"0b4d8029-adf3-4763-8717-59771507d935","execution":{"iopub.status.busy":"2024-10-19T18:57:11.475899Z","iopub.execute_input":"2024-10-19T18:57:11.476232Z","iopub.status.idle":"2024-10-19T18:57:11.666614Z","shell.execute_reply.started":"2024-10-19T18:57:11.476199Z","shell.execute_reply":"2024-10-19T18:57:11.665626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"votes_columns = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\ndf_votes = df[votes_columns].melt(var_name='Brain Activity', value_name='Votes')","metadata":{"id":"1a36b1f0","execution":{"iopub.status.busy":"2024-10-19T18:57:11.667703Z","iopub.execute_input":"2024-10-19T18:57:11.668408Z","iopub.status.idle":"2024-10-19T18:57:11.698762Z","shell.execute_reply.started":"2024-10-19T18:57:11.668362Z","shell.execute_reply":"2024-10-19T18:57:11.697864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. **`votes_columns`**: This is a list containing the names of columns in your original DataFrame (`df`) that represent the counts of votes for different brain activity classes. Each element in the list corresponds to a specific brain activity class.\n\n2. **`df_votes`**: This line creates a new DataFrame (`df_votes`) by selecting only the columns specified in `votes_columns` from the original DataFrame (`df`). The selected columns are essentially the counts of votes for different brain activity classes.\n\n3. **`.melt(var_name='Brain Activity', value_name='Votes')`**: This part of the code transforms the DataFrame from wide format to long format using the `melt` function. It takes the selected columns (`votes_columns`) and \"melts\" or unpivots them, creating two new columns:\n   - `Brain Activity`: This column stores the variable names from the original DataFrame (`df`) that represent different brain activity classes (e.g., 'seizure_vote', 'lpd_vote', etc.).\n   - `Votes`: This column stores the corresponding values (vote counts) for each brain activity class.","metadata":{"id":"b9ac5e00"}},{"cell_type":"code","source":"df_votes","metadata":{"id":"02fb6640","outputId":"2d2a2ef8-bca4-4356-faac-6bc70453b862","execution":{"iopub.status.busy":"2024-10-19T18:57:11.700272Z","iopub.execute_input":"2024-10-19T18:57:11.700741Z","iopub.status.idle":"2024-10-19T18:57:11.715219Z","shell.execute_reply.started":"2024-10-19T18:57:11.700678Z","shell.execute_reply":"2024-10-19T18:57:11.714165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.barplot(x='Brain Activity', y='Votes', hue='Brain Activity', data=df_votes, estimator=sum)\nplt.title('Distribution of Votes for Brain Activity Classes')\nplt.show()","metadata":{"id":"00cbff0d","outputId":"0dbf0b0b-1a6e-4a1b-c46d-d8b2a8c25fd0","execution":{"iopub.status.busy":"2024-10-19T18:57:11.723920Z","iopub.execute_input":"2024-10-19T18:57:11.724416Z","iopub.status.idle":"2024-10-19T18:59:01.412996Z","shell.execute_reply.started":"2024-10-19T18:57:11.724384Z","shell.execute_reply":"2024-10-19T18:59:01.412020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"votes_sum = df_votes.groupby('Brain Activity')['Votes'].sum()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:01.414284Z","iopub.execute_input":"2024-10-19T18:59:01.414673Z","iopub.status.idle":"2024-10-19T18:59:01.477979Z","shell.execute_reply.started":"2024-10-19T18:59:01.414628Z","shell.execute_reply":"2024-10-19T18:59:01.477253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nplt.figure(figsize=(10, 6))\nplt.pie(votes_sum, labels=votes_sum.index, autopct='%1.1f%%', startangle=90)\nplt.title('Distribution of Votes for Brain Activity Classes')\nplt.show()","metadata":{"id":"d49ea695","outputId":"b7441d14-6bf1-40aa-c2fb-cb0f9941b20b","execution":{"iopub.status.busy":"2024-10-19T18:59:01.479183Z","iopub.execute_input":"2024-10-19T18:59:01.479580Z","iopub.status.idle":"2024-10-19T18:59:01.661921Z","shell.execute_reply.started":"2024-10-19T18:59:01.479533Z","shell.execute_reply":"2024-10-19T18:59:01.660754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_vote_columns = df[votes_columns]","metadata":{"id":"8c8f7746","execution":{"iopub.status.busy":"2024-10-19T18:59:01.663789Z","iopub.execute_input":"2024-10-19T18:59:01.665069Z","iopub.status.idle":"2024-10-19T18:59:01.672886Z","shell.execute_reply.started":"2024-10-19T18:59:01.665002Z","shell.execute_reply":"2024-10-19T18:59:01.671510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_votes","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:01.674867Z","iopub.execute_input":"2024-10-19T18:59:01.676304Z","iopub.status.idle":"2024-10-19T18:59:01.691053Z","shell.execute_reply.started":"2024-10-19T18:59:01.676242Z","shell.execute_reply":"2024-10-19T18:59:01.689782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_vote_columns","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:01.692797Z","iopub.execute_input":"2024-10-19T18:59:01.693847Z","iopub.status.idle":"2024-10-19T18:59:01.707268Z","shell.execute_reply.started":"2024-10-19T18:59:01.693787Z","shell.execute_reply":"2024-10-19T18:59:01.706423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pivot_table_continuous = pd.pivot_table(df, values=continuous, index='expert_consensus', aggfunc='mean', fill_value=0)\npivot_table_continuous","metadata":{"id":"f1db71b7","outputId":"ff712249-e1be-4e04-94c9-4e69f291f0c4","execution":{"iopub.status.busy":"2024-10-19T18:59:01.708333Z","iopub.execute_input":"2024-10-19T18:59:01.708599Z","iopub.status.idle":"2024-10-19T18:59:01.750867Z","shell.execute_reply.started":"2024-10-19T18:59:01.708569Z","shell.execute_reply":"2024-10-19T18:59:01.750010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. **eeg_id:**\n   - GPD has an average EEG recording identifier of approximately 2.1126e+09.\n   - Other classes have varying average EEG recording identifiers.\n\n   **Interpretation:** The average EEG recording identifier may not provide direct insights into the characteristics of brain activity classes.\n\n2. **eeg_label_offset_seconds:**\n   - GPD has an average EEG label offset of approximately 271.79 seconds.\n   - Other classes have varying average EEG label offsets.\n\n   **Interpretation:** GPD tends to have EEG signals labeled around 271.79 seconds, suggesting a specific temporal pattern associated with this class.\n\n3. **spectrogram_id:**\n   - GPD has an average spectrogram identifier of approximately 1.01299e+09.\n   - Other classes have varying average spectrogram identifiers.\n\n   **Interpretation:** The average spectrogram identifier for GPD may indicate a specific pattern or set of characteristics in the spectrogram associated with this class.\n\n4. **seizure_vote:**\n   - Seizure has the highest average vote count, indicating a higher level of agreement among annotators for this class.\n   - Other classes have lower average vote counts.\n\n   **Interpretation:** Seizure tends to have a higher level of agreement among annotators, suggesting that it might be a more easily identifiable brain activity class compared to others.\n\nThese interpretations provide a high-level understanding of the average characteristics associated with different brain activity classes based on the provided features. It's important to note that further in-depth analysis and domain expertise may be required to draw more specific conclusions about the significance of these averages in the context of EEG signal processing and brain activity classification.","metadata":{"id":"63af5038"}},{"cell_type":"code","source":"df['Brain Activity'] = df[votes_columns].idxmax(axis=1).apply(lambda x: x.replace('_vote', ''))","metadata":{"id":"0848ac82","execution":{"iopub.status.busy":"2024-10-19T18:59:01.751902Z","iopub.execute_input":"2024-10-19T18:59:01.752177Z","iopub.status.idle":"2024-10-19T18:59:01.821126Z","shell.execute_reply.started":"2024-10-19T18:59:01.752146Z","shell.execute_reply":"2024-10-19T18:59:01.820453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"id":"92ae5e6a","outputId":"1d5e4935-ebc9-453a-840d-992c71e2c638","execution":{"iopub.status.busy":"2024-10-19T18:59:01.822225Z","iopub.execute_input":"2024-10-19T18:59:01.822573Z","iopub.status.idle":"2024-10-19T18:59:01.841577Z","shell.execute_reply.started":"2024-10-19T18:59:01.822530Z","shell.execute_reply":"2024-10-19T18:59:01.840782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grouped_continuous = df.groupby('eeg_id')[continuous].mean().reset_index(drop=True)\ngrouped_continuous","metadata":{"id":"664598db","outputId":"c5c884f4-d62f-44df-c665-194756e8b155","execution":{"iopub.status.busy":"2024-10-19T18:59:01.842634Z","iopub.execute_input":"2024-10-19T18:59:01.842969Z","iopub.status.idle":"2024-10-19T18:59:01.892336Z","shell.execute_reply.started":"2024-10-19T18:59:01.842938Z","shell.execute_reply":"2024-10-19T18:59:01.891400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']","metadata":{"id":"bb435de0","execution":{"iopub.status.busy":"2024-10-19T18:59:01.893593Z","iopub.execute_input":"2024-10-19T18:59:01.893920Z","iopub.status.idle":"2024-10-19T18:59:01.898117Z","shell.execute_reply.started":"2024-10-19T18:59:01.893887Z","shell.execute_reply":"2024-10-19T18:59:01.897234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"total_votes_per_pat = df.groupby('patient_id')[targets].sum().sum(axis=1)\nnormalized_votes = df.groupby('patient_id')[targets].sum().div(total_votes_per_pat, axis=0)\nmean_vote_ratio = normalized_votes.mean()\nprint( mean_vote_ratio )","metadata":{"id":"b0d882f9","outputId":"9a7dd695-7240-4202-a0b5-9724612ae1b0","execution":{"iopub.status.busy":"2024-10-19T18:59:01.899348Z","iopub.execute_input":"2024-10-19T18:59:01.899689Z","iopub.status.idle":"2024-10-19T18:59:01.920057Z","shell.execute_reply.started":"2024-10-19T18:59:01.899646Z","shell.execute_reply":"2024-10-19T18:59:01.918912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_eeg_id = 1000913311\ndf_filtered = df[df['eeg_id'] == target_eeg_id]\ndf_filtered","metadata":{"id":"60ea4718","outputId":"5b8b6e0a-3c15-4f41-b417-33e8b7b97653","execution":{"iopub.status.busy":"2024-10-19T18:59:01.921173Z","iopub.execute_input":"2024-10-19T18:59:01.921476Z","iopub.status.idle":"2024-10-19T18:59:01.938304Z","shell.execute_reply.started":"2024-10-19T18:59:01.921442Z","shell.execute_reply":"2024-10-19T18:59:01.937429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1000086677.parquet","metadata":{}},{"cell_type":"code","source":"df_filtered = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/1000913311.parquet')","metadata":{"id":"b9cf22d8","execution":{"iopub.status.busy":"2024-10-19T18:59:01.939640Z","iopub.execute_input":"2024-10-19T18:59:01.940611Z","iopub.status.idle":"2024-10-19T18:59:01.952905Z","shell.execute_reply.started":"2024-10-19T18:59:01.940564Z","shell.execute_reply":"2024-10-19T18:59:01.951667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_filtered","metadata":{"id":"bba88c2d","outputId":"321f66ac-a5f5-4b46-a328-8576ba2dd679","execution":{"iopub.status.busy":"2024-10-19T18:59:01.954329Z","iopub.execute_input":"2024-10-19T18:59:01.954742Z","iopub.status.idle":"2024-10-19T18:59:01.984515Z","shell.execute_reply.started":"2024-10-19T18:59:01.954681Z","shell.execute_reply":"2024-10-19T18:59:01.983679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(20, figsize=(10, 100))\n\nfor i, column in enumerate(df_filtered.columns):\n    ax[i].plot(df_filtered.index, df_filtered[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"id":"cfc53fde","outputId":"e6c4a3db-fd83-46c4-cc36-0ebde449f60c","execution":{"iopub.status.busy":"2024-10-19T18:59:01.985765Z","iopub.execute_input":"2024-10-19T18:59:01.986156Z","iopub.status.idle":"2024-10-19T18:59:08.277783Z","shell.execute_reply.started":"2024-10-19T18:59:01.986113Z","shell.execute_reply":"2024-10-19T18:59:08.276783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Data has a lot of peaks and troughs, indicating a high variation\nThe scale of most of the plots ranging from -200 to 100, excluding EKG chart\nThe line seems to fluctuate above and below a central value, which seems to be around zero value","metadata":{"id":"2941807e"}},{"cell_type":"code","source":"df_filtered.shape","metadata":{"id":"f3a46e3b","outputId":"c4c7097c-4eec-4e00-e8c7-73491fdaa36e","execution":{"iopub.status.busy":"2024-10-19T18:59:08.279046Z","iopub.execute_input":"2024-10-19T18:59:08.279367Z","iopub.status.idle":"2024-10-19T18:59:08.285168Z","shell.execute_reply.started":"2024-10-19T18:59:08.279332Z","shell.execute_reply":"2024-10-19T18:59:08.284252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_filtered.columns","metadata":{"id":"7263fcdc","outputId":"f4a9b66d-a092-4c8a-f328-d326e0deaf52","execution":{"iopub.status.busy":"2024-10-19T18:59:08.286579Z","iopub.execute_input":"2024-10-19T18:59:08.287308Z","iopub.status.idle":"2024-10-19T18:59:08.297388Z","shell.execute_reply.started":"2024-10-19T18:59:08.287254Z","shell.execute_reply":"2024-10-19T18:59:08.296515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The list of labels you provided appears to represent different electrode locations in an electroencephalography (EEG) recording setup. Each label corresponds to a specific position on the scalp where electrodes are placed to capture electrical activity from different regions of the brain. Let's break down each label:\n\n1. **Fp1 and Fp2:**\n   - These are frontal polar electrodes.\n   - Positioned on the forehead, they capture electrical activity from the frontal lobes, which are associated with cognitive functions, decision-making, and personality.\n\n2. **F3, F4, F7, and F8:**\n   - These electrodes are located in the frontal region.\n   - They capture electrical signals from the frontal lobes, contributing to the analysis of cognitive processes and emotional states.\n\n3. **C3 and C4:**\n   - These are central electrodes.\n   - Positioned over the central region of the scalp, they record activity from the central sulcus area, involved in motor control and sensory processing.\n\n4. **P3 and P4:**\n   - These electrodes are located in the parietal region.\n   - They capture signals related to sensory processing, spatial awareness, and interpretation of visual information.\n\n5. **Fz, Cz, and Pz:**\n   - These are midline electrodes located at the vertex (top) of the head.\n   - Fz captures activity from the frontal midline, Cz from the central midline, and Pz from the parietal midline. These are important for analyzing midline brain functions.\n\n6. **T3, T4, T5, and T6:**\n   - These electrodes are positioned in the temporal region.\n   - They record electrical activity from the temporal lobes, which are involved in auditory processing, memory, and language.\n\n7. **O1 and O2:**\n   - These electrodes are located in the occipital region.\n   - They capture signals from the occipital lobes, primarily responsible for visual processing.\n\n8. **EKG:**\n   - This electrode records electrocardiogram (EKG or ECG) signals from the heart.\n   - While not directly related to brain activity, it is often included in EEG setups to monitor cardiac activity, allowing for correlation with potential physiological effects on the EEG.\n\nThese electrode locations collectively provide a comprehensive view of brain activity, enabling the analysis of different cognitive and sensory functions. EEG recordings from these electrodes help researchers and clinicians understand brain dynamics and detect abnormalities in various neurological conditions.","metadata":{"id":"da039373"}},{"cell_type":"code","source":"df1 = df_filtered.copy()","metadata":{"id":"c16102a6","execution":{"iopub.status.busy":"2024-10-19T18:59:08.298458Z","iopub.execute_input":"2024-10-19T18:59:08.298786Z","iopub.status.idle":"2024-10-19T18:59:08.305652Z","shell.execute_reply.started":"2024-10-19T18:59:08.298749Z","shell.execute_reply":"2024-10-19T18:59:08.304777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1['Brain Activity'] = 'other'","metadata":{"id":"fcca1d68","execution":{"iopub.status.busy":"2024-10-19T18:59:08.306708Z","iopub.execute_input":"2024-10-19T18:59:08.307008Z","iopub.status.idle":"2024-10-19T18:59:08.314500Z","shell.execute_reply.started":"2024-10-19T18:59:08.306978Z","shell.execute_reply":"2024-10-19T18:59:08.313623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1","metadata":{"id":"61f24a03","outputId":"adf4dae4-2e5d-4cb1-9db9-bd2454a57093","execution":{"iopub.status.busy":"2024-10-19T18:59:08.315591Z","iopub.execute_input":"2024-10-19T18:59:08.315881Z","iopub.status.idle":"2024-10-19T18:59:08.353251Z","shell.execute_reply.started":"2024-10-19T18:59:08.315851Z","shell.execute_reply":"2024-10-19T18:59:08.352343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"electrodes = ['Fp1', 'F3', 'C3', 'P3', 'F7', 'T3', 'T5', 'O1', 'Fz', 'Cz', 'Pz', 'Fp2', 'F4', 'C4', 'P4', 'F8', 'T4', 'T6', 'O2', 'EKG']","metadata":{"id":"7054d7b8","execution":{"iopub.status.busy":"2024-10-19T18:59:08.354161Z","iopub.execute_input":"2024-10-19T18:59:08.354406Z","iopub.status.idle":"2024-10-19T18:59:08.359016Z","shell.execute_reply.started":"2024-10-19T18:59:08.354378Z","shell.execute_reply":"2024-10-19T18:59:08.358195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15, 8))\nsns.boxplot(data=df1[electrodes])\nplt.xlabel('Electrode')\nplt.ylabel('EEG Activity')\nplt.show()","metadata":{"id":"f41f6694","outputId":"b90be219-c83d-4713-dfc2-1e4bdf852be2","execution":{"iopub.status.busy":"2024-10-19T18:59:08.360208Z","iopub.execute_input":"2024-10-19T18:59:08.360477Z","iopub.status.idle":"2024-10-19T18:59:08.766066Z","shell.execute_reply.started":"2024-10-19T18:59:08.360446Z","shell.execute_reply":"2024-10-19T18:59:08.765160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15, 8))\nsns.boxenplot(data=df1[electrodes])\nplt.xlabel('Electrode')\nplt.ylabel('EEG Activity')\nplt.show()","metadata":{"id":"00842e94","outputId":"4b9ad3ef-144b-4fe5-b857-88031068ad11","execution":{"iopub.status.busy":"2024-10-19T18:59:08.767481Z","iopub.execute_input":"2024-10-19T18:59:08.767873Z","iopub.status.idle":"2024-10-19T18:59:09.317098Z","shell.execute_reply.started":"2024-10-19T18:59:08.767829Z","shell.execute_reply":"2024-10-19T18:59:09.316208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns","metadata":{"id":"b7743fe7","outputId":"f078f509-54f9-4c34-be10-fbcdf447d837","execution":{"iopub.status.busy":"2024-10-19T18:59:09.318784Z","iopub.execute_input":"2024-10-19T18:59:09.319146Z","iopub.status.idle":"2024-10-19T18:59:09.325686Z","shell.execute_reply.started":"2024-10-19T18:59:09.319102Z","shell.execute_reply":"2024-10-19T18:59:09.324801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_eeg_ids = df.drop_duplicates(subset='Brain Activity')['eeg_id'].tolist()\nprint(unique_eeg_ids)","metadata":{"id":"408eb9f9","outputId":"282f7989-027c-4b5e-c0ec-62d24c4c23bf","execution":{"iopub.status.busy":"2024-10-19T18:59:09.326937Z","iopub.execute_input":"2024-10-19T18:59:09.327245Z","iopub.status.idle":"2024-10-19T18:59:09.344401Z","shell.execute_reply.started":"2024-10-19T18:59:09.327207Z","shell.execute_reply":"2024-10-19T18:59:09.343410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"eeg_activity_dict = {}\n\nfor eeg_id in unique_eeg_ids:\n    subset_df = df[df['eeg_id'] == eeg_id]\n\n    unique_brain_activity = subset_df['Brain Activity'].unique()\n\n    if len(unique_brain_activity) == 1:\n        eeg_activity_dict[eeg_id] = unique_brain_activity[0]\n\nprint(eeg_activity_dict)","metadata":{"id":"086559d2","outputId":"0f3f1d28-a8f8-4875-8acd-b61fe9f0429c","execution":{"iopub.status.busy":"2024-10-19T18:59:09.345406Z","iopub.execute_input":"2024-10-19T18:59:09.345661Z","iopub.status.idle":"2024-10-19T18:59:09.357319Z","shell.execute_reply.started":"2024-10-19T18:59:09.345632Z","shell.execute_reply":"2024-10-19T18:59:09.356362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. `eeg_activity_dict = {}`: Initializes an empty dictionary to store EEG IDs and their corresponding brain activities.\n\n2. `for eeg_id in unique_eeg_ids:`: Iterates through each unique EEG ID.\n\n3. `subset_df = df[df['eeg_id'] == eeg_id]`: Creates a subset DataFrame (`subset_df`) containing only rows where the 'eeg_id' column matches the current `eeg_id`.\n\n4. `unique_brain_activity = subset_df['Brain Activity'].unique()`: Extracts unique values from the 'Brain Activity' column within the subset DataFrame.\n\n5. `if len(unique_brain_activity) == 1:`: Checks if there is only one unique brain activity in the subset.\n\n6. `eeg_activity_dict[eeg_id] = unique_brain_activity[0]`: If there is only one unique brain activity, it adds an entry to the dictionary with the EEG ID as the key and the unique brain activity as the value.\n\n7. Finally, the dictionary `eeg_activity_dict` contains mappings from unique EEG IDs to their corresponding unique brain activities.","metadata":{"id":"69d43134"}},{"cell_type":"code","source":" df[df['eeg_id'] == 1628180742]","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:09.358670Z","iopub.execute_input":"2024-10-19T18:59:09.359291Z","iopub.status.idle":"2024-10-19T18:59:09.380963Z","shell.execute_reply.started":"2024-10-19T18:59:09.359247Z","shell.execute_reply":"2024-10-19T18:59:09.380142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1 = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/1628180742.parquet')","metadata":{"id":"f4b7ba21","execution":{"iopub.status.busy":"2024-10-19T18:59:09.382236Z","iopub.execute_input":"2024-10-19T18:59:09.382870Z","iopub.status.idle":"2024-10-19T18:59:09.395854Z","shell.execute_reply.started":"2024-10-19T18:59:09.382824Z","shell.execute_reply":"2024-10-19T18:59:09.395184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:09.396812Z","iopub.execute_input":"2024-10-19T18:59:09.397069Z","iopub.status.idle":"2024-10-19T18:59:09.417325Z","shell.execute_reply.started":"2024-10-19T18:59:09.397039Z","shell.execute_reply":"2024-10-19T18:59:09.416339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2 = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/736446371.parquet')","metadata":{"id":"69c9e9c3","execution":{"iopub.status.busy":"2024-10-19T18:59:09.418526Z","iopub.execute_input":"2024-10-19T18:59:09.418879Z","iopub.status.idle":"2024-10-19T18:59:09.431564Z","shell.execute_reply.started":"2024-10-19T18:59:09.418838Z","shell.execute_reply":"2024-10-19T18:59:09.430570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df3 = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/722738444.parquet')","metadata":{"id":"fb78275e","execution":{"iopub.status.busy":"2024-10-19T18:59:09.432791Z","iopub.execute_input":"2024-10-19T18:59:09.433284Z","iopub.status.idle":"2024-10-19T18:59:09.445127Z","shell.execute_reply.started":"2024-10-19T18:59:09.433240Z","shell.execute_reply":"2024-10-19T18:59:09.444110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df4 = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/1202099836.parquet')","metadata":{"id":"ca0ddabe","execution":{"iopub.status.busy":"2024-10-19T18:59:09.446184Z","iopub.execute_input":"2024-10-19T18:59:09.446507Z","iopub.status.idle":"2024-10-19T18:59:09.456196Z","shell.execute_reply.started":"2024-10-19T18:59:09.446476Z","shell.execute_reply":"2024-10-19T18:59:09.455286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df5 = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/2578018731.parquet')","metadata":{"id":"9e26b139","execution":{"iopub.status.busy":"2024-10-19T18:59:09.457408Z","iopub.execute_input":"2024-10-19T18:59:09.458270Z","iopub.status.idle":"2024-10-19T18:59:09.467951Z","shell.execute_reply.started":"2024-10-19T18:59:09.458237Z","shell.execute_reply":"2024-10-19T18:59:09.467032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df6 = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/2277392603.parquet')","metadata":{"id":"856d5019","execution":{"iopub.status.busy":"2024-10-19T18:59:09.469113Z","iopub.execute_input":"2024-10-19T18:59:09.472342Z","iopub.status.idle":"2024-10-19T18:59:09.480547Z","shell.execute_reply.started":"2024-10-19T18:59:09.472310Z","shell.execute_reply":"2024-10-19T18:59:09.479683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1['Brain Activity'] = 'seizure'","metadata":{"id":"e2294b63","execution":{"iopub.status.busy":"2024-10-19T18:59:09.481525Z","iopub.execute_input":"2024-10-19T18:59:09.482289Z","iopub.status.idle":"2024-10-19T18:59:09.487078Z","shell.execute_reply.started":"2024-10-19T18:59:09.482257Z","shell.execute_reply":"2024-10-19T18:59:09.486020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:09.502633Z","iopub.execute_input":"2024-10-19T18:59:09.502912Z","iopub.status.idle":"2024-10-19T18:59:09.524432Z","shell.execute_reply.started":"2024-10-19T18:59:09.502883Z","shell.execute_reply":"2024-10-19T18:59:09.523553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2['Brain Activity'] = 'lpd'","metadata":{"id":"5d718c16","execution":{"iopub.status.busy":"2024-10-19T18:59:09.525763Z","iopub.execute_input":"2024-10-19T18:59:09.526110Z","iopub.status.idle":"2024-10-19T18:59:09.532747Z","shell.execute_reply.started":"2024-10-19T18:59:09.526078Z","shell.execute_reply":"2024-10-19T18:59:09.531780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df3['Brain Activity'] = 'lrda'","metadata":{"id":"1c986241","execution":{"iopub.status.busy":"2024-10-19T18:59:09.533964Z","iopub.execute_input":"2024-10-19T18:59:09.534260Z","iopub.status.idle":"2024-10-19T18:59:09.543339Z","shell.execute_reply.started":"2024-10-19T18:59:09.534216Z","shell.execute_reply":"2024-10-19T18:59:09.542348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df4['Brain Activity'] = 'other'","metadata":{"id":"cdad1af0","execution":{"iopub.status.busy":"2024-10-19T18:59:09.544343Z","iopub.execute_input":"2024-10-19T18:59:09.544612Z","iopub.status.idle":"2024-10-19T18:59:09.554796Z","shell.execute_reply.started":"2024-10-19T18:59:09.544582Z","shell.execute_reply":"2024-10-19T18:59:09.553703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df5['Brain Activity'] = 'grda'","metadata":{"id":"7b577241","execution":{"iopub.status.busy":"2024-10-19T18:59:09.555988Z","iopub.execute_input":"2024-10-19T18:59:09.556351Z","iopub.status.idle":"2024-10-19T18:59:09.563316Z","shell.execute_reply.started":"2024-10-19T18:59:09.556304Z","shell.execute_reply":"2024-10-19T18:59:09.562296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df6['Brain Activity'] = 'gpd'","metadata":{"id":"7f07eae6","execution":{"iopub.status.busy":"2024-10-19T18:59:09.564820Z","iopub.execute_input":"2024-10-19T18:59:09.565109Z","iopub.status.idle":"2024-10-19T18:59:09.571866Z","shell.execute_reply.started":"2024-10-19T18:59:09.565066Z","shell.execute_reply":"2024-10-19T18:59:09.570789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1","metadata":{"id":"1b0cf6f7","outputId":"9c2d1f82-52df-4da6-b667-ef91aa04993f","execution":{"iopub.status.busy":"2024-10-19T18:59:09.573048Z","iopub.execute_input":"2024-10-19T18:59:09.573429Z","iopub.status.idle":"2024-10-19T18:59:09.604550Z","shell.execute_reply.started":"2024-10-19T18:59:09.573397Z","shell.execute_reply":"2024-10-19T18:59:09.603685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2","metadata":{"id":"81b74305","outputId":"4bfa1f4a-e46b-42e4-f866-c7f999a4b269","execution":{"iopub.status.busy":"2024-10-19T18:59:09.605609Z","iopub.execute_input":"2024-10-19T18:59:09.605904Z","iopub.status.idle":"2024-10-19T18:59:09.635599Z","shell.execute_reply.started":"2024-10-19T18:59:09.605873Z","shell.execute_reply":"2024-10-19T18:59:09.634748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df3","metadata":{"id":"e3aacf93","outputId":"6d6ee9b4-9d72-430a-9ef2-184393918c48","execution":{"iopub.status.busy":"2024-10-19T18:59:09.636683Z","iopub.execute_input":"2024-10-19T18:59:09.636990Z","iopub.status.idle":"2024-10-19T18:59:09.665483Z","shell.execute_reply.started":"2024-10-19T18:59:09.636960Z","shell.execute_reply":"2024-10-19T18:59:09.664652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df4","metadata":{"id":"45f3c0a3","outputId":"1d2db6cf-5336-4aa8-bca8-ee570a7eee2a","execution":{"iopub.status.busy":"2024-10-19T18:59:09.666678Z","iopub.execute_input":"2024-10-19T18:59:09.667080Z","iopub.status.idle":"2024-10-19T18:59:09.708660Z","shell.execute_reply.started":"2024-10-19T18:59:09.667036Z","shell.execute_reply":"2024-10-19T18:59:09.707538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df5","metadata":{"id":"a049602a","outputId":"c52642b9-b0d0-4d62-b77e-9bb44e16946a","execution":{"iopub.status.busy":"2024-10-19T18:59:09.710610Z","iopub.execute_input":"2024-10-19T18:59:09.711027Z","iopub.status.idle":"2024-10-19T18:59:09.750090Z","shell.execute_reply.started":"2024-10-19T18:59:09.710980Z","shell.execute_reply":"2024-10-19T18:59:09.749237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df6","metadata":{"id":"864b0cd8","outputId":"216af17d-fb35-4914-924a-64f42f2dfa50","execution":{"iopub.status.busy":"2024-10-19T18:59:09.751434Z","iopub.execute_input":"2024-10-19T18:59:09.752039Z","iopub.status.idle":"2024-10-19T18:59:09.783010Z","shell.execute_reply.started":"2024-10-19T18:59:09.752007Z","shell.execute_reply":"2024-10-19T18:59:09.782177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.shape","metadata":{"id":"25acce66","outputId":"7566c175-bb48-46ae-bdff-1ca8e066c9ee","execution":{"iopub.status.busy":"2024-10-19T18:59:09.784136Z","iopub.execute_input":"2024-10-19T18:59:09.784410Z","iopub.status.idle":"2024-10-19T18:59:09.790538Z","shell.execute_reply.started":"2024-10-19T18:59:09.784379Z","shell.execute_reply":"2024-10-19T18:59:09.789551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2.shape","metadata":{"id":"1beb682e","outputId":"5d29ec3b-ef54-4b4c-e9ce-9800ce74adc4","execution":{"iopub.status.busy":"2024-10-19T18:59:09.791895Z","iopub.execute_input":"2024-10-19T18:59:09.792597Z","iopub.status.idle":"2024-10-19T18:59:09.801544Z","shell.execute_reply.started":"2024-10-19T18:59:09.792558Z","shell.execute_reply":"2024-10-19T18:59:09.800608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df3.shape","metadata":{"id":"c1209c59","outputId":"1f416845-8269-4bd2-a4bf-d05a8a1903ad","execution":{"iopub.status.busy":"2024-10-19T18:59:09.802814Z","iopub.execute_input":"2024-10-19T18:59:09.803084Z","iopub.status.idle":"2024-10-19T18:59:09.809633Z","shell.execute_reply.started":"2024-10-19T18:59:09.803055Z","shell.execute_reply":"2024-10-19T18:59:09.808790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df4.shape","metadata":{"id":"82d871da","outputId":"b8f1983e-2930-4c35-b442-a058fefc3a2f","execution":{"iopub.status.busy":"2024-10-19T18:59:09.811124Z","iopub.execute_input":"2024-10-19T18:59:09.811397Z","iopub.status.idle":"2024-10-19T18:59:09.818653Z","shell.execute_reply.started":"2024-10-19T18:59:09.811367Z","shell.execute_reply":"2024-10-19T18:59:09.817782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df5.shape","metadata":{"id":"9e042aa4","outputId":"e4a84087-638d-4014-d2c1-1923278f5893","execution":{"iopub.status.busy":"2024-10-19T18:59:09.819913Z","iopub.execute_input":"2024-10-19T18:59:09.820182Z","iopub.status.idle":"2024-10-19T18:59:09.827545Z","shell.execute_reply.started":"2024-10-19T18:59:09.820153Z","shell.execute_reply":"2024-10-19T18:59:09.826588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df6.shape","metadata":{"id":"d1d81661","outputId":"7ab7dd20-a998-4478-b73f-63ffc7b69d0d","execution":{"iopub.status.busy":"2024-10-19T18:59:09.828768Z","iopub.execute_input":"2024-10-19T18:59:09.829108Z","iopub.status.idle":"2024-10-19T18:59:09.837642Z","shell.execute_reply.started":"2024-10-19T18:59:09.829071Z","shell.execute_reply":"2024-10-19T18:59:09.836775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.columns","metadata":{"id":"22bef282","outputId":"3a7ed9c0-7e4b-4d8c-ff40-acffd3acb523","execution":{"iopub.status.busy":"2024-10-19T18:59:09.838944Z","iopub.execute_input":"2024-10-19T18:59:09.839287Z","iopub.status.idle":"2024-10-19T18:59:09.847548Z","shell.execute_reply.started":"2024-10-19T18:59:09.839255Z","shell.execute_reply":"2024-10-19T18:59:09.846529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2.columns","metadata":{"id":"d3f2dc70","outputId":"187de739-db81-4d7a-acdb-682f546af214","execution":{"iopub.status.busy":"2024-10-19T18:59:09.848677Z","iopub.execute_input":"2024-10-19T18:59:09.848984Z","iopub.status.idle":"2024-10-19T18:59:09.856621Z","shell.execute_reply.started":"2024-10-19T18:59:09.848954Z","shell.execute_reply":"2024-10-19T18:59:09.855746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df3.columns","metadata":{"id":"fc3344c6","outputId":"97d373fe-8fe6-45e3-c8ce-01a34e163bd8","execution":{"iopub.status.busy":"2024-10-19T18:59:09.857890Z","iopub.execute_input":"2024-10-19T18:59:09.858166Z","iopub.status.idle":"2024-10-19T18:59:09.866344Z","shell.execute_reply.started":"2024-10-19T18:59:09.858137Z","shell.execute_reply":"2024-10-19T18:59:09.865397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df4.columns","metadata":{"id":"01f73849","outputId":"9913a6af-42f7-463e-ca73-7632149f0157","execution":{"iopub.status.busy":"2024-10-19T18:59:09.867871Z","iopub.execute_input":"2024-10-19T18:59:09.868168Z","iopub.status.idle":"2024-10-19T18:59:09.876445Z","shell.execute_reply.started":"2024-10-19T18:59:09.868129Z","shell.execute_reply":"2024-10-19T18:59:09.875534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df5.columns","metadata":{"id":"33826c8e","outputId":"415f8d47-c995-47cf-ac7d-f4119b545cde","execution":{"iopub.status.busy":"2024-10-19T18:59:09.878111Z","iopub.execute_input":"2024-10-19T18:59:09.878451Z","iopub.status.idle":"2024-10-19T18:59:09.897395Z","shell.execute_reply.started":"2024-10-19T18:59:09.878404Z","shell.execute_reply":"2024-10-19T18:59:09.896336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df6.columns","metadata":{"id":"fbf3c75b","outputId":"f9eb2ef4-575d-400a-d506-65a4cef56078","execution":{"iopub.status.busy":"2024-10-19T18:59:09.898590Z","iopub.execute_input":"2024-10-19T18:59:09.898928Z","iopub.status.idle":"2024-10-19T18:59:09.910345Z","shell.execute_reply.started":"2024-10-19T18:59:09.898897Z","shell.execute_reply":"2024-10-19T18:59:09.909409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(21, figsize=(10, 100))\n\nfor i, column in enumerate(df1.columns):\n    ax[i].plot(df1.index, df1[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"id":"f117655e","outputId":"b83d745f-7326-44b1-b127-46c2aa2922fe","execution":{"iopub.status.busy":"2024-10-19T18:59:09.911538Z","iopub.execute_input":"2024-10-19T18:59:09.911830Z","iopub.status.idle":"2024-10-19T18:59:14.714564Z","shell.execute_reply.started":"2024-10-19T18:59:09.911800Z","shell.execute_reply":"2024-10-19T18:59:14.713407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(21, figsize=(10, 100))\n\nfor i, column in enumerate(df2.columns):\n    ax[i].plot(df2.index, df2[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"id":"cbaaeaea","outputId":"6f10cf44-9fc1-4e0b-cb51-2ab5587b6948","execution":{"iopub.status.busy":"2024-10-19T18:59:14.715940Z","iopub.execute_input":"2024-10-19T18:59:14.716232Z","iopub.status.idle":"2024-10-19T18:59:19.381704Z","shell.execute_reply.started":"2024-10-19T18:59:14.716200Z","shell.execute_reply":"2024-10-19T18:59:19.380782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(21, figsize=(10, 100))\n\nfor i, column in enumerate(df3.columns):\n    ax[i].plot(df3.index, df3[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"id":"0ac85785","outputId":"a4d320f1-0e65-416a-b629-0a86235290b8","execution":{"iopub.status.busy":"2024-10-19T18:59:19.382858Z","iopub.execute_input":"2024-10-19T18:59:19.383144Z","iopub.status.idle":"2024-10-19T18:59:23.797570Z","shell.execute_reply.started":"2024-10-19T18:59:19.383113Z","shell.execute_reply":"2024-10-19T18:59:23.796334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(21, figsize=(10, 100))\n\nfor i, column in enumerate(df4.columns):\n    ax[i].plot(df4.index, df4[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"id":"baa21178","outputId":"b9e30126-d028-467e-878e-d206153b4275","execution":{"iopub.status.busy":"2024-10-19T18:59:23.798809Z","iopub.execute_input":"2024-10-19T18:59:23.799769Z","iopub.status.idle":"2024-10-19T18:59:28.205112Z","shell.execute_reply.started":"2024-10-19T18:59:23.799704Z","shell.execute_reply":"2024-10-19T18:59:28.204208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(21, figsize=(10, 100))\n\nfor i, column in enumerate(df5.columns):\n    ax[i].plot(df5.index, df5[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"id":"f96ec0e2","outputId":"66cdf7cf-ca88-4bc8-8b33-b3d569c898f0","execution":{"iopub.status.busy":"2024-10-19T18:59:28.206524Z","iopub.execute_input":"2024-10-19T18:59:28.206897Z","iopub.status.idle":"2024-10-19T18:59:32.206004Z","shell.execute_reply.started":"2024-10-19T18:59:28.206857Z","shell.execute_reply":"2024-10-19T18:59:32.205104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(21, figsize=(10, 100))\n\nfor i, column in enumerate(df6.columns):\n    ax[i].plot(df6.index, df6[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"id":"510a358a","outputId":"344d1a95-16a4-4b05-ca89-be222a1b4aa8","execution":{"iopub.status.busy":"2024-10-19T18:59:32.207486Z","iopub.execute_input":"2024-10-19T18:59:32.208291Z","iopub.status.idle":"2024-10-19T18:59:36.778941Z","shell.execute_reply.started":"2024-10-19T18:59:32.208247Z","shell.execute_reply":"2024-10-19T18:59:36.777625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfs = [df1, df2, df3, df4, df5, df6]","metadata":{"id":"532afdf9","execution":{"iopub.status.busy":"2024-10-19T18:59:36.780225Z","iopub.execute_input":"2024-10-19T18:59:36.780568Z","iopub.status.idle":"2024-10-19T18:59:36.784994Z","shell.execute_reply.started":"2024-10-19T18:59:36.780530Z","shell.execute_reply":"2024-10-19T18:59:36.784058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df = pd.concat(dfs, ignore_index=True)","metadata":{"id":"a554c74d","execution":{"iopub.status.busy":"2024-10-19T18:59:36.786520Z","iopub.execute_input":"2024-10-19T18:59:36.786855Z","iopub.status.idle":"2024-10-19T18:59:36.797477Z","shell.execute_reply.started":"2024-10-19T18:59:36.786816Z","shell.execute_reply":"2024-10-19T18:59:36.796560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df","metadata":{"id":"0fd5d527","outputId":"8f5a12b4-35db-436d-e1a3-fcead81acf80","execution":{"iopub.status.busy":"2024-10-19T18:59:36.798604Z","iopub.execute_input":"2024-10-19T18:59:36.798904Z","iopub.status.idle":"2024-10-19T18:59:36.840762Z","shell.execute_reply.started":"2024-10-19T18:59:36.798867Z","shell.execute_reply":"2024-10-19T18:59:36.839867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.shape","metadata":{"id":"1adace43","outputId":"973674c5-2dd6-4394-c289-6548ba893609","execution":{"iopub.status.busy":"2024-10-19T18:59:36.841964Z","iopub.execute_input":"2024-10-19T18:59:36.842237Z","iopub.status.idle":"2024-10-19T18:59:36.848129Z","shell.execute_reply.started":"2024-10-19T18:59:36.842206Z","shell.execute_reply":"2024-10-19T18:59:36.847132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.columns","metadata":{"id":"6ddf044b","outputId":"9b535970-c95d-4770-9fd4-a4c4d2e12cdb","execution":{"iopub.status.busy":"2024-10-19T18:59:36.849432Z","iopub.execute_input":"2024-10-19T18:59:36.849806Z","iopub.status.idle":"2024-10-19T18:59:36.858025Z","shell.execute_reply.started":"2024-10-19T18:59:36.849763Z","shell.execute_reply":"2024-10-19T18:59:36.857117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.duplicated().sum()","metadata":{"id":"f40ad41e","outputId":"972639e6-0b73-42c7-c73a-08d5037a5f1d","execution":{"iopub.status.busy":"2024-10-19T18:59:36.859237Z","iopub.execute_input":"2024-10-19T18:59:36.859512Z","iopub.status.idle":"2024-10-19T18:59:36.922605Z","shell.execute_reply.started":"2024-10-19T18:59:36.859481Z","shell.execute_reply":"2024-10-19T18:59:36.921761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new = combined_df.copy()","metadata":{"id":"257c0c54","execution":{"iopub.status.busy":"2024-10-19T18:59:36.923768Z","iopub.execute_input":"2024-10-19T18:59:36.924141Z","iopub.status.idle":"2024-10-19T18:59:36.929967Z","shell.execute_reply.started":"2024-10-19T18:59:36.924098Z","shell.execute_reply":"2024-10-19T18:59:36.929028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new.set_index('Brain Activity', inplace=True)","metadata":{"id":"4e3de61f","execution":{"iopub.status.busy":"2024-10-19T18:59:36.930977Z","iopub.execute_input":"2024-10-19T18:59:36.931268Z","iopub.status.idle":"2024-10-19T18:59:36.938629Z","shell.execute_reply.started":"2024-10-19T18:59:36.931237Z","shell.execute_reply":"2024-10-19T18:59:36.937776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new","metadata":{"id":"3a564285","outputId":"5a2722d6-5d8a-4ef9-cd15-41366160444f","execution":{"iopub.status.busy":"2024-10-19T18:59:36.939763Z","iopub.execute_input":"2024-10-19T18:59:36.940064Z","iopub.status.idle":"2024-10-19T18:59:36.975528Z","shell.execute_reply.started":"2024-10-19T18:59:36.940032Z","shell.execute_reply":"2024-10-19T18:59:36.974747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new.duplicated().sum()","metadata":{"id":"1011d9e1","outputId":"ebbacf4f-76dc-4323-fee2-97de4b5a3492","execution":{"iopub.status.busy":"2024-10-19T18:59:36.976615Z","iopub.execute_input":"2024-10-19T18:59:36.976893Z","iopub.status.idle":"2024-10-19T18:59:37.029056Z","shell.execute_reply.started":"2024-10-19T18:59:36.976863Z","shell.execute_reply":"2024-10-19T18:59:37.028179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.isnull().sum()","metadata":{"id":"e8b7d9c2","outputId":"01ec23d4-ba53-42a0-a611-3ed40892421f","execution":{"iopub.status.busy":"2024-10-19T18:59:37.030121Z","iopub.execute_input":"2024-10-19T18:59:37.030493Z","iopub.status.idle":"2024-10-19T18:59:37.047566Z","shell.execute_reply.started":"2024-10-19T18:59:37.030461Z","shell.execute_reply":"2024-10-19T18:59:37.046769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.info()","metadata":{"id":"fee50dce","outputId":"681de9a2-74e6-4642-d62c-effed8368453","execution":{"iopub.status.busy":"2024-10-19T18:59:37.048698Z","iopub.execute_input":"2024-10-19T18:59:37.049053Z","iopub.status.idle":"2024-10-19T18:59:37.070088Z","shell.execute_reply.started":"2024-10-19T18:59:37.049012Z","shell.execute_reply":"2024-10-19T18:59:37.069251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.describe()","metadata":{"id":"f1f3eb18","outputId":"f9d4fa70-0fb3-44fa-a9fb-93a403fef840","execution":{"iopub.status.busy":"2024-10-19T18:59:37.071234Z","iopub.execute_input":"2024-10-19T18:59:37.071508Z","iopub.status.idle":"2024-10-19T18:59:37.162316Z","shell.execute_reply.started":"2024-10-19T18:59:37.071478Z","shell.execute_reply":"2024-10-19T18:59:37.161427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.nunique()","metadata":{"id":"1f652c7a","outputId":"6886e9dd-38ad-4be6-f64e-4c5865d7b29c","execution":{"iopub.status.busy":"2024-10-19T18:59:37.163387Z","iopub.execute_input":"2024-10-19T18:59:37.163667Z","iopub.status.idle":"2024-10-19T18:59:37.209311Z","shell.execute_reply.started":"2024-10-19T18:59:37.163636Z","shell.execute_reply":"2024-10-19T18:59:37.208423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def classify_features(df):\n    categorical_features = []\n    non_categorical_features = []\n    discrete_features = []\n    continuous_features = []\n\n    for column in df.columns:\n        if df[column].dtype in ['object', 'bool']:\n            if df[column].nunique() < 15:\n                categorical_features.append(column)\n            else:\n                non_categorical_features.append(column)\n        elif df[column].dtype in ['int64', 'float32']:\n            if df[column].nunique() < 10:\n                discrete_features.append(column)\n            else:\n                continuous_features.append(column)\n\n    return categorical_features, non_categorical_features, discrete_features, continuous_features","metadata":{"id":"eb682adc","execution":{"iopub.status.busy":"2024-10-19T18:59:37.210422Z","iopub.execute_input":"2024-10-19T18:59:37.210689Z","iopub.status.idle":"2024-10-19T18:59:37.217633Z","shell.execute_reply.started":"2024-10-19T18:59:37.210660Z","shell.execute_reply":"2024-10-19T18:59:37.216614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical, non_categorical, discrete, continuous = classify_features(combined_df)","metadata":{"id":"42b6f17e","execution":{"iopub.status.busy":"2024-10-19T18:59:37.218782Z","iopub.execute_input":"2024-10-19T18:59:37.219103Z","iopub.status.idle":"2024-10-19T18:59:37.263337Z","shell.execute_reply.started":"2024-10-19T18:59:37.219048Z","shell.execute_reply":"2024-10-19T18:59:37.262656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Categorical Features:\", categorical)\nprint(\"Non-Categorical Features:\", non_categorical)\nprint(\"Discrete Features:\", discrete)\nprint(\"Continuous Features:\", continuous)","metadata":{"id":"eaad5548","outputId":"43a71cc7-56dc-4e1f-e898-69c4b077b7a1","execution":{"iopub.status.busy":"2024-10-19T18:59:37.264217Z","iopub.execute_input":"2024-10-19T18:59:37.264481Z","iopub.status.idle":"2024-10-19T18:59:37.269649Z","shell.execute_reply.started":"2024-10-19T18:59:37.264452Z","shell.execute_reply":"2024-10-19T18:59:37.268661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in categorical:\n    plt.figure(figsize=(15, 6))\n    ax = sns.countplot(x=i, data=combined_df, palette='hls')\n    for p in ax.patches:\n        ax.annotate(f'{p.get_height()}', (p.get_x() + p.get_width() / 2., p.get_height()),\n                    ha='center', va='center', xytext=(0, 10), textcoords='offset points')\n    plt.show()","metadata":{"id":"76b7ee47","outputId":"9e9270a4-ba00-42d1-cd04-96641ca0be7e","execution":{"iopub.status.busy":"2024-10-19T18:59:37.270580Z","iopub.execute_input":"2024-10-19T18:59:37.270910Z","iopub.status.idle":"2024-10-19T18:59:37.618217Z","shell.execute_reply.started":"2024-10-19T18:59:37.270873Z","shell.execute_reply":"2024-10-19T18:59:37.617219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in categorical:\n    plt.figure(figsize=(20,10))\n    plt.pie(combined_df[i].value_counts(), labels=combined_df[i].value_counts().index, autopct='%1.1f%%', textprops={'fontsize': 15,\n                                           'color': 'black',\n                                           'weight': 'bold',\n                                           'family': 'serif' })\n    hfont = {'fontname':'serif', 'weight': 'bold'}\n    plt.title(i, size=20, **hfont)\n    plt.show()","metadata":{"id":"8b8666e1","outputId":"31df97a8-fd0b-4cde-ef67-04d1861bc37b","execution":{"iopub.status.busy":"2024-10-19T18:59:37.619539Z","iopub.execute_input":"2024-10-19T18:59:37.620333Z","iopub.status.idle":"2024-10-19T18:59:37.861267Z","shell.execute_reply.started":"2024-10-19T18:59:37.620288Z","shell.execute_reply":"2024-10-19T18:59:37.860356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## FEATURE ENGINEERING","metadata":{}},{"cell_type":"code","source":"numeric_df = combined_df.select_dtypes(include=[np.number])\ncorrelation_matrix = numeric_df.corr()\n\nprint(\"Correlation Matrix:\")\ncorrelation_matrix","metadata":{"id":"9709ee57","outputId":"b2ae4310-c386-41af-ee10-23cfe7f4fbfe","execution":{"iopub.status.busy":"2024-10-19T18:59:37.862361Z","iopub.execute_input":"2024-10-19T18:59:37.862644Z","iopub.status.idle":"2024-10-19T18:59:37.977653Z","shell.execute_reply.started":"2024-10-19T18:59:37.862612Z","shell.execute_reply":"2024-10-19T18:59:37.976739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15, 8))\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', fmt=\".2f\", linewidths=.5)\nplt.title('Correlation Heatmap')\nplt.show()","metadata":{"id":"9f1c74d9","outputId":"01698fd4-f927-4787-93a9-c3edf0578dc5","execution":{"iopub.status.busy":"2024-10-19T18:59:37.978862Z","iopub.execute_input":"2024-10-19T18:59:37.979223Z","iopub.status.idle":"2024-10-19T18:59:39.154293Z","shell.execute_reply.started":"2024-10-19T18:59:37.979165Z","shell.execute_reply":"2024-10-19T18:59:39.153392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.columns","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.155409Z","iopub.execute_input":"2024-10-19T18:59:39.155683Z","iopub.status.idle":"2024-10-19T18:59:39.161841Z","shell.execute_reply.started":"2024-10-19T18:59:39.155653Z","shell.execute_reply":"2024-10-19T18:59:39.160960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ensure you are using the correct column names\n# Example:\ncombined_df = combined_df[['C3', 'Cz', 'Fp2', 'O2', 'Brain Activity']]\n","metadata":{"id":"da8e28eb","execution":{"iopub.status.busy":"2024-10-19T18:59:39.162956Z","iopub.execute_input":"2024-10-19T18:59:39.163238Z","iopub.status.idle":"2024-10-19T18:59:39.173479Z","shell.execute_reply.started":"2024-10-19T18:59:39.163207Z","shell.execute_reply":"2024-10-19T18:59:39.172545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assuming 'corr_matrix' is the correlation matrix you've created\n# Extract correlation values for the 'EKG' column\nekg_corr = correlation_matrix['EKG']\n\n# Filter values less than -0.16\nfiltered_features = ekg_corr[ekg_corr < -0.16]\nprint(filtered_features.index)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.174830Z","iopub.execute_input":"2024-10-19T18:59:39.175195Z","iopub.status.idle":"2024-10-19T18:59:39.182934Z","shell.execute_reply.started":"2024-10-19T18:59:39.175150Z","shell.execute_reply":"2024-10-19T18:59:39.182070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"brain_activities = combined_df['Brain Activity'].unique()\nprint(brain_activities)","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.184150Z","iopub.execute_input":"2024-10-19T18:59:39.184515Z","iopub.status.idle":"2024-10-19T18:59:39.200427Z","shell.execute_reply.started":"2024-10-19T18:59:39.184474Z","shell.execute_reply":"2024-10-19T18:59:39.199474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" \nmapping = {activity: idx for idx, activity in enumerate(brain_activities)}\n\ncombined_df['Brain Activity Numerical'] = combined_df['Brain Activity'].map(mapping)","metadata":{"id":"c11642e7","execution":{"iopub.status.busy":"2024-10-19T18:59:39.201742Z","iopub.execute_input":"2024-10-19T18:59:39.202137Z","iopub.status.idle":"2024-10-19T18:59:39.214920Z","shell.execute_reply.started":"2024-10-19T18:59:39.202049Z","shell.execute_reply":"2024-10-19T18:59:39.213989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df","metadata":{"id":"459c4698","outputId":"0ced002f-6bad-4cd4-a617-4efe960c6d62","execution":{"iopub.status.busy":"2024-10-19T18:59:39.215911Z","iopub.execute_input":"2024-10-19T18:59:39.216227Z","iopub.status.idle":"2024-10-19T18:59:39.232120Z","shell.execute_reply.started":"2024-10-19T18:59:39.216196Z","shell.execute_reply":"2024-10-19T18:59:39.231213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### DEFINE X AND Y","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score, classification_report , confusion_matrix","metadata":{"id":"b5ca0392","execution":{"iopub.status.busy":"2024-10-19T18:59:39.233418Z","iopub.execute_input":"2024-10-19T18:59:39.233708Z","iopub.status.idle":"2024-10-19T18:59:39.242439Z","shell.execute_reply.started":"2024-10-19T18:59:39.233677Z","shell.execute_reply":"2024-10-19T18:59:39.241550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = combined_df.drop(['Brain Activity Numerical','Brain Activity'], axis = 1)","metadata":{"id":"35cc5f6d","execution":{"iopub.status.busy":"2024-10-19T18:59:39.243535Z","iopub.execute_input":"2024-10-19T18:59:39.244273Z","iopub.status.idle":"2024-10-19T18:59:39.253151Z","shell.execute_reply.started":"2024-10-19T18:59:39.244239Z","shell.execute_reply":"2024-10-19T18:59:39.252385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.254165Z","iopub.execute_input":"2024-10-19T18:59:39.254535Z","iopub.status.idle":"2024-10-19T18:59:39.267086Z","shell.execute_reply.started":"2024-10-19T18:59:39.254493Z","shell.execute_reply":"2024-10-19T18:59:39.266072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = combined_df['Brain Activity Numerical']","metadata":{"id":"7f42ab02","execution":{"iopub.status.busy":"2024-10-19T18:59:39.268254Z","iopub.execute_input":"2024-10-19T18:59:39.268563Z","iopub.status.idle":"2024-10-19T18:59:39.277097Z","shell.execute_reply.started":"2024-10-19T18:59:39.268519Z","shell.execute_reply":"2024-10-19T18:59:39.276290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, stratify = y,\n                                                    random_state=42)","metadata":{"id":"465754bf","execution":{"iopub.status.busy":"2024-10-19T18:59:39.278265Z","iopub.execute_input":"2024-10-19T18:59:39.278596Z","iopub.status.idle":"2024-10-19T18:59:39.315954Z","shell.execute_reply.started":"2024-10-19T18:59:39.278565Z","shell.execute_reply":"2024-10-19T18:59:39.315256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.columns","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.316854Z","iopub.execute_input":"2024-10-19T18:59:39.317123Z","iopub.status.idle":"2024-10-19T18:59:39.322963Z","shell.execute_reply.started":"2024-10-19T18:59:39.317095Z","shell.execute_reply":"2024-10-19T18:59:39.321984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## MODEL CREATION ","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers","metadata":{"id":"af1d5826","execution":{"iopub.status.busy":"2024-10-19T18:59:39.324133Z","iopub.execute_input":"2024-10-19T18:59:39.324478Z","iopub.status.idle":"2024-10-19T18:59:39.329611Z","shell.execute_reply.started":"2024-10-19T18:59:39.324438Z","shell.execute_reply":"2024-10-19T18:59:39.328617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = StandardScaler()\nX_train_new = scaler.fit_transform(X_train)\nX_test_new= scaler.transform(X_test)\n","metadata":{"id":"b8bbbd70","execution":{"iopub.status.busy":"2024-10-19T18:59:39.330954Z","iopub.execute_input":"2024-10-19T18:59:39.331252Z","iopub.status.idle":"2024-10-19T18:59:39.350873Z","shell.execute_reply.started":"2024-10-19T18:59:39.331220Z","shell.execute_reply":"2024-10-19T18:59:39.349966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The StandardScaler is a preprocessing technique used in machine learning to standardize the features of a dataset. Standardization involves transforming the data so that it has a mean of 0 and a standard deviation of 1. This process is also known as Z-score normalization.\n\nHere's how the StandardScaler works:\n\nCalculate Mean and Standard Deviation:\n\nFor each feature (column) in the dataset, calculate the mean and standard deviation.\nStandardization Formula:\n\nFor each data point in a feature, apply the following formula:\nStandardized Value\n=\nOriginal Value\n−\nMean\nStandard Deviation\nStandardized Value=\nStandard Deviation\nOriginal Value−Mean\n​\n\nResult:\n\nAfter standardization, each feature will have a mean of 0 and a standard deviation of 1.\nThe StandardScaler is beneficial in scenarios where features have different scales or units. Standardizing the features ensures that they are on a similar scale, which can be important for algorithms that are sensitive to the scale of the input features. Some machine learning algorithms, such as support vector machines, k-means clustering, and principal component analysis, work more effectively when features are standardized.","metadata":{"id":"4245db66"}},{"cell_type":"code","source":"X_train_new.shape,","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.351833Z","iopub.execute_input":"2024-10-19T18:59:39.352152Z","iopub.status.idle":"2024-10-19T18:59:39.358040Z","shell.execute_reply.started":"2024-10-19T18:59:39.352121Z","shell.execute_reply":"2024-10-19T18:59:39.357112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_new","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.359050Z","iopub.execute_input":"2024-10-19T18:59:39.359385Z","iopub.status.idle":"2024-10-19T18:59:39.367872Z","shell.execute_reply.started":"2024-10-19T18:59:39.359355Z","shell.execute_reply":"2024-10-19T18:59:39.366978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.368912Z","iopub.execute_input":"2024-10-19T18:59:39.369208Z","iopub.status.idle":"2024-10-19T18:59:39.379967Z","shell.execute_reply.started":"2024-10-19T18:59:39.369164Z","shell.execute_reply":"2024-10-19T18:59:39.379208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from keras.utils import to_categorical\ny_train_one_hot = to_categorical(y_train)","metadata":{"id":"6f3e04eb","execution":{"iopub.status.busy":"2024-10-19T18:59:39.380990Z","iopub.execute_input":"2024-10-19T18:59:39.381234Z","iopub.status.idle":"2024-10-19T18:59:39.390093Z","shell.execute_reply.started":"2024-10-19T18:59:39.381206Z","shell.execute_reply":"2024-10-19T18:59:39.389205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train_one_hot\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.391171Z","iopub.execute_input":"2024-10-19T18:59:39.391451Z","iopub.status.idle":"2024-10-19T18:59:39.400786Z","shell.execute_reply.started":"2024-10-19T18:59:39.391421Z","shell.execute_reply":"2024-10-19T18:59:39.399923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test_one_hot = to_categorical(y_test)","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.401850Z","iopub.execute_input":"2024-10-19T18:59:39.402187Z","iopub.status.idle":"2024-10-19T18:59:39.409214Z","shell.execute_reply.started":"2024-10-19T18:59:39.402147Z","shell.execute_reply":"2024-10-19T18:59:39.408348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test_one_hot","metadata":{"execution":{"iopub.status.busy":"2024-10-19T18:59:39.410251Z","iopub.execute_input":"2024-10-19T18:59:39.410584Z","iopub.status.idle":"2024-10-19T18:59:39.420633Z","shell.execute_reply.started":"2024-10-19T18:59:39.410543Z","shell.execute_reply":"2024-10-19T18:59:39.419782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = keras.Sequential([\n    layers.Dense(64, activation='relu', input_shape=(X_train_new.shape[1],)),\n    layers.Dense(32, activation='relu'),\n    layers.Dense(6, activation='softmax')\n])","metadata":{"id":"137aee49","execution":{"iopub.status.busy":"2024-10-19T18:59:39.421791Z","iopub.execute_input":"2024-10-19T18:59:39.422094Z","iopub.status.idle":"2024-10-19T18:59:39.455965Z","shell.execute_reply.started":"2024-10-19T18:59:39.422062Z","shell.execute_reply":"2024-10-19T18:59:39.455306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"id":"619041d4","outputId":"89ce51a1-1fed-43e0-a8b0-4a4cf6d2dd65","execution":{"iopub.status.busy":"2024-10-19T18:59:39.456891Z","iopub.execute_input":"2024-10-19T18:59:39.457143Z","iopub.status.idle":"2024-10-19T18:59:39.474098Z","shell.execute_reply.started":"2024-10-19T18:59:39.457115Z","shell.execute_reply":"2024-10-19T18:59:39.473273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The provided model summary belongs to a neural network model built using the Keras library with a Sequential API. Let's break down the key components of the model summary:\n\n### Model Architecture:\n- **Model Type:** Sequential\n  - The Sequential model is a linear stack of layers, where you can simply add one layer at a time.\n  \n### Layers:\n1. **Dense Layer (dense_6):**\n   - **Output Shape:** (None, 64)\n   - **Number of Parameters:** 256\n   - This layer is a fully connected (dense) layer with 64 units. The output shape is (None, 64), indicating that the layer produces an output vector with 64 values for each input.\n\n2. **Dense Layer (dense_7):**\n   - **Output Shape:** (None, 32)\n   - **Number of Parameters:** 2080\n   - Another fully connected layer with 32 units. The output shape is (None, 32).\n\n3. **Dense Layer (dense_8):**\n   - **Output Shape:** (None, 6)\n   - **Number of Parameters:** 198\n   - Final dense layer with 6 units, representing the output layer for a classification task with 6 classes. The output shape is (None, 6).\n\n### Total Parameters:\n- **Total Parameters:** 2,534\n- **Trainable Parameters:** 2,534\n- **Non-trainable Parameters:** 0\n  - Trainable parameters are the weights and biases in the layers that the model learns during training. Non-trainable parameters could be, for example, the parameters in a BatchNormalization layer, but in this case, there are none.\n\n### Note:\n- The activation function used in each dense layer is not specified in the provided summary. By default, Keras uses the linear activation function.\n- The model is designed for a classification task with 6 classes based on the last dense layer having 6 units.\n- The model architecture is relatively simple with one input layer, two hidden layers, and one output layer. Adjustments to the architecture, such as activation functions and layer sizes, can be made based on the specific requirements of your task.","metadata":{"id":"ea4ab6cc"}},{"cell_type":"code","source":"model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])","metadata":{"id":"815894c5","execution":{"iopub.status.busy":"2024-10-19T18:59:39.475164Z","iopub.execute_input":"2024-10-19T18:59:39.475468Z","iopub.status.idle":"2024-10-19T18:59:39.483545Z","shell.execute_reply.started":"2024-10-19T18:59:39.475437Z","shell.execute_reply":"2024-10-19T18:59:39.482706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In Keras, the model.compile method is used to configure the learning process of a neural network model before training. It essentially sets up the model for training by specifying three key components:\n\nOptimizer: The optimizer is the algorithm responsible for updating the weights of the neural network during training. It controls how the model adjusts its internal parameters to minimize the loss function. Common optimizers include SGD (Stochastic Gradient Descent), Adam, RMSprop, etc.\n\nLoss Function: The loss function, or objective function, quantifies how well the model is performing. It measures the difference between the predicted output and the actual target values. The goal during training is to minimize this loss. The choice of the loss function depends on the type of problem (e.g., classification, regression) and the nature of the output.\n\nMetrics: Metrics are used to evaluate the performance of the model. While the loss function is used during training to guide the optimization process, metrics provide additional measures of model performance that are easier to interpret. Common metrics for classification tasks include accuracy, precision, recall, and F1-score.","metadata":{"id":"877b34ec"}},{"cell_type":"code","source":"history = model.fit(X_train_new, y_train_one_hot, epochs=10, batch_size=32, validation_split=0.2)","metadata":{"id":"4d9f6635","outputId":"a705c52d-a380-46af-8081-f0983f73c68a","execution":{"iopub.status.busy":"2024-10-19T18:59:39.484654Z","iopub.execute_input":"2024-10-19T18:59:39.485020Z","iopub.status.idle":"2024-10-19T19:00:04.486114Z","shell.execute_reply.started":"2024-10-19T18:59:39.484977Z","shell.execute_reply":"2024-10-19T19:00:04.485312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_probs = model.predict(X_test_new)\n\n# Round probabilities to 2 decimal places\ny_pred_rounded = np.round(y_pred_probs, 2)\n\nprint(y_pred_rounded)\n","metadata":{"id":"f90fe7fb","execution":{"iopub.status.busy":"2024-10-19T19:00:04.487421Z","iopub.execute_input":"2024-10-19T19:00:04.487742Z","iopub.status.idle":"2024-10-19T19:00:05.820187Z","shell.execute_reply.started":"2024-10-19T19:00:04.487696Z","shell.execute_reply":"2024-10-19T19:00:05.819123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"[i.sum() for i in y_pred_probs[:10]]","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:00:05.821321Z","iopub.execute_input":"2024-10-19T19:00:05.821770Z","iopub.status.idle":"2024-10-19T19:00:05.829194Z","shell.execute_reply.started":"2024-10-19T19:00:05.821707Z","shell.execute_reply":"2024-10-19T19:00:05.828480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_loss, test_accuracy = model.evaluate(X_test_new, y_test_one_hot)\nprint(\"Test Accuracy: {:.2f}%\".format(test_accuracy * 100))","metadata":{"id":"fb0dd593","outputId":"de35185b-ac4a-4a11-c4b6-6498939cb53e","execution":{"iopub.status.busy":"2024-10-19T19:00:05.830239Z","iopub.execute_input":"2024-10-19T19:00:05.830905Z","iopub.status.idle":"2024-10-19T19:00:06.793353Z","shell.execute_reply.started":"2024-10-19T19:00:05.830854Z","shell.execute_reply":"2024-10-19T19:00:06.792489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Conv1D, MaxPooling1D, Flatten, Dense, Dropout\n","metadata":{"id":"d8cbc19e","execution":{"iopub.status.busy":"2024-10-19T19:00:06.794619Z","iopub.execute_input":"2024-10-19T19:00:06.795019Z","iopub.status.idle":"2024-10-19T19:00:06.799351Z","shell.execute_reply.started":"2024-10-19T19:00:06.794984Z","shell.execute_reply":"2024-10-19T19:00:06.798468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_new","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:00:06.800491Z","iopub.execute_input":"2024-10-19T19:00:06.800802Z","iopub.status.idle":"2024-10-19T19:00:06.811458Z","shell.execute_reply.started":"2024-10-19T19:00:06.800763Z","shell.execute_reply":"2024-10-19T19:00:06.810475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"from tensorflow.keras import models, layers\n\n# Reshape X_train_new: each sample is now a sequence of 3 timesteps with 1 feature\nX_train_new_reshaped = X_train_new.reshape((X_train_new.shape[0], X_train_new.shape[1], 1))\n\n# Define the LSTM model\nmodel_lstm = models.Sequential([\n    # LSTM layer to capture temporal patterns\n   layers.Conv1D(64, kernel_size=3, activation='relu', input_shape=(X_train_new.shape[1], 1)),\n    layers.MaxPooling1D(pool_size=2),\n    \n    layers.Conv1D(128, kernel_size=3, activation='relu'),\n    layers.MaxPooling1D(pool_size=2),\n\n    # LSTM layer to capture temporal patterns\n    layers.LSTM(64, return_sequences=True),\n    layers.LSTM(32),\n\n    # Dense layer for further feature extraction\n    layers.Dense(32, activation='relu'),\n\n    # Output layer for seizure prediction (assuming 6 classes)\n    layers.Dense(6, activation='softmax')\n])\n\n# Model summary to show architecture\nmodel_lstm.summary()\n","metadata":{}},{"cell_type":"markdown","source":"from tensorflow.keras import models, layers\n\n# Reshape X_train_new: each sample is now a sequence of 3 timesteps with 1 feature\nX_train_new_reshaped = X_train_new.reshape((X_train_new.shape[0], X_train_new.shape[1], 1))\n\n# Define the LSTM model\nmodel_lstm = models.Sequential([\n    # LSTM layer to capture temporal patterns\n    layers.LSTM(64, activation='tanh', input_shape=(X_train_new.shape[1], 1)),\n    \n    # Dense layer for further feature extraction\n    layers.Dense(32, activation='relu'),\n    \n    # Output layer for seizure prediction (assuming 6 classes)\n    layers.Dense(6, activation='softmax')\n])\n\n\n","metadata":{}},{"cell_type":"code","source":"import tensorflow.keras.backend as K\nfrom tensorflow.keras.layers import Layer\n\n# Define a simple attention mechanism\nclass Attention(Layer):\n    def __init__(self, **kwargs):\n        super(Attention, self).__init__(**kwargs)\n\n    def build(self, input_shape):\n        self.W = self.add_weight(shape=(input_shape[-1], input_shape[-1]), initializer=\"glorot_uniform\", trainable=True)\n        self.b = self.add_weight(shape=(input_shape[-1],), initializer=\"zeros\", trainable=True)\n        self.u = self.add_weight(shape=(input_shape[-1], 1), initializer=\"glorot_uniform\", trainable=True)\n        super(Attention, self).build(input_shape)\n\n    def call(self, x):\n        # Calculate attention score\n        score = K.tanh(K.dot(x, self.W) + self.b)\n\n        # Compute attention weights and apply softmax\n        attention_weights = K.softmax(K.dot(score, self.u), axis=1)\n\n        # Compute the context vector as the weighted sum\n        context_vector = attention_weights * x\n        context_vector = K.sum(context_vector, axis=1)\n\n        return context_vector\n\n    def compute_output_shape(self, input_shape):\n        return (input_shape[0], input_shape[-1])\n\n# Add this custom attention layer to your model\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:01:46.716850Z","iopub.execute_input":"2024-10-19T19:01:46.717281Z","iopub.status.idle":"2024-10-19T19:01:46.729528Z","shell.execute_reply.started":"2024-10-19T19:01:46.717242Z","shell.execute_reply":"2024-10-19T19:01:46.728578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models\n\n# Reshape X_train_new: each sample is now a sequence of 3 timesteps with 1 feature\nX_train_new_reshaped = X_train_new.reshape((X_train_new.shape[0], X_train_new.shape[1], 1))\n\n# Define a more complex model\nmodel_complex = models.Sequential([\n\n    # 1D Convolution layer to capture local temporal patterns\n    layers.Conv1D(128, kernel_size=3, activation='relu', input_shape=(X_train_new.shape[1], 1)),\n    layers.MaxPooling1D(pool_size=2),\n    \n    # BiLSTM layer to capture temporal patterns in both directions\n    layers.Bidirectional(layers.LSTM(64, return_sequences=True)),\n\n    # Attention mechanism to focus on important timesteps\n    Attention(),\n    \n    # Dense layer for further feature extraction\n    layers.Dense(64, activation='relu'),\n\n    # Output layer for seizure prediction (assuming 6 classes)\n    layers.Dense(6, activation='softmax')\n])\n\n# Compile the model\nmodel_complex.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])\n\n# Model summary to show architecture\nmodel_complex.summary()\n\n","metadata":{"id":"7330f221","execution":{"iopub.status.busy":"2024-10-19T19:02:05.897598Z","iopub.execute_input":"2024-10-19T19:02:05.898004Z","iopub.status.idle":"2024-10-19T19:02:05.994497Z","shell.execute_reply.started":"2024-10-19T19:02:05.897967Z","shell.execute_reply":"2024-10-19T19:02:05.993606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. **Input Shape Definition:**\n   - The input data shape is defined, indicating the number of features and channels.\n\n2. **Model Initialization:**\n   - A sequential model is initialized. In a sequential model, layers are added one by one in a linear sequence.\n\n3. **Convolutional Layer Addition:**\n   - A 1D convolutional layer is added to the model.\n     - `filters`: Specifies the number of filters in the convolutional layer.\n     - `kernel_size`: Sets the size of the convolutional kernel.\n\n4. **Flatten Layer Addition:**\n   - A flatten layer is added to the model.\n     - This layer is used to flatten the multi-dimensional data into a one-dimensional array.\n\n5. **Dense Layer Addition (1st):**\n   - A dense layer with ReLU activation is added to the model.\n     - This layer is fully connected, meaning each neuron in the layer is connected to every neuron in the previous layer.\n\n6. **Dense Layer Addition (2nd):**\n   - Another dense layer with softmax activation is added to the model.\n     - Softmax activation is often used in the output layer for multi-class classification problems.","metadata":{"id":"443fee4a"}},{"cell_type":"code","source":"# Train the model\nhistory_complex = model_complex.fit(X_train_new_reshaped, y_train_one_hot, epochs=20, batch_size=32, validation_split=0.2)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:02:10.561678Z","iopub.execute_input":"2024-10-19T19:02:10.562107Z","iopub.status.idle":"2024-10-19T19:05:15.258382Z","shell.execute_reply.started":"2024-10-19T19:02:10.562068Z","shell.execute_reply":"2024-10-19T19:05:15.257538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_probs = model_complex.predict(X_test_new)\n\n# Round probabilities to 2 decimal places\ny_pred_rounded = np.round(y_pred_probs, 2)\n\nprint(y_pred_rounded)","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:05:26.185636Z","iopub.execute_input":"2024-10-19T19:05:26.186052Z","iopub.status.idle":"2024-10-19T19:05:28.376037Z","shell.execute_reply.started":"2024-10-19T19:05:26.186011Z","shell.execute_reply":"2024-10-19T19:05:28.375094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. **Conv1D Layer:**\n   - Type: Conv1D\n   - Output Shape: (None, 1, 64)\n   - Parameters: 256\n   - Explanation: This layer applies 64 filters of size 3 to the input data, resulting in an output shape of (1, 64) for each filter.\n\n2. **Flatten Layer:**\n   - Type: Flatten\n   - Output Shape: (None, 64)\n   - Parameters: 0\n   - Explanation: This layer flattens the output from the Conv1D layer into a one-dimensional array with 64 elements.\n\n3. **Dense Layer (1st):**\n   - Type: Dense\n   - Output Shape: (None, 128)\n   - Parameters: 8,320\n   - Explanation: This fully connected layer has 128 neurons with a ReLU activation function. It takes the flattened input from the previous layer.\n\n4. **Dense Layer (2nd):**\n   - Type: Dense\n   - Output Shape: (None, 6)\n   - Parameters: 774\n   - Explanation: The final layer is another fully connected layer with 6 neurons, representing the number of classes in the classification task. It uses the softmax activation function.\n\n5. **Total Parameters:**\n   - Total trainable parameters in the model: 9,350\n   - Non-trainable parameters: 0\n   - Trainable parameters are the weights and biases in the model that are updated during training to learn from the data. Non-trainable parameters are fixed and not updated during training.\n\nThis model architecture is suitable for a multi-class classification problem where the input data has sequential features. The convolutional layer helps capture local patterns in the input sequence.","metadata":{"id":"62e1d75e"}},{"cell_type":"markdown","source":"- **`X_train_new.reshape(X_train_new.shape[0], X_train_new.shape[1], 1)`**: This reshapes the input data (`X_train_new`) to be compatible with the Conv1D layer's input shape. It adds an additional dimension with size 1, indicating that the data has a single channel. The reshaped data is then used for training.\n\n- **`y_train_one_hot`**: This represents the one-hot encoded labels corresponding to the training data.\n\n- **`epochs=20`**: This specifies the number of times the entire dataset will be passed forward and backward through the neural network during training.\n\n- **`batch_size=32`**: It defines the number of samples that will be used in each iteration of training.\n\n- **`validation_split=0.2`**: This parameter is used for validation during training. It indicates that 20% of the training data will be used as a validation set, and the model's performance will be evaluated on this set after each epoch.\n\nThe `fit` method will train the model for the specified number of epochs, and the training/validation loss and accuracy metrics for each epoch will be stored in the `history_cnn` variable. This information can be used to analyze the training process and evaluate the model's performance.","metadata":{"id":"16d7c77a"}},{"cell_type":"code","source":"pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:13:01.013517Z","iopub.execute_input":"2024-10-19T19:13:01.014250Z","iopub.status.idle":"2024-10-19T19:13:01.030769Z","shell.execute_reply.started":"2024-10-19T19:13:01.014207Z","shell.execute_reply":"2024-10-19T19:13:01.029821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.read_parquet('/kaggle/input/hms-harmful-brain-activity-classification/test_eegs/3911565283.parquet')\ntest_df\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:05:39.269428Z","iopub.execute_input":"2024-10-19T19:05:39.270309Z","iopub.status.idle":"2024-10-19T19:05:39.326683Z","shell.execute_reply.started":"2024-10-19T19:05:39.270265Z","shell.execute_reply":"2024-10-19T19:05:39.325770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(21, figsize=(10, 100))\n\nfor i, column in enumerate(test_df.columns):\n    ax[i].plot(test_df.index, test_df[column], label=column)\n    ax[i].grid(True)\n    ax[i].set_title(str(column))\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:05:39.499817Z","iopub.execute_input":"2024-10-19T19:05:39.500128Z","iopub.status.idle":"2024-10-19T19:05:44.062131Z","shell.execute_reply.started":"2024-10-19T19:05:39.500096Z","shell.execute_reply":"2024-10-19T19:05:44.061109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaled_test_df=scaler.transform(test_df[['C3', 'Cz', 'Fp2', 'O2']])\ny_pred_probs = model_complex.predict(scaled_test_df)\n\n# Round probabilities to 2 decimal places\ny_pred_rounded = np.round(y_pred_probs, 2)\n\nprint(y_pred_rounded[0])\nprint(y_pred_rounded[0].sum( ))","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:05:44.063979Z","iopub.execute_input":"2024-10-19T19:05:44.064341Z","iopub.status.idle":"2024-10-19T19:05:44.941226Z","shell.execute_reply.started":"2024-10-19T19:05:44.064301Z","shell.execute_reply":"2024-10-19T19:05:44.940322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\n\n# Set the max probability in each row to 1, others to 0\ny_pred_max = np.zeros_like(y_pred_rounded)\ny_pred_max[np.arange(len(y_pred_rounded)), y_pred_rounded.argmax(axis=1)] = 1\n\n# Sum column-wise and divide by the total number of rows to get the class probability\nclass_probabilities = y_pred_max.sum(axis=0) / len(y_pred_max)\n\nprint(class_probabilities)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:06:13.906341Z","iopub.execute_input":"2024-10-19T19:06:13.907040Z","iopub.status.idle":"2024-10-19T19:06:13.914460Z","shell.execute_reply.started":"2024-10-19T19:06:13.906996Z","shell.execute_reply":"2024-10-19T19:06:13.913435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count=0\nfor i in class_probabilities:\n    print(i)","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:09:37.551099Z","iopub.execute_input":"2024-10-19T19:09:37.552027Z","iopub.status.idle":"2024-10-19T19:09:37.556842Z","shell.execute_reply.started":"2024-10-19T19:09:37.551984Z","shell.execute_reply":"2024-10-19T19:09:37.555755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_data=pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:52:39.322387Z","iopub.execute_input":"2024-10-19T19:52:39.323194Z","iopub.status.idle":"2024-10-19T19:52:39.337048Z","shell.execute_reply.started":"2024-10-19T19:52:39.323151Z","shell.execute_reply":"2024-10-19T19:52:39.336134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_data\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:52:40.062569Z","iopub.execute_input":"2024-10-19T19:52:40.063283Z","iopub.status.idle":"2024-10-19T19:52:40.074649Z","shell.execute_reply.started":"2024-10-19T19:52:40.063245Z","shell.execute_reply":"2024-10-19T19:52:40.073623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assign the values to the appropriate columns\nsub_data.loc[0, ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']] = class_probabilities\n","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:57:03.704446Z","iopub.execute_input":"2024-10-19T19:57:03.704828Z","iopub.status.idle":"2024-10-19T19:57:03.712808Z","shell.execute_reply.started":"2024-10-19T19:57:03.704788Z","shell.execute_reply":"2024-10-19T19:57:03.711787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_data","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:58:51.936580Z","iopub.execute_input":"2024-10-19T19:58:51.936983Z","iopub.status.idle":"2024-10-19T19:58:51.948443Z","shell.execute_reply.started":"2024-10-19T19:58:51.936944Z","shell.execute_reply":"2024-10-19T19:58:51.947493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_data.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2024-10-19T19:58:52.390660Z","iopub.execute_input":"2024-10-19T19:58:52.391648Z","iopub.status.idle":"2024-10-19T19:58:52.398053Z","shell.execute_reply.started":"2024-10-19T19:58:52.391611Z","shell.execute_reply":"2024-10-19T19:58:52.396078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}