{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1>Introduction</h1>","metadata":{}},{"cell_type":"markdown","source":"<h3>In this notebook I will attempt to fit electroencephalogram data in order to predict harmful brain activity. I am using a convolutional neural network from the TensorFlow framework.</h3>","metadata":{}},{"cell_type":"markdown","source":"<h1>Prepare Training Data</h1>","metadata":{}},{"cell_type":"markdown","source":"<h2>Define Imports and Set Global Variables</h2>","metadata":{}},{"cell_type":"code","source":"import os # Using for directory operations\nimport tqdm # progress bars\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt # plot it\nimport librosa\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier # For initial testing\n\n# Competition directory for that houses all data.\nCOMPDIR = '/kaggle/input/hms-harmful-brain-activity-classification'\n\n# Reading the CSV file 'train.csv' located in the directory specified by COMPDIR\ntrain = pd.read_csv(os.path.join(COMPDIR, 'train.csv'))\n\n# Reading the CSV file 'test.csv' located in the directory specified by COMPDIR\ntest = pd.read_csv(os.path.join(COMPDIR, 'test.csv'))","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:24:55.638295Z","iopub.execute_input":"2024-04-22T19:24:55.639292Z","iopub.status.idle":"2024-04-22T19:25:13.134700Z","shell.execute_reply.started":"2024-04-22T19:24:55.639231Z","shell.execute_reply":"2024-04-22T19:25:13.133403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Load CSV Meta Data</h2>","metadata":{}},{"cell_type":"code","source":"# Displaying the first and last few rows of the training DataFrame.\ndisplay(train.head())\ndisplay(train.tail())\n\nprint(\"Train shape:\", train.shape, \"\\n\")\nprint(\"Unique eeg_ids: \", train.eeg_id.nunique())\nprint(train.groupby(\"eeg_id\")[\"eeg_sub_id\"].count().describe(), \"\\n\")\nprint(\"Unique spectrogram_ids: \", train.spectrogram_id.nunique())\nprint(\"Unique patient_ids: \", train.patient_id.nunique(), \"\\n\")","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:13.136594Z","iopub.execute_input":"2024-04-22T19:25:13.136956Z","iopub.status.idle":"2024-04-22T19:25:13.212415Z","shell.execute_reply.started":"2024-04-22T19:25:13.136926Z","shell.execute_reply":"2024-04-22T19:25:13.211191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(9):\n    print(train.loc[i], \"\\n\")\n    \n# Example of EEG data that we care about from the Cz electrode\neeg_data = pd.read_parquet(os.path.join(COMPDIR, 'train_eegs', f'1628180742.parquet'))\n\n# Extracting data for Cz electrode example\ncz_electrode_data = eeg_data['Cz'][:]\n\nprint(cz_electrode_data)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:13.219504Z","iopub.execute_input":"2024-04-22T19:25:13.220321Z","iopub.status.idle":"2024-04-22T19:25:13.437011Z","shell.execute_reply.started":"2024-04-22T19:25:13.220283Z","shell.execute_reply":"2024-04-22T19:25:13.436020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Displaying the testing DataFrame.\ndisplay(test)\n\nprint(\"Test shape:\", test.shape, \"\\n\")\nprint(\"Unique eeg_ids: \", test.eeg_id.nunique())\nprint(\"Unique spectrogram_ids: \", test.spectrogram_id.nunique())\nprint(\"Unique patient_ids: \", test.patient_id.nunique(), \"\\n\")","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:13.438545Z","iopub.execute_input":"2024-04-22T19:25:13.439223Z","iopub.status.idle":"2024-04-22T19:25:13.451556Z","shell.execute_reply.started":"2024-04-22T19:25:13.439190Z","shell.execute_reply":"2024-04-22T19:25:13.450246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example of EEG data that we care about from the Cz electrode\neeg_data = pd.read_parquet(os.path.join(COMPDIR, 'test_eegs', f'3911565283.parquet'))\n\n# Extracting data for Cz electrode example\ncz_electrode_data = eeg_data['Cz'][:]\n\nprint(cz_electrode_data)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:13.453575Z","iopub.execute_input":"2024-04-22T19:25:13.454385Z","iopub.status.idle":"2024-04-22T19:25:13.511083Z","shell.execute_reply.started":"2024-04-22T19:25:13.454347Z","shell.execute_reply":"2024-04-22T19:25:13.510174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Load EEG and Spectrogram Data</h2>","metadata":{}},{"cell_type":"code","source":"# Setting the sampling frequency and duration for EEG data collection\nsampling_frequency = 200  # Sampling frequency in Hz\ndata_collection_duration = 50  # Duration of EEG data collection in seconds\ntotal_samples = sampling_frequency * data_collection_duration  # Total number of samples in the duration\ntime_values = np.linspace(0, data_collection_duration, total_samples)  # Time values from 0 to 50 seconds\n\n# Setting the number of training data points\nnum_train_data_points = 500  # 106800, all of the data\n\n# Creating an empty DataFrame to store training data\ntraining_data_df = pd.DataFrame()\n\n# Iterating over each training data point\nfor i in tqdm.tqdm(range(num_train_data_points)):\n    # Loading EEG data and spectrogram data for specified ids\n    eeg_id = train.loc[i, 'eeg_id']\n    spec_id = train.loc[i, 'spectrogram_id']\n    eeg_data = pd.read_parquet(os.path.join(COMPDIR, 'train_eegs', f'{eeg_id}.parquet'))\n    spec_data = pd.read_parquet(os.path.join(COMPDIR, 'train_spectrograms', f'{spec_id}.parquet'))\n    last_eeg_id = eeg_id\n\n    # Extracting EEG data from the Cz electrode for 50 seconds\n    label_offset_time = train.loc[i, 'eeg_label_offset_seconds']  # Offset time for the EEG label\n    label_offset_index = int(sampling_frequency * label_offset_time)  # Calculating offset index\n    cz_electrode_data = eeg_data['Cz'][label_offset_index:label_offset_index + total_samples]  # Extracting data for Cz electrode\n\n    # Adding the extracted data as a row to the training DataFrame\n    training_data_df = pd.concat([training_data_df, cz_electrode_data.reset_index(drop=True).to_frame().transpose()], axis=0)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:13.512579Z","iopub.execute_input":"2024-04-22T19:25:13.513224Z","iopub.status.idle":"2024-04-22T19:25:42.376514Z","shell.execute_reply.started":"2024-04-22T19:25:13.513191Z","shell.execute_reply":"2024-04-22T19:25:42.374962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>EEG Data Displayed</h2>","metadata":{}},{"cell_type":"code","source":"display(eeg_data.head())\ndisplay(eeg_data.tail())\nprint(eeg_data.shape)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:42.379284Z","iopub.execute_input":"2024-04-22T19:25:42.380046Z","iopub.status.idle":"2024-04-22T19:25:42.430062Z","shell.execute_reply.started":"2024-04-22T19:25:42.380001Z","shell.execute_reply":"2024-04-22T19:25:42.428871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Spectrogram Data Displayed</h2>","metadata":{}},{"cell_type":"code","source":"display(spec_data.head())\ndisplay(spec_data.tail())\nprint(spec_data.shape)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:42.431882Z","iopub.execute_input":"2024-04-22T19:25:42.432309Z","iopub.status.idle":"2024-04-22T19:25:42.492023Z","shell.execute_reply.started":"2024-04-22T19:25:42.432271Z","shell.execute_reply":"2024-04-22T19:25:42.490823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Training Data Displayed</h2>","metadata":{}},{"cell_type":"code","source":"display(training_data_df)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:42.496147Z","iopub.execute_input":"2024-04-22T19:25:42.496536Z","iopub.status.idle":"2024-04-22T19:25:42.535287Z","shell.execute_reply.started":"2024-04-22T19:25:42.496506Z","shell.execute_reply":"2024-04-22T19:25:42.534093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Prepare Features (X_train) and Target Variable (y_train)</h2>","metadata":{}},{"cell_type":"code","source":"# Adding diagnosis results, some data was dropped previously, account for that with len(training_data_df)\ntraining_data_df['expert_consensus'] = train[:len(training_data_df)]['expert_consensus'].values\n\ndisplay(training_data_df['expert_consensus'].head())\ndisplay(training_data_df['expert_consensus'].tail())\n\n# Removing rows with missing values\ntraining_data_df = training_data_df.dropna()\ntraining_data_df = training_data_df.reset_index(drop=True)\n\n# Separating data into features and target\ny_train = training_data_df['expert_consensus']\nX_train = training_data_df.drop('expert_consensus', axis=1)\n\n# Displaying the first and last few rows of the feature dataset\ndisplay(X_train.head())\ndisplay(X_train.tail())\nprint(X_train.shape)\nprint(y_train)\nprint(y_train.shape)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:42.536667Z","iopub.execute_input":"2024-04-22T19:25:42.537057Z","iopub.status.idle":"2024-04-22T19:25:42.636611Z","shell.execute_reply.started":"2024-04-22T19:25:42.537024Z","shell.execute_reply":"2024-04-22T19:25:42.635111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>EEG Data Visualized</h2>","metadata":{}},{"cell_type":"code","source":"X_train = X_train.T # transpose for plotting\n\n# Plot each set of samples separately\nfor i in range(int(num_train_data_points/50)):\n    plt.plot(time_values, X_train[i], label=f'Set {i+1}')\n\nplt.title('Samples over Time')\nplt.xlabel('Time (seconds)')\nplt.ylabel('Sampled Values')\nplt.legend()  # Show legend with labels for each set\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:42.638452Z","iopub.execute_input":"2024-04-22T19:25:42.638810Z","iopub.status.idle":"2024-04-22T19:25:43.634253Z","shell.execute_reply.started":"2024-04-22T19:25:42.638761Z","shell.execute_reply":"2024-04-22T19:25:43.633086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create subplots for each set of samples\nfig, axs = plt.subplots(int(num_train_data_points/50), 1, figsize=(8, 10), sharex=True, sharey=True)\n\n# Plot each set of samples on its own subplot\nfor i in range(int(num_train_data_points/50)):\n    axs[i].plot(time_values, X_train[i])\n    axs[i].set_title(f'Set {i+1}')\n    axs[i].grid(True)\n\n# Set common labels for x-axis and y-axis\nfig.text(0.5, 0, 'Time (seconds)', ha='center', va='center')\nfig.text(0, 0.5, 'Sampled Values', ha='center', va='center', rotation='vertical')\n\nplt.subplots_adjust(hspace=0.5)  # Increase space between subplots\nplt.tight_layout()  # Adjust layout to prevent overlap\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:43.639769Z","iopub.execute_input":"2024-04-22T19:25:43.640360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Preprocess Data</h2>","metadata":{}},{"cell_type":"markdown","source":"<h4>Since X_train is a dataset of floats and y_train is a dataset of strings, we need to encode the labels into numeric form and split your data into training and validation sets.</h4>","metadata":{}},{"cell_type":"code","source":"X_train = X_train.T\n\n# Encode labels into numerical form\nlabel_encoder = LabelEncoder()\ny_train_encoded = label_encoder.fit_transform(y_train)\n\n# Split data into training and validation sets\nX_train, X_val, y_train_encoded, y_val_encoded = train_test_split(X_train, y_train_encoded, test_size=0.2, random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Define and Compile CNN Model</h2>","metadata":{}},{"cell_type":"code","source":"input_shape = X_train.shape[1:]\n\nmodel = models.Sequential([\n    layers.Reshape(input_shape=input_shape + (1,), target_shape=input_shape + (1,)),  # Reshape for 1D Conv input\n    layers.Conv1D(32, kernel_size=3, activation='relu'),\n    layers.MaxPooling1D(pool_size=2),\n    layers.Conv1D(64, kernel_size=3, activation='relu'),\n    layers.MaxPooling1D(pool_size=2),\n    layers.Flatten(),\n    layers.Dense(128, activation='relu'),\n    layers.Dense(len(np.unique(y_train_encoded)), activation='softmax')  # Output layer\n])\n\nmodel.summary()\n\nmodel.compile(optimizer='adam',\n          loss='sparse_categorical_crossentropy',\n          metrics=['accuracy'])","metadata":{"execution":{"iopub.status.idle":"2024-04-22T19:25:46.238663Z","shell.execute_reply.started":"2024-04-22T19:25:45.936239Z","shell.execute_reply":"2024-04-22T19:25:46.237335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Train Model</h2>","metadata":{}},{"cell_type":"code","source":"history = model.fit(X_train, y_train_encoded, epochs=10, batch_size=32, validation_data=(X_val, y_val_encoded))","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:25:46.240279Z","iopub.execute_input":"2024-04-22T19:25:46.240683Z","iopub.status.idle":"2024-04-22T19:27:16.777088Z","shell.execute_reply.started":"2024-04-22T19:25:46.240646Z","shell.execute_reply":"2024-04-22T19:27:16.776109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Loss and Accuracy Visualized</h2>","metadata":{}},{"cell_type":"code","source":"# Plot training & validation loss values\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend(['Train', 'Validation'], loc='upper right')\nplt.show()\n\n# Plot training & validation accuracy values\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend(['Train', 'Validation'], loc='lower right')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:27:16.780821Z","iopub.execute_input":"2024-04-22T19:27:16.781309Z","iopub.status.idle":"2024-04-22T19:27:17.288376Z","shell.execute_reply.started":"2024-04-22T19:27:16.781256Z","shell.execute_reply":"2024-04-22T19:27:17.286867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1>Prepare Test Data</h1>","metadata":{}},{"cell_type":"markdown","source":"<h2>Prepare Features (X_test)</h2>","metadata":{}},{"cell_type":"code","source":"# Creating an empty DataFrame to store testing data\nX_test = pd.DataFrame()\n\n# Iterating over each test data point\nfor i in tqdm.tqdm(range(len(test))):\n    # Loading EEG data for a specified eeg_id\n    eeg_id_ = test.loc[i, 'eeg_id']\n    tmp = pd.read_parquet(os.path.join(COMPDIR, 'test_eegs', f'{eeg_id_}.parquet'))\n    \n    # Extracting EEG data from the Cz electrode\n    cz_electrode_data = tmp['Cz']\n    \n    # Adding the extracted data as a row to the testing DataFrame\n    X_test = pd.concat([X_test, cz_electrode_data.reset_index(drop=True).to_frame().transpose()], axis=0)\n    \ndisplay(cz_electrode_data)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:27:17.289970Z","iopub.execute_input":"2024-04-22T19:27:17.290426Z","iopub.status.idle":"2024-04-22T19:27:17.327051Z","shell.execute_reply.started":"2024-04-22T19:27:17.290382Z","shell.execute_reply":"2024-04-22T19:27:17.325845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Predict and Submit</h2>","metadata":{}},{"cell_type":"code","source":"# Assuming label_encoder is your LabelEncoder instance used during training\nclass_labels = label_encoder.inverse_transform(np.arange(len(label_encoder.classes_)))\n\n# Calculate predictions using the trained RandomForestClassifier model\npredictions = model.predict(X_test)\n\n# Read the sample submission file\nsubmission = pd.read_csv(f'{COMPDIR}/sample_submission.csv')\n\n# Iterate over each test data point\nfor i in tqdm.tqdm(range(len(test))):\n    # Set the 'eeg_id' in the submission DataFrame\n    submission.loc[i, 'eeg_id'] = test.loc[i, 'eeg_id']\n    \n    # Set the probability for each class in the submission DataFrame\n    for j, cls_name in enumerate(class_labels):\n        # Note: This actually performs better when introducing pseudo-random data.\n        #submission.loc[i, f'{cls_name.lower()}_vote'] = random_variables[j]        \n        submission.loc[i, f'{cls_name.lower()}_vote'] = predictions[i, j]","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:27:17.329306Z","iopub.execute_input":"2024-04-22T19:27:17.330248Z","iopub.status.idle":"2024-04-22T19:27:17.530633Z","shell.execute_reply.started":"2024-04-22T19:27:17.330196Z","shell.execute_reply":"2024-04-22T19:27:17.529470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the submission DataFrame\ndisplay(submission)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:27:17.532403Z","iopub.execute_input":"2024-04-22T19:27:17.532893Z","iopub.status.idle":"2024-04-22T19:27:17.551533Z","shell.execute_reply.started":"2024-04-22T19:27:17.532849Z","shell.execute_reply":"2024-04-22T19:27:17.549857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Saving the submission DataFrame to a CSV file without including the index\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:27:17.552881Z","iopub.execute_input":"2024-04-22T19:27:17.553353Z","iopub.status.idle":"2024-04-22T19:27:17.564058Z","shell.execute_reply.started":"2024-04-22T19:27:17.553281Z","shell.execute_reply":"2024-04-22T19:27:17.562249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Past Testing</h2>","metadata":{}},{"cell_type":"code","source":"# X_train = X_train.T # transpose for fitting\n\n# # Initializing a RandomForestClassifier with a random state of 42\n# forest = RandomForestClassifier(random_state=42)\n\n# # Fitting the classifier to the training data\n# forest.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:27:17.565513Z","iopub.execute_input":"2024-04-22T19:27:17.566039Z","iopub.status.idle":"2024-04-22T19:27:17.572487Z","shell.execute_reply.started":"2024-04-22T19:27:17.565977Z","shell.execute_reply":"2024-04-22T19:27:17.571025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3>Generating random variables actually performs better than the random forest.</h3>","metadata":{}},{"cell_type":"code","source":"# Generate 6 random positive variables that add up to 1 using the Dirichlet distribution\n# random_variables = np.random.dirichlet(np.ones(6), size=1)[0]\n# print(\"Random variables:\", random_variables)\n# print(\"Sum of variables:\", sum(random_variables))","metadata":{"execution":{"iopub.status.busy":"2024-04-22T19:27:17.574472Z","iopub.execute_input":"2024-04-22T19:27:17.574978Z","iopub.status.idle":"2024-04-22T19:27:17.583469Z","shell.execute_reply.started":"2024-04-22T19:27:17.574937Z","shell.execute_reply":"2024-04-22T19:27:17.582269Z"},"trusted":true},"execution_count":null,"outputs":[]}]}