{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Harmful Brain Activity Classification - Multi Layer Perceptron**\n\n## **Written by:** Aarish Asif Khan\n\n## **Date:** 6 March 2024","metadata":{}},{"cell_type":"markdown","source":"# **Multi Layer Perceptron**\n\nA Multi-Layer Perceptron (MLP) is a type of feedforward artificial neural network that consists of multiple layers of nodes, also known as neurons. It is a supervised learning algorithm used for classification and regression tasks. MLPs are capable of learning non-linear relationships in data and are widely used in various fields such as image recognition, natural language processing, and financial forecasting.\n\n## **Structure of a Multi-Layer Perceptron**\n\nAn MLP consists of three types of layers:\n\n1. **Input Layer**: This layer receives the input features and passes them to the next layer. The number of neurons in the input layer corresponds to the number of features in the input data.\n\n2. **Hidden Layers**: These are intermediate layers between the input and output layers. Each hidden layer consists of multiple neurons that perform computations on the input data. The number of hidden layers and the number of neurons in each layer are hyperparameters that need to be determined based on the complexity of the problem.\n\n3. **Output Layer**: The output layer produces the final output of the network. The number of neurons in the output layer depends on the type of problem being solved (e.g., regression, classification). For example, in binary classification tasks, there is usually one neuron in the output layer, whereas in multi-class classification tasks, the number of neurons equals the number of classes.\n\nEach neuron in an MLP is connected to every neuron in the subsequent layer, and each connection is associated with a weight that determines the strength of the connection.\n\n## **Activation Function**\n\nAn activation function is applied to the output of each neuron in the hidden layers to introduce non-linearity into the network, allowing it to learn complex patterns. Common activation functions used in MLPs include:\n\n- **ReLU (Rectified Linear Unit)**: Returns the input if it is positive, otherwise returns zero.\n- **Sigmoid**: Maps the input to a range between 0 and 1.\n- **Tanh**: Similar to the sigmoid function but maps the input to a range between -1 and 1.\n\nThe choice of activation function can impact the performance and training speed of the MLP.\n\n## **Training a Multi-Layer Perceptron**\n\nTraining an MLP involves presenting input data along with the corresponding target outputs and adjusting the weights of the connections between neurons to minimize the error between the predicted and actual outputs. This is typically done using optimization algorithms such as stochastic gradient descent (SGD), Adam, or RMSprop.\n","metadata":{}},{"cell_type":"code","source":"# Remove warnings\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:06.295309Z","iopub.execute_input":"2024-03-27T16:27:06.296022Z","iopub.status.idle":"2024-03-27T16:27:06.333231Z","shell.execute_reply.started":"2024-03-27T16:27:06.295985Z","shell.execute_reply":"2024-03-27T16:27:06.331976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import libraries\nimport pandas as pd\nimport numpy as np\n\nimport tensorflow as tf\nimport matplotlib.pyplot as plt\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler, LabelEncoder\n\nfrom tensorflow.keras import layers, models","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:06.335849Z","iopub.execute_input":"2024-03-27T16:27:06.336285Z","iopub.status.idle":"2024-03-27T16:27:22.007973Z","shell.execute_reply.started":"2024-03-27T16:27:06.336245Z","shell.execute_reply":"2024-03-27T16:27:22.006684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.009286Z","iopub.execute_input":"2024-03-27T16:27:22.009899Z","iopub.status.idle":"2024-03-27T16:27:22.297031Z","shell.execute_reply.started":"2024-03-27T16:27:22.009867Z","shell.execute_reply":"2024-03-27T16:27:22.296119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print the 5 rows of the dataset\ntrain_data.head(5)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.299317Z","iopub.execute_input":"2024-03-27T16:27:22.299667Z","iopub.status.idle":"2024-03-27T16:27:22.329474Z","shell.execute_reply.started":"2024-03-27T16:27:22.299639Z","shell.execute_reply":"2024-03-27T16:27:22.328136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split data into features and labels\nX = train_data.drop(columns=[\"eeg_id\", \"eeg_sub_id\", \"spectrogram_id\", \"spectrogram_sub_id\", \"label_id\", \"patient_id\", \"expert_consensus\"])\ny = train_data[\"expert_consensus\"]","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.330805Z","iopub.execute_input":"2024-03-27T16:27:22.331705Z","iopub.status.idle":"2024-03-27T16:27:22.344174Z","shell.execute_reply.started":"2024-03-27T16:27:22.331671Z","shell.execute_reply":"2024-03-27T16:27:22.343104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split data into training and validation sets\nX_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.345731Z","iopub.execute_input":"2024-03-27T16:27:22.346084Z","iopub.status.idle":"2024-03-27T16:27:22.370612Z","shell.execute_reply.started":"2024-03-27T16:27:22.346032Z","shell.execute_reply":"2024-03-27T16:27:22.369535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pre-processing the dataset\n\n# Initialize LabelEncoder\nlabel_encoder = LabelEncoder()\n\n# Fit LabelEncoder on the labels and transform them to numerical format\ny_train_encoded = label_encoder.fit_transform(y_train)\ny_val_encoded = label_encoder.transform(y_val)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.371573Z","iopub.execute_input":"2024-03-27T16:27:22.371879Z","iopub.status.idle":"2024-03-27T16:27:22.409424Z","shell.execute_reply.started":"2024-03-27T16:27:22.371853Z","shell.execute_reply":"2024-03-27T16:27:22.408532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Normalize features\nscaler = StandardScaler()\nX_train = scaler.fit_transform(X_train)\nX_val = scaler.transform(X_val)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.410471Z","iopub.execute_input":"2024-03-27T16:27:22.411748Z","iopub.status.idle":"2024-03-27T16:27:22.442534Z","shell.execute_reply.started":"2024-03-27T16:27:22.411705Z","shell.execute_reply":"2024-03-27T16:27:22.441420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Build the MLP Model\nmodel = models.Sequential([\n    layers.Dense(64, activation='relu', input_shape=(X_train.shape[1],)),\n    layers.Dense(64, activation='relu'),\n    layers.Dense(1, activation='sigmoid')\n])","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.444282Z","iopub.execute_input":"2024-03-27T16:27:22.444697Z","iopub.status.idle":"2024-03-27T16:27:22.736308Z","shell.execute_reply.started":"2024-03-27T16:27:22.444660Z","shell.execute_reply":"2024-03-27T16:27:22.735194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compile the model\nmodel.compile(optimizer='adam',\n              loss='binary_crossentropy',\n              metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.739748Z","iopub.execute_input":"2024-03-27T16:27:22.740096Z","iopub.status.idle":"2024-03-27T16:27:22.753613Z","shell.execute_reply.started":"2024-03-27T16:27:22.740069Z","shell.execute_reply":"2024-03-27T16:27:22.752485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Unique classes after encoding:\", label_encoder.classes_)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.754653Z","iopub.execute_input":"2024-03-27T16:27:22.758403Z","iopub.status.idle":"2024-03-27T16:27:22.763788Z","shell.execute_reply.started":"2024-03-27T16:27:22.758359Z","shell.execute_reply":"2024-03-27T16:27:22.762725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train_encoded = y_train_encoded.astype(float)\ny_val_encoded = y_val_encoded.astype(float)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.765293Z","iopub.execute_input":"2024-03-27T16:27:22.765582Z","iopub.status.idle":"2024-03-27T16:27:22.773828Z","shell.execute_reply.started":"2024-03-27T16:27:22.765558Z","shell.execute_reply":"2024-03-27T16:27:22.772835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the Model\nhistory = model.fit(X_train, y_train_encoded, epochs=10, batch_size=32, validation_data=(X_val, y_val_encoded))","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:27:22.775147Z","iopub.execute_input":"2024-03-27T16:27:22.775437Z","iopub.status.idle":"2024-03-27T16:28:12.104786Z","shell.execute_reply.started":"2024-03-27T16:27:22.775412Z","shell.execute_reply":"2024-03-27T16:28:12.103932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize Training History\nplt.plot(history.history['accuracy'], label='Training Accuracy')\nplt.plot(history.history['val_accuracy'], label='Validation Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\nplt.show()\n\nplt.plot(history.history['loss'], label='Training Loss')\nplt.plot(history.history['val_loss'], label='Validation Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:28:12.106041Z","iopub.execute_input":"2024-03-27T16:28:12.106916Z","iopub.status.idle":"2024-03-27T16:28:12.703079Z","shell.execute_reply.started":"2024-03-27T16:28:12.106878Z","shell.execute_reply":"2024-03-27T16:28:12.701909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the model\nmodel.save(\"trained_model.h5\")","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:28:12.704325Z","iopub.execute_input":"2024-03-27T16:28:12.704660Z","iopub.status.idle":"2024-03-27T16:28:12.740218Z","shell.execute_reply.started":"2024-03-27T16:28:12.704633Z","shell.execute_reply":"2024-03-27T16:28:12.739124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model.predict(X_val)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:28:51.402616Z","iopub.execute_input":"2024-03-27T16:28:51.403016Z","iopub.status.idle":"2024-03-27T16:28:52.548030Z","shell.execute_reply.started":"2024-03-27T16:28:51.402975Z","shell.execute_reply":"2024-03-27T16:28:52.547147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.DataFrame(predictions, columns=[\"prediction\"])\nsubmission_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T16:28:55.835893Z","iopub.execute_input":"2024-03-27T16:28:55.836535Z","iopub.status.idle":"2024-03-27T16:28:55.883384Z","shell.execute_reply.started":"2024-03-27T16:28:55.836480Z","shell.execute_reply":"2024-03-27T16:28:55.882473Z"},"trusted":true},"execution_count":null,"outputs":[]}]}