{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7392733,"sourceType":"datasetVersion","datasetId":4297749},{"sourceId":7414022,"sourceType":"datasetVersion","datasetId":4312784},{"sourceId":7526248,"sourceType":"datasetVersion","datasetId":4308295},{"sourceId":7898593,"sourceType":"datasetVersion","datasetId":4638484},{"sourceId":8121200,"sourceType":"datasetVersion","datasetId":4798711},{"sourceId":6125,"sourceType":"modelInstanceVersion","modelInstanceId":4596},{"sourceId":6127,"sourceType":"modelInstanceVersion","modelInstanceId":4598}],"dockerImageVersionId":30698,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n<h1 style=\"color: blue;\"><strong> Harmful Brain Activity Classification </strong></h1>\n    <h2 style=\"color: blue;\">Harvard Medical School</h2> </div>\n\n\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n  <a href=\"https://imgur.com/J14fUpT\"><img src=\"https://i.imgur.com/J14fUpT.jpg\" title=\"source: imgur.com\" /></a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\"> <strong> Objectives and Goal</strong></h1></div>\n\n* To create a model that can automates EEG analysis will help doctors and brain researchers detect seizures and other types of brain activity that can cause brain damage, so that they can give treatments more quickly and accurately. \n* May also help researchers who are working to develop drugs to treat and prevent seizures.\n* To detect and classify seizures and other types of harmful brain activity in critical ill patients.\n* To develop a model trained on [EEG (Electroencephalography)](https://en.wikipedia.org/wiki/Electroencephalography) signals recorded from critically ill hospital patients.\n\n","metadata":{}},{"cell_type":"markdown","source":"## Patterns for Classification:\n1. Seizure (SZ)\n2. Generalized periodic discharges (GPD)\n3. Lateralized periodic discharges (LPD)\n4. Lateralized rhythmic delta activity (LRDA)\n5. Generalized rhythmic delta activity (GRDA)\n6. Other\n\n##### Detailed explanation of above patterns:\n* https://www.acns.org/UserFiles/file/ACNSStandardizedCriticalCareEEGTerminology_rev2021.pdf \n\n\n* In some cases experts completely agree about the correct label. On other cases the experts disagree. They call segments where there are high levels of agreement “idealized” patterns. Cases where ~1/2 of experts give a label as “other” and ~1/2 give one of the remaining five labels, we call “proto patterns”. Cases where experts are approximately split between 2 of the 5 named patterns, we call “edge cases”.","metadata":{"execution":{"iopub.status.busy":"2024-04-15T01:53:01.194450Z","iopub.execute_input":"2024-04-15T01:53:01.195051Z","iopub.status.idle":"2024-04-15T01:53:01.204939Z","shell.execute_reply.started":"2024-04-15T01:53:01.195021Z","shell.execute_reply":"2024-04-15T01:53:01.203625Z"}}},{"cell_type":"markdown","source":"1. High level of agreement                        - Idealized pattern\n2. Half of experts give a label                    - Other\n3. Half of experts give one of the remaining five label      - proto patterns\n4. Experts split between 2 of the 5 named patterns - edge cases","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n<h1 style=\"color: blue;\"> Dataset Flowchart </h1>\n</div>\n\n<a href=\"https://imgur.com/JuRILfJ\"><img src=\"https://i.imgur.com/JuRILfJ.jpg\" title=\"source: imgur.com\" /></a>","metadata":{"execution":{"iopub.status.busy":"2024-04-15T01:53:13.422210Z","iopub.execute_input":"2024-04-15T01:53:13.422580Z","iopub.status.idle":"2024-04-15T01:53:13.429019Z","shell.execute_reply.started":"2024-04-15T01:53:13.422551Z","shell.execute_reply":"2024-04-15T01:53:13.427145Z"}}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\">  <strong>Train Metadata </strong> </h1></div>\n ","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\"><a href=\"https://imgur.com/2PkCp9A\"><img src=\"https://i.imgur.com/2PkCp9A.jpg\" title=\"source: imgur.com\" /></a></div>","metadata":{"execution":{"iopub.status.busy":"2024-04-23T00:03:43.959023Z","iopub.execute_input":"2024-04-23T00:03:43.959835Z","iopub.status.idle":"2024-04-23T00:03:44.379349Z","shell.execute_reply.started":"2024-04-23T00:03:43.959794Z","shell.execute_reply":"2024-04-23T00:03:44.377618Z"}}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\">  <strong> Spectrogram Data  </strong> </h1>\n<a href=\"https://imgur.com/G5R4xMt\"><img src=\"https://i.imgur.com/G5R4xMt.jpg\" title=\"source: imgur.com\" /></a>\n<a href=\"https://imgur.com/azRNvZv\"><img src=\"https://i.imgur.com/azRNvZv.jpg\" title=\"source: imgur.com\" /></a></div>","metadata":{"execution":{"iopub.status.busy":"2024-04-21T21:08:37.246516Z","iopub.execute_input":"2024-04-21T21:08:37.247430Z","iopub.status.idle":"2024-04-21T21:08:37.268175Z","shell.execute_reply.started":"2024-04-21T21:08:37.247395Z","shell.execute_reply":"2024-04-21T21:08:37.267284Z"}}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\">  <strong> EEG Data  </strong> </h1>\n    <a href=\"https://imgur.com/m1XQpNX\"><img src=\"https://i.imgur.com/m1XQpNX.jpg\" title=\"source: imgur.com\" /></a>\n    <a href=\"https://imgur.com/gwa6RqI\"><img src=\"https://i.imgur.com/gwa6RqI.jpg\" title=\"source: imgur.com\" /></a></div>","metadata":{}},{"cell_type":"markdown","source":"<div style = \"text-align: center;\">\n    <a href=\"https://imgur.com/vyZ6e0x\"><img src=\"https://i.imgur.com/vyZ6e0x.jpg\" title=\"source: imgur.com\" /></a>\n   </div>","metadata":{}},{"cell_type":"markdown","source":"\n<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\">  <strong> Spectrograms  </strong> </h1>\n<a href=\"https://imgur.com/ELn7UOW\"><img src=\"https://i.imgur.com/ELn7UOW.jpg\" title=\"source: imgur.com\" /></a></div>","metadata":{}},{"cell_type":"markdown","source":"\n\n<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\">  <strong> Feature Engineering to Understand the Data  </strong> </h1></div>\n\n In some cases experts completely agree about the correct label. On other cases the experts disagree. They call segments where there are high levels of agreement “idealized” patterns. Cases where ~1/2 of experts give a label as “other” and ~1/2 give one of the remaining five labels, we call “proto patterns”. Cases where experts are approximately split between 2 of the 5 named patterns, we call “edge cases”.\nadd Codeadd Markdown\n* High level of agreement - Idealized pattern\n* half of the experts assign the label \"other\" and the remaining half assign one of the five named patterns -  proto patterns\n* Half of experts give one of the remaining five label - proto patterns\n* Experts split between 2 of the 5 named patterns - edge cases\n\n<div style=\"text-align:center;\">\n<a href=\"https://imgur.com/MqLoFfW\"><img src=\"https://i.imgur.com/MqLoFfW.jpg\" title=\"source: imgur.com\" /></a>\n    </div>\n    \n ### This explains how experts analyze EEG signals. Most patterns are Idealized, then Proto, and finally Edge. Some patterns remain unidentified. Understanding this helps interpret EEG data accurately for the model.","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:58:20.306062Z","iopub.execute_input":"2024-04-23T01:58:20.306481Z","iopub.status.idle":"2024-04-23T01:58:20.317903Z","shell.execute_reply.started":"2024-04-23T01:58:20.306448Z","shell.execute_reply":"2024-04-23T01:58:20.316106Z"}}},{"cell_type":"markdown","source":"\n\n<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\">  <strong> Class Distribution  </strong> </h1>\n    <a href=\"https://imgur.com/K9jbfYf\"><img src=\"https://i.imgur.com/K9jbfYf.jpg\" title=\"source: imgur.com\" /></a></div>\n    \n    \n   \n    \n   ### In the class distribution pie chart, the \"other\" category accounts for 42% of the dataset, while seizures make up 10%, LPD 16%, GPD 11%, LRDA 7.2%, and GRDA 9.4%. This chart helps us understand the proportion of different patterns in the dataset.\n\n### For a CNN model, this information is crucial because it guides how the model is trained to recognize different patterns. By knowing the distribution of classes, the model can be trained more effectively to accurately identify seizures and other brain activities.","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:19:52.702198Z","iopub.execute_input":"2024-04-23T02:19:52.702664Z","iopub.status.idle":"2024-04-23T02:19:52.710306Z","shell.execute_reply.started":"2024-04-23T02:19:52.702632Z","shell.execute_reply":"2024-04-23T02:19:52.708786Z"}}},{"cell_type":"markdown","source":"\n<div style=\"text-align:center;\">\n    <h1 style=\"color: blue;\">  <strong> Handle Imbalanced Data </strong> </h1>\n      </div>\n\n   <div style=\"text-align:center;\"> <a href=\"https://imgur.com/Jy90yQ2\"><img src=\"https://i.imgur.com/Jy90yQ2.jpg\" title=\"source: imgur.com\" /></a>\n  </div>\n  \n  \n### To address the imbalance in classification, I utilize stratified group k-fold cross-validation ***StratifiedGroupKFold*** from ***sklearn.model_selection***. This method ensures that each fold of the data contains a proportional representation of each class, helping the model to learn from all classes equally.","metadata":{}},{"cell_type":"markdown","source":"\n## Other Techniques Used:\n     To address class imbalance, I implemented Stratified Group K-Fold cross-validation, ensuring fair representation of each class in the training process.\n\n## Data Augmentation:\n     I applied augmentation techniques such as RandomCutout, RandomFlip, and MixUp to enhance the training data. However, I limited augmentation to the training set only, preserving the originality of patterns. After experimenting with extensive augmentation, I found that reducing some techniques improved model performance.","metadata":{"execution":{"iopub.status.busy":"2024-04-23T11:56:07.848047Z","iopub.execute_input":"2024-04-23T11:56:07.848452Z","iopub.status.idle":"2024-04-23T11:56:07.944728Z","shell.execute_reply.started":"2024-04-23T11:56:07.848421Z","shell.execute_reply":"2024-04-23T11:56:07.942991Z"}}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\" >\n  <h1 style = \"color: blue;\">Models Architecture</h1>\n    </div>\n    \n \n  <h1 style = \"color: blue;\">List of Models:</h1>\n  <ul>\n    <li>EfficientNetV2-B2</li>\n    <li>ResNet50</li>\n    <li>EfficientNetV2-B0</li>\n  </ul>\n<a href=\"https://imgur.com/7YBE8Wd\"><img src=\"https://i.imgur.com/7YBE8Wd.jpg\" title=\"source: imgur.com\" /></a>\n\n\n* Global Pooling 2D: A pooling operation that reduces the spatial dimensions of the input by taking the maximum or average value over each feature map.\n* Dense Layer: A fully connected layer where each neuron is connected to every neuron in the previous layer.\n* Kernel Regularization L2 0.01: Regularization technique used to prevent overfitting by penalizing large weights.\n* KL Divergence: A measure of how one probability distribution diverges from a second, expected probability distribution, commonly used in probabilistic models.\n* Adam Optimizer: An optimization algorithm used for training deep learning models, known for its adaptive learning rate and momentum.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n  <a href=\"https://imgur.com/LBRa958\">\n    <img src=\"https://i.imgur.com/LBRa958.jpg\" title=\"source: imgur.com\" />\n  </a>   \n</div>\n\n\n## Model 1: EfficientNetV2B2\n### Epochs 1-15:\n* Training Accuracy: Ranges from 31.45% to 71.70%\n* Training Loss: Decreases from 1.4558 to 0.6863\n* Validation Accuracy: Ranges from 46.57% to 68.36%\n* Validation Loss: Decreases from 1.4106 to 0.8878\n\n## Model 2: ResNet50\n### Epochs 1-15:\n* Training Accuracy: Ranges from 33.67% to 74.24%\n* Training Loss: Decreases from 1.7303 to 0.6641\n* Validation Accuracy: Ranges from 6.33% to 67.32%\n* Validation Loss: Increases from 2.0393 to 1.0153\n\n## Model 3: EfficientNetV2B0\n### Epochs 1-15:\n* Training Accuracy: Ranges from 31.76% to 71.03%\n* Training Loss: Decreases from 1.4545 to 0.7109\n* Validation Accuracy: Ranges from 46.57% to 66.98%\n* Validation Loss: Decreases from 1.4142 to 0.9212\n\n## Analysis:\n* All models show an improvement in both training and validation accuracy over the epochs.\n* Model 1 and Model 3 exhibit a steady decrease in both training and validation loss, indicating improved performance.\n* Model 2, however, shows erratic behavior with a significant increase in validation loss from epochs 1-15, which suggests overfitting.\n* Overall, Model 1 and Model 3 appear to perform better compared to Model 2, as they demonstrate a more consistent decrease in loss and increase in accuracy over the epochs.","metadata":{}},{"cell_type":"markdown","source":"\n\n\n<div style=\"text-align:center;\">\n    <h1 style = \"color: blue;\"> Classification Report of Ensemble Models</h1>\n  <a href=\"https://imgur.com/Qo1Dkaq\"><img src=\"https://i.imgur.com/Qo1Dkaq.jpg\" title=\"source: imgur.com\" /></a>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"### In this classification report, we're evaluating the performance of a model in recognizing different classes. Precision tells us how many of the instances predicted as belonging to a certain class actually belong to that class. Recall, on the other hand, indicates how many of the actual instances of a class were correctly predicted by the model. F1-score is a balance between precision and recall.\n\n### Looking at the report, we can see that the classes \"Seizure,\" \"LPD,\" \"GPD,\" and \"Other\" have relatively high precision scores, indicating that when the model predicts these classes, it's usually correct. However, when it comes to recall, which measures how well the model is capturing instances of each class, \"Other\" stands out with the highest score, indicating that the model is better at correctly identifying instances of this class compared to the others.\n\n### In simpler terms, while the model performs well overall, it's particularly good at recognizing instances of the \"Other\" class. This might mean that the data for this class is more distinguishable or that the model has been trained more effectively on this class.","metadata":{}},{"cell_type":"markdown","source":"\n<div style=\"text-align:center;\">\n    <h1 style = \"color: blue;\"><strong> Test Data </strong></h1>\n\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n\n  <a href=\"https://imgur.com/BtiMs1r\">\n      <img src=\"https://i.imgur.com/BtiMs1r.jpg\" title=\"source: imgur.com\" />\n    </a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"\n<div style=\"text-align:center;\">\n    <h1 style = \"color: blue;\"> Spectrogram Image of Test Data</h1>\n\n\n<a href=\"https://imgur.com/3k1AFu9\">\n<img src=\"https://i.imgur.com/3k1AFu9.jpg\" title=\"source: imgur.com\" />\n</a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"\n<div style=\"text-align:center;\">\n    <h1 style = \"color: blue;\"> Predicted Probability on Test Data Using Average Ensemble Technique</h1>\n<a href=\"https://imgur.com/gw62K0R\"><img src=\"https://i.imgur.com/gw62K0R.jpg\" title=\"source: imgur.com\" /></a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"\n\n<div style=\"text-align:center;\">\n    <h1 style = \"color: blue;\"> My Submission Scores Table </h1>\n    <a href=\"https://imgur.com/WjaVNhp\"><img src=\"https://i.imgur.com/WjaVNhp.jpg\" title=\"source: imgur.com\"/></a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# Notebook Link:\n[Original Code of this notebook](https://www.kaggle.com/code/sameenamujawar/sm-harmful-brain-activity-classification)","metadata":{}}]}