{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Harmful Brain Activity Classification - Creating a Model with K-Nearest Neighbour**\n\n## **Written by:** Aarish Asif Khan\n\n## **Date:** 17th February 2024","metadata":{}},{"cell_type":"markdown","source":"# K-Nearest Neighbors (KNN)\n\nK-Nearest Neighbors (KNN) is a popular supervised machine learning algorithm used for classification and regression tasks. It is a simple and intuitive algorithm that classifies data points based on the majority class of their neighboring data points. \n\n## **How KNN Works**\n\n1. **Training Phase**:\n   - During the training phase, KNN stores all available data points and their corresponding class labels or continuous values.\n   \n2. **Prediction Phase**:\n   - For a given unlabeled data point, KNN calculates the distances between this point and all other points in the dataset.\n   - It then selects the K nearest data points (closest neighbors) based on a distance metric such as Euclidean distance, Manhattan distance, or Minkowski distance.\n   - Finally, KNN assigns the class label (for classification) or the average/median value (for regression) of the majority of its K nearest neighbors to the unlabeled point.\n\n## **Parameters of KNN**\n\n- **K**: The number of nearest neighbors to consider. This is a hyperparameter that needs to be chosen beforehand.\n\n- **Distance Metric**: The method used to measure the distance between data points, such as Euclidean, Manhattan, or Minkowski distance.\n\n- **Weight Function**: An optional parameter used to give more weight to closer neighbors while making predictions.\n\n## **Advantages**\n\n- Simple to implement and understand.\n- No assumptions about the underlying data distribution.\n- Effective for multi-class classification tasks.\n- Works well with small datasets.\n\n## **Disadvantages**\n\n- Computationally expensive during the prediction phase, especially with large datasets.\n- Sensitive to irrelevant features or noisy data.\n- Optimal choice of K is crucial and may require experimentation.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:24.524271Z","iopub.execute_input":"2024-03-09T16:53:24.524832Z","iopub.status.idle":"2024-03-09T16:53:27.895990Z","shell.execute_reply.started":"2024-03-09T16:53:24.524788Z","shell.execute_reply":"2024-03-09T16:53:27.894657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:27.898205Z","iopub.execute_input":"2024-03-09T16:53:27.898848Z","iopub.status.idle":"2024-03-09T16:53:28.238404Z","shell.execute_reply.started":"2024-03-09T16:53:27.898807Z","shell.execute_reply":"2024-03-09T16:53:28.237256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select relevant features and target variable\nX = train_data[['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds', 'patient_id']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:28.239715Z","iopub.execute_input":"2024-03-09T16:53:28.240053Z","iopub.status.idle":"2024-03-09T16:53:28.257322Z","shell.execute_reply.started":"2024-03-09T16:53:28.240023Z","shell.execute_reply":"2024-03-09T16:53:28.255860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:28.260544Z","iopub.execute_input":"2024-03-09T16:53:28.261504Z","iopub.status.idle":"2024-03-09T16:53:28.289988Z","shell.execute_reply.started":"2024-03-09T16:53:28.261458Z","shell.execute_reply":"2024-03-09T16:53:28.288583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize the KNN classifier\nknn = KNeighborsClassifier(n_neighbors=5)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:28.291591Z","iopub.execute_input":"2024-03-09T16:53:28.292179Z","iopub.status.idle":"2024-03-09T16:53:28.298273Z","shell.execute_reply.started":"2024-03-09T16:53:28.292128Z","shell.execute_reply":"2024-03-09T16:53:28.296838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier\nknn.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:28.300091Z","iopub.execute_input":"2024-03-09T16:53:28.300901Z","iopub.status.idle":"2024-03-09T16:53:28.643684Z","shell.execute_reply.started":"2024-03-09T16:53:28.300840Z","shell.execute_reply":"2024-03-09T16:53:28.642189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions\npredictions = knn.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:28.645410Z","iopub.execute_input":"2024-03-09T16:53:28.645904Z","iopub.status.idle":"2024-03-09T16:53:30.136756Z","shell.execute_reply.started":"2024-03-09T16:53:28.645863Z","shell.execute_reply":"2024-03-09T16:53:30.135153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate accuracy\naccuracy = accuracy_score(y_test, predictions)\nprint(\"Accuracy:\", accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.138553Z","iopub.execute_input":"2024-03-09T16:53:30.138980Z","iopub.status.idle":"2024-03-09T16:53:30.203691Z","shell.execute_reply.started":"2024-03-09T16:53:30.138940Z","shell.execute_reply":"2024-03-09T16:53:30.202560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"An accuracy of 0.7686 (76.86%) is decent, but whether it's considered good or moderate depends on the specific context of your problem and the desired level of performance.\n\n#### **Improving the Accuracy of the model by the steps below:**\n\n# **Feature Engineering:**\n\nExplore and engineer new features from the existing ones that might provide more information to the model. This could involve transforming existing features, creating interaction terms, or incorporating domain knowledge.\n\n# **Hyperparameter Tuning:** \n\nExperiment with different values of hyperparameters for the KNN algorithm, such as the number of neighbors (n_neighbors), the distance metric, and the weight function. Grid search or random search can be used to find the optimal hyperparameters.\n\n# **Data Preprocessing:** \n\nEnsure that the data is preprocessed appropriately before training the model. This may involve handling missing values, scaling numerical features, encoding categorical variables, and dealing with outliers.","metadata":{}},{"cell_type":"markdown","source":"1. # **Feature Engineering**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.205574Z","iopub.execute_input":"2024-03-09T16:53:30.205935Z","iopub.status.idle":"2024-03-09T16:53:30.211382Z","shell.execute_reply.started":"2024-03-09T16:53:30.205893Z","shell.execute_reply":"2024-03-09T16:53:30.210074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the data\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.216836Z","iopub.execute_input":"2024-03-09T16:53:30.217989Z","iopub.status.idle":"2024-03-09T16:53:30.447308Z","shell.execute_reply.started":"2024-03-09T16:53:30.217950Z","shell.execute_reply":"2024-03-09T16:53:30.445859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a new feature, Total duration\ntrain_data['total_duration'] = train_data['eeg_label_offset_seconds'] + train_data['spectrogram_label_offset_seconds']","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.448850Z","iopub.execute_input":"2024-03-09T16:53:30.449324Z","iopub.status.idle":"2024-03-09T16:53:30.457656Z","shell.execute_reply.started":"2024-03-09T16:53:30.449264Z","shell.execute_reply":"2024-03-09T16:53:30.456396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transforming existing features: Scaling numerical features\nscaler = StandardScaler()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.459140Z","iopub.execute_input":"2024-03-09T16:53:30.459591Z","iopub.status.idle":"2024-03-09T16:53:30.477573Z","shell.execute_reply.started":"2024-03-09T16:53:30.459560Z","shell.execute_reply":"2024-03-09T16:53:30.476364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['eeg_label_offset_seconds_scaled'] = scaler.fit_transform(train_data[['eeg_label_offset_seconds']])","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.479145Z","iopub.execute_input":"2024-03-09T16:53:30.479566Z","iopub.status.idle":"2024-03-09T16:53:30.503727Z","shell.execute_reply.started":"2024-03-09T16:53:30.479537Z","shell.execute_reply":"2024-03-09T16:53:30.502357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Encoding categorical variables: One-hot encoding for 'patient_id'\ntrain_data = pd.get_dummies(train_data, columns=['patient_id'])","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.505517Z","iopub.execute_input":"2024-03-09T16:53:30.505889Z","iopub.status.idle":"2024-03-09T16:53:30.747267Z","shell.execute_reply.started":"2024-03-09T16:53:30.505853Z","shell.execute_reply":"2024-03-09T16:53:30.746297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['interaction_term'] = train_data['eeg_label_offset_seconds'] * train_data['spectrogram_label_offset_seconds']","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.748490Z","iopub.execute_input":"2024-03-09T16:53:30.749324Z","iopub.status.idle":"2024-03-09T16:53:30.755691Z","shell.execute_reply.started":"2024-03-09T16:53:30.749263Z","shell.execute_reply":"2024-03-09T16:53:30.754478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop original columns if necessary\ntrain_data.drop(['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.757173Z","iopub.execute_input":"2024-03-09T16:53:30.757539Z","iopub.status.idle":"2024-03-09T16:53:30.865923Z","shell.execute_reply.started":"2024-03-09T16:53:30.757508Z","shell.execute_reply":"2024-03-09T16:53:30.864400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. # **Hyperparameter Tuning**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split, GridSearchCV\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.867302Z","iopub.execute_input":"2024-03-09T16:53:30.867638Z","iopub.status.idle":"2024-03-09T16:53:30.873402Z","shell.execute_reply.started":"2024-03-09T16:53:30.867611Z","shell.execute_reply":"2024-03-09T16:53:30.872219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the data\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:30.875024Z","iopub.execute_input":"2024-03-09T16:53:30.875478Z","iopub.status.idle":"2024-03-09T16:53:31.101118Z","shell.execute_reply.started":"2024-03-09T16:53:30.875440Z","shell.execute_reply":"2024-03-09T16:53:31.100015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data[['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds', 'patient_id']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:31.104184Z","iopub.execute_input":"2024-03-09T16:53:31.104680Z","iopub.status.idle":"2024-03-09T16:53:31.114508Z","shell.execute_reply.started":"2024-03-09T16:53:31.104640Z","shell.execute_reply":"2024-03-09T16:53:31.113110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:31.116497Z","iopub.execute_input":"2024-03-09T16:53:31.116974Z","iopub.status.idle":"2024-03-09T16:53:31.141172Z","shell.execute_reply.started":"2024-03-09T16:53:31.116926Z","shell.execute_reply":"2024-03-09T16:53:31.139789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the hyperparameters to tune\nparam_grid = {\n    'n_neighbors': [3, 5, 7, 9],  # Number of neighbors\n    'weights': ['uniform', 'distance'],  # Weight function\n    'metric': ['euclidean', 'manhattan']  # Distance metric\n}","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:31.144609Z","iopub.execute_input":"2024-03-09T16:53:31.145020Z","iopub.status.idle":"2024-03-09T16:53:31.150195Z","shell.execute_reply.started":"2024-03-09T16:53:31.144977Z","shell.execute_reply":"2024-03-09T16:53:31.149189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize the KNN classifier\nknn = KNeighborsClassifier()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:31.151160Z","iopub.execute_input":"2024-03-09T16:53:31.151504Z","iopub.status.idle":"2024-03-09T16:53:31.163742Z","shell.execute_reply.started":"2024-03-09T16:53:31.151475Z","shell.execute_reply":"2024-03-09T16:53:31.162653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Perform grid search with cross-validation\ngrid_search = GridSearchCV(knn, param_grid, cv=5, scoring='accuracy')\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:53:31.165044Z","iopub.execute_input":"2024-03-09T16:53:31.165396Z","iopub.status.idle":"2024-03-09T16:54:41.838106Z","shell.execute_reply.started":"2024-03-09T16:53:31.165367Z","shell.execute_reply":"2024-03-09T16:54:41.836736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the best hyperparameters\nbest_params = grid_search.best_params_\nprint(\"Best Hyperparameters:\", best_params)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:41.839766Z","iopub.execute_input":"2024-03-09T16:54:41.840256Z","iopub.status.idle":"2024-03-09T16:54:41.846274Z","shell.execute_reply.started":"2024-03-09T16:54:41.840221Z","shell.execute_reply":"2024-03-09T16:54:41.845090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the best model\nbest_model = grid_search.best_estimator_\n\n# Make predictions\npredictions = best_model.predict(X_test)\n\n# Calculate accuracy\naccuracy = accuracy_score(y_test, predictions)\nprint(\"Accuracy:\", accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:41.847858Z","iopub.execute_input":"2024-03-09T16:54:41.848637Z","iopub.status.idle":"2024-03-09T16:54:42.026019Z","shell.execute_reply.started":"2024-03-09T16:54:41.848593Z","shell.execute_reply":"2024-03-09T16:54:42.024681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. # **Data Pre-processing**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.027582Z","iopub.execute_input":"2024-03-09T16:54:42.028496Z","iopub.status.idle":"2024-03-09T16:54:42.033302Z","shell.execute_reply.started":"2024-03-09T16:54:42.028443Z","shell.execute_reply":"2024-03-09T16:54:42.032043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.034682Z","iopub.execute_input":"2024-03-09T16:54:42.035018Z","iopub.status.idle":"2024-03-09T16:54:42.241297Z","shell.execute_reply.started":"2024-03-09T16:54:42.034990Z","shell.execute_reply":"2024-03-09T16:54:42.239757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data[['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds', 'patient_id']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.250195Z","iopub.execute_input":"2024-03-09T16:54:42.250624Z","iopub.status.idle":"2024-03-09T16:54:42.258460Z","shell.execute_reply.started":"2024-03-09T16:54:42.250591Z","shell.execute_reply":"2024-03-09T16:54:42.257083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.260365Z","iopub.execute_input":"2024-03-09T16:54:42.260779Z","iopub.status.idle":"2024-03-09T16:54:42.286052Z","shell.execute_reply.started":"2024-03-09T16:54:42.260734Z","shell.execute_reply":"2024-03-09T16:54:42.284496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Data preprocessing steps\n\n# 1. Handling missing values (if any)\nprint(\"Missing values before preprocessing:\")\nprint(X_train.isnull().sum())\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.288394Z","iopub.execute_input":"2024-03-09T16:54:42.288780Z","iopub.status.idle":"2024-03-09T16:54:42.297716Z","shell.execute_reply.started":"2024-03-09T16:54:42.288748Z","shell.execute_reply":"2024-03-09T16:54:42.296539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 2. Scaling numerical features\nscaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.298841Z","iopub.execute_input":"2024-03-09T16:54:42.299193Z","iopub.status.idle":"2024-03-09T16:54:42.324805Z","shell.execute_reply.started":"2024-03-09T16:54:42.299165Z","shell.execute_reply":"2024-03-09T16:54:42.323706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_encoded = pd.get_dummies(X_train, columns=['patient_id'])\nX_test_encoded = pd.get_dummies(X_test, columns=['patient_id'])","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.326577Z","iopub.execute_input":"2024-03-09T16:54:42.327095Z","iopub.status.idle":"2024-03-09T16:54:42.548326Z","shell.execute_reply.started":"2024-03-09T16:54:42.327052Z","shell.execute_reply":"2024-03-09T16:54:42.547134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now 'X_train_encoded' and 'X_test_encoded' contain the preprocessed features","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.549816Z","iopub.execute_input":"2024-03-09T16:54:42.550297Z","iopub.status.idle":"2024-03-09T16:54:42.555492Z","shell.execute_reply.started":"2024-03-09T16:54:42.550257Z","shell.execute_reply":"2024-03-09T16:54:42.553907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Visualizing the Data distribution**","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Define custom colors for each bar\ncustom_colors = [\"#1f77b4\", \"#ff7f0e\", \"#2ca02c\", \"#d62728\", \"#9467bd\", \"#8c564b\"]\n\n# Visualize distribution of target variable\nplt.figure(figsize=(8, 6))\nsns.countplot(x='expert_consensus', data=train_data, palette=custom_colors)\nplt.title('Distribution of Expert Consensus Labels')\nplt.xlabel('Expert Consensus Labels')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:42.556902Z","iopub.execute_input":"2024-03-09T16:54:42.557221Z","iopub.status.idle":"2024-03-09T16:54:43.290724Z","shell.execute_reply.started":"2024-03-09T16:54:42.557195Z","shell.execute_reply":"2024-03-09T16:54:43.289217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Visualizing the Model performance**","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import confusion_matrix\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Assuming you have initialized and trained your KNeighborsClassifier\nknn = KNeighborsClassifier()\nknn.fit(X_train, y_train)  # Make sure you have fitted your classifier\n\n# Make predictions\npredictions = knn.predict(X_test)\n\n# Calculate confusion matrix\ncm = confusion_matrix(y_test, predictions)\n\n# Visualize confusion matrix\nplt.figure(figsize=(8, 6))\nsns.heatmap(cm, annot=True, cmap='Blues', fmt='g', \n            xticklabels=knn.classes_, yticklabels=knn.classes_)\nplt.title('Confusion Matrix')\nplt.xlabel('Predicted Labels')\nplt.ylabel('True Labels')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:43.292643Z","iopub.execute_input":"2024-03-09T16:54:43.293454Z","iopub.status.idle":"2024-03-09T16:54:45.603443Z","shell.execute_reply.started":"2024-03-09T16:54:43.293416Z","shell.execute_reply":"2024-03-09T16:54:45.602156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Creating a model with Decision-Tree Algorithm**\n\n## **Date:** 17 February 2024","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:45.605205Z","iopub.execute_input":"2024-03-09T16:54:45.605696Z","iopub.status.idle":"2024-03-09T16:54:45.651643Z","shell.execute_reply.started":"2024-03-09T16:54:45.605656Z","shell.execute_reply":"2024-03-09T16:54:45.650340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:45.653354Z","iopub.execute_input":"2024-03-09T16:54:45.653846Z","iopub.status.idle":"2024-03-09T16:54:45.853799Z","shell.execute_reply.started":"2024-03-09T16:54:45.653794Z","shell.execute_reply":"2024-03-09T16:54:45.852386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select features and target variable\nX = train_data[['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:45.855378Z","iopub.execute_input":"2024-03-09T16:54:45.855834Z","iopub.status.idle":"2024-03-09T16:54:45.866101Z","shell.execute_reply.started":"2024-03-09T16:54:45.855790Z","shell.execute_reply":"2024-03-09T16:54:45.864459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:45.867541Z","iopub.execute_input":"2024-03-09T16:54:45.868405Z","iopub.status.idle":"2024-03-09T16:54:45.894742Z","shell.execute_reply.started":"2024-03-09T16:54:45.868358Z","shell.execute_reply":"2024-03-09T16:54:45.893343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the decision tree classifier\ndt_classifier = DecisionTreeClassifier()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:45.896137Z","iopub.execute_input":"2024-03-09T16:54:45.896476Z","iopub.status.idle":"2024-03-09T16:54:45.901441Z","shell.execute_reply.started":"2024-03-09T16:54:45.896449Z","shell.execute_reply":"2024-03-09T16:54:45.900003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier on the training data\ndt_classifier.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:45.902748Z","iopub.execute_input":"2024-03-09T16:54:45.903178Z","iopub.status.idle":"2024-03-09T16:54:46.419791Z","shell.execute_reply.started":"2024-03-09T16:54:45.903140Z","shell.execute_reply":"2024-03-09T16:54:46.418476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Evaluate the performance of the trained model on the testing data\npredictions = dt_classifier.predict(X_test)\n\naccuracy = accuracy_score(y_test, predictions)\nprint(\"Accuracy:\", accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:46.422362Z","iopub.execute_input":"2024-03-09T16:54:46.422719Z","iopub.status.idle":"2024-03-09T16:54:46.486187Z","shell.execute_reply.started":"2024-03-09T16:54:46.422692Z","shell.execute_reply":"2024-03-09T16:54:46.484730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.metrics import confusion_matrix","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:46.487792Z","iopub.execute_input":"2024-03-09T16:54:46.488297Z","iopub.status.idle":"2024-03-09T16:54:46.493906Z","shell.execute_reply.started":"2024-03-09T16:54:46.488233Z","shell.execute_reply":"2024-03-09T16:54:46.493038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Bar chart for the distribution of the target variable (expert_consensus)\nplt.figure(figsize=(8, 6))\nsns.countplot(x='expert_consensus', data=train_data)\nplt.title('Distribution of Expert Consensus')\nplt.xlabel('Expert Consensus')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:46.495405Z","iopub.execute_input":"2024-03-09T16:54:46.495779Z","iopub.status.idle":"2024-03-09T16:54:46.937711Z","shell.execute_reply.started":"2024-03-09T16:54:46.495725Z","shell.execute_reply":"2024-03-09T16:54:46.936350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Confusion matrix to visualize the performance of the classifier\nconf_matrix = confusion_matrix(y_test, predictions)\nplt.figure(figsize=(8, 6))\nsns.heatmap(conf_matrix, annot=True, cmap='Blues', fmt='g')\nplt.title('Confusion Matrix')\nplt.xlabel('Predicted Labels')\nplt.ylabel('True Labels')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:46.939302Z","iopub.execute_input":"2024-03-09T16:54:46.940554Z","iopub.status.idle":"2024-03-09T16:54:47.555542Z","shell.execute_reply.started":"2024-03-09T16:54:46.940515Z","shell.execute_reply":"2024-03-09T16:54:47.554245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Applying Entropy, Information gain and Gini Impurity**","metadata":{}},{"cell_type":"markdown","source":"1. # **Entropy**","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:47.557215Z","iopub.execute_input":"2024-03-09T16:54:47.557613Z","iopub.status.idle":"2024-03-09T16:54:47.563656Z","shell.execute_reply.started":"2024-03-09T16:54:47.557583Z","shell.execute_reply":"2024-03-09T16:54:47.562416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:47.564945Z","iopub.execute_input":"2024-03-09T16:54:47.565269Z","iopub.status.idle":"2024-03-09T16:54:47.771570Z","shell.execute_reply.started":"2024-03-09T16:54:47.565243Z","shell.execute_reply":"2024-03-09T16:54:47.770079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data.drop('expert_consensus', axis=1)  # Features\ny = train_data['expert_consensus']  # Target variable","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:47.773981Z","iopub.execute_input":"2024-03-09T16:54:47.774468Z","iopub.status.idle":"2024-03-09T16:54:47.783642Z","shell.execute_reply.started":"2024-03-09T16:54:47.774435Z","shell.execute_reply":"2024-03-09T16:54:47.782190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:47.786228Z","iopub.execute_input":"2024-03-09T16:54:47.786684Z","iopub.status.idle":"2024-03-09T16:54:47.822249Z","shell.execute_reply.started":"2024-03-09T16:54:47.786651Z","shell.execute_reply":"2024-03-09T16:54:47.820997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the decision tree classifier with entropy as criterion\ndt_classifier = DecisionTreeClassifier(criterion='entropy')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:47.823635Z","iopub.execute_input":"2024-03-09T16:54:47.823978Z","iopub.status.idle":"2024-03-09T16:54:47.830530Z","shell.execute_reply.started":"2024-03-09T16:54:47.823950Z","shell.execute_reply":"2024-03-09T16:54:47.829200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier on the training data\ndt_classifier.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:47.831769Z","iopub.execute_input":"2024-03-09T16:54:47.832710Z","iopub.status.idle":"2024-03-09T16:54:48.818042Z","shell.execute_reply.started":"2024-03-09T16:54:47.832661Z","shell.execute_reply":"2024-03-09T16:54:48.817085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions on the testing data\npredictions = dt_classifier.predict(X_test)\n\n# Calculate the accuracy of the model\naccuracy_entropy = accuracy_score(y_test, predictions)\nprint(\"Accuracy:\", accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:48.819407Z","iopub.execute_input":"2024-03-09T16:54:48.819965Z","iopub.status.idle":"2024-03-09T16:54:48.879059Z","shell.execute_reply.started":"2024-03-09T16:54:48.819933Z","shell.execute_reply":"2024-03-09T16:54:48.878103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. # **Information gain**","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:48.880374Z","iopub.execute_input":"2024-03-09T16:54:48.880934Z","iopub.status.idle":"2024-03-09T16:54:48.885808Z","shell.execute_reply.started":"2024-03-09T16:54:48.880903Z","shell.execute_reply":"2024-03-09T16:54:48.884370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:48.887472Z","iopub.execute_input":"2024-03-09T16:54:48.888308Z","iopub.status.idle":"2024-03-09T16:54:49.098296Z","shell.execute_reply.started":"2024-03-09T16:54:48.888238Z","shell.execute_reply":"2024-03-09T16:54:49.097125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data.drop('expert_consensus', axis=1)  # Features\ny = train_data['expert_consensus']  # Target variable","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:49.099576Z","iopub.execute_input":"2024-03-09T16:54:49.101984Z","iopub.status.idle":"2024-03-09T16:54:49.110786Z","shell.execute_reply.started":"2024-03-09T16:54:49.101947Z","shell.execute_reply":"2024-03-09T16:54:49.109548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:49.112633Z","iopub.execute_input":"2024-03-09T16:54:49.113112Z","iopub.status.idle":"2024-03-09T16:54:49.146212Z","shell.execute_reply.started":"2024-03-09T16:54:49.113070Z","shell.execute_reply":"2024-03-09T16:54:49.144773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the decision tree classifier with information gain as criterion\ndt_classifier = DecisionTreeClassifier(criterion='entropy')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:49.147794Z","iopub.execute_input":"2024-03-09T16:54:49.148296Z","iopub.status.idle":"2024-03-09T16:54:49.154351Z","shell.execute_reply.started":"2024-03-09T16:54:49.148238Z","shell.execute_reply":"2024-03-09T16:54:49.152750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier on the training data\ndt_classifier.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:49.156083Z","iopub.execute_input":"2024-03-09T16:54:49.156553Z","iopub.status.idle":"2024-03-09T16:54:50.115202Z","shell.execute_reply.started":"2024-03-09T16:54:49.156514Z","shell.execute_reply":"2024-03-09T16:54:50.113727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions on the testing data\npredictions = dt_classifier.predict(X_test)\n\n# Calculate the accuracy of the model\naccuracy_information_gain = accuracy_score(y_test, predictions)\nprint(\"Accuracy:\", accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:50.117090Z","iopub.execute_input":"2024-03-09T16:54:50.117553Z","iopub.status.idle":"2024-03-09T16:54:50.177975Z","shell.execute_reply.started":"2024-03-09T16:54:50.117521Z","shell.execute_reply":"2024-03-09T16:54:50.176693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. # **Gini Impurity**","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:50.179244Z","iopub.execute_input":"2024-03-09T16:54:50.179584Z","iopub.status.idle":"2024-03-09T16:54:50.188440Z","shell.execute_reply.started":"2024-03-09T16:54:50.179557Z","shell.execute_reply":"2024-03-09T16:54:50.186654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ndata = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:50.190137Z","iopub.execute_input":"2024-03-09T16:54:50.190560Z","iopub.status.idle":"2024-03-09T16:54:50.384853Z","shell.execute_reply.started":"2024-03-09T16:54:50.190519Z","shell.execute_reply":"2024-03-09T16:54:50.383404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = data.drop('expert_consensus', axis=1)  # Features\ny = data['expert_consensus']  # Target variable","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:50.386130Z","iopub.execute_input":"2024-03-09T16:54:50.386483Z","iopub.status.idle":"2024-03-09T16:54:50.395029Z","shell.execute_reply.started":"2024-03-09T16:54:50.386454Z","shell.execute_reply":"2024-03-09T16:54:50.393779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:50.396707Z","iopub.execute_input":"2024-03-09T16:54:50.397376Z","iopub.status.idle":"2024-03-09T16:54:50.427327Z","shell.execute_reply.started":"2024-03-09T16:54:50.397329Z","shell.execute_reply":"2024-03-09T16:54:50.425490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the decision tree classifier with Gini impurity as criterion\ndt_classifier = DecisionTreeClassifier(criterion='gini')","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:50.429057Z","iopub.execute_input":"2024-03-09T16:54:50.429639Z","iopub.status.idle":"2024-03-09T16:54:50.436428Z","shell.execute_reply.started":"2024-03-09T16:54:50.429595Z","shell.execute_reply":"2024-03-09T16:54:50.435210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier on the training data\ndt_classifier.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:50.438419Z","iopub.execute_input":"2024-03-09T16:54:50.439236Z","iopub.status.idle":"2024-03-09T16:54:51.439264Z","shell.execute_reply.started":"2024-03-09T16:54:50.439194Z","shell.execute_reply":"2024-03-09T16:54:51.437682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions on the testing data\npredictions = dt_classifier.predict(X_test)\n\n# Calculate the accuracy of the model\naccuracy_gini = accuracy_score(y_test, predictions)\nprint(\"Accuracy:\", accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:51.440983Z","iopub.execute_input":"2024-03-09T16:54:51.441512Z","iopub.status.idle":"2024-03-09T16:54:51.501950Z","shell.execute_reply.started":"2024-03-09T16:54:51.441471Z","shell.execute_reply":"2024-03-09T16:54:51.501048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define colors for the bars\ncolors = ['purple', 'yellow', 'cyan']\n\n# Plot the accuracies with custom colors\nplt.bar(['Entropy', 'Information Gain', 'Gini Impurity'], [accuracy_entropy, accuracy_information_gain, accuracy_gini],\n        color=colors)\nplt.title('Accuracy of Decision Tree with Different Splitting Criteria')\nplt.xlabel('Splitting Criterion')\nplt.ylabel('Accuracy')\nplt.ylim(0, 1)  # Set y-axis limit to range between 0 and 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-09T16:54:51.503194Z","iopub.execute_input":"2024-03-09T16:54:51.503993Z","iopub.status.idle":"2024-03-09T16:54:51.763294Z","shell.execute_reply.started":"2024-03-09T16:54:51.503952Z","shell.execute_reply":"2024-03-09T16:54:51.761998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **What the above Visualization shows:**","metadata":{}},{"cell_type":"markdown","source":"This code snippet defines custom colors for the bars in the bar chart and then plots the accuracies of decision trees with different splitting criteria (entropy, information gain, and Gini impurity) using these custom colors, purple, yellow and cyan.\n\n","metadata":{}}]}