{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Harmful Brain Activity Classification - HMS Dataset**\n\n## **What we will do in this Notebook!**\n\n> **- Removing Outliers by IQR, Z-Score & K-Means Clustering**\n\n> **- Scaling to Pre-process the data**\n\n> **- Applying Transformations etc**\n\n\n## **Project by:** [Aarish Asif Khan](https://www.kaggle.com/aarishasifkhan)\n\n## **Date:** 7th February 2024\n\n## **Dataset:** [HMS - Harmful Brain Activity Dataset](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification)","metadata":{}},{"cell_type":"markdown","source":"# **Outliers**\n\nOutliers are data points that significantly differ from other observations in a dataset. They can occur due to various reasons such as measurement errors, experimental variability, or natural variance in the phenomenon being studied.\n\nOutliers can greatly affect statistical analyses and machine learning algorithms by skewing results and distorting interpretations if not appropriately handled.","metadata":{}},{"cell_type":"markdown","source":"# **IQR method**","metadata":{}},{"cell_type":"markdown","source":"1. ### **Import the libraries**","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport numpy as np ","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.004884Z","iopub.execute_input":"2024-03-18T08:38:49.005318Z","iopub.status.idle":"2024-03-18T08:38:49.010622Z","shell.execute_reply.started":"2024-03-18T08:38:49.005288Z","shell.execute_reply":"2024-03-18T08:38:49.009258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. ### **Load the HMC dataset**","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\ntest_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.022496Z","iopub.execute_input":"2024-03-18T08:38:49.022906Z","iopub.status.idle":"2024-03-18T08:38:49.211091Z","shell.execute_reply.started":"2024-03-18T08:38:49.022879Z","shell.execute_reply":"2024-03-18T08:38:49.209807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.213041Z","iopub.execute_input":"2024-03-18T08:38:49.213443Z","iopub.status.idle":"2024-03-18T08:38:49.235035Z","shell.execute_reply.started":"2024-03-18T08:38:49.213413Z","shell.execute_reply":"2024-03-18T08:38:49.233556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. ### **Calculate Q1 & Q3**","metadata":{}},{"cell_type":"code","source":"# Calculate Q1 and Q3\nQ1 = train_data['eeg_label_offset_seconds'].quantile(0.25)\nQ3 = train_data['eeg_label_offset_seconds'].quantile(0.75)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.236864Z","iopub.execute_input":"2024-03-18T08:38:49.237998Z","iopub.status.idle":"2024-03-18T08:38:49.250076Z","shell.execute_reply.started":"2024-03-18T08:38:49.237928Z","shell.execute_reply":"2024-03-18T08:38:49.248681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"4. ### **Compile the IQR**","metadata":{}},{"cell_type":"code","source":"# Compute the IQR\nIQR = Q3 - Q1","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.253636Z","iopub.execute_input":"2024-03-18T08:38:49.254204Z","iopub.status.idle":"2024-03-18T08:38:49.260532Z","shell.execute_reply.started":"2024-03-18T08:38:49.254170Z","shell.execute_reply":"2024-03-18T08:38:49.258932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"5. ### **Define lower and upper bounds**","metadata":{}},{"cell_type":"code","source":"# Define the lower and upper bounds for outliers\nlower_bound = Q1 - 1.5 * IQR\nupper_bound = Q3 + 1.5 * IQR","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.262184Z","iopub.execute_input":"2024-03-18T08:38:49.262574Z","iopub.status.idle":"2024-03-18T08:38:49.273545Z","shell.execute_reply.started":"2024-03-18T08:38:49.262545Z","shell.execute_reply":"2024-03-18T08:38:49.272108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"6. ### **Filter the data**","metadata":{}},{"cell_type":"code","source":"# Filter the DataFrame to exclude outliers\nfiltered_data = train_data[(train_data['eeg_label_offset_seconds'] >= lower_bound) & (train_data['eeg_label_offset_seconds'] <= upper_bound)]","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.275512Z","iopub.execute_input":"2024-03-18T08:38:49.276249Z","iopub.status.idle":"2024-03-18T08:38:49.296043Z","shell.execute_reply.started":"2024-03-18T08:38:49.276201Z","shell.execute_reply":"2024-03-18T08:38:49.294423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now 'filtered_data' contains rows without outliers\n# You can use this DataFrame for further analysis","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.299213Z","iopub.execute_input":"2024-03-18T08:38:49.299779Z","iopub.status.idle":"2024-03-18T08:38:49.305152Z","shell.execute_reply.started":"2024-03-18T08:38:49.299734Z","shell.execute_reply":"2024-03-18T08:38:49.303729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"7. ### **Print the no of outliers removed**","metadata":{}},{"cell_type":"code","source":"# Print the number of removed outliers\nnum_outliers_removed = len(train_data) - len(filtered_data)\nprint(f\"Removed {num_outliers_removed} outliers from the 'eeg_label_offset_seconds' column.\")","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.306795Z","iopub.execute_input":"2024-03-18T08:38:49.307875Z","iopub.status.idle":"2024-03-18T08:38:49.318829Z","shell.execute_reply.started":"2024-03-18T08:38:49.307791Z","shell.execute_reply":"2024-03-18T08:38:49.317546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Z-Score method**","metadata":{}},{"cell_type":"markdown","source":"1. ### **Import Libraries**","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport numpy as np \nfrom scipy import stats","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.320488Z","iopub.execute_input":"2024-03-18T08:38:49.320962Z","iopub.status.idle":"2024-03-18T08:38:49.334140Z","shell.execute_reply.started":"2024-03-18T08:38:49.320918Z","shell.execute_reply":"2024-03-18T08:38:49.333173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. ### **Load the Dataset**","metadata":{}},{"cell_type":"code","source":"# Load your data\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.338642Z","iopub.execute_input":"2024-03-18T08:38:49.339585Z","iopub.status.idle":"2024-03-18T08:38:49.529096Z","shell.execute_reply.started":"2024-03-18T08:38:49.339537Z","shell.execute_reply":"2024-03-18T08:38:49.527827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. ### **Calculate the Z-Scores for the specific column**","metadata":{}},{"cell_type":"code","source":"# Calculate Z-scores for the specified column\nz_scores = np.abs(stats.zscore(train_data['eeg_label_offset_seconds']))","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.531390Z","iopub.execute_input":"2024-03-18T08:38:49.532053Z","iopub.status.idle":"2024-03-18T08:38:49.543123Z","shell.execute_reply.started":"2024-03-18T08:38:49.532009Z","shell.execute_reply":"2024-03-18T08:38:49.541926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"4. ### **Set a threshold**","metadata":{}},{"cell_type":"code","source":"# Set a threshold (e.g., 3 standard deviations)\nthreshold = 3","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.545181Z","iopub.execute_input":"2024-03-18T08:38:49.545722Z","iopub.status.idle":"2024-03-18T08:38:49.550781Z","shell.execute_reply.started":"2024-03-18T08:38:49.545679Z","shell.execute_reply":"2024-03-18T08:38:49.549840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"5. ### **Filter out rows with Z-Scores beyond the threshold above**","metadata":{}},{"cell_type":"code","source":"# Filter out rows with any Z-score beyond the threshold\nfiltered_data = train_data[(z_scores < threshold)]","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.552107Z","iopub.execute_input":"2024-03-18T08:38:49.552729Z","iopub.status.idle":"2024-03-18T08:38:49.570253Z","shell.execute_reply.started":"2024-03-18T08:38:49.552696Z","shell.execute_reply":"2024-03-18T08:38:49.568834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now 'filtered_data' contains rows without outliers\n# You can use this DataFrame for further analysis","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.572023Z","iopub.execute_input":"2024-03-18T08:38:49.572417Z","iopub.status.idle":"2024-03-18T08:38:49.577617Z","shell.execute_reply.started":"2024-03-18T08:38:49.572355Z","shell.execute_reply":"2024-03-18T08:38:49.576330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"6. ### **Print the number of removed Outliers**","metadata":{}},{"cell_type":"code","source":"# Print the number of removed outliers\nnum_outliers_removed = len(train_data) - len(filtered_data)\nprint(f\"Removed {num_outliers_removed} outliers using Z-scores.\")","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.579413Z","iopub.execute_input":"2024-03-18T08:38:49.579799Z","iopub.status.idle":"2024-03-18T08:38:49.588502Z","shell.execute_reply.started":"2024-03-18T08:38:49.579771Z","shell.execute_reply":"2024-03-18T08:38:49.587218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **K-Means Clustering method**","metadata":{}},{"cell_type":"markdown","source":"1. ### **Import the libraries**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.cluster import KMeans","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.589685Z","iopub.execute_input":"2024-03-18T08:38:49.590395Z","iopub.status.idle":"2024-03-18T08:38:49.803374Z","shell.execute_reply.started":"2024-03-18T08:38:49.590335Z","shell.execute_reply":"2024-03-18T08:38:49.802045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. ### **Load the dataset**","metadata":{}},{"cell_type":"code","source":"# Load your data\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.804751Z","iopub.execute_input":"2024-03-18T08:38:49.805193Z","iopub.status.idle":"2024-03-18T08:38:49.996383Z","shell.execute_reply.started":"2024-03-18T08:38:49.805163Z","shell.execute_reply":"2024-03-18T08:38:49.994928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. ### **Choosing a column for clustering**","metadata":{}},{"cell_type":"code","source":"# Choose the column for clustering\nX = train_data[['eeg_label_offset_seconds']]","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:49.998091Z","iopub.execute_input":"2024-03-18T08:38:49.998483Z","iopub.status.idle":"2024-03-18T08:38:50.005441Z","shell.execute_reply.started":"2024-03-18T08:38:49.998453Z","shell.execute_reply":"2024-03-18T08:38:50.004430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"4. ### **Specify the number of Clusters**","metadata":{}},{"cell_type":"code","source":"# Specify the number of clusters (K)\nK = 3  # Adjust as needed","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.006990Z","iopub.execute_input":"2024-03-18T08:38:50.008255Z","iopub.status.idle":"2024-03-18T08:38:50.015928Z","shell.execute_reply.started":"2024-03-18T08:38:50.008204Z","shell.execute_reply":"2024-03-18T08:38:50.014705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"5. ### **Fit the Model**","metadata":{}},{"cell_type":"code","source":"# Fit K-means model\nkmeans = KMeans(n_clusters=K, random_state=42)\nkmeans.fit(X)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.017822Z","iopub.execute_input":"2024-03-18T08:38:50.018310Z","iopub.status.idle":"2024-03-18T08:38:50.395079Z","shell.execute_reply.started":"2024-03-18T08:38:50.018267Z","shell.execute_reply":"2024-03-18T08:38:50.394089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"6. ### **Get cluster centroids**","metadata":{}},{"cell_type":"code","source":"# Get cluster centroids\ncentroids = kmeans.cluster_centers_","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.396788Z","iopub.execute_input":"2024-03-18T08:38:50.397495Z","iopub.status.idle":"2024-03-18T08:38:50.402805Z","shell.execute_reply.started":"2024-03-18T08:38:50.397455Z","shell.execute_reply":"2024-03-18T08:38:50.401380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"7. ### **Calculate distances from each point to the nearest centroid**","metadata":{}},{"cell_type":"code","source":"# Calculate distances from each point to its nearest centroid\ndistances = np.linalg.norm(X - centroids[kmeans.labels_], axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.404807Z","iopub.execute_input":"2024-03-18T08:38:50.405228Z","iopub.status.idle":"2024-03-18T08:38:50.416635Z","shell.execute_reply.started":"2024-03-18T08:38:50.405197Z","shell.execute_reply":"2024-03-18T08:38:50.414916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"8. ### **Set a threshold for the Outlier detection**","metadata":{}},{"cell_type":"code","source":"# Set a threshold for outlier detection\nthreshold = 3 * np.std(distances)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.418418Z","iopub.execute_input":"2024-03-18T08:38:50.418922Z","iopub.status.idle":"2024-03-18T08:38:50.426114Z","shell.execute_reply.started":"2024-03-18T08:38:50.418881Z","shell.execute_reply":"2024-03-18T08:38:50.424743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"9. ### **Filter out rows with distances**","metadata":{}},{"cell_type":"code","source":"# Filter out rows with distances beyond the threshold\nfiltered_data = train_data[distances < threshold]","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.428176Z","iopub.execute_input":"2024-03-18T08:38:50.428653Z","iopub.status.idle":"2024-03-18T08:38:50.443110Z","shell.execute_reply.started":"2024-03-18T08:38:50.428611Z","shell.execute_reply":"2024-03-18T08:38:50.441564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now 'filtered_data' contains rows without outliers\n# You can use this DataFrame for further analysis\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.445013Z","iopub.execute_input":"2024-03-18T08:38:50.445542Z","iopub.status.idle":"2024-03-18T08:38:50.452862Z","shell.execute_reply.started":"2024-03-18T08:38:50.445499Z","shell.execute_reply":"2024-03-18T08:38:50.450665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"10. ### **Print the number of removed Outliers**","metadata":{}},{"cell_type":"code","source":"# Print the number of removed outliers\nnum_outliers_removed = len(train_data) - len(filtered_data)\nprint(f\"Removed {num_outliers_removed} outliers using K-means clustering.\")","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.455550Z","iopub.execute_input":"2024-03-18T08:38:50.455962Z","iopub.status.idle":"2024-03-18T08:38:50.464161Z","shell.execute_reply.started":"2024-03-18T08:38:50.455931Z","shell.execute_reply":"2024-03-18T08:38:50.462429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Using various Scalers to Pre-process the data for Machine Learning models**","metadata":{}},{"cell_type":"markdown","source":"# **Standard Scaler**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.466097Z","iopub.execute_input":"2024-03-18T08:38:50.466628Z","iopub.status.idle":"2024-03-18T08:38:50.480513Z","shell.execute_reply.started":"2024-03-18T08:38:50.466585Z","shell.execute_reply":"2024-03-18T08:38:50.479520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv') ","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.492884Z","iopub.execute_input":"2024-03-18T08:38:50.494576Z","iopub.status.idle":"2024-03-18T08:38:50.693746Z","shell.execute_reply.started":"2024-03-18T08:38:50.494518Z","shell.execute_reply":"2024-03-18T08:38:50.692480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Choose feature\nX = train_data[['eeg_label_offset_seconds']]\n\n# Choose target column\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.697233Z","iopub.execute_input":"2024-03-18T08:38:50.697633Z","iopub.status.idle":"2024-03-18T08:38:50.705671Z","shell.execute_reply.started":"2024-03-18T08:38:50.697604Z","shell.execute_reply":"2024-03-18T08:38:50.704024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.707301Z","iopub.execute_input":"2024-03-18T08:38:50.708014Z","iopub.status.idle":"2024-03-18T08:38:50.731748Z","shell.execute_reply.started":"2024-03-18T08:38:50.707969Z","shell.execute_reply":"2024-03-18T08:38:50.730488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize StandardScaler\nscaler = StandardScaler()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.733351Z","iopub.execute_input":"2024-03-18T08:38:50.733739Z","iopub.status.idle":"2024-03-18T08:38:50.740717Z","shell.execute_reply.started":"2024-03-18T08:38:50.733708Z","shell.execute_reply":"2024-03-18T08:38:50.739584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit the scaler to the training data and transform it\nX_train_scaled = scaler.fit_transform(X_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.742579Z","iopub.execute_input":"2024-03-18T08:38:50.743839Z","iopub.status.idle":"2024-03-18T08:38:50.755150Z","shell.execute_reply.started":"2024-03-18T08:38:50.743697Z","shell.execute_reply":"2024-03-18T08:38:50.753782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transform the test data using the fitted scaler\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.756675Z","iopub.execute_input":"2024-03-18T08:38:50.757135Z","iopub.status.idle":"2024-03-18T08:38:50.765725Z","shell.execute_reply.started":"2024-03-18T08:38:50.757104Z","shell.execute_reply":"2024-03-18T08:38:50.764448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled training features\nprint(\"Scaled Training Features (X_train_scaled):\")\nprint(X_train_scaled[:5])","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.767186Z","iopub.execute_input":"2024-03-18T08:38:50.768291Z","iopub.status.idle":"2024-03-18T08:38:50.777781Z","shell.execute_reply.started":"2024-03-18T08:38:50.768253Z","shell.execute_reply":"2024-03-18T08:38:50.776492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled testing features\nprint(\"\\nScaled Testing Features (X_test_scaled):\")\nprint(X_test_scaled[:5])","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.779317Z","iopub.execute_input":"2024-03-18T08:38:50.780136Z","iopub.status.idle":"2024-03-18T08:38:50.792002Z","shell.execute_reply.started":"2024-03-18T08:38:50.780099Z","shell.execute_reply":"2024-03-18T08:38:50.790456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Min-max Scaler**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MinMaxScaler","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.794077Z","iopub.execute_input":"2024-03-18T08:38:50.794711Z","iopub.status.idle":"2024-03-18T08:38:50.800854Z","shell.execute_reply.started":"2024-03-18T08:38:50.794666Z","shell.execute_reply":"2024-03-18T08:38:50.799927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv') ","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.802116Z","iopub.execute_input":"2024-03-18T08:38:50.803214Z","iopub.status.idle":"2024-03-18T08:38:50.991896Z","shell.execute_reply.started":"2024-03-18T08:38:50.803175Z","shell.execute_reply":"2024-03-18T08:38:50.990793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Choose feature\nX = train_data[['eeg_label_offset_seconds']]\n\n# Choose target column\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:50.993379Z","iopub.execute_input":"2024-03-18T08:38:50.993713Z","iopub.status.idle":"2024-03-18T08:38:51.001015Z","shell.execute_reply.started":"2024-03-18T08:38:50.993686Z","shell.execute_reply":"2024-03-18T08:38:50.999845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.002705Z","iopub.execute_input":"2024-03-18T08:38:51.003010Z","iopub.status.idle":"2024-03-18T08:38:51.024908Z","shell.execute_reply.started":"2024-03-18T08:38:51.002986Z","shell.execute_reply":"2024-03-18T08:38:51.023431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize MinMaxScaler\nscaler = MinMaxScaler()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.028681Z","iopub.execute_input":"2024-03-18T08:38:51.029113Z","iopub.status.idle":"2024-03-18T08:38:51.034627Z","shell.execute_reply.started":"2024-03-18T08:38:51.029082Z","shell.execute_reply":"2024-03-18T08:38:51.033191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit the scaler to the training data and transform it\nX_train_scaled = scaler.fit_transform(X_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.036464Z","iopub.execute_input":"2024-03-18T08:38:51.036881Z","iopub.status.idle":"2024-03-18T08:38:51.049774Z","shell.execute_reply.started":"2024-03-18T08:38:51.036847Z","shell.execute_reply":"2024-03-18T08:38:51.048479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transform the test data using the fitted scaler\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.051743Z","iopub.execute_input":"2024-03-18T08:38:51.052448Z","iopub.status.idle":"2024-03-18T08:38:51.062836Z","shell.execute_reply.started":"2024-03-18T08:38:51.052385Z","shell.execute_reply":"2024-03-18T08:38:51.061435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled training features\nprint(\"Scaled Training Features (X_train_scaled):\")\nprint(X_train_scaled[:5])","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.064495Z","iopub.execute_input":"2024-03-18T08:38:51.066866Z","iopub.status.idle":"2024-03-18T08:38:51.075920Z","shell.execute_reply.started":"2024-03-18T08:38:51.066821Z","shell.execute_reply":"2024-03-18T08:38:51.074972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled testing features\nprint(\"\\nScaled Testing Features (X_test_scaled):\")\nprint(X_test_scaled[:5])","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.077494Z","iopub.execute_input":"2024-03-18T08:38:51.078352Z","iopub.status.idle":"2024-03-18T08:38:51.093120Z","shell.execute_reply.started":"2024-03-18T08:38:51.078319Z","shell.execute_reply":"2024-03-18T08:38:51.092124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Robust Scaler**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import RobustScaler","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.094226Z","iopub.execute_input":"2024-03-18T08:38:51.094620Z","iopub.status.idle":"2024-03-18T08:38:51.102331Z","shell.execute_reply.started":"2024-03-18T08:38:51.094583Z","shell.execute_reply":"2024-03-18T08:38:51.101405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.103695Z","iopub.execute_input":"2024-03-18T08:38:51.104235Z","iopub.status.idle":"2024-03-18T08:38:51.294130Z","shell.execute_reply.started":"2024-03-18T08:38:51.104206Z","shell.execute_reply":"2024-03-18T08:38:51.292683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Choose feature\nX = train_data[['eeg_label_offset_seconds']]\n\n# Choose target column\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.295831Z","iopub.execute_input":"2024-03-18T08:38:51.296790Z","iopub.status.idle":"2024-03-18T08:38:51.304276Z","shell.execute_reply.started":"2024-03-18T08:38:51.296743Z","shell.execute_reply":"2024-03-18T08:38:51.302700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.305902Z","iopub.execute_input":"2024-03-18T08:38:51.306266Z","iopub.status.idle":"2024-03-18T08:38:51.329159Z","shell.execute_reply.started":"2024-03-18T08:38:51.306238Z","shell.execute_reply":"2024-03-18T08:38:51.327806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize RobustScaler\nscaler = RobustScaler()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.330693Z","iopub.execute_input":"2024-03-18T08:38:51.331059Z","iopub.status.idle":"2024-03-18T08:38:51.335955Z","shell.execute_reply.started":"2024-03-18T08:38:51.331032Z","shell.execute_reply":"2024-03-18T08:38:51.334736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit the scaler to the training data and transform it\nX_train_scaled = scaler.fit_transform(X_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.337796Z","iopub.execute_input":"2024-03-18T08:38:51.338250Z","iopub.status.idle":"2024-03-18T08:38:51.353809Z","shell.execute_reply.started":"2024-03-18T08:38:51.338210Z","shell.execute_reply":"2024-03-18T08:38:51.352553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transform the test data using the fitted scaler\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.355183Z","iopub.execute_input":"2024-03-18T08:38:51.355634Z","iopub.status.idle":"2024-03-18T08:38:51.363110Z","shell.execute_reply.started":"2024-03-18T08:38:51.355594Z","shell.execute_reply":"2024-03-18T08:38:51.361951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled training features\nprint(\"Scaled Training Features (X_train_scaled):\")\nprint(X_train_scaled[:5])","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.364545Z","iopub.execute_input":"2024-03-18T08:38:51.364946Z","iopub.status.idle":"2024-03-18T08:38:51.374807Z","shell.execute_reply.started":"2024-03-18T08:38:51.364914Z","shell.execute_reply":"2024-03-18T08:38:51.373692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled testing features\nprint(\"\\nScaled Testing Features (X_test_scaled):\")\nprint(X_test_scaled[:5])","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.376394Z","iopub.execute_input":"2024-03-18T08:38:51.377006Z","iopub.status.idle":"2024-03-18T08:38:51.389272Z","shell.execute_reply.started":"2024-03-18T08:38:51.376974Z","shell.execute_reply":"2024-03-18T08:38:51.387758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Max-Abs Scaler**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MaxAbsScaler","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.391329Z","iopub.execute_input":"2024-03-18T08:38:51.391751Z","iopub.status.idle":"2024-03-18T08:38:51.399341Z","shell.execute_reply.started":"2024-03-18T08:38:51.391719Z","shell.execute_reply":"2024-03-18T08:38:51.398432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.400538Z","iopub.execute_input":"2024-03-18T08:38:51.401765Z","iopub.status.idle":"2024-03-18T08:38:51.592095Z","shell.execute_reply.started":"2024-03-18T08:38:51.401727Z","shell.execute_reply":"2024-03-18T08:38:51.590788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Choose feature\nX = train_data[['eeg_label_offset_seconds']]\n\n# Choose target column\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.593820Z","iopub.execute_input":"2024-03-18T08:38:51.594425Z","iopub.status.idle":"2024-03-18T08:38:51.602189Z","shell.execute_reply.started":"2024-03-18T08:38:51.594392Z","shell.execute_reply":"2024-03-18T08:38:51.600895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.603751Z","iopub.execute_input":"2024-03-18T08:38:51.604160Z","iopub.status.idle":"2024-03-18T08:38:51.627680Z","shell.execute_reply.started":"2024-03-18T08:38:51.604128Z","shell.execute_reply":"2024-03-18T08:38:51.626304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize MaxAbsScaler\nscaler = MaxAbsScaler()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.629101Z","iopub.execute_input":"2024-03-18T08:38:51.629551Z","iopub.status.idle":"2024-03-18T08:38:51.634621Z","shell.execute_reply.started":"2024-03-18T08:38:51.629519Z","shell.execute_reply":"2024-03-18T08:38:51.633332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit the scaler to the training data and transform it\nX_train_scaled = scaler.fit_transform(X_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.636319Z","iopub.execute_input":"2024-03-18T08:38:51.636737Z","iopub.status.idle":"2024-03-18T08:38:51.649818Z","shell.execute_reply.started":"2024-03-18T08:38:51.636697Z","shell.execute_reply":"2024-03-18T08:38:51.648421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transform the test data using the fitted scaler\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.651885Z","iopub.execute_input":"2024-03-18T08:38:51.652269Z","iopub.status.idle":"2024-03-18T08:38:51.660045Z","shell.execute_reply.started":"2024-03-18T08:38:51.652239Z","shell.execute_reply":"2024-03-18T08:38:51.658579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled training features\nprint(\"Scaled Training Features (X_train_scaled):\")\nprint(X_train_scaled[:5])\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.661497Z","iopub.execute_input":"2024-03-18T08:38:51.662281Z","iopub.status.idle":"2024-03-18T08:38:51.677868Z","shell.execute_reply.started":"2024-03-18T08:38:51.662155Z","shell.execute_reply":"2024-03-18T08:38:51.676799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the scaled testing features\nprint(\"\\nScaled Testing Features (X_test_scaled):\")\nprint(X_test_scaled[:5])","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.679128Z","iopub.execute_input":"2024-03-18T08:38:51.680494Z","iopub.status.idle":"2024-03-18T08:38:51.690272Z","shell.execute_reply.started":"2024-03-18T08:38:51.680456Z","shell.execute_reply":"2024-03-18T08:38:51.688984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Using Gaussian Distribution to normalize the Data**","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport matplotlib.pyplot as plt \nfrom sklearn.datasets import make_classification\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.691699Z","iopub.execute_input":"2024-03-18T08:38:51.692174Z","iopub.status.idle":"2024-03-18T08:38:51.830455Z","shell.execute_reply.started":"2024-03-18T08:38:51.692132Z","shell.execute_reply":"2024-03-18T08:38:51.829302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assuming you have already split your data into X_train, X_test, y_train, y_test","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.832546Z","iopub.execute_input":"2024-03-18T08:38:51.832935Z","iopub.status.idle":"2024-03-18T08:38:51.837680Z","shell.execute_reply.started":"2024-03-18T08:38:51.832904Z","shell.execute_reply":"2024-03-18T08:38:51.836464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a Gaussian Naive Bayes classifier\ngnb = GaussianNB()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.839166Z","iopub.execute_input":"2024-03-18T08:38:51.840108Z","iopub.status.idle":"2024-03-18T08:38:51.849108Z","shell.execute_reply.started":"2024-03-18T08:38:51.840063Z","shell.execute_reply":"2024-03-18T08:38:51.847898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier on the training data\ngnb.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:51.850713Z","iopub.execute_input":"2024-03-18T08:38:51.851705Z","iopub.status.idle":"2024-03-18T08:38:52.136404Z","shell.execute_reply.started":"2024-03-18T08:38:51.851665Z","shell.execute_reply":"2024-03-18T08:38:52.135242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions on the test data\ny_pred = gnb.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.137914Z","iopub.execute_input":"2024-03-18T08:38:52.138270Z","iopub.status.idle":"2024-03-18T08:38:52.151161Z","shell.execute_reply.started":"2024-03-18T08:38:52.138241Z","shell.execute_reply":"2024-03-18T08:38:52.149972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate accuracy\naccuracy = accuracy_score(y_test, y_pred)\nprint(\"Accuracy:\", accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.152797Z","iopub.execute_input":"2024-03-18T08:38:52.153140Z","iopub.status.idle":"2024-03-18T08:38:52.194027Z","shell.execute_reply.started":"2024-03-18T08:38:52.153111Z","shell.execute_reply":"2024-03-18T08:38:52.192784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Generalized approach to using Gaussian Naive Bayes**","metadata":{}},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.195442Z","iopub.execute_input":"2024-03-18T08:38:52.195803Z","iopub.status.idle":"2024-03-18T08:38:52.375126Z","shell.execute_reply.started":"2024-03-18T08:38:52.195773Z","shell.execute_reply":"2024-03-18T08:38:52.374130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# choose feature\nX = train_data[['eeg_label_offset_seconds']]","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.376463Z","iopub.execute_input":"2024-03-18T08:38:52.377014Z","iopub.status.idle":"2024-03-18T08:38:52.383743Z","shell.execute_reply.started":"2024-03-18T08:38:52.376982Z","shell.execute_reply":"2024-03-18T08:38:52.382067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# choose target \ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.385916Z","iopub.execute_input":"2024-03-18T08:38:52.386411Z","iopub.status.idle":"2024-03-18T08:38:52.397447Z","shell.execute_reply.started":"2024-03-18T08:38:52.386354Z","shell.execute_reply":"2024-03-18T08:38:52.395871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.399116Z","iopub.execute_input":"2024-03-18T08:38:52.399525Z","iopub.status.idle":"2024-03-18T08:38:52.421215Z","shell.execute_reply.started":"2024-03-18T08:38:52.399494Z","shell.execute_reply":"2024-03-18T08:38:52.419809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a Gaussian Naive Bayes classifier\ngnb = GaussianNB()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.422720Z","iopub.execute_input":"2024-03-18T08:38:52.423607Z","iopub.status.idle":"2024-03-18T08:38:52.429905Z","shell.execute_reply.started":"2024-03-18T08:38:52.423559Z","shell.execute_reply":"2024-03-18T08:38:52.428468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier on the training data\ngnb.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.431240Z","iopub.execute_input":"2024-03-18T08:38:52.432109Z","iopub.status.idle":"2024-03-18T08:38:52.716033Z","shell.execute_reply.started":"2024-03-18T08:38:52.432073Z","shell.execute_reply":"2024-03-18T08:38:52.714848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions on the test data\ny_pred = gnb.predict(X_test)\n\n# Calculate accuracy\naccuracy = accuracy_score(y_test, y_pred)\nprint(\"Accuracy:\", accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.717592Z","iopub.execute_input":"2024-03-18T08:38:52.717945Z","iopub.status.idle":"2024-03-18T08:38:52.761646Z","shell.execute_reply.started":"2024-03-18T08:38:52.717915Z","shell.execute_reply":"2024-03-18T08:38:52.760214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns \n\n# Plot the PDFs for each class\nplt.figure(figsize=(10, 8))\nsns.histplot(data=train_data, x='eeg_label_offset_seconds', hue='expert_consensus', kde=True, stat='density', common_norm=False)\nplt.title('Probability Density Functions (PDFs) for Each Class')\nplt.xlabel('EEG Label Offset Seconds')\nplt.ylabel('Density')\nplt.legend(title='Class')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:38:52.763221Z","iopub.execute_input":"2024-03-18T08:38:52.763715Z","iopub.status.idle":"2024-03-18T08:39:09.735125Z","shell.execute_reply.started":"2024-03-18T08:38:52.763670Z","shell.execute_reply":"2024-03-18T08:39:09.733786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Using different Transformations to normalize the Dataset**\n\n##### - **Yeo-Johnson transformation**\n\n##### - **Quantile transformation**","metadata":{}},{"cell_type":"markdown","source":"1. # **Yeo-Johnson Transformation:**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np \nimport matplotlib.pyplot as plt \nimport seaborn as sns \nfrom scipy.stats import boxcox\n\nfrom sklearn.preprocessing import PowerTransformer","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:09.736585Z","iopub.execute_input":"2024-03-18T08:39:09.736945Z","iopub.status.idle":"2024-03-18T08:39:09.743576Z","shell.execute_reply.started":"2024-03-18T08:39:09.736912Z","shell.execute_reply":"2024-03-18T08:39:09.742143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:09.744920Z","iopub.execute_input":"2024-03-18T08:39:09.745291Z","iopub.status.idle":"2024-03-18T08:39:09.948571Z","shell.execute_reply.started":"2024-03-18T08:39:09.745263Z","shell.execute_reply":"2024-03-18T08:39:09.947154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select the EEG signal feature for transformation\nselected_feature = train_data['eeg_label_offset_seconds']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:09.950114Z","iopub.execute_input":"2024-03-18T08:39:09.950523Z","iopub.status.idle":"2024-03-18T08:39:09.957263Z","shell.execute_reply.started":"2024-03-18T08:39:09.950491Z","shell.execute_reply":"2024-03-18T08:39:09.955877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize Yeo-Johnson transformer\nyeo_johnson_transformer = PowerTransformer(method='yeo-johnson', standardize=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:09.959045Z","iopub.execute_input":"2024-03-18T08:39:09.959699Z","iopub.status.idle":"2024-03-18T08:39:09.971685Z","shell.execute_reply.started":"2024-03-18T08:39:09.959662Z","shell.execute_reply":"2024-03-18T08:39:09.969910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reshape the selected feature to a 2D array (required by PowerTransformer)\nselected_feature_2d = selected_feature.values.reshape(-1, 1)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:09.973444Z","iopub.execute_input":"2024-03-18T08:39:09.973979Z","iopub.status.idle":"2024-03-18T08:39:09.984320Z","shell.execute_reply.started":"2024-03-18T08:39:09.973935Z","shell.execute_reply":"2024-03-18T08:39:09.982977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit and transform the feature using Yeo-Johnson transformation\ntransformed_feature_yeo_johnson = yeo_johnson_transformer.fit_transform(selected_feature_2d)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:10.005578Z","iopub.execute_input":"2024-03-18T08:39:10.006804Z","iopub.status.idle":"2024-03-18T08:39:10.201707Z","shell.execute_reply.started":"2024-03-18T08:39:10.006750Z","shell.execute_reply":"2024-03-18T08:39:10.200064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print the transformed feature\nprint(\"Yeo-Johnson Transformed Feature:\")\nprint(transformed_feature_yeo_johnson)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:10.203548Z","iopub.execute_input":"2024-03-18T08:39:10.203923Z","iopub.status.idle":"2024-03-18T08:39:10.211237Z","shell.execute_reply.started":"2024-03-18T08:39:10.203894Z","shell.execute_reply":"2024-03-18T08:39:10.209802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a figure with subplots\nfig, axs = plt.subplots(1, 2, figsize=(12, 6))\n\n# Plot the histogram of the original feature\naxs[0].hist(selected_feature, bins=30, color='blue', alpha=0.7)\naxs[0].set_title('Histogram of Original Feature')\naxs[0].set_xlabel('Feature Values')\naxs[0].set_ylabel('Frequency')\n\n# Plot the histogram of the transformed feature\naxs[1].hist(transformed_feature_yeo_johnson, bins=30, color='green', alpha=0.7)\naxs[1].set_title('Histogram of Transformed Feature (Yeo-Johnson)')\naxs[1].set_xlabel('Transformed Values')\naxs[1].set_ylabel('Frequency')\n\n# Show the plot\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:10.212900Z","iopub.execute_input":"2024-03-18T08:39:10.213257Z","iopub.status.idle":"2024-03-18T08:39:11.247560Z","shell.execute_reply.started":"2024-03-18T08:39:10.213230Z","shell.execute_reply":"2024-03-18T08:39:11.246220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. # **Quantile transformation**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nfrom sklearn.preprocessing import QuantileTransformer","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:11.248972Z","iopub.execute_input":"2024-03-18T08:39:11.249343Z","iopub.status.idle":"2024-03-18T08:39:11.256395Z","shell.execute_reply.started":"2024-03-18T08:39:11.249313Z","shell.execute_reply":"2024-03-18T08:39:11.254388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:11.258010Z","iopub.execute_input":"2024-03-18T08:39:11.258394Z","iopub.status.idle":"2024-03-18T08:39:11.448064Z","shell.execute_reply.started":"2024-03-18T08:39:11.258349Z","shell.execute_reply":"2024-03-18T08:39:11.446819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select the EEG signal feature\nselected_feature = train_data['eeg_label_offset_seconds']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:11.449568Z","iopub.execute_input":"2024-03-18T08:39:11.449904Z","iopub.status.idle":"2024-03-18T08:39:11.455783Z","shell.execute_reply.started":"2024-03-18T08:39:11.449877Z","shell.execute_reply":"2024-03-18T08:39:11.454305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize Quantile transformer\nquantile_transformer = QuantileTransformer(output_distribution='normal')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:11.457471Z","iopub.execute_input":"2024-03-18T08:39:11.457837Z","iopub.status.idle":"2024-03-18T08:39:11.466551Z","shell.execute_reply.started":"2024-03-18T08:39:11.457808Z","shell.execute_reply":"2024-03-18T08:39:11.465208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reshape the selected feature to a 2D array (required by QuantileTransformer)\nselected_feature_2d = selected_feature.values.reshape(-1, 1)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:11.468698Z","iopub.execute_input":"2024-03-18T08:39:11.469266Z","iopub.status.idle":"2024-03-18T08:39:11.477523Z","shell.execute_reply.started":"2024-03-18T08:39:11.469223Z","shell.execute_reply":"2024-03-18T08:39:11.476024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit and transform the feature using Quantile transformation\ntransformed_feature_quantile = quantile_transformer.fit_transform(selected_feature_2d)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:11.480661Z","iopub.execute_input":"2024-03-18T08:39:11.481498Z","iopub.status.idle":"2024-03-18T08:39:11.514805Z","shell.execute_reply.started":"2024-03-18T08:39:11.481457Z","shell.execute_reply":"2024-03-18T08:39:11.513176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a figure with subplots\nfig, axs = plt.subplots(1, 2, figsize=(12, 6))\n\n# Plot the histogram of the original feature\naxs[0].hist(selected_feature, bins=30, color='blue', alpha=0.7)\naxs[0].set_title('Histogram of Original Feature')\naxs[0].set_xlabel('Feature Values')\naxs[0].set_ylabel('Frequency')\n\n# Plot the histogram of the transformed feature\naxs[1].hist(transformed_feature_quantile, bins=30, color='orange', alpha=0.7)\naxs[1].set_title('Histogram of Transformed Feature (Quantile)')\naxs[1].set_xlabel('Transformed Values')\naxs[1].set_ylabel('Frequency')\n\n# Show the plot\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:11.516522Z","iopub.execute_input":"2024-03-18T08:39:11.517479Z","iopub.status.idle":"2024-03-18T08:39:12.523183Z","shell.execute_reply.started":"2024-03-18T08:39:11.517437Z","shell.execute_reply":"2024-03-18T08:39:12.521399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Normalization, L1 & L2**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import Normalizer\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.524769Z","iopub.execute_input":"2024-03-18T08:39:12.525139Z","iopub.status.idle":"2024-03-18T08:39:12.531338Z","shell.execute_reply.started":"2024-03-18T08:39:12.525109Z","shell.execute_reply":"2024-03-18T08:39:12.529760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load your dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.532770Z","iopub.execute_input":"2024-03-18T08:39:12.533117Z","iopub.status.idle":"2024-03-18T08:39:12.725718Z","shell.execute_reply.started":"2024-03-18T08:39:12.533081Z","shell.execute_reply":"2024-03-18T08:39:12.724451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Choose the features you want to normalize\nX = train_data[['eeg_label_offset_seconds']] ","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.727573Z","iopub.execute_input":"2024-03-18T08:39:12.727950Z","iopub.status.idle":"2024-03-18T08:39:12.734795Z","shell.execute_reply.started":"2024-03-18T08:39:12.727920Z","shell.execute_reply":"2024-03-18T08:39:12.733416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize the Normalizer with L1 normalization\nnormalizer_l1 = Normalizer(norm='l1')\nX_normalized_l1 = normalizer_l1.fit_transform(X)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.736541Z","iopub.execute_input":"2024-03-18T08:39:12.737060Z","iopub.status.idle":"2024-03-18T08:39:12.750758Z","shell.execute_reply.started":"2024-03-18T08:39:12.737017Z","shell.execute_reply":"2024-03-18T08:39:12.749259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize the Normalizer with L2 normalization\nnormalizer_l2 = Normalizer(norm='l2')\nX_normalized_l2 = normalizer_l2.fit_transform(X)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.751855Z","iopub.execute_input":"2024-03-18T08:39:12.752199Z","iopub.status.idle":"2024-03-18T08:39:12.767645Z","shell.execute_reply.started":"2024-03-18T08:39:12.752171Z","shell.execute_reply":"2024-03-18T08:39:12.766126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print the normalized data\nprint(\"Normalized Data (L1):\")\nprint(X_normalized_l1)\n\n# Print the normalized data (L2)\nprint(\"Normalized Data (L2):\")\nprint(X_normalized_l2)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.769135Z","iopub.execute_input":"2024-03-18T08:39:12.769495Z","iopub.status.idle":"2024-03-18T08:39:12.778042Z","shell.execute_reply.started":"2024-03-18T08:39:12.769466Z","shell.execute_reply":"2024-03-18T08:39:12.777087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Performing Square root transformation**","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\n# Perform square root transformation on the chosen column\ntrain_data['eeg_label_offset_seconds_sqrt'] = np.sqrt(train_data['eeg_label_offset_seconds'])\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.779488Z","iopub.execute_input":"2024-03-18T08:39:12.779874Z","iopub.status.idle":"2024-03-18T08:39:12.810649Z","shell.execute_reply.started":"2024-03-18T08:39:12.779844Z","shell.execute_reply":"2024-03-18T08:39:12.809325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns \nimport matplotlib.pyplot as plt\n\n# Create a figure with two subplots\nfig, axes = plt.subplots(1, 2, figsize=(14, 6))\n\n# Plot histogram of the original data\nsns.histplot(data=train_data, x='eeg_label_offset_seconds', kde=True, bins=30, ax=axes[0])\naxes[0].set_title('Histogram of EEG Label Offset Seconds (Original)')\naxes[0].set_xlabel('EEG Label Offset Seconds')\naxes[0].set_ylabel('Frequency')\n\n# Plot histogram of the square root-transformed data\nsns.histplot(data=train_data, x='eeg_label_offset_seconds_sqrt', kde=True, bins=30, ax=axes[1])\naxes[1].set_title('Histogram of Square Root-transformed EEG Label Offset Seconds')\naxes[1].set_xlabel('Square Root-transformed EEG Label Offset Seconds')\naxes[1].set_ylabel('Frequency')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:12.812343Z","iopub.execute_input":"2024-03-18T08:39:12.812879Z","iopub.status.idle":"2024-03-18T08:39:15.036675Z","shell.execute_reply.started":"2024-03-18T08:39:12.812838Z","shell.execute_reply":"2024-03-18T08:39:15.035183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Performing Log transformation**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\n# Choose the column for log transformation\ncolumn_name = 'eeg_label_offset_seconds'\n\n# Perform log transformation\ntrain_data[column_name + '_log'] = train_data[column_name].apply(lambda x: np.log1p(x))\n\n# Display the transformed column\nprint(train_data[[column_name, column_name + '_log']].head())\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:15.038330Z","iopub.execute_input":"2024-03-18T08:39:15.038722Z","iopub.status.idle":"2024-03-18T08:39:15.478631Z","shell.execute_reply.started":"2024-03-18T08:39:15.038693Z","shell.execute_reply":"2024-03-18T08:39:15.477117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a figure with two subplots\nfig, axes = plt.subplots(1, 2, figsize=(14, 6))\n\n# Plot histogram of the original data\nsns.histplot(data=train_data, x='eeg_label_offset_seconds', kde=True, bins=30, ax=axes[0])\naxes[0].set_title('Histogram of EEG Label Offset Seconds (Original)')\naxes[0].set_xlabel('EEG Label Offset Seconds')\naxes[0].set_ylabel('Frequency')\n\n# Plot histogram of the log-transformed data\nsns.histplot(data=train_data, x='eeg_label_offset_seconds_log', kde=True, bins=30, ax=axes[1])\naxes[1].set_title('Histogram of Log-transformed EEG Label Offset Seconds')\naxes[1].set_xlabel('Log-transformed EEG Label Offset Seconds')\naxes[1].set_ylabel('Frequency')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:15.480137Z","iopub.execute_input":"2024-03-18T08:39:15.480623Z","iopub.status.idle":"2024-03-18T08:39:17.539657Z","shell.execute_reply.started":"2024-03-18T08:39:15.480579Z","shell.execute_reply":"2024-03-18T08:39:17.538469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Encoders, Using the four-types of Encoders on the HMC Dataset**\n\n- **Binary Encoder**\n- **Label Encoder**\n- **Ordinal Encoder**\n- **One-Hot Encoder**","metadata":{}},{"cell_type":"markdown","source":"1. # **Label Encoder**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\n# Initialize LabelEncoder\nle = LabelEncoder()\n\n# Apply label encoding to the 'expert_consensus' column\ntrain_data['labelencoder_expert_consensus'] = le.fit_transform(train_data['expert_consensus'])\ntrain_data\n\ntrain_data['expert_consensus'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.541245Z","iopub.execute_input":"2024-03-18T08:39:17.542719Z","iopub.status.idle":"2024-03-18T08:39:17.605461Z","shell.execute_reply.started":"2024-03-18T08:39:17.542670Z","shell.execute_reply":"2024-03-18T08:39:17.603816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_data['encoded_expert_consensus'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.606823Z","iopub.execute_input":"2024-03-18T08:39:17.607183Z","iopub.status.idle":"2024-03-18T08:39:17.612771Z","shell.execute_reply.started":"2024-03-18T08:39:17.607155Z","shell.execute_reply":"2024-03-18T08:39:17.611381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.614239Z","iopub.execute_input":"2024-03-18T08:39:17.615183Z","iopub.status.idle":"2024-03-18T08:39:17.642850Z","shell.execute_reply.started":"2024-03-18T08:39:17.615147Z","shell.execute_reply":"2024-03-18T08:39:17.641494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. # **Binary Encoder**","metadata":{}},{"cell_type":"code","source":"import category_encoders as ce\nfrom category_encoders import BinaryEncoder\n\nbe = BinaryEncoder()\n\ndf_binary = be.fit_transform(train_data['expert_consensus'])\ndf_binary\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.645583Z","iopub.execute_input":"2024-03-18T08:39:17.646113Z","iopub.status.idle":"2024-03-18T08:39:17.772957Z","shell.execute_reply.started":"2024-03-18T08:39:17.646068Z","shell.execute_reply":"2024-03-18T08:39:17.771504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. # **Ordianl Encoder**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import OrdinalEncoder\n\n# Initialize OrdinalEncoder\noe = OrdinalEncoder()\n\ntrain_data['encoded_expert_consensus'] = oe.fit_transform(train_data[['expert_consensus']])\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.774567Z","iopub.execute_input":"2024-03-18T08:39:17.774970Z","iopub.status.idle":"2024-03-18T08:39:17.840806Z","shell.execute_reply.started":"2024-03-18T08:39:17.774937Z","shell.execute_reply":"2024-03-18T08:39:17.839565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.842215Z","iopub.execute_input":"2024-03-18T08:39:17.842684Z","iopub.status.idle":"2024-03-18T08:39:17.850029Z","shell.execute_reply.started":"2024-03-18T08:39:17.842652Z","shell.execute_reply":"2024-03-18T08:39:17.849008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **One-Hot Encoder**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import OneHotEncoder\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.851143Z","iopub.execute_input":"2024-03-18T08:39:17.851548Z","iopub.status.idle":"2024-03-18T08:39:17.858819Z","shell.execute_reply.started":"2024-03-18T08:39:17.851510Z","shell.execute_reply":"2024-03-18T08:39:17.857506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\ntrain_data.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:17.860594Z","iopub.execute_input":"2024-03-18T08:39:17.862059Z","iopub.status.idle":"2024-03-18T08:39:18.053728Z","shell.execute_reply.started":"2024-03-18T08:39:17.862015Z","shell.execute_reply":"2024-03-18T08:39:18.052430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.preprocessing import OneHotEncoder\n\ncolumn_to_encode = 'expert_consensus'\n\n# Initialize OneHotEncoder\nencoder = OneHotEncoder(sparse=False, drop='first')\n\n# Reshape the column to a 2D array\nencoded_features = encoder.fit_transform(train_data[column_to_encode].values.reshape(-1, 1))\n\n# Extract categories from the encoder\ncategories = encoder.categories_[0]\n\n# Generate column names for the encoded features\nencoded_column_names = [f'{column_to_encode}_{category}' for category in categories[1:]]\n\n# Create a DataFrame for the encoded features\nencoded_df = pd.DataFrame(encoded_features, columns=encoded_column_names)\n\n# Concatenate the original DataFrame with the encoded DataFrame\ntrain_data_encoded = pd.concat([train_data, encoded_df], axis=1)\n\n# Drop the original categorical column\ntrain_data_encoded.drop(columns=[column_to_encode], inplace=True)\n\n# Display the encoded DataFrame\nprint(train_data_encoded.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:18.056502Z","iopub.execute_input":"2024-03-18T08:39:18.056991Z","iopub.status.idle":"2024-03-18T08:39:18.140483Z","shell.execute_reply.started":"2024-03-18T08:39:18.056948Z","shell.execute_reply":"2024-03-18T08:39:18.139248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_data.dtypes)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:18.142207Z","iopub.execute_input":"2024-03-18T08:39:18.142567Z","iopub.status.idle":"2024-03-18T08:39:18.149845Z","shell.execute_reply.started":"2024-03-18T08:39:18.142537Z","shell.execute_reply":"2024-03-18T08:39:18.148439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Reduction**","metadata":{}},{"cell_type":"markdown","source":"# **Data Discretization with K-Bins Discretizer**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import KBinsDiscretizer\nimport pandas as pd\n\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\ncontinuous_column = train_data['eeg_label_offset_seconds'].values.reshape(-1, 1)\n\nn_bins = 2  # Number of bins\nstrategy = 'quantile'  # Strategy for binning, e.g., 'uniform' or 'quantile'\ndiscretizer = KBinsDiscretizer(n_bins=n_bins, encode='ordinal', strategy=strategy)\n\n# Fit and transform the continuous data into discrete bins\ndiscretized_data = discretizer.fit_transform(continuous_column)\n\n# Convert the result to a DataFrame for easier visualization\ndiscretized_df = pd.DataFrame(discretized_data, columns=['eeg_label_offset_seconds_discretized'])\n\n# Display the discretized data\nprint(discretized_df.head())\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:18.151934Z","iopub.execute_input":"2024-03-18T08:39:18.153152Z","iopub.status.idle":"2024-03-18T08:39:18.348328Z","shell.execute_reply.started":"2024-03-18T08:39:18.153110Z","shell.execute_reply":"2024-03-18T08:39:18.347119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Plot histogram before discretization\nplt.figure(figsize=(10, 5))\nplt.subplot(1, 2, 1)\nplt.hist(train_data['eeg_label_offset_seconds'], bins=20, color='blue', alpha=0.7)\nplt.title('Histogram before Discretization')\nplt.xlabel('eeg_label_offset_seconds')\nplt.ylabel('Frequency')\n\n# Plot histogram after discretization\nplt.subplot(1, 2, 2)\nplt.hist(discretized_df['eeg_label_offset_seconds_discretized'], bins=n_bins, color='yellow', alpha=0.7)\nplt.title('Histogram after Discretization')\nplt.xlabel('Discretized bins')\nplt.ylabel('Frequency')\n\n# Adjust layout and show plot\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:18.349755Z","iopub.execute_input":"2024-03-18T08:39:18.350112Z","iopub.status.idle":"2024-03-18T08:39:19.176872Z","shell.execute_reply.started":"2024-03-18T08:39:18.350082Z","shell.execute_reply":"2024-03-18T08:39:19.175456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n\n# Set the style of the seaborn plot\nsns.set_style(\"whitegrid\")\n\n# Plot histogram before discretization\nplt.figure(figsize=(10, 5))\nplt.subplot(1, 2, 1)\nsns.histplot(train_data['eeg_label_offset_seconds'], bins=20, color='blue', alpha=0.7)\nplt.title('Histogram before Discretization')\nplt.xlabel('eeg_label_offset_seconds')\nplt.ylabel('Frequency')\n\n# Plot histogram after discretization\nplt.subplot(1, 2, 2)\nsns.histplot(discretized_df['eeg_label_offset_seconds_discretized'], bins=n_bins, color='red', alpha=0.7)\nplt.title('Histogram after Discretization')\nplt.xlabel('Discretized bins')\nplt.ylabel('Frequency')\n\n# Adjust layout and show plot\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:19.179427Z","iopub.execute_input":"2024-03-18T08:39:19.180692Z","iopub.status.idle":"2024-03-18T08:39:20.272732Z","shell.execute_reply.started":"2024-03-18T08:39:19.180644Z","shell.execute_reply":"2024-03-18T08:39:20.270883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import KBinsDiscretizer\n\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\n\ncontinuous_column = train_data['spectrogram_label_offset_seconds'].values.reshape(-1, 1)\n\n# Initialize the KBinsDiscretizer object\nn_bins = 5  # Number of bins\nstrategy = 'uniform'  # Strategy for binning, e.g., 'uniform' or 'quantile'\ndiscretizer = KBinsDiscretizer(n_bins=n_bins, encode='ordinal', strategy=strategy)\n\n# Fit and transform the continuous data into discrete bins\ndiscretized_data = discretizer.fit_transform(continuous_column)\n\n# Convert the result to a DataFrame for easier visualization\ndiscretized_df = pd.DataFrame(discretized_data, columns=['spectrogram_label_offset_seconds_discretized'])\n\n# Display the first few rows of the DataFrame with the discretized column\nprint(discretized_df.head())\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:20.274722Z","iopub.execute_input":"2024-03-18T08:39:20.275197Z","iopub.status.idle":"2024-03-18T08:39:20.469959Z","shell.execute_reply.started":"2024-03-18T08:39:20.275153Z","shell.execute_reply":"2024-03-18T08:39:20.468738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Create subplots with 1 row and 2 columns\nfig, axes = plt.subplots(1, 2, figsize=(16, 6))\n\n# Plot histogram of the original continuous data\nsns.histplot(train_data['spectrogram_label_offset_seconds'], color='blue', kde=False, ax=axes[0])\naxes[0].set_title('Histogram of Original Data')\naxes[0].set_xlabel('Continuous Values')\naxes[0].set_ylabel('Frequency')\n\n# Plot histogram of the discretized data\nsns.histplot(discretized_df['spectrogram_label_offset_seconds_discretized'], color='red', kde=False, ax=axes[1])\naxes[1].set_title('Histogram of Discretized Data')\naxes[1].set_xlabel('Discretized Values')\naxes[1].set_ylabel('Frequency')\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:20.471601Z","iopub.execute_input":"2024-03-18T08:39:20.472311Z","iopub.status.idle":"2024-03-18T08:39:23.259432Z","shell.execute_reply.started":"2024-03-18T08:39:20.472278Z","shell.execute_reply":"2024-03-18T08:39:23.258030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Create subplots with 1 row and 2 columns\nfig, axes = plt.subplots(1, 2, figsize=(16, 6))\n\n# Plot histogram of the original continuous data\naxes[0].hist(train_data['spectrogram_label_offset_seconds'], color='blue', edgecolor='black', bins=20)\naxes[0].set_title('Histogram of Original Data')\naxes[0].set_xlabel('Continuous Values')\naxes[0].set_ylabel('Frequency')\n\n# Plot histogram of the discretized data\naxes[1].hist(discretized_df['spectrogram_label_offset_seconds_discretized'], color='black', edgecolor='black', bins=n_bins)\naxes[1].set_title('Histogram of Discretized Data')\naxes[1].set_xlabel('Discretized Values')\naxes[1].set_ylabel('Frequency')\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:23.260998Z","iopub.execute_input":"2024-03-18T08:39:23.261811Z","iopub.status.idle":"2024-03-18T08:39:24.070683Z","shell.execute_reply.started":"2024-03-18T08:39:23.261769Z","shell.execute_reply":"2024-03-18T08:39:24.068857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Pandas method for Manual Binning**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\nbin_edges = [0, 10, 20, 30, 40, 50] \n\n# Perform manual binning using pd.cut()\ntrain_data['eeg_label_offset_seconds_binned'] = pd.cut(train_data['eeg_label_offset_seconds'], bins=bin_edges)\n\n# Display the first few rows of the DataFrame with the binned column\nprint(train_data[['eeg_label_offset_seconds', 'eeg_label_offset_seconds_binned']].head())\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:24.072579Z","iopub.execute_input":"2024-03-18T08:39:24.073067Z","iopub.status.idle":"2024-03-18T08:39:24.272161Z","shell.execute_reply.started":"2024-03-18T08:39:24.073025Z","shell.execute_reply":"2024-03-18T08:39:24.270589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Convert Interval objects to midpoints\nbinned_midpoints = train_data['eeg_label_offset_seconds_binned'].apply(lambda x: x.mid)\n\n# Plot histogram of the binned data\nplt.figure(figsize=(8, 5))\nplt.bar(binned_midpoints.value_counts().index, binned_midpoints.value_counts(), color='purple')\nplt.title('Histogram of Binned Data')\nplt.xlabel('Binned Values')\nplt.ylabel('Frequency')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:24.273971Z","iopub.execute_input":"2024-03-18T08:39:24.274356Z","iopub.status.idle":"2024-03-18T08:39:24.734197Z","shell.execute_reply.started":"2024-03-18T08:39:24.274324Z","shell.execute_reply":"2024-03-18T08:39:24.733134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n\n# Convert Interval objects to midpoints\nbinned_midpoints = train_data['eeg_label_offset_seconds_binned'].apply(lambda x: x.mid)\n\n# Plot histogram of the binned data using seaborn\nplt.figure(figsize=(8, 5))\nsns.histplot(binned_midpoints, color='black', kde=False)  # Set kde=False to disable kernel density estimation\nplt.title('Histogram of Binned Data')\nplt.xlabel('Binned Values')\nplt.ylabel('Frequency')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:24.736078Z","iopub.execute_input":"2024-03-18T08:39:24.736585Z","iopub.status.idle":"2024-03-18T08:39:25.354053Z","shell.execute_reply.started":"2024-03-18T08:39:24.736535Z","shell.execute_reply":"2024-03-18T08:39:25.352683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Removing Dummies from the Dataset**","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\ndummy_columns = [col for col in train_data.columns if train_data[col].dtype == 'uint8']\n\n# Drop dummy variable columns from the dataset\ntrain_data = train_data.drop(dummy_columns, axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.355462Z","iopub.execute_input":"2024-03-18T08:39:25.355838Z","iopub.status.idle":"2024-03-18T08:39:25.552711Z","shell.execute_reply.started":"2024-03-18T08:39:25.355789Z","shell.execute_reply":"2024-03-18T08:39:25.551323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Identify dummy variable columns\ndummy_columns = [col for col in train_data.columns if train_data[col].nunique() <= 2]\n\n# Drop dummy variable columns from the dataset\ntrain_data = train_data.drop(dummy_columns, axis=1)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.554177Z","iopub.execute_input":"2024-03-18T08:39:25.554634Z","iopub.status.idle":"2024-03-18T08:39:25.595979Z","shell.execute_reply.started":"2024-03-18T08:39:25.554595Z","shell.execute_reply":"2024-03-18T08:39:25.594695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Performing Logistic Regression on the HMC Dataset**","metadata":{}},{"cell_type":"markdown","source":"`Logistic regression` is a classification algorithm used to assign observations to a discrete set of classes. Unlike linear regression which outputs continuous number values, logistic regression transforms its output using the `logistic sigmoid function `to return a probability value which can then be mapped to two or more discrete classes.\n\n## **Logistic regression can be used for:**\n\n* Binary Classification\n* Multi-class Classification\n* One-vs-Rest Classification","metadata":{}},{"cell_type":"markdown","source":"# **Assumptions of Logistic regression**\n\n* The dependent variable must be categorical in nature.\n* The independent variables(features) must be independent.\n* There should be no outliers in the data. Check for outliers.\n* There should be no high correlations among the independent variables. This can be checked using a correlation matrix.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.metrics import recall_score, precision_score, accuracy_score, classification_report, confusion_matrix","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.597845Z","iopub.execute_input":"2024-03-18T08:39:25.598246Z","iopub.status.idle":"2024-03-18T08:39:25.604756Z","shell.execute_reply.started":"2024-03-18T08:39:25.598216Z","shell.execute_reply":"2024-03-18T08:39:25.603454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the HMC dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.606074Z","iopub.execute_input":"2024-03-18T08:39:25.606464Z","iopub.status.idle":"2024-03-18T08:39:25.803351Z","shell.execute_reply.started":"2024-03-18T08:39:25.606433Z","shell.execute_reply":"2024-03-18T08:39:25.801992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data[['eeg_id', 'eeg_sub_id', 'spectrogram_id', 'label_id', 'patient_id']]","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.804937Z","iopub.execute_input":"2024-03-18T08:39:25.805319Z","iopub.status.idle":"2024-03-18T08:39:25.811838Z","shell.execute_reply.started":"2024-03-18T08:39:25.805290Z","shell.execute_reply":"2024-03-18T08:39:25.810698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.813249Z","iopub.execute_input":"2024-03-18T08:39:25.813668Z","iopub.status.idle":"2024-03-18T08:39:25.823708Z","shell.execute_reply.started":"2024-03-18T08:39:25.813637Z","shell.execute_reply":"2024-03-18T08:39:25.822660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.825346Z","iopub.execute_input":"2024-03-18T08:39:25.826566Z","iopub.status.idle":"2024-03-18T08:39:25.851446Z","shell.execute_reply.started":"2024-03-18T08:39:25.826531Z","shell.execute_reply":"2024-03-18T08:39:25.849951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Preprocess the data using StandardScaler\nscaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_test_scaled = scaler.transform(X_test)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.853389Z","iopub.execute_input":"2024-03-18T08:39:25.854128Z","iopub.status.idle":"2024-03-18T08:39:25.873468Z","shell.execute_reply.started":"2024-03-18T08:39:25.854084Z","shell.execute_reply":"2024-03-18T08:39:25.871476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create and train the logistic regression model\nmodel = LogisticRegression()\nmodel.fit(X_train_scaled, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:25.876132Z","iopub.execute_input":"2024-03-18T08:39:25.877203Z","iopub.status.idle":"2024-03-18T08:39:27.225941Z","shell.execute_reply.started":"2024-03-18T08:39:25.877152Z","shell.execute_reply":"2024-03-18T08:39:27.224753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions\ny_pred = model.predict(X_test_scaled)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:27.227539Z","iopub.execute_input":"2024-03-18T08:39:27.228252Z","iopub.status.idle":"2024-03-18T08:39:27.241045Z","shell.execute_reply.started":"2024-03-18T08:39:27.228206Z","shell.execute_reply":"2024-03-18T08:39:27.238820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate evaluation metrics\nrecall = recall_score(y_test, y_pred, average='weighted')\nprecision = precision_score(y_test, y_pred, average='weighted')\naccuracy = accuracy_score(y_test, y_pred)\nclassification_report_output = classification_report(y_test, y_pred)\nconf_matrix = confusion_matrix(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:27.247077Z","iopub.execute_input":"2024-03-18T08:39:27.253555Z","iopub.status.idle":"2024-03-18T08:39:29.547299Z","shell.execute_reply.started":"2024-03-18T08:39:27.253469Z","shell.execute_reply":"2024-03-18T08:39:29.546203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print evaluation metrics\nprint(\"Recall Score:\", recall)\nprint(\"Precision Score:\", precision)\nprint(\"Accuracy Score:\", accuracy)\nprint(\"Classification Report:\\n\", classification_report_output)\nprint(\"Confusion Matrix:\\n\", conf_matrix)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:29.548696Z","iopub.execute_input":"2024-03-18T08:39:29.549024Z","iopub.status.idle":"2024-03-18T08:39:29.555455Z","shell.execute_reply.started":"2024-03-18T08:39:29.548996Z","shell.execute_reply":"2024-03-18T08:39:29.554149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Plot confusion matrix\nplt.figure(figsize=(8, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel(\"Predicted Labels\")\nplt.ylabel(\"True Labels\")\nplt.title(\"Confusion Matrix\")\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:29.557081Z","iopub.execute_input":"2024-03-18T08:39:29.557536Z","iopub.status.idle":"2024-03-18T08:39:29.926586Z","shell.execute_reply.started":"2024-03-18T08:39:29.557496Z","shell.execute_reply":"2024-03-18T08:39:29.925195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from joblib import dump\n\n# Assuming 'model' is the trained logistic regression model\n# Save the model to a file\ndump(model, 'logistic_regression_model.joblib')","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:29.928183Z","iopub.execute_input":"2024-03-18T08:39:29.928676Z","iopub.status.idle":"2024-03-18T08:39:29.939783Z","shell.execute_reply.started":"2024-03-18T08:39:29.928632Z","shell.execute_reply":"2024-03-18T08:39:29.938346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Creating a Machine Learning model Using Support Vector Machines (SVM)**","metadata":{}},{"cell_type":"code","source":"# import libraries\nimport pandas as pd\nimport matplotlib.pyplot as plt \nimport seaborn as sns\n\n# import machine learning libraries\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.svm import SVC\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.metrics import recall_score, precision_score, accuracy_score, classification_report, confusion_matrix, r2_score","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:29.941460Z","iopub.execute_input":"2024-03-18T08:39:29.941880Z","iopub.status.idle":"2024-03-18T08:39:29.952433Z","shell.execute_reply.started":"2024-03-18T08:39:29.941849Z","shell.execute_reply":"2024-03-18T08:39:29.951069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# load the HMC dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:29.953964Z","iopub.execute_input":"2024-03-18T08:39:29.954381Z","iopub.status.idle":"2024-03-18T08:39:30.159415Z","shell.execute_reply.started":"2024-03-18T08:39:29.954332Z","shell.execute_reply":"2024-03-18T08:39:30.158182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data[['eeg_id', 'eeg_sub_id', 'spectrogram_id', 'label_id', 'patient_id']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:30.161111Z","iopub.execute_input":"2024-03-18T08:39:30.162272Z","iopub.status.idle":"2024-03-18T08:39:30.169940Z","shell.execute_reply.started":"2024-03-18T08:39:30.162231Z","shell.execute_reply":"2024-03-18T08:39:30.168652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# split the data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:30.171198Z","iopub.execute_input":"2024-03-18T08:39:30.171635Z","iopub.status.idle":"2024-03-18T08:39:30.194077Z","shell.execute_reply.started":"2024-03-18T08:39:30.171590Z","shell.execute_reply":"2024-03-18T08:39:30.192748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# pre-process the data using StandardScaler\nscaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:30.195672Z","iopub.execute_input":"2024-03-18T08:39:30.196306Z","iopub.status.idle":"2024-03-18T08:39:30.213292Z","shell.execute_reply.started":"2024-03-18T08:39:30.196271Z","shell.execute_reply":"2024-03-18T08:39:30.211985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# build the model\nmodel = SVC(kernel='rbf')\n\n# train the model\nmodel.fit(X_train_scaled, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:39:30.215123Z","iopub.execute_input":"2024-03-18T08:39:30.215489Z","iopub.status.idle":"2024-03-18T08:47:25.645230Z","shell.execute_reply.started":"2024-03-18T08:39:30.215459Z","shell.execute_reply":"2024-03-18T08:47:25.643965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# make predictions\ny_pred = model.predict(X_test_scaled)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:47:25.646621Z","iopub.execute_input":"2024-03-18T08:47:25.646984Z","iopub.status.idle":"2024-03-18T08:49:08.236260Z","shell.execute_reply.started":"2024-03-18T08:47:25.646952Z","shell.execute_reply":"2024-03-18T08:49:08.234889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calculate evaluation metrics\nrecall = recall_score(y_test, y_pred, average='weighted')\nprecision = precision_score(y_test, y_pred, average='weighted')\naccuracy = accuracy_score(y_test, y_pred)\nclassification_report_output = classification_report(y_test, y_pred)\nconf_matrix = confusion_matrix(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:08.237665Z","iopub.execute_input":"2024-03-18T08:49:08.238061Z","iopub.status.idle":"2024-03-18T08:49:10.461454Z","shell.execute_reply.started":"2024-03-18T08:49:08.238030Z","shell.execute_reply":"2024-03-18T08:49:10.459886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print evaluation metrics\nprint(\"Recall Score:\", recall)\nprint(\"Precision Score:\", precision)\nprint(\"Accuracy Score:\", accuracy)\nprint(\"Classification Report:\\n\", classification_report_output)\nprint(\"Confusion Matrix:\\n\", conf_matrix)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:10.463276Z","iopub.execute_input":"2024-03-18T08:49:10.463704Z","iopub.status.idle":"2024-03-18T08:49:10.471947Z","shell.execute_reply.started":"2024-03-18T08:49:10.463673Z","shell.execute_reply":"2024-03-18T08:49:10.470459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot confusion matrix\nplt.figure(figsize=(8, 6))\nsns.heatmap(conf_matrix, annot=True, cmap='Blues', fmt='g', \n            xticklabels=['Predicted Negative', 'Predicted Positive'],\n            yticklabels=['Actual Negative', 'Actual Positive'])\nplt.xlabel('Predicted label')\nplt.ylabel('True label')\nplt.title('Confusion Matrix')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:10.473295Z","iopub.execute_input":"2024-03-18T08:49:10.473724Z","iopub.status.idle":"2024-03-18T08:49:11.462024Z","shell.execute_reply.started":"2024-03-18T08:49:10.473692Z","shell.execute_reply":"2024-03-18T08:49:11.460734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **What this Confusion matrix shows**\nThe confusion matrix provides a summary of the performance of a classification model. It showcases the counts of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) predictions made by the model.\n\n> Here's what each part of the confusion matrix represents:\n\n* True Positive (TP): The number of instances correctly predicted as positive.\n\n* True Negative (TN): The number of instances correctly predicted as negative.\n\n* False Positive (FP): The number of instances incorrectly predicted as positive (actually negative).\n\n* False Negative (FN): The number of instances incorrectly predicted as negative (actually positive).\n\nIn the confusion matrix visualization, the rows represent the actual classes, while the columns represent the predicted classes. This allows for easy interpretation of the model's performance in terms of correctly and incorrectly classified instances.","metadata":{}},{"cell_type":"markdown","source":"# **Creating a Machine Learning model with K-Nearest Neighbour**","metadata":{}},{"cell_type":"code","source":"# import libraries\nimport pandas as pd\nimport numpy as np \nimport matplotlib.pyplot as plt \nimport seaborn as sns \n\n# ml libraries\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import classification_report\nfrom sklearn.model_selection import GridSearchCV","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:11.463353Z","iopub.execute_input":"2024-03-18T08:49:11.463742Z","iopub.status.idle":"2024-03-18T08:49:11.469965Z","shell.execute_reply.started":"2024-03-18T08:49:11.463712Z","shell.execute_reply":"2024-03-18T08:49:11.468972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# load the HMC dataset\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:11.471135Z","iopub.execute_input":"2024-03-18T08:49:11.471540Z","iopub.status.idle":"2024-03-18T08:49:11.697216Z","shell.execute_reply.started":"2024-03-18T08:49:11.471508Z","shell.execute_reply":"2024-03-18T08:49:11.695282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# rrop irrelevant columns\ntrain_data.drop(['eeg_id', 'eeg_sub_id', 'spectrogram_id', 'spectrogram_sub_id', 'label_id', 'patient_id'], axis=1, inplace=True)\n\n# handle missing data if any\ntrain_data.dropna(inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:11.698920Z","iopub.execute_input":"2024-03-18T08:49:11.699481Z","iopub.status.idle":"2024-03-18T08:49:11.728262Z","shell.execute_reply.started":"2024-03-18T08:49:11.699422Z","shell.execute_reply":"2024-03-18T08:49:11.726439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# split features and target variable\nX = train_data.drop(['expert_consensus'], axis=1)\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:11.729840Z","iopub.execute_input":"2024-03-18T08:49:11.731307Z","iopub.status.idle":"2024-03-18T08:49:11.739414Z","shell.execute_reply.started":"2024-03-18T08:49:11.731251Z","shell.execute_reply":"2024-03-18T08:49:11.737643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# split the data into training and testing\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:11.740999Z","iopub.execute_input":"2024-03-18T08:49:11.741398Z","iopub.status.idle":"2024-03-18T08:49:11.771338Z","shell.execute_reply.started":"2024-03-18T08:49:11.741352Z","shell.execute_reply":"2024-03-18T08:49:11.769787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# build the model\nknn = KNeighborsClassifier(n_neighbors=5) \n\n# train the model\nknn.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:11.773211Z","iopub.execute_input":"2024-03-18T08:49:11.774340Z","iopub.status.idle":"2024-03-18T08:49:12.158717Z","shell.execute_reply.started":"2024-03-18T08:49:11.774287Z","shell.execute_reply":"2024-03-18T08:49:12.157449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# evaluate the model\ny_pred = knn.predict(X_test)\nprint(classification_report(y_test, y_pred))","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:12.160219Z","iopub.execute_input":"2024-03-18T08:49:12.160743Z","iopub.status.idle":"2024-03-18T08:49:15.501946Z","shell.execute_reply.started":"2024-03-18T08:49:12.160701Z","shell.execute_reply":"2024-03-18T08:49:15.500195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# perform model tuning\nparam_grid = {'n_neighbors': [3, 5, 7, 9, 11]}\ngrid_search = GridSearchCV(KNeighborsClassifier(), param_grid, cv=5)\ngrid_search.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:49:15.504577Z","iopub.execute_input":"2024-03-18T08:49:15.505104Z","iopub.status.idle":"2024-03-18T08:50:03.043824Z","shell.execute_reply.started":"2024-03-18T08:49:15.505063Z","shell.execute_reply":"2024-03-18T08:50:03.041070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_k = grid_search.best_params_['n_neighbors']\nprint(\"Best K value:\", best_k)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:50:03.045535Z","iopub.execute_input":"2024-03-18T08:50:03.045924Z","iopub.status.idle":"2024-03-18T08:50:03.053128Z","shell.execute_reply.started":"2024-03-18T08:50:03.045894Z","shell.execute_reply":"2024-03-18T08:50:03.051249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model deployment\nbest_knn = KNeighborsClassifier(n_neighbors=best_k)\nbest_knn.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:50:03.055326Z","iopub.execute_input":"2024-03-18T08:50:03.055789Z","iopub.status.idle":"2024-03-18T08:50:03.447597Z","shell.execute_reply.started":"2024-03-18T08:50:03.055757Z","shell.execute_reply":"2024-03-18T08:50:03.445698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# visualize the performance of different values of K during tuning\nresults = grid_search.cv_results_\nplt.figure(figsize=(10, 6))\nplt.plot(param_grid['n_neighbors'], results['mean_test_score'], marker='o', linestyle='-')\nplt.title('Grid Search Results')\nplt.xlabel('Number of Neighbors (K)')\nplt.ylabel('Mean Test Score')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:50:03.449208Z","iopub.execute_input":"2024-03-18T08:50:03.449860Z","iopub.status.idle":"2024-03-18T08:50:03.886217Z","shell.execute_reply.started":"2024-03-18T08:50:03.449811Z","shell.execute_reply":"2024-03-18T08:50:03.884875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# visualize the classification report\nreport = classification_report(y_test, y_pred, output_dict=True)\ndf_report = pd.DataFrame(report).transpose()\nplt.figure(figsize=(10, 6))\nsns.heatmap(df_report.iloc[:-1, :-1], annot=True, cmap='YlGnBu')\nplt.title('Classification Report')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-18T08:50:03.888105Z","iopub.execute_input":"2024-03-18T08:50:03.888916Z","iopub.status.idle":"2024-03-18T08:50:05.785756Z","shell.execute_reply.started":"2024-03-18T08:50:03.888788Z","shell.execute_reply":"2024-03-18T08:50:05.784310Z"},"trusted":true},"execution_count":null,"outputs":[]}]}