{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Harmful Brain Activity \n#  Decision Trees (DTs) : \nWhat is a decision tree?\nA decision tree is a non-parametric supervised learning algorithm, which is utilized for both classification and regression tasks. It has a hierarchical, tree structure, which consists of a root node, branches, internal nodes and leaf nodes.","metadata":{}},{"cell_type":"markdown","source":"# exploratory Data Analysis (EDA):\nis crucial for developing effective machine learning models. This article discusses the various techniques and methods used in EDA, such as scatter plots, histograms, box plots, and descriptive statistics, to identify trends and patterns in datasets. It also explains how to handle missing values and transform categorical and numerical data for optimal data preparation.\n\nAdditionally, the article covers the process of model selection and optimization, including hyperparameter tuning using GridSearchCV, evaluating model performance using metrics like accuracy, classification report, confusion matrix, and ROC-AUC score, and cross-validation. By following these steps, readers can create accurate and effective machine learning solutions.\n\n# Why is it important to perform EDA?\n# Methods and techniques of EDA:\nThere are several techniques and methods for performing an EDA, such as scatter plots, histograms, box plots, and descriptive statistics. The choice of techniques depends on the nature of the data and the goal of the analysis.\n\n# Data visualization:\nData visualization is a powerful tool for identifying trends and patterns in datasets. Charts such as lines, bars, scatter, and box plots facilitate the identification of relationships between variables, frequency distributions, and the presence of outliers.","metadata":{}},{"cell_type":"markdown","source":"# Import libraries\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport numpy as np \n\nimport matplotlib.pyplot as plt \nimport seaborn as sns \n\nimport tensorflow as tf \n\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\n\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout\n\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\n\nfrom sklearn.svm import SVC\nfrom sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:35.260473Z","iopub.execute_input":"2024-03-20T08:25:35.261052Z","iopub.status.idle":"2024-03-20T08:25:35.273750Z","shell.execute_reply.started":"2024-03-20T08:25:35.260994Z","shell.execute_reply":"2024-03-20T08:25:35.271403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the `train.csv` dataset","metadata":{}},{"cell_type":"code","source":"\ntrain_data = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:35.275883Z","iopub.execute_input":"2024-03-20T08:25:35.276734Z","iopub.status.idle":"2024-03-20T08:25:35.602877Z","shell.execute_reply.started":"2024-03-20T08:25:35.276683Z","shell.execute_reply":"2024-03-20T08:25:35.601533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Print the 10 rows of the dataset\n","metadata":{}},{"cell_type":"code","source":"train_data.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:35.604275Z","iopub.execute_input":"2024-03-20T08:25:35.604760Z","iopub.status.idle":"2024-03-20T08:25:35.641979Z","shell.execute_reply.started":"2024-03-20T08:25:35.604717Z","shell.execute_reply":"2024-03-20T08:25:35.640495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Remove all the warnings\n","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:35.644802Z","iopub.execute_input":"2024-03-20T08:25:35.645171Z","iopub.status.idle":"2024-03-20T08:25:35.650739Z","shell.execute_reply.started":"2024-03-20T08:25:35.645125Z","shell.execute_reply":"2024-03-20T08:25:35.649295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check for Missing values and Duplicated columns","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:35.652224Z","iopub.execute_input":"2024-03-20T08:25:35.652580Z","iopub.status.idle":"2024-03-20T08:25:35.707730Z","shell.execute_reply.started":"2024-03-20T08:25:35.652542Z","shell.execute_reply":"2024-03-20T08:25:35.706574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot the missing values \n","metadata":{}},{"cell_type":"code","source":"train_data.isna().sum().plot(kind='bar')\nplt.title('Missing values in the dataset')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:35.709520Z","iopub.execute_input":"2024-03-20T08:25:35.710252Z","iopub.status.idle":"2024-03-20T08:25:36.156687Z","shell.execute_reply.started":"2024-03-20T08:25:35.710208Z","shell.execute_reply":"2024-03-20T08:25:36.155357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Missing values before preprocessing:\n","metadata":{}},{"cell_type":"code","source":"# Data preprocessing steps\n\n# 1. Handling missing values (if any)\nprint(\"Missing values before preprocessing:\")\nprint(train_data.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:36.158344Z","iopub.execute_input":"2024-03-20T08:25:36.159188Z","iopub.status.idle":"2024-03-20T08:25:36.182606Z","shell.execute_reply.started":"2024-03-20T08:25:36.159122Z","shell.execute_reply":"2024-03-20T08:25:36.181201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check for duplicated columns\n","metadata":{}},{"cell_type":"code","source":"train_data.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:36.184433Z","iopub.execute_input":"2024-03-20T08:25:36.184782Z","iopub.status.idle":"2024-03-20T08:25:36.239380Z","shell.execute_reply.started":"2024-03-20T08:25:36.184754Z","shell.execute_reply":"2024-03-20T08:25:36.238243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing the Data","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Define custom colors for each bar\ncustom_colors = [\"#1f77b4\", \"#ff7f0e\", \"#2ca02c\", \"#d62728\", \"#9467bd\", \"#8c564b\"]\n\n# Visualize distribution of target variable\nplt.figure(figsize=(8, 6))\nsns.countplot(x='expert_consensus', data=train_data, palette=custom_colors)\nplt.title('Distribution of Expert Consensus Labels')\nplt.xlabel('Expert Consensus Labels')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:36.241062Z","iopub.execute_input":"2024-03-20T08:25:36.241567Z","iopub.status.idle":"2024-03-20T08:25:36.690842Z","shell.execute_reply.started":"2024-03-20T08:25:36.241525Z","shell.execute_reply":"2024-03-20T08:25:36.688600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encode the `expert_consensus` column\n","metadata":{}},{"cell_type":"code","source":"# Encode the `expert_consensus` column\nlabel_encoder = LabelEncoder()\ntrain_data['expert_consensus'] = label_encoder.fit_transform(train_data['expert_consensus'])","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:36.697087Z","iopub.execute_input":"2024-03-20T08:25:36.697632Z","iopub.status.idle":"2024-03-20T08:25:36.739510Z","shell.execute_reply.started":"2024-03-20T08:25:36.697585Z","shell.execute_reply":"2024-03-20T08:25:36.738242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:36.741224Z","iopub.execute_input":"2024-03-20T08:25:36.741808Z","iopub.status.idle":"2024-03-20T08:25:36.763606Z","shell.execute_reply.started":"2024-03-20T08:25:36.741767Z","shell.execute_reply":"2024-03-20T08:25:36.762254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating   RandomForestClassifier and Decision Tree Models and comparing them\n\nRandomForestClassifier - это популярный алгоритм машинного обучения из библиотеки scikit-learn в Python, основанный на ансамблевом методе. Он использует множество решающих деревьев для улучшения предсказательной способности и уменьшения риска переобучения.,\n\nМодели решающих деревьев (Decision Tree Models) представляют собой важный и широко используемый метод машинного обучения. Они основываются на структуре дерева, где узлы представляют собой признаки, ветви - результаты проверки этих признаков, а листья - конечные решения или прогнозы.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:36.765245Z","iopub.execute_input":"2024-03-20T08:25:36.765708Z","iopub.status.idle":"2024-03-20T08:25:36.776276Z","shell.execute_reply.started":"2024-03-20T08:25:36.765665Z","shell.execute_reply":"2024-03-20T08:25:36.774829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the dataset\n","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:36.778332Z","iopub.execute_input":"2024-03-20T08:25:36.778703Z","iopub.status.idle":"2024-03-20T08:25:37.020381Z","shell.execute_reply.started":"2024-03-20T08:25:36.778673Z","shell.execute_reply":"2024-03-20T08:25:37.018732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data[['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds', 'patient_id']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:37.022057Z","iopub.execute_input":"2024-03-20T08:25:37.022547Z","iopub.status.idle":"2024-03-20T08:25:37.034085Z","shell.execute_reply.started":"2024-03-20T08:25:37.022508Z","shell.execute_reply":"2024-03-20T08:25:37.032653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Split the dataset into training and testing sets\n","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:37.036188Z","iopub.execute_input":"2024-03-20T08:25:37.036887Z","iopub.status.idle":"2024-03-20T08:25:37.069688Z","shell.execute_reply.started":"2024-03-20T08:25:37.036820Z","shell.execute_reply":"2024-03-20T08:25:37.067935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Scaling numerical features\n","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:37.071297Z","iopub.execute_input":"2024-03-20T08:25:37.071700Z","iopub.status.idle":"2024-03-20T08:25:37.093469Z","shell.execute_reply.started":"2024-03-20T08:25:37.071665Z","shell.execute_reply":"2024-03-20T08:25:37.091702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_encoded = pd.get_dummies(X_train, columns=['patient_id'])\nX_test_encoded = pd.get_dummies(X_test, columns=['patient_id'])","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:37.095279Z","iopub.execute_input":"2024-03-20T08:25:37.095749Z","iopub.status.idle":"2024-03-20T08:25:37.373493Z","shell.execute_reply.started":"2024-03-20T08:25:37.095704Z","shell.execute_reply":"2024-03-20T08:25:37.372130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KNeighborsClassifier()\n\nKNeighborsClassifier - это алгоритм машинного обучения из библиотеки scikit-learn в Python, который относится к методам обучения с запоминанием. Он не строит явную модель, а хранит тренировочные данные и использует их для предсказания новых данных на основе ближайших соседей.","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import confusion_matrix\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Assuming you have initialized and trained your KNeighborsClassifier\nknn = KNeighborsClassifier()\nknn.fit(X_train, y_train)  # Make sure you have fitted your classifier\n\n# Make predictions\npredictions = knn.predict(X_test)\n\n# Calculate confusion matrix\ncm = confusion_matrix(y_test, predictions)\n\n# Visualize confusion matrix\nplt.figure(figsize=(8, 6))\nsns.heatmap(cm, annot=True, cmap='Blues', fmt='g', \n            xticklabels=knn.classes_, yticklabels=knn.classes_)\nplt.title('Confusion Matrix')\nplt.xlabel('Predicted Labels')\nplt.ylabel('True Labels')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:37.374987Z","iopub.execute_input":"2024-03-20T08:25:37.375459Z","iopub.status.idle":"2024-03-20T08:25:39.797391Z","shell.execute_reply.started":"2024-03-20T08:25:37.375409Z","shell.execute_reply":"2024-03-20T08:25:39.795983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a model with Decision-Tree Algorithm","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:39.798980Z","iopub.execute_input":"2024-03-20T08:25:39.799360Z","iopub.status.idle":"2024-03-20T08:25:39.804669Z","shell.execute_reply.started":"2024-03-20T08:25:39.799328Z","shell.execute_reply":"2024-03-20T08:25:39.803622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:39.805973Z","iopub.execute_input":"2024-03-20T08:25:39.806940Z","iopub.status.idle":"2024-03-20T08:25:40.030737Z","shell.execute_reply.started":"2024-03-20T08:25:39.806903Z","shell.execute_reply":"2024-03-20T08:25:40.029725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select features and target variable\nX = train_data[['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.032331Z","iopub.execute_input":"2024-03-20T08:25:40.032948Z","iopub.status.idle":"2024-03-20T08:25:40.039760Z","shell.execute_reply.started":"2024-03-20T08:25:40.032915Z","shell.execute_reply":"2024-03-20T08:25:40.038873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.040958Z","iopub.execute_input":"2024-03-20T08:25:40.042326Z","iopub.status.idle":"2024-03-20T08:25:40.069850Z","shell.execute_reply.started":"2024-03-20T08:25:40.042287Z","shell.execute_reply":"2024-03-20T08:25:40.068681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the decision tree classifier\ndt_classifier = DecisionTreeClassifier()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.071578Z","iopub.execute_input":"2024-03-20T08:25:40.071923Z","iopub.status.idle":"2024-03-20T08:25:40.077658Z","shell.execute_reply.started":"2024-03-20T08:25:40.071894Z","shell.execute_reply":"2024-03-20T08:25:40.076037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the classifier on the training data\ndt_classifier.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.079861Z","iopub.execute_input":"2024-03-20T08:25:40.080325Z","iopub.status.idle":"2024-03-20T08:25:40.584023Z","shell.execute_reply.started":"2024-03-20T08:25:40.080281Z","shell.execute_reply":"2024-03-20T08:25:40.583128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Evaluate the performance of the trained model on the testing data\npredictions = dt_classifier.predict(X_test)\n\naccuracy = accuracy_score(y_test, predictions)\nprint(\"Accuracy:\", accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.585651Z","iopub.execute_input":"2024-03-20T08:25:40.586344Z","iopub.status.idle":"2024-03-20T08:25:40.646561Z","shell.execute_reply.started":"2024-03-20T08:25:40.586312Z","shell.execute_reply":"2024-03-20T08:25:40.645216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# RandomForestClassifier","metadata":{}},{"cell_type":"code","source":"import os\nimport tqdm\nimport pandas as pd\nimport numpy as np\nfrom sklearn.ensemble import RandomForestClassifier","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.647850Z","iopub.execute_input":"2024-03-20T08:25:40.648223Z","iopub.status.idle":"2024-03-20T08:25:40.660518Z","shell.execute_reply.started":"2024-03-20T08:25:40.648192Z","shell.execute_reply":"2024-03-20T08:25:40.659108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# parent directory\nPDIR = '/kaggle/input/hms-harmful-brain-activity-classification'","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.662389Z","iopub.execute_input":"2024-03-20T08:25:40.662769Z","iopub.status.idle":"2024-03-20T08:25:40.668467Z","shell.execute_reply.started":"2024-03-20T08:25:40.662737Z","shell.execute_reply":"2024-03-20T08:25:40.667048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading the CSV file 'train.csv' located in the directory specified by PDIR\ndf = pd.read_csv(os.path.join(PDIR, 'train.csv'))\n# Load the dataset\ntrain_data = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\n# Displaying the first few rows of the DataFrame\ndisplay(df.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:27:00.557978Z","iopub.execute_input":"2024-03-20T08:27:00.558399Z","iopub.status.idle":"2024-03-20T08:27:00.956468Z","shell.execute_reply.started":"2024-03-20T08:27:00.558360Z","shell.execute_reply":"2024-03-20T08:27:00.955196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select features and target variable\nX = train_data[['eeg_label_offset_seconds', 'spectrogram_label_offset_seconds', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote']]\ny = train_data['expert_consensus']","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:27:28.205422Z","iopub.execute_input":"2024-03-20T08:27:28.205857Z","iopub.status.idle":"2024-03-20T08:27:28.214095Z","shell.execute_reply.started":"2024-03-20T08:27:28.205823Z","shell.execute_reply":"2024-03-20T08:27:28.212448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:27:57.378691Z","iopub.execute_input":"2024-03-20T08:27:57.379097Z","iopub.status.idle":"2024-03-20T08:27:57.401770Z","shell.execute_reply.started":"2024-03-20T08:27:57.379066Z","shell.execute_reply":"2024-03-20T08:27:57.400386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_Model = RandomForestClassifier()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:40:34.970225Z","iopub.execute_input":"2024-03-20T08:40:34.970650Z","iopub.status.idle":"2024-03-20T08:40:34.976972Z","shell.execute_reply.started":"2024-03-20T08:40:34.970619Z","shell.execute_reply":"2024-03-20T08:40:34.975368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_Model.fit(X_train,y_train)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:40:44.348539Z","iopub.execute_input":"2024-03-20T08:40:44.348963Z","iopub.status.idle":"2024-03-20T08:40:54.351904Z","shell.execute_reply.started":"2024-03-20T08:40:44.348924Z","shell.execute_reply":"2024-03-20T08:40:54.350666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print (f'Train Accuracy - : {rf_Model.score(X_train,y_train):.3f}')\nprint (f'Test Accuracy - : {rf_Model.score(X_test,y_test):.3f}')","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:41:05.478028Z","iopub.execute_input":"2024-03-20T08:41:05.478483Z","iopub.status.idle":"2024-03-20T08:41:07.922547Z","shell.execute_reply.started":"2024-03-20T08:41:05.478445Z","shell.execute_reply":"2024-03-20T08:41:07.921013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Setting the sampling frequency and duration for EEG data collection","metadata":{}},{"cell_type":"code","source":"\nsampling_frequency = 200  # Sampling frequency in Hz\ndata_collection_duration = 50  # Duration of EEG data collection in seconds\ntotal_samples = sampling_frequency * data_collection_duration  # Total number of samples in the duration\n\n# Setting the number of training data points\nnum_train_data_points = 500  \n\n# Creating an empty DataFrame to store training data\ntraining_data_df = pd.DataFrame()\n\n# Iterating over each training data point\nfor i in tqdm.tqdm(range(num_train_data_points)):\n    # Loading EEG data for a specified eeg_id\n    eeg_id = df.loc[i, 'eeg_id']\n    eeg_data = pd.read_parquet(os.path.join(PDIR, 'train_eegs', f'{eeg_id}.parquet'))\n    \n    # Extracting EEG data from the Cz electrode for 50 seconds\n    label_offset_time = df.loc[i, 'eeg_label_offset_seconds']  # Offset time for the EEG label\n    label_offset_index = int(sampling_frequency * label_offset_time)  # Calculating offset index\n    cz_electrode_data = eeg_data['Cz'][label_offset_index:label_offset_index + total_samples]  # Extracting data for Cz electrode\n    \n    # Adding the extracted data as a row to the training DataFrame\n    training_data_df = pd.concat([training_data_df, cz_electrode_data.reset_index(drop=True).to_frame().transpose()], axis=0)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:40.897246Z","iopub.execute_input":"2024-03-20T08:25:40.897611Z","iopub.status.idle":"2024-03-20T08:25:50.698391Z","shell.execute_reply.started":"2024-03-20T08:25:40.897579Z","shell.execute_reply":"2024-03-20T08:25:50.696934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare features (X_train) and target variable (y_train)","metadata":{}},{"cell_type":"code","source":"# Adding diagnosis results\ntraining_data_df['expert_consensus'] = df[:num_train_data_points]['expert_consensus'].values\n\n# Removing rows with missing values\ntraining_data_df = training_data_df.dropna()\ntraining_data_df = training_data_df.reset_index(drop=True)\n\n# Separating data into features and target\ny_train = training_data_df['expert_consensus']\nX_train = training_data_df.drop('expert_consensus', axis=1)\n\n# Displaying the first few rows of the feature  dataset\ndisplay(X_train.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:50.699836Z","iopub.execute_input":"2024-03-20T08:25:50.700183Z","iopub.status.idle":"2024-03-20T08:25:50.774857Z","shell.execute_reply.started":"2024-03-20T08:25:50.700140Z","shell.execute_reply":"2024-03-20T08:25:50.773454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train RandomForestClassifier()\n","metadata":{}},{"cell_type":"code","source":"# Initializing a RandomForestClassifier with a random state of 0\nforest = RandomForestClassifier(random_state=0)\n\n# Fitting the classifier to the training data\nforest.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:50.776797Z","iopub.execute_input":"2024-03-20T08:25:50.777710Z","iopub.status.idle":"2024-03-20T08:25:53.563648Z","shell.execute_reply.started":"2024-03-20T08:25:50.777671Z","shell.execute_reply":"2024-03-20T08:25:53.562446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading the CSV file 'test.csv' located in the directory specified by PDIR","metadata":{}},{"cell_type":"code","source":"\ndf_test = pd.read_csv(os.path.join(PDIR, 'test.csv'))\n\n# Displaying the first few rows of the DataFrame\ndisplay(df_test.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:53.565123Z","iopub.execute_input":"2024-03-20T08:25:53.565498Z","iopub.status.idle":"2024-03-20T08:25:53.582627Z","shell.execute_reply.started":"2024-03-20T08:25:53.565467Z","shell.execute_reply":"2024-03-20T08:25:53.581428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating an empty DataFrame to store testing data","metadata":{}},{"cell_type":"code","source":"\nX_test = pd.DataFrame()\n\n# Iterating over each test data point\nfor i in tqdm.tqdm(range(len(df_test))):\n    # Loading EEG data for a specified eeg_id\n    eeg_id_ = df_test.loc[i, 'eeg_id']\n    tmp = pd.read_parquet(os.path.join(PDIR, 'test_eegs', f'{eeg_id_}.parquet'))\n    \n    # Extracting EEG data from the Cz electrode\n    cz_electrode_data = tmp['Cz']\n    \n    # Adding the extracted data as a row to the testing DataFrame\n    X_test = pd.concat([X_test, cz_electrode_data.reset_index(drop=True).to_frame().transpose()], axis=0)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:53.584249Z","iopub.execute_input":"2024-03-20T08:25:53.584606Z","iopub.status.idle":"2024-03-20T08:25:53.630013Z","shell.execute_reply.started":"2024-03-20T08:25:53.584576Z","shell.execute_reply":"2024-03-20T08:25:53.628474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calculate predictions using the trained RandomForestClassifier model\n","metadata":{}},{"cell_type":"code","source":"predictions = forest.predict_proba(X_test)\n\n# Read the sample submission file\nsubmission = pd.read_csv(f'{PDIR}/sample_submission.csv')\n\n# Iterate over each test data point\nfor i in tqdm.tqdm(range(len(df_test))):\n    # Set the 'eeg_id' in the submission DataFrame\n    submission.loc[i, 'eeg_id'] = df_test.loc[i, 'eeg_id']\n    \n    # Set the probability for each class in the submission DataFrame\n    for j, cls_name in enumerate(forest.classes_):\n        submission.loc[i, f'{cls_name.lower()}_vote'] = predictions[i, j]","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:53.631505Z","iopub.execute_input":"2024-03-20T08:25:53.631852Z","iopub.status.idle":"2024-03-20T08:25:53.800808Z","shell.execute_reply.started":"2024-03-20T08:25:53.631823Z","shell.execute_reply":"2024-03-20T08:25:53.799477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Display the submission DataFrame","metadata":{}},{"cell_type":"code","source":"display(submission)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:53.802443Z","iopub.execute_input":"2024-03-20T08:25:53.802769Z","iopub.status.idle":"2024-03-20T08:25:53.815995Z","shell.execute_reply.started":"2024-03-20T08:25:53.802741Z","shell.execute_reply":"2024-03-20T08:25:53.814838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Saving the submission DataFrame to a CSV file without including the index\n","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T08:25:53.817703Z","iopub.execute_input":"2024-03-20T08:25:53.818113Z","iopub.status.idle":"2024-03-20T08:25:53.833088Z","shell.execute_reply.started":"2024-03-20T08:25:53.818077Z","shell.execute_reply":"2024-03-20T08:25:53.831741Z"},"trusted":true},"execution_count":null,"outputs":[]}]}