{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![ML](https://raw.githubusercontent.com/AniMilina/Parkinson-s-Freezing-of-Gait-Prediction/main/ML.jpg)\n","metadata":{}},{"cell_type":"markdown","source":"To detect episodes of freezing of gait (FOG) and classify their types along the timeline, we have the following dataset available:\n\n* Unique patient identifier (stored in the 'Id' column)  \n* Patient number (stored in the 'Subject' column)  \n* Visit number (stored in the 'Visit' column)  \n* Medications administered (stored in the 'Medication' column)  \n* Time of FOG measurement (stored in the 'Time' column)  \n* Start time of FOG event (recorded in the 'Init' column)  \n* End time of FOG event (recorded in the 'Completion' column)  \n* Vertical axis acceleration (captured in the 'AccV' column)  \n* Mediolateral axis acceleration (captured in the 'AccML' column)  \n* Anteroposterior axis acceleration (captured in the 'AccAP' column)  \n* Initial uncertainty at the start of the event (stored in the 'StartHesitation' column)  \n* Uncertainty during turning (stored in the 'Turn' column)  \n* Movement delays (captured in the 'Walking' column)  \n\nOur goal is to develop a model capable of predicting episodes of freezing of gait and their corresponding types using this dataset. To evaluate the model's performance, we will utilize the average sum of AP (mean Average Precision) across all three event classes.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport sklearn\nimport matplotlib.pyplot as plt\nimport os\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MinMaxScaler, StandardScaler, LabelEncoder\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom sklearn.preprocessing import OneHotEncoder, LabelEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.linear_model import LogisticRegression\nimport xgboost as xgb\nimport joblib\n\nfrom sklearn.multioutput import MultiOutputClassifier\nfrom sklearn.pipeline import Pipeline,make_pipeline\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.metrics import average_precision_score,accuracy_score, confusion_matrix, roc_auc_score, f1_score","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2023-05-22T09:40:32.074035Z","iopub.execute_input":"2023-05-22T09:40:32.074397Z","iopub.status.idle":"2023-05-22T09:40:33.709041Z","shell.execute_reply.started":"2023-05-22T09:40:32.074368Z","shell.execute_reply":"2023-05-22T09:40:33.707904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Settings","metadata":{}},{"cell_type":"code","source":"# # Set formatting option\n\npd.options.display.float_format = '{:.2f}'.format","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:40:34.526445Z","iopub.execute_input":"2023-05-22T09:40:34.527165Z","iopub.status.idle":"2023-05-22T09:40:34.534726Z","shell.execute_reply.started":"2023-05-22T09:40:34.527127Z","shell.execute_reply":"2023-05-22T09:40:34.533677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fill_missing_values(df):\n    \"\"\"\n    Replaces missing values with the median value for each numerical column.\n\n    :param df: pandas DataFrame, the dataset in which missing values need to be replaced\n    :return: pandas DataFrame, the dataset with replaced missing values\n    \"\"\"\n    numeric_cols = df.select_dtypes(include=['float64', 'int64']).columns  # Selecting all numerical columns\n    for col in numeric_cols:\n        median = df[col].median()  # Finding the Median Value of a Column\n        df[col].fillna(median, inplace=True)  # Replacing Missing Values with the Median Value\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:40:36.723825Z","iopub.execute_input":"2023-05-22T09:40:36.724190Z","iopub.status.idle":"2023-05-22T09:40:36.731975Z","shell.execute_reply.started":"2023-05-22T09:40:36.724159Z","shell.execute_reply":"2023-05-22T09:40:36.729009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to Get Data Information\n\ndef explore_dataframe(df):\n    print(\"Shape of dataframe:\", df.shape)\n    display(df.head())\n    print(\"Info of dataframe:\\n\")\n    df.info()\n    print(\"Summary statistics of dataframe:\\n\", df.describe())\n    print(\"Missing values in dataframe:\\n\", df.isnull().sum())\n    print(\"Duplicate rows in dataframe:\", df.duplicated().sum())","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:40:38.440569Z","iopub.execute_input":"2023-05-22T09:40:38.441285Z","iopub.status.idle":"2023-05-22T09:40:38.447176Z","shell.execute_reply.started":"2023-05-22T09:40:38.441247Z","shell.execute_reply":"2023-05-22T09:40:38.446018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking for Missing Values in Each Column\n\ndef check_missing_values(df):\n    \"\"\"\n    Checks the count of missing values in each column of a DataFrame.\n\n    :param df: pandas.DataFrame, the DataFrame to check for missing values.\n    :return: pandas.DataFrame, the DataFrame with information about missing values.\n    \"\"\"\n    return df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:40:40.861844Z","iopub.execute_input":"2023-05-22T09:40:40.862526Z","iopub.status.idle":"2023-05-22T09:40:40.868541Z","shell.execute_reply.started":"2023-05-22T09:40:40.862492Z","shell.execute_reply":"2023-05-22T09:40:40.867603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Reading Data","metadata":{}},{"cell_type":"code","source":"# Reading Data from a CSV File\n\ndata = pd.read_csv('/kaggle/input/eda-parkinson/EDA_Parkinson.csv', low_memory=False)","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:40:43.434566Z","iopub.execute_input":"2023-05-22T09:40:43.435253Z","iopub.status.idle":"2023-05-22T09:41:30.960772Z","shell.execute_reply.started":"2023-05-22T09:40:43.435211Z","shell.execute_reply":"2023-05-22T09:41:30.959849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:41:38.059321Z","iopub.execute_input":"2023-05-22T09:41:38.059894Z","iopub.status.idle":"2023-05-22T09:41:38.088332Z","shell.execute_reply.started":"2023-05-22T09:41:38.059858Z","shell.execute_reply":"2023-05-22T09:41:38.087346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Removing unnecessary columns\ndata.drop(['Id', 'Subject', 'Visit', 'Medication'], axis=1, inplace=True)\n\n# Splitting into features (X) and target variable (y)\nX = data.drop(['StartHesitation', 'Turn', 'Walking'], axis=1)\ny = data[['Turn', 'Walking']]","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:41:44.596948Z","iopub.execute_input":"2023-05-22T09:41:44.597312Z","iopub.status.idle":"2023-05-22T09:41:45.431522Z","shell.execute_reply.started":"2023-05-22T09:41:44.597284Z","shell.execute_reply":"2023-05-22T09:41:45.430465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Splitting into training and test sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Creating a pipeline for each column\npipelines = {}\nhyperparameters = {\n    'max_depth': 3,\n    'n_estimators': 100,\n    'learning_rate': 0.1\n}\n\n# Processing each column separately\nfor column in y.columns:\n    pipeline = make_pipeline(StandardScaler(), xgb.XGBClassifier(**hyperparameters))\n    pipeline.fit(X_train, y_train[column])\n    pipelines[column] = pipeline\n\n# Model evaluation\nfor column, pipeline in pipelines.items():\n    train_score = pipeline.score(X_train, y_train[column])\n    test_score = pipeline.score(X_test, y_test[column])\n    print(f\"{column}: Train Score - {train_score}, Test Score - {test_score}\")\n\n# Predictions on the test set\npredictions = pd.DataFrame()\nfor column, pipeline in pipelines.items():\n    predictions[column] = pipeline.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-05-22T09:41:51.423260Z","iopub.execute_input":"2023-05-22T09:41:51.423836Z","iopub.status.idle":"2023-05-22T10:06:40.659239Z","shell.execute_reply.started":"2023-05-22T09:41:51.423804Z","shell.execute_reply":"2023-05-22T10:06:40.658268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Metrics evaluation for each column\n\nfor column in predictions.columns:  \n    y_train_pred = pipeline.predict(X_train)  \n    y_test_pred = pipeline.predict(X_test) \n\n    train_accuracy = accuracy_score(y_train[column], y_train_pred)\n    test_accuracy = accuracy_score(y_test[column], y_test_pred)\n\n    average_precision = average_precision_score(y_test[column], y_test_pred)\n\n    confusion = confusion_matrix(y_test[column], y_test_pred)\n\n    try:\n        roc_auc = roc_auc_score(y_test[column], y_test_pred)\n    except ValueError:\n        roc_auc = \"Not defined\"\n\n    f1 = f1_score(y_test[column], y_test_pred)\n\n    print(f\"Metrics for column '{column}':\")\n    print(f\"Train Accuracy: {train_accuracy}\")\n    print(f\"Test Accuracy: {test_accuracy}\")\n    print(f\"Average Precision: {average_precision}\")\n    print(\"Confusion Matrix:\")\n    print(confusion)\n    print(f\"ROC AUC Score: {roc_auc}\")\n    print(f\"F1 Score: {f1}\")\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:20:10.358743Z","iopub.execute_input":"2023-05-22T10:20:10.360972Z","iopub.status.idle":"2023-05-22T10:20:52.281998Z","shell.execute_reply.started":"2023-05-22T10:20:10.360928Z","shell.execute_reply":"2023-05-22T10:20:52.280912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading the sample_submission.csv file\ntest_data = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/sample_submission.csv')\n\n# Creating the submission dataframe\nsubmission_df = pd.DataFrame({'Id': test_data['Id'], 'StartHesitation': test_data['StartHesitation'].astype(int), 'Turn': predictions['Turn'], 'Walking': predictions['Walking']})\n\n# Trimming excess rows if necessary\nsubmission_df = submission_df[:len(test_data)]\n\n# Formatting StartHesitation column\nsubmission_df['StartHesitation'] = submission_df['StartHesitation'].astype(int)\n\n# Saving the submission file as submission.csv\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:31:37.500339Z","iopub.execute_input":"2023-05-22T10:31:37.500706Z","iopub.status.idle":"2023-05-22T10:31:38.607623Z","shell.execute_reply.started":"2023-05-22T10:31:37.500677Z","shell.execute_reply":"2023-05-22T10:31:38.606701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explore_dataframe(submission_df)","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:31:42.383644Z","iopub.execute_input":"2023-05-22T10:31:42.384347Z","iopub.status.idle":"2023-05-22T10:31:42.658549Z","shell.execute_reply.started":"2023-05-22T10:31:42.384314Z","shell.execute_reply":"2023-05-22T10:31:42.657590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import joblib\n\n# Saving the pipelines\n\njoblib.dump(pipelines, 'pipeline.pkl')\n\n# Loading the pipelines\n\npredictions = joblib.load('pipeline.pkl')","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:31:48.142715Z","iopub.execute_input":"2023-05-22T10:31:48.143558Z","iopub.status.idle":"2023-05-22T10:31:48.166073Z","shell.execute_reply.started":"2023-05-22T10:31:48.143522Z","shell.execute_reply":"2023-05-22T10:31:48.165361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def check_submission(submission_df):\n    # Check values in columns\n    valid_values = {0, 1}\n    for column in [\"StartHesitation\", \"Turn\", \"Walking\"]:\n        values = submission_df[column].unique()\n        if set(values) != valid_values:\n            return f\"Invalid values found in column '{column}': {values}\"\n\n    # Check data types\n    expected_data_types = {int, bool}\n    for column in [\"StartHesitation\", \"Turn\", \"Walking\"]:\n        data_type = submission_df[column].dtype\n        if data_type not in expected_data_types:\n            return f\"Invalid data type found in column '{column}': {data_type}\"\n\n    # Check for missing or invalid values\n    if submission_df.isnull().values.any():\n        return \"Missing values found in the submission\"\n    \n    return \"Submission check passed successfully\"","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:32:50.893114Z","iopub.execute_input":"2023-05-22T10:32:50.893493Z","iopub.status.idle":"2023-05-22T10:32:50.899842Z","shell.execute_reply.started":"2023-05-22T10:32:50.893456Z","shell.execute_reply":"2023-05-22T10:32:50.898956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result = check_submission(submission_df)\nprint(result)","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:32:54.514821Z","iopub.execute_input":"2023-05-22T10:32:54.515180Z","iopub.status.idle":"2023-05-22T10:32:54.522544Z","shell.execute_reply.started":"2023-05-22T10:32:54.515151Z","shell.execute_reply":"2023-05-22T10:32:54.521627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def check_submission(test_data):\n    # Check values in columns\n    valid_values = {0, 1}\n    for column in [\"StartHesitation\", \"Turn\", \"Walking\"]:\n        values = test_data[column].unique()\n        if set(values) != valid_values:\n            return f\"Invalid values found in column '{column}': {values}\"\n\n    # Check data types\n    expected_data_types = {int, bool}\n    for column in [\"StartHesitation\", \"Turn\", \"Walking\"]:\n        data_type = test_data[column].dtype\n        if data_type not in expected_data_types:\n            return f\"Invalid data type found in column '{column}': {data_type}\"\n\n    # Check for missing or invalid values\n    if test_data.isnull().values.any():\n        return \"Missing values found in the submission\"\n    \n    return \"Submission check passed successfully\"","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:33:16.652747Z","iopub.execute_input":"2023-05-22T10:33:16.653444Z","iopub.status.idle":"2023-05-22T10:33:16.660867Z","shell.execute_reply.started":"2023-05-22T10:33:16.653409Z","shell.execute_reply":"2023-05-22T10:33:16.659878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result = check_submission(test_data)\nprint(result)","metadata":{"execution":{"iopub.status.busy":"2023-05-22T10:33:19.954561Z","iopub.execute_input":"2023-05-22T10:33:19.955799Z","iopub.status.idle":"2023-05-22T10:33:19.965104Z","shell.execute_reply.started":"2023-05-22T10:33:19.955741Z","shell.execute_reply":"2023-05-22T10:33:19.963912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion:\n\nOur model for predicting episodes of mobility impairment and their types is based on utilizing various features from the dataset. We incorporated patient information, including their identifiers, numbers, and visits, as well as data on administered medications. Additionally, we utilized information on the timing of FOG measurements, the start and completion times of FOG events, and acceleration data from the vertical, mediolateral, and anteroposterior axes.  \n\nTo train the model, we employed a pipeline consisting of several stages.  \n\n* Firstly, we performed the separation of data into **features (X)** and the **target variable (y)**, where X contains all columns except those related to the target variable, and y contains the `'StartHesitation'`, `'Turn'`, and `'Walking'` columns.  \n\n* Next, we applied the train-test split method, using the train_test_split function with a **test_size of 0.2**.\n\nFor each target variable column, we created a separate pipeline comprising data scaling using StandardScaler and training an XGBClassifier with specific hyperparameters. We chose **XGBClassifier** due to its powerful gradient boosting capabilities, enabling it to handle complex dependencies in the data.\n\nTraining the model for each target variable class was necessary as each class represents a distinct type of mobility impairment episode. Our model can account for the differences and characteristics of each class, resulting in more accurate predictions and classification of FOG episodes.\n\nIn conclusion, we obtained a model capable of predicting episodes of mobility impairment and their types based on patient data and acceleration along different axes. Incorporating various features and training for each target variable class enhances the prediction quality and enables personalized approaches to the treatment and monitoring of Parkinson's disease patients.  \n","metadata":{}}]}