{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Motor and Non-Motor Manifestations in Patients with Parkinson's Disease\n\n**Parkinson's disease** is characterized by various `motor` and `non-motor` manifestations that can significantly impact the quality of life of patients. Let's examine what these manifestations entail and how they are studied.\n\nMotor Manifestations in Patients with Parkinson's Disease:\n\n**Motor impairments:**  \n\n* Patients with Parkinson's disease often experience rigidity (difficulty with mobility), tremors, and bradykinesia (slowed movements).  \n* Akinesia and blockages: Akinesia refers to the absence or delay in initiating movements, while blockages are temporary pauses or \"freezing\" during movement. These manifestations can lead to difficulties in performing simple daily tasks, such as getting up from a chair or initiating walking.  \n* Freezing of Gait (FoG): This refers to brief episodes of gait blockage while walking, where patients feel unable to continue or get \"stuck\" in place.  \n\n**Non-Motor Manifestations in Patients with Parkinson's Disease:**  \n\n* Sleep disturbances: Patients may experience insomnia, restless legs during sleep, nightmares, and daytime sleepiness.  \n* Depression and anxiety: Parkinson's disease is associated with an increased risk of developing depression and anxiety. Patients may experience feelings of sadness, loss of interest in previously enjoyable activities, as well as anxiety and restlessness.  \n* Cognitive impairments: Some patients may experience problems with memory, attention, concentration, and other cognitive functions.  \n\nThe study of motor and non-motor manifestations in patients with Parkinson's disease is conducted using various methods and tools:\n\n* **Clinical assessment:**  Neurologists observe and evaluate patients using standardized scales and questionnaires, such as the Unified Parkinson's Disease Rating Scale `(UPDRS)`, which assess the severity of symptoms and their impact on the patient's functioning.  \n* **Accelerometry:**  The use of accelerometers to measure movements allows recording and analysis of data on patients' motor manifestations in everyday life, including walking, tremors, and activity levels.  \n* **Electroencephalography (EEG):** This method allows the study of the electrical activity of the brain in patients and the identification of changes associated with motor and non-motor manifestations.  \n* **Neuroimaging:**  Research using neurovisualization techniques such as functional magnetic resonance imaging (fMRI) or positron emission tomography (PET) can help identify changes in brain activity and associated pathological processes.  \n\nData for the study of motor and non-motor manifestations of Parkinson's disease can include accelerometric data, clinical assessments, sleep information, electroencephalographic data, and neuroimaging results. These data enable scientists and physicians to gain a more comprehensive understanding of the nature and impact of Parkinson's disease symptoms on patients.  \n\n**Based on the provided data, we can utilize the following datasets for the investigation of motor and non-motor manifestations in patients with Parkinson's disease:**  \n\nThe file `\"tdcsfog_metadata.csv\"` contains information about each data series in the \"tdcsfog/\" dataset, including unique identifiers, visit data, test types, and information about administered medications.  \nThe file `\"defog_metadata.csv\"` contains information about each data series in the \"defog/\" dataset, including unique identifiers, visit data, and information about administered medications.  \nThe file `\"subjects.csv\"` contains metadata about each patient, including age, gender, visit data, and other information that may be useful for analyzing Parkinson's disease manifestations.  \nThe files `\"events.csv\"` and  `\"tasks.csv\"` contain metadata about FoG events and tasks, which can be used for analysis and identification of motor and non-motor manifestations.  \n\nWith these data, we can conduct analyses of motor and non-motor manifestations of Parkinson's disease, including studying accelerometric data related to movements and events, as well as analyzing metadata about patients, visits, and medication treatments. These data allow us to gain a better understanding and explore various aspects of Parkinson's disease and its impact on patients.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nfrom IPython.display import display\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport seaborn as sns\n\nfrom sklearn.metrics import mutual_info_score\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import OneHotEncoder, StandardScaler\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import average_precision_score, f1_score, accuracy_score, confusion_matrix, classification_report, roc_auc_score\nfrom sklearn.multioutput import MultiOutputClassifier\nfrom sklearn.impute import SimpleImputer","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:30:43.061535Z","iopub.execute_input":"2023-05-23T12:30:43.061923Z","iopub.status.idle":"2023-05-23T12:30:43.069371Z","shell.execute_reply.started":"2023-05-23T12:30:43.061892Z","shell.execute_reply":"2023-05-23T12:30:43.068232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set formatting option\n\npd.options.display.float_format = '{:.2f}'.format","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:30:46.140866Z","iopub.execute_input":"2023-05-23T12:30:46.141218Z","iopub.status.idle":"2023-05-23T12:30:46.145622Z","shell.execute_reply.started":"2023-05-23T12:30:46.141189Z","shell.execute_reply":"2023-05-23T12:30:46.144737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fill_missing_values(df):\n    \"\"\"\n    Replaces missing values with the median value for each numerical column.\n\n    :param df: pandas DataFrame, the dataset in which missing values need to be replaced\n    :return: pandas DataFrame, the dataset with replaced missing values\n    \"\"\"\n    numeric_cols = df.select_dtypes(include=['float64', 'int64']).columns  # Selecting all numerical columns\n    for col in numeric_cols:\n        median = df[col].median()  # Finding the Median Value of a Column\n        df[col].fillna(median, inplace=True)   # Replacing Missing Values with the Median Value\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:30:47.887759Z","iopub.execute_input":"2023-05-23T12:30:47.888399Z","iopub.status.idle":"2023-05-23T12:30:47.895404Z","shell.execute_reply.started":"2023-05-23T12:30:47.888363Z","shell.execute_reply":"2023-05-23T12:30:47.894387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to Get Data Information\n\ndef explore_dataframe(df):\n    print(\"Shape of dataframe:\", df.shape)\n    \n    display(df.head())\n    \n    print(\"Info of dataframe:\\n\")\n    df.info()\n         \n    print(\"Missing values in dataframe:\\n\", df.isnull().sum())\n    \n    print(\"Duplicate rows in dataframe:\", df.duplicated().sum())","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:30:50.165625Z","iopub.execute_input":"2023-05-23T12:30:50.166106Z","iopub.status.idle":"2023-05-23T12:30:50.172294Z","shell.execute_reply.started":"2023-05-23T12:30:50.166065Z","shell.execute_reply":"2023-05-23T12:30:50.171221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking for Missing Values in Each Column\n\ndef check_missing_values(df):\n    \"\"\"\n    Checks the count of missing values in each column of a DataFrame.\n\n    :param df: pandas.DataFrame, the DataFrame to check for missing values.\n    :return: pandas.DataFrame, the DataFrame with information about missing values.\n    \"\"\"\n    return df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:30:52.175939Z","iopub.execute_input":"2023-05-23T12:30:52.176298Z","iopub.status.idle":"2023-05-23T12:30:52.181934Z","shell.execute_reply.started":"2023-05-23T12:30:52.176269Z","shell.execute_reply":"2023-05-23T12:30:52.180910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Reading and Preparation of a New Table","metadata":{}},{"cell_type":"code","source":"def check_submission(df):\n    # Check values in columns\n    valid_values = {0, 1}\n    for column in [\"\"]:\n        values = df[column].unique()\n        if set(values) != valid_values:\n            return f\"Invalid values found in column '{column}': {values}\"\n\n    # Check data types\n    expected_data_types = {int, bool}\n    for column in [\"\"]:\n        data_type = df[column].dtype\n        if data_type not in expected_data_types:\n            return f\"Invalid data type found in column '{column}': {data_type}\"\n\n    # Check for missing or invalid values\n    if submission_df.isnull().values.any():\n        return \"Missing values found in the submission\"\n    \n    return \"Submission check passed successfully\"","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:30:55.452505Z","iopub.execute_input":"2023-05-23T12:30:55.452880Z","iopub.status.idle":"2023-05-23T12:30:55.459283Z","shell.execute_reply.started":"2023-05-23T12:30:55.452849Z","shell.execute_reply":"2023-05-23T12:30:55.458322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading Data from tdcsfog_metadata.csv\n\ntdcsfog = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/tdcsfog_metadata.csv')\n\n# Reading Data from defog_metadata.csv\n\ndefog = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/defog_metadata.csv')\n\n# Reading Data from events.csv\n\nevents = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/events.csv')\n\n# Reading Data from train & tdcsfog & daily_metadata.csv\n\ntask = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/tasks.csv')\n\n# Reading Data from subjects.csv\n\nsubject = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/subjects.csv')","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:30:57.895118Z","iopub.execute_input":"2023-05-23T12:30:57.895465Z","iopub.status.idle":"2023-05-23T12:30:57.925891Z","shell.execute_reply.started":"2023-05-23T12:30:57.895436Z","shell.execute_reply":"2023-05-23T12:30:57.925054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explore_dataframe(tdcsfog)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:00.484786Z","iopub.execute_input":"2023-05-23T12:31:00.485429Z","iopub.status.idle":"2023-05-23T12:31:00.507220Z","shell.execute_reply.started":"2023-05-23T12:31:00.485392Z","shell.execute_reply":"2023-05-23T12:31:00.506216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explore_dataframe(defog)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:03.253969Z","iopub.execute_input":"2023-05-23T12:31:03.254328Z","iopub.status.idle":"2023-05-23T12:31:03.275353Z","shell.execute_reply.started":"2023-05-23T12:31:03.254299Z","shell.execute_reply":"2023-05-23T12:31:03.274367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explore_dataframe(events)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:06.072417Z","iopub.execute_input":"2023-05-23T12:31:06.072785Z","iopub.status.idle":"2023-05-23T12:31:06.097275Z","shell.execute_reply.started":"2023-05-23T12:31:06.072751Z","shell.execute_reply":"2023-05-23T12:31:06.096276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explore_dataframe(task)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:08.757526Z","iopub.execute_input":"2023-05-23T12:31:08.758639Z","iopub.status.idle":"2023-05-23T12:31:08.784206Z","shell.execute_reply.started":"2023-05-23T12:31:08.758600Z","shell.execute_reply":"2023-05-23T12:31:08.783157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explore_dataframe(subject)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:13.068852Z","iopub.execute_input":"2023-05-23T12:31:13.069522Z","iopub.status.idle":"2023-05-23T12:31:13.092308Z","shell.execute_reply.started":"2023-05-23T12:31:13.069488Z","shell.execute_reply":"2023-05-23T12:31:13.091394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# List of all datasets\n\ndatasets = [tdcsfog, defog, events, task, subject]\n\n# Dataset merging\n\nmerged_data = pd.concat(datasets, axis=1, join='outer')\n\n# Removing duplicate columns\n\nmerged_data = merged_data.loc[:, ~merged_data.columns.duplicated()]\n\n# Merging rows with the same column names\n\nmerged_data['Subject'] = merged_data['Subject'].apply(lambda x: ', '.join(x) if isinstance(x, list) else x)\n\n# Removing 'Id' and 'Subject' columns\n\nmerged_data = merged_data.drop(['Id', 'Subject'], axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:15.594438Z","iopub.execute_input":"2023-05-23T12:31:15.595134Z","iopub.status.idle":"2023-05-23T12:31:15.613488Z","shell.execute_reply.started":"2023-05-23T12:31:15.595097Z","shell.execute_reply":"2023-05-23T12:31:15.612620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Replacing values in the 'Sex' column\n\nmerged_data['Sex'] = merged_data['Sex'].replace({'M': 1, 'F': 0})\n\n# Replacing values in the 'Medication' column\n\nmerged_data['Medication'] = merged_data['Medication'].replace({'on': 1, 'off': 0})","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:19.858801Z","iopub.execute_input":"2023-05-23T12:31:19.859893Z","iopub.status.idle":"2023-05-23T12:31:19.870417Z","shell.execute_reply.started":"2023-05-23T12:31:19.859849Z","shell.execute_reply":"2023-05-23T12:31:19.869469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"merged_data['Visit'] = merged_data['Visit'].fillna(0).astype(int)\nmerged_data['Test'] = merged_data['Test'].fillna(0).astype(int)\nmerged_data['Kinetic'] = merged_data['Kinetic'].fillna(0).astype(int)\nmerged_data['Age'] = merged_data['Age'].fillna(0).astype(int)\nmerged_data['YearsSinceDx'] = merged_data['YearsSinceDx'].fillna(0).astype(int)\nmerged_data['UPDRSIII_On'] = merged_data['UPDRSIII_On'].fillna(0).astype(int)\nmerged_data['UPDRSIII_Off'] = merged_data['UPDRSIII_Off'].fillna(0).astype(int)\nmerged_data['NFOGQ'] = merged_data['NFOGQ'].fillna(0).astype(int)\nmerged_data['Sex'] = merged_data['Sex'].fillna(0).astype(int)\nmerged_data['Medication'] = merged_data['Medication'].fillna(0).astype(int)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:22.450360Z","iopub.execute_input":"2023-05-23T12:31:22.450734Z","iopub.status.idle":"2023-05-23T12:31:22.466224Z","shell.execute_reply.started":"2023-05-23T12:31:22.450688Z","shell.execute_reply":"2023-05-23T12:31:22.465179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Mode imputation\n\ncolumns_mode = ['Sex', 'Medication', 'Age', 'YearsSinceDx', 'UPDRSIII_On', 'UPDRSIII_Off', 'NFOGQ']\nmerged_data[columns_mode] = merged_data[columns_mode].fillna(merged_data[columns_mode].mode().iloc[0])\n\n# Median imputation\n\ncolumns_median = ['Begin', 'End', 'Init', 'Completion', 'Kinetic', 'Visit', 'Test']\nmerged_data[columns_median] = merged_data[columns_median].fillna(merged_data[columns_median].median())\n\n# Dropping rows with missing values in columns 'Type' and 'Task'\n\nmerged_data = merged_data.dropna(subset=['Type', 'Task'])","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:24.790132Z","iopub.execute_input":"2023-05-23T12:31:24.790476Z","iopub.status.idle":"2023-05-23T12:31:24.816774Z","shell.execute_reply.started":"2023-05-23T12:31:24.790448Z","shell.execute_reply":"2023-05-23T12:31:24.815974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Outputting the final dataset\n\nexplore_dataframe(merged_data.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:27.689439Z","iopub.execute_input":"2023-05-23T12:31:27.689950Z","iopub.status.idle":"2023-05-23T12:31:27.718246Z","shell.execute_reply.started":"2023-05-23T12:31:27.689916Z","shell.execute_reply":"2023-05-23T12:31:27.717371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The process of merging datasets involved combining data from the following files: `\"tdcsfog_metadata.csv\"`, `\"defog_metadata.csv\"`, `\"events\"`, `\"task\"`, `\"subject\"`.\n\nThese datasets contain information about time series recorded at a sampling rate of 128 and 100 time steps per second, respectively. These data allow us to study motor manifestations in patients, such as accelerometer measurements (vertical, medio-lateral, and antero-posterior accelerations) and event indicators (onsets, turns, walking).\n\nWe also utilized the \"subjects.csv\" file, which contains metadata about each patient, including age, sex, visit information, and other useful details.\n\nThe aim of the analysis was to investigate motor and non-motor symptoms of Parkinson's disease using the available data. By merging the datasets and employing various analysis methods, we can gain a better understanding of the nature and impact of these symptoms on patients.\n\nThe application of mode and median in data analysis is driven by their practical utility. Mode allows us to identify the most frequently occurring values in a sample, which can be important for identifying characteristic features of Parkinson's disease manifestations. On the other hand, the median represents a central value that is not strongly influenced by outliers and extreme values. Thus, using the median allows us to capture typical values in the sample without being distorted by anomalies.\n\nThe overall data analysis and merging process provide us with the opportunity to gain deeper insights into the motor and non-motor manifestations of Parkinson's disease and their impact on patients. This knowledge can contribute to the development of more effective diagnostic and treatment methods for this condition.","metadata":{}},{"cell_type":"markdown","source":"## Features: UPDRSIII_On, UPDRSIII_Off, and NFOGQ\n\nThese features are standardized scales and questionnaires used to assess the symptoms of Parkinson's disease and their impact on patients. Let's take a closer look at each of them:\n\n**UPDRSIII_On:** This is the Unified Parkinson's Disease Rating Scale, which evaluates motor symptoms of Parkinson's disease in the \"on\" state. The scale includes various aspects of motor impairments, such as rigidity (difficulty in muscle movement), tremors, and bradykinesia (slowed movements). Higher values on the scale indicate more pronounced symptoms and a more severe condition of the patient.\n\n**UPDRSIII_Off:** This is the same Unified Parkinson's Disease Rating Scale, but it assesses motor symptoms in the \"off\" state, meaning when there is no effect of medication or outside the medication's duration of action. The assessment is based on the same parameters as the UPDRSIII_On scale and allows for comparison of symptom changes before and after medication intervention.\n\n**NFOGQ:** This is the Freezing of Gait Questionnaire, designed to evaluate the symptom of freezing of gait (FoG). Freezing of gait is a characteristic symptom of Parkinson's disease, where patients experience temporary blockages or \"freezing\" during walking, having difficulty in continuing the movement. The NFOGQ questionnaire consists of a series of questions aimed at identifying and assessing the frequency, duration, and impact of freezing of gait on patients.\n\nThe use of these scales and questionnaires enables physicians and researchers to systematically evaluate and measure the symptoms of Parkinson's disease, their severity, and changes in response to treatment.","metadata":{}},{"cell_type":"code","source":"merged_data[['UPDRSIII_On', 'UPDRSIII_Off', 'NFOGQ']].describe()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:34.093990Z","iopub.execute_input":"2023-05-23T12:31:34.094698Z","iopub.status.idle":"2023-05-23T12:31:34.123909Z","shell.execute_reply.started":"2023-05-23T12:31:34.094653Z","shell.execute_reply":"2023-05-23T12:31:34.122683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**UPDRSIII_On** and **UPDRSIII_Off**: The average value for both columns is around `2.9`, indicating the presence of some motor symptoms in patients. However, the standard deviation for both columns (`10.33` for **UPDRSIII_On** and `11.18` for **UPDRSIII_Off**) indicates a significant variation in values.\n\nThe minimum values are `0`, and the maximum values are `79` for **UPDRSIII_On** and `91` for **UPDRSIII_Off**. This suggests that some patients may have no or very mild symptoms, while others may have more pronounced symptoms.\n\n**NFOGQ**: The average value for the **NFOGQ** column is `1.49`, and the standard deviation is `5.31`. This indicates the presence of some freezing of gait symptoms in patients. However, most of the values `(25%, 50%, 75%)` are `0`, which may indicate that freezing of gait is absent or very weakly expressed in the majority of patients.\n\nFrom this data, we can conclude that the overall average value suggests the presence of some Parkinson's disease symptoms in patients, but there is significant variation among patients.\n\n**Further analysis and visualization of the data will provide a more detailed understanding of the distribution and relationships between these features.**","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Create a figure and axes for the plot\n\nfig, ax = plt.subplots()\n\n# Plot the histogram for the 'UPDRSIII_On' column and set the color\n\nax.hist(merged_data['UPDRSIII_On'], bins=6, color='lightblue', alpha=0.8, label='UPDRSIII_On')\n\n# Plot the histogram for the 'UPDRSIII_Off' column and set the color\n\nax.hist(merged_data['UPDRSIII_Off'], bins=6, color='yellow', alpha=0.5, label='UPDRSIII_Off')\n\n# Plot the histogram for the 'NFOGQ' column and set the color\n\nax.hist(merged_data['NFOGQ'], bins=6, color='lightpink', alpha=0.8, label='NFOGQ')\n\n# Set the labels for the x and y axes, and the title\nax.set_xlabel('Value')\nax.set_ylabel('Frequency')\nax.set_title('Distribution of UPDRSIII_On, UPDRSIII_Off, and NFOGQ')\n\n# Add a legend\nax.legend()\n\n# Display the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:37.908825Z","iopub.execute_input":"2023-05-23T12:31:37.909794Z","iopub.status.idle":"2023-05-23T12:31:38.190515Z","shell.execute_reply.started":"2023-05-23T12:31:37.909746Z","shell.execute_reply":"2023-05-23T12:31:38.189623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distribution of values in the columns **'UPDRSIII_On'**, **'UPDRSIII_Off'**, and **'NFOGQ'** is characterized by high frequencies (up to 1750) for values not exceeding 20. This indicates that the majority of observations in the data fall into the low value range for these columns.\n\nIn the **'UPDRSIII_On'** and **'UPDRSIII_Off'** columns, the highest frequencies are observed in the range from 0 to 20, indicating relatively low expression of Parkinson's disease symptoms during the medication test (**'UPDRSIII_On'**) and after its cessation (**'UPDRSIII_Off'**). This may indicate a positive effect of medication therapy.\n\nIn the **'NFOGQ'** column, there is also a high frequency of values up to 20, indicating a low expression of problems related to movement slowing and freezing of gait (non-freezing of gait) in patients with Parkinson's disease. This may suggest relatively good control of symptoms in this aspect of the disease.\n\nThus, the data shows that the majority of patients have a low expression of Parkinson's disease symptoms in the aspects considered.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Creating a figure with three subplots in a single line\n\nfig, axes = plt.subplots(1, 3, figsize=(15, 5))\n\n# Plot for the 'UPDRSIII_On' column\n\naxes[0].boxplot(merged_data['UPDRSIII_On'])\naxes[0].set_xlabel('UPDRSIII_On')\naxes[0].set_title('Box-plot UPDRSIII_On')\n\n# Plot for the 'UPDRSIII_Off' column\n\naxes[1].boxplot(merged_data['UPDRSIII_Off'])\naxes[1].set_xlabel('UPDRSIII_Off')\naxes[1].set_title('Box-plot UPDRSIII_Off')\n\n# Plot for the 'NFOGQ' column\n\naxes[2].boxplot(merged_data['NFOGQ'])\naxes[2].set_xlabel('NFOGQ')\naxes[2].set_title('Box-plot NFOGQ')\n\n# Display the plots\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:41.584625Z","iopub.execute_input":"2023-05-23T12:31:41.586249Z","iopub.status.idle":"2023-05-23T12:31:42.163084Z","shell.execute_reply.started":"2023-05-23T12:31:41.586206Z","shell.execute_reply":"2023-05-23T12:31:42.162188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**UPDRSIII_On:**  \n\nThe majority of values in the UPDRSIII_On column are located in the range of 10 to 50, indicating the presence of pronounced Parkinson's disease symptoms in patients. The box representing the interquartile range appears as a strip, indicating a small dispersion of values within this range. There are also some outliers beyond this range, which may suggest the presence of patients with a higher degree of symptoms.  \n\n**UPDRSIII_Off:**  \n\nA significant clustering of values in the UPDRSIII_Off column is observed in the range of 20 to 55, indicating an improvement in Parkinson's disease symptoms after treatment. The box also appears as a strip, indicating a small dispersion of values within this range. However, there are outliers both above and below this range, which may indicate the presence of patients with higher or lower degrees of symptoms despite treatment.  \n\n**NFOGQ:**  \n\nThe values (circles) in the NFOGQ column are few and gradually increase vertically from 10 to 30. This may indicate the presence of some gait freezing symptoms in patients, but in a milder form compared to other symptoms. The box also appears as a strip, indicating a limited dispersion of values within this range.  \n\n**Overall Conclusion:**  \n\nThe box-plot graphs allow for an assessment of the distribution of values and the range of Parkinson's disease symptoms in the UPDRSIII_On, UPDRSIII_Off, and NFOGQ columns. These characteristics indicate the presence of symptoms and changes in symptoms after treatment. However, it is important to consider the outliers, which may indicate the presence of patients with higher or lower degrees of symptoms, different from the main group.  ","metadata":{}},{"cell_type":"code","source":"task_counts = merged_data['Task'].value_counts()\n\ncolors = ['lightblue', 'yellow', 'lightpink']\n\nplt.bar(task_counts.index, task_counts.values, color=colors)\nplt.xlabel('Task')\nplt.ylabel('Frequency')\nplt.title('Bar Chart of Task')\nplt.xticks(rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:46.564426Z","iopub.execute_input":"2023-05-23T12:31:46.565666Z","iopub.status.idle":"2023-05-23T12:31:46.942973Z","shell.execute_reply.started":"2023-05-23T12:31:46.565621Z","shell.execute_reply":"2023-05-23T12:31:46.942074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_values_task = merged_data['Task'].unique().tolist()\nprint(\"Unique values in the 'Task' column:\")\nprint(unique_values_task)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:50.361548Z","iopub.execute_input":"2023-05-23T12:31:50.362217Z","iopub.status.idle":"2023-05-23T12:31:50.369372Z","shell.execute_reply.started":"2023-05-23T12:31:50.362175Z","shell.execute_reply":"2023-05-23T12:31:50.368223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the values in the 'Task' column, several groups or categories can be identified:\n\n| Group/Task | Description |\n| --- | --- |\n| Rest | Rest 1, Rest 2 |\n| Walking (4MW) | 4-Minute Walk, 4-Minute Walk with counting |\n| Marching Band (MB) | Marching Band 1, Marching Band 2a, Marching Band 2b, Marching Band 3 (right side), Marching Band 3 (left side), Marching Band 4, Marching Band 5, Marching Band 6 (right side), Marching Band 6 (left side), Marching Band 7, Marching Band 8, Marching Band 9, Marching Band 10, Marching Band 11, Marching Band 12, Marching Band 13, Marching Band 6 |\n| Timed Up and Go (TUG) Test | TUG (static), TUG (dynamic), TUG with counting |\n| Turning | Turning (static), Turning (dynamic), Turning with counting |\n| Hotspot | Hotspot 1, Hotspot 1 with counting, Hotspot 2, Hotspot 2 with counting |\n\nEach group corresponds to a specific type of task or test used in the study.","metadata":{}},{"cell_type":"markdown","source":"Based on the data in the 'Task' column, the following observations and descriptions can be made for each group:\n\n**TUG-ST (Timed Up and Go - Straight Line):**\n\nTUG-ST tasks have the highest frequency in the Task column, indicating their widespread use in Parkinson's disease research.\nTUG-ST is a test that evaluates the time taken by a patient to walk a certain distance in a straight line. This test measures the patient's mobility and strength.\n\n**TUG-DT (Timed Up and Go - Dual Task):**\n\nTUG-DT tasks also have a high frequency in the Task column, indicating their significance in Parkinson's disease research.\nTUG-DT is an extension of the TUG-ST test that involves performing an additional cognitive or motor task while walking. It assesses the patient's ability to perform two tasks simultaneously and can provide information about cognitive and motor deficits.\n\n**Turning-DT (Turning - Dual Task):**\n\nTurning-DT tasks have a moderate frequency in the Task column.\nTurning-DT assesses the patient's ability to turn in place while performing an additional cognitive or motor task. It provides information about the patient's balance and coordination under dual-task conditions.\n\n**4MW (4-Meter Walk):**\n\n4MW tasks also have a moderate frequency in the Task column.\n4MW is a test that measures the time taken by a patient to walk a distance of 4 meters. It assesses walking speed and provides information about the patient's mobility and coordination.\n\n**Turning-ST (Turning - Straight Line):**\n\nTurning-ST tasks have an average frequency in the Task column.\nTurning-ST assesses the patient's ability to turn in place in a straight line. It provides information about the patient's balance and coordination.\n\n**Hotspot1 and Hotspot2:**\n\nHotspot1 and Hotspot2 tasks have a low frequency in the Task column.\nThese tasks likely relate to specific measurements or data analysis.\n\n**Group MB6:**\n\nGroup MB6 has a very low frequency in the Task column.\nMB6 likely represents a specific subgroup of tasks or studies that may be unique and less commonly used. \n ","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\n\n# Calculation of correlation matrix\n\ncorrelation_matrix = merged_data[['UPDRSIII_On', 'UPDRSIII_Off', 'NFOGQ', 'Visit', 'Test', 'Medication', 'Init', 'Completion', 'Type', 'Kinetic', 'Begin', 'End', 'Task', 'Age', 'Sex', 'YearsSinceDx']].corr()\n\n# Creating a heatmap of the correlation matrix\n\nplt.figure(figsize=(10, 8))\nsns.heatmap(correlation_matrix, cmap='Blues', annot=True, fmt=\".2f\", annot_kws={\"fontsize\": 10})\nplt.title('Correlation Matrix Heatmap')\nplt.xticks(rotation=45)\nplt.yticks(rotation=0)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:31:55.624185Z","iopub.execute_input":"2023-05-23T12:31:55.625051Z","iopub.status.idle":"2023-05-23T12:31:56.493451Z","shell.execute_reply.started":"2023-05-23T12:31:55.625013Z","shell.execute_reply":"2023-05-23T12:31:56.492561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the computed correlation coefficients, the following observations can be made about the relationships between different columns:\n\n- Age and YearsSinceDx have a strong positive correlation (0.84), indicating that as the age of the patient increases, so do the years since diagnosis.\n\n- The NFOGQ score and YearsSinceDx also have a strong positive correlation (0.83), suggesting that with increasing years since diagnosis, there is a higher likelihood of experiencing freezing of gait phenomena.\n\n- YearsSinceDx has a strong positive correlation with UPDRSIII_Off (0.79) and UPDRSIII_On (0.85). This implies that as the years since diagnosis increase, the severity of symptoms measured by UPDRSIII worsens, both with and without medication.\n\n- UPDRSIII_Off and NFOGQ have a strong positive correlation (0.87), indicating that higher values of UPDRSIII_Off are associated with higher values of the NFOGQ index.\n\n- UPDRSIII_On and NFOGQ also have a strong positive correlation (0.89), indicating a relationship between higher values of UPDRSIII_On (considering medication) and higher values of the NFOGQ index.\n\n- UPDRSIII_On and UPDRSIII_Off have a strong positive correlation (0.83), indicating that these two measures are closely related, and an increase in one is usually accompanied by an increase in the other.","metadata":{}},{"cell_type":"code","source":"merged_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:00.365754Z","iopub.execute_input":"2023-05-23T12:32:00.366687Z","iopub.status.idle":"2023-05-23T12:32:00.381761Z","shell.execute_reply.started":"2023-05-23T12:32:00.366650Z","shell.execute_reply":"2023-05-23T12:32:00.380783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's try to train a model that can predict the set of features in patients most predisposed to motor manifestations","metadata":{}},{"cell_type":"code","source":"X = merged_data[['Medication', 'Kinetic', 'Task', 'Age', 'Sex', 'YearsSinceDx']]\ny = merged_data[['UPDRSIII_On', 'UPDRSIII_Off', 'NFOGQ']]\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:08.475464Z","iopub.execute_input":"2023-05-23T12:32:08.477815Z","iopub.status.idle":"2023-05-23T12:32:08.487562Z","shell.execute_reply.started":"2023-05-23T12:32:08.477779Z","shell.execute_reply":"2023-05-23T12:32:08.486705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Encoding categorical features using OneHotEncoder\nfrom sklearn.preprocessing import OneHotEncoder\n\ncategorical_features = ['Task']\nencoder = OneHotEncoder(sparse=False, handle_unknown='ignore')\n\nX_train_encoded = encoder.fit_transform(X_train[categorical_features])\nX_test_encoded = encoder.transform(X_test[categorical_features])","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:10.695511Z","iopub.execute_input":"2023-05-23T12:32:10.695884Z","iopub.status.idle":"2023-05-23T12:32:10.706849Z","shell.execute_reply.started":"2023-05-23T12:32:10.695853Z","shell.execute_reply":"2023-05-23T12:32:10.705766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Standardizing numerical features\nnumeric_features = ['Medication', 'Kinetic', 'Age', 'Sex', 'YearsSinceDx']\nscaler = StandardScaler()\nX_train_scaled = pd.DataFrame(scaler.fit_transform(X_train[numeric_features]), columns=numeric_features)\nX_test_scaled = pd.DataFrame(scaler.transform(X_test[numeric_features]), columns=numeric_features)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:15.819466Z","iopub.execute_input":"2023-05-23T12:32:15.819859Z","iopub.status.idle":"2023-05-23T12:32:15.833302Z","shell.execute_reply.started":"2023-05-23T12:32:15.819825Z","shell.execute_reply":"2023-05-23T12:32:15.832357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concatenating the transformed features\n\nX_train_encoded_df = pd.DataFrame(X_train_encoded, columns=encoder.get_feature_names_out(categorical_features))\nX_test_encoded_df = pd.DataFrame(X_test_encoded, columns=encoder.get_feature_names_out(categorical_features))\nX_train_processed = pd.concat([X_train_encoded_df, X_train_scaled], axis=1)\nX_test_processed = pd.concat([X_test_encoded_df, X_test_scaled], axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:19.660453Z","iopub.execute_input":"2023-05-23T12:32:19.660814Z","iopub.status.idle":"2023-05-23T12:32:19.668687Z","shell.execute_reply.started":"2023-05-23T12:32:19.660783Z","shell.execute_reply":"2023-05-23T12:32:19.667618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Handling missing values\n\nimputer = SimpleImputer(strategy='mean')\nX_train_processed = pd.DataFrame(imputer.fit_transform(X_train_processed), columns=X_train_processed.columns)\nX_test_processed = pd.DataFrame(imputer.transform(X_test_processed), columns=X_test_processed.columns)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:22.169580Z","iopub.execute_input":"2023-05-23T12:32:22.170563Z","iopub.status.idle":"2023-05-23T12:32:22.184257Z","shell.execute_reply.started":"2023-05-23T12:32:22.170515Z","shell.execute_reply":"2023-05-23T12:32:22.183305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating and training the model for UPDRSIII_On\n\nlogistic_regression_on = LogisticRegression(max_iter=5000)\nlogistic_regression_on.fit(X_train_processed, y_train['UPDRSIII_On'])\n\n# Creating and training the model for UPDRSIII_Off\n\nlogistic_regression_off = LogisticRegression(max_iter=5000)\nlogistic_regression_off.fit(X_train_processed, y_train['UPDRSIII_Off'])\n\n# Creating and training the model for NFOGQ\n\nlogistic_regression_nfogq = LogisticRegression(max_iter=5000)\nlogistic_regression_nfogq.fit(X_train_processed, y_train['NFOGQ'])\n\ny_train_pred_on = logistic_regression_on.predict(X_train_processed)\ny_test_pred_on = logistic_regression_on.predict(X_test_processed)\n\ny_train_pred_off = logistic_regression_off.predict(X_train_processed)\ny_test_pred_off = logistic_regression_off.predict(X_test_processed)\n\ny_train_pred_nfogq = logistic_regression_nfogq.predict(X_train_processed)\ny_test_pred_nfogq = logistic_regression_nfogq.predict(X_test_processed)","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:25.154636Z","iopub.execute_input":"2023-05-23T12:32:25.155634Z","iopub.status.idle":"2023-05-23T12:32:27.576415Z","shell.execute_reply.started":"2023-05-23T12:32:25.155589Z","shell.execute_reply":"2023-05-23T12:32:27.575172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from joblib import dump\n\n# Creating and training the combined model\n\ncombined_model = {\n    'UPDRSIII_On': logistic_regression_on,\n    'UPDRSIII_Off': logistic_regression_off,\n    'NFOGQ': logistic_regression_nfogq\n}\n\n# Saving the combined model\n\ndump(combined_model, 'combined_model.joblib')","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:31.077967Z","iopub.execute_input":"2023-05-23T12:32:31.078311Z","iopub.status.idle":"2023-05-23T12:32:31.091488Z","shell.execute_reply.started":"2023-05-23T12:32:31.078282Z","shell.execute_reply":"2023-05-23T12:32:31.090425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics_results = []\n\n# Metrics for the model UPDRSIII_On\nclass_name = 'UPDRSIII_On'\ntrain_accuracy = accuracy_score(y_train[class_name], y_train_pred_on)\ntest_accuracy = accuracy_score(y_test[class_name], y_test_pred_on)\nf1 = f1_score(y_test[class_name], y_test_pred_on, average='weighted')\n\nmetrics_results.append({\n    'Class': class_name,\n    'Train Accuracy': train_accuracy,\n    'Test Accuracy': test_accuracy,\n    'F1 Score': f1\n})\n\n# Metrics for the model UPDRSIII_Off\nclass_name = 'UPDRSIII_Off'\ntrain_accuracy = accuracy_score(y_train[class_name], y_train_pred_off)\ntest_accuracy = accuracy_score(y_test[class_name], y_test_pred_off)\nf1 = f1_score(y_test[class_name], y_test_pred_off, average='weighted')\n\nmetrics_results.append({\n    'Class': class_name,\n    'Train Accuracy': train_accuracy,\n    'Test Accuracy': test_accuracy,\n    'F1 Score': f1\n})\n\n# Metrics for the model NFOGQ\nclass_name = 'NFOGQ'\ntrain_accuracy = accuracy_score(y_train[class_name], y_train_pred_nfogq)\ntest_accuracy = accuracy_score(y_test[class_name], y_test_pred_nfogq)\nf1 = f1_score(y_test[class_name], y_test_pred_nfogq, average='weighted')\n\nmetrics_results.append({\n    'Class': class_name,\n    'Train Accuracy': train_accuracy,\n    'Test Accuracy': test_accuracy,\n    'F1 Score': f1\n})\n\n# Output of the results\nfor metrics in metrics_results:\n    print(f\"Metrics for class '{metrics['Class']}':\")\n    print(f\"Train Accuracy: {metrics['Train Accuracy']}\")\n    print(f\"Test Accuracy: {metrics['Test Accuracy']}\")\n    print(f\"F1 Score: {metrics['F1 Score']}\")\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:44.055339Z","iopub.execute_input":"2023-05-23T12:32:44.055690Z","iopub.status.idle":"2023-05-23T12:32:44.079621Z","shell.execute_reply.started":"2023-05-23T12:32:44.055659Z","shell.execute_reply":"2023-05-23T12:32:44.078702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a list of feature labels for the x-axis\n\nx_labels = X_train_processed.columns\n\n# Creating the plot\n\nplt.figure(figsize=(10, 6))\nplt.bar(x_labels, logistic_regression_on.coef_[0], color='lightblue', label='UPDRSIII_On')\nplt.bar(x_labels, logistic_regression_off.coef_[0], color='yellow', label='UPDRSIII_Off')\nplt.bar(x_labels, logistic_regression_nfogq.coef_[0], color='lightpink', label='NFOGQ')\n\n# Plot settings\n\nplt.title(\"Feature Importance\")\nplt.xlabel(\"Features\")\nplt.ylabel(\"Importance\")\nplt.xticks(rotation=90)\nplt.legend()\n\n# Displaying the plot\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-23T12:32:49.418569Z","iopub.execute_input":"2023-05-23T12:32:49.419210Z","iopub.status.idle":"2023-05-23T12:32:50.036476Z","shell.execute_reply.started":"2023-05-23T12:32:49.419176Z","shell.execute_reply":"2023-05-23T12:32:50.035647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusions:  \n\nThe model allows predicting the category of patients with Parkinson's disease and their most likely motor manifestations.\nFeatures UPDRSIII_On, UPDRSIII_Off, and NFOGQ, which are standardized scales and questionnaires, play an important role in assessing the symptoms of Parkinson's disease and their impact on patients.\n\nBased on the conducted research and provided information, the following conclusions can be drawn:\n\n* The trained model enables predicting the category of patients with Parkinson's disease and assessing the probability of their motor manifestations.  \n* Age is a significant factor influencing the severity of Parkinson's disease symptoms. The Age variable has distinct values for the UPDRSIII_On and Off targets.  \n* YearsSinceDx also have an impact on the motor manifestations of Parkinson's disease. With an increase in years since diagnosis, there is an increased probability of more pronounced symptoms, both with and without medication treatment.  \n* The UPDRSIII_On and UPDRSIII_Off scores have a strong positive correlation, indicating their interrelation and accompanying symptom changes before and after medication therapy.","metadata":{}}]}