{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Slee_GaussianHMM_undersample : 3 Acc variables\n\n# Introduction\n \nThe goal of this competition is to **detect freezing of gait (FOG)**, a debilitating symptom that afflicts many people **with Parkinson’s disease**. It is requred to **develop a machine learning model trained on data collected from a wearable 3D lower back sensor** to better understand **when and why FOG episodes occur**.","metadata":{"papermill":{"duration":0.019614,"end_time":"2023-05-03T03:15:41.562672","exception":false,"start_time":"2023-05-03T03:15:41.543058","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{"papermill":{"duration":0.017722,"end_time":"2023-05-03T03:15:41.601086","exception":false,"start_time":"2023-05-03T03:15:41.583364","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nfrom hmmlearn import hmm","metadata":{"papermill":{"duration":1.321847,"end_time":"2023-05-03T03:15:42.940889","exception":false,"start_time":"2023-05-03T03:15:41.619042","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:14:05.013250Z","iopub.execute_input":"2023-05-30T21:14:05.013584Z","iopub.status.idle":"2023-05-30T21:14:07.175161Z","shell.execute_reply.started":"2023-05-30T21:14:05.013550Z","shell.execute_reply":"2023-05-30T21:14:07.174153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Take All the CSV Files in the Train tdcsfog Folder","metadata":{"papermill":{"duration":0.019501,"end_time":"2023-05-03T03:15:44.857769","exception":false,"start_time":"2023-05-03T03:15:44.838268","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ntdcsfog_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/tdcsfog'\n\n# Initialize an empty list to store the dataframes.\ntdcsfog_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(tdcsfog_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(tdcsfog_path, file_name)\n        file = pd.read_csv(file_path)\n        file.Time = file.Time / (len(file) - 1)\n        tdcsfog_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ntdcsfog = pd.concat(tdcsfog_list, axis = 0)\n\n# Show the concatenated dataframe.\ntdcsfog","metadata":{"papermill":{"duration":19.002796,"end_time":"2023-05-03T03:16:03.880383","exception":false,"start_time":"2023-05-03T03:15:44.877587","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:14:07.177031Z","iopub.execute_input":"2023-05-30T21:14:07.177428Z","iopub.status.idle":"2023-05-30T21:14:21.348090Z","shell.execute_reply.started":"2023-05-30T21:14:07.177394Z","shell.execute_reply":"2023-05-30T21:14:21.347256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is better to reduce the memory usage. Reference: [Reducing DataFrame memory size by ~65%](https://www.kaggle.com/code/arjanso/reducing-dataframe-memory-size-by-65)","metadata":{"papermill":{"duration":0.020431,"end_time":"2023-05-03T03:16:03.922090","exception":false,"start_time":"2023-05-03T03:16:03.901659","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Show the numbers by graph.\ndata = pd.DataFrame(\n    np.concatenate([\n        ['Total'] * len(tdcsfog),\n        ['StartHesitation'] * int(np.ceil(len(tdcsfog) / 2 * tdcsfog['StartHesitation'].mean())),\n        ['Turn'] * int(np.ceil(len(tdcsfog) / 2 * tdcsfog['Turn'].mean())),\n        ['Walking'] * int(np.ceil(len(tdcsfog) / 2 * tdcsfog['Walking'].mean()))\n    ]),\n    columns = [\"The Number of 1\"]\n)\n\nsns.countplot(x = 'The Number of 1', data = data)","metadata":{"papermill":{"duration":8.320927,"end_time":"2023-05-03T03:16:16.993658","exception":false,"start_time":"2023-05-03T03:16:08.672731","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:14:21.349084Z","iopub.execute_input":"2023-05-30T21:14:21.349373Z","iopub.status.idle":"2023-05-30T21:14:27.166642Z","shell.execute_reply.started":"2023-05-30T21:14:21.349349Z","shell.execute_reply":"2023-05-30T21:14:27.165802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Most of the target variables are 0.**","metadata":{"papermill":{"duration":0.022327,"end_time":"2023-05-03T03:16:17.038616","exception":false,"start_time":"2023-05-03T03:16:17.016289","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Visualize \"AccV\", \"AccML\", and \"AccAP.\"\nsns.pairplot(tdcsfog[['AccV', 'AccML', 'AccAP']])","metadata":{"papermill":{"duration":195.940937,"end_time":"2023-05-03T03:19:33.002515","exception":false,"start_time":"2023-05-03T03:16:17.061578","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:14:27.168880Z","iopub.execute_input":"2023-05-30T21:14:27.169420Z","iopub.status.idle":"2023-05-30T21:16:11.455026Z","shell.execute_reply.started":"2023-05-30T21:14:27.169389Z","shell.execute_reply":"2023-05-30T21:16:11.454041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Take All the CSV Files in the Train defog Folder","metadata":{"papermill":{"duration":0.030832,"end_time":"2023-05-03T03:20:47.310476","exception":false,"start_time":"2023-05-03T03:20:47.279644","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ndefog_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/defog'\n\n# Initialize an empty list to store the dataframes.\ndefog_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(defog_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(defog_path, file_name)\n        file = pd.read_csv(file_path)\n        file.Time = file.Time / (len(file) - 1)\n        defog_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ndefog = pd.concat(defog_list, axis = 0)\n\n# Show the concatenated dataframe.\ndefog","metadata":{"papermill":{"duration":28.116394,"end_time":"2023-05-03T03:21:15.457334","exception":false,"start_time":"2023-05-03T03:20:47.340940","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:11.456107Z","iopub.execute_input":"2023-05-30T21:16:11.457050Z","iopub.status.idle":"2023-05-30T21:16:29.297646Z","shell.execute_reply.started":"2023-05-30T21:16:11.457021Z","shell.execute_reply":"2023-05-30T21:16:29.296376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We are going to use data with (Task=1 and Valid=1) only.**","metadata":{"papermill":{"duration":0.027633,"end_time":"2023-05-03T03:21:16.913824","exception":false,"start_time":"2023-05-03T03:21:16.886191","status":"completed"},"tags":[]}},{"cell_type":"code","source":"defog = defog[(defog['Task'] == 1) & (defog['Valid'] == 1)]","metadata":{"papermill":{"duration":0.506509,"end_time":"2023-05-03T03:21:17.448140","exception":false,"start_time":"2023-05-03T03:21:16.941631","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:29.299564Z","iopub.execute_input":"2023-05-30T21:16:29.300687Z","iopub.status.idle":"2023-05-30T21:16:29.760628Z","shell.execute_reply.started":"2023-05-30T21:16:29.300641Z","shell.execute_reply":"2023-05-30T21:16:29.759772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog.describe()","metadata":{"papermill":{"duration":2.058443,"end_time":"2023-05-03T03:21:19.643530","exception":false,"start_time":"2023-05-03T03:21:17.585087","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:29.761619Z","iopub.execute_input":"2023-05-30T21:16:29.761853Z","iopub.status.idle":"2023-05-30T21:16:30.454003Z","shell.execute_reply.started":"2023-05-30T21:16:29.761833Z","shell.execute_reply":"2023-05-30T21:16:30.453260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog = defog.iloc[:, :7]","metadata":{"papermill":{"duration":0.07985,"end_time":"2023-05-03T03:21:17.556087","exception":false,"start_time":"2023-05-03T03:21:17.476237","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:30.454948Z","iopub.execute_input":"2023-05-30T21:16:30.455642Z","iopub.status.idle":"2023-05-30T21:16:30.517986Z","shell.execute_reply.started":"2023-05-30T21:16:30.455613Z","shell.execute_reply":"2023-05-30T21:16:30.516796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We concatenate tdcsfog and defog datasets into one concatenated dataset.**","metadata":{"papermill":{"duration":0.028782,"end_time":"2023-05-03T03:21:19.703070","exception":false,"start_time":"2023-05-03T03:21:19.674288","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Concatenate the dataframes vertically using pd.concat().\nconcated = pd.concat([tdcsfog, defog], axis = 0)\n\n# Show the concatenated dataframe.\nconcated","metadata":{"papermill":{"duration":0.189226,"end_time":"2023-05-03T03:21:19.920330","exception":false,"start_time":"2023-05-03T03:21:19.731104","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:30.519372Z","iopub.execute_input":"2023-05-30T21:16:30.519731Z","iopub.status.idle":"2023-05-30T21:16:30.724980Z","shell.execute_reply.started":"2023-05-30T21:16:30.519689Z","shell.execute_reply":"2023-05-30T21:16:30.723910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Dataset\n\nFirst, we need to **split the data into input features (i.e., \"Time\", \"AccV\", \"AccML\", and \"AccAP\") and event variables (i.e., \"StartHesitation\", \"Turn\", and \"Walking\")**. We can do this using the .iloc method to select the appropriate columns.","metadata":{"papermill":{"duration":0.030405,"end_time":"2023-05-03T03:21:20.041453","exception":false,"start_time":"2023-05-03T03:21:20.011048","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# input features\n# add one new variable, \"None\" for NO_FOG\nconcated['None']= (concated['StartHesitation'] + concated['Turn'] \n                     + concated['Walking'] == 0).astype(int)\n\npd.crosstab (index= concated['None'], columns='count')","metadata":{"papermill":{"duration":0.31837,"end_time":"2023-05-03T03:21:20.390159","exception":false,"start_time":"2023-05-03T03:21:20.071789","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:30.729273Z","iopub.execute_input":"2023-05-30T21:16:30.729553Z","iopub.status.idle":"2023-05-30T21:16:32.511482Z","shell.execute_reply.started":"2023-05-30T21:16:30.729530Z","shell.execute_reply":"2023-05-30T21:16:32.510217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"concated['y'] = 0\nconcated.loc[concated['None']==1, 'y'] = 0  \nconcated.loc[concated['StartHesitation']==1, 'y'] = 1\nconcated.loc[concated['Turn']==1, 'y'] = 2 \nconcated.loc[concated['Walking']==1, 'y'] = 3 \npd.crosstab (index= concated['y'], columns='count')","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:32.512693Z","iopub.execute_input":"2023-05-30T21:16:32.513070Z","iopub.status.idle":"2023-05-30T21:16:34.257408Z","shell.execute_reply.started":"2023-05-30T21:16:32.513038Z","shell.execute_reply":"2023-05-30T21:16:34.256383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#defining feature matrix\nX = concated.iloc[:, 1:4]  # Acc_V, Acc_ML, Acc_AP only\n#X = concated.iloc[:, 0:4]  # Time, Acc_V, Acc_ML, Acc_AP only\n\n#corr_mtx = X.corr()\n#print (corr_mtx.round(3))","metadata":{}},{"cell_type":"code","source":"concated.info()","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:34.258616Z","iopub.execute_input":"2023-05-30T21:16:34.259640Z","iopub.status.idle":"2023-05-30T21:16:34.274997Z","shell.execute_reply.started":"2023-05-30T21:16:34.259594Z","shell.execute_reply":"2023-05-30T21:16:34.274033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"concated['y'].tail()","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:34.276217Z","iopub.execute_input":"2023-05-30T21:16:34.276502Z","iopub.status.idle":"2023-05-30T21:16:34.290367Z","shell.execute_reply.started":"2023-05-30T21:16:34.276480Z","shell.execute_reply":"2023-05-30T21:16:34.288863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the initial state probabilities\nStart_p = np.array([1,0,0,0])","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:34.291677Z","iopub.execute_input":"2023-05-30T21:16:34.292078Z","iopub.status.idle":"2023-05-30T21:16:34.300972Z","shell.execute_reply.started":"2023-05-30T21:16:34.292050Z","shell.execute_reply":"2023-05-30T21:16:34.300001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Will's Transistion Matrix - this is fixed\nT_m = np.array([[4870130, 100,    936,     99]\n                , [89,    304690, 11,      0]\n                , [947,   1,      1677824, 10]\n                , [96,    0,      12,      207730]])\n#print(T_m)\nrow_sums = np.sum(T_m, axis=1)\n#print(row_sums)\nTrans_p = T_m / row_sums[:, np.newaxis]\nprint(Trans_p)","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:34.301877Z","iopub.execute_input":"2023-05-30T21:16:34.302151Z","iopub.status.idle":"2023-05-30T21:16:34.315299Z","shell.execute_reply.started":"2023-05-30T21:16:34.302129Z","shell.execute_reply":"2023-05-30T21:16:34.314441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Nathanils Emissions Matrix - added 1's where there were 0's\nE_m = np.array([[654763, 164, 36468, 2399]\n                , [1,    1,   1,     1]\n                , [3738, 1,   45631, 2602]\n                , [318,  1,   1,     653]])\nE_m_T = E_m.T\n#print(E_m_T)\nrow_sums_e = np.sum(E_m_T, axis=1)\n#print(row_sums_e)\nEmiss_p = E_m_T / row_sums_e[:, np.newaxis]\nprint(Emiss_p)","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:34.316502Z","iopub.execute_input":"2023-05-30T21:16:34.317243Z","iopub.status.idle":"2023-05-30T21:16:34.328455Z","shell.execute_reply.started":"2023-05-30T21:16:34.317209Z","shell.execute_reply":"2023-05-30T21:16:34.327411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preparing X and y for training\nX = concated.iloc[:, 0:4]\ny_dat = concated['y']\nX.info()","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:34.329837Z","iopub.execute_input":"2023-05-30T21:16:34.330164Z","iopub.status.idle":"2023-05-30T21:16:34.438381Z","shell.execute_reply.started":"2023-05-30T21:16:34.330139Z","shell.execute_reply":"2023-05-30T21:16:34.436999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Since the result was so poor, we need to do balancing: under, over, hybrid\n#1 randomundersampling\nfrom collections import Counter\nfrom imblearn.under_sampling import RandomUnderSampler\nrus = RandomUnderSampler(random_state=11)\nX_under, y_under = rus.fit_resample(X, y_dat)\nprint(sorted(Counter(y_under).items()))\n","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:16:34.439991Z","iopub.execute_input":"2023-05-30T21:16:34.440283Z","iopub.status.idle":"2023-05-30T21:16:37.545537Z","shell.execute_reply.started":"2023-05-30T21:16:34.440259Z","shell.execute_reply":"2023-05-30T21:16:37.544513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" We will split the data of X and y into training and testing sets using train_test_split function","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_test, y_train, y_test = train_test_split(X_under,y_under , test_size = 0.2, random_state = 111)\n","metadata":{"papermill":{"duration":0.974759,"end_time":"2023-05-03T03:21:24.166856","exception":false,"start_time":"2023-05-03T03:21:23.192097","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:37.547128Z","iopub.execute_input":"2023-05-30T21:16:37.547420Z","iopub.status.idle":"2023-05-30T21:16:37.813798Z","shell.execute_reply.started":"2023-05-30T21:16:37.547394Z","shell.execute_reply":"2023-05-30T21:16:37.812732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then, we **standardize the independent variables**.","metadata":{"papermill":{"duration":0.028866,"end_time":"2023-05-03T03:21:24.224800","exception":false,"start_time":"2023-05-03T03:21:24.195934","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\n# Standardize the independent variables.\nscaler1 = StandardScaler()\nX_train = scaler1.fit_transform(X_train)\nX_test = scaler1.transform(X_test)","metadata":{"papermill":{"duration":0.829221,"end_time":"2023-05-03T03:21:25.083391","exception":false,"start_time":"2023-05-03T03:21:24.254170","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:37.815066Z","iopub.execute_input":"2023-05-30T21:16:37.815327Z","iopub.status.idle":"2023-05-30T21:16:37.867512Z","shell.execute_reply.started":"2023-05-30T21:16:37.815308Z","shell.execute_reply":"2023-05-30T21:16:37.865945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create, Train, and Evaluate Model\n\nFinally, we can **create and train three separate models**, one for each target variable, using a suitable algorithm, such as **hidden Markov model**: Ross, S. M.(2014). Introduction to Probability models, 11th Ed. NY: Academic Press.\n\n","metadata":{"papermill":{"duration":0.028215,"end_time":"2023-05-03T03:21:25.141519","exception":false,"start_time":"2023-05-03T03:21:25.113304","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Create GaussianHMM model and initial probs.\nmodel = hmm.GaussianHMM(n_components=4,covariance_type='full',\n                       algorithm='viterbi', n_iter=100, random_state=22)\n\nmodel.startprob_ = Start_p\nmodel.transmat_ = Trans_p\nmodel.emissionprob_ = Emiss_p\n\n# Train the model on the training data: model estimation stage.\nmodel.fit(X_train) # model estimation stage, no need of y_train","metadata":{"papermill":{"duration":4.449156,"end_time":"2023-05-03T03:21:29.618950","exception":false,"start_time":"2023-05-03T03:21:25.169794","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:16:37.868923Z","iopub.execute_input":"2023-05-30T21:16:37.869220Z","iopub.status.idle":"2023-05-30T21:19:04.229770Z","shell.execute_reply.started":"2023-05-30T21:16:37.869197Z","shell.execute_reply":"2023-05-30T21:19:04.228530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Prediction stage\ny_pred = model.predict(X_test)\npd.crosstab (index= y_pred, columns='count')\n#states = pd.unique(y_pred)\n\n#np.unique (states, return_counts = True)","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:19:04.231617Z","iopub.execute_input":"2023-05-30T21:19:04.231909Z","iopub.status.idle":"2023-05-30T21:19:04.352069Z","shell.execute_reply.started":"2023-05-30T21:19:04.231886Z","shell.execute_reply":"2023-05-30T21:19:04.351369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nIt is important to consider various metrics like **accuracy, precision, recall, and F1-score** to evaluate the model's performance in more detail.\n\n\\begin{equation*}\nAccuracy: Accuracy = \\frac{TP + TN}{TP + TN + FP + FN}\n\\end{equation*}\n\n\\begin{equation*}\nPrecision: Precision = \\frac{TP}{TP + FP}\n\\end{equation*}\n\n\\begin{equation*}\nRecall: Recall = \\frac{TP}{TP + FN}\n\\end{equation*}\n\n\\begin{equation*}\nF1-Score: F1 = 2 \\cdot \\frac{precision \\cdot recall}{precision + recall}\n\\end{equation*}\n\nwhere TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives.","metadata":{"papermill":{"duration":0.037315,"end_time":"2023-05-03T03:21:29.732896","exception":false,"start_time":"2023-05-03T03:21:29.695581","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### To create a classification report and confusion matrix, we will need to use the predictions made by the models with the test data.","metadata":{"papermill":{"duration":0.029698,"end_time":"2023-05-03T03:21:29.792601","exception":false,"start_time":"2023-05-03T03:21:29.762903","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.metrics import classification_report, confusion_matrix\n\n# Create a classification report for the model.\nprint('Classification Report')\nprint(classification_report(y_test, y_pred))\n\n# Create a confusion matrix for the model.\nprint('Confusion Matrix')\nprint(confusion_matrix(y_test, y_pred))","metadata":{"papermill":{"duration":2.357069,"end_time":"2023-05-03T03:21:32.179164","exception":false,"start_time":"2023-05-03T03:21:29.822095","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:04.353228Z","iopub.execute_input":"2023-05-30T21:19:04.353643Z","iopub.status.idle":"2023-05-30T21:19:04.690566Z","shell.execute_reply.started":"2023-05-30T21:19:04.353619Z","shell.execute_reply":"2023-05-30T21:19:04.689646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Confusion Matrix ","metadata":{"papermill":{"duration":0.02978,"end_time":"2023-05-03T03:21:32.297500","exception":false,"start_time":"2023-05-03T03:21:32.267720","status":"completed"},"tags":[]}},{"cell_type":"code","source":"cm = confusion_matrix(y_test, y_pred)\n\n# Define the class names.\nclass_names = ['None', 'StartHesitation','Turn', 'Walking']\n\n# # Create the heatmap with class names as tick labels.\nax = sns.heatmap(cm, annot = True, fmt = '.0f', cmap = \"Blues\", annot_kws = {\"size\": 16},\\\n           xticklabels = class_names, yticklabels = class_names)\n\n# # Set the axis labels.\nax.set_xlabel(\"Prediction\")\nax.set_ylabel(\"Truth\")\n\n# Vicky's code below\n# labs = [0,1,2,3]\n# cm = confusion_matrix(y_test, y_pred,labels=labs)\n# cm_df = pd.DataFrame(cm,index = labs, columns = labs)\n# sns.heatmap(cm_df, annot = True)\n# plt.ylabel(\"True Sequence\")\n# plt.xlabel(\"Hidden State Sequence\")\n","metadata":{"papermill":{"duration":0.314667,"end_time":"2023-05-03T03:21:32.643381","exception":false,"start_time":"2023-05-03T03:21:32.328714","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:04.691633Z","iopub.execute_input":"2023-05-30T21:19:04.691934Z","iopub.status.idle":"2023-05-30T21:19:04.979008Z","shell.execute_reply.started":"2023-05-30T21:19:04.691907Z","shell.execute_reply":"2023-05-30T21:19:04.977697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Test Dataset","metadata":{"papermill":{"duration":0.032959,"end_time":"2023-05-03T03:21:33.750656","exception":false,"start_time":"2023-05-03T03:21:33.717697","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ntdcsfog_test_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/tdcsfog'\n\n# Initialize an empty list to store the dataframes.\ntdcsfog_test_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(tdcsfog_test_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(tdcsfog_test_path, file_name)\n        file = pd.read_csv(file_path)\n        file['Id'] = file_name[:-4] + '_' + file['Time'].apply(str)\n        file.Time = file.Time / (len(file) - 1)\n        tdcsfog_test_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ntdcsfog_test = pd.concat(tdcsfog_test_list, axis = 0)\n\n# Show the concatenated dataframe.\ntdcsfog_test","metadata":{"papermill":{"duration":0.080681,"end_time":"2023-05-03T03:21:33.864773","exception":false,"start_time":"2023-05-03T03:21:33.784092","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:04.980634Z","iopub.execute_input":"2023-05-30T21:19:04.980968Z","iopub.status.idle":"2023-05-30T21:19:05.020425Z","shell.execute_reply.started":"2023-05-30T21:19:04.980943Z","shell.execute_reply":"2023-05-30T21:19:05.019324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ndefog_test_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/defog'\n\n# Initialize an empty list to store the dataframes.\ndefog_test_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(defog_test_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(defog_test_path, file_name)\n        file = pd.read_csv(file_path)\n        file['Id'] = file_name[:-4] + '_' + file['Time'].apply(str)\n        file.Time = file.Time / (len(file) - 1)\n        defog_test_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ndefog_test = pd.concat(defog_test_list, axis = 0)\n\n# Show the concatenated dataframe.\ndefog_test","metadata":{"papermill":{"duration":0.602419,"end_time":"2023-05-03T03:21:34.589276","exception":false,"start_time":"2023-05-03T03:21:33.986857","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:05.021815Z","iopub.execute_input":"2023-05-30T21:19:05.022089Z","iopub.status.idle":"2023-05-30T21:19:05.443140Z","shell.execute_reply.started":"2023-05-30T21:19:05.022067Z","shell.execute_reply":"2023-05-30T21:19:05.441783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.concat([tdcsfog_test, defog_test], axis = 0).reset_index(drop = True)\ntest","metadata":{"papermill":{"duration":0.201613,"end_time":"2023-05-03T03:21:35.308538","exception":false,"start_time":"2023-05-03T03:21:35.106925","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:05.444405Z","iopub.execute_input":"2023-05-30T21:19:05.444686Z","iopub.status.idle":"2023-05-30T21:19:05.471133Z","shell.execute_reply.started":"2023-05-30T21:19:05.444663Z","shell.execute_reply":"2023-05-30T21:19:05.469877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference","metadata":{"papermill":{"duration":0.033653,"end_time":"2023-05-03T03:21:35.376460","exception":false,"start_time":"2023-05-03T03:21:35.342807","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Separate the dataset for the independent variables.\ntest_X = test.iloc[:, 0:4]\n\n# Standardize the independent variables by a new scaler.\nscaler = StandardScaler()\nscaler.fit(test_X)\ntest_X = scaler.transform(test_X)\n\n# Get the predictions for the model on the test data.\npred_y = model.predict(test_X)\npd.crosstab(index= pred_y, columns='count')\n","metadata":{"papermill":{"duration":0.208078,"end_time":"2023-05-03T03:21:35.618421","exception":false,"start_time":"2023-05-03T03:21:35.410343","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:05.475342Z","iopub.execute_input":"2023-05-30T21:19:05.475678Z","iopub.status.idle":"2023-05-30T21:19:05.622795Z","shell.execute_reply.started":"2023-05-30T21:19:05.475655Z","shell.execute_reply":"2023-05-30T21:19:05.621742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # select SH, Turn, Walking from pred_y and create pred_y1, pred_y2,pred_y3\npred_y_df=pd.DataFrame ({'pred_y': pred_y})\npred_y1=pred_y_df['pred_y'].apply(lambda x:1 if x==1 else 0)\npred_y2=pred_y_df['pred_y'].apply(lambda x:1 if x==2 else 0)\npred_y3=pred_y_df['pred_y'].apply(lambda x:1 if x==3 else 0)\n\npd.crosstab(index= pred_y2, columns='count')","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:19:05.625679Z","iopub.execute_input":"2023-05-30T21:19:05.626407Z","iopub.status.idle":"2023-05-30T21:19:05.948163Z","shell.execute_reply.started":"2023-05-30T21:19:05.626381Z","shell.execute_reply":"2023-05-30T21:19:05.947343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test['StartHesitation'] = pred_y1\ntest['Turn'] = pred_y2\ntest['Walking'] = pred_y3","metadata":{"execution":{"iopub.status.busy":"2023-05-30T21:19:05.949960Z","iopub.execute_input":"2023-05-30T21:19:05.950356Z","iopub.status.idle":"2023-05-30T21:19:05.957551Z","shell.execute_reply.started":"2023-05-30T21:19:05.950324Z","shell.execute_reply":"2023-05-30T21:19:05.956466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{"papermill":{"duration":0.034728,"end_time":"2023-05-03T03:21:35.735065","exception":false,"start_time":"2023-05-03T03:21:35.700337","status":"completed"},"tags":[]}},{"cell_type":"code","source":"submission = test.iloc[:, 4:].fillna(0.0)\nsubmission","metadata":{"papermill":{"duration":0.088712,"end_time":"2023-05-03T03:21:35.858102","exception":false,"start_time":"2023-05-03T03:21:35.769390","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:05.958864Z","iopub.execute_input":"2023-05-30T21:19:05.959186Z","iopub.status.idle":"2023-05-30T21:19:06.025572Z","shell.execute_reply.started":"2023-05-30T21:19:05.959163Z","shell.execute_reply":"2023-05-30T21:19:06.024553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index = False)","metadata":{"papermill":{"duration":0.483855,"end_time":"2023-05-03T03:21:36.376732","exception":false,"start_time":"2023-05-03T03:21:35.892877","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-05-30T21:19:06.026541Z","iopub.execute_input":"2023-05-30T21:19:06.026792Z","iopub.status.idle":"2023-05-30T21:19:06.421186Z","shell.execute_reply.started":"2023-05-30T21:19:06.026773Z","shell.execute_reply":"2023-05-30T21:19:06.420236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\n\nIt is possible that **more features or more advanced machine learning algorithms** could improve the precision of the models. Additionally, it may be useful to **investigate other factors** that contribute to the occurrence of freezing of gait events, such as cognitive or environmental factors.","metadata":{"papermill":{"duration":0.046307,"end_time":"2023-05-03T03:21:36.747360","exception":false,"start_time":"2023-05-03T03:21:36.701053","status":"completed"},"tags":[]}}]}