{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom sklearn import *\nimport glob\nimport gc\nfrom pathlib import Path","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:31.055433Z","iopub.execute_input":"2023-04-17T02:15:31.055887Z","iopub.status.idle":"2023-04-17T02:15:33.799225Z","shell.execute_reply.started":"2023-04-17T02:15:31.055841Z","shell.execute_reply":"2023-04-17T02:15:33.797653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Understanding the Data (Train, Unlabeled, and Metadata)**","metadata":{}},{"cell_type":"markdown","source":"* The **tDCS FOG** (tdcsfog) dataset, comprising data series collected in the lab, as subjects completed a FOG-provoking protocol.\n* The **DeFOG** (defog) dataset, comprising data series collected in the subject's home, as subjects completed a FOG-provoking protocol\n* The **Daily Living** (daily) dataset, comprising one week of continuous 24/7 recordings from sixty-five subjects. Forty-five subjects exhibit FOG symptoms and also have series in the defog dataset, while the other twenty subjects do not exhibit FOG symptoms and do not have series elsewhere in the data.","metadata":{}},{"cell_type":"markdown","source":"# tdcsfog Meta Data","metadata":{}},{"cell_type":"markdown","source":"Identifies each series in the tdcsfog dataset by a unique Subject, Visit, Test, Medication condition.\n* Visit Lab visits consist of a baseline assessment, two post-treatment assessments for different treatment stages, and one follow-up assessment.\n* Test Which of three test types was performed, with 3 the most challenging.\n* Medication Subjects may have been either off or on anti-parkinsonian medication during the recording","metadata":{}},{"cell_type":"markdown","source":"# Daily MetaData","metadata":{}},{"cell_type":"markdown","source":"Each series in the daily dataset is identified by the Subject id. This file also contains the time of day the recording began.","metadata":{}},{"cell_type":"markdown","source":"# Events","metadata":{}},{"cell_type":"markdown","source":"Metadata for each FoG event in all data series. The event times agree with the labels in the data series.\n* Id The data series the event occured in.\n* Init Time (s) the event began.\n* Completion Time (s) the event ended.\n* Type Whether StartHesitation, Turn, or Walking.\n* Kinetic Whether the event was kinetic (1) and involved movement, or akinetic (0) and static.","metadata":{}},{"cell_type":"markdown","source":"# Subject","metadata":{}},{"cell_type":"markdown","source":"Metadata for each Subject in the study, including their Age and Sex as well as:\n* Visit Only available for subjects in the daily and defog datasets.\n* YearsSinceDx Years since Parkinson's diagnosis.\n* UPDRSIIIOn/UPDRSIIIOff Unified Parkinson's Disease Rating Scale score during on/off medication respectively.\n* NFOGQ Self-report FoG questionnaire score. See:\nhttps://pubmed.ncbi.nlm.nih.gov/19660949/","metadata":{}},{"cell_type":"markdown","source":"# Tasks","metadata":{}},{"cell_type":"markdown","source":"Task metadata for series in the defog dataset. (Not relevant for the series in the fog or daily datasets.)\n* Id The data series where the task was measured.\n* Begin Time (s) the task began.\n* End Time (s) the task ended.\n* Task One of seven tasks types in the DeFOG protocol, described on this page.","metadata":{}},{"cell_type":"markdown","source":"# Merging Metadata and Subject data","metadata":{}},{"cell_type":"code","source":"defog_metadata_df = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/defog_metadata.csv')\ntdcsfog_meta_df = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/tdcsfog_metadata.csv')\nsubject_df = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/subjects.csv')","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.802078Z","iopub.execute_input":"2023-04-17T02:15:33.803208Z","iopub.status.idle":"2023-04-17T02:15:33.837116Z","shell.execute_reply.started":"2023-04-17T02:15:33.803145Z","shell.execute_reply":"2023-04-17T02:15:33.835870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata=pd.concat([defog_metadata_df,tdcsfog_meta_df])","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.839317Z","iopub.execute_input":"2023-04-17T02:15:33.840072Z","iopub.status.idle":"2023-04-17T02:15:33.854996Z","shell.execute_reply.started":"2023-04-17T02:15:33.840018Z","shell.execute_reply":"2023-04-17T02:15:33.853468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata['Medication'] = np.where(metadata['Medication'] != 'on', 0, 1)\nmetadata_new=metadata.merge(subject_df,how='left',on='Subject').copy()\nmetadata_new.fillna(0,inplace=True)\nmetadata_new['Sex'] = np.where(metadata_new['Sex'] != 'M', 0, 1)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.859071Z","iopub.execute_input":"2023-04-17T02:15:33.860183Z","iopub.status.idle":"2023-04-17T02:15:33.893646Z","shell.execute_reply.started":"2023-04-17T02:15:33.860115Z","shell.execute_reply":"2023-04-17T02:15:33.892333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bins = [0,10,20,30,40,50,60,70,80,90,100]\nlabels = [0,1,2,3,4,5,6,7,8,9]\nmetadata_new['Age_Category'] = pd.cut(metadata_new['Age'],bins,labels=labels)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.895813Z","iopub.execute_input":"2023-04-17T02:15:33.896249Z","iopub.status.idle":"2023-04-17T02:15:33.908413Z","shell.execute_reply.started":"2023-04-17T02:15:33.896200Z","shell.execute_reply":"2023-04-17T02:15:33.907379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bins = [0,5,10,15,20,25,30,35]\nlabels = [0,1,2,3,4,5,6]\nmetadata_new['YearsSinceDx_Category'] = pd.cut(metadata_new['YearsSinceDx'],bins,labels=labels)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.910049Z","iopub.execute_input":"2023-04-17T02:15:33.910921Z","iopub.status.idle":"2023-04-17T02:15:33.931301Z","shell.execute_reply.started":"2023-04-17T02:15:33.910876Z","shell.execute_reply":"2023-04-17T02:15:33.929721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata_new.drop(['UPDRSIII_On', 'UPDRSIII_Off', 'NFOGQ', 'Visit_x', 'Visit_y', 'Test'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.933320Z","iopub.execute_input":"2023-04-17T02:15:33.934224Z","iopub.status.idle":"2023-04-17T02:15:33.951763Z","shell.execute_reply.started":"2023-04-17T02:15:33.934166Z","shell.execute_reply":"2023-04-17T02:15:33.950256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata_new = metadata_new.drop_duplicates()\nmetadata_new = metadata_new.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.953261Z","iopub.execute_input":"2023-04-17T02:15:33.954047Z","iopub.status.idle":"2023-04-17T02:15:33.976412Z","shell.execute_reply.started":"2023-04-17T02:15:33.954001Z","shell.execute_reply":"2023-04-17T02:15:33.974733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Defog and Tdcfog","metadata":{}},{"cell_type":"code","source":"path=\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/\"\n\ntrain_defog = glob.glob(path+'train/defog/**')\ntrain_tdcsfog = glob.glob(path+'train/tdcsfog/**')","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:33.978230Z","iopub.execute_input":"2023-04-17T02:15:33.979042Z","iopub.status.idle":"2023-04-17T02:15:34.072999Z","shell.execute_reply.started":"2023-04-17T02:15:33.978989Z","shell.execute_reply":"2023-04-17T02:15:34.071454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_data(f):\n    df = pd.read_csv(f)\n    df['Id'] = f.split('/')[-1].split('.')[0]\n    df['data_type'] = f.split('/')[-2]\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:34.079526Z","iopub.execute_input":"2023-04-17T02:15:34.080045Z","iopub.status.idle":"2023-04-17T02:15:34.087167Z","shell.execute_reply.started":"2023-04-17T02:15:34.079989Z","shell.execute_reply":"2023-04-17T02:15:34.085703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_defog = pd.concat([get_data(f) for f in train_defog])\ndf_train_tdcsfog = pd.concat([get_data(f) for f in train_tdcsfog])","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:15:34.088730Z","iopub.execute_input":"2023-04-17T02:15:34.090014Z","iopub.status.idle":"2023-04-17T02:16:21.065828Z","shell.execute_reply.started":"2023-04-17T02:15:34.089967Z","shell.execute_reply":"2023-04-17T02:16:21.064802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Converting times to make dbs the same","metadata":{}},{"cell_type":"code","source":"def convert_time(df, freq):\n    if freq == 100:\n        df['Time'] /= 100\n        df['Time'] = round(df['Time'], 3)\n    elif freq == 128:\n        df['Time'] = round(df['Time'] / 128, )\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:16:21.067376Z","iopub.execute_input":"2023-04-17T02:16:21.068357Z","iopub.status.idle":"2023-04-17T02:16:21.074098Z","shell.execute_reply.started":"2023-04-17T02:16:21.068314Z","shell.execute_reply":"2023-04-17T02:16:21.072920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_defog = convert_time(df_train_defog, 100)\ndf_train_tdcsfog = convert_time(df_train_tdcsfog, 128)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:16:21.075524Z","iopub.execute_input":"2023-04-17T02:16:21.075861Z","iopub.status.idle":"2023-04-17T02:16:21.499573Z","shell.execute_reply.started":"2023-04-17T02:16:21.075830Z","shell.execute_reply":"2023-04-17T02:16:21.498105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def convert_acc(df):\n    df['AccV'] *= 9.80665\n    df['AccML'] *= 9.80665\n    df['AccAP'] *= 9.80665\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:16:21.501627Z","iopub.execute_input":"2023-04-17T02:16:21.502319Z","iopub.status.idle":"2023-04-17T02:16:21.512512Z","shell.execute_reply.started":"2023-04-17T02:16:21.502240Z","shell.execute_reply":"2023-04-17T02:16:21.511096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_defog = convert_acc(df_train_defog)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:16:21.514118Z","iopub.execute_input":"2023-04-17T02:16:21.514674Z","iopub.status.idle":"2023-04-17T02:16:21.859133Z","shell.execute_reply.started":"2023-04-17T02:16:21.514611Z","shell.execute_reply":"2023-04-17T02:16:21.858009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df=pd.concat([df_train_defog,df_train_tdcsfog])","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:16:21.860666Z","iopub.execute_input":"2023-04-17T02:16:21.861349Z","iopub.status.idle":"2023-04-17T02:16:26.204853Z","shell.execute_reply.started":"2023-04-17T02:16:21.861298Z","shell.execute_reply":"2023-04-17T02:16:26.203616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.fillna(0)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:16:26.206233Z","iopub.execute_input":"2023-04-17T02:16:26.206673Z","iopub.status.idle":"2023-04-17T02:16:41.913795Z","shell.execute_reply.started":"2023-04-17T02:16:26.206634Z","shell.execute_reply":"2023-04-17T02:16:41.912443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Combining w/ Metadata","metadata":{}},{"cell_type":"code","source":"combined_df = df.merge(metadata_new,how='left',on='Id').copy()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:16:41.915730Z","iopub.execute_input":"2023-04-17T02:16:41.916526Z","iopub.status.idle":"2023-04-17T02:17:04.573331Z","shell.execute_reply.started":"2023-04-17T02:16:41.916473Z","shell.execute_reply":"2023-04-17T02:17:04.571937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df['data_type'] = np.where(combined_df['data_type'] != 'defog', 0, 1)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:04.574870Z","iopub.execute_input":"2023-04-17T02:17:04.575245Z","iopub.status.idle":"2023-04-17T02:17:07.692504Z","shell.execute_reply.started":"2023-04-17T02:17:04.575207Z","shell.execute_reply":"2023-04-17T02:17:07.691116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# combined_df","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:07.694125Z","iopub.execute_input":"2023-04-17T02:17:07.694686Z","iopub.status.idle":"2023-04-17T02:17:07.700167Z","shell.execute_reply.started":"2023-04-17T02:17:07.694628Z","shell.execute_reply":"2023-04-17T02:17:07.698659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features=['Time', 'AccV', 'AccML', 'AccAP', 'Age_Category', 'Medication', 'YearsSinceDx_Category', 'data_type']\nTargets=['StartHesitation', 'Turn' , 'Walking']","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:07.702171Z","iopub.execute_input":"2023-04-17T02:17:07.703003Z","iopub.status.idle":"2023-04-17T02:17:07.712788Z","shell.execute_reply.started":"2023-04-17T02:17:07.702948Z","shell.execute_reply":"2023-04-17T02:17:07.711485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tuning Hyperparams","metadata":{}},{"cell_type":"code","source":"# from sklearn.model_selection import RandomizedSearchCV\n# from pprint import pprint\n\n# n_estimators = [int(x) for x in np.linspace(start = 50, stop = 150, num = 5)]\n# max_depth = [int(x) for x in np.linspace(5, 15, num = 3)]\n# max_depth.append(None)\n# min_samples_split = [2, 5, 10]\n# min_samples_leaf = [1, 2, 4]\n# max_features = ['sqrt', 'log2']\n\n# random_grid = {'n_estimators': n_estimators,\n#                 'max_depth': max_depth,\n#                 'min_samples_split': min_samples_split,\n#                 'min_samples_leaf': min_samples_leaf,\n#                 'max_features': max_features\n#               }\n# pprint(random_grid)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:07.714154Z","iopub.execute_input":"2023-04-17T02:17:07.714582Z","iopub.status.idle":"2023-04-17T02:17:07.724368Z","shell.execute_reply.started":"2023-04-17T02:17:07.714544Z","shell.execute_reply":"2023-04-17T02:17:07.723168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# X, x, Y, y = model_selection.train_test_split(subset_df_train[features], subset_df_train[Targets], test_size=.30, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:07.725735Z","iopub.execute_input":"2023-04-17T02:17:07.726137Z","iopub.status.idle":"2023-04-17T02:17:07.734173Z","shell.execute_reply.started":"2023-04-17T02:17:07.726086Z","shell.execute_reply":"2023-04-17T02:17:07.733119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#print(y.dtypes)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:07.735505Z","iopub.execute_input":"2023-04-17T02:17:07.736536Z","iopub.status.idle":"2023-04-17T02:17:07.747260Z","shell.execute_reply.started":"2023-04-17T02:17:07.736495Z","shell.execute_reply":"2023-04-17T02:17:07.746291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from sklearn.model_selection import GridSearchCV\n\n# rf = ensemble.RandomForestRegressor(random_state=42)\n\n# grid_search = GridSearchCV(rf, param_grid=random_grid, cv=5, n_jobs=-1)\n\n# grid_search.fit(X, Y)\n\n# print(\"Best hyperparameters: \", grid_search.best_params_)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:07.748775Z","iopub.execute_input":"2023-04-17T02:17:07.749564Z","iopub.status.idle":"2023-04-17T02:17:07.759402Z","shell.execute_reply.started":"2023-04-17T02:17:07.749518Z","shell.execute_reply":"2023-04-17T02:17:07.758222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Best hyperparameters:  {'max_depth': 10, 'max_features': 'sqrt', 'min_samples_leaf': 1, 'min_samples_split': 2, 'n_estimators': 125}","metadata":{}},{"cell_type":"markdown","source":"# Combined Rf Training and Testing ","metadata":{}},{"cell_type":"code","source":"X_train, X_valid, y_train, y_valid = model_selection.train_test_split(combined_df[features], combined_df[Targets], test_size=.30, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:07.761089Z","iopub.execute_input":"2023-04-17T02:17:07.762210Z","iopub.status.idle":"2023-04-17T02:17:22.383407Z","shell.execute_reply.started":"2023-04-17T02:17:07.762169Z","shell.execute_reply":"2023-04-17T02:17:22.380026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf = ensemble.RandomForestRegressor(max_depth= 10, min_samples_leaf=1, min_samples_split=2, n_estimators=125, max_features='sqrt', random_state=42, n_jobs=-1)\nrf.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:17:22.389217Z","iopub.execute_input":"2023-04-17T02:17:22.391064Z","iopub.status.idle":"2023-04-17T02:45:43.694875Z","shell.execute_reply.started":"2023-04-17T02:17:22.390924Z","shell.execute_reply":"2023-04-17T02:45:43.691917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = rf.predict(X_valid)\nprint(metrics.average_precision_score(y_valid, y_pred.clip(0.0,1.0)))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:45:43.706341Z","iopub.execute_input":"2023-04-17T02:45:43.707498Z","iopub.status.idle":"2023-04-17T02:46:25.593688Z","shell.execute_reply.started":"2023-04-17T02:45:43.707438Z","shell.execute_reply":"2023-04-17T02:46:25.592200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# importances = rf.feature_importances_\n\n# # Plot feature importances\n# plt.bar(X_train.columns, importances)\n# plt.xticks(rotation=90)\n# plt.ylabel('Importance')\n# plt.xlabel('Feature')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:46:25.595742Z","iopub.execute_input":"2023-04-17T02:46:25.596124Z","iopub.status.idle":"2023-04-17T02:46:25.601694Z","shell.execute_reply.started":"2023-04-17T02:46:25.596086Z","shell.execute_reply":"2023-04-17T02:46:25.600466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluation ","metadata":{}},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"test_files = glob.glob(path + 'test/**/*.csv', recursive=True)\nsub = pd.read_csv(path+'sample_submission.csv')\nsub['t'] = 0\ntargets = ['Id', 'StartHesitation', 'Turn' , 'Walking']\n\nsubmission = []\n\nfor file in test_files:\n    test_df = pd.read_csv(file)\n    id_str = file.split('/')[-1].split('.')[0]\n\n    data_type = 1 if 'defog' in id_str else 0\n    df = test_df.assign(Id=id_str, data_type=data_type)\n    df = convert_time(df, 100) if data_type else convert_time(df, 128)\n    df = convert_acc(df) if data_type else df\n    df = df.merge(metadata_new, how='left', on='Id')\n    df = df.fillna(0).reset_index(drop=True)\n\n    results = pd.DataFrame(np.round(rf.predict(df[features]),3), columns=Targets)\n    \n    df_ = pd.concat([df[['Id']], results], axis=1)\n    \n    df_['Id'] = df_['Id'].astype(str) + '_' + df_.index.astype(str)\n \n    submission.append(df_[targets])    \n    \nsubmission = pd.concat(submission)\nsubmission = pd.merge(sub[['Id']], submission, how='left', on='Id').fillna(0.0)\nsubmission[targets].to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:47:16.777299Z","iopub.execute_input":"2023-04-17T02:47:16.777774Z","iopub.status.idle":"2023-04-17T02:47:19.522594Z","shell.execute_reply.started":"2023-04-17T02:47:16.777730Z","shell.execute_reply":"2023-04-17T02:47:19.521434Z"},"trusted":true},"execution_count":null,"outputs":[]}]}