{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.10","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":41880,"databundleVersionId":5677426,"sourceType":"competition"}],"dockerImageVersionId":30474,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nThe goal of this competition is to **detect freezing of gait (FOG)**, a debilitating symptom that afflicts many people **with Parkinson’s disease**. It is required to **develop a machine learning model trained on data collected from a wearable 3D lower back sensor** to better understand **when and why FOG episodes occur**.","metadata":{"papermill":{"duration":0.021368,"end_time":"2023-05-31T09:27:57.316377","exception":false,"start_time":"2023-05-31T09:27:57.295009","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{"papermill":{"duration":0.021034,"end_time":"2023-05-31T09:27:57.358461","exception":false,"start_time":"2023-05-31T09:27:57.337427","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os","metadata":{"papermill":{"duration":1.279093,"end_time":"2023-05-31T09:27:58.658747","exception":false,"start_time":"2023-05-31T09:27:57.379654","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:12.455637Z","iopub.execute_input":"2025-05-20T10:01:12.456047Z","iopub.status.idle":"2025-05-20T10:01:14.415347Z","shell.execute_reply.started":"2025-05-20T10:01:12.456016Z","shell.execute_reply":"2025-05-20T10:01:14.414029Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Read CSV Files\n\nThe train dataset is made up of the defog, notype, and tdcsfog folders. **The tdcsfog folder includes** more than 800 csv files and accounts for **the majority of the information contained in the whole dataset for this competition**. Moreover, csv files in the edcsfog folder include annotation: StartHesitation, Turn, and Walking.","metadata":{"papermill":{"duration":0.020632,"end_time":"2023-05-31T09:27:58.700670","exception":false,"start_time":"2023-05-31T09:27:58.680038","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# one tdcsfog file\ntdcsfog_003f117e14 = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/tdcsfog/003f117e14.csv')\ntdcsfog_003f117e14.head(5)","metadata":{"papermill":{"duration":0.086349,"end_time":"2023-05-31T09:27:58.808648","exception":false,"start_time":"2023-05-31T09:27:58.722299","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.417943Z","iopub.execute_input":"2025-05-20T10:01:14.418389Z","iopub.status.idle":"2025-05-20T10:01:14.489560Z","shell.execute_reply.started":"2025-05-20T10:01:14.418348Z","shell.execute_reply":"2025-05-20T10:01:14.488487Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tdcsfog_003f117e14.tail(5)","metadata":{"papermill":{"duration":0.03836,"end_time":"2023-05-31T09:27:58.870417","exception":false,"start_time":"2023-05-31T09:27:58.832057","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.494639Z","iopub.execute_input":"2025-05-20T10:01:14.495050Z","iopub.status.idle":"2025-05-20T10:01:14.512546Z","shell.execute_reply.started":"2025-05-20T10:01:14.495011Z","shell.execute_reply":"2025-05-20T10:01:14.511330Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This looks like a table of motion sensor data, possibly from a wearable device, with each row representing a particular moment in time (as indicated by the \"Time\" column) and containing measurements of acceleration in three dimensions (**\"AccV\", \"AccML\", \"AccAP\"**). **The values in these columns indicate the acceleration of the device along the vertical, medial-lateral, and anterior-posterior axes, respectively.**\n\n**The \"StartHesitation\", \"Turn\", and \"Walking\" columns appear to be binary variables indicating whether or not the corresponding activity is taking place at a given moment in time.** However, in this table, all values in these columns are 0, so it is not possible to predict the status of these variables based on the data in this table alone.","metadata":{"papermill":{"duration":0.021119,"end_time":"2023-05-31T09:27:58.912904","exception":false,"start_time":"2023-05-31T09:27:58.891785","status":"completed"},"tags":[]}},{"cell_type":"code","source":"tdcsfog_003f117e14.describe()","metadata":{"papermill":{"duration":0.065167,"end_time":"2023-05-31T09:27:58.999676","exception":false,"start_time":"2023-05-31T09:27:58.934509","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.514176Z","iopub.execute_input":"2025-05-20T10:01:14.514610Z","iopub.status.idle":"2025-05-20T10:01:14.566952Z","shell.execute_reply.started":"2025-05-20T10:01:14.514553Z","shell.execute_reply":"2025-05-20T10:01:14.565703Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The \"mean\", \"std\", \"min\", \"25%\", \"50%\", \"75%\", and \"max\" values for each variable provide information about the distribution and range of the data. For example, the mean value for \"AccV\" is -9.15, indicating that the device was tilted slightly downwards on average. The standard deviation of 1.38 for \"AccV\" suggests that the device orientation varied quite a bit over time. The maximum and minimum values for each variable provide an idea of the range of motion that was captured by the sensor over the course of the data collection period.\n\nHere, **only Turn has the max value of 1**. Thus, **only a turn event occurred** in this experiment.","metadata":{"papermill":{"duration":0.02174,"end_time":"2023-05-31T09:27:59.043303","exception":false,"start_time":"2023-05-31T09:27:59.021563","status":"completed"},"tags":[]}},{"cell_type":"code","source":"tdcsfog_003f117e14[tdcsfog_003f117e14['Turn'] == 1]","metadata":{"papermill":{"duration":0.046011,"end_time":"2023-05-31T09:27:59.111311","exception":false,"start_time":"2023-05-31T09:27:59.065300","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.568688Z","iopub.execute_input":"2025-05-20T10:01:14.569159Z","iopub.status.idle":"2025-05-20T10:01:14.592088Z","shell.execute_reply.started":"2025-05-20T10:01:14.569117Z","shell.execute_reply":"2025-05-20T10:01:14.589953Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**The sensor data was collected during this period of time (from 1103 to 1890) when the subject was making turns.**","metadata":{"papermill":{"duration":0.022084,"end_time":"2023-05-31T09:27:59.155643","exception":false,"start_time":"2023-05-31T09:27:59.133559","status":"completed"},"tags":[]}},{"cell_type":"code","source":"tdcsfog_003f117e14.info()","metadata":{"papermill":{"duration":0.048458,"end_time":"2023-05-31T09:27:59.226487","exception":false,"start_time":"2023-05-31T09:27:59.178029","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.597533Z","iopub.execute_input":"2025-05-20T10:01:14.597960Z","iopub.status.idle":"2025-05-20T10:01:14.648866Z","shell.execute_reply.started":"2025-05-20T10:01:14.597929Z","shell.execute_reply":"2025-05-20T10:01:14.646893Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here, we **normalize \"Time\"**.","metadata":{"papermill":{"duration":0.022312,"end_time":"2023-05-31T09:27:59.271301","exception":false,"start_time":"2023-05-31T09:27:59.248989","status":"completed"},"tags":[]}},{"cell_type":"code","source":"tdcsfog_003f117e14.Time = tdcsfog_003f117e14.Time / (len(tdcsfog_003f117e14) -1)\ntdcsfog_003f117e14","metadata":{"papermill":{"duration":0.043935,"end_time":"2023-05-31T09:27:59.337588","exception":false,"start_time":"2023-05-31T09:27:59.293653","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.651609Z","iopub.execute_input":"2025-05-20T10:01:14.652203Z","iopub.status.idle":"2025-05-20T10:01:14.696077Z","shell.execute_reply.started":"2025-05-20T10:01:14.652157Z","shell.execute_reply":"2025-05-20T10:01:14.693666Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In addition, **tdcsfog_metadata.csv identifies** each series in the tdcsfog dataset by **a unique Subject, Visit, Test, and Medication condition**.","metadata":{"papermill":{"duration":0.022638,"end_time":"2023-05-31T09:27:59.383135","exception":false,"start_time":"2023-05-31T09:27:59.360497","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# tdcsfog metadata file\ntdcsfog_metadata = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/defog_metadata.csv\")\ntdcsfog_metadata.head(5)","metadata":{"papermill":{"duration":0.042975,"end_time":"2023-05-31T09:27:59.448954","exception":false,"start_time":"2023-05-31T09:27:59.405979","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.698031Z","iopub.execute_input":"2025-05-20T10:01:14.698481Z","iopub.status.idle":"2025-05-20T10:01:14.719704Z","shell.execute_reply.started":"2025-05-20T10:01:14.698439Z","shell.execute_reply":"2025-05-20T10:01:14.717228Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This table appears to be a dataset containing information about patients, including their **unique ID (\"Id\"), subject ID (\"Subject\"), visit number (\"Visit\"), and medication status (\"Medication\")**.\n\nThe **\"Id\"** column likely represents **a unique identifier assigned to each experiment** in the dataset. The \"Subject\" column represents **a unique identifier assigned to each patient**. The \"Visit\" column may represent a particular visit or assessment of the patient, and could be used to track changes in medication status or other variables over time.\n\nThe \"Medication\" column indicates **whether or not the patient is taking medication at the time of the visit**, and could potentially be used as a predictor variable in a machine learning model to predict patient outcomes or response to treatment. However, without additional information about the context and purpose of this dataset, it is difficult to draw any further conclusions.\n\n**Next, let's see the test and sample submission files.**\n\n### Do Not Run This Cell in Case of Submission!","metadata":{"papermill":{"duration":0.022749,"end_time":"2023-05-31T09:27:59.495311","exception":false,"start_time":"2023-05-31T09:27:59.472562","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# test file for tdcsfog\n#tdcsfog_003f117e14 = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/tdcsfog/003f117e14.csv')\n#tdcsfog_003f117e14","metadata":{"papermill":{"duration":0.031118,"end_time":"2023-05-31T09:27:59.549312","exception":false,"start_time":"2023-05-31T09:27:59.518194","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.721144Z","iopub.execute_input":"2025-05-20T10:01:14.721514Z","iopub.status.idle":"2025-05-20T10:01:14.736552Z","shell.execute_reply.started":"2025-05-20T10:01:14.721485Z","shell.execute_reply":"2025-05-20T10:01:14.734449Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This table appears to be a time series dataset containing motion sensor data, possibly from a wearable device, for a single subject. Each row represents a particular moment in time (as indicated by the **\"Time\"** column) and contains measurements of acceleration in three dimensions (**\"AccV\", \"AccML\", and \"AccAP\"**).","metadata":{"papermill":{"duration":0.022644,"end_time":"2023-05-31T09:27:59.594841","exception":false,"start_time":"2023-05-31T09:27:59.572197","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# sample submission file\nsample_submission = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/sample_submission.csv')\nsample_submission","metadata":{"papermill":{"duration":0.347526,"end_time":"2023-05-31T09:27:59.965210","exception":false,"start_time":"2023-05-31T09:27:59.617684","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:14.738154Z","iopub.execute_input":"2025-05-20T10:01:14.738639Z","iopub.status.idle":"2025-05-20T10:01:15.105881Z","shell.execute_reply.started":"2025-05-20T10:01:14.738575Z","shell.execute_reply":"2025-05-20T10:01:15.104668Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This table appears to contain information as to **\"StartHesitation\", \"Turn\", and \"Walking.\"**\n\nTherefore, we have to make a model to **predict  \"StartHesitation\", \"Turn\", and \"Walking\" from  \"Time\",  \"AccV\", \"AccML\", and \"AccAP\"** in this competition. Thus, at the moment, we do **not use tdcsfog_metadata**.\n\n**Next, let's see one file in the defog folder.**","metadata":{"papermill":{"duration":0.023118,"end_time":"2023-05-31T09:28:00.011888","exception":false,"start_time":"2023-05-31T09:27:59.988770","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# one defog file\ndefog_02ea782681 = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/defog/02ea782681.csv')\ndefog_02ea782681","metadata":{"papermill":{"duration":0.315265,"end_time":"2023-05-31T09:28:00.350763","exception":false,"start_time":"2023-05-31T09:28:00.035498","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:15.107276Z","iopub.execute_input":"2025-05-20T10:01:15.107734Z","iopub.status.idle":"2025-05-20T10:01:15.386265Z","shell.execute_reply.started":"2025-05-20T10:01:15.107600Z","shell.execute_reply":"2025-05-20T10:01:15.384847Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here, we **normalize \"Time\"**.","metadata":{"papermill":{"duration":0.023283,"end_time":"2023-05-31T09:28:00.398379","exception":false,"start_time":"2023-05-31T09:28:00.375096","status":"completed"},"tags":[]}},{"cell_type":"code","source":"defog_02ea782681.Time = defog_02ea782681.Time / (len(defog_02ea782681) - 1)\ndefog_02ea782681","metadata":{"papermill":{"duration":0.050272,"end_time":"2023-05-31T09:28:00.472227","exception":false,"start_time":"2023-05-31T09:28:00.421955","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:15.387864Z","iopub.execute_input":"2025-05-20T10:01:15.388232Z","iopub.status.idle":"2025-05-20T10:01:15.411020Z","shell.execute_reply.started":"2025-05-20T10:01:15.388203Z","shell.execute_reply":"2025-05-20T10:01:15.409712Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This table appears to contain sensor data and annotations related to a task performed by a subject. The columns **\"Time\", \"AccV\", \"AccML\", and \"AccAP\" likely contain motion sensor data similar to the previous examples**.\n\nThe columns \"StartHesitation\", \"Turn\", and \"Walking\" appear to correspond to whether the subject is performing certain activities at each time point, similar to the previous example. However, in this case, there are **additional columns \"Valid\" and \"Task\" that provide information about the validity of the annotations and the task performed**.\n\nThe **\"Valid\"** column appears to indicate whether the annotations for a given time point are considered to be **reliable or not**. The **\"Task\"** column appears to indicate whether the task was performed during a given time point. Portions of the data marked **\"False\"** in this column should be considered **unannotated and not used in analysis**.\n\nIt's possible that this dataset was collected in the context of studying movement disorders or other neurological conditions, where accurate annotation of sensor data is important for clinical diagnosis and treatment planning. **This dataset could potentially be used to develop machine learning models to predict task performance or identify patterns of sensor data associated with certain movements or behaviors.** However, given the additional complexity introduced by the \"Valid\" and \"Task\" columns, **careful attention** would need to be paid to cleaning and preprocessing the data before using it for analysis or modeling.\n\n**Next, let's see its test file.**\n\n### Do Not Run This Cell in Case of Submission!","metadata":{"papermill":{"duration":0.023691,"end_time":"2023-05-31T09:28:00.520251","exception":false,"start_time":"2023-05-31T09:28:00.496560","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# test file for defog\n#defog_02ab235146 = pd.read_csv('/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/defog/02ab235146.csv')\n#defog_02ab235146","metadata":{"papermill":{"duration":0.032622,"end_time":"2023-05-31T09:28:00.576953","exception":false,"start_time":"2023-05-31T09:28:00.544331","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:15.412549Z","iopub.execute_input":"2025-05-20T10:01:15.413491Z","iopub.status.idle":"2025-05-20T10:01:15.420527Z","shell.execute_reply.started":"2025-05-20T10:01:15.413445Z","shell.execute_reply":"2025-05-20T10:01:15.418463Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This table appears to be a time series dataset containing motion sensor data, possibly from a wearable device, for a single subject. Each row represents a particular moment in time (as indicated by the **\"Time\"** column) and contains measurements of acceleration in three dimensions (**\"AccV\", \"AccML\", and \"AccAP\"**).\n\nTherefore, we have to make a model to **predict  \"StartHesitation\", \"Turn\", and \"Walking\" from  \"Time\",  \"AccV\", \"AccML\", and \"AccAP\"** in this competition, **which is exactly the same as the tdcsfog folder**.\n\n**The tdcsfog folder includes more than 800 csv files and accounts for the majority of the information contained in the whole dataset for this competition.** Therefore, **the training files in the tdcsfog folder might be sufficient to create the prediction model to carry out the tasks**.\n\nLet's **take all the csv files in the tdcsfog folder and combine them into one dataset**.","metadata":{"papermill":{"duration":0.024211,"end_time":"2023-05-31T09:28:00.627192","exception":false,"start_time":"2023-05-31T09:28:00.602981","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Take All the CSV Files in the Train tdcsfog Folder","metadata":{"papermill":{"duration":0.023696,"end_time":"2023-05-31T09:28:00.675122","exception":false,"start_time":"2023-05-31T09:28:00.651426","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ntdcsfog_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/tdcsfog'\n\n# Initialize an empty list to store the dataframes.\ntdcsfog_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(tdcsfog_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(tdcsfog_path, file_name)\n        file = pd.read_csv(file_path)\n        file.Time = file.Time / (len(file) - 1)\n        tdcsfog_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ntdcsfog = pd.concat(tdcsfog_list, axis = 0)\n\n# Show the concatenated dataframe.\ntdcsfog","metadata":{"papermill":{"duration":17.996849,"end_time":"2023-05-31T09:28:18.695906","exception":false,"start_time":"2023-05-31T09:28:00.699057","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:15.422837Z","iopub.execute_input":"2025-05-20T10:01:15.423227Z","iopub.status.idle":"2025-05-20T10:01:33.964897Z","shell.execute_reply.started":"2025-05-20T10:01:15.423188Z","shell.execute_reply":"2025-05-20T10:01:33.963783Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It is better to reduce the memory usage. Reference: [Reducing DataFrame memory size by ~65%](https://www.kaggle.com/code/arjanso/reducing-dataframe-memory-size-by-65)","metadata":{"papermill":{"duration":0.024577,"end_time":"2023-05-31T09:28:18.745302","exception":false,"start_time":"2023-05-31T09:28:18.720725","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def reduce_memory_usage(df):\n    \n    start_mem = df.memory_usage().sum() / 1024 ** 2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n    \n    for col in df.columns:\n        col_type = df[col].dtype.name\n        if ((col_type != 'datetime64[ns]') & (col_type != 'category')):\n            if (col_type != 'object'):\n                c_min = df[col].min()\n                c_max = df[col].max()\n\n                if str(col_type)[:3] == 'int':\n                    if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                        df[col] = df[col].astype(np.int8)\n                    elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                        df[col] = df[col].astype(np.int16)\n                    elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                        df[col] = df[col].astype(np.int32)\n                    elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                        df[col] = df[col].astype(np.int64)\n\n                else:\n                    if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                        df[col] = df[col].astype(np.float16)\n                    elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                        df[col] = df[col].astype(np.float32)\n                    else:\n                        pass\n            else:\n                df[col] = df[col].astype('category')\n    mem_usg = df.memory_usage().sum() / 1024 ** 2 \n    print(\"Memory usage became: \",mem_usg,\" MB\")\n    \n    return df","metadata":{"papermill":{"duration":0.043675,"end_time":"2023-05-31T09:28:18.813661","exception":false,"start_time":"2023-05-31T09:28:18.769986","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:33.966227Z","iopub.execute_input":"2025-05-20T10:01:33.966514Z","iopub.status.idle":"2025-05-20T10:01:33.978960Z","shell.execute_reply.started":"2025-05-20T10:01:33.966490Z","shell.execute_reply":"2025-05-20T10:01:33.977635Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tdcsfog = reduce_memory_usage(tdcsfog)","metadata":{"papermill":{"duration":0.592113,"end_time":"2023-05-31T09:28:19.430649","exception":false,"start_time":"2023-05-31T09:28:18.838536","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:33.980434Z","iopub.execute_input":"2025-05-20T10:01:33.980883Z","iopub.status.idle":"2025-05-20T10:01:34.594246Z","shell.execute_reply.started":"2025-05-20T10:01:33.980842Z","shell.execute_reply":"2025-05-20T10:01:34.593067Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tdcsfog.describe()","metadata":{"papermill":{"duration":3.387785,"end_time":"2023-05-31T09:28:22.843322","exception":false,"start_time":"2023-05-31T09:28:19.455537","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:34.595549Z","iopub.execute_input":"2025-05-20T10:01:34.595929Z","iopub.status.idle":"2025-05-20T10:01:38.134395Z","shell.execute_reply.started":"2025-05-20T10:01:34.595900Z","shell.execute_reply":"2025-05-20T10:01:38.133242Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tdcsfog['StartHesitation'].mean()","metadata":{"papermill":{"duration":0.04158,"end_time":"2023-05-31T09:28:22.910624","exception":false,"start_time":"2023-05-31T09:28:22.869044","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:38.136501Z","iopub.execute_input":"2025-05-20T10:01:38.137004Z","iopub.status.idle":"2025-05-20T10:01:38.151625Z","shell.execute_reply.started":"2025-05-20T10:01:38.136962Z","shell.execute_reply":"2025-05-20T10:01:38.149510Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tdcsfog['Turn'].mean()","metadata":{"papermill":{"duration":0.042566,"end_time":"2023-05-31T09:28:22.978553","exception":false,"start_time":"2023-05-31T09:28:22.935987","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:38.154008Z","iopub.execute_input":"2025-05-20T10:01:38.154395Z","iopub.status.idle":"2025-05-20T10:01:38.181796Z","shell.execute_reply.started":"2025-05-20T10:01:38.154363Z","shell.execute_reply":"2025-05-20T10:01:38.179939Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tdcsfog['Walking'].mean()","metadata":{"papermill":{"duration":0.041547,"end_time":"2023-05-31T09:28:23.046791","exception":false,"start_time":"2023-05-31T09:28:23.005244","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:38.184355Z","iopub.execute_input":"2025-05-20T10:01:38.184883Z","iopub.status.idle":"2025-05-20T10:01:38.203373Z","shell.execute_reply.started":"2025-05-20T10:01:38.184839Z","shell.execute_reply":"2025-05-20T10:01:38.201936Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(tdcsfog)","metadata":{"papermill":{"duration":0.036599,"end_time":"2023-05-31T09:28:23.109693","exception":false,"start_time":"2023-05-31T09:28:23.073094","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:38.209922Z","iopub.execute_input":"2025-05-20T10:01:38.210355Z","iopub.status.idle":"2025-05-20T10:01:38.221297Z","shell.execute_reply.started":"2025-05-20T10:01:38.210324Z","shell.execute_reply":"2025-05-20T10:01:38.219275Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Show the numbers by graph.\ndata = pd.DataFrame(\n    np.concatenate([\n        ['Total'] * len(tdcsfog),\n        ['StartHesitation'] * int(np.ceil(len(tdcsfog) / 2 * tdcsfog['StartHesitation'].mean())),\n        ['Turn'] * int(np.ceil(len(tdcsfog) / 2 * tdcsfog['Turn'].mean())),\n        ['Walking'] * int(np.ceil(len(tdcsfog) / 2 * tdcsfog['Walking'].mean()))\n    ]),\n    columns = [\"The Number of 1\"]\n)\n\nsns.countplot(x = 'The Number of 1', data = data)","metadata":{"papermill":{"duration":13.72787,"end_time":"2023-05-31T09:28:36.863404","exception":false,"start_time":"2023-05-31T09:28:23.135534","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:38.223095Z","iopub.execute_input":"2025-05-20T10:01:38.223447Z","iopub.status.idle":"2025-05-20T10:01:47.119935Z","shell.execute_reply.started":"2025-05-20T10:01:38.223418Z","shell.execute_reply":"2025-05-20T10:01:47.117368Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Most of the target variables are 0.**","metadata":{"papermill":{"duration":0.026102,"end_time":"2023-05-31T09:28:36.915774","exception":false,"start_time":"2023-05-31T09:28:36.889672","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Visualize \"AccV\", \"AccML\", and \"AccAP.\"\nsns.pairplot(tdcsfog[['AccV', 'AccML', 'AccAP']])","metadata":{"papermill":{"duration":207.424293,"end_time":"2023-05-31T09:32:04.366293","exception":false,"start_time":"2023-05-31T09:28:36.942000","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:01:47.123229Z","iopub.execute_input":"2025-05-20T10:01:47.123665Z","iopub.status.idle":"2025-05-20T10:05:07.965338Z","shell.execute_reply.started":"2025-05-20T10:01:47.123623Z","shell.execute_reply":"2025-05-20T10:05:07.964191Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# The Relationship between Time and the 3 Events\n\nWe should analyze **to what extent the \"Time\" is relevant to occurrence of \"StartHesitation\", \"Turn\", and \"Walking,\" respectively**.  Thus, we make a figure to show the relationships between \"Time\" (x-axis) and \"StartHesitation\", \"Turn\", and \"Walking\" (y-axis).","metadata":{"papermill":{"duration":0.029415,"end_time":"2023-05-31T09:32:04.425239","exception":false,"start_time":"2023-05-31T09:32:04.395824","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Create a figure and axis object.\nfig, ax = plt.subplots(figsize = (10, 6))\n\n# Plot the StartHesitation, Turn, and Walking columns against Time.\nax.plot(tdcsfog['Time'], tdcsfog['StartHesitation'], label = 'StartHesitation')\nax.plot(tdcsfog['Time'], tdcsfog['Turn'], label = 'Turn')\nax.plot(tdcsfog['Time'], tdcsfog['Walking'], label = 'Walking')\n\n# Add axis labels and title.\nax.set_xlabel('Time')\nax.set_ylabel('Binary Status')\nax.set_title('Relationship between Time and Movement Status')\n\n# Add a legend to the plot.\nax.legend()\n\n# Show the plot.\nplt.show()","metadata":{"papermill":{"duration":71.519416,"end_time":"2023-05-31T09:33:15.974028","exception":false,"start_time":"2023-05-31T09:32:04.454612","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:05:07.966649Z","iopub.execute_input":"2025-05-20T10:05:07.967008Z","iopub.status.idle":"2025-05-20T10:06:07.823496Z","shell.execute_reply.started":"2025-05-20T10:05:07.966979Z","shell.execute_reply":"2025-05-20T10:06:07.822259Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Unfortunately, this figure above does not clearly show the relationships.\n\nWe create a line plot by **taking the means of the three values by time**. This code groups the tdcsfog dataframe by the \"Time\" column, takes the mean of the other columns for each group, and then resets the index to a sequential integer index. This code creates a figure with the mean values of \"StartHesitation\", \"Turn\", and \"Walking\" on the y-axis and \"Time\" on the x-axis. The rest of the code is the same as before.","metadata":{"papermill":{"duration":0.02984,"end_time":"2023-05-31T09:33:16.034268","exception":false,"start_time":"2023-05-31T09:33:16.004428","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the figure size.\nplt.figure(figsize = (12,6))\n\n# Calculates the mean values of all other columns for each unique value of Time.\n# Reset the index of the resulting dataframe.\ntdcsfog_means = tdcsfog.groupby('Time').mean().reset_index()\n\n# Plot the mean StartHesitation values over time.\nplt.plot(tdcsfog_means['Time'], tdcsfog_means['StartHesitation'], label = 'StartHesitation')\n\n# Plot the mean Turn values over time.\nplt.plot(tdcsfog_means['Time'], tdcsfog_means['Turn'], label = 'Turn')\n\n# Plot the mean Walking values over time.\nplt.plot(tdcsfog_means['Time'], tdcsfog_means['Walking'], label = 'Walking')\n\n# Add a legend to the plot.\nplt.legend()\n\n# Set the x-label of the plot.\nplt.xlabel('Time')\n\n# Set the y-label of the plot.\nplt.ylabel('Mean Value')\n\n# Set the title of the plot.\nplt.title('Mean Values of StartHesitation, Turn, and Walking over Time')\n\n# Display the plot.\nplt.show()","metadata":{"papermill":{"duration":0.894101,"end_time":"2023-05-31T09:33:16.958718","exception":false,"start_time":"2023-05-31T09:33:16.064617","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:07.825122Z","iopub.execute_input":"2025-05-20T10:06:07.825513Z","iopub.status.idle":"2025-05-20T10:06:08.666506Z","shell.execute_reply.started":"2025-05-20T10:06:07.825478Z","shell.execute_reply":"2025-05-20T10:06:08.664956Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There may be some relationship between \"Time\" and the 3 Events. Particularly, **\"StartHesitation\" was more likely to occur at specific times**. **Positive correlation may exist between \"Time\" and \"Walking.\"** However, **the relationship is a bit irregular**. It is not clear whether a prediction model had better include \"Time\" as an independent variable.\n\nFor the moment, we create **a model with \"AccV\", \"AccML\", and \"AccAP,\" including \"Time\" to predict \"StartHesitation\", \"Turn\", and \"Walking,\" respectively**.","metadata":{"papermill":{"duration":0.031286,"end_time":"2023-05-31T09:33:17.021709","exception":false,"start_time":"2023-05-31T09:33:16.990423","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Take All the CSV Files in the Train defog Folder","metadata":{"papermill":{"duration":0.031123,"end_time":"2023-05-31T09:33:17.084621","exception":false,"start_time":"2023-05-31T09:33:17.053498","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ndefog_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/defog'\n\n# Initialize an empty list to store the dataframes.\ndefog_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(defog_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(defog_path, file_name)\n        file = pd.read_csv(file_path)\n        file.Time = file.Time / (len(file) - 1)\n        defog_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ndefog = pd.concat(defog_list, axis = 0)\n\n# Show the concatenated dataframe.\ndefog","metadata":{"papermill":{"duration":24.680545,"end_time":"2023-05-31T09:33:41.796686","exception":false,"start_time":"2023-05-31T09:33:17.116141","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:08.668664Z","iopub.execute_input":"2025-05-20T10:06:08.669131Z","iopub.status.idle":"2025-05-20T10:06:31.298058Z","shell.execute_reply.started":"2025-05-20T10:06:08.669090Z","shell.execute_reply":"2025-05-20T10:06:31.296907Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"defog = reduce_memory_usage(defog)","metadata":{"papermill":{"duration":1.213096,"end_time":"2023-05-31T09:33:43.041621","exception":false,"start_time":"2023-05-31T09:33:41.828525","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:31.299397Z","iopub.execute_input":"2025-05-20T10:06:31.299771Z","iopub.status.idle":"2025-05-20T10:06:32.590151Z","shell.execute_reply.started":"2025-05-20T10:06:31.299742Z","shell.execute_reply":"2025-05-20T10:06:32.588499Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**We are going to use valid data only.**","metadata":{"papermill":{"duration":0.031648,"end_time":"2023-05-31T09:33:43.105167","exception":false,"start_time":"2023-05-31T09:33:43.073519","status":"completed"},"tags":[]}},{"cell_type":"code","source":"defog = defog[(defog['Task'] == 1) & (defog['Valid'] == 1)]","metadata":{"papermill":{"duration":0.652616,"end_time":"2023-05-31T09:33:43.789828","exception":false,"start_time":"2023-05-31T09:33:43.137212","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:32.591394Z","iopub.execute_input":"2025-05-20T10:06:32.591774Z","iopub.status.idle":"2025-05-20T10:06:33.238464Z","shell.execute_reply.started":"2025-05-20T10:06:32.591744Z","shell.execute_reply":"2025-05-20T10:06:33.236472Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"defog = defog.iloc[:, :7]","metadata":{"papermill":{"duration":0.0794,"end_time":"2023-05-31T09:33:43.902287","exception":false,"start_time":"2023-05-31T09:33:43.822887","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:33.240244Z","iopub.execute_input":"2025-05-20T10:06:33.241212Z","iopub.status.idle":"2025-05-20T10:06:33.291517Z","shell.execute_reply.started":"2025-05-20T10:06:33.241170Z","shell.execute_reply":"2025-05-20T10:06:33.289082Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"defog.describe()","metadata":{"papermill":{"duration":2.00688,"end_time":"2023-05-31T09:33:45.941278","exception":false,"start_time":"2023-05-31T09:33:43.934398","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:33.293753Z","iopub.execute_input":"2025-05-20T10:06:33.294216Z","iopub.status.idle":"2025-05-20T10:06:35.346399Z","shell.execute_reply.started":"2025-05-20T10:06:33.294174Z","shell.execute_reply":"2025-05-20T10:06:35.345266Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**We merge tdcsfog and defog datasets into one merged dataset.**","metadata":{"papermill":{"duration":0.032544,"end_time":"2023-05-31T09:33:46.006198","exception":false,"start_time":"2023-05-31T09:33:45.973654","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Concatenate the dataframes vertically using pd.concat().\nmerged = pd.concat([tdcsfog, defog], axis = 0)\n\n# Show the concatenated dataframe.\nmerged","metadata":{"papermill":{"duration":0.12561,"end_time":"2023-05-31T09:33:46.164044","exception":false,"start_time":"2023-05-31T09:33:46.038434","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:35.347730Z","iopub.execute_input":"2025-05-20T10:06:35.348063Z","iopub.status.idle":"2025-05-20T10:06:35.444167Z","shell.execute_reply.started":"2025-05-20T10:06:35.348037Z","shell.execute_reply":"2025-05-20T10:06:35.443041Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create Dataset\n\nFirst, we need to **split the data into input features (i.e. \"Time\", \"AccV\", \"AccML\", and \"AccAP\") and target variables (i.e. \"StartHesitation\", \"Turn\", and \"Walking\")**. We can do this using the .iloc method to select the appropriate columns.","metadata":{"papermill":{"duration":0.032463,"end_time":"2023-05-31T09:33:46.229212","exception":false,"start_time":"2023-05-31T09:33:46.196749","status":"completed"},"tags":[]}},{"cell_type":"code","source":"X_merged = merged.iloc[:, 0:4]  # input features\ny1 = merged['StartHesitation']  # target variable for StartHesitation\ny2 = merged['Turn']  # target variable for Turn\ny3 = merged['Walking']  # target variable for Walking","metadata":{"papermill":{"duration":0.207904,"end_time":"2023-05-31T09:33:46.469640","exception":false,"start_time":"2023-05-31T09:33:46.261736","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:35.445392Z","iopub.execute_input":"2025-05-20T10:06:35.445733Z","iopub.status.idle":"2025-05-20T10:06:35.631990Z","shell.execute_reply.started":"2025-05-20T10:06:35.445706Z","shell.execute_reply":"2025-05-20T10:06:35.630138Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Most of the target variables are 0. We had better **create each balanced dataset with the target variables of 0 and 1 equally**.","metadata":{"papermill":{"duration":0.032717,"end_time":"2023-05-31T09:33:46.535339","exception":false,"start_time":"2023-05-31T09:33:46.502622","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Find the positions of y1 where it equals 0.\ny1_zeros = np.where(y1 == 0)[0]\ny1_ones = np.where(y1 == 1)[0]\n\n# Choose the same number of samples with y1 == 1 as there are with y1 == 0.\nnum1_ones = (y1 == 1).sum()\nnp.random.seed(42)\ny1_zeros = np.random.choice(np.where(y1 == 0)[0], size = num1_ones, replace = False)\n\n# Combine the positions of y1 == 0 and y1 == 1.\ny1_balanced_idxs = np.sort(np.concatenate([y1_zeros, y1_ones]))\n\n# Use the balanced indices to get the corresponding rows of X and y1.\nX1_balanced = X_merged.iloc[y1_balanced_idxs, :]\ny1_balanced = y1.iloc[y1_balanced_idxs]","metadata":{"papermill":{"duration":0.727806,"end_time":"2023-05-31T09:33:47.296034","exception":false,"start_time":"2023-05-31T09:33:46.568228","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:35.633695Z","iopub.execute_input":"2025-05-20T10:06:35.634062Z","iopub.status.idle":"2025-05-20T10:06:36.425583Z","shell.execute_reply.started":"2025-05-20T10:06:35.634033Z","shell.execute_reply":"2025-05-20T10:06:36.424234Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Find the positions of y2 where it equals 0.\ny2_zeros = np.where(y2 == 0)[0]\ny2_ones = np.where(y2 == 1)[0]\n\n# Choose the same number of samples with y2 == 1 as there are with y2 == 0.\nnum2_ones = (y2 == 1).sum()\nnp.random.seed(42)\ny2_zeros = np.random.choice(np.where(y2 == 0)[0], size = num2_ones, replace = False)\n\n# Combine the positions of y2 == 0 and y2 == 1.\ny2_balanced_idxs = np.sort(np.concatenate([y2_zeros, y2_ones]))\n\n# Use the balanced indices to get the corresponding rows of X and y1.\nX2_balanced = X_merged.iloc[y2_balanced_idxs, :]\ny2_balanced = y2.iloc[y2_balanced_idxs]","metadata":{"papermill":{"duration":1.286369,"end_time":"2023-05-31T09:33:48.616849","exception":false,"start_time":"2023-05-31T09:33:47.330480","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:36.426838Z","iopub.execute_input":"2025-05-20T10:06:36.427262Z","iopub.status.idle":"2025-05-20T10:06:37.861471Z","shell.execute_reply.started":"2025-05-20T10:06:36.427223Z","shell.execute_reply":"2025-05-20T10:06:37.860179Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Find the positions of y3 where it equals 0.\ny3_zeros = np.where(y3 == 0)[0]\ny3_ones = np.where(y3 == 1)[0]\n\n# Choose the same number of samples with y3 == 1 as there are with y3 == 0.\nnum3_ones = (y3 == 1).sum()\nnp.random.seed(42)\ny3_zeros = np.random.choice(np.where(y3 == 0)[0], size = num3_ones, replace = False)\n\n# Combine the positions of y3 == 0 and y3 == 1.\ny3_balanced_idxs = np.sort(np.concatenate([y3_zeros, y3_ones]))\n\n# Use the balanced indices to get the corresponding rows of X and y3.\nX3_balanced = X_merged.iloc[y3_balanced_idxs, :]\ny3_balanced = y3.iloc[y3_balanced_idxs]","metadata":{"papermill":{"duration":0.76629,"end_time":"2023-05-31T09:33:49.416428","exception":false,"start_time":"2023-05-31T09:33:48.650138","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:37.863075Z","iopub.execute_input":"2025-05-20T10:06:37.863407Z","iopub.status.idle":"2025-05-20T10:06:38.683536Z","shell.execute_reply.started":"2025-05-20T10:06:37.863379Z","shell.execute_reply":"2025-05-20T10:06:38.682294Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Next, we can **split the data into training and testing sets using the train_test_split function from scikit-learn**.","metadata":{"papermill":{"duration":0.03229,"end_time":"2023-05-31T09:33:49.481666","exception":false,"start_time":"2023-05-31T09:33:49.449376","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX1_train, X1_test, y1_train, y1_test = train_test_split(X1_balanced, y1_balanced, test_size = 0.2, random_state = 42)\nX2_train, X2_test, y2_train, y2_test = train_test_split(X2_balanced, y2_balanced, test_size = 0.2, random_state = 42)\nX3_train, X3_test, y3_train, y3_test = train_test_split(X3_balanced, y3_balanced, test_size = 0.2, random_state = 42)","metadata":{"papermill":{"duration":1.351479,"end_time":"2023-05-31T09:33:50.865887","exception":false,"start_time":"2023-05-31T09:33:49.514408","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:38.684882Z","iopub.execute_input":"2025-05-20T10:06:38.685208Z","iopub.status.idle":"2025-05-20T10:06:40.466823Z","shell.execute_reply.started":"2025-05-20T10:06:38.685181Z","shell.execute_reply":"2025-05-20T10:06:40.464986Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Then, we **standardize the independent variables**.","metadata":{"papermill":{"duration":0.032626,"end_time":"2023-05-31T09:33:50.931767","exception":false,"start_time":"2023-05-31T09:33:50.899141","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\n# Standardize the independent variables.\nscaler1 = StandardScaler()\nX1_train = scaler1.fit_transform(X1_train)\nX1_test = scaler1.transform(X1_test)\n\nscaler2 = StandardScaler()\nX2_train = scaler2.fit_transform(X2_train)\nX2_test = scaler2.transform(X2_test)\n\nscaler3 = StandardScaler()\nX3_train = scaler3.fit_transform(X3_train)\nX3_test = scaler3.transform(X3_test)","metadata":{"papermill":{"duration":1.158502,"end_time":"2023-05-31T09:33:52.122963","exception":false,"start_time":"2023-05-31T09:33:50.964461","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:40.468561Z","iopub.execute_input":"2025-05-20T10:06:40.469025Z","iopub.status.idle":"2025-05-20T10:06:41.629004Z","shell.execute_reply.started":"2025-05-20T10:06:40.468992Z","shell.execute_reply":"2025-05-20T10:06:41.627339Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create, Train, and Evaluate Model\n\nFinally, we can **create and train three separate models**, one for each target variable, using a suitable algorithm.\n       \n### This time we use **Random Forest Regressor instead of the Logistic Regression model**.\n\n**For a Logistic Regression model, please see [PD FOG Prediction Baseline by Logistic Regression](https://www.kaggle.com/code/gokifujiya/pd-fog-prediction-baseline-by-logistic-regression).**\n\nGridSearchCV is used to perform **a grid search with cross-validation to find the best hyperparameters for each Random Forest Regressor model**. The param_grid dictionary defines the hyperparameter values to be explored. The cv parameter specifies **the number of cross-validation folds** to use during the grid search.\n\n### However, it may lead to time out to carry out a grid search sufficiently on this platform.","metadata":{"papermill":{"duration":0.033122,"end_time":"2023-05-31T09:33:52.189128","exception":false,"start_time":"2023-05-31T09:33:52.156006","status":"completed"},"tags":[]}},{"cell_type":"code","source":"#from sklearn.linear_model import LogisticRegression\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.ensemble import RandomForestRegressor\n\n# Create three separate logistic regression models.\n#model1 = LogisticRegression()\n#model2 = LogisticRegression()\n#model3 = LogisticRegression()\n\n# Define the parameter grid for hyperparameter tuning.\nparam_grid = {\n    'n_estimators': [200, 300],\n    'max_depth': [8],\n    'n_jobs': [-1],\n    'random_state': [42]\n}\n\n# Create three separate Random Forest Regressor models.\nmodel1 = GridSearchCV(RandomForestRegressor(), param_grid, cv = 5)\nmodel2 = GridSearchCV(RandomForestRegressor(), param_grid, cv = 5)\nmodel3 = GridSearchCV(RandomForestRegressor(), param_grid, cv = 5)\n\n# Train the models on the training data.\nmodel1.fit(X1_train, y1_train)\nmodel2.fit(X2_train, y2_train)\nmodel3.fit(X3_train, y3_train)\n\n# Evaluate the models on the test data.\nprint('R2 for StartHesitation:', model1.best_score_)\nprint('R2 for Turn:', model2.best_score_)\nprint('R2 for Walking:', model3.best_score_)\n\n# Print the best parameters for each model\nprint('Best parameters for StartHesitation:', model1.best_params_)\nprint('Best parameters for Turn:', model2.best_params_)\nprint('Best parameters for Walking:', model3.best_params_)","metadata":{"papermill":{"duration":11383.773866,"end_time":"2023-05-31T12:43:35.996399","exception":false,"start_time":"2023-05-31T09:33:52.222533","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T10:06:41.630768Z","iopub.execute_input":"2025-05-20T10:06:41.631353Z","iopub.status.idle":"2025-05-20T13:10:17.023632Z","shell.execute_reply.started":"2025-05-20T10:06:41.631304Z","shell.execute_reply":"2025-05-20T13:10:17.021642Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Attention!\n\nPlease note that the grid search was only conducted for the number of estimators: \"'n_estimators': [200, 300],\" and there were **only two options**. Nevertheless, this notebook took as much as around **3 hours** to run.\n\nYou will find grid search extremely time-consuming. Random search is also time-consuming and does **not guarantee the best combination of these hyperparameters**.\n\nWe should also keep in mind that **\"best_params_\" does not necessarily mean the true best combination of these hyperparameters**. It only means **the best combination among the selected combinations** of these hyperparameters **by grid or random search**. Therefore, **there is generally a better combination than \"best_params_,\" which is very important for competition**.","metadata":{}},{"cell_type":"markdown","source":"# Create Test Dataset","metadata":{"papermill":{"duration":0.033117,"end_time":"2023-05-31T12:43:36.065254","exception":false,"start_time":"2023-05-31T12:43:36.032137","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ntdcsfog_test_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/tdcsfog'\n\n# Initialize an empty list to store the dataframes.\ntdcsfog_test_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(tdcsfog_test_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(tdcsfog_test_path, file_name)\n        file = pd.read_csv(file_path)\n        file['Id'] = file_name[:-4] + '_' + file['Time'].apply(str)\n        file.Time = file.Time / (len(file) - 1)\n        tdcsfog_test_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ntdcsfog_test = pd.concat(tdcsfog_test_list, axis = 0)\n\n# Show the concatenated dataframe.\ntdcsfog_test","metadata":{"papermill":{"duration":0.097727,"end_time":"2023-05-31T12:43:36.196496","exception":false,"start_time":"2023-05-31T12:43:36.098769","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:17.027063Z","iopub.execute_input":"2025-05-20T13:10:17.027624Z","iopub.status.idle":"2025-05-20T13:10:17.095253Z","shell.execute_reply.started":"2025-05-20T13:10:17.027544Z","shell.execute_reply":"2025-05-20T13:10:17.093878Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tdcsfog_test = reduce_memory_usage(tdcsfog_test)","metadata":{"papermill":{"duration":0.061764,"end_time":"2023-05-31T12:43:36.292334","exception":false,"start_time":"2023-05-31T12:43:36.230570","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:17.097483Z","iopub.execute_input":"2025-05-20T13:10:17.098754Z","iopub.status.idle":"2025-05-20T13:10:17.140793Z","shell.execute_reply.started":"2025-05-20T13:10:17.098701Z","shell.execute_reply":"2025-05-20T13:10:17.137956Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Set the directory path to the folder containing the CSV files.\ndefog_test_path = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/defog'\n\n# Initialize an empty list to store the dataframes.\ndefog_test_list = []\n\n# Loop through each file in the directory and read it into a dataframe.\nfor file_name in os.listdir(defog_test_path):\n    if file_name.endswith('.csv'):\n        file_path = os.path.join(defog_test_path, file_name)\n        file = pd.read_csv(file_path)\n        file['Id'] = file_name[:-4] + '_' + file['Time'].apply(str)\n        file.Time = file.Time / (len(file) - 1)\n        defog_test_list.append(file)\n\n# Concatenate the dataframes vertically using pd.concat().\ndefog_test = pd.concat(defog_test_list, axis = 0)\n\n# Show the concatenated dataframe.\ndefog_test","metadata":{"papermill":{"duration":0.647709,"end_time":"2023-05-31T12:43:36.974966","exception":false,"start_time":"2023-05-31T12:43:36.327257","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:17.143060Z","iopub.execute_input":"2025-05-20T13:10:17.143481Z","iopub.status.idle":"2025-05-20T13:10:17.788328Z","shell.execute_reply.started":"2025-05-20T13:10:17.143447Z","shell.execute_reply":"2025-05-20T13:10:17.786736Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"defog_test = reduce_memory_usage(defog_test)","metadata":{"papermill":{"duration":0.669049,"end_time":"2023-05-31T12:43:37.680029","exception":false,"start_time":"2023-05-31T12:43:37.010980","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:17.790249Z","iopub.execute_input":"2025-05-20T13:10:17.791056Z","iopub.status.idle":"2025-05-20T13:10:18.223458Z","shell.execute_reply.started":"2025-05-20T13:10:17.791012Z","shell.execute_reply":"2025-05-20T13:10:18.222130Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test = pd.concat([tdcsfog_test, defog_test], axis = 0).reset_index(drop = True)\ntest","metadata":{"papermill":{"duration":0.238634,"end_time":"2023-05-31T12:43:37.953181","exception":false,"start_time":"2023-05-31T12:43:37.714547","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:18.225527Z","iopub.execute_input":"2025-05-20T13:10:18.226016Z","iopub.status.idle":"2025-05-20T13:10:18.375913Z","shell.execute_reply.started":"2025-05-20T13:10:18.225975Z","shell.execute_reply":"2025-05-20T13:10:18.374007Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Inference","metadata":{"papermill":{"duration":0.035043,"end_time":"2023-05-31T12:43:38.023649","exception":false,"start_time":"2023-05-31T12:43:37.988606","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Separate the dataset for the independent variables.\ntest_X = test.iloc[:, 0:4]\n\n# Standardize the independent variables by a new scaler.\nscaler = StandardScaler()\ntest_X = scaler.fit_transform(test_X)\n\n# Get the predictions for the three models on the test data.\npred_y1 = model1.predict(test_X)\npred_y2 = model2.predict(test_X)\npred_y3 = model3.predict(test_X)\n\ntest['StartHesitation'] = pred_y1 # target variable for StartHesitation\ntest['Turn'] = pred_y2 # target variable for Turn\ntest['Walking'] = pred_y3 # target variable for Walking\n\ntest","metadata":{"papermill":{"duration":2.760256,"end_time":"2023-05-31T12:43:40.818519","exception":false,"start_time":"2023-05-31T12:43:38.058263","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:18.378459Z","iopub.execute_input":"2025-05-20T13:10:18.379718Z","iopub.status.idle":"2025-05-20T13:10:21.546077Z","shell.execute_reply.started":"2025-05-20T13:10:18.379655Z","shell.execute_reply":"2025-05-20T13:10:21.544686Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Submission","metadata":{"papermill":{"duration":0.034841,"end_time":"2023-05-31T12:43:40.888329","exception":false,"start_time":"2023-05-31T12:43:40.853488","status":"completed"},"tags":[]}},{"cell_type":"code","source":"submission = test.iloc[:, 4:].fillna(0.0)\nsubmission","metadata":{"papermill":{"duration":0.168213,"end_time":"2023-05-31T12:43:41.091624","exception":false,"start_time":"2023-05-31T12:43:40.923411","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:21.547735Z","iopub.execute_input":"2025-05-20T13:10:21.548149Z","iopub.status.idle":"2025-05-20T13:10:21.635543Z","shell.execute_reply.started":"2025-05-20T13:10:21.548120Z","shell.execute_reply":"2025-05-20T13:10:21.634070Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index = False)","metadata":{"papermill":{"duration":2.550445,"end_time":"2023-05-31T12:43:43.677460","exception":false,"start_time":"2023-05-31T12:43:41.127015","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:21.636710Z","iopub.execute_input":"2025-05-20T13:10:21.637079Z","iopub.status.idle":"2025-05-20T13:10:23.413280Z","shell.execute_reply.started":"2025-05-20T13:10:21.637051Z","shell.execute_reply":"2025-05-20T13:10:23.411838Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Save, Load, and Use Model\n\nTo save the trained Logistic Regression model, you can use the joblib library from the sklearn.externals module. This will save the model to a file in the current working directory. **To load the saved model later**, we can use the joblib.load() function.","metadata":{"papermill":{"duration":0.034656,"end_time":"2023-05-31T12:43:43.747787","exception":false,"start_time":"2023-05-31T12:43:43.713131","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import joblib\n\n# Save the model to disk.\njoblib.dump(model1, 'model1.joblib')\njoblib.dump(model2, 'model2.joblib')\njoblib.dump(model3, 'model3.joblib')\n\n# Load the saved models from disk.\nmodel1_loaded = joblib.load('model1.joblib')\nmodel2_loaded = joblib.load('model2.joblib')\nmodel3_loaded = joblib.load('model3.joblib')\n\n# Use the loaded models to make predictions on test data.\ny1_pred_loaded = model1_loaded.predict(test_X)\ny2_pred_loaded = model2_loaded.predict(test_X)\ny3_pred_loaded = model3_loaded.predict(test_X)","metadata":{"papermill":{"duration":3.462789,"end_time":"2023-05-31T12:43:47.334749","exception":false,"start_time":"2023-05-31T12:43:43.871960","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-05-20T13:10:23.414790Z","iopub.execute_input":"2025-05-20T13:10:23.415119Z","iopub.status.idle":"2025-05-20T13:10:26.857196Z","shell.execute_reply.started":"2025-05-20T13:10:23.415091Z","shell.execute_reply":"2025-05-20T13:10:26.855950Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Conclusion\n\nIt is possible that **more features or more advanced machine learning algorithms** could improve the accuracy of the models. Additionally, it may be useful to **investigate other factors** that contribute to the occurrence of freezing gait events, such as cognitive or environmental factors.","metadata":{"papermill":{"duration":0.035057,"end_time":"2023-05-31T12:43:47.406144","exception":false,"start_time":"2023-05-31T12:43:47.371087","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"I am a medical doctor working on **artificial intelligence (AI) for medicine**. At present AI is also widely used in the medical field. Particularly, AI performs in the healthcare sector following tasks: **image classification, object detection, semantic segmentation, GANs, text classification, etc**. **If you are interested in AI for medicine, please see my other notebooks.**","metadata":{"papermill":{"duration":0.035022,"end_time":"2023-05-31T12:43:47.476523","exception":false,"start_time":"2023-05-31T12:43:47.441501","status":"completed"},"tags":[]}}]}