{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# import modules","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import os\nimport glob\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:54.594251Z","iopub.execute_input":"2023-04-17T02:56:54.594670Z","iopub.status.idle":"2023-04-17T02:56:54.628991Z","shell.execute_reply.started":"2023-04-17T02:56:54.594634Z","shell.execute_reply":"2023-04-17T02:56:54.627943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Overview\n\n<p style=\"font-size: 16px;\"><b>Objective</b> : detect the start and stop of freezing of gait and types of FOG events: Start Hesitation, Turn, and Walking.</p>\n\n\n\n<p style=\"font-size: 16px;\">- The competition dataset includes time-series data from 3-D accelerometer on subjects lower-back.</p>\n<p style=\"font-size: 16px;\">- The data series include three dataset, collected under distinct circumstances.</p>\n<p style=\"font-size: 16px; text-indent: 2rem;\">- The <b>tDCS FOG</b> (<code>tdcsfog</code>) : Collected in the lab. Subjects completed a FOG-provoking protcol.</p>\n<p style=\"font-size: 16px; text-indent: 2rem;\">- The <b>DeFOG</b> (<code>defog</code>) : Collected in the subject's home. Subject completed a FOG-provoking protcol.</p>\n<p style=\"font-size: 16px; text-indent: 2rem;\">- The <b>Daily Living</b> (<code>daily</code>) : Comprising one week of continuous 24/7 recordings from **sixty-five** subjects. Forty-five subjects exhibit FOG symptoms (series in the defog dataset), while the other twenty subjects exhibit no FOG symptoms (no series).</p>\n<p style=\"font-size: 16px;\">- The tdcsfog and defog dataset include annotation by expert reviewers.</p>\n<p style=\"font-size: 16px;\">- The daily dataset are *unannotated*. This dataset may be useful for modeling unsupervised or semi-supervised learning.</p>","metadata":{}},{"cell_type":"markdown","source":"# Train data\n\n<ul>\n<li><strong>train/</strong> Folder containing the data series in the training set within three subfolders: <strong>tdcsfog/</strong>, <strong>defog/</strong>, and <strong>notype/</strong>. Series in the <em>notype</em> folder are from the <code>defog</code> dataset but lack event-type annotations. The fields present in these series vary by folder.<ul>\n<li><code>Time</code> An integer timestep. Series from the <code>tdcsfog</code> dataset are recorded at 128Hz (128 timesteps per second), while series from the <code>defog</code> and <code>daily</code> series are recorded at 100Hz (100 timesteps per second).</li>\n<li><code>AccV</code>, <code>AccML</code>, and <code>AccAP</code> Acceleration from a lower-back sensor on three axes: V - vertical, ML - mediolateral, AP - anteroposterior. Data is in units of <em>m/s^2</em> for <code>tdcsfog/</code> and <em>g</em> for <code>defog/</code> and <code>notype/</code>.</li>\n<li><code>StartHesitation</code>, <code>Turn</code>, <code>Walking</code> Indicator variables for the occurrence of each of the event types.</li>\n<li><code>Event</code> Indicator variable for the occurrence of <em>any</em> FOG-type event. Present only in the <strong>notype</strong> series, which lack type-level annotations.</li>\n<li><code>Valid</code> There were cases during the video annotation that were hard for the annotator to decide if there was an Akinetic (i.e., essentially no movement) FoG or the subject stopped voluntarily. Only event annotations where the series is marked <code>true</code> should be considered as unambiguous.</li>\n<li><code>Task</code> Series were only annotated where this value is <code>true</code>. Portions marked <code>false</code> should be considered unannotated.</li></ul></li>\n</ul>","metadata":{}},{"cell_type":"markdown","source":"# Explore tdcsfog train data\n","metadata":{}},{"cell_type":"code","source":"pdir = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/'","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:54.631108Z","iopub.execute_input":"2023-04-17T02:56:54.631836Z","iopub.status.idle":"2023-04-17T02:56:54.636540Z","shell.execute_reply.started":"2023-04-17T02:56:54.631797Z","shell.execute_reply":"2023-04-17T02:56:54.635485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The number of csv files\n\n<ul>\n    <li>There are 833 csv files for 65 subjects.</li><ul>\n    <li>Subjects' acceleration and event series may have been subdivided.</li>\n</ul>","metadata":{}},{"cell_type":"code","source":"tdcsfog_train_files = os.listdir(os.path.join(pdir, 'train/tdcsfog'))\nprint(f'a number of files: {len(tdcsfog_train_files)}')","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:54.638506Z","iopub.execute_input":"2023-04-17T02:56:54.638949Z","iopub.status.idle":"2023-04-17T02:56:54.732505Z","shell.execute_reply.started":"2023-04-17T02:56:54.638913Z","shell.execute_reply":"2023-04-17T02:56:54.731531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore the details of the data","metadata":{}},{"cell_type":"code","source":"csv_fname = tdcsfog_train_files[0]\ndf_tdcs_train = pd.read_csv(os.path.join(pdir, 'train/tdcsfog', csv_fname))\ndisplay(df_tdcs_train.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:54.735089Z","iopub.execute_input":"2023-04-17T02:56:54.735833Z","iopub.status.idle":"2023-04-17T02:56:54.790671Z","shell.execute_reply.started":"2023-04-17T02:56:54.735792Z","shell.execute_reply":"2023-04-17T02:56:54.789783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# descriptive statistics\ndisplay(df_tdcs_train.describe())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:54.792237Z","iopub.execute_input":"2023-04-17T02:56:54.792917Z","iopub.status.idle":"2023-04-17T02:56:54.843816Z","shell.execute_reply.started":"2023-04-17T02:56:54.792880Z","shell.execute_reply":"2023-04-17T02:56:54.841989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check NaN values\ndisplay(df_tdcs_train.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:54.845682Z","iopub.execute_input":"2023-04-17T02:56:54.846107Z","iopub.status.idle":"2023-04-17T02:56:54.858456Z","shell.execute_reply.started":"2023-04-17T02:56:54.846062Z","shell.execute_reply":"2023-04-17T02:56:54.856936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the DataFrame\n\n- The visualization example below suggests the occurrence of a FOG event in Turn.","metadata":{}},{"cell_type":"code","source":"# Show ['AccV', 'AccML', 'AccAP', 'StartHesitation', 'Turn', 'Walking']\nrows = 3\ncolumns = 2\nnames = ['AccV', 'AccML', 'AccAP', 'StartHesitation', 'Turn', 'Walking']\n\ni = 0\nfig, ax = plt.subplots(3, 2, figsize=(15, 8))\nfor c in range(columns):    \n    for r in range(rows):        \n        ax[r, c].plot(df_tdcs_train[names[i]])\n        ax[r, c].set_title(names[i])\n        i += 1\nfig.suptitle(csv_fname, fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:54.860107Z","iopub.execute_input":"2023-04-17T02:56:54.860528Z","iopub.status.idle":"2023-04-17T02:56:56.125792Z","shell.execute_reply.started":"2023-04-17T02:56:54.860490Z","shell.execute_reply":"2023-04-17T02:56:56.124519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compare Turn is False vs. True\n\nrows = 3\ncolumns = 2\nnames = ['AccV', 'AccML', 'AccAP']\n\n\nfig, ax = plt.subplots(3, 2, figsize=(15, 8))\n\n# plot acceleration while Turn is False\nfor r in range(rows):\n    df_trun_is_false = df_tdcs_train[df_tdcs_train['Turn'] == False].reset_index(drop=True)\n    ax[r, 0].plot(df_trun_is_false[names[r]])\n    ax[r, 0].set_title(f'Turn is False: {names[r]}')\n    # Match the Y-axis display range between True and False.\n    ymin, ymax = min(df_tdcs_train[names[r]]), max(df_tdcs_train[names[r]])\n    ax[r, 0].set_ylim(ymin-0.1*abs(ymin), ymax+0.1*ymax)\n    \n# plot acceleration while Turn is True\nfor r in range(rows):\n    df_trun_is_true = df_tdcs_train[df_tdcs_train['Turn'] == True].reset_index(drop=True)\n    ax[r, 1].plot(df_trun_is_true[names[r]])\n    ax[r, 1].set_title(f'Turn is True: {names[r]}')\n    # Match the Y-axis display range between True and False.\n    ymin, ymax = min(df_tdcs_train[names[r]]), max(df_tdcs_train[names[r]])\n    ax[r, 1].set_ylim(ymin-0.1*abs(ymin), ymax+0.1*ymax)\n    \nfig.suptitle(f'Turn is False vs. True in {csv_fname}', fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:56.127378Z","iopub.execute_input":"2023-04-17T02:56:56.127756Z","iopub.status.idle":"2023-04-17T02:56:57.173791Z","shell.execute_reply.started":"2023-04-17T02:56:56.127686Z","shell.execute_reply":"2023-04-17T02:56:57.172433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- When considering only the data in a171e61840.csv, there is a tendency for the acceleration sensor values to have greater variability when Turn is True compared to when it is False.","metadata":{}},{"cell_type":"markdown","source":"# Explore tdcsfog metadata\n<ul>\n<li><p><strong>tdcsfog_metadata.csv</strong> Identifies each series in the <code>tdcsfog</code> dataset by a unique <code>Subject, Visit, Test, Medication</code> condition.</p>\n<ul>\n<li><code>Visit</code> Lab visits consist of a baseline assessment, two post-treatment assessments for different treatment stages, and one follow-up assessment.</li>\n<li><code>Test</code> Which of three test types was performed, with <code>3</code> the most challenging.</li>\n<li><code>Medication</code> Subjects may have been either <code>off</code> or <code>on</code> anti-parkinsonian medication during the recording.</li></ul></li></ul>","metadata":{}},{"cell_type":"code","source":"df_tdcs_meta = pd.read_csv(os.path.join(pdir, 'tdcsfog_metadata.csv'))\ndisplay(df_tdcs_meta.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.175493Z","iopub.execute_input":"2023-04-17T02:56:57.176371Z","iopub.status.idle":"2023-04-17T02:56:57.196527Z","shell.execute_reply.started":"2023-04-17T02:56:57.176327Z","shell.execute_reply":"2023-04-17T02:56:57.195133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore the details of the data","metadata":{"execution":{"iopub.status.busy":"2023-04-12T21:27:14.456147Z","iopub.execute_input":"2023-04-12T21:27:14.456572Z","iopub.status.idle":"2023-04-12T21:27:14.469695Z","shell.execute_reply.started":"2023-04-12T21:27:14.456537Z","shell.execute_reply":"2023-04-12T21:27:14.468499Z"}}},{"cell_type":"code","source":"# descriptive statistics\ndisplay(df_tdcs_meta.describe(include='all'))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.202482Z","iopub.execute_input":"2023-04-17T02:56:57.202909Z","iopub.status.idle":"2023-04-17T02:56:57.237434Z","shell.execute_reply.started":"2023-04-17T02:56:57.202871Z","shell.execute_reply":"2023-04-17T02:56:57.235979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Since the number of unique values in the Subject column is 62, there is data available for at least 62 subjects in the tdcsfog dataset.\n- Since all Id values are different, it can be inferred that the data is subdivided into smaller segments for the 62 subjects.\n\n\n\n","metadata":{}},{"cell_type":"code","source":"# check NaN values\ndisplay(df_tdcs_meta.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.239347Z","iopub.execute_input":"2023-04-17T02:56:57.239877Z","iopub.status.idle":"2023-04-17T02:56:57.252609Z","shell.execute_reply.started":"2023-04-17T02:56:57.239823Z","shell.execute_reply":"2023-04-17T02:56:57.251032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the number of data corresponding to each subject\nsubject_count = df_tdcs_meta[['Id', 'Subject']].groupby('Subject').count()\n\ndisplay('with a small number of data',subject_count.sort_values(by='Id').head(2))\ndisplay('with a large number of data',subject_count.sort_values(by='Id').tail(2))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.254435Z","iopub.execute_input":"2023-04-17T02:56:57.254942Z","iopub.status.idle":"2023-04-17T02:56:57.285084Z","shell.execute_reply.started":"2023-04-17T02:56:57.254890Z","shell.execute_reply":"2023-04-17T02:56:57.284071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There is variation in the number of data for each subject, with subject 39306b having only two data, while subject f5586 has 24 data.","metadata":{}},{"cell_type":"code","source":"# Display the unique subjects data\nsubjects_id = 'f5586f'\ndisplay(df_tdcs_meta[df_tdcs_meta['Subject'] == subjects_id])","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.286555Z","iopub.execute_input":"2023-04-17T02:56:57.287965Z","iopub.status.idle":"2023-04-17T02:56:57.306366Z","shell.execute_reply.started":"2023-04-17T02:56:57.287913Z","shell.execute_reply":"2023-04-17T02:56:57.305431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- I expected the Medication column to always be \"on\" if a participant was undergoing treatment for Parkinson's disease.\n- However, in reality, the Medication column can contain both \"on\" and \"off\" values for specific subjects. This column indicates whether or not tha subject took their medication during the laboratory measurement.\n- There may be a possibility that there are changes in the trend of acceleration sensor values depending on whether medication is taken or not during the measurement.","metadata":{}},{"cell_type":"markdown","source":"## Visualize the Dataframe","metadata":{}},{"cell_type":"code","source":"# bar chart for 'Visit'\n\nvisit_counts = df_tdcs_meta['Visit'].value_counts()\n\nplt.figure()\nplt.bar(x=visit_counts.index, height=visit_counts.values)\nplt.title('counts of Visit')\nplt.xlabel('Visit')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.308136Z","iopub.execute_input":"2023-04-17T02:56:57.309025Z","iopub.status.idle":"2023-04-17T02:56:57.547570Z","shell.execute_reply.started":"2023-04-17T02:56:57.308975Z","shell.execute_reply":"2023-04-17T02:56:57.546278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Test'\n\ntest_counts = df_tdcs_meta['Test'].value_counts()\n\nplt.figure()\nplt.bar(x=test_counts.index, height=test_counts.values)\nplt.title('counts of Test')\nplt.xlabel('Test')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.549113Z","iopub.execute_input":"2023-04-17T02:56:57.550158Z","iopub.status.idle":"2023-04-17T02:56:57.780752Z","shell.execute_reply.started":"2023-04-17T02:56:57.550103Z","shell.execute_reply":"2023-04-17T02:56:57.779632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Medication'\n\nmed_counts = df_tdcs_meta['Medication'].value_counts()\n\nplt.figure()\nplt.bar(x=med_counts.index, height=med_counts.values)\nplt.title('counts of Medication')\nplt.xlabel('Medication')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.782296Z","iopub.execute_input":"2023-04-17T02:56:57.783522Z","iopub.status.idle":"2023-04-17T02:56:57.980816Z","shell.execute_reply.started":"2023-04-17T02:56:57.783473Z","shell.execute_reply":"2023-04-17T02:56:57.979775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore defog train data\n\n<li><strong>train/</strong> Folder containing the data series in the training set within three subfolders: <strong>tdcsfog/</strong>, <strong>defog/</strong>, and <strong>notype/</strong>. Series in the <em>notype</em> folder are from the <code>defog</code> dataset but lack event-type annotations. The fields present in these series vary by folder.<ul>\n<li><code>Time</code> An integer timestep. Series from the <code>tdcsfog</code> dataset are recorded at 128Hz (128 timesteps per second), while series from the <code>defog</code> and <code>daily</code> series are recorded at 100Hz (100 timesteps per second).</li>\n<li><code>AccV</code>, <code>AccML</code>, and <code>AccAP</code> Acceleration from a lower-back sensor on three axes: V - vertical, ML - mediolateral, AP - anteroposterior. Data is in units of <em>m/s^2</em> for <code>tdcsfog/</code> and <em>g</em> for <code>defog/</code> and <code>notype/</code>.</li>\n<li><code>StartHesitation</code>, <code>Turn</code>, <code>Walking</code> Indicator variables for the occurrence of each of the event types.</li>\n<li><code>Event</code> Indicator variable for the occurrence of <em>any</em> FOG-type event. Present only in the <strong>notype</strong> series, which lack type-level annotations.</li>\n<li><code>Valid</code> There were cases during the video annotation that were hard for the annotator to decide if there was an Akinetic (i.e., essentially no movement) FoG or the subject stopped voluntarily. Only event annotations where the series is marked <code>true</code> should be considered as unambiguous.</li>\n<li><code>Task</code> Series were only annotated where this value is <code>true</code>. Portions marked <code>false</code> should be considered unannotated.</li></ul></li>","metadata":{}},{"cell_type":"markdown","source":"## The number of csv files","metadata":{}},{"cell_type":"code","source":"defog_train_files = os.listdir(os.path.join(pdir, 'train/defog'))\nprint(f'a number of files: {len(defog_train_files)}')","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.982456Z","iopub.execute_input":"2023-04-17T02:56:57.983600Z","iopub.status.idle":"2023-04-17T02:56:57.996661Z","shell.execute_reply.started":"2023-04-17T02:56:57.983546Z","shell.execute_reply":"2023-04-17T02:56:57.995178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check a CSV file\n\n- The column \"Time\" is recorded as integer values.\n    - The sampling frequency is 100Hz, so it was assumed that each row represents a time interval of 1/100 seconds, but it may instead represent simple measurement points contrary to my expectations.","metadata":{}},{"cell_type":"code","source":"csv_fname = defog_train_files[0]\ndf_de_train = pd.read_csv(os.path.join(pdir, 'train/defog', csv_fname))\ndisplay(df_de_train.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:57.998157Z","iopub.execute_input":"2023-04-17T02:56:57.998552Z","iopub.status.idle":"2023-04-17T02:56:58.278914Z","shell.execute_reply.started":"2023-04-17T02:56:57.998516Z","shell.execute_reply":"2023-04-17T02:56:58.277944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# descriptive statistics\ndisplay(df_de_train.describe(include='all'))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:58.280269Z","iopub.execute_input":"2023-04-17T02:56:58.281214Z","iopub.status.idle":"2023-04-17T02:56:58.349070Z","shell.execute_reply.started":"2023-04-17T02:56:58.281176Z","shell.execute_reply":"2023-04-17T02:56:58.347823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check NaN values\ndisplay(df_de_train.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:58.350734Z","iopub.execute_input":"2023-04-17T02:56:58.351113Z","iopub.status.idle":"2023-04-17T02:56:58.365554Z","shell.execute_reply.started":"2023-04-17T02:56:58.351076Z","shell.execute_reply":"2023-04-17T02:56:58.364184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the DataFrame\n\n- The visualization example below suggests the occurrence of a FOG events in Turn and Waling.","metadata":{}},{"cell_type":"code","source":"# show 'AccV', 'AccML', 'AccAP', 'StartHesitation', 'Turn', 'Walking', 'Valid', 'Task'\nrows = 3\ncolumns = 3\nnames = ['AccV', 'AccML', 'AccAP', 'StartHesitation', 'Turn', 'Walking', 'Valid', 'Task', 'Pass']\n\ni = 0\nfig, ax = plt.subplots(3, 3, figsize=(15, 8))\nfor c in range(columns):    \n    for r in range(rows):      \n        if names[i] == 'Pass':\n            pass\n        else:\n            ax[r, c].plot(df_de_train[names[i]])\n            ax[r, c].set_title(names[i])\n        i += 1\nfig.suptitle(csv_fname, fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:58.367166Z","iopub.execute_input":"2023-04-17T02:56:58.368229Z","iopub.status.idle":"2023-04-17T02:56:59.931202Z","shell.execute_reply.started":"2023-04-17T02:56:58.368177Z","shell.execute_reply":"2023-04-17T02:56:59.929740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# show 'AccV', 'AccML', 'AccAP', 'StartHesitation', 'Turn', 'Walking' while 'Valid' and 'Task' is True\nrows = 3\ncolumns = 3\nnames = ['AccV', 'AccML', 'AccAP', 'StartHesitation', 'Turn', 'Walking', 'Valid', 'Task', 'Pass']\n\n# get data while Valid and Task is True because marked false should be considered unannotated.\ndf_de_is_Valid_Task = df_de_train[(df_de_train['Valid']==True) & (df_de_train['Task']==True)].reset_index(drop=True)\n\ni = 0\nfig, ax = plt.subplots(3, 3, figsize=(15, 8))\nfor c in range(columns):    \n    for r in range(rows):      \n        if names[i] == 'Pass':\n            pass\n        else:\n            ax[r, c].plot(df_de_is_Valid_Task[names[i]])\n            ax[r, c].set_title(names[i])\n        i += 1\nfig.suptitle(f'Valid and Task are both True @{csv_fname}', fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:56:59.933428Z","iopub.execute_input":"2023-04-17T02:56:59.934541Z","iopub.status.idle":"2023-04-17T02:57:01.600747Z","shell.execute_reply.started":"2023-04-17T02:56:59.934477Z","shell.execute_reply":"2023-04-17T02:57:01.599531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- After extracting the data where Valid and Task columns are both True, the Walking events disappeared.","metadata":{}},{"cell_type":"code","source":"# Compare Turn is False vs. True\n\nrows = 3\ncolumns = 2\nnames = ['AccV', 'AccML', 'AccAP']\n\n\nfig, ax = plt.subplots(3, 2, figsize=(15, 8))\n\n# plot acceleration while Turn is False\nfor r in range(rows):\n    df_trun_is_false = df_de_is_Valid_Task[df_de_is_Valid_Task['Turn'] == False].reset_index(drop=True)\n    ax[r, 0].plot(df_trun_is_false[names[r]])\n    ax[r, 0].set_title(f'Turn is False: {names[r]}')\n    # Match the Y-axis display range between True and False.\n    ymin, ymax = min(df_de_is_Valid_Task[names[r]]), max(df_de_is_Valid_Task[names[r]])\n    ax[r, 0].set_ylim(ymin-0.1*abs(ymin), ymax+0.1*ymax)\n    \n# plot acceleration while Turn is True\nfor r in range(rows):\n    df_trun_is_true = df_de_is_Valid_Task[df_de_is_Valid_Task['Turn'] == True].reset_index(drop=True)\n    ax[r, 1].plot(df_trun_is_true[names[r]])\n    ax[r, 1].set_title(f'Turn is True: {names[r]}')\n    # Match the Y-axis display range between True and False.\n    ymin, ymax = min(df_de_is_Valid_Task[names[r]]), max(df_de_is_Valid_Task[names[r]])\n    ax[r, 1].set_ylim(ymin-0.1*abs(ymin), ymax+0.1*ymax)\n    \nfig.suptitle(f'Turn is False vs. True in {csv_fname}', fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:01.602643Z","iopub.execute_input":"2023-04-17T02:57:01.603442Z","iopub.status.idle":"2023-04-17T02:57:02.862815Z","shell.execute_reply.started":"2023-04-17T02:57:01.603394Z","shell.execute_reply":"2023-04-17T02:57:02.861660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Unlike in the case of a171e61840.csv in tdcsfog dataset, in be9d33541d.csv, the variance of the acceleration series when Turn is True is smaller than the variance when Turn is False.","metadata":{}},{"cell_type":"markdown","source":"# Explore defog metadata\n\n<p><strong>defog_metadata.csv</strong> Identifies each series in the <code>defog</code> dataset by a unique <code>Subject, Visit, Medication</code> condition.</p>\n<ul>\n<li><code>Visit</code> Lab visits consist of a baseline assessment, two post-treatment assessments for different treatment stages, and one follow-up assessment.</li>\n<li><code>Medication</code> Subjects may have been either <code>off</code> or <code>on</code> anti-parkinsonian medication during the recording.</li></ul>","metadata":{}},{"cell_type":"code","source":"df_de_meta = pd.read_csv(os.path.join(pdir, 'defog_metadata.csv'))\ndisplay(df_de_meta.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:02.864837Z","iopub.execute_input":"2023-04-17T02:57:02.865676Z","iopub.status.idle":"2023-04-17T02:57:02.887202Z","shell.execute_reply.started":"2023-04-17T02:57:02.865619Z","shell.execute_reply":"2023-04-17T02:57:02.886070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore the datails of the data","metadata":{}},{"cell_type":"code","source":"# descriptive statistics\ndisplay(df_de_meta.describe(include='all'))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:02.897259Z","iopub.execute_input":"2023-04-17T02:57:02.898305Z","iopub.status.idle":"2023-04-17T02:57:02.927166Z","shell.execute_reply.started":"2023-04-17T02:57:02.898239Z","shell.execute_reply":"2023-04-17T02:57:02.926021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Since the number of unique values in the Subject column is 45, there is data available for at least 45 subjects in the defog dataset.\n- Since all Id values are different, it can be inferred that the data is subdivided into smaller segments for the 45 subjects.","metadata":{}},{"cell_type":"code","source":"# check NaN values\ndisplay(df_de_meta.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:02.929175Z","iopub.execute_input":"2023-04-17T02:57:02.930043Z","iopub.status.idle":"2023-04-17T02:57:02.941638Z","shell.execute_reply.started":"2023-04-17T02:57:02.929996Z","shell.execute_reply":"2023-04-17T02:57:02.940515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the number of data corresponding to each subject\nsubject_count = df_de_meta[['Id', 'Subject']].groupby('Subject').count()\n\ndisplay('with a small number of data',subject_count.sort_values(by='Id').head(2))\ndisplay('with a large number of data',subject_count.sort_values(by='Id').tail(2))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:02.943630Z","iopub.execute_input":"2023-04-17T02:57:02.944500Z","iopub.status.idle":"2023-04-17T02:57:02.975982Z","shell.execute_reply.started":"2023-04-17T02:57:02.944453Z","shell.execute_reply":"2023-04-17T02:57:02.974855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the unique subjects data\nsubjects_id = 'fe5d84'\ndisplay(df_de_meta[df_de_meta['Subject'] == subjects_id])","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:02.977956Z","iopub.execute_input":"2023-04-17T02:57:02.978811Z","iopub.status.idle":"2023-04-17T02:57:02.996093Z","shell.execute_reply.started":"2023-04-17T02:57:02.978764Z","shell.execute_reply":"2023-04-17T02:57:02.994949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the DataFrame","metadata":{}},{"cell_type":"code","source":"# bar chart for 'Visit'\n\nvisit_counts = df_de_meta['Visit'].value_counts()\n\nplt.figure()\nplt.bar(x=visit_counts.index, height=visit_counts.values)\nplt.title('counts of Visit')\nplt.xlabel('Visit')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:02.998156Z","iopub.execute_input":"2023-04-17T02:57:02.999015Z","iopub.status.idle":"2023-04-17T02:57:03.278761Z","shell.execute_reply.started":"2023-04-17T02:57:02.998970Z","shell.execute_reply":"2023-04-17T02:57:03.277295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Medication'\n\nmed_counts = df_tdcs_meta['Medication'].value_counts()\n\nplt.figure()\nplt.bar(x=med_counts.index, height=med_counts.values)\nplt.title('counts of Medication')\nplt.xlabel('Medication')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:03.280211Z","iopub.execute_input":"2023-04-17T02:57:03.281052Z","iopub.status.idle":"2023-04-17T02:57:03.478146Z","shell.execute_reply.started":"2023-04-17T02:57:03.281003Z","shell.execute_reply":"2023-04-17T02:57:03.476796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore subjects metadata\n\n<ul>\n<li><p><strong>subjects.csv</strong> Metadata for each <code>Subject</code> in the study, including their <code>Age</code> and <code>Sex</code> as well as:</p>\n<ul>\n<li><code>Visit</code> Only available for subjects in the <code>daily</code> and <code>defog</code> datasets.</li>\n<li><code>YearsSinceDx</code> Years since Parkinson's diagnosis.</li>\n<li><code>UPDRSIIIOn</code>/<code>UPDRSIIIOff</code> Unified Parkinson's Disease Rating Scale score during on/off medication respectively.</li>\n<li><code>NFOGQ</code> Self-report FoG questionnaire score. See: <br>\n<a rel=\"noreferrer nofollow\" target=\"_blank\" href=\"https://pubmed.ncbi.nlm.nih.gov/19660949/\">https://pubmed.ncbi.nlm.nih.gov/19660949/</a></li></ul></li></ul>","metadata":{}},{"cell_type":"code","source":"df_subjects = pd.read_csv(os.path.join(pdir, 'subjects.csv'))\ndisplay(df_subjects.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:03.479796Z","iopub.execute_input":"2023-04-17T02:57:03.481032Z","iopub.status.idle":"2023-04-17T02:57:03.510472Z","shell.execute_reply.started":"2023-04-17T02:57:03.480981Z","shell.execute_reply":"2023-04-17T02:57:03.509093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore the details of the data","metadata":{}},{"cell_type":"code","source":"# descriptive statistics\ndisplay(df_subjects.describe(include='all'))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:03.512632Z","iopub.execute_input":"2023-04-17T02:57:03.513764Z","iopub.status.idle":"2023-04-17T02:57:03.564099Z","shell.execute_reply.started":"2023-04-17T02:57:03.513689Z","shell.execute_reply":"2023-04-17T02:57:03.562782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check NaN values\ndisplay(df_subjects.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:03.565735Z","iopub.execute_input":"2023-04-17T02:57:03.566480Z","iopub.status.idle":"2023-04-17T02:57:03.579679Z","shell.execute_reply.started":"2023-04-17T02:57:03.566432Z","shell.execute_reply":"2023-04-17T02:57:03.578061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the DataFrame","metadata":{}},{"cell_type":"code","source":"# bar chart for 'Visit'\n\nvisit_counts = df_subjects['Visit'].value_counts()\n\nplt.figure()\nplt.bar(x=visit_counts.index, height=visit_counts.values)\nplt.title('counts of Visit')\nplt.xlabel('Visit')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:03.582049Z","iopub.execute_input":"2023-04-17T02:57:03.582552Z","iopub.status.idle":"2023-04-17T02:57:03.819456Z","shell.execute_reply.started":"2023-04-17T02:57:03.582474Z","shell.execute_reply":"2023-04-17T02:57:03.818158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Age'\n\nage_counts = df_subjects['Age'].value_counts()\n\nplt.figure()\nplt.bar(x=age_counts.index, height=age_counts.values)\nplt.title('counts of Age')\nplt.xlabel('Age')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:03.820964Z","iopub.execute_input":"2023-04-17T02:57:03.822001Z","iopub.status.idle":"2023-04-17T02:57:04.102530Z","shell.execute_reply.started":"2023-04-17T02:57:03.821959Z","shell.execute_reply":"2023-04-17T02:57:04.101622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Sex'\n\nsex_counts = df_subjects['Sex'].value_counts()\n\nplt.figure()\nplt.bar(x=sex_counts.index, height=sex_counts.values)\nplt.title('counts of Sex')\nplt.xlabel('Sex')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:04.104077Z","iopub.execute_input":"2023-04-17T02:57:04.104456Z","iopub.status.idle":"2023-04-17T02:57:04.296034Z","shell.execute_reply.started":"2023-04-17T02:57:04.104420Z","shell.execute_reply":"2023-04-17T02:57:04.294822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'YearsSinceDx'\n\nysd_counts = df_subjects['YearsSinceDx'].value_counts()\n\nplt.figure()\nplt.bar(x=ysd_counts.index, height=ysd_counts.values)\nplt.title('counts of YearsSinceDx')\nplt.xlabel('YearsSinceDx')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:04.297373Z","iopub.execute_input":"2023-04-17T02:57:04.297748Z","iopub.status.idle":"2023-04-17T02:57:04.583042Z","shell.execute_reply.started":"2023-04-17T02:57:04.297687Z","shell.execute_reply":"2023-04-17T02:57:04.581736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'UPDRSIII_On'\n\nupdon_counts = df_subjects['UPDRSIII_On'].value_counts()\n\nplt.figure()\nplt.bar(x=updon_counts.index, height=updon_counts.values)\nplt.title('counts of UPDRSIII_On')\nplt.xlabel('UPDRSIII_On')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:04.584547Z","iopub.execute_input":"2023-04-17T02:57:04.584973Z","iopub.status.idle":"2023-04-17T02:57:04.903873Z","shell.execute_reply.started":"2023-04-17T02:57:04.584935Z","shell.execute_reply":"2023-04-17T02:57:04.902502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'UPDRSIII_Off'\n\nupdoff_counts = df_subjects['UPDRSIII_Off'].value_counts()\n\nplt.figure()\nplt.bar(x=updoff_counts.index, height=updoff_counts.values)\nplt.title('counts of UPDRSIII_Off')\nplt.xlabel('UPDRSIII_Off')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:04.905730Z","iopub.execute_input":"2023-04-17T02:57:04.906163Z","iopub.status.idle":"2023-04-17T02:57:05.229458Z","shell.execute_reply.started":"2023-04-17T02:57:04.906117Z","shell.execute_reply":"2023-04-17T02:57:05.228243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'NFOGQ'\n\nnfogq_counts = df_subjects['NFOGQ'].value_counts()\n\nplt.figure()\nplt.bar(x=nfogq_counts.index, height=nfogq_counts.values)\nplt.title('counts of NFOGQ')\nplt.xlabel('NFOGQ')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:05.230757Z","iopub.execute_input":"2023-04-17T02:57:05.231077Z","iopub.status.idle":"2023-04-17T02:57:05.717181Z","shell.execute_reply.started":"2023-04-17T02:57:05.231046Z","shell.execute_reply":"2023-04-17T02:57:05.716042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore tasks data\n\n<ul>\n<li><p><strong>tasks.csv</strong> Task metadata for series in the <code>defog</code> dataset. (Not relevant for the series in the <code>fog</code> or <code>daily</code> datasets.)</p>\n<ul>\n<li><code>Id</code> The data series where the task was measured.</li>\n<li><code>Begin</code> Time (s) the task began.</li>\n<li><code>End</code> Time (s) the task ended.</li>\n<li><code>Task</code> One of seven tasks types in the DeFOG protocol, described on <a target=\"_blank\" href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/additional-data-documentation\">this page</a>.</li></ul></li></ul>\n    \n- By subtracting the value in the Begin column from the value in the End column for each row, the duration time per series can be calculated (?).","metadata":{}},{"cell_type":"code","source":"df_tasks = pd.read_csv(os.path.join(pdir, 'tasks.csv'))\ndisplay(df_tasks.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:05.718834Z","iopub.execute_input":"2023-04-17T02:57:05.719198Z","iopub.status.idle":"2023-04-17T02:57:05.741242Z","shell.execute_reply.started":"2023-04-17T02:57:05.719163Z","shell.execute_reply":"2023-04-17T02:57:05.739781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore the details of the data","metadata":{}},{"cell_type":"code","source":"print(f'length of tasks.csv: {len(df_tasks)}')","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:05.742816Z","iopub.execute_input":"2023-04-17T02:57:05.743606Z","iopub.status.idle":"2023-04-17T02:57:05.748314Z","shell.execute_reply.started":"2023-04-17T02:57:05.743569Z","shell.execute_reply":"2023-04-17T02:57:05.747432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# descriptive statistics\ndisplay(df_tasks.describe(include='all'))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:05.749673Z","iopub.execute_input":"2023-04-17T02:57:05.750305Z","iopub.status.idle":"2023-04-17T02:57:05.777939Z","shell.execute_reply.started":"2023-04-17T02:57:05.750269Z","shell.execute_reply":"2023-04-17T02:57:05.776546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check NaN values\ndisplay(df_tasks.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:05.779887Z","iopub.execute_input":"2023-04-17T02:57:05.780351Z","iopub.status.idle":"2023-04-17T02:57:05.790266Z","shell.execute_reply.started":"2023-04-17T02:57:05.780305Z","shell.execute_reply":"2023-04-17T02:57:05.789186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# types of Tasks\nprint(df_tasks['Task'].unique())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:05.792105Z","iopub.execute_input":"2023-04-17T02:57:05.793203Z","iopub.status.idle":"2023-04-17T02:57:05.801249Z","shell.execute_reply.started":"2023-04-17T02:57:05.793149Z","shell.execute_reply":"2023-04-17T02:57:05.799690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There are various types of tasks.","metadata":{"execution":{"iopub.status.busy":"2023-04-13T09:33:32.217329Z","iopub.execute_input":"2023-04-13T09:33:32.218205Z","iopub.status.idle":"2023-04-13T09:33:32.226287Z","shell.execute_reply.started":"2023-04-13T09:33:32.218151Z","shell.execute_reply":"2023-04-13T09:33:32.224275Z"}}},{"cell_type":"markdown","source":"## Visualize the DataFrame","metadata":{"execution":{"iopub.status.busy":"2023-04-13T09:35:37.930704Z","iopub.execute_input":"2023-04-13T09:35:37.931796Z","iopub.status.idle":"2023-04-13T09:35:37.937742Z","shell.execute_reply.started":"2023-04-13T09:35:37.931724Z","shell.execute_reply":"2023-04-13T09:35:37.936364Z"}}},{"cell_type":"code","source":"# bar chart for 'Task'\n\ntask_counts = df_tasks['Task'].value_counts()\n\nplt.figure()\nplt.bar(x=task_counts.index, height=task_counts.values)\nplt.title('counts of Task')\nplt.xticks(rotation=-90)\nplt.xlabel('Task')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:05.803144Z","iopub.execute_input":"2023-04-17T02:57:05.803488Z","iopub.status.idle":"2023-04-17T02:57:06.171370Z","shell.execute_reply.started":"2023-04-17T02:57:05.803456Z","shell.execute_reply":"2023-04-17T02:57:06.170298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Duration'\ndf_tasks['Duration'] = df_tasks['End'] - df_tasks['Begin']\n\nplt.figure()\nplt.hist(df_tasks['Duration'], bins=50)\nplt.title('Histgram of Duration')\nplt.xlabel('Duration')\nplt.ylabel('frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:06.172421Z","iopub.execute_input":"2023-04-17T02:57:06.172768Z","iopub.status.idle":"2023-04-17T02:57:06.509726Z","shell.execute_reply.started":"2023-04-17T02:57:06.172734Z","shell.execute_reply":"2023-04-17T02:57:06.508790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize defog acceleration data by task","metadata":{}},{"cell_type":"code","source":"target_id = '02ea782681'","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:06.510846Z","iopub.execute_input":"2023-04-17T02:57:06.511408Z","iopub.status.idle":"2023-04-17T02:57:06.516360Z","shell.execute_reply.started":"2023-04-17T02:57:06.511373Z","shell.execute_reply":"2023-04-17T02:57:06.514984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get target_id task data\ndf_tasks_tmp = df_tasks[df_tasks['Id'] == target_id]\n\n# Multiply 100 to Begin, End, and Duration to adjust to a measurement frequency of 100Hz.\ndf_tasks_tmp.at[:, ['Begin', 'End', 'Duration']] = (df_tasks_tmp[['Begin', 'End', 'Duration']] * 100).astype(int)\n\ndisplay(df_tasks_tmp)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:06.517853Z","iopub.execute_input":"2023-04-17T02:57:06.518290Z","iopub.status.idle":"2023-04-17T02:57:06.546305Z","shell.execute_reply.started":"2023-04-17T02:57:06.518257Z","shell.execute_reply":"2023-04-17T02:57:06.545158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_defog_tmp = pd.read_csv(os.path.join(pdir, 'train/defog', target_id+'.csv'))\ndisplay(df_defog_tmp)","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:06.549847Z","iopub.execute_input":"2023-04-17T02:57:06.550193Z","iopub.status.idle":"2023-04-17T02:57:06.816202Z","shell.execute_reply.started":"2023-04-17T02:57:06.550160Z","shell.execute_reply":"2023-04-17T02:57:06.814851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check the length of Task data\nlen(df_tasks_tmp['Task'])","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:06.817871Z","iopub.execute_input":"2023-04-17T02:57:06.818363Z","iopub.status.idle":"2023-04-17T02:57:06.827374Z","shell.execute_reply.started":"2023-04-17T02:57:06.818311Z","shell.execute_reply":"2023-04-17T02:57:06.826004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compare AccV by Task\n\nrows = 11\ncolumns = 3\nTask = df_tasks_tmp['Task'].to_list()\nBegin = df_tasks_tmp['Begin'].to_list()\nEnd = df_tasks_tmp['End'].to_list()\n\ni = 0\nfig, ax = plt.subplots(rows, columns, figsize=(15, 15))\nfor c in range(columns):    \n    for r in range(rows):   \n        try:\n            ax[r, c].plot(df_defog_tmp.loc[Begin[i]:End[i], 'AccV'])\n            ax[r, c].set_title(Task[i])\n        except:\n            pass\n        i += 1\n\nfig.suptitle(f'AccV by type @ {target_id}', fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:06.829415Z","iopub.execute_input":"2023-04-17T02:57:06.830304Z","iopub.status.idle":"2023-04-17T02:57:11.515152Z","shell.execute_reply.started":"2023-04-17T02:57:06.830265Z","shell.execute_reply":"2023-04-17T02:57:11.514006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compare AccML by Task\n\nrows = 11\ncolumns = 3\nTask = df_tasks_tmp['Task'].to_list()\nBegin = df_tasks_tmp['Begin'].to_list()\nEnd = df_tasks_tmp['End'].to_list()\n\ni = 0\nfig, ax = plt.subplots(rows, columns, figsize=(15, 15))\nfor c in range(columns):    \n    for r in range(rows):   \n        try:\n            ax[r, c].plot(df_defog_tmp.loc[Begin[i]:End[i], 'AccML'])\n            ax[r, c].set_title(Task[i])\n        except:\n            pass\n        i += 1\n\nfig.suptitle(f'AccML by type @ {target_id}', fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:11.516873Z","iopub.execute_input":"2023-04-17T02:57:11.517256Z","iopub.status.idle":"2023-04-17T02:57:16.233908Z","shell.execute_reply.started":"2023-04-17T02:57:11.517220Z","shell.execute_reply":"2023-04-17T02:57:16.232632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compare AccAP by Task\n\nrows = 11\ncolumns = 3\nTask = df_tasks_tmp['Task'].to_list()\nBegin = df_tasks_tmp['Begin'].to_list()\nEnd = df_tasks_tmp['End'].to_list()\n\ni = 0\nfig, ax = plt.subplots(rows, columns, figsize=(15, 15))\nfor c in range(columns):    \n    for r in range(rows):   \n        try:\n            ax[r, c].plot(df_defog_tmp.loc[Begin[i]:End[i], 'AccAP'])\n            ax[r, c].set_title(Task[i])\n        except:\n            pass\n        i += 1\n\nfig.suptitle(f'AccAP by type @ {target_id}', fontsize=20)\nplt.tight_layout(rect=[0, 0, 1, 0.99])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:16.235527Z","iopub.execute_input":"2023-04-17T02:57:16.236008Z","iopub.status.idle":"2023-04-17T02:57:21.102119Z","shell.execute_reply.started":"2023-04-17T02:57:16.235962Z","shell.execute_reply":"2023-04-17T02:57:21.100707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore events data\n<ul>\n<li><p><strong>events.csv</strong> Metadata for each FoG event in all data series. The event times agree with the labels in the data series.</p>\n<ul>\n<li><code>Id</code> The data series the event occured in.</li>\n<li><code>Init</code> Time (s) the event began.</li>\n<li><code>Completion</code> Time (s) the event ended.</li>\n<li><code>Type</code> Whether <code>StartHesitation</code>, <code>Turn</code>, or <code>Walking</code>.</li>\n<li><code>Kinetic</code> Whether the event was <em>kinetic</em> (<code>1</code>) and involved movement, or <em>akinetic</em> (<code>0</code>) and static.</li></ul></li></ul>","metadata":{}},{"cell_type":"code","source":"df_events = pd.read_csv(os.path.join(pdir, 'events.csv'))\ndisplay(df_events.head(5))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.103817Z","iopub.execute_input":"2023-04-17T02:57:21.104857Z","iopub.status.idle":"2023-04-17T02:57:21.140914Z","shell.execute_reply.started":"2023-04-17T02:57:21.104814Z","shell.execute_reply":"2023-04-17T02:57:21.139714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The datailes of the data","metadata":{"execution":{"iopub.status.busy":"2023-04-15T22:44:16.492799Z","iopub.execute_input":"2023-04-15T22:44:16.493189Z","iopub.status.idle":"2023-04-15T22:44:16.499270Z","shell.execute_reply.started":"2023-04-15T22:44:16.493157Z","shell.execute_reply":"2023-04-15T22:44:16.497812Z"}}},{"cell_type":"code","source":"# Length of the data\nprint(f'Length of the data: {len(df_events)}')","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.142893Z","iopub.execute_input":"2023-04-17T02:57:21.143269Z","iopub.status.idle":"2023-04-17T02:57:21.149328Z","shell.execute_reply.started":"2023-04-17T02:57:21.143225Z","shell.execute_reply":"2023-04-17T02:57:21.148011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# descriptive statistics\ndisplay(df_events.describe(include='all'))","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.150996Z","iopub.execute_input":"2023-04-17T02:57:21.151467Z","iopub.status.idle":"2023-04-17T02:57:21.184504Z","shell.execute_reply.started":"2023-04-17T02:57:21.151420Z","shell.execute_reply":"2023-04-17T02:57:21.183236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check NaN values\ndisplay(df_events.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.185913Z","iopub.execute_input":"2023-04-17T02:57:21.186341Z","iopub.status.idle":"2023-04-17T02:57:21.196639Z","shell.execute_reply.started":"2023-04-17T02:57:21.186307Z","shell.execute_reply":"2023-04-17T02:57:21.195284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the data","metadata":{}},{"cell_type":"code","source":"# histogram of 'Init'\n\nplt.figure()\nplt.hist(df_events['Init'], bins=50)\nplt.title('Histogram of Init')\nplt.xlabel('Init')\nplt.ylabel('frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.198480Z","iopub.execute_input":"2023-04-17T02:57:21.199249Z","iopub.status.idle":"2023-04-17T02:57:21.438562Z","shell.execute_reply.started":"2023-04-17T02:57:21.199204Z","shell.execute_reply":"2023-04-17T02:57:21.437268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# histogram of duration time\n\nduration = df_events['Completion'] - df_events['Init']\n\nplt.figure()\nplt.hist(duration, bins=50)\nplt.title('Histogram of Duration')\nplt.xlabel('Init')\nplt.ylabel('frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.440191Z","iopub.execute_input":"2023-04-17T02:57:21.440552Z","iopub.status.idle":"2023-04-17T02:57:21.682429Z","shell.execute_reply.started":"2023-04-17T02:57:21.440519Z","shell.execute_reply":"2023-04-17T02:57:21.680951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Type'\n\ncounts = df_events['Type'].value_counts()\n\nplt.figure()\nplt.bar(x=counts.index, height=counts.values)\nplt.title('counts of Type')\nplt.xlabel('Type')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.683715Z","iopub.execute_input":"2023-04-17T02:57:21.684107Z","iopub.status.idle":"2023-04-17T02:57:21.819080Z","shell.execute_reply.started":"2023-04-17T02:57:21.684072Z","shell.execute_reply":"2023-04-17T02:57:21.817819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bar chart for 'Kinetic'\n\ncounts = df_events['Kinetic'].value_counts()\n\nplt.figure()\nplt.bar(x=counts.index, height=counts.values)\nplt.title('counts of Kinetic')\nplt.xlabel('Kinetic')\nplt.ylabel('counts')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-17T02:57:21.820595Z","iopub.execute_input":"2023-04-17T02:57:21.821063Z","iopub.status.idle":"2023-04-17T02:57:21.971042Z","shell.execute_reply.started":"2023-04-17T02:57:21.821013Z","shell.execute_reply":"2023-04-17T02:57:21.969591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}