{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Intro","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we delve into the temporal patterns and relationships between preceding **tasks** and specific Types of FOG episodes.\n\n*Our goal* is to investigate if there are any `Task` values that have the most influence on the Type of FoG episodes or reveal any other temporal patterns that can help in predicting FOG episodes and their types using information about the tasks. By analyzing these patterns and identifying tasks that consistently precede FOG episodes, we aim to extract meaningful features that can enhance our machine learning model. This feature extraction process will provide valuable insights into the context of FoG and potential triggers for FOG episodes.","metadata":{}},{"cell_type":"code","source":"# Import libraries\nimport pandas as pd\nimport numpy as np\nimport glob\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:00.696043Z","iopub.execute_input":"2023-05-19T18:39:00.696983Z","iopub.status.idle":"2023-05-19T18:39:02.273032Z","shell.execute_reply.started":"2023-05-19T18:39:00.696897Z","shell.execute_reply":"2023-05-19T18:39:02.271566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Loading the data","metadata":{}},{"cell_type":"code","source":"data_dir = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/'\n\ndefog_metadata_file = f'{data_dir}defog_metadata.csv'\nevents_data_file = f'{data_dir}events.csv'\ntasks_data_file = f'{data_dir}tasks.csv'\n\nevents_data = pd.read_csv(events_data_file)\ntasks_data = pd.read_csv(tasks_data_file)\ndefog_metadata = pd.read_csv(defog_metadata_file)","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:02.285404Z","iopub.execute_input":"2023-05-19T18:39:02.285981Z","iopub.status.idle":"2023-05-19T18:39:02.352642Z","shell.execute_reply.started":"2023-05-19T18:39:02.285911Z","shell.execute_reply":"2023-05-19T18:39:02.350911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's get started\n\nAs we know from the data descriptions, `tasks.csv` is a task metadata for series in the defog dataset. (Not relevant for the series in the tdcsfog or daily datasets). Let's take a look at first 5 rows of this dataset.","metadata":{}},{"cell_type":"code","source":"tasks_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:02.363296Z","iopub.execute_input":"2023-05-19T18:39:02.364659Z","iopub.status.idle":"2023-05-19T18:39:02.406554Z","shell.execute_reply.started":"2023-05-19T18:39:02.364614Z","shell.execute_reply":"2023-05-19T18:39:02.405064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tasks_data.Task.unique()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:02.408146Z","iopub.execute_input":"2023-05-19T18:39:02.408737Z","iopub.status.idle":"2023-05-19T18:39:02.421292Z","shell.execute_reply.started":"2023-05-19T18:39:02.408703Z","shell.execute_reply":"2023-05-19T18:39:02.419723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For further analysis we also need to use the `events_data` dataset. Let's create a filtered `events_data` dataset with metadata for *defog* events only.","metadata":{}},{"cell_type":"code","source":"# Creating a list of indices that correspond to the notype series IDs in the defog_metadata\nnotype_idx = [0, 4, 13, 15, 17, 19, 20, 22, 25, 27, 30, 31, 32, 34, 36, 40, 42,\n              53, 54, 55, 58, 60, 62, 66, 67, 73, 74, 76, 79, 83, 84, 88, 89, 94,\n              96, 97, 98, 101, 102, 105, 113, 114, 115, 121, 124, 126]\n\n# Separating the defog_metadata\ndefog_metadata_only = defog_metadata[~defog_metadata.index.isin(notype_idx)].copy()\n\n# Filtering events_data dataset (for defog events only)\nevents_data_defog = events_data[events_data.Id.isin(defog_metadata_only.Id.unique())]","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:02.423280Z","iopub.execute_input":"2023-05-19T18:39:02.423644Z","iopub.status.idle":"2023-05-19T18:39:02.451165Z","shell.execute_reply.started":"2023-05-19T18:39:02.423616Z","shell.execute_reply":"2023-05-19T18:39:02.450208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Analyzing part","metadata":{}},{"cell_type":"markdown","source":"Let's make a new dataframe from `events_data` with additional columns of:\n- `Preceding_Tasks` - the list of Task values that precede corresponding FoG episode,\n- `Min_Duration` - the minimum duration between Begin of the Task and Init time of FoG episode,\n- `Task_Min_Duration` - the task corresponding to the minimum duration.\n\nAfter that we will explore each new column to gain insights into the temporal patterns of task occurrences before specific events, as well as the relationship between preceding tasks and event types.","metadata":{}},{"cell_type":"code","source":"result = []\n\nfor _, row in events_data_defog.iterrows():\n    id = row['Id']\n    init = row['Init']\n    completion = row['Completion']\n    type_val = row['Type']\n    \n    filtered_df = tasks_data[(tasks_data['Id'] == id) & (tasks_data['Begin'] < init) & (tasks_data['End'] < completion)]\n    task_values = filtered_df['Task'].unique()\n    if not filtered_df.empty:\n        # Calculate the duration between 'Begin' of Task and 'Init' of FoG episode\n        durations = init - filtered_df['Begin']\n        min_duration = np.min(durations)\n        \n        # Get the task corresponding to the minimum duration\n        task_min_duration = filtered_df.loc[durations.idxmin(), 'Task']\n    else:\n        min_duration = None\n        task_min_duration = None    \n    \n    # Append the result as a dictionary\n    result.append({'Id': id, 'Init': init, 'Completion': completion, 'Type': type_val, \\\n                   'Preceding_Tasks': task_values, 'Min_Duration': min_duration, 'Task_Min_Duration': task_min_duration})\n    \nresult_df = pd.DataFrame(result)","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:02.453274Z","iopub.execute_input":"2023-05-19T18:39:02.453640Z","iopub.status.idle":"2023-05-19T18:39:05.169864Z","shell.execute_reply.started":"2023-05-19T18:39:02.453611Z","shell.execute_reply":"2023-05-19T18:39:05.168783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:05.175757Z","iopub.execute_input":"2023-05-19T18:39:05.177280Z","iopub.status.idle":"2023-05-19T18:39:05.199262Z","shell.execute_reply.started":"2023-05-19T18:39:05.177216Z","shell.execute_reply":"2023-05-19T18:39:05.198054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The `Min_Duration` column\nLet's analyze the distributions of the minimum durations between begins of the tasks and the occurrence of FoG episodes for each event Type.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(7, 5))\n\n# Plot the distributions of Min_Duration column for each event Type\nfor event_type in result_df['Type'].unique():\n    data = result_df[result_df['Type'] == event_type]['Min_Duration']\n    sns.histplot(data, label=event_type)\n\nplt.xlabel('Minimum Duration (s)')\nplt.ylabel('The number of episodes')\nplt.title('Distributions of min durations between task and FoG for each event Type')\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:05.218879Z","iopub.execute_input":"2023-05-19T18:39:05.219996Z","iopub.status.idle":"2023-05-19T18:39:05.787216Z","shell.execute_reply.started":"2023-05-19T18:39:05.219934Z","shell.execute_reply":"2023-05-19T18:39:05.786014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that there is no significant difference in the distributions of minimum durations between tasks and FoG episodes for different event Types.","metadata":{}},{"cell_type":"markdown","source":"### The `Task_Min_Duration` column\nLet's analyze the frequency of task values with minimum durations for each event Type by creating barplots.","metadata":{}},{"cell_type":"code","source":"# Count the frequancy of the occurrence of each Task_Min_Duration value for each event Type\ntask_min_duration_counts = result_df.groupby(['Type'])['Task_Min_Duration'].value_counts(normalize=True).sort_values(ascending=False)\ntask_min_duration_counts = task_min_duration_counts.rename('Normalized_Count').reset_index()\ntask_min_duration_counts.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:05.788620Z","iopub.execute_input":"2023-05-19T18:39:05.789006Z","iopub.status.idle":"2023-05-19T18:39:05.810959Z","shell.execute_reply.started":"2023-05-19T18:39:05.788962Z","shell.execute_reply":"2023-05-19T18:39:05.810048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 2, figsize=(20, 15))\naxes = axes.flatten()\n\n# Plot the frequency of Task_Min_Duration occurrences for each event Type\nfor i, event_type in enumerate(task_min_duration_counts['Type'].unique()):\n    ax = axes[i]\n    data = task_min_duration_counts[task_min_duration_counts['Type'] == event_type]\n    sns.barplot(x='Task_Min_Duration', y='Normalized_Count', data=data, order=data['Task_Min_Duration'].unique(), ax=ax)\n    ax.set_xlabel('Task')\n    ax.set_ylabel('Frequency')\n    ax.set_title(f'Frequency of Task with min duration before FoG episode occurrences for \"{event_type}\"')\n    ax.tick_params(axis='x', rotation=40)\n    \n# Remove empty 4th subplot\nfig.delaxes(axes[3])\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:05.812551Z","iopub.execute_input":"2023-05-19T18:39:05.813578Z","iopub.status.idle":"2023-05-19T18:39:07.094082Z","shell.execute_reply.started":"2023-05-19T18:39:05.813543Z","shell.execute_reply":"2023-05-19T18:39:07.093202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the first subplot, which corresponds to the `StartHesitation` event Type, we observe that the most frequent task value with the minimum duration before FoG episodes is `Turning-DT`, followed by `MB12`, `Rest2`, `TUG-DT`.\n\nIn the second subplot, representing the `Turn` event Type, the most frequent task value with the minimum duration before FoG episodes is `Turning-ST`, followed by `TUG-DT`.\n\nThe third subplot corresponds to the `Walking` event type, where the most frequent task value with the minimum duration before FoG episodes is `Hotspot1`, followed by `TUG-ST`, `TUG-DT`, `Hotspot1-C`.","metadata":{}},{"cell_type":"markdown","source":"### Exploring the influence of `Preceding_Tasks` on FoG episodes\nIn this chapter we will try to identify the most common tasks that precede each event Type based on their frequency. First of all, let's calculate the normalized counts of the occurrence of each Task value in the `Preceding_Tasks` column for each event Type and make barplots that will show the frequency of preceding tasks for each event Type. ","metadata":{}},{"cell_type":"code","source":"# Explode the 'Preceding_Tasks' column and calculate the frequency of Preceding_Tasks for each event Type\npreceding_tasks_frequency = result_df.explode('Preceding_Tasks') \\\n                .groupby(['Type'])['Preceding_Tasks'].value_counts(normalize=True).sort_values(ascending=False)\n\npreceding_tasks_frequency = preceding_tasks_frequency.rename('Frequency').reset_index()\npreceding_tasks_frequency.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:07.095173Z","iopub.execute_input":"2023-05-19T18:39:07.095469Z","iopub.status.idle":"2023-05-19T18:39:07.139134Z","shell.execute_reply.started":"2023-05-19T18:39:07.095442Z","shell.execute_reply":"2023-05-19T18:39:07.138141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=2, ncols=2, figsize=(20, 15))\naxes = axes.flatten()\n\n# Plot the frequency of preceding tasks for each event Type\nfor i, event_type in enumerate(preceding_tasks_frequency['Type'].unique()):\n    data = preceding_tasks_frequency[preceding_tasks_frequency['Type'] == event_type]\n    ax = axes[i]\n    sns.barplot(x='Preceding_Tasks', y='Frequency', data=data, order=data['Preceding_Tasks'].unique(), ax=ax)\n    ax.set_xlabel('Task')\n    ax.set_ylabel('Frequency')\n    ax.set_title(f'Frequency of Preceding Tasks for \"{event_type}\" Type')\n    ax.tick_params(axis='x', rotation=50)\n\n# Remove empty 4th subplot\nfig.delaxes(axes[3])\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:07.140731Z","iopub.execute_input":"2023-05-19T18:39:07.141357Z","iopub.status.idle":"2023-05-19T18:39:09.023402Z","shell.execute_reply.started":"2023-05-19T18:39:07.141316Z","shell.execute_reply":"2023-05-19T18:39:09.021777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's retrieve the first five preceding tasks for each event type, which correspond to the most prevalent tasks that occur before each specific event type. By analyzing these prevalent tasks, we can gain insights into the common patterns and associations between tasks and event types.","metadata":{}},{"cell_type":"code","source":"result_dict = {'Turn': [], 'StartHesitation': [], 'Walking': []}\n\nfor i, type_val in enumerate(preceding_tasks_frequency['Type'].unique()):\n    data = preceding_tasks_frequency[preceding_tasks_frequency['Type'] == type_val]\n    result_dict[type_val].append(data.Preceding_Tasks[:6].values)\ntop_tasks_df = pd.DataFrame(result_dict)\ntop_tasks_df","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:09.025542Z","iopub.execute_input":"2023-05-19T18:39:09.026213Z","iopub.status.idle":"2023-05-19T18:39:09.042933Z","shell.execute_reply.started":"2023-05-19T18:39:09.026178Z","shell.execute_reply":"2023-05-19T18:39:09.041802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in top_tasks_df.columns:\n    print(col)\n    print(top_tasks_df[col].values)","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:09.044261Z","iopub.execute_input":"2023-05-19T18:39:09.044601Z","iopub.status.idle":"2023-05-19T18:39:09.059730Z","shell.execute_reply.started":"2023-05-19T18:39:09.044565Z","shell.execute_reply":"2023-05-19T18:39:09.058662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's compare the arrays of top preceding tasks for different event types. As we can see, the tasks `TUG-DT`, `4MW`, `TUG-ST`, `Turning-ST` consistently appear in the arrays for all event Types, which may indicate their significance in triggering freezing episodes overall. The `Turning-DT` task more frequently occurs before `StartHesitation` and `Walking` event types. Additionally, the task `Hotspot1` more often associated with the `Walking` event Type, while tasks`MB9`, `MB7` are more often frequently observed before `StartHesitation` event Type.","metadata":{}},{"cell_type":"markdown","source":"## Tasks feature extraction","metadata":{}},{"cell_type":"code","source":"# Read the data and append the Id\ndef read_data(path):\n    df = pd.read_csv(path)\n    df['Id'] = path.split(\"/\")[-1].split(\".\")[0]\n    \n    return df\n\n\n# Read and concatenate all files of the train data from specified dataset\ndef create_full_data(dataset_name):\n    \n    paths = glob.glob(data_dir + f'train/{dataset_name}/*')\n    final_df = pd.concat([read_data(p) for p in paths])\n    final_df['dataset'] = dataset_name\n    \n    return final_df","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:09.061044Z","iopub.execute_input":"2023-05-19T18:39:09.061376Z","iopub.status.idle":"2023-05-19T18:39:09.072148Z","shell.execute_reply.started":"2023-05-19T18:39:09.061349Z","shell.execute_reply":"2023-05-19T18:39:09.070826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a dataframe with all defog data\ndefog_data = create_full_data('defog')\n\n# Filter a dataframe by saving only valid defog data\ndefog_data_valid = defog_data.loc[ \\\n                        (defog_data.Valid == True) & (defog_data.Task == True)].copy()\n\ndefog_data_valid.reset_index(drop=True, inplace=True)\ndefog_data_valid.drop(['Valid', 'Task'], axis=1, inplace=True)\n\ndel defog_data","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:09.073584Z","iopub.execute_input":"2023-05-19T18:39:09.073897Z","iopub.status.idle":"2023-05-19T18:39:40.678068Z","shell.execute_reply.started":"2023-05-19T18:39:09.073871Z","shell.execute_reply":"2023-05-19T18:39:40.676797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_data_valid.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.679840Z","iopub.execute_input":"2023-05-19T18:39:40.680367Z","iopub.status.idle":"2023-05-19T18:39:40.707126Z","shell.execute_reply.started":"2023-05-19T18:39:40.680321Z","shell.execute_reply":"2023-05-19T18:39:40.705445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's make a list of tasks that are more prevalent before *all* Types of events.","metadata":{}},{"cell_type":"code","source":"top_freezing_tasks = ['TUG-DT', '4MW', 'TUG-ST', 'Turning-ST']","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.708845Z","iopub.execute_input":"2023-05-19T18:39:40.710009Z","iopub.status.idle":"2023-05-19T18:39:40.721993Z","shell.execute_reply.started":"2023-05-19T18:39:40.709969Z","shell.execute_reply":"2023-05-19T18:39:40.721090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below we are creating a function that can be used to check the association of a task with freezing episode based on top freezing tasks (that are common for all event types) list.\n","metadata":{}},{"cell_type":"code","source":"def is_freezing_task(row):\n    return row['Preceding_Tasks'] in top_freezing_tasks","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.723317Z","iopub.execute_input":"2023-05-19T18:39:40.724815Z","iopub.status.idle":"2023-05-19T18:39:40.738772Z","shell.execute_reply.started":"2023-05-19T18:39:40.724765Z","shell.execute_reply":"2023-05-19T18:39:40.737543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And here we are creating functions that can be used to check the association of a task with specific event Types.","metadata":{}},{"cell_type":"code","source":"def is_start_hes_task(row):\n    # Check if the task is associated with a Start Hesitation event\n    start_hesitation_tasks = ['MB6-R', 'Turning-DT', 'MB12', 'Rest2']\n    return row['Preceding_Tasks'] in start_hesitation_tasks\n\ndef is_turn_task(row):\n    # Check if the task is associated with a Turn event\n    turn_tasks = ['MB7', 'MB9']\n    return row['Preceding_Tasks'] in turn_tasks\n\ndef is_walking_task(row):\n    # Check if the task is associated with a Walking event\n    walking_tasks = ['Hotspot1', 'Hotspot1-C']\n    return row['Preceding_Tasks'] in walking_tasks","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.740722Z","iopub.execute_input":"2023-05-19T18:39:40.741254Z","iopub.status.idle":"2023-05-19T18:39:40.755314Z","shell.execute_reply.started":"2023-05-19T18:39:40.741211Z","shell.execute_reply":"2023-05-19T18:39:40.753898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exploded_tasks_df = result_df.explode('Preceding_Tasks')","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.756384Z","iopub.execute_input":"2023-05-19T18:39:40.756675Z","iopub.status.idle":"2023-05-19T18:39:40.782934Z","shell.execute_reply.started":"2023-05-19T18:39:40.756648Z","shell.execute_reply":"2023-05-19T18:39:40.782114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exploded_tasks_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.789069Z","iopub.execute_input":"2023-05-19T18:39:40.789422Z","iopub.status.idle":"2023-05-19T18:39:40.796880Z","shell.execute_reply.started":"2023-05-19T18:39:40.789392Z","shell.execute_reply":"2023-05-19T18:39:40.795396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exploded_tasks_df[['Id', 'Type', 'Preceding_Tasks']].duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.798620Z","iopub.execute_input":"2023-05-19T18:39:40.799034Z","iopub.status.idle":"2023-05-19T18:39:40.823646Z","shell.execute_reply.started":"2023-05-19T18:39:40.798996Z","shell.execute_reply":"2023-05-19T18:39:40.822763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exploded_tasks_df_drop_dupl = exploded_tasks_df[~(exploded_tasks_df[['Id', 'Type', 'Preceding_Tasks']].duplicated())]","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.824784Z","iopub.execute_input":"2023-05-19T18:39:40.825144Z","iopub.status.idle":"2023-05-19T18:39:40.841041Z","shell.execute_reply.started":"2023-05-19T18:39:40.825113Z","shell.execute_reply":"2023-05-19T18:39:40.839691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_data_valid['is_freezing_task'] = 0\ndefog_data_valid['is_start_hes_task'] = 0\ndefog_data_valid['is_turn_task'] = 0\ndefog_data_valid['is_walking_task'] = 0","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.842534Z","iopub.execute_input":"2023-05-19T18:39:40.843793Z","iopub.status.idle":"2023-05-19T18:39:40.869718Z","shell.execute_reply.started":"2023-05-19T18:39:40.843732Z","shell.execute_reply":"2023-05-19T18:39:40.868703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's iterate over each non-duplicated row in the `exploded_tasks_df_drop_dupl` dataset and determine whether a given preceding task is one of the top freezing tasks or one of the three specific task types, and update the `defog_data_valid` dataframe by setting flags to indicate the occurrence of specific task types during specific event timings.","metadata":{}},{"cell_type":"code","source":"for _, row in exploded_tasks_df_drop_dupl.iterrows():\n    event_id = row['Id']\n    start_time = round(row['Init'] * 100)\n    end_time = round(row['Completion'] * 100)\n    \n    # Check for freezing task based on top_freezing_tasks list\n    if is_freezing_task(row):\n        defog_data_valid['is_freezing_task'] = np.where(\n            (defog_data_valid.Id == event_id) & \n            (defog_data_valid['Time'] >= start_time) &\n            (defog_data_valid['Time'] <= end_time),\n            1,\n            defog_data_valid['is_freezing_task']\n        )\n    \n    # Check if task is associated with one of the three types of FoG events\n    if is_start_hes_task(row):    \n        defog_data_valid['is_start_hes_task'] = np.where(\n            (defog_data_valid.Id == event_id) & \n            (defog_data_valid['Time'] >= start_time) &\n            (defog_data_valid['Time'] <= end_time),\n            1,\n            defog_data_valid['is_start_hes_task']\n        )\n    \n    if is_turn_task(row):\n        defog_data_valid['is_turn_task'] = np.where(\n            (defog_data_valid.Id == event_id) & \n            (defog_data_valid['Time'] >= start_time) &\n            (defog_data_valid['Time'] <= end_time),\n            1,\n            defog_data_valid['is_turn_task']\n        )\n    \n    if is_walking_task(row):\n        defog_data_valid['is_walking_task'] = np.where(\n            (defog_data_valid.Id == event_id) & \n            (defog_data_valid['Time'] >= start_time) &\n            (defog_data_valid['Time'] <= end_time),\n            1,\n            defog_data_valid['is_walking_task']\n        )","metadata":{"execution":{"iopub.status.busy":"2023-05-19T18:39:40.871146Z","iopub.execute_input":"2023-05-19T18:39:40.872215Z","iopub.status.idle":"2023-05-19T18:52:01.886274Z","shell.execute_reply.started":"2023-05-19T18:39:40.872168Z","shell.execute_reply":"2023-05-19T18:52:01.884654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's validate our features.","metadata":{}},{"cell_type":"code","source":"defog_data_valid[['is_freezing_task', 'Turn', 'Walking', 'StartHesitation']].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T19:00:07.866870Z","iopub.execute_input":"2023-05-19T19:00:07.868243Z","iopub.status.idle":"2023-05-19T19:00:08.278667Z","shell.execute_reply.started":"2023-05-19T19:00:07.868198Z","shell.execute_reply":"2023-05-19T19:00:08.277658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_data_valid[['is_turn_task', 'Turn']].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T19:01:28.094662Z","iopub.execute_input":"2023-05-19T19:01:28.095637Z","iopub.status.idle":"2023-05-19T19:01:28.312428Z","shell.execute_reply.started":"2023-05-19T19:01:28.095596Z","shell.execute_reply":"2023-05-19T19:01:28.311357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that features `is_freezing_task`, `is_turn_task` are providing valuable information about the target values. Their performance is not entirely perfect, but incorporating these features enables us to make more accurate predictions.","metadata":{}},{"cell_type":"code","source":"defog_data_valid[['is_walking_task', 'Walking']].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T19:04:43.042683Z","iopub.execute_input":"2023-05-19T19:04:43.043208Z","iopub.status.idle":"2023-05-19T19:04:43.273855Z","shell.execute_reply.started":"2023-05-19T19:04:43.043172Z","shell.execute_reply":"2023-05-19T19:04:43.272524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_data_valid[['is_start_hes_task', 'StartHesitation']].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-19T19:05:24.329261Z","iopub.execute_input":"2023-05-19T19:05:24.329921Z","iopub.status.idle":"2023-05-19T19:05:24.577296Z","shell.execute_reply.started":"2023-05-19T19:05:24.329869Z","shell.execute_reply":"2023-05-19T19:05:24.575891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears that features `is_walking_task`, `is_start_hes_task` are not producing the expected results as the values of those features are not align well with the actual target values. ","metadata":{}}]}