{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# # This Python 3 environment comes with many helpful analytics libraries installed\n# # It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# # For example, here's several helpful packages to load\n\n# import numpy as np # linear algebra\n# import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# # Input data files are available in the read-only \"../input/\" directory\n# # For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n# import os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# # You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# # You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-14T11:39:20.127062Z","iopub.execute_input":"2023-03-14T11:39:20.127538Z","iopub.status.idle":"2023-03-14T11:39:20.134826Z","shell.execute_reply.started":"2023-03-14T11:39:20.127494Z","shell.execute_reply":"2023-03-14T11:39:20.133405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dataset Description and checking the files in each folder\n\n## Dataset Description\nThis competition dataset comprises lower-back 3D accelerometer data from subjects exhibiting freezing of gait episodes, a disabling symptom that is common among people with Parkinson's disease. Freezing of gait (FOG) negatively impacts walking abilities and impinges locomotion and independence.\n\nYour objective is to detect the start and stop of each freezing episode and the occurrence in these series of three types of freezing of gait events: Start Hesitation, Turn, and Walking.\n\n# The Datasets\nThe data series include three datasets, collected under distinct circumstances:\n\n- The **tDCS FOG (tdcsfog) dataset**, comprising **data series collected in the lab**, as subjects completed a FOG-provoking protocol.\n- The **DeFOG (defog) dataset**, comprising **data series collected in the subject's home**, as subjects completed a FOG-provoking protocol\n- The **Daily Living (daily) dataset**, comprising **one week of continuous 24/7 recordings from sixty-five subjects**. Forty-five subjects exhibit FOG symptoms and also have series in the defog dataset, while the other twenty subjects do not exhibit FOG symptoms and do not have series elsewhere in the data.\n- Trials from the **tdcsfog and defog datasets** were videotaped and **annotated** by expert reviewers documented the freezing of gait episodes. That is, the start, end and type of each episode were marked by the experts. Series in the **daily dataset are unannotated**. You will be detecting FOG episodes for the tdcsfog and defog series. You may wish to apply unsupervised or semi-supervised methods to the series in the daily dataset to support your detection modelling.\n\n## Checking the files in each folder\n There are 3 folders :\n*  train (defog/tdcsfog/notype) -> csv files\n*  test (defog/tdcsfog) -> csv files\n*  unlabeled -> parquest files\n\nThere are also other files on metadata/sample submission ['sample_submission.csv','subjects.csv', 'tasks.csv', 'defog_metadata.csv', 'daily_metadata.csv',  'events.csv', 'tdcsfog_metadata.csv']\n","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\n\npath = \"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/\"\nprint(f'The file/subfolders in {path} are: {os.listdir(path)}')\n\nfolders = [\"train\", \"test\", \"unlabeled\"]\nfor folder in folders:\n    subfolder_list = os.listdir(os.path.join(path, folder))\n    print(\"Folder:\",folder, subfolder_list)\n    for subfolder in subfolder_list: \n        if os.path.isdir(os.path.join(path, folder,subfolder)):\n            print(\"Subfolder:\", subfolder)\n            print(f'The no. of files in {os.path.join(path, folder,subfolder)} are: {len(os.listdir(os.path.join(path, folder,subfolder)))}')\n#             print(f'The files in {os.path.join(path, folder,subfolder)} are: {os.listdir(os.path.join(path, folder,subfolder))}')\n","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:39:20.137517Z","iopub.execute_input":"2023-03-14T11:39:20.138514Z","iopub.status.idle":"2023-03-14T11:39:20.223758Z","shell.execute_reply.started":"2023-03-14T11:39:20.138454Z","shell.execute_reply":"2023-03-14T11:39:20.222558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Check how the submission file looks like\n- sample_submission.csv: A submission file in the correct format.\n- Contains id and predictions for \"StartHesitation\",\"Turn\",\"Walking\"\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\npath = \"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/\"\nsubmission = pd.read_csv(path + \"sample_submission.csv\")\ndisplay(submission.head())\nprint(submission[\"StartHesitation\"].value_counts())\nprint(submission[\"Turn\"].value_counts())\nprint(submission[\"Walking\"].value_counts())","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:39:20.225311Z","iopub.execute_input":"2023-03-14T11:39:20.225979Z","iopub.status.idle":"2023-03-14T11:39:20.452654Z","shell.execute_reply.started":"2023-03-14T11:39:20.225941Z","shell.execute_reply":"2023-03-14T11:39:20.451478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### EDA of train data\n* **train/ Folder** containing the data series in the training set within three subfolders: tdcsfog/, defog/, and notype/. Series in the notype folder are from the defog dataset but lack event-type annotations. The fields present in these series vary by folder.\n* **Time** An integer timestep. Series from the tdcsfog dataset are recorded at 128Hz (128 timesteps per second), while series from the defog and daily series are recorded at 100Hz (100 timesteps per second).\n* **AccV, AccML, and AccAP** Acceleration in units of g, from a lower-back sensor on three axes: V - vertical, ML - mediolateral, AP - anteroposterior.\n* **StartHesitation, Turn, Walking Indicator** variables for the occurrence of each of the event types.\n* **Event Indicator** variable for the occurrence of any FOG-type event. Present only in the notype series, which lack type-level annotations.\n* **Valid** There were cases during the video annotation that were hard for the annotator to decide if there was an Akinetic (i.e., essentially no movement) FoG or the subject stopped voluntarily. Only event annotations where the series is marked true should be considered as unambiguous.\n* **Task** Series were only annotated where this value is true. Portions marked false should be considered unannotated.\n* Note that the **Valid and Task fields are only present in the defog dataset**. They are not relevant for the tdcsfog data.","metadata":{}},{"cell_type":"markdown","source":"#### a) EDA of defog train","metadata":{}},{"cell_type":"code","source":"#read all the files in train defog folder and concat them\n#create column \"File Name\" which is the file name of each row\ntrain_path = path +\"train/\"\n\nimport pandas as pd\nimport os\nimport glob\n\n# Set the directory path\ndirectory_path = os.path.join(train_path,'defog')\n\n# Get all the CSV files in the directory\nfiles = glob.glob(os.path.join(directory_path, \"*.csv\"))\n\n# Create an empty DataFrame\ntrain_defog = pd.DataFrame()\n\n# Loop through each file and append the data to the DataFrame\nfor file in files:\n    # Read the CSV file into a DataFrame\n    temp_df = pd.read_csv(file)\n\n    # Add a new column to the DataFrame with the file name\n    temp_df[\"File Name\"] = os.path.splitext(os.path.basename(file))[0]\n\n    # Append the data to the main DataFrame\n    train_defog = train_defog.append(temp_df, ignore_index=True)\n\n# Print the DataFrame\ndisplay(train_defog.head())\n","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:39:20.455608Z","iopub.execute_input":"2023-03-14T11:39:20.455969Z","iopub.status.idle":"2023-03-14T11:39:59.791638Z","shell.execute_reply.started":"2023-03-14T11:39:20.455933Z","shell.execute_reply":"2023-03-14T11:39:59.790129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Split train_defog data into unambiguos (valid=True) and ambiguous (valid=False) data to study them separately. If valid is True, \"Task\" value should be True as well which means that there is annotation","metadata":{}},{"cell_type":"code","source":"print(\"Train_defog distribution for columns Valid\",train_defog[\"Valid\"].value_counts())\ntrain_defog_unambiguos = train_defog[train_defog[\"Valid\"]== True]\ntrain_defog_ambiguous = train_defog[train_defog[\"Valid\"]== False]\nprint(\"train_defog_unambiguos\",train_defog_unambiguos[\"Valid\"].value_counts())\nprint(\"train_defog_ambiguous\",train_defog_ambiguous[\"Valid\"].value_counts())","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:39:59.793834Z","iopub.execute_input":"2023-03-14T11:39:59.794261Z","iopub.status.idle":"2023-03-14T11:40:01.371981Z","shell.execute_reply.started":"2023-03-14T11:39:59.794223Z","shell.execute_reply":"2023-03-14T11:40:01.370674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### EDA on train_defog unambiguous data using pandas-profiling. \n* ##### There are no duplicates or missing values. \n* ##### Features are not highly correlated with each other\n* ##### \"StartHesitation\",\"Turn\",\"Walking\" are highly imbalanced with more 0 than 1","metadata":{}},{"cell_type":"code","source":"#EDA on unambiguous data using pandas-profiling\nfrom pandas_profiling import ProfileReport\ntrain_defog_unambiguos_report = ProfileReport(train_defog_unambiguos, title = \"train_defog_unambiguous\",explorative=True)\ntrain_defog_unambiguos_report","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:40:01.373664Z","iopub.execute_input":"2023-03-14T11:40:01.374503Z","iopub.status.idle":"2023-03-14T11:42:16.731592Z","shell.execute_reply.started":"2023-03-14T11:40:01.374462Z","shell.execute_reply":"2023-03-14T11:42:16.729852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the correlation between features\nimport matplotlib.pyplot as plt\nimport seaborn as sns\ncorr_matrix = train_defog_unambiguos.drop([\"Valid\",\"Task\"],axis=1).corr()\n\n# Plot the correlation heatmap with values\nplt.figure(figsize=(10, 8))\nsns.heatmap(corr_matrix, annot=True, cmap=\"coolwarm\")\nplt.title(\"Correlation Heatmap of train_defog_unambiguos\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:42:16.733664Z","iopub.execute_input":"2023-03-14T11:42:16.734526Z","iopub.status.idle":"2023-03-14T11:42:18.165540Z","shell.execute_reply.started":"2023-03-14T11:42:16.734483Z","shell.execute_reply":"2023-03-14T11:42:18.164227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Check distribution of \"StartHesitation\",\"Turn\",\"Walking\"\nprint(train_defog_unambiguos[\"StartHesitation\"].value_counts())\nprint(train_defog_unambiguos[\"Turn\"].value_counts())\nprint(train_defog_unambiguos[\"Walking\"].value_counts())","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:42:18.167406Z","iopub.execute_input":"2023-03-14T11:42:18.167904Z","iopub.status.idle":"2023-03-14T11:42:18.282082Z","shell.execute_reply.started":"2023-03-14T11:42:18.167851Z","shell.execute_reply":"2023-03-14T11:42:18.280931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### EDA of DeFOG (defog) dataset, comprising data series collected in the subject's home, as subjects completed a FOG-provoking protocol\n\nTo be continued.. kindly give an upvote if you find this notebook useful","metadata":{"execution":{"iopub.status.busy":"2023-03-12T05:31:40.620054Z","iopub.execute_input":"2023-03-12T05:31:40.621235Z","iopub.status.idle":"2023-03-12T05:31:40.626183Z","shell.execute_reply.started":"2023-03-12T05:31:40.621171Z","shell.execute_reply":"2023-03-12T05:31:40.624974Z"}}},{"cell_type":"markdown","source":"## Metadata\n 1) subjects.csv Metadata for each Subject in the study, including their Age and Sex as well as:\n* Visit - Only available for subjects in the daily and defog datasets.\n* Years - SinceDx Years since Parkinson's diagnosis.\n* UPDRSIIIOn/UPDRSIIIOff -  Unified Parkinson's Disease Rating Scale score during on/off medication respectively.\n* NFOGQ  - Self-report FoG questionnaire score. See: https://pubmed.ncbi.nlm.nih.gov/19660949/\n\n2) tasks.csv Task metadata for series in the defog dataset. (Not relevant for the series in the fog or daily datasets.)\n\n* Id -  The data series where the task was measured.\n* Begin Time (s)  -  the task began.\n* End Time (s) -  the task ended.\n* Task -  One of seven tasks types in the DeFOG protocol, described on this page.\n* Description -  Description of the task.\n\n3) defog_metadata.csv Identifies each series in the defog dataset by a unique Subject, Visit, Medication condition.\n\n4) daily_metadata.csv Each series in the daily dataset is identified by the Subject id. This file also contains the time of day the recording began.\n\n5) events.csv Metadata for each FoG event in all data series. The event times agree with the labels in the data series.\n\n* Id -   The data series the event occured in.\n* Init Time (s) -   the event began.\n* Completion Time (s) -   the event ended.\n* Type -   Whether StartHesitation, Turn, or Walking.\n* Kinetic  -  Whether the event was kinetic (1) and involved movement, or akinetic (0) and static.","metadata":{}},{"cell_type":"markdown","source":"#### Overview of metadata ","metadata":{}},{"cell_type":"code","source":"#Overview of data \nsubjects = pd.read_csv(path + \"subjects.csv\")\ntasks = pd.read_csv(path + \"tasks.csv\")\ndefog_metadata = pd.read_csv(path + \"defog_metadata.csv\")\ndaily_metadata = pd.read_csv(path + \"daily_metadata.csv\")\nevents = pd.read_csv(path + \"events.csv\")\n\nprint(\"1) subjects:\")\ndisplay(subjects.head())\nprint(subjects.shape)\nprint(\"2) tasks:\")\ndisplay(tasks.head())\nprint(tasks.shape)\nprint(\"3) defog_metadata:\")\ndisplay(defog_metadata.head())\nprint(defog_metadata.shape)\nprint(\"4) daily_metadata:\")\ndisplay(daily_metadata.head())\nprint(daily_metadata.shape)\nprint(\"5) events:\")\ndisplay(events.head())\nprint(events.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-14T11:42:18.283409Z","iopub.execute_input":"2023-03-14T11:42:18.283988Z","iopub.status.idle":"2023-03-14T11:42:18.359018Z","shell.execute_reply.started":"2023-03-14T11:42:18.283951Z","shell.execute_reply":"2023-03-14T11:42:18.358139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### EDA of metadata","metadata":{}}]}