{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# goal:\nThis notebook will perform some basic EDA insights about datasets. \nApart from profile report, fast_eda gives a better and fast insight into the basic data structure. \nCheck it out...","metadata":{"execution":{"iopub.status.busy":"2023-05-05T08:25:57.009083Z","iopub.execute_input":"2023-05-05T08:25:57.009481Z","iopub.status.idle":"2023-05-05T08:26:01.254813Z","shell.execute_reply.started":"2023-05-05T08:25:57.009454Z","shell.execute_reply":"2023-05-05T08:26:01.253604Z"}}},{"cell_type":"code","source":"%%capture\n#!pip install ydata_profiling\n!pip install fasteda","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:37.034425Z","iopub.execute_input":"2023-05-06T13:05:37.034925Z","iopub.status.idle":"2023-05-06T13:05:50.455176Z","shell.execute_reply.started":"2023-05-06T13:05:37.034888Z","shell.execute_reply":"2023-05-06T13:05:50.453701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport ydata_profiling\n#from ydata_profiling import ProfileReport\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom ydata_profiling import ProfileReport\nfrom fasteda import fast_eda\nfrom phik.phik import phik_matrix\nfrom phik.report import plot_correlation_matrix","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:50.458363Z","iopub.execute_input":"2023-05-06T13:05:50.458882Z","iopub.status.idle":"2023-05-06T13:05:50.815072Z","shell.execute_reply.started":"2023-05-06T13:05:50.458836Z","shell.execute_reply":"2023-05-06T13:05:50.814041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#loading data\ndata_dir = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/'\n# \ndaily_metadata_file = f'{data_dir}daily_metadata.csv'\ndefog_metadata_file = f'{data_dir}defog_metadata.csv'\ntdcsfog_metadata_file = f'{data_dir}tdcsfog_metadata.csv'\nevents_data_file = f'{data_dir}events.csv'\nsubjects_data_file = f'{data_dir}subjects.csv'\ntasks_data_file = f'{data_dir}tasks.csv'\n\n# Read the meta data\ndaily_metadata = pd.read_csv(daily_metadata_file)\ndefog_metadata = pd.read_csv(defog_metadata_file)\ntdcsfog_metadata = pd.read_csv(tdcsfog_metadata_file)\n\nevents_data = pd.read_csv(events_data_file)\nsubjects_data = pd.read_csv(subjects_data_file)\ntasks_data = pd.read_csv(tasks_data_file)","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:50.816454Z","iopub.execute_input":"2023-05-06T13:05:50.817458Z","iopub.status.idle":"2023-05-06T13:05:50.879479Z","shell.execute_reply.started":"2023-05-06T13:05:50.817425Z","shell.execute_reply":"2023-05-06T13:05:50.878357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Events data\n\nEvents data has 5 features and 3544 instances\n\n  1). Id - 535 unique ids.\n  \n  2). Init-  Time (s) the event began.\n  \n  3). Completion Time-  (s) the event ended.\n  \n  4). Type- Three types: StartHesitation, Turn, or Walking.\n  \n  5). Kinetic-  Whether the event was kinetic (1) and involved movement, or akinetic (0) and static.\n\nThe Fast_eda analysis of events data have the following finding:\n1. The number of FoG events of 'Turn' type is higher than other types.\n2. The number of 'Kinetic' events involving movement out number 'STATIC'.\n3. Missing values 30%  (1045/3544) in both types and Kinetics columns\n","metadata":{}},{"cell_type":"code","source":"events_data.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:50.882225Z","iopub.execute_input":"2023-05-06T13:05:50.882555Z","iopub.status.idle":"2023-05-06T13:05:50.903485Z","shell.execute_reply.started":"2023-05-06T13:05:50.882526Z","shell.execute_reply":"2023-05-06T13:05:50.902406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"events_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:50.904832Z","iopub.execute_input":"2023-05-06T13:05:50.905118Z","iopub.status.idle":"2023-05-06T13:05:50.911971Z","shell.execute_reply.started":"2023-05-06T13:05:50.905094Z","shell.execute_reply":"2023-05-06T13:05:50.910922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fast_eda(events_data)","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:50.913256Z","iopub.execute_input":"2023-05-06T13:05:50.913585Z","iopub.status.idle":"2023-05-06T13:05:55.544955Z","shell.execute_reply.started":"2023-05-06T13:05:50.913557Z","shell.execute_reply":"2023-05-06T13:05:55.543821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# subjects_data\nMetadata for each Subject in the study, has 178 instances and 8 variables:\n\n1. Subject - 136 unique ids of subjects\n2. Visit- 1.0 or 2.0, only available for subjects in the daily and defog datasets.\n3. Years- since time of diagnosis.\n4. Age- from 30 years to 90 years.\n5. SEX- Majprity are males (121 males vs 52 females).\n6. UPDRSIIIOn- score\n7. UPDRSIIIOff Unified Parkinson's Disease Rating Scale score during on/off medication respectively.\n8. NFOGQ Self-report FoG questionnaire score.\n\nAs can be seen from the fast_eda report, \n1. the number of male subjects is greater than female subjects.\n2. The age column seems to have a normal distribution based on the histogram, skewness and kurtosis values. \n3. The majority of patients in our dataset are aged between 56 and 79 years old (the average age is around 68). \n4. Column UPDRSIII_On is highly correlated with UPDRSIII_Off column. \n5. Columns Visit and UPDRSIII_Off have missing values. \n6. NFOGQ column has 12.1% zeros.","metadata":{}},{"cell_type":"code","source":"subjects_data.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:55.546658Z","iopub.execute_input":"2023-05-06T13:05:55.547715Z","iopub.status.idle":"2023-05-06T13:05:55.563041Z","shell.execute_reply.started":"2023-05-06T13:05:55.547670Z","shell.execute_reply":"2023-05-06T13:05:55.561566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fast_eda(subjects_data)","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:05:55.564351Z","iopub.execute_input":"2023-05-06T13:05:55.565697Z","iopub.status.idle":"2023-05-06T13:06:06.787753Z","shell.execute_reply.started":"2023-05-06T13:05:55.565653Z","shell.execute_reply":"2023-05-06T13:06:06.786660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# tasks_data:\nTask metadata for series in the defog dataset. (Not relevant for the series in the tdcsfog or daily datasets.) It has 4 varaibles and 2817 instances.\n\n1. Id: The data series where the task was measured.\n2. Begin Time (s) the task began.\n3. End Time (s) the task ended.\n4. Task One of seven tasks types in the DeFOG protocol, described on this page.\n","metadata":{}},{"cell_type":"code","source":"tasks_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:06:06.789455Z","iopub.execute_input":"2023-05-06T13:06:06.790004Z","iopub.status.idle":"2023-05-06T13:06:06.797022Z","shell.execute_reply.started":"2023-05-06T13:06:06.789965Z","shell.execute_reply":"2023-05-06T13:06:06.796122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tasks_data.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:06:06.800157Z","iopub.execute_input":"2023-05-06T13:06:06.800993Z","iopub.status.idle":"2023-05-06T13:06:06.833873Z","shell.execute_reply.started":"2023-05-06T13:06:06.800963Z","shell.execute_reply":"2023-05-06T13:06:06.832538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fast_eda(tasks_data)","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:06:06.835858Z","iopub.execute_input":"2023-05-06T13:06:06.836169Z","iopub.status.idle":"2023-05-06T13:06:09.075187Z","shell.execute_reply.started":"2023-05-06T13:06:06.836143Z","shell.execute_reply":"2023-05-06T13:06:09.073893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Full metadata\nLet's concatenate the tdcsfog metadata with the defog metadata, and then merge the result with the subjects dataset for further analysis.","metadata":{}},{"cell_type":"code","source":"tdcsfog_metadata['dataset'] = 'tdcsfog'\ndefog_metadata['dataset'] = 'defog'\n\nfull_metadata = pd.concat([tdcsfog_metadata, defog_metadata])","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:06:09.077348Z","iopub.execute_input":"2023-05-06T13:06:09.077814Z","iopub.status.idle":"2023-05-06T13:06:09.086882Z","shell.execute_reply.started":"2023-05-06T13:06:09.077771Z","shell.execute_reply":"2023-05-06T13:06:09.085584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(full_metadata.shape)\nfull_metadata.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:06:09.088987Z","iopub.execute_input":"2023-05-06T13:06:09.089297Z","iopub.status.idle":"2023-05-06T13:06:09.111974Z","shell.execute_reply.started":"2023-05-06T13:06:09.089271Z","shell.execute_reply":"2023-05-06T13:06:09.110912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_metadata.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:33:58.882286Z","iopub.execute_input":"2023-05-06T13:33:58.882782Z","iopub.status.idle":"2023-05-06T13:33:58.893786Z","shell.execute_reply.started":"2023-05-06T13:33:58.882750Z","shell.execute_reply":"2023-05-06T13:33:58.892945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fast_eda(full_metadata)","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:06:09.113561Z","iopub.execute_input":"2023-05-06T13:06:09.113911Z","iopub.status.idle":"2023-05-06T13:06:12.392614Z","shell.execute_reply.started":"2023-05-06T13:06:09.113877Z","shell.execute_reply":"2023-05-06T13:06:12.391099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The train Datasets \n1. The lab data (tdcsfog Dataset)\nThe tDCS FOG (tdcsfog) dataset comprises data series collected in the lab, as subjects completed a FOG-provoking protocol.\n\nLet's select only 500 files from the tdcsfog folder and take a look at the dataset profile report and correlation matrix.","metadata":{}},{"cell_type":"code","source":"# Read 500 files of the tdcsfog metadata (the total number of files is 833)\ntdcsfog_files = [f'{data_dir}train/tdcsfog/{id}.csv' for id in \\\n                                         tdcsfog_metadata.Id.to_list()[:501]]\n\ntdcsfog_data = pd.concat([pd.read_csv(file) for file in tdcsfog_files])","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:37:40.762326Z","iopub.execute_input":"2023-05-06T13:37:40.763482Z","iopub.status.idle":"2023-05-06T13:37:50.452166Z","shell.execute_reply.started":"2023-05-06T13:37:40.763441Z","shell.execute_reply":"2023-05-06T13:37:50.451036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tdcsfog_data.nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fast_eda(tdcsfog_data)","metadata":{"execution":{"iopub.status.busy":"2023-05-06T13:38:00.676773Z","iopub.execute_input":"2023-05-06T13:38:00.677156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}