{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color:\n    #f78102;\"> Parkinson's Freezing of Gait Prediction</span></b></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction</h2>\n\n[Previously](https://www.kaggle.com/code/dinowun/eda-simplified-amp-pd-progression-prediction), we went through how Parkinson's disease was first described as \"shaking palsy\" by James Parkinson, how Parkinson's disease worsen over time, and how it was uncurable in the near future. However, once the new competition based on Parkinson's disease was released, we discovered that one of the symptoms based on this disease is Freezing of Gait (aka FOG for short). Speaking of that symptom, Once a patient has a \"FOG\" episode, they were glued in the ground, remaining stationary and despite their attempts, they can't move forward. Furthermore, FOG episodes has a profound negative impact on health-related quality of life—people who suffer from FOG, which makes them often depressed, as well as having an increased risk of falling, which are likelier to be confined to wheelchair use, and have restricted independence. On the bright side, there are many possibilities on evaluating FOG although most of them were FOG-provoking protocols. Basically, a person with FOG were filmed while performing everyday tasks, then an expert review each frame to score them, whether it occurred to them. Nonetheless, scoring each frame takes time and requires specific expertise. In spite of that, machine-learning models can boost the accuracy of detecting FOG from a lower back accelerometer is relatively high, as it can lead to potential treatments. Without a doubt, The Michael J. Fox Foundation hosted this competition for detecting the FOG events from wearable sensor data. And with that being said, let's take off into data visualization on the FOG predictions!\n\n<center>\n    <img src=\"https://www.frontiersin.org/files/Articles/828355/fnhum-16-828355-HTML/image_m/fnhum-16-828355-g001.jpg\" width=500>\n    <figcaption style=\"color: gray;\">A diagram about the progressing Freezing of Gait pattern in a person. Credit: Frontiers</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Setting Up</h2>\n\nTo titillate on our data analysis in our typical EDA notebook, we first import the important pandas module as pd and the numpy module as np (just in case) for data analysis and possibly linear algebra. After that, we then load the modules for plotting such as the plotly module as px and the altair module as alt (along with disabling Altair's MaxRowsError with the alt module's data_transformers attribute's enable function that is set to 'default' as well as None for the max_rows parameter). ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport plotly.express as px\nimport altair as alt\n\nalt.data_transformers.enable('default', max_rows=None)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:23.012829Z","iopub.execute_input":"2023-04-28T06:19:23.013764Z","iopub.status.idle":"2023-04-28T06:19:25.986160Z","shell.execute_reply.started":"2023-04-28T06:19:23.013715Z","shell.execute_reply":"2023-04-28T06:19:25.984774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After loading the modules, we realized that just like some other competitions ([like the ASL one](https://www.kaggle.com/competitions/asl-signs)), the data in the FOG prediction one contains csv and parquet files, but as we get back into the topic of creating dataframes, let's create several dataframes from the CSV files and from the parquet files later on in the next, next section. Without saying this further, we create three dataframes: daily_df, defog_df, and tdcsfog_df from reading the csv files: daily_metadata.csv, defog_metadata.csv, and tdcsfog_metadata.csv with the pd module's read_csv function, as well as using it to read out the events.csv, subjects.csv, and tasks.csv files for creating additional three dataframes: events_df, subjects_df, and tasks_df. With all six dataframes created, we display the top five rows of each dataframe with the head function.","metadata":{}},{"cell_type":"code","source":"# Daily, Defog, and Tdcsfog data:\ndaily_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/daily_metadata.csv\")\ndefog_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/defog_metadata.csv\")\ntdcsfog_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/tdcsfog_metadata.csv\")\n\n# Other three dataframes:\nevents_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/events.csv\")\nsubjects_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/subjects.csv\")\ntasks_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/tasks.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:25.989130Z","iopub.execute_input":"2023-04-28T06:19:25.990070Z","iopub.status.idle":"2023-04-28T06:19:26.055864Z","shell.execute_reply.started":"2023-04-28T06:19:25.990011Z","shell.execute_reply":"2023-04-28T06:19:26.054543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"daily_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.057403Z","iopub.execute_input":"2023-04-28T06:19:26.057893Z","iopub.status.idle":"2023-04-28T06:19:26.086820Z","shell.execute_reply.started":"2023-04-28T06:19:26.057841Z","shell.execute_reply":"2023-04-28T06:19:26.085329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.092105Z","iopub.execute_input":"2023-04-28T06:19:26.092527Z","iopub.status.idle":"2023-04-28T06:19:26.109083Z","shell.execute_reply.started":"2023-04-28T06:19:26.092485Z","shell.execute_reply":"2023-04-28T06:19:26.107887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tdcsfog_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.110899Z","iopub.execute_input":"2023-04-28T06:19:26.111997Z","iopub.status.idle":"2023-04-28T06:19:26.128958Z","shell.execute_reply.started":"2023-04-28T06:19:26.111931Z","shell.execute_reply":"2023-04-28T06:19:26.127968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"events_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.132778Z","iopub.execute_input":"2023-04-28T06:19:26.133194Z","iopub.status.idle":"2023-04-28T06:19:26.149850Z","shell.execute_reply.started":"2023-04-28T06:19:26.133155Z","shell.execute_reply":"2023-04-28T06:19:26.148474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subjects_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.151573Z","iopub.execute_input":"2023-04-28T06:19:26.153030Z","iopub.status.idle":"2023-04-28T06:19:26.171350Z","shell.execute_reply.started":"2023-04-28T06:19:26.152967Z","shell.execute_reply":"2023-04-28T06:19:26.169811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tasks_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.172918Z","iopub.execute_input":"2023-04-28T06:19:26.173342Z","iopub.status.idle":"2023-04-28T06:19:26.188829Z","shell.execute_reply.started":"2023-04-28T06:19:26.173289Z","shell.execute_reply":"2023-04-28T06:19:26.187378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Analysis on Dataframe Metadata</h3>\n\nOnce we generate our six dataframes, let's visualize their basic surface metadata! To begin with our metadata analysis, let's find the number of the overall data entities in each dataframe by encasing them with the len function.","metadata":{}},{"cell_type":"code","source":"len(daily_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.190635Z","iopub.execute_input":"2023-04-28T06:19:26.191075Z","iopub.status.idle":"2023-04-28T06:19:26.199762Z","shell.execute_reply.started":"2023-04-28T06:19:26.191034Z","shell.execute_reply":"2023-04-28T06:19:26.198615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(defog_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.205300Z","iopub.execute_input":"2023-04-28T06:19:26.206137Z","iopub.status.idle":"2023-04-28T06:19:26.214485Z","shell.execute_reply.started":"2023-04-28T06:19:26.206066Z","shell.execute_reply":"2023-04-28T06:19:26.213133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(tdcsfog_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.216812Z","iopub.execute_input":"2023-04-28T06:19:26.217551Z","iopub.status.idle":"2023-04-28T06:19:26.226534Z","shell.execute_reply.started":"2023-04-28T06:19:26.217500Z","shell.execute_reply":"2023-04-28T06:19:26.225587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(events_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.228239Z","iopub.execute_input":"2023-04-28T06:19:26.229455Z","iopub.status.idle":"2023-04-28T06:19:26.238254Z","shell.execute_reply.started":"2023-04-28T06:19:26.229414Z","shell.execute_reply":"2023-04-28T06:19:26.237060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(subjects_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.240273Z","iopub.execute_input":"2023-04-28T06:19:26.241270Z","iopub.status.idle":"2023-04-28T06:19:26.249392Z","shell.execute_reply.started":"2023-04-28T06:19:26.241227Z","shell.execute_reply":"2023-04-28T06:19:26.248129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(tasks_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.251005Z","iopub.execute_input":"2023-04-28T06:19:26.251367Z","iopub.status.idle":"2023-04-28T06:19:26.262406Z","shell.execute_reply.started":"2023-04-28T06:19:26.251331Z","shell.execute_reply":"2023-04-28T06:19:26.261112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From each data entity counts calculated from 6 dataframes, we counted 65 data units in the daily_df dataframe, 137 in the defog_df dataframe, 833 in the tdcsfog_df dataframe, 3712 in the the events_df dataframe, 173 in the subjects_df dataframe, and 2817 in the tasks_df dataframe. In other words, the less data entities in the daily_df, defog_df, tdcsfog_df, and subjects_df dataframes hinted us that these dataframes cover the background data for the machine learning model to predict the FOG progression.","metadata":{}},{"cell_type":"markdown","source":"Now let's count the missing values in each dataframe-at-a-time! We simply plug the isna function to each of 6 dataframes to find the overall nan values, thus summing the counted NaN values up with the sum function.","metadata":{}},{"cell_type":"code","source":"daily_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.263865Z","iopub.execute_input":"2023-04-28T06:19:26.264335Z","iopub.status.idle":"2023-04-28T06:19:26.279272Z","shell.execute_reply.started":"2023-04-28T06:19:26.264283Z","shell.execute_reply":"2023-04-28T06:19:26.277967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.281325Z","iopub.execute_input":"2023-04-28T06:19:26.282746Z","iopub.status.idle":"2023-04-28T06:19:26.295522Z","shell.execute_reply.started":"2023-04-28T06:19:26.282667Z","shell.execute_reply":"2023-04-28T06:19:26.294516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tdcsfog_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.298311Z","iopub.execute_input":"2023-04-28T06:19:26.299151Z","iopub.status.idle":"2023-04-28T06:19:26.312250Z","shell.execute_reply.started":"2023-04-28T06:19:26.299088Z","shell.execute_reply":"2023-04-28T06:19:26.310688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"events_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.314590Z","iopub.execute_input":"2023-04-28T06:19:26.315677Z","iopub.status.idle":"2023-04-28T06:19:26.327799Z","shell.execute_reply.started":"2023-04-28T06:19:26.315610Z","shell.execute_reply":"2023-04-28T06:19:26.326321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subjects_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.329582Z","iopub.execute_input":"2023-04-28T06:19:26.330527Z","iopub.status.idle":"2023-04-28T06:19:26.343078Z","shell.execute_reply.started":"2023-04-28T06:19:26.330462Z","shell.execute_reply":"2023-04-28T06:19:26.341756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tasks_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.344866Z","iopub.execute_input":"2023-04-28T06:19:26.346057Z","iopub.status.idle":"2023-04-28T06:19:26.358992Z","shell.execute_reply.started":"2023-04-28T06:19:26.346006Z","shell.execute_reply":"2023-04-28T06:19:26.357503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there are no missing values in the daily_df, defog_df, tdcsfog_df, and tasks_df dataframes. However, we counted 1043 missing values in both Type and Kinetic data columns in the events_df dataframe as well as 62 missing values in the Visit column, 41 in UPDRSIII_Off column, and 1 in UPDRSIII_On column in the subjects_df dataframe. Specifically, the number of missing values that were present in the events_df and subjects_df dataframes make us think that some patients anonymously doesn't want to log their visit data, or there were corrupted data with possible human errors.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's count how many columns were there in all six dataframes separately! We pull out each dataframe to find the shape of them with the shape attribute, as well as finding their last index with the slice index of 1.","metadata":{}},{"cell_type":"code","source":"daily_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.361065Z","iopub.execute_input":"2023-04-28T06:19:26.362271Z","iopub.status.idle":"2023-04-28T06:19:26.371441Z","shell.execute_reply.started":"2023-04-28T06:19:26.362207Z","shell.execute_reply":"2023-04-28T06:19:26.370254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"defog_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.373079Z","iopub.execute_input":"2023-04-28T06:19:26.373541Z","iopub.status.idle":"2023-04-28T06:19:26.384192Z","shell.execute_reply.started":"2023-04-28T06:19:26.373492Z","shell.execute_reply":"2023-04-28T06:19:26.382302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tdcsfog_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.388813Z","iopub.execute_input":"2023-04-28T06:19:26.389565Z","iopub.status.idle":"2023-04-28T06:19:26.406202Z","shell.execute_reply.started":"2023-04-28T06:19:26.389469Z","shell.execute_reply":"2023-04-28T06:19:26.403147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"events_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.411573Z","iopub.execute_input":"2023-04-28T06:19:26.412538Z","iopub.status.idle":"2023-04-28T06:19:26.422197Z","shell.execute_reply.started":"2023-04-28T06:19:26.412476Z","shell.execute_reply":"2023-04-28T06:19:26.421021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subjects_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.424610Z","iopub.execute_input":"2023-04-28T06:19:26.425489Z","iopub.status.idle":"2023-04-28T06:19:26.434560Z","shell.execute_reply.started":"2023-04-28T06:19:26.425435Z","shell.execute_reply":"2023-04-28T06:19:26.432987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tasks_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.437381Z","iopub.execute_input":"2023-04-28T06:19:26.438174Z","iopub.status.idle":"2023-04-28T06:19:26.454542Z","shell.execute_reply.started":"2023-04-28T06:19:26.438125Z","shell.execute_reply":"2023-04-28T06:19:26.451389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running six data columns, we counted 4 data columns in the daily_df, defog_df, and tasks_df dataframes, 5 data columns in tdcsfog_df and events_df dataframes, and 8 columns in the subjects_df dataframe. Specifically, the less average number of the data columns of all six dataframes showed us that they wanted to keep the data concise about the data based on the patient's information and their FOG episodes. ","metadata":{}},{"cell_type":"markdown","source":"And with that visualized, we completed our basic investigation on the six dataframe's metadata! What's next for our data visualization voyage is that we're going to envisage the data series based on daily_df, defog_df, and tdcsfog_df dataframes all in one chapter on our investigative study on that data we're in.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 1: Daily Living, DeFog, tDCS Fog, All at Once</h2>\n\nAfter we arrived to this chapter of our data analysis, we found out that the data series based on the FOG data were split into three datasets: Daily FOG, DeFog, and tDCS Fog. And since we created the three dataframes out of the three FOG data series, let's summarize their columns concisely with detail!\n* **Id**: Identifier specified from each FOG patient.\n* **Subject**: Identifier specified for each subject.\n* **Visit**: Lab visits consisting on the baseline assessment.\n* **Medication**: Whether the patient has taken medication to alleviate the FOG symptoms.\n* **Test**: Specifies of which the three test types was performed, from easy (1) to hard (3) (for the tdcsfog_df dataframe).\n* **Beginning of recording [00:00-23:59]**: The time of day the recording begin (for the daily_df dataframe).\n\nWith all of the columns explained briefly, let's go visualize the three FOG data series, one by one from the daily_df dataframe towards the tdcsfog_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">daily_df</h3>","metadata":{}},{"cell_type":"markdown","source":"First, as we land into the chapter's first section of the daily_df dataframe, let's distribute and visualize the data from the Visit column into Plotly's histogram graph! We characterize the fig variable to the px module's histogram function for configuring a histogram, placing the daily_df dataframe as our data for the histogram, as well as arranging the x parameter to the Visit data column. Once completed, we display our graph to the notebook output with the show function plugged to the fig variable figure. ","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(daily_df, x=\"Visit\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:26.456638Z","iopub.execute_input":"2023-04-28T06:19:26.457502Z","iopub.status.idle":"2023-04-28T06:19:28.386959Z","shell.execute_reply.started":"2023-04-28T06:19:26.457447Z","shell.execute_reply":"2023-04-28T06:19:28.385892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the histogram we compiled, we realized that the data labeled as \"1\" was counted more than \"2\". Specifically, there are 59 entities for the data marked as \"1\" whereas there are 6 data entities for the data marked as \"2\". In addition, the data marked as \"1\" gave us a clue that most FOG patients has follow-up assessments than those who has post-treatment assessments as the MICHAEL J. FOX FOUNDATION stated:\n> ...two post-treatment assessments for different treatment stages, and one follow-up assessment.","metadata":{}},{"cell_type":"markdown","source":"Let's then distribute the recording data column into another histogram graph, but this time, include a boxplot with Plotly! Beforehand into plotting, we create another column in the daily_df dataframe, Recording, to convert to string with the str attribute that was plugged to the daily_df dataframe's column that was named as \"Beginning of recording [00:00-23:59]\" and then replace the colon with an empty string with the replace function thus converting to integer type with the astype function, containing the \"int\" type.\n\nAfter we created our new column, we characterize our one and only fig variable again, towards the px module's histogram function, setting the daily_df dataframe for the histogram's data, alongside with the x parameter to the Recording data column for specifying the x-axes of the histogram and the marginal parameter to box for placing the boxplot on the top of the graph. And just like that, we use the show function and plug it towards the fig variable figure for displaying the figure in the notebook output.","metadata":{}},{"cell_type":"code","source":"daily_df[\"Recording\"] = daily_df[\"Beginning of recording [00:00-23:59]\"].str.replace(\":\", \"\").astype(int)\n\nfig = px.histogram(daily_df, x=\"Recording\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:28.396518Z","iopub.execute_input":"2023-04-28T06:19:28.397358Z","iopub.status.idle":"2023-04-28T06:19:28.733214Z","shell.execute_reply.started":"2023-04-28T06:19:28.397297Z","shell.execute_reply":"2023-04-28T06:19:28.731876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the histogram above, we found out that the peak of data bins has formed on the left of the diagram, indicating this distribution of data as a right-skew. Besides, the highest data range in the recording data listed in the histogram is 800 to 890, with 34 data entities, while the data ranges from 1000 to 1190 and from 1400 to 1490 are counted the least, with 2 data entities. Meanwhile in the box plot, we spotted a few outliers on the right of the diagram as well as finding the malformed left side of the box part in the graph. Other than that, the first and quartiles of the recording data specified in the boxplot is 800 and 1021.75, the median is 830, and the interquartile range is 221.75. Specifically, the data distribution that displayed a right-skew in the histogram hinted us that most recordings based on the FOG data from patients start at around 8AM.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's visualize the relations between the Visit data column and the column that held the recording data in the 2D histogram! Once again, we characterize the fig variable figure to the px module's density_heatmap function from the px module for creating a 2D histogram, setting the daily_df dataframe as the data for the 2D histogram plot, along with the x and y parameters to the \"Visit\" and \"Recording\" columns for configuring the x and y axes of the graph as well as the text_auto parameter to True for displaying the numbers in each box of data density. Finally, we once again use the show function to display the fig variable's graph.","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(daily_df, x=\"Visit\", y=\"Recording\", text_auto=True)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:28.735041Z","iopub.execute_input":"2023-04-28T06:19:28.735714Z","iopub.status.idle":"2023-04-28T06:19:28.841490Z","shell.execute_reply.started":"2023-04-28T06:19:28.735636Z","shell.execute_reply":"2023-04-28T06:19:28.839938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After the 2D histogram was compiled, we found out that there's a peak of data mapped at the bottom-left of the histogram. Specifically, the most counted data shown by the 2D histogram is between 800 to 990 in the Recording column as well as under the value of 1 in the Visit Column, while box located at the range between 1400 and 1590 in the Recording column under the value of 1 is counted the least, as there were 2 data entities inside. In addition, the yellow box shown in the bottom-left of the 2D histogram made us explain that most patients has follow-up assessments while the FOG recording starts around 8 to 9 AM.","metadata":{}},{"cell_type":"markdown","source":"Now that we have the data analysis about the Daily Living FOG data that was held in the daily_df dataframe fully visualized, let's now proceed to visualize the DeFog data in the defog_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">defog_df</h3>\n\nOnce we head into this section over visualizing the data from the defog_df dataframe, let's first distribute and observe the Visit data into a histogram made from Plotly! We characterize the fig variable to create a histogram graph with the px module's histogram function, in which we place the defog_df dataframe as the data for the histogram, followed by configuring the x parameter to the Visit column for specifying the histogram's x-axis. Last but not least, we apply the show function to the fig variable figure for displaying the graph into the notebook output.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(defog_df, x=\"Visit\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:28.843850Z","iopub.execute_input":"2023-04-28T06:19:28.844991Z","iopub.status.idle":"2023-04-28T06:19:28.909547Z","shell.execute_reply.started":"2023-04-28T06:19:28.844922Z","shell.execute_reply":"2023-04-28T06:19:28.908100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just like the previous distribution of the Visit data column in the daily_df dataframe, we realized that the data marked as \"1\" was counted more than the other data that was marked as \"2\". However, the data marked as \"2\" was counted higher than the data marked as \"2\" in the daily_df dataframe. Furthermore, there are 70 entities for the data labeled as \"1\", while there are 67 entities for the data labeled as \"2\". Specifically, the data marked as \"1\" hinted us that most FOG patients has follow-up assessments than those who has post-treatment assessments just like the data in the daily_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"Let's now visualize the Medication column into Plotly's pie chart! Beforehand, we characterize the medication_df dataframe to count the values of the defog_df dataframe's Medication column with the value_counts function. Following from that, we then characterize the fig variable to create a pie chart with the pie function from the px module, as we place the medication_df dataframe as the data for the pie chart alongside configuring the names parameter to the indexes from the medication_df dataframe with the index attribute and the values parameter to the values extracted from the medication_df dataframe with the values attribute. As we finish creating the pie chart, we now display it into the notebook output cell with the show function plugged to the fig variable figure.","metadata":{}},{"cell_type":"code","source":"medication_df = defog_df[\"Medication\"].value_counts()\n\nfig = px.pie(medication_df, names=medication_df.index, values=medication_df.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:28.911309Z","iopub.execute_input":"2023-04-28T06:19:28.912281Z","iopub.status.idle":"2023-04-28T06:19:28.995265Z","shell.execute_reply.started":"2023-04-28T06:19:28.912222Z","shell.execute_reply":"2023-04-28T06:19:28.993899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the pie chart that was created out from the \"Medication\" column data in the defog_df dataframe, we witnessed almost half of the data were labeled as \"on\" and others that were labeled as \"off\". Additionally, the data marked as \"on\" is hardly counted the most as there are 69 entities whereas the other data labeled as \"off\" is barely the least counted since there are 68 entities. In other words, the nearly semi-counted data based on the \"Medication\" column hinted us that half of the patients took medication to alleviate the FOG and Parkinson's symptoms while the other half of the patients didn't took medication because of their mild conditions on FOG.","metadata":{}},{"cell_type":"markdown","source":"With just two data analysis explained in the defog_df dataframe since there are less columns to visualize, let's head onwards to the last section of the first chapter in which we're going to analyze the tdcsfog_df dataframe, as it has more columns than both daily_df and defog_df dataframes.","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">tdcsfog_df</h3>\n\nIn our final section we're jumping into, let's first distribute and visualize the data from the tdcsfog_df dataframe's \"Visit\" column into the histogram! As always, we build our variable, fig, to create the histogram graph with the px module's histogram function, placing the tdcsfog_df dataframe as the data for plotting the histogram and then configure the x parameter to the \"Visit\" column for specifying the histogram's x-axes based on the column from a specific dataframe. After we create our histogram model, we typically use the show function to the fig variable figure for displaying the graph on the notebook output.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(tdcsfog_df, x=\"Visit\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:28.997552Z","iopub.execute_input":"2023-04-28T06:19:28.998058Z","iopub.status.idle":"2023-04-28T06:19:29.077142Z","shell.execute_reply.started":"2023-04-28T06:19:28.998011Z","shell.execute_reply":"2023-04-28T06:19:29.075782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unlike the \"Visit\" data distributed from the daily_df and the defog_df dataframes, we realized that there's not only ones and twos displayed in the diagram, there are more data distributed in the tdcsfog_df as the data plotted in the tdcsfog_df dataframe's Visit column displayed a right-skew distribution. Specifically, the value 2 in the Visit column were counted the most containing 245 data entities, while the value 20 is counted the least as it held 55 data entities. In addition, the right-skewed data distribution of the \"Visit\" column in the tdcsfog_df hinted us that most FOG patients visited to the labs for a baseline assessment frequently, while others visited to the labs sometimes.","metadata":{}},{"cell_type":"markdown","source":"Let's then proceed to plot down the \"Test\" column into a pie chart! Before we begin plotting this out, we characterize the test_tdcsfog_df dataframe to count the values with the value_counts function of the tdcsfog_df dataframe's \"Test\" column. We then plot our pie chart by creating our fig variable to define it into the px module's pie function, setting the test_tdcsfog_df dataframe as the pie chart's data, as well as configuring the names parameter to the indexes of the test_tdcsfog_df dataframe with the index attribute and the values parameter to the values of the test_tdcsfog_df dataframe with the values attribute. Now that we have our pie graph completed, we plug the show function to the fig variable figure to display the graph.","metadata":{}},{"cell_type":"code","source":"test_tdcsfog_df = tdcsfog_df[\"Test\"].value_counts()\n\nfig = px.pie(test_tdcsfog_df, names=test_tdcsfog_df.index, values=test_tdcsfog_df.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.078573Z","iopub.execute_input":"2023-04-28T06:19:29.079670Z","iopub.status.idle":"2023-04-28T06:19:29.144378Z","shell.execute_reply.started":"2023-04-28T06:19:29.079626Z","shell.execute_reply":"2023-04-28T06:19:29.143047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we create our pie chart, we glimpsed how the pie chart were nearly split evenly into three parts, as each slice shows a data label for the \"Test\" column from the tdcsfog_df dataframe. In other words, the data marked as \"1\" is counted the most with 286 data entities, while the data label \"3\" is counted the least, containing 265 data entities. Specifically, the highest data label \"1\" made us explain that most patient's tests were least challenging, while other patient's tests were most challenging hence the lowest data label \"3\".","metadata":{}},{"cell_type":"markdown","source":"Since we have plotted the standalone graphs for the \"Visit\" and \"Test\" columns, let's plot them into our density heatmap graph! Once again, we characterize the fig variable figure to the px module's density_heatmap function for creating the density heatmap figure, setting the tdcsfog_df dataframe as the data for the density heatmap, as well as the x and y parameters to the \"Test\" and \"Visit\" columns for specifying the x and y axes of the diagram along with configuring the text_auto parameter to True for displaying the value counts in each square. With that completed, we apply the show function to the fig variable figure for displaying our diagram to the cell output below.","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(tdcsfog_df, x=\"Test\", y=\"Visit\", text_auto=True)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.146306Z","iopub.execute_input":"2023-04-28T06:19:29.147519Z","iopub.status.idle":"2023-04-28T06:19:29.213963Z","shell.execute_reply.started":"2023-04-28T06:19:29.147347Z","shell.execute_reply":"2023-04-28T06:19:29.212548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the density heatmap that held the \"Test\" and \"Visit\" columns, we visualized that there's more data in the bottom of the diagram, while there's few data on the top of the diagram. Nevertheless, the \"Test\" column data labeled as \"1\" under the range from 0 to 4 in the \"Visit\" column is counted the most with 185 entities listed, while the \"Test\" column data labeled as \"3\" under the range from 20 to 24 in the \"Visit\" column is counted the least with 15 entities calculated. Additionally, the density heatmap we plotted made us explain that the tests from most patients that visited the labs frequently were least challenging while other tests from patients who visited the labs sometimes vary from least challenging to most challenging.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's visualize and sort out the \"Medication\" column into the pie chart! Prior to plotting this pie graph out, we inherit the medication_df dataframe into tallying the values of the tdcsfog_df dataframe's \"Medication\" column with the value_counts function. After that, we create our fig variable and define it to the pie function from the px module for configuring our pie graph, setting the medication_df dataframe as the data for the pie chart, followed by arranging the names parameter to the medication_df dataframe's indexes specified by the index attribute and the values parameter to the values specifcation of the medication_df dataframe with the values attribute. Finally, we use the show function into the fig variable figure for exhibiting the graph into the output cell on the bottom of the code cell.","metadata":{}},{"cell_type":"code","source":"medication_df = tdcsfog_df[\"Medication\"].value_counts()\n\nfig = px.pie(medication_df, names=medication_df.index, values=medication_df.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.215999Z","iopub.execute_input":"2023-04-28T06:19:29.216828Z","iopub.status.idle":"2023-04-28T06:19:29.281399Z","shell.execute_reply.started":"2023-04-28T06:19:29.216773Z","shell.execute_reply":"2023-04-28T06:19:29.280109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw in the pie chart we configured above, we found out that the data marked as \"on\" has more values than the data labeled as \"off\" unlike the pie chart we graphed in the defog_df dataframe. In other words, 64.3% of the data were marked as \"on\" with 536 entities calculated, while 35.7% of the data were labeled as \"off\" with 297 data entities listed. Furthermore, the data that was marked as \"on\" mostly in the tdcsfog_df dataframe's Medication column made us explain that most patients took the anti-Parkinsonian medication during the recording than the ones who didn't, since it was based on the severity of the symptoms they experienced with.","metadata":{}},{"cell_type":"markdown","source":"With our analysis finished in all of the three dataframes, daily_df, defog_df, and tdcsfog_df, we completed the first chapter of our data analysis based on Freezing of Gait! What's next for visualizing the data is that we are going to explore the FOG events of each patient in the events_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 2: Throwback to the FOG-gy Event</h2>\n\nPreviously, we explored through the FoG data from the daily_df, defog_df, and tdcsfog_df dataframes all at the same time separately, and we were familiar about the medication the patients have as well as the visits they appointed and the tests they took. Now, we're into the visualizing and exploring the metadata of FOG events inside the events_df dataframe, because it was gathered by the sensor data that patients wore for experts to track the FoG events. Besides, let's go into the columns inside the events_df dataframe!\n* **Id**: Identifier for when the event occurred in.\n* **Init**: Time specified for when the event began.\n* **Completion**: Time specified for when the event ended.\n* **Type**: Specified whether it has \"StartHestitation\", \"Turn\", and \"Walking\".\n* **Kinetic**: Specified whether the event is kinetic (1) or akinetic (0).","metadata":{}},{"cell_type":"markdown","source":"With all of the columns detailed precisely, let's head on to distribute data from the Init column into the histogram and box-plot combined with Altair! First, we create our variable, fig1, to characterize it to Altair's chart configuration with the alt module's Chart function that contains the events_df dataframe for loading the data in the graph, followed by marking the bars into our graph with the mark_bar function along with encoding the parameters of our chart with the encode function, setting the x-axis configuration with the X function from the alt module inside, in which it held the Init data column and the bin parameter that is set to True for creating the bins of the histogram, followed by configuring the y parameter to the count function that was encased in strings for counting the values of a specific data column.\n\nAfter we create our histogram graph in the fig1 variable, we characterize another variable, fig2, to build another graph with the Chart function from the alt module that contained the events_df dataframe for the graph's data but this time, we use the mark_boxplot to create our box-plot graph, setting the extent parameter to 'min-max' for enabling the whiskers into the box-plot, and then we encode our box-plot parameters with the encode function, setting the x parameter to the Init column for configuring the x-axes of our box-plot graph followed by adjusting the properties of our box-plot variable with the properties function, setting the height parameter to 300 for sizing the box-plot's height. Thenceforth, we use the alt module's concat function to add two graphs into the subplot, placing the fig1 and fig2 variables inside.","metadata":{}},{"cell_type":"code","source":"fig1 = alt.Chart(events_df).mark_bar().encode(\n    alt.X(\"Init\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(events_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"Init\"\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.283090Z","iopub.execute_input":"2023-04-28T06:19:29.284564Z","iopub.status.idle":"2023-04-28T06:19:29.452659Z","shell.execute_reply.started":"2023-04-28T06:19:29.284519Z","shell.execute_reply":"2023-04-28T06:19:29.451557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the histogram we created on the left of our subplot, we noted that there's a peak of data shown in the left of the diagram, as we noticed that the data in the Init column displayed a right-skew distribution. Additionally, the range in the Init column that has the highest data is from 0 to 500, with around 1490 entities listed, while the other range between 4000 and 4500 is counted the least, with approxmiately 50 to 60 entities listed.\n\nMeanwhile in the boxplot we graphed, we glimpsed on how the width of the minimum whiskers were shorter than the width of the maximum whiskers, which we spotted how there's a right-skew in the box-plot. Specificallly, the first and third quartiles of the Init data is 45.02 and 1608.29, the median is 849.58, and the interquartile range is 1563.27. Furthermore, the right-skewed data distribution from the Init data column implyed us that most events were logged earlier to determine whether patients had the symptoms of FoG.","metadata":{}},{"cell_type":"markdown","source":"After we plot down the Init column into the previous histogram and the box-plot, let's plot another histogram and the box-plot for distributing the data from the Completion column! As always, we create our variable, fig1, and define it to the chart configuration with the alt module's Chart function that held the events_df dataframe for the graph's data, as well as placing the mark_bar to plot the bars towards our graph, followed by encoding our graph's parameters with the encode function, loading the alt module's X function to configure our graph's x-axes in which it held the Completion column thus having the bin parameter set to true indicating that we're creating our histogram, followed by setting the y parameter to the count function that was covered by the strings for counting the values of the specific data column to the y-axis. \n\nNext, we configure another variable, fig2, to create another graph with the alt module's Chart function that contained the events_df dataframe for the data of the graph, then we plug the mark_boxplot function for plotting our box-plot graph setting the extent parameter to \"min-max\" to enable whiskers in the box-plot, as well as applying the encode function for adjusting our graph's parameters, just setting the x parameter to the Completion column for arranging our graph's x-axes followed by constructing our graph's properties with the properties function, arranging the height parameter to 300 for adjuting the height of our graph. With two graphs completed, we use the alt module's concat function to concatenate two graphs into a single subplot, placing the fig1 and fig2 variables inside.","metadata":{}},{"cell_type":"code","source":"fig1 = alt.Chart(events_df).mark_bar().encode(\n    alt.X(\"Completion\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(events_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"Completion\"\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.453824Z","iopub.execute_input":"2023-04-28T06:19:29.454164Z","iopub.status.idle":"2023-04-28T06:19:29.595994Z","shell.execute_reply.started":"2023-04-28T06:19:29.454129Z","shell.execute_reply":"2023-04-28T06:19:29.594518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Likewise to the data distribution from the Init column to the previous subplot that held the histogram and the box-plot, we found out that the data distributed from the Completion column were nearly the same as the Init column distributed, as both distributions displayed a right-skew distribution. In other words, the data range that has the highest data is from 0 to 500 with approximately 1460 entities, whereas the data range from 4000 to 4500 is counted the least with roughly 50 or 60 entities. On the boxplot chart, we realized that the box portion is shifted to the left of the diagram, as the width of the minimum whiskers is shorter than the width of the maximum whiskers. Specifically, the first and third quartiles of the Completion column shown in the boxplot is 55.69 and 1611.91, the median is 852.58, and the interquartile range is 1556.22. Furthermore, the right-skewed distribution based from the Completion column implied to us that most recording events of patients end earlier, similarly to the data distribution in the Init column.","metadata":{}},{"cell_type":"markdown","source":"Now let's create the 2D histogram scatter plot consisting over the distributions of the Init and Completion columns! Once again, we use the alt module's Chart function to create our chart with Altair that contains the events_df dataframe inside for our graph's data, then we place the mark_circle function to mark scatter plots around our graph, and then configure our graph's parameters with the encode function, applying the alt module's X and Y columns to create our graph's x and y-axes in which it held the Init and Completion columns as both bin parameters were set to True, followed by configuring the size parameter to the count function to count the values of two columns.","metadata":{}},{"cell_type":"code","source":"alt.Chart(events_df).mark_circle().encode(\n    alt.X(\"Init\", bin=True),\n    alt.Y(\"Completion\", bin=True),\n    size=\"count()\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.597850Z","iopub.execute_input":"2023-04-28T06:19:29.599245Z","iopub.status.idle":"2023-04-28T06:19:29.713478Z","shell.execute_reply.started":"2023-04-28T06:19:29.599188Z","shell.execute_reply":"2023-04-28T06:19:29.712121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the 2d histogram scatter-plot graph we assembled, we realized that the circles were shifting diagonally to the upper-right of our diagram as the size of the circles shrink. Specifically, the graph we plotted out from the data of the Init and Completion columns hinted us that some recordings start earlier and then end earlier, making the duration of them shorter. ","metadata":{}},{"cell_type":"markdown","source":"Speaking about the duration of each recording, let's now visualize the distribution of the Duration column in both histogram and boxplot nested in a subplot! Before we begin graphing this, we create a new column in the events_df dataframe, Duration, to subtract the values from the Completion column by the values in the Init column. \n\nAfter we built the Duration column, we characterize our fig1 variable to create a new graph with the Chart function from the alt module in which we place in the events_df dataframe for loading the data to our graph, followed by inserting the mark_bar function to add the bars in our graph thus encoding the parameters of our graph with the encode function, setting the alt module's X function that held the Duration column we created as well as the bin parameter set to True for loading the x-axes into our graph, along with setting the y parameter to the count function under the strings for counting a specific column of data into our graph's y-axes.\n\nThenceforth, we build another variable, fig2, to generate another graph with the Chart function from the alt module with the events_df dataframe held inside again, but this time, we use the mark_boxplot function to generate our box plot inside our graph in which we set the extent parameter to min-max for enabling whiskers, as well as configuring our parameters of our graph with the encode function, setting the x parameter to the Duration column for specifying the x-axes to our graph, followed by inserting the properties function on the outside of the encode function just for adjusting the appearance of our graph, arranging the height parameter to 300. Lastly, we use the alt module's concat function for creating our subplot of two graphs we created in the fig1 and fig2 variables.","metadata":{}},{"cell_type":"code","source":"events_df[\"Duration\"] = events_df[\"Completion\"] - events_df[\"Init\"]\n\nfig1 = alt.Chart(events_df).mark_bar().encode(\n    alt.X(\"Duration\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(events_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"Duration\"\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.715349Z","iopub.execute_input":"2023-04-28T06:19:29.716429Z","iopub.status.idle":"2023-04-28T06:19:29.866423Z","shell.execute_reply.started":"2023-04-28T06:19:29.716381Z","shell.execute_reply":"2023-04-28T06:19:29.865091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the data distribution based on the Duration column, we realized that there's a data peak on the left of our histogram graph that displayed a left-skew distribution type, meaning that the range from 0 to 100 held the most data with around 3600 data entities. On the box-plot we graphed on the right of our subplot, we glimpsed of how the box portion was narrow to see, followed by noticing how the whisker from the box portion to the maximum of the Duration column was longer than the whisker from the minimum of the Duration column to the box portion of the box plot. In other words, the first and third quartiles of the Duration column noted by the box plot is 1.159 and 6.453, the median is 2.52, and the interquartile range is 5.294. Likewise to the distribution of Init and Completion columns we created in the 2D histogram scatter plot, most durations of the recordings specified from the Duration column distribution into our histogram and box plot start earlier and then end earlier, making the duration of them shorter. ","metadata":{}},{"cell_type":"markdown","source":"After we visualized the data on the Duration, Init, and Completion columns, let's sort the Type column into a pie chart! Prior to plotting the pie chart out, we create another dataframe, fog_type_df, to calculate the overall values with the value_counts function placed into the events_df dataframe's Type column and then redefine it to create a dataframe with the pd module's DataFrame function, containing a dictionary that has the type key assigned to the indexes of the fog_type_df dataframe specified by the index attribute and the count key to the fog_type_df dataframe's values recorded by the values attribute. \n\nFollowing from creating the fog_type_df dataframe, we use it to the alt module's Chart function to create our typical chart, then we use the mark_arc function to assemble the pie chart, and then we encode our graph's parameters with the encode function, setting the theta parameter to the alt module's Theta function that contained the field parameter set to the count column and the type parameter set to quantitative for specifying the circle's theta slice, the color parameter to the alt module's Color function that accommodates the field parameter configured to the type column and the type parameter configured to nominal for arranging the color of each pie slice, and the tooltip parameter to a list containing the type and count columns for making our pie graph interactive.","metadata":{}},{"cell_type":"code","source":"fog_type_df = events_df[\"Type\"].value_counts()\n\nfog_type_df = pd.DataFrame({\n    'type': fog_type_df.index,\n    'count': fog_type_df.values\n})\n\nalt.Chart(fog_type_df).mark_arc().encode(\n    theta=alt.Theta(field='count', type=\"quantitative\"),\n    color=alt.Color(field='type', type=\"nominal\"),\n    tooltip=[\"type\", \"count\"]\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.867908Z","iopub.execute_input":"2023-04-28T06:19:29.868902Z","iopub.status.idle":"2023-04-28T06:19:29.922341Z","shell.execute_reply.started":"2023-04-28T06:19:29.868858Z","shell.execute_reply":"2023-04-28T06:19:29.921129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As a result from plotting data in the pie graph based on the events_df dataframe's Type column, we realized that approxmiately 80% of the data were marked as \"Turn\" with 2147 data entities calculated, while 16% of the data was labeled as \"Walking\" with 415 entities listed and 4% of the data was listed as \"StartHesitiation\" with 107 units logged. Specifically, the data that was mostly marked as \"Turn\" hinted us that it reveals the characteristics of FoG better than walking forward and backward with Parkinson's disease, [according to a research paper](https://www.sciencedirect.com/science/article/abs/pii/S0966636222000807?via%3Dihub).","metadata":{}},{"cell_type":"markdown","source":"Lastly for our chapter of visualizing the data in the events_df dataframe, let's visualize the Kinetic column into an ordinary bar chart! Basically, we use the alt module's Chart function to create our new chart as we place the events_df dataframe inside for loading the data for our graph, followed by using the mark_bar function to plot the bars into our graph, as well as encoding the parameters of our graph with the encode function, setting the x parameter to the Kinetic column for specifying our graph's x-axes, and the y parameter to the count function encased in strings for counting the values for a specific column.","metadata":{}},{"cell_type":"code","source":"alt.Chart(events_df).mark_bar().encode(\n    x=\"Kinetic\",\n    y=\"count()\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:29.923743Z","iopub.execute_input":"2023-04-28T06:19:29.924098Z","iopub.status.idle":"2023-04-28T06:19:30.056168Z","shell.execute_reply.started":"2023-04-28T06:19:29.924056Z","shell.execute_reply":"2023-04-28T06:19:30.054904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the graph we compiled, we found out that almost 2190 data entities in the Kinetic column were labeled as 1 while the other 500 data entities were labeled as 0. Additionally, the data that was marked as 1 implied us that most Parkinson's FoG patients [were mostly kinetic and had their involved movement and didn't experience freezing episodes sometimes, whether they walk forward or backward or turn around in circles](https://www.parkinson.org/library/fact-sheets/freezing) as the other data marked as 0 [showed how they were temporary glued to their ground while moving](https://www.apdaparkinson.org/article/freezing-gait-and-parkinsons-disease/#:~:text=Freezing%20of%20gait%20is%20an,sense%2C%20you're%20stuck.).","metadata":{}},{"cell_type":"markdown","source":"Now that we completed visualizing most of the data held in the columns from the events_df dataframe in this section, let's now proceed to visualize the data of the subjects gathered in the subjects_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 3: Studying the Subject's FoG Data</h2>\n\nAs we approach into the third chapter of our data analysis on the Parkinson's Freezing of Gait Competition, we came across the data full of patient subjects and their background information in the subjects_df dataframe, as it was used to identify whether they have any episodes of freezing. Besides from the overview on the subjects_df dataframe, here are the subjects_df dataframe's columns explained with precise details!\n* **Subject**: Identifier for each patient.\n* **Visit**: Specifies how frequent each patient visit to a lab.\n* **Age**: Specifies the age of each patient.\n* **Sex**: Specifies whether each patient is male or female.\n* **YearsSinceDx**: Specifies how many years since the patient was diagnosed with Parkinson's.\n* **UPDRSIIIOn/UPDRSIIIOff**: Specifies the rating of the Unified Parkinson's Disease Scale during the on or off medication all at once.\n* **NFOGQ**: Specifies the FoG self-report questionaire score, from [NIH](https://pubmed.ncbi.nlm.nih.gov/19660949/).","metadata":{}},{"cell_type":"markdown","source":"With all the columns expounded with concise details, let's begin visualizing the subjects_df dataframe by distributing the Visit column into a histogram graph! Basically, we generate the fig variable figure, and define it to the px module's histogram function for creating our histogram graph, placing the subjects_df dataframe as the data needed for the graph, followed by configuring the x parameter to the Visit column for specifying the graph's x-axes. With that completed, we use the show function to the fig variable figure for displaying the histogram graph into the output cell.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(subjects_df, x=\"Visit\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.058043Z","iopub.execute_input":"2023-04-28T06:19:30.058866Z","iopub.status.idle":"2023-04-28T06:19:30.122572Z","shell.execute_reply.started":"2023-04-28T06:19:30.058800Z","shell.execute_reply":"2023-04-28T06:19:30.121388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although the Visit column has 1s and 2s in the histogram, we glimpsed on how the histogram based on the Visit column data distribution displayed a right-skew type, as the data that is labeled as 1 is counted more than the other data that was marked as 2. Additionally, the data that was tagged as 1 has 70 entities, while the other data that was characterized as 2 has 41 data entities. Moreover, the data that was mostly classified as 1 implied us that most FOG patients has more follow-up assessments than those who has post-treatment assessments, likewise to the defog_df and daily_df dataframes.","metadata":{}},{"cell_type":"markdown","source":"Let's now then distribute the data from the Age column into plotting them with a histogram-box chart! Once again, we characterize the fig variable figure to create another histogram model with the px module's histogram function, setting the subjects_df dataframe as the histogram's data to plot, alongside configuring the x parameter to the Age column for specifying the x-axes of the column and the marginal parameter to box for plotting a box-plot on the top of the histogram model. After we created our histogram-box graph, we display the fig variable's graph with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(subjects_df, x=\"Age\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.123827Z","iopub.execute_input":"2023-04-28T06:19:30.124167Z","iopub.status.idle":"2023-04-28T06:19:30.205360Z","shell.execute_reply.started":"2023-04-28T06:19:30.124133Z","shell.execute_reply":"2023-04-28T06:19:30.204155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As a result of compiling this histogram-box graph based on the Age column, we espied that there's a peak of data on the middle-right of the graph, as it clearly displayed a left-skew distribution. In addition, the range from 70 to 74 held the most data with 42 entities, while the other ranges 25-29 and 90-94 has one data entity, which they had the least data. On the box-plot, we spotted two outliers on the far left and right of the diagram, followed by observing the whiskers begin and end on nearly the right. In other words, the first and third quartile ranges is 62 and 73, the median is 68, and the interquartile range is 11. Furthermore, the left-skewed distribution in the middle of the histogram made us receive a clue that most patients were aged 60 or older, [as there are most cases seen in older people](https://parkinsonsdisease.net/elderly-population).","metadata":{}},{"cell_type":"markdown","source":"Once we distributed the data from the Age and Visit columns separately, let's visualize both data columns into a 2D histogram graph! We create the fig variable figure and define it to the px module's density_heatmap function for creating the 2D histogram heatmap diagram, setting the subjects_df dataframe as the data for the 2D histogram heatmap graph, followed by configuring the x parameter to the Visit column for specifying the x-axes of the graph, and the y parameter to the Age column for specifying the y-axes of the graph. And following from creating the 2D histogram heatmap, we use the show function and plug it to the fig variable figure for displaying the graph into the notebook output.","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(subjects_df, x=\"Visit\", y=\"Age\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.207308Z","iopub.execute_input":"2023-04-28T06:19:30.207643Z","iopub.status.idle":"2023-04-28T06:19:30.266329Z","shell.execute_reply.started":"2023-04-28T06:19:30.207610Z","shell.execute_reply":"2023-04-28T06:19:30.265074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the Visit-Age data column distribution, we spotted the hotspots in the middle-top of the histogram, as the Age column range from 70 to 74 that is under the Visit column label 1 has the most common data, with 18 data entities. In other words, most patients that were aged 60 or older had more follow-up assessments than the ones that had two post-treatment assessments.","metadata":{}},{"cell_type":"markdown","source":"Aside from distributing and visualizing the Visit and Age columns, let's now proceed to distribute the Sex data into the pie chart! Prior to plotting this out, we inherit a new dataframe, subjects_sex_df, to tally the values in the subjects_df dataframe's Sex column by plugging the value_counts function into it. Following from this step, we characterize the fig variable figure to create a pie chart with the px module's pie function, setting the subjects_sex_df dataframe as the data for the pie chart, followed by the names parameter to the indexes specified by the index attribute plugged into the subjects_sex_df dataframe for configuring the names and labels of the pie chart, and the values parameter to the values gathered from the values attribute that is applied into the subjects_sex_df dataframe. With our graph finished, we use the show function into the fig variable figure for displaying our graph towards the code cell's output below.","metadata":{}},{"cell_type":"code","source":"subjects_sex_df = subjects_df[\"Sex\"].value_counts()\n\nfig = px.pie(subjects_sex_df, names=subjects_sex_df.index, values=subjects_sex_df.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.267749Z","iopub.execute_input":"2023-04-28T06:19:30.268180Z","iopub.status.idle":"2023-04-28T06:19:30.321996Z","shell.execute_reply.started":"2023-04-28T06:19:30.268145Z","shell.execute_reply":"2023-04-28T06:19:30.320681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the pie chart generated based on the \"Sex\" data column distribution in the subjects_df dataframe, we realized that 69.9% of the data was marked as \"M\" (represents male) with 121 entities listed, while 30.1% of the other data was labeled as \"F\" (represents female) with 52 entities calculated. Specifically, the most data that is labeled as \"M\" in the subjects_df dataframe's Sex column implied us that [most male patients were at risk from contracting the Parkinson's disease, as the relative risk was 1.5 times greater than in female patients](https://pubmed.ncbi.nlm.nih.gov/15026515/).","metadata":{}},{"cell_type":"markdown","source":"Now let's distribute and then visualize the YearsSinceDx column into another histogram-box graph! As always, we define our figure variable, fig, to create a histogram graph template with the histogram function from the px module, loading the subjects_df dataframe as the data for the graph, as well as configuring the x parameter to the YearsSinceDx column for loading the x-axes in the histogram and the marginal parameter to box for creating our box chart on the top of the histogram. Finally, we use the show function and apply it to the fig variable figure for presenting our graph into the notebook output.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(subjects_df, x=\"YearsSinceDx\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.323470Z","iopub.execute_input":"2023-04-28T06:19:30.323851Z","iopub.status.idle":"2023-04-28T06:19:30.408819Z","shell.execute_reply.started":"2023-04-28T06:19:30.323812Z","shell.execute_reply":"2023-04-28T06:19:30.407461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw in the YearsSinceDx data distribution, we realized that there's a data peak on the left of the histogram as it shifts to the left, making it showing a right-skewed distribution type. Besides, the highest data counted is in the range between 6 to 7.9 with entities while the range from 30 to 31.9 has one data entity, making it as the lowest data counted. Meanwhile in the box-plot on the top of the histogram graph, we spotted one outlier on the right of the diagram thus seeing the whisker to the maximum of the data greater than the whisker to the minimum of the data. Additionally, the first and third quartiles seen in the box plot is 5.75 and 15, the median is 9, and the interquartile range is 9.25. Funilly enough, the right skewed data distribution in the YearsSinceDx column gave us a clue [that some patients was diagnosed with Parkinson's disease too early, as about 5 to 10% of them had \"early-onset\" disease.](https://www.ncoa.org/article/parkinsons-disease-early-signs-symptoms-and-what-to-do-when-diagnosed.)","metadata":{}},{"cell_type":"markdown","source":"Now let's graph a subplot containing the two data distributions of UPDRSIIIOn and UPDRSIIIOff columns in two histograms! Before we get into graphing the subplot out, we import the plotly module's graph_objects attribute as go and the make_subplots function from the plotly module's subplots attribute. \n\nAfter we import the two modules needed for creating a subplot graph, we characterize the fig variable figure to create our subplots with the make_subplots function, setting the rows parameter to 1 and the cols parameter to 2 just for configuring a single row of a subplot with two columns. Now that we have the subplot template ready, we add two graphs with the add_trace function plugged into the fig variable figure, inserting the Histogram function from the go module two times separately in which the x parameter inside was configured to the subjects_df dataframe's UPDRSIII_On and UPDRSIII_Off columns separately as well as the name parameter set to the column names of UPDRSIII_On and UPDRSIII_Off individually, followed by configuring the row parameter in both of the add_trace functions to 1 for placing two graphs in the same row but we configured the col parameter to 1 in the first add_trace function for placing the first chart into the first column and 2 in the second add_trace function for placing the last chart into the second column. And with our subplot that contains two histograms being finalized, we use the show function into the fig variable figure for displaying the graph.","metadata":{}},{"cell_type":"code","source":"import plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\nfig = make_subplots(rows=1, cols=2)\n\nfig.append_trace(go.Histogram(x=subjects_df[\"UPDRSIII_On\"], name=\"UPDRSIII_On\"), row=1, col=1)\nfig.append_trace(go.Histogram(x=subjects_df[\"UPDRSIII_Off\"], name=\"UPDRSIII_Off\"), row=1, col=2)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.410213Z","iopub.execute_input":"2023-04-28T06:19:30.411167Z","iopub.status.idle":"2023-04-28T06:19:30.451510Z","shell.execute_reply.started":"2023-04-28T06:19:30.411125Z","shell.execute_reply":"2023-04-28T06:19:30.450215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the UPDRSIII_On and UPDRSIII_Off columns, we found out that both graphs from left and right of the subplot showed a right-skew distribution type on the left of each chart. Aside from the look of each two graphs, the range that has the highest data is from 35 to 39 in the UPDRSIII_On column with 34 entities and from 45 to 49 in the UPDRSIII_Off column with 27 entities, and the range that has the lowest data is in between 5 to 9, 60 to 64, 65 to 69, and 75 to 79 in UPDRSIII_On column as well as from 70 to 74 and 90 to 94 in UPDRSIII_Off, with a single entity listed. Additionally, the right-skew distribution shown in both graphs based on the UPDRSIII_On and UPDRSIII_Off data columns implied us that most patients that were diagosed with Parkinson's disease [experienced mild symptoms as they took medication to alleviate their severity of their symptoms.](https://en.wikipedia.org/wiki/Unified_Parkinson%27s_disease_rating_scale)","metadata":{}},{"cell_type":"markdown","source":"Finally for our data analysis on the subjects_df dataframe, let's dispense the data from the NFOGQ column into our histogram-box graph! All we have to do is to define the fig variable into creating our histogram graph with the px module's histogram function, setting the subjects_df dataframe as the data for the graph, followed by configuring the x parameter to the NFOGQ column for specifying the graph's x-axes, and the marginal parameter to box for creating our box-plot on the top of the histogram. With our histogram-box graph finalized, we use the show function into the fig variable figure for displaying it below the code cell.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(subjects_df, x=\"NFOGQ\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.453210Z","iopub.execute_input":"2023-04-28T06:19:30.453642Z","iopub.status.idle":"2023-04-28T06:19:30.542079Z","shell.execute_reply.started":"2023-04-28T06:19:30.453595Z","shell.execute_reply":"2023-04-28T06:19:30.540712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we observe in the NFOGQ data column distributed into the histogram, we realized that there's a medium peak on the far left of the diagram, however we noticed there's a left-skew distribution displayed on the middle-right of the graph. Specifically, the highest data counted is in the range from 20 to 21 with 37 entities, while the range between 6 and 7 held one data entity, making it as least-counted. And from the box-plot we graph above the histogram, we espied one outlier on the left of the graph as well as seeing how the whiskers in the box plot start in the further left away from the outlier and ends at the far-right of the graph thus catching sight of the asymmetrical box-plot. Other than the look of the box-plot, the first and third quartiles of the NFOGQ column according to the box-plot we graphed is 15 and 22, the median is 19, and the interquartile range is 7. Furthermore, the left-skew distribution shown in the middle-right of the diagram made us explain that [most patients showed their most noticeable FOG symptoms based on their experiences on 6 items that are related to FOG during the previous week, while other patients didn't had their noticeable FOG symptoms.](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8114865/)","metadata":{}},{"cell_type":"markdown","source":"After we visualized the data from the NFOGQ column into our histogram-box graph model, we finished our chapter of analyzing almost all of the data columns in the subjects_df dataframe! What's following from analyzing the subjects_df dataframe is that we're going to visualize some few columns in the tasks_df dataframe, since it contained the metadata full of recorded tasks.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 4: The Tasks to be Envisaged</h2>\n\nOnce we got into envisaging the data from the tasks_df dataframe with Altair, we realized that there are less columns to visualize, making this chapter on the FOG data analysis short. Nevertheless, here are the tasks_df dataframe's columns detailed one-by-one!\n* **Id**: The identifier for each subject.\n* **Begin**: Time specified for when the task began.\n* **End**: Time specified for when the task ended.\n* **Task**: Specifies the seven tasks types in the DeFOG protocol.","metadata":{}},{"cell_type":"markdown","source":"With all of the columns explained concisely, let's first visualize the Id column into a bar chart! We use the alt module's Chart function to create our chart object that contains the tasks_df dataframe for loading the data into the graph, then we use the mark_bar function into the alt module's Chart function to plot the bars into our graph, and then encode our graph's characteristics with the encode function, setting the x parameter to the Id column for specifying the x-axes of our graph, followed by configuring the y parameter to the count function for counting the values of a specfic column into the y-axes.","metadata":{}},{"cell_type":"code","source":"alt.Chart(tasks_df).mark_bar().encode(\n    x=\"Id\",\n    y=\"count()\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.544196Z","iopub.execute_input":"2023-04-28T06:19:30.544586Z","iopub.status.idle":"2023-04-28T06:19:30.635852Z","shell.execute_reply.started":"2023-04-28T06:19:30.544546Z","shell.execute_reply":"2023-04-28T06:19:30.634533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we plotted down the Id column data from the tasks_df dataframe, we espied a lot of data based on the patient's identifier, as the width of the graph we created is longer than some graphs we created in the previous chapters. In other words, the data labels \"2acdf5a450\" and \"af02b83cbf\" are counted the most, with approximately 36 entities, while the data label \"850748a138\" is counted the least, with around 6 entities. Funilly enough, the long graph of data labels extracted from the tasks_df dataframe's Id column hinted us that there's variation of data gathered from the FOG patients in numbers, as it was used for gathering the tasks needed for each FOG patient.","metadata":{}},{"cell_type":"markdown","source":"Now let's then visualize the Begin and End columns individually into a histogram and box-chart graph combined into a subplot! To get started, we characterize the fig1 variable into creating our chart with the alt module's Chart function in which we place the tasks_df dataframe as the data for our chart, then we use the mark_bar function to graph some bars into our chart, followed by encoding our chart's parameters with the encode function, setting the alt module's X function for loading the x-axes of our graph that contains the Begin and End columns separately and the bin parameter set to True (which it implied us that we're creating our histogram), followed by configuring the y parameter to the count function to count the values of a specific column into the y-axis of our chart.\n\nAfter we created our histogram graph in the fig1 variable figure, we characterize another variable, fig2, into creating another chart with the Chart function from the alt module in which we place the tasks_df dataframe into it for inputting the data for the graph, followed by marking the boxplot into our graph with the mark_boxplot function in which we configure the extent parameter to \"min-max\" for enabling whiskers of the box plot, as well as encoding the parameters of our graph with the encode function, setting the x parameter to the Begin and End columns separately and the appearance of the box chart with the properties function that contains the height parameter set to 300 for adjusting the height of our graph. Last but not least, we concatenate our two graphs specified by the fig1 and fig2 variables into a subplot with the concat module from the alt module. ","metadata":{}},{"cell_type":"code","source":"fig1 = alt.Chart(tasks_df).mark_bar().encode(\n    alt.X(\"Begin\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(tasks_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"Begin\",\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.637419Z","iopub.execute_input":"2023-04-28T06:19:30.637827Z","iopub.status.idle":"2023-04-28T06:19:30.750830Z","shell.execute_reply.started":"2023-04-28T06:19:30.637789Z","shell.execute_reply":"2023-04-28T06:19:30.749532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw from distributing the tasks_df dataframe's Begin column into our histogram chart, we realized that the bins peaked from the left side of our diagram, displaying a right-skew distribution. Although we're uncertain about the range that has the lowest data of the Begin column seen in the histogram, the range that has the highest data is between 0 and 500, with around 965 or 970 entities. On the other graph, we noticed that the whisker to the maximum of the data in the Begin column was longer than the whisker from the minimum of the data in the Begin column, as we espied on how the box portion was shifted to the left of the diagram. Specifically, the first and third quartiles of the Begin column specified from the box plot is 341.99 and 1230.84, the median is 742.466, and the interquartile range is 888.85. Additionally, the right-skew display in the histogram and in the box plot we noticed implied us that most of the FOG tasks began earlier than the ones that began later.","metadata":{}},{"cell_type":"code","source":"fig1 = alt.Chart(tasks_df).mark_bar().encode(\n    alt.X(\"End\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(tasks_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"End\",\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.752940Z","iopub.execute_input":"2023-04-28T06:19:30.753385Z","iopub.status.idle":"2023-04-28T06:19:30.861173Z","shell.execute_reply.started":"2023-04-28T06:19:30.753340Z","shell.execute_reply":"2023-04-28T06:19:30.859868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Likewise to the data distribution of the Begin column, we noticed how there's a data peak on the left of the diagram, implying that it displayed a right-skew distribution type. And as we can't determine the range that has the lowest data in the End column distribution, we know that the data range from 0 to 500 has the highest data with around 940 to 945 entities listed. And from the box-plot we created on the right of our histogram, we noticed that just like the box-plot we generated while distributing the Begin column previously, we found out that the right whisker from the box portion to the maximum of the End column data is longer than the left whisker from the minimum of the End column data to the box portion, as the box portion was shifted into the left of the diagram. In other words, the first and third quartiles of the End column gathered from the box plot is 360.209 and 1260.44, the median is 762.44, and the interquartile range is 900.231. Additionally, the right-skew distribution we saw in the histogram gave us a clue that most tasks were finished earlier, while others were completed after a while.","metadata":{}},{"cell_type":"markdown","source":"Now that we distributed the Begin and End columns separately, let's visualize the relationship of the two columns into a 2D histogram scatter plot! Basically, we use the alt module's Chart function to compile a new graph in which we place the tasks_df dataframe inside for loading the data into the graph, followed by marking some circles into our graph with the mark_circle function, as well as encoding the parameters of our graph with the encode function, placing the alt module's X and Y functions that contained the Begin and End columns and the bin parameter set to True for configuring the x and y axes, followed by arranging the size parameter to the count function in strings for counting the values of the two columns.","metadata":{}},{"cell_type":"code","source":"alt.Chart(tasks_df).mark_circle().encode(\n    alt.X(\"Begin\", bin=True),\n    alt.Y(\"End\", bin=True),\n    size=\"count()\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.862681Z","iopub.execute_input":"2023-04-28T06:19:30.863096Z","iopub.status.idle":"2023-04-28T06:19:30.948040Z","shell.execute_reply.started":"2023-04-28T06:19:30.863055Z","shell.execute_reply":"2023-04-28T06:19:30.946647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Similarly to the 2D histogram graph we plotted previously in the data distribution of the events_df dataframe's Init and Completion columns, we found out that the circles were shifted diagonally into the upper-right of the diagram as their size was decreasing in the process. Specifically, the shrinking circles that hauled to the top-right of the diagram implied to us that some tasks begin earlier and then end earlier, while others begin late and then end late.","metadata":{}},{"cell_type":"markdown","source":"As we were dubious about the shrinking diagonal circles made from the Begin and End column distributions, let's dig deeper into the relations between the two columns by visualizing the duration of each task into the histogram-box graph! To begin this process, we create a new column in the tasks_df dataframe, Duration, to find the difference between the End and Begin columns. Thenceforth, we characterize the fig1 varaible to use the alt module's Chart function with the tasks_df dataframe inside to create a new graph with Altair followed by setting the data for our first graph, then using the mark_bar function to graph the bars into our graph, as well as encoding the parameters of our graph with the encode function, setting the alt module's X function that contained the Duration column and the bin parameter set to True for configuring our graph's x-axes, and the y parameter set to the count function for counting the values of a specific column into our graph.\n\nNext, we generate another variable, fig2, to create another graph with the Chart function from the alt module that contains the tasks_df dataframe for loading the data for another graph, then apply the mark_boxplot function to graph a box-chart into our chart in which we configure the extent parameter to \"min-max\" for enabling the whiskers of the box-plot, followed by encoding the characteristics of our graph with the encode function, just simply setting the x parameter to the Duration column as well as configuring the appearance of our graph with the properties function that has the height parameter set to 300 for arranging the height of our graph. And since we completed graphing the fig1 and fig2 variables, we concatenate them together into a subplot graph with the concat function from the alt module.","metadata":{}},{"cell_type":"code","source":"tasks_df[\"Duration\"] = tasks_df[\"End\"] - tasks_df[\"Begin\"]\n\nfig1 = alt.Chart(tasks_df).mark_bar().encode(\n    alt.X(\"Duration\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(tasks_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"Duration\"\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:30.949584Z","iopub.execute_input":"2023-04-28T06:19:30.949996Z","iopub.status.idle":"2023-04-28T06:19:31.072238Z","shell.execute_reply.started":"2023-04-28T06:19:30.949953Z","shell.execute_reply":"2023-04-28T06:19:31.070926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just like the distributions of Begin and End columns we visualized, we found out that the histogram for the Duration column we made in the tasks_df dataframe displayed a right-skew distribution since there's a data peak on the left of the diagram. Aside from that, the range that has the highest data is from 0 to 50, with 2480 entities, while the other range that is between 350 and 400 held the least data, with approximately 5 entities. And from the box plot we graphed on the right of our histogram, we came across on how the width of the right whisker is greater than the width of the left whisker, since the box is shifted to the left of the diagram. Besides, the first and third quartiles we saw in the Duration column according to the box-plot we graphed is 8.75 and 30.28, the median is 15.04, and the interquartile range is 21.53. Additionally, the right-skewed distribution we saw in the histogram based on the Duration data column implied us that most tasks lasted for a shorter time, while a few others lasted for a longer time.","metadata":{}},{"cell_type":"markdown","source":"Finally for this data analysis of the tasks_df dataframe we're in, let's distribute the data in the Task column into a bar graph! Prior to plotting the data from the Task column out, we create another dataframe, task, to tally the total values of the tasks_df dataframe's task column, and then redefine it to create a dataframe with the pd module's DataFrame function, setting the dictionary in which the task key is assigned to the indexes specified from the index attribute that is plugged into the task dataframe and the count key to the values specified from the values attribute that is plugged into the task dataframe.\n\nFollowing from this step, we use the alt module's Chart function to generate our graph in which we set the task dataframe inside for loading the data for our graph, followed by marking the bars into our graph with the mark_bar function as well as adjusting the parameters of our graph with the encode function, setting the x parameter to the task column for loading the x-axes of our graph, and the y parameter to the count column for counting the values of a specific column into our graph's y-axes.","metadata":{}},{"cell_type":"code","source":"task = tasks_df[\"Task\"].value_counts()\ntask = pd.DataFrame({\n    'task': task.index,\n    'count': task.values\n})\n\nalt.Chart(task).mark_bar().encode(\n    x='task',\n    y='count'\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:31.074163Z","iopub.execute_input":"2023-04-28T06:19:31.074539Z","iopub.status.idle":"2023-04-28T06:19:31.130796Z","shell.execute_reply.started":"2023-04-28T06:19:31.074502Z","shell.execute_reply":"2023-04-28T06:19:31.129518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we sorted the data from the tasks_df dataframe's Task column, we found out that there are subvariants of the seven task types, which was specified from 4MW, Hotspot{1-2}, MB{0-9}, Rest{1-2}, TUG, and Turning as there are some tall and short bars seen in the diagram. Out of the high and low bars shown in the bar chart, the data label \"TUG-ST\" is counted the most with almost 280 entities, while the other data label, \"MB6\" is counted the least with around 10 data entities. Specifically, the bars that are counted high and low in the tasks_df dataframe's Task column implied to us that the task types vary in each DeFOG protocol, whether each patient had severe or less-severe FoG symptoms.","metadata":{}},{"cell_type":"markdown","source":"And as we complete the data visualization in the Task column, we are now finished with the data analysis in the tasks_df dataframe! What we envisage our end-of-sight in our data analysis of the FoG Prediction competition is that we're going to read and analyze a specific parquet file in the \"unlabeled\" folder followed by visualizing each csv file in the \"train\" folder.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Chapter 5: CSVs and Parquets, Annotating Forthwith</h2>\n\nAfter we trekked into visualizing five dataframes, including the ones in the daily, defog, and tdcsfog dataframes, we came across some parquet files that was included in the FoG Prediction competition (as it was stored under the \"unlabeled\" folder), followed by visualizing some csv files under the train folder, as some were held under the \"defog\", \"notype\", and \"tdcsfog\" folders.","metadata":{}},{"cell_type":"markdown","source":"In other words, let's visualize one of the parquet file from the \"unlabeled\" directory and each csv file that is stored by each folder in the \"train\" directory! To get started, we create four additional dataframes, the unlabeled_df dataframe from reading one of the parquets in the \"unlabeled\" directory with the read_parquet function, followed by creating the train_defog_df, train_tdcsfog_df, and train_notype_df dataframes from reading each csv file under each folder with the read_csv function. And since we created all four dataframes, we use the head function into each four dataframes for displaying the first five rows.","metadata":{}},{"cell_type":"code","source":"unlabeled_df = pd.read_parquet(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/unlabeled/28e6c306ba.parquet\")\n\ntrain_defog_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/defog/02ea782681.csv\")\ntrain_notype_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/notype/60dfb26b2c.csv\")\ntrain_tdcsfog_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/tdcsfog/02e8454f57.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:31.132398Z","iopub.execute_input":"2023-04-28T06:19:31.133372Z","iopub.status.idle":"2023-04-28T06:19:49.114118Z","shell.execute_reply.started":"2023-04-28T06:19:31.133333Z","shell.execute_reply":"2023-04-28T06:19:49.112813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unlabeled_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.115870Z","iopub.execute_input":"2023-04-28T06:19:49.116312Z","iopub.status.idle":"2023-04-28T06:19:49.129932Z","shell.execute_reply.started":"2023-04-28T06:19:49.116269Z","shell.execute_reply":"2023-04-28T06:19:49.128485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_notype_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.131671Z","iopub.execute_input":"2023-04-28T06:19:49.132062Z","iopub.status.idle":"2023-04-28T06:19:49.148323Z","shell.execute_reply.started":"2023-04-28T06:19:49.132027Z","shell.execute_reply":"2023-04-28T06:19:49.146943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_defog_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.150181Z","iopub.execute_input":"2023-04-28T06:19:49.151325Z","iopub.status.idle":"2023-04-28T06:19:49.169426Z","shell.execute_reply.started":"2023-04-28T06:19:49.151268Z","shell.execute_reply":"2023-04-28T06:19:49.168029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_tdcsfog_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.171366Z","iopub.execute_input":"2023-04-28T06:19:49.171801Z","iopub.status.idle":"2023-04-28T06:19:49.193664Z","shell.execute_reply.started":"2023-04-28T06:19:49.171760Z","shell.execute_reply":"2023-04-28T06:19:49.192290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Analysis on the Four Dataframe's Metadata</h3>\n\nAs we finish setting up the four dataframes from three csv files and one parquet file, let's take a look at the basic metadata of them by looking at the overall number of data entities by using the len function to each of the four dataframes!","metadata":{}},{"cell_type":"code","source":"len(unlabeled_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.195814Z","iopub.execute_input":"2023-04-28T06:19:49.196294Z","iopub.status.idle":"2023-04-28T06:19:49.208570Z","shell.execute_reply.started":"2023-04-28T06:19:49.196244Z","shell.execute_reply":"2023-04-28T06:19:49.207510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_notype_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.209869Z","iopub.execute_input":"2023-04-28T06:19:49.211233Z","iopub.status.idle":"2023-04-28T06:19:49.218631Z","shell.execute_reply.started":"2023-04-28T06:19:49.211150Z","shell.execute_reply":"2023-04-28T06:19:49.217605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_defog_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.220183Z","iopub.execute_input":"2023-04-28T06:19:49.220832Z","iopub.status.idle":"2023-04-28T06:19:49.230642Z","shell.execute_reply.started":"2023-04-28T06:19:49.220780Z","shell.execute_reply":"2023-04-28T06:19:49.229446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_tdcsfog_df)","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.232155Z","iopub.execute_input":"2023-04-28T06:19:49.232729Z","iopub.status.idle":"2023-04-28T06:19:49.242932Z","shell.execute_reply.started":"2023-04-28T06:19:49.232678Z","shell.execute_reply":"2023-04-28T06:19:49.241893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As for the number of data entities in the unlabeled_df dataframe, we noticed that there are 60479559 entities inside, meaning that it held a lot of unannotated data from the \"daily\" dataset (indicates as daily_df dataframe). On the other hand, we counted 384660 entities in the train_notype_df dataframe, 162907 entities in the train_defog_df dataframe, and 3749 data entities in the train_tdcsfog_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"Let's now visualize the number of missing values in all four dataframes! To find the missing values, we use the isna function and plug it into each of the four dataframes, then we apply the sum function to sum up all missing values that are present in each four dataframes.","metadata":{}},{"cell_type":"code","source":"unlabeled_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.244392Z","iopub.execute_input":"2023-04-28T06:19:49.245013Z","iopub.status.idle":"2023-04-28T06:19:49.806808Z","shell.execute_reply.started":"2023-04-28T06:19:49.244969Z","shell.execute_reply":"2023-04-28T06:19:49.805427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_defog_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.808720Z","iopub.execute_input":"2023-04-28T06:19:49.809493Z","iopub.status.idle":"2023-04-28T06:19:49.822706Z","shell.execute_reply.started":"2023-04-28T06:19:49.809440Z","shell.execute_reply":"2023-04-28T06:19:49.821455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_notype_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.824228Z","iopub.execute_input":"2023-04-28T06:19:49.824989Z","iopub.status.idle":"2023-04-28T06:19:49.839901Z","shell.execute_reply.started":"2023-04-28T06:19:49.824952Z","shell.execute_reply":"2023-04-28T06:19:49.838737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_tdcsfog_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.841430Z","iopub.execute_input":"2023-04-28T06:19:49.841886Z","iopub.status.idle":"2023-04-28T06:19:49.852449Z","shell.execute_reply.started":"2023-04-28T06:19:49.841839Z","shell.execute_reply":"2023-04-28T06:19:49.851004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unlike our previous calculations of missing values from the previous six dataframes, we found no missing values in the four dataframes we compiled. Specifically, it gave us a hint that the data series in each of the four dataframes was gathered accurately from each patient that has the Freezing of Gait symptoms of Parkinson's.","metadata":{}},{"cell_type":"markdown","source":"Lastly for our short surface analysis of the four dataframes, let's visualize the number of columns in each of them! We simply use the shape attribute to each of the four dataframes and then extract the last index of it with the slice index of 1.","metadata":{}},{"cell_type":"code","source":"unlabeled_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.853958Z","iopub.execute_input":"2023-04-28T06:19:49.854826Z","iopub.status.idle":"2023-04-28T06:19:49.860973Z","shell.execute_reply.started":"2023-04-28T06:19:49.854785Z","shell.execute_reply":"2023-04-28T06:19:49.860062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_defog_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.862204Z","iopub.execute_input":"2023-04-28T06:19:49.863272Z","iopub.status.idle":"2023-04-28T06:19:49.873193Z","shell.execute_reply.started":"2023-04-28T06:19:49.863221Z","shell.execute_reply":"2023-04-28T06:19:49.871802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_notype_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.874730Z","iopub.execute_input":"2023-04-28T06:19:49.875119Z","iopub.status.idle":"2023-04-28T06:19:49.885825Z","shell.execute_reply.started":"2023-04-28T06:19:49.875076Z","shell.execute_reply":"2023-04-28T06:19:49.884546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_tdcsfog_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.887515Z","iopub.execute_input":"2023-04-28T06:19:49.887929Z","iopub.status.idle":"2023-04-28T06:19:49.896973Z","shell.execute_reply.started":"2023-04-28T06:19:49.887889Z","shell.execute_reply":"2023-04-28T06:19:49.895685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the unlabeled_df dataframe, we counted four columns, while in the train_defog_df dataframe, we counted nine columns. However, we counted seven columns in the train_notype_df and train_tdcsfog_df dataframes. Funnily enough, the number of columns we counted in all four dataframes hinted us that they store the different types of data based from the acceleration of the FoG patient's kinetic movement, their turning, walking, and hesitiating, as well as the patients' tasks.","metadata":{}},{"cell_type":"markdown","source":"With all the basic metadata of the four dataframes explained with detail, let's proceed into visualizing the data from each dataframe at a time in the next four sections of the chapter we're in!","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">unlabeled_df</h3>\n\nAs we look into the unlabeled_df dataframe, we understood that it contain the data that doesn't have the annotations recorded thus containing huge numbers of data and was stored in a parquet file format, so that it will not gauge up the memory in the Jupyter Notebook. Besides, let's detail the columns in the unlabeled_df dataframe one by one!\n* **Time**: Data specified for each timestep.\n* **AccV**: Specifies vertical acceleration.\n* **AccML**: Specifies mediolateral acceleration.\n* **AccAP**: Specifies anteroposterior acceleration.","metadata":{}},{"cell_type":"markdown","source":"With all four columns explained in concise detail, let's start our short data analysis in the unlabeled_df dataframe by distributing the AccV, AccML, and AccAP columns into three histograms in a subplot! First, we characterize the fig variable figure into creating our subplot chart with the make_subplots function, setting the rows parameter to 1 and the cols parameter to 3 for configuring our subplot with three columns in one row. \n\nFollowing from creating our subplot, we use the append_trace function into the fig variable three times, setting the go module's Histogram function for configuring our histogram model into each three subplots as it contained the x parameter being set to the unlabeled_df dataframe's AccV, AccML, and AccAP columns separately but has the head function with the value 50000 just to limit the first 50000 rows of data for the histogram's x-axes as well as the name parameter to the individual strings \"AccV\", \"AccML\", and AccAP\" for labelling the graph in each subplot, followed by configuring the row parameter to 1 and the col parameter to 1, 2, and 3 separately for placing the histograms into each subplot. Finally, we use the show function to display our subplot graph into the output cell.","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=3)\n\nfig.append_trace(go.Histogram(x=unlabeled_df[\"AccV\"].head(50000), name=\"AccV\"), row=1, col=1)\nfig.append_trace(go.Histogram(x=unlabeled_df[\"AccML\"].head(50000), name=\"AccML\"), row=1, col=2)\nfig.append_trace(go.Histogram(x=unlabeled_df[\"AccAP\"].head(50000), name=\"AccAP\"), row=1, col=3)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.900636Z","iopub.execute_input":"2023-04-28T06:19:49.901420Z","iopub.status.idle":"2023-04-28T06:19:49.997861Z","shell.execute_reply.started":"2023-04-28T06:19:49.901378Z","shell.execute_reply":"2023-04-28T06:19:49.994701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the data distribution of the AccV column, we found out that the there's a data peak on the right of the diagram, as it displayed a left-skew distribution since one thin bar was placed on the near-right of the tall thin bar. Not only that, the highest data counted is in the range from 0.21 to 0.215 with 29058 entities, while there are some ranges that had one data entity, making them having the least data counted. In other words, the right-skewed distribution shown in the AccV column hinted us that [most of the vertical accelerations are slow, as it represents how patients freeze while walking with their knees trembling.](https://jneuroengrehab.biomedcentral.com/articles/10.1186/1743-0003-10-19)\n\nMeanwhile from the AccML data distribution, we clearly see two separate peaks on the left on the histogram as the far left bar is taller than the one on the right of the left bar, making this distribution as a right-skew type. Besides, the range that has the highest data counted is between -0.9 and -0.88, with 32156 entities, while there are some other ranges that contain one data entity, as they has lowest data counted. Furthermore, the right-skewed distribution we saw in the AccML column gave us some clues [that most of the FOG patients showed significantly slower gait velocities and shorter step strides, as there are signs of severe bradykinesia during gaits.](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9470496/)\n\nLastly for the data distribution of the AccAP column, we espied two thin data peaks as the peak on the right is taller than the one at the left, indicating this as a left-skew distribution. Specifically, the most-counted data is in the range between 0.4 and 0.42 with 27817 entities, while there's one data counted in some ranges as it was the least-counted data likewise to the data distributions of AccV and AccML columns. Additionally, the left-skew distribution shown in the AccAP column implied to us that the [almost all patients with Parkinson's FOG presented decreased acceleration immensity and slower turns.](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6729092/)","metadata":{}},{"cell_type":"markdown","source":"After we distributed the AccV, AccML, and AccAP columns separately, let's then visualize the relation between the three columns into the 3D scatter plot by the Time column! We characterize the fig variable figure into creating the 3D scatter plot with the px module's scatter_3d function, setting the unlabeled_df dataframe as our data for the graph but truncated to 50000 for not crashing our notebook, followed by configuring the x, y, and z parameters to the AccV, AccML, and AccAP columns for arranging the x, y, and z axes, and the color parameter to the Time column for color coding the plots. With that completed, we use the show function into the fig variable graph figure for displaying our graph down from our code cell below.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_3d(unlabeled_df.head(50000), x=\"AccV\", y=\"AccML\", z=\"AccAP\", color=\"Time\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:49.999596Z","iopub.execute_input":"2023-04-28T06:19:50.000025Z","iopub.status.idle":"2023-04-28T06:19:50.190220Z","shell.execute_reply.started":"2023-04-28T06:19:49.999983Z","shell.execute_reply":"2023-04-28T06:19:50.189160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since we distributed the AccML, AccAP, and AccV columns all at once into the 3D scatter plot we created, we realized that most data plots clustered in the left of the AccML axes, while we spotted one outlier on the right of the AccV and AccML columns. Not only that, we noticed that most of the plots we graphed were colored in orange although we saw some of the purple colorred plots, meaning that the tasks gathered in the unlabeled_df was recorded on an interval between approximately 35k to 40k. In other words, the cluster of data shown on the left of the AccML axes hinted us that most of the Parkinson's FoG patients showed slower vertical accelerations as well as the gait velocities in medio-lateral speeds.","metadata":{}},{"cell_type":"markdown","source":"As we finished visualizing all four columns in the unlabeled_df dataframe, let's proceed into visualizing the data columns inside the train_defog_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">train_defog_df</h3>\n\nAs we got into visualizing the train_defog_df dataframe, it contains most of the event-type annotations that has the comprising data series collected in the subject's home as they completed a FOG-provoking protocol, similarly to the defog_df dataframe we visualized previously. Aside from the prologue overview of the train_defog_df dataframe, let's go over the columns one by one!\n* **Time**, **AccV**, **AccML**, **AccAP**: See previous visualizations of the unlabeled_df dataframe.\n* **StartHesitation**, **Turn**, **Walking**: Specifies the indicator variables of the occurances of each event types.\n* **Valid**: Specifies whether the event annotations is unambiguous.\n* **Task**: Specifies whether the event was annotated.","metadata":{}},{"cell_type":"markdown","source":"Without further ado, let's jump into visualizing the train_defog_df dataframe by distributing the AccV, AccML, and AccAP data columns all at once into our subplot graph! First, we define our variable, fig, into creating our subplots with the make_subplots function, setting the rows parameter to 1 and the cols parameter to 3 for configuring our subplot to have one row and three columns. After we created our subplot template, we add the traces into our subplot with the add_trace function plugged into the fig variable figure, setting the go module's Histogram function to generate our histogram, in which the x parameter is set to the train_defog_df dataframe's AccV, AccML, and AccAP columns separately as well as the name parameter set to the columns specfied from the configuration of the x parameter, and on the outside of the go module's Histogram setup, we configure our row parameter to 1, and the col parameter to 1, 2, and 3 separately for placing each histogram graph we created into each subplot's part. Finally, we use the show function into the fig variable figure for displaying our graph into the code cell output.","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=3)\n\nfig.add_trace(go.Histogram(x=train_defog_df[\"AccV\"], name=\"AccV\"), row=1, col=1)\nfig.add_trace(go.Histogram(x=train_defog_df[\"AccML\"], name=\"AccML\"), row=1, col=2)\nfig.add_trace(go.Histogram(x=train_defog_df[\"AccAP\"], name=\"AccAP\"), row=1, col=3)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:50.191459Z","iopub.execute_input":"2023-04-28T06:19:50.193399Z","iopub.status.idle":"2023-04-28T06:19:50.379107Z","shell.execute_reply.started":"2023-04-28T06:19:50.193301Z","shell.execute_reply":"2023-04-28T06:19:50.377665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the AccV data distribution we visualized, we found out that there's a tall data peak on the mid-right of our diagram, as it narrowly showed the right-skew distribution since we observed a little data peak on the right of the tall data spike. Additionally, the range that has the highest-data counted is between -1.001 and -0.999 with 68402 entities, while there are some ranges that has one data entity, making them have the least counted data. Specifically, the right-skew data distribution in -1 of the x-axes from the AccV column hinted us that there are most events in which FoG patients walked slowly in the opposite directions when they freeze while they had their knees trembling.\n\nMeanwhile in the AccML data distribution, we noticed that there's another tall data peak on the near middle-left of the histogram graph, as it merely showed a left-skew distribution type because of the two tall peaks on the left of the tallest bin. In other way of explaination, the highest data is in the range from 0.046 to 0.048 with 27586 entities, while we mostly see some ranges that has one entity, as they have the least counted data. Furthermore, the left-skewed data distribution we spotted from the AccML column gave us clues that there are most events of how FoG patients showed slower gait velocities and step strides possibly because of bradykinesia.\n\nLastly for the distribution of the AccAP column, we caught glimpse of the three peaks of data on the far right of the histogram chart, as the highest data peak displayed a left-skew distribution because of the small ascending bins of data we noticed from the left of the highest peak. Besides, the range -0.25 to -0.245 is counted the most, with 28083 data entities, while likewise to what we saw from the data distributions of AccV and AccML, some ranges in the AccAP column is counted the least because of one data entity. Funilly enough, the left-skewed distribution of the AccAP column we glimpsed at the right of the histogram implied to us that there are most events of how FoG patients showed decreased acceleration and slower turns as we noted how there's a data peak in the negative ranges from 0 on the left of the histogram.","metadata":{}},{"cell_type":"markdown","source":"Speaking about the separate distributions of the AccV, AccML, and AccAP data columns, let's visualize the relation of them into distributing them into the 3D scatter plot all at once by the Time column! We simply use the fig variable figure into creating our three-dimensional scatter plot with the px module's scatter_3d function, setting the train_defog_df dataframe as the data for the 3D scatter plot, followed by configuring the x, y, and z parameters into the AccV, AccML, and AccAP columns for arranging the x, y, and z axes as well as the color parameter set to the Time column for specifying the color legend. Once we complete our scatter plot figure, we use the show function into the fig variable figure for displaying our graph into the code output below.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_3d(train_defog_df, x=\"AccV\", y=\"AccML\", z=\"AccAP\", color=\"Time\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:50.380766Z","iopub.execute_input":"2023-04-28T06:19:50.381172Z","iopub.status.idle":"2023-04-28T06:19:50.638572Z","shell.execute_reply.started":"2023-04-28T06:19:50.381129Z","shell.execute_reply":"2023-04-28T06:19:50.637270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw in the 3D scatter plot we graphed from the relations between the data of the AccAP, AccML, and AccV columns, we glimpsed on how there's a cluster of plots in the middle of the diagram, as some of them were color-coded as violet, purple, orange, and yellow. Additionally, the cluster of data plots in the middle of the 3D scatter plot we graphed hinted us that there are variations of each timestep throughout the recordings of each event, as they record on how the FoG patients showed slow gait velocities as well as decreased turns and acceleration, as they [showed the signs of bradykinesia symptoms in Parkinson's disease.](https://www.parkinson.org/understanding-parkinsons/movement-symptoms/bradykinesia)","metadata":{}},{"cell_type":"markdown","source":"Let's then visualize the distributions of the StartHesitation, Turn, and Walking columns all at once into three histograms in a subplot! For creating our subplot chart, we characterize the fig variable into the make_subplots function, setting the rows and cols parameters to 1 and 3 for configuring our row with three columns into our subplot. Following from creating our subplot, we use the add_trace function into the fig variable figure for placing our graphs into our subplot, setting the go module's Histogram function for creating our histogram in which the x parameter is set to the train_defog_df dataframe's StartHesitation, Turn, and Walking columns separately for configuring the x-axes of the histogram as well as the name parameter set into the same name of the specified columns in the x parameter, and on outside of the go module's Histogram function setup, we configure our row parameter to 1 and the col parameter to 1, 2, and 3 separately for placing our graphs into the first row and each of three columns. And as we finish setting up our graph with three subplots, we use the show function into the fig variable figure for displaying the graph down below the code.","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=3)\n\nfig.add_trace(go.Histogram(x=train_defog_df[\"StartHesitation\"], name=\"StartHesitation\"), row=1, col=1)\nfig.add_trace(go.Histogram(x=train_defog_df[\"Turn\"], name=\"Turn\"), row=1, col=2)\nfig.add_trace(go.Histogram(x=train_defog_df[\"Walking\"], name=\"Walking\"), row=1, col=3)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:50.640021Z","iopub.execute_input":"2023-04-28T06:19:50.640642Z","iopub.status.idle":"2023-04-28T06:19:50.724082Z","shell.execute_reply.started":"2023-04-28T06:19:50.640598Z","shell.execute_reply":"2023-04-28T06:19:50.722877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the StartHesitation and Walking columns we distributed, we realized that all 162907 data entities in the two specified columns (StartHesitation and Walking) were marked as 0, as there are no event occurences of the FoG patients hesitating and walking. However for the Turn column data distribution, we glimpsed that the data in the Turn column distribution was labeled as 0 more than the data that is marked as 1, since there are 162777 data entities for the data marked as 0 and 130 data entities for the data that is labeled as 1. Specifically, the data that is mostly labeled as 0 in the Turn column gave us some clues that there are most events in which the FoG patient didn't turn, while there are a few events in which the FoG patient turns.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the data relations between the StartHesitation, Turn, and Walking columns into the 3D scatter plot! We define the fig variable to create our 3D scatter plot with the scatter_3d function from the px module, setting the train_defog_df dataframe as the data for the three-dimensional scatter plot, followed by configuring the x, y, and z parameters into the StartHesitation, Turn, and Walking columns for arranging the x, y, and z axes in the 3D scatter plot. With all of our 3D scatter plot completed, we display our graph down below by plugging the show function into the fig variable figure.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter_3d(train_defog_df, x=\"StartHesitation\", y=\"Turn\", z=\"Walking\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:50.725544Z","iopub.execute_input":"2023-04-28T06:19:50.726525Z","iopub.status.idle":"2023-04-28T06:19:50.833666Z","shell.execute_reply.started":"2023-04-28T06:19:50.726473Z","shell.execute_reply":"2023-04-28T06:19:50.832339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, we glimpsed that there's two plots shown on the scatter plot, as one was at the left of the Turn column axes and the other was at the right of the Turn column axes while they were centered at the middle of the Walking and StartHesitation column axes. Specifically, the two plots shown in the 3D scatter plot hinted us that there are some events in which the FoG patient didn't turn, walk, or hesitate, while there are other events in which the FoG patient only turn. However, that was one of the events we visualized in one of the patients that had Parkinson's and may differ from what we visualized in the train_defog_df dataframe, as the values in the Walking, Turn, and StartHesitation columns may have more or less zeros or ones.","metadata":{}},{"cell_type":"markdown","source":"Lastly for our visualization in the train_defog_df dataframe, let's envisage the Valid and Task columns into both pie charts in a subplot! Before we begin plotting both pie charts into our subplot, we characterize the defog_valid_df and defog_task_df dataframes into counting the values inside the train_defog_df dataframe's Valid and Task columns with the value_counts function for extracting the values and indexes from each column for plotting our graphs out.\n\nThenceforth, we create and assign the fig variable to create our subplots with the make_subplots function, setting the rows and cols parameter to 1 and 2 for configuring our subplot with one row and two columns, and the specs parameter to a list containing two dictionaries as the type key in strings is assigned to \"pie\" in strings so that Plotly will not throw an error about placing pie charts into the subplot. We then add the graphs into our subplot by applying the add_trace function into the fig variable figure, setting the go module's Pie function for creating our pie graph with the values parameter set to the values specified from the values attribute plugged into the defog_valid_df and defog_task_df dataframes separately, the labels parameter set to the indexes specified from the index attribute that is applied into the defog_valid_df and the defog_task_df dataframes individually, and the name parameter to the column names Valid and Task independently for configuring the labels, values of each pie chart as well as naming each graph so that it'll not confuse a person, and on outside from arranging the go module's Pie function, we configure the row parameter to 1 and the col parameter to 1 and 2 one by one! Finally, we use the show function into the fig variable figure for exhibiting our graph.","metadata":{}},{"cell_type":"code","source":"defog_valid_df = train_defog_df[\"Valid\"].value_counts()\ndefog_task_df = train_defog_df[\"Task\"].value_counts()\n\nfig = make_subplots(rows=1, cols=2, specs=[[{\"type\": \"pie\"}, {\"type\": \"pie\"}]])\nfig.add_trace(go.Pie(values=defog_valid_df.values, labels=defog_valid_df.index, name=\"Valid\"), row=1, col=1)\nfig.add_trace(go.Pie(values=defog_task_df.values, labels=defog_task_df.index, name=\"Task\"), row=1, col=2)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:50.835278Z","iopub.execute_input":"2023-04-28T06:19:50.835766Z","iopub.status.idle":"2023-04-28T06:19:50.868990Z","shell.execute_reply.started":"2023-04-28T06:19:50.835725Z","shell.execute_reply":"2023-04-28T06:19:50.867868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we compiled our subplot graph to have two pie charts, we realized that both of them held 77.4% of the data that is labeled as false, as well as 22.6% of the data being marked as true. However, from the Valid column, there are 126034 data entities of the \"false\" label and 36873 data entities of the \"true\" label, while in the Task column, there are 126013 data entities of the \"true\" label and 36894 data entities of the \"false\" label. Specifically, the number of the data that is classified as \"false\" hinted us that most of the events isn't umambiguous as well as having most of the events annotated. And likewise to the EventAnnotation, Turn, and Walking columns distributions we visualized, the trues and falses of the Valid and Task column from each patient varies, as some had more falses than trues in either Valid or Task columns or both. ","metadata":{}},{"cell_type":"markdown","source":"As we finished visualizing all of the columns from the train_defog_df dataframe, let's jump into visualizing the train_notype_df dataframe data but we switch into static plotting so that we'll not get our notebook RAM laggy!","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">train_notype_df</h3>\n\nOnce we head into the train_notype_df dataframe to prepare visualizing the data inside, we understood that it contained the data full of event recordings that was from the defog dataset, but it lacks the event-type annotations. In summary of the background information, let's detail the columns inside the train_notype_df dataframe one-by-one!\n* **Time**, **AccV**, **AccAP**, **AccML**, **Valid**, **Task**: Same columns goes to the train_defog_df and unlabeled_df dataframes. See previous visualizations of both.\n* **Event**: Specifies the indicator variables for each occurence of any FoG-type event.","metadata":{}},{"cell_type":"markdown","source":"Now with our explaination in each column done, let's switch to plotting the data in the train_notype_df dataframe with Seaborn and Matplotlib by importing the matplotlib module's pyplot attribute to plt and the seaborn module as sns, so that we'll never get our notebook glitchy and not fill up more RAM in the notebook.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:50.872759Z","iopub.execute_input":"2023-04-28T06:19:50.873863Z","iopub.status.idle":"2023-04-28T06:19:51.442651Z","shell.execute_reply.started":"2023-04-28T06:19:50.873816Z","shell.execute_reply":"2023-04-28T06:19:51.441492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following from importing the modules, let's begin analyzing the train_notype dataframe by envisaging and distributing the AccV, AccML, and AccAP data columns into the three histogram-kde charts in a subplot! To get started, we characterize the fig and axes variables into creating our subplot graph with the plt module's subplots function, setting the values 1 and 3 for configuring a row and three columns into our subplot as well as the figsize parameter to 15 and 5 inside the parentheses for adjusting the width and height of the subplot, and the sharey parameter to True for letting the subplot share the y-axes of each graph. We then create three histograms with the sns module's histplot, setting the train_notype_df dataframe as the data for the graph, the x parameter to the AccV, AccAP, and AccML columns separately for configuring the x-axes, the ax parameter to the axes variable's indexes of 0, 1, and 2 individually for placing the graphs into each different subplot.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(15, 5), sharey=True)\n\nsns.histplot(train_notype_df, x=\"AccV\", kde=True, ax=axes[0])\nsns.histplot(train_notype_df, x=\"AccAP\", kde=True, ax=axes[1])\nsns.histplot(train_notype_df, x=\"AccML\", kde=True, ax=axes[2])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:19:51.443994Z","iopub.execute_input":"2023-04-28T06:19:51.445278Z","iopub.status.idle":"2023-04-28T06:20:03.314677Z","shell.execute_reply.started":"2023-04-28T06:19:51.445229Z","shell.execute_reply":"2023-04-28T06:20:03.313554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the AccV column we visualized in the left of our subplot, we glimpsed a data peak on the right of the diagram, as it showed a right-skewed distribution since the data peak was shifted to the left according to the kernel-density estimation outline. In other ways of explaination, the highest data count is in somewhere between -1 and -0.8, with around approximately 15100 entities, while there are some ranges that held one data entity, as they store the least counted data. Additionally, the right-skewed data peak shown on the AccV column hinted us that there's most events in which FoG patients walked slowly in the opposite directions when freezing while having their knees trembling.\n\nMeanwhile from the AccAP data distribution in the middle of the subplot, we glimpsed on how there are around 4 data peaks in the left of the histogram, as the second tallest data peak that is on the right of the tallest data peak resembled a right-skew data distribution type. Besides from the display of the AccAP distribution, the range that has the highest data is in approximately between -0.7 to -0.4, with around 20950 entities, while there are some ranges that has one data entity around the AccAP data distribution, as they contained the lowest data. Furthermore, the right-skewed data we found from the AccAP column hinted us that there are most events that showed decreased acceleration as well as slower gait turns, as it hinted the signs of a freezing of gait event.\n\nAnd from the AccML distribution, we espied a single data peak in the near middle-left of the histogram, and it displayed a near symmetrical distribution with no left or right skewing. Other than that, the highest data counts is in the data range from approximately -0.11 to 0.04, with roughly 6060 entities, while the least counted data is in some ranges of the AccML column because of only one entity inside of them, likewise to the AccV and AccAP columns. In addition, the symmetrical distribution we saw in the AccML column implied to us that it showed the events of a specific FoG patient showing the symptoms of bradykinesia because of their slower gait velocities and step strides.","metadata":{}},{"cell_type":"markdown","source":"Now let's then visualize the Event and Valid columns into two bar charts in a subplot! Before we begin plotting this out, we create two dataframes, notype_valid_df and notype_event_df, to count the values of the train_notype_df dataframe's Event and Valid columns with the value_counts function. We then redefine the notype_valid_df and notype_event_df dataframes into creating an actual dataframe with the pd module's DataFrame function, setting a dictionary in which the index key is assigned to the indexes of the notype_valid_df and the notype_event_df dataframes with the index attribute and the count key is assigned to the values specified by the values attribute plugged into the notype_event_df and notype_valid_df dataframes.\n\nThenceforth, we characterize the fig and axes variables into creating new subplots with the plt module's subplots function, setting 1 and 2 for configuring one row and two columns as well as the figsize parameter to 10 and 5 in parentheses for adjusting the size of our subplot to 10x5 and the sharey parameter to True for making our plots share the y-axes in the subplot. We then add the bar plots into our subplot with the sns module's barplot function, setting the notype_valid_df and notype_event_df dataframes as the data for the bar plot separately, followed by configuring the x and y parameters to the indexes and values of the individual notype_valid_df and notype_event_df dataframes with the index and values attributes for arranging the x and y axes of the bar chart and the ax parameter to the axes variable's indexes of 0 and 1.","metadata":{}},{"cell_type":"code","source":"notype_valid_df = train_notype_df[\"Valid\"].value_counts()\nnotype_event_df = train_notype_df[\"Event\"].value_counts()\n\nnotype_valid_df = pd.DataFrame({\n    \"index\": notype_valid_df.index,\n    \"count\": notype_valid_df.values\n})\n\nnotype_event_df = pd.DataFrame({\n    \"index\": notype_event_df.index,\n    \"count\": notype_event_df.values\n})\n\nfig, axes = plt.subplots(1, 2, figsize=(10,5), sharey=True)\nsns.barplot(notype_event_df, x=\"index\", y=\"count\", ax=axes[0])\nsns.barplot(notype_valid_df, x=\"index\", y=\"count\", ax=axes[1])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:20:03.316253Z","iopub.execute_input":"2023-04-28T06:20:03.316860Z","iopub.status.idle":"2023-04-28T06:20:03.613165Z","shell.execute_reply.started":"2023-04-28T06:20:03.316822Z","shell.execute_reply":"2023-04-28T06:20:03.611784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the data distribution of the train_notype_df dataframe's Event column, we found out that most of the data is marked as 0 than the data being marked as 1. In other words of explaination, the data label 0 has almost 250000 entities, while the data label 1 has approximately 141800 entities. Additionally, the data that is marked as 0 in the train_notype_df dataframe's Event column hinted us that there are mostly no event occurences of any FoG-type event from a specific patient.\n\nAnd from what we saw from the train_notype_df dataframe's Valid column data distribution, we found out that the data being marked as True is almost greater than the data being labeled as False, in which the data being marked as False is counted the most than the data that is marked as True. Specifically, the data that is marked as False has almost 200100 entities, while the data marked as True has approximately 158500 entities. Furthermore, the data that is marked as False in the train_notype_df dataframe's Valid column gave us some clues that almost some of the events from a specific patient we analyzed aren't unambiguous, compared to what we saw in the previous distributions of the train_defog_df dataframe's Valid column.","metadata":{}},{"cell_type":"markdown","source":"Lastly, let's visualize and distribute the data from the train_notype_df dataframe's Task column into a standalone bar-plot! Before we begin plotting this data out, we characterize another dataframe, notype_task_df, into tallying the values in the train_notype_df dataframe's Task column with the value_counts function. After that, we redefine it to create the actual dataframe with the pd module's DataFrame function, setting a dictionary in which it has the index key being assigned into the notype_task_df dataframe's index specified with the index attribute and the count key being defined into the values specfied by the values attribute being plugged into the notype_task_df dataframe.\n\nWe then create our bar plot with the sns module's barplot function, setting the notype_task_df dataframe as the data for the bar chart, followed by configuring the x and y parameters to the index and count columns for specifying our bar chart's x and y axes.","metadata":{}},{"cell_type":"code","source":"notype_task_df = train_notype_df[\"Task\"].value_counts()\nnotype_task_df = pd.DataFrame({\n    \"index\": notype_task_df.index,\n    \"count\": notype_task_df.values\n})\n\nsns.barplot(notype_task_df, x=\"index\", y=\"count\")","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:20:03.625263Z","iopub.execute_input":"2023-04-28T06:20:03.626378Z","iopub.status.idle":"2023-04-28T06:20:03.834686Z","shell.execute_reply.started":"2023-04-28T06:20:03.626323Z","shell.execute_reply":"2023-04-28T06:20:03.833424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the data distribution of the train_notype_df dataframe's Task column, we found out that some of the data was marked as False, as it was barely counted more than the other data being marked as True. Besides, there are almost 200000 entities for the data marked as False, while there are approximately 179600 data entities for the data marked as True. Additionally, the data that is marked as False in the train_notype_df dataframe's Task column implied to us that almost some of the events from the recordings of a specific FoG patient aren't annotated, unlike from what we saw from the data distribution of the train_defog_df dataframe's Task column.","metadata":{}},{"cell_type":"markdown","source":"Now that we finalized our data analysis in the train_notype_df dataframe, let's wrap up the chapter by proceeding into visualizing the data from the train_tdcsfog_df dataframe as we still use matplotlib and seaborn for plotting them out just to save our notebook RAM memory!","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">train_tdcsfog_df</h3>\n\nAfter we rummaged through visualizing the data from the unlabeled_df, train_defog_df, and train_notype_df dataframes, we came across the train_tdcsfog_df dataframe in which it stores the data gathered from the tdcsfog dataset we visualized previously. And with the data columns of the train_tdcsfog_df being similar to the train_defog_df dataframe, let's begin our visualization inside this section!","metadata":{}},{"cell_type":"markdown","source":"First of all, let's distribute and then visualize the AccV, AccAP, and AccML columns into three histograms into one subplot chart! To get started on plotting this out, we characterize the fig and axes variable into configuring our subplot chart with the plt module's subplots function, setting 1 and 3 inside as the rows and columns of our subplot, as well as arranging the figsize parameter to 15 by 5 for adjusting the width and height of our subplot, and the sharey parameter to True for enabling the y-axes sharing to the graphs inside.\n\nWe then proceed to create our three histogram graphs with the sns module's histplot function, setting the train_tdcsfog_df dataframe as the graph's data, followed by arranging the x parameter to the AccV, AccAP, and AccML columns individually for specifying the graph's x-axes, the kde parameter to True for enabling kernel-density estimation plots inside the graph, and the ax parameter to the axes variable's slice index of 0, 1, and 2 separately for placing our graphs into three different subplots.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(15, 5), sharey=True)\n\nsns.histplot(train_tdcsfog_df, x=\"AccV\", kde=True, ax=axes[0])\nsns.histplot(train_tdcsfog_df, x=\"AccAP\", kde=True, ax=axes[1])\nsns.histplot(train_tdcsfog_df, x=\"AccML\", kde=True, ax=axes[2])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:20:03.836436Z","iopub.execute_input":"2023-04-28T06:20:03.836812Z","iopub.status.idle":"2023-04-28T06:20:04.782487Z","shell.execute_reply.started":"2023-04-28T06:20:03.836777Z","shell.execute_reply":"2023-04-28T06:20:04.781563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the AccV column, we glimpsed a nearly symmetrical distribution with no left or right skews traced by the kde line, though the data peak was formed on the near right of the histogram. In other words, the range from approximately -9.2 to -8.8 has the highest data, with around 545 entities, while there are least-counted data in some ranges outside from the data peak since it has one data entity. Specifically, the symmetrical data on the right we saw from the AccV column hinted us that there are most events in which a specific FoG patient walked slowly in the opposite directions when they freeze while they had their knees trembling, hence observing the negative values we saw in the x-axes of the AccV column.\n\nOn the AccAP column, we clearly see the right-skew distribution in the middle of the diagram, as the kde line we plotted displayed a trace of how the bins in the right of the data peak ascending sequentially into the left of the histogram. Not only that, the highest data in the AccAP column is in the range roughly between -0.3 and 0 with 328 estimated entities, while the lowest data is in some ranges outside the data peak on the left and right, with only one entity. Furthermore, the right skewed distribution we noticed from the AccAP column implied to us that there are some events that showed how there's decreased acceleration and slower gait turns made by a FoG person, as it clearly showed the signs of slow movement symptoms in Parkinson's.\n\nFinally for the AccML column data visualization, we espied a left-skew distribution in the middle of the histogram plot that was mapped by the kernel-density estimation line when we noticed on how the bins' height on the middle-left increased consecutively into the data peak on the right. Nevertheless, the highest data we saw in the AccML column is around 0.7 to 0.9 with almost 495 entities, while once again the lowest data is in some ranges outside the left-skewed peak because of a single entity. In summary, the left-skewed data we noticed from the train_tdcsfog_df dataframe's AccML column made us explain that there are some events of how a specific FoG patient experienced bradykinesia, as they had slow-moving strides and gait velocities.","metadata":{}},{"cell_type":"markdown","source":"Lastly for our data visualization in the train_tdcsfog_df dataframe, let's visualize the StartHesitation, Turn, and Walking columns into three histograms in a subplot! We do the same thing from what we did for graphing the AccV, AccAP, and AccML columns, but we use the train_tdcsfog_df dataframe's StartHesitation, Turn, and Walking columns separately when creating each histogram in a subplot.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(15, 5), sharey=True)\n\nsns.histplot(train_tdcsfog_df, x=\"StartHesitation\", kde=True, ax=axes[0])\nsns.histplot(train_tdcsfog_df, x=\"Turn\", kde=True, ax=axes[1])\nsns.histplot(train_tdcsfog_df, x=\"Walking\", kde=True, ax=axes[2])","metadata":{"execution":{"iopub.status.busy":"2023-04-28T06:20:04.784128Z","iopub.execute_input":"2023-04-28T06:20:04.785281Z","iopub.status.idle":"2023-04-28T06:20:05.328630Z","shell.execute_reply.started":"2023-04-28T06:20:04.785231Z","shell.execute_reply":"2023-04-28T06:20:05.327367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How strange. Unlike what we saw from visualizing the train_defog_df dataframe's Turn, StartHesitation, and Walking columns, we espied that all three columns we plotted from the train_tdcsfog_df dataframe showed all of the zeros, with around 3800 data entities each. Additionally, the zeros we founded from the train_tdcsfog_df dataframe's StartHesitation, Turn, and Walking columns hinted us that all of the events recorded on this dataframe showed how the patient that has FoG didn't turn, hesitate, or even walk.","metadata":{}},{"cell_type":"markdown","source":"As we finalize our short two-graphed visualization in the train_tdcsfog_df dataframe, we completed the long section of visualizating the four dataframes from the train and unlabeled data folders as well as completing our whole data analysis in the Parkinson's FoG prediction competition!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #f78102; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion</h2>\n\nSo what are the key takeaways from the visualizing the data i**n** the P**a**rkinson's Freezing of Gait competitio**n**? For our visualization of the defog_df, tdcsfog_df, and daily_df dataframes**,** we learned that most patients has more follow-up assessments than post-treatment assessments while reflecting that there are some least and most challenging tests for each FoG patient. And from the subjects_df, tasks_df, and events_df dataframes, we recalled that most patients are male that were mostly aged 60 to 70, sometimes known for taking Parkinson's medication, and mostly showed signs of FoG symptoms as well as visualizing each task type and the duration of it. Lastly, we reflected on how most events from one of the Parkinson's FoG patients showed bradykinesia and slower gait turns and acceleration, along with seeing no event occurences in their hesitation, walking, or turning from the unlabeled_df, train_defog_df, train_notype_df, and train_tdcsfog_df dataframes. Aside from that, we will foresee our advancement in the evaluatio**n**s and tre**a**tments of the **n**efarious FoG**,** alongside with improving the lives of the people who ha**d** the devastating Parkins**on'**s disease symp**t**oms. Not only that, we wil**l** prop**e**l forw**a**rd into en**v**isaging our n**e**xt generations of breaking through the problems ahead of our **m**edical and h**e**alth researches and findings in the near future**!**","metadata":{}}]}