{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color: #4285F4;\"> Google</span> <span style=\"color: #DB4437\"> Isolated</span><span style=\"color: #F4B400\"> Sign Language</span><span style=\"color: #0F9D58;\"> Recognition</span></b></h1>\n\n<center><img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/46105/logos/header.png\" width=900></center>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #4285F4; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction</h2>\n\nAround the 18th century in 1760, Abbé de l'Épée, who is a philanthropic educator in France, stumbled across the two sisters communicating in signs, which they chatted together by showing hand gestures to each other. As a result, he became aware that they represented a signing community of 200 deaf Parisians, making this as the Old French Sign Language. 57 years later, Thomas Gallaudet compromised the Old French Sign Language, home sign systems, and various village sign languages, making this emerged as a language in the ASD (aka American School for the Deaf) as it is known as the standard American Sign Language (or ASL for short).\n\n<center>\n    <img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/3/3c/American_School_for_the_Deaf%2C_main_building%2C_August_10%2C_2008.jpg/1920px-American_School_for_the_Deaf%2C_main_building%2C_August_10%2C_2008.jpg\" width=500>\n    <figcaption style=\"color: gray;\">The main building of the American School of the Deaf in West side of Hartford, Connecticut. Credit: Wikimedia Commons</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"However, learning ASL language is too complicated to learn for the parents that have a child that is deaf because around thirty-three babies are born with permanent hearing loss in the United States daily, as ninety percent of them are birthed with their hearing parents who don't fully understand the American Sign Language. But on the bright side, interactive games, even with interpreting live video, can guide the hearing parents how to learn about the ASL language. And as we landed into this competition [after trekking through our data analysis with Jo Wilder](https://www.kaggle.com/code/dinowun/eda-simplified-jo-wilder-student-performance), our main focus is to classify isolated ASL signs by amplifying PopSign's educational games based on this complicated sign language, since it is a smartphone game app that makes learning American Sign Language fun, interactive, and accessible. Thus, with our presence of envisaging the data in this competition, let's go off for our session of EDA!\n\n<center>\n    <img src=\"https://play-lh.googleusercontent.com/jbIoAiVbHD8VaHhUxWbTM81SjSxpDtPL5s8WeIWKg2gLu-n72LQS_roKRMqWS7rkM9gN=w5120-h2880-rw\" width=500>\n    <figcaption style=\"color: gray;\">A snapshot of the PopSign gameplay. Credit: Google Play Store/Cat Sign Developer.</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"Prerequisite: Before we begin our analysis, we recommend you to read the Jo Wilder Student Performance notebook here:\nhttps://www.kaggle.com/code/dinowun/eda-simplified-jo-wilder-student-performance","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #DB4437; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Module Imports and Data Setup</h2>\n\nTo set our motion into initiating the data examination based on sign, lets haul the obvious pandas module as pd and numpy module as np for intergating data analysis and linear algebra. Afterwards, we haul off the modules for plotting, which is the plotly module's express attribute as px along with installing and importing pandas_bokeh module and outputting the graphs with the output_notebook function plugged to the pandas_bokeh module.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\n!pip install pandas_bokeh\nimport plotly.express as px\nimport pandas_bokeh\n\npandas_bokeh.output_notebook()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-12T07:11:19.041517Z","iopub.execute_input":"2023-03-12T07:11:19.042053Z","iopub.status.idle":"2023-03-12T07:11:37.647418Z","shell.execute_reply.started":"2023-03-12T07:11:19.042007Z","shell.execute_reply":"2023-03-12T07:11:37.645972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we brought in our principal modules for our data analysis, we understood that there is a csv file, a json file (sign_to_prediction_index_map.json), and lots of parquet files in the competition data. Nonetheless, we create two dataframes, train_df and landmark_file_df, to read out the train.csv file and any parquet under the train_landmark_files as well as any participant_id folder separately with the read_csv, read_json, and read_parquet functions from the pd module. We then place the head function to the two dataframes that we created for displaying the first five rows of them.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/asl-signs/train.csv\")\nlandmark_file_df = pd.read_parquet(\"/kaggle/input/asl-signs/train_landmark_files/16069/1007343357.parquet\")","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:37.650426Z","iopub.execute_input":"2023-03-12T07:11:37.650899Z","iopub.status.idle":"2023-03-12T07:11:38.049924Z","shell.execute_reply.started":"2023-03-12T07:11:37.650852Z","shell.execute_reply":"2023-03-12T07:11:38.048692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.053621Z","iopub.execute_input":"2023-03-12T07:11:38.054217Z","iopub.status.idle":"2023-03-12T07:11:38.085099Z","shell.execute_reply.started":"2023-03-12T07:11:38.054157Z","shell.execute_reply":"2023-03-12T07:11:38.083744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.088763Z","iopub.execute_input":"2023-03-12T07:11:38.089197Z","iopub.status.idle":"2023-03-12T07:11:38.107200Z","shell.execute_reply.started":"2023-03-12T07:11:38.089158Z","shell.execute_reply":"2023-03-12T07:11:38.105702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #DB4437; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Basic Dataframe Analysis</h3>","metadata":{}},{"cell_type":"markdown","source":"Since we generated the two dataframes, let's take the fundemental analysis of them! First of all, let's find out the number of data in each dataframe with the len function.","metadata":{}},{"cell_type":"code","source":"len(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.109463Z","iopub.execute_input":"2023-03-12T07:11:38.110474Z","iopub.status.idle":"2023-03-12T07:11:38.119352Z","shell.execute_reply.started":"2023-03-12T07:11:38.110383Z","shell.execute_reply":"2023-03-12T07:11:38.118265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(landmark_file_df)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.121077Z","iopub.execute_input":"2023-03-12T07:11:38.121451Z","iopub.status.idle":"2023-03-12T07:11:38.134235Z","shell.execute_reply.started":"2023-03-12T07:11:38.121411Z","shell.execute_reply":"2023-03-12T07:11:38.133138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the train_df dataframe, we counted 94417 data entities, while in the landmark_file_df dataframe, we listed a total 8145 data units in all. Specifically, the number of data in the train_df dataframe is huge, meaning that there's a lot of hand signing to be gathered to classify what each ASL gesture mean.","metadata":{}},{"cell_type":"markdown","source":"Now let's find and calculate how many missing values are present in the train_df and the landmark_file_df dataframes! We plug the isna function to the two dataframes specified for finding the present missing values and then place the sum function to sum all the present missing values.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.136343Z","iopub.execute_input":"2023-03-12T07:11:38.137411Z","iopub.status.idle":"2023-03-12T07:11:38.161815Z","shell.execute_reply.started":"2023-03-12T07:11:38.137350Z","shell.execute_reply":"2023-03-12T07:11:38.160707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.163425Z","iopub.execute_input":"2023-03-12T07:11:38.164104Z","iopub.status.idle":"2023-03-12T07:11:38.176218Z","shell.execute_reply.started":"2023-03-12T07:11:38.164064Z","shell.execute_reply":"2023-03-12T07:11:38.175234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the train_df dataframe, we found no missing values that is present in the data, but however for the landmark_file_df dataframe, we counted 441 missing values in the x, y, and z columns. Nevertheless, the 441 NaN values in the three columns of x, y, and z in the landmark_file_df dataframe showed us that there are invalid or unclear data over the hand gesture movements that replicate the sign language.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's compute the entire number of columns in each two dataframes! All we must do is to attach the shape attribute to the train_df and the landmark_file_df dataframes to find the shape of each dataframe, then we pull out the first last segment of the dataframe's shape by applying the slice index of 1.","metadata":{}},{"cell_type":"code","source":"train_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.177576Z","iopub.execute_input":"2023-03-12T07:11:38.178149Z","iopub.status.idle":"2023-03-12T07:11:38.186005Z","shell.execute_reply.started":"2023-03-12T07:11:38.178115Z","shell.execute_reply":"2023-03-12T07:11:38.184708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.189766Z","iopub.execute_input":"2023-03-12T07:11:38.190619Z","iopub.status.idle":"2023-03-12T07:11:38.203458Z","shell.execute_reply.started":"2023-03-12T07:11:38.190579Z","shell.execute_reply":"2023-03-12T07:11:38.201891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the two dataframe's last shape index we obtained, there are 4 columns in the train_df dataframe while there's 7 columns in the landmark_file_df dataframe. In addition, the fewer number of columns we counted in the train_df dataframe made us understand that the less number of columns displayed the background information based on the meaning of the sign language pose, participant, or sequence.","metadata":{}},{"cell_type":"markdown","source":"Now that we have our two dataframes visualized by the basics, let's propel forward into our first quick data analysis of the train_df dataframe with hints of interactive plotting by Plotly!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #F4B400; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Visualizing the train_df Dataframe</h2>\n\nAs we now enter the passage of visualizing the train_df dataframe we created, Let's make a short summary about each column we found in the train_df dataframe!\n* **path**: Path directory specified to the parquet files.\n* **participant_id**: ID specified for each participant that does the sign language.\n* **sequence_id**: ID specified for each sign language sequence.\n* **sign**: Specfies what is the meaning of each hand gesture.\n\nWith all of each column explained concisely, let's visualize the participant_id data column into Plotly's histogram and box-plot all at once! To get started, we characterize the one-and-only fig variable to create our histogram graph with the px module's histogram function, setting the train_df dataframe as the data used for the histogram, along with the x parameter to the participant_id data column for loading the x-axes with the specified column, and the marginal parameter to box so that we are going to plot the box-plot on the top of the histogram graph. Lastly, as you all know the process once we finished creating our graph, we display it with the show function that is plugged to the fig variable figure.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"participant_id\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:38.205305Z","iopub.execute_input":"2023-03-12T07:11:38.205804Z","iopub.status.idle":"2023-03-12T07:11:40.124030Z","shell.execute_reply.started":"2023-03-12T07:11:38.205759Z","shell.execute_reply":"2023-03-12T07:11:40.122368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the participant_id graph we created, we found out that most of the bins inside the histogram isolated themselves, as some were near to the surrounding bin, while others were farther from other bins. Not only that, 49445 was the most counted data in the participant_id column with 4968 data entities listed, while 30680 was the least counted since it has 3338 entities. On the boxplot shown above, we noticed that the box portion in the graph showed an asymmetrical display as well as the distance from the lower fence to the first quartile and the distance from the third quartile to the upper fence. Funnily enough, the first and third quartiles of the participant_id data is 25571 and 49445, the median is 32319, and the interquartile range is 23874. In addition, the isolated data counts big and small in the participant_id that was recorded on the histogram we created hinted us that each participant that was recording the hand gestures indicate an identifier for making each hand gesture that was needed for the computer model to classify what each sign language showed.","metadata":{}},{"cell_type":"markdown","source":"Now let's distribute the sequence_id data column into another histogram and boxplot combined into one graph with help from Plotly! Again, we characterize the fig variable figure to create another histogram and boxplot fusion graph with the histogram provided by the px module, loading the train_df as the data for plotting the histogram-boxplot graph, plus arranging the x parameter to the sequence_id column for plotting the specified data column into the graph's x-axes and the marginal parameter to box per stated that we're going to plot the participant_id data into the histogram as well as plotting the box-plot above the histogram. With that completed, we exhibit our fig variable figure by plugging the show function into it. ","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"sequence_id\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:11:40.126266Z","iopub.execute_input":"2023-03-12T07:11:40.126744Z","iopub.status.idle":"2023-03-12T07:11:40.271268Z","shell.execute_reply.started":"2023-03-12T07:11:40.126694Z","shell.execute_reply":"2023-03-12T07:11:40.269627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our distribution on the sequence_id data column to our histogram and box-plot combined, we noted that the bins on this histogram displayed a uniform distribution, meaning that each number of distributions of the sequence_id data were estimately the same towards other distributions. But aside from that, the data range from 2250000000 to 2300000000 was counted the most since it has 1206 data entities while the range between 4250000000 and 4300000000 was counted the least because it has 987 data entities. And for the boxplot we saw above, we glimpsed that the box portion and the distance from it to the upper or lower fences nearly exhibited a symmetrical display. Besides, the first and third quartiles of the sequence_id data specified by the box plot we created is 1078076000 and 3218845000, the median is 2154240000, and the interquartile range is 2140769000. As regards to the data based on the sequence_id column, the uniform distribution of the data made us explain that there are variate number of sequence_id data for indicating an identifier for resembling a landmark sequence, so that the computer model will identify which hand gesture really signify in plain English.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's assemble a wordcloud consisting on the sign column from the train_df dataframe, but with style by stylecloud! Before we begin this process, we install the stylecloud module via pip and then import it followed by the Image module from the IPython module with the display attribute.","metadata":{}},{"cell_type":"code","source":"!pip install stylecloud\n\nimport stylecloud\nfrom IPython.display import Image","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-12T07:11:40.272873Z","iopub.execute_input":"2023-03-12T07:11:40.273288Z","iopub.status.idle":"2023-03-12T07:12:00.276047Z","shell.execute_reply.started":"2023-03-12T07:11:40.273247Z","shell.execute_reply":"2023-03-12T07:12:00.274567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we import the urgent modules for creating our styled wordcloud, we characterize the concat_signs variable to just an empty string followed by joining every translation of each sign gesture with the join function, containing a tuple in which it held the i variable that was looping over each data in the train_df dataframe's sign data column attribute and convert to str with the astype function that contained the str type.","metadata":{}},{"cell_type":"code","source":"concat_signs = ' '.join([i for i in train_df.sign.astype(str)])","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:00.278076Z","iopub.execute_input":"2023-03-12T07:12:00.278592Z","iopub.status.idle":"2023-03-12T07:12:00.304511Z","shell.execute_reply.started":"2023-03-12T07:12:00.278541Z","shell.execute_reply":"2023-03-12T07:12:00.302761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And now's the time to plot out the cloud of words consisting on what it means for each for each hand gesture! We use the gen_stylecloud function from the stylecloud attribute for creating our styled wordcloud, configuring the text parameter to the concat_signs for loading the text data, the icon parameter to \"fas fa-hand\" or any other font awesome icon for shaping the words into the shape of a icon, the palatte parameter to the colorbrewer module with the diverging attribute followed by the Spectral_11 attribute that was encased in strings or other palette format in strings for coloring the shaped wordcloud, the background_color parameter to black or any other color for specifying the background color of our background of the styled wordcloud, the gradient parameter to vertical or horizontal for specifying which gradient runs from, and the size parameter to 1024 or to any value for sizing the width and height of our styled wordcloud image.","metadata":{}},{"cell_type":"code","source":"stylecloud.gen_stylecloud(text=concat_signs,\n                          icon_name='fas fa-hand-paper',\n                          palette='colorbrewer.diverging.Spectral_11',\n                          background_color='black',\n                          gradient='vertical',\n                          size=1024)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:00.306589Z","iopub.execute_input":"2023-03-12T07:12:00.307997Z","iopub.status.idle":"2023-03-12T07:12:05.247759Z","shell.execute_reply.started":"2023-03-12T07:12:00.307945Z","shell.execute_reply":"2023-03-12T07:12:05.246258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we created our wordcloud, the gen_stylecloud function saves as the image file known as \"stylecloud.png\" in the output files of our notebook. In other words, we use the Image function to display the specified image file, setting the filename parameter to \"./stylecloud.png\" for indicating that we're going to view the image of a styled wordcloud in the output files alongside with the width and height parameters to 1024 for specifying the width and height of displaying an image.","metadata":{}},{"cell_type":"code","source":"Image(filename=\"./stylecloud.png\", width=1024, height=1024)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:05.249944Z","iopub.execute_input":"2023-03-12T07:12:05.250404Z","iopub.status.idle":"2023-03-12T07:12:05.266300Z","shell.execute_reply.started":"2023-03-12T07:12:05.250357Z","shell.execute_reply":"2023-03-12T07:12:05.264825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the wordcloud we generated, we created the shape of a hand out of the data we gathered from the train_df dataframe's sign column, which is seen in the vertical rainbow gradient. In other words, \"listen\", \"look\", and \"shhh\" are the common words as it was shown in big size, while small-sized words in our styled wordcloud like \"now\", \"dad\", \"ride\", or \"give\" are the uncommon words in the sign column from the train_df dataframe. By way of explaination, the data full of words that were extracted from the sign column made us understood that these are the words that were used for the computer model to translate each hand gesture, as some words were repeated more or less for making the compute model translate each meaning of each hand gesture clearly while training.","metadata":{}},{"cell_type":"markdown","source":"And with that, we completed our short section of visualizing the train_df dataframe! As for what's next in our ASL data analysis journey is we're going to envisage the data from each sign language hand gesture specified in an individual parquet file from the competition data.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #0F9D58; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Envisioning the Parquet data of Each Sign Language</h2>\n\nWhen we head off to another section, we learned that there are the parquet files stored in each folder consisting over the landmark files in each folder. In other words, we created the landmark_file_df dataframe from reading one of the parquet files that was held in the train_landmark_files folder. And with that, let's detail the columns that were stored in each parquet file!\n* **frame**: The frame number specified in each raw video.\n* **row_id**: The ID specified for each row.\n* **type**: The type specified for each landmark.\n* **landmark_index**: The number specified for each landmark index.\n* **x/y/z**: The normalized spatial coords of the landmark.\n\nAnd with the columns in the landmark_file_df dataframe detailed concisely, let's march onwards to the threshold of this section for visualizing deep into one of the parquet data that was specified by the landmark_file_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"In the first place, let's distribute and visualize the frame data into the histogram created with Bokeh! We use the landmark_file_df dataframe's frame data column to plot a Bokeh plot with the plot_bokeh function, setting the kind parameter to hist for creating a histogram.","metadata":{}},{"cell_type":"code","source":"landmark_file_df[\"frame\"].plot_bokeh(\n    kind=\"hist\",\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:05.267810Z","iopub.execute_input":"2023-03-12T07:12:05.268837Z","iopub.status.idle":"2023-03-12T07:12:05.385193Z","shell.execute_reply.started":"2023-03-12T07:12:05.268783Z","shell.execute_reply":"2023-03-12T07:12:05.383437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the histogram we compiled out of Bokeh, we spotted that there are 5 peaks of data seen in the left and right of the histogram, as the bins displayed a mulitmodal distribution. In other words, the ranges 24-25.4, 26.8-28.2, 31-32.4, 33.8-35.2, and 36.6-38 were counted the most as each range has 1086 data entities, while the other ranges were counted the least since each of them has 543 data entities. Specifically, the mulimodal display in the landmark_file_df dataframe's frame data column varies and depends in each parquet file, as some ranges in each frame column show more or less multimodal bins than the one we plotted.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the type data column into a pie chart from Bokeh! Prior to plotting this data out, we inherit another dataframe, type_df, to counting the values in the landmark_file_df dataframe's type data column with the value_counts function. After we created the type_df dataframe, we use the plot_bokeh function to the type_df dataframe for creating our diagram with Bokeh, setting the kind parameter to pie for creating a pie chart and the x parameter to the type_df dataframe's indexes specified by the index attribute for setting the names of the pie chart.","metadata":{}},{"cell_type":"code","source":"type_df = landmark_file_df[\"type\"].value_counts()\n\ntype_df.plot_bokeh(\n    kind=\"pie\",\n    x=type_df.index,\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:05.387106Z","iopub.execute_input":"2023-03-12T07:12:05.387630Z","iopub.status.idle":"2023-03-12T07:12:05.497940Z","shell.execute_reply.started":"2023-03-12T07:12:05.387589Z","shell.execute_reply":"2023-03-12T07:12:05.496164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the pie chart we generated out of the type data column, we realized that most of the data in this column were labeled as \"face\" as there are 7020 data entities of them, while right_hand and left_hand were counted the least, as both of them has 315 data entities listed. Specifically, the number of the data labeled as \"face\" in the type data column from the landmark_file_df dataframe allow us explain that most participants recording down the ASL hand gesture showed their faces in most frames in the video, as this may occur in other parquet files that were stored in the landmark_files data folder.","metadata":{}},{"cell_type":"markdown","source":"Now let's distribute and visualize the data from the x, y, and z columns separately in Bokeh's histogram! We pull out the landmark_file_df dataframe's x, y, and z data columns separately to plot a typical Bokeh plot by plugging the plot_bokeh function in each of them, and simply set the kind parameter to hist for making our graph type to a histogram.","metadata":{}},{"cell_type":"code","source":"landmark_file_df[\"x\"].plot_bokeh(\n    kind=\"hist\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:05.499503Z","iopub.execute_input":"2023-03-12T07:12:05.500034Z","iopub.status.idle":"2023-03-12T07:12:05.624018Z","shell.execute_reply.started":"2023-03-12T07:12:05.499979Z","shell.execute_reply":"2023-03-12T07:12:05.622077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file_df[\"y\"].plot_bokeh(\n    kind=\"hist\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:05.626322Z","iopub.execute_input":"2023-03-12T07:12:05.626853Z","iopub.status.idle":"2023-03-12T07:12:05.762546Z","shell.execute_reply.started":"2023-03-12T07:12:05.626809Z","shell.execute_reply":"2023-03-12T07:12:05.761011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file_df[\"z\"].plot_bokeh(\n    kind=\"hist\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:05.764514Z","iopub.execute_input":"2023-03-12T07:12:05.766068Z","iopub.status.idle":"2023-03-12T07:12:05.910160Z","shell.execute_reply.started":"2023-03-12T07:12:05.766003Z","shell.execute_reply":"2023-03-12T07:12:05.908741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Despite the warnings emitted by Bokeh over plotting the nan values in the x, y, z columns from the landmark_file_df dataframe, we visualized the right-skew display on the x and y data that was plotted on the first two histograms, but for the histogram based on the z column distribution, we noted that the peak of data was generated on the near right of the graph, indicating this distribution display as a left-skew. In other words, the highest data counted is from 0.308 to 0.425 in the x column with 3377 data entities listed, 0.239 to 0.467 in the y column containing 6388 data entities, and -0.306 to 0.113 in the z column that held 7159 data entities, while the lowest data counted is between 1.013 and 1.131 for the x column with 9 data values counted, 0.922 to 1.149 with 15 data values in the y column, and for the z column, -1.979 to -1.56 with 4 data values. Specifically, the data distributions based on the x, y, and z columns varies, as they can be left or right skewed, bimodal or multimodal, or possible uniform or symmetrical, as their distribution differs on each parquet file specified in the landmark_files folder.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's graph out the x, y, and z data separately with bokeh's scatter plot! We use the landmark_file_df dataframe to plug in the plot_bokeh attribute for creating a bokeh plot followed by applying the scatter function to configure a scatter plot, setting the x parameter to the standalone x or y columns for specifying the x-axes of the scatter graph, the y parameter to the y or z columns separately for specifying the y-axes of the scatter graph.","metadata":{}},{"cell_type":"code","source":"landmark_file_df.plot_bokeh.scatter(\n    x=\"x\",\n    y=\"y\",\n    category=\"type\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:05.911832Z","iopub.execute_input":"2023-03-12T07:12:05.912274Z","iopub.status.idle":"2023-03-12T07:12:06.191979Z","shell.execute_reply.started":"2023-03-12T07:12:05.912234Z","shell.execute_reply":"2023-03-12T07:12:06.190413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file_df.plot_bokeh.scatter(\n    x=\"x\",\n    y=\"z\",\n    category=\"type\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:06.194132Z","iopub.execute_input":"2023-03-12T07:12:06.194710Z","iopub.status.idle":"2023-03-12T07:12:06.485501Z","shell.execute_reply.started":"2023-03-12T07:12:06.194652Z","shell.execute_reply":"2023-03-12T07:12:06.484433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file_df.plot_bokeh.scatter(\n    x=\"y\",\n    y=\"z\",\n    category=\"type\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:06.487186Z","iopub.execute_input":"2023-03-12T07:12:06.488368Z","iopub.status.idle":"2023-03-12T07:12:06.817388Z","shell.execute_reply.started":"2023-03-12T07:12:06.488316Z","shell.execute_reply":"2023-03-12T07:12:06.816222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From within the three data angles plotted on a scatter plot, we got a faint visualization of how the x, y, and z data columns resembled a person making a sign language gesture, although we cannot visualize how the hand gesture change over time in each frame. In other words, the data plotted in the x, y, and z columns varies as each parquet file outlined other sign language gesture.","metadata":{}},{"cell_type":"markdown","source":"And after we visualized the x, y, and z data columns in the landmark_file_df dataframe, we were skeptical about how each parquet file outlined the media consisting on each sign language gesture based on the type and frame columns we visualized through. So, we dived into another, small section of this chapter for visualizing the sign language gesture clearly by drawing the hand connections with help from mediapipe.","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #0F9D58; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Depicting the Sign Language Movements with Mediapipe</h3>\n\nAs we previously plotted down the three scatter plots from the landmark_file_df dataframe's x, y, and z data columns, we clearly didn't understand about how the hands, face, and posing resembled a specific hand gesture. However, the mediapipe module can help us envisage a specific sign language, meaning that you can observe a specific movement with people or hand movements.","metadata":{}},{"cell_type":"markdown","source":"Without further ado, we install the mediapipe module with pip and then import it as mp along with importing the matplotlib module with the pyplot attribute as plt. Thus, we characterize the hand_movement variable to the mp module's solutions attribute and then the hands attribute for enabling hand and finger tracking solution.","metadata":{}},{"cell_type":"code","source":"!pip3 install mediapipe\nimport mediapipe as mp\nimport matplotlib.pyplot as plt\n\nhand_movement = mp.solutions.hands","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-12T07:12:06.819175Z","iopub.execute_input":"2023-03-12T07:12:06.820207Z","iopub.status.idle":"2023-03-12T07:12:22.090389Z","shell.execute_reply.started":"2023-03-12T07:12:06.820162Z","shell.execute_reply":"2023-03-12T07:12:22.088624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we installed and configured the mediapipe module as well as importing the matplotlib module, we then create another data column in the landmark_file_df dataframe, y_, to multiply the landmark_file_df dataframe's y column with -1 thus defining the fig and ax variables to the plt module's subplots function for creating our graph, setting the figsize parameter to 6 by 6 for adjusting the size of our graph. Thenceforth, we create a for loop statement, looping the defined gesture variable to a list containing the left_hand name from the data labels in the type column because we didn't see any data labeled as right_hand in the landmark_file_df dataframe's type column.\n\nInside the for loop we created, we characterize the frame variable to the landmark_file_df dataframe and search the specific data in it with the query function, setting a string that contained how the frame data column is equal to 25 and the type is equal to the frame variable. We now get to plot down the scatter plot with the scatter function that is plugged to the ax variable, setting the x and y axes with the landmark_file_df dataframe's x and y_ columns. Following from that, we create another for loop inside the for loop we're in, defining the line_connect variable and loop it to the hand_movement variable that has the HAND_CONNECTIONS attribute. While we're inside in another for loop we generated inside a for loop, we characterize the a and b variables to the first and second indexes indicated as 0 and 1, and then we define the x1, y1, x2, and y2 variables separately in pairs to search the specific object in the frame variable with the query function, setting it to a string that showed the landmark_index data column equal to a and b variables as well as having an index containing the list that held the x and y_ columns from the landmark_file_df dataframe along with getting out the values of it with the values attribute and extract the first value of it with the slice index of 0. Finally while inside this for loop, we plot down the line plot with the plt module's plot function, setting the list containing x1 and x2 as the graph's x-axis, the y1 and y2 as the graph's y-axis, and the color parameter to green for coloring the lines.","metadata":{}},{"cell_type":"code","source":"landmark_file_df[\"y_\"] = landmark_file_df[\"y\"] * -1\nfig, ax = plt.subplots(figsize=(6, 6))\n\nfor gesture in [\"left_hand\"]:\n    frame = landmark_file_df.query(\"frame == 25 and type == @gesture\")\n    ax.scatter(frame[\"x\"], frame[\"y_\"])\n    \n    for line_connect in hand_movement.HAND_CONNECTIONS:\n        a = line_connect[0]\n        b = line_connect[1]\n        \n        x1, y1 = frame.query(\"landmark_index == @a\")[[\"x\", \"y_\"]].values[0]\n        x2, y2 = frame.query(\"landmark_index == @b\")[[\"x\", \"y_\"]].values[0]\n        \n        plt.plot([x1, x2], [y1, y2], color=\"green\")","metadata":{"execution":{"iopub.status.busy":"2023-03-12T07:12:22.092284Z","iopub.execute_input":"2023-03-12T07:12:22.092694Z","iopub.status.idle":"2023-03-12T07:12:22.599767Z","shell.execute_reply.started":"2023-03-12T07:12:22.092638Z","shell.execute_reply":"2023-03-12T07:12:22.597899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the scatter and line graph plotted with mediapipe, we clearly glimpsed of how this diagram displayed one of the gestures that formed a specific sign language. Other than that, it also displayed an outline of a hand that was mapped out of dots and lines, and the segment of the asl gesture resembled the word \"boat\". Additionally, the plot that plotted one of the hand gesture with mediapipe made us understand that most hand gestures vary in each parquet file, as each showed a specific word like chocolate, milk, or shhh.\n\n<center>\n    <img src=\"https://developers.google.com/static/mediapipe/images/solutions/hand-landmarks.png\" width=500>\n    <figcaption style=\"color: gray;\">A diagram consisting on the hand landmarks once outlined by mediapipe. Credit: Google Developers</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"And with help from mediapipe, we completed visualizing the hand gestures in one of the parquet files in the landmark_files folder as well as finishing our whole, full chapter of tabular data analysis in the landmark_file_df dataframe with graphs!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #4285F4; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion</h2>\n\nSo what did we learned from this competition? We recalled of how we graphed out the data in the train_df dataframe like the IDs for the sequences and participants and the sign translations in wordclouds alongside with graphing out one of the parquet files that was specified in the landmark_file_df dataframe such as plotting down the x, y, and z data in a scatterplot, sorting and counting the type data, and last but not least, tracing and outlining the fragment of a hand gesture with the mediapipe module. But after we amplify the educational games of PopSign based on learning the sign language, we'll help the parents connect with deaf children and then break down language barriers, so that the communication between them will be diversely linked together ever after, even with those who speak another language other than English.","metadata":{}}]}