{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color: #d50a0a;\"> NFL</span> <span style=\"color: #013369\"> 1st and Future Player Contact Detection</span> </b></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d50a0a; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction</h2>\n\nIn each game of NFL's Gridiron football, they have one problem that is looming in every match from a team against each other: Safety. Football is a violent sport, in which players from both sides of two teams bashed their football helmets against each other, often tackling them, and even pinning them to the ground. And yet, the way how the NFL's football players tackle and pinning each other sparks a lots of injuries, some broke their bones or got head concussions. And as a result from this, NFL and the Amazon Web Services teamed up together to improve the prediction of the player injuries with the machine learning models created by data scientists and machine learners. And once the models created by them were accurately enough to predict player injury from each football match, then the injuries inflicted to players will be obivous, thus increasing safetiness to each player playing the game. Fast-forward to this notebook you're in, we've done our EDA analysis on soccer moves from the DFL Bundesliga Shootout Data, and yet, let's embark our journey to the 1st and Future EDA on Player Contact Detection in Football!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #013369; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Imports + Dataframe Setup</h2>\n\nTo begin our EDA analysis, we load our pandas module as pd and the numpy module as np for data science and linear algebra mathematics, just like how an average python notebook starts their data analysis with these. Next, we import the plotting modules for plotting graphs from the data given in this notebook like the plotly module with the express submodule as px, the graph_objects submodule as go (from the plotly module again), and the matplotlib module with the pyplot submodule as plt for image and video visualization. Lastly, we install the moviepy module via pip and then import the ffmpeg_extract_subclip function from the module we installed with the video submodule followed by the io submodule and then the ffmpeg_tools submodule, followed by importing the Video function from the IPython module with the display submodule.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport plotly.express as px\nimport plotly.graph_objects as go\n\n!pip3 install moviepy\nfrom moviepy.video.io.ffmpeg_tools import ffmpeg_extract_subclip\nfrom IPython.display import Video","metadata":{"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we imported our urgent, must-use modules, we create four dataframes to read the training (indicated with \"train_\" prefix) csv files with the pd module's read_csv function, which is base_helm_df, labels_df, player_tracking_df, and video_metadata_df. Once finished, we display their first five rows with the head function.","metadata":{}},{"cell_type":"code","source":"base_helm_df = pd.read_csv('/kaggle/input/nfl-player-contact-detection/train_baseline_helmets.csv')\nlabels_df = pd.read_csv('/kaggle/input/nfl-player-contact-detection/train_labels.csv')\nplayer_tracking_df = pd.read_csv('/kaggle/input/nfl-player-contact-detection/train_player_tracking.csv')\nvideo_metadata_df = pd.read_csv('/kaggle/input/nfl-player-contact-detection/train_video_metadata.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_helm_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"player_tracking_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"video_metadata_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d50a0a; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Basic Dataframe Analysis</h2>\n\nFollowing from our creation from our four dataframe from the training csv files, let's visualize the surface basics to our dataframes! First, let's observe how many data were in each four dataframes by using the len function to each of the four dataframes.","metadata":{}},{"cell_type":"code","source":"len(base_helm_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(labels_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(player_tracking_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(video_metadata_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After the four separate cells was compiled, we've found out that there are 3783616 data inside the base_helm_df, 4721618 in labels_df dataframe, 1353053 in playing_tracking_df dataframe, and 480 in video_metadata_df dataframe. In other words, the huge number of data inside the first three dataframes shows us that there are a lot of games for the data scientists and machine learners to analyze, since there is each injury inflicted to each player going on the defense and offense.","metadata":{}},{"cell_type":"markdown","source":"Now, let's move onwards to finding the NaN values in each four dataframes! We simply use the isna function to find any NaN values in each data followed by the sum function to count all NaN values counted.","metadata":{}},{"cell_type":"code","source":"base_helm_df.isna().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df.isna().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"player_tracking_df.isna().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"video_metadata_df.isna().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, there are no NaN values throughout the four dataframes we've mentioned. That explains us that AWS and NFL carefully logged through each move a football player made in each game logged in the data.","metadata":{}},{"cell_type":"markdown","source":"Finally, for this section, let's visualize the number of columns in each of four dataframes! We find the shape of the four dataframes first with the shape attribute and then we display the last index of the four dataframes' shape with the slice index specified to 1.","metadata":{}},{"cell_type":"code","source":"base_helm_df.shape[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df.shape[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"player_tracking_df.shape[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"video_metadata_df.shape[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once the four code cells has successfully compiled each output separately, we found out that there are 12 columns in the base_helm_df dataframe, 7 columns in the labels_df dataframe, 17 columns in the player_tracking_df dataframe, and 7 columns in the video_metadata_df dataframe. Particularly, the number of columns from each dataframe shows us that the data on detecting player contact is organized through game footage from start to end time, player moves, tactics, position, speed, distance, and contact_id.","metadata":{}},{"cell_type":"markdown","source":"Now that we have our four dataframes assembled and visualized in a basic way, let's dive deep into data analysis on each dataframe at a time into each section of our EDA!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #013369; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Section 1: base_helm_df</h2>\n\nIn this section we're in, here's a little background information about the base_helm_df dataframe. It contains the data regarding of imperfect baseline predictions for helmet boxes and player assignments for the Sideline and Endzone video view and was created after the model was created from the winning solution from last year's competition and can be used to leverage their predictions. Without further ado, let's go on to analyzing this first dataframe!\n\nFirst, let's visualize the game_key column by distributing them into the histogram! We characterize the fig variable to the histogram function from the px module, setting the base_helm_df dataframe as our data for the histogram and the x parameter to the game_key column for configuring our histogram's x-axis. Finally, we exhibit our fig variable that contains the histogram graph with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(base_helm_df, x=\"game_key\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our histogram we've plotted, we visualized a lot of thin lines around the histogram, in which the highest data counted is 58204 in the game_key column, with 88533 data entities, while the least data counted is 58326 in the game_key column, with 9185 data entities. In other words, the data distribution over the game_key data is counted based on how long the game lasted, in which some lasted for short time while others lasted for long time.","metadata":{}},{"cell_type":"markdown","source":"Let's proceed towards visualizing the play_id column from the base_helm_df dataframe by plotting them into the histogram! We generate our variable, fig, to configuring our histogram graph with the px module's histogram function, setting the base_helm_df dataframe as the data for the histogram and the x parameter to the play_id column for configuring the x-axis in the histogram. Finally, let's show out our fig variable figure with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(base_helm_df, x=\"play_id\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just like the histogram graph consisting on the base_helm_df dataframe's game_key column, we see thin lines around the histogram, in which the 750 to 759 range is counted the most, with 89587 units of data counted, while the play_id range from 4110 to 4119 is counted the least, with 9185 entities of data calculated in all. Moreover, the data in the play_id column shows us that the variant distribution of them matches to the variant distributions of the game_key column data, as we noticed a lot of videos regardless of each video named as a pair with the game_key and game_id columns.","metadata":{}},{"cell_type":"markdown","source":"Let's find out the distribution count of the view column to a pie chart! Before that, we create another dataframe, view_point, to the base_helm_df dataframe's view data column counted by values with the value_counts function. Thenceforth, we characterize the fig variable to the pie function from the px module, setting the view_point dataframe as our data for the pie chart, the names parameter to the indexes of the view_point dataframe specified by the index attribute, and the values parameter to the values of the view_point dataframe listed by the values attribute. With all setup completed on our pie graph, we display out our fig variable figure with the show function.","metadata":{}},{"cell_type":"code","source":"view_point = base_helm_df[\"view\"].value_counts()\n\nfig = px.pie(view_point, names=view_point.index, values=view_point.values)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our pie chart we've generated, we found out that 55.6% of the data in the view column from the base_helm_df dataframe is \"Sideline\", while the other 44.2% of the data is Endzone and 0.172% of it is \"Endzone2\". Specifically, the \"Sideline\" data in the view column from the base_helm_df dataframe explains how the camera view can vary on each angle from the football stadium. Not only that, we realized that the NFL video labels in the competition data is named with the game_key, play_id, and view columns combined with an underscore (e.g: `58326_000750_Sideline.mp4`).","metadata":{}},{"cell_type":"markdown","source":"Once we finished our view data analysis, let's propel towards analyzing the frame data into our histogram! To proceed into that, we create our fig variable to the histogram configuration with the px module's histogram function, setting the base_helm_df dataframe as our data for the graph and the x parameter to the frame column as we configure it for the histogram's x-axis. With that completed, we use the fig variable to the show function for displaying our newly-created graph.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(base_helm_df, x=\"frame\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we waited for the code cell above to generate our histogram, we found out that there's a downward curve regarding about the frame data from the base_helm_df dataframe as we couldn't determine the most and least values around the diagram. Additionally, the associated frame distributed towards each data in the base_helm_df dataframe varies on each video's frame, whether the frame of them is short or long.","metadata":{}},{"cell_type":"markdown","source":"Let's create another histogram based on our distribution and analysis over the nfl_player_id data column! All we need to do is to simply create our fig variable and define it to the px module's histogram function just for creating our histogram, setting the base_helm_df dataframe as our data for the histogram and the x parameter to the nfl_player_id column for setting the x-axis for the histogram. With that completed, we make our histogram appear by using the show function to the fig variable graph.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(base_helm_df, x=\"nfl_player_id\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we compiled our histogram based on the base_helm_df dataframe's nfl_player_id column, we spotted a lot of spikes throughout the diagram, in which some were short and others were tall. Nevertheless, the most range of data counted in the histogram is in between 47800 and 47849, with 130,323 total entities listed, while the range between 31400 and 31449 is counted the least, in which it contained 515 entities. Additionally, the data based on the nfl_player_id column hints that there are lot of player's ID code has been predicted imperfectly, as there's missing or incomplete data on their moves and tactics.","metadata":{}},{"cell_type":"markdown","source":"Thereafter from our nfl_player_id column distribution, let's move on towards analyzing the player_label data and distribute them into our bar graph! But prior to graphing this bar chart out of Plotly, we characterize the dataframe, player_labels, to add up the values of the base_helm_df dataframe's player_label column with the value_counts function. Thenceforth, we create our fig variable figure to the bar chart configuration with the px module's bar chart, setting the player_labels dataframe as the bar chart's data, the x parameter along with the color parameter to the indexes of the player_labels dataframe specified by the index attribute for configuring the bar chart's x-axis and color legend, and the y parameter to the values of the player_labels dataframe specified by the values dataframe for configuring the bar chart's y-axis. Afterwards, we reveal our bar chart with the show function into the fig variable figure.","metadata":{}},{"cell_type":"code","source":"player_labels = base_helm_df[\"player_label\"].value_counts()\n\nfig = px.bar(player_labels, x=player_labels.index, color=player_labels.index, y=player_labels.values)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see from this bar chart, we found out that there's a lot of players in the player_label column and not only that, we spotted thin lines in our bar chart. In other words, \"H54\" is counted the most, as there are 39469 data entities regarding to this data, while \"V39\" is counted the least, with 4048 data entities listed in this player_label graph. Additionally, the player_labels data that is specified in the bar graph explains to us that most players that were involved in every NFL game is evaluated whether one of them made foul contact or not.","metadata":{}},{"cell_type":"markdown","source":"Lastly, let's visualize and distribute the left, top, width, and height columns into the histograms separately, but in our subplot! Before that, we import the make_subplots function from the plotly module with the subplots submodule and the go from the plotly module with the graph_objects submodule. After that, we create the fig variable to make our subplots with the make_subplots function, setting the rows parameter to 2 and the cols parameter to 2 for configuring the two rows and two columns of each subplot. Once we created the subplots, we add the traces to the fig variable subplots with the add_trace function, each containing a histogram configuration with the go module's Histogram function, with the x parameter set to the base_helm_df dataframe's left, top, width, and height columns for the x-axis configuration and the name parameter to the same names of the base_helm_df dataframe's columns specified for indicating the label of the graphs, setting the row and col parameters to 1, 2, 1, 2 and 1, 1, 2, 2 for placing our histograms to the subplots. Thenceforth, we show off the fig variable with the show function!","metadata":{}},{"cell_type":"code","source":"from plotly.subplots import make_subplots\nimport plotly.graph_objects as go\n\nfig = make_subplots(rows=2, cols=2)\n\nfig.add_trace(\n    go.Histogram(x=base_helm_df[\"left\"], name=\"left\"),\n    row=1, col=1\n)\n\nfig.add_trace(\n    go.Histogram(x=base_helm_df[\"top\"], name=\"top\"),\n    row=2, col=1\n)\n\nfig.add_trace(\n    go.Histogram(x=base_helm_df[\"width\"], name=\"width\"),\n    row=1, col=2\n)\n\nfig.add_trace(\n    go.Histogram(x=base_helm_df[\"height\"], name=\"height\"),\n    row=2, col=2\n)\n\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the four histograms we compiled into our subplot, we spotted a peak of data mostly in the middle of the top and left columns distributed in a histogram graph thus seeing the long bar of data in the first part of the histogram too while on the other hand, we spot the two peaks of data conjoined from the width and height columns. Specifically, the four histograms plotted in the subplot shows us that the data based on the top, width, height, and left columns all form into a bounding box.","metadata":{}},{"cell_type":"markdown","source":"And just like that, we mastered through the first section of analyzing the base_helm_df dataframe, though we skipped some columns for our analysis because we've encountered a \"SIGILL\" crash and irresponsive notebook editor loads. Without a doubt, let's proceed towards analyzing the train_labels dataframe, as long as we keep our analysis minimal.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d50a0a; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Section 2: labels_df</h2>\n\nAs we enter this section based on the train_labels dataframe, we knew that the train_labels dataframe contains every row for every possible combinations of players, including those who are with the ground for every point one second timestamp in a play. Without further ado, let's visualize the step column into a box plot! To get started, we use the fig variable to the box function from the px module, setting the train_labels dataframe as our data for the box chart and the y parameter to the step column for configuring the values of the box chart. With that completed, we use the show function to the fig variable for showing off our chart.","metadata":{}},{"cell_type":"code","source":"fig = px.box(labels_df, y=\"step\") \nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the box chart, we spotted that the median of the view column is 38, the first and third quartiles is 19 and 59, and the min, max of it is 0 and 172. Additionally, the view column in the labels_df dataframe shows the timestep intervals from each play of the games, as it increased by 1 or 0.1 seconds from 0, the starting point of a play.","metadata":{}},{"cell_type":"markdown","source":"With the step column visualized, let's proceed to envisage the nfl_player_id_1 and nfl_player_id_2 in two subplots! First, let's create our fig variable to the subplot configuration with the make_subplots function, setting the rows parameter to 1 and the cols parameter to 2 for configuring the rows and columns of our subplot. Secondly, we configure the traces two times with the add_trace function to the fig variable subplots, and for the first trace, we apply the histogram configuration with the go module's Histogram function with the x parameter set to the labels_df dataframe's nfl_player_id_1 for the x-axis configuration of the graph and the name to the same name of the column, setting the row parameter to 1 and the col parameter to 1 for placing the histogram to the first row and column and for the second trace, we apply the bar chart with the Bar function from the go module, in which the x parameter is set to the indexes of the labels_df dataframe's nfl_player_id_2 column counted by values with the value_counts function from the index attribute for the x-axis of the bar chart, the y-axis to the same thing as the x parameter, but specified by values with the values attribute, and the name parameter to the same name of the bar chart column thus setting the row parameter to 1 and col parameter to 2 for placing the bar chart to the second column and in the first row. Finally, we use the show parameter to the fig subplot graph for displaying the two graphs in a subplot.","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=2)\n\nfig.add_trace(\n    go.Histogram(x=labels_df[\"nfl_player_id_1\"], name=\"nfl_player_id_1\"),\n    row=1, col=1\n)\n\nfig.add_trace(\n    go.Bar(x=labels_df[\"nfl_player_id_2\"].value_counts().index, y=labels_df[\"nfl_player_id_2\"].value_counts().values, name=\"nfl_player_id_2\"),\n    row=1, col=2\n)\n\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As for the histogram consisting over the nfl_player_id_1 column, we espied several spikes of data, as some were short and some were tall just like the previous graph we visualized in our base_helm_df dataframe analysis. On the other hand, we visualized that the bars from the nfl_player_id_2 column data were faintly small, but we see that the \"G\" data is counted the most with 410633 data entities. Nevertheless, the data behind the nfl_player_id_1 and nfl_player_id_2 columns shows how the low and large numbers were listed to the id for gathering the contact pair whether any NFL player made contact with someone or even with the ground, hence we saw the \"G\" in the nfl_player_id_2 column.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's look over the contact column into our pie chart! Before we begin this, we create the contact dataframe from counting the values with the value_counts function to the labels_df dataframe with the contact column. From that time, we create our variable fig to the pie chart configuration with the pie function from the px module, setting the contact dataframe as the pie chart's data, the names parameter to the index specification with the contact dataframe by the index attribute for configuring the labels of the pie chart, and the values parameter to the value specification of the contact dataframe with the values attribute for configuring the pie chart's values. With this configuration finished, we manifest our fig variable figure with the show function.","metadata":{}},{"cell_type":"code","source":"contact = labels_df[\"contact\"].value_counts()\n\nfig = px.pie(contact, names=contact.index, values=contact.values)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our pie chart we created, we found out that 98.6% of the data inside the contact column from the labels_df were marked as False (indicated as zero), while the other 1.37% of the data were marked as True (indicated as one). Specifically, that explains us that in most cases from each video, each player contact or ground contact can occur sometimes by rough play or tripping and trampling.","metadata":{}},{"cell_type":"markdown","source":"And so, we've finished our analysis in the second section based on the labels_df dataframe! Doubtlessly, let's zoom over to the third section in our EDA analysis on the Player Contact Detection competition, which we're going to analyze the player_tracking_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #013369; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Section 3: player_tracking_df</h2>\n\nAs we got into the section about our upcoming analysis to the player_tracking_df dataframe, here's the background information behind this dataframe. the player_tracking_df dataframe shows the data that is gathered from the sensors players wear, so that the NFL can locate them for spotting their contact with each other and possibly injuries. So before we begin our section analysis, we import the seaborn module as sns to keep our analysis minimal because we may expect notebook crashes.","metadata":{}},{"cell_type":"code","source":"import seaborn as sns","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we've imported the seaborn module as sns, let's visualize the nfl_player_id in a displot! To do that, we use the sns module's displot function, setting the player_tracking_df as our data for the displot, and the x parameter to the nfl_player_id for the displot's x-axis. ","metadata":{}},{"cell_type":"code","source":"sns.displot(player_tracking_df, x=\"nfl_player_id\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the displot plotted after the above code cell compiled it, we spotted high spikes of data that was distributed from the nfl_player_id column, in which the highest data listed in the displot is roughly in 45600, with estimated 80200 data entities listed. Specifically just like our previous analysis based on the labels_df and base_helm_df dataframes, this data gathered in the nfl_player_id dataframe shows the NFL player's id code for tracking their play whether contact or injuries occur in each game. ","metadata":{}},{"cell_type":"markdown","source":"Next, let's move onwards to visualize the step column into the same displot in Seaborn! Once again, we use the sns module's displot function, setting the player_tracking_df dataframe as our data input for the displot, and the x parameter to the step column for specifying the x-axis of the displot but this time, we set the kind parameter to kde for visualizing our displot in kernel-density estimation.","metadata":{}},{"cell_type":"code","source":"sns.displot(player_tracking_df, x=\"step\", kind=\"kde\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the displot that showed the kernel density estimation on the step column, we visualized that the data distributed in the diagram is skewed to the right though it stayed in the left portion of it. Specifically, the range from 0 to 100ish is counted the most, as we espied the high peak of data in this range. In addition, the data distributed from the step column noted us that they represented the timesteps that were within the relative starting time of the play.","metadata":{}},{"cell_type":"markdown","source":"As we still use seaborn for our data analysis, let's visualize them into the barplot graph! Prior to graphing this out, we create the team dataframe from counting the player_tracking_df dataframe's team column with the value_counts function. After that, we use the sns module's barplot function to configure our barplot graph, setting the x parameter to the indexes of the team dataframe specified by the index attribute, the y parameter to the values of the team dataframe specified by the values attribute, and the data parameter set to the player_tracking_df dataframe for specifying the data for the barplot.","metadata":{}},{"cell_type":"code","source":"team = player_tracking_df[\"team\"].value_counts()\nsns.barplot(x=team.index, y=team.values, data=player_tracking_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see after the barplot is generated, we visualized that the bar heights of home and away in the player_tracking_df dataframe's team column are the same. Specifically, the number of data counted from the player_tracking_df dataframe's team column gave us a clear understanding that there are lots of games logged from the players that participated the football match, as each game contained 22 players: 11 from home and 11 from away.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the player's position and categorize them into the catplot graph! Before we begin plotting them in, we create the position dataframe to counting the values of the player_tracking_df dataframe with the value_counts function and convert it to the dataframe with the to_frame function thus resetting the indexes with the reset_index function and then renaming them with the rename function, setting the columns parameter to a dictionary in which the \"index\" key is assigned to \"position\" and the \"position\" key is assigned to \"count\" for specifying the columns when renaming it.\n\nAfter we create the position dataframe, we create our catplot with the sns module's catplot function, setting the x parameter to the position column for specifying the x-axis of the catplot, the y parameter to the count column for specifying the y-axis of the catplot, the data parameter to the position dataframe for importing the data to the catplot, the aspect parameter set to 2 for specifying the width of the catplot diagram, and the kind parameter to the bar function for specifying the type of catplot.","metadata":{}},{"cell_type":"code","source":"position = player_tracking_df[\"position\"].value_counts().to_frame().reset_index().rename(columns={'index': 'position', 'position': 'count'})\nsns.catplot(x=\"position\", y=\"count\", data=position, aspect=2, kind=\"bar\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw from the catplot diagram consisting over the data from the position column in the player_tracking_df dataframe, we found out that CB is counted the most, as there are roughly 160100 occurences to this data, while HB is counted the least, as there are about 200 occurences regarding to this data. Additionally, the CB data in the position column from the player_tracking_df dataframe, which is the cornerback position, is listed the most when there are occurences of rough plays throughout the position, as they blitzed and defended against extremely offensive running plays by sweeping and reversing against them.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the jersey_number column by distributing them into the displot and the boxplots and put them into the two subplots! Preparatory to plotting them into two plots, we import the matplotlib module with the pyplot submodule as plt, as this module is used for plotting graphs and most importantly, subplots. Following from importing this module, we characterize the fig and axes variables to the subplot configuration with the plt module's subplots function, setting 1 and 2 inside of it for configuring the subplot's one row and two columns, thus setting the figsize parameter to 10 by 5 for configuring the subplot's width and height. Afterwards, we add the suptitle for the subplot configuration with the suptitle plugged into the fig variable, in which we name it as \"Jersey Number Distribution\". \n\nNow it's time to plot the Seaborn figures into our subplot! As we configure the first plot, we use the histplot function from the sns module for configuring our seaborn-made histogram, setting the player_tracking_df dataframe as our data for the histplot graph, the x parameter to the jersey_number column for configuring the x-axis of the histplot diagram, and the ax parameter to the first index of the axes variable (0), just for placing the graph into the first subplot. Afterwards, we configure our box plot out from seaborn with the boxplot function from the sns module, setting the data parameter to the player_tracking_df dataframe as our data for the box plot diagram, setting the x parameter to the jersey_number for configuring the x-axes for the box plot, and the ax parameter to the last index of the axes variable (1) for placing our box plot into the second column of the subplot. Eventually, we launch our subplot graph with the show function from the plt module.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nfig, axes = plt.subplots(1, 2, figsize=(10,5))\nfig.suptitle(\"Jersey Number Distribution\")\n\nsns.histplot(player_tracking_df, x=\"jersey_number\", ax=axes[0])\nsns.boxplot(data=player_tracking_df, x=\"jersey_number\", ax=axes[1])\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we finished creating our histogram and box plot out from seaborn, we visualized that the histogram on the left shows variant distribution based out of the jersey numbers from the football players, as we espied several peaks of data in the diagram. Nevertheless, the highest number of distributions based on the jersey number data is roughly 23 to 25, as we counted 25550 data entities from that range. Meanwhile on the other graph, we found out that the median of the jersey number data from the box plot is roughly 55, while the first and third interquartiles are about 25 and 75 thus the interquartile range is 50 when we find the [difference between the two of them](https://www.varsitytutors.com/algebra_1-help/how-to-find-interquartile-range).\n\nSpecifically, the distribution based out of the jersey numbers from the participating football players hinted us that there's a lot of variation on every football player in NFL depending on the team they played for (e.g. Las Vegas Raiders, San Francisco 49ers, or Miami Dolphins).","metadata":{}},{"cell_type":"markdown","source":"Let's visualize the x and y positions in the displot based on the team columns! We define the fig and axes variables to the plt module's subplots function for configuring the subplots of our graph, setting 1 and 2 as one row and two columns thus setting the figsize parameter to 10 by 5 for specifying the width and height of the graph and the sharey parameter to True for enabling our subplots to share the y-axes together all at once. Thus, we configure our title for the subplot graph with the suptitle module from the plt module, as we set it to \"X and Y Distribution by Team\".\n\nAfter we configure the subplots and the title for the graph, we plot down two histograms and kernel density estimations with the sns module's histplot function, setting the player_tracking_df dataframe as our data for the histogram, the hue parameter to team for specifiying the data categorization in the histograms, and the kde parameter to True for displaying the kernel-density estimation in the graph. However, for the x parameter as for configuring the x-axes of the histograms, the first graph is configured to the x_position column while the second graph is configured to the y_position column, thus for the ax parameter when placing the graph in, the first graph is configured to the first index of the axes variable (0) while the second graph is configured to the second index of the axes variable (1). Afterwards, we use the plt module's show function to display out the graph subplots.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(10,5), sharey=True)\nfig.suptitle(\"X and Y Distribution by Team\")\n\nsns.histplot(player_tracking_df, x=\"x_position\", hue=\"team\", kde=True, ax=axes[0])\nsns.histplot(player_tracking_df, x=\"y_position\", hue=\"team\", kde=True, ax=axes[1])\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following from a long run to compile this two histograms that laid in two subplots, we espied that both the x_position and the y_position data distributions showed a nearly symmetrical shape, though the x_position data distribution has high counts of data than the y_position data. Meanwhile for the kde visualization, we found out that they are nearly in symmetrical shape too, just like the distributed bins of data gathered from both x_position and y_position columns. Not to mention, the distributed data based on the x and y positions gathered in the player_tracking_df dataframe indicates the players' movement from the sensor data, as they were assembled and positioning together before the game starts and tackle together during the game, depending on the roster from home or away.","metadata":{}},{"cell_type":"markdown","source":"Speaking about the x and y positions that is logged in the player_tracking_df dataframe, let's distribute them into a scatter plot made with displot! We use the displot from the sns module for creating the scatter plot, setting the data parameter to the player_tracking_df dataframe for setting the data in the scatter plot, the x parameter to the x_position column for setting the x-axes of the graph, the y parameter to the y_position column for setting the y-axes of the graph, and the hue parameter to the team column for specifying the color legend based out of another distributive data.","metadata":{}},{"cell_type":"code","source":"sns.displot(data=player_tracking_df, x=\"x_position\", y=\"y_position\", hue=\"team\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the x_position and y_position graphs we gathered into the scatter displot, we found out that there are lot of scattered data shown around the diagram, as the away plots were graphed in the center and around the border of the graph, while the home plots were also graphed in the center, overlapping the away plots and somehow got plotted inside and around the border. Additionally, the cluttered points from the x_position and y_positions columns by away and home category gave us a hint that almost all players scatter around the football field, trying to tackle each other just for getting the ball to the goals, as the overlapped plots in the middle of the diagram represents the hotspot of where players made rough play against each other.","metadata":{}},{"cell_type":"markdown","source":"And just like that, let's plow into creating two subplots regarding about the data distribution and visualization of speed and distance columns into two box plots. Once again, we create the fig and axes variable to the subplot configuration with the plt module's subplots function, setting 1 and 2 inside of it for setting the subplot's one row and two columns thus arranging the figsize parameter to 10 by 5 for specifying the width and height of the subplots. Moreover, we apply the title to the subplots graph with the plt module's suptitle graph naming it as, \"Speed and Distance Box Plot Distribution\".\n\nFollowing from configuring our subplots including the title, we configure two box plot figures with the sns module's boxplot function, setting the x parameter to the player_tracking_df dataframe's speed and distance columns separately, as we configure the data and x-axes to the boxplot graph all at once thus setting the ax parameter to the axes variable's first and last slice index separately (0 and 1), so that we can place the plots inside the subplot. With that completed, we present out our two box graphs with the show function given from the plt module.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(10,5))\nplt.suptitle(\"Speed and Distance Box Plot Distribution\")\n\nsns.boxplot(x=player_tracking_df[\"speed\"], ax=axes[0])\nsns.boxplot(x=player_tracking_df[\"distance\"], ax=axes[1])\n\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we create our two subplots based on the box plot distribution of speed and distances columns from the player_tracking_df dataframe, we found out that the quartiles, median, and the range is short than the outliers on the further right of the two diagrams. Nevertheless, the median for speed and distance data is roughly 0.6 and 0.3, their first and third quartiles is around 0.1 and 1.5 and 0.1 and 0.23, and their interquartile range is about 1.4 and 0.13. Specifically, the outliers shown in both two graphs hinted us that some players ran fast in some number of yards, as some of their speeds are measured up to 15 and some distances measured up to 2.5ish.","metadata":{}},{"cell_type":"markdown","source":"Let's proceed to create the four subplots based on the data from the direction, orientation, acceleration, and sa columns with the kdeplot! Basically, we setup our fig and axes variables to the subplot configuration with the plt module's subplots function, in which we set up the subplot's 2 rows and 2 columns thus configuring the figsize parameter to 10 by 5 for specifying the subplot's width and height. Furthermore, we configure the subplot's title with the plt module's suptitle function, titling it as \"Direction, Orientation, Acceleration, and Sa Distribution\".\n\nNow that we have the subplots configured, we each add our \"kernel-density estimation\" plots into each subplot with the sns module's kdeplot function, setting the data parameter to the player_tracking_df dataframe for loading the data to all four kdeplots, but we configure the x parameter to the direction, orientation, acceleration, and sa columns separately for specifying the x-axis of the kdeplots and the ax parameter to the axes variable's slice indexes of arrays consisting of [0,0], [0,1], [1,0], and [1,1] separately as we place the graphs to four subplots. After we configure our subplots and the kde graphs, we use the show function from the plt module for displaying off the kde graphs, from each of the four subplots.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 2, figsize=(10,5))\nplt.suptitle(\"Direction, Orientation, Acceleration, and Sa Distribution\")\n\nsns.kdeplot(data=player_tracking_df, x=\"direction\", ax=axes[0,0])\nsns.kdeplot(data=player_tracking_df, x=\"orientation\", ax=axes[0,1])\nsns.kdeplot(data=player_tracking_df, x=\"acceleration\", ax=axes[1,0])\nsns.kdeplot(data=player_tracking_df, x=\"sa\", ax=axes[1,1])\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw in each of four subplots consisting over the data distribution from the direction, orientation, acceleration and sa columns, we noticed that the kernel-density plots based on the distribution of the direction and orientation data columns were symmetrical, along with the sa data distribution though it started in the left direction. On the other hand, the acceleration data in the kdeplot is skewed to the left, as it occured on the range from 0 to 5. Additionally, the data that is collected from direction and orientation reflects how football players position themselves in different angles, whether they go on the offensive or in defensive positions, while for the acceleration and the sa data, players might travel fast and slow, whether they made contact with each other or dash towards the touchdown goal, as their acceleration vary up to 4.49 yards per second in a 40 yard dash.","metadata":{}},{"cell_type":"markdown","source":"As we finished the direction, orientation, acceleration, and sa data visualization above, we completed the section based on analyzing the data from the player_tracking_df dataframe! As we look forward towards the next one, we're going to expect visualizing the video_metadata_df dataframe thus one of the videos from the competition data because of the less number of columns in this dataframe we're going to visualize, which it will be a two-in-one section.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d50a0a; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Section 4: video_metadata_df / Video Visualization</h2>\n\nAs we got into our data analysis based on the video_metadata_df dataframe and the videos, we are going to visualize the distribution based on the start_time, end_time, and snap_time columns and then analyze the video footage from the competition files, as we separate the section into two parts. \n\n<h3 style=\"background-color: #013369; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Part 1: video_metadata_df Analysis</h3>\n\nNevertheless, we start converting the start_time, end_time, and snap_time columns into the standard datetime format by redefining the columns to itself that is applied by the apply function with the formatting the datetime with the pd module's to_datetime module. Following that, we create the duration column in the video_metadata_df dataframe to calculating the total seconds of the video_metadata_df dataframe's end_time column that is subtracted from the start_time column with the dt attribute followed by the total_seconds function. Moreover, we define the video_metadata_df dataframe's year_snap, month_snap, and day_snap columns to the column conversion of the video_metadata_df dataframe's snap_time column to year, month, and day with the dt attribute followed by the year, month, and day attributes.","metadata":{}},{"cell_type":"code","source":"video_metadata_df[\"start_time\"] = video_metadata_df.start_time.apply(pd.to_datetime)\nvideo_metadata_df[\"end_time\"] = video_metadata_df.end_time.apply(pd.to_datetime)\nvideo_metadata_df[\"snap_time\"] = video_metadata_df.snap_time.apply(pd.to_datetime)\nvideo_metadata_df[\"duration\"] = (video_metadata_df[\"end_time\"] - video_metadata_df[\"start_time\"]).dt.total_seconds()\n\nvideo_metadata_df[\"year_snap\"] = video_metadata_df[\"snap_time\"].dt.year\nvideo_metadata_df[\"month_snap\"] = video_metadata_df[\"snap_time\"].dt.month\nvideo_metadata_df[\"day_snap\"] = video_metadata_df[\"snap_time\"].dt.day","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we create our duration, year_snap, month_snap, and day_snap columns from modifying the start_time, end_time, and snap_time columns in the video_metadata_df dataframe, let's create our duration data visualization consisting of two graphs in two subplots! We configure our subplot graphs by defining the fig and axes variable to the plt module's subplots function, setting 1 and 2 as our row and column configuration, thus the figsize parameter to 8 by 5 for configuring the width and height of the subplot graph. Moreover, we title our subplot graph with the plt module's suptitle function, setting it to \"Time Duration Distribution\".\n\nOnce we finish configuring our subplots, we add the histogram with the sns module's histplot function, setting the video_metadata_df dataframe as our data for the histplot, the x parameter to the duration column we created previously for setting the x-axes of the histogram, the kde parameter to True for visualizing the kernel-density estimation of the histogram, and the ax parameter to the first slice index of the axes variable (indicated as 0) for placing the histogram to the first column.\n\nFollowing from creating our histogram into the first subplot column, we characterize the box plot graph with the boxplot function given from the sns module, setting the x parameter to the video_metadata_df dataframe's duration column for configuring the x-axes specified in the box-plot, and the ax parameter to the last slice index of the axes variable (indicated as 1) for placing the box plot to the second column.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(8,5))\nplt.suptitle(\"Time Duration Distribution\")\n\nsns.histplot(data=video_metadata_df, x=\"duration\", kde=True, ax=axes[0])\nsns.boxplot(x=video_metadata_df[\"duration\"], ax=axes[1])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based from the two plots consisting of the duration distribution, we found out that the histogram in the left of us was skewed left, along with the kde line chart. Nevertheless, the most counted data specified in the duration column by the histogram chart is around 13, with almost 70 entities, while around 10.5 is the least counted, as there are around 2 entities listed. Meanwhile, in the box chart, we realized that the box plot is shifted to the left of the graph, with some outliers to the right. In other words, the first quartile is around 12, the median is around 12.7, the third quartile is around 13, and the interquartile range is around 1. Specifically, the time duration distribution explains us that the duration of the videos specified in the competition data may vary, as some videos lasted for 11 to 15 seconds, while others lasted short as 9 seconds and long as 22 seconds, depending on how long the two teams played their moves.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the year, month, and day snap distributions into three subplots with our histplot! Again, we create our subplots by defining our fig and axes variables to the plt module's subplots function, setting the values 1 and 3 as our rows and columns for the subplots, the figsize parameter to 12 by 5 for configuring the width and height of the subplots, and the sharey parameter to True for sharing the y-axes of the graphs thus configuring the title of our subplot graph with the subplot's title with the plt module's suptitle function, setting it to \"Year, Month, and Day Visualization\". \n\nNow that we have the subplots configured, we arrange the histograms with the histplot function from the sns mdoule, setting the data parameter to the video_metadata_df dataframe as the histogram's data, the x parameter to year_snap, month_snap, and day_snap columns separately for configuring the x-axes for the histogram, the kde parameter to True for visualizing the kernel-density estimation of the histogram, and the ax parameter to the axes variable's slice indexes of 0, 1, and 2 separately for placing the histograms into the subplots.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(12,5), sharey=True)\nplt.suptitle(\"Year, Month, and Day Distribution\")\n\nsns.histplot(data=video_metadata_df, x=\"year_snap\", kde=True, ax=axes[0])\nsns.histplot(data=video_metadata_df, x=\"month_snap\", kde=True, ax=axes[1])\nsns.histplot(data=video_metadata_df, x=\"day_snap\", kde=True, ax=axes[2])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we visualize based on the data distribution of year_snap, month_snap, and day_snap columns, we found out that for the year_snap distribution, we saw two bars separate from each other, and for the month_snap distribution, we espied the left skewed bars that start on the right, and lastly for the day_snap distribution, we envisaged the nearly symmetrical and bimodal bars, as what we saw is similar to the three graphs' kde plots. Nevertheless, 2020 is counted the most in the year_snap data distribution, along 9 in month_snap, and roughly 10-11 in the day_snap distributions. In other words, the distributions based out of the years, months, and days from the video_metadata_df's created year_snap, month_snap, and day_snap columns explains us that the NFL videos in the competition data were recorded during the 2020 NFL season, as the games played contained most player contacts that may be injurious.\n\nWith our short, short part of our video_metadata_df dataframe analysis finished, let's proceed to visualize the videos gathered into the competition data, to see whether the contacts made by the players is injurious or not.","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"background-color: #013369; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Part 2: Video Analysis</h3>\n\nAs we approach into our video analysis, here's our game plan for this section. We first process and slice the videos from the competition data, and then visualize them by analyzing the whole part of the video and by image frame.\n\nTo get started, we import the cv2 module first, then define the vidcap variable to the cv2 module's VideoCapture function to capture a video slice from any path directory leading to the video files in the competition data such as \"58177_004239_Endzone.mp4\". With that done, we characterize the success and image variables to read out the given video.","metadata":{}},{"cell_type":"code","source":"import cv2\nvidcap = cv2.VideoCapture(\"/kaggle/input/nfl-player-contact-detection/train/58177_004239_Endzone.mp4\")\nsuccess, image = vidcap.read()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T01:00:03.373533Z","iopub.execute_input":"2023-01-19T01:00:03.374030Z","iopub.status.idle":"2023-01-19T01:00:03.791702Z","shell.execute_reply.started":"2023-01-19T01:00:03.373937Z","shell.execute_reply":"2023-01-19T01:00:03.790423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have our success and image variables configured, we redefine the image variable to convert the color of the video frame with the cv2 module's cvtColor function, setting the image variable inside as the given image, and the cv2 module's COLOR_BGR2RGB attribute for converting the bgr color to rgb. Thus, we configure our figure graph with the figure function from the plt module, setting the figsize parameter to 16 by 8 for configuring the width and height of the plot and show the image from the image variable by using the plt module's imshow function.","metadata":{}},{"cell_type":"code","source":"image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\nplt.figure(figsize=(16, 8))\nplt.imshow(image)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T01:00:03.794137Z","iopub.execute_input":"2023-01-19T01:00:03.794552Z","iopub.status.idle":"2023-01-19T01:00:04.515755Z","shell.execute_reply.started":"2023-01-19T01:00:03.794517Z","shell.execute_reply":"2023-01-19T01:00:04.514614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, the image shown above in the matplotlib graph shows us that the players were prepared for a match of football, as we understand that they might tackle and play rough to each other. Specifically, the image frame from one of the competition videos hinted us that in every round during a game of football, players assemble in their positions, bracing themselves from injuries, as they might make foul player contact.","metadata":{}},{"cell_type":"markdown","source":"Now let's create our plot based out of the frames in this video we analyzed! First of all, we create our figure plot with the plt module's figure function, setting the figsize parameter to 16 by 48 for configuring the width and height of our figure, thus defining the np_video variable to an empty list for listing the image frames of the video data, the count variable to 0 for the counter of the figure plot, and the step variable to 40 for configuring the skips of the video data. \n\nFollowing from configurating our figure and variables, we create a try-except statement, as we create a while loop when the success variable is True inside the try statement, then the image variable is defined to converting the given image's colors with the cv2 module's cvtColor function, setting the image variable as the specified image and the cv2 module's COLOR_BGR2RGB attribute for converting the BGR to RGB color scale thus appending the values from the image variable to the np_video list with the append function. Afterwards, an if-statement is summoned whether it is not the modulus division between the count and step values, then the subplot is created to the plot from the plt module's subplot function, setting 6 and 2 as our values for the subplot's rows and columns, and the division of count and step variables added by 1 for iterating the subplots for placing the image frames in. Moreover, the image will be displayed to each subplot with the plt module's imshow function, containing the image variable for loading the given image, then the axes of the subplot will turn off with the axis function from the plt module that contained the \"off\" string. Afterwards, the success and image variables is redefined to read the frames of the video with the read function from the vidcap module, thus incrementing the count variable value by 1. And as for the except statement, we added the ValueError inside of it (num must be 1 <= num <= 12 not 13 error) and use the pass statement to nullify that error we had.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(16, 48))\n\nnp_video = []\nvidcap = cv2.VideoCapture(\"/kaggle/input/nfl-player-contact-detection/train/58177_004239_Endzone.mp4\")\nsuccess, image = vidcap.read()\n\ncount = 0\nstep = 40\n\ntry:\n    while success:\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        np_video.append(image)\n        if count % step:\n            plt.subplot(6, 2, count // step + 1)\n            plt.imshow(image)\n            plt.axis(\"off\")\n        success, image = vidcap.read()\n\n        count += 1\nexcept ValueError:\n    pass","metadata":{"execution":{"iopub.status.busy":"2023-01-19T01:00:04.517238Z","iopub.execute_input":"2023-01-19T01:00:04.517587Z","iopub.status.idle":"2023-01-19T01:01:37.655431Z","shell.execute_reply.started":"2023-01-19T01:00:04.517557Z","shell.execute_reply":"2023-01-19T01:01:37.654213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Upon every frame at a time we analyzed inside the subplot we generated, we can see how the football players started to make contact to each other, meaning that there may be injurious occurances while tackling and shoving roughly.","metadata":{}},{"cell_type":"markdown","source":"However, that was some samples from the video frames we visualized, which led us creating our animation video. Without further ado, let's generate the whole video with Matplotlib! Before we begin generating our video animation, we import the animation and rc functions from the matplotlib module, thus calling the rc function to configure our video display, setting 'animation' for animating the video display, and the html parameter to 'jshtml' for configuring the port for displaying the specified frames in a video. \n\nAfter that, we characterize the create_animation function, setting the ims variable as the input parameter of this function. Inside the create_animation function, we characterize the fig variable to configure the figure of the graph, setting the figsize parameter to 6 by 6 for specifying the width and height of the figure, thus turning off the figure's axis with the plt module's axis function that contains the \"off\" string. Thenceforth, we characterize the im variable to the plt module's imshow function to show images into the plot, setting the ims parameter input's first slice index as 0, and the cmap parameter to gray for configuring the colormap of the image. We then create another function and name it as animate_func, thus setting the i variable as the input parameter of this function. While we're in the animate_func function, we configure the array into the im variable with the set_array function, containing the ims input parameter's index indicated as the i variable, and then we return the im variable inside the square brackets. After we configured the inside of the animate_func function, we return our configured video animation with the animation module's FuncAnimation function (for creating the animation into the video) out off the create_animation function, in which we set the fig variable as the specified plot graph, the animate_func function for animation configuration, the frames parameter to the number of entities in the ims variable specified by the len function for arranging the number of frames, and the interval parameter to 1000 divided integrally with 24 for configuring the interval value.","metadata":{}},{"cell_type":"code","source":"from matplotlib import animation, rc\nrc('animation', html='jshtml')\n\ndef create_animation(ims):\n    fig = plt.figure(figsize=(6, 6))\n    plt.axis('off')\n    im = plt.imshow(ims[0], cmap='gray')\n    \n    def animate_func(i):\n        im.set_array(ims[i])\n        return [im]\n    \n    return animation.FuncAnimation(fig, animate_func, frames=len(ims), interval=1000//24)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T01:01:37.657633Z","iopub.execute_input":"2023-01-19T01:01:37.657973Z","iopub.status.idle":"2023-01-19T01:01:37.669831Z","shell.execute_reply.started":"2023-01-19T01:01:37.657941Z","shell.execute_reply":"2023-01-19T01:01:37.668606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we characterized our create_animation function, we generate our animation video with the create_animation function, setting the np_video variable as the specified video frame input to animate with!","metadata":{}},{"cell_type":"code","source":"create_animation(np_video)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T01:01:37.671220Z","iopub.execute_input":"2023-01-19T01:01:37.671550Z","iopub.status.idle":"2023-01-19T01:02:21.884403Z","shell.execute_reply.started":"2023-01-19T01:01:37.671520Z","shell.execute_reply":"2023-01-19T01:02:21.883292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As for the video, we visualized how the football players postition themselves just for preparing to tackle for the football but we cannot see the whole scene in where they tackled or played rough to each other. However, the image frame below the video animation shows how the football players shoved or pushed each other brutally, as we hint that they may get injured from each move they made.","metadata":{}},{"cell_type":"markdown","source":"And as our finale, let's visualize the color map of the football players specified in the video! As we follow the steps from [Towards Data Science](https://towardsdatascience.com/football-players-tracking-identifying-players-team-based-on-their-jersey-colors-using-opencv-7eed1b8a1095), we characterize a variable which we name it as color_list, to the list containing three strings: \"red\", \"white\", and \"blue\". We then configure the boundaries variable to another list, containing three objects, containing two tuples, which is: (17, 15, 75), (50, 56, 200) for red, (43, 31, 4), (250, 88, 50) for blue, and (187, 169, 112) and (255, 255, 255) for white. Thus, we characterize the frame_img variable to the np_video list variable's any slice index.","metadata":{}},{"cell_type":"code","source":"color_list = [\"red\", \"white\", \"blue\"]\n\nboundaries = [\n    ((17, 15, 75), (50, 56, 200)), # RED\n    ((187, 169, 112), (255, 255, 255)), # WHITE\n    ((43, 31, 4), (250, 88, 50)), # BLUE\n]\n\nframe_img = np_video[4]","metadata":{"execution":{"iopub.status.busy":"2023-01-19T01:02:21.885711Z","iopub.execute_input":"2023-01-19T01:02:21.886050Z","iopub.status.idle":"2023-01-19T01:02:21.892083Z","shell.execute_reply.started":"2023-01-19T01:02:21.886019Z","shell.execute_reply":"2023-01-19T01:02:21.890967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have our color_list and boundaries variables configured and ready to go, we create the color and boundary variables to loop the zip objects of the color_list and boundaries variables specified by the zip function. While we're in the for loop we created, we characterize the mask variable to the cv2 module's inRange function for detecting the objects based on the range of the pixel values in the HSV colorspace, which is the frame_img variable, and the first and second slice indexes of the boundary variable, along with the output variable to the cv2 module's bitwise_and function to compute the bit-wise and of two array element-wise objects, placing the frame_img variable twice, thus setting the mask parameter to the mask variable for outlining the specified object. Thenceforth, we plot our figure with the plt module's figure variable, setting the figsize parameter to 12 by 12 for configuring the width and height of our figure. Following from that, we add our subplot with the plt module's subplot function, setting the values, 2, 2, and 1 as the rows, columns, and index specification for the subplot thus displaying the specified image with the plt module's imshow function, which is the frame_img variable. Same goes to the output and mask variables, but the indexes for each subplot is 3 and 2.","metadata":{}},{"cell_type":"code","source":"for color, boundary in zip(color_list, boundaries):\n    mask = cv2.inRange(frame_img, boundary[0], boundary[1])\n    output = cv2.bitwise_and(frame_img, frame_img, mask=mask)\n    \n    plt.figure(figsize=(12,12))\n    plt.subplot(2,2,1)\n    plt.imshow(frame_img)\n    \n    plt.subplot(2,2,3)\n    plt.imshow(output)\n    \n    plt.subplot(2,2,2)\n    plt.imshow(mask)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T01:02:21.893651Z","iopub.execute_input":"2023-01-19T01:02:21.893953Z","iopub.status.idle":"2023-01-19T01:02:24.827043Z","shell.execute_reply.started":"2023-01-19T01:02:21.893925Z","shell.execute_reply":"2023-01-19T01:02:24.825956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we compiled the three subplots separately based on the image frames we visualized, we saw the white color and the red color masks outlined in the last two subplots. However, we didn't saw the blue outlined in the first subplot, as there are no color matches to the frame of the image frame we observed. Not only that, we observed the foreign objects in the last two subplots that consisted the masks of the red and white colors, like the football pole and the grass shadows from the football players along with the outline of the text in the football field. In other words, the color masks on each frame of the videos in the competition data varies, depending on the jersey color football players wore from one team they played, as each showed two colors (red white, red blue, blue white), one color, or no colors.\n\nAnd just like that, we finished the part where we process and analyze the videos in this section thus finishing our whole data analysis in this competition data we're in!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #d50a0a; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion</h2>\n\nFrom our data analysis we've journeyed through after a hard series of page crashes (SIGILL encounters), we highlighted that there's mostly less contact throughout each of the whole game football players played, meaning that there's less likely for injurious reports from their brutal moves like shoving or tackling. And hopefully once we train our computer model to detect any player injuries from this competition, we hopefully will decrease the number of injurious reports from player contacts. And w","metadata":{}}]}