{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nA change in possession represents a small victory that must be continued with each additional play. Punt returns are key in building on this momentum and ensuring the receiving team is in a good position to give their offensive squad an advantageous field position. During a punt, the returner is the focal point, and every decision he makes is crucial to gain maximum yardage. A punt returner’s ability to make split second decisions, correctly assess nearby opponents, evaluate blocking leverages, and identify open paths can separate him from the rest and give their team a winning edge. In quantifying these metrics, we determine the best punt returners throughout the league by evaluating their performances. \n\nFocusing on attempted punt returns:\n\n1. **We construct an algorithm to determine the punt returner’s optimal path instantaneously as the play progresses, conditional on the position of all players for that frame, and evaluate their decision-making through traffic.** Our method allows for review immediately after the play to visualize where the optimal path lies in comparison to the path taken. By updating at every frame, we can see the exact moments where the runner went astray and look to improve their performance for future punt returns. \n2. **We develop various metrics to evaluate returns and predict expected yards remaining for each punt return at every frame**. Our main developed metric is an *Adaptive Stochastic Convex Hull* that quantifies field control for each team. We then evaluate returners’ performances above expected and provide detailed rankings. These rankings are unique as they use measures that quantify both the mental and physical abilities of the punt returners and their interactions with their surroundings.  \n","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle"}},{"cell_type":"markdown","source":"# Data Preparation\n\nData for the analysis consist of all punts that resulted in a return from the 2018-2020 NFL seasons. Specifically, we include all frames from the moment the punt was received until the final event is complete (tackle, out-of-bounds, touchdown, fumble, etc.). Additionally, any plays that resulted in a penalty were removed *with the exception of penalties incurred after the completion of the play*. The data was then standardized using [Micheal Lopez's notebook](https://www.kaggle.com/statsbymichaellopez/nfl-tracking-wrangling-voronoi-and-sonars). We trained our models on the 2018-2019 seasons and tested on the 2020 season.\n\nThroughout this notebook, we use the punt return from the Dec. 26, 2020 game between Detroit and Tampa Bay where Jamal Agnew (Detroit Lions) completed a 74 yard return for a touchdown.","metadata":{}},{"cell_type":"markdown","source":"# Part I: Find the Best Path\n\nPunt returners must be quick on their feet and even quicker mentally. They must quickly process complex and chaotic information in their surroundings while simultaneously choosing a direction for their next step. In addition to different physical talents, some returners are undoubtedly better at processing the spatial information than others. We look to **quantify the decision-making ability of the punt returner in each play by constructing an algorithm to determine their optimal path** at any given frame. We define \"*optimal*\" as the path that has the largest potential increase in yards while posing minimal risk of being tackled. **We can then use the various metrics relating to optimal decision-making in Part II to evaluate whether punt returners perform as expected**.\n\n## Methods\n\nAssume that the returner identifies all potential tacklers at the moment of the reception, and tries to find a path that avoids them for as long as possible to maximize yardage. Rather than running straight into his opponents, he will search for the 'windows’ between them that he can run through to evade the tackle. However, the positions and windows are constantly changing, so the returner must continually re-evaluate the state of each window and anticipate the creation of new ones based on opponent directions, speeds and potential blocking from teammates.  \n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/Optimal%20Path%20Infographic.png?raw=true\">\n\nTo determine the optimal path at any given frame, we start by defining ‘*windows*’ as the potential gaps between players or along the sideline that the returner can pass through that are constructed using Delaunay Triangulation (Lee & Schachter, 1980). \n\nNext, we assess the relative pressure of the tacklers at evenly spaced points along each window by calculating the *expected arrival time of all tacklers* to the targeted point at 7 yards/sec and *incorporating a time penalty for each blocker* that may impede their path to the target. Intuitively speaking, this time penalty is:\n\n* **1-5 seconds** when the blocker is <5 yards to the tackler and directly in his path,\n* **0.1-1 seconds** when the blocker is >5 yards from the tackler but in the neighbourhood of his path, or\n* **0-0.1 seconds** when he is far enough to the side of the tackler’s path that he will likely not be able to block him. \n\nThis is meant to quantify the likely delay to be imparted to the player by the blockers. Selecting the minimum time of all penalized expected arrival times describes how quickly an opponent can arrive at that point and thus potentially complete a tackle should the punt returner arrive at that point. **Our goal is for the punt returner to pass through windows at points with a larger penalized expected arrival time where he can follow his blockers and gain as many yards as possible**.\n\nFinally, using an *adapted A * (A star) search algorithm* (Hart et al., 1968), we connect the point with the lowest risk in one window to the next and find the path to the end zone is optimal. This process updates constantly in discretized continuous time to project where the new best path is in each moment.\n\n### Additional specifics related to Step 2:\n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/Arrival%20Time%20Infographic.png?raw=true\" width=\"1000\">\n\n## Evaluation\n\nTo evaluate how much the punt returner’s path deviates from our optimal path we use the *Fréchet distance* (Aronov et al., 2006). This distance calculates the shortest connection required to bridge two forward moving curves. At each frame, **we compare the punt returner’s next 5-yard ** increase with the optimal path to gain a sense of how they are evaluating their current situation and whether they are continuing in the best direction available**.\n\n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/Frechet%20Infographic.png?raw=true\" width=\"700\">\n\nFinally, we take the median of the Fréchet path deviations of all frames starting with the point the ball was caught until either the punt returner passes the final kicking team player or the conclusion of the play. **The average of these performances across all punts returned by each returner gives us a measure of how well they find the optimal path in high pressure situations**.\n\n## Results\n\nUsing our selected play, we see Agnew’s path in light grey and the optimal path with the black arrows. For this play, **Agnew’s average Fréchet path deviation is 1.906 yards**. This implies that Agnew’s path in this play typically follows what we consider to be the optimal path and he did a good job at evaluating the situation he faced at each frame.\n\n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/Agnew%20Path.gif?raw=true\" width=\"800\">\n\nNext, we display the **average path deviation vs. the average yards gained** for all punt returners with a minimum of 10 attempts on a per season basis. We suggest that a lower path deviation can be interpreted as a returner having good decision-making and the foresight to move in the optimal direction. In comparing the log-scaled yards and path deviation on a per play basis, we found a **correlation of -0.44** reaffirming that our model provides a strong estimate of the optimal path available to the returner.","metadata":{}},{"cell_type":"code","source":"options(warn=-1)\nload.libraries = c(\"tidyverse\", \"plotly\", \"htmlwidgets\", \"IRdisplay\")\ninstall.lib = load.libraries[!load.libraries %in% installed.packages()]\nfor (libs in install.lib) {suppressMessages(install.packages(libs, dependencies = TRUE))}\nsuppressMessages(require(tidyverse)); require(plotly) %>% suppressMessages()\nrequire(htmlwidgets) %>% suppressMessages(); require(IRdisplay) %>% suppressMessages()\n\npath_deviation_by_carrier = read.csv(\"../input/path-deviations-by-carrier/path_deviation_by_carrier2.csv\")\n\n# Plot the average path deviation by the average yards on the play\np2 = path_deviation_by_carrier %>%\n  mutate(season = factor(season)) %>%\n  rename(`Punt Returner` = carrier_name,\n         `Season` = season,\n         `Average Path Deviation` = path_dev_med,\n         `Average Yards` = yards,\n         `Number of Attempts` = attempts,\n         `Team` = team\n  ) %>%\n  mutate(`Line of Best Fit` = \"\") %>%\n  ggplot() +\n  geom_point(aes(label2 = `Punt Returner`, label3 = `Team`, label4 = `Number of Attempts`, fill = `Season`, x = `Average Path Deviation`, y = `Average Yards`), size = 3.25, shape = 21, stroke = 0.25) +\n  geom_smooth(formula = \"y~x\", aes(x = `Average Path Deviation`, y = `Average Yards`, label2 = `Line of Best Fit`), method = \"lm\", se = F, colour = \"black\") +\n  labs(x = \"Average Path Deviation\", y = \"Average Yards\", title = \"Path Deviation by Punt Returner\", subtitle = \"Minimum 10 attempted returns\") +\n  geom_text(aes(x = 4.5, y = 19.5, label = \"Good Decision-Making,\\nHigh Skill\"), size = 3) +\n  geom_text(aes(x = 13.5, y = 19.5, label = \"Poor Decision-Making,\\nHigh Skill\"), size = 3) +\n  geom_text(aes(x = 4.5, y = 6, label = \"Good Decision-Making,\\nLow Skill\"), size = 3) +\n  geom_text(aes(x = 13.5, y = 6, label = \"Poor Decision-Making,\\nLow Skill\"), size = 3) +\n  theme_bw()\n\np = ggplotly(p2, tooltip = c(\"label2\", \"label3\", \"label4\", \"fill\"), height = 600, width = 662)\n\ndir.create(file.path(\"plots/\"), showWarnings = FALSE)\nf <-\"plots/table1.html\"\nsaveWidget(p, file.path(normalizePath(dirname(f)),basename(f)))\ndisplay_html('<iframe src=\"plots/table1.html\" align=\"center\" width=\"100%\" height=\"500\" frameBorder=\"0\"></iframe>')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T03:49:34.054956Z","iopub.execute_input":"2022-01-06T03:49:34.056498Z","iopub.status.idle":"2022-01-06T03:49:34.979707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Part II: Evaluate Plays\n\nClassically, metrics like *Expected Points Added* and *Quarterback Rating* are used to evaluate players and teams in the NFL. These metrics use the ending yardline of the play to determine what that play is worth. Frame-by-frame analysis is limited and few metrics are able to quantify its information. For some preliminary related research see Yurko et al. (2020). **We update our estimator at each frame using the information from the previous one and, therefore, update our metrics as the play progresses**.\n\n## Methods\nTo estimate the ability of punt returners, we look to estimate the *Return Yards Above Expected (RYAE)* by constructing a model for the expected yards on the return. A common metric to describe this is the tackle probability on each frame. Therefore, **we model the probability that a frame will end in a tackle using a Gradient Boosting Machine classifier** tuned using the CARET package (Kuhn, 2008). The features included in this model are: \n\n\n1. **Defender Distance** - The distance between each defender on the kicking team and the returner. \n2. **Blocker Leverage** - Using the expected position of the kicking team player relative to the blocker on the receiving team and the punt returner, *we estimated the blocker’s leverage on his nearest kicking team player* (Rumsey, 2020). With this, we estimate the likelihood of a blocker impeding their opponent’s path. We considered a player to be blocked if they are within 2 yards and have sufficient leverage.\n3. **Penalized Expected Arrival Time to the Returner** - see Part I.\n\nNext to estimate the distribution of the amount of return yardage remaining after each frame, we use a *Random Forest - Conditional Density Estimation (RF-CDE) model* (Pospisil & Lee, 2018). **The model’s output is a probability distribution that associates a probability to each potential return yardage** As the play progresses, the distribution and the probabilities are updated to reflect the varying conditions the returner faces. We can use the conditional median of the distribution as a point estimate for the return yards remaining on the play at any frame. This model utilizes the following metrics,\n\n1. **Tackle Probabilities** - see above.\n2. **Fréchet Path deviation** - see Part I.\n3. **The Adaptive Stochastic Convex Hull** - the simple *Convex Hull* creates a polygon around the players and is often used to quantify the space on the field that each team occupies (Guan et al., 2021). We adapt it to **exclude players based on the blocking leverage** or if they are 5+ yards behind the punt returner. \n\n\n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/Agnew%20Hull.gif?raw=true\" width=800>\n\nFinally, the **RYAE is defined as the average distance between the conditional median of the estimated probability density and the true remaining return yards on the play**. \n \n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/model_workflow.png?raw=true\"> \n\nBy dynamically updating our model with each frame we achieve a better estimate for the RYAE since we are:\n\n1. accounting for the previous information while the play progresses, and\n2. averaging estimates throughout the play rather than making one estimate per play.\n\nCalculating *the mean absolute error (MAE)* of the testing data, the predictions are off by an average of **2 yards for returns between 0-10 yards, 5 yards for returns between 10-20 yards, and 15 yards for returns between 20-30 yards**. It is clear that our model performs best for short returns, and that our accuracy decreases as the return yardage increases. This is to be expected, as the average return length for a punt is closer to 5-10 yards.\n\n## Results\n\nThe following displays the *conditional density* starting with the frame when Agnew receives the punt until he reaches the end zone. In the beginning, our model places very little weight on him returning the ball 70+ yards, but **as the play develops the conditional density updates to get closer to the true value**. The model could be improved if we had more plays of longer length to more accurately predict plays of this type.\n\n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/conditional_density.gif?raw=true\" width=\"700\"> ","metadata":{}},{"cell_type":"markdown","source":"# Rankings\n\nUsing our newly calculated metrics, we are able to **rank the 2020 punt returners that had 10 or more returns** (displaying the top 20). \n\n<img src=\"https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path/blob/main/Figures/player_rankings.png?raw=true\">","metadata":{}},{"cell_type":"markdown","source":"# How is this innovative and useful to the NFL?\n\n* **New innovative metrics** for evaluating special teams, including: *Adaptive Stochastic Convex Hull, Penalized Expected Arrival Time to Target, and Fréchet Path Deviation*. \n* **Visualization tools** for *post-play analysis, mid-week film sessions, in-game feedback, and TV broadcasts*. ","metadata":{}},{"cell_type":"markdown","source":"# Links\n\n [Code, data, additional visualisations and references.](https://github.com/ritchi12/punt_returns_using_the_math_to_find_the_path)","metadata":{}},{"cell_type":"markdown","source":"Word Count: 1923","metadata":{}}]}