{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# DEEPS Dive: A Process-Oriented Approach to Evaluating Punt Returners\n### By Zach Bradlow, Zach Drapkin, & Sarah Hu\n### The Wharton School, University of Pennsylvania","metadata":{}},{"cell_type":"markdown","source":"# 1. Introduction\n\nIn Philly, we have something of a motto: [Trust the Process](https://www.si.com/nba/2018/05/02/trust-the-process-meaning-philadelphia-76ers-team-motto). We like to evaluate decisions by the processes used to arrive at them, not by their outcomes. In football, when we evaluate punt returners, we often get distracted by results-oriented metrics such as yards per return. Metrics like these only look at what *actually* happened on the play and ignore what *could or should* have happened; they do not capture the value of the decisions made by punt returners, nor do they quantify the difficulty of the yards gained on a return. The goal of our analysis is to fill this gap by evaluating the decision-making and returning skills of punt returners through a process-oriented approach. \n\nUsing tracking data from all catchable punts over the course of the 2018-2020 NFL seasons, we created metrics to assess punt returner skill in both decision-making and returning. Decision Expected Expected Points Surrendered (DEEPS) assigns an expected point (EP) value to the field position we expect from a return, from a fair catch, and from letting the ball bounce and calculates whether the returner made the optimal decision between the three, as well as how large each error was in terms of EP. Punt Return Expected Points Added (PREPA) models the expected result of a punt return continuously throughout the play and estimates the frame-by-frame contributions of the returner to his team’s expected points relative to expectation. Evaluating players along these two dimensions allows NFL teams to grade punt returners with process-based metrics and enrich the insight they derive from quantifying special teams performance.","metadata":{}},{"cell_type":"markdown","source":"# 2. DEEPS Methodology\n## 2.1 Data\nWe built DEEPS using the tracking and scouting data provided, play-by-play data & expected points values from nflfastR, and weather data. To subset the data for only catchable punts, we imputed approximate z-coordinate data with [NFL3D’s methodology](https://dutta.github.io/nfl3d.html) and removed punts whose trajectory went out of bounds or punts listed in the data as out of bounds. With the same z-coordinate data, we were also able to roughly identify the apex of each punt; this was assumed (after consulting with NFL special teams staffers) to be the frame at which the return decision was made. For the decision frame (apex), we generated the following features:\n\n- **Returner**: speed, acceleration, orientation, direction, distance from the ball’s projected trajectory, distance from the ball’s apex, and NFL position in addition to their closest three teammates/opponents distances, speeds, accelerations, orientations, and directions\n- **Football**: location, velocities, launch angle, and height\n- **Punt Information**: number of gunners, vises, and punt rushers, and operation time from snap to punt\n- **Game Situation**: quarter, time remaining in the game, and point differential\n- **Weather Data**: temperature, humidity,  precipitation, wind speed, and pressure \n- **Field Control**: the voronoi area around the returner\n\n## 2.2 Model Components\nDEEPS consists of five separate models trained on the aforementioned features: an expected field position model for each punt return option (return, fair catch, let it bounce), a muff probability model for fair catches and returns, and a fumble probability model for returns. Predictions were then converted into expected points (making them expected expected points) and the expectations for the three return options were compared to determine the optimal decision. The equations used to calculated expected EP were as follows:\n- **Expected Let It Bounce EP**: $E$[Let It Bounce EP$]$\n- **Expected Fair Catch EP**: $(1 - (1 - 0.674) * Pr[$Muff$]) * E[$Fair Catch EP$] + (1 - 0.674) * Pr[$Muff$] * E[$Muff EP$]$\n- **Expected Return EP**: $(1 - (1 - 0.674) * Pr[$Fumble$]) * E[$Return EP$] +  (1 - 0.674) * Pr[$Fumble$] * E[$Fumble EP$]$","metadata":{}},{"cell_type":"markdown","source":"The number 0.674 in these equations represents the probability that a muffed or fumbled ball will be recovered by the receiving team, as explained by [Football Perspective’s Chase Stuart](https://www.footballperspective.com/more-on-fumbling-and-recovery-rates-defense-and-special-teams/). For simplicity, we assume this probability to be constant and we assume that the probability of a fumble is uniform throughout a play. To predict the likelihood of a fumble or muff on a given play, we implemented 10,000-tree random forest classification models with 10-fold cross validation. \n\nTo model the end yard line on returns and bounces, we generated conditional density estimates using 10,000-tree random forests within the [RFCDE package in R](https://arxiv.org/abs/1804.05753). We implemented 10-fold cross validation to assess model fit and generate out-of-fold predictions for our training data. For both our out-of-fold predictions and our predictions for plays where we did not observe the modeled decision, we then mapped expected points to the density at each yard line and computed the expected value of the predictions in terms of EP. For the fair catch model, we simply mapped the yard line of a potential fair catch to EP.\n\nOnce we have the expected EP for each option, DEEPS is calculated as the difference between the maximum expected EP of the three options and the EP of the option chosen by the returner. Thus, optimal decision-making would result in zero decision-expected expected points surrendered.\n\n## 2.3 Example\nOn the following play from the 2020 season, Rams returner Cooper Kupp lined up to receive a punt from Giants punter Riley Dixon. Our model recommended a fair catch at 1.85 expected EP; however, Kupp elected to return the punt, a decision which we expected at the apex of the punt to yield 1.20 EP on average. Thus, his DEEPS for the play was 1.85 expected EP - 1.20 expected EP, or 0.65 Decision Expected EP Surrendered.\n\n![image](https://imgur.com/JbPlvYN.png)\n![image](https://i.imgur.com/9Ks2sQn.png)\n\n## 2.4 Feature Importance\nAnalyzing feature importance from our muff and models models, we find that having defenders close by greatly increases the chances of miscues and lowers the expected yards from a return. In addition to the distance from nearest defenders, punt velocity and launch angle also show up as important predictors of miscues. This means that punts with greater hang time tend to lean toward fair catch being the optimal decision over a return, both because gunners have more time to get downfield and because the ball will be traveling faster vertically when it arrives, making it harder to catch. Meanwhile, on returns, we see that location of the ball and returner are evidently important variables for predicting yardage, as well as field control and distance from nearest defenders (which are evidently related).\n\n## 2.5 Decision Insights\nOf the 4,869 punts analyzed with DEEPS, we find return to be the optimal decision 74.9% of the time, compared to 20.7% for fair catch and 4.4% for letting the ball bounce. Most of the time, the upside of a return seems to outweigh the corresponding muff and fumble risk. Now, these recommendations do not account for injury risk on returns, and hence likely lean too optimistic for returning punts, but directionally they point to the fact that muff rates are low and returns almost always improve field position. On the flip side, letting the ball bounce in hopes of a touchback appears not to be a worthwhile gamble; most of the time, a returner should usually try to field punts and limit how close the ball can get to the goal line.\n\nBased on our model insights, the following is a basic decision guide for returners:\n- *Default to returning the punt*\n- *Do not field the ball if you expect it to land within the ~5 yard line*\n- *If a defender can reach you before the ball arrives (usually on punts of shorter distance or longer hang time), call for a fair catch*\n","metadata":{}},{"cell_type":"markdown","source":"# 3. DEEPS Results\n## 3.1 Player Rankings\nThe following table shows the players who surrendered the fewest expected EP per play over the 2018-2020 seasons, led by Chiefs receiver Tyreek Hill.\n![image](https://imgur.com/wjtjwHT.png)\n## 3.2 Metric Stability\nWe measured the stability of DEEPS by looking at the correlation between the metric in the year N and the year N + 1. We found that there is a weak negative year-to-year correlation of -0.125. It is, however, important to note that there is a relationship between the two variables and could be indicative of decreased decision-making as a returner ages.\n","metadata":{}},{"cell_type":"markdown","source":"# 4. PREPA Methodology\n## 4.1 Model Explanation\nTo evaluate post-decision return skill, we built a continuous-time model for punt return yardage inspired by Yurko et al.’s [Going Deep](https://arxiv.org/abs/1906.01760) paper. At every frame within each play, we predict the remaining return yardage on the play; aggregating over all frames, we evaluate player performance by taking the average difference per frame between the predicted and actual area under the curves with frame on the x-axis and EP on the y-axis. To generate frame-by-frame predictions, we trained a Least Absolute Shrinkage and Selection Operator (LASSO) model with 10-fold cross validation to predict the logarithm of the remaining return yardage. We used similar features to those mentioned for DEEPS as our covariates. We intentionally did not choose to model the change in frame-to-frame yardage as our response variable since it lacks predictive practicality and can be explained incredibly well by simple tracking features like speed, acceleration, and direction of movement.\n\nOnce we have the predicted remaining return yardage for each frame, we can map the implied final yard line to expected points to get PREPA – Punt Return Expected Points Added. The example below shows how our prediction changes over the course of a punt return:\n","metadata":{"execution":{"iopub.status.busy":"2022-01-06T21:19:58.877689Z","iopub.execute_input":"2022-01-06T21:19:58.883397Z","iopub.status.idle":"2022-01-06T21:19:59.035198Z"}}},{"cell_type":"markdown","source":"# 4.2 Example\nReturning (no pun intended) to the Cooper Kupp example from 2.3, we see that Kupp elected to return the punt despite our recommendation to call for a fair catch. We recommended a decision to fair catch due to the higher probability of a fumble or penalty. When the ball actually lands, we can see that Kupp has more space to run than originally predicted. In addition, his teammates were able to successfully block, creating more space for Kupp to return the punt for 7 yards. \nWe show each team’s field control, adapted from [Adam Sonty’s notebook](https://www.kaggle.com/adamsonty/nfl-big-data-bowl-a-basic-field-control-model). \n\n![image](https://i.imgur.com/wnIBIK7.gif)\n\n# 4.3 Feature Importance Insights\n![image](https://imgur.com/gZml1De.png)\n\nReturner speed and distance from the nearest three defenders are important predictors of return yardage, as one would expect. Moving fast in open space is the best way to gain yards. Interestingly, yards to go on the play had a substantive impact on return yardage predictions, perhaps due to return teams focusing less on blocking on short-yardage plays to avoid falling victim to fake punts. The relevance of player speeds and distances from one another clearly shows the value of tracking data for modeling play outcomes in the NFL. ","metadata":{}},{"cell_type":"markdown","source":"# 5. PREPA Results\n## 5.1 Player Rankings\nThe following table shows the players who led the NFL in PREPA per play over the 2018-2020 seasons, led by Ravens receiver James Proche II and Cowboys receiver CeeDee Lamb:\n\n![image](https://imgur.com/HUB6Ofm.png)\n\nDue to low sample size, we felt it would be improper to measure the stability of PREPA. There are not enough punt returns per player each season, especially across multiple consecutive seasons, to get a meaningful sense of the year-to-year correlation of an individual’s PREPA. While we believe PREPA is a more holistic measure of punt return ability than other measures of return skill, it certainly remains more results-oriented than DEEPS. \n","metadata":{}},{"cell_type":"markdown","source":"# 6. Final Results & Takeaways\n## 6.1 Player Evaluation\n![image](https://imgur.com/BZD6iHG.png)\nThe plot above shows how players with at least 20 returns from 2018-2020 performed on both DEEPS and PREPA and separates players into four quadrants based on their combination of metric scores. Kalif Raymond and James Proche II appear to exhibit the strongest combined decision-making and punt returning skills, while players in the top-left quadrant may need to be re-evaluated for the returning role by their teams. We find no strong correlation (r = -0.096) between the two metrics, meaning these skills are relatively independent and our metrics are indeed measuring two distinct returner abilities. \n\n## 6.2 Discussion & Conclusions\nWe have created a pair of metrics for evaluating the decision-making and punt returning skills of NFL players using tracking data from the 2018-2020 seasons. Through this process-oriented approach, we are now able to quantify and rank players based on these abilities, as well as assess important factors for successful returns. If returners were to follow our decision guide and make optimal decisions according to DEEPS, they would maximize the value they bring to their team on punts. Teams can use these metrics to continually assess the performance of their designated returners, evaluate potential changes at the position, and measure the value of a player’s punt return skills for future compensation.\n\n## 6.3 Limitations & Future Research\nIn our analysis, we simplify the definition of a catchable punt, excluding some out of bounds punts that likely bounced inbounds at a catchable spot. These assumptions could be eliminated in future work by modeling the distribution of fumble outcomes and building a physics-based model to measure the catchability of a punt. Additionally, there are a number of considerations for punt return decisions that our model can’t quantify. Injury risk is a big driver of fair catch decisions, but there is no expected points value for that. Also, while expected points is well suited for valuing play outcomes in aggregate, it may not capture the objective functions of teams late in games; as returns are high-variance, a losing team may choose to return more punts than a team trying to secure a victory. An interesting future approach to this problem would be mapping play outcomes to win probability rather than EP and seeing how the variance of each decision affects the model’s prescription. In future models, we would ideally like to model all aspects of the play as distributions rather than point estimates, implementing RFCDE or another density estimator to better capture the uncertainty in play outcomes. \n","metadata":{}},{"cell_type":"markdown","source":"# Appendix\nOur project code can be found on [Github](https://github.com/husarah/big-data-bowl-2022).\n\nContact us!\n\nZach Bradlow: [Twitter](https://twitter.com/zachbradlow) | [Email ](mailto:zbradlow@wharton.upenn.edu)| [Linkedin](https://www.linkedin.com/in/zach-bradlow/)\n\nZach Drapkin: [Twitter](https://twitter.com/ZachDrapkin) | [Email](mailto:zdrapkin@wharton.upenn.edu) | [LinkedIn](https://www.linkedin.com/in/zach-drapkin-3614a8174)\n\nSarah Hu: [Twitter](https://twitter.com/sarahhuuuu) | [Email ](mailto:husarah@wharton.upenn.edu)| [LinkedIn](https://www.linkedin.com/in/sarah-hu1/)\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}