{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 2022 NFL Big Data Bowl\n## Determining the Optimal Decision Making for Punt Returners\n\nAuthor: Tej Seth\n\nCode: https://github.com/tejseth/bdb-optimal-returner-decision","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Introduction\n\nWhen the punter boots a punt away and the ball is hanging in the air, the punt returner usually has a few seconds to make up their decision on what they want to do with that punt. They could:\n1. **Return It**: The punt returner can catch the ball in the air and attempt to advance it before being tackled by the punting team.\n2. **Fair Catch**: The punt returner can wave their hand as the ball is in the air, catch it and their team will start with the ball on offense at the yardline in which they catch the ball at.\n3. **Bounce**: The returner can also choose not to catch the punt out of the air and instead choose to let it bounce\n\nThese three (3) decisions can be seen in Figure 1.\n\n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/decisions.png)\n\nOne of the best things a punt returner can do for their team is make the right decision so that they are maximizing their team's field position. I set out to create models that would tell the punt returner what the optimal decision they could make is. To do that, we can create three different metrics that are all made at the frame right before the punt arrives at the returner: The first is the expected return yardline ($ER_y$), the second is the expected fair catch yardline ($FC_y$) and the third is the expected yardline if the returner decides to let it bounce ($B_y$). The maximum of ($ER_y$), ($FC_y$) and ($B_y$) will then be taken to see what the \"optimal\" decision is for a punt returner on any given play. \n\n","metadata":{}},{"cell_type":"markdown","source":"## Early Data Analysis\n\nThe main data source used that was used in this project was the tracking data provided by the NFL which tracks where each player is on the field every tenth of a second (known as a \"frame). Another dataset that was also provided was the Plays spreadsheet which had an overarching descrption for what happened on each play. Inside that dataset, `specialTeamsPlayType` was sent to be only punts as field goals and kickoffs were also in the dataset but not needed for the scope of this project. Using the ID's from just the punts, the tracking dataset was filtered to just included those ID's. An additional dataset was provided by Pro Football Focus and had critical information such as direction of the punt, number of gunners, hang time and more. \n\nFigure 2 shows some early data analysis of the punts that *were* returned and what the net return was for those punts when factoring in penalty yards\n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/net-return.png)","metadata":{}},{"cell_type":"markdown","source":"## Building Models\n\nThe first model we could build was an expected return yards model which we could then add onto where the returner was catching the punt to give an expected yardline if returned ($ER_y$). The model was trained using 50% of the punts that were actually returned in the dataset to avoid overfitting. Since this model will be applied to all punts, those returned and those not returned, it should be noted that there could be some selection bias as punt returners are more likely to return punts in which they feel like they have space to return and more likely to fair catch punts in which they feel like they don't have space. A random forest technique was taken so that the splits at the nodes for the individual decision trees could hopefully account for that. The feature importance for expected return yards can be seen in Figure 3:\n* `Distance_i`: The i stands for what order the player on the punting team is distance-wise from the returner on the frame before the ball is about to arrive with distance_1 being the player closest to the returner and distance_11 being the player furtherest from the returner.\n* `returner_o`, `returner_s`, `returner_x`, `returner_y`: the returner's orientation, speed, x coordinate and y coordinate at the frame before the field the punt.\n* `hangTime`: The amount of time the punt was in the air, charted by PFF\n* `num_vises`, `num_gunners`, `num_rushers`: The number of vises blocking the gunners for the returner, the number of gunners trying to tackle the punt returner and the number of rushers trying to block the punt.\n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/vip-return.png)\n\nThe next model that can be made is the expected yardline of where the ball will be placed if the returner decides not to field the punt and to let it bounce instead. There are two components to what could happen here: the ball could be downed at a certain yardline by either bouncing out of bounds or being touched by the punting team *or* the ball bounces into the endzone for a touchback and is given at the 20 yardline to the offense. An expected touchback model can be built with the feature importance seen in Figure 4. \n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/touchback-exp.png)\n\nA similar model can made to model what yardline the ball will end up at *if* there isn't a touchback with that feature importance shown in Figure 5.\n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/bounce-exp.png)\n\nWith both of those models made, the expected yardline if the returner lets it bounce can be made using the formula below.\n\n$(P(touchback)*20 + (1-(P(touchback))*(E(yardline))$\n\nA model for the expected yardline for if the returner fair catches a punt isn't needed because it would just be the x coordinate of the football on the frame right before it arrives at the returner. \n\n","metadata":{}},{"cell_type":"markdown","source":"## Player Evaluation\n\nBy taking the maximum of the expected yardline if returned ($ER_y$), expected yardline if bounced ($B_y$) and yardline of the fair catch ($FC_y$), we can get the optimal decision for the returner. It should also be noted that the 3% chance of a muff is baked into the decision as that factors into ($ER_y$) and ($FC_y$). Figure 6 shows a matrix of what the actual decisions have been on the x-axis and the optimal decisions on the y-axis. On average, punt returners make the \"optimal\" decision 49.6% of the time, which is better than the 33% we would expect if they just randomly picked what they were going to do. \n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/matrix.png)\n\nFigure 6 shows us that punt returners aren't returning as much as they should be according to the models. (Again it should be mentioned that there might be survivorship and selection bias given that the expected return yards model was trained based on so it might be thinking that the expected return yards for all punts is really high). \n\nFigure 7 shows the top 10 punt returners at making the optimal decision (with at least 25 punt returns).\n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/top_10_gtt.png)\n\nAs seen, Jakeem Grant tops the board by making the optimal decision 64% of the time. He returned it when he should have 65% of the time, fair caught it when he should have 0% of the time and let it bounce when he should have 100% of the time. \n\nOn the flip side, Figure 8 shows what bottom 10 punt returners in optimal decision rate. \n![](https://raw.githubusercontent.com/tejseth/bdb-optimal-returner-decision/master/images/bottom_10_gtt.png)","metadata":{}},{"cell_type":"markdown","source":"## Additional Notes\n\n**Usefulness to Teams**: The goal of this project was to create something that could be helpful to NFL teams as they attempt to gain edges to win football games. We believe that this could provide a framework for coaches to show punt returners what their actual decisions were on certain punts and why or why not that was the optimal decision.\n\n**Future Additions** This project entailed the decision happening on the frame right before the ball arrived to the punt returner but that could be considered too late of a time to make a decision. Because of that, in the future we can build on top of this and have the decision happen either when the punt leaves the punter's foot or ~2-3 seconds before it arrives.\n\n**Find Me on Twitter**\n* Tej Seth: @tejfbanalytics\n\nThank you to Michael Lopez, Thomas Bliss and the rest of the team for giving me the opportunity to compete in this competition!","metadata":{}}]}