{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Decision-Making in Punt Returns\n\n## Analyzing Decision-Making on Returnable Punts\n\nBy Tabor Alemu and Vinesh Kannan\n","metadata":{}},{"cell_type":"markdown","source":"# Introduction\n\nWhen a punt returner sees that a punt is returnable, they have three choices:\n\n- Return: Catch the ball and try to advance it for better field position.\n- Fair Catch: Make the catch and accept that field position.\n- Bail: Let the ball land and hope for good field position.\n\nPut yourself in the shoes of Corey Clement, punt returner for the Philadelphia Eagles. Which action would you choose on this punt?\n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/Clement%20Return%20Decision.gif?raw=true\" />\n\nWhile the ball is still in the air, Clement signals to his teammate to stay away, so that they do not accidentally touch the ball and give Trenton Cannon, the nearby gunner for the New York Jets, a chance to recover.\n\nBut after the ball takes a bounce in favor of the punting team, threatening to pin the Eagles back even deeper, Clement changes his decision.\n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/Clement%20Return%20Muff.gif?raw=true\" />\n\nThe result is disaster. Clement fails to secure the ball and Cannon recovers it, setting up the Jets’ only touchdown of the game. Even the announcer remarks that it was a bad decision, but it is easier to say that as an observer than from in the heat of the play.\n\nWe use the NFL’s Next Gen Stats (NGS) tracking data to form a better picture of what the returner faced on the field to analyze punt returner decision-making.\n\nWe developed a model that uses the pre-arrival frame to predict whether a returnable punt will net the return team a loss or zero yards from the end of the kick, a measure of return difficulty. This Jets-Eagles play was in the hold-out set, meaning that our model was never trained or tuned on it. Our model predicted a 76% chance of this play resulting in a loss or no gain from the end of the kick. This indicates not only that Clement's original decision to stay away was correct, but that his second decision to try and limit the losses was a big risk.\n\nOur analysis provides:\n\n- Heatmaps to help punting teams decide where to punt.\n- Average return yards plots to help receiving teams decide whether to return or bail.\n- Ranking of punt returners who turn the most difficult punts into gains.\n- Ranking of returning teams who turn the most difficult punts into gains.\n","metadata":{}},{"cell_type":"markdown","source":"# Part 1: Tendencies and Trade-Offs\n\nThroughout this analysis, we focus only on returnable punts, which we define as a punt that lands or is caught in the field, including in the endzone. More details on how we filtered returnable punts can be found in the methodology section.\n\nFirst, we formed a baseline for what parts of the field punt returners generally return, fair catch, or bail by using tracking data from 2018 to 2020 returnable punts \n\n> Kicking teams can use this information to decide where to punt if they want to force a returner to fair catch or bail.\n\nIn the following plots, the receiving team’s end zone is on the left. Kicking team yard line is relative to the kick team’s  endzone, for example, a kicking yard line of 60 means the receiving team’s 40 yard line.\n\nWe constructed multiple “heatmap” models depicting the most telling parts of the field that would yield better results for punt specialists. By balancing averages of previous punt results, we have made diagrams that will tell returners where on the field they should call a fair catch, bail in the hopes for a touchback, or return the ball. \n\n## Figures 1-3. Heatmap of returnable punt decisions based on field position.\n\nThis figure separates three punt return categories (fair catch, bail/downed, and return) and is plotted to show possible destinations for which these categories occur at when they first hit the ground. These plots are subject to where the line of scrimmage is located prior to the punt, separated by 10 yard bucket increments. \n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/Return_plot.png?raw=true\" />\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/fc_plot.png?raw=true\" />\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/Bail_plot.png?raw=true\" />\n\nThis can be handy for punters to understand ball placement when neutralizing dangerous and explosive punt returns. Therefore forcing desired outcomes from certain punts. The model above shows destinations of punts snapped between the punting team’s own 30 to 40 yard line resulting in fair catches throughout the 3 season dataset. Tendencies can be described, telling us that the “sweet spot” for forcing fair catches is between the 15 to 25 yardline. \n\n## Figure 4. Average return yards gained for returnable punt decisions based on field position.\n\nNext, we analyzed the average return yards gained for each type of decision, depending on field position, based on returnable punts from the 2018 to 2020 season.\n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/mix_plot.png?raw=true\" />\n\n\nThe second model also looks at the same three categories and examines different positional situations to obtain the highest net yards. While examining, the main goal is to get the most net yards from when the punt first reaches ground level to when the offense starts their drive. The model is also independently evaluated based on the line of scrimmage prior to the punt. This model is extremely useful for punt returners and special team coordinators in understanding what decisions should be made before fielding a punt. For example, if a team is punting to a returner from their own 35 yardline, this model tells specialists to almost always attempt to return the ball when comfortable and outside their own 5 yard line. On the contrary, if the returner does not feel comfortable to return a punt due to pressure from the punt coverage team, the model tells coaches that calling for a fair catch inside of the returner’s own 14 yardline will yield a lower average net yardage than bailing from catching the ball with the intent for the ball to be a touchback. This gives returners a step up on the competition, using analytics to get the best field position for his offense.\n","metadata":{}},{"cell_type":"markdown","source":"# Part 2: Ranking Difficult Returns\n\nTo evaluate punt returner decision-making, we developed an estimate of return difficulty  before the outcome is known. More specific details can be found in the methodology section.\n\nFor each punt, we identify the decision frame, the point in the tracking data one second before the ball lands. We assume that this is the last point when the returner can use information about their surroundings to decide whether to return, fair catch, or bail. Returners may decide earlier (or later, at their peril), but we will use this frame to evaluate the difficulty.\n\nWe train a machine learning model to classify returnable punts that result in zero or negative return yards gained by the returning team, from the end of the kick. This model estimates which punts result in losses for the returning team or that punt returners choose not to return. Our justification for combining both loss and neutral plays into one target is that both outcomes represent situations where the expected value of returning is low.\n\nWe rank punt returners based on how often they face, return, and gain from returnable punts that our model flags as difficult to return for a gain.\n\n## Figures 5-6. Top Returners by Difficulty\n\nHere are the top returners, ranked based on how often they turn a difficult punt into a gain.\n\n<img style=\"max-width: 700px;\" src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/player_ranks.png?raw=true\" />\n\nThis plot compares how often returners return difficult returnable punts with how often those returns gain yards from the end of the kick.\n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/player_plot.png?raw=true\" />\n\n## Figures 7-8. Top Teams by Difficulty\n\nHere are the top teams, ranked based on how often their returners turn a difficult punt into a gain.\n\nOriginally, the Raiders were the top team, but after merging their Las Vegas (LV) and Oakland (OAK) records, the Cleveland Browns (CLE) edged them out for the top spot.\n\n<img style=\"max-width: 700px;\" src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/team_ranks.png?raw=true\" />\n\nThis plot compares how often teams return difficult returnable punts with how often those returns gain yards from the end of the kick.\n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/team_plot.png?raw=true\" />","metadata":{}},{"cell_type":"markdown","source":"# Code\n\n- [Returnable Punt Data Processing](https://www.kaggle.com/vingkan/process-punt-return-decision-data)\n- [Model Training and Evaluation](https://www.kaggle.com/vingkan/model-training-returns-for-loss)\n- [Difficult Returnable Punts Analytics](https://www.kaggle.com/vingkan/analytics-gutsy-returners)\n- [Tendencies and Trade-Offs Analytics](https://www.kaggle.com/talemu/analytics-returns-for-loss)\n\nThank you for organizing this challenge! Please see the appendix for methodology details and model scores.","metadata":{}},{"cell_type":"markdown","source":"# Appendix: Methodology\n\nThe first half of our analysis uses the processed tracking data to provide a baseline for what parts of the field punt returner can expect to return, fair catch, or bail.\n\nThe second half of our analysis predicts which punts will be difficult to return for a gain and applies the predictions to rank players and punts.\n\n## Data Processing\n\nPrior to our analysis, we applied the following data processing steps:\n\n1. Reorient tracking data so that the receiving team always advances from the left of the screen to the right.\n1. Filter plays to only include punts that are returnable, with no penalties.\n1. Assign a decision based on the result of the play: return, fair catch, or bail.\n\n## Machine Learning\n\nFor each returnable punt:\n\nStep 1: Identify the decision frame.\n\n- The decision frame is the latest point by which the returner has to decide whether or not to return the ball, that comes before the ball has actually arrived.\n- Identify the frame of the first returnable event, the first time the punt lands or is caught in the field of play.\n- Move back one second (or ten frames) from that frame to get the decision frame.\n- Only train the model on tracking data from the decision frame for model inputs.\n\nStep 2: Split the punts into cross-validation sets.\n\n- Keep 50% of the data for training, keep 25% for tuning (validation), and keep 25% as a hold out set (test).\n- Only one frame per play is used, the decision frame.\n- Any play can be in any split. We decided not to stratify by season because the rates of different decisions were comparable across seasons.\n\nStep 3: Create a target variable for model training.\n\n- Create a target variable isZeroOrLoss that is false if the receiving team gained yards from the end of the kick, and true if they lost yards or gained no yards.\n- The majority of returnable punts result in zero gain, including many fair catches.\n- Our justification for combining fair catches with returns for zero or loss is returners fair catch in a situation where they do not want to return or let the ball land, so both fair catches and negative returns represent situations where returners face difficult punts and the expected value of returning is low.\n\nStep 4: Create features based on tracking data to use as model inputs.\n\nFor all input features, we use the ball location from the decision frame, when it is still in the air.\n\nIt could be valid to use the ball’s eventual landing spot, under the assumption that the returner and other players are able to estimate where the ball will land, but we choose to use the decision frame location for a more conservative estimate.\n\nWe derive these ten features from the tracking data, during the decision frame:\n\n- Ball Yard Line: Location of the ball, in yards from the receiving team goal line.\n- Closest Defender Distance: Distance in yards from the ball to the closest member of the kicking team/\n- Defenders Within Radius: Number of members of the kicking team within a two yard radius of the ball.\n- Blockers Within Radius: Number of members of the receiving team besides the returner within a five yard radius of the ball and who are ahead of the returner.\n- Closest Defender Speed Upfield: The component of the closest defender’s speed that goes from goal line to goal line, negative if heading towards the receiving team goal line, in yards per second.\n- Closest Defender Speed Lateral: The component of the closest defender’s speed that goes from sideline to sideline, always positive, in yards per second.\n- Distance to Sideline: Distance in yards from the ball to the closest sideline, with a value of zero if the ball is out of bounds.\n- Is Inside Own Endzone: Value of one if the ball is over the receiving team endzone during the decision frame, zero otherwise.\n- Is Inside Own 10: Value of one if the ball is inside the receiving team 10 yard line during the decision frame, zero otherwise.\n- Is Inside Own 20: Value of one if the ball is inside the receiving team 20 yard line during the decision frame, zero otherwise.\n\nStep 5: Train and evaluate models.\n\n- We employed three types of models: logistic regression, random forest classifier, and support vector machine classifier.\n- All models are binary classifiers that support a prediction probability score.\n- We scored and compared models on the unseen validation data.\n- We evaluated models using precision, recall, F1-score, accuracy, and area under the receiver operating characteristic (ROC) curve.\n- We did not perform any automatic hyperparameter tuning.\n- The random forest model performed best, but we chose the class-balanced logistic regression model because it had strong, similar precision and recall scores, as well as comparable ordering power (by area under ROC curve) to the random forest. We wanted to prioritize a high-bias model over a high-variance model, and a model whose results would be more explainable to coaches and players. The visual simplity of the heatmap of the logistic regression model compared to the random forest model bears this out.\n\n## Figure 9: Final Model Performance\n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/model_scores.png?raw=true\" />\n\nFinal model scores on validation data:\n\n```\nAccuracy  = 0.729\nPrecision = 0.776\nRecall    = 0.768\nF1-Score  = 0.772\nROC AUC   = 0.720\n```\n\n## Figure 10. Heatmap for Clement Play\n\nHeat map of where the ball could have landed on the Clement return play and the predicted probability of a loss or no gain (pink) vs a gain (green).\n\n<img src=\"https://github.com/vingkan/nfl-big-data-bowl-2022-supplementals/blob/main/clement_map.png?raw=true\" />","metadata":{}}]}