---
title: "An Evaluation of NFL Special Teams Players"
author: 'Connor Nickol, email: can2hr@virginia.edu'
date: "Big Data Bowl 2022"
output: 
  html_document:
    number_sections: true
    toc: true
    
---

# Introduction

As a student equipment manager the last three years with the UVA Football Team under Coach Bronco Mendenhall, I've seen firsthand the time, effort, and preparation required to have quality special teams play at a high level of football. With the NFL moving towards safer kickoffs and rules being proposed to eliminate them entirely, I chose to focus on punts and punt returns. In this paper I will discuss how I went about the task of evaluating punt returner performance, gunner performance, and the optimal number of rushers for a punt. 
 
# Data Preparation

The data for the NFL Big Data Bowl on Kaggle contains game and play data, player information data and player specific tracking data from AWS Next Gen Stats throughout all 17 weeks of the 2018, 2019, and 2020 NFL seasons^[https://www.kaggle.com/c/nfl-big-data-bowl-2022]. 

For my evaluation of punt returners and gunners I used the 2018 and 2019 seasons as training data in my regressions and machine learning algorithms. The 2020 season was used as testing data. I filtered the plays data set to only include punts and then split it by season to be joined with that season's tracking data. Further preparation will be explained in the specific analysis sections. 

# Expected Yards Per Punt Return

On any given punt return, is it possible to tell if the returner did a poor or good job? To try and answer this question I decided to run a regression to predict how many yards a returner was expected to get versus how many he actually got. Using an animation inspired and adapted from former fellow UVA student Ella Summer's 2021 Big Data Bowl Submission^[https://www.kaggle.com/es4vr17/random-player-effects-defensive-backs/] I show an example punt return play. On this play from Week 1 of the 2020 NFL season, Hunter Renfrow (Jersery number 13) receives the punt and returns it for 27 yards. My model predicted that the expected return on this play would be 8 yards. Renfrow was able to exceed that by 19 yards, setting the Raiders up in excellent field position against the Panthers.   

```{r, include=FALSE}
library(tidyverse)
library(gganimate)
library(ggplot2)
library(cowplot)
library(repr)
library (formattable)

games<-read.csv("../input/nfl-big-data-bowl-2022/games.csv")

players<-read.csv("../input/nfl-big-data-bowl-2022/players.csv")

plays<-read.csv("../input/nfl-big-data-bowl-2022/plays.csv")

track_20<-read.csv("../input/nfl-big-data-bowl-2022/tracking2020.csv")

example_play<-plays%>%filter(gameId==2020091303 & playId == 2143)%>%select(gameId,playId,playDescription)

example_play <- inner_join(example_play,
                           games,
                           by = c("gameId" = "gameId"))

example_play <- inner_join(example_play,
                           track_20,
                           by = c("gameId" = "gameId",
                                  "playId" = "playId"))

cols_fill <- c("#999999", "#663300", "#3d85c6")
cols_col <- c("#000000", "#663300", "#000000")

xmin <- 0
xmax <- 160/3
hash.right <- 38.35
hash.left <- 12
hash.width <- 3.3

plot_title <- str_trim(gsub("\\s*\\([^\\)]+\\)","",as.character(example_play$playDescription[1])))

ymin <- 0
ymax <- 120

#hash marks
df.hash <- expand.grid(x = c(0, 23.36667, 29.96667, xmax), y = (10:110))
df.hash <- df.hash %>% filter(!(floor(y %% 5) == 0))
df.hash <- df.hash %>% filter(y < ymax, y > ymin)


#plotting
p<- ggplot() +
  
  #setting size and color parameters
  scale_size_manual(values = c(6, 4, 6), guide = FALSE) + 
  scale_shape_manual(values = c(21, 16, 21), guide = FALSE) +
  scale_fill_manual(values = cols_fill, guide = FALSE) + 
  scale_colour_manual(values = cols_col, guide = FALSE) +
  
  #adding hash marks
  annotate("text", x = df.hash$x[df.hash$x < 55/2], 
           y = df.hash$y[df.hash$x < 55/2], label = "_", hjust = 0, vjust = -0.2) + 
  annotate("text", x = df.hash$x[df.hash$x > 55/2], 
           y = df.hash$y[df.hash$x > 55/2], label = "_", hjust = 1, vjust = -0.2) + 
  
  #adding yard lines
  annotate("segment", x = xmin, 
           y = seq(max(10, ymin), min(ymax, 110), by = 5), 
           xend =  xmax, 
           yend = seq(max(10, ymin), min(ymax, 110), by = 5)) + 
  
  #adding field yardline text
  annotate("text", x = rep(hash.left, 11), y = seq(10, 110, by = 10), 
           label = c("G   ", seq(10, 50, by = 10), rev(seq(10, 40, by = 10)), "   G"), 
           angle = 270, size = 4) + 
  annotate("text", x = rep((xmax - hash.left), 11), y = seq(10, 110, by = 10), 
           label = c("   G", seq(10, 50, by = 10), rev(seq(10, 40, by = 10)), "G   "), 
           angle = 90, size = 4) + 
  
  #adding field exterior
  annotate("segment", x = c(xmin, xmin, xmax, xmax), 
           y = c(ymin, ymax, ymax, ymin), 
           xend = c(xmin, xmax, xmax, xmin), 
           yend = c(ymax, ymax, ymin, ymin), colour = "black") +
  
  #adding field interior 
  annotate("rect", xmin = xmin, xmax = xmax, ymin = ymin, ymax = ymax, 
           fill = "dark green", alpha = 0.2) + 
  
  #adding players
  geom_point(data = example_play, aes(x = (xmax-y),
                                      y = x, 
                                      shape = team,
                                      fill = team,
                                      group = nflId,
                                      size = team,
                                      colour = team), 
             alpha = 0.7) +  
  
  #adding jersey numbers
  geom_text(data = example_play, aes(x = (xmax-y), y = x, label = jerseyNumber), colour = "white", 
            vjust = 0.36, size = 3.5) + 
  
  #applying plot limits
  ylim(ymin, ymax) + 
  coord_fixed() +
  
  #applying theme
  theme_nothing() +
  
  #titling plot with play description
  labs(title = plot_title) +
  
  #setting animation parameters
  transition_time(frameId)  +
  ease_aes('linear')

```

```{r, warning=FALSE, echo=FALSE}
animate(p)
```

## Regression Decisions

In order to calculate expected yards for a given return, I made two critical decisions in setting up my variables for regression. I chose to examine player positioning at the time of the punt returner catching the ball. I also chose to only consider punt returns in which the return yardage was less than 33 yards. I made my decision based off of the average return of 9.5 yards and a standard deviation of 11.0 yards. Because of the infrequency of punt return touchdowns and longer returns in general, I felt it would be better to try and evaluate the more normal return.  I took any players from the 2020 season who did not have at least 5 return attempts and instead placed these observations in the 2018-2019 training data.

## Regression Variables
I first ran a linear regression with the given variable kickReturnLength as my dependent variable.

I considered the returning team to be on offense since I was evaluating the return aspect of a given punt.

My independent variables were the provided variables

* kickLength (Kick length in air of punt)
* operationTime (Timing from snap to kick on punt plays in seconds)
* hangTime (Hangtime of player's punt or kickoff attempt in seconds)

and then my created variables 

* def_1 (Euclidean distance of the closest defender to the punt returner)
* def_2 (Distance of second defender)
* def_3 (Distance of third defender)
* off_1 (Euclidean distance of the closest blocker to the punt returner)
* off_2 (Distance of second blocker)
* off_3 (Distance of third blocker)
* dist_sdl (Euclidean distance of the punt returner to the closest side line)
* dist_goall (Euclidean distance of the punt returner to the closest goal line)
* def_1_blkd (Dummy variable: 1 if the closest defender is actively being blocked, 0 otherwise)
* def_2_blkd (1 if second defender is blocked)
* def_3_blkd (1 if third defender is blocked)

## Variable Creation
 
In order to define a player as being actively blocked, they had to have an offensive player within 1 yard of them, the offensive player also needed to be between the defender and the returner, and the offensive player needed to be looking at the defensive player. 

This was done using the provided player tracking variables
* x (Player position along the long axis of the field, 0 - 120 yards)
* y (Player position along the short axis of the field, 0 - 53.3 yards)
* o (Player orientation (deg), 0 - 360 degrees)


```{r, echo = F,fig.cap='Figure 1. Coverage Player Angle Calculation',fig.align='center'}
library(knitr)
include_graphics("https://i.imgur.com/lHzJjXC.jpg")
```

The image above provides an example of how the blocked dummy variables were calculated. First I subtracted the defensive player's x and y variables from the offensive players x and y to figure out what "quadrant" the two players were in. Using that knowledge, I was able to use trigonometry to calculate the angle of the defensive player from the origin in comparison to the offensive player's orientation with either invere sin or inverse cosine. I then I added and subtracted 60 degrees to the orientation variable of the offensive player and checked if the angle of the defensive player was within that range. 

## Regression Methods

I first began by running a linear regression and found the variables kickLength, hangTime, def_1, and def_3 to be significant. I then calculated my test mean squared error for my linear regression by calculating the mean of the actual values minus the predicted values squared. I found the test MSE to be 38.6. 

I decided to move to tree based regression methods to see if I may be able to obtain better results. I reran the regression using Recursive Binary Splitting, Bagging, Random Forest, and Boosting and chose Random Forest as it gave me the lowest mean squared error of 38.4. All variables were left in for my tree based methods. I added my Random Forest predictions to the original data in order to compare with the actual return values. For each player I averaged their actual return yards minus their predicted return yards. 

## Results Table

```{r,echo=FALSE}

customGreen0 = "#DeF7E9"
 
customGreen = "#71CA97"
 
customRed = "#ff7f7f"

returner<-read.csv("../input/data-for-tables/returner_data.csv")
returner<-returner%>%mutate("Average of Actual-Predicted (Yards)"=avg,"Name"=displayName,"Attempts"=n)%>%select("Name","Attempts","Average of Actual-Predicted (Yards)")
 
improvement_formatter <-
  formatter("span",
            style = x ~ style(
              font.weight = "bold",
              color = ifelse(x > 0, customGreen, ifelse(x < 0, customRed, "black"))))
 
 
formattable(returner, align =c("c","c","c"), list(
  `Name` = formatter("span",
                     style = ~ style(color = "grey",font.weight = "bold"),
                     x ~ icontext(ifelse(x == "Hunter Renfrow", "star", ""), x)),
  `Attempts`= color_tile(customGreen0, customGreen),
  `Average of Actual-Predicted (Yards)` = improvement_formatter
))
```
            
## Returner Results Analysis

Hunter Renfrow and Jabril Peppers were clearly the most over performing returners according to my model, returning the average kick 5 yards further than expected. Mecole Hardman was a notable poor performer as he averaged almost 3 yards less than expected on his 16 attempts. If I was a team looking to improve in the return game by signing or trading for a new returner, I would take a look at Trent Taylor. He averaged over 3 yards more than expected last year and is currently on the Bengals' practice squad so he could be signed to any team's active roster. Players such as Renfrow and Peppers would not be as easily obtained as Taylor would be. 

# Gunner Evaluation

In the most common punt formation, the punting team lines up a player to the left and right of the formation on the outside. These players are called "gunners". The players the return team lines up across from these players to cover them are called "vises". One aspect of a gunner's job is to try and force the returner to fair catch the ball which is a win for the punting team with a good kick from the punter. I decided to evaluate a gunner's ability to force a fair catch.  

## Technique for Evaluation

I chose to make this evaluation a classification problem with the response variable being a dummy variable of whether a 1 if the gunner was within 8 yards of the returner at the time of the returner catching the ball and a 0 otherwise. I chose 8 yards because the average distance of the closest defender to someone who called fair catch was 5 yards and one standard deviation was 3 yards. I again took any gunner's with less than 5 attempts in 2020 and added those to the training data.

## Variables

The response variable was a variable called under_8 which was a 1 if the gunner was under 8 yards at the time of the returner catching the ball and a 0 otherwise. 

The independent variables consisted of 

* kickLength 
* operationTime 
* hangTime 

and my created variables of 

* doubled (Dummy variable: 1 if the gunner is being double teamed, 0 otherwise)
* euc_dist_snap (The euclidean distance of the gunner at the snap to the point on the field where the punt returner will catch the ball)


```{r, echo = F,fig.cap='Figure 2. Jamel Dean Double Team',out.width = '70%',fig.align='center'}
library(knitr)
include_graphics("https://i.imgur.com/Jj9A0XM.jpg")
```


On this play from week 2 of the 2020 season, Buccaneers' gunner Jamel Dean (Number 35) is being double teamed and starts on the far side of the field from where the punt returner catches the ball. He beats the double team by running to the inside and covers a euclidean distance of 53 yards to make it to 3 yards from the returner when he catches the ball. While the Panther's returner breaks Dean's tackle, it was his initial contact that prevents a bigger return even if he should have made the tackle. My model gave Dean a 9% chance of being within 8 yards of the returner at the time of the catch so I would say this play was a win for Dean even with the missed tackle.

```{r, include=FALSE}
library(tidyverse)
library(gganimate)
library(ggplot2)
library(cowplot)
library(repr)

games<-read.csv("../input/nfl-big-data-bowl-2022/games.csv")

players<-read.csv("../input/nfl-big-data-bowl-2022/players.csv")

plays<-read.csv("../input/nfl-big-data-bowl-2022/plays.csv")

track_20<-read.csv("../input/nfl-big-data-bowl-2022/tracking2020.csv")

example_play<-plays%>%filter(gameId==2020092008 & playId == 1998)%>%select(gameId,playId,playDescription)

example_play <- inner_join(example_play,
                           games,
                           by = c("gameId" = "gameId"))

example_play <- inner_join(example_play,
                           track_20,
                           by = c("gameId" = "gameId",
                                  "playId" = "playId"))

cols_fill <- c("#3d85c6", "#663300", "#783f04")
cols_col <- c("#000000", "#663300", "#000000")

xmin <- 0
xmax <- 160/3
hash.right <- 38.35
hash.left <- 12
hash.width <- 3.3

plot_title <- str_trim(gsub("\\s*\\([^\\)]+\\)","",as.character(example_play$playDescription[1])))

ymin <- max(round(min(example_play$x, na.rm = TRUE) - 10, -1), 0)
ymax <- min(round(max(example_play$x, na.rm = TRUE) + 10, -1), 120)

#hash marks
df.hash <- expand.grid(x = c(0, 23.36667, 29.96667, xmax), y = (10:110))
df.hash <- df.hash %>% filter(!(floor(y %% 5) == 0))
df.hash <- df.hash %>% filter(y < ymax, y > ymin)


#plotting
g<- ggplot() +
  
  #setting size and color parameters
  scale_size_manual(values = c(6, 4, 6), guide = FALSE) + 
  scale_shape_manual(values = c(21, 16, 21), guide = FALSE) +
  scale_fill_manual(values = cols_fill, guide = FALSE) + 
  scale_colour_manual(values = cols_col, guide = FALSE) +
  
  #adding hash marks
  annotate("text", x = df.hash$x[df.hash$x < 55/2], 
           y = df.hash$y[df.hash$x < 55/2], label = "_", hjust = 0, vjust = -0.2) + 
  annotate("text", x = df.hash$x[df.hash$x > 55/2], 
           y = df.hash$y[df.hash$x > 55/2], label = "_", hjust = 1, vjust = -0.2) + 
  
  #adding yard lines
  annotate("segment", x = xmin, 
           y = seq(max(10, ymin), min(ymax, 110), by = 5), 
           xend =  xmax, 
           yend = seq(max(10, ymin), min(ymax, 110), by = 5)) + 
  
  #adding field yardline text
  annotate("text", x = rep(hash.left, 11), y = seq(10, 110, by = 10), 
           label = c("G   ", seq(10, 50, by = 10), rev(seq(10, 40, by = 10)), "   G"), 
           angle = 270, size = 4) + 
  annotate("text", x = rep((xmax - hash.left), 11), y = seq(10, 110, by = 10), 
           label = c("   G", seq(10, 50, by = 10), rev(seq(10, 40, by = 10)), "G   "), 
           angle = 90, size = 4) + 
  
  #adding field exterior
  annotate("segment", x = c(xmin, xmin, xmax, xmax), 
           y = c(ymin, ymax, ymax, ymin), 
           xend = c(xmin, xmax, xmax, xmin), 
           yend = c(ymax, ymax, ymin, ymin), colour = "black") +
  
  #adding field interior 
  annotate("rect", xmin = xmin, xmax = xmax, ymin = ymin, ymax = ymax, 
           fill = "dark green", alpha = 0.2) + 
  
  #adding players
  geom_point(data = example_play, aes(x = (xmax-y),
                                      y = x, 
                                      shape = team,
                                      fill = team,
                                      group = nflId,
                                      size = team,
                                      colour = team), 
             alpha = 0.7) +  
  
  #adding jersey numbers
  geom_text(data = example_play, aes(x = (xmax-y), y = x, label = jerseyNumber), colour = "white", 
            vjust = 0.36, size = 3.5) + 
  
  #applying plot limits
  ylim(ymin, ymax) + 
  coord_fixed() +
  
  #applying theme
  theme_nothing() +
  
  #titling plot with play description
  labs(title = plot_title) +
  
  #setting animation parameters
  transition_time(frameId)  +
  ease_aes('linear')

```

```{r, warning=FALSE, echo=FALSE}
animate(g)
```

## Classification Methods

I first used the variables mentioned above to create a logistic regression that modeled the probability a gunner would be within 8 yards of the returner at the time of the catch. Some notable statistics from this model include an accuracy of 73.3%, a False Positive Rate of 16.1% and a False Negative rate of 42.8%.

I then moved onto tree based classification methods including Bagging, Random Forest, and Boosting. Boosting had the highest accuracy of 77.2% with a False Positive Rate of 20.0% and a False Negative rate of 26.8%. I chose boosting as my classification method and proceeded to add my predictions to my testing data set to compare with the actual values. 

Actual Rate is the total number of time a gunner made it within 8 yards of the returner at the time of the catch divided by the total number of attempts he had. Predicted Rate is the total number of predicted times a gunner projected to have been within 8 yards at the time of the catch divided by the total number of attempts. I then subtracted predicted rate from the actual rate to create Actual - Predicted. 

## Results Table

```{r,echo=FALSE}
library(DT)
gunner<-read.csv("../input/data-for-tables/gunner_eval_data.csv")
gunner<-gunner%>%select(displayName,n,actual_rate,predicted_rate,rate_dif)

datatable(gunner,caption = "Comparing Predicted Success Rates vs. Actual Rates for NFL Gunners",filter = 'top', options = list(
  pageLength = 15, autoWidth = TRUE),colnames = c("Name","Attempts","Actual Rate (%)","Predicted Rate (%)","Actual - Predicted (%)"))
```
            
## Gunner Results Analysis

Some interesting results include Jamel Dean, who had the greatest positive differential, making it within 8 yards 60% of the time when he was only predicted to have 10% of the time. If I was a team looking for a boost at the gunner position, I would again try to pick out over achieving players who are currently on practice squads, as they can be easily signed away to my team. A notable player who is currently on a roster but plays sparingly barring injuries, is Nick Westbrook-Ikhine. Westbrook-Ikhine has an actual rate of 60% even though his predicted rate is on 27%.  

# Optimal Number of Punt Rushers 

On a given play, how many rushers should the receiving team bring against the punter? Teams are a little restricted in this decision as the answer lies somewhere in between 1 and 8. If you don't at least send 1 rusher, a punter could just hold the ball and let all his players run down the field to cover. If you bring more than 8, either you have no returner, or potentially could be leaving gunners uncovered. A punter could quickly receive the snap and throw the ball to a gunner for a first down. 

## Method

In order to figure out the optimal number of rushers, I performed some basic filters. I first took out plays listed as "Non-special teams results", punts that ended in a touchback, and punts that were blocked. Punts that had penalties were also removed. I decided to remove blocked punts because I did not want these numbers to negatively influence my averages. Touchbacks were removed because playResult was always listed as kickLength - 20, which makes sense. The problem lies is this is exactly what a play with a 20 yard return would look like, so it, in a sense, invalidates the net yardage variable. I considered artificially changing the kicklength distance to account for this, but decided that it would be better just to remove touchbacks entirely. 

I next created a variable I called "Number_of_Rushers" which counted the number of punt rushers on a given punt. In the PFF Scouting Data data set^[https://www.kaggle.com/c/nfl-big-data-bowl-2022/data], Pro Football Focus has a variable where they list all of the players on a given punt "actively trying to block the punt" and does not include players who "cross the line of scrimmage to engage in punt coverage players in a "Hold Up" role". My variable totaled the numbers of rushers on the list provided by PFF for a particular punt. 

## Results Table

```{r,echo=FALSE}
library(reactable)
rushers<-read.csv("../input/data-for-tables/rushers_data.csv")
rushers<-rushers%>%select(Number_of_Rushers,Avg_Length,Avg_Net,Attempts,Num_Blocked)

reactable(
  rushers,
  defaultColDef = colDef(
    header = function(value) gsub(".", " ", value, fixed = TRUE),
    cell = function(value) format(value, nsmall = 1),
    align = "center",
    minWidth = 70,
    headerStyle = list(background = "#f7f7f8")
  ),
  columns = list(
    Number_of_Rushers=colDef(name="Rushers"),
    Avg_Length=colDef(name = "Average Punt Length"),
    Avg_Net=colDef(name = "Average Net Yards For Punting Team"),
    Attempts=colDef(name="Punts",filterable = FALSE),
    Num_Blocked=colDef(name = "Number Blocked",filterable = FALSE)
  ),
  bordered = TRUE,
  highlight = TRUE,
  filterable = TRUE,
  striped = TRUE
)
```

## Number of Rushers Results Analysis

Looking at the created table, I believe the optimal number of rushers to bring on each punt is 5. Rushing more than 5 does not result in any gain in either net yardage or the number of blocked punts. Rushing 9 provides an interesting case as the average punt length and net yardage is only 36, without including any blocked punt yardage, of which there were 0 anyway. However, we do not have enough attempts to say if the shorter punt length and net is significant or not and rushing 9 players is impractical, as discussed earlier. Unless a player or coach notices a vulnerability in the punt blocking scheme that would require more than 5 rushers to exploit, 5 rushers is the optimal number to bring based on this analysis.

# Conclusions 

In this paper I discussed my techniques for evaluating punt returners by calculating expected return yards, evaluating gunners using classification methods, and found the optimal number of rushers for a punt. If I were to use either my expected return results or classification results to try and improve the special teams of a particular NFL team, I would look to players currently on NFL practice squads, as mentioned with Trent Taylor earlier. These players are realistically obtainable, as compared to players such as Jabril Peppers, Hunter Renfrow, and Jamel Dean. 

# Bonus Content

## Can you "Out Kick" your punt coverage?

Should a punter always punt the ball as hard as he can or is there some scenarios where doing that would result in less yardage than if he kicked it with less power? To attempt to answer this question I split up punts from all 3 seasons into 2 yard intervals or "Bins". 

## Method

Similar to the optimal number of rushers analysis, I first took out punts that had penalties, punts that end in a "Non-special teams result", touchbacks, and blocked punts. I again did this in order to not have these influence my averages. 

I filtered the punt data to be punts that were greater than 45 yards but less than 70 yards. Only 2 punts over 70 yards were returned. The other few were all punts that had been muffed or had bounced down the field to make it 70 yards. 

I then proceeded to split the data into 12 bins of ~2 yards each. Next, I created the mean and median variables by calculating the mean and median of playResult. I created the Average Return variable by calculating the mean of kickReturnYardage. The number of each punts in each bin was totaled and this is the punts variable.

## Results Table

```{r,echo=FALSE}
library(formattable)
out_kick<-read.csv("../input/data-for-tables/out_kick_data.csv")

out_kick<-out_kick%>%mutate("Punt Length"=kickLengthbins,"Median Net Yardage"=Median_Net,"Average Net Yardage"=Mean_Net,"Average Return Yardage"=Avg_return,"Punts"=n)%>%select("Punt Length","Median Net Yardage","Average Net Yardage","Average Return Yardage","Punts")

customGreen0 = "#DeF7E9"

customGreen = "#71CA97"

customRed = "#ff7f7f"

improvement_formatter <-
  formatter("span",
            style = x ~ style(
              font.weight = "bold",
              color = ifelse(x < 9, customGreen, ifelse(x > 12, customRed, "black"))))


formattable(out_kick, align =c("c","c","c"), list(
  `Punt Length` = formatter("span",
                     style = ~ style(color = "black",font.weight = "bold"),
                     x ~ icontext(ifelse(x == "(63.2,65.2]", "star", ""), x)),
  `Punts`= color_tile(customGreen0, customGreen),
  `Median Net Yardage`=color_tile(customGreen0,customGreen),
  `Average Net Yardage`=color_tile(customGreen0,customGreen),
  `Average Return Yardage` = improvement_formatter
))
```

## Results Analysis

The most interesting result observed from this analysis is that average return yardage on the 17 punts that went between 63 and 65 yards was over 18 yards. This was enough to drive the average net yardage gained all the way down to 46, which is less than punts that went between 61 and 63 yards. While this is a small sample size of only 17 punts over 3 seasons, this is indicating that it is possible to out kick your punt coverage.  

# Future Work

Advanced statistics and metrics discussed in this paper are great tools to be used for things like player evaluation and strategy. However, they are most effective when combined with traditional approaches. A smart General Manager or Coach will not draft a player just because he ran a fast 40 yard dash at the combine. When a player puts up an unexpected number, that's a signal to say "Hey, I need to go back and take another look at his tape to see if that speed shows up on the actual field". The same idea rings true with statistics. Players who perform well in the advanced metrics are players that should be reevaluated, but we shouldn't rush to make decisions without first confirming that what the metric says is actually true on the field. 

My future work will consist of watching film to see for example, was Hunter Renfrow actually the best returner in the league last year or was his high performance in my regression the result of outside factors? Why wasn't Matthew Slater, who has made the Pro Bowl the last 10 years as a special teamer, higher on my list of the best gunners? 

# Appendix

All code used in this analysis is available <a href = "https://github.com/cnickol26/Big-Data-Bowl-2022">here</a>.