{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this report, we would like to figure out **if there exist some differences between good teams and bad teams with respect to punts**, and **what factors could influence the punt results**.\n\n","metadata":{}},{"cell_type":"code","source":"# library\nlibrary(readr)\nlibrary(tidyverse)\nlibrary(ggplot2)\nlibrary(stargazer)\nlibrary(lessR)\nlibrary(gtsummary)\nlibrary(lubridate)\nlibrary(ggalluvial)\nlibrary(psych)\nlibrary(olsrr)\nlibrary(MASS)\nlibrary(dplyr)\n\n# import datasets\ngames <- read_csv(\"../input/nfl-big-data-bowl-2022/games.csv\")\nPFFScoutingData <- read_csv(\"../input/nfl-big-data-bowl-2022/PFFScoutingData.csv\")\nplayers <- read_csv(\"../input/nfl-big-data-bowl-2022/players.csv\")\nplays <- read_csv(\"../input/nfl-big-data-bowl-2022/plays.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-01-05T12:33:21.068628Z","iopub.execute_input":"2022-01-05T12:33:21.101012Z","iopub.status.idle":"2022-01-05T12:33:23.367767Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n**1. How to define good or bad teams**\n\n\nFor regular season, both AFC and NFC have 16 teams, and each team play 16 games, therefore, we use **winning percentage** to range teams and select the best or worst team (for example the best team from AFC \"KC\" has 12 wins and 4 losses in 2018)\n\n\n**Good Teams:**\n\n\n2018: KC (AFC) ,  NO (NFC)\n\n\n2019: BAL (AFC) , SF (NFC)\n\n\n2020: KC (AFC) ,  GB (NFC)\n\n\n**Bad Teams:**\n\n\n2018: OAK (AFC) , ARI (NFC)\n\n\n2019: CIN (AFC) , WAS (NFC)\n\n\n2020: JAX (AFC) , ATL (NFC)\n\n\n\nIn order to select the worst team of that specific year, we have 2 conditions, which are \n\n* possessionTeam = team's abbreviation\n\n\n* season = specific year\n\n","metadata":{}},{"cell_type":"code","source":"game_play_data <- merge(games,plays,by=\"gameId\")\npunttotal <- game_play_data %>% filter(specialTeamsPlayType == \"Punt\")\n\n# adjust datatype\npunttotal$homeTeamAbbr <- as.factor(punttotal$homeTeamAbbr)\npunttotal$visitorTeamAbbr <- as.factor(punttotal$visitorTeamAbbr)\npunttotal$possessionTeam <- as.factor(punttotal$possessionTeam)\npunttotal$specialTeamsPlayType <- as.factor(punttotal$specialTeamsPlayType)\npunttotal$specialTeamsResult <- as.factor(punttotal$specialTeamsResult)\npunttotal$penaltyCodes <- as.factor(punttotal$penaltyCodes)\npunttotal$yardlineSide <- as.factor(punttotal$yardlineSide)\n\n# subset only \"good\" and \"bad\" teams from plays\npunttotalgood <- punttotal %>% \n  filter(possessionTeam == \"KC\" & season == \"2018\"| possessionTeam == \"NO\" & season == \"2018\"|\n           possessionTeam == \"BAL\"& season == \"2019\"| possessionTeam == \"SF\"& season == \"2019\"|\n           possessionTeam == \"KC\"& season == \"2020\"| possessionTeam == \"GB\"& season == \"2020\")\n\npunttotalbad <-  punttotal %>% \n  filter(possessionTeam == \"ARI\" & season == \"2018\"| possessionTeam == \"OAK\" & season == \"2018\"|\n           possessionTeam == \"CIN\" & season == \"2019\"| possessionTeam == \"WAS\" & season == \"2019\"|\n           possessionTeam == \"JAX\" & season == \"2020\"| possessionTeam == \"ATL\" & season == \"2020\")\n\npunttotalgood <- punttotalgood %>% add_column(goodorbad=\"Good\")\npunttotalbad <- punttotalbad %>% add_column(goodorbad=\"Bad\")\npunttotal_data <- rbind(punttotalgood,punttotalbad)\npunttotal_data$goodorbad <- as.factor(punttotal_data$goodorbad)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:23.370182Z","iopub.execute_input":"2022-01-05T12:33:23.371954Z","iopub.status.idle":"2022-01-05T12:33:23.709798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**2. Is there any difference between good and bad teams？**","metadata":{}},{"cell_type":"markdown","source":"In the following analysis, we would only focus on the above mentioned 12 teams on corresponding year.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"* **Question 1: is there any difference with regard to their special teams punt results?**","metadata":{}},{"cell_type":"code","source":"# First, we remove the non-special team result\npunttotal_data <- punttotal_data %>% filter(specialTeamsResult!=\"Non-Special Teams Result\")\nggplot(punttotal_data, aes(specialTeamsResult)) + geom_bar(aes(fill = goodorbad), position = \"dodge\") ","metadata":{"execution":{"iopub.status.busy":"2022-01-05T12:33:23.712169Z","iopub.execute_input":"2022-01-05T12:33:23.713587Z","iopub.status.idle":"2022-01-05T12:33:24.260369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By plotting this graph, we found good teams seem to \"make more mistakes\". \n\n**Muffed**: Punts observations of bad teams are much larger than good teams, while their muffed times are similar, which suggests a higher muffed rate of good teams. \n\n**Blocked Punt**: There's no blocked punt from bad teams.\n\nLet's look deeper into the percentage of special teams result.","metadata":{}},{"cell_type":"code","source":"punttotalgood <- punttotalgood %>% filter(specialTeamsResult!=\"Non-Special Teams Result\")\npunttotalbad <- punttotalbad %>% filter(specialTeamsResult!=\"Non-Special Teams Result\")\nBarChart(specialTeamsResult,data=punttotalgood,main=\"Good teams' special team result of punts\")\nBarChart(specialTeamsResult,data=punttotalbad,main=\"Bad teams' special team result of punts\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:24.26274Z","iopub.execute_input":"2022-01-05T12:33:24.264146Z","iopub.status.idle":"2022-01-05T12:33:24.379186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings are as follows:\n\n* **Downed**, **Fair Catch**, **Touchback**: Similar for good and bad teams;\n\n* **Return**: Bad teams > Good teams;\n\n* **Out of Bounds**: Good teams > Bad teams\n","metadata":{}},{"cell_type":"markdown","source":"* **Question 2: do good teams punt longer than bad teams?**","metadata":{}},{"cell_type":"markdown","source":"**Not really**...In the following graph, we found bad teams' kick length seems to be higher than good teams, which is quite surprising.","metadata":{}},{"cell_type":"code","source":"boxplot(punttotalgood$kickLength,punttotalbad$kickLength, col=c(\"red\",\"blue\"),xlab =\"kick length of punt by Good Teams (red) vs Bad Team (blue)\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:24.381608Z","iopub.execute_input":"2022-01-05T12:33:24.382973Z","iopub.status.idle":"2022-01-05T12:33:24.484296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One possible reason could be:\n\nGood teams are already in better position (bigger yardlineNumber), which is shown in the following graph.","metadata":{}},{"cell_type":"code","source":"boxplot(punttotalgood$yardlineNumber,punttotalbad$yardlineNumber, col=c(\"red\",\"blue\"),xlab =\"yard line number of punt by Good Teams (red) vs Bad Team (blue)\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:24.48666Z","iopub.execute_input":"2022-01-05T12:33:24.488077Z","iopub.status.idle":"2022-01-05T12:33:24.565883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems confusing that good teams punt worse than bad teams...\n\nIn order to figure out the reason and influential factors of punt result, we may ask ourselves \"**How to measure a punt's result?**\"\n\nWe first consider that we should focus on punters' kicklength (because kicklength reflects punters' ability).\n\nBut then we figure that punts' kicklength shall be influenced by the punt's location and their strategy. \nFor the reason that when punters' team is close to endzone, a long-distance punt could lead to \"touchback\" (so as long as the ball crosses the line, it doesn't matter if it's 40 yards' punt or 60 yards'). \nAnd when the punt is short, it may be difficult for the punt returner to return, and could be more likely lead to a \"fair catch\".\n\nTherefore, we would like to use the \"**Play Result**\" to measure the punt's result, which *in most cases* equal to \n\n  **kicklength - return yardage - penalty (if any)**\n\nThe above is our hypothesis, and let's take a look **whether different special team results** (eg. fair catch, return, out of bounds, etc.) **varies with punt's play result** :D","metadata":{}},{"cell_type":"code","source":"# distribution of playResult with respect to different special team results\nggplot(punttotal_data,aes(specialTeamsResult,playResult)) +\n  geom_boxplot() ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:24.569607Z","iopub.execute_input":"2022-01-05T12:33:24.571047Z","iopub.status.idle":"2022-01-05T12:33:24.8937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's not surprise that when punt is blocked, the play result is negative. And touchback also results in shorter punt distance.\n\nAs \"muffed\" accounts for only 1%-2% of observations, it seems to have small reliability.\n\n\nWhat we want to highlight here is: in both good and bad teams **Fair Catch have smaller variation than Return**.","metadata":{}},{"cell_type":"markdown","source":"If we only focus on fair catch and return (which are two most frequent results of punts), but this time analysis the play results with respect to the ratio between kicklength and hangtime. \n\nThe intuition of kicklength/hangtime is: hangtime may differ with various strategy but it highly correlates with kicklength, and in order to reflect the potential strategy characteristics, we use their ratio.","metadata":{}},{"cell_type":"code","source":"punttotal_dataPFF <- merge(punttotal_data, PFFScoutingData,by=c(\"gameId\",\"playId\"))\npunttotal_dataPFF %>% filter(specialTeamsResult == \"Fair Catch\"|\n                               specialTeamsResult == \"Return\") %>% \n  ggplot(aes(kickLength/hangTime, playResult, col=specialTeamsResult)) +\n  geom_point()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:24.895901Z","iopub.execute_input":"2022-01-05T12:33:24.897284Z","iopub.status.idle":"2022-01-05T12:33:25.413576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The graph of playResult with respect to kickLength/hangtime has shown that:\n\n* **Fair Catch's play result has better linear correlation with the ratio between kicklength and hangtime**;\n\n* **The ratio between KickLength and hangtime: Fair catch < Return**;\n\n* **Return seems to have no obvious correlation between kicklength/hangtime and play result**.","metadata":{}},{"cell_type":"markdown","source":"But, the result of comparing good and bad teams' punts is disappointing. Good teams punt not as good as bad teams.\n","metadata":{}},{"cell_type":"code","source":"ggplot(punttotal_dataPFF,aes(goodorbad,playResult)) +\n  geom_boxplot()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:25.416127Z","iopub.execute_input":"2022-01-05T12:33:25.417543Z","iopub.status.idle":"2022-01-05T12:33:25.619489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tracking2018 <- read_csv(\"../input/nfl-big-data-bowl-2022/tracking2018.csv\")\ntracking2019 <- read_csv(\"../input/nfl-big-data-bowl-2022/tracking2019.csv\")\ntracking2020 <- read_csv(\"../input/nfl-big-data-bowl-2022/tracking2020.csv\")","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:25.62337Z","iopub.execute_input":"2022-01-05T12:33:25.624854Z","iopub.status.idle":"2022-01-05T12:33:53.020159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"punttotal_dataPFF_18 <- merge(punttotal_dataPFF,tracking2018,by=c(\"gameId\",\"playId\"))\npunttotal_dataPFF_19 <- merge(punttotal_dataPFF,tracking2019,by=c(\"gameId\",\"playId\"))\npunttotal_dataPFF_20 <- merge(punttotal_dataPFF,tracking2020,by=c(\"gameId\",\"playId\"))\npunttotal_PFF_track <- rbind(punttotal_dataPFF_18,punttotal_dataPFF_19)\npunttotal_PFF_track <- rbind(punttotal_PFF_track,punttotal_dataPFF_20)\n punttotal_PFF_tk = punttotal_PFF_track %>% filter(position == \"P\") %>%\n  mutate(\n         snapDetail = as.factor(snapDetail),\n         kickType = as.factor(kickType),\n         kickDirectionIntended = as.factor(kickDirectionIntended),\n         kickDirectionActual = as.factor(kickDirectionActual),\n         playDirection  = as.factor(playDirection))\n\npunttotal_PFF_tk$penaltyYards[is.na(punttotal_PFF_tk$penaltyYards)] <- 0 \npunttotal_PFF_tk$kickReturnYardage[is.na(punttotal_PFF_tk$kickReturnYardage)] <- 0","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:33:53.024943Z","iopub.execute_input":"2022-01-05T12:33:53.027059Z","iopub.status.idle":"2022-01-05T12:37:09.796499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, we try to include most likely related variables to build model 1.","metadata":{}},{"cell_type":"code","source":"m1 <- lm(playResult~yardsToGo+specialTeamsResult+yardlineNumber+\n           penaltyYards+kickLength+kickReturnYardage+goodorbad+snapTime+operationTime+hangTime+\n           kickType+kickDirectionIntended+kickDirectionActual+\n           x+y+s+a+dis+o+dir+playDirection,\n          data=punttotal_PFF_tk)\nsummary(m1)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:37:09.79905Z","iopub.execute_input":"2022-01-05T12:37:09.800458Z","iopub.status.idle":"2022-01-05T12:37:10.09274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above model seems to be \"perfect\", because it contains variables such as \"kicklength\",\"penalty yards\", and \"kick return yardage\" which directly defines the \"play result\". Therefore, in the following model, we would remove these variables and find other influential factors.","metadata":{}},{"cell_type":"code","source":"m2 <- lm(playResult~yardsToGo+specialTeamsResult+yardlineNumber+\n           goodorbad+snapTime+operationTime+hangTime+\n           kickType+kickDirectionIntended+kickDirectionActual+\n           x+y+s+a+dis+o+dir+playDirection,\n          data=punttotal_PFF_tk)\nsummary(m2)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-05T12:37:10.099787Z","iopub.execute_input":"2022-01-05T12:37:10.104777Z","iopub.status.idle":"2022-01-05T12:37:10.383017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From result of the second linear model m2, we found more variables become significant (although the overall model does not have a good explanation of observations, adjusted R-squared is only 0.18) .\n\n\nWe have following findings:\n\n* **Different special teams results have different estimated play result (in yards).**\nTouchbacks tend to have lowest play results, and then it follows with muffed, out of bounds. \nAnd being Fair Catched are on average 0.71 yards better than being returned.\n\n\n* **The smaller yardline number is, the better the punt play results.**\n\n\n* **Good team seems to be perform worse than bad teams with regard to punts.**\n\n\n* **While Snap Time deteriorates play result, longer operation time and longer hang time have positive influence on play result.**\n\n\n* **Normal punts tend to have better results than Aussie-style punts.**\n\n\n* **An intended left punt and an actual right kick tend to have better play results.**\nA right play direction could lead to longer punts result than left. \n\n\n* **The higher speed the punter has, the better the play result.**\n\n\n* **The more distance punter traveled from prior time point, the better result.**\n\n\n","metadata":{}},{"cell_type":"markdown","source":"*Some Addtional:*\n\nLet's try to analysis the play result by machine learning regression tree.\n","metadata":{}},{"cell_type":"code","source":"library(mlr3)\nlibrary(mlr3learners)\nlibrary(mlr3verse)\npuntimp <- punttotal_PFF_tk %>% dplyr::select(c(\"playResult\",\"yardsToGo\",\"specialTeamsResult\",\"yardlineNumber\",\n                                                   \"goodorbad\",\"snapTime\", \"operationTime\", \"hangTime\",\n                                                   \"kickType\",\"kickDirectionIntended\", \"kickDirectionActual\",\n                                                   \"x\" , \"y\" , \"s\",\"a\" ,\"dis\" ,\"o\" , \"dir\",\"playDirection\")) \n\npuntimp <- puntimp[which(complete.cases(puntimp)),]\ntask = as_task_regr(puntimp,target=\"playResult\")\nlearner <- lrn(\"regr.rpart\")  # regression tree\ntrain_set <- sample(task$nrow, 0.8*task$nrow)\ntest_set <- setdiff(seq_len(task$nrow),train_set)\nlearner$train(task,row_ids = train_set)\nprint(learner$model)\nprediction = learner$predict(task, row_ids = test_set)\nprint(prediction) \nautoplot(prediction)\nmeasure = msr(\"regr.mse\")\nprint(measure)\nprediction$score(measure) ","metadata":{"execution":{"iopub.status.busy":"2022-01-05T12:37:10.389923Z","iopub.execute_input":"2022-01-05T12:37:10.394814Z","iopub.status.idle":"2022-01-05T12:37:12.706496Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*From the above graph, by simply using these factors to estimate the true value seems to be problematic, therefore, we can only suggest the postive/negative effect of various factors (as writting above), rather than predicting the exact punt result.*\n\n\n**Thank you for reading this report!** :D","metadata":{}}]}