{"cells":[{"metadata":{"_uuid":"29e846eeb270450ba1ea7e74b2e1f84c34a30c22"},"cell_type":"markdown","source":"Hi, welcome to my first  Kaggle kernel! As implied by the title, this kernel will be focusing on the NFL Punt Analytics Competiton specifically on graphing some basic numbers regarding concussions. Any feedack would be appreciated :)\nLets's get started! "},{"metadata":{"_uuid":"94b70a6753678f7d64e7cd2fdd4e88bb9d7d5901"},"cell_type":"markdown","source":"Firstly, we will start by loading in a few useful libraries. These libraries will be useful for reading in our data manipulation."},{"metadata":{"trusted":true,"_uuid":"26d700c30da6ae180b9431a9e1eb337362da44c4"},"cell_type":"code","source":"library(\"tidyverse\")\nlibrary(\"stringr\")\nlibrary(\"gridExtra\")\nlibrary(\"dplyr\")\nlibrary(\"lubridate\")\nlibrary(\"utils\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"69ff770cce1b9c8f7fd9678ef2141f3faedc5a83"},"cell_type":"markdown","source":"Next we will load in all the data. Notice, there is a significant number of files beginning with \"NGS\", I'll keep those separated for now."},{"metadata":{"_kg_hide-output":false,"_kg_hide-input":true,"trusted":true,"_uuid":"c43e33c6087d7d9bff76777ed708b91a83cb41e5"},"cell_type":"code","source":"#Read in the game data\ngame_data<- read.csv('../input/game_data.csv')\nNGS_2016_post <- read.csv('../input/NGS-2016-post.csv')\nNGS_2016_pre <- read.csv('../input/NGS-2016-pre.csv')\nNGS_2016_reg_wk1_6 <- read.csv('../input/NGS-2016-reg-wk1-6.csv')\nNGS_2016_reg_wk13_17 <- read.csv('../input/NGS-2016-reg-wk13-17.csv')\nNGS_2016_reg_wk7_12 <- read.csv('../input/NGS-2016-reg-wk7-12.csv')\nNGS_2017_post <- read.csv('../input/NGS-2017-post.csv')\nNGS_2017_pre <- read.csv('../input/NGS-2017-pre.csv')\nNGS_2017_reg_wk1_6 <- read.csv('../input/NGS-2017-reg-wk1-6.csv')\nNGS_2017_reg_wk13_17 <- read.csv('../input/NGS-2017-reg-wk13-17.csv')\nNGS_2017_reg_wk7_12 <- read.csv('../input/NGS-2017-reg-wk7-12.csv')\nplay_information <- read.csv('../input/play_information.csv')\nplay_player_role_data <- read.csv('../input/play_player_role_data.csv')\nplayer_punt_data <- read.csv('../input/player_punt_data.csv')\nvideo_review <- read.csv('../input/video_review.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"127a5b97fa3cab6ea11747656945727f0ddf4854"},"cell_type":"markdown","source":"I have taken a look at the data types of all the variables usings **glimpse()** and there are a few character variables which can be factor variables or dates. Let's make those changes to make potential future calculations a lot easier.....\n"},{"metadata":{"trusted":true,"_uuid":"d2af69696a0dbe0b385b4ff49d8fc40d478d5ee6"},"cell_type":"code","source":"#First, let me make a few changes to the exisitng data\n###In this script, I will change some of the variables\n\n#Video_Review\nvideo_review$Primary_Impact_Type <- as.factor(video_review$Primary_Impact_Type)\nvideo_review$Player_Activity_Derived <- as.factor(video_review$Player_Activity_Derived)\nvideo_review$Friendly_Fire <- as.factor(video_review$Friendly_Fire)\nvideo_review$Primary_Partner_Activity_Derived <- as.factor(video_review$Primary_Partner_Activity_Derived)\nvideo_review$Turnover_Related <- as.factor(video_review$Turnover_Related)\nvideo_review$Primary_Partner_GSISID <- as.numeric(video_review$Primary_Partner_GSISID)\n\n#Play_Infomatation\nplay_information$Game_Date <- mdy(play_information$Game_Date)\nplay_information$Play_Type <- as.factor(play_information$Play_Type)\nplay_information$Poss_Team <- as.factor(play_information$Poss_Team)\nplay_information$Season_Type <- as.factor(play_information$Season_Type)\n\n#Play_Player_Role\nplay_player_role_data$Role <- as.factor(play_player_role_data$Role)\n\n#PlayerPuntData\nplayer_punt_data$Position <- as.factor(player_punt_data$Position)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"161e3316d34c83adcba27bb6bd668545f9447025"},"cell_type":"markdown","source":"Now, let's take a look at the NGS file for week 7 to week 12 during the 2017 regular season....."},{"metadata":{"trusted":true,"_uuid":"2df58a7f787ff0d2f25806bdf52e3037ffa61c9b"},"cell_type":"code","source":"head(NGS_2017_reg_wk7_12)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"edfbbb95b286deb1ec16ea3df4912a0b973ef39f"},"cell_type":"markdown","source":"Next, I want to extract \"NGS\" data about players who suffered concussions. Since there are plenty of NGS files to go through, I am going to use nested for loops to automate the process. Prior, to making the loops , I am going to make vectors with the relevant years, seasons and weeks to help the loop."},{"metadata":{"trusted":true,"_uuid":"5a02f043def166a63e5e092281c563cd5c398347"},"cell_type":"code","source":"#For regular season\nfyears =c(\"_2016\",\"_2017\")\nseason = \"_reg\"\nweeks = c(\"_wk1_6\",\"_wk7_12\",\"_wk13_17\")\n\n#For post and pre season\nfyears =c(\"_2016\",\"_2017\")\nseasons = c(\"_pre\",\"_post\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f3515236ff57831471b7c79ec27c16a6db40210"},"cell_type":"markdown","source":"It's a little bit difficult for me to loop through regular and non-regular season data in the same loop , so I am going to have 2 separate loops for both regular season and non-regular season. Let's start by extracting the regular season NGS data for concussion players..."},{"metadata":{"_uuid":"22a29dcaf78a703877aedd86c70d6a14bc625abc","_execution_state":"idle","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"for (fyear in fyears){\n    for (week in weeks){\n        \n        ##Step1, Make inputs specifically for the NGS file\n        #Inputs for the NGS File\n        ngsfilename = eval(parse(text= paste(\"NGS\",fyear,season,week,sep=\"\")))\n       \n        \n        \n        #Step2, make separate files of game ids, play ids and player ids for scenarios were players were injured\n        assign(paste(\"injure_game_id\",fyear,sep=\"\"),subset(video_review, Season_Year == as.numeric(str_sub(fyear,start=2,end=5)))$GameKey)\n        assign(paste(\"injure_play_id\",fyear,sep=\"\"),subset(video_review, Season_Year == as.numeric(str_sub(fyear,start=2,end=5)))$PlayID)\n        assign(paste(\"injure_player_id\",fyear,sep=\"\"),subset(video_review, Season_Year == as.numeric(str_sub(fyear,start=2,end=5)))$GSISID)\n\n        ####In this script, I am going to get the NGS for players who suffered concussions\n\n\n        #Getting relevant games\n        a = tibble()\n        for (i in eval(parse(text = paste(\"injure_game_id\",fyear,sep=\"\")))) {\n          a <- rbind(a, ngsfilename %>%\n                       select(everything()) %>%\n                       filter(GameKey == i))\n        }\n\n\n        #Getting relevant plays\n        b = tibble()\n        for (i in eval(parse(text = paste(\"injure_play_id\",fyear,sep=\"\")))) {\n          b <- rbind(b,a %>%\n            select(everything()) %>%\n            filter(a$PlayID %in% i ))\n        }\n\n\n        #Getting relevant players\n        c = tibble()\n        for (i in eval(parse(text = paste(\"injure_player_id\",fyear,sep=\"\")))) {\n          c <- rbind(c,b %>%\n            select(everything()) %>%\n            filter(b$GSISID %in% i))\n        }\n\n        assign(paste(\"NGS\",fyear,week,\"injure\",sep = \"\"),c)\n    }\n}\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"12108216d5b3f256481a5bb012ba53d866544abd"},"cell_type":"markdown","source":"Next, let's do the post-season and the pre-season data....."},{"metadata":{"trusted":true,"_uuid":"a5e6221eb840658d9bb2c9f86ff507c6722e53f9"},"cell_type":"code","source":"for (fyear in fyears){\n    for (ppp in seasons){\n        \n        ##Step1, Make inputs specifically for the NGS file\n        #Inputs for the NGS File\n        ngsfilename = eval(parse(text= paste(\"NGS\",fyear,ppp,sep=\"\")))\n       \n        \n        \n        #Step2, make separate files of game ids, play ids and player ids for scenarios were players were injured\n        assign(paste(\"injure_game_id\",fyear,sep=\"\"),subset(video_review, Season_Year == as.numeric(str_sub(fyear,start=2,end=5)))$GameKey)\n        assign(paste(\"injure_play_id\",fyear,sep=\"\"),subset(video_review, Season_Year == as.numeric(str_sub(fyear,start=2,end=5)))$PlayID)\n        assign(paste(\"injure_player_id\",fyear,sep=\"\"),subset(video_review, Season_Year == as.numeric(str_sub(fyear,start=2,end=5)))$GSISID)\n\n        ####In this script, I am going to get the NGS for players who suffered concussions\n\n\n        #Getting relevant games\n        a = tibble()\n        for (i in eval(parse(text = paste(\"injure_game_id\",fyear,sep=\"\")))) {\n          a <- rbind(a, ngsfilename %>%\n                       select(everything()) %>%\n                       filter(GameKey == i))\n        }\n\n\n        #Getting relevant plays\n        b = tibble()\n        for (i in eval(parse(text = paste(\"injure_play_id\",fyear,sep=\"\")))) {\n          b <- rbind(b,a %>%\n            select(everything()) %>%\n            filter(a$PlayID %in% i ))\n        }\n\n\n        #Getting relevant players\n        c = tibble()\n        for (i in eval(parse(text = paste(\"injure_player_id\",fyear,sep=\"\")))) {\n          c <- rbind(c,b %>%\n            select(everything()) %>%\n            filter(b$GSISID %in% i))\n        }\n\n        assign(paste(\"NGS\",fyear,ppp,\"injure\",sep = \"\"),c)\n    }\n}\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1f2a88e0c0491bd29232f2fef15761e859889b04"},"cell_type":"markdown","source":"Now to put them all together..."},{"metadata":{"trusted":true,"_uuid":"2841e0b78ebc16f9cfa8457ca5221e39b2894567"},"cell_type":"code","source":"NGS_total_injure <- rbind(NGS_2016_preinjure,NGS_2016_wk1_6injure,NGS_2016_wk7_12injure,NGS_2016_wk13_17injure,NGS_2016_postinjure,NGS_2017_preinjure,NGS_2017_wk1_6injure,NGS_2017_wk7_12injure,NGS_2017_wk13_17injure,NGS_2017_postinjure)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"254f419705810d9e3d0319422a571393fdd2e2f3"},"cell_type":"markdown","source":"Now that we have done that, let's make some graphs!\nI want to graph the following:\n* 1. Where concussion players are at the start of plays\n* 2. Where concussion players are during plays\n* 3. Where concussion players are at the end of plays\n* 4. Which direction are concussion players facing at the start of plays\n* 5. Which direction are concussion players facing during plays\n* 6. Which direction are concussion players facing at the end of plays\n"},{"metadata":{"trusted":true,"_uuid":"e5e7ca24486b933ad71901fd8f3e19a80186b4ca","scrolled":false},"cell_type":"code","source":"#Where were they on the field at the start of the play ?\nNGS_total_injure %>%\n  filter(Event==\"punt_play\")%>%\n  ggplot(mapping=aes(x=x,y=y)) + geom_bin2d(binwidth=c(10,5),color = \"black\") + coord_cartesian(xlim =c(0,120), ylim = c(0,53.3)) +\n  ggtitle(\"Where are concussion players on the field at the start of plays\")\n\n#Where are players on the field during plays ?\nNGS_total_injure %>%\n  filter(Event %in% unique(NGS_total_injure$Event)[1])%>%\n  ggplot(mapping=aes(x=x,y=y)) + geom_bin2d(binwidth=c(8,5),color = \"black\") + coord_cartesian(xlim =c(0,120), ylim = c(0,53.3)) +\n  ggtitle(\"On the field, where are concussion players during plays ?\")\n  \n \n#Where are concussion players on the field at the end of the play ?\nNGS_total_injure %>%\n  filter(Event==\"play_submit\")%>%\n  ggplot(mapping=aes(x=x,y=y)) + geom_bin2d(binwidth=c(10,5),color = \"black\") + coord_cartesian(xlim =c(0,120), ylim = c(0,53.3)) +\n  ggtitle(\"On the field, where are concussion players at the end of plays ?\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d76efd818d0bc57b5bb94000c61aaa338a61dda1"},"cell_type":"markdown","source":"Looking at these graphs , the pattern's look as expected given the setup of punt plays.\nThere is an area with the  coordinates (85,28) seems to be a hive of activity for concussion punt plays. I'll later compare it with non-concussion punt plays and see if there's any significant takeaways.\n\nNext ,let's examine the direciton players are facing throughout punt plays"},{"metadata":{"trusted":true,"_uuid":"7293eb14ea3e7428d86ce39da946c6f40c58d576"},"cell_type":"code","source":"#What direction are players facing at the start of plays ?\nNGS_total_injure %>%\n  filter(Event==\"punt_play\")%>%\n  ggplot(mapping=aes(x=o)) + geom_histogram(bins=11,binwidth = 10) + \n  ggtitle(\"On the field, which direction are players facing at the start of plays ?\")\n\n#What direction are players facing the during of plays\nNGS_total_injure %>%\n  filter(Event %in% unique(NGS_total_injure$Event)[1])%>%\n  ggplot(mapping=aes(x=o)) + geom_histogram(bins=11,binwidth = 10) +\n  ggtitle(\"On the field, which direction are players facing during of plays ?\")\n\n#What direcition are players facing a the end of plays ?\nNGS_total_injure %>%\n  filter(Event==\"play_submit\")%>%\n  ggplot(mapping=aes(x=o)) + geom_histogram(bins=11,binwidth = 10) +\n  ggtitle(\"On the field, which direcition are players facing at the end of plays ?\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f4f4fcbbea912d5f5b4d7c6d4800920842b1d77c"},"cell_type":"markdown","source":"These directions seem uniformly random and don't provide any useful insight . We can try counting for whether a player is part of the punt-coverage or punt-return to hopeful gain something useful.\nUnfortunately , I have run out of RAM , but I will continue in the next kernel :)"}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}