{"cells":[{"metadata":{"_uuid":"5956280dcc957d676d75382ebb1d2c1aa8085e50"},"cell_type":"markdown","source":"# NFL Punt Approximate Entropy Analysis\n## by Shawn Hainsworth\n\n### **Overview**\nConcussions are more likely to occur during kick-off returns and punts than other play types\nSource:  Twelve Years of National Football League Concussion Data [^1]\n\n![Concussions by Play Type](https://legalbiguy.com/wp-content/uploads/2019/01/ConcussionByPlayType.png)\n[^1] Authors: Ira R. Casson, MD, Dr med David C. Viano, PhD, John W. Powell, PhD, and Elliot J. Pellman, MD\n\n**Approximate Entropy**\nIn statistics, approximate entropy (ApEn) is a technique used to quantify the amount of unpredictability over time-series data.\n\n**Hypothesis:**\nThe additional distance covered by each player during Kick-off returns and punt plays result in greater player speeds as well as more unpredictable player direction and orientation changes.  In addition, less predictable ball position, and resulting ad-hoc blocking and tackling formations result in greater entropy for these play types than for running and passing plays originating from the line of scrimmage.  Reacting to less predictable situations at greater speeds creates a greater risk of concussions.\n\n**Methodology**\nApproximate Entropy will be calculated for each Game/Play/Player for three dimensions: distance, direction and orientation.  In addition, player entropies will be summed for each play.  Averages of all player entropy values by play (distance, direction, orientation, total) will be calculated.  Statistical correlations between player and play entropy and concussions will be evaluated.  Finally, a rule change will be proposed to reduce entropy for both punts and kick offs.\n\n**Technical**\nThis analysis uses the [Microsoft Machine Learning Server](https://docs.microsoft.com/en-us/machine-learning-server/what-is-machine-learning-server) and [RevoScaleR package](https://docs.microsoft.com/en-us/machine-learning-server/r-reference/revoscaler/revoscaler) to speed up player entropy calculations.  This parallel processing solution also makes it possible to run this analysis on a dataset of any size.  However, the first block of code is not runnable directly in this notebook, but must be run on a Machine Learning Server instance.  The resulting file has been uploaded and subsequent code blocks run within the notebook using the entropy results generated on the Machine Learning server.  In addition, the Min / Max XY data was also calculated on the machine learning server and the data uploaded for use in this notebook.  Two entropy packages were tested.  The results included use the empirical entropy function from [Package \"entropy\"](https://cran.r-project.org/web/packages/entropy/entropy.pdf).  In addition, results were generated using [Package 'pracma'](https://cran.r-project.org/web/packages/pracma/pracma.pdf).  As the pracma calculations take longer to run, the entropy package was used.\n\n**Limitations**\nFor this competition, the data provided represents only punt plays.  Therefore, within the context of the competition, it is not possible to compare the entropy of punt plays vs. other play types. "},{"metadata":{"_uuid":"22e799dda4a40af6b76e43fc162c502b1153c575"},"cell_type":"markdown","source":"```r\n# NOTE:  This cell will not run in the Jupyter notebook.  It must be run on the Microsoft Machine Learnining Server\n#        This code uses the RevoScaleR package to process large files in parallel\n#        For the competition, this code was run on a Microsoft Azure Standard A8m v2 (8 vcpus, 64 GB memory)\n\n# Install standard R Packages\ninstall.packages(\"dplyr\")\ninstall.packages(\"lubridate\")\ninstall.packages(\"entropy\")\n\n# Install dplpyrXdf, requires devtools\n# devtools requires openssl and curl: sudo yum install openssl-devel, sudo yum install libcurl-devel\ninstall.packages(\"devtools\")\ninstall.packages(c('devtools', 'curl'))\nSys.setenv(CURL_CA_BUNDLE = file.path(Sys.getenv(\"R_HOME\"), \"lib/microsoft-r-cacert.pem\"))\ndevtools::install_github(\"RevolutionAnalytics/dplyrXdf\")\n\nlibrary(RevoScaleR)\nlibrary(dplyrXdf)\nlibrary(lubridate)\nlibrary(entropy)\n\n\".CalcEntropy\" <- function(keys, data) {\n    library(dplyrXdf)\n    library(entropy)\n\n    gppNGSDF <- rxImport(inData = data)\n\n    gppNGSDF <- gppNGSDF[order(gppNGSDF$Time),]\n\n    disEntropy <- entropy.empirical(gppNGSDF$dis)\n    oEntropy <- entropy.empirical(gppNGSDF$o)\n    dirEntropy <- entropy.empirical(gppNGSDF$dir)\n    totEntropy <- disEntropy + oEntropy + dirEntropy\n\n    return(list(\"Season_Year\" = gppNGSDF$Season_Year[1],\n                          \"GameKey\" = gppNGSDF$GameKey[1],\n                          \"PlayID\" = gppNGSDF$PlayID[1],\n                          \"GSISID\" = gppNGSDF$GSISID[1],\n                          \"disEntropy\" = disEntropy,\n                          \"oEntropy\" = oEntropy,\n                          \"dirEntropy\" = dirEntropy,\n                          \"totEntropy\" = totEntropy))\n}\n\n\".ProcessNGSFile\" <- function(name) {\n    inFile <- file.path(\"/opt/nfl\", paste(name, \".csv\", sep = \"\"))\n    inXdfDS <- RxXdfData(file = inFile)\n\n    # Create a partitioned Xdf data source\n    outFile <- file.path(tempdir(), name)\n    outPartXdfDataSource <- RxXdfData(file = outFile, createPartitionSet = TRUE)\n    partDF <- rxPartition(inData = inXdfDS, outData = outPartXdfDataSource, varsToPartition = c(\"GameKey\", \"PlayID\", \"GSISID\"))\n\n    # Local compute context\n    parallelContext <- RxLocalParallel()\n    rxSetComputeContext(parallelContext)\n    rxOptions(numCoresToUse = 8)\n\n    # Call rxExecBy to calculate entropy on partitions\n    entropyResults <- rxExecBy(inData = outPartXdfDataSource, keys = c(\"GameKey\", \"PlayID\", \"GSISID\"), func = .CalcEntropy)\n\n    # Transform the results and write to XDF\n    entropyList <- lapply(entropyResults, function(x) data.frame(x$result))\n    entropyDF <- as.data.frame(do.call(\"rbind\", entropyList))\n    entropyXdfFileName <- file.path(\"/opt/nfl\", paste(name, \"Entropy.xdf\", sep = \"\"))\n    rxImport(entropyDF, outFile = entropyXdfFileName)\n\n    # Clean-up and delete the partitioned Xdf\n    unlink(outFile, recursive = TRUE, force = TRUE)\n\n    return(entropyDF)\n}\n\n#2016\nrxImport(.ProcessNGSFile(\"NGS-2016-pre\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2016-reg-wk1-6\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2016-reg-wk7-12\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2016-reg-wk13-17\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2016-post\"))\n\n#2017\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2017-pre\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2017-reg-wk1-6\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2017-reg-wk7-12\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2017-reg-wk13-17\"))\nresults <- rbind(results, .ProcessNGSFile(\"NGS-2017-post\"))\n\n# Combine results\nresultsCsvFileName <- file.path(\"/opt/nfl/CombinedEntropy.csv\")\nrxXdfToText(inData = results, outFile = resultsCsvFileName, overwrite = TRUE)\n```                     "},{"metadata":{"trusted":true,"_uuid":"ee16c68ad5c23f6695021456cb597798624b8517"},"cell_type":"code","source":"# Load Data\ngame <- read.csv(\"../input/NFL-Punt-Analytics-Competition/game_data.csv\")\nplay_information <- read.csv(\"../input/NFL-Punt-Analytics-Competition/play_information.csv\")\nvideo_review <- read.csv(\"../input/NFL-Punt-Analytics-Competition/video_review.csv\")\nplay_player_role <- read.csv(\"../input/NFL-Punt-Analytics-Competition/play_player_role_data.csv\")\nplayer_punt <- read.csv(\"../input/NFL-Punt-Analytics-Competition/player_punt_data.csv\")\ngpp_entropy <- read.csv(\"../input/nfl-ngs-entropy/CombinedEntropy.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"07436ebe69b6e64333d3aee0cf514289265f5c6c"},"cell_type":"code","source":"# Feature Engineering\n# Create punt_play_entropy file with one entry per GameKey/PlayId/GSISID (Player) with entropy features\nlibrary(dplyr)\nlibrary(stringr)\n\n# Create Punt Play Entropy DataSet\npunt_play_entropy <- play_information[,c(1:3,6)]\ngpp_entropy <- gpp_entropy[complete.cases(gpp_entropy),]\ngpp_entropy_mean <- gpp_entropy %>% group_by(GameKey, PlayID) %>% summarize_at(.vars = vars(disEntropy, oEntropy, dirEntropy, totEntropy), .funs = c(mean))\npunt_play_entropy <- left_join(punt_play_entropy, gpp_entropy_mean, by = c(\"GameKey\", \"PlayID\"))\n\n# Add Concussion data from video_review\nvideo_review$Concussion <- TRUE\nvideo_review <- video_review[ , !names(video_review) %in% c(\"Season_Year\")]\npunt_play_entropy <- left_join(punt_play_entropy, video_review, by = c(\"GameKey\", \"PlayID\"))\n\n# Add Concussion Player Entropy\ngpp_entropy <- gpp_entropy[ , !names(gpp_entropy) %in% c(\"Season_Year\")]\n\npunt_play_entropy <- left_join(punt_play_entropy, gpp_entropy, by = c(\"GameKey\", \"PlayID\", \"GSISID\"))\n# Rename .y columns\npunt_play_entropy <- punt_play_entropy %>% rename(disEntropy.GSISID = disEntropy.y)\npunt_play_entropy <- punt_play_entropy %>% rename(oEntropy.GSISID = oEntropy.y)\npunt_play_entropy <- punt_play_entropy %>% rename(dirEntropy.GSISID = dirEntropy.y)\npunt_play_entropy <- punt_play_entropy %>% rename(totEntropy.GSISID = totEntropy.y)\n\n# Add Primary Partner Entropy\npunt_play_entropy$Primary_Partner_GSISID <- as.integer(str_extract(punt_play_entropy$Primary_Partner_GSISID, \"[[:digit:]]+\"))\npunt_play_entropy <- left_join(punt_play_entropy, gpp_entropy, by = c(\"GameKey\", \"PlayID\", \"Primary_Partner_GSISID\" = \"GSISID\"))\n# Rename columns\npunt_play_entropy <- punt_play_entropy %>% rename(disEntropy.Primary_Partner_GSISID = disEntropy)\npunt_play_entropy <- punt_play_entropy %>% rename(oEntropy.Primary_Partner_GSISID = oEntropy)\npunt_play_entropy <- punt_play_entropy %>% rename(dirEntropy.Primary_Partner_GSISID = dirEntropy)\npunt_play_entropy <- punt_play_entropy %>% rename(totEntropy.Primary_Partner_GSISID = totEntropy)\n\n# Rename columns\npunt_play_entropy <- punt_play_entropy %>% rename(disEntropy.Play = disEntropy.x)\npunt_play_entropy <- punt_play_entropy %>% rename(oEntropy.Play = oEntropy.x)\npunt_play_entropy <- punt_play_entropy %>% rename(dirEntropy.Play = dirEntropy.x)\npunt_play_entropy <- punt_play_entropy %>% rename(totEntropy.Play = totEntropy.x)\n\n# Role Entropy\nplayer_play_role_entropy <- play_player_role\n# Remove unused columns\nplayer_play_role_entropy <- player_play_role_entropy[ , !names(player_play_role_entropy) %in% c(\"Season_Year\")]\nplayer_play_role_entropy <- player_play_role_entropy[ , !names(player_play_role_entropy) %in% c(\"Number\")]\n# Join Entropy\nplayer_play_role_entropy <- left_join(player_play_role_entropy, gpp_entropy, by = c(\"GameKey\", \"PlayID\", \"GSISID\"))\nplayer_play_role_entropy <- player_play_role_entropy[complete.cases(player_play_role_entropy),]\nplayer_role_entropy_mean <- player_play_role_entropy %>% group_by(Role) %>% summarize_at(.vars = vars(disEntropy, oEntropy, dirEntropy, totEntropy), .funs = c(mean))\n# Create Role Mean Entropy columns\nplayer_play_role_entropy$disEntropyRoleMean <- NA\nplayer_play_role_entropy$oEntropyRoleMean <- NA\nplayer_play_role_entropy$dirEntropyRoleMean <- NA\nplayer_play_role_entropy$totEntropyRoleMean <- NA\n# Set Role Mean Entropy values\nfor(role in player_role_entropy_mean$Role) {\n  index <- player_play_role_entropy$Role %in% role\n  roleMean <- player_role_entropy_mean %>% filter(Role == role)\n  player_play_role_entropy[index, c(\"disEntropyRoleMean\")] <- roleMean$disEntropy\n  player_play_role_entropy[index, c(\"oEntropyRoleMean\")] <- roleMean$oEntropy\n  player_play_role_entropy[index, c(\"dirEntropyRoleMean\")] <- roleMean$dirEntropy\n  player_play_role_entropy[index, c(\"totEntropyRoleMean\")] <- roleMean$totEntropy\n}\n# Remove unused columns\nplayer_play_role_entropy <- player_play_role_entropy[ , !names(player_play_role_entropy) %in% c(\"disEntropy\")]\nplayer_play_role_entropy <- player_play_role_entropy[ , !names(player_play_role_entropy) %in% c(\"oEntropy\")]\nplayer_play_role_entropy <- player_play_role_entropy[ , !names(player_play_role_entropy) %in% c(\"dirEntropy\")]\nplayer_play_role_entropy <- player_play_role_entropy[ , !names(player_play_role_entropy) %in% c(\"totEntropy\")]\n# Join Player Role Entropy\npunt_play_entropy <- left_join(punt_play_entropy, player_play_role_entropy, by = c(\"GameKey\", \"PlayID\", \"GSISID\"))\n# Join Primary Partner Role Entropy\npunt_play_entropy <- left_join(punt_play_entropy, player_play_role_entropy, by = c(\"GameKey\", \"PlayID\", \"Primary_Partner_GSISID\" = \"GSISID\"))\n# Rename Columns\npunt_play_entropy <- punt_play_entropy %>% rename(Role.GSISID = Role.x)\npunt_play_entropy <- punt_play_entropy %>% rename(disEntropyRoleMean.GSISID = disEntropyRoleMean.x)\npunt_play_entropy <- punt_play_entropy %>% rename(oEntropyRoleMean.GSISID = oEntropyRoleMean.x)\npunt_play_entropy <- punt_play_entropy %>% rename(dirEntropyRoleMean.GSISID = dirEntropyRoleMean.x)\npunt_play_entropy <- punt_play_entropy %>% rename(totEntropyRoleMean.GSISID = totEntropyRoleMean.x)\npunt_play_entropy <- punt_play_entropy %>% rename(Role.PrimaryPartnerGSISID = Role.y)\npunt_play_entropy <- punt_play_entropy %>% rename(disEntropyRoleMean.PrimaryPartnerGSISID = disEntropyRoleMean.y)\npunt_play_entropy <- punt_play_entropy %>% rename(oEntropyRoleMean.PrimaryPartnerGSISID = oEntropyRoleMean.y)\npunt_play_entropy <- punt_play_entropy %>% rename(dirEntropyRoleMean.PrimaryPartnerGSISID = dirEntropyRoleMean.y)\npunt_play_entropy <- punt_play_entropy %>% rename(totEntropyRoleMean.PrimaryPartnerGSISID = totEntropyRoleMean.y)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a21247688816312bce28bd43affb4ca03a39e7ad"},"cell_type":"markdown","source":"## Analysis ##\nThe following cells will analyze the entropy data.  This analysis will be used to support the recommendations for rule changes below."},{"metadata":{"trusted":true,"_uuid":"9c0433dc0f7e71605e3052e240db92ad0dec9ba4"},"cell_type":"code","source":"# Compare Total Entropy of Concussed Player (GSISID) and Primary Partner with Mean Entropy of Other Players in Play\nconcussions_plays <- punt_play_entropy %>% filter(Concussion == TRUE)\nconcussions_plays <- concussions_plays %>% mutate(GSISID_pct_diff = ((totEntropy.GSISID - totEntropy.Play) / totEntropy.Play * 100))\nconcussions_plays <- concussions_plays %>% mutate(Primary_Partner_GSISID_pct_diff = ((totEntropy.Primary_Partner_GSISID - totEntropy.Play) / totEntropy.Play * 100))\nconcussions_plays[,c(\"totEntropy.Play\", \"totEntropy.GSISID\", \"GSISID_pct_diff\", \"totEntropy.Primary_Partner_GSISID\", \"Primary_Partner_GSISID_pct_diff\")]\nprint(paste(\"Mean GSISID Pct Diff: \", mean(concussions_plays$GSISID_pct_diff)))\nprint(paste(\"Mean Primary Player GSISID Pct Diff: \", mean(na.omit(concussions_plays$Primary_Partner_GSISID_pct_diff))))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"db1d6f8b1712c90a1980560a4b12d044a69606ec"},"cell_type":"markdown","source":"## Analysis ##\nIn the above table, in almost all cases, the player and primary parter (if identified) have consistently higher total entropy values than the average player entropy for the play.  This data supports the hypothesis that there is a correlation between entropy and concussions. "},{"metadata":{"trusted":true,"_uuid":"3bebcac0b858dc2478a9502ba20eae252b444af1"},"cell_type":"code","source":"library(ggplot2)\nconcussions_plays$row <- seq(1:nrow(concussions_plays))\nggplot(concussions_plays, aes(row)) + \n  geom_line(aes(y = totEntropy.Play), color = \"blue\") + \n  geom_line(aes(y = totEntropy.GSISID), color = \"red\") +\n  ggtitle(\"Total Entropy Comparison (Concussion Player = Red, Mean Player = Blue)\") +\n  xlab('row') +\n  ylab('Total Entropy')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"583c02bc3daca73330d5362c942b78cf7574b372"},"cell_type":"code","source":"# Compare concussion player with the mean entropy for players in the same role.\nconcussions_plays <- concussions_plays %>% mutate(GSISID_Role_pct_diff = ((totEntropy.GSISID - totEntropyRoleMean.GSISID) / totEntropyRoleMean.GSISID * 100))\nconcussions_plays[,c(\"Role.GSISID\", \"totEntropyRoleMean.GSISID\", \"totEntropy.GSISID\", \"GSISID_Role_pct_diff\")]\nprint(paste(\"Mean Role GSISID Pct Diff: \", mean(concussions_plays$GSISID_Role_pct_diff)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2bf38cbbcd7594183276732d77b7a3460cf6a28c"},"cell_type":"code","source":"# Compare primary partner entropy with the mean entropy for players in the same role.\nconcussions_plays <- concussions_plays %>% mutate(Primary_Partner_GSISID_Role_pct_diff = ((totEntropy.Primary_Partner_GSISID - totEntropyRoleMean.PrimaryPartnerGSISID) / totEntropyRoleMean.PrimaryPartnerGSISID * 100))\nconcussions_plays[,c(\"Role.PrimaryPartnerGSISID\", \"totEntropyRoleMean.PrimaryPartnerGSISID\", \"totEntropy.Primary_Partner_GSISID\", \"Primary_Partner_GSISID_Role_pct_diff\")]\nprint(paste(\"Mean Primary Player Role GSISID Pct Diff: \", mean(na.omit(concussions_plays$Primary_Partner_GSISID_Role_pct_diff))))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b33c1243b8bf54445e4647573407e7f90c020845"},"cell_type":"markdown","source":"## Analysis ##\nThe entropy for players involved in concussion plays is very close to the average entropy for the role they are playing.  As indicated above, the data supports the assertion that entropy is related to concussion events.  However, by looking at role data, we can conclude that players involved in concussions are playing higher entropy roles, making them more susceptible to concussion events.  As we cannot alter the roles, we must look for a rule change which lowers the entropy for all players on the field."},{"metadata":{"trusted":false,"_uuid":"81cfdb4cc962f92ec6a0757e8c3a7d7238108a8a"},"cell_type":"markdown","source":"## Hypothesis ##\nThere is a correlation between the total playing area and the mean player entropy by play.  To test this hypothesis, the minimum and maximum X and Y will be calculated for all players for a given play.  The resulting values describe a rectangle of a given area in which a given play occurred (the XY area).  These values will be plotted against total entropy."},{"metadata":{"_uuid":"831c7f4b4afd68adfcbc3d818ad6880d3c43a4c1"},"cell_type":"markdown","source":"```r\n# NOTE:  The code in this cell was also run on the Microsoft Machine Learnining Server for performance\n#        However, the logic is simple dplyr group_by and summarize and can be run in any R environment\n\nlibrary(RevoScaleR)\nlibrary(dplyrXdf)\n\n\".MinMaxNGSFile\" <- function(name) {\n\tinFile <- file.path(\"/opt/nfl\", paste(name, \".xdf\", sep = \"\"))\n\tprint(inFile)\n\tinXdfDS <- RxXdfData(file = inFile)\n\n\tngsDF <- rxImport(inXdfDS)\n\tngsDF <- ngsDF %>% group_by(GameKey, PlayID)\n\tminX <- ngsDF %>% summarize(min_x = min(x, na.rm = TRUE))\n\tmaxX <- ngsDF %>% summarize(max_x = max(x, na.rm = TRUE))\n\tminY <- ngsDF %>% summarize(min_y = min(y, na.rm = TRUE))\n\tmaxY <- ngsDF %>% summarize(max_y = max(y, na.rm = TRUE))\n\tminMaxDF <- minX\n\tminMaxDF$max_x <- maxX$max_x\n\tminMaxDF$min_y <- minY$min_y\n\tminMaxDF$max_y <- maxY$max_y\n\tminMaxXdfFileName <- file.path(\"/opt/nfl\", paste(name, \"MinMaxXY.xdf\", sep = \"\"))\n\trxImport(minMaxDF, outFile = minMaxXdfFileName)\n}\n\n.MinMaxNGSFile(\"NGS-2016-pre\")\n.MinMaxNGSFile(\"NGS-2016-reg-wk1-6\")\n.MinMaxNGSFile(\"NGS-2016-reg-wk7-12\")\n.MinMaxNGSFile(\"NGS-2016-reg-wk13-17\")\n.MinMaxNGSFile(\"NGS-2016-post\")\n\n.MinMaxNGSFile(\"NGS-2017-pre\")\n.MinMaxNGSFile(\"NGS-2017-reg-wk1-6\")\n.MinMaxNGSFile(\"NGS-2017-reg-wk7-12\")\n.MinMaxNGSFile(\"NGS-2017-reg-wk13-17\")\n.MinMaxNGSFile(\"NGS-2017-post\")\n\n\n# Combine results\ny2016PreXdfFileName <- file.path(\"/opt/nfl\", \"NGS-2016-preMinMaxXY.xdf\")\nresults <- rxImport(y2016PreXdfFileName)\ny2016W1_6XdfFileName <- file.path(\"/opt/nfl\", \"NGS-2016-reg-wk1-6MinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2016W1_6XdfFileName))\ny2016W7_12XdfFileName <- file.path(\"/opt/nfl\", \"NGS-2016-reg-wk7-12MinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2016W7_12XdfFileName))\ny2016W13_17XdfFileName <- file.path(\"/opt/nfl\", \"NGS-2016-reg-wk13-17MinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2016W13_17XdfFileName))\ny2016PostXdfFileName <- file.path(\"/opt/nfl\", \"NGS-2016-postMinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2016PostXdfFileName))\n\ny2017PreXdfFileName <- file.path(\"/opt/nfl\", \"NGS-2017-preMinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2017PreXdfFileName))\ny2017W1_6XdfFileName <- file.path(\"/opt/nfl\", \"NGS-2017-reg-wk1-6MinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2017W1_6XdfFileName))\ny2017W7_12XdfFileName <- file.path(\"/opt/nfl\", \"NGS-2017-reg-wk7-12MinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2017W7_12XdfFileName))\ny2017W13_17XdfFileName <- file.path(\"/opt/nfl\", \"NGS-2017-reg-wk13-17MinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2017W13_17XdfFileName))\ny2017PostXdfFileName <- file.path(\"/opt/nfl\", \"NGS-2017-postMinMaxXY.xdf\")\nresults <- rbind(results, rxImport(y2017PostXdfFileName))\n\nresultsXdfFileName <- file.path(\"/opt/nfl/CombinedMinMaxXY.xdf\")\nrxImport(results, outFile = resultsXdfFileName)\nresultsCsvFileName <- file.path(\"/opt/nfl/CombinedMinMaxXY.csv\")\nrxXdfToText(inFile = resultsXdfFileName, outFile = resultsCsvFileName, overwrite = TRUE)\n```"},{"metadata":{"trusted":true,"_uuid":"89f59376a79eb0f3fda9c11c2793c932c0181649"},"cell_type":"code","source":"gp_min_max <- read.csv(\"../input/ngs-combined-min-max-x-y/CombinedMinMaxXY.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"69bbe17f6b3cf8bd57e87f2494a53220b036c939"},"cell_type":"code","source":"library(dplyr)\ngp_min_max <- gp_min_max %>% mutate(xRange = max_x - min_x, yRange = max_y - min_y)\ngp_min_max <- gp_min_max %>% mutate(xyArea = xRange * yRange)\npunt_play_entropy <- left_join(punt_play_entropy, gp_min_max, by = c(\"GameKey\", \"PlayID\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"e585ecd208260196a22545908aedf46451ee7b89"},"cell_type":"code","source":"# Create a plot Total Entropy vs. XY Area\nlibrary(ggplot2)\nggplot(punt_play_entropy, aes(xyArea, totEntropy.Play)) +\n    geom_point(shape=1)  +\n    geom_smooth(method = lm) +\n    ggtitle(\"Total Entropy By XY Area\") +\n  xlab(\"XY Area\") + ylab(\"Total Entropy\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a61bc8244f7eebdcf097583aae5b4d98472c34b"},"cell_type":"markdown","source":"## Analysis ##\nThere is a clear correlation between XY Area (the total area of the field used by a given play) and the Total Entropy of that play."},{"metadata":{"trusted":true,"_uuid":"8712003a1ac532c50eb81e19ca9e6a084e4b3378"},"cell_type":"code","source":"# There are no concussions on out of bounds plays\nout_of_bounds_plays <- play_information[grepl(\"\\\\<out of bounds\\\\>\", play_information$PlayDescription, ignore.case = T),c(3,6)]\nout_of_bounds_plays$Out_Of_Bounds = TRUE\npunt_play_entropy <- left_join(punt_play_entropy, out_of_bounds_plays, by = c(\"GameKey\", \"PlayID\"))\nconcussions_out_of_bounds <- punt_play_entropy %>% filter(Out_Of_Bounds == TRUE & Concussion == TRUE)\nnrow(concussions_out_of_bounds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"76393bf6763a10ae003f02c480e73bf27f825c1d"},"cell_type":"markdown","source":"## Analysis ##\nThere are no concussions on out of bounds plays, which re-inforces the idea that reducing the XY Area will reduce the likelihood of concussions."},{"metadata":{"trusted":true,"_uuid":"d900a0f461b1c3b940abff9b3fc4ccce09297ec5"},"cell_type":"code","source":"# Review the statistical summary of the X and Y range\nsummary(punt_play_entropy$xRange)\nsummary(punt_play_entropy$yRange)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b8f2d3ce971009daa5af706a0e38219a437f9b83"},"cell_type":"markdown","source":"## Analysis ##\nThe mean X Range is 69.49 yards.  The length of the field is 120 yards.  Therefore, on average a punt play uses just over half of the field.  There is no reasonable way to constrain the X range, since we cannot add a reasonable rule to limit the distance of punts.\n\nThe mean Y Range is 58.52 yards.  The width of the field is 53.3 yards.  This means players generally use all of the width of the field, plus 5 yards of the sidelines.\n\nTo reduce the XY area, we must reduce the width of the field for punts (and kickoff returns)"},{"metadata":{"_uuid":"416e82293f18c3bb551d5823b09e8f712d336d5d"},"cell_type":"markdown","source":"## Proposed Rule Change ##\nAdd two additional sidelines (of a different color) which are observed as the sidelines only on punt returns (and kick-offs).\n\n![Proposed Solution](https://legalbiguy.com/wp-content/uploads/2019/01/ProposedSolution.jpg)\n\nThis rule changes has the following advantages:\n1.  Reduces the total XY area, and therefore total player entropy, and therefore the likelihood of concussions.\n2. Does not fundamentally alter the nature of these plays.  There can still be blocked punts, run backs, on-side kicks for kick-offs, etc.  This rule change does not affect the excitement or athleticism of these plays\n3.  Does not constrain coaching.  Coaches can develop special teams strategies, formations, etc. for the new Y dimension.\n4.  It is very easy to implement.  Players, coaches, referees and fans will rapidly be able to incorporate the change into the game.\n"}],"metadata":{"celltoolbar":"Raw Cell Format","kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}