{"cells":[{"metadata":{},"cell_type":"markdown","source":"<font size=\"+3\" color=\"black\"><strong> Injury Analysis and Training with H2O AutoML (PART I)</strong></font>"},{"metadata":{},"cell_type":"markdown","source":"![H20AutoML](https://miro.medium.com/max/543/1*S6kE_nwoge5m7ok1onsjsQ.png)"},{"metadata":{},"cell_type":"markdown","source":"> **Introduction**\n\n(PART I) \nThis Notebook is all about NFL Injuries Analysis and training. In 1st part we will conduct a detailed analysis by joining the data in one big dataset and will select only necessary pieces for our trainging in H2O AutoML.\n\n\n**So, basically we have 3 csv data files with data related to:**\n*  Injury Record: The injury record file in .csv format contains information on 105 lower-limb injuries that occurred during regular season games over the two seasons. Injuries can be linked to specific records in a player history using the PlayerKey, GameID, and PlayKey fields.\n\n* Play List: – The play list file contains the details for the 267,005 player-plays that make up the dataset. Each play is indexed by PlayerKey, GameID, and PlayKey fields. Details about the game and play include the player’s assigned roster position, stadium type, field type, weather, play type, position for the play, and position group.\n\n* Player Track Data (76M rows): player level data that describes the location, orientation, speed, and direction of each player during a play recorded at 10 Hz (i.e. 10 observations recorded per second).\n\n**Next Steps:**\n* 1. Data Sources joining and selecting data for further analysis.\n* 2. Data selection based on \"Injury\" dataset.\n* 3. Uploading new data as a separate Kaggle dataset.\n* 4. H2O AutoML data modeling/training and results analysis.\n* 5. Data Export as new Kaggle dataset to be detailed analized in PART II.\n\n"},{"metadata":{},"cell_type":"markdown","source":"![NFL PLayers](https://mikedropsports.com/wp-content/uploads/2019/07/50_2-1-e1564969810435-678x381.png)"},{"metadata":{},"cell_type":"markdown","source":"> **1. Data Sources joining and selecting data for further analysis**\n\nNow we are going to join all data in order to get a dataset with 76M rows of data.\nBasically now we have all Players details, PLays data and Injury data in one place.\n\n**Player's with injuries and without**\n\nAs we know all injuries are related to players and way of playing (positions).\nI selected ONLY players wich encounter at least 1 injury and dataset has been reduced to 25M rows.\nIn fact we removed all players without injuries and we can consider that probability is very low to get injuried.\nThere is a huge spectre of injuries not present in our dataset, so basically we will consider only these selected by the organizers.\n\n**Type of the users or Position Group spplit**\n\nAssumption: For different type of players there could be different types and frequence of injuries that might occur. Basically I selected the fastest player types and the players which interact mostly at the scrimmage line. In this case we selected 3 major types of players by Position Group:\n\n* DB (8M rows): defensive backs (DBs) are the players on the defensive team who take positions somewhat back from the line of scrimmage; they are distinguished from the defensive line players and linebackers, who take positions directly behind or close to the line of scrimmage.(Wikipedia)\n* LB (5M rows): A linebacker (LB or backer) is a playing position in gridiron football. Linebackers are members of the defensive team, and line up approximately three to five yards (4 m) behind the line of scrimmage, behind the defensive linemen, and therefore \"back up the line\". Linebackers generally align themselves before the ball is snapped by standing upright in a \"two-point stance\" (as opposed to the defensive linemen, who put one or two hands on the ground for a \"three-point stance\" or \"four-point stance\" before the ball is snapped).(Wikipedia)\n* WR (5M rows):A wide receiver, also referred to as wideouts or simply receivers, is an offensive position in gridiron football, and is a key player. They get their name because they are split out \"wide\" (near the sidelines), farthest away from the rest of the team. Wide receivers are among the fastest players on the field.(Wikipedia)\n\n\n**Field Type Importance of Natural and Synthetic turf**\nIn order to train data and to a comparison analysis we need to dublicate each dataset and \"flip\" field type Natural with Synthetic turf and viceverca. \n\n**Main idea is to train injury data and after to apply the model to normal test data and with \"flipped\" field type **\nIdea fix: All players have beed trained all their non professional life or school mostly on Natural turf.\n\n"},{"metadata":{},"cell_type":"markdown","source":"**Train Injury data selection**\n\nFor the Training dataset we selected all rows with injuries (22915 rows). Also to extend the data set we need to include data without injuries. After many days of analysis and logic deduction I found out that the perfect data to add will be each full Play without injuries played before the Play with injury. This will give us the opportunity to understand also what might be the causes of the injury in the next play. Final Trainig data - 31035 rows and in adition to Body Part injury I added NoInjury value.\n\n"},{"metadata":{},"cell_type":"markdown","source":"> **3. Uploading new data as a separate Kaggle dataset.**\n\nAll final data uploaded as Kaggel public dataset: [nfl-analysis-training-data](https://www.kaggle.com/zinovadr/nfl-analysis-training-data)\n"},{"metadata":{},"cell_type":"markdown","source":"> ** 4. H2O AutoML data modeling/training and results analysis.**\n\nNow the magic part of the analysis and training wiht H2O AutoML.\nSteps:\n* Train on \"Train.csv\" data.\n* Predict \"Body Part\" for 6 uploaded test files.\n* Summary charts on results.\n* Export Prediction to new Kaggle dataset.\n"},{"metadata":{},"cell_type":"markdown","source":""},{"metadata":{"trusted":true},"cell_type":"code","source":"library(tidyverse) # metapackage with lots of helpful functions\nlist.files(path = \"../input/nfl-analysis-training-data\")\n# 'TestDB.csv' 'TestDBInv.csv' 'TestLB.csv' 'TestLBInv.csv' 'TestWR.csv' 'TestWRInv.csv' 'Train.csv' \n\nlibrary(h2o)\nh2o.init()\n\n# Import a sample binary outcome train/test set into H2O\ninj <- h2o.importFile(\"../input/nfl-analysis-training-data/Train.csv\")\n\n\ninj$PlayKeyIN<-NULL\n#removing this column - too manu NUll and will confuse the results\n\ninj.split <- h2o.splitFrame(data = inj,ratios = 0.8, seed = 1234)\n\ntrain <- inj.split[[1]]\nvalid <- inj.split[[2]]\n\n\n# Identify predictors and response\ny <- \"Body Part\"\nx <- setdiff(names(train), y)\n\n# For binary classification, response should be a factor\ntrain[,y] <- as.factor(train[,y])\nvalid[,y] <- as.factor(valid[,y])\n\n\naml <- h2o.automl(x = x, y = y,\n                  training_frame = train,\n                  validation_frame = valid,\n                  max_models = 20,\n                  seed = 1)\n\n# AutoML Leaderboard\nlb <- aml@leaderboard\n\n# Print all rows (instead of default 6 rows)\nprint(lb, n = nrow(lb))\n\n# The leader model is stored here\naml@leader\n#h2o.saveModel(aml)\n\n# If you need to generate predictions on a test set, you can make\n# predictions directly on the `\"H2OAutoML\"` object, or on the leader\n# model object directly\n\n#1# DB train \n\nforpred <- h2o.importFile(\"../input/nfl-analysis-training-data/TestDB.csv\")\nforpred$PlayKeyIN<-NULL\npred <- h2o.predict(aml, forpred)# predict(aml, test) also works\npred_sub <- h2o.cbind(forpred, pred)\nsummary(pred_sub)\nh2o.exportFile(pred_sub, path=\"predDB.csv\")\nforpred2 <- h2o.importFile(\"../input/nfl-analysis-training-data/TestDBInv.csv\")\nforpred2$PlayKeyIN<-NULL\npred2 <- h2o.predict(aml, forpred2)# predict(aml, test) also works\npred_sub2 <- h2o.cbind(forpred2, pred2)\nsummary(pred_sub2)\nh2o.exportFile(pred_sub2, path=\"predDBInv.csv\")\n\n\n#2# LB train \n\n#forpredLB <- h2o.importFile(\"../input/nfl-analysis-training-data/TestLB.csv\")\n#forpredLB$PlayKeyIN<-NULL\n#predLB <- h2o.predict(aml, forpredLB)# predict(aml, test) also works\n#pred_subLB <- h2o.cbind(forpredLB, predLB)\n#summary(pred_subLB)\n#h2o.exportFile(pred_subLB, path=\"predLB.csv\")\n#forpred2LB <- h2o.importFile(\"../input/nfl-analysis-training-data/TestLBInv.csv\")\n#forpred2LB$PlayKeyIN<-NULL\n#pred2LB <- h2o.predict(aml, forpred2LB)# predict(aml, test) also works\n#pred_sub2LB <- h2o.cbind(forpred2LB, pred2LB)\n#summary(pred_sub2LB)\n#h2o.exportFile(pred_sub2LB, path=\"predLBInv.csv\")\n\n\n#3# WR train \n\n#forpredWR <- h2o.importFile(\"../input/nfl-analysis-training-data/TestWR.csv\")\n#forpredWR$PlayKeyIN<-NULL\n#predWR <- h2o.predict(aml, forpredWR)# predict(aml, test) also works\n#pred_subWR <- h2o.cbind(forpredWR, predWR)\n#summary(pred_subWR)\n#h2o.exportFile(pred_subWR, path=\"predWR.csv\")\n#forpred2WR <- h2o.importFile(\"../input/nfl-analysis-training-data/TestWRInv.csv\")\n#forpred2WR$PlayKeyIN<-NULL\n#pred2WR <- h2o.predict(aml, forpred2WR)# predict(aml, test) also works\n#pred_sub2WR <- h2o.cbind(forpred2WR, pred2WR)\n#summary(pred_sub2WR)\n#h2o.exportFile(pred_sub2WR, path=\"predWRInv.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The limit is 5 GB, so all files should be processed individually.\n\n> **Analysis to be extended in PART II**\n"}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}