{"cells":[{"metadata":{},"cell_type":"markdown","source":"I am new to data science :) Happy to learn from mistakes. Please let me know your thoughts.\n\n1. Let's get the libraries required for the kernel\n2. Read the injury records file\n3. Check the summary, structure, and few rows of the injury records\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"library(ggplot2)\nlibrary(dplyr)\nlibrary(nnet)\nlibrary(scales)\nlibrary(MASS)\nlibrary(caret)\nlibrary(car)\nlibrary(DMwR)\nlibrary(glmnet)\nlibrary(readr)\nlibrary(Hmisc)\ninjury_records = read.csv(\"../input/nfl-playing-surface-analytics/InjuryRecord.csv\")\nsummary(injury_records)\nstr(injury_records)\nhead(injury_records,20)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Key takes aways from summary of injury records:\n\n* There are total of 105 rows with 9 columns.\n* Synthetic surface observations are slightly more than natural surface i.e. 57 vs 48. \n* Knee and Ankle injuries more compared to Toes and Heel. We may want to find the relation between surface and injury body part next.\n* But, before evaluating injury body part vs surface, let's work around 'days missed due to injury'\n* DM_ columns indicates number of days missed due to injury, which indirectly potrays severity of the injury. Let's add these columns.\n* Also, Heel, Toes, and foot injuries are few in number compared to rest. Let's treat them as 'other' "},{"metadata":{"trusted":true},"cell_type":"code","source":"injury_records$severity = injury_records$DM_M1+injury_records$DM_M28+injury_records$DM_M7+injury_records$DM_M42\ncolnames(injury_records)\ninjury_records = injury_records[,-c(6,7,8,9)]\ninjury_records$BodyPart = as.character(injury_records$BodyPart)\nfor(i in 1:nrow(injury_records)){\n  if(injury_records$BodyPart[i]==\"Foot\"||injury_records$BodyPart[i]==\"Heel\"||injury_records$BodyPart[i]==\"Toes\"){\n    injury_records$BodyPart[i]=\"Other\"\n  }\n}\ninjury_records$BodyPart = as.factor(injury_records$BodyPart)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's conduct fisher test to understand relation between surface and body part.\nNull-hypothesis of fisher test is that the variables are independent from each other."},{"metadata":{"trusted":true},"cell_type":"code","source":"fisher.test(injury_records$BodyPart,injury_records$Surface)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Considering a significant level of 0.05, the test failed to reject null hypothesis (based on p-value)\n* Therefore, surface and injurity type(BodyPart) are indepedent variables.\n\nHowever, surface can be a factor in severity of injuries. Can it be ? Let's check as we progress further in this kernel \nAt this point, let's examine play list by merging injury data with it. \n\nShall we start with visualization ?"},{"metadata":{"trusted":true},"cell_type":"code","source":"injury_records %>%\n  ggplot(aes(x=BodyPart, fill = Surface, colour = Surface)) + \n  geom_bar(stat = \"count\", position = \"fill\")\ninjury_records %>%\n  ggplot(aes(x=severity, fill = Surface, colour = Surface)) + \n  geom_bar(stat = \"count\", position = \"fill\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* The number injuries on synthetic surface are bit more compared to natural surface.\n* Importantly, synthetic surface has greater share at severity levels 1,3 and 4. Especially, severity level 3. Interesting insight.\n\nLet's seperate all the three severity levels and plot"},{"metadata":{"trusted":true},"cell_type":"code","source":"severity_1 = injury_records %>% filter(severity==1)\nseverity_2 = injury_records %>% filter(severity==2)\nseverity_3 = injury_records %>% filter(severity==3)\nseverity_4 = injury_records %>% filter(severity==4)\nseverity_1 %>%\nggplot(aes(x=BodyPart, fill = Surface, colour = Surface)) + \n  geom_bar(stat = \"count\", position = \"fill\")\nseverity_2 %>%\n  ggplot(aes(x=BodyPart, fill = Surface, colour = Surface)) + \n  geom_bar(stat = \"count\", position = \"fill\")\nseverity_3 %>%\n  ggplot(aes(x=BodyPart, fill = Surface, colour = Surface)) + \n  geom_bar(stat = \"count\", position = \"fill\")\nseverity_4 %>%\n  ggplot(aes(x=BodyPart, fill = Surface, colour = Surface)) + \n  geom_bar(stat = \"count\", position = \"fill\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Great! our guess looks like correct. Synthetic surface has strong influence at all the severity levels.\n\nNow, let's get the player list file and merge with injury records. Some more interesting factors contributing to the severity of injuries may appear? who knows.. let's check"},{"metadata":{"trusted":true},"cell_type":"code","source":"play_list = read.csv(\"../input/nfl-playing-surface-analytics/PlayList.csv\")\nsummary(play_list)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Temperature has some influencial points/outliers. Let's check."},{"metadata":{"trusted":true},"cell_type":"code","source":"plot(play_list$Temperature)\nnrow(play_list)\nnrow(subset(play_list,play_list$Temperature<0))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's remove the temperature outliers. Also, weather and stadium type contains duplicates. Therefore, let's read manually created data to simplify the weather and stadium type. Furthermore, let's remove the empty playkey rows from injury records and then merge with playlist.\n\nLet's create two new datasets via merge. 1: 'play_list_full': full joint merge which includes all the data points in play list and injury records. 2: 'play_list_injured': inner joint merge which is an intersection of play list and injury records."},{"metadata":{"trusted":true},"cell_type":"code","source":"play_list = subset(play_list,play_list$Temperature>0)\nsimplified_weather = data.frame(read.csv(\"../input/weatherpoints/weatherpoints.csv\"))\nsummary(simplified_weather)\nplay_list$Weather = as.character(play_list$Weather)\nsimplified_weather$x = as.character(simplified_weather$x)\nsimplified_weather$simplified = as.character(simplified_weather$simplified)\nsimplified_stadium_type = data.frame(read.csv(\"../input/stadiumtype/stadiumtype.csv\"))\nsummary(simplified_stadium_type)\nplay_list$StadiumType = as.character(play_list$StadiumType)\nsimplified_stadium_type$x = as.character(simplified_stadium_type$x)\nsimplified_stadium_type$simplified = as.character(simplified_stadium_type$simplified)\nsum(is.na(play_list$Weather))\nsum(is.na(simplified_weather$weather))\nfor (i in 1:nrow(play_list)){\n  j=1\n  for(j in 1:nrow(simplified_weather)){\n    if(play_list$Weather[i]==simplified_weather$x[j]){\n      play_list$Weather[i]=simplified_weather$simplified[j]\n    }\n  }\n  \n  for(j in 1:nrow(simplified_stadium_type)){\n    if(play_list$StadiumType[i]==simplified_stadium_type$x[j]){\n      play_list$StadiumType[i]=simplified_stadium_type$simplified[j]\n    }\n  }\n}\ninjury_records_final = injury_records %>% filter(!PlayKey==\"\")\ncolnames(injury_records_final)[1] = \"PlayerKey\"\ncolnames(injury_records_final)\ncolnames(play_list)[8] = \"Surface\"\ncolnames(play_list)\nplay_list_full = merge(x=play_list,y=injury_records_final,by = c(\"PlayKey\",\"GameID\",\"PlayerKey\",\"Surface\"), all = TRUE)\nplay_list_injured = merge(x=play_list,y=injury_records_final,by = c(\"PlayKey\",\"GameID\",\"PlayerKey\",\"Surface\"), all = FALSE)\nsummary(play_list_injured)\nsummary(play_list_full)\nplay_list_full$BodyPart =as.character(play_list_full$BodyPart)\nplay_list_full$severity =as.numeric(play_list_full$severity)\nplay_list_full$BodyPart = ifelse(is.na(play_list_full$BodyPart),\"No\",play_list_full$BodyPart)\nplay_list_full$severity =ifelse(is.na(play_list_full$severity),0,play_list_full$severity)\nplay_list_full$BodyPart = as.factor(play_list_full$BodyPart)\nsummary(play_list_full)\nplay_list_full = na.omit(play_list_full)\nstr(play_list_full)\nplay_list_full$PlayerKey = as.factor(play_list_full$PlayerKey)\nplay_list_full$severity = as.numeric(play_list_full$severity)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Excellent. We have combined the injury records and playlist.\n\nKey take-aways so far:\n* Surface isn't a siginificant contributing factor for body part (based on fisher's test)\n* However, when we look at surface role in serverity, Synthetic has upper hand over Natural. \n* At this stage, let's create a regression model for severity as there many factors involved. The goal of regression models is to understand/indentify key factors/variables behind severity of injuries"},{"metadata":{},"cell_type":"markdown","source":"**Model building**\n\nWe may need dimentionality reduction technique here as there are more number of variable involved. However, before that, let's create a simple linear regression model. As you already know by know, severity is our target/interested variable."},{"metadata":{"trusted":true},"cell_type":"code","source":"data = play_list_full[,c(4:16)]\ncolnames(data)\nstr(data)\nmodel_injury_lm = lm(severity~.,data = data)\nsummary(model_injury_lm)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Wow! Significant model with good adjusted R-squared value. We have some significant factors such as stadium type based on p-value in the model. However, let's conduct PCA (dimentionality reduction technique) on the numerical data."},{"metadata":{"trusted":true},"cell_type":"code","source":"data_numerical = data[,c(3,4,6,9)]\nsummary(data_numerical)\ndata_categorical = data[,-c(3,4,6,9,13)]\nseverity = data[,c(13)]\ndata_numerical_obj =preProcess(data_numerical,method = \"scale\")\ndata_numerical=predict(data_numerical_obj,data_numerical)\npca = princomp(data_numerical)\ncolnames(data_categorical)\nsummary(pca)\nplot(pca)\np_variance_explained = pca$sdev^2 / sum(pca$sdev^2)\nbarplot(100*p_variance_explained, las=2, xlab='', ylab='% Variance Explained')\npca_data_numerical = predict(pca, data_numerical)\npca_data =data.frame(pca_data_numerical[,c(1:3)],data_categorical,severity)\nstr(pca_data)\ncolnames(pca_data)\ncolSums(is.na(play_list_full))\nmodel_injury = lm(severity~.,data = pca_data)\nsummary(model_injury)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Principal components looks have no siginficance. Let's run StepAIC(to choose better variables) on this model:"},{"metadata":{"trusted":true},"cell_type":"code","source":"step = stepAIC(model_injury,direction = \"both\")\nmodel_injury_simplified =lm(severity~Surface + StadiumType + BodyPart +Weather,data = data)\nsummary(model_injury_simplified)\nmodel_injury_simplified_preds=model_injury_simplified$fitted.values\nregr.eval(trues = data$severity, preds = model_injury_simplified_preds)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Key takeaways:\n* Model is significant. As mentioned in the beginning, the model is intended to understand the impact of variables on severity\n* Key variables based on p-values: Surface, Stadiumtype,BodyPart, and Weather \n* Quite an interesting factors! aren't they ?\n* Let's retain these variables and remove the remaining variables"},{"metadata":{"trusted":true},"cell_type":"code","source":"colnames(data)\ndata_new = data[,c(1,5,7,13)]\nsummary(data_new)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, let's study 'play_list_injured' data. The inner joint data set we created during the merge process between injury records and play list."},{"metadata":{"trusted":true},"cell_type":"code","source":"str(play_list_injured)\nsummary(play_list_injured)\npar(mfrow=c(2,3))\nplot(play_list_injured$Temperature,play_list_injured$severity, xlim = c(20,100),main=\"severity vs temp\")\nplot(play_list_injured$PlayerDay,play_list_injured$severity,main=\"severity vs PlayerDay\")\nplot(play_list_injured$PlayerGame,play_list_injured$severity,main=\"severity vs PlayerGame\")\nplot(play_list_injured$Temperature,play_list_injured$BodyPart, xlim = c(20,100), main=\"BodyPart vs temp\")\nplot(play_list_injured$PlayerDay,play_list_injured$BodyPart,main=\"BodyPart vs PlayerDay\")\nplot(play_list_injured$PlayerGame,play_list_injured$BodyPart,main=\"BodyPart vs playerGame\")\nstr(play_list_injured)\nplay_list_injured$severity = as.numeric(play_list_injured$severity)\nplay_list_injured$PlayerKey = as.integer(play_list_injured$PlayerKey)\ncolSums(is.na(play_list_injured))\ncolnames(play_list_injured)\ndata_injured = play_list_injured[,-c(1,2,3,5,8,11,14)]\nsummary(data_injured)\ncolnames(is.na(data_injured))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Key takeaways:Differences can be observed at:\n* Temparature between 60 to 80\n* PlayerDay<100 vs PlayerDay>300\n* PlayerGame<15 and PlayerGame>15\n\nLet's create model for this data to further understand the fsctors/variables influencing severity of injuries"},{"metadata":{"trusted":true},"cell_type":"code","source":"data_injured = play_list_injured[,-c(1,2,3,5,8,11,14)]\nstr(data_injured)\nmodel_injured_lm = lm(severity~.,data=data_injured)\nsummary(model_injured_lm)\nvif(model_injured_lm)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model_injured_lm2 = lm(severity~Surface + Position + BodyPart,data=data_injured)\nsummary(model_injured_lm2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Model is significant. Key factors based on p-values: Surface, Bodypart, and Position."},{"metadata":{},"cell_type":"markdown","source":"Summary:\n\nSurface, Stadiumtype, BodyPart and Weather are significant factors behind severity of injuries in a combined dataset of injured and non-injured players.\n\nSurface, BodyPart, and Position are significant factors behind severity of injuries among injured players. \n\n\nI had to pause/stop the data analysis at this point. With limited machine power, reading/manipulating/analyzing the player track data was complicated because of the huge data size. Also, speed and direction are some of the factors which may not be regulated in a given match? \n\nThank you. All the kernels in the competition are amazing. Kaggle is helping me alot to learn data science."}],"metadata":{"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}