---
title: "Injury Risk Model for NFL Non-Contact Injuries"
output:
    html_document:
    code_folding: hide
    df_print: paged
    fig_height: 8
    fig_width: 11
    highlight: breezedark
    number_sections: yes
    theme: journal
    toc: yes
    toc_depth: 3
---
```{r Libraries, include = FALSE}
library(rmarkdown)
library(dplyr)
library(caret)
```
# Other Notebooks

This analysis is split between four notebooks. The other three notebooks can be found at the following links:

[Exploratory Analysis of NFL Non-Contact Injuries](https://www.kaggle.com/erinpsajdl/exploratory-analysis-of-nfl-non-contact-injuries) \
This notebook takes an in-depth look at the variables provided in the given datasets and how they intereact with the others.

[Player Movement Effects](https://www.kaggle.com/erinpsajdl/player-movement-effects) \
This notebook takes a look at the effects that various game scenarios have on overall player movement.

[Player Movement Patterns](https://www.kaggle.com/erinpsajdl/player-movement-patterns) \ 
This notebook looks at the player movement patterns for injured players in various scenarios.


# Injury Risk Model
    
Now for the interesting part: the injury risk model! 
    
For the sake of saving memory space, I manipulated these files locally on my own R software. I cleaned the Weather and Stadium Type variables, and split the data into a training data set and a testing data set, to build the model and test the model, respectively. 

```{r bring back playlist injury record and track data, include = FALSE}
down_train <- data.table::fread("../input/injuryrisk/down_train.csv",stringsAsFactors = F)
testData <- data.table::fread("../input/testdata/testData.csv", stringsAsFactors = F)
```

``` {r building logistic model, warning = FALSE}
    # Building The Logistic Regression Model
    logitmod <- glm(Injury ~ FieldType + StadiumType + RosterPosition + Weather + Average_Speed + Max_Speed, family = "binomial", data = down_train)
    summary(logitmod)
    rm(down_train)
```
``` {r predict on test data}
    # Predict on Test Data Set
    pred <- predict(logitmod, newdata = testData, type = "response")
    rm(logitmod)
    y_pred_num <- ifelse(pred > 0.5, 1,0)
    y_pred <- factor(y_pred_num, levels = c(0,1))
    y_act <- testData$Injury
```
``` {r compute accuracy}
    # compute accuracy
    mean(y_pred == y_act)
        # 66.72% accurate
```

The above model is 66.72% accurate at predicting injuries with the following variables:

- Field Type
- Stadium Type
- Roster Position
- Weather
- Average Speed on Play
- Maximum Speed on Play