{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Q1: Cost of Turf Field\n\n## Abstract\n\nIn this notebook, we examine the costs of playing on artificial turf as compared natural turf (grass). Our assumption is that playing on artificial turf causes relatively more injuries than playing on grass. We would like to calculate how many more games players are expected to miss per season by playing on artificial turf as compared to grass, then translate that difference in games played into a dollar amount by using per game player salaries.\n\nOur conclusion is based on NFL teams playing 8 home games and 8 away games, meaning that a team's surface decision will impact their players for 8 games per year.\n\nOur analysis indicates that injuries are about 1.6 times more likely to occur on artifical turf as compared to grass. We find the total cost of having an artifical turf at home to be about $2.95m / season in missed game salaries more than having a grass field at home. "},{"metadata":{},"cell_type":"markdown","source":"## Preparation\n\n### Import data\n\nFor this analysis, we do not need player track data, as do not look at specific intra-play movement."},{"metadata":{"trusted":false},"cell_type":"code","source":"import pandas as pd\nimport scipy.stats as stats\nimport numpy as np\nfrom tqdm import tqdm\nimport statsmodels.api as sm\nimport statsmodels.formula.api as smf\nimport matplotlib.pyplot as plt\n\n\"\"\"\nImport datasets \nhttps://www.kaggle.com/c/nfl-playing-surface-analytics/data\n\"\"\"\n\nKAGGLE = True\nif not KAGGLE:\n    IR_data = pd.read_csv('InjuryRecord.csv')\n    PL_data = pd.read_csv('PlayList.csv')\nelse:\n    IR_data = pd.read_csv('../input/nfl-playing-surface-analytics/InjuryRecord.csv')\n    PL_data = pd.read_csv('../input/nfl-playing-surface-analytics/PlayList.csv')\n\npd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 500)\n\nplt.rcParams['figure.dpi'] = 200\nplt.rcParams.update({'font.size': 14})","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Clean IR data\n\nSome of the injuries do not have a play associated with them. Lets use the last play they played as the injury play. For this analysis, it does not matter which play they were injured on, as our analysis is based on game-level variables."},{"metadata":{"trusted":false},"cell_type":"code","source":"def cleanIRdata(IR_data):\n\n    # Mark whether we need to set a play for each row\n    IR_data['PlayBool'] = [False if pd.isnull(i) else True for i in IR_data['PlayKey']] \n\n    # Get maximum play count for each play\n    IR_data['HighPlay'] = [PL_data[PL_data['GameID'] == i]['PlayerGamePlay'].max() for i in IR_data['GameID']] \n\n    # Create a new playkey\n    PlayKey_adj = []\n    PlayKey = list(IR_data['PlayKey'])\n    HighPlay = list(IR_data['HighPlay'])\n    PlayBool = list(IR_data['PlayBool'])\n    GameID = list(IR_data['GameID'])\n    for i in range(len(IR_data)):\n        if PlayBool[i]:\n            PlayKey_adj.append(PlayKey[i])\n        else:\n            PlayKey_adj.append(GameID[i] + '-' + str(HighPlay[i]))\n\n    IR_data['PlayKey_adj'] = PlayKey_adj\n    return IR_data\n    \nIR_data = cleanIRdata(IR_data)\nIR_data.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Merge data\n\nWe start with two datasets:\n\n- PL_data: lists every player-play for the 2017 and 2018 season for 250 players\n- IR_data: lists every injury at a play level (or if not available, a game level)\n\nWe will merge these datasets to create one dataset with every player play, and an indication of whether that player was injured on that play or not."},{"metadata":{"trusted":false},"cell_type":"code","source":"# Merge data\nPL_data['PlayKey_adj'] = PL_data['PlayKey']  \ndf = pd.merge(PL_data, IR_data, on = 'PlayKey_adj', how='left').fillna(0)\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### In terms of \"games missed\"\n\nOur original dataset tells us whether the player missed 1, 7, 28, or 42 days of gametime. For simplicity, we are going to create a new variable called \"games missed.\" We assume that missing 1 day means that the player got injured during the game and did not return, but was able to return the following week. We therefore assign this as 0.5 games missed. \n\nSince all injuries happened during a game, we assume that an injury results in 0.5 games missed for the game that caused the injury, as well as X number of games missed for the DM/7 following days missed.\n\n- DM_M1   = 0.5 games missed\n- DM_M7\t  = They missed this game and the next game. 1.5 games missed\n- DM_M28  = 4 weeks, so 4.5 games.\n- DM_M42  = 6 weeks, so 6.5 games"},{"metadata":{"trusted":false},"cell_type":"code","source":"# This function calculates the number of games missed and appends it to our dataframe\ndef countGamesMissed(df):\n    \n    GM_array = []\n    for i, row in df.iterrows():\n\n        DM1  = row['DM_M1']\n        DM7  = row['DM_M7']\n        DM28 = row['DM_M28']\n        DM42 = row['DM_M42']\n\n        if DM42 == 1:\n            GM_array.append(6.5)\n        elif DM28 == 1:\n            GM_array.append(4.5)\n        elif DM7 == 1:\n            GM_array.append(1.5)\n        elif DM1 == 1:\n            GM_array.append(.5)\n        else:\n            GM_array.append(0)\n\n    df['GamesMissed'] = GM_array\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"df = countGamesMissed(df)\ndf[df['GamesMissed'] > 0][['PlayKey_adj', 'DM_M1','DM_M7','DM_M28','DM_M42','GamesMissed']].head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data Analysis\n\nWe start by taking a look at our dataset to see how many plays we have, how many injuries we have, and what the breakdown looks like between position group."},{"metadata":{"trusted":false},"cell_type":"code","source":"PGs = ['LB', 'QB', 'DL', 'OL', 'SPEC', 'TE', 'WR', 'RB', 'DB']\ninjurySummary = pd.DataFrame(columns=['PG', 'Plays', 'Injuries'])\nfor i in range(len(PGs)):\n    PG = PGs[i]\n    injurySummary.loc[i] = [PG, len(df[df['PositionGroup'] == PG]), len(df[(df['PositionGroup'] == PG) & (df['DM_M1'] == 1)])]\n    \ninjurySummary","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We do not have any injuries in our dataset for QBs and SPEC -- therefore we will use the average injury rate for our analysis for those position groups."},{"metadata":{},"cell_type":"markdown","source":"### Natural v. Artificial Injury Rates\n\nAt a play level, is there a higher rate of injuries for plays on artifical terf?"},{"metadata":{"trusted":false},"cell_type":"code","source":"syntheticPlays    = df[df['FieldType'] == 'Synthetic']\nsyntheticInjuries = df[(df['DM_M1'] == 1) & (df['FieldType'] == 'Synthetic')]\nnaturalPlays      = df[df['FieldType'] == 'Natural']\nnaturalInjuries   = df[(df['DM_M1'] == 1) & (df['FieldType'] == 'Natural')]\nplays             = df\ninjuries          = df[df['DM_M1'] == 1]\n\nprint('There are', len(IR_data), 'plays in the full dataset')\nprint('There are', len(plays), 'plays in the merged dataset')\nprint(len(injuries), 'of', len(plays), \"(%.4f\" % (len(injuries) / len(plays) * 100) + '%)', 'of plays in the merged dataset are injuries')\nprint(len(syntheticInjuries), 'of', len(syntheticPlays), \"(%.4f\" % (len(syntheticInjuries) / len(syntheticPlays) * 100) + '%)', 'sythetic injuries')\nprint(len(naturalInjuries), 'of', len(naturalPlays), \"(%.4f\" % (len(naturalInjuries) / len(naturalPlays) * 100) + '%)', 'natural injuries')\n\n# P-Test\nnatural_sample = [1] * len(naturalInjuries) + [0] * (len(naturalPlays) - len(naturalInjuries))\nsynthetic_sample = [1] * len(syntheticInjuries) + [0] * (len(syntheticPlays) - len(syntheticInjuries))\nt_stat, p_val = stats.ttest_ind(natural_sample, synthetic_sample, equal_var=False)\nprint('The p-value that synthetic is worse than natural is', \"%.5f\" % p_val)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The answer is yes -- our p-value is significantly lower than 0.05.\n\nLet's run the same analysis at a game level."},{"metadata":{"trusted":false},"cell_type":"code","source":"SYN_INJURIES = len(set(df[(df['DM_M1'] == 1) & (df['FieldType'] == 'Synthetic')]['GameID_x']))\nSYN_GAMES = len(set(df[(df['FieldType'] == 'Synthetic')]['GameID_x']))\nNAT_INJURIES = len(set(df[(df['DM_M1'] == 1) & (df['FieldType'] == 'Natural')]['GameID_x']))\nNAT_GAMES = len(set(df[(df['FieldType'] == 'Natural')]['GameID_x']))\n\nprint(SYN_INJURIES, 'of', SYN_GAMES, \"(%.4f\" % (SYN_INJURIES / SYN_GAMES * 100) + '%)', 'of synthetic games had injuries')\nprint(NAT_INJURIES, 'of', NAT_GAMES, \"(%.4f\" % (NAT_INJURIES / NAT_GAMES * 100) + '%)', 'of natural games had injuries')\nprint(\"%.4f\" %  ((SYN_INJURIES / SYN_GAMES) / (NAT_INJURIES / NAT_GAMES)), 'higher injury rate on turf')\n\n\n# P-Test\nnatural_sample = [1] * NAT_INJURIES + [0] * (NAT_GAMES - NAT_INJURIES)\nsynthetic_sample = [1] * SYN_INJURIES + [0] * (SYN_GAMES - SYN_INJURIES)\nt_stat, p_val = stats.ttest_ind(natural_sample, synthetic_sample, equal_var=False)\nprint('The p-value that synthetic is worse than natural (on a game level) is', \"%.5f\" % p_val)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Taking a similar approach to above, we see that (as expected) synthetic games have a higher rate of injury than natural grass games."},{"metadata":{},"cell_type":"markdown","source":"### Comparison in terms of games missed\n\nOur previous test shows that artifical term is associated with a higher risk of injury. Is one set of injuries worse than an other set in terms of resulting games missed?"},{"metadata":{"trusted":false},"cell_type":"code","source":"syntheticGM       = sum(df[df['FieldType'] == 'Synthetic']['GamesMissed'])\nnaturalGM         = sum(df[df['FieldType'] == 'Natural']['GamesMissed'])\nSYN_INJURY_AVG = round(syntheticGM / SYN_INJURIES,2)\nNAT_INJURY_AVG = round(naturalGM / NAT_INJURIES,2)\nprint('Each synthetic injury averages', SYN_INJURY_AVG, 'games missed')\nprint('Each natural injury averages', NAT_INJURY_AVG, 'games missed')\nprint('Synthetic injuries miss', round(SYN_INJURY_AVG / NAT_INJURY_AVG,3), 'more games')\n\nGM_NAT_LIST = list(df[(df['DM_M1'] == 1) & (df['FieldType'] == 'Natural')]['GamesMissed'])\nGM_SYN_LIST = list(df[(df['DM_M1'] == 1) & (df['FieldType'] == 'Synthetic')]['GamesMissed'])\n\nt_stat, p_val = stats.ttest_ind(GM_NAT_LIST, GM_SYN_LIST, equal_var=False)\nprint('The p-value that synthetic injuries are worse (in terms of games missed) than natural injuries is', \"%.3f\" % p_val)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It appears that synthetic injuries are no more dangerous than injuries that happened on grass."},{"metadata":{},"cell_type":"markdown","source":"## Regression\n\nWe will now run a regression to determine the expected number of games missed per play. We will break players down by their position group to account for signficiantly different injury probabilities for different groups.\n\n### Clean data\n\nTo hopefully mitigate any other explanatory variables, we create a few variables for weather such as\n\n- Is_Rain: whether it was raining or had a chance of rain\n- Is_Snow: whether it was snowing\n- Temperature_Adj: The temperature during the game. This is set as room temperature (72F) for games that happened indoors."},{"metadata":{"scrolled":false,"trusted":false},"cell_type":"code","source":"def convertVariablesForRegression(tmp):\n\n    def isIn(item, word):\n        return item in str(word).lower()\n\n    # Convert some fields to boolean\n    tmp['Is_Synthetic'] = tmp['FieldType'] == 'Synthetic'\n    tmp['Is_Rain'] = [True if isIn('rain', i) or isIn('shower', i) else False for i in tmp['Weather']]\n    tmp['Is_Snow'] = [True if isIn('snow', i) else False for i in tmp['Weather']]\n\n    # Clean some variables\n    tmp['Temperature_Adj'] = [i if i != -999 else 72 for i in tmp['Temperature']]\n    \n    return tmp\n    \ndf = convertVariablesForRegression(df)\n\ndisplay(df[['PlayKey_adj', 'PositionGroup', 'PlayerDay', 'PlayerGamePlay', 'Is_Synthetic', 'Temperature_Adj', 'Is_Rain', \\\n    'Is_Snow', 'GamesMissed']].head())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Run regression\n\nWe can now run a generalized linear model using a Poisson distribution. We want to make a seperate model for each position group.\n\n\nWe have no data for some position groups -- SPEC and QB. We will just use the average value for their position."},{"metadata":{"trusted":false},"cell_type":"code","source":"X = 'GamesMissed'\nY = ['C(PositionGroup) * Is_Synthetic', 'Is_Synthetic', 'PlayerDay', 'PlayerGamePlay', 'Temperature_Adj', 'Is_Rain', 'Is_Snow']\nY = ' + '.join(Y)\n\nmodel = sm.GLM.from_formula(X + ' ~ ' + Y, data=df, family=sm.families.Poisson()).fit()\n# model = smf.ols(X + ' ~ ' + Y, data = df).fit()\nprint(model.summary())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Predictions\n\nWe can use our regression results to predict how many expected games a player in each position group will miss per play on synthetic and natural turf."},{"metadata":{"trusted":false},"cell_type":"code","source":"# Create testing dataframe to predict on\ntd = pd.DataFrame()\nPG = list(set(df['PositionGroup']))\n\n# Get average plays per game by position\nPPG = []\nfor i in PG:\n    POS_PLAYS = len(df[df['PositionGroup'] == i])\n    POS_GAMES = len(set(df[df['PositionGroup'] == i]['GameID_x']))\n    PPG.append(int(POS_PLAYS / POS_GAMES))\n\n# Use mean for each group\ndef getMeanForPG(stat, PG=PG, df=df):\n    return [df[df['PositionGroup'] == pos_group][stat].median() for pos_group in PG]\n\ntd['PositionGroup']     = PG\ntd['Is_Synthetic']      = [False] * len(PG)\ntd['PlayerDay']         = getMeanForPG('PlayerDay')\ntd['PlayerGamePlay']    = getMeanForPG('PlayerGamePlay')\ntd['Temperature_Adj']   = getMeanForPG('Temperature_Adj')\ntd['Is_Rain']           = [False] * len(PG)\ntd['Is_Snow']           = [False] * len(PG)\ntd['PPG']               = PPG\n                    \n# Predict\ntd['GM_Natural']        = model.predict(td)\n\n# Predict on artificial\ntd['Is_Synthetic']      = [True]  * len(PG)\ntd['GM_Synthetic']      = model.predict(td)\n\n# Set average values for QB and SPEC since insufficient data (0 injuries)\nGM_Avg = {}\nfor fieldType in ['Natural', 'Synthetic']:\n    GM_Avg[fieldType] = sum(df[df['FieldType'] == fieldType]['GamesMissed']) / len(df[df['FieldType'] == fieldType])\n    for PG in ['QB', 'SPEC']:\n        td.at[list(td[td['PositionGroup'] == PG].index)[0], 'GM_' + fieldType] = GM_Avg[fieldType]\n\n# Differences\ntd['SyntheticDelta']    = td['GM_Synthetic'] - td['GM_Natural']\ntd['SyntheticRatio']    = td['GM_Synthetic'] / td['GM_Natural']      \n\n# GM / Game\ntd['DeltaGMperGame'] = list(td['SyntheticDelta'] * td['PPG'])\n\n# GM / other units\ntd['GM_Natural_Game'] = list(td['GM_Natural'] * td['PPG'])\ntd['GM_Synthetic_Game'] = list(td['GM_Synthetic'] * td['PPG'])\ntd['GM_Natural_HomeSeason'] = [i*8 for i in list(td['GM_Natural'] * td['PPG'])]\ntd['GM_Synthetic_HomeSeason'] = [i*8 for i in list(td['GM_Synthetic'] * td['PPG'])]  \n        \n# Remove \"missing data\" player group\ntd = td[td['PositionGroup'] != 'Missing Data'].reset_index(drop=True)\n\n# Display\ntd.sort_values('SyntheticRatio', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Results\n\nFrom here, we can see which position groups are most likely to get injured per play and which position groups are more  (or less) likely to get injured on synthetic turf as compared to grass."},{"metadata":{"scrolled":false,"trusted":false},"cell_type":"code","source":"def plotGamesMissedPG(PerGame = False, PerHomeSeason=False):\n\n    # Example from here https://pythonspot.com/matplotlib-bar-chart/\n    fig, ax = plt.subplots(figsize=(10,5))\n    index = np.arange(len(td))\n    bar_width = 0.35\n    opacity = 0.75\n    \n    nat_bars = list(td['GM_Natural'])\n    syn_bars = list(td['GM_Synthetic'])\n    timeframe = 'play'\n        \n    if PerGame:\n        nat_bars = td['GM_Natural_Game']\n        syn_bars = td['GM_Synthetic_Game']\n        timeframe = 'game'\n        \n    if PerHomeSeason:\n        nat_bars = td['GM_Natural_HomeSeason']\n        syn_bars = td['GM_Synthetic_HomeSeason']\n        timeframe = 'home season'\n    \n    \n    plt.bar(\n                x      = index, \n                height = nat_bars, \n                width  = bar_width,\n                alpha  = opacity,\n                color  = '#48BB78', # color from tailwindcss\n                label  = 'Natural'\n           )\n\n    plt.bar(\n                x      = index + bar_width, \n                height = syn_bars, \n                width  = bar_width,\n                alpha  = opacity,\n                color  = '#2B6CB0', # color from tailwindcss\n                label  = 'Synthetic'\n            )\n\n    ax.spines['top'].set_color('white')\n    ax.spines['right'].set_color('white')\n    plt.xlabel('Position')\n    plt.ylabel('Expected Games Missed (per %s)' % timeframe)\n    plt.title('Natural v. Synthetic Turf - Games Missed Due to Injury', fontsize=14, fontweight='bold', pad=20)\n    plt.xticks(index + bar_width, list(td['PositionGroup']))\n    plt.legend()\n    plt.tight_layout()\n    plt.show()\n    \nplotGamesMissedPG(PerGame=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Home-Season Level\n\nLet's make the same chart that shows the expected games missed difference in playing 8 games at home on natural grass compared to 8 games at home on artifical turf."},{"metadata":{"trusted":false},"cell_type":"code","source":"plotGamesMissedPG(PerHomeSeason=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Salary-Injury Dollars\n\nWe have already converted expected missed games per play into expected games missed per season. We can then use that result along with player salary to determine the salary value of the missed games.\n\nFor example, if a QB is paid \\\\$16m / season, or \\\\$1m per game, having 0.5 additional expected games has a salary loss value of \\\\$500,000.\n\nThis somewhat represents the cost of replacing a player, or \"lost productivity cost\" associated with the injury. Of course, there are certainly other non quantifiable costs associated with injuries, such as emotional and distress costs to that player. There are also quantifiable costs, such as healthcare costs, and increased chances of reaggravating the injury.\n\n\n### Methodology\n\nWe start with our first finding of the change in missed games per play.\n\n<br>\n<br>\n$$\\:\\frac{Missed\\:Games\\:Artificial}{Play}\\:- \\:\\frac{Missed\\:Games\\:Grass}{Play}\\:=\\:\\frac{\\Delta Missed\\:Games}{Play}\\:$$\n<br>\n<br>\nWe can then use this $\\frac{\\Delta Missed Games}{Play}$ number to calcuate the effect over a season. As an example, we will look at the WR position group.\n<br>\n<br>\n$$\\:\\frac{\\Delta Missed\\:Games}{Play}*\\:\\frac{Plays}{WR}*\\:\\frac{WR}{Game}*\\:\\frac{\\\\$}{Game}*\\:\\frac{Games}{Season}\\:=\\:\\frac{Cost}{Season}\\:$$\n<br>\n<br>\nWe can reduce simplify this equation a bit by just finding the average number of WRs plays per game, instead of the number of WRs multiplied by average plays \n<br>\n<br>\n$$\\frac{Plays}{WR}*\\:\\frac{WR}{Game}=\\:\\frac{WR\\:Plays}{Game}\\:$$\n<br>\n<br>\nAdditionally, we can just look at $\\frac{Dollar}{Season}$ salary, instead of at a per-game level, resulting in the following equation that calculates the associated injury-cost for the WR group that playing 16 games on artificial turf would as compared to grass\n<br>\n<br>\n$$\\:\\frac{\\Delta Missed\\:Games}{WR\\:Play}*\\:\\:\\frac{WR\\:Plays}{Game}\\:*\\:\\frac{\\$}{Season}\\:=\\:\\frac{\\$}{Season}\\:$$\n<br>\n<br>\nWe then want to convert this value to a per-season-team level. We can do this by summing the $\\frac{Dollar}{Season}$ cost $C$ for each position group\n<br>\n<br>\n$${C}_{LB}\\;+\\;{C}_{QB}\\;+{C}_{DL}\\;+\\;{C}_{OL}\\;+{C}_{SPEC}\\;+\\;{C}_{TE}\\;+\\:{C}_{WR}\\;+\\;{C}_{RB}\\;+\\;{C}_{DB}\\:=\\:{C}_{Team}$$\n<br>\n<br>\nFinally, we can convert the increased injury cost per season into a more meaningful stat, the injury cost for having 8 home games on artificial turf.\n<br>\n<br>\n$$\\:\\:\\frac{{\\$}_{\\:Team}}{Season}\\:*\\:\\frac{8\\:Home\\:Games}{16\\:Games}\\:=\\:\\frac{\\$_{\\:Team\\:Home}}{Season}\\:$$"},{"metadata":{},"cell_type":"markdown","source":"### Static variables\n\nTo convert expected missed games into a dollar amount, we have two variables that we need to calculate at a group level. We'll need to obtain some NFL data from outside our given dataset for this. The variables we need are $\\frac{Plays}{Game}\\:$ and $\\frac{$}{Game}\\:$ at a player group level.\n\n#### Plays Per Game \n\nAs we have expected games missed per play (by position group), we can convert this into expected games missed per game (by position group) by multiplying $\\frac{Missed\\:Games}{Play}\\:*\\frac{Plays}{Game}\\:$\n\nWe want to find how many snaps per game players of each position group play. This is going to be higher for [DBs (defensive backs)](https://en.wikipedia.org/wiki/American_football_positions) as each normal defensive play has 5 DBs, unlike QBs which only have 1 player per play. \n\nWe found snap count data for all players on 8 teams from the 2018 season on Pro-Football-Reference.com. We then collapse then collapsed the table by player position, summing snap counts, and average the data out to find the per game / per team snap count per position group.\n<br><br>\n*Sports Reference LLC. [\"New England Patriots Snap.\"](https://www.pro-football-reference.com/teams/nwe/2018-snap-counts.htm#snap_counts::none) Pro-Football-Reference.com - Pro Football Statistics and History. [PRF](https://www.pro-football-reference.com/). Dec 25, 2019. [Citation Link](https://www.pro-football-reference.com/about/contact.htm)*"},{"metadata":{"trusted":false},"cell_type":"code","source":"# Read in all snap data\nteams = ['car', 'cle', 'crd', 'gnb', 'nwe', 'pit', 'tam', 'was']\n\nif KAGGLE:\n    PATH = '/kaggle/input/nflpfrdata/PFR_data/'\nelse:\n    PATH = 'PFR_DATA/'\n\nsnaps = pd.read_csv(PATH + 'snap_count/pit.csv', header=1)\nfor team in teams[1:]:\n    teamSnaps = pd.read_csv(PATH + '/snap_count/%s.csv' % team, header=1)\n    snaps = pd.concat([snaps, teamSnaps])\n    \nsnaps[['Player', 'Pos','Num']].head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Data Cleaning\n\nThe position group labels are more granular then the groups we want to work with. So we have to manually map the specific positions to their position groups "},{"metadata":{"trusted":false},"cell_type":"code","source":"# We need to conert all positions to our position groups\n# https://en.wikipedia.org/wiki/American_football_positions\npositionGroupConvert = {\n    'C'    : 'OL', # C\n    'CB'   : 'DB', # C\n    'CBDB' : 'DB', # C\n    'DB'   : 'DB', # C\n    'DE'   : 'DL', # Could also be LB\n    'DT'   : 'DL', # C\n    'FB'   : 'RB', # C\n    'FS'   : 'DB', # C\n    'FSS'  : 'DB', # C\n    'FSSS' : 'DB', # C\n    'FSSSS': 'DB', # C\n    'G'    : 'OL', # C\n    'K'    : 'SPEC', # c\n    'LB'   : 'LB', # C\n    'LS'   : 'SPEC', # Long snapper\n    'NT'   : 'DL', # C\n    'P'    : 'SPEC', # C\n    'QB'   : 'QB', # C\n    'RB'   : 'RB', # C\n    'S'    : 'DB', # C\n    'SS'   : 'DB', # C\n    'SSS'  : 'DB', # C\n    'T'    : 'OL', # C\n    'TE'   : 'TE', # C\n    'WR'   : 'WR', # C\n    \n    # Extras from salary dataset\n    'DL'   : 'DL', # C\n    'EDGE' : 'DL', # Same as defensive end\n    'HB'   : 'RB', # Halfback\n    'ILB'  : 'LB', # C\n    'LB-DE': 'DL', # Suggs (DE)\n    'LG'   : 'OL', # C\n    'LT'   : 'OL', #C\n    'NT'   : 'DL', # Nose tackle, this is center on defense\n    'OG'   : 'OL', # C\n    'OL'   : 'OL', # C\n    'OLB'  : 'LB', # C \n    'OT'   : 'OL', # C \n    'QB/TE': 'TE', # Logan Thomas (TE)\n    'RB-WR': 'RB', # Ty Montgomery (RB)\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Findings\n\nAfter summing the table and averaging it out to a per game level, we were able to find the number of position groups snaps per game. The numbers intuitively make sense -- there are about 65-70 offensive players per team per game, and for QB we have 67 team snaps per game."},{"metadata":{"trusted":false},"cell_type":"code","source":"# Group and count\nsnaps['PositionGroup'] = [positionGroupConvert[i] for i in snaps['Pos']]\nsnaps.rename(columns={'Num': 'Off', 'Num.1': 'Def', 'Num.2':'Spec'}, inplace=True)\nsnaps['Snaps'] = snaps['Off'] + snaps['Def'] + snaps['Spec']\nsnaps[['Pos', 'Snaps', 'PositionGroup']]\nsnaps['TeamSnapsPerGame'] = round(snaps['Snaps'] / 16 / 8, 2)\ngroupedSnaps = snaps.groupby('PositionGroup').sum()\ngroupedSnaps","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Salary data\n\nWe want to calculate the average salary for a player in each position group play. More accurately, we want a weighted average salary by position group snap count. Since higher paid players are more likely to play more snaps per game (and therefore more likely to get injured), we don't really want an average of all players in a position group, but rather we want to weight it towards players that are playing significant snap counts.\n\nTo further clarify, a team could have 3 quarterbacks, paid \\\\$20m, \\\\$1m, and \\\\$1m each, but playing 90\\%-9\\%-1\\% snaps. If our model showed that QBs miss one game per average per season, the correct calulation would be closer to \\$20m * 1/16 as compared to (\\\\$20m + \\\\$1m + \\\\$1m)/3 * 1/16.\n\n#### Dataset\n\nWe found salary data on the 2019 season from Pro-Football-Reference.com. \n\nObservations:\n- Data is from 2019, while we would ideally want to average 2017-2018 data\n- Salary does does not appear to include signing bonuses: Le'Veon Bell is listed at \\\\$2m\n\n*Sports Reference LLC. [\"NFL 2019 Player Salaries.\"](https://www.pro-football-reference.com/players/salary.htm) Pro-Football-Reference.com - Pro Football Statistics and History. [PRF](https://www.pro-football-reference.com/). (Dec 25, 2019). [Citation Link](https://www.pro-football-reference.com/about/contact.htm)*"},{"metadata":{"trusted":false},"cell_type":"code","source":"# Load data, preview\nif KAGGLE:\n    PATH = '/kaggle/input/nflpfrdata/PFR_data/'\nelse:\n    PATH = 'PFR_DATA/'\nsalary = pd.read_csv(PATH + 'salary/salary.csv', header=0)\nsalary['PositionGroup'] = [i if pd.isnull(i) else positionGroupConvert[i] for i in salary['Pos']]\nsalary['Salary'] = [int(i.replace('$', '')) for i in salary['Salary']]\nsalary[['Player', 'Salary', 'PositionGroup']].head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Average position group salary \n\nWe want to calculate the average position group salary for players who are playing meaningful snap counts.\n\nFor QBs, we assume that about 32 different players will play 80%+ of all QB snaps in a season. This would be 100% if QBs didn't get injured or benched. Taking the average salary of the top 32 QBs would probably be appropriate -- or at least relatively close to the average position group salary.\n\nHowever, a team might have 4 WRs on the field at one time. So it wouldn't really make sense to take the average salary of the top 32 WRs -- that would likely be a top 20% salary for starting WRs. If indeed teams do play an average of 4 WRs, we could reasonably then take the average of the top 32 * 4 highest paid WRs.\n\nThis means we first have to calculate how many players are played at each position group. What we can do is look at the total snap counts on offense, defense, and special teams to see what percent each position group had. We can then look at what proportion of players on the field were in that position group, and multiply that by 11 to find how many players are typically on the field from that group at one time"},{"metadata":{"trusted":false},"cell_type":"code","source":"# Calculate Number of Men from each group on field\nfor i in ['Off', 'Def', 'Spec']:\n    groupedSnaps[i + '_men'] = round(11 * groupedSnaps[i] / sum(groupedSnaps[i]), 2)\n    \ngroupedSnaps['Men'] = groupedSnaps[['Off_men','Def_men','Spec_men']].max(axis=1)\ngroupedSnaps","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This tells us that the average number of men in each group on average. It makes sense that QB is always 1, WR is about 2.5 and OL is about 5.\n\nWe can now use these numbers to calculate how many top salaries we should look at in each group. To do this, we will take 32 teams * the average number of players that play in that group (i.e. 'Men')."},{"metadata":{"trusted":false},"cell_type":"code","source":"# Calculate average position salary\npositionSalary = []\nfor pos in list(groupedSnaps.index):\n    numPosition = int(32 * groupedSnaps['Men'][pos])\n    positionSalary.append(salary[salary['PositionGroup'] == pos][0:numPosition]['Salary'].mean())\n    \ngroupedSnaps['Salary'] = positionSalary\ngroupedSnaps['SalaryFormatted'] = ['${:,.2f}m'.format(i/1000000) for i in positionSalary]\n\ngroupedSnaps.sort_values(by=['Salary'], ascending=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"These results show that QB have the highest average salary while SPEC (punters, kickers) have the lowest."},{"metadata":{},"cell_type":"markdown","source":"### Calculations\n\nWe now have all the values needed to make our calculations. To review, we have the following data:"},{"metadata":{"trusted":false},"cell_type":"code","source":"res = pd.merge(td, groupedSnaps, on = 'PositionGroup', how='left').fillna(0)\nres[['PositionGroup', 'SyntheticDelta', 'TeamSnapsPerGame', 'Salary']]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And can use the following equation to find the injury lost salary per game by position group:\n<br>\n<br>\n$$\\:\\frac{\\Delta Missed\\:Games}{Play}*\\:\\:\\frac{Plays}{Game}\\:*\\:\\frac{$}{Season}\\:=\\:\\frac{$}{Season}\\:$$\n<br>\n<br>\nAnd then convert this to home games:\n<br><br>\n$$\\:\\:\\frac{{$}_{\\:Group}}{Season}\\:*\\:\\frac{8\\:Home\\:Games}{16\\:Games}\\:=\\:\\frac{$_{\\:Home\\:Group}}{Season}\\:$$"},{"metadata":{"trusted":false},"cell_type":"code","source":"# Calculations\nHOME_GAMES = 8 / 16\nres['InjuryHomeCostNatural'] = res['GM_Natural'] * res['TeamSnapsPerGame'] * res['Salary'] * HOME_GAMES\nres['InjuryHomeCostSynthetic'] = res['GM_Synthetic'] * res['TeamSnapsPerGame'] * res['Salary'] * HOME_GAMES\nres['InjuryHomeCostDelta'] = res['SyntheticDelta'] * res['TeamSnapsPerGame'] * res['Salary'] * HOME_GAMES\n\nres['TeamSnapsPerGame_F'] = [int(i) for i in res['TeamSnapsPerGame']]\nres['Salary_F'] = ['${:,.2f}m'.format(i/1000000) for i in res['Salary']]\n\nres['InjuryHomeCostDelta_F'] = ['${:,.0f}k'.format(i/1000) for i in res['InjuryHomeCostDelta']]\n\n\nres[['PositionGroup', 'TeamSnapsPerGame_F', 'Salary_F', \\\n    'SyntheticDelta', 'InjuryHomeCostDelta_F' ]]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Results\n\nWe can plot our results to look at the cost breakdown by position group"},{"metadata":{"trusted":false},"cell_type":"code","source":"# Example from here https://pythonspot.com/matplotlib-bar-chart/\nfig, ax = plt.subplots(figsize=(10,5))\nindex = np.arange(len(td))\nbar_width = 0.35\nopacity = 0.75\n\ntotalNaturalCost = '${:,.2f}m'.format(float(sum(res['InjuryHomeCostNatural']) / 1000000))\ntotalSyntheticCost = '${:,.2f}m'.format(float(sum(res['InjuryHomeCostSynthetic']) / 1000000))\nplt.bar(\n            x      = index, \n            height = list(res['InjuryHomeCostNatural']), \n            width  = bar_width,\n            alpha  = opacity,\n            color  = '#48BB78', # color from tailwindcss\n            label  = 'Natural (Total Cost of Injuries: ' + totalNaturalCost + ')'\n       )\n\nplt.bar(\n            x      = index + bar_width, \n            height = list(res['InjuryHomeCostSynthetic']), \n            width  = bar_width,\n            alpha  = opacity,\n            color  = '#2B6CB0', # color from tailwindcss\n            label  = 'Sythetic (Total Cost: ' + totalSyntheticCost + ')'\n        )\n\nax.set_yticklabels(['${:,}'.format(int(x)) for x in ax.get_yticks().tolist()])\nax.spines['top'].set_color('white')\nax.spines['right'].set_color('white')\nplt.xlabel('Position Group')\nplt.ylabel('Cost to Team')\nplt.title('Home Games on Natural v. Synthetic Turf - Injury Cost', fontsize=14, fontweight='bold', pad=20)\nplt.xticks(index + bar_width, list(res['PositionGroup']))\nplt.legend()\n\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.5"}},"nbformat":4,"nbformat_minor":1}