{"cells":[{"metadata":{"_uuid":"f9720139a2a6887a5060df806ffbd015090f4061"},"cell_type":"markdown","source":"## NFL Punt Analytics Proposal  ##\n\nThirty seven known punt-related concussions occurred during the 2016 and 2017 NFL seasons. By reviewing the evidence from all 6,681 punt plays that are recorded from these two seasons, I seek to show that the following three rule changes will improve player safety while being implementable and preserving the excitement and unpredictability of the great game of football.\n\n* [Proposal #1](#proposal_1): Award a five yard bonus to the receiving team on a fair catch. \n* [Proposal #2](#proposal_2): Require single coverage of gunners by the receiving team.\n* [Proposal #3](#proposal_3): Install helmet sensors to monitor deceleration.\n\nIn support of proposal #1, **I attempt to quantify the risk-reward tradeoff of reducing the number of punt returns and increasing the number of fair catches**. I find that 32% of punt returns conditional on a punt received yield fewer than five yards on the return. In the current framework, punt returners have an incentive to attempt a return whenever they think their expected value of returning is greater than 0 yards. In the new framework, punt returns will have an incentive to make a return only when they believe the expected value of returning is greater than 5 yards. The concussion rate is 10 injuries per 1000 plays for a fielded return versus 1.8 injuries per 1000 plays for a fair catch, so we expect a 80% decrease in concussions for each incremental fair catch that is called. On the other hand, punt returners will continue to be able to try to go for a return and make a play if they see an opportunity to go for it.  \n\nIn support of proposal #2, **I analyze the injury rate conditioning on both the choice of coverage and on the yards between the line of scrimmage and the end zone, as these factors are not independent of one another**. Teams are more likely to choose double coverage when there is more open field, i.e. greater yards to go between the kicking team and the receiving team's end zone. On the other hand, plays where the field is longer also have higher injury rates, perhaps because the possibility of a touchback or coffin corner is reduced and the coverage team has to run farther to reach the spot of punt reception, resulting in a more open field and higher speed play. I show that even when one conditions on yards-to-go, double coverage generates somewhat higher injury rates than single coverage in these long field situations. On the other hand, double coverage does not appear to help the receiving team make longer punt returns, i.e. the punt returner is able to generate just as much excitement whether or not the gunners were single covered or not. Enforcing single coverage would then appear to reduce injuries without limiting excitement, and so seems like a costless design proposal, although the statistical significance of these findings (recall there are only 37 concussion events in this entire sample) is limited by the small sample size.\n\nIn support of proposal #3, **I show that deceleration appears more important than velocity in determining injury, but that the NGS data is inherently limited in its ability to measure deceleration.**. The code below attempts to measure each player's velocity and deceleration at the time of a tackle event. I find that players who are not injured on a play have velocity as great as or greater than players involved in an injury generating tackle. On the other hand, the injured player experiences much greater deceleration than uninjured players. This suggests that deceleration is a more important factor in injury than velocity. The NGS data provides (x, y) coordinates at 100 millisecond resolution, so a player who decelerates from 9.8 m/s to 0.0 m/s would be recorded as having experienced at most 10 g's in deceleration, since 10g = (9.8m/s - 0.0m/s) / 0.1s. In fact the player could have experienced much greater deceleration, e.g. he could have decelerated to 0 in 20 milliseconds, but the resolution of the NGS system is not granular enough to show this. For the purpose of measuring head injury, the deceleration experienced at the head would seem to be a critical indicator of whether injury is likely to have occurred. This data would be collected not to take players out of the game or to limit their playing time, but rather to better study the conditions and plays that result in greater head deceleration and likely injury.\n\nThe notebook below provides the analysis and figures that support each of these three proposals. There is an extensive amount of formatting code and data preparation code, which I have attempted to hide, so that the focus can be on the interesting parts - the output and figures.\n\n[Additional observations](#additional_observations) are found at the end.\n\n[Final thoughts](#final_thoughts) concludes.\n"},{"metadata":{"trusted":true,"_uuid":"b5d09b8551015a9514fbe256c0f89f66a42ed564","_kg_hide-input":true},"cell_type":"code","source":"%matplotlib inline\nimport os\nimport sys\nimport re\nimport pandas as pd\nimport numpy as np\nimport glob\nimport os\nimport logging\nimport sys\nimport re\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings  \nwarnings.filterwarnings('ignore')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"26a7ab62e81ebd1a20ea43a5697eeb63102e3553","_kg_hide-input":true},"cell_type":"code","source":"ngs_files = ['../input/NGS-2016-pre.csv',\n             '../input/NGS-2016-reg-wk1-6.csv',\n             '../input/NGS-2016-reg-wk7-12.csv',\n             '../input/NGS-2016-reg-wk13-17.csv',\n             '../input/NGS-2016-post.csv',\n             '../input/NGS-2017-pre.csv',\n             '../input/NGS-2017-reg-wk1-6.csv',\n             '../input/NGS-2017-reg-wk7-12.csv',\n             '../input/NGS-2017-reg-wk13-17.csv',\n             '../input/NGS-2017-post.csv']\nPUNT_TEAM = set(['GL', 'PLW', 'PLT', 'PLG', 'PLS', 'PRG', 'PRT', 'PRW', 'PC',\n                 'PPR', 'P', 'GR'])\nRECV_TEAM = set(['VR', 'PDR', 'PDL', 'PLR', 'PLM', 'PLL', 'VL', 'PFB', 'PR'])\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b5d560a7de211cf78607b15d2ea1a6d722204b4e"},"cell_type":"markdown","source":"### **Proposal 1: **Award five yard bonus on fair catches ###\n\n<a id=\"proposal_1\"></a>\n\n**Fair catches are 80% safer than punt received events.**\nThe graph below shows that the concussion rate is 10.0 per 1000 returns on punt received events vs. 1.8 on fair catch events. While fair catches are not as exciting as punt returns, they are certainly much safer. Our goal should not be to ban all punt returns wholesale, but rather to reduce the number of dangerous but less exciting (i.e. short) punt returns while preserving the potential for long punt returns (ESPN highlight material). There is always risk in each play (even fair catches), and the analysis below helps us to find the best trade off between risk and reward.\n"},{"metadata":{"trusted":true,"_uuid":"01ff2a828b29aa66fd308521cf9ba1f4c70fb501","_kg_hide-input":true},"cell_type":"code","source":"plays_df = pd.read_csv('../input/play_information.csv')\n\ndef get_return_yards(s):\n    m = re.search('for ([0-9]+) yards', s)\n    if m:\n        return int(m.group(1))\n    elif re.search('for no gain', s):\n        return 0\n    else:\n        return np.nan\n\nplays_df['Return'] = plays_df['PlayDescription'].map(\n        lambda x: get_return_yards(x))\n\nvideo_review = pd.read_csv('../input/video_review.csv')\nvideo_review = video_review.rename(columns={'GSISID': 'InjuredGSISID'})\n\nplays_df= plays_df.merge(video_review, how='left',\n                         on=['Season_Year', 'GameKey', 'PlayID'])\n\nplays_df['InjuryOnPlay'] = 0\nplays_df.loc[plays_df['InjuredGSISID'].notnull(), 'InjuryOnPlay'] = 1\n\nplays_df = plays_df[['Season_Year', 'GameKey', 'PlayID', 'Return', 'InjuryOnPlay']]\n\nngs_df = []\nfor filename in ngs_files:\n    df = pd.read_csv(filename, parse_dates=['Time'])\n    df = df.loc[df['Event'].isin(['fair_catch', 'punt_received'])]\n    df = pd.concat([df, pd.get_dummies(df['Event'])], axis=1)\n    df = df.groupby(['Season_Year', 'GameKey', 'PlayID'])[['fair_catch', 'punt_received']].max()\n    ngs_df.append(df.reset_index())\nngs_df = pd.concat(ngs_df)\n\nplays_df = plays_df.merge(ngs_df, on=['Season_Year', 'GameKey', 'PlayID'])","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_kg_hide-input":true,"_uuid":"753d4d13ce9a0d56b0d940758a15fa1c11deb7bb"},"cell_type":"code","source":"injury_per_1000_fair_catch = 1000 * plays_df.loc[plays_df['fair_catch']==1,\n                                          'InjuryOnPlay'].mean()\ninjury_per_1000_punt_received = 1000 * plays_df.loc[plays_df['punt_received']==1,\n                                           'InjuryOnPlay'].mean()\nfig = plt.figure()\nax = plt.subplot2grid((1, 1), (0, 0))\nplt.bar([0, 1], [injury_per_1000_fair_catch, injury_per_1000_punt_received])\nax.set_xticks([0, 1])\nax.set_xticklabels(['Fair Catch', 'Punt Received'])\nplt.text(0, injury_per_1000_fair_catch+0.2, '{:.1f}'.format(injury_per_1000_fair_catch))\nplt.text(1, injury_per_1000_punt_received+0.2, '{:.1f}'.format(injury_per_1000_punt_received))\nplt.title(\"Concussion Rate\")\nplt.ylabel(\"Injuries per 1000 Events\")\nsns.despine(top=True, right=True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"10848f900bea78eea93c86e83ec27c7e829577a9"},"cell_type":"markdown","source":"#### ** 32% of punt returns result in fewer than 5 yards ** ####\nThe graph below shows that 32% of punt returns yield fewer than 5 yards. These seem like prime candidates for events that we could replace with an awarded bonus of 5 yards, with little effect on field position but with substantial decrease in injury. \n\n**Nudge Theory**\n\nAs an additional alternative, we could consider a proposal that makes fair catch the default, and require a returner to make a signal or gesture in order to indicate that he will not choose a fair catch. Evidence in support of setting an optimal default when giving people choices include research by Richard Thaler, Nobel prize-winning economist at the University of Chicago, who has written extensively on the notion of 'nudges.' If we nudge receivers into choosing a fair catch, then they will have to make the conscious decision to elect to return after assessing that the situation presents an opportunity to make a long return."},{"metadata":{"scrolled":true,"trusted":true,"_kg_hide-input":true,"_uuid":"cb7194f6e48f890ab4bf2b37fd1196a19c88e46e"},"cell_type":"code","source":"x_groups = ['0-3 yds', '3-5 yds', '5-7 yds', '7-9 yds',\n            '9-12 yds', '12-15 yds', '15-20 yds', '20+ yds']\nrec = plays_df.loc[(plays_df['punt_received']==1) \n                   &(plays_df['Return'].notnull())]\n\ny_groups = [sum(rec['Return']<=3) / len(rec),\n            sum((rec['Return']>3) & (rec['Return']<=5)) / len(rec),\n            sum((rec['Return']>5) & (rec['Return']<=7)) / len(rec),\n            sum((rec['Return']>7) & (rec['Return']<=9)) / len(rec),\n            sum((rec['Return']>9) & (rec['Return']<=12)) / len(rec),\n            sum((rec['Return']>12) & (rec['Return']<=15)) / len(rec),\n            sum((rec['Return']>15) & (rec['Return']<=20))/ len(rec),\n            sum(rec['Return']>20) / len(rec)]\n\ny_bottoms = [0,\n             sum(rec['Return']<=3) / len(rec),\n             sum(rec['Return']<=5) / len(rec),\n             sum(rec['Return']<=7) / len(rec),\n             sum(rec['Return']<=9) / len(rec),\n             sum(rec['Return']<=12) / len(rec),\n             sum(rec['Return']<=15) / len(rec),\n             sum(rec['Return']<=20) / len(rec)]\n\nfig = plt.figure(figsize=(8.5,4.5))\nax = plt.subplot2grid((1, 1), (0, 0))\nplt.bar(range(len(x_groups)), y_groups, bottom=y_bottoms)\nax.set_xticks(range(len(x_groups)))\nax.set_xticklabels(x_groups)\nax.set_yticks([0, 0.2, 0.4, 0.6, 0.8, 1.0])\nax.set_yticklabels(['0%', '20%', '40%', '60%', '80%', '100%'])\nfor i in range(len(x_groups)):\n    plt.text(i-0.2, y_bottoms[i]+y_groups[i]+0.02, '{:.0f}%'.format(100*y_groups[i]))\nsns.despine(top=True, right=True)\nplt.title(\"Distribution of Punt Returns by Length\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a950ef8e0b7e473ab1472d95e17689c7e4336907"},"cell_type":"markdown","source":"#### ** Example 1 of a play where the punt returner should probably have called a fair catch ** ####\nA NFL punter can kick a punt with hang time of 4.4 seconds on average, which gives NFL level gunners time to advance 40 yards. In practice there is usually a 1-2 second gap between the time of ball snap and the time of the punt, allowing the gunner to advance 50 or 60 yards down the field before the ball is caught That's a lot of time, and sometimes I cringe while watching a punt returner (aka sitting duck) about to get nailed by a gunner - often the punt returner is too busy focusing on the ball to even notice that a player is about to hit him on his blind side. This appears to be one of those plays, although as often happens it's the gunner who suffers the injury and not the returner."},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"0135ffc744eb60211b1e960b0806f2f837b085d0"},"cell_type":"code","source":"play_player_role_data = pd.read_csv('../input/play_player_role_data.csv')\ngfp = video_review.merge(play_player_role_data, on=['GameKey', 'PlayID', 'Season_Year'])\n","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"866fd5586dd0171c2e79a8e94a7a7fb7f125698b"},"cell_type":"code","source":"df_29 = pd.read_csv('../input/NGS-2016-pre.csv', parse_dates=['Time'])\ndf_29 = df_29.loc[(df_29['GameKey']==29) & (df_29['PlayID']==538)]\ndf_29 = df_29.merge(gfp, on=['GameKey', 'PlayID', 'Season_Year', 'GSISID'])\ndf_29 = df_29.sort_values(['GameKey', 'PlayID', 'Season_Year', 'GSISID', 'Time'])\n\nfig = plt.figure(figsize = (10, 4.5))\nax = plt.subplot2grid((1, 1), (0, 0))\nline_set = df_29.loc[df_29['Event'].isin(['ball_snap'])]\nline_set_time = line_set['Time'].min()\nline_of_scrimmage = line_set.loc[line_set['Role'].isin(['PLS', 'PLG', 'PRG']), 'x'].median()\n\nrecv_df = df_29.loc[df_29['Event']=='punt_received']\nrecv_time = recv_df['Time'].min()\nevent_df = df_29.loc[df_29['Time'] <= recv_time]\nevent_df = event_df.loc[event_df['Time'] >= recv_time + pd.Timedelta('-2s')]\n\ninjured = df_29['InjuredGSISID'].values[0]\npartner = float(df_29['Primary_Partner_GSISID'].values[0])\n            \nplayers = event_df['GSISID'].unique()\nfor player in players:\n    player_df = event_df.loc[event_df['GSISID'] == player]\n    role = str(player_df['Role'].values[0])\n    if re.sub('[io0-9]', '', str(role)) in PUNT_TEAM:\n        color = '#fdc086'\n        marker = 'x'\n        linewidth = 6\n    else:\n        color = '#beaed4'\n        marker = 'o'\n        linewidth = 6\n    if player == injured:\n        marker = '*'\n        linewidth = 10\n        linestyle = '-'\n        color = '#f0027f'\n    elif player == partner:\n        marker = '*'\n        linewidth = 10\n        linestyle = '-'\n        color = '#ffff99'\n    else:\n        linestyle = '--'\n    alphas = np.ones(len(player_df))\n    alphas = alphas.cumsum() / alphas.sum()\n    px = player_df['x'].values\n    py = player_df['y'].values\n    for k in range(len(px)):\n        plt.plot(px[k:], py[k:], color=color,\n                 linewidth=linewidth*(k+1+4)/(4+len(px)),\n                 linestyle=linestyle,\n                 alpha=(k+1)/len(px))\n        player_df = player_df.reset_index(drop=True)\n        x = player_df['x'].iloc[-1]\n        y = player_df['y'].iloc[-1]\n\n        marker = (3, 0, 90 + player_df['o'].iloc[-1])\n        plt.scatter(player_df['x'].iloc[-1],\n                    player_df['y'].iloc[-1],\n                    marker=marker,\n                    s=linewidth*60,\n                    color=color)\n        if (role == 'PR'):\n            circ = plt.Circle((player_df['x'].iloc[-1],\n                                player_df['y'].iloc[-1]),\n                                5, color=color,\n                                fill=False)\nax.set_xlim(0, 120)\nax.set_ylim(0, 53.3)\nplt.axvline(x=0, color='w', linewidth=2)\nplt.axvline(x=10, color='w', linewidth=2)\nplt.axvline(x=15, color='w', linewidth=2)\nplt.axvline(x=20, color='w', linewidth=2)\nplt.axvline(x=25, color='w', linewidth=2)\nplt.axvline(x=30, color='w', linewidth=2)\nplt.axvline(x=35, color='w', linewidth=2)\nplt.axvline(x=40, color='w', linewidth=2)\nplt.axvline(x=45, color='w', linewidth=2)\nplt.axvline(x=50, color='w', linewidth=2)\nplt.axvline(x=55, color='w', linewidth=2)\nplt.axvline(x=60, color='w', linewidth=2)\nplt.axvline(x=65, color='w', linewidth=2)\nplt.axvline(x=70, color='w', linewidth=2)\nplt.axvline(x=75, color='w', linewidth=2)\nplt.axvline(x=80, color='w', linewidth=2)\nplt.axvline(x=85, color='w', linewidth=2)\nplt.axvline(x=90, color='w', linewidth=2)\nplt.axvline(x=95, color='w', linewidth=2)\nplt.axvline(x=100, color='w', linewidth=2)\nplt.axvline(x=105, color='w', linewidth=2)\nplt.axvline(x=110, color='w', linewidth=2)\nplt.axvline(x=120, color='w', linewidth=2)\nplt.axvline(x=line_of_scrimmage, color='y', linewidth=3)\nplt.axhline(y=0, color='w', linewidth=2)\nplt.axhline(y=53.3, color='w', linewidth=2)\nplt.text(x=18, y=2, s= '1', color='w')\nplt.text(x=21, y=2, s= '0', color='w')\nplt.text(x=28, y=2, s= '2', color='w')\nplt.text(x=31, y=2, s= '0', color='w')\nplt.text(x=38, y=2, s= '3', color='w')\nplt.text(x=41, y=2, s= '0', color='w')\nplt.text(x=48, y=2, s= '4', color='w')\nplt.text(x=51, y=2, s= '0', color='w')\nplt.text(x=58, y=2, s= '5', color='w')\nplt.text(x=61, y=2, s= '0', color='w')\nplt.text(x=68, y=2, s= '4', color='w')\nplt.text(x=71, y=2, s= '0', color='w')\nplt.text(x=78, y=2, s= '3', color='w')\nplt.text(x=81, y=2, s= '0', color='w')\nplt.text(x=88, y=2, s= '2', color='w')\nplt.text(x=91, y=2, s= '0', color='w')\nplt.text(x=98, y=2, s= '1', color='w')\nplt.text(x=101, y=2, s= '0', color='w')\n\nax.set_xticks([0, 120])\nax.set_yticks([0, 53.3])\nax.set_xticklabels(['', ''])\nax.set_yticklabels(['', ''])\nax.tick_params(axis=u'both', which=u'both', length=0)\nax.add_artist(circ)\nax.set_facecolor(\"#2ca25f\")\nplt.title(\"GR is injured while tackling PR (helmet to body).\\nGameKey 29, Play ID 538. NYJ Punting to WAS.\")\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"60c901cdab45330bfb9b2f45da29f6e05f69e74e"},"cell_type":"markdown","source":"#### Example 2 of a play where the punt returner should probably have called a fair catch ####\nHere  the punt returner isn't even directly involved in the injury, but rather two members of the coverage team collide with each other. When multiple players have a chance to converge at a high rate of speed to the same stationary target, there is increased risk. "},{"metadata":{"scrolled":false,"trusted":true,"_uuid":"0d0bbaebcfcff5cd81539fc23c7923f73455397a","_kg_hide-input":true},"cell_type":"code","source":"df_296 = pd.read_csv('../input/NGS-2016-reg-wk13-17.csv', parse_dates=['Time'])\ndf_296 = df_296.loc[(df_296['GameKey']==296) & (df_296['PlayID']==2667)]\ndf_296 = df_296.merge(gfp, on=['GameKey', 'PlayID', 'Season_Year', 'GSISID'])\ndf_296 = df_296.sort_values(['GameKey', 'PlayID', 'Season_Year', 'GSISID', 'Time'])\nfig = plt.figure(figsize = (10, 4.5))\nax = plt.subplot2grid((1, 1), (0, 0))\nline_set = df_296.loc[df_296['Event'].isin(['ball_snap'])]\nline_set_time = line_set['Time'].min()\nline_of_scrimmage = line_set.loc[line_set['Role'].isin(['PLS', 'PLG', 'PRG']), 'x'].median()\n\nrecv_df = df_296.loc[df_296['Event']=='punt_received']\nrecv_time = recv_df['Time'].min()\nevent_df = df_296.loc[df_296['Time'] <= recv_time]\nevent_df = event_df.loc[event_df['Time'] >= recv_time + pd.Timedelta('-2s')]\n\ninjured = df_296['InjuredGSISID'].values[0]\npartner = float(df_296['Primary_Partner_GSISID'].values[0])\n            \nplayers = event_df['GSISID'].unique()\nfor player in players:\n    player_df = event_df.loc[event_df['GSISID'] == player]\n    role = str(player_df['Role'].values[0])\n    if re.sub('[io0-9]', '', str(role)) in PUNT_TEAM:\n        color = '#fdc086'\n        marker = 'x'\n        linewidth = 6\n    else:\n        color = '#beaed4'\n        marker = 'o'\n        linewidth = 6\n    if player == injured:\n        marker = '*'\n        linewidth = 10\n        linestyle = '-'\n        color = '#f0027f'\n    elif player == partner:\n        marker = '*'\n        linewidth = 10\n        linestyle = '-'\n        color = '#ffff99'\n    else:\n        linestyle = '--'\n    alphas = np.ones(len(player_df))\n    alphas = alphas.cumsum() / alphas.sum()\n    px = player_df['x'].values\n    py = player_df['y'].values\n    for k in range(len(px)):\n        plt.plot(px[k:], py[k:], color=color,\n                 linewidth=linewidth*(k+1+4)/(4+len(px)),\n                 linestyle=linestyle,\n                 alpha=(k+1)/len(px))\n        player_df = player_df.reset_index(drop=True)\n        x = player_df['x'].iloc[-1]\n        y = player_df['y'].iloc[-1]\n\n        marker = (3, 0, 90 + player_df['o'].iloc[-1])\n        plt.scatter(player_df['x'].iloc[-1],\n                    player_df['y'].iloc[-1],\n                    marker=marker,\n                    s=linewidth*60,\n                    color=color)\n        if (role == 'PR'):\n            circ = plt.Circle((player_df['x'].iloc[-1],\n                                player_df['y'].iloc[-1]),\n                                5, color=color,\n                                fill=False)\nax.set_xlim(0, 120)\nax.set_ylim(0, 53.3)\nplt.axvline(x=0, color='w', linewidth=2)\nplt.axvline(x=10, color='w', linewidth=2)\nplt.axvline(x=15, color='w', linewidth=2)\nplt.axvline(x=20, color='w', linewidth=2)\nplt.axvline(x=25, color='w', linewidth=2)\nplt.axvline(x=30, color='w', linewidth=2)\nplt.axvline(x=35, color='w', linewidth=2)\nplt.axvline(x=40, color='w', linewidth=2)\nplt.axvline(x=45, color='w', linewidth=2)\nplt.axvline(x=50, color='w', linewidth=2)\nplt.axvline(x=55, color='w', linewidth=2)\nplt.axvline(x=60, color='w', linewidth=2)\nplt.axvline(x=65, color='w', linewidth=2)\nplt.axvline(x=70, color='w', linewidth=2)\nplt.axvline(x=75, color='w', linewidth=2)\nplt.axvline(x=80, color='w', linewidth=2)\nplt.axvline(x=85, color='w', linewidth=2)\nplt.axvline(x=90, color='w', linewidth=2)\nplt.axvline(x=95, color='w', linewidth=2)\nplt.axvline(x=100, color='w', linewidth=2)\nplt.axvline(x=105, color='w', linewidth=2)\nplt.axvline(x=110, color='w', linewidth=2)\nplt.axvline(x=120, color='w', linewidth=2)\nplt.axvline(x=line_of_scrimmage, color='y', linewidth=3)\nplt.axhline(y=0, color='w', linewidth=2)\nplt.axhline(y=53.3, color='w', linewidth=2)\nplt.text(x=18, y=2, s= '1', color='w')\nplt.text(x=21, y=2, s= '0', color='w')\nplt.text(x=28, y=2, s= '2', color='w')\nplt.text(x=31, y=2, s= '0', color='w')\nplt.text(x=38, y=2, s= '3', color='w')\nplt.text(x=41, y=2, s= '0', color='w')\nplt.text(x=48, y=2, s= '4', color='w')\nplt.text(x=51, y=2, s= '0', color='w')\nplt.text(x=58, y=2, s= '5', color='w')\nplt.text(x=61, y=2, s= '0', color='w')\nplt.text(x=68, y=2, s= '4', color='w')\nplt.text(x=71, y=2, s= '0', color='w')\nplt.text(x=78, y=2, s= '3', color='w')\nplt.text(x=81, y=2, s= '0', color='w')\nplt.text(x=88, y=2, s= '2', color='w')\nplt.text(x=91, y=2, s= '0', color='w')\nplt.text(x=98, y=2, s= '1', color='w')\nplt.text(x=101, y=2, s= '0', color='w')\n\nax.set_xticks([0, 120])\nax.set_yticks([0, 53.3])\nax.set_xticklabels(['', ''])\nax.set_yticklabels(['', ''])\nax.tick_params(axis=u'both', which=u'both', length=0)\nax.add_artist(circ)\nax.set_facecolor(\"#2ca25f\")\nplt.title(\"GL and GR collide, injuring GL (helmet to helmet friendly fire).\\nGameKey 296, Play ID 2667. TEN punting to JAX.\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7914cc0c9a70aaa549c8f8a3ba8906c0cf37a67"},"cell_type":"markdown","source":"#### **Going for it on fourth down**\nWhile somewhat out of the scope of this study, there is considerable outside evidence that teams are overly conservative on fourth down, choosing to punt when they would be better off attempting to move the chains. If the fair catch bonus of 5 yards were to encourage more teams to go for it rather than punt, that would seem to a positive side effect."},{"metadata":{"_uuid":"5b0d8bcedfbe60aa2cccd6da45ce6907b07c4056"},"cell_type":"markdown","source":"### **Proposal 2** : require single coverage on gunners.\n <a id=\"proposal_2\"></a>\nDoes single coverage result in fewer injuries? To start to analyze this, we classify plays as having single, hybrid, or double coverage, and see whether or not an injury occurred. Because teams are more likely to opt for double coverage in plays where the field is long (i.e. the punt is less likely to be coffin cornered or result in touchback), we need to bucket our data in two dimensions: (a) choice of coverage and (b) distance from line of scrimmage to the end zone.\n "},{"metadata":{"trusted":true,"_uuid":"9fe5bbf66eaeab91ea623912d3065b3b37b1b857","_kg_hide-input":true},"cell_type":"code","source":"ppr = pd.read_csv('../input/play_player_role_data.csv')\nppr['Role'] = ppr['Role'].map(lambda x: re.sub('[oi0-9]', '', x))\nroles = ppr['Role'].unique()\n\nppr = pd.concat([ppr, pd.get_dummies(ppr['Role'])], axis=1)\nppr = ppr.groupby(['Season_Year', 'GameKey', 'PlayID'])[roles].sum()\nppr = ppr.reset_index()\n\nvi = pd.read_csv('../input/video_review.csv')\nvi = vi[['Season_Year', 'GameKey', 'PlayID', 'GSISID']]\n\nppr = ppr.merge(vi, on=['Season_Year', 'GameKey', 'PlayID'], how='left')\nppr['Injury'] = 0\nppr['const'] = 1\n\nppr.loc[ppr['GSISID'].notnull(), 'Injury'] = 1\n\nplay_information = pd.read_csv('../input/play_information.csv')\ndef extract_recv_yards(s):\n    m = re.search('for ([0-9]+) yards', s)\n    if m:\n        return int(m.group(1))\n    elif re.search('for no gain', s):\n        return 0\n    else:\n        return np.nan\nplay_information['recv_length'] = play_information['PlayDescription'].map(\n        lambda x: extract_recv_yards(x))\nplay_information = play_information[['Season_Year', 'GameKey', 'PlayID', 'YardLine', 'Poss_Team', 'recv_length']]\nplay_information['yards_to_go'] = play_information['YardLine'].map(lambda x: int(x[-2:]))\nplay_information['back_half'] = play_information.apply(lambda x:\n    x['YardLine'].startswith(x['Poss_Team']), axis = 1)\nplay_information.loc[play_information['back_half']==1, 'yards_to_go'] = 100 - (\n    play_information.loc[play_information['back_half']==1, 'yards_to_go'])\n\nplay_information = play_information[['Season_Year', 'GameKey', 'PlayID', 'yards_to_go', 'recv_length']]\nppr = ppr.merge(play_information, on=['Season_Year', 'GameKey', 'PlayID'], how='inner')\n\ncol = 'VR_VL'\nppr[col] = ppr.loc[:, ['VR', 'VL']].sum(axis=1)\nptiles = np.percentile(ppr[col], [0, 25, 50, 75, 100])\n\nppr['yards'] = ''\nppr['coverage'] = ''\nppr.loc[ppr[col]==2, 'coverage'] = 'single'\nppr.loc[ppr[col]==3, 'coverage'] = 'hybrid'\nppr.loc[ppr[col]==4, 'coverage'] = 'double'\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e39bc966b3f7053d05ef94f1b2f30ea0eb28a1ec"},"cell_type":"markdown","source":"This graph below shows that receiving teams choose single coverage 84% of the time when the field of play is effectively short, e.g. there are fewer than 50 yards from the line of scrimmage to the end zone. On the other hand, receiving teams choose double coverage, single coverage, and hybrid coverage with more or less even probability when there are more than 70 yards between the line of scrimmage and the goal line. "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"7a45c4db165d23ac47103a620d5264e072833d71"},"cell_type":"code","source":"### For graphing, keep track of counts of plays by single, hybrid, or double coverage\n### Sort by yards-to-go.\nx_single = [0.00, 1.00, 2.00]\nx_hybrid = [0.25, 1.25, 2.25]\nx_double = [0.50, 1.50, 2.50]\ny_single = []\ny_hybrid = []\ny_double = []\ninjury_rate = {'Yards to Go': [], '30-50': [], '50-70': [], '70-100': []}\nfor i in range(2, 5):\n    if i == 2:\n        mode = 'Single'\n    elif i == 3:\n        mode = 'Hybrid'\n    elif i == 4:\n        mode = 'Double'\n    injury_rate['Yards to Go'].append(mode)\n    for r in ([30, 50], [50, 70], [70, 100]):\n        # Edit : condition on punts that are returned, \n        # ie ignore fair catch\n        ii = ((ppr['yards_to_go']>=r[0])\n              &(ppr['yards_to_go']<r[1])\n              &(ppr['recv_length'].notnull())\n              &(ppr['recv_length']!=0))\n        if (r[0]==30) & (r[1]==50):\n            ppr.loc[ii, 'yards'] = '30 to 50'\n        elif (r[0]==50) & (r[1]==70):\n            ppr.loc[ii, 'yards'] = '50 to 70'\n        elif (r[0]==70) & (r[1]==100):\n            ppr.loc[ii, 'yards'] = '70 to 100'\n\n        pprt = ppr.loc[ii]\n        if len(pprt) == 0:\n            pass\n        iii = (pprt[col] == i)\n        if sum(iii) == 0:\n            continue\n        # Keep track of coverage choice by yards to go\n        if mode == 'Single':\n            y_single.append(sum(iii) / sum(ii)) # Ratio of times single coverage is elected\n        elif mode == 'Hybrid':\n            y_hybrid.append(sum(iii) / sum(ii)) # Ratio of times hybrid coverage is elected\n        elif mode == 'Double':\n            y_double.append(sum(iii) / sum(ii)) # Ratio of times double coverage is elected\n        if r[0] == 30:\n            injury_rate['30-50'].append(1000 * pprt.loc[iii, 'Injury'].mean())\n        elif r[0] == 50:\n            injury_rate['50-70'].append(1000 * pprt.loc[iii, 'Injury'].mean())\n        elif r[0] == 70:\n            injury_rate['70-100'].append(1000 * pprt.loc[iii, 'Injury'].mean())\n            \ninjury_rate['Yards to Go'].append(\"All\")\ninjury_rate['30-50'].append(1000 * ppr.loc[\n    (ppr['yards_to_go']>=30) & (ppr['yards_to_go']<50) &(ppr['recv_length'].notnull())\n              &(ppr['recv_length']!=0) , 'Injury'].mean())\ninjury_rate['50-70'].append(1000 * ppr.loc[\n    (ppr['yards_to_go']>=50) & (ppr['yards_to_go']<70) &(ppr['recv_length'].notnull())\n              &(ppr['recv_length']!=0), 'Injury'].mean())\ninjury_rate['70-100'].append(1000 * ppr.loc[\n    (ppr['yards_to_go']>=70) & (ppr['yards_to_go']<100) &(ppr['recv_length'].notnull())\n              &(ppr['recv_length']!=0), 'Injury'].mean())\ninjury_rate = pd.DataFrame(injury_rate)\ninjury_rate = injury_rate[['Yards to Go', '30-50', '50-70', '70-100']]\n\nfig = plt.figure(figsize = (6.5, 4.5))\nax = plt.subplot2grid((1, 1), (0, 0))\nplt.bar(x_single, y_single, color='#7fc97f', width=.25, label='Single Coverage')\nplt.bar(x_hybrid, y_hybrid, color='#beaed4', width=.25, label='Hybrid Coverage')\nplt.bar(x_double, y_double, color='#fdc086', width=.25, label='Double Coverage')\nfor i in range(3):\n    plt.text(x_single[i]-0.06, y_single[i]+0.02, '{:.0f}%'.format(y_single[i] * 100))\n    plt.text(x_hybrid[i]-0.06, y_hybrid[i]+0.02, '{:.0f}%'.format(y_hybrid[i] * 100))\n    plt.text(x_double[i]-0.06, y_double[i]+0.02, '{:.0f}%'.format(y_double[i] * 100))\n    \nplt.legend()\nax.set_yticks([0, 0.25, 0.5, 0.75, 1.00])\nax.set_yticklabels(['0%', '25%', '50%', '75%', '100%'])\nax.set_xticks([0.25, 1.25, 2.25])\nax.set_xticklabels(['30-50 Yards', '50-70 Yards', '70-100 Yards'])\nax.set_xlabel('Yards from Line of Scrimmage to End Zone')\nax.set_ylabel('Coverage Choice (%)')\nsns.despine(top=True, right=True)\n\nplt.title(\"Coverage Choice vs. Yards between Line of Scrimmage and End Zone\\n\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"970a6743723d90e838ef32326509f1889fd22392"},"cell_type":"markdown","source":"The injury rate is higher on double coverage and hybrid coverage vs. single coverage, at least in plays with longer fields. E.g. looking at all coverage types (the last row of the table below), there are 3 injuries per 1000 plays with 30-50 yards-to-go versus 9 injuries per 1000 plays when there are 70-100 yards-to-go. Because choice of coverage is correlated with yards-to-go, we need to control for both effects when trying to evaluate whether coverage is related to injury rate.\n\nBy subsetting plays on both dimensions, we can see that even when controlling for yards to go, double coverage appears to result in substantially higher injury rates than single and hybrid coverage. That said, the statistical significance of this finding is limited by the small sample size, as there are too few observed injuries in the provided data to draw sharp inferences. Given that we are not strongly confident that double coverage increases the injury rate, we need at least to be sure that prohibiting double coverage does not reduce the quality of game play.\n\n** Table of Injury Rate vs. Yards-to-Go and Coverage Choice **"},{"metadata":{"trusted":true,"_uuid":"9f5e84b812c233b46c07f65f581fc7df83bb96bf","_kg_hide-input":true},"cell_type":"code","source":"injury_rate.style.set_precision(2).hide_index()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a72902e77e37e71c2bf3ef298e336ba2407b4905"},"cell_type":"markdown","source":"One measure of quality of play is the length of punt return, with longer punt returns being (in my mind) associated with more exciting games. Let us look at the distribution of punt return length for punts received with single, hybrid, and double coverage. Again, we group by yards-to-go, since coverage decisions appear partly conditioned on the length of the playable field. The \"violin plots\" below show the density and distribution of punt returns, condtioned on coverage choice and yards to go. At first glance, return length appears to be indistinguishable between single, hybrid, and double coverage, suggesting that the return team would do just as well electing single coverage.\n\n** Distribution of Punt Return Length Conditioned on Yards-to-Go and Coveage Choice **"},{"metadata":{"trusted":true,"_uuid":"119919ce7a496faf6ddd5fd65f59bbf9adb0f6d4","_kg_hide-input":true},"cell_type":"code","source":"fig = plt.figure(figsize = (8.5, 5.5))\nax = plt.subplot2grid((1, 1), (0, 0))\npal = {'single': '#7fc97f', 'hybrid': '#beaed4', 'double': '#fdc086'}\nax = sns.violinplot(x=\"yards\", y=\"recv_length\", hue=\"coverage\",\n                    data=ppr, palette=pal,\n                    order=['30 to 50', '50 to 70', '70 to 100'],\n                    hue_order=['single', 'hybrid', 'double'],\n                    cut=0)\nax.legend().remove()\nsns.despine(top=True, right=True)\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eedeaf0c6a52c4ff50939462c312dffb2d8e9aba"},"cell_type":"markdown","source":"### **Proposal 3**: Require helmet monitors to better measure deceleration at time of tackle.\n<a id=\"proposal_3\"></a>\nAs the plots below show, deceleration is greater for tackles generating injuries than for tackles without injuries. On the other hand, velocities are somewhat higher on non-injury players and plays. Intuitively, there is nothing dangerous about running fast, unless you have to stop suddenly. \n\nBecause NGS data is provided only at 100 millisecond resolution, the maximum observable deceleration is approximately 100 m/s/s (assuming a person running flat out at 10 m/s at time t is moving at 0 m/s  in the next NGS data point, 100 milli seconds later) , or less than 10g's. On the other hand, it is possible for players to have experienced much greater g-forces, and for those forces to be stronger at the helmet than in the rest of the body. Without helmet sensor data, there is a finite upper bound on the measurable deceleration forces, one that is too coarse to be really useful at distinguishing injury from safe and routine play.\n"},{"metadata":{"trusted":true,"_uuid":"fff1a109abcfc956943fcfb2fa2e45940e90a2d3","_kg_hide-input":true},"cell_type":"code","source":"# Collate data\ngame_data = pd.read_csv('../input/game_data.csv')\nplay_information = pd.read_csv('../input/play_information.csv')\nplayer_punt_data = pd.read_csv('../input/player_punt_data.csv')\nplayer_punt_data = player_punt_data.groupby('GSISID').head(1).reset_index()\nplay_player_role_data = pd.read_csv('../input/play_player_role_data.csv')\nvideo_review = pd.read_csv('../input/video_review.csv')\n\ncombined = game_data.merge(play_information.drop(['Game_Date'], axis=1),\n                        on=['GameKey', 'Season_Year', 'Season_Type', 'Week'])\ncombined = combined.merge(play_player_role_data,\n                        on=['GameKey', 'Season_Year', 'PlayID'])\ncombined = combined.merge(player_punt_data, on=['GSISID'])\n\ncombined = combined.merge(video_review, how='left',\n                          on=['Season_Year', 'GameKey', 'PlayID', 'GSISID'])\ncombined['injury'] = 0\ncombined.loc[combined['Player_Activity_Derived'].notnull(), 'injury'] = 1\n\nngs_files = ['../input/NGS-2016-pre.csv',\n             '../input/NGS-2016-reg-wk1-6.csv',\n             '../input/NGS-2016-reg-wk7-12.csv',\n             '../input/NGS-2016-reg-wk13-17.csv',\n             '../input/NGS-2016-post.csv',\n             '../input/NGS-2017-pre.csv',\n             '../input/NGS-2017-reg-wk1-6.csv',\n             '../input/NGS-2017-reg-wk7-12.csv',\n             '../input/NGS-2017-reg-wk13-17.csv',\n             '../input/NGS-2017-post.csv']\n\nmax_decel_df = []\nfor filename in ngs_files:\n    logging.info(\"Loading file \" + filename)\n    group_keys = ['Season_Year', 'GameKey', 'PlayID', 'GSISID']\n    df = pd.read_csv(filename, parse_dates=['Time'])\n    logging.info(\"Read file \" + filename)\n\n    df = df.sort_values(group_keys + ['Time'])\n    df['dx'] = df.groupby(group_keys)['x'].diff(1)\n    df['dy'] = df.groupby(group_keys)['y'].diff(1)\n    df['dis'] = (df['dx']**2 + df['dy']**2)**0.5\n    df['dt'] = df.groupby(group_keys)['Time'].diff(1).dt.total_seconds()\n    df['velocity'] = 0\n    ii = (df['dis'].notnull() & df['dt'].notnull() & (df['dt']>0))\n    df.loc[ii, 'velocity'] = df.loc[ii, 'dis'] / df.loc[ii, 'dt']\n    df['velocity'] *= 0.9144 # Convert yards to meters\n    df['deceleration'] = -1 * df.groupby(group_keys)['velocity'].diff(1)\n    df['velocity'] = df.groupby(group_keys)['velocity'].shift(1)\n\n    # Only look at the one second window around each tackle\n    df['Event'] = df.groupby(group_keys)['Event'].ffill(limit=5)\n    df['Event'] = df.groupby(group_keys)['Event'].bfill(limit=5)\n\n    t_df = df.loc[df['Event']=='tackle']\n\n    t_max_decel = t_df.loc[t_df.groupby(['Season_Year', 'GameKey', 'PlayID', 'GSISID'])['deceleration'].idxmax()]\n    t_max_decel = t_max_decel[['Season_Year', 'GameKey', 'PlayID', 'GSISID', 'deceleration']].rename(columns={'deceleration': 'deceleration_at_tackle'})\n    \n    t_max_velocity = t_df.loc[t_df.groupby(['Season_Year', 'GameKey', 'PlayID', 'GSISID'])['velocity'].idxmax()]\n    t_max_velocity = t_max_velocity[['Season_Year', 'GameKey', 'PlayID', 'GSISID', 'velocity']].rename(columns={'velocity': 'velocity_at_tackle'})\n\n    max_decel = t_max_velocity.merge(t_max_decel, on=['Season_Year', 'GameKey', 'PlayID', 'GSISID'],\n                                      how='outer')\n    max_decel_df.append(max_decel)\n\nmax_decel_df = pd.concat(max_decel_df)\ncombined = combined.merge(max_decel_df, on=['Season_Year', 'GameKey', 'PlayID', 'GSISID'], how='left')\n\n\ncombined['tackle_injury'] = combined['Player_Activity_Derived'].isin(['Tackled', 'Tackling'])\n### Original work by Halla Yang","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9ba75450cd71e8b20a29afd0809833311629eaa3"},"cell_type":"markdown","source":"#### Distribution of Player Velocities and Deceleration at time of Event Tackle, Injury Plays vs. non-Injury Plays "},{"metadata":{"trusted":true,"_uuid":"eef097d7cacf017ee3398315f3e4f7072ac6287e","_kg_hide-input":true},"cell_type":"code","source":"fig = plt.figure(figsize = (6.5, 5.0))\nax = plt.subplot2grid((1, 1), (0, 0))\ninj = combined.loc[(combined['injury']==1)&(combined['velocity_at_tackle'].notnull())\n                  &(combined['deceleration_at_tackle'].notnull())]\nax = sns.kdeplot(inj.velocity_at_tackle, inj.deceleration_at_tackle,\n                 cmap=\"Reds\")\nnotinj = combined.loc[(combined['injury']==0)&(combined['velocity_at_tackle'].notnull())\n                  &(combined['deceleration_at_tackle'].notnull())]\nax = sns.kdeplot(notinj.velocity_at_tackle, notinj.deceleration_at_tackle,\n                 cmap=\"Blues\")\nax.set_xlim(0, 10)\nax.set_ylim(0, 2)\nplt.xlabel(\"Velocity at Tackle\")\nplt.ylabel(\"Deceleration at Tackle\")\nsns.despine(top=True, right=True)\nplt.title(\"Velocity/Deceleration at Time of Tackle for Injuries shown in Red, Non-Injuries in Blue\")\nplt.subplots_adjust(top=0.9)\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a5d0fb28cac1be81d62284cde88f71d11d9f6943"},"cell_type":"markdown","source":"We show that players with less than median deceleration (approximately 0.5) have much lower injury rates than players with greater than median deceleration. Similar results can be found using logistic regressions, that deceleration is a significant predictor of injury. If we required players to wear mouthguard or helmet sensors that measured deceleration, we would have a better way to empirically measure whether rule changes, formation changes, etc. reduce the likelihood of injury. "},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"797c3973bb889420871e569e5452d16f37ea54d4"},"cell_type":"code","source":"# Multiply by 22 since there are 22 players per play.\nii = (combined['deceleration_at_tackle']>=0.0) & (combined['deceleration_at_tackle']<0.5)\njj = (combined['deceleration_at_tackle']>=0.5) & (combined['deceleration_at_tackle']<100)\nrate1 = 22*1000*combined.loc[ii, 'tackle_injury'].mean()\nrate2 = 22*1000*combined.loc[jj, 'tackle_injury'].mean()\nprint(\"Injury rate per 1000 plays when deceleration at tackle is less than 0.5: {:.2f} (N={:.0f})\".format(rate1,\n                                                                                                                sum(ii)))\nprint(\"Injury rate per 1000 plays when deceleration at tackle is greater than 0.5: {:.2f} (N={:.0f})\".format(rate2,\n                                                                                                                sum(jj)))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bc03b5d4eff7758c3e0f4862f24759a77c45c147"},"cell_type":"markdown","source":"### **Additional Observations and Opinions**\n<a id=\"additional_observations\"></a>\n#### The Helmet Rule of 2018 \n* Many of the concussions on punt returns could possibly have been avoided if the new Helmet Rule introduced in 2018 had been in effect. If data from the 2018 season were available, it could be used to quantify the effect of the Helmet Rule in reducing the number of concussions. \n\n#### Making gunners wait at the line of scrimmage\n* In my opinion, the possibility of a fake punt and a throw by the punter adds excitement to a punt play. Requiring gunners to wait at the line of scrimmage until the ball is kicked would make these fake punt plays impractical and overly reduce the unpredictability and excitement of punt plays. \n\n#### The CFL Rule\n* The CFL rule giving the returner a 5 yard buffer of protection seems to pre-assume that a punt is to be made. What if instead the punter rolls out of the pocket and throws a long pass on the run to a 'gunner'? Unpredictability, e.g. not knowing whether the punting team will really punt, is for me a key factor in enjoying the game, and I love watching the games of great coaches who are not afraid to gamble and take risks in 4th and short situations. \n\n### ** Final Thoughts **\n<a id=\"final_thoughts\"></a>\nIn summary, I have proposed three rule changes that all strive to increase safety while importantly preserving unpredictability, excitement, and the possibility for players to make great plays. The NFL is about great individual and team performance, and we need to protect the players' ability to showcase their talent and to be rewarded for their hard work, both in the near term when they are on the field and in the long term when they serve as ambassadors for the game, long after they have hung up their cleats. Thank you for the opportunity to work on this important and interesting challenge.\n\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}