{"cells":[{"metadata":{"_uuid":"a440f4a5456c7bd9b90d5bd0ac22086a7da0992b"},"cell_type":"markdown","source":"# NFL Punt Analytics Competition\n(Last Modified: 2019/01/09) \n\n*Charlie Bonfield* <br/>\n\nThis notebook is organized into the following sections: \n1. Background \n2. Proposed Changes \n3. Exploratory Data Analysis\n    * NGS Data (Velocity/Acceleration)\n    * NGS Data (Orientation/Direction) \n    * Game Conditions \n4. Summary\n\nGitHub repo (additional code): https://github.com/cbonfield/nfl_punt_analysis"},{"metadata":{"_uuid":"b66d13f93786bdd6bee5a7a9b67071e1f722978d"},"cell_type":"markdown","source":"## Background\n\nIn recent years, there has been a more concerted effort to reduce the number of player concussions in the NFL. In light of the changes made to the helmet rule and on kickoffs prior to the 2018 season, it is clear that times are changin'! However, there have yet to be any changes made to the rules surrounding punt plays to reduce the risk of concussions. In the kernel below, I present a number of fixes that I feel would go a long way in mitigating the risk of concussions on punts, and with it, I perform some analysis to motivate the proposed changes. \n\nA few notes from the author about **preprocessing**:\n* The process of calculating player velocities/accelerations from the NGS data is not overly complicated, but it takes a while to run given the size of the datasets. Consequently, I did so offline so that I would not need to run all of those steps from within the kernel, and the scripts that I used to do all of the preprocessing are in my public GitHub repo for this project (here's my [script](https://github.com/cbonfield/nfl_punt_analysis/blob/master/code/preprocess_ngs_data.py)). \n* Tying into the last point, I further trimmed the NGS data down to what I would actually need to run my analysis end-to-end from within the kernel. When coding, my personal preference is to use Atom+Hydrogen (here's a nice blog post about it: https://acarril.github.io/posts/atom-hydrogen) and it was simpler to do all of the work there before sticking it into a Kaggle kernel. I will make said datasets publicly available when this kernel is published.  "},{"metadata":{"_uuid":"8011da600c80bf8a9f0c159188cde437ff242e61"},"cell_type":"markdown","source":"## **Proposed Changes**\n\nI decided to put my set of proposed changes right at the top of my kernel for those interested in reading the changes that I would propose for punt returns without viewing all that I did for my analysis (as I'm sure the sheer volume of kernels will make this a daunting task!). Without further ado, I propose the following changes that should/could be implemented: \n\n### Proposed Change 1: \n*Limit the number of players allowed to move across the line of scrimmage prior to the punt.*\n* Since concussions that occur on punt plays occur after the ball is punted, there is more of a need to create conditions post punt that reduce player speed on all parts of the field. \n* Decreasing the number of players crossing the line will (1) naturally move players downfield sooner and (2) prevent spaces from opening between the first and second wave of coverage that would enable players on either team to speed up and/or continue to move at high speeds.  \n* In light of the proposed formation change in PC #2, I would recommend the limit to be five players. This would allow the punting team to keep an additional blocker to protect the punter while also moving players outside to decrease the amount of unoccupied space on the return.   \n\n### Proposed Change 2: \n*Require a player to move off each side of the line for both teams that must be positioned at the line of scrimmage and between the hashes and numbers.*\n* When outside players on the punting team (gunners) move downfield when the ball is snapped, almost all of the space off the line (around the outside of the hashes) is available for players to use to gain speed prior to impacting players on the opposing team. \n* Forcing players to move off the line should, in theory, not make the rest of the field available for players to cut out wide and move downfield without resistance.  \n\n### Proposed Change 3\n*Add a ten yard \"no blocking\" zone from the line of scrimmage to enable players to move downfield.*\n* Preventing blocking close to the line of scrimmage will force players to move downfield quicker and consequently, reduce the risk of players moving without interactions with members of the opposing team for extended periods of time. This is especially important given how quickly players are able to accelerate to their top speed. \n* This probably goes without saying, but this rule is only enforced immediately following the punt - once the play starts to develop on the return, players (should) have moved beyond this zone regardless.\n* This is in the same spirit as the fifteen yard \"no blocking\" zone imposed on kickoffs prior to the 2018 season. While the length of the playing field is shorter than on kickoffs due to the nature of punts, there is still much more available to players than on typical offensive/defensive series. \n* While the purpose of this change is not to intentionally reduce the number of returned punts, it stands to reason that could, in fact, be part of the outcome. \n\n### Solution Efficacy \n* In the kernel below, I present player speed as the singlemost (and likely the only) factor that contributes to concussions on punt plays.  \n    * It is worth noting that all concussions involving two players contained either helmet-to-helmet or helmet-to-body impact. The changes to the helmet rule should do a lot to decrease the likelihood of those types of plays, but regardless, you have to be moving quickly to generate enough force to cause a concussion.  \n    * While I also considered at the effect of relative player-partner orientation/direction, the only knowledge that I gleaned from it is that players are more likely to be moving in the same direction prior to impact. In my opinion, this is expected due to the fact that concussions occur during the return. \n    * I also include a plug about the use of helmet sensors for concussion research in the NFL - the instantaneous accelerations experienced by players are far lower than we would expect for concussions, and while this is undoubtedly a consequence the frequency of sampling for NGS data, the motion of the head/neck relative to the rest of the body is really what one needs to consider for concussions. (As an added bonus, this could even present a route to automatically identifying in-game concussion risk for players - I digress though). \n* Towards the end of the kernel, I took a cursory look to see if there were gameflow specific changes that could be made to punts (i.e., when they would be allowed) that would decrease the risk of concussion, but in truth, I did not see anything statistically significant. In retrospect, while proposing something like limiting when teams can punt would be exciting, it may cross the line with respect to game integrity at the present time. \n\n### Game Integrity \n* Since my proposed changes relate to formation changes, I feel that the NFL would be able to implement them without much resistance from players, coaches, owners, and fans. Fundamentally, my proposed changes are all intended to reduce concussion risk by impeding player motion over the course of a punt play.   \n* While player safety is paramount, I believe my proposed changes also strike a balance between decreasing the risk of concussions and preventing players from ever actually being able to return the ball. Although it is too early to tell if the changes to the kickoff will have a lasting impact on the length of returns and/or the number of return touchdowns, I can say that as a causal football fan, I did not see any marked changes in the complexity associated with returning or defending kickoffs this season.   \n* In the process of devising my set of proposed changes, I could not think of any particular player(s) that would be at increased risk - there is still an extra player back to block the punter (and for the returner), there are less players allowed to come at the punter, and players will be less likely to go full throttle across the entire length of the field.  \n"},{"metadata":{"trusted":true,"_uuid":"a3cb69e3743e3babf5f5bfe8b32e14e4908ca26f","_kg_hide-input":true},"cell_type":"code","source":"## IMPORTS \nimport glob \nimport numpy as np\nimport pandas as pd\n\nfrom scipy import stats\nfrom scipy.stats import norm\nfrom scipy import interpolate\nfrom sklearn.neighbors import KernelDensity\n\nimport plotly.io as pio\nfrom plotly import tools\nimport plotly.graph_objs as go\nfrom plotly.offline import download_plotlyjs, init_notebook_mode, plot, iplot\n\npd.set_option('display.max_rows', 5000)\npd.set_option('display.max_columns', 500)\ninit_notebook_mode(connected=True)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"eff3959979fbaa0e7107ba49eaae1f7c4224b229"},"cell_type":"code","source":"## FUNCTIONS (utility/preprocessing)\nDDIR = '../input/NFL-Punt-Analytics-Competition/'\n\ndef collect_outcomes(data):\n    \"\"\"\n    Extract the punt outcome from the PlayDescription field.\n\n    Parameters:\n        data: dict (keys: labels, values: DataFrames)\n            Data dictionary - likely the output from load_data().\n    \"\"\"\n\n    play_info =  data['play_info']\n\n    def _process_description(row):\n        tmp_desc = row.PlayDescription\n        outcome = ''\n\n        if 'touchback' in tmp_desc.lower():\n            outcome = 'touchback'\n        elif 'fair catch' in tmp_desc.lower():\n            outcome = 'fair catch'\n        elif 'out of bounds' in tmp_desc.lower():\n            outcome = 'out of bounds'\n        elif 'muff' in tmp_desc.lower():\n            outcome = 'muffed punt'\n        elif 'downed' in tmp_desc.lower():\n            outcome = 'downed'\n        elif 'no play' in tmp_desc.lower():\n            outcome = 'no play'\n        elif 'blocked' in tmp_desc.lower():\n            outcome = 'blocked punt'\n        elif 'fumble' in tmp_desc.lower():\n            outcome = 'fumble'\n        elif 'pass' in tmp_desc.lower():\n            outcome = 'pass'\n        elif 'declined' in tmp_desc.lower():\n            outcome = 'declined penalty'\n        elif 'direct snap' in tmp_desc.lower():\n            outcome = 'direct snap'\n        elif 'safety' in tmp_desc.lower():\n            outcome = 'safety'\n        else:\n            if 'punts' in tmp_desc.lower():\n                outcome = 'return'\n            else:\n                outcome = 'SPECIAL'\n\n        return outcome\n\n    play_info.loc[:, 'Punt_Outcome'] = play_info.apply(_process_description, axis=1)\n\n    def _identify_penalties(row):\n        if 'penalty' in row.PlayDescription.lower():\n            return 1\n        else:\n            return 0\n\n    play_info.loc[:, 'Penalty_on_Punt'] = play_info.apply(_identify_penalties, axis=1)\n\n    # Update dictionary to include additional set of features.\n    data.update({'play_info': play_info})\n\n    return data\n\ndef expand_play_description(data):\n    \"\"\"\n    Expand the PlayDescription field in a standardized fashion. This function\n    extracts a number of relevant additional features from PlayDescription,\n    including punt distance, post-punt field location, and a few other derived\n    features.\n\n    Parameters:\n        data: dict (keys: labels, values: DataFrames)\n            Data dictionary - likely the output from load_data().\n    \"\"\"\n\n    play_info = data['play_info']\n\n    def _split_punt_distance(row):\n        try:\n            return int(row.PlayDescription.split('punts ')[1].split('yard')[0])\n        except IndexError:\n            return np.nan\n\n    def _split_field_position(row):\n        try:\n            return row.PlayDescription.split(',')[0].split('to ')[1]\n        except IndexError:\n            return ''\n\n    def _post_punt_territory(row):\n        if row.Poss_Team == row.Post_Punt_FieldSide:\n            return 0\n        else:\n            return 1\n\n    def _start_punt_field_position(row):\n        try:\n            field_position = int(row.YardLine.split(' ')[1])\n        except:\n            print(row.YardLine)\n\n        if row.Poss_Team in row.YardLine:\n            return field_position\n        else:\n            return 100 - field_position\n\n    def _field_position_punt(row):\n        if 'end zone' in row.Post_Punt_YardLine:\n            return 0\n        elif '50' in row.Post_Punt_YardLine:\n            return 50\n        else:\n            try:\n                yard_line = int(row.Post_Punt_YardLine.split(' ')[1])\n                own_field = int(row.Post_Punt_Own_Territory)\n\n                if not own_field:\n                    return 100 - yard_line\n                else:\n                    return yard_line\n            except:\n                return -999\n\n    play_info.loc[:, 'Punt_Distance'] = play_info.apply(_split_punt_distance, axis=1)\n    play_info.loc[:, 'Post_Punt_YardLine'] = play_info.apply(_split_field_position, axis=1)\n    play_info.loc[:, 'Post_Punt_FieldSide'] = play_info.Post_Punt_YardLine.apply(lambda x: x.split(' ')[0])\n    play_info.loc[:, 'Post_Punt_Own_Territory'] = play_info.apply(_post_punt_territory, axis=1)\n    play_info.loc[:, 'Pre_Punt_RelativeYardLine'] = play_info.apply(_start_punt_field_position, axis=1)\n    play_info.loc[:, 'Post_Punt_RelativeYardLine'] = play_info.apply(_field_position_punt, axis=1)\n\n    # Extract additional information from play info (home team, away team, score\n    # differential, home/away punt identifier).\n    play_info.loc[:, 'Home_Team'] = play_info.Home_Team_Visit_Team.apply(lambda x: x.split('-')[0])\n    play_info.loc[:, 'Away_Team'] = play_info.Home_Team_Visit_Team.apply(lambda x: x.split('-')[1])\n    play_info.loc[:, 'Home_Points'] = play_info.Score_Home_Visiting.apply(lambda x: x.split('-')[0]).astype(int)\n    play_info.loc[:, 'Away_Points'] = play_info.Score_Home_Visiting.apply(lambda x: x.split('-')[1]).astype(int)\n\n    def _home_away_punt_bool(row):\n        if row.Home_Team == row.Poss_Team:\n            return 1\n        else:\n            return 0\n\n    play_info.loc[:, 'Home_Visit_Team_Punt'] = play_info.apply(_home_away_punt_bool, axis=1)\n\n    def _get_score_differential(row):\n        if not row.Home_Visit_Team_Punt:\n            return int(row.Away_Points - row.Home_Points)\n        else:\n            return int(row.Home_Points - row.Away_Points)\n\n    play_info.loc[:, 'Score_Differential'] = play_info.apply(_get_score_differential, axis=1)\n\n    # Update dictionary to include additional set of features.\n    data.update({'play_info': play_info})\n\n    return data\n\ndef load_data(raw_bool=False):\n    \"\"\"\n    When called, this function will load in all of the data and do the relevant\n    preprocessing (mainly just a series of merges to link the injury data with a\n    few of the other data sources). The output of this function is a dictionary\n    with key/value pairs that are labels/DataFrames, respectively.\n\n    Parameters:\n        raw_bool: bool (default False)\n            Boolean indicating whether you wish to perform the necessary\n            preprocessing steps (False) or not (True).\n    \"\"\"\n\n    # Load data.\n    game_data = pd.read_csv(f'{DDIR}game_data.csv')\n    play_info = pd.read_csv(f'{DDIR}play_information.csv')\n    play_role = pd.read_csv(f'{DDIR}play_player_role_data.csv')\n    punt_data = pd.read_csv(f'{DDIR}player_punt_data.csv')\n\n    video_injury = pd.read_csv(f'{DDIR}video_footage-injury.csv')\n    video_review = pd.read_csv(f'{DDIR}video_review.csv')\n    video_control = pd.read_csv(f'{DDIR}video_footage-control.csv')\n\n    if raw_bool:\n        pass\n    else:\n        # Rename columns to match format (between video_injury/video_control and\n        # everything else).\n        ren_dict = {\n            'season': 'Season_Year',\n            'Type': 'Season_Type',\n            'Home_team': 'Home_Team',\n            'gamekey': 'GameKey',\n            'playid': 'PlayId'\n        }\n\n        video_injury.rename(index=str, columns=ren_dict, inplace=True)\n        video_control.rename(index=str, columns=ren_dict, inplace=True)\n        video_review.rename(index=str, columns={'PlayID':'PlayId'}, inplace=True)\n\n        # Join video_review to video_injury.\n        video_injury = video_injury.merge(video_review, how='outer',\n                                          left_on=['Season_Year', 'GameKey', 'PlayId'],\n                                          right_on=['Season_Year', 'GameKey', 'PlayId'])\n\n        # Process punt_data - it's possible to have multiple numbers for the same\n        # player, so we'll drop number to get rid of duplicates.\n        punt_data.drop('Number', axis=1, inplace=True)\n        punt_data.drop_duplicates(inplace=True)\n\n        # Add player primary position to video_injury.\n        video_injury = video_injury.merge(punt_data, how='inner', on=['GSISID'])\n        video_injury.rename(index=str, columns={'Position':'Player_Position'},\n                            inplace=True)\n\n        # Fix a few values in Primary_Partner_GSISID that will cause the next\n        # merge to barf (one nan, one 'Unclear').\n        video_injury.replace(to_replace={'Primary_Partner_GSISID':'Unclear'},\n                             value=99999, inplace=True)\n        video_injury.replace(to_replace={'Primary_Partner_GSISID':np.nan},\n                             value=99999, inplace=True)\n        video_injury.loc[:, 'Primary_Partner_GSISID'] = video_injury.Primary_Partner_GSISID.astype(int)\n\n        # Add primary partner primary position to video_injury.\n        video_injury = video_injury.merge(punt_data, how='left',\n                                          left_on=['Primary_Partner_GSISID'],\n                                          right_on=['GSISID'])\n        video_injury.drop('GSISID_y', axis=1, inplace=True)\n        video_injury.rename(index=str, columns={'GSISID_x':'GSISID'}, inplace=True)\n        video_injury.rename(index=str, columns={'Position':'Primary_Partner_Position'},\n                            inplace=True)\n\n        # Add punt specific play role for players to video_injury.\n        play_role.rename(index=str, columns={'PlayID':'PlayId'}, inplace=True)\n        video_injury = video_injury.merge(play_role, how='left',\n                                          left_on=['Season_Year', 'GameKey', 'PlayId', 'GSISID'],\n                                          right_on=['Season_Year', 'GameKey', 'PlayId', 'GSISID'])\n        video_injury.rename(index=str, columns={'Role':'Player_Punt_Role'}, inplace=True)\n\n        # Add punt specific play role for primary partners to video_injury.\n        video_injury = video_injury.merge(play_role, how='left',\n                                          left_on=['Season_Year', 'GameKey', 'PlayId', 'Primary_Partner_GSISID'],\n                                          right_on=['Season_Year', 'GameKey', 'PlayId', 'GSISID'])\n        video_injury.drop('GSISID_y', axis=1, inplace=True)\n        video_injury.rename(index=str, columns={'GSISID_x':'GSISID'}, inplace=True)\n        video_injury.rename(index=str, columns={'Role':'Primary_Partner_Punt_Role'},\n                            inplace=True)\n\n    # Stick everything in a dictionary to return as output.\n    out_dict = {\n        'game_data': game_data,\n        'play_info': play_info,\n        'play_role': play_role,\n        'punt_data': punt_data,\n        'video_injury': video_injury,\n        'video_control': video_control,\n        'video_review': video_review\n    }\n\n    return out_dict\n\ndef parse_penalties(play_info_df):\n    \"\"\"\n    Extract penalty types for plays on which we had penalties.\n\n    Parameters:\n        play_info_df: pd.DataFrame\n            DataFrame containing play information.\n    \"\"\"\n\n    pen_df = play_info_df.loc[play_info_df.Penalty_on_Punt == 1].reset_index(drop=True)\n\n    def _extract_penalty_type(row):\n        try:\n            tmp_desc = row.PlayDescription.lower()\n            pen_suff = tmp_desc.split('penalty on ')[1]\n            drop_plr = pen_suff.split(', ')[1]\n\n            penalty = drop_plr.split(',')[0]\n        except:\n            penalty = 'EXCEPTION'\n\n        return penalty\n\n    pen_df.loc[:, 'Penalty_Type'] = pen_df.apply(_extract_penalty_type, axis=1)\n\n    return pen_df\n\ndef trim_player_partner_data(ngs_df):\n    \"\"\"\n    Given a DataFrame with NGS data for player/partner on punt play, cut out\n    the relevant NGS data.\n\n    Parameters:\n        ngs_df: pd.DataFrame\n            DataFrame containing NGS data.\n    \"\"\"\n\n    # Isolate player/partner data.\n    play_df = ngs_df.loc[ngs_df.Identifier == 'PLAYER'].dropna().reset_index(drop=True)\n    part_df = ngs_df.loc[ngs_df.Identifier == 'PARTNER'].dropna().reset_index(drop=True)\n\n    # Figure out where the ball snap occurred and get the index so that we can\n    # discard all data prior to that instant.\n    try:\n        play_st = play_df.loc[play_df.Event == 'punt'].index[0]\n        part_st = part_df.loc[part_df.Event == 'punt'].index[0]\n    except IndexError:\n        try:\n            play_st = play_df.loc[play_df.Event == 'ball_snap'].index[0]\n            part_st = part_df.loc[part_df.Event == 'ball_snap'].index[0]\n        except IndexError:\n            play_st = play_df.index.min()\n            part_st = part_df.index.min()\n\n    # Figure out where the play \"ended\" so that we can discard all data after\n    # that. For simplicity, we assume that any concussion event would have occured\n    # prior to a penalty flag being thrown or within five seconds of a tackle.\n    try:\n        play_ei = play_df.loc[play_df.Event == 'tackle'].index[0] + 50\n        part_ei = part_df.loc[part_df.Event == 'tackle'].index[0] + 50\n        \n        play_ps = play_df.loc[play_df.Event == 'play_submit'].index[0]\n        part_ps = part_df.loc[part_df.Event == 'play_submit'].index[0]\n        \n        while play_ei > play_ps:\n            play_ei -= 10\n\n        while part_ei > part_ps:\n            part_ei -= 10\n    except IndexError:\n        try:\n            play_ei = play_df.loc[play_df.Event == 'penalty_flag'].index[0]\n            part_ei = part_df.loc[part_df.Event == 'penalty_flag'].index[0]\n        except IndexError:\n            play_ei = play_df.index.max()\n            part_ei = part_df.index.max()\n\n    # Slice out the data that we actually need.\n    play_df = play_df.iloc[play_st:play_ei]\n    part_df = part_df.iloc[part_st:part_ei]\n\n    return play_df, part_df","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"84832709dc71426322a0ce82783e5a425f2d4d07"},"cell_type":"markdown","source":"## Exploratory Analysis \n\nIn the course of making some snappy visualizations for my slide deck, I tallied up a few statistics by hand that I figured would be worth including here. Most notably, I saw that:\n* **27** out of **37** concussed players were on the punting team. \n* Of the ten concussed players on the receiving team, the punt returner was concussed **five** times.\n* Punt returners were involved in concussion events more than any other player on the field (**13** out of **37** times, **five** as a \"concussee\" and **eight** as a \"concusser\"). \n\nWithout doing anything fancy, we can already start to get a sense for what sorts of things will give rise to increased risk of concussion - players that move the fastest and cover the most ground, for instance, are the most likely to be concussed. This aligned with my gut instinct, but it was interesting to see all the same. "},{"metadata":{"trusted":true,"_uuid":"2ed5fe674144775dfbfdf132ad7ef3ddd9ec20c9"},"cell_type":"markdown","source":"To get a sense for what players were actually doing on field, however, I decided to make a plot of the player-partner trajectories superposed on a football field. The interactive plot below shows just that - the path of the player is in red, the path of the partner is in blue, and the colorbars indicate the speed of each respective player (in m/s). For ease of presentation, I indexed the punt plays (hence the number displayed by the slider), and I would suggest looking through it one frame at a time."},{"metadata":{"_uuid":"c5e4d2ca084c7dc7f5724d3f87180971d7dc00c2","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Load data from plays with concussions.\nWDIR = '../input/ngs-dataset-playerpartnerinjuries/'\nngs_data = pd.read_csv(f'{WDIR}injury_ngs_data.csv')\n\n# Add column for easy indexing.\nmerge_cols = ['Season_Year', 'GameKey', 'PlayID']\nind_df = ngs_data.drop_duplicates(merge_cols).reset_index(drop=True)\nind_df.loc[:, 'eventIndex'] = ind_df.index.values\nplay_indexes = ind_df.eventIndex.tolist()\n\nind_df = ind_df.loc[:, ['eventIndex', 'Season_Year', 'GameKey', 'PlayID']]\nngs_data = ngs_data.merge(ind_df, how='inner', left_on=merge_cols, right_on=merge_cols)\n\n# Run to generate animated figure.\nfigure = {\n    'data': [],\n    'layout': {},\n    'frames': []\n}\n\n## CUSTOM\nfield_xaxis=dict(\n        range=[0,120],\n        linecolor='black',\n        linewidth=2,\n        mirror=True,\n        showticklabels=False\n)\nfield_yaxis=dict(\n        range=[0,53.3],\n        linecolor='black',\n        linewidth=2,\n        mirror=True,\n        showticklabels=False\n)\nfield_annotations=[\n        dict(\n            x=0,\n            y=0.5,\n            showarrow=False,\n            text='HOME ENDZONE',\n            textangle=270,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=24,\n                color='white'\n            )\n        ),\n        dict(\n            x=1,\n            y=0.5,\n            showarrow=False,\n            text='AWAY ENDZONE',\n            textangle=90,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=24,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(17./120.),\n            y=1,\n            showarrow=False,\n            text='10',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(27./120.),\n            y=1,\n            showarrow=False,\n            text='20',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(37./120.),\n            y=1,\n            showarrow=False,\n            text='30',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(50./120.),\n            y=1,\n            showarrow=False,\n            text='40',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(60./120.),\n            y=1,\n            showarrow=False,\n            text='50',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(70./120.),\n            y=1,\n            showarrow=False,\n            text='40',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(80./120.),\n            y=1,\n            showarrow=False,\n            text='30',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(93./120.),\n            y=1,\n            showarrow=False,\n            text='20',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        ),\n        dict(\n            x=float(103./120.),\n            y=1,\n            showarrow=False,\n            text='10',\n            textangle=180,\n            xref='paper',\n            yref='paper',\n            font=dict(\n                family='sans serif',\n                size=20,\n                color='white'\n            )\n        )\n]\nfield_shapes=[\n        {\n            'type': 'line',\n            'x0': 10,\n            'y0': 0,\n            'x1': 10,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 2\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 110,\n            'y0': 0,\n            'x1': 110,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 2\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 20,\n            'y0': 0,\n            'x1': 20,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 30,\n            'y0': 0,\n            'x1': 30,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 40,\n            'y0': 0,\n            'x1': 40,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 50,\n            'y0': 0,\n            'x1': 50,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 60,\n            'y0': 0,\n            'x1': 60,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 70,\n            'y0': 0,\n            'x1': 70,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 80,\n            'y0': 0,\n            'x1': 80,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 90,\n            'y0': 0,\n            'x1': 90,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 100,\n            'y0': 0,\n            'x1': 100,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 15,\n            'y0': 0,\n            'x1': 15,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 25,\n            'y0': 0,\n            'x1': 25,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 35,\n            'y0': 0,\n            'x1': 35,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 45,\n            'y0': 0,\n            'x1': 45,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 55,\n            'y0': 0,\n            'x1': 55,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 65,\n            'y0': 0,\n            'x1': 65,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 75,\n            'y0': 0,\n            'x1': 75,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 85,\n            'y0': 0,\n            'x1': 85,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 95,\n            'y0': 0,\n            'x1': 95,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        },\n        {\n            'type': 'line',\n            'x0': 105,\n            'y0': 0,\n            'x1': 105,\n            'y1': 53.3,\n            'line': {\n                'color': 'white',\n                'width': 1,\n                'dash':'dot'\n            },\n        }\n]\n\n\n# fill in most of layout\nfigure['layout']['autosize'] = True\nfigure['layout']['showlegend'] = False\nfigure['layout']['plot_bgcolor'] = '#008000'\n\nfigure['layout']['xaxis'] = field_xaxis\nfigure['layout']['yaxis'] = field_yaxis\nfigure['layout']['annotations'] = field_annotations\nfigure['layout']['shapes'] = field_shapes\n\nfigure['layout']['hovermode'] = 'closest'\nfigure['layout']['sliders'] = {\n    'args': [\n        'transition', {\n            'duration': 400,\n            'easing': 'cubic-in-out'\n        }\n    ],\n    'initialValue': 0,\n    'plotlycommand': 'animate',\n    'values': play_indexes,\n    'visible': True\n}\nfigure['layout']['updatemenus'] = [\n    {\n        'buttons': [\n            {\n                'args': [None, {'frame': {'duration': 500, 'redraw': False},\n                         'fromcurrent': True, 'transition': {'duration': 300, 'easing': 'quadratic-in-out'}}],\n                'label': 'Play',\n                'method': 'animate'\n            },\n            {\n                'args': [[None], {'frame': {'duration': 0, 'redraw': False}, 'mode': 'immediate',\n                'transition': {'duration': 0}}],\n                'label': 'Pause',\n                'method': 'animate'\n            }\n        ],\n        'direction': 'left',\n        'pad': {'r': 10, 't': 87},\n        'showactive': False,\n        'type': 'buttons',\n        'x': 0.1,\n        'xanchor': 'right',\n        'y': 0,\n        'yanchor': 'top'\n    }\n]\n\nsliders_dict = {\n    'active': 0,\n    'yanchor': 'top',\n    'xanchor': 'left',\n    'currentvalue': {\n        'font': {'size': 20},\n        'prefix': 'Play Index: ',\n        'visible': True,\n        'xanchor': 'right'\n    },\n    'transition': {'duration': 300, 'easing': 'cubic-in-out'},\n    'pad': {'b': 10, 't': 50},\n    'len': 0.9,\n    'x': 0.1,\n    'y': 0,\n    'steps': []\n}\n\n# Make data (for single play).\nplt_dicts = []\npidx = 0\n\nsp_data = ngs_data.loc[ngs_data.eventIndex == pidx].reset_index(drop=True)\n\n# Grab some stuff for labeling saved figure.\nsy = sp_data.Season_Year.values[0]\ngk = sp_data.GameKey.values[0]\npi = sp_data.PlayID.values[0]\n\nplt_dict = {}\nplt_dict['playIndex'] = pi\nplt_dict['seasonYear'] = sy\nplt_dict['gameKey'] = gk\nplt_dict['playID'] = pi\nplt_dicts.append(plt_dict)\n\n# Get player/partner data (reduced).\nrd_play_df, rd_part_df = trim_player_partner_data(sp_data)\n\nfor i in range(2):\n    if not i:\n        ngs_dataset = rd_play_df.copy()\n        plt_name = 'Player'\n        color_scale = 'Reds'\n        cb_loc = 1.0\n    else:\n        ngs_dataset = rd_part_df.copy()\n        plt_name = 'Partner'\n        color_scale = 'Blues'\n        cb_loc = 1.1\n\n    data_dict = {\n        'x': list(ngs_dataset['x']),\n        'y': list(ngs_dataset['y']),\n        'mode': 'markers',\n        'marker': {\n            'color': list(ngs_dataset['s']),\n            'colorbar': {'x':cb_loc},\n            'colorscale':color_scale,\n            'size':12\n        },\n        'name':plt_name\n    }\n    \n    if i == 1:\n        data_dict['marker']['reversescale'] = True\n        \n    figure['data'].append(data_dict)\n\n# Add data for yardline trace.\ndata_dict = {\n    'x': [20, 30, 40, 50, 60, 70, 80, 90, 100],\n    'y': [1, 1, 1, 1, 1, 1, 1, 1, 1],\n    'mode': 'text',\n    'text': ['10','20','30','40','50','40','30','20','10'],\n    'textposition': 'top center',\n    'textfont': {\n        'family': 'sans serif',\n        'size': 20,\n        'color': 'white'\n    }\n}\nfigure['data'].append(data_dict)\n\n# Make frames.\nfor pidx in play_indexes:\n    frame = {'data': [], 'name': pidx}\n    try:\n        sp_data = ngs_data.loc[ngs_data.eventIndex == pidx].reset_index(drop=True)\n\n        # Grab some stuff for labeling saved figure.\n        sy = sp_data.Season_Year.values[0]\n        gk = sp_data.GameKey.values[0]\n        pi = sp_data.PlayID.values[0]\n\n        plt_dict = {}\n        plt_dict['playIndex'] = pi\n        plt_dict['seasonYear'] = sy\n        plt_dict['gameKey'] = gk\n        plt_dict['playID'] = pi\n        plt_dicts.append(plt_dict)\n\n        # Get player/partner data (reduced).\n        rd_play_df, rd_part_df = trim_player_partner_data(sp_data)\n\n        for i in range(2):\n            if not i:\n                ngs_dataset = rd_play_df.copy()\n                plt_name = 'Player'\n                color_scale = 'Reds'\n                cb_loc = 1.0\n            else:\n                ngs_dataset = rd_part_df.copy()\n                plt_name = 'Partner'\n                color_scale = 'Blues'\n                cb_loc = 1.1\n\n            data_dict = {\n                'x': list(ngs_dataset['x']),\n                'y': list(ngs_dataset['y']),\n                'mode': 'markers',\n                'marker': {\n                    'color': list(ngs_dataset['s']),\n                    'colorbar': {'x':cb_loc},\n                    'colorscale':color_scale,\n                    'size':12\n                },\n                'name':plt_name\n            }\n            \n            if i == 1:\n                data_dict['marker']['reversescale'] = True\n                \n            frame['data'].append(data_dict)\n\n        # Add data for yardline trace.\n        data_dict = {\n            'x': [20, 30, 40, 50, 60, 70, 80, 90, 100],\n            'y': [1, 1, 1, 1, 1, 1, 1, 1, 1],\n            'mode': 'text',\n            'text': ['10','20','30','40','50','40','30','20','10'],\n            'textposition': 'top center',\n            'textfont': {\n                'family': 'sans serif',\n                'size': 20,\n                'color': 'white'\n            }\n        }\n        frame['data'].append(data_dict)\n\n        # Add frames.\n        figure['frames'].append(frame)\n        slider_step = {'args': [\n            [pidx],\n            {'frame': {'duration': 300, 'redraw': False},\n             'mode': 'immediate',\n           'transition': {'duration': 300}}\n         ],\n         'label': pidx,\n         'method': 'animate'}\n        sliders_dict['steps'].append(slider_step)\n    except TypeError:\n        continue\n\nfigure['layout']['sliders'] = [sliders_dict]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4ffec0e7cc409449ced154057bf4b9723f3eb88f"},"cell_type":"code","source":"iplot(figure)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3a1d6d006f5e38fd39b67528725c400822d3bbb9"},"cell_type":"markdown","source":"**Note**: *I did what I could to trim off irrelevant parts of the NGS data when making these visualizations - I sliced out data that occured prior to the punt and after either the tackle on the ball carrier was made or a penalty flag was thrown*. \n\nClicking through the figures, we can see a few things just by eye:\n* It's common for players involved in concussions to cover large distances on the field, thus allowing them to reach top speed, prior to making contact with members of the other team (examples: play index 0, 13, 28, 36). \n* For the most part, players look to move without much interaction with other players prior to making contact. While I would have liked to done a more thorough analysis to prove systematically that this is indeed the case, this is evident from the lack of abrupt changes to player velocities when looking through player trajectories (we would expect to see lighter patches of color in the case of significant interactions). \n* There does not appear to be much rhyme or reason to how players move on the field, which is in line with what we would expect for punt plays - players move towards the location of the ball carrier and take whichever route would appear to minimize that distance.\n* Since each point is a time step (0.1 s), the fact that we see mostly dark trajectories indicates that players are able to accelerate quickly to their maximum (presumably sprint) speeds. Again, while this would've been nice to analyze further, it highlights the need for more prolonged player interaction over the course of the playing field to effectively inhibit that kind of rapid acceleration and sustained speed. \n\nHowever, let's dive a little deeper into the NGS data to see if we can find the most likely conditions that will lead to concussions."},{"metadata":{"_uuid":"8742b9fb92569b8dfdc5006560444f112a71c50c"},"cell_type":"markdown","source":"### NGS Data: Velocity/Acceleration\n\nFor the first part of the analysis, I chose to use the NGS position/time data for players (and partners) involved in documented concussion events to try to determine if we could identify conditions leading to increased risk. Since concussions are tied to large impacts/accelerations, I started there - for each player on punt plays, I looked to the maximum acceleration experienced by each player (independent of player role):"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"58c45d2585bdc081ac62a096a106f8260ea2a2cc"},"cell_type":"code","source":"## FUNCTIONS \ndef ecdf(data):\n    \"\"\"\n    Compute empirical cumulative distribution function for set of 1D data.\n\n    Parameters:\n        data: np.array\n    Returns:\n        x: np.array\n            Sorted data.\n        y: np.array\n            ECDF(x).\n    \"\"\"\n\n    n = len(data)\n    x = np.sort(data)\n    y = np.arange(1, n+1) / n\n\n    return x, y\n\ndef make_histogram(ngs_df, col_to_plot, title, add_lines=False):\n    \"\"\"\n    Generate plotly histogram using maximum accelerations experienced by\n    players.\n\n    Parameters:\n        ngs_df: pd.DataFrame\n            DataFrame containing NGS data.\n        col_to_plot: str\n            Name of column to plot.\n        title: str \n            Title to display over histogram. \n        add_lines: bool (default False)\n            Boolean indicating whether we wish to superpose lines for values\n            from concussed players.\n    \"\"\"\n\n    trace = (\n        go.Histogram(\n            x=ngs_df[col_to_plot].values,\n            histnorm='probability',\n            nbinsx=100\n        )\n    )\n\n    layout = go.Layout(\n                title=title, \n                autosize=True,\n                xaxis=dict(\n                    range=[0,100],\n                    title='Maximum Acceleration (m/s^2)'\n                ),\n                yaxis=dict(\n                    title='Normalized Density'\n                )\n             )\n\n    if add_lines:\n        _ = 'placeholder'\n    else:\n        data = [trace]\n\n    fig = go.Figure(data=data, layout=layout)\n\n    return fig\n\ndef plot_ecdf(int_ecdf, inj_ecdf, title, unit, max_x):\n    \"\"\"\n    Plot ECDF from entire set of play data and from the subset where a concussion\n    occurred.\n\n    Parameters:\n        int_ecdf: np.array (values: [quantity, ecdf])\n            Values from interpolated ECDF from entire dataset.\n        inj_ecdf: np.array (values: [quantity, ecdf])\n            ECDF/values for players in the injury set. \n        title: str \n            Title to display over figure. \n        unit: str \n            Unit to stick on horizontal axis. \n        max_x: int\n            Max x-value (used for plot range).\n    \"\"\"\n\n    int_trace = go.Scatter(\n                    x = int_ecdf[:,0],\n                    y = int_ecdf[:,1],\n                    mode = 'markers',\n                    marker = dict(opacity=0.3)\n                )\n\n    inj_trace = go.Scatter(\n                    x = inj_ecdf[:,0],\n                    y = inj_ecdf[:,1],\n                    mode = 'markers',\n                    marker = dict(color='red')\n                )\n\n    data = [int_trace, inj_trace]\n\n    layout = go.Layout(\n                title=title,\n                autosize=True,\n                showlegend=False,\n                xaxis=dict(\n                    range=[0,max_x],\n                    title=f'{title} ({unit})'\n                ),\n                yaxis=dict(\n                    title='CDF'\n                )\n             )\n\n    fig = go.Figure(data=data, layout=layout)\n\n    return fig","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"3dcad6c34920c9bd1324ffe31af6932b589b25e6"},"cell_type":"code","source":"# Load data.\nWDIR = '../input/ngs-data-summary-statistics/sumdynamics/sumdynamics/'\nfiles = glob.glob(f'{WDIR}*.csv')\nflist = []\n\nfor f in files:\n    tmp_df = pd.read_csv(f)\n    flist.append(tmp_df)\n\nsum_df = pd.concat(flist, ignore_index=True)\n\n# Drop unreasonable accelerations.\nsum_df = sum_df.loc[sum_df.max_a <= 150.]\n\n# Construct ECDF for players.\nsrt_spds, spd_ecdf = ecdf(sum_df.max_s.values)\nint_s_ecdf = interpolate.interp1d(srt_spds, spd_ecdf)\nsrt_accs, acc_ecdf = ecdf(sum_df.max_a.values)\nint_a_ecdf = interpolate.interp1d(srt_accs, acc_ecdf)\n\n# Load in set of summary statistics for players involved in concussions.\nWDIR = '../input/ngs-data-speedacceleration-summary-stats/'\ninj_df = pd.read_csv(f'{WDIR}spd_acc_summary.csv')\n\nren_dict = {'season_year':'Season_Year', 'game_key':'GameKey',\n            'play_id': 'PlayID'}\ninj_df.rename(index=str, columns=ren_dict, inplace=True)\n\n# Add column for player/partner action.\nWDIR = '../input/NFL-Punt-Analytics-Competition/'\nppa_df = pd.read_csv(f'{WDIR}video_review.csv')\nppa_df = ppa_df.loc[:, ['Season_Year', 'GameKey', 'PlayID', 'Player_Activity_Derived', 'Primary_Partner_Activity_Derived']]\n\nmer_cols = ['Season_Year', 'GameKey', 'PlayID']\ninj_df = inj_df.merge(ppa_df, how='inner', left_on=mer_cols, right_on=mer_cols)\n\ndef _identify_moving_pp(row):\n    player_activity = row.Player_Activity_Derived\n\n    if 'ing' in player_activity:\n        return 1\n    elif 'ed' in player_activity:\n        return 0\n    else:\n        raise ValueError('Check derived activity!')\n\ndef _moving_play_s(row, dyn_opt):\n    pp_activity = row.PP_Activity\n\n    if dyn_opt == 's':\n        return pp_activity * row.max_play_s + abs(pp_activity - 1) * row.max_part_s\n    elif dyn_opt == 'a':\n        return pp_activity * row.max_play_a + abs(pp_activity - 1) * row.max_part_a\n    else:\n        raise ValueError('Check derived activity!')\n\ninj_df.loc[:, 'PP_Activity'] = inj_df.apply(_identify_moving_pp, axis=1)\ninj_df.loc[:, 'max_move_s'] = inj_df.apply(_moving_play_s, args=('s'), axis=1)\ninj_df.loc[:, 'max_move_a'] = inj_df.apply(_moving_play_s, args=('a'), axis=1)\n\ndef _get_s_cdf(x):\n    return int_s_ecdf(x)\n\ndef _get_a_cdf(x):\n    return int_a_ecdf(x)\n\ninj_df.loc[:, 's_cumprob'] = inj_df.max_play_s.apply(_get_s_cdf)\ninj_df.loc[:, 'a_cumprob'] = inj_df.max_play_a.apply(_get_a_cdf)\ninj_df.loc[:, 's_part_cumprob'] = inj_df.max_part_s.apply(_get_s_cdf)\n#inj_df.loc[:, 'a_part_cumprob'] = inj_df.max_part_a.apply(_get_a_cdf)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"55c90d4961883560d2c0448a1612456007dd1b28"},"cell_type":"code","source":"# Make figure (histogram), then plot.\nplotcol = 'max_a'\n\nfigure = make_histogram(sum_df, plotcol, 'Maximum Acceleration')\niplot(figure, filename='acc-histogram')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"32e23c3227a00f64797046be2b0ef859a8edea4e"},"cell_type":"markdown","source":"And so, we see that it's typical for players to experience accelerations on the order of *2g* (*g* = 9.81 m/s<sup>2</sup>; in physics parlance, *g* is the acceleration due to gravity and can be thought of as a standard unit for acceleration), and while rare, they can extend well above *10g*. However, for concussions, we would expect to see something on the order of about *100g*. Nevertheless, let's see if players who are concussed experience atypically large accelerations:"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"671c5acd836e98def5f91d29c785038116b7819b"},"cell_type":"code","source":"# Make figure (acceleration ECDF), then plot.\na_vals = np.linspace(min(srt_accs), max(srt_accs),200)\na_ecdf = np.array([int_a_ecdf(x) for x in a_vals])\na_ecdf = np.vstack([a_vals,a_ecdf]).T\n\ninj_a_data = inj_df.loc[:, ['max_play_a', 'a_cumprob']].values\n\nfigure = plot_ecdf(a_ecdf, inj_a_data, 'Maximum Player Acceleration', 'm/s^2', 60)\niplot(figure, filename='acc-ecdf-play')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1cb5f44b8b80cf3996d0bf4999f4abf250443c65"},"cell_type":"markdown","source":"On the plot shown above, the blue points correspond to data from all players on punts, while the red points are for data from concussed players. While there are some players who experience accelerations larger than a few *g*, these accelerations are nowhere near large enough to be commeasurate with a concussion and their spread within the empirical CDF suggests that there's not much there either. \n\nIn my opinion, using accelerations directly from the NGS data to characterize concussions is not viable - the relative motion of the head and neck is really what matters, and since the sensors used to register all of the data are located in the shoulder pads (and the neck/head is free to move independent of the shoulders, of course), we cannot definitively look to accelerations alone. That fact notwithstanding, we also may be missing part of the picture due to the rate at which we are receiving NGS data (0.1 s time steps), as the typical time over which players exchange momentum is smaller than that. Instead, it may be more fruitful to look at maximum player speed, as players need to be moving quickly in order to impart enough energy/momentum to cause a concussion."},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"c357c466cf24ad1e10f3fc622844be1f1099a6b2"},"cell_type":"code","source":"# Make figure (speed ECDF), then plot.\ns_vals = np.linspace(min(srt_spds), max(srt_spds),200)\ns_ecdf = np.array([int_s_ecdf(x) for x in s_vals])\ns_ecdf = np.vstack([s_vals,s_ecdf]).T\n\ninj_s_data = inj_df.loc[:, ['max_play_s', 's_cumprob']].values\n\nfigure = plot_ecdf(s_ecdf, inj_s_data, 'Maximum Player Speed', 'm/s', 10)\niplot(figure, filename='spd-ecdf')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4943fc3dcf040086416674cd94e5c7b7d6fceeea"},"cell_type":"code","source":"# Make figure (speed ECDF), then plot.\ns_vals = np.linspace(min(srt_spds), max(srt_spds),200)\ns_ecdf = np.array([int_s_ecdf(x) for x in s_vals])\ns_ecdf = np.vstack([s_vals,s_ecdf]).T\n\ninj_s_data = inj_df.loc[:, ['max_part_s', 's_part_cumprob']].values\n\nfigure = plot_ecdf(s_ecdf, inj_s_data, 'Maximum Partner Speed', 'm/s', 10)\niplot(figure, filename='spd-ecdf')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e025be86201102283c239928b18c2ec6cd184c02"},"cell_type":"markdown","source":"*For context, 8 m/s clocks in at right around 18 mph. The top speed recorded in the NFL during the 2018 season was 22.09 mph, registered on a run by Matt Breida of the 49ers (courtesy of NFL's Next Gen Stats website).*\n\nWhile it's still not a smoking gun per se, we see a little more of a tendency towards higher maximum speeds. Thus, as we might expect, both the players that are concussed and the partners doing the concussing are moving quickly on the play prior to the concussion."},{"metadata":{"trusted":true,"_uuid":"0d78c9f0cd71117d81c22037526bf6dfc30f06e0"},"cell_type":"markdown","source":"### NGS Data: Direction/Orientation\n\nTo complement the velocities and accelerations derived from the NGS data, I decided to dig into the angular (orientation/direction) data a little bit as well. As mentioned in the data description, orientation tells us which way the players are facing on the field at any given moment, while direction tells us the way in which the players are moving on the field. \n\nFor each player-partner combination on plays with concussions, I took the difference of the player angle and partner angle for both orientation and direction - the thought here is that if we are able to identify conditions related to the orientation of players prior to the concussion, we can provide guidance into formation changes that will ameliorate the risk of concussions. \n\nSince it's difficult to visualize what actual playing conditions relative angles correspond with, here are a few examples (note that player/partner roles can be reversed in any of these scenarios):\n* 0 degrees:\n    * Direction: Player and partner are moving in precisely the **same** direction. A simple example to visualize is the player chasing the partner (or vice versa), while something more subtle could be the player backpedaling while the partner runs in the same direction towards the player. \n    * Orientation: Player is perfectly aligned with and facing the back of the partner. \n* 90 (or 270) degrees:\n    * Direction: Player and partner are moving **perpendicularly** to each other. As an example, think about a player moving in a straight line and being approached by the partner from a right angle. \n    * Orientation: Player is facing forwards while the partner is facing one of the player's shoulders. \n* 180 degrees: \n    * Direction: Player and partner moving in **opposite** directions. This is probably the worst case scenario for a tackle, as the player and partner would be moving straight towards each other. \n    * Orientation: Player and partner are facing each other. \n  \n*One final note before visualizing the data - I took the time at which the player-partner distance was at a minimum as the most probable point of contact for each player-partner combination (thus leading to the concussion). There is probably a more rigorous way to identify when the contact occurred, but I figured this would be good enough to first order.*"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"b20a0c06b221335d079165b73b41fa72e50e5d7d","_kg_hide-output":true},"cell_type":"code","source":"## FUNCTIONS\ndef calculate_pp_distance(ngs_df):\n    \"\"\"\n    Given the NGS data for the injury set, calculate player-partner distance.\n\n    Parameters:\n        ngs_df: pd.DataFrame\n            DataFrame containing NGS data for player/partner on plays with a\n            concussion.\n    \"\"\"\n\n    # Split player/partner data.\n    play_df = ngs_df.loc[ngs_df.Identifier == 'PLAYER']\n    part_df = ngs_df.loc[ngs_df.Identifier == 'PARTNER']\n    play_df = play_df.drop(['Identifier'], axis=1)\n    part_df = part_df.drop(['Identifier'], axis=1)\n\n    # Rename columns for ease of merge.\n    ren_cols = ['GSISID', 'x', 'y', 'dis', 'o', 'dir', 'vx', 'vy', 's', 'ax',\n                'ay', 'a', 't']\n    play_cols = {x:f'play_{x}' for x in ren_cols}\n    part_cols = {x:f'part_{x}' for x in ren_cols}\n    play_df.rename(index=str, columns=play_cols, inplace=True)\n    part_df.rename(index=str, columns=part_cols, inplace=True)\n\n    # Perform merge.\n    mer_cols = ['Season_Year', 'GameKey', 'PlayID', 'Event',\n                'eventIndex', 'Time']\n    pp_df = play_df.merge(part_df, how='inner', left_on=mer_cols, right_on=mer_cols)\n\n    # Add extra columns that will assist with determining when tackle was made.\n    pp_df.loc[:, 'diff_x'] = pp_df.part_x - pp_df.play_x\n    pp_df.loc[:, 'diff_y'] = pp_df.part_y - pp_df.play_y\n    pp_df.loc[:, 'pp_dis'] = np.sqrt(np.square(pp_df.diff_x.values) + np.square(pp_df.diff_y.values))\n\n    return pp_df\n\ndef find_impact(pp_df):\n    \"\"\"\n    Given a DataFrame containing player-parter NGS data (for a single play),\n    identify the most likely time of impact and return all data from one second\n    around that time of impact.\n\n    Parameters:\n        pp_df: pd.DataFrame\n            NGS data for player/partner pair.\n    \"\"\"\n\n    # Figure out where the ball snap occurred and get the index so that we can\n    # discard all data prior to that instant.\n    try:\n        play_st = pp_df.loc[pp_df.Event == 'punt'].index[0]\n    except IndexError:\n        try:\n            play_st = pp_df.loc[pp_df.Event == 'ball_snap'].index[0]\n        except IndexError:\n            play_st = pp_df.index.min()\n\n    # Figure out where the play \"ended\" so that we can discard all data after\n    # that. For simplicity, we assume that any concussion event would have occured\n    # prior to a penalty flag being thrown or within five seconds of a tackle.\n    try:\n        play_ei = pp_df.loc[pp_df.Event == 'penalty_flag'].index[0]\n    except IndexError:\n        try:\n            play_ei = pp_df.loc[pp_df.Event == 'tackle'].index[0] + 50\n\n            play_ps = pp_df.loc[pp_df.Event == 'play_submit'].index[0]\n\n            while play_ei > play_ps:\n                play_ei -= 10\n\n        except IndexError:\n            play_ei = pp_df.index.max()\n\n    # Slice out the data that we actually need.\n    pp_df = pp_df.iloc[play_st:play_ei].reset_index(drop=True)\n\n    # Find the DataFrame index corresponding to the minimum player-partner\n    # distance.\n    min_dis_index = int(pp_df.loc[pp_df.pp_dis == pp_df.pp_dis.min()].index[0])\n\n    # Grab a few rows around that index.\n    min_dis_df = pp_df.iloc[(min_dis_index-5):(min_dis_index+5)]\n    min_dis_df.loc[min_dis_index,'impact'] = 1\n    min_dis_df.fillna(0, inplace=True)\n    min_dis_df.reset_index(drop=True, inplace=True)\n\n    return min_dis_df\n\ndef make_radial_plot(plot_df, angle_opt, title):\n    \"\"\"\n    Make a radial plot using a DataFrame built specifically for this purpose\n    (i.e., expected column names are hard-coded).\n\n    Parameters:\n        plot_df: pd.DataFrame\n            DataFrame containing data to be plotted.\n        angle_opt: str\n            Angle to plot ('o', 'dir', 'dir_tt').\n        title: str \n            Title to display over figure. \n    \"\"\"\n\n    if angle_opt == 'o':\n        plt_col = 'pp_o_diff'\n        color = 'orange'\n    elif angle_opt == 'dir':\n        plt_col = 'pp_dir_diff'\n        color = 'green'\n    elif angle_opt == 'dir_tt':\n        plt_col = 'pp_dir_diff'\n\n        impact_types = plot_df.Primary_Impact_Type.tolist()\n        imp_col_dict = {'Helmet-to-helmet':'red', 'Helmet-to-body':'blue'}\n        color = [imp_col_dict[x] for x in impact_types]\n\n        #pa_types = plot_df.Player_Activity_Derived.tolist()\n        #pa_col_dict = {\n        #    'Blocking': 'red', 'Blocked': 'orange', 'Tackling': 'green',\n        #    'Tackled': 'blue'\n        #}\n        #color = [pa_col_dict[x] for x in pa_types]\n    else:\n        raise ValueError('Not a valid option!')\n\n    angles = plot_df.loc[:,plt_col].values\n\n    def _fix_domain(ang):\n        if ang < 0:\n            return 360. - abs(ang)\n        else:\n            return ang\n\n    radii = plot_df.acc_rank.astype(float).tolist()\n    fix_angles = [_fix_domain(x) for x in angles]\n\n    data = [\n        go.Scatterpolar(\n            r = radii,\n            theta = fix_angles,\n            mode = 'markers',\n            marker = dict(\n                color = color\n            )\n        )\n    ]\n\n    layout = go.Layout(\n        title=title,\n        showlegend = False,\n        font=dict(\n            family='sans serif',\n            size=24,\n            color='black'\n        ),\n        polar = dict(\n            radialaxis = dict(\n                showticklabels = False,\n                showline=False,\n                ticklen=0\n            )\n        )\n    )\n\n    fig = go.Figure(data=data,layout=layout)\n\n    return fig","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"abd72ed34573f266131ce22d542166c4d82e7493"},"cell_type":"code","source":"# Load data.\nWDIR = '../input/ngs-dataset-playerpartnerinjuries/'\ninj_df = pd.read_csv(f'{WDIR}injury_ngs_data.csv')\n\n# Add column for easy indexing.\nmerge_cols = ['Season_Year', 'GameKey', 'PlayID']\nind_df = inj_df.drop_duplicates(merge_cols).reset_index(drop=True)\nind_df.loc[:, 'eventIndex'] = ind_df.index.values\n\nind_df = ind_df.loc[:, ['eventIndex', 'Season_Year', 'GameKey', 'PlayID']]\ninj_df = inj_df.merge(ind_df, how='outer', left_on=merge_cols, right_on=merge_cols)\n\n# Get player-partner processed DataFrame.\nplay_part_df = calculate_pp_distance(inj_df)\n\n# Step through each play, identifying the most likely point of impact and\n# grabbing a few rows around it.\nimpacts = []\n\nfor play_idx in range(len(inj_df.eventIndex.unique())):\n    try:\n        sp_df = play_part_df.loc[play_part_df.eventIndex == play_idx].reset_index(drop=True)\n\n        # Find most probable time for impact between player/partner.\n        impact_df = find_impact(sp_df)\n        impacts.append(impact_df)\n    except TypeError:\n        continue\n\npp_impact_df = pd.concat(impacts, ignore_index=True)\n\n# Add angle difference columns.\npp_impact_df.loc[:, 'pp_dir_diff'] = pp_impact_df.play_dir - pp_impact_df.part_dir\npp_impact_df.loc[:, 'pp_o_diff'] = pp_impact_df.play_o - pp_impact_df.part_o\n\n# Isolate most likely time step for impact.\njust_impact = pp_impact_df.loc[pp_impact_df.impact == 1]\n\n# Bring in summary stats for speed/acceleration (to be used on plot).\nWDIR = '../input/ngs-data-speedacceleration-summary-stats/'\nspd_acc_ss = pd.read_csv(f'{WDIR}spd_acc_summary.csv')\nspd_acc_ss.rename(index=str, columns={'season_year': 'Season_Year', 'game_key': 'GameKey', 'play_id': 'PlayID'}, inplace=True)\n\njust_impact = just_impact.merge(spd_acc_ss, how='inner', left_on=merge_cols,\n                                right_on=merge_cols)\n\n# Bring in impact types/player activity.\nWDIR = '../input/NFL-Punt-Analytics-Competition/'\nimpact_type = pd.read_csv(f'{WDIR}video_review.csv')\nimpact_type = impact_type.loc[:, ['Season_Year', 'GameKey', 'PlayID', 'Player_Activity_Derived', 'Primary_Impact_Type']]\n\njust_impact = just_impact.merge(impact_type, how='inner', left_on=merge_cols,\n                                right_on=merge_cols)\n\n# Pluck out the columns relevant for plotting.\n#keep_columns = ['max_play_a', 'pp_dir_diff', 'pp_o_diff']\n#plot_df = just_impact.loc[:, keep_columns].sort_values(by='max_play_a').reset_index(drop=True)\nplot_df = just_impact.sort_values(by='max_play_a').reset_index(drop=True)\nplot_df.loc[:, 'acc_rank'] = (plot_df.index.values+1)/(plot_df.index.max())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3643e8bcc004655503a134392ef0cf78c9dc21a7","_kg_hide-input":true},"cell_type":"code","source":"# Make orientation plot. \nplt_opt = 'o'\nfigure = make_radial_plot(plot_df, plt_opt, 'Player-Partner Orientation')\niplot(figure, filename='pp-orientation')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec627018dce3eda31f8b5439f79dd95f34647da5"},"cell_type":"markdown","source":"When looking at relative orientation, we really don't see any sort of clear trend in the data - the relative orientations are scattered fairly uniformly, and there's not a clear relation between relative orientation and maximum player acceleration (which is captured in the radial direction - I suppressed the axis label for clarity). If we think about how players move on punt plays, this really comes as no surprise - once the punt has been caught by the punt returner, the objective becomes either blocking for the PR or blocking players on the return, regardless of where they started or which way they are facing.  "},{"metadata":{"trusted":true,"_uuid":"7ad5ae87a454a6857d477d3803fa9ba01d3b9e17","_kg_hide-input":true},"cell_type":"code","source":"# Make direction plot. \nplt_opt = 'dir'\nfigure = make_radial_plot(plot_df, plt_opt, 'Player-Partner Direction')\niplot(figure, filename='pp-direction')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"676d73c8869981f62c5466f6907643df5c115496"},"cell_type":"markdown","source":"With relative direction, we do see a bit more of a tendency towards angles between +/- 90 degrees, which is also to be expected. Since much of the action between players of opposing teams occurs after the ball is in the hands of the returner, players of both teams will likely be heading in the same direction (towards the ball) to either block or try to make a play on the PR. Unfortunately, the few exceptions that fall closer to 180 degrees are the more gruesome plays in the dataset (e.g., the Darren Sproles concussion) and consequently, those are the ones where the concussed player experiences the largest acceleration (because the momentum exchange occurs along the same axis while the players are moving faster in opposite directions along that axis). "},{"metadata":{"trusted":true,"_uuid":"fe4a303247d4eab8105cece7d9997697520f3995","_kg_hide-input":true},"cell_type":"code","source":"# Make direction plot that utilizes tackle type. \nplt_opt = 'dir_tt'\nfigure = make_radial_plot(plot_df, plt_opt, 'Player-Partner Direction')\niplot(figure, filename='pp-direction-tackletype')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e8be501cfaf84cf791d93a0e53c813b4422e3f9f"},"cell_type":"markdown","source":"Out of curiosity, I also decided to consider relative orientation as a function of the type of contact (helmet-to-helmet: red, helmet-to-body: blue) that occurred between the players and partners on the plays. Although the helmet rule should go a long way at mitigating the risk of concussions from these types of impacts, it was interesting to see that players are often more likely to target the body of the opposing player (conciously or not) when they are moving primarily in antiparallel directions. However, without having ever played an organized game of tackle football in my life, this may just be how football players are taught to tackle early on in life.  "},{"metadata":{"_uuid":"5a1523d23060c465b39b42b3ddeeb05764cb587d"},"cell_type":"markdown","source":"On the whole, the angular information provided to us also did not provide a clear view of whether players were approaching tackles on punt plays from odd angles. With the exception of the clear, hard-to-watch plays (fast speeds, head-on collisions), players generally seem to be moving in a similar direction to their partners prior to impact and do not appear to orient themselves in any particular way before the tackle. "},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"### Game Conditions\n\nWith the NGS data really just serving to validate a couple of my gut assumptions about the nature of player dynamics on punt plays with concussions, I decided to do a little exploration into the game conditions as well to determine if I could glean any sort of actionable insights. Prior to looking into the data, I whittled down the list of possible game conditions to ones that seemed reasonable, and they were as follows:\n* Quarter: are concussions more likely to occur in some times of the game than others?\n* Score Difference: are concussions more likely to occur when a team is leading/losing by a certain score? \n* Punting Team Field Position (prior to punt): does where the punt is being kicked from have any influence on concussions? \n* Receiving Team Field Position (after catch): does the field position of the punt returner have any bearing on concussions? \n\nTo see if there was anything to the game conditions listed above, I chose to perform an (ad hoc) Kolmogorov-Smirnov test to determine if the distributions of quantities associated with each game condition varied significantly between all punt plays with some kind of play downfield (return, muffed punt, and downed punt) and those with concussions. This is, of course, a far cry from any sort of proper causal inference, but I figured that given the size of our dataset, this would be a quick and dirty way to get a zeroth-order approximation of the truth. "},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"10a0dab58fe0f551cad8dbba051ab5362665d67e"},"cell_type":"code","source":"## FUNCTIONS\ndef perform_ks_test(pop_df, sam_df, col_of_interest):\n    \"\"\"\n    Perform Kolmogorov-Smirnov test for population (all punts) and sample\n    (punts with concussions) distribution of provided quantity (col_of_interest).\n\n    Parameters:\n        pop_df: pd.DataFrame\n            DataFrame containing data from all selected punts (excluding the\n            ones in the concussion set).\n        sam_df: pd.DataFrame\n            DataFrame containing data from concussion set.\n        col_of_interest: str\n            Name of column/quantity for which you'd like to conduct the KS test.\n    \"\"\"\n\n    pdata = pop_df.loc[:, col_of_interest].values\n    sdata = sam_df.loc[:, col_of_interest].values\n\n    stat, p = stats.ks_2samp(pdata, sdata)\n\n    return (stat, p)\n\ndef plot_distribution(pop_df, sam_df, col_of_interest, plot_hp, title):\n    \"\"\"\n    Plot distribution of quantity (col_of_interest) for population/sample.\n\n    Parameters:\n        pop_df: pd.DataFrame\n            DataFrame containing data from all selected punts (excluding the\n            ones in the concussion set).\n        sam_df: pd.DataFrame\n            DataFrame containing data from concussion set.\n        col_of_interest: str\n            Name of column/quantity that you'd like to plot.\n        plot_hp: tuple (ints/floats)\n            Hyperparameter for plotting (bandwidth for KDE, number of bins for\n            histogram). Index 0 contains population hyperparameter, Index 1\n            contains sample hyperparameter. \n        title: str \n            Title to display over figure. \n    \"\"\"\n\n    pdata = pop_df.loc[:, col_of_interest].values\n    sdata = sam_df.loc[:, col_of_interest].values\n\n    # Uncomment for histograms.\n    pop_trace = go.Histogram(\n                    x=pdata,\n                    name='Population',\n                    opacity=0.75,\n                    marker=dict(color='red'),\n                    histnorm='probability',\n                    nbinsx=plot_hp[0]\n                )\n    sam_trace = go.Histogram(\n                    x=sdata,\n                    name='Sample (concussions)',\n                    opacity=0.75,\n                    marker=dict(color='blue'),\n                    histnorm='probability',\n                    nbinsx=plot_hp[1]\n                )\n\n    \"\"\"\n    # Uncomment for KDEs.\n    # Make KDE for each sample.\n    pop_kde = KernelDensity(kernel='gaussian', bandwidth=bw).fit(pdata[:, np.newaxis])\n    pop_plot = np.linspace(np.min(pdata), np.max(pdata), 1000)[:, np.newaxis]\n    plog_dens = pop_kde.score_samples(pop_plot)\n\n    sam_kde = KernelDensity(kernel='gaussian', bandwidth=bw).fit(sdata[:, np.newaxis])\n    sam_plot = np.linspace(np.min(sdata), np.max(sdata), 1000)[:, np.newaxis]\n    slog_dens = sam_kde.score_samples(sam_plot)\n\n    pop_trace = go.Scatter(\n                    x=pop_plot[:, 0],\n                    y=np.exp(plog_dens),\n                    mode='lines',\n                    fill='tozeroy',\n                    line=dict(color='red', width=2)\n                )\n\n    sam_trace = go.Scatter(\n                    x=sam_plot[:, 0],\n                    y=np.exp(slog_dens),\n                    mode='lines',\n                    fill='tozeroy',\n                    line=dict(color='blue', width=2)\n                )\n    \"\"\"\n\n    # Make figure.\n    fig = tools.make_subplots(rows=2,cols=1,shared_xaxes=True)\n    fig.append_trace(pop_trace, 1, 1)\n    fig.append_trace(sam_trace, 2, 1)\n    fig['layout'].update(\n        barmode='overlay', \n        title=title, \n        xaxis=dict(title='Receiving Team Points - Punting Team Points'),\n        yaxis1=dict(title='Normalized Density'),\n        yaxis2=dict(title='Normalized Density')\n    )\n\n    #data = [pop_trace, sam_trace]\n    #layout = go.Layout(barmode='overlay')\n    #fig = go.Figure(data=data, layout=layout)\n\n    return fig","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"dba1067894cbf92e596d60fcd166c56766baa1c0"},"cell_type":"code","source":"# Load data.\ndata_dict = load_data()\ndata_dict = collect_outcomes(data_dict)\ndata_dict = expand_play_description(data_dict)\n\n# Focus on play information.\nplay_info = data_dict['play_info']\noutcomes = ['return', 'downed', 'muffed punt']\nplay_info = play_info.loc[play_info.Punt_Outcome.isin(outcomes)].reset_index(drop=True)\nplay_info.loc[:, 'playIndex'] = play_info.index.values\n\n# Split out the plays on which we had an identified concussion.\ninj_df = data_dict['video_injury']\ninj_df.loc[:, 'concussionPlay'] = 1\ndrop_cols = ['Home_Team', 'Visit_Team', 'Qtr', 'PlayDescription', 'Week']\ninj_df.drop(drop_cols, axis=1, inplace=True)\ninj_df.rename(index=str, columns={'PlayId':'PlayID'}, inplace=True)\n\n# Join onto play_info.\nmer_cols = ['Season_Year', 'Season_Type', 'GameKey', 'PlayID']\n\ninj_play_info = play_info.merge(inj_df, how='inner', left_on=mer_cols,\n                                right_on=mer_cols)\n\n# Exclude plays from injury set from population set.\nplay_info = play_info.loc[~play_info.playIndex.isin(inj_play_info.playIndex.tolist())].reset_index(drop=True)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"cc7e04f7f42619fad4796f2510ddcbd748bea1b6"},"cell_type":"code","source":"cols_of_interest = ['Quarter', 'Score_Differential', 'Pre_Punt_RelativeYardLine', 'Post_Punt_RelativeYardLine']\nhp_dict = {'Quarter':(5,4), 'Score_Differential':(20,20), 'Pre_Punt_RelativeYardLine': (20,20),\n           'Post_Punt_RelativeYardLine': (20,20)}\n\n# Perform KS Test for quantities of interest. \nprint('Kolmogorov-Smirnov Test Results:')\n\nfor col in cols_of_interest:\n    # Drop some values.\n    if col == 'Post_Punt_RelativeYardLine':\n        pop_info = play_info.loc[play_info.Post_Punt_RelativeYardLine != -999]\n    else:\n        pop_info = play_info.copy()\n\n    ks_stat, pval = perform_ks_test(pop_info, inj_play_info, col)\n    print(f'quantity: {col}, test-statistic: {ks_stat}, p-value:{pval}')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec9babb16c71c0dbea48c1742f89489da1aabca5"},"cell_type":"markdown","source":"Woof. Depending on the strength of your stomach (i.e., level of statistical significance), we *might* be able to get away with saying that there's something more to score differential. However, it's clear that the other three (quarter, pre-punt yard line, post-punt yard line) aren't relevant here.\n\nTo see the nature of the distributions of score difference, let's plot them:"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"4352ba8459772089459bcb45449c6917963e74ab"},"cell_type":"code","source":"# Generate plot for score difference. \ncol = 'Score_Differential'\nfigure = plot_distribution(pop_info, inj_play_info, col, hp_dict[col], 'Score Difference')\niplot(figure, filename='score-difference-dist')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4b1e1fe534e7b27397cb863df3259b0735e01abf"},"cell_type":"markdown","source":"Qualms about the sample size aside, we can see that there is much more weight in the side that's less than zero (and none beyond the +14 margin), suggesting that teams/returners may be more willing to gamble on returning a punt if they need to dig their team out of a hole. However, if we look at the relative frequency of possible outcomes on all plays with \"action\" (returns, muffed punts, downed punts), we see:"},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"2182a3f27f987b3c184515a7e9a7557964b79051"},"cell_type":"code","source":"# Make heatmap for score difference using population set. \n# Split score differential into bins.\ndef _bin_score_differential(row):\n    row_sd = row.Score_Differential\n\n    if row_sd <= -21:\n        return '< -21'\n    elif (row_sd > -21) & (row_sd <= -14):\n        return '-20 to -14'\n    elif (row_sd > -14) & (row_sd <= -7):\n        return '-13 to -7'\n    elif (row_sd > -7) & (row_sd <= -1):\n        return '-6 to -1'\n    elif row_sd == 0:\n        return 'TIE'\n    elif (row_sd > 0) & (row_sd < 7):\n        return '+1 to +6'\n    elif (row_sd >= 7) & (row_sd < 14):\n        return '+7 to +13'\n    elif (row_sd >= 14) & (row_sd < 21):\n        return '+14 to +20'\n    elif row_sd >= 21:\n        return '> +21'\n    else:\n        return 'ERROR'\n\nplay_info.loc[:, 'binSD'] = play_info.apply(_bin_score_differential, axis=1)\n\n# Tally up outcomes as a function of score differential.\nreorder_indices = ['< -21', '-20 to -14', '-13 to -7', '-6 to -1', 'TIE',\n                   '+1 to +6', '+7 to +13', '+14 to +20', '> +21']\nbinned_outcomes = play_info.pivot_table(index='binSD', columns='Punt_Outcome',\n                                        aggfunc='size', fill_value=0)\ntest = binned_outcomes.copy()\nbinned_outcomes = binned_outcomes.div(binned_outcomes.sum(axis=1), axis=0)\nbinned_outcomes = binned_outcomes.reindex(reorder_indices)\n\n# Plot!\ntrace = go.Heatmap(\n            z=binned_outcomes.values.T,\n            x=binned_outcomes.index.tolist(),\n            y=binned_outcomes.columns.tolist())\ndata=[trace]\n\nlayout = go.Layout(\n            title='Score Difference (Punt Outcomes)',\n            xaxis=dict(title='Receiving Team Points - Punting Team Points')\n        )\n\nfigure = go.Figure(data=data, layout=layout)\n\niplot(data, filename='sd-heatmap')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d5746e134f65608a0081fe9373d3c45eadbf4e2b"},"cell_type":"markdown","source":"Gliding over the panes in the heatmap above, we see that there's only about a 4% increase in return probability (at the cost of downed punts) moving from TIE games to the games with the largest deficits (-21+ <), with the major caveat being that those situations are far less probable than close games. That said, a quick hypothesis test confirms that this is not a statistically significant difference. \n\nIn conclusion, although it would be exciting to entertain the idea of forcing teams with large leads to go for it on fourth down independent of their position on the field, there is not much of a case for it if we're looking to do it on the basis of reducing the likelihood of concussions. "},{"metadata":{"_uuid":"42a97d35c818de6528939fe3aae72d6e9b658bd4"},"cell_type":"markdown","source":"## Summary\nIn summary, we have seen the following from our analysis: \n* Players that move the fastest and/or are allowed to move about the field unimpeded are at the greatest risk of concussion. \n* There is a greater risk of concussion for players on the punting team, indicating that the increased risk is rooted in the return. \n* There were no significant trends in how players orient/angle themselves when defending punts that would lead to any actionable change in the way of how players are tackling on punts. \n* Outside of score difference (only marginally), there is no clear game condition that points to increased risk of concussion on punt plays.\n\nWith all that in mind, the changes that I have proposed above are intended to increase the likelihood of smaller impacts at the cost of players being able to move freely about the field uninhibited during punt plays. While I would like to do more to estimate the actual *impact* of my proposed changes (for instance, I'd love to see/hear about anyone in the world who has done anything related to simulating player dynamics, as I feel that would be necessary to truly characterize the effect that formation changes would have on punts prior to going live with them), it is my hope that the analysis that I've performed is sufficient to motivate the suggested changes. "},{"metadata":{"trusted":true,"_uuid":"0efc9cca109ee91babd2244075081f6cbceb2bef"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}