{"cells":[{"metadata":{"_uuid":"c687434390dbb7d91a5ef7e261b0f6b7aab0f9ad"},"cell_type":"markdown","source":"# **Data Exploration, Extraction, and Visualization using Pandas, RE, and Matplotlib**"},{"metadata":{"_uuid":"1ea542218dc0bc6d970397de930db6b4c7118cf2"},"cell_type":"markdown","source":"I wanted to try and make the biggest and most informational dataset using the NGS and play information provided in this competition. Using pandas, matplotlib, and some regular expressions, I attempted to extract as much data form the PlayDescription field as I could and join this new table with more data provided through NGS analysis tables. Hopefully it proves to be helpful. Enjoy."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# workspace prep \nimport pandas as pd \nimport matplotlib.pyplot as plt\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67aca42bf5fe90a2ba7e71cd73cdd5510b30ebba"},"cell_type":"markdown","source":"First, let's get our data in here."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# import data & look at data\nplay = pd.read_csv('../input/NFL-Punt-Analytics-Competition/play_information.csv')\nplayer_role =pd.read_csv(\"../input/NFL-Punt-Analytics-Competition/play_player_role_data.csv\") \nplayer = pd.read_csv('../input/NFL-Punt-Analytics-Competition/player_punt_data.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a8f6682d252da6e7850fd75a6bf020422aae4639"},"cell_type":"markdown","source":"Next, we need to check our data quality. Missing values and other data imperfections can prove to be quite the pain."},{"metadata":{"trusted":true,"_uuid":"b400d9caa8962f33084d9f8b553ba58bef92bd06"},"cell_type":"code","source":"# data quality check\ndfs = ([play, player_role, player])\n\nfor df in dfs:\n    \n    print(\"dataframe information\")\n    nan_count = df.apply(lambda x: x.count(), axis=0)\n    if sum(nan_count) == len(df)*len(df.columns):\n        print('No Missing Values')\n    elif nan_count != len(df):\n        print(nan_count)\n    \n    print(df.shape)\n    print(df.info())\n    print(df.head())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d693fce467af86d4926b246dbe213561e9b8a49c"},"cell_type":"markdown","source":"Great! Seems we were provided with some high quality data. Let's move on. Next, we will conduct some simple joins using the all powerful pandas package. Using the unique fields og GSISID, GameKey, and PlayID, we can create a very imformative data set from which we can gain some insights on the injuries during punt plays.  "},{"metadata":{"trusted":true,"_uuid":"67610e6add5dd356af9bec9f0d3cf071dd3ba153"},"cell_type":"code","source":"# join on proper keys - no missing data\n# first two on GSISID\n# then on GameKey and PlayID to get the full data set for each player    \nfull_players = player.merge(player_role, left_on='GSISID',right_on='GSISID',how = 'left')\nfull_set = full_players.merge(play, left_on=['GameKey','PlayID'],\n                              right_on = ['GameKey','PlayID'],\n                              how = 'left')\nprint(full_set.info())\nprint(full_set.isna().sum())\nfull_set.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ff2e0b7b5d2e4fda22208949c5370f1fbccb13ca"},"cell_type":"markdown","source":"Now we can drop the null values to ensure data quality."},{"metadata":{"trusted":true,"_uuid":"a22348387ef03f6c41c0a2190d09e64d75fe4c41"},"cell_type":"code","source":"#drop the null values\ndf=full_set.dropna()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"766cbbe48eeab0dfcbbb82eb9bad56ef928db559"},"cell_type":"markdown","source":"## Working with Play Description"},{"metadata":{"_uuid":"05982b0ed205b14e5011ea082ce40fee6923819d"},"cell_type":"markdown","source":"Since I wanted to work with the text in the certain columns, some coloumns need to become strings to do so."},{"metadata":{"trusted":true,"_uuid":"646a447e90a5c12e2003e35b250de1476c5b2103"},"cell_type":"code","source":"# need to split the score column and also the home and away column into 4 \n# diff columns \ndf['Home_Team_Visit_Team'] = df['Home_Team_Visit_Team'].astype(str)\ndf['Score_Home_Visiting'] = df['Score_Home_Visiting'].astype(str)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3ad9a3648b87c892d82d4cbeb36ee935f58d6a75"},"cell_type":"markdown","source":"We can now split the comlums we turned into strings into seperate variables using string split, and we can also quickly change the date format."},{"metadata":{"trusted":true,"_uuid":"d82079a0ccf62ee931b4cc92c61f3b8283893102"},"cell_type":"code","source":"# splits\ndf=df.join(df['Home_Team_Visit_Team'].str.split('-', 1, expand=True).rename(columns={0:'Home',1:'Away'}))\ndf=df.join(df['Score_Home_Visiting'].str.split(' - ', 1, expand=True).rename(columns={0:'Home_score',1:'Away_score'}))\n\n# Date\ndf[\"Game_Date\"] = pd.to_datetime(df[\"Game_Date\"], format = '%m/%d/%Y')\n\n# drop columns that were split\ndf = df.drop(['Home_Team_Visit_Team'], axis = 1)\ndf = df.drop(['Score_Home_Visiting'], axis = 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"de4479b8ef365c7737d1c1c4e86b0e76974d53d6"},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c690c850402219b91a3e0db4149dda01219faf69"},"cell_type":"markdown","source":"With those few lines of code, we now have more data than we did before that we can isolate and begin working with. Let's do the same for Play Description to get dummy variables and numerical variables for punt length, return length, fair catch, injury, penatly, a downed punt, fumbles, muffed punts, touchdowns, and touchbacks. This will allow for us to have the most amount of data possible for analysis later on. We can do this using a combination of for loops, regular expression, and basic string selection in Python. "},{"metadata":{"trusted":true,"_uuid":"0c531a480ad622eaf36b66544b66bda36d609999"},"cell_type":"code","source":"# Extract key information from the Play Description string variable using re\ndf['PlayDescription'] = df['PlayDescription'].astype(str)\n\n# punt length \nimport re \npunt_length = []\nfor row in df['PlayDescription']:\n    match = re.search('punts (\\d+)', row)\n    if match:\n        punt_length.append(match.group(1))\n    elif match is None:\n        punt_length.append(0)\n        \n# return length\nreturn_length = []\nfor row in df['PlayDescription']:\n    match = re.search('for (\\d+)', row)\n    if match:\n        return_length.append(match.group(1))\n    elif match is None:\n        return_length.append(0)\n            \n# fair catch\nfair_catch = []\nfor row in df['PlayDescription']:\n    match = re.search('fair catch', row)\n    if match:\n        fair_catch.append(1)\n    elif match is None:\n        fair_catch.append(0)\n\n# injury\ninjury = []\nfor row in df['PlayDescription']:\n    match = re.search('injured', row)\n    if match:\n        injury.append(1)\n    elif match is None:\n            injury.append(0)\n\n# penalty         \npenalty = []\nfor row in df['PlayDescription']:\n    if 'Penalty' in row.split():\n        penalty.append(1)\n    elif 'PENALTY' in row.split():\n        penalty.append(1)\n    elif 'Penalty' not in row.split():\n        penalty.append(0)\n    elif 'PENALTY' not in row.split():\n        penalty.append(0)\n        \n\n# downed\ndowned = []\nfor row in df['PlayDescription']:\n    match = re.search('downed', row)\n    if match:\n        downed.append(1)\n    elif match is None:\n        downed.append(0)\n\n# fumble\nfumble = []\nfor row in df['PlayDescription']:\n    match = re.search('FUMBLES', row)\n    if match:\n        fumble.append(1)\n    elif match is None:\n        fumble.append(0)\n\n# muff\nmuff = []\nfor row in df['PlayDescription']:\n    match = re.search('MUFFS', row)\n    if match:\n        muff.append(1)\n    elif match is None:\n        muff.append(0)\n\n# Touchback\ntouchback = []\nfor row in df['PlayDescription']:\n    match = re.search('Touchback', row)\n    if match:\n        touchback.append(1)\n    elif match is None:\n        touchback.append(0)\n\n# Touchdown\ntouchdown = []\nfor row in df['PlayDescription']:\n    match = re.search('TOUCHDOWN', row)\n    if match:\n        touchdown.append(1)\n    elif match is None:\n        touchdown.append(0)\n\n# add new columns to the df \ndf[\"punt_length\"] = punt_length\ndf[\"return_length\"] = return_length\ndf[\"fair_catch\"] = fair_catch\ndf[\"injury\"] = injury\ndf[\"penalty\"] = penalty\ndf[\"downed\"] = downed\ndf[\"fumble\"] = fumble\ndf['muff'] = muff\ndf['touchback'] = touchback\ndf['touchdown'] = touchdown","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8319db9fbe5f942520772a89075d9fde1d07e293"},"cell_type":"markdown","source":"Take a look at your new and imporved dataframe, ready to tell the full story about each player for every punt play provided in the data."},{"metadata":{"trusted":true,"_uuid":"1890ea7ef1b822307f74cf9e34e0b9f17a1798f2"},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"996ceac88b463899637de767c5fcfd1386b577ba"},"cell_type":"markdown","source":"I wanted to make this new dataset as robust and informational as possible so it seemed right to add the corresponding NGS data as well. Due to the sheer size of all the data, I used this amazing and helpful [kernal](http://www.kaggle.com/kmader/convert-to-feather-for-use-in-other-kernels/) by the great [Kevin Mader](http://www.kaggle.com/kmader). Using Apache Feather by the Pandas Father Wes McKinney, the NGS data file is cut nearly into a quarter of its original size, reducing its impact on the disk when read in by pandas, allowing for the kernal to survive the import. The NGS file contains all the NGS data as the serperate files share column names, making the concat process seamless."},{"metadata":{"_uuid":"8cefe8eccc83f31762c348c9fdfff25a9a7d9c2f"},"cell_type":"markdown","source":"## More Joining using the feathered data file"},{"metadata":{"trusted":true,"_uuid":"18cef6169c84ab76359340e81a85daf6fc82488d"},"cell_type":"code","source":"import feather\ndf_final = feather.read_dataframe('../input/feathered-ngs/ngs.feather')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"12a1058ceba29c06e22a349744b2952a74e54e6c"},"cell_type":"markdown","source":"Let's take a peak at the data we just loaded in."},{"metadata":{"trusted":true,"_uuid":"ce0b46737578c92f36e3a0cc321b7b046e53412b"},"cell_type":"code","source":"print(df_final.shape)\ndf_final.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bfb5b99d57c7fb96285ea5dea898784339bf8c88"},"cell_type":"markdown","source":"Lots of data there but we will soon trim it down. Let's do some joining with pandas again."},{"metadata":{"trusted":true,"_uuid":"56ac1f6d78ff84343d2534be64d1484ba3e7a03b"},"cell_type":"code","source":"new_df = df.merge(df_final.drop_duplicates(subset=['GSISID','GameKey','PlayID']), how='left',\n                  left_on=['GSISID','GameKey','PlayID','Season_Year_x'], right_on = ['GSISID','GameKey','PlayID','Season_Year'])\ndel df_final","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5ce9ddf3812cb50f008e3e8bf2692f4b2e0e3d19"},"cell_type":"code","source":"new_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"afb83ec9c35a3f4fe5c9fcf2fc1f693131bd5de9"},"cell_type":"markdown","source":"We are beginning to see the power of proper joining, giving us a very insightful data set. We are not done yet though. Let do some final touches and trim down the columns."},{"metadata":{"trusted":true,"_uuid":"937586cbd3280af08553c124b953946713d059fa"},"cell_type":"code","source":"# game data\ngame = pd.read_csv('../input/NFL-Punt-Analytics-Competition/game_data.csv')\n\ngame_with_new = new_df.merge(game, how = 'left', left_on = \"GameKey\", \n                             right_on = \"GameKey\")\n# columns to keep \nkeep = ['GSISID', 'Number', 'Position','Season_Year_x', 'GameKey', 'PlayID',\n       'Role', 'Game_Date_x', 'Week_x',\n       'Game_Clock', 'YardLine', 'Quarter', 'Play_Type', 'Poss_Team',\n       'Home', 'Away', 'Home_score', 'Away_score',\n       'punt_length', 'return_length', 'fair_catch', 'injury', 'penalty',\n       'downed', 'fumble', 'muff', 'touchback','touchdown','x',\n       'y', 'dis', 'o', 'dir', 'Event', 'Season_Type_y',\n       'Game_Day', 'Game_Site', 'Start_Time',\n       'Home_Team', 'Visit_Team', 'Stadium',\n       'StadiumType', 'Turf', 'GameWeather', 'Temperature', 'OutdoorWeather'\n       ]\ndf_clean = game_with_new[keep]\ndel game_with_new\n\n# rename columns\nheaders = ['GSISID', 'Number', 'Position', 'Season_Year','Season_Year_x', 'GameKey', 'PlayID',\n       'Role', 'Game_Date', 'Week',\n       'Game_Clock', 'YardLine', 'Quarter', 'Play_Type', 'Poss_Team',\n       'Home', 'Away', 'Home_score', 'Away_score',\n       'punt_length', 'return_length', 'fair_catch', 'injury', 'penalty',\n       'downed', 'fumble', 'muff', 'touchback','touchdown','x',\n       'y', 'dis', 'o', 'dir', 'Event', 'Season_Type',\n       'Game_Day', 'Game_Site', 'Start_Time',\n       'Home_Team', 'Visit_Team', 'Stadium',\n       'StadiumType', 'Turf', 'GameWeather', 'Temperature', 'OutdoorWeather'\n       ]\ndf_clean.columns = headers","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0b57c010f6aa439629504630efe978c35e544fa8"},"cell_type":"code","source":"print(df_clean.dtypes)\ndf_clean[[\"punt_length\", \"return_length\"]] = df_clean[[\"punt_length\", \"return_length\"]].apply(pd.to_numeric)\ndf_clean = df_clean.drop(columns='Season_Year_x')\ndf_clean.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5e7e8e3223dc312f7432b5e4d1ad89ddabe2cb1c"},"cell_type":"markdown","source":"We now have a data set that gives us in depth looks at every play a player was out for a punt for the last two years. We can now begin to subset this data based on injuries and begin to gain some insights to injuries on punt plays in the NFL."},{"metadata":{"_uuid":"d4472c05886c68ba620b1b5b508377b5e670e880"},"cell_type":"markdown","source":"## Injury Exploration"},{"metadata":{"trusted":true,"_uuid":"c0a535531bd78c9bb1615d8891449d81dc3f3074"},"cell_type":"code","source":"# lets look at how many games and punts there are \ngames = len(df_clean['GameKey'].unique().tolist())\nprint('There are ' + str(games) + ' games in the dataset.')\npunts = len(df_clean['PlayID'].unique().tolist())\nprint('There are ' + str(punts) + ' punts in the dataset.')\nprint('On average, there are ' + str(punts/games) + ' punts per game.')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9c80e682bcd2ad78f662a65d38e537ba43c8adf2"},"cell_type":"code","source":"# let's start with the injury field\nno_injuries = df_clean.loc[df_clean['injury'] == 0]\ninjuries = df_clean.loc[df_clean['injury'] == 1]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b4149f44ac8ddf244190e3d75a8c3e489487efd"},"cell_type":"markdown","source":"We now have two differnt data set that contain injury plays and none injury plays and the data for each player on the field. We will need to be careful when doing aggregation to use unique vales when appropriate. We can need to quickly define an avergae function for lists since .nunique() will be returning lists."},{"metadata":{"trusted":true,"_uuid":"fa4014ebd454f4cd8f7b774e6ed62b08ccab8eda"},"cell_type":"code","source":"# average function \ndef avg(lst):\n    return sum(lst)/len(lst)\n\n# Number of injuries\nprint('There are ' + str(len(injuries['PlayID'].unique().tolist())) + ' injuries in the dataset.')\n\n# lets look at the average punt length and return lenth for both new dfs\nprint('The average punt length for a play with an injury is ' + str(avg(injuries['punt_length'].unique().tolist())))\nprint('The average punt length for a play without an injury is ' + str(avg(no_injuries['punt_length'].unique().tolist())))\nprint('The average punt return for a play with an injury is ' + str(avg(injuries['return_length'].unique().tolist())))\nprint('The average punt return for a play without an injury is ' + str(avg(no_injuries['return_length'].unique().tolist())))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d0d7fea62854a91018ba65113dc555e8b8d02ded"},"cell_type":"markdown","source":"The difference between the average punt return lengths for plays with and without injuries is jarring. This is likely due to the returner being stopped abruptly by the coverage team with a bone-crushing tackle."},{"metadata":{"trusted":true,"_uuid":"aa9318503906edba3f7085268f9046bafc5647f1"},"cell_type":"code","source":"#injuries by gameday\ntotal_injuries = injuries.groupby('Game_Day')['PlayID'].nunique()\ntotal_no_injuries = no_injuries.groupby('Game_Day')['PlayID'].nunique()\nprint('On Fridays, injuires occured on ' + str(3/203) + ' percent of punt plays.')\nprint('On Mondays, injuires occured on ' + str(4/312) + ' percent of punt plays.')\nprint('On Saturdays, injuires occured on ' + str(7/602) + ' percent of punt plays.')\nprint('On Sundays, injuires occured on ' + str(56/2648) + ' percent of punt plays.')\nprint('On Thursdays, injuires occured on ' + str(16/871) + ' percent of punt plays.')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0d0e0e618ea0953aa8d5e05e8f57e29f305c9767"},"cell_type":"markdown","source":"While Sunday is the most represented day in the data set, it still seems the odds of being hurt on a punt play on Sunday are still the highest. Keep your head on a swivel!"},{"metadata":{"_uuid":"5ca0f0f54d3b765983c2aec94c573d4817db280a"},"cell_type":"markdown","source":"Let's now do some quick but fun bar charts and histograms to wrap up."},{"metadata":{"trusted":true,"_uuid":"d00543f3f5e03e451388371cec5e9f66a208009c"},"cell_type":"code","source":"# injuries by game site\ninjuries.groupby('Game_Site')['PlayID'].nunique().plot(kind='bar',figsize=(18, 16))\nplt.xlabel('Week')\nplt.ylabel('Injuries')\nplt.title('Injuries per Location')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bf700763f24ba8c724b5c789956cc726a0036e4c"},"cell_type":"markdown","source":"Miami Garden seems to be a pretty rough place to play or the Dolphins are getting it doen on special teams. This probably deserves more analysis."},{"metadata":{"trusted":true,"_uuid":"c14dd5cb003c38c4cdcb741437536c96495ee9e6"},"cell_type":"code","source":"# injuries by season year\ninjuries.groupby('Season_Year')['PlayID'].nunique().plot(kind='bar',figsize=(12, 10))\nplt.xlabel('Year')\nplt.ylabel('Injuries')\nplt.title('Injuries (2016-2017)')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0c6c646523bd9f3f6db16221626865503c534d41"},"cell_type":"markdown","source":"Even with increased awareness of player safety, there were more injuries in 2017 than 2016."},{"metadata":{"trusted":true,"_uuid":"c86c81f89a73f0dcb496ace8ec745f594277f11a"},"cell_type":"code","source":"# injuries by muff\ndata = injuries.groupby('muff')['PlayID'].nunique().plot(kind='bar', figsize=(12, 10))\nplt.xlabel('muff')\nplt.ylabel('Injuries')\nplt.title('Injuries on Muffs')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7681305d9116e74685323b5c84842b5a5e8bf23f"},"cell_type":"markdown","source":"Muffs do not seem to be a cause of injury."},{"metadata":{"trusted":true,"_uuid":"34461e9ff58614b63019cb5e936ec6ff20556775"},"cell_type":"code","source":"# injuries by fumble\ndata = injuries.groupby('fumble')['PlayID'].nunique().plot(kind='bar', figsize=(12, 10))\nplt.xlabel('Fumble')\nplt.ylabel('Injuries')\nplt.title('Injuries on Fumbles')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"545db1bf7d2b00df2d083bbbf8cea3cbcd203f22"},"cell_type":"markdown","source":"On fumble plays, the focus shifts from the player to the ball on the ground, therefore decreasing injuries dramatically."},{"metadata":{"trusted":true,"_uuid":"42bb017b5f5af93a6c8d6db56d9c3215ffcead89"},"cell_type":"code","source":"# injuries by touchdown\ndata = injuries.groupby('touchdown')['PlayID'].nunique().plot(kind='bar',figsize=(12, 10))\nplt.xlabel('Touchdowns')\nplt.ylabel('Injuries')\nplt.title('Injuries on Touchdowns')\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"18c8e877917a29b2e1cfebc66310255c58229c1e"},"cell_type":"markdown","source":"6 points don't seem to be a cause for injury."},{"metadata":{"trusted":true,"_uuid":"6ce4fcfb0b1c90e7486563b6d7fff12e426a5465"},"cell_type":"code","source":"# injuries by week\ndata = injuries.groupby('Week')['PlayID'].nunique().plot(kind='bar',figsize=(18, 16))\nplt.xlabel('Week')\nplt.ylabel('Injuries')\nplt.title('Injuries per Week')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e8f2511bef05e9995c838326eab976ded4b03310"},"cell_type":"markdown","source":"Week 5 has more injuries than any other week during both seasons. This deserves further analysis."},{"metadata":{"trusted":true,"_uuid":"f8262a44ed0053e43482908e11b74671ba6f2e99"},"cell_type":"code","source":"# injuries by quarter\ninjuries.groupby('Quarter')['PlayID'].nunique().plot(kind='bar',figsize=(12, 10))\nplt.xlabel('Week')\nplt.ylabel('Injuries')\nplt.title('Injuries per Quarter')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f6d180d9cfcee1c6bea78d59dbc73d7aa004efa"},"cell_type":"markdown","source":"As the game carries on, injuries begin to happen more. This can be attributed to several different factors. As the game wears on player's become tired, fatigued and often crucial plays are made by these world class athletes on coverage teams resulting big hits."},{"metadata":{"trusted":true,"_uuid":"fd957d8e04ac332d65d4bea7dc8704c025c16f64"},"cell_type":"code","source":"# injuries per season type\ninjuries.groupby('Season_Type')['PlayID'].nunique().plot(kind='bar', figsize=(12, 10))\nplt.xlabel('Season Type')\nplt.ylabel('Injuries')\nplt.title('Injuries in Pre, Post, and Regular Season Games')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b7992f345953d2b0b777a7b4335affbea5eaecf0"},"cell_type":"markdown","source":"There are more regular season games through the year, therefore seeing most injuires occuring in the regular season makes sense. That pregame number looks high and probably deserves further analysis."},{"metadata":{"trusted":true,"_uuid":"bcf24d9e0e9a5d819b5b115c2a21545d8bbac416"},"cell_type":"code","source":"# lets look at teams who have the most injuries\nfig, axes = plt.subplots(nrows=1, ncols=2, sharey = True)\n\ninjuries.groupby('Home')['PlayID'].nunique().plot(figsize=(18, 16),ax=axes[0],kind='bar')\nplt.ylabel('Injuries')\nplt.suptitle('Frequency of Injuries by Home (Left) and Away (Right)')\ninjuries.groupby('Away')['PlayID'].nunique().plot(figsize=(18, 16),ax=axes[1],kind='bar')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9a5db5eebfaa79d196f3c50d63c9c757f8076edf"},"cell_type":"markdown","source":"As we saw earlier Miami Gardens has a high number of injuries and we now see that most of those injuries are suffered by the home team, the Dolphins :/"},{"metadata":{"trusted":true,"_uuid":"e4b35ba1dc3b07f135c395be4b0a5e777f1efc66"},"cell_type":"code","source":"# lets look at punt length\ncols = ['GameKey', 'PlayID','punt_length','injury']\npunt_length = df_clean[cols]\npunt_length = punt_length.drop_duplicates()\n\n# histogram for punt length on injuires\nfig, axes = plt.subplots(nrows=1, ncols=2)\n\npunt_length['punt_length'].loc[punt_length['injury']==1].plot(ax=axes[0],kind='hist', bins = 10, color = 'red', edgecolor = 'black', figsize=(18, 16))\npunt_length['punt_length'].loc[punt_length['injury']==0].plot(ax=axes[1],kind='hist', bins = 10, edgecolor = 'black', figsize=(18, 16))\nplt.suptitle('Frequency of Injuries (Red) and Non-Injuries (Blue) by Return length')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d9c9f83de5d4274a0a19f2afc8f2767d1f8ee143"},"cell_type":"markdown","source":"We see that most of the innjuries happen right around the 50-55 yard range. This is likely due to the fact that punts of this length give sufficent time for theplayers on the coverage team to cover the play properly, and deliver blows to the punt returner almost immediately. Punts of this length also allow for the returner to sometimes return the ball if blocked correctly, leading to blindside hits and more opportunites for coverage players to deliver blows as well."},{"metadata":{"trusted":true,"_uuid":"a5ca62da2ff452bdd2373c0e098b9fc9fe6e1507"},"cell_type":"code","source":"# same process for return length\ncols = ['GameKey', 'PlayID','return_length','injury']\nreturn_length= df_clean[cols]\nreturn_length= return_length.drop_duplicates()\n\n# histogram for return length on injuires\nfig, axes = plt.subplots(nrows=1, ncols=2)\n\nreturn_length['return_length'].loc[return_length['injury']==1].plot(ax=axes[0],kind='hist', bins = 15, color = 'red', edgecolor='black',figsize=(18, 16))\nreturn_length['return_length'].loc[return_length['injury']==0].plot(ax=axes[1],kind='hist', bins = 15, edgecolor='black',figsize=(18, 16))\nplt.suptitle('Frequency of Injuries (Red) and Non-Injuries (Blue) by Return length')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8191a9355d4476679d5d77c1e1709451fe8db92b"},"cell_type":"markdown","source":"Concurrent with our prior analysis of punt length, it seems that most of the injuries take place almost immedately after the return process begins. Let's double check with a quick bar graph of fair catches just to be certain."},{"metadata":{"trusted":true,"_uuid":"e1cfaf156c3ae03ae3a68d281e06d5b658cdcb81"},"cell_type":"code","source":"# injuries by fair catch\ninjuries.groupby('fair_catch')['PlayID'].nunique().plot(kind='bar', figsize=(12, 10))\nplt.xlabel('Fair Catch')\nplt.ylabel('Injuries')\nplt.title('Injuries on Fair Catches')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0201f8dd6445397de7685a01cec5644e82a268f4"},"cell_type":"markdown","source":"It seems we were corrent, most injures are happening directly after the returner begins the return process."},{"metadata":{"_uuid":"a7a7c4999b6bcdc92886f2eff1ebf02380bb07e7"},"cell_type":"markdown","source":"## Conclusion "},{"metadata":{"_uuid":"107bd83c611dad4edec2ec87bb228a3341054b66"},"cell_type":"markdown","source":"Further analysis is needed but it seems that most of the injuries that are taking place on punt plays are happening during the regular season, late in the game, on punts of rougly 50-55 yards, immedately after the returner catches the ball and starts the return process. I will post another kernal if time permits of my analysis of the video data provided as well. I hope you enjoyed this analysis."},{"metadata":{"_uuid":"45561e24af7ac63c3c66eb2c1fe6b827bb05ce00"},"cell_type":"markdown","source":"## Cheers!"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}