{"cells":[{"metadata":{"_uuid":"26626034-8e0d-4d8f-985f-9e4cf10e2c53","_cell_guid":"0339a3a3-a165-4d70-9141-b50367197346","trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \nimport timeit          #library to check the time to run a code\nimport json            #library to read the json formats\nimport pandas as pd    #library to execute Dataframe\nimport numpy as np     #library to execute Numerical/Statistical Calculations\nfrom pandas.io.json import json_normalize # library to normalize other JSON formats\nimport numpy as np\nfrom functools import reduce\nfrom math import radians, cos, sin, asin, sqrt\nfrom datetime import datetime \nimport datetime as dt\nfrom collections import Counter\nimport gc\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Data Visualization Tools\nimport seaborn as sns\nfrom matplotlib import pyplot\nimport matplotlib.pyplot as plt\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.offline as offline\noffline.init_notebook_mode()\nfrom plotly import tools\nimport plotly.tools as tls\nimport plotly.express as px\nimport plotly.figure_factory as ff\n\n\n#Libraries for Modeling\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.metrics import accuracy_score, confusion_matrix\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn import svm, tree\nimport xgboost\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.tree import DecisionTreeRegressor\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Loading_Raw_Data:"},{"metadata":{"trusted":true},"cell_type":"code","source":"# p_list_df_raw = pd.read_csv('../input/nfl-playing-surface-analytics/PlayList.csv')\n# p_trk_df_raw = pd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv')\ninjury_df_raw = pd.read_csv('../input/nfl-playing-surface-analytics/InjuryRecord.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Exploratory Analysis\n\n## EDA for Player Injury Data Frame:"},{"metadata":{},"cell_type":"markdown","source":"Let's Check the Shape of the file first!"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# injury_df = injury_df_raw.copy()\ninjury_df_raw.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As per problem Statement it is said that a total of** 100** unique players' injury data is provided. However, from the injury data-set it can be seen there are 105 records. \n\nAre there players with multiple injury entry?\nLet's Check!\n\n#### Players with multiple injury"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"## Checking Repeated injuries in same players in the tournament\ninjury_df = injury_df_raw\nRepeat_injury = injury_df[injury_df['PlayerKey'].duplicated()]\nprint(Repeat_injury['PlayerKey'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The above are 5 players with multiple entries. Lets see their injury history."},{"metadata":{"trusted":true},"cell_type":"code","source":"Repeat_injury = injury_df.loc[injury_df['PlayerKey'].isin([43540,45950,44449,33337,47307])].sort_values(by=['PlayerKey'])\nRepeat_injury[['PlayerKey','GameID','BodyPart','Surface']]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the above We can see that these players have mulitple injury records over the 2 seasons. \n\n\nNow let's Calculate the Imapct of the injury , Given that the recovery time is mentioned. \n\nCategorizing the Injury Impact :\n1+ : Low(1), \n7+ : Moderate(2), \n28+ : High(3) , \n42+: Extreme(4)"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Categorizing the Injury Impact [1+ : Low(1) , 7+ : Moderate(2) , 28+ : High(3) , 42+: Extreme(4)]\ninjury_df = injury_df_raw\ninjury_df['Injury_Impact']  = 0\n\nfor i in range(len(injury_df)):\n    if (injury_df['DM_M42'][i] == 1):\n        injury_df['Injury_Impact'][i] = 4\n        \n    elif (injury_df['DM_M28'][i] == 1) and (injury_df['DM_M42'][i] == 0):\n        injury_df['Injury_Impact'][i] = 3\n        \n    elif (injury_df['DM_M7'][i] == 1) and (injury_df['DM_M28'][i] == 0):\n        injury_df['Injury_Impact'][i] = 2\n    \n    else:\n        injury_df['Injury_Impact'][i] = 1\n        \ninjury_df_Final = injury_df.drop(columns=['DM_M1', 'DM_M7', 'DM_M28', 'DM_M42'])\n\n\ninjury_df_Final['Injured'] = 1 #also adding the tag as injured players\ninjury_df_Final.tail()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Histogram for Injury Impact Distribution and Injury Type"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"for i in range(len(injury_df)):\n    if (injury_df['DM_M42'][i] == 1):\n        injury_df['Injury_Impact'][i] = '6+_Weeks'\n        \n    elif (injury_df['DM_M28'][i] == 1) and (injury_df['DM_M42'][i] == 0):\n        injury_df['Injury_Impact'][i] = '4+_Weeks'\n        \n    elif (injury_df['DM_M7'][i] == 1) and (injury_df['DM_M28'][i] == 0):\n        injury_df['Injury_Impact'][i] = '1+_Week'\n    \n    else:\n        injury_df['Injury_Impact'][i] = '1+_Day'\n        \ninjury_df_Final = injury_df.drop(columns=['DM_M1', 'DM_M7', 'DM_M28', 'DM_M42'])\n\n\n############################################################################################################################\n#                                        Injury Impact DISTRIBUTION\n############################################################################################################################\n\n\nimport plotly.express as px\n\ndata = px.histogram(injury_df_Final, x=\"Injury_Impact\", title='<b>Histogram of Injury Impact: Based on recovery Days</b>',\n                      opacity=0.7, color_discrete_sequence=['indianred'])\n\nlayout = go.Layout(\n)\n\nfig = go.Figure(data=data, layout=layout)\n\nfig.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Recovery Days (Injury Impact)</b>\",\n    yaxis_title=\"<b>Count of Players</b>\",\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='LightPink')\nfig.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n############################################################################################################################\n#                                        Injury Type DISTRIBUTION\n############################################################################################################################\n\nimport plotly.express as px\n\ndata = px.histogram(injury_df_Final, x=\"BodyPart\", histnorm='percent', title='<b>Histogram of Body Part Distribution</b>',\n                      opacity=0.7, color_discrete_sequence=['darkkhaki'])\nlayout = go.Layout()\nfig0 = go.Figure(data=data, layout=layout)\n\n\nfig0.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig0.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Injured Body part</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig0.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig0.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig0.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='khaki')\nfig0.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\nfig.show()\nfig0.show()\n\ndel fig\ndel fig0\ndel data\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"injury_df = injury_df_raw.copy()\ninjury_df['Injury_Impact']  = 0\n\nfor i in range(len(injury_df)):\n    if (injury_df['DM_M42'][i] == 1):\n        injury_df['Injury_Impact'][i] = 4\n        \n    elif (injury_df['DM_M28'][i] == 1) and (injury_df['DM_M42'][i] == 0):\n        injury_df['Injury_Impact'][i] = 3\n        \n    elif (injury_df['DM_M7'][i] == 1) and (injury_df['DM_M28'][i] == 0):\n        injury_df['Injury_Impact'][i] = 2\n    \n    else:\n        injury_df['Injury_Impact'][i] = 1\n        \ninjury_df_Final = injury_df.drop(columns=['DM_M1', 'DM_M7', 'DM_M28', 'DM_M42'])\n\n\ninjury_df_Final['Injured'] = 1 #also adding the tag as injured players\n\n\n# injury_df_Final.tail()\ndel injury_df_raw\ndel injury_df\ndel Repeat_injury\n#collect residual garbage\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## EDA for Play List Data Frame:"},{"metadata":{},"cell_type":"markdown","source":"Feature Size Reduction on :\n1. Weather Type : Bucketed all the weather types into 6 major category\n2. Stadium type : Bucketed into 2 major category\n3. Roster position : Bucketed in 9 Position Category\n\nFeature Created\n1. Roster position Retained : This feature is extracted by comparing Roster position and Player Position feature. This is extracted to know, if the player is currely in its designated position or not during the play. "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"p_list_df_raw = pd.read_csv('../input/nfl-playing-surface-analytics/PlayList.csv')\np_list_df = p_list_df_raw\n#delete when no longer needed\ndel p_list_df_raw\n#collect residual garbage\ngc.collect()\n\n#########################################################################################################################\n#                                              Weather Tagging\n#########################################################################################################################\n\nCloudy = ['Cloudy','Hazy','Cloudy and Cool', 'Overcast', 'Rain Chance 40%', 'cloudy', 'Cloudy, fog started developing in 2nd quarter', \n          'Cloudy with periods of rain, thunder possible. Winds shifting to WNW',\n          'Mostly Cloudy', 'Mostly cloudy', 'Partly Sunny', 'Partly sunny','Mostly Coudy', 'Mostly_Cloudy',\n          'Partly Cloudy','Partly cloudy', 'Party Cloudy','Partly Clouidy', 'Partly_Cloudy']  \n\nSunny = ['Sunny', 'Sunny and warm', 'Mostly Sunny', 'Sunny and clear', 'Sunny Skies','Heat Index 95', 'Sunny, highs to upper 80s', 'Sun & clouds', 'Mostly sunny', 'Sunny, Windy', 'Mostly Sunny Skies']\nClear = ['Clear and warm', 'Fair', 'Clear', 'Clear and Cool', 'Clear Skies', 'Clear skies', 'Partly clear', '10% Chance of Rain', 'Clear and sunny','Clear to Partly Cloudy']          \nRain = ['Rain', 'Showers', 'Scattered Showers', 'Light Rain', 'Cloudy, Rain', 'Rainy','30% Chance of Rain', 'Rain shower']          \nIndoor = ['Controlled Climate', 'Indoor', 'Indoors', 'N/A (Indoors)','N/A Indoor']\nSnow_Cold = ['Snow', 'Heavy lake effect snow', 'Cloudy, light snow accumulating 1-3\"', 'Cloudy and cold',\n        'Clear and cold', 'Sunny and cold', 'Cold', 'Rain likely, temps in low 40s.']\n    \np_list_df['Weather'] = p_list_df['Weather'].astype(str)   \n          \ndef assign_feature_problem(data):\n    \n    if any(word in data for word in Cloudy):\n          data = \"Cloudy\"\n          return data\n    elif any(word in data for word in Sunny):\n          data = \"Sunny\"\n          return data\n    elif any(word in data for word in Clear):\n          data = \"Clear\"\n          return data\n    elif any(word in data for word in Rain):\n          data = \"Rain\"\n          return data\n    elif any(word in data for word in Indoor):\n          data = \"Indoor\"\n          return data\n    elif any(word in data for word in Snow_Cold):\n          data = \"Snow_Cold\"\n          return data\n \np_list_df['Weather'] = p_list_df['Weather'].apply(lambda x : assign_feature_problem(x))\n\n\n#########################################################################################################################\n#                                              Stadium Type Tagging\n#########################################################################################################################\n\nOutdoor = ['Outdoor', 'Outdoors', 'Open', 'Domed, open', 'Oudoor', 'Domed, Open', 'Ourdoor', 'Outdoor Retr Roof-Open', 'Outddors',\n           'Retr. Roof-Open', 'Retr. Roof - Open', 'Indoor, Open Roof', 'Outdor', 'Outside', 'Cloudy', 'Heinz Field',\n           'Retractable Roof']\nIndoor = ['Indoors', 'Dome', 'Indoor', 'Domed, closed', 'Dome, closed', 'Closed Dome', 'Domed', 'Indoor, Roof Closed', \n          'Retr. Roof Closed', 'Retr. Roof - Closed', 'Retr. Roof-Closed']\n\n\np_list_df['StadiumType'] = p_list_df['StadiumType'].astype(str)   \n          \ndef assign_feature_problem(data):\n    \n    if any(word in data for word in Outdoor):\n          data = \"Outdoor\"\n          return data\n    elif any(word in data for word in Indoor):\n          data = \"Indoor\"\n          return data\n \np_list_df['StadiumType'] = p_list_df['StadiumType'].apply(lambda x : assign_feature_problem(x))\n\n# Hence adjusting the Weather Column as well\np_list_df['Weather'] = np.where(p_list_df['StadiumType']=='Indoor', 'Indoor', p_list_df['Weather'])\n    \n#########################################################################################################################\n#                                              Roster Position Tagging\n#########################################################################################################################\n\nQB = 'Quarterback'\nWR = 'Wide Receiver'\nLB = 'Linebacker'\nRB = 'Running Back'\nDL = 'Defensive Lineman'\nTE = 'Tight End'\nDB = ['Safety', 'Cornerback']\nOL = 'Offensive Lineman'\nSPEC = 'Kicker' \n\n\np_list_df['RosterPosition'] = p_list_df['RosterPosition'].astype(str)   \n          \ndef assign_feature_problem(data):\n    \n    if any(word in data for word in QB):\n          data = \"QB\"\n          return data\n    elif any(word in data for word in WR):\n          data = \"WR\"\n          return data\n    elif any(word in data for word in LB):\n          data = \"LB\"\n          return data\n    elif any(word in data for word in RB):\n          data = \"RB\"\n          return data\n    elif any(word in data for word in DL):\n          data = \"DL\"\n          return data\n    elif any(word in data for word in TE):\n          data = \"TE\"\n          return data\n    elif any(word in data for word in DB):\n          data = \"DB\"\n          return data\n    elif any(word in data for word in OL):\n          data = \"OL\"\n          return data\n    elif any(word in data for word in SPEC):\n          data = \"SPEC\"\n          return data  \n \np_list_df['RosterPosition'] = p_list_df['RosterPosition'].apply(lambda x : assign_feature_problem(x))\n\np_list_df.tail()\n\n#########################################################################################################################\n#                                              Roster_Position_Retained Tagging\n#########################################################################################################################\n\np_list_df['PositionRetained'] = 'N'\n\n\nfor i in range(len(p_list_df)):   \n    if p_list_df['RosterPosition'][i] == p_list_df['PositionGroup'][i]:\n        p_list_df['PositionRetained'][i] = 'Y'\n    else:\n        continue\n        \n  \n#########################################################################################################################\n#                                              Final DataFrame\n#########################################################################################################################\n\np_list_df_Final = p_list_df.drop(columns=['Temperature', 'Position'])\np_list_df_Final.tail()\n#delete when no longer needed\ndel p_list_df\n#collect residual garbage\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Distribution plots for Weather type, Stadium type and Field type"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"############################################################################################################################\n#                                         Stadium TYPE DISTRIBUTION\n############################################################################################################################\n\nimport plotly.express as px\n    \ndata = px.histogram(p_list_df_Final, x=\"StadiumType\", histnorm='percent',\n                   title='<b>Distribution of Stadium Type</b>', opacity=0.7)\n\nfig1 = go.Figure(data=data, layout=layout)\n\n\nfig1.update_layout(\n    autosize=False,\n    width=500,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig1.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Stadium Type</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig1.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig1.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig1.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='LightPink')\nfig1.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n# ############################################################################################################################\n# #                                         FIELD TYPE DISTRIBUTION\n# ############################################################################################################################\n\ndata = px.histogram(p_list_df_Final, x=\"FieldType\", histnorm='percent',\n                title='<b>Distribution of Field Type</b>', opacity=0.6, color_discrete_sequence=['olivedrab'])\n\nfig2 = go.Figure(data=data, layout=layout)\n\nfig2.update_layout(\n   autosize=False,\n    width=500,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig2.update_layout(\n#     title=\"'<b>Distribution of Stadium Type v/s Field Type</b>'\",\n    title_x=0.5,\n    xaxis_title=\" <b>Field Type</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig2.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig2.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig2.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='LightPink')\nfig2.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n############################################################################################################################\n#                                         Weather DISTRIBUTION\n############################################################################################################################\n\nimport plotly.express as px\n    \ndata = px.histogram(p_list_df_Final, x=\"Weather\", histnorm='percent',\n                   title='<b>Distribution of Weather</b>', opacity=0.8, color_discrete_sequence=['crimson'])\n\nfig3 = go.Figure(data=data, layout=layout)\n\n\nfig3.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig3.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Weather Type</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig3.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig3.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig3.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='LightPink')\nfig3.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n\nfig1.show()\nfig2.show()\nfig3.show()\n\n\ndel fig1\ndel fig2\ndel fig3\ndel data\n#collect residual garbage\ngc.collect()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Inferences from the plots above:\n1. Most of the games played are in outdoors --> Hence weather might be important factor\n2. The field type is almost equally distributed between Natural and Sythetic.\n3. The weather is fairly distributed among all categories , with the exception of Snow/Cold being an outlier."},{"metadata":{},"cell_type":"markdown","source":"#### Joining the Injury Table with Play List Table"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"p_list_df = p_list_df_Final.merge(injury_df_Final, how = 'left', left_on='GameID', right_on='GameID')\np_list_df = p_list_df.drop(columns=['PlayerKey_y', 'PlayKey_y', 'Surface'])\np_list_df = p_list_df.fillna(0)\np_list_df = p_list_df.rename(columns={\"PlayerKey_x\": \"PlayerKey\", \"PlayKey_x\": \"PlayKey\"})\np_list_df = p_list_df.drop_duplicates(subset=['PlayKey']).reset_index().drop(columns=['index'])\np_list_df['Injury_Impact'] = p_list_df['Injury_Impact'].astype(int)\np_list_df['Injured'] = p_list_df['Injured'].astype(int)\np_list_df.to_csv('Merged_Injury_PlayerList.csv', index=False)\np_list_df\n\n\ndel p_list_df_Final\ndel injury_df_Final\n#collect residual garbage\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Feature Engineering the Player Position table"},{"metadata":{},"cell_type":"markdown","source":"Feature Extraction 1:\n**Non Alignment Scoring Column:** This column is created to identify the player body instanteneous stress level. If the Orientation and Direction of a player at that moment is deviated by a large margin it can cause injury. for example. If you are facing North and Running North, you are not causing body Stress, However, If you are facing East and Running North, you are stressing your Body.\n\nCalculation:\n\nStep 1: Bucket the Direction angle and orientation angle into 8 pie of 45 degrees\n\nStep 2: Non-Alignment Score = Direction Angle - Orientation Angle\n\nStep 3: Categorization"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"p_trk_df_raw = pd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv')\np_trk_df = p_trk_df_raw.drop(columns=['x', 'y', 's'])\n#delete when no longer needed\ndel p_trk_df_raw\n#collect residual garbage\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now As we can see that The Position Data frame is One-Tenth of a second Granualar Level. However, Our 'PlayKey' is Minute Level. Hence, I would first Convert this raw data into Scalable format by groupby. However, certail features are grouped with certain criteria:\n\nDistance: Distance is added over the minute\n\nSeconds: Number of Seconds per minute is counted to add weightage to the Speed Calculation (as the speed is not correctly given here and all the minutes do not have 60 sec records)\n\nEvent : Only considering the first non-Nan values in a given Minute\n\nNon-Alignment Score: It is averaged over the minute. Note, we might lose some instanteneous information but as a base line this is good to go.\n\nAverage_Speed : Distance / Seconds\n\n"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Grouping By the individual Column\np_trk_df['Non_Alignment_Score'] = np.absolute((p_trk_df['dir'] - p_trk_df['o'])/45).apply(np.ceil)\np_trk_df['Non_Alignment_Score'] = p_trk_df['Non_Alignment_Score'].replace([5, 6, 7, 8], [3, 2, 1,0])\nprint(p_trk_df['Non_Alignment_Score'].unique())\n\np_trk_df = p_trk_df.drop(columns=['dir','o'])\nDistance = p_trk_df.groupby(['PlayKey']).sum()[['dis']].reset_index().rename({'dis': 'Distance'}, axis=1)\np_trk_df = p_trk_df.drop(columns=['dis'])\nEvent = p_trk_df.dropna(subset=['event']).groupby(['PlayKey']).first()[['event']].reset_index().rename({'event':'Event'}, axis=1)\np_trk_df = p_trk_df.drop(columns=['event'])\nCount = round((p_trk_df.groupby(['PlayKey']).size())/10).reset_index().rename({0: 'Seconds'}, axis=1)\nNon_Alignment_Score = round(p_trk_df.groupby(['PlayKey']).mean()[['Non_Alignment_Score']]).reset_index()\np_trk_df = p_trk_df.drop(columns=['Non_Alignment_Score'])\n\n\ndata_frames = [Distance, Count, Event, Non_Alignment_Score]\nplr_pos_Final = reduce(lambda  left,right: pd.merge(left,right,on=['PlayKey'],how='inner'), data_frames)\n\ndel Distance\ndel Event\ndel Count\ndel Non_Alignment_Score\ngc.collect()\n\nplr_pos_Final['Avg_Spd'] = round(plr_pos_Final['Distance']/plr_pos_Final['Seconds'])\nplr_pos_Final","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the shape above we can see we have lost almost 45 rows (after inner joining) out of 267000 rows. We can Ignore that for the simplicity\n\n\n### Now that our Position table is created, Let's merge it with the other table in order to create our Final Raw DataFramethat can be used for Modeling and Feature Importance Scoring"},{"metadata":{},"cell_type":"markdown","source":"## Creating the Feature-Testing DataFrame\n#### Merging the Player position and Player Scenario DataFrame"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plr_pos_Final.to_csv('Downsized_PlrPos.csv', index=False)\ndata_frames = [p_list_df, plr_pos_Final]\nFtest_Df = reduce(lambda left,right: pd.merge(left,right,on=['PlayKey'], how='inner'), data_frames)\nFtest_Df.to_csv('Merged_Injury_PlayerList_PlayerPos.csv', index=False)\nFtest_Df.tail()\n\n# del p_list_df\n# del plr_pos_Final\n# gc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Feature Categorizing the Main Events to reduce the feature variablility."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Ftest_Df = pd.read_csv('Merged_Injury_PlayerList_PlayerPos.csv')\nFtest_Df['Event'].nunique()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"Huddle = ['huddle_start_offense', 'huddle_break_offense']\nPass = ['pass_forward', 'pass_arrived', 'pass_outcome_incomplete', 'pass_outcome_caught', \n        'pass_tipped', 'pass_shovel', 'pass_outcome_touchdown', 'pass_outcome_interception', 'pass_tipped']          \nPunt = ['punt_play', 'punt', 'punt_land', 'punt_downed', 'punt_received', 'punt_fake','punt_muffed', 'punt_blocked']          \nKickoff = ['kickoff', 'kickoff_land', 'kickoff_play']\nline_set = 'line_set'\nball_snap = 'ball_snap'\nman_in_motion = 'man_in_motion'\nshift = 'shift'\npoint_play = ['two_point_conversion', 'extra_point_attempt', 'extra_point']\nhandoff = 'handoff'\ntimeout = ['timeout_tv', 'timeout', 'timeout_quarter', 'timeout_home']\nkick = ['onside_kick', 'drop_kick', 'free_kick']\nOthers = ['penalty_flag','field_goal_play', 'play_action', 'free_kick_play', 'two_point_play', 'qb_kneel', 'qb_sack',\n               'snap_direct', 'run', 'two_minute_warning']\n\n\nFtest_Df['Event'] = Ftest_Df['Event'].astype(str)   \n          \ndef assign_feature_problem(data):\n    \n    if any(word in data for word in Huddle):\n          data = \"Huddle\"\n          return data\n    elif any(word in data for word in Pass):\n          data = \"Pass\"\n          return data\n    elif any(word in data for word in Punt):\n          data = \"Punt\"\n          return data\n    elif any(word in data for word in Kickoff):\n          data = \"Kickoff\"\n          return data\n    elif any(word in data for word in line_set):\n          data = \"line_set\"\n          return data\n    elif any(word in data for word in ball_snap):\n          data = \"ball_snap\"\n          return data\n    elif any(word in data for word in man_in_motion):\n          data = \"man_in_motion\"\n          return data    \n    elif any(word in data for word in shift):\n          data = \"shift\"\n          return data\n    elif any(word in data for word in point_play):\n          data = \"point_play\"\n          return data\n    elif any(word in data for word in handoff):\n          data = \"handoff\"\n          return data\n    elif any(word in data for word in timeout):\n          data = \"timeout\"\n          return data\n    elif any(word in data for word in kick):\n          data = \"kick\"\n          return data\n    elif any(word in data for word in Others):\n          data = \"Others\"\n          return data  \n\nFtest_Df['Event'] = Ftest_Df['Event'].apply(lambda x : assign_feature_problem(x))\nFtest_Df['Event'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Plotting the Derived Features "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df = Ftest_Df.copy()\n# df['Non_Alignment_Score'] = df['Non_Alignment_Score']\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '0.0', 'Aligned (0 Deg)', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '0.5',  '0-45 deg', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '1.5', '45-90 deg', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '2.5',  '90-135 deg', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '1.0',  '0-45 deg', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '2.0', '45-90 deg', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '3.0',  '90-135 deg', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '3.5', '135-180 deg', df['Non_Alignment_Score'])\ndf['Non_Alignment_Score'] = np.where(df['Non_Alignment_Score'] == '4.0', 'Opposite Aligned (180 deg)', df['Non_Alignment_Score'])\n\n################################################################################################################################\n                                                # Non-Alignment Score : For Injured Players Only\n################################################################################################################################\n\nimport plotly.express as px\n    \n    \ndf = df[df['Injured']== 1]\n\ndata= px.histogram(df, x=\"Non_Alignment_Score\", histnorm='percent',\n                   title='<b>Angular differences b/w Player Direction and Player Orientation at Injury Minute</b>', \n                opacity=0.8, color_discrete_sequence=['paleturquoise'])\n\nfig4b = go.Figure(data=data, layout=layout)\n\n\nfig4b.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig4b.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Non-Alignment Angle</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig4b.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig4b.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig4b.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='Blue')\nfig4b.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n\n################################################################################################################################\n                                                # Average Speed\n################################################################################################################################\n\nimport plotly.express as px\n    \n    \ndf = df[df['Injured']== 1]\n\ndata= px.histogram(df, x=\"Avg_Spd\", histnorm='percent',\n                   title='<b>Distribution of Average_Speed at Injury Minute : For Injured players</b>', \n                opacity=0.8, color_discrete_sequence=['rosybrown'])\n\nfig5 = go.Figure(data=data, layout=layout)\n\n\nfig5.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig5.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Avg_Speed (Yards/Second)</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig5.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig5.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig5.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='Brown')\nfig5.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n################################################################################################################################\n                                                # Event When Injured\n################################################################################################################################\n\nimport plotly.express as px\n    \n    \ndf = df[df['Injured']== 1]\n\ndata= px.histogram(df, x=\"Event\", histnorm='percent',\n                   title='<b>Distribution of Event at Injury Minute : For Injured players</b>', \n                opacity=0.8, color_discrete_sequence=['orchid'])\n\nfig6 = go.Figure(data=data, layout=layout)\n\n\nfig6.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig6.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Event Type</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig6.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig6.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig6.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='Red')\nfig6.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n\n################################################################################################################################\n                                                # player position When Injured\n################################################################################################################################\nimport plotly.express as px\n    \n    \ndf = df[df['Injured']== 1]\n\ndata= px.histogram(df, x=\"PositionGroup\", histnorm='percent',\n                   title='<b>Distribution of Player Position at Injury Minute : For Injured players</b>', \n                opacity=0.8, color_discrete_sequence=['orangered'])\n\nfig7 = go.Figure(data=data, layout=layout)\n\n\nfig7.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig7.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Player Position</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig7.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig7.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig7.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='Brown')\nfig7.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n\n\n################################################################################################################################\n                                                # player position Retained when Injured\n################################################################################################################################\nimport plotly.express as px\n    \n    \ndf = df[df['Injured']== 1]\n\ndata= px.histogram(df, x=\"PositionRetained\", histnorm='percent',\n                   title='<b>Roster Position retained at Injury Time</b>', \n                opacity=0.8, color_discrete_sequence=['crimson'])\n\nfig8 = go.Figure(data=data, layout=layout)\n\n\nfig8.update_layout(\n    autosize=False,\n    width=500,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig8.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Y: Yes, N:No</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig8.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig8.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig8.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='Black')\nfig8.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n\nfig4b.show()\nfig5.show()\nfig6.show()\nfig7.show()\nfig8.show()\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Plot Inferences:\n\n\n1. It can be seen that 100% of the injury occurred to players outside of their Roster position\n2. A 45-90 Degree Difference Betrween Direction and Orientation is likely to end in Non-Contact Injury\n3. A speed of 1-2 yards/sec causes most injuries\n4. Most of the injuries occurs during Huddling and While in 'Defense' Position\n"},{"metadata":{},"cell_type":"markdown","source":"### Feature Extraction : Body Stress\n\nBased on the PlayerDay Timeline, i have calculated the resting days for each Player Between the Two Games ( Keeping in mind the 2 Seasons)\n\nAlso, I have broken each season into 4 parts to calculated the fatigue of a player over multiple games in a season.\n\nIn all total Each season has 16 games each and few players have even played all games in each of these season!\n\nHence a Weighted Scoring Metrics has been set to measure the Body Fatigue and then categorically binned"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"F_Modeling_Df = Ftest_Df.copy()\nF_Modeling_Df = F_Modeling_Df.drop(['RosterPosition', 'PlayKey',\n                                   'Seconds', 'PlayerGamePlay', 'Distance'], axis=1)\nF_Modeling_Df = Ftest_Df.copy()\nF_Modeling_Df = F_Modeling_Df.drop(['RosterPosition', 'PlayKey',\n                                   'Seconds', 'PlayerGamePlay', 'Distance'], axis=1)\n\nF_Modeling_Df = F_Modeling_Df.sort_values(['GameID'], ascending=[True])\n\n# The min value in player day is -62. Hence shifting the timeline to make it timeline in positve\nF_Modeling_Df['PlayerDay'] = F_Modeling_Df['PlayerDay'] + 63\n\n# Also finding the start day of the second season ad 336th day in the new Timeline.  \nF_Modeling_Df['PlayerDay'] = np.where(F_Modeling_Df['PlayerDay'] > 336, F_Modeling_Df['PlayerDay']-336, F_Modeling_Df['PlayerDay'])\n\n\n# Setting up a new feature to calculate the Resting period between the next game\nF_Modeling_Df['Rest_Days'] = 0\n\nfor i in range(len(F_Modeling_Df)-1):\n        if F_Modeling_Df['GameID'][i+1] != F_Modeling_Df['GameID'][i]:\n\n            F_Modeling_Df['Rest_Days'][i+1] = F_Modeling_Df['PlayerDay'][i+1] - F_Modeling_Df['PlayerDay'][i]\n\n        else:\n            F_Modeling_Df['Rest_Days'][i+1] = F_Modeling_Df['Rest_Days'][i] \n\n# Hence adjusting the season 2 timeline\nF_Modeling_Df['Rest_Days'] = np.where(F_Modeling_Df['Rest_Days'] < 0, 0, F_Modeling_Df['Rest_Days'])\n\n# Creating the Player Stress Factor Column\nF_Modeling_Df['D1'] = np.where((F_Modeling_Df['Rest_Days']== 0) , 'Fresh','')\nF_Modeling_Df['D2'] = np.where((F_Modeling_Df['Rest_Days']> 0)& (F_Modeling_Df['Rest_Days']<4.5) , 'High','')\nF_Modeling_Df['D3'] = np.where((F_Modeling_Df['Rest_Days']> 4.5)& (F_Modeling_Df['Rest_Days']<9.5) , 'Medium','')\nF_Modeling_Df['D4'] = np.where((F_Modeling_Df['Rest_Days']> 9.5)& (F_Modeling_Df['Rest_Days']<16.6) , 'Normal','')\nF_Modeling_Df['D5'] = np.where((F_Modeling_Df['Rest_Days']> 16.5)& (F_Modeling_Df['Rest_Days']<30.5) , 'Low','')\nF_Modeling_Df['D6'] = np.where((F_Modeling_Df['Rest_Days']> 30.5) , 'Fresh','')\n\n\nF_Modeling_Df['Player_Stress'] =  (F_Modeling_Df['D1'] + F_Modeling_Df['D3'] + F_Modeling_Df['D4'] + F_Modeling_Df['D5'] +\n                           F_Modeling_Df['D6']  + F_Modeling_Df['D2'] )\n\nF_Modeling_Df = F_Modeling_Df.drop(columns = ['D1','D2', 'D3', 'D4', 'D5', 'D6','PlayerDay','Rest_Days', 'PlayerKey', 'GameID', 'PlayerGame'])\nF_Modeling_Df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Plotting Distribution of Player Stress level when Injured"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df = F_Modeling_Df.copy()\n################################################################################################################################\n                                                # Player Body_Stress at time of Injury\n################################################################################################################################\nimport plotly.express as px\n    \ndf = df[df['Injured']== 1]\n\ndata= px.histogram(df, x=\"Player_Stress\", histnorm='percent',\n                   title='<b>Distribution of Player Stress at Injury Game : For Injured players</b>', \n                opacity=0.8, color_discrete_sequence=['mediumspringgreen'])\n\nfig9 = go.Figure(data=data, layout=layout)\n\n\nfig9.update_layout(\n    autosize=False,\n    width=1000,\n    height=500,\n    margin=go.layout.Margin(\n        l=50,\n        r=50,\n        b=100,\n        t=100,\n        pad=4),\n    paper_bgcolor='rgba(0,0,0,0)',\n    plot_bgcolor='rgba(0,0,0,0)'\n)\n\nfig9.update_layout(\n#     title=\"Plot Title\",\n    title_x=0.5,\n    xaxis_title=\" <b>Player Stress</b>\",\n    yaxis_title=\"<b>Distribution Percentage</b>\",\n    yaxis_range=[0,100],\n    font=dict(\n        family=\"Courier New, monospace\",\n        size=14,\n        color=\"#7f7f7f\"\n    )\n)\n\nfig9.update_xaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig9.update_yaxes(ticks=\"outside\", tickwidth=2, tickcolor='crimson', ticklen=5)\nfig9.update_xaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=True, gridcolor='Green')\nfig9.update_yaxes(showline=True, linewidth=2, linecolor='black', zeroline=True, showgrid=False)\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Feature Description From the Processed Model Table"},{"metadata":{},"cell_type":"markdown","source":"**Key:** 'PlayerKey', 'GameID', 'PlayKey'\n\n**Categorical:** 'BodyPart','RosterPosition','StadiumType','FieldType','Weather','PlayType','Event','PositionGroup'\n\n**Time_Category:** 'Seconds', 'PlayerGamePlay'\n\n**Ordinal:** 'PlayerDay','PlayerGame', 'Injury_Impact', 'Distance', 'Non_Alignment_Score','Avg_Spd'\n\n**Boolean:** 'PositionRetained'\n\n**Target_Variable:** 'Injured'"},{"metadata":{},"cell_type":"markdown","source":"### Feature Elimination by Definition:\n\n\n**'RosterPosition'** : Provides the roster position of player. Good For Analysis not enough for Modeling\n\n**'GameID', 'PlayKey' , 'PlayerKey'**: These are Keys. Good for basic analysis. Not useful for Modeling.\n**\n'Seconds'** : Good for Feature Engineering. But Nothing Else.\n\n**'PlayerGame' , 'PlayerGamePlay'**: Good for analysis. Might be useful to analyse which game-minutes & game Number did they get hurt. Not actually useful for Modeling.\n\n**'Distance':** Useful for analysis. To identify which player has covered more distance. But this essence is already captured by Avg_Spd.\n\n**'BodyPart' **: It is not the cause for injury\n\n\n\n### Final Features For Feature Testing:\n\n**Key:** None\n\n**Categorical:** 'StadiumType','FieldType','Weather','PlayType','Event', 'PlayerStress'\n\n**Time_Category:**\n\n**Ordinal:** 'Injury_Impact', 'Non_Alignment_Score','Avg_Spd'\n\n**Boolean:** 'PositionRetained'\n\n**Target_Variable:** 'Injured'"},{"metadata":{},"cell_type":"markdown","source":"## Categorical-Data Labeling \n\nSince we have many features as categories, we need to label these features for normalization and easier interpretation"},{"metadata":{},"cell_type":"markdown","source":"As we prepare our dataset for modeling, out first step is to check the 'Balance' of the dataframe.\nSInce we have a Binary Target Variable, We can see that the distribution is very Skewed. \nIt has 3385 rows for injured data and 263,000 rows of uninjured tags. Thus in order to balance(50-50 ratio) this dataset, we can either undersample or Over sample.\n\nUndersampling usually causes loss of information. And because we have very high skewed data set Under sampling will perform poorly.\n\nOversampling usually encapsulates all the original information by generating synthetic data for the minority group. However, the run time is very high due to increased data. \n\n### Creating Sample dataframe to overcome Model OVERFIT (Undersampling)"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"F_Modeling_Df_Injured = F_Modeling_Df[F_Modeling_Df['Injured']==1]\nF_Modeling_Df_non_Injured = F_Modeling_Df[F_Modeling_Df['Injured']==0].sample(n = 5000)\nU_Sample_df = pd.concat([F_Modeling_Df_Injured, F_Modeling_Df_non_Injured], ignore_index=True)\nU_Sample_df\n\nX = U_Sample_df.drop(columns=['Injured','Injury_Impact', 'BodyPart', 'PlayType', 'StadiumType', 'PositionRetained']).copy()\ny = U_Sample_df['Injured']\ncategorical_feature_mask = X.dtypes==object\ncategorical_cols = X.columns[categorical_feature_mask].tolist()\nle = LabelEncoder()\nX[categorical_cols] = X[categorical_cols].apply(lambda col: le.fit_transform(col.astype('str')))\nX.tail()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Creating SMOTE (Synthetic Minority Over-sampling Technique): Oversampling the data set"},{"metadata":{"trusted":true},"cell_type":"code","source":"X = F_Modeling_Df.drop(columns=['Injured','Injury_Impact', 'BodyPart', 'PlayType', 'StadiumType', 'PositionRetained']).copy()\ny = F_Modeling_Df['Injured']\n# Categorical boolean mask\ncategorical_feature_mask = X.dtypes==object\n# filter categorical columns using mask and turn it into a list\ncategorical_cols = X.columns[categorical_feature_mask].tolist()\n# instantiate labelencoder object\nle = LabelEncoder()\n# apply le on categorical feature columns\nX[categorical_cols] = X[categorical_cols].apply(lambda col: le.fit_transform(col.astype('str')))\n\n\nfrom sklearn.linear_model.base import MultiOutputMixin\nfrom imblearn.over_sampling import (RandomOverSampler, SMOTE, ADASYN)\n# Resample the minority class. You can change the strategy to 'auto' if you are not sure.\nsm = SMOTE(sampling_strategy='minority', random_state=7)\n# Fit the model to generate the data.\noversampled_trainX, oversampled_trainY = sm.fit_sample(X, y)\nO_Sample_df = pd.concat([pd.DataFrame(oversampled_trainY), pd.DataFrame(oversampled_trainX)], axis=1)\nO_Sample_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Labeling of the Sample data-frame:"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"### For UnderSampling\n\n## MULTI LABEL\n# instantiate OneHotEncoder\nfrom sklearn.preprocessing import OneHotEncoder\nMultilabel = OneHotEncoder(categories='auto', sparse=False)\n# apply OneHotEncoder on categorical feature columns\nU_X_Multilabel = Multilabel.fit_transform(X) # It returns an numpy array\n\n\n## BINARY LABELING\n# instantiate OneHotEncoder\nBinaryLabel = OneHotEncoder(categories='auto', sparse=True)\n# apply OneHotEncoder on categorical feature columns\nU_X_Binarylabel = BinaryLabel.fit_transform(X) # It returns an numpy array\nU_X_Binarylabel","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Feature Testing: Pearson's Correlation Plot"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import seaborn as sns\nsns.set(style=\"white\")\n\n# Compute the correlation matrix\ncorr = X.corr()\n\n# Generate a mask for the upper triangle\nmask = np.zeros_like(corr, dtype=np.bool)\nmask[np.triu_indices_from(mask)] = True\n\n# Set up the matplotlib figure\nf, ax = plt.subplots(figsize=(11, 9))\n\n# Generate a custom diverging colormap\ncmap = sns.diverging_palette(220, 10, as_cmap=True)\n\n# Draw the heatmap with the mask and correct aspect ratio\nsns.heatmap(corr, mask=mask, cmap=cmap, vmax=.3, center=0,\n            square=True, linewidths=.5, cbar_kws={\"shrink\": .5})","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Since we do not see a very high co-relation among features we will not eliminate anyone. \n#### This also implies that our output will not heavily depend on any one features. "},{"metadata":{},"cell_type":"markdown","source":"### Checking the Distribution of Individual features:"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = plt.figure(figsize = (15,20))\nax = fig.gca()\nX.hist(ax = ax)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Cheking the various model Accuracies ( for base level validity): With MultiLabeling"},{"metadata":{"trusted":true},"cell_type":"code","source":"x = U_X_Multilabel\n\nX_train, X_test, y_train, y_test = train_test_split(x,y,test_size=0.30, shuffle=True)\n\nclassifiers=[]\nmodel1 = xgboost.XGBClassifier()\nclassifiers.append(model1)\n# model2 = svm.SVC()\n# classifiers.append(model2)\nmodel3 = tree.DecisionTreeClassifier()\nclassifiers.append(model3)\nmodel4 = RandomForestClassifier()\nclassifiers.append(model4)\n\n\nfor clf in classifiers:\n    clf.fit(X_train, y_train)\n    y_pred= clf.predict(X_test)\n    acc = accuracy_score(y_test, y_pred)\n    print(\"Accuracy of %s is %s\"%(clf, acc))\n    cm = confusion_matrix(y_test, y_pred)\n    print(\"Confusion Matrix of %s is %s\"%(clf, cm))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Cheking the various model Accuracies ( for base level validity): With Binary-Multi Labeling"},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.metrics import precision_recall_fscore_support as score\nfrom sklearn.metrics import f1_score\n\nx = U_X_Binarylabel\n\nX_train, X_test, y_train, y_test = train_test_split(x,y,test_size=0.30, shuffle=True)\n\nclassifiers=[]\nmodel1 = xgboost.XGBClassifier()\nclassifiers.append(model1)\n# model2 = svm.SVC()\n# classifiers.append(model2)\nmodel3 = tree.DecisionTreeClassifier()\nclassifiers.append(model3)\nmodel4 = RandomForestClassifier()\nclassifiers.append(model4)\n\n\nfor clf in classifiers:\n    clf.fit(X_train, y_train)\n    y_pred= clf.predict(X_test)\n    acc = accuracy_score(y_test, y_pred)\n    print(\"Accuracy of %s is %s\"%(clf, acc))\n    cm = confusion_matrix(y_test, y_pred)\n    print(\"Confusion Matrix of %s is %s\"%(clf, cm))\n    f1 = f1_score(y_test, y_pred, average='micro')\n    print(\"F1 Score of %s is %s\"%(clf, f1))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Feature Importance: Random Forest Classifier Model Selected from the above Models\n\nFrom the above f1 scores, we can see that the RF-Classifier score better compared to the rest of the three. Hence RF is chosen for Feature association scoring.\n\n#### Printing the scoring matrix for each Features "},{"metadata":{"trusted":true},"cell_type":"code","source":"X=X\ny=y\n\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import precision_recall_fscore_support as score\nfrom sklearn.model_selection import train_test_split\nX_train,X_test,y_train,y_test = train_test_split(X,y,test_size = 0.3)\n\nrf = RandomForestClassifier(n_estimators= 50, max_depth= 20, n_jobs= -1)\nrf_model = rf.fit(X_train,y_train)\nsorted(zip((rf.feature_importances_)*100,X_train.columns), reverse = True)[0:7]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Hence The top 3 associative feature that contributes to non Contact Injury are:\n1. Position Of the Player\n2. Weather Condition\n3. Player Stress Level"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}