{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Prediction of Punt Location using Player and Ball Positions at Time of Snap\n### &emsp;When a team chooses to punt, the goal of the punting team is usually to give the receiving team poor field position. The outcome of the play depends on many factors: where the punter kicks the ball, the efficacy of the gunners getting downfield, the efficacy of the returner, etc... An important part of this is how the kicking and receiving teams are aligned prior to the play (e.g. if there are multiple jammers on each gunner, if the receiving team is going for a block, if the returner is anticipating a fair catch). The goal of this analysis is to use player and ball positioning in order to estimate where the punter will kick the ball, as this could be beneficial for special teams strategy.<br>&emsp;This notebook uses provided tracking data of players and the football from punt plays in order to build two models that predict where the punter is most likely to kick the ball on a given play. The models accept inputs of the x and y coordinates of all 22 players and the ball at the time of snap (i.e. when \"event\" equals \"ball_snap\") and output the estimated x and y coordinates of where the ensuing punt will land. Two models were developed: one using traditional linear regression and the other a neural network for regression. The models perform about the same against the test data, both having around 12 yards of error on average. The models both have similar error/score/loss values against train and test data which indicates good generalization.<br>&emsp;The notebook is organized into three parts:\n## Part 1: Data Cleansing and Exploratory Visualization\n#### Organization of data for model building and visualization\n## Part 2: Model Implementation\n#### The definition and fitting/training of both the linear regression model and neural network for regression\n## Part 3: Analysis of Model Results\n#### Testing of models\n<br>","metadata":{}},{"cell_type":"code","source":"from types import SimpleNamespace\n\nimport matplotlib.animation as anm\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nfrom IPython.display import display, Markdown, HTML\nfrom keras.models import Sequential\nfrom keras.layers import Dense\nfrom matplotlib.lines import Line2D\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.model_selection import train_test_split\n\nplt.rcParams['animation.html'] = 'jshtml'\nplt.rcParams['animation.embed_limit'] = 100","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:15:57.987952Z","iopub.execute_input":"2022-01-06T04:15:57.988337Z","iopub.status.idle":"2022-01-06T04:16:05.399998Z","shell.execute_reply.started":"2022-01-06T04:15:57.988230Z","shell.execute_reply":"2022-01-06T04:16:05.399056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Part 1: Data Cleansing and Exploratory Visualization","metadata":{}},{"cell_type":"code","source":"# Define constants\nCOLUMN = SimpleNamespace(**{\n    'PLAY_TYPE': 'specialTeamsPlayType',\n    'SPECIAL_TEAMS_RESULT': 'specialTeamsResult',\n    'POSITION': 'position',\n    'PLAY_EVENT': 'event',\n    'GAME_ID': 'gameId',\n    'PLAY_ID': 'playId',\n    'HOME_TEAM': 'homeTeamAbbr',\n    'VISITOR_TEAM': 'visitorTeamAbbr',\n    'TEAM': 'team',\n    'POSSESSION_TEAM': 'possessionTeam',\n    'TIME': 'time',\n    'NFL_ID': 'nflId',\n    'RETURNER_ID': 'returnerId',\n    'PRIMARY_RETURNER_ID': 'primaryReturnerId',\n    'X': 'x',\n    'Y': 'y',\n    'X_X': 'x_x',\n    'Y_X': 'y_x',\n    'X_Y': 'x_y',\n    'Y_Y': 'y_y',\n    'COLOR': 'color',\n    'FRAME': 'frame',\n    'SORT_POSITION': 'sortPosition',\n    'PLAYER_TEAM_ABBR': 'playerTeamAbbr',\n    'IS_PUNT_TEAM': 'isPuntTeam',\n    'PUNT_TEAM_Y_POSITION': 'puntTeamYPosition',\n    'REMAINING_SORT_LOCATION': 'remainingSortLocation',\n    'SORT_VALUE': 'sortValue',\n    'FEATURES': 'features',\n    'LINEAR_REG_PREDICTION': 'linearRegressionPrediction',\n    'NEURAL_NET_PREDICTION': 'neuralNetworkPrediction',\n    'LINEAR_REG_ABS_ERR': 'linearRegressionAbsoluteError',\n    'NEURAL_NET_ABS_ERR': 'neuralNetworkAbsoluteError',\n})\n\nPOSITION = SimpleNamespace(**{\n    'PUNTER': 'P',\n})\n\nPLAY_TYPE = SimpleNamespace(**{\n    'PUNT': 'Punt',\n})\n\nPLAY_EVENT = SimpleNamespace(**{\n    'BALL_SNAP': 'ball_snap',\n    'PUNT': 'punt',\n    'PUNT_RECEIVED': 'punt_received',\n    'OUT_OF_BOUNDS': 'out_of_bounds',\n    'PUNT_LAND': 'punt_land',\n    'FAIR_CATCH': 'fair_catch',\n    'PUNT_MUFFED': 'punt_muffed',\n})\n\nSPECIAL_TEAMS_RESULT = SimpleNamespace(**{\n    'BLOCKED': 'Blocked Kick Attempt',\n    'DOWNED': 'Downed',\n    'RETURN': 'Return',\n    'TOUCHBACK': 'Touchback',\n    'FAIR_CATCH': 'Fair Catch',\n    'MUFFED': 'Muffed',\n    'OUT_OF_BOUNDS': 'Out of Bounds',\n})\n\nTEAM = SimpleNamespace(**{\n    'HOME': 'home',\n    'AWAY': 'away',\n    'FOOTBALL': 'football',\n})\n\nCOLOR = SimpleNamespace(**{\n    'GREEN': 'green',\n    'BLUE': 'blue',\n    'RED': 'red',\n    'PURPLE': 'purple',\n    'ORANGE': 'orange',\n})","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:16:05.402178Z","iopub.execute_input":"2022-01-06T04:16:05.402478Z","iopub.status.idle":"2022-01-06T04:16:05.414714Z","shell.execute_reply.started":"2022-01-06T04:16:05.402439Z","shell.execute_reply":"2022-01-06T04:16:05.414094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load data into dataframes\ngames_df = pd.read_csv('../input/nfl-big-data-bowl-2022/games.csv')\nplays_df = pd.read_csv('../input/nfl-big-data-bowl-2022/plays.csv')\npff_scouting_df = pd.read_csv('../input/nfl-big-data-bowl-2022/PFFScoutingData.csv')\ntracking_2018_df = pd.read_csv('../input/nfl-big-data-bowl-2022/tracking2018.csv')\ntracking_2019_df = pd.read_csv('../input/nfl-big-data-bowl-2022/tracking2019.csv')\ntracking_2020_df = pd.read_csv('../input/nfl-big-data-bowl-2022/tracking2020.csv')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:16:05.415614Z","iopub.execute_input":"2022-01-06T04:16:05.416529Z","iopub.status.idle":"2022-01-06T04:18:17.438883Z","shell.execute_reply.started":"2022-01-06T04:16:05.416480Z","shell.execute_reply":"2022-01-06T04:18:17.437858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop unused columns\ncolumns = list(COLUMN.__dict__.values())\ngames_df = games_df[games_df.columns.intersection(columns)]\nplays_df = plays_df[plays_df.columns.intersection(columns)]\npff_scouting_df = pff_scouting_df[pff_scouting_df.columns.intersection(columns)]\ntracking_2018_df = tracking_2018_df[tracking_2018_df.columns.intersection(columns)]\ntracking_2019_df = tracking_2019_df[tracking_2019_df.columns.intersection(columns)]\ntracking_2020_df = tracking_2020_df[tracking_2020_df.columns.intersection(columns)]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:18:17.440795Z","iopub.execute_input":"2022-01-06T04:18:17.441019Z","iopub.status.idle":"2022-01-06T04:18:19.255018Z","shell.execute_reply.started":"2022-01-06T04:18:17.440993Z","shell.execute_reply":"2022-01-06T04:18:19.254177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initial consolidating and tidying\ngame_plays_df = pd.merge(games_df, plays_df, left_on=COLUMN.GAME_ID, right_on=COLUMN.GAME_ID)\ntracking_df = pd.concat([tracking_2018_df, tracking_2019_df, tracking_2020_df], ignore_index=True)\ntracking_df[COLUMN.TIME] = pd.to_datetime(tracking_df[COLUMN.TIME])\ntracking_df[COLUMN.NFL_ID] = tracking_df[COLUMN.NFL_ID].astype('Int64')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:18:19.256232Z","iopub.execute_input":"2022-01-06T04:18:19.256469Z","iopub.status.idle":"2022-01-06T04:18:34.787478Z","shell.execute_reply.started":"2022-01-06T04:18:19.256425Z","shell.execute_reply":"2022-01-06T04:18:34.786449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add a column indicating the primary/first returner for each play.\n# This will be used as a feature in the models later on.\ndef getPrimaryReturnerId(value):\n    if pd.isnull(value):\n        return np.NaN\n    returners = str(value).split(';')\n    return int(returners[0])\n\ngame_plays_df[COLUMN.PRIMARY_RETURNER_ID] = game_plays_df[COLUMN.RETURNER_ID] \\\n    .apply(getPrimaryReturnerId) \\\n    .astype('Int64')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:18:34.788863Z","iopub.execute_input":"2022-01-06T04:18:34.789125Z","iopub.status.idle":"2022-01-06T04:18:34.825126Z","shell.execute_reply.started":"2022-01-06T04:18:34.789095Z","shell.execute_reply":"2022-01-06T04:18:34.824281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter down to punts only\npunt_tracking_df = pd.merge(\n    pff_scouting_df,\n    tracking_df,\n    left_on=[COLUMN.GAME_ID, COLUMN.PLAY_ID],\n    right_on=[COLUMN.GAME_ID, COLUMN.PLAY_ID],\n)\npunt_plays_df = game_plays_df[game_plays_df[COLUMN.PLAY_TYPE] == PLAY_TYPE.PUNT]\npunt_tracking_df = pd.merge(\n    punt_tracking_df,\n    punt_plays_df,\n    left_on=[COLUMN.GAME_ID, COLUMN.PLAY_ID],\n    right_on=[COLUMN.GAME_ID, COLUMN.PLAY_ID],\n)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:18:34.826363Z","iopub.execute_input":"2022-01-06T04:18:34.826947Z","iopub.status.idle":"2022-01-06T04:18:55.749922Z","shell.execute_reply.started":"2022-01-06T04:18:34.826907Z","shell.execute_reply":"2022-01-06T04:18:55.748918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot player and ball positions over the course of each punt for one game\n\nplot_punts_df = punt_tracking_df[punt_tracking_df[COLUMN.GAME_ID] == 2018090600].copy()\n\nplot_punts_df.loc[:, COLUMN.FRAME] = plot_punts_df[[COLUMN.GAME_ID, COLUMN.TIME]] \\\n    .apply(tuple, axis=1) \\\n    .rank(ascending=True, method='dense')\n\nconditions = [\n    plot_punts_df[COLUMN.TEAM] == TEAM.HOME,\n    plot_punts_df[COLUMN.TEAM] == TEAM.AWAY,\n    plot_punts_df[COLUMN.TEAM] == TEAM.FOOTBALL,\n]\nvalues = [\n    COLOR.GREEN,\n    COLOR.BLUE,\n    COLOR.RED,\n]\nplot_punts_df.loc[:, COLUMN.COLOR] = np.select(conditions, values)\n\nfig, ax = plt.subplots(figsize=(7.5, 4.5))\nax.set(xlim=(-10, 110), ylim=(-10, 63))\nfirst_frame = plot_punts_df[plot_punts_df[COLUMN.FRAME] == 1]\nc = first_frame[COLUMN.COLOR]\nx = first_frame[COLUMN.X]\ny = first_frame[COLUMN.Y]\nscatter = ax.scatter(x, y, c=c)\nlegend_elements = [\n    Line2D([0], [0], marker='o', color='w', label='Home Team', markerfacecolor=COLOR.GREEN),\n    Line2D([0], [0], marker='o', color='w', label='Away Team', markerfacecolor=COLOR.BLUE),\n    Line2D([0], [0], marker='o', color='w', label='Football', markerfacecolor=COLOR.RED),\n]\nax.legend(handles=legend_elements, loc='lower right')\nplt.title('Player and Ball Positions over Time')\n\ndef plot_punts_anim_fn(i):\n    filtered = plot_punts_df[plot_punts_df[COLUMN.FRAME] == i+1]\n    c = filtered[COLUMN.COLOR]\n    x = filtered[COLUMN.X]\n    y = filtered[COLUMN.Y]\n    data = np.c_[x, y]\n    scatter.set_offsets(data)\n    scatter.set_color(c)\n\nplot_punts_anim = anm.FuncAnimation(\n    fig,\n    plot_punts_anim_fn,\n    interval=50,\n    frames=int(plot_punts_df[COLUMN.FRAME].max()),\n    repeat=True,\n)\nplt.close()\nplot_punts_anim","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:18:55.751329Z","iopub.execute_input":"2022-01-06T04:18:55.751552Z","iopub.status.idle":"2022-01-06T04:20:34.230295Z","shell.execute_reply.started":"2022-01-06T04:18:55.751527Z","shell.execute_reply":"2022-01-06T04:20:34.229584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Part 2: Model Implementation","metadata":{}},{"cell_type":"code","source":"'''\nPredict punt landing position based on pre-snap inputs (e.g. ball and player positions)\n\nInputs:\n  - Positions of all players and ball at time of snap (22 players + ball, x and y positions)\nOutputs:\n  - Ball landing position x, y\n  \nA, B, C, D, E, F\n\nA * x0 + B * x1 + C * x2 + D * x3 ... = ball landing position\n\n                      SORT\nx0 = ball x           1\nx1 = ball y           2\nx2 = punter x\nx3 = punter y\nx4 = returner x\nx5 = returner y\nx6 = player3 x\nx7 = player3 y\netc...\n\ny0 = ball landing x\ny1 = ball landing y\n'''\n\nspecial_teams_results = [\n    SPECIAL_TEAMS_RESULT.RETURN,\n    SPECIAL_TEAMS_RESULT.TOUCHBACK,\n    SPECIAL_TEAMS_RESULT.FAIR_CATCH,\n    SPECIAL_TEAMS_RESULT.DOWNED,\n    SPECIAL_TEAMS_RESULT.MUFFED,\n    SPECIAL_TEAMS_RESULT.OUT_OF_BOUNDS,\n]\nball_land_events = [\n    PLAY_EVENT.PUNT_RECEIVED,\n    PLAY_EVENT.PUNT_LAND,\n    PLAY_EVENT.FAIR_CATCH,\n    PLAY_EVENT.PUNT_MUFFED,\n]\npunt_predict_df = punt_tracking_df[\n    punt_tracking_df[COLUMN.SPECIAL_TEAMS_RESULT].isin(special_teams_results)\n]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:34.231485Z","iopub.execute_input":"2022-01-06T04:20:34.231840Z","iopub.status.idle":"2022-01-06T04:20:37.166655Z","shell.execute_reply.started":"2022-01-06T04:20:34.231810Z","shell.execute_reply":"2022-01-06T04:20:37.165675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Label positions of football, punter and returner as features 1, 2 and 3\nis_football = punt_predict_df[COLUMN.TEAM] == TEAM.FOOTBALL\nis_punter = punt_predict_df[COLUMN.POSITION] == POSITION.PUNTER\nis_returner = (~np.isnan(punt_predict_df[COLUMN.PRIMARY_RETURNER_ID])) \\\n    & (~np.isnan(punt_predict_df[COLUMN.NFL_ID])) \\\n    & (punt_predict_df[COLUMN.PRIMARY_RETURNER_ID] == punt_predict_df[COLUMN.NFL_ID])\ninputs_df = punt_predict_df[punt_predict_df[COLUMN.PLAY_EVENT] == PLAY_EVENT.BALL_SNAP].copy()\n\nconditions = [\n    (inputs_df[COLUMN.TEAM] == TEAM.FOOTBALL),\n    (inputs_df[COLUMN.POSITION] == POSITION.PUNTER),\n    np.where(inputs_df[COLUMN.PRIMARY_RETURNER_ID].fillna(-1) == inputs_df[COLUMN.NFL_ID].fillna(-2), True, False),\n]\nvalues = [1, 2, 3]\ninputs_df.loc[:, COLUMN.SORT_POSITION] = np.select(conditions, values)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:37.169568Z","iopub.execute_input":"2022-01-06T04:20:37.169812Z","iopub.status.idle":"2022-01-06T04:20:43.626109Z","shell.execute_reply.started":"2022-01-06T04:20:37.169784Z","shell.execute_reply":"2022-01-06T04:20:43.625502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Use the remaining player positions sorted by team and y position for the rest of our features\ninputs_df.loc[:, COLUMN.PLAYER_TEAM_ABBR] = np.where(\n    inputs_df[COLUMN.TEAM] == TEAM.HOME,\n    inputs_df[COLUMN.HOME_TEAM],\n    inputs_df[COLUMN.VISITOR_TEAM],\n)\ninputs_df.loc[:, COLUMN.IS_PUNT_TEAM] = np.where(\n    inputs_df[COLUMN.PLAYER_TEAM_ABBR] == inputs_df[COLUMN.POSSESSION_TEAM],\n    True,\n    False,\n)\ninputs_df.loc[:, COLUMN.PUNT_TEAM_Y_POSITION] = inputs_df[[COLUMN.IS_PUNT_TEAM, COLUMN.Y]].apply(tuple, axis=1)\ninputs_df.loc[:, COLUMN.REMAINING_SORT_LOCATION] = inputs_df.groupby([COLUMN.GAME_ID, COLUMN.PLAY_ID])[COLUMN.PUNT_TEAM_Y_POSITION] \\\n    .rank(method='first')\ninputs_df.loc[:, COLUMN.REMAINING_SORT_LOCATION] = inputs_df[COLUMN.REMAINING_SORT_LOCATION] + 3\ninputs_df.loc[:, COLUMN.SORT_VALUE] = np.where(\n    inputs_df[COLUMN.SORT_POSITION] > 0,\n    inputs_df[COLUMN.SORT_POSITION],\n    inputs_df[COLUMN.REMAINING_SORT_LOCATION],\n)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:43.627320Z","iopub.execute_input":"2022-01-06T04:20:43.627707Z","iopub.status.idle":"2022-01-06T04:20:45.877854Z","shell.execute_reply.started":"2022-01-06T04:20:43.627665Z","shell.execute_reply":"2022-01-06T04:20:45.877060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assemble sorted tracking data into model input a.k.a features\ninputs_df = inputs_df.sort_values(by=[COLUMN.GAME_ID, COLUMN.PLAY_ID, COLUMN.SORT_VALUE])\n\naggs = { COLUMN.X: lambda x: x.to_list(), COLUMN.Y: lambda y: y.to_list() }\ninputs_df = inputs_df.groupby([COLUMN.GAME_ID, COLUMN.PLAY_ID]).agg(aggs).reset_index()\n\n# Remove plays where we have missing tracking data for the ball and/or one or more players\ninputs_df = inputs_df[\n    (inputs_df[COLUMN.X].map(len) == 23)\n    & (inputs_df[COLUMN.Y].map(len) == 23)\n]\n\n# Remove plays where returner is near line of scrimmage as this is not indicative of a typical punt play\ninputs_df = inputs_df[\n    np.abs(inputs_df[COLUMN.X].str[0] - inputs_df[COLUMN.X].str[2]) > 10\n]\n\n# Combine coordinates into x and y pairs\ndef merge_coordinates(df):\n    merged = []\n    for i in range(23):\n        merged.append(df[COLUMN.X][i])\n        merged.append(df[COLUMN.Y][i])\n    return merged\n\ninputs_df[COLUMN.FEATURES] = inputs_df.apply(merge_coordinates, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:45.879191Z","iopub.execute_input":"2022-01-06T04:20:45.879434Z","iopub.status.idle":"2022-01-06T04:20:47.092036Z","shell.execute_reply.started":"2022-01-06T04:20:45.879406Z","shell.execute_reply":"2022-01-06T04:20:47.091246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter to ball landing position for model output\noutputs_df = punt_predict_df[\n    (punt_predict_df[COLUMN.PLAY_EVENT].isin(ball_land_events))\n    & (punt_predict_df[COLUMN.TEAM] == TEAM.FOOTBALL)\n].sort_values(by=[COLUMN.GAME_ID, COLUMN.PLAY_ID]).copy()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:47.093045Z","iopub.execute_input":"2022-01-06T04:20:47.093844Z","iopub.status.idle":"2022-01-06T04:20:49.927861Z","shell.execute_reply.started":"2022-01-06T04:20:47.093805Z","shell.execute_reply":"2022-01-06T04:20:49.927198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge model inputs/outputs and use 80/20 train test split\nmodel_data_df = pd.merge(\n    inputs_df,\n    outputs_df,\n    left_on=[COLUMN.GAME_ID, COLUMN.PLAY_ID],\n    right_on=[COLUMN.GAME_ID, COLUMN.PLAY_ID],\n)\n\nX = np.array(model_data_df[COLUMN.FEATURES].tolist())\ny = model_data_df[[COLUMN.X_Y, COLUMN.Y_Y]].to_numpy()\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.2, random_state = 1)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:49.929450Z","iopub.execute_input":"2022-01-06T04:20:49.929873Z","iopub.status.idle":"2022-01-06T04:20:49.984049Z","shell.execute_reply.started":"2022-01-06T04:20:49.929826Z","shell.execute_reply":"2022-01-06T04:20:49.983071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit our linear regression model\nreg = LinearRegression().fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:49.985386Z","iopub.execute_input":"2022-01-06T04:20:49.985619Z","iopub.status.idle":"2022-01-06T04:20:50.010562Z","shell.execute_reply.started":"2022-01-06T04:20:49.985593Z","shell.execute_reply":"2022-01-06T04:20:50.009523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate linear regression model predictions\nmodel_data_df[COLUMN.LINEAR_REG_PREDICTION] = reg.predict(X).tolist()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:50.012087Z","iopub.execute_input":"2022-01-06T04:20:50.012924Z","iopub.status.idle":"2022-01-06T04:20:50.037422Z","shell.execute_reply.started":"2022-01-06T04:20:50.012879Z","shell.execute_reply":"2022-01-06T04:20:50.036008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot linear regression model predictions for a handful of plays\nplays = 50\nfig, ax = plt.subplots(figsize=(7.5, 4.5))\nax.set(xlim=(-10, 110), ylim=(-10, 63))\n\n# Football\nfootball_x = model_data_df.iloc[0][COLUMN.X_X][0]\nfootball_y = model_data_df.iloc[0][COLUMN.Y_X][0]\n\n# Punter\npunter_x = model_data_df.iloc[0][COLUMN.X_X][1]\npunter_y = model_data_df.iloc[0][COLUMN.Y_X][1]\n\n# Returner\nreturner_x = model_data_df.iloc[0][COLUMN.X_X][2]\nreturner_y = model_data_df.iloc[0][COLUMN.Y_X][2]\n\n# Actual Punt Placement\nactual_x = model_data_df.iloc[0][COLUMN.X_Y]\nactual_y = model_data_df.iloc[0][COLUMN.Y_Y]\n\n# Predicted Punt Placement\npredict_x = model_data_df.iloc[0][COLUMN.LINEAR_REG_PREDICTION][0]\npredict_y = model_data_df.iloc[0][COLUMN.LINEAR_REG_PREDICTION][1]\n\n# Plot the initial data points\ninitial_x = [\n    football_x,\n    punter_x,\n    returner_x,\n    actual_x,\n    predict_x,\n]\ninitial_y = [\n    football_y,\n    punter_y,\n    returner_y,\n    actual_y,\n    predict_y,\n]\ncolors = [\n    COLOR.RED,\n    COLOR.GREEN,\n    COLOR.BLUE,\n    COLOR.PURPLE,\n    COLOR.ORANGE,\n]\nscatter = ax.scatter(x=initial_x, y=initial_y, c=colors)\n\n# Add legend/title\nlegend_elements = [\n    Line2D([0], [0], marker='o', color='w', label='Football At Snap', markerfacecolor=COLOR.RED),\n    Line2D([0], [0], marker='o', color='w', label='Punter', markerfacecolor=COLOR.GREEN),\n    Line2D([0], [0], marker='o', color='w', label='Returner', markerfacecolor=COLOR.BLUE),\n    Line2D([0], [0], marker='o', color='w', label='Actual Punt Placement', markerfacecolor=COLOR.PURPLE),\n    Line2D([0], [0], marker='o', color='w', label='Predicted Punt Placement', markerfacecolor=COLOR.ORANGE),\n]\nax.legend(handles=legend_elements, loc='lower right')\nplt.title('Linear Regression Punt Placement Prediction (Play by Play)')\n\ndef linear_reg_anim_fn(i):\n    row = model_data_df.iloc[i+1]\n    c = [\n        COLOR.RED,\n        COLOR.GREEN,\n        COLOR.BLUE,\n        COLOR.PURPLE,\n        COLOR.ORANGE,\n    ]\n    x = np.array([\n        row[COLUMN.X_X][0],\n        row[COLUMN.X_X][1],\n        row[COLUMN.X_X][2],\n        row[COLUMN.X_Y],\n        row[COLUMN.LINEAR_REG_PREDICTION][0],\n    ])\n    y = np.array([\n        row[COLUMN.Y_X][0],\n        row[COLUMN.Y_X][1],\n        row[COLUMN.Y_X][2],\n        row[COLUMN.Y_Y],\n        row[COLUMN.LINEAR_REG_PREDICTION][1],\n    ])\n    scatter.set_offsets(np.c_[x, y])\n    scatter.set_color(c)\n\nlinear_reg_anim = anm.FuncAnimation(\n    fig,\n    linear_reg_anim_fn,\n    interval=1000,\n    frames=plays,\n    repeat=True,\n)\nplt.close()\nlinear_reg_anim","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:50.043927Z","iopub.execute_input":"2022-01-06T04:20:50.044366Z","iopub.status.idle":"2022-01-06T04:20:54.222387Z","shell.execute_reply.started":"2022-01-06T04:20:50.044318Z","shell.execute_reply":"2022-01-06T04:20:54.221398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predict X\nmodel_data_df[COLUMN.LINEAR_REG_PREDICTION].str[0].describe()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:54.224029Z","iopub.execute_input":"2022-01-06T04:20:54.224949Z","iopub.status.idle":"2022-01-06T04:20:54.246406Z","shell.execute_reply.started":"2022-01-06T04:20:54.224903Z","shell.execute_reply":"2022-01-06T04:20:54.245501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Actual X\nmodel_data_df[COLUMN.X_Y].describe()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:54.247749Z","iopub.execute_input":"2022-01-06T04:20:54.248535Z","iopub.status.idle":"2022-01-06T04:20:54.258658Z","shell.execute_reply.started":"2022-01-06T04:20:54.248490Z","shell.execute_reply":"2022-01-06T04:20:54.258040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predict Y\nmodel_data_df[COLUMN.LINEAR_REG_PREDICTION].str[1].describe()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:54.259929Z","iopub.execute_input":"2022-01-06T04:20:54.260199Z","iopub.status.idle":"2022-01-06T04:20:54.282274Z","shell.execute_reply.started":"2022-01-06T04:20:54.260167Z","shell.execute_reply":"2022-01-06T04:20:54.281474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Standard deviation of predicted y position is much smaller than actual\n# Actual Y\nmodel_data_df[COLUMN.Y_Y].describe()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:54.283824Z","iopub.execute_input":"2022-01-06T04:20:54.284265Z","iopub.status.idle":"2022-01-06T04:20:54.294019Z","shell.execute_reply.started":"2022-01-06T04:20:54.284234Z","shell.execute_reply":"2022-01-06T04:20:54.293183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the neural network regression model\ndef get_model(n_inputs, n_outputs):\n    model = Sequential()\n    model.add(Dense(23 * 2 * 10, input_dim=n_inputs, kernel_initializer='he_uniform', activation='relu'))\n    model.add(Dense(n_outputs))\n    model.compile(loss='mae', optimizer='adam')\n    return model\n\nmodel = get_model(X.shape[1], y.shape[1])","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:54.295487Z","iopub.execute_input":"2022-01-06T04:20:54.295969Z","iopub.status.idle":"2022-01-06T04:20:54.456956Z","shell.execute_reply.started":"2022-01-06T04:20:54.295935Z","shell.execute_reply":"2022-01-06T04:20:54.456100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the neural network regression model\nmodel.fit(X_train, y_train, verbose=0, epochs=1000)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:20:54.458439Z","iopub.execute_input":"2022-01-06T04:20:54.458745Z","iopub.status.idle":"2022-01-06T04:23:54.965263Z","shell.execute_reply.started":"2022-01-06T04:20:54.458705Z","shell.execute_reply":"2022-01-06T04:23:54.964097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate neural network model predictions\nmodel_data_df[COLUMN.NEURAL_NET_PREDICTION] = model.predict(X).tolist()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:23:54.966934Z","iopub.execute_input":"2022-01-06T04:23:54.967509Z","iopub.status.idle":"2022-01-06T04:23:55.239492Z","shell.execute_reply.started":"2022-01-06T04:23:54.967461Z","shell.execute_reply":"2022-01-06T04:23:55.238223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot neural network regression model predictions for a handful of plays\nplays = 50\nfig, ax = plt.subplots(figsize=(7.5, 4.5))\nax.set(xlim=(-10, 110), ylim=(-10, 63))\n\n# Football\nfootball_x = model_data_df.iloc[0][COLUMN.X_X][0]\nfootball_y = model_data_df.iloc[0][COLUMN.Y_X][0]\n\n# Punter\npunter_x = model_data_df.iloc[0][COLUMN.X_X][1]\npunter_y = model_data_df.iloc[0][COLUMN.Y_X][1]\n\n# Returner\nreturner_x = model_data_df.iloc[0][COLUMN.X_X][2]\nreturner_y = model_data_df.iloc[0][COLUMN.Y_X][2]\n\n# Actual Punt Placement\nactual_x = model_data_df.iloc[0][COLUMN.X_Y]\nactual_y = model_data_df.iloc[0][COLUMN.Y_Y]\n\n# Predicted Punt Placement\npredict_x = model_data_df.iloc[0][COLUMN.NEURAL_NET_PREDICTION][0]\npredict_y = model_data_df.iloc[0][COLUMN.NEURAL_NET_PREDICTION][1]\n\n# Plot the initial data points\ninitial_x = [\n    football_x,\n    punter_x,\n    returner_x,\n    actual_x,\n    predict_x,\n]\ninitial_y = [\n    football_y,\n    punter_y,\n    returner_y,\n    actual_y,\n    predict_y,\n]\ncolors = [\n    COLOR.RED,\n    COLOR.GREEN,\n    COLOR.BLUE,\n    COLOR.PURPLE,\n    COLOR.ORANGE,\n]\nscatter = ax.scatter(x=initial_x, y=initial_y, c=colors)\n\n# Add legend/title\nlegend_elements = [\n    Line2D([0], [0], marker='o', color='w', label='Football At Snap', markerfacecolor=COLOR.RED),\n    Line2D([0], [0], marker='o', color='w', label='Punter', markerfacecolor=COLOR.GREEN),\n    Line2D([0], [0], marker='o', color='w', label='Returner', markerfacecolor=COLOR.BLUE),\n    Line2D([0], [0], marker='o', color='w', label='Actual Punt Placement', markerfacecolor=COLOR.PURPLE),\n    Line2D([0], [0], marker='o', color='w', label='Predicted Punt Placement', markerfacecolor=COLOR.ORANGE),\n]\nax.legend(handles=legend_elements, loc='lower right')\nplt.title('Neural Network Punt Placement Prediction (Play by Play)')\n\ndef neural_net_anim_fn(i):\n    row = model_data_df.iloc[i+1]\n    c = [\n        COLOR.RED,\n        COLOR.GREEN,\n        COLOR.BLUE,\n        COLOR.PURPLE,\n        COLOR.ORANGE,\n    ]\n    x = np.array([\n        row[COLUMN.X_X][0],\n        row[COLUMN.X_X][1],\n        row[COLUMN.X_X][2],\n        row[COLUMN.X_Y],\n        row[COLUMN.NEURAL_NET_PREDICTION][0],\n    ])\n    y = np.array([\n        row[COLUMN.Y_X][0],\n        row[COLUMN.Y_X][1],\n        row[COLUMN.Y_X][2],\n        row[COLUMN.Y_Y],\n        row[COLUMN.NEURAL_NET_PREDICTION][1],\n    ])\n    scatter.set_offsets(np.c_[x, y])\n    scatter.set_color(c)\n\nneural_net_anim = anm.FuncAnimation(\n    fig,\n    neural_net_anim_fn,\n    interval=1000,\n    frames=plays,\n    repeat=True,\n)\nplt.close()\nneural_net_anim","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:23:55.241115Z","iopub.execute_input":"2022-01-06T04:23:55.241380Z","iopub.status.idle":"2022-01-06T04:23:59.401331Z","shell.execute_reply.started":"2022-01-06T04:23:55.241348Z","shell.execute_reply":"2022-01-06T04:23:59.400589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Part 3: Analysis of Model Results","metadata":{}},{"cell_type":"code","source":"# Calculate mean absolute error for both models against all data.\n# Use Pythagorean Theorem to calculate the \"direct\" yardage between estimate and actual.\nmodel_data_df[COLUMN.LINEAR_REG_ABS_ERR] = (\n    (model_data_df[COLUMN.X_Y] - model_data_df[COLUMN.LINEAR_REG_PREDICTION].str[0]) ** 2 +\n    (model_data_df[COLUMN.Y_Y] - model_data_df[COLUMN.LINEAR_REG_PREDICTION].str[1]) ** 2\n) ** 0.5\nmodel_data_df[COLUMN.NEURAL_NET_ABS_ERR] = (\n    (model_data_df[COLUMN.X_Y] - model_data_df[COLUMN.NEURAL_NET_PREDICTION].str[0]) ** 2 +\n    (model_data_df[COLUMN.Y_Y] - model_data_df[COLUMN.NEURAL_NET_PREDICTION].str[1]) ** 2\n) ** 0.5","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:23:59.402803Z","iopub.execute_input":"2022-01-06T04:23:59.403457Z","iopub.status.idle":"2022-01-06T04:23:59.438710Z","shell.execute_reply.started":"2022-01-06T04:23:59.403409Z","shell.execute_reply":"2022-01-06T04:23:59.437742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate mean absolute error for both models against test data.\n# Use Pythagorean Theorem to calculate the \"direct\" yardage between estimate and actual.\ny_test_predict = reg.predict(X_test)\nlinearRegressionAbsoluteErrorTest = (\n    (y_test[:, 0] - y_test_predict[:, 0]) ** 2 +\n    (y_test[:, 1] - y_test_predict[:, 1]) ** 2\n) ** 0.5\ny_nn_test_predict = model.predict(X_test)\nneuralNetworkAbsoluteErrorTest = (\n    (y_test[:, 0] - y_nn_test_predict[:, 0]) ** 2 +\n    (y_test[:, 1] - y_nn_test_predict[:, 1]) ** 2\n) ** 0.5","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:23:59.440487Z","iopub.execute_input":"2022-01-06T04:23:59.440973Z","iopub.status.idle":"2022-01-06T04:23:59.523569Z","shell.execute_reply.started":"2022-01-06T04:23:59.440935Z","shell.execute_reply":"2022-01-06T04:23:59.522641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compare absolute error w/ all data vs test data to see if linear regression model is overfitting\nall_data_mean_error = model_data_df[COLUMN.LINEAR_REG_ABS_ERR].mean()\ntest_data_mean_error = np.mean(linearRegressionAbsoluteErrorTest, axis=0)\ndisplay(Markdown(f'**Linear Regression Model Absolute Error w/ All Data:** {all_data_mean_error}'))\ndisplay(Markdown(f'**Linear Regression Model Absolute Error w/ Test Data:** {test_data_mean_error}'))\n# Score the linear regression model on train and test data to see if model is overfitting\ntrain_data_score = reg.score(X_train, y_train)\ntest_data_score = reg.score(X_test, y_test)\ndisplay(Markdown(f'**Linear Regression Model Score w/ Train Data:** {train_data_score}'))\ndisplay(Markdown(f'**Linear Regression Model Score w/ Test Data:** {test_data_score}'))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:23:59.528486Z","iopub.execute_input":"2022-01-06T04:23:59.528760Z","iopub.status.idle":"2022-01-06T04:23:59.564341Z","shell.execute_reply.started":"2022-01-06T04:23:59.528731Z","shell.execute_reply":"2022-01-06T04:23:59.563344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compare absolute error w/ all data vs test data to see if neural net regression is overfitting\nall_data_mean_error = model_data_df[COLUMN.NEURAL_NET_ABS_ERR].mean()\ntest_data_mean_error = np.mean(neuralNetworkAbsoluteErrorTest, axis=0)\ndisplay(Markdown(f'**Neural Network Absolute Error w/ All Data:** {all_data_mean_error}'))\ndisplay(Markdown(f'**Neural Network Absolute Error w/ Test Data:** {test_data_mean_error}'))\n# Evaluate the neural net regression on train and test data to see if model is overfitting\ndisplay(Markdown('**Neural Network Evaluated on Train Data:**'))\nmodel.evaluate(X_train, y_train)\ndisplay(Markdown('**Neural Network Evaluated on Test Data:**'))\n_ = model.evaluate(X_test, y_test)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T04:23:59.567439Z","iopub.execute_input":"2022-01-06T04:23:59.568852Z","iopub.status.idle":"2022-01-06T04:24:00.009919Z","shell.execute_reply.started":"2022-01-06T04:23:59.568790Z","shell.execute_reply":"2022-01-06T04:24:00.009237Z"},"trusted":true},"execution_count":null,"outputs":[]}]}