{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.12"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Github for all code and supporting notebooks: <a href='https://github.com/pschlais/kaggle_nfl_bdb_2022'>https://github.com/pschlais/kaggle_nfl_bdb_2022</a>","metadata":{}},{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"Punts are one of the most volatile plays in football due to the change in possession, high speeds, and large distances covered during the play. From the receiving team's point of view, attempting a return or calling for a fair catch is a split-second decision that can have a large impact on the result of the play. When a return is attempted, the outcome of that risk-reward decision is easy to quantify with multiple metrics: yards gained, possession retained, and player injury. However, it's hard to quantify the outcome of a fair catch with existing metrics because the play has been cut off; the yards that would've been gained if a return was attempted are an opportunity cost for the returning team. From a fan's perspective, determining how \"good\" or \"bad\" a fair catch decision is on any given punt is an educated guess at best.\n\nAs an example, look at the play below. Was a fair catch a good decision on this play? The catch is at frame 76. If no fair catch was called, the closest defender was a few yards away but closing at high speed, resulting in no return yards and a forceful tackle. However, if that tackle was missed, there was a crease towards the bottom sideline and blockers in the area. \n\nWithout a framework to quantify fair catches, any comparisons or judgments are subjective.","metadata":{}},{"cell_type":"code","source":"# add local directory to import path\nimport os\nimport sys\nmodule_path = os.path.abspath(os.path.join('.'))\nif module_path not in sys.path:\n    sys.path.append(module_path)\n    \nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n    \n    \n#### --- Standard imports ------\nimport pickle\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport xgboost\nfrom sklearn.metrics import mean_absolute_error, r2_score\n\nfrom IPython.display import HTML, display\nfrom matplotlib.colors import LinearSegmentedColormap\n\n# local import\nsys.path.insert(1, '/kaggle/input/nfl-bdb22-pyfiles')\nimport nflplot\nimport nflutil\nimport nfl_bdb22","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-09-05T17:05:30.248979Z","iopub.execute_input":"2023-09-05T17:05:30.250116Z","iopub.status.idle":"2023-09-05T17:05:32.922140Z","shell.execute_reply.started":"2023-09-05T17:05:30.250086Z","shell.execute_reply":"2023-09-05T17:05:32.920913Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"year = 2020\ntrack_df = pd.read_csv(f'/kaggle/input/nfl-big-data-bowl-2022/tracking{year}.csv')\nplay_df = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/plays.csv')\ngame_df = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/games.csv')\nplayer_df = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/players.csv')\npff_df = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/PFFScoutingData.csv')","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:05:32.927033Z","iopub.execute_input":"2023-09-05T17:05:32.929450Z","iopub.status.idle":"2023-09-05T17:06:03.384859Z","shell.execute_reply.started":"2023-09-05T17:05:32.929411Z","shell.execute_reply":"2023-09-05T17:06:03.384004Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"game_id = 2020121312\nplay_id = 2829\nplayvis = nflplot.PlayAnimation(track_df, play_df, game_df, game_id, play_id)\nplt.close()\n\nprint(play_df.loc[(play_df.gameId==game_id) & (play_df.playId==play_id), 'playDescription'].iloc[0])\n\nHTML(playvis.animation.to_jshtml())","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:03.386134Z","iopub.execute_input":"2023-09-05T17:06:03.386522Z","iopub.status.idle":"2023-09-05T17:06:23.959100Z","shell.execute_reply.started":"2023-09-05T17:06:03.386498Z","shell.execute_reply":"2023-09-05T17:06:23.958038Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The goal of this analysis is to provide a quantitative basis for evaluating teams and players on their ability to determine when calling for a fair catch is a good decision relative to the risks of attempting a return. Since the severity of risks are subjective, an objective way to rank fair catches against one another is to focus on the reward side of the tradeoff. This analysis quantifies the reward by estimating return yardage if a return was attempted instead of a fair catch, which will be called **Predicted Yards Forfeited (PYF)** for the remainder of this report.\n\nIn other words, a fair catch on a play with low Predicted Yards Forfeited is a better decision than a fair catch on a play with high Predicted Yards Forfeited.","metadata":{}},{"cell_type":"markdown","source":"# Analysis/Modeling","metadata":{}},{"cell_type":"markdown","source":"## Premise for model to predict return yardage on fair catch plays","metadata":{}},{"cell_type":"markdown","source":"* Use the state of the game at the moment a fair catch is signaled to predict the return yardage if a fair catch had not been signaled (PYF). This **must** be the prediction point because defender behavior changes beyond this time, due to protections given to the returner by the rule book (e.g. slowing down and/or altering path to avoid contact to avoid a penalty).\n* Training a model to predict the eventual return yardage at any point during an actual return will provide coverage and generality for any fair catch play. This can generate PYF for any arbitrary fair catch signal timing during the play.\n* Only generating features based on the state of the players without consideration for the teams or individual players provides generality across time, as the general structure and rules of punt plays have not changed substantially in recent years. This allows the model to be applied to fair catches in the past with a very low risk of target leakage.","metadata":{}},{"cell_type":"markdown","source":"## Quantifying factors that may influence return yardage","metadata":{}},{"cell_type":"markdown","source":"At each frame:\n* **Time until the catch**\n* **Expected defenders in proximity**: at current distances and speeds of defenders, how many will reach the player? How many within 5, 10, 20 yards?\n* **Distance** between defenders and returner\n* **Time required by defender to reach the returner** at given speed\n* **\"Will-Reach Factor\"**: Expected distance covered by the defender at current speed compared to the current distance to the returner, normalized by time until the catch. Higher magnitudes mean more certainty that the defender will (or won't) reach the returner and therefore more impact on the play result. As time to catch decreases, certainty increases and this reflects as higher magnitudes due to a decreasing denominator.\n* **\"Leverage\"**: lateral distance between defender and returner. This is a proxy for the returner having space or increased chance to run past the defender, force a poor tackle attempt, or break a tackle.\n* **Current speed of the returner**: absolute, lateral, and downfield\n* **Distance from the nearest sideline** - how much open space is available to the returner laterally on both sides\n\nOther potential factors that could be investigated but were not included as part of this analysis:\n* Weather conditions\n* Field/surface type\n* \"Zones of control\" as demonstrated by soccer analytics and past NFL Big Data Bowl projects","metadata":{}},{"cell_type":"markdown","source":"### Determining how many defenders to include in per-defender features","metadata":{}},{"cell_type":"markdown","source":"Intuition says that the closer a defender is to the returner at the time of the catch, the more likely the defender will be the one who tackles the returner. The tracking data and PFF data contain all of the information needed to confirm this theory, as well as calculate frequencies of what defender makes the tackle, where the defenders are ordered based on their distance from the returner at the time of the catch.\n\nThe two plots below are the frequencies and CDF for the relative position of the eventual tackler on the play for the 2020 season.","metadata":{}},{"cell_type":"code","source":"# limit to punt plays\nplay_punt_mask = ((play_df['specialTeamsPlayType']=='Punt') & \n                  (play_df['specialTeamsResult']=='Return') &\n                  ~play_df['returnerId'].astype(str).str.contains(';')\n                 )\n\nrpac_df = (play_df.loc[play_punt_mask, ['gameId','playId','returnerId']]\n .astype({\"returnerId\": float})\n # attach the returner ID and frame ID of the \"punt_received\" event to each play\n .merge(nflutil.get_frame_of_event(track_df, 'punt_received'),\n        how='inner',\n        on=['gameId','playId'])\n # get the returner position on the field when the punt is caught\n .merge(track_df[['gameId','playId','frameId','nflId','team','x','y']],\n        how='inner',\n        left_on=['gameId','playId','frameId','returnerId'],\n        right_on=['gameId','playId','frameId','nflId'])\n .drop(columns=['nflId'])\n .rename(columns={'x':'x_returner', 'y':'y_returner', 'team':'teamReturner'})\n # attach home and away team labels for each game\n .merge(game_df[['gameId','homeTeamAbbr','visitorTeamAbbr']],\n        how='inner',\n        on='gameId')\n # attach the returner location to each opposing team player's position\n .merge(track_df[['gameId','playId','frameId','nflId','team','jerseyNumber','x','y']],\n        how='inner',\n        on=['gameId','playId','frameId'])\n .query('teamReturner != team and team != \"football\"')\n .astype({'jerseyNumber': int})\n # assemble identifier for punting team players\n .assign(puntTeamAbbr=lambda df_: np.where(df_['team']=='home', df_['homeTeamAbbr'], df_['visitorTeamAbbr']),\n         jerseyID=lambda df_: df_['puntTeamAbbr'] + ' ' + df_['jerseyNumber'].astype(str).str.rjust(2, '0'))\n .drop(columns=['teamReturner', 'team', 'homeTeamAbbr', 'visitorTeamAbbr'])\n # calculate distance to returner\n .assign(dist=lambda df_: np.sqrt((df_['x']-df_['x_returner'])**2 + (df_['y']-df_['y_returner'])**2))\n # attach distance order within given play\n .sort_values(['gameId','playId','dist'])\n .assign(dist_order=lambda df_: df_.groupby(['gameId','playId']).cumcount()+1)\n # merge with PFF tackler data\n .merge(pff_df[['gameId','playId','tackler','kickContactType']].dropna(subset=['tackler']),\n        how='inner',\n        left_on=['gameId','playId','jerseyID'],\n        right_on=['gameId','playId','tackler'])\n .merge(player_df[['nflId','displayName']],\n        how='left',\n        on='nflId')\n .query('kickContactType == \"CC\"')  # air clean catch\n)","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:23.962621Z","iopub.execute_input":"2023-09-05T17:06:23.963740Z","iopub.status.idle":"2023-09-05T17:06:28.712800Z","shell.execute_reply.started":"2023-09-05T17:06:23.963703Z","shell.execute_reply":"2023-09-05T17:06:28.711785Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.countplot(data=rpac_df, x='dist_order', color='blue')\n_ = ax.set(xlabel='Position at Catch (1 = closest to returner)', title=\"Position of Eventual Tackler at Time of Catch, Relative to Teammates\")","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:28.714121Z","iopub.execute_input":"2023-09-05T17:06:28.714382Z","iopub.status.idle":"2023-09-05T17:06:28.981583Z","shell.execute_reply.started":"2023-09-05T17:06:28.714361Z","shell.execute_reply":"2023-09-05T17:06:28.979985Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, the closer the defender is to the returner relative to the defender's teammates, the more likely the defender is to make the tackle. A closer look at the CDF is needed to determine a suitable cutoff \"N\" for generating features for the N-closest defenders.","metadata":{}},{"cell_type":"code","source":"rpac_cdf = rpac_df.dist_order.value_counts().sort_index().cumsum() / len(rpac_df)\nprint('CDF:')\nprint(rpac_cdf)\nax = sns.lineplot(x=rpac_cdf.index, y=rpac_cdf.values, marker='o', color='blue')\n_ = ax.set(xlabel='Eventual Tackler Position <= Given Position at Catch', title=\"CDF: Position of Eventual Tackler at Time of Catch, Relative to Teammates\",\n           ylabel='CDF (proportion)')\nax.set_ylim([0,1])\nax.minorticks_on()\nax.tick_params(axis='x', which='minor', bottom=False)\nax.set_xticks(range(1, np.max(rpac_cdf.index) + 1))\nax.grid(visible=True, which='major', axis='y', alpha=0.5)\nax.grid(visible=True, which='minor', axis='y', alpha=0.15)","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:28.982678Z","iopub.execute_input":"2023-09-05T17:06:28.983621Z","iopub.status.idle":"2023-09-05T17:06:29.748420Z","shell.execute_reply.started":"2023-09-05T17:06:28.983570Z","shell.execute_reply":"2023-09-05T17:06:29.747499Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For model simplicity, the closest **5** defenders will be used for per-defender features. The closest 5 defenders account for the tackle in 71% of plays. This is half of the players defending (excluding the punter, who trails the play), and even if these 5 players do not make the tackle, it's reasonable to assume they have a major influence on the returner's path, speed, and direction which indirectly determines the outcome of the play.","metadata":{}},{"cell_type":"markdown","source":"## Model Training and Evaluation","metadata":{}},{"cell_type":"markdown","source":"An XGBoost Regression model is trained on clean-catch punt returns where a return is attempted, for all frames between the punt and the catch, to predict the return yardage based on the game state features during that frame.\n\nTo prevent model overfitting/memorization in the train-test split and cross validation, all frames for a given play were grouped together for the purposes of determining the data splits. In other words, a play is exclusively in either the train, cross-validation, or test split at any point in the model fitting process. The tracking data is 10 hertz, so successive frames are highly correlated with each other and would likely overfit the model if frames of the same play were spread across the data splits.\n\nThe training set is relevant frames from 2018, 2019, and Weeks 1-8 of 2020, using 5-fold cross validation with play grouping as described above. The test set is relevant frames from Weeks 9-17 of 2020 to simulate the model being applied to future data.","metadata":{}},{"cell_type":"code","source":"# get data and generated features\ngame_df = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/games.csv')\nplay_df = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/plays.csv')\ndf_features = pd.read_csv('/kaggle/input/nfl-bdb22-pyfiles/model_features.csv', index_col=0)\n\n# -- split games for train and test -------------------------\n# Train = 2018, 2019, and Weeks 1-8 of 2020\n# Test = Weeks 9-17 of 2020\nweeks = range(9,18)\ntest_games = game_df.loc[game_df.week.isin(weeks) & (game_df.season==2020), ['gameId','season','week']]\ntrain_games = game_df.loc[~(game_df.week.isin(weeks) & (game_df.season==2020)), ['gameId','season','week']]\n\n# -- TRAIN SET -------------------------------------------\nmodel_df = df_features.loc[df_features.gameId.isin(train_games['gameId']), :]\nX_train = model_df.drop(columns='kickReturnYardage')\nY_train = model_df[['gameId','playId','frameId','kickReturnYardage']]\n\n# -- TEST SET --------------------------------------------\nmodel_test_df = df_features.loc[df_features.gameId.isin(test_games.gameId)]\nX_test = model_test_df.drop(columns='kickReturnYardage')\nY_test = model_test_df[['gameId', 'playId', 'frameId', 'kickReturnYardage']]\n\n# -- LOAD MODEL (see github notebooks for details) ----------\nwith open('/kaggle/input/nfl-bdb22-pyfiles/yd_predict_model.pickle', 'rb') as f:\n    model = pickle.load(f)","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:29.749348Z","iopub.execute_input":"2023-09-05T17:06:29.749594Z","iopub.status.idle":"2023-09-05T17:06:30.992729Z","shell.execute_reply.started":"2023-09-05T17:06:29.749571Z","shell.execute_reply":"2023-09-05T17:06:30.991100Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Quantitative Performance","metadata":{}},{"cell_type":"code","source":"# Training set\nmae_train = mean_absolute_error(Y_train.kickReturnYardage, model.predict(X=X_train.drop(columns=['gameId','playId','frameId'])))\nprint('Training set results')\nprint(f'MAE: {mae_train}')\nprint()\n\n# Test set\nmae_test = mean_absolute_error(Y_test.kickReturnYardage, model.predict(X=X_test.drop(columns=['gameId','playId','frameId'])))\nprint('Test set results')\nprint(f'MAE: {mae_test}')\nprint()","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:30.994638Z","iopub.execute_input":"2023-09-05T17:06:30.995013Z","iopub.status.idle":"2023-09-05T17:06:31.145455Z","shell.execute_reply.started":"2023-09-05T17:06:30.994982Z","shell.execute_reply":"2023-09-05T17:06:31.143791Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"An average error of over 6 yards is not particularly good for predicting return yardage on a single play, given that most returns are under 10 yards. This is not entirely unexpected because: \n\n***Punt returns are a high-variance play.*** Since punts span roughly 50 yards from the line of scrimmage, players are spread out and the impact of the initial defender making a tackle versus missing a tackle can be the difference between a 25 yard gain and a stop for a loss.\n\n***This model is predicting far into the future.*** The yardage gain on plays are already difficult to predict when a defined player has the ball and every defender is attempting to tackle that player; this observation was the problem statement for the <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020\">2020 Big Data Bowl</a>. Predicting eventual return yardage seconds before the ball gets to the returner adds even more uncertainty about player positioning and speeds that have a big impact on the end result.\n\nEven though the model may not predict the absolute return yardage correctly on a given play, if it captures enough of the signal, it can still be used as a prediction for relative ranking. To investigate this, the actual return yardage in the test set was binned and the predicted yardage for those returns was plotted using a box plot.\n\nIn general, looking at the medians in each actual return yardage bin below, an increase in the actual return yardage correlates to an increase in the predicted return yardage. There appears to be enough signal in this model to do a relative analysis.","metadata":{}},{"cell_type":"code","source":"# add predictions and metadata to Test set targets\nY_test_w_pred = Y_test.copy()\nY_test_w_pred['kickReturnYardage_pred'] = model.predict(X=X_test.drop(columns=['gameId','playId','frameId']))\nY_test_w_pred['diff'] = Y_test_w_pred.kickReturnYardage_pred - Y_test_w_pred.kickReturnYardage\nY_test_w_pred['yardGroup'] = pd.cut(Y_test_w_pred.kickReturnYardage, bins=[-10,0,5,10,20,30,100], right=False)\nY_test_w_pred['timeToCatch'] = (\n        # catch frame for the play\n        Y_test_w_pred.merge(Y_test_w_pred.groupby(['gameId','playId'])['frameId'].max().rename('catch_frame'), how='inner', on=['gameId','playId'])['catch_frame'].to_numpy()\n        -\n        Y_test_w_pred['frameId'].to_numpy()\n) / 10 # convert to seconds\n\n# create boxplot of actual return yardage bins vs. predicted return yardage\n_, ax = plt.subplots()\nsns.boxplot(data=Y_test_w_pred, x='yardGroup',y='kickReturnYardage_pred', ax=ax)\nax.set_title('Test Set: Actual Return Yardage Bins vs. Predicted Return Yardage')\nax.set_ylabel('Predicted')\nax.set_xlabel('Actual');","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:31.147081Z","iopub.execute_input":"2023-09-05T17:06:31.147367Z","iopub.status.idle":"2023-09-05T17:06:31.452358Z","shell.execute_reply.started":"2023-09-05T17:06:31.147345Z","shell.execute_reply":"2023-09-05T17:06:31.450992Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Qualitative Performance","metadata":{}},{"cell_type":"markdown","source":"Visualizing the tracking data for the largest and smallest predictions on ***fair catch*** plays (i.e. Predicted Yards Forfeited - PYF) during the 2020 season can give an indication if the model is worthwhile to use as the basis for a relative ranking.\n\nThe trained model is applied to the frame where the returner signals for a fair catch to generate PYF for that decision. In this dataset, the fair catch signal event is not labeled; the \"fair_catch\" event identifier is the actual catch. For the purpose of model demonstration, the fair catch signal frame is assumed to be **1.5 seconds (15 frames)** before the catch on all fair catch plays. This event could be labeled in future datasets to more accurately select the proper frame for a play.","metadata":{}},{"cell_type":"code","source":"# -- generate features for fair catch plays, and downselect to \"fair catch signal\" frames ----------------------\nn_defenders = 5\nyears = [2020]\ndfs = []\nfor year in years:\n    track_df = nflutil.transform_tracking_data(pd.read_csv(f'/kaggle/input/nfl-big-data-bowl-2022/tracking{year}.csv'))\n\n    temp_feature_df = (\n        nfl_bdb22.prep_get_modeling_frames(track_df, \n                                    play_df=play_df, \n                                    pff_df=pff_df, \n                                    play_end_event_name='fair_catch'\n                                    )\n        .pipe(nfl_bdb22.model_create_features, play_df=play_df, game_df=game_df, n_defenders=n_defenders, catch_type='fair_catch')\n        .query('timeToCatch==1.5')\n    )\n    dfs.append(temp_feature_df)\n\n# turn into single dataframe (if multiple years were transformed)\nfeature_df = pd.concat(dfs)\n\n# -- add predictions ------------------------------------------------------------\npredict_df = (\n                feature_df[['gameId','playId','frameId']]\n                .assign(pry_pred=model.predict(X=feature_df.drop(columns=['gameId','playId','frameId'])))\n                .reset_index(drop=True)\n            )\n\n# -- get base dataframe for punt returns --------------------------------------------\npuntreturn_df = (\n    predict_df\n    # get returner nflId\n    .merge(play_df[['gameId','playId','possessionTeam','returnerId']],\n           how='inner',\n           on=['gameId', 'playId'])\n    .assign(returnerId=lambda df_: pd.to_numeric(df_.returnerId))\n    # get team information\n    .merge(game_df[['gameId','homeTeamAbbr','visitorTeamAbbr','season','week']],\n           how='inner',\n           on=['gameId'])\n    # set the returning team - the opposite of the possesion team (defined as the one kicking/punting the ball)\n    .assign(receivingTeam=lambda df_: np.where(df_.possessionTeam==df_.homeTeamAbbr, df_.visitorTeamAbbr, df_.homeTeamAbbr))\n    .drop(columns=['homeTeamAbbr', 'visitorTeamAbbr'])\n    # add returner name\n    .merge(player_df[['nflId', 'displayName']],\n           how='left',\n           left_on='returnerId',\n           right_on='nflId')\n    .drop(columns='returnerId')\n)\n\npuntreturn_df.sort_values('pry_pred', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:31.455368Z","iopub.execute_input":"2023-09-05T17:06:31.455661Z","iopub.status.idle":"2023-09-05T17:06:53.909779Z","shell.execute_reply.started":"2023-09-05T17:06:31.455639Z","shell.execute_reply":"2023-09-05T17:06:53.908312Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### High predicted return yardage on fair catch","metadata":{}},{"cell_type":"markdown","source":"Looking at the highest predicted value on a fair catch by Jabrill Peppers (gameId = 2020112902, play = 3007, PYF = 11.4 yards), from a visual perspective this is a very poor fair catch decision. Assuming a fair catch signal at frame 62, the ball is ahead of the closest gunner (Wilson, FS #40) and the ball will clearly arrive before the gunner arrives. Beyond Wilson, there is no other defender close by. Peppers has the option to go towards the bottom sideline to make the first gunner miss a tackle, or run angled towards the wide open field.\n\nIf a return was attempted on this play, there was an opportunity for a big return, with a high floor of a few yard gain. This fair catch was a poor decision, and the model aligns with that assessment.","metadata":{}},{"cell_type":"code","source":"game_id = 2020112902\nplay_id = 3007\nplayvis = nflplot.PlayAnimation(track_df, play_df, game_df, game_id, play_id)\nplt.close()\n\nprint(play_df.loc[(play_df.gameId==game_id) & (play_df.playId==play_id), 'playDescription'].iloc[0])\n\nHTML(playvis.animation.to_jshtml())","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:06:53.911195Z","iopub.execute_input":"2023-09-05T17:06:53.911599Z","iopub.status.idle":"2023-09-05T17:07:14.688502Z","shell.execute_reply.started":"2023-09-05T17:06:53.911569Z","shell.execute_reply":"2023-09-05T17:07:14.687185Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Low predicted return yardage on fair catch","metadata":{}},{"cell_type":"markdown","source":"Looking at the lowest predicted value on a fair catch by Jakeem Grant (gameId = 2020110106, play = 2722, PYF = 1.2 yards), from a visual perspective this appears to be a good fair catch decision. Assuming a fair catch signal at frame 68, Grant has identified this as a short punt relative to where he lined up to receive it, and he is running forward to get into position to catch the ball. The closest gunner (Jefferson, WR #12) is clearly going to arrive before the ball reaches Grant, which increases the likelihood of an immediate tackle, or a big hit in an unprotected position with a fumble or an injury. Beyond Jefferson, the other Rams defenders are closing in so there will be multiple defenders in position to tackle.\n\nIf a return was attempted on this play, there was very little opportunity for a big return, with a high likelihood of a poor return result. This fair catch was a good decision, and the model aligns with that assessment.","metadata":{}},{"cell_type":"code","source":"game_id = 2020110106\nplay_id = 2722\nplayvis = nflplot.PlayAnimation(track_df, play_df, game_df, game_id, play_id)\nplt.close()\n\nprint(play_df.loc[(play_df.gameId==game_id) & (play_df.playId==play_id), 'playDescription'].iloc[0])\n\nHTML(playvis.animation.to_jshtml())","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:07:14.689735Z","iopub.execute_input":"2023-09-05T17:07:14.690056Z","iopub.status.idle":"2023-09-05T17:07:36.809973Z","shell.execute_reply.started":"2023-09-05T17:07:14.690033Z","shell.execute_reply":"2023-09-05T17:07:36.808720Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Results","metadata":{}},{"cell_type":"markdown","source":"The model can be used to compare performance of teams and individual returners at a season level by looking at their PYF per fair catch (PYF/FC), simply the average PYF over all fair catches. The results use the 2020 season as an example.","metadata":{}},{"cell_type":"markdown","source":"## Team Performance","metadata":{}},{"cell_type":"code","source":"# --- get punt return performance by team ---------------------------------------------\n# punt plays with return yardage\nplay_filter = (play_df.specialTeamsPlayType=='Punt') & (play_df.specialTeamsResult=='Return') & (~play_df.kickReturnYardage.isna()) & (play_df.gameId.astype(str).str[:4].astype(int).isin(years))\n\nprbyteam_actual_df = (\n    play_df.loc[play_filter, ['gameId','playId','possessionTeam','kickReturnYardage']]\n    # get team information\n    .merge(game_df[['gameId','homeTeamAbbr','visitorTeamAbbr','season','week']],\n           how='inner',\n           on=['gameId'])\n    # set the returning team - the opposite of the possesion team (defined as the one kicking/punting the ball)\n    .assign(receivingTeam=lambda df_: np.where(df_.possessionTeam==df_.homeTeamAbbr, df_.visitorTeamAbbr, df_.homeTeamAbbr))\n    .drop(columns=['homeTeamAbbr', 'visitorTeamAbbr'])\n    .groupby('receivingTeam')['kickReturnYardage']\n    .aggregate(['count','mean'])\n    .reset_index()\n    .rename(columns={'mean':'mean_returnyardage', 'count':'returns'})\n)\n\n# --- add Predicted Yards Forfeited (PYF) and team colors to dataframe ------------------\nteam_colors = {abbr: c['main'] for abbr, c in nflutil.TEAM_COLORS.items()}\nteam_color_df = pd.DataFrame(zip(team_colors.keys(), team_colors.values()), columns=['teamAbbr','teamColor'])\n\nprbyteam_df = (\n    # get predicted yards by team\n    puntreturn_df.groupby('receivingTeam')['pry_pred'].agg(['count','mean']).sort_values('mean', ascending=False).reset_index().rename(columns={'mean':'mean_faircatchlost'})\n    # attach actual yards by team\n    .merge(prbyteam_actual_df,\n           how='inner',\n           on='receivingTeam')\n    # add colors for each team\n    .merge(team_color_df,\n           how='inner',\n           left_on='receivingTeam',\n           right_on='teamAbbr')\n    .sort_values('mean_faircatchlost', ascending=True)\n    )","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:07:36.812552Z","iopub.execute_input":"2023-09-05T17:07:36.813624Z","iopub.status.idle":"2023-09-05T17:07:36.866666Z","shell.execute_reply.started":"2023-09-05T17:07:36.813575Z","shell.execute_reply":"2023-09-05T17:07:36.864939Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create figure\nf, ax = plt.subplots(figsize=(14,5))\n# create bar plot\nax = nflplot.create_team_bar_plot(ax, prbyteam_df['teamAbbr'], prbyteam_df['mean_faircatchlost'], asset_folder_location='/kaggle/input/nfl-bdb22-pyfiles/assets/logos')\n# set labels and grid\nax.set_ylabel('PYF/FC')\nax.set_ylim([0, 7])\nax.set_title('Predicted Yards Forfeited per Fair Catch (PYF/FC) by Team, 2020 Season')\nax.minorticks_on()\nax.tick_params(axis='x', which='minor', bottom=False)\nax.grid(visible=True, which='major', axis='y', alpha=0.5)\nax.grid(visible=True, which='minor', axis='y', alpha=0.15)","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:07:36.867980Z","iopub.execute_input":"2023-09-05T17:07:36.868867Z","iopub.status.idle":"2023-09-05T17:07:39.010803Z","shell.execute_reply.started":"2023-09-05T17:07:36.868836Z","shell.execute_reply":"2023-09-05T17:07:39.009351Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the chart above, special teams units can be ranked against each other to see where each team ranks in the quality of fair catch decision-making. After calculating the PYF/FC for each team for the 2020 season, the 5 best teams are:\n* Miami Dolphins\n* Houston Texans\n* Los Angeles Chargers\n* Jacksonville Jaguars\n* Washington Commanders (name was Football Team in 2020)\n\nThe 5 worst teams are:\n* Carolina Panthers\n* New York Giants\n* San Francisco 49ers\n* Pittsburgh Steelers\n* New England Patriots\n\nDecision-making can also be explored in 2 dimensions, both for return decisions and fair catch decisions. The ideal team would have a high punt return average and low PYF/FC. This comparison is shown in the plot below.","metadata":{}},{"cell_type":"code","source":"#create figure\nf, ax = plt.subplots(figsize=(7,7))\n# create scatter plot\nax = nflplot.create_team_scatter_plot(ax=ax, x=prbyteam_df['mean_returnyardage'], y=prbyteam_df['mean_faircatchlost'], team_labels=prbyteam_df['teamAbbr'], asset_folder_location='/kaggle/input/nfl-bdb22-pyfiles/assets/logos')\n# set labels and grid\nax.set_xlabel('Yards/Return')\nax.set_ylabel('PYF/FC')\nax.set_title('Predicted Yards Forfeited per Fair Catch (PYF/FC) vs. Yards per Return by Team, 2020 Season')\nax.grid(visible=True, which='major', axis='both', alpha=0.4)","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:07:39.012008Z","iopub.execute_input":"2023-09-05T17:07:39.012302Z","iopub.status.idle":"2023-09-05T17:07:39.687615Z","shell.execute_reply.started":"2023-09-05T17:07:39.012279Z","shell.execute_reply":"2023-09-05T17:07:39.686622Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The best decision making is in the bottom right: high yards per return and low PYF/FC. The worst decision making is in the top left: low yards per return and high PYF/FC. A few observations:\n* MIA is far to the bottom right, which means great decisions in both returns and fair catches. They are 1st in PYF/FC and 5th in yards per return.\n* NE is the worst in PYF/FC, but they also have the highest return average. They are very good at returning when they elect to return, so they may benefit from being less conservative about signaling for a fair catch.\n* CAR is near the bottom in PYF/FC and also near the bottom in return average. Their decision making and return ability needs improvement to catch up to the rest of the league.\n","metadata":{}},{"cell_type":"markdown","source":"## Individual Returners","metadata":{}},{"cell_type":"markdown","source":"The same analysis can be applied to individual returners. To remove small sample sizes, the minimum threshold for inclusion is 8 fair catches over the course of the season. This equates to about one fair catch per two games.","metadata":{}},{"cell_type":"code","source":"# set the minimum number of attempts\nmin_faircatch_threshold = 8  # minimum ~1 fair catch per 2 games average\n\n# get relevant player predicted fair catch data\nplayer_analysis_df = (\n    puntreturn_df\n    .groupby('displayName')\n     .filter(lambda x: x.displayName.count() >= min_faircatch_threshold)\n)\n\n# punt plays with return yardage\nplay_filter = (play_df.specialTeamsPlayType=='Punt') & (play_df.specialTeamsResult=='Return') & (~play_df.kickReturnYardage.isna()) & (~play_df.returnerId.astype(str).str.contains(';') & (play_df.gameId.astype(str).str[:4].astype(int).isin(years)))\n\n# add actual return performance per player\nplayer_means = (\n    player_analysis_df.groupby(['displayName','nflId'])['pry_pred'].agg(['count','mean']).sort_values('mean', ascending=True).reset_index()\n    .merge(\n        (\n            play_df.loc[play_filter, ['returnerId','kickReturnYardage']]\n            .assign(returnerId=lambda df_: pd.to_numeric(df_.returnerId))\n            .rename(columns={'returnerId': 'nflId'})\n            .groupby('nflId')['kickReturnYardage'].aggregate(['count','mean'])\n        ),\n        how='left',\n        on='nflId'\n    )\n    .rename(columns={'count_x': 'fairCatches', 'mean_x': 'PYFperFC', 'count_y': 'returns', 'mean_y': 'yardsPerReturn'})\n)\n\nplayer_means.style.background_gradient(cmap='viridis', subset=['PYFperFC', 'yardsPerReturn'])","metadata":{"execution":{"iopub.status.busy":"2023-09-05T17:07:39.688628Z","iopub.execute_input":"2023-09-05T17:07:39.689589Z","iopub.status.idle":"2023-09-05T17:07:39.832572Z","shell.execute_reply.started":"2023-09-05T17:07:39.689563Z","shell.execute_reply":"2023-09-05T17:07:39.830899Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A few observations from the individual returner statistics:\n* **Jakeem Grant** (who was the returner on the lowest PYF play) had a great season for fair catches and returns. He has the lowest PYF/FC and is near the top for Yards per Return.\n* **Jabrill Peppers** (who was the returner on the highest PYF play) did not have a good season for fair catches and returns. He has the 2nd highest PYF/FC and is in the bottom half of Yards per Return.\n* **Gunner Olszewski** and **Diontae Spencer** were both explosive returners with the highest Yards per Return, but were in the bottom half of PYF/FC. Their fair catch decision making could be developed further to become premier punt returners.\n* **Pharoh Cooper** (CAR) is near the bottom for PYF/FC and last in Yards per Return. The Panthers may want to consider making a change at punt returner due to his poor performance in both aspects of returning.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion","metadata":{}},{"cell_type":"markdown","source":"Predicted Yards Forfeited (PYF) provides a relative metric to evaluate fair catch quality on individual plays and at an aggregate level. Here are a few application ideas for this metric:\n\n* **Team-level special teams unit evaluation**: fair catch plays can be quantified and performance evaluated at a unit level, both as the returning unit and the punting unit\n* **Player evaluation**: fair catch performance as another factor for roster construction or determining the depth chart\n* **Enhanced TV viewer experience**: if the player data collection and processing are near-real time, the resulting fair catch and PYF value could be classified into bin-based historical quantiles to give a general indication on the quality of the decision (poor, good, excellent, etc.). This would be similar to using Next-Gen stats in current TV broadcasts.","metadata":{}}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}