{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Logistic Regression vs Random Forest to Predict if a Punt will be Returned**","metadata":{}},{"cell_type":"markdown","source":"In this report, we would like to figure out if we can predict if a punt will be returned or not using two different methods of classification. We will compare Logistic Regression and Random Forest.","metadata":{}},{"cell_type":"markdown","source":"# **1. Import libraries**","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\n\nplt.style.use('fivethirtyeight')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-01-05T22:19:57.541589Z","iopub.execute_input":"2022-01-05T22:19:57.542422Z","iopub.status.idle":"2022-01-05T22:19:58.61301Z","shell.execute_reply.started":"2022-01-05T22:19:57.542313Z","shell.execute_reply":"2022-01-05T22:19:58.611806Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **2. Bringing in data / Cleaning data**","metadata":{}},{"cell_type":"code","source":"#import plays data\nplays = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/plays.csv')","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:58.617054Z","iopub.execute_input":"2022-01-05T22:19:58.617448Z","iopub.status.idle":"2022-01-05T22:19:58.764581Z","shell.execute_reply.started":"2022-01-05T22:19:58.61741Z","shell.execute_reply":"2022-01-05T22:19:58.763807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#filter out punt plays from all plays\npunts = plays[plays['specialTeamsPlayType'] == 'Punt']","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:58.765609Z","iopub.execute_input":"2022-01-05T22:19:58.765847Z","iopub.status.idle":"2022-01-05T22:19:58.78243Z","shell.execute_reply.started":"2022-01-05T22:19:58.765814Z","shell.execute_reply":"2022-01-05T22:19:58.781603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import PFF data\nPFF = pd.read_csv('/kaggle/input/nfl-big-data-bowl-2022/PFFScoutingData.csv')","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:58.784826Z","iopub.execute_input":"2022-01-05T22:19:58.785078Z","iopub.status.idle":"2022-01-05T22:19:58.874391Z","shell.execute_reply.started":"2022-01-05T22:19:58.785043Z","shell.execute_reply":"2022-01-05T22:19:58.873404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#combine punt plays with PFF data\nPuntPFF = PFF.merge(punts, how='inner',on=['gameId','playId'])\nPuntPFF.sort_values(by=['gameId'], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:58.875822Z","iopub.execute_input":"2022-01-05T22:19:58.876149Z","iopub.status.idle":"2022-01-05T22:19:58.912741Z","shell.execute_reply.started":"2022-01-05T22:19:58.876109Z","shell.execute_reply":"2022-01-05T22:19:58.912059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#drop some variables\nPuntPlays = PuntPFF.drop(['passResult', 'penaltyCodes', 'penaltyJerseyNumbers', 'gameClock','preSnapHomeScore','preSnapVisitorScore','kickerId','returnerId','kickBlockerId','down','specialTeamsPlayType','specialTeamsSafeties','vises','tackler','kickoffReturnFormation','gunners','puntRushers','missedTackler','assistTackler'], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:58.915227Z","iopub.execute_input":"2022-01-05T22:19:58.915573Z","iopub.status.idle":"2022-01-05T22:19:58.926563Z","shell.execute_reply.started":"2022-01-05T22:19:58.915529Z","shell.execute_reply":"2022-01-05T22:19:58.925713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nWe began by trimming the data provided down to only punt plays. We removed unnecessary variables to make the data more manageable and allow for our model to run cleanly later. The variables that we decided to continue studying heavily focused on the punt’s context rather than the actions of the players. To restate, the purpose of our model is to predict whether a punt will be returned based on game factors including the quarter in which the play is occurring, the yardline from which the ball is punted, the length of the kick, and other variables.\n\nThis lead us to our final data frame of punt plays.","metadata":{}},{"cell_type":"code","source":"#make punt dataframe and drop more variables\nPuntDF = PuntPlays.drop(['returnDirectionIntended','returnDirectionActual','kickContactType'], axis=1)\nPuntDF","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:58.928334Z","iopub.execute_input":"2022-01-05T22:19:58.928669Z","iopub.status.idle":"2022-01-05T22:19:58.978203Z","shell.execute_reply.started":"2022-01-05T22:19:58.928612Z","shell.execute_reply":"2022-01-05T22:19:58.977173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **3. Data Manipulation**","metadata":{}},{"cell_type":"markdown","source":"### Changing Categorical Variables to Binary","metadata":{}},{"cell_type":"markdown","source":"To better manipulate our variables, we adjusted them from text to boolean and numerical values. ","metadata":{}},{"cell_type":"code","source":"# Create 'returned' boolean column to indicate whether a punt was returned or not\nPuntDF.loc[(PuntDF['specialTeamsResult'] == 'Return'), 'returned'] = 1\nPuntDF['returned'] = PuntDF['returned'].fillna(0)\nPuntDF.reset_index(inplace=True)\n\n# Create 'yardlineNumber' variable which accounts for teams punting over the midway 50 yard line\nPuntDF.loc[(PuntDF['possessionTeam'] == PuntDF['yardlineSide']), 'totalYardline'] = PuntDF['yardlineNumber']\nPuntDF.loc[(PuntDF['possessionTeam'] != PuntDF['yardlineSide']), 'totalYardline'] = ((50 - PuntDF['yardlineNumber']) + 50)\n\n# Re factor kickDirectionActual into numeric\nPuntDF.loc[(PuntDF['kickDirectionActual'] == 'L'), 'kickDirectionActual'] = 0\nPuntDF.loc[(PuntDF['kickDirectionActual'] == 'C'), 'kickDirectionActual'] = 1\nPuntDF.loc[(PuntDF['kickDirectionActual'] == 'R'), 'kickDirectionActual'] = 2\n\n# Re factor kickDirectionIntended into numeric\nPuntDF.loc[(PuntDF['kickDirectionIntended'] == 'L'), 'kickDirectionIntended'] = 0\nPuntDF.loc[(PuntDF['kickDirectionIntended'] == 'C'), 'kickDirectionIntended'] = 1\nPuntDF.loc[(PuntDF['kickDirectionIntended'] == 'R'), 'kickDirectionIntended'] = 2\n\n# Re factor kick type into two groups: normal = 0, non-normal (Aussie/rugby) = 1\nPuntDF.loc[(PuntDF['kickType'] != 'N'), 'kickType'] = 1\nPuntDF.loc[(PuntDF['kickType'] == 'N'), 'kickType'] = 0\n\n# Re factor snapDetail into two groups: 0 = OK, 1 = Left/right/high/low\nPuntDF.loc[(PuntDF['snapDetail'] != 'OK'), 'snapDetail'] = 1\nPuntDF.loc[(PuntDF['snapDetail'] == 'OK'), 'snapDetail'] = 0\n\n","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:58.980005Z","iopub.execute_input":"2022-01-05T22:19:58.980313Z","iopub.status.idle":"2022-01-05T22:19:59.023046Z","shell.execute_reply.started":"2022-01-05T22:19:58.980263Z","shell.execute_reply":"2022-01-05T22:19:59.022045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Data Frame used for analysis with booleans.","metadata":{}},{"cell_type":"code","source":"PuntDF","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:25:11.469201Z","iopub.execute_input":"2022-01-05T22:25:11.46954Z","iopub.status.idle":"2022-01-05T22:25:11.510608Z","shell.execute_reply.started":"2022-01-05T22:25:11.469503Z","shell.execute_reply":"2022-01-05T22:25:11.509746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Check how evenly weighted the two classes are","metadata":{}},{"cell_type":"markdown","source":"Unreturned kicks, represented by 0.0, totaled at 3,705. Returned kicks, represented by 1.0, reached a total of 2286.","metadata":{}},{"cell_type":"code","source":"# Check how evenly weighted the two classes are\nPuntDF['returned'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:59.025075Z","iopub.execute_input":"2022-01-05T22:19:59.025871Z","iopub.status.idle":"2022-01-05T22:19:59.036964Z","shell.execute_reply.started":"2022-01-05T22:19:59.025816Z","shell.execute_reply":"2022-01-05T22:19:59.036052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#change type of variables\nPuntDF['snapDetail'] = PuntDF['snapDetail'].astype('int64')\nPuntDF['kickType'] = PuntDF['kickType'].astype('int64')","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:59.038526Z","iopub.execute_input":"2022-01-05T22:19:59.038942Z","iopub.status.idle":"2022-01-05T22:19:59.05538Z","shell.execute_reply.started":"2022-01-05T22:19:59.038911Z","shell.execute_reply":"2022-01-05T22:19:59.054717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **4. Data Visualization**","metadata":{}},{"cell_type":"markdown","source":"### Checking Correlation","metadata":{}},{"cell_type":"code","source":"# Correlation matrix\n\ncorr = PuntDF.corr()\n\nplt.rcParams[\"figure.figsize\"] = (15,10)\nsns.heatmap(corr, annot=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:19:59.059167Z","iopub.execute_input":"2022-01-05T22:19:59.062885Z","iopub.status.idle":"2022-01-05T22:20:00.87569Z","shell.execute_reply.started":"2022-01-05T22:19:59.062813Z","shell.execute_reply":"2022-01-05T22:20:00.875042Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We found that kickReturnYardage and playResult have high, negative correlation and that playResult and penaltyYards have a significant correlation. All others do not demonstrate a remarkable level of correlation with playResult. ","metadata":{}},{"cell_type":"markdown","source":"### Create target df with features and outcome variable","metadata":{}},{"cell_type":"code","source":"# df with features + target variable\n\nPuntDF_Target = PuntDF[['snapTime','operationTime','hangTime','kickDirectionActual','kickDirectionIntended','kickType','snapDetail','quarter','yardsToGo','kickLength','totalYardline','returned']]\nPuntDF_Target = PuntDF_Target.dropna()\n\nPuntDF_Target['kickDirectionIntended'] = PuntDF_Target['kickDirectionIntended'].astype('int64')\nPuntDF_Target['kickDirectionActual'] = PuntDF_Target['kickDirectionActual'].astype('int64')\nprint(PuntDF_Target.shape[0])\n\n# check correlation between variables\n\nplt.figure(figsize=(12,10))\ncor_target = PuntDF_Target.corr()\nsns.heatmap(cor_target, annot=True, cmap=plt.cm.Reds)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:00.876835Z","iopub.execute_input":"2022-01-05T22:20:00.877162Z","iopub.status.idle":"2022-01-05T22:20:02.051726Z","shell.execute_reply.started":"2022-01-05T22:20:00.877134Z","shell.execute_reply":"2022-01-05T22:20:02.051064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"kickDirectionInteded and kickDirectionActual have very high correlation. totalYardline and kickType have a high correlation as well. Everything else seems to be fine.","metadata":{}},{"cell_type":"code","source":"# Drop kickDirectionIntended\n\nPuntDF_Target = PuntDF_Target.drop('kickDirectionIntended', axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:02.065371Z","iopub.execute_input":"2022-01-05T22:20:02.065869Z","iopub.status.idle":"2022-01-05T22:20:02.072336Z","shell.execute_reply.started":"2022-01-05T22:20:02.065834Z","shell.execute_reply":"2022-01-05T22:20:02.07152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"kickDirectionInteded is dropped due to high correlation and the belief that kickDirectionActual has more influence on a punt return.","metadata":{}},{"cell_type":"markdown","source":"### Let's explore the distribution of the data","metadata":{}},{"cell_type":"code","source":"# Pairplot of the target dataframe\n\nsns.pairplot(PuntDF_Target, hue='returned')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:02.073352Z","iopub.execute_input":"2022-01-05T22:20:02.073563Z","iopub.status.idle":"2022-01-05T22:20:57.41917Z","shell.execute_reply.started":"2022-01-05T22:20:02.073538Z","shell.execute_reply":"2022-01-05T22:20:57.41823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Zooming in on a few of the most revealing plots","metadata":{}},{"cell_type":"code","source":"# Scatter plot of operationTime vs. kickLength with returned punts colored orange\n\nsns.scatterplot(x='operationTime', y='kickLength', hue='returned', data=PuntDF)\nplt.title('Operation Time vs. Kick Length of Punts')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Scatter plot of hangTime vs. kickLength with returned punts colored in orange\n\nsns.scatterplot(PuntDF['hangTime'], PuntDF['kickLength'], hue=PuntDF['returned'])\nplt.title('Hang Time vs. Kick Length of Punts')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Break down totalYardline into returned punts and non-returned punts. Check distribution of where those punts are being\n# kicked from on the field\n\nyardlineNotReturned = PuntDF[PuntDF['returned'] == 0]['totalYardline']\nyardlineReturned = PuntDF[PuntDF['returned'] == 1]['totalYardline']\n\nsns.distplot(yardlineNotReturned, label='No Return')\nsns.distplot(yardlineReturned, label='Return')\nplt.xlabel('Yards from own endzone to LOS')\nplt.title('Location of LOS for Returned and Non-returned Punts')\nplt.legend()\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **5. Building Models**","metadata":{}},{"cell_type":"markdown","source":"### Logistic Regression","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn import metrics\nfrom sklearn.linear_model import LogisticRegression\n\nX_train, X_test, y_train, y_test = train_test_split(PuntDF_Target[['snapTime','operationTime','hangTime','kickDirectionActual','kickType','snapDetail','quarter','yardsToGo','kickLength','totalYardline']], PuntDF_Target[['returned']] , test_size=0.2, random_state=123)\n\nlogreg = LogisticRegression(max_iter=500)\nlogreg.fit(X_train, y_train.values.ravel())\npredictions = logreg.predict(X_test)\nscore = logreg.score(X_test, y_test)\nprint(score)","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:57.420776Z","iopub.execute_input":"2022-01-05T22:20:57.421027Z","iopub.status.idle":"2022-01-05T22:20:57.906788Z","shell.execute_reply.started":"2022-01-05T22:20:57.421Z","shell.execute_reply":"2022-01-05T22:20:57.905702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Logistic Regression has a moderate score for classification.","metadata":{}},{"cell_type":"code","source":"# Classification report for logistic regression model\n\nprint(metrics.classification_report(y_test, predictions))","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:57.90888Z","iopub.execute_input":"2022-01-05T22:20:57.909725Z","iopub.status.idle":"2022-01-05T22:20:57.931613Z","shell.execute_reply.started":"2022-01-05T22:20:57.909661Z","shell.execute_reply":"2022-01-05T22:20:57.930713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We found that the logistic regression model had a strong recall rate for 0 (punts not returned). The model was less successful predicting 1 (punts that are returned).","metadata":{}},{"cell_type":"code","source":"# Heat map for actual and predicted values for logistic regression\n\ncm = metrics.confusion_matrix(y_test, predictions)\nplt.figure(figsize=(9,9))\nsns.heatmap(cm, annot=True, fmt=\".3f\", linewidths=.5, square=True, cmap='Blues_r');\nplt.ylabel('Actual label');\nplt.xlabel('Predicted label');\nall_sample_title = 'Accuracy Score: {0}'.format(score)\nplt.title(all_sample_title, size = 15);","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:57.933366Z","iopub.execute_input":"2022-01-05T22:20:57.933979Z","iopub.status.idle":"2022-01-05T22:20:58.243329Z","shell.execute_reply.started":"2022-01-05T22:20:57.933935Z","shell.execute_reply":"2022-01-05T22:20:58.242439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As demonstrated above, this model is effective for punts not returned (high correct predictions). It lacks when predicting true punt returns.","metadata":{}},{"cell_type":"code","source":"# ROC curve for logistic regression\n\nimport scikitplot as skplt\n\nplt.rcParams['figure.figsize'] = [10, 10]\n\npredicted_probas = logreg.predict_proba(X_test)\n\nskplt.metrics.plot_roc(y_test, predicted_probas)\nplt.title('ROC Curves: Logistic Regression Classifier')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:58.245123Z","iopub.execute_input":"2022-01-05T22:20:58.245744Z","iopub.status.idle":"2022-01-05T22:20:58.559755Z","shell.execute_reply.started":"2022-01-05T22:20:58.245696Z","shell.execute_reply":"2022-01-05T22:20:58.558694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Random Forest Classification","metadata":{}},{"cell_type":"code","source":"# Random Forest classification\n\nfrom sklearn.ensemble import RandomForestClassifier\n\nrf = RandomForestClassifier(n_estimators=500, oob_score=True, random_state=123)\nrf.fit(X_train, y_train.values.ravel())","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:20:58.561371Z","iopub.execute_input":"2022-01-05T22:20:58.561693Z","iopub.status.idle":"2022-01-05T22:21:03.148435Z","shell.execute_reply.started":"2022-01-05T22:20:58.561658Z","shell.execute_reply":"2022-01-05T22:21:03.147468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# OOB score and accuracy score of RF model\n\npredicted = rf.predict(X_test)\naccuracy = metrics.accuracy_score(y_test, predicted)\nprint(f'Out-of-bag score estimate: {rf.oob_score_:.3}')\nprint(f'Mean accuracy score: {accuracy:.3}')","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:03.150275Z","iopub.execute_input":"2022-01-05T22:21:03.150852Z","iopub.status.idle":"2022-01-05T22:21:03.328002Z","shell.execute_reply.started":"2022-01-05T22:21:03.150768Z","shell.execute_reply":"2022-01-05T22:21:03.327048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Classification report for RF model\n\nprint(metrics.classification_report(y_test, predicted))","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:03.329195Z","iopub.execute_input":"2022-01-05T22:21:03.329424Z","iopub.status.idle":"2022-01-05T22:21:03.342578Z","shell.execute_reply.started":"2022-01-05T22:21:03.329398Z","shell.execute_reply":"2022-01-05T22:21:03.341547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We found that the Random Forest classifcation was much better at predicting the punt returns than the logistic regression model. It is also slightly better at predicting the punts not returned.","metadata":{}},{"cell_type":"code","source":"cm = metrics.confusion_matrix(y_test, predicted)\nplt.figure(figsize=(9,9))\nsns.heatmap(cm, annot=True, fmt=\".3f\", linewidths=.5, square=True, cmap='Blues_r');\nplt.ylabel('Actual label');\nplt.xlabel('Predicted label');\nall_sample_title = 'Accuracy Score: {0}'.format(accuracy)\nplt.title(all_sample_title, size = 15);","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:03.344366Z","iopub.execute_input":"2022-01-05T22:21:03.345181Z","iopub.status.idle":"2022-01-05T22:21:03.586807Z","shell.execute_reply.started":"2022-01-05T22:21:03.34511Z","shell.execute_reply":"2022-01-05T22:21:03.58585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As mentioned above, Random Forest is still good at predicting punts not returned and better at making accurate predictions of punt returns. The overall accuracy score saw an increase of .043 from the logistic regression model we created. ","metadata":{}},{"cell_type":"code","source":"# ROC curve for RF model\n\nplt.rcParams['figure.figsize'] = [10, 10]\n\npredicted_probas = rf.predict_proba(X_test)\n\nskplt.metrics.plot_roc(y_test, predicted_probas)\nplt.title('ROC Curves: Random Forest Classifier')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:03.58846Z","iopub.execute_input":"2022-01-05T22:21:03.58884Z","iopub.status.idle":"2022-01-05T22:21:03.99558Z","shell.execute_reply.started":"2022-01-05T22:21:03.588793Z","shell.execute_reply":"2022-01-05T22:21:03.994539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We decided here that we would use Random Forest for the rest of our model, as it is more accurate in predicitng correctly. Our next step was to determine feature importance for our variables. ","metadata":{}},{"cell_type":"markdown","source":"# **6. Feature Importance**","metadata":{}},{"cell_type":"code","source":"# RF feature importances using built-in functionality\nrf.feature_importances_","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:03.99732Z","iopub.execute_input":"2022-01-05T22:21:03.997868Z","iopub.status.idle":"2022-01-05T22:21:04.065869Z","shell.execute_reply.started":"2022-01-05T22:21:03.997822Z","shell.execute_reply":"2022-01-05T22:21:04.064847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Feature importance for RF model\n\nfeature_names = np.array(['snapTime','operationTime','hangTime','kickDirectionActual','kickType','snapDetail','quarter','yardsToGo','kickLength','totalYardline'])\nsorted_idx = rf.feature_importances_.argsort()\nplt.barh(feature_names[sorted_idx], rf.feature_importances_[sorted_idx], color='midnightblue')\nplt.xlabel(\"Random Forest Feature Importance\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:04.067362Z","iopub.execute_input":"2022-01-05T22:21:04.068417Z","iopub.status.idle":"2022-01-05T22:21:04.376498Z","shell.execute_reply.started":"2022-01-05T22:21:04.068361Z","shell.execute_reply":"2022-01-05T22:21:04.375599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Two variables that stood out immediately were snapDetail and totalYardline. Snap detail is very low in importance while totalYardline is exceedingly important. The rest of our variables fall between the two. ","metadata":{}},{"cell_type":"code","source":"# Permutation importance values\n# Similar to the in-model importances with a few exceptions\n\nfrom sklearn.inspection import permutation_importance\n\nperm_importance = permutation_importance(rf, X_test, y_test)\nsorted_idx = perm_importance.importances_mean.argsort()\nplt.barh(feature_names[sorted_idx], perm_importance.importances_mean[sorted_idx], color='midnightblue')\nplt.xlabel(\"Permutation Importance\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:04.378318Z","iopub.execute_input":"2022-01-05T22:21:04.378588Z","iopub.status.idle":"2022-01-05T22:21:12.81361Z","shell.execute_reply.started":"2022-01-05T22:21:04.378557Z","shell.execute_reply":"2022-01-05T22:21:12.812656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When looking at permutation importance, we found kicklength took the top spot, but totalYardline is still highly important. snapTime and operationTime are least important.\n\nWe noticed that on both importance checks, snap detail is low. We decide that it may be worth dropping next time.","metadata":{}},{"cell_type":"code","source":"# SHAP Value estimating how much each feature contributes to the prediction\n\nimport shap\nexplainer = shap.TreeExplainer(rf)\nshap_values = explainer.shap_values(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:21:12.814767Z","iopub.execute_input":"2022-01-05T22:21:12.814999Z","iopub.status.idle":"2022-01-05T22:24:45.529728Z","shell.execute_reply.started":"2022-01-05T22:21:12.814973Z","shell.execute_reply":"2022-01-05T22:24:45.528332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"shap.summary_plot(shap_values, X_test, plot_type=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-01-05T22:24:45.530869Z","iopub.status.idle":"2022-01-05T22:24:45.53121Z","shell.execute_reply.started":"2022-01-05T22:24:45.531039Z","shell.execute_reply":"2022-01-05T22:24:45.531057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **7. Summary**","metadata":{}},{"cell_type":"markdown","source":"Ultimately, from the two models we set out to create, we found that the Random Forest was best for predicting whether or not a punt will be returned. The most influential variables that impact the decision of a punt returner are totalYardline, kickLength, and hangTime (not in a specific order). To further polish our model, there were variables that could have been removed to focus more squarely on the factors that impact the decision of the returner. We were happy to close with a model that has shown success in predicting the outcome of a given punt.","metadata":{}}]}