{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Goal\nThis notebook analyzes data on kickoff and punt plays.  \n\nA player in opposing team who receives the ball on a punt or a kickoff play runs in order to return the ball, while the kickoff team run up to the returner and tackle to prevent them from returning.\n\nFor better strategy, it is useful to predict the return distance of a punt from the player tracking data.\nHowever, it is not possible to predict the exact return distance from temporal position data.\nThe return distance increases significantly when the returner slips through the tackle, but this does not always happen.\n\nIn this note, we will obtain the probability distribution of the distance the returner travels from the current position, instead of predicting the return distance and considers the success rate of the tackle and the position of all players.","metadata":{}},{"cell_type":"markdown","source":"# 1. Load the Libraries and Datasets\nWe use the following datasets and the parameters.  \nTo compress data size, downcast function which copied from this page $\\downarrow$ is applied. <br>\nhttps://www.kaggle.com/werooring/nfl-big-data-bowl-basic-eda-for-beginner\n\n* **games.csv** : gameId, homeTeamAbbr\n\n* **plays.csv** : gameId, playId, possessionTeam, specialTeamsPlayType, specialTeamsResult, returnerId, kickReturnYardage\n\n* **tracking20\\*\\*.csv** : x, y, s, dir, event, nflId, displayName, team, frameId, gameId, playId, playDirection","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T20:02:27.298384Z","iopub.execute_input":"2022-01-06T20:02:27.299018Z","iopub.status.idle":"2022-01-06T20:02:27.312169Z","shell.execute_reply.started":"2022-01-06T20:02:27.298854Z","shell.execute_reply":"2022-01-06T20:02:27.311401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# いつものおまじない\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# データをインポートします\n#players = pd.read_csv('../input/nfl-big-data-bowl-2022/players.csv')\n#pffs = pd.read_csv('../input/nfl-big-data-bowl-2022/PFFScoutingData.csv')\ntrack18 = pd.read_csv('../input/nfl-big-data-bowl-2022/tracking2018.csv')\n#track19 = pd.read_csv('../input/nfl-big-data-bowl-2022/tracking2019.csv')\n#track20 = pd.read_csv('../input/nfl-big-data-bowl-2022/tracking2020.csv')\ngames = pd.read_csv('../input/nfl-big-data-bowl-2022/games.csv')\nplays = pd.read_csv('../input/nfl-big-data-bowl-2022/plays.csv')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:02:27.313949Z","iopub.execute_input":"2022-01-06T20:02:27.314420Z","iopub.status.idle":"2022-01-06T20:03:06.762283Z","shell.execute_reply.started":"2022-01-06T20:02:27.314377Z","shell.execute_reply":"2022-01-06T20:03:06.760904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Copied from https://www.kaggle.com/werooring/nfl-big-data-bowl-basic-eda-for-beginner\n\ndef downcast(df, verbose=True):\n    start_mem = df.memory_usage().sum() / 1024**2\n    for col in df.columns:\n        dtype_name = df[col].dtype.name\n        if dtype_name == 'object':\n            pass\n        elif dtype_name == 'bool':\n            df[col] = df[col].astype('int8')\n        elif dtype_name.startswith('int') or (df[col].round() == df[col]).all():\n            df[col] = pd.to_numeric(df[col], downcast='integer')\n        else:\n            df[col] = pd.to_numeric(df[col], downcast='float')\n    end_mem = df.memory_usage().sum() / 1024**2\n#    if verbose:\n#        print('{:.1f}% Compressed'.format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df\n\n\n# 不要な列の削除\ngames = games.drop(['season', 'week', 'gameDate', 'gameTimeEastern','visitorTeamAbbr'], axis=1)\ngames = downcast(games)\n\nplays = plays.drop(['playDescription', 'quarter', 'down', 'yardsToGo', \n                    'kickerId', 'kickBlockerId', 'yardlineSide', 'yardlineNumber', \n                    'gameClock', 'penaltyCodes', 'penaltyJerseyNumbers', 'penaltyYards', \n                    'preSnapHomeScore', 'preSnapVisitorScore', 'passResult', \n                    'kickLength', 'playResult', 'absoluteYardlineNumber'], axis=1)\nplays = downcast(plays)\n\ndf_list = []\nfor df_temp in [track18]:#, track19, track20]:\n    df_temp = df_temp.drop(['time','dis','o','jerseyNumber','position'], axis=1)\n    df_temp = downcast(df_temp)\n    df_list.append(df_temp)\ntrack_data = pd.concat(df_list, axis=0).reset_index(drop=True)\n\n# メモリの節約用\ndel track18#, track19, track20","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:03:06.764815Z","iopub.execute_input":"2022-01-06T20:03:06.765222Z","iopub.status.idle":"2022-01-06T20:03:16.016624Z","shell.execute_reply.started":"2022-01-06T20:03:06.765165Z","shell.execute_reply":"2022-01-06T20:03:16.015350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Data Preparation\nLooking at parameter \"event\" in tracking data, the flow of the game should be in the following order in a normal punt play.\n\n* ball_snap\n* punt\n* punt_received\n* first_contact\n* tackle\n  \nWe will examine return yardage from punt_recieved (or kick_received in kickoff play) to tackle.\n\n\n### Modifying \"plays.csv\" data (+ \"games.csv\")\n1. Limit plays.csv data to only when a return occurs.\n1. Add home team information (\"homeTeamAbbr\") from games.csv data.\n1. Add returner's id (\"returnerId\") as an explanatory variable.\n\n### Modifying \"tracking20\\*\\*.csv\" data\n\n1. Restrict tracking data to only when a return occurs.\n1. Modify the coordinates according to the \"playDirection\".\n1. Convert the coordinate data team from (home / away) to (offense / diffense).\n1. Restrict to the time between receive and tackle.","metadata":{}},{"cell_type":"code","source":"# Return 距離の頻度を調べる\nreturn_plays = plays[(plays['specialTeamsResult']== 'Return')]\n\n# game データから home のチームの情報を追加\nreturn_plays = pd.merge(return_plays, games, on=['gameId'], how='inner')\n\n# returner の id を説明変数に加える\nreturn_plays = return_plays.dropna(subset=['returnerId'])\nreturn_plays['returnerId'] = return_plays['returnerId'].str[:5]\nreturn_plays['returnerId'] = return_plays['returnerId'].astype(int)\n\n\n###\n# 追跡データを return が起きた場合のみに制限\ntrack_data = pd.merge(track_data, return_plays, on=['gameId','playId'], how='inner')\n\n# コートの入れ替えとかあるっぽいので攻撃方向(playDirection)に応じて座標を修正\ntrack_data['x'] = (120-track_data['x']) - (120-track_data['x']*2) * (track_data['playDirection'] == 'right')\ntrack_data['y'] = (53.3-track_data['y']) - (53.3-track_data['y']*2) * (track_data['playDirection'] == 'right')\ntrack_data['dir'] = np.mod(track_data['dir'] + 180 * (track_data['playDirection'] == 'left'), 360)\n\n# 座標データの team を home / away から offense / diffense に変換\nhomeplayer = (track_data['team'] == 'home')                                         # 'football'は'away'のもの\nhometeam   = (track_data['possessionTeam'] == track_data['homeTeamAbbr'])\ntrack_data['team'] = (homeplayer == hometeam).replace({True:'offense', False:'diffense'})\n\n\n###\n# さらにreceive されてからtackleするまでの間に制限\nreceive = track_data[(track_data['displayName']=='football')&((track_data['event']=='punt_received') | (track_data['event']=='kick_received'))]\nreceive = receive[['gameId', 'playId','frameId']]\nreceive = receive.rename(columns={'frameId': 'frame_start'})\n\ntackle = track_data[(track_data['displayName']=='football')&(track_data['event']=='tackle')]\ntackle = tackle[['frameId','gameId', 'playId','x', 'returnerId']]\ntackle = tackle.rename(columns={'frameId': 'frame_end', 'x': 'x_end'})\n\nreturn_frame = pd.merge(receive, tackle, on=['gameId', 'playId'], how='inner')\n#return_frame.head()\n\n\n###\n# いらない変数のお掃除\ndel games, plays, receive, tackle\n\ntrack_data = track_data[track_data['displayName'] != 'football']\ntrack_data = track_data[['x', 'y', 's', 'a', 'dir', 'nflId', 'team', 'frameId', 'gameId', 'playId']]\n#track_data.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:03:16.018481Z","iopub.execute_input":"2022-01-06T20:03:16.018858Z","iopub.status.idle":"2022-01-06T20:03:29.882579Z","shell.execute_reply.started":"2022-01-06T20:03:16.018817Z","shell.execute_reply":"2022-01-06T20:03:29.881358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Add New Feature\nFrom the original dataset,\n\nThe following variables will be adopted as features.\n\n* ball_neighbour(df_track, dist) : number of enemies and allies within (dist) yards around the returner when the player placement follows (df_track)\n* pos_neighbour(df_track, dist) : number of enemies and allies within (dist) yards **in front of** the returner when the player placement follows (df_track)\n\nIn addition, the average of the number of people in the last five frames is also added as an explanatory variable.","metadata":{}},{"cell_type":"code","source":"# 特徴量として、ボールから一定距離の味方・相手選手の人数を用いる\ndef ball_neighbour(df_track, dist):\n    offense_temp = df_track[(df_track['distance'] < dist)&(df_track['team']=='offense')]\n    diffense_temp = df_track[(df_track['distance'] < dist)&(df_track['team']=='diffense')]\n    return [offense_temp.shape[0], diffense_temp.shape[0]]\n\ndef pos_neighbour(df_track, dist):\n    offense_temp = df_track[(df_track['distance'] < dist)&(df_track['distance_x'] >= 0)&(df_track['team']=='offense')]\n    diffense_temp = df_track[(df_track['distance'] < dist)&(df_track['distance_x'] >= 0)&(df_track['team']=='diffense')]\n    return [offense_temp.shape[0], diffense_temp.shape[0]]\n\n# 選手の位置データから特徴量を作成\nresult_list = []\nfor i_index in return_frame.index:\n    game_id = return_frame.loc[i_index,'gameId']\n    play_id = return_frame.loc[i_index,'playId']\n    frame_start = return_frame.loc[i_index,'frame_start']\n    frame_end = return_frame.loc[i_index,'frame_end']\n    x_end = return_frame.loc[i_index,'x_end']\n    returner = return_frame.loc[i_index, 'returnerId']\n\n    # 中間のフレームをすべてデータとして用いる\n    for frame_index in range(frame_start, frame_end, 3):\n        df_temp = track_data[(track_data['gameId']==game_id)&(track_data['playId']==play_id)&(track_data['frameId']==frame_index)].copy()\n        df_temp1 = df_temp[df_temp['team']=='offense'].copy()\n        df_temp2 = df_temp[(df_temp['team']=='diffense')&(df_temp['nflId']!=returner)].copy()\n\n        returner_loc = df_temp[df_temp['nflId']==returner][['x','y','s','a','dir']].values[0]\n        df_temp['distance'] = ((df_temp['x'] - returner_loc[0]) ** 2 + (df_temp['y'] - returner_loc[1]) ** 2) ** 0.5\n        df_temp['distance_x'] = -(df_temp['x'] - returner_loc[0])     # returner より前に人がいるなら正値\n\n        df_temp1['distance'] = 999999\n        for diffense_index in df_temp2.index:\n            diffense_loc = df_temp.loc[diffense_index,['x','y']].values\n            df_temp1['distance'] = np.minimum(df_temp1['distance'], ((df_temp1['x'] - diffense_loc[0]) ** 2 + (df_temp2['y'] - diffense_loc[1]) ** 2) ** 0.5)\n        df_temp1 = df_temp1[df_temp1['distance'] > 0.5]\n        df_temp1['distance'] = ((df_temp1['x'] - returner_loc[0]) ** 2 + (df_temp1['y'] - returner_loc[1]) ** 2) ** 0.5\n\n        feature_item = [i_index, game_id, play_id, frame_index] + returner_loc.tolist()\n        feature_item += ball_neighbour(df_temp, 7) + ball_neighbour(df_temp, 5) + ball_neighbour(df_temp, 3) + ball_neighbour(df_temp, 1) + ball_neighbour(df_temp, 0.5)\n        feature_item += pos_neighbour(df_temp, 7) + pos_neighbour(df_temp, 5) + pos_neighbour(df_temp, 3) + pos_neighbour(df_temp, 1) + pos_neighbour(df_temp, 0.5)\n        feature_item += [ball_neighbour(df_temp1, 7)[0]] + [ball_neighbour(df_temp1, 3)[0]] + [ball_neighbour(df_temp1, 1)[0]]\n        feature_item += [-(x_end-returner_loc[0])]\n        \n        result_list.append(feature_item)\n\ndf_train = pd.DataFrame(result_list, columns=['gameplayId','gameId','playId','frameId', \n                                              'returner_x','returner_y','returner_s','returner_a','returner_dir', \n                                              'offense7','diffense7','offense5','diffense5','offense3','diffense3','offense1','diffense1','offense.5','diffense.5',\n                                              'offense+7','diffense+7','offense+5','diffense+5','offense+3','diffense+3','offense+1','diffense+1','offense+.5','diffense+.5',\n                                              'free-offense7','free-offense3','free-offense1','return_yard'])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:03:29.885174Z","iopub.execute_input":"2022-01-06T20:03:29.885777Z","iopub.status.idle":"2022-01-06T20:27:47.714751Z","shell.execute_reply.started":"2022-01-06T20:03:29.885720Z","shell.execute_reply":"2022-01-06T20:27:47.713543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.to_csv('intermediate_file.csv')\n#df_train = pd.read_csv('../input/nfl2022-app/intermediate_file.csv', index_col=0)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T20:27:47.717160Z","iopub.execute_input":"2022-01-06T20:27:47.717413Z","iopub.status.idle":"2022-01-06T20:27:48.174757Z","shell.execute_reply.started":"2022-01-06T20:27:47.717385Z","shell.execute_reply":"2022-01-06T20:27:48.173366Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 機械学習で使う用のデータ\ndf_list = []\nfor i_index in return_frame.index:\n    df_temp = df_train[df_train['gameplayId']==i_index]\n    df_temp = df_temp.rolling(5, min_periods=1).sum() / 5    # 前の 5 フレーム間のデータの平均\n    df_list.append(df_temp)\n\ndf_train_mean = pd.concat(df_list, axis=0)\ndf_train_mean = df_train_mean.drop(['gameplayId','gameId','playId','frameId','returner_s','returner_a','returner_dir','return_yard'], axis=1)\ndf_train_mean = df_train_mean.add_suffix('_mean')\n\ndf_train = pd.concat([df_train, df_train_mean], axis=1)\ndf_train = df_train.astype(float)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T20:27:48.176462Z","iopub.execute_input":"2022-01-06T20:27:48.176815Z","iopub.status.idle":"2022-01-06T20:27:55.218713Z","shell.execute_reply.started":"2022-01-06T20:27:48.176769Z","shell.execute_reply":"2022-01-06T20:27:55.217887Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Predict Success Rate of the Tackle\nFrom the data above, we will predict the success rate of return rejection.  \nFor easy understanding, to block a return is defined as the returner not advancing more than 5 yards in the x-axis direction from the moment.\n* 1 : Returner will stop in 5 yards.\n* 0 : Returner will advance at least 5 yards.\n\nLogistic regression is used for prediction.  \nThe data from October and November 2018 will be used as the teacher data and Descember 2018 will be used as the test data.","metadata":{}},{"cell_type":"code","source":"# 訓練データ、検証データの分割\nfrom sklearn.model_selection import train_test_split\n\ny_train = (df_train['return_yard'] <= 5) * 1\nX_train = df_train.drop(['gameplayId','gameId','playId','frameId','return_yard'], axis=1)\n\nmean_x = X_train.mean()\nstd_x = X_train.std()\n#X_train = (X_train - mean_x) / std_x\n\n# テストデータ、検証データのとりわけ\nX_test = X_train[df_train['gameId'] >= 2018120000]\ny_test = y_train[df_train['gameId'] >= 2018120000]\nX_train = X_train[df_train['gameId'] < 2018120000]\ny_train = y_train[df_train['gameId'] < 2018120000]\n\n# 学習開始\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import confusion_matrix, f1_score\nfrom sklearn.metrics import precision_score, recall_score\n\nlr5 = LogisticRegression(solver='liblinear', penalty='l1')\nlr5.fit(X_train, y_train)\ny_train_pred = lr5.predict(X_train)\ny_test_pred = lr5.predict(X_test)\nprint('F-score(train data): ', f1_score(y_train.values, y_train_pred))\nprint('F-score(test data): ', f1_score(y_test.values, y_test_pred))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T20:27:55.220542Z","iopub.execute_input":"2022-01-06T20:27:55.220847Z","iopub.status.idle":"2022-01-06T20:28:04.024658Z","shell.execute_reply.started":"2022-01-06T20:27:55.220808Z","shell.execute_reply":"2022-01-06T20:28:04.023750Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This model gives us not only a prediction of blocking success / failure, but also the success rate.   \nThe calibration curve below shows the relationship between the predicted probability (x axis) and actual success rate in each bins.","metadata":{}},{"cell_type":"code","source":"# Calibration curves の作図\nfrom sklearn.calibration import calibration_curve\n\nprob = lr5.predict_proba(X_test)[:, 1] # 目的変数が1である確率を予測\nprob_true, prob_pred = calibration_curve(y_true=y_test, y_prob=prob, n_bins=20)\n\nfig, ax1 = plt.subplots()\nax1.plot(prob_pred, prob_true, marker='s', label='calibration plot', color='skyblue') # キャリプレーションプロットを作成\nax1.plot([0, 1], [0, 1], linestyle='--', label='ideal', color='limegreen') # 45度線をプロット\nax1.legend(bbox_to_anchor=(1.12, 1), loc='upper left')\nax1.set_xlabel(\"Mean Predicted Probability\")\nax1.set_ylabel(\"Fraction of positives\")\n#ax2 = ax1.twinx() # 2軸を追加\n#ax2.hist(prob, bins=20, histtype='step', color='orangered') # スコアのヒストグラムも併せてプロット\n#ax2.set_ylabel(\"Counts\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:28:04.026446Z","iopub.execute_input":"2022-01-06T20:28:04.027173Z","iopub.status.idle":"2022-01-06T20:28:04.303263Z","shell.execute_reply.started":"2022-01-06T20:28:04.027085Z","shell.execute_reply":"2022-01-06T20:28:04.302398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Similarly, calculate the probability for each distance of how far the returner can go before being stopped.","metadata":{}},{"cell_type":"code","source":"# リターンの阻止の定義をすこしずつ変えて計算\nyard_list = np.arange(0.5, 15, 0.5)\nmodel_list = []\n\nfor yard_index in yard_list:\n    y_train = (df_train['return_yard'] <= yard_index) * 1\n\n    # テストデータ、検証データのとりわけ\n    y_test = y_train[df_train['gameId'] >= 2018120000]\n    y_train = y_train[df_train['gameId'] < 2018120000]\n\n    # 学習開始\n    lr = LogisticRegression(solver='liblinear', penalty='l1')\n    lr.fit(X_train, y_train)\n    y_train_pred = lr.predict(X_train)\n    y_test_pred = lr.predict(X_test)\n    model_list.append(lr)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:28:04.305303Z","iopub.execute_input":"2022-01-06T20:28:04.305942Z","iopub.status.idle":"2022-01-06T20:32:22.856528Z","shell.execute_reply.started":"2022-01-06T20:28:04.305889Z","shell.execute_reply":"2022-01-06T20:32:22.854629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Output Example\nAs an example, we will predict the distance the returner will advance in a given frame.\nThe horizontal axis represents the distance the returner being stopped and the vertical axis represents the corresponding probability.\nThe figure below shows the player's position. The blue dots represent the attackers, and the orange dots represent the defenders.","metadata":{}},{"cell_type":"code","source":"X_sample = X_test.iloc[240:241,:]\nprob_list = []\nfor lr in model_list:\n    prob = lr.predict_proba(X_sample)[:, 1] # 目的変数が1である確率を予測\n    prob_list.append(prob[0])\nplt.scatter(yard_list, prob_list)\nplt.plot(yard_list, prob_list)\nplt.ylim(0,1)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:32:22.858694Z","iopub.execute_input":"2022-01-06T20:32:22.859279Z","iopub.status.idle":"2022-01-06T20:32:23.177169Z","shell.execute_reply.started":"2022-01-06T20:32:22.859227Z","shell.execute_reply":"2022-01-06T20:32:23.175943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_sample = df_train.iloc[240,:]\ngameId = X_sample['gameId']\nplayId = X_sample['playId']\nframeId = X_sample['frameId']\n\ndf_temp = track_data[(track_data['gameId']==gameId)&(track_data['playId']==playId)&(track_data['frameId']==frameId)]\ndf_temp1 = df_temp[df_temp['team']=='offense']\ndf_temp2 = df_temp[df_temp['team']=='diffense']\n\nplt.scatter(df_temp1['x'], df_temp1['y'], alpha=0.9)\nplt.scatter(df_temp2['x'], df_temp2['y'], alpha=0.9)\nplt.ylim(0,53.3)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-01-06T20:32:23.178873Z","iopub.execute_input":"2022-01-06T20:32:23.179176Z","iopub.status.idle":"2022-01-06T20:32:23.413070Z","shell.execute_reply.started":"2022-01-06T20:32:23.179139Z","shell.execute_reply":"2022-01-06T20:32:23.412192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}