{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"The goal of this competition is to detect external contact experienced by players during an NFL football game. I will use video and player tracking data to identify moments with contact to help improve player safety.\n\nThe National Football League (NFL) has teamed up with Amazon Web Services (AWS) to strengthen its commitment to predict player injuries. The NFL aspires to have the best injury surveillance and mitigation program in any sport. With your machine learning and computer vision skills, you can help the NFL accurately identify when players experience contact throughout a football play.\n\nIn prior years, the NFL challenged the Kaggle community to create helmet impact detection and identification algorithms. This year the NFL looks to automatically identify all moments when players experience contact. This competition will be successful if we can reliably detect moments when players are in contact with one another and when a player’s body is in contact with the ground.\n\nCurrently, the NFL uses its tracking system to monitor a large number of statistics about players’ load during the season. The league has a solution that predicts contact between players, but it only leverages the player tracking data. This competition hopes to improve the predictive power by including video in addition to tracking data. Categorizing ground contact will also provide a more comprehensive view of impacts, improving analysis for player health and safety.\n\nMore accurate data is an important step toward the NFL’s injury surveillance and mitigation goals. With complete contact detection, the league can identify correlations between certain types of contact and injury, a contributor to future prevention. My efforts could help mitigate unsafe situations to reduce injury to all players.\n\nThe National Football League is America's most popular sports league. Founded in 1920, the NFL developed the model for the successful modern sports league and is committed to advancing progress in the diagnosis, prevention, and treatment of sports-related injuries. This competition is part of the Digital Athlete, a joint effort between the NFL and AWS to build a virtual, 360-degree representation of an NFL player’s experience. The Digital Athlete hopes to generate a precise picture of what they need when it comes to preventing and recovering from injuries while performing at their best. Health and safety efforts include support for independent medical research and engineering advancements as well as a commitment to work to better protect players and make the game safer, including enhancements to medical protocols and improvements to how our game is taught and played. \n\nFor more information about the NFL's health and safety efforts, please visit the NFL Player Health and Safety website.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:50:08.697996Z","iopub.execute_input":"2023-02-20T02:50:08.698619Z","iopub.status.idle":"2023-02-20T02:50:09.299878Z","shell.execute_reply.started":"2023-02-20T02:50:08.698525Z","shell.execute_reply":"2023-02-20T02:50:09.298809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport sys\nsys.path.append('/kaggle/input/timm-0-6-9/pytorch-image-models-master')\nimport glob\nimport random\nimport math\nimport gc\nimport cv2\nfrom tqdm import tqdm\nimport time\nfrom functools import lru_cache\nimport torch\nfrom torch import nn\nfrom torch.nn import functional as F\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.cuda.amp import autocast, GradScaler\nimport timm\nimport albumentations as A\nfrom albumentations.pytorch import ToTensorV2\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import matthews_corrcoef","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:50:09.301836Z","iopub.execute_input":"2023-02-20T02:50:09.302469Z","iopub.status.idle":"2023-02-20T02:50:15.522278Z","shell.execute_reply.started":"2023-02-20T02:50:09.302432Z","shell.execute_reply":"2023-02-20T02:50:15.521222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CFG = {\n    'seed': 42,\n    'model': 'resnet50',\n    'img_size': 256,\n    'epochs': 10,\n    'train_bs': 100, \n    'valid_bs': 64,\n    'lr': 1e-3, \n    'weight_decay': 1e-6,\n    'num_workers': 2\n}","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:50:15.523943Z","iopub.execute_input":"2023-02-20T02:50:15.524322Z","iopub.status.idle":"2023-02-20T02:50:15.531436Z","shell.execute_reply.started":"2023-02-20T02:50:15.524285Z","shell.execute_reply":"2023-02-20T02:50:15.529561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.cuda.manual_seed_all(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n\nseed_everything(CFG['seed'])\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:50:15.534294Z","iopub.execute_input":"2023-02-20T02:50:15.535425Z","iopub.status.idle":"2023-02-20T02:50:15.616872Z","shell.execute_reply.started":"2023-02-20T02:50:15.535386Z","shell.execute_reply":"2023-02-20T02:50:15.615883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def expand_contact_id(df):\n    \"\"\"\n    Splits out contact_id into seperate columns.\n    \"\"\"\n    df[\"game_play\"] = df[\"contact_id\"].str[:12]\n    df[\"step\"] = df[\"contact_id\"].str.split(\"_\").str[-3].astype(\"int\")\n    df[\"nfl_player_id_1\"] = df[\"contact_id\"].str.split(\"_\").str[-2]\n    df[\"nfl_player_id_2\"] = df[\"contact_id\"].str.split(\"_\").str[-1]\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:50:15.619159Z","iopub.execute_input":"2023-02-20T02:50:15.619921Z","iopub.status.idle":"2023-02-20T02:50:15.627047Z","shell.execute_reply.started":"2023-02-20T02:50:15.619871Z","shell.execute_reply":"2023-02-20T02:50:15.626018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = expand_contact_id(pd.read_csv(\"/kaggle/input/nfl-player-contact-detection/sample_submission.csv\"))\n\ntest_tracking = pd.read_csv(\"/kaggle/input/nfl-player-contact-detection/test_player_tracking.csv\")\n\ntest_helmets = pd.read_csv(\"/kaggle/input/nfl-player-contact-detection/test_baseline_helmets.csv\")\n\ntest_video_metadata = pd.read_csv(\"/kaggle/input/nfl-player-contact-detection/test_video_metadata.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:50:15.629647Z","iopub.execute_input":"2023-02-20T02:50:15.630300Z","iopub.status.idle":"2023-02-20T02:50:16.375925Z","shell.execute_reply.started":"2023-02-20T02:50:15.630263Z","shell.execute_reply":"2023-02-20T02:50:16.374948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir -p ../work/frames\n\nfor video in tqdm(test_helmets.video.unique()):\n    if 'Endzone2' not in video:\n        !ffmpeg -i /kaggle/input/nfl-player-contact-detection/test/{video} -q:v 2 -f image2 /kaggle/work/frames/{video}_%04d.jpg -hide_banner -loglevel error","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:50:16.377240Z","iopub.execute_input":"2023-02-20T02:50:16.377609Z","iopub.status.idle":"2023-02-20T02:51:03.999915Z","shell.execute_reply.started":"2023-02-20T02:50:16.377575Z","shell.execute_reply":"2023-02-20T02:51:03.998649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def create_features(df, tr_tracking, merge_col=\"step\", use_cols=[\"x_position\", \"y_position\"]):\n    output_cols = []\n    df_combo = (\n        df.astype({\"nfl_player_id_1\": \"str\"})\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", merge_col, \"nfl_player_id\",] + use_cols\n            ],\n            left_on=[\"game_play\", merge_col, \"nfl_player_id_1\"],\n            right_on=[\"game_play\", merge_col, \"nfl_player_id\"],\n            how=\"left\",\n        )\n        .rename(columns={c: c+\"_1\" for c in use_cols})\n        .drop(\"nfl_player_id\", axis=1)\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", merge_col, \"nfl_player_id\"] + use_cols\n            ],\n            left_on=[\"game_play\", merge_col, \"nfl_player_id_2\"],\n            right_on=[\"game_play\", merge_col, \"nfl_player_id\"],\n            how=\"left\",\n        )\n        .drop(\"nfl_player_id\", axis=1)\n        .rename(columns={c: c+\"_2\" for c in use_cols})\n        .sort_values([\"game_play\", merge_col, \"nfl_player_id_1\", \"nfl_player_id_2\"])\n        .reset_index(drop=True)\n    )\n    output_cols += [c+\"_1\" for c in use_cols]\n    output_cols += [c+\"_2\" for c in use_cols]\n    \n    if (\"x_position\" in use_cols) & (\"y_position\" in use_cols):\n        index = df_combo['x_position_2'].notnull()\n        \n        distance_arr = np.full(len(index), np.nan)\n        tmp_distance_arr = np.sqrt(\n            np.square(df_combo.loc[index, \"x_position_1\"] - df_combo.loc[index, \"x_position_2\"])\n            + np.square(df_combo.loc[index, \"y_position_1\"]- df_combo.loc[index, \"y_position_2\"])\n        )\n        \n        distance_arr[index] = tmp_distance_arr\n        df_combo['distance'] = distance_arr\n        output_cols += [\"distance\"]\n        \n    df_combo['G_flug'] = (df_combo['nfl_player_id_2']==\"G\")\n    output_cols += [\"G_flug\"]\n    return df_combo, output_cols","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.001741Z","iopub.execute_input":"2023-02-20T02:51:04.002048Z","iopub.status.idle":"2023-02-20T02:51:04.014797Z","shell.execute_reply.started":"2023-02-20T02:51:04.002020Z","shell.execute_reply":"2023-02-20T02:51:04.013407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"use_cols = [\n    'x_position', 'y_position', 'speed', 'distance',\n    'direction', 'orientation', 'acceleration', 'sa'\n]\n\ntest, feature_cols = create_features(labels, test_tracking, use_cols=use_cols)\ntest","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.016717Z","iopub.execute_input":"2023-02-20T02:51:04.017244Z","iopub.status.idle":"2023-02-20T02:51:04.312181Z","shell.execute_reply.started":"2023-02-20T02:51:04.017208Z","shell.execute_reply":"2023-02-20T02:51:04.311047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_filtered = test.query('not distance>2').reset_index(drop=True)\ntest_filtered['frame'] = (test_filtered['step']/10*59.94+5*59.94).astype('int')+1\ntest_filtered","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.316614Z","iopub.execute_input":"2023-02-20T02:51:04.316993Z","iopub.status.idle":"2023-02-20T02:51:04.371142Z","shell.execute_reply.started":"2023-02-20T02:51:04.316961Z","shell.execute_reply":"2023-02-20T02:51:04.370149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del test, labels, test_tracking\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.372845Z","iopub.execute_input":"2023-02-20T02:51:04.373264Z","iopub.status.idle":"2023-02-20T02:51:04.565213Z","shell.execute_reply.started":"2023-02-20T02:51:04.373225Z","shell.execute_reply":"2023-02-20T02:51:04.564069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_aug = A.Compose([\n    A.HorizontalFlip(p=0.5),\n    A.ShiftScaleRotate(p=0.5),\n    A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.1), contrast_limit=(-0.1, 0.1), p=0.5),\n    A.Normalize(mean=[0.], std=[1.]),\n    ToTensorV2()\n])\n\nvalid_aug = A.Compose([\n    A.Normalize(mean=[0.], std=[1.]),\n    ToTensorV2()\n])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.566886Z","iopub.execute_input":"2023-02-20T02:51:04.567329Z","iopub.status.idle":"2023-02-20T02:51:04.577775Z","shell.execute_reply.started":"2023-02-20T02:51:04.567289Z","shell.execute_reply":"2023-02-20T02:51:04.576563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"video2helmets = {}\ntest_helmets_new = test_helmets.set_index('video')\nfor video in tqdm(test_helmets.video.unique()):\n    video2helmets[video] = test_helmets_new.loc[video].reset_index(drop=True)\n    \ndel test_helmets, test_helmets_new\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.579079Z","iopub.execute_input":"2023-02-20T02:51:04.580290Z","iopub.status.idle":"2023-02-20T02:51:04.764448Z","shell.execute_reply.started":"2023-02-20T02:51:04.580231Z","shell.execute_reply":"2023-02-20T02:51:04.763383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"video2frames = {}\n\nfor game_play in tqdm(test_video_metadata.game_play.unique()):\n    for view in ['Endzone', 'Sideline']:\n        video = game_play + f'_{view}.mp4'\n        video2frames[video] = max(list(map(lambda x:int(x.split('_')[-1].split('.')[0]), \\\n                                           glob.glob(f'/kaggle/work/frames/{video}*'))))","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.765861Z","iopub.execute_input":"2023-02-20T02:51:04.766446Z","iopub.status.idle":"2023-02-20T02:51:04.810633Z","shell.execute_reply.started":"2023-02-20T02:51:04.766409Z","shell.execute_reply":"2023-02-20T02:51:04.809527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class MyDataset(Dataset):\n    def __init__(self, df, aug=valid_aug, mode='train'):\n        self.df = df\n        self.frame = df.frame.values\n        self.feature = df[feature_cols].fillna(-1).values\n        self.players = df[['nfl_player_id_1','nfl_player_id_2']].values\n        self.game_play = df.game_play.values\n        self.aug = aug\n        self.mode = mode\n        \n    def __len__(self):\n        return len(self.df)\n    \n    # @lru_cache(1024)\n    # def read_img(self, path):\n    #     return cv2.imread(path, 0)\n   \n    def __getitem__(self, idx):   \n        window = 24\n        frame = self.frame[idx]\n        \n        if self.mode == 'train':\n            frame = frame + random.randint(-6, 6)\n\n        players = []\n        for p in self.players[idx]:\n            if p == 'G':\n                players.append(p)\n            else:\n                players.append(int(p))\n        \n        imgs = []\n        for view in ['Endzone', 'Sideline']:\n            video = self.game_play[idx] + f'_{view}.mp4'\n\n            tmp = video2helmets[video]\n#             tmp = tmp.query('@frame-@window<=frame<=@frame+@window')\n            tmp[tmp['frame'].between(frame-window, frame+window)]\n            tmp = tmp[tmp.nfl_player_id.isin(players)]#.sort_values(['nfl_player_id', 'frame'])\n            tmp_frames = tmp.frame.values\n            tmp = tmp.groupby('frame')[['left','width','top','height']].mean()\n#0.002s\n\n            bboxes = []\n            for f in range(frame-window, frame+window+1, 1):\n                if f in tmp_frames:\n                    x, w, y, h = tmp.loc[f][['left','width','top','height']]\n                    bboxes.append([x, w, y, h])\n                else:\n                    bboxes.append([np.nan, np.nan, np.nan, np.nan])\n            bboxes = pd.DataFrame(bboxes).interpolate(limit_direction='both').values\n            bboxes = bboxes[::4]\n\n            if bboxes.sum() > 0:\n                flag = 1\n            else:\n                flag = 0\n#0.03s\n                    \n            for i, f in enumerate(range(frame-window, frame+window+1, 4)):\n                img_new = np.zeros((256, 256), dtype=np.float32)\n\n                if flag == 1 and f <= video2frames[video]:\n                    img = cv2.imread(f'/kaggle/work/frames/{video}_{f:04d}.jpg', 0)\n\n                    x, w, y, h = bboxes[i]\n\n                    img = img[int(y+h/2)-128:int(y+h/2)+128,int(x+w/2)-128:int(x+w/2)+128].copy()\n                    img_new[:img.shape[0], :img.shape[1]] = img\n                    \n                imgs.append(img_new)\n#0.06s\n                \n        feature = np.float32(self.feature[idx])\n\n        img = np.array(imgs).transpose(1, 2, 0)    \n        img = self.aug(image=img)[\"image\"]\n        label = np.float32(self.df.contact.values[idx])\n\n        return img, feature, label","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.812308Z","iopub.execute_input":"2023-02-20T02:51:04.812708Z","iopub.status.idle":"2023-02-20T02:51:04.830033Z","shell.execute_reply.started":"2023-02-20T02:51:04.812670Z","shell.execute_reply":"2023-02-20T02:51:04.828940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img, feature, label = MyDataset(test_filtered, valid_aug, 'test')[0]\nplt.imshow(img.permute(1,2,0)[:,:,7])\nplt.show()\nimg.shape, feature, label","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:04.831929Z","iopub.execute_input":"2023-02-20T02:51:04.832585Z","iopub.status.idle":"2023-02-20T02:51:05.293861Z","shell.execute_reply.started":"2023-02-20T02:51:04.832548Z","shell.execute_reply":"2023-02-20T02:51:05.292833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.backbone = timm.create_model(CFG['model'], pretrained=False, num_classes=500, in_chans=13)\n        self.mlp = nn.Sequential(\n            nn.Linear(18, 64),\n            nn.LayerNorm(64),\n            nn.ReLU(),\n            nn.Dropout(0.2),\n            # nn.Linear(64, 64),\n            # nn.LayerNorm(64),\n            # nn.ReLU(),\n            # nn.Dropout(0.2)\n        )\n        self.fc = nn.Linear(64+500*2, 1)\n\n    def forward(self, img, feature):\n        b, c, h, w = img.shape\n        img = img.reshape(b*2, c//2, h, w)\n        img = self.backbone(img).reshape(b, -1)\n        feature = self.mlp(feature)\n        y = self.fc(torch.cat([img, feature], dim=1))\n        return y","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:05.295334Z","iopub.execute_input":"2023-02-20T02:51:05.296462Z","iopub.status.idle":"2023-02-20T02:51:05.308254Z","shell.execute_reply.started":"2023-02-20T02:51:05.296421Z","shell.execute_reply":"2023-02-20T02:51:05.307317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_set = MyDataset(test_filtered, valid_aug, 'test')\ntest_loader = DataLoader(test_set, batch_size=CFG['valid_bs'], shuffle=False, num_workers=CFG['num_workers'], pin_memory=True)\n\nmodel = Model().to(device)\nmodel.load_state_dict(torch.load('/kaggle/input/nfl-exp1/resnet50_fold0.pt'))\n\nmodel.eval()\n    \ny_pred = []\nwith torch.no_grad():\n    tk = tqdm(test_loader, total=len(test_loader))\n    for step, batch in enumerate(tk):\n        if(step % 4 != 3):\n            img, feature, label = [x.to(device) for x in batch]\n            output1 = model(img, feature).squeeze(-1)\n            output2 = model(img.flip(-1), feature).squeeze(-1)\n            \n            y_pred.extend(0.2*(output1.sigmoid().cpu().numpy()) + 0.8*(output2.sigmoid().cpu().numpy()))\n        else:\n            img, feature, label = [x.to(device) for x in batch]\n            output = model(img.flip(-1), feature).squeeze(-1)\n            y_pred.extend(output.sigmoid().cpu().numpy())    \n\ny_pred = np.array(y_pred)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:51:05.309960Z","iopub.execute_input":"2023-02-20T02:51:05.310710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"th = 0.29\n\ntest_filtered['contact'] = (y_pred >= th).astype('int')\n\nsub = pd.read_csv('/kaggle/input/nfl-player-contact-detection/sample_submission.csv')\n\nsub = sub.drop(\"contact\", axis=1).merge(test_filtered[['contact_id', 'contact']], how='left', on='contact_id')\nsub['contact'] = sub['contact'].fillna(0).astype('int')\n\nsub[[\"contact_id\", \"contact\"]].to_csv(\"submission.csv\", index=False)\n\nsub.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this competition, my task is to predict moments of contact between player pairs, as well as when players make non-foot contact with the ground using game footage and tracking data. Each play has four associated videos. Two videos, showing a sideline and endzone view, are time synced and aligned with each other. Additionally, an All29 view is provided but not guaranteed to be time synced. The training set videos are in train/ with corresponding labels in train_labels.csv, while the videos for which you must predict are in the test/ folder.\n\nThis year we are also providing baseline helmet detection and assignment boxes for the training and test set. train_baseline_helmets.csv is the output from last year's winning player assignment model.","metadata":{}}]}