{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [1st and Future - NFL Player Contact Detection](https://www.kaggle.com/competitions/nfl-player-contact-detection)\n\n> NOTE : **This is the inference notebook.**\nIt is used to run and score your trained model against both test vides AND unseen videos that the competition host holds back.\n\n## UC Berkeley W207 - Applied Machine Learning\n> Track player contacts in NFL games to monitor player injury load and health and safety. This project was undertaken as part of W207 - Applied Machine Learning for MIDS (Masters In Data Science) @ UC Berkeley in Spring 2023. Our goal was primarily to study and learn applied machine larning approaches via the design of a new model and CNN using TensorFlow that was based on successful PyTorch approaches such as [zzy's example](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference). \n\n> Thanks to everyone that shared their code and models. We tried to comment and reference all used code.\n\n> We implemented the model and training completely in TensorFlow.\n\nBerkeley W207 Team\n\n* [Mick Dreeling](https://www.linkedin.com/in/mdreeling/)\n* [Shruthi Jaganathan](https://www.linkedin.com/in/shruthi-jaganathan-251188142/)\n* [Saurabh Naurain](https://www.linkedin.com/in/narainsaurabh/)\n\n<img src=\"https://s3.amazonaws.com/nonwebstorage/headstrong/fieldbanner.jpg\">","metadata":{}},{"cell_type":"markdown","source":"# A. Methodology  🎯\n* In this notebook, we use **2.5D** image training on **Video Frames** from each play with `tf.data.DataSet`, using `Tensorflow`.\n* In a nutshell, **2.5D Image Training** is training of a **3D** image like a **2D** Image.  More about **2.5D** training is discussed later. \n* In this notebook, several frames are pulled, per contact instance, from each video play and then 'stacked' on top of each other and fed into the model. The idea being that we want to train the model to recognise what happens BEFORE contact occurs, and AFTER contact occurs.\n* This notebook is compatible for both **GPU** and **TPU**. Device is automatically selected so you won't have to do anything to allocate device.","metadata":{}},{"cell_type":"markdown","source":"# B. Notebooks 📒\n📌 **2.5D-NFL Contact**:\n* Scoring and Inference: [This notebook](https://www.kaggle.com/mdreeling/nfl-player-contact-detection-train-tf-berkeley)\n* Training: [Use this notebook to tune and train the model](https://www.kaggle.com/code/mickles/nfl-for-learning-train-tf-berkeley-w207)\n\n📌 **Data/Dataset**:\n* Contact-Data: [NFL Competition Contact Data](https://www.kaggle.com/competitions/nfl-player-contact-detection/data)\n* [nfl-trained-model-keras](https://www.kaggle.com/datasets/mickles/nfl-trained-model-keras) A pre-trained model that has been trained on only a few plays","metadata":{}},{"cell_type":"markdown","source":"# 2. Import Libraries 📚\nLet's import necessary libraries.","metadata":{}},{"cell_type":"code","source":"import os\nimport sys\nimport glob\nimport numpy as np\nimport pandas as pd\nimport random\nimport math\nimport gc\nimport cv2\nfrom tqdm import tqdm\nimport time\nfrom functools import lru_cache\nimport albumentations as A\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split\n\nimport tensorflow as tf\nfrom tensorflow.keras import datasets, layers, models\nimport matplotlib.pyplot as plt\n\n# For GCP Access to data if necessary\nfrom kaggle_datasets import KaggleDatasets","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-04-20T02:34:27.504811Z","iopub.execute_input":"2023-04-20T02:34:27.505237Z","iopub.status.idle":"2023-04-20T02:34:39.784206Z","shell.execute_reply.started":"2023-04-20T02:34:27.505199Z","shell.execute_reply":"2023-04-20T02:34:39.783012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Configuration ⚙️\n\n* On each run you should change the value of **wandb_run_name**","metadata":{}},{"cell_type":"code","source":"!mkdir -p /kaggle/working/frames\n\nclass CFG:\n    base_data_path=\"/kaggle/input/nfl-player-contact-detection\"\n    ffmpeg_output=\"/kaggle/working/frames\"\n    \n    img_width = 256\n    img_height = 256\n    input_channel = 26\n    output_channel = 1\n    \n    wandb = True\n    competition = \"nfl-player-contact-detection\"\n    _wandb_kernel = \"awsaf49\"\n    debug = False\n    exp_name = \"v4\"\n    wandb_project_name=\"nfl-contact-1\"\n    wandb_run_name= \"mick-kaggle-infer-version-1\"\n\n    # Use verbose=0 for silent, 1 for interactive\n    verbose = 0\n    display_plot = True\n\n    # Device for training\n    device = None  # device is automatically selected\n\n    # Batch Size & Epochs\n    batch_size = 1\n    batch_prefetch = 1\n    epochs = 3\n    \n    # Seeding for reproducibility\n    seed = 101\n    \n    # Output debug info\n    debugging = False","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:39.786571Z","iopub.execute_input":"2023-04-20T02:34:39.787349Z","iopub.status.idle":"2023-04-20T02:34:40.902012Z","shell.execute_reply.started":"2023-04-20T02:34:39.787311Z","shell.execute_reply":"2023-04-20T02:34:40.900435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Reproducibility ♻️\nSets value for random seed to produce similar result in each run.","metadata":{}},{"cell_type":"code","source":"def seeding(SEED):\n    \"\"\"\n    Sets all random seeds for the program (Python, NumPy, and TensorFlow).\n    \"\"\"\n    np.random.seed(SEED)\n    random.seed(SEED)\n    os.environ[\"PYTHONHASHSEED\"] = str(SEED)\n    os.environ[\"TF_CUDNN_DETERMINISTIC\"] = str(SEED)\n    tf.random.set_seed(SEED)\n    print(\"seeding done!!!\")\n\n\nseeding(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:40.904256Z","iopub.execute_input":"2023-04-20T02:34:40.904674Z","iopub.status.idle":"2023-04-20T02:34:40.913408Z","shell.execute_reply.started":"2023-04-20T02:34:40.904627Z","shell.execute_reply":"2023-04-20T02:34:40.911958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Set Up Device 📱\nFollowing codes automatically detects hardware(tpu or gpu or cpu). ","metadata":{}},{"cell_type":"markdown","source":"# 5.5. Physical GPU Setup and Meta Data Setup 📖\n\nThis will set you up for GPU if it is available","metadata":{}},{"cell_type":"code","source":"physical_devices = tf.config.experimental.list_physical_devices('GPU')\nif len(physical_devices) > 0:\n    tf.config.experimental.set_memory_growth(physical_devices[0], True)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:40.917089Z","iopub.execute_input":"2023-04-20T02:34:40.917445Z","iopub.status.idle":"2023-04-20T02:34:40.930624Z","shell.execute_reply.started":"2023-04-20T02:34:40.917392Z","shell.execute_reply":"2023-04-20T02:34:40.929259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.5. Feature Engineering 1\n\nThis is a method to split the contact field into several other useful fields, as it is a composite.","metadata":{}},{"cell_type":"code","source":"def expand_contact_id(df):\n    \"\"\"\n    Splits out contact_id into seperate columns.\n    \"\"\"\n    df[\"game_play\"] = df[\"contact_id\"].str[:12]\n    df[\"step\"] = df[\"contact_id\"].str.split(\"_\").str[-3].astype(\"int\")\n    df[\"nfl_player_id_1\"] = df[\"contact_id\"].str.split(\"_\").str[-2]\n    df[\"nfl_player_id_2\"] = df[\"contact_id\"].str.split(\"_\").str[-1]\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:40.932314Z","iopub.execute_input":"2023-04-20T02:34:40.933189Z","iopub.status.idle":"2023-04-20T02:34:40.941516Z","shell.execute_reply.started":"2023-04-20T02:34:40.933152Z","shell.execute_reply":"2023-04-20T02:34:40.940174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Meta Data 📖\n\nYou can read a lot of information about this data [here](https://www.kaggle.com/competitions/nfl-player-contact-detection/data) but what we will do below is try and give a general overview of the important parts of the data.\n\n* **Training Video Files**\n\nThere are 3 views of each play, and there are 360 plays, for a total of 1020 videos.\n\n\n<img src=\"https://i.ibb.co/rtCCm38/ezgif-2-8540360a8e.gif\" style=\"display:inline;margin:1px;width:25%\">\n<img src=\"https://i.ibb.co/f9Zs5Q0/ezgif-2-71cd4d2d16.gif\" style=\"display:inline;margin:1px;width:25%\">\n<img src=\"https://i.ibb.co/0VJtffG/ezgif-2-2b462a86dd.gif\" style=\"display:inline;margin:1px;width:25%\">\n\nFrom left to right\n\n  * `Sideline` - A view showing the sideline view of the play\n  * `Endzone` - A view showing the view from the endzone\n  * `All 29` - A view showing the entire field (most people did not use this in their training)\n\n* **Training CSV Files**\n    * `train_labels.csv` - These are the labelled contacts for the videos in the /train folder for every player combination. What this means is that for every single frame of video we have an indicator saying whether player 1 has EVERY other player, player 2 has contacted EVERY other player and so on. This is why this file is over 400MB,it is basically a registry of (player_x * num_players) * num_frames + contacted_occured \n\n    * `train_baseline_helmets.csv` - These are baseline helmet detection and assignment boxes for the training and test set. These are useful when predicting contacts. It provides the bounding boxes for all detected helmets. Not all helmets are detected in every frame.\n\n    * `train_player_tracking.csv` -  This is 10 Hz tracking data for each player on the field during the provided plays. What this means is that for every 1/10th of a second, we have the location, acceleration and direction of each player. This is useful for numerous reasons, including figuring out exactly how close players are to each other.\n\n    * `train_video_metadata.csv` - contains timestamps associated with each Sideline and Endzone view for syncing with the player tracking data.","metadata":{}},{"cell_type":"markdown","source":"## 6.1 Load CSV Data 📖","metadata":{}},{"cell_type":"code","source":"base_path = CFG.base_data_path\nos.chdir(base_path)\n\nlabels = expand_contact_id(pd.read_csv(\"sample_submission.csv\"))\nprint(\"Number of labelled contacts : \",len(labels))\ntrain_tracking = pd.read_csv(\"test_player_tracking.csv\")\nprint(\"Number of tracking records : \",len(train_tracking))\ntrain_helmets = pd.read_csv(\"test_baseline_helmets.csv\")\nprint(\"Number of helmet detections : \",len(train_helmets))\ntrain_video_metadata = pd.read_csv(\"test_video_metadata.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:40.943088Z","iopub.execute_input":"2023-04-20T02:34:40.943495Z","iopub.status.idle":"2023-04-20T02:34:41.770266Z","shell.execute_reply.started":"2023-04-20T02:34:40.943445Z","shell.execute_reply":"2023-04-20T02:34:41.769363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import re\ndef count_data_items(filenames):\n    n = [int(re.compile(r\"-([0-9]*)\\.\").search(filename).group(1)) for filename in filenames]\n    return np.sum(n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-20T02:34:41.771349Z","iopub.execute_input":"2023-04-20T02:34:41.772006Z","iopub.status.idle":"2023-04-20T02:34:41.779009Z","shell.execute_reply.started":"2023-04-20T02:34:41.771969Z","shell.execute_reply":"2023-04-20T02:34:41.777410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To run code on **TPU** we need our data to be stored on **Google Cloud Storage**. Hence, we'll be needing **GCS_PATH** of our stored data. Worried about how we will get our data stored on **GCS**? \"Kaggle to the Rescue\" Kaggle provides a **GCS_PATH** for public datasets. Hence we can use it for training our model on **TPU**. Simply we have to use `KaggleDatasets()` to get `GCS_PATH` of our dataset.","metadata":{}},{"cell_type":"markdown","source":"## 6.2 Create Features📖\n\nOne of the first important steps is to create the list of features that will travel into the model with each set of images.\n\nBelow, we create a master set of features by joining data from various CSV files into one dataframe.\n\nFor example, for any specific contact record, we will also have the x and y position, and acceleration ofthe 2 players, all in the same row! yay!","metadata":{}},{"cell_type":"code","source":"# To create the features, we join the labelled data of players (whether players are contacting each \n# other) with the detailed tracking data of the direction and distance that they are travelling\n#\n# Algorithm for creating dataset\n# ================================\n# tr_tracking is the file train_player_tracking.csv which has a single row per player position\n# df is train_labels.csv which contains a row for BOTH players and whether they are in contact with each other or the ground\n\n# The feature dataset is a combination of both files, with the contact and tracking data all flattened into 1 row\n\n# 1. Primary merge column is the step number in the video\n# 2. We merge players into nfl_player_id_1 and nfl_player_id_2 based on the step and game_play that they are both in\n# 3. We use merging to create _2 versions of each players tracking stats and add that to the dataframe\n# 4. We square each x and y player location difference, sum it, and then square it again to creatre a distance column\n\ndef create_features(df, tr_tracking, merge_col=\"step\", use_cols=[\"x_position\", \"y_position\"]):\n    output_cols = []\n    df_combo = (\n        df.astype({\"nfl_player_id_1\": \"str\"})\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", merge_col, \"nfl_player_id\",] + use_cols\n            ],\n            left_on=[\"game_play\", merge_col, \"nfl_player_id_1\"],\n            right_on=[\"game_play\", merge_col, \"nfl_player_id\"],\n            how=\"left\",\n        )\n        .rename(columns={c: c+\"_1\" for c in use_cols})\n        .drop(\"nfl_player_id\", axis=1)\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", merge_col, \"nfl_player_id\"] + use_cols\n            ],\n            left_on=[\"game_play\", merge_col, \"nfl_player_id_2\"],\n            right_on=[\"game_play\", merge_col, \"nfl_player_id\"],\n            how=\"left\",\n        )\n        .drop(\"nfl_player_id\", axis=1)\n        .rename(columns={c: c+\"_2\" for c in use_cols})\n        .sort_values([\"game_play\", merge_col, \"nfl_player_id_1\", \"nfl_player_id_2\"])\n        .reset_index(drop=True)\n    )\n    output_cols += [c+\"_1\" for c in use_cols]\n    output_cols += [c+\"_2\" for c in use_cols]\n    \n    if (\"x_position\" in use_cols) & (\"y_position\" in use_cols):\n        index = df_combo['x_position_2'].notnull()\n        \n        distance_arr = np.full(len(index), np.nan)\n        tmp_distance_arr = np.sqrt(\n            np.square(df_combo.loc[index, \"x_position_1\"] - df_combo.loc[index, \"x_position_2\"])\n            + np.square(df_combo.loc[index, \"y_position_1\"]- df_combo.loc[index, \"y_position_2\"])\n        )\n        \n        distance_arr[index] = tmp_distance_arr\n        df_combo['distance'] = distance_arr\n        output_cols += [\"distance\"]\n        \n    df_combo['G_flug'] = (df_combo['nfl_player_id_2']==\"G\")\n    output_cols += [\"G_flug\"]\n    return df_combo, output_cols\n\n\nuse_cols = [\n    'x_position', 'y_position', 'speed', 'distance',\n    'direction', 'orientation', 'acceleration', 'sa'\n]\n\ntrain, feature_cols = create_features(labels, train_tracking, use_cols=use_cols)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:41.781011Z","iopub.execute_input":"2023-04-20T02:34:41.781622Z","iopub.status.idle":"2023-04-20T02:34:42.054652Z","shell.execute_reply.started":"2023-04-20T02:34:41.781582Z","shell.execute_reply":"2023-04-20T02:34:42.053001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.3 Additional Filtering📖\n\nBelow, we do some additional filtering such as removing records where the distance between players is greater than 2 yards, as it is unlikely they are contacting at that distance. \n\nReference interesting comments on this [here](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference/comments#2071677)\n\nWe also create a 'frame' column that will become useful later. (In the data it does not label activity as frames, but rather as 'steps' of 1/10th of a second).","metadata":{}},{"cell_type":"code","source":"# This section does some additional filtering\n# First it removes any 'distance' values greater than 2 yards\n# The distance is defined as => distance: distance traveled from prior time point, in yards.\n\ntrain_filtered = train.query('not distance>2').reset_index(drop=True)\n\n# We now convert the step value into a 'frame' number in the video and add that as a column\n# 59.94 is average frame rate, 5*59.94 is because step=0 starts\n# from 5s not 0s (The data is labelled starting AFTER the snap)\ntrain_filtered['frame'] = (train_filtered['step']/10*59.94+5*59.94).astype('int')+1\ntrain_filtered.head(100)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:42.056404Z","iopub.execute_input":"2023-04-20T02:34:42.056787Z","iopub.status.idle":"2023-04-20T02:34:42.132011Z","shell.execute_reply.started":"2023-04-20T02:34:42.056739Z","shell.execute_reply":"2023-04-20T02:34:42.130816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.5 - Video Frame Conversion\n\nWe need to convert all of the plays which have been selected into their repsective video frames. For this we use [ffmeg](https://ffmpeg.org/).\n\n> **Note:** This conversion step is expensive. Typically while we were training, we would cache this step locally or on our instance, and then skip it in subsequent runs. Once you have converted the videos to images, that is not something you need to continue to do.\n\nThis step generates a **LOT** of JPG files. The videos themselves are 60 frames per second, so a 14 second play yields 840 files.\n\nIf you convert the entire dataset, this will take several hours, and the output will be roughly 60 Gigbytes.\n\nAn example showing some conversion stats is below.\n\n| Video File | Video File Length | Video Size | No. Converted JPGS | Size of converted JPGS |\n| --- | --- | --- | -- | -- |\n| 58168_003392_Endzone.mp4 | 11 seconds | 5.95 Mb | 711 | 97.1 Mb","metadata":{}},{"cell_type":"code","source":"for video in tqdm(train_helmets.video.unique()):\n    if 'Endzone2' not in video:\n        !ffmpeg -i test/{video} -q:v 2 -f image2 {CFG.ffmpeg_output}/{video}_%04d.jpg -hide_banner -loglevel error","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:34:42.136011Z","iopub.execute_input":"2023-04-20T02:34:42.136362Z","iopub.status.idle":"2023-04-20T02:35:14.329507Z","shell.execute_reply.started":"2023-04-20T02:34:42.136328Z","shell.execute_reply":"2023-04-20T02:35:14.327976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.5 Build Lookup Tables📖\n\nDuring the creation of our `tf.data.DataSet` records, we will some utilities to perform some lookups on video metadata","metadata":{}},{"cell_type":"markdown","source":"### 6.5.1 - Lookup Video to Helmet Bounding Boxes\n\nProvides an easy lookup of all of the helmet bounding boxes associated with a particular video file","metadata":{}},{"cell_type":"code","source":"# This code creates a new lookup dataset called video2helmets which gives you all players that \n# are taking part in a particular play / video\n\nvideo2helmets = {}\ntrain_helmets_new = train_helmets.set_index('video')\nfor video in tqdm(train_helmets.video.unique()):\n    video2helmets[video] = train_helmets_new.loc[video].reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:14.331938Z","iopub.execute_input":"2023-04-20T02:35:14.332494Z","iopub.status.idle":"2023-04-20T02:35:14.380126Z","shell.execute_reply.started":"2023-04-20T02:35:14.332443Z","shell.execute_reply":"2023-04-20T02:35:14.378810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.5.2 - Lookup Video Frames as JPG\n\nProvides an easy lookup of all of the converted JPG files that belong to a particular video","metadata":{}},{"cell_type":"code","source":"# This code creates a new lookup dataset called video2frames \n# which gives the number of frames that per video\n\nvideo2frames = {}\n\nfor game_play in tqdm(train_video_metadata.game_play.unique()):\n    for view in ['Endzone', 'Sideline']:\n        video = game_play + f'_{view}.mp4'\n        video2frames[video] = max(list(map(lambda x:int(x.split('_')[-1].split('.')[0]), \\\n                                           glob.glob(f'{CFG.ffmpeg_output}/{video}*'))))","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:14.382342Z","iopub.execute_input":"2023-04-20T02:35:14.382856Z","iopub.status.idle":"2023-04-20T02:35:14.434030Z","shell.execute_reply.started":"2023-04-20T02:35:14.382805Z","shell.execute_reply":"2023-04-20T02:35:14.432790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Data Augmentation 🌈\n> **Note:** Unlike a traditional 2.5D CNN application (such as taking multiple 'slices' of a brain scan and compositing into multiple images), the approach here is that the slices are actually the frames of the video.\n","metadata":{}},{"cell_type":"markdown","source":"## 7.1. Used Augmentations\n\nSome Augmentations that were used are from the albumentations library. Augmentations (also known as TTA) are used to provide even more information to the model, by providing slightly different versions of the same image.\n\n* HorizontalFlip\n* ShiftScaleRotate\n* RandomBrightnessContrast\n\n<img src=\"https://i.ibb.co/ZgJ6CxY/augmentations.png\">","metadata":{}},{"cell_type":"markdown","source":"### 7.1.1. Augmentation Utility","metadata":{}},{"cell_type":"code","source":"# This is an implementation of TTA\n# https://www.kaggle.com/code/andrewkh/test-time-augmentation-tta-worth-it\n\n# The image is flipped, transposed and contrast adjusted\ntrain_aug = A.Compose([\n    A.HorizontalFlip(p=0.75),\n    A.ShiftScaleRotate(p=0.5),\n    A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.1), contrast_limit=(-0.1, 0.1), p=0.25),\n    A.Normalize(mean=[0.], std=[1.])\n])\n\nvalid_aug = A.Compose([\n    A.Normalize(mean=[0.], std=[1.])\n])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-20T02:35:14.435749Z","iopub.execute_input":"2023-04-20T02:35:14.436097Z","iopub.status.idle":"2023-04-20T02:35:14.444986Z","shell.execute_reply.started":"2023-04-20T02:35:14.436063Z","shell.execute_reply":"2023-04-20T02:35:14.443421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. Data Pipeline 🍚","metadata":{}},{"cell_type":"markdown","source":"## 8.1 - 2.5D Training\n\n**What is 2.5D Training?**\n\nTypically when people talk about 2.5D training, they are talking about feeding multiple 'similar' images into a model, which when stacked on top of one another provide the model a much better idea of the situation that we are trying to predict, so for instance, in the example with the GI Tract scan below, we overlay multple 'slices' of the scan, to create a composite image that the model can use to undertstand what healthy organs look like. More detail on this is available [here](https://www.kaggle.com/code/awsaf49/uwmgi-transunet-2-5d-train-tf), which is the notebook on which the style and narrative of our own notebook is based.\n\n<div align=center><img src=\"https://i.ibb.co/KKtZ7Gn/Picture1-3d.png\" width=500></div>\n\nNormally, when you feed color image data into a model, you feed it a single image, with 3 RGB 'channels'. These RGB channels are combined in order to create the correct color image.\n\nA 2.5DCNN modifies this approach, by instead utilising the channels to feed 'slices' of a similar image, so instead of receiving the same image 3 times, but with slightly different color profiles, it receives 3 images, all of which are slightly different. So you are basically 'building up' multiple views of an image 'as if' you were trying to create it in 3D, but with less information.\n\nWhy do we do this? Because training 3D CNN's is very compute intensive, requires a lot of data and is quite complex. You will already see when training this particular model that it takes a long time, so if we had even more data to try and represent the space in 3D, it would start to get very complex. There is also more reference material and models in 2D readily available for our use.\n\nThis method has some cool advantages over 3D training for instance,\n* Low GPU/memory cost.\n* Simpler pipeline.\n* Easier augmentation.\n* Quick inference.\n* Many open-source models.\n\n## 8.2 - 2.5D for this competition - Stacking Video Frames\n\nUsing a 2.5D CNN for this competition creates a little bit more complexity. \n\nWhat we are dealing with here are hundreds of videos of American Football plays, where we already have all of the contacts labelled, so we know to a fairly high level of accuracy when contacts occur in the video. Our job is to feed 'what contact between players looks like' into the model.\n\nBelow you can see a play starting ('the snap'), and there is contact between at least 2 players within a few frames. This contact will have already been labelled for us in `train_labels.csv`.\n\n<div align=center><img src=\"https://i.ibb.co/CH9QS3S/ezgif-1-cda7bef7dd.gif\" width=500></div>\n\nWhat this notebook does, is a little different than what we have seen in other 2.5D CNN examples. Normally, we would do as we see above in the 'GI Tract' example, which is to virtually stack different slices of the same image on top of each other.\n\nInstead what we do is to feed frames into the model which occur 'around the same time' as the point of contact. This means that for each contact that we are training on, we are using quite a few images. In the case of this notebook, we look back before the point of contact and then after the contact to see if we can feed data to the model so that it understands more clearly what exactly contact looks like. We will use many more frames/images that what are shown below, but it gives you an idea of what we are trying to do.\n\n<div align=center><img src=\"https://i.ibb.co/gtCCfXq/Stacked-labelled-process.png\" width=500></div>\n\n> **Note:** It should be noted that while in the image above we are showing the stacking of just the sideline view (for illustration purposes), frames from the endzone view are also stacked when the model is trained. This allows the model to be able to see the contact occur from **2 angles**.","metadata":{}},{"cell_type":"markdown","source":"## 8.3 - Building our pipeline with **tf.data.DataSet**\n\nYou will notice from most other notebooks for this competition, they tend to use `torch.utils.data.Dataset` which has a batching mechanism that allows us to dripfeed our massive image dataset into our model during training without the GPU and System running out of memory. As we used TensorFlow instead of PyTorch, we are using `tf.data.DataSet` which is a suitable equiqalent.\n\nBasically, we feed batches of 'frame windows' into the model along with the features for a particular contact. So for a single contact between two players, we might provide 11 images,5 preceding the contact,the contact itself, and 5 afterwards. We go into more detail on features in subsequent sections.\n\n<div align=center> <img src=\"https://i.ibb.co/4YCMzTN/tf-data-new.png\" width=900></div>\n\nTo build our data pipeline, we need to use `tf.data` API\n\nCheckout this [doc](https://www.tensorflow.org/guide/data) if you want to learn more about `tf.data`.\n\n### 8.3.1 - Creating records for **tf.data.DataSet**\n\nEach `getitem()` call to the `TFDataset` object returns\n\n* **26** 256x256 pixel JPGS\n    * 13 x (256x256) pixel **EndZone** JPGS + 13 x (256x256) pixel **Sideline** JPGS\n* **18** float values representing the features of each contact record\n\n**Packaging Images and Features**\n\nThis section is probably the most interesting piece of this codebase, as it explains exactly how you present data on which the model is going to be trained. Tweaking elements of this code can have dramatic effects on how well your model recnogises contact, and for good reason.\n\nWe have commented the code in detail below, and so you should be able to understand what is happening, but you can also switch `debugging=true` in the config section and it will print out what it is doing.\n\nWe will briefly describe the  key steps towards creating a record to be fed to the model\n\n* First off remember that **for every single record** in the filtered version of `train_labels.csv`, you are going to be creating a set of `tf.Tensor` objects that represent a quite a bit of data on that record, whether it is an actual contact or not.\n    * i.e if record 1 represents player 1 and player 7 NOT contacting (but they are actually less than 2 yards from each other), you will be passing in a group of sideline and endzone JPG files representing frames **before** and **after** the contact, even when if didn't even happen (contact=0).\n    * This obviously is to help the model figure out what is *not* a contact, as well as what *is*.\n\nAs mentioned above, the features for that record (the speed, acceleration etc) of each player, are also returned as a `tf.Tensor`\n\n**Record creation pseudocode**\n\n* For each frame in the play\n    * Find frame number\n    * Identify which two players are potentially contacting in this record\n        * Set frame window look-ahead and look-back (default 24 frames)\n            * For EndZone\n                * Zoom in and crop on player 1's helmet for every 4th frame of 24 (yielding 12 jpgs)\n            * For SideLine\n                * Zoom in and crop on player 1's helmet for every 4th frame of 24 (yielding 12 jpgs)\n    * Attach all zoomed in images as 256x256 float arrays\n    * Attach 18 features as float array\n    * Attach label (contact or no contact)\n    * Augment image\n    * return record\n    \n<div align=center> <img src=\"https://i.ibb.co/xYXRpSs/Screenshot-2023-04-13-165423.png\" width=500></div>\n<h6 align=\"center\">Image records returned from getitem()</h6> \n\n\n**Detailed commented code is below**","metadata":{}},{"cell_type":"code","source":"class TFDataset():\n    def __init__(self, df, aug=train_aug, mode='train'):\n        self.df = df\n        self.frame = df.frame.values\n        self.feature = df[feature_cols].fillna(-1).values\n        self.players = df[['nfl_player_id_1','nfl_player_id_2']].values\n        self.game_play = df.game_play.values\n        self.aug = aug\n        self.mode = mode\n    \n    def __len__(self):\n        return len(self.df)\n    \n    # @lru_cache(1024)\n    # def read_img(self, path):\n    #     return cv2.imread(path, 0)\n   \n    def getitem(self):\n         # 1. Get a list of labelled contacts for this play\n        allIndexes = np.arange(len(self.df))\n        #np.random.default_rng(seed).shuffle(allIndexes)\n        if CFG.debugging:\n            print(\"Processing %i labelled contacts\"  % (len(self.df)))\n        # 2. Loop through all contacts for every player-player and player-ground combination\n        for idx in allIndexes:\n\n          # 3. Set the look-back and the look-forward from this frame-record that you are going to use to train the model\n          window = 24\n\n          # 4. Pull the first frame.\n          # NOTE : The data is labelled starting right when the 'snap' occurs\n          # The snap occurs at 5 seconds into every video meaning that the lowest frame number will always be 300\n          # NOTE : Each 'step' in our data is 0.1s or 10hz. In other words, our tracking and labelling data is not as granular as our frame data\n          # This means that we will be first pulling frame 300, then 306, then 312 etc as we do not have tracking or labelling data for the in-between frames\n          frame = self.frame[idx]\n\n          if CFG.debugging:\n            print(\"     Working on frame number %i (Index %s) \"  % (frame, idx))\n\n          # 5. If we are training the model, randomize the first pulled frame to one step forward or back (+/- 6 frames)\n          if self.mode == 'train':\n              frame = frame + random.randint(-6, 6)\n              if CFG.debugging:\n                print(\"     Frame randomly shifted to %i (Index %s) \"  % (frame, idx))\n\n          # 6. Add the two players in this contact to an array\n          players = []\n          for p in self.players[idx]:\n              if CFG.debugging:\n                  print(\"       Adding player %s\"  % (p))\n\n              if p == 'G':\n                  players.append(p)\n              else:\n                  players.append(int(p))\n          \n          imgs = []\n\n          # 7. Process this frame for both the EndZone and Sideline videos of this gameplay\n          # Here we will now start pulling frames within the frame window defined above\n          for view in ['Endzone', 'Sideline']:\n              video = self.game_play[idx] + f'_{view}.mp4'\n\n              # 8. Retrieve a mapping of all the helmets that are present in this entire video\n              tmp = video2helmets[video]\n              # 9. Shrink the available frame-set down to be only the relevant frames\n              # (+24 and -24 frames from the currently processed frame)\n              tmp[tmp['frame'].between(frame-window, frame+window)]\n              # 10. Now shrink the frame-set again to only include frames in which\n              # The first NFL player of interest is present\n              tmp = tmp[tmp.nfl_player_id.isin(players)]\n              tmp_frames = tmp.frame.values\n              # 11. Now, group our remaining frames, and get the mean value of the bounding box dimensions\n              tmp = tmp.groupby('frame')[['left','width','top','height']].mean()\n  #0.002s\n\n              bboxes = []\n              # 12. Iterate through every frame in our frame window\n              for f in range(frame-window, frame+window+1, 1):\n                  if f in tmp_frames:\n                      # 13. Locate the frame in our frame list, and return the bounding boxes for player 1's helmet\n                      x, w, y, h = tmp.loc[f][['left','width','top','height']]\n                      bboxes.append([x, w, y, h])\n                  else:\n                      bboxes.append([np.nan, np.nan, np.nan, np.nan])\n              # 14. We may not have found bounding boxes for all frames, so interpolate the ones that were missing\n              # using a standard interpolation method\n              bboxes = pd.DataFrame(bboxes).interpolate(limit_direction='both').values\n              # 15. Of all collected bounding boxes, take every 4th one\n              bboxes = bboxes[::4]\n\n              # 16. Check if our remaining bounding boxes were valid\n              if bboxes.sum() > 0:\n                  flag = 1\n              else:\n                  flag = 0\n  #0.03s\n              frame_sampled = False\n\n              # 17. Now we will be doing the same iteration as above but now reading these frames from disk\n              # and this time we will be reading every 4th frame. This is why we took every 4th bounding box above.\n              for i, f in enumerate(range(frame-window, frame+window+1, 4)):\n                  if CFG.debugging:\n                    print(\"        Pulling from %s frame window %d to %d (at iteration %d)\" % (view ,frame-window,frame+window+1,i))\n                  img_new = np.zeros((256, 256), dtype=np.float32)\n\n                  # 18. Read this frame if it had a valid player 1 helmet bounding box and is within the frame limit of the video\n                  if flag == 1 and f <= video2frames[video]:\n                      img = cv2.imread(f'{CFG.ffmpeg_output}/{video}_{f:04d}.jpg', 0)\n                      # 19. Get the bounding box for this frame, and use it to zoom in around the two players\n                      x, w, y, h = bboxes[i]\n                      img = img[int(y+h/2)-128:int(y+h/2)+128,int(x+w/2)-128:int(x+w/2)+128].copy()\n                      # 20. Create a new image from the zoomed in version\n                      img_new[:img.shape[0], :img.shape[1]] = img\n                  # 21. Add the zoomed in version to a running list of images\n                  imgs.append(img_new)\n  #0.06s\n          # 22. Cast all of the features for this record to float32 and store them\n          feature = np.float32(self.feature[idx])\n          if CFG.debugging:\n            print(\"        Pulled %s frame images in total into this record\" % (len(imgs)))\n          # 23. Transpose the image list and store to the record\n          img = np.array(imgs).transpose(1, 2, 0)\n          #plt.imshow(img[:,:,5])\n          if CFG.debugging:\n            print(\"        Augmenting images...\")\n          img = self.aug(image=img)[\"image\"]\n          record_label = np.float32(self.df.contact.values[idx])\n          if CFG.debugging:\n            print(\"        Adding label %i \" % record_label)\n\n          yield (img, feature), record_label","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:14.446669Z","iopub.execute_input":"2023-04-20T02:35:14.447061Z","iopub.status.idle":"2023-04-20T02:35:14.475063Z","shell.execute_reply.started":"2023-04-20T02:35:14.447027Z","shell.execute_reply":"2023-04-20T02:35:14.473560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8.5. Create generator function \n\nOur generator function is what actually feeds the training process.\n\n> **Note:** It is important to not that this is where we set up our `tf.Tensor` shapes to be fed in during training. The **CFG.input_channel value** here is very very important and is governed by the `window` value in our `TFDataSet` class. By changing the look-back/look-forward window, you are also changing the number of channels that are input in your model because it changes the number of frames that fall under the window, and input_frames = channels in this notebook.","metadata":{}},{"cell_type":"code","source":"tfTrainData = TFDataset(train_filtered, train_aug, 'test')\n\ndataset = tf.data.Dataset.from_generator(\n    tfTrainData.getitem,\n    args=[],\n    output_signature=((\n        tf.TensorSpec(shape=(256,256,CFG.input_channel), dtype=tf.float32),\n        tf.TensorSpec(shape=(18,), dtype=tf.float32)),\n        tf.TensorSpec(shape=(), dtype=tf.uint8)))","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:14.477103Z","iopub.execute_input":"2023-04-20T02:35:14.477598Z","iopub.status.idle":"2023-04-20T02:35:14.643591Z","shell.execute_reply.started":"2023-04-20T02:35:14.477545Z","shell.execute_reply":"2023-04-20T02:35:14.642518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9. Visualization 🔭\nTo ensure our pipeline is generating **images** and **augmentations** correctly, we'll read a sample from the training set.","metadata":{}},{"cell_type":"code","source":"seti = tfTrainData.getitem()\nprint(seti)\n\nimport itertools\nimg_feature, label = next(itertools.islice(tfTrainData.getitem(), 1))\n\nfig, ax_list = plt.subplots(4,6, figsize=(15, 15))\n\ni = 0\n\nfor ax in ax_list.flatten():\n    ax.imshow(img_feature[0][:,:,i])\n    i = i + 1\n\nfor ax in ax_list.flatten():\n    ax.set_xticks([])\n    ax.set_yticks([])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:14.645393Z","iopub.execute_input":"2023-04-20T02:35:14.645842Z","iopub.status.idle":"2023-04-20T02:35:16.437906Z","shell.execute_reply.started":"2023-04-20T02:35:14.645801Z","shell.execute_reply":"2023-04-20T02:35:16.436418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10.1 Imports","metadata":{}},{"cell_type":"code","source":"# import the necessary packages\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import BatchNormalization\nfrom tensorflow.keras.layers import Conv2D\nfrom tensorflow.keras.layers import MaxPooling2D\nfrom tensorflow.keras.layers import Activation\nfrom tensorflow.keras.layers import Dropout\nfrom tensorflow.keras.layers import Dense\nfrom tensorflow.keras.layers import Flatten\nfrom tensorflow.keras.layers import Input\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.layers import concatenate\nfrom tensorflow.keras.applications import ResNet50V2\nfrom tensorflow.keras.layers import Dropout, MaxPooling2D, Conv2D, Conv2DTranspose, UpSampling2D, Dense\nfrom tensorflow.keras.layers import Input, concatenate\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.regularizers import l2\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.metrics import MeanIoU\nfrom tensorflow.keras import backend as K\nfrom tensorflow.keras.layers import BatchNormalization\nfrom tensorflow.keras.layers import Dropout, MaxPooling2D, Conv2D, AveragePooling2D, LayerNormalization\nfrom tensorflow.keras.layers import Conv2DTranspose, UpSampling2D, GlobalAveragePooling2D\nfrom tensorflow.keras.layers import LeakyReLU, ReLU","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:16.439472Z","iopub.execute_input":"2023-04-20T02:35:16.439860Z","iopub.status.idle":"2023-04-20T02:35:16.454290Z","shell.execute_reply.started":"2023-04-20T02:35:16.439821Z","shell.execute_reply":"2023-04-20T02:35:16.452491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10.4. Load Model","metadata":{}},{"cell_type":"code","source":"from tensorflow import keras\nmodel = keras.models.load_model('/kaggle/input/nfl-trained-model-keras/models')","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:16.455706Z","iopub.execute_input":"2023-04-20T02:35:16.456178Z","iopub.status.idle":"2023-04-20T02:35:18.624670Z","shell.execute_reply.started":"2023-04-20T02:35:16.456134Z","shell.execute_reply":"2023-04-20T02:35:18.623417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10.4.1 Evaluate the model\n\nEvaluate our model on the test video dataset","metadata":{}},{"cell_type":"code","source":"# Evaluate the model on the test data using `evaluate`\nprint(\"Evaluate on test data\")\nresults = model.evaluate(dataset.batch(CFG.batch_size).prefetch(CFG.batch_prefetch))\nprint(\"test loss, test acc:\", results)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:35:18.629290Z","iopub.execute_input":"2023-04-20T02:35:18.629659Z","iopub.status.idle":"2023-04-20T02:42:47.264247Z","shell.execute_reply.started":"2023-04-20T02:35:18.629621Z","shell.execute_reply":"2023-04-20T02:42:47.256887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10.4.2 Make predictions\n\nFor the test video dataset, make predictions","metadata":{}},{"cell_type":"code","source":"# Generate predictions (probabilities -- the output of the last layer)\nprint(\"Generate predictions..\")\npredictions = model.predict(dataset.batch(CFG.batch_size).prefetch(CFG.batch_prefetch))\nprint(\"predictions shape:\", predictions.shape)\n\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:42:47.265651Z","iopub.status.idle":"2023-04-20T02:42:47.266813Z","shell.execute_reply.started":"2023-04-20T02:42:47.266505Z","shell.execute_reply":"2023-04-20T02:42:47.266539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = np.array(predictions)\nprint(predictions )","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:42:47.268190Z","iopub.status.idle":"2023-04-20T02:42:47.268628Z","shell.execute_reply.started":"2023-04-20T02:42:47.268415Z","shell.execute_reply":"2023-04-20T02:42:47.268446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Clean up and remove the frames from the output so that we can actually find and submit the submission.csv\n!rm -fr /kaggle/working/frames","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10.4.3 Prepare submission\n\nAdd our predictions to the submission file","metadata":{}},{"cell_type":"code","source":"th = 0.29\n\ntrain_filtered['contact'] = (predictions >= th).astype('int')\n\nsub = pd.read_csv('/kaggle/input/nfl-player-contact-detection/sample_submission.csv')\n\nsub = sub.drop(\"contact\", axis=1).merge(train_filtered[['contact_id', 'contact']], how='left', on='contact_id')\nsub['contact'] = sub['contact'].fillna(0).astype('int')\n\nsub[[\"contact_id\", \"contact\"]].to_csv(\"/kaggle/working/submission.csv\", index=False)\n\nprint(len(sub))\nsub.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T02:42:47.270565Z","iopub.status.idle":"2023-04-20T02:42:47.270999Z","shell.execute_reply.started":"2023-04-20T02:42:47.270797Z","shell.execute_reply":"2023-04-20T02:42:47.270820Z"},"trusted":true},"execution_count":null,"outputs":[]}]}