{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# NFL Player Contact Detection - MultiShot MultiPose Estimation using YOLOv8 & Mediapipe","metadata":{}},{"cell_type":"markdown","source":"# Problem Statement","metadata":{}},{"cell_type":"markdown","source":"## Goal of the Competition\nThe goal of this competition is to detect external contact experienced by players during an NFL football game. You will use video and player tracking data to identify moments with contact to help improve player safety.\n​\n## Context\nThe National Football League (NFL) has teamed up with Amazon Web Services (AWS) to strengthen its commitment to predict player injuries. The NFL aspires to have the best injury surveillance and mitigation program in any sport. With your machine learning and computer vision skills, you can help the NFL accurately identify when players experience contact throughout a football play.\n​\nIn prior years, the NFL challenged the Kaggle community to create helmet impact detection and identification algorithms. This year the NFL looks to automatically identify all moments when players experience contact. This competition will be successful if we can reliably detect moments when players are in contact with one another and when a player’s body is in contact with the ground.\n​\nCurrently, the NFL uses its tracking system to monitor a large number of statistics about players’ load during the season. The league has a solution that predicts contact between players, but it only leverages the player tracking data. This competition hopes to improve the predictive power by including video in addition to tracking data. Categorizing ground contact will also provide a more comprehensive view of impacts, improving analysis for player health and safety.\n​\nMore accurate data is an important step toward the NFL’s injury surveillance and mitigation goals. With complete contact detection, the league can identify correlations between certain types of contact and injury, a contributor to future prevention. Your efforts could help mitigate unsafe situations to reduce injury to all players.\n​\nThe National Football League is America's most popular sports league. Founded in 1920, the NFL developed the model for the successful modern sports league and is committed to advancing progress in the diagnosis, prevention, and treatment of sports-related injuries. This competition is part of the Digital Athlete, a joint effort between the NFL and AWS to build a virtual, 360-degree representation of an NFL player’s experience. The Digital Athlete hopes to generate a precise picture of what they need when it comes to preventing and recovering from injuries while performing at their best. Health and safety efforts include support for independent medical research and engineering advancements as well as a commitment to work to better protect players and make the game safer, including enhancements to medical protocols and improvements to how our game is taught and played. For more information about the NFL's health and safety efforts, please visit the [NFL Player Health and Safety website](https://www.nfl.com/playerhealthandsafety/).\n​\n## Evaluation\nSubmissions are evaluated on [Matthews Correlation Coefficient](https://en.wikipedia.org/wiki/Phi_coefficient) between the predicted and actual contact events.\n​\n​\n$$MCC = \\frac{TP * TN - FP * FN}{\\sqrt{(TP + FP)(TP + FN)(TN + FP)(TN + FN)}}$$\n​\nFor every allowable `1contact_id` (formed by concatenating the `game_play_step_player1_player2`), you must predict whether the involved players are in contact at that moment in time. sample_submission.csv provides the exhaustive list of `contact_ids`. Note that the ground, denoted as player G, is included as a possible contact in place of player2. The player with the lower id is always listed first in the `contact_id`.\n​\nThe file should contain a header and have the following format:\n​\n```\ncontact_id,contact\n58168_003392_0_38590_43854,0\n58168_003392_0_38590_41257,1\n58168_003392_0_38590_41944,0\netc.\n```","metadata":{}},{"cell_type":"markdown","source":"# Import dependencies","metadata":{}},{"cell_type":"code","source":"# install mediapipe\n!pip install mediapipe\n!pip install ultralytics","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-28T09:31:13.094501Z","iopub.execute_input":"2023-01-28T09:31:13.095225Z","iopub.status.idle":"2023-01-28T09:31:51.33595Z","shell.execute_reply.started":"2023-01-28T09:31:13.095114Z","shell.execute_reply":"2023-01-28T09:31:51.334749Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import ultralytics\nultralytics.checks()","metadata":{"execution":{"iopub.status.busy":"2023-01-28T09:31:51.338922Z","iopub.execute_input":"2023-01-28T09:31:51.339365Z","iopub.status.idle":"2023-01-28T09:31:53.942687Z","shell.execute_reply.started":"2023-01-28T09:31:51.339324Z","shell.execute_reply":"2023-01-28T09:31:53.941346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import dependencies\nimport os\nimport subprocess\nimport IPython\nfrom IPython.display import Video, display\n\nimport numpy as np\nimport pandas as pd\n\nimport cv2\nimport mediapipe as  mp","metadata":{"execution":{"iopub.status.busy":"2023-01-28T09:31:53.944666Z","iopub.execute_input":"2023-01-28T09:31:53.945429Z","iopub.status.idle":"2023-01-28T09:31:54.112823Z","shell.execute_reply.started":"2023-01-28T09:31:53.94538Z","shell.execute_reply":"2023-01-28T09:31:54.111491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# YOLOv8 + Mediapipe MultiShot MultiPose Estimation","metadata":{}},{"cell_type":"markdown","source":"## Approach\n- We use YOLOv8 object detection to first find the bounding boxes of all the persons in the video.\n- We use the bounding boxes on to indiviudally crop each frame to include only one person.\n- We run MediaPipe single pose estimation on the cropped frame to detect the pose estimations of the person.\n- We carry out this single pose estimation for each cropped bounding box in each frame, and finally annotate the video with the estimated pose landmarks.","metadata":{}},{"cell_type":"code","source":"def play_video(video_path: str):\n    frac = 0.65 # scaling factor for display \n    display(\n        Video(data=video_path, embed=True, height=int(720*frac), width=int(1280*frac))\n    )","metadata":{"execution":{"iopub.status.busy":"2023-01-28T09:31:54.11967Z","iopub.execute_input":"2023-01-28T09:31:54.120347Z","iopub.status.idle":"2023-01-28T09:31:54.13076Z","shell.execute_reply.started":"2023-01-28T09:31:54.120277Z","shell.execute_reply":"2023-01-28T09:31:54.129621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def detect_player_bbox(video_path:str) -> str:\n    \"\"\"\n    Annotates video with object detection bounding boxes using YOLOv8 model.\n    \"\"\"\n    \n    video_name = video_path.split('/')[-1]\n    folder_name = 'predict'\n    \n    # YOLOv8 Object Detection\n    # you can edit the model if you want to use a different yolo8 model\n    os.system(f\"yolo predict model=yolov8x.pt source={video_path} save_txt=True\")\n    \n    # path to latest output \n    max = 0\n    for dir_name in os.listdir('/kaggle/working/runs/detect'):\n        if dir_name!='predict':\n            if int(dir_name[-1])>max:\n                max = int(dir_name[-1])\n    if max>0:\n        folder_name += str(max)\n    output_path = f'/kaggle/working/runs/detect/{folder_name}/labels'\n    \n    return output_path","metadata":{"execution":{"iopub.status.busy":"2023-01-28T09:31:54.135832Z","iopub.execute_input":"2023-01-28T09:31:54.138539Z","iopub.status.idle":"2023-01-28T09:31:54.149082Z","shell.execute_reply.started":"2023-01-28T09:31:54.138501Z","shell.execute_reply":"2023-01-28T09:31:54.148028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def bbox_labels_to_dataframe(bbox_labels_path: str) -> pd.DataFrame:\n\n    bbox_labels = {\n        'video_name':[],\n        'frame':[],\n        'class_id':[],\n        'center_x':[],\n        'center_y':[],\n        'width':[],\n        'height':[]\n    }\n\n    for filename in os.listdir(bbox_labels_path):\n        video_name = \"_\".join(filename.split('_')[0:3]) + '.mp4'\n        frame = filename.split('_')[-1]\n        frame = int(frame.split('.')[0])\n\n        with open(bbox_labels_path + '/' + filename, 'r') as f:\n            for line in f:\n                line = line.split(\" \")\n                class_id = int(line[0])\n                center_x = float(line[1])\n                center_y = float(line[2])\n                width = float(line[3])\n                height = float(line[4])\n\n                if class_id == 0: # if person\n                    # append to dict\n                    bbox_labels['video_name'].append(video_name)\n                    bbox_labels['frame'].append(frame)\n                    bbox_labels['class_id'].append(class_id)\n                    bbox_labels['center_x'].append(center_x)\n                    bbox_labels['center_y'].append(center_y)\n                    bbox_labels['width'].append(width)\n                    bbox_labels['height'].append(height)\n                    \n    \n    return pd.DataFrame(bbox_labels)","metadata":{"execution":{"iopub.status.busy":"2023-01-28T09:31:54.154806Z","iopub.execute_input":"2023-01-28T09:31:54.157384Z","iopub.status.idle":"2023-01-28T09:31:54.174565Z","shell.execute_reply.started":"2023-01-28T09:31:54.157327Z","shell.execute_reply":"2023-01-28T09:31:54.17337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def multi_pose_estimation(video_path:str, bbox_labels: pd.DataFrame, verbose=True) -> str:\n    \"\"\"\n    Performs multi-shot multi-pose estimation by obtianing person bbox from YOLOv8\n    and performing single pose estimation on the bbox crop.\n    \"\"\"\n    \n    # intializing mediapipe utils\n    mp_drawing = mp.solutions.drawing_utils\n    mp_drawing_styles = mp.solutions.drawing_styles\n    mp_pose = mp.solutions.pose\n    \n    # video name\n    video_name = video_path.split('/')[-1]\n\n    # VideoCapture Object\n    cap = cv2.VideoCapture(video_path)\n    \n    # video variables\n    width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))\n    height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))\n    fps = cap.get(cv2.CAP_PROP_FPS)\n    total_frames = bbox_labels['frame'].max()\n    \n    # VideoWriter Object\n    output_path  = \"labeled_\" + video_name\n    tmp_output_path = 'tmp_' + output_path\n    out = cv2.VideoWriter(tmp_output_path, cv2.VideoWriter_fourcc(*'MP4V'), fps, (width, height))\n\n    # check if camera opened successfully\n    if (cap.isOpened()==False):\n        print('Error opening video file!')\n\n    # multipose estimation\n    with mp_pose.Pose(\n        min_detection_confidence = 0.8,\n        min_tracking_confidence = 0.8) as pose:\n        frame = 1\n        while (cap.isOpened()):\n            success, image = cap.read()\n            if success:\n                # selecting the frame\n                bbox_set = bbox_labels.query('frame==@frame')\n                # iterating through bboxs in the frame\n                for idx, annot in bbox_set.iterrows():\n                    bbox_center_x = annot['center_x'] * width\n                    bbox_center_y = annot['center_y'] * height\n                    bbox_width = annot['width'] * width\n                    bbox_height = annot['height'] * height\n\n                    # finding top-left and bottom-right bbox cooridnates\n                    bbox_top_left_x = int(bbox_center_x - (bbox_width/2))\n                    bbox_top_left_y = int(bbox_center_y - (bbox_height/2))\n                    bbox_bottom_right_x = int(bbox_center_x + (bbox_width/2))\n                    bbox_bottom_right_y = int(bbox_center_y + (bbox_height/2))\n\n                    # cropping image to bbox\n                    image_crop = image[bbox_top_left_y:bbox_bottom_right_y, bbox_top_left_x:bbox_bottom_right_x]\n\n                    # pose estimation\n                    # set image as not writeable to improve perfromance\n                    image_crop.flags.writeable = False\n                    image_crop = cv2.cvtColor(image_crop, cv2.COLOR_BGR2RGB)\n                    results = pose.process(image_crop)\n\n                    # transposing results to be drawn on the original image\n                    if results.pose_landmarks != None:\n                        for landmark in results.pose_landmarks.landmark:\n                            landmark.x = ((abs(bbox_bottom_right_x - bbox_top_left_x) / width) * landmark.x) + (bbox_top_left_x/width)\n                            landmark.y = ((abs(bbox_bottom_right_y - bbox_top_left_y) / height) * landmark.y) + (bbox_top_left_y/height)\n\n                        # draw the pose annotations on the image\n                        # set image as writeable\n                        image.flags.writeable = True\n                        mp_drawing.draw_landmarks(\n                            image,\n                            results.pose_landmarks,\n                            mp_pose.POSE_CONNECTIONS,\n                            landmark_drawing_spec = mp_drawing_styles.get_default_pose_landmarks_style())\n\n\n\n                # save video\n                out.write(image)\n                if verbose:\n                    print(f'Frame: {frame}/{total_frames}')\n                frame += 1\n            else:\n                break\n\n\n        cap.release()\n        out.release()\n        \n\n    # Not all browsers support the codec, we will re-load the file at tmp_output_path\n    # and convert to a codec that is more broadly readable using ffmpeg\n    if os.path.exists(output_path):\n        os.remove(output_path)\n    subprocess.run(\n            [\n                \"ffmpeg\",\n                \"-i\",\n                tmp_output_path,\n                \"-crf\",\n                \"18\",\n                \"-preset\",\n                \"veryfast\",\n                \"-hide_banner\",\n                \"-loglevel\",\n                \"error\",\n                \"-vcodec\",\n                \"libx264\",\n                output_path,\n            ]\n        )\n    os.remove(tmp_output_path)\n    \n    return output_path","metadata":{"execution":{"iopub.status.busy":"2023-01-28T09:49:41.903097Z","iopub.execute_input":"2023-01-28T09:49:41.903476Z","iopub.status.idle":"2023-01-28T09:49:41.923484Z","shell.execute_reply.started":"2023-01-28T09:49:41.903446Z","shell.execute_reply":"2023-01-28T09:49:41.922287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"video_path = '/kaggle/input/nfl-player-contact-detection/train/58168_003392_Endzone.mp4'\nbbox_labels_path = detect_player_bbox(video_path)\nbbox_labels = bbox_labels_to_dataframe(bbox_labels_path)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-28T09:31:54.214414Z","iopub.execute_input":"2023-01-28T09:31:54.217103Z","iopub.status.idle":"2023-01-28T09:33:03.040849Z","shell.execute_reply.started":"2023-01-28T09:31:54.217042Z","shell.execute_reply":"2023-01-28T09:33:03.039821Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output_path = multi_pose_estimation(video_path, bbox_labels)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-28T09:49:49.797583Z","iopub.execute_input":"2023-01-28T09:49:49.797982Z","iopub.status.idle":"2023-01-28T09:54:37.322905Z","shell.execute_reply.started":"2023-01-28T09:49:49.797934Z","shell.execute_reply":"2023-01-28T09:54:37.321591Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(output_path)","metadata":{"execution":{"iopub.status.busy":"2023-01-28T09:54:37.863938Z","iopub.execute_input":"2023-01-28T09:54:37.865156Z","iopub.status.idle":"2023-01-28T09:54:39.252401Z","shell.execute_reply.started":"2023-01-28T09:54:37.865117Z","shell.execute_reply":"2023-01-28T09:54:39.251281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# References\n- https://www.kaggle.com/code/dariussingh/player-contact-detection-eda\n- https://www.kaggle.com/code/dariussingh/nfl-yolov8-object-detection-and-segmentation\n- https://github.com/ultralytics/ultralytics\n- https://google.github.io/mediapipe/solutions/pose","metadata":{}}]}