{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Use MediaPipe to Render Parquet Files\n\nThis competition uses data derived from MediaPipe to extract landmark poses from video.  For each frame the landmarks have been stored in a parquet file.  In this notebook, I show how to extract the landmarks from an image using mediapipe and then I demonstrate how to render a parquet file from the competition dataset using mediapipe.  Which might be useful if you try to generate your own training data.","metadata":{}},{"cell_type":"code","source":"!pip install mediapipe --quiet","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-16T20:02:16.462675Z","iopub.execute_input":"2023-03-16T20:02:16.463065Z","iopub.status.idle":"2023-03-16T20:02:28.010002Z","shell.execute_reply.started":"2023-03-16T20:02:16.463030Z","shell.execute_reply":"2023-03-16T20:02:28.008468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import cv2\n\nimport numpy as np\nimport pandas as pd\nfrom path import Path\nfrom PIL import Image\nfrom fastai.vision.all import show_image\n\nfrom ipywidgets import interact, interactive, fixed, interact_manual\nimport ipywidgets as widgets\n\nimport mediapipe as mp\nfrom mediapipe.framework.formats import landmark_pb2\nmp_drawing = mp.solutions.drawing_utils\nmp_drawing_styles = mp.solutions.drawing_styles\nmp_holistic = mp.solutions.holistic","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:28.013138Z","iopub.execute_input":"2023-03-16T20:02:28.013705Z","iopub.status.idle":"2023-03-16T20:02:28.022154Z","shell.execute_reply.started":"2023-03-16T20:02:28.013652Z","shell.execute_reply":"2023-03-16T20:02:28.020846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Mediapipe Landmark Detection\n\nHere I'll give a demonstration of downloading a sample image from the Internet and running mediapipe on and image and then demonstrate how to render the landmark results interactively using mediapipe to do the rendering.","metadata":{}},{"cell_type":"code","source":"# Image credit: https://www.pexels.com/photo/woman-in-black-blouse-showing-sign-language-gesture-10029275/\n!curl \"https://images.pexels.com/photos/10029275/pexels-photo-10029275.jpeg?cs=srgb&dl=pexels-rodnae-productions-10029275.jpg&fm=jpg&w=640&h=960\" > /kaggle/working/sign.jpg","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:28.024020Z","iopub.execute_input":"2023-03-16T20:02:28.024625Z","iopub.status.idle":"2023-03-16T20:02:29.439802Z","shell.execute_reply.started":"2023-03-16T20:02:28.024575Z","shell.execute_reply":"2023-03-16T20:02:29.438323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load and show the sample image\nim = Image.open('/kaggle/working/sign.jpg')\nshow_image(im,figsize=(6,17))","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:29.443876Z","iopub.execute_input":"2023-03-16T20:02:29.444313Z","iopub.status.idle":"2023-03-16T20:02:29.889305Z","shell.execute_reply.started":"2023-03-16T20:02:29.444269Z","shell.execute_reply":"2023-03-16T20:02:29.887671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Renering Pose Landmarks using MediaPipe\nThis function will take a set of landmarks as returned from the mediapipe model and render the landmarks onto an image in the form of a numpy array.","metadata":{}},{"cell_type":"code","source":"def draw_landmarks(landmarks,image,show_pose=True,show_face_contour=True,show_face_tesselation=True,show_left_hand=True,show_right_hand=True):\n    annotated_image = image.copy()\n    results = landmarks\n    if show_face_tesselation:\n        mp_drawing.draw_landmarks(\n            annotated_image,\n            results.face_landmarks,\n            mp_holistic.FACEMESH_TESSELATION,\n            landmark_drawing_spec=None,\n            connection_drawing_spec=mp_drawing_styles\n            .get_default_face_mesh_tesselation_style())\n    if show_face_contour:\n        mp_drawing.draw_landmarks(\n            annotated_image,\n            results.face_landmarks,\n            mp_holistic.FACEMESH_CONTOURS,\n            landmark_drawing_spec=None,\n            connection_drawing_spec=mp_drawing_styles\n            .get_default_face_mesh_contours_style())\n    if show_pose:\n        mp_drawing.draw_landmarks(\n            annotated_image,\n            results.pose_landmarks,\n            mp_holistic.POSE_CONNECTIONS,\n            landmark_drawing_spec=mp_drawing_styles.\n            get_default_pose_landmarks_style())\n    if show_left_hand:\n        mp_drawing.draw_landmarks(\n            annotated_image,\n            results.left_hand_landmarks,\n            mp_holistic.HAND_CONNECTIONS,\n            landmark_drawing_spec=mp_drawing_styles\n            .get_default_hand_landmarks_style())\n    if show_right_hand:\n        mp_drawing.draw_landmarks(\n            annotated_image,\n            results.right_hand_landmarks,\n            mp_holistic.HAND_CONNECTIONS,\n            landmark_drawing_spec=mp_drawing_styles\n            .get_default_hand_landmarks_style())\n    return annotated_image","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:29.891525Z","iopub.execute_input":"2023-03-16T20:02:29.893028Z","iopub.status.idle":"2023-03-16T20:02:29.903128Z","shell.execute_reply.started":"2023-03-16T20:02:29.892952Z","shell.execute_reply":"2023-03-16T20:02:29.902099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Using MediaPipe to Predict Pose Landmarks\nThe following function take an image in the form of a PIL image, runs MediaPipe on it it to detect the landmark features for face, pose, left hand and right hand and then renders the landmarks using the draw_landmarks function.","metadata":{}},{"cell_type":"code","source":"def static_image(image,show_pose=True,show_face_contour=True,show_face_tesselation=True,show_left_hand=True,show_right_hand=True,show_image=True):\n    BG_COLOR = (192, 192, 192) # gray\n    with mp_holistic.Holistic(\n        static_image_mode=True,\n        model_complexity=2,\n        enable_segmentation=True,\n        refine_face_landmarks=True) as holistic:\n            image = np.array(image)\n            image_height, image_width, _ = image.shape\n            # Convert the BGR image to RGB before processing.\n            #results = holistic.process(cv2.cvtColor(image, cv2.COLOR_BGR2RGB))\n            results = holistic.process(image)\n\n            if show_image:\n                annotated_image = image.copy()\n            else:\n                annotated_image = np.zeros(image.shape,dtype=np.uint8)\n\n            annotated_image = draw_landmarks(results,annotated_image,show_pose,show_face_contour,show_face_tesselation,show_left_hand,show_right_hand)\n    return Image.fromarray(annotated_image)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:29.904839Z","iopub.execute_input":"2023-03-16T20:02:29.906289Z","iopub.status.idle":"2023-03-16T20:02:29.922021Z","shell.execute_reply.started":"2023-03-16T20:02:29.906172Z","shell.execute_reply":"2023-03-16T20:02:29.920422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Interactive Viewer Using MediaPipe to Render the Pose Landmarks\nYou can use the controls to toggle on and off rendering of the different landmarks recognized by mediapipe.","metadata":{}},{"cell_type":"code","source":"def f(show_pose,show_face_contour,show_face_tesselation,show_left_hand,show_right_hand,show_img):\n    show_image(static_image(im,show_pose,show_face_contour,show_face_tesselation,show_left_hand,show_right_hand,show_img),figsize=(6,17))\n    \ni = interact(f,show_pose=True,show_face_contour=True,show_face_tesselation=True,show_left_hand=True,show_right_hand=True,show_img=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:29.924031Z","iopub.execute_input":"2023-03-16T20:02:29.925051Z","iopub.status.idle":"2023-03-16T20:02:30.917474Z","shell.execute_reply.started":"2023-03-16T20:02:29.924997Z","shell.execute_reply":"2023-03-16T20:02:30.916091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Using MediaPipe to Render Parquet Files","metadata":{"execution":{"iopub.status.busy":"2023-03-16T19:06:13.762178Z","iopub.execute_input":"2023-03-16T19:06:13.762636Z","iopub.status.idle":"2023-03-16T19:06:13.768161Z","shell.execute_reply.started":"2023-03-16T19:06:13.762596Z","shell.execute_reply":"2023-03-16T19:06:13.766094Z"}}},{"cell_type":"code","source":"# load up the train csv\npath = Path('../input/asl-signs')","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:30.919168Z","iopub.execute_input":"2023-03-16T20:02:30.919518Z","iopub.status.idle":"2023-03-16T20:02:30.924487Z","shell.execute_reply.started":"2023-03-16T20:02:30.919479Z","shell.execute_reply":"2023-03-16T20:02:30.923118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The train.csv file contains one row for each training sample.  The sign column show the name of the asl sign represented in the training samples and the path columns identifies a parquet file.  ","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(path/'train.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:30.926476Z","iopub.execute_input":"2023-03-16T20:02:30.926846Z","iopub.status.idle":"2023-03-16T20:02:31.073904Z","shell.execute_reply.started":"2023-03-16T20:02:30.926811Z","shell.execute_reply":"2023-03-16T20:02:31.072060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each parquet file contains some number of frames of landmark data for each sign. We'll grab a sample parquet file and take a look at it.","metadata":{}},{"cell_type":"code","source":"parquet_file = df.iloc[0].path\nprint(f'Sample parquet file: {parquet_file}')\npf = pd.read_parquet(path/parquet_file)\npf.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.078971Z","iopub.execute_input":"2023-03-16T20:02:31.079953Z","iopub.status.idle":"2023-03-16T20:02:31.118047Z","shell.execute_reply.started":"2023-03-16T20:02:31.079896Z","shell.execute_reply":"2023-03-16T20:02:31.116567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each parquet file contains some number of frames of landmark data which would have been generated using mediapipe from some original video of a person making the sign.  The video data is not available in the training dataset.","metadata":{}},{"cell_type":"code","source":"print('This parquet file contains the following frames:')\nprint(pf.frame.unique())","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.119949Z","iopub.execute_input":"2023-03-16T20:02:31.120472Z","iopub.status.idle":"2023-03-16T20:02:31.128955Z","shell.execute_reply.started":"2023-03-16T20:02:31.120420Z","shell.execute_reply":"2023-03-16T20:02:31.127513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each frame contains the following types of landmarks","metadata":{}},{"cell_type":"code","source":"print('Frame types: ')\nprint(pf.type.unique())","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.131295Z","iopub.execute_input":"2023-03-16T20:02:31.131869Z","iopub.status.idle":"2023-03-16T20:02:31.143299Z","shell.execute_reply.started":"2023-03-16T20:02:31.131819Z","shell.execute_reply":"2023-03-16T20:02:31.141748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each frame has the same number of landmark points (x,y,z).  Landmarks are points in 3D space. But beware some of these are missing (NANs)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T19:18:46.689177Z","iopub.execute_input":"2023-03-16T19:18:46.689595Z","iopub.status.idle":"2023-03-16T19:18:46.696692Z","shell.execute_reply.started":"2023-03-16T19:18:46.689561Z","shell.execute_reply":"2023-03-16T19:18:46.695067Z"}}},{"cell_type":"code","source":"print('Total number of features per frame: ')\npf.shape[0]/pf.frame.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.145946Z","iopub.execute_input":"2023-03-16T20:02:31.146439Z","iopub.status.idle":"2023-03-16T20:02:31.156586Z","shell.execute_reply.started":"2023-03-16T20:02:31.146391Z","shell.execute_reply":"2023-03-16T20:02:31.155408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The features for each frame are broken out as follows for each type.","metadata":{}},{"cell_type":"code","source":"print('Number of features per frame per type: ')\npf.type.value_counts()/pf.frame.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.157822Z","iopub.execute_input":"2023-03-16T20:02:31.158279Z","iopub.status.idle":"2023-03-16T20:02:31.174627Z","shell.execute_reply.started":"2023-03-16T20:02:31.158208Z","shell.execute_reply":"2023-03-16T20:02:31.173122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can use mediapipe to render our annotations on top of a background image.  Since we don't have the images for the video frames we'll just create a black background image to render on top of (all zeros)","metadata":{}},{"cell_type":"code","source":"# the size is arbitrary the landmark locations will be rendered relative to the dimensions of the background image provided\nannotated_image = np.zeros((1024,1024,3),dtype=np.uint8)\nshow_image(annotated_image)  # show empty blackground image for reference","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.175908Z","iopub.execute_input":"2023-03-16T20:02:31.176363Z","iopub.status.idle":"2023-03-16T20:02:31.541708Z","shell.execute_reply.started":"2023-03-16T20:02:31.176308Z","shell.execute_reply":"2023-03-16T20:02:31.539458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following function will take a parquet file and a frame number and will create a \"landmarks\" object that is consistent with the prediction results returned back from mediapipe so that we can use the same landmark rendering function from above.","metadata":{}},{"cell_type":"code","source":"# simple holder for our landmarks that will mimic the results that we get back from mediapipe but using our data\nclass Landmarks(object):\n    pass\n\ndef get_landmarks_from_parquet(pf,frame):\n    f = pf[pf.frame == frame]\n    face = landmark_pb2.NormalizedLandmarkList()\n    for t in f[f.type=='face'][['x','y','z']].itertuples(index=False):\n        face.landmark.add(x=t.x,y=t.y,z=t.z)\n    pose = landmark_pb2.NormalizedLandmarkList()\n    for t in f[f.type=='pose'][['x','y','z']].itertuples(index=False):\n        pose.landmark.add(x=t.x,y=t.y,z=t.z)\n    left_hand = landmark_pb2.NormalizedLandmarkList()\n    for t in f[f.type=='left_hand'][['x','y','z']].itertuples(index=False):\n        left_hand.landmark.add(x=t.x,y=t.y,z=t.z)\n    right_hand = landmark_pb2.NormalizedLandmarkList()\n    for t in f[f.type=='right_hand'][['x','y','z']].itertuples(index=False):\n        right_hand.landmark.add(x=t.x,y=t.y,z=t.z)    \n    result = Landmarks()\n    result.face_landmarks = face\n    result.pose_landmarks = pose\n    result.left_hand_landmarks = left_hand\n    result.right_hand_landmarks = right_hand\n    return result\n    ","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.543068Z","iopub.execute_input":"2023-03-16T20:02:31.543430Z","iopub.status.idle":"2023-03-16T20:02:31.555056Z","shell.execute_reply.started":"2023-03-16T20:02:31.543391Z","shell.execute_reply":"2023-03-16T20:02:31.554019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Take our sample parquet file from earlier translate it to a set of landmarks that mediapipe can use and then render them using mediapipe.","metadata":{}},{"cell_type":"code","source":"landmarks = get_landmarks_from_parquet(pf,20)\nshow_image(draw_landmarks(landmarks,annotated_image))","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.556384Z","iopub.execute_input":"2023-03-16T20:02:31.557669Z","iopub.status.idle":"2023-03-16T20:02:31.933933Z","shell.execute_reply.started":"2023-03-16T20:02:31.557627Z","shell.execute_reply":"2023-03-16T20:02:31.932457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Interactive MediaPipe Viewer for a Parquet File\nyou can drag the slider interactively to scrub through the landmark frames","metadata":{"execution":{"iopub.status.busy":"2023-03-16T18:30:04.587124Z","iopub.execute_input":"2023-03-16T18:30:04.588486Z","iopub.status.idle":"2023-03-16T18:30:04.595654Z","shell.execute_reply.started":"2023-03-16T18:30:04.588412Z","shell.execute_reply":"2023-03-16T18:30:04.594186Z"}}},{"cell_type":"code","source":"frames = pf.frame.unique()  # get the frames from the parquet file\n\ndef show_frame(frame):\n    landmarks = get_landmarks_from_parquet(pf,frames[frame])\n    show_image(draw_landmarks(landmarks,annotated_image),figsize=(9,9),title=f'frame: {frames[frame]} [{frame+1} of {len(frames)}]')\n    #print(f'showing frame: {frames[frame]}')\n    \ni = interact(show_frame,frame=widgets.IntSlider(min=0, max=len(frames)-1, step=1, value=0))","metadata":{"execution":{"iopub.status.busy":"2023-03-16T20:02:31.935599Z","iopub.execute_input":"2023-03-16T20:02:31.936057Z","iopub.status.idle":"2023-03-16T20:02:32.343888Z","shell.execute_reply.started":"2023-03-16T20:02:31.936019Z","shell.execute_reply":"2023-03-16T20:02:32.342628Z"},"trusted":true},"execution_count":null,"outputs":[]}]}