{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Mediapipe Hands - How landmark data is generated?","metadata":{}},{"cell_type":"markdown","source":"This notebook is a modified version of the Colab notebook given here so that it can be viewed on Kaggle.<br>\nhttps://colab.research.google.com/drive/1FvH5eTiZqayZBOHZsFm-i7D-JvoB9DVz#scrollTo=nW2TjFyhLvVH<br>\nhttps://github.com/google/mediapipe","metadata":{}},{"cell_type":"markdown","source":"## What is Mediapipe?\nMediapipe is a framework developed by Google that provides a set of tools and building blocks for building machine learning models for various tasks such as object detection, face detection, pose estimation, and hand tracking. It is an open-source, cross-platform framework designed to be flexible and scalable, and it can be used for a wide range of applications such as robotics, augmented reality, virtual reality, and video analysis.","metadata":{}},{"cell_type":"code","source":"!pip install mediapipe","metadata":{"id":"AuQfLvpuJkb0","outputId":"66d1f28a-8f45-46c8-fa86-b2ba383ab169","_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-14T08:16:59.163621Z","iopub.execute_input":"2023-03-14T08:16:59.164834Z","iopub.status.idle":"2023-03-14T08:17:16.128251Z","shell.execute_reply.started":"2023-03-14T08:16:59.164785Z","shell.execute_reply":"2023-03-14T08:17:16.126329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import cv2\nimport math\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport mediapipe as mp","metadata":{"execution":{"iopub.status.busy":"2023-03-14T08:17:16.130733Z","iopub.execute_input":"2023-03-14T08:17:16.131278Z","iopub.status.idle":"2023-03-14T08:17:16.457346Z","shell.execute_reply.started":"2023-03-14T08:17:16.131231Z","shell.execute_reply":"2023-03-14T08:17:16.456366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path0='/kaggle/input/rock-paper-scissors-dataset/train/paper/glu_172.png'\nimage=cv2.imread(path0)\nimage=cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\nprint(image.shape)\nimages=[image]","metadata":{"execution":{"iopub.status.busy":"2023-03-14T08:17:16.458900Z","iopub.execute_input":"2023-03-14T08:17:16.460374Z","iopub.status.idle":"2023-03-14T08:17:16.531241Z","shell.execute_reply.started":"2023-03-14T08:17:16.460312Z","shell.execute_reply":"2023-03-14T08:17:16.529721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"name=path0.split('/')[-1]\nprint(name)","metadata":{"execution":{"iopub.status.busy":"2023-03-14T08:17:16.534350Z","iopub.execute_input":"2023-03-14T08:17:16.534819Z","iopub.status.idle":"2023-03-14T08:17:16.541182Z","shell.execute_reply.started":"2023-03-14T08:17:16.534763Z","shell.execute_reply":"2023-03-14T08:17:16.539675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DESIRED_HEIGHT = 480\nDESIRED_WIDTH = 480\n\ndef resize_and_show(image):\n    h, w = image.shape[:2]\n    if h < w:\n        img = cv2.resize(image, (DESIRED_WIDTH, math.floor(h/(w/DESIRED_WIDTH))))\n    else:\n        img = cv2.resize(image, (math.floor(w/(h/DESIRED_HEIGHT)), DESIRED_HEIGHT))\n    plt.imshow(img)\n\nresize_and_show(image)","metadata":{"execution":{"iopub.status.busy":"2023-03-14T08:17:16.543193Z","iopub.execute_input":"2023-03-14T08:17:16.543629Z","iopub.status.idle":"2023-03-14T08:17:16.977943Z","shell.execute_reply.started":"2023-03-14T08:17:16.543586Z","shell.execute_reply":"2023-03-14T08:17:16.977013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* mp.solutions.hands is a module that provides tools for hand tracking and pose estimation. It includes pre-trained models for detecting and tracking hand landmarks in real-time video streams.\n\n* mp.solutions.drawing_utils is a module that provides utility functions for drawing 2D and 3D annotations on images and videos. It includes functions for drawing landmarks, connections, and bounding boxes.\n\n* mp.solutions.drawing_styles is a module that provides a set of predefined styles for drawing annotations with mp.solutions.drawing_utils. These styles define the colors, line thickness, and other visual properties of the annotations.\n\n* Together, these modules provide a powerful toolkit for building hand tracking and pose estimation applications with the Mediapipe framework.","metadata":{}},{"cell_type":"code","source":"mp_hands = mp.solutions.hands\nmp_drawing = mp.solutions.drawing_utils\nmp_drawing_styles = mp.solutions.drawing_styles","metadata":{"id":"BboTB-FAMfPo","outputId":"4709bce5-d4ae-464c-dfb4-b6b69e34dbc8","execution":{"iopub.status.busy":"2023-03-14T08:17:16.979061Z","iopub.execute_input":"2023-03-14T08:17:16.980184Z","iopub.status.idle":"2023-03-14T08:17:16.984837Z","shell.execute_reply.started":"2023-03-14T08:17:16.980143Z","shell.execute_reply":"2023-03-14T08:17:16.983518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet uses the mp_hands module to detect and draw landmarks on hands in an image. The mp_hands.Hands() function to create a hands object which can be used to process an image and detect hand landmarks. The results variable is assigned the output of processing the image using the hands.process() method, which detects the hand landmarks in the image.","metadata":{}},{"cell_type":"code","source":"with mp_hands.Hands(\n    static_image_mode=True,\n    max_num_hands=2,\n    min_detection_confidence=0.7) as hands:\n    \n    results = hands.process(cv2.flip(image,1))\n\n    if not results.multi_hand_landmarks:\n        print(\"No hand detected in the image.\")\n        \n    else:\n        image_hight, image_width, _ = image.shape\n        annotated_image = cv2.flip(image.copy(),1)\n\n        for hand_landmarks in results.multi_hand_landmarks:\n            mp_drawing.draw_landmarks(\n                annotated_image,\n                hand_landmarks,\n                mp_hands.HAND_CONNECTIONS,\n                mp_drawing_styles.get_default_hand_landmarks_style(),\n                mp_drawing_styles.get_default_hand_connections_style())\n            \n        resize_and_show(cv2.flip(annotated_image,1))","metadata":{"id":"BAivyQ_xOtFp","outputId":"512b4dcc-c26c-4f73-fed9-991cabb7d42d","execution":{"iopub.status.busy":"2023-03-14T08:17:16.986197Z","iopub.execute_input":"2023-03-14T08:17:16.986530Z","iopub.status.idle":"2023-03-14T08:17:17.511881Z","shell.execute_reply.started":"2023-03-14T08:17:16.986497Z","shell.execute_reply":"2023-03-14T08:17:17.510650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The mp_drawing.plot_landmarks() function is called to draw the landmarks on the image using the mp_hands.HAND_CONNECTIONS parameter to draw lines between the landmarks. The azimuth=5 parameter sets the azimuth angle (in degrees) of the camera used to capture the image.","metadata":{}},{"cell_type":"code","source":"with mp_hands.Hands(\n    static_image_mode=True,\n    max_num_hands=2,\n    min_detection_confidence=0.7) as hands:\n\n    results = hands.process(image)\n\n    for hand_world_landmarks in results.multi_hand_world_landmarks:\n        if not hand_world_landmarks:\n            continue\n\n        mp_drawing.plot_landmarks(\n            hand_world_landmarks, mp_hands.HAND_CONNECTIONS, azimuth=5)","metadata":{"id":"LAchzK23Uabf","outputId":"e63f897a-12bf-48d8-a7e0-591aa078e32c","execution":{"iopub.status.busy":"2023-03-14T08:17:17.513187Z","iopub.execute_input":"2023-03-14T08:17:17.513508Z","iopub.status.idle":"2023-03-14T08:17:18.399400Z","shell.execute_reply.started":"2023-03-14T08:17:17.513477Z","shell.execute_reply":"2023-03-14T08:17:18.397882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With mp_hands.Hands(...) an instance of the Hands class is created. An instance of the Hands class is required to execute the code inside the with block, so you cannot create a hands instance outside of the with block and then use it. Also, exiting the with block automatically frees resources. So to detect the hands again, you have to do with mp_hands.Hands(...) again.","metadata":{}},{"cell_type":"markdown","source":"## The MediaPipe Hands API implements algorithms for estimating 3D hand poses from 2D images using deep learning.","metadata":{}},{"cell_type":"markdown","source":"The landmark data obtained from the code above is stored in the multi_hand_landmarks attribute of the results object. This attribute contains a list of landmarks for each hand in the image. Each landmark contains x, y, z coordinates and is labeled using values from the mp_hands.HandLandmark enumeration.　It is because The MediaPipe Hands API implements algorithms for estimating 3D hand poses from 2D images using deep learning.","metadata":{}},{"cell_type":"markdown","source":"## This is the resason why the competition data has x, y, z coordinates.  ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Basics for 3D Plot","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom mpl_toolkits.mplot3d import Axes3D\n\nx = [1, 2, 3, 4, 5]\ny = [5, 4, 3, 2, 1]\nz = [0, 0, 0, 0, 0]\n\nfig = plt.figure()\nax = fig.add_subplot(111, projection='3d')\nax.scatter(x, y, z)\n\nax.set_xlabel('X Label')\nax.set_ylabel('Y Label')\nax.set_zlabel('Z Label')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-14T08:17:18.401465Z","iopub.execute_input":"2023-03-14T08:17:18.401977Z","iopub.status.idle":"2023-03-14T08:17:18.646320Z","shell.execute_reply.started":"2023-03-14T08:17:18.401906Z","shell.execute_reply":"2023-03-14T08:17:18.644960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}