{
  "id": 398113,
  "title": "Doubt regarding how mediapipe landmarks were obtained",
  "url": "/competitions/asl-signs/discussion/398113",
  "author_name": "",
  "post_date": "2023-03-28T15:44:58.148101Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello Kagglers! I was trying to make a WebApp that would extract the landmarks from a video and the tflite model would be used to predict the sign. I used the following code for extracting the landmarks on this <a href=\"cat\" target=\"_blank\">https://www.youtube.com/watch?v=iwJujoYRSMg</a> ASL video from youtube.</p>\n<pre><code> cv2\n mediapipe  mp\n pandas  pd\n numpy  np\n\nmp_hands = mp.solutions.hands\nmp_pose = mp.solutions.pose\nmp_face_mesh = mp.solutions.face_mesh\n\nmp_drawing = mp.solutions.drawing_utils\ndf={:[],:[],:[],:[]}\n\ncap = cv2.VideoCapture()\n mp_hands.Hands(static_image_mode=, max_num_hands=, min_detection_confidence=)  hands, \\\n     mp_pose.Pose(static_image_mode=, min_detection_confidence=)  pose, \\\n     mp_face_mesh.FaceMesh(static_image_mode=, max_num_faces=, min_detection_confidence=)  face_mesh:\n    =\n     cap.isOpened():\n        success, image = cap.read()\n          success:\n            \n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        hands_results = hands.process(image)\n        pose_results = pose.process(image)\n        face_mesh_results = face_mesh.process(image)\n        face_list =[i  i  ()]\n        pose_list =[i  i  ()]\n        left_list =[i  i  ()]\n        right_list = [i  i  ()]\n        landmark_list = face_list+left_list+pose_list+right_list\n        final_df={:landmark_list,:[np.nan  i  ()],:[np.nan  i  ()],:[np.nan  i  ()]}\n         ():\n            (result):\n                  isPose:\n                     frame_number,body_landmarks  (result):\n                         l_id, landmark  (body_landmarks.landmark):\n                            final_df[][l_id+offset]=landmark.x\n                            final_df[][l_id+offset]=landmark.y\n                            final_df[][l_id+offset]=landmark.z\n                 isPose:\n                     l_id, landmark  (result.landmark):\n                            final_df[][l_id+offset]=landmark.x\n                            final_df[][l_id+offset]=landmark.y\n                            final_df[][l_id+offset]=landmark.z\n        (hands_results):\n            (hands_results.multi_handedness  (hands_results.multi_handedness)==):\n                (hands_results.multi_handedness[].classification[].label==):\n                      insert_into_df(hands_results.multi_hand_landmarks[],final_df,,)\n                      insert_into_df(hands_results.multi_hand_landmarks[],final_df,++,)\n                :\n                     insert_into_df(hands_results.multi_hand_landmarks[],final_df,++,)\n                     insert_into_df(hands_results.multi_hand_landmarks[],final_df,,)\n            (hands_results.multi_handedness  (hands_results.multi_handedness)==):\n                  (hands_results.multi_handedness[].classification[].label==):\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,)\n                  :\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,++)\n        insert_into_df(face_mesh_results.multi_face_landmarks,final_df)\n        insert_into_df(pose_results.pose_landmarks,final_df,+,)\n        df[].extend(final_df[])\n        df[].extend(final_df[])\n        df[].extend(final_df[])\n        df[].extend(final_df[])\n\n\n    df=pd.DataFrame(df).reset_index(drop=)\n    df.to_csv()\n</code></pre>\n<p>But even the top performing models like [<a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution](this\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution](this</a> one) by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> detect the sign wrongly(I tried for 5-6 signs all of them were not predicted correctly). The \"cat\" sign was not even in top25 predictions. So I was wondering how exactly the landmarks were calculated from videos to create dataset for the competition. Any help on this will be highly appreciated!</p>",
  "messages": [
    {
      "id": "2200516",
      "postDate": "03/28/2023 15:44:58",
      "content": "<p>Hello Kagglers! I was trying to make a WebApp that would extract the landmarks from a video and the tflite model would be used to predict the sign. I used the following code for extracting the landmarks on this <a href=\"cat\" target=\"_blank\">https://www.youtube.com/watch?v=iwJujoYRSMg</a> ASL video from youtube.</p>\n<pre><code> cv2\n mediapipe  mp\n pandas  pd\n numpy  np\n\nmp_hands = mp.solutions.hands\nmp_pose = mp.solutions.pose\nmp_face_mesh = mp.solutions.face_mesh\n\nmp_drawing = mp.solutions.drawing_utils\ndf={:[],:[],:[],:[]}\n\ncap = cv2.VideoCapture()\n mp_hands.Hands(static_image_mode=, max_num_hands=, min_detection_confidence=)  hands, \\\n     mp_pose.Pose(static_image_mode=, min_detection_confidence=)  pose, \\\n     mp_face_mesh.FaceMesh(static_image_mode=, max_num_faces=, min_detection_confidence=)  face_mesh:\n    =\n     cap.isOpened():\n        success, image = cap.read()\n          success:\n            \n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        hands_results = hands.process(image)\n        pose_results = pose.process(image)\n        face_mesh_results = face_mesh.process(image)\n        face_list =[i  i  ()]\n        pose_list =[i  i  ()]\n        left_list =[i  i  ()]\n        right_list = [i  i  ()]\n        landmark_list = face_list+left_list+pose_list+right_list\n        final_df={:landmark_list,:[np.nan  i  ()],:[np.nan  i  ()],:[np.nan  i  ()]}\n         ():\n            (result):\n                  isPose:\n                     frame_number,body_landmarks  (result):\n                         l_id, landmark  (body_landmarks.landmark):\n                            final_df[][l_id+offset]=landmark.x\n                            final_df[][l_id+offset]=landmark.y\n                            final_df[][l_id+offset]=landmark.z\n                 isPose:\n                     l_id, landmark  (result.landmark):\n                            final_df[][l_id+offset]=landmark.x\n                            final_df[][l_id+offset]=landmark.y\n                            final_df[][l_id+offset]=landmark.z\n        (hands_results):\n            (hands_results.multi_handedness  (hands_results.multi_handedness)==):\n                (hands_results.multi_handedness[].classification[].label==):\n                      insert_into_df(hands_results.multi_hand_landmarks[],final_df,,)\n                      insert_into_df(hands_results.multi_hand_landmarks[],final_df,++,)\n                :\n                     insert_into_df(hands_results.multi_hand_landmarks[],final_df,++,)\n                     insert_into_df(hands_results.multi_hand_landmarks[],final_df,,)\n            (hands_results.multi_handedness  (hands_results.multi_handedness)==):\n                  (hands_results.multi_handedness[].classification[].label==):\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,)\n                  :\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,++)\n        insert_into_df(face_mesh_results.multi_face_landmarks,final_df)\n        insert_into_df(pose_results.pose_landmarks,final_df,+,)\n        df[].extend(final_df[])\n        df[].extend(final_df[])\n        df[].extend(final_df[])\n        df[].extend(final_df[])\n\n\n    df=pd.DataFrame(df).reset_index(drop=)\n    df.to_csv()\n</code></pre>\n<p>But even the top performing models like [<a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution](this\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution](this</a> one) by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> detect the sign wrongly(I tried for 5-6 signs all of them were not predicted correctly). The \"cat\" sign was not even in top25 predictions. So I was wondering how exactly the landmarks were calculated from videos to create dataset for the competition. Any help on this will be highly appreciated!</p>",
      "rawMarkdown": "Hello Kagglers! I was trying to make a WebApp that would extract the landmarks from a video and the tflite model would be used to predict the sign. I used the following code for extracting the landmarks on this [https://www.youtube.com/watch?v=iwJujoYRSMg](cat) ASL video from youtube.\n\n\n```python\nimport cv2\nimport mediapipe as mp\nimport pandas as pd\nimport numpy as np\n\nmp_hands = mp.solutions.hands\nmp_pose = mp.solutions.pose\nmp_face_mesh = mp.solutions.face_mesh\n\nmp_drawing = mp.solutions.drawing_utils\ndf={'landmark':[],'x':[],'y':[],'z':[]}\n\ncap = cv2.VideoCapture('cat.mp4')\nwith mp_hands.Hands(static_image_mode=False, max_num_hands=2, min_detection_confidence=0.5) as hands, \\\n     mp_pose.Pose(static_image_mode=False, min_detection_confidence=0.5) as pose, \\\n     mp_face_mesh.FaceMesh(static_image_mode=False, max_num_faces=1, min_detection_confidence=0.5) as face_mesh:\n    iter=0\n    while cap.isOpened():\n        success, image = cap.read()\n        if not success:\n            break\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        hands_results = hands.process(image)\n        pose_results = pose.process(image)\n        face_mesh_results = face_mesh.process(image)\n        face_list =[i for i in range(468)]\n        pose_list =[i for i in range(33)]\n        left_list =[i for i in range(21)]\n        right_list = [i for i in range(21)]\n        landmark_list = face_list+left_list+pose_list+right_list\n        final_df={'landmark':landmark_list,'x':[np.nan for i in range(543)],'y':[np.nan for i in range(543)],'z':[np.nan for i in range(543)]}\n        def insert_into_df(result,final_df,offset=0,isPose=False):\n            if(result):\n                if not isPose:\n                    for frame_number,body_landmarks in enumerate(result):\n                        for l_id, landmark in enumerate(body_landmarks.landmark):\n                            final_df['x'][l_id+offset]=landmark.x\n                            final_df['y'][l_id+offset]=landmark.y\n                            final_df['z'][l_id+offset]=landmark.z\n                elif isPose:\n                    for l_id, landmark in enumerate(result.landmark):\n                            final_df['x'][l_id+offset]=landmark.x\n                            final_df['y'][l_id+offset]=landmark.y\n                            final_df['z'][l_id+offset]=landmark.z\n        if(hands_results):\n            if(hands_results.multi_handedness and len(hands_results.multi_handedness)==2):\n                if(hands_results.multi_handedness[0].classification[0].label=='Left'):\n                      insert_into_df(hands_results.multi_hand_landmarks[0],final_df,468,True)\n                      insert_into_df(hands_results.multi_hand_landmarks[1],final_df,468+21+33,True)\n                else:\n                     insert_into_df(hands_results.multi_hand_landmarks[0],final_df,468+21+33,True)\n                     insert_into_df(hands_results.multi_hand_landmarks[1],final_df,468,True)\n            elif(hands_results.multi_handedness and len(hands_results.multi_handedness)==1):\n                  if(hands_results.multi_handedness[0].classification[0].label=='Left'):\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,468)\n                  else:\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,468+21+33)\n        insert_into_df(face_mesh_results.multi_face_landmarks,final_df)\n        insert_into_df(pose_results.pose_landmarks,final_df,468+21,True)\n        df['landmark'].extend(final_df['landmark'])\n        df['x'].extend(final_df['x'])\n        df['y'].extend(final_df['y'])\n        df['z'].extend(final_df['z'])\n    \n\n    df=pd.DataFrame(df).reset_index(drop=True)\n    df.to_csv('Cat_sign.csv')\n```\n\nBut even the top performing models like [https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution](this one) by @hengck23 detect the sign wrongly(I tried for 5-6 signs all of them were not predicted correctly). The \"cat\" sign was not even in top25 predictions. So I was wondering how exactly the landmarks were calculated from videos to create dataset for the competition. Any help on this will be highly appreciated!",
      "votes": null
    },
    {
      "id": "2200555",
      "postDate": "03/28/2023 16:01:38",
      "content": "<p>Hi,<br>\nThis is a cool idea. I'm curious what does the distribution of the coordinates you've collected look like? The points in the training dataset have been normalized to the video frame from what I've seen.<br>\nRegards,<br>\nPeter</p>",
      "rawMarkdown": "Hi,\nThis is a cool idea. I'm curious what does the distribution of the coordinates you've collected look like? The points in the training dataset have been normalized to the video frame from what I've seen.\nRegards,\nPeter",
      "votes": null
    },
    {
      "id": "2200607",
      "postDate": "03/28/2023 16:35:00",
      "content": "<p>You could have a look at the raw logits output of the model and check whether the predictions are looking random, or whether the model predicts a sign with high probability (75+%). In the first case, if the predictions are looking random, this could be an indication the data pre-processing is different from the training data as <a href=\"https://www.kaggle.com/koehlepe\" target=\"_blank\">@koehlepe</a> mentioned.</p>\n<p>Top 25 does not mean much if the predicted probabilities are all random'ish and all ~0.4% (1/250 random guessing chance).</p>",
      "rawMarkdown": "You could have a look at the raw logits output of the model and check whether the predictions are looking random, or whether the model predicts a sign with high probability (75+%). In the first case, if the predictions are looking random, this could be an indication the data pre-processing is different from the training data as @koehlepe mentioned.\n\nTop 25 does not mean much if the predicted probabilities are all random'ish and all ~0.4% (1/250 random guessing chance).",
      "votes": null
    },
    {
      "id": "2200616",
      "postDate": "03/28/2023 16:49:00",
      "content": "<p>A couple of points:<br>\n1) the video is relatively low resolution: 360p.  Mediapipe likes much higher resolution. I'd go with 1080p minimum.<br>\n2) Make sure you are using the right version of Mediapipe Holistic. <br>\n3) the youtube video you link is of a child while the training videos in the competition are of adults</p>",
      "rawMarkdown": "A couple of points:\n1) the video is relatively low resolution: 360p.  Mediapipe likes much higher resolution. I'd go with 1080p minimum.\n2) Make sure you are using the right version of Mediapipe Holistic. \n3) the youtube video you link is of a child while the training videos in the competition are of adults",
      "votes": null
    },
    {
      "id": "2200648",
      "postDate": "03/28/2023 17:22:47",
      "content": "<p>Hi, thanks for your inputs. Just to clarify, I used mediapipe version 0.9.2.1, Is that the right version?</p>",
      "rawMarkdown": "Hi, thanks for your inputs. Just to clarify, I used mediapipe version 0.9.2.1, Is that the right version?",
      "votes": null
    },
    {
      "id": "2200653",
      "postDate": "03/28/2023 17:32:31",
      "content": "<p>According to the Data Card the correct Mediapipe version is 0.9.0.1.</p>\n<pre><code>The Isolated Sign Language Recognition corpus (version 1.0) is a collection of hand and facial landmarks generated by Mediapipe version 0.9.0.1 on ~100k videos of isolated signs performed by 21 Deaf signers from a 250-sign vocabulary.\n</code></pre>\n<p>Source: <a href=\"https://www.kaggle.com/competitions/asl-signs/overview/data-card\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/overview/data-card</a></p>",
      "rawMarkdown": "According to the Data Card the correct Mediapipe version is 0.9.0.1.\n\n```\nThe Isolated Sign Language Recognition corpus (version 1.0) is a collection of hand and facial landmarks generated by Mediapipe version 0.9.0.1 on ~100k videos of isolated signs performed by 21 Deaf signers from a 250-sign vocabulary.\n```\n\nSource: https://www.kaggle.com/competitions/asl-signs/overview/data-card",
      "votes": null
    },
    {
      "id": "2200987",
      "postDate": "03/29/2023 00:26:10",
      "content": "<p>note that there is some \"flaw\" in the dataset<br>\n\"signs performed by 21 Deaf signers\"</p>",
      "rawMarkdown": "note that there is some \"flaw\" in the dataset\n\"signs performed by 21 Deaf signers\"",
      "votes": null
    },
    {
      "id": "2201430",
      "postDate": "03/29/2023 09:50:13",
      "content": "<p>Have you considered using MediaPipe holistic instead of hands, pose and face separately? The hand model expects cropped hand images, not a full frame as you are giving here.</p>\n<p>Plus, you should try visualizing the keypoints to see if the output looks correct.</p>",
      "rawMarkdown": "Have you considered using MediaPipe holistic instead of hands, pose and face separately? The hand model expects cropped hand images, not a full frame as you are giving here.\n\nPlus, you should try visualizing the keypoints to see if the output looks correct.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2200555,
      "author_name": "koehlepe",
      "author_url": "",
      "post_date": "03/28/2023 16:01:38",
      "content": "<p>Hi,<br>\nThis is a cool idea. I'm curious what does the distribution of the coordinates you've collected look like? The points in the training dataset have been normalized to the video frame from what I've seen.<br>\nRegards,<br>\nPeter</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2200607,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "03/28/2023 16:35:00",
      "content": "<p>You could have a look at the raw logits output of the model and check whether the predictions are looking random, or whether the model predicts a sign with high probability (75+%). In the first case, if the predictions are looking random, this could be an indication the data pre-processing is different from the training data as <a href=\"https://www.kaggle.com/koehlepe\" target=\"_blank\">@koehlepe</a> mentioned.</p>\n<p>Top 25 does not mean much if the predicted probabilities are all random'ish and all ~0.4% (1/250 random guessing chance).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2200616,
      "author_name": "thadstarner",
      "author_url": "",
      "post_date": "03/28/2023 16:49:00",
      "content": "<p>A couple of points:<br>\n1) the video is relatively low resolution: 360p.  Mediapipe likes much higher resolution. I'd go with 1080p minimum.<br>\n2) Make sure you are using the right version of Mediapipe Holistic. <br>\n3) the youtube video you link is of a child while the training videos in the competition are of adults</p>",
      "votes": null,
      "replies": [
        {
          "id": 2200648,
          "author_name": "dhan12369",
          "author_url": "",
          "post_date": "03/28/2023 17:22:47",
          "content": "<p>Hi, thanks for your inputs. Just to clarify, I used mediapipe version 0.9.2.1, Is that the right version?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2200653,
              "author_name": "gregorlied",
              "author_url": "",
              "post_date": "03/28/2023 17:32:31",
              "content": "<p>According to the Data Card the correct Mediapipe version is 0.9.0.1.</p>\n<pre><code>The Isolated Sign Language Recognition corpus (version 1.0) is a collection of hand and facial landmarks generated by Mediapipe version 0.9.0.1 on ~100k videos of isolated signs performed by 21 Deaf signers from a 250-sign vocabulary.\n</code></pre>\n<p>Source: <a href=\"https://www.kaggle.com/competitions/asl-signs/overview/data-card\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/overview/data-card</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2200987,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/29/2023 00:26:10",
      "content": "<p>note that there is some \"flaw\" in the dataset<br>\n\"signs performed by 21 Deaf signers\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2201430,
      "author_name": "mdecoster",
      "author_url": "",
      "post_date": "03/29/2023 09:50:13",
      "content": "<p>Have you considered using MediaPipe holistic instead of hands, pose and face separately? The hand model expects cropped hand images, not a full frame as you are giving here.</p>\n<p>Plus, you should try visualizing the keypoints to see if the output looks correct.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2200516": "Hello Kagglers! I was trying to make a WebApp that would extract the landmarks from a video and the tflite model would be used to predict the sign. I used the following code for extracting the landmarks on this [https://www.youtube.com/watch?v=iwJujoYRSMg](cat) ASL video from youtube.\n\n\n```python\nimport cv2\nimport mediapipe as mp\nimport pandas as pd\nimport numpy as np\n\nmp_hands = mp.solutions.hands\nmp_pose = mp.solutions.pose\nmp_face_mesh = mp.solutions.face_mesh\n\nmp_drawing = mp.solutions.drawing_utils\ndf={'landmark':[],'x':[],'y':[],'z':[]}\n\ncap = cv2.VideoCapture('cat.mp4')\nwith mp_hands.Hands(static_image_mode=False, max_num_hands=2, min_detection_confidence=0.5) as hands, \\\n     mp_pose.Pose(static_image_mode=False, min_detection_confidence=0.5) as pose, \\\n     mp_face_mesh.FaceMesh(static_image_mode=False, max_num_faces=1, min_detection_confidence=0.5) as face_mesh:\n    iter=0\n    while cap.isOpened():\n        success, image = cap.read()\n        if not success:\n            break\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        hands_results = hands.process(image)\n        pose_results = pose.process(image)\n        face_mesh_results = face_mesh.process(image)\n        face_list =[i for i in range(468)]\n        pose_list =[i for i in range(33)]\n        left_list =[i for i in range(21)]\n        right_list = [i for i in range(21)]\n        landmark_list = face_list+left_list+pose_list+right_list\n        final_df={'landmark':landmark_list,'x':[np.nan for i in range(543)],'y':[np.nan for i in range(543)],'z':[np.nan for i in range(543)]}\n        def insert_into_df(result,final_df,offset=0,isPose=False):\n            if(result):\n                if not isPose:\n                    for frame_number,body_landmarks in enumerate(result):\n                        for l_id, landmark in enumerate(body_landmarks.landmark):\n                            final_df['x'][l_id+offset]=landmark.x\n                            final_df['y'][l_id+offset]=landmark.y\n                            final_df['z'][l_id+offset]=landmark.z\n                elif isPose:\n                    for l_id, landmark in enumerate(result.landmark):\n                            final_df['x'][l_id+offset]=landmark.x\n                            final_df['y'][l_id+offset]=landmark.y\n                            final_df['z'][l_id+offset]=landmark.z\n        if(hands_results):\n            if(hands_results.multi_handedness and len(hands_results.multi_handedness)==2):\n                if(hands_results.multi_handedness[0].classification[0].label=='Left'):\n                      insert_into_df(hands_results.multi_hand_landmarks[0],final_df,468,True)\n                      insert_into_df(hands_results.multi_hand_landmarks[1],final_df,468+21+33,True)\n                else:\n                     insert_into_df(hands_results.multi_hand_landmarks[0],final_df,468+21+33,True)\n                     insert_into_df(hands_results.multi_hand_landmarks[1],final_df,468,True)\n            elif(hands_results.multi_handedness and len(hands_results.multi_handedness)==1):\n                  if(hands_results.multi_handedness[0].classification[0].label=='Left'):\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,468)\n                  else:\n                      insert_into_df(hands_results.multi_hand_landmarks,final_df,468+21+33)\n        insert_into_df(face_mesh_results.multi_face_landmarks,final_df)\n        insert_into_df(pose_results.pose_landmarks,final_df,468+21,True)\n        df['landmark'].extend(final_df['landmark'])\n        df['x'].extend(final_df['x'])\n        df['y'].extend(final_df['y'])\n        df['z'].extend(final_df['z'])\n    \n\n    df=pd.DataFrame(df).reset_index(drop=True)\n    df.to_csv('Cat_sign.csv')\n```\n\nBut even the top performing models like [https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution](this one) by @hengck23 detect the sign wrongly(I tried for 5-6 signs all of them were not predicted correctly). The \"cat\" sign was not even in top25 predictions. So I was wondering how exactly the landmarks were calculated from videos to create dataset for the competition. Any help on this will be highly appreciated!",
    "2200555": "Hi,\nThis is a cool idea. I'm curious what does the distribution of the coordinates you've collected look like? The points in the training dataset have been normalized to the video frame from what I've seen.\nRegards,\nPeter",
    "2200607": "You could have a look at the raw logits output of the model and check whether the predictions are looking random, or whether the model predicts a sign with high probability (75+%). In the first case, if the predictions are looking random, this could be an indication the data pre-processing is different from the training data as @koehlepe mentioned.\n\nTop 25 does not mean much if the predicted probabilities are all random'ish and all ~0.4% (1/250 random guessing chance).",
    "2200616": "A couple of points:\n1) the video is relatively low resolution: 360p.  Mediapipe likes much higher resolution. I'd go with 1080p minimum.\n2) Make sure you are using the right version of Mediapipe Holistic. \n3) the youtube video you link is of a child while the training videos in the competition are of adults",
    "2200648": "Hi, thanks for your inputs. Just to clarify, I used mediapipe version 0.9.2.1, Is that the right version?",
    "2200653": "According to the Data Card the correct Mediapipe version is 0.9.0.1.\n\n```\nThe Isolated Sign Language Recognition corpus (version 1.0) is a collection of hand and facial landmarks generated by Mediapipe version 0.9.0.1 on ~100k videos of isolated signs performed by 21 Deaf signers from a 250-sign vocabulary.\n```\n\nSource: https://www.kaggle.com/competitions/asl-signs/overview/data-card",
    "2200987": "note that there is some \"flaw\" in the dataset\n\"signs performed by 21 Deaf signers\"",
    "2201430": "Have you considered using MediaPipe holistic instead of hands, pose and face separately? The hand model expects cropped hand images, not a full frame as you are giving here.\n\nPlus, you should try visualizing the keypoints to see if the output looks correct."
  },
  "source": "meta"
}