{
  "id": 390715,
  "title": "How the data were recorded? Can all 250 words be completed with one hand?",
  "url": "/competitions/asl-signs/discussion/390715",
  "author_name": "",
  "post_date": "2023-02-26T22:50:03.414317500Z",
  "votes": 17,
  "comment_count": 11,
  "views": 0,
  "content": "<blockquote>\n  <p>Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign. </p>\n</blockquote>\n<p>It says here that signers must press and hold the screen continuously while recording, and I have the following questions.</p>\n<ol>\n<li><p>Does this mean that all 250 words can be done with one hand?</p></li>\n<li><p>Why does it sometimes appear that both hands are present at the same time? For example,  <a href=\"https://www.kaggle.com/code/robikscube/sign-language-recognition-eda-twitch-stream\" target=\"_blank\">here</a>, the shhh action in frames 17 and 19 is done by two hands, is this because mediapipe incorrectly identifies one hand as two hands in these two frames?</p></li>\n<li><p>What device was the data recorded with? Was it taken by the phone's front camera that the participant had to hold down? Or was it a different camera?</p></li>\n<li><p>What angle was the data recorded with? Do all participants take the selfie with one hand on the phone and sign with the other hand? Or will the phone be set up in the proper position with a stand and then have to reach out and hold the screen while recording?</p></li>\n</ol>\n<p>Since the 33 keypoints of the pose are usually there, I assume there is another camera set up in front of the participant to capture the whole body. But then why would there usually only be one hand? I would like to get further confirmation.</p>",
  "messages": [
    {
      "id": "2160703",
      "postDate": "02/26/2023 22:50:03",
      "content": "<blockquote>\n  <p>Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign. </p>\n</blockquote>\n<p>It says here that signers must press and hold the screen continuously while recording, and I have the following questions.</p>\n<ol>\n<li><p>Does this mean that all 250 words can be done with one hand?</p></li>\n<li><p>Why does it sometimes appear that both hands are present at the same time? For example,  <a href=\"https://www.kaggle.com/code/robikscube/sign-language-recognition-eda-twitch-stream\" target=\"_blank\">here</a>, the shhh action in frames 17 and 19 is done by two hands, is this because mediapipe incorrectly identifies one hand as two hands in these two frames?</p></li>\n<li><p>What device was the data recorded with? Was it taken by the phone's front camera that the participant had to hold down? Or was it a different camera?</p></li>\n<li><p>What angle was the data recorded with? Do all participants take the selfie with one hand on the phone and sign with the other hand? Or will the phone be set up in the proper position with a stand and then have to reach out and hold the screen while recording?</p></li>\n</ol>\n<p>Since the 33 keypoints of the pose are usually there, I assume there is another camera set up in front of the participant to capture the whole body. But then why would there usually only be one hand? I would like to get further confirmation.</p>",
      "rawMarkdown": ">Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign. \n\nIt says here that signers must press and hold the screen continuously while recording, and I have the following questions.\n\n1. Does this mean that all 250 words can be done with one hand?\n\n2. Why does it sometimes appear that both hands are present at the same time? For example,  [here](https://www.kaggle.com/code/robikscube/sign-language-recognition-eda-twitch-stream), the shhh action in frames 17 and 19 is done by two hands, is this because mediapipe incorrectly identifies one hand as two hands in these two frames?\n\n3. What device was the data recorded with? Was it taken by the phone's front camera that the participant had to hold down? Or was it a different camera?\n\n4. What angle was the data recorded with? Do all participants take the selfie with one hand on the phone and sign with the other hand? Or will the phone be set up in the proper position with a stand and then have to reach out and hold the screen while recording?\n\nSince the 33 keypoints of the pose are usually there, I assume there is another camera set up in front of the participant to capture the whole body. But then why would there usually only be one hand? I would like to get further confirmation.",
      "votes": null
    },
    {
      "id": "2160743",
      "postDate": "02/27/2023 00:02:49",
      "content": "<p>Great question. My only thought would be they had someone else record it?</p>\n<p>Definitely interested to hear what the hosts have to say though. Thanks for raising this!</p>",
      "rawMarkdown": "Great question. My only thought would be they had someone else record it?\n\nDefinitely interested to hear what the hosts have to say though. Thanks for raising this!",
      "votes": null
    },
    {
      "id": "2160782",
      "postDate": "02/27/2023 01:31:38",
      "content": "<p>All videos were recorded with the selfie camera.  We asked that all of them be recorded with one hand.</p>",
      "rawMarkdown": "All videos were recorded with the selfie camera.  We asked that all of them be recorded with one hand.",
      "votes": null
    },
    {
      "id": "2161629",
      "postDate": "02/27/2023 16:07:11",
      "content": "<p>Randomly picking various \"signs\" and reviewing some of the recorded key-points shows to me that the dataset is very \"raw\". The \"hands\" landmarks are missing a lot. The recording FPS seem to vary a lot even for the same participants.</p>\n<p>Random sample: </p>\n<ul>\n<li>sign \"TV\"    </li>\n<li>number of samples for participant: 21</li>\n<li>shortest recording size: 97385 bytes</li>\n<li>mean recording size: 760217.38 bytes</li>\n<li>longest recording size: 2402799 bytes</li>\n</ul>\n<p>That's 24x difference in recording size from smallest to longest recording for the same sign!! </p>\n<p>I'm personally curious if the videos dataset, test-set had been reviewed and if the labels had been confirmed. Or, is it \"self\" curated and the labels where assigned by persons that recorded the signs. </p>",
      "rawMarkdown": "Randomly picking various \"signs\" and reviewing some of the recorded key-points shows to me that the dataset is very \"raw\". The \"hands\" landmarks are missing a lot. The recording FPS seem to vary a lot even for the same participants.\n\nRandom sample: \n- sign \"TV\"\t\n- number of samples for participant: 21\n- shortest recording size: 97385 bytes\n- mean recording size: 760217.38 bytes\n- longest recording size: 2402799 bytes\n\nThat's 24x difference in recording size from smallest to longest recording for the same sign!! \n\nI'm personally curious if the videos dataset, test-set had been reviewed and if the labels had been confirmed. Or, is it \"self\" curated and the labels where assigned by persons that recorded the signs.",
      "votes": null
    },
    {
      "id": "2163375",
      "postDate": "02/28/2023 19:05:43",
      "content": "<p>\"Since the 33 keypoints of the pose are usually there\"</p>\n<p>I haven't done any EDA myself, but I got the impression from others EDA that each of the four types are all-or-nothing in terms of NaN values(?) If so, there might be some MediaPipe logic that extrapolates the 'off-camera position' of all the pose elements that aren't on camera?</p>\n<p>In the end, I think this is another important question. Although I doubt we care about the foot, using it as an example, how is the right foot handled (when off camera) by MediaPipe pose landmark processing?</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> care to run some EDA on the dataset given, and/or maybe on picture-to-mediapipe results, in terms of NaN vs off-camera coordinates?</p>",
      "rawMarkdown": "\"Since the 33 keypoints of the pose are usually there\"\n\nI haven't done any EDA myself, but I got the impression from others EDA that each of the four types are all-or-nothing in terms of NaN values(?) If so, there might be some MediaPipe logic that extrapolates the 'off-camera position' of all the pose elements that aren't on camera?\n\nIn the end, I think this is another important question. Although I doubt we care about the foot, using it as an example, how is the right foot handled (when off camera) by MediaPipe pose landmark processing?\n\n@dschettler8845 care to run some EDA on the dataset given, and/or maybe on picture-to-mediapipe results, in terms of NaN vs off-camera coordinates?",
      "votes": null
    },
    {
      "id": "2163423",
      "postDate": "02/28/2023 19:49:26",
      "content": "<p>Not sure if this helps and more of an FYI:</p>\n<p>I looked at around 10,000 pieces of the data just to gauge… it looks like around 4-5% of the examples have both hands present. (407 out of the 10000 had non-zero sums for LH or RH).</p>\n<hr>\n<p><strong>Here are some examples showing the raw idx (--idx--) and the sum of the LH array and the sum of RH array (LH – RH)</strong></p>\n<pre><code>--68562--\n76.44444  –  317.19916\n\n--59308--\n102.50776  –  202.10736\n\n--84350--\n325.70483  –  32.586136\n\n--44107--\n176.67221  –  57.18132\n\n--70501--\n44.46574  –  199.6751\n\n--75306--\n91.02222  –  244.8219\n\n--37792--\n763.42896  –  203.41107\n\n--65315--\n643.9147  –  182.17331\n\n--25109--\n404.81876  –  111.70086\n\n--79022--\n267.5301  –  25.0708\n</code></pre>",
      "rawMarkdown": "Not sure if this helps and more of an FYI:\n\nI looked at around 10,000 pieces of the data just to gauge... it looks like around 4-5% of the examples have both hands present. (407 out of the 10000 had non-zero sums for LH or RH).\n\n---\n\n**Here are some examples showing the raw idx (--idx--) and the sum of the LH array and the sum of RH array (LH – RH)**\n\n```\n--68562--\n76.44444  –  317.19916\n\n--59308--\n102.50776  –  202.10736\n\n--84350--\n325.70483  –  32.586136\n\n--44107--\n176.67221  –  57.18132\n\n--70501--\n44.46574  –  199.6751\n\n--75306--\n91.02222  –  244.8219\n\n--37792--\n763.42896  –  203.41107\n\n--65315--\n643.9147  –  182.17331\n\n--25109--\n404.81876  –  111.70086\n\n--79022--\n267.5301  –  25.0708\n```",
      "votes": null
    },
    {
      "id": "2163434",
      "postDate": "02/28/2023 19:59:20",
      "content": "<p>I have looked at some EDAs done by others and my impression is that all samples have all the face and pose, Nan is only present in the left and right hands.</p>\n<p>I downloaded a notebook that uses MediaPipe and my laptop's camera to inference keypoints and I only got the keypoints for all that appeared in the camera frame, the pose part, most of the time it was only two points on the shoulders and 11 points on the face, if the hands appeared in the camera there were also a few points on the hands. I'm not sure if the notebook I'm using is predicting all the pose and just not showing all the keypoints, or if it's only predicting the keypoints in the frame.</p>\n<p>If all the data is done as selfies, then the keypoints data for the rest of the body is really strange and I think we need to experiment to see what difference the results are after removing the points from those parts.</p>",
      "rawMarkdown": "I have looked at some EDAs done by others and my impression is that all samples have all the face and pose, Nan is only present in the left and right hands.\n\nI downloaded a notebook that uses MediaPipe and my laptop's camera to inference keypoints and I only got the keypoints for all that appeared in the camera frame, the pose part, most of the time it was only two points on the shoulders and 11 points on the face, if the hands appeared in the camera there were also a few points on the hands. I'm not sure if the notebook I'm using is predicting all the pose and just not showing all the keypoints, or if it's only predicting the keypoints in the frame.\n\nIf all the data is done as selfies, then the keypoints data for the rest of the body is really strange and I think we need to experiment to see what difference the results are after removing the points from those parts.",
      "votes": null
    },
    {
      "id": "2163441",
      "postDate": "02/28/2023 20:03:51",
      "content": "<p>I have used mediapipe for a project. I have seen NaN and extrapolated values for both off image and within image.  Maybe not the case with all video, but I've seen it do this with frames in middle of both low and high resolution video at various fps even though the video looks really sharp and seems like the values should have been accurate.</p>",
      "rawMarkdown": "I have used mediapipe for a project. I have seen NaN and extrapolated values for both off image and within image.  Maybe not the case with all video, but I've seen it do this with frames in middle of both low and high resolution video at various fps even though the video looks really sharp and seems like the values should have been accurate.",
      "votes": null
    },
    {
      "id": "2163450",
      "postDate": "02/28/2023 20:11:20",
      "content": "<p>I guess this is because the MediaPipe prediction is wrong and identifies one hand as two, perhaps because the signer is moving too fast, resulting in a virtual image. In the example I gave above, there are only two discrete frames in the video where the two hands appear, and Rob has drawn the hand in one of the frames. If you look carefully you will see that the two hands are oriented exactly the same way, rather than symmetrically left and right, and I don't think it would be possible to flip the other hand anyway, which would require the fingers of the other hand to make a reverse fist😂</p>",
      "rawMarkdown": "I guess this is because the MediaPipe prediction is wrong and identifies one hand as two, perhaps because the signer is moving too fast, resulting in a virtual image. In the example I gave above, there are only two discrete frames in the video where the two hands appear, and Rob has drawn the hand in one of the frames. If you look carefully you will see that the two hands are oriented exactly the same way, rather than symmetrically left and right, and I don't think it would be possible to flip the other hand anyway, which would require the fingers of the other hand to make a reverse fist😂",
      "votes": null
    },
    {
      "id": "2163547",
      "postDate": "02/28/2023 22:59:38",
      "content": "<p>Sure. Let me add it to the eda and I’ll calculate it for every example.</p>",
      "rawMarkdown": "Sure. Let me add it to the eda and I’ll calculate it for every example.",
      "votes": null
    },
    {
      "id": "2164564",
      "postDate": "03/01/2023 16:59:33",
      "content": "<p>I updated my EDA to extract NaN counts and calculate the percentage per example as well (i.e. 5 NaN of 100 vs 5 NaN of 1000).</p>\n<p>The following plot shows the general distributions (independently plotted). </p>\n<ul>\n<li><b>Note the Y axis is log scale.</b></li>\n<li>0.0 (Left Side) - Indicates that no values are missing</li>\n<li>1.0 (Right Side) - Indicates that all values are missing</li>\n</ul>\n<p><img src=\"https://i.ibb.co/8KNd8zK/Screenshot-2023-03-01-at-11-53-16-AM.png\"></p>\n<p><b>Quick Takeaways</b></p>\n<ul>\n<li>Face points can be NaN although it is less common than in the Hand data</li>\n<li>Pose points are never NaN</li>\n<li>Left and Right hand distributions are similar but Right Hand is full NaN less than Left Hand</li>\n<li>Pose, Left-Hand, and Right-Hand all have intermediate (not all missing or all present) sequences, however, they are less common than the case where all points are NaN or valid.</li>\n</ul>",
      "rawMarkdown": "I updated my EDA to extract NaN counts and calculate the percentage per example as well (i.e. 5 NaN of 100 vs 5 NaN of 1000).\n\nThe following plot shows the general distributions (independently plotted). \n* <b>Note the Y axis is log scale.</b>\n* 0.0 (Left Side) - Indicates that no values are missing\n* 1.0 (Right Side) - Indicates that all values are missing\n\n<img src=\"https://i.ibb.co/8KNd8zK/Screenshot-2023-03-01-at-11-53-16-AM.png\" width=100%>\n\n<b>Quick Takeaways</b>\n* Face points can be NaN although it is less common than in the Hand data\n* Pose points are never NaN\n* Left and Right hand distributions are similar but Right Hand is full NaN less than Left Hand\n* Pose, Left-Hand, and Right-Hand all have intermediate (not all missing or all present) sequences, however, they are less common than the case where all points are NaN or valid.",
      "votes": null
    },
    {
      "id": "2164600",
      "postDate": "03/01/2023 17:27:30",
      "content": "<p>This is great analysis, thanks!</p>",
      "rawMarkdown": "This is great analysis, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2160743,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "02/27/2023 00:02:49",
      "content": "<p>Great question. My only thought would be they had someone else record it?</p>\n<p>Definitely interested to hear what the hosts have to say though. Thanks for raising this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2160782,
      "author_name": "thadstarner",
      "author_url": "",
      "post_date": "02/27/2023 01:31:38",
      "content": "<p>All videos were recorded with the selfie camera.  We asked that all of them be recorded with one hand.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2161629,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "02/27/2023 16:07:11",
      "content": "<p>Randomly picking various \"signs\" and reviewing some of the recorded key-points shows to me that the dataset is very \"raw\". The \"hands\" landmarks are missing a lot. The recording FPS seem to vary a lot even for the same participants.</p>\n<p>Random sample: </p>\n<ul>\n<li>sign \"TV\"    </li>\n<li>number of samples for participant: 21</li>\n<li>shortest recording size: 97385 bytes</li>\n<li>mean recording size: 760217.38 bytes</li>\n<li>longest recording size: 2402799 bytes</li>\n</ul>\n<p>That's 24x difference in recording size from smallest to longest recording for the same sign!! </p>\n<p>I'm personally curious if the videos dataset, test-set had been reviewed and if the labels had been confirmed. Or, is it \"self\" curated and the labels where assigned by persons that recorded the signs. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2163375,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "02/28/2023 19:05:43",
      "content": "<p>\"Since the 33 keypoints of the pose are usually there\"</p>\n<p>I haven't done any EDA myself, but I got the impression from others EDA that each of the four types are all-or-nothing in terms of NaN values(?) If so, there might be some MediaPipe logic that extrapolates the 'off-camera position' of all the pose elements that aren't on camera?</p>\n<p>In the end, I think this is another important question. Although I doubt we care about the foot, using it as an example, how is the right foot handled (when off camera) by MediaPipe pose landmark processing?</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> care to run some EDA on the dataset given, and/or maybe on picture-to-mediapipe results, in terms of NaN vs off-camera coordinates?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2163434,
          "author_name": "chenlin1999",
          "author_url": "",
          "post_date": "02/28/2023 19:59:20",
          "content": "<p>I have looked at some EDAs done by others and my impression is that all samples have all the face and pose, Nan is only present in the left and right hands.</p>\n<p>I downloaded a notebook that uses MediaPipe and my laptop's camera to inference keypoints and I only got the keypoints for all that appeared in the camera frame, the pose part, most of the time it was only two points on the shoulders and 11 points on the face, if the hands appeared in the camera there were also a few points on the hands. I'm not sure if the notebook I'm using is predicting all the pose and just not showing all the keypoints, or if it's only predicting the keypoints in the frame.</p>\n<p>If all the data is done as selfies, then the keypoints data for the rest of the body is really strange and I think we need to experiment to see what difference the results are after removing the points from those parts.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2163441,
          "author_name": "timsereno",
          "author_url": "",
          "post_date": "02/28/2023 20:03:51",
          "content": "<p>I have used mediapipe for a project. I have seen NaN and extrapolated values for both off image and within image.  Maybe not the case with all video, but I've seen it do this with frames in middle of both low and high resolution video at various fps even though the video looks really sharp and seems like the values should have been accurate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2163547,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "02/28/2023 22:59:38",
          "content": "<p>Sure. Let me add it to the eda and I’ll calculate it for every example.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2164564,
              "author_name": "dschettler8845",
              "author_url": "",
              "post_date": "03/01/2023 16:59:33",
              "content": "<p>I updated my EDA to extract NaN counts and calculate the percentage per example as well (i.e. 5 NaN of 100 vs 5 NaN of 1000).</p>\n<p>The following plot shows the general distributions (independently plotted). </p>\n<ul>\n<li><b>Note the Y axis is log scale.</b></li>\n<li>0.0 (Left Side) - Indicates that no values are missing</li>\n<li>1.0 (Right Side) - Indicates that all values are missing</li>\n</ul>\n<p><img src=\"https://i.ibb.co/8KNd8zK/Screenshot-2023-03-01-at-11-53-16-AM.png\"></p>\n<p><b>Quick Takeaways</b></p>\n<ul>\n<li>Face points can be NaN although it is less common than in the Hand data</li>\n<li>Pose points are never NaN</li>\n<li>Left and Right hand distributions are similar but Right Hand is full NaN less than Left Hand</li>\n<li>Pose, Left-Hand, and Right-Hand all have intermediate (not all missing or all present) sequences, however, they are less common than the case where all points are NaN or valid.</li>\n</ul>",
              "votes": null,
              "replies": [
                {
                  "id": 2164600,
                  "author_name": "roberthatch",
                  "author_url": "",
                  "post_date": "03/01/2023 17:27:30",
                  "content": "<p>This is great analysis, thanks!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2163423,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "02/28/2023 19:49:26",
      "content": "<p>Not sure if this helps and more of an FYI:</p>\n<p>I looked at around 10,000 pieces of the data just to gauge… it looks like around 4-5% of the examples have both hands present. (407 out of the 10000 had non-zero sums for LH or RH).</p>\n<hr>\n<p><strong>Here are some examples showing the raw idx (--idx--) and the sum of the LH array and the sum of RH array (LH – RH)</strong></p>\n<pre><code>--68562--\n76.44444  –  317.19916\n\n--59308--\n102.50776  –  202.10736\n\n--84350--\n325.70483  –  32.586136\n\n--44107--\n176.67221  –  57.18132\n\n--70501--\n44.46574  –  199.6751\n\n--75306--\n91.02222  –  244.8219\n\n--37792--\n763.42896  –  203.41107\n\n--65315--\n643.9147  –  182.17331\n\n--25109--\n404.81876  –  111.70086\n\n--79022--\n267.5301  –  25.0708\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2163450,
          "author_name": "chenlin1999",
          "author_url": "",
          "post_date": "02/28/2023 20:11:20",
          "content": "<p>I guess this is because the MediaPipe prediction is wrong and identifies one hand as two, perhaps because the signer is moving too fast, resulting in a virtual image. In the example I gave above, there are only two discrete frames in the video where the two hands appear, and Rob has drawn the hand in one of the frames. If you look carefully you will see that the two hands are oriented exactly the same way, rather than symmetrically left and right, and I don't think it would be possible to flip the other hand anyway, which would require the fingers of the other hand to make a reverse fist😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2160703": ">Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign. \n\nIt says here that signers must press and hold the screen continuously while recording, and I have the following questions.\n\n1. Does this mean that all 250 words can be done with one hand?\n\n2. Why does it sometimes appear that both hands are present at the same time? For example,  [here](https://www.kaggle.com/code/robikscube/sign-language-recognition-eda-twitch-stream), the shhh action in frames 17 and 19 is done by two hands, is this because mediapipe incorrectly identifies one hand as two hands in these two frames?\n\n3. What device was the data recorded with? Was it taken by the phone's front camera that the participant had to hold down? Or was it a different camera?\n\n4. What angle was the data recorded with? Do all participants take the selfie with one hand on the phone and sign with the other hand? Or will the phone be set up in the proper position with a stand and then have to reach out and hold the screen while recording?\n\nSince the 33 keypoints of the pose are usually there, I assume there is another camera set up in front of the participant to capture the whole body. But then why would there usually only be one hand? I would like to get further confirmation.",
    "2160743": "Great question. My only thought would be they had someone else record it?\n\nDefinitely interested to hear what the hosts have to say though. Thanks for raising this!",
    "2160782": "All videos were recorded with the selfie camera.  We asked that all of them be recorded with one hand.",
    "2161629": "Randomly picking various \"signs\" and reviewing some of the recorded key-points shows to me that the dataset is very \"raw\". The \"hands\" landmarks are missing a lot. The recording FPS seem to vary a lot even for the same participants.\n\nRandom sample: \n- sign \"TV\"\t\n- number of samples for participant: 21\n- shortest recording size: 97385 bytes\n- mean recording size: 760217.38 bytes\n- longest recording size: 2402799 bytes\n\nThat's 24x difference in recording size from smallest to longest recording for the same sign!! \n\nI'm personally curious if the videos dataset, test-set had been reviewed and if the labels had been confirmed. Or, is it \"self\" curated and the labels where assigned by persons that recorded the signs.",
    "2163375": "\"Since the 33 keypoints of the pose are usually there\"\n\nI haven't done any EDA myself, but I got the impression from others EDA that each of the four types are all-or-nothing in terms of NaN values(?) If so, there might be some MediaPipe logic that extrapolates the 'off-camera position' of all the pose elements that aren't on camera?\n\nIn the end, I think this is another important question. Although I doubt we care about the foot, using it as an example, how is the right foot handled (when off camera) by MediaPipe pose landmark processing?\n\n@dschettler8845 care to run some EDA on the dataset given, and/or maybe on picture-to-mediapipe results, in terms of NaN vs off-camera coordinates?",
    "2163423": "Not sure if this helps and more of an FYI:\n\nI looked at around 10,000 pieces of the data just to gauge... it looks like around 4-5% of the examples have both hands present. (407 out of the 10000 had non-zero sums for LH or RH).\n\n---\n\n**Here are some examples showing the raw idx (--idx--) and the sum of the LH array and the sum of RH array (LH – RH)**\n\n```\n--68562--\n76.44444  –  317.19916\n\n--59308--\n102.50776  –  202.10736\n\n--84350--\n325.70483  –  32.586136\n\n--44107--\n176.67221  –  57.18132\n\n--70501--\n44.46574  –  199.6751\n\n--75306--\n91.02222  –  244.8219\n\n--37792--\n763.42896  –  203.41107\n\n--65315--\n643.9147  –  182.17331\n\n--25109--\n404.81876  –  111.70086\n\n--79022--\n267.5301  –  25.0708\n```",
    "2163434": "I have looked at some EDAs done by others and my impression is that all samples have all the face and pose, Nan is only present in the left and right hands.\n\nI downloaded a notebook that uses MediaPipe and my laptop's camera to inference keypoints and I only got the keypoints for all that appeared in the camera frame, the pose part, most of the time it was only two points on the shoulders and 11 points on the face, if the hands appeared in the camera there were also a few points on the hands. I'm not sure if the notebook I'm using is predicting all the pose and just not showing all the keypoints, or if it's only predicting the keypoints in the frame.\n\nIf all the data is done as selfies, then the keypoints data for the rest of the body is really strange and I think we need to experiment to see what difference the results are after removing the points from those parts.",
    "2163441": "I have used mediapipe for a project. I have seen NaN and extrapolated values for both off image and within image.  Maybe not the case with all video, but I've seen it do this with frames in middle of both low and high resolution video at various fps even though the video looks really sharp and seems like the values should have been accurate.",
    "2163450": "I guess this is because the MediaPipe prediction is wrong and identifies one hand as two, perhaps because the signer is moving too fast, resulting in a virtual image. In the example I gave above, there are only two discrete frames in the video where the two hands appear, and Rob has drawn the hand in one of the frames. If you look carefully you will see that the two hands are oriented exactly the same way, rather than symmetrically left and right, and I don't think it would be possible to flip the other hand anyway, which would require the fingers of the other hand to make a reverse fist😂",
    "2163547": "Sure. Let me add it to the eda and I’ll calculate it for every example.",
    "2164564": "I updated my EDA to extract NaN counts and calculate the percentage per example as well (i.e. 5 NaN of 100 vs 5 NaN of 1000).\n\nThe following plot shows the general distributions (independently plotted). \n* <b>Note the Y axis is log scale.</b>\n* 0.0 (Left Side) - Indicates that no values are missing\n* 1.0 (Right Side) - Indicates that all values are missing\n\n<img src=\"https://i.ibb.co/8KNd8zK/Screenshot-2023-03-01-at-11-53-16-AM.png\" width=100%>\n\n<b>Quick Takeaways</b>\n* Face points can be NaN although it is less common than in the Hand data\n* Pose points are never NaN\n* Left and Right hand distributions are similar but Right Hand is full NaN less than Left Hand\n* Pose, Left-Hand, and Right-Hand all have intermediate (not all missing or all present) sequences, however, they are less common than the case where all points are NaN or valid.",
    "2164600": "This is great analysis, thanks!"
  },
  "source": "meta"
}