{
  "id": 392286,
  "title": "Why are x and y coordinates are not normalized in 0 to 1 range?",
  "url": "/competitions/asl-signs/discussion/392286",
  "author_name": "",
  "post_date": "2023-03-04T14:45:42.911426300Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I saw this effect in multiple parquet files. In the media-pipe documentation (<a href=\"https://google.github.io/mediapipe/solutions/hands.html#multi_hand_landmarks\" target=\"_blank\">https://google.github.io/mediapipe/solutions/hands.html#multi_hand_landmarks</a>) its written that the x and y coordinates will be  normalized between 0 and 1, by the input image width and height.</p>\n<p>I loaded this parquet file: <strong>train_landmark_files/26734/1000035562.parquet</strong>, and observed the below output where max value of x and y is more than 1:    </p>\n<pre><code>Shapes of x, y, z are:  (543,) (543,) (543,)\nmin and max of x:  0.1792217493057251 1.181456446647644\nmin and max of y:  0.22510389983654022 2.046358585357666\nmin and max of z:  -2.329070806503296 1.4482454061508179\n</code></pre>",
  "messages": [
    {
      "id": "2168810",
      "postDate": "03/04/2023 14:45:42",
      "content": "<p>I saw this effect in multiple parquet files. In the media-pipe documentation (<a href=\"https://google.github.io/mediapipe/solutions/hands.html#multi_hand_landmarks\" target=\"_blank\">https://google.github.io/mediapipe/solutions/hands.html#multi_hand_landmarks</a>) its written that the x and y coordinates will be  normalized between 0 and 1, by the input image width and height.</p>\n<p>I loaded this parquet file: <strong>train_landmark_files/26734/1000035562.parquet</strong>, and observed the below output where max value of x and y is more than 1:    </p>\n<pre><code>Shapes of x, y, z are:  (543,) (543,) (543,)\nmin and max of x:  0.1792217493057251 1.181456446647644\nmin and max of y:  0.22510389983654022 2.046358585357666\nmin and max of z:  -2.329070806503296 1.4482454061508179\n</code></pre>",
      "rawMarkdown": "I saw this effect in multiple parquet files. In the media-pipe documentation (https://google.github.io/mediapipe/solutions/hands.html#multi_hand_landmarks) its written that the x and y coordinates will be  normalized between 0 and 1, by the input image width and height.\n\nI loaded this parquet file: **train_landmark_files/26734/1000035562.parquet**, and observed the below output where max value of x and y is more than 1:\t\n```\nShapes of x, y, z are:  (543,) (543,) (543,)\nmin and max of x:  0.1792217493057251 1.181456446647644\nmin and max of y:  0.22510389983654022 2.046358585357666\nmin and max of z:  -2.329070806503296 1.4482454061508179\n```",
      "votes": null
    },
    {
      "id": "2168836",
      "postDate": "03/04/2023 15:13:19",
      "content": "<p>&lt; 0 and &gt; 1 might mean extrapolated data off-camera. (Especially for 'pose' landmarks?) But that's just speculation. </p>",
      "rawMarkdown": "< 0 and > 1 might mean extrapolated data off-camera. (Especially for 'pose' landmarks?) But that's just speculation.",
      "votes": null
    },
    {
      "id": "2168870",
      "postDate": "03/04/2023 15:43:05",
      "content": "<p>Thanks, extrapolated off-camera means predicted by the media-pose model ? So, something like regression model but since its not bounded, in few cases it predicts out of training data range ? </p>\n<p>Also, I found this on their github: <a href=\"https://github.com/google/mediapipe/issues/2414#issuecomment-907587575\" target=\"_blank\">https://github.com/google/mediapipe/issues/2414#issuecomment-907587575</a></p>\n<pre><code>no problem，because the input of blazepose is the affined image that centered on mid hip，so back to original image, x or y may be greater than 1\n</code></pre>\n<p>But to be honest, I am a bit new to these terms and I was not able to understand it properly. </p>",
      "rawMarkdown": "Thanks, extrapolated off-camera means predicted by the media-pose model ? So, something like regression model but since its not bounded, in few cases it predicts out of training data range ? \n\nAlso, I found this on their github: https://github.com/google/mediapipe/issues/2414#issuecomment-907587575\n```\nno problem，because the input of blazepose is the affined image that centered on mid hip，so back to original image, x or y may be greater than 1\n```\n\nBut to be honest, I am a bit new to these terms and I was not able to understand it properly.",
      "votes": null
    },
    {
      "id": "2168964",
      "postDate": "03/04/2023 16:59:56",
      "content": "<p>It's just my guess, I don't really know</p>",
      "rawMarkdown": "It's just my guess, I don't really know",
      "votes": null
    },
    {
      "id": "2169062",
      "postDate": "03/04/2023 18:52:16",
      "content": "<p>Got it, makes sense. Thanks for your time. </p>",
      "rawMarkdown": "Got it, makes sense. Thanks for your time.",
      "votes": null
    },
    {
      "id": "2171542",
      "postDate": "03/06/2023 21:36:40",
      "content": "<p>We used the raw mediapipe landmarks so any points that are outside of [0, 1] are Mediapipe artifacts. I believe those are relatively rare, but unfortunately I don't have any more information about the root causes. </p>",
      "rawMarkdown": "We used the raw mediapipe landmarks so any points that are outside of [0, 1] are Mediapipe artifacts. I believe those are relatively rare, but unfortunately I don't have any more information about the root causes.",
      "votes": null
    },
    {
      "id": "2171814",
      "postDate": "03/07/2023 05:36:51",
      "content": "<p>Got it, thanks. </p>",
      "rawMarkdown": "Got it, thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2168836,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "03/04/2023 15:13:19",
      "content": "<p>&lt; 0 and &gt; 1 might mean extrapolated data off-camera. (Especially for 'pose' landmarks?) But that's just speculation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2168870,
          "author_name": "yogendrayatnalkar",
          "author_url": "",
          "post_date": "03/04/2023 15:43:05",
          "content": "<p>Thanks, extrapolated off-camera means predicted by the media-pose model ? So, something like regression model but since its not bounded, in few cases it predicts out of training data range ? </p>\n<p>Also, I found this on their github: <a href=\"https://github.com/google/mediapipe/issues/2414#issuecomment-907587575\" target=\"_blank\">https://github.com/google/mediapipe/issues/2414#issuecomment-907587575</a></p>\n<pre><code>no problem，because the input of blazepose is the affined image that centered on mid hip，so back to original image, x or y may be greater than 1\n</code></pre>\n<p>But to be honest, I am a bit new to these terms and I was not able to understand it properly. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2168964,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "03/04/2023 16:59:56",
              "content": "<p>It's just my guess, I don't really know</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2169062,
                  "author_name": "yogendrayatnalkar",
                  "author_url": "",
                  "post_date": "03/04/2023 18:52:16",
                  "content": "<p>Got it, makes sense. Thanks for your time. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2171542,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "03/06/2023 21:36:40",
      "content": "<p>We used the raw mediapipe landmarks so any points that are outside of [0, 1] are Mediapipe artifacts. I believe those are relatively rare, but unfortunately I don't have any more information about the root causes. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2171814,
          "author_name": "yogendrayatnalkar",
          "author_url": "",
          "post_date": "03/07/2023 05:36:51",
          "content": "<p>Got it, thanks. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2168810": "I saw this effect in multiple parquet files. In the media-pipe documentation (https://google.github.io/mediapipe/solutions/hands.html#multi_hand_landmarks) its written that the x and y coordinates will be  normalized between 0 and 1, by the input image width and height.\n\nI loaded this parquet file: **train_landmark_files/26734/1000035562.parquet**, and observed the below output where max value of x and y is more than 1:\t\n```\nShapes of x, y, z are:  (543,) (543,) (543,)\nmin and max of x:  0.1792217493057251 1.181456446647644\nmin and max of y:  0.22510389983654022 2.046358585357666\nmin and max of z:  -2.329070806503296 1.4482454061508179\n```",
    "2168836": "< 0 and > 1 might mean extrapolated data off-camera. (Especially for 'pose' landmarks?) But that's just speculation.",
    "2168870": "Thanks, extrapolated off-camera means predicted by the media-pose model ? So, something like regression model but since its not bounded, in few cases it predicts out of training data range ? \n\nAlso, I found this on their github: https://github.com/google/mediapipe/issues/2414#issuecomment-907587575\n```\nno problem，because the input of blazepose is the affined image that centered on mid hip，so back to original image, x or y may be greater than 1\n```\n\nBut to be honest, I am a bit new to these terms and I was not able to understand it properly.",
    "2168964": "It's just my guess, I don't really know",
    "2169062": "Got it, makes sense. Thanks for your time.",
    "2171542": "We used the raw mediapipe landmarks so any points that are outside of [0, 1] are Mediapipe artifacts. I believe those are relatively rare, but unfortunately I don't have any more information about the root causes.",
    "2171814": "Got it, thanks."
  },
  "source": "meta"
}