{
  "id": 391394,
  "title": "Do we really need all the face marks?",
  "url": "/competitions/asl-signs/discussion/391394",
  "author_name": "",
  "post_date": "2023-03-01T10:45:13.716539Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I agree that the information about where the head is can be helpful, but aren't the landmarks too many? <br>\nWe don't actually need information about facial expressions, the data are not sentences but words, and are recorded only for the training, which means there is no emotional expression in them. I'm carrying out an experiment only with the mean coordinate of face landmarks and it had better performance during training &amp; validation step, not to mention that it was much faster.<br>\nI'm having a error in the submission step so I can't check the test result.<br>\nI just wanted to inform you that you may try using fewer face data.<br>\nI'm not good at English. so please understand that the sentences can be somewhat unnatural.</p>",
  "messages": [
    {
      "id": "2164154",
      "postDate": "03/01/2023 10:45:13",
      "content": "<p>I agree that the information about where the head is can be helpful, but aren't the landmarks too many? <br>\nWe don't actually need information about facial expressions, the data are not sentences but words, and are recorded only for the training, which means there is no emotional expression in them. I'm carrying out an experiment only with the mean coordinate of face landmarks and it had better performance during training &amp; validation step, not to mention that it was much faster.<br>\nI'm having a error in the submission step so I can't check the test result.<br>\nI just wanted to inform you that you may try using fewer face data.<br>\nI'm not good at English. so please understand that the sentences can be somewhat unnatural.</p>",
      "rawMarkdown": "I agree that the information about where the head is can be helpful, but aren't the landmarks too many? \nWe don't actually need information about facial expressions, the data are not sentences but words, and are recorded only for the training, which means there is no emotional expression in them. I'm carrying out an experiment only with the mean coordinate of face landmarks and it had better performance during training & validation step, not to mention that it was much faster.\nI'm having a error in the submission step so I can't check the test result.\nI just wanted to inform you that you may try using fewer face data.\nI'm not good at English. so please understand that the sentences can be somewhat unnatural.",
      "votes": null
    },
    {
      "id": "2164800",
      "postDate": "03/01/2023 19:57:28",
      "content": "<p>I agree. Mostly.</p>\n<p>My current model takes the mean of all 468 and reduces to a single (x,y,z) value.</p>\n<p>That said, if memory and training budget are unlimited, using more data generally won't hurt and can give small improvements. So the main reason to throw away data - even supposedly useless data - is to optimize for any data size constraints, and/or to speed up training. The risk, especially early on, is it might seem fine (no reduction in CV and LB), but much later as you try to refine the model and other features, it's technically possible that the stripped out features could've been helpful with the more refined model.</p>\n<p>It could also be that the data is <em>almost</em> useless, but is completely critical for very specific corner cases. Maybe aligning/correcting the pose face data is important, even though you don't need MORE data points. Or maybe detecting 'hand touches face at face location N' is totally critical even though it rarely comes up. And perhaps hand touches face can benefit from just about any facial landmark, since the touch could be at any of them?</p>",
      "rawMarkdown": "I agree. Mostly.\n\nMy current model takes the mean of all 468 and reduces to a single (x,y,z) value.\n\nThat said, if memory and training budget are unlimited, using more data generally won't hurt and can give small improvements. So the main reason to throw away data - even supposedly useless data - is to optimize for any data size constraints, and/or to speed up training. The risk, especially early on, is it might seem fine (no reduction in CV and LB), but much later as you try to refine the model and other features, it's technically possible that the stripped out features could've been helpful with the more refined model.\n\nIt could also be that the data is *almost* useless, but is completely critical for very specific corner cases. Maybe aligning/correcting the pose face data is important, even though you don't need MORE data points. Or maybe detecting 'hand touches face at face location N' is totally critical even though it rarely comes up. And perhaps hand touches face can benefit from just about any facial landmark, since the touch could be at any of them?",
      "votes": null
    },
    {
      "id": "2165390",
      "postDate": "03/02/2023 07:09:28",
      "content": "<p>Many people who sign also say/mouth the words.  If the signers here are doing it (and at least some of them are by the looks of the animations) then the models might pick up extra clues by lip reading.  However, I appreciate that it might just be noise at the moment.</p>",
      "rawMarkdown": "Many people who sign also say/mouth the words.  If the signers here are doing it (and at least some of them are by the looks of the animations) then the models might pick up extra clues by lip reading.  However, I appreciate that it might just be noise at the moment.",
      "votes": null
    },
    {
      "id": "2166297",
      "postDate": "03/02/2023 18:12:44",
      "content": "<p>I've just run some experiments.  Removing all the face data increases my accuracy by 2 percentages points and adding just the lips back again increases it by a further 1 percentage point.</p>\n<p>Taking a mean of the all the face points (like <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> does) didn't help me, even when I wasn't using the lip points.</p>",
      "rawMarkdown": "I've just run some experiments.  Removing all the face data increases my accuracy by 2 percentages points and adding just the lips back again increases it by a further 1 percentage point.\n\nTaking a mean of the all the face points (like @roberthatch does) didn't help me, even when I wasn't using the lip points.",
      "votes": null
    },
    {
      "id": "2166795",
      "postDate": "03/03/2023 02:57:07",
      "content": "<p>I am still not sure if the sign language in this game uses other body parts besides hands but it sounds rational to assume that face data can be essential in specific cases. Thanks for your response!</p>",
      "rawMarkdown": "I am still not sure if the sign language in this game uses other body parts besides hands but it sounds rational to assume that face data can be essential in specific cases. Thanks for your response!",
      "votes": null
    },
    {
      "id": "2166805",
      "postDate": "03/03/2023 03:17:15",
      "content": "<p>Since pop signs game is also designed to educate non-deaf adults who have deaf children, it is likely that those who play it will also say the word while signing to memorize it better. Adding just the lips seems like a great idea to figure out the data. I should run the model only with face data or use emotion detection to get emotion data and run with it so that i can get better understanding of the data.<br>\nThanks for your response!</p>",
      "rawMarkdown": "Since pop signs game is also designed to educate non-deaf adults who have deaf children, it is likely that those who play it will also say the word while signing to memorize it better. Adding just the lips seems like a great idea to figure out the data. I should run the model only with face data or use emotion detection to get emotion data and run with it so that i can get better understanding of the data.\nThanks for your response!",
      "votes": null
    },
    {
      "id": "2173284",
      "postDate": "03/08/2023 09:11:43",
      "content": "<p>Ten landmarks about the face are alreay contained in the <a href=\"https://google.github.io/mediapipe/solutions/pose.html#pose-landmark-model-blazepose-ghum-3d\" target=\"_blank\">pose data</a>. If you keep those and don't care about lip movement I think you can safely drop all face landmarks.</p>",
      "rawMarkdown": "Ten landmarks about the face are alreay contained in the [pose data](https://google.github.io/mediapipe/solutions/pose.html#pose-landmark-model-blazepose-ghum-3d). If you keep those and don't care about lip movement I think you can safely drop all face landmarks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2164800,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "03/01/2023 19:57:28",
      "content": "<p>I agree. Mostly.</p>\n<p>My current model takes the mean of all 468 and reduces to a single (x,y,z) value.</p>\n<p>That said, if memory and training budget are unlimited, using more data generally won't hurt and can give small improvements. So the main reason to throw away data - even supposedly useless data - is to optimize for any data size constraints, and/or to speed up training. The risk, especially early on, is it might seem fine (no reduction in CV and LB), but much later as you try to refine the model and other features, it's technically possible that the stripped out features could've been helpful with the more refined model.</p>\n<p>It could also be that the data is <em>almost</em> useless, but is completely critical for very specific corner cases. Maybe aligning/correcting the pose face data is important, even though you don't need MORE data points. Or maybe detecting 'hand touches face at face location N' is totally critical even though it rarely comes up. And perhaps hand touches face can benefit from just about any facial landmark, since the touch could be at any of them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2166795,
          "author_name": "zckoaxg",
          "author_url": "",
          "post_date": "03/03/2023 02:57:07",
          "content": "<p>I am still not sure if the sign language in this game uses other body parts besides hands but it sounds rational to assume that face data can be essential in specific cases. Thanks for your response!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2165390,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "03/02/2023 07:09:28",
      "content": "<p>Many people who sign also say/mouth the words.  If the signers here are doing it (and at least some of them are by the looks of the animations) then the models might pick up extra clues by lip reading.  However, I appreciate that it might just be noise at the moment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2166297,
          "author_name": "andrewrrose",
          "author_url": "",
          "post_date": "03/02/2023 18:12:44",
          "content": "<p>I've just run some experiments.  Removing all the face data increases my accuracy by 2 percentages points and adding just the lips back again increases it by a further 1 percentage point.</p>\n<p>Taking a mean of the all the face points (like <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> does) didn't help me, even when I wasn't using the lip points.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2166805,
              "author_name": "zckoaxg",
              "author_url": "",
              "post_date": "03/03/2023 03:17:15",
              "content": "<p>Since pop signs game is also designed to educate non-deaf adults who have deaf children, it is likely that those who play it will also say the word while signing to memorize it better. Adding just the lips seems like a great idea to figure out the data. I should run the model only with face data or use emotion detection to get emotion data and run with it so that i can get better understanding of the data.<br>\nThanks for your response!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2173284,
      "author_name": "lcdrdata",
      "author_url": "",
      "post_date": "03/08/2023 09:11:43",
      "content": "<p>Ten landmarks about the face are alreay contained in the <a href=\"https://google.github.io/mediapipe/solutions/pose.html#pose-landmark-model-blazepose-ghum-3d\" target=\"_blank\">pose data</a>. If you keep those and don't care about lip movement I think you can safely drop all face landmarks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2164154": "I agree that the information about where the head is can be helpful, but aren't the landmarks too many? \nWe don't actually need information about facial expressions, the data are not sentences but words, and are recorded only for the training, which means there is no emotional expression in them. I'm carrying out an experiment only with the mean coordinate of face landmarks and it had better performance during training & validation step, not to mention that it was much faster.\nI'm having a error in the submission step so I can't check the test result.\nI just wanted to inform you that you may try using fewer face data.\nI'm not good at English. so please understand that the sentences can be somewhat unnatural.",
    "2164800": "I agree. Mostly.\n\nMy current model takes the mean of all 468 and reduces to a single (x,y,z) value.\n\nThat said, if memory and training budget are unlimited, using more data generally won't hurt and can give small improvements. So the main reason to throw away data - even supposedly useless data - is to optimize for any data size constraints, and/or to speed up training. The risk, especially early on, is it might seem fine (no reduction in CV and LB), but much later as you try to refine the model and other features, it's technically possible that the stripped out features could've been helpful with the more refined model.\n\nIt could also be that the data is *almost* useless, but is completely critical for very specific corner cases. Maybe aligning/correcting the pose face data is important, even though you don't need MORE data points. Or maybe detecting 'hand touches face at face location N' is totally critical even though it rarely comes up. And perhaps hand touches face can benefit from just about any facial landmark, since the touch could be at any of them?",
    "2165390": "Many people who sign also say/mouth the words.  If the signers here are doing it (and at least some of them are by the looks of the animations) then the models might pick up extra clues by lip reading.  However, I appreciate that it might just be noise at the moment.",
    "2166297": "I've just run some experiments.  Removing all the face data increases my accuracy by 2 percentages points and adding just the lips back again increases it by a further 1 percentage point.\n\nTaking a mean of the all the face points (like @roberthatch does) didn't help me, even when I wasn't using the lip points.",
    "2166795": "I am still not sure if the sign language in this game uses other body parts besides hands but it sounds rational to assume that face data can be essential in specific cases. Thanks for your response!",
    "2166805": "Since pop signs game is also designed to educate non-deaf adults who have deaf children, it is likely that those who play it will also say the word while signing to memorize it better. Adding just the lips seems like a great idea to figure out the data. I should run the model only with face data or use emotion detection to get emotion data and run with it so that i can get better understanding of the data.\nThanks for your response!",
    "2173284": "Ten landmarks about the face are alreay contained in the [pose data](https://google.github.io/mediapipe/solutions/pose.html#pose-landmark-model-blazepose-ghum-3d). If you keep those and don't care about lip movement I think you can safely drop all face landmarks."
  },
  "source": "meta"
}