{
  "id": 392335,
  "title": "EDA: Handedness by participant ID",
  "url": "/competitions/asl-signs/discussion/392335",
  "author_name": "",
  "post_date": "2023-03-04T19:24:48.375602200Z",
  "votes": 27,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hey all! Let me give you a hand…</p>\n<p>Here is my hand-generated (pun definitely intended) categorization of each signer by which hand they tend to use:</p>\n<pre><code>right_handed_signer = [, , , , , \n                       , , ,  ,  , \n                       , ]\nleft_handed_signer  = [, , , , , \n                       , , , ]\nboth_hands_signer   = [, ]\n\nmessy = [, ]\n</code></pre>\n<p>Notes:</p>\n<ul>\n<li>I counted non-nan for each parquet file in right hand region and left hand region, and printed out rh_count / (rh_count+lh_count), the percentage of right_hand. Usually it's 1.0 or 0.0.</li>\n<li>Most signers use their dominant hand 85% plus, and get noise pretty infrequently (anything higher than 0.0 but lower than 1.0), also maybe 80% without noise, 5-20%(?) of parquet files having noisier result.</li>\n<li>\"Both_hands_signer\" (37055) used each hand to demonstrate each sign, pretty close to 50/50. Maybe 2-6 other signers also clearly switched hands, but visually seemed to use one hand at least 2/3rds of the time, so I still bucketed them as right or as left. Note: I didn't run the program to generate full statistics, so it's technically possible for me to have reached an incorrect conclusion. Or even to have manually put someone in the wrong bucket. Use at your own risk.</li>\n<li>messy: besides categorizing all 21 signers by handedness, I noticed that this signer had a HUGE percentage (75%?) of 'noisy' data. There should be only two possibilities(?): either this signer had very bad video quality so consistently getting mislabeled hands, OR this signer (and only this signer) did two-handed signing, in spite of the directions to use one hand.</li>\n</ul>",
  "messages": [
    {
      "id": "2169105",
      "postDate": "03/04/2023 19:24:48",
      "content": "<p>Hey all! Let me give you a hand…</p>\n<p>Here is my hand-generated (pun definitely intended) categorization of each signer by which hand they tend to use:</p>\n<pre><code>right_handed_signer = [, , , , , \n                       , , ,  ,  , \n                       , ]\nleft_handed_signer  = [, , , , , \n                       , , , ]\nboth_hands_signer   = [, ]\n\nmessy = [, ]\n</code></pre>\n<p>Notes:</p>\n<ul>\n<li>I counted non-nan for each parquet file in right hand region and left hand region, and printed out rh_count / (rh_count+lh_count), the percentage of right_hand. Usually it's 1.0 or 0.0.</li>\n<li>Most signers use their dominant hand 85% plus, and get noise pretty infrequently (anything higher than 0.0 but lower than 1.0), also maybe 80% without noise, 5-20%(?) of parquet files having noisier result.</li>\n<li>\"Both_hands_signer\" (37055) used each hand to demonstrate each sign, pretty close to 50/50. Maybe 2-6 other signers also clearly switched hands, but visually seemed to use one hand at least 2/3rds of the time, so I still bucketed them as right or as left. Note: I didn't run the program to generate full statistics, so it's technically possible for me to have reached an incorrect conclusion. Or even to have manually put someone in the wrong bucket. Use at your own risk.</li>\n<li>messy: besides categorizing all 21 signers by handedness, I noticed that this signer had a HUGE percentage (75%?) of 'noisy' data. There should be only two possibilities(?): either this signer had very bad video quality so consistently getting mislabeled hands, OR this signer (and only this signer) did two-handed signing, in spite of the directions to use one hand.</li>\n</ul>",
      "rawMarkdown": "Hey all! Let me give you a hand...\n\nHere is my hand-generated (pun definitely intended) categorization of each signer by which hand they tend to use:\n\n```python\nright_handed_signer = [26734, 28656, 25571, 62590, 29302, \n                       49445, 53618, 18796,  4718,  2044, \n                       37779, 30680]\nleft_handed_signer  = [16069, 32319, 36257, 22343, 27610, \n                       61333, 34503, 55372, ]\nboth_hands_signer   = [37055, ]\n\nmessy = [29302, ]\n```\n\nNotes:\n* I counted non-nan for each parquet file in right hand region and left hand region, and printed out rh_count / (rh_count+lh_count), the percentage of right_hand. Usually it's 1.0 or 0.0.\n* Most signers use their dominant hand 85% plus, and get noise pretty infrequently (anything higher than 0.0 but lower than 1.0), also maybe 80% without noise, 5-20%(?) of parquet files having noisier result.\n* \"Both_hands_signer\" (37055) used each hand to demonstrate each sign, pretty close to 50/50. Maybe 2-6 other signers also clearly switched hands, but visually seemed to use one hand at least 2/3rds of the time, so I still bucketed them as right or as left. Note: I didn't run the program to generate full statistics, so it's technically possible for me to have reached an incorrect conclusion. Or even to have manually put someone in the wrong bucket. Use at your own risk.\n* messy: besides categorizing all 21 signers by handedness, I noticed that this signer had a HUGE percentage (75%?) of 'noisy' data. There should be only two possibilities(?): either this signer had very bad video quality so consistently getting mislabeled hands, OR this signer (and only this signer) did two-handed signing, in spite of the directions to use one hand.",
      "votes": null
    },
    {
      "id": "2169356",
      "postDate": "03/05/2023 03:43:50",
      "content": "<p>I opened a topic on an idea: how about using only one hand in the model? Since in most cases one of the hands in NaN, how about instead of processing LH and RH we only process ACTIVE_HAND by picking the one hand with data? What do you think?</p>",
      "rawMarkdown": "I opened a topic on an idea: how about using only one hand in the model? Since in most cases one of the hands in NaN, how about instead of processing LH and RH we only process ACTIVE_HAND by picking the one hand with data? What do you think?",
      "votes": null
    },
    {
      "id": "2174095",
      "postDate": "03/08/2023 20:38:09",
      "content": "<p>As a signer would get tired providing data, they might switch hands. I would suggest mirroring all data and training on that. That way you have an equal number of right and left handed.   </p>",
      "rawMarkdown": "As a signer would get tired providing data, they might switch hands. I would suggest mirroring all data and training on that. That way you have an equal number of right and left handed.",
      "votes": null
    },
    {
      "id": "2174165",
      "postDate": "03/08/2023 22:32:42",
      "content": "<p>And probably also use TTA and run inference on both original and mirrored. I agree :)</p>",
      "rawMarkdown": "And probably also use TTA and run inference on both original and mirrored. I agree :)",
      "votes": null
    },
    {
      "id": "2174346",
      "postDate": "03/09/2023 04:45:35",
      "content": "<p>i do some basic read up on ASL. indeed ASL are symmetrical (i.e. either left or right hand can be used).<br>\nBut if i trained with mirror as augmentation, performance dropped. did anyone has the some experience?</p>",
      "rawMarkdown": "i do some basic read up on ASL. indeed ASL are symmetrical (i.e. either left or right hand can be used).\nBut if i trained with mirror as augmentation, performance dropped. did anyone has the some experience?",
      "votes": null
    },
    {
      "id": "2177452",
      "postDate": "03/11/2023 13:53:14",
      "content": "<p>i make a mistake.</p>\n<p>flip x cannot be just -x (assume points are zero mean normalised)<br>\ni need to match the landmark index as well.</p>",
      "rawMarkdown": "i make a mistake.\n\nflip x cannot be just -x (assume points are zero mean normalised)\ni need to match the landmark index as well.",
      "votes": null
    },
    {
      "id": "2192833",
      "postDate": "03/23/2023 00:04:39",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thanks for sharing. Can you explain what you mean by matching the landmark index as well?</p>\n<p>Can be found in the forward function <a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution?scriptVersionId=122239639&amp;cellId=4\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Hi @hengck23, thanks for sharing. Can you explain what you mean by matching the landmark index as well?\n\nCan be found in the forward function [here](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution?scriptVersionId=122239639&cellId=4)",
      "votes": null
    },
    {
      "id": "2194201",
      "postDate": "03/23/2023 19:55:05",
      "content": "<p>If you look at your hands, the right and left hands are mirrors of each other. I think he means, if you mirror the right hand, the reflected x points would no longer match a right hand - they'd match a left hand. So you'd need to change all your right hand landmark points to left hand ones. And vice-a-versa if you reflect a left hand</p>",
      "rawMarkdown": "If you look at your hands, the right and left hands are mirrors of each other. I think he means, if you mirror the right hand, the reflected x points would no longer match a right hand - they'd match a left hand. So you'd need to change all your right hand landmark points to left hand ones. And vice-a-versa if you reflect a left hand",
      "votes": null
    },
    {
      "id": "2238238",
      "postDate": "04/28/2023 10:38:31",
      "content": "<p>Try making things more generalized by using Feature Engineering techniques to use distances and angles between landmark indices of a hand and using only the hand that was active in that particular video as each video should be having only a single hand.</p>",
      "rawMarkdown": "Try making things more generalized by using Feature Engineering techniques to use distances and angles between landmark indices of a hand and using only the hand that was active in that particular video as each video should be having only a single hand.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2169356,
      "author_name": "juanolano",
      "author_url": "",
      "post_date": "03/05/2023 03:43:50",
      "content": "<p>I opened a topic on an idea: how about using only one hand in the model? Since in most cases one of the hands in NaN, how about instead of processing LH and RH we only process ACTIVE_HAND by picking the one hand with data? What do you think?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2174095,
      "author_name": "thadstarner",
      "author_url": "",
      "post_date": "03/08/2023 20:38:09",
      "content": "<p>As a signer would get tired providing data, they might switch hands. I would suggest mirroring all data and training on that. That way you have an equal number of right and left handed.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 2174165,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "03/08/2023 22:32:42",
          "content": "<p>And probably also use TTA and run inference on both original and mirrored. I agree :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2174346,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/09/2023 04:45:35",
              "content": "<p>i do some basic read up on ASL. indeed ASL are symmetrical (i.e. either left or right hand can be used).<br>\nBut if i trained with mirror as augmentation, performance dropped. did anyone has the some experience?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2177452,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "03/11/2023 13:53:14",
                  "content": "<p>i make a mistake.</p>\n<p>flip x cannot be just -x (assume points are zero mean normalised)<br>\ni need to match the landmark index as well.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2192833,
                      "author_name": "brendanartley",
                      "author_url": "",
                      "post_date": "03/23/2023 00:04:39",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thanks for sharing. Can you explain what you mean by matching the landmark index as well?</p>\n<p>Can be found in the forward function <a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution?scriptVersionId=122239639&amp;cellId=4\" target=\"_blank\">here</a></p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2194201,
                          "author_name": "mattt394",
                          "author_url": "",
                          "post_date": "03/23/2023 19:55:05",
                          "content": "<p>If you look at your hands, the right and left hands are mirrors of each other. I think he means, if you mirror the right hand, the reflected x points would no longer match a right hand - they'd match a left hand. So you'd need to change all your right hand landmark points to left hand ones. And vice-a-versa if you reflect a left hand</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2238238,
              "author_name": "uttammittal02",
              "author_url": "",
              "post_date": "04/28/2023 10:38:31",
              "content": "<p>Try making things more generalized by using Feature Engineering techniques to use distances and angles between landmark indices of a hand and using only the hand that was active in that particular video as each video should be having only a single hand.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2169105": "Hey all! Let me give you a hand...\n\nHere is my hand-generated (pun definitely intended) categorization of each signer by which hand they tend to use:\n\n```python\nright_handed_signer = [26734, 28656, 25571, 62590, 29302, \n                       49445, 53618, 18796,  4718,  2044, \n                       37779, 30680]\nleft_handed_signer  = [16069, 32319, 36257, 22343, 27610, \n                       61333, 34503, 55372, ]\nboth_hands_signer   = [37055, ]\n\nmessy = [29302, ]\n```\n\nNotes:\n* I counted non-nan for each parquet file in right hand region and left hand region, and printed out rh_count / (rh_count+lh_count), the percentage of right_hand. Usually it's 1.0 or 0.0.\n* Most signers use their dominant hand 85% plus, and get noise pretty infrequently (anything higher than 0.0 but lower than 1.0), also maybe 80% without noise, 5-20%(?) of parquet files having noisier result.\n* \"Both_hands_signer\" (37055) used each hand to demonstrate each sign, pretty close to 50/50. Maybe 2-6 other signers also clearly switched hands, but visually seemed to use one hand at least 2/3rds of the time, so I still bucketed them as right or as left. Note: I didn't run the program to generate full statistics, so it's technically possible for me to have reached an incorrect conclusion. Or even to have manually put someone in the wrong bucket. Use at your own risk.\n* messy: besides categorizing all 21 signers by handedness, I noticed that this signer had a HUGE percentage (75%?) of 'noisy' data. There should be only two possibilities(?): either this signer had very bad video quality so consistently getting mislabeled hands, OR this signer (and only this signer) did two-handed signing, in spite of the directions to use one hand.",
    "2169356": "I opened a topic on an idea: how about using only one hand in the model? Since in most cases one of the hands in NaN, how about instead of processing LH and RH we only process ACTIVE_HAND by picking the one hand with data? What do you think?",
    "2174095": "As a signer would get tired providing data, they might switch hands. I would suggest mirroring all data and training on that. That way you have an equal number of right and left handed.",
    "2174165": "And probably also use TTA and run inference on both original and mirrored. I agree :)",
    "2174346": "i do some basic read up on ASL. indeed ASL are symmetrical (i.e. either left or right hand can be used).\nBut if i trained with mirror as augmentation, performance dropped. did anyone has the some experience?",
    "2177452": "i make a mistake.\n\nflip x cannot be just -x (assume points are zero mean normalised)\ni need to match the landmark index as well.",
    "2192833": "Hi @hengck23, thanks for sharing. Can you explain what you mean by matching the landmark index as well?\n\nCan be found in the forward function [here](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution?scriptVersionId=122239639&cellId=4)",
    "2194201": "If you look at your hands, the right and left hands are mirrors of each other. I think he means, if you mirror the right hand, the reflected x points would no longer match a right hand - they'd match a left hand. So you'd need to change all your right hand landmark points to left hand ones. And vice-a-versa if you reflect a left hand",
    "2238238": "Try making things more generalized by using Feature Engineering techniques to use distances and angles between landmark indices of a hand and using only the hand that was active in that particular video as each video should be having only a single hand."
  },
  "source": "meta"
}