{
  "id": 395380,
  "title": "Let's Merge All Landmarks!",
  "url": "/competitions/asl-signs/discussion/395380",
  "author_name": "Bilzard",
  "post_date": "2023-03-17T04:05:28.577000",
  "votes": 34,
  "comment_count": 7,
  "views": 0,
  "content": "<p>To see more intuitive picture of the landmark data, I created another visualization: combining all provided landmarks  of face, pose, and both hands all-in-one.</p>\n<h2>Sample Animations</h2>\n<p>sign=\"apple\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6c7768b93264233622d5f73ce8719b58%2Fapple.gif?generation=1679040207051065&amp;alt=media\" alt=\"\"></p>\n<p>sign=\"airplane\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fa2f507345a2b077e6da0387ff5489719%2Fairplane.gif?generation=1679040223251685&amp;alt=media\" alt=\"\"></p>\n<p>sign=\"TV\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F48be1e13ac29805efec507e502949b0a%2FTV.gif?generation=1679040234516593&amp;alt=media\" alt=\"\"></p>\n<h2>Reference</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/islr-let-s-get-more-intuitive-pictures/notebook\" target=\"_blank\">noteoobk</a> to generate all-in-one animation</li>\n<li><a href=\"https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs-intuitive\" target=\"_blank\">ISLR: Intuitive Animations of 250 signs</a> - you can fully access to the sampled 250 signs where landmarks are combined together</li>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/islr-eda-let-s-get-animated\" target=\"_blank\">previous notebook</a> - where each landmark data are animated separately</li>\n<li><a href=\"https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs\" target=\"_blank\">ISLR: Animated 250 Sampled Signs</a> - fully access to the animations of sampled 250 signs where landmarks are visualized separately</li>\n</ul>",
  "messages": [
    {
      "id": 2185429,
      "postDate": "2023-03-17T04:05:28.577Z",
      "content": "<p>To see more intuitive picture of the landmark data, I created another visualization: combining all provided landmarks  of face, pose, and both hands all-in-one.</p>\n<h2>Sample Animations</h2>\n<p>sign=\"apple\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6c7768b93264233622d5f73ce8719b58%2Fapple.gif?generation=1679040207051065&amp;alt=media\" alt=\"\"></p>\n<p>sign=\"airplane\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fa2f507345a2b077e6da0387ff5489719%2Fairplane.gif?generation=1679040223251685&amp;alt=media\" alt=\"\"></p>\n<p>sign=\"TV\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F48be1e13ac29805efec507e502949b0a%2FTV.gif?generation=1679040234516593&amp;alt=media\" alt=\"\"></p>\n<h2>Reference</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/islr-let-s-get-more-intuitive-pictures/notebook\" target=\"_blank\">noteoobk</a> to generate all-in-one animation</li>\n<li><a href=\"https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs-intuitive\" target=\"_blank\">ISLR: Intuitive Animations of 250 signs</a> - you can fully access to the sampled 250 signs where landmarks are combined together</li>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/islr-eda-let-s-get-animated\" target=\"_blank\">previous notebook</a> - where each landmark data are animated separately</li>\n<li><a href=\"https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs\" target=\"_blank\">ISLR: Animated 250 Sampled Signs</a> - fully access to the animations of sampled 250 signs where landmarks are visualized separately</li>\n</ul>",
      "rawMarkdown": "To see more intuitive picture of the landmark data, I created another visualization: combining all provided landmarks  of face, pose, and both hands all-in-one.\n\n## Sample Animations\n\nsign=\"apple\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6c7768b93264233622d5f73ce8719b58%2Fapple.gif?generation=1679040207051065&alt=media)\n\nsign=\"airplane\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fa2f507345a2b077e6da0387ff5489719%2Fairplane.gif?generation=1679040223251685&alt=media)\n\nsign=\"TV\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F48be1e13ac29805efec507e502949b0a%2FTV.gif?generation=1679040234516593&alt=media)\n\n## Reference\n\n- [noteoobk](https://www.kaggle.com/code/tatamikenn/islr-let-s-get-more-intuitive-pictures/notebook) to generate all-in-one animation\n- [ISLR: Intuitive Animations of 250 signs](https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs-intuitive) - you can fully access to the sampled 250 signs where landmarks are combined together\n- [previous notebook](https://www.kaggle.com/code/tatamikenn/islr-eda-let-s-get-animated) - where each landmark data are animated separately\n- [ISLR: Animated 250 Sampled Signs](https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs) - fully access to the animations of sampled 250 signs where landmarks are visualized separately\n",
      "votes": 34
    },
    {
      "id": 2185503,
      "postDate": "2023-03-17T05:38:18.317Z",
      "content": "<p>actually, it is not just for visualization.<br>\ni think some kaggler has already noted that you would get worst results if you use all 543 landmarks.</p>\n<p>some magic subset work the best, yet it is difficult to select the best subset.</p>\n<p>besides, there is global and local motion. some subset uses local feature (and invariant  to global), some are vice versa.<br>\ni.e. does it look better to show the landmarks in one image or shows them in separate crops (like what you have previously did)</p>\n<p>further, some sign focus on shape only, yet some focus on motion, and others on both shape+motion</p>\n<p>(i am thinking of making a gate/weight/attention to automatically select the best points)</p>",
      "rawMarkdown": "actually, it is not just for visualization.\ni think some kaggler has already noted that you would get worst results if you use all 543 landmarks.\n\nsome magic subset work the best, yet it is difficult to select the best subset.\n\nbesides, there is global and local motion. some subset uses local feature (and invariant  to global), some are vice versa.\ni.e. does it look better to show the landmarks in one image or shows them in separate crops (like what you have previously did)\n\nfurther, some sign focus on shape only, yet some focus on motion, and others on both shape+motion\n\n(i am thinking of making a gate/weight/attention to automatically select the best points)\n ",
      "votes": 3,
      "replies": [
        {
          "id": 2185540,
          "postDate": "2023-03-17T06:23:27.887Z",
          "content": "<p>In my point of view, this competition's labels are really noisy at least on train data.<br>\nI compared actual hand signs and this competitions sequence with my eyes, and found:</p>\n<ol>\n<li>Motions of legs are sometimes really noisy and looks totally meaningless (so I dropped them from visualizations)</li>\n<li>In almost all sequences, only landmarks of main hand is captured. This is the same for even hand signs where both hands should be used.</li>\n<li>Motions of sub hand sometimes becomes really noisy and that shows totally meaningless movement.</li>\n<li>Sometimes motions looks far from the motions of true ASL signs (however, that may be because of lack of my knowledge on ALS signs).</li>\n<li>Some sequences seems incomplete. They sometimes looks like being cut in the middle of the whole sequence: e.g. \"backyard\" consists of the signs of \"back\" and \"yard\", but some clips only shows \"back\" or \"yard\" separately. (I guess it may be automatically trimmed by some inaccurate algorithms.)</li>\n</ol>\n<p>Thus, as you said, I can understand using whole set of landmark for the training easily overfit to the noisy motions or hand shapes. <br>\nPersonally, I even doubt test data might have partly similar data to train data, and LB score just showing the overfitted score. If it isn't, it seems to me really strange that the model could get top1 accuracy over 0.7 by trained with this noisy labels.</p>",
          "rawMarkdown": "In my point of view, this competition's labels are really noisy at least on train data.\nI compared actual hand signs and this competitions sequence with my eyes, and found:\n\n1. Motions of legs are sometimes really noisy and looks totally meaningless (so I dropped them from visualizations)\n2. In almost all sequences, only landmarks of main hand is captured. This is the same for even hand signs where both hands should be used.\n3. Motions of sub hand sometimes becomes really noisy and that shows totally meaningless movement.\n4. Sometimes motions looks far from the motions of true ASL signs (however, that may be because of lack of my knowledge on ALS signs).\n5. Some sequences seems incomplete. They sometimes looks like being cut in the middle of the whole sequence: e.g. \"backyard\" consists of the signs of \"back\" and \"yard\", but some clips only shows \"back\" or \"yard\" separately. (I guess it may be automatically trimmed by some inaccurate algorithms.)\n\nThus, as you said, I can understand using whole set of landmark for the training easily overfit to the noisy motions or hand shapes. \nPersonally, I even doubt test data might have partly similar data to train data, and LB score just showing the overfitted score. If it isn't, it seems to me really strange that the model could get top1 accuracy over 0.7 by trained with this noisy labels.",
          "votes": 3,
          "replies": [
            {
              "id": 2185556,
              "postDate": "2023-03-17T06:45:00.627Z",
              "content": "<p>\"could get top1 accuracy over 0.7 by trained with this noisy labels.\"</p>\n<p>local cv is about 0.80 for pure random split (i.e. same participant in train and valid)</p>",
              "rawMarkdown": "\"could get top1 accuracy over 0.7 by trained with this noisy labels.\"\n\nlocal cv is about 0.80 for pure random split (i.e. same participant in train and valid)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2191731,
      "postDate": "2023-03-22T07:05:02.723Z",
      "content": "<p>According to \"Data Card\" tab, maybe we can safely discard landmark data of one hand.</p>\n<blockquote>\n  <p>Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. The app prompted the signer with the concept in English to sign, randomly selected from the 250-sign vocabulary. <strong>Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign</strong>. The video of the sign is extracted with a buffer 0.5 seconds before the press of the button and 0.5 seconds after the release of the button.</p>\n</blockquote>",
      "rawMarkdown": "According to \"Data Card\" tab, maybe we can safely discard landmark data of one hand.\n\n> Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. The app prompted the signer with the concept in English to sign, randomly selected from the 250-sign vocabulary. **Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign**. The video of the sign is extracted with a buffer 0.5 seconds before the press of the button and 0.5 seconds after the release of the button.",
      "votes": 4,
      "replies": [
        {
          "id": 2206044,
          "postDate": "2023-04-02T07:40:28.350Z",
          "content": "<p>I find it hard to choose which hand to give up. Maybe it would be good to use both hands?</p>",
          "rawMarkdown": "I find it hard to choose which hand to give up. Maybe it would be good to use both hands?"
        }
      ]
    },
    {
      "id": 2212524,
      "postDate": "2023-04-06T20:39:58.780Z",
      "content": "<p>Your visualizations were not only aesthetically pleasing, but they also effectively conveyed complex information in a clear and concise manner. The attention to detail and the creativity that you showed in your work were truly impressive!!</p>",
      "rawMarkdown": "Your visualizations were not only aesthetically pleasing, but they also effectively conveyed complex information in a clear and concise manner. The attention to detail and the creativity that you showed in your work were truly impressive!!",
      "votes": -2
    },
    {
      "id": 2185446,
      "postDate": "2023-03-17T04:27:07.570Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2185503,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-17T05:38:18.317000",
      "content": "<p>actually, it is not just for visualization.<br>\ni think some kaggler has already noted that you would get worst results if you use all 543 landmarks.</p>\n<p>some magic subset work the best, yet it is difficult to select the best subset.</p>\n<p>besides, there is global and local motion. some subset uses local feature (and invariant  to global), some are vice versa.<br>\ni.e. does it look better to show the landmarks in one image or shows them in separate crops (like what you have previously did)</p>\n<p>further, some sign focus on shape only, yet some focus on motion, and others on both shape+motion</p>\n<p>(i am thinking of making a gate/weight/attention to automatically select the best points)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2185540,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-03-17T06:23:27.887000",
          "content": "<p>In my point of view, this competition's labels are really noisy at least on train data.<br>\nI compared actual hand signs and this competitions sequence with my eyes, and found:</p>\n<ol>\n<li>Motions of legs are sometimes really noisy and looks totally meaningless (so I dropped them from visualizations)</li>\n<li>In almost all sequences, only landmarks of main hand is captured. This is the same for even hand signs where both hands should be used.</li>\n<li>Motions of sub hand sometimes becomes really noisy and that shows totally meaningless movement.</li>\n<li>Sometimes motions looks far from the motions of true ASL signs (however, that may be because of lack of my knowledge on ALS signs).</li>\n<li>Some sequences seems incomplete. They sometimes looks like being cut in the middle of the whole sequence: e.g. \"backyard\" consists of the signs of \"back\" and \"yard\", but some clips only shows \"back\" or \"yard\" separately. (I guess it may be automatically trimmed by some inaccurate algorithms.)</li>\n</ol>\n<p>Thus, as you said, I can understand using whole set of landmark for the training easily overfit to the noisy motions or hand shapes. <br>\nPersonally, I even doubt test data might have partly similar data to train data, and LB score just showing the overfitted score. If it isn't, it seems to me really strange that the model could get top1 accuracy over 0.7 by trained with this noisy labels.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2185556,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-17T06:45:00.627000",
              "content": "<p>\"could get top1 accuracy over 0.7 by trained with this noisy labels.\"</p>\n<p>local cv is about 0.80 for pure random split (i.e. same participant in train and valid)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2191731,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2023-03-22T07:05:02.723000",
      "content": "<p>According to \"Data Card\" tab, maybe we can safely discard landmark data of one hand.</p>\n<blockquote>\n  <p>Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. The app prompted the signer with the concept in English to sign, randomly selected from the 250-sign vocabulary. <strong>Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign</strong>. The video of the sign is extracted with a buffer 0.5 seconds before the press of the button and 0.5 seconds after the release of the button.</p>\n</blockquote>",
      "votes": 4,
      "replies": [
        {
          "id": 2206044,
          "author_name": "Jackson You",
          "author_url": "",
          "post_date": "2023-04-02T07:40:28.350000",
          "content": "<p>I find it hard to choose which hand to give up. Maybe it would be good to use both hands?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2212524,
      "author_name": "Othman Bakria",
      "author_url": "",
      "post_date": "2023-04-06T20:39:58.780000",
      "content": "<p>Your visualizations were not only aesthetically pleasing, but they also effectively conveyed complex information in a clear and concise manner. The attention to detail and the creativity that you showed in your work were truly impressive!!</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2185446,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-17T04:27:07.570000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2185429": "To see more intuitive picture of the landmark data, I created another visualization: combining all provided landmarks  of face, pose, and both hands all-in-one.\n\n## Sample Animations\n\nsign=\"apple\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6c7768b93264233622d5f73ce8719b58%2Fapple.gif?generation=1679040207051065&alt=media)\n\nsign=\"airplane\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fa2f507345a2b077e6da0387ff5489719%2Fairplane.gif?generation=1679040223251685&alt=media)\n\nsign=\"TV\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F48be1e13ac29805efec507e502949b0a%2FTV.gif?generation=1679040234516593&alt=media)\n\n## Reference\n\n- [noteoobk](https://www.kaggle.com/code/tatamikenn/islr-let-s-get-more-intuitive-pictures/notebook) to generate all-in-one animation\n- [ISLR: Intuitive Animations of 250 signs](https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs-intuitive) - you can fully access to the sampled 250 signs where landmarks are combined together\n- [previous notebook](https://www.kaggle.com/code/tatamikenn/islr-eda-let-s-get-animated) - where each landmark data are animated separately\n- [ISLR: Animated 250 Sampled Signs](https://www.kaggle.com/datasets/tatamikenn/islr-animation-250-signs) - fully access to the animations of sampled 250 signs where landmarks are visualized separately\n",
    "2185503": "actually, it is not just for visualization.\ni think some kaggler has already noted that you would get worst results if you use all 543 landmarks.\n\nsome magic subset work the best, yet it is difficult to select the best subset.\n\nbesides, there is global and local motion. some subset uses local feature (and invariant  to global), some are vice versa.\ni.e. does it look better to show the landmarks in one image or shows them in separate crops (like what you have previously did)\n\nfurther, some sign focus on shape only, yet some focus on motion, and others on both shape+motion\n\n(i am thinking of making a gate/weight/attention to automatically select the best points)\n ",
    "2191731": "According to \"Data Card\" tab, maybe we can safely discard landmark data of one hand.\n\n> Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. The app prompted the signer with the concept in English to sign, randomly selected from the 250-sign vocabulary. **Signers pressed and held an on-screen button on the phone to record video while signing each concept, releasing the button after each sign**. The video of the sign is extracted with a buffer 0.5 seconds before the press of the button and 0.5 seconds after the release of the button.",
    "2212524": "Your visualizations were not only aesthetically pleasing, but they also effectively conveyed complex information in a clear and concise manner. The attention to detail and the creativity that you showed in your work were truly impressive!!",
    "2185446": ""
  }
}