{
  "id": 406457,
  "title": "Model generalization and perfomance plateau",
  "url": "/competitions/asl-signs/discussion/406457",
  "author_name": "Vadim Timakin",
  "post_date": "2023-05-02T12:45:45.713000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was using a transformer based model during this competition. I started with creating basic features, improving my model's architecture, setting up training process and optimization which led to a high results in the first half of the competition. </p>\n<p>I delayed the most promising (as I thought) ideas closer to the end of the competition. Those ideas include: advanced normalization (moving coordinate system to the center and rotating, and scaling it), joint angles, OX-angles, External distances and constant number of frames, etc.</p>\n<p>I was surprised when none of those ideas improved the score, which was interesting enough since all of them must extract more useful information from the data and all of them also led to a significantly higher score on the first part of the training but then all came to the same metric. This observation brings a thought that the model might have very good generalization and learn many patterns by itself. I would be glad to here your thoughts or results of your experiments.</p>",
  "messages": [
    {
      "id": 2242674,
      "postDate": "2023-05-02T12:45:45.713Z",
      "content": "<p>I was using a transformer based model during this competition. I started with creating basic features, improving my model's architecture, setting up training process and optimization which led to a high results in the first half of the competition. </p>\n<p>I delayed the most promising (as I thought) ideas closer to the end of the competition. Those ideas include: advanced normalization (moving coordinate system to the center and rotating, and scaling it), joint angles, OX-angles, External distances and constant number of frames, etc.</p>\n<p>I was surprised when none of those ideas improved the score, which was interesting enough since all of them must extract more useful information from the data and all of them also led to a significantly higher score on the first part of the training but then all came to the same metric. This observation brings a thought that the model might have very good generalization and learn many patterns by itself. I would be glad to here your thoughts or results of your experiments.</p>",
      "rawMarkdown": "I was using a transformer based model during this competition. I started with creating basic features, improving my model's architecture, setting up training process and optimization which led to a high results in the first half of the competition. \n\nI delayed the most promising (as I thought) ideas closer to the end of the competition. Those ideas include: advanced normalization (moving coordinate system to the center and rotating, and scaling it), joint angles, OX-angles, External distances and constant number of frames, etc.\n\nI was surprised when none of those ideas improved the score, which was interesting enough since all of them must extract more useful information from the data and all of them also led to a significantly higher score on the first part of the training but then all came to the same metric. This observation brings a thought that the model might have very good generalization and learn many patterns by itself. I would be glad to here your thoughts or results of your experiments.",
      "votes": 6
    },
    {
      "id": 2242787,
      "postDate": "2023-05-02T14:01:41.723Z",
      "content": "<p>I used the shared transformer with my original hand-crafted features.<br>\nIt resulted in higher score than original shared notebook just by feature engineering. (+ 2~3%)</p>\n<p>Also, robustness of the model is so important because data from human may be so noisy, I think.</p>\n<p>Here are my hand-crafted features.</p>\n<ul>\n<li>distances<br>\nbetween wrist and each fingertip, face and wrist, elbow and wrist.</li>\n<li>angles in dominant hand</li>\n<li>velocity of all raw coordinates</li>\n<li>acceleration of all raw coordinates</li>\n</ul>\n<p>I calculate them in preprocessing layer and input all into Embedding layer in Transformer model.<br>\nMy embedding of transformer consists of 7 Landmark Embedding.</p>\n<pre><code>x_embedded = my_Embedding(lips_, left_hand_, pose_, distance, angle, velocity, acceleration, non_empty_frame_idxs)\n</code></pre>\n<p>each embedding weights result</p>\n<pre><code>lips_embedding weight: %\nleft_hand_embedding weight: %\npose_embedding weight: %\ndistance_embedding weight: %\nangle_embedding weight: %\nvelocity_embedding weight: %\nacceleration_embedding weight: %\n</code></pre>\n<p>thank you.</p>",
      "rawMarkdown": "I used the shared transformer with my original hand-crafted features.\nIt resulted in higher score than original shared notebook just by feature engineering. (+ 2~3%)\n\n\nAlso, robustness of the model is so important because data from human may be so noisy, I think.\n\nHere are my hand-crafted features.\n- distances\nbetween wrist and each fingertip, face and wrist, elbow and wrist.\n- angles in dominant hand\n- velocity of all raw coordinates\n- acceleration of all raw coordinates\n\nI calculate them in preprocessing layer and input all into Embedding layer in Transformer model.\nMy embedding of transformer consists of 7 Landmark Embedding.\n\n\n```python\nx_embedded = my_Embedding(lips_, left_hand_, pose_, distance, angle, velocity, acceleration, non_empty_frame_idxs)\n```\n\n\neach embedding weights result\n```python\nlips_embedding weight: 14.9%\nleft_hand_embedding weight: 24.1%\npose_embedding weight: 13.7%\ndistance_embedding weight: 15.9%\nangle_embedding weight: 13.6%\nvelocity_embedding weight: 10.7%\nacceleration_embedding weight: 7.1%\n```\n  \n\nthank you.\n",
      "votes": 2,
      "replies": [
        {
          "id": 2242901,
          "postDate": "2023-05-02T15:11:40.157Z",
          "content": "<p>Thank you for sharing embedding weights, that's really useful!</p>",
          "rawMarkdown": "Thank you for sharing embedding weights, that's really useful!"
        }
      ]
    },
    {
      "id": 2242778,
      "postDate": "2023-05-02T13:54:11.783Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2242787,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2023-05-02T14:01:41.723000",
      "content": "<p>I used the shared transformer with my original hand-crafted features.<br>\nIt resulted in higher score than original shared notebook just by feature engineering. (+ 2~3%)</p>\n<p>Also, robustness of the model is so important because data from human may be so noisy, I think.</p>\n<p>Here are my hand-crafted features.</p>\n<ul>\n<li>distances<br>\nbetween wrist and each fingertip, face and wrist, elbow and wrist.</li>\n<li>angles in dominant hand</li>\n<li>velocity of all raw coordinates</li>\n<li>acceleration of all raw coordinates</li>\n</ul>\n<p>I calculate them in preprocessing layer and input all into Embedding layer in Transformer model.<br>\nMy embedding of transformer consists of 7 Landmark Embedding.</p>\n<pre><code>x_embedded = my_Embedding(lips_, left_hand_, pose_, distance, angle, velocity, acceleration, non_empty_frame_idxs)\n</code></pre>\n<p>each embedding weights result</p>\n<pre><code>lips_embedding weight: %\nleft_hand_embedding weight: %\npose_embedding weight: %\ndistance_embedding weight: %\nangle_embedding weight: %\nvelocity_embedding weight: %\nacceleration_embedding weight: %\n</code></pre>\n<p>thank you.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2242901,
          "author_name": "Vadim Timakin",
          "author_url": "",
          "post_date": "2023-05-02T15:11:40.157000",
          "content": "<p>Thank you for sharing embedding weights, that's really useful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2242778,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-02T13:54:11.783000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2242674": "I was using a transformer based model during this competition. I started with creating basic features, improving my model's architecture, setting up training process and optimization which led to a high results in the first half of the competition. \n\nI delayed the most promising (as I thought) ideas closer to the end of the competition. Those ideas include: advanced normalization (moving coordinate system to the center and rotating, and scaling it), joint angles, OX-angles, External distances and constant number of frames, etc.\n\nI was surprised when none of those ideas improved the score, which was interesting enough since all of them must extract more useful information from the data and all of them also led to a significantly higher score on the first part of the training but then all came to the same metric. This observation brings a thought that the model might have very good generalization and learn many patterns by itself. I would be glad to here your thoughts or results of your experiments.",
    "2242787": "I used the shared transformer with my original hand-crafted features.\nIt resulted in higher score than original shared notebook just by feature engineering. (+ 2~3%)\n\n\nAlso, robustness of the model is so important because data from human may be so noisy, I think.\n\nHere are my hand-crafted features.\n- distances\nbetween wrist and each fingertip, face and wrist, elbow and wrist.\n- angles in dominant hand\n- velocity of all raw coordinates\n- acceleration of all raw coordinates\n\nI calculate them in preprocessing layer and input all into Embedding layer in Transformer model.\nMy embedding of transformer consists of 7 Landmark Embedding.\n\n\n```python\nx_embedded = my_Embedding(lips_, left_hand_, pose_, distance, angle, velocity, acceleration, non_empty_frame_idxs)\n```\n\n\neach embedding weights result\n```python\nlips_embedding weight: 14.9%\nleft_hand_embedding weight: 24.1%\npose_embedding weight: 13.7%\ndistance_embedding weight: 15.9%\nangle_embedding weight: 13.6%\nvelocity_embedding weight: 10.7%\nacceleration_embedding weight: 7.1%\n```\n  \n\nthank you.\n",
    "2242778": ""
  }
}