{
  "id": 406659,
  "title": "26th place solution",
  "url": "/competitions/asl-signs/writeups/danylo-kasianenko-26th-place-solution",
  "author_name": "",
  "post_date": "2023-05-03T08:45:07.616155600Z",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<h2><a href=\"https://github.com/dajes/asl-signs\" target=\"_blank\">Github with all the code used</a></h2>\n<h1>Summary</h1>\n<p>The most important part of the solution is the data utilization. Major improvements were from keypoints choice and mixup. External data does not help because it is from a very different distribution. Given data amount does not benefit larger models so ensembles of small models is the way to utilize given constraints to the fullest.</p>\n<p>Most augmentations are not helpful, because they prevent model from learning the true data distribution. So only used mirroring and mixup (0.5).</p>\n<h1>Inputs to the model</h1>\n<p>All models are trained to support sequences of up to 512 frames.</p>\n<h2>Preprocessing</h2>\n<ul>\n<li>Only 2d coordinates are used as 3rd dimension leads to unstable training.</li>\n<li>To normalize inputs all keypoints are shifted so that head is located at the origin.</li>\n<li>Scaling did not provide any benefit so not used.</li>\n<li>All nans are replaced with 0 after normalization.</li>\n</ul>\n<h2>Chosen keypoints</h2>\n<ul>\n<li>All (21) hand keypoints</li>\n<li>26 face keypoints</li>\n<li>17 pose keypoints<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Ff5ab58ff94d879f12b65cdd7b06e5d57%2Fbody_keypoints.png?generation=1683103187741075&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Fb8b0058e8ce4f760fb3a0d229ffef1e3%2Fface_keypoints.png?generation=1683103204162958&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h1>Architecture</h1>\n<ul>\n<li><p>LLaMa-inspired architecture. Most notable improvement comes from much better normalization RMSNorm.</p></li>\n<li><p>For all models head dimensions are set to 64</p></li>\n<li><p>Single model (Private/Public LB: 0.8543689/0.7702471)</p>\n<p>6 heads 5 layers 9.2M parameters</p></li>\n<li><p>Ensemble of 3 models (Private/Public LB: 0.8584568/0.7725324)</p>\n<p>2 heads 6 layers 1.7M parameters per model</p></li>\n</ul>\n<p>Larger models could be fit into file size limit, but it would time out during submission.</p>\n<h1>Augmentations</h1>\n<ul>\n<li>The most important is mixup. Significantly improves accuracy and allows for a much longer train before overfitting. Using with alpha=0.5.</li>\n<li>Mirroring, does not show much benefit on CV but can't hurt to use.</li>\n</ul>\n<h2>Augmentations that did not work:</h2>\n<ul>\n<li>Interpolation over time. Tried using linear and spline interpolation to change the speed of the video and sample at sub-frame coordinates.</li>\n<li>Shifting keypoints. Tried shifting keypoints in groups and individually by a small amounts.</li>\n<li>Rotating keypoints. Tried rotating keypoints around the head.</li>\n<li>Scaling keypoints. Tried scaling keypoint distance to the head.</li>\n<li>Mixup of 3+ datapoints. Did not have time to tune the alpha but from the start it did not show any benefit. So stayed with default 2 datapoints mixing.</li>\n<li>Rotating and scaling datapoints that are to be mixed up so that it would only mix relative pose and not absolute position.</li>\n</ul>\n<h1>Training</h1>\n<ul>\n<li>Lion optimizer is used. Faster convergence and better result (+1%) at the end.</li>\n<li>LR is 1e-4 with cosine decay to 1e-5 at the end of training.</li>\n<li>Batch size 128</li>\n<li>Dropout 0.1</li>\n<li>Exponential moving average of weights is used to improve stability of training. Gives 2-4% accuracy boost on CV.</li>\n</ul>\n<h2>Did not work:</h2>\n<ul>\n<li>Training model of the data (i.e. trying to predict the next keypoint locations). Tried using different frequency of training the next token predictor and the best frequency turned out to be 0.</li>\n<li>Label smoothing. Slower convergence and worse final result on CV.</li>\n</ul>",
  "messages": [
    {
      "id": "2243875",
      "postDate": "05/03/2023 08:45:07",
      "content": "<h2><a href=\"https://github.com/dajes/asl-signs\" target=\"_blank\">Github with all the code used</a></h2>\n<h1>Summary</h1>\n<p>The most important part of the solution is the data utilization. Major improvements were from keypoints choice and mixup. External data does not help because it is from a very different distribution. Given data amount does not benefit larger models so ensembles of small models is the way to utilize given constraints to the fullest.</p>\n<p>Most augmentations are not helpful, because they prevent model from learning the true data distribution. So only used mirroring and mixup (0.5).</p>\n<h1>Inputs to the model</h1>\n<p>All models are trained to support sequences of up to 512 frames.</p>\n<h2>Preprocessing</h2>\n<ul>\n<li>Only 2d coordinates are used as 3rd dimension leads to unstable training.</li>\n<li>To normalize inputs all keypoints are shifted so that head is located at the origin.</li>\n<li>Scaling did not provide any benefit so not used.</li>\n<li>All nans are replaced with 0 after normalization.</li>\n</ul>\n<h2>Chosen keypoints</h2>\n<ul>\n<li>All (21) hand keypoints</li>\n<li>26 face keypoints</li>\n<li>17 pose keypoints<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Ff5ab58ff94d879f12b65cdd7b06e5d57%2Fbody_keypoints.png?generation=1683103187741075&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Fb8b0058e8ce4f760fb3a0d229ffef1e3%2Fface_keypoints.png?generation=1683103204162958&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h1>Architecture</h1>\n<ul>\n<li><p>LLaMa-inspired architecture. Most notable improvement comes from much better normalization RMSNorm.</p></li>\n<li><p>For all models head dimensions are set to 64</p></li>\n<li><p>Single model (Private/Public LB: 0.8543689/0.7702471)</p>\n<p>6 heads 5 layers 9.2M parameters</p></li>\n<li><p>Ensemble of 3 models (Private/Public LB: 0.8584568/0.7725324)</p>\n<p>2 heads 6 layers 1.7M parameters per model</p></li>\n</ul>\n<p>Larger models could be fit into file size limit, but it would time out during submission.</p>\n<h1>Augmentations</h1>\n<ul>\n<li>The most important is mixup. Significantly improves accuracy and allows for a much longer train before overfitting. Using with alpha=0.5.</li>\n<li>Mirroring, does not show much benefit on CV but can't hurt to use.</li>\n</ul>\n<h2>Augmentations that did not work:</h2>\n<ul>\n<li>Interpolation over time. Tried using linear and spline interpolation to change the speed of the video and sample at sub-frame coordinates.</li>\n<li>Shifting keypoints. Tried shifting keypoints in groups and individually by a small amounts.</li>\n<li>Rotating keypoints. Tried rotating keypoints around the head.</li>\n<li>Scaling keypoints. Tried scaling keypoint distance to the head.</li>\n<li>Mixup of 3+ datapoints. Did not have time to tune the alpha but from the start it did not show any benefit. So stayed with default 2 datapoints mixing.</li>\n<li>Rotating and scaling datapoints that are to be mixed up so that it would only mix relative pose and not absolute position.</li>\n</ul>\n<h1>Training</h1>\n<ul>\n<li>Lion optimizer is used. Faster convergence and better result (+1%) at the end.</li>\n<li>LR is 1e-4 with cosine decay to 1e-5 at the end of training.</li>\n<li>Batch size 128</li>\n<li>Dropout 0.1</li>\n<li>Exponential moving average of weights is used to improve stability of training. Gives 2-4% accuracy boost on CV.</li>\n</ul>\n<h2>Did not work:</h2>\n<ul>\n<li>Training model of the data (i.e. trying to predict the next keypoint locations). Tried using different frequency of training the next token predictor and the best frequency turned out to be 0.</li>\n<li>Label smoothing. Slower convergence and worse final result on CV.</li>\n</ul>",
      "rawMarkdown": "## [Github with all the code used](https://github.com/dajes/asl-signs)\n# Summary\nThe most important part of the solution is the data utilization. Major improvements were from keypoints choice and mixup. External data does not help because it is from a very different distribution. Given data amount does not benefit larger models so ensembles of small models is the way to utilize given constraints to the fullest.\n\nMost augmentations are not helpful, because they prevent model from learning the true data distribution. So only used mirroring and mixup (0.5).\n\n# Inputs to the model\nAll models are trained to support sequences of up to 512 frames.\n\n## Preprocessing\n- Only 2d coordinates are used as 3rd dimension leads to unstable training.\n- To normalize inputs all keypoints are shifted so that head is located at the origin.\n- Scaling did not provide any benefit so not used.\n- All nans are replaced with 0 after normalization.\n## Chosen keypoints\n- All (21) hand keypoints\n- 26 face keypoints\n- 17 pose keypoints\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Ff5ab58ff94d879f12b65cdd7b06e5d57%2Fbody_keypoints.png?generation=1683103187741075&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Fb8b0058e8ce4f760fb3a0d229ffef1e3%2Fface_keypoints.png?generation=1683103204162958&alt=media)\n\n# Architecture\n- LLaMa-inspired architecture. Most notable improvement comes from much better normalization RMSNorm.\n\n- For all models head dimensions are set to 64\n\n- Single model (Private/Public LB: 0.8543689/0.7702471)\n\n    6 heads 5 layers 9.2M parameters\n\n- Ensemble of 3 models (Private/Public LB: 0.8584568/0.7725324)\n\n    2 heads 6 layers 1.7M parameters per model\n\nLarger models could be fit into file size limit, but it would time out during submission.\n\n# Augmentations\n- The most important is mixup. Significantly improves accuracy and allows for a much longer train before overfitting. Using with alpha=0.5.\n- Mirroring, does not show much benefit on CV but can't hurt to use.\n## Augmentations that did not work:\n- Interpolation over time. Tried using linear and spline interpolation to change the speed of the video and sample at sub-frame coordinates.\n- Shifting keypoints. Tried shifting keypoints in groups and individually by a small amounts.\n- Rotating keypoints. Tried rotating keypoints around the head.\n- Scaling keypoints. Tried scaling keypoint distance to the head.\n- Mixup of 3+ datapoints. Did not have time to tune the alpha but from the start it did not show any benefit. So stayed with default 2 datapoints mixing.\n- Rotating and scaling datapoints that are to be mixed up so that it would only mix relative pose and not absolute position.\n# Training\n- Lion optimizer is used. Faster convergence and better result (+1%) at the end.\n- LR is 1e-4 with cosine decay to 1e-5 at the end of training.\n- Batch size 128\n- Dropout 0.1\n- Exponential moving average of weights is used to improve stability of training. Gives 2-4% accuracy boost on CV.\n## Did not work:\n- Training model of the data (i.e. trying to predict the next keypoint locations). Tried using different frequency of training the next token predictor and the best frequency turned out to be 0.\n- Label smoothing. Slower convergence and worse final result on CV.",
      "votes": null
    },
    {
      "id": "2243881",
      "postDate": "05/03/2023 08:52:59",
      "content": "<p>I tried RMSNorm, but it worked way worse than LayerNorm (~-0.02), probably not compatible with other hyperparameters</p>",
      "rawMarkdown": "I tried RMSNorm, but it worked way worse than LayerNorm (~-0.02), probably not compatible with other hyperparameters",
      "votes": null
    },
    {
      "id": "2243925",
      "postDate": "05/03/2023 09:45:32",
      "content": "<p>Did you try BatchNorm1d? It was comparable with RMSNorm and also better than the layer norm. Maybe the problem is in batch size.</p>",
      "rawMarkdown": "Did you try BatchNorm1d? It was comparable with RMSNorm and also better than the layer norm. Maybe the problem is in batch size.",
      "votes": null
    },
    {
      "id": "2243936",
      "postDate": "05/03/2023 09:59:44",
      "content": "<p>We didnt experiment on the architecture much, - most of the work we did was on how to preprocess/augment the data and the loss functions, so no I didn't try it, though it worked for me in the past</p>",
      "rawMarkdown": "We didnt experiment on the architecture much, - most of the work we did was on how to preprocess/augment the data and the loss functions, so no I didn't try it, though it worked for me in the past",
      "votes": null
    },
    {
      "id": "2244780",
      "postDate": "05/03/2023 21:53:24",
      "content": "<p>Congratulations. How did you perform mixup? Spatially, temporary, mixed landmarks, etc? Can you provide more details</p>",
      "rawMarkdown": "Congratulations. How did you perform mixup? Spatially, temporary, mixed landmarks, etc? Can you provide more details",
      "votes": null
    },
    {
      "id": "2248068",
      "postDate": "05/06/2023 13:22:24",
      "content": "<p>Spatially. Picking the longest common interval between 2 datapoints at a random location if one of the datapoints is longer than the other. Masking out attention from parts which are not mixed as a result</p>",
      "rawMarkdown": "Spatially. Picking the longest common interval between 2 datapoints at a random location if one of the datapoints is longer than the other. Masking out attention from parts which are not mixed as a result",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2243881,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "05/03/2023 08:52:59",
      "content": "<p>I tried RMSNorm, but it worked way worse than LayerNorm (~-0.02), probably not compatible with other hyperparameters</p>",
      "votes": null,
      "replies": [
        {
          "id": 2243925,
          "author_name": "kasyanenko",
          "author_url": "",
          "post_date": "05/03/2023 09:45:32",
          "content": "<p>Did you try BatchNorm1d? It was comparable with RMSNorm and also better than the layer norm. Maybe the problem is in batch size.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2243936,
              "author_name": "martynoveduard",
              "author_url": "",
              "post_date": "05/03/2023 09:59:44",
              "content": "<p>We didnt experiment on the architecture much, - most of the work we did was on how to preprocess/augment the data and the loss functions, so no I didn't try it, though it worked for me in the past</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2244780,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/03/2023 21:53:24",
      "content": "<p>Congratulations. How did you perform mixup? Spatially, temporary, mixed landmarks, etc? Can you provide more details</p>",
      "votes": null,
      "replies": [
        {
          "id": 2248068,
          "author_name": "kasyanenko",
          "author_url": "",
          "post_date": "05/06/2023 13:22:24",
          "content": "<p>Spatially. Picking the longest common interval between 2 datapoints at a random location if one of the datapoints is longer than the other. Masking out attention from parts which are not mixed as a result</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2243875": "## [Github with all the code used](https://github.com/dajes/asl-signs)\n# Summary\nThe most important part of the solution is the data utilization. Major improvements were from keypoints choice and mixup. External data does not help because it is from a very different distribution. Given data amount does not benefit larger models so ensembles of small models is the way to utilize given constraints to the fullest.\n\nMost augmentations are not helpful, because they prevent model from learning the true data distribution. So only used mirroring and mixup (0.5).\n\n# Inputs to the model\nAll models are trained to support sequences of up to 512 frames.\n\n## Preprocessing\n- Only 2d coordinates are used as 3rd dimension leads to unstable training.\n- To normalize inputs all keypoints are shifted so that head is located at the origin.\n- Scaling did not provide any benefit so not used.\n- All nans are replaced with 0 after normalization.\n## Chosen keypoints\n- All (21) hand keypoints\n- 26 face keypoints\n- 17 pose keypoints\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Ff5ab58ff94d879f12b65cdd7b06e5d57%2Fbody_keypoints.png?generation=1683103187741075&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2609709%2Fb8b0058e8ce4f760fb3a0d229ffef1e3%2Fface_keypoints.png?generation=1683103204162958&alt=media)\n\n# Architecture\n- LLaMa-inspired architecture. Most notable improvement comes from much better normalization RMSNorm.\n\n- For all models head dimensions are set to 64\n\n- Single model (Private/Public LB: 0.8543689/0.7702471)\n\n    6 heads 5 layers 9.2M parameters\n\n- Ensemble of 3 models (Private/Public LB: 0.8584568/0.7725324)\n\n    2 heads 6 layers 1.7M parameters per model\n\nLarger models could be fit into file size limit, but it would time out during submission.\n\n# Augmentations\n- The most important is mixup. Significantly improves accuracy and allows for a much longer train before overfitting. Using with alpha=0.5.\n- Mirroring, does not show much benefit on CV but can't hurt to use.\n## Augmentations that did not work:\n- Interpolation over time. Tried using linear and spline interpolation to change the speed of the video and sample at sub-frame coordinates.\n- Shifting keypoints. Tried shifting keypoints in groups and individually by a small amounts.\n- Rotating keypoints. Tried rotating keypoints around the head.\n- Scaling keypoints. Tried scaling keypoint distance to the head.\n- Mixup of 3+ datapoints. Did not have time to tune the alpha but from the start it did not show any benefit. So stayed with default 2 datapoints mixing.\n- Rotating and scaling datapoints that are to be mixed up so that it would only mix relative pose and not absolute position.\n# Training\n- Lion optimizer is used. Faster convergence and better result (+1%) at the end.\n- LR is 1e-4 with cosine decay to 1e-5 at the end of training.\n- Batch size 128\n- Dropout 0.1\n- Exponential moving average of weights is used to improve stability of training. Gives 2-4% accuracy boost on CV.\n## Did not work:\n- Training model of the data (i.e. trying to predict the next keypoint locations). Tried using different frequency of training the next token predictor and the best frequency turned out to be 0.\n- Label smoothing. Slower convergence and worse final result on CV.",
    "2243881": "I tried RMSNorm, but it worked way worse than LayerNorm (~-0.02), probably not compatible with other hyperparameters",
    "2243925": "Did you try BatchNorm1d? It was comparable with RMSNorm and also better than the layer norm. Maybe the problem is in batch size.",
    "2243936": "We didnt experiment on the architecture much, - most of the work we did was on how to preprocess/augment the data and the loss functions, so no I didn't try it, though it worked for me in the past",
    "2244780": "Congratulations. How did you perform mixup? Spatially, temporary, mixed landmarks, etc? Can you provide more details",
    "2248068": "Spatially. Picking the longest common interval between 2 datapoints at a random location if one of the datapoints is longer than the other. Masking out attention from parts which are not mixed as a result"
  },
  "source": "meta"
}