{
  "id": 406426,
  "title": "49th place silver solution ",
  "url": "/competitions/asl-signs/writeups/rb-49th-place-silver-solution",
  "author_name": "",
  "post_date": "2023-05-02T09:45:01.501621800Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thank you Kaggle and Pop sign for hosting this competition. I enjoyed working on this problem and learned quite a bit from really good <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265\" target=\"_blank\">discussion posts</a>  and <a href=\"https://www.kaggle.com/datasets/hengck23/asl-demo\" target=\"_blank\">code</a> shared - Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<ul>\n<li>Pytorch 2.0  (pure pytorch no training library)→ ONNX → TFLITE</li>\n<li>Max length i.e. number of frames per video = 256 (Center crop)</li>\n<li>Normalization by LIP, SPOSE, Lhand and Rhand (total 90 kepyoints)</li>\n<li>Flip augmentation (p=0.5)</li>\n<li>Distance, motion and acceleration features  (Total 1050)</li>\n<li>Model almost same as shared <a href=\"https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture]\" target=\"_blank\">here</a>  - with positional encoding 16, added layernorm in FeedForwardNetwork and 2 fully connected layers with batchnorm, CLS token + mean pooling for final fully connected layer</li>\n<li>Loss function cross entropy with label smoothing = 0.75</li>\n<li>Epochs 50 - dropout 0.0 for 15 epochs, dropout 0.4 from 15-35 epochs, dropout 0.2 for the remaining</li>\n<li>FP16 quantization reduced model size significantly - final size ~26MB and ~60 mins for 5 folds</li>\n</ul>\n<p>Things that didn’t work </p>\n<ul>\n<li>External data - It improved by 3-4 ranks but decided to not use it , too messy data and no commercial license</li>\n<li>Adding more features, reducing max length, removing data with too high/low number of frames</li>\n<li>SWA - couldn’t get it to work</li>\n<li>Augment Affine improved CV and private LB but not public LB.</li>\n<li>Various hyper parameters like # of encoder blocks, embedding dimension</li>\n<li>Arcface Loss - Based on this <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406301\" target=\"_blank\">post</a>  it seems Arcface loss + label smoothing didn’t work . Should have tried just Arcface and no label smoothing.</li>\n<li>Also tried to convert TF model <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">here</a> to PyTorch</li>\n</ul>\n<p>Based on private LB scores it seems that data in private LB is much less noisy than training set. </p>\n<hr>\n<p>My code here - <a href=\"https://github.com/rashmibanthia/ASL\" target=\"_blank\">https://github.com/rashmibanthia/ASL</a></p>",
  "messages": [
    {
      "id": "2242453",
      "postDate": "05/02/2023 09:45:01",
      "content": "<p>Thank you Kaggle and Pop sign for hosting this competition. I enjoyed working on this problem and learned quite a bit from really good <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265\" target=\"_blank\">discussion posts</a>  and <a href=\"https://www.kaggle.com/datasets/hengck23/asl-demo\" target=\"_blank\">code</a> shared - Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<ul>\n<li>Pytorch 2.0  (pure pytorch no training library)→ ONNX → TFLITE</li>\n<li>Max length i.e. number of frames per video = 256 (Center crop)</li>\n<li>Normalization by LIP, SPOSE, Lhand and Rhand (total 90 kepyoints)</li>\n<li>Flip augmentation (p=0.5)</li>\n<li>Distance, motion and acceleration features  (Total 1050)</li>\n<li>Model almost same as shared <a href=\"https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture]\" target=\"_blank\">here</a>  - with positional encoding 16, added layernorm in FeedForwardNetwork and 2 fully connected layers with batchnorm, CLS token + mean pooling for final fully connected layer</li>\n<li>Loss function cross entropy with label smoothing = 0.75</li>\n<li>Epochs 50 - dropout 0.0 for 15 epochs, dropout 0.4 from 15-35 epochs, dropout 0.2 for the remaining</li>\n<li>FP16 quantization reduced model size significantly - final size ~26MB and ~60 mins for 5 folds</li>\n</ul>\n<p>Things that didn’t work </p>\n<ul>\n<li>External data - It improved by 3-4 ranks but decided to not use it , too messy data and no commercial license</li>\n<li>Adding more features, reducing max length, removing data with too high/low number of frames</li>\n<li>SWA - couldn’t get it to work</li>\n<li>Augment Affine improved CV and private LB but not public LB.</li>\n<li>Various hyper parameters like # of encoder blocks, embedding dimension</li>\n<li>Arcface Loss - Based on this <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406301\" target=\"_blank\">post</a>  it seems Arcface loss + label smoothing didn’t work . Should have tried just Arcface and no label smoothing.</li>\n<li>Also tried to convert TF model <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">here</a> to PyTorch</li>\n</ul>\n<p>Based on private LB scores it seems that data in private LB is much less noisy than training set. </p>\n<hr>\n<p>My code here - <a href=\"https://github.com/rashmibanthia/ASL\" target=\"_blank\">https://github.com/rashmibanthia/ASL</a></p>",
      "rawMarkdown": "Thank you Kaggle and Pop sign for hosting this competition. I enjoyed working on this problem and learned quite a bit from really good [discussion posts](https://www.kaggle.com/competitions/asl-signs/discussion/391265)  and [code](https://www.kaggle.com/datasets/hengck23/asl-demo) shared - Thank you @hengck23 \n\n\n- Pytorch 2.0  (pure pytorch no training library)→ ONNX → TFLITE\n- Max length i.e. number of frames per video = 256 (Center crop)\n- Normalization by LIP, SPOSE, Lhand and Rhand (total 90 kepyoints)\n- Flip augmentation (p=0.5)\n- Distance, motion and acceleration features  (Total 1050)\n- Model almost same as shared [here](https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture])  - with positional encoding 16, added layernorm in FeedForwardNetwork and 2 fully connected layers with batchnorm, CLS token + mean pooling for final fully connected layer\n- Loss function cross entropy with label smoothing = 0.75\n- Epochs 50 - dropout 0.0 for 15 epochs, dropout 0.4 from 15-35 epochs, dropout 0.2 for the remaining\n- FP16 quantization reduced model size significantly - final size ~26MB and ~60 mins for 5 folds\n\nThings that didn’t work \n\n- External data - It improved by 3-4 ranks but decided to not use it , too messy data and no commercial license\n- Adding more features, reducing max length, removing data with too high/low number of frames\n- SWA - couldn’t get it to work\n- Augment Affine improved CV and private LB but not public LB.\n- Various hyper parameters like # of encoder blocks, embedding dimension\n- Arcface Loss - Based on this [post](https://www.kaggle.com/competitions/asl-signs/discussion/406301)  it seems Arcface loss + label smoothing didn’t work . Should have tried just Arcface and no label smoothing.\n- Also tried to convert TF model [here](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training) to PyTorch\n\nBased on private LB scores it seems that data in private LB is much less noisy than training set. \n\n----\n\nMy code here - https://github.com/rashmibanthia/ASL",
      "votes": null
    },
    {
      "id": "2242597",
      "postDate": "05/02/2023 11:44:43",
      "content": "<p>Superb Finish ☺️</p>",
      "rawMarkdown": "Superb Finish ☺️",
      "votes": null
    },
    {
      "id": "2242700",
      "postDate": "05/02/2023 13:07:56",
      "content": "<p>Thank you :) </p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "2242927",
      "postDate": "05/02/2023 15:25:31",
      "content": "<p><a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> congrats with silver!🎉 Thank you for sharing your code and explanation! 👍 It helps a lot🙂 </p>",
      "rawMarkdown": "rashmibanthia congrats with silver!🎉 Thank you for sharing your code and explanation! 👍 It helps a lot🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242597,
      "author_name": "aman1391",
      "author_url": "",
      "post_date": "05/02/2023 11:44:43",
      "content": "<p>Superb Finish ☺️</p>",
      "votes": null,
      "replies": [
        {
          "id": 2242700,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "05/02/2023 13:07:56",
          "content": "<p>Thank you :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2242927,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/02/2023 15:25:31",
      "content": "<p><a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> congrats with silver!🎉 Thank you for sharing your code and explanation! 👍 It helps a lot🙂 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2242453": "Thank you Kaggle and Pop sign for hosting this competition. I enjoyed working on this problem and learned quite a bit from really good [discussion posts](https://www.kaggle.com/competitions/asl-signs/discussion/391265)  and [code](https://www.kaggle.com/datasets/hengck23/asl-demo) shared - Thank you @hengck23 \n\n\n- Pytorch 2.0  (pure pytorch no training library)→ ONNX → TFLITE\n- Max length i.e. number of frames per video = 256 (Center crop)\n- Normalization by LIP, SPOSE, Lhand and Rhand (total 90 kepyoints)\n- Flip augmentation (p=0.5)\n- Distance, motion and acceleration features  (Total 1050)\n- Model almost same as shared [here](https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture])  - with positional encoding 16, added layernorm in FeedForwardNetwork and 2 fully connected layers with batchnorm, CLS token + mean pooling for final fully connected layer\n- Loss function cross entropy with label smoothing = 0.75\n- Epochs 50 - dropout 0.0 for 15 epochs, dropout 0.4 from 15-35 epochs, dropout 0.2 for the remaining\n- FP16 quantization reduced model size significantly - final size ~26MB and ~60 mins for 5 folds\n\nThings that didn’t work \n\n- External data - It improved by 3-4 ranks but decided to not use it , too messy data and no commercial license\n- Adding more features, reducing max length, removing data with too high/low number of frames\n- SWA - couldn’t get it to work\n- Augment Affine improved CV and private LB but not public LB.\n- Various hyper parameters like # of encoder blocks, embedding dimension\n- Arcface Loss - Based on this [post](https://www.kaggle.com/competitions/asl-signs/discussion/406301)  it seems Arcface loss + label smoothing didn’t work . Should have tried just Arcface and no label smoothing.\n- Also tried to convert TF model [here](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training) to PyTorch\n\nBased on private LB scores it seems that data in private LB is much less noisy than training set. \n\n----\n\nMy code here - https://github.com/rashmibanthia/ASL",
    "2242597": "Superb Finish ☺️",
    "2242700": "Thank you :)",
    "2242927": "rashmibanthia congrats with silver!🎉 Thank you for sharing your code and explanation! 👍 It helps a lot🙂"
  },
  "source": "meta"
}