{
  "id": 406657,
  "title": "11th place solution with code",
  "url": "/competitions/asl-signs/writeups/camaro-11th-place-solution-with-code",
  "author_name": "",
  "post_date": "2023-05-03T23:24:56.077Z",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thank you to the organizer and Kaggle for hosting this interesting challenge.<br>\nEspecially I enjoyed this strict inference time restriction. It keeps model size reasonable and requires us for some practical technique.</p>\n<h2>TL;DR</h2>\n<ul>\n<li>Ensemble 5 transformer models</li>\n<li>Strong augmentation</li>\n<li>Manual model conversion from pytroch to tensorflow</li>\n<li>Code is available here -&gt; <a href=\"https://github.com/bamps53/kaggle-asl-11th-place-solution\" target=\"_blank\">https://github.com/bamps53/kaggle-asl-11th-place-solution</a></li>\n</ul>\n<h2>Overview</h2>\n<p>I started from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ‘s <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265\" target=\"_blank\">great discussion</a> and <a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution\" target=\"_blank\">notebook</a>. Thanks for sharing a lot of useful tricks as always!</p>\n<p>The changes I made are following;</p>\n<ul>\n<li>Change model architecture to CLIP transformer in HuggingFace</li>\n<li>Decrease parameter size to maximize latency within the range of same accuracy</li>\n<li>Some strong augmentations<ul>\n<li>Horizontal flip(p=0.5)</li>\n<li>Random 3d rotation(p=1, -45~45)</li>\n<li>Random scale(p=1, 0.5~1.5)</li>\n<li>Random shift(p=1, 0.7~1.3)</li>\n<li>Random mask frames(p=1, mask_ratio=0.5)</li>\n<li>Random resize (p=1, 0.5~1.5)</li></ul></li>\n<li>Add motion features<ul>\n<li>current - prev</li>\n<li>next - current</li>\n<li>Velocity  </li></ul></li>\n<li>Longer epoch, 250 for 5 fold and 300 for all data</li>\n</ul>\n<p>For the details, please refer to the code.(planning to upload)</p>\n<h2>Model conversion</h2>\n<p>I’m too lazy to implement augmentations in tensorflow dataset, so I keep using pytorch training pipeline. But I’ve realized that inference time of the model converted by onnx_tf is way slower than bare tensorflow models. Then I’ve tried some model conversion framework like nobuco, but there were too many errors for some reason. Finally I’ve decided to write the model architecture both in pytorch and tensorflow, then manually port the weight. Thankfully HuggingFace has both pytorch and tensorflow CLIP implementation, this work is easier than I thought. This significantly speeds up inference time and I can put more models when ensembling.</p>\n<h2>Ensemble</h2>\n<p>I’ve tried to diversify the models as much as possible within the same accuracy range.</p>\n<table>\n<thead>\n<tr>\n<th>seed</th>\n<th>layers</th>\n<th>dim</th>\n<th>act</th>\n<th>max_len</th>\n<th>features</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>2</td>\n<td>384</td>\n<td>relu</td>\n<td>64</td>\n<td>lip/hand</td>\n</tr>\n<tr>\n<td>1</td>\n<td>3</td>\n<td>256</td>\n<td>relu</td>\n<td>48</td>\n<td>lip/hand/eye/motion</td>\n</tr>\n<tr>\n<td>2</td>\n<td>2</td>\n<td>384</td>\n<td>geru</td>\n<td>64</td>\n<td>lip/hand</td>\n</tr>\n<tr>\n<td>3</td>\n<td>2</td>\n<td>384</td>\n<td>geru</td>\n<td>64</td>\n<td>lip/hand/eye/motion</td>\n</tr>\n<tr>\n<td>4</td>\n<td>3</td>\n<td>256</td>\n<td>geru</td>\n<td>48</td>\n<td>lip/hand/eye/motion</td>\n</tr>\n</tbody>\n</table>\n<p>The scores are as follows;</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Single best model</td>\n<td>0.786</td>\n<td>0.865</td>\n</tr>\n<tr>\n<td>Ensemble 5 models</td>\n<td>0.794</td>\n<td>0.87</td>\n</tr>\n</tbody>\n</table>\n<h2>What didn’t work;</h2>\n<ul>\n<li>Normalization matters a lot, so I’ve tried a lot of variants but couldn’t get a better result than simple video mean/std normalization.</li>\n<li>More landmarks(like arm, ear, nose or pose)</li>\n<li>Conv1d</li>\n<li>Knowledge distillation</li>\n<li>Model soup</li>\n<li>Bigger model</li>\n<li>Pretrained model (ex. CLIP pretrained model or pretrain with NTU dataset)</li>\n<li>And so on..</li>\n</ul>\n<h2>Code</h2>\n<p><a href=\"https://github.com/bamps53/kaggle-asl-11th-place-solution\" target=\"_blank\">https://github.com/bamps53/kaggle-asl-11th-place-solution</a></p>",
  "messages": [
    {
      "id": "2243873",
      "postDate": "05/03/2023 08:39:14",
      "content": "<p>Thank you to the organizer and Kaggle for hosting this interesting challenge.<br>\nEspecially I enjoyed this strict inference time restriction. It keeps model size reasonable and requires us for some practical technique.</p>\n<h2>TL;DR</h2>\n<ul>\n<li>Ensemble 5 transformer models</li>\n<li>Strong augmentation</li>\n<li>Manual model conversion from pytroch to tensorflow</li>\n<li>Code is available here -&gt; <a href=\"https://github.com/bamps53/kaggle-asl-11th-place-solution\" target=\"_blank\">https://github.com/bamps53/kaggle-asl-11th-place-solution</a></li>\n</ul>\n<h2>Overview</h2>\n<p>I started from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ‘s <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265\" target=\"_blank\">great discussion</a> and <a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution\" target=\"_blank\">notebook</a>. Thanks for sharing a lot of useful tricks as always!</p>\n<p>The changes I made are following;</p>\n<ul>\n<li>Change model architecture to CLIP transformer in HuggingFace</li>\n<li>Decrease parameter size to maximize latency within the range of same accuracy</li>\n<li>Some strong augmentations<ul>\n<li>Horizontal flip(p=0.5)</li>\n<li>Random 3d rotation(p=1, -45~45)</li>\n<li>Random scale(p=1, 0.5~1.5)</li>\n<li>Random shift(p=1, 0.7~1.3)</li>\n<li>Random mask frames(p=1, mask_ratio=0.5)</li>\n<li>Random resize (p=1, 0.5~1.5)</li></ul></li>\n<li>Add motion features<ul>\n<li>current - prev</li>\n<li>next - current</li>\n<li>Velocity  </li></ul></li>\n<li>Longer epoch, 250 for 5 fold and 300 for all data</li>\n</ul>\n<p>For the details, please refer to the code.(planning to upload)</p>\n<h2>Model conversion</h2>\n<p>I’m too lazy to implement augmentations in tensorflow dataset, so I keep using pytorch training pipeline. But I’ve realized that inference time of the model converted by onnx_tf is way slower than bare tensorflow models. Then I’ve tried some model conversion framework like nobuco, but there were too many errors for some reason. Finally I’ve decided to write the model architecture both in pytorch and tensorflow, then manually port the weight. Thankfully HuggingFace has both pytorch and tensorflow CLIP implementation, this work is easier than I thought. This significantly speeds up inference time and I can put more models when ensembling.</p>\n<h2>Ensemble</h2>\n<p>I’ve tried to diversify the models as much as possible within the same accuracy range.</p>\n<table>\n<thead>\n<tr>\n<th>seed</th>\n<th>layers</th>\n<th>dim</th>\n<th>act</th>\n<th>max_len</th>\n<th>features</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>2</td>\n<td>384</td>\n<td>relu</td>\n<td>64</td>\n<td>lip/hand</td>\n</tr>\n<tr>\n<td>1</td>\n<td>3</td>\n<td>256</td>\n<td>relu</td>\n<td>48</td>\n<td>lip/hand/eye/motion</td>\n</tr>\n<tr>\n<td>2</td>\n<td>2</td>\n<td>384</td>\n<td>geru</td>\n<td>64</td>\n<td>lip/hand</td>\n</tr>\n<tr>\n<td>3</td>\n<td>2</td>\n<td>384</td>\n<td>geru</td>\n<td>64</td>\n<td>lip/hand/eye/motion</td>\n</tr>\n<tr>\n<td>4</td>\n<td>3</td>\n<td>256</td>\n<td>geru</td>\n<td>48</td>\n<td>lip/hand/eye/motion</td>\n</tr>\n</tbody>\n</table>\n<p>The scores are as follows;</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Single best model</td>\n<td>0.786</td>\n<td>0.865</td>\n</tr>\n<tr>\n<td>Ensemble 5 models</td>\n<td>0.794</td>\n<td>0.87</td>\n</tr>\n</tbody>\n</table>\n<h2>What didn’t work;</h2>\n<ul>\n<li>Normalization matters a lot, so I’ve tried a lot of variants but couldn’t get a better result than simple video mean/std normalization.</li>\n<li>More landmarks(like arm, ear, nose or pose)</li>\n<li>Conv1d</li>\n<li>Knowledge distillation</li>\n<li>Model soup</li>\n<li>Bigger model</li>\n<li>Pretrained model (ex. CLIP pretrained model or pretrain with NTU dataset)</li>\n<li>And so on..</li>\n</ul>\n<h2>Code</h2>\n<p><a href=\"https://github.com/bamps53/kaggle-asl-11th-place-solution\" target=\"_blank\">https://github.com/bamps53/kaggle-asl-11th-place-solution</a></p>",
      "rawMarkdown": "Thank you to the organizer and Kaggle for hosting this interesting challenge.\nEspecially I enjoyed this strict inference time restriction. It keeps model size reasonable and requires us for some practical technique.\n\n## TL;DR\n- Ensemble 5 transformer models\n- Strong augmentation\n- Manual model conversion from pytroch to tensorflow\n- Code is available here -> https://github.com/bamps53/kaggle-asl-11th-place-solution\n\n## Overview\nI started from @hengck23 ‘s [great discussion](https://www.kaggle.com/competitions/asl-signs/discussion/391265) and [notebook](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution). Thanks for sharing a lot of useful tricks as always!\n\n\nThe changes I made are following;\n- Change model architecture to CLIP transformer in HuggingFace\n- Decrease parameter size to maximize latency within the range of same accuracy\n- Some strong augmentations\n    -  Horizontal flip(p=0.5)\n    -  Random 3d rotation(p=1, -45~45)\n    -  Random scale(p=1, 0.5~1.5)\n    -  Random shift(p=1, 0.7~1.3)\n    -  Random mask frames(p=1, mask_ratio=0.5)\n    -  Random resize (p=1, 0.5~1.5)\n- Add motion features\n    - current - prev\n    - next - current\n    - Velocity  \n- Longer epoch, 250 for 5 fold and 300 for all data\n\nFor the details, please refer to the code.(planning to upload)\n\n## Model conversion\nI’m too lazy to implement augmentations in tensorflow dataset, so I keep using pytorch training pipeline. But I’ve realized that inference time of the model converted by onnx_tf is way slower than bare tensorflow models. Then I’ve tried some model conversion framework like nobuco, but there were too many errors for some reason. Finally I’ve decided to write the model architecture both in pytorch and tensorflow, then manually port the weight. Thankfully HuggingFace has both pytorch and tensorflow CLIP implementation, this work is easier than I thought. This significantly speeds up inference time and I can put more models when ensembling.\n\n## Ensemble\nI’ve tried to diversify the models as much as possible within the same accuracy range.\n\n| seed | layers | dim | act  | max_len | features            |\n|------|--------|-----|------|---------|---------------------|\n|    0 |      2 | 384 | relu |      64 | lip/hand            |\n|    1 |      3 | 256 | relu |      48 | lip/hand/eye/motion |\n|    2 |      2 | 384 | geru |      64 | lip/hand            |\n|    3 |      2 | 384 | geru |      64 | lip/hand/eye/motion |\n|    4 |      3 | 256 | geru |      48 | lip/hand/eye/motion |\n\n\nThe scores are as follows;\n\n|                   | public | private |\n|-------------------|--------|---------|\n| Single best model |  0.786 |   0.865 |\n| Ensemble 5 models |  0.794 |    0.87 |\n\n\n## What didn’t work;\n- Normalization matters a lot, so I’ve tried a lot of variants but couldn’t get a better result than simple video mean/std normalization.\n- More landmarks(like arm, ear, nose or pose)\n- Conv1d\n- Knowledge distillation\n- Model soup\n- Bigger model\n- Pretrained model (ex. CLIP pretrained model or pretrain with NTU dataset)\n- And so on..\n\n## Code\nhttps://github.com/bamps53/kaggle-asl-11th-place-solution",
      "votes": null
    },
    {
      "id": "2244845",
      "postDate": "05/03/2023 23:25:44",
      "content": "<p>Updated to attach a link to the code:)<br>\n<a href=\"https://github.com/bamps53/kaggle-asl-11th-place-solution\" target=\"_blank\">https://github.com/bamps53/kaggle-asl-11th-place-solution</a></p>",
      "rawMarkdown": "Updated to attach a link to the code:)\nhttps://github.com/bamps53/kaggle-asl-11th-place-solution",
      "votes": null
    },
    {
      "id": "2246980",
      "postDate": "05/05/2023 15:35:36",
      "content": "<p>11th place is a great result!!!<br>\nUsing other people's notebooks really speeds up the workflow</p>",
      "rawMarkdown": "11th place is a great result!!!\nUsing other people's notebooks really speeds up the workflow",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2244845,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "05/03/2023 23:25:44",
      "content": "<p>Updated to attach a link to the code:)<br>\n<a href=\"https://github.com/bamps53/kaggle-asl-11th-place-solution\" target=\"_blank\">https://github.com/bamps53/kaggle-asl-11th-place-solution</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2246980,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/05/2023 15:35:36",
      "content": "<p>11th place is a great result!!!<br>\nUsing other people's notebooks really speeds up the workflow</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2243873": "Thank you to the organizer and Kaggle for hosting this interesting challenge.\nEspecially I enjoyed this strict inference time restriction. It keeps model size reasonable and requires us for some practical technique.\n\n## TL;DR\n- Ensemble 5 transformer models\n- Strong augmentation\n- Manual model conversion from pytroch to tensorflow\n- Code is available here -> https://github.com/bamps53/kaggle-asl-11th-place-solution\n\n## Overview\nI started from @hengck23 ‘s [great discussion](https://www.kaggle.com/competitions/asl-signs/discussion/391265) and [notebook](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution). Thanks for sharing a lot of useful tricks as always!\n\n\nThe changes I made are following;\n- Change model architecture to CLIP transformer in HuggingFace\n- Decrease parameter size to maximize latency within the range of same accuracy\n- Some strong augmentations\n    -  Horizontal flip(p=0.5)\n    -  Random 3d rotation(p=1, -45~45)\n    -  Random scale(p=1, 0.5~1.5)\n    -  Random shift(p=1, 0.7~1.3)\n    -  Random mask frames(p=1, mask_ratio=0.5)\n    -  Random resize (p=1, 0.5~1.5)\n- Add motion features\n    - current - prev\n    - next - current\n    - Velocity  \n- Longer epoch, 250 for 5 fold and 300 for all data\n\nFor the details, please refer to the code.(planning to upload)\n\n## Model conversion\nI’m too lazy to implement augmentations in tensorflow dataset, so I keep using pytorch training pipeline. But I’ve realized that inference time of the model converted by onnx_tf is way slower than bare tensorflow models. Then I’ve tried some model conversion framework like nobuco, but there were too many errors for some reason. Finally I’ve decided to write the model architecture both in pytorch and tensorflow, then manually port the weight. Thankfully HuggingFace has both pytorch and tensorflow CLIP implementation, this work is easier than I thought. This significantly speeds up inference time and I can put more models when ensembling.\n\n## Ensemble\nI’ve tried to diversify the models as much as possible within the same accuracy range.\n\n| seed | layers | dim | act  | max_len | features            |\n|------|--------|-----|------|---------|---------------------|\n|    0 |      2 | 384 | relu |      64 | lip/hand            |\n|    1 |      3 | 256 | relu |      48 | lip/hand/eye/motion |\n|    2 |      2 | 384 | geru |      64 | lip/hand            |\n|    3 |      2 | 384 | geru |      64 | lip/hand/eye/motion |\n|    4 |      3 | 256 | geru |      48 | lip/hand/eye/motion |\n\n\nThe scores are as follows;\n\n|                   | public | private |\n|-------------------|--------|---------|\n| Single best model |  0.786 |   0.865 |\n| Ensemble 5 models |  0.794 |    0.87 |\n\n\n## What didn’t work;\n- Normalization matters a lot, so I’ve tried a lot of variants but couldn’t get a better result than simple video mean/std normalization.\n- More landmarks(like arm, ear, nose or pose)\n- Conv1d\n- Knowledge distillation\n- Model soup\n- Bigger model\n- Pretrained model (ex. CLIP pretrained model or pretrain with NTU dataset)\n- And so on..\n\n## Code\nhttps://github.com/bamps53/kaggle-asl-11th-place-solution",
    "2244845": "Updated to attach a link to the code:)\nhttps://github.com/bamps53/kaggle-asl-11th-place-solution",
    "2246980": "11th place is a great result!!!\nUsing other people's notebooks really speeds up the workflow"
  },
  "source": "meta"
}