{
  "id": 406301,
  "title": "14th place solution: publicly shared Transformer architecture was so strong!",
  "url": "/competitions/asl-signs/writeups/bilzard-14th-place-solution-publicly-shared-transf",
  "author_name": "",
  "post_date": "2024-04-22T11:36:01.220Z",
  "votes": 20,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Summary</h1>\n<p>I thank Kaggle administraror &amp; host for holding this competition. Although struggling with Tensorflow's unfriendly errors and lots of try-and-errors for failing TFLite conversion was really, really tough, these low-layer experience was valuable for me.<br>\nBelow is my solution writeup of this competition.</p>\n<h2>Pipeline</h2>\n<p>I tried over 377 different patterns of training models for this competition, however, the best architecture is only minor-changed one from <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">Mark Wijkhuizen's great public notebook</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F980bafe93f77c3d2418f5512ebfb2aa3%2Fpipeline.png?generation=1682985827786639&amp;alt=media\"></p>\n<h3>Mixed PostLN &amp; PreLN Architecture</h3>\n<p>I tested a) PostLN, b) PreLN, c) Mixed architectures, and found Mixed architecture provides the best result. This architecture was originally (perhaps unintentionally) implemented in the Mark Wijkhuizen's public notebook (in earlier version).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe045ca21bb34a220006c06d9306446aa%2Ftransformer_block.png?generation=1682985844553627&amp;alt=media\"></p>\n<h3>Other modifications</h3>\n<ul>\n<li>keep frames with no hands (instead of dropping) in pre-processing.</li>\n<li>increase number of layers in the keypoint encoder. With more layers, the accuracy gets better. The 4x layer seetting is the best tradeoff for accuracy and inference time.</li>\n<li>set number of hidden units in keypoint encoder independently for each parts (lips=192, left_hand=256, right_hand=256, pose=128). This reduces inference time without losing accuracy.</li>\n<li>attach ArcFace layer on training.</li>\n</ul>\n<h2>Training Setting</h2>\n<ul>\n<li>loss function: <code>0.5 * ArcFace + 0.5 * CrossEntropy</code> (This setting was shared by <a href=\"https://www.kaggle.com/code/medali1992/gislr-nn-arcface-baseline\" target=\"_blank\">Med Ali Bouchhioua</a>). Using ArcFace loss together with cross entropy loss converges faster, as well as the accuracy gets better.</li>\n<li>I tested 50, 80, 100, 120 epochs, but 100 epoch is the best on LB.</li>\n<li>Label smoothing (0.20-0.25) can avoid overfitting, but ArcFace is better. Using both label smoothing and ArcFace didn't increase CV/LB.</li>\n</ul>\n<h3>Data Augmentation</h3>\n<p>The augmentation strategy is almost the same as <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265\" target=\"_blank\">that of </a><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> was shared.</p>\n<ul>\n<li>handedness swapping (p=0.5)</li>\n<li>global 2D Affine transformation (p=0.5, shift=(-0.1, 0.1), rotation=(-30, 30), scale=(0.9, 1.1), shear=(-1.5, 1.5))</li>\n<li>random frame masking (p=0.5, mask_ratio=0.75).</li>\n</ul>\n<p>Frame masking simulates the mis-detection of keypoins. It also act as cutout augmentation in image tasks.</p>\n<h2>TFLite Conversion</h2>\n<p>When FP16 quantization is applied for the Mark's original implementation, the Transformer block outputs NaN value. So I rewrote transformer block based on <a href=\"https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture/notebook\" target=\"_blank\">@henck23's implementation</a>.<br>\nFP16 quantization made model size about 1/2, without losing model accuracy.</p>\n<h3>TTA</h3>\n<ul>\n<li>apply handedness swapping augmentation to 2/4 of ensemble-seed models</li>\n</ul>\n<h3>The output TFLite model</h3>\n<ul>\n<li>model size=26MB</li>\n<li>scoring time=56-59 min</li>\n</ul>\n<h2>Not-worked Experiments</h2>\n<ul>\n<li>distance/velocity features didn't contribute to increase accuracy.</li>\n<li>using synthesized hand pose dataset, I trained z-axis prediction model. I used this model's prediction result for a) 3D-affine transform augmentation and b) sub-task to predict z-axis from key-point encoder's output, but both trials didn't contribute to increase accuracy.</li>\n</ul>",
  "messages": [
    {
      "id": "2241888",
      "postDate": "05/02/2023 00:05:25",
      "content": "<h1>Summary</h1>\n<p>I thank Kaggle administraror &amp; host for holding this competition. Although struggling with Tensorflow's unfriendly errors and lots of try-and-errors for failing TFLite conversion was really, really tough, these low-layer experience was valuable for me.<br>\nBelow is my solution writeup of this competition.</p>\n<h2>Pipeline</h2>\n<p>I tried over 377 different patterns of training models for this competition, however, the best architecture is only minor-changed one from <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">Mark Wijkhuizen's great public notebook</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F980bafe93f77c3d2418f5512ebfb2aa3%2Fpipeline.png?generation=1682985827786639&amp;alt=media\"></p>\n<h3>Mixed PostLN &amp; PreLN Architecture</h3>\n<p>I tested a) PostLN, b) PreLN, c) Mixed architectures, and found Mixed architecture provides the best result. This architecture was originally (perhaps unintentionally) implemented in the Mark Wijkhuizen's public notebook (in earlier version).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe045ca21bb34a220006c06d9306446aa%2Ftransformer_block.png?generation=1682985844553627&amp;alt=media\"></p>\n<h3>Other modifications</h3>\n<ul>\n<li>keep frames with no hands (instead of dropping) in pre-processing.</li>\n<li>increase number of layers in the keypoint encoder. With more layers, the accuracy gets better. The 4x layer seetting is the best tradeoff for accuracy and inference time.</li>\n<li>set number of hidden units in keypoint encoder independently for each parts (lips=192, left_hand=256, right_hand=256, pose=128). This reduces inference time without losing accuracy.</li>\n<li>attach ArcFace layer on training.</li>\n</ul>\n<h2>Training Setting</h2>\n<ul>\n<li>loss function: <code>0.5 * ArcFace + 0.5 * CrossEntropy</code> (This setting was shared by <a href=\"https://www.kaggle.com/code/medali1992/gislr-nn-arcface-baseline\" target=\"_blank\">Med Ali Bouchhioua</a>). Using ArcFace loss together with cross entropy loss converges faster, as well as the accuracy gets better.</li>\n<li>I tested 50, 80, 100, 120 epochs, but 100 epoch is the best on LB.</li>\n<li>Label smoothing (0.20-0.25) can avoid overfitting, but ArcFace is better. Using both label smoothing and ArcFace didn't increase CV/LB.</li>\n</ul>\n<h3>Data Augmentation</h3>\n<p>The augmentation strategy is almost the same as <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265\" target=\"_blank\">that of </a><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> was shared.</p>\n<ul>\n<li>handedness swapping (p=0.5)</li>\n<li>global 2D Affine transformation (p=0.5, shift=(-0.1, 0.1), rotation=(-30, 30), scale=(0.9, 1.1), shear=(-1.5, 1.5))</li>\n<li>random frame masking (p=0.5, mask_ratio=0.75).</li>\n</ul>\n<p>Frame masking simulates the mis-detection of keypoins. It also act as cutout augmentation in image tasks.</p>\n<h2>TFLite Conversion</h2>\n<p>When FP16 quantization is applied for the Mark's original implementation, the Transformer block outputs NaN value. So I rewrote transformer block based on <a href=\"https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture/notebook\" target=\"_blank\">@henck23's implementation</a>.<br>\nFP16 quantization made model size about 1/2, without losing model accuracy.</p>\n<h3>TTA</h3>\n<ul>\n<li>apply handedness swapping augmentation to 2/4 of ensemble-seed models</li>\n</ul>\n<h3>The output TFLite model</h3>\n<ul>\n<li>model size=26MB</li>\n<li>scoring time=56-59 min</li>\n</ul>\n<h2>Not-worked Experiments</h2>\n<ul>\n<li>distance/velocity features didn't contribute to increase accuracy.</li>\n<li>using synthesized hand pose dataset, I trained z-axis prediction model. I used this model's prediction result for a) 3D-affine transform augmentation and b) sub-task to predict z-axis from key-point encoder's output, but both trials didn't contribute to increase accuracy.</li>\n</ul>",
      "rawMarkdown": "# Summary\n\nI thank Kaggle administraror & host for holding this competition. Although struggling with Tensorflow's unfriendly errors and lots of try-and-errors for failing TFLite conversion was really, really tough, these low-layer experience was valuable for me.\nBelow is my solution writeup of this competition.\n\n## Pipeline\n\nI tried over 377 different patterns of training models for this competition, however, the best architecture is only minor-changed one from [Mark Wijkhuizen's great public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F980bafe93f77c3d2418f5512ebfb2aa3%2Fpipeline.png?generation=1682985827786639&alt=media)\n\n### Mixed PostLN & PreLN Architecture\n\nI tested a) PostLN, b) PreLN, c) Mixed architectures, and found Mixed architecture provides the best result. This architecture was originally (perhaps unintentionally) implemented in the Mark Wijkhuizen's public notebook (in earlier version).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe045ca21bb34a220006c06d9306446aa%2Ftransformer_block.png?generation=1682985844553627&alt=media)\n\n### Other modifications\n\n* keep frames with no hands (instead of dropping) in pre-processing.\n* increase number of layers in the keypoint encoder. With more layers, the accuracy gets better. The 4x layer seetting is the best tradeoff for accuracy and inference time.\n* set number of hidden units in keypoint encoder independently for each parts (lips=192, left_hand=256, right_hand=256, pose=128). This reduces inference time without losing accuracy.\n* attach ArcFace layer on training.\n\n## Training Setting\n\n* loss function: `0.5 * ArcFace + 0.5 * CrossEntropy` (This setting was shared by [Med Ali Bouchhioua](https://www.kaggle.com/code/medali1992/gislr-nn-arcface-baseline)). Using ArcFace loss together with cross entropy loss converges faster, as well as the accuracy gets better.\n* I tested 50, 80, 100, 120 epochs, but 100 epoch is the best on LB.\n* Label smoothing (0.20-0.25) can avoid overfitting, but ArcFace is better. Using both label smoothing and ArcFace didn't increase CV/LB.\n\n### Data Augmentation\n\nThe augmentation strategy is almost the same as [that of @hengck23 was shared](https://www.kaggle.com/competitions/asl-signs/discussion/391265).\n\n* handedness swapping (p=0.5)\n* global 2D Affine transformation (p=0.5, shift=(-0.1, 0.1), rotation=(-30, 30), scale=(0.9, 1.1), shear=(-1.5, 1.5))\n* random frame masking (p=0.5, mask_ratio=0.75).\n\nFrame masking simulates the mis-detection of keypoins. It also act as cutout augmentation in image tasks.\n\n## TFLite Conversion\n\nWhen FP16 quantization is applied for the Mark's original implementation, the Transformer block outputs NaN value. So I rewrote transformer block based on [@henck23's implementation](https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture/notebook).\nFP16 quantization made model size about 1/2, without losing model accuracy.\n\n### TTA\n\n* apply handedness swapping augmentation to 2/4 of ensemble-seed models\n\n### The output TFLite model\n\n* model size=26MB\n* scoring time=56-59 min\n\n## Not-worked Experiments\n\n* distance/velocity features didn't contribute to increase accuracy.\n* using synthesized hand pose dataset, I trained z-axis prediction model. I used this model's prediction result for a) 3D-affine transform augmentation and b) sub-task to predict z-axis from key-point encoder's output, but both trials didn't contribute to increase accuracy.",
      "votes": null
    },
    {
      "id": "2241961",
      "postDate": "05/02/2023 01:18:37",
      "content": "<p>Fantastic approach note and excellent result <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, hearty congratulations for the result! All the best and regards!</p>",
      "rawMarkdown": "Fantastic approach note and excellent result @tatamikenn, hearty congratulations for the result! All the best and regards!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2241961,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/02/2023 01:18:37",
      "content": "<p>Fantastic approach note and excellent result <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, hearty congratulations for the result! All the best and regards!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2241888": "# Summary\n\nI thank Kaggle administraror & host for holding this competition. Although struggling with Tensorflow's unfriendly errors and lots of try-and-errors for failing TFLite conversion was really, really tough, these low-layer experience was valuable for me.\nBelow is my solution writeup of this competition.\n\n## Pipeline\n\nI tried over 377 different patterns of training models for this competition, however, the best architecture is only minor-changed one from [Mark Wijkhuizen's great public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F980bafe93f77c3d2418f5512ebfb2aa3%2Fpipeline.png?generation=1682985827786639&alt=media)\n\n### Mixed PostLN & PreLN Architecture\n\nI tested a) PostLN, b) PreLN, c) Mixed architectures, and found Mixed architecture provides the best result. This architecture was originally (perhaps unintentionally) implemented in the Mark Wijkhuizen's public notebook (in earlier version).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe045ca21bb34a220006c06d9306446aa%2Ftransformer_block.png?generation=1682985844553627&alt=media)\n\n### Other modifications\n\n* keep frames with no hands (instead of dropping) in pre-processing.\n* increase number of layers in the keypoint encoder. With more layers, the accuracy gets better. The 4x layer seetting is the best tradeoff for accuracy and inference time.\n* set number of hidden units in keypoint encoder independently for each parts (lips=192, left_hand=256, right_hand=256, pose=128). This reduces inference time without losing accuracy.\n* attach ArcFace layer on training.\n\n## Training Setting\n\n* loss function: `0.5 * ArcFace + 0.5 * CrossEntropy` (This setting was shared by [Med Ali Bouchhioua](https://www.kaggle.com/code/medali1992/gislr-nn-arcface-baseline)). Using ArcFace loss together with cross entropy loss converges faster, as well as the accuracy gets better.\n* I tested 50, 80, 100, 120 epochs, but 100 epoch is the best on LB.\n* Label smoothing (0.20-0.25) can avoid overfitting, but ArcFace is better. Using both label smoothing and ArcFace didn't increase CV/LB.\n\n### Data Augmentation\n\nThe augmentation strategy is almost the same as [that of @hengck23 was shared](https://www.kaggle.com/competitions/asl-signs/discussion/391265).\n\n* handedness swapping (p=0.5)\n* global 2D Affine transformation (p=0.5, shift=(-0.1, 0.1), rotation=(-30, 30), scale=(0.9, 1.1), shear=(-1.5, 1.5))\n* random frame masking (p=0.5, mask_ratio=0.75).\n\nFrame masking simulates the mis-detection of keypoins. It also act as cutout augmentation in image tasks.\n\n## TFLite Conversion\n\nWhen FP16 quantization is applied for the Mark's original implementation, the Transformer block outputs NaN value. So I rewrote transformer block based on [@henck23's implementation](https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture/notebook).\nFP16 quantization made model size about 1/2, without losing model accuracy.\n\n### TTA\n\n* apply handedness swapping augmentation to 2/4 of ensemble-seed models\n\n### The output TFLite model\n\n* model size=26MB\n* scoring time=56-59 min\n\n## Not-worked Experiments\n\n* distance/velocity features didn't contribute to increase accuracy.\n* using synthesized hand pose dataset, I trained z-axis prediction model. I used this model's prediction result for a) 3D-affine transform augmentation and b) sub-task to predict z-axis from key-point encoder's output, but both trials didn't contribute to increase accuracy.",
    "2241961": "Fantastic approach note and excellent result @tatamikenn, hearty congratulations for the result! All the best and regards!"
  },
  "source": "meta"
}