{
  "id": 406538,
  "title": "119th Place Solution: Transformer is much much better than GRU!",
  "url": "/competitions/asl-signs/writeups/nghi-huynh-119th-place-solution-transformer-is-muc",
  "author_name": "",
  "post_date": "2023-05-02T17:57:06.333Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First, I want to thank <strong>Kaggle</strong>, <strong>PopSign</strong>, <strong>Google</strong>, and <strong>Partners</strong> for organizing this interesting competition! It provided me a great opportunity to learn and compete with other participants. Second, I want to congratulate to all the winners and other Kagglers for participating in this competition.  </p>\n<p>To document my learning journey in this competition, I want to share my approach towards this competition. The full inference code is publicly shared as <a href=\"https://www.kaggle.com/code/nghihuynh/gislr-inference-transformer\" target=\"_blank\">GISLR: Inference Transformer</a>.</p>\n<hr>\n<p><strong>Data Preprocessing:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc5c9a1b33b3c547327ed99a16323a36f%2Ffeature_processing.jpg?generation=1683047881721695&amp;alt=media\" alt=\"\"><br>\nI applied a simple data preprocessing pipeline with a small modification (without data normalization) from my shared notebook: <a href=\"https://www.kaggle.com/code/nghihuynh/gislr-eda-feature-processing\" target=\"_blank\">GISLR: EDA + Feature Processing</a>. Then, I saved it as featured_preprocessed data for further training.</p>\n<p><strong>Model Architecture:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Ffdf7168215528bb5e1e1289537440306%2Fmodel_architecture.jpg?generation=1683047933576113&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th><strong>Model</strong></th>\n<th><strong>Architecture</strong></th>\n<th><strong>Original Size</strong></th>\n<th><strong>Post Quantization (TFLite FP16 )</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GRUNet</td>\n<td>Landmark embedding:128; Combined embedding:396; GRU units:368; Conv1D: filter: 256, kernel_size: 5</td>\n<td>8.78MB</td>\n<td>4.4MB</td>\n</tr>\n<tr>\n<td>BiGRUNet</td>\n<td>Landmark embedding: 128; GRU units: 768; BiGRU units: 396; Conv1D: filter:396, kernel_size:5</td>\n<td>30.84MB</td>\n<td>15.4MB</td>\n</tr>\n<tr>\n<td>Transformer</td>\n<td>Landmark embedding: 256; Combined embedding: 396; GRU units: 368; Transformer: MHA units: 368, n_head:8, num_block:1; Conv1D: filters:396, kernel_size:5</td>\n<td>13.87MB</td>\n<td>6.9MB</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Cross-Validation</strong>: Participant K-Fold=4<br>\n2 sets of validation from fold 3= [62590, 36257, 22343, 37779]</p>\n<ul>\n<li>GRU: val_set= [22343, 37779]</li>\n<li>BiGRU &amp; Transformer: val_set=[37779]</li>\n</ul>\n<p><strong>Training:</strong><br>\nReduce LR on plateau: factor=0.30, patient=3<br>\nAdam optimizer: lr=4.0e-4<br>\nSparse Categorical Loss Entropy <br>\nEpochs:100</p>\n<p><strong>Results:</strong></p>\n<table>\n<thead>\n<tr>\n<th><strong>Model (Ensemble)</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Private LB</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GRU-BiGRU-Transformer (selected as final submission)</td>\n<td>0.7452</td>\n<td>0.82339</td>\n</tr>\n<tr>\n<td>Transformer (not-selected 🙁)</td>\n<td>0.73918</td>\n<td>0.82484</td>\n</tr>\n<tr>\n<td>Transformer (x2)-GRU (not-selected 🙁)</td>\n<td>0.74217</td>\n<td>0.82457</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Conclusion:</strong></p>\n<p>Special thanks to the following public notebooks, which helped me a lot during the competition.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">GISLR TF Data Processing &amp; Transformer Training by Mark Wijkhuizen</a></li>\n<li><a href=\"https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-69\" target=\"_blank\">GISLR TF: On the Shoulders ENSAMBLE V2 0.69 by Andrij</a></li>\n<li><a href=\"https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language\" target=\"_blank\">LSTM Baseline for Starters - Sign Language by JvThunder</a></li>\n<li><a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution\" target=\"_blank\">[LB 0.67] one pytorch transformer solution by hengck23</a></li>\n</ul>\n<p>Things I tried but didn't work:</p>\n<ul>\n<li>Different preprocessing (increase number of frames, include more pose, and eyes data, normalize data)</li>\n<li>Positional embedding(overfitting)</li>\n<li>Attention mask in GRU, BiGRU(overfitting)</li>\n</ul>",
  "messages": [
    {
      "id": "2243158",
      "postDate": "05/02/2023 17:53:40",
      "content": "<p>First, I want to thank <strong>Kaggle</strong>, <strong>PopSign</strong>, <strong>Google</strong>, and <strong>Partners</strong> for organizing this interesting competition! It provided me a great opportunity to learn and compete with other participants. Second, I want to congratulate to all the winners and other Kagglers for participating in this competition.  </p>\n<p>To document my learning journey in this competition, I want to share my approach towards this competition. The full inference code is publicly shared as <a href=\"https://www.kaggle.com/code/nghihuynh/gislr-inference-transformer\" target=\"_blank\">GISLR: Inference Transformer</a>.</p>\n<hr>\n<p><strong>Data Preprocessing:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc5c9a1b33b3c547327ed99a16323a36f%2Ffeature_processing.jpg?generation=1683047881721695&amp;alt=media\" alt=\"\"><br>\nI applied a simple data preprocessing pipeline with a small modification (without data normalization) from my shared notebook: <a href=\"https://www.kaggle.com/code/nghihuynh/gislr-eda-feature-processing\" target=\"_blank\">GISLR: EDA + Feature Processing</a>. Then, I saved it as featured_preprocessed data for further training.</p>\n<p><strong>Model Architecture:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Ffdf7168215528bb5e1e1289537440306%2Fmodel_architecture.jpg?generation=1683047933576113&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th><strong>Model</strong></th>\n<th><strong>Architecture</strong></th>\n<th><strong>Original Size</strong></th>\n<th><strong>Post Quantization (TFLite FP16 )</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GRUNet</td>\n<td>Landmark embedding:128; Combined embedding:396; GRU units:368; Conv1D: filter: 256, kernel_size: 5</td>\n<td>8.78MB</td>\n<td>4.4MB</td>\n</tr>\n<tr>\n<td>BiGRUNet</td>\n<td>Landmark embedding: 128; GRU units: 768; BiGRU units: 396; Conv1D: filter:396, kernel_size:5</td>\n<td>30.84MB</td>\n<td>15.4MB</td>\n</tr>\n<tr>\n<td>Transformer</td>\n<td>Landmark embedding: 256; Combined embedding: 396; GRU units: 368; Transformer: MHA units: 368, n_head:8, num_block:1; Conv1D: filters:396, kernel_size:5</td>\n<td>13.87MB</td>\n<td>6.9MB</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Cross-Validation</strong>: Participant K-Fold=4<br>\n2 sets of validation from fold 3= [62590, 36257, 22343, 37779]</p>\n<ul>\n<li>GRU: val_set= [22343, 37779]</li>\n<li>BiGRU &amp; Transformer: val_set=[37779]</li>\n</ul>\n<p><strong>Training:</strong><br>\nReduce LR on plateau: factor=0.30, patient=3<br>\nAdam optimizer: lr=4.0e-4<br>\nSparse Categorical Loss Entropy <br>\nEpochs:100</p>\n<p><strong>Results:</strong></p>\n<table>\n<thead>\n<tr>\n<th><strong>Model (Ensemble)</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Private LB</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GRU-BiGRU-Transformer (selected as final submission)</td>\n<td>0.7452</td>\n<td>0.82339</td>\n</tr>\n<tr>\n<td>Transformer (not-selected 🙁)</td>\n<td>0.73918</td>\n<td>0.82484</td>\n</tr>\n<tr>\n<td>Transformer (x2)-GRU (not-selected 🙁)</td>\n<td>0.74217</td>\n<td>0.82457</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Conclusion:</strong></p>\n<p>Special thanks to the following public notebooks, which helped me a lot during the competition.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">GISLR TF Data Processing &amp; Transformer Training by Mark Wijkhuizen</a></li>\n<li><a href=\"https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-69\" target=\"_blank\">GISLR TF: On the Shoulders ENSAMBLE V2 0.69 by Andrij</a></li>\n<li><a href=\"https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language\" target=\"_blank\">LSTM Baseline for Starters - Sign Language by JvThunder</a></li>\n<li><a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution\" target=\"_blank\">[LB 0.67] one pytorch transformer solution by hengck23</a></li>\n</ul>\n<p>Things I tried but didn't work:</p>\n<ul>\n<li>Different preprocessing (increase number of frames, include more pose, and eyes data, normalize data)</li>\n<li>Positional embedding(overfitting)</li>\n<li>Attention mask in GRU, BiGRU(overfitting)</li>\n</ul>",
      "rawMarkdown": "First, I want to thank **Kaggle**, **PopSign**, **Google**, and **Partners** for organizing this interesting competition! It provided me a great opportunity to learn and compete with other participants. Second, I want to congratulate to all the winners and other Kagglers for participating in this competition.  \n\nTo document my learning journey in this competition, I want to share my approach towards this competition. The full inference code is publicly shared as [GISLR: Inference Transformer](https://www.kaggle.com/code/nghihuynh/gislr-inference-transformer).\n\n---\n\n**Data Preprocessing:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc5c9a1b33b3c547327ed99a16323a36f%2Ffeature_processing.jpg?generation=1683047881721695&alt=media)\nI applied a simple data preprocessing pipeline with a small modification (without data normalization) from my shared notebook: [GISLR: EDA + Feature Processing](https://www.kaggle.com/code/nghihuynh/gislr-eda-feature-processing). Then, I saved it as featured_preprocessed data for further training.\n\n**Model Architecture:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Ffdf7168215528bb5e1e1289537440306%2Fmodel_architecture.jpg?generation=1683047933576113&alt=media)\n\n| **Model**   | **Architecture**                                                                                                                                        | **Original Size** | **Post Quantization (TFLite FP16 )** |\n|-------------|---------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|------------------------------|\n| GRUNet      | Landmark embedding:128; Combined embedding:396; GRU units:368; Conv1D: filter: 256, kernel_size: 5                                                       | 8.78MB            | 4.4MB                        |\n| BiGRUNet    | Landmark embedding: 128; GRU units: 768; BiGRU units: 396; Conv1D: filter:396, kernel_size:5                                                             | 30.84MB           | 15.4MB                       |\n| Transformer | Landmark embedding: 256; Combined embedding: 396; GRU units: 368; Transformer: MHA units: 368, n_head:8, num_block:1; Conv1D: filters:396, kernel_size:5 | 13.87MB           | 6.9MB                        |\n\n\n**Cross-Validation**: Participant K-Fold=4\n2 sets of validation from fold 3= [62590, 36257, 22343, 37779]\n* GRU: val_set= [22343, 37779]\n* BiGRU & Transformer: val_set=[37779]\n\n**Training:**\nReduce LR on plateau: factor=0.30, patient=3\nAdam optimizer: lr=4.0e-4\nSparse Categorical Loss Entropy \nEpochs:100\n\n**Results:**\n\n| **Model (Ensemble)**                                | **Public LB** | **Private LB** |\n|------------------------------------------|---------------|----------------|\n| GRU-BiGRU-Transformer (selected as final submission) | 0.7452        | 0.82339        |\n| Transformer (not-selected 🙁)            | 0.73918       | 0.82484        |\n| Transformer (x2)-GRU (not-selected 🙁)   | 0.74217       | 0.82457        |\n\n--- \n\n**Conclusion:**\n\nSpecial thanks to the following public notebooks, which helped me a lot during the competition.\n* [GISLR TF Data Processing & Transformer Training by Mark Wijkhuizen](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training)\n* [GISLR TF: On the Shoulders ENSAMBLE V2 0.69 by Andrij](https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-69)\n* [LSTM Baseline for Starters - Sign Language by JvThunder](https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language)\n* [[LB 0.67] one pytorch transformer solution by hengck23](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution)\n\nThings I tried but didn't work:\n* Different preprocessing (increase number of frames, include more pose, and eyes data, normalize data)\n* Positional embedding(overfitting)\n* Attention mask in GRU, BiGRU(overfitting)",
      "votes": null
    },
    {
      "id": "2243473",
      "postDate": "05/03/2023 00:10:13",
      "content": "<p>A huge congratulations Nghi, <br>\nyou're an amazing talented professional. Thanks for sharing your worthy solution.</p>",
      "rawMarkdown": "A huge congratulations Nghi, \nyou're an amazing talented professional. Thanks for sharing your worthy solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2243473,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "05/03/2023 00:10:13",
      "content": "<p>A huge congratulations Nghi, <br>\nyou're an amazing talented professional. Thanks for sharing your worthy solution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2243158": "First, I want to thank **Kaggle**, **PopSign**, **Google**, and **Partners** for organizing this interesting competition! It provided me a great opportunity to learn and compete with other participants. Second, I want to congratulate to all the winners and other Kagglers for participating in this competition.  \n\nTo document my learning journey in this competition, I want to share my approach towards this competition. The full inference code is publicly shared as [GISLR: Inference Transformer](https://www.kaggle.com/code/nghihuynh/gislr-inference-transformer).\n\n---\n\n**Data Preprocessing:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Fc5c9a1b33b3c547327ed99a16323a36f%2Ffeature_processing.jpg?generation=1683047881721695&alt=media)\nI applied a simple data preprocessing pipeline with a small modification (without data normalization) from my shared notebook: [GISLR: EDA + Feature Processing](https://www.kaggle.com/code/nghihuynh/gislr-eda-feature-processing). Then, I saved it as featured_preprocessed data for further training.\n\n**Model Architecture:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6261540%2Ffdf7168215528bb5e1e1289537440306%2Fmodel_architecture.jpg?generation=1683047933576113&alt=media)\n\n| **Model**   | **Architecture**                                                                                                                                        | **Original Size** | **Post Quantization (TFLite FP16 )** |\n|-------------|---------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|------------------------------|\n| GRUNet      | Landmark embedding:128; Combined embedding:396; GRU units:368; Conv1D: filter: 256, kernel_size: 5                                                       | 8.78MB            | 4.4MB                        |\n| BiGRUNet    | Landmark embedding: 128; GRU units: 768; BiGRU units: 396; Conv1D: filter:396, kernel_size:5                                                             | 30.84MB           | 15.4MB                       |\n| Transformer | Landmark embedding: 256; Combined embedding: 396; GRU units: 368; Transformer: MHA units: 368, n_head:8, num_block:1; Conv1D: filters:396, kernel_size:5 | 13.87MB           | 6.9MB                        |\n\n\n**Cross-Validation**: Participant K-Fold=4\n2 sets of validation from fold 3= [62590, 36257, 22343, 37779]\n* GRU: val_set= [22343, 37779]\n* BiGRU & Transformer: val_set=[37779]\n\n**Training:**\nReduce LR on plateau: factor=0.30, patient=3\nAdam optimizer: lr=4.0e-4\nSparse Categorical Loss Entropy \nEpochs:100\n\n**Results:**\n\n| **Model (Ensemble)**                                | **Public LB** | **Private LB** |\n|------------------------------------------|---------------|----------------|\n| GRU-BiGRU-Transformer (selected as final submission) | 0.7452        | 0.82339        |\n| Transformer (not-selected 🙁)            | 0.73918       | 0.82484        |\n| Transformer (x2)-GRU (not-selected 🙁)   | 0.74217       | 0.82457        |\n\n--- \n\n**Conclusion:**\n\nSpecial thanks to the following public notebooks, which helped me a lot during the competition.\n* [GISLR TF Data Processing & Transformer Training by Mark Wijkhuizen](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training)\n* [GISLR TF: On the Shoulders ENSAMBLE V2 0.69 by Andrij](https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-69)\n* [LSTM Baseline for Starters - Sign Language by JvThunder](https://www.kaggle.com/code/jvthunder/lstm-baseline-for-starters-sign-language)\n* [[LB 0.67] one pytorch transformer solution by hengck23](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution)\n\nThings I tried but didn't work:\n* Different preprocessing (increase number of frames, include more pose, and eyes data, normalize data)\n* Positional embedding(overfitting)\n* Attention mask in GRU, BiGRU(overfitting)",
    "2243473": "A huge congratulations Nghi, \nyou're an amazing talented professional. Thanks for sharing your worthy solution."
  },
  "source": "meta"
}