{
  "id": 406537,
  "title": "6th Place Solution - Transformer Tweaking Madness",
  "url": "/competitions/asl-signs/writeups/6th-place-solution-transformer-tweaking-madness",
  "author_name": "",
  "post_date": "2023-05-03T14:20:15.910Z",
  "votes": 49,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thanks to both, the organizers of this competition who offered a fun yet challenging problem as well as all of the other competitors - well done to everyone who worked hard for small incremental increases. </p>\n<p>Although I am the one posting the topic, this is the result of a great team effort, so big shoutout to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>\n<h1>Brief Summary</h1>\n<p>Our solution is a 2 model ensemble of a MLP-encoder-frame-transformer model. We pushed our transformer models close to the limit and implemented a lot of tricks to climb up to 6th place. <br>\nI have 1403 hours of experiment monitoring time in April (that’s 48h per day :)).</p>\n<p><strong>Update :</strong> Code is available here : <a href=\"https://github.com/TheoViel/kaggle_islr\" target=\"_blank\">https://github.com/TheoViel/kaggle_islr</a></p>\n<h1>Detailed Summary</h1>\n<h2>Preprocessing &amp; Model</h2>\n<p><a href=\"https://ibb.co/g3np4jQ\"><img src=\"https://i.ibb.co/80hpYKG/ISLR-drawio-1.png\" alt=\"ISLR-drawio-1\"></a></p>\n<h3>Preprocessing</h3>\n<ul>\n<li>Remove frames without fingers</li>\n<li>Stride the sequence (use 1 every n frames) such that the sequence size is <code>&lt;= max_len</code>. We used <code>max_len=25</code> and <code>80</code> in the final ensemble</li>\n<li>Normalization is done for the whole sequence to have 0 mean and 1 std. We do an extra centering before the specific MLP. Nan values are set to 0</li>\n</ul>\n<h3>Embedding</h3>\n<ul>\n<li>2x 1D convolutions (<code>k=5</code>) to smooth the positions</li>\n<li>Embed the landmark id and type (e.g. lips, right hand, …), <code>embed_dim=dense_dim=16</code></li>\n</ul>\n<h3>Feature extractor</h3>\n<ul>\n<li>One MLP combining all the features, and 4 for specific landmark types (2x hands, face, lips)</li>\n<li>Max aggregation for hands is to take into account that signers use one hand</li>\n<li><code>dim=192</code>, <code>dropout=0.25</code></li>\n</ul>\n<h3>Transformer</h3>\n<ul>\n<li>Deberta was better than Bert, but we had to rewrite the attention layer for it to be efficient</li>\n<li>To reduce the number of parameters, we use a smaller first transformer, and modify the output layer to upscale/downscale the features. This was key to enable blending 2 models</li>\n<li><code>d_in=512</code>, <code>δ=256</code> for <code>max_len=25</code>, <code>δ=64</code> for <code>max_len=80</code>, <code>num_heads=16</code>, <code>dropout=0.05</code> for the first layer, <code>0.1</code> for the other 2</li>\n<li>Unfortunately, we did not have any luck with using a pre-trained version of Deberta, for example by importing some of the pretrained weights</li>\n</ul>\n<h2>Training strategy</h2>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/VvT0pVx/ISLR-Train-drawio-1.png\" alt=\"ISLR-Train-drawio-1\"></a></p>\n<h3>Augmentations:</h3>\n<ul>\n<li>Horizontal Flip (<code>p=0.5</code>)</li>\n<li>Rotate around <code>(0, 0, 0)</code> by an angle between -60 and 60°  (<code>p=0.5</code>)</li>\n<li>Resizing by a factor in <code>(0.7, 1.3)</code> (<code>p=0.5</code>). We also allow for distortion (<code>p=0.5</code>)</li>\n<li>Crop 20% of the start or end (<code>p=0.5</code>)</li>\n<li>Interpolate to fill missing values (<code>p=0.5</code>)</li>\n<li>Manifold Mixup <a href=\"https://arxiv.org/abs/1806.05236\" target=\"_blank\">[1]</a> (scheduled, <code>p=0.5 * epoch / (0.9 * n_epochs)</code>) : random apply mixup to the features before one of the transformer layer</li>\n<li>Only during the first half of the training, since it improved convergence<ul>\n<li>Fill the value of the missing hand with those of the existing one (<code>p=0.25</code>)</li>\n<li>Face CutMix : replace the face landmarks with those of another signer doing the same sign (<code>p=0.25</code>)</li></ul></li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>100 epochs, <code>lr=3e-4</code>, 25% warmup, linear schedule</li>\n<li>Cross entropy with smoothing (<code>eps=0.3</code>)</li>\n<li><code>weight_decay=0.4</code>, <code>batch_size=32</code> (x8 GPUs)</li>\n<li>OUSM <a href=\"https://arxiv.org/pdf/1901.07759.pdf\" target=\"_blank\">[2]</a>, i.e. exclude the top k (<code>k=3</code>) samples with the highest loss from the computation</li>\n<li>Mean teacher <a href=\"https://arxiv.org/abs/1703.01780\" target=\"_blank\">[3]</a> &amp; Knowledge Distillation (see image above). We train 3 models at the same time, and use the distilled one for inference</li>\n<li>Model soup <a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">[4]</a> of the last 10 epochs checkpoints</li>\n</ul>\n<h2>Experiments</h2>\n<p><a href=\"https://ibb.co/MV343Xx\"><img src=\"https://i.ibb.co/jM121pP/cvlb.png\" alt=\"cvlb\"></a></p>\n<h3>Validation strategy</h3>\n<p>We use a Stratified 4-fold, grouped by patient. We had a great CV - LB correlation (see figure), and our best model achieved CV 0.749 - Public 0.795 - Private 0.877. We submitted single models trained on the full dataset, and 2 models for ensembles.</p>\n<h3>What worked &amp; reported improvements</h3>\n<p>Starting from early April :</p>\n<ul>\n<li>Baseline : MLP + Bert : LB 0.72</li>\n<li>Deberta instead of Bert : CV +0.015  <strong>-&gt; LB 0.73</strong></li>\n<li>Improve Feature extractor : CV +0.01  <strong>-&gt; LB 0.74</strong></li>\n<li>Flip augmentation : CV +0.015  <strong>-&gt; LB 0.75</strong></li>\n<li>Increase model size : CV +0.006 </li>\n<li>Two stage training + interp aug : CV + 0.002 ?</li>\n<li>Increase model size  : CV +0.005  <strong>-&gt; LB 0.76</strong></li>\n<li>Deberta Rework : CV +0.005   <strong>-&gt; LB 0.77</strong></li>\n<li>OUSM : CV +0.003</li>\n<li>Mean Teacher : CV +0.005</li>\n<li>Ensemble two distilled models : +0.005 <strong>-&gt; LB 0.78</strong></li>\n<li>Conv layers : CV +0.001</li>\n<li>Mish <a href=\"https://arxiv.org/abs/1908.08681\" target=\"_blank\">[5]</a> activation instead of ReLU/LeakyReLU : CV +0.001</li>\n<li>Warmup 0.1 -&gt; 0.25 : CV +0.001</li>\n<li>Model soup : CV +0.001</li>\n<li><a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">Nobuco</a> instead of onnx : 45ms/it -&gt; 38ms/it, this enabled to improve the ensemble </li>\n<li>Weight decay 0 -&gt; 0.4 : CV +0.004</li>\n<li>Truncation aug : CV +0.001</li>\n<li>Manifold mixup : CV +0.002</li>\n<li>Different input, model and teacher size for ensemble : CV +0.005 <strong>-&gt; LB 0.79</strong></li>\n<li>Centering before the MLP layers : CV +0.004  <strong>-&gt; LB 0.795  (final)</strong></li>\n</ul>\n<h3>What did not work (for us)</h3>\n<ul>\n<li>Unfortunately we could not make CNNs work</li>\n<li>GCNs, or other architectures that are “sota” for ASL</li>\n<li>Heng’s architecture, although others had great results with it</li>\n<li>Getting a successful 3-model sub</li>\n<li>Relabel the data to reduce label noise</li>\n<li>Other augmentations such as noise, dropout, drop frame, shift</li>\n<li>Dropout scheduling, custom lr per layer</li>\n<li>Pretraining part of the model</li>\n<li>Handcrafted graph features (adjacency matrices, edge features)</li>\n<li>Adding more landmarks (eyes, eyebrows)</li>\n</ul>\n<p><em>Thanks for reading !</em></p>",
  "messages": [
    {
      "id": "2243152",
      "postDate": "05/02/2023 17:51:55",
      "content": "<p>Thanks to both, the organizers of this competition who offered a fun yet challenging problem as well as all of the other competitors - well done to everyone who worked hard for small incremental increases. </p>\n<p>Although I am the one posting the topic, this is the result of a great team effort, so big shoutout to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>\n<h1>Brief Summary</h1>\n<p>Our solution is a 2 model ensemble of a MLP-encoder-frame-transformer model. We pushed our transformer models close to the limit and implemented a lot of tricks to climb up to 6th place. <br>\nI have 1403 hours of experiment monitoring time in April (that’s 48h per day :)).</p>\n<p><strong>Update :</strong> Code is available here : <a href=\"https://github.com/TheoViel/kaggle_islr\" target=\"_blank\">https://github.com/TheoViel/kaggle_islr</a></p>\n<h1>Detailed Summary</h1>\n<h2>Preprocessing &amp; Model</h2>\n<p><a href=\"https://ibb.co/g3np4jQ\"><img src=\"https://i.ibb.co/80hpYKG/ISLR-drawio-1.png\" alt=\"ISLR-drawio-1\"></a></p>\n<h3>Preprocessing</h3>\n<ul>\n<li>Remove frames without fingers</li>\n<li>Stride the sequence (use 1 every n frames) such that the sequence size is <code>&lt;= max_len</code>. We used <code>max_len=25</code> and <code>80</code> in the final ensemble</li>\n<li>Normalization is done for the whole sequence to have 0 mean and 1 std. We do an extra centering before the specific MLP. Nan values are set to 0</li>\n</ul>\n<h3>Embedding</h3>\n<ul>\n<li>2x 1D convolutions (<code>k=5</code>) to smooth the positions</li>\n<li>Embed the landmark id and type (e.g. lips, right hand, …), <code>embed_dim=dense_dim=16</code></li>\n</ul>\n<h3>Feature extractor</h3>\n<ul>\n<li>One MLP combining all the features, and 4 for specific landmark types (2x hands, face, lips)</li>\n<li>Max aggregation for hands is to take into account that signers use one hand</li>\n<li><code>dim=192</code>, <code>dropout=0.25</code></li>\n</ul>\n<h3>Transformer</h3>\n<ul>\n<li>Deberta was better than Bert, but we had to rewrite the attention layer for it to be efficient</li>\n<li>To reduce the number of parameters, we use a smaller first transformer, and modify the output layer to upscale/downscale the features. This was key to enable blending 2 models</li>\n<li><code>d_in=512</code>, <code>δ=256</code> for <code>max_len=25</code>, <code>δ=64</code> for <code>max_len=80</code>, <code>num_heads=16</code>, <code>dropout=0.05</code> for the first layer, <code>0.1</code> for the other 2</li>\n<li>Unfortunately, we did not have any luck with using a pre-trained version of Deberta, for example by importing some of the pretrained weights</li>\n</ul>\n<h2>Training strategy</h2>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/VvT0pVx/ISLR-Train-drawio-1.png\" alt=\"ISLR-Train-drawio-1\"></a></p>\n<h3>Augmentations:</h3>\n<ul>\n<li>Horizontal Flip (<code>p=0.5</code>)</li>\n<li>Rotate around <code>(0, 0, 0)</code> by an angle between -60 and 60°  (<code>p=0.5</code>)</li>\n<li>Resizing by a factor in <code>(0.7, 1.3)</code> (<code>p=0.5</code>). We also allow for distortion (<code>p=0.5</code>)</li>\n<li>Crop 20% of the start or end (<code>p=0.5</code>)</li>\n<li>Interpolate to fill missing values (<code>p=0.5</code>)</li>\n<li>Manifold Mixup <a href=\"https://arxiv.org/abs/1806.05236\" target=\"_blank\">[1]</a> (scheduled, <code>p=0.5 * epoch / (0.9 * n_epochs)</code>) : random apply mixup to the features before one of the transformer layer</li>\n<li>Only during the first half of the training, since it improved convergence<ul>\n<li>Fill the value of the missing hand with those of the existing one (<code>p=0.25</code>)</li>\n<li>Face CutMix : replace the face landmarks with those of another signer doing the same sign (<code>p=0.25</code>)</li></ul></li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>100 epochs, <code>lr=3e-4</code>, 25% warmup, linear schedule</li>\n<li>Cross entropy with smoothing (<code>eps=0.3</code>)</li>\n<li><code>weight_decay=0.4</code>, <code>batch_size=32</code> (x8 GPUs)</li>\n<li>OUSM <a href=\"https://arxiv.org/pdf/1901.07759.pdf\" target=\"_blank\">[2]</a>, i.e. exclude the top k (<code>k=3</code>) samples with the highest loss from the computation</li>\n<li>Mean teacher <a href=\"https://arxiv.org/abs/1703.01780\" target=\"_blank\">[3]</a> &amp; Knowledge Distillation (see image above). We train 3 models at the same time, and use the distilled one for inference</li>\n<li>Model soup <a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">[4]</a> of the last 10 epochs checkpoints</li>\n</ul>\n<h2>Experiments</h2>\n<p><a href=\"https://ibb.co/MV343Xx\"><img src=\"https://i.ibb.co/jM121pP/cvlb.png\" alt=\"cvlb\"></a></p>\n<h3>Validation strategy</h3>\n<p>We use a Stratified 4-fold, grouped by patient. We had a great CV - LB correlation (see figure), and our best model achieved CV 0.749 - Public 0.795 - Private 0.877. We submitted single models trained on the full dataset, and 2 models for ensembles.</p>\n<h3>What worked &amp; reported improvements</h3>\n<p>Starting from early April :</p>\n<ul>\n<li>Baseline : MLP + Bert : LB 0.72</li>\n<li>Deberta instead of Bert : CV +0.015  <strong>-&gt; LB 0.73</strong></li>\n<li>Improve Feature extractor : CV +0.01  <strong>-&gt; LB 0.74</strong></li>\n<li>Flip augmentation : CV +0.015  <strong>-&gt; LB 0.75</strong></li>\n<li>Increase model size : CV +0.006 </li>\n<li>Two stage training + interp aug : CV + 0.002 ?</li>\n<li>Increase model size  : CV +0.005  <strong>-&gt; LB 0.76</strong></li>\n<li>Deberta Rework : CV +0.005   <strong>-&gt; LB 0.77</strong></li>\n<li>OUSM : CV +0.003</li>\n<li>Mean Teacher : CV +0.005</li>\n<li>Ensemble two distilled models : +0.005 <strong>-&gt; LB 0.78</strong></li>\n<li>Conv layers : CV +0.001</li>\n<li>Mish <a href=\"https://arxiv.org/abs/1908.08681\" target=\"_blank\">[5]</a> activation instead of ReLU/LeakyReLU : CV +0.001</li>\n<li>Warmup 0.1 -&gt; 0.25 : CV +0.001</li>\n<li>Model soup : CV +0.001</li>\n<li><a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">Nobuco</a> instead of onnx : 45ms/it -&gt; 38ms/it, this enabled to improve the ensemble </li>\n<li>Weight decay 0 -&gt; 0.4 : CV +0.004</li>\n<li>Truncation aug : CV +0.001</li>\n<li>Manifold mixup : CV +0.002</li>\n<li>Different input, model and teacher size for ensemble : CV +0.005 <strong>-&gt; LB 0.79</strong></li>\n<li>Centering before the MLP layers : CV +0.004  <strong>-&gt; LB 0.795  (final)</strong></li>\n</ul>\n<h3>What did not work (for us)</h3>\n<ul>\n<li>Unfortunately we could not make CNNs work</li>\n<li>GCNs, or other architectures that are “sota” for ASL</li>\n<li>Heng’s architecture, although others had great results with it</li>\n<li>Getting a successful 3-model sub</li>\n<li>Relabel the data to reduce label noise</li>\n<li>Other augmentations such as noise, dropout, drop frame, shift</li>\n<li>Dropout scheduling, custom lr per layer</li>\n<li>Pretraining part of the model</li>\n<li>Handcrafted graph features (adjacency matrices, edge features)</li>\n<li>Adding more landmarks (eyes, eyebrows)</li>\n</ul>\n<p><em>Thanks for reading !</em></p>",
      "rawMarkdown": "Thanks to both, the organizers of this competition who offered a fun yet challenging problem as well as all of the other competitors - well done to everyone who worked hard for small incremental increases. \n\nAlthough I am the one posting the topic, this is the result of a great team effort, so big shoutout to @christofhenkel.\n\n# Brief Summary\n\nOur solution is a 2 model ensemble of a MLP-encoder-frame-transformer model. We pushed our transformer models close to the limit and implemented a lot of tricks to climb up to 6th place. \nI have 1403 hours of experiment monitoring time in April (that’s 48h per day :)).\n\n**Update :** Code is available here : https://github.com/TheoViel/kaggle_islr\n\n# Detailed Summary\n\n## Preprocessing & Model\n\n<a href=\"https://ibb.co/g3np4jQ\"><img src=\"https://i.ibb.co/80hpYKG/ISLR-drawio-1.png\" alt=\"ISLR-drawio-1\" border=\"0\"></a>\n\n### Preprocessing\n- Remove frames without fingers\n- Stride the sequence (use 1 every n frames) such that the sequence size is `<= max_len`. We used `max_len=25` and `80` in the final ensemble\n- Normalization is done for the whole sequence to have 0 mean and 1 std. We do an extra centering before the specific MLP. Nan values are set to 0\n\n### Embedding\n- 2x 1D convolutions (`k=5`) to smooth the positions\n- Embed the landmark id and type (e.g. lips, right hand, ...), `embed_dim=dense_dim=16`\n\n### Feature extractor\n- One MLP combining all the features, and 4 for specific landmark types (2x hands, face, lips)\n- Max aggregation for hands is to take into account that signers use one hand\n- `dim=192`, `dropout=0.25`\n\n### Transformer\n- Deberta was better than Bert, but we had to rewrite the attention layer for it to be efficient\n- To reduce the number of parameters, we use a smaller first transformer, and modify the output layer to upscale/downscale the features. This was key to enable blending 2 models\n- `d_in=512`, `δ=256` for `max_len=25`, `δ=64` for `max_len=80`, `num_heads=16`, `dropout=0.05` for the first layer, `0.1` for the other 2\n- Unfortunately, we did not have any luck with using a pre-trained version of Deberta, for example by importing some of the pretrained weights\n\n## Training strategy\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/VvT0pVx/ISLR-Train-drawio-1.png\" alt=\"ISLR-Train-drawio-1\" border=\"0\"></a>\n\n### Augmentations:\n- Horizontal Flip (`p=0.5`)\n- Rotate around `(0, 0, 0)` by an angle between -60 and 60°  (`p=0.5`)\n- Resizing by a factor in `(0.7, 1.3)` (`p=0.5`). We also allow for distortion (`p=0.5`)\n- Crop 20% of the start or end (`p=0.5`)\n- Interpolate to fill missing values (`p=0.5`)\n- Manifold Mixup [[1]](https://arxiv.org/abs/1806.05236) (scheduled, `p=0.5 * epoch / (0.9 * n_epochs)`) : random apply mixup to the features before one of the transformer layer\n- Only during the first half of the training, since it improved convergence\n  - Fill the value of the missing hand with those of the existing one (`p=0.25`)\n  - Face CutMix : replace the face landmarks with those of another signer doing the same sign (`p=0.25`)\n\n### Training\n- 100 epochs, `lr=3e-4`, 25% warmup, linear schedule\n- Cross entropy with smoothing (`eps=0.3`)\n- `weight_decay=0.4`, `batch_size=32` (x8 GPUs)\n- OUSM [[2]](https://arxiv.org/pdf/1901.07759.pdf), i.e. exclude the top k (`k=3`) samples with the highest loss from the computation\n- Mean teacher [[3]](https://arxiv.org/abs/1703.01780) & Knowledge Distillation (see image above). We train 3 models at the same time, and use the distilled one for inference\n- Model soup [[4]](https://arxiv.org/abs/2203.05482) of the last 10 epochs checkpoints\n\n## Experiments \n\n<a href=\"https://ibb.co/MV343Xx\"><img src=\"https://i.ibb.co/jM121pP/cvlb.png\" alt=\"cvlb\" border=\"0\"></a>\n\n###  Validation strategy\nWe use a Stratified 4-fold, grouped by patient. We had a great CV - LB correlation (see figure), and our best model achieved CV 0.749 - Public 0.795 - Private 0.877. We submitted single models trained on the full dataset, and 2 models for ensembles.\n\n###  What worked & reported improvements\nStarting from early April :\n- Baseline : MLP + Bert : LB 0.72\n- Deberta instead of Bert : CV +0.015  **-> LB 0.73**\n- Improve Feature extractor : CV +0.01  **-> LB 0.74**\n- Flip augmentation : CV +0.015  **-> LB 0.75**\n- Increase model size : CV +0.006 \n- Two stage training + interp aug : CV + 0.002 ?\n- Increase model size  : CV +0.005  **-> LB 0.76**\n- Deberta Rework : CV +0.005   **-> LB 0.77**\n- OUSM : CV +0.003\n- Mean Teacher : CV +0.005\n- Ensemble two distilled models : +0.005 **-> LB 0.78**\n- Conv layers : CV +0.001\n- Mish [[5]](https://arxiv.org/abs/1908.08681) activation instead of ReLU/LeakyReLU : CV +0.001\n- Warmup 0.1 -> 0.25 : CV +0.001\n- Model soup : CV +0.001\n- [Nobuco](https://github.com/AlexanderLutsenko/nobuco) instead of onnx : 45ms/it -> 38ms/it, this enabled to improve the ensemble \n- Weight decay 0 -> 0.4 : CV +0.004\n- Truncation aug : CV +0.001\n- Manifold mixup : CV +0.002\n- Different input, model and teacher size for ensemble : CV +0.005 **-> LB 0.79**\n- Centering before the MLP layers : CV +0.004  **-> LB 0.795  (final)**\n\n###  What did not work (for us)\n- Unfortunately we could not make CNNs work\n- GCNs, or other architectures that are “sota” for ASL\n- Heng’s architecture, although others had great results with it\n- Getting a successful 3-model sub\n- Relabel the data to reduce label noise\n- Other augmentations such as noise, dropout, drop frame, shift\n- Dropout scheduling, custom lr per layer\n- Pretraining part of the model\n- Handcrafted graph features (adjacency matrices, edge features)\n- Adding more landmarks (eyes, eyebrows)\n\n\n*Thanks for reading !*",
      "votes": null
    },
    {
      "id": "2243199",
      "postDate": "05/02/2023 18:24:50",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    },
    {
      "id": "2243438",
      "postDate": "05/02/2023 22:44:00",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks for sharing a very detailed solution!<br>\nI have 2 questions:</p>\n<blockquote>\n  <p>Manifold Mixup [1] (scheduled, p=0.5 * epoch / (0.9 * n_epochs)) : random apply mixup to the features before one of the transformer layer  </p>\n</blockquote>\n<p>First one is about mixup. If you apply mixup before the transformer layer, there should be a temporal dim and attention mask. How did you handle the mask?</p>\n<p>Second question is about model soup. It always brings worse results than a single model every time I try. Is there any secret to making it work? </p>",
      "rawMarkdown": "theoviel Thanks for sharing a very detailed solution!\nI have 2 questions:\n\n> Manifold Mixup [1] (scheduled, p=0.5 * epoch / (0.9 * n_epochs)) : random apply mixup to the features before one of the transformer layer  \n\nFirst one is about mixup. If you apply mixup before the transformer layer, there should be a temporal dim and attention mask. How did you handle the mask?\n\nSecond question is about model soup. It always brings worse results than a single model every time I try. Is there any secret to making it work?",
      "votes": null
    },
    {
      "id": "2243818",
      "postDate": "05/03/2023 07:28:17",
      "content": "<p>You're welcome !</p>\n<ul>\n<li>The mixed mask is the maximum of the two initial masks, in order to allow the transformer to attend everywhere.</li>\n<li>You need to have cps close to the same local minimum. In our case lr is quite low during the last 10 epochs (linearly decreased from 4e-5 to 0), and performance does not change a lot. Also, the models are constrained by the teacher model, which weights do not change a lot either as they are updated with an EMA </li>\n</ul>",
      "rawMarkdown": "You're welcome !\n- The mixed mask is the maximum of the two initial masks, in order to allow the transformer to attend everywhere.\n- You need to have cps close to the same local minimum. In our case lr is quite low during the last 10 epochs (linearly decreased from 4e-5 to 0), and performance does not change a lot. Also, the models are constrained by the teacher model, which weights do not change a lot either as they are updated with an EMA",
      "votes": null
    },
    {
      "id": "2243829",
      "postDate": "05/03/2023 07:49:12",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks, good to know the details! I love mixup as it works every time, but in this competition I discarded it as I couldn't come up with how to handle attention masks. But now I can use mixup in future competitions with transformers!</p>",
      "rawMarkdown": "theoviel Thanks, good to know the details! I love mixup as it works every time, but in this competition I discarded it as I couldn't come up with how to handle attention masks. But now I can use mixup in future competitions with transformers!",
      "votes": null
    },
    {
      "id": "2243872",
      "postDate": "05/03/2023 08:39:09",
      "content": "<p>By the way, we used mix-up as well in our solution, but leaved masks as-is, and it worked fine somehow, one can say that if apply mix-up like this, it's a mix-up with sequence masking.</p>\n<p>I wonder what would happen if we changed masks like Theo did.</p>",
      "rawMarkdown": "By the way, we used mix-up as well in our solution, but leaved masks as-is, and it worked fine somehow, one can say that if apply mix-up like this, it's a mix-up with sequence masking.\n\nI wonder what would happen if we changed masks like Theo did.",
      "votes": null
    },
    {
      "id": "2244656",
      "postDate": "05/03/2023 19:29:34",
      "content": "<p>wow..  the number of tricks is out of this world.. <br>\nHow long did it take you to train all of it? How did you experiment with it? didn't each training session take too long?  <br>\nAgain, amazing work, thank you for sharing!<br>\n[And thank you for sharing all the code, this is a gold mine 🙏 ]</p>\n<p><strong>Edit:</strong> One more question, how did you come up with the layerwise lr values? simple trial and error? </p>",
      "rawMarkdown": "wow..  the number of tricks is out of this world.. \nHow long did it take you to train all of it? How did you experiment with it? didn't each training session take too long?  \nAgain, amazing work, thank you for sharing!\n[And thank you for sharing all the code, this is a gold mine 🙏 ]\n\n\n**Edit:** One more question, how did you come up with the layerwise lr values? simple trial and error?",
      "votes": null
    },
    {
      "id": "2244804",
      "postDate": "05/03/2023 22:25:10",
      "content": "<p>Most tricks were tested on a single fold which took around 1h15 on 8 GPUs (1h without mean teacher). <br>\nThen I ran the full 4 folds + fullfit overnight. With distillation it took 2h per fold. </p>\n<p>I am lucky to have access to a lot of compute power, which helps a lot :)</p>\n<p>Regarding layerwise lr, I couldn't make it work. The overall idea is to give less lr to transformer layers since they have a lot of params improve stability. I tried a few multipliers (x0.5, x0.25), so yes simply trial and error.</p>",
      "rawMarkdown": "Most tricks were tested on a single fold which took around 1h15 on 8 GPUs (1h without mean teacher). \nThen I ran the full 4 folds + fullfit overnight. With distillation it took 2h per fold. \n\nI am lucky to have access to a lot of compute power, which helps a lot :)\n\nRegarding layerwise lr, I couldn't make it work. The overall idea is to give less lr to transformer layers since they have a lot of params improve stability. I tried a few multipliers (x0.5, x0.25), so yes simply trial and error.",
      "votes": null
    },
    {
      "id": "2244925",
      "postDate": "05/04/2023 02:38:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, <br>\nI didn't find a perfect library for GCN, GNN tflite conversion, Would you mind provide your implementation details, <br>\nThanks</p>",
      "rawMarkdown": "Hi @theoviel, \nI didn't find a perfect library for GCN, GNN tflite conversion, Would you mind provide your implementation details, \nThanks",
      "votes": null
    },
    {
      "id": "2246901",
      "postDate": "05/05/2023 14:39:30",
      "content": "<p>Wow that's a great job!! What exactly did you fail with CNN?</p>",
      "rawMarkdown": "Wow that's a great job!! What exactly did you fail with CNN?",
      "votes": null
    },
    {
      "id": "2247654",
      "postDate": "05/06/2023 07:28:18",
      "content": "<p>Amazing! Congrats <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. Thanks for sharing </p>",
      "rawMarkdown": "Amazing! Congrats @theoviel. Thanks for sharing",
      "votes": null
    },
    {
      "id": "2247713",
      "postDate": "05/06/2023 08:23:23",
      "content": "<p>It is really good work! may I asking how to make the image? your model structure figure looks very well.</p>",
      "rawMarkdown": "It is really good work! may I asking how to make the image? your model structure figure looks very well.",
      "votes": null
    },
    {
      "id": "2251194",
      "postDate": "05/09/2023 06:47:15",
      "content": "<p>Thanks ! I used draw.io </p>",
      "rawMarkdown": "Thanks ! I used draw.io",
      "votes": null
    },
    {
      "id": "2261100",
      "postDate": "05/16/2023 05:31:53",
      "content": "<p>This is awesome! </p>",
      "rawMarkdown": "This is awesome!",
      "votes": null
    },
    {
      "id": "3185123",
      "postDate": "04/22/2025 21:33:49",
      "content": "<p>But please couldn't colab also attain compute power as well or no?<br>\nI mean are there limitations to using google colabs?</p>",
      "rawMarkdown": "But please couldn't colab also attain compute power as well or no?\nI mean are there limitations to using google colabs?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2243199,
      "author_name": "smnuruzzaman",
      "author_url": "",
      "post_date": "05/02/2023 18:24:50",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2243438,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "05/02/2023 22:44:00",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks for sharing a very detailed solution!<br>\nI have 2 questions:</p>\n<blockquote>\n  <p>Manifold Mixup [1] (scheduled, p=0.5 * epoch / (0.9 * n_epochs)) : random apply mixup to the features before one of the transformer layer  </p>\n</blockquote>\n<p>First one is about mixup. If you apply mixup before the transformer layer, there should be a temporal dim and attention mask. How did you handle the mask?</p>\n<p>Second question is about model soup. It always brings worse results than a single model every time I try. Is there any secret to making it work? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2243818,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/03/2023 07:28:17",
          "content": "<p>You're welcome !</p>\n<ul>\n<li>The mixed mask is the maximum of the two initial masks, in order to allow the transformer to attend everywhere.</li>\n<li>You need to have cps close to the same local minimum. In our case lr is quite low during the last 10 epochs (linearly decreased from 4e-5 to 0), and performance does not change a lot. Also, the models are constrained by the teacher model, which weights do not change a lot either as they are updated with an EMA </li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2243829,
              "author_name": "bamps53",
              "author_url": "",
              "post_date": "05/03/2023 07:49:12",
              "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks, good to know the details! I love mixup as it works every time, but in this competition I discarded it as I couldn't come up with how to handle attention masks. But now I can use mixup in future competitions with transformers!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2243872,
                  "author_name": "martynoveduard",
                  "author_url": "",
                  "post_date": "05/03/2023 08:39:09",
                  "content": "<p>By the way, we used mix-up as well in our solution, but leaved masks as-is, and it worked fine somehow, one can say that if apply mix-up like this, it's a mix-up with sequence masking.</p>\n<p>I wonder what would happen if we changed masks like Theo did.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2244656,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "05/03/2023 19:29:34",
      "content": "<p>wow..  the number of tricks is out of this world.. <br>\nHow long did it take you to train all of it? How did you experiment with it? didn't each training session take too long?  <br>\nAgain, amazing work, thank you for sharing!<br>\n[And thank you for sharing all the code, this is a gold mine 🙏 ]</p>\n<p><strong>Edit:</strong> One more question, how did you come up with the layerwise lr values? simple trial and error? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2244804,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/03/2023 22:25:10",
          "content": "<p>Most tricks were tested on a single fold which took around 1h15 on 8 GPUs (1h without mean teacher). <br>\nThen I ran the full 4 folds + fullfit overnight. With distillation it took 2h per fold. </p>\n<p>I am lucky to have access to a lot of compute power, which helps a lot :)</p>\n<p>Regarding layerwise lr, I couldn't make it work. The overall idea is to give less lr to transformer layers since they have a lot of params improve stability. I tried a few multipliers (x0.5, x0.25), so yes simply trial and error.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3185123,
              "author_name": "retroace",
              "author_url": "",
              "post_date": "04/22/2025 21:33:49",
              "content": "<p>But please couldn't colab also attain compute power as well or no?<br>\nI mean are there limitations to using google colabs?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2244925,
      "author_name": "gowrishankarp",
      "author_url": "",
      "post_date": "05/04/2023 02:38:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, <br>\nI didn't find a perfect library for GCN, GNN tflite conversion, Would you mind provide your implementation details, <br>\nThanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2246901,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/05/2023 14:39:30",
      "content": "<p>Wow that's a great job!! What exactly did you fail with CNN?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2247654,
      "author_name": "szebiniso",
      "author_url": "",
      "post_date": "05/06/2023 07:28:18",
      "content": "<p>Amazing! Congrats <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. Thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2247713,
      "author_name": "dwchen",
      "author_url": "",
      "post_date": "05/06/2023 08:23:23",
      "content": "<p>It is really good work! may I asking how to make the image? your model structure figure looks very well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2251194,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/09/2023 06:47:15",
          "content": "<p>Thanks ! I used draw.io </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2261100,
      "author_name": "ihsanzami",
      "author_url": "",
      "post_date": "05/16/2023 05:31:53",
      "content": "<p>This is awesome! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2243152": "Thanks to both, the organizers of this competition who offered a fun yet challenging problem as well as all of the other competitors - well done to everyone who worked hard for small incremental increases. \n\nAlthough I am the one posting the topic, this is the result of a great team effort, so big shoutout to @christofhenkel.\n\n# Brief Summary\n\nOur solution is a 2 model ensemble of a MLP-encoder-frame-transformer model. We pushed our transformer models close to the limit and implemented a lot of tricks to climb up to 6th place. \nI have 1403 hours of experiment monitoring time in April (that’s 48h per day :)).\n\n**Update :** Code is available here : https://github.com/TheoViel/kaggle_islr\n\n# Detailed Summary\n\n## Preprocessing & Model\n\n<a href=\"https://ibb.co/g3np4jQ\"><img src=\"https://i.ibb.co/80hpYKG/ISLR-drawio-1.png\" alt=\"ISLR-drawio-1\" border=\"0\"></a>\n\n### Preprocessing\n- Remove frames without fingers\n- Stride the sequence (use 1 every n frames) such that the sequence size is `<= max_len`. We used `max_len=25` and `80` in the final ensemble\n- Normalization is done for the whole sequence to have 0 mean and 1 std. We do an extra centering before the specific MLP. Nan values are set to 0\n\n### Embedding\n- 2x 1D convolutions (`k=5`) to smooth the positions\n- Embed the landmark id and type (e.g. lips, right hand, ...), `embed_dim=dense_dim=16`\n\n### Feature extractor\n- One MLP combining all the features, and 4 for specific landmark types (2x hands, face, lips)\n- Max aggregation for hands is to take into account that signers use one hand\n- `dim=192`, `dropout=0.25`\n\n### Transformer\n- Deberta was better than Bert, but we had to rewrite the attention layer for it to be efficient\n- To reduce the number of parameters, we use a smaller first transformer, and modify the output layer to upscale/downscale the features. This was key to enable blending 2 models\n- `d_in=512`, `δ=256` for `max_len=25`, `δ=64` for `max_len=80`, `num_heads=16`, `dropout=0.05` for the first layer, `0.1` for the other 2\n- Unfortunately, we did not have any luck with using a pre-trained version of Deberta, for example by importing some of the pretrained weights\n\n## Training strategy\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/VvT0pVx/ISLR-Train-drawio-1.png\" alt=\"ISLR-Train-drawio-1\" border=\"0\"></a>\n\n### Augmentations:\n- Horizontal Flip (`p=0.5`)\n- Rotate around `(0, 0, 0)` by an angle between -60 and 60°  (`p=0.5`)\n- Resizing by a factor in `(0.7, 1.3)` (`p=0.5`). We also allow for distortion (`p=0.5`)\n- Crop 20% of the start or end (`p=0.5`)\n- Interpolate to fill missing values (`p=0.5`)\n- Manifold Mixup [[1]](https://arxiv.org/abs/1806.05236) (scheduled, `p=0.5 * epoch / (0.9 * n_epochs)`) : random apply mixup to the features before one of the transformer layer\n- Only during the first half of the training, since it improved convergence\n  - Fill the value of the missing hand with those of the existing one (`p=0.25`)\n  - Face CutMix : replace the face landmarks with those of another signer doing the same sign (`p=0.25`)\n\n### Training\n- 100 epochs, `lr=3e-4`, 25% warmup, linear schedule\n- Cross entropy with smoothing (`eps=0.3`)\n- `weight_decay=0.4`, `batch_size=32` (x8 GPUs)\n- OUSM [[2]](https://arxiv.org/pdf/1901.07759.pdf), i.e. exclude the top k (`k=3`) samples with the highest loss from the computation\n- Mean teacher [[3]](https://arxiv.org/abs/1703.01780) & Knowledge Distillation (see image above). We train 3 models at the same time, and use the distilled one for inference\n- Model soup [[4]](https://arxiv.org/abs/2203.05482) of the last 10 epochs checkpoints\n\n## Experiments \n\n<a href=\"https://ibb.co/MV343Xx\"><img src=\"https://i.ibb.co/jM121pP/cvlb.png\" alt=\"cvlb\" border=\"0\"></a>\n\n###  Validation strategy\nWe use a Stratified 4-fold, grouped by patient. We had a great CV - LB correlation (see figure), and our best model achieved CV 0.749 - Public 0.795 - Private 0.877. We submitted single models trained on the full dataset, and 2 models for ensembles.\n\n###  What worked & reported improvements\nStarting from early April :\n- Baseline : MLP + Bert : LB 0.72\n- Deberta instead of Bert : CV +0.015  **-> LB 0.73**\n- Improve Feature extractor : CV +0.01  **-> LB 0.74**\n- Flip augmentation : CV +0.015  **-> LB 0.75**\n- Increase model size : CV +0.006 \n- Two stage training + interp aug : CV + 0.002 ?\n- Increase model size  : CV +0.005  **-> LB 0.76**\n- Deberta Rework : CV +0.005   **-> LB 0.77**\n- OUSM : CV +0.003\n- Mean Teacher : CV +0.005\n- Ensemble two distilled models : +0.005 **-> LB 0.78**\n- Conv layers : CV +0.001\n- Mish [[5]](https://arxiv.org/abs/1908.08681) activation instead of ReLU/LeakyReLU : CV +0.001\n- Warmup 0.1 -> 0.25 : CV +0.001\n- Model soup : CV +0.001\n- [Nobuco](https://github.com/AlexanderLutsenko/nobuco) instead of onnx : 45ms/it -> 38ms/it, this enabled to improve the ensemble \n- Weight decay 0 -> 0.4 : CV +0.004\n- Truncation aug : CV +0.001\n- Manifold mixup : CV +0.002\n- Different input, model and teacher size for ensemble : CV +0.005 **-> LB 0.79**\n- Centering before the MLP layers : CV +0.004  **-> LB 0.795  (final)**\n\n###  What did not work (for us)\n- Unfortunately we could not make CNNs work\n- GCNs, or other architectures that are “sota” for ASL\n- Heng’s architecture, although others had great results with it\n- Getting a successful 3-model sub\n- Relabel the data to reduce label noise\n- Other augmentations such as noise, dropout, drop frame, shift\n- Dropout scheduling, custom lr per layer\n- Pretraining part of the model\n- Handcrafted graph features (adjacency matrices, edge features)\n- Adding more landmarks (eyes, eyebrows)\n\n\n*Thanks for reading !*",
    "2243199": "Thank you for sharing",
    "2243438": "theoviel Thanks for sharing a very detailed solution!\nI have 2 questions:\n\n> Manifold Mixup [1] (scheduled, p=0.5 * epoch / (0.9 * n_epochs)) : random apply mixup to the features before one of the transformer layer  \n\nFirst one is about mixup. If you apply mixup before the transformer layer, there should be a temporal dim and attention mask. How did you handle the mask?\n\nSecond question is about model soup. It always brings worse results than a single model every time I try. Is there any secret to making it work?",
    "2243818": "You're welcome !\n- The mixed mask is the maximum of the two initial masks, in order to allow the transformer to attend everywhere.\n- You need to have cps close to the same local minimum. In our case lr is quite low during the last 10 epochs (linearly decreased from 4e-5 to 0), and performance does not change a lot. Also, the models are constrained by the teacher model, which weights do not change a lot either as they are updated with an EMA",
    "2243829": "theoviel Thanks, good to know the details! I love mixup as it works every time, but in this competition I discarded it as I couldn't come up with how to handle attention masks. But now I can use mixup in future competitions with transformers!",
    "2243872": "By the way, we used mix-up as well in our solution, but leaved masks as-is, and it worked fine somehow, one can say that if apply mix-up like this, it's a mix-up with sequence masking.\n\nI wonder what would happen if we changed masks like Theo did.",
    "2244656": "wow..  the number of tricks is out of this world.. \nHow long did it take you to train all of it? How did you experiment with it? didn't each training session take too long?  \nAgain, amazing work, thank you for sharing!\n[And thank you for sharing all the code, this is a gold mine 🙏 ]\n\n\n**Edit:** One more question, how did you come up with the layerwise lr values? simple trial and error?",
    "2244804": "Most tricks were tested on a single fold which took around 1h15 on 8 GPUs (1h without mean teacher). \nThen I ran the full 4 folds + fullfit overnight. With distillation it took 2h per fold. \n\nI am lucky to have access to a lot of compute power, which helps a lot :)\n\nRegarding layerwise lr, I couldn't make it work. The overall idea is to give less lr to transformer layers since they have a lot of params improve stability. I tried a few multipliers (x0.5, x0.25), so yes simply trial and error.",
    "2244925": "Hi @theoviel, \nI didn't find a perfect library for GCN, GNN tflite conversion, Would you mind provide your implementation details, \nThanks",
    "2246901": "Wow that's a great job!! What exactly did you fail with CNN?",
    "2247654": "Amazing! Congrats @theoviel. Thanks for sharing",
    "2247713": "It is really good work! may I asking how to make the image? your model structure figure looks very well.",
    "2251194": "Thanks ! I used draw.io",
    "2261100": "This is awesome!",
    "3185123": "But please couldn't colab also attain compute power as well or no?\nI mean are there limitations to using google colabs?"
  },
  "source": "meta"
}