{
  "id": 406343,
  "title": "9th place solution",
  "url": "/competitions/asl-signs/writeups/str-9th-place-solution",
  "author_name": "",
  "post_date": "2023-05-03T18:01:09.943Z",
  "votes": 26,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks Google and Kaggle for organizing this competition, I hope our solutions will help to create a PopSign product and lower the entering barrier to ASL. </p>\n<p>Below I'll give a brief overview of our solution, but before that I want to congratulate my teammate <a href=\"https://www.kaggle.com/timriggins\" target=\"_blank\">@timriggins</a>, with this competition he became Competition Grandmaster! The fun part is that we met during Kaggle Days World Championship in Barcelona last October, once again thanks HP and Kaggle Days for such an opportunity :D</p>\n<h2>Model</h2>\n<p>We used a transformer model which is very similar to the one <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared<br>\nThings we changed in architecture/features:</p>\n<ul>\n<li>1 -&gt; 2 transformer layers (in the final ensemble we used 3 1-layer models and 1 2-layer model)</li>\n<li>aside from motion and distance features for hands, our models use lip and eye distances, <br>\nabsolute hand coordinates <code>abdcoords_r = (rhand - np.nanmean(rhand, 0))</code> <br>\nand acceleration <code>racc = np.pad(rh_dist[1:] - rh_dist[:-1], [[0,1], [0,0], [0,0]])</code></li>\n</ul>\n<p>In two models we used absolute learnable (init. from sinusoidal) positional encoding, <br>\nin other two we used learnable 1d-depthwise convolution as pe encoding </p>\n<pre><code>self.pe = nn.Conv1d(emb_dim, emb_dim, kernel_size=5, padding=2, stride=1, groups=emb_dim)\n## x.shape (B, L, E)\nx = self.pe(x.permute(0, 2, 1)).permute(0, 2, 1) + x\n</code></pre>\n<h2>Augmentations</h2>\n<p>Empirically we found following augmentations to be effective:</p>\n<ul>\n<li>flip-aug</li>\n<li>random fps drop <code>xyz = xyz[::2]</code></li>\n<li>affine 3d - rotate around spine</li>\n<li>early mixup (keypoint space) &amp; manifold mixup (embedding space)</li>\n<li>random cuts from start-end of the frame series</li>\n<li>interpolation</li>\n</ul>\n<p>Mix-up was an important augmentation, it allowed us to use longer training schedule w/ cosine scheduler without overfit, all of our models were trained for 150 epochs and converged on the last epoch.</p>\n<h2>Pytorch to TFLite conversion</h2>\n<p>We used <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">nobuco</a>, check it out, it is very intuitive </p>\n<h2>Fun bug</h2>\n<p>I made a bug in a code that led to a better performance, when implementing a mixup augmentation I did this in the model forward pass function, - </p>\n<pre><code>if np.random.randn() &lt; 0.5 and mode_flag == 'train':\n    x, labels_ohe = self.mixup(x,labels_ohe)\n</code></pre>\n<p>Later I found out that I'm using normal distribution instead of a uniform one thus if we imagine a PDF of normal disribution, we can say that we're using mixup in ~70% of cases, and when I corrected it, the performance dropped! (so it's better to use mixup with higher probability) </p>\n<h2>Results</h2>\n<p>Best single model reaches 0.7316 on the hengck participant split, 0.754 if using 4-model ensemble, CV/LB correlation was perfect and each CV improvement translated to both public and private LB. Single model gets scored in approx. 30 mins and ensemble gets scored very close to an hour. I was very worried, since we made last submission on the final day of the competition, but thankfully it didn't timeout :D</p>\n<h2>What we tried, but didn't work</h2>\n<ul>\n<li>External data, - WLASL-pretrain/adding WLASL to train/pretrain on greek sign language dataset, neither of those gave performance boost, even decreased compared to original initialization</li>\n<li>Continous MLM/deep predictive coding pretraining - I think that the model is too shallow to benefit for pre-training, I would like to hear from other participants if they used it in their solution</li>\n<li>Using 1d-CNN model, - we tried 1d-cnn with depthwise-separable convolutions </li>\n<li>Arcface</li>\n<li>Train-time overparametrization (for linear layers), concept description can be found in MobileOne paper </li>\n</ul>",
  "messages": [
    {
      "id": "2242089",
      "postDate": "05/02/2023 03:16:36",
      "content": "<p>Thanks Google and Kaggle for organizing this competition, I hope our solutions will help to create a PopSign product and lower the entering barrier to ASL. </p>\n<p>Below I'll give a brief overview of our solution, but before that I want to congratulate my teammate <a href=\"https://www.kaggle.com/timriggins\" target=\"_blank\">@timriggins</a>, with this competition he became Competition Grandmaster! The fun part is that we met during Kaggle Days World Championship in Barcelona last October, once again thanks HP and Kaggle Days for such an opportunity :D</p>\n<h2>Model</h2>\n<p>We used a transformer model which is very similar to the one <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared<br>\nThings we changed in architecture/features:</p>\n<ul>\n<li>1 -&gt; 2 transformer layers (in the final ensemble we used 3 1-layer models and 1 2-layer model)</li>\n<li>aside from motion and distance features for hands, our models use lip and eye distances, <br>\nabsolute hand coordinates <code>abdcoords_r = (rhand - np.nanmean(rhand, 0))</code> <br>\nand acceleration <code>racc = np.pad(rh_dist[1:] - rh_dist[:-1], [[0,1], [0,0], [0,0]])</code></li>\n</ul>\n<p>In two models we used absolute learnable (init. from sinusoidal) positional encoding, <br>\nin other two we used learnable 1d-depthwise convolution as pe encoding </p>\n<pre><code>self.pe = nn.Conv1d(emb_dim, emb_dim, kernel_size=5, padding=2, stride=1, groups=emb_dim)\n## x.shape (B, L, E)\nx = self.pe(x.permute(0, 2, 1)).permute(0, 2, 1) + x\n</code></pre>\n<h2>Augmentations</h2>\n<p>Empirically we found following augmentations to be effective:</p>\n<ul>\n<li>flip-aug</li>\n<li>random fps drop <code>xyz = xyz[::2]</code></li>\n<li>affine 3d - rotate around spine</li>\n<li>early mixup (keypoint space) &amp; manifold mixup (embedding space)</li>\n<li>random cuts from start-end of the frame series</li>\n<li>interpolation</li>\n</ul>\n<p>Mix-up was an important augmentation, it allowed us to use longer training schedule w/ cosine scheduler without overfit, all of our models were trained for 150 epochs and converged on the last epoch.</p>\n<h2>Pytorch to TFLite conversion</h2>\n<p>We used <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">nobuco</a>, check it out, it is very intuitive </p>\n<h2>Fun bug</h2>\n<p>I made a bug in a code that led to a better performance, when implementing a mixup augmentation I did this in the model forward pass function, - </p>\n<pre><code>if np.random.randn() &lt; 0.5 and mode_flag == 'train':\n    x, labels_ohe = self.mixup(x,labels_ohe)\n</code></pre>\n<p>Later I found out that I'm using normal distribution instead of a uniform one thus if we imagine a PDF of normal disribution, we can say that we're using mixup in ~70% of cases, and when I corrected it, the performance dropped! (so it's better to use mixup with higher probability) </p>\n<h2>Results</h2>\n<p>Best single model reaches 0.7316 on the hengck participant split, 0.754 if using 4-model ensemble, CV/LB correlation was perfect and each CV improvement translated to both public and private LB. Single model gets scored in approx. 30 mins and ensemble gets scored very close to an hour. I was very worried, since we made last submission on the final day of the competition, but thankfully it didn't timeout :D</p>\n<h2>What we tried, but didn't work</h2>\n<ul>\n<li>External data, - WLASL-pretrain/adding WLASL to train/pretrain on greek sign language dataset, neither of those gave performance boost, even decreased compared to original initialization</li>\n<li>Continous MLM/deep predictive coding pretraining - I think that the model is too shallow to benefit for pre-training, I would like to hear from other participants if they used it in their solution</li>\n<li>Using 1d-CNN model, - we tried 1d-cnn with depthwise-separable convolutions </li>\n<li>Arcface</li>\n<li>Train-time overparametrization (for linear layers), concept description can be found in MobileOne paper </li>\n</ul>",
      "rawMarkdown": "Thanks Google and Kaggle for organizing this competition, I hope our solutions will help to create a PopSign product and lower the entering barrier to ASL. \n\nBelow I'll give a brief overview of our solution, but before that I want to congratulate my teammate @timriggins, with this competition he became Competition Grandmaster! The fun part is that we met during Kaggle Days World Championship in Barcelona last October, once again thanks HP and Kaggle Days for such an opportunity :D\n\n## Model\n\nWe used a transformer model which is very similar to the one @hengck23 shared\nThings we changed in architecture/features:\n* 1 -> 2 transformer layers (in the final ensemble we used 3 1-layer models and 1 2-layer model)\n* aside from motion and distance features for hands, our models use lip and eye distances, \nabsolute hand coordinates ```abdcoords_r = (rhand - np.nanmean(rhand, 0))``` \nand acceleration ``` racc = np.pad(rh_dist[1:] - rh_dist[:-1], [[0,1], [0,0], [0,0]]) ```\n\nIn two models we used absolute learnable (init. from sinusoidal) positional encoding, \nin other two we used learnable 1d-depthwise convolution as pe encoding \n```\nself.pe = nn.Conv1d(emb_dim, emb_dim, kernel_size=5, padding=2, stride=1, groups=emb_dim)\n## x.shape (B, L, E)\nx = self.pe(x.permute(0, 2, 1)).permute(0, 2, 1) + x\n```\n## Augmentations\n\nEmpirically we found following augmentations to be effective:\n* flip-aug\n* random fps drop ```xyz = xyz[::2]```\n* affine 3d - rotate around spine\n* early mixup (keypoint space) & manifold mixup (embedding space)\n* random cuts from start-end of the frame series\n* interpolation\n\nMix-up was an important augmentation, it allowed us to use longer training schedule w/ cosine scheduler without overfit, all of our models were trained for 150 epochs and converged on the last epoch.\n\n## Pytorch to TFLite conversion\n\nWe used [nobuco](https://github.com/AlexanderLutsenko/nobuco), check it out, it is very intuitive \n\n## Fun bug\n\nI made a bug in a code that led to a better performance, when implementing a mixup augmentation I did this in the model forward pass function, - \n``` \nif np.random.randn() < 0.5 and mode_flag == 'train':\n\tx, labels_ohe = self.mixup(x,labels_ohe)\n   ```\nLater I found out that I'm using normal distribution instead of a uniform one thus if we imagine a PDF of normal disribution, we can say that we're using mixup in ~70% of cases, and when I corrected it, the performance dropped! (so it's better to use mixup with higher probability) \n\n\n## Results\n\nBest single model reaches 0.7316 on the hengck participant split, 0.754 if using 4-model ensemble, CV/LB correlation was perfect and each CV improvement translated to both public and private LB. Single model gets scored in approx. 30 mins and ensemble gets scored very close to an hour. I was very worried, since we made last submission on the final day of the competition, but thankfully it didn't timeout :D\n\n## What we tried, but didn't work\n* External data, - WLASL-pretrain/adding WLASL to train/pretrain on greek sign language dataset, neither of those gave performance boost, even decreased compared to original initialization\n* Continous MLM/deep predictive coding pretraining - I think that the model is too shallow to benefit for pre-training, I would like to hear from other participants if they used it in their solution\n* Using 1d-CNN model, - we tried 1d-cnn with depthwise-separable convolutions \n* Arcface\n* Train-time overparametrization (for linear layers), concept description can be found in MobileOne paper",
      "votes": null
    },
    {
      "id": "2242145",
      "postDate": "05/02/2023 04:16:03",
      "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> thanks for topic. And congrats with 9th place!</p>",
      "rawMarkdown": "martynoveduard thanks for topic. And congrats with 9th place!",
      "votes": null
    },
    {
      "id": "2242807",
      "postDate": "05/02/2023 14:20:37",
      "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> great result! 👍 Congratulations!🎉 </p>",
      "rawMarkdown": "martynoveduard great result! 👍 Congratulations!🎉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242145,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "05/02/2023 04:16:03",
      "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> thanks for topic. And congrats with 9th place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2242807,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/02/2023 14:20:37",
      "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> great result! 👍 Congratulations!🎉 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2242089": "Thanks Google and Kaggle for organizing this competition, I hope our solutions will help to create a PopSign product and lower the entering barrier to ASL. \n\nBelow I'll give a brief overview of our solution, but before that I want to congratulate my teammate @timriggins, with this competition he became Competition Grandmaster! The fun part is that we met during Kaggle Days World Championship in Barcelona last October, once again thanks HP and Kaggle Days for such an opportunity :D\n\n## Model\n\nWe used a transformer model which is very similar to the one @hengck23 shared\nThings we changed in architecture/features:\n* 1 -> 2 transformer layers (in the final ensemble we used 3 1-layer models and 1 2-layer model)\n* aside from motion and distance features for hands, our models use lip and eye distances, \nabsolute hand coordinates ```abdcoords_r = (rhand - np.nanmean(rhand, 0))``` \nand acceleration ``` racc = np.pad(rh_dist[1:] - rh_dist[:-1], [[0,1], [0,0], [0,0]]) ```\n\nIn two models we used absolute learnable (init. from sinusoidal) positional encoding, \nin other two we used learnable 1d-depthwise convolution as pe encoding \n```\nself.pe = nn.Conv1d(emb_dim, emb_dim, kernel_size=5, padding=2, stride=1, groups=emb_dim)\n## x.shape (B, L, E)\nx = self.pe(x.permute(0, 2, 1)).permute(0, 2, 1) + x\n```\n## Augmentations\n\nEmpirically we found following augmentations to be effective:\n* flip-aug\n* random fps drop ```xyz = xyz[::2]```\n* affine 3d - rotate around spine\n* early mixup (keypoint space) & manifold mixup (embedding space)\n* random cuts from start-end of the frame series\n* interpolation\n\nMix-up was an important augmentation, it allowed us to use longer training schedule w/ cosine scheduler without overfit, all of our models were trained for 150 epochs and converged on the last epoch.\n\n## Pytorch to TFLite conversion\n\nWe used [nobuco](https://github.com/AlexanderLutsenko/nobuco), check it out, it is very intuitive \n\n## Fun bug\n\nI made a bug in a code that led to a better performance, when implementing a mixup augmentation I did this in the model forward pass function, - \n``` \nif np.random.randn() < 0.5 and mode_flag == 'train':\n\tx, labels_ohe = self.mixup(x,labels_ohe)\n   ```\nLater I found out that I'm using normal distribution instead of a uniform one thus if we imagine a PDF of normal disribution, we can say that we're using mixup in ~70% of cases, and when I corrected it, the performance dropped! (so it's better to use mixup with higher probability) \n\n\n## Results\n\nBest single model reaches 0.7316 on the hengck participant split, 0.754 if using 4-model ensemble, CV/LB correlation was perfect and each CV improvement translated to both public and private LB. Single model gets scored in approx. 30 mins and ensemble gets scored very close to an hour. I was very worried, since we made last submission on the final day of the competition, but thankfully it didn't timeout :D\n\n## What we tried, but didn't work\n* External data, - WLASL-pretrain/adding WLASL to train/pretrain on greek sign language dataset, neither of those gave performance boost, even decreased compared to original initialization\n* Continous MLM/deep predictive coding pretraining - I think that the model is too shallow to benefit for pre-training, I would like to hear from other participants if they used it in their solution\n* Using 1d-CNN model, - we tried 1d-cnn with depthwise-separable convolutions \n* Arcface\n* Train-time overparametrization (for linear layers), concept description can be found in MobileOne paper",
    "2242145": "martynoveduard thanks for topic. And congrats with 9th place!",
    "2242807": "martynoveduard great result! 👍 Congratulations!🎉"
  },
  "source": "meta"
}