{
  "id": 406306,
  "title": "2nd Place Solution | Google - Isolated Sign Language Recognition",
  "url": "/competitions/asl-signs/discussion/406306",
  "author_name": "Kolya Forrat",
  "post_date": "2023-05-02T00:16:47.837000",
  "votes": 118,
  "comment_count": 48,
  "views": 0,
  "content": "<h2>TLDR</h2>\n<p>We used an approach similar to audio spectrogram classification using the EfficientNet-B0 model, with numerous augmentations and transformer models such as BERT and DeBERTa as helper models. The final solution consists of one EfficientNet-B0 with an input size of 160x80, trained on a single fold from 8 randomly split folds, as well as DeBERTa and BERT trained on the full dataset. A single fold model using EfficientNet has a CV score of 0.898 and a leaderboard score of ~0.8.</p>\n<p>We used only competition data.</p>\n<h2>1. Data Preprocessing</h2>\n<h3>1.1 CNN Preprocessing</h3>\n<ul>\n<li>We extracted 18 lip points, 20 pose points (including arms, shoulders, eyebrows, and nose), and all hand points, resulting in a total of 80 points.</li>\n<li>During training, we applied various augmentations.</li>\n<li>We implemented standard normalization.</li>\n<li>Instead of dropping NaN values, we filled them with zeros after normalization.</li>\n<li>We interpolated the time axis to a size of 160 using 'nearest' interpolation: <code>yy = F.interpolate(yy[None, None, :], size=self.new_size, mode='nearest')</code>.</li>\n<li>Finally, we obtained a tensor with dimensions 160x80x3, where 3 represents the <code>(X, Y, Z)</code> axes. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Fc47290891a7ac6497a6a0c296973f071%2Fdata_prep.jpg?generation=1682986293067532&amp;alt=media\" alt=\"Preprocessing\"></li>\n</ul>\n<h3>1.2 Transformer Preprocessing</h3>\n<ul>\n<li><p>Only 61 points were kept, including 40 lip points and 21 hand points. For left and right hand, the one with less NaN was kept. If right hand was kept, mirror it to left hand.</p></li>\n<li><p>Augmentations, normalization and NaN-filling were applied sequentially.</p></li>\n<li><p>Sequences longer than 96 were interpolated to 96. Sequences shorter than 96 were unchanged.</p></li>\n<li><p>Apart from raw positions, hand-crafted features were also used, including motion, distances, and cosine of angles.</p></li>\n<li><p>Motion features consist of future motion and history motion, which can be denoted as:</p></li>\n</ul>\n<p>$$<br>\n  Motion_{future} = position_{t+1} - position_{t}<br>\n$$<br>\n$$<br>\n  Motion_{history} = position_{t} - position_{t-1}<br>\n$$</p>\n<ul>\n<li><p>Full 210 pairwise distances among 21 hand points were included. </p></li>\n<li><p>There are 5 vertices in a finger (e.g. thumb is <code>[0,1,2,3,4]</code>), and therefore, there are 3 angles: <code>&lt;0,1,2&gt;, &lt;1,2,3&gt;, &lt;2,3,4&gt;</code>. So 15 angles of 5 fingers were included.</p></li>\n<li><p>Randomly selected 190 pairwise distances and randomly selected 8 angles among 40 lip points were included.</p></li>\n</ul>\n<h2>2. Augmentation</h2>\n<h3>2.1 Common Augmentations</h3>\n<blockquote>\n  <p>These augmentations are used in both CNN training and transformer training</p>\n</blockquote>\n<ol>\n<li><p><code>Random affine</code>: Same as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared. In CNN, after global affine, shift-scale-rotate was also applied to each part separately (e.g. hand, lip, body-pose).</p></li>\n<li><p><code>Random interpolation</code>: Slightly scale and shift the time dimension.</p></li>\n<li><p><code>Flip pose</code>: Flip the x-coordinates of all points. In CNN, <code>x_new = x_max - x_old</code>. In transformer, <code>x_new = 2 * frame[:,0,0] - x_old</code>.</p></li>\n<li><p><code>Finger tree rotate</code>: There are 4 root-children pairs in a finger with 5-vertices. E.g. in thumb (<code>[0,1,2,3,4]</code>), these 4 root-children pairs are: <code>0-[1,2,3,4]</code>,<code>1-[2,3,4]</code>,<code>2-[3,4]</code>,<code>3-[4]</code>. We randomly choose some of these pairs, and rotate the children points around root point with a small random angle.</p></li>\n</ol>\n<h3>2.2 CNN Specific Augmentations</h3>\n<ul>\n<li><code>Mixup</code>: Implement basic mixup augmentation (only works with CNNs, not transformers).</li>\n<li><code>Replace augmentation</code>: Replace some random parts from other samples of the same class.</li>\n<li><code>Time and frequence masking</code>: This basic torchaudio augmentation works exceptionally well.</li>\n</ul>\n<pre><code>freq_m = torchaudio.transforms.FrequencyMasking()  \ntime_m = torchaudio.transforms.TimeMasking()       \n</code></pre>\n<h3>2.3 Augmented Sample Example</h3>\n<p>Before augmentation:</p>\n<p><img alt=\"aug1\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F5d2b97cf6754c5f4063724181bfe7172%2Fbefore_aug.png?generation=1682986332091937&amp;alt=media\"> </p>\n<p>After augmentation:</p>\n<p> <img alt=\"aug2\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F516f0790f6f9903ad8943977bc392965%2Fafter_aug.png?generation=1682986413818912&amp;alt=media\"> </p>\n<h2>3. Training</h2>\n<h3>3.1 CNN Training</h3>\n<ul>\n<li>Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters</li>\n<li>Onecycle scheduler with 0.1 warmup.</li>\n<li>Use weighted <code>CrossEntropyLoss</code>. Increase the weights for poorly predicted classes and classes with semantically similar pairs (such as kitty and cat)</li>\n<li>Implement a hypercolumn for EfficientNet with 5 blocks</li>\n</ul>\n<h3>3.2 Transformer Training</h3>\n<ul>\n<li>Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters</li>\n<li>Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule.</li>\n<li>A 4-layer, 256 hidden-size, 512 intermediate-size transformer were trained.</li>\n<li>A 3-layer model was initialized with 4-layer model's first 3 layers. Knowledge distillation were used in 3-layer model training, in which the 4-layer model is the teacher.</li>\n</ul>\n<h3>3.3 Hyperparameter Tuning</h3>\n<p>Since we trained only one fold and used smaller models, we decided to tune most parameters with Optuna. </p>\n<p>Here is the parameters list of CNN training (transformer training has a similar param-list):</p>\n<ul>\n<li><p>All augmentations probabilities (0.1 - 0.5+)</p></li>\n<li><p>Learning rate (2e-3 - 3e-3)</p></li>\n<li><p>Drop out (0.1 - 0.25)</p></li>\n<li><p>Num of epochs (170-185)</p></li>\n<li><p>Loss weights powers (0.75 - 2)</p></li>\n<li><p>Optimizer (<code>Lookahead_RAdam</code>, <code>RAdam</code>)</p></li>\n<li><p>Label smoothing (0.5 - 0.7)</p></li>\n</ul>\n<h2>4. Submissions, Conversion and Ensemble</h2>\n<ol>\n<li><p>We rewrote all our models in Keras and transferred PyTorch weights to them, resulting in a speed boost of around 30%. For transformer model, pytorch-onnx-tf-tflite will generate too much useless tensor shape operations, a fully rewriting can reduce these manually. For CNN model, we rewrote DepthwiseConv2D with a hard-coded way, whose speed is 200%~300% of its original version of tflite DepthwiseConv2D.</p></li>\n<li><p>After that, we aggregated all these models in the <code>tf.Module</code> class. Converting directly from Keras resulted in lower speed (don't know why).</p></li>\n<li><p>We calculated ensemble weights for models trained on fold 0 using the local fold 0 score and applied these weights to the full dataset models.</p></li>\n</ol>\n<p>EfficientNet-B0 achieved a leaderboard score of approximately 0.8, and transformers improved the score to 0.81. The final ensemble included:</p>\n<ol>\n<li>Efficientnet-B0, fold 0</li>\n<li>BERT, full data train</li>\n<li>DeBERTa, full data train</li>\n</ol>\n<p>Interestingly, a key feature was using the ensemble without softmax, which consistently provided a boost of around 0.01.</p>\n<h2>5. PS. Need <strong>BETTER</strong> TFlite DepthwiseConv2D</h2>\n<p>Depthwise convolution models performed very well for these tasks, outperforming other CNN and ViT models (rexnet_100 was also good).<br>\nWe spent a lot of time dealing with the conversion of DepthwiseConv2D operation. Here are some strange results:</p>\n<p>Given a input image with 82x42x32 (HWC), there are two ways to do a 3x3 depthwise convolution in Keras. One is <code>Conv2D(32, 3, groups = 32)</code>, the other is <code>DepthwiseConv2D(3)</code>. However, after converting these two to tflite, the running time of the <code>Conv2D</code> is 5.05ms, and the running time of <code>DepthwiseConv2D</code> is 3.70ms. More strangely, a full convolution <code>Conv2D(32, 3, groups = 1)</code> with FLOPs = HWC^2 only takes 2.09ms, even faster than previous two with FLOPs = HWC.</p>\n<p>Then we rewrote the depthwise-conv like this:</p>\n<pre><code>     ():\n        out = x[:,:self.H_out:self.strides,:self.W_out:self.strides] * self.weight[,]\n         i  (self.kernel_size):\n             j  (self.kernel_size):\n                 i ==   j == :\n                    \n                out += x[:,i:self.H_out + i:self.strides,j:self.W_out + j:self.strides] * self.weight[i,j]\n         self.bias   :\n            out = out + self.bias\n         out\n</code></pre>\n<p>The running time of this is 1.24 ms.</p>\n<p>In summary, our version (1.24ms) &gt; full <code>Conv2D</code> with larger FLOPs (2.09ms) &gt; <code>DepthwiseConv2D</code> (3.70ms) &gt; <code>Conv2D(C, groups = C)</code> (5.05ms).</p>\n<p>However, our version introduced too much nodes in tflite graph, which is not stable in running time. If the tensorflow team has a better implementation of DepthwiseConv2D, we can even ensemble two CNN models, which is expected to reach 0.82 LB.</p>\n<p>By the way, EfficientNet with ONNX was ~5 times faster than TFLite.</p>\n<h3>Big thanks to my teammates <a href=\"https://www.kaggle.com/artemtprv\" target=\"_blank\">@artemtprv</a> and <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> and congrats with new tiers, Master and GrandMaster!</h3>\n<p><a href=\"https://github.com/ffs333/2nd_place_GISLR\" target=\"_blank\">github code</a></p>",
  "messages": [
    {
      "id": 2241899,
      "postDate": "2023-05-02T00:16:47.837Z",
      "content": "<h2>TLDR</h2>\n<p>We used an approach similar to audio spectrogram classification using the EfficientNet-B0 model, with numerous augmentations and transformer models such as BERT and DeBERTa as helper models. The final solution consists of one EfficientNet-B0 with an input size of 160x80, trained on a single fold from 8 randomly split folds, as well as DeBERTa and BERT trained on the full dataset. A single fold model using EfficientNet has a CV score of 0.898 and a leaderboard score of ~0.8.</p>\n<p>We used only competition data.</p>\n<h2>1. Data Preprocessing</h2>\n<h3>1.1 CNN Preprocessing</h3>\n<ul>\n<li>We extracted 18 lip points, 20 pose points (including arms, shoulders, eyebrows, and nose), and all hand points, resulting in a total of 80 points.</li>\n<li>During training, we applied various augmentations.</li>\n<li>We implemented standard normalization.</li>\n<li>Instead of dropping NaN values, we filled them with zeros after normalization.</li>\n<li>We interpolated the time axis to a size of 160 using 'nearest' interpolation: <code>yy = F.interpolate(yy[None, None, :], size=self.new_size, mode='nearest')</code>.</li>\n<li>Finally, we obtained a tensor with dimensions 160x80x3, where 3 represents the <code>(X, Y, Z)</code> axes. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Fc47290891a7ac6497a6a0c296973f071%2Fdata_prep.jpg?generation=1682986293067532&amp;alt=media\" alt=\"Preprocessing\"></li>\n</ul>\n<h3>1.2 Transformer Preprocessing</h3>\n<ul>\n<li><p>Only 61 points were kept, including 40 lip points and 21 hand points. For left and right hand, the one with less NaN was kept. If right hand was kept, mirror it to left hand.</p></li>\n<li><p>Augmentations, normalization and NaN-filling were applied sequentially.</p></li>\n<li><p>Sequences longer than 96 were interpolated to 96. Sequences shorter than 96 were unchanged.</p></li>\n<li><p>Apart from raw positions, hand-crafted features were also used, including motion, distances, and cosine of angles.</p></li>\n<li><p>Motion features consist of future motion and history motion, which can be denoted as:</p></li>\n</ul>\n<p>$$<br>\n  Motion_{future} = position_{t+1} - position_{t}<br>\n$$<br>\n$$<br>\n  Motion_{history} = position_{t} - position_{t-1}<br>\n$$</p>\n<ul>\n<li><p>Full 210 pairwise distances among 21 hand points were included. </p></li>\n<li><p>There are 5 vertices in a finger (e.g. thumb is <code>[0,1,2,3,4]</code>), and therefore, there are 3 angles: <code>&lt;0,1,2&gt;, &lt;1,2,3&gt;, &lt;2,3,4&gt;</code>. So 15 angles of 5 fingers were included.</p></li>\n<li><p>Randomly selected 190 pairwise distances and randomly selected 8 angles among 40 lip points were included.</p></li>\n</ul>\n<h2>2. Augmentation</h2>\n<h3>2.1 Common Augmentations</h3>\n<blockquote>\n  <p>These augmentations are used in both CNN training and transformer training</p>\n</blockquote>\n<ol>\n<li><p><code>Random affine</code>: Same as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared. In CNN, after global affine, shift-scale-rotate was also applied to each part separately (e.g. hand, lip, body-pose).</p></li>\n<li><p><code>Random interpolation</code>: Slightly scale and shift the time dimension.</p></li>\n<li><p><code>Flip pose</code>: Flip the x-coordinates of all points. In CNN, <code>x_new = x_max - x_old</code>. In transformer, <code>x_new = 2 * frame[:,0,0] - x_old</code>.</p></li>\n<li><p><code>Finger tree rotate</code>: There are 4 root-children pairs in a finger with 5-vertices. E.g. in thumb (<code>[0,1,2,3,4]</code>), these 4 root-children pairs are: <code>0-[1,2,3,4]</code>,<code>1-[2,3,4]</code>,<code>2-[3,4]</code>,<code>3-[4]</code>. We randomly choose some of these pairs, and rotate the children points around root point with a small random angle.</p></li>\n</ol>\n<h3>2.2 CNN Specific Augmentations</h3>\n<ul>\n<li><code>Mixup</code>: Implement basic mixup augmentation (only works with CNNs, not transformers).</li>\n<li><code>Replace augmentation</code>: Replace some random parts from other samples of the same class.</li>\n<li><code>Time and frequence masking</code>: This basic torchaudio augmentation works exceptionally well.</li>\n</ul>\n<pre><code>freq_m = torchaudio.transforms.FrequencyMasking()  \ntime_m = torchaudio.transforms.TimeMasking()       \n</code></pre>\n<h3>2.3 Augmented Sample Example</h3>\n<p>Before augmentation:</p>\n<p><img alt=\"aug1\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F5d2b97cf6754c5f4063724181bfe7172%2Fbefore_aug.png?generation=1682986332091937&amp;alt=media\"> </p>\n<p>After augmentation:</p>\n<p> <img alt=\"aug2\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F516f0790f6f9903ad8943977bc392965%2Fafter_aug.png?generation=1682986413818912&amp;alt=media\"> </p>\n<h2>3. Training</h2>\n<h3>3.1 CNN Training</h3>\n<ul>\n<li>Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters</li>\n<li>Onecycle scheduler with 0.1 warmup.</li>\n<li>Use weighted <code>CrossEntropyLoss</code>. Increase the weights for poorly predicted classes and classes with semantically similar pairs (such as kitty and cat)</li>\n<li>Implement a hypercolumn for EfficientNet with 5 blocks</li>\n</ul>\n<h3>3.2 Transformer Training</h3>\n<ul>\n<li>Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters</li>\n<li>Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule.</li>\n<li>A 4-layer, 256 hidden-size, 512 intermediate-size transformer were trained.</li>\n<li>A 3-layer model was initialized with 4-layer model's first 3 layers. Knowledge distillation were used in 3-layer model training, in which the 4-layer model is the teacher.</li>\n</ul>\n<h3>3.3 Hyperparameter Tuning</h3>\n<p>Since we trained only one fold and used smaller models, we decided to tune most parameters with Optuna. </p>\n<p>Here is the parameters list of CNN training (transformer training has a similar param-list):</p>\n<ul>\n<li><p>All augmentations probabilities (0.1 - 0.5+)</p></li>\n<li><p>Learning rate (2e-3 - 3e-3)</p></li>\n<li><p>Drop out (0.1 - 0.25)</p></li>\n<li><p>Num of epochs (170-185)</p></li>\n<li><p>Loss weights powers (0.75 - 2)</p></li>\n<li><p>Optimizer (<code>Lookahead_RAdam</code>, <code>RAdam</code>)</p></li>\n<li><p>Label smoothing (0.5 - 0.7)</p></li>\n</ul>\n<h2>4. Submissions, Conversion and Ensemble</h2>\n<ol>\n<li><p>We rewrote all our models in Keras and transferred PyTorch weights to them, resulting in a speed boost of around 30%. For transformer model, pytorch-onnx-tf-tflite will generate too much useless tensor shape operations, a fully rewriting can reduce these manually. For CNN model, we rewrote DepthwiseConv2D with a hard-coded way, whose speed is 200%~300% of its original version of tflite DepthwiseConv2D.</p></li>\n<li><p>After that, we aggregated all these models in the <code>tf.Module</code> class. Converting directly from Keras resulted in lower speed (don't know why).</p></li>\n<li><p>We calculated ensemble weights for models trained on fold 0 using the local fold 0 score and applied these weights to the full dataset models.</p></li>\n</ol>\n<p>EfficientNet-B0 achieved a leaderboard score of approximately 0.8, and transformers improved the score to 0.81. The final ensemble included:</p>\n<ol>\n<li>Efficientnet-B0, fold 0</li>\n<li>BERT, full data train</li>\n<li>DeBERTa, full data train</li>\n</ol>\n<p>Interestingly, a key feature was using the ensemble without softmax, which consistently provided a boost of around 0.01.</p>\n<h2>5. PS. Need <strong>BETTER</strong> TFlite DepthwiseConv2D</h2>\n<p>Depthwise convolution models performed very well for these tasks, outperforming other CNN and ViT models (rexnet_100 was also good).<br>\nWe spent a lot of time dealing with the conversion of DepthwiseConv2D operation. Here are some strange results:</p>\n<p>Given a input image with 82x42x32 (HWC), there are two ways to do a 3x3 depthwise convolution in Keras. One is <code>Conv2D(32, 3, groups = 32)</code>, the other is <code>DepthwiseConv2D(3)</code>. However, after converting these two to tflite, the running time of the <code>Conv2D</code> is 5.05ms, and the running time of <code>DepthwiseConv2D</code> is 3.70ms. More strangely, a full convolution <code>Conv2D(32, 3, groups = 1)</code> with FLOPs = HWC^2 only takes 2.09ms, even faster than previous two with FLOPs = HWC.</p>\n<p>Then we rewrote the depthwise-conv like this:</p>\n<pre><code>     ():\n        out = x[:,:self.H_out:self.strides,:self.W_out:self.strides] * self.weight[,]\n         i  (self.kernel_size):\n             j  (self.kernel_size):\n                 i ==   j == :\n                    \n                out += x[:,i:self.H_out + i:self.strides,j:self.W_out + j:self.strides] * self.weight[i,j]\n         self.bias   :\n            out = out + self.bias\n         out\n</code></pre>\n<p>The running time of this is 1.24 ms.</p>\n<p>In summary, our version (1.24ms) &gt; full <code>Conv2D</code> with larger FLOPs (2.09ms) &gt; <code>DepthwiseConv2D</code> (3.70ms) &gt; <code>Conv2D(C, groups = C)</code> (5.05ms).</p>\n<p>However, our version introduced too much nodes in tflite graph, which is not stable in running time. If the tensorflow team has a better implementation of DepthwiseConv2D, we can even ensemble two CNN models, which is expected to reach 0.82 LB.</p>\n<p>By the way, EfficientNet with ONNX was ~5 times faster than TFLite.</p>\n<h3>Big thanks to my teammates <a href=\"https://www.kaggle.com/artemtprv\" target=\"_blank\">@artemtprv</a> and <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> and congrats with new tiers, Master and GrandMaster!</h3>\n<p><a href=\"https://github.com/ffs333/2nd_place_GISLR\" target=\"_blank\">github code</a></p>",
      "rawMarkdown": "## TLDR\nWe used an approach similar to audio spectrogram classification using the EfficientNet-B0 model, with numerous augmentations and transformer models such as BERT and DeBERTa as helper models. The final solution consists of one EfficientNet-B0 with an input size of 160x80, trained on a single fold from 8 randomly split folds, as well as DeBERTa and BERT trained on the full dataset. A single fold model using EfficientNet has a CV score of 0.898 and a leaderboard score of ~0.8.\n\nWe used only competition data.\n\n\n\n## 1. Data Preprocessing\n\n### 1.1 CNN Preprocessing\n\n* We extracted 18 lip points, 20 pose points (including arms, shoulders, eyebrows, and nose), and all hand points, resulting in a total of 80 points.\n* During training, we applied various augmentations.\n* We implemented standard normalization.\n* Instead of dropping NaN values, we filled them with zeros after normalization.\n* We interpolated the time axis to a size of 160 using 'nearest' interpolation: `yy = F.interpolate(yy[None, None, :], size=self.new_size, mode='nearest')`.\n* Finally, we obtained a tensor with dimensions 160x80x3, where 3 represents the `(X, Y, Z)` axes. ![Preprocessing](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Fc47290891a7ac6497a6a0c296973f071%2Fdata_prep.jpg?generation=1682986293067532&alt=media)\n\n### 1.2 Transformer Preprocessing\n\n* Only 61 points were kept, including 40 lip points and 21 hand points. For left and right hand, the one with less NaN was kept. If right hand was kept, mirror it to left hand.\n\n* Augmentations, normalization and NaN-filling were applied sequentially.\n\n* Sequences longer than 96 were interpolated to 96. Sequences shorter than 96 were unchanged.\n\n* Apart from raw positions, hand-crafted features were also used, including motion, distances, and cosine of angles.\n\n* Motion features consist of future motion and history motion, which can be denoted as:\n  \n$$\n  Motion_{future} = position_{t+1} - position_{t}\n$$\n$$\n  Motion_{history} = position_{t} - position_{t-1}\n$$\n\n* Full 210 pairwise distances among 21 hand points were included. \n* There are 5 vertices in a finger (e.g. thumb is `[0,1,2,3,4]`), and therefore, there are 3 angles: `<0,1,2>, <1,2,3>, <2,3,4>`. So 15 angles of 5 fingers were included.\n\n* Randomly selected 190 pairwise distances and randomly selected 8 angles among 40 lip points were included.\n\n## 2. Augmentation\n\n### 2.1 Common Augmentations\n\n> These augmentations are used in both CNN training and transformer training\n\n1. `Random affine`: Same as @hengck23 shared. In CNN, after global affine, shift-scale-rotate was also applied to each part separately (e.g. hand, lip, body-pose).\n\n2. `Random interpolation`: Slightly scale and shift the time dimension.\n\n3. `Flip pose`: Flip the x-coordinates of all points. In CNN, `x_new = x_max - x_old`. In transformer, `x_new = 2 * frame[:,0,0] - x_old`.\n\n4. `Finger tree rotate`: There are 4 root-children pairs in a finger with 5-vertices. E.g. in thumb (`[0,1,2,3,4]`), these 4 root-children pairs are: `0-[1,2,3,4]`,`1-[2,3,4]`,`2-[3,4]`,`3-[4]`. We randomly choose some of these pairs, and rotate the children points around root point with a small random angle.\n\n### 2.2 CNN Specific Augmentations \n\n* `Mixup`: Implement basic mixup augmentation (only works with CNNs, not transformers).\n* `Replace augmentation`: Replace some random parts from other samples of the same class.\n* `Time and frequence masking`: This basic torchaudio augmentation works exceptionally well.\n\n```python\nfreq_m = torchaudio.transforms.FrequencyMasking(80)  # it's time axis\ntime_m = torchaudio.transforms.TimeMasking(18)       # it's points axis\n```\n\n### 2.3 Augmented Sample Example\n\nBefore augmentation:\n\n<p align=\"center\"><img alt=\"aug1\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F5d2b97cf6754c5f4063724181bfe7172%2Fbefore_aug.png?generation=1682986332091937&alt=media\" width=\"600\" height=\"200\"/> </p>\n\nAfter augmentation:\n\n<p align=\"center\"> <img alt=\"aug2\" height=\"200\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F516f0790f6f9903ad8943977bc392965%2Fafter_aug.png?generation=1682986413818912&alt=media\" width=\"600\"/> </p>\n\n## 3. Training\n\n### 3.1 CNN Training\n\n* Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters\n* Onecycle scheduler with 0.1 warmup.\n* Use weighted `CrossEntropyLoss`. Increase the weights for poorly predicted classes and classes with semantically similar pairs (such as kitty and cat)\n* Implement a hypercolumn for EfficientNet with 5 blocks\n\n\n\n### 3.2 Transformer Training\n\n* Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters\n* Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule.\n* A 4-layer, 256 hidden-size, 512 intermediate-size transformer were trained.\n* A 3-layer model was initialized with 4-layer model's first 3 layers. Knowledge distillation were used in 3-layer model training, in which the 4-layer model is the teacher.\n\n### 3.3 Hyperparameter Tuning\n\nSince we trained only one fold and used smaller models, we decided to tune most parameters with Optuna. \n\nHere is the parameters list of CNN training (transformer training has a similar param-list):\n\n* All augmentations probabilities (0.1 - 0.5+)\n\n* Learning rate (2e-3 - 3e-3)\n\n* Drop out (0.1 - 0.25)\n\n* Num of epochs (170-185)\n\n* Loss weights powers (0.75 - 2)\n\n* Optimizer (`Lookahead_RAdam`, `RAdam`)\n\n* Label smoothing (0.5 - 0.7)\n\n## 4. Submissions, Conversion and Ensemble\n\n1. We rewrote all our models in Keras and transferred PyTorch weights to them, resulting in a speed boost of around 30%. For transformer model, pytorch-onnx-tf-tflite will generate too much useless tensor shape operations, a fully rewriting can reduce these manually. For CNN model, we rewrote DepthwiseConv2D with a hard-coded way, whose speed is 200%~300% of its original version of tflite DepthwiseConv2D.\n\n2. After that, we aggregated all these models in the `tf.Module` class. Converting directly from Keras resulted in lower speed (don't know why).\n\n3. We calculated ensemble weights for models trained on fold 0 using the local fold 0 score and applied these weights to the full dataset models.\n\nEfficientNet-B0 achieved a leaderboard score of approximately 0.8, and transformers improved the score to 0.81. The final ensemble included:\n1. Efficientnet-B0, fold 0\n2. BERT, full data train\n3. DeBERTa, full data train\n\nInterestingly, a key feature was using the ensemble without softmax, which consistently provided a boost of around 0.01.\n\n## 5. PS. Need **BETTER** TFlite DepthwiseConv2D\n\nDepthwise convolution models performed very well for these tasks, outperforming other CNN and ViT models (rexnet_100 was also good).\nWe spent a lot of time dealing with the conversion of DepthwiseConv2D operation. Here are some strange results:\n\nGiven a input image with 82x42x32 (HWC), there are two ways to do a 3x3 depthwise convolution in Keras. One is `Conv2D(32, 3, groups = 32)`, the other is `DepthwiseConv2D(3)`. However, after converting these two to tflite, the running time of the `Conv2D` is 5.05ms, and the running time of `DepthwiseConv2D` is 3.70ms. More strangely, a full convolution `Conv2D(32, 3, groups = 1)` with FLOPs = HWC^2 only takes 2.09ms, even faster than previous two with FLOPs = HWC.\n\nThen we rewrote the depthwise-conv like this:\n\n```python\n    def call(self, x):\n        out = x[:,0:self.H_out:self.strides,0:self.W_out:self.strides] * self.weight[0,0]\n        for i in range(self.kernel_size):\n            for j in range(self.kernel_size):\n                if i == 0 and j == 0:\n                    continue\n                out += x[:,i:self.H_out + i:self.strides,j:self.W_out + j:self.strides] * self.weight[i,j]\n        if self.bias is not None:\n            out = out + self.bias\n        return out\n```\n\nThe running time of this is 1.24 ms.\n\nIn summary, our version (1.24ms) > full `Conv2D` with larger FLOPs (2.09ms) > `DepthwiseConv2D` (3.70ms) > `Conv2D(C, groups = C)` (5.05ms).\n\nHowever, our version introduced too much nodes in tflite graph, which is not stable in running time. If the tensorflow team has a better implementation of DepthwiseConv2D, we can even ensemble two CNN models, which is expected to reach 0.82 LB.\n\nBy the way, EfficientNet with ONNX was ~5 times faster than TFLite.\n\n\n### Big thanks to my teammates @artemtprv and @carnozhao and congrats with new tiers, Master and GrandMaster!\n\n\n[github code](https://github.com/ffs333/2nd_place_GISLR)",
      "votes": 118
    },
    {
      "id": 2242254,
      "postDate": "2023-05-02T06:31:09.200Z",
      "content": "<p>Congrats for the 2nd place! <br>\nI haven't even imagine treating this data as image… (I like audio spectrogram analogy by the way!)<br>\nWell deserved!</p>",
      "rawMarkdown": "Congrats for the 2nd place! \nI haven't even imagine treating this data as image... (I like audio spectrogram analogy by the way!)\nWell deserved!",
      "votes": 3,
      "replies": [
        {
          "id": 2242855,
          "postDate": "2023-05-02T14:41:52.940Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 2242329,
      "postDate": "2023-05-02T07:31:49.030Z",
      "content": "<p>Very nice work for both branches (CNN, Transformer) of the solution. How did you identify Depthwise convolution particularly as the bottleneck ? I tried efficientnet but it was too slow for me… however, I didnt rewrite the base model in to tf, only the prepreprocessing. I see that transferring the base model to keras gave a 30% speed up. </p>\n<p>Also, was the ordering of the xyz points important in the CNN. I guess you would want adjacent points beside each other in the CNN. But some points, like wrist, have multiple adjacent points. </p>",
      "rawMarkdown": "Very nice work for both branches (CNN, Transformer) of the solution. How did you identify Depthwise convolution particularly as the bottleneck ? I tried efficientnet but it was too slow for me... however, I didnt rewrite the base model in to tf, only the prepreprocessing. I see that transferring the base model to keras gave a 30% speed up. \n\nAlso, was the ordering of the xyz points important in the CNN. I guess you would want adjacent points beside each other in the CNN. But some points, like wrist, have multiple adjacent points. ",
      "votes": 4,
      "replies": [
        {
          "id": 2242520,
          "postDate": "2023-05-02T10:36:24.830Z",
          "content": "<p>About identification I hope Carno will answer to you later (he found that)</p>\n<p>About the ordering:<br>\nWe arranged them in groups (lips, left hand, pose, right hand), and within groups, in the order in which they are arranged in the raw data. Didn't tried any other orderings. But tried different amount of points and this combination of 80 points was the best</p>",
          "rawMarkdown": "About identification I hope Carno will answer to you later (he found that)\n\nAbout the ordering:\nWe arranged them in groups (lips, left hand, pose, right hand), and within groups, in the order in which they are arranged in the raw data. Didn't tried any other orderings. But tried different amount of points and this combination of 80 points was the best",
          "votes": 2
        },
        {
          "id": 2242611,
          "postDate": "2023-05-02T11:51:44.417Z",
          "content": "<p>I use Tensorflow's benchmark tool to analyze Node-wise running time. But DepthwiseConv is slower than what I expected.</p>",
          "rawMarkdown": "I use Tensorflow's benchmark tool to analyze Node-wise running time. But DepthwiseConv is slower than what I expected.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2241922,
      "postDate": "2023-05-02T00:44:29.260Z",
      "content": "<p>By the way, for us in public LB score for full data train CNN versions was worse and we haven't select it. But in private they were better. But still lower than <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> . That's a pure victory on last hours, GZ!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F10ba2123e7c100e364688e095a50b569%2Ftg_image_455350557.jpeg?generation=1682988215699105&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "By the way, for us in public LB score for full data train CNN versions was worse and we haven't select it. But in private they were better. But still lower than @hoyso48 . That's a pure victory on last hours, GZ!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F10ba2123e7c100e364688e095a50b569%2Ftg_image_455350557.jpeg?generation=1682988215699105&alt=media)",
      "votes": 4
    },
    {
      "id": 2242681,
      "postDate": "2023-05-02T12:50:19.213Z",
      "content": "<p>Congrats on the great result. Thanks for the detailed write-up. Found a lot of useful tricks that are gonna help in the future 🔥</p>",
      "rawMarkdown": "Congrats on the great result. Thanks for the detailed write-up. Found a lot of useful tricks that are gonna help in the future 🔥",
      "votes": 1
    },
    {
      "id": 2242511,
      "postDate": "2023-05-02T10:30:59.163Z",
      "content": "<p>Congrats great approach and good idea to use it as image 🎉🎉🎉🔥</p>",
      "rawMarkdown": "Congrats great approach and good idea to use it as image 🎉🎉🎉🔥",
      "votes": 1
    },
    {
      "id": 2241912,
      "postDate": "2023-05-02T00:34:15.397Z",
      "content": "<p>Congratz with the result! <br>\nAbout DepthwiseConv2D, - I had the same problem when depthwise conv1d was converted not efficiently when using pytorch-onnx-tf pipeline, I used it for positional encoding in transformer.</p>\n<p>You can try to check output conversion nodes via tensorflow benchmark that hengck shared, for me it appeared that whole transformer takes 200 ops and single depthwise conv1d takes 1800 ops (???) when using pytorch-&gt;onnx-&gt;tf conversion pipeline, and it was indeed very slow.</p>\n<p>I fixed it when I used <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">nobuco</a>, - depthwise convolution converted succesfully using built-in tf function.</p>",
      "rawMarkdown": "Congratz with the result! \nAbout DepthwiseConv2D, - I had the same problem when depthwise conv1d was converted not efficiently when using pytorch-onnx-tf pipeline, I used it for positional encoding in transformer.\n\nYou can try to check output conversion nodes via tensorflow benchmark that hengck shared, for me it appeared that whole transformer takes 200 ops and single depthwise conv1d takes 1800 ops (???) when using pytorch->onnx->tf conversion pipeline, and it was indeed very slow.\n\nI fixed it when I used [nobuco](https://github.com/AlexanderLutsenko/nobuco), - depthwise convolution converted succesfully using built-in tf function.",
      "votes": 1,
      "replies": [
        {
          "id": 2241917,
          "postDate": "2023-05-02T00:37:31.773Z",
          "content": "<p>Thanks. Congrats to you too!</p>\n<p>Yeah, we tried 'nobuco' and some other converters with ONNX. But they are all slower than our keras version. For example with nobuco we was able to add only one transformer model, with our current two transformers</p>",
          "rawMarkdown": "Thanks. Congrats to you too!\n\nYeah, we tried 'nobuco' and some other converters with ONNX. But they are all slower than our keras version. For example with nobuco we was able to add only one transformer model, with our current two transformers",
          "votes": 2
        },
        {
          "id": 2241925,
          "postDate": "2023-05-02T00:48:17.833Z",
          "content": "<p>I think what causes DepthwiseConv2D slow is the low-level code of tflite. We can not solve this unless we can use custom ops. In our solution, we use many other ops like Mul, Add, StrideSlice to replace DWConv. </p>",
          "rawMarkdown": "I think what causes DepthwiseConv2D slow is the low-level code of tflite. We can not solve this unless we can use custom ops. In our solution, we use many other ops like Mul, Add, StrideSlice to replace DWConv. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2243391,
      "postDate": "2023-05-02T21:49:33.037Z",
      "content": "<p>Congratz on the result and on the overall impressive performance !<br>\n0.81 public was impressive, we guessed you were using CNNs but could not get them to match our transformer performances.</p>",
      "rawMarkdown": "Congratz on the result and on the overall impressive performance !\n0.81 public was impressive, we guessed you were using CNNs but could not get them to match our transformer performances.",
      "votes": 2,
      "replies": [
        {
          "id": 2245326,
          "postDate": "2023-05-04T09:30:03.727Z",
          "content": "<p>Thanks. Congratulations to you too!</p>\n<blockquote>\n  <p>we guessed you were using CNNs but could not get them to match our transformer performances.</p>\n</blockquote>\n<p>I guess it because we have started CNN at early stage, when transformers didn't get 0.78+ yet, and with simply implementation it has 0.74 LB score (top1 at that moment)</p>",
          "rawMarkdown": "Thanks. Congratulations to you too!\n\n>we guessed you were using CNNs but could not get them to match our transformer performances.\n\nI guess it because we have started CNN at early stage, when transformers didn't get 0.78+ yet, and with simply implementation it has 0.74 LB score (top1 at that moment)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2242007,
      "postDate": "2023-05-02T01:59:03.923Z",
      "content": "<p>Big congratulations, and thanks for write-ups. It's amazing to read your struggle to speed up the inference 👍</p>",
      "rawMarkdown": "Big congratulations, and thanks for write-ups. It's amazing to read your struggle to speed up the inference 👍",
      "votes": 2,
      "replies": [
        {
          "id": 2242012,
          "postDate": "2023-05-02T02:03:15.260Z",
          "content": "<p>Thanks to you 🙃</p>",
          "rawMarkdown": "Thanks to you 🙃",
          "votes": 2
        }
      ]
    },
    {
      "id": 2433850,
      "postDate": "2023-09-11T22:29:49.607Z",
      "content": "<p>Thank you so much for posting your solution. Great job!!</p>",
      "rawMarkdown": "Thank you so much for posting your solution. Great job!!"
    },
    {
      "id": 2392704,
      "postDate": "2023-08-15T20:08:21.507Z",
      "content": "<p>Hi,<br>\nCongrats on second place!</p>\n<p>We are trying to run the code on the colab file training_notebook_example, and I saw that there is a possibility to run TREF model, which for what I understand run both CNN and transformers model. the problem I faced was choosing this model because I got an error when the name was sent to get_model() and I noticed that there is no option there. Can you help me with that.</p>\n<p>Thank you and congrats again!</p>",
      "rawMarkdown": "Hi,\nCongrats on second place!\n\nWe are trying to run the code on the colab file training_notebook_example, and I saw that there is a possibility to run TREF model, which for what I understand run both CNN and transformers model. the problem I faced was choosing this model because I got an error when the name was sent to get_model() and I noticed that there is no option there. Can you help me with that.\n\nThank you and congrats again!"
    },
    {
      "id": 2381121,
      "postDate": "2023-08-09T02:25:50.227Z",
      "content": "<p>Could someone explain this part for the normalization with how you found those values to be replaced by  (lign 237 to 243 of data.py) :   <br>\nself.interesting_idx = np.array(LIPS + l_hand + pose + r_hand)<br>\n # norm params<br>\n        self.x_mean = torch.tensor(0.48282188177108765).float()<br>\n        self.y_mean = torch.tensor(0.6022918224334717).float()<br>\n        self.z_mean = torch.tensor(-0.4806228578090668).float()<br>\n        self.x_std = torch.tensor(0.22193588316440582).float()<br>\n        self.y_std = torch.tensor(0.252733051776886).float()<br>\n        self.z_std = torch.tensor(0.68902188539505).float()</p>",
      "rawMarkdown": "Could someone explain this part for the normalization with how you found those values to be replaced by  (lign 237 to 243 of data.py) :   \nself.interesting_idx = np.array(LIPS + l_hand + pose + r_hand)\n # norm params\n        self.x_mean = torch.tensor(0.48282188177108765).float()\n        self.y_mean = torch.tensor(0.6022918224334717).float()\n        self.z_mean = torch.tensor(-0.4806228578090668).float()\n        self.x_std = torch.tensor(0.22193588316440582).float()\n        self.y_std = torch.tensor(0.252733051776886).float()\n        self.z_std = torch.tensor(0.68902188539505).float()"
    },
    {
      "id": 2355782,
      "postDate": "2023-07-23T15:41:40.637Z",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> great work. I could not find the transfer_transformer function used in inference code. Could you help me find it? </p>",
      "rawMarkdown": "@kolyaforrat great work. I could not find the transfer_transformer function used in inference code. Could you help me find it? "
    },
    {
      "id": 2246172,
      "postDate": "2023-05-05T00:08:55.600Z",
      "content": "<p>Great work, 2nd place is impressive. Hope to learn from your approach and use it to solve similar problems. </p>",
      "rawMarkdown": "Great work, 2nd place is impressive. Hope to learn from your approach and use it to solve similar problems. "
    },
    {
      "id": 2245705,
      "postDate": "2023-05-04T14:34:39.260Z",
      "content": "<p>Great work! I'm very inspired by your spectrogram classification approach. Very clever! I've been trying to recreate it locally from scratch using your code on github a a guide. Any chance you could share a screenshot of the training curves for the train/val loss and accuracy and learning rate?</p>",
      "rawMarkdown": "Great work! I'm very inspired by your spectrogram classification approach. Very clever! I've been trying to recreate it locally from scratch using your code on github a a guide. Any chance you could share a screenshot of the training curves for the train/val loss and accuracy and learning rate?",
      "replies": [
        {
          "id": 2245792,
          "postDate": "2023-05-04T15:25:01.003Z",
          "content": "<p>Thanks!<br>\nSure, that's my last attempts to train, here is 16 random fold split train (accuracy ~0.005 higher than 8 folds split)</p>\n<p>LR:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F0d2b6a6b595b8c5e323fb27a16bc74c8%2Ftg_image_762402010.jpeg?generation=1683213719007236&amp;alt=media\" alt=\"\"></p>\n<p>Train/Val Loss<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F16063ab6e8101a1f6f88c08be4f7548f%2Ftg_image_1545167309.jpeg?generation=1683213729689919&amp;alt=media\" alt=\"\"></p>\n<p>Accuracy:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Feee8989b2b68e37bb0279e44984fa101%2Ftg_image_2756323583.jpeg?generation=1683213753844440&amp;alt=media\" alt=\"\"></p>\n<p>TopK3:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F622f783745ea25768373fdd9d23a1f36%2Ftg_image_3839074706.jpeg?generation=1683213791599040&amp;alt=media\" alt=\"\"></p>\n<p>Epochs-accuracy-LR:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F0cd11d2e85c7097bff3c898d1e598aaa%2Ftg_image_3735457178.jpeg?generation=1683213779294885&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thanks!\nSure, that's my last attempts to train, here is 16 random fold split train (accuracy ~0.005 higher than 8 folds split)\n\nLR:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F0d2b6a6b595b8c5e323fb27a16bc74c8%2Ftg_image_762402010.jpeg?generation=1683213719007236&alt=media)\n\nTrain/Val Loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F16063ab6e8101a1f6f88c08be4f7548f%2Ftg_image_1545167309.jpeg?generation=1683213729689919&alt=media)\n\nAccuracy:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Feee8989b2b68e37bb0279e44984fa101%2Ftg_image_2756323583.jpeg?generation=1683213753844440&alt=media)\n\nTopK3:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F622f783745ea25768373fdd9d23a1f36%2Ftg_image_3839074706.jpeg?generation=1683213791599040&alt=media)\n\nEpochs-accuracy-LR:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F0cd11d2e85c7097bff3c898d1e598aaa%2Ftg_image_3735457178.jpeg?generation=1683213779294885&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 2246053,
              "postDate": "2023-05-04T19:57:43.583Z",
              "content": "<p>This is great! Thanks for sharing. I'm noticing your loss function uses <code>bad_weights.npy</code> and <code>common_weights.npy</code> which you'ver provided in your repo- do you have any more detail about how those were created and how you decided to implement the loss function like this? Thanks!</p>\n<pre><code>full_wgt = torch.tensor(bad_wgt ** cfg.pw_bad + common_wgt ** cfg.pw_com, dtype=torch.float32)\n</code></pre>\n<p>here: <a href=\"https://github.com/ffs333/2nd_place_GISLR/blob/main/GISLR_utils/losses.py#L13\" target=\"_blank\">https://github.com/ffs333/2nd_place_GISLR/blob/main/GISLR_utils/losses.py#L13</a></p>\n<p>It looks like your best config used:</p>\n<pre><code>use_loss_wgt:True\npw_bad:0.8219710950094926\npw_com:1.3979208248868304\n</code></pre>",
              "rawMarkdown": "This is great! Thanks for sharing. I'm noticing your loss function uses `bad_weights.npy` and `common_weights.npy` which you'ver provided in your repo- do you have any more detail about how those were created and how you decided to implement the loss function like this? Thanks!\n```\nfull_wgt = torch.tensor(bad_wgt ** cfg.pw_bad + common_wgt ** cfg.pw_com, dtype=torch.float32)\n```\nhere: https://github.com/ffs333/2nd_place_GISLR/blob/main/GISLR_utils/losses.py#L13\n\nIt looks like your best config used:\n```\nuse_loss_wgt:True\npw_bad:0.8219710950094926\npw_com:1.3979208248868304\n```"
            },
            {
              "id": 2246061,
              "postDate": "2023-05-04T20:14:44.060Z",
              "content": "<p>First we trained models with basic CE loss. And then we analyzed the errors of our models as provided <a href=\"https://www.kaggle.com/kolyaforrat/error-analyse\" target=\"_blank\">here</a></p>\n<ul>\n<li>Bad weights - increase weights of bad predicted classes</li>\n<li>Common weights - increase weights of classes which has similar pairs (like kitty and cat)</li>\n</ul>\n<p>Then just convert it from json to numpy arrays<br>\nPowers applied to these weights were just hyper parameter we tuned in optuna. At the beginning we started with wide range from 0 to 3</p>\n<p>Also we had problems with samples with low lengths (&lt;8). But we didn't found the solution to predict them better. Tried to train or finetune two models:</p>\n<ol>\n<li>On train data with increased amount of low length data (&lt;=10)</li>\n<li>On only long data (&gt;10)</li>\n</ol>\n<p>And predict it with if else statements in submission. But it has ~the same score and even little worse</p>",
              "rawMarkdown": "First we trained models with basic CE loss. And then we analyzed the errors of our models as provided [here](https://www.kaggle.com/kolyaforrat/error-analyse)\n\n* Bad weights - increase weights of bad predicted classes\n* Common weights - increase weights of classes which has similar pairs (like kitty and cat)\n\nThen just convert it from json to numpy arrays\nPowers applied to these weights were just hyper parameter we tuned in optuna. At the beginning we started with wide range from 0 to 3\n\nAlso we had problems with samples with low lengths (<8). But we didn't found the solution to predict them better. Tried to train or finetune two models:\n1. On train data with increased amount of low length data (<=10)\n2. On only long data (>10)\n\nAnd predict it with if else statements in submission. But it has ~the same score and even little worse",
              "votes": 1
            },
            {
              "id": 2246073,
              "postDate": "2023-05-04T20:33:58.303Z",
              "content": "<p>Thanks again for answering. Very interesting. I did something similar with looking at common words - but I did the opposite and added additional label smoothing for them. Maybe I should have been doing like you did an increased their weights. </p>\n<p>I have another quick question, because your <code>pw_bad</code> value is &lt; 1 then you are effectively reducing the weights of the \"bad\" words. Am I understanding that correctly?</p>",
              "rawMarkdown": "Thanks again for answering. Very interesting. I did something similar with looking at common words - but I did the opposite and added additional label smoothing for them. Maybe I should have been doing like you did an increased their weights. \n\nI have another quick question, because your `pw_bad` value is < 1 then you are effectively reducing the weights of the \"bad\" words. Am I understanding that correctly?"
            },
            {
              "id": 2246078,
              "postDate": "2023-05-04T20:49:42.530Z",
              "content": "<p>Yeah, they become lower than calculated by errors, but higher that just 1 :)</p>\n<p>For example weight for class cat = 2, and it becomes 2**0.75~=1.68 in current run if pw_bad = 0.75</p>\n<p>As I said we tuned range of loss weights powers from 0, because we didn't know will they help or not</p>",
              "rawMarkdown": "Yeah, they become lower than calculated by errors, but higher that just 1 :)\n\nFor example weight for class cat = 2, and it becomes 2**0.75~=1.68 in current run if pw_bad = 0.75\n\nAs I said we tuned range of loss weights powers from 0, because we didn't know will they help or not",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2244071,
      "postDate": "2023-05-03T12:14:36.877Z",
      "content": "<p>So cool to read about your approach, thank you for sharing! Interestingly, as probably many of us, I could not get CNN to work on this problem, now I see why. \"Devil is in the details\". Congrats on the 2nd place, well deserved! </p>",
      "rawMarkdown": "So cool to read about your approach, thank you for sharing! Interestingly, as probably many of us, I could not get CNN to work on this problem, now I see why. \"Devil is in the details\". Congrats on the 2nd place, well deserved! "
    },
    {
      "id": 2243854,
      "postDate": "2023-05-03T08:23:05.040Z",
      "content": "<p>Congrats ! You did a great work</p>",
      "rawMarkdown": "Congrats ! You did a great work"
    },
    {
      "id": 2243335,
      "postDate": "2023-05-02T20:32:58.400Z",
      "content": "<p>Now just tried to submit single fold solo CNN model and got this score:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F867040d30c2f4535cd33fd19989de3d3%2Ftg_image_380692410.jpeg?generation=1683059371986310&amp;alt=media\" alt=\"\"></p>\n<p>It's 5th place in public and 4th in private with 17.5 MB and 42 minutes</p>",
      "rawMarkdown": "Now just tried to submit single fold solo CNN model and got this score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F867040d30c2f4535cd33fd19989de3d3%2Ftg_image_380692410.jpeg?generation=1683059371986310&alt=media)\n\nIt's 5th place in public and 4th in private with 17.5 MB and 42 minutes"
    },
    {
      "id": 2243334,
      "postDate": "2023-05-02T20:32:33.493Z",
      "content": "<p>wow great !this will be very helpful this is what in summary i understood-<br>\nThey used an approach similar to audio spectrogram classification for their solution.<br>\nThey used EfficientNet-B0 as the main model and applied various augmentations and transformer models such as BERT and DeBERTa as helper models.<br>\nThe team extracted 80 points for CNN preprocessing and 61 points for Transformer preprocessing.<br>\nThey used various augmentations such as random affine, random interpolation, flip pose, finger tree rotate, mixup, replace augmentation, time and frequency masking.<br>\nThey trained on one fold with a random split (8 folds in total) or the full dataset using the best parameters.<br>\nThey used a onecycle scheduler with 0.1 warmup and a weighted CrossEntropyLoss for CNN training, and a Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule for Transformer training.<br>\nThey implemented hyperparameter tuning with Optuna for most parameters.<br>\nThey rewrote all their models in Keras and transferred PyTorch models to Keras models.<br>\nThey used an ensemble of different models to improve their performance.</p>",
      "rawMarkdown": "wow great !this will be very helpful this is what in summary i understood-\nThey used an approach similar to audio spectrogram classification for their solution.\nThey used EfficientNet-B0 as the main model and applied various augmentations and transformer models such as BERT and DeBERTa as helper models.\nThe team extracted 80 points for CNN preprocessing and 61 points for Transformer preprocessing.\nThey used various augmentations such as random affine, random interpolation, flip pose, finger tree rotate, mixup, replace augmentation, time and frequency masking.\nThey trained on one fold with a random split (8 folds in total) or the full dataset using the best parameters.\nThey used a onecycle scheduler with 0.1 warmup and a weighted CrossEntropyLoss for CNN training, and a Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule for Transformer training.\nThey implemented hyperparameter tuning with Optuna for most parameters.\nThey rewrote all their models in Keras and transferred PyTorch models to Keras models.\nThey used an ensemble of different models to improve their performance."
    },
    {
      "id": 2243059,
      "postDate": "2023-05-02T16:56:26.467Z",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a>  <a href=\"https://www.kaggle.com/artemtprv\" target=\"_blank\">@artemtprv</a> <a href=\"https://www.kaggle.com/carnozhaocongrats\" target=\"_blank\">@carnozhaocongrats</a> with the second highest place in so competitive contest!💥🎉 The 5-th paragraph of you post is especially interesting as well as Efficientnet+BERT+DeBERTa ensemble approach👍💪🙂</p>",
      "rawMarkdown": "@kolyaforrat  @artemtprv @carnozhaocongrats with the second highest place in so competitive contest!💥🎉 The 5-th paragraph of you post is especially interesting as well as Efficientnet+BERT+DeBERTa ensemble approach👍💪🙂"
    },
    {
      "id": 2243019,
      "postDate": "2023-05-02T16:32:14.707Z",
      "content": "<p>Congrats on the perfect result! </p>\n<p>May I ask you to clarify a way you use cross-validation? You mentioned <code>Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters</code>. Do you mean you run all 8 folds to validate your model? Or the single fold was used most of the time? Am I also wondering, how do you validate an ensemble of models? I trained all the folds for each model, collect OOF predictions and calculate the ensemble using CV over OOF. It was a time-consuming process and I am not sure whether this is the best way to follow. <br>\nWould love to learn from your experience</p>",
      "rawMarkdown": "Congrats on the perfect result! \n\nMay I ask you to clarify a way you use cross-validation? You mentioned `Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters`. Do you mean you run all 8 folds to validate your model? Or the single fold was used most of the time? Am I also wondering, how do you validate an ensemble of models? I trained all the folds for each model, collect OOF predictions and calculate the ensemble using CV over OOF. It was a time-consuming process and I am not sure whether this is the best way to follow. \nWould love to learn from your experience",
      "replies": [
        {
          "id": 2243053,
          "postDate": "2023-05-02T16:54:04.137Z",
          "content": "<p>Thanks!</p>\n<p>We tested few times for 5 folds, and with random split OOF scores was almost the same for all folds. That's why after this we trained only fold 0 for all splits. We tried 5 folds, 8 folds, 16 folds. 8 folds version was the best</p>\n<p>After finding of some best checkpoint for fold 0 we trained full data checkpoint with the same params.</p>\n<p>For getting best ensemble weights we trained all models (eff-b0, bert, deberta) with fold 0 and calculate best weights for ensemble. It was 4 numbers: <br>\n<code>cnn_wgt * CNN + trans_wgt * (bert * bert_wgt + deberta * deberta_wgt)</code><br>\n Also we have tried manually tune weights on LB score but always that fold0 weights was the best.</p>\n<p>You can check how we did it in last 2 commented cells in <a href=\"https://github.com/ffs333/2nd_place_GISLR/blob/main/kaggle_inference_notebook_example.ipynb\" target=\"_blank\">our inference notebook</a></p>\n<p>Also sorry about your shakedown, I know that feeling 🙃</p>",
          "rawMarkdown": "Thanks!\n\nWe tested few times for 5 folds, and with random split OOF scores was almost the same for all folds. That's why after this we trained only fold 0 for all splits. We tried 5 folds, 8 folds, 16 folds. 8 folds version was the best\n\nAfter finding of some best checkpoint for fold 0 we trained full data checkpoint with the same params.\n\nFor getting best ensemble weights we trained all models (eff-b0, bert, deberta) with fold 0 and calculate best weights for ensemble. It was 4 numbers: \n```cnn_wgt * CNN + trans_wgt * (bert * bert_wgt + deberta * deberta_wgt)```\n Also we have tried manually tune weights on LB score but always that fold0 weights was the best.\n\nYou can check how we did it in last 2 commented cells in [our inference notebook](https://github.com/ffs333/2nd_place_GISLR/blob/main/kaggle_inference_notebook_example.ipynb)\n\nAlso sorry about your shakedown, I know that feeling 🙃",
          "votes": 1,
          "replies": [
            {
              "id": 2243169,
              "postDate": "2023-05-02T18:02:02.593Z",
              "content": "<p>Thank you so much for your answer!</p>\n<blockquote>\n  <p>Also sorry about your shakedown, I know that feeling 🙃</p>\n</blockquote>\n<p>Ah, that is painful, taking into account that it was my first real competition I participate from start to end almost evening. Nonetheless, I am looking forward to meeting you in one of the next competitions! </p>",
              "rawMarkdown": "Thank you so much for your answer!\n\n> Also sorry about your shakedown, I know that feeling 🙃\n\nAh, that is painful, taking into account that it was my first real competition I participate from start to end almost evening. Nonetheless, I am looking forward to meeting you in one of the next competitions! "
            }
          ]
        }
      ]
    },
    {
      "id": 2242218,
      "postDate": "2023-05-02T06:01:22.063Z",
      "content": "<p>Could you further explain the meaning of your xyz heatmap? I do not understand why you put it in the repo…</p>",
      "rawMarkdown": "Could you further explain the meaning of your xyz heatmap? I do not understand why you put it in the repo...",
      "replies": [
        {
          "id": 2242237,
          "postDate": "2023-05-02T06:15:41.597Z",
          "content": "<p>data is:</p>\n<pre><code>sample = train.loc[i] \nyy = load_relevant_data_subset(base_dir + sample[])  \n</code></pre>\n<p>In training dataset class we just combine all these <code>yy</code> to list and iterate through it. We keep it RAM for faster training.<br>\nAnd load it like this <code>data = np.load(_cfg.base_path + 'gen_xyz/data.npy', allow_pickle=True)</code><br>\nYou can see it in repo in <code>GISLR_utils/mixup_data.py</code> <code>DatasetImageSmall80Mixup</code></p>",
          "rawMarkdown": "data is:\n```python\nsample = train.loc[i] \nyy = load_relevant_data_subset(base_dir + sample['path'])  # function provided by hosts\n```\n\nIn training dataset class we just combine all these `yy` to list and iterate through it. We keep it RAM for faster training.\nAnd load it like this `data = np.load(_cfg.base_path + 'gen_xyz/data.npy', allow_pickle=True)`\nYou can see it in repo in `GISLR_utils/mixup_data.py ` `DatasetImageSmall80Mixup`",
          "replies": [
            {
              "id": 2290887,
              "postDate": "2023-06-07T07:16:53.420Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2290951,
              "postDate": "2023-06-07T08:13:20.387Z",
              "content": "<p>Answer <a href=\"https://github.com/ffs333/2nd_place_GISLR/issues/1#issuecomment-1555364919\" target=\"_blank\">here</a></p>",
              "rawMarkdown": "Answer [here](https://github.com/ffs333/2nd_place_GISLR/issues/1#issuecomment-1555364919)"
            }
          ]
        }
      ]
    },
    {
      "id": 2242200,
      "postDate": "2023-05-02T05:40:05.650Z",
      "content": "<p>Thanks for sharing, and congratulations!</p>",
      "rawMarkdown": "Thanks for sharing, and congratulations!"
    },
    {
      "id": 2242198,
      "postDate": "2023-05-02T05:32:59.060Z",
      "content": "<p>Amazing!Thanks for sharing the detail approaches here.</p>",
      "rawMarkdown": "Amazing!Thanks for sharing the detail approaches here."
    },
    {
      "id": 2242141,
      "postDate": "2023-05-02T04:06:32.953Z",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> cool solution! Congrats with 2nd place!</p>",
      "rawMarkdown": "@kolyaforrat cool solution! Congrats with 2nd place!"
    },
    {
      "id": 2242098,
      "postDate": "2023-05-02T03:25:04.227Z",
      "content": "<p>Thank you for the explanation! Is the source code going to be published publicly? That would be great.</p>",
      "rawMarkdown": "Thank you for the explanation! Is the source code going to be published publicly? That would be great.",
      "replies": [
        {
          "id": 2242150,
          "postDate": "2023-05-02T04:20:45.850Z",
          "content": "<p>Sure, it posted <a href=\"https://github.com/ffs333/2nd_place_GISLR\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Sure, it posted [here](https://github.com/ffs333/2nd_place_GISLR)",
          "votes": 2
        }
      ]
    },
    {
      "id": 2241955,
      "postDate": "2023-05-02T01:13:55.517Z",
      "content": "<p>Hearty congratulations to you for the fantastic result and associated gold medal and best regards <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> </p>",
      "rawMarkdown": "Hearty congratulations to you for the fantastic result and associated gold medal and best regards @kolyaforrat "
    },
    {
      "id": 2241929,
      "postDate": "2023-05-02T00:50:25.557Z",
      "content": "<p>Very impressive.  Thanks for sharing the detail approaches here.   </p>",
      "rawMarkdown": "Very impressive.  Thanks for sharing the detail approaches here.   "
    },
    {
      "id": 2970705,
      "postDate": "2024-08-26T12:17:03.667Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2242422,
      "postDate": "2023-05-02T09:07:43Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2246448,
      "postDate": "2023-05-05T06:47:22.663Z",
      "content": "<p>Amazing work!<br>\nThanks for sharing.</p>",
      "rawMarkdown": "Amazing work!\nThanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 2242254,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2023-05-02T06:31:09.200000",
      "content": "<p>Congrats for the 2nd place! <br>\nI haven't even imagine treating this data as image… (I like audio spectrogram analogy by the way!)<br>\nWell deserved!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2242855,
          "author_name": "Artem Toporov",
          "author_url": "",
          "post_date": "2023-05-02T14:41:52.940000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2242329,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2023-05-02T07:31:49.030000",
      "content": "<p>Very nice work for both branches (CNN, Transformer) of the solution. How did you identify Depthwise convolution particularly as the bottleneck ? I tried efficientnet but it was too slow for me… however, I didnt rewrite the base model in to tf, only the prepreprocessing. I see that transferring the base model to keras gave a 30% speed up. </p>\n<p>Also, was the ordering of the xyz points important in the CNN. I guess you would want adjacent points beside each other in the CNN. But some points, like wrist, have multiple adjacent points. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2242520,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-02T10:36:24.830000",
          "content": "<p>About identification I hope Carno will answer to you later (he found that)</p>\n<p>About the ordering:<br>\nWe arranged them in groups (lips, left hand, pose, right hand), and within groups, in the order in which they are arranged in the raw data. Didn't tried any other orderings. But tried different amount of points and this combination of 80 points was the best</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2242611,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-05-02T11:51:44.417000",
          "content": "<p>I use Tensorflow's benchmark tool to analyze Node-wise running time. But DepthwiseConv is slower than what I expected.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2241922,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-02T00:44:29.260000",
      "content": "<p>By the way, for us in public LB score for full data train CNN versions was worse and we haven't select it. But in private they were better. But still lower than <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> . That's a pure victory on last hours, GZ!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F10ba2123e7c100e364688e095a50b569%2Ftg_image_455350557.jpeg?generation=1682988215699105&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2242681,
      "author_name": "CroDoc",
      "author_url": "",
      "post_date": "2023-05-02T12:50:19.213000",
      "content": "<p>Congrats on the great result. Thanks for the detailed write-up. Found a lot of useful tricks that are gonna help in the future 🔥</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2242511,
      "author_name": "mhdaw",
      "author_url": "",
      "post_date": "2023-05-02T10:30:59.163000",
      "content": "<p>Congrats great approach and good idea to use it as image 🎉🎉🎉🔥</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2241912,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2023-05-02T00:34:15.397000",
      "content": "<p>Congratz with the result! <br>\nAbout DepthwiseConv2D, - I had the same problem when depthwise conv1d was converted not efficiently when using pytorch-onnx-tf pipeline, I used it for positional encoding in transformer.</p>\n<p>You can try to check output conversion nodes via tensorflow benchmark that hengck shared, for me it appeared that whole transformer takes 200 ops and single depthwise conv1d takes 1800 ops (???) when using pytorch-&gt;onnx-&gt;tf conversion pipeline, and it was indeed very slow.</p>\n<p>I fixed it when I used <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">nobuco</a>, - depthwise convolution converted succesfully using built-in tf function.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2241917,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-02T00:37:31.773000",
          "content": "<p>Thanks. Congrats to you too!</p>\n<p>Yeah, we tried 'nobuco' and some other converters with ONNX. But they are all slower than our keras version. For example with nobuco we was able to add only one transformer model, with our current two transformers</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2241925,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-05-02T00:48:17.833000",
          "content": "<p>I think what causes DepthwiseConv2D slow is the low-level code of tflite. We can not solve this unless we can use custom ops. In our solution, we use many other ops like Mul, Add, StrideSlice to replace DWConv. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2243391,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2023-05-02T21:49:33.037000",
      "content": "<p>Congratz on the result and on the overall impressive performance !<br>\n0.81 public was impressive, we guessed you were using CNNs but could not get them to match our transformer performances.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2245326,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-04T09:30:03.727000",
          "content": "<p>Thanks. Congratulations to you too!</p>\n<blockquote>\n  <p>we guessed you were using CNNs but could not get them to match our transformer performances.</p>\n</blockquote>\n<p>I guess it because we have started CNN at early stage, when transformers didn't get 0.78+ yet, and with simply implementation it has 0.74 LB score (top1 at that moment)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2242007,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-05-02T01:59:03.923000",
      "content": "<p>Big congratulations, and thanks for write-ups. It's amazing to read your struggle to speed up the inference 👍</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2242012,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-02T02:03:15.260000",
          "content": "<p>Thanks to you 🙃</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2433850,
      "author_name": "Johnpaul Nwagwu",
      "author_url": "",
      "post_date": "2023-09-11T22:29:49.607000",
      "content": "<p>Thank you so much for posting your solution. Great job!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2392704,
      "author_name": "Adir Eliza",
      "author_url": "",
      "post_date": "2023-08-15T20:08:21.507000",
      "content": "<p>Hi,<br>\nCongrats on second place!</p>\n<p>We are trying to run the code on the colab file training_notebook_example, and I saw that there is a possibility to run TREF model, which for what I understand run both CNN and transformers model. the problem I faced was choosing this model because I got an error when the name was sent to get_model() and I noticed that there is no option there. Can you help me with that.</p>\n<p>Thank you and congrats again!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2381121,
      "author_name": "Quentin Parrenin",
      "author_url": "",
      "post_date": "2023-08-09T02:25:50.227000",
      "content": "<p>Could someone explain this part for the normalization with how you found those values to be replaced by  (lign 237 to 243 of data.py) :   <br>\nself.interesting_idx = np.array(LIPS + l_hand + pose + r_hand)<br>\n # norm params<br>\n        self.x_mean = torch.tensor(0.48282188177108765).float()<br>\n        self.y_mean = torch.tensor(0.6022918224334717).float()<br>\n        self.z_mean = torch.tensor(-0.4806228578090668).float()<br>\n        self.x_std = torch.tensor(0.22193588316440582).float()<br>\n        self.y_std = torch.tensor(0.252733051776886).float()<br>\n        self.z_std = torch.tensor(0.68902188539505).float()</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2355782,
      "author_name": "Heysem Ismail",
      "author_url": "",
      "post_date": "2023-07-23T15:41:40.637000",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> great work. I could not find the transfer_transformer function used in inference code. Could you help me find it? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246172,
      "author_name": "Hasan Patel Rodriguez",
      "author_url": "",
      "post_date": "2023-05-05T00:08:55.600000",
      "content": "<p>Great work, 2nd place is impressive. Hope to learn from your approach and use it to solve similar problems. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2245705,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2023-05-04T14:34:39.260000",
      "content": "<p>Great work! I'm very inspired by your spectrogram classification approach. Very clever! I've been trying to recreate it locally from scratch using your code on github a a guide. Any chance you could share a screenshot of the training curves for the train/val loss and accuracy and learning rate?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2245792,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-04T15:25:01.003000",
          "content": "<p>Thanks!<br>\nSure, that's my last attempts to train, here is 16 random fold split train (accuracy ~0.005 higher than 8 folds split)</p>\n<p>LR:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F0d2b6a6b595b8c5e323fb27a16bc74c8%2Ftg_image_762402010.jpeg?generation=1683213719007236&amp;alt=media\" alt=\"\"></p>\n<p>Train/Val Loss<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F16063ab6e8101a1f6f88c08be4f7548f%2Ftg_image_1545167309.jpeg?generation=1683213729689919&amp;alt=media\" alt=\"\"></p>\n<p>Accuracy:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Feee8989b2b68e37bb0279e44984fa101%2Ftg_image_2756323583.jpeg?generation=1683213753844440&amp;alt=media\" alt=\"\"></p>\n<p>TopK3:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F622f783745ea25768373fdd9d23a1f36%2Ftg_image_3839074706.jpeg?generation=1683213791599040&amp;alt=media\" alt=\"\"></p>\n<p>Epochs-accuracy-LR:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F0cd11d2e85c7097bff3c898d1e598aaa%2Ftg_image_3735457178.jpeg?generation=1683213779294885&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2246053,
              "author_name": "Rob Mulla",
              "author_url": "",
              "post_date": "2023-05-04T19:57:43.583000",
              "content": "<p>This is great! Thanks for sharing. I'm noticing your loss function uses <code>bad_weights.npy</code> and <code>common_weights.npy</code> which you'ver provided in your repo- do you have any more detail about how those were created and how you decided to implement the loss function like this? Thanks!</p>\n<pre><code>full_wgt = torch.tensor(bad_wgt ** cfg.pw_bad + common_wgt ** cfg.pw_com, dtype=torch.float32)\n</code></pre>\n<p>here: <a href=\"https://github.com/ffs333/2nd_place_GISLR/blob/main/GISLR_utils/losses.py#L13\" target=\"_blank\">https://github.com/ffs333/2nd_place_GISLR/blob/main/GISLR_utils/losses.py#L13</a></p>\n<p>It looks like your best config used:</p>\n<pre><code>use_loss_wgt:True\npw_bad:0.8219710950094926\npw_com:1.3979208248868304\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2246061,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-04T20:14:44.060000",
              "content": "<p>First we trained models with basic CE loss. And then we analyzed the errors of our models as provided <a href=\"https://www.kaggle.com/kolyaforrat/error-analyse\" target=\"_blank\">here</a></p>\n<ul>\n<li>Bad weights - increase weights of bad predicted classes</li>\n<li>Common weights - increase weights of classes which has similar pairs (like kitty and cat)</li>\n</ul>\n<p>Then just convert it from json to numpy arrays<br>\nPowers applied to these weights were just hyper parameter we tuned in optuna. At the beginning we started with wide range from 0 to 3</p>\n<p>Also we had problems with samples with low lengths (&lt;8). But we didn't found the solution to predict them better. Tried to train or finetune two models:</p>\n<ol>\n<li>On train data with increased amount of low length data (&lt;=10)</li>\n<li>On only long data (&gt;10)</li>\n</ol>\n<p>And predict it with if else statements in submission. But it has ~the same score and even little worse</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2246073,
              "author_name": "Rob Mulla",
              "author_url": "",
              "post_date": "2023-05-04T20:33:58.303000",
              "content": "<p>Thanks again for answering. Very interesting. I did something similar with looking at common words - but I did the opposite and added additional label smoothing for them. Maybe I should have been doing like you did an increased their weights. </p>\n<p>I have another quick question, because your <code>pw_bad</code> value is &lt; 1 then you are effectively reducing the weights of the \"bad\" words. Am I understanding that correctly?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2246078,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-05-04T20:49:42.530000",
              "content": "<p>Yeah, they become lower than calculated by errors, but higher that just 1 :)</p>\n<p>For example weight for class cat = 2, and it becomes 2**0.75~=1.68 in current run if pw_bad = 0.75</p>\n<p>As I said we tuned range of loss weights powers from 0, because we didn't know will they help or not</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2244071,
      "author_name": "Siarhei T",
      "author_url": "",
      "post_date": "2023-05-03T12:14:36.877000",
      "content": "<p>So cool to read about your approach, thank you for sharing! Interestingly, as probably many of us, I could not get CNN to work on this problem, now I see why. \"Devil is in the details\". Congrats on the 2nd place, well deserved! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2243854,
      "author_name": "Hemangi atale",
      "author_url": "",
      "post_date": "2023-05-03T08:23:05.040000",
      "content": "<p>Congrats ! You did a great work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2243335,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-02T20:32:58.400000",
      "content": "<p>Now just tried to submit single fold solo CNN model and got this score:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F867040d30c2f4535cd33fd19989de3d3%2Ftg_image_380692410.jpeg?generation=1683059371986310&amp;alt=media\" alt=\"\"></p>\n<p>It's 5th place in public and 4th in private with 17.5 MB and 42 minutes</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2243334,
      "author_name": "asjad2024",
      "author_url": "",
      "post_date": "2023-05-02T20:32:33.493000",
      "content": "<p>wow great !this will be very helpful this is what in summary i understood-<br>\nThey used an approach similar to audio spectrogram classification for their solution.<br>\nThey used EfficientNet-B0 as the main model and applied various augmentations and transformer models such as BERT and DeBERTa as helper models.<br>\nThe team extracted 80 points for CNN preprocessing and 61 points for Transformer preprocessing.<br>\nThey used various augmentations such as random affine, random interpolation, flip pose, finger tree rotate, mixup, replace augmentation, time and frequency masking.<br>\nThey trained on one fold with a random split (8 folds in total) or the full dataset using the best parameters.<br>\nThey used a onecycle scheduler with 0.1 warmup and a weighted CrossEntropyLoss for CNN training, and a Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule for Transformer training.<br>\nThey implemented hyperparameter tuning with Optuna for most parameters.<br>\nThey rewrote all their models in Keras and transferred PyTorch models to Keras models.<br>\nThey used an ensemble of different models to improve their performance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2243059,
      "author_name": "Ivan Isaev",
      "author_url": "",
      "post_date": "2023-05-02T16:56:26.467000",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a>  <a href=\"https://www.kaggle.com/artemtprv\" target=\"_blank\">@artemtprv</a> <a href=\"https://www.kaggle.com/carnozhaocongrats\" target=\"_blank\">@carnozhaocongrats</a> with the second highest place in so competitive contest!💥🎉 The 5-th paragraph of you post is especially interesting as well as Efficientnet+BERT+DeBERTa ensemble approach👍💪🙂</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2243019,
      "author_name": "Mykola",
      "author_url": "",
      "post_date": "2023-05-02T16:32:14.707000",
      "content": "<p>Congrats on the perfect result! </p>\n<p>May I ask you to clarify a way you use cross-validation? You mentioned <code>Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters</code>. Do you mean you run all 8 folds to validate your model? Or the single fold was used most of the time? Am I also wondering, how do you validate an ensemble of models? I trained all the folds for each model, collect OOF predictions and calculate the ensemble using CV over OOF. It was a time-consuming process and I am not sure whether this is the best way to follow. <br>\nWould love to learn from your experience</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2243053,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-02T16:54:04.137000",
          "content": "<p>Thanks!</p>\n<p>We tested few times for 5 folds, and with random split OOF scores was almost the same for all folds. That's why after this we trained only fold 0 for all splits. We tried 5 folds, 8 folds, 16 folds. 8 folds version was the best</p>\n<p>After finding of some best checkpoint for fold 0 we trained full data checkpoint with the same params.</p>\n<p>For getting best ensemble weights we trained all models (eff-b0, bert, deberta) with fold 0 and calculate best weights for ensemble. It was 4 numbers: <br>\n<code>cnn_wgt * CNN + trans_wgt * (bert * bert_wgt + deberta * deberta_wgt)</code><br>\n Also we have tried manually tune weights on LB score but always that fold0 weights was the best.</p>\n<p>You can check how we did it in last 2 commented cells in <a href=\"https://github.com/ffs333/2nd_place_GISLR/blob/main/kaggle_inference_notebook_example.ipynb\" target=\"_blank\">our inference notebook</a></p>\n<p>Also sorry about your shakedown, I know that feeling 🙃</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2243169,
              "author_name": "Mykola",
              "author_url": "",
              "post_date": "2023-05-02T18:02:02.593000",
              "content": "<p>Thank you so much for your answer!</p>\n<blockquote>\n  <p>Also sorry about your shakedown, I know that feeling 🙃</p>\n</blockquote>\n<p>Ah, that is painful, taking into account that it was my first real competition I participate from start to end almost evening. Nonetheless, I am looking forward to meeting you in one of the next competitions! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2242218,
      "author_name": "zj yang",
      "author_url": "",
      "post_date": "2023-05-02T06:01:22.063000",
      "content": "<p>Could you further explain the meaning of your xyz heatmap? I do not understand why you put it in the repo…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2242237,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-02T06:15:41.597000",
          "content": "<p>data is:</p>\n<pre><code>sample = train.loc[i] \nyy = load_relevant_data_subset(base_dir + sample[])  \n</code></pre>\n<p>In training dataset class we just combine all these <code>yy</code> to list and iterate through it. We keep it RAM for faster training.<br>\nAnd load it like this <code>data = np.load(_cfg.base_path + 'gen_xyz/data.npy', allow_pickle=True)</code><br>\nYou can see it in repo in <code>GISLR_utils/mixup_data.py</code> <code>DatasetImageSmall80Mixup</code></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2290887,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-07T07:16:53.420000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2290951,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-06-07T08:13:20.387000",
              "content": "<p>Answer <a href=\"https://github.com/ffs333/2nd_place_GISLR/issues/1#issuecomment-1555364919\" target=\"_blank\">here</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2242200,
      "author_name": "JimmyWang",
      "author_url": "",
      "post_date": "2023-05-02T05:40:05.650000",
      "content": "<p>Thanks for sharing, and congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2242198,
      "author_name": "gutendzx",
      "author_url": "",
      "post_date": "2023-05-02T05:32:59.060000",
      "content": "<p>Amazing!Thanks for sharing the detail approaches here.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2242141,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "2023-05-02T04:06:32.953000",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> cool solution! Congrats with 2nd place!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2242098,
      "author_name": "Emre Kurtoglu",
      "author_url": "",
      "post_date": "2023-05-02T03:25:04.227000",
      "content": "<p>Thank you for the explanation! Is the source code going to be published publicly? That would be great.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2242150,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-05-02T04:20:45.850000",
          "content": "<p>Sure, it posted <a href=\"https://github.com/ffs333/2nd_place_GISLR\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2241955,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2023-05-02T01:13:55.517000",
      "content": "<p>Hearty congratulations to you for the fantastic result and associated gold medal and best regards <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2241929,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2023-05-02T00:50:25.557000",
      "content": "<p>Very impressive.  Thanks for sharing the detail approaches here.   </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2970705,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-08-26T12:17:03.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2242422,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-02T09:07:43",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246448,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-05T06:47:22.663000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2241899": "## TLDR\nWe used an approach similar to audio spectrogram classification using the EfficientNet-B0 model, with numerous augmentations and transformer models such as BERT and DeBERTa as helper models. The final solution consists of one EfficientNet-B0 with an input size of 160x80, trained on a single fold from 8 randomly split folds, as well as DeBERTa and BERT trained on the full dataset. A single fold model using EfficientNet has a CV score of 0.898 and a leaderboard score of ~0.8.\n\nWe used only competition data.\n\n\n\n## 1. Data Preprocessing\n\n### 1.1 CNN Preprocessing\n\n* We extracted 18 lip points, 20 pose points (including arms, shoulders, eyebrows, and nose), and all hand points, resulting in a total of 80 points.\n* During training, we applied various augmentations.\n* We implemented standard normalization.\n* Instead of dropping NaN values, we filled them with zeros after normalization.\n* We interpolated the time axis to a size of 160 using 'nearest' interpolation: `yy = F.interpolate(yy[None, None, :], size=self.new_size, mode='nearest')`.\n* Finally, we obtained a tensor with dimensions 160x80x3, where 3 represents the `(X, Y, Z)` axes. ![Preprocessing](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2Fc47290891a7ac6497a6a0c296973f071%2Fdata_prep.jpg?generation=1682986293067532&alt=media)\n\n### 1.2 Transformer Preprocessing\n\n* Only 61 points were kept, including 40 lip points and 21 hand points. For left and right hand, the one with less NaN was kept. If right hand was kept, mirror it to left hand.\n\n* Augmentations, normalization and NaN-filling were applied sequentially.\n\n* Sequences longer than 96 were interpolated to 96. Sequences shorter than 96 were unchanged.\n\n* Apart from raw positions, hand-crafted features were also used, including motion, distances, and cosine of angles.\n\n* Motion features consist of future motion and history motion, which can be denoted as:\n  \n$$\n  Motion_{future} = position_{t+1} - position_{t}\n$$\n$$\n  Motion_{history} = position_{t} - position_{t-1}\n$$\n\n* Full 210 pairwise distances among 21 hand points were included. \n* There are 5 vertices in a finger (e.g. thumb is `[0,1,2,3,4]`), and therefore, there are 3 angles: `<0,1,2>, <1,2,3>, <2,3,4>`. So 15 angles of 5 fingers were included.\n\n* Randomly selected 190 pairwise distances and randomly selected 8 angles among 40 lip points were included.\n\n## 2. Augmentation\n\n### 2.1 Common Augmentations\n\n> These augmentations are used in both CNN training and transformer training\n\n1. `Random affine`: Same as @hengck23 shared. In CNN, after global affine, shift-scale-rotate was also applied to each part separately (e.g. hand, lip, body-pose).\n\n2. `Random interpolation`: Slightly scale and shift the time dimension.\n\n3. `Flip pose`: Flip the x-coordinates of all points. In CNN, `x_new = x_max - x_old`. In transformer, `x_new = 2 * frame[:,0,0] - x_old`.\n\n4. `Finger tree rotate`: There are 4 root-children pairs in a finger with 5-vertices. E.g. in thumb (`[0,1,2,3,4]`), these 4 root-children pairs are: `0-[1,2,3,4]`,`1-[2,3,4]`,`2-[3,4]`,`3-[4]`. We randomly choose some of these pairs, and rotate the children points around root point with a small random angle.\n\n### 2.2 CNN Specific Augmentations \n\n* `Mixup`: Implement basic mixup augmentation (only works with CNNs, not transformers).\n* `Replace augmentation`: Replace some random parts from other samples of the same class.\n* `Time and frequence masking`: This basic torchaudio augmentation works exceptionally well.\n\n```python\nfreq_m = torchaudio.transforms.FrequencyMasking(80)  # it's time axis\ntime_m = torchaudio.transforms.TimeMasking(18)       # it's points axis\n```\n\n### 2.3 Augmented Sample Example\n\nBefore augmentation:\n\n<p align=\"center\"><img alt=\"aug1\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F5d2b97cf6754c5f4063724181bfe7172%2Fbefore_aug.png?generation=1682986332091937&alt=media\" width=\"600\" height=\"200\"/> </p>\n\nAfter augmentation:\n\n<p align=\"center\"> <img alt=\"aug2\" height=\"200\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F516f0790f6f9903ad8943977bc392965%2Fafter_aug.png?generation=1682986413818912&alt=media\" width=\"600\"/> </p>\n\n## 3. Training\n\n### 3.1 CNN Training\n\n* Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters\n* Onecycle scheduler with 0.1 warmup.\n* Use weighted `CrossEntropyLoss`. Increase the weights for poorly predicted classes and classes with semantically similar pairs (such as kitty and cat)\n* Implement a hypercolumn for EfficientNet with 5 blocks\n\n\n\n### 3.2 Transformer Training\n\n* Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters\n* Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule.\n* A 4-layer, 256 hidden-size, 512 intermediate-size transformer were trained.\n* A 3-layer model was initialized with 4-layer model's first 3 layers. Knowledge distillation were used in 3-layer model training, in which the 4-layer model is the teacher.\n\n### 3.3 Hyperparameter Tuning\n\nSince we trained only one fold and used smaller models, we decided to tune most parameters with Optuna. \n\nHere is the parameters list of CNN training (transformer training has a similar param-list):\n\n* All augmentations probabilities (0.1 - 0.5+)\n\n* Learning rate (2e-3 - 3e-3)\n\n* Drop out (0.1 - 0.25)\n\n* Num of epochs (170-185)\n\n* Loss weights powers (0.75 - 2)\n\n* Optimizer (`Lookahead_RAdam`, `RAdam`)\n\n* Label smoothing (0.5 - 0.7)\n\n## 4. Submissions, Conversion and Ensemble\n\n1. We rewrote all our models in Keras and transferred PyTorch weights to them, resulting in a speed boost of around 30%. For transformer model, pytorch-onnx-tf-tflite will generate too much useless tensor shape operations, a fully rewriting can reduce these manually. For CNN model, we rewrote DepthwiseConv2D with a hard-coded way, whose speed is 200%~300% of its original version of tflite DepthwiseConv2D.\n\n2. After that, we aggregated all these models in the `tf.Module` class. Converting directly from Keras resulted in lower speed (don't know why).\n\n3. We calculated ensemble weights for models trained on fold 0 using the local fold 0 score and applied these weights to the full dataset models.\n\nEfficientNet-B0 achieved a leaderboard score of approximately 0.8, and transformers improved the score to 0.81. The final ensemble included:\n1. Efficientnet-B0, fold 0\n2. BERT, full data train\n3. DeBERTa, full data train\n\nInterestingly, a key feature was using the ensemble without softmax, which consistently provided a boost of around 0.01.\n\n## 5. PS. Need **BETTER** TFlite DepthwiseConv2D\n\nDepthwise convolution models performed very well for these tasks, outperforming other CNN and ViT models (rexnet_100 was also good).\nWe spent a lot of time dealing with the conversion of DepthwiseConv2D operation. Here are some strange results:\n\nGiven a input image with 82x42x32 (HWC), there are two ways to do a 3x3 depthwise convolution in Keras. One is `Conv2D(32, 3, groups = 32)`, the other is `DepthwiseConv2D(3)`. However, after converting these two to tflite, the running time of the `Conv2D` is 5.05ms, and the running time of `DepthwiseConv2D` is 3.70ms. More strangely, a full convolution `Conv2D(32, 3, groups = 1)` with FLOPs = HWC^2 only takes 2.09ms, even faster than previous two with FLOPs = HWC.\n\nThen we rewrote the depthwise-conv like this:\n\n```python\n    def call(self, x):\n        out = x[:,0:self.H_out:self.strides,0:self.W_out:self.strides] * self.weight[0,0]\n        for i in range(self.kernel_size):\n            for j in range(self.kernel_size):\n                if i == 0 and j == 0:\n                    continue\n                out += x[:,i:self.H_out + i:self.strides,j:self.W_out + j:self.strides] * self.weight[i,j]\n        if self.bias is not None:\n            out = out + self.bias\n        return out\n```\n\nThe running time of this is 1.24 ms.\n\nIn summary, our version (1.24ms) > full `Conv2D` with larger FLOPs (2.09ms) > `DepthwiseConv2D` (3.70ms) > `Conv2D(C, groups = C)` (5.05ms).\n\nHowever, our version introduced too much nodes in tflite graph, which is not stable in running time. If the tensorflow team has a better implementation of DepthwiseConv2D, we can even ensemble two CNN models, which is expected to reach 0.82 LB.\n\nBy the way, EfficientNet with ONNX was ~5 times faster than TFLite.\n\n\n### Big thanks to my teammates @artemtprv and @carnozhao and congrats with new tiers, Master and GrandMaster!\n\n\n[github code](https://github.com/ffs333/2nd_place_GISLR)",
    "2242254": "Congrats for the 2nd place! \nI haven't even imagine treating this data as image... (I like audio spectrogram analogy by the way!)\nWell deserved!",
    "2242329": "Very nice work for both branches (CNN, Transformer) of the solution. How did you identify Depthwise convolution particularly as the bottleneck ? I tried efficientnet but it was too slow for me... however, I didnt rewrite the base model in to tf, only the prepreprocessing. I see that transferring the base model to keras gave a 30% speed up. \n\nAlso, was the ordering of the xyz points important in the CNN. I guess you would want adjacent points beside each other in the CNN. But some points, like wrist, have multiple adjacent points. ",
    "2241922": "By the way, for us in public LB score for full data train CNN versions was worse and we haven't select it. But in private they were better. But still lower than @hoyso48 . That's a pure victory on last hours, GZ!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F10ba2123e7c100e364688e095a50b569%2Ftg_image_455350557.jpeg?generation=1682988215699105&alt=media)",
    "2242681": "Congrats on the great result. Thanks for the detailed write-up. Found a lot of useful tricks that are gonna help in the future 🔥",
    "2242511": "Congrats great approach and good idea to use it as image 🎉🎉🎉🔥",
    "2241912": "Congratz with the result! \nAbout DepthwiseConv2D, - I had the same problem when depthwise conv1d was converted not efficiently when using pytorch-onnx-tf pipeline, I used it for positional encoding in transformer.\n\nYou can try to check output conversion nodes via tensorflow benchmark that hengck shared, for me it appeared that whole transformer takes 200 ops and single depthwise conv1d takes 1800 ops (???) when using pytorch->onnx->tf conversion pipeline, and it was indeed very slow.\n\nI fixed it when I used [nobuco](https://github.com/AlexanderLutsenko/nobuco), - depthwise convolution converted succesfully using built-in tf function.",
    "2243391": "Congratz on the result and on the overall impressive performance !\n0.81 public was impressive, we guessed you were using CNNs but could not get them to match our transformer performances.",
    "2242007": "Big congratulations, and thanks for write-ups. It's amazing to read your struggle to speed up the inference 👍",
    "2433850": "Thank you so much for posting your solution. Great job!!",
    "2392704": "Hi,\nCongrats on second place!\n\nWe are trying to run the code on the colab file training_notebook_example, and I saw that there is a possibility to run TREF model, which for what I understand run both CNN and transformers model. the problem I faced was choosing this model because I got an error when the name was sent to get_model() and I noticed that there is no option there. Can you help me with that.\n\nThank you and congrats again!",
    "2381121": "Could someone explain this part for the normalization with how you found those values to be replaced by  (lign 237 to 243 of data.py) :   \nself.interesting_idx = np.array(LIPS + l_hand + pose + r_hand)\n # norm params\n        self.x_mean = torch.tensor(0.48282188177108765).float()\n        self.y_mean = torch.tensor(0.6022918224334717).float()\n        self.z_mean = torch.tensor(-0.4806228578090668).float()\n        self.x_std = torch.tensor(0.22193588316440582).float()\n        self.y_std = torch.tensor(0.252733051776886).float()\n        self.z_std = torch.tensor(0.68902188539505).float()",
    "2355782": "@kolyaforrat great work. I could not find the transfer_transformer function used in inference code. Could you help me find it? ",
    "2246172": "Great work, 2nd place is impressive. Hope to learn from your approach and use it to solve similar problems. ",
    "2245705": "Great work! I'm very inspired by your spectrogram classification approach. Very clever! I've been trying to recreate it locally from scratch using your code on github a a guide. Any chance you could share a screenshot of the training curves for the train/val loss and accuracy and learning rate?",
    "2244071": "So cool to read about your approach, thank you for sharing! Interestingly, as probably many of us, I could not get CNN to work on this problem, now I see why. \"Devil is in the details\". Congrats on the 2nd place, well deserved! ",
    "2243854": "Congrats ! You did a great work",
    "2243335": "Now just tried to submit single fold solo CNN model and got this score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F867040d30c2f4535cd33fd19989de3d3%2Ftg_image_380692410.jpeg?generation=1683059371986310&alt=media)\n\nIt's 5th place in public and 4th in private with 17.5 MB and 42 minutes",
    "2243334": "wow great !this will be very helpful this is what in summary i understood-\nThey used an approach similar to audio spectrogram classification for their solution.\nThey used EfficientNet-B0 as the main model and applied various augmentations and transformer models such as BERT and DeBERTa as helper models.\nThe team extracted 80 points for CNN preprocessing and 61 points for Transformer preprocessing.\nThey used various augmentations such as random affine, random interpolation, flip pose, finger tree rotate, mixup, replace augmentation, time and frequency masking.\nThey trained on one fold with a random split (8 folds in total) or the full dataset using the best parameters.\nThey used a onecycle scheduler with 0.1 warmup and a weighted CrossEntropyLoss for CNN training, and a Ranger optimizer with 60% flat and 40% cosine annealing learning rate schedule for Transformer training.\nThey implemented hyperparameter tuning with Optuna for most parameters.\nThey rewrote all their models in Keras and transferred PyTorch models to Keras models.\nThey used an ensemble of different models to improve their performance.",
    "2243059": "@kolyaforrat  @artemtprv @carnozhaocongrats with the second highest place in so competitive contest!💥🎉 The 5-th paragraph of you post is especially interesting as well as Efficientnet+BERT+DeBERTa ensemble approach👍💪🙂",
    "2243019": "Congrats on the perfect result! \n\nMay I ask you to clarify a way you use cross-validation? You mentioned `Train on one fold with a random split (8 folds in total) or the full dataset using the best parameters`. Do you mean you run all 8 folds to validate your model? Or the single fold was used most of the time? Am I also wondering, how do you validate an ensemble of models? I trained all the folds for each model, collect OOF predictions and calculate the ensemble using CV over OOF. It was a time-consuming process and I am not sure whether this is the best way to follow. \nWould love to learn from your experience",
    "2242218": "Could you further explain the meaning of your xyz heatmap? I do not understand why you put it in the repo...",
    "2242200": "Thanks for sharing, and congratulations!",
    "2242198": "Amazing!Thanks for sharing the detail approaches here.",
    "2242141": "@kolyaforrat cool solution! Congrats with 2nd place!",
    "2242098": "Thank you for the explanation! Is the source code going to be published publicly? That would be great.",
    "2241955": "Hearty congratulations to you for the fantastic result and associated gold medal and best regards @kolyaforrat ",
    "2241929": "Very impressive.  Thanks for sharing the detail approaches here.   ",
    "2970705": "",
    "2242422": "",
    "2246448": "Amazing work!\nThanks for sharing."
  }
}