{
  "id": 406684,
  "title": "1st place solution - 1DCNN combined with Transformer",
  "url": "/competitions/asl-signs/discussion/406684",
  "author_name": "hoyso48",
  "post_date": "2023-05-03T11:54:48.673000",
  "votes": 191,
  "comment_count": 48,
  "views": 0,
  "content": "<p>First of all, I would like to express my gratitude to the Google for hosting this amazing competition. I have always been a big fan of the services, frameworks, and platforms provided by Google.(Colab, GCP, TensorFlow.. all of them are amazing). Without access to these offerings from Google, I wouldn't have been able to win this competition. I would also like to thank all the other participants who shared their ideas. In particular, I gained valuable insights from the ideas shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>\n<h1>TL;DR</h1>\n<p>My solution involved a combination of a 1D CNN and a Transformer, trained from scratch using all training data(competition data only), and used 4x seed ensemble for submission. and I initially started with PyTorch + GPU but later switched to TensorFlow + Colab TPU(tpuv2-8) to ensure compatibility with TensorFlow Lite.</p>\n<h1>1D CNN vs. Transformer?</h1>\n<p>My hypothesis was that in modeling sequential data, <br>\n<strong>if there is a strong inter-frame correlation, 1D CNNs would be more efficient than Transformers.</strong></p>\n<p>As in my experiment, the pure 1D CNN easily outperformed the Transformer and as a result, I was able to achieve a public LB score of 0.80 using only the 1D CNN at the end. </p>\n<p>However, there still were roles for the Transformer, which could be used on top of the 1D CNN(we can view 1d cnn as some kind of trainable tokenizer).</p>\n<h1>Model</h1>\n<pre><code> ():\n    inp = tf.keras.Input((max_len,CHANNELS))\n    x = tf.keras.layers.Masking(mask_value=PAD,input_shape=(max_len,CHANNELS))(inp)\n    ksize = \n    x = tf.keras.layers.Dense(dim, use_bias=,name=)(x)\n    x = tf.keras.layers.BatchNormalization(momentum=,name=)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = TransformerBlock(dim,expand=)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = TransformerBlock(dim,expand=)(x)\n\n     dim == : \n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = TransformerBlock(dim,expand=)(x)\n\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = TransformerBlock(dim,expand=)(x)\n\n    x = tf.keras.layers.Dense(dim*,activation=,name=)(x)\n    x = tf.keras.layers.GlobalAveragePooling1D()(x)\n    x = LateDropout(, start_step=dropout_step)(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES,name=)(x)\n     tf.keras.Model(inp, x)\n</code></pre>\n<p>Combining CNN and Transformer is a prevalent idea in recent state-of-the-art models(coatnet, conformer, Maxvit, nextvit…).  I started with an 192d 8-layer 1D CNN, then switched to a 192d (3+1)x2 conv-transformer structure, which yielded a +0.01 CV and LB improvement. </p>\n<p>The 1D CNN model employed depthwise convolution and causal padding. The Transformer used BatchNorm + Swish instead of the typical LayerNorm + GELU, due to slightly(negligible) lighter inference with the same accuracy.<br>\nsingle model has around 1.85M parameters.</p>\n<h1>Masking</h1>\n<p><strong>Handling variable-length input correctly was very crucial for ensuring train-test consistency and efficient inference</strong> , as we do not necessarily have to pad the short videos. During training, I used a max_len=384 with padding and truncation, while for inference, I only applied truncation. This approach provided sufficient inference speed and allowed the use of reasonably large models. To accurately  apply masking to the 1D CNN, I used causal padding to maintain the mask index. In TensorFlow, masking can be easily implemented using tf.keras.layers.Masking at the beginning of the model. Plus, It is essential to ensure that masking is accurately applied to operations like batch normalization and global average pooling which can be affected by masking.</p>\n<h1>Regularization</h1>\n<ol>\n<li>Drop Path(stochastic depth, p=0.2)</li>\n<li>high rate of Dropout (p=0.8)</li>\n<li>AWP(Adversarial Weight Perturbation, with lambda = 0.2)</li>\n</ol>\n<h6> </h6>\n<p>As we need to train the model from scratch, regularization technics played a significant role. I used drop_path(0.2, applied after each block), dropout(0.8, applied after GAP) and AWP(Adversarial Weight Perturbation, lambda=0.2) AWP and the dropout applied after epoch 15. All these methods were very crucial for preventing overfitting when training with long epochs(&gt;300). All three methods had a significant impact on both CV and leaderboard scores, and removing any one of them led to noticeable performance drops.</p>\n<h1>Preprocessing</h1>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n        ().__init__(**kwargs)\n        self.max_len = max_len\n        self.point_landmarks = point_landmarks\n\n     ():\n         tf.rank(inputs) == :\n            x = inputs[,...]\n        :\n            x = inputs\n\n        mean = tf_nan_mean(tf.gather(x, [], axis=), axis=[,], keepdims=)\n        mean = tf.where(tf.math.is_nan(mean), tf.constant(,x.dtype), mean)\n        x = tf.gather(x, self.point_landmarks, axis=) \n        std = tf_nan_std(x, center=mean, axis=[,], keepdims=)\n\n        x = (x - mean)/std\n\n         self.max_len   :\n            x = x[:,:self.max_len]\n        length = tf.shape(x)[]\n        x = x[...,:]\n\n        dx = tf.cond(tf.shape(x)[]&gt;,:tf.pad(x[:,:] - x[:,:-], [[,],[,],[,],[,]]),:tf.zeros_like(x))\n\n        dx2 = tf.cond(tf.shape(x)[]&gt;,:tf.pad(x[:,:] - x[:,:-], [[,],[,],[,],[,]]),:tf.zeros_like(x))\n\n        x = tf.concat([\n            tf.reshape(x, (-,length,*(self.point_landmarks))),\n            tf.reshape(dx, (-,length,*(self.point_landmarks))),\n            tf.reshape(dx2, (-,length,*(self.point_landmarks))),\n        ], axis = -)\n\n        x = tf.where(tf.math.is_nan(x),tf.constant(,x.dtype),x)\n\n         x\n</code></pre>\n<p>I used left-right hand, eye, nose, and lips landmarks. For normalization, I used the 17th landmark located in the nose as a reference point, since it is usually located close to the center ([0.5, 0.5]). I used motion feature of lag1 x[:1] - x[1:], and lag2 x[:2] - x[2:](lag &gt; 2 did not help much).</p>\n<h1>Augmentation</h1>\n<ul>\n<li>temporal augmentation</li>\n</ul>\n<ol>\n<li>Random resample (0.5x ~ 1.5x to original length)</li>\n<li>Random masking</li>\n</ol>\n<ul>\n<li>Spatial augmentation</li>\n</ul>\n<ol>\n<li>hflip</li>\n<li>Random Affine(Scale, shift, rotate, shear)</li>\n<li>Random Cutout</li>\n</ol>\n<h1>Training</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Ff23fef7388aa47b7a1e26f6261e6c0a4%2F.png?generation=1683114294544440&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Epoch = 400</li>\n<li>Lr = 5e-4 * num_replicas = 4e-3</li>\n<li>Schedule = CosineDecay with no warmup</li>\n<li>Optimizer = RAdam with Lookahead(better than AdamW with optimal parameters)</li>\n<li>Loss = CCE with label smoothing=0.1 or just plain CCE</li>\n</ul>\n<h6> </h6>\n<p>final single model result<br>\nCV(participant split 5fold): 0.80<br>\npublic LB: 0.80<br>\nprivate LB: 0.88</p>\n<h6> </h6>\n<p>Training takes around 4 hours with colab TPUv2-8.</p>\n<p>Single model CV was around 0.80 with participant split(5fold) at the end. I ensemble 4 different seed(with some minor differences in training configurations) for the final model and got LB 0.81. By the way, I could see slightly worse score when I submitted 4x size(384d, 16layers) single model with the same settings. I think it can achieve same or better score with better configurations.</p>\n<h1>Tried but not worked</h1>\n<ul>\n<li>GCNs</li>\n<li>More Complex augmentations: augmentation based on angle and the distance between the landmarks, grid distortion on temporal, spatial axis, etc.</li>\n<li>CutMix, MixUp: Not worked. Main problem was how to define new label with two different length of inputs. I could not find the correct way to implement it.</li>\n<li>Knowledge Distillation: I tried to use single 4x sized model and distill it with 4x seed 4x sized model, but I could not manage it to work due to lack of time.</li>\n</ul>\n<h6> </h6>\n<p>I'm always amazed by the fact that we did try each other’s methods, but we came up with different result. I also gave the 2D-CNN(similar to <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> 's brilliant solution) and pure transformer approaches a try, but due to their initial weak performance and my confidence in my own hypothesis, I didn't dig any deeper. Seeing other teams succeed with the ideas I had trouble with, through their skill and hard work, has been truly inspiring. I've learned a lot through this competition, and I'm grateful for the experience. Thank you all!!! :)</p>\n<p>Edit: I made my code public, check <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406978</a></p>",
  "messages": [
    {
      "id": 2244052,
      "postDate": "2023-05-03T11:54:48.673Z",
      "content": "<p>First of all, I would like to express my gratitude to the Google for hosting this amazing competition. I have always been a big fan of the services, frameworks, and platforms provided by Google.(Colab, GCP, TensorFlow.. all of them are amazing). Without access to these offerings from Google, I wouldn't have been able to win this competition. I would also like to thank all the other participants who shared their ideas. In particular, I gained valuable insights from the ideas shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</p>\n<h1>TL;DR</h1>\n<p>My solution involved a combination of a 1D CNN and a Transformer, trained from scratch using all training data(competition data only), and used 4x seed ensemble for submission. and I initially started with PyTorch + GPU but later switched to TensorFlow + Colab TPU(tpuv2-8) to ensure compatibility with TensorFlow Lite.</p>\n<h1>1D CNN vs. Transformer?</h1>\n<p>My hypothesis was that in modeling sequential data, <br>\n<strong>if there is a strong inter-frame correlation, 1D CNNs would be more efficient than Transformers.</strong></p>\n<p>As in my experiment, the pure 1D CNN easily outperformed the Transformer and as a result, I was able to achieve a public LB score of 0.80 using only the 1D CNN at the end. </p>\n<p>However, there still were roles for the Transformer, which could be used on top of the 1D CNN(we can view 1d cnn as some kind of trainable tokenizer).</p>\n<h1>Model</h1>\n<pre><code> ():\n    inp = tf.keras.Input((max_len,CHANNELS))\n    x = tf.keras.layers.Masking(mask_value=PAD,input_shape=(max_len,CHANNELS))(inp)\n    ksize = \n    x = tf.keras.layers.Dense(dim, use_bias=,name=)(x)\n    x = tf.keras.layers.BatchNormalization(momentum=,name=)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = TransformerBlock(dim,expand=)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n    x = TransformerBlock(dim,expand=)(x)\n\n     dim == : \n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = TransformerBlock(dim,expand=)(x)\n\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=)(x)\n        x = TransformerBlock(dim,expand=)(x)\n\n    x = tf.keras.layers.Dense(dim*,activation=,name=)(x)\n    x = tf.keras.layers.GlobalAveragePooling1D()(x)\n    x = LateDropout(, start_step=dropout_step)(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES,name=)(x)\n     tf.keras.Model(inp, x)\n</code></pre>\n<p>Combining CNN and Transformer is a prevalent idea in recent state-of-the-art models(coatnet, conformer, Maxvit, nextvit…).  I started with an 192d 8-layer 1D CNN, then switched to a 192d (3+1)x2 conv-transformer structure, which yielded a +0.01 CV and LB improvement. </p>\n<p>The 1D CNN model employed depthwise convolution and causal padding. The Transformer used BatchNorm + Swish instead of the typical LayerNorm + GELU, due to slightly(negligible) lighter inference with the same accuracy.<br>\nsingle model has around 1.85M parameters.</p>\n<h1>Masking</h1>\n<p><strong>Handling variable-length input correctly was very crucial for ensuring train-test consistency and efficient inference</strong> , as we do not necessarily have to pad the short videos. During training, I used a max_len=384 with padding and truncation, while for inference, I only applied truncation. This approach provided sufficient inference speed and allowed the use of reasonably large models. To accurately  apply masking to the 1D CNN, I used causal padding to maintain the mask index. In TensorFlow, masking can be easily implemented using tf.keras.layers.Masking at the beginning of the model. Plus, It is essential to ensure that masking is accurately applied to operations like batch normalization and global average pooling which can be affected by masking.</p>\n<h1>Regularization</h1>\n<ol>\n<li>Drop Path(stochastic depth, p=0.2)</li>\n<li>high rate of Dropout (p=0.8)</li>\n<li>AWP(Adversarial Weight Perturbation, with lambda = 0.2)</li>\n</ol>\n<h6> </h6>\n<p>As we need to train the model from scratch, regularization technics played a significant role. I used drop_path(0.2, applied after each block), dropout(0.8, applied after GAP) and AWP(Adversarial Weight Perturbation, lambda=0.2) AWP and the dropout applied after epoch 15. All these methods were very crucial for preventing overfitting when training with long epochs(&gt;300). All three methods had a significant impact on both CV and leaderboard scores, and removing any one of them led to noticeable performance drops.</p>\n<h1>Preprocessing</h1>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n        ().__init__(**kwargs)\n        self.max_len = max_len\n        self.point_landmarks = point_landmarks\n\n     ():\n         tf.rank(inputs) == :\n            x = inputs[,...]\n        :\n            x = inputs\n\n        mean = tf_nan_mean(tf.gather(x, [], axis=), axis=[,], keepdims=)\n        mean = tf.where(tf.math.is_nan(mean), tf.constant(,x.dtype), mean)\n        x = tf.gather(x, self.point_landmarks, axis=) \n        std = tf_nan_std(x, center=mean, axis=[,], keepdims=)\n\n        x = (x - mean)/std\n\n         self.max_len   :\n            x = x[:,:self.max_len]\n        length = tf.shape(x)[]\n        x = x[...,:]\n\n        dx = tf.cond(tf.shape(x)[]&gt;,:tf.pad(x[:,:] - x[:,:-], [[,],[,],[,],[,]]),:tf.zeros_like(x))\n\n        dx2 = tf.cond(tf.shape(x)[]&gt;,:tf.pad(x[:,:] - x[:,:-], [[,],[,],[,],[,]]),:tf.zeros_like(x))\n\n        x = tf.concat([\n            tf.reshape(x, (-,length,*(self.point_landmarks))),\n            tf.reshape(dx, (-,length,*(self.point_landmarks))),\n            tf.reshape(dx2, (-,length,*(self.point_landmarks))),\n        ], axis = -)\n\n        x = tf.where(tf.math.is_nan(x),tf.constant(,x.dtype),x)\n\n         x\n</code></pre>\n<p>I used left-right hand, eye, nose, and lips landmarks. For normalization, I used the 17th landmark located in the nose as a reference point, since it is usually located close to the center ([0.5, 0.5]). I used motion feature of lag1 x[:1] - x[1:], and lag2 x[:2] - x[2:](lag &gt; 2 did not help much).</p>\n<h1>Augmentation</h1>\n<ul>\n<li>temporal augmentation</li>\n</ul>\n<ol>\n<li>Random resample (0.5x ~ 1.5x to original length)</li>\n<li>Random masking</li>\n</ol>\n<ul>\n<li>Spatial augmentation</li>\n</ul>\n<ol>\n<li>hflip</li>\n<li>Random Affine(Scale, shift, rotate, shear)</li>\n<li>Random Cutout</li>\n</ol>\n<h1>Training</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Ff23fef7388aa47b7a1e26f6261e6c0a4%2F.png?generation=1683114294544440&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Epoch = 400</li>\n<li>Lr = 5e-4 * num_replicas = 4e-3</li>\n<li>Schedule = CosineDecay with no warmup</li>\n<li>Optimizer = RAdam with Lookahead(better than AdamW with optimal parameters)</li>\n<li>Loss = CCE with label smoothing=0.1 or just plain CCE</li>\n</ul>\n<h6> </h6>\n<p>final single model result<br>\nCV(participant split 5fold): 0.80<br>\npublic LB: 0.80<br>\nprivate LB: 0.88</p>\n<h6> </h6>\n<p>Training takes around 4 hours with colab TPUv2-8.</p>\n<p>Single model CV was around 0.80 with participant split(5fold) at the end. I ensemble 4 different seed(with some minor differences in training configurations) for the final model and got LB 0.81. By the way, I could see slightly worse score when I submitted 4x size(384d, 16layers) single model with the same settings. I think it can achieve same or better score with better configurations.</p>\n<h1>Tried but not worked</h1>\n<ul>\n<li>GCNs</li>\n<li>More Complex augmentations: augmentation based on angle and the distance between the landmarks, grid distortion on temporal, spatial axis, etc.</li>\n<li>CutMix, MixUp: Not worked. Main problem was how to define new label with two different length of inputs. I could not find the correct way to implement it.</li>\n<li>Knowledge Distillation: I tried to use single 4x sized model and distill it with 4x seed 4x sized model, but I could not manage it to work due to lack of time.</li>\n</ul>\n<h6> </h6>\n<p>I'm always amazed by the fact that we did try each other’s methods, but we came up with different result. I also gave the 2D-CNN(similar to <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> 's brilliant solution) and pure transformer approaches a try, but due to their initial weak performance and my confidence in my own hypothesis, I didn't dig any deeper. Seeing other teams succeed with the ideas I had trouble with, through their skill and hard work, has been truly inspiring. I've learned a lot through this competition, and I'm grateful for the experience. Thank you all!!! :)</p>\n<p>Edit: I made my code public, check <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406978</a></p>",
      "rawMarkdown": "First of all, I would like to express my gratitude to the Google for hosting this amazing competition. I have always been a big fan of the services, frameworks, and platforms provided by Google.(Colab, GCP, TensorFlow.. all of them are amazing). Without access to these offerings from Google, I wouldn't have been able to win this competition. I would also like to thank all the other participants who shared their ideas. In particular, I gained valuable insights from the ideas shared by @hengck23.\n\n# TL;DR\nMy solution involved a combination of a 1D CNN and a Transformer, trained from scratch using all training data(competition data only), and used 4x seed ensemble for submission. and I initially started with PyTorch + GPU but later switched to TensorFlow + Colab TPU(tpuv2-8) to ensure compatibility with TensorFlow Lite.\n\n# 1D CNN vs. Transformer?\nMy hypothesis was that in modeling sequential data, \n**if there is a strong inter-frame correlation, 1D CNNs would be more efficient than Transformers.**\n\nAs in my experiment, the pure 1D CNN easily outperformed the Transformer and as a result, I was able to achieve a public LB score of 0.80 using only the 1D CNN at the end. \n\nHowever, there still were roles for the Transformer, which could be used on top of the 1D CNN(we can view 1d cnn as some kind of trainable tokenizer).\n\n# Model\n```python\ndef get_model(max_len=64, dropout_step=0, dim=192):\n    inp = tf.keras.Input((max_len,CHANNELS))\n    x = tf.keras.layers.Masking(mask_value=PAD,input_shape=(max_len,CHANNELS))(inp)\n    ksize = 17\n    x = tf.keras.layers.Dense(dim, use_bias=False,name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(momentum=0.95,name='stem_bn')(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    if dim == 384: #for the 4x sized model\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = TransformerBlock(dim,expand=2)(x)\n\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = TransformerBlock(dim,expand=2)(x)\n\n    x = tf.keras.layers.Dense(dim*2,activation=None,name='top_conv')(x)\n    x = tf.keras.layers.GlobalAveragePooling1D()(x)\n    x = LateDropout(0.8, start_step=dropout_step)(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES,name='classifier')(x)\n    return tf.keras.Model(inp, x)\n\n```\nCombining CNN and Transformer is a prevalent idea in recent state-of-the-art models(coatnet, conformer, Maxvit, nextvit…).  I started with an 192d 8-layer 1D CNN, then switched to a 192d (3+1)x2 conv-transformer structure, which yielded a +0.01 CV and LB improvement. \n\n\nThe 1D CNN model employed depthwise convolution and causal padding. The Transformer used BatchNorm + Swish instead of the typical LayerNorm + GELU, due to slightly(negligible) lighter inference with the same accuracy.\nsingle model has around 1.85M parameters.\n\n# Masking\n**Handling variable-length input correctly was very crucial for ensuring train-test consistency and efficient inference** , as we do not necessarily have to pad the short videos. During training, I used a max_len=384 with padding and truncation, while for inference, I only applied truncation. This approach provided sufficient inference speed and allowed the use of reasonably large models. To accurately  apply masking to the 1D CNN, I used causal padding to maintain the mask index. In TensorFlow, masking can be easily implemented using tf.keras.layers.Masking at the beginning of the model. Plus, It is essential to ensure that masking is accurately applied to operations like batch normalization and global average pooling which can be affected by masking.\n\n# Regularization\n1. Drop Path(stochastic depth, p=0.2)\n2. high rate of Dropout (p=0.8)\n3. AWP(Adversarial Weight Perturbation, with lambda = 0.2)\n###### \nAs we need to train the model from scratch, regularization technics played a significant role. I used drop_path(0.2, applied after each block), dropout(0.8, applied after GAP) and AWP(Adversarial Weight Perturbation, lambda=0.2) AWP and the dropout applied after epoch 15. All these methods were very crucial for preventing overfitting when training with long epochs(>300). All three methods had a significant impact on both CV and leaderboard scores, and removing any one of them led to noticeable performance drops.\n\n# Preprocessing\n```python\nclass Preprocess(tf.keras.layers.Layer):\n    def __init__(self, max_len=MAX_LEN, point_landmarks=POINT_LANDMARKS, **kwargs):\n        super().__init__(**kwargs)\n        self.max_len = max_len\n        self.point_landmarks = point_landmarks\n\n    def call(self, inputs):\n        if tf.rank(inputs) == 3:\n            x = inputs[None,...]\n        else:\n            x = inputs\n        \n        mean = tf_nan_mean(tf.gather(x, [17], axis=2), axis=[1,2], keepdims=True)\n        mean = tf.where(tf.math.is_nan(mean), tf.constant(0.5,x.dtype), mean)\n        x = tf.gather(x, self.point_landmarks, axis=2) #N,T,P,C\n        std = tf_nan_std(x, center=mean, axis=[1,2], keepdims=True)\n        \n        x = (x - mean)/std\n\n        if self.max_len is not None:\n            x = x[:,:self.max_len]\n        length = tf.shape(x)[1]\n        x = x[...,:2]\n\n        dx = tf.cond(tf.shape(x)[1]>1,lambda:tf.pad(x[:,1:] - x[:,:-1], [[0,0],[0,1],[0,0],[0,0]]),lambda:tf.zeros_like(x))\n\n        dx2 = tf.cond(tf.shape(x)[1]>2,lambda:tf.pad(x[:,2:] - x[:,:-2], [[0,0],[0,2],[0,0],[0,0]]),lambda:tf.zeros_like(x))\n\n        x = tf.concat([\n            tf.reshape(x, (-1,length,2*len(self.point_landmarks))),\n            tf.reshape(dx, (-1,length,2*len(self.point_landmarks))),\n            tf.reshape(dx2, (-1,length,2*len(self.point_landmarks))),\n        ], axis = -1)\n        \n        x = tf.where(tf.math.is_nan(x),tf.constant(0.,x.dtype),x)\n        \n        return x\n```\nI used left-right hand, eye, nose, and lips landmarks. For normalization, I used the 17th landmark located in the nose as a reference point, since it is usually located close to the center ([0.5, 0.5]). I used motion feature of lag1 x[:1] - x[1:], and lag2 x[:2] - x[2:](lag > 2 did not help much).\n\n#Augmentation\n\n- temporal augmentation\n1. Random resample (0.5x ~ 1.5x to original length)\n2. Random masking\n\n- Spatial augmentation\n1. hflip\n2. Random Affine(Scale, shift, rotate, shear)\n3. Random Cutout\n\n#Training\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Ff23fef7388aa47b7a1e26f6261e6c0a4%2F.png?generation=1683114294544440&alt=media)\n- Epoch = 400\n- Lr = 5e-4 * num_replicas = 4e-3\n- Schedule = CosineDecay with no warmup\n- Optimizer = RAdam with Lookahead(better than AdamW with optimal parameters)\n- Loss = CCE with label smoothing=0.1 or just plain CCE\n###### \nfinal single model result\nCV(participant split 5fold): 0.80\npublic LB: 0.80\nprivate LB: 0.88\n###### \nTraining takes around 4 hours with colab TPUv2-8.\n\nSingle model CV was around 0.80 with participant split(5fold) at the end. I ensemble 4 different seed(with some minor differences in training configurations) for the final model and got LB 0.81. By the way, I could see slightly worse score when I submitted 4x size(384d, 16layers) single model with the same settings. I think it can achieve same or better score with better configurations.\n\n#Tried but not worked\n- GCNs\n- More Complex augmentations: augmentation based on angle and the distance between the landmarks, grid distortion on temporal, spatial axis, etc.\n- CutMix, MixUp: Not worked. Main problem was how to define new label with two different length of inputs. I could not find the correct way to implement it.\n- Knowledge Distillation: I tried to use single 4x sized model and distill it with 4x seed 4x sized model, but I could not manage it to work due to lack of time.\n###### \nI'm always amazed by the fact that we did try each other’s methods, but we came up with different result. I also gave the 2D-CNN(similar to @kolyaforrat 's brilliant solution) and pure transformer approaches a try, but due to their initial weak performance and my confidence in my own hypothesis, I didn't dig any deeper. Seeing other teams succeed with the ideas I had trouble with, through their skill and hard work, has been truly inspiring. I've learned a lot through this competition, and I'm grateful for the experience. Thank you all!!! :)\n\n\nEdit: I made my code public, check https://www.kaggle.com/competitions/asl-signs/discussion/406978",
      "votes": 191
    },
    {
      "id": 2245152,
      "postDate": "2023-05-04T07:24:13.913Z",
      "content": "<p>amazing solution. congrats and thanks for sharing.</p>",
      "rawMarkdown": "amazing solution. congrats and thanks for sharing.",
      "votes": 1
    },
    {
      "id": 2244928,
      "postDate": "2023-05-04T02:45:16.953Z",
      "content": "<p>Thank you for sharing. can the code be implemented on EEG signals for classification?</p>",
      "rawMarkdown": "Thank you for sharing. can the code be implemented on EEG signals for classification?",
      "votes": 1,
      "replies": [
        {
          "id": 2245901,
          "postDate": "2023-05-04T16:54:30.433Z",
          "content": "<p>I believe yes, with some modifications.</p>",
          "rawMarkdown": "I believe yes, with some modifications."
        }
      ]
    },
    {
      "id": 2244632,
      "postDate": "2023-05-03T19:08:02.177Z",
      "content": "<p>Many congratulations. Thank you for sharing!</p>",
      "rawMarkdown": "Many congratulations. Thank you for sharing!",
      "votes": 1
    },
    {
      "id": 2244625,
      "postDate": "2023-05-03T18:59:22.287Z",
      "content": "<p>Really cool solution (Cool approach in general, there weren't a lot of conv + transformer architectures to this point!).</p>\n<p>Quick question, Is the masking <code>PAD</code> for ignoring gradients computed for shorter sequences in the batch? <br>\n[I assume the causal masking is inside the transformer block itself, am I correct?]</p>\n<p>Congrats, and thank you again for sharing! <br>\nAmazing work. I am a bit shocked by how simple it is.. </p>",
      "rawMarkdown": "Really cool solution (Cool approach in general, there weren't a lot of conv + transformer architectures to this point!).\n\nQuick question, Is the masking `PAD` for ignoring gradients computed for shorter sequences in the batch? \n[I assume the causal masking is inside the transformer block itself, am I correct?]\n\n\nCongrats, and thank you again for sharing! \nAmazing work. I am a bit shocked by how simple it is.. ",
      "votes": 1,
      "replies": [
        {
          "id": 2244693,
          "postDate": "2023-05-03T20:18:15.437Z",
          "content": "<p>Masking of PAD is literally for the masking of PAD tokens in the sequence. for example, if the data have the length of 256, 128 PAD tokens are added to match the fixed length of 384 during the training. But we don't want to add the PAD tokens during the inference, as it will significantly increase the inference time. Maybe someone can suggest to use PAD=0 and no masking during training + no padding during inference, this method drops accuracy significantly due to the train-test time inconsistency. So the only choice for us is to mask the PAD part of the inputs during training so that they do not contribute to the calculation of the gradients.</p>\n<p>TransformerBlock is originally designed to support variable length inputs, as they use 'attention_mask' when training with PAD tokens, and making it invisible during training. So yes, Transformer block also supports Masking in my implementation, which is trivial.</p>\n<p>But the 'masking by causal padding' in the Conv1DBlock I mentioned is bit more trickier to understand. Conv1DLayer does not internally supports masking, as when they used with kernel size &gt; 1, the PAD token will be calculated  together in the kernel with neighboring frames. So if we use typical padding='same' argument, the PAD tokens will intrude into the original input sequence's time frames, and makes it much harder to implement masking. On the other hand, the causal padding prevents kernels to see the future frames, and it always preserves the time frame information if strides=1, thus we do not need to additionally handle the position of the mask which is originally generated by the initial tf.keras.layers.Masking layer.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fc658c6591b0eec8bbd97400ddf9d9e7d%2F2023-05-04%20%205.14.16.png?generation=1683144873769874&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Masking of PAD is literally for the masking of PAD tokens in the sequence. for example, if the data have the length of 256, 128 PAD tokens are added to match the fixed length of 384 during the training. But we don't want to add the PAD tokens during the inference, as it will significantly increase the inference time. Maybe someone can suggest to use PAD=0 and no masking during training + no padding during inference, this method drops accuracy significantly due to the train-test time inconsistency. So the only choice for us is to mask the PAD part of the inputs during training so that they do not contribute to the calculation of the gradients.\n\nTransformerBlock is originally designed to support variable length inputs, as they use 'attention_mask' when training with PAD tokens, and making it invisible during training. So yes, Transformer block also supports Masking in my implementation, which is trivial.\n\nBut the 'masking by causal padding' in the Conv1DBlock I mentioned is bit more trickier to understand. Conv1DLayer does not internally supports masking, as when they used with kernel size > 1, the PAD token will be calculated  together in the kernel with neighboring frames. So if we use typical padding='same' argument, the PAD tokens will intrude into the original input sequence's time frames, and makes it much harder to implement masking. On the other hand, the causal padding prevents kernels to see the future frames, and it always preserves the time frame information if strides=1, thus we do not need to additionally handle the position of the mask which is originally generated by the initial tf.keras.layers.Masking layer.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fc658c6591b0eec8bbd97400ddf9d9e7d%2F2023-05-04%20%205.14.16.png?generation=1683144873769874&alt=media)\n",
          "votes": 21
        }
      ]
    },
    {
      "id": 2244476,
      "postDate": "2023-05-03T16:41:11.870Z",
      "content": "<p>Congrats on 1st place solo cash gold finish :)</p>",
      "rawMarkdown": "Congrats on 1st place solo cash gold finish :)",
      "votes": 1
    },
    {
      "id": 2244154,
      "postDate": "2023-05-03T13:28:16.837Z",
      "content": "<p>Congratulations! Good job!</p>",
      "rawMarkdown": "Congratulations! Good job!",
      "votes": 1,
      "replies": [
        {
          "id": 2244194,
          "postDate": "2023-05-03T13:56:49.133Z",
          "content": "<p>congratulations for you too! thanks a lot! :)</p>",
          "rawMarkdown": "congratulations for you too! thanks a lot! :)"
        }
      ]
    },
    {
      "id": 2244092,
      "postDate": "2023-05-03T12:38:50.443Z",
      "content": "<p>Congratulations with solo win! Great solution<br>\nDo you plan to post your submission preparing code with model's weights? I'm really curious about trying 1d CNN + 2d CNN approach</p>\n<p>Also have you tried to compare ensemble with and without softmax?</p>",
      "rawMarkdown": "Congratulations with solo win! Great solution\nDo you plan to post your submission preparing code with model's weights? I'm really curious about trying 1d CNN + 2d CNN approach\n\nAlso have you tried to compare ensemble with and without softmax?",
      "votes": 1,
      "replies": [
        {
          "id": 2244206,
          "postDate": "2023-05-03T14:04:01.387Z",
          "content": "<p>Im cleaning up my code right now… but right now im not sure which part of my code to be available on the public since it contains bunch of things included in my other projects. </p>\n<p>about the softmax, yes I compared it on public LB and without softmax was better so I used it, but it turned out that with softmax was better on private LB and their difference was negligible(around +-0.0004).</p>\n<p>thank you and it was an honor for me to compete with you and your team!:)</p>",
          "rawMarkdown": "Im cleaning up my code right now… but right now im not sure which part of my code to be available on the public since it contains bunch of things included in my other projects. \n\nabout the softmax, yes I compared it on public LB and without softmax was better so I used it, but it turned out that with softmax was better on private LB and their difference was negligible(around +-0.0004).\n\nthank you and it was an honor for me to compete with you and your team!:)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2254519,
      "postDate": "2023-05-11T03:57:40.500Z",
      "content": "<p>you ate it solo, my congrats</p>",
      "rawMarkdown": "you ate it solo, my congrats",
      "votes": 2
    },
    {
      "id": 2253855,
      "postDate": "2023-05-10T13:40:40.653Z",
      "content": "<p>wow thats great…</p>",
      "rawMarkdown": "wow thats great...",
      "votes": 2
    },
    {
      "id": 2245711,
      "postDate": "2023-05-04T14:36:59.383Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": 2
    },
    {
      "id": 2245495,
      "postDate": "2023-05-04T12:07:45.653Z",
      "content": "<p>Congratulations with solo win! I was glad to see your solution. Thanks!</p>",
      "rawMarkdown": "Congratulations with solo win! I was glad to see your solution. Thanks!",
      "votes": 2
    },
    {
      "id": 2245172,
      "postDate": "2023-05-04T07:41:17.980Z",
      "content": "<p>Congratulation for the 1st prize. This writeup is really impressive.</p>\n<p>Now, I am trying to reproduce your pipeline in my environment, but mask tensor fails to propagate when saving to tflite although original model can produce valid output. </p>\n<p>How did you convert your tf.keras.Model into TFLite model?</p>\n<p>Error message</p>\n<pre><code>in user code:\n\n    File \"/ml-tf/working/kaggle-ISLR/train/model_v41.py\", line 43, in call  *\n        mask = tf.cast(mask, tf.bool)\n\n    ValueError: None values not supported.\n</code></pre>\n<p>c.f. My failed code is below:</p>\n<pre><code>             (tf.Module):\n                 ():\n                    ().__init__()\n\n                    \n                    self.preprocess_layer = preprocess_layer\n                    self.model = model\n\n\n                 ():\n                    \n                    x, non_empty_frame_idxs = self.preprocess_layer(x)\n                    \n                    x = tf.expand_dims(x, axis=)\n                    non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=)\n                    dummy_labels = tf.zeros((x.shape[]), dtype=tf.int32)\n                    dummy_z_preds = tf.zeros(\n                        (\n                            x.shape[],\n                            cfg.input_size,\n                            (cfg.left_hand_idxs) + (cfg.right_hand_idxs),\n                            ,\n                        ),\n                        dtype=tf.float32,\n                    )\n                    outputs = model(\n                        {\n                            : x,\n                            : non_empty_frame_idxs,\n                            : dummy_labels,\n                            : dummy_z_preds,\n                        }\n                    )\n                    outputs = tf.squeeze(outputs, axis=)\n                     {: outputs}\n\n            \n            submit_model = TestSubmitModel()\n            \n            seq_idx = \n            raw_data = load_relevant_data_subset(train.iloc[seq_idx][])\n            ()\n            demo_output = submit_model(raw_data)[]\n            ()\n            demo_prediction = demo_output.numpy().argmax()\n            (\n                \n            )\n\n            tf.saved_model.save(submit_model, )\n            converter = tf.lite.TFLiteConverter.from_saved_model()  \n            tflite_model = converter.convert()\n</code></pre>",
      "rawMarkdown": "Congratulation for the 1st prize. This writeup is really impressive.\n\nNow, I am trying to reproduce your pipeline in my environment, but mask tensor fails to propagate when saving to tflite although original model can produce valid output. \n\nHow did you convert your tf.keras.Model into TFLite model?\n\nError message\n```\nin user code:\n\n    File \"/ml-tf/working/kaggle-ISLR/train/model_v41.py\", line 43, in call  *\n        mask = tf.cast(mask, tf.bool)\n\n    ValueError: None values not supported.\n```\n\n\nc.f. My failed code is below:\n\n```python\n            class TestSubmitModel(tf.Module):\n                def __init__(self):\n                    super().__init__()\n\n                    # Load the feature generation and main models\n                    self.preprocess_layer = preprocess_layer\n                    self.model = model\n\n                @tf.function(\n                    input_signature=[\n                        tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name=\"inputs\")\n                    ]\n                )\n                def __call__(self, x):\n                    # Preprocess Data\n                    x, non_empty_frame_idxs = self.preprocess_layer(x)\n                    # Return a dictionary with the output tensor\n                    x = tf.expand_dims(x, axis=0)\n                    non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=0)\n                    dummy_labels = tf.zeros((x.shape[0]), dtype=tf.int32)\n                    dummy_z_preds = tf.zeros(\n                        (\n                            x.shape[0],\n                            cfg.input_size,\n                            len(cfg.left_hand_idxs) + len(cfg.right_hand_idxs),\n                            1,\n                        ),\n                        dtype=tf.float32,\n                    )\n                    outputs = model(\n                        {\n                            \"frames\": x,\n                            \"non_empty_frame_idxs\": non_empty_frame_idxs,\n                            \"labels\": dummy_labels,\n                            \"z_preds\": dummy_z_preds,\n                        }\n                    )\n                    outputs = tf.squeeze(outputs, axis=0)\n                    return {\"outputs\": outputs}\n\n            # Define TF Lite Model\n            submit_model = TestSubmitModel()\n            # Sanity Check\n            seq_idx = 25\n            raw_data = load_relevant_data_subset(train.iloc[seq_idx][\"file_path\"])\n            print(f\"demo_raw_data shape: {raw_data.shape}, dtype: {raw_data.dtype}\")\n            demo_output = submit_model(raw_data)[\"outputs\"]\n            print(f\"demo_output shape: {demo_output.shape}, dtype: {demo_output.dtype}\")\n            demo_prediction = demo_output.numpy().argmax()\n            print(\n                f'demo_prediction: {demo_prediction}, correct: {train.iloc[seq_idx][\"sign_ord\"]}'\n            )\n\n            tf.saved_model.save(submit_model, \"submit_model\")\n            converter = tf.lite.TFLiteConverter.from_saved_model(\"submit_model\")  # <- failed at this line\n            tflite_model = converter.convert()\n```",
      "votes": 2,
      "replies": [
        {
          "id": 2245898,
          "postDate": "2023-05-04T16:53:51.037Z",
          "content": "<p>I released my inference notebook. check it if you want! :)</p>",
          "rawMarkdown": "I released my inference notebook. check it if you want! :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2775323,
      "postDate": "2024-04-25T16:09:03.527Z",
      "content": "<p>I love how you explained your thought process through your code and explanation. Somehow I understand them, I just need to understand the concepts you mentioned. That was lovely and I look forward to seeing more.</p>",
      "rawMarkdown": "I love how you explained your thought process through your code and explanation. Somehow I understand them, I just need to understand the concepts you mentioned. That was lovely and I look forward to seeing more."
    },
    {
      "id": 2419416,
      "postDate": "2023-09-02T02:40:05.687Z",
      "content": "<p>I think your approach is very inspiring! Especially the way you see the 1D CNN modules as a \"trainable tokenizer\".   That broaden my perspective.<br>\nBy the way, can I ask you some questions about the <strong>input layer</strong> you implemented?</p>\n<ol>\n<li>What's the reason behind setting the input layer \"<strong>CHANNEL</strong>\" size as \"<strong>6x118</strong>\"? not <strong>3x118(xyz per landmarks)</strong></li>\n<li>And, why you set the <strong>max_len</strong> as 64? Is it because althoguh there were various lengths of data, max_len 64 can cover majority of your input data?</li>\n<li>If 2. is true, did you also thought about the information loss of long sequence data?</li>\n</ol>\n<p>I will be very pleased if you answer my questions :)</p>",
      "rawMarkdown": "I think your approach is very inspiring! Especially the way you see the 1D CNN modules as a \"trainable tokenizer\".   That broaden my perspective.\nBy the way, can I ask you some questions about the **input layer** you implemented?\n\n1. What's the reason behind setting the input layer \"**CHANNEL**\" size as \"**6x118**\"? not **3x118(xyz per landmarks)**\n2. And, why you set the **max_len** as 64? Is it because althoguh there were various lengths of data, max_len 64 can cover majority of your input data?\n3. If 2. is true, did you also thought about the information loss of long sequence data?\n\nI will be very pleased if you answer my questions :)"
    },
    {
      "id": 2410514,
      "postDate": "2023-08-27T03:31:11.110Z",
      "content": "<p>Thank you for sharing this  crucial information . </p>",
      "rawMarkdown": "Thank you for sharing this  crucial information . "
    },
    {
      "id": 2391619,
      "postDate": "2023-08-15T08:49:22.560Z",
      "content": "<p>English is not my mother tongue. Could you explain what <strong>4x seed</strong> is?🥺<br>\nIs that mean CFG.seed = 42 or 43 or 44 or 45?</p>",
      "rawMarkdown": "English is not my mother tongue. Could you explain what **4x seed** is?🥺\nIs that mean CFG.seed = 42 or 43 or 44 or 45?",
      "replies": [
        {
          "id": 2416301,
          "postDate": "2023-08-30T23:37:50.220Z",
          "content": "<p>In hoyso48's notebook(shared code),<br>\n\" you should run this notebook four time (for each seed=42,43,44,45) to get all 4 seed weights of the model. \"<br>\nSo I think you got the point.</p>",
          "rawMarkdown": "In hoyso48's notebook(shared code),\n\" you should run this notebook four time (for each seed=42,43,44,45) to get all 4 seed weights of the model. \"\nSo I think you got the point."
        }
      ]
    },
    {
      "id": 2371367,
      "postDate": "2023-08-03T05:07:21.650Z",
      "content": "<p>Kudos to you!</p>",
      "rawMarkdown": "Kudos to you!"
    },
    {
      "id": 2363187,
      "postDate": "2023-07-28T14:53:09.607Z",
      "content": "<p>Thank you for sharing your solution. Your notes and explanations are very helpful.</p>",
      "rawMarkdown": "Thank you for sharing your solution. Your notes and explanations are very helpful."
    },
    {
      "id": 2350628,
      "postDate": "2023-07-19T11:11:08.963Z",
      "content": "<p>this is great solution!</p>",
      "rawMarkdown": "this is great solution!"
    },
    {
      "id": 2272709,
      "postDate": "2023-05-24T17:12:36.363Z",
      "content": "<p>Greetings, <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> !</p>\n<p>Amazing solution, congratulations!</p>\n<p>Have you already looked at the latest <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling\" target=\"_blank\">Google ASL competition</a>? Eager to see how will you solve this task… Keep up with the great work💪🔥</p>",
      "rawMarkdown": "Greetings, @hoyso48 !\n\nAmazing solution, congratulations!\n\nHave you already looked at the latest [Google ASL competition](https://www.kaggle.com/competitions/asl-fingerspelling)? Eager to see how will you solve this task... Keep up with the great work💪🔥"
    },
    {
      "id": 2259317,
      "postDate": "2023-05-14T21:20:08.137Z",
      "content": "<p>Congratulations! <br>\nThank you so much for sharing :) </p>",
      "rawMarkdown": "Congratulations! \nThank you so much for sharing :) "
    },
    {
      "id": 2257600,
      "postDate": "2023-05-13T14:03:22.933Z",
      "content": "<p>Kudos <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> ! </p>",
      "rawMarkdown": "Kudos @hoyso48 ! "
    },
    {
      "id": 2255463,
      "postDate": "2023-05-11T18:20:21.473Z",
      "content": "<p>Thats a great solution</p>",
      "rawMarkdown": "Thats a great solution"
    },
    {
      "id": 2251549,
      "postDate": "2023-05-09T13:11:01.477Z",
      "content": "<p>Congrats!!! Thanks for sharing your solutions.</p>",
      "rawMarkdown": "Congrats!!! Thanks for sharing your solutions."
    },
    {
      "id": 2249772,
      "postDate": "2023-05-08T05:29:01.843Z",
      "content": "<p>👍👏🎉 Congrats on winning the competition with your innovative approach! It's great to see how you combined a 1D CNN and a Transformer to achieve such impressive results. Your explanation of the model architecture and the techniques you used for handling variable-length inputs and regularization are also very insightful. Thanks for sharing your experience!</p>",
      "rawMarkdown": "👍👏🎉 Congrats on winning the competition with your innovative approach! It's great to see how you combined a 1D CNN and a Transformer to achieve such impressive results. Your explanation of the model architecture and the techniques you used for handling variable-length inputs and regularization are also very insightful. Thanks for sharing your experience!"
    },
    {
      "id": 2249735,
      "postDate": "2023-05-08T04:48:24.313Z",
      "content": "<p>That's an amazing solution!</p>",
      "rawMarkdown": "That's an amazing solution!"
    },
    {
      "id": 2249134,
      "postDate": "2023-05-07T13:31:20.510Z",
      "content": "<p>congratulations for excellent work@</p>",
      "rawMarkdown": "congratulations for excellent work@"
    },
    {
      "id": 2248649,
      "postDate": "2023-05-07T04:26:01.033Z",
      "content": "<p>Congratulations !!!</p>",
      "rawMarkdown": "Congratulations !!!"
    },
    {
      "id": 2247710,
      "postDate": "2023-05-06T08:20:11.507Z",
      "content": "<p>Glad to know you using a CNN model and better than transformer :)</p>",
      "rawMarkdown": "Glad to know you using a CNN model and better than transformer :)"
    },
    {
      "id": 2247618,
      "postDate": "2023-05-06T06:58:02.497Z",
      "content": "<p>congratulations to you! <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> </p>",
      "rawMarkdown": "congratulations to you! @hoyso48 "
    },
    {
      "id": 2247507,
      "postDate": "2023-05-06T04:36:55.447Z",
      "content": "<p>Thanks for sharing a great solution.</p>\n<p>Would I ask one question?</p>\n<blockquote>\n  <p>if there is a strong inter-frame correlation, 1D CNNs would be more efficient than Transformers</p>\n</blockquote>\n<p>How did you come up with above hypothesis? I would like to know the rationale behind the hypothesis.</p>",
      "rawMarkdown": "Thanks for sharing a great solution.\n\nWould I ask one question?\n\n>if there is a strong inter-frame correlation, 1D CNNs would be more efficient than Transformers\n\nHow did you come up with above hypothesis? I would like to know the rationale behind the hypothesis."
    },
    {
      "id": 2247456,
      "postDate": "2023-05-06T02:55:04.803Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations"
    },
    {
      "id": 2247421,
      "postDate": "2023-05-06T01:59:14.987Z",
      "content": "<p>Congratulations. Thanks for posting detailed code with documentation and explanation.</p>",
      "rawMarkdown": "Congratulations. Thanks for posting detailed code with documentation and explanation."
    },
    {
      "id": 2247406,
      "postDate": "2023-05-06T01:23:59.273Z",
      "content": "<p>Congratulations! Interesting solution.</p>",
      "rawMarkdown": "Congratulations! Interesting solution."
    },
    {
      "id": 2247277,
      "postDate": "2023-05-05T20:52:15.830Z",
      "content": "<p>A very interesting solution, but it was not clear to me if the length of each record was normalized or they were worked as irregular records.</p>",
      "rawMarkdown": "A very interesting solution, but it was not clear to me if the length of each record was normalized or they were worked as irregular records."
    },
    {
      "id": 2247203,
      "postDate": "2023-05-05T19:51:34.780Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>, great job!</p>",
      "rawMarkdown": "Congrats @hoyso48, great job!"
    },
    {
      "id": 2246881,
      "postDate": "2023-05-05T14:29:56.107Z",
      "content": "<p>Congratulations on your win, Tensorflow is indeed a great library, although I prefer Pytorch</p>",
      "rawMarkdown": "Congratulations on your win, Tensorflow is indeed a great library, although I prefer Pytorch"
    },
    {
      "id": 2246242,
      "postDate": "2023-05-05T02:28:37.657Z",
      "content": "<p>Congratulations with solo win! I was glad to see your solution. Thanks!</p>",
      "rawMarkdown": "Congratulations with solo win! I was glad to see your solution. Thanks!"
    },
    {
      "id": 2246146,
      "postDate": "2023-05-04T23:21:14.943Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!"
    },
    {
      "id": 2246135,
      "postDate": "2023-05-04T22:56:09.033Z",
      "content": "<p>Congratulations great work🎉</p>",
      "rawMarkdown": "Congratulations great work🎉"
    },
    {
      "id": 2356409,
      "postDate": "2023-07-24T07:20:19.380Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2258742,
      "postDate": "2023-05-14T12:54:10.883Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2245152,
      "author_name": "huyidao",
      "author_url": "",
      "post_date": "2023-05-04T07:24:13.913000",
      "content": "<p>amazing solution. congrats and thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2244928,
      "author_name": "Muchamad Arif Hana Sasono",
      "author_url": "",
      "post_date": "2023-05-04T02:45:16.953000",
      "content": "<p>Thank you for sharing. can the code be implemented on EEG signals for classification?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2245901,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-05-04T16:54:30.433000",
          "content": "<p>I believe yes, with some modifications.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2244632,
      "author_name": "Kshitiz Kumar",
      "author_url": "",
      "post_date": "2023-05-03T19:08:02.177000",
      "content": "<p>Many congratulations. Thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2244625,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-05-03T18:59:22.287000",
      "content": "<p>Really cool solution (Cool approach in general, there weren't a lot of conv + transformer architectures to this point!).</p>\n<p>Quick question, Is the masking <code>PAD</code> for ignoring gradients computed for shorter sequences in the batch? <br>\n[I assume the causal masking is inside the transformer block itself, am I correct?]</p>\n<p>Congrats, and thank you again for sharing! <br>\nAmazing work. I am a bit shocked by how simple it is.. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244693,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-05-03T20:18:15.437000",
          "content": "<p>Masking of PAD is literally for the masking of PAD tokens in the sequence. for example, if the data have the length of 256, 128 PAD tokens are added to match the fixed length of 384 during the training. But we don't want to add the PAD tokens during the inference, as it will significantly increase the inference time. Maybe someone can suggest to use PAD=0 and no masking during training + no padding during inference, this method drops accuracy significantly due to the train-test time inconsistency. So the only choice for us is to mask the PAD part of the inputs during training so that they do not contribute to the calculation of the gradients.</p>\n<p>TransformerBlock is originally designed to support variable length inputs, as they use 'attention_mask' when training with PAD tokens, and making it invisible during training. So yes, Transformer block also supports Masking in my implementation, which is trivial.</p>\n<p>But the 'masking by causal padding' in the Conv1DBlock I mentioned is bit more trickier to understand. Conv1DLayer does not internally supports masking, as when they used with kernel size &gt; 1, the PAD token will be calculated  together in the kernel with neighboring frames. So if we use typical padding='same' argument, the PAD tokens will intrude into the original input sequence's time frames, and makes it much harder to implement masking. On the other hand, the causal padding prevents kernels to see the future frames, and it always preserves the time frame information if strides=1, thus we do not need to additionally handle the position of the mask which is originally generated by the initial tf.keras.layers.Masking layer.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Fc658c6591b0eec8bbd97400ddf9d9e7d%2F2023-05-04%20%205.14.16.png?generation=1683144873769874&amp;alt=media\" alt=\"\"></p>",
          "votes": 21,
          "replies": []
        }
      ]
    },
    {
      "id": 2244476,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2023-05-03T16:41:11.870000",
      "content": "<p>Congrats on 1st place solo cash gold finish :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2244154,
      "author_name": "Artem Toporov",
      "author_url": "",
      "post_date": "2023-05-03T13:28:16.837000",
      "content": "<p>Congratulations! Good job!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244194,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-05-03T13:56:49.133000",
          "content": "<p>congratulations for you too! thanks a lot! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2244092,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-05-03T12:38:50.443000",
      "content": "<p>Congratulations with solo win! Great solution<br>\nDo you plan to post your submission preparing code with model's weights? I'm really curious about trying 1d CNN + 2d CNN approach</p>\n<p>Also have you tried to compare ensemble with and without softmax?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244206,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-05-03T14:04:01.387000",
          "content": "<p>Im cleaning up my code right now… but right now im not sure which part of my code to be available on the public since it contains bunch of things included in my other projects. </p>\n<p>about the softmax, yes I compared it on public LB and without softmax was better so I used it, but it turned out that with softmax was better on private LB and their difference was negligible(around +-0.0004).</p>\n<p>thank you and it was an honor for me to compete with you and your team!:)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2254519,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-11T03:57:40.500000",
      "content": "<p>you ate it solo, my congrats</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2253855,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-05-10T13:40:40.653000",
      "content": "<p>wow thats great…</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2245711,
      "author_name": "Jaydev",
      "author_url": "",
      "post_date": "2023-05-04T14:36:59.383000",
      "content": "<p>congratulations</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2245495,
      "author_name": "Gleb Papchikhin",
      "author_url": "",
      "post_date": "2023-05-04T12:07:45.653000",
      "content": "<p>Congratulations with solo win! I was glad to see your solution. Thanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2245172,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2023-05-04T07:41:17.980000",
      "content": "<p>Congratulation for the 1st prize. This writeup is really impressive.</p>\n<p>Now, I am trying to reproduce your pipeline in my environment, but mask tensor fails to propagate when saving to tflite although original model can produce valid output. </p>\n<p>How did you convert your tf.keras.Model into TFLite model?</p>\n<p>Error message</p>\n<pre><code>in user code:\n\n    File \"/ml-tf/working/kaggle-ISLR/train/model_v41.py\", line 43, in call  *\n        mask = tf.cast(mask, tf.bool)\n\n    ValueError: None values not supported.\n</code></pre>\n<p>c.f. My failed code is below:</p>\n<pre><code>             (tf.Module):\n                 ():\n                    ().__init__()\n\n                    \n                    self.preprocess_layer = preprocess_layer\n                    self.model = model\n\n\n                 ():\n                    \n                    x, non_empty_frame_idxs = self.preprocess_layer(x)\n                    \n                    x = tf.expand_dims(x, axis=)\n                    non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=)\n                    dummy_labels = tf.zeros((x.shape[]), dtype=tf.int32)\n                    dummy_z_preds = tf.zeros(\n                        (\n                            x.shape[],\n                            cfg.input_size,\n                            (cfg.left_hand_idxs) + (cfg.right_hand_idxs),\n                            ,\n                        ),\n                        dtype=tf.float32,\n                    )\n                    outputs = model(\n                        {\n                            : x,\n                            : non_empty_frame_idxs,\n                            : dummy_labels,\n                            : dummy_z_preds,\n                        }\n                    )\n                    outputs = tf.squeeze(outputs, axis=)\n                     {: outputs}\n\n            \n            submit_model = TestSubmitModel()\n            \n            seq_idx = \n            raw_data = load_relevant_data_subset(train.iloc[seq_idx][])\n            ()\n            demo_output = submit_model(raw_data)[]\n            ()\n            demo_prediction = demo_output.numpy().argmax()\n            (\n                \n            )\n\n            tf.saved_model.save(submit_model, )\n            converter = tf.lite.TFLiteConverter.from_saved_model()  \n            tflite_model = converter.convert()\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2245898,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "2023-05-04T16:53:51.037000",
          "content": "<p>I released my inference notebook. check it if you want! :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2775323,
      "author_name": "Drishti B.📚📝",
      "author_url": "",
      "post_date": "2024-04-25T16:09:03.527000",
      "content": "<p>I love how you explained your thought process through your code and explanation. Somehow I understand them, I just need to understand the concepts you mentioned. That was lovely and I look forward to seeing more.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2419416,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-02T02:40:05.687000",
      "content": "<p>I think your approach is very inspiring! Especially the way you see the 1D CNN modules as a \"trainable tokenizer\".   That broaden my perspective.<br>\nBy the way, can I ask you some questions about the <strong>input layer</strong> you implemented?</p>\n<ol>\n<li>What's the reason behind setting the input layer \"<strong>CHANNEL</strong>\" size as \"<strong>6x118</strong>\"? not <strong>3x118(xyz per landmarks)</strong></li>\n<li>And, why you set the <strong>max_len</strong> as 64? Is it because althoguh there were various lengths of data, max_len 64 can cover majority of your input data?</li>\n<li>If 2. is true, did you also thought about the information loss of long sequence data?</li>\n</ol>\n<p>I will be very pleased if you answer my questions :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2410514,
      "author_name": "Mariam Akter",
      "author_url": "",
      "post_date": "2023-08-27T03:31:11.110000",
      "content": "<p>Thank you for sharing this  crucial information . </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2391619,
      "author_name": "Tim Liu",
      "author_url": "",
      "post_date": "2023-08-15T08:49:22.560000",
      "content": "<p>English is not my mother tongue. Could you explain what <strong>4x seed</strong> is?🥺<br>\nIs that mean CFG.seed = 42 or 43 or 44 or 45?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2416301,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-30T23:37:50.220000",
          "content": "<p>In hoyso48's notebook(shared code),<br>\n\" you should run this notebook four time (for each seed=42,43,44,45) to get all 4 seed weights of the model. \"<br>\nSo I think you got the point.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2371367,
      "author_name": "Ardian Bramantyo",
      "author_url": "",
      "post_date": "2023-08-03T05:07:21.650000",
      "content": "<p>Kudos to you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2363187,
      "author_name": "Davide Boeker",
      "author_url": "",
      "post_date": "2023-07-28T14:53:09.607000",
      "content": "<p>Thank you for sharing your solution. Your notes and explanations are very helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2350628,
      "author_name": "Skinny Lich",
      "author_url": "",
      "post_date": "2023-07-19T11:11:08.963000",
      "content": "<p>this is great solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2272709,
      "author_name": "Vladimir Simões da Luz Junior",
      "author_url": "",
      "post_date": "2023-05-24T17:12:36.363000",
      "content": "<p>Greetings, <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> !</p>\n<p>Amazing solution, congratulations!</p>\n<p>Have you already looked at the latest <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling\" target=\"_blank\">Google ASL competition</a>? Eager to see how will you solve this task… Keep up with the great work💪🔥</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2259317,
      "author_name": "Rasha Salim",
      "author_url": "",
      "post_date": "2023-05-14T21:20:08.137000",
      "content": "<p>Congratulations! <br>\nThank you so much for sharing :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2257600,
      "author_name": "Suraj",
      "author_url": "",
      "post_date": "2023-05-13T14:03:22.933000",
      "content": "<p>Kudos <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> ! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2255463,
      "author_name": "apranav_19",
      "author_url": "",
      "post_date": "2023-05-11T18:20:21.473000",
      "content": "<p>Thats a great solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2251549,
      "author_name": "june",
      "author_url": "",
      "post_date": "2023-05-09T13:11:01.477000",
      "content": "<p>Congrats!!! Thanks for sharing your solutions.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2249772,
      "author_name": "Arslan3x5",
      "author_url": "",
      "post_date": "2023-05-08T05:29:01.843000",
      "content": "<p>👍👏🎉 Congrats on winning the competition with your innovative approach! It's great to see how you combined a 1D CNN and a Transformer to achieve such impressive results. Your explanation of the model architecture and the techniques you used for handling variable-length inputs and regularization are also very insightful. Thanks for sharing your experience!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2249735,
      "author_name": "Harshiv Chandra",
      "author_url": "",
      "post_date": "2023-05-08T04:48:24.313000",
      "content": "<p>That's an amazing solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2249134,
      "author_name": "Mariam Akter",
      "author_url": "",
      "post_date": "2023-05-07T13:31:20.510000",
      "content": "<p>congratulations for excellent work@</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2248649,
      "author_name": "Juzar Kagdi",
      "author_url": "",
      "post_date": "2023-05-07T04:26:01.033000",
      "content": "<p>Congratulations !!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247710,
      "author_name": "Dewei Chen",
      "author_url": "",
      "post_date": "2023-05-06T08:20:11.507000",
      "content": "<p>Glad to know you using a CNN model and better than transformer :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247618,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-06T06:58:02.497000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247507,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-06T04:36:55.447000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247456,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-06T02:55:04.803000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247421,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-06T01:59:14.987000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247406,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-06T01:23:59.273000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247277,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-05T20:52:15.830000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2247203,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-05T19:51:34.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246881,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-05T14:29:56.107000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246242,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-05T02:28:37.657000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246146,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-04T23:21:14.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246135,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-04T22:56:09.033000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2356409,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-24T07:20:19.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2258742,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-14T12:54:10.883000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2244052": "First of all, I would like to express my gratitude to the Google for hosting this amazing competition. I have always been a big fan of the services, frameworks, and platforms provided by Google.(Colab, GCP, TensorFlow.. all of them are amazing). Without access to these offerings from Google, I wouldn't have been able to win this competition. I would also like to thank all the other participants who shared their ideas. In particular, I gained valuable insights from the ideas shared by @hengck23.\n\n# TL;DR\nMy solution involved a combination of a 1D CNN and a Transformer, trained from scratch using all training data(competition data only), and used 4x seed ensemble for submission. and I initially started with PyTorch + GPU but later switched to TensorFlow + Colab TPU(tpuv2-8) to ensure compatibility with TensorFlow Lite.\n\n# 1D CNN vs. Transformer?\nMy hypothesis was that in modeling sequential data, \n**if there is a strong inter-frame correlation, 1D CNNs would be more efficient than Transformers.**\n\nAs in my experiment, the pure 1D CNN easily outperformed the Transformer and as a result, I was able to achieve a public LB score of 0.80 using only the 1D CNN at the end. \n\nHowever, there still were roles for the Transformer, which could be used on top of the 1D CNN(we can view 1d cnn as some kind of trainable tokenizer).\n\n# Model\n```python\ndef get_model(max_len=64, dropout_step=0, dim=192):\n    inp = tf.keras.Input((max_len,CHANNELS))\n    x = tf.keras.layers.Masking(mask_value=PAD,input_shape=(max_len,CHANNELS))(inp)\n    ksize = 17\n    x = tf.keras.layers.Dense(dim, use_bias=False,name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(momentum=0.95,name='stem_bn')(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    if dim == 384: #for the 4x sized model\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = TransformerBlock(dim,expand=2)(x)\n\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = Conv1DBlock(dim,ksize,drop_rate=0.2)(x)\n        x = TransformerBlock(dim,expand=2)(x)\n\n    x = tf.keras.layers.Dense(dim*2,activation=None,name='top_conv')(x)\n    x = tf.keras.layers.GlobalAveragePooling1D()(x)\n    x = LateDropout(0.8, start_step=dropout_step)(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES,name='classifier')(x)\n    return tf.keras.Model(inp, x)\n\n```\nCombining CNN and Transformer is a prevalent idea in recent state-of-the-art models(coatnet, conformer, Maxvit, nextvit…).  I started with an 192d 8-layer 1D CNN, then switched to a 192d (3+1)x2 conv-transformer structure, which yielded a +0.01 CV and LB improvement. \n\n\nThe 1D CNN model employed depthwise convolution and causal padding. The Transformer used BatchNorm + Swish instead of the typical LayerNorm + GELU, due to slightly(negligible) lighter inference with the same accuracy.\nsingle model has around 1.85M parameters.\n\n# Masking\n**Handling variable-length input correctly was very crucial for ensuring train-test consistency and efficient inference** , as we do not necessarily have to pad the short videos. During training, I used a max_len=384 with padding and truncation, while for inference, I only applied truncation. This approach provided sufficient inference speed and allowed the use of reasonably large models. To accurately  apply masking to the 1D CNN, I used causal padding to maintain the mask index. In TensorFlow, masking can be easily implemented using tf.keras.layers.Masking at the beginning of the model. Plus, It is essential to ensure that masking is accurately applied to operations like batch normalization and global average pooling which can be affected by masking.\n\n# Regularization\n1. Drop Path(stochastic depth, p=0.2)\n2. high rate of Dropout (p=0.8)\n3. AWP(Adversarial Weight Perturbation, with lambda = 0.2)\n###### \nAs we need to train the model from scratch, regularization technics played a significant role. I used drop_path(0.2, applied after each block), dropout(0.8, applied after GAP) and AWP(Adversarial Weight Perturbation, lambda=0.2) AWP and the dropout applied after epoch 15. All these methods were very crucial for preventing overfitting when training with long epochs(>300). All three methods had a significant impact on both CV and leaderboard scores, and removing any one of them led to noticeable performance drops.\n\n# Preprocessing\n```python\nclass Preprocess(tf.keras.layers.Layer):\n    def __init__(self, max_len=MAX_LEN, point_landmarks=POINT_LANDMARKS, **kwargs):\n        super().__init__(**kwargs)\n        self.max_len = max_len\n        self.point_landmarks = point_landmarks\n\n    def call(self, inputs):\n        if tf.rank(inputs) == 3:\n            x = inputs[None,...]\n        else:\n            x = inputs\n        \n        mean = tf_nan_mean(tf.gather(x, [17], axis=2), axis=[1,2], keepdims=True)\n        mean = tf.where(tf.math.is_nan(mean), tf.constant(0.5,x.dtype), mean)\n        x = tf.gather(x, self.point_landmarks, axis=2) #N,T,P,C\n        std = tf_nan_std(x, center=mean, axis=[1,2], keepdims=True)\n        \n        x = (x - mean)/std\n\n        if self.max_len is not None:\n            x = x[:,:self.max_len]\n        length = tf.shape(x)[1]\n        x = x[...,:2]\n\n        dx = tf.cond(tf.shape(x)[1]>1,lambda:tf.pad(x[:,1:] - x[:,:-1], [[0,0],[0,1],[0,0],[0,0]]),lambda:tf.zeros_like(x))\n\n        dx2 = tf.cond(tf.shape(x)[1]>2,lambda:tf.pad(x[:,2:] - x[:,:-2], [[0,0],[0,2],[0,0],[0,0]]),lambda:tf.zeros_like(x))\n\n        x = tf.concat([\n            tf.reshape(x, (-1,length,2*len(self.point_landmarks))),\n            tf.reshape(dx, (-1,length,2*len(self.point_landmarks))),\n            tf.reshape(dx2, (-1,length,2*len(self.point_landmarks))),\n        ], axis = -1)\n        \n        x = tf.where(tf.math.is_nan(x),tf.constant(0.,x.dtype),x)\n        \n        return x\n```\nI used left-right hand, eye, nose, and lips landmarks. For normalization, I used the 17th landmark located in the nose as a reference point, since it is usually located close to the center ([0.5, 0.5]). I used motion feature of lag1 x[:1] - x[1:], and lag2 x[:2] - x[2:](lag > 2 did not help much).\n\n#Augmentation\n\n- temporal augmentation\n1. Random resample (0.5x ~ 1.5x to original length)\n2. Random masking\n\n- Spatial augmentation\n1. hflip\n2. Random Affine(Scale, shift, rotate, shear)\n3. Random Cutout\n\n#Training\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5003978%2Ff23fef7388aa47b7a1e26f6261e6c0a4%2F.png?generation=1683114294544440&alt=media)\n- Epoch = 400\n- Lr = 5e-4 * num_replicas = 4e-3\n- Schedule = CosineDecay with no warmup\n- Optimizer = RAdam with Lookahead(better than AdamW with optimal parameters)\n- Loss = CCE with label smoothing=0.1 or just plain CCE\n###### \nfinal single model result\nCV(participant split 5fold): 0.80\npublic LB: 0.80\nprivate LB: 0.88\n###### \nTraining takes around 4 hours with colab TPUv2-8.\n\nSingle model CV was around 0.80 with participant split(5fold) at the end. I ensemble 4 different seed(with some minor differences in training configurations) for the final model and got LB 0.81. By the way, I could see slightly worse score when I submitted 4x size(384d, 16layers) single model with the same settings. I think it can achieve same or better score with better configurations.\n\n#Tried but not worked\n- GCNs\n- More Complex augmentations: augmentation based on angle and the distance between the landmarks, grid distortion on temporal, spatial axis, etc.\n- CutMix, MixUp: Not worked. Main problem was how to define new label with two different length of inputs. I could not find the correct way to implement it.\n- Knowledge Distillation: I tried to use single 4x sized model and distill it with 4x seed 4x sized model, but I could not manage it to work due to lack of time.\n###### \nI'm always amazed by the fact that we did try each other’s methods, but we came up with different result. I also gave the 2D-CNN(similar to @kolyaforrat 's brilliant solution) and pure transformer approaches a try, but due to their initial weak performance and my confidence in my own hypothesis, I didn't dig any deeper. Seeing other teams succeed with the ideas I had trouble with, through their skill and hard work, has been truly inspiring. I've learned a lot through this competition, and I'm grateful for the experience. Thank you all!!! :)\n\n\nEdit: I made my code public, check https://www.kaggle.com/competitions/asl-signs/discussion/406978",
    "2245152": "amazing solution. congrats and thanks for sharing.",
    "2244928": "Thank you for sharing. can the code be implemented on EEG signals for classification?",
    "2244632": "Many congratulations. Thank you for sharing!",
    "2244625": "Really cool solution (Cool approach in general, there weren't a lot of conv + transformer architectures to this point!).\n\nQuick question, Is the masking `PAD` for ignoring gradients computed for shorter sequences in the batch? \n[I assume the causal masking is inside the transformer block itself, am I correct?]\n\n\nCongrats, and thank you again for sharing! \nAmazing work. I am a bit shocked by how simple it is.. ",
    "2244476": "Congrats on 1st place solo cash gold finish :)",
    "2244154": "Congratulations! Good job!",
    "2244092": "Congratulations with solo win! Great solution\nDo you plan to post your submission preparing code with model's weights? I'm really curious about trying 1d CNN + 2d CNN approach\n\nAlso have you tried to compare ensemble with and without softmax?",
    "2254519": "you ate it solo, my congrats",
    "2253855": "wow thats great...",
    "2245711": "congratulations",
    "2245495": "Congratulations with solo win! I was glad to see your solution. Thanks!",
    "2245172": "Congratulation for the 1st prize. This writeup is really impressive.\n\nNow, I am trying to reproduce your pipeline in my environment, but mask tensor fails to propagate when saving to tflite although original model can produce valid output. \n\nHow did you convert your tf.keras.Model into TFLite model?\n\nError message\n```\nin user code:\n\n    File \"/ml-tf/working/kaggle-ISLR/train/model_v41.py\", line 43, in call  *\n        mask = tf.cast(mask, tf.bool)\n\n    ValueError: None values not supported.\n```\n\n\nc.f. My failed code is below:\n\n```python\n            class TestSubmitModel(tf.Module):\n                def __init__(self):\n                    super().__init__()\n\n                    # Load the feature generation and main models\n                    self.preprocess_layer = preprocess_layer\n                    self.model = model\n\n                @tf.function(\n                    input_signature=[\n                        tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name=\"inputs\")\n                    ]\n                )\n                def __call__(self, x):\n                    # Preprocess Data\n                    x, non_empty_frame_idxs = self.preprocess_layer(x)\n                    # Return a dictionary with the output tensor\n                    x = tf.expand_dims(x, axis=0)\n                    non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=0)\n                    dummy_labels = tf.zeros((x.shape[0]), dtype=tf.int32)\n                    dummy_z_preds = tf.zeros(\n                        (\n                            x.shape[0],\n                            cfg.input_size,\n                            len(cfg.left_hand_idxs) + len(cfg.right_hand_idxs),\n                            1,\n                        ),\n                        dtype=tf.float32,\n                    )\n                    outputs = model(\n                        {\n                            \"frames\": x,\n                            \"non_empty_frame_idxs\": non_empty_frame_idxs,\n                            \"labels\": dummy_labels,\n                            \"z_preds\": dummy_z_preds,\n                        }\n                    )\n                    outputs = tf.squeeze(outputs, axis=0)\n                    return {\"outputs\": outputs}\n\n            # Define TF Lite Model\n            submit_model = TestSubmitModel()\n            # Sanity Check\n            seq_idx = 25\n            raw_data = load_relevant_data_subset(train.iloc[seq_idx][\"file_path\"])\n            print(f\"demo_raw_data shape: {raw_data.shape}, dtype: {raw_data.dtype}\")\n            demo_output = submit_model(raw_data)[\"outputs\"]\n            print(f\"demo_output shape: {demo_output.shape}, dtype: {demo_output.dtype}\")\n            demo_prediction = demo_output.numpy().argmax()\n            print(\n                f'demo_prediction: {demo_prediction}, correct: {train.iloc[seq_idx][\"sign_ord\"]}'\n            )\n\n            tf.saved_model.save(submit_model, \"submit_model\")\n            converter = tf.lite.TFLiteConverter.from_saved_model(\"submit_model\")  # <- failed at this line\n            tflite_model = converter.convert()\n```",
    "2775323": "I love how you explained your thought process through your code and explanation. Somehow I understand them, I just need to understand the concepts you mentioned. That was lovely and I look forward to seeing more.",
    "2419416": "I think your approach is very inspiring! Especially the way you see the 1D CNN modules as a \"trainable tokenizer\".   That broaden my perspective.\nBy the way, can I ask you some questions about the **input layer** you implemented?\n\n1. What's the reason behind setting the input layer \"**CHANNEL**\" size as \"**6x118**\"? not **3x118(xyz per landmarks)**\n2. And, why you set the **max_len** as 64? Is it because althoguh there were various lengths of data, max_len 64 can cover majority of your input data?\n3. If 2. is true, did you also thought about the information loss of long sequence data?\n\nI will be very pleased if you answer my questions :)",
    "2410514": "Thank you for sharing this  crucial information . ",
    "2391619": "English is not my mother tongue. Could you explain what **4x seed** is?🥺\nIs that mean CFG.seed = 42 or 43 or 44 or 45?",
    "2371367": "Kudos to you!",
    "2363187": "Thank you for sharing your solution. Your notes and explanations are very helpful.",
    "2350628": "this is great solution!",
    "2272709": "Greetings, @hoyso48 !\n\nAmazing solution, congratulations!\n\nHave you already looked at the latest [Google ASL competition](https://www.kaggle.com/competitions/asl-fingerspelling)? Eager to see how will you solve this task... Keep up with the great work💪🔥",
    "2259317": "Congratulations! \nThank you so much for sharing :) ",
    "2257600": "Kudos @hoyso48 ! ",
    "2255463": "Thats a great solution",
    "2251549": "Congrats!!! Thanks for sharing your solutions.",
    "2249772": "👍👏🎉 Congrats on winning the competition with your innovative approach! It's great to see how you combined a 1D CNN and a Transformer to achieve such impressive results. Your explanation of the model architecture and the techniques you used for handling variable-length inputs and regularization are also very insightful. Thanks for sharing your experience!",
    "2249735": "That's an amazing solution!",
    "2249134": "congratulations for excellent work@",
    "2248649": "Congratulations !!!",
    "2247710": "Glad to know you using a CNN model and better than transformer :)",
    "2247618": "congratulations to you! @hoyso48 ",
    "2247507": "Thanks for sharing a great solution.\n\nWould I ask one question?\n\n>if there is a strong inter-frame correlation, 1D CNNs would be more efficient than Transformers\n\nHow did you come up with above hypothesis? I would like to know the rationale behind the hypothesis.",
    "2247456": "congratulations",
    "2247421": "Congratulations. Thanks for posting detailed code with documentation and explanation.",
    "2247406": "Congratulations! Interesting solution.",
    "2247277": "A very interesting solution, but it was not clear to me if the length of each record was normalized or they were worked as irregular records.",
    "2247203": "Congrats @hoyso48, great job!",
    "2246881": "Congratulations on your win, Tensorflow is indeed a great library, although I prefer Pytorch",
    "2246242": "Congratulations with solo win! I was glad to see your solution. Thanks!",
    "2246146": "congratulations!",
    "2246135": "Congratulations great work🎉",
    "2356409": "",
    "2258742": ""
  }
}