{
  "id": 406441,
  "title": "🥈18th Place Solution🥈 ",
  "url": "/competitions/asl-signs/discussion/406441",
  "author_name": "siwooyong",
  "post_date": "2023-05-02T11:19:29.683000",
  "votes": 17,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>TLDR</h1>\n<p>Based on my experiment, it seems that the most significant factors contributing to improved performance were <strong>regularization</strong>, <strong>feature processing</strong> and <strong>embedding_layer</strong>. The final ensemble model consisted of various input features and models, including Transformer, MLP, and GRU.<br>\n<br></p>\n<h1>Data Processing</h1>\n<p>Out of the 543 landmarks in the provided raw data, a total of 115 landmarks were used for hands, poses, and lips. Input features were constructed by concatenating xy(z), motion, and distance, resulting in a final input feature size of 1196. Initially, using only the distance feature of hand data was employed, but the significant performance improvement (CV +0.05) was achieved by adding the distance feature of pose and lip data. </p>\n<pre><code>feature = tf.concat([\n        tf.reshape(xyz_hand[:, :21, :3], [-1, 21 * 3]), \n        tf.reshape(xyz_pose[:, 21:46, :2], [-1, 25 * 2]), \n        tf.reshape(xyz_lip[:, 46:66, :2], [-1, 20 * 2]), \n        tf.reshape(motion_hand[:, :21, :3], [-1, 21 * 3]), \n        tf.reshape(motion_pose[:, 21:46, :2], [-1, 25 * 2]), \n        tf.reshape(motion_lip[:, 46:66, :2], [-1, 20 * 2]), \n        tf.reshape(distance_hand, [-1, 210]),\n        tf.reshape(distance_pose, [-1, 300]),\n        tf.reshape(distance_outlip, [-1, 190]),\n        tf.reshape(distance_inlip, [-1, 190]),\n    ], axis=-1)\n</code></pre>\n<p>Additionally, the hand with fewer NaN values was utilized, and the leg landmarks were removed from the pose landmarks. Based on this, two versions of inputs were created. The first version only utilized frames with non-NaN hand data, while the second version included frames with NaN hand data. The former had a max_length of 100, and the latter had a max_length of 200.</p>\n<pre><code>cond = lefth_sum &gt; righth_sum\nh_x = tf.where(cond, lefth_x, righth_x)\nxfeat = tf.where(cond, tf.concat([lefth_x, pose_x, lip_x], axis = 1), tf.concat([righth_x, pose_x, lip_x], axis = 1))\n</code></pre>\n<p><br></p>\n<h1>Augmentation</h1>\n<p>The hand with fewer NaN values was utilized, and both hands were flipped to be recognized as right hands in the model, which actually contributed to a performance improvement of about 0.01 in CV.</p>\n<pre><code># x-axis mirroring\ncond = lefth_sum &gt; righth_sum\nxfeat_xcoordi = xfeat[:, :, 0]\nxfeat_else = xfeat[:, :, 1:]\nxfeat_xcoordi = tf.where(cond, -xfeat_xcoordi, xfeat_xcoordi)\nxfeat = tf.concat([xfeat_xcoordi[:, :, tf.newaxis], xfeat_else], axis = -1)\n</code></pre>\n<p>I did not observe that flip, rotate, mixup, and other augmentation techniques contributed to performance improvement in CV, so I supplemented this by ensembling the models of the two versions of inputs mentioned earlier.<br>\n<br></p>\n<h1>Model</h1>\n<p>Prior to being utilized as inputs for the transformer model, the input features, namely xy(z), motion, and distance, underwent individual processing through dedicated embedding layers. Compared to the scenario where features were not processed independently, a performance improvement of 0.01 was observed in the CV score when the features were treated independently.</p>\n<pre><code># embedding layer\nxy = xy_embeddings(xy)\nmotion = motion_embeddings(motion)\ndistance_hand = distance_hand_embeddings(distance_hand)\ndistance_pose = distance_pose_embeddings(distance_pose)\ndistance_outlip = distance_outlip_embeddings(distance_outlip)\ndistance_inlip = distance_inlip_embeddings(distance_inlip)\n\nx = tf.concat([xy, motion, distance_hand, distance_pose, distance_outlip, distance_inlip], axis=-1)\nx = relu(x)\nx = fc_layer(x)\nx = TransformerModel(input_ids = None, inputs_embeds=x, attention_mask=x_mask).last_hidden_state\n</code></pre>\n<p><br></p>\n<p>For Transformer models, I used huggingface's RoBERTa-PreLayerNorm, DeBERTaV2, and GPT2. The input was processed independently for xyz, motion, and distance, and then concatenated to form a 300-dimensional transformer input. The mean, max, and std values of the Transformer output were then concatenated to obtain the final output.</p>\n<pre><code># pooling function\ndef get_pool(self, x, x_mask):\n    x = x * tf.expand_dims(x_mask, axis=-1)  # apply mask\n    nonzero_count = tf.reduce_sum(x_mask, axis=1, keepdims=True)  # count nonzero elements\n    max_discount = (1-x_mask)*1e10\n\n    apool = tf.reduce_sum(x, axis=1) / nonzero_count\n    mpool = tf.reduce_max(x - tf.expand_dims(max_discount, axis=-1), axis=1)\n    spool = tf.sqrt((tf.reduce_sum(((x - tf.expand_dims(apool, axis=1)) ** 2) * tf.expand_dims(x_mask, axis=-1), axis=1) / nonzero_count) + 1e-9)\n    return tf.concat([apool, mpool, spool], axis=-1) \n</code></pre>\n<p>In addition to the Transformer model, simple linear models and GRU models also achieved similar performance to the Transformer model, so I ensembled these three types of models.<br>\n<br></p>\n<h1>Training</h1>\n<ul>\n<li>Scheduler : lr_warmup_cosine_decay </li>\n<li>Warmup Ratio : 0.2 </li>\n<li>Optimizer : AdamW </li>\n<li>Weight Decay : 0.01</li>\n<li>Epoch : 40</li>\n<li>Learning Rate : 1e-3 </li>\n<li>Loss Function : CrossEntropyLoss </li>\n<li>Smoothing Value : 0.65 ~ 0.75</li>\n</ul>\n<p><br></p>\n<h1>Regularization</h1>\n<p>During model training, there are three primary regularization techniques that have made significant contributions to both improving convergence speed and final performance</p>\n<ul>\n<li><p>Weight normalization applied to the final linear layer(<a href=\"https://arxiv.org/abs/1602.07868\" target=\"_blank\">paper</a>)</p>\n<pre><code>final_layer = torch.nn.utils.weight_norm(nn.Linear(hidden_size, 250))\n</code></pre></li>\n<li><p>Batch normalization applied before the final linear layer</p>\n<pre><code># head \nx = fc_layer(x)\nx = batchnorm1d(x)\nx = relu(x)\nx = dropout(x)\nx = final_layer(x)\n</code></pre></li>\n<li><p>High weight decay value with the AdamW<br>\n<br></p></li>\n</ul>\n<h1>TFLite Conversion</h1>\n<p>In the early stages of the competition, I worked with PyTorch, which meant I had to deal with numerous errors when converting to TFLite, and ultimately failed to handle dynamic input shapes. In the latter stages of the competition, I began working with Keras, and the number of errors when converting to TFLite was significantly reduced.<br>\n<br></p>\n<h1>Didn't Work</h1>\n<ul>\n<li>Adding the distance feature contributed to performance improvement, but adding angle and direction features did not.</li>\n<li>Increasing the number of transformer layers did not contribute to performance improvement.</li>\n<li>I attempted to model the relationships between landmark points using Transformers or GATs, but the inference speed of the model became slower, and the performance actually decreased.</li>\n<li>Bert-like pretraining (MLM) for XYZ coordinates did not improve performance with the provided data in the competition.<br>\n<br></li>\n</ul>\n<h1>Learned</h1>\n<p>I am currently serving in the military in South Korea. During my army training, I learned a lot while coding in between. I was amazed by the culture of Kagglers sharing various ideas and discussing them. It's really cool that something bigger and better something can be born through this community, and I also want to participate in it diligently in the future. Thank you for reading my post.<br>\n<br></p>\n<h1>Code</h1>\n<p><a href=\"https://github.com/siwooyong/Google-Isolated-Sign-Language-Recognition\" target=\"_blank\">https://github.com/siwooyong/Google-Isolated-Sign-Language-Recognition</a><br>\n<br></p>",
  "messages": [
    {
      "id": 2242571,
      "postDate": "2023-05-02T11:19:29.683Z",
      "content": "<h1>TLDR</h1>\n<p>Based on my experiment, it seems that the most significant factors contributing to improved performance were <strong>regularization</strong>, <strong>feature processing</strong> and <strong>embedding_layer</strong>. The final ensemble model consisted of various input features and models, including Transformer, MLP, and GRU.<br>\n<br></p>\n<h1>Data Processing</h1>\n<p>Out of the 543 landmarks in the provided raw data, a total of 115 landmarks were used for hands, poses, and lips. Input features were constructed by concatenating xy(z), motion, and distance, resulting in a final input feature size of 1196. Initially, using only the distance feature of hand data was employed, but the significant performance improvement (CV +0.05) was achieved by adding the distance feature of pose and lip data. </p>\n<pre><code>feature = tf.concat([\n        tf.reshape(xyz_hand[:, :21, :3], [-1, 21 * 3]), \n        tf.reshape(xyz_pose[:, 21:46, :2], [-1, 25 * 2]), \n        tf.reshape(xyz_lip[:, 46:66, :2], [-1, 20 * 2]), \n        tf.reshape(motion_hand[:, :21, :3], [-1, 21 * 3]), \n        tf.reshape(motion_pose[:, 21:46, :2], [-1, 25 * 2]), \n        tf.reshape(motion_lip[:, 46:66, :2], [-1, 20 * 2]), \n        tf.reshape(distance_hand, [-1, 210]),\n        tf.reshape(distance_pose, [-1, 300]),\n        tf.reshape(distance_outlip, [-1, 190]),\n        tf.reshape(distance_inlip, [-1, 190]),\n    ], axis=-1)\n</code></pre>\n<p>Additionally, the hand with fewer NaN values was utilized, and the leg landmarks were removed from the pose landmarks. Based on this, two versions of inputs were created. The first version only utilized frames with non-NaN hand data, while the second version included frames with NaN hand data. The former had a max_length of 100, and the latter had a max_length of 200.</p>\n<pre><code>cond = lefth_sum &gt; righth_sum\nh_x = tf.where(cond, lefth_x, righth_x)\nxfeat = tf.where(cond, tf.concat([lefth_x, pose_x, lip_x], axis = 1), tf.concat([righth_x, pose_x, lip_x], axis = 1))\n</code></pre>\n<p><br></p>\n<h1>Augmentation</h1>\n<p>The hand with fewer NaN values was utilized, and both hands were flipped to be recognized as right hands in the model, which actually contributed to a performance improvement of about 0.01 in CV.</p>\n<pre><code># x-axis mirroring\ncond = lefth_sum &gt; righth_sum\nxfeat_xcoordi = xfeat[:, :, 0]\nxfeat_else = xfeat[:, :, 1:]\nxfeat_xcoordi = tf.where(cond, -xfeat_xcoordi, xfeat_xcoordi)\nxfeat = tf.concat([xfeat_xcoordi[:, :, tf.newaxis], xfeat_else], axis = -1)\n</code></pre>\n<p>I did not observe that flip, rotate, mixup, and other augmentation techniques contributed to performance improvement in CV, so I supplemented this by ensembling the models of the two versions of inputs mentioned earlier.<br>\n<br></p>\n<h1>Model</h1>\n<p>Prior to being utilized as inputs for the transformer model, the input features, namely xy(z), motion, and distance, underwent individual processing through dedicated embedding layers. Compared to the scenario where features were not processed independently, a performance improvement of 0.01 was observed in the CV score when the features were treated independently.</p>\n<pre><code># embedding layer\nxy = xy_embeddings(xy)\nmotion = motion_embeddings(motion)\ndistance_hand = distance_hand_embeddings(distance_hand)\ndistance_pose = distance_pose_embeddings(distance_pose)\ndistance_outlip = distance_outlip_embeddings(distance_outlip)\ndistance_inlip = distance_inlip_embeddings(distance_inlip)\n\nx = tf.concat([xy, motion, distance_hand, distance_pose, distance_outlip, distance_inlip], axis=-1)\nx = relu(x)\nx = fc_layer(x)\nx = TransformerModel(input_ids = None, inputs_embeds=x, attention_mask=x_mask).last_hidden_state\n</code></pre>\n<p><br></p>\n<p>For Transformer models, I used huggingface's RoBERTa-PreLayerNorm, DeBERTaV2, and GPT2. The input was processed independently for xyz, motion, and distance, and then concatenated to form a 300-dimensional transformer input. The mean, max, and std values of the Transformer output were then concatenated to obtain the final output.</p>\n<pre><code># pooling function\ndef get_pool(self, x, x_mask):\n    x = x * tf.expand_dims(x_mask, axis=-1)  # apply mask\n    nonzero_count = tf.reduce_sum(x_mask, axis=1, keepdims=True)  # count nonzero elements\n    max_discount = (1-x_mask)*1e10\n\n    apool = tf.reduce_sum(x, axis=1) / nonzero_count\n    mpool = tf.reduce_max(x - tf.expand_dims(max_discount, axis=-1), axis=1)\n    spool = tf.sqrt((tf.reduce_sum(((x - tf.expand_dims(apool, axis=1)) ** 2) * tf.expand_dims(x_mask, axis=-1), axis=1) / nonzero_count) + 1e-9)\n    return tf.concat([apool, mpool, spool], axis=-1) \n</code></pre>\n<p>In addition to the Transformer model, simple linear models and GRU models also achieved similar performance to the Transformer model, so I ensembled these three types of models.<br>\n<br></p>\n<h1>Training</h1>\n<ul>\n<li>Scheduler : lr_warmup_cosine_decay </li>\n<li>Warmup Ratio : 0.2 </li>\n<li>Optimizer : AdamW </li>\n<li>Weight Decay : 0.01</li>\n<li>Epoch : 40</li>\n<li>Learning Rate : 1e-3 </li>\n<li>Loss Function : CrossEntropyLoss </li>\n<li>Smoothing Value : 0.65 ~ 0.75</li>\n</ul>\n<p><br></p>\n<h1>Regularization</h1>\n<p>During model training, there are three primary regularization techniques that have made significant contributions to both improving convergence speed and final performance</p>\n<ul>\n<li><p>Weight normalization applied to the final linear layer(<a href=\"https://arxiv.org/abs/1602.07868\" target=\"_blank\">paper</a>)</p>\n<pre><code>final_layer = torch.nn.utils.weight_norm(nn.Linear(hidden_size, 250))\n</code></pre></li>\n<li><p>Batch normalization applied before the final linear layer</p>\n<pre><code># head \nx = fc_layer(x)\nx = batchnorm1d(x)\nx = relu(x)\nx = dropout(x)\nx = final_layer(x)\n</code></pre></li>\n<li><p>High weight decay value with the AdamW<br>\n<br></p></li>\n</ul>\n<h1>TFLite Conversion</h1>\n<p>In the early stages of the competition, I worked with PyTorch, which meant I had to deal with numerous errors when converting to TFLite, and ultimately failed to handle dynamic input shapes. In the latter stages of the competition, I began working with Keras, and the number of errors when converting to TFLite was significantly reduced.<br>\n<br></p>\n<h1>Didn't Work</h1>\n<ul>\n<li>Adding the distance feature contributed to performance improvement, but adding angle and direction features did not.</li>\n<li>Increasing the number of transformer layers did not contribute to performance improvement.</li>\n<li>I attempted to model the relationships between landmark points using Transformers or GATs, but the inference speed of the model became slower, and the performance actually decreased.</li>\n<li>Bert-like pretraining (MLM) for XYZ coordinates did not improve performance with the provided data in the competition.<br>\n<br></li>\n</ul>\n<h1>Learned</h1>\n<p>I am currently serving in the military in South Korea. During my army training, I learned a lot while coding in between. I was amazed by the culture of Kagglers sharing various ideas and discussing them. It's really cool that something bigger and better something can be born through this community, and I also want to participate in it diligently in the future. Thank you for reading my post.<br>\n<br></p>\n<h1>Code</h1>\n<p><a href=\"https://github.com/siwooyong/Google-Isolated-Sign-Language-Recognition\" target=\"_blank\">https://github.com/siwooyong/Google-Isolated-Sign-Language-Recognition</a><br>\n<br></p>",
      "rawMarkdown": "# TLDR\nBased on my experiment, it seems that the most significant factors contributing to improved performance were **regularization**, **feature processing** and **embedding_layer**. The final ensemble model consisted of various input features and models, including Transformer, MLP, and GRU.\n<br>\n\n# Data Processing\nOut of the 543 landmarks in the provided raw data, a total of 115 landmarks were used for hands, poses, and lips. Input features were constructed by concatenating xy(z), motion, and distance, resulting in a final input feature size of 1196. Initially, using only the distance feature of hand data was employed, but the significant performance improvement (CV +0.05) was achieved by adding the distance feature of pose and lip data. \n    \n    feature = tf.concat([\n            tf.reshape(xyz_hand[:, :21, :3], [-1, 21 * 3]), \n            tf.reshape(xyz_pose[:, 21:46, :2], [-1, 25 * 2]), \n            tf.reshape(xyz_lip[:, 46:66, :2], [-1, 20 * 2]), \n            tf.reshape(motion_hand[:, :21, :3], [-1, 21 * 3]), \n            tf.reshape(motion_pose[:, 21:46, :2], [-1, 25 * 2]), \n            tf.reshape(motion_lip[:, 46:66, :2], [-1, 20 * 2]), \n            tf.reshape(distance_hand, [-1, 210]),\n            tf.reshape(distance_pose, [-1, 300]),\n            tf.reshape(distance_outlip, [-1, 190]),\n            tf.reshape(distance_inlip, [-1, 190]),\n        ], axis=-1)\n\nAdditionally, the hand with fewer NaN values was utilized, and the leg landmarks were removed from the pose landmarks. Based on this, two versions of inputs were created. The first version only utilized frames with non-NaN hand data, while the second version included frames with NaN hand data. The former had a max_length of 100, and the latter had a max_length of 200.\n\n    cond = lefth_sum > righth_sum\n    h_x = tf.where(cond, lefth_x, righth_x)\n    xfeat = tf.where(cond, tf.concat([lefth_x, pose_x, lip_x], axis = 1), tf.concat([righth_x, pose_x, lip_x], axis = 1))\n\n<br>\n\n# Augmentation\nThe hand with fewer NaN values was utilized, and both hands were flipped to be recognized as right hands in the model, which actually contributed to a performance improvement of about 0.01 in CV.\n    \n    # x-axis mirroring\n    cond = lefth_sum > righth_sum\n    xfeat_xcoordi = xfeat[:, :, 0]\n    xfeat_else = xfeat[:, :, 1:]\n    xfeat_xcoordi = tf.where(cond, -xfeat_xcoordi, xfeat_xcoordi)\n    xfeat = tf.concat([xfeat_xcoordi[:, :, tf.newaxis], xfeat_else], axis = -1)\n\nI did not observe that flip, rotate, mixup, and other augmentation techniques contributed to performance improvement in CV, so I supplemented this by ensembling the models of the two versions of inputs mentioned earlier.\n<br>\n\n# Model\nPrior to being utilized as inputs for the transformer model, the input features, namely xy(z), motion, and distance, underwent individual processing through dedicated embedding layers. Compared to the scenario where features were not processed independently, a performance improvement of 0.01 was observed in the CV score when the features were treated independently.\n    \n    # embedding layer\n    xy = xy_embeddings(xy)\n    motion = motion_embeddings(motion)\n    distance_hand = distance_hand_embeddings(distance_hand)\n    distance_pose = distance_pose_embeddings(distance_pose)\n    distance_outlip = distance_outlip_embeddings(distance_outlip)\n    distance_inlip = distance_inlip_embeddings(distance_inlip)\n\n    x = tf.concat([xy, motion, distance_hand, distance_pose, distance_outlip, distance_inlip], axis=-1)\n    x = relu(x)\n    x = fc_layer(x)\n    x = TransformerModel(input_ids = None, inputs_embeds=x, attention_mask=x_mask).last_hidden_state\n<br>\n\nFor Transformer models, I used huggingface's RoBERTa-PreLayerNorm, DeBERTaV2, and GPT2. The input was processed independently for xyz, motion, and distance, and then concatenated to form a 300-dimensional transformer input. The mean, max, and std values of the Transformer output were then concatenated to obtain the final output.\n    \n    # pooling function\n    def get_pool(self, x, x_mask):\n        x = x * tf.expand_dims(x_mask, axis=-1)  # apply mask\n        nonzero_count = tf.reduce_sum(x_mask, axis=1, keepdims=True)  # count nonzero elements\n        max_discount = (1-x_mask)*1e10\n\n        apool = tf.reduce_sum(x, axis=1) / nonzero_count\n        mpool = tf.reduce_max(x - tf.expand_dims(max_discount, axis=-1), axis=1)\n        spool = tf.sqrt((tf.reduce_sum(((x - tf.expand_dims(apool, axis=1)) ** 2) * tf.expand_dims(x_mask, axis=-1), axis=1) / nonzero_count) + 1e-9)\n        return tf.concat([apool, mpool, spool], axis=-1) \n\nIn addition to the Transformer model, simple linear models and GRU models also achieved similar performance to the Transformer model, so I ensembled these three types of models.\n<br>\n\n# Training\n\n* Scheduler : lr_warmup_cosine_decay \n* Warmup Ratio : 0.2 \n* Optimizer : AdamW \n* Weight Decay : 0.01\n* Epoch : 40\n* Learning Rate : 1e-3 \n* Loss Function : CrossEntropyLoss \n* Smoothing Value : 0.65 ~ 0.75\n\n<br>\n\n# Regularization\nDuring model training, there are three primary regularization techniques that have made significant contributions to both improving convergence speed and final performance\n\n* Weight normalization applied to the final linear layer([paper](https://arxiv.org/abs/1602.07868))\n    \n        final_layer = torch.nn.utils.weight_norm(nn.Linear(hidden_size, 250))\n\n* Batch normalization applied before the final linear layer\n       \n        # head \n        x = fc_layer(x)\n        x = batchnorm1d(x)\n        x = relu(x)\n        x = dropout(x)\n        x = final_layer(x)\n\n* High weight decay value with the AdamW\n<br>\n\n# TFLite Conversion\nIn the early stages of the competition, I worked with PyTorch, which meant I had to deal with numerous errors when converting to TFLite, and ultimately failed to handle dynamic input shapes. In the latter stages of the competition, I began working with Keras, and the number of errors when converting to TFLite was significantly reduced.\n<br>\n\n# Didn't Work\n* Adding the distance feature contributed to performance improvement, but adding angle and direction features did not.\n* Increasing the number of transformer layers did not contribute to performance improvement.\n* I attempted to model the relationships between landmark points using Transformers or GATs, but the inference speed of the model became slower, and the performance actually decreased.\n* Bert-like pretraining (MLM) for XYZ coordinates did not improve performance with the provided data in the competition.\n<br>\n\n# Learned\nI am currently serving in the military in South Korea. During my army training, I learned a lot while coding in between. I was amazed by the culture of Kagglers sharing various ideas and discussing them. It's really cool that something bigger and better something can be born through this community, and I also want to participate in it diligently in the future. Thank you for reading my post.\n<br>\n\n# Code\nhttps://github.com/siwooyong/Google-Isolated-Sign-Language-Recognition\n<br>",
      "votes": 17
    },
    {
      "id": 2264527,
      "postDate": "2023-05-18T13:53:00.577Z",
      "content": "<p>Congratssss !!!</p>",
      "rawMarkdown": "Congratssss !!!\n",
      "votes": 1
    },
    {
      "id": 2247709,
      "postDate": "2023-05-06T08:19:17.143Z",
      "content": "<p>well done, congrats!!!!</p>",
      "rawMarkdown": "well done, congrats!!!!",
      "votes": 1
    },
    {
      "id": 2247222,
      "postDate": "2023-05-05T20:01:10.230Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a>, keep going!</p>",
      "rawMarkdown": "Congratulations @siwooyong, keep going!",
      "votes": 1,
      "replies": [
        {
          "id": 2247389,
          "postDate": "2023-05-06T00:42:54.273Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2247395,
          "postDate": "2023-05-06T00:53:01.617Z",
          "content": "<p>I appreciate it. Thank you!!!🔥🔥🔥</p>",
          "rawMarkdown": "I appreciate it. Thank you!!!🔥🔥🔥",
          "votes": 1
        }
      ]
    },
    {
      "id": 2246970,
      "postDate": "2023-05-05T15:25:22.433Z",
      "content": "<p>Congratulations!!!<br>\nDo you know why there are many errors when converting Pytorch to TFlite?</p>",
      "rawMarkdown": "Congratulations!!!\nDo you know why there are many errors when converting Pytorch to TFlite?",
      "votes": 1,
      "replies": [
        {
          "id": 2247393,
          "postDate": "2023-05-06T00:50:41.403Z",
          "content": "<p>In my opinion, the reason why there are many errors in the process of converting PyTorch to TFLite is that the process itself is quite lengthy. Unlike Keras models, which can be directly converted using specific code, </p>\n<pre><code>converter = tf.lite.TFLiteConverter.from_keras_model(KerasModel)\ntflite_model = converter.convert()\n</code></pre>\n<p>PyTorch models go through multiple steps, such as conversion to ONNX, ONNX to TensorFlow, and then TensorFlow to TFLite. Additionally, there are often cases where operations used in PyTorch are not supported in TFLite, resulting in frequent errors during the conversion process.</p>",
          "rawMarkdown": "In my opinion, the reason why there are many errors in the process of converting PyTorch to TFLite is that the process itself is quite lengthy. Unlike Keras models, which can be directly converted using specific code, \n\n    converter = tf.lite.TFLiteConverter.from_keras_model(KerasModel)\n    tflite_model = converter.convert()\n\n\nPyTorch models go through multiple steps, such as conversion to ONNX, ONNX to TensorFlow, and then TensorFlow to TFLite. Additionally, there are often cases where operations used in PyTorch are not supported in TFLite, resulting in frequent errors during the conversion process."
        }
      ]
    },
    {
      "id": 2246831,
      "postDate": "2023-05-05T13:53:55.853Z",
      "content": "<p>Congrats bro</p>",
      "rawMarkdown": "Congrats bro",
      "votes": 1,
      "replies": [
        {
          "id": 2246932,
          "postDate": "2023-05-05T14:52:24.547Z",
          "content": "<p>Thank you!!! bro👍👍👍</p>",
          "rawMarkdown": "Thank you!!! bro👍👍👍"
        }
      ]
    },
    {
      "id": 2265609,
      "postDate": "2023-05-19T11:30:02.840Z",
      "content": "<p><code>data = np.load('/content/drive/MyDrive/Kaggle/aggregation/data_m_with_lip.npy')\nlabel = np.load('/content/drive/MyDrive/Kaggle/aggregation/label.npy')\nframe = np.load('/content/drive/MyDrive/Kaggle/aggregation/frame.npy')\nbatch_id = np.load('/content/drive/MyDrive/Kaggle/aggregation/batch_id.npy')</code><br>\nhi i'm a new kaggler and i didn't understand this part.<br>\ncan u please explain to me where did u get this npy files?</p>",
      "rawMarkdown": "`data = np.load('/content/drive/MyDrive/Kaggle/aggregation/data_m_with_lip.npy')\nlabel = np.load('/content/drive/MyDrive/Kaggle/aggregation/label.npy')\nframe = np.load('/content/drive/MyDrive/Kaggle/aggregation/frame.npy')\nbatch_id = np.load('/content/drive/MyDrive/Kaggle/aggregation/batch_id.npy')`\nhi i'm a new kaggler and i didn't understand this part.\ncan u please explain to me where did u get this npy files?"
    },
    {
      "id": 2244160,
      "postDate": "2023-05-03T13:30:20.063Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2264527,
      "author_name": "Khadija Benihich",
      "author_url": "",
      "post_date": "2023-05-18T13:53:00.577000",
      "content": "<p>Congratssss !!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2247709,
      "author_name": "Aiana Abdyrakhmanova",
      "author_url": "",
      "post_date": "2023-05-06T08:19:17.143000",
      "content": "<p>well done, congrats!!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2247222,
      "author_name": "szebiniso",
      "author_url": "",
      "post_date": "2023-05-05T20:01:10.230000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a>, keep going!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2247389,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-05-06T00:42:54.273000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2247395,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "2023-05-06T00:53:01.617000",
          "content": "<p>I appreciate it. Thank you!!!🔥🔥🔥</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2246970,
      "author_name": "Ivan Isaev",
      "author_url": "",
      "post_date": "2023-05-05T15:25:22.433000",
      "content": "<p>Congratulations!!!<br>\nDo you know why there are many errors when converting Pytorch to TFlite?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2247393,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "2023-05-06T00:50:41.403000",
          "content": "<p>In my opinion, the reason why there are many errors in the process of converting PyTorch to TFLite is that the process itself is quite lengthy. Unlike Keras models, which can be directly converted using specific code, </p>\n<pre><code>converter = tf.lite.TFLiteConverter.from_keras_model(KerasModel)\ntflite_model = converter.convert()\n</code></pre>\n<p>PyTorch models go through multiple steps, such as conversion to ONNX, ONNX to TensorFlow, and then TensorFlow to TFLite. Additionally, there are often cases where operations used in PyTorch are not supported in TFLite, resulting in frequent errors during the conversion process.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2246831,
      "author_name": "navyrutada",
      "author_url": "",
      "post_date": "2023-05-05T13:53:55.853000",
      "content": "<p>Congrats bro</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2246932,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "2023-05-05T14:52:24.547000",
          "content": "<p>Thank you!!! bro👍👍👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2265609,
      "author_name": "Khadija Benihich",
      "author_url": "",
      "post_date": "2023-05-19T11:30:02.840000",
      "content": "<p><code>data = np.load('/content/drive/MyDrive/Kaggle/aggregation/data_m_with_lip.npy')\nlabel = np.load('/content/drive/MyDrive/Kaggle/aggregation/label.npy')\nframe = np.load('/content/drive/MyDrive/Kaggle/aggregation/frame.npy')\nbatch_id = np.load('/content/drive/MyDrive/Kaggle/aggregation/batch_id.npy')</code><br>\nhi i'm a new kaggler and i didn't understand this part.<br>\ncan u please explain to me where did u get this npy files?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2244160,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-03T13:30:20.063000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2242571": "# TLDR\nBased on my experiment, it seems that the most significant factors contributing to improved performance were **regularization**, **feature processing** and **embedding_layer**. The final ensemble model consisted of various input features and models, including Transformer, MLP, and GRU.\n<br>\n\n# Data Processing\nOut of the 543 landmarks in the provided raw data, a total of 115 landmarks were used for hands, poses, and lips. Input features were constructed by concatenating xy(z), motion, and distance, resulting in a final input feature size of 1196. Initially, using only the distance feature of hand data was employed, but the significant performance improvement (CV +0.05) was achieved by adding the distance feature of pose and lip data. \n    \n    feature = tf.concat([\n            tf.reshape(xyz_hand[:, :21, :3], [-1, 21 * 3]), \n            tf.reshape(xyz_pose[:, 21:46, :2], [-1, 25 * 2]), \n            tf.reshape(xyz_lip[:, 46:66, :2], [-1, 20 * 2]), \n            tf.reshape(motion_hand[:, :21, :3], [-1, 21 * 3]), \n            tf.reshape(motion_pose[:, 21:46, :2], [-1, 25 * 2]), \n            tf.reshape(motion_lip[:, 46:66, :2], [-1, 20 * 2]), \n            tf.reshape(distance_hand, [-1, 210]),\n            tf.reshape(distance_pose, [-1, 300]),\n            tf.reshape(distance_outlip, [-1, 190]),\n            tf.reshape(distance_inlip, [-1, 190]),\n        ], axis=-1)\n\nAdditionally, the hand with fewer NaN values was utilized, and the leg landmarks were removed from the pose landmarks. Based on this, two versions of inputs were created. The first version only utilized frames with non-NaN hand data, while the second version included frames with NaN hand data. The former had a max_length of 100, and the latter had a max_length of 200.\n\n    cond = lefth_sum > righth_sum\n    h_x = tf.where(cond, lefth_x, righth_x)\n    xfeat = tf.where(cond, tf.concat([lefth_x, pose_x, lip_x], axis = 1), tf.concat([righth_x, pose_x, lip_x], axis = 1))\n\n<br>\n\n# Augmentation\nThe hand with fewer NaN values was utilized, and both hands were flipped to be recognized as right hands in the model, which actually contributed to a performance improvement of about 0.01 in CV.\n    \n    # x-axis mirroring\n    cond = lefth_sum > righth_sum\n    xfeat_xcoordi = xfeat[:, :, 0]\n    xfeat_else = xfeat[:, :, 1:]\n    xfeat_xcoordi = tf.where(cond, -xfeat_xcoordi, xfeat_xcoordi)\n    xfeat = tf.concat([xfeat_xcoordi[:, :, tf.newaxis], xfeat_else], axis = -1)\n\nI did not observe that flip, rotate, mixup, and other augmentation techniques contributed to performance improvement in CV, so I supplemented this by ensembling the models of the two versions of inputs mentioned earlier.\n<br>\n\n# Model\nPrior to being utilized as inputs for the transformer model, the input features, namely xy(z), motion, and distance, underwent individual processing through dedicated embedding layers. Compared to the scenario where features were not processed independently, a performance improvement of 0.01 was observed in the CV score when the features were treated independently.\n    \n    # embedding layer\n    xy = xy_embeddings(xy)\n    motion = motion_embeddings(motion)\n    distance_hand = distance_hand_embeddings(distance_hand)\n    distance_pose = distance_pose_embeddings(distance_pose)\n    distance_outlip = distance_outlip_embeddings(distance_outlip)\n    distance_inlip = distance_inlip_embeddings(distance_inlip)\n\n    x = tf.concat([xy, motion, distance_hand, distance_pose, distance_outlip, distance_inlip], axis=-1)\n    x = relu(x)\n    x = fc_layer(x)\n    x = TransformerModel(input_ids = None, inputs_embeds=x, attention_mask=x_mask).last_hidden_state\n<br>\n\nFor Transformer models, I used huggingface's RoBERTa-PreLayerNorm, DeBERTaV2, and GPT2. The input was processed independently for xyz, motion, and distance, and then concatenated to form a 300-dimensional transformer input. The mean, max, and std values of the Transformer output were then concatenated to obtain the final output.\n    \n    # pooling function\n    def get_pool(self, x, x_mask):\n        x = x * tf.expand_dims(x_mask, axis=-1)  # apply mask\n        nonzero_count = tf.reduce_sum(x_mask, axis=1, keepdims=True)  # count nonzero elements\n        max_discount = (1-x_mask)*1e10\n\n        apool = tf.reduce_sum(x, axis=1) / nonzero_count\n        mpool = tf.reduce_max(x - tf.expand_dims(max_discount, axis=-1), axis=1)\n        spool = tf.sqrt((tf.reduce_sum(((x - tf.expand_dims(apool, axis=1)) ** 2) * tf.expand_dims(x_mask, axis=-1), axis=1) / nonzero_count) + 1e-9)\n        return tf.concat([apool, mpool, spool], axis=-1) \n\nIn addition to the Transformer model, simple linear models and GRU models also achieved similar performance to the Transformer model, so I ensembled these three types of models.\n<br>\n\n# Training\n\n* Scheduler : lr_warmup_cosine_decay \n* Warmup Ratio : 0.2 \n* Optimizer : AdamW \n* Weight Decay : 0.01\n* Epoch : 40\n* Learning Rate : 1e-3 \n* Loss Function : CrossEntropyLoss \n* Smoothing Value : 0.65 ~ 0.75\n\n<br>\n\n# Regularization\nDuring model training, there are three primary regularization techniques that have made significant contributions to both improving convergence speed and final performance\n\n* Weight normalization applied to the final linear layer([paper](https://arxiv.org/abs/1602.07868))\n    \n        final_layer = torch.nn.utils.weight_norm(nn.Linear(hidden_size, 250))\n\n* Batch normalization applied before the final linear layer\n       \n        # head \n        x = fc_layer(x)\n        x = batchnorm1d(x)\n        x = relu(x)\n        x = dropout(x)\n        x = final_layer(x)\n\n* High weight decay value with the AdamW\n<br>\n\n# TFLite Conversion\nIn the early stages of the competition, I worked with PyTorch, which meant I had to deal with numerous errors when converting to TFLite, and ultimately failed to handle dynamic input shapes. In the latter stages of the competition, I began working with Keras, and the number of errors when converting to TFLite was significantly reduced.\n<br>\n\n# Didn't Work\n* Adding the distance feature contributed to performance improvement, but adding angle and direction features did not.\n* Increasing the number of transformer layers did not contribute to performance improvement.\n* I attempted to model the relationships between landmark points using Transformers or GATs, but the inference speed of the model became slower, and the performance actually decreased.\n* Bert-like pretraining (MLM) for XYZ coordinates did not improve performance with the provided data in the competition.\n<br>\n\n# Learned\nI am currently serving in the military in South Korea. During my army training, I learned a lot while coding in between. I was amazed by the culture of Kagglers sharing various ideas and discussing them. It's really cool that something bigger and better something can be born through this community, and I also want to participate in it diligently in the future. Thank you for reading my post.\n<br>\n\n# Code\nhttps://github.com/siwooyong/Google-Isolated-Sign-Language-Recognition\n<br>",
    "2264527": "Congratssss !!!\n",
    "2247709": "well done, congrats!!!!",
    "2247222": "Congratulations @siwooyong, keep going!",
    "2246970": "Congratulations!!!\nDo you know why there are many errors when converting Pytorch to TFlite?",
    "2246831": "Congrats bro",
    "2265609": "`data = np.load('/content/drive/MyDrive/Kaggle/aggregation/data_m_with_lip.npy')\nlabel = np.load('/content/drive/MyDrive/Kaggle/aggregation/label.npy')\nframe = np.load('/content/drive/MyDrive/Kaggle/aggregation/frame.npy')\nbatch_id = np.load('/content/drive/MyDrive/Kaggle/aggregation/batch_id.npy')`\nhi i'm a new kaggler and i didn't understand this part.\ncan u please explain to me where did u get this npy files?",
    "2244160": ""
  }
}