{
  "id": 406122,
  "title": "Skeleton Based Action Recognition: A failed attempt",
  "url": "/competitions/asl-signs/discussion/406122",
  "author_name": "",
  "post_date": "2023-05-01T00:39:46.004743700Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Since this competition is about to end, I decided to document my learnings. Why a failed attempt? Well, I am not topping the leaderboard. Why?</p>\n<ul>\n<li>Lack of computing</li>\n<li>Could not focus on the competition consistently</li>\n</ul>\n<p>Now that I am done sulking about not reaching the top of LB and preparing myself for another competition, it's important to share what didn't work in this competition. Even though these ideas didn't turn out to be great for this competition, I am positive that more can be squeezed out of them. By sharing my failed attempts, I am illuminating a corner of the overall space of potential solutions, IMO.</p>\n<h1>Setup</h1>\n<h2>Data</h2>\n<p>I converted the provided data (in .parquet) to TFRecords to improve the I/O bottleneck of the data pipeline. It was also done to leverage the TPUs.</p>\n<ul>\n<li>Stratified split of the dataset: <a href=\"https://www.kaggle.com/datasets/ayuraj/asl-tfrecords\" target=\"_blank\">https://www.kaggle.com/datasets/ayuraj/asl-tfrecords</a></li>\n<li>TFRecords by <code>participant_id</code>: <a href=\"https://www.kaggle.com/datasets/ayuraj/asl-tfrecords-participants\" target=\"_blank\">https://www.kaggle.com/datasets/ayuraj/asl-tfrecords-participants</a></li>\n</ul>\n<p>All the training were done with the stratified split (20 splits for training and 4 splits for validation).</p>\n<h2>Framework</h2>\n<p>Well, I decided to use TensorFlow/Keras entirely for my workflow. This resulted in a few limitations - no SOTA action recognition model, had to implement the papers from scratch. Why stick with it then? It's my core area of strength + one can leverage the same pipeline to do quantised aware training (I didn't go to this point though).</p>\n<h2>Experiment Tracking</h2>\n<p>I used Weights and Biases for tracking everything. Here is my W&amp;B dashboard: <a href=\"https://wandb.ai/ayush-thakur/kaggle-asl\" target=\"_blank\">https://wandb.ai/ayush-thakur/kaggle-asl</a></p>\n<p>There is a lot of unfiltered information on this dashboard but you might still find it useful especially to organise your own experiments.</p>\n<h1>1. ConvLSTM1D</h1>\n<p>One of the first initial things I tried was good-old <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/layers/ConvLSTM1D\" target=\"_blank\"><code>ConvLSTM1D</code></a> based model. The implementation can be found in this <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/03_Train%20Baseline.ipynb\" target=\"_blank\">notebook</a>.</p>\n<p>The Data pipeline uses the <code>tf.cond</code> method to zero pad the sequence if num of frames were lower than certain threshold. If the num frames was greater than threshold, the sequence was sliced.</p>\n<p>I used Pose and Right and Left hands for training. Adding extra body types didn't help as reported in multiple notebooks and discussions.</p>\n<h1>2. Tubelet Embedding</h1>\n<p>This was inspired by the <a href=\"https://arxiv.org/abs/2103.15691\" target=\"_blank\">Video ViT paper</a>. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2Fa66298421b178fa887022e93194583ab%2FScreenshot%202023-05-01%20at%205.33.08%20AM.png?generation=1682899415191246&amp;alt=media\" alt=\"\"></p>\n<p>Aritra and I <a href=\"https://keras.io/examples/vision/vivit/\" target=\"_blank\">implemented this architecture in Keras</a> and the key idea is to use a Tubelet Embedding that basically extract patches from a sequence and embeds them. It's works really well for videos but didn't translate to skeleton data sequence.</p>\n<h1>3. DDNet (Double-motion Network)</h1>\n<p>This idea was introduced in <a href=\"https://arxiv.org/abs/1907.09658\" target=\"_blank\">Make Skeleton-based Action Recognition Model Smaller, Faster and Better</a> paper. The idea is to represent the skeleton as Joint Collection Distances (JCD) and use motion features. You can find the implementation in this <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/11_JCD_and_Motion_Features.ipynb\" target=\"_blank\">notebook</a>.</p>\n<p>Having done a few experimentations I noticed that motion features worked well and JCD wasn't working at all.</p>\n<h1>4. Graph based networks</h1>\n<p>I did a lot of research to use graph based neural networks. I didn't have much experience with graph based nets thus I moved to # 5 but I would still like to documents the papers I read and my what imo should work in this competition. I will not be surprised if any of the top solution is using one of these papers for inspiration.</p>\n<ul>\n<li>STGCN: <a href=\"https://arxiv.org/abs/1801.07455\" target=\"_blank\">https://arxiv.org/abs/1801.07455</a></li>\n<li>InfoGCN: <a href=\"https://openaccess.thecvf.com/content/CVPR2022/papers/Chi_InfoGCN_Representation_Learning_for_Human_Skeleton-Based_Action_Recognition_CVPR_2022_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/CVPR2022/papers/Chi_InfoGCN_Representation_Learning_for_Human_Skeleton-Based_Action_Recognition_CVPR_2022_paper.pdf</a> One could have used this network to learn a meaningful representation from extra data (pretrain in a sense) and use the encoder to be fine-tuned on this competition's data.</li>\n<li>Hyperformer: <a href=\"https://arxiv.org/abs/2211.09590\" target=\"_blank\">https://arxiv.org/abs/2211.09590</a> This was the most promising paper imo. I was able to find an official PyTorch <a href=\"https://github.com/ZhouYuxuanYX/Hyperformer\" target=\"_blank\">implementation</a> but was too late for me to start porting it to TF.</li>\n</ul>\n<h1>5. PoseConv3D</h1>\n<p>The paper <a href=\"https://arxiv.org/pdf/2104.13586.pdf\" target=\"_blank\">Revisiting Skeleton-based Action Recognition</a> introduces the idea of converting graph data (skeleton) to heat map and use <code>Conv3D</code> based image-style model (ResNet). This paper though old, is still the SOTA in skeleton based action recognition.</p>\n<p>I spent the last two-three weeks implementing this paper. The hardest part was managing the dataset size. You can find my dataset creation step <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/scripts/create_heatmap.py\" target=\"_blank\">here</a>. <strong>The resulting dataset was visualised here</strong>: <a href=\"https://wandb.ai/ayush-thakur/kaggle-asl/runs/nkn0h3sw?workspace=user-ayush-thakur\" target=\"_blank\">https://wandb.ai/ayush-thakur/kaggle-asl/runs/nkn0h3sw?workspace=user-ayush-thakur</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2F9d7b9a8bc270fec4e9b0e74d69dbd0f4%2FScreenshot%202023-05-01%20at%205.57.35%20AM.png?generation=1682900877237841&amp;alt=media\" alt=\"\"></p>\n<p>I created a simple <code>Conv3D</code> based model and was able to reach the same score as that of <code>ConvLSTM1D</code>. I think this method has a lot of promise but requires compute to train it well.</p>\n<p>You can find the training code <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/train_poseconv3d.py\" target=\"_blank\">here</a>.</p>\n<h1>Conclusion</h1>\n<p>It was a fun competition. Thanks to the organizers for arranging this and congrats to everyone who participated in this competition. The impact of ML to solve problem statement like this is why we all are trying to push this technology ahead. What could I have done better?</p>\n<ul>\n<li>Collaborated - I am in search of a team with whom I can participate more seriously.</li>\n<li>Experimented more with a idea before moving ahead.</li>\n</ul>",
  "messages": [
    {
      "id": "2240883",
      "postDate": "05/01/2023 00:39:46",
      "content": "<p>Since this competition is about to end, I decided to document my learnings. Why a failed attempt? Well, I am not topping the leaderboard. Why?</p>\n<ul>\n<li>Lack of computing</li>\n<li>Could not focus on the competition consistently</li>\n</ul>\n<p>Now that I am done sulking about not reaching the top of LB and preparing myself for another competition, it's important to share what didn't work in this competition. Even though these ideas didn't turn out to be great for this competition, I am positive that more can be squeezed out of them. By sharing my failed attempts, I am illuminating a corner of the overall space of potential solutions, IMO.</p>\n<h1>Setup</h1>\n<h2>Data</h2>\n<p>I converted the provided data (in .parquet) to TFRecords to improve the I/O bottleneck of the data pipeline. It was also done to leverage the TPUs.</p>\n<ul>\n<li>Stratified split of the dataset: <a href=\"https://www.kaggle.com/datasets/ayuraj/asl-tfrecords\" target=\"_blank\">https://www.kaggle.com/datasets/ayuraj/asl-tfrecords</a></li>\n<li>TFRecords by <code>participant_id</code>: <a href=\"https://www.kaggle.com/datasets/ayuraj/asl-tfrecords-participants\" target=\"_blank\">https://www.kaggle.com/datasets/ayuraj/asl-tfrecords-participants</a></li>\n</ul>\n<p>All the training were done with the stratified split (20 splits for training and 4 splits for validation).</p>\n<h2>Framework</h2>\n<p>Well, I decided to use TensorFlow/Keras entirely for my workflow. This resulted in a few limitations - no SOTA action recognition model, had to implement the papers from scratch. Why stick with it then? It's my core area of strength + one can leverage the same pipeline to do quantised aware training (I didn't go to this point though).</p>\n<h2>Experiment Tracking</h2>\n<p>I used Weights and Biases for tracking everything. Here is my W&amp;B dashboard: <a href=\"https://wandb.ai/ayush-thakur/kaggle-asl\" target=\"_blank\">https://wandb.ai/ayush-thakur/kaggle-asl</a></p>\n<p>There is a lot of unfiltered information on this dashboard but you might still find it useful especially to organise your own experiments.</p>\n<h1>1. ConvLSTM1D</h1>\n<p>One of the first initial things I tried was good-old <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/layers/ConvLSTM1D\" target=\"_blank\"><code>ConvLSTM1D</code></a> based model. The implementation can be found in this <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/03_Train%20Baseline.ipynb\" target=\"_blank\">notebook</a>.</p>\n<p>The Data pipeline uses the <code>tf.cond</code> method to zero pad the sequence if num of frames were lower than certain threshold. If the num frames was greater than threshold, the sequence was sliced.</p>\n<p>I used Pose and Right and Left hands for training. Adding extra body types didn't help as reported in multiple notebooks and discussions.</p>\n<h1>2. Tubelet Embedding</h1>\n<p>This was inspired by the <a href=\"https://arxiv.org/abs/2103.15691\" target=\"_blank\">Video ViT paper</a>. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2Fa66298421b178fa887022e93194583ab%2FScreenshot%202023-05-01%20at%205.33.08%20AM.png?generation=1682899415191246&amp;alt=media\" alt=\"\"></p>\n<p>Aritra and I <a href=\"https://keras.io/examples/vision/vivit/\" target=\"_blank\">implemented this architecture in Keras</a> and the key idea is to use a Tubelet Embedding that basically extract patches from a sequence and embeds them. It's works really well for videos but didn't translate to skeleton data sequence.</p>\n<h1>3. DDNet (Double-motion Network)</h1>\n<p>This idea was introduced in <a href=\"https://arxiv.org/abs/1907.09658\" target=\"_blank\">Make Skeleton-based Action Recognition Model Smaller, Faster and Better</a> paper. The idea is to represent the skeleton as Joint Collection Distances (JCD) and use motion features. You can find the implementation in this <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/11_JCD_and_Motion_Features.ipynb\" target=\"_blank\">notebook</a>.</p>\n<p>Having done a few experimentations I noticed that motion features worked well and JCD wasn't working at all.</p>\n<h1>4. Graph based networks</h1>\n<p>I did a lot of research to use graph based neural networks. I didn't have much experience with graph based nets thus I moved to # 5 but I would still like to documents the papers I read and my what imo should work in this competition. I will not be surprised if any of the top solution is using one of these papers for inspiration.</p>\n<ul>\n<li>STGCN: <a href=\"https://arxiv.org/abs/1801.07455\" target=\"_blank\">https://arxiv.org/abs/1801.07455</a></li>\n<li>InfoGCN: <a href=\"https://openaccess.thecvf.com/content/CVPR2022/papers/Chi_InfoGCN_Representation_Learning_for_Human_Skeleton-Based_Action_Recognition_CVPR_2022_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/CVPR2022/papers/Chi_InfoGCN_Representation_Learning_for_Human_Skeleton-Based_Action_Recognition_CVPR_2022_paper.pdf</a> One could have used this network to learn a meaningful representation from extra data (pretrain in a sense) and use the encoder to be fine-tuned on this competition's data.</li>\n<li>Hyperformer: <a href=\"https://arxiv.org/abs/2211.09590\" target=\"_blank\">https://arxiv.org/abs/2211.09590</a> This was the most promising paper imo. I was able to find an official PyTorch <a href=\"https://github.com/ZhouYuxuanYX/Hyperformer\" target=\"_blank\">implementation</a> but was too late for me to start porting it to TF.</li>\n</ul>\n<h1>5. PoseConv3D</h1>\n<p>The paper <a href=\"https://arxiv.org/pdf/2104.13586.pdf\" target=\"_blank\">Revisiting Skeleton-based Action Recognition</a> introduces the idea of converting graph data (skeleton) to heat map and use <code>Conv3D</code> based image-style model (ResNet). This paper though old, is still the SOTA in skeleton based action recognition.</p>\n<p>I spent the last two-three weeks implementing this paper. The hardest part was managing the dataset size. You can find my dataset creation step <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/scripts/create_heatmap.py\" target=\"_blank\">here</a>. <strong>The resulting dataset was visualised here</strong>: <a href=\"https://wandb.ai/ayush-thakur/kaggle-asl/runs/nkn0h3sw?workspace=user-ayush-thakur\" target=\"_blank\">https://wandb.ai/ayush-thakur/kaggle-asl/runs/nkn0h3sw?workspace=user-ayush-thakur</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2F9d7b9a8bc270fec4e9b0e74d69dbd0f4%2FScreenshot%202023-05-01%20at%205.57.35%20AM.png?generation=1682900877237841&amp;alt=media\" alt=\"\"></p>\n<p>I created a simple <code>Conv3D</code> based model and was able to reach the same score as that of <code>ConvLSTM1D</code>. I think this method has a lot of promise but requires compute to train it well.</p>\n<p>You can find the training code <a href=\"https://github.com/ayulockin/kaggle-asl/blob/main/train_poseconv3d.py\" target=\"_blank\">here</a>.</p>\n<h1>Conclusion</h1>\n<p>It was a fun competition. Thanks to the organizers for arranging this and congrats to everyone who participated in this competition. The impact of ML to solve problem statement like this is why we all are trying to push this technology ahead. What could I have done better?</p>\n<ul>\n<li>Collaborated - I am in search of a team with whom I can participate more seriously.</li>\n<li>Experimented more with a idea before moving ahead.</li>\n</ul>",
      "rawMarkdown": "Since this competition is about to end, I decided to document my learnings. Why a failed attempt? Well, I am not topping the leaderboard. Why?\n* Lack of computing\n* Could not focus on the competition consistently\n\nNow that I am done sulking about not reaching the top of LB and preparing myself for another competition, it's important to share what didn't work in this competition. Even though these ideas didn't turn out to be great for this competition, I am positive that more can be squeezed out of them. By sharing my failed attempts, I am illuminating a corner of the overall space of potential solutions, IMO.\n\n# Setup\n\n## Data\n\nI converted the provided data (in .parquet) to TFRecords to improve the I/O bottleneck of the data pipeline. It was also done to leverage the TPUs.\n* Stratified split of the dataset: https://www.kaggle.com/datasets/ayuraj/asl-tfrecords\n* TFRecords by `participant_id`: https://www.kaggle.com/datasets/ayuraj/asl-tfrecords-participants\n\nAll the training were done with the stratified split (20 splits for training and 4 splits for validation).\n\n## Framework\n\nWell, I decided to use TensorFlow/Keras entirely for my workflow. This resulted in a few limitations - no SOTA action recognition model, had to implement the papers from scratch. Why stick with it then? It's my core area of strength + one can leverage the same pipeline to do quantised aware training (I didn't go to this point though).\n\n## Experiment Tracking\n\nI used Weights and Biases for tracking everything. Here is my W&B dashboard: https://wandb.ai/ayush-thakur/kaggle-asl\n\nThere is a lot of unfiltered information on this dashboard but you might still find it useful especially to organise your own experiments.\n\n# 1. ConvLSTM1D\n\nOne of the first initial things I tried was good-old [`ConvLSTM1D`](https://www.tensorflow.org/api_docs/python/tf/keras/layers/ConvLSTM1D) based model. The implementation can be found in this [notebook](https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/03_Train%20Baseline.ipynb).\n\nThe Data pipeline uses the `tf.cond` method to zero pad the sequence if num of frames were lower than certain threshold. If the num frames was greater than threshold, the sequence was sliced.\n\nI used Pose and Right and Left hands for training. Adding extra body types didn't help as reported in multiple notebooks and discussions.\n\n# 2. Tubelet Embedding\n\nThis was inspired by the [Video ViT paper](https://arxiv.org/abs/2103.15691). ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2Fa66298421b178fa887022e93194583ab%2FScreenshot%202023-05-01%20at%205.33.08%20AM.png?generation=1682899415191246&alt=media)\n\nAritra and I [implemented this architecture in Keras](https://keras.io/examples/vision/vivit/) and the key idea is to use a Tubelet Embedding that basically extract patches from a sequence and embeds them. It's works really well for videos but didn't translate to skeleton data sequence.\n\n# 3. DDNet (Double-motion Network)\n\nThis idea was introduced in [Make Skeleton-based Action Recognition Model Smaller, Faster and Better](https://arxiv.org/abs/1907.09658) paper. The idea is to represent the skeleton as Joint Collection Distances (JCD) and use motion features. You can find the implementation in this [notebook](https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/11_JCD_and_Motion_Features.ipynb).\n\nHaving done a few experimentations I noticed that motion features worked well and JCD wasn't working at all.\n\n# 4. Graph based networks\n\nI did a lot of research to use graph based neural networks. I didn't have much experience with graph based nets thus I moved to # 5 but I would still like to documents the papers I read and my what imo should work in this competition. I will not be surprised if any of the top solution is using one of these papers for inspiration.\n\n* STGCN: https://arxiv.org/abs/1801.07455\n* InfoGCN: https://openaccess.thecvf.com/content/CVPR2022/papers/Chi_InfoGCN_Representation_Learning_for_Human_Skeleton-Based_Action_Recognition_CVPR_2022_paper.pdf One could have used this network to learn a meaningful representation from extra data (pretrain in a sense) and use the encoder to be fine-tuned on this competition's data.\n* Hyperformer: https://arxiv.org/abs/2211.09590 This was the most promising paper imo. I was able to find an official PyTorch [implementation](https://github.com/ZhouYuxuanYX/Hyperformer) but was too late for me to start porting it to TF.\n\n# 5. PoseConv3D\n\nThe paper [Revisiting Skeleton-based Action Recognition](https://arxiv.org/pdf/2104.13586.pdf) introduces the idea of converting graph data (skeleton) to heat map and use `Conv3D` based image-style model (ResNet). This paper though old, is still the SOTA in skeleton based action recognition.\n\nI spent the last two-three weeks implementing this paper. The hardest part was managing the dataset size. You can find my dataset creation step [here](https://github.com/ayulockin/kaggle-asl/blob/main/scripts/create_heatmap.py). **The resulting dataset was visualised here**: https://wandb.ai/ayush-thakur/kaggle-asl/runs/nkn0h3sw?workspace=user-ayush-thakur\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2F9d7b9a8bc270fec4e9b0e74d69dbd0f4%2FScreenshot%202023-05-01%20at%205.57.35%20AM.png?generation=1682900877237841&alt=media)\n\nI created a simple `Conv3D` based model and was able to reach the same score as that of `ConvLSTM1D`. I think this method has a lot of promise but requires compute to train it well.\n\nYou can find the training code [here](https://github.com/ayulockin/kaggle-asl/blob/main/train_poseconv3d.py).\n\n# Conclusion\n\nIt was a fun competition. Thanks to the organizers for arranging this and congrats to everyone who participated in this competition. The impact of ML to solve problem statement like this is why we all are trying to push this technology ahead. What could I have done better?\n\n* Collaborated - I am in search of a team with whom I can participate more seriously.\n* Experimented more with a idea before moving ahead.",
      "votes": null
    },
    {
      "id": "2242489",
      "postDate": "05/02/2023 10:17:47",
      "content": "<p>Your willingness to share your journey and learnings is inspiring, and I'm sure you'll do great in the next competition. Keep pushing forward! </p>",
      "rawMarkdown": "Your willingness to share your journey and learnings is inspiring, and I'm sure you'll do great in the next competition. Keep pushing forward!",
      "votes": null
    },
    {
      "id": "2243368",
      "postDate": "05/02/2023 21:23:49",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/drindramanisinha\" target=\"_blank\">@drindramanisinha</a>.</p>",
      "rawMarkdown": "Thanks @drindramanisinha.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242489,
      "author_name": "drindramanisinha",
      "author_url": "",
      "post_date": "05/02/2023 10:17:47",
      "content": "<p>Your willingness to share your journey and learnings is inspiring, and I'm sure you'll do great in the next competition. Keep pushing forward! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2243368,
          "author_name": "ayuraj",
          "author_url": "",
          "post_date": "05/02/2023 21:23:49",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/drindramanisinha\" target=\"_blank\">@drindramanisinha</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2240883": "Since this competition is about to end, I decided to document my learnings. Why a failed attempt? Well, I am not topping the leaderboard. Why?\n* Lack of computing\n* Could not focus on the competition consistently\n\nNow that I am done sulking about not reaching the top of LB and preparing myself for another competition, it's important to share what didn't work in this competition. Even though these ideas didn't turn out to be great for this competition, I am positive that more can be squeezed out of them. By sharing my failed attempts, I am illuminating a corner of the overall space of potential solutions, IMO.\n\n# Setup\n\n## Data\n\nI converted the provided data (in .parquet) to TFRecords to improve the I/O bottleneck of the data pipeline. It was also done to leverage the TPUs.\n* Stratified split of the dataset: https://www.kaggle.com/datasets/ayuraj/asl-tfrecords\n* TFRecords by `participant_id`: https://www.kaggle.com/datasets/ayuraj/asl-tfrecords-participants\n\nAll the training were done with the stratified split (20 splits for training and 4 splits for validation).\n\n## Framework\n\nWell, I decided to use TensorFlow/Keras entirely for my workflow. This resulted in a few limitations - no SOTA action recognition model, had to implement the papers from scratch. Why stick with it then? It's my core area of strength + one can leverage the same pipeline to do quantised aware training (I didn't go to this point though).\n\n## Experiment Tracking\n\nI used Weights and Biases for tracking everything. Here is my W&B dashboard: https://wandb.ai/ayush-thakur/kaggle-asl\n\nThere is a lot of unfiltered information on this dashboard but you might still find it useful especially to organise your own experiments.\n\n# 1. ConvLSTM1D\n\nOne of the first initial things I tried was good-old [`ConvLSTM1D`](https://www.tensorflow.org/api_docs/python/tf/keras/layers/ConvLSTM1D) based model. The implementation can be found in this [notebook](https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/03_Train%20Baseline.ipynb).\n\nThe Data pipeline uses the `tf.cond` method to zero pad the sequence if num of frames were lower than certain threshold. If the num frames was greater than threshold, the sequence was sliced.\n\nI used Pose and Right and Left hands for training. Adding extra body types didn't help as reported in multiple notebooks and discussions.\n\n# 2. Tubelet Embedding\n\nThis was inspired by the [Video ViT paper](https://arxiv.org/abs/2103.15691). ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2Fa66298421b178fa887022e93194583ab%2FScreenshot%202023-05-01%20at%205.33.08%20AM.png?generation=1682899415191246&alt=media)\n\nAritra and I [implemented this architecture in Keras](https://keras.io/examples/vision/vivit/) and the key idea is to use a Tubelet Embedding that basically extract patches from a sequence and embeds them. It's works really well for videos but didn't translate to skeleton data sequence.\n\n# 3. DDNet (Double-motion Network)\n\nThis idea was introduced in [Make Skeleton-based Action Recognition Model Smaller, Faster and Better](https://arxiv.org/abs/1907.09658) paper. The idea is to represent the skeleton as Joint Collection Distances (JCD) and use motion features. You can find the implementation in this [notebook](https://github.com/ayulockin/kaggle-asl/blob/main/notebooks/11_JCD_and_Motion_Features.ipynb).\n\nHaving done a few experimentations I noticed that motion features worked well and JCD wasn't working at all.\n\n# 4. Graph based networks\n\nI did a lot of research to use graph based neural networks. I didn't have much experience with graph based nets thus I moved to # 5 but I would still like to documents the papers I read and my what imo should work in this competition. I will not be surprised if any of the top solution is using one of these papers for inspiration.\n\n* STGCN: https://arxiv.org/abs/1801.07455\n* InfoGCN: https://openaccess.thecvf.com/content/CVPR2022/papers/Chi_InfoGCN_Representation_Learning_for_Human_Skeleton-Based_Action_Recognition_CVPR_2022_paper.pdf One could have used this network to learn a meaningful representation from extra data (pretrain in a sense) and use the encoder to be fine-tuned on this competition's data.\n* Hyperformer: https://arxiv.org/abs/2211.09590 This was the most promising paper imo. I was able to find an official PyTorch [implementation](https://github.com/ZhouYuxuanYX/Hyperformer) but was too late for me to start porting it to TF.\n\n# 5. PoseConv3D\n\nThe paper [Revisiting Skeleton-based Action Recognition](https://arxiv.org/pdf/2104.13586.pdf) introduces the idea of converting graph data (skeleton) to heat map and use `Conv3D` based image-style model (ResNet). This paper though old, is still the SOTA in skeleton based action recognition.\n\nI spent the last two-three weeks implementing this paper. The hardest part was managing the dataset size. You can find my dataset creation step [here](https://github.com/ayulockin/kaggle-asl/blob/main/scripts/create_heatmap.py). **The resulting dataset was visualised here**: https://wandb.ai/ayush-thakur/kaggle-asl/runs/nkn0h3sw?workspace=user-ayush-thakur\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2212893%2F9d7b9a8bc270fec4e9b0e74d69dbd0f4%2FScreenshot%202023-05-01%20at%205.57.35%20AM.png?generation=1682900877237841&alt=media)\n\nI created a simple `Conv3D` based model and was able to reach the same score as that of `ConvLSTM1D`. I think this method has a lot of promise but requires compute to train it well.\n\nYou can find the training code [here](https://github.com/ayulockin/kaggle-asl/blob/main/train_poseconv3d.py).\n\n# Conclusion\n\nIt was a fun competition. Thanks to the organizers for arranging this and congrats to everyone who participated in this competition. The impact of ML to solve problem statement like this is why we all are trying to push this technology ahead. What could I have done better?\n\n* Collaborated - I am in search of a team with whom I can participate more seriously.\n* Experimented more with a idea before moving ahead.",
    "2242489": "Your willingness to share your journey and learnings is inspiring, and I'm sure you'll do great in the next competition. Keep pushing forward!",
    "2243368": "Thanks @drindramanisinha."
  },
  "source": "meta"
}