{
  "id": 406568,
  "title": "3rd Place Solution",
  "url": "/competitions/asl-signs/writeups/sabaisabai-3rd-place-solution",
  "author_name": "",
  "post_date": "2023-05-02T20:14:12.267Z",
  "votes": 37,
  "comment_count": 5,
  "views": 0,
  "content": "<p>We used an <strong>ensemble of six conv1d models and two versions of transformers</strong> based on the <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">public notebook</a>. The key points are <strong>data preprocessing, hard augmentation and ensemble</strong>.</p>\n<p>I initially tried to develop my own solution based on the aforementioned public transformer but then I found that the architecture of a model didn’t matter as much as proper work with data and ensembling, so I switched to more simple architectures like <strong>conv1d multilayer models</strong>. I used <code>Keras</code> and utilized all the advantages of native tflite layers like <code>DepthwiseConv1D</code>. My typical model looked like that:</p>\n<pre><code>do = \n\nmodel = Sequential()\nmodel.add(InputLayer(input_shape=(max_len, , )))\nmodel.add(Reshape((max_len, *)))\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(MaxPool1D(, ))\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(GlobalAvgPool1D())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(, activation=))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(, activation=))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(, activation=))\n</code></pre>\n<p>Then two of my colleagues joined me. We had a lot of ideas such as using synthetic data. We found this <a href=\"https://upcommons.upc.edu/bitstream/handle/2117/94313/TEPP1de1.pdf\" target=\"_blank\">paper</a>, created a script which generated synthetic 3D points of a hand parametrized by inner parameters (joints rotations) and camera model and tried to predict the inner parameters of hand by a model, based on the 3d points. And then use this pretrained model to preprocess data. But unfortunately we didn’t have enough time to finish it.</p>\n<p>Also we tried to train on additional data (WLASL), found that this dataset gave ~0.005-0.007 increase in local CV score but then we abandoned this idea and haven’t used additional data in our LB submissions.</p>\n<h2>Normalization</h2>\n<p>We used normalization by this reference points <code>ref = coords[:, [500, 501, 512, 513, 159,  386, 13,]]</code><br>\nWe didn’t use depth dimension.</p>\n<p>We didn’t throw away frames without hands but we took the presence of a hand into account during TTA. </p>\n<h2>Augmentation</h2>\n<ul>\n<li>Random rotation, shift, scale applied globally to each frame</li>\n<li>Random small shift added to every point </li>\n<li>Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score</li>\n<li>CutMix - took parts of samples from different classes. For hand we used 0.7 as a label value, for other parts - 0.3</li>\n</ul>\n<h2>TTA</h2>\n<ul>\n<li>Random left/right padding for short sequences</li>\n<li>For large sequences of frames we throw away frames with different probability taking into account presence of hands on a frame</li>\n</ul>\n<p>For different models we used different combinations of points: both hands / only active hand, lips, eyes, top part of pose. For some models we used only every second point of lips/eyes.<br>\nFor models where we used only one hand we determined which hand is more presented in the video and if it was the left hand we mirrored all points and their coordinates.<br>\nWe took only <strong>32 frames</strong> for all models except for one where we used 96 frames.</p>\n<p>These conv1d models were fast and lightweight so we could ensemble up to six models and still had space and time for additional models and TTA. Ensemble of six such models got <strong>0.7948</strong> on public LB and <strong>0.8711</strong> on private LB.<br>\nThen we decided to collaborate with <a href=\"https://www.kaggle.com/sqqqqy\" target=\"_blank\">@sqqqqy</a> who had been working on improving of transformer from a public kernel. His solution looked like that:</p>\n<h2>1. Transformer Model</h2>\n<ul>\n<li>Our transformer model is based on the <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">public notebook</a>, using different sequence normalization and smaller UNITS size to reduce model parameters.</li>\n<li>We implement two kinds of transformer-based model, the first model learn sperate embeddings for each part (just like public notebook), the second learn one embedding with input whole xyz sequence.</li>\n<li>1seed of first model score <strong>LB 0.77+</strong>, 1seed of second model score <strong>LB 0.768</strong></li>\n</ul>\n<h2>2. Preprocessing</h2>\n<ul>\n<li>We use 20 lip points, 32 eyes points, 42 hands points(left hand and right hand) and 8 pose points.</li>\n<li>The input sequence is normalized with shoulder, hip, lip and eyes points.</li>\n<li>Filling the NaN values with 0.0</li>\n<li>Learn a motion embedding by input $(d_x, d_y, \\sqrt{(d_x)^2+(d_y)^2})<em>t$ sequence, \n$$\n(d_x, d_y)_t = xyz_t - xyz</em>{t-1}<br>\n$$</li>\n<li>The final embedding is the concat of motion embedding and xyz embedding</li>\n</ul>\n<h2>3.Augmentation</h2>\n<ul>\n<li><p>Global augmentation (apply same aug for all frames), including rotation(-10,10), shift(-0.1,0.1), scale(0.8,1.2), shear(-1.0,1.0), flip(apply for some signs)</p></li>\n<li><p>Time-based augmentation (apply aug for some frames), random select some frames(1-8) do affine augmentations, random drop frames (fill with 0.0)</p></li>\n<li><p>The augmentations greatly improve the score</p></li>\n</ul>\n<h2>4. HyperParameters</h2>\n<ul>\n<li>NUM_BLOCKS, 2</li>\n<li>NUM_HEAD 8</li>\n<li>lr 1e-3 </li>\n<li>Optimizer AdamW</li>\n<li>Epoch 100</li>\n<li>LateDropout 0.2-0.3</li>\n<li>Label Smoothing 0.5</li>\n</ul>\n<p>Six conv1d models with TTA combined with two transformers gave us third place in this competition.</p>",
  "messages": [
    {
      "id": "2243318",
      "postDate": "05/02/2023 20:12:22",
      "content": "<p>We used an <strong>ensemble of six conv1d models and two versions of transformers</strong> based on the <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">public notebook</a>. The key points are <strong>data preprocessing, hard augmentation and ensemble</strong>.</p>\n<p>I initially tried to develop my own solution based on the aforementioned public transformer but then I found that the architecture of a model didn’t matter as much as proper work with data and ensembling, so I switched to more simple architectures like <strong>conv1d multilayer models</strong>. I used <code>Keras</code> and utilized all the advantages of native tflite layers like <code>DepthwiseConv1D</code>. My typical model looked like that:</p>\n<pre><code>do = \n\nmodel = Sequential()\nmodel.add(InputLayer(input_shape=(max_len, , )))\nmodel.add(Reshape((max_len, *)))\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(MaxPool1D(, ))\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(, , strides=, padding=, activation=))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(, strides=, padding=, depth_multiplier=, activation=))\nmodel.add(BatchNormalization())\n\nmodel.add(GlobalAvgPool1D())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(, activation=))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(, activation=))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(, activation=))\n</code></pre>\n<p>Then two of my colleagues joined me. We had a lot of ideas such as using synthetic data. We found this <a href=\"https://upcommons.upc.edu/bitstream/handle/2117/94313/TEPP1de1.pdf\" target=\"_blank\">paper</a>, created a script which generated synthetic 3D points of a hand parametrized by inner parameters (joints rotations) and camera model and tried to predict the inner parameters of hand by a model, based on the 3d points. And then use this pretrained model to preprocess data. But unfortunately we didn’t have enough time to finish it.</p>\n<p>Also we tried to train on additional data (WLASL), found that this dataset gave ~0.005-0.007 increase in local CV score but then we abandoned this idea and haven’t used additional data in our LB submissions.</p>\n<h2>Normalization</h2>\n<p>We used normalization by this reference points <code>ref = coords[:, [500, 501, 512, 513, 159,  386, 13,]]</code><br>\nWe didn’t use depth dimension.</p>\n<p>We didn’t throw away frames without hands but we took the presence of a hand into account during TTA. </p>\n<h2>Augmentation</h2>\n<ul>\n<li>Random rotation, shift, scale applied globally to each frame</li>\n<li>Random small shift added to every point </li>\n<li>Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score</li>\n<li>CutMix - took parts of samples from different classes. For hand we used 0.7 as a label value, for other parts - 0.3</li>\n</ul>\n<h2>TTA</h2>\n<ul>\n<li>Random left/right padding for short sequences</li>\n<li>For large sequences of frames we throw away frames with different probability taking into account presence of hands on a frame</li>\n</ul>\n<p>For different models we used different combinations of points: both hands / only active hand, lips, eyes, top part of pose. For some models we used only every second point of lips/eyes.<br>\nFor models where we used only one hand we determined which hand is more presented in the video and if it was the left hand we mirrored all points and their coordinates.<br>\nWe took only <strong>32 frames</strong> for all models except for one where we used 96 frames.</p>\n<p>These conv1d models were fast and lightweight so we could ensemble up to six models and still had space and time for additional models and TTA. Ensemble of six such models got <strong>0.7948</strong> on public LB and <strong>0.8711</strong> on private LB.<br>\nThen we decided to collaborate with <a href=\"https://www.kaggle.com/sqqqqy\" target=\"_blank\">@sqqqqy</a> who had been working on improving of transformer from a public kernel. His solution looked like that:</p>\n<h2>1. Transformer Model</h2>\n<ul>\n<li>Our transformer model is based on the <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">public notebook</a>, using different sequence normalization and smaller UNITS size to reduce model parameters.</li>\n<li>We implement two kinds of transformer-based model, the first model learn sperate embeddings for each part (just like public notebook), the second learn one embedding with input whole xyz sequence.</li>\n<li>1seed of first model score <strong>LB 0.77+</strong>, 1seed of second model score <strong>LB 0.768</strong></li>\n</ul>\n<h2>2. Preprocessing</h2>\n<ul>\n<li>We use 20 lip points, 32 eyes points, 42 hands points(left hand and right hand) and 8 pose points.</li>\n<li>The input sequence is normalized with shoulder, hip, lip and eyes points.</li>\n<li>Filling the NaN values with 0.0</li>\n<li>Learn a motion embedding by input $(d_x, d_y, \\sqrt{(d_x)^2+(d_y)^2})<em>t$ sequence, \n$$\n(d_x, d_y)_t = xyz_t - xyz</em>{t-1}<br>\n$$</li>\n<li>The final embedding is the concat of motion embedding and xyz embedding</li>\n</ul>\n<h2>3.Augmentation</h2>\n<ul>\n<li><p>Global augmentation (apply same aug for all frames), including rotation(-10,10), shift(-0.1,0.1), scale(0.8,1.2), shear(-1.0,1.0), flip(apply for some signs)</p></li>\n<li><p>Time-based augmentation (apply aug for some frames), random select some frames(1-8) do affine augmentations, random drop frames (fill with 0.0)</p></li>\n<li><p>The augmentations greatly improve the score</p></li>\n</ul>\n<h2>4. HyperParameters</h2>\n<ul>\n<li>NUM_BLOCKS, 2</li>\n<li>NUM_HEAD 8</li>\n<li>lr 1e-3 </li>\n<li>Optimizer AdamW</li>\n<li>Epoch 100</li>\n<li>LateDropout 0.2-0.3</li>\n<li>Label Smoothing 0.5</li>\n</ul>\n<p>Six conv1d models with TTA combined with two transformers gave us third place in this competition.</p>",
      "rawMarkdown": "We used an **ensemble of six conv1d models and two versions of transformers** based on the [public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training). The key points are **data preprocessing, hard augmentation and ensemble**.\n\nI initially tried to develop my own solution based on the aforementioned public transformer but then I found that the architecture of a model didn’t matter as much as proper work with data and ensembling, so I switched to more simple architectures like **conv1d multilayer models**. I used `Keras` and utilized all the advantages of native tflite layers like `DepthwiseConv1D`. My typical model looked like that:\n```python\ndo = 0.5\n\nmodel = Sequential()\nmodel.add(InputLayer(input_shape=(max_len, 61, 2)))\nmodel.add(Reshape((max_len, 61*2)))\n\nmodel.add(Conv1D(64, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(3, strides=1, padding='valid', depth_multiplier=1, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(64, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(5, strides=2, padding='valid', depth_multiplier=4, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(MaxPool1D(2, 2))\n\nmodel.add(Conv1D(256, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(3, strides=1, padding='valid', depth_multiplier=1, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(256, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(5, strides=2, padding='valid', depth_multiplier=4, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(GlobalAvgPool1D())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(1024, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(1024, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(250, activation='softmax'))\n```\n\nThen two of my colleagues joined me. We had a lot of ideas such as using synthetic data. We found this [paper](https://upcommons.upc.edu/bitstream/handle/2117/94313/TEPP1de1.pdf), created a script which generated synthetic 3D points of a hand parametrized by inner parameters (joints rotations) and camera model and tried to predict the inner parameters of hand by a model, based on the 3d points. And then use this pretrained model to preprocess data. But unfortunately we didn’t have enough time to finish it.\n\nAlso we tried to train on additional data (WLASL), found that this dataset gave ~0.005-0.007 increase in local CV score but then we abandoned this idea and haven’t used additional data in our LB submissions.\n\n## Normalization\nWe used normalization by this reference points `ref = coords[:, [500, 501, 512, 513, 159,  386, 13,]]`\nWe didn’t use depth dimension.\n\nWe didn’t throw away frames without hands but we took the presence of a hand into account during TTA. \n\n## Augmentation\n- Random rotation, shift, scale applied globally to each frame\n- Random small shift added to every point \n- Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score\n- CutMix - took parts of samples from different classes. For hand we used 0.7 as a label value, for other parts - 0.3\n\n## TTA\n- Random left/right padding for short sequences\n- For large sequences of frames we throw away frames with different probability taking into account presence of hands on a frame\n\nFor different models we used different combinations of points: both hands / only active hand, lips, eyes, top part of pose. For some models we used only every second point of lips/eyes.\nFor models where we used only one hand we determined which hand is more presented in the video and if it was the left hand we mirrored all points and their coordinates.\nWe took only **32 frames** for all models except for one where we used 96 frames.\n\nThese conv1d models were fast and lightweight so we could ensemble up to six models and still had space and time for additional models and TTA. Ensemble of six such models got **0.7948** on public LB and **0.8711** on private LB.\nThen we decided to collaborate with @sqqqqy who had been working on improving of transformer from a public kernel. His solution looked like that:\n\n## 1. Transformer Model\n- Our transformer model is based on the [public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training), using different sequence normalization and smaller UNITS size to reduce model parameters.\n- We implement two kinds of transformer-based model, the first model learn sperate embeddings for each part (just like public notebook), the second learn one embedding with input whole xyz sequence.\n- 1seed of first model score **LB 0.77+**, 1seed of second model score **LB 0.768**\n\n## 2. Preprocessing\n- We use 20 lip points, 32 eyes points, 42 hands points(left hand and right hand) and 8 pose points.\n- The input sequence is normalized with shoulder, hip, lip and eyes points.\n- Filling the NaN values with 0.0\n- Learn a motion embedding by input $(d_x, d_y, \\sqrt{(d_x)^2+(d_y)^2})_t$ sequence, \n$$\n(d_x, d_y)_t = xyz_t - xyz_{t-1}\n$$\n- The final embedding is the concat of motion embedding and xyz embedding\n\n## 3.Augmentation\n- Global augmentation (apply same aug for all frames), including rotation(-10,10), shift(-0.1,0.1), scale(0.8,1.2), shear(-1.0,1.0), flip(apply for some signs)\n\n- Time-based augmentation (apply aug for some frames), random select some frames(1-8) do affine augmentations, random drop frames (fill with 0.0)\n\n- The augmentations greatly improve the score\n\n## 4. HyperParameters\n- NUM_BLOCKS, 2\n- NUM_HEAD 8\n- lr 1e-3 \n- Optimizer AdamW\n- Epoch 100\n- LateDropout 0.2-0.3\n- Label Smoothing 0.5\n\n\nSix conv1d models with TTA combined with two transformers gave us third place in this competition.",
      "votes": null
    },
    {
      "id": "2243321",
      "postDate": "05/02/2023 20:16:20",
      "content": "<p>Congratulations! I like your data augmentation</p>\n<blockquote>\n  <p>Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score</p>\n</blockquote>\n<p>That is very creative augmentation! </p>",
      "rawMarkdown": "Congratulations! I like your data augmentation\n>Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score\n\nThat is very creative augmentation!",
      "votes": null
    },
    {
      "id": "2243327",
      "postDate": "05/02/2023 20:24:08",
      "content": "<p>It was surprising for me that it worked. Because we didn't align coordinates of points from different parts so that it resemble a real pose.</p>",
      "rawMarkdown": "It was surprising for me that it worked. Because we didn't align coordinates of points from different parts so that it resemble a real pose.",
      "votes": null
    },
    {
      "id": "2243467",
      "postDate": "05/02/2023 23:49:33",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": null
    },
    {
      "id": "2246927",
      "postDate": "05/05/2023 14:50:10",
      "content": "<p>Congrats on third place!!<br>\nUsing synthetic data is a very cool idea. I don't know why I didn't think of this before</p>",
      "rawMarkdown": "Congrats on third place!!\nUsing synthetic data is a very cool idea. I don't know why I didn't think of this before",
      "votes": null
    },
    {
      "id": "2269532",
      "postDate": "05/22/2023 14:20:05",
      "content": "<p>Congratulations on 3rd place and thanks for sharing the solution <a href=\"https://www.kaggle.com/mrnnnn\" target=\"_blank\">@mrnnnn</a> </p>",
      "rawMarkdown": "Congratulations on 3rd place and thanks for sharing the solution @mrnnnn",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2243321,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/02/2023 20:16:20",
      "content": "<p>Congratulations! I like your data augmentation</p>\n<blockquote>\n  <p>Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score</p>\n</blockquote>\n<p>That is very creative augmentation! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2243327,
          "author_name": "mrnnnn",
          "author_url": "",
          "post_date": "05/02/2023 20:24:08",
          "content": "<p>It was surprising for me that it worked. Because we didn't align coordinates of points from different parts so that it resemble a real pose.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2243467,
      "author_name": "artemtprv",
      "author_url": "",
      "post_date": "05/02/2023 23:49:33",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2246927,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/05/2023 14:50:10",
      "content": "<p>Congrats on third place!!<br>\nUsing synthetic data is a very cool idea. I don't know why I didn't think of this before</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2269532,
      "author_name": "tamannaakterswarna",
      "author_url": "",
      "post_date": "05/22/2023 14:20:05",
      "content": "<p>Congratulations on 3rd place and thanks for sharing the solution <a href=\"https://www.kaggle.com/mrnnnn\" target=\"_blank\">@mrnnnn</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2243318": "We used an **ensemble of six conv1d models and two versions of transformers** based on the [public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training). The key points are **data preprocessing, hard augmentation and ensemble**.\n\nI initially tried to develop my own solution based on the aforementioned public transformer but then I found that the architecture of a model didn’t matter as much as proper work with data and ensembling, so I switched to more simple architectures like **conv1d multilayer models**. I used `Keras` and utilized all the advantages of native tflite layers like `DepthwiseConv1D`. My typical model looked like that:\n```python\ndo = 0.5\n\nmodel = Sequential()\nmodel.add(InputLayer(input_shape=(max_len, 61, 2)))\nmodel.add(Reshape((max_len, 61*2)))\n\nmodel.add(Conv1D(64, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(3, strides=1, padding='valid', depth_multiplier=1, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(64, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(5, strides=2, padding='valid', depth_multiplier=4, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(MaxPool1D(2, 2))\n\nmodel.add(Conv1D(256, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(3, strides=1, padding='valid', depth_multiplier=1, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(Conv1D(256, 1, strides=1, padding='valid', activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(DepthwiseConv1D(5, strides=2, padding='valid', depth_multiplier=4, activation='relu'))\nmodel.add(BatchNormalization())\n\nmodel.add(GlobalAvgPool1D())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(1024, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(1024, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(rate=do))\n\nmodel.add(Dense(250, activation='softmax'))\n```\n\nThen two of my colleagues joined me. We had a lot of ideas such as using synthetic data. We found this [paper](https://upcommons.upc.edu/bitstream/handle/2117/94313/TEPP1de1.pdf), created a script which generated synthetic 3D points of a hand parametrized by inner parameters (joints rotations) and camera model and tried to predict the inner parameters of hand by a model, based on the 3d points. And then use this pretrained model to preprocess data. But unfortunately we didn’t have enough time to finish it.\n\nAlso we tried to train on additional data (WLASL), found that this dataset gave ~0.005-0.007 increase in local CV score but then we abandoned this idea and haven’t used additional data in our LB submissions.\n\n## Normalization\nWe used normalization by this reference points `ref = coords[:, [500, 501, 512, 513, 159,  386, 13,]]`\nWe didn’t use depth dimension.\n\nWe didn’t throw away frames without hands but we took the presence of a hand into account during TTA. \n\n## Augmentation\n- Random rotation, shift, scale applied globally to each frame\n- Random small shift added to every point \n- Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score\n- CutMix - took parts of samples from different classes. For hand we used 0.7 as a label value, for other parts - 0.3\n\n## TTA\n- Random left/right padding for short sequences\n- For large sequences of frames we throw away frames with different probability taking into account presence of hands on a frame\n\nFor different models we used different combinations of points: both hands / only active hand, lips, eyes, top part of pose. For some models we used only every second point of lips/eyes.\nFor models where we used only one hand we determined which hand is more presented in the video and if it was the left hand we mirrored all points and their coordinates.\nWe took only **32 frames** for all models except for one where we used 96 frames.\n\nThese conv1d models were fast and lightweight so we could ensemble up to six models and still had space and time for additional models and TTA. Ensemble of six such models got **0.7948** on public LB and **0.8711** on private LB.\nThen we decided to collaborate with @sqqqqy who had been working on improving of transformer from a public kernel. His solution looked like that:\n\n## 1. Transformer Model\n- Our transformer model is based on the [public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training), using different sequence normalization and smaller UNITS size to reduce model parameters.\n- We implement two kinds of transformer-based model, the first model learn sperate embeddings for each part (just like public notebook), the second learn one embedding with input whole xyz sequence.\n- 1seed of first model score **LB 0.77+**, 1seed of second model score **LB 0.768**\n\n## 2. Preprocessing\n- We use 20 lip points, 32 eyes points, 42 hands points(left hand and right hand) and 8 pose points.\n- The input sequence is normalized with shoulder, hip, lip and eyes points.\n- Filling the NaN values with 0.0\n- Learn a motion embedding by input $(d_x, d_y, \\sqrt{(d_x)^2+(d_y)^2})_t$ sequence, \n$$\n(d_x, d_y)_t = xyz_t - xyz_{t-1}\n$$\n- The final embedding is the concat of motion embedding and xyz embedding\n\n## 3.Augmentation\n- Global augmentation (apply same aug for all frames), including rotation(-10,10), shift(-0.1,0.1), scale(0.8,1.2), shear(-1.0,1.0), flip(apply for some signs)\n\n- Time-based augmentation (apply aug for some frames), random select some frames(1-8) do affine augmentations, random drop frames (fill with 0.0)\n\n- The augmentations greatly improve the score\n\n## 4. HyperParameters\n- NUM_BLOCKS, 2\n- NUM_HEAD 8\n- lr 1e-3 \n- Optimizer AdamW\n- Epoch 100\n- LateDropout 0.2-0.3\n- Label Smoothing 0.5\n\n\nSix conv1d models with TTA combined with two transformers gave us third place in this competition.",
    "2243321": "Congratulations! I like your data augmentation\n>Combinations of parts - for example we took a hand from one sample and lips from another sample with the same label. This augmentation gave a significant increase in score\n\nThat is very creative augmentation!",
    "2243327": "It was surprising for me that it worked. Because we didn't align coordinates of points from different parts so that it resemble a real pose.",
    "2243467": "Congratulations and thanks for sharing!",
    "2246927": "Congrats on third place!!\nUsing synthetic data is a very cool idea. I don't know why I didn't think of this before",
    "2269532": "Congratulations on 3rd place and thanks for sharing the solution @mrnnnn"
  },
  "source": "meta"
}