{
  "id": 406356,
  "title": " 5x th place solutions (Silver) - Single model and hand made feature approach.(public : 0.84469)",
  "url": "/competitions/asl-signs/discussion/406356",
  "author_name": "",
  "post_date": "2023-05-02T04:05:34.990346900Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<ul>\n<li>Update<br>\nInformation sharing to our member just before the team merge has led to deprivation. <br>\nWe will be careful from now on.<br>\nIm sorry for all.</li>\n</ul>\n<hr>\n<p>Thanks to Kaggle for hosting this interesting competition!!!!<br>\nVery enjoyable competition!</p>\n<h1>Team Member</h1>\n<p><a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a>, <a href=\"https://www.kaggle.com/hatakee\" target=\"_blank\">@hatakee</a>, <a href=\"https://www.kaggle.com/kfuji\" target=\"_blank\">@kfuji</a><br>\nco-workers!!</p>\n<h2>SUMMARY</h2>\n<ul>\n<li>My solution is based on this notebook. Thank you <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a><ul>\n<li>Link : <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training</a></li></ul></li>\n<li>The changes made are as follows:<ol>\n<li>Added more features. Add features based on the following ideas.<ul>\n<li>Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)</li></ul></li>\n<li>Applied coordinate normalization processing for each frame.</li>\n<li>Changed the parameters(epoch, MLP ratio, dropout).</li></ol></li>\n<li>Our code is in Appendix.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F2e070943f27cf2c0eb206ef8e258a0ec%2F3.png?generation=1683000565860768&amp;alt=media\" alt=\"\"></p>\n<h2>Scores transitions</h2>\n<table>\n<thead>\n<tr>\n<th>No.</th>\n<th>Modifications</th>\n<th>CV</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>Baseline</td>\n<td>0.8327</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>Augmentation</td>\n<td>0.8338</td>\n<td></td>\n</tr>\n<tr>\n<td>3</td>\n<td>Add feature (Vector)</td>\n<td>0.8367</td>\n<td></td>\n</tr>\n<tr>\n<td>4</td>\n<td>Add feature (velocity, distance)</td>\n<td>0.8408</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>Add feature (acceleration, angle)</td>\n<td>0.8444</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>Change MLP ratio</td>\n<td>0.8496</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>Normalize position</td>\n<td>0.8518</td>\n<td></td>\n</tr>\n<tr>\n<td>7</td>\n<td>Epoch 300</td>\n<td>0.8575</td>\n<td></td>\n</tr>\n<tr>\n<td>8</td>\n<td>Add feature (shape)</td>\n<td>0.8643</td>\n<td></td>\n</tr>\n<tr>\n<td>9</td>\n<td>Use All data (epoch100)</td>\n<td>----</td>\n<td>0.85</td>\n</tr>\n</tbody>\n</table>\n<h1>Progress</h1>\n<ul>\n<li>Early Stages<ul>\n<li>What was done<ul>\n<li>Understanding the competition</li>\n<li>Understanding the baseline</li>\n<li>Studying transformers</li></ul></li></ul></li>\n<li>Middle Stages<ul>\n<li>What was done<ul>\n<li>Changing network parameters</li>\n<li>Implementing discussions</li>\n<li>random split validation</li></ul></li>\n<li>Insights gained in the middle stage<ul>\n<li>Realized that adding features is key, from studying the basics and experimenting</li></ul></li></ul></li>\n<li>Final Stages<ul>\n<li>What was done<ul>\n<li>Adding features</li>\n<li>Adjusting parameters</li>\n<li>all data training</li></ul></li></ul></li>\n</ul>\n<h2>Preprocessing</h2>\n<ul>\n<li>Adjust all coordinates so that the center of numbers 11 and 12 is set to 0.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Ffa394dbd96e33fb875aef6ed1f25c757%2F1.png?generation=1682987922298156&amp;alt=media\" alt=\"\"></li>\n<li>Agumentations (it runs outside of training loop).<ul>\n<li>NaN interpolation</li>\n<li>3D scaling</li>\n<li>time direction scaling</li></ul></li>\n</ul>\n<p>Augmentation layer is in Appendix.</p>\n<p>In our case, we added more features. Add features based on the following ideas.</p>\n<ul>\n<li>Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)</li>\n</ul>\n<p>Here is our features.</p>\n<ul>\n<li>lip,body,hand : position, distance, velocity, accelaration, Angle, Angle velocity</li>\n<li>body,hand : Shape (Calculate below distance.)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb739316822c577ef9b4be204325f569e%2F2.png?generation=1682987881314581&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Modeling</h2>\n<p>if you want to know the details, please access to the base notebook.</p>\n<ul>\n<li>Network:<ul>\n<li>MLP ratio : 4</li>\n<li>Embeddings : 384</li>\n<li>Units : 512</li>\n<li>Transfomer Block : 2</li></ul></li>\n<li>Input <ul>\n<li>input size : 32</li></ul></li>\n</ul>\n<h2>CV</h2>\n<ul>\n<li>random split(8:2). <ul>\n<li>Seed 4949 is the best!!</li></ul></li>\n<li>The final submission was using all data</li>\n</ul>\n<h2>other</h2>\n<ul>\n<li>Chatgpt(GPT-4) was very helpful for writing code.</li>\n</ul>\n<h2>not worked for me</h2>\n<ul>\n<li>Augmentation<ul>\n<li>Local affine</li>\n<li>Noise</li></ul></li>\n<li>Bigger Paramer<ul>\n<li>MLP Ratio &gt; 4</li>\n<li>Epoch &gt; 300</li>\n<li>Length &gt; 32</li></ul></li>\n</ul>\n<h2>Appendix:</h2>\n<h3>Augmentation layer code.</h3>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n        (NanInterpolation, self).__init__(**kwargs)\n        self.order = \n        self.limit = \n\n     ():\n         training:\n            \n            data = inputs.numpy()\n\n            \n            interpolated_data = []\n             i  (data.shape[-]):\n                df = pd.DataFrame(data[..., i])\n                \n                \n                df = df.interpolate(limit_direction=)\n                \n                \n                interpolated_data.append(df.to_numpy())\n\n            \n            result = np.stack(interpolated_data, axis=-)\n            inputs = tf.convert_to_tensor(result, dtype=inputs.dtype)\n\n         inputs\n\n\n (tf.keras.layers.Layer):\n     ():\n        (Scaling3D, self).__init__(**kwargs)\n        self.scale_range = scale_range\n\n     ():\n         training:\n            \n            scale_factor = tf.random.uniform(\n                (), minval=self.scale_range[], maxval=self.scale_range[]\n            )\n\n            \n            inputs = inputs * scale_factor\n\n         inputs \n\n (tf.keras.layers.Layer):\n     ():\n        (TimeSeriesAugmentation, self).__init__(**kwargs)\n        self.framerate_factor_range = framerate_factor_range\n\n     ():\n         training:\n            \n            framerate_factor = tf.random.uniform(\n                (), minval=self.framerate_factor_range[], maxval=self.framerate_factor_range[]\n            )\n            new_length = tf.cast(tf.cast(tf.shape(inputs)[], tf.float32) * framerate_factor, tf.int32)\n            inputs_expanded = tf.expand_dims(inputs, axis=)\n            resized_inputs = tf.image.resize(inputs_expanded, (new_length, tf.shape(inputs)[-]))\n            inputs = resized_inputs[]\n\n         inputs\n</code></pre>",
  "messages": [
    {
      "id": "2242140",
      "postDate": "05/02/2023 04:05:34",
      "content": "<ul>\n<li>Update<br>\nInformation sharing to our member just before the team merge has led to deprivation. <br>\nWe will be careful from now on.<br>\nIm sorry for all.</li>\n</ul>\n<hr>\n<p>Thanks to Kaggle for hosting this interesting competition!!!!<br>\nVery enjoyable competition!</p>\n<h1>Team Member</h1>\n<p><a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a>, <a href=\"https://www.kaggle.com/hatakee\" target=\"_blank\">@hatakee</a>, <a href=\"https://www.kaggle.com/kfuji\" target=\"_blank\">@kfuji</a><br>\nco-workers!!</p>\n<h2>SUMMARY</h2>\n<ul>\n<li>My solution is based on this notebook. Thank you <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a><ul>\n<li>Link : <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training</a></li></ul></li>\n<li>The changes made are as follows:<ol>\n<li>Added more features. Add features based on the following ideas.<ul>\n<li>Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)</li></ul></li>\n<li>Applied coordinate normalization processing for each frame.</li>\n<li>Changed the parameters(epoch, MLP ratio, dropout).</li></ol></li>\n<li>Our code is in Appendix.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F2e070943f27cf2c0eb206ef8e258a0ec%2F3.png?generation=1683000565860768&amp;alt=media\" alt=\"\"></p>\n<h2>Scores transitions</h2>\n<table>\n<thead>\n<tr>\n<th>No.</th>\n<th>Modifications</th>\n<th>CV</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>Baseline</td>\n<td>0.8327</td>\n<td></td>\n</tr>\n<tr>\n<td>2</td>\n<td>Augmentation</td>\n<td>0.8338</td>\n<td></td>\n</tr>\n<tr>\n<td>3</td>\n<td>Add feature (Vector)</td>\n<td>0.8367</td>\n<td></td>\n</tr>\n<tr>\n<td>4</td>\n<td>Add feature (velocity, distance)</td>\n<td>0.8408</td>\n<td></td>\n</tr>\n<tr>\n<td>5</td>\n<td>Add feature (acceleration, angle)</td>\n<td>0.8444</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>Change MLP ratio</td>\n<td>0.8496</td>\n<td></td>\n</tr>\n<tr>\n<td>6</td>\n<td>Normalize position</td>\n<td>0.8518</td>\n<td></td>\n</tr>\n<tr>\n<td>7</td>\n<td>Epoch 300</td>\n<td>0.8575</td>\n<td></td>\n</tr>\n<tr>\n<td>8</td>\n<td>Add feature (shape)</td>\n<td>0.8643</td>\n<td></td>\n</tr>\n<tr>\n<td>9</td>\n<td>Use All data (epoch100)</td>\n<td>----</td>\n<td>0.85</td>\n</tr>\n</tbody>\n</table>\n<h1>Progress</h1>\n<ul>\n<li>Early Stages<ul>\n<li>What was done<ul>\n<li>Understanding the competition</li>\n<li>Understanding the baseline</li>\n<li>Studying transformers</li></ul></li></ul></li>\n<li>Middle Stages<ul>\n<li>What was done<ul>\n<li>Changing network parameters</li>\n<li>Implementing discussions</li>\n<li>random split validation</li></ul></li>\n<li>Insights gained in the middle stage<ul>\n<li>Realized that adding features is key, from studying the basics and experimenting</li></ul></li></ul></li>\n<li>Final Stages<ul>\n<li>What was done<ul>\n<li>Adding features</li>\n<li>Adjusting parameters</li>\n<li>all data training</li></ul></li></ul></li>\n</ul>\n<h2>Preprocessing</h2>\n<ul>\n<li>Adjust all coordinates so that the center of numbers 11 and 12 is set to 0.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Ffa394dbd96e33fb875aef6ed1f25c757%2F1.png?generation=1682987922298156&amp;alt=media\" alt=\"\"></li>\n<li>Agumentations (it runs outside of training loop).<ul>\n<li>NaN interpolation</li>\n<li>3D scaling</li>\n<li>time direction scaling</li></ul></li>\n</ul>\n<p>Augmentation layer is in Appendix.</p>\n<p>In our case, we added more features. Add features based on the following ideas.</p>\n<ul>\n<li>Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)</li>\n</ul>\n<p>Here is our features.</p>\n<ul>\n<li>lip,body,hand : position, distance, velocity, accelaration, Angle, Angle velocity</li>\n<li>body,hand : Shape (Calculate below distance.)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb739316822c577ef9b4be204325f569e%2F2.png?generation=1682987881314581&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Modeling</h2>\n<p>if you want to know the details, please access to the base notebook.</p>\n<ul>\n<li>Network:<ul>\n<li>MLP ratio : 4</li>\n<li>Embeddings : 384</li>\n<li>Units : 512</li>\n<li>Transfomer Block : 2</li></ul></li>\n<li>Input <ul>\n<li>input size : 32</li></ul></li>\n</ul>\n<h2>CV</h2>\n<ul>\n<li>random split(8:2). <ul>\n<li>Seed 4949 is the best!!</li></ul></li>\n<li>The final submission was using all data</li>\n</ul>\n<h2>other</h2>\n<ul>\n<li>Chatgpt(GPT-4) was very helpful for writing code.</li>\n</ul>\n<h2>not worked for me</h2>\n<ul>\n<li>Augmentation<ul>\n<li>Local affine</li>\n<li>Noise</li></ul></li>\n<li>Bigger Paramer<ul>\n<li>MLP Ratio &gt; 4</li>\n<li>Epoch &gt; 300</li>\n<li>Length &gt; 32</li></ul></li>\n</ul>\n<h2>Appendix:</h2>\n<h3>Augmentation layer code.</h3>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n        (NanInterpolation, self).__init__(**kwargs)\n        self.order = \n        self.limit = \n\n     ():\n         training:\n            \n            data = inputs.numpy()\n\n            \n            interpolated_data = []\n             i  (data.shape[-]):\n                df = pd.DataFrame(data[..., i])\n                \n                \n                df = df.interpolate(limit_direction=)\n                \n                \n                interpolated_data.append(df.to_numpy())\n\n            \n            result = np.stack(interpolated_data, axis=-)\n            inputs = tf.convert_to_tensor(result, dtype=inputs.dtype)\n\n         inputs\n\n\n (tf.keras.layers.Layer):\n     ():\n        (Scaling3D, self).__init__(**kwargs)\n        self.scale_range = scale_range\n\n     ():\n         training:\n            \n            scale_factor = tf.random.uniform(\n                (), minval=self.scale_range[], maxval=self.scale_range[]\n            )\n\n            \n            inputs = inputs * scale_factor\n\n         inputs \n\n (tf.keras.layers.Layer):\n     ():\n        (TimeSeriesAugmentation, self).__init__(**kwargs)\n        self.framerate_factor_range = framerate_factor_range\n\n     ():\n         training:\n            \n            framerate_factor = tf.random.uniform(\n                (), minval=self.framerate_factor_range[], maxval=self.framerate_factor_range[]\n            )\n            new_length = tf.cast(tf.cast(tf.shape(inputs)[], tf.float32) * framerate_factor, tf.int32)\n            inputs_expanded = tf.expand_dims(inputs, axis=)\n            resized_inputs = tf.image.resize(inputs_expanded, (new_length, tf.shape(inputs)[-]))\n            inputs = resized_inputs[]\n\n         inputs\n</code></pre>",
      "rawMarkdown": "Update\nInformation sharing to our member just before the team merge has led to deprivation. \nWe will be careful from now on.\nIm sorry for all.\n\n---\n\nThanks to Kaggle for hosting this interesting competition!!!!\nVery enjoyable competition!\n\n# Team Member\n@sugupoko, @hatakee, @kfuji\nco-workers!!\n\n## SUMMARY\n- My solution is based on this notebook. Thank you @markwijkhuizen\n     - Link : https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\n- The changes made are as follows:\n    1. Added more features. Add features based on the following ideas.\n        - Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)\n    2. Applied coordinate normalization processing for each frame.\n    3. Changed the parameters(epoch, MLP ratio, dropout).\n- Our code is in Appendix.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F2e070943f27cf2c0eb206ef8e258a0ec%2F3.png?generation=1683000565860768&alt=media)\n\n## Scores transitions\n\n| No. | Modifications                | CV     |Private|\n|-----|------------------------------|--------|----|\n| 1   | Baseline                     | 0.8327 ||\n| 2   | Augmentation                 | 0.8338 | |\n| 3   | Add feature (Vector)         | 0.8367 ||\n| 4   | Add feature (velocity, distance) | 0.8408 ||\n| 5   | Add feature (acceleration, angle) | 0.8444 | |\n| 6   | Change MLP ratio             | 0.8496 | |\n| 6   | Normalize position           | 0.8518 ||\n| 7   | Epoch 300                    | 0.8575 ||\n| 8   | Add feature (shape)          | 0.8643 | |\n| 9   | Use All data (epoch100)      | ----   | 0.85     |\n\n# Progress\n- Early Stages\n    - What was done\n        - Understanding the competition\n        - Understanding the baseline\n        - Studying transformers\n- Middle Stages\n    - What was done\n        - Changing network parameters\n        - Implementing discussions\n        - random split validation\n    - Insights gained in the middle stage\n        - Realized that adding features is key, from studying the basics and experimenting\n- Final Stages\n    - What was done\n        - Adding features\n        - Adjusting parameters\n        - all data training\n\n\n## Preprocessing\n- Adjust all coordinates so that the center of numbers 11 and 12 is set to 0.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Ffa394dbd96e33fb875aef6ed1f25c757%2F1.png?generation=1682987922298156&alt=media)\n- Agumentations (it runs outside of training loop).\n    - NaN interpolation\n    - 3D scaling\n    - time direction scaling\n\nAugmentation layer is in Appendix.\n\nIn our case, we added more features. Add features based on the following ideas.\n- Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)\n\nHere is our features.\n- lip,body,hand : position, distance, velocity, accelaration, Angle, Angle velocity\n- body,hand : Shape (Calculate below distance.)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb739316822c577ef9b4be204325f569e%2F2.png?generation=1682987881314581&alt=media)\n\n\n## Modeling\n\nif you want to know the details, please access to the base notebook.\n- Network:\n  - MLP ratio : 4\n  - Embeddings : 384\n  - Units : 512\n  - Transfomer Block : 2\n- Input \n  - input size : 32\n\n## CV\n- random split(8:2). \n    - Seed 4949 is the best!!\n- The final submission was using all data\n\n## other\n- Chatgpt(GPT-4) was very helpful for writing code.\n\n## not worked for me\n- Augmentation\n    - Local affine\n    - Noise\n- Bigger Paramer\n    - MLP Ratio > 4\n    - Epoch > 300\n    - Length > 32\n\n\n## Appendix:\n### Augmentation layer code.\n\n``` python\nclass NanInterpolation(tf.keras.layers.Layer):\n    def __init__(self,  **kwargs):\n        super(NanInterpolation, self).__init__(**kwargs)\n        self.order = 3\n        self.limit = 3\n        \n    def call(self, inputs, training=False):\n        if training:\n            # 入力データをNumpy配列に変換\n            data = inputs.numpy()\n\n            # 補間処理\n            interpolated_data = []\n            for i in range(data.shape[-1]):\n                df = pd.DataFrame(data[..., i])\n                # df = df.interpolate(method=\"spline\", order=self.order, limit=self.limit, limit_direction='both')\n                # df = df.interpolate(method=\"spline\", order=self.order, limit_direction='both')\n                df = df.interpolate(limit_direction='both')\n                # df.fillna(method=\"ffill\", inplace=True)   \n                # df.fillna(method=\"bfill\", inplace=True)\n                interpolated_data.append(df.to_numpy())\n\n            # 補間後のデータをテンソルに変換\n            result = np.stack(interpolated_data, axis=-1)\n            inputs = tf.convert_to_tensor(result, dtype=inputs.dtype)\n            \n        return inputs\n\n    \nclass Scaling3D(tf.keras.layers.Layer):\n    def __init__(self, scale_range=(0.9, 1.1), **kwargs):\n        super(Scaling3D, self).__init__(**kwargs)\n        self.scale_range = scale_range\n\n    def call(self, inputs, training=False):\n        if training:\n            # ランダムなスケーリング係数を生成\n            scale_factor = tf.random.uniform(\n                (), minval=self.scale_range[0], maxval=self.scale_range[1]\n            )\n\n            # ポーズデータにスケーリング係数を適用\n            inputs = inputs * scale_factor\n\n        return inputs \n\nclass TimeSeriesAugmentation(tf.keras.layers.Layer):\n    def __init__(self, framerate_factor_range=(0.8, 1.2),  **kwargs):\n        super(TimeSeriesAugmentation, self).__init__(**kwargs)\n        self.framerate_factor_range = framerate_factor_range\n\n    def call(self, inputs, training=False):\n        if training:\n            # フレームレート変更\n            framerate_factor = tf.random.uniform(\n                (), minval=self.framerate_factor_range[0], maxval=self.framerate_factor_range[1]\n            )\n            new_length = tf.cast(tf.cast(tf.shape(inputs)[0], tf.float32) * framerate_factor, tf.int32)\n            inputs_expanded = tf.expand_dims(inputs, axis=0)\n            resized_inputs = tf.image.resize(inputs_expanded, (new_length, tf.shape(inputs)[-2]))\n            inputs = resized_inputs[0]\n\n        return inputs\n\n```",
      "votes": null
    },
    {
      "id": "2246986",
      "postDate": "05/05/2023 15:45:43",
      "content": "<p>Congratulations!!!<br>\nUsing the shape of the hands, the orientation of the hands, the movements, the position of the hands and the facial expressions is an amazing idea.</p>",
      "rawMarkdown": "Congratulations!!!\nUsing the shape of the hands, the orientation of the hands, the movements, the position of the hands and the facial expressions is an amazing idea.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2246986,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/05/2023 15:45:43",
      "content": "<p>Congratulations!!!<br>\nUsing the shape of the hands, the orientation of the hands, the movements, the position of the hands and the facial expressions is an amazing idea.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2242140": "Update\nInformation sharing to our member just before the team merge has led to deprivation. \nWe will be careful from now on.\nIm sorry for all.\n\n---\n\nThanks to Kaggle for hosting this interesting competition!!!!\nVery enjoyable competition!\n\n# Team Member\n@sugupoko, @hatakee, @kfuji\nco-workers!!\n\n## SUMMARY\n- My solution is based on this notebook. Thank you @markwijkhuizen\n     - Link : https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\n- The changes made are as follows:\n    1. Added more features. Add features based on the following ideas.\n        - Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)\n    2. Applied coordinate normalization processing for each frame.\n    3. Changed the parameters(epoch, MLP ratio, dropout).\n- Our code is in Appendix.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2F2e070943f27cf2c0eb206ef8e258a0ec%2F3.png?generation=1683000565860768&alt=media)\n\n## Scores transitions\n\n| No. | Modifications                | CV     |Private|\n|-----|------------------------------|--------|----|\n| 1   | Baseline                     | 0.8327 ||\n| 2   | Augmentation                 | 0.8338 | |\n| 3   | Add feature (Vector)         | 0.8367 ||\n| 4   | Add feature (velocity, distance) | 0.8408 ||\n| 5   | Add feature (acceleration, angle) | 0.8444 | |\n| 6   | Change MLP ratio             | 0.8496 | |\n| 6   | Normalize position           | 0.8518 ||\n| 7   | Epoch 300                    | 0.8575 ||\n| 8   | Add feature (shape)          | 0.8643 | |\n| 9   | Use All data (epoch100)      | ----   | 0.85     |\n\n# Progress\n- Early Stages\n    - What was done\n        - Understanding the competition\n        - Understanding the baseline\n        - Studying transformers\n- Middle Stages\n    - What was done\n        - Changing network parameters\n        - Implementing discussions\n        - random split validation\n    - Insights gained in the middle stage\n        - Realized that adding features is key, from studying the basics and experimenting\n- Final Stages\n    - What was done\n        - Adding features\n        - Adjusting parameters\n        - all data training\n\n\n## Preprocessing\n- Adjust all coordinates so that the center of numbers 11 and 12 is set to 0.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Ffa394dbd96e33fb875aef6ed1f25c757%2F1.png?generation=1682987922298156&alt=media)\n- Agumentations (it runs outside of training loop).\n    - NaN interpolation\n    - 3D scaling\n    - time direction scaling\n\nAugmentation layer is in Appendix.\n\nIn our case, we added more features. Add features based on the following ideas.\n- Sign language is constructed from five elements: handshapes, hand orientation, movements, hand positions, and facial expressions. (We Asked Japanease Sign Language Professionals.)\n\nHere is our features.\n- lip,body,hand : position, distance, velocity, accelaration, Angle, Angle velocity\n- body,hand : Shape (Calculate below distance.)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930242%2Fb739316822c577ef9b4be204325f569e%2F2.png?generation=1682987881314581&alt=media)\n\n\n## Modeling\n\nif you want to know the details, please access to the base notebook.\n- Network:\n  - MLP ratio : 4\n  - Embeddings : 384\n  - Units : 512\n  - Transfomer Block : 2\n- Input \n  - input size : 32\n\n## CV\n- random split(8:2). \n    - Seed 4949 is the best!!\n- The final submission was using all data\n\n## other\n- Chatgpt(GPT-4) was very helpful for writing code.\n\n## not worked for me\n- Augmentation\n    - Local affine\n    - Noise\n- Bigger Paramer\n    - MLP Ratio > 4\n    - Epoch > 300\n    - Length > 32\n\n\n## Appendix:\n### Augmentation layer code.\n\n``` python\nclass NanInterpolation(tf.keras.layers.Layer):\n    def __init__(self,  **kwargs):\n        super(NanInterpolation, self).__init__(**kwargs)\n        self.order = 3\n        self.limit = 3\n        \n    def call(self, inputs, training=False):\n        if training:\n            # 入力データをNumpy配列に変換\n            data = inputs.numpy()\n\n            # 補間処理\n            interpolated_data = []\n            for i in range(data.shape[-1]):\n                df = pd.DataFrame(data[..., i])\n                # df = df.interpolate(method=\"spline\", order=self.order, limit=self.limit, limit_direction='both')\n                # df = df.interpolate(method=\"spline\", order=self.order, limit_direction='both')\n                df = df.interpolate(limit_direction='both')\n                # df.fillna(method=\"ffill\", inplace=True)   \n                # df.fillna(method=\"bfill\", inplace=True)\n                interpolated_data.append(df.to_numpy())\n\n            # 補間後のデータをテンソルに変換\n            result = np.stack(interpolated_data, axis=-1)\n            inputs = tf.convert_to_tensor(result, dtype=inputs.dtype)\n            \n        return inputs\n\n    \nclass Scaling3D(tf.keras.layers.Layer):\n    def __init__(self, scale_range=(0.9, 1.1), **kwargs):\n        super(Scaling3D, self).__init__(**kwargs)\n        self.scale_range = scale_range\n\n    def call(self, inputs, training=False):\n        if training:\n            # ランダムなスケーリング係数を生成\n            scale_factor = tf.random.uniform(\n                (), minval=self.scale_range[0], maxval=self.scale_range[1]\n            )\n\n            # ポーズデータにスケーリング係数を適用\n            inputs = inputs * scale_factor\n\n        return inputs \n\nclass TimeSeriesAugmentation(tf.keras.layers.Layer):\n    def __init__(self, framerate_factor_range=(0.8, 1.2),  **kwargs):\n        super(TimeSeriesAugmentation, self).__init__(**kwargs)\n        self.framerate_factor_range = framerate_factor_range\n\n    def call(self, inputs, training=False):\n        if training:\n            # フレームレート変更\n            framerate_factor = tf.random.uniform(\n                (), minval=self.framerate_factor_range[0], maxval=self.framerate_factor_range[1]\n            )\n            new_length = tf.cast(tf.cast(tf.shape(inputs)[0], tf.float32) * framerate_factor, tf.int32)\n            inputs_expanded = tf.expand_dims(inputs, axis=0)\n            resized_inputs = tf.image.resize(inputs_expanded, (new_length, tf.shape(inputs)[-2]))\n            inputs = resized_inputs[0]\n\n        return inputs\n\n```",
    "2246986": "Congratulations!!!\nUsing the shape of the hands, the orientation of the hands, the movements, the position of the hands and the facial expressions is an amazing idea."
  },
  "source": "meta"
}