{
  "id": 406346,
  "title": "Solution - Single transformer without val dataset",
  "url": "/competitions/asl-signs/writeups/solution-single-transformer-without-val-dataset",
  "author_name": "",
  "post_date": "2023-05-08T03:26:33.690Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to the organisers of the PopSign Games and all of the Kaggle community for running such a great competition.</p>\n<p>More importantly, my heartfelt thanks go to both <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>  and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. MARK contributed a very strong transformer framework, which is the machine learning pipeline for me in this competition. And I will find daily inspiration in <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s discussion post. </p>\n<p>As a learner, this competition has taken my machine learning skills a step further. Therefore, I would like to <strong>record the results of my experiment here</strong>.</p>\n<h1>Model Architecture</h1>\n<p>Shallow networks are the best choice.</p>\n<ul>\n<li>Single shallow transformer base (Refer to <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">version 5</a>)<ul>\n<li>Added <code>label smoothing</code>.</li>\n<li>Used warm up to <code>lr</code>.</li>\n<li>Remove all LayerNorm.</li></ul></li>\n</ul>\n<p>The main changes I have made that have led to improvements:</p>\n<ul>\n<li>Increase <code>Embedding</code> to 600.</li>\n<li>Add a <code>Gaussian noise (0.01)</code> to the FC layer of the Embedding layer.</li>\n<li>Increase <code>epoch</code> to 230</li>\n<li>Deepens the depth of the last FC layer in the transformer layer.</li>\n<li>Late Dropout. 0---0.3---0.1</li>\n</ul>\n<h1>Feature Engineering</h1>\n<p>Feature engineering was the most difficult part of my entire experiment. I came up with a lot of ideas, but most of them ended up being submission errors.</p>\n<p>Here is my summary of options that can bring about an uplift and are feasible.</p>\n<ul>\n<li><p>Use of <code>lips</code>, <code>eyes</code> (regardless of left or right eye), <code>main hand</code> (21 landmarks), and <code>main pose</code> (6 landmarks)</p>\n<ul>\n<li>Main hand means normalize the right hand to the left hand. The formula is <code>x = 1 - x</code>.</li>\n<li>The formula of main pose is <code>x = 1.2 -x</code>. </li></ul></li>\n<li><p>Set the input frame count <code>threshold</code> to <strong>48</strong>.</p>\n<ul>\n<li>Less than that is filled with 0.</li>\n<li>More than that is compressed using the <code>Bilinear Interpolation</code> method. (I think it only work well to data where the frame rate is a multiple of 48)</li>\n<li>Removal of data with only <code>one frame</code>.</li></ul></li>\n<li><p>Add <code>motion feature</code> for coordinate X.</p></li>\n<li><p>Add <code>Hand distance feature</code>. (Get inspired by <a href=\"https://www.kaggle.com/josephzahar\" target=\"_blank\">@josephzahar</a>'s <a href=\"https://www.kaggle.com/code/josephzahar/distance-angle-based-features-using-keras\" target=\"_blank\">post</a>)</p>\n<ul>\n<li>Calculate the Euclidean distance between the discontinuous joints of the hand.</li></ul></li>\n</ul>\n<p>Hence, the dimensionality of each data in my pipeline is 3 dimensions <code>(x, y, motion)</code>. For the Hand, it is 213<code>(x, y, motion, distance)</code>. </p>\n<h1>Augmentation</h1>\n<p>There is no need for mirror(flip) augmentation in this pipeline, as the hand is already normalised in the pre-processing.</p>\n<p>In augmentation, I only found rotation to be little useful for me. Specifically, I rotated the data with a small sample size (<code>data[-42:]</code>) by 15 degrees up and down. But I think my approach is very poor.</p>\n<h1>Future Work</h1>\n<ul>\n<li>How to distinguish similar words, like awake and wake?</li>\n<li>Solve the problem of <code>keras_model_converter.target_spec.supported_types = [tf.float16]</code> not working for me.</li>\n<li>How to do rotation correctly?</li>\n</ul>\n<h1>Summary of Best Solutions (In processing)</h1>\n<h3>\"How To Improve Best Public transformer basded Notebook\" from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></h3>\n<ul>\n<li><p>Original Post</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406302\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406302</a></li></ul></li>\n<li><p>Corresponding Notebook</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76</a></li></ul></li>\n<li><p>Keypoints: </p>\n<ul>\n<li>Reduce batch size and increse transformer block.</li>\n<li>Ensemble with different model seeds.</li></ul></li>\n<li><p>All changes</p>\n<ol>\n<li><p>Train 1 model =&gt; Train 4 models</p></li>\n<li><p>Add Time Scale augmentation</p></li>\n<li><p>Ensemble and apply TFLite FP16 quantization</p></li></ol>\n<p>Change the following parameters:</p>\n<ol>\n<li><p>INPUT_SIZE, 64 =&gt; 12</p></li>\n<li><p>BATCH_ALL_SIGNS_N, 4 =&gt; 1</p></li>\n<li><p>N_EPOCHS, 250 =&gt; 120</p></li>\n<li><p>LANDMARK_UNITS, 384 =&gt; 224</p></li>\n<li><p>UNITS, 512 =&gt; 376</p></li>\n<li><p>NUM_BLOCKS, 2 =&gt; 3</p></li>\n<li><p>MLP_RATIO, 4 =&gt; 3</p></li>\n<li><p>MLP_DROPOUT_RATIO, 0.40 =&gt; 0.30</p></li>\n<li><p>Remove random frame masking</p></li></ol></li>\n<li><p>The results of the experiments combined with those tricks</p>\n<ul>\n<li><p>Just use a-i parameters changes in my solution (Step i is necessary, otherwise it would be by gradient messages.). </p>\n<ul>\n<li>Public: 0.7525.      </li>\n<li>Private: 0.8340</li></ul></li>\n<li><p>Ensemble 2 seed models without quantization. (Surprised by this decline, I am looking into the reasons for it.)</p>\n<ul>\n<li>Ensemble Public: 0.7443    <ul>\n<li>Model0 Public: 0.73</li>\n<li>Model1 Public: 0.7314</li></ul></li>\n<li>Ensemble Private: 0.8276<ul>\n<li>Model0 Private: 0.814</li>\n<li>Model1 Private: 0.816</li></ul></li>\n<li><strong>I have found the reason. I change <code>lr</code> to <code>0.4*lr</code>, cause <code>BATCH_ALL_SIGNS_N</code> reduces from 4 to 1. This action was wrong, because I use rotation augmentation before training. In other words, all changing variables can affect the choice of <code>lr</code>. If those variables stay same but <code>batch_size</code>, We need to make the change in <code>batch_size</code> proportional to the change in <code>lr</code>.</strong> (Answer from this <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248253\" target=\"_blank\">post</a>)<br>\nAfter cancelling the change of <code>lr</code>. The result is shown below</li>\n<li>Ensemble Public: 0.7683   </li>\n<li>Ensemble Private: 0.8445</li></ul></li></ul></li>\n<li><p>To be resolved</p>\n<ul>\n<li><p>What is Time Scale augmentation? and  How to remove random frame masking?</p></li>\n<li><p>Why the learning rate is proportional to the batch size?</p>\n<ul>\n<li>Answer: This is correct with the other model parameters held constant. However, this is wrong if other variables (epoch, dropout rate, etc.) change.</li></ul></li>\n<li><p>Where this from \"Accurate local validation scheme which estimates LB score\" ?</p></li></ul></li>\n</ul>\n<h1>Rule of thumbs</h1>\n<h2>How to choose learning rate</h2>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook\" target=\"_blank\">https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook</a></p></li>\n<li><p>From this <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248264\" target=\"_blank\">post</a></p></li>\n</ul>\n<blockquote>\n  <p>If training requires only a few epochs (i.e. 1, 2, 3, 4, or 5 epochs) like NLP models, I usually just go with constant LR. Then i look at validation metric and if it increases in the last epoch or two, perhaps I will use <code>LR</code> for the first epochs and <code>LR * 0.1</code> (or 0.3) for the last epochs.</p>\n  <p>When a model requires many epochs like 20, 40, 80, 160. I usually go with cosine schedule (like this comp). These two techniques work well in most cases.</p>\n</blockquote>",
  "messages": [
    {
      "id": "2242097",
      "postDate": "05/02/2023 03:24:28",
      "content": "<p>Thanks to the organisers of the PopSign Games and all of the Kaggle community for running such a great competition.</p>\n<p>More importantly, my heartfelt thanks go to both <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>  and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. MARK contributed a very strong transformer framework, which is the machine learning pipeline for me in this competition. And I will find daily inspiration in <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s discussion post. </p>\n<p>As a learner, this competition has taken my machine learning skills a step further. Therefore, I would like to <strong>record the results of my experiment here</strong>.</p>\n<h1>Model Architecture</h1>\n<p>Shallow networks are the best choice.</p>\n<ul>\n<li>Single shallow transformer base (Refer to <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">version 5</a>)<ul>\n<li>Added <code>label smoothing</code>.</li>\n<li>Used warm up to <code>lr</code>.</li>\n<li>Remove all LayerNorm.</li></ul></li>\n</ul>\n<p>The main changes I have made that have led to improvements:</p>\n<ul>\n<li>Increase <code>Embedding</code> to 600.</li>\n<li>Add a <code>Gaussian noise (0.01)</code> to the FC layer of the Embedding layer.</li>\n<li>Increase <code>epoch</code> to 230</li>\n<li>Deepens the depth of the last FC layer in the transformer layer.</li>\n<li>Late Dropout. 0---0.3---0.1</li>\n</ul>\n<h1>Feature Engineering</h1>\n<p>Feature engineering was the most difficult part of my entire experiment. I came up with a lot of ideas, but most of them ended up being submission errors.</p>\n<p>Here is my summary of options that can bring about an uplift and are feasible.</p>\n<ul>\n<li><p>Use of <code>lips</code>, <code>eyes</code> (regardless of left or right eye), <code>main hand</code> (21 landmarks), and <code>main pose</code> (6 landmarks)</p>\n<ul>\n<li>Main hand means normalize the right hand to the left hand. The formula is <code>x = 1 - x</code>.</li>\n<li>The formula of main pose is <code>x = 1.2 -x</code>. </li></ul></li>\n<li><p>Set the input frame count <code>threshold</code> to <strong>48</strong>.</p>\n<ul>\n<li>Less than that is filled with 0.</li>\n<li>More than that is compressed using the <code>Bilinear Interpolation</code> method. (I think it only work well to data where the frame rate is a multiple of 48)</li>\n<li>Removal of data with only <code>one frame</code>.</li></ul></li>\n<li><p>Add <code>motion feature</code> for coordinate X.</p></li>\n<li><p>Add <code>Hand distance feature</code>. (Get inspired by <a href=\"https://www.kaggle.com/josephzahar\" target=\"_blank\">@josephzahar</a>'s <a href=\"https://www.kaggle.com/code/josephzahar/distance-angle-based-features-using-keras\" target=\"_blank\">post</a>)</p>\n<ul>\n<li>Calculate the Euclidean distance between the discontinuous joints of the hand.</li></ul></li>\n</ul>\n<p>Hence, the dimensionality of each data in my pipeline is 3 dimensions <code>(x, y, motion)</code>. For the Hand, it is 213<code>(x, y, motion, distance)</code>. </p>\n<h1>Augmentation</h1>\n<p>There is no need for mirror(flip) augmentation in this pipeline, as the hand is already normalised in the pre-processing.</p>\n<p>In augmentation, I only found rotation to be little useful for me. Specifically, I rotated the data with a small sample size (<code>data[-42:]</code>) by 15 degrees up and down. But I think my approach is very poor.</p>\n<h1>Future Work</h1>\n<ul>\n<li>How to distinguish similar words, like awake and wake?</li>\n<li>Solve the problem of <code>keras_model_converter.target_spec.supported_types = [tf.float16]</code> not working for me.</li>\n<li>How to do rotation correctly?</li>\n</ul>\n<h1>Summary of Best Solutions (In processing)</h1>\n<h3>\"How To Improve Best Public transformer basded Notebook\" from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></h3>\n<ul>\n<li><p>Original Post</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406302\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406302</a></li></ul></li>\n<li><p>Corresponding Notebook</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76</a></li></ul></li>\n<li><p>Keypoints: </p>\n<ul>\n<li>Reduce batch size and increse transformer block.</li>\n<li>Ensemble with different model seeds.</li></ul></li>\n<li><p>All changes</p>\n<ol>\n<li><p>Train 1 model =&gt; Train 4 models</p></li>\n<li><p>Add Time Scale augmentation</p></li>\n<li><p>Ensemble and apply TFLite FP16 quantization</p></li></ol>\n<p>Change the following parameters:</p>\n<ol>\n<li><p>INPUT_SIZE, 64 =&gt; 12</p></li>\n<li><p>BATCH_ALL_SIGNS_N, 4 =&gt; 1</p></li>\n<li><p>N_EPOCHS, 250 =&gt; 120</p></li>\n<li><p>LANDMARK_UNITS, 384 =&gt; 224</p></li>\n<li><p>UNITS, 512 =&gt; 376</p></li>\n<li><p>NUM_BLOCKS, 2 =&gt; 3</p></li>\n<li><p>MLP_RATIO, 4 =&gt; 3</p></li>\n<li><p>MLP_DROPOUT_RATIO, 0.40 =&gt; 0.30</p></li>\n<li><p>Remove random frame masking</p></li></ol></li>\n<li><p>The results of the experiments combined with those tricks</p>\n<ul>\n<li><p>Just use a-i parameters changes in my solution (Step i is necessary, otherwise it would be by gradient messages.). </p>\n<ul>\n<li>Public: 0.7525.      </li>\n<li>Private: 0.8340</li></ul></li>\n<li><p>Ensemble 2 seed models without quantization. (Surprised by this decline, I am looking into the reasons for it.)</p>\n<ul>\n<li>Ensemble Public: 0.7443    <ul>\n<li>Model0 Public: 0.73</li>\n<li>Model1 Public: 0.7314</li></ul></li>\n<li>Ensemble Private: 0.8276<ul>\n<li>Model0 Private: 0.814</li>\n<li>Model1 Private: 0.816</li></ul></li>\n<li><strong>I have found the reason. I change <code>lr</code> to <code>0.4*lr</code>, cause <code>BATCH_ALL_SIGNS_N</code> reduces from 4 to 1. This action was wrong, because I use rotation augmentation before training. In other words, all changing variables can affect the choice of <code>lr</code>. If those variables stay same but <code>batch_size</code>, We need to make the change in <code>batch_size</code> proportional to the change in <code>lr</code>.</strong> (Answer from this <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248253\" target=\"_blank\">post</a>)<br>\nAfter cancelling the change of <code>lr</code>. The result is shown below</li>\n<li>Ensemble Public: 0.7683   </li>\n<li>Ensemble Private: 0.8445</li></ul></li></ul></li>\n<li><p>To be resolved</p>\n<ul>\n<li><p>What is Time Scale augmentation? and  How to remove random frame masking?</p></li>\n<li><p>Why the learning rate is proportional to the batch size?</p>\n<ul>\n<li>Answer: This is correct with the other model parameters held constant. However, this is wrong if other variables (epoch, dropout rate, etc.) change.</li></ul></li>\n<li><p>Where this from \"Accurate local validation scheme which estimates LB score\" ?</p></li></ul></li>\n</ul>\n<h1>Rule of thumbs</h1>\n<h2>How to choose learning rate</h2>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook\" target=\"_blank\">https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook</a></p></li>\n<li><p>From this <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248264\" target=\"_blank\">post</a></p></li>\n</ul>\n<blockquote>\n  <p>If training requires only a few epochs (i.e. 1, 2, 3, 4, or 5 epochs) like NLP models, I usually just go with constant LR. Then i look at validation metric and if it increases in the last epoch or two, perhaps I will use <code>LR</code> for the first epochs and <code>LR * 0.1</code> (or 0.3) for the last epochs.</p>\n  <p>When a model requires many epochs like 20, 40, 80, 160. I usually go with cosine schedule (like this comp). These two techniques work well in most cases.</p>\n</blockquote>",
      "rawMarkdown": "Thanks to the organisers of the PopSign Games and all of the Kaggle community for running such a great competition.\n\nMore importantly, my heartfelt thanks go to both @markwijkhuizen  and @hengck23. MARK contributed a very strong transformer framework, which is the machine learning pipeline for me in this competition. And I will find daily inspiration in @hengck23's discussion post. \n\nAs a learner, this competition has taken my machine learning skills a step further. Therefore, I would like to **record the results of my experiment here**.\n\n\n# Model Architecture\n\nShallow networks are the best choice.\n\n- Single shallow transformer base (Refer to [version 5](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training))\n  - Added `label smoothing`.\n  - Used warm up to `lr`.\n  - Remove all LayerNorm.\n\nThe main changes I have made that have led to improvements:\n\n- Increase `Embedding` to 600.\n- Add a `Gaussian noise (0.01)` to the FC layer of the Embedding layer.\n- Increase ` epoch` to 230\n- Deepens the depth of the last FC layer in the transformer layer.\n- Late Dropout. 0---0.3---0.1\n\n\n# Feature Engineering\n\nFeature engineering was the most difficult part of my entire experiment. I came up with a lot of ideas, but most of them ended up being submission errors.\n\nHere is my summary of options that can bring about an uplift and are feasible.\n\n- Use of `lips`, `eyes` (regardless of left or right eye), `main hand` (21 landmarks), and `main pose` (6 landmarks)\n  - Main hand means normalize the right hand to the left hand. The formula is `x = 1 - x`.\n  - The formula of main pose is `x = 1.2 -x`. \n\n- Set the input frame count `threshold` to **48**.\n  - Less than that is filled with 0.\n  - More than that is compressed using the `Bilinear Interpolation` method. (I think it only work well to data where the frame rate is a multiple of 48)\n  - Removal of data with only `one frame`.\n- Add `motion feature` for coordinate X.\n- Add `Hand distance feature`. (Get inspired by @josephzahar's [post](https://www.kaggle.com/code/josephzahar/distance-angle-based-features-using-keras))\n  - Calculate the Euclidean distance between the discontinuous joints of the hand.\n\nHence, the dimensionality of each data in my pipeline is 3 dimensions `(x, y, motion)`. For the Hand, it is 213` (x, y, motion, distance)`. \n\n# Augmentation\nThere is no need for mirror(flip) augmentation in this pipeline, as the hand is already normalised in the pre-processing.\n\nIn augmentation, I only found rotation to be little useful for me. Specifically, I rotated the data with a small sample size (`data[-42:]`) by 15 degrees up and down. But I think my approach is very poor.\n\n# Future Work\n\n- How to distinguish similar words, like awake and wake?\n- Solve the problem of `keras_model_converter.target_spec.supported_types = [tf.float16]` not working for me.\n- How to do rotation correctly?\n\n\n\n\n\n# Summary of Best Solutions (In processing)\n\n###  \"How To Improve Best Public transformer basded Notebook\" from @cdeotte \n\n- Original Post\n\n  - https://www.kaggle.com/competitions/asl-signs/discussion/406302\n\n- Corresponding Notebook\n\n  - https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\n\n- Keypoints: \n\n  - Reduce batch size and increse transformer block.\n  - Ensemble with different model seeds.\n\n- All changes\n\n  1. Train 1 model => Train 4 models\n\n  2. Add Time Scale augmentation\n\n  3. Ensemble and apply TFLite FP16 quantization\n\n  Change the following parameters:\n\n  4. INPUT_SIZE, 64 => 12\n\n  5. BATCH_ALL_SIGNS_N, 4 => 1\n\n  6. N_EPOCHS, 250 => 120\n\n  7. LANDMARK_UNITS, 384 => 224\n\n  8. UNITS, 512 => 376\n\n  9. NUM_BLOCKS, 2 => 3\n\n  10. MLP_RATIO, 4 => 3\n\n  11. MLP_DROPOUT_RATIO, 0.40 => 0.30\n\n  12. Remove random frame masking\n\n- The results of the experiments combined with those tricks\n\n  - Just use a-i parameters changes in my solution (Step i is necessary, otherwise it would be by gradient messages.). \n      - Public: 0.7525.      \n      - Private: 0.8340\n\n  - Ensemble 2 seed models without quantization. (Surprised by this decline, I am looking into the reasons for it.)\n      - Ensemble Public: 0.7443    \n         - Model0 Public: 0.73\n         - Model1 Public: 0.7314\n      - Ensemble Private: 0.8276\n         - Model0 Private: 0.814\n         - Model1 Private: 0.816\n      - **I have found the reason. I change `lr` to `0.4*lr`, cause `BATCH_ALL_SIGNS_N` reduces from 4 to 1. This action was wrong, because I use rotation augmentation before training. In other words, all changing variables can affect the choice of `lr`. If those variables stay same but `batch_size`, We need to make the change in `batch_size` proportional to the change in `lr`.** (Answer from this [post](https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248253))\nAfter cancelling the change of `lr`. The result is shown below\n      - Ensemble Public: 0.7683   \n      - Ensemble Private: 0.8445\n\n- To be resolved\n\n  - What is Time Scale augmentation? and  How to remove random frame masking?\n\n   - Why the learning rate is proportional to the batch size?\n      - Answer: This is correct with the other model parameters held constant. However, this is wrong if other variables (epoch, dropout rate, etc.) change.\n\n   - Where this from \"Accurate local validation scheme which estimates LB score\" ?\n\n\n\n# Rule of thumbs\n\n## How to choose learning rate\n\n- https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook\n\n- From this [post](https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248264)\n> If training requires only a few epochs (i.e. 1, 2, 3, 4, or 5 epochs) like NLP models, I usually just go with constant LR. Then i look at validation metric and if it increases in the last epoch or two, perhaps I will use `LR` for the first epochs and `LR * 0.1` (or 0.3) for the last epochs.\n> \n> When a model requires many epochs like 20, 40, 80, 160. I usually go with cosine schedule (like this comp). These two techniques work well in most cases.",
      "votes": null
    },
    {
      "id": "2242344",
      "postDate": "05/02/2023 07:47:17",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": null
    },
    {
      "id": "2244888",
      "postDate": "05/04/2023 01:19:44",
      "content": "<p>Thank you. Your thoughts bring me a lot.</p>",
      "rawMarkdown": "Thank you. Your thoughts bring me a lot.",
      "votes": null
    },
    {
      "id": "2246982",
      "postDate": "05/05/2023 15:38:52",
      "content": "<p>Great work!<br>\nWhy shallow networks is a best choice?</p>",
      "rawMarkdown": "Great work!\nWhy shallow networks is a best choice?",
      "votes": null
    },
    {
      "id": "2247400",
      "postDate": "05/06/2023 01:05:44",
      "content": "<p>Firstly, there are time constraints and size constraints, and you need to make a lightweight model.<br>\nSecondly, there is the large scope for manipulation that a shallow network brings. You can do operations such as ensemble.<br>\nHowever, that's only relative to the context of this competition.</p>",
      "rawMarkdown": "Firstly, there are time constraints and size constraints, and you need to make a lightweight model.\nSecondly, there is the large scope for manipulation that a shallow network brings. You can do operations such as ensemble.\nHowever, that's only relative to the context of this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242344,
      "author_name": "josephzahar",
      "author_url": "",
      "post_date": "05/02/2023 07:47:17",
      "content": "<p>Great Work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2244888,
          "author_name": "jacksonyou",
          "author_url": "",
          "post_date": "05/04/2023 01:19:44",
          "content": "<p>Thank you. Your thoughts bring me a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2246982,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/05/2023 15:38:52",
      "content": "<p>Great work!<br>\nWhy shallow networks is a best choice?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2247400,
          "author_name": "jacksonyou",
          "author_url": "",
          "post_date": "05/06/2023 01:05:44",
          "content": "<p>Firstly, there are time constraints and size constraints, and you need to make a lightweight model.<br>\nSecondly, there is the large scope for manipulation that a shallow network brings. You can do operations such as ensemble.<br>\nHowever, that's only relative to the context of this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2242097": "Thanks to the organisers of the PopSign Games and all of the Kaggle community for running such a great competition.\n\nMore importantly, my heartfelt thanks go to both @markwijkhuizen  and @hengck23. MARK contributed a very strong transformer framework, which is the machine learning pipeline for me in this competition. And I will find daily inspiration in @hengck23's discussion post. \n\nAs a learner, this competition has taken my machine learning skills a step further. Therefore, I would like to **record the results of my experiment here**.\n\n\n# Model Architecture\n\nShallow networks are the best choice.\n\n- Single shallow transformer base (Refer to [version 5](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training))\n  - Added `label smoothing`.\n  - Used warm up to `lr`.\n  - Remove all LayerNorm.\n\nThe main changes I have made that have led to improvements:\n\n- Increase `Embedding` to 600.\n- Add a `Gaussian noise (0.01)` to the FC layer of the Embedding layer.\n- Increase ` epoch` to 230\n- Deepens the depth of the last FC layer in the transformer layer.\n- Late Dropout. 0---0.3---0.1\n\n\n# Feature Engineering\n\nFeature engineering was the most difficult part of my entire experiment. I came up with a lot of ideas, but most of them ended up being submission errors.\n\nHere is my summary of options that can bring about an uplift and are feasible.\n\n- Use of `lips`, `eyes` (regardless of left or right eye), `main hand` (21 landmarks), and `main pose` (6 landmarks)\n  - Main hand means normalize the right hand to the left hand. The formula is `x = 1 - x`.\n  - The formula of main pose is `x = 1.2 -x`. \n\n- Set the input frame count `threshold` to **48**.\n  - Less than that is filled with 0.\n  - More than that is compressed using the `Bilinear Interpolation` method. (I think it only work well to data where the frame rate is a multiple of 48)\n  - Removal of data with only `one frame`.\n- Add `motion feature` for coordinate X.\n- Add `Hand distance feature`. (Get inspired by @josephzahar's [post](https://www.kaggle.com/code/josephzahar/distance-angle-based-features-using-keras))\n  - Calculate the Euclidean distance between the discontinuous joints of the hand.\n\nHence, the dimensionality of each data in my pipeline is 3 dimensions `(x, y, motion)`. For the Hand, it is 213` (x, y, motion, distance)`. \n\n# Augmentation\nThere is no need for mirror(flip) augmentation in this pipeline, as the hand is already normalised in the pre-processing.\n\nIn augmentation, I only found rotation to be little useful for me. Specifically, I rotated the data with a small sample size (`data[-42:]`) by 15 degrees up and down. But I think my approach is very poor.\n\n# Future Work\n\n- How to distinguish similar words, like awake and wake?\n- Solve the problem of `keras_model_converter.target_spec.supported_types = [tf.float16]` not working for me.\n- How to do rotation correctly?\n\n\n\n\n\n# Summary of Best Solutions (In processing)\n\n###  \"How To Improve Best Public transformer basded Notebook\" from @cdeotte \n\n- Original Post\n\n  - https://www.kaggle.com/competitions/asl-signs/discussion/406302\n\n- Corresponding Notebook\n\n  - https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\n\n- Keypoints: \n\n  - Reduce batch size and increse transformer block.\n  - Ensemble with different model seeds.\n\n- All changes\n\n  1. Train 1 model => Train 4 models\n\n  2. Add Time Scale augmentation\n\n  3. Ensemble and apply TFLite FP16 quantization\n\n  Change the following parameters:\n\n  4. INPUT_SIZE, 64 => 12\n\n  5. BATCH_ALL_SIGNS_N, 4 => 1\n\n  6. N_EPOCHS, 250 => 120\n\n  7. LANDMARK_UNITS, 384 => 224\n\n  8. UNITS, 512 => 376\n\n  9. NUM_BLOCKS, 2 => 3\n\n  10. MLP_RATIO, 4 => 3\n\n  11. MLP_DROPOUT_RATIO, 0.40 => 0.30\n\n  12. Remove random frame masking\n\n- The results of the experiments combined with those tricks\n\n  - Just use a-i parameters changes in my solution (Step i is necessary, otherwise it would be by gradient messages.). \n      - Public: 0.7525.      \n      - Private: 0.8340\n\n  - Ensemble 2 seed models without quantization. (Surprised by this decline, I am looking into the reasons for it.)\n      - Ensemble Public: 0.7443    \n         - Model0 Public: 0.73\n         - Model1 Public: 0.7314\n      - Ensemble Private: 0.8276\n         - Model0 Private: 0.814\n         - Model1 Private: 0.816\n      - **I have found the reason. I change `lr` to `0.4*lr`, cause `BATCH_ALL_SIGNS_N` reduces from 4 to 1. This action was wrong, because I use rotation augmentation before training. In other words, all changing variables can affect the choice of `lr`. If those variables stay same but `batch_size`, We need to make the change in `batch_size` proportional to the change in `lr`.** (Answer from this [post](https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248253))\nAfter cancelling the change of `lr`. The result is shown below\n      - Ensemble Public: 0.7683   \n      - Ensemble Private: 0.8445\n\n- To be resolved\n\n  - What is Time Scale augmentation? and  How to remove random frame masking?\n\n   - Why the learning rate is proportional to the batch size?\n      - Answer: This is correct with the other model parameters held constant. However, this is wrong if other variables (epoch, dropout rate, etc.) change.\n\n   - Where this from \"Accurate local validation scheme which estimates LB score\" ?\n\n\n\n# Rule of thumbs\n\n## How to choose learning rate\n\n- https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook\n\n- From this [post](https://www.kaggle.com/competitions/asl-signs/discussion/406302#2248264)\n> If training requires only a few epochs (i.e. 1, 2, 3, 4, or 5 epochs) like NLP models, I usually just go with constant LR. Then i look at validation metric and if it increases in the last epoch or two, perhaps I will use `LR` for the first epochs and `LR * 0.1` (or 0.3) for the last epochs.\n> \n> When a model requires many epochs like 20, 40, 80, 160. I usually go with cosine schedule (like this comp). These two techniques work well in most cases.",
    "2242344": "Great Work!",
    "2244888": "Thank you. Your thoughts bring me a lot.",
    "2246982": "Great work!\nWhy shallow networks is a best choice?",
    "2247400": "Firstly, there are time constraints and size constraints, and you need to make a lightweight model.\nSecondly, there is the large scope for manipulation that a shallow network brings. You can do operations such as ensemble.\nHowever, that's only relative to the context of this competition."
  },
  "source": "meta"
}