{
  "id": 406302,
  "title": "44th Place Silver - How To Improve Best Public Notebook",
  "url": "/competitions/asl-signs/discussion/406302",
  "author_name": "Chris Deotte",
  "post_date": "2023-05-02T00:08:37.798000",
  "votes": 73,
  "comment_count": 41,
  "views": 0,
  "content": "<p>Thank you Kaggle, Kagglers, PopSign, and Partners for a fun competition! Public notebooks in this competition are very strong! The host said that they need <code>LB &gt;= 0.6</code> to be helpful and the best public notebook scores <code>LB = 0.73</code>! Great job Kagglers!</p>\n<p>I joined this competition one week ago, so I didn't have much time to build my own model. Instead, i read the discussions and public notebooks, and attempted to improve the best public notebook. I boosted the best public notebook from public <code>LB = 0.73</code> to <code>LB = 0.77</code> and achieved 44th place Silver.</p>\n<h1>Praise for Best Public Notebooks</h1>\n<p>I am very impressed with the shared public notebooks</p>\n<ul>\n<li>Great Transformer model architecture</li>\n<li>Selected which 66 Landmarks out of 543 Landmarks are important (hands, lips, pose)</li>\n<li>Great preprocessing to convert videos of variable length to small fixed length</li>\n<li>Accurate local validation scheme which estimates LB score</li>\n</ul>\n<h1>Improvements to Best Public Notebook</h1>\n<p>After reading the best public notebook, I made the following 11 changes to boost CV and LB <code>+0.03</code> from <code>LB = 0.73</code> to <code>LB = 0.76</code>:</p>\n<ul>\n<li>Train 1 model =&gt; Train 4 models</li>\n<li>Add Time Scale augmentation</li>\n<li>Ensemble and apply TFLite FP16 quantization</li>\n<li>Change the following parameters:</li>\n<li>INPUT_SIZE, 64 =&gt; 12</li>\n<li>BATCH_ALL_SIGNS_N, 4 =&gt; 1</li>\n<li>N_EPOCHS, 250 =&gt; 120</li>\n<li>LANDMARK_UNITS, 384 =&gt; 224</li>\n<li>UNITS, 512 =&gt; 376</li>\n<li>NUM_BLOCKS, 2 =&gt; 3</li>\n<li>MLP_RATIO, 4 =&gt; 3</li>\n<li>MLP_DROPOUT_RATIO, 0.40 =&gt; 0.30</li>\n<li>remove random frame masking</li>\n</ul>\n<h1>Published Code</h1>\n<p>I published my code <a href=\"https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\" target=\"_blank\">here</a> for those curious to learn \"How To Improve Best Public Notebook\" from <code>LB = 0.73</code> to <code>LB = 0.76</code></p>\n<h1>My Thought Process</h1>\n<p>Most Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.</p>\n<p>The first thing i did was reduce <code>LANDMARK_UNITS</code>, <code>UNITS</code>, and <code>MLP_RATIO</code> to reduce the model size. I slowly reduced these parameters to see how small I could go without reducing CV score. It is important to reduce the model to less than 5M (i.e. 5 million) parameters because then we can infer 4 copies of the model within the 40MB size restriction when using FP16. With 40MB size limit, the maximum number of parameters during inference with FP16 is 20M. I also tried less than 4M parameters and 3.3M parameters which allows 5 copies and 6 copies respectively (when using FP16 quantization) </p>\n<p>Second, I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about <code>+0.01</code> or <code>'0.02</code>. Reducing the <code>MLP_RATIO</code> from 4 to 3 significantly reduced the number of model parameters without affecting CV LB so this gave me extra parameters to allow increasing the transformer blocks.</p>\n<p>Next, I reduced <code>INPUT_SIZE</code> to speed up the model to make it infer (and train) faster. The parameter <code>INPUT_SIZE</code> is the transformer sequence length. This competition has a inference time constraint of 1 hour (in addition to a 40MB size constraint). Reducing <code>INPUT_SIZE</code> will decrease inference time (and does not affect model size). With sequence length 12, we can easily infer 5+ models under 1 hour.</p>\n<p>At this point, i got lucky. Reducing <code>INPUT_SIZE</code> actually boosted the CV and LB by <code>+0.01</code> or <code>+0.02</code> or so. After this I tuned the learning rate, batch size, and learning schedule and found <code>N_EPOCHS = 120</code> and <code>BATCH_ALL_SIGNS_N = 1</code> to be best. Many Kagglers overlook the fact that changing batch size can make a big difference for models. We should always try 0.25x, 0.5x, 2x, 4x batch size and change the learning rate for those experiments to be 0.25x, 0.5x, 2x, 4x respectively. Experiments demonstrate that <code>Batch Size = 256</code> is better than the original <code>Batch Size = 1024</code>. This helped <code>+0.0005</code> or so.</p>\n<p>Lastly NN always benefit from more data and data augmentation. I tried external data and many different data augmentation. Using external data did not help me. Regarding augmentation i only got benefit from <code>frame drop out</code> and <code>time scale augmentation</code>. This helped about <code>+0.0005</code> or so. I also tried rotation, scale, shift but this did not help. Flip was not needed since preprocessing normalized all videos to be left handed.</p>",
  "messages": [
    {
      "id": 2241890,
      "postDate": "2023-05-02T00:08:37.797Z",
      "content": "<p>Thank you Kaggle, Kagglers, PopSign, and Partners for a fun competition! Public notebooks in this competition are very strong! The host said that they need <code>LB &gt;= 0.6</code> to be helpful and the best public notebook scores <code>LB = 0.73</code>! Great job Kagglers!</p>\n<p>I joined this competition one week ago, so I didn't have much time to build my own model. Instead, i read the discussions and public notebooks, and attempted to improve the best public notebook. I boosted the best public notebook from public <code>LB = 0.73</code> to <code>LB = 0.77</code> and achieved 44th place Silver.</p>\n<h1>Praise for Best Public Notebooks</h1>\n<p>I am very impressed with the shared public notebooks</p>\n<ul>\n<li>Great Transformer model architecture</li>\n<li>Selected which 66 Landmarks out of 543 Landmarks are important (hands, lips, pose)</li>\n<li>Great preprocessing to convert videos of variable length to small fixed length</li>\n<li>Accurate local validation scheme which estimates LB score</li>\n</ul>\n<h1>Improvements to Best Public Notebook</h1>\n<p>After reading the best public notebook, I made the following 11 changes to boost CV and LB <code>+0.03</code> from <code>LB = 0.73</code> to <code>LB = 0.76</code>:</p>\n<ul>\n<li>Train 1 model =&gt; Train 4 models</li>\n<li>Add Time Scale augmentation</li>\n<li>Ensemble and apply TFLite FP16 quantization</li>\n<li>Change the following parameters:</li>\n<li>INPUT_SIZE, 64 =&gt; 12</li>\n<li>BATCH_ALL_SIGNS_N, 4 =&gt; 1</li>\n<li>N_EPOCHS, 250 =&gt; 120</li>\n<li>LANDMARK_UNITS, 384 =&gt; 224</li>\n<li>UNITS, 512 =&gt; 376</li>\n<li>NUM_BLOCKS, 2 =&gt; 3</li>\n<li>MLP_RATIO, 4 =&gt; 3</li>\n<li>MLP_DROPOUT_RATIO, 0.40 =&gt; 0.30</li>\n<li>remove random frame masking</li>\n</ul>\n<h1>Published Code</h1>\n<p>I published my code <a href=\"https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\" target=\"_blank\">here</a> for those curious to learn \"How To Improve Best Public Notebook\" from <code>LB = 0.73</code> to <code>LB = 0.76</code></p>\n<h1>My Thought Process</h1>\n<p>Most Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.</p>\n<p>The first thing i did was reduce <code>LANDMARK_UNITS</code>, <code>UNITS</code>, and <code>MLP_RATIO</code> to reduce the model size. I slowly reduced these parameters to see how small I could go without reducing CV score. It is important to reduce the model to less than 5M (i.e. 5 million) parameters because then we can infer 4 copies of the model within the 40MB size restriction when using FP16. With 40MB size limit, the maximum number of parameters during inference with FP16 is 20M. I also tried less than 4M parameters and 3.3M parameters which allows 5 copies and 6 copies respectively (when using FP16 quantization) </p>\n<p>Second, I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about <code>+0.01</code> or <code>'0.02</code>. Reducing the <code>MLP_RATIO</code> from 4 to 3 significantly reduced the number of model parameters without affecting CV LB so this gave me extra parameters to allow increasing the transformer blocks.</p>\n<p>Next, I reduced <code>INPUT_SIZE</code> to speed up the model to make it infer (and train) faster. The parameter <code>INPUT_SIZE</code> is the transformer sequence length. This competition has a inference time constraint of 1 hour (in addition to a 40MB size constraint). Reducing <code>INPUT_SIZE</code> will decrease inference time (and does not affect model size). With sequence length 12, we can easily infer 5+ models under 1 hour.</p>\n<p>At this point, i got lucky. Reducing <code>INPUT_SIZE</code> actually boosted the CV and LB by <code>+0.01</code> or <code>+0.02</code> or so. After this I tuned the learning rate, batch size, and learning schedule and found <code>N_EPOCHS = 120</code> and <code>BATCH_ALL_SIGNS_N = 1</code> to be best. Many Kagglers overlook the fact that changing batch size can make a big difference for models. We should always try 0.25x, 0.5x, 2x, 4x batch size and change the learning rate for those experiments to be 0.25x, 0.5x, 2x, 4x respectively. Experiments demonstrate that <code>Batch Size = 256</code> is better than the original <code>Batch Size = 1024</code>. This helped <code>+0.0005</code> or so.</p>\n<p>Lastly NN always benefit from more data and data augmentation. I tried external data and many different data augmentation. Using external data did not help me. Regarding augmentation i only got benefit from <code>frame drop out</code> and <code>time scale augmentation</code>. This helped about <code>+0.0005</code> or so. I also tried rotation, scale, shift but this did not help. Flip was not needed since preprocessing normalized all videos to be left handed.</p>",
      "rawMarkdown": "Thank you Kaggle, Kagglers, PopSign, and Partners for a fun competition! Public notebooks in this competition are very strong! The host said that they need `LB >= 0.6` to be helpful and the best public notebook scores `LB = 0.73`! Great job Kagglers!\n\nI joined this competition one week ago, so I didn't have much time to build my own model. Instead, i read the discussions and public notebooks, and attempted to improve the best public notebook. I boosted the best public notebook from public `LB = 0.73` to `LB = 0.77` and achieved 44th place Silver.\n\n# Praise for Best Public Notebooks\nI am very impressed with the shared public notebooks\n* Great Transformer model architecture\n* Selected which 66 Landmarks out of 543 Landmarks are important (hands, lips, pose)\n* Great preprocessing to convert videos of variable length to small fixed length\n* Accurate local validation scheme which estimates LB score\n\n# Improvements to Best Public Notebook\nAfter reading the best public notebook, I made the following 11 changes to boost CV and LB `+0.03` from `LB = 0.73` to `LB = 0.76`:\n\n* Train 1 model => Train 4 models\n* Add Time Scale augmentation\n* Ensemble and apply TFLite FP16 quantization\n* Change the following parameters:\n* INPUT_SIZE, 64 => 12\n* BATCH_ALL_SIGNS_N, 4 => 1\n* N_EPOCHS, 250 => 120\n* LANDMARK_UNITS, 384 => 224\n* UNITS, 512 => 376\n* NUM_BLOCKS, 2 => 3\n* MLP_RATIO, 4 => 3\n* MLP_DROPOUT_RATIO, 0.40 => 0.30\n* remove random frame masking\n\n# Published Code\nI published my code [here][1] for those curious to learn \"How To Improve Best Public Notebook\" from `LB = 0.73` to `LB = 0.76`\n\n# My Thought Process\nMost Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.\n\nThe first thing i did was reduce `LANDMARK_UNITS`, `UNITS`, and `MLP_RATIO` to reduce the model size. I slowly reduced these parameters to see how small I could go without reducing CV score. It is important to reduce the model to less than 5M (i.e. 5 million) parameters because then we can infer 4 copies of the model within the 40MB size restriction when using FP16. With 40MB size limit, the maximum number of parameters during inference with FP16 is 20M. I also tried less than 4M parameters and 3.3M parameters which allows 5 copies and 6 copies respectively (when using FP16 quantization) \n\nSecond, I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about `+0.01` or `'0.02`. Reducing the `MLP_RATIO` from 4 to 3 significantly reduced the number of model parameters without affecting CV LB so this gave me extra parameters to allow increasing the transformer blocks.\n\nNext, I reduced `INPUT_SIZE` to speed up the model to make it infer (and train) faster. The parameter `INPUT_SIZE` is the transformer sequence length. This competition has a inference time constraint of 1 hour (in addition to a 40MB size constraint). Reducing `INPUT_SIZE` will decrease inference time (and does not affect model size). With sequence length 12, we can easily infer 5+ models under 1 hour.\n\nAt this point, i got lucky. Reducing `INPUT_SIZE` actually boosted the CV and LB by `+0.01` or `+0.02` or so. After this I tuned the learning rate, batch size, and learning schedule and found `N_EPOCHS = 120` and `BATCH_ALL_SIGNS_N = 1` to be best. Many Kagglers overlook the fact that changing batch size can make a big difference for models. We should always try 0.25x, 0.5x, 2x, 4x batch size and change the learning rate for those experiments to be 0.25x, 0.5x, 2x, 4x respectively. Experiments demonstrate that `Batch Size = 256` is better than the original `Batch Size = 1024`. This helped `+0.0005` or so.\n\nLastly NN always benefit from more data and data augmentation. I tried external data and many different data augmentation. Using external data did not help me. Regarding augmentation i only got benefit from `frame drop out` and `time scale augmentation`. This helped about `+0.0005` or so. I also tried rotation, scale, shift but this did not help. Flip was not needed since preprocessing normalized all videos to be left handed.\n\n[1]: https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76",
      "votes": 72
    },
    {
      "id": 2244633,
      "postDate": "2023-05-03T19:09:54.453Z",
      "content": "<p>I still remember how last year, during the ventilator competition, you casually mentioned at the end, \"If you just take the top public notebook and reduce the batch size, it gets to the silver zone.\"</p>\n<p><strong>Reducing the batch size never gets old!</strong></p>\n<p>Amazing work, as always!</p>",
      "rawMarkdown": "I still remember how last year, during the ventilator competition, you casually mentioned at the end, \"If you just take the top public notebook and reduce the batch size, it gets to the silver zone.\"\n\n**Reducing the batch size never gets old!**\n\nAmazing work, as always!",
      "votes": 5,
      "replies": [
        {
          "id": 2244872,
          "postDate": "2023-05-04T00:36:50.063Z",
          "content": "<p>Thanks Devastator!</p>",
          "rawMarkdown": "Thanks Devastator!"
        }
      ]
    },
    {
      "id": 2242849,
      "postDate": "2023-05-02T14:38:33.180Z",
      "content": "<p><strong>UPDATE:</strong> I shared my code in notebook <a href=\"https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "**UPDATE:** I shared my code in notebook [here][1]\n\n[1]: https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76",
      "votes": 3
    },
    {
      "id": 2400466,
      "postDate": "2023-08-21T04:54:50.517Z",
      "content": "<blockquote>\n  <p>Most Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.</p>\n</blockquote>\n<p>Great Insight. Thank you.</p>",
      "rawMarkdown": ">Most Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.\n\nGreat Insight. Thank you.",
      "votes": 1
    },
    {
      "id": 2247677,
      "postDate": "2023-05-06T07:57:36.553Z",
      "content": "<p>I was not active in this competition but reading solution description to learn as much as possible. Your approach for improving public notebook is brilliant. We have many things to touch … many to introduce but great people manage to find appropriate things which lead to great results. Thank you for sharing your workflow. </p>",
      "rawMarkdown": "I was not active in this competition but reading solution description to learn as much as possible. Your approach for improving public notebook is brilliant. We have many things to touch … many to introduce but great people manage to find appropriate things which lead to great results. Thank you for sharing your workflow. ",
      "votes": 1
    },
    {
      "id": 2247434,
      "postDate": "2023-05-06T02:11:05.763Z",
      "content": "<p>Thanks for sharing great insight of the competition. </p>\n<p>Would I ask one question?</p>\n<blockquote>\n  <p>I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about +0.01 or '0.02. Reducing the MLP_RATIO from 4 to 3 significantly reduced the number of model parameters without affecting CV LB.</p>\n</blockquote>\n<p>I don't understand why increasing the number of blocks and decreasing MLP_RATIO leads to boosting the CV and LB. What made you think of increasing block size while decreasing MLP_RATIO?</p>",
      "rawMarkdown": "Thanks for sharing great insight of the competition. \n\nWould I ask one question?\n\n> I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about +0.01 or '0.02. Reducing the MLP_RATIO from 4 to 3 significantly reduced the number of model parameters without affecting CV LB.\n\nI don't understand why increasing the number of blocks and decreasing MLP_RATIO leads to boosting the CV and LB. What made you think of increasing block size while decreasing MLP_RATIO?",
      "votes": 1,
      "replies": [
        {
          "id": 2248251,
          "postDate": "2023-05-06T15:56:16.093Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ju7on9\" target=\"_blank\">@ju7on9</a> . The original network architecture is defined by</p>\n<pre><code># Dense layer units for landmarks\nLIPS_UNITS = 224 # WAS 384\nHANDS_UNITS = 224 # WAS 384\nPOSE_UNITS = 224 # WAS 384\n# final embedding and transformer embedding size\nUNITS = 376 # WAS 512\n\n# Transformer\nNUM_BLOCKS = 3 # WAS 2\nMLP_RATIO = 3 # WAS 4\n</code></pre>\n<p>I changed each for a different reason. I lowered <code>MLP_RATIO</code> to shrink the model and I raised <code>NUM_BLOCKS</code> to improve CV and LB. In my discussion, i should say reducing <code>LANDMARK_UNITS, UNITS, and MLP_RATIO</code> gave me extra room to increase <code>NUM_BLOCKS</code>. The point is that increasing <code>NUM_BLOCKS</code> will add 50% more parameters! So we must find a way to significantly reduce the size of the model before we can do that. </p>\n<p>The original network had 7.4M parameters and our goal is to get one NN under 5M. We can reduce <code>MLP_RATIO</code> from <code>4 =&gt; 3</code> without degrading CV nor LB. This reduces the network to 6.3M parameters. Next decreasing <code>LANDMARK_UNITS</code> from <code>384 =&gt; 224</code> reduces to 5.9M parameters. And decreasing <code>UNITS</code> from <code>512 =&gt; 384</code> reduces to 3.4M parameters. And neither of these reductions hurt CV nor LB either.</p>\n<p>We now have extra room to increase <code>NUM_BLOCKS</code>. Increasing <code>NUM_BLOCKS</code> from <code>2 =&gt; 3</code> increases parameters to 4.9M parameters. And boosts CV and LB. We are now under 5M parameters and can safely ensemble 4xNN at <code>FP16</code> within the restriction of 40MB size.</p>",
          "rawMarkdown": "Hi @ju7on9 . The original network architecture is defined by\n\n    # Dense layer units for landmarks\n    LIPS_UNITS = 224 # WAS 384\n    HANDS_UNITS = 224 # WAS 384\n    POSE_UNITS = 224 # WAS 384\n    # final embedding and transformer embedding size\n    UNITS = 376 # WAS 512\n\n    # Transformer\n    NUM_BLOCKS = 3 # WAS 2\n    MLP_RATIO = 3 # WAS 4\n\nI changed each for a different reason. I lowered `MLP_RATIO` to shrink the model and I raised `NUM_BLOCKS` to improve CV and LB. In my discussion, i should say reducing `LANDMARK_UNITS, UNITS, and MLP_RATIO` gave me extra room to increase `NUM_BLOCKS`. The point is that increasing `NUM_BLOCKS` will add 50% more parameters! So we must find a way to significantly reduce the size of the model before we can do that. \n\nThe original network had 7.4M parameters and our goal is to get one NN under 5M. We can reduce `MLP_RATIO` from `4 => 3` without degrading CV nor LB. This reduces the network to 6.3M parameters. Next decreasing `LANDMARK_UNITS` from `384 => 224` reduces to 5.9M parameters. And decreasing `UNITS` from `512 => 384` reduces to 3.4M parameters. And neither of these reductions hurt CV nor LB either.\n\nWe now have extra room to increase `NUM_BLOCKS`. Increasing `NUM_BLOCKS` from `2 => 3` increases parameters to 4.9M parameters. And boosts CV and LB. We are now under 5M parameters and can safely ensemble 4xNN at `FP16` within the restriction of 40MB size.",
          "votes": 2,
          "replies": [
            {
              "id": 2249000,
              "postDate": "2023-05-07T12:02:56.277Z",
              "content": "<p>Thank you for the detailed answer! It really helped me a lot!</p>",
              "rawMarkdown": "Thank you for the detailed answer! It really helped me a lot!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2245288,
      "postDate": "2023-05-04T09:13:42.523Z",
      "content": "<p>Congrats and thanks for sharing your experience. I learn a lot from your experience.</p>\n<p>I'm a bit confused by the phrase \"trained with different seeds\". Do you mean the SEED for each batch of data read or the SEED for the initialization of parameters inside the model? In other words, is this SEED that we don't have to set manually?</p>",
      "rawMarkdown": "Congrats and thanks for sharing your experience. I learn a lot from your experience.\n\nI'm a bit confused by the phrase \"trained with different seeds\". Do you mean the SEED for each batch of data read or the SEED for the initialization of parameters inside the model? In other words, is this SEED that we don't have to set manually?",
      "votes": 1,
      "replies": [
        {
          "id": 2245578,
          "postDate": "2023-05-04T13:00:03.393Z",
          "content": "<blockquote>\n  <p>is this SEED that we don't have to set manually?</p>\n</blockquote>\n<p>Yes, there is nothing for us to change. The \"SEED\" is for the initialization of NN layers. And the shuffle of batches during training. And the randomness of augmentation etc etc. It is all the random elements during training. So if this \"SEED\" changes then each NN will be trained differently.</p>",
          "rawMarkdown": ">is this SEED that we don't have to set manually?\n\nYes, there is nothing for us to change. The \"SEED\" is for the initialization of NN layers. And the shuffle of batches during training. And the randomness of augmentation etc etc. It is all the random elements during training. So if this \"SEED\" changes then each NN will be trained differently.",
          "votes": 1,
          "replies": [
            {
              "id": 2247793,
              "postDate": "2023-05-06T09:28:24.313Z",
              "content": "<blockquote>\n  <p>We should always try 0.25x, 0.5x, 2x, 4x batch size and change the learning rate for those experiments to be 0.25x, 0.5x, 2x, 4x respectively.</p>\n</blockquote>\n<p>I found that this was not followed in your code. In my subsequent experiments, following that also had a negative impact. I wonder if this applies to large models rather than small nerwork?</p>",
              "rawMarkdown": ">We should always try 0.25x, 0.5x, 2x, 4x batch size and change the learning rate for those experiments to be 0.25x, 0.5x, 2x, 4x respectively.\n\nI found that this was not followed in your code. In my subsequent experiments, following that also had a negative impact. I wonder if this applies to large models rather than small nerwork?",
              "votes": 1
            },
            {
              "id": 2248253,
              "postDate": "2023-05-06T16:05:22.333Z",
              "content": "<p>The original notebook (that I forked) trains batch size 1024 for 250 epochs at <code>LR_MAX = 0.001</code> with <code>MLP_DROPOUT_RATIO = 0.40</code>. After adding data augmentation, reducing epochs to 120, and using <code>MLP_DROPOUT_RATIO = 0.30</code> the optimal <code>LR_MAX</code> may be <code>0.004</code> for the original notebook, i forget. Then we reduce the batch size to 256 and reduce the <code>LR_MAX</code> to <code>0.001</code>. So both batch size and learning rate get decreased by <code>0.25x</code>.</p>\n<p>Note that other things affect learning rate too like data augmentation, drop out rate, number of epochs, the presence of batch normalization or layer normalization in model architecture, shape of learning schedule, etc etc. If all of these things stay the same that usually adjusting <code>LR_MAX</code> proportionally to adjusting batch size works, however it is only a rule of thumb. </p>\n<p>Sometimes when changing batch size we also change other things, and then we might need to search for a new optimal LR. </p>",
              "rawMarkdown": "The original notebook (that I forked) trains batch size 1024 for 250 epochs at `LR_MAX = 0.001` with `MLP_DROPOUT_RATIO = 0.40`. After adding data augmentation, reducing epochs to 120, and using `MLP_DROPOUT_RATIO = 0.30` the optimal `LR_MAX` may be `0.004` for the original notebook, i forget. Then we reduce the batch size to 256 and reduce the `LR_MAX` to `0.001`. So both batch size and learning rate get decreased by `0.25x`.\n\nNote that other things affect learning rate too like data augmentation, drop out rate, number of epochs, the presence of batch normalization or layer normalization in model architecture, shape of learning schedule, etc etc. If all of these things stay the same that usually adjusting `LR_MAX` proportionally to adjusting batch size works, however it is only a rule of thumb. \n\nSometimes when changing batch size we also change other things, and then we might need to search for a new optimal LR. ",
              "votes": 1
            },
            {
              "id": 2249629,
              "postDate": "2023-05-08T00:49:46.123Z",
              "content": "<p>Thank you very much for being so specific in your explanations. I carelessly ignore the other changing variables. Btw, I like this term <strong>rule of thumb</strong>. This is something I am currently pursuing.</p>",
              "rawMarkdown": "Thank you very much for being so specific in your explanations. I carelessly ignore the other changing variables. Btw, I like this term **rule of thumb**. This is something I am currently pursuing.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2245179,
      "postDate": "2023-05-04T07:47:08.257Z",
      "content": "<p>Many thanks for the detailed solution chris!</p>",
      "rawMarkdown": "Many thanks for the detailed solution chris!",
      "votes": 1
    },
    {
      "id": 2245130,
      "postDate": "2023-05-04T06:59:01.860Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thank you very much for your sharing the detail of your approach. I could learn a lot from that and this might be good hint for my next challenge!</p>",
      "rawMarkdown": "@cdeotte , thank you very much for your sharing the detail of your approach. I could learn a lot from that and this might be good hint for my next challenge!",
      "votes": 1,
      "replies": [
        {
          "id": 2247060,
          "postDate": "2023-05-05T17:07:45.910Z",
          "content": "<p>Thank you HechtJP!</p>",
          "rawMarkdown": "Thank you HechtJP!"
        }
      ]
    },
    {
      "id": 2244565,
      "postDate": "2023-05-03T18:00:12.273Z",
      "content": "<p>thank you for this post, i learned about you, true kaggle GM, grateful! hope to do awesome things like you &lt;3</p>",
      "rawMarkdown": "thank you for this post, i learned about you, true kaggle GM, grateful! hope to do awesome things like you <3",
      "votes": 1
    },
    {
      "id": 2244088,
      "postDate": "2023-05-03T12:35:52.210Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <br>\nOne week is really less, have you used any library to tune hyperparams or done manually by changing every value.<br>\nThanks</p>",
      "rawMarkdown": "Hi @cdeotte, \nOne week is really less, have you used any library to tune hyperparams or done manually by changing every value.\nThanks",
      "votes": 1,
      "replies": [
        {
          "id": 2244212,
          "postDate": "2023-05-03T14:09:29.283Z",
          "content": "<p>Thank you. I tune everything manually. I change one value and run again. When <code>INPUT_SIZE = 12</code> and <code>N_EPOCHS = 120</code>, each experiment takes about 45 minutes on 1xV100 GPU.</p>",
          "rawMarkdown": "Thank you. I tune everything manually. I change one value and run again. When `INPUT_SIZE = 12` and `N_EPOCHS = 120`, each experiment takes about 45 minutes on 1xV100 GPU.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2243837,
      "postDate": "2023-05-03T07:58:23.750Z",
      "content": "<p>Congrats! Most respect from the guy doing the same job but much less successful. <br>\nWondering how could you manage to have so many experiments within one week due to the 30h GPU quota(new kaggler here and had big trouble after running out of GPU time haha)</p>",
      "rawMarkdown": "Congrats! Most respect from the guy doing the same job but much less successful. \nWondering how could you manage to have so many experiments within one week due to the 30h GPU quota(new kaggler here and had big trouble after running out of GPU time haha)",
      "votes": 1,
      "replies": [
        {
          "id": 2244217,
          "postDate": "2023-05-03T14:13:40.583Z",
          "content": "<p>Thanks Junjie. I use GPU outside of Kaggle which gives a big advantage. Locally, I used 8xV100 GPU. So each day, I had 24x8 = 192 hours of GPU compute. In one week I ran about 300 experiments.</p>",
          "rawMarkdown": "Thanks Junjie. I use GPU outside of Kaggle which gives a big advantage. Locally, I used 8xV100 GPU. So each day, I had 24x8 = 192 hours of GPU compute. In one week I ran about 300 experiments.",
          "votes": 4,
          "replies": [
            {
              "id": 2245201,
              "postDate": "2023-05-04T08:02:17.597Z",
              "content": "<p>Oh I see now, that explains a lot👀. Thanks!</p>",
              "rawMarkdown": "Oh I see now, that explains a lot👀. Thanks!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2242997,
      "postDate": "2023-05-02T16:19:19.720Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> congrats with the medal and great appreciation for sharing your explanation and code!👍🤝 Especially your thoughts about parameters tuning is very useful.  Your posts and notebooks in this and Jo Wilder competitions help me a lot!🙂</p>",
      "rawMarkdown": "@cdeotte congrats with the medal and great appreciation for sharing your explanation and code!👍🤝 Especially your thoughts about parameters tuning is very useful.  Your posts and notebooks in this and Jo Wilder competitions help me a lot!🙂",
      "votes": 1
    },
    {
      "id": 2242939,
      "postDate": "2023-05-02T15:31:07.790Z",
      "content": "<p>Thanks for sharing the though process <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>\n<p>Can you provide some more thoughts on this line? I am trying to understand why this works so well</p>\n<blockquote>\n  <p>\"easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds).\"</p>\n</blockquote>\n<p>In my case, I tried to ensemble w/ 2 different transformer structures (embedding dimensions, units, etc), but found that an identical model structure trained on 2 different seeds worked best. </p>",
      "rawMarkdown": "Thanks for sharing the though process @cdeotte.\n\nCan you provide some more thoughts on this line? I am trying to understand why this works so well\n\n>\"easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds).\"\n\nIn my case, I tried to ensemble w/ 2 different transformer structures (embedding dimensions, units, etc), but found that an identical model structure trained on 2 different seeds worked best. ",
      "votes": 1,
      "replies": [
        {
          "id": 2242982,
          "postDate": "2023-05-02T16:06:13.703Z",
          "content": "<p>Hi. To clarify, I am saying \"ensemble with itself\" is better than \"submit single model\". Note that I wasn't saying that \"ensemble with itself\" is better than \"ensemble with different NN\". </p>\n<p>Regarding NN, there is a big difference between submitting 1 NN and 4 NN. The choice for the 4 NN can be same or different, but we get a significant boost using 4 vs. 1. (This is not the case for XGB for example because each XGB is already composed of many trees).</p>\n<p>The reason I say \"with itself\" is \"easiest\" is because it doesn't require us to design a new NN. However, ensemble with other strong NN will usually perform better than \"ensemble with itself\".</p>\n<p>In actuality, my final sub is 4 models with <code>INPUT_SIZE = 24, 16, 12, 12</code>. Using different sizes versus using <code>INPUT_SIZE = 12, 12, 12, 12</code> achieved slightly better CV and better LB. (And my second final sub is <code>INPUT_SIZE = 8, 10, 12, 16, 24</code>).</p>",
          "rawMarkdown": "Hi. To clarify, I am saying \"ensemble with itself\" is better than \"submit single model\". Note that I wasn't saying that \"ensemble with itself\" is better than \"ensemble with different NN\". \n\nRegarding NN, there is a big difference between submitting 1 NN and 4 NN. The choice for the 4 NN can be same or different, but we get a significant boost using 4 vs. 1. (This is not the case for XGB for example because each XGB is already composed of many trees).\n\nThe reason I say \"with itself\" is \"easiest\" is because it doesn't require us to design a new NN. However, ensemble with other strong NN will usually perform better than \"ensemble with itself\".\n\nIn actuality, my final sub is 4 models with `INPUT_SIZE = 24, 16, 12, 12`. Using different sizes versus using `INPUT_SIZE = 12, 12, 12, 12` achieved slightly better CV and better LB. (And my second final sub is `INPUT_SIZE = 8, 10, 12, 16, 24`).",
          "votes": 3,
          "replies": [
            {
              "id": 2243166,
              "postDate": "2023-05-02T17:58:13.077Z",
              "content": "<p>Ok that makes sense. Thanks for the follow up!</p>\n<p>Found a good article <a href=\"https://pechyonkin.me/stochastic-weight-averaging/\" target=\"_blank\">here</a> by Max Pechyonkin that talks about Ensembling NNs if anyone wants more info</p>",
              "rawMarkdown": "Ok that makes sense. Thanks for the follow up!\n\nFound a good article [here](https://pechyonkin.me/stochastic-weight-averaging/) by Max Pechyonkin that talks about Ensembling NNs if anyone wants more info",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2242374,
      "postDate": "2023-05-02T08:11:20.300Z",
      "content": "<blockquote>\n  <p>INPUT_SIZE, 64 =&gt; 12</p>\n</blockquote>\n<p>This blows my mins 🤯 your code interpolates all frames with a visible dominant hand to just 12 frames?</p>\n<p>Thanks for explaining your thought process. Very interesting to read. Congrats on the nice result in a short time :)</p>",
      "rawMarkdown": ">INPUT_SIZE, 64 => 12\n\nThis blows my mins 🤯 your code interpolates all frames with a visible dominant hand to just 12 frames?\n\nThanks for explaining your thought process. Very interesting to read. Congrats on the nice result in a short time :)",
      "votes": 1,
      "replies": [
        {
          "id": 2242989,
          "postDate": "2023-05-02T16:11:00.420Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> Congrats on your result.</p>\n<p>It is surprising that using only 12 frames works well. I also have a model with <code>INPUT_SIZE = 8</code> and it achieves <code>LB = 0.76</code>. I didn't try smaller but i suspect that 6, 4, 3, 2, 1 might work too with good result.</p>\n<p>I guess if we knew sign language then perhaps we (as humans) could look at a few frames and predict which of the 250 words is being signed.</p>",
          "rawMarkdown": "Thanks @crodoc Congrats on your result.\n\nIt is surprising that using only 12 frames works well. I also have a model with `INPUT_SIZE = 8` and it achieves `LB = 0.76`. I didn't try smaller but i suspect that 6, 4, 3, 2, 1 might work too with good result.\n\nI guess if we knew sign language then perhaps we (as humans) could look at a few frames and predict which of the 250 words is being signed.",
          "votes": 1,
          "replies": [
            {
              "id": 2243247,
              "postDate": "2023-05-02T18:48:04.773Z",
              "content": "<p>Which interpolation method did you use? The code from the public kernel or?</p>",
              "rawMarkdown": "Which interpolation method did you use? The code from the public kernel or?",
              "votes": 1
            },
            {
              "id": 2243253,
              "postDate": "2023-05-02T18:53:28.113Z",
              "content": "<p>Yes the code from the public notebook with one important change. The public notebook truncates the original video to <code>INPUT_SIZE ** 2</code> frames before interpolation. So to make small <code>INPUT_SIZE&lt;12</code> work we need to increase the truncation. </p>\n<p>When <code>INPUT_SIZE = 12</code> then it first truncates all videos to 144 frames before applying interpolation. To make INPUT_SIZE =8 and 10 to work, i do not truncate to 64 and 100. Instead I truncate to 160 and 200 before interpolation.</p>",
              "rawMarkdown": "Yes the code from the public notebook with one important change. The public notebook truncates the original video to `INPUT_SIZE ** 2` frames before interpolation. So to make small `INPUT_SIZE<12` work we need to increase the truncation. \n\nWhen `INPUT_SIZE = 12` then it first truncates all videos to 144 frames before applying interpolation. To make INPUT_SIZE =8 and 10 to work, i do not truncate to 64 and 100. Instead I truncate to 160 and 200 before interpolation.",
              "votes": 1
            },
            {
              "id": 2243259,
              "postDate": "2023-05-02T19:02:01.947Z",
              "content": "<p>Note that many of the videos are short. I think 40% are less than 12 frames!</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2023/frames.png\" alt=\"\"></p>",
              "rawMarkdown": "Note that many of the videos are short. I think 40% are less than 12 frames!\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2023/frames.png)",
              "votes": 2
            },
            {
              "id": 2243381,
              "postDate": "2023-05-02T21:37:09.567Z",
              "content": "<p>Still blows my mind :)</p>",
              "rawMarkdown": "Still blows my mind :)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2242143,
      "postDate": "2023-05-02T04:13:15.407Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the topic and detailed info about improvements!</p>",
      "rawMarkdown": "@cdeotte thanks for the topic and detailed info about improvements!",
      "votes": 1
    },
    {
      "id": 2241901,
      "postDate": "2023-05-02T00:18:10.597Z",
      "content": "<p>Congrats Chris 👏<br>\nWould you mind giving some details about the choice of these parameters values? Or it is just based on experiments and cv? Thanks</p>",
      "rawMarkdown": "Congrats Chris 👏\nWould you mind giving some details about the choice of these parameters values? Or it is just based on experiments and cv? Thanks",
      "votes": 2,
      "replies": [
        {
          "id": 2241911,
          "postDate": "2023-05-02T00:33:00.103Z",
          "content": "<p>Great question. I added my \"thought process\" to the discussion post above</p>",
          "rawMarkdown": "Great question. I added my \"thought process\" to the discussion post above",
          "votes": 3,
          "replies": [
            {
              "id": 2248226,
              "postDate": "2023-05-06T15:45:53.003Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks, it was really helpful.<br>\nJust have one more question. What do you think are good rules of thumb for choosing the type of learning rate scheduler?<br>\nI found a lot of them in notebooks like this: <a href=\"https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook\" target=\"_blank\">https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook</a></p>",
              "rawMarkdown": "@cdeotte Thanks, it was really helpful.\nJust have one more question. What do you think are good rules of thumb for choosing the type of learning rate scheduler?\nI found a lot of them in notebooks like this: https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook",
              "votes": 1
            },
            {
              "id": 2248264,
              "postDate": "2023-05-06T16:36:06.060Z",
              "content": "<p>If training requires only a few epochs (i.e. 1, 2, 3, 4, or 5 epochs) like NLP models, I usually just go with constant LR. Then i look at validation metric and if it increases in the last epoch or two, perhaps I will use <code>LR</code> for the first epochs and <code>LR * 0.1</code> (or 0.3) for the last epochs.</p>\n<p>When a model requires many epochs like 20, 40, 80, 160. I usually go with cosine schedule (like this comp). These two techniques work well in most cases.</p>",
              "rawMarkdown": "If training requires only a few epochs (i.e. 1, 2, 3, 4, or 5 epochs) like NLP models, I usually just go with constant LR. Then i look at validation metric and if it increases in the last epoch or two, perhaps I will use `LR` for the first epochs and `LR * 0.1` (or 0.3) for the last epochs.\n\nWhen a model requires many epochs like 20, 40, 80, 160. I usually go with cosine schedule (like this comp). These two techniques work well in most cases.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2257110,
      "postDate": "2023-05-13T03:44:30.647Z",
      "content": "<p>you had a great expirience</p>",
      "rawMarkdown": "you had a great expirience"
    },
    {
      "id": 2254657,
      "postDate": "2023-05-11T06:32:20.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2242668,
      "postDate": "2023-05-02T12:41:57.680Z",
      "content": "<p>thanks for the explanation!</p>",
      "rawMarkdown": "thanks for the explanation!",
      "votes": 1
    },
    {
      "id": 2254648,
      "postDate": "2023-05-11T06:27:05.573Z",
      "content": "<p>thanks for detailed info</p>",
      "rawMarkdown": "thanks for detailed info"
    }
  ],
  "comments": [
    {
      "id": 2244633,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-05-03T19:09:54.453000",
      "content": "<p>I still remember how last year, during the ventilator competition, you casually mentioned at the end, \"If you just take the top public notebook and reduce the batch size, it gets to the silver zone.\"</p>\n<p><strong>Reducing the batch size never gets old!</strong></p>\n<p>Amazing work, as always!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2244872,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-04T00:36:50.063000",
          "content": "<p>Thanks Devastator!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2242849,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-05-02T14:38:33.180000",
      "content": "<p><strong>UPDATE:</strong> I shared my code in notebook <a href=\"https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76\" target=\"_blank\">here</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2400466,
      "author_name": "DongYK",
      "author_url": "",
      "post_date": "2023-08-21T04:54:50.517000",
      "content": "<blockquote>\n  <p>Most Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.</p>\n</blockquote>\n<p>Great Insight. Thank you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2247677,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-05-06T07:57:36.553000",
      "content": "<p>I was not active in this competition but reading solution description to learn as much as possible. Your approach for improving public notebook is brilliant. We have many things to touch … many to introduce but great people manage to find appropriate things which lead to great results. Thank you for sharing your workflow. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2247434,
      "author_name": "Ju7on9",
      "author_url": "",
      "post_date": "2023-05-06T02:11:05.763000",
      "content": "<p>Thanks for sharing great insight of the competition. </p>\n<p>Would I ask one question?</p>\n<blockquote>\n  <p>I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about +0.01 or '0.02. Reducing the MLP_RATIO from 4 to 3 significantly reduced the number of model parameters without affecting CV LB.</p>\n</blockquote>\n<p>I don't understand why increasing the number of blocks and decreasing MLP_RATIO leads to boosting the CV and LB. What made you think of increasing block size while decreasing MLP_RATIO?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2248251,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-06T15:56:16.093000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ju7on9\" target=\"_blank\">@ju7on9</a> . The original network architecture is defined by</p>\n<pre><code># Dense layer units for landmarks\nLIPS_UNITS = 224 # WAS 384\nHANDS_UNITS = 224 # WAS 384\nPOSE_UNITS = 224 # WAS 384\n# final embedding and transformer embedding size\nUNITS = 376 # WAS 512\n\n# Transformer\nNUM_BLOCKS = 3 # WAS 2\nMLP_RATIO = 3 # WAS 4\n</code></pre>\n<p>I changed each for a different reason. I lowered <code>MLP_RATIO</code> to shrink the model and I raised <code>NUM_BLOCKS</code> to improve CV and LB. In my discussion, i should say reducing <code>LANDMARK_UNITS, UNITS, and MLP_RATIO</code> gave me extra room to increase <code>NUM_BLOCKS</code>. The point is that increasing <code>NUM_BLOCKS</code> will add 50% more parameters! So we must find a way to significantly reduce the size of the model before we can do that. </p>\n<p>The original network had 7.4M parameters and our goal is to get one NN under 5M. We can reduce <code>MLP_RATIO</code> from <code>4 =&gt; 3</code> without degrading CV nor LB. This reduces the network to 6.3M parameters. Next decreasing <code>LANDMARK_UNITS</code> from <code>384 =&gt; 224</code> reduces to 5.9M parameters. And decreasing <code>UNITS</code> from <code>512 =&gt; 384</code> reduces to 3.4M parameters. And neither of these reductions hurt CV nor LB either.</p>\n<p>We now have extra room to increase <code>NUM_BLOCKS</code>. Increasing <code>NUM_BLOCKS</code> from <code>2 =&gt; 3</code> increases parameters to 4.9M parameters. And boosts CV and LB. We are now under 5M parameters and can safely ensemble 4xNN at <code>FP16</code> within the restriction of 40MB size.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2249000,
              "author_name": "Ju7on9",
              "author_url": "",
              "post_date": "2023-05-07T12:02:56.277000",
              "content": "<p>Thank you for the detailed answer! It really helped me a lot!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2245288,
      "author_name": "Jackson You",
      "author_url": "",
      "post_date": "2023-05-04T09:13:42.523000",
      "content": "<p>Congrats and thanks for sharing your experience. I learn a lot from your experience.</p>\n<p>I'm a bit confused by the phrase \"trained with different seeds\". Do you mean the SEED for each batch of data read or the SEED for the initialization of parameters inside the model? In other words, is this SEED that we don't have to set manually?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2245578,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-04T13:00:03.393000",
          "content": "<blockquote>\n  <p>is this SEED that we don't have to set manually?</p>\n</blockquote>\n<p>Yes, there is nothing for us to change. The \"SEED\" is for the initialization of NN layers. And the shuffle of batches during training. And the randomness of augmentation etc etc. It is all the random elements during training. So if this \"SEED\" changes then each NN will be trained differently.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2247793,
              "author_name": "Jackson You",
              "author_url": "",
              "post_date": "2023-05-06T09:28:24.313000",
              "content": "<blockquote>\n  <p>We should always try 0.25x, 0.5x, 2x, 4x batch size and change the learning rate for those experiments to be 0.25x, 0.5x, 2x, 4x respectively.</p>\n</blockquote>\n<p>I found that this was not followed in your code. In my subsequent experiments, following that also had a negative impact. I wonder if this applies to large models rather than small nerwork?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2248253,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-05-06T16:05:22.333000",
              "content": "<p>The original notebook (that I forked) trains batch size 1024 for 250 epochs at <code>LR_MAX = 0.001</code> with <code>MLP_DROPOUT_RATIO = 0.40</code>. After adding data augmentation, reducing epochs to 120, and using <code>MLP_DROPOUT_RATIO = 0.30</code> the optimal <code>LR_MAX</code> may be <code>0.004</code> for the original notebook, i forget. Then we reduce the batch size to 256 and reduce the <code>LR_MAX</code> to <code>0.001</code>. So both batch size and learning rate get decreased by <code>0.25x</code>.</p>\n<p>Note that other things affect learning rate too like data augmentation, drop out rate, number of epochs, the presence of batch normalization or layer normalization in model architecture, shape of learning schedule, etc etc. If all of these things stay the same that usually adjusting <code>LR_MAX</code> proportionally to adjusting batch size works, however it is only a rule of thumb. </p>\n<p>Sometimes when changing batch size we also change other things, and then we might need to search for a new optimal LR. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2249629,
              "author_name": "Jackson You",
              "author_url": "",
              "post_date": "2023-05-08T00:49:46.123000",
              "content": "<p>Thank you very much for being so specific in your explanations. I carelessly ignore the other changing variables. Btw, I like this term <strong>rule of thumb</strong>. This is something I am currently pursuing.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2245179,
      "author_name": "Punch Line",
      "author_url": "",
      "post_date": "2023-05-04T07:47:08.257000",
      "content": "<p>Many thanks for the detailed solution chris!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2245130,
      "author_name": "HechtJP",
      "author_url": "",
      "post_date": "2023-05-04T06:59:01.860000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thank you very much for your sharing the detail of your approach. I could learn a lot from that and this might be good hint for my next challenge!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2247060,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-05T17:07:45.910000",
          "content": "<p>Thank you HechtJP!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2244565,
      "author_name": "Anirban Chatterjee",
      "author_url": "",
      "post_date": "2023-05-03T18:00:12.273000",
      "content": "<p>thank you for this post, i learned about you, true kaggle GM, grateful! hope to do awesome things like you &lt;3</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2244088,
      "author_name": "Gowri Shankar Penugonda",
      "author_url": "",
      "post_date": "2023-05-03T12:35:52.210000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <br>\nOne week is really less, have you used any library to tune hyperparams or done manually by changing every value.<br>\nThanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244212,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-03T14:09:29.283000",
          "content": "<p>Thank you. I tune everything manually. I change one value and run again. When <code>INPUT_SIZE = 12</code> and <code>N_EPOCHS = 120</code>, each experiment takes about 45 minutes on 1xV100 GPU.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2243837,
      "author_name": "Junjie Huang",
      "author_url": "",
      "post_date": "2023-05-03T07:58:23.750000",
      "content": "<p>Congrats! Most respect from the guy doing the same job but much less successful. <br>\nWondering how could you manage to have so many experiments within one week due to the 30h GPU quota(new kaggler here and had big trouble after running out of GPU time haha)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244217,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-03T14:13:40.583000",
          "content": "<p>Thanks Junjie. I use GPU outside of Kaggle which gives a big advantage. Locally, I used 8xV100 GPU. So each day, I had 24x8 = 192 hours of GPU compute. In one week I ran about 300 experiments.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2245201,
              "author_name": "Junjie Huang",
              "author_url": "",
              "post_date": "2023-05-04T08:02:17.597000",
              "content": "<p>Oh I see now, that explains a lot👀. Thanks!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2242997,
      "author_name": "Ivan Isaev",
      "author_url": "",
      "post_date": "2023-05-02T16:19:19.720000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> congrats with the medal and great appreciation for sharing your explanation and code!👍🤝 Especially your thoughts about parameters tuning is very useful.  Your posts and notebooks in this and Jo Wilder competitions help me a lot!🙂</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2242939,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2023-05-02T15:31:07.790000",
      "content": "<p>Thanks for sharing the though process <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>\n<p>Can you provide some more thoughts on this line? I am trying to understand why this works so well</p>\n<blockquote>\n  <p>\"easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds).\"</p>\n</blockquote>\n<p>In my case, I tried to ensemble w/ 2 different transformer structures (embedding dimensions, units, etc), but found that an identical model structure trained on 2 different seeds worked best. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2242982,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-02T16:06:13.703000",
          "content": "<p>Hi. To clarify, I am saying \"ensemble with itself\" is better than \"submit single model\". Note that I wasn't saying that \"ensemble with itself\" is better than \"ensemble with different NN\". </p>\n<p>Regarding NN, there is a big difference between submitting 1 NN and 4 NN. The choice for the 4 NN can be same or different, but we get a significant boost using 4 vs. 1. (This is not the case for XGB for example because each XGB is already composed of many trees).</p>\n<p>The reason I say \"with itself\" is \"easiest\" is because it doesn't require us to design a new NN. However, ensemble with other strong NN will usually perform better than \"ensemble with itself\".</p>\n<p>In actuality, my final sub is 4 models with <code>INPUT_SIZE = 24, 16, 12, 12</code>. Using different sizes versus using <code>INPUT_SIZE = 12, 12, 12, 12</code> achieved slightly better CV and better LB. (And my second final sub is <code>INPUT_SIZE = 8, 10, 12, 16, 24</code>).</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2243166,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2023-05-02T17:58:13.077000",
              "content": "<p>Ok that makes sense. Thanks for the follow up!</p>\n<p>Found a good article <a href=\"https://pechyonkin.me/stochastic-weight-averaging/\" target=\"_blank\">here</a> by Max Pechyonkin that talks about Ensembling NNs if anyone wants more info</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2242374,
      "author_name": "CroDoc",
      "author_url": "",
      "post_date": "2023-05-02T08:11:20.300000",
      "content": "<blockquote>\n  <p>INPUT_SIZE, 64 =&gt; 12</p>\n</blockquote>\n<p>This blows my mins 🤯 your code interpolates all frames with a visible dominant hand to just 12 frames?</p>\n<p>Thanks for explaining your thought process. Very interesting to read. Congrats on the nice result in a short time :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2242989,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-02T16:11:00.420000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> Congrats on your result.</p>\n<p>It is surprising that using only 12 frames works well. I also have a model with <code>INPUT_SIZE = 8</code> and it achieves <code>LB = 0.76</code>. I didn't try smaller but i suspect that 6, 4, 3, 2, 1 might work too with good result.</p>\n<p>I guess if we knew sign language then perhaps we (as humans) could look at a few frames and predict which of the 250 words is being signed.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2243247,
              "author_name": "CroDoc",
              "author_url": "",
              "post_date": "2023-05-02T18:48:04.773000",
              "content": "<p>Which interpolation method did you use? The code from the public kernel or?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2243253,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-05-02T18:53:28.113000",
              "content": "<p>Yes the code from the public notebook with one important change. The public notebook truncates the original video to <code>INPUT_SIZE ** 2</code> frames before interpolation. So to make small <code>INPUT_SIZE&lt;12</code> work we need to increase the truncation. </p>\n<p>When <code>INPUT_SIZE = 12</code> then it first truncates all videos to 144 frames before applying interpolation. To make INPUT_SIZE =8 and 10 to work, i do not truncate to 64 and 100. Instead I truncate to 160 and 200 before interpolation.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2243259,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-05-02T19:02:01.947000",
              "content": "<p>Note that many of the videos are short. I think 40% are less than 12 frames!</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2023/frames.png\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2243381,
              "author_name": "CroDoc",
              "author_url": "",
              "post_date": "2023-05-02T21:37:09.567000",
              "content": "<p>Still blows my mind :)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2242143,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "2023-05-02T04:13:15.407000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the topic and detailed info about improvements!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2241901,
      "author_name": "Mohamed Eltayeb",
      "author_url": "",
      "post_date": "2023-05-02T00:18:10.597000",
      "content": "<p>Congrats Chris 👏<br>\nWould you mind giving some details about the choice of these parameters values? Or it is just based on experiments and cv? Thanks</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2241911,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-05-02T00:33:00.103000",
          "content": "<p>Great question. I added my \"thought process\" to the discussion post above</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2248226,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2023-05-06T15:45:53.003000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks, it was really helpful.<br>\nJust have one more question. What do you think are good rules of thumb for choosing the type of learning rate scheduler?<br>\nI found a lot of them in notebooks like this: <a href=\"https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook\" target=\"_blank\">https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling/notebook</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2248264,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-05-06T16:36:06.060000",
              "content": "<p>If training requires only a few epochs (i.e. 1, 2, 3, 4, or 5 epochs) like NLP models, I usually just go with constant LR. Then i look at validation metric and if it increases in the last epoch or two, perhaps I will use <code>LR</code> for the first epochs and <code>LR * 0.1</code> (or 0.3) for the last epochs.</p>\n<p>When a model requires many epochs like 20, 40, 80, 160. I usually go with cosine schedule (like this comp). These two techniques work well in most cases.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2257110,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-13T03:44:30.647000",
      "content": "<p>you had a great expirience</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2254657,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-11T06:32:20.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2242668,
      "author_name": "zhuhx666",
      "author_url": "",
      "post_date": "2023-05-02T12:41:57.680000",
      "content": "<p>thanks for the explanation!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2254648,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-11T06:27:05.573000",
      "content": "<p>thanks for detailed info</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2241890": "Thank you Kaggle, Kagglers, PopSign, and Partners for a fun competition! Public notebooks in this competition are very strong! The host said that they need `LB >= 0.6` to be helpful and the best public notebook scores `LB = 0.73`! Great job Kagglers!\n\nI joined this competition one week ago, so I didn't have much time to build my own model. Instead, i read the discussions and public notebooks, and attempted to improve the best public notebook. I boosted the best public notebook from public `LB = 0.73` to `LB = 0.77` and achieved 44th place Silver.\n\n# Praise for Best Public Notebooks\nI am very impressed with the shared public notebooks\n* Great Transformer model architecture\n* Selected which 66 Landmarks out of 543 Landmarks are important (hands, lips, pose)\n* Great preprocessing to convert videos of variable length to small fixed length\n* Accurate local validation scheme which estimates LB score\n\n# Improvements to Best Public Notebook\nAfter reading the best public notebook, I made the following 11 changes to boost CV and LB `+0.03` from `LB = 0.73` to `LB = 0.76`:\n\n* Train 1 model => Train 4 models\n* Add Time Scale augmentation\n* Ensemble and apply TFLite FP16 quantization\n* Change the following parameters:\n* INPUT_SIZE, 64 => 12\n* BATCH_ALL_SIGNS_N, 4 => 1\n* N_EPOCHS, 250 => 120\n* LANDMARK_UNITS, 384 => 224\n* UNITS, 512 => 376\n* NUM_BLOCKS, 2 => 3\n* MLP_RATIO, 4 => 3\n* MLP_DROPOUT_RATIO, 0.40 => 0.30\n* remove random frame masking\n\n# Published Code\nI published my code [here][1] for those curious to learn \"How To Improve Best Public Notebook\" from `LB = 0.73` to `LB = 0.76`\n\n# My Thought Process\nMost Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.\n\nThe first thing i did was reduce `LANDMARK_UNITS`, `UNITS`, and `MLP_RATIO` to reduce the model size. I slowly reduced these parameters to see how small I could go without reducing CV score. It is important to reduce the model to less than 5M (i.e. 5 million) parameters because then we can infer 4 copies of the model within the 40MB size restriction when using FP16. With 40MB size limit, the maximum number of parameters during inference with FP16 is 20M. I also tried less than 4M parameters and 3.3M parameters which allows 5 copies and 6 copies respectively (when using FP16 quantization) \n\nSecond, I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about `+0.01` or `'0.02`. Reducing the `MLP_RATIO` from 4 to 3 significantly reduced the number of model parameters without affecting CV LB so this gave me extra parameters to allow increasing the transformer blocks.\n\nNext, I reduced `INPUT_SIZE` to speed up the model to make it infer (and train) faster. The parameter `INPUT_SIZE` is the transformer sequence length. This competition has a inference time constraint of 1 hour (in addition to a 40MB size constraint). Reducing `INPUT_SIZE` will decrease inference time (and does not affect model size). With sequence length 12, we can easily infer 5+ models under 1 hour.\n\nAt this point, i got lucky. Reducing `INPUT_SIZE` actually boosted the CV and LB by `+0.01` or `+0.02` or so. After this I tuned the learning rate, batch size, and learning schedule and found `N_EPOCHS = 120` and `BATCH_ALL_SIGNS_N = 1` to be best. Many Kagglers overlook the fact that changing batch size can make a big difference for models. We should always try 0.25x, 0.5x, 2x, 4x batch size and change the learning rate for those experiments to be 0.25x, 0.5x, 2x, 4x respectively. Experiments demonstrate that `Batch Size = 256` is better than the original `Batch Size = 1024`. This helped `+0.0005` or so.\n\nLastly NN always benefit from more data and data augmentation. I tried external data and many different data augmentation. Using external data did not help me. Regarding augmentation i only got benefit from `frame drop out` and `time scale augmentation`. This helped about `+0.0005` or so. I also tried rotation, scale, shift but this did not help. Flip was not needed since preprocessing normalized all videos to be left handed.\n\n[1]: https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76",
    "2244633": "I still remember how last year, during the ventilator competition, you casually mentioned at the end, \"If you just take the top public notebook and reduce the batch size, it gets to the silver zone.\"\n\n**Reducing the batch size never gets old!**\n\nAmazing work, as always!",
    "2242849": "**UPDATE:** I shared my code in notebook [here][1]\n\n[1]: https://www.kaggle.com/code/cdeotte/improve-best-public-notebook-lb-0-76",
    "2400466": ">Most Kagglers are probably curious why and how I discovered these modifications. So let me share my thinking. First, the easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds). So the easiest way to boost public notebook CV LB is to make the model smaller and make the model faster (because this competition has size and time constraints). Then we can include multiple copies of the NN during inference. And each NN we will train with 100% train data.\n\nGreat Insight. Thank you.",
    "2247677": "I was not active in this competition but reading solution description to learn as much as possible. Your approach for improving public notebook is brilliant. We have many things to touch … many to introduce but great people manage to find appropriate things which lead to great results. Thank you for sharing your workflow. ",
    "2247434": "Thanks for sharing great insight of the competition. \n\nWould I ask one question?\n\n> I made the transformer deeper with more transformer blocks. I increased blocks from 2 to 3 which boosted the CV and LB about +0.01 or '0.02. Reducing the MLP_RATIO from 4 to 3 significantly reduced the number of model parameters without affecting CV LB.\n\nI don't understand why increasing the number of blocks and decreasing MLP_RATIO leads to boosting the CV and LB. What made you think of increasing block size while decreasing MLP_RATIO?",
    "2245288": "Congrats and thanks for sharing your experience. I learn a lot from your experience.\n\nI'm a bit confused by the phrase \"trained with different seeds\". Do you mean the SEED for each batch of data read or the SEED for the initialization of parameters inside the model? In other words, is this SEED that we don't have to set manually?",
    "2245179": "Many thanks for the detailed solution chris!",
    "2245130": "@cdeotte , thank you very much for your sharing the detail of your approach. I could learn a lot from that and this might be good hint for my next challenge!",
    "2244565": "thank you for this post, i learned about you, true kaggle GM, grateful! hope to do awesome things like you <3",
    "2244088": "Hi @cdeotte, \nOne week is really less, have you used any library to tune hyperparams or done manually by changing every value.\nThanks",
    "2243837": "Congrats! Most respect from the guy doing the same job but much less successful. \nWondering how could you manage to have so many experiments within one week due to the 30h GPU quota(new kaggler here and had big trouble after running out of GPU time haha)",
    "2242997": "@cdeotte congrats with the medal and great appreciation for sharing your explanation and code!👍🤝 Especially your thoughts about parameters tuning is very useful.  Your posts and notebooks in this and Jo Wilder competitions help me a lot!🙂",
    "2242939": "Thanks for sharing the though process @cdeotte.\n\nCan you provide some more thoughts on this line? I am trying to understand why this works so well\n\n>\"easiest way to boost CV LB for NN is to ensemble NN with multiple copies of itself (trained with different seeds).\"\n\nIn my case, I tried to ensemble w/ 2 different transformer structures (embedding dimensions, units, etc), but found that an identical model structure trained on 2 different seeds worked best. ",
    "2242374": ">INPUT_SIZE, 64 => 12\n\nThis blows my mins 🤯 your code interpolates all frames with a visible dominant hand to just 12 frames?\n\nThanks for explaining your thought process. Very interesting to read. Congrats on the nice result in a short time :)",
    "2242143": "@cdeotte thanks for the topic and detailed info about improvements!",
    "2241901": "Congrats Chris 👏\nWould you mind giving some details about the choice of these parameters values? Or it is just based on experiments and cv? Thanks",
    "2257110": "you had a great expirience",
    "2254657": "",
    "2242668": "thanks for the explanation!",
    "2254648": "thanks for detailed info"
  }
}