{
  "id": 160896,
  "title": "34th Place Solution - Layerwise and Sampling",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/jun-34th-place-solution-layerwise-and-sampling",
  "author_name": "",
  "post_date": "2020-06-23T19:19:40.627Z",
  "votes": 15,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I want to thank the organizers for this amazing competition and Kaggle for providing v3 TPU where people without hardwares like myself were able to compete without much disadvantage. I also want to congratulate the winners. This was my first Kaggle competition and I learned a lot about SOTA NLP models/techniques through the process. It was tough at first, but it was exciting to see my ideas work and move up the leaderboard. </p>\n\n<h2><strong>Summary</strong></h2>\n\n<ol>\n<li>ensemble of XLM-R models</li>\n<li>layer-wise learning rate/weight decay </li>\n<li>pseudo-labelling</li>\n<li>sampling strategy </li>\n</ol>\n\n<h2><strong>Ensemble of XLM-R models</strong></h2>\n\n<p>For all of my models, I trained on the training data for 2 epochs using [CLS] token without any extra layers. I tried different heads on top but I saw a drop in the score for all of them. I believe this is because the last-layer of a pretrained deep NN is too specific for the previous task (MLM objective in our case). Adding more layers on top just made the training more unstable. </p>\n\n<p>I further finetuned on the validation set for 2 epochs using 5-fold then averaged the oof predictions. This 5-fold CV AUC score was a good indicator of the public LB score. I ensembled only XLM-R models because other multilingual models (like XLM-17, m-BERT) were not very useful. I also trained each model 3 times and averaged the predictions for stability. </p>\n\n<p>Mainly, I trained \n1. XLM-R-large english only model (~0.9356 on public LB)\n2. XLM-R-large translated only model - pretrained more with MLM using test data for 4 epochs (~0.9471 on public LB)\n3. XLM-R-large translated only model with last 50 (~0.9449 on public LB)</p>\n\n<h2><strong>Layer-Wise Learning Rate/Weight Decay</strong></h2>\n\n<p>I tried two different layer-wise strategies: layer-wise learning rate and layer-wise weight decay. I saw that layer-wise learning rate was effective in finetuning BERT in this article  <a href=\"https://arxiv.org/abs/1905.05583\">https://arxiv.org/abs/1905.05583</a> and found it to be very useful for this competition as well. From this result, I also tried out a variant of this layerwise training strategy where I regularized the model more in the lower layers. I found that this worked much better than the layer-wise learning rate (about ~0.008 boost in public LB). It works just like the layer-wise learning rate strategy but instead of multiplying the learning rate by the layerwise coefficient alpha^(total_num_layers-layer_number) (where layer_number = 0 for the first layer), you multiply the weight decay rate by the coefficient alpha^layer_number. Thus, you are regularizing the lower layers more than the upper layers. This made sense because a wide exploration of the parameter space was needed for the upper layers in order to adapt to the new objective while the lower layers didnt need to change much. I used 0.01 for the weight decay and 0.99 for the layer-wise coefficient (alpha). Layer-wise learning rate strategy could actually be better but it was very sensitive to hyperparameter choices, so with limited TPU time I worked with layer-wise weight decay. </p>\n\n<h2><strong>Pseudo-labelling</strong></h2>\n\n<p>I used my best model's predictions on the test set as the pseudo-labels and included in the training set. Even using pseudo-labels of a relatively worse model (scoring ~0.935) was effective. It was interesting to observe the instability of the XLM-R-large model where switching only the pseudo-labels had resulted in diverging of the model. I had to lower the learning rate from 1e-5 to 6e-6 in order to converge.</p>\n\n<h2><strong>Sampling-strategy</strong></h2>\n\n<p>I downsampled training data to 1:1 for all models. For the XLM-R translated only models, I mixed yandex and google translation and saw a little boost in local CV and public LB. I also made sure that when sampling for non-toxic there were no duplicates across different languages for non-toxic.</p>\n\n<h2><strong>Other Ideas that Didn't Work</strong></h2>\n\n<ol>\n<li>text pre-processing</li>\n<li>label-smoothing</li>\n<li>extending the test dataset by translating back from english</li>\n<li>making multiple predictions - predicting at val finetune epoch 1 + val finetune epoch 2 and combining</li>\n<li>meta-learning </li>\n<li>embeddings - I tried using bilingual Attract-Repel embeddings</li>\n</ol>\n\n<p>Lastly, I learned that strictly planning your ideas by ranking them by their importance was really crucial especially when working with only 30 hours of TPU time. I found myself spending majority of my TPU time trying to tweak the working model to make it better rather than trying more of other novel ideas: striking a good balance between Exploitation and Exploration is the key! It was quite sad to see one of my ideas (making expert models using monolingual transformer models) in the middle of my untried idea list in the solution for the 1st place solution, but I have learned my lesson :)</p>\n\n<p>I want to thank the authors of these wonderful kernels for their work. I have learned a lot from them.\n<a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a>\n<a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a>\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>",
  "messages": [
    {
      "id": "897735",
      "postDate": "06/23/2020 04:15:08",
      "content": "<p>I want to thank the organizers for this amazing competition and Kaggle for providing v3 TPU where people without hardwares like myself were able to compete without much disadvantage. I also want to congratulate the winners. This was my first Kaggle competition and I learned a lot about SOTA NLP models/techniques through the process. It was tough at first, but it was exciting to see my ideas work and move up the leaderboard. </p>\n\n<h2><strong>Summary</strong></h2>\n\n<ol>\n<li>ensemble of XLM-R models</li>\n<li>layer-wise learning rate/weight decay </li>\n<li>pseudo-labelling</li>\n<li>sampling strategy </li>\n</ol>\n\n<h2><strong>Ensemble of XLM-R models</strong></h2>\n\n<p>For all of my models, I trained on the training data for 2 epochs using [CLS] token without any extra layers. I tried different heads on top but I saw a drop in the score for all of them. I believe this is because the last-layer of a pretrained deep NN is too specific for the previous task (MLM objective in our case). Adding more layers on top just made the training more unstable. </p>\n\n<p>I further finetuned on the validation set for 2 epochs using 5-fold then averaged the oof predictions. This 5-fold CV AUC score was a good indicator of the public LB score. I ensembled only XLM-R models because other multilingual models (like XLM-17, m-BERT) were not very useful. I also trained each model 3 times and averaged the predictions for stability. </p>\n\n<p>Mainly, I trained \n1. XLM-R-large english only model (~0.9356 on public LB)\n2. XLM-R-large translated only model - pretrained more with MLM using test data for 4 epochs (~0.9471 on public LB)\n3. XLM-R-large translated only model with last 50 (~0.9449 on public LB)</p>\n\n<h2><strong>Layer-Wise Learning Rate/Weight Decay</strong></h2>\n\n<p>I tried two different layer-wise strategies: layer-wise learning rate and layer-wise weight decay. I saw that layer-wise learning rate was effective in finetuning BERT in this article  <a href=\"https://arxiv.org/abs/1905.05583\">https://arxiv.org/abs/1905.05583</a> and found it to be very useful for this competition as well. From this result, I also tried out a variant of this layerwise training strategy where I regularized the model more in the lower layers. I found that this worked much better than the layer-wise learning rate (about ~0.008 boost in public LB). It works just like the layer-wise learning rate strategy but instead of multiplying the learning rate by the layerwise coefficient alpha^(total_num_layers-layer_number) (where layer_number = 0 for the first layer), you multiply the weight decay rate by the coefficient alpha^layer_number. Thus, you are regularizing the lower layers more than the upper layers. This made sense because a wide exploration of the parameter space was needed for the upper layers in order to adapt to the new objective while the lower layers didnt need to change much. I used 0.01 for the weight decay and 0.99 for the layer-wise coefficient (alpha). Layer-wise learning rate strategy could actually be better but it was very sensitive to hyperparameter choices, so with limited TPU time I worked with layer-wise weight decay. </p>\n\n<h2><strong>Pseudo-labelling</strong></h2>\n\n<p>I used my best model's predictions on the test set as the pseudo-labels and included in the training set. Even using pseudo-labels of a relatively worse model (scoring ~0.935) was effective. It was interesting to observe the instability of the XLM-R-large model where switching only the pseudo-labels had resulted in diverging of the model. I had to lower the learning rate from 1e-5 to 6e-6 in order to converge.</p>\n\n<h2><strong>Sampling-strategy</strong></h2>\n\n<p>I downsampled training data to 1:1 for all models. For the XLM-R translated only models, I mixed yandex and google translation and saw a little boost in local CV and public LB. I also made sure that when sampling for non-toxic there were no duplicates across different languages for non-toxic.</p>\n\n<h2><strong>Other Ideas that Didn't Work</strong></h2>\n\n<ol>\n<li>text pre-processing</li>\n<li>label-smoothing</li>\n<li>extending the test dataset by translating back from english</li>\n<li>making multiple predictions - predicting at val finetune epoch 1 + val finetune epoch 2 and combining</li>\n<li>meta-learning </li>\n<li>embeddings - I tried using bilingual Attract-Repel embeddings</li>\n</ol>\n\n<p>Lastly, I learned that strictly planning your ideas by ranking them by their importance was really crucial especially when working with only 30 hours of TPU time. I found myself spending majority of my TPU time trying to tweak the working model to make it better rather than trying more of other novel ideas: striking a good balance between Exploitation and Exploration is the key! It was quite sad to see one of my ideas (making expert models using monolingual transformer models) in the middle of my untried idea list in the solution for the 1st place solution, but I have learned my lesson :)</p>\n\n<p>I want to thank the authors of these wonderful kernels for their work. I have learned a lot from them.\n<a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a>\n<a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a>\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>",
      "rawMarkdown": "I want to thank the organizers for this amazing competition and Kaggle for providing v3 TPU where people without hardwares like myself were able to compete without much disadvantage. I also want to congratulate the winners. This was my first Kaggle competition and I learned a lot about SOTA NLP models/techniques through the process. It was tough at first, but it was exciting to see my ideas work and move up the leaderboard. \n\n## **Summary**\n1. ensemble of XLM-R models\n2. layer-wise learning rate/weight decay \n3. pseudo-labelling\n4. sampling strategy \n\n## **Ensemble of XLM-R models**\nFor all of my models, I trained on the training data for 2 epochs using [CLS] token without any extra layers. I tried different heads on top but I saw a drop in the score for all of them. I believe this is because the last-layer of a pretrained deep NN is too specific for the previous task (MLM objective in our case). Adding more layers on top just made the training more unstable. \n\nI further finetuned on the validation set for 2 epochs using 5-fold then averaged the oof predictions. This 5-fold CV AUC score was a good indicator of the public LB score. I ensembled only XLM-R models because other multilingual models (like XLM-17, m-BERT) were not very useful. I also trained each model 3 times and averaged the predictions for stability. \n\nMainly, I trained \n1. XLM-R-large english only model (~0.9356 on public LB)\n2. XLM-R-large translated only model - pretrained more with MLM using test data for 4 epochs (~0.9471 on public LB)\n3. XLM-R-large translated only model with last 50 (~0.9449 on public LB)\n\n## **Layer-Wise Learning Rate/Weight Decay**\nI tried two different layer-wise strategies: layer-wise learning rate and layer-wise weight decay. I saw that layer-wise learning rate was effective in finetuning BERT in this article  https://arxiv.org/abs/1905.05583 and found it to be very useful for this competition as well. From this result, I also tried out a variant of this layerwise training strategy where I regularized the model more in the lower layers. I found that this worked much better than the layer-wise learning rate (about ~0.008 boost in public LB). It works just like the layer-wise learning rate strategy but instead of multiplying the learning rate by the layerwise coefficient alpha^(total_num_layers-layer_number) (where layer_number = 0 for the first layer), you multiply the weight decay rate by the coefficient alpha^layer_number. Thus, you are regularizing the lower layers more than the upper layers. This made sense because a wide exploration of the parameter space was needed for the upper layers in order to adapt to the new objective while the lower layers didnt need to change much. I used 0.01 for the weight decay and 0.99 for the layer-wise coefficient (alpha). Layer-wise learning rate strategy could actually be better but it was very sensitive to hyperparameter choices, so with limited TPU time I worked with layer-wise weight decay. \n\n##**Pseudo-labelling**\nI used my best model's predictions on the test set as the pseudo-labels and included in the training set. Even using pseudo-labels of a relatively worse model (scoring ~0.935) was effective. It was interesting to observe the instability of the XLM-R-large model where switching only the pseudo-labels had resulted in diverging of the model. I had to lower the learning rate from 1e-5 to 6e-6 in order to converge.\n\n##**Sampling-strategy**\n I downsampled training data to 1:1 for all models. For the XLM-R translated only models, I mixed yandex and google translation and saw a little boost in local CV and public LB. I also made sure that when sampling for non-toxic there were no duplicates across different languages for non-toxic.\n\n##**Other Ideas that Didn't Work**\n1. text pre-processing\n2. label-smoothing\n3. extending the test dataset by translating back from english\n4. making multiple predictions - predicting at val finetune epoch 1 + val finetune epoch 2 and combining\n5. meta-learning \n6. embeddings - I tried using bilingual Attract-Repel embeddings\n\nLastly, I learned that strictly planning your ideas by ranking them by their importance was really crucial especially when working with only 30 hours of TPU time. I found myself spending majority of my TPU time trying to tweak the working model to make it better rather than trying more of other novel ideas: striking a good balance between Exploitation and Exploration is the key! It was quite sad to see one of my ideas (making expert models using monolingual transformer models) in the middle of my untried idea list in the solution for the 1st place solution, but I have learned my lesson :)\n \nI want to thank the authors of these wonderful kernels for their work. I have learned a lot from them.\nhttps://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\nhttps://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta",
      "votes": null
    },
    {
      "id": "897746",
      "postDate": "06/23/2020 04:26:30",
      "content": "<p>Well done !</p>",
      "rawMarkdown": "Well done !",
      "votes": null
    },
    {
      "id": "897772",
      "postDate": "06/23/2020 04:55:31",
      "content": "<p>Thank you for your solution!  i have two question:\n1: About Pseudo-labelling, Did you use the model output(continuous values ​​from 0 to 1) concat with 0/1 training data without any process? For me i add a threshold(like 0.5) to change pseudo-labelling to discrete value, which lead better cv and lb, but i think it's not a good idea.\n2:Did you fix seed when training model? I am confused with this for a long time , since i saw <a href=\"/philippsinger\">@philippsinger</a>  say he never fix seed and embrace random. If we do not fix seed , any idea to try will run for 5 to 10 times to do ensemble? </p>",
      "rawMarkdown": "Thank you for your solution!  i have two question:\n1: About Pseudo-labelling, Did you use the model output(continuous values ​​from 0 to 1) concat with 0/1 training data without any process? For me i add a threshold(like 0.5) to change pseudo-labelling to discrete value, which lead better cv and lb, but i think it's not a good idea.\n2:Did you fix seed when training model? I am confused with this for a long time , since i saw @philippsinger  say he never fix seed and embrace random. If we do not fix seed , any idea to try will run for 5 to 10 times to do ensemble?",
      "votes": null
    },
    {
      "id": "897782",
      "postDate": "06/23/2020 05:03:57",
      "content": "<ol>\n<li>I trained everything with 1s and 0s ( &lt; 0.5 = 0, &gt;= 0.5 = 1)</li>\n<li>I fixed seed for everything when I was training the models initially to test out new ideas, but for my final model I simply trained the model 3 times then averaged the predictions. However, even with fixing  it was not possible to replicate the result entirely because of TPU, so maybe that's why some people opted to just never fixing the seed.</li>\n</ol>",
      "rawMarkdown": "1. I trained everything with 1s and 0s ( &lt; 0.5 = 0, &gt;= 0.5 = 1)\n2. I fixed seed for everything when I was training the models initially to test out new ideas, but for my final model I simply trained the model 3 times then averaged the predictions. However, even with fixing  it was not possible to replicate the result entirely because of TPU, so maybe that's why some people opted to just never fixing the seed.",
      "votes": null
    },
    {
      "id": "897812",
      "postDate": "06/23/2020 05:34:52",
      "content": "<p>Thank you ! Do you remember how much improvement with  Pseudo-labelling ?</p>",
      "rawMarkdown": "Thank you ! Do you remember how much improvement with  Pseudo-labelling ?",
      "votes": null
    },
    {
      "id": "898054",
      "postDate": "06/23/2020 08:51:00",
      "content": "<p>Nice work! Which lr schedule did you use? </p>",
      "rawMarkdown": "Nice work! Which lr schedule did you use?",
      "votes": null
    },
    {
      "id": "898334",
      "postDate": "06/23/2020 13:10:48",
      "content": "<p><a href=\"/alphaecho\">@alphaecho</a> The idea is to ensemble NNs with different seeds. I have never seen a NN that did not get a boost from that.</p>",
      "rawMarkdown": "alphaecho The idea is to ensemble NNs with different seeds. I have never seen a NN that did not get a boost from that.",
      "votes": null
    },
    {
      "id": "898596",
      "postDate": "06/23/2020 16:02:05",
      "content": "<p>I used simple constant decay with warm up. I got the hyperparameter setting from the ROBERTA paper and found it to work well.</p>",
      "rawMarkdown": "I used simple constant decay with warm up. I got the hyperparameter setting from the ROBERTA paper and found it to work well.",
      "votes": null
    },
    {
      "id": "899140",
      "postDate": "06/24/2020 02:54:40",
      "content": "<p>Congrats with a great performance, <a href=\"/hansungj\">@hansungj</a> ! Great ideas and you made them work!\nCould you show a snippet of code of how you accessed layers' weight decays or url reference of an example? What optimizer did you use ?</p>",
      "rawMarkdown": "Congrats with a great performance, @hansungj ! Great ideas and you made them work!\nCould you show a snippet of code of how you accessed layers' weight decays or url reference of an example? What optimizer did you use ?",
      "votes": null
    },
    {
      "id": "899143",
      "postDate": "06/24/2020 03:00:33",
      "content": "<p>Thank you. Seems it is time consuming but also effective</p>",
      "rawMarkdown": "Thank you. Seems it is time consuming but also effective",
      "votes": null
    },
    {
      "id": "899170",
      "postDate": "06/24/2020 03:55:02",
      "content": "<p>I simply modified AdamW implementation by huggingface\n<code>\ndef _decay_weights_op(self, var, learning_rate, apply_state):\n        do_decay = self._do_use_weight_decay(var.name)\n        #added\n        layer_coef = self._get_layer_wise_coefficient(var.name)\n        if do_decay:\n            return var.assign_sub(\n                learning_rate * var * layer_coef * apply_state[(var.device, var.dtype.base_dtype)][\"weight_decay_rate\"],\n                use_locking=self._use_locking,\n            )\n        return tf.no_op()\n</code></p>\n\n<p>Then i added this function\n```</p>\n\n<pre><code>def _get_layer_wise_coefficient(self, param_name):\n    \"\"\"gets layer wise coefficient for the `param_name`.\"\"\"\n\n    if self._layer_wise_coefficient == 1:\n        return tf.math.pow(self._layer_wise_coefficient, 0)\n\n    if self._exclude_from_layer_wise:\n        for r in self._exclude_from_layer_wise:\n            if re.search(r, param_name) is not None:\n                return tf.math.pow(self._layer_wise_coefficient, 0)\n\n    for l in range(self._number_of_layers):\n        if re.search(self._layer_re_name + f'{l}/', param_name) is not None:\n            return tf.math.pow(self._layer_wise_coefficient, l)\n\n    return tf.math.pow(self._layer_wise_coefficient, self._number_of_layers )\n</code></pre>\n\n<p><code>``\nXLM-R-large model by jplu, the layers could be distinguished by</code>'layer_._#/' <code>so I had</code>self._layer_re_name = 'layer_._`</p>",
      "rawMarkdown": "I simply modified AdamW implementation by huggingface\n```\ndef _decay_weights_op(self, var, learning_rate, apply_state):\n        do_decay = self._do_use_weight_decay(var.name)\n        #added\n        layer_coef = self._get_layer_wise_coefficient(var.name)\n        if do_decay:\n            return var.assign_sub(\n                learning_rate * var * layer_coef * apply_state[(var.device, var.dtype.base_dtype)][\"weight_decay_rate\"],\n                use_locking=self._use_locking,\n            )\n        return tf.no_op()\n```\n\nThen i added this function\n```\n\n    def _get_layer_wise_coefficient(self, param_name):\n        \"\"\"gets layer wise coefficient for the `param_name`.\"\"\"\n\n        if self._layer_wise_coefficient == 1:\n            return tf.math.pow(self._layer_wise_coefficient, 0)\n\n        if self._exclude_from_layer_wise:\n            for r in self._exclude_from_layer_wise:\n                if re.search(r, param_name) is not None:\n                    return tf.math.pow(self._layer_wise_coefficient, 0)\n        \n        for l in range(self._number_of_layers):\n            if re.search(self._layer_re_name + f'{l}/', param_name) is not None:\n                return tf.math.pow(self._layer_wise_coefficient, l)\n\n        return tf.math.pow(self._layer_wise_coefficient, self._number_of_layers )\n```\nXLM-R-large model by jplu, the layers could be distinguished by `'layer_._#/' `so I had `self._layer_re_name = 'layer_._`",
      "votes": null
    },
    {
      "id": "937856",
      "postDate": "07/21/2020 07:50:27",
      "content": "<p>Thanks a lot. Very helpful code. </p>",
      "rawMarkdown": "Thanks a lot. Very helpful code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 897746,
      "author_name": "haythemtellili5",
      "author_url": "",
      "post_date": "06/23/2020 04:26:30",
      "content": "<p>Well done !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 897772,
      "author_name": "alphaecho",
      "author_url": "",
      "post_date": "06/23/2020 04:55:31",
      "content": "<p>Thank you for your solution!  i have two question:\n1: About Pseudo-labelling, Did you use the model output(continuous values ​​from 0 to 1) concat with 0/1 training data without any process? For me i add a threshold(like 0.5) to change pseudo-labelling to discrete value, which lead better cv and lb, but i think it's not a good idea.\n2:Did you fix seed when training model? I am confused with this for a long time , since i saw <a href=\"/philippsinger\">@philippsinger</a>  say he never fix seed and embrace random. If we do not fix seed , any idea to try will run for 5 to 10 times to do ensemble? </p>",
      "votes": null,
      "replies": [
        {
          "id": 897782,
          "author_name": "hansungj",
          "author_url": "",
          "post_date": "06/23/2020 05:03:57",
          "content": "<ol>\n<li>I trained everything with 1s and 0s ( &lt; 0.5 = 0, &gt;= 0.5 = 1)</li>\n<li>I fixed seed for everything when I was training the models initially to test out new ideas, but for my final model I simply trained the model 3 times then averaged the predictions. However, even with fixing  it was not possible to replicate the result entirely because of TPU, so maybe that's why some people opted to just never fixing the seed.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 897812,
          "author_name": "alphaecho",
          "author_url": "",
          "post_date": "06/23/2020 05:34:52",
          "content": "<p>Thank you ! Do you remember how much improvement with  Pseudo-labelling ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898334,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/23/2020 13:10:48",
          "content": "<p><a href=\"/alphaecho\">@alphaecho</a> The idea is to ensemble NNs with different seeds. I have never seen a NN that did not get a boost from that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899143,
          "author_name": "alphaecho",
          "author_url": "",
          "post_date": "06/24/2020 03:00:33",
          "content": "<p>Thank you. Seems it is time consuming but also effective</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898054,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/23/2020 08:51:00",
      "content": "<p>Nice work! Which lr schedule did you use? </p>",
      "votes": null,
      "replies": [
        {
          "id": 898596,
          "author_name": "hansungj",
          "author_url": "",
          "post_date": "06/23/2020 16:02:05",
          "content": "<p>I used simple constant decay with warm up. I got the hyperparameter setting from the ROBERTA paper and found it to work well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 899140,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "06/24/2020 02:54:40",
      "content": "<p>Congrats with a great performance, <a href=\"/hansungj\">@hansungj</a> ! Great ideas and you made them work!\nCould you show a snippet of code of how you accessed layers' weight decays or url reference of an example? What optimizer did you use ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 899170,
          "author_name": "hansungj",
          "author_url": "",
          "post_date": "06/24/2020 03:55:02",
          "content": "<p>I simply modified AdamW implementation by huggingface\n<code>\ndef _decay_weights_op(self, var, learning_rate, apply_state):\n        do_decay = self._do_use_weight_decay(var.name)\n        #added\n        layer_coef = self._get_layer_wise_coefficient(var.name)\n        if do_decay:\n            return var.assign_sub(\n                learning_rate * var * layer_coef * apply_state[(var.device, var.dtype.base_dtype)][\"weight_decay_rate\"],\n                use_locking=self._use_locking,\n            )\n        return tf.no_op()\n</code></p>\n\n<p>Then i added this function\n```</p>\n\n<pre><code>def _get_layer_wise_coefficient(self, param_name):\n    \"\"\"gets layer wise coefficient for the `param_name`.\"\"\"\n\n    if self._layer_wise_coefficient == 1:\n        return tf.math.pow(self._layer_wise_coefficient, 0)\n\n    if self._exclude_from_layer_wise:\n        for r in self._exclude_from_layer_wise:\n            if re.search(r, param_name) is not None:\n                return tf.math.pow(self._layer_wise_coefficient, 0)\n\n    for l in range(self._number_of_layers):\n        if re.search(self._layer_re_name + f'{l}/', param_name) is not None:\n            return tf.math.pow(self._layer_wise_coefficient, l)\n\n    return tf.math.pow(self._layer_wise_coefficient, self._number_of_layers )\n</code></pre>\n\n<p><code>``\nXLM-R-large model by jplu, the layers could be distinguished by</code>'layer_._#/' <code>so I had</code>self._layer_re_name = 'layer_._`</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937856,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "07/21/2020 07:50:27",
          "content": "<p>Thanks a lot. Very helpful code. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "897735": "I want to thank the organizers for this amazing competition and Kaggle for providing v3 TPU where people without hardwares like myself were able to compete without much disadvantage. I also want to congratulate the winners. This was my first Kaggle competition and I learned a lot about SOTA NLP models/techniques through the process. It was tough at first, but it was exciting to see my ideas work and move up the leaderboard. \n\n## **Summary**\n1. ensemble of XLM-R models\n2. layer-wise learning rate/weight decay \n3. pseudo-labelling\n4. sampling strategy \n\n## **Ensemble of XLM-R models**\nFor all of my models, I trained on the training data for 2 epochs using [CLS] token without any extra layers. I tried different heads on top but I saw a drop in the score for all of them. I believe this is because the last-layer of a pretrained deep NN is too specific for the previous task (MLM objective in our case). Adding more layers on top just made the training more unstable. \n\nI further finetuned on the validation set for 2 epochs using 5-fold then averaged the oof predictions. This 5-fold CV AUC score was a good indicator of the public LB score. I ensembled only XLM-R models because other multilingual models (like XLM-17, m-BERT) were not very useful. I also trained each model 3 times and averaged the predictions for stability. \n\nMainly, I trained \n1. XLM-R-large english only model (~0.9356 on public LB)\n2. XLM-R-large translated only model - pretrained more with MLM using test data for 4 epochs (~0.9471 on public LB)\n3. XLM-R-large translated only model with last 50 (~0.9449 on public LB)\n\n## **Layer-Wise Learning Rate/Weight Decay**\nI tried two different layer-wise strategies: layer-wise learning rate and layer-wise weight decay. I saw that layer-wise learning rate was effective in finetuning BERT in this article  https://arxiv.org/abs/1905.05583 and found it to be very useful for this competition as well. From this result, I also tried out a variant of this layerwise training strategy where I regularized the model more in the lower layers. I found that this worked much better than the layer-wise learning rate (about ~0.008 boost in public LB). It works just like the layer-wise learning rate strategy but instead of multiplying the learning rate by the layerwise coefficient alpha^(total_num_layers-layer_number) (where layer_number = 0 for the first layer), you multiply the weight decay rate by the coefficient alpha^layer_number. Thus, you are regularizing the lower layers more than the upper layers. This made sense because a wide exploration of the parameter space was needed for the upper layers in order to adapt to the new objective while the lower layers didnt need to change much. I used 0.01 for the weight decay and 0.99 for the layer-wise coefficient (alpha). Layer-wise learning rate strategy could actually be better but it was very sensitive to hyperparameter choices, so with limited TPU time I worked with layer-wise weight decay. \n\n##**Pseudo-labelling**\nI used my best model's predictions on the test set as the pseudo-labels and included in the training set. Even using pseudo-labels of a relatively worse model (scoring ~0.935) was effective. It was interesting to observe the instability of the XLM-R-large model where switching only the pseudo-labels had resulted in diverging of the model. I had to lower the learning rate from 1e-5 to 6e-6 in order to converge.\n\n##**Sampling-strategy**\n I downsampled training data to 1:1 for all models. For the XLM-R translated only models, I mixed yandex and google translation and saw a little boost in local CV and public LB. I also made sure that when sampling for non-toxic there were no duplicates across different languages for non-toxic.\n\n##**Other Ideas that Didn't Work**\n1. text pre-processing\n2. label-smoothing\n3. extending the test dataset by translating back from english\n4. making multiple predictions - predicting at val finetune epoch 1 + val finetune epoch 2 and combining\n5. meta-learning \n6. embeddings - I tried using bilingual Attract-Repel embeddings\n\nLastly, I learned that strictly planning your ideas by ranking them by their importance was really crucial especially when working with only 30 hours of TPU time. I found myself spending majority of my TPU time trying to tweak the working model to make it better rather than trying more of other novel ideas: striking a good balance between Exploitation and Exploration is the key! It was quite sad to see one of my ideas (making expert models using monolingual transformer models) in the middle of my untried idea list in the solution for the 1st place solution, but I have learned my lesson :)\n \nI want to thank the authors of these wonderful kernels for their work. I have learned a lot from them.\nhttps://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\nhttps://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta",
    "897746": "Well done !",
    "897772": "Thank you for your solution!  i have two question:\n1: About Pseudo-labelling, Did you use the model output(continuous values ​​from 0 to 1) concat with 0/1 training data without any process? For me i add a threshold(like 0.5) to change pseudo-labelling to discrete value, which lead better cv and lb, but i think it's not a good idea.\n2:Did you fix seed when training model? I am confused with this for a long time , since i saw @philippsinger  say he never fix seed and embrace random. If we do not fix seed , any idea to try will run for 5 to 10 times to do ensemble?",
    "897782": "1. I trained everything with 1s and 0s ( &lt; 0.5 = 0, &gt;= 0.5 = 1)\n2. I fixed seed for everything when I was training the models initially to test out new ideas, but for my final model I simply trained the model 3 times then averaged the predictions. However, even with fixing  it was not possible to replicate the result entirely because of TPU, so maybe that's why some people opted to just never fixing the seed.",
    "897812": "Thank you ! Do you remember how much improvement with  Pseudo-labelling ?",
    "898054": "Nice work! Which lr schedule did you use?",
    "898334": "alphaecho The idea is to ensemble NNs with different seeds. I have never seen a NN that did not get a boost from that.",
    "898596": "I used simple constant decay with warm up. I got the hyperparameter setting from the ROBERTA paper and found it to work well.",
    "899140": "Congrats with a great performance, @hansungj ! Great ideas and you made them work!\nCould you show a snippet of code of how you accessed layers' weight decays or url reference of an example? What optimizer did you use ?",
    "899143": "Thank you. Seems it is time consuming but also effective",
    "899170": "I simply modified AdamW implementation by huggingface\n```\ndef _decay_weights_op(self, var, learning_rate, apply_state):\n        do_decay = self._do_use_weight_decay(var.name)\n        #added\n        layer_coef = self._get_layer_wise_coefficient(var.name)\n        if do_decay:\n            return var.assign_sub(\n                learning_rate * var * layer_coef * apply_state[(var.device, var.dtype.base_dtype)][\"weight_decay_rate\"],\n                use_locking=self._use_locking,\n            )\n        return tf.no_op()\n```\n\nThen i added this function\n```\n\n    def _get_layer_wise_coefficient(self, param_name):\n        \"\"\"gets layer wise coefficient for the `param_name`.\"\"\"\n\n        if self._layer_wise_coefficient == 1:\n            return tf.math.pow(self._layer_wise_coefficient, 0)\n\n        if self._exclude_from_layer_wise:\n            for r in self._exclude_from_layer_wise:\n                if re.search(r, param_name) is not None:\n                    return tf.math.pow(self._layer_wise_coefficient, 0)\n        \n        for l in range(self._number_of_layers):\n            if re.search(self._layer_re_name + f'{l}/', param_name) is not None:\n                return tf.math.pow(self._layer_wise_coefficient, l)\n\n        return tf.math.pow(self._layer_wise_coefficient, self._number_of_layers )\n```\nXLM-R-large model by jplu, the layers could be distinguished by `'layer_._#/' `so I had `self._layer_re_name = 'layer_._`",
    "937856": "Thanks a lot. Very helpful code."
  },
  "source": "meta"
}