{
  "id": 156622,
  "title": "Fine tuning with masked language modelling",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/156622",
  "author_name": "",
  "post_date": "2020-06-07T01:27:03.636870600Z",
  "votes": 44,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Fine tuning BERT models with masked language modelling (MLM) is a common way of improving these models for specific tasks. It is also quite often mentioned in previous kaggle NLP competition writeups. </p>\n\n<p>Surprisingly no public kernels used this idea, so I created one where I finetune XLM-R large on the competition test dataset.\n<a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\">https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm</a></p>\n\n<p>As a demonstration of the uselfullness of MLM finetuning I further train this model and reach 0.9422 in less than half an hour using the translated training dataset. \n<a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a></p>\n\n<p>Generally I found that MLM fine tuned models are more stable, train faster, and more often reach higher scores, than training from the original pretrained model.</p>\n\n<p>I hope you will find these notebooks useful.</p>",
  "messages": [
    {
      "id": "876731",
      "postDate": "06/07/2020 01:27:03",
      "content": "<p>Fine tuning BERT models with masked language modelling (MLM) is a common way of improving these models for specific tasks. It is also quite often mentioned in previous kaggle NLP competition writeups. </p>\n\n<p>Surprisingly no public kernels used this idea, so I created one where I finetune XLM-R large on the competition test dataset.\n<a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\">https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm</a></p>\n\n<p>As a demonstration of the uselfullness of MLM finetuning I further train this model and reach 0.9422 in less than half an hour using the translated training dataset. \n<a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a></p>\n\n<p>Generally I found that MLM fine tuned models are more stable, train faster, and more often reach higher scores, than training from the original pretrained model.</p>\n\n<p>I hope you will find these notebooks useful.</p>",
      "rawMarkdown": "Fine tuning BERT models with masked language modelling (MLM) is a common way of improving these models for specific tasks. It is also quite often mentioned in previous kaggle NLP competition writeups. \n\nSurprisingly no public kernels used this idea, so I created one where I finetune XLM-R large on the competition test dataset.\nhttps://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\n\nAs a demonstration of the uselfullness of MLM finetuning I further train this model and reach 0.9422 in less than half an hour using the translated training dataset. \nhttps://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\n\nGenerally I found that MLM fine tuned models are more stable, train faster, and more often reach higher scores, than training from the original pretrained model.\n\nI hope you will find these notebooks useful.",
      "votes": null
    },
    {
      "id": "876791",
      "postDate": "06/07/2020 03:40:26",
      "content": "<p>Thanks for your kind sharing. \nI just download output and will blend it directly later. \nMore kernels are also welcome</p>",
      "rawMarkdown": "Thanks for your kind sharing. \nI just download output and will blend it directly later. \nMore kernels are also welcome",
      "votes": null
    },
    {
      "id": "877237",
      "postDate": "06/07/2020 12:31:45",
      "content": "<p>Cool notebook(s)!</p>",
      "rawMarkdown": "Cool notebook(s)!",
      "votes": null
    },
    {
      "id": "877249",
      "postDate": "06/07/2020 12:45:40",
      "content": "<p>Have you used this in the final ensemble for your score?  BTW loving your series of cool notebooks thanks for those </p>",
      "rawMarkdown": "Have you used this in the final ensemble for your score?  BTW loving your series of cool notebooks thanks for those",
      "votes": null
    },
    {
      "id": "877407",
      "postDate": "06/07/2020 15:14:37",
      "content": "<p>First of all thanks for the notebook. I am using pytorch, so I wonder is it the hugging face's implementation that you mentioned in the notebook? <a href=\"https://colab.research.google.com/github/huggingface/blog/blob/master/notebooks/01_how_to_train.ipynb\">https://colab.research.google.com/github/huggingface/blog/blob/master/notebooks/01_how_to_train.ipynb</a></p>",
      "rawMarkdown": "First of all thanks for the notebook. I am using pytorch, so I wonder is it the hugging face's implementation that you mentioned in the notebook? https://colab.research.google.com/github/huggingface/blog/blob/master/notebooks/01_how_to_train.ipynb",
      "votes": null
    },
    {
      "id": "877421",
      "postDate": "06/07/2020 15:29:53",
      "content": "<p>Very useful things to learn.  Thanks <a href=\"/riblidezso\">@riblidezso</a>.</p>\n\n<p>Please I wonder:\n1) how did you find 2000 steps to train is good: did you first used validation split ? \n2) why maxlen=128 and not larger ?\n3) what do you think is  the MLM finetuning contribution to the PL and CV score in your classification notebook (excluding the improvements due to the differential learning rate and using translation data)</p>",
      "rawMarkdown": "Very useful things to learn.  Thanks @riblidezso.\n\nPlease I wonder:\n1) how did you find 2000 steps to train is good: did you first used validation split ? \n2) why maxlen=128 and not larger ?\n3) what do you think is  the MLM finetuning contribution to the PL and CV score in your classification notebook (excluding the improvements due to the differential learning rate and using translation data)",
      "votes": null
    },
    {
      "id": "877440",
      "postDate": "06/07/2020 15:52:17",
      "content": "<p>Thank I'm glad they are useful. I use one submission which is trained this way, but not exactly this one.</p>",
      "rawMarkdown": "Thank I'm glad they are useful. I use one submission which is trained this way, but not exactly this one.",
      "votes": null
    },
    {
      "id": "877443",
      "postDate": "06/07/2020 15:53:57",
      "content": "<p>Yes I meant that one, or more specifically the script is explains:\n<a href=\"https://github.com/huggingface/transformers/blob/master/examples/language-modeling/run_language_modeling.py\">https://github.com/huggingface/transformers/blob/master/examples/language-modeling/run_language_modeling.py</a></p>",
      "rawMarkdown": "Yes I meant that one, or more specifically the script is explains:\nhttps://github.com/huggingface/transformers/blob/master/examples/language-modeling/run_language_modeling.py",
      "votes": null
    },
    {
      "id": "877458",
      "postDate": "06/07/2020 16:04:25",
      "content": "<p>Thanks!</p>\n\n<p>1, If you train longer you will see that the AUC on the validation data does not really increase after 2000 steps in the first stage.\n2, I used 128 for the MLM, but 192 for the model. Most of the comments are shorter than 128 token, so it makes the MLM faster, and we do not lose too much real text. Thats all.\n3,  It trains quicker and training is more stable after MLM. Still there is quite large scatter in the scores. But on average I think it brings a 0.001-0.002 improvement for me. I reached 94.2+ scores without MLM too, just it took many many more tries to get lucky. Here, if the stage1 validation score reaches 94.7 then the LB score will almost always reach 94.15+.</p>",
      "rawMarkdown": "Thanks!\n\n1, If you train longer you will see that the AUC on the validation data does not really increase after 2000 steps in the first stage.\n2, I used 128 for the MLM, but 192 for the model. Most of the comments are shorter than 128 token, so it makes the MLM faster, and we do not lose too much real text. Thats all.\n3,  It trains quicker and training is more stable after MLM. Still there is quite large scatter in the scores. But on average I think it brings a 0.001-0.002 improvement for me. I reached 94.2+ scores without MLM too, just it took many many more tries to get lucky. Here, if the stage1 validation score reaches 94.7 then the LB score will almost always reach 94.15+.",
      "votes": null
    },
    {
      "id": "877537",
      "postDate": "06/07/2020 17:22:54",
      "content": "<p>Thank you <a href=\"/riblidezso\">@riblidezso</a> . That's a significant boost from pretuning MLM!</p>",
      "rawMarkdown": "Thank you @riblidezso . That's a significant boost from pretuning MLM!",
      "votes": null
    },
    {
      "id": "878253",
      "postDate": "06/08/2020 11:39:27",
      "content": "<p>Really well done !</p>",
      "rawMarkdown": "Really well done !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 876791,
      "author_name": "mcggood",
      "author_url": "",
      "post_date": "06/07/2020 03:40:26",
      "content": "<p>Thanks for your kind sharing. \nI just download output and will blend it directly later. \nMore kernels are also welcome</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 877237,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "06/07/2020 12:31:45",
      "content": "<p>Cool notebook(s)!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 877249,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "06/07/2020 12:45:40",
      "content": "<p>Have you used this in the final ensemble for your score?  BTW loving your series of cool notebooks thanks for those </p>",
      "votes": null,
      "replies": [
        {
          "id": 877440,
          "author_name": "riblidezso",
          "author_url": "",
          "post_date": "06/07/2020 15:52:17",
          "content": "<p>Thank I'm glad they are useful. I use one submission which is trained this way, but not exactly this one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 877407,
      "author_name": "ajinomoto132",
      "author_url": "",
      "post_date": "06/07/2020 15:14:37",
      "content": "<p>First of all thanks for the notebook. I am using pytorch, so I wonder is it the hugging face's implementation that you mentioned in the notebook? <a href=\"https://colab.research.google.com/github/huggingface/blog/blob/master/notebooks/01_how_to_train.ipynb\">https://colab.research.google.com/github/huggingface/blog/blob/master/notebooks/01_how_to_train.ipynb</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 877443,
          "author_name": "riblidezso",
          "author_url": "",
          "post_date": "06/07/2020 15:53:57",
          "content": "<p>Yes I meant that one, or more specifically the script is explains:\n<a href=\"https://github.com/huggingface/transformers/blob/master/examples/language-modeling/run_language_modeling.py\">https://github.com/huggingface/transformers/blob/master/examples/language-modeling/run_language_modeling.py</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 877421,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "06/07/2020 15:29:53",
      "content": "<p>Very useful things to learn.  Thanks <a href=\"/riblidezso\">@riblidezso</a>.</p>\n\n<p>Please I wonder:\n1) how did you find 2000 steps to train is good: did you first used validation split ? \n2) why maxlen=128 and not larger ?\n3) what do you think is  the MLM finetuning contribution to the PL and CV score in your classification notebook (excluding the improvements due to the differential learning rate and using translation data)</p>",
      "votes": null,
      "replies": [
        {
          "id": 877458,
          "author_name": "riblidezso",
          "author_url": "",
          "post_date": "06/07/2020 16:04:25",
          "content": "<p>Thanks!</p>\n\n<p>1, If you train longer you will see that the AUC on the validation data does not really increase after 2000 steps in the first stage.\n2, I used 128 for the MLM, but 192 for the model. Most of the comments are shorter than 128 token, so it makes the MLM faster, and we do not lose too much real text. Thats all.\n3,  It trains quicker and training is more stable after MLM. Still there is quite large scatter in the scores. But on average I think it brings a 0.001-0.002 improvement for me. I reached 94.2+ scores without MLM too, just it took many many more tries to get lucky. Here, if the stage1 validation score reaches 94.7 then the LB score will almost always reach 94.15+.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 877537,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "06/07/2020 17:22:54",
          "content": "<p>Thank you <a href=\"/riblidezso\">@riblidezso</a> . That's a significant boost from pretuning MLM!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 878253,
      "author_name": "haythemtellili5",
      "author_url": "",
      "post_date": "06/08/2020 11:39:27",
      "content": "<p>Really well done !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "876731": "Fine tuning BERT models with masked language modelling (MLM) is a common way of improving these models for specific tasks. It is also quite often mentioned in previous kaggle NLP competition writeups. \n\nSurprisingly no public kernels used this idea, so I created one where I finetune XLM-R large on the competition test dataset.\nhttps://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\n\nAs a demonstration of the uselfullness of MLM finetuning I further train this model and reach 0.9422 in less than half an hour using the translated training dataset. \nhttps://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\n\nGenerally I found that MLM fine tuned models are more stable, train faster, and more often reach higher scores, than training from the original pretrained model.\n\nI hope you will find these notebooks useful.",
    "876791": "Thanks for your kind sharing. \nI just download output and will blend it directly later. \nMore kernels are also welcome",
    "877237": "Cool notebook(s)!",
    "877249": "Have you used this in the final ensemble for your score?  BTW loving your series of cool notebooks thanks for those",
    "877407": "First of all thanks for the notebook. I am using pytorch, so I wonder is it the hugging face's implementation that you mentioned in the notebook? https://colab.research.google.com/github/huggingface/blog/blob/master/notebooks/01_how_to_train.ipynb",
    "877421": "Very useful things to learn.  Thanks @riblidezso.\n\nPlease I wonder:\n1) how did you find 2000 steps to train is good: did you first used validation split ? \n2) why maxlen=128 and not larger ?\n3) what do you think is  the MLM finetuning contribution to the PL and CV score in your classification notebook (excluding the improvements due to the differential learning rate and using translation data)",
    "877440": "Thank I'm glad they are useful. I use one submission which is trained this way, but not exactly this one.",
    "877443": "Yes I meant that one, or more specifically the script is explains:\nhttps://github.com/huggingface/transformers/blob/master/examples/language-modeling/run_language_modeling.py",
    "877458": "Thanks!\n\n1, If you train longer you will see that the AUC on the validation data does not really increase after 2000 steps in the first stage.\n2, I used 128 for the MLM, but 192 for the model. Most of the comments are shorter than 128 token, so it makes the MLM faster, and we do not lose too much real text. Thats all.\n3,  It trains quicker and training is more stable after MLM. Still there is quite large scatter in the scores. But on average I think it brings a 0.001-0.002 improvement for me. I reached 94.2+ scores without MLM too, just it took many many more tries to get lucky. Here, if the stage1 validation score reaches 94.7 then the LB score will almost always reach 94.15+.",
    "877537": "Thank you @riblidezso . That's a significant boost from pretuning MLM!",
    "878253": "Really well done !"
  },
  "source": "meta"
}