{
  "id": 176662,
  "title": "An improved MLM finetuning implementation in TensorFlow",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/176662",
  "author_name": "Yih-Dar SHIEH",
  "post_date": "2020-08-22T20:08:46.580000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>Some participants in this competition might have already seen <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s useful notebook <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\" target=\"_blank\">Finetune XLM-Roberta on Jigsaw test data with MLM</a>.</p>\n<p>Recently, I published a notebook <a href=\"https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu\" target=\"_blank\">Masked, My Dear Watson - MLM with TPU</a> in a getting started competition <a href=\"https://www.kaggle.com/c/contradictory-my-dear-watson\" target=\"_blank\">Contradictory, My Dear Watson</a>. In this notebook, I implement token masking in pure <a href=\"https://www.tensorflow.org/\" target=\"_blank\">TensorFlow</a> operations, which is used for training models like Bert and (XLM-)Roberta.</p>\n<p>If you want to skip all the code irrelevant to this <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/overview\" target=\"_blank\">Jigsaw Multilingual Toxic Comment Classification</a> competition, the main implementation is <a href=\"https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu#mask-tokens\" target=\"_blank\">here</a>.</p>\n<p>The implemention is a translation into <a href=\"https://www.tensorflow.org/\" target=\"_blank\">TensorFlow</a> from <a href=\"https://github.com/huggingface/transformers/blob/390c1285925dd119705e69a266202ef04490d012/examples/distillation/distiller.py\" target=\"_blank\">Hugging Face</a>'s code, which is originally in <a href=\"https://pytorch.org/\" target=\"_blank\">PyTorch</a>.</p>\n<p>Here are some features in this implementation:</p>\n<ul>\n<li>dynamic token masking in pure TensorFlow operations, therefore it could be used as a tf.data.Dataset transformation.</li>\n<li>the number of tokens to mask is calculated based on non-padding tokens (and optionally, excluding other special tokens)</li>\n<li>including a smoothing option to mask rare tokens more frequently</li>\n</ul>\n<p>The differences to <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s notebook <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\" target=\"_blank\">Finetune XLM-Roberta on Jigsaw test data with MLM</a> are:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s notebook implements token masking in numpy operations, and the tokens are statically masked before MLM finetuning. In this notebook, the maksing is implemented in pure TensorFlow operations, and the tokens are masked dynamically during MLM finetuning as tf.data.Dataset transformation.</li>\n<li>In <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s notebook, the number of tokens to be maksed is calculated before ignoring special tokens (in particular, the [PAD] token). However, in the case where we have a lot of padding tokens due to short texts, we will get a higher ratio of token being masked. This notebook carefully determines the number of tokens to be maksed based on the number the actual (non-special) tokens.</li>\n</ul>\n<p>A visualization of token masking:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F81d6249dfbaba0fcb0da05c7109f855c%2Fmlm-preview.png?generation=1598043639234956&amp;alt=media\" alt=\"\"></p>\n<p>I use token masking to further fine tune the loaded model on the competition datasets and external datasets (MNLI + XNLI) with MLM loss as objective, before training the NLI classifier on the labeled dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F7b075b70bebfdfa487ce840322df01b8%2Fpretrain%200-vs-pretrain%203.png?generation=1598043493428883&amp;alt=media\" alt=\"\"></p>\n<p>Here are the observations:</p>\n<ul>\n<li>After a few epochs (usually 3-6 epochs), the classifier using a (further) pretrained LM model outperforms the one without using further pretrained LM model by 1 ~ 2 %.</li>\n<li>At the first few epochs, further pretraining gives worse results. A possible explanation is that the further pretraining makes a LM model lose some of its overall language knowledge, therefore in an early stage of training on labeled datasets, it can't extract good features for predicting labels. However, once the training epochs increases, the knowledge gained from further pretraining can help it to better predict a downstream task labels.</li>\n<li>However, in some different notebook runnings, I saw the same performance for training with/without using further pretraining. Probably the training/validation split and the sampling from external datasets (for pretraining) also play some role here.</li>\n</ul>\n<p>The training is written in a customized way. I would probably try to convert it to <code>model.fit</code> style later, but not sure for now.</p>\n<p>Despite the notebook is written for another competition, I still hope you find my notebook helpful and enjoy it.</p>",
  "messages": [
    {
      "id": 981898,
      "postDate": "2020-08-22T20:08:46.580Z",
      "content": "<p>Hi,</p>\n<p>Some participants in this competition might have already seen <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s useful notebook <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\" target=\"_blank\">Finetune XLM-Roberta on Jigsaw test data with MLM</a>.</p>\n<p>Recently, I published a notebook <a href=\"https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu\" target=\"_blank\">Masked, My Dear Watson - MLM with TPU</a> in a getting started competition <a href=\"https://www.kaggle.com/c/contradictory-my-dear-watson\" target=\"_blank\">Contradictory, My Dear Watson</a>. In this notebook, I implement token masking in pure <a href=\"https://www.tensorflow.org/\" target=\"_blank\">TensorFlow</a> operations, which is used for training models like Bert and (XLM-)Roberta.</p>\n<p>If you want to skip all the code irrelevant to this <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/overview\" target=\"_blank\">Jigsaw Multilingual Toxic Comment Classification</a> competition, the main implementation is <a href=\"https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu#mask-tokens\" target=\"_blank\">here</a>.</p>\n<p>The implemention is a translation into <a href=\"https://www.tensorflow.org/\" target=\"_blank\">TensorFlow</a> from <a href=\"https://github.com/huggingface/transformers/blob/390c1285925dd119705e69a266202ef04490d012/examples/distillation/distiller.py\" target=\"_blank\">Hugging Face</a>'s code, which is originally in <a href=\"https://pytorch.org/\" target=\"_blank\">PyTorch</a>.</p>\n<p>Here are some features in this implementation:</p>\n<ul>\n<li>dynamic token masking in pure TensorFlow operations, therefore it could be used as a tf.data.Dataset transformation.</li>\n<li>the number of tokens to mask is calculated based on non-padding tokens (and optionally, excluding other special tokens)</li>\n<li>including a smoothing option to mask rare tokens more frequently</li>\n</ul>\n<p>The differences to <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s notebook <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\" target=\"_blank\">Finetune XLM-Roberta on Jigsaw test data with MLM</a> are:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s notebook implements token masking in numpy operations, and the tokens are statically masked before MLM finetuning. In this notebook, the maksing is implemented in pure TensorFlow operations, and the tokens are masked dynamically during MLM finetuning as tf.data.Dataset transformation.</li>\n<li>In <a href=\"https://www.kaggle.com/riblidezso\" target=\"_blank\">Dezső Ribli</a>'s notebook, the number of tokens to be maksed is calculated before ignoring special tokens (in particular, the [PAD] token). However, in the case where we have a lot of padding tokens due to short texts, we will get a higher ratio of token being masked. This notebook carefully determines the number of tokens to be maksed based on the number the actual (non-special) tokens.</li>\n</ul>\n<p>A visualization of token masking:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F81d6249dfbaba0fcb0da05c7109f855c%2Fmlm-preview.png?generation=1598043639234956&amp;alt=media\" alt=\"\"></p>\n<p>I use token masking to further fine tune the loaded model on the competition datasets and external datasets (MNLI + XNLI) with MLM loss as objective, before training the NLI classifier on the labeled dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F7b075b70bebfdfa487ce840322df01b8%2Fpretrain%200-vs-pretrain%203.png?generation=1598043493428883&amp;alt=media\" alt=\"\"></p>\n<p>Here are the observations:</p>\n<ul>\n<li>After a few epochs (usually 3-6 epochs), the classifier using a (further) pretrained LM model outperforms the one without using further pretrained LM model by 1 ~ 2 %.</li>\n<li>At the first few epochs, further pretraining gives worse results. A possible explanation is that the further pretraining makes a LM model lose some of its overall language knowledge, therefore in an early stage of training on labeled datasets, it can't extract good features for predicting labels. However, once the training epochs increases, the knowledge gained from further pretraining can help it to better predict a downstream task labels.</li>\n<li>However, in some different notebook runnings, I saw the same performance for training with/without using further pretraining. Probably the training/validation split and the sampling from external datasets (for pretraining) also play some role here.</li>\n</ul>\n<p>The training is written in a customized way. I would probably try to convert it to <code>model.fit</code> style later, but not sure for now.</p>\n<p>Despite the notebook is written for another competition, I still hope you find my notebook helpful and enjoy it.</p>",
      "rawMarkdown": "Hi,\n\nSome participants in this competition might have already seen [Dezső Ribli](https://www.kaggle.com/riblidezso)'s useful notebook [Finetune XLM-Roberta on Jigsaw test data with MLM](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm).\n\nRecently, I published a notebook [Masked, My Dear Watson - MLM with TPU](https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu) in a getting started competition [Contradictory, My Dear Watson](https://www.kaggle.com/c/contradictory-my-dear-watson). In this notebook, I implement token masking in pure [TensorFlow](https://www.tensorflow.org/) operations, which is used for training models like Bert and (XLM-)Roberta.\n\nIf you want to skip all the code irrelevant to this [Jigsaw Multilingual Toxic Comment Classification](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/overview) competition, the main implementation is [here](https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu#mask-tokens).\n\nThe implemention is a translation into [TensorFlow](https://www.tensorflow.org/) from [Hugging Face](https://github.com/huggingface/transformers/blob/390c1285925dd119705e69a266202ef04490d012/examples/distillation/distiller.py)'s code, which is originally in [PyTorch](https://pytorch.org/).\n\nHere are some features in this implementation:\n  - dynamic token masking in pure TensorFlow operations, therefore it could be used as a tf.data.Dataset transformation.\n  - the number of tokens to mask is calculated based on non-padding tokens (and optionally, excluding other special tokens)\n  - including a smoothing option to mask rare tokens more frequently\n\nThe differences to [Dezső Ribli](https://www.kaggle.com/riblidezso)'s notebook [Finetune XLM-Roberta on Jigsaw test data with MLM](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm) are:\n\n  - [Dezső Ribli](https://www.kaggle.com/riblidezso)'s notebook implements token masking in numpy operations, and the tokens are statically masked before MLM finetuning. In this notebook, the maksing is implemented in pure TensorFlow operations, and the tokens are masked dynamically during MLM finetuning as tf.data.Dataset transformation.\n  - In [Dezső Ribli](https://www.kaggle.com/riblidezso)'s notebook, the number of tokens to be maksed is calculated before ignoring special tokens (in particular, the [PAD] token). However, in the case where we have a lot of padding tokens due to short texts, we will get a higher ratio of token being masked. This notebook carefully determines the number of tokens to be maksed based on the number the actual (non-special) tokens.\n\nA visualization of token masking:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F81d6249dfbaba0fcb0da05c7109f855c%2Fmlm-preview.png?generation=1598043639234956&alt=media)\n\nI use token masking to further fine tune the loaded model on the competition datasets and external datasets (MNLI + XNLI) with MLM loss as objective, before training the NLI classifier on the labeled dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F7b075b70bebfdfa487ce840322df01b8%2Fpretrain%200-vs-pretrain%203.png?generation=1598043493428883&alt=media)\n\nHere are the observations:\n\n  - After a few epochs (usually 3-6 epochs), the classifier using a (further) pretrained LM model outperforms the one without using further pretrained LM model by 1 ~ 2 %.\n  - At the first few epochs, further pretraining gives worse results. A possible explanation is that the further pretraining makes a LM model lose some of its overall language knowledge, therefore in an early stage of training on labeled datasets, it can't extract good features for predicting labels. However, once the training epochs increases, the knowledge gained from further pretraining can help it to better predict a downstream task labels.\n  - However, in some different notebook runnings, I saw the same performance for training with/without using further pretraining. Probably the training/validation split and the sampling from external datasets (for pretraining) also play some role here.\n\nThe training is written in a customized way. I would probably try to convert it to `model.fit` style later, but not sure for now.\n\nDespite the notebook is written for another competition, I still hope you find my notebook helpful and enjoy it.",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "981898": "Hi,\n\nSome participants in this competition might have already seen [Dezső Ribli](https://www.kaggle.com/riblidezso)'s useful notebook [Finetune XLM-Roberta on Jigsaw test data with MLM](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm).\n\nRecently, I published a notebook [Masked, My Dear Watson - MLM with TPU](https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu) in a getting started competition [Contradictory, My Dear Watson](https://www.kaggle.com/c/contradictory-my-dear-watson). In this notebook, I implement token masking in pure [TensorFlow](https://www.tensorflow.org/) operations, which is used for training models like Bert and (XLM-)Roberta.\n\nIf you want to skip all the code irrelevant to this [Jigsaw Multilingual Toxic Comment Classification](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/overview) competition, the main implementation is [here](https://www.kaggle.com/yihdarshieh/masked-my-dear-watson-mlm-with-tpu#mask-tokens).\n\nThe implemention is a translation into [TensorFlow](https://www.tensorflow.org/) from [Hugging Face](https://github.com/huggingface/transformers/blob/390c1285925dd119705e69a266202ef04490d012/examples/distillation/distiller.py)'s code, which is originally in [PyTorch](https://pytorch.org/).\n\nHere are some features in this implementation:\n  - dynamic token masking in pure TensorFlow operations, therefore it could be used as a tf.data.Dataset transformation.\n  - the number of tokens to mask is calculated based on non-padding tokens (and optionally, excluding other special tokens)\n  - including a smoothing option to mask rare tokens more frequently\n\nThe differences to [Dezső Ribli](https://www.kaggle.com/riblidezso)'s notebook [Finetune XLM-Roberta on Jigsaw test data with MLM](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm) are:\n\n  - [Dezső Ribli](https://www.kaggle.com/riblidezso)'s notebook implements token masking in numpy operations, and the tokens are statically masked before MLM finetuning. In this notebook, the maksing is implemented in pure TensorFlow operations, and the tokens are masked dynamically during MLM finetuning as tf.data.Dataset transformation.\n  - In [Dezső Ribli](https://www.kaggle.com/riblidezso)'s notebook, the number of tokens to be maksed is calculated before ignoring special tokens (in particular, the [PAD] token). However, in the case where we have a lot of padding tokens due to short texts, we will get a higher ratio of token being masked. This notebook carefully determines the number of tokens to be maksed based on the number the actual (non-special) tokens.\n\nA visualization of token masking:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F81d6249dfbaba0fcb0da05c7109f855c%2Fmlm-preview.png?generation=1598043639234956&alt=media)\n\nI use token masking to further fine tune the loaded model on the competition datasets and external datasets (MNLI + XNLI) with MLM loss as objective, before training the NLI classifier on the labeled dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1533864%2F7b075b70bebfdfa487ce840322df01b8%2Fpretrain%200-vs-pretrain%203.png?generation=1598043493428883&alt=media)\n\nHere are the observations:\n\n  - After a few epochs (usually 3-6 epochs), the classifier using a (further) pretrained LM model outperforms the one without using further pretrained LM model by 1 ~ 2 %.\n  - At the first few epochs, further pretraining gives worse results. A possible explanation is that the further pretraining makes a LM model lose some of its overall language knowledge, therefore in an early stage of training on labeled datasets, it can't extract good features for predicting labels. However, once the training epochs increases, the knowledge gained from further pretraining can help it to better predict a downstream task labels.\n  - However, in some different notebook runnings, I saw the same performance for training with/without using further pretraining. Probably the training/validation split and the sampling from external datasets (for pretraining) also play some role here.\n\nThe training is written in a customized way. I would probably try to convert it to `model.fit` style later, but not sure for now.\n\nDespite the notebook is written for another competition, I still hope you find my notebook helpful and enjoy it."
  }
}