{
  "id": 161103,
  "title": "43rd place. Sampling with Adversarial Validation",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/bptt-dsmlkz-43rd-place-sampling-with-adversarial-v",
  "author_name": "",
  "post_date": "2020-06-24T01:39:37.937Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to Kaggle and all kagglers who were generous enough to share their expertise. The learning was rich thanks to:  Dezso Ribli's <a href=\"/riblidezso\">@riblidezso</a> -   <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\">Finetune XLM-Roberta on Jigsaw test data with MLM</a>, Alex Shonenkov <a href=\"/shonenkov\">@shonenkov</a> with Pytorch <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">[TPU-Training] Super Fast XLMRoberta</a>, DimitreOliveira <a href=\"/dimitreoliveira\">@dimitreoliveira</a> - <a href=\"https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops\">Jigsaw - TPU optimized training loops</a>, Abhishek <a href=\"/abhishek\">@abhishek</a> with his invaluable 'real-time coding' youtube videos and <a href=\"https://www.kaggle.com/abhishek/i-like-clean-tpu-training-kernels-i-can-not-lie\">I Like Clean TPU Training Kernels &amp; I Can Not Lie</a>, Xhlulu <a href=\"/xhlulu\">@xhlulu</a> - <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">Jigsaw TPU: XLM-Roberta</a>, Michael Kazachok’s <a href=\"/miklgr500\">@miklgr500</a> translations <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">dataset</a> and  <a href=\"https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-with-huggingface-and-keras\">Jigsaw TPU: BERT with Huggingface and Keras</a></p>\n\n<h3>Adversarial Validation.</h3>\n\n<p>The main difference of the solution is the way the training and validation sets were sampled. I used adversarial validation in this <a href=\"https://www.kaggle.com/isakev/jigsaw-adversarial-validation-folds-0-1\">kernel (previous version)</a> to sample <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">translations</a> by Michael Kazachok to pick samples ‘most similar to test set’ (280k-480k samples for training and 4k samples for validation (+8k original validation.csv). Adversarial Validation, the idea successfully used often across Kaggle (e.g. <a href=\"https://www.kaggle.com/tunguz/quora-adversarial-validation\">Quora Adversarial Validation\n</a> and <a href=\"https://www.kaggle.com/konradb/adversarial-validation\">Adversarial validation</a></p>\n\n<ul>\n<li>selecting samples this way, at least, did not make the performance worse than lucky picks of random sampling</li>\n</ul>\n\n<h3>The rest.</h3>\n\n<p>The best submission is the blend of predictions of 6 models, all Roberta-XLM-large MLM, 3 of them are based on the MLM finetuned to test set by  <a href=\"/riblidezso\">@riblidezso</a>  Dezso Ribli (link)[https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm].\nDiversity to models comes from different loss functions (apart from BCE, used focal loss and MSE (with soft labels)), 3 models with 4 top hidden layers’ outputs concatenated, 1 model sums those 4 layers. 3 models use differential learning rates for head and transformers as in <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">this kernel</a>, different lengths 192 and 256, for those samples that exceed max_len=192 - concatenating the last 25%*max_len of text to the beginning of text,,  Alex Shonenkov’s <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">kernel</a> with addition of opensubtitles data,  training on english text and without, using training data from 200k to 480k filtered by similarity to test set and/or similarity by language ditribution and/or similarity by ‘toxic’ target distribution.</p>\n\n<p>All were run in tensorflow and keras on TPU only. The GPU had been used for predictions only, or as in the case of Adversarial Validation when smaller Roberata-base model was employed.</p>",
  "messages": [
    {
      "id": "898771",
      "postDate": "06/23/2020 18:21:42",
      "content": "<p>Thanks to Kaggle and all kagglers who were generous enough to share their expertise. The learning was rich thanks to:  Dezso Ribli's <a href=\"/riblidezso\">@riblidezso</a> -   <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\">Finetune XLM-Roberta on Jigsaw test data with MLM</a>, Alex Shonenkov <a href=\"/shonenkov\">@shonenkov</a> with Pytorch <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">[TPU-Training] Super Fast XLMRoberta</a>, DimitreOliveira <a href=\"/dimitreoliveira\">@dimitreoliveira</a> - <a href=\"https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops\">Jigsaw - TPU optimized training loops</a>, Abhishek <a href=\"/abhishek\">@abhishek</a> with his invaluable 'real-time coding' youtube videos and <a href=\"https://www.kaggle.com/abhishek/i-like-clean-tpu-training-kernels-i-can-not-lie\">I Like Clean TPU Training Kernels &amp; I Can Not Lie</a>, Xhlulu <a href=\"/xhlulu\">@xhlulu</a> - <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">Jigsaw TPU: XLM-Roberta</a>, Michael Kazachok’s <a href=\"/miklgr500\">@miklgr500</a> translations <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">dataset</a> and  <a href=\"https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-with-huggingface-and-keras\">Jigsaw TPU: BERT with Huggingface and Keras</a></p>\n\n<h3>Adversarial Validation.</h3>\n\n<p>The main difference of the solution is the way the training and validation sets were sampled. I used adversarial validation in this <a href=\"https://www.kaggle.com/isakev/jigsaw-adversarial-validation-folds-0-1\">kernel (previous version)</a> to sample <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">translations</a> by Michael Kazachok to pick samples ‘most similar to test set’ (280k-480k samples for training and 4k samples for validation (+8k original validation.csv). Adversarial Validation, the idea successfully used often across Kaggle (e.g. <a href=\"https://www.kaggle.com/tunguz/quora-adversarial-validation\">Quora Adversarial Validation\n</a> and <a href=\"https://www.kaggle.com/konradb/adversarial-validation\">Adversarial validation</a></p>\n\n<ul>\n<li>selecting samples this way, at least, did not make the performance worse than lucky picks of random sampling</li>\n</ul>\n\n<h3>The rest.</h3>\n\n<p>The best submission is the blend of predictions of 6 models, all Roberta-XLM-large MLM, 3 of them are based on the MLM finetuned to test set by  <a href=\"/riblidezso\">@riblidezso</a>  Dezso Ribli (link)[https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm].\nDiversity to models comes from different loss functions (apart from BCE, used focal loss and MSE (with soft labels)), 3 models with 4 top hidden layers’ outputs concatenated, 1 model sums those 4 layers. 3 models use differential learning rates for head and transformers as in <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">this kernel</a>, different lengths 192 and 256, for those samples that exceed max_len=192 - concatenating the last 25%*max_len of text to the beginning of text,,  Alex Shonenkov’s <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">kernel</a> with addition of opensubtitles data,  training on english text and without, using training data from 200k to 480k filtered by similarity to test set and/or similarity by language ditribution and/or similarity by ‘toxic’ target distribution.</p>\n\n<p>All were run in tensorflow and keras on TPU only. The GPU had been used for predictions only, or as in the case of Adversarial Validation when smaller Roberata-base model was employed.</p>",
      "rawMarkdown": "Thanks to Kaggle and all kagglers who were generous enough to share their expertise. The learning was rich thanks to:  Dezso Ribli's @riblidezso -   [Finetune XLM-Roberta on Jigsaw test data with MLM](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm), Alex Shonenkov @shonenkov with Pytorch [[TPU-Training] Super Fast XLMRoberta](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta), DimitreOliveira @dimitreoliveira - [Jigsaw - TPU optimized training loops](https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops), Abhishek @abhishek with his invaluable 'real-time coding' youtube videos and [I Like Clean TPU Training Kernels &amp; I Can Not Lie](https://www.kaggle.com/abhishek/i-like-clean-tpu-training-kernels-i-can-not-lie), Xhlulu @xhlulu - [Jigsaw TPU: XLM-Roberta](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta), Michael Kazachok’s @miklgr500 translations [dataset](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api) and  [Jigsaw TPU: BERT with Huggingface and Keras](https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-with-huggingface-and-keras)\n\n### Adversarial Validation.\n\nThe main difference of the solution is the way the training and validation sets were sampled. I used adversarial validation in this [kernel (previous version)](https://www.kaggle.com/isakev/jigsaw-adversarial-validation-folds-0-1) to sample [translations](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api) by Michael Kazachok to pick samples ‘most similar to test set’ (280k-480k samples for training and 4k samples for validation (+8k original validation.csv). Adversarial Validation, the idea successfully used often across Kaggle (e.g. [Quora Adversarial Validation\n](https://www.kaggle.com/tunguz/quora-adversarial-validation) and [Adversarial validation](https://www.kaggle.com/konradb/adversarial-validation)\n\n- selecting samples this way, at least, did not make the performance worse than lucky picks of random sampling\n\n### The rest.\n\nThe best submission is the blend of predictions of 6 models, all Roberta-XLM-large MLM, 3 of them are based on the MLM finetuned to test set by  @riblidezso  Dezso Ribli (link)[https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm].\nDiversity to models comes from different loss functions (apart from BCE, used focal loss and MSE (with soft labels)), 3 models with 4 top hidden layers’ outputs concatenated, 1 model sums those 4 layers. 3 models use differential learning rates for head and transformers as in [this kernel](https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large), different lengths 192 and 256, for those samples that exceed max_len=192 - concatenating the last 25%*max_len of text to the beginning of text,,  Alex Shonenkov’s [kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta) with addition of opensubtitles data,  training on english text and without, using training data from 200k to 480k filtered by similarity to test set and/or similarity by language ditribution and/or similarity by ‘toxic’ target distribution.\n\nAll were run in tensorflow and keras on TPU only. The GPU had been used for predictions only, or as in the case of Adversarial Validation when smaller Roberata-base model was employed.",
      "votes": null
    },
    {
      "id": "898964",
      "postDate": "06/23/2020 21:33:10",
      "content": "<p>Hey <a href=\"/isakev\">@isakev</a> , thanks for the mention, and congratulations on your finish, great job!</p>",
      "rawMarkdown": "Hey @isakev , thanks for the mention, and congratulations on your finish, great job!",
      "votes": null
    },
    {
      "id": "899017",
      "postDate": "06/23/2020 22:38:19",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 898964,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "06/23/2020 21:33:10",
      "content": "<p>Hey <a href=\"/isakev\">@isakev</a> , thanks for the mention, and congratulations on your finish, great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899017,
      "author_name": "harisankarsivankutty",
      "author_url": "",
      "post_date": "06/23/2020 22:38:19",
      "content": "<p>Congrats</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "898771": "Thanks to Kaggle and all kagglers who were generous enough to share their expertise. The learning was rich thanks to:  Dezso Ribli's @riblidezso -   [Finetune XLM-Roberta on Jigsaw test data with MLM](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm), Alex Shonenkov @shonenkov with Pytorch [[TPU-Training] Super Fast XLMRoberta](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta), DimitreOliveira @dimitreoliveira - [Jigsaw - TPU optimized training loops](https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops), Abhishek @abhishek with his invaluable 'real-time coding' youtube videos and [I Like Clean TPU Training Kernels &amp; I Can Not Lie](https://www.kaggle.com/abhishek/i-like-clean-tpu-training-kernels-i-can-not-lie), Xhlulu @xhlulu - [Jigsaw TPU: XLM-Roberta](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta), Michael Kazachok’s @miklgr500 translations [dataset](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api) and  [Jigsaw TPU: BERT with Huggingface and Keras](https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-with-huggingface-and-keras)\n\n### Adversarial Validation.\n\nThe main difference of the solution is the way the training and validation sets were sampled. I used adversarial validation in this [kernel (previous version)](https://www.kaggle.com/isakev/jigsaw-adversarial-validation-folds-0-1) to sample [translations](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api) by Michael Kazachok to pick samples ‘most similar to test set’ (280k-480k samples for training and 4k samples for validation (+8k original validation.csv). Adversarial Validation, the idea successfully used often across Kaggle (e.g. [Quora Adversarial Validation\n](https://www.kaggle.com/tunguz/quora-adversarial-validation) and [Adversarial validation](https://www.kaggle.com/konradb/adversarial-validation)\n\n- selecting samples this way, at least, did not make the performance worse than lucky picks of random sampling\n\n### The rest.\n\nThe best submission is the blend of predictions of 6 models, all Roberta-XLM-large MLM, 3 of them are based on the MLM finetuned to test set by  @riblidezso  Dezso Ribli (link)[https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm].\nDiversity to models comes from different loss functions (apart from BCE, used focal loss and MSE (with soft labels)), 3 models with 4 top hidden layers’ outputs concatenated, 1 model sums those 4 layers. 3 models use differential learning rates for head and transformers as in [this kernel](https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large), different lengths 192 and 256, for those samples that exceed max_len=192 - concatenating the last 25%*max_len of text to the beginning of text,,  Alex Shonenkov’s [kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta) with addition of opensubtitles data,  training on english text and without, using training data from 200k to 480k filtered by similarity to test set and/or similarity by language ditribution and/or similarity by ‘toxic’ target distribution.\n\nAll were run in tensorflow and keras on TPU only. The GPU had been used for predictions only, or as in the case of Adversarial Validation when smaller Roberata-base model was employed.",
    "898964": "Hey @isakev , thanks for the mention, and congratulations on your finish, great job!",
    "899017": "Congrats"
  },
  "source": "meta"
}