{
  "id": 160862,
  "title": "1st place solution overview",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160862",
  "author_name": "Chun Ming Lee",
  "post_date": "2020-06-23T00:48:59.557000",
  "votes": 183,
  "comment_count": 39,
  "views": 0,
  "content": "<p>We’d like to start off by thanking Kaggle/Jigsaw for a drama-free competition and also by congratulating the other medallists. </p>\n\n<h2>TL;DR</h2>\n\n<ol>\n<li>Ensemble, ensemble, ensemble</li>\n<li>Pseudo-labelling</li>\n<li>Bootstrap with multilingual models, refine with monolingual models</li>\n</ol>\n\n<p>I haven't competed on Kaggle since I co-won the 2018 Toxic Comments competition. The state-of-the-art for NLP classification at the time was non-contextual word embeddings (e.g., FastText) and I was curious if I’d make a good showing in a world of Transformers. I'm pleasantly surprised to have done so well with my team-mate @rafiko1. </p>\n\n<h2>Our public LB milestones</h2>\n\n<ul>\n<li>Baseline XLM-Roberta model (Public LB: 0.93XX - 31 March)</li>\n<li>Average ensemble of XLM-R models (0.942X - 3 April)</li>\n<li>Blending with monolingual Transformer models (0.9510 - 23 April)</li>\n<li>Team merger - weighted average of our individual best subs (0.9529 - 28 April)</li>\n<li>Blending with monolingual FastText classifiers (0.9544 - 5 May)</li>\n<li>Post-processing and misc. optimizations (0.9556 - 13 June)</li>\n</ul>\n\n<p>You’ll note we hit 1st place on the <strong>FINAL</strong> public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)</p>\n\n<h2>CV strategy</h2>\n\n<p>We initially used a mix of k-fold CV and validation set as hold-out but as we refined our test predictions and used pseudo-labels + validation set for training, the validation metric became noisy to the point where we relied primarily on the public LB score. </p>\n\n<h2>Insights</h2>\n\n<h3>Ensembling to mitigate Transformer training variability</h3>\n\n<p>It’s been noted that the performance of Transformer models is impacted heavily by initialization and data order (<a href=\"https://arxiv.org/pdf/2002.06305.pdf\">https://arxiv.org/pdf/2002.06305.pdf</a>,  <a href=\"https://www.aclweb.org/anthology/2020.trac-1.9.pdf\">https://www.aclweb.org/anthology/2020.trac-1.9.pdf</a>). To mitigate that, we emphasized the ensembling and bagging our models. This included temporal self-ensembling. Given the public test set, we went with an iterative blending approach, refining the test set predictions across submissions with a weighted average of the previous best submission and the current model’s predictions. We began with a simple average, and gradually increased the weight of the previous best submission. For the training data, we largely used sub-samples of the translations of the 2018 toxic comments for each model run. </p>\n\n<h3>Pseudo-labels (PL)</h3>\n\n<p>We observed a performance improvement when we used test-set predictions as training data - the intuition being that it helps models learn the test set distribution. Using all test-set predictions as soft-labels worked better than any other version of pseudo-labelling (e.g., hard labels, confidence thresholded PLs etc.). Towards the end of the competition, we discovered a minor but material boost in LB when we upsampled the PLs. </p>\n\n<h3>Multilingual XLM-Roberta models</h3>\n\n<p>As with most teams, we began with a vanilla XLM-R model, incorporating translations of the 2018 dataset in the 6 test-set languages as training data. We used a vanilla classification head on the CLS token of the last layer with the Adam optimizer and binary cross entropy loss function, and finetuned the entire model with a low learning rate. Given Transformer models have several hundred million trainable weights put to the relatively simple task of making a binary prediction, we didn’t believe the default architecture required much tweaking to get a good signal through the model.  Consequently, we didn't spend too much time on hyper-parameter optimization, architectural tweaks, or preprocessing. </p>\n\n<h3>Foreign language monolingual Transformer models</h3>\n\n<p>@rafiko1 and I stumbled on this technique independently and prior to forming a team. We were both inspired by the MultiFiT paper (<a href=\"https://arxiv.org/pdf/1909.04761.pdf\">https://arxiv.org/pdf/1909.04761.pdf</a>) - specifically:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1029053%2Fb40975c9453fba6f640836e4fe297746%2FCapture.PNG?generation=1592872625708960&amp;alt=media\" alt=\"\"></p>\n\n<p>We observed a dramatic performance boost when we used pretrained foreign language monolingual Transformer models from HuggingFace for the test-set languages(e.g., Camembert for french samples, Rubert for russian, BerTurk for turkish, BETO for spanish etc.). </p>\n\n<p>We finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language (initially from the XLM-R multilingual models), training the corresponding monolingual model, predicting the same samples then blending it back with the “main branch” of all predictions. It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements. </p>\n\n<p>For each model run, we’d reload weight initalizations from the pretrained models to prevent overfitting. In other words, the continuing improvements we saw were being driven by refinements in the pseudo-labels we were providing to the models as training data. </p>\n\n<p>For a given monolingual model, predicting only test-set samples in that language worked best. Translating test-set samples in other languages to the model's language and predicting them worsened performance. </p>\n\n<h3>Finetuning pre-trained foreign language monolingual FastText models</h3>\n\n<p>After we exhausted the HuggingFace monolingual model library, we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, to continue refining the test set predictions (albeit with a lower weight when combined with the main branch of predictions) and saw a small but meaningful performance boost (0.9536 to 0.9544). </p>\n\n<h3>Post-processing</h3>\n\n<p>Given the large number of submissions we were making, @rafiko1 came up with the novel idea to make use of the history of submissions to tweak our test set predictions. We tracked the delta of predictions for each sample for successful submissions, averaged them and nudged the predictions in the same direction. We saw a minor but material boost in performance (~0.0005)</p>\n\n<p>We were concerned about the risk of overfitting with this post-processing technique so for our final 2 submissions, selected one that incorporated this post-processing and one that didn't. It ended up working very well on the private LB. </p>\n\n<h2>Misc -</h2>\n\n<h3>Training setup</h3>\n\n<p>@rafiko1 and I kept separate code-bases. He trained on Kaggle TPU instances with Tensorflow code derived from the public kernels by @xhlulu and @shonenkov while I trained on my own hardware (dual RTX Titans) using from-scratch PyTorch code. I’d attribute part of our outsized merger ensemble boost (0.9510 -&gt; 0.9529) to this diversity in training. </p>\n\n<h3>What didn’t work</h3>\n\n<p>We went through a checklist of contemporary NLP classification techniques and most of them didn’t work probably due to the simple problem objective. Here is a selection of what didn’t work :- \n        - Document-level embedders (e.g., LASER)\n        - Further MLM pretraining of Transformer models using task data\n        - Alternative ensembling mechanisms (rank-averaging, stochastic weight averaging)\n        - Alternative loss functions (e.g., focal loss, histogram loss)\n        - Alternative pooling mechanisms for the classification head (e.g., max-pool CNN across tokens, using multiple hidden layers etc.)\n        - Non-FastText pretrained embeddings (e.g., Flair, glove, bpem)\n        - Freeze-finetuning for the Transformer models\n        - Regularization (e.g., multisample dropout, input mixup, manifold mixup, sentencepiece-dropout)\n        - Backtranslation as data augmentation\n        - English translations as train/test-time augmentation\n        - Self-distilling to relabel the 2018 data \n        - Adversarial training by perturbing the embeddings layer using FGM\n        - Multi-task learning\n        - Temperature scaling on pseudo-labels\n        - Semi-supervised learning using the test data <br>\n        - Composing two models into an encoder-decoder model \n        - Use of translation focused pretrained models (e.g., mBart)</p>\n\n<p>We'll release our code later this week.</p>",
  "messages": [
    {
      "id": 897564,
      "postDate": "2020-06-23T00:48:59.557Z",
      "content": "<p>We’d like to start off by thanking Kaggle/Jigsaw for a drama-free competition and also by congratulating the other medallists. </p>\n\n<h2>TL;DR</h2>\n\n<ol>\n<li>Ensemble, ensemble, ensemble</li>\n<li>Pseudo-labelling</li>\n<li>Bootstrap with multilingual models, refine with monolingual models</li>\n</ol>\n\n<p>I haven't competed on Kaggle since I co-won the 2018 Toxic Comments competition. The state-of-the-art for NLP classification at the time was non-contextual word embeddings (e.g., FastText) and I was curious if I’d make a good showing in a world of Transformers. I'm pleasantly surprised to have done so well with my team-mate @rafiko1. </p>\n\n<h2>Our public LB milestones</h2>\n\n<ul>\n<li>Baseline XLM-Roberta model (Public LB: 0.93XX - 31 March)</li>\n<li>Average ensemble of XLM-R models (0.942X - 3 April)</li>\n<li>Blending with monolingual Transformer models (0.9510 - 23 April)</li>\n<li>Team merger - weighted average of our individual best subs (0.9529 - 28 April)</li>\n<li>Blending with monolingual FastText classifiers (0.9544 - 5 May)</li>\n<li>Post-processing and misc. optimizations (0.9556 - 13 June)</li>\n</ul>\n\n<p>You’ll note we hit 1st place on the <strong>FINAL</strong> public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)</p>\n\n<h2>CV strategy</h2>\n\n<p>We initially used a mix of k-fold CV and validation set as hold-out but as we refined our test predictions and used pseudo-labels + validation set for training, the validation metric became noisy to the point where we relied primarily on the public LB score. </p>\n\n<h2>Insights</h2>\n\n<h3>Ensembling to mitigate Transformer training variability</h3>\n\n<p>It’s been noted that the performance of Transformer models is impacted heavily by initialization and data order (<a href=\"https://arxiv.org/pdf/2002.06305.pdf\">https://arxiv.org/pdf/2002.06305.pdf</a>,  <a href=\"https://www.aclweb.org/anthology/2020.trac-1.9.pdf\">https://www.aclweb.org/anthology/2020.trac-1.9.pdf</a>). To mitigate that, we emphasized the ensembling and bagging our models. This included temporal self-ensembling. Given the public test set, we went with an iterative blending approach, refining the test set predictions across submissions with a weighted average of the previous best submission and the current model’s predictions. We began with a simple average, and gradually increased the weight of the previous best submission. For the training data, we largely used sub-samples of the translations of the 2018 toxic comments for each model run. </p>\n\n<h3>Pseudo-labels (PL)</h3>\n\n<p>We observed a performance improvement when we used test-set predictions as training data - the intuition being that it helps models learn the test set distribution. Using all test-set predictions as soft-labels worked better than any other version of pseudo-labelling (e.g., hard labels, confidence thresholded PLs etc.). Towards the end of the competition, we discovered a minor but material boost in LB when we upsampled the PLs. </p>\n\n<h3>Multilingual XLM-Roberta models</h3>\n\n<p>As with most teams, we began with a vanilla XLM-R model, incorporating translations of the 2018 dataset in the 6 test-set languages as training data. We used a vanilla classification head on the CLS token of the last layer with the Adam optimizer and binary cross entropy loss function, and finetuned the entire model with a low learning rate. Given Transformer models have several hundred million trainable weights put to the relatively simple task of making a binary prediction, we didn’t believe the default architecture required much tweaking to get a good signal through the model.  Consequently, we didn't spend too much time on hyper-parameter optimization, architectural tweaks, or preprocessing. </p>\n\n<h3>Foreign language monolingual Transformer models</h3>\n\n<p>@rafiko1 and I stumbled on this technique independently and prior to forming a team. We were both inspired by the MultiFiT paper (<a href=\"https://arxiv.org/pdf/1909.04761.pdf\">https://arxiv.org/pdf/1909.04761.pdf</a>) - specifically:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1029053%2Fb40975c9453fba6f640836e4fe297746%2FCapture.PNG?generation=1592872625708960&amp;alt=media\" alt=\"\"></p>\n\n<p>We observed a dramatic performance boost when we used pretrained foreign language monolingual Transformer models from HuggingFace for the test-set languages(e.g., Camembert for french samples, Rubert for russian, BerTurk for turkish, BETO for spanish etc.). </p>\n\n<p>We finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language (initially from the XLM-R multilingual models), training the corresponding monolingual model, predicting the same samples then blending it back with the “main branch” of all predictions. It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements. </p>\n\n<p>For each model run, we’d reload weight initalizations from the pretrained models to prevent overfitting. In other words, the continuing improvements we saw were being driven by refinements in the pseudo-labels we were providing to the models as training data. </p>\n\n<p>For a given monolingual model, predicting only test-set samples in that language worked best. Translating test-set samples in other languages to the model's language and predicting them worsened performance. </p>\n\n<h3>Finetuning pre-trained foreign language monolingual FastText models</h3>\n\n<p>After we exhausted the HuggingFace monolingual model library, we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, to continue refining the test set predictions (albeit with a lower weight when combined with the main branch of predictions) and saw a small but meaningful performance boost (0.9536 to 0.9544). </p>\n\n<h3>Post-processing</h3>\n\n<p>Given the large number of submissions we were making, @rafiko1 came up with the novel idea to make use of the history of submissions to tweak our test set predictions. We tracked the delta of predictions for each sample for successful submissions, averaged them and nudged the predictions in the same direction. We saw a minor but material boost in performance (~0.0005)</p>\n\n<p>We were concerned about the risk of overfitting with this post-processing technique so for our final 2 submissions, selected one that incorporated this post-processing and one that didn't. It ended up working very well on the private LB. </p>\n\n<h2>Misc -</h2>\n\n<h3>Training setup</h3>\n\n<p>@rafiko1 and I kept separate code-bases. He trained on Kaggle TPU instances with Tensorflow code derived from the public kernels by @xhlulu and @shonenkov while I trained on my own hardware (dual RTX Titans) using from-scratch PyTorch code. I’d attribute part of our outsized merger ensemble boost (0.9510 -&gt; 0.9529) to this diversity in training. </p>\n\n<h3>What didn’t work</h3>\n\n<p>We went through a checklist of contemporary NLP classification techniques and most of them didn’t work probably due to the simple problem objective. Here is a selection of what didn’t work :- \n        - Document-level embedders (e.g., LASER)\n        - Further MLM pretraining of Transformer models using task data\n        - Alternative ensembling mechanisms (rank-averaging, stochastic weight averaging)\n        - Alternative loss functions (e.g., focal loss, histogram loss)\n        - Alternative pooling mechanisms for the classification head (e.g., max-pool CNN across tokens, using multiple hidden layers etc.)\n        - Non-FastText pretrained embeddings (e.g., Flair, glove, bpem)\n        - Freeze-finetuning for the Transformer models\n        - Regularization (e.g., multisample dropout, input mixup, manifold mixup, sentencepiece-dropout)\n        - Backtranslation as data augmentation\n        - English translations as train/test-time augmentation\n        - Self-distilling to relabel the 2018 data \n        - Adversarial training by perturbing the embeddings layer using FGM\n        - Multi-task learning\n        - Temperature scaling on pseudo-labels\n        - Semi-supervised learning using the test data <br>\n        - Composing two models into an encoder-decoder model \n        - Use of translation focused pretrained models (e.g., mBart)</p>\n\n<p>We'll release our code later this week.</p>",
      "rawMarkdown": "We’d like to start off by thanking Kaggle/Jigsaw for a drama-free competition and also by congratulating the other medallists. \n\n## TL;DR\n1. Ensemble, ensemble, ensemble\n2. Pseudo-labelling\n3. Bootstrap with multilingual models, refine with monolingual models\n\n\nI haven't competed on Kaggle since I co-won the 2018 Toxic Comments competition. The state-of-the-art for NLP classification at the time was non-contextual word embeddings (e.g., FastText) and I was curious if I’d make a good showing in a world of Transformers. I'm pleasantly surprised to have done so well with my team-mate @rafiko1. \n\n## Our public LB milestones\n- Baseline XLM-Roberta model (Public LB: 0.93XX - 31 March)\n- Average ensemble of XLM-R models (0.942X - 3 April)\n- Blending with monolingual Transformer models (0.9510 - 23 April)\n- Team merger - weighted average of our individual best subs (0.9529 - 28 April)\n- Blending with monolingual FastText classifiers (0.9544 - 5 May)\n- Post-processing and misc. optimizations (0.9556 - 13 June)\n\nYou’ll note we hit 1st place on the **FINAL** public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)\n\n## CV strategy\nWe initially used a mix of k-fold CV and validation set as hold-out but as we refined our test predictions and used pseudo-labels + validation set for training, the validation metric became noisy to the point where we relied primarily on the public LB score. \n\n## Insights\n### Ensembling to mitigate Transformer training variability\nIt’s been noted that the performance of Transformer models is impacted heavily by initialization and data order (https://arxiv.org/pdf/2002.06305.pdf,  https://www.aclweb.org/anthology/2020.trac-1.9.pdf). To mitigate that, we emphasized the ensembling and bagging our models. This included temporal self-ensembling. Given the public test set, we went with an iterative blending approach, refining the test set predictions across submissions with a weighted average of the previous best submission and the current model’s predictions. We began with a simple average, and gradually increased the weight of the previous best submission. For the training data, we largely used sub-samples of the translations of the 2018 toxic comments for each model run. \n\n### Pseudo-labels (PL)\n We observed a performance improvement when we used test-set predictions as training data - the intuition being that it helps models learn the test set distribution. Using all test-set predictions as soft-labels worked better than any other version of pseudo-labelling (e.g., hard labels, confidence thresholded PLs etc.). Towards the end of the competition, we discovered a minor but material boost in LB when we upsampled the PLs. \n\n### Multilingual XLM-Roberta models\nAs with most teams, we began with a vanilla XLM-R model, incorporating translations of the 2018 dataset in the 6 test-set languages as training data. We used a vanilla classification head on the CLS token of the last layer with the Adam optimizer and binary cross entropy loss function, and finetuned the entire model with a low learning rate. Given Transformer models have several hundred million trainable weights put to the relatively simple task of making a binary prediction, we didn’t believe the default architecture required much tweaking to get a good signal through the model.  Consequently, we didn't spend too much time on hyper-parameter optimization, architectural tweaks, or preprocessing. \n\n### Foreign language monolingual Transformer models\n@rafiko1 and I stumbled on this technique independently and prior to forming a team. We were both inspired by the MultiFiT paper (https://arxiv.org/pdf/1909.04761.pdf) - specifically:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1029053%2Fb40975c9453fba6f640836e4fe297746%2FCapture.PNG?generation=1592872625708960&amp;alt=media)\n\n\n\nWe observed a dramatic performance boost when we used pretrained foreign language monolingual Transformer models from HuggingFace for the test-set languages(e.g., Camembert for french samples, Rubert for russian, BerTurk for turkish, BETO for spanish etc.). \n\nWe finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language (initially from the XLM-R multilingual models), training the corresponding monolingual model, predicting the same samples then blending it back with the “main branch” of all predictions. It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements. \n\nFor each model run, we’d reload weight initalizations from the pretrained models to prevent overfitting. In other words, the continuing improvements we saw were being driven by refinements in the pseudo-labels we were providing to the models as training data. \n\nFor a given monolingual model, predicting only test-set samples in that language worked best. Translating test-set samples in other languages to the model's language and predicting them worsened performance. \n\n### Finetuning pre-trained foreign language monolingual FastText models\nAfter we exhausted the HuggingFace monolingual model library, we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, to continue refining the test set predictions (albeit with a lower weight when combined with the main branch of predictions) and saw a small but meaningful performance boost (0.9536 to 0.9544). \n\n### Post-processing\nGiven the large number of submissions we were making, @rafiko1 came up with the novel idea to make use of the history of submissions to tweak our test set predictions. We tracked the delta of predictions for each sample for successful submissions, averaged them and nudged the predictions in the same direction. We saw a minor but material boost in performance (~0.0005)\n\nWe were concerned about the risk of overfitting with this post-processing technique so for our final 2 submissions, selected one that incorporated this post-processing and one that didn't. It ended up working very well on the private LB. \n\n## Misc - \n### Training setup \n@rafiko1 and I kept separate code-bases. He trained on Kaggle TPU instances with Tensorflow code derived from the public kernels by @xhlulu and @shonenkov while I trained on my own hardware (dual RTX Titans) using from-scratch PyTorch code. I’d attribute part of our outsized merger ensemble boost (0.9510 -&gt; 0.9529) to this diversity in training. \n\n### What didn’t work\nWe went through a checklist of contemporary NLP classification techniques and most of them didn’t work probably due to the simple problem objective. Here is a selection of what didn’t work :- \n        - Document-level embedders (e.g., LASER)\n        - Further MLM pretraining of Transformer models using task data\n        - Alternative ensembling mechanisms (rank-averaging, stochastic weight averaging)\n        - Alternative loss functions (e.g., focal loss, histogram loss)\n        - Alternative pooling mechanisms for the classification head (e.g., max-pool CNN across tokens, using multiple hidden layers etc.)\n        - Non-FastText pretrained embeddings (e.g., Flair, glove, bpem)\n        - Freeze-finetuning for the Transformer models\n        - Regularization (e.g., multisample dropout, input mixup, manifold mixup, sentencepiece-dropout)\n        - Backtranslation as data augmentation\n        - English translations as train/test-time augmentation\n        - Self-distilling to relabel the 2018 data \n        - Adversarial training by perturbing the embeddings layer using FGM\n        - Multi-task learning\n        - Temperature scaling on pseudo-labels\n        - Semi-supervised learning using the test data  \n        - Composing two models into an encoder-decoder model \n        - Use of translation focused pretrained models (e.g., mBart)\n\nWe'll release our code later this week.",
      "votes": 183
    },
    {
      "id": 1590612,
      "postDate": "2021-11-21T13:22:19.493Z",
      "content": "<p>日本語訳<br>\nTL;DR<br>\nアンサンブル、アンサンブル、アンサンブル<br>\n疑似ラベリング<br>\n多言語モデルでブートストラップし、単一言語モデルで洗練する<br>\n2018 Toxic Commentsコンテストで共同優勝して以来、Kaggleに参加していません。 当時のNLP分類の最先端は、文脈に依存しない単語の埋め込み（FastTextなど）でした。トランスフォーマーの世界で上手く見せられるかどうか興味がありました。 チームメイトの@ rafiko1とうまくやってくれて、うれしい驚きです。</p>\n<p>Our public LB milestones<br>\nベースラインXLM-Robertaモデル（公開LB：0.93XX- 3月31日）<br>\nXLM-Rモデルの平均アンサンブル（0.942X- 4月3日）<br>\n単一言語のTransformerモデルとのブレンド（0.9510- 4月23日）<br>\nチームの合併-私たちの個々の最高のsubの加重平均（0.9529- 4月28日）<br>\n単一言語のFastText分類子とのブレンド（0.9544- 5月5日）<br>\n後処理およびその他。 最適化（0.9556- 6月13日）<br>\n4月末までにFINALパブリックLBで1位になりました。 言い換えれば、私たちは競争が終わるのを2ヶ月待っていました:)</p>\n<p>CV strategy<br>\n最初はk-foldCVとバリデーションセットを組み合わせてホールドアウトとして使用しましたが、テスト予測を改良し、トレーニングに疑似ラベル+バリデーションセットを使用すると、検証メトリックがノイズになり、主に一般の人々に依存するようになりました。 LBスコア。</p>\n<p>Insights<br>\nTransformerトレーニングの変動を緩和するためのアンサンブル<br>\nTransformerモデルのパフォーマンスは、初期化とデータの順序によって大きく影響を受けることに注意してください（<a href=\"https://arxiv.org/pdf/2002.06305.pdf、https://www.aclweb.org/anthology/2020.trac-1.9。\" target=\"_blank\">https://arxiv.org/pdf/2002.06305.pdf、https://www.aclweb.org/anthology/2020.trac-1.9。</a> pdf）。 これを軽減するために、モデルのアンサンブルとバギングを強調しました。 これには、一時的な自己アンサンブルが含まれていました。 公開テストセットを前提として、反復ブレンディングアプローチを採用し、以前の最良の提出と現在のモデルの予測の加重平均を使用して、提出全体のテストセットの予測を改良しました。 単純な平均から始めて、前回の最高の提出物の重みを徐々に増やしていきました。 トレーニングデータには、モデルの実行ごとに2018年の有毒なコメントの翻訳のサブサンプルを主に使用しました。</p>\n<p>Pseudo-labels (PL)<br>\nテストセットの予測をトレーニングデータとして使用すると、パフォーマンスの向上が見られました。これは、モデルがテストセットの分布を学習するのに役立つという直感です。 すべてのテストセット予測をソフトラベルとして使用すると、他のバージョンの疑似ラベル（ハードラベル、信頼度しきい値PLなど）よりもうまく機能しました。 競争の終わりに向かって、PLをアップサンプリングしたときに、LBでマイナーではあるが重要なブーストを発見しました。</p>\n<p>Multilingual XLM-Roberta models<br>\nほとんどのチームと同様に、トレーニングデータとして6つのテストセット言語での2018データセットの翻訳を組み込んだバニラXLM-Rモデルから始めました。 アダムオプティマイザーとバイナリクロスエントロピー損失関数を使用して、最後のレイヤーのCLSトークンにバニラ分類ヘッドを使用し、低い学習率でモデル全体を微調整しました。 Transformerモデルには、バイナリ予測を行うという比較的単純なタスクに数億のトレーニング可能な重みが設定されているため、デフォルトのアーキテクチャでは、モデルを通じて適切な信号を取得するために多くの調整が必要であるとは考えていませんでした。 その結果、ハイパーパラメータの最適化、アーキテクチャの調整、または前処理にあまり時間をかけませんでした。</p>\n<p>Foreign language monolingual Transformer models<br>\n@ rafiko1と私は、チームを結成する前に、このテクニックを独自に見つけました。 私たちは両方ともMultiFiTペーパー（<a href=\"https://arxiv.org/pdf/1909.04761.pdf）に触発されました-具体的には：\" target=\"_blank\">https://arxiv.org/pdf/1909.04761.pdf）に触発されました-具体的には：</a></p>\n<p>テストセット言語にHuggingFaceの事前トレーニング済み外国語単一言語Transformerモデルを使用すると、パフォーマンスが劇的に向上することがわかりました（たとえば、フランス語のサンプルにはCamembert、ロシア語にはRubert、トルコ語にはBerTurk、スペイン語にはBETOなど）。</p>\n<p>2018 Toxic Commentsの翻訳とその特定の言語のサンプルの疑似ラベル（最初はXLM-R多言語モデルから）を組み合わせ、対応する単一言語モデルをトレーニングし、同じサンプルを予測することで、6つの言語のそれぞれのモデルを微調整しました次に、それをすべての予測の「メインブランチ」とブレンドし直します。トレーニングモデルA-&gt;トレーニングモデルB-&gt;トレーニングモデルAなどが継続的なパフォーマンスの向上につながるという点で相乗効果がありました。</p>\n<p>モデルの実行ごとに、過剰適合を防ぐために、事前にトレーニングされたモデルからウェイトの初期化をリロードします。言い換えれば、私たちが見た継続的な改善は、トレーニングデータとしてモデルに提供していた疑似ラベルの改良によって推進されていました。</p>\n<p>特定の単一言語モデルでは、その言語のテストセットサンプルのみを予測するのが最も効果的でした。他の言語のテストセットサンプルをモデルの言語に翻訳し、それらを予測すると、パフォーマンスが低下しました。</p>\n<p>事前にトレーニングされた外国語の単一言語FastTextモデルの微調整<br>\nHuggingFaceモノリンガルモデルライブラリを使い果たした後、テストセット言語の事前トレーニング済み埋め込みを使用して、2018年の有毒コメントごとにモノリンガルFastText双方向GRUモデルをトレーニングし、テストセット予測の改良を続けました（メインと組み合わせると重みは小さくなりますが）予測のブランチ）、小さいながらも意味のあるパフォーマンスの向上（0.9536から0.9544）が見られました。</p>\n<p>Post-processing<br>\n多数の提出があったため、@ rafiko1は、提出の履歴を利用してテストセットの予測を微調整するという斬新なアイデアを思いつきました。 提出が成功するまで、各サンプルの予測のデルタを追跡し、それらを平均して、同じ方向に予測を微調整しました。 パフォーマンスはわずかですが大幅に向上しました（〜0.0005）</p>\n<p>この後処理手法に過剰適合するリスクが懸念されたため、最後の2つの提出では、この後処理を組み込んだものと組み込んでいないものを選択しました。 プライベートLBで非常にうまく機能するようになりました。</p>\n<p>Misc -<br>\nトレーニングのセットアップ<br>\n@ rafiko1と私は別々のコードベースを保持していました。 彼は、@ xhluluと@shonenkovによってパブリックカーネルから派生したTensorflowコードを使用してKaggleTPUインスタンスをトレーニングし、私はゼロからのPyTorchコードを使用して独自のハードウェア（デュアルRTX Titans）をトレーニングしました。 特大の合併アンサンブルブースト（0.9510-&gt; 0.9529）の一部は、トレーニングのこの多様性に起因すると思います。</p>\n<p>What didn’t work<br>\n現代のNLP分類手法のチェックリストを確認しましたが、おそらく単純な問題の目的のために、それらのほとんどは機能しませんでした。これがうまくいかなかったものの選択です：-<br>\n-ドキュメントレベルのエンベッダー（例：レーザー）<br>\n-タスクデータを使用したTransformerモデルのさらなるMLM事前トレーニング<br>\n-代替アンサンブルメカニズム（ランク平均、確率的重み平均）<br>\n-代替損失関数（例：焦点損失、ヒストグラム損失）<br>\n-分類ヘッドの代替プーリングメカニズム（たとえば、トークン間での最大プールCNN、複数の非表示レイヤーの使用など）<br>\n-FastText以外の事前トレーニング済みの埋め込み（例：フレア、グローブ、bpem）<br>\n-Transformerモデルのフリーズ微調整<br>\n-正則化（例：マルチサンプルドロップアウト、入力ミックスアップ、多様体ミックスアップ、センテンスピースドロップアウト）<br>\n-データ拡張としての逆翻訳<br>\n-列車/テスト時間の拡張としての英語の翻訳<br>\n-2018年のデータにラベルを付け直すための自己蒸留<br>\n-FGMを使用して埋め込み層を混乱させることによる敵対的訓練<br>\n-マルチタスク学習<br>\n-疑似ラベルの温度スケーリング<br>\n-テストデータを使用した半教師あり学習<br>\n-2つのモデルをエンコーダー-デコーダーモデルに構成する<br>\n-翻訳に焦点を当てた事前トレーニング済みモデルの使用（例：mBart）</p>\n<p>We'll release our code later this week.</p>",
      "rawMarkdown": "日本語訳\nTL;DR\nアンサンブル、アンサンブル、アンサンブル\n疑似ラベリング\n多言語モデルでブートストラップし、単一言語モデルで洗練する\n2018 Toxic Commentsコンテストで共同優勝して以来、Kaggleに参加していません。 当時のNLP分類の最先端は、文脈に依存しない単語の埋め込み（FastTextなど）でした。トランスフォーマーの世界で上手く見せられるかどうか興味がありました。 チームメイトの@ rafiko1とうまくやってくれて、うれしい驚きです。\n\nOur public LB milestones\nベースラインXLM-Robertaモデル（公開LB：0.93XX- 3月31日）\nXLM-Rモデルの平均アンサンブル（0.942X- 4月3日）\n単一言語のTransformerモデルとのブレンド（0.9510- 4月23日）\nチームの合併-私たちの個々の最高のsubの加重平均（0.9529- 4月28日）\n単一言語のFastText分類子とのブレンド（0.9544- 5月5日）\n後処理およびその他。 最適化（0.9556- 6月13日）\n4月末までにFINALパブリックLBで1位になりました。 言い換えれば、私たちは競争が終わるのを2ヶ月待っていました:)\n\nCV strategy\n最初はk-foldCVとバリデーションセットを組み合わせてホールドアウトとして使用しましたが、テスト予測を改良し、トレーニングに疑似ラベル+バリデーションセットを使用すると、検証メトリックがノイズになり、主に一般の人々に依存するようになりました。 LBスコア。\n\nInsights\nTransformerトレーニングの変動を緩和するためのアンサンブル\nTransformerモデルのパフォーマンスは、初期化とデータの順序によって大きく影響を受けることに注意してください（https://arxiv.org/pdf/2002.06305.pdf、https://www.aclweb.org/anthology/2020.trac-1.9。 pdf）。 これを軽減するために、モデルのアンサンブルとバギングを強調しました。 これには、一時的な自己アンサンブルが含まれていました。 公開テストセットを前提として、反復ブレンディングアプローチを採用し、以前の最良の提出と現在のモデルの予測の加重平均を使用して、提出全体のテストセットの予測を改良しました。 単純な平均から始めて、前回の最高の提出物の重みを徐々に増やしていきました。 トレーニングデータには、モデルの実行ごとに2018年の有毒なコメントの翻訳のサブサンプルを主に使用しました。\n\nPseudo-labels (PL)\nテストセットの予測をトレーニングデータとして使用すると、パフォーマンスの向上が見られました。これは、モデルがテストセットの分布を学習するのに役立つという直感です。 すべてのテストセット予測をソフトラベルとして使用すると、他のバージョンの疑似ラベル（ハードラベル、信頼度しきい値PLなど）よりもうまく機能しました。 競争の終わりに向かって、PLをアップサンプリングしたときに、LBでマイナーではあるが重要なブーストを発見しました。\n\nMultilingual XLM-Roberta models\nほとんどのチームと同様に、トレーニングデータとして6つのテストセット言語での2018データセットの翻訳を組み込んだバニラXLM-Rモデルから始めました。 アダムオプティマイザーとバイナリクロスエントロピー損失関数を使用して、最後のレイヤーのCLSトークンにバニラ分類ヘッドを使用し、低い学習率でモデル全体を微調整しました。 Transformerモデルには、バイナリ予測を行うという比較的単純なタスクに数億のトレーニング可能な重みが設定されているため、デフォルトのアーキテクチャでは、モデルを通じて適切な信号を取得するために多くの調整が必要であるとは考えていませんでした。 その結果、ハイパーパラメータの最適化、アーキテクチャの調整、または前処理にあまり時間をかけませんでした。\n\nForeign language monolingual Transformer models\n@ rafiko1と私は、チームを結成する前に、このテクニックを独自に見つけました。 私たちは両方ともMultiFiTペーパー（https://arxiv.org/pdf/1909.04761.pdf）に触発されました-具体的には：\n\n\nテストセット言語にHuggingFaceの事前トレーニング済み外国語単一言語Transformerモデルを使用すると、パフォーマンスが劇的に向上することがわかりました（たとえば、フランス語のサンプルにはCamembert、ロシア語にはRubert、トルコ語にはBerTurk、スペイン語にはBETOなど）。\n\n2018 Toxic Commentsの翻訳とその特定の言語のサンプルの疑似ラベル（最初はXLM-R多言語モデルから）を組み合わせ、対応する単一言語モデルをトレーニングし、同じサンプルを予測することで、6つの言語のそれぞれのモデルを微調整しました次に、それをすべての予測の「メインブランチ」とブレンドし直します。トレーニングモデルA->トレーニングモデルB->トレーニングモデルAなどが継続的なパフォーマンスの向上につながるという点で相乗効果がありました。\n\nモデルの実行ごとに、過剰適合を防ぐために、事前にトレーニングされたモデルからウェイトの初期化をリロードします。言い換えれば、私たちが見た継続的な改善は、トレーニングデータとしてモデルに提供していた疑似ラベルの改良によって推進されていました。\n\n特定の単一言語モデルでは、その言語のテストセットサンプルのみを予測するのが最も効果的でした。他の言語のテストセットサンプルをモデルの言語に翻訳し、それらを予測すると、パフォーマンスが低下しました。\n\n事前にトレーニングされた外国語の単一言語FastTextモデルの微調整\nHuggingFaceモノリンガルモデルライブラリを使い果たした後、テストセット言語の事前トレーニング済み埋め込みを使用して、2018年の有毒コメントごとにモノリンガルFastText双方向GRUモデルをトレーニングし、テストセット予測の改良を続けました（メインと組み合わせると重みは小さくなりますが）予測のブランチ）、小さいながらも意味のあるパフォーマンスの向上（0.9536から0.9544）が見られました。\n\nPost-processing\n多数の提出があったため、@ rafiko1は、提出の履歴を利用してテストセットの予測を微調整するという斬新なアイデアを思いつきました。 提出が成功するまで、各サンプルの予測のデルタを追跡し、それらを平均して、同じ方向に予測を微調整しました。 パフォーマンスはわずかですが大幅に向上しました（〜0.0005）\n\nこの後処理手法に過剰適合するリスクが懸念されたため、最後の2つの提出では、この後処理を組み込んだものと組み込んでいないものを選択しました。 プライベートLBで非常にうまく機能するようになりました。\n\nMisc -\nトレーニングのセットアップ\n@ rafiko1と私は別々のコードベースを保持していました。 彼は、@ xhluluと@shonenkovによってパブリックカーネルから派生したTensorflowコードを使用してKaggleTPUインスタンスをトレーニングし、私はゼロからのPyTorchコードを使用して独自のハードウェア（デュアルRTX Titans）をトレーニングしました。 特大の合併アンサンブルブースト（0.9510-> 0.9529）の一部は、トレーニングのこの多様性に起因すると思います。\n\nWhat didn’t work\n現代のNLP分類手法のチェックリストを確認しましたが、おそらく単純な問題の目的のために、それらのほとんどは機能しませんでした。これがうまくいかなかったものの選択です：-\n-ドキュメントレベルのエンベッダー（例：レーザー）\n-タスクデータを使用したTransformerモデルのさらなるMLM事前トレーニング\n-代替アンサンブルメカニズム（ランク平均、確率的重み平均）\n-代替損失関数（例：焦点損失、ヒストグラム損失）\n-分類ヘッドの代替プーリングメカニズム（たとえば、トークン間での最大プールCNN、複数の非表示レイヤーの使用など）\n-FastText以外の事前トレーニング済みの埋め込み（例：フレア、グローブ、bpem）\n-Transformerモデルのフリーズ微調整\n-正則化（例：マルチサンプルドロップアウト、入力ミックスアップ、多様体ミックスアップ、センテンスピースドロップアウト）\n-データ拡張としての逆翻訳\n-列車/テスト時間の拡張としての英語の翻訳\n-2018年のデータにラベルを付け直すための自己蒸留\n-FGMを使用して埋め込み層を混乱させることによる敵対的訓練\n-マルチタスク学習\n-疑似ラベルの温度スケーリング\n-テストデータを使用した半教師あり学習\n-2つのモデルをエンコーダー-デコーダーモデルに構成する\n-翻訳に焦点を当てた事前トレーニング済みモデルの使用（例：mBart）\n\nWe'll release our code later this week."
    },
    {
      "id": 897946,
      "postDate": "2020-06-23T07:32:13.233Z",
      "content": "<p>Congrats on the the thorough solution and on the win.  Glad to see you back on Kaggle, I remember when we fought in WTF competition.  Yours was so simple and effective!  Same here for effectiveness.</p>",
      "rawMarkdown": "Congrats on the the thorough solution and on the win.  Glad to see you back on Kaggle, I remember when we fought in WTF competition.  Yours was so simple and effective!  Same here for effectiveness.",
      "votes": 3,
      "replies": [
        {
          "id": 899031,
          "postDate": "2020-06-23T22:58:06.740Z",
          "content": "<p>Thanks! WTF was my first gold and a solo gold at that so I'll alway remember it :)</p>\n\n<p>I'm hoping to hit GM by end of the year so hopefully there'll be opportunities for us to team up in another competition.</p>",
          "rawMarkdown": "Thanks! WTF was my first gold and a solo gold at that so I'll alway remember it :)\n\nI'm hoping to hit GM by end of the year so hopefully there'll be opportunities for us to team up in another competition.",
          "votes": 1
        }
      ]
    },
    {
      "id": 963336,
      "postDate": "2020-08-08T23:11:07.153Z",
      "content": "<p><a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a> <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> </p>\n<p>Congrats for being the first place. Your solution is really nice and inspired me a lot. I still have a few questions.</p>\n<blockquote>\n  <p>When training a foreign monolingual foreign language model,  you only use samples in that specific language from the test/val dataset + translations from the orig 2018 training.</p>\n</blockquote>\n<ul>\n<li>Are you using hard labels for training/val set and soft labels for test set? </li>\n<li>Are you using two seperate loss functions during training as mentioned in the LASER paper?  Or you just used a single loss?</li>\n<li>Are losses calculated from hard-target and soft-target weighted differently?</li>\n</ul>\n<blockquote>\n  <p>It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements.</p>\n</blockquote>\n<ul>\n<li>What are models A and B here? Are they both monolingual models or one of them is XLM-R? If one of them is XLM-R, how do you improve the performance of it continuously?  </li>\n</ul>",
      "rawMarkdown": "@leecming @rafiko1 \n\nCongrats for being the first place. Your solution is really nice and inspired me a lot. I still have a few questions.\n\n&gt; When training a foreign monolingual foreign language model,  you only use samples in that specific language from the test/val dataset + translations from the orig 2018 training.\n\n* Are you using hard labels for training/val set and soft labels for test set? \n* Are you using two seperate loss functions during training as mentioned in the LASER paper?  Or you just used a single loss?\n* Are losses calculated from hard-target and soft-target weighted differently?\n\n&gt; It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements.\n\n* What are models A and B here? Are they both monolingual models or one of them is XLM-R? If one of them is XLM-R, how do you improve the performance of it continuously?  \n",
      "votes": 1,
      "replies": [
        {
          "id": 985918,
          "postDate": "2020-08-26T05:18:43.537Z",
          "content": "<ul>\n<li>Yes, mostly. <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a> also worked with soft labels of train/val (distillation)</li>\n<li>Just a single loss for most of our models</li>\n<li>Same weights. What did work is to repeat (i.e. upsample) the test pseudolabels during training - to give them a more even representation with the train set.<br>\n<br></li>\n<li>You can say that model A represents XLM-R as multilingual model. Model B would be a monolingual model specific to the language (see <a href=\"https://github.com/leecming/jigsaw-multilingual#notes\" target=\"_blank\">list</a>). The performance is improved continuously because we refine our (distilled) train labels together with the new test pseudolabels. </li>\n</ul>",
          "rawMarkdown": "- Yes, mostly. @leecming also worked with soft labels of train/val (distillation)\n- Just a single loss for most of our models\n- Same weights. What did work is to repeat (i.e. upsample) the test pseudolabels during training - to give them a more even representation with the train set.\n</br>\n- You can say that model A represents XLM-R as multilingual model. Model B would be a monolingual model specific to the language (see [list](https://github.com/leecming/jigsaw-multilingual#notes)). The performance is improved continuously because we refine our (distilled) train labels together with the new test pseudolabels. "
        }
      ]
    },
    {
      "id": 903414,
      "postDate": "2020-06-26T20:16:31.830Z",
      "content": "<p>Excellent write-up!  Easy to read and sounds like the language specific modeling was a unique approach that might have really helped separate you from the pack.  Well done! 👍 💯 🎉 </p>",
      "rawMarkdown": "Excellent write-up!  Easy to read and sounds like the language specific modeling was a unique approach that might have really helped separate you from the pack.  Well done! 👍 💯 🎉 ",
      "votes": 1
    },
    {
      "id": 900594,
      "postDate": "2020-06-24T23:05:02.320Z",
      "content": "<p>Thanks for sharing the methodology, a very well-deserved win, can't wait for the code release. </p>",
      "rawMarkdown": "Thanks for sharing the methodology, a very well-deserved win, can't wait for the code release. ",
      "votes": 1
    },
    {
      "id": 899154,
      "postDate": "2020-06-24T03:16:46Z",
      "content": "<p>Congrats ! Thank you for the amazing overview ! i have learnt a lot from you ..  </p>",
      "rawMarkdown": "Congrats ! Thank you for the amazing overview ! i have learnt a lot from you ..  ",
      "votes": 1
    },
    {
      "id": 899037,
      "postDate": "2020-06-23T23:03:19.043Z",
      "content": "<p>Congratulations!  I have some questions regarding the monolingual training set up. </p>\n\n<blockquote>\n  <p>We finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language</p>\n</blockquote>\n\n<p>I assume you are using soft labels for the pseudo-labels samples. How do you set up the validation strategy here, particularly for those languages not in the validation set? </p>",
      "rawMarkdown": "Congratulations!  I have some questions regarding the monolingual training set up. \n&gt;We finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language\n\nI assume you are using soft labels for the pseudo-labels samples. How do you set up the validation strategy here, particularly for those languages not in the validation set? \n",
      "votes": 1,
      "replies": [
        {
          "id": 899042,
          "postDate": "2020-06-23T23:17:50.753Z",
          "content": "<p>Soft labels. </p>\n\n<p>Pretty much public LB. To reduce the need for CV, we also used temporal self-ensembling: predicting the test set every epoch and averaging them for the final predictions (<a href=\"/rafiko1\">@rafiko1</a> went beyond and predicted multiple times an epoch). This removed the need to pick a best epoch for predictions.</p>\n\n<p>I did monitor training loss - for larger models that would occassionally fail due to instability, I'd stop and rerun the training if I saw training loss jump up.  </p>",
          "rawMarkdown": "Soft labels. \n\nPretty much public LB. To reduce the need for CV, we also used temporal self-ensembling: predicting the test set every epoch and averaging them for the final predictions (@rafiko1 went beyond and predicted multiple times an epoch). This removed the need to pick a best epoch for predictions.\n\nI did monitor training loss - for larger models that would occassionally fail due to instability, I'd stop and rerun the training if I saw training loss jump up.  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 898240,
      "postDate": "2020-06-23T11:51:20.323Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a>, <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> for the win and thank you so much for sharing your well thought out solution. A well deserved win.</p>\n<blockquote>\n  <p>We began with a simple average, and gradually increased the weight of the previous best submission.</p>\n</blockquote>\n<p>I go out of my way to avoid doing this hence I am surprised the above did not overfit to the public LB. Why do you think that is?</p>\n<blockquote>\n  <p>we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, … and saw a small but meaningful performance boost (0.9536 to 0.9544).</p>\n</blockquote>\n<p>I was surprised at how my low scoring pooled GRU model contributed to my ensemble. I used it more for diversity and just did not think about using it the way you did. Bravo!</p>",
      "rawMarkdown": "Congratulations @leecming, @rafiko1 for the win and thank you so much for sharing your well thought out solution. A well deserved win.\n\n&gt; We began with a simple average, and gradually increased the weight of the previous best submission.\n\nI go out of my way to avoid doing this hence I am surprised the above did not overfit to the public LB. Why do you think that is?\n\n&gt; we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, ... and saw a small but meaningful performance boost (0.9536 to 0.9544).\n\nI was surprised at how my low scoring pooled GRU model contributed to my ensemble. I used it more for diversity and just did not think about using it the way you did. Bravo!",
      "votes": 1,
      "replies": [
        {
          "id": 898262,
          "postDate": "2020-06-23T12:05:12.933Z",
          "content": "<p>The predictions are an exponential moving average of all past model predictions and the current model's prediction. Increasing the weight of the previous best submission means you're putting more weight on past model predictions versus the current model.</p>",
          "rawMarkdown": "The predictions are an exponential moving average of all past model predictions and the current model's prediction. Increasing the weight of the previous best submission means you're putting more weight on past model predictions versus the current model.",
          "votes": 1
        }
      ]
    },
    {
      "id": 898205,
      "postDate": "2020-06-23T11:22:15.137Z",
      "content": "<p>Congratulations ! </p>",
      "rawMarkdown": "Congratulations ! ",
      "votes": 1
    },
    {
      "id": 897957,
      "postDate": "2020-06-23T07:39:32.177Z",
      "content": "<p>Congrats! And thanks for sharing the solution.</p>\n\n<p>I wonder if you also used samples from jigsaw-unintended-bias dataset, in particular, the positive examples?</p>",
      "rawMarkdown": "Congrats! And thanks for sharing the solution.\n\nI wonder if you also used samples from jigsaw-unintended-bias dataset, in particular, the positive examples?",
      "votes": 1,
      "replies": [
        {
          "id": 897971,
          "postDate": "2020-06-23T07:49:32.397Z",
          "content": "<p>Yup, my team-mate did leverage the bias dataset in some of his training runs on TPU. \nI didn't have any success with it. </p>",
          "rawMarkdown": "Yup, my team-mate did leverage the bias dataset in some of his training runs on TPU. \nI didn't have any success with it. ",
          "votes": 2
        },
        {
          "id": 898043,
          "postDate": "2020-06-23T08:41:04.633Z",
          "content": "<p>We didn't have much success with the bias dataset. Only the Spanish-translated bias gave an improvement (~0.0004), using a subset of the negative examples, and positive examples from both the bias dataset and 2018 Spanish-translated</p>",
          "rawMarkdown": "We didn't have much success with the bias dataset. Only the Spanish-translated bias gave an improvement (~0.0004), using a subset of the negative examples, and positive examples from both the bias dataset and 2018 Spanish-translated",
          "votes": 3
        }
      ]
    },
    {
      "id": 897854,
      "postDate": "2020-06-23T06:29:03.930Z",
      "content": "<p>&gt; You’ll note we hit 1st place on the FINAL public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)</p>\n\n<p>I can tell, It was really boring for you guys 😄 </p>\n\n<p>Anyway, congratulation. 🎉 </p>",
      "rawMarkdown": "&gt; You’ll note we hit 1st place on the FINAL public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)\n\nI can tell, It was really boring for you guys 😄 \n\nAnyway, congratulation. 🎉 ",
      "votes": 1
    },
    {
      "id": 897632,
      "postDate": "2020-06-23T02:11:56.053Z",
      "content": "<p>Congrats! you guys have been doing so great and thanks for the detailed summary of your amazing solution, learned a lot:) I got one follow-up, in the post-processing part, what do you mean by 'delta successful submissions'? is it like using high-score subs and calculate the diff between each pair? it would be nice if you can release more details, thanks!</p>",
      "rawMarkdown": "Congrats! you guys have been doing so great and thanks for the detailed summary of your amazing solution, learned a lot:) I got one follow-up, in the post-processing part, what do you mean by 'delta successful submissions'? is it like using high-score subs and calculate the diff between each pair? it would be nice if you can release more details, thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 897940,
          "postDate": "2020-06-23T07:27:32.997Z",
          "content": "<p>I'll ask my team-mate <a href=\"/rafiko1\">@rafiko1</a> to outline the post-processing trick in detail. Think it's useful as a generic PP technique in any competition. </p>",
          "rawMarkdown": "I'll ask my team-mate @rafiko1 to outline the post-processing trick in detail. Think it's useful as a generic PP technique in any competition. "
        }
      ]
    },
    {
      "id": 897623,
      "postDate": "2020-06-23T02:04:11.107Z",
      "content": "<p>Congrats! Amazing performance, quite a gap with second place. Thanks for sharing details of your solution.</p>\n\n<p>The post-processing idea is quite smart and significant as well, good tip for future competitions.</p>\n\n<p>In the \"synergistic\" improvement between cross-lingual and monolingual models, how many steps to convergence typically for how much improvement?</p>",
      "rawMarkdown": "Congrats! Amazing performance, quite a gap with second place. Thanks for sharing details of your solution.\n\nThe post-processing idea is quite smart and significant as well, good tip for future competitions.\n\nIn the \"synergistic\" improvement between cross-lingual and monolingual models, how many steps to convergence typically for how much improvement?",
      "votes": 1,
      "replies": [
        {
          "id": 897937,
          "postDate": "2020-06-23T07:26:42.463Z",
          "content": "<p>We didn't train to convergence - typically trained something along the lines of 3 epochs on the 2018 training dataset (subsample) + PL and 3 epochs on the validation data. </p>",
          "rawMarkdown": "We didn't train to convergence - typically trained something along the lines of 3 epochs on the 2018 training dataset (subsample) + PL and 3 epochs on the validation data. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 897614,
      "postDate": "2020-06-23T01:51:52.340Z",
      "content": "<p>Congratulations !   Your solution is very strong.  Thanks for sharing. </p>",
      "rawMarkdown": "Congratulations !   Your solution is very strong.  Thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 897611,
      "postDate": "2020-06-23T01:51:11.683Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 897589,
      "postDate": "2020-06-23T01:14:36.173Z",
      "content": "<p>Amazing work. I knew there was some value to the monolingual models but couldnt think of a way to integrate them so the distributions of all of the predictions were totally misaligned. I even trained camembert models and rubert, but never got them to play nicely with everything else we had done. Very elegant solution. Almost like pseudolabeling, right?</p>",
      "rawMarkdown": "Amazing work. I knew there was some value to the monolingual models but couldnt think of a way to integrate them so the distributions of all of the predictions were totally misaligned. I even trained camembert models and rubert, but never got them to play nicely with everything else we had done. Very elegant solution. Almost like pseudolabeling, right?",
      "votes": 1
    },
    {
      "id": 898255,
      "postDate": "2020-06-23T11:59:48.883Z",
      "content": "<p>Hi, Congratulations !</p>\n\n<p>Thanks you for sharing some tips about your works.\nI have few questions, : \n- Have you used only large Transformers or also base or small one ?\n- \"- Semi-supervised learning using the test data \" what semi supervised learning did not work exactly ? as you used pseudo labelling which I think can be considered as semi supervised. Are you referring about using hard labels instead of soft labels or is it some others approachs ?\n- about FastText, have you take the embedding and add few GRU layers then pretrained on english data ? And then use a regular training on translated + PL data ? OR did you create one per language ?</p>",
      "rawMarkdown": "Hi, Congratulations !\n\nThanks you for sharing some tips about your works.\nI have few questions, : \n- Have you used only large Transformers or also base or small one ?\n- \"- Semi-supervised learning using the test data \" what semi supervised learning did not work exactly ? as you used pseudo labelling which I think can be considered as semi supervised. Are you referring about using hard labels instead of soft labels or is it some others approachs ?\n- about FastText, have you take the embedding and add few GRU layers then pretrained on english data ? And then use a regular training on translated + PL data ? OR did you create one per language ?",
      "votes": 2,
      "replies": [
        {
          "id": 898265,
          "postDate": "2020-06-23T12:07:47.190Z",
          "content": "<ul>\n<li>We tried everything on HuggingFace. Larger Transformers tended to perform better but also tended to be more unstable (we'd have to re-run training when we observed losses increasing - i saw this especially with models like Flaubert-large)</li>\n<li>I consider SSL distinct from PL in the NLP context, referring more to stuff like the Unsupervised Data Augmentation paper which uses a consistency loss to handle unlabelled samples</li>\n<li>For FT, we used official pretrained embeddings in the 6 test set languages so didn't touch English data</li>\n</ul>",
          "rawMarkdown": "- We tried everything on HuggingFace. Larger Transformers tended to perform better but also tended to be more unstable (we'd have to re-run training when we observed losses increasing - i saw this especially with models like Flaubert-large)\n- I consider SSL distinct from PL in the NLP context, referring more to stuff like the Unsupervised Data Augmentation paper which uses a consistency loss to handle unlabelled samples\n- For FT, we used official pretrained embeddings in the 6 test set languages so didn't touch English data",
          "votes": 2
        }
      ]
    },
    {
      "id": 897651,
      "postDate": "2020-06-23T02:41:42.837Z",
      "content": "<p>Congrats on the 1st place, very worthy. Your team solution is really powerful and impressive!</p>",
      "rawMarkdown": "Congrats on the 1st place, very worthy. Your team solution is really powerful and impressive!",
      "votes": 2
    },
    {
      "id": 897575,
      "postDate": "2020-06-23T00:58:59.753Z",
      "content": "<p>Congratulation on the 1st place, very good overview, it is interesting to see that you achieved top place with a simple model architecture (just the CLS token), as you said such a big model might already have enough parameters, I did some experimentations on that but also did not saw significant improvements.</p>",
      "rawMarkdown": "Congratulation on the 1st place, very good overview, it is interesting to see that you achieved top place with a simple model architecture (just the CLS token), as you said such a big model might already have enough parameters, I did some experimentations on that but also did not saw significant improvements.",
      "votes": 2
    },
    {
      "id": 1480940,
      "postDate": "2021-08-19T08:33:33.167Z",
      "content": "<p>Congratulations on the first place!!  I have a question regarding the low F1, Precision, and Recall score. How to deal with this low score since the class is highly imbalanced.</p>",
      "rawMarkdown": "Congratulations on the first place!!  I have a question regarding the low F1, Precision, and Recall score. How to deal with this low score since the class is highly imbalanced."
    },
    {
      "id": 975061,
      "postDate": "2020-08-18T06:09:30.023Z",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a>, you trained monolingual FastText Bidirectional GRU models in six test languages then concatenate or only in english(translate all data into english) ?</p>",
      "rawMarkdown": "Hello, @leecming, you trained monolingual FastText Bidirectional GRU models in six test languages then concatenate or only in english(translate all data into english) ?"
    },
    {
      "id": 899332,
      "postDate": "2020-06-24T07:05:34.880Z",
      "content": "<p>Hi first of all Congrats on winning the competition and for an amazing approach. Can you please point me to the papers of above monolingual models. I was able to find the below.</p>\n\n<p>[1] : Camembert (<a href=\"https://arxiv.org/pdf/1911.03894.pdf?fbclid=IwAR03h_fUIfSqPlaqTkMew4mybw-LjBrh8dZ3lpTgQeldLdsZFfr54E7nzk4\">https://arxiv.org/pdf/1911.03894.pdf?fbclid=IwAR03h_fUIfSqPlaqTkMew4mybw-LjBrh8dZ3lpTgQeldLdsZFfr54E7nzk4</a>) </p>\n\n<p>[2] : Beto (<a href=\"https://arxiv.org/abs/1904.09077?fbclid=IwAR1mieFlVqpl0iHZq2DAA0dSJ7uDCBpLoCeX_p6CIY69n7RTQQrFuDFNlm4\">https://arxiv.org/abs/1904.09077?fbclid=IwAR1mieFlVqpl0iHZq2DAA0dSJ7uDCBpLoCeX_p6CIY69n7RTQQrFuDFNlm4</a>)</p>\n\n<p>I wasn't able to find BerTurk and Rubert. Any help would be appreciated. Also the corresponding usage of such models. Thanks.</p>",
      "rawMarkdown": "Hi first of all Congrats on winning the competition and for an amazing approach. Can you please point me to the papers of above monolingual models. I was able to find the below.\n\n[1] : Camembert (https://arxiv.org/pdf/1911.03894.pdf?fbclid=IwAR03h_fUIfSqPlaqTkMew4mybw-LjBrh8dZ3lpTgQeldLdsZFfr54E7nzk4) \n\n[2] : Beto (https://arxiv.org/abs/1904.09077?fbclid=IwAR1mieFlVqpl0iHZq2DAA0dSJ7uDCBpLoCeX_p6CIY69n7RTQQrFuDFNlm4)\n\nI wasn't able to find BerTurk and Rubert. Any help would be appreciated. Also the corresponding usage of such models. Thanks.",
      "replies": [
        {
          "id": 899866,
          "postDate": "2020-06-24T13:45:27.320Z",
          "content": "<p>Not all of the models have their separate papers. BertTurk just has a <a href=\"https://github.com/stefan-it/turkish-bert\">github</a> for example. You can find most models and their usage in <a href=\"https://huggingface.co/models\">huggingface</a></p>",
          "rawMarkdown": "Not all of the models have their separate papers. BertTurk just has a [github](https://github.com/stefan-it/turkish-bert) for example. You can find most models and their usage in [huggingface](https://huggingface.co/models)"
        }
      ]
    },
    {
      "id": 898701,
      "postDate": "2020-06-23T17:20:48.637Z",
      "content": "<p>Hi, I got another question.</p>\n\n<p>For fine-tuning <code>a foreign language monolingual Transformer models</code>, what is the training dataset?\nDid you use only the samples in that specific language from test dataset? (If so, the training samples are quite few).\nOr you also included the translated samples in the target lang from the original training/validation dataset?</p>",
      "rawMarkdown": "Hi, I got another question.\n\nFor fine-tuning `a foreign language monolingual Transformer models`, what is the training dataset?\nDid you use only the samples in that specific language from test dataset? (If so, the training samples are quite few).\nOr you also included the translated samples in the target lang from the original training/validation dataset?\n",
      "replies": [
        {
          "id": 899020,
          "postDate": "2020-06-23T22:44:29.387Z",
          "content": "<p>Only samples in that specific language from the test/val dataset + translations from the orig 2018 training.</p>\n\n<p>A training run would for example have 10K test + 70K subsampled 2018 translations + 2.5K val.\nThese models are quite good at few-shot learning so &lt;100K is sufficient to learn. </p>",
          "rawMarkdown": "Only samples in that specific language from the test/val dataset + translations from the orig 2018 training.\n\nA training run would for example have 10K test + 70K subsampled 2018 translations + 2.5K val.\nThese models are quite good at few-shot learning so &lt;100K is sufficient to learn. ",
          "votes": 1
        },
        {
          "id": 899029,
          "postDate": "2020-06-23T22:55:10.897Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 897647,
      "postDate": "2020-06-23T02:36:54.403Z",
      "content": "<p>Is your Private LB score also has similar improvement as your Public LB milestones? Or any of the techniques (in your milestones) provide much lesser contribution after taking Private LB score into consideration? </p>",
      "rawMarkdown": "Is your Private LB score also has similar improvement as your Public LB milestones? Or any of the techniques (in your milestones) provide much lesser contribution after taking Private LB score into consideration? ",
      "replies": [
        {
          "id": 897935,
          "postDate": "2020-06-23T07:25:21.583Z",
          "content": "<p>Private more or less tracked with public. Not sure about lesser contribution but the two biggest contributors were: 1) monolingual models, and 2) our team merger ensemble boost.</p>",
          "rawMarkdown": "Private more or less tracked with public. Not sure about lesser contribution but the two biggest contributors were: 1) monolingual models, and 2) our team merger ensemble boost.",
          "votes": 1
        }
      ]
    },
    {
      "id": 897574,
      "postDate": "2020-06-23T00:58:36.217Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 973735,
      "postDate": "2020-08-17T13:37:22.660Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 1590612,
      "author_name": "pixyz0130",
      "author_url": "",
      "post_date": "2021-11-21T13:22:19.493000",
      "content": "<p>日本語訳<br>\nTL;DR<br>\nアンサンブル、アンサンブル、アンサンブル<br>\n疑似ラベリング<br>\n多言語モデルでブートストラップし、単一言語モデルで洗練する<br>\n2018 Toxic Commentsコンテストで共同優勝して以来、Kaggleに参加していません。 当時のNLP分類の最先端は、文脈に依存しない単語の埋め込み（FastTextなど）でした。トランスフォーマーの世界で上手く見せられるかどうか興味がありました。 チームメイトの@ rafiko1とうまくやってくれて、うれしい驚きです。</p>\n<p>Our public LB milestones<br>\nベースラインXLM-Robertaモデル（公開LB：0.93XX- 3月31日）<br>\nXLM-Rモデルの平均アンサンブル（0.942X- 4月3日）<br>\n単一言語のTransformerモデルとのブレンド（0.9510- 4月23日）<br>\nチームの合併-私たちの個々の最高のsubの加重平均（0.9529- 4月28日）<br>\n単一言語のFastText分類子とのブレンド（0.9544- 5月5日）<br>\n後処理およびその他。 最適化（0.9556- 6月13日）<br>\n4月末までにFINALパブリックLBで1位になりました。 言い換えれば、私たちは競争が終わるのを2ヶ月待っていました:)</p>\n<p>CV strategy<br>\n最初はk-foldCVとバリデーションセットを組み合わせてホールドアウトとして使用しましたが、テスト予測を改良し、トレーニングに疑似ラベル+バリデーションセットを使用すると、検証メトリックがノイズになり、主に一般の人々に依存するようになりました。 LBスコア。</p>\n<p>Insights<br>\nTransformerトレーニングの変動を緩和するためのアンサンブル<br>\nTransformerモデルのパフォーマンスは、初期化とデータの順序によって大きく影響を受けることに注意してください（<a href=\"https://arxiv.org/pdf/2002.06305.pdf、https://www.aclweb.org/anthology/2020.trac-1.9。\" target=\"_blank\">https://arxiv.org/pdf/2002.06305.pdf、https://www.aclweb.org/anthology/2020.trac-1.9。</a> pdf）。 これを軽減するために、モデルのアンサンブルとバギングを強調しました。 これには、一時的な自己アンサンブルが含まれていました。 公開テストセットを前提として、反復ブレンディングアプローチを採用し、以前の最良の提出と現在のモデルの予測の加重平均を使用して、提出全体のテストセットの予測を改良しました。 単純な平均から始めて、前回の最高の提出物の重みを徐々に増やしていきました。 トレーニングデータには、モデルの実行ごとに2018年の有毒なコメントの翻訳のサブサンプルを主に使用しました。</p>\n<p>Pseudo-labels (PL)<br>\nテストセットの予測をトレーニングデータとして使用すると、パフォーマンスの向上が見られました。これは、モデルがテストセットの分布を学習するのに役立つという直感です。 すべてのテストセット予測をソフトラベルとして使用すると、他のバージョンの疑似ラベル（ハードラベル、信頼度しきい値PLなど）よりもうまく機能しました。 競争の終わりに向かって、PLをアップサンプリングしたときに、LBでマイナーではあるが重要なブーストを発見しました。</p>\n<p>Multilingual XLM-Roberta models<br>\nほとんどのチームと同様に、トレーニングデータとして6つのテストセット言語での2018データセットの翻訳を組み込んだバニラXLM-Rモデルから始めました。 アダムオプティマイザーとバイナリクロスエントロピー損失関数を使用して、最後のレイヤーのCLSトークンにバニラ分類ヘッドを使用し、低い学習率でモデル全体を微調整しました。 Transformerモデルには、バイナリ予測を行うという比較的単純なタスクに数億のトレーニング可能な重みが設定されているため、デフォルトのアーキテクチャでは、モデルを通じて適切な信号を取得するために多くの調整が必要であるとは考えていませんでした。 その結果、ハイパーパラメータの最適化、アーキテクチャの調整、または前処理にあまり時間をかけませんでした。</p>\n<p>Foreign language monolingual Transformer models<br>\n@ rafiko1と私は、チームを結成する前に、このテクニックを独自に見つけました。 私たちは両方ともMultiFiTペーパー（<a href=\"https://arxiv.org/pdf/1909.04761.pdf）に触発されました-具体的には：\" target=\"_blank\">https://arxiv.org/pdf/1909.04761.pdf）に触発されました-具体的には：</a></p>\n<p>テストセット言語にHuggingFaceの事前トレーニング済み外国語単一言語Transformerモデルを使用すると、パフォーマンスが劇的に向上することがわかりました（たとえば、フランス語のサンプルにはCamembert、ロシア語にはRubert、トルコ語にはBerTurk、スペイン語にはBETOなど）。</p>\n<p>2018 Toxic Commentsの翻訳とその特定の言語のサンプルの疑似ラベル（最初はXLM-R多言語モデルから）を組み合わせ、対応する単一言語モデルをトレーニングし、同じサンプルを予測することで、6つの言語のそれぞれのモデルを微調整しました次に、それをすべての予測の「メインブランチ」とブレンドし直します。トレーニングモデルA-&gt;トレーニングモデルB-&gt;トレーニングモデルAなどが継続的なパフォーマンスの向上につながるという点で相乗効果がありました。</p>\n<p>モデルの実行ごとに、過剰適合を防ぐために、事前にトレーニングされたモデルからウェイトの初期化をリロードします。言い換えれば、私たちが見た継続的な改善は、トレーニングデータとしてモデルに提供していた疑似ラベルの改良によって推進されていました。</p>\n<p>特定の単一言語モデルでは、その言語のテストセットサンプルのみを予測するのが最も効果的でした。他の言語のテストセットサンプルをモデルの言語に翻訳し、それらを予測すると、パフォーマンスが低下しました。</p>\n<p>事前にトレーニングされた外国語の単一言語FastTextモデルの微調整<br>\nHuggingFaceモノリンガルモデルライブラリを使い果たした後、テストセット言語の事前トレーニング済み埋め込みを使用して、2018年の有毒コメントごとにモノリンガルFastText双方向GRUモデルをトレーニングし、テストセット予測の改良を続けました（メインと組み合わせると重みは小さくなりますが）予測のブランチ）、小さいながらも意味のあるパフォーマンスの向上（0.9536から0.9544）が見られました。</p>\n<p>Post-processing<br>\n多数の提出があったため、@ rafiko1は、提出の履歴を利用してテストセットの予測を微調整するという斬新なアイデアを思いつきました。 提出が成功するまで、各サンプルの予測のデルタを追跡し、それらを平均して、同じ方向に予測を微調整しました。 パフォーマンスはわずかですが大幅に向上しました（〜0.0005）</p>\n<p>この後処理手法に過剰適合するリスクが懸念されたため、最後の2つの提出では、この後処理を組み込んだものと組み込んでいないものを選択しました。 プライベートLBで非常にうまく機能するようになりました。</p>\n<p>Misc -<br>\nトレーニングのセットアップ<br>\n@ rafiko1と私は別々のコードベースを保持していました。 彼は、@ xhluluと@shonenkovによってパブリックカーネルから派生したTensorflowコードを使用してKaggleTPUインスタンスをトレーニングし、私はゼロからのPyTorchコードを使用して独自のハードウェア（デュアルRTX Titans）をトレーニングしました。 特大の合併アンサンブルブースト（0.9510-&gt; 0.9529）の一部は、トレーニングのこの多様性に起因すると思います。</p>\n<p>What didn’t work<br>\n現代のNLP分類手法のチェックリストを確認しましたが、おそらく単純な問題の目的のために、それらのほとんどは機能しませんでした。これがうまくいかなかったものの選択です：-<br>\n-ドキュメントレベルのエンベッダー（例：レーザー）<br>\n-タスクデータを使用したTransformerモデルのさらなるMLM事前トレーニング<br>\n-代替アンサンブルメカニズム（ランク平均、確率的重み平均）<br>\n-代替損失関数（例：焦点損失、ヒストグラム損失）<br>\n-分類ヘッドの代替プーリングメカニズム（たとえば、トークン間での最大プールCNN、複数の非表示レイヤーの使用など）<br>\n-FastText以外の事前トレーニング済みの埋め込み（例：フレア、グローブ、bpem）<br>\n-Transformerモデルのフリーズ微調整<br>\n-正則化（例：マルチサンプルドロップアウト、入力ミックスアップ、多様体ミックスアップ、センテンスピースドロップアウト）<br>\n-データ拡張としての逆翻訳<br>\n-列車/テスト時間の拡張としての英語の翻訳<br>\n-2018年のデータにラベルを付け直すための自己蒸留<br>\n-FGMを使用して埋め込み層を混乱させることによる敵対的訓練<br>\n-マルチタスク学習<br>\n-疑似ラベルの温度スケーリング<br>\n-テストデータを使用した半教師あり学習<br>\n-2つのモデルをエンコーダー-デコーダーモデルに構成する<br>\n-翻訳に焦点を当てた事前トレーニング済みモデルの使用（例：mBart）</p>\n<p>We'll release our code later this week.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 897946,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-06-23T07:32:13.233000",
      "content": "<p>Congrats on the the thorough solution and on the win.  Glad to see you back on Kaggle, I remember when we fought in WTF competition.  Yours was so simple and effective!  Same here for effectiveness.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 899031,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T22:58:06.740000",
          "content": "<p>Thanks! WTF was my first gold and a solo gold at that so I'll alway remember it :)</p>\n\n<p>I'm hoping to hit GM by end of the year so hopefully there'll be opportunities for us to team up in another competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 963336,
      "author_name": "Liyan Tang",
      "author_url": "",
      "post_date": "2020-08-08T23:11:07.153000",
      "content": "<p><a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a> <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> </p>\n<p>Congrats for being the first place. Your solution is really nice and inspired me a lot. I still have a few questions.</p>\n<blockquote>\n  <p>When training a foreign monolingual foreign language model,  you only use samples in that specific language from the test/val dataset + translations from the orig 2018 training.</p>\n</blockquote>\n<ul>\n<li>Are you using hard labels for training/val set and soft labels for test set? </li>\n<li>Are you using two seperate loss functions during training as mentioned in the LASER paper?  Or you just used a single loss?</li>\n<li>Are losses calculated from hard-target and soft-target weighted differently?</li>\n</ul>\n<blockquote>\n  <p>It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements.</p>\n</blockquote>\n<ul>\n<li>What are models A and B here? Are they both monolingual models or one of them is XLM-R? If one of them is XLM-R, how do you improve the performance of it continuously?  </li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 985918,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2020-08-26T05:18:43.537000",
          "content": "<ul>\n<li>Yes, mostly. <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a> also worked with soft labels of train/val (distillation)</li>\n<li>Just a single loss for most of our models</li>\n<li>Same weights. What did work is to repeat (i.e. upsample) the test pseudolabels during training - to give them a more even representation with the train set.<br>\n<br></li>\n<li>You can say that model A represents XLM-R as multilingual model. Model B would be a monolingual model specific to the language (see <a href=\"https://github.com/leecming/jigsaw-multilingual#notes\" target=\"_blank\">list</a>). The performance is improved continuously because we refine our (distilled) train labels together with the new test pseudolabels. </li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 903414,
      "author_name": "Matt Yates",
      "author_url": "",
      "post_date": "2020-06-26T20:16:31.830000",
      "content": "<p>Excellent write-up!  Easy to read and sounds like the language specific modeling was a unique approach that might have really helped separate you from the pack.  Well done! 👍 💯 🎉 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 900594,
      "author_name": "Huan Vo",
      "author_url": "",
      "post_date": "2020-06-24T23:05:02.320000",
      "content": "<p>Thanks for sharing the methodology, a very well-deserved win, can't wait for the code release. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 899154,
      "author_name": "Ashraf Mahdhi",
      "author_url": "",
      "post_date": "2020-06-24T03:16:46",
      "content": "<p>Congrats ! Thank you for the amazing overview ! i have learnt a lot from you ..  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 899037,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2020-06-23T23:03:19.043000",
      "content": "<p>Congratulations!  I have some questions regarding the monolingual training set up. </p>\n\n<blockquote>\n  <p>We finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language</p>\n</blockquote>\n\n<p>I assume you are using soft labels for the pseudo-labels samples. How do you set up the validation strategy here, particularly for those languages not in the validation set? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 899042,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T23:17:50.753000",
          "content": "<p>Soft labels. </p>\n\n<p>Pretty much public LB. To reduce the need for CV, we also used temporal self-ensembling: predicting the test set every epoch and averaging them for the final predictions (<a href=\"/rafiko1\">@rafiko1</a> went beyond and predicted multiple times an epoch). This removed the need to pick a best epoch for predictions.</p>\n\n<p>I did monitor training loss - for larger models that would occassionally fail due to instability, I'd stop and rerun the training if I saw training loss jump up.  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 898240,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-06-23T11:51:20.323000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a>, <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> for the win and thank you so much for sharing your well thought out solution. A well deserved win.</p>\n<blockquote>\n  <p>We began with a simple average, and gradually increased the weight of the previous best submission.</p>\n</blockquote>\n<p>I go out of my way to avoid doing this hence I am surprised the above did not overfit to the public LB. Why do you think that is?</p>\n<blockquote>\n  <p>we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, … and saw a small but meaningful performance boost (0.9536 to 0.9544).</p>\n</blockquote>\n<p>I was surprised at how my low scoring pooled GRU model contributed to my ensemble. I used it more for diversity and just did not think about using it the way you did. Bravo!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 898262,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T12:05:12.933000",
          "content": "<p>The predictions are an exponential moving average of all past model predictions and the current model's prediction. Increasing the weight of the previous best submission means you're putting more weight on past model predictions versus the current model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 898205,
      "author_name": "Felipe Jardim Fiorentino",
      "author_url": "",
      "post_date": "2020-06-23T11:22:15.137000",
      "content": "<p>Congratulations ! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 897957,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-06-23T07:39:32.177000",
      "content": "<p>Congrats! And thanks for sharing the solution.</p>\n\n<p>I wonder if you also used samples from jigsaw-unintended-bias dataset, in particular, the positive examples?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 897971,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T07:49:32.397000",
          "content": "<p>Yup, my team-mate did leverage the bias dataset in some of his training runs on TPU. \nI didn't have any success with it. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 898043,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2020-06-23T08:41:04.633000",
          "content": "<p>We didn't have much success with the bias dataset. Only the Spanish-translated bias gave an improvement (~0.0004), using a subset of the negative examples, and positive examples from both the bias dataset and 2018 Spanish-translated</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 897854,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-06-23T06:29:03.930000",
      "content": "<p>&gt; You’ll note we hit 1st place on the FINAL public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)</p>\n\n<p>I can tell, It was really boring for you guys 😄 </p>\n\n<p>Anyway, congratulation. 🎉 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 897632,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-23T02:11:56.053000",
      "content": "<p>Congrats! you guys have been doing so great and thanks for the detailed summary of your amazing solution, learned a lot:) I got one follow-up, in the post-processing part, what do you mean by 'delta successful submissions'? is it like using high-score subs and calculate the diff between each pair? it would be nice if you can release more details, thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 897940,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T07:27:32.997000",
          "content": "<p>I'll ask my team-mate <a href=\"/rafiko1\">@rafiko1</a> to outline the post-processing trick in detail. Think it's useful as a generic PP technique in any competition. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 897623,
      "author_name": "Seba",
      "author_url": "",
      "post_date": "2020-06-23T02:04:11.107000",
      "content": "<p>Congrats! Amazing performance, quite a gap with second place. Thanks for sharing details of your solution.</p>\n\n<p>The post-processing idea is quite smart and significant as well, good tip for future competitions.</p>\n\n<p>In the \"synergistic\" improvement between cross-lingual and monolingual models, how many steps to convergence typically for how much improvement?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 897937,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T07:26:42.463000",
          "content": "<p>We didn't train to convergence - typically trained something along the lines of 3 epochs on the 2018 training dataset (subsample) + PL and 3 epochs on the validation data. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 897614,
      "author_name": "huiqin",
      "author_url": "",
      "post_date": "2020-06-23T01:51:52.340000",
      "content": "<p>Congratulations !   Your solution is very strong.  Thanks for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 897611,
      "author_name": "Hieu Nguyen",
      "author_url": "",
      "post_date": "2020-06-23T01:51:11.683000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 897589,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-06-23T01:14:36.173000",
      "content": "<p>Amazing work. I knew there was some value to the monolingual models but couldnt think of a way to integrate them so the distributions of all of the predictions were totally misaligned. I even trained camembert models and rubert, but never got them to play nicely with everything else we had done. Very elegant solution. Almost like pseudolabeling, right?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 898255,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2020-06-23T11:59:48.883000",
      "content": "<p>Hi, Congratulations !</p>\n\n<p>Thanks you for sharing some tips about your works.\nI have few questions, : \n- Have you used only large Transformers or also base or small one ?\n- \"- Semi-supervised learning using the test data \" what semi supervised learning did not work exactly ? as you used pseudo labelling which I think can be considered as semi supervised. Are you referring about using hard labels instead of soft labels or is it some others approachs ?\n- about FastText, have you take the embedding and add few GRU layers then pretrained on english data ? And then use a regular training on translated + PL data ? OR did you create one per language ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 898265,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T12:07:47.190000",
          "content": "<ul>\n<li>We tried everything on HuggingFace. Larger Transformers tended to perform better but also tended to be more unstable (we'd have to re-run training when we observed losses increasing - i saw this especially with models like Flaubert-large)</li>\n<li>I consider SSL distinct from PL in the NLP context, referring more to stuff like the Unsupervised Data Augmentation paper which uses a consistency loss to handle unlabelled samples</li>\n<li>For FT, we used official pretrained embeddings in the 6 test set languages so didn't touch English data</li>\n</ul>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 897651,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-06-23T02:41:42.837000",
      "content": "<p>Congrats on the 1st place, very worthy. Your team solution is really powerful and impressive!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 897575,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-06-23T00:58:59.753000",
      "content": "<p>Congratulation on the 1st place, very good overview, it is interesting to see that you achieved top place with a simple model architecture (just the CLS token), as you said such a big model might already have enough parameters, I did some experimentations on that but also did not saw significant improvements.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1480940,
      "author_name": "Shubhesh Swain",
      "author_url": "",
      "post_date": "2021-08-19T08:33:33.167000",
      "content": "<p>Congratulations on the first place!!  I have a question regarding the low F1, Precision, and Recall score. How to deal with this low score since the class is highly imbalanced.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975061,
      "author_name": "README",
      "author_url": "",
      "post_date": "2020-08-18T06:09:30.023000",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/leecming\" target=\"_blank\">@leecming</a>, you trained monolingual FastText Bidirectional GRU models in six test languages then concatenate or only in english(translate all data into english) ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 899332,
      "author_name": "JamshaidSohail",
      "author_url": "",
      "post_date": "2020-06-24T07:05:34.880000",
      "content": "<p>Hi first of all Congrats on winning the competition and for an amazing approach. Can you please point me to the papers of above monolingual models. I was able to find the below.</p>\n\n<p>[1] : Camembert (<a href=\"https://arxiv.org/pdf/1911.03894.pdf?fbclid=IwAR03h_fUIfSqPlaqTkMew4mybw-LjBrh8dZ3lpTgQeldLdsZFfr54E7nzk4\">https://arxiv.org/pdf/1911.03894.pdf?fbclid=IwAR03h_fUIfSqPlaqTkMew4mybw-LjBrh8dZ3lpTgQeldLdsZFfr54E7nzk4</a>) </p>\n\n<p>[2] : Beto (<a href=\"https://arxiv.org/abs/1904.09077?fbclid=IwAR1mieFlVqpl0iHZq2DAA0dSJ7uDCBpLoCeX_p6CIY69n7RTQQrFuDFNlm4\">https://arxiv.org/abs/1904.09077?fbclid=IwAR1mieFlVqpl0iHZq2DAA0dSJ7uDCBpLoCeX_p6CIY69n7RTQQrFuDFNlm4</a>)</p>\n\n<p>I wasn't able to find BerTurk and Rubert. Any help would be appreciated. Also the corresponding usage of such models. Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 899866,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2020-06-24T13:45:27.320000",
          "content": "<p>Not all of the models have their separate papers. BertTurk just has a <a href=\"https://github.com/stefan-it/turkish-bert\">github</a> for example. You can find most models and their usage in <a href=\"https://huggingface.co/models\">huggingface</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 898701,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-06-23T17:20:48.637000",
      "content": "<p>Hi, I got another question.</p>\n\n<p>For fine-tuning <code>a foreign language monolingual Transformer models</code>, what is the training dataset?\nDid you use only the samples in that specific language from test dataset? (If so, the training samples are quite few).\nOr you also included the translated samples in the target lang from the original training/validation dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 899020,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T22:44:29.387000",
          "content": "<p>Only samples in that specific language from the test/val dataset + translations from the orig 2018 training.</p>\n\n<p>A training run would for example have 10K test + 70K subsampled 2018 translations + 2.5K val.\nThese models are quite good at few-shot learning so &lt;100K is sufficient to learn. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 899029,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-06-23T22:55:10.897000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 897647,
      "author_name": "Chew Kok Wah",
      "author_url": "",
      "post_date": "2020-06-23T02:36:54.403000",
      "content": "<p>Is your Private LB score also has similar improvement as your Public LB milestones? Or any of the techniques (in your milestones) provide much lesser contribution after taking Private LB score into consideration? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 897935,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-23T07:25:21.583000",
          "content": "<p>Private more or less tracked with public. Not sure about lesser contribution but the two biggest contributors were: 1) monolingual models, and 2) our team merger ensemble boost.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 897574,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-23T00:58:36.217000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 973735,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-17T13:37:22.660000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "897564": "We’d like to start off by thanking Kaggle/Jigsaw for a drama-free competition and also by congratulating the other medallists. \n\n## TL;DR\n1. Ensemble, ensemble, ensemble\n2. Pseudo-labelling\n3. Bootstrap with multilingual models, refine with monolingual models\n\n\nI haven't competed on Kaggle since I co-won the 2018 Toxic Comments competition. The state-of-the-art for NLP classification at the time was non-contextual word embeddings (e.g., FastText) and I was curious if I’d make a good showing in a world of Transformers. I'm pleasantly surprised to have done so well with my team-mate @rafiko1. \n\n## Our public LB milestones\n- Baseline XLM-Roberta model (Public LB: 0.93XX - 31 March)\n- Average ensemble of XLM-R models (0.942X - 3 April)\n- Blending with monolingual Transformer models (0.9510 - 23 April)\n- Team merger - weighted average of our individual best subs (0.9529 - 28 April)\n- Blending with monolingual FastText classifiers (0.9544 - 5 May)\n- Post-processing and misc. optimizations (0.9556 - 13 June)\n\nYou’ll note we hit 1st place on the **FINAL** public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)\n\n## CV strategy\nWe initially used a mix of k-fold CV and validation set as hold-out but as we refined our test predictions and used pseudo-labels + validation set for training, the validation metric became noisy to the point where we relied primarily on the public LB score. \n\n## Insights\n### Ensembling to mitigate Transformer training variability\nIt’s been noted that the performance of Transformer models is impacted heavily by initialization and data order (https://arxiv.org/pdf/2002.06305.pdf,  https://www.aclweb.org/anthology/2020.trac-1.9.pdf). To mitigate that, we emphasized the ensembling and bagging our models. This included temporal self-ensembling. Given the public test set, we went with an iterative blending approach, refining the test set predictions across submissions with a weighted average of the previous best submission and the current model’s predictions. We began with a simple average, and gradually increased the weight of the previous best submission. For the training data, we largely used sub-samples of the translations of the 2018 toxic comments for each model run. \n\n### Pseudo-labels (PL)\n We observed a performance improvement when we used test-set predictions as training data - the intuition being that it helps models learn the test set distribution. Using all test-set predictions as soft-labels worked better than any other version of pseudo-labelling (e.g., hard labels, confidence thresholded PLs etc.). Towards the end of the competition, we discovered a minor but material boost in LB when we upsampled the PLs. \n\n### Multilingual XLM-Roberta models\nAs with most teams, we began with a vanilla XLM-R model, incorporating translations of the 2018 dataset in the 6 test-set languages as training data. We used a vanilla classification head on the CLS token of the last layer with the Adam optimizer and binary cross entropy loss function, and finetuned the entire model with a low learning rate. Given Transformer models have several hundred million trainable weights put to the relatively simple task of making a binary prediction, we didn’t believe the default architecture required much tweaking to get a good signal through the model.  Consequently, we didn't spend too much time on hyper-parameter optimization, architectural tweaks, or preprocessing. \n\n### Foreign language monolingual Transformer models\n@rafiko1 and I stumbled on this technique independently and prior to forming a team. We were both inspired by the MultiFiT paper (https://arxiv.org/pdf/1909.04761.pdf) - specifically:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1029053%2Fb40975c9453fba6f640836e4fe297746%2FCapture.PNG?generation=1592872625708960&amp;alt=media)\n\n\n\nWe observed a dramatic performance boost when we used pretrained foreign language monolingual Transformer models from HuggingFace for the test-set languages(e.g., Camembert for french samples, Rubert for russian, BerTurk for turkish, BETO for spanish etc.). \n\nWe finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language (initially from the XLM-R multilingual models), training the corresponding monolingual model, predicting the same samples then blending it back with the “main branch” of all predictions. It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements. \n\nFor each model run, we’d reload weight initalizations from the pretrained models to prevent overfitting. In other words, the continuing improvements we saw were being driven by refinements in the pseudo-labels we were providing to the models as training data. \n\nFor a given monolingual model, predicting only test-set samples in that language worked best. Translating test-set samples in other languages to the model's language and predicting them worsened performance. \n\n### Finetuning pre-trained foreign language monolingual FastText models\nAfter we exhausted the HuggingFace monolingual model library, we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, to continue refining the test set predictions (albeit with a lower weight when combined with the main branch of predictions) and saw a small but meaningful performance boost (0.9536 to 0.9544). \n\n### Post-processing\nGiven the large number of submissions we were making, @rafiko1 came up with the novel idea to make use of the history of submissions to tweak our test set predictions. We tracked the delta of predictions for each sample for successful submissions, averaged them and nudged the predictions in the same direction. We saw a minor but material boost in performance (~0.0005)\n\nWe were concerned about the risk of overfitting with this post-processing technique so for our final 2 submissions, selected one that incorporated this post-processing and one that didn't. It ended up working very well on the private LB. \n\n## Misc - \n### Training setup \n@rafiko1 and I kept separate code-bases. He trained on Kaggle TPU instances with Tensorflow code derived from the public kernels by @xhlulu and @shonenkov while I trained on my own hardware (dual RTX Titans) using from-scratch PyTorch code. I’d attribute part of our outsized merger ensemble boost (0.9510 -&gt; 0.9529) to this diversity in training. \n\n### What didn’t work\nWe went through a checklist of contemporary NLP classification techniques and most of them didn’t work probably due to the simple problem objective. Here is a selection of what didn’t work :- \n        - Document-level embedders (e.g., LASER)\n        - Further MLM pretraining of Transformer models using task data\n        - Alternative ensembling mechanisms (rank-averaging, stochastic weight averaging)\n        - Alternative loss functions (e.g., focal loss, histogram loss)\n        - Alternative pooling mechanisms for the classification head (e.g., max-pool CNN across tokens, using multiple hidden layers etc.)\n        - Non-FastText pretrained embeddings (e.g., Flair, glove, bpem)\n        - Freeze-finetuning for the Transformer models\n        - Regularization (e.g., multisample dropout, input mixup, manifold mixup, sentencepiece-dropout)\n        - Backtranslation as data augmentation\n        - English translations as train/test-time augmentation\n        - Self-distilling to relabel the 2018 data \n        - Adversarial training by perturbing the embeddings layer using FGM\n        - Multi-task learning\n        - Temperature scaling on pseudo-labels\n        - Semi-supervised learning using the test data  \n        - Composing two models into an encoder-decoder model \n        - Use of translation focused pretrained models (e.g., mBart)\n\nWe'll release our code later this week.",
    "1590612": "日本語訳\nTL;DR\nアンサンブル、アンサンブル、アンサンブル\n疑似ラベリング\n多言語モデルでブートストラップし、単一言語モデルで洗練する\n2018 Toxic Commentsコンテストで共同優勝して以来、Kaggleに参加していません。 当時のNLP分類の最先端は、文脈に依存しない単語の埋め込み（FastTextなど）でした。トランスフォーマーの世界で上手く見せられるかどうか興味がありました。 チームメイトの@ rafiko1とうまくやってくれて、うれしい驚きです。\n\nOur public LB milestones\nベースラインXLM-Robertaモデル（公開LB：0.93XX- 3月31日）\nXLM-Rモデルの平均アンサンブル（0.942X- 4月3日）\n単一言語のTransformerモデルとのブレンド（0.9510- 4月23日）\nチームの合併-私たちの個々の最高のsubの加重平均（0.9529- 4月28日）\n単一言語のFastText分類子とのブレンド（0.9544- 5月5日）\n後処理およびその他。 最適化（0.9556- 6月13日）\n4月末までにFINALパブリックLBで1位になりました。 言い換えれば、私たちは競争が終わるのを2ヶ月待っていました:)\n\nCV strategy\n最初はk-foldCVとバリデーションセットを組み合わせてホールドアウトとして使用しましたが、テスト予測を改良し、トレーニングに疑似ラベル+バリデーションセットを使用すると、検証メトリックがノイズになり、主に一般の人々に依存するようになりました。 LBスコア。\n\nInsights\nTransformerトレーニングの変動を緩和するためのアンサンブル\nTransformerモデルのパフォーマンスは、初期化とデータの順序によって大きく影響を受けることに注意してください（https://arxiv.org/pdf/2002.06305.pdf、https://www.aclweb.org/anthology/2020.trac-1.9。 pdf）。 これを軽減するために、モデルのアンサンブルとバギングを強調しました。 これには、一時的な自己アンサンブルが含まれていました。 公開テストセットを前提として、反復ブレンディングアプローチを採用し、以前の最良の提出と現在のモデルの予測の加重平均を使用して、提出全体のテストセットの予測を改良しました。 単純な平均から始めて、前回の最高の提出物の重みを徐々に増やしていきました。 トレーニングデータには、モデルの実行ごとに2018年の有毒なコメントの翻訳のサブサンプルを主に使用しました。\n\nPseudo-labels (PL)\nテストセットの予測をトレーニングデータとして使用すると、パフォーマンスの向上が見られました。これは、モデルがテストセットの分布を学習するのに役立つという直感です。 すべてのテストセット予測をソフトラベルとして使用すると、他のバージョンの疑似ラベル（ハードラベル、信頼度しきい値PLなど）よりもうまく機能しました。 競争の終わりに向かって、PLをアップサンプリングしたときに、LBでマイナーではあるが重要なブーストを発見しました。\n\nMultilingual XLM-Roberta models\nほとんどのチームと同様に、トレーニングデータとして6つのテストセット言語での2018データセットの翻訳を組み込んだバニラXLM-Rモデルから始めました。 アダムオプティマイザーとバイナリクロスエントロピー損失関数を使用して、最後のレイヤーのCLSトークンにバニラ分類ヘッドを使用し、低い学習率でモデル全体を微調整しました。 Transformerモデルには、バイナリ予測を行うという比較的単純なタスクに数億のトレーニング可能な重みが設定されているため、デフォルトのアーキテクチャでは、モデルを通じて適切な信号を取得するために多くの調整が必要であるとは考えていませんでした。 その結果、ハイパーパラメータの最適化、アーキテクチャの調整、または前処理にあまり時間をかけませんでした。\n\nForeign language monolingual Transformer models\n@ rafiko1と私は、チームを結成する前に、このテクニックを独自に見つけました。 私たちは両方ともMultiFiTペーパー（https://arxiv.org/pdf/1909.04761.pdf）に触発されました-具体的には：\n\n\nテストセット言語にHuggingFaceの事前トレーニング済み外国語単一言語Transformerモデルを使用すると、パフォーマンスが劇的に向上することがわかりました（たとえば、フランス語のサンプルにはCamembert、ロシア語にはRubert、トルコ語にはBerTurk、スペイン語にはBETOなど）。\n\n2018 Toxic Commentsの翻訳とその特定の言語のサンプルの疑似ラベル（最初はXLM-R多言語モデルから）を組み合わせ、対応する単一言語モデルをトレーニングし、同じサンプルを予測することで、6つの言語のそれぞれのモデルを微調整しました次に、それをすべての予測の「メインブランチ」とブレンドし直します。トレーニングモデルA->トレーニングモデルB->トレーニングモデルAなどが継続的なパフォーマンスの向上につながるという点で相乗効果がありました。\n\nモデルの実行ごとに、過剰適合を防ぐために、事前にトレーニングされたモデルからウェイトの初期化をリロードします。言い換えれば、私たちが見た継続的な改善は、トレーニングデータとしてモデルに提供していた疑似ラベルの改良によって推進されていました。\n\n特定の単一言語モデルでは、その言語のテストセットサンプルのみを予測するのが最も効果的でした。他の言語のテストセットサンプルをモデルの言語に翻訳し、それらを予測すると、パフォーマンスが低下しました。\n\n事前にトレーニングされた外国語の単一言語FastTextモデルの微調整\nHuggingFaceモノリンガルモデルライブラリを使い果たした後、テストセット言語の事前トレーニング済み埋め込みを使用して、2018年の有毒コメントごとにモノリンガルFastText双方向GRUモデルをトレーニングし、テストセット予測の改良を続けました（メインと組み合わせると重みは小さくなりますが）予測のブランチ）、小さいながらも意味のあるパフォーマンスの向上（0.9536から0.9544）が見られました。\n\nPost-processing\n多数の提出があったため、@ rafiko1は、提出の履歴を利用してテストセットの予測を微調整するという斬新なアイデアを思いつきました。 提出が成功するまで、各サンプルの予測のデルタを追跡し、それらを平均して、同じ方向に予測を微調整しました。 パフォーマンスはわずかですが大幅に向上しました（〜0.0005）\n\nこの後処理手法に過剰適合するリスクが懸念されたため、最後の2つの提出では、この後処理を組み込んだものと組み込んでいないものを選択しました。 プライベートLBで非常にうまく機能するようになりました。\n\nMisc -\nトレーニングのセットアップ\n@ rafiko1と私は別々のコードベースを保持していました。 彼は、@ xhluluと@shonenkovによってパブリックカーネルから派生したTensorflowコードを使用してKaggleTPUインスタンスをトレーニングし、私はゼロからのPyTorchコードを使用して独自のハードウェア（デュアルRTX Titans）をトレーニングしました。 特大の合併アンサンブルブースト（0.9510-> 0.9529）の一部は、トレーニングのこの多様性に起因すると思います。\n\nWhat didn’t work\n現代のNLP分類手法のチェックリストを確認しましたが、おそらく単純な問題の目的のために、それらのほとんどは機能しませんでした。これがうまくいかなかったものの選択です：-\n-ドキュメントレベルのエンベッダー（例：レーザー）\n-タスクデータを使用したTransformerモデルのさらなるMLM事前トレーニング\n-代替アンサンブルメカニズム（ランク平均、確率的重み平均）\n-代替損失関数（例：焦点損失、ヒストグラム損失）\n-分類ヘッドの代替プーリングメカニズム（たとえば、トークン間での最大プールCNN、複数の非表示レイヤーの使用など）\n-FastText以外の事前トレーニング済みの埋め込み（例：フレア、グローブ、bpem）\n-Transformerモデルのフリーズ微調整\n-正則化（例：マルチサンプルドロップアウト、入力ミックスアップ、多様体ミックスアップ、センテンスピースドロップアウト）\n-データ拡張としての逆翻訳\n-列車/テスト時間の拡張としての英語の翻訳\n-2018年のデータにラベルを付け直すための自己蒸留\n-FGMを使用して埋め込み層を混乱させることによる敵対的訓練\n-マルチタスク学習\n-疑似ラベルの温度スケーリング\n-テストデータを使用した半教師あり学習\n-2つのモデルをエンコーダー-デコーダーモデルに構成する\n-翻訳に焦点を当てた事前トレーニング済みモデルの使用（例：mBart）\n\nWe'll release our code later this week.",
    "897946": "Congrats on the the thorough solution and on the win.  Glad to see you back on Kaggle, I remember when we fought in WTF competition.  Yours was so simple and effective!  Same here for effectiveness.",
    "963336": "@leecming @rafiko1 \n\nCongrats for being the first place. Your solution is really nice and inspired me a lot. I still have a few questions.\n\n&gt; When training a foreign monolingual foreign language model,  you only use samples in that specific language from the test/val dataset + translations from the orig 2018 training.\n\n* Are you using hard labels for training/val set and soft labels for test set? \n* Are you using two seperate loss functions during training as mentioned in the LASER paper?  Or you just used a single loss?\n* Are losses calculated from hard-target and soft-target weighted differently?\n\n&gt; It was synergistic in that training model A -&gt; training model B -&gt; training model A etc. lead to continual performance improvements.\n\n* What are models A and B here? Are they both monolingual models or one of them is XLM-R? If one of them is XLM-R, how do you improve the performance of it continuously?  \n",
    "903414": "Excellent write-up!  Easy to read and sounds like the language specific modeling was a unique approach that might have really helped separate you from the pack.  Well done! 👍 💯 🎉 ",
    "900594": "Thanks for sharing the methodology, a very well-deserved win, can't wait for the code release. ",
    "899154": "Congrats ! Thank you for the amazing overview ! i have learnt a lot from you ..  ",
    "899037": "Congratulations!  I have some questions regarding the monolingual training set up. \n&gt;We finetuned models for each of the 6 languages - by combining translations of the 2018 Toxic Comments together with pseudo-labels for samples in that specific language\n\nI assume you are using soft labels for the pseudo-labels samples. How do you set up the validation strategy here, particularly for those languages not in the validation set? \n",
    "898240": "Congratulations @leecming, @rafiko1 for the win and thank you so much for sharing your well thought out solution. A well deserved win.\n\n&gt; We began with a simple average, and gradually increased the weight of the previous best submission.\n\nI go out of my way to avoid doing this hence I am surprised the above did not overfit to the public LB. Why do you think that is?\n\n&gt; we trained monolingual FastText Bidirectional GRU models a la 2018 Toxic Comments, using pretrained embeddings for the test-set languages, ... and saw a small but meaningful performance boost (0.9536 to 0.9544).\n\nI was surprised at how my low scoring pooled GRU model contributed to my ensemble. I used it more for diversity and just did not think about using it the way you did. Bravo!",
    "898205": "Congratulations ! ",
    "897957": "Congrats! And thanks for sharing the solution.\n\nI wonder if you also used samples from jigsaw-unintended-bias dataset, in particular, the positive examples?",
    "897854": "&gt; You’ll note we hit 1st place on the FINAL public LB by end April. Put another way, we were waiting for 2 months for the competition to end :)\n\nI can tell, It was really boring for you guys 😄 \n\nAnyway, congratulation. 🎉 ",
    "897632": "Congrats! you guys have been doing so great and thanks for the detailed summary of your amazing solution, learned a lot:) I got one follow-up, in the post-processing part, what do you mean by 'delta successful submissions'? is it like using high-score subs and calculate the diff between each pair? it would be nice if you can release more details, thanks!",
    "897623": "Congrats! Amazing performance, quite a gap with second place. Thanks for sharing details of your solution.\n\nThe post-processing idea is quite smart and significant as well, good tip for future competitions.\n\nIn the \"synergistic\" improvement between cross-lingual and monolingual models, how many steps to convergence typically for how much improvement?",
    "897614": "Congratulations !   Your solution is very strong.  Thanks for sharing. ",
    "897611": "Congratulations!",
    "897589": "Amazing work. I knew there was some value to the monolingual models but couldnt think of a way to integrate them so the distributions of all of the predictions were totally misaligned. I even trained camembert models and rubert, but never got them to play nicely with everything else we had done. Very elegant solution. Almost like pseudolabeling, right?",
    "898255": "Hi, Congratulations !\n\nThanks you for sharing some tips about your works.\nI have few questions, : \n- Have you used only large Transformers or also base or small one ?\n- \"- Semi-supervised learning using the test data \" what semi supervised learning did not work exactly ? as you used pseudo labelling which I think can be considered as semi supervised. Are you referring about using hard labels instead of soft labels or is it some others approachs ?\n- about FastText, have you take the embedding and add few GRU layers then pretrained on english data ? And then use a regular training on translated + PL data ? OR did you create one per language ?",
    "897651": "Congrats on the 1st place, very worthy. Your team solution is really powerful and impressive!",
    "897575": "Congratulation on the 1st place, very good overview, it is interesting to see that you achieved top place with a simple model architecture (just the CLS token), as you said such a big model might already have enough parameters, I did some experimentations on that but also did not saw significant improvements.",
    "1480940": "Congratulations on the first place!!  I have a question regarding the low F1, Precision, and Recall score. How to deal with this low score since the class is highly imbalanced.",
    "975061": "Hello, @leecming, you trained monolingual FastText Bidirectional GRU models in six test languages then concatenate or only in english(translate all data into english) ?",
    "899332": "Hi first of all Congrats on winning the competition and for an amazing approach. Can you please point me to the papers of above monolingual models. I was able to find the below.\n\n[1] : Camembert (https://arxiv.org/pdf/1911.03894.pdf?fbclid=IwAR03h_fUIfSqPlaqTkMew4mybw-LjBrh8dZ3lpTgQeldLdsZFfr54E7nzk4) \n\n[2] : Beto (https://arxiv.org/abs/1904.09077?fbclid=IwAR1mieFlVqpl0iHZq2DAA0dSJ7uDCBpLoCeX_p6CIY69n7RTQQrFuDFNlm4)\n\nI wasn't able to find BerTurk and Rubert. Any help would be appreciated. Also the corresponding usage of such models. Thanks.",
    "898701": "Hi, I got another question.\n\nFor fine-tuning `a foreign language monolingual Transformer models`, what is the training dataset?\nDid you use only the samples in that specific language from test dataset? (If so, the training samples are quite few).\nOr you also included the translated samples in the target lang from the original training/validation dataset?\n",
    "897647": "Is your Private LB score also has similar improvement as your Public LB milestones? Or any of the techniques (in your milestones) provide much lesser contribution after taking Private LB score into consideration? ",
    "897574": "",
    "973735": "Thanks for sharing."
  }
}