{
  "id": 161100,
  "title": "10th Place Solution",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/up-10th-place-solution",
  "author_name": "",
  "post_date": "2020-06-28T04:05:44.230Z",
  "votes": 48,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Thanks to my wonderful teamates <a href=\"/naivelamb\">@naivelamb</a> <a href=\"/wuyhbb\">@wuyhbb</a> <a href=\"/terenceliu4444\">@terenceliu4444</a> <a href=\"/hughshaoqz\">@hughshaoqz</a> , Thank you for helping me to get my 5th gold medal!</p>\n\n<p>Congratulations to all winners!</p>\n\n<h1>Summary</h1>\n\n<p>Here's main ideas that worked for us.</p>\n\n<ul>\n<li>Translated Data</li>\n<li>Pseudo Labelling</li>\n<li>Multi-Stage Training</li>\n<li>Freeze Embed Layer</li>\n<li>UDA (<a href=\"https://arxiv.org/pdf/1904.12848.pdf\">https://arxiv.org/pdf/1904.12848.pdf</a>)</li>\n<li>Mono Language Modeling</li>\n<li>Test Time Augmentation</li>\n</ul>\n\n<h1>Details</h1>\n\n<h3>Translated Data</h3>\n\n<p>We use translated data listed here for training and testing (for TTA)</p>\n\n<p><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159888\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159888</a></p>\n\n<p>Thank you guys for publishing this excellent datasets!</p>\n\n<h3>Pseudo Labelling</h3>\n\n<p>We trained some baseline models then blend them with some public submission files to get to LB <code>0.9481</code> then use it as Pseudo Label till the end.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta</a></li>\n<li><a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a></li>\n</ul>\n\n<p>When training on this test data with pseudo label, we found that use all test data by KL-Div loss with soft label gave us best performance.</p>\n\n<h3>Multi-Stage Training</h3>\n\n<p>We all know that fine-tuning the model on validation set after training on train set can boost the LB score, and we call it <code>2-stage training</code>. (<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>)\nBy adding more stages on our training pipeline we were able to further boost our LB score. Some of our pipeline is like:</p>\n\n<ul>\n<li>Pseudo Labelling (5epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)</li>\n<li>Train2 (1epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)</li>\n<li>Train1 (1epo) -&gt; Pseudo Labelling (3epo) -&gt; Valid (3epo)</li>\n<li>...</li>\n</ul>\n\n<p>(Train1: training data of the 1st jigsaw competition on 2018)\n(Train2: training data of the 2nd jigsaw competition on 2019)\n(Pseudo Labelling: test data of this competition)\n(Valid: validation data of this competition)</p>\n\n<p>Note that we usually don't train on the full dataset when using Train1 and Train2, but a subset of it.</p>\n\n<h3>Freeze Embed Layer</h3>\n\n<p>Freeze Embed Layer of transformers can save the GPU memory so that we were able to use bigger batchsize when training, while speeding up the training process.</p>\n\n<p>By Multi-Stage Training, Pseudo Labelling and Freeze Embed Layer we were able to get to LB <code>0.9487</code> without blending any public submission file.</p>\n\n<h3>Test Time Augmentation</h3>\n\n<p>We do the inference on all 6 languages, and the blending weight between the original language and the other 5 languages is 8:2\nWe got to LB <code>0.9492</code> by this TTA.</p>\n\n<h3>UDA</h3>\n\n<p>reference: <a href=\"https://arxiv.org/pdf/1904.12848.pdf\">https://arxiv.org/pdf/1904.12848.pdf</a></p>\n\n<p>Then we blend some models that used UDA while training to get to LB <code>0.9498</code>, by some further tuning we got to LB <code>0.9500</code>.</p>\n\n<h3>Mono Language Modeling</h3>\n\n<p>At the last week, we found that blend with models that only trained on single language data can also boost the CV score.\nBut since in validation data we only got <code>es</code>, <code>it</code>, <code>tr</code>, so we only do <code>Mono Language Modeling</code> on these 3 languages.</p>\n\n<p>By blending models that only trained on <code>es</code> or <code>it</code> or <code>tr</code>, we finally got to our best public LB <code>0.9504</code> and this is our best submission on private LB as well. </p>\n\n<hr>\n\n<h1>Thank you for reading!</h1>",
  "messages": [
    {
      "id": "898761",
      "postDate": "06/23/2020 18:07:54",
      "content": "<p>Thanks to my wonderful teamates <a href=\"/naivelamb\">@naivelamb</a> <a href=\"/wuyhbb\">@wuyhbb</a> <a href=\"/terenceliu4444\">@terenceliu4444</a> <a href=\"/hughshaoqz\">@hughshaoqz</a> , Thank you for helping me to get my 5th gold medal!</p>\n\n<p>Congratulations to all winners!</p>\n\n<h1>Summary</h1>\n\n<p>Here's main ideas that worked for us.</p>\n\n<ul>\n<li>Translated Data</li>\n<li>Pseudo Labelling</li>\n<li>Multi-Stage Training</li>\n<li>Freeze Embed Layer</li>\n<li>UDA (<a href=\"https://arxiv.org/pdf/1904.12848.pdf\">https://arxiv.org/pdf/1904.12848.pdf</a>)</li>\n<li>Mono Language Modeling</li>\n<li>Test Time Augmentation</li>\n</ul>\n\n<h1>Details</h1>\n\n<h3>Translated Data</h3>\n\n<p>We use translated data listed here for training and testing (for TTA)</p>\n\n<p><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159888\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159888</a></p>\n\n<p>Thank you guys for publishing this excellent datasets!</p>\n\n<h3>Pseudo Labelling</h3>\n\n<p>We trained some baseline models then blend them with some public submission files to get to LB <code>0.9481</code> then use it as Pseudo Label till the end.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta</a></li>\n<li><a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a></li>\n</ul>\n\n<p>When training on this test data with pseudo label, we found that use all test data by KL-Div loss with soft label gave us best performance.</p>\n\n<h3>Multi-Stage Training</h3>\n\n<p>We all know that fine-tuning the model on validation set after training on train set can boost the LB score, and we call it <code>2-stage training</code>. (<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>)\nBy adding more stages on our training pipeline we were able to further boost our LB score. Some of our pipeline is like:</p>\n\n<ul>\n<li>Pseudo Labelling (5epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)</li>\n<li>Train2 (1epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)</li>\n<li>Train1 (1epo) -&gt; Pseudo Labelling (3epo) -&gt; Valid (3epo)</li>\n<li>...</li>\n</ul>\n\n<p>(Train1: training data of the 1st jigsaw competition on 2018)\n(Train2: training data of the 2nd jigsaw competition on 2019)\n(Pseudo Labelling: test data of this competition)\n(Valid: validation data of this competition)</p>\n\n<p>Note that we usually don't train on the full dataset when using Train1 and Train2, but a subset of it.</p>\n\n<h3>Freeze Embed Layer</h3>\n\n<p>Freeze Embed Layer of transformers can save the GPU memory so that we were able to use bigger batchsize when training, while speeding up the training process.</p>\n\n<p>By Multi-Stage Training, Pseudo Labelling and Freeze Embed Layer we were able to get to LB <code>0.9487</code> without blending any public submission file.</p>\n\n<h3>Test Time Augmentation</h3>\n\n<p>We do the inference on all 6 languages, and the blending weight between the original language and the other 5 languages is 8:2\nWe got to LB <code>0.9492</code> by this TTA.</p>\n\n<h3>UDA</h3>\n\n<p>reference: <a href=\"https://arxiv.org/pdf/1904.12848.pdf\">https://arxiv.org/pdf/1904.12848.pdf</a></p>\n\n<p>Then we blend some models that used UDA while training to get to LB <code>0.9498</code>, by some further tuning we got to LB <code>0.9500</code>.</p>\n\n<h3>Mono Language Modeling</h3>\n\n<p>At the last week, we found that blend with models that only trained on single language data can also boost the CV score.\nBut since in validation data we only got <code>es</code>, <code>it</code>, <code>tr</code>, so we only do <code>Mono Language Modeling</code> on these 3 languages.</p>\n\n<p>By blending models that only trained on <code>es</code> or <code>it</code> or <code>tr</code>, we finally got to our best public LB <code>0.9504</code> and this is our best submission on private LB as well. </p>\n\n<hr>\n\n<h1>Thank you for reading!</h1>",
      "rawMarkdown": "Thanks to my wonderful teamates @naivelamb @wuyhbb @terenceliu4444 @hughshaoqz , Thank you for helping me to get my 5th gold medal!\n\nCongratulations to all winners!\n\n# Summary\n\nHere's main ideas that worked for us.\n\n* Translated Data\n* Pseudo Labelling\n* Multi-Stage Training\n* Freeze Embed Layer\n* UDA (https://arxiv.org/pdf/1904.12848.pdf)\n* Mono Language Modeling\n* Test Time Augmentation\n\n# Details\n\n### Translated Data\n\nWe use translated data listed here for training and testing (for TTA)\n\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159888\n\nThank you guys for publishing this excellent datasets!\n\n### Pseudo Labelling\n\nWe trained some baseline models then blend them with some public submission files to get to LB `0.9481` then use it as Pseudo Label till the end.\n\n* https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\n* https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\n\nWhen training on this test data with pseudo label, we found that use all test data by KL-Div loss with soft label gave us best performance.\n\n### Multi-Stage Training\n\nWe all know that fine-tuning the model on validation set after training on train set can boost the LB score, and we call it `2-stage training`. (https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta)\nBy adding more stages on our training pipeline we were able to further boost our LB score. Some of our pipeline is like:\n\n* Pseudo Labelling (5epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)\n* Train2 (1epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)\n* Train1 (1epo) -&gt; Pseudo Labelling (3epo) -&gt; Valid (3epo)\n* ...\n\n(Train1: training data of the 1st jigsaw competition on 2018)\n(Train2: training data of the 2nd jigsaw competition on 2019)\n(Pseudo Labelling: test data of this competition)\n(Valid: validation data of this competition)\n\nNote that we usually don't train on the full dataset when using Train1 and Train2, but a subset of it.\n\n\n\n### Freeze Embed Layer\n\nFreeze Embed Layer of transformers can save the GPU memory so that we were able to use bigger batchsize when training, while speeding up the training process.\n\nBy Multi-Stage Training, Pseudo Labelling and Freeze Embed Layer we were able to get to LB `0.9487` without blending any public submission file.\n\n\n### Test Time Augmentation\n\nWe do the inference on all 6 languages, and the blending weight between the original language and the other 5 languages is 8:2\nWe got to LB `0.9492` by this TTA.\n\n\n### UDA\n\nreference: https://arxiv.org/pdf/1904.12848.pdf\n\nThen we blend some models that used UDA while training to get to LB `0.9498`, by some further tuning we got to LB `0.9500`.\n\n\n### Mono Language Modeling\n\nAt the last week, we found that blend with models that only trained on single language data can also boost the CV score.\nBut since in validation data we only got `es`, `it`, `tr`, so we only do `Mono Language Modeling` on these 3 languages.\n\nBy blending models that only trained on `es` or `it` or `tr`, we finally got to our best public LB `0.9504` and this is our best submission on private LB as well. \n\n\n---\n\n# Thank you for reading!",
      "votes": null
    },
    {
      "id": "898766",
      "postDate": "06/23/2020 18:16:07",
      "content": "<p>Thanks for wrapping up! I really enjoyed working with you and all other teammates! And congrats again for becoming a GrandMaster!</p>",
      "rawMarkdown": "Thanks for wrapping up! I really enjoyed working with you and all other teammates! And congrats again for becoming a GrandMaster!",
      "votes": null
    },
    {
      "id": "899148",
      "postDate": "06/24/2020 03:09:16",
      "content": "<p>Thanks for sharing such a nice set of insights with us all. Looking forward to learn those techniques myself too.. </p>",
      "rawMarkdown": "Thanks for sharing such a nice set of insights with us all. Looking forward to learn those techniques myself too..",
      "votes": null
    },
    {
      "id": "899172",
      "postDate": "06/24/2020 03:56:04",
      "content": "<p>Thanks for sharing </p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "899335",
      "postDate": "06/24/2020 07:08:13",
      "content": "<p>Congrats！</p>",
      "rawMarkdown": "Congrats！",
      "votes": null
    },
    {
      "id": "899496",
      "postDate": "06/24/2020 09:18:42",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "899662",
      "postDate": "06/24/2020 11:14:40",
      "content": "<p>Congratulation.</p>",
      "rawMarkdown": "Congratulation.",
      "votes": null
    },
    {
      "id": "899683",
      "postDate": "06/24/2020 11:40:51",
      "content": "<p>Congrats!   thanks for sharing the UDA paper, it looks interesting.</p>",
      "rawMarkdown": "Congrats!   thanks for sharing the UDA paper, it looks interesting.",
      "votes": null
    },
    {
      "id": "900099",
      "postDate": "06/24/2020 16:10:11",
      "content": "<p>Thanks for wrapping up. Very nice working with you and all other teammates. Here is a snippet of UDA we implemented in tf.keras, where transformer is a Huggingface Pretrained model (XLM-R in our case).\n```\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\")\n    orig_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"orig_word_ids\")\n    aug_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"aug_word_ids\")</p>\n\n<pre><code>sequence_output = transformer(input_word_ids)[0]\ncls_token = sequence_output[:, 0, :]\nout = Dense(1,  activation='sigmoid', name=\"out\")(cls_token)\n\norig_sequence_output = transformer(orig_word_ids)[0]\norig_sequence_output = tf.keras.layers.Lambda(lambda x: K.stop_gradient(x))(orig_sequence_output)\norig_cls_token = orig_sequence_output[:, 0, :]\norig_out = Dense(1,  activation='sigmoid', name='orig_out')(orig_cls_token)\n\naug_sequence_output = transformer(aug_word_ids)[0]\naug_cls_token = aug_sequence_output[:, 0, :]\naug_out = Dense(1,  activation='sigmoid', name=\"aug_out\")(aug_cls_token)\nmodel = Model(inputs=[input_word_ids, orig_word_ids, aug_word_ids],\n              outputs=out)\nunsup_loss = tf.keras.losses.KLDivergence(reduction=tf.keras.losses.Reduction.SUM)(orig_out, aug_out) * (1./BATCH_SIZE)\nmodel.add_loss(unsup_loss)\nmodel.compile(Adam(lr=1e-5), loss='binary_crossentropy',\n              metrics=['accuracy', tf.keras.metrics.AUC()])\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Thanks for wrapping up. Very nice working with you and all other teammates. Here is a snippet of UDA we implemented in tf.keras, where transformer is a Huggingface Pretrained model (XLM-R in our case).\n```\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\")\n    orig_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"orig_word_ids\")\n    aug_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"aug_word_ids\")\n\n    sequence_output = transformer(input_word_ids)[0]\n    cls_token = sequence_output[:, 0, :]\n    out = Dense(1,  activation='sigmoid', name=\"out\")(cls_token)\n\n    orig_sequence_output = transformer(orig_word_ids)[0]\n    orig_sequence_output = tf.keras.layers.Lambda(lambda x: K.stop_gradient(x))(orig_sequence_output)\n    orig_cls_token = orig_sequence_output[:, 0, :]\n    orig_out = Dense(1,  activation='sigmoid', name='orig_out')(orig_cls_token)\n\n    aug_sequence_output = transformer(aug_word_ids)[0]\n    aug_cls_token = aug_sequence_output[:, 0, :]\n    aug_out = Dense(1,  activation='sigmoid', name=\"aug_out\")(aug_cls_token)\n    model = Model(inputs=[input_word_ids, orig_word_ids, aug_word_ids],\n                  outputs=out)\n    unsup_loss = tf.keras.losses.KLDivergence(reduction=tf.keras.losses.Reduction.SUM)(orig_out, aug_out) * (1./BATCH_SIZE)\n    model.add_loss(unsup_loss)\n    model.compile(Adam(lr=1e-5), loss='binary_crossentropy',\n                  metrics=['accuracy', tf.keras.metrics.AUC()])\n```",
      "votes": null
    },
    {
      "id": "900217",
      "postDate": "06/24/2020 17:30:27",
      "content": "<p>Congratulations! Interesting approach and thanks for sharing the UDA paper and code. Just one query here. What approach did your team tried with aug_word that improved your UDA? Is \"Word replacing with TF-IDF\" approach used here? Thanks. Dr.</p>",
      "rawMarkdown": "Congratulations! Interesting approach and thanks for sharing the UDA paper and code. Just one query here. What approach did your team tried with aug_word that improved your UDA? Is \"Word replacing with TF-IDF\" approach used here? Thanks. Dr.",
      "votes": null
    },
    {
      "id": "900311",
      "postDate": "06/24/2020 18:24:47",
      "content": "<p>Congratulations as well! For XLM-R, we used the translated and round-trip (source-en-source) translated text as augmented text. For mono lingual model, we use round-trip translated text. We did not try \"Word replacing with TF-IDF\".</p>",
      "rawMarkdown": "Congratulations as well! For XLM-R, we used the translated and round-trip (source-en-source) translated text as augmented text. For mono lingual model, we use round-trip translated text. We did not try \"Word replacing with TF-IDF\".",
      "votes": null
    },
    {
      "id": "900546",
      "postDate": "06/24/2020 22:06:55",
      "content": "<p>Congratulations! Great to see it turn out so well for you :) 👍 🥇</p>\n\n<p>I have recently tried my hands into Audio Processing and Classification for the competition <a href=\"https://www.kaggle.com/c/birdsong-recognition\">Cornell Birdcall Identification</a>. Do check out <a href=\"https://www.kaggle.com/navinmundhra/birdcall-extensive-audio-representn-extractn\">my extensive work on Audio feature extraction in my notebook</a>. I hope you like it!</p>\n\n<p>Also, check out, review, and upvote <a href=\"https://www.kaggle.com/navinmundhra/notebooks\">my notebooks</a>. They involve a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you. </p>",
      "rawMarkdown": "Congratulations! Great to see it turn out so well for you :) 👍 🥇\n\nI have recently tried my hands into Audio Processing and Classification for the competition [Cornell Birdcall Identification](https://www.kaggle.com/c/birdsong-recognition). Do check out [my extensive work on Audio feature extraction in my notebook](https://www.kaggle.com/navinmundhra/birdcall-extensive-audio-representn-extractn). I hope you like it!\n\nAlso, check out, review, and upvote [my notebooks](https://www.kaggle.com/navinmundhra/notebooks). They involve a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you.",
      "votes": null
    },
    {
      "id": "903883",
      "postDate": "06/27/2020 07:28:52",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "904347",
      "postDate": "06/27/2020 15:01:36",
      "content": "<p>Congratulations to all!</p>\n\n<p>Thanks for sharing</p>",
      "rawMarkdown": "Congratulations to all!\n\nThanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 898766,
      "author_name": "hughshaoqz",
      "author_url": "",
      "post_date": "06/23/2020 18:16:07",
      "content": "<p>Thanks for wrapping up! I really enjoyed working with you and all other teammates! And congrats again for becoming a GrandMaster!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899148,
      "author_name": "redwankarimsony",
      "author_url": "",
      "post_date": "06/24/2020 03:09:16",
      "content": "<p>Thanks for sharing such a nice set of insights with us all. Looking forward to learn those techniques myself too.. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899172,
      "author_name": "muralidhar123",
      "author_url": "",
      "post_date": "06/24/2020 03:56:04",
      "content": "<p>Thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899335,
      "author_name": "machinelp",
      "author_url": "",
      "post_date": "06/24/2020 07:08:13",
      "content": "<p>Congrats！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899496,
      "author_name": "gauravdahiya",
      "author_url": "",
      "post_date": "06/24/2020 09:18:42",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899662,
      "author_name": "koshirosato",
      "author_url": "",
      "post_date": "06/24/2020 11:14:40",
      "content": "<p>Congratulation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899683,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/24/2020 11:40:51",
      "content": "<p>Congrats!   thanks for sharing the UDA paper, it looks interesting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 900099,
      "author_name": "terenceliu4444",
      "author_url": "",
      "post_date": "06/24/2020 16:10:11",
      "content": "<p>Thanks for wrapping up. Very nice working with you and all other teammates. Here is a snippet of UDA we implemented in tf.keras, where transformer is a Huggingface Pretrained model (XLM-R in our case).\n```\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\")\n    orig_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"orig_word_ids\")\n    aug_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"aug_word_ids\")</p>\n\n<pre><code>sequence_output = transformer(input_word_ids)[0]\ncls_token = sequence_output[:, 0, :]\nout = Dense(1,  activation='sigmoid', name=\"out\")(cls_token)\n\norig_sequence_output = transformer(orig_word_ids)[0]\norig_sequence_output = tf.keras.layers.Lambda(lambda x: K.stop_gradient(x))(orig_sequence_output)\norig_cls_token = orig_sequence_output[:, 0, :]\norig_out = Dense(1,  activation='sigmoid', name='orig_out')(orig_cls_token)\n\naug_sequence_output = transformer(aug_word_ids)[0]\naug_cls_token = aug_sequence_output[:, 0, :]\naug_out = Dense(1,  activation='sigmoid', name=\"aug_out\")(aug_cls_token)\nmodel = Model(inputs=[input_word_ids, orig_word_ids, aug_word_ids],\n              outputs=out)\nunsup_loss = tf.keras.losses.KLDivergence(reduction=tf.keras.losses.Reduction.SUM)(orig_out, aug_out) * (1./BATCH_SIZE)\nmodel.add_loss(unsup_loss)\nmodel.compile(Adam(lr=1e-5), loss='binary_crossentropy',\n              metrics=['accuracy', tf.keras.metrics.AUC()])\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": [
        {
          "id": 900217,
          "author_name": "drpatrickchan",
          "author_url": "",
          "post_date": "06/24/2020 17:30:27",
          "content": "<p>Congratulations! Interesting approach and thanks for sharing the UDA paper and code. Just one query here. What approach did your team tried with aug_word that improved your UDA? Is \"Word replacing with TF-IDF\" approach used here? Thanks. Dr.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 900311,
          "author_name": "terenceliu4444",
          "author_url": "",
          "post_date": "06/24/2020 18:24:47",
          "content": "<p>Congratulations as well! For XLM-R, we used the translated and round-trip (source-en-source) translated text as augmented text. For mono lingual model, we use round-trip translated text. We did not try \"Word replacing with TF-IDF\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 900546,
      "author_name": "navinmundhra",
      "author_url": "",
      "post_date": "06/24/2020 22:06:55",
      "content": "<p>Congratulations! Great to see it turn out so well for you :) 👍 🥇</p>\n\n<p>I have recently tried my hands into Audio Processing and Classification for the competition <a href=\"https://www.kaggle.com/c/birdsong-recognition\">Cornell Birdcall Identification</a>. Do check out <a href=\"https://www.kaggle.com/navinmundhra/birdcall-extensive-audio-representn-extractn\">my extensive work on Audio feature extraction in my notebook</a>. I hope you like it!</p>\n\n<p>Also, check out, review, and upvote <a href=\"https://www.kaggle.com/navinmundhra/notebooks\">my notebooks</a>. They involve a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903883,
      "author_name": "mnk812",
      "author_url": "",
      "post_date": "06/27/2020 07:28:52",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 904347,
      "author_name": "mdselimreza",
      "author_url": "",
      "post_date": "06/27/2020 15:01:36",
      "content": "<p>Congratulations to all!</p>\n\n<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "898761": "Thanks to my wonderful teamates @naivelamb @wuyhbb @terenceliu4444 @hughshaoqz , Thank you for helping me to get my 5th gold medal!\n\nCongratulations to all winners!\n\n# Summary\n\nHere's main ideas that worked for us.\n\n* Translated Data\n* Pseudo Labelling\n* Multi-Stage Training\n* Freeze Embed Layer\n* UDA (https://arxiv.org/pdf/1904.12848.pdf)\n* Mono Language Modeling\n* Test Time Augmentation\n\n# Details\n\n### Translated Data\n\nWe use translated data listed here for training and testing (for TTA)\n\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/159888\n\nThank you guys for publishing this excellent datasets!\n\n### Pseudo Labelling\n\nWe trained some baseline models then blend them with some public submission files to get to LB `0.9481` then use it as Pseudo Label till the end.\n\n* https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\n* https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\n\nWhen training on this test data with pseudo label, we found that use all test data by KL-Div loss with soft label gave us best performance.\n\n### Multi-Stage Training\n\nWe all know that fine-tuning the model on validation set after training on train set can boost the LB score, and we call it `2-stage training`. (https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta)\nBy adding more stages on our training pipeline we were able to further boost our LB score. Some of our pipeline is like:\n\n* Pseudo Labelling (5epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)\n* Train2 (1epo) -&gt; Train1 (1epo) -&gt; Valid (3epo)\n* Train1 (1epo) -&gt; Pseudo Labelling (3epo) -&gt; Valid (3epo)\n* ...\n\n(Train1: training data of the 1st jigsaw competition on 2018)\n(Train2: training data of the 2nd jigsaw competition on 2019)\n(Pseudo Labelling: test data of this competition)\n(Valid: validation data of this competition)\n\nNote that we usually don't train on the full dataset when using Train1 and Train2, but a subset of it.\n\n\n\n### Freeze Embed Layer\n\nFreeze Embed Layer of transformers can save the GPU memory so that we were able to use bigger batchsize when training, while speeding up the training process.\n\nBy Multi-Stage Training, Pseudo Labelling and Freeze Embed Layer we were able to get to LB `0.9487` without blending any public submission file.\n\n\n### Test Time Augmentation\n\nWe do the inference on all 6 languages, and the blending weight between the original language and the other 5 languages is 8:2\nWe got to LB `0.9492` by this TTA.\n\n\n### UDA\n\nreference: https://arxiv.org/pdf/1904.12848.pdf\n\nThen we blend some models that used UDA while training to get to LB `0.9498`, by some further tuning we got to LB `0.9500`.\n\n\n### Mono Language Modeling\n\nAt the last week, we found that blend with models that only trained on single language data can also boost the CV score.\nBut since in validation data we only got `es`, `it`, `tr`, so we only do `Mono Language Modeling` on these 3 languages.\n\nBy blending models that only trained on `es` or `it` or `tr`, we finally got to our best public LB `0.9504` and this is our best submission on private LB as well. \n\n\n---\n\n# Thank you for reading!",
    "898766": "Thanks for wrapping up! I really enjoyed working with you and all other teammates! And congrats again for becoming a GrandMaster!",
    "899148": "Thanks for sharing such a nice set of insights with us all. Looking forward to learn those techniques myself too..",
    "899172": "Thanks for sharing",
    "899335": "Congrats！",
    "899496": "Thanks for sharing!",
    "899662": "Congratulation.",
    "899683": "Congrats!   thanks for sharing the UDA paper, it looks interesting.",
    "900099": "Thanks for wrapping up. Very nice working with you and all other teammates. Here is a snippet of UDA we implemented in tf.keras, where transformer is a Huggingface Pretrained model (XLM-R in our case).\n```\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\")\n    orig_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"orig_word_ids\")\n    aug_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"aug_word_ids\")\n\n    sequence_output = transformer(input_word_ids)[0]\n    cls_token = sequence_output[:, 0, :]\n    out = Dense(1,  activation='sigmoid', name=\"out\")(cls_token)\n\n    orig_sequence_output = transformer(orig_word_ids)[0]\n    orig_sequence_output = tf.keras.layers.Lambda(lambda x: K.stop_gradient(x))(orig_sequence_output)\n    orig_cls_token = orig_sequence_output[:, 0, :]\n    orig_out = Dense(1,  activation='sigmoid', name='orig_out')(orig_cls_token)\n\n    aug_sequence_output = transformer(aug_word_ids)[0]\n    aug_cls_token = aug_sequence_output[:, 0, :]\n    aug_out = Dense(1,  activation='sigmoid', name=\"aug_out\")(aug_cls_token)\n    model = Model(inputs=[input_word_ids, orig_word_ids, aug_word_ids],\n                  outputs=out)\n    unsup_loss = tf.keras.losses.KLDivergence(reduction=tf.keras.losses.Reduction.SUM)(orig_out, aug_out) * (1./BATCH_SIZE)\n    model.add_loss(unsup_loss)\n    model.compile(Adam(lr=1e-5), loss='binary_crossentropy',\n                  metrics=['accuracy', tf.keras.metrics.AUC()])\n```",
    "900217": "Congratulations! Interesting approach and thanks for sharing the UDA paper and code. Just one query here. What approach did your team tried with aug_word that improved your UDA? Is \"Word replacing with TF-IDF\" approach used here? Thanks. Dr.",
    "900311": "Congratulations as well! For XLM-R, we used the translated and round-trip (source-en-source) translated text as augmented text. For mono lingual model, we use round-trip translated text. We did not try \"Word replacing with TF-IDF\".",
    "900546": "Congratulations! Great to see it turn out so well for you :) 👍 🥇\n\nI have recently tried my hands into Audio Processing and Classification for the competition [Cornell Birdcall Identification](https://www.kaggle.com/c/birdsong-recognition). Do check out [my extensive work on Audio feature extraction in my notebook](https://www.kaggle.com/navinmundhra/birdcall-extensive-audio-representn-extractn). I hope you like it!\n\nAlso, check out, review, and upvote [my notebooks](https://www.kaggle.com/navinmundhra/notebooks). They involve a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you.",
    "903883": "Congratulations!",
    "904347": "Congratulations to all!\n\nThanks for sharing"
  },
  "source": "meta"
}