{
  "id": 161095,
  "title": "6th place solution",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/toxu-nzholmes-6th-place-solution",
  "author_name": "",
  "post_date": "2020-06-23T17:58:54.289965Z",
  "votes": 22,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to Jigsaw to hold such amazing competition! And also thanks to my teammate <a href=\"/tonyxu\">@tonyxu</a> for diligence and brilliant ideas. This is our second jigsaw competition and also the first time to join such multilingual nlp competition. </p>\n\n<h2>Brief summary</h2>\n\n<p>Before team merger, I mainly focused on how to train a single model with decent performance while my teammate worked on training diverse models, especially training on single language data and combine them together. Before merger, we both achieved 0.9487 on public leaderboard. As you can see, we worked in relatively different directions and achieved 0.9499 right after team merger. After team merger, we focused on training more diverse models for diversity, which gave us great boost but also caused us to ignore the post processing trick most top winners found.</p>\n\n<h2>Detailed summary</h2>\n\n<h3>Data sampling</h3>\n\n<p>We used both english training data and its translation shared publicly. Since there were tremendous amount of training data available, the main issue was not too little but too much. I mainly adopted hard sampling technique to sample the training data. </p>\n\n<p>First, I included all positives and randomly sampled negatives to train a XLM-Roberta Large model and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment. </p>\n\n<p>Then for english data, we set a threshold for negatives while we took all positives due to the data unbalance. The data with gap between the predicted probability and label greater than the threshold were all taken and the data with gap smaller than that were sampled. The same logic for translated data except the threshold setting. A single threshold may cause issues because translation led to information loss. So I set a value range, say [0.3, 0.7] and any data in this range were all taken. I didn't convert soft label to hard 0/1 label because it worked better in \nthe last jigsaw competition. The number of training data is ~934k. XLM-Roberta Large trained on this data could reach 0.9469 on public lb and 0.9458 on private lb. </p>\n\n<h3>Multilingual Models</h3>\n\n<p>Our multilingual models mainly consist of <code>XLM-Roberta-Large</code>, <code>XLM-Roberta-base</code>, <code>Bert-multilingual-base-uncased</code>. My teammate also did language model finetuning on <code>jigsaw-unintended-bias</code> data. We don't know the exact number for each of them used in the final ensemble but the best public scores are 0.9469, 0.9364 and 0.9325 respectively. </p>\n\n<p>In terms of model structure, I used a vanilla classification head on the CLS token of the second to last layer or the concatenation of last four layers while my teammate took the <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/103280\">architecture of winning model</a> in the last jigsaw competition.</p>\n\n<h3>Monlingual Models</h3>\n\n<p>Just one week before the deadline, we found the monolingual language models in <a href=\"https://huggingface.co/models\">huggingface communities</a>. We trained the models listed below on corresponding translated language, which gave us boost from 0.9507 to 0.9512 on public lb. </p>\n\n<p>&gt; * es:\n        - <a href=\"https://huggingface.co/dccuchile/bert-base-spanish-wwm-uncased#\">https://huggingface.co/dccuchile/bert-base-spanish-wwm-uncased#</a>\n        - <a href=\"https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased#\">https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased#</a>\n        - <a href=\"https://github.com/dccuchile/beto\">https://github.com/dccuchile/beto</a> <br>\n* fr:\n        - <a href=\"https://huggingface.co/camembert/camembert-large\">https://huggingface.co/camembert/camembert-large</a>\n        - <a href=\"https://huggingface.co/camembert-base\">https://huggingface.co/camembert-base</a>\n        - <a href=\"https://huggingface.co/flaubert/flaubert-large-cased\">https://huggingface.co/flaubert/flaubert-large-cased</a>\n        - <a href=\"https://huggingface.co/flaubert/flaubert-base-uncased\">https://huggingface.co/flaubert/flaubert-base-uncased</a>\n        - <a href=\"https://huggingface.co/flaubert/flaubert-base-cased\">https://huggingface.co/flaubert/flaubert-base-cased</a>\n* tr:\n        - (128k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-128k-uncased\">https://huggingface.co/dbmdz/bert-base-turkish-128k-uncased</a>\n        - (128k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-128k-cased#\">https://huggingface.co/dbmdz/bert-base-turkish-128k-cased#</a>\n        - (32k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-uncased#\">https://huggingface.co/dbmdz/bert-base-turkish-uncased#</a>\n        - (32k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-cased#\">https://huggingface.co/dbmdz/bert-base-turkish-cased#</a> <br>\n* it\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-xxl-uncased#\">https://huggingface.co/dbmdz/bert-base-italian-xxl-uncased#</a>\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-xxl-cased#\">https://huggingface.co/dbmdz/bert-base-italian-xxl-cased#</a>\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-uncased#\">https://huggingface.co/dbmdz/bert-base-italian-uncased#</a>\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-cased#\">https://huggingface.co/dbmdz/bert-base-italian-cased#</a> <br>\n* pt\n        - <a href=\"https://huggingface.co/neuralmind/bert-large-portuguese-cased#\">https://huggingface.co/neuralmind/bert-large-portuguese-cased#</a>\n        - <a href=\"https://huggingface.co/neuralmind/bert-base-portuguese-cased#\">https://huggingface.co/neuralmind/bert-base-portuguese-cased#</a> <br>\n* ru\n        - <a href=\"https://huggingface.co/DeepPavlov/bert-base-bg-cs-pl-ru-cased\">https://huggingface.co/DeepPavlov/bert-base-bg-cs-pl-ru-cased</a>\n        - <a href=\"https://huggingface.co/DeepPavlov/rubert-base-cased#\">https://huggingface.co/DeepPavlov/rubert-base-cased#</a></p>\n\n<h3>English models</h3>\n\n<p>To add more diversities, We also trained <code>roberta-large, bert-large, albert</code> on english training data and made prediction on data translated into English and to our surprise, this did improve the performance! And we also had all cross-lingual models predict on english-translated test data and added into our final ensemble.</p>\n\n<h3>Training</h3>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1905.05583\">dynamic learning rate decay</a></li>\n<li>Adam optimizer</li>\n<li>binary cross entropy loss function</li>\n<li>learning rate of 1e-6 to 3e-6 for 2 epochs for large models and 3 epochs for base models</li>\n<li>after training on english and translated data, we both finetuned further on validation dataset for 1-2 epochs.</li>\n</ul>\n\n<p>My teammate had two different training schemes:\n* took <code>jigsaw-toxic-comment-train.csv</code> and its translated datasets to train 7 models, one for each language and combine prediction together in the way below</p>\n\n<p><code>\ntest = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv\")\nsub = pd.read_csv('/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission.csv')\ntest['toxic'] = 0.0\nfor lang in ['en', 'fr', 'es', 'it', 'pt', 'ru', 'tr']:\n    lang_df = pd.read_csv(f'/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission_{lang}.csv')\n    idx = test[test['lang'] == lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight\n    idx = test[test['lang'] != lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic']\ntest['toxic'] /= 7\n</code>\n*  took <code>jigsaw-unintended-bias-train.csv</code> and its translation to train in the first epoch and used the same data in the above point</p>\n\n<h3>Postprocessing</h3>\n\n<p>After seeing the trick used in nearly all the top solutions, we felt it was really a pity that we missed it because once my teammate forgot to average predictions for 7 language-specific models, causing some predictions to be greater and 0.0001 improvement. We didn't delve deeper into this due to the tiny improvement!</p>\n\n<h3>Ensemble</h3>\n\n<p>Power averaging with power 3 in the <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100661\">2nd place solution</a> in the last jigsaw competition but didn't bring improvement compared with normal weighted average.</p>\n\n<h3>What didn't work</h3>\n\n<ul>\n<li>auxiliary task with 5 other labels in the training data</li>\n<li>sample weights for different languages</li>\n<li>extra labelled hatespeech data</li>\n<li>pesudo labelling</li>\n<li>multi-sample dropout</li>\n<li>cyclic learning rate</li>\n<li>checkpoint ensemble</li>\n<li>training multilingual bert from scratch on multilingual wikipedia data dump</li>\n</ul>",
  "messages": [
    {
      "id": "898754",
      "postDate": "06/23/2020 17:58:54",
      "content": "<p>Thanks to Jigsaw to hold such amazing competition! And also thanks to my teammate <a href=\"/tonyxu\">@tonyxu</a> for diligence and brilliant ideas. This is our second jigsaw competition and also the first time to join such multilingual nlp competition. </p>\n\n<h2>Brief summary</h2>\n\n<p>Before team merger, I mainly focused on how to train a single model with decent performance while my teammate worked on training diverse models, especially training on single language data and combine them together. Before merger, we both achieved 0.9487 on public leaderboard. As you can see, we worked in relatively different directions and achieved 0.9499 right after team merger. After team merger, we focused on training more diverse models for diversity, which gave us great boost but also caused us to ignore the post processing trick most top winners found.</p>\n\n<h2>Detailed summary</h2>\n\n<h3>Data sampling</h3>\n\n<p>We used both english training data and its translation shared publicly. Since there were tremendous amount of training data available, the main issue was not too little but too much. I mainly adopted hard sampling technique to sample the training data. </p>\n\n<p>First, I included all positives and randomly sampled negatives to train a XLM-Roberta Large model and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment. </p>\n\n<p>Then for english data, we set a threshold for negatives while we took all positives due to the data unbalance. The data with gap between the predicted probability and label greater than the threshold were all taken and the data with gap smaller than that were sampled. The same logic for translated data except the threshold setting. A single threshold may cause issues because translation led to information loss. So I set a value range, say [0.3, 0.7] and any data in this range were all taken. I didn't convert soft label to hard 0/1 label because it worked better in \nthe last jigsaw competition. The number of training data is ~934k. XLM-Roberta Large trained on this data could reach 0.9469 on public lb and 0.9458 on private lb. </p>\n\n<h3>Multilingual Models</h3>\n\n<p>Our multilingual models mainly consist of <code>XLM-Roberta-Large</code>, <code>XLM-Roberta-base</code>, <code>Bert-multilingual-base-uncased</code>. My teammate also did language model finetuning on <code>jigsaw-unintended-bias</code> data. We don't know the exact number for each of them used in the final ensemble but the best public scores are 0.9469, 0.9364 and 0.9325 respectively. </p>\n\n<p>In terms of model structure, I used a vanilla classification head on the CLS token of the second to last layer or the concatenation of last four layers while my teammate took the <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/103280\">architecture of winning model</a> in the last jigsaw competition.</p>\n\n<h3>Monlingual Models</h3>\n\n<p>Just one week before the deadline, we found the monolingual language models in <a href=\"https://huggingface.co/models\">huggingface communities</a>. We trained the models listed below on corresponding translated language, which gave us boost from 0.9507 to 0.9512 on public lb. </p>\n\n<p>&gt; * es:\n        - <a href=\"https://huggingface.co/dccuchile/bert-base-spanish-wwm-uncased#\">https://huggingface.co/dccuchile/bert-base-spanish-wwm-uncased#</a>\n        - <a href=\"https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased#\">https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased#</a>\n        - <a href=\"https://github.com/dccuchile/beto\">https://github.com/dccuchile/beto</a> <br>\n* fr:\n        - <a href=\"https://huggingface.co/camembert/camembert-large\">https://huggingface.co/camembert/camembert-large</a>\n        - <a href=\"https://huggingface.co/camembert-base\">https://huggingface.co/camembert-base</a>\n        - <a href=\"https://huggingface.co/flaubert/flaubert-large-cased\">https://huggingface.co/flaubert/flaubert-large-cased</a>\n        - <a href=\"https://huggingface.co/flaubert/flaubert-base-uncased\">https://huggingface.co/flaubert/flaubert-base-uncased</a>\n        - <a href=\"https://huggingface.co/flaubert/flaubert-base-cased\">https://huggingface.co/flaubert/flaubert-base-cased</a>\n* tr:\n        - (128k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-128k-uncased\">https://huggingface.co/dbmdz/bert-base-turkish-128k-uncased</a>\n        - (128k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-128k-cased#\">https://huggingface.co/dbmdz/bert-base-turkish-128k-cased#</a>\n        - (32k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-uncased#\">https://huggingface.co/dbmdz/bert-base-turkish-uncased#</a>\n        - (32k vocabulary) <a href=\"https://huggingface.co/dbmdz/bert-base-turkish-cased#\">https://huggingface.co/dbmdz/bert-base-turkish-cased#</a> <br>\n* it\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-xxl-uncased#\">https://huggingface.co/dbmdz/bert-base-italian-xxl-uncased#</a>\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-xxl-cased#\">https://huggingface.co/dbmdz/bert-base-italian-xxl-cased#</a>\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-uncased#\">https://huggingface.co/dbmdz/bert-base-italian-uncased#</a>\n        - <a href=\"https://huggingface.co/dbmdz/bert-base-italian-cased#\">https://huggingface.co/dbmdz/bert-base-italian-cased#</a> <br>\n* pt\n        - <a href=\"https://huggingface.co/neuralmind/bert-large-portuguese-cased#\">https://huggingface.co/neuralmind/bert-large-portuguese-cased#</a>\n        - <a href=\"https://huggingface.co/neuralmind/bert-base-portuguese-cased#\">https://huggingface.co/neuralmind/bert-base-portuguese-cased#</a> <br>\n* ru\n        - <a href=\"https://huggingface.co/DeepPavlov/bert-base-bg-cs-pl-ru-cased\">https://huggingface.co/DeepPavlov/bert-base-bg-cs-pl-ru-cased</a>\n        - <a href=\"https://huggingface.co/DeepPavlov/rubert-base-cased#\">https://huggingface.co/DeepPavlov/rubert-base-cased#</a></p>\n\n<h3>English models</h3>\n\n<p>To add more diversities, We also trained <code>roberta-large, bert-large, albert</code> on english training data and made prediction on data translated into English and to our surprise, this did improve the performance! And we also had all cross-lingual models predict on english-translated test data and added into our final ensemble.</p>\n\n<h3>Training</h3>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1905.05583\">dynamic learning rate decay</a></li>\n<li>Adam optimizer</li>\n<li>binary cross entropy loss function</li>\n<li>learning rate of 1e-6 to 3e-6 for 2 epochs for large models and 3 epochs for base models</li>\n<li>after training on english and translated data, we both finetuned further on validation dataset for 1-2 epochs.</li>\n</ul>\n\n<p>My teammate had two different training schemes:\n* took <code>jigsaw-toxic-comment-train.csv</code> and its translated datasets to train 7 models, one for each language and combine prediction together in the way below</p>\n\n<p><code>\ntest = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv\")\nsub = pd.read_csv('/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission.csv')\ntest['toxic'] = 0.0\nfor lang in ['en', 'fr', 'es', 'it', 'pt', 'ru', 'tr']:\n    lang_df = pd.read_csv(f'/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission_{lang}.csv')\n    idx = test[test['lang'] == lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight\n    idx = test[test['lang'] != lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic']\ntest['toxic'] /= 7\n</code>\n*  took <code>jigsaw-unintended-bias-train.csv</code> and its translation to train in the first epoch and used the same data in the above point</p>\n\n<h3>Postprocessing</h3>\n\n<p>After seeing the trick used in nearly all the top solutions, we felt it was really a pity that we missed it because once my teammate forgot to average predictions for 7 language-specific models, causing some predictions to be greater and 0.0001 improvement. We didn't delve deeper into this due to the tiny improvement!</p>\n\n<h3>Ensemble</h3>\n\n<p>Power averaging with power 3 in the <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100661\">2nd place solution</a> in the last jigsaw competition but didn't bring improvement compared with normal weighted average.</p>\n\n<h3>What didn't work</h3>\n\n<ul>\n<li>auxiliary task with 5 other labels in the training data</li>\n<li>sample weights for different languages</li>\n<li>extra labelled hatespeech data</li>\n<li>pesudo labelling</li>\n<li>multi-sample dropout</li>\n<li>cyclic learning rate</li>\n<li>checkpoint ensemble</li>\n<li>training multilingual bert from scratch on multilingual wikipedia data dump</li>\n</ul>",
      "rawMarkdown": "Thanks to Jigsaw to hold such amazing competition! And also thanks to my teammate @tonyxu for diligence and brilliant ideas. This is our second jigsaw competition and also the first time to join such multilingual nlp competition. \n\n## Brief summary \nBefore team merger, I mainly focused on how to train a single model with decent performance while my teammate worked on training diverse models, especially training on single language data and combine them together. Before merger, we both achieved 0.9487 on public leaderboard. As you can see, we worked in relatively different directions and achieved 0.9499 right after team merger. After team merger, we focused on training more diverse models for diversity, which gave us great boost but also caused us to ignore the post processing trick most top winners found.\n\n## Detailed summary\n\n### Data sampling\nWe used both english training data and its translation shared publicly. Since there were tremendous amount of training data available, the main issue was not too little but too much. I mainly adopted hard sampling technique to sample the training data. \n\nFirst, I included all positives and randomly sampled negatives to train a XLM-Roberta Large model and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment. \n\nThen for english data, we set a threshold for negatives while we took all positives due to the data unbalance. The data with gap between the predicted probability and label greater than the threshold were all taken and the data with gap smaller than that were sampled. The same logic for translated data except the threshold setting. A single threshold may cause issues because translation led to information loss. So I set a value range, say [0.3, 0.7] and any data in this range were all taken. I didn't convert soft label to hard 0/1 label because it worked better in \nthe last jigsaw competition. The number of training data is ~934k. XLM-Roberta Large trained on this data could reach 0.9469 on public lb and 0.9458 on private lb. \n\n### Multilingual Models\nOur multilingual models mainly consist of `XLM-Roberta-Large`, `XLM-Roberta-base`, `Bert-multilingual-base-uncased`. My teammate also did language model finetuning on `jigsaw-unintended-bias` data. We don't know the exact number for each of them used in the final ensemble but the best public scores are 0.9469, 0.9364 and 0.9325 respectively. \n\nIn terms of model structure, I used a vanilla classification head on the CLS token of the second to last layer or the concatenation of last four layers while my teammate took the [architecture of winning model](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/103280) in the last jigsaw competition.\n\n### Monlingual Models\nJust one week before the deadline, we found the monolingual language models in [huggingface communities](https://huggingface.co/models). We trained the models listed below on corresponding translated language, which gave us boost from 0.9507 to 0.9512 on public lb. \n\n&gt; * es:\n        - https://huggingface.co/dccuchile/bert-base-spanish-wwm-uncased#\n        - https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased#\n        - https://github.com/dccuchile/beto        \n* fr:\n        - https://huggingface.co/camembert/camembert-large\n        - https://huggingface.co/camembert-base\n        - https://huggingface.co/flaubert/flaubert-large-cased\n        - https://huggingface.co/flaubert/flaubert-base-uncased\n        - https://huggingface.co/flaubert/flaubert-base-cased\n* tr:\n        - (128k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-128k-uncased\n        - (128k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-128k-cased#\n        - (32k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-uncased#\n        - (32k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-cased#     \n* it\n        - https://huggingface.co/dbmdz/bert-base-italian-xxl-uncased#\n        - https://huggingface.co/dbmdz/bert-base-italian-xxl-cased#\n        - https://huggingface.co/dbmdz/bert-base-italian-uncased#\n        - https://huggingface.co/dbmdz/bert-base-italian-cased#     \n* pt\n        - https://huggingface.co/neuralmind/bert-large-portuguese-cased#\n        - https://huggingface.co/neuralmind/bert-base-portuguese-cased#      \n* ru\n        - https://huggingface.co/DeepPavlov/bert-base-bg-cs-pl-ru-cased\n        - https://huggingface.co/DeepPavlov/rubert-base-cased#\n\n### English models\nTo add more diversities, We also trained `roberta-large, bert-large, albert` on english training data and made prediction on data translated into English and to our surprise, this did improve the performance! And we also had all cross-lingual models predict on english-translated test data and added into our final ensemble.\n\n### Training \n* [dynamic learning rate decay](https://arxiv.org/abs/1905.05583)\n* Adam optimizer\n* binary cross entropy loss function\n* learning rate of 1e-6 to 3e-6 for 2 epochs for large models and 3 epochs for base models\n* after training on english and translated data, we both finetuned further on validation dataset for 1-2 epochs.\n\nMy teammate had two different training schemes:\n* took `jigsaw-toxic-comment-train.csv` and its translated datasets to train 7 models, one for each language and combine prediction together in the way below\n\n```\ntest = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv\")\nsub = pd.read_csv('/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission.csv')\ntest['toxic'] = 0.0\nfor lang in ['en', 'fr', 'es', 'it', 'pt', 'ru', 'tr']:\n    lang_df = pd.read_csv(f'/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission_{lang}.csv')\n    idx = test[test['lang'] == lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight\n    idx = test[test['lang'] != lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic']\ntest['toxic'] /= 7\n```\n*  took `jigsaw-unintended-bias-train.csv` and its translation to train in the first epoch and used the same data in the above point\n\n### Postprocessing\nAfter seeing the trick used in nearly all the top solutions, we felt it was really a pity that we missed it because once my teammate forgot to average predictions for 7 language-specific models, causing some predictions to be greater and 0.0001 improvement. We didn't delve deeper into this due to the tiny improvement!\n\n### Ensemble\nPower averaging with power 3 in the [2nd place solution](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100661) in the last jigsaw competition but didn't bring improvement compared with normal weighted average.\n\n### What didn't work\n* auxiliary task with 5 other labels in the training data\n* sample weights for different languages\n* extra labelled hatespeech data\n* pesudo labelling\n* multi-sample dropout\n* cyclic learning rate\n* checkpoint ensemble\n* training multilingual bert from scratch on multilingual wikipedia data dump",
      "votes": null
    },
    {
      "id": "899767",
      "postDate": "06/24/2020 12:35:14",
      "content": "<p>Congratulations on your first gold! Surprised you haven't had more comments to say that. One thing that is really clear from your write up is the benefit of teaming up, for Kagglers who haven't tried that before I'd definitely recommend it.</p>\n\n<p>One thing I like about your write up is the list of what didn't work. It shows the effort needed in top solutions! Once again congrats.</p>\n\n<p>This solution and many other top solutions are listed <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161066?rvi=1\">here</a>.</p>",
      "rawMarkdown": "Congratulations on your first gold! Surprised you haven't had more comments to say that. One thing that is really clear from your write up is the benefit of teaming up, for Kagglers who haven't tried that before I'd definitely recommend it.\n\nOne thing I like about your write up is the list of what didn't work. It shows the effort needed in top solutions! Once again congrats.\n\nThis solution and many other top solutions are listed [here](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161066?rvi=1).",
      "votes": null
    },
    {
      "id": "899992",
      "postDate": "06/24/2020 15:06:27",
      "content": "<p>Thanks! I am also very excited to win my first gold medal! And thanks for the compilation of top solutions!</p>",
      "rawMarkdown": "Thanks! I am also very excited to win my first gold medal! And thanks for the compilation of top solutions!",
      "votes": null
    },
    {
      "id": "900892",
      "postDate": "06/25/2020 06:08:13",
      "content": "<p>Congrats <a href=\"/nzholmes\">@nzholmes</a>, great write up!</p>",
      "rawMarkdown": "Congrats @nzholmes, great write up!",
      "votes": null
    },
    {
      "id": "901595",
      "postDate": "06/25/2020 15:36:19",
      "content": "<p>congrats, nice write-up.</p>",
      "rawMarkdown": "congrats, nice write-up.",
      "votes": null
    },
    {
      "id": "902642",
      "postDate": "06/26/2020 09:27:01",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "918768",
      "postDate": "07/07/2020 13:33:56",
      "content": "<p>Congrats！\nI have some little questions.</p>\n\n<p>1.\n&gt; and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment.</p>\n\n<p>When you did this prediction, did you use the model to predict out-of-fold data, or you didn't care whether each sample  is seen by this model during training?</p>\n\n<p>2.</p>\n\n<blockquote>\n  <p>test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight</p>\n</blockquote>\n\n<p>Is this a weight bigger than 1?</p>",
      "rawMarkdown": "Congrats！\nI have some little questions.\n\n1.\n&gt; and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment.\n\nWhen you did this prediction, did you use the model to predict out-of-fold data, or you didn't care whether each sample  is seen by this model during training?\n\n2.\n&gt; test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight\n\nIs this a weight bigger than 1?",
      "votes": null
    },
    {
      "id": "919719",
      "postDate": "07/08/2020 04:32:37",
      "content": "<p>Thanks! \n1.  I didn't care if the randomly selected samples were seen in the training set. This was not done for any logic reasons but for simpleness. And another reason was that there were huge numbers of negatives, which made the omission of some samples negligible.\n2. <code>weight</code> could be greater than 1. It was 1 in the final submission due to the better performance than other values.</p>",
      "rawMarkdown": "Thanks! \n1.  I didn't care if the randomly selected samples were seen in the training set. This was not done for any logic reasons but for simpleness. And another reason was that there were huge numbers of negatives, which made the omission of some samples negligible.\n2. `weight` could be greater than 1. It was 1 in the final submission due to the better performance than other values.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 899767,
      "author_name": "datahobbit",
      "author_url": "",
      "post_date": "06/24/2020 12:35:14",
      "content": "<p>Congratulations on your first gold! Surprised you haven't had more comments to say that. One thing that is really clear from your write up is the benefit of teaming up, for Kagglers who haven't tried that before I'd definitely recommend it.</p>\n\n<p>One thing I like about your write up is the list of what didn't work. It shows the effort needed in top solutions! Once again congrats.</p>\n\n<p>This solution and many other top solutions are listed <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161066?rvi=1\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 899992,
          "author_name": "nzholmes",
          "author_url": "",
          "post_date": "06/24/2020 15:06:27",
          "content": "<p>Thanks! I am also very excited to win my first gold medal! And thanks for the compilation of top solutions!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 900892,
      "author_name": "ankitsajwan",
      "author_url": "",
      "post_date": "06/25/2020 06:08:13",
      "content": "<p>Congrats <a href=\"/nzholmes\">@nzholmes</a>, great write up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 901595,
      "author_name": "nachiket273",
      "author_url": "",
      "post_date": "06/25/2020 15:36:19",
      "content": "<p>congrats, nice write-up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 902642,
      "author_name": "rashidulhasanhridoy",
      "author_url": "",
      "post_date": "06/26/2020 09:27:01",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 918768,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "07/07/2020 13:33:56",
      "content": "<p>Congrats！\nI have some little questions.</p>\n\n<p>1.\n&gt; and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment.</p>\n\n<p>When you did this prediction, did you use the model to predict out-of-fold data, or you didn't care whether each sample  is seen by this model during training?</p>\n\n<p>2.</p>\n\n<blockquote>\n  <p>test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight</p>\n</blockquote>\n\n<p>Is this a weight bigger than 1?</p>",
      "votes": null,
      "replies": [
        {
          "id": 919719,
          "author_name": "nzholmes",
          "author_url": "",
          "post_date": "07/08/2020 04:32:37",
          "content": "<p>Thanks! \n1.  I didn't care if the randomly selected samples were seen in the training set. This was not done for any logic reasons but for simpleness. And another reason was that there were huge numbers of negatives, which made the omission of some samples negligible.\n2. <code>weight</code> could be greater than 1. It was 1 in the final submission due to the better performance than other values.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "898754": "Thanks to Jigsaw to hold such amazing competition! And also thanks to my teammate @tonyxu for diligence and brilliant ideas. This is our second jigsaw competition and also the first time to join such multilingual nlp competition. \n\n## Brief summary \nBefore team merger, I mainly focused on how to train a single model with decent performance while my teammate worked on training diverse models, especially training on single language data and combine them together. Before merger, we both achieved 0.9487 on public leaderboard. As you can see, we worked in relatively different directions and achieved 0.9499 right after team merger. After team merger, we focused on training more diverse models for diversity, which gave us great boost but also caused us to ignore the post processing trick most top winners found.\n\n## Detailed summary\n\n### Data sampling\nWe used both english training data and its translation shared publicly. Since there were tremendous amount of training data available, the main issue was not too little but too much. I mainly adopted hard sampling technique to sample the training data. \n\nFirst, I included all positives and randomly sampled negatives to train a XLM-Roberta Large model and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment. \n\nThen for english data, we set a threshold for negatives while we took all positives due to the data unbalance. The data with gap between the predicted probability and label greater than the threshold were all taken and the data with gap smaller than that were sampled. The same logic for translated data except the threshold setting. A single threshold may cause issues because translation led to information loss. So I set a value range, say [0.3, 0.7] and any data in this range were all taken. I didn't convert soft label to hard 0/1 label because it worked better in \nthe last jigsaw competition. The number of training data is ~934k. XLM-Roberta Large trained on this data could reach 0.9469 on public lb and 0.9458 on private lb. \n\n### Multilingual Models\nOur multilingual models mainly consist of `XLM-Roberta-Large`, `XLM-Roberta-base`, `Bert-multilingual-base-uncased`. My teammate also did language model finetuning on `jigsaw-unintended-bias` data. We don't know the exact number for each of them used in the final ensemble but the best public scores are 0.9469, 0.9364 and 0.9325 respectively. \n\nIn terms of model structure, I used a vanilla classification head on the CLS token of the second to last layer or the concatenation of last four layers while my teammate took the [architecture of winning model](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/103280) in the last jigsaw competition.\n\n### Monlingual Models\nJust one week before the deadline, we found the monolingual language models in [huggingface communities](https://huggingface.co/models). We trained the models listed below on corresponding translated language, which gave us boost from 0.9507 to 0.9512 on public lb. \n\n&gt; * es:\n        - https://huggingface.co/dccuchile/bert-base-spanish-wwm-uncased#\n        - https://huggingface.co/dccuchile/bert-base-spanish-wwm-cased#\n        - https://github.com/dccuchile/beto        \n* fr:\n        - https://huggingface.co/camembert/camembert-large\n        - https://huggingface.co/camembert-base\n        - https://huggingface.co/flaubert/flaubert-large-cased\n        - https://huggingface.co/flaubert/flaubert-base-uncased\n        - https://huggingface.co/flaubert/flaubert-base-cased\n* tr:\n        - (128k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-128k-uncased\n        - (128k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-128k-cased#\n        - (32k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-uncased#\n        - (32k vocabulary) https://huggingface.co/dbmdz/bert-base-turkish-cased#     \n* it\n        - https://huggingface.co/dbmdz/bert-base-italian-xxl-uncased#\n        - https://huggingface.co/dbmdz/bert-base-italian-xxl-cased#\n        - https://huggingface.co/dbmdz/bert-base-italian-uncased#\n        - https://huggingface.co/dbmdz/bert-base-italian-cased#     \n* pt\n        - https://huggingface.co/neuralmind/bert-large-portuguese-cased#\n        - https://huggingface.co/neuralmind/bert-base-portuguese-cased#      \n* ru\n        - https://huggingface.co/DeepPavlov/bert-base-bg-cs-pl-ru-cased\n        - https://huggingface.co/DeepPavlov/rubert-base-cased#\n\n### English models\nTo add more diversities, We also trained `roberta-large, bert-large, albert` on english training data and made prediction on data translated into English and to our surprise, this did improve the performance! And we also had all cross-lingual models predict on english-translated test data and added into our final ensemble.\n\n### Training \n* [dynamic learning rate decay](https://arxiv.org/abs/1905.05583)\n* Adam optimizer\n* binary cross entropy loss function\n* learning rate of 1e-6 to 3e-6 for 2 epochs for large models and 3 epochs for base models\n* after training on english and translated data, we both finetuned further on validation dataset for 1-2 epochs.\n\nMy teammate had two different training schemes:\n* took `jigsaw-toxic-comment-train.csv` and its translated datasets to train 7 models, one for each language and combine prediction together in the way below\n\n```\ntest = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv\")\nsub = pd.read_csv('/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission.csv')\ntest['toxic'] = 0.0\nfor lang in ['en', 'fr', 'es', 'it', 'pt', 'ru', 'tr']:\n    lang_df = pd.read_csv(f'/kaggle/input/toxu-submissions/xlmroberta-large-lm-all/submission_{lang}.csv')\n    idx = test[test['lang'] == lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight\n    idx = test[test['lang'] != lang].index\n    test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic']\ntest['toxic'] /= 7\n```\n*  took `jigsaw-unintended-bias-train.csv` and its translation to train in the first epoch and used the same data in the above point\n\n### Postprocessing\nAfter seeing the trick used in nearly all the top solutions, we felt it was really a pity that we missed it because once my teammate forgot to average predictions for 7 language-specific models, causing some predictions to be greater and 0.0001 improvement. We didn't delve deeper into this due to the tiny improvement!\n\n### Ensemble\nPower averaging with power 3 in the [2nd place solution](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100661) in the last jigsaw competition but didn't bring improvement compared with normal weighted average.\n\n### What didn't work\n* auxiliary task with 5 other labels in the training data\n* sample weights for different languages\n* extra labelled hatespeech data\n* pesudo labelling\n* multi-sample dropout\n* cyclic learning rate\n* checkpoint ensemble\n* training multilingual bert from scratch on multilingual wikipedia data dump",
    "899767": "Congratulations on your first gold! Surprised you haven't had more comments to say that. One thing that is really clear from your write up is the benefit of teaming up, for Kagglers who haven't tried that before I'd definitely recommend it.\n\nOne thing I like about your write up is the list of what didn't work. It shows the effort needed in top solutions! Once again congrats.\n\nThis solution and many other top solutions are listed [here](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161066?rvi=1).",
    "899992": "Thanks! I am also very excited to win my first gold medal! And thanks for the compilation of top solutions!",
    "900892": "Congrats @nzholmes, great write up!",
    "901595": "congrats, nice write-up.",
    "902642": "Congrats!",
    "918768": "Congrats！\nI have some little questions.\n\n1.\n&gt; and then used it to make predictions on all training data including all english data and translated jigsaw toxic comment.\n\nWhen you did this prediction, did you use the model to predict out-of-fold data, or you didn't care whether each sample  is seen by this model during training?\n\n2.\n&gt; test.loc[idx, 'toxic'] += lang_df.loc[idx, 'toxic'] * weight\n\nIs this a weight bigger than 1?",
    "919719": "Thanks! \n1.  I didn't care if the randomly selected samples were seen in the training set. This was not done for any logic reasons but for simpleness. And another reason was that there were huge numbers of negatives, which made the omission of some samples negligible.\n2. `weight` could be greater than 1. It was 1 in the final submission due to the better performance than other values."
  },
  "source": "meta"
}