{
  "id": 160980,
  "title": "4th place solution",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/rapids-nlp-4th-place-solution",
  "author_name": "",
  "post_date": "2020-06-23T10:48:17.360872100Z",
  "votes": 106,
  "comment_count": 38,
  "views": 0,
  "content": "<p>Thanks to kaggle for hosting such an interesting and challenging competition! Multilingual NLP is something very interesting yet difficult. In hindsight I (Dieter) wish I’d spend less time on the tweet sentiment extraction competition and more on this one. </p>\n\n<h1>Brief Summary</h1>\n\n<p>Our solution is a simple blend of several transformer models (mainly xlm-roberta-large) paired with some post-processing. The architecture was the same as in public available kernels with the classification head taking either max+mean pooling of hidden states or the hidden state of the CLS token. We used a 3-step approach for training our models, where starting from 7 languages we fine-tuned to 3 languages and finished at a single language. As this 3-step approach results in distortion of global predictions we shifted languages individually by a factor as a post-processing step. </p>\n\n<h1>Detailed Summary</h1>\n\n<p>We are still astonished by the great result in such a short time. While @aerdem4 worked a bit longer on this competition and already had a good understanding of specific challenges @cpmpml and @christofhenkel joined right after the tweet sentiment competition a few hours before team merger deadline. Two days ago we were still at a 100+ spot and were hoping for a silver medal at best. But with persistence and the right amount of intuition what ideas to proceed with, and of course some luck we managed to climb right to the upper gold position.</p>\n\n<h2>Preprocessing</h2>\n\n<p>None. That's NLP in 2020 :D</p>\n\n<h2>Speeding up training</h2>\n\n<p>We did two things.  One was to downsample negative samples to get a more balanced dataset and a smaller dataset at the same time.  If you only use as many negative as positive then dataset is reduced 5x roughly, same for time to run one epoch.  Some of our models were trained that way.</p>\n\n<p>Another speedup came from padding by batch.  The main idea is to limit the amount of padding to what is necessary for the current batch.  Inputs are padded to the length of the longest input in the batch rather than a fixed length, say 512.  This is now a well known technique, and it has been used in previous competitions.  It accelerates training significantly. <br>\nWe refined the idea by sorting the samples by their length so that the samples in a given batch have similar length.  This reduces even further the need for padding.  If all inputs in the batch have the same length then there is no padding at all.  </p>\n\n<p>Given samples are sorted, we cannot shuffle them in training mode.  We rather shuffled batches. This yields an extra 2x speedup compared to the original padding by batch.  Training one epoch for xlm-roberta-large on the first train set takes about 17 minutes on a V100 GPU.</p>\n\n<h2>Cross-validation</h2>\n\n<p>As time was short and hence we did not have many submissions to spare for getting LB feedback we worked on getting a reliable cross-validation. We used a mix of Group-3fold per language and simple 3fold of the validation data to represent the unknown languages ru,pt,fr as well as the known languages it, es and tr. which results in a 6fold scheme. So do for example fold1 represent what if es would not be in valid.csv which should behave the same as pt is not in valid. Taking the mean AUC of all folds was a very good proxy for Public and Private LB.</p>\n\n<h2>Architectures</h2>\n\n<p>As most teams we used xlm-roberta large as the main backbone with a simple classification head that either uses max+mean pooling or the hidden states of the CLS token. In total we used the following backbones with their approximate weighting in our final ensemble in brackets.</p>\n\n<ul>\n<li>5x xlm-roberta-large (85%)</li>\n<li>1x xlm-roberta-base (5%)</li>\n<li>1x mBart-large (10%)</li>\n</ul>\n\n<h2>Training strategy</h2>\n\n<p>One main ingredient for our good result is a stepwise finetuning towards single language models for tr, it and es. Let us illustrate the whole approach before explaining:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd03ded105462fb5ba610393ba97e4ad8%2Fjigsaw%20arch2.001.jpeg?generation=1592909197036996&amp;alt=media\" alt=\"\"></p>\n\n<p>As input data we used the english  jigsaw-toxic-comment-train.csv and its six translations available as public datasets which combined are roughly 1.4M comments. I realized that training a model directly on the combined dataset is not optimal, as each comment appears 7x in each epoch (although in different languages) and the model overfits on that. So I divided the combined dataset into 7 stratified parts, where each part contains a comment only once. For each fold we then finetuned a transformer in a 3step manner:</p>\n\n<ul>\n<li>Step 1: finetune to all 7 languages for 2 epochs</li>\n<li>Step 2: finetune only to the full valid.csv which only has tr, it and es</li>\n<li>Step 3: 3x finetune to each language in valid resulting in 3 models</li>\n</ul>\n\n<p>We then use the step1 model for predicting ru, the step2 model for predicting pt and fr and the respective step3 models for tr, it and es. Using the step2 model for pt and fr gave a significant boost compared to using step1 model for those due. Most likely due to the language similarity between it, es, fr and pt.</p>\n\n<p>We used mainly a batchsize of 32 using gradient accumulation and a learning rate of 3e-6 with linear decay with AdamW optimizer. One thing worth mentioning is that we train on a max sequence length of 512 due to the dataloader mentioned above which sorts the data by length and then randomly serves batches of same length comments. Only the huge speed-up paired with the non padding of shorter comments made using a 512 max length reasonable. </p>\n\n<p>Apart from one model which was built on top of a public kernel, all models were trained using pytorch and GPU.</p>\n\n<h2>Ensembling</h2>\n\n<p>As for ensembling we did team member individual ensembling and then combined the models of the team members and public kernels by weighted sum of rank percentile </p>\n\n<h2>Post-processing</h2>\n\n<p>Another key ingredient, which we luckily found on the last day is post processing of the final predictions. Although languages are derived from individual models, the competition metric is sensitive to the global rank of all predictions and not language specific. So we took care of the languages having a correct relation to each other, by shifting the predictions of each language individually. In our final submission for example we used the following factors </p>\n\n<p><code>\ntest_df.loc[test_df[\"lang\"] == \"es\", \"toxic\"] *= 1.06\ntest_df.loc[test_df[\"lang\"] == \"fr\", \"toxic\"] *= 1.04\ntest_df.loc[test_df[\"lang\"] == \"it\", \"toxic\"] *= 0.97\ntest_df.loc[test_df[\"lang\"] == \"pt\", \"toxic\"] *= 0.96\ntest_df.loc[test_df[\"lang\"] == \"tr\", \"toxic\"] *= 0.98\n</code></p>\n\n<p>We derived the factors by matching the test prediction mean with the public LB mean for each language individually. We had already made 6 submissions for probing the test set by setting one language to 1 and the rest to 0. Then we made sure that means are aligned to give similar experimental AUC. That posprocessing moved us from 14th place to 4th place on public LB! </p>\n\n<p>While this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.</p>\n\n<p>Thanks for reading. Questions welcome.</p>",
  "messages": [
    {
      "id": "898155",
      "postDate": "06/23/2020 10:48:17",
      "content": "<p>Thanks to kaggle for hosting such an interesting and challenging competition! Multilingual NLP is something very interesting yet difficult. In hindsight I (Dieter) wish I’d spend less time on the tweet sentiment extraction competition and more on this one. </p>\n\n<h1>Brief Summary</h1>\n\n<p>Our solution is a simple blend of several transformer models (mainly xlm-roberta-large) paired with some post-processing. The architecture was the same as in public available kernels with the classification head taking either max+mean pooling of hidden states or the hidden state of the CLS token. We used a 3-step approach for training our models, where starting from 7 languages we fine-tuned to 3 languages and finished at a single language. As this 3-step approach results in distortion of global predictions we shifted languages individually by a factor as a post-processing step. </p>\n\n<h1>Detailed Summary</h1>\n\n<p>We are still astonished by the great result in such a short time. While @aerdem4 worked a bit longer on this competition and already had a good understanding of specific challenges @cpmpml and @christofhenkel joined right after the tweet sentiment competition a few hours before team merger deadline. Two days ago we were still at a 100+ spot and were hoping for a silver medal at best. But with persistence and the right amount of intuition what ideas to proceed with, and of course some luck we managed to climb right to the upper gold position.</p>\n\n<h2>Preprocessing</h2>\n\n<p>None. That's NLP in 2020 :D</p>\n\n<h2>Speeding up training</h2>\n\n<p>We did two things.  One was to downsample negative samples to get a more balanced dataset and a smaller dataset at the same time.  If you only use as many negative as positive then dataset is reduced 5x roughly, same for time to run one epoch.  Some of our models were trained that way.</p>\n\n<p>Another speedup came from padding by batch.  The main idea is to limit the amount of padding to what is necessary for the current batch.  Inputs are padded to the length of the longest input in the batch rather than a fixed length, say 512.  This is now a well known technique, and it has been used in previous competitions.  It accelerates training significantly. <br>\nWe refined the idea by sorting the samples by their length so that the samples in a given batch have similar length.  This reduces even further the need for padding.  If all inputs in the batch have the same length then there is no padding at all.  </p>\n\n<p>Given samples are sorted, we cannot shuffle them in training mode.  We rather shuffled batches. This yields an extra 2x speedup compared to the original padding by batch.  Training one epoch for xlm-roberta-large on the first train set takes about 17 minutes on a V100 GPU.</p>\n\n<h2>Cross-validation</h2>\n\n<p>As time was short and hence we did not have many submissions to spare for getting LB feedback we worked on getting a reliable cross-validation. We used a mix of Group-3fold per language and simple 3fold of the validation data to represent the unknown languages ru,pt,fr as well as the known languages it, es and tr. which results in a 6fold scheme. So do for example fold1 represent what if es would not be in valid.csv which should behave the same as pt is not in valid. Taking the mean AUC of all folds was a very good proxy for Public and Private LB.</p>\n\n<h2>Architectures</h2>\n\n<p>As most teams we used xlm-roberta large as the main backbone with a simple classification head that either uses max+mean pooling or the hidden states of the CLS token. In total we used the following backbones with their approximate weighting in our final ensemble in brackets.</p>\n\n<ul>\n<li>5x xlm-roberta-large (85%)</li>\n<li>1x xlm-roberta-base (5%)</li>\n<li>1x mBart-large (10%)</li>\n</ul>\n\n<h2>Training strategy</h2>\n\n<p>One main ingredient for our good result is a stepwise finetuning towards single language models for tr, it and es. Let us illustrate the whole approach before explaining:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd03ded105462fb5ba610393ba97e4ad8%2Fjigsaw%20arch2.001.jpeg?generation=1592909197036996&amp;alt=media\" alt=\"\"></p>\n\n<p>As input data we used the english  jigsaw-toxic-comment-train.csv and its six translations available as public datasets which combined are roughly 1.4M comments. I realized that training a model directly on the combined dataset is not optimal, as each comment appears 7x in each epoch (although in different languages) and the model overfits on that. So I divided the combined dataset into 7 stratified parts, where each part contains a comment only once. For each fold we then finetuned a transformer in a 3step manner:</p>\n\n<ul>\n<li>Step 1: finetune to all 7 languages for 2 epochs</li>\n<li>Step 2: finetune only to the full valid.csv which only has tr, it and es</li>\n<li>Step 3: 3x finetune to each language in valid resulting in 3 models</li>\n</ul>\n\n<p>We then use the step1 model for predicting ru, the step2 model for predicting pt and fr and the respective step3 models for tr, it and es. Using the step2 model for pt and fr gave a significant boost compared to using step1 model for those due. Most likely due to the language similarity between it, es, fr and pt.</p>\n\n<p>We used mainly a batchsize of 32 using gradient accumulation and a learning rate of 3e-6 with linear decay with AdamW optimizer. One thing worth mentioning is that we train on a max sequence length of 512 due to the dataloader mentioned above which sorts the data by length and then randomly serves batches of same length comments. Only the huge speed-up paired with the non padding of shorter comments made using a 512 max length reasonable. </p>\n\n<p>Apart from one model which was built on top of a public kernel, all models were trained using pytorch and GPU.</p>\n\n<h2>Ensembling</h2>\n\n<p>As for ensembling we did team member individual ensembling and then combined the models of the team members and public kernels by weighted sum of rank percentile </p>\n\n<h2>Post-processing</h2>\n\n<p>Another key ingredient, which we luckily found on the last day is post processing of the final predictions. Although languages are derived from individual models, the competition metric is sensitive to the global rank of all predictions and not language specific. So we took care of the languages having a correct relation to each other, by shifting the predictions of each language individually. In our final submission for example we used the following factors </p>\n\n<p><code>\ntest_df.loc[test_df[\"lang\"] == \"es\", \"toxic\"] *= 1.06\ntest_df.loc[test_df[\"lang\"] == \"fr\", \"toxic\"] *= 1.04\ntest_df.loc[test_df[\"lang\"] == \"it\", \"toxic\"] *= 0.97\ntest_df.loc[test_df[\"lang\"] == \"pt\", \"toxic\"] *= 0.96\ntest_df.loc[test_df[\"lang\"] == \"tr\", \"toxic\"] *= 0.98\n</code></p>\n\n<p>We derived the factors by matching the test prediction mean with the public LB mean for each language individually. We had already made 6 submissions for probing the test set by setting one language to 1 and the rest to 0. Then we made sure that means are aligned to give similar experimental AUC. That posprocessing moved us from 14th place to 4th place on public LB! </p>\n\n<p>While this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.</p>\n\n<p>Thanks for reading. Questions welcome.</p>",
      "rawMarkdown": "Thanks to kaggle for hosting such an interesting and challenging competition! Multilingual NLP is something very interesting yet difficult. In hindsight I (Dieter) wish I’d spend less time on the tweet sentiment extraction competition and more on this one. \n\n# Brief Summary\nOur solution is a simple blend of several transformer models (mainly xlm-roberta-large) paired with some post-processing. The architecture was the same as in public available kernels with the classification head taking either max+mean pooling of hidden states or the hidden state of the CLS token. We used a 3-step approach for training our models, where starting from 7 languages we fine-tuned to 3 languages and finished at a single language. As this 3-step approach results in distortion of global predictions we shifted languages individually by a factor as a post-processing step. \n\n# Detailed Summary\nWe are still astonished by the great result in such a short time. While @aerdem4 worked a bit longer on this competition and already had a good understanding of specific challenges @cpmpml and @christofhenkel joined right after the tweet sentiment competition a few hours before team merger deadline. Two days ago we were still at a 100+ spot and were hoping for a silver medal at best. But with persistence and the right amount of intuition what ideas to proceed with, and of course some luck we managed to climb right to the upper gold position.\n\n## Preprocessing\nNone. That's NLP in 2020 :D\n\n## Speeding up training\nWe did two things.  One was to downsample negative samples to get a more balanced dataset and a smaller dataset at the same time.  If you only use as many negative as positive then dataset is reduced 5x roughly, same for time to run one epoch.  Some of our models were trained that way.\n\nAnother speedup came from padding by batch.  The main idea is to limit the amount of padding to what is necessary for the current batch.  Inputs are padded to the length of the longest input in the batch rather than a fixed length, say 512.  This is now a well known technique, and it has been used in previous competitions.  It accelerates training significantly.  \nWe refined the idea by sorting the samples by their length so that the samples in a given batch have similar length.  This reduces even further the need for padding.  If all inputs in the batch have the same length then there is no padding at all.  \n\nGiven samples are sorted, we cannot shuffle them in training mode.  We rather shuffled batches. This yields an extra 2x speedup compared to the original padding by batch.  Training one epoch for xlm-roberta-large on the first train set takes about 17 minutes on a V100 GPU.\n\n## Cross-validation\nAs time was short and hence we did not have many submissions to spare for getting LB feedback we worked on getting a reliable cross-validation. We used a mix of Group-3fold per language and simple 3fold of the validation data to represent the unknown languages ru,pt,fr as well as the known languages it, es and tr. which results in a 6fold scheme. So do for example fold1 represent what if es would not be in valid.csv which should behave the same as pt is not in valid. Taking the mean AUC of all folds was a very good proxy for Public and Private LB.\n\n## Architectures\nAs most teams we used xlm-roberta large as the main backbone with a simple classification head that either uses max+mean pooling or the hidden states of the CLS token. In total we used the following backbones with their approximate weighting in our final ensemble in brackets.\n \n- 5x xlm-roberta-large (85%)\n- 1x xlm-roberta-base (5%)\n- 1x mBart-large (10%)\n\n## Training strategy\nOne main ingredient for our good result is a stepwise finetuning towards single language models for tr, it and es. Let us illustrate the whole approach before explaining:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd03ded105462fb5ba610393ba97e4ad8%2Fjigsaw%20arch2.001.jpeg?generation=1592909197036996&amp;alt=media)\n\nAs input data we used the english  jigsaw-toxic-comment-train.csv and its six translations available as public datasets which combined are roughly 1.4M comments. I realized that training a model directly on the combined dataset is not optimal, as each comment appears 7x in each epoch (although in different languages) and the model overfits on that. So I divided the combined dataset into 7 stratified parts, where each part contains a comment only once. For each fold we then finetuned a transformer in a 3step manner:\n\n\n- Step 1: finetune to all 7 languages for 2 epochs\n- Step 2: finetune only to the full valid.csv which only has tr, it and es\n- Step 3: 3x finetune to each language in valid resulting in 3 models\n\nWe then use the step1 model for predicting ru, the step2 model for predicting pt and fr and the respective step3 models for tr, it and es. Using the step2 model for pt and fr gave a significant boost compared to using step1 model for those due. Most likely due to the language similarity between it, es, fr and pt.\n\nWe used mainly a batchsize of 32 using gradient accumulation and a learning rate of 3e-6 with linear decay with AdamW optimizer. One thing worth mentioning is that we train on a max sequence length of 512 due to the dataloader mentioned above which sorts the data by length and then randomly serves batches of same length comments. Only the huge speed-up paired with the non padding of shorter comments made using a 512 max length reasonable. \n\nApart from one model which was built on top of a public kernel, all models were trained using pytorch and GPU.\n\n## Ensembling\nAs for ensembling we did team member individual ensembling and then combined the models of the team members and public kernels by weighted sum of rank percentile \n\n## Post-processing\nAnother key ingredient, which we luckily found on the last day is post processing of the final predictions. Although languages are derived from individual models, the competition metric is sensitive to the global rank of all predictions and not language specific. So we took care of the languages having a correct relation to each other, by shifting the predictions of each language individually. In our final submission for example we used the following factors \n\n```\ntest_df.loc[test_df[\"lang\"] == \"es\", \"toxic\"] *= 1.06\ntest_df.loc[test_df[\"lang\"] == \"fr\", \"toxic\"] *= 1.04\ntest_df.loc[test_df[\"lang\"] == \"it\", \"toxic\"] *= 0.97\ntest_df.loc[test_df[\"lang\"] == \"pt\", \"toxic\"] *= 0.96\ntest_df.loc[test_df[\"lang\"] == \"tr\", \"toxic\"] *= 0.98\n```\n\n\nWe derived the factors by matching the test prediction mean with the public LB mean for each language individually. We had already made 6 submissions for probing the test set by setting one language to 1 and the rest to 0. Then we made sure that means are aligned to give similar experimental AUC. That posprocessing moved us from 14th place to 4th place on public LB! \n\nWhile this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.\n\nThanks for reading. Questions welcome.",
      "votes": null
    },
    {
      "id": "898189",
      "postDate": "06/23/2020 11:16:22",
      "content": "<p>Post Proc ... Genius! Thank you for your exp!</p>",
      "rawMarkdown": "Post Proc ... Genius! Thank you for your exp!",
      "votes": null
    },
    {
      "id": "898201",
      "postDate": "06/23/2020 11:21:02",
      "content": "<p>Question: so, how was the architecture graph done? 👀 \nMore seriously: what hardware have you used, i.e. your own or cloud? Also, have you tried TPU (I guess not since you only mention V100)? \nGreat achievement and a lot to learn from!</p>",
      "rawMarkdown": "Question: so, how was the architecture graph done? 👀 \nMore seriously: what hardware have you used, i.e. your own or cloud? Also, have you tried TPU (I guess not since you only mention V100)? \nGreat achievement and a lot to learn from!",
      "votes": null
    },
    {
      "id": "898221",
      "postDate": "06/23/2020 11:32:46",
      "content": "<p>Thanks for sharing your great approach!\nI used the normal 1 model to predict 6 languages approach with xlmr-large models.  I got +0.008 improvement with your post processing. <br>\n|  |  public score| private score |\n| --- | --- | --- |\n|xlmr-large my blending  | 0.9487 |0.9471 |\n| xlmr-large my blending + your post processing | 0.9494|0.9479| </p>",
      "rawMarkdown": "Thanks for sharing your great approach!\nI used the normal 1 model to predict 6 languages approach with xlmr-large models.  I got +0.008 improvement with your post processing.   \n|  |  public score| private score |\n| --- | --- | --- |\n|xlmr-large my blending  | 0.9487 |0.9471 |\n| xlmr-large my blending + your post processing | 0.9494|0.9479|",
      "votes": null
    },
    {
      "id": "898225",
      "postDate": "06/23/2020 11:37:27",
      "content": "<p>Amazing work...I never thought your team would boost from 9500 to 9522 in the last day😜 . Can you share some details about how your team found this post-processing trick?</p>",
      "rawMarkdown": "Amazing work...I never thought your team would boost from 9500 to 9522 in the last day😜 . Can you share some details about how your team found this post-processing trick?",
      "votes": null
    },
    {
      "id": "898233",
      "postDate": "06/23/2020 11:46:23",
      "content": "<p>To me it made sense that global ranking matters, and especially the 'ru' prediction <em>should</em> be off as the only thing the step1 model has seen were translations from en, which I think are weaker in terms of toxicity as original russian comments. Hence we expected those predictions to be in average too low. Then we checked shifting the predictions on the validation set, and it gave a significant improvement in score. So we worked on a method to shift test predictions too...</p>",
      "rawMarkdown": "To me it made sense that global ranking matters, and especially the 'ru' prediction *should* be off as the only thing the step1 model has seen were translations from en, which I think are weaker in terms of toxicity as original russian comments. Hence we expected those predictions to be in average too low. Then we checked shifting the predictions on the validation set, and it gave a significant improvement in score. So we worked on a method to shift test predictions too...",
      "votes": null
    },
    {
      "id": "898247",
      "postDate": "06/23/2020 11:54:54",
      "content": "<p>Congrats for one week journey and anther gold medal!\nWe also find each language need calibration in the last day, but we have got no idea about the coef in only 5 submissions, we just try some lucky coef and get a small boost on post processing.\nAmazing work and thanks for detailed writeup!</p>",
      "rawMarkdown": "Congrats for one week journey and anther gold medal!\nWe also find each language need calibration in the last day, but we have got no idea about the coef in only 5 submissions, we just try some lucky coef and get a small boost on post processing.\nAmazing work and thanks for detailed writeup!",
      "votes": null
    },
    {
      "id": "898248",
      "postDate": "06/23/2020 11:55:52",
      "content": "<p>Thank you for your explanation!!!\nHope to learn more from you</p>",
      "rawMarkdown": "Thank you for your explanation!!!\nHope to learn more from you",
      "votes": null
    },
    {
      "id": "898250",
      "postDate": "06/23/2020 11:56:01",
      "content": "<p>The post-processing is real magic! My model got +0.006  and +0.011  improvement when I just multiple Spanish by 1.2 .</p>",
      "rawMarkdown": "The post-processing is real magic! My model got +0.006  and +0.011  improvement when I just multiple Spanish by 1.2 .",
      "votes": null
    },
    {
      "id": "898275",
      "postDate": "06/23/2020 12:20:12",
      "content": "<p>Congrats <a href=\"/christofhenkel\">@christofhenkel</a>, <a href=\"/cpmpml\">@cpmpml</a>, <a href=\"/aerdem4\">@aerdem4</a> and thanks for sharing a very good solution writeup.</p>",
      "rawMarkdown": "Congrats @christofhenkel, @cpmpml, @aerdem4 and thanks for sharing a very good solution writeup.",
      "votes": null
    },
    {
      "id": "898285",
      "postDate": "06/23/2020 12:23:45",
      "content": "<p>So many brilliant ideas! Thanks for sharing.</p>\n\n<p>On the ensembling, is the <code>weighted sum of rank percentile</code> the method of choice to try first, from experience, or just one from a toolbox ? </p>",
      "rawMarkdown": "So many brilliant ideas! Thanks for sharing.\n\nOn the ensembling, is the `weighted sum of rank percentile` the method of choice to try first, from experience, or just one from a toolbox ?",
      "votes": null
    },
    {
      "id": "898287",
      "postDate": "06/23/2020 12:25:04",
      "content": "<p>You forget <a href=\"/aerdem4\">@aerdem4</a> ;)</p>",
      "rawMarkdown": "You forget @aerdem4 ;)",
      "votes": null
    },
    {
      "id": "898290",
      "postDate": "06/23/2020 12:27:53",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, amended :-)</p>",
      "rawMarkdown": "cpmpml, amended :-)",
      "votes": null
    },
    {
      "id": "898292",
      "postDate": "06/23/2020 12:30:07",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, it must be the Tweet and Jigsaw exhaustion. <a href=\"/aerdem4\">@aerdem4</a> may agree :-)</p>",
      "rawMarkdown": "cpmpml, it must be the Tweet and Jigsaw exhaustion. @aerdem4 may agree :-)",
      "votes": null
    },
    {
      "id": "898373",
      "postDate": "06/23/2020 13:41:24",
      "content": "<p>Dieter and team, very discipline and intelligent move indeed.  Watching you and your team come after us from behind is akin watching a horror movie. What a great chase. Thanks very much for sharing the insights. Dr.</p>",
      "rawMarkdown": "Dieter and team, very discipline and intelligent move indeed.  Watching you and your team come after us from behind is akin watching a horror movie. What a great chase. Thanks very much for sharing the insights. Dr.",
      "votes": null
    },
    {
      "id": "898388",
      "postDate": "06/23/2020 13:50:51",
      "content": "<p>Ahmet used TPU, Dieter and I used GPU (V100).  The former used Keras/TF while the latter used Pytorch.  There may be a correlation there ;)</p>",
      "rawMarkdown": "Ahmet used TPU, Dieter and I used GPU (V100).  The former used Keras/TF while the latter used Pytorch.  There may be a correlation there ;)",
      "votes": null
    },
    {
      "id": "898529",
      "postDate": "06/23/2020 15:14:41",
      "content": "<p>Great model, congrats on building such a great model in a short time. </p>\n\n<p>Wow, nice post post processing trick. I just applied it to my final model. It increased both public and private LB by 0.008 and increased my private rank from 61st to 24th. Wow.</p>",
      "rawMarkdown": "Great model, congrats on building such a great model in a short time. \n\nWow, nice post post processing trick. I just applied it to my final model. It increased both public and private LB by 0.008 and increased my private rank from 61st to 24th. Wow.",
      "votes": null
    },
    {
      "id": "898586",
      "postDate": "06/23/2020 15:52:59",
      "content": "<p>Makes sense. :)</p>",
      "rawMarkdown": "Makes sense. :)",
      "votes": null
    },
    {
      "id": "898597",
      "postDate": "06/23/2020 16:02:47",
      "content": "<p>I am sure if you would optimize it further on public part you could get even higher. The optimal scaling factors can differ quite a bit depending on how your predictions look like.</p>",
      "rawMarkdown": "I am sure if you would optimize it further on public part you could get even higher. The optimal scaling factors can differ quite a bit depending on how your predictions look like.",
      "votes": null
    },
    {
      "id": "899129",
      "postDate": "06/24/2020 02:29:35",
      "content": "<p>Great Post-processing Guys ! </p>",
      "rawMarkdown": "Great Post-processing Guys !",
      "votes": null
    },
    {
      "id": "899200",
      "postDate": "06/24/2020 04:44:14",
      "content": "<p>Congratulations sir.</p>",
      "rawMarkdown": "Congratulations sir.",
      "votes": null
    },
    {
      "id": "899414",
      "postDate": "06/24/2020 08:18:06",
      "content": "<pre><code>Congrats and thanks for sharing!  All of the methods you've shared are so cool and effective,  which have opened my eyes!&amp;nbsp; I also tried to apply your post-processing to my final model.  It increased my private rank from 23rd to 13th！I'd never thought of such a fast but effective trick to improve score before.  That's really awesome! \nThere's so much more I need to learn.\n</code></pre>",
      "rawMarkdown": "Congrats and thanks for sharing!  All of the methods you've shared are so cool and effective,  which have opened my eyes!&nbsp; I also tried to apply your post-processing to my final model.  It increased my private rank from 23rd to 13th！I'd never thought of such a fast but effective trick to improve score before.  That's really awesome! \n    There's so much more I need to learn.",
      "votes": null
    },
    {
      "id": "899656",
      "postDate": "06/24/2020 11:08:09",
      "content": "<p>My takeway from recent deep learning competitions is that postprocessing is the best way to improve score, more than NN architecture.  I'm serious.</p>",
      "rawMarkdown": "My takeway from recent deep learning competitions is that postprocessing is the best way to improve score, more than NN architecture.  I'm serious.",
      "votes": null
    },
    {
      "id": "899668",
      "postDate": "06/24/2020 11:20:26",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> happened for both jigsaw and twitter. And suprisingly both were expected to have big shakeups and ended up having nothing.</p>",
      "rawMarkdown": "cpmpml happened for both jigsaw and twitter. And suprisingly both were expected to have big shakeups and ended up having nothing.",
      "votes": null
    },
    {
      "id": "899676",
      "postDate": "06/24/2020 11:33:56",
      "content": "<p>When everyone is working off the same set of pretrained models and knowledge of common architectural tweaks, postprocessing is one of the few ways to distinguish your models. </p>\n\n<p>The main downside with PP is that it's easy  to do it at the last minute so people can see their leads evaporate literally overnight. </p>",
      "rawMarkdown": "When everyone is working off the same set of pretrained models and knowledge of common architectural tweaks, postprocessing is one of the few ways to distinguish your models. \n\nThe main downside with PP is that it's easy  to do it at the last minute so people can see their leads evaporate literally overnight.",
      "votes": null
    },
    {
      "id": "899680",
      "postDate": "06/24/2020 11:38:00",
      "content": "<p><a href=\"/shahules\">@shahules</a> it also happened in Bengali</p>\n\n<p><a href=\"/leecming\">@leecming</a> sometimes the pp is tricky.  In Tweet sentiment I didn't found it nor did <a href=\"/philippsinger\">@philippsinger</a> for instance.</p>",
      "rawMarkdown": "shahules it also happened in Bengali\n\n@leecming sometimes the pp is tricky.  In Tweet sentiment I didn't found it nor did @philippsinger for instance.",
      "votes": null
    },
    {
      "id": "899682",
      "postDate": "06/24/2020 11:38:53",
      "content": "<p>You shoudl try many methods and see their effect.  But when roc-auc is the metric, only rank matter, hence using rank makes sense.</p>",
      "rawMarkdown": "You shoudl try many methods and see their effect.  But when roc-auc is the metric, only rank matter, hence using rank makes sense.",
      "votes": null
    },
    {
      "id": "899704",
      "postDate": "06/24/2020 11:58:00",
      "content": "<p>Yeah, I found/profited from it here, but did not in Tweets.</p>",
      "rawMarkdown": "Yeah, I found/profited from it here, but did not in Tweets.",
      "votes": null
    },
    {
      "id": "899713",
      "postDate": "06/24/2020 12:03:48",
      "content": "<p>And to be fair, if we would have had more time, I would have tried to incorporate fixing the disturbed individual language distributions in the modelling, instead of fixing it via pp. Often post-processing fixes the symptoms but does not solve the root cause. </p>",
      "rawMarkdown": "And to be fair, if we would have had more time, I would have tried to incorporate fixing the disturbed individual language distributions in the modelling, instead of fixing it via pp. Often post-processing fixes the symptoms but does not solve the root cause.",
      "votes": null
    },
    {
      "id": "899741",
      "postDate": "06/24/2020 12:24:07",
      "content": "<p>You still need to adjust for test distributions. So in both ways you are doing some form of test pp, even if it is from a different perspective.</p>",
      "rawMarkdown": "You still need to adjust for test distributions. So in both ways you are doing some form of test pp, even if it is from a different perspective.",
      "votes": null
    },
    {
      "id": "899747",
      "postDate": "06/24/2020 12:28:51",
      "content": "<p>We wouldn't postprocess without LB probing. If Kaggle splits the public/private not random, it could be much different. In the past, Kaggle had more tricky splits.</p>",
      "rawMarkdown": "We wouldn't postprocess without LB probing. If Kaggle splits the public/private not random, it could be much different. In the past, Kaggle had more tricky splits.",
      "votes": null
    },
    {
      "id": "899884",
      "postDate": "06/24/2020 13:56:34",
      "content": "<p>Yes, I wanted to make a longer post about it but pub LB should probably not have included all languages, and private test should have been completely hidden. I still dont understand the technical argument of why TPUs need to have the full test data available.</p>",
      "rawMarkdown": "Yes, I wanted to make a longer post about it but pub LB should probably not have included all languages, and private test should have been completely hidden. I still dont understand the technical argument of why TPUs need to have the full test data available.",
      "votes": null
    },
    {
      "id": "904323",
      "postDate": "06/27/2020 14:31:15",
      "content": "<p>I got 0.9499 on private LB by just applying the PP trick to my (not chosen) best model ( from 0.9478  to 0.9499 )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F974295%2Fb81459e076acd5500bd0df2ce56ea854%2F11Capture.PNG?generation=1593268260582472&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I got 0.9499 on private LB by just applying the PP trick to my (not chosen) best model ( from 0.9478  to 0.9499 )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F974295%2Fb81459e076acd5500bd0df2ce56ea854%2F11Capture.PNG?generation=1593268260582472&amp;alt=media)",
      "votes": null
    },
    {
      "id": "904747",
      "postDate": "06/27/2020 21:44:51",
      "content": "<p>yeah, its a good pp trick :D</p>",
      "rawMarkdown": "yeah, its a good pp trick :D",
      "votes": null
    },
    {
      "id": "910784",
      "postDate": "07/01/2020 11:22:15",
      "content": "<p>Congratulation! Super solid solution and really loved it, may I ask you how you implemented the padding by batch trick in detail? It would be great if you can share any links for refs, sort of new to NLP and BERT so need some help. By the way around the last 2 days, we were at 14th and then you guys jumped up over us, it makes losing a gold medal not that suffering if it was you guys who outplayed our team LOL, thanks for sharing your cool ideas.</p>",
      "rawMarkdown": "Congratulation! Super solid solution and really loved it, may I ask you how you implemented the padding by batch trick in detail? It would be great if you can share any links for refs, sort of new to NLP and BERT so need some help. By the way around the last 2 days, we were at 14th and then you guys jumped up over us, it makes losing a gold medal not that suffering if it was you guys who outplayed our team LOL, thanks for sharing your cool ideas.",
      "votes": null
    },
    {
      "id": "916408",
      "postDate": "07/05/2020 15:50:36",
      "content": "<p>Congrats! \nAnd sorry for late asking.\nI'm still confused about the way from 6 single language submissions to pp factors, could you sharing more details?</p>",
      "rawMarkdown": "Congrats! \nAnd sorry for late asking.\nI'm still confused about the way from 6 single language submissions to pp factors, could you sharing more details?",
      "votes": null
    },
    {
      "id": "1495839",
      "postDate": "08/29/2021 20:37:19",
      "content": "<p>Thanks for sharing!! very helpful</p>",
      "rawMarkdown": "Thanks for sharing!! very helpful",
      "votes": null
    },
    {
      "id": "1574214",
      "postDate": "11/07/2021 10:43:22",
      "content": "<p>Is your BERT model able to detect negation and sarcasm?</p>",
      "rawMarkdown": "Is your BERT model able to detect negation and sarcasm?",
      "votes": null
    },
    {
      "id": "1590621",
      "postDate": "11/21/2021 13:37:26",
      "content": "<p>日本語訳です<br>\nBrief Summary<br>\n私たちのソリューションは、いくつかの後処理と組み合わせたいくつかのトランスモデル（主にxlm-roberta-large）の単純なブレンドです。 アーキテクチャは公開されているカーネルと同じで、分類ヘッドは非表示状態の最大+平均プーリングまたはCLSトークンの非表示状態のいずれかを取ります。 モデルのトレーニングには3段階のアプローチを使用しました。ここでは、7つの言語から始めて、3つの言語に微調整し、1つの言語で終了しました。 この3ステップのアプローチではグローバル予測が歪むため、後処理ステップとして言語を1つの要因で個別にシフトしました。</p>\n<p>Detailed Summary<br>\nこんなに短い時間で素晴らしい結果が得られたことに、私たちは今でも驚いています。 @ aerdem4はこのコンテストでもう少し長く働き、特定の課題をすでによく理解していましたが、@ cpmpmlと@christofhenkelは、チームの合併期限の数時間前のツイート感情コンテストの直後に参加しました。 2日前、私たちはまだ100以上の場所にいて、せいぜい銀メダルを望んでいました。 しかし、粘り強さと適切な量の直感で、どのアイデアを進めるか、そしてもちろん運が良ければ、私たちは金の上位に登ることができました。</p>\n<p>Preprocessing<br>\nNone. That's NLP in 2020 :D</p>\n<p>Speeding up training<br>\n私たちは2つのことをしました。 1つは、ネガティブサンプルをダウンサンプリングして、よりバランスの取れたデータセットとより小さなデータセットを同時に取得することでした。 ポジティブとネガティブの数だけを使用する場合、データセットは約5分の1に削減されます。これは、1つのエポックを実行する時間と同じです。 私たちのモデルのいくつかはそのように訓練されました。</p>\n<p>もう1つのスピードアップは、バッチによるパディングによるものです。 主なアイデアは、パディングの量を現在のバッチに必要な量に制限することです。 入力は、固定長（512など）ではなく、バッチ内の最長入力の長さにパディングされます。これは現在よく知られている手法であり、以前のコンテストで使用されています。 トレーニングを大幅に加速します。<br>\n特定のバッチのサンプルの長さが同じになるように、サンプルを長さで並べ替えることでアイデアを洗練しました。 これにより、パディングの必要性がさらに減少します。 バッチ内のすべての入力の長さが同じである場合、パディングはまったくありません。</p>\n<p>サンプルがソートされている場合、トレーニングモードでそれらをシャッフルすることはできません。 むしろバッチをシャッフルしました。 これにより、バッチによる元のパディングと比較して、さらに2倍のスピードアップが得られます。 最初のトレインセットでxlm-roberta-largeの1つのエポックをトレーニングするには、V100GPUで約17分かかります。</p>\n<p>Cross-validation<br>\n時間が短く、LBフィードバックを取得するための提出物があまりなかったため、信頼性の高い相互検証の取得に取り組みました。 未知の言語ru、pt、fr、および既知の言語it、es、trを表すために、言語ごとに3倍のグループと検証データの単純な3倍の組み合わせを使用しました。 その結果、6倍のスキームになります。 したがって、たとえばfold1は、esがvalid.csvにない場合、ptが無効であるのと同じように動作するはずです。 すべてのフォールドの平均AUCを取ることは、パブリックおよびプライベートLBの非常に優れたプロキシでした。</p>\n<p>Architectures<br>\nほとんどのチームとして、max + meanプーリングまたはCLSトークンの非表示状態のいずれかを使用する単純な分類ヘッドを備えたメインバックボーンとしてxlm-robertalargeを使用しました。合計で、次のバックボーンを使用し、括弧内の最終的なアンサンブルでおおよその重みを付けました。</p>\n<p>5x xlm-roberta-large（85％）<br>\n1x xlm-roberta-base（5％）<br>\n1x mBart-large（10％）<br>\nトレーニング戦略<br>\n私たちの良い結果の主な要素の1つは、tr、it、およびesの単一言語モデルに向けた段階的な微調整です。説明する前に、アプローチ全体を説明しましょう。</p>\n<p>入力データとして、英語のjigsaw-toxic-comment-train.csvを使用し、公開データセットとして利用可能な6つの翻訳を組み合わせると、約140万件のコメントになります。結合されたデータセットでモデルを直接トレーニングすることは最適ではないことに気付きました。各コメントは各エポックで7倍に表示され（言語は異なりますが）、モデルはそれに適合しすぎます。そこで、結合したデータセットを7つの階層化された部分に分割しました。各部分には、コメントが1回だけ含まれています。次に、折り畳みごとに、トランスフォーマーを3段階で微調整しました。</p>\n<p>ステップ1：2つのエポックで7つの言語すべてに微調整する<br>\nステップ2：tr、it、esのみを含む完全なvalid.csvにのみ微調整します<br>\nステップ3：有効な各言語への3倍の微調整により、3つのモデルが作成されます<br>\n次に、ruを予測するためのstep1モデル、ptとfrを予測するためのstep2モデル、およびtr、it、esのそれぞれのstep3モデルを使用します。 ptとfrにstep2モデルを使用すると、期限のあるものにstep1モデルを使用する場合に比べて大幅に向上しました。おそらく、es、fr、ptの間の言語の類似性が原因です。</p>\n<p>勾配累積を使用した32のバッチサイズと、AdamWオプティマイザーによる線形減衰を伴う3e-6の学習率を主に使用しました。言及する価値のあることの1つは、データを長さでソートし、同じ長さのコメントのバッチをランダムに提供する上記のデータローダーのために、最大シーケンス長512でトレーニングすることです。最大512の長さを使用して作成された短いコメントの非パディングと組み合わせた大幅なスピードアップのみが妥当です。</p>\n<p>パブリックカーネル上に構築された1つのモデルを除いて、すべてのモデルはpytorchとGPUを使用してトレーニングされました。</p>\n<p>Ensembling<br>\nアンサンブルについては、チームメンバーの個別アンサンブルを行い、ランクパーセンタイルの加重和でチームメンバーとパブリックカーネルのモデルを組み合わせました。</p>\n<p>Post-processing<br>\n最終日に幸運にも見つけたもう1つの重要な要素は、最終予測の後処理です。言語は個々のモデルから派生していますが、競合指標はすべての予測のグローバルランクに敏感であり、言語固有ではありません。そこで、各言語の予測を個別にシフトすることで、相互に正しい関係を持つ言語を処理しました。たとえば、最終的な提出では、次の要素を使用しました</p>\n<p>test_df.loc [test_df [\"lang\"] == \"es\"、 \"toxic\"] * = 1.06<br>\ntest_df.loc [test_df [\"lang\"] == \"fr\"、 \"toxic\"] * = 1.04<br>\ntest_df.loc [test_df [\"lang\"] == \"it\"、 \"toxic\"] * = 0.97<br>\ntest_df.loc [test_df [\"lang\"] == \"pt\"、 \"toxic\"] * = 0.96<br>\ntest_df.loc [test_df [\"lang\"] == \"tr\"、 \"toxic\"] * = 0.98<br>\nテスト予測平均を各言語の公開LB平均と個別に照合することにより、因子を導き出しました。 1つの言語を1に設定し、残りを0に設定して、テストセットを調査するために、すでに6つの提出を行いました。次に、同様の実験的なAUCが得られるように平均が調整されていることを確認しました。そのposprocessingは私達を公共LBの14位から4位に動かしました！</p>\n<p>この後処理は他のチームにも役立つ可能性がありますが、6つの言語を予測するために5つのモデルを持つことによって導入されたグローバル予測分布の相違を具体的に修正し、チームが単一のモデルアプローチを使用した場合はあまり役に立たない可能性があります。</p>\n<p>Thanks for reading. Questions welcome.</p>",
      "rawMarkdown": "日本語訳です\nBrief Summary\n私たちのソリューションは、いくつかの後処理と組み合わせたいくつかのトランスモデル（主にxlm-roberta-large）の単純なブレンドです。 アーキテクチャは公開されているカーネルと同じで、分類ヘッドは非表示状態の最大+平均プーリングまたはCLSトークンの非表示状態のいずれかを取ります。 モデルのトレーニングには3段階のアプローチを使用しました。ここでは、7つの言語から始めて、3つの言語に微調整し、1つの言語で終了しました。 この3ステップのアプローチではグローバル予測が歪むため、後処理ステップとして言語を1つの要因で個別にシフトしました。\n\nDetailed Summary\nこんなに短い時間で素晴らしい結果が得られたことに、私たちは今でも驚いています。 @ aerdem4はこのコンテストでもう少し長く働き、特定の課題をすでによく理解していましたが、@ cpmpmlと@christofhenkelは、チームの合併期限の数時間前のツイート感情コンテストの直後に参加しました。 2日前、私たちはまだ100以上の場所にいて、せいぜい銀メダルを望んでいました。 しかし、粘り強さと適切な量の直感で、どのアイデアを進めるか、そしてもちろん運が良ければ、私たちは金の上位に登ることができました。\n\nPreprocessing\nNone. That's NLP in 2020 :D\n\nSpeeding up training\n私たちは2つのことをしました。 1つは、ネガティブサンプルをダウンサンプリングして、よりバランスの取れたデータセットとより小さなデータセットを同時に取得することでした。 ポジティブとネガティブの数だけを使用する場合、データセットは約5分の1に削減されます。これは、1つのエポックを実行する時間と同じです。 私たちのモデルのいくつかはそのように訓練されました。\n\nもう1つのスピードアップは、バッチによるパディングによるものです。 主なアイデアは、パディングの量を現在のバッチに必要な量に制限することです。 入力は、固定長（512など）ではなく、バッチ内の最長入力の長さにパディングされます。これは現在よく知られている手法であり、以前のコンテストで使用されています。 トレーニングを大幅に加速します。\n特定のバッチのサンプルの長さが同じになるように、サンプルを長さで並べ替えることでアイデアを洗練しました。 これにより、パディングの必要性がさらに減少します。 バッチ内のすべての入力の長さが同じである場合、パディングはまったくありません。\n\nサンプルがソートされている場合、トレーニングモードでそれらをシャッフルすることはできません。 むしろバッチをシャッフルしました。 これにより、バッチによる元のパディングと比較して、さらに2倍のスピードアップが得られます。 最初のトレインセットでxlm-roberta-largeの1つのエポックをトレーニングするには、V100GPUで約17分かかります。\n\nCross-validation\n時間が短く、LBフィードバックを取得するための提出物があまりなかったため、信頼性の高い相互検証の取得に取り組みました。 未知の言語ru、pt、fr、および既知の言語it、es、trを表すために、言語ごとに3倍のグループと検証データの単純な3倍の組み合わせを使用しました。 その結果、6倍のスキームになります。 したがって、たとえばfold1は、esがvalid.csvにない場合、ptが無効であるのと同じように動作するはずです。 すべてのフォールドの平均AUCを取ることは、パブリックおよびプライベートLBの非常に優れたプロキシでした。\n\nArchitectures\nほとんどのチームとして、max + meanプーリングまたはCLSトークンの非表示状態のいずれかを使用する単純な分類ヘッドを備えたメインバックボーンとしてxlm-robertalargeを使用しました。合計で、次のバックボーンを使用し、括弧内の最終的なアンサンブルでおおよその重みを付けました。\n\n5x xlm-roberta-large（85％）\n1x xlm-roberta-base（5％）\n1x mBart-large（10％）\nトレーニング戦略\n私たちの良い結果の主な要素の1つは、tr、it、およびesの単一言語モデルに向けた段階的な微調整です。説明する前に、アプローチ全体を説明しましょう。\n\n\n\n入力データとして、英語のjigsaw-toxic-comment-train.csvを使用し、公開データセットとして利用可能な6つの翻訳を組み合わせると、約140万件のコメントになります。結合されたデータセットでモデルを直接トレーニングすることは最適ではないことに気付きました。各コメントは各エポックで7倍に表示され（言語は異なりますが）、モデルはそれに適合しすぎます。そこで、結合したデータセットを7つの階層化された部分に分割しました。各部分には、コメントが1回だけ含まれています。次に、折り畳みごとに、トランスフォーマーを3段階で微調整しました。\n\nステップ1：2つのエポックで7つの言語すべてに微調整する\nステップ2：tr、it、esのみを含む完全なvalid.csvにのみ微調整します\nステップ3：有効な各言語への3倍の微調整により、3つのモデルが作成されます\n次に、ruを予測するためのstep1モデル、ptとfrを予測するためのstep2モデル、およびtr、it、esのそれぞれのstep3モデルを使用します。 ptとfrにstep2モデルを使用すると、期限のあるものにstep1モデルを使用する場合に比べて大幅に向上しました。おそらく、es、fr、ptの間の言語の類似性が原因です。\n\n勾配累積を使用した32のバッチサイズと、AdamWオプティマイザーによる線形減衰を伴う3e-6の学習率を主に使用しました。言及する価値のあることの1つは、データを長さでソートし、同じ長さのコメントのバッチをランダムに提供する上記のデータローダーのために、最大シーケンス長512でトレーニングすることです。最大512の長さを使用して作成された短いコメントの非パディングと組み合わせた大幅なスピードアップのみが妥当です。\n\nパブリックカーネル上に構築された1つのモデルを除いて、すべてのモデルはpytorchとGPUを使用してトレーニングされました。\n\nEnsembling\nアンサンブルについては、チームメンバーの個別アンサンブルを行い、ランクパーセンタイルの加重和でチームメンバーとパブリックカーネルのモデルを組み合わせました。\n\nPost-processing\n最終日に幸運にも見つけたもう1つの重要な要素は、最終予測の後処理です。言語は個々のモデルから派生していますが、競合指標はすべての予測のグローバルランクに敏感であり、言語固有ではありません。そこで、各言語の予測を個別にシフトすることで、相互に正しい関係を持つ言語を処理しました。たとえば、最終的な提出では、次の要素を使用しました\n\ntest_df.loc [test_df [\"lang\"] == \"es\"、 \"toxic\"] * = 1.06\ntest_df.loc [test_df [\"lang\"] == \"fr\"、 \"toxic\"] * = 1.04\ntest_df.loc [test_df [\"lang\"] == \"it\"、 \"toxic\"] * = 0.97\ntest_df.loc [test_df [\"lang\"] == \"pt\"、 \"toxic\"] * = 0.96\ntest_df.loc [test_df [\"lang\"] == \"tr\"、 \"toxic\"] * = 0.98\nテスト予測平均を各言語の公開LB平均と個別に照合することにより、因子を導き出しました。 1つの言語を1に設定し、残りを0に設定して、テストセットを調査するために、すでに6つの提出を行いました。次に、同様の実験的なAUCが得られるように平均が調整されていることを確認しました。そのposprocessingは私達を公共LBの14位から4位に動かしました！\n\nこの後処理は他のチームにも役立つ可能性がありますが、6つの言語を予測するために5つのモデルを持つことによって導入されたグローバル予測分布の相違を具体的に修正し、チームが単一のモデルアプローチを使用した場合はあまり役に立たない可能性があります。\n\nThanks for reading. Questions welcome.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1495839,
      "author_name": "top10chi3nthan",
      "author_url": "",
      "post_date": "08/29/2021 20:37:19",
      "content": "<p>Thanks for sharing!! very helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1574214,
      "author_name": "shubheshswain",
      "author_url": "",
      "post_date": "11/07/2021 10:43:22",
      "content": "<p>Is your BERT model able to detect negation and sarcasm?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1590621,
      "author_name": "pixyz0130",
      "author_url": "",
      "post_date": "11/21/2021 13:37:26",
      "content": "<p>日本語訳です<br>\nBrief Summary<br>\n私たちのソリューションは、いくつかの後処理と組み合わせたいくつかのトランスモデル（主にxlm-roberta-large）の単純なブレンドです。 アーキテクチャは公開されているカーネルと同じで、分類ヘッドは非表示状態の最大+平均プーリングまたはCLSトークンの非表示状態のいずれかを取ります。 モデルのトレーニングには3段階のアプローチを使用しました。ここでは、7つの言語から始めて、3つの言語に微調整し、1つの言語で終了しました。 この3ステップのアプローチではグローバル予測が歪むため、後処理ステップとして言語を1つの要因で個別にシフトしました。</p>\n<p>Detailed Summary<br>\nこんなに短い時間で素晴らしい結果が得られたことに、私たちは今でも驚いています。 @ aerdem4はこのコンテストでもう少し長く働き、特定の課題をすでによく理解していましたが、@ cpmpmlと@christofhenkelは、チームの合併期限の数時間前のツイート感情コンテストの直後に参加しました。 2日前、私たちはまだ100以上の場所にいて、せいぜい銀メダルを望んでいました。 しかし、粘り強さと適切な量の直感で、どのアイデアを進めるか、そしてもちろん運が良ければ、私たちは金の上位に登ることができました。</p>\n<p>Preprocessing<br>\nNone. That's NLP in 2020 :D</p>\n<p>Speeding up training<br>\n私たちは2つのことをしました。 1つは、ネガティブサンプルをダウンサンプリングして、よりバランスの取れたデータセットとより小さなデータセットを同時に取得することでした。 ポジティブとネガティブの数だけを使用する場合、データセットは約5分の1に削減されます。これは、1つのエポックを実行する時間と同じです。 私たちのモデルのいくつかはそのように訓練されました。</p>\n<p>もう1つのスピードアップは、バッチによるパディングによるものです。 主なアイデアは、パディングの量を現在のバッチに必要な量に制限することです。 入力は、固定長（512など）ではなく、バッチ内の最長入力の長さにパディングされます。これは現在よく知られている手法であり、以前のコンテストで使用されています。 トレーニングを大幅に加速します。<br>\n特定のバッチのサンプルの長さが同じになるように、サンプルを長さで並べ替えることでアイデアを洗練しました。 これにより、パディングの必要性がさらに減少します。 バッチ内のすべての入力の長さが同じである場合、パディングはまったくありません。</p>\n<p>サンプルがソートされている場合、トレーニングモードでそれらをシャッフルすることはできません。 むしろバッチをシャッフルしました。 これにより、バッチによる元のパディングと比較して、さらに2倍のスピードアップが得られます。 最初のトレインセットでxlm-roberta-largeの1つのエポックをトレーニングするには、V100GPUで約17分かかります。</p>\n<p>Cross-validation<br>\n時間が短く、LBフィードバックを取得するための提出物があまりなかったため、信頼性の高い相互検証の取得に取り組みました。 未知の言語ru、pt、fr、および既知の言語it、es、trを表すために、言語ごとに3倍のグループと検証データの単純な3倍の組み合わせを使用しました。 その結果、6倍のスキームになります。 したがって、たとえばfold1は、esがvalid.csvにない場合、ptが無効であるのと同じように動作するはずです。 すべてのフォールドの平均AUCを取ることは、パブリックおよびプライベートLBの非常に優れたプロキシでした。</p>\n<p>Architectures<br>\nほとんどのチームとして、max + meanプーリングまたはCLSトークンの非表示状態のいずれかを使用する単純な分類ヘッドを備えたメインバックボーンとしてxlm-robertalargeを使用しました。合計で、次のバックボーンを使用し、括弧内の最終的なアンサンブルでおおよその重みを付けました。</p>\n<p>5x xlm-roberta-large（85％）<br>\n1x xlm-roberta-base（5％）<br>\n1x mBart-large（10％）<br>\nトレーニング戦略<br>\n私たちの良い結果の主な要素の1つは、tr、it、およびesの単一言語モデルに向けた段階的な微調整です。説明する前に、アプローチ全体を説明しましょう。</p>\n<p>入力データとして、英語のjigsaw-toxic-comment-train.csvを使用し、公開データセットとして利用可能な6つの翻訳を組み合わせると、約140万件のコメントになります。結合されたデータセットでモデルを直接トレーニングすることは最適ではないことに気付きました。各コメントは各エポックで7倍に表示され（言語は異なりますが）、モデルはそれに適合しすぎます。そこで、結合したデータセットを7つの階層化された部分に分割しました。各部分には、コメントが1回だけ含まれています。次に、折り畳みごとに、トランスフォーマーを3段階で微調整しました。</p>\n<p>ステップ1：2つのエポックで7つの言語すべてに微調整する<br>\nステップ2：tr、it、esのみを含む完全なvalid.csvにのみ微調整します<br>\nステップ3：有効な各言語への3倍の微調整により、3つのモデルが作成されます<br>\n次に、ruを予測するためのstep1モデル、ptとfrを予測するためのstep2モデル、およびtr、it、esのそれぞれのstep3モデルを使用します。 ptとfrにstep2モデルを使用すると、期限のあるものにstep1モデルを使用する場合に比べて大幅に向上しました。おそらく、es、fr、ptの間の言語の類似性が原因です。</p>\n<p>勾配累積を使用した32のバッチサイズと、AdamWオプティマイザーによる線形減衰を伴う3e-6の学習率を主に使用しました。言及する価値のあることの1つは、データを長さでソートし、同じ長さのコメントのバッチをランダムに提供する上記のデータローダーのために、最大シーケンス長512でトレーニングすることです。最大512の長さを使用して作成された短いコメントの非パディングと組み合わせた大幅なスピードアップのみが妥当です。</p>\n<p>パブリックカーネル上に構築された1つのモデルを除いて、すべてのモデルはpytorchとGPUを使用してトレーニングされました。</p>\n<p>Ensembling<br>\nアンサンブルについては、チームメンバーの個別アンサンブルを行い、ランクパーセンタイルの加重和でチームメンバーとパブリックカーネルのモデルを組み合わせました。</p>\n<p>Post-processing<br>\n最終日に幸運にも見つけたもう1つの重要な要素は、最終予測の後処理です。言語は個々のモデルから派生していますが、競合指標はすべての予測のグローバルランクに敏感であり、言語固有ではありません。そこで、各言語の予測を個別にシフトすることで、相互に正しい関係を持つ言語を処理しました。たとえば、最終的な提出では、次の要素を使用しました</p>\n<p>test_df.loc [test_df [\"lang\"] == \"es\"、 \"toxic\"] * = 1.06<br>\ntest_df.loc [test_df [\"lang\"] == \"fr\"、 \"toxic\"] * = 1.04<br>\ntest_df.loc [test_df [\"lang\"] == \"it\"、 \"toxic\"] * = 0.97<br>\ntest_df.loc [test_df [\"lang\"] == \"pt\"、 \"toxic\"] * = 0.96<br>\ntest_df.loc [test_df [\"lang\"] == \"tr\"、 \"toxic\"] * = 0.98<br>\nテスト予測平均を各言語の公開LB平均と個別に照合することにより、因子を導き出しました。 1つの言語を1に設定し、残りを0に設定して、テストセットを調査するために、すでに6つの提出を行いました。次に、同様の実験的なAUCが得られるように平均が調整されていることを確認しました。そのposprocessingは私達を公共LBの14位から4位に動かしました！</p>\n<p>この後処理は他のチームにも役立つ可能性がありますが、6つの言語を予測するために5つのモデルを持つことによって導入されたグローバル予測分布の相違を具体的に修正し、チームが単一のモデルアプローチを使用した場合はあまり役に立たない可能性があります。</p>\n<p>Thanks for reading. Questions welcome.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898189,
      "author_name": "andrilko",
      "author_url": "",
      "post_date": "06/23/2020 11:16:22",
      "content": "<p>Post Proc ... Genius! Thank you for your exp!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898201,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/23/2020 11:21:02",
      "content": "<p>Question: so, how was the architecture graph done? 👀 \nMore seriously: what hardware have you used, i.e. your own or cloud? Also, have you tried TPU (I guess not since you only mention V100)? \nGreat achievement and a lot to learn from!</p>",
      "votes": null,
      "replies": [
        {
          "id": 898388,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/23/2020 13:50:51",
          "content": "<p>Ahmet used TPU, Dieter and I used GPU (V100).  The former used Keras/TF while the latter used Pytorch.  There may be a correlation there ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898586,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/23/2020 15:52:59",
          "content": "<p>Makes sense. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898221,
      "author_name": "sai11fkaneko",
      "author_url": "",
      "post_date": "06/23/2020 11:32:46",
      "content": "<p>Thanks for sharing your great approach!\nI used the normal 1 model to predict 6 languages approach with xlmr-large models.  I got +0.008 improvement with your post processing. <br>\n|  |  public score| private score |\n| --- | --- | --- |\n|xlmr-large my blending  | 0.9487 |0.9471 |\n| xlmr-large my blending + your post processing | 0.9494|0.9479| </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898225,
      "author_name": "tonyxu",
      "author_url": "",
      "post_date": "06/23/2020 11:37:27",
      "content": "<p>Amazing work...I never thought your team would boost from 9500 to 9522 in the last day😜 . Can you share some details about how your team found this post-processing trick?</p>",
      "votes": null,
      "replies": [
        {
          "id": 898233,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/23/2020 11:46:23",
          "content": "<p>To me it made sense that global ranking matters, and especially the 'ru' prediction <em>should</em> be off as the only thing the step1 model has seen were translations from en, which I think are weaker in terms of toxicity as original russian comments. Hence we expected those predictions to be in average too low. Then we checked shifting the predictions on the validation set, and it gave a significant improvement in score. So we worked on a method to shift test predictions too...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898248,
          "author_name": "tonyxu",
          "author_url": "",
          "post_date": "06/23/2020 11:55:52",
          "content": "<p>Thank you for your explanation!!!\nHope to learn more from you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898247,
      "author_name": "mzr2017",
      "author_url": "",
      "post_date": "06/23/2020 11:54:54",
      "content": "<p>Congrats for one week journey and anther gold medal!\nWe also find each language need calibration in the last day, but we have got no idea about the coef in only 5 submissions, we just try some lucky coef and get a small boost on post processing.\nAmazing work and thanks for detailed writeup!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898250,
      "author_name": "medrau",
      "author_url": "",
      "post_date": "06/23/2020 11:56:01",
      "content": "<p>The post-processing is real magic! My model got +0.006  and +0.011  improvement when I just multiple Spanish by 1.2 .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898275,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/23/2020 12:20:12",
      "content": "<p>Congrats <a href=\"/christofhenkel\">@christofhenkel</a>, <a href=\"/cpmpml\">@cpmpml</a>, <a href=\"/aerdem4\">@aerdem4</a> and thanks for sharing a very good solution writeup.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898287,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/23/2020 12:25:04",
          "content": "<p>You forget <a href=\"/aerdem4\">@aerdem4</a> ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898290,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "06/23/2020 12:27:53",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, amended :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898292,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "06/23/2020 12:30:07",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, it must be the Tweet and Jigsaw exhaustion. <a href=\"/aerdem4\">@aerdem4</a> may agree :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898285,
      "author_name": "sebastienm",
      "author_url": "",
      "post_date": "06/23/2020 12:23:45",
      "content": "<p>So many brilliant ideas! Thanks for sharing.</p>\n\n<p>On the ensembling, is the <code>weighted sum of rank percentile</code> the method of choice to try first, from experience, or just one from a toolbox ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 899682,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/24/2020 11:38:53",
          "content": "<p>You shoudl try many methods and see their effect.  But when roc-auc is the metric, only rank matter, hence using rank makes sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898373,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "06/23/2020 13:41:24",
      "content": "<p>Dieter and team, very discipline and intelligent move indeed.  Watching you and your team come after us from behind is akin watching a horror movie. What a great chase. Thanks very much for sharing the insights. Dr.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898529,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/23/2020 15:14:41",
      "content": "<p>Great model, congrats on building such a great model in a short time. </p>\n\n<p>Wow, nice post post processing trick. I just applied it to my final model. It increased both public and private LB by 0.008 and increased my private rank from 61st to 24th. Wow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898597,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/23/2020 16:02:47",
          "content": "<p>I am sure if you would optimize it further on public part you could get even higher. The optimal scaling factors can differ quite a bit depending on how your predictions look like.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899656,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/24/2020 11:08:09",
          "content": "<p>My takeway from recent deep learning competitions is that postprocessing is the best way to improve score, more than NN architecture.  I'm serious.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899668,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "06/24/2020 11:20:26",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> happened for both jigsaw and twitter. And suprisingly both were expected to have big shakeups and ended up having nothing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899676,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "06/24/2020 11:33:56",
          "content": "<p>When everyone is working off the same set of pretrained models and knowledge of common architectural tweaks, postprocessing is one of the few ways to distinguish your models. </p>\n\n<p>The main downside with PP is that it's easy  to do it at the last minute so people can see their leads evaporate literally overnight. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899680,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/24/2020 11:38:00",
          "content": "<p><a href=\"/shahules\">@shahules</a> it also happened in Bengali</p>\n\n<p><a href=\"/leecming\">@leecming</a> sometimes the pp is tricky.  In Tweet sentiment I didn't found it nor did <a href=\"/philippsinger\">@philippsinger</a> for instance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899704,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/24/2020 11:58:00",
          "content": "<p>Yeah, I found/profited from it here, but did not in Tweets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899713,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/24/2020 12:03:48",
          "content": "<p>And to be fair, if we would have had more time, I would have tried to incorporate fixing the disturbed individual language distributions in the modelling, instead of fixing it via pp. Often post-processing fixes the symptoms but does not solve the root cause. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899741,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/24/2020 12:24:07",
          "content": "<p>You still need to adjust for test distributions. So in both ways you are doing some form of test pp, even if it is from a different perspective.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899747,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "06/24/2020 12:28:51",
          "content": "<p>We wouldn't postprocess without LB probing. If Kaggle splits the public/private not random, it could be much different. In the past, Kaggle had more tricky splits.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899884,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/24/2020 13:56:34",
          "content": "<p>Yes, I wanted to make a longer post about it but pub LB should probably not have included all languages, and private test should have been completely hidden. I still dont understand the technical argument of why TPUs need to have the full test data available.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 899129,
      "author_name": "mahdhiashraf",
      "author_url": "",
      "post_date": "06/24/2020 02:29:35",
      "content": "<p>Great Post-processing Guys ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899200,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "06/24/2020 04:44:14",
      "content": "<p>Congratulations sir.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899414,
      "author_name": "xj609210970",
      "author_url": "",
      "post_date": "06/24/2020 08:18:06",
      "content": "<pre><code>Congrats and thanks for sharing!  All of the methods you've shared are so cool and effective,  which have opened my eyes!&amp;nbsp; I also tried to apply your post-processing to my final model.  It increased my private rank from 23rd to 13th！I'd never thought of such a fast but effective trick to improve score before.  That's really awesome! \nThere's so much more I need to learn.\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 904323,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "06/27/2020 14:31:15",
      "content": "<p>I got 0.9499 on private LB by just applying the PP trick to my (not chosen) best model ( from 0.9478  to 0.9499 )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F974295%2Fb81459e076acd5500bd0df2ce56ea854%2F11Capture.PNG?generation=1593268260582472&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 904747,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/27/2020 21:44:51",
          "content": "<p>yeah, its a good pp trick :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 910784,
      "author_name": "",
      "author_url": "",
      "post_date": "07/01/2020 11:22:15",
      "content": "<p>Congratulation! Super solid solution and really loved it, may I ask you how you implemented the padding by batch trick in detail? It would be great if you can share any links for refs, sort of new to NLP and BERT so need some help. By the way around the last 2 days, we were at 14th and then you guys jumped up over us, it makes losing a gold medal not that suffering if it was you guys who outplayed our team LOL, thanks for sharing your cool ideas.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 916408,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "07/05/2020 15:50:36",
      "content": "<p>Congrats! \nAnd sorry for late asking.\nI'm still confused about the way from 6 single language submissions to pp factors, could you sharing more details?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "898155": "Thanks to kaggle for hosting such an interesting and challenging competition! Multilingual NLP is something very interesting yet difficult. In hindsight I (Dieter) wish I’d spend less time on the tweet sentiment extraction competition and more on this one. \n\n# Brief Summary\nOur solution is a simple blend of several transformer models (mainly xlm-roberta-large) paired with some post-processing. The architecture was the same as in public available kernels with the classification head taking either max+mean pooling of hidden states or the hidden state of the CLS token. We used a 3-step approach for training our models, where starting from 7 languages we fine-tuned to 3 languages and finished at a single language. As this 3-step approach results in distortion of global predictions we shifted languages individually by a factor as a post-processing step. \n\n# Detailed Summary\nWe are still astonished by the great result in such a short time. While @aerdem4 worked a bit longer on this competition and already had a good understanding of specific challenges @cpmpml and @christofhenkel joined right after the tweet sentiment competition a few hours before team merger deadline. Two days ago we were still at a 100+ spot and were hoping for a silver medal at best. But with persistence and the right amount of intuition what ideas to proceed with, and of course some luck we managed to climb right to the upper gold position.\n\n## Preprocessing\nNone. That's NLP in 2020 :D\n\n## Speeding up training\nWe did two things.  One was to downsample negative samples to get a more balanced dataset and a smaller dataset at the same time.  If you only use as many negative as positive then dataset is reduced 5x roughly, same for time to run one epoch.  Some of our models were trained that way.\n\nAnother speedup came from padding by batch.  The main idea is to limit the amount of padding to what is necessary for the current batch.  Inputs are padded to the length of the longest input in the batch rather than a fixed length, say 512.  This is now a well known technique, and it has been used in previous competitions.  It accelerates training significantly.  \nWe refined the idea by sorting the samples by their length so that the samples in a given batch have similar length.  This reduces even further the need for padding.  If all inputs in the batch have the same length then there is no padding at all.  \n\nGiven samples are sorted, we cannot shuffle them in training mode.  We rather shuffled batches. This yields an extra 2x speedup compared to the original padding by batch.  Training one epoch for xlm-roberta-large on the first train set takes about 17 minutes on a V100 GPU.\n\n## Cross-validation\nAs time was short and hence we did not have many submissions to spare for getting LB feedback we worked on getting a reliable cross-validation. We used a mix of Group-3fold per language and simple 3fold of the validation data to represent the unknown languages ru,pt,fr as well as the known languages it, es and tr. which results in a 6fold scheme. So do for example fold1 represent what if es would not be in valid.csv which should behave the same as pt is not in valid. Taking the mean AUC of all folds was a very good proxy for Public and Private LB.\n\n## Architectures\nAs most teams we used xlm-roberta large as the main backbone with a simple classification head that either uses max+mean pooling or the hidden states of the CLS token. In total we used the following backbones with their approximate weighting in our final ensemble in brackets.\n \n- 5x xlm-roberta-large (85%)\n- 1x xlm-roberta-base (5%)\n- 1x mBart-large (10%)\n\n## Training strategy\nOne main ingredient for our good result is a stepwise finetuning towards single language models for tr, it and es. Let us illustrate the whole approach before explaining:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1424766%2Fd03ded105462fb5ba610393ba97e4ad8%2Fjigsaw%20arch2.001.jpeg?generation=1592909197036996&amp;alt=media)\n\nAs input data we used the english  jigsaw-toxic-comment-train.csv and its six translations available as public datasets which combined are roughly 1.4M comments. I realized that training a model directly on the combined dataset is not optimal, as each comment appears 7x in each epoch (although in different languages) and the model overfits on that. So I divided the combined dataset into 7 stratified parts, where each part contains a comment only once. For each fold we then finetuned a transformer in a 3step manner:\n\n\n- Step 1: finetune to all 7 languages for 2 epochs\n- Step 2: finetune only to the full valid.csv which only has tr, it and es\n- Step 3: 3x finetune to each language in valid resulting in 3 models\n\nWe then use the step1 model for predicting ru, the step2 model for predicting pt and fr and the respective step3 models for tr, it and es. Using the step2 model for pt and fr gave a significant boost compared to using step1 model for those due. Most likely due to the language similarity between it, es, fr and pt.\n\nWe used mainly a batchsize of 32 using gradient accumulation and a learning rate of 3e-6 with linear decay with AdamW optimizer. One thing worth mentioning is that we train on a max sequence length of 512 due to the dataloader mentioned above which sorts the data by length and then randomly serves batches of same length comments. Only the huge speed-up paired with the non padding of shorter comments made using a 512 max length reasonable. \n\nApart from one model which was built on top of a public kernel, all models were trained using pytorch and GPU.\n\n## Ensembling\nAs for ensembling we did team member individual ensembling and then combined the models of the team members and public kernels by weighted sum of rank percentile \n\n## Post-processing\nAnother key ingredient, which we luckily found on the last day is post processing of the final predictions. Although languages are derived from individual models, the competition metric is sensitive to the global rank of all predictions and not language specific. So we took care of the languages having a correct relation to each other, by shifting the predictions of each language individually. In our final submission for example we used the following factors \n\n```\ntest_df.loc[test_df[\"lang\"] == \"es\", \"toxic\"] *= 1.06\ntest_df.loc[test_df[\"lang\"] == \"fr\", \"toxic\"] *= 1.04\ntest_df.loc[test_df[\"lang\"] == \"it\", \"toxic\"] *= 0.97\ntest_df.loc[test_df[\"lang\"] == \"pt\", \"toxic\"] *= 0.96\ntest_df.loc[test_df[\"lang\"] == \"tr\", \"toxic\"] *= 0.98\n```\n\n\nWe derived the factors by matching the test prediction mean with the public LB mean for each language individually. We had already made 6 submissions for probing the test set by setting one language to 1 and the rest to 0. Then we made sure that means are aligned to give similar experimental AUC. That posprocessing moved us from 14th place to 4th place on public LB! \n\nWhile this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.\n\nThanks for reading. Questions welcome.",
    "898189": "Post Proc ... Genius! Thank you for your exp!",
    "898201": "Question: so, how was the architecture graph done? 👀 \nMore seriously: what hardware have you used, i.e. your own or cloud? Also, have you tried TPU (I guess not since you only mention V100)? \nGreat achievement and a lot to learn from!",
    "898221": "Thanks for sharing your great approach!\nI used the normal 1 model to predict 6 languages approach with xlmr-large models.  I got +0.008 improvement with your post processing.   \n|  |  public score| private score |\n| --- | --- | --- |\n|xlmr-large my blending  | 0.9487 |0.9471 |\n| xlmr-large my blending + your post processing | 0.9494|0.9479|",
    "898225": "Amazing work...I never thought your team would boost from 9500 to 9522 in the last day😜 . Can you share some details about how your team found this post-processing trick?",
    "898233": "To me it made sense that global ranking matters, and especially the 'ru' prediction *should* be off as the only thing the step1 model has seen were translations from en, which I think are weaker in terms of toxicity as original russian comments. Hence we expected those predictions to be in average too low. Then we checked shifting the predictions on the validation set, and it gave a significant improvement in score. So we worked on a method to shift test predictions too...",
    "898247": "Congrats for one week journey and anther gold medal!\nWe also find each language need calibration in the last day, but we have got no idea about the coef in only 5 submissions, we just try some lucky coef and get a small boost on post processing.\nAmazing work and thanks for detailed writeup!",
    "898248": "Thank you for your explanation!!!\nHope to learn more from you",
    "898250": "The post-processing is real magic! My model got +0.006  and +0.011  improvement when I just multiple Spanish by 1.2 .",
    "898275": "Congrats @christofhenkel, @cpmpml, @aerdem4 and thanks for sharing a very good solution writeup.",
    "898285": "So many brilliant ideas! Thanks for sharing.\n\nOn the ensembling, is the `weighted sum of rank percentile` the method of choice to try first, from experience, or just one from a toolbox ?",
    "898287": "You forget @aerdem4 ;)",
    "898290": "cpmpml, amended :-)",
    "898292": "cpmpml, it must be the Tweet and Jigsaw exhaustion. @aerdem4 may agree :-)",
    "898373": "Dieter and team, very discipline and intelligent move indeed.  Watching you and your team come after us from behind is akin watching a horror movie. What a great chase. Thanks very much for sharing the insights. Dr.",
    "898388": "Ahmet used TPU, Dieter and I used GPU (V100).  The former used Keras/TF while the latter used Pytorch.  There may be a correlation there ;)",
    "898529": "Great model, congrats on building such a great model in a short time. \n\nWow, nice post post processing trick. I just applied it to my final model. It increased both public and private LB by 0.008 and increased my private rank from 61st to 24th. Wow.",
    "898586": "Makes sense. :)",
    "898597": "I am sure if you would optimize it further on public part you could get even higher. The optimal scaling factors can differ quite a bit depending on how your predictions look like.",
    "899129": "Great Post-processing Guys !",
    "899200": "Congratulations sir.",
    "899414": "Congrats and thanks for sharing!  All of the methods you've shared are so cool and effective,  which have opened my eyes!&nbsp; I also tried to apply your post-processing to my final model.  It increased my private rank from 23rd to 13th！I'd never thought of such a fast but effective trick to improve score before.  That's really awesome! \n    There's so much more I need to learn.",
    "899656": "My takeway from recent deep learning competitions is that postprocessing is the best way to improve score, more than NN architecture.  I'm serious.",
    "899668": "cpmpml happened for both jigsaw and twitter. And suprisingly both were expected to have big shakeups and ended up having nothing.",
    "899676": "When everyone is working off the same set of pretrained models and knowledge of common architectural tweaks, postprocessing is one of the few ways to distinguish your models. \n\nThe main downside with PP is that it's easy  to do it at the last minute so people can see their leads evaporate literally overnight.",
    "899680": "shahules it also happened in Bengali\n\n@leecming sometimes the pp is tricky.  In Tweet sentiment I didn't found it nor did @philippsinger for instance.",
    "899682": "You shoudl try many methods and see their effect.  But when roc-auc is the metric, only rank matter, hence using rank makes sense.",
    "899704": "Yeah, I found/profited from it here, but did not in Tweets.",
    "899713": "And to be fair, if we would have had more time, I would have tried to incorporate fixing the disturbed individual language distributions in the modelling, instead of fixing it via pp. Often post-processing fixes the symptoms but does not solve the root cause.",
    "899741": "You still need to adjust for test distributions. So in both ways you are doing some form of test pp, even if it is from a different perspective.",
    "899747": "We wouldn't postprocess without LB probing. If Kaggle splits the public/private not random, it could be much different. In the past, Kaggle had more tricky splits.",
    "899884": "Yes, I wanted to make a longer post about it but pub LB should probably not have included all languages, and private test should have been completely hidden. I still dont understand the technical argument of why TPUs need to have the full test data available.",
    "904323": "I got 0.9499 on private LB by just applying the PP trick to my (not chosen) best model ( from 0.9478  to 0.9499 )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F974295%2Fb81459e076acd5500bd0df2ce56ea854%2F11Capture.PNG?generation=1593268260582472&amp;alt=media)",
    "904747": "yeah, its a good pp trick :D",
    "910784": "Congratulation! Super solid solution and really loved it, may I ask you how you implemented the padding by batch trick in detail? It would be great if you can share any links for refs, sort of new to NLP and BERT so need some help. By the way around the last 2 days, we were at 14th and then you guys jumped up over us, it makes losing a gold medal not that suffering if it was you guys who outplayed our team LOL, thanks for sharing your cool ideas.",
    "916408": "Congrats! \nAnd sorry for late asking.\nI'm still confused about the way from 6 single language submissions to pp factors, could you sharing more details?",
    "1495839": "Thanks for sharing!! very helpful",
    "1574214": "Is your BERT model able to detect negation and sarcasm?",
    "1590621": "日本語訳です\nBrief Summary\n私たちのソリューションは、いくつかの後処理と組み合わせたいくつかのトランスモデル（主にxlm-roberta-large）の単純なブレンドです。 アーキテクチャは公開されているカーネルと同じで、分類ヘッドは非表示状態の最大+平均プーリングまたはCLSトークンの非表示状態のいずれかを取ります。 モデルのトレーニングには3段階のアプローチを使用しました。ここでは、7つの言語から始めて、3つの言語に微調整し、1つの言語で終了しました。 この3ステップのアプローチではグローバル予測が歪むため、後処理ステップとして言語を1つの要因で個別にシフトしました。\n\nDetailed Summary\nこんなに短い時間で素晴らしい結果が得られたことに、私たちは今でも驚いています。 @ aerdem4はこのコンテストでもう少し長く働き、特定の課題をすでによく理解していましたが、@ cpmpmlと@christofhenkelは、チームの合併期限の数時間前のツイート感情コンテストの直後に参加しました。 2日前、私たちはまだ100以上の場所にいて、せいぜい銀メダルを望んでいました。 しかし、粘り強さと適切な量の直感で、どのアイデアを進めるか、そしてもちろん運が良ければ、私たちは金の上位に登ることができました。\n\nPreprocessing\nNone. That's NLP in 2020 :D\n\nSpeeding up training\n私たちは2つのことをしました。 1つは、ネガティブサンプルをダウンサンプリングして、よりバランスの取れたデータセットとより小さなデータセットを同時に取得することでした。 ポジティブとネガティブの数だけを使用する場合、データセットは約5分の1に削減されます。これは、1つのエポックを実行する時間と同じです。 私たちのモデルのいくつかはそのように訓練されました。\n\nもう1つのスピードアップは、バッチによるパディングによるものです。 主なアイデアは、パディングの量を現在のバッチに必要な量に制限することです。 入力は、固定長（512など）ではなく、バッチ内の最長入力の長さにパディングされます。これは現在よく知られている手法であり、以前のコンテストで使用されています。 トレーニングを大幅に加速します。\n特定のバッチのサンプルの長さが同じになるように、サンプルを長さで並べ替えることでアイデアを洗練しました。 これにより、パディングの必要性がさらに減少します。 バッチ内のすべての入力の長さが同じである場合、パディングはまったくありません。\n\nサンプルがソートされている場合、トレーニングモードでそれらをシャッフルすることはできません。 むしろバッチをシャッフルしました。 これにより、バッチによる元のパディングと比較して、さらに2倍のスピードアップが得られます。 最初のトレインセットでxlm-roberta-largeの1つのエポックをトレーニングするには、V100GPUで約17分かかります。\n\nCross-validation\n時間が短く、LBフィードバックを取得するための提出物があまりなかったため、信頼性の高い相互検証の取得に取り組みました。 未知の言語ru、pt、fr、および既知の言語it、es、trを表すために、言語ごとに3倍のグループと検証データの単純な3倍の組み合わせを使用しました。 その結果、6倍のスキームになります。 したがって、たとえばfold1は、esがvalid.csvにない場合、ptが無効であるのと同じように動作するはずです。 すべてのフォールドの平均AUCを取ることは、パブリックおよびプライベートLBの非常に優れたプロキシでした。\n\nArchitectures\nほとんどのチームとして、max + meanプーリングまたはCLSトークンの非表示状態のいずれかを使用する単純な分類ヘッドを備えたメインバックボーンとしてxlm-robertalargeを使用しました。合計で、次のバックボーンを使用し、括弧内の最終的なアンサンブルでおおよその重みを付けました。\n\n5x xlm-roberta-large（85％）\n1x xlm-roberta-base（5％）\n1x mBart-large（10％）\nトレーニング戦略\n私たちの良い結果の主な要素の1つは、tr、it、およびesの単一言語モデルに向けた段階的な微調整です。説明する前に、アプローチ全体を説明しましょう。\n\n\n\n入力データとして、英語のjigsaw-toxic-comment-train.csvを使用し、公開データセットとして利用可能な6つの翻訳を組み合わせると、約140万件のコメントになります。結合されたデータセットでモデルを直接トレーニングすることは最適ではないことに気付きました。各コメントは各エポックで7倍に表示され（言語は異なりますが）、モデルはそれに適合しすぎます。そこで、結合したデータセットを7つの階層化された部分に分割しました。各部分には、コメントが1回だけ含まれています。次に、折り畳みごとに、トランスフォーマーを3段階で微調整しました。\n\nステップ1：2つのエポックで7つの言語すべてに微調整する\nステップ2：tr、it、esのみを含む完全なvalid.csvにのみ微調整します\nステップ3：有効な各言語への3倍の微調整により、3つのモデルが作成されます\n次に、ruを予測するためのstep1モデル、ptとfrを予測するためのstep2モデル、およびtr、it、esのそれぞれのstep3モデルを使用します。 ptとfrにstep2モデルを使用すると、期限のあるものにstep1モデルを使用する場合に比べて大幅に向上しました。おそらく、es、fr、ptの間の言語の類似性が原因です。\n\n勾配累積を使用した32のバッチサイズと、AdamWオプティマイザーによる線形減衰を伴う3e-6の学習率を主に使用しました。言及する価値のあることの1つは、データを長さでソートし、同じ長さのコメントのバッチをランダムに提供する上記のデータローダーのために、最大シーケンス長512でトレーニングすることです。最大512の長さを使用して作成された短いコメントの非パディングと組み合わせた大幅なスピードアップのみが妥当です。\n\nパブリックカーネル上に構築された1つのモデルを除いて、すべてのモデルはpytorchとGPUを使用してトレーニングされました。\n\nEnsembling\nアンサンブルについては、チームメンバーの個別アンサンブルを行い、ランクパーセンタイルの加重和でチームメンバーとパブリックカーネルのモデルを組み合わせました。\n\nPost-processing\n最終日に幸運にも見つけたもう1つの重要な要素は、最終予測の後処理です。言語は個々のモデルから派生していますが、競合指標はすべての予測のグローバルランクに敏感であり、言語固有ではありません。そこで、各言語の予測を個別にシフトすることで、相互に正しい関係を持つ言語を処理しました。たとえば、最終的な提出では、次の要素を使用しました\n\ntest_df.loc [test_df [\"lang\"] == \"es\"、 \"toxic\"] * = 1.06\ntest_df.loc [test_df [\"lang\"] == \"fr\"、 \"toxic\"] * = 1.04\ntest_df.loc [test_df [\"lang\"] == \"it\"、 \"toxic\"] * = 0.97\ntest_df.loc [test_df [\"lang\"] == \"pt\"、 \"toxic\"] * = 0.96\ntest_df.loc [test_df [\"lang\"] == \"tr\"、 \"toxic\"] * = 0.98\nテスト予測平均を各言語の公開LB平均と個別に照合することにより、因子を導き出しました。 1つの言語を1に設定し、残りを0に設定して、テストセットを調査するために、すでに6つの提出を行いました。次に、同様の実験的なAUCが得られるように平均が調整されていることを確認しました。そのposprocessingは私達を公共LBの14位から4位に動かしました！\n\nこの後処理は他のチームにも役立つ可能性がありますが、6つの言語を予測するために5つのモデルを持つことによって導入されたグローバル予測分布の相違を具体的に修正し、チームが単一のモデルアプローチを使用した場合はあまり役に立たない可能性があります。\n\nThanks for reading. Questions welcome."
  },
  "source": "meta"
}