{
  "id": 160964,
  "title": "3rd Place Solution",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/ace-team-3rd-place-solution",
  "author_name": "",
  "post_date": "2020-06-23T09:28:33.434624200Z",
  "votes": 35,
  "comment_count": 19,
  "views": 0,
  "content": "<p>First of all we are grateful to Jigsaw for organizing such an interesting competition, and flawless execution in terms of data quality, train / validation/ test split etc. Also very thankful to Kaggle for providing the free TPUs which made quick experimentation possible. This is the first time I used just Kaggle Notebook (kernels) for training / inference and was impressed with how well they work most of the time, with a slight exception of lacking full blown IDE functionality, love everything about the Kernels. </p>\n\n<p>Last, but not the least, I want to thank my amazing teammates - @brightertiger (Ujjwal), @soloway (Igor) and@drpatrickchan (Dr Patrick) - who made our LB position possible and were super fun to work with. Our best submission was made in the last 15 min of the competition, so this was a relentless team, which never stopped improving.  We were constantly looking and re-evaluating new angles to improve our solution.  Although, I am writing the solution description, but I speak on everyone’s behalf here. </p>\n\n<p>Following were the major parts of our best submission:</p>\n\n<p><strong>RoBERTa XLM pre-trained</strong>: Like we saw with all the publicly shared kernels, RoBERTa XLM was the workhorse, which just delivered great performance with minimum effort / training. Used a lot of things discussed in the forum like translated data, open subtitles data along with averaging across multiple folds. Typically about ~200K or so observations were enough to train a single model, so there was a lot of scope to do multi-fold averaging given the huge data available at hand. A big thanks to @shonenkov and @xhlulu for their excellent kernels. </p>\n\n<p><strong>Language Specific Pre-trained Bert Models</strong>: Used language specific pre-trained Bert models for all the 6 test languages. A given model with a given capacity is any day more powerful for a single language vs multiple languages. The only challenge was to merge the probabilities coming out of different models (which could have different distributions) into a single ranking which could be especially problematic for a ROC evaluation metric (Even if you got the rank ordering perfectly right within each language you can mess things up when combining across languages). Our teammate Igor could make it work magically. Big thanks to @shonenkov for his great <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">kernel</a> that is used for training these Bert Models.</p>\n\n<p><strong>RoBERTa XLM MLM</strong>: Borrowed the concept of training the model on domain specific data using this excellent <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\">kernel</a> by @riblidezso. While the performance from these models were similar to our existing XLM-Roberta models, they worked well in the blend. We added more data to the MLM step by using translations of training data, validation data and test dataset. This gave us a couple versions of the pre-trained model that we averaged in the blend. </p>\n\n<p><strong>Post Processing</strong>: We observed that around 5-10% of the comments looked like automated / template based messages. E.g - Look at Test ID 39482. What’s really happening here is - the username / page title / attachment name is quoted by the bot here, which may contain profanities. But most likely due to the way data labels would have been generated, these are likely marked as non-toxic - our models unfortunately get confused by this. We applied regular expressions, clustering and other heuristics to adjust scores coming directly out of the models for these comments. This gave us close to additional 0.001  on LB</p>\n\n<p>Other things which impacted results a tiny bit:\n<strong>TTA</strong>: Some of our models did test time augmentation - where we did a weighted average of the prediction over foreign and english language. <br>\n<strong>Label smoothing</strong>: Was used in most of our models including the mono lang models.</p>\n\n<p>A weighted blend of the components # 1-3 put together and post processing described in # 4 gave us our 0.9523 on Public and 0.9509 on Private LB. </p>",
  "messages": [
    {
      "id": "898079",
      "postDate": "06/23/2020 09:28:33",
      "content": "<p>First of all we are grateful to Jigsaw for organizing such an interesting competition, and flawless execution in terms of data quality, train / validation/ test split etc. Also very thankful to Kaggle for providing the free TPUs which made quick experimentation possible. This is the first time I used just Kaggle Notebook (kernels) for training / inference and was impressed with how well they work most of the time, with a slight exception of lacking full blown IDE functionality, love everything about the Kernels. </p>\n\n<p>Last, but not the least, I want to thank my amazing teammates - @brightertiger (Ujjwal), @soloway (Igor) and@drpatrickchan (Dr Patrick) - who made our LB position possible and were super fun to work with. Our best submission was made in the last 15 min of the competition, so this was a relentless team, which never stopped improving.  We were constantly looking and re-evaluating new angles to improve our solution.  Although, I am writing the solution description, but I speak on everyone’s behalf here. </p>\n\n<p>Following were the major parts of our best submission:</p>\n\n<p><strong>RoBERTa XLM pre-trained</strong>: Like we saw with all the publicly shared kernels, RoBERTa XLM was the workhorse, which just delivered great performance with minimum effort / training. Used a lot of things discussed in the forum like translated data, open subtitles data along with averaging across multiple folds. Typically about ~200K or so observations were enough to train a single model, so there was a lot of scope to do multi-fold averaging given the huge data available at hand. A big thanks to @shonenkov and @xhlulu for their excellent kernels. </p>\n\n<p><strong>Language Specific Pre-trained Bert Models</strong>: Used language specific pre-trained Bert models for all the 6 test languages. A given model with a given capacity is any day more powerful for a single language vs multiple languages. The only challenge was to merge the probabilities coming out of different models (which could have different distributions) into a single ranking which could be especially problematic for a ROC evaluation metric (Even if you got the rank ordering perfectly right within each language you can mess things up when combining across languages). Our teammate Igor could make it work magically. Big thanks to @shonenkov for his great <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">kernel</a> that is used for training these Bert Models.</p>\n\n<p><strong>RoBERTa XLM MLM</strong>: Borrowed the concept of training the model on domain specific data using this excellent <a href=\"https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm\">kernel</a> by @riblidezso. While the performance from these models were similar to our existing XLM-Roberta models, they worked well in the blend. We added more data to the MLM step by using translations of training data, validation data and test dataset. This gave us a couple versions of the pre-trained model that we averaged in the blend. </p>\n\n<p><strong>Post Processing</strong>: We observed that around 5-10% of the comments looked like automated / template based messages. E.g - Look at Test ID 39482. What’s really happening here is - the username / page title / attachment name is quoted by the bot here, which may contain profanities. But most likely due to the way data labels would have been generated, these are likely marked as non-toxic - our models unfortunately get confused by this. We applied regular expressions, clustering and other heuristics to adjust scores coming directly out of the models for these comments. This gave us close to additional 0.001  on LB</p>\n\n<p>Other things which impacted results a tiny bit:\n<strong>TTA</strong>: Some of our models did test time augmentation - where we did a weighted average of the prediction over foreign and english language. <br>\n<strong>Label smoothing</strong>: Was used in most of our models including the mono lang models.</p>\n\n<p>A weighted blend of the components # 1-3 put together and post processing described in # 4 gave us our 0.9523 on Public and 0.9509 on Private LB. </p>",
      "rawMarkdown": "First of all we are grateful to Jigsaw for organizing such an interesting competition, and flawless execution in terms of data quality, train / validation/ test split etc. Also very thankful to Kaggle for providing the free TPUs which made quick experimentation possible. This is the first time I used just Kaggle Notebook (kernels) for training / inference and was impressed with how well they work most of the time, with a slight exception of lacking full blown IDE functionality, love everything about the Kernels. \n\nLast, but not the least, I want to thank my amazing teammates - @brightertiger (Ujjwal), @soloway (Igor) and@drpatrickchan (Dr Patrick) - who made our LB position possible and were super fun to work with. Our best submission was made in the last 15 min of the competition, so this was a relentless team, which never stopped improving.  We were constantly looking and re-evaluating new angles to improve our solution.  Although, I am writing the solution description, but I speak on everyone’s behalf here. \n\nFollowing were the major parts of our best submission:\n\n**RoBERTa XLM pre-trained**: Like we saw with all the publicly shared kernels, RoBERTa XLM was the workhorse, which just delivered great performance with minimum effort / training. Used a lot of things discussed in the forum like translated data, open subtitles data along with averaging across multiple folds. Typically about ~200K or so observations were enough to train a single model, so there was a lot of scope to do multi-fold averaging given the huge data available at hand. A big thanks to @shonenkov and @xhlulu for their excellent kernels. \n\n**Language Specific Pre-trained Bert Models**: Used language specific pre-trained Bert models for all the 6 test languages. A given model with a given capacity is any day more powerful for a single language vs multiple languages. The only challenge was to merge the probabilities coming out of different models (which could have different distributions) into a single ranking which could be especially problematic for a ROC evaluation metric (Even if you got the rank ordering perfectly right within each language you can mess things up when combining across languages). Our teammate Igor could make it work magically. Big thanks to @shonenkov for his great [kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta ) that is used for training these Bert Models.\n\n**RoBERTa XLM MLM**: Borrowed the concept of training the model on domain specific data using this excellent [kernel](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm) by @riblidezso. While the performance from these models were similar to our existing XLM-Roberta models, they worked well in the blend. We added more data to the MLM step by using translations of training data, validation data and test dataset. This gave us a couple versions of the pre-trained model that we averaged in the blend. \n\n**Post Processing**: We observed that around 5-10% of the comments looked like automated / template based messages. E.g - Look at Test ID 39482. What’s really happening here is - the username / page title / attachment name is quoted by the bot here, which may contain profanities. But most likely due to the way data labels would have been generated, these are likely marked as non-toxic - our models unfortunately get confused by this. We applied regular expressions, clustering and other heuristics to adjust scores coming directly out of the models for these comments. This gave us close to additional 0.001  on LB\n\nOther things which impacted results a tiny bit:\n**TTA**: Some of our models did test time augmentation - where we did a weighted average of the prediction over foreign and english language.  \n**Label smoothing**: Was used in most of our models including the mono lang models.\n\nA weighted blend of the components # 1-3 put together and post processing described in # 4 gave us our 0.9523 on Public and 0.9509 on Private LB.",
      "votes": null
    },
    {
      "id": "898165",
      "postDate": "06/23/2020 10:58:52",
      "content": "<p>Great write-up. Thank you for sharing</p>",
      "rawMarkdown": "Great write-up. Thank you for sharing",
      "votes": null
    },
    {
      "id": "898259",
      "postDate": "06/23/2020 12:02:46",
      "content": "<p>Congrats <a href=\"/moizsaifee\">@moizsaifee</a> and team. Thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @moizsaifee and team. Thanks for sharing your solution overview.",
      "votes": null
    },
    {
      "id": "898407",
      "postDate": "06/23/2020 13:59:49",
      "content": "<p>Brief and clear. Thank you for summarizing your team's solution and congratulations for 3rd place <a href=\"/moizsaifee\">@moizsaifee</a> and your team!</p>",
      "rawMarkdown": "Brief and clear. Thank you for summarizing your team's solution and congratulations for 3rd place @moizsaifee and your team!",
      "votes": null
    },
    {
      "id": "898471",
      "postDate": "06/23/2020 14:38:06",
      "content": "<p>Thanks Dieter,  Great performance from your team as well. </p>",
      "rawMarkdown": "Thanks Dieter,  Great performance from your team as well.",
      "votes": null
    },
    {
      "id": "898503",
      "postDate": "06/23/2020 14:59:11",
      "content": "<p>Congrats and amzing sharing. Then how to merge the probabilities coming out of different single language models?</p>",
      "rawMarkdown": "Congrats and amzing sharing. Then how to merge the probabilities coming out of different single language models?",
      "votes": null
    },
    {
      "id": "898507",
      "postDate": "06/23/2020 15:00:49",
      "content": "<p>Thank you so much. </p>",
      "rawMarkdown": "Thank you so much.",
      "votes": null
    },
    {
      "id": "898508",
      "postDate": "06/23/2020 15:01:11",
      "content": "<p>Thanks you, appreciate it. </p>",
      "rawMarkdown": "Thanks you, appreciate it.",
      "votes": null
    },
    {
      "id": "898511",
      "postDate": "06/23/2020 15:03:53",
      "content": "<p>Thanks Dieter. You guys had a meteoric rise after entering late in the competition (kudos for doing it so consistently!). You almost pipped us of # 3 on the last day :-)</p>",
      "rawMarkdown": "Thanks Dieter. You guys had a meteoric rise after entering late in the competition (kudos for doing it so consistently!). You almost pipped us of # 3 on the last day :-)",
      "votes": null
    },
    {
      "id": "898517",
      "postDate": "06/23/2020 15:08:40",
      "content": "<p>congrats.  What pretrained BERT models did you use?  The detection of automated labels is very nice.</p>",
      "rawMarkdown": "congrats.  What pretrained BERT models did you use?  The detection of automated labels is very nice.",
      "votes": null
    },
    {
      "id": "898568",
      "postDate": "06/23/2020 15:35:00",
      "content": "<p>Thanks! All models are from <a href=\"https://huggingface.co/models\">https://huggingface.co/models</a>\nru - DeepPavlov/rubert-base-cased-conversational\nes - dccuchile/bert-base-spanish-wwm-cased\nit - dbmdz/bert-base-italian-xxl-uncased\ntr - dbmdz/bert-base-turkish-cased\npt - neuralmind/bert-large-portuguese-cased (didn't help, not included in our blend)\nfr - camembert/camembert-large</p>",
      "rawMarkdown": "Thanks! All models are from https://huggingface.co/models\nru - DeepPavlov/rubert-base-cased-conversational\nes - dccuchile/bert-base-spanish-wwm-cased\nit - dbmdz/bert-base-italian-xxl-uncased\ntr - dbmdz/bert-base-turkish-cased\npt - neuralmind/bert-large-portuguese-cased (didn't help, not included in our blend)\nfr - camembert/camembert-large",
      "votes": null
    },
    {
      "id": "898699",
      "postDate": "06/23/2020 17:19:36",
      "content": "<p>Congratulations and thank you for sharing your approach.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing your approach.",
      "votes": null
    },
    {
      "id": "899502",
      "postDate": "06/24/2020 09:22:28",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": null
    },
    {
      "id": "900339",
      "postDate": "06/24/2020 18:47:47",
      "content": "<p>congrats. How you merge XLM-R and BERT- model?</p>",
      "rawMarkdown": "congrats. How you merge XLM-R and BERT- model?",
      "votes": null
    },
    {
      "id": "903486",
      "postDate": "06/26/2020 21:41:40",
      "content": "<p>Many congratulations to you and your team. 😄 \nI have few questions regarding your solution and techniques that you have used here.</p>\n\n<p><strong>RoBERTa XLM pre-trained</strong></p>\n\n<blockquote>\n  <p>so there was a lot of scope to do multi-fold averaging given the huge data available </p>\n</blockquote>\n\n<p>Does multi-fold averaging means OOF prediction here or is it something else?</p>\n\n<p><strong>Language Specific Pre-trained Bert Models</strong></p>\n\n<blockquote>\n  <p>The only challenge was to merge the probabilities coming out of different models</p>\n</blockquote>\n\n<p>How did you all managed to solved this challenge? </p>\n\n<p><strong>TTA</strong></p>\n\n<blockquote>\n  <p>weighted average of the prediction over foreign and english language`</p>\n</blockquote>\n\n<p>This means you did prediction on both test and translated test data during TTA? Then you used weighted average on TTA predictions amirite.</p>\n\n<p>Please don't mind if I have understood anything wrong. Many thanks.</p>",
      "rawMarkdown": "Many congratulations to you and your team. 😄 \nI have few questions regarding your solution and techniques that you have used here.\n\n**RoBERTa XLM pre-trained**\n&gt;  so there was a lot of scope to do multi-fold averaging given the huge data available \n\nDoes multi-fold averaging means OOF prediction here or is it something else?\n\n**Language Specific Pre-trained Bert Models**\n&gt; The only challenge was to merge the probabilities coming out of different models\n\nHow did you all managed to solved this challenge? \n\n**TTA**\n&gt;  weighted average of the prediction over foreign and english language`\n\nThis means you did prediction on both test and translated test data during TTA? Then you used weighted average on TTA predictions amirite.\n\nPlease don't mind if I have understood anything wrong. Many thanks.",
      "votes": null
    },
    {
      "id": "910738",
      "postDate": "07/01/2020 10:29:50",
      "content": "<p>We did weighted average of rank. Weights were decided based on individual performance on LB. The weights may not have been optimal</p>",
      "rawMarkdown": "We did weighted average of rank. Weights were decided based on individual performance on LB. The weights may not have been optimal",
      "votes": null
    },
    {
      "id": "910740",
      "postDate": "07/01/2020 10:31:47",
      "content": "<p>As per our teammate Igor who worked on this, just using raw probabilities coming out of the model worked pretty well. We could have done scaling based on validation data / LB, but wanted to avoid the risk of over-fitting. </p>",
      "rawMarkdown": "As per our teammate Igor who worked on this, just using raw probabilities coming out of the model worked pretty well. We could have done scaling based on validation data / LB, but wanted to avoid the risk of over-fitting.",
      "votes": null
    },
    {
      "id": "910743",
      "postDate": "07/01/2020 10:35:54",
      "content": "<p>Re Multi-fold: I mean training the model on different subset of data, and then taking average of predictions of test data using the resulting models.</p>\n\n<p>Re Scaling: There was the option of scaling probabilities based on validation data / LB. We ended up using just the raw probabilities as they were working up pretty well and wanted to avoid over-fitting risk. </p>\n\n<p>Re: TTA - yes, you are right. </p>",
      "rawMarkdown": "Re Multi-fold: I mean training the model on different subset of data, and then taking average of predictions of test data using the resulting models.\n\nRe Scaling: There was the option of scaling probabilities based on validation data / LB. We ended up using just the raw probabilities as they were working up pretty well and wanted to avoid over-fitting risk. \n\nRe: TTA - yes, you are right.",
      "votes": null
    },
    {
      "id": "910950",
      "postDate": "07/01/2020 13:18:19",
      "content": "<p>Amazing.! Congrats again and thanks for explanation.</p>",
      "rawMarkdown": "Amazing.! Congrats again and thanks for explanation.",
      "votes": null
    },
    {
      "id": "912795",
      "postDate": "07/02/2020 18:18:44",
      "content": "<p>nice. can you introduce me some source. I am new hear about this. Thanks for sharing</p>",
      "rawMarkdown": "nice. can you introduce me some source. I am new hear about this. Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 898165,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "06/23/2020 10:58:52",
      "content": "<p>Great write-up. Thank you for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 898471,
          "author_name": "drpatrickchan",
          "author_url": "",
          "post_date": "06/23/2020 14:38:06",
          "content": "<p>Thanks Dieter,  Great performance from your team as well. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898511,
          "author_name": "moizsaifee",
          "author_url": "",
          "post_date": "06/23/2020 15:03:53",
          "content": "<p>Thanks Dieter. You guys had a meteoric rise after entering late in the competition (kudos for doing it so consistently!). You almost pipped us of # 3 on the last day :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898259,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/23/2020 12:02:46",
      "content": "<p>Congrats <a href=\"/moizsaifee\">@moizsaifee</a> and team. Thanks for sharing your solution overview.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898508,
          "author_name": "moizsaifee",
          "author_url": "",
          "post_date": "06/23/2020 15:01:11",
          "content": "<p>Thanks you, appreciate it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898407,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "06/23/2020 13:59:49",
      "content": "<p>Brief and clear. Thank you for summarizing your team's solution and congratulations for 3rd place <a href=\"/moizsaifee\">@moizsaifee</a> and your team!</p>",
      "votes": null,
      "replies": [
        {
          "id": 898507,
          "author_name": "moizsaifee",
          "author_url": "",
          "post_date": "06/23/2020 15:00:49",
          "content": "<p>Thank you so much. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898503,
      "author_name": "mcggood",
      "author_url": "",
      "post_date": "06/23/2020 14:59:11",
      "content": "<p>Congrats and amzing sharing. Then how to merge the probabilities coming out of different single language models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 910740,
          "author_name": "moizsaifee",
          "author_url": "",
          "post_date": "07/01/2020 10:31:47",
          "content": "<p>As per our teammate Igor who worked on this, just using raw probabilities coming out of the model worked pretty well. We could have done scaling based on validation data / LB, but wanted to avoid the risk of over-fitting. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898517,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/23/2020 15:08:40",
      "content": "<p>congrats.  What pretrained BERT models did you use?  The detection of automated labels is very nice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898568,
          "author_name": "soloway",
          "author_url": "",
          "post_date": "06/23/2020 15:35:00",
          "content": "<p>Thanks! All models are from <a href=\"https://huggingface.co/models\">https://huggingface.co/models</a>\nru - DeepPavlov/rubert-base-cased-conversational\nes - dccuchile/bert-base-spanish-wwm-cased\nit - dbmdz/bert-base-italian-xxl-uncased\ntr - dbmdz/bert-base-turkish-cased\npt - neuralmind/bert-large-portuguese-cased (didn't help, not included in our blend)\nfr - camembert/camembert-large</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898699,
      "author_name": "deepchatterjeevns",
      "author_url": "",
      "post_date": "06/23/2020 17:19:36",
      "content": "<p>Congratulations and thank you for sharing your approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899502,
      "author_name": "gauravdahiya",
      "author_url": "",
      "post_date": "06/24/2020 09:22:28",
      "content": "<p>Congratulations!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 900339,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "06/24/2020 18:47:47",
      "content": "<p>congrats. How you merge XLM-R and BERT- model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 910738,
          "author_name": "moizsaifee",
          "author_url": "",
          "post_date": "07/01/2020 10:29:50",
          "content": "<p>We did weighted average of rank. Weights were decided based on individual performance on LB. The weights may not have been optimal</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912795,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "07/02/2020 18:18:44",
          "content": "<p>nice. can you introduce me some source. I am new hear about this. Thanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903486,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "06/26/2020 21:41:40",
      "content": "<p>Many congratulations to you and your team. 😄 \nI have few questions regarding your solution and techniques that you have used here.</p>\n\n<p><strong>RoBERTa XLM pre-trained</strong></p>\n\n<blockquote>\n  <p>so there was a lot of scope to do multi-fold averaging given the huge data available </p>\n</blockquote>\n\n<p>Does multi-fold averaging means OOF prediction here or is it something else?</p>\n\n<p><strong>Language Specific Pre-trained Bert Models</strong></p>\n\n<blockquote>\n  <p>The only challenge was to merge the probabilities coming out of different models</p>\n</blockquote>\n\n<p>How did you all managed to solved this challenge? </p>\n\n<p><strong>TTA</strong></p>\n\n<blockquote>\n  <p>weighted average of the prediction over foreign and english language`</p>\n</blockquote>\n\n<p>This means you did prediction on both test and translated test data during TTA? Then you used weighted average on TTA predictions amirite.</p>\n\n<p>Please don't mind if I have understood anything wrong. Many thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 910743,
          "author_name": "moizsaifee",
          "author_url": "",
          "post_date": "07/01/2020 10:35:54",
          "content": "<p>Re Multi-fold: I mean training the model on different subset of data, and then taking average of predictions of test data using the resulting models.</p>\n\n<p>Re Scaling: There was the option of scaling probabilities based on validation data / LB. We ended up using just the raw probabilities as they were working up pretty well and wanted to avoid over-fitting risk. </p>\n\n<p>Re: TTA - yes, you are right. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 910950,
          "author_name": "rhtsingh",
          "author_url": "",
          "post_date": "07/01/2020 13:18:19",
          "content": "<p>Amazing.! Congrats again and thanks for explanation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "898079": "First of all we are grateful to Jigsaw for organizing such an interesting competition, and flawless execution in terms of data quality, train / validation/ test split etc. Also very thankful to Kaggle for providing the free TPUs which made quick experimentation possible. This is the first time I used just Kaggle Notebook (kernels) for training / inference and was impressed with how well they work most of the time, with a slight exception of lacking full blown IDE functionality, love everything about the Kernels. \n\nLast, but not the least, I want to thank my amazing teammates - @brightertiger (Ujjwal), @soloway (Igor) and@drpatrickchan (Dr Patrick) - who made our LB position possible and were super fun to work with. Our best submission was made in the last 15 min of the competition, so this was a relentless team, which never stopped improving.  We were constantly looking and re-evaluating new angles to improve our solution.  Although, I am writing the solution description, but I speak on everyone’s behalf here. \n\nFollowing were the major parts of our best submission:\n\n**RoBERTa XLM pre-trained**: Like we saw with all the publicly shared kernels, RoBERTa XLM was the workhorse, which just delivered great performance with minimum effort / training. Used a lot of things discussed in the forum like translated data, open subtitles data along with averaging across multiple folds. Typically about ~200K or so observations were enough to train a single model, so there was a lot of scope to do multi-fold averaging given the huge data available at hand. A big thanks to @shonenkov and @xhlulu for their excellent kernels. \n\n**Language Specific Pre-trained Bert Models**: Used language specific pre-trained Bert models for all the 6 test languages. A given model with a given capacity is any day more powerful for a single language vs multiple languages. The only challenge was to merge the probabilities coming out of different models (which could have different distributions) into a single ranking which could be especially problematic for a ROC evaluation metric (Even if you got the rank ordering perfectly right within each language you can mess things up when combining across languages). Our teammate Igor could make it work magically. Big thanks to @shonenkov for his great [kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta ) that is used for training these Bert Models.\n\n**RoBERTa XLM MLM**: Borrowed the concept of training the model on domain specific data using this excellent [kernel](https://www.kaggle.com/riblidezso/finetune-xlm-roberta-on-jigsaw-test-data-with-mlm) by @riblidezso. While the performance from these models were similar to our existing XLM-Roberta models, they worked well in the blend. We added more data to the MLM step by using translations of training data, validation data and test dataset. This gave us a couple versions of the pre-trained model that we averaged in the blend. \n\n**Post Processing**: We observed that around 5-10% of the comments looked like automated / template based messages. E.g - Look at Test ID 39482. What’s really happening here is - the username / page title / attachment name is quoted by the bot here, which may contain profanities. But most likely due to the way data labels would have been generated, these are likely marked as non-toxic - our models unfortunately get confused by this. We applied regular expressions, clustering and other heuristics to adjust scores coming directly out of the models for these comments. This gave us close to additional 0.001  on LB\n\nOther things which impacted results a tiny bit:\n**TTA**: Some of our models did test time augmentation - where we did a weighted average of the prediction over foreign and english language.  \n**Label smoothing**: Was used in most of our models including the mono lang models.\n\nA weighted blend of the components # 1-3 put together and post processing described in # 4 gave us our 0.9523 on Public and 0.9509 on Private LB.",
    "898165": "Great write-up. Thank you for sharing",
    "898259": "Congrats @moizsaifee and team. Thanks for sharing your solution overview.",
    "898407": "Brief and clear. Thank you for summarizing your team's solution and congratulations for 3rd place @moizsaifee and your team!",
    "898471": "Thanks Dieter,  Great performance from your team as well.",
    "898503": "Congrats and amzing sharing. Then how to merge the probabilities coming out of different single language models?",
    "898507": "Thank you so much.",
    "898508": "Thanks you, appreciate it.",
    "898511": "Thanks Dieter. You guys had a meteoric rise after entering late in the competition (kudos for doing it so consistently!). You almost pipped us of # 3 on the last day :-)",
    "898517": "congrats.  What pretrained BERT models did you use?  The detection of automated labels is very nice.",
    "898568": "Thanks! All models are from https://huggingface.co/models\nru - DeepPavlov/rubert-base-cased-conversational\nes - dccuchile/bert-base-spanish-wwm-cased\nit - dbmdz/bert-base-italian-xxl-uncased\ntr - dbmdz/bert-base-turkish-cased\npt - neuralmind/bert-large-portuguese-cased (didn't help, not included in our blend)\nfr - camembert/camembert-large",
    "898699": "Congratulations and thank you for sharing your approach.",
    "899502": "Congratulations!!",
    "900339": "congrats. How you merge XLM-R and BERT- model?",
    "903486": "Many congratulations to you and your team. 😄 \nI have few questions regarding your solution and techniques that you have used here.\n\n**RoBERTa XLM pre-trained**\n&gt;  so there was a lot of scope to do multi-fold averaging given the huge data available \n\nDoes multi-fold averaging means OOF prediction here or is it something else?\n\n**Language Specific Pre-trained Bert Models**\n&gt; The only challenge was to merge the probabilities coming out of different models\n\nHow did you all managed to solved this challenge? \n\n**TTA**\n&gt;  weighted average of the prediction over foreign and english language`\n\nThis means you did prediction on both test and translated test data during TTA? Then you used weighted average on TTA predictions amirite.\n\nPlease don't mind if I have understood anything wrong. Many thanks.",
    "910738": "We did weighted average of rank. Weights were decided based on individual performance on LB. The weights may not have been optimal",
    "910740": "As per our teammate Igor who worked on this, just using raw probabilities coming out of the model worked pretty well. We could have done scaling based on validation data / LB, but wanted to avoid the risk of over-fitting.",
    "910743": "Re Multi-fold: I mean training the model on different subset of data, and then taking average of predictions of test data using the resulting models.\n\nRe Scaling: There was the option of scaling probabilities based on validation data / LB. We ended up using just the raw probabilities as they were working up pretty well and wanted to avoid over-fitting risk. \n\nRe: TTA - yes, you are right.",
    "910950": "Amazing.! Congrats again and thanks for explanation.",
    "912795": "nice. can you introduce me some source. I am new hear about this. Thanks for sharing"
  },
  "source": "meta"
}