{
  "id": 160927,
  "title": "19th place solution",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160927",
  "author_name": "Nickil Maveli",
  "post_date": "2020-06-23T07:02:02.361000",
  "votes": 27,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Congratulations to <a href=\"/rafiko1\">@rafiko1</a> and <a href=\"/leecming\">@leecming</a> for winning the first place, besides being in the top of the public LB for the longest time, and <a href=\"https://jigsaw.google.com/\">Google Jigsaw</a> for hosting the third edition of toxicity classification challenge.</p>\n\n<p>We decided to divide our models into 3 parts. Number is listed alongside each model to denote what part they belong to.</p>\n\n<h2>Models:</h2>\n\n<ul>\n<li>XLM_RoBERTa_Base with different architectures/epochs/learning_rate/augmentation (I)</li>\n<li>XLM_RoBERTa_Large with different architectures/epochs/learning_rate/augmentation (I)</li>\n<li>XLM_RoBERTa_Large_MLM_training (I)</li>\n<li>Multilingual_BERT (cased and uncased variants) (II)</li>\n<li>BERT (cased and uncased variants) (II)</li>\n<li>RNN: pooled Bidirectional GRU with fasttext aligned word vector embeddings on translated train (III)</li>\n<li>RNN: pooled Bidirectional LSTM + GRU + attention with fasttext + glove + paragram embedding averaged on translated test (III)</li>\n<li>WeakLearner: NB-SVM on translated train as well as translated test averaged (III)</li>\n<li>WeakLearner: SGD on translated train as well as translated test averaged (III)</li>\n</ul>\n\n<p>Overall, there were a combination of about 50 models with the base models belonging to one type (parent) averaged. So, all varieties of XLM_RoBERTa_Base models were averaged and so on.</p>\n\n<p><strong>First-Level Blending procedure:</strong></p>\n\n<p>&gt; (I) -&gt; (XLM_RoBERTa_Base * 0.1)  + (XLM_RoBERTa_Large * 0.5) + (XLM_RoBERTa_Large_MLM_training * 0.4)</p>\n\n<p>&gt; (II) -&gt; (Multilingual_BERT * 0.2) + (BERT  * 0.8)</p>\n\n<p>&gt; (III) -&gt; (RNN * 0.6) + (WeakLearner * 0.4)</p>\n\n<p><strong>Final Blending procedure:</strong></p>\n\n<p>&gt; <code>toxic</code> = (I) * 0.8 + (II) * 0.1 + (III) * 0.1</p>\n\n<h2>Post Processing: (0.0005 boost)</h2>\n\n<p>I think this is the major highlight of our approach. Just a day before the competition finish, <a href=\"/veryrobustperson\">@veryrobustperson</a> discovered that most of our lower scoring submissions were over-predicting while compared to the higher scoring submissions. Based on this hypothesis, we tried to introduce a re-scaling factor for predictions &gt;0.8 as well as &lt;0.01 through the use of probabilistic random noise that introduces a small penalty. I am sure there was a lot to explore here, but due to shortage of submissions, and lack of time, we couldn't delve deeper into it and optimize further. For two different submissions that we tried, the score increased by about 0.0005, so we thought it would generalize well even on the private LB test data. Luckily, it did.  </p>\n\n<h2>Things that didn't / couldn't make it to work:</h2>\n\n<ul>\n<li>pseudo labelling by making &lt; 0.2 -&gt; 0 and &gt; 0.8 -&gt; 1.</li>\n<li>training a gradient boosting algorithm on the 1024 laser embeddings.</li>\n<li>hard-coding test probabilities to ground truth in validation as we found about 1000 samples overlapping based on cosine similarity of laser embeddings.</li>\n<li>changing the optimal threshold of toxic probabilities in <code>jigsaw-unintended-bias-train.csv</code> to a value ranging from 0.2 to 0.3, as 0.5 (default rounding) looked extremely non-toxic centric.</li>\n<li>using external datasets for hatespeech, toxic word list for 6 languages.</li>\n<li>label smoothing. </li>\n<li>Wordbatch FM_FTRL.</li>\n<li>power averaging of base models.</li>\n</ul>\n\n<p>Lastly, I thank my team-mates <a href=\"/veryrobustperson\">@veryrobustperson</a> and <a href=\"/ipythonx\">@ipythonx</a> for the wonderful discussions we've had and coming up with newer ideas from time to time. We almost made it to the gold zone :)</p>\n\n<p>Hope you all had fun competing! </p>\n\n<p>cheers,\nNickil</p>",
  "messages": [
    {
      "id": 897890,
      "postDate": "2020-06-23T07:02:02.360Z",
      "content": "<p>Congratulations to <a href=\"/rafiko1\">@rafiko1</a> and <a href=\"/leecming\">@leecming</a> for winning the first place, besides being in the top of the public LB for the longest time, and <a href=\"https://jigsaw.google.com/\">Google Jigsaw</a> for hosting the third edition of toxicity classification challenge.</p>\n\n<p>We decided to divide our models into 3 parts. Number is listed alongside each model to denote what part they belong to.</p>\n\n<h2>Models:</h2>\n\n<ul>\n<li>XLM_RoBERTa_Base with different architectures/epochs/learning_rate/augmentation (I)</li>\n<li>XLM_RoBERTa_Large with different architectures/epochs/learning_rate/augmentation (I)</li>\n<li>XLM_RoBERTa_Large_MLM_training (I)</li>\n<li>Multilingual_BERT (cased and uncased variants) (II)</li>\n<li>BERT (cased and uncased variants) (II)</li>\n<li>RNN: pooled Bidirectional GRU with fasttext aligned word vector embeddings on translated train (III)</li>\n<li>RNN: pooled Bidirectional LSTM + GRU + attention with fasttext + glove + paragram embedding averaged on translated test (III)</li>\n<li>WeakLearner: NB-SVM on translated train as well as translated test averaged (III)</li>\n<li>WeakLearner: SGD on translated train as well as translated test averaged (III)</li>\n</ul>\n\n<p>Overall, there were a combination of about 50 models with the base models belonging to one type (parent) averaged. So, all varieties of XLM_RoBERTa_Base models were averaged and so on.</p>\n\n<p><strong>First-Level Blending procedure:</strong></p>\n\n<p>&gt; (I) -&gt; (XLM_RoBERTa_Base * 0.1)  + (XLM_RoBERTa_Large * 0.5) + (XLM_RoBERTa_Large_MLM_training * 0.4)</p>\n\n<p>&gt; (II) -&gt; (Multilingual_BERT * 0.2) + (BERT  * 0.8)</p>\n\n<p>&gt; (III) -&gt; (RNN * 0.6) + (WeakLearner * 0.4)</p>\n\n<p><strong>Final Blending procedure:</strong></p>\n\n<p>&gt; <code>toxic</code> = (I) * 0.8 + (II) * 0.1 + (III) * 0.1</p>\n\n<h2>Post Processing: (0.0005 boost)</h2>\n\n<p>I think this is the major highlight of our approach. Just a day before the competition finish, <a href=\"/veryrobustperson\">@veryrobustperson</a> discovered that most of our lower scoring submissions were over-predicting while compared to the higher scoring submissions. Based on this hypothesis, we tried to introduce a re-scaling factor for predictions &gt;0.8 as well as &lt;0.01 through the use of probabilistic random noise that introduces a small penalty. I am sure there was a lot to explore here, but due to shortage of submissions, and lack of time, we couldn't delve deeper into it and optimize further. For two different submissions that we tried, the score increased by about 0.0005, so we thought it would generalize well even on the private LB test data. Luckily, it did.  </p>\n\n<h2>Things that didn't / couldn't make it to work:</h2>\n\n<ul>\n<li>pseudo labelling by making &lt; 0.2 -&gt; 0 and &gt; 0.8 -&gt; 1.</li>\n<li>training a gradient boosting algorithm on the 1024 laser embeddings.</li>\n<li>hard-coding test probabilities to ground truth in validation as we found about 1000 samples overlapping based on cosine similarity of laser embeddings.</li>\n<li>changing the optimal threshold of toxic probabilities in <code>jigsaw-unintended-bias-train.csv</code> to a value ranging from 0.2 to 0.3, as 0.5 (default rounding) looked extremely non-toxic centric.</li>\n<li>using external datasets for hatespeech, toxic word list for 6 languages.</li>\n<li>label smoothing. </li>\n<li>Wordbatch FM_FTRL.</li>\n<li>power averaging of base models.</li>\n</ul>\n\n<p>Lastly, I thank my team-mates <a href=\"/veryrobustperson\">@veryrobustperson</a> and <a href=\"/ipythonx\">@ipythonx</a> for the wonderful discussions we've had and coming up with newer ideas from time to time. We almost made it to the gold zone :)</p>\n\n<p>Hope you all had fun competing! </p>\n\n<p>cheers,\nNickil</p>",
      "rawMarkdown": "Congratulations to @rafiko1 and @leecming for winning the first place, besides being in the top of the public LB for the longest time, and [Google Jigsaw](https://jigsaw.google.com/) for hosting the third edition of toxicity classification challenge.\n\nWe decided to divide our models into 3 parts. Number is listed alongside each model to denote what part they belong to.\n\n## Models:\n\n* XLM\\_RoBERTa\\_Base with different architectures/epochs/learning_rate/augmentation (I)\n* XLM\\_RoBERTa\\_Large with different architectures/epochs/learning_rate/augmentation (I)\n* XLM\\_RoBERTa\\_Large\\_MLM\\_training (I)\n* Multilingual\\_BERT (cased and uncased variants) (II)\n* BERT (cased and uncased variants) (II)\n* RNN: pooled Bidirectional GRU with fasttext aligned word vector embeddings on translated train (III)\n* RNN: pooled Bidirectional LSTM + GRU + attention with fasttext + glove + paragram embedding averaged on translated test (III)\n* WeakLearner: NB-SVM on translated train as well as translated test averaged (III)\n* WeakLearner: SGD on translated train as well as translated test averaged (III)\n\nOverall, there were a combination of about 50 models with the base models belonging to one type (parent) averaged. So, all varieties of XLM\\_RoBERTa\\_Base models were averaged and so on.\n\n**First-Level Blending procedure:**\n\n&gt; (I) -&gt; (XLM\\_RoBERTa\\_Base * 0.1)  + (XLM\\_RoBERTa\\_Large * 0.5) + (XLM_RoBERTa\\_Large\\_MLM\\_training * 0.4)\n\n&gt; (II) -&gt; (Multilingual\\_BERT * 0.2) + (BERT  * 0.8)\n\n&gt; (III) -&gt; (RNN * 0.6) + (WeakLearner * 0.4)\n\n**Final Blending procedure:**\n\n&gt; `toxic` = (I) * 0.8 + (II) * 0.1 + (III) * 0.1\n\n## Post Processing: (0.0005 boost)\n\nI think this is the major highlight of our approach. Just a day before the competition finish, @veryrobustperson discovered that most of our lower scoring submissions were over-predicting while compared to the higher scoring submissions. Based on this hypothesis, we tried to introduce a re-scaling factor for predictions &gt;0.8 as well as &lt;0.01 through the use of probabilistic random noise that introduces a small penalty. I am sure there was a lot to explore here, but due to shortage of submissions, and lack of time, we couldn't delve deeper into it and optimize further. For two different submissions that we tried, the score increased by about 0.0005, so we thought it would generalize well even on the private LB test data. Luckily, it did.  \n\n## Things that didn't / couldn't make it to work:\n\n* pseudo labelling by making &lt; 0.2 -&gt; 0 and &gt; 0.8 -&gt; 1.\n* training a gradient boosting algorithm on the 1024 laser embeddings.\n* hard-coding test probabilities to ground truth in validation as we found about 1000 samples overlapping based on cosine similarity of laser embeddings.\n* changing the optimal threshold of toxic probabilities in `jigsaw-unintended-bias-train.csv` to a value ranging from 0.2 to 0.3, as 0.5 (default rounding) looked extremely non-toxic centric.\n* using external datasets for hatespeech, toxic word list for 6 languages.\n* label smoothing. \n* Wordbatch FM_FTRL.\n* power averaging of base models.\n\nLastly, I thank my team-mates @veryrobustperson and @ipythonx for the wonderful discussions we've had and coming up with newer ideas from time to time. We almost made it to the gold zone :)\n\nHope you all had fun competing! \n\n\ncheers,\nNickil",
      "votes": 27
    },
    {
      "id": 897950,
      "postDate": "2020-06-23T07:36:23.200Z",
      "content": "<p>We worked so hard for this competition!! Cheers!!</p>",
      "rawMarkdown": "We worked so hard for this competition!! Cheers!!",
      "votes": 3,
      "replies": [
        {
          "id": 897986,
          "postDate": "2020-06-23T07:55:51.367Z",
          "content": "<p><a href=\"/nickil21\">@nickil21</a> <a href=\"/veryrobustperson\">@veryrobustperson</a> almost gold 😄 </p>",
          "rawMarkdown": "@nickil21 @veryrobustperson almost gold 😄 ",
          "votes": 1
        },
        {
          "id": 899341,
          "postDate": "2020-06-24T07:12:40.290Z",
          "content": "<p>nb</p>",
          "rawMarkdown": "nb"
        }
      ]
    },
    {
      "id": 903401,
      "postDate": "2020-06-26T19:58:42.563Z",
      "content": "<p>Great review and very easy to read!  Thanks for sharing! 👍 🤘 💯 </p>",
      "rawMarkdown": "Great review and very easy to read!  Thanks for sharing! 👍 🤘 💯 ",
      "votes": 1
    },
    {
      "id": 898549,
      "postDate": "2020-06-23T15:26:55.660Z",
      "content": "<p>Congrats well done. 0.0005 isn't much increase do you mean 0.005 instead?</p>",
      "rawMarkdown": "Congrats well done. 0.0005 isn't much increase do you mean 0.005 instead?",
      "votes": 1,
      "replies": [
        {
          "id": 898717,
          "postDate": "2020-06-23T17:28:24.043Z",
          "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a>. I wish it was 0.005, but no it's indeed 0.0005.</p>",
          "rawMarkdown": "Thank you @cdeotte. I wish it was 0.005, but no it's indeed 0.0005."
        }
      ]
    },
    {
      "id": 898047,
      "postDate": "2020-06-23T08:44:45.953Z",
      "content": "<p>great! i learned awesome thing</p>",
      "rawMarkdown": "great! i learned awesome thing",
      "votes": 1
    },
    {
      "id": 898041,
      "postDate": "2020-06-23T08:40:37.273Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 903551,
      "postDate": "2020-06-27T00:06:00.883Z",
      "content": "<p>great insights for NLP tasks.. Thanks for sharing them <a href=\"/nickil21\">@nickil21</a> </p>",
      "rawMarkdown": "great insights for NLP tasks.. Thanks for sharing them @nickil21 ",
      "votes": 2
    },
    {
      "id": 898006,
      "postDate": "2020-06-23T08:14:56.793Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 2
    },
    {
      "id": 897974,
      "postDate": "2020-06-23T07:50:59.410Z",
      "content": "<p>congrats!</p>",
      "rawMarkdown": "congrats!",
      "votes": 2
    },
    {
      "id": 897959,
      "postDate": "2020-06-23T07:41:22.117Z",
      "content": "<p>Well done !</p>",
      "rawMarkdown": "Well done !",
      "votes": 2
    },
    {
      "id": 899336,
      "postDate": "2020-06-24T07:08:27.460Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 897952,
      "postDate": "2020-06-23T07:37:10.447Z",
      "content": "<p>Thank for sharing!!!!</p>",
      "rawMarkdown": "Thank for sharing!!!!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 897950,
      "author_name": "cat",
      "author_url": "",
      "post_date": "2020-06-23T07:36:23.200000",
      "content": "<p>We worked so hard for this competition!! Cheers!!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 897986,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-06-23T07:55:51.367000",
          "content": "<p><a href=\"/nickil21\">@nickil21</a> <a href=\"/veryrobustperson\">@veryrobustperson</a> almost gold 😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 899341,
          "author_name": "MachineLP",
          "author_url": "",
          "post_date": "2020-06-24T07:12:40.290000",
          "content": "<p>nb</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 903401,
      "author_name": "Matt Yates",
      "author_url": "",
      "post_date": "2020-06-26T19:58:42.563000",
      "content": "<p>Great review and very easy to read!  Thanks for sharing! 👍 🤘 💯 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 898549,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-23T15:26:55.660000",
      "content": "<p>Congrats well done. 0.0005 isn't much increase do you mean 0.005 instead?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 898717,
          "author_name": "Nickil Maveli",
          "author_url": "",
          "post_date": "2020-06-23T17:28:24.043000",
          "content": "<p>Thank you <a href=\"/cdeotte\">@cdeotte</a>. I wish it was 0.005, but no it's indeed 0.0005.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 898047,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2020-06-23T08:44:45.953000",
      "content": "<p>great! i learned awesome thing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 898041,
      "author_name": "PAB97",
      "author_url": "",
      "post_date": "2020-06-23T08:40:37.273000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 903551,
      "author_name": "Redwan Sony",
      "author_url": "",
      "post_date": "2020-06-27T00:06:00.883000",
      "content": "<p>great insights for NLP tasks.. Thanks for sharing them <a href=\"/nickil21\">@nickil21</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 898006,
      "author_name": "zr",
      "author_url": "",
      "post_date": "2020-06-23T08:14:56.793000",
      "content": "<p>Congrats!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 897974,
      "author_name": "Dracarys",
      "author_url": "",
      "post_date": "2020-06-23T07:50:59.410000",
      "content": "<p>congrats!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 897959,
      "author_name": "Haythem Tellili",
      "author_url": "",
      "post_date": "2020-06-23T07:41:22.117000",
      "content": "<p>Well done !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 899336,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-24T07:08:27.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 897952,
      "author_name": "Rajeev",
      "author_url": "",
      "post_date": "2020-06-23T07:37:10.447000",
      "content": "<p>Thank for sharing!!!!</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "897890": "Congratulations to @rafiko1 and @leecming for winning the first place, besides being in the top of the public LB for the longest time, and [Google Jigsaw](https://jigsaw.google.com/) for hosting the third edition of toxicity classification challenge.\n\nWe decided to divide our models into 3 parts. Number is listed alongside each model to denote what part they belong to.\n\n## Models:\n\n* XLM\\_RoBERTa\\_Base with different architectures/epochs/learning_rate/augmentation (I)\n* XLM\\_RoBERTa\\_Large with different architectures/epochs/learning_rate/augmentation (I)\n* XLM\\_RoBERTa\\_Large\\_MLM\\_training (I)\n* Multilingual\\_BERT (cased and uncased variants) (II)\n* BERT (cased and uncased variants) (II)\n* RNN: pooled Bidirectional GRU with fasttext aligned word vector embeddings on translated train (III)\n* RNN: pooled Bidirectional LSTM + GRU + attention with fasttext + glove + paragram embedding averaged on translated test (III)\n* WeakLearner: NB-SVM on translated train as well as translated test averaged (III)\n* WeakLearner: SGD on translated train as well as translated test averaged (III)\n\nOverall, there were a combination of about 50 models with the base models belonging to one type (parent) averaged. So, all varieties of XLM\\_RoBERTa\\_Base models were averaged and so on.\n\n**First-Level Blending procedure:**\n\n&gt; (I) -&gt; (XLM\\_RoBERTa\\_Base * 0.1)  + (XLM\\_RoBERTa\\_Large * 0.5) + (XLM_RoBERTa\\_Large\\_MLM\\_training * 0.4)\n\n&gt; (II) -&gt; (Multilingual\\_BERT * 0.2) + (BERT  * 0.8)\n\n&gt; (III) -&gt; (RNN * 0.6) + (WeakLearner * 0.4)\n\n**Final Blending procedure:**\n\n&gt; `toxic` = (I) * 0.8 + (II) * 0.1 + (III) * 0.1\n\n## Post Processing: (0.0005 boost)\n\nI think this is the major highlight of our approach. Just a day before the competition finish, @veryrobustperson discovered that most of our lower scoring submissions were over-predicting while compared to the higher scoring submissions. Based on this hypothesis, we tried to introduce a re-scaling factor for predictions &gt;0.8 as well as &lt;0.01 through the use of probabilistic random noise that introduces a small penalty. I am sure there was a lot to explore here, but due to shortage of submissions, and lack of time, we couldn't delve deeper into it and optimize further. For two different submissions that we tried, the score increased by about 0.0005, so we thought it would generalize well even on the private LB test data. Luckily, it did.  \n\n## Things that didn't / couldn't make it to work:\n\n* pseudo labelling by making &lt; 0.2 -&gt; 0 and &gt; 0.8 -&gt; 1.\n* training a gradient boosting algorithm on the 1024 laser embeddings.\n* hard-coding test probabilities to ground truth in validation as we found about 1000 samples overlapping based on cosine similarity of laser embeddings.\n* changing the optimal threshold of toxic probabilities in `jigsaw-unintended-bias-train.csv` to a value ranging from 0.2 to 0.3, as 0.5 (default rounding) looked extremely non-toxic centric.\n* using external datasets for hatespeech, toxic word list for 6 languages.\n* label smoothing. \n* Wordbatch FM_FTRL.\n* power averaging of base models.\n\nLastly, I thank my team-mates @veryrobustperson and @ipythonx for the wonderful discussions we've had and coming up with newer ideas from time to time. We almost made it to the gold zone :)\n\nHope you all had fun competing! \n\n\ncheers,\nNickil",
    "897950": "We worked so hard for this competition!! Cheers!!",
    "903401": "Great review and very easy to read!  Thanks for sharing! 👍 🤘 💯 ",
    "898549": "Congrats well done. 0.0005 isn't much increase do you mean 0.005 instead?",
    "898047": "great! i learned awesome thing",
    "898041": "Congratulations!",
    "903551": "great insights for NLP tasks.. Thanks for sharing them @nickil21 ",
    "898006": "Congrats!",
    "897974": "congrats!",
    "897959": "Well done !",
    "899336": "",
    "897952": "Thank for sharing!!!!"
  }
}