{
  "id": 161998,
  "title": "13th Place Solution Part II",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/161998",
  "author_name": "DHZM",
  "post_date": "2020-06-27T00:52:23.828000",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I want to thank all my wonderful teammates. Without your effort, we won't be able to make this happen. <a href=\"/euclidean\">@euclidean</a> <a href=\"/nullrecurrent\">@nullrecurrent</a> <a href=\"/strider1125\">@strider1125</a> <a href=\"/sherryli94\">@sherryli94</a> </p>\n\n<p>Congratulations to all the medal winners, it's been a tough competition for all of us.</p>\n\n<p>For the overview of our solution, you can refer to Part I written by <a href=\"/strider1125\">@strider1125</a> here:\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161974\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161974</a></p>\n\n<p>I will talk more about some details of our solution and where they came from:\n- CV\n- Stacking / Blending\n- Exotic Experiments\n- Things I Wish We Tried\n- Final Thoughts</p>\n\n<h1>CV</h1>\n\n<ul>\n<li>We started the competition using the validation set as purely holdout. In this phase, we reached around 9380 public.</li>\n<li>Later it became clear that we need to use the validation set and we first tried to do group-3fold by language, which gave significantly underestimated CV score and we reached 939x public.</li>\n<li>Then the different distribution among languages seemed strange when building OOF predictions and we decided to switch to random 4-fold split, which helped improve the stability of OOF predictions. </li>\n<li>Finally, we settled for 5fold with (language x toxic) stratified and used this CV setup until the end of the competition.</li>\n<li>It seems that <a href=\"/christofhenkel\">@christofhenkel</a> found some very reliable and elegant CV setup:\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980</a></li>\n</ul>\n\n<h1>Stacking / Blending</h1>\n\n<ul>\n<li>The reason we spent so much time on CV was that the CV score being so unstable, especially in stacking/blending.</li>\n<li>Stacking / Blending methods we tried\nHillClimb / Greedy\n<a href=\"https://www.kaggle.com/hhstrand/hillclimb-ensembling/\">https://www.kaggle.com/hhstrand/hillclimb-ensembling/</a>\nUnintended 3rd place <a href=\"/sakami\">@sakami</a> 's blending method\n<a href=\"https://github.com/sakami0000/kaggle_jigsaw/blob/master/compute_blending_weights.py\">https://github.com/sakami0000/kaggle_jigsaw/blob/master/compute_blending_weights.py</a>\nUnintended 9th place <a href=\"/khyeh0719\">@khyeh0719</a> 's stacking using ExtraTreeClassifier\n<a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100530\">https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100530</a>\nToxic (and this time!) 1st place <a href=\"/leecming\">@leecming</a> 's discussion\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557</a>\nGM <a href=\"/kazanova\">@kazanova</a> 's StackNet\n<a href=\"https://github.com/h2oai/pystacknet\">https://github.com/h2oai/pystacknet</a></li>\n<li>In terms of the second level, we mainly tried logistic regression, LGBM, and ExtraTreeClassifier. LGBM stacking brought us back to silver-zone with 9464 at some point, but then it seemed too overfitting. Eventually, we chose ExtraTreeClassifier as it was more stable and gave a better lb score. There is also some intuition here that our models seemed to have high variance and ExtraTreeClassifier is a good candidate to reduce that variance while keeping the predictive power of all base models.</li>\n<li>Our file structures allowed for a good enough stacking workflow (i.e. saving all OOF predictions and test predictions per model), but I have to say StackNet works very well and I highly recommend people to give it a try.</li>\n<li>The most painful part of stacking/blending is that every time you get a higher CV score from parameter tuning or feature selection, thinking you are going to save the world, the LB score will beat your ass. Eventually, after tons of failures, I realized that doing blind parameter optimization and feature selection under such an unstable CV setup was not going to help. </li>\n<li>The consistent trend we observed was that whenever we add more diverse models to the stack it improved our LB score and the CV-LB gap reduced, so we were pretty confident that our solution would be stable enough to survive the shake-up (but you might notice I'm saying this after getting the medal :-).</li>\n<li>Although our CV was not stable if we look at each individual score/experiment, on average our distribution of models has a good correlation with lb score and since we are using something like ExtraTreeClassifier, we can to some extent guarantee that it converges in probability to our LB score (I usually refer to this as my faith to Law of Large Numbers). Then it became clear that we should create more diverse models that have good scores (that's where the 1900 oof predictions came from).</li>\n</ul>\n\n<h1>Exotic Experiments</h1>\n\n<ul>\n<li>Averaging multiple XLMR's weights (similar to SWA), this gave our best single model (lb 9471).</li>\n<li>Adding the language tag (tr/fr/en/es etc.) as the first token after CLS in XLMR. So your tokens would look like [CLS] + [tr/fr/en/es...] + [other sentence tokens], this simple trick helped create one of our best single models. One could also do it in the way 2nd place winner <a href=\"/xiwuhan\">@xiwuhan</a> suggested: <a href=\"https://www.kaggle.com/xiwuhan/jmtc-fine-tune\">https://www.kaggle.com/xiwuhan/jmtc-fine-tune</a></li>\n<li>Random masking of tokens for data augmentation. </li>\n<li>Using aux columns in stacking, quite a few aux predictions were selected when doing blending(forward selection) and we believe improved stacking diversity as well.</li>\n<li>Null permutation and RFECV for stacking feature selection (didn't help).</li>\n<li>Adding more 2nd level OOF in StackNet (didn't help much).</li>\n<li>Pseudo Labeling in stacking (didn't help).</li>\n<li>Roberta large on English (score low, probably that's why we missed the monolingual models most other top teams touched: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862</a>)</li>\n</ul>\n\n<h1>Things I Wish We Tried</h1>\n\n<ul>\n<li>Freeze embedding layers / different learning rate per layer</li>\n<li>PL with soft labels and more iterations</li>\n<li>Putting monolingual models into our stacking</li>\n<li>Trying GPT-2 / XLNET etc.</li>\n<li>Post-processing mentioned by other teams (adjusting average toxicity per language)</li>\n</ul>\n\n<h1>Final Thoughts</h1>\n\n<ul>\n<li>It's not easy to get a gold medal. I read the winning solutions from Unintended so many times that I almost remember who used which trick/technique. All those discussions were super helpful and I learned a lot from them. And now we finally get the chance to write our own gold solution.</li>\n<li>Keep trying new things, keep doing new experiments, come up with new hypotheses, and validate them with a proper CV/LB. Whenever I get stuck I go back reading winning solutions, and they always gave me some new thoughts.</li>\n<li>I love this game!</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F703066%2Fad5806b3804a0d6ec895f8a8497feeeb%2FWechatIMG309.jpeg?generation=1593219737189824&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 903579,
      "postDate": "2020-06-27T00:52:23.830Z",
      "content": "<p>First of all, I want to thank all my wonderful teammates. Without your effort, we won't be able to make this happen. <a href=\"/euclidean\">@euclidean</a> <a href=\"/nullrecurrent\">@nullrecurrent</a> <a href=\"/strider1125\">@strider1125</a> <a href=\"/sherryli94\">@sherryli94</a> </p>\n\n<p>Congratulations to all the medal winners, it's been a tough competition for all of us.</p>\n\n<p>For the overview of our solution, you can refer to Part I written by <a href=\"/strider1125\">@strider1125</a> here:\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161974\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161974</a></p>\n\n<p>I will talk more about some details of our solution and where they came from:\n- CV\n- Stacking / Blending\n- Exotic Experiments\n- Things I Wish We Tried\n- Final Thoughts</p>\n\n<h1>CV</h1>\n\n<ul>\n<li>We started the competition using the validation set as purely holdout. In this phase, we reached around 9380 public.</li>\n<li>Later it became clear that we need to use the validation set and we first tried to do group-3fold by language, which gave significantly underestimated CV score and we reached 939x public.</li>\n<li>Then the different distribution among languages seemed strange when building OOF predictions and we decided to switch to random 4-fold split, which helped improve the stability of OOF predictions. </li>\n<li>Finally, we settled for 5fold with (language x toxic) stratified and used this CV setup until the end of the competition.</li>\n<li>It seems that <a href=\"/christofhenkel\">@christofhenkel</a> found some very reliable and elegant CV setup:\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980</a></li>\n</ul>\n\n<h1>Stacking / Blending</h1>\n\n<ul>\n<li>The reason we spent so much time on CV was that the CV score being so unstable, especially in stacking/blending.</li>\n<li>Stacking / Blending methods we tried\nHillClimb / Greedy\n<a href=\"https://www.kaggle.com/hhstrand/hillclimb-ensembling/\">https://www.kaggle.com/hhstrand/hillclimb-ensembling/</a>\nUnintended 3rd place <a href=\"/sakami\">@sakami</a> 's blending method\n<a href=\"https://github.com/sakami0000/kaggle_jigsaw/blob/master/compute_blending_weights.py\">https://github.com/sakami0000/kaggle_jigsaw/blob/master/compute_blending_weights.py</a>\nUnintended 9th place <a href=\"/khyeh0719\">@khyeh0719</a> 's stacking using ExtraTreeClassifier\n<a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100530\">https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100530</a>\nToxic (and this time!) 1st place <a href=\"/leecming\">@leecming</a> 's discussion\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557</a>\nGM <a href=\"/kazanova\">@kazanova</a> 's StackNet\n<a href=\"https://github.com/h2oai/pystacknet\">https://github.com/h2oai/pystacknet</a></li>\n<li>In terms of the second level, we mainly tried logistic regression, LGBM, and ExtraTreeClassifier. LGBM stacking brought us back to silver-zone with 9464 at some point, but then it seemed too overfitting. Eventually, we chose ExtraTreeClassifier as it was more stable and gave a better lb score. There is also some intuition here that our models seemed to have high variance and ExtraTreeClassifier is a good candidate to reduce that variance while keeping the predictive power of all base models.</li>\n<li>Our file structures allowed for a good enough stacking workflow (i.e. saving all OOF predictions and test predictions per model), but I have to say StackNet works very well and I highly recommend people to give it a try.</li>\n<li>The most painful part of stacking/blending is that every time you get a higher CV score from parameter tuning or feature selection, thinking you are going to save the world, the LB score will beat your ass. Eventually, after tons of failures, I realized that doing blind parameter optimization and feature selection under such an unstable CV setup was not going to help. </li>\n<li>The consistent trend we observed was that whenever we add more diverse models to the stack it improved our LB score and the CV-LB gap reduced, so we were pretty confident that our solution would be stable enough to survive the shake-up (but you might notice I'm saying this after getting the medal :-).</li>\n<li>Although our CV was not stable if we look at each individual score/experiment, on average our distribution of models has a good correlation with lb score and since we are using something like ExtraTreeClassifier, we can to some extent guarantee that it converges in probability to our LB score (I usually refer to this as my faith to Law of Large Numbers). Then it became clear that we should create more diverse models that have good scores (that's where the 1900 oof predictions came from).</li>\n</ul>\n\n<h1>Exotic Experiments</h1>\n\n<ul>\n<li>Averaging multiple XLMR's weights (similar to SWA), this gave our best single model (lb 9471).</li>\n<li>Adding the language tag (tr/fr/en/es etc.) as the first token after CLS in XLMR. So your tokens would look like [CLS] + [tr/fr/en/es...] + [other sentence tokens], this simple trick helped create one of our best single models. One could also do it in the way 2nd place winner <a href=\"/xiwuhan\">@xiwuhan</a> suggested: <a href=\"https://www.kaggle.com/xiwuhan/jmtc-fine-tune\">https://www.kaggle.com/xiwuhan/jmtc-fine-tune</a></li>\n<li>Random masking of tokens for data augmentation. </li>\n<li>Using aux columns in stacking, quite a few aux predictions were selected when doing blending(forward selection) and we believe improved stacking diversity as well.</li>\n<li>Null permutation and RFECV for stacking feature selection (didn't help).</li>\n<li>Adding more 2nd level OOF in StackNet (didn't help much).</li>\n<li>Pseudo Labeling in stacking (didn't help).</li>\n<li>Roberta large on English (score low, probably that's why we missed the monolingual models most other top teams touched: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862</a>)</li>\n</ul>\n\n<h1>Things I Wish We Tried</h1>\n\n<ul>\n<li>Freeze embedding layers / different learning rate per layer</li>\n<li>PL with soft labels and more iterations</li>\n<li>Putting monolingual models into our stacking</li>\n<li>Trying GPT-2 / XLNET etc.</li>\n<li>Post-processing mentioned by other teams (adjusting average toxicity per language)</li>\n</ul>\n\n<h1>Final Thoughts</h1>\n\n<ul>\n<li>It's not easy to get a gold medal. I read the winning solutions from Unintended so many times that I almost remember who used which trick/technique. All those discussions were super helpful and I learned a lot from them. And now we finally get the chance to write our own gold solution.</li>\n<li>Keep trying new things, keep doing new experiments, come up with new hypotheses, and validate them with a proper CV/LB. Whenever I get stuck I go back reading winning solutions, and they always gave me some new thoughts.</li>\n<li>I love this game!</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F703066%2Fad5806b3804a0d6ec895f8a8497feeeb%2FWechatIMG309.jpeg?generation=1593219737189824&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "First of all, I want to thank all my wonderful teammates. Without your effort, we won't be able to make this happen. @euclidean @nullrecurrent @strider1125 @sherryli94 \n\nCongratulations to all the medal winners, it's been a tough competition for all of us.\n\nFor the overview of our solution, you can refer to Part I written by @strider1125 here:\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161974\n\nI will talk more about some details of our solution and where they came from:\n- CV\n- Stacking / Blending\n- Exotic Experiments\n- Things I Wish We Tried\n- Final Thoughts\n\n\n# CV\n- We started the competition using the validation set as purely holdout. In this phase, we reached around 9380 public.\n- Later it became clear that we need to use the validation set and we first tried to do group-3fold by language, which gave significantly underestimated CV score and we reached 939x public.\n- Then the different distribution among languages seemed strange when building OOF predictions and we decided to switch to random 4-fold split, which helped improve the stability of OOF predictions. \n- Finally, we settled for 5fold with (language x toxic) stratified and used this CV setup until the end of the competition.\n- It seems that @christofhenkel found some very reliable and elegant CV setup:\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\n\n# Stacking / Blending\n- The reason we spent so much time on CV was that the CV score being so unstable, especially in stacking/blending.\n- Stacking / Blending methods we tried\nHillClimb / Greedy\nhttps://www.kaggle.com/hhstrand/hillclimb-ensembling/\nUnintended 3rd place @sakami 's blending method\nhttps://github.com/sakami0000/kaggle_jigsaw/blob/master/compute_blending_weights.py\nUnintended 9th place @khyeh0719 's stacking using ExtraTreeClassifier\nhttps://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100530\nToxic (and this time!) 1st place @leecming 's discussion\nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\nGM @kazanova 's StackNet\nhttps://github.com/h2oai/pystacknet\n- In terms of the second level, we mainly tried logistic regression, LGBM, and ExtraTreeClassifier. LGBM stacking brought us back to silver-zone with 9464 at some point, but then it seemed too overfitting. Eventually, we chose ExtraTreeClassifier as it was more stable and gave a better lb score. There is also some intuition here that our models seemed to have high variance and ExtraTreeClassifier is a good candidate to reduce that variance while keeping the predictive power of all base models.\n- Our file structures allowed for a good enough stacking workflow (i.e. saving all OOF predictions and test predictions per model), but I have to say StackNet works very well and I highly recommend people to give it a try.\n- The most painful part of stacking/blending is that every time you get a higher CV score from parameter tuning or feature selection, thinking you are going to save the world, the LB score will beat your ass. Eventually, after tons of failures, I realized that doing blind parameter optimization and feature selection under such an unstable CV setup was not going to help. \n- The consistent trend we observed was that whenever we add more diverse models to the stack it improved our LB score and the CV-LB gap reduced, so we were pretty confident that our solution would be stable enough to survive the shake-up (but you might notice I'm saying this after getting the medal :-).\n- Although our CV was not stable if we look at each individual score/experiment, on average our distribution of models has a good correlation with lb score and since we are using something like ExtraTreeClassifier, we can to some extent guarantee that it converges in probability to our LB score (I usually refer to this as my faith to Law of Large Numbers). Then it became clear that we should create more diverse models that have good scores (that's where the 1900 oof predictions came from).\n\n# Exotic Experiments\n\n- Averaging multiple XLMR's weights (similar to SWA), this gave our best single model (lb 9471).\n- Adding the language tag (tr/fr/en/es etc.) as the first token after CLS in XLMR. So your tokens would look like [CLS] + [tr/fr/en/es...] + [other sentence tokens], this simple trick helped create one of our best single models. One could also do it in the way 2nd place winner @xiwuhan suggested: https://www.kaggle.com/xiwuhan/jmtc-fine-tune\n- Random masking of tokens for data augmentation. \n- Using aux columns in stacking, quite a few aux predictions were selected when doing blending(forward selection) and we believe improved stacking diversity as well.\n- Null permutation and RFECV for stacking feature selection (didn't help).\n- Adding more 2nd level OOF in StackNet (didn't help much).\n- Pseudo Labeling in stacking (didn't help).\n- Roberta large on English (score low, probably that's why we missed the monolingual models most other top teams touched: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862)\n\n# Things I Wish We Tried\n\n- Freeze embedding layers / different learning rate per layer\n- PL with soft labels and more iterations\n- Putting monolingual models into our stacking\n- Trying GPT-2 / XLNET etc.\n- Post-processing mentioned by other teams (adjusting average toxicity per language)\n\n# Final Thoughts\n\n- It's not easy to get a gold medal. I read the winning solutions from Unintended so many times that I almost remember who used which trick/technique. All those discussions were super helpful and I learned a lot from them. And now we finally get the chance to write our own gold solution.\n- Keep trying new things, keep doing new experiments, come up with new hypotheses, and validate them with a proper CV/LB. Whenever I get stuck I go back reading winning solutions, and they always gave me some new thoughts.\n- I love this game!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F703066%2Fad5806b3804a0d6ec895f8a8497feeeb%2FWechatIMG309.jpeg?generation=1593219737189824&amp;alt=media)\n",
      "votes": 11
    },
    {
      "id": 903583,
      "postDate": "2020-06-27T00:56:44.397Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 903583,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-27T00:56:44.397000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "903579": "First of all, I want to thank all my wonderful teammates. Without your effort, we won't be able to make this happen. @euclidean @nullrecurrent @strider1125 @sherryli94 \n\nCongratulations to all the medal winners, it's been a tough competition for all of us.\n\nFor the overview of our solution, you can refer to Part I written by @strider1125 here:\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/161974\n\nI will talk more about some details of our solution and where they came from:\n- CV\n- Stacking / Blending\n- Exotic Experiments\n- Things I Wish We Tried\n- Final Thoughts\n\n\n# CV\n- We started the competition using the validation set as purely holdout. In this phase, we reached around 9380 public.\n- Later it became clear that we need to use the validation set and we first tried to do group-3fold by language, which gave significantly underestimated CV score and we reached 939x public.\n- Then the different distribution among languages seemed strange when building OOF predictions and we decided to switch to random 4-fold split, which helped improve the stability of OOF predictions. \n- Finally, we settled for 5fold with (language x toxic) stratified and used this CV setup until the end of the competition.\n- It seems that @christofhenkel found some very reliable and elegant CV setup:\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\n\n# Stacking / Blending\n- The reason we spent so much time on CV was that the CV score being so unstable, especially in stacking/blending.\n- Stacking / Blending methods we tried\nHillClimb / Greedy\nhttps://www.kaggle.com/hhstrand/hillclimb-ensembling/\nUnintended 3rd place @sakami 's blending method\nhttps://github.com/sakami0000/kaggle_jigsaw/blob/master/compute_blending_weights.py\nUnintended 9th place @khyeh0719 's stacking using ExtraTreeClassifier\nhttps://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/discussion/100530\nToxic (and this time!) 1st place @leecming 's discussion\nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\nGM @kazanova 's StackNet\nhttps://github.com/h2oai/pystacknet\n- In terms of the second level, we mainly tried logistic regression, LGBM, and ExtraTreeClassifier. LGBM stacking brought us back to silver-zone with 9464 at some point, but then it seemed too overfitting. Eventually, we chose ExtraTreeClassifier as it was more stable and gave a better lb score. There is also some intuition here that our models seemed to have high variance and ExtraTreeClassifier is a good candidate to reduce that variance while keeping the predictive power of all base models.\n- Our file structures allowed for a good enough stacking workflow (i.e. saving all OOF predictions and test predictions per model), but I have to say StackNet works very well and I highly recommend people to give it a try.\n- The most painful part of stacking/blending is that every time you get a higher CV score from parameter tuning or feature selection, thinking you are going to save the world, the LB score will beat your ass. Eventually, after tons of failures, I realized that doing blind parameter optimization and feature selection under such an unstable CV setup was not going to help. \n- The consistent trend we observed was that whenever we add more diverse models to the stack it improved our LB score and the CV-LB gap reduced, so we were pretty confident that our solution would be stable enough to survive the shake-up (but you might notice I'm saying this after getting the medal :-).\n- Although our CV was not stable if we look at each individual score/experiment, on average our distribution of models has a good correlation with lb score and since we are using something like ExtraTreeClassifier, we can to some extent guarantee that it converges in probability to our LB score (I usually refer to this as my faith to Law of Large Numbers). Then it became clear that we should create more diverse models that have good scores (that's where the 1900 oof predictions came from).\n\n# Exotic Experiments\n\n- Averaging multiple XLMR's weights (similar to SWA), this gave our best single model (lb 9471).\n- Adding the language tag (tr/fr/en/es etc.) as the first token after CLS in XLMR. So your tokens would look like [CLS] + [tr/fr/en/es...] + [other sentence tokens], this simple trick helped create one of our best single models. One could also do it in the way 2nd place winner @xiwuhan suggested: https://www.kaggle.com/xiwuhan/jmtc-fine-tune\n- Random masking of tokens for data augmentation. \n- Using aux columns in stacking, quite a few aux predictions were selected when doing blending(forward selection) and we believe improved stacking diversity as well.\n- Null permutation and RFECV for stacking feature selection (didn't help).\n- Adding more 2nd level OOF in StackNet (didn't help much).\n- Pseudo Labeling in stacking (didn't help).\n- Roberta large on English (score low, probably that's why we missed the monolingual models most other top teams touched: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862)\n\n# Things I Wish We Tried\n\n- Freeze embedding layers / different learning rate per layer\n- PL with soft labels and more iterations\n- Putting monolingual models into our stacking\n- Trying GPT-2 / XLNET etc.\n- Post-processing mentioned by other teams (adjusting average toxicity per language)\n\n# Final Thoughts\n\n- It's not easy to get a gold medal. I read the winning solutions from Unintended so many times that I almost remember who used which trick/technique. All those discussions were super helpful and I learned a lot from them. And now we finally get the chance to write our own gold solution.\n- Keep trying new things, keep doing new experiments, come up with new hypotheses, and validate them with a proper CV/LB. Whenever I get stuck I go back reading winning solutions, and they always gave me some new thoughts.\n- I love this game!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F703066%2Fad5806b3804a0d6ec895f8a8497feeeb%2FWechatIMG309.jpeg?generation=1593219737189824&amp;alt=media)\n",
    "903583": ""
  }
}