{
  "id": 160876,
  "title": "21st Place solution : Magic of Ensemble",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/burning-tpu-21st-place-solution-magic-of-ensemble",
  "author_name": "",
  "post_date": "2020-06-23T01:51:39.055658600Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Big thanks to <a href=\"/riblidezso\">@riblidezso</a> who share this amazing notebook: <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a> .</p>\n\n<p>In short, ensemble and second step fine-tuned by language is magic (and overfitting too).</p>\n\n<h1>PB score timeline</h1>\n\n<p>Ensemble Gmean of top score public submissions -&gt; 0.9473 (8 June)\nEnsemble LGBM solution and Gmean submission -&gt; 0.9478 (10 June)\n(Magic) Add external data to validate dataset and change to validate by language in second step training of XLM and ensemble by language -&gt; 0.9500 (21 June)\nUnfortunately, we boosted the score to 0.9500 in the last 2 days but no more TPU left to train new external data.</p>\n\n<h1>Validate Dataset</h1>\n\n<p>We separated validate data to 3 set by language and find another 3 languages from external data to use in second step training in <a href=\"/riblidezso\">@riblidezso</a> notebook. The out of fold score looks promising so we blend them to our main submission and boost the score to 0.9500.</p>\n\n<h1>Ref</h1>\n\n<p>Gmean notebook: <a href=\"https://www.kaggle.com/paulorzp/gmean-of-low-correlation-lb-0-952x\">https://www.kaggle.com/paulorzp/gmean-of-low-correlation-lb-0-952x</a>\nXLM-Roberta: <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a>\nLGBM solution: <a href=\"https://www.kaggle.com/miklgr500/lgbm-solution\">https://www.kaggle.com/miklgr500/lgbm-solution</a>\nExternal data: <a href=\"https://www.kaggle.com/blackmoon/russian-language-toxic-comments\">https://www.kaggle.com/blackmoon/russian-language-toxic-comments</a>  , <a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling</a>\n(Magic) Validate and fine-tuned by language: <a href=\"https://www.kaggle.com/medrau/train-from-mlm-finetuned-val-per-lang\">https://www.kaggle.com/medrau/train-from-mlm-finetuned-val-per-lang</a></p>",
  "messages": [
    {
      "id": "897613",
      "postDate": "06/23/2020 01:51:39",
      "content": "<p>Big thanks to <a href=\"/riblidezso\">@riblidezso</a> who share this amazing notebook: <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a> .</p>\n\n<p>In short, ensemble and second step fine-tuned by language is magic (and overfitting too).</p>\n\n<h1>PB score timeline</h1>\n\n<p>Ensemble Gmean of top score public submissions -&gt; 0.9473 (8 June)\nEnsemble LGBM solution and Gmean submission -&gt; 0.9478 (10 June)\n(Magic) Add external data to validate dataset and change to validate by language in second step training of XLM and ensemble by language -&gt; 0.9500 (21 June)\nUnfortunately, we boosted the score to 0.9500 in the last 2 days but no more TPU left to train new external data.</p>\n\n<h1>Validate Dataset</h1>\n\n<p>We separated validate data to 3 set by language and find another 3 languages from external data to use in second step training in <a href=\"/riblidezso\">@riblidezso</a> notebook. The out of fold score looks promising so we blend them to our main submission and boost the score to 0.9500.</p>\n\n<h1>Ref</h1>\n\n<p>Gmean notebook: <a href=\"https://www.kaggle.com/paulorzp/gmean-of-low-correlation-lb-0-952x\">https://www.kaggle.com/paulorzp/gmean-of-low-correlation-lb-0-952x</a>\nXLM-Roberta: <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a>\nLGBM solution: <a href=\"https://www.kaggle.com/miklgr500/lgbm-solution\">https://www.kaggle.com/miklgr500/lgbm-solution</a>\nExternal data: <a href=\"https://www.kaggle.com/blackmoon/russian-language-toxic-comments\">https://www.kaggle.com/blackmoon/russian-language-toxic-comments</a>  , <a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling</a>\n(Magic) Validate and fine-tuned by language: <a href=\"https://www.kaggle.com/medrau/train-from-mlm-finetuned-val-per-lang\">https://www.kaggle.com/medrau/train-from-mlm-finetuned-val-per-lang</a></p>",
      "rawMarkdown": "Big thanks to @riblidezso who share this amazing notebook: https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large .\n\nIn short, ensemble and second step fine-tuned by language is magic (and overfitting too).\n\n# PB score timeline\n\nEnsemble Gmean of top score public submissions -&gt; 0.9473 (8 June)\nEnsemble LGBM solution and Gmean submission -&gt; 0.9478 (10 June)\n(Magic) Add external data to validate dataset and change to validate by language in second step training of XLM and ensemble by language -&gt; 0.9500 (21 June)\nUnfortunately, we boosted the score to 0.9500 in the last 2 days but no more TPU left to train new external data.\n\n# Validate Dataset \n\nWe separated validate data to 3 set by language and find another 3 languages from external data to use in second step training in @riblidezso notebook. The out of fold score looks promising so we blend them to our main submission and boost the score to 0.9500.\n\n# Ref\n\nGmean notebook: https://www.kaggle.com/paulorzp/gmean-of-low-correlation-lb-0-952x\nXLM-Roberta: https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\nLGBM solution: https://www.kaggle.com/miklgr500/lgbm-solution\nExternal data: https://www.kaggle.com/blackmoon/russian-language-toxic-comments  , https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\n(Magic) Validate and fine-tuned by language: https://www.kaggle.com/medrau/train-from-mlm-finetuned-val-per-lang",
      "votes": null
    },
    {
      "id": "897662",
      "postDate": "06/23/2020 03:00:55",
      "content": "<p>Hey, this is a great solution. but I guess you am referring to an incorrect the gmean notebook. it refers to a different competition</p>",
      "rawMarkdown": "Hey, this is a great solution. but I guess you am referring to an incorrect the gmean notebook. it refers to a different competition",
      "votes": null
    },
    {
      "id": "897670",
      "postDate": "06/23/2020 03:06:56",
      "content": "<p>The idea is based on this notebook so just change submission files to this competition should work.</p>",
      "rawMarkdown": "The idea is based on this notebook so just change submission files to this competition should work.",
      "votes": null
    },
    {
      "id": "899145",
      "postDate": "06/24/2020 03:00:53",
      "content": "<p>Nice man,I now regret not trying the second level language specific training because I thought it will overfit.Thanks for sharing <a href=\"/medrau\">@medrau</a> </p>",
      "rawMarkdown": "Nice man,I now regret not trying the second level language specific training because I thought it will overfit.Thanks for sharing @medrau",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 897662,
      "author_name": "roydatascience",
      "author_url": "",
      "post_date": "06/23/2020 03:00:55",
      "content": "<p>Hey, this is a great solution. but I guess you am referring to an incorrect the gmean notebook. it refers to a different competition</p>",
      "votes": null,
      "replies": [
        {
          "id": 897670,
          "author_name": "medrau",
          "author_url": "",
          "post_date": "06/23/2020 03:06:56",
          "content": "<p>The idea is based on this notebook so just change submission files to this competition should work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 899145,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "06/24/2020 03:00:53",
      "content": "<p>Nice man,I now regret not trying the second level language specific training because I thought it will overfit.Thanks for sharing <a href=\"/medrau\">@medrau</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "897613": "Big thanks to @riblidezso who share this amazing notebook: https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large .\n\nIn short, ensemble and second step fine-tuned by language is magic (and overfitting too).\n\n# PB score timeline\n\nEnsemble Gmean of top score public submissions -&gt; 0.9473 (8 June)\nEnsemble LGBM solution and Gmean submission -&gt; 0.9478 (10 June)\n(Magic) Add external data to validate dataset and change to validate by language in second step training of XLM and ensemble by language -&gt; 0.9500 (21 June)\nUnfortunately, we boosted the score to 0.9500 in the last 2 days but no more TPU left to train new external data.\n\n# Validate Dataset \n\nWe separated validate data to 3 set by language and find another 3 languages from external data to use in second step training in @riblidezso notebook. The out of fold score looks promising so we blend them to our main submission and boost the score to 0.9500.\n\n# Ref\n\nGmean notebook: https://www.kaggle.com/paulorzp/gmean-of-low-correlation-lb-0-952x\nXLM-Roberta: https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\nLGBM solution: https://www.kaggle.com/miklgr500/lgbm-solution\nExternal data: https://www.kaggle.com/blackmoon/russian-language-toxic-comments  , https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\n(Magic) Validate and fine-tuned by language: https://www.kaggle.com/medrau/train-from-mlm-finetuned-val-per-lang",
    "897662": "Hey, this is a great solution. but I guess you am referring to an incorrect the gmean notebook. it refers to a different competition",
    "897670": "The idea is based on this notebook so just change submission files to this competition should work.",
    "899145": "Nice man,I now regret not trying the second level language specific training because I thought it will overfit.Thanks for sharing @medrau"
  },
  "source": "meta"
}