{
  "id": 161014,
  "title": "🏅 Best single models 🏅",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/la-compa-a-easy-best-single-models",
  "author_name": "",
  "post_date": "2020-06-23T13:38:28.790555700Z",
  "votes": 17,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I wonder what were the best singles models around.</p>\n\n<p>From the write ups it seems like most people got XLM-R around 0.942* and heavily rely on ensembling, post-processing and other tricks (not diminishing their amazing achievement).</p>\n\n<p>Anyway, I'd like to share a <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\">single model</a> (weights and code) with pseudo labels and knowledge distillation that got to <strong>LB 0.9475</strong> (and <strong>0.9487</strong> after applying the hardcoded multipliers from @christofhenkel thx!) here <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\">https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model</a></p>\n\n<p>Unfortunately we failed to ensemble as well as others and did not explore the language modifiers beforehand. This comp was an amazing learning experience for me doing NLP for the first time in my life :D</p>\n\n<p>The main take out for me was that relying to heavily on pseudo labels seem to have made blending a much harder task (knowledge distillation seem to have worsen that issue). It took me too long to realize that but lesson learned for the next comps. I'll extend my report when I find some more time, but please let me know your thoughts on that! \n(and please share your models if you can)</p>\n\n<p>Cheers!</p>",
  "messages": [
    {
      "id": "898366",
      "postDate": "06/23/2020 13:38:28",
      "content": "<p>I wonder what were the best singles models around.</p>\n\n<p>From the write ups it seems like most people got XLM-R around 0.942* and heavily rely on ensembling, post-processing and other tricks (not diminishing their amazing achievement).</p>\n\n<p>Anyway, I'd like to share a <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\">single model</a> (weights and code) with pseudo labels and knowledge distillation that got to <strong>LB 0.9475</strong> (and <strong>0.9487</strong> after applying the hardcoded multipliers from @christofhenkel thx!) here <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\">https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model</a></p>\n\n<p>Unfortunately we failed to ensemble as well as others and did not explore the language modifiers beforehand. This comp was an amazing learning experience for me doing NLP for the first time in my life :D</p>\n\n<p>The main take out for me was that relying to heavily on pseudo labels seem to have made blending a much harder task (knowledge distillation seem to have worsen that issue). It took me too long to realize that but lesson learned for the next comps. I'll extend my report when I find some more time, but please let me know your thoughts on that! \n(and please share your models if you can)</p>\n\n<p>Cheers!</p>",
      "rawMarkdown": "I wonder what were the best singles models around.\n\nFrom the write ups it seems like most people got XLM-R around 0.942* and heavily rely on ensembling, post-processing and other tricks (not diminishing their amazing achievement).\n\nAnyway, I'd like to share a [single model](https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model) (weights and code) with pseudo labels and knowledge distillation that got to **LB 0.9475** (and **0.9487** after applying the hardcoded multipliers from @christofhenkel thx!) here https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\n\nUnfortunately we failed to ensemble as well as others and did not explore the language modifiers beforehand. This comp was an amazing learning experience for me doing NLP for the first time in my life :D\n\nThe main take out for me was that relying to heavily on pseudo labels seem to have made blending a much harder task (knowledge distillation seem to have worsen that issue). It took me too long to realize that but lesson learned for the next comps. I'll extend my report when I find some more time, but please let me know your thoughts on that! \n(and please share your models if you can)\n\nCheers!",
      "votes": null
    },
    {
      "id": "898580",
      "postDate": "06/23/2020 15:44:40",
      "content": "<p>Very interesting share Henrique. Thanks. Dr.</p>",
      "rawMarkdown": "Very interesting share Henrique. Thanks. Dr.",
      "votes": null
    },
    {
      "id": "898629",
      "postDate": "06/23/2020 16:34:37",
      "content": "<p>Thanks <a href=\"/drpatrickchan\">@drpatrickchan</a> Congrats on your first gold medal and the Comp Master! :)</p>",
      "rawMarkdown": "Thanks @drpatrickchan Congrats on your first gold medal and the Comp Master! :)",
      "votes": null
    },
    {
      "id": "898790",
      "postDate": "06/23/2020 18:38:49",
      "content": "<p>wondering what is hardcoded multipliers. thanks</p>",
      "rawMarkdown": "wondering what is hardcoded multipliers. thanks",
      "votes": null
    },
    {
      "id": "898919",
      "postDate": "06/23/2020 20:43:53",
      "content": "<p>0.9431 for us using XLM-Roberta, pseudo labelling, uniform language distribution, 50/50 toxicity distribution.\nLR of 1e-5 and finetuning LM on train dataset</p>",
      "rawMarkdown": "0.9431 for us using XLM-Roberta, pseudo labelling, uniform language distribution, 50/50 toxicity distribution.\nLR of 1e-5 and finetuning LM on train dataset",
      "votes": null
    },
    {
      "id": "899018",
      "postDate": "06/23/2020 22:41:08",
      "content": "<p>It's fairly trivial to distill an ensemble into a single model. As an experiment during the competition, I distilled my team's predictions (at public LB 0.9552) into a single XLM-R model and hit 0.9487 without PP. </p>\n\n<p>The problem with single multilingual models is you hit the curse of multilinguality and the problem of a limited size vocabulary covering multiple languages. Even with XLM-R which acknowledges and tries to address this problem - they have I believe a 250K shared vocab covering 100 languages. There's a loss of fidelity at the embeddings level.</p>",
      "rawMarkdown": "It's fairly trivial to distill an ensemble into a single model. As an experiment during the competition, I distilled my team's predictions (at public LB 0.9552) into a single XLM-R model and hit 0.9487 without PP. \n\nThe problem with single multilingual models is you hit the curse of multilinguality and the problem of a limited size vocabulary covering multiple languages. Even with XLM-R which acknowledges and tries to address this problem - they have I believe a 250K shared vocab covering 100 languages. There's a loss of fidelity at the embeddings level.",
      "votes": null
    },
    {
      "id": "899107",
      "postDate": "06/24/2020 01:51:42",
      "content": "<p>see the 4th place discussion, Post-Processing section, there <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980</a></p>",
      "rawMarkdown": "see the 4th place discussion, Post-Processing section, there https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980",
      "votes": null
    },
    {
      "id": "899361",
      "postDate": "06/24/2020 07:26:32",
      "content": "<p>thanks for let us know <a href=\"/leecming\">@leecming</a> your <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862\">write-up</a> is very inspiring!\nI assume from your msg that you haven't used your distilled model in your final ensemble, is that right?\nMoreover, how do you prevent your use of PL to interfere with the ensemble diversification?</p>\n\n<p>Indeed, I agree you that the vocab embeddings seem to be the bottle neck of the current models, both in learning capacity as well as in number of parameters, which hits me as a possible improvement to the NLP models. However, that might not be a major issue on the new huge models like GPT3, we'll see...\nMy team mate <a href=\"/ratthachat\">@ratthachat</a> and I tried to mitigate the issue adding more layers at the top of the model instead, but only with some moderate success.</p>",
      "rawMarkdown": "thanks for let us know @leecming your [write-up](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862) is very inspiring!\nI assume from your msg that you haven't used your distilled model in your final ensemble, is that right?\nMoreover, how do you prevent your use of PL to interfere with the ensemble diversification?\n\nIndeed, I agree you that the vocab embeddings seem to be the bottle neck of the current models, both in learning capacity as well as in number of parameters, which hits me as a possible improvement to the NLP models. However, that might not be a major issue on the new huge models like GPT3, we'll see...\nMy team mate @ratthachat and I tried to mitigate the issue adding more layers at the top of the model instead, but only with some moderate success.",
      "votes": null
    },
    {
      "id": "899390",
      "postDate": "06/24/2020 07:54:35",
      "content": "<p><a href=\"/konohayui\">@konohayui</a> you can see it at the last section of the model <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model#Apply-language-multipliers\">https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model#Apply-language-multipliers</a> </p>",
      "rawMarkdown": "konohayui you can see it at the last section of the model https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model#Apply-language-multipliers",
      "votes": null
    },
    {
      "id": "899457",
      "postDate": "06/24/2020 08:53:35",
      "content": "<p>Wonderful 👊</p>",
      "rawMarkdown": "Wonderful 👊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 898580,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "06/23/2020 15:44:40",
      "content": "<p>Very interesting share Henrique. Thanks. Dr.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898629,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "06/23/2020 16:34:37",
          "content": "<p>Thanks <a href=\"/drpatrickchan\">@drpatrickchan</a> Congrats on your first gold medal and the Comp Master! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898790,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "06/23/2020 18:38:49",
      "content": "<p>wondering what is hardcoded multipliers. thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 899107,
          "author_name": "sebastienm",
          "author_url": "",
          "post_date": "06/24/2020 01:51:42",
          "content": "<p>see the 4th place discussion, Post-Processing section, there <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 899390,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "06/24/2020 07:54:35",
          "content": "<p><a href=\"/konohayui\">@konohayui</a> you can see it at the last section of the model <a href=\"https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model#Apply-language-multipliers\">https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model#Apply-language-multipliers</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898919,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/23/2020 20:43:53",
      "content": "<p>0.9431 for us using XLM-Roberta, pseudo labelling, uniform language distribution, 50/50 toxicity distribution.\nLR of 1e-5 and finetuning LM on train dataset</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899018,
      "author_name": "leecming",
      "author_url": "",
      "post_date": "06/23/2020 22:41:08",
      "content": "<p>It's fairly trivial to distill an ensemble into a single model. As an experiment during the competition, I distilled my team's predictions (at public LB 0.9552) into a single XLM-R model and hit 0.9487 without PP. </p>\n\n<p>The problem with single multilingual models is you hit the curse of multilinguality and the problem of a limited size vocabulary covering multiple languages. Even with XLM-R which acknowledges and tries to address this problem - they have I believe a 250K shared vocab covering 100 languages. There's a loss of fidelity at the embeddings level.</p>",
      "votes": null,
      "replies": [
        {
          "id": 899361,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "06/24/2020 07:26:32",
          "content": "<p>thanks for let us know <a href=\"/leecming\">@leecming</a> your <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862\">write-up</a> is very inspiring!\nI assume from your msg that you haven't used your distilled model in your final ensemble, is that right?\nMoreover, how do you prevent your use of PL to interfere with the ensemble diversification?</p>\n\n<p>Indeed, I agree you that the vocab embeddings seem to be the bottle neck of the current models, both in learning capacity as well as in number of parameters, which hits me as a possible improvement to the NLP models. However, that might not be a major issue on the new huge models like GPT3, we'll see...\nMy team mate <a href=\"/ratthachat\">@ratthachat</a> and I tried to mitigate the issue adding more layers at the top of the model instead, but only with some moderate success.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 899457,
      "author_name": "ahmetfurkandemr",
      "author_url": "",
      "post_date": "06/24/2020 08:53:35",
      "content": "<p>Wonderful 👊</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "898366": "I wonder what were the best singles models around.\n\nFrom the write ups it seems like most people got XLM-R around 0.942* and heavily rely on ensembling, post-processing and other tricks (not diminishing their amazing achievement).\n\nAnyway, I'd like to share a [single model](https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model) (weights and code) with pseudo labels and knowledge distillation that got to **LB 0.9475** (and **0.9487** after applying the hardcoded multipliers from @christofhenkel thx!) here https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model\n\nUnfortunately we failed to ensemble as well as others and did not explore the language modifiers beforehand. This comp was an amazing learning experience for me doing NLP for the first time in my life :D\n\nThe main take out for me was that relying to heavily on pseudo labels seem to have made blending a much harder task (knowledge distillation seem to have worsen that issue). It took me too long to realize that but lesson learned for the next comps. I'll extend my report when I find some more time, but please let me know your thoughts on that! \n(and please share your models if you can)\n\nCheers!",
    "898580": "Very interesting share Henrique. Thanks. Dr.",
    "898629": "Thanks @drpatrickchan Congrats on your first gold medal and the Comp Master! :)",
    "898790": "wondering what is hardcoded multipliers. thanks",
    "898919": "0.9431 for us using XLM-Roberta, pseudo labelling, uniform language distribution, 50/50 toxicity distribution.\nLR of 1e-5 and finetuning LM on train dataset",
    "899018": "It's fairly trivial to distill an ensemble into a single model. As an experiment during the competition, I distilled my team's predictions (at public LB 0.9552) into a single XLM-R model and hit 0.9487 without PP. \n\nThe problem with single multilingual models is you hit the curse of multilinguality and the problem of a limited size vocabulary covering multiple languages. Even with XLM-R which acknowledges and tries to address this problem - they have I believe a 250K shared vocab covering 100 languages. There's a loss of fidelity at the embeddings level.",
    "899107": "see the 4th place discussion, Post-Processing section, there https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980",
    "899361": "thanks for let us know @leecming your [write-up](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160862) is very inspiring!\nI assume from your msg that you haven't used your distilled model in your final ensemble, is that right?\nMoreover, how do you prevent your use of PL to interfere with the ensemble diversification?\n\nIndeed, I agree you that the vocab embeddings seem to be the bottle neck of the current models, both in learning capacity as well as in number of parameters, which hits me as a possible improvement to the NLP models. However, that might not be a major issue on the new huge models like GPT3, we'll see...\nMy team mate @ratthachat and I tried to mitigate the issue adding more layers at the top of the model instead, but only with some moderate success.",
    "899390": "konohayui you can see it at the last section of the model https://www.kaggle.com/hmendonca/jigsaw20-xlm-r-lb0-9487-singel-model#Apply-language-multipliers",
    "899457": "Wonderful 👊"
  },
  "source": "meta"
}