{
  "id": 160892,
  "title": "16th Place Solution: Simple Calibrations",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/ods-ai-sxwat-16th-place-solution-simple-calibratio",
  "author_name": "",
  "post_date": "2020-06-23T03:25:27.260331Z",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Congratulations to winner! \nThanks to Kaggle team, organizers, educative and interesting notebooks ( @riblidezso , @shonenkov , @jazivxt ).\nThat was my first nlp competition, here I acquainted with SOTA models in NLP (GloVe/FastText + LSTMs -&gt; Transformer -&gt; BERT -&gt; XLM-ROBERTa, excited journey).\nTake a gold medal will be too good to be true (I am the one who lost the most positions in top10), but anyway I glad to take my silver medal (the only one my medal 😃 ).</p>\n\n<h1>Things that worked for me</h1>\n\n<ol>\n<li>XLM-ROBERTa with LSTM + MAXPOOLING head lead to <strong>0.9328 private</strong> (0.9346 public)</li>\n<li>Ensemble models from 1st step <strong>0.9453 private</strong> (0.9474 public)</li>\n<li>Find bot comments by popular templates <strong>0.9458 private</strong> (0.9481 public)</li>\n<li>Make simple calibration to the every lang in test set (convert pd from 3rd step to logit and then add some constants for every lang to logit, so it occurs that I need to add positive constant for fr and es, negative to it and pt), such additions give my final <strong>0.9482 private</strong> (0.9505 public) </li>\n</ol>\n\n<p>As I can see now I was overfitted from 2nd step on ensembles, didn't find good CV strategy</p>",
  "messages": [
    {
      "id": "897692",
      "postDate": "06/23/2020 03:25:27",
      "content": "<p>Congratulations to winner! \nThanks to Kaggle team, organizers, educative and interesting notebooks ( @riblidezso , @shonenkov , @jazivxt ).\nThat was my first nlp competition, here I acquainted with SOTA models in NLP (GloVe/FastText + LSTMs -&gt; Transformer -&gt; BERT -&gt; XLM-ROBERTa, excited journey).\nTake a gold medal will be too good to be true (I am the one who lost the most positions in top10), but anyway I glad to take my silver medal (the only one my medal 😃 ).</p>\n\n<h1>Things that worked for me</h1>\n\n<ol>\n<li>XLM-ROBERTa with LSTM + MAXPOOLING head lead to <strong>0.9328 private</strong> (0.9346 public)</li>\n<li>Ensemble models from 1st step <strong>0.9453 private</strong> (0.9474 public)</li>\n<li>Find bot comments by popular templates <strong>0.9458 private</strong> (0.9481 public)</li>\n<li>Make simple calibration to the every lang in test set (convert pd from 3rd step to logit and then add some constants for every lang to logit, so it occurs that I need to add positive constant for fr and es, negative to it and pt), such additions give my final <strong>0.9482 private</strong> (0.9505 public) </li>\n</ol>\n\n<p>As I can see now I was overfitted from 2nd step on ensembles, didn't find good CV strategy</p>",
      "rawMarkdown": "Congratulations to winner! \nThanks to Kaggle team, organizers, educative and interesting notebooks ( @riblidezso , @shonenkov , @jazivxt ).\nThat was my first nlp competition, here I acquainted with SOTA models in NLP (GloVe/FastText + LSTMs -&gt; Transformer -&gt; BERT -&gt; XLM-ROBERTa, excited journey).\nTake a gold medal will be too good to be true (I am the one who lost the most positions in top10), but anyway I glad to take my silver medal (the only one my medal 😃 ).\n\n# Things that worked for me \n1. XLM-ROBERTa with LSTM + MAXPOOLING head lead to **0.9328 private** (0.9346 public)\n2. Ensemble models from 1st step **0.9453 private** (0.9474 public)\n3. Find bot comments by popular templates **0.9458 private** (0.9481 public)\n4. Make simple calibration to the every lang in test set (convert pd from 3rd step to logit and then add some constants for every lang to logit, so it occurs that I need to add positive constant for fr and es, negative to it and pt), such additions give my final **0.9482 private** (0.9505 public) \n\nAs I can see now I was overfitted from 2nd step on ensembles, didn't find good CV strategy",
      "votes": null
    },
    {
      "id": "897697",
      "postDate": "06/23/2020 03:29:22",
      "content": "<p><a href=\"/aybatov\">@aybatov</a> congrats!</p>",
      "rawMarkdown": "aybatov congrats!",
      "votes": null
    },
    {
      "id": "897701",
      "postDate": "06/23/2020 03:35:08",
      "content": "<p>Well done !</p>",
      "rawMarkdown": "Well done !",
      "votes": null
    },
    {
      "id": "897712",
      "postDate": "06/23/2020 03:45:56",
      "content": "<p>Well done!  Very nice solution.    If you create more diversity models and ensemble them, then you should have a gold medal .  </p>",
      "rawMarkdown": "Well done!  Very nice solution.    If you create more diversity models and ensemble them, then you should have a gold medal .",
      "votes": null
    },
    {
      "id": "901032",
      "postDate": "06/25/2020 08:04:19",
      "content": "<p>Nice. Congratulations to you.</p>",
      "rawMarkdown": "Nice. Congratulations to you.",
      "votes": null
    },
    {
      "id": "906744",
      "postDate": "06/29/2020 14:21:01",
      "content": "<p>Great Work!!</p>",
      "rawMarkdown": "Great Work!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 897697,
      "author_name": "rohitsingh9990",
      "author_url": "",
      "post_date": "06/23/2020 03:29:22",
      "content": "<p><a href=\"/aybatov\">@aybatov</a> congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 897701,
      "author_name": "haythemtellili5",
      "author_url": "",
      "post_date": "06/23/2020 03:35:08",
      "content": "<p>Well done !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 897712,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "06/23/2020 03:45:56",
      "content": "<p>Well done!  Very nice solution.    If you create more diversity models and ensemble them, then you should have a gold medal .  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 901032,
      "author_name": "winterbreeze",
      "author_url": "",
      "post_date": "06/25/2020 08:04:19",
      "content": "<p>Nice. Congratulations to you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 906744,
      "author_name": "shubhamksingh",
      "author_url": "",
      "post_date": "06/29/2020 14:21:01",
      "content": "<p>Great Work!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "897692": "Congratulations to winner! \nThanks to Kaggle team, organizers, educative and interesting notebooks ( @riblidezso , @shonenkov , @jazivxt ).\nThat was my first nlp competition, here I acquainted with SOTA models in NLP (GloVe/FastText + LSTMs -&gt; Transformer -&gt; BERT -&gt; XLM-ROBERTa, excited journey).\nTake a gold medal will be too good to be true (I am the one who lost the most positions in top10), but anyway I glad to take my silver medal (the only one my medal 😃 ).\n\n# Things that worked for me \n1. XLM-ROBERTa with LSTM + MAXPOOLING head lead to **0.9328 private** (0.9346 public)\n2. Ensemble models from 1st step **0.9453 private** (0.9474 public)\n3. Find bot comments by popular templates **0.9458 private** (0.9481 public)\n4. Make simple calibration to the every lang in test set (convert pd from 3rd step to logit and then add some constants for every lang to logit, so it occurs that I need to add positive constant for fr and es, negative to it and pt), such additions give my final **0.9482 private** (0.9505 public) \n\nAs I can see now I was overfitted from 2nd step on ensembles, didn't find good CV strategy",
    "897697": "aybatov congrats!",
    "897701": "Well done !",
    "897712": "Well done!  Very nice solution.    If you create more diversity models and ensemble them, then you should have a gold medal .",
    "901032": "Nice. Congratulations to you.",
    "906744": "Great Work!!"
  },
  "source": "meta"
}