{
  "id": 160885,
  "title": "75th place solution ",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/the-underfit-team-75th-place-solution",
  "author_name": "",
  "post_date": "2020-06-23T03:02:15.580634800Z",
  "votes": 21,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Congratulations everyone!\nFirst of all my teammates  @rftexas  @tanulsingh077 and  @ibtesama  was really awesome and supportive.</p>\n\n<p><strong>Overview of our models</strong></p>\n\n<ol>\n<li><p>Fold CV with 250k translated data and 8k validation dataset using custom pre-trained XLM Roberta large version. ( score : 0.9433). </p></li>\n<li><p>Single fold normal training XLM large and prediction on Test Time Augmentation (TTA). Test translated to 6 languages and avg predictions.( score : 0.9427)</p></li>\n<li><p>Single fold training on pre-trained XLM large with Multisample dropout architecture.(score 0.9417)</p></li>\n<li><p>Three Stage training on training data.\n                            - ( Train on English -&gt; inference)\n                            - (Train on translated  lang -&gt; inference)\n                            - (Train on validation -&gt; inference)\n                            Finally weighted average of these predictions gave us (0.9399)</p></li>\n<li><p>Two-stage XLM Base with 400k samples and TTA. (0.9363)</p></li>\n</ol>\n\n<p><strong>Tricks that worked</strong></p>\n\n<ul>\n<li>Token length = 224</li>\n<li>Balanced the dataset in language distribution.</li>\n<li>Balanced the dataset in target class distribution ( 125k for each class)</li>\n<li>Head + tail encoding.</li>\n<li>Learning Rate scheduling.</li>\n<li>Label Smoothing.</li>\n</ul>\n\n<p>Finally blending using <a href=\"https://www.kaggle.com/paulorzp/gmean-of-light-gbm-models-lb-0-951x\">GMEAN blend.</a></p>",
  "messages": [
    {
      "id": "897664",
      "postDate": "06/23/2020 03:02:15",
      "content": "<p>Congratulations everyone!\nFirst of all my teammates  @rftexas  @tanulsingh077 and  @ibtesama  was really awesome and supportive.</p>\n\n<p><strong>Overview of our models</strong></p>\n\n<ol>\n<li><p>Fold CV with 250k translated data and 8k validation dataset using custom pre-trained XLM Roberta large version. ( score : 0.9433). </p></li>\n<li><p>Single fold normal training XLM large and prediction on Test Time Augmentation (TTA). Test translated to 6 languages and avg predictions.( score : 0.9427)</p></li>\n<li><p>Single fold training on pre-trained XLM large with Multisample dropout architecture.(score 0.9417)</p></li>\n<li><p>Three Stage training on training data.\n                            - ( Train on English -&gt; inference)\n                            - (Train on translated  lang -&gt; inference)\n                            - (Train on validation -&gt; inference)\n                            Finally weighted average of these predictions gave us (0.9399)</p></li>\n<li><p>Two-stage XLM Base with 400k samples and TTA. (0.9363)</p></li>\n</ol>\n\n<p><strong>Tricks that worked</strong></p>\n\n<ul>\n<li>Token length = 224</li>\n<li>Balanced the dataset in language distribution.</li>\n<li>Balanced the dataset in target class distribution ( 125k for each class)</li>\n<li>Head + tail encoding.</li>\n<li>Learning Rate scheduling.</li>\n<li>Label Smoothing.</li>\n</ul>\n\n<p>Finally blending using <a href=\"https://www.kaggle.com/paulorzp/gmean-of-light-gbm-models-lb-0-951x\">GMEAN blend.</a></p>",
      "rawMarkdown": "Congratulations everyone!\nFirst of all my teammates  @rftexas  @tanulsingh077 and  @ibtesama  was really awesome and supportive.\n\n**Overview of our models**\n\n1.  Fold CV with 250k translated data and 8k validation dataset using custom pre-trained XLM Roberta large version. ( score : 0.9433). \n\n2. Single fold normal training XLM large and prediction on Test Time Augmentation (TTA). Test translated to 6 languages and avg predictions.( score : 0.9427)\n\n3. Single fold training on pre-trained XLM large with Multisample dropout architecture.(score 0.9417)\n\n4. Three Stage training on training data.\n                                - ( Train on English -&gt; inference)\n                                - (Train on translated  lang -&gt; inference)\n                                - (Train on validation -&gt; inference)\n                                Finally weighted average of these predictions gave us (0.9399)\n\n5.  Two-stage XLM Base with 400k samples and TTA. (0.9363)\n\n\n**Tricks that worked**\n\n- Token length = 224\n- Balanced the dataset in language distribution.\n- Balanced the dataset in target class distribution ( 125k for each class)\n- Head + tail encoding.\n- Learning Rate scheduling.\n- Label Smoothing.\n\nFinally blending using [GMEAN blend.](https://www.kaggle.com/paulorzp/gmean-of-light-gbm-models-lb-0-951x)",
      "votes": null
    },
    {
      "id": "897665",
      "postDate": "06/23/2020 03:03:49",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "897666",
      "postDate": "06/23/2020 03:04:35",
      "content": "<p>congrachulations <a href=\"/shahules\">@shahules</a>.</p>",
      "rawMarkdown": "congrachulations @shahules.",
      "votes": null
    },
    {
      "id": "897667",
      "postDate": "06/23/2020 03:04:43",
      "content": "<p>Congrats bro :)</p>",
      "rawMarkdown": "Congrats bro :)",
      "votes": null
    },
    {
      "id": "897668",
      "postDate": "06/23/2020 03:05:04",
      "content": "<p>Thanks buddy <a href=\"/rohitsingh9990\">@rohitsingh9990</a> </p>",
      "rawMarkdown": "Thanks buddy @rohitsingh9990",
      "votes": null
    },
    {
      "id": "897786",
      "postDate": "06/23/2020 05:05:58",
      "content": "<p>Congratulations <a href=\"/shahules\">@shahules</a> </p>",
      "rawMarkdown": "Congratulations @shahules",
      "votes": null
    },
    {
      "id": "897791",
      "postDate": "06/23/2020 05:15:47",
      "content": "<p>Thanks man <a href=\"/kurianbenoy\">@kurianbenoy</a> </p>",
      "rawMarkdown": "Thanks man @kurianbenoy",
      "votes": null
    },
    {
      "id": "897796",
      "postDate": "06/23/2020 05:24:57",
      "content": "<p>Congratz to the team, interesting read!</p>",
      "rawMarkdown": "Congratz to the team, interesting read!",
      "votes": null
    },
    {
      "id": "898060",
      "postDate": "06/23/2020 08:52:59",
      "content": "<p>Well done bro! </p>",
      "rawMarkdown": "Well done bro!",
      "votes": null
    },
    {
      "id": "898154",
      "postDate": "06/23/2020 10:48:15",
      "content": "<p>we did every thing was similar but gmean man well deserved <a href=\"/shahules\">@shahules</a> and team</p>",
      "rawMarkdown": "we did every thing was similar but gmean man well deserved @shahules and team",
      "votes": null
    },
    {
      "id": "900055",
      "postDate": "06/24/2020 15:37:04",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "901132",
      "postDate": "06/25/2020 09:13:34",
      "content": "<p>Congrats mate!</p>",
      "rawMarkdown": "Congrats mate!",
      "votes": null
    },
    {
      "id": "904337",
      "postDate": "06/27/2020 14:48:09",
      "content": "<p>congratulations!!</p>",
      "rawMarkdown": "congratulations!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 897665,
      "author_name": "haythemtellili5",
      "author_url": "",
      "post_date": "06/23/2020 03:03:49",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 897667,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "06/23/2020 03:04:43",
          "content": "<p>Congrats bro :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897666,
      "author_name": "rohitsingh9990",
      "author_url": "",
      "post_date": "06/23/2020 03:04:35",
      "content": "<p>congrachulations <a href=\"/shahules\">@shahules</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 897668,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "06/23/2020 03:05:04",
          "content": "<p>Thanks buddy <a href=\"/rohitsingh9990\">@rohitsingh9990</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897786,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "06/23/2020 05:05:58",
      "content": "<p>Congratulations <a href=\"/shahules\">@shahules</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 897791,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "06/23/2020 05:15:47",
          "content": "<p>Thanks man <a href=\"/kurianbenoy\">@kurianbenoy</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897796,
      "author_name": "frtgnn",
      "author_url": "",
      "post_date": "06/23/2020 05:24:57",
      "content": "<p>Congratz to the team, interesting read!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898060,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/23/2020 08:52:59",
      "content": "<p>Well done bro! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898154,
      "author_name": "pranshu29",
      "author_url": "",
      "post_date": "06/23/2020 10:48:15",
      "content": "<p>we did every thing was similar but gmean man well deserved <a href=\"/shahules\">@shahules</a> and team</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 900055,
      "author_name": "mdselimreza",
      "author_url": "",
      "post_date": "06/24/2020 15:37:04",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 901132,
      "author_name": "namanj27",
      "author_url": "",
      "post_date": "06/25/2020 09:13:34",
      "content": "<p>Congrats mate!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 904337,
      "author_name": "rahulharlalka",
      "author_url": "",
      "post_date": "06/27/2020 14:48:09",
      "content": "<p>congratulations!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "897664": "Congratulations everyone!\nFirst of all my teammates  @rftexas  @tanulsingh077 and  @ibtesama  was really awesome and supportive.\n\n**Overview of our models**\n\n1.  Fold CV with 250k translated data and 8k validation dataset using custom pre-trained XLM Roberta large version. ( score : 0.9433). \n\n2. Single fold normal training XLM large and prediction on Test Time Augmentation (TTA). Test translated to 6 languages and avg predictions.( score : 0.9427)\n\n3. Single fold training on pre-trained XLM large with Multisample dropout architecture.(score 0.9417)\n\n4. Three Stage training on training data.\n                                - ( Train on English -&gt; inference)\n                                - (Train on translated  lang -&gt; inference)\n                                - (Train on validation -&gt; inference)\n                                Finally weighted average of these predictions gave us (0.9399)\n\n5.  Two-stage XLM Base with 400k samples and TTA. (0.9363)\n\n\n**Tricks that worked**\n\n- Token length = 224\n- Balanced the dataset in language distribution.\n- Balanced the dataset in target class distribution ( 125k for each class)\n- Head + tail encoding.\n- Learning Rate scheduling.\n- Label Smoothing.\n\nFinally blending using [GMEAN blend.](https://www.kaggle.com/paulorzp/gmean-of-light-gbm-models-lb-0-951x)",
    "897665": "Congratulations!",
    "897666": "congrachulations @shahules.",
    "897667": "Congrats bro :)",
    "897668": "Thanks buddy @rohitsingh9990",
    "897786": "Congratulations @shahules",
    "897791": "Thanks man @kurianbenoy",
    "897796": "Congratz to the team, interesting read!",
    "898060": "Well done bro!",
    "898154": "we did every thing was similar but gmean man well deserved @shahules and team",
    "900055": "Congratulations!",
    "901132": "Congrats mate!",
    "904337": "congratulations!!"
  },
  "source": "meta"
}