{
  "id": 160861,
  "title": "100th place solution (+ GitHub)",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/dimitreoliveira-100th-place-solution-github",
  "author_name": "",
  "post_date": "2020-07-04T16:45:18.960Z",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I did not place so well this time, but still got a medal, so I think it still worth it to share my solution. Congratulation to all winners and new GMs.\nThis was a very interesting competition, being able to work with huge datasets and models was very challenging. Luckily my best submission was the last one that I made 😄. I had many more Ideas to try including using external data but had no time.\nI also have a <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification\">Git repository</a> with my experiments.</p>\n\n<h2>Quick summary</h2>\n\n<ul>\n<li>Model: 5-Fold <code>XLM_RoBERTa large</code> using the last layer <code>CLS</code> token.\n<ul><li>5-Fold <code>XLM_RoBERTa large</code> concatenated <code>AVG</code> and <code>MAX</code> last layer pooled with 8 multi-sample dropout.</li></ul></li>\n<li>Inference: I have used <code>TTA</code> for each sample, I predicted on <code>head</code>, <code>tail</code>, and a mix of both tokens.</li>\n<li>Labels: <code>toxic</code> cast to <code>int</code> then pseudo-labels on the test set.</li>\n<li>Dataset: For the 1st model I used a 1:2 ratio between toxic and not toxic samples, for the 2nd model use 1:1 ratio, also did some basic cleaning. (details below)</li>\n<li>Framework: <code>Tensorflow</code> with TPU.</li>\n</ul>\n\n<h2>Detailed summary</h2>\n\n<h3>Model &amp; training</h3>\n\n<ul>\n<li>1st model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/99-jigsaw-fold1-xlm-roberta-large-best.ipynb\">link for the 1st fold training</a>\n<ul><li>5-Fold XLM_RoBERTa large with the last layer <code>CLS</code> token</li>\n<li>Sequence length: 192</li>\n<li>Batch size: 128</li>\n<li>Epochs: 4</li>\n<li>Learning rate: 1e-5</li>\n<li>Training schedule: Exponential decay with 1 epoch warm-up</li>\n<li>Losses: BinaryCrossentropy on the labels cast to <code>int</code></li></ul></li>\n</ul>\n\n<p>Trained the model 4 epochs on the train set then 1 epoch on the validation set, after that 2 epochs on the test set with pseudo-labels.</p>\n\n<ul>\n<li>2nd model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/136-jigsaw-fold1-xlm-roberta-ratio-1-8-sample-drop.ipynb\">link for the 1st fold training</a>\n<ul><li>5-Fold XLM_RoBERTa large with the last layer <code>AVG</code> and<code>MAX</code> pooled concatenated then fed to 8 multi-sample dropout</li>\n<li>Sequence length: 192</li>\n<li>Batch size: 128</li>\n<li>Epochs: 3</li>\n<li>Learning rate: 1e-5</li>\n<li>Training schedule: Exponential decay with 10% of steps warm-up</li>\n<li>Losses: BinaryCrossentropy on the labels cast to <code>int</code></li></ul></li>\n</ul>\n\n<p>Trained the model 3 epochs on the train set then 2 epochs on the validation set, after that 2 epochs on the test set with pseudo-labels.</p>\n\n<h3>Datasets</h3>\n\n<ul>\n<li>1st model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-2-upper.ipynb\">dataset creation</a> 1:2 toxic to non-toxic samples, 400830 total samples.</li>\n<li>2nd model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-1-clean-polish.ipynb\">dataset creation</a> 1:1 toxic to non-toxic samples, 267220  total samples.</li>\n</ul>\n\n<p>Both models used a similar dataset, upper case text, just a sample of the negative data, data cleaning was just removal of <code>numbers</code>, <code>#hash-tags</code>, <code>@mentios</code>, <code>links</code>, and <code>multiple white spaces</code>. Tokenizer was <code>AutoTokenizer.from_pretrained(</code>jplu/tf-xlm-roberta-large', lowercase=False)` like many more.</p>\n\n<h3>Inference</h3>\n\n<p>For inference, I predicted the <code>head</code>, <code>tail</code>, and a mix of both <code>(50% head &amp;amp; 50%tail)</code> sentence tokens.</p>",
  "messages": [
    {
      "id": "897562",
      "postDate": "06/23/2020 00:46:54",
      "content": "<p>I did not place so well this time, but still got a medal, so I think it still worth it to share my solution. Congratulation to all winners and new GMs.\nThis was a very interesting competition, being able to work with huge datasets and models was very challenging. Luckily my best submission was the last one that I made 😄. I had many more Ideas to try including using external data but had no time.\nI also have a <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification\">Git repository</a> with my experiments.</p>\n\n<h2>Quick summary</h2>\n\n<ul>\n<li>Model: 5-Fold <code>XLM_RoBERTa large</code> using the last layer <code>CLS</code> token.\n<ul><li>5-Fold <code>XLM_RoBERTa large</code> concatenated <code>AVG</code> and <code>MAX</code> last layer pooled with 8 multi-sample dropout.</li></ul></li>\n<li>Inference: I have used <code>TTA</code> for each sample, I predicted on <code>head</code>, <code>tail</code>, and a mix of both tokens.</li>\n<li>Labels: <code>toxic</code> cast to <code>int</code> then pseudo-labels on the test set.</li>\n<li>Dataset: For the 1st model I used a 1:2 ratio between toxic and not toxic samples, for the 2nd model use 1:1 ratio, also did some basic cleaning. (details below)</li>\n<li>Framework: <code>Tensorflow</code> with TPU.</li>\n</ul>\n\n<h2>Detailed summary</h2>\n\n<h3>Model &amp; training</h3>\n\n<ul>\n<li>1st model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/99-jigsaw-fold1-xlm-roberta-large-best.ipynb\">link for the 1st fold training</a>\n<ul><li>5-Fold XLM_RoBERTa large with the last layer <code>CLS</code> token</li>\n<li>Sequence length: 192</li>\n<li>Batch size: 128</li>\n<li>Epochs: 4</li>\n<li>Learning rate: 1e-5</li>\n<li>Training schedule: Exponential decay with 1 epoch warm-up</li>\n<li>Losses: BinaryCrossentropy on the labels cast to <code>int</code></li></ul></li>\n</ul>\n\n<p>Trained the model 4 epochs on the train set then 1 epoch on the validation set, after that 2 epochs on the test set with pseudo-labels.</p>\n\n<ul>\n<li>2nd model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/136-jigsaw-fold1-xlm-roberta-ratio-1-8-sample-drop.ipynb\">link for the 1st fold training</a>\n<ul><li>5-Fold XLM_RoBERTa large with the last layer <code>AVG</code> and<code>MAX</code> pooled concatenated then fed to 8 multi-sample dropout</li>\n<li>Sequence length: 192</li>\n<li>Batch size: 128</li>\n<li>Epochs: 3</li>\n<li>Learning rate: 1e-5</li>\n<li>Training schedule: Exponential decay with 10% of steps warm-up</li>\n<li>Losses: BinaryCrossentropy on the labels cast to <code>int</code></li></ul></li>\n</ul>\n\n<p>Trained the model 3 epochs on the train set then 2 epochs on the validation set, after that 2 epochs on the test set with pseudo-labels.</p>\n\n<h3>Datasets</h3>\n\n<ul>\n<li>1st model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-2-upper.ipynb\">dataset creation</a> 1:2 toxic to non-toxic samples, 400830 total samples.</li>\n<li>2nd model <a href=\"https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-1-clean-polish.ipynb\">dataset creation</a> 1:1 toxic to non-toxic samples, 267220  total samples.</li>\n</ul>\n\n<p>Both models used a similar dataset, upper case text, just a sample of the negative data, data cleaning was just removal of <code>numbers</code>, <code>#hash-tags</code>, <code>@mentios</code>, <code>links</code>, and <code>multiple white spaces</code>. Tokenizer was <code>AutoTokenizer.from_pretrained(</code>jplu/tf-xlm-roberta-large', lowercase=False)` like many more.</p>\n\n<h3>Inference</h3>\n\n<p>For inference, I predicted the <code>head</code>, <code>tail</code>, and a mix of both <code>(50% head &amp;amp; 50%tail)</code> sentence tokens.</p>",
      "rawMarkdown": "I did not place so well this time, but still got a medal, so I think it still worth it to share my solution. Congratulation to all winners and new GMs.\nThis was a very interesting competition, being able to work with huge datasets and models was very challenging. Luckily my best submission was the last one that I made 😄. I had many more Ideas to try including using external data but had no time.\nI also have a [Git repository](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification) with my experiments.\n\n## Quick summary\n- Model: 5-Fold `XLM_RoBERTa large` using the last layer `CLS` token.\n  - 5-Fold `XLM_RoBERTa large` concatenated `AVG` and `MAX` last layer pooled with 8 multi-sample dropout.\n- Inference: I have used `TTA` for each sample, I predicted on `head`, `tail`, and a mix of both tokens.\n- Labels: `toxic` cast to `int` then pseudo-labels on the test set.\n- Dataset: For the 1st model I used a 1:2 ratio between toxic and not toxic samples, for the 2nd model use 1:1 ratio, also did some basic cleaning. (details below)\n- Framework: `Tensorflow` with TPU.\n\n## Detailed summary\n\n### Model &amp; training\n- 1st model [link for the 1st fold training](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/99-jigsaw-fold1-xlm-roberta-large-best.ipynb)\n  - 5-Fold XLM_RoBERTa large with the last layer `CLS` token\n  - Sequence length: 192\n  - Batch size: 128\n  - Epochs: 4\n  - Learning rate: 1e-5\n  - Training schedule: Exponential decay with 1 epoch warm-up\n  - Losses: BinaryCrossentropy on the labels cast to `int`\n\nTrained the model 4 epochs on the train set then 1 epoch on the validation set, after that 2 epochs on the test set with pseudo-labels.\n\n- 2nd model [link for the 1st fold training](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/136-jigsaw-fold1-xlm-roberta-ratio-1-8-sample-drop.ipynb)\n  - 5-Fold XLM_RoBERTa large with the last layer `AVG` and`MAX` pooled concatenated then fed to 8 multi-sample dropout\n  - Sequence length: 192\n  - Batch size: 128\n  - Epochs: 3\n  - Learning rate: 1e-5\n  - Training schedule: Exponential decay with 10% of steps warm-up\n  - Losses: BinaryCrossentropy on the labels cast to `int`\n\nTrained the model 3 epochs on the train set then 2 epochs on the validation set, after that 2 epochs on the test set with pseudo-labels.\n\n### Datasets\n- 1st model [dataset creation](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-2-upper.ipynb) 1:2 toxic to non-toxic samples, 400830 total samples.\n- 2nd model [dataset creation](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-1-clean-polish.ipynb) 1:1 toxic to non-toxic samples, 267220  total samples.\n\nBoth models used a similar dataset, upper case text, just a sample of the negative data, data cleaning was just removal of `numbers`, `#hash-tags`, `@mentios`, `links`, and `multiple white spaces`. Tokenizer was `AutoTokenizer.from_pretrained(`jplu/tf-xlm-roberta-large', lowercase=False)` like many more.\n\n### Inference\nFor inference, I predicted the `head`, `tail`, and a mix of both `(50% head &amp; 50%tail)` sentence tokens.",
      "votes": null
    },
    {
      "id": "897663",
      "postDate": "06/23/2020 03:01:42",
      "content": "<p>Congratulations Dimitre !</p>",
      "rawMarkdown": "Congratulations Dimitre !",
      "votes": null
    },
    {
      "id": "897689",
      "postDate": "06/23/2020 03:23:55",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> congrats!</p>",
      "rawMarkdown": "dimitreoliveira congrats!",
      "votes": null
    },
    {
      "id": "897894",
      "postDate": "06/23/2020 07:06:05",
      "content": "<p>Congratulations! Good job done.</p>\n\n<p>I have some questions. First, what is TTA you mentioned. Second, could you specify how you do pseudo labelling? Finally, what is 8 multi sample dropout? Thanks.</p>",
      "rawMarkdown": "Congratulations! Good job done.\n\nI have some questions. First, what is TTA you mentioned. Second, could you specify how you do pseudo labelling? Finally, what is 8 multi sample dropout? Thanks.",
      "votes": null
    },
    {
      "id": "898231",
      "postDate": "06/23/2020 11:46:06",
      "content": "<p>Thanks <a href=\"/haythemtellili5\">@haythemtellili5</a> </p>",
      "rawMarkdown": "Thanks @haythemtellili5",
      "votes": null
    },
    {
      "id": "898232",
      "postDate": "06/23/2020 11:46:16",
      "content": "<p>Thank you <a href=\"/rohitsingh9990\">@rohitsingh9990</a> </p>",
      "rawMarkdown": "Thank you @rohitsingh9990",
      "votes": null
    },
    {
      "id": "898249",
      "postDate": "06/23/2020 11:56:00",
      "content": "<p>Thanks <a href=\"/yihdarshieh\">@yihdarshieh</a> </p>\n\n<p>TTA is <code>test time augmentation</code> it is more common on image data, where you rotate, zoom or do some other transformations on inference time, to get an average prediction, in my case I used 192 tokens from the original sentence, so in case of very long sentences, I was predicting, in the first 192, last 192, and on the first 96 + last 96, this way I can have a better context of the whole text.\nMulti-sample dropout is a technique to that is supposed to decrease training time and increase generalization of models by combining random masks of the upper model, here is the <a href=\"https://arxiv.org/pdf/1905.09788.pdf\">paper</a> and an <a href=\"https://towardsdatascience.com/multi-sample-dropout-in-keras-ea8b8a9bfd83\">article</a> about it, I can say that this was not very significant in my case, and the <code>8</code> comes because I used 8 samples of it, on the link for the 2nd model you can check my Tensorflow implementation, is very simple.</p>",
      "rawMarkdown": "Thanks @yihdarshieh \n\nTTA is `test time augmentation` it is more common on image data, where you rotate, zoom or do some other transformations on inference time, to get an average prediction, in my case I used 192 tokens from the original sentence, so in case of very long sentences, I was predicting, in the first 192, last 192, and on the first 96 + last 96, this way I can have a better context of the whole text.\nMulti-sample dropout is a technique to that is supposed to decrease training time and increase generalization of models by combining random masks of the upper model, here is the [paper](https://arxiv.org/pdf/1905.09788.pdf) and an [article](https://towardsdatascience.com/multi-sample-dropout-in-keras-ea8b8a9bfd83) about it, I can say that this was not very significant in my case, and the `8` comes because I used 8 samples of it, on the link for the 2nd model you can check my Tensorflow implementation, is very simple.",
      "votes": null
    },
    {
      "id": "900859",
      "postDate": "06/25/2020 05:40:06",
      "content": "<p>Congrats Buddy</p>",
      "rawMarkdown": "Congrats Buddy",
      "votes": null
    },
    {
      "id": "902758",
      "postDate": "06/26/2020 10:44:32",
      "content": "<p>Congrats for the medal and thanks for sharing! such contributions inspires new data scientists like us!</p>",
      "rawMarkdown": "Congrats for the medal and thanks for sharing! such contributions inspires new data scientists like us!",
      "votes": null
    },
    {
      "id": "904246",
      "postDate": "06/27/2020 13:18:09",
      "content": "<p>Thanks <a href=\"/mdselimreza\">@mdselimreza</a> </p>",
      "rawMarkdown": "Thanks @mdselimreza",
      "votes": null
    },
    {
      "id": "904247",
      "postDate": "06/27/2020 13:18:20",
      "content": "<p>Thank you <a href=\"/gauravdahiya\">@gauravdahiya</a> </p>",
      "rawMarkdown": "Thank you @gauravdahiya",
      "votes": null
    },
    {
      "id": "904336",
      "postDate": "06/27/2020 14:46:46",
      "content": "<p>You are most welcome!</p>",
      "rawMarkdown": "You are most welcome!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 897663,
      "author_name": "haythemtellili5",
      "author_url": "",
      "post_date": "06/23/2020 03:01:42",
      "content": "<p>Congratulations Dimitre !</p>",
      "votes": null,
      "replies": [
        {
          "id": 898231,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "06/23/2020 11:46:06",
          "content": "<p>Thanks <a href=\"/haythemtellili5\">@haythemtellili5</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897689,
      "author_name": "rohitsingh9990",
      "author_url": "",
      "post_date": "06/23/2020 03:23:55",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 898232,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "06/23/2020 11:46:16",
          "content": "<p>Thank you <a href=\"/rohitsingh9990\">@rohitsingh9990</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897894,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "06/23/2020 07:06:05",
      "content": "<p>Congratulations! Good job done.</p>\n\n<p>I have some questions. First, what is TTA you mentioned. Second, could you specify how you do pseudo labelling? Finally, what is 8 multi sample dropout? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898249,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "06/23/2020 11:56:00",
          "content": "<p>Thanks <a href=\"/yihdarshieh\">@yihdarshieh</a> </p>\n\n<p>TTA is <code>test time augmentation</code> it is more common on image data, where you rotate, zoom or do some other transformations on inference time, to get an average prediction, in my case I used 192 tokens from the original sentence, so in case of very long sentences, I was predicting, in the first 192, last 192, and on the first 96 + last 96, this way I can have a better context of the whole text.\nMulti-sample dropout is a technique to that is supposed to decrease training time and increase generalization of models by combining random masks of the upper model, here is the <a href=\"https://arxiv.org/pdf/1905.09788.pdf\">paper</a> and an <a href=\"https://towardsdatascience.com/multi-sample-dropout-in-keras-ea8b8a9bfd83\">article</a> about it, I can say that this was not very significant in my case, and the <code>8</code> comes because I used 8 samples of it, on the link for the 2nd model you can check my Tensorflow implementation, is very simple.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 900859,
      "author_name": "mdselimreza",
      "author_url": "",
      "post_date": "06/25/2020 05:40:06",
      "content": "<p>Congrats Buddy</p>",
      "votes": null,
      "replies": [
        {
          "id": 904246,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "06/27/2020 13:18:09",
          "content": "<p>Thanks <a href=\"/mdselimreza\">@mdselimreza</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904336,
          "author_name": "mdselimreza",
          "author_url": "",
          "post_date": "06/27/2020 14:46:46",
          "content": "<p>You are most welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 902758,
      "author_name": "gauravdahiya",
      "author_url": "",
      "post_date": "06/26/2020 10:44:32",
      "content": "<p>Congrats for the medal and thanks for sharing! such contributions inspires new data scientists like us!</p>",
      "votes": null,
      "replies": [
        {
          "id": 904247,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "06/27/2020 13:18:20",
          "content": "<p>Thank you <a href=\"/gauravdahiya\">@gauravdahiya</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "897562": "I did not place so well this time, but still got a medal, so I think it still worth it to share my solution. Congratulation to all winners and new GMs.\nThis was a very interesting competition, being able to work with huge datasets and models was very challenging. Luckily my best submission was the last one that I made 😄. I had many more Ideas to try including using external data but had no time.\nI also have a [Git repository](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification) with my experiments.\n\n## Quick summary\n- Model: 5-Fold `XLM_RoBERTa large` using the last layer `CLS` token.\n  - 5-Fold `XLM_RoBERTa large` concatenated `AVG` and `MAX` last layer pooled with 8 multi-sample dropout.\n- Inference: I have used `TTA` for each sample, I predicted on `head`, `tail`, and a mix of both tokens.\n- Labels: `toxic` cast to `int` then pseudo-labels on the test set.\n- Dataset: For the 1st model I used a 1:2 ratio between toxic and not toxic samples, for the 2nd model use 1:1 ratio, also did some basic cleaning. (details below)\n- Framework: `Tensorflow` with TPU.\n\n## Detailed summary\n\n### Model &amp; training\n- 1st model [link for the 1st fold training](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/99-jigsaw-fold1-xlm-roberta-large-best.ipynb)\n  - 5-Fold XLM_RoBERTa large with the last layer `CLS` token\n  - Sequence length: 192\n  - Batch size: 128\n  - Epochs: 4\n  - Learning rate: 1e-5\n  - Training schedule: Exponential decay with 1 epoch warm-up\n  - Losses: BinaryCrossentropy on the labels cast to `int`\n\nTrained the model 4 epochs on the train set then 1 epoch on the validation set, after that 2 epochs on the test set with pseudo-labels.\n\n- 2nd model [link for the 1st fold training](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Model%20backlog/Train/136-jigsaw-fold1-xlm-roberta-ratio-1-8-sample-drop.ipynb)\n  - 5-Fold XLM_RoBERTa large with the last layer `AVG` and`MAX` pooled concatenated then fed to 8 multi-sample dropout\n  - Sequence length: 192\n  - Batch size: 128\n  - Epochs: 3\n  - Learning rate: 1e-5\n  - Training schedule: Exponential decay with 10% of steps warm-up\n  - Losses: BinaryCrossentropy on the labels cast to `int`\n\nTrained the model 3 epochs on the train set then 2 epochs on the validation set, after that 2 epochs on the test set with pseudo-labels.\n\n### Datasets\n- 1st model [dataset creation](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-2-upper.ipynb) 1:2 toxic to non-toxic samples, 400830 total samples.\n- 2nd model [dataset creation](https://github.com/dimitreOliveira/Jigsaw-Multilingual-Toxic-Comment-Classification/blob/master/Datasets/jigsaw-data-split-roberta-192-ratio-1-clean-polish.ipynb) 1:1 toxic to non-toxic samples, 267220  total samples.\n\nBoth models used a similar dataset, upper case text, just a sample of the negative data, data cleaning was just removal of `numbers`, `#hash-tags`, `@mentios`, `links`, and `multiple white spaces`. Tokenizer was `AutoTokenizer.from_pretrained(`jplu/tf-xlm-roberta-large', lowercase=False)` like many more.\n\n### Inference\nFor inference, I predicted the `head`, `tail`, and a mix of both `(50% head &amp; 50%tail)` sentence tokens.",
    "897663": "Congratulations Dimitre !",
    "897689": "dimitreoliveira congrats!",
    "897894": "Congratulations! Good job done.\n\nI have some questions. First, what is TTA you mentioned. Second, could you specify how you do pseudo labelling? Finally, what is 8 multi sample dropout? Thanks.",
    "898231": "Thanks @haythemtellili5",
    "898232": "Thank you @rohitsingh9990",
    "898249": "Thanks @yihdarshieh \n\nTTA is `test time augmentation` it is more common on image data, where you rotate, zoom or do some other transformations on inference time, to get an average prediction, in my case I used 192 tokens from the original sentence, so in case of very long sentences, I was predicting, in the first 192, last 192, and on the first 96 + last 96, this way I can have a better context of the whole text.\nMulti-sample dropout is a technique to that is supposed to decrease training time and increase generalization of models by combining random masks of the upper model, here is the [paper](https://arxiv.org/pdf/1905.09788.pdf) and an [article](https://towardsdatascience.com/multi-sample-dropout-in-keras-ea8b8a9bfd83) about it, I can say that this was not very significant in my case, and the `8` comes because I used 8 samples of it, on the link for the 2nd model you can check my Tensorflow implementation, is very simple.",
    "900859": "Congrats Buddy",
    "902758": "Congrats for the medal and thanks for sharing! such contributions inspires new data scientists like us!",
    "904246": "Thanks @mdselimreza",
    "904247": "Thank you @gauravdahiya",
    "904336": "You are most welcome!"
  },
  "source": "meta"
}