{
  "id": 146877,
  "title": "Hack with Parallel Corpus",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/146877",
  "author_name": "",
  "post_date": "2020-04-28T19:10:35.366596200Z",
  "votes": 12,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>I think you are tired of my kernels.</p>\n\n<p>But this kernel you should see ;)</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/hack-with-parallel-corpus\">https://www.kaggle.com/shonenkov/hack-with-parallel-corpus</a></p>",
  "messages": [
    {
      "id": "825082",
      "postDate": "04/28/2020 19:10:35",
      "content": "<p>Hi everyone!</p>\n\n<p>I think you are tired of my kernels.</p>\n\n<p>But this kernel you should see ;)</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/hack-with-parallel-corpus\">https://www.kaggle.com/shonenkov/hack-with-parallel-corpus</a></p>",
      "rawMarkdown": "Hi everyone!\n\n\nI think you are tired of my kernels.\n\nBut this kernel you should see ;)\n\nhttps://www.kaggle.com/shonenkov/hack-with-parallel-corpus",
      "votes": null
    },
    {
      "id": "825103",
      "postDate": "04/28/2020 19:30:21",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "830706",
      "postDate": "05/02/2020 19:57:07",
      "content": "<p>Have you had any success with pseudo labelling? I tried it extracting 1m samples. There are very few toxic samples, so I collected 8,000 samples from each class, but it slighly decreases my LB score...</p>",
      "rawMarkdown": "Have you had any success with pseudo labelling? I tried it extracting 1m samples. There are very few toxic samples, so I collected 8,000 samples from each class, but it slighly decreases my LB score...",
      "votes": null
    },
    {
      "id": "834105",
      "postDate": "05/05/2020 09:58:53",
      "content": "<p><a href=\"/rftexas\">@rftexas</a> I suppose you direct used this corpus for training. I did't get success with this approach. I think that token length distribution is very important. In corpus 99% samples have less &lt; 80 token length. </p>\n\n<p>I have used technique with mixing pseudo-labeled text for increasing token length. You can see it in <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">my training pipeline kernel</a>.</p>\n\n<p>You are welcome!</p>",
      "rawMarkdown": "rftexas I suppose you direct used this corpus for training. I did't get success with this approach. I think that token length distribution is very important. In corpus 99% samples have less &lt; 80 token length. \n\nI have used technique with mixing pseudo-labeled text for increasing token length. You can see it in [my training pipeline kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta).\n\nYou are welcome!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 825103,
      "author_name": "albeffe",
      "author_url": "",
      "post_date": "04/28/2020 19:30:21",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 830706,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "05/02/2020 19:57:07",
      "content": "<p>Have you had any success with pseudo labelling? I tried it extracting 1m samples. There are very few toxic samples, so I collected 8,000 samples from each class, but it slighly decreases my LB score...</p>",
      "votes": null,
      "replies": [
        {
          "id": 834105,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "05/05/2020 09:58:53",
          "content": "<p><a href=\"/rftexas\">@rftexas</a> I suppose you direct used this corpus for training. I did't get success with this approach. I think that token length distribution is very important. In corpus 99% samples have less &lt; 80 token length. </p>\n\n<p>I have used technique with mixing pseudo-labeled text for increasing token length. You can see it in <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">my training pipeline kernel</a>.</p>\n\n<p>You are welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "825082": "Hi everyone!\n\n\nI think you are tired of my kernels.\n\nBut this kernel you should see ;)\n\nhttps://www.kaggle.com/shonenkov/hack-with-parallel-corpus",
    "825103": "Thanks for sharing :)",
    "830706": "Have you had any success with pseudo labelling? I tried it extracting 1m samples. There are very few toxic samples, so I collected 8,000 samples from each class, but it slighly decreases my LB score...",
    "834105": "rftexas I suppose you direct used this corpus for training. I did't get success with this approach. I think that token length distribution is very important. In corpus 99% samples have less &lt; 80 token length. \n\nI have used technique with mixing pseudo-labeled text for increasing token length. You can see it in [my training pipeline kernel](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta).\n\nYou are welcome!"
  },
  "source": "meta"
}