{
  "id": 138118,
  "title": "Official External Data Thread",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138118",
  "author_name": "Julia Elliott",
  "post_date": "2020-03-23T21:07:47.559000",
  "votes": 14,
  "comment_count": 83,
  "views": 0,
  "content": "<p>Per the <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/rules\">Competition Rules</a>, freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).</p>\n\n<p>Note that if you wish to use Kaggle's TPU integration, you cannot use privately-held external datasets. Instead, you'll want to upload a public dataset to Kaggle to use and declare those here.</p>\n\n<p>Once someone posts an external dataset to this thread, you do not need to re-post it if you are using the same one.</p>\n\n<p>You only need to declare the original dataset used; you do not need to declare re-labeled or augmented or otherwise processed versions of datasets. Pre-trained models can be declared, as well; however, models resulting from your own original work, that you have trained yourself offline do not need to be shared/declared.</p>",
  "messages": [
    {
      "id": 783984,
      "postDate": "2020-03-23T21:07:47.560Z",
      "content": "<p>Per the <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/rules\">Competition Rules</a>, freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).</p>\n\n<p>Note that if you wish to use Kaggle's TPU integration, you cannot use privately-held external datasets. Instead, you'll want to upload a public dataset to Kaggle to use and declare those here.</p>\n\n<p>Once someone posts an external dataset to this thread, you do not need to re-post it if you are using the same one.</p>\n\n<p>You only need to declare the original dataset used; you do not need to declare re-labeled or augmented or otherwise processed versions of datasets. Pre-trained models can be declared, as well; however, models resulting from your own original work, that you have trained yourself offline do not need to be shared/declared.</p>",
      "rawMarkdown": "Per the [Competition Rules](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/rules), freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).\n\nNote that if you wish to use Kaggle's TPU integration, you cannot use privately-held external datasets. Instead, you'll want to upload a public dataset to Kaggle to use and declare those here.\n\nOnce someone posts an external dataset to this thread, you do not need to re-post it if you are using the same one.\n\nYou only need to declare the original dataset used; you do not need to declare re-labeled or augmented or otherwise processed versions of datasets. Pre-trained models can be declared, as well; however, models resulting from your own original work, that you have trained yourself offline do not need to be shared/declared.",
      "votes": 14
    },
    {
      "id": 786136,
      "postDate": "2020-03-25T16:46:30.540Z",
      "content": "<p>About translations, we checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify.</p>",
      "rawMarkdown": "About translations, we checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify.",
      "votes": 10,
      "replies": [
        {
          "id": 853420,
          "postDate": "2020-05-19T07:11:32.720Z",
          "content": "<p>hi, I have one question. in the rules saying\n<em>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</em>\nIf I post my external data just a few seconds ahead of Entry Deadline, then no team will have time to use this external data. However this external data is still available to all participants and also posted in time(prior to the Entry Deadline), that is to say, not breaking the rules. \nThen what's the meaning of posting external data here? is that fair?</p>",
          "rawMarkdown": "hi, I have one question. in the rules saying\n*C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.*\nIf I post my external data just a few seconds ahead of Entry Deadline, then no team will have time to use this external data. However this external data is still available to all participants and also posted in time(prior to the Entry Deadline), that is to say, not breaking the rules. \nThen what's the meaning of posting external data here? is that fair?",
          "votes": 1
        },
        {
          "id": 854383,
          "postDate": "2020-05-20T02:11:09.520Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"/martingorner\">@martingorner</a> sorry to @ you.</p>",
          "rawMarkdown": "@mgornergoogle @martingorner sorry to @ you.\n\n\n\n"
        },
        {
          "id": 857107,
          "postDate": "2020-05-22T09:54:56.170Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> need help</p>",
          "rawMarkdown": "@juliaelliott need help"
        },
        {
          "id": 857367,
          "postDate": "2020-05-22T14:56:00.467Z",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> Doing what you describe is within the rules and therefore permitted. The data you use should already be publicly available and licensed for that purpose.</p>",
          "rawMarkdown": "@mcggood Doing what you describe is within the rules and therefore permitted. The data you use should already be publicly available and licensed for that purpose."
        },
        {
          "id": 857482,
          "postDate": "2020-05-22T16:36:55.667Z",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> the Entry deadline (along with Team Merger deadline) is one week prior to the Final submission deadline. With one week, there should be enough time to use the external data posted here.</p>",
          "rawMarkdown": "@mcggood the Entry deadline (along with Team Merger deadline) is one week prior to the Final submission deadline. With one week, there should be enough time to use the external data posted here.",
          "votes": -1
        },
        {
          "id": 858614,
          "postDate": "2020-05-23T16:17:31.403Z",
          "content": "<p>I don't think so. one week is not enough even for downloading some big dataset for me:))\nSo could you post your team's dataset as soon as possible? </p>",
          "rawMarkdown": "I don't think so. one week is not enough even for downloading some big dataset for me:))\nSo could you post your team's dataset as soon as possible? "
        },
        {
          "id": 858756,
          "postDate": "2020-05-23T19:22:59.130Z",
          "content": "<p>My comment was not made for personal interest, rather I assumed you mistook the Entry deadline from the Final submission deadline (happened to me as well). \nWho said our team is using (big) external datasets? 😏 </p>",
          "rawMarkdown": "My comment was not made for personal interest, rather I assumed you mistook the Entry deadline from the Final submission deadline (happened to me as well). \nWho said our team is using (big) external datasets? 😏 "
        }
      ]
    },
    {
      "id": 856481,
      "postDate": "2020-05-21T18:31:22.370Z",
      "content": "<p>french :\n<a href=\"https://github.com/marcoguerini/CONAN\">https://github.com/marcoguerini/CONAN</a>\n<a href=\"https://github.com/HKUST-KnowComp/MLMA_hate_speech\">https://github.com/HKUST-KnowComp/MLMA_hate_speech</a></p>\n\n<p>Turkish\n<a href=\"https://coltekin.github.io/offensive-turkish/\">https://coltekin.github.io/offensive-turkish/</a></p>\n\n<p>Spanish \n<a href=\"https://competitions.codalab.org/competitions/19935#learn_the_details\">https://competitions.codalab.org/competitions/19935#learn_the_details</a>\n<a href=\"https://zenodo.org/record/2592149#.Xrhi2GhKiUk\">https://zenodo.org/record/2592149#.Xrhi2GhKiUk</a></p>\n\n<p>Italian</p>\n\n<p><a href=\"http://www.di.unito.it/~tutreeb/haspeede-evalita18/index.html#\">http://www.di.unito.it/~tutreeb/haspeede-evalita18/index.html#</a>\n<a href=\"https://github.com/msang/hate-speech-corpus\">https://github.com/msang/hate-speech-corpus</a></p>\n\n<p>Russian : \n<a href=\"https://www.kaggle.com/blackmoon/russian-language-toxic-comments\">https://www.kaggle.com/blackmoon/russian-language-toxic-comments</a></p>\n\n<p>Portugesh:\n<a href=\"https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\">https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset</a>\n<a href=\"https://github.com/rogersdepelle/OffComBR\">https://github.com/rogersdepelle/OffComBR</a></p>\n\n<p><a href=\"https://huggingface.co/\">https://huggingface.co/</a>\n<a href=\"https://github.com/n-waves/multifit\">https://github.com/n-waves/multifit</a></p>",
      "rawMarkdown": "french :\nhttps://github.com/marcoguerini/CONAN\nhttps://github.com/HKUST-KnowComp/MLMA_hate_speech\n\n\nTurkish\nhttps://coltekin.github.io/offensive-turkish/\n\n\nSpanish \nhttps://competitions.codalab.org/competitions/19935#learn_the_details\nhttps://zenodo.org/record/2592149#.Xrhi2GhKiUk\n\nItalian\n\nhttp://www.di.unito.it/~tutreeb/haspeede-evalita18/index.html#\nhttps://github.com/msang/hate-speech-corpus\n\nRussian : \nhttps://www.kaggle.com/blackmoon/russian-language-toxic-comments\n\n\nPortugesh:\nhttps://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\nhttps://github.com/rogersdepelle/OffComBR\n\nhttps://huggingface.co/\nhttps://github.com/n-waves/multifit",
      "votes": 7,
      "replies": [
        {
          "id": 875803,
          "postDate": "2020-06-06T06:55:38.840Z",
          "content": "<p>I created a Kaggle dataset with all of these files that you mentioned, not sure if this is helpful but here it is:\n<a href=\"https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended\">https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended</a></p>",
          "rawMarkdown": "I created a Kaggle dataset with all of these files that you mentioned, not sure if this is helpful but here it is:\nhttps://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended",
          "votes": 2
        },
        {
          "id": 880386,
          "postDate": "2020-06-10T08:48:56.493Z",
          "content": "<p>Translated Russian comments into English\n<a href=\"https://www.kaggle.com/aybatov/toxic-russian-comments-from-pikabu-and-2ch\">https://www.kaggle.com/aybatov/toxic-russian-comments-from-pikabu-and-2ch</a></p>",
          "rawMarkdown": "Translated Russian comments into English\nhttps://www.kaggle.com/aybatov/toxic-russian-comments-from-pikabu-and-2ch"
        },
        {
          "id": 880414,
          "postDate": "2020-06-10T09:19:38.140Z",
          "content": "<p>Thanks, I will try to preprocess them this week end. I am also on another competition so I did not have the time to do that. I was more into architecture and parameters until now. </p>",
          "rawMarkdown": "Thanks, I will try to preprocess them this week end. I am also on another competition so I did not have the time to do that. I was more into architecture and parameters until now. ",
          "votes": 1
        },
        {
          "id": 886340,
          "postDate": "2020-06-15T00:20:16.273Z",
          "content": "<p>sorry for the delay, the data are available in one .csv</p>\n\n<p><a href=\"https://www.kaggle.com/ludovick/hatespeechmulti\">https://www.kaggle.com/ludovick/hatespeechmulti</a></p>",
          "rawMarkdown": "sorry for the delay, the data are available in one .csv\n\nhttps://www.kaggle.com/ludovick/hatespeechmulti",
          "votes": 1
        },
        {
          "id": 888712,
          "postDate": "2020-06-16T14:17:50.613Z",
          "content": "<p>Thank you. Much appreciated 👍 👍 💪 </p>",
          "rawMarkdown": "Thank you. Much appreciated 👍 👍 💪 "
        }
      ]
    },
    {
      "id": 784578,
      "postDate": "2020-03-24T10:41:11.930Z",
      "content": "<p><a href=\"https://github.com/ssut/py-googletrans\">https://github.com/ssut/py-googletrans</a>\n<a href=\"https://github.com/terryyin/translate-python\">https://github.com/terryyin/translate-python</a></p>\n\n<p><a href=\"https://mymemory.translated.net/\">https://mymemory.translated.net/</a>\n<a href=\"https://cloud.google.com/translate/\">https://cloud.google.com/translate/</a>\n<a href=\"https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\">https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/</a>\n<a href=\"https://tech.yandex.com/translate/\">https://tech.yandex.com/translate/</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/notebooks\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/notebooks</a> any kernel + output here (just naming a few next) + of course all associated datasets\n<a href=\"https://www.kaggle.com/hamditarek/ensemble\">https://www.kaggle.com/hamditarek/ensemble</a>\n<a href=\"https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta</a>\n<a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a>\n<a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a>\n<a href=\"https://www.kaggle.com/yeayates21/xlm-roberta-augmentation-ssl-0-9417-pub-lb\">https://www.kaggle.com/yeayates21/xlm-roberta-augmentation-ssl-0-9417-pub-lb</a>\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>\n<a href=\"https://www.kaggle.com/mobassir/understanding-cross-lingual-models\">https://www.kaggle.com/mobassir/understanding-cross-lingual-models</a>\n<a href=\"https://drive.google.com/drive/folders/1hbcSRfvtTTlERs7remsRST2amIWAFVry\">https://drive.google.com/drive/folders/1hbcSRfvtTTlERs7remsRST2amIWAFVry</a></p>\n\n<p><a href=\"https://huggingface.co/models\">https://huggingface.co/models</a> any model here</p>",
      "rawMarkdown": "https://github.com/ssut/py-googletrans\nhttps://github.com/terryyin/translate-python\n\nhttps://mymemory.translated.net/\nhttps://cloud.google.com/translate/\nhttps://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\nhttps://tech.yandex.com/translate/\n\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/notebooks any kernel + output here (just naming a few next) + of course all associated datasets\nhttps://www.kaggle.com/hamditarek/ensemble\nhttps://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\nhttps://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\nhttps://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\nhttps://www.kaggle.com/yeayates21/xlm-roberta-augmentation-ssl-0-9417-pub-lb\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\nhttps://www.kaggle.com/mobassir/understanding-cross-lingual-models\nhttps://drive.google.com/drive/folders/1hbcSRfvtTTlERs7remsRST2amIWAFVry\n\nhttps://huggingface.co/models any model here",
      "votes": 5
    },
    {
      "id": 887823,
      "postDate": "2020-06-15T23:01:25.033Z",
      "content": "<p>Notebook from previous competition\n<a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part1-eda\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part1-eda</a>\n<a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part2-usage\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part2-usage</a>\n<a href=\"https://www.kaggle.com/haqishen/jigsaw-predict\">https://www.kaggle.com/haqishen/jigsaw-predict</a></p>\n\n<p>Word embedding\n<a href=\"https://fasttext.cc/docs/en/english-vectors.html\">https://fasttext.cc/docs/en/english-vectors.html</a>\n<a href=\"https://fasttext.cc/docs/en/crawl-vectors.html\">https://fasttext.cc/docs/en/crawl-vectors.html</a></p>",
      "rawMarkdown": "Notebook from previous competition\nhttps://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part1-eda\nhttps://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part2-usage\nhttps://www.kaggle.com/haqishen/jigsaw-predict\n\nWord embedding\nhttps://fasttext.cc/docs/en/english-vectors.html\nhttps://fasttext.cc/docs/en/crawl-vectors.html\n",
      "votes": 1
    },
    {
      "id": 886328,
      "postDate": "2020-06-14T23:49:53.723Z",
      "content": "<p>may or may not useful : \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">https://www.kaggle.com/c/quora-insincere-questions-classification</a></p>\n\n<p>various langs : <a href=\"https://github.com/valeriobasile/hurtlex/tree/master/lexica\">https://github.com/valeriobasile/hurtlex/tree/master/lexica</a></p>\n\n<p>Portugese : <a href=\"https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\">https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset</a>\nSpanish : <a href=\"https://github.com/msang/hateval\">https://github.com/msang/hateval</a> , <a href=\"https://zenodo.org/record/2592149\">https://zenodo.org/record/2592149</a>\nFrench : <a href=\"https://github.com/HKUST-KnowComp/MLMA_hate_speech\">https://github.com/HKUST-KnowComp/MLMA_hate_speech</a>\nRussian : <a href=\"https://www.kaggle.com/blackmoon/russian-language-toxic-comments\">https://www.kaggle.com/blackmoon/russian-language-toxic-comments</a>\nTurkish : <a href=\"https://www.kaggle.com/ahmetax/hury-dataset\">https://www.kaggle.com/ahmetax/hury-dataset</a>\nPortugese : <a href=\"https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\">https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset</a> , <a href=\"https://github.com/rogersdepelle/OffComBR\">https://github.com/rogersdepelle/OffComBR</a></p>",
      "rawMarkdown": "may or may not useful : \nhttps://www.kaggle.com/c/quora-insincere-questions-classification\n\nvarious langs : https://github.com/valeriobasile/hurtlex/tree/master/lexica\n\nPortugese : https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\nSpanish : https://github.com/msang/hateval , https://zenodo.org/record/2592149\nFrench : https://github.com/HKUST-KnowComp/MLMA_hate_speech\nRussian : https://www.kaggle.com/blackmoon/russian-language-toxic-comments\nTurkish : https://www.kaggle.com/ahmetax/hury-dataset\nPortugese : https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset , https://github.com/rogersdepelle/OffComBR",
      "votes": 1
    },
    {
      "id": 886262,
      "postDate": "2020-06-14T20:56:03.183Z",
      "content": "<p><a href=\"https://github.com/leondz/hatespeechdata\">https://github.com/leondz/hatespeechdata</a></p>",
      "rawMarkdown": "https://github.com/leondz/hatespeechdata",
      "votes": 1
    },
    {
      "id": 885814,
      "postDate": "2020-06-14T13:52:44.927Z",
      "content": "<p>I haven't used this yet, but may give it a shot. They have pretrained models.</p>\n\n<p>LASER Language-Agnostic SEntence Representations\n<a href=\"https://github.com/facebookresearch/LASER\">https://github.com/facebookresearch/LASER</a></p>",
      "rawMarkdown": "I haven't used this yet, but may give it a shot. They have pretrained models.\n\nLASER Language-Agnostic SEntence Representations\nhttps://github.com/facebookresearch/LASER",
      "votes": 1
    },
    {
      "id": 828148,
      "postDate": "2020-04-30T19:37:27.907Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": 1
    },
    {
      "id": 887834,
      "postDate": "2020-06-15T23:10:57.367Z",
      "content": "<p><a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api</a>\n<a href=\"https://www.kaggle.com/rafiko1/translated-train-bias-all-langs\">https://www.kaggle.com/rafiko1/translated-train-bias-all-langs</a></p>\n\n<p><a href=\"https://github.com/allenai/allennlp\">https://github.com/allenai/allennlp</a>\n<a href=\"https://github.com/pytorch/fairseq\">https://github.com/pytorch/fairseq</a>\n<a href=\"https://github.com/flairNLP/flair\">https://github.com/flairNLP/flair</a>\n<a href=\"https://github.com/facebookresearch/LASER\">https://github.com/facebookresearch/LASER</a></p>",
      "rawMarkdown": "https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/rafiko1/translated-train-bias-all-langs\n\nhttps://github.com/allenai/allennlp\nhttps://github.com/pytorch/fairseq\nhttps://github.com/flairNLP/flair\nhttps://github.com/facebookresearch/LASER",
      "votes": 2
    },
    {
      "id": 823388,
      "postDate": "2020-04-27T15:50:15.050Z",
      "content": "<p><a href=\"http://opus.nlpl.eu/OpenSubtitles-v2018.php\">http://opus.nlpl.eu/OpenSubtitles-v2018.php</a> </p>",
      "rawMarkdown": "http://opus.nlpl.eu/OpenSubtitles-v2018.php ",
      "votes": 2
    },
    {
      "id": 792277,
      "postDate": "2020-03-31T03:56:54.020Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038",
      "votes": 2
    },
    {
      "id": 784227,
      "postDate": "2020-03-24T02:51:04.673Z",
      "content": "<p>Tbh I'm surprised that external data is allowed here. </p>\n\n<p>Please correct me if I'm wrong (I haven't really done any EDA on the data yet) but, do you think that stuff like abusing Google's translation API may ruin the purposes of this competition? The challenge is to make a model that generalizes to many languages using only English as a starting point, now if you can just eyeball the test set to see which languages are there and translate your train set accordingly, what makes it different?</p>",
      "rawMarkdown": "Tbh I'm surprised that external data is allowed here. \n\nPlease correct me if I'm wrong (I haven't really done any EDA on the data yet) but, do you think that stuff like abusing Google's translation API may ruin the purposes of this competition? The challenge is to make a model that generalizes to many languages using only English as a starting point, now if you can just eyeball the test set to see which languages are there and translate your train set accordingly, what makes it different?",
      "votes": 2,
      "replies": [
        {
          "id": 784581,
          "postDate": "2020-03-24T10:42:55.720Z",
          "content": "<p>If there is Internet access on in scoring kernels, you can just translate the test data to English. Hope for clarification here.</p>",
          "rawMarkdown": "If there is Internet access on in scoring kernels, you can just translate the test data to English. Hope for clarification here.",
          "votes": 2
        },
        {
          "id": 785529,
          "postDate": "2020-03-25T06:15:02.313Z",
          "content": "<p>We don't even need internet access in the inference kernel to do that. as we already have access to entire test set which we can translate beforehand.</p>",
          "rawMarkdown": "We don't even need internet access in the inference kernel to do that. as we already have access to entire test set which we can translate beforehand.",
          "votes": 2
        },
        {
          "id": 785645,
          "postDate": "2020-03-25T08:22:38.020Z",
          "content": "<p>Really? Oh...</p>",
          "rawMarkdown": "Really? Oh..."
        },
        {
          "id": 785653,
          "postDate": "2020-03-25T08:32:12.823Z",
          "content": "<p>All my will to commit just vanished</p>",
          "rawMarkdown": "All my will to commit just vanished"
        },
        {
          "id": 785663,
          "postDate": "2020-03-25T08:42:52.347Z",
          "content": "<p>I have a bad feeling about this comp;\n- First Many possibilites to cheat i can think of;\n- Secondly training with 5k examples is so close to .80 AUC's benchmark which was trained with whole data [wrt TF TPU kernel]?\n- Plus PyTorch XLA is not stable yet\n- Plus the comp is purely overfit the data to win type as well (as it seems for now)</p>",
          "rawMarkdown": "I have a bad feeling about this comp;\n- First Many possibilites to cheat i can think of;\n- Secondly training with 5k examples is so close to .80 AUC's benchmark which was trained with whole data [wrt TF TPU kernel]?\n- Plus PyTorch XLA is not stable yet\n- Plus the comp is purely overfit the data to win type as well (as it seems for now)",
          "votes": 2
        },
        {
          "id": 785807,
          "postDate": "2020-03-25T12:03:01.827Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 785811,
          "postDate": "2020-03-25T12:10:04.233Z",
          "content": "<p>Hand-labeling is actually not allowed based on the rules:</p>\n\n<p>&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n\n<p>But I don't see any rule forbidding automatic translation. Actually this had been done in past competitions.</p>",
          "rawMarkdown": "Hand-labeling is actually not allowed based on the rules:\n\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nBut I don't see any rule forbidding automatic translation. Actually this had been done in past competitions."
        },
        {
          "id": 786134,
          "postDate": "2020-03-25T16:43:50.387Z",
          "content": "<p>&gt; Plus PyTorch XLA is not stable yet\nYes, that has been clearly announced multiple times. PyTorch TPU support on Kaggle is experimental at this point.</p>\n\n<p>Tensorflow is the way to go for TPUs for now: <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started</a></p>",
          "rawMarkdown": "&gt; Plus PyTorch XLA is not stable yet\nYes, that has been clearly announced multiple times. PyTorch TPU support on Kaggle is experimental at this point.\n\nTensorflow is the way to go for TPUs for now: https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started",
          "votes": 1
        },
        {
          "id": 786505,
          "postDate": "2020-03-26T00:07:56.613Z",
          "content": "<blockquote>\n  <p>One could also label the test set by hand. Nobody reads that many languages, but a team could. How do you rule that out?</p>\n</blockquote>\n\n<p>'Submissions to this competition must be made by an output from a Kaggle Notebook/Script.' - From the code requirements</p>\n\n<p>I believe this will weed out submissions that were hand labeled, as you could easily run their notebook to verify.</p>",
          "rawMarkdown": "&gt; One could also label the test set by hand. Nobody reads that many languages, but a team could. How do you rule that out?\n\n'Submissions to this competition must be made by an output from a Kaggle Notebook/Script.' - From the code requirements\n\nI believe this will weed out submissions that were hand labeled, as you could easily run their notebook to verify."
        },
        {
          "id": 786633,
          "postDate": "2020-03-26T04:21:26.677Z",
          "content": "<p>There are smart ways to hide it;  Like what about this, I trained my model on the test set as well (after hand-labelling) -&gt; Cached the weights -&gt; External Dataset Import -&gt; Use it freely :) ;\nThoughts?</p>",
          "rawMarkdown": "There are smart ways to hide it;  Like what about this, I trained my model on the test set as well (after hand-labelling) -&gt; Cached the weights -&gt; External Dataset Import -&gt; Use it freely :) ;\nThoughts?\n",
          "votes": 2
        },
        {
          "id": 811121,
          "postDate": "2020-04-17T16:00:14.820Z",
          "content": "<p>But again problem with translation is that, many times the text gets de toxified.</p>",
          "rawMarkdown": "But again problem with translation is that, many times the text gets de toxified."
        },
        {
          "id": 857679,
          "postDate": "2020-05-22T20:17:52.540Z",
          "content": "<p>does that mean we can use the test set if not hand labeled ?</p>",
          "rawMarkdown": "does that mean we can use the test set if not hand labeled ?"
        },
        {
          "id": 888030,
          "postDate": "2020-06-16T04:33:48.957Z",
          "content": "<p>The test set <strong>should not be hand-labeled</strong>. <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138118#786136\">As stated previously by Martin</a>, datasets translated into the languages that are represented in the test set is legitimate, but if you are using the test set directly in any way that includes manual labels, then that would be a prohibited use of the test set. If hand-labeling of the test set is discovered in the review of a winning solution, it is subject to disqualification. </p>",
          "rawMarkdown": "The test set **should not be hand-labeled**. [As stated previously by Martin](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138118#786136), datasets translated into the languages that are represented in the test set is legitimate, but if you are using the test set directly in any way that includes manual labels, then that would be a prohibited use of the test set. If hand-labeling of the test set is discovered in the review of a winning solution, it is subject to disqualification. ",
          "votes": 1
        },
        {
          "id": 888684,
          "postDate": "2020-06-16T14:01:12.910Z",
          "content": "<p>@Julia\nI want to make it clear whether the pseudo label legal ? e.g.  I use my model to predict the test data and get the toxic target.  Then I use the predicted toxic target as the label of test and add the test data into my training data for my model training.  Does it allowed ? </p>",
          "rawMarkdown": "@Julia\nI want to make it clear whether the pseudo label legal ? e.g.  I use my model to predict the test data and get the toxic target.  Then I use the predicted toxic target as the label of test and add the test data into my training data for my model training.  Does it allowed ? "
        },
        {
          "id": 891237,
          "postDate": "2020-06-18T03:26:25.467Z",
          "content": "<p>I think so. Pseudo labelling is used in every NLP competition.</p>",
          "rawMarkdown": "I think so. Pseudo labelling is used in every NLP competition."
        },
        {
          "id": 897396,
          "postDate": "2020-06-22T20:37:06.950Z",
          "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> Pseudo-labeling is permitted. Manual hand-labeling of the test set is not.</p>",
          "rawMarkdown": "@qinhui1999 Pseudo-labeling is permitted. Manual hand-labeling of the test set is not."
        }
      ]
    },
    {
      "id": 890027,
      "postDate": "2020-06-17T09:08:10.380Z",
      "content": "<p>all dataset has been declaired by others</p>",
      "rawMarkdown": "all dataset has been declaired by others",
      "votes": -1
    },
    {
      "id": 898391,
      "postDate": "2020-06-23T13:53:50.537Z",
      "content": "<p><strong>Added myself:</strong>\n<a href=\"https://www.kaggle.com/yeayates21/jigsawmultilingualrobertaavgblender\">https://www.kaggle.com/yeayates21/jigsawmultilingualrobertaavgblender</a>\n<a href=\"https://www.kaggle.com/yeayates21/jigsawtpuxlmrobertacopypickledata\">https://www.kaggle.com/yeayates21/jigsawtpuxlmrobertacopypickledata</a></p>\n\n<p><strong>Added by others:</strong>\n<a href=\"https://www.kaggle.com/hamditarek/blendings\">https://www.kaggle.com/hamditarek/blendings</a>\n<a href=\"https://www.kaggle.com/hamditarek/ensemble-version-96\">https://www.kaggle.com/hamditarek/ensemble-version-96</a>\n<a href=\"https://www.kaggle.com/hamditarek/tfidf\">https://www.kaggle.com/hamditarek/tfidf</a>\n<a href=\"https://www.kaggle.com/hamditarek/009248\">https://www.kaggle.com/hamditarek/009248</a>\n<a href=\"https://www.kaggle.com/hamditarek/009259\">https://www.kaggle.com/hamditarek/009259</a>\n<a href=\"https://www.kaggle.com/hamditarek/009354\">https://www.kaggle.com/hamditarek/009354</a>\n<a href=\"https://www.kaggle.com/hamditarek/009383\">https://www.kaggle.com/hamditarek/009383</a>\n<a href=\"https://www.kaggle.com/hamditarek/009406\">https://www.kaggle.com/hamditarek/009406</a>\n<a href=\"https://www.kaggle.com/hamditarek/009423\">https://www.kaggle.com/hamditarek/009423</a>\n<a href=\"https://www.kaggle.com/hamditarek/abhishek\">https://www.kaggle.com/hamditarek/abhishek</a>\n<a href=\"https://www.kaggle.com/hamditarek/blending\">https://www.kaggle.com/hamditarek/blending</a>\n<a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling</a></p>",
      "rawMarkdown": "**Added myself:**\nhttps://www.kaggle.com/yeayates21/jigsawmultilingualrobertaavgblender\nhttps://www.kaggle.com/yeayates21/jigsawtpuxlmrobertacopypickledata\n\n**Added by others:**\nhttps://www.kaggle.com/hamditarek/blendings\nhttps://www.kaggle.com/hamditarek/ensemble-version-96\nhttps://www.kaggle.com/hamditarek/tfidf\nhttps://www.kaggle.com/hamditarek/009248\nhttps://www.kaggle.com/hamditarek/009259\nhttps://www.kaggle.com/hamditarek/009354\nhttps://www.kaggle.com/hamditarek/009383\nhttps://www.kaggle.com/hamditarek/009406\nhttps://www.kaggle.com/hamditarek/009423\nhttps://www.kaggle.com/hamditarek/abhishek\nhttps://www.kaggle.com/hamditarek/blending\nhttps://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling"
    },
    {
      "id": 895766,
      "postDate": "2020-06-21T15:36:17.180Z",
      "content": "<p><a href=\"https://www.kaggle.com/ibtesama/hatespeech\">https://www.kaggle.com/ibtesama/hatespeech</a></p>",
      "rawMarkdown": "https://www.kaggle.com/ibtesama/hatespeech"
    },
    {
      "id": 894826,
      "postDate": "2020-06-20T19:31:10.967Z",
      "content": "<p>I guess that any dataset that is marked as \"research-only\", \"not for commercial use\", etc. can't be used in any way to train the model?</p>",
      "rawMarkdown": "I guess that any dataset that is marked as \"research-only\", \"not for commercial use\", etc. can't be used in any way to train the model?",
      "replies": [
        {
          "id": 897394,
          "postDate": "2020-06-22T20:34:41.930Z",
          "content": "<p>In this competition, external data licensed for non-commercial, research, and/or academic-use is permitted.</p>",
          "rawMarkdown": "In this competition, external data licensed for non-commercial, research, and/or academic-use is permitted."
        },
        {
          "id": 897405,
          "postDate": "2020-06-22T20:43:38.480Z",
          "content": "<p>Thank you! I was a bit late asking this, so have used only fully open data.</p>",
          "rawMarkdown": "Thank you! I was a bit late asking this, so have used only fully open data."
        }
      ]
    },
    {
      "id": 893839,
      "postDate": "2020-06-20T02:15:33.207Z",
      "content": "<p><a href=\"https://www.kaggle.com/ma7555/jigsaw-train-translated\">https://www.kaggle.com/ma7555/jigsaw-train-translated</a>\n<a href=\"https://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api\">https://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api</a></p>",
      "rawMarkdown": "https://www.kaggle.com/ma7555/jigsaw-train-translated\nhttps://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api"
    },
    {
      "id": 893781,
      "postDate": "2020-06-19T22:09:34.867Z",
      "content": "<p>I addition to the list I may use the following:-\n<a href=\"https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\">https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated</a>\n<a href=\"https://www.kaggle.com/bamps53/val-en-df\">https://www.kaggle.com/bamps53/val-en-df</a>\n<a href=\"https://www.kaggle.com/bamps53/test-en-df\">https://www.kaggle.com/bamps53/test-en-df</a></p>",
      "rawMarkdown": "I addition to the list I may use the following:-\nhttps://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\nhttps://www.kaggle.com/bamps53/val-en-df\nhttps://www.kaggle.com/bamps53/test-en-df"
    },
    {
      "id": 891384,
      "postDate": "2020-06-18T06:31:49.893Z",
      "content": "<p><a href=\"https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended\">https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended</a>\n<a href=\"https://huggingface.co/models\">https://huggingface.co/models</a></p>",
      "rawMarkdown": "[https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended](https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended)\n[https://huggingface.co/models](https://huggingface.co/models)"
    },
    {
      "id": 887850,
      "postDate": "2020-06-15T23:36:51.243Z",
      "content": "<p>May or may not use:</p>\n\n<p><a href=\"https://github.com/uliontse/translators\">https://github.com/uliontse/translators</a>\n<a href=\"https://github.com/sloria/textblob\">https://github.com/sloria/textblob</a>\n<a href=\"https://github.com/littlecodersh/translation\">https://github.com/littlecodersh/translation</a>\n<a href=\"https://github.com/soimort/translate-shell\">https://github.com/soimort/translate-shell</a></p>\n\n<p><a href=\"https://github.com/Rayraegah/warui\">https://github.com/Rayraegah/warui</a>\n<a href=\"https://github.com/CRomano31415/SpanishProfanity\">https://github.com/CRomano31415/SpanishProfanity</a>\n<a href=\"https://github.com/mmcclarty/lyrics_melange\">https://github.com/mmcclarty/lyrics_melange</a>\n<a href=\"https://github.com/TheSoma300/random-italian-curse-generator\">https://github.com/TheSoma300/random-italian-curse-generator</a>\n<a href=\"https://github.com/ChaseFlorell/jQuery.ProfanityFilter\">https://github.com/ChaseFlorell/jQuery.ProfanityFilter</a>\n<a href=\"https://github.com/Zeindelf/badwords\">https://github.com/Zeindelf/badwords</a>\n<a href=\"https://github.com/BotanUA/russian_swear_words\">https://github.com/BotanUA/russian_swear_words</a>\n<a href=\"https://github.com/voyula/turkish-bad-words\">https://github.com/voyula/turkish-bad-words</a></p>",
      "rawMarkdown": "May or may not use:\n\nhttps://github.com/uliontse/translators\nhttps://github.com/sloria/textblob\nhttps://github.com/littlecodersh/translation\nhttps://github.com/soimort/translate-shell\n\nhttps://github.com/Rayraegah/warui\nhttps://github.com/CRomano31415/SpanishProfanity\nhttps://github.com/mmcclarty/lyrics_melange\nhttps://github.com/TheSoma300/random-italian-curse-generator\nhttps://github.com/ChaseFlorell/jQuery.ProfanityFilter\nhttps://github.com/Zeindelf/badwords\nhttps://github.com/BotanUA/russian_swear_words\nhttps://github.com/voyula/turkish-bad-words"
    },
    {
      "id": 887631,
      "postDate": "2020-06-15T19:23:58.120Z",
      "content": "<p><a href=\"https://www.kaggle.com/ludovick/jigsawtanslatedgoogle\">Google translated data</a>\noutputs of public notebooks in this competition\npre-trained models in <a href=\"https://huggingface.co/models\">huggingface</a></p>",
      "rawMarkdown": "[Google translated data](https://www.kaggle.com/ludovick/jigsawtanslatedgoogle)\noutputs of public notebooks in this competition\npre-trained models in [huggingface](https://huggingface.co/models)"
    },
    {
      "id": 887586,
      "postDate": "2020-06-15T18:49:44.310Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">https://www.kaggle.com/c/quora-insincere-questions-classification</a>\n<a href=\"https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\">https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large</a>\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api</a>\n<a href=\"https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\">https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data</a>\n<a href=\"https://github.com/google-research/xtreme\">https://github.com/google-research/xtreme</a>\n<a href=\"https://github.com/google-research-datasets/paws/tree/master/pawsx\">https://github.com/google-research-datasets/paws/tree/master/pawsx</a>\n<a href=\"https://github.com/facebookresearch/XNLI\">https://github.com/facebookresearch/XNLI</a>\n<a href=\"https://mymemory.translated.net/\">https://mymemory.translated.net/</a>\n<a href=\"https://cloud.google.com/translate/\">https://cloud.google.com/translate/</a>\n<a href=\"https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\">https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/</a>\n<a href=\"https://commoncrawl.org\">https://commoncrawl.org</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/quora-insincere-questions-classification\nhttps://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\nhttps://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\nhttps://github.com/google-research/xtreme\nhttps://github.com/google-research-datasets/paws/tree/master/pawsx\nhttps://github.com/facebookresearch/XNLI\nhttps://mymemory.translated.net/\nhttps://cloud.google.com/translate/\nhttps://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\nhttps://commoncrawl.org"
    },
    {
      "id": 887387,
      "postDate": "2020-06-15T16:37:34.527Z",
      "content": "<p><a href=\"https://huggingface.co/novinsh/xlm-roberta-large-toxicomments-12k\">https://huggingface.co/novinsh/xlm-roberta-large-toxicomments-12k</a></p>",
      "rawMarkdown": "https://huggingface.co/novinsh/xlm-roberta-large-toxicomments-12k"
    },
    {
      "id": 887362,
      "postDate": "2020-06-15T16:21:19.770Z",
      "content": "<p>Thanks to the creators of the datasets below (might be used for the final submission):\n<a href=\"https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\">https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large</a>\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api</a>\n<a href=\"https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\">https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data</a>\n<a href=\"https://www.kaggle.com/shonenkov/jigsaw-public-baseline-results\">https://www.kaggle.com/shonenkov/jigsaw-public-baseline-results</a>\n<a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling</a>\n<a href=\"https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\">https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated</a></p>",
      "rawMarkdown": "Thanks to the creators of the datasets below (might be used for the final submission):\nhttps://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\nhttps://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\nhttps://www.kaggle.com/shonenkov/jigsaw-public-baseline-results\nhttps://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\nhttps://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated"
    },
    {
      "id": 887238,
      "postDate": "2020-06-15T15:05:22.700Z",
      "content": "<p>Data for previous competition  <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification\">https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification</a></p>",
      "rawMarkdown": "Data for previous competition  https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification"
    },
    {
      "id": 887093,
      "postDate": "2020-06-15T13:26:09.237Z",
      "content": "<p>ELECTRA (monolingual) model: <a href=\"https://huggingface.co/google/electra-large-discriminator\">https://huggingface.co/google/electra-large-discriminator</a>\nWe also extracted the output of a public ensembling kernel <a href=\"https://www.kaggle.com/hamditarek/ensemble?scriptVersionId=35925815\">https://www.kaggle.com/hamditarek/ensemble?scriptVersionId=35925815</a> to a separate dataset: <a href=\"https://www.kaggle.com/buzhinsky/hamditarekensemblejun12\">https://www.kaggle.com/buzhinsky/hamditarekensemblejun12</a></p>",
      "rawMarkdown": "ELECTRA (monolingual) model: https://huggingface.co/google/electra-large-discriminator\nWe also extracted the output of a public ensembling kernel https://www.kaggle.com/hamditarek/ensemble?scriptVersionId=35925815 to a separate dataset: https://www.kaggle.com/buzhinsky/hamditarekensemblejun12"
    },
    {
      "id": 886916,
      "postDate": "2020-06-15T11:08:39.780Z",
      "content": "<p>4chan data repository: <a href=\"https://zenodo.org/record/3603292#.XudWL0UzaUk\">https://zenodo.org/record/3603292#.XudWL0UzaUk</a>\nReddit comment history since 2005: <a href=\"https://console.cloud.google.com/bigquery?project=fh-bigquery&amp;redirect_from_classic=true&amp;p=fh-bigquery&amp;d=reddit_comments&amp;page=dataset\">https://console.cloud.google.com/bigquery?project=fh-bigquery&amp;redirect_from_classic=true&amp;p=fh-bigquery&amp;d=reddit_comments&amp;page=dataset</a></p>",
      "rawMarkdown": "4chan data repository: https://zenodo.org/record/3603292#.XudWL0UzaUk\nReddit comment history since 2005: https://console.cloud.google.com/bigquery?project=fh-bigquery&amp;redirect_from_classic=true&amp;p=fh-bigquery&amp;d=reddit_comments&amp;page=dataset"
    },
    {
      "id": 886244,
      "postDate": "2020-06-14T20:28:18.070Z",
      "content": "<p>some additional Wikipedia talk comments here: <br>\n<a href=\"https://figshare.com/projects/Wikipedia_Talk/16731\">https://figshare.com/projects/Wikipedia_Talk/16731</a></p>\n\n<p>multi-lingual embedding: \n<a href=\"https://github.com/facebookresearch/MUSE\">https://github.com/facebookresearch/MUSE</a>\n<a href=\"https://fasttext.cc/docs/en/aligned-vectors.html\">https://fasttext.cc/docs/en/aligned-vectors.html</a>\n<a href=\"https://nlp.stanford.edu/projects/glove/\">https://nlp.stanford.edu/projects/glove/</a></p>",
      "rawMarkdown": "some additional Wikipedia talk comments here:  \nhttps://figshare.com/projects/Wikipedia_Talk/16731\n\nmulti-lingual embedding: \nhttps://github.com/facebookresearch/MUSE\nhttps://fasttext.cc/docs/en/aligned-vectors.html\nhttps://nlp.stanford.edu/projects/glove/"
    },
    {
      "id": 884674,
      "postDate": "2020-06-13T14:41:28.263Z",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Do the submission results from the public notebook count as external data? For example this one: <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>  Are public submission results like this one allowed to be used in final submissions?</p>",
      "rawMarkdown": "@juliaelliott Do the submission results from the public notebook count as external data? For example this one: https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta  Are public submission results like this one allowed to be used in final submissions?",
      "replies": [
        {
          "id": 884676,
          "postDate": "2020-06-13T14:42:49.750Z",
          "content": "<p>You can always submit public notebooks.</p>",
          "rawMarkdown": "You can always submit public notebooks."
        },
        {
          "id": 884699,
          "postDate": "2020-06-13T15:01:30.940Z",
          "content": "<p>Some results in public notebook cannot be replicated like the example I list.😂 </p>",
          "rawMarkdown": "Some results in public notebook cannot be replicated like the example I list.😂 "
        },
        {
          "id": 884905,
          "postDate": "2020-06-13T17:58:37.393Z",
          "content": "<p>It is the same with the runs you are doing yourself, no difference between public kernels and your own. There is randomness involved in running NNs.</p>\n\n<p>Submitting external submission.csv files was explicitly allowed in this competition.</p>",
          "rawMarkdown": "It is the same with the runs you are doing yourself, no difference between public kernels and your own. There is randomness involved in running NNs.\n\nSubmitting external submission.csv files was explicitly allowed in this competition."
        },
        {
          "id": 888016,
          "postDate": "2020-06-16T04:23:52.220Z",
          "content": "<p><a href=\"/nzholmes\">@nzholmes</a> - all public notebooks are automatically licensed under Apache 2.0 license. As a result, its contents and output are all permitted to be used by any user. </p>",
          "rawMarkdown": "@nzholmes - all public notebooks are automatically licensed under Apache 2.0 license. As a result, its contents and output are all permitted to be used by any user. ",
          "votes": 1
        },
        {
          "id": 891539,
          "postDate": "2020-06-18T08:44:54.230Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> - I am a bit concerned about people using inference only notebooks that do not include the training scripts. Is it allowed to use the output of these kernels?</p>",
          "rawMarkdown": "@juliaelliott - I am a bit concerned about people using inference only notebooks that do not include the training scripts. Is it allowed to use the output of these kernels?",
          "votes": 1
        },
        {
          "id": 897512,
          "postDate": "2020-06-22T23:29:17.390Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I'm just seeing your question, so apologies for the delayed response. I may not be clearly comprehending your concern. Based on what I do understand, the outputs have been made available on Kaggle to use by the notebook publisher, as a result of making the notebook public. That said, competition winners would still be expected to have license and rights to open-source the code used to achieve their result.</p>",
          "rawMarkdown": "@philippsinger I'm just seeing your question, so apologies for the delayed response. I may not be clearly comprehending your concern. Based on what I do understand, the outputs have been made available on Kaggle to use by the notebook publisher, as a result of making the notebook public. That said, competition winners would still be expected to have license and rights to open-source the code used to achieve their result."
        },
        {
          "id": 897781,
          "postDate": "2020-06-23T05:03:40.813Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I think the issue is that you do not know how the results were generated in these kernels.</p>",
          "rawMarkdown": "@juliaelliott I think the issue is that you do not know how the results were generated in these kernels.",
          "votes": 1
        }
      ]
    },
    {
      "id": 884475,
      "postDate": "2020-06-13T12:01:56.467Z",
      "content": "<p><a href=\"https://dumps.wikimedia.org/\">https://dumps.wikimedia.org/</a>\n<a href=\"https://www.kaggle.com/ludovick/jigsawtanslatedgoogle\">https://www.kaggle.com/ludovick/jigsawtanslatedgoogle</a>\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api</a>\n<a href=\"https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\">https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated</a>\n<a href=\"https://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api\">https://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api</a>\nall outputs from public kernels</p>",
      "rawMarkdown": "https://dumps.wikimedia.org/\nhttps://www.kaggle.com/ludovick/jigsawtanslatedgoogle\nhttps://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\nhttps://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api\nall outputs from public kernels",
      "replies": [
        {
          "id": 887792,
          "postDate": "2020-06-15T21:59:42.737Z",
          "content": "<p>add some extra data:\npretrained models in <a href=\"https://huggingface.co/models\">huggingface community</a>\n<a href=\"https://github.com/leondz/hatespeechdata\">https://github.com/leondz/hatespeechdata</a>\n<a href=\"http://hatespeechdata.com/\">http://hatespeechdata.com/</a>\n<a href=\"https://github.com/HKUST-KnowComp/MLMA_hate_speech\">https://github.com/HKUST-KnowComp/MLMA_hate_speech</a>\n<a href=\"http://www.di.unito.it/~tutreeb/haspeede-evalita18/data.html\">http://www.di.unito.it/~tutreeb/haspeede-evalita18/data.html</a>\n<a href=\"https://github.com/rogersdepelle/OffComBR\">https://github.com/rogersdepelle/OffComBR</a>\n<a href=\"https://rdm.inesctec.pt/dataset/cs-2017-008\">https://rdm.inesctec.pt/dataset/cs-2017-008</a>\n<a href=\"https://coltekin.github.io/offensive-turkish/\">https://coltekin.github.io/offensive-turkish/</a> \n<a href=\"https://github.com/cicl2018/HateEvalTeam\">https://github.com/cicl2018/HateEvalTeam</a>\n<a href=\"https://competitions.codalab.org/competitions/19935#participate\">https://competitions.codalab.org/competitions/19935#participate</a></p>\n\n<p>and all external data posted in this thread</p>",
          "rawMarkdown": "add some extra data:\npretrained models in [huggingface community](https://huggingface.co/models)\nhttps://github.com/leondz/hatespeechdata\nhttp://hatespeechdata.com/\nhttps://github.com/HKUST-KnowComp/MLMA_hate_speech\nhttp://www.di.unito.it/~tutreeb/haspeede-evalita18/data.html\nhttps://github.com/rogersdepelle/OffComBR\nhttps://rdm.inesctec.pt/dataset/cs-2017-008\nhttps://coltekin.github.io/offensive-turkish/ \nhttps://github.com/cicl2018/HateEvalTeam\nhttps://competitions.codalab.org/competitions/19935#participate\n\nand all external data posted in this thread"
        }
      ]
    },
    {
      "id": 882185,
      "postDate": "2020-06-11T16:26:04.413Z",
      "content": "<p><a href=\"https://www.kaggle.com/ishivinal/contractions\">https://www.kaggle.com/ishivinal/contractions</a>\n<a href=\"https://github.com/OuassimADNANE/datasets/blob/master/badwords.csv\">https://github.com/OuassimADNANE/datasets/blob/master/badwords.csv</a></p>",
      "rawMarkdown": "https://www.kaggle.com/ishivinal/contractions\nhttps://github.com/OuassimADNANE/datasets/blob/master/badwords.csv"
    },
    {
      "id": 881919,
      "postDate": "2020-06-11T13:27:03.430Z",
      "content": "<p>Great datasets from shonenkov, his kernel is great too (<a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a>).\n- <a href=\"http://opus.nlpl.eu/OpenSubtitles-v2018.php\">http://opus.nlpl.eu/OpenSubtitles-v2018.php</a> -&gt; <a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling</a></p>",
      "rawMarkdown": "Great datasets from shonenkov, his kernel is great too (https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta).\n- http://opus.nlpl.eu/OpenSubtitles-v2018.php -&gt; https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling"
    },
    {
      "id": 869502,
      "postDate": "2020-06-01T03:57:58.413Z",
      "content": "<p>nice!</p>",
      "rawMarkdown": "nice!"
    },
    {
      "id": 868704,
      "postDate": "2020-05-31T12:30:38.533Z",
      "content": "<p><a href=\"https://www.textgain.com/portfolio/4chan-8chan-embeddings-textgain-technical-report-1/\">https://www.textgain.com/portfolio/4chan-8chan-embeddings-textgain-technical-report-1/</a></p>\n\n<p>They collected over 30 million messages from the publicly available /pol/ message boards on 4chan and 8chan, and compiled them into a model of toxic language use. The trained word embeddings (±0.4GB) are released for free.</p>",
      "rawMarkdown": "https://www.textgain.com/portfolio/4chan-8chan-embeddings-textgain-technical-report-1/\n\nThey collected over 30 million messages from the publicly available /pol/ message boards on 4chan and 8chan, and compiled them into a model of toxic language use. The trained word embeddings (±0.4GB) are released for free."
    },
    {
      "id": 848781,
      "postDate": "2020-05-15T08:14:43.880Z",
      "content": "<p><a href=\"https://commoncrawl.org\">https://commoncrawl.org</a>\n<a href=\"http://files.pushshift.io/\">http://files.pushshift.io/</a>\n<a href=\"http://files.pushshift.io/twitter/\">http://files.pushshift.io/twitter/</a>\n<a href=\"http://files.pushshift.io/reddit/comments/\">http://files.pushshift.io/reddit/comments/</a></p>",
      "rawMarkdown": "https://commoncrawl.org\nhttp://files.pushshift.io/\nhttp://files.pushshift.io/twitter/\nhttp://files.pushshift.io/reddit/comments/"
    },
    {
      "id": 830780,
      "postDate": "2020-05-02T22:20:20.133Z",
      "content": "<p>Guys, I do not understand quite clearly whether translating test set to English than inference is allowed or not?</p>",
      "rawMarkdown": "Guys, I do not understand quite clearly whether translating test set to English than inference is allowed or not?",
      "replies": [
        {
          "id": 868733,
          "postDate": "2020-05-31T12:52:00.723Z",
          "content": "<p>it is  allowed</p>",
          "rawMarkdown": "it is  allowed"
        },
        {
          "id": 888019,
          "postDate": "2020-06-16T04:24:57.033Z",
          "content": "<p>Translating datasets is permitted (as long as the underlying dataset is public). The datasets should be shared on this thread.</p>",
          "rawMarkdown": "Translating datasets is permitted (as long as the underlying dataset is public). The datasets should be shared on this thread."
        }
      ]
    },
    {
      "id": 788270,
      "postDate": "2020-03-27T14:28:22.970Z",
      "content": "<p>Harvard dictionaries, available for download here: \n<a href=\"http://www.wjh.harvard.edu/~inquirer/spreadsheet_guide.htm\">http://www.wjh.harvard.edu/~inquirer/spreadsheet_guide.htm</a></p>\n\n<p>Happiness scores, available here:\n<a href=\"http://hedonometer.org/api/v1/words/?format=json\">http://hedonometer.org/api/v1/words/?format=json</a></p>",
      "rawMarkdown": "Harvard dictionaries, available for download here: \nhttp://www.wjh.harvard.edu/~inquirer/spreadsheet_guide.htm\n\nHappiness scores, available here:\nhttp://hedonometer.org/api/v1/words/?format=json"
    },
    {
      "id": 785804,
      "postDate": "2020-03-25T12:01:32.040Z",
      "content": "<p>Can we train a model on external data that is not in english?</p>",
      "rawMarkdown": "Can we train a model on external data that is not in english?",
      "replies": [
        {
          "id": 786626,
          "postDate": "2020-03-26T03:59:36.960Z",
          "content": "<p><a href=\"/davidbnn92\">@davidbnn92</a> Great question. Yes you may! We don’t have a large enough labeled non-English datasets to be robustly inclusive of all of the languages represented in the test set. So bringing in external datasets that are non-English are very welcome. However it should be repeated that these datasets must be made available for public use and shared on this thread.</p>",
          "rawMarkdown": "@davidbnn92 Great question. Yes you may! We don’t have a large enough labeled non-English datasets to be robustly inclusive of all of the languages represented in the test set. So bringing in external datasets that are non-English are very welcome. However it should be repeated that these datasets must be made available for public use and shared on this thread.",
          "votes": 2
        }
      ]
    },
    {
      "id": 887205,
      "postDate": "2020-06-15T14:37:00.963Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 858717,
      "postDate": "2020-05-23T18:15:20.350Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 820241,
      "postDate": "2020-04-25T09:12:38.880Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 801514,
      "postDate": "2020-04-08T15:02:27Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 786136,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-03-25T16:46:30.540000",
      "content": "<p>About translations, we checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 853420,
          "author_name": "MaChaogong",
          "author_url": "",
          "post_date": "2020-05-19T07:11:32.720000",
          "content": "<p>hi, I have one question. in the rules saying\n<em>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</em>\nIf I post my external data just a few seconds ahead of Entry Deadline, then no team will have time to use this external data. However this external data is still available to all participants and also posted in time(prior to the Entry Deadline), that is to say, not breaking the rules. \nThen what's the meaning of posting external data here? is that fair?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 854383,
          "author_name": "MaChaogong",
          "author_url": "",
          "post_date": "2020-05-20T02:11:09.520000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"/martingorner\">@martingorner</a> sorry to @ you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857107,
          "author_name": "MaChaogong",
          "author_url": "",
          "post_date": "2020-05-22T09:54:56.170000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> need help</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857367,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-05-22T14:56:00.467000",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> Doing what you describe is within the rules and therefore permitted. The data you use should already be publicly available and licensed for that purpose.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857482,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2020-05-22T16:36:55.667000",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> the Entry deadline (along with Team Merger deadline) is one week prior to the Final submission deadline. With one week, there should be enough time to use the external data posted here.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 858614,
          "author_name": "MaChaogong",
          "author_url": "",
          "post_date": "2020-05-23T16:17:31.403000",
          "content": "<p>I don't think so. one week is not enough even for downloading some big dataset for me:))\nSo could you post your team's dataset as soon as possible? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 858756,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2020-05-23T19:22:59.130000",
          "content": "<p>My comment was not made for personal interest, rather I assumed you mistook the Entry deadline from the Final submission deadline (happened to me as well). \nWho said our team is using (big) external datasets? 😏 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 856481,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2020-05-21T18:31:22.370000",
      "content": "<p>french :\n<a href=\"https://github.com/marcoguerini/CONAN\">https://github.com/marcoguerini/CONAN</a>\n<a href=\"https://github.com/HKUST-KnowComp/MLMA_hate_speech\">https://github.com/HKUST-KnowComp/MLMA_hate_speech</a></p>\n\n<p>Turkish\n<a href=\"https://coltekin.github.io/offensive-turkish/\">https://coltekin.github.io/offensive-turkish/</a></p>\n\n<p>Spanish \n<a href=\"https://competitions.codalab.org/competitions/19935#learn_the_details\">https://competitions.codalab.org/competitions/19935#learn_the_details</a>\n<a href=\"https://zenodo.org/record/2592149#.Xrhi2GhKiUk\">https://zenodo.org/record/2592149#.Xrhi2GhKiUk</a></p>\n\n<p>Italian</p>\n\n<p><a href=\"http://www.di.unito.it/~tutreeb/haspeede-evalita18/index.html#\">http://www.di.unito.it/~tutreeb/haspeede-evalita18/index.html#</a>\n<a href=\"https://github.com/msang/hate-speech-corpus\">https://github.com/msang/hate-speech-corpus</a></p>\n\n<p>Russian : \n<a href=\"https://www.kaggle.com/blackmoon/russian-language-toxic-comments\">https://www.kaggle.com/blackmoon/russian-language-toxic-comments</a></p>\n\n<p>Portugesh:\n<a href=\"https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\">https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset</a>\n<a href=\"https://github.com/rogersdepelle/OffComBR\">https://github.com/rogersdepelle/OffComBR</a></p>\n\n<p><a href=\"https://huggingface.co/\">https://huggingface.co/</a>\n<a href=\"https://github.com/n-waves/multifit\">https://github.com/n-waves/multifit</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 875803,
          "author_name": "Alan Sun",
          "author_url": "",
          "post_date": "2020-06-06T06:55:38.840000",
          "content": "<p>I created a Kaggle dataset with all of these files that you mentioned, not sure if this is helpful but here it is:\n<a href=\"https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended\">https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 880386,
          "author_name": "Sxwat",
          "author_url": "",
          "post_date": "2020-06-10T08:48:56.493000",
          "content": "<p>Translated Russian comments into English\n<a href=\"https://www.kaggle.com/aybatov/toxic-russian-comments-from-pikabu-and-2ch\">https://www.kaggle.com/aybatov/toxic-russian-comments-from-pikabu-and-2ch</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880414,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-06-10T09:19:38.140000",
          "content": "<p>Thanks, I will try to preprocess them this week end. I am also on another competition so I did not have the time to do that. I was more into architecture and parameters until now. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 886340,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-06-15T00:20:16.273000",
          "content": "<p>sorry for the delay, the data are available in one .csv</p>\n\n<p><a href=\"https://www.kaggle.com/ludovick/hatespeechmulti\">https://www.kaggle.com/ludovick/hatespeechmulti</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 888712,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-16T14:17:50.613000",
          "content": "<p>Thank you. Much appreciated 👍 👍 💪 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 784578,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-03-24T10:41:11.930000",
      "content": "<p><a href=\"https://github.com/ssut/py-googletrans\">https://github.com/ssut/py-googletrans</a>\n<a href=\"https://github.com/terryyin/translate-python\">https://github.com/terryyin/translate-python</a></p>\n\n<p><a href=\"https://mymemory.translated.net/\">https://mymemory.translated.net/</a>\n<a href=\"https://cloud.google.com/translate/\">https://cloud.google.com/translate/</a>\n<a href=\"https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\">https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/</a>\n<a href=\"https://tech.yandex.com/translate/\">https://tech.yandex.com/translate/</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/notebooks\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/notebooks</a> any kernel + output here (just naming a few next) + of course all associated datasets\n<a href=\"https://www.kaggle.com/hamditarek/ensemble\">https://www.kaggle.com/hamditarek/ensemble</a>\n<a href=\"https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta</a>\n<a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large</a>\n<a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a>\n<a href=\"https://www.kaggle.com/yeayates21/xlm-roberta-augmentation-ssl-0-9417-pub-lb\">https://www.kaggle.com/yeayates21/xlm-roberta-augmentation-ssl-0-9417-pub-lb</a>\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>\n<a href=\"https://www.kaggle.com/mobassir/understanding-cross-lingual-models\">https://www.kaggle.com/mobassir/understanding-cross-lingual-models</a>\n<a href=\"https://drive.google.com/drive/folders/1hbcSRfvtTTlERs7remsRST2amIWAFVry\">https://drive.google.com/drive/folders/1hbcSRfvtTTlERs7remsRST2amIWAFVry</a></p>\n\n<p><a href=\"https://huggingface.co/models\">https://huggingface.co/models</a> any model here</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 887823,
      "author_name": "Fumihiro Kaneko",
      "author_url": "",
      "post_date": "2020-06-15T23:01:25.033000",
      "content": "<p>Notebook from previous competition\n<a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part1-eda\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part1-eda</a>\n<a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part2-usage\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part2-usage</a>\n<a href=\"https://www.kaggle.com/haqishen/jigsaw-predict\">https://www.kaggle.com/haqishen/jigsaw-predict</a></p>\n\n<p>Word embedding\n<a href=\"https://fasttext.cc/docs/en/english-vectors.html\">https://fasttext.cc/docs/en/english-vectors.html</a>\n<a href=\"https://fasttext.cc/docs/en/crawl-vectors.html\">https://fasttext.cc/docs/en/crawl-vectors.html</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 886328,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2020-06-14T23:49:53.723000",
      "content": "<p>may or may not useful : \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">https://www.kaggle.com/c/quora-insincere-questions-classification</a></p>\n\n<p>various langs : <a href=\"https://github.com/valeriobasile/hurtlex/tree/master/lexica\">https://github.com/valeriobasile/hurtlex/tree/master/lexica</a></p>\n\n<p>Portugese : <a href=\"https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\">https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset</a>\nSpanish : <a href=\"https://github.com/msang/hateval\">https://github.com/msang/hateval</a> , <a href=\"https://zenodo.org/record/2592149\">https://zenodo.org/record/2592149</a>\nFrench : <a href=\"https://github.com/HKUST-KnowComp/MLMA_hate_speech\">https://github.com/HKUST-KnowComp/MLMA_hate_speech</a>\nRussian : <a href=\"https://www.kaggle.com/blackmoon/russian-language-toxic-comments\">https://www.kaggle.com/blackmoon/russian-language-toxic-comments</a>\nTurkish : <a href=\"https://www.kaggle.com/ahmetax/hury-dataset\">https://www.kaggle.com/ahmetax/hury-dataset</a>\nPortugese : <a href=\"https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\">https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset</a> , <a href=\"https://github.com/rogersdepelle/OffComBR\">https://github.com/rogersdepelle/OffComBR</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 886262,
      "author_name": "Gena",
      "author_url": "",
      "post_date": "2020-06-14T20:56:03.183000",
      "content": "<p><a href=\"https://github.com/leondz/hatespeechdata\">https://github.com/leondz/hatespeechdata</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 885814,
      "author_name": "Dezső Ribli",
      "author_url": "",
      "post_date": "2020-06-14T13:52:44.927000",
      "content": "<p>I haven't used this yet, but may give it a shot. They have pretrained models.</p>\n\n<p>LASER Language-Agnostic SEntence Representations\n<a href=\"https://github.com/facebookresearch/LASER\">https://github.com/facebookresearch/LASER</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 828148,
      "author_name": "Deepak Rajpurohit",
      "author_url": "",
      "post_date": "2020-04-30T19:37:27.907000",
      "content": "<p>nice</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 887834,
      "author_name": "Chun Ming Lee",
      "author_url": "",
      "post_date": "2020-06-15T23:10:57.367000",
      "content": "<p><a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api</a>\n<a href=\"https://www.kaggle.com/rafiko1/translated-train-bias-all-langs\">https://www.kaggle.com/rafiko1/translated-train-bias-all-langs</a></p>\n\n<p><a href=\"https://github.com/allenai/allennlp\">https://github.com/allenai/allennlp</a>\n<a href=\"https://github.com/pytorch/fairseq\">https://github.com/pytorch/fairseq</a>\n<a href=\"https://github.com/flairNLP/flair\">https://github.com/flairNLP/flair</a>\n<a href=\"https://github.com/facebookresearch/LASER\">https://github.com/facebookresearch/LASER</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 823388,
      "author_name": "Alex Shonenkov",
      "author_url": "",
      "post_date": "2020-04-27T15:50:15.050000",
      "content": "<p><a href=\"http://opus.nlpl.eu/OpenSubtitles-v2018.php\">http://opus.nlpl.eu/OpenSubtitles-v2018.php</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 792277,
      "author_name": "takuoko",
      "author_url": "",
      "post_date": "2020-03-31T03:56:54.020000",
      "content": "<p><a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 784227,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2020-03-24T02:51:04.673000",
      "content": "<p>Tbh I'm surprised that external data is allowed here. </p>\n\n<p>Please correct me if I'm wrong (I haven't really done any EDA on the data yet) but, do you think that stuff like abusing Google's translation API may ruin the purposes of this competition? The challenge is to make a model that generalizes to many languages using only English as a starting point, now if you can just eyeball the test set to see which languages are there and translate your train set accordingly, what makes it different?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 784581,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-24T10:42:55.720000",
          "content": "<p>If there is Internet access on in scoring kernels, you can just translate the test data to English. Hope for clarification here.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 785529,
          "author_name": "Dhananjay Raut",
          "author_url": "",
          "post_date": "2020-03-25T06:15:02.313000",
          "content": "<p>We don't even need internet access in the inference kernel to do that. as we already have access to entire test set which we can translate beforehand.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 785645,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-25T08:22:38.020000",
          "content": "<p>Really? Oh...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 785653,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2020-03-25T08:32:12.823000",
          "content": "<p>All my will to commit just vanished</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 785663,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-03-25T08:42:52.347000",
          "content": "<p>I have a bad feeling about this comp;\n- First Many possibilites to cheat i can think of;\n- Secondly training with 5k examples is so close to .80 AUC's benchmark which was trained with whole data [wrt TF TPU kernel]?\n- Plus PyTorch XLA is not stable yet\n- Plus the comp is purely overfit the data to win type as well (as it seems for now)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 785807,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-25T12:03:01.827000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 785811,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-25T12:10:04.233000",
          "content": "<p>Hand-labeling is actually not allowed based on the rules:</p>\n\n<p>&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n\n<p>But I don't see any rule forbidding automatic translation. Actually this had been done in past competitions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 786134,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-03-25T16:43:50.387000",
          "content": "<p>&gt; Plus PyTorch XLA is not stable yet\nYes, that has been clearly announced multiple times. PyTorch TPU support on Kaggle is experimental at this point.</p>\n\n<p>Tensorflow is the way to go for TPUs for now: <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 786505,
          "author_name": "Jerry Qu",
          "author_url": "",
          "post_date": "2020-03-26T00:07:56.613000",
          "content": "<blockquote>\n  <p>One could also label the test set by hand. Nobody reads that many languages, but a team could. How do you rule that out?</p>\n</blockquote>\n\n<p>'Submissions to this competition must be made by an output from a Kaggle Notebook/Script.' - From the code requirements</p>\n\n<p>I believe this will weed out submissions that were hand labeled, as you could easily run their notebook to verify.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 786633,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-03-26T04:21:26.677000",
          "content": "<p>There are smart ways to hide it;  Like what about this, I trained my model on the test set as well (after hand-labelling) -&gt; Cached the weights -&gt; External Dataset Import -&gt; Use it freely :) ;\nThoughts?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 811121,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-17T16:00:14.820000",
          "content": "<p>But again problem with translation is that, many times the text gets de toxified.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857679,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-05-22T20:17:52.540000",
          "content": "<p>does that mean we can use the test set if not hand labeled ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 888030,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-16T04:33:48.957000",
          "content": "<p>The test set <strong>should not be hand-labeled</strong>. <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138118#786136\">As stated previously by Martin</a>, datasets translated into the languages that are represented in the test set is legitimate, but if you are using the test set directly in any way that includes manual labels, then that would be a prohibited use of the test set. If hand-labeling of the test set is discovered in the review of a winning solution, it is subject to disqualification. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 888684,
          "author_name": "huiqin",
          "author_url": "",
          "post_date": "2020-06-16T14:01:12.910000",
          "content": "<p>@Julia\nI want to make it clear whether the pseudo label legal ? e.g.  I use my model to predict the test data and get the toxic target.  Then I use the predicted toxic target as the label of test and add the test data into my training data for my model training.  Does it allowed ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 891237,
          "author_name": "cat",
          "author_url": "",
          "post_date": "2020-06-18T03:26:25.467000",
          "content": "<p>I think so. Pseudo labelling is used in every NLP competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 897396,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-22T20:37:06.950000",
          "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> Pseudo-labeling is permitted. Manual hand-labeling of the test set is not.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 890027,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2020-06-17T09:08:10.380000",
      "content": "<p>all dataset has been declaired by others</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 898391,
      "author_name": "Matt Yates",
      "author_url": "",
      "post_date": "2020-06-23T13:53:50.537000",
      "content": "<p><strong>Added myself:</strong>\n<a href=\"https://www.kaggle.com/yeayates21/jigsawmultilingualrobertaavgblender\">https://www.kaggle.com/yeayates21/jigsawmultilingualrobertaavgblender</a>\n<a href=\"https://www.kaggle.com/yeayates21/jigsawtpuxlmrobertacopypickledata\">https://www.kaggle.com/yeayates21/jigsawtpuxlmrobertacopypickledata</a></p>\n\n<p><strong>Added by others:</strong>\n<a href=\"https://www.kaggle.com/hamditarek/blendings\">https://www.kaggle.com/hamditarek/blendings</a>\n<a href=\"https://www.kaggle.com/hamditarek/ensemble-version-96\">https://www.kaggle.com/hamditarek/ensemble-version-96</a>\n<a href=\"https://www.kaggle.com/hamditarek/tfidf\">https://www.kaggle.com/hamditarek/tfidf</a>\n<a href=\"https://www.kaggle.com/hamditarek/009248\">https://www.kaggle.com/hamditarek/009248</a>\n<a href=\"https://www.kaggle.com/hamditarek/009259\">https://www.kaggle.com/hamditarek/009259</a>\n<a href=\"https://www.kaggle.com/hamditarek/009354\">https://www.kaggle.com/hamditarek/009354</a>\n<a href=\"https://www.kaggle.com/hamditarek/009383\">https://www.kaggle.com/hamditarek/009383</a>\n<a href=\"https://www.kaggle.com/hamditarek/009406\">https://www.kaggle.com/hamditarek/009406</a>\n<a href=\"https://www.kaggle.com/hamditarek/009423\">https://www.kaggle.com/hamditarek/009423</a>\n<a href=\"https://www.kaggle.com/hamditarek/abhishek\">https://www.kaggle.com/hamditarek/abhishek</a>\n<a href=\"https://www.kaggle.com/hamditarek/blending\">https://www.kaggle.com/hamditarek/blending</a>\n<a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 895766,
      "author_name": "Ouassim Adnane",
      "author_url": "",
      "post_date": "2020-06-21T15:36:17.180000",
      "content": "<p><a href=\"https://www.kaggle.com/ibtesama/hatespeech\">https://www.kaggle.com/ibtesama/hatespeech</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 894826,
      "author_name": "Gena",
      "author_url": "",
      "post_date": "2020-06-20T19:31:10.967000",
      "content": "<p>I guess that any dataset that is marked as \"research-only\", \"not for commercial use\", etc. can't be used in any way to train the model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 897394,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-22T20:34:41.930000",
          "content": "<p>In this competition, external data licensed for non-commercial, research, and/or academic-use is permitted.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 897405,
          "author_name": "Gena",
          "author_url": "",
          "post_date": "2020-06-22T20:43:38.480000",
          "content": "<p>Thank you! I was a bit late asking this, so have used only fully open data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 893839,
      "author_name": "ma7555",
      "author_url": "",
      "post_date": "2020-06-20T02:15:33.207000",
      "content": "<p><a href=\"https://www.kaggle.com/ma7555/jigsaw-train-translated\">https://www.kaggle.com/ma7555/jigsaw-train-translated</a>\n<a href=\"https://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api\">https://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 893781,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-06-19T22:09:34.867000",
      "content": "<p>I addition to the list I may use the following:-\n<a href=\"https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\">https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated</a>\n<a href=\"https://www.kaggle.com/bamps53/val-en-df\">https://www.kaggle.com/bamps53/val-en-df</a>\n<a href=\"https://www.kaggle.com/bamps53/test-en-df\">https://www.kaggle.com/bamps53/test-en-df</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 891384,
      "author_name": "Urvish",
      "author_url": "",
      "post_date": "2020-06-18T06:31:49.893000",
      "content": "<p><a href=\"https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended\">https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended</a>\n<a href=\"https://huggingface.co/models\">https://huggingface.co/models</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887850,
      "author_name": "vecxoz",
      "author_url": "",
      "post_date": "2020-06-15T23:36:51.243000",
      "content": "<p>May or may not use:</p>\n\n<p><a href=\"https://github.com/uliontse/translators\">https://github.com/uliontse/translators</a>\n<a href=\"https://github.com/sloria/textblob\">https://github.com/sloria/textblob</a>\n<a href=\"https://github.com/littlecodersh/translation\">https://github.com/littlecodersh/translation</a>\n<a href=\"https://github.com/soimort/translate-shell\">https://github.com/soimort/translate-shell</a></p>\n\n<p><a href=\"https://github.com/Rayraegah/warui\">https://github.com/Rayraegah/warui</a>\n<a href=\"https://github.com/CRomano31415/SpanishProfanity\">https://github.com/CRomano31415/SpanishProfanity</a>\n<a href=\"https://github.com/mmcclarty/lyrics_melange\">https://github.com/mmcclarty/lyrics_melange</a>\n<a href=\"https://github.com/TheSoma300/random-italian-curse-generator\">https://github.com/TheSoma300/random-italian-curse-generator</a>\n<a href=\"https://github.com/ChaseFlorell/jQuery.ProfanityFilter\">https://github.com/ChaseFlorell/jQuery.ProfanityFilter</a>\n<a href=\"https://github.com/Zeindelf/badwords\">https://github.com/Zeindelf/badwords</a>\n<a href=\"https://github.com/BotanUA/russian_swear_words\">https://github.com/BotanUA/russian_swear_words</a>\n<a href=\"https://github.com/voyula/turkish-bad-words\">https://github.com/voyula/turkish-bad-words</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887631,
      "author_name": "godelscat",
      "author_url": "",
      "post_date": "2020-06-15T19:23:58.120000",
      "content": "<p><a href=\"https://www.kaggle.com/ludovick/jigsawtanslatedgoogle\">Google translated data</a>\noutputs of public notebooks in this competition\npre-trained models in <a href=\"https://huggingface.co/models\">huggingface</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887586,
      "author_name": "Chew Kok Wah",
      "author_url": "",
      "post_date": "2020-06-15T18:49:44.310000",
      "content": "<p><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">https://www.kaggle.com/c/quora-insincere-questions-classification</a>\n<a href=\"https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\">https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large</a>\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api</a>\n<a href=\"https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\">https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data</a>\n<a href=\"https://github.com/google-research/xtreme\">https://github.com/google-research/xtreme</a>\n<a href=\"https://github.com/google-research-datasets/paws/tree/master/pawsx\">https://github.com/google-research-datasets/paws/tree/master/pawsx</a>\n<a href=\"https://github.com/facebookresearch/XNLI\">https://github.com/facebookresearch/XNLI</a>\n<a href=\"https://mymemory.translated.net/\">https://mymemory.translated.net/</a>\n<a href=\"https://cloud.google.com/translate/\">https://cloud.google.com/translate/</a>\n<a href=\"https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\">https://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/</a>\n<a href=\"https://commoncrawl.org\">https://commoncrawl.org</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887387,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2020-06-15T16:37:34.527000",
      "content": "<p><a href=\"https://huggingface.co/novinsh/xlm-roberta-large-toxicomments-12k\">https://huggingface.co/novinsh/xlm-roberta-large-toxicomments-12k</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887362,
      "author_name": "Issagali Konysbayev [dsmlkz]",
      "author_url": "",
      "post_date": "2020-06-15T16:21:19.770000",
      "content": "<p>Thanks to the creators of the datasets below (might be used for the final submission):\n<a href=\"https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\">https://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large</a>\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\">https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api</a>\n<a href=\"https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\">https://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data</a>\n<a href=\"https://www.kaggle.com/shonenkov/jigsaw-public-baseline-results\">https://www.kaggle.com/shonenkov/jigsaw-public-baseline-results</a>\n<a href=\"https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\">https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling</a>\n<a href=\"https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\">https://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887238,
      "author_name": "arseny-n",
      "author_url": "",
      "post_date": "2020-06-15T15:05:22.700000",
      "content": "<p>Data for previous competition  <a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification\">https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887093,
      "author_name": "Igor Buzhinsky",
      "author_url": "",
      "post_date": "2020-06-15T13:26:09.237000",
      "content": "<p>ELECTRA (monolingual) model: <a href=\"https://huggingface.co/google/electra-large-discriminator\">https://huggingface.co/google/electra-large-discriminator</a>\nWe also extracted the output of a public ensembling kernel <a href=\"https://www.kaggle.com/hamditarek/ensemble?scriptVersionId=35925815\">https://www.kaggle.com/hamditarek/ensemble?scriptVersionId=35925815</a> to a separate dataset: <a href=\"https://www.kaggle.com/buzhinsky/hamditarekensemblejun12\">https://www.kaggle.com/buzhinsky/hamditarekensemblejun12</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 886916,
      "author_name": "Rishabh Jha",
      "author_url": "",
      "post_date": "2020-06-15T11:08:39.780000",
      "content": "<p>4chan data repository: <a href=\"https://zenodo.org/record/3603292#.XudWL0UzaUk\">https://zenodo.org/record/3603292#.XudWL0UzaUk</a>\nReddit comment history since 2005: <a href=\"https://console.cloud.google.com/bigquery?project=fh-bigquery&amp;redirect_from_classic=true&amp;p=fh-bigquery&amp;d=reddit_comments&amp;page=dataset\">https://console.cloud.google.com/bigquery?project=fh-bigquery&amp;redirect_from_classic=true&amp;p=fh-bigquery&amp;d=reddit_comments&amp;page=dataset</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 886244,
      "author_name": "Striderl",
      "author_url": "",
      "post_date": "2020-06-14T20:28:18.070000",
      "content": "<p>some additional Wikipedia talk comments here: <br>\n<a href=\"https://figshare.com/projects/Wikipedia_Talk/16731\">https://figshare.com/projects/Wikipedia_Talk/16731</a></p>\n\n<p>multi-lingual embedding: \n<a href=\"https://github.com/facebookresearch/MUSE\">https://github.com/facebookresearch/MUSE</a>\n<a href=\"https://fasttext.cc/docs/en/aligned-vectors.html\">https://fasttext.cc/docs/en/aligned-vectors.html</a>\n<a href=\"https://nlp.stanford.edu/projects/glove/\">https://nlp.stanford.edu/projects/glove/</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 884674,
      "author_name": "nzholmes",
      "author_url": "",
      "post_date": "2020-06-13T14:41:28.263000",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Do the submission results from the public notebook count as external data? For example this one: <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>  Are public submission results like this one allowed to be used in final submissions?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 884676,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-06-13T14:42:49.750000",
          "content": "<p>You can always submit public notebooks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884699,
          "author_name": "nzholmes",
          "author_url": "",
          "post_date": "2020-06-13T15:01:30.940000",
          "content": "<p>Some results in public notebook cannot be replicated like the example I list.😂 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884905,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-06-13T17:58:37.393000",
          "content": "<p>It is the same with the runs you are doing yourself, no difference between public kernels and your own. There is randomness involved in running NNs.</p>\n\n<p>Submitting external submission.csv files was explicitly allowed in this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 888016,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-16T04:23:52.220000",
          "content": "<p><a href=\"/nzholmes\">@nzholmes</a> - all public notebooks are automatically licensed under Apache 2.0 license. As a result, its contents and output are all permitted to be used by any user. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 891539,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-06-18T08:44:54.230000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> - I am a bit concerned about people using inference only notebooks that do not include the training scripts. Is it allowed to use the output of these kernels?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 897512,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-22T23:29:17.390000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I'm just seeing your question, so apologies for the delayed response. I may not be clearly comprehending your concern. Based on what I do understand, the outputs have been made available on Kaggle to use by the notebook publisher, as a result of making the notebook public. That said, competition winners would still be expected to have license and rights to open-source the code used to achieve their result.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 897781,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-06-23T05:03:40.813000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I think the issue is that you do not know how the results were generated in these kernels.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 884475,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-13T12:01:56.467000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 887792,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-15T21:59:42.737000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 882185,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-11T16:26:04.413000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 881919,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-11T13:27:03.430000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 869502,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-01T03:57:58.413000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 868704,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-31T12:30:38.533000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 848781,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-15T08:14:43.880000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 830780,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-02T22:20:20.133000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 868733,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-31T12:52:00.723000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 888019,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-16T04:24:57.033000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 788270,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-27T14:28:22.970000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 785804,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-25T12:01:32.040000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 786626,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-26T03:59:36.960000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 887205,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-15T14:37:00.963000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 858717,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-23T18:15:20.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820241,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T09:12:38.880000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 801514,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-08T15:02:27",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "783984": "Per the [Competition Rules](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/rules), freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).\n\nNote that if you wish to use Kaggle's TPU integration, you cannot use privately-held external datasets. Instead, you'll want to upload a public dataset to Kaggle to use and declare those here.\n\nOnce someone posts an external dataset to this thread, you do not need to re-post it if you are using the same one.\n\nYou only need to declare the original dataset used; you do not need to declare re-labeled or augmented or otherwise processed versions of datasets. Pre-trained models can be declared, as well; however, models resulting from your own original work, that you have trained yourself offline do not need to be shared/declared.",
    "786136": "About translations, we checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify.",
    "856481": "french :\nhttps://github.com/marcoguerini/CONAN\nhttps://github.com/HKUST-KnowComp/MLMA_hate_speech\n\n\nTurkish\nhttps://coltekin.github.io/offensive-turkish/\n\n\nSpanish \nhttps://competitions.codalab.org/competitions/19935#learn_the_details\nhttps://zenodo.org/record/2592149#.Xrhi2GhKiUk\n\nItalian\n\nhttp://www.di.unito.it/~tutreeb/haspeede-evalita18/index.html#\nhttps://github.com/msang/hate-speech-corpus\n\nRussian : \nhttps://www.kaggle.com/blackmoon/russian-language-toxic-comments\n\n\nPortugesh:\nhttps://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\nhttps://github.com/rogersdepelle/OffComBR\n\nhttps://huggingface.co/\nhttps://github.com/n-waves/multifit",
    "784578": "https://github.com/ssut/py-googletrans\nhttps://github.com/terryyin/translate-python\n\nhttps://mymemory.translated.net/\nhttps://cloud.google.com/translate/\nhttps://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\nhttps://tech.yandex.com/translate/\n\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/notebooks any kernel + output here (just naming a few next) + of course all associated datasets\nhttps://www.kaggle.com/hamditarek/ensemble\nhttps://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\nhttps://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\nhttps://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\nhttps://www.kaggle.com/yeayates21/xlm-roberta-augmentation-ssl-0-9417-pub-lb\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\nhttps://www.kaggle.com/mobassir/understanding-cross-lingual-models\nhttps://drive.google.com/drive/folders/1hbcSRfvtTTlERs7remsRST2amIWAFVry\n\nhttps://huggingface.co/models any model here",
    "887823": "Notebook from previous competition\nhttps://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part1-eda\nhttps://www.kaggle.com/christofhenkel/how-to-preprocessing-for-glove-part2-usage\nhttps://www.kaggle.com/haqishen/jigsaw-predict\n\nWord embedding\nhttps://fasttext.cc/docs/en/english-vectors.html\nhttps://fasttext.cc/docs/en/crawl-vectors.html\n",
    "886328": "may or may not useful : \nhttps://www.kaggle.com/c/quora-insincere-questions-classification\n\nvarious langs : https://github.com/valeriobasile/hurtlex/tree/master/lexica\n\nPortugese : https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset\nSpanish : https://github.com/msang/hateval , https://zenodo.org/record/2592149\nFrench : https://github.com/HKUST-KnowComp/MLMA_hate_speech\nRussian : https://www.kaggle.com/blackmoon/russian-language-toxic-comments\nTurkish : https://www.kaggle.com/ahmetax/hury-dataset\nPortugese : https://github.com/paulafortuna/Portuguese-Hate-Speech-Dataset , https://github.com/rogersdepelle/OffComBR",
    "886262": "https://github.com/leondz/hatespeechdata",
    "885814": "I haven't used this yet, but may give it a shot. They have pretrained models.\n\nLASER Language-Agnostic SEntence Representations\nhttps://github.com/facebookresearch/LASER",
    "828148": "nice",
    "887834": "https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/rafiko1/translated-train-bias-all-langs\n\nhttps://github.com/allenai/allennlp\nhttps://github.com/pytorch/fairseq\nhttps://github.com/flairNLP/flair\nhttps://github.com/facebookresearch/LASER",
    "823388": "http://opus.nlpl.eu/OpenSubtitles-v2018.php ",
    "792277": "https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038",
    "784227": "Tbh I'm surprised that external data is allowed here. \n\nPlease correct me if I'm wrong (I haven't really done any EDA on the data yet) but, do you think that stuff like abusing Google's translation API may ruin the purposes of this competition? The challenge is to make a model that generalizes to many languages using only English as a starting point, now if you can just eyeball the test set to see which languages are there and translate your train set accordingly, what makes it different?",
    "890027": "all dataset has been declaired by others",
    "898391": "**Added myself:**\nhttps://www.kaggle.com/yeayates21/jigsawmultilingualrobertaavgblender\nhttps://www.kaggle.com/yeayates21/jigsawtpuxlmrobertacopypickledata\n\n**Added by others:**\nhttps://www.kaggle.com/hamditarek/blendings\nhttps://www.kaggle.com/hamditarek/ensemble-version-96\nhttps://www.kaggle.com/hamditarek/tfidf\nhttps://www.kaggle.com/hamditarek/009248\nhttps://www.kaggle.com/hamditarek/009259\nhttps://www.kaggle.com/hamditarek/009354\nhttps://www.kaggle.com/hamditarek/009383\nhttps://www.kaggle.com/hamditarek/009406\nhttps://www.kaggle.com/hamditarek/009423\nhttps://www.kaggle.com/hamditarek/abhishek\nhttps://www.kaggle.com/hamditarek/blending\nhttps://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling",
    "895766": "https://www.kaggle.com/ibtesama/hatespeech",
    "894826": "I guess that any dataset that is marked as \"research-only\", \"not for commercial use\", etc. can't be used in any way to train the model?",
    "893839": "https://www.kaggle.com/ma7555/jigsaw-train-translated\nhttps://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api",
    "893781": "I addition to the list I may use the following:-\nhttps://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\nhttps://www.kaggle.com/bamps53/val-en-df\nhttps://www.kaggle.com/bamps53/test-en-df",
    "891384": "[https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended](https://www.kaggle.com/alansun17904/toxic-comment-detection-multilingual-extended)\n[https://huggingface.co/models](https://huggingface.co/models)",
    "887850": "May or may not use:\n\nhttps://github.com/uliontse/translators\nhttps://github.com/sloria/textblob\nhttps://github.com/littlecodersh/translation\nhttps://github.com/soimort/translate-shell\n\nhttps://github.com/Rayraegah/warui\nhttps://github.com/CRomano31415/SpanishProfanity\nhttps://github.com/mmcclarty/lyrics_melange\nhttps://github.com/TheSoma300/random-italian-curse-generator\nhttps://github.com/ChaseFlorell/jQuery.ProfanityFilter\nhttps://github.com/Zeindelf/badwords\nhttps://github.com/BotanUA/russian_swear_words\nhttps://github.com/voyula/turkish-bad-words",
    "887631": "[Google translated data](https://www.kaggle.com/ludovick/jigsawtanslatedgoogle)\noutputs of public notebooks in this competition\npre-trained models in [huggingface](https://huggingface.co/models)",
    "887586": "https://www.kaggle.com/c/quora-insincere-questions-classification\nhttps://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\nhttps://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\nhttps://github.com/google-research/xtreme\nhttps://github.com/google-research-datasets/paws/tree/master/pawsx\nhttps://github.com/facebookresearch/XNLI\nhttps://mymemory.translated.net/\nhttps://cloud.google.com/translate/\nhttps://azure.microsoft.com/en-us/services/cognitive-services/translator-text-api/\nhttps://commoncrawl.org",
    "887387": "https://huggingface.co/novinsh/xlm-roberta-large-toxicomments-12k",
    "887362": "Thanks to the creators of the datasets below (might be used for the final submission):\nhttps://www.kaggle.com/riblidezso/jigsaw-mlm-finetuned-xlm-r-large\nhttps://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/shonenkov/jigsaw-public-baseline-train-data\nhttps://www.kaggle.com/shonenkov/jigsaw-public-baseline-results\nhttps://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling\nhttps://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated",
    "887238": "Data for previous competition  https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification",
    "887093": "ELECTRA (monolingual) model: https://huggingface.co/google/electra-large-discriminator\nWe also extracted the output of a public ensembling kernel https://www.kaggle.com/hamditarek/ensemble?scriptVersionId=35925815 to a separate dataset: https://www.kaggle.com/buzhinsky/hamditarekensemblejun12",
    "886916": "4chan data repository: https://zenodo.org/record/3603292#.XudWL0UzaUk\nReddit comment history since 2005: https://console.cloud.google.com/bigquery?project=fh-bigquery&amp;redirect_from_classic=true&amp;p=fh-bigquery&amp;d=reddit_comments&amp;page=dataset",
    "886244": "some additional Wikipedia talk comments here:  \nhttps://figshare.com/projects/Wikipedia_Talk/16731\n\nmulti-lingual embedding: \nhttps://github.com/facebookresearch/MUSE\nhttps://fasttext.cc/docs/en/aligned-vectors.html\nhttps://nlp.stanford.edu/projects/glove/",
    "884674": "@juliaelliott Do the submission results from the public notebook count as external data? For example this one: https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta  Are public submission results like this one allowed to be used in final submissions?",
    "884475": "https://dumps.wikimedia.org/\nhttps://www.kaggle.com/ludovick/jigsawtanslatedgoogle\nhttps://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api\nhttps://www.kaggle.com/kashnitsky/jigsaw-multilingual-toxic-test-translated\nhttps://www.kaggle.com/ma7555/jigsaw-train-translated-yandex-api\nall outputs from public kernels",
    "882185": "https://www.kaggle.com/ishivinal/contractions\nhttps://github.com/OuassimADNANE/datasets/blob/master/badwords.csv",
    "881919": "Great datasets from shonenkov, his kernel is great too (https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta).\n- http://opus.nlpl.eu/OpenSubtitles-v2018.php -&gt; https://www.kaggle.com/shonenkov/open-subtitles-toxic-pseudo-labeling",
    "869502": "nice!",
    "868704": "https://www.textgain.com/portfolio/4chan-8chan-embeddings-textgain-technical-report-1/\n\nThey collected over 30 million messages from the publicly available /pol/ message boards on 4chan and 8chan, and compiled them into a model of toxic language use. The trained word embeddings (±0.4GB) are released for free.",
    "848781": "https://commoncrawl.org\nhttp://files.pushshift.io/\nhttp://files.pushshift.io/twitter/\nhttp://files.pushshift.io/reddit/comments/",
    "830780": "Guys, I do not understand quite clearly whether translating test set to English than inference is allowed or not?",
    "788270": "Harvard dictionaries, available for download here: \nhttp://www.wjh.harvard.edu/~inquirer/spreadsheet_guide.htm\n\nHappiness scores, available here:\nhttp://hedonometer.org/api/v1/words/?format=json",
    "785804": "Can we train a model on external data that is not in english?",
    "887205": "",
    "858717": "",
    "820241": "",
    "801514": ""
  }
}