{
  "id": 141377,
  "title": "Multilingual training dataset ",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141377",
  "author_name": "",
  "post_date": "2020-04-05T19:04:46.659756600Z",
  "votes": 82,
  "comment_count": 30,
  "views": 0,
  "content": "<p>I created a multilingual <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api#jigsaw-toxic-comment-train-google-fr-cleaned.csv\">train dataset</a> using <code>translates</code> with Google API. And suggest sharing other variations of training dataset on this topic. </p>",
  "messages": [
    {
      "id": "798715",
      "postDate": "04/05/2020 19:04:46",
      "content": "<p>I created a multilingual <a href=\"https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api#jigsaw-toxic-comment-train-google-fr-cleaned.csv\">train dataset</a> using <code>translates</code> with Google API. And suggest sharing other variations of training dataset on this topic. </p>",
      "rawMarkdown": "I created a multilingual [train dataset](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api#jigsaw-toxic-comment-train-google-fr-cleaned.csv) using ```translates``` with Google API. And suggest sharing other variations of training dataset on this topic.",
      "votes": null
    },
    {
      "id": "798852",
      "postDate": "04/05/2020 23:13:21",
      "content": "<p>Hey Michael! Amazing work!\nThanks for sharing</p>",
      "rawMarkdown": "Hey Michael! Amazing work!\nThanks for sharing",
      "votes": null
    },
    {
      "id": "798863",
      "postDate": "04/05/2020 23:40:23",
      "content": "<p>I've only eyeballed the data but it looks like some of the translated comments aren't aligned with their IDs?</p>",
      "rawMarkdown": "I've only eyeballed the data but it looks like some of the translated comments aren't aligned with their IDs?",
      "votes": null
    },
    {
      "id": "799202",
      "postDate": "04/06/2020 09:10:02",
      "content": "<p>Nice job, thanks for sharing <a href=\"/miklgr500\">@miklgr500</a>! I am wondering which library did you use, <a href=\"https://github.com/terryyin/translate-python/\">translate-python </a> or <a href=\"https://github.com/ssut/py-googletrans\">googletrans</a>?</p>",
      "rawMarkdown": "Nice job, thanks for sharing @miklgr500! I am wondering which library did you use, [translate-python ](https://github.com/terryyin/translate-python/) or [googletrans](https://github.com/ssut/py-googletrans)?",
      "votes": null
    },
    {
      "id": "799280",
      "postDate": "04/06/2020 10:35:06",
      "content": "<p>Yep, thank you. I need 3 hours to fix it.</p>",
      "rawMarkdown": "Yep, thank you. I need 3 hours to fix it.",
      "votes": null
    },
    {
      "id": "799436",
      "postDate": "04/06/2020 12:52:08",
      "content": "<p>Hi Michael, a huge thanks for your effort during this competition .\nPlease did u fix the ID of comments and i wonder how much this dataset improve your score ?</p>",
      "rawMarkdown": "Hi Michael, a huge thanks for your effort during this competition .\nPlease did u fix the ID of comments and i wonder how much this dataset improve your score ?",
      "votes": null
    },
    {
      "id": "799503",
      "postDate": "04/06/2020 13:49:36",
      "content": "<p><a href=\"/haythemtellili5\">@haythemtellili5</a> <a href=\"/leecming\">@leecming</a> the new version of the dataset is available.</p>",
      "rawMarkdown": "haythemtellili5 @leecming the new version of the dataset is available.",
      "votes": null
    },
    {
      "id": "799508",
      "postDate": "04/06/2020 13:53:12",
      "content": "<p>Wow, how did you overcome api limits? Awesome work</p>",
      "rawMarkdown": "Wow, how did you overcome api limits? Awesome work",
      "votes": null
    },
    {
      "id": "799511",
      "postDate": "04/06/2020 13:54:51",
      "content": "<p>I used <a href=\"https://pypi.org/project/translators/\">translators</a> library with some modification for accelerating translation and <a href=\"https://docs.python.org/3/library/multiprocessing.html\">multiprocessing</a>.</p>",
      "rawMarkdown": "I used [translators](https://pypi.org/project/translators/) library with some modification for accelerating translation and [multiprocessing](https://docs.python.org/3/library/multiprocessing.html).",
      "votes": null
    },
    {
      "id": "799531",
      "postDate": "04/06/2020 14:20:26",
      "content": "<p>Thank, it'll be interesting to see how this shakes things up!</p>",
      "rawMarkdown": "Thank, it'll be interesting to see how this shakes things up!",
      "votes": null
    },
    {
      "id": "799552",
      "postDate": "04/06/2020 14:33:33",
      "content": "<p>It feature of the<a href=\"https://pypi.org/project/translators/\"><code>translators</code></a> library.</p>",
      "rawMarkdown": "It feature of the[```translators```](https://pypi.org/project/translators/) library.",
      "votes": null
    },
    {
      "id": "799597",
      "postDate": "04/06/2020 15:03:44",
      "content": "<p>And which API did you use there?</p>",
      "rawMarkdown": "And which API did you use there?",
      "votes": null
    },
    {
      "id": "799642",
      "postDate": "04/06/2020 15:52:53",
      "content": "<p>Nice job! Upvoting the post and the dataset as well.</p>",
      "rawMarkdown": "Nice job! Upvoting the post and the dataset as well.",
      "votes": null
    },
    {
      "id": "799894",
      "postDate": "04/06/2020 21:41:58",
      "content": "<p>Thanks for sharing.\nWhat is the difference between the cleaned and not clean version ?</p>",
      "rawMarkdown": "Thanks for sharing.\nWhat is the difference between the cleaned and not clean version ?",
      "votes": null
    },
    {
      "id": "799970",
      "postDate": "04/07/2020 00:05:17",
      "content": "<p>I tried translators for translate but I cannot translate all.(it does not include sleep).\nI used .google method.</p>\n\n<p>Did you use sleep or other waiting function?</p>",
      "rawMarkdown": "I tried translators for translate but I cannot translate all.(it does not include sleep).\nI used .google method.\n\nDid you use sleep or other waiting function?",
      "votes": null
    },
    {
      "id": "800097",
      "postDate": "04/07/2020 04:49:35",
      "content": "<p><a href=\"https://www.kaggle.com/miklgr500/how-to-use-translators-for-comments-translation/log?scriptVersionId=31557415\">How to use \"translators\" for comments translation </a></p>",
      "rawMarkdown": "[How to use \"translators\" for comments translation ](https://www.kaggle.com/miklgr500/how-to-use-translators-for-comments-translation/log?scriptVersionId=31557415)",
      "votes": null
    },
    {
      "id": "800524",
      "postDate": "04/07/2020 13:40:08",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "800536",
      "postDate": "04/07/2020 13:54:39",
      "content": "<p>Cleaned files dosn't have NaN values. </p>",
      "rawMarkdown": "Cleaned files dosn't have NaN values.",
      "votes": null
    },
    {
      "id": "804845",
      "postDate": "04/12/2020 03:11:51",
      "content": "<p>Would you mind sharing the modified function? Would make it easier to share results.</p>",
      "rawMarkdown": "Would you mind sharing the modified function? Would make it easier to share results.",
      "votes": null
    },
    {
      "id": "806710",
      "postDate": "04/14/2020 01:23:17",
      "content": "<p>Thanks for sharing! Wondering if training on this translated dataset helped? I tried training mBert on the concatenated version of translated version but I only get a score of <code>0.868</code> on the test set.</p>",
      "rawMarkdown": "Thanks for sharing! Wondering if training on this translated dataset helped? I tried training mBert on the concatenated version of translated version but I only get a score of `0.868` on the test set.",
      "votes": null
    },
    {
      "id": "806799",
      "postDate": "04/14/2020 04:40:41",
      "content": "<p><a href=\"/aroraaman\">@aroraaman</a>  There is a obvious difference between translated text and natural text, but it helped for me.</p>",
      "rawMarkdown": "aroraaman  There is a obvious difference between translated text and natural text, but it helped for me.",
      "votes": null
    },
    {
      "id": "806811",
      "postDate": "04/14/2020 05:06:12",
      "content": "<p>It’s not as simple as training a mBert or XLM-R on translated text right? I am having trouble being able to replicate any of the notebook scores that show 0.9389 but when I submit its 0.8582 or something. </p>\n\n<p>There are obvious smarter tricks involved I assume?</p>",
      "rawMarkdown": "It’s not as simple as training a mBert or XLM-R on translated text right? I am having trouble being able to replicate any of the notebook scores that show 0.9389 but when I submit its 0.8582 or something. \n\nThere are obvious smarter tricks involved I assume?",
      "votes": null
    },
    {
      "id": "808598",
      "postDate": "04/15/2020 14:30:06",
      "content": "<p>Hey <a href=\"/leecming\">@leecming</a> , did the multilingual training data help you to improve your score ?</p>",
      "rawMarkdown": "Hey @leecming , did the multilingual training data help you to improve your score ?",
      "votes": null
    },
    {
      "id": "808684",
      "postDate": "04/15/2020 14:59:00",
      "content": "<p><a href=\"/aroraaman\">@aroraaman</a> You mean you couldn't replicate the result even if you just re-run the public kernel? If so it's a bit strange..</p>",
      "rawMarkdown": "aroraaman You mean you couldn't replicate the result even if you just re-run the public kernel? If so it's a bit strange..",
      "votes": null
    },
    {
      "id": "811785",
      "postDate": "04/18/2020 08:23:56",
      "content": "<p>Great dataset, Michael ! \nAdding this dataset as it is to the training data did not improve my cv/PL scores.\nAccording to  <a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152\">comment</a>  \" ...gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify ...\"  and also <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521\">here:</a>      \"...all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.\"</p>\n\n<p>Do you see any ways to translate without losing toxicity in text ? Do you think it is feasible that some of us could build such \"unbiased\" translator in-house ?</p>\n\n<p>N.B. Just noticed alternative observation: according to <a href=\"/andypenrose\">@andypenrose</a>  <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521\">here</a> \"...it looks like toxicity does survive the automatic translation process...\"</p>",
      "rawMarkdown": "Great dataset, Michael ! \nAdding this dataset as it is to the training data did not improve my cv/PL scores.\nAccording to  @mgornergoogle [comment](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152)  \" ...gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify ...\"  and also [here:](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521)      \"...all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.\"\n\nDo you see any ways to translate without losing toxicity in text ? Do you think it is feasible that some of us could build such \"unbiased\" translator in-house ?\n\nN.B. Just noticed alternative observation: according to @andypenrose  [here](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521) \"...it looks like toxicity does survive the automatic translation process...\"",
      "votes": null
    },
    {
      "id": "812341",
      "postDate": "04/18/2020 17:12:02",
      "content": "<p>One more dataset, which may be useful for someone.\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-multilingual-swear-profanity\">[Jigsaw] Multilingual swear profanity</a></p>",
      "rawMarkdown": "One more dataset, which may be useful for someone.\n[[Jigsaw] Multilingual swear profanity](https://www.kaggle.com/miklgr500/jigsaw-multilingual-swear-profanity)",
      "votes": null
    },
    {
      "id": "812439",
      "postDate": "04/18/2020 18:47:26",
      "content": "<p><a href=\"/aroraaman\">@aroraaman</a> : maybe you cannot replicate the result due to the multi-cores TPU training, in that it can give you varying results due to the data sharding. You can refer to <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140706\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140706</a> </p>",
      "rawMarkdown": "aroraaman : maybe you cannot replicate the result due to the multi-cores TPU training, in that it can give you varying results due to the data sharding. You can refer to https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140706",
      "votes": null
    },
    {
      "id": "812589",
      "postDate": "04/18/2020 20:44:18",
      "content": "<p>I obtain a little better results on my models, but it needs the right multilingual data splitting.\nBut, I must agree with you, that translating a bit de-toxify comments. The real multilingual comments obtain specific words and patterns, which translation distorted. \n Because of it, I start to think about:\n1. Additional features, which can help save some patterns\n2. Scraped the real data from some websites (i.e. youtube...)</p>",
      "rawMarkdown": "I obtain a little better results on my models, but it needs the right multilingual data splitting.\nBut, I must agree with you, that translating a bit de-toxify comments. The real multilingual comments obtain specific words and patterns, which translation distorted. \n Because of it, I start to think about:\n1. Additional features, which can help save some patterns\n2. Scraped the real data from some websites (i.e. youtube...)",
      "votes": null
    },
    {
      "id": "813250",
      "postDate": "04/19/2020 14:11:51",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "819223",
      "postDate": "04/24/2020 12:30:49",
      "content": "<p>Hi, <a href=\"/miklgr500\">@miklgr500</a> thanks for your works and sharing. Do you also translate <code>jigsaw-unintended-bias-train.csv</code> set as well?</p>",
      "rawMarkdown": "Hi, @miklgr500 thanks for your works and sharing. Do you also translate `jigsaw-unintended-bias-train.csv` set as well?",
      "votes": null
    },
    {
      "id": "843703",
      "postDate": "05/12/2020 08:00:07",
      "content": "<p>yet to try this in comp .. But yeah for personal use ... Very useful </p>",
      "rawMarkdown": "yet to try this in comp .. But yeah for personal use ... Very useful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 798852,
      "author_name": "ma7555",
      "author_url": "",
      "post_date": "04/05/2020 23:13:21",
      "content": "<p>Hey Michael! Amazing work!\nThanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 798863,
      "author_name": "leecming",
      "author_url": "",
      "post_date": "04/05/2020 23:40:23",
      "content": "<p>I've only eyeballed the data but it looks like some of the translated comments aren't aligned with their IDs?</p>",
      "votes": null,
      "replies": [
        {
          "id": 799280,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/06/2020 10:35:06",
          "content": "<p>Yep, thank you. I need 3 hours to fix it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799202,
      "author_name": "rafiko1",
      "author_url": "",
      "post_date": "04/06/2020 09:10:02",
      "content": "<p>Nice job, thanks for sharing <a href=\"/miklgr500\">@miklgr500</a>! I am wondering which library did you use, <a href=\"https://github.com/terryyin/translate-python/\">translate-python </a> or <a href=\"https://github.com/ssut/py-googletrans\">googletrans</a>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 799511,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/06/2020 13:54:51",
          "content": "<p>I used <a href=\"https://pypi.org/project/translators/\">translators</a> library with some modification for accelerating translation and <a href=\"https://docs.python.org/3/library/multiprocessing.html\">multiprocessing</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804845,
          "author_name": "aroraaman",
          "author_url": "",
          "post_date": "04/12/2020 03:11:51",
          "content": "<p>Would you mind sharing the modified function? Would make it easier to share results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799436,
      "author_name": "haythemtellili5",
      "author_url": "",
      "post_date": "04/06/2020 12:52:08",
      "content": "<p>Hi Michael, a huge thanks for your effort during this competition .\nPlease did u fix the ID of comments and i wonder how much this dataset improve your score ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 799503,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/06/2020 13:49:36",
          "content": "<p><a href=\"/haythemtellili5\">@haythemtellili5</a> <a href=\"/leecming\">@leecming</a> the new version of the dataset is available.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799531,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "04/06/2020 14:20:26",
          "content": "<p>Thank, it'll be interesting to see how this shakes things up!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 808598,
          "author_name": "haythemtellili5",
          "author_url": "",
          "post_date": "04/15/2020 14:30:06",
          "content": "<p>Hey <a href=\"/leecming\">@leecming</a> , did the multilingual training data help you to improve your score ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799508,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/06/2020 13:53:12",
      "content": "<p>Wow, how did you overcome api limits? Awesome work</p>",
      "votes": null,
      "replies": [
        {
          "id": 799552,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/06/2020 14:33:33",
          "content": "<p>It feature of the<a href=\"https://pypi.org/project/translators/\"><code>translators</code></a> library.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799597,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/06/2020 15:03:44",
          "content": "<p>And which API did you use there?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799970,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/07/2020 00:05:17",
          "content": "<p>I tried translators for translate but I cannot translate all.(it does not include sleep).\nI used .google method.</p>\n\n<p>Did you use sleep or other waiting function?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800097,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/07/2020 04:49:35",
          "content": "<p><a href=\"https://www.kaggle.com/miklgr500/how-to-use-translators-for-comments-translation/log?scriptVersionId=31557415\">How to use \"translators\" for comments translation </a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800524,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/07/2020 13:40:08",
          "content": "<p>Thank you for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799642,
      "author_name": "sachinprabhu",
      "author_url": "",
      "post_date": "04/06/2020 15:52:53",
      "content": "<p>Nice job! Upvoting the post and the dataset as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 799894,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "04/06/2020 21:41:58",
      "content": "<p>Thanks for sharing.\nWhat is the difference between the cleaned and not clean version ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 800536,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/07/2020 13:54:39",
          "content": "<p>Cleaned files dosn't have NaN values. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 806710,
      "author_name": "aroraaman",
      "author_url": "",
      "post_date": "04/14/2020 01:23:17",
      "content": "<p>Thanks for sharing! Wondering if training on this translated dataset helped? I tried training mBert on the concatenated version of translated version but I only get a score of <code>0.868</code> on the test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 806799,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/14/2020 04:40:41",
          "content": "<p><a href=\"/aroraaman\">@aroraaman</a>  There is a obvious difference between translated text and natural text, but it helped for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 806811,
          "author_name": "aroraaman",
          "author_url": "",
          "post_date": "04/14/2020 05:06:12",
          "content": "<p>It’s not as simple as training a mBert or XLM-R on translated text right? I am having trouble being able to replicate any of the notebook scores that show 0.9389 but when I submit its 0.8582 or something. </p>\n\n<p>There are obvious smarter tricks involved I assume?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 808684,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/15/2020 14:59:00",
          "content": "<p><a href=\"/aroraaman\">@aroraaman</a> You mean you couldn't replicate the result even if you just re-run the public kernel? If so it's a bit strange..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 812439,
          "author_name": "glorian",
          "author_url": "",
          "post_date": "04/18/2020 18:47:26",
          "content": "<p><a href=\"/aroraaman\">@aroraaman</a> : maybe you cannot replicate the result due to the multi-cores TPU training, in that it can give you varying results due to the data sharding. You can refer to <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140706\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140706</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 811785,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "04/18/2020 08:23:56",
      "content": "<p>Great dataset, Michael ! \nAdding this dataset as it is to the training data did not improve my cv/PL scores.\nAccording to  <a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152\">comment</a>  \" ...gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify ...\"  and also <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521\">here:</a>      \"...all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.\"</p>\n\n<p>Do you see any ways to translate without losing toxicity in text ? Do you think it is feasible that some of us could build such \"unbiased\" translator in-house ?</p>\n\n<p>N.B. Just noticed alternative observation: according to <a href=\"/andypenrose\">@andypenrose</a>  <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521\">here</a> \"...it looks like toxicity does survive the automatic translation process...\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 812589,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "04/18/2020 20:44:18",
          "content": "<p>I obtain a little better results on my models, but it needs the right multilingual data splitting.\nBut, I must agree with you, that translating a bit de-toxify comments. The real multilingual comments obtain specific words and patterns, which translation distorted. \n Because of it, I start to think about:\n1. Additional features, which can help save some patterns\n2. Scraped the real data from some websites (i.e. youtube...)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 812341,
      "author_name": "miklgr500",
      "author_url": "",
      "post_date": "04/18/2020 17:12:02",
      "content": "<p>One more dataset, which may be useful for someone.\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-multilingual-swear-profanity\">[Jigsaw] Multilingual swear profanity</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 843703,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "05/12/2020 08:00:07",
          "content": "<p>yet to try this in comp .. But yeah for personal use ... Very useful </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 813250,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "04/19/2020 14:11:51",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819223,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "04/24/2020 12:30:49",
      "content": "<p>Hi, <a href=\"/miklgr500\">@miklgr500</a> thanks for your works and sharing. Do you also translate <code>jigsaw-unintended-bias-train.csv</code> set as well?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "798715": "I created a multilingual [train dataset](https://www.kaggle.com/miklgr500/jigsaw-train-multilingual-coments-google-api#jigsaw-toxic-comment-train-google-fr-cleaned.csv) using ```translates``` with Google API. And suggest sharing other variations of training dataset on this topic.",
    "798852": "Hey Michael! Amazing work!\nThanks for sharing",
    "798863": "I've only eyeballed the data but it looks like some of the translated comments aren't aligned with their IDs?",
    "799202": "Nice job, thanks for sharing @miklgr500! I am wondering which library did you use, [translate-python ](https://github.com/terryyin/translate-python/) or [googletrans](https://github.com/ssut/py-googletrans)?",
    "799280": "Yep, thank you. I need 3 hours to fix it.",
    "799436": "Hi Michael, a huge thanks for your effort during this competition .\nPlease did u fix the ID of comments and i wonder how much this dataset improve your score ?",
    "799503": "haythemtellili5 @leecming the new version of the dataset is available.",
    "799508": "Wow, how did you overcome api limits? Awesome work",
    "799511": "I used [translators](https://pypi.org/project/translators/) library with some modification for accelerating translation and [multiprocessing](https://docs.python.org/3/library/multiprocessing.html).",
    "799531": "Thank, it'll be interesting to see how this shakes things up!",
    "799552": "It feature of the[```translators```](https://pypi.org/project/translators/) library.",
    "799597": "And which API did you use there?",
    "799642": "Nice job! Upvoting the post and the dataset as well.",
    "799894": "Thanks for sharing.\nWhat is the difference between the cleaned and not clean version ?",
    "799970": "I tried translators for translate but I cannot translate all.(it does not include sleep).\nI used .google method.\n\nDid you use sleep or other waiting function?",
    "800097": "[How to use \"translators\" for comments translation ](https://www.kaggle.com/miklgr500/how-to-use-translators-for-comments-translation/log?scriptVersionId=31557415)",
    "800524": "Thank you for sharing!",
    "800536": "Cleaned files dosn't have NaN values.",
    "804845": "Would you mind sharing the modified function? Would make it easier to share results.",
    "806710": "Thanks for sharing! Wondering if training on this translated dataset helped? I tried training mBert on the concatenated version of translated version but I only get a score of `0.868` on the test set.",
    "806799": "aroraaman  There is a obvious difference between translated text and natural text, but it helped for me.",
    "806811": "It’s not as simple as training a mBert or XLM-R on translated text right? I am having trouble being able to replicate any of the notebook scores that show 0.9389 but when I submit its 0.8582 or something. \n\nThere are obvious smarter tricks involved I assume?",
    "808598": "Hey @leecming , did the multilingual training data help you to improve your score ?",
    "808684": "aroraaman You mean you couldn't replicate the result even if you just re-run the public kernel? If so it's a bit strange..",
    "811785": "Great dataset, Michael ! \nAdding this dataset as it is to the training data did not improve my cv/PL scores.\nAccording to  @mgornergoogle [comment](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152)  \" ...gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify ...\"  and also [here:](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521)      \"...all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.\"\n\nDo you see any ways to translate without losing toxicity in text ? Do you think it is feasible that some of us could build such \"unbiased\" translator in-house ?\n\nN.B. Just noticed alternative observation: according to @andypenrose  [here](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138671#786521) \"...it looks like toxicity does survive the automatic translation process...\"",
    "812341": "One more dataset, which may be useful for someone.\n[[Jigsaw] Multilingual swear profanity](https://www.kaggle.com/miklgr500/jigsaw-multilingual-swear-profanity)",
    "812439": "aroraaman : maybe you cannot replicate the result due to the multi-cores TPU training, in that it can give you varying results due to the data sharding. You can refer to https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/140706",
    "812589": "I obtain a little better results on my models, but it needs the right multilingual data splitting.\nBut, I must agree with you, that translating a bit de-toxify comments. The real multilingual comments obtain specific words and patterns, which translation distorted. \n Because of it, I start to think about:\n1. Additional features, which can help save some patterns\n2. Scraped the real data from some websites (i.e. youtube...)",
    "813250": "Thanks for sharing.",
    "819223": "Hi, @miklgr500 thanks for your works and sharing. Do you also translate `jigsaw-unintended-bias-train.csv` set as well?",
    "843703": "yet to try this in comp .. But yeah for personal use ... Very useful"
  },
  "source": "meta"
}