{
  "id": 141629,
  "title": "Training dataset translated to 6 languages (Microsoft API)",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141629",
  "author_name": "",
  "post_date": "2020-04-06T22:25:53.826807400Z",
  "votes": 25,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have created a dataset with Microsoft translation for the training dataset into:\n1. tr\n2. ru\n3. it\n4. fr\n5. pt\n6. es</p>\n\n<p><a href=\"https://www.kaggle.com/ma7555/jigsaw-train-translated\">https://www.kaggle.com/ma7555/jigsaw-train-translated</a></p>\n\n<p>Please note that due to API limits and cost (used only free credit), I had to skip long comments and didn't get to finish the whole dataset. You would find some NaNs which you would need to take care of.</p>\n\n<p>A snippet to load it quickly and disregard the NaNs:</p>\n\n<p>```\ntrain1 = pd.read_csv(\"/kaggle/input/jigsaw-train-translated/train_mic.csv\")\ntrain2 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-unintended-bias-train.csv\")\ntrain2.toxic = train2.toxic.round().astype(int)</p>\n\n<p>train = pd.concat([\n    train1[['comment_text', 'toxic']],\n    train1[['tr', 'toxic']].rename(columns={'tr': 'comment_text'}).dropna(),\n    train1[['ru', 'toxic']].rename(columns={'ru': 'comment_text'}).dropna(),\n    train1[['it', 'toxic']].rename(columns={'it': 'comment_text'}).dropna(),\n    train1[['fr', 'toxic']].rename(columns={'fr': 'comment_text'}).dropna(),\n    train1[['pt', 'toxic']].rename(columns={'pt': 'comment_text'}).dropna(),\n    train1[['es', 'toxic']].rename(columns={'es': 'comment_text'}).dropna(),\n    train2[['comment_text', 'toxic']].query('toxic==1'),\n    train2[['comment_text', 'toxic']].query('toxic==0').sample(n=100000, random_state=0)\n]).reset_index(drop=True)</p>\n\n<p>y_train = train.toxic.values\n```</p>\n\n<p>I might push more updates but for the time being, feel free to use what has already been translated. </p>\n\n<p>I am willing to share the code I used to translate with microsoft API if someone is willing to continue translating and add him as a colaborator in this dataset!</p>",
  "messages": [
    {
      "id": "799910",
      "postDate": "04/06/2020 22:25:53",
      "content": "<p>I have created a dataset with Microsoft translation for the training dataset into:\n1. tr\n2. ru\n3. it\n4. fr\n5. pt\n6. es</p>\n\n<p><a href=\"https://www.kaggle.com/ma7555/jigsaw-train-translated\">https://www.kaggle.com/ma7555/jigsaw-train-translated</a></p>\n\n<p>Please note that due to API limits and cost (used only free credit), I had to skip long comments and didn't get to finish the whole dataset. You would find some NaNs which you would need to take care of.</p>\n\n<p>A snippet to load it quickly and disregard the NaNs:</p>\n\n<p>```\ntrain1 = pd.read_csv(\"/kaggle/input/jigsaw-train-translated/train_mic.csv\")\ntrain2 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-unintended-bias-train.csv\")\ntrain2.toxic = train2.toxic.round().astype(int)</p>\n\n<p>train = pd.concat([\n    train1[['comment_text', 'toxic']],\n    train1[['tr', 'toxic']].rename(columns={'tr': 'comment_text'}).dropna(),\n    train1[['ru', 'toxic']].rename(columns={'ru': 'comment_text'}).dropna(),\n    train1[['it', 'toxic']].rename(columns={'it': 'comment_text'}).dropna(),\n    train1[['fr', 'toxic']].rename(columns={'fr': 'comment_text'}).dropna(),\n    train1[['pt', 'toxic']].rename(columns={'pt': 'comment_text'}).dropna(),\n    train1[['es', 'toxic']].rename(columns={'es': 'comment_text'}).dropna(),\n    train2[['comment_text', 'toxic']].query('toxic==1'),\n    train2[['comment_text', 'toxic']].query('toxic==0').sample(n=100000, random_state=0)\n]).reset_index(drop=True)</p>\n\n<p>y_train = train.toxic.values\n```</p>\n\n<p>I might push more updates but for the time being, feel free to use what has already been translated. </p>\n\n<p>I am willing to share the code I used to translate with microsoft API if someone is willing to continue translating and add him as a colaborator in this dataset!</p>",
      "rawMarkdown": "I have created a dataset with Microsoft translation for the training dataset into:\n1. tr\n2. ru\n3. it\n4. fr\n5. pt\n6. es\n\nhttps://www.kaggle.com/ma7555/jigsaw-train-translated\n\nPlease note that due to API limits and cost (used only free credit), I had to skip long comments and didn't get to finish the whole dataset. You would find some NaNs which you would need to take care of.\n\nA snippet to load it quickly and disregard the NaNs:\n\n```\ntrain1 = pd.read_csv(\"/kaggle/input/jigsaw-train-translated/train_mic.csv\")\ntrain2 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-unintended-bias-train.csv\")\ntrain2.toxic = train2.toxic.round().astype(int)\n\ntrain = pd.concat([\n    train1[['comment_text', 'toxic']],\n    train1[['tr', 'toxic']].rename(columns={'tr': 'comment_text'}).dropna(),\n    train1[['ru', 'toxic']].rename(columns={'ru': 'comment_text'}).dropna(),\n    train1[['it', 'toxic']].rename(columns={'it': 'comment_text'}).dropna(),\n    train1[['fr', 'toxic']].rename(columns={'fr': 'comment_text'}).dropna(),\n    train1[['pt', 'toxic']].rename(columns={'pt': 'comment_text'}).dropna(),\n    train1[['es', 'toxic']].rename(columns={'es': 'comment_text'}).dropna(),\n    train2[['comment_text', 'toxic']].query('toxic==1'),\n    train2[['comment_text', 'toxic']].query('toxic==0').sample(n=100000, random_state=0)\n]).reset_index(drop=True)\n\ny_train = train.toxic.values\n```\n\nI might push more updates but for the time being, feel free to use what has already been translated. \n\nI am willing to share the code I used to translate with microsoft API if someone is willing to continue translating and add him as a colaborator in this dataset!",
      "votes": null
    },
    {
      "id": "1541201",
      "postDate": "10/11/2021 09:22:11",
      "content": "<p>I am thinking to translate it to the German language but not sure if it will have the same meaning as the original languag.</p>",
      "rawMarkdown": "I am thinking to translate it to the German language but not sure if it will have the same meaning as the original languag.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1541201,
      "author_name": "shubheshswain",
      "author_url": "",
      "post_date": "10/11/2021 09:22:11",
      "content": "<p>I am thinking to translate it to the German language but not sure if it will have the same meaning as the original languag.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "799910": "I have created a dataset with Microsoft translation for the training dataset into:\n1. tr\n2. ru\n3. it\n4. fr\n5. pt\n6. es\n\nhttps://www.kaggle.com/ma7555/jigsaw-train-translated\n\nPlease note that due to API limits and cost (used only free credit), I had to skip long comments and didn't get to finish the whole dataset. You would find some NaNs which you would need to take care of.\n\nA snippet to load it quickly and disregard the NaNs:\n\n```\ntrain1 = pd.read_csv(\"/kaggle/input/jigsaw-train-translated/train_mic.csv\")\ntrain2 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-unintended-bias-train.csv\")\ntrain2.toxic = train2.toxic.round().astype(int)\n\ntrain = pd.concat([\n    train1[['comment_text', 'toxic']],\n    train1[['tr', 'toxic']].rename(columns={'tr': 'comment_text'}).dropna(),\n    train1[['ru', 'toxic']].rename(columns={'ru': 'comment_text'}).dropna(),\n    train1[['it', 'toxic']].rename(columns={'it': 'comment_text'}).dropna(),\n    train1[['fr', 'toxic']].rename(columns={'fr': 'comment_text'}).dropna(),\n    train1[['pt', 'toxic']].rename(columns={'pt': 'comment_text'}).dropna(),\n    train1[['es', 'toxic']].rename(columns={'es': 'comment_text'}).dropna(),\n    train2[['comment_text', 'toxic']].query('toxic==1'),\n    train2[['comment_text', 'toxic']].query('toxic==0').sample(n=100000, random_state=0)\n]).reset_index(drop=True)\n\ny_train = train.toxic.values\n```\n\nI might push more updates but for the time being, feel free to use what has already been translated. \n\nI am willing to share the code I used to translate with microsoft API if someone is willing to continue translating and add him as a colaborator in this dataset!",
    "1541201": "I am thinking to translate it to the German language but not sure if it will have the same meaning as the original languag."
  },
  "source": "meta"
}