{
  "id": 138671,
  "title": "Validation with translated text",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138671",
  "author_name": "",
  "post_date": "2020-03-25T22:27:55.100609400Z",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Now it's officially allowed to use translated test set.(see <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152\">this post</a>)  </p>\n\n<p>Even though test set is now translated to English, it looks still different from original English text. To evaluate your model performance, it makes more sense if you use translated validation set which is translated by same way as test set.(google translate api)\nSo, I also created translated validation set and notebook as a example.\nPlease check it out!\n- <a href=\"https://www.kaggle.com/bamps53/pytorch-tpu-training-with-translated-validation?scriptVersionId=30841456\">translated valid notebook</a>\n- <a href=\"https://www.kaggle.com/bamps53/val-en-df\">translated valid set</a>\n- <a href=\"https://www.kaggle.com/bamps53/test-en-df\">translated test set</a></p>",
  "messages": [
    {
      "id": "786440",
      "postDate": "03/25/2020 22:27:55",
      "content": "<p>Now it's officially allowed to use translated test set.(see <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152\">this post</a>)  </p>\n\n<p>Even though test set is now translated to English, it looks still different from original English text. To evaluate your model performance, it makes more sense if you use translated validation set which is translated by same way as test set.(google translate api)\nSo, I also created translated validation set and notebook as a example.\nPlease check it out!\n- <a href=\"https://www.kaggle.com/bamps53/pytorch-tpu-training-with-translated-validation?scriptVersionId=30841456\">translated valid notebook</a>\n- <a href=\"https://www.kaggle.com/bamps53/val-en-df\">translated valid set</a>\n- <a href=\"https://www.kaggle.com/bamps53/test-en-df\">translated test set</a></p>",
      "rawMarkdown": "Now it's officially allowed to use translated test set.(see [this post](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152))  \n\nEven though test set is now translated to English, it looks still different from original English text. To evaluate your model performance, it makes more sense if you use translated validation set which is translated by same way as test set.(google translate api)\nSo, I also created translated validation set and notebook as a example.\nPlease check it out!\n- [translated valid notebook](https://www.kaggle.com/bamps53/pytorch-tpu-training-with-translated-validation?scriptVersionId=30841456)\n- [translated valid set](https://www.kaggle.com/bamps53/val-en-df)\n- [translated test set](https://www.kaggle.com/bamps53/test-en-df)",
      "votes": null
    },
    {
      "id": "786521",
      "postDate": "03/26/2020 00:29:39",
      "content": "<p>With this validation set based on translations, you might be optimizing your model for the data + the biases of the translation engine. I'm pretty sure all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.</p>",
      "rawMarkdown": "With this validation set based on translations, you might be optimizing your model for the data + the biases of the translation engine. I'm pretty sure all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.",
      "votes": null
    },
    {
      "id": "786539",
      "postDate": "03/26/2020 00:59:33",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Yes, I know what you said. I'll look for another creative solution other than straight translation.  But so far it's better to use translated test set and I guess most of the current top score is using translated test set. So, I just wanted to share the way to measure the performance on translated test set. I hope this won't be the best way to deal with multilingual text classification.</p>",
      "rawMarkdown": "mgornergoogle Yes, I know what you said. I'll look for another creative solution other than straight translation.  But so far it's better to use translated test set and I guess most of the current top score is using translated test set. So, I just wanted to share the way to measure the performance on translated test set. I hope this won't be the best way to deal with multilingual text classification.",
      "votes": null
    },
    {
      "id": "787262",
      "postDate": "03/26/2020 16:45:51",
      "content": "<p>That makes sense. Although, experimenting with Abhishek's latest notebook, <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation\">https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation</a>, I find that with the same multilingual model, the translated text scores very slightly better on the LB than the original text. So it looks like toxicity does survive the automatic translation process. A quick test with Google Translate also suggests profanity in -&gt; profanity out.</p>\n\n<p>PS Thanks to <a href=\"/abhishek\">@abhishek</a> for sharing the notebook and others for sharing the translated data.</p>",
      "rawMarkdown": "That makes sense. Although, experimenting with Abhishek's latest notebook, https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation, I find that with the same multilingual model, the translated text scores very slightly better on the LB than the original text. So it looks like toxicity does survive the automatic translation process. A quick test with Google Translate also suggests profanity in -&gt; profanity out.\n\nPS Thanks to @abhishek for sharing the notebook and others for sharing the translated data.",
      "votes": null
    },
    {
      "id": "787585",
      "postDate": "03/26/2020 22:39:17",
      "content": "<p>Thanks for sharing! <a href=\"/bamps53\">@bamps53</a> </p>",
      "rawMarkdown": "Thanks for sharing! @bamps53",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 786521,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/26/2020 00:29:39",
      "content": "<p>With this validation set based on translations, you might be optimizing your model for the data + the biases of the translation engine. I'm pretty sure all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.</p>",
      "votes": null,
      "replies": [
        {
          "id": 786539,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "03/26/2020 00:59:33",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Yes, I know what you said. I'll look for another creative solution other than straight translation.  But so far it's better to use translated test set and I guess most of the current top score is using translated test set. So, I just wanted to share the way to measure the performance on translated test set. I hope this won't be the best way to deal with multilingual text classification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787262,
          "author_name": "andypenrose",
          "author_url": "",
          "post_date": "03/26/2020 16:45:51",
          "content": "<p>That makes sense. Although, experimenting with Abhishek's latest notebook, <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation\">https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation</a>, I find that with the same multilingual model, the translated text scores very slightly better on the LB than the original text. So it looks like toxicity does survive the automatic translation process. A quick test with Google Translate also suggests profanity in -&gt; profanity out.</p>\n\n<p>PS Thanks to <a href=\"/abhishek\">@abhishek</a> for sharing the notebook and others for sharing the translated data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 787585,
      "author_name": "hamditarek",
      "author_url": "",
      "post_date": "03/26/2020 22:39:17",
      "content": "<p>Thanks for sharing! <a href=\"/bamps53\">@bamps53</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "786440": "Now it's officially allowed to use translated test set.(see [this post](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138576#786152))  \n\nEven though test set is now translated to English, it looks still different from original English text. To evaluate your model performance, it makes more sense if you use translated validation set which is translated by same way as test set.(google translate api)\nSo, I also created translated validation set and notebook as a example.\nPlease check it out!\n- [translated valid notebook](https://www.kaggle.com/bamps53/pytorch-tpu-training-with-translated-validation?scriptVersionId=30841456)\n- [translated valid set](https://www.kaggle.com/bamps53/val-en-df)\n- [translated test set](https://www.kaggle.com/bamps53/test-en-df)",
    "786521": "With this validation set based on translations, you might be optimizing your model for the data + the biases of the translation engine. I'm pretty sure all public translation engines are biased against strong toxic language. It would be bad if Google translate was outputting profanity-laden sentences without a very good reason.",
    "786539": "mgornergoogle Yes, I know what you said. I'll look for another creative solution other than straight translation.  But so far it's better to use translated test set and I guess most of the current top score is using translated test set. So, I just wanted to share the way to measure the performance on translated test set. I hope this won't be the best way to deal with multilingual text classification.",
    "787262": "That makes sense. Although, experimenting with Abhishek's latest notebook, https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation, I find that with the same multilingual model, the translated text scores very slightly better on the LB than the original text. So it looks like toxicity does survive the automatic translation process. A quick test with Google Translate also suggests profanity in -&gt; profanity out.\n\nPS Thanks to @abhishek for sharing the notebook and others for sharing the translated data.",
    "787585": "Thanks for sharing! @bamps53"
  },
  "source": "meta"
}