{
  "id": 138273,
  "title": "Is Translating Texts Allowed (or Viable) ?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138273",
  "author_name": "",
  "post_date": "2020-03-24T11:13:38.759736700Z",
  "votes": 23,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I am assuming the goal of the competition is to train on english data only, and see how they adapt to other languages. </p>\n\n<p>To quote the description page :</p>\n\n<p>&gt; This year, we're taking advantage of Kaggle's new TPU support and challenging you to build multilingual models with English-only training data.</p>\n\n<p>However, a straightforward approach to this problem is to translate texts in the test data to english. \nI have the feeling that this removes the point of the competition but perhaps this was discussed. Hopefully a clever approach will outperform this approach but still it looks like a strong baseline to me.</p>\n\n<p>Any thoughts ?</p>",
  "messages": [
    {
      "id": "784600",
      "postDate": "03/24/2020 11:13:38",
      "content": "<p>I am assuming the goal of the competition is to train on english data only, and see how they adapt to other languages. </p>\n\n<p>To quote the description page :</p>\n\n<p>&gt; This year, we're taking advantage of Kaggle's new TPU support and challenging you to build multilingual models with English-only training data.</p>\n\n<p>However, a straightforward approach to this problem is to translate texts in the test data to english. \nI have the feeling that this removes the point of the competition but perhaps this was discussed. Hopefully a clever approach will outperform this approach but still it looks like a strong baseline to me.</p>\n\n<p>Any thoughts ?</p>",
      "rawMarkdown": "I am assuming the goal of the competition is to train on english data only, and see how they adapt to other languages. \n\nTo quote the description page :\n\n&gt; This year, we're taking advantage of Kaggle's new TPU support and challenging you to build multilingual models with English-only training data.\n\nHowever, a straightforward approach to this problem is to translate texts in the test data to english. \nI have the feeling that this removes the point of the competition but perhaps this was discussed. Hopefully a clever approach will outperform this approach but still it looks like a strong baseline to me.\n\nAny thoughts ?",
      "votes": null
    },
    {
      "id": "785348",
      "postDate": "03/25/2020 01:33:10",
      "content": "<p>This is a fine approach and permitted as long as it’s used with the training/validation data. That said, you shouldn’t be using any hand labeled version of the test set in any way (i.e. as training or processing input). This includes versions of the test set that have been translated and hand labeled. As this would be violation of the rule against hand-labeling or hand annotation of the test set.</p>",
      "rawMarkdown": "This is a fine approach and permitted as long as it’s used with the training/validation data. That said, you shouldn’t be using any hand labeled version of the test set in any way (i.e. as training or processing input). This includes versions of the test set that have been translated and hand labeled. As this would be violation of the rule against hand-labeling or hand annotation of the test set.",
      "votes": null
    },
    {
      "id": "785353",
      "postDate": "03/25/2020 01:41:56",
      "content": "<p>What about the use of external data for translation as soon as posted in the Official External Data Thread?</p>\n\n<p>&gt; Per the Competition Rules, freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).</p>\n\n<p>By allowing the use of external data and with internet allowed because of TPUs, there are many ways to translate test data into English, either by doing it with any translation service and sharing the translated test data, or by translating the test set directly in the kernel with internet.</p>",
      "rawMarkdown": "What about the use of external data for translation as soon as posted in the Official External Data Thread?\n\n&gt; Per the Competition Rules, freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).\n\nBy allowing the use of external data and with internet allowed because of TPUs, there are many ways to translate test data into English, either by doing it with any translation service and sharing the translated test data, or by translating the test set directly in the kernel with internet.",
      "votes": null
    },
    {
      "id": "785449",
      "postDate": "03/25/2020 04:10:41",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thanks for the explanation, but it's a bit surprising and still unclear for me. \n1. What is the difference between normal preprocess or tokenization and translation? \n2. How about using language model to translate online and then tokenize?\n3. How do you judge and check if the solution is using translated test set or not?\n4. How about using translated test set only for unsupervised training such as pseudo labeling? Is it possible to detect this?</p>\n\n<p>If you truly want to keep test set as it is, you should have kept it as a private dataset like other 2 stage kernel competitions. (At first I thought so, but surprisingly all test set is disclosed.)\nWith this situation, it seems very difficult to prohibit to use translated test set and keep fairness because obviously it brings benefit to use translated test set and you cannot check all participant's solution (even if checked it's difficult to detect).</p>\n\n<p>To make this competition useful, it's better to have another test set and use it as a private test set, but it's too late and seems almost impossible. So, I suggest to allow translation to keep fairness. If not, I definitely believe someone who secretly violates the rule will finish in higher position.</p>",
      "rawMarkdown": "juliaelliott Thanks for the explanation, but it's a bit surprising and still unclear for me. \n1. What is the difference between normal preprocess or tokenization and translation? \n2. How about using language model to translate online and then tokenize?\n3. How do you judge and check if the solution is using translated test set or not?\n4. How about using translated test set only for unsupervised training such as pseudo labeling? Is it possible to detect this?\n\nIf you truly want to keep test set as it is, you should have kept it as a private dataset like other 2 stage kernel competitions. (At first I thought so, but surprisingly all test set is disclosed.)\nWith this situation, it seems very difficult to prohibit to use translated test set and keep fairness because obviously it brings benefit to use translated test set and you cannot check all participant's solution (even if checked it's difficult to detect).\n\nTo make this competition useful, it's better to have another test set and use it as a private test set, but it's too late and seems almost impossible. So, I suggest to allow translation to keep fairness. If not, I definitely believe someone who secretly violates the rule will finish in higher position.",
      "votes": null
    },
    {
      "id": "785640",
      "postDate": "03/25/2020 08:21:03",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> So does this mean that I am allowed to automatically translate the test data to English using an external API? I see no rule preventing this as it is not hand labeling. I am sure many people will do this and you need strict guidelines for exactly this issue.</p>\n\n<p>Actually, I just saw that whole test data is available, so I could even do this offline.</p>",
      "rawMarkdown": "juliaelliott So does this mean that I am allowed to automatically translate the test data to English using an external API? I see no rule preventing this as it is not hand labeling. I am sure many people will do this and you need strict guidelines for exactly this issue.\n\nActually, I just saw that whole test data is available, so I could even do this offline.",
      "votes": null
    },
    {
      "id": "785771",
      "postDate": "03/25/2020 11:19:14",
      "content": "<p>Competition should be renamed to\n<strong>Jigsaw Google Translate API Toxic Comment Classification</strong></p>",
      "rawMarkdown": "Competition should be renamed to\n**Jigsaw Google Translate API Toxic Comment Classification**",
      "votes": null
    },
    {
      "id": "785852",
      "postDate": "03/25/2020 12:58:56",
      "content": "<p>Or Jigsaw DeepL Toxic Comment Classification</p>",
      "rawMarkdown": "Or Jigsaw DeepL Toxic Comment Classification",
      "votes": null
    },
    {
      "id": "786062",
      "postDate": "03/25/2020 15:47:58",
      "content": "<p>We checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify. </p>",
      "rawMarkdown": "We checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify.",
      "votes": null
    },
    {
      "id": "786153",
      "postDate": "03/25/2020 16:59:23",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> It's interesting to know that host assumes translation doesn't help! <br>\nI just want to double check, does that mean it's OK to use any translation approach such as google translate api to test set?</p>",
      "rawMarkdown": "mgornergoogle It's interesting to know that host assumes translation doesn't help!  \nI just want to double check, does that mean it's OK to use any translation approach such as google translate api to test set?",
      "votes": null
    },
    {
      "id": "786229",
      "postDate": "03/25/2020 18:05:44",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the answer :)</p>",
      "rawMarkdown": "mgornergoogle Thanks for the answer :)",
      "votes": null
    },
    {
      "id": "786324",
      "postDate": "03/25/2020 19:59:45",
      "content": "<p>Seems worth a try. One thing that might make this difficult is that the test and validation sets have a different distribution and subset of languages, as shown below. For example, the test set is 17% Russian and the validation set is 0% Russian. If toxicity translates differently for different languages, there could be a significant difference in performance on the validation versus test set, right? </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2280345%2Fbf8ed2c23e24a6d0431e68478fc4ab46%2FCapture.PNG?generation=1585166249933594&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Seems worth a try. One thing that might make this difficult is that the test and validation sets have a different distribution and subset of languages, as shown below. For example, the test set is 17% Russian and the validation set is 0% Russian. If toxicity translates differently for different languages, there could be a significant difference in performance on the validation versus test set, right? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2280345%2Fbf8ed2c23e24a6d0431e68478fc4ab46%2FCapture.PNG?generation=1585166249933594&amp;alt=media)",
      "votes": null
    },
    {
      "id": "787593",
      "postDate": "03/26/2020 22:51:15",
      "content": "<p>Yes <a href=\"/bamps53\">@bamps53</a> especially DeepL, but DeepL has a limited set of languages, also I don't know if it has an API.</p>",
      "rawMarkdown": "Yes @bamps53 especially DeepL, but DeepL has a limited set of languages, also I don't know if it has an API.",
      "votes": null
    },
    {
      "id": "797747",
      "postDate": "04/04/2020 21:29:09",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "803947",
      "postDate": "04/11/2020 03:35:31",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/mgornergoogle\">@mgornergoogle</a> Hi Julia/Martin, as also mentioned by others previously, it's still not clear whether it's ok to predict on machine(for example Google API)-translated test data? Would you please help clarify? Thanks.</p>",
      "rawMarkdown": "juliaelliott @mgornergoogle Hi Julia/Martin, as also mentioned by others previously, it's still not clear whether it's ok to predict on machine(for example Google API)-translated test data? Would you please help clarify? Thanks.",
      "votes": null
    },
    {
      "id": "895613",
      "postDate": "06/21/2020 13:53:03",
      "content": "<p>It has via its pro version: <a href=\"https://www.deepl.com/en/docs-api/\">https://www.deepl.com/en/docs-api/</a>. It isn't free but I have noticed that DeepL translation quality is often better than Google. </p>",
      "rawMarkdown": "It has via its pro version: https://www.deepl.com/en/docs-api/. It isn't free but I have noticed that DeepL translation quality is often better than Google.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 785348,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "03/25/2020 01:33:10",
      "content": "<p>This is a fine approach and permitted as long as it’s used with the training/validation data. That said, you shouldn’t be using any hand labeled version of the test set in any way (i.e. as training or processing input). This includes versions of the test set that have been translated and hand labeled. As this would be violation of the rule against hand-labeling or hand annotation of the test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 785353,
          "author_name": "mika30",
          "author_url": "",
          "post_date": "03/25/2020 01:41:56",
          "content": "<p>What about the use of external data for translation as soon as posted in the Official External Data Thread?</p>\n\n<p>&gt; Per the Competition Rules, freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).</p>\n\n<p>By allowing the use of external data and with internet allowed because of TPUs, there are many ways to translate test data into English, either by doing it with any translation service and sharing the translated test data, or by translating the test set directly in the kernel with internet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785449,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "03/25/2020 04:10:41",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thanks for the explanation, but it's a bit surprising and still unclear for me. \n1. What is the difference between normal preprocess or tokenization and translation? \n2. How about using language model to translate online and then tokenize?\n3. How do you judge and check if the solution is using translated test set or not?\n4. How about using translated test set only for unsupervised training such as pseudo labeling? Is it possible to detect this?</p>\n\n<p>If you truly want to keep test set as it is, you should have kept it as a private dataset like other 2 stage kernel competitions. (At first I thought so, but surprisingly all test set is disclosed.)\nWith this situation, it seems very difficult to prohibit to use translated test set and keep fairness because obviously it brings benefit to use translated test set and you cannot check all participant's solution (even if checked it's difficult to detect).</p>\n\n<p>To make this competition useful, it's better to have another test set and use it as a private test set, but it's too late and seems almost impossible. So, I suggest to allow translation to keep fairness. If not, I definitely believe someone who secretly violates the rule will finish in higher position.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785640,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/25/2020 08:21:03",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> So does this mean that I am allowed to automatically translate the test data to English using an external API? I see no rule preventing this as it is not hand labeling. I am sure many people will do this and you need strict guidelines for exactly this issue.</p>\n\n<p>Actually, I just saw that whole test data is available, so I could even do this offline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 797747,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "04/04/2020 21:29:09",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 803947,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/11/2020 03:35:31",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/mgornergoogle\">@mgornergoogle</a> Hi Julia/Martin, as also mentioned by others previously, it's still not clear whether it's ok to predict on machine(for example Google API)-translated test data? Would you please help clarify? Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 785771,
      "author_name": "maxjeblick",
      "author_url": "",
      "post_date": "03/25/2020 11:19:14",
      "content": "<p>Competition should be renamed to\n<strong>Jigsaw Google Translate API Toxic Comment Classification</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 785852,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "03/25/2020 12:58:56",
          "content": "<p>Or Jigsaw DeepL Toxic Comment Classification</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787593,
          "author_name": "hamditarek",
          "author_url": "",
          "post_date": "03/26/2020 22:51:15",
          "content": "<p>Yes <a href=\"/bamps53\">@bamps53</a> especially DeepL, but DeepL has a limited set of languages, also I don't know if it has an API.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 895613,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/21/2020 13:53:03",
          "content": "<p>It has via its pro version: <a href=\"https://www.deepl.com/en/docs-api/\">https://www.deepl.com/en/docs-api/</a>. It isn't free but I have noticed that DeepL translation quality is often better than Google. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 786062,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/25/2020 15:47:58",
      "content": "<p>We checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify. </p>",
      "votes": null,
      "replies": [
        {
          "id": 786153,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "03/25/2020 16:59:23",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> It's interesting to know that host assumes translation doesn't help! <br>\nI just want to double check, does that mean it's OK to use any translation approach such as google translate api to test set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 786229,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/25/2020 18:05:44",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the answer :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 786324,
          "author_name": "davelo12",
          "author_url": "",
          "post_date": "03/25/2020 19:59:45",
          "content": "<p>Seems worth a try. One thing that might make this difficult is that the test and validation sets have a different distribution and subset of languages, as shown below. For example, the test set is 17% Russian and the validation set is 0% Russian. If toxicity translates differently for different languages, there could be a significant difference in performance on the validation versus test set, right? </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2280345%2Fbf8ed2c23e24a6d0431e68478fc4ab46%2FCapture.PNG?generation=1585166249933594&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "784600": "I am assuming the goal of the competition is to train on english data only, and see how they adapt to other languages. \n\nTo quote the description page :\n\n&gt; This year, we're taking advantage of Kaggle's new TPU support and challenging you to build multilingual models with English-only training data.\n\nHowever, a straightforward approach to this problem is to translate texts in the test data to english. \nI have the feeling that this removes the point of the competition but perhaps this was discussed. Hopefully a clever approach will outperform this approach but still it looks like a strong baseline to me.\n\nAny thoughts ?",
    "785348": "This is a fine approach and permitted as long as it’s used with the training/validation data. That said, you shouldn’t be using any hand labeled version of the test set in any way (i.e. as training or processing input). This includes versions of the test set that have been translated and hand labeled. As this would be violation of the rule against hand-labeling or hand annotation of the test set.",
    "785353": "What about the use of external data for translation as soon as posted in the Official External Data Thread?\n\n&gt; Per the Competition Rules, freely and publicly available external data is permitted in this competition, but must be posted to this forum thread no later than the Entry Deadline (one week before competition close).\n\nBy allowing the use of external data and with internet allowed because of TPUs, there are many ways to translate test data into English, either by doing it with any translation service and sharing the translated test data, or by translating the test set directly in the kernel with internet.",
    "785449": "juliaelliott Thanks for the explanation, but it's a bit surprising and still unclear for me. \n1. What is the difference between normal preprocess or tokenization and translation? \n2. How about using language model to translate online and then tokenize?\n3. How do you judge and check if the solution is using translated test set or not?\n4. How about using translated test set only for unsupervised training such as pseudo labeling? Is it possible to detect this?\n\nIf you truly want to keep test set as it is, you should have kept it as a private dataset like other 2 stage kernel competitions. (At first I thought so, but surprisingly all test set is disclosed.)\nWith this situation, it seems very difficult to prohibit to use translated test set and keep fairness because obviously it brings benefit to use translated test set and you cannot check all participant's solution (even if checked it's difficult to detect).\n\nTo make this competition useful, it's better to have another test set and use it as a private test set, but it's too late and seems almost impossible. So, I suggest to allow translation to keep fairness. If not, I definitely believe someone who secretly violates the rule will finish in higher position.",
    "785640": "juliaelliott So does this mean that I am allowed to automatically translate the test data to English using an external API? I see no rule preventing this as it is not hand labeling. I am sure many people will do this and you need strict guidelines for exactly this issue.\n\nActually, I just saw that whole test data is available, so I could even do this offline.",
    "785771": "Competition should be renamed to\n**Jigsaw Google Translate API Toxic Comment Classification**",
    "785852": "Or Jigsaw DeepL Toxic Comment Classification",
    "786062": "We checked with the Jigsaw team. They are interested in seeing if there is something useful that can be done through automatic translations. Their gut feeling however is that it will not be a very useful approach because automatic translations tend to de-toxify.",
    "786153": "mgornergoogle It's interesting to know that host assumes translation doesn't help!  \nI just want to double check, does that mean it's OK to use any translation approach such as google translate api to test set?",
    "786229": "mgornergoogle Thanks for the answer :)",
    "786324": "Seems worth a try. One thing that might make this difficult is that the test and validation sets have a different distribution and subset of languages, as shown below. For example, the test set is 17% Russian and the validation set is 0% Russian. If toxicity translates differently for different languages, there could be a significant difference in performance on the validation versus test set, right? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2280345%2Fbf8ed2c23e24a6d0431e68478fc4ab46%2FCapture.PNG?generation=1585166249933594&amp;alt=media)",
    "787593": "Yes @bamps53 especially DeepL, but DeepL has a limited set of languages, also I don't know if it has an API.",
    "797747": "",
    "803947": "juliaelliott @mgornergoogle Hi Julia/Martin, as also mentioned by others previously, it's still not clear whether it's ok to predict on machine(for example Google API)-translated test data? Would you please help clarify? Thanks.",
    "895613": "It has via its pro version: https://www.deepl.com/en/docs-api/. It isn't free but I have noticed that DeepL translation quality is often better than Google."
  },
  "source": "meta"
}