{
  "id": 160951,
  "title": "What is human-level performance for this problem?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160951",
  "author_name": "",
  "post_date": "2020-06-23T08:35:34.970041500Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>It'd be interesting to see how \"naive\" annotations by a group of native speakers (unaffliated with the original annotators and without using the same annotation methodology) would score as a baseline to compare ML models with.</p>\n\n<p>There was a fair bit of concern about the potential for cheating but looking at the validation labels - quite a few samples were noisily labelled. I would be surprised if humans would perform better than a ML model. </p>",
  "messages": [
    {
      "id": "898031",
      "postDate": "06/23/2020 08:35:34",
      "content": "<p>It'd be interesting to see how \"naive\" annotations by a group of native speakers (unaffliated with the original annotators and without using the same annotation methodology) would score as a baseline to compare ML models with.</p>\n\n<p>There was a fair bit of concern about the potential for cheating but looking at the validation labels - quite a few samples were noisily labelled. I would be surprised if humans would perform better than a ML model. </p>",
      "rawMarkdown": "It'd be interesting to see how \"naive\" annotations by a group of native speakers (unaffliated with the original annotators and without using the same annotation methodology) would score as a baseline to compare ML models with.\n\nThere was a fair bit of concern about the potential for cheating but looking at the validation labels - quite a few samples were noisily labelled. I would be surprised if humans would perform better than a ML model.",
      "votes": null
    },
    {
      "id": "898046",
      "postDate": "06/23/2020 08:44:20",
      "content": "<p>IMHO, it is quite difficult for such a task to assess human-level performance. Toxicity can be subjective sometimes depending on the context and the general tone of the conversation. Some humans might find one text toxic while others not, it depends on their perception of irony, cynicism... \nTo be honest, I find quite scary the perspective of letting a ML algorithm decide which messages are toxic. Humans are already highly imperfect. I guess it is pretty hard to find a commonly adopted benchmark for human performance.</p>",
      "rawMarkdown": "IMHO, it is quite difficult for such a task to assess human-level performance. Toxicity can be subjective sometimes depending on the context and the general tone of the conversation. Some humans might find one text toxic while others not, it depends on their perception of irony, cynicism... \nTo be honest, I find quite scary the perspective of letting a ML algorithm decide which messages are toxic. Humans are already highly imperfect. I guess it is pretty hard to find a commonly adopted benchmark for human performance.",
      "votes": null
    },
    {
      "id": "898053",
      "postDate": "06/23/2020 08:49:51",
      "content": "<p>Yeah, there was clearly noise in the labels. I wonder if they used mechanical turk or some other external labeling service. The definition they give us for toxic is quite vague and they didnt give us labeling methodology like they did in the last competition. </p>",
      "rawMarkdown": "Yeah, there was clearly noise in the labels. I wonder if they used mechanical turk or some other external labeling service. The definition they give us for toxic is quite vague and they didnt give us labeling methodology like they did in the last competition.",
      "votes": null
    },
    {
      "id": "898071",
      "postDate": "06/23/2020 09:11:42",
      "content": "<p>Then there will be Unintended Bias in Multilingual Toxic challenge comp next year ;) </p>",
      "rawMarkdown": "Then there will be Unintended Bias in Multilingual Toxic challenge comp next year ;)",
      "votes": null
    },
    {
      "id": "898159",
      "postDate": "06/23/2020 10:50:42",
      "content": "<p><a href=\"/leecming\">@leecming</a> I speak/read pt, es and it (to some level at least)\nI had a look at the validation data and the samples with worst prediction errors, where my model was supposedly completely wrong, and to be honest, I often would agree with my model better than of the annotated labels! 😆 </p>\n\n<p>Some of the labels seem to be automated, e.g. toxic if contains certain words, independent of context.\nHowever, in Italian some weren't tagged toxic, as curse words can be culturally accepted in non-toxic text (apparently some people used a hardcoded negative scale for Italian and Pt).</p>\n\n<p>In short, I'd do a very bad job ;)</p>",
      "rawMarkdown": "leecming I speak/read pt, es and it (to some level at least)\nI had a look at the validation data and the samples with worst prediction errors, where my model was supposedly completely wrong, and to be honest, I often would agree with my model better than of the annotated labels! 😆 \n\nSome of the labels seem to be automated, e.g. toxic if contains certain words, independent of context.\nHowever, in Italian some weren't tagged toxic, as curse words can be culturally accepted in non-toxic text (apparently some people used a hardcoded negative scale for Italian and Pt).\n\nIn short, I'd do a very bad job ;)",
      "votes": null
    },
    {
      "id": "898552",
      "postDate": "06/23/2020 15:29:24",
      "content": "<p>Hahaha true</p>",
      "rawMarkdown": "Hahaha true",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 898046,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/23/2020 08:44:20",
      "content": "<p>IMHO, it is quite difficult for such a task to assess human-level performance. Toxicity can be subjective sometimes depending on the context and the general tone of the conversation. Some humans might find one text toxic while others not, it depends on their perception of irony, cynicism... \nTo be honest, I find quite scary the perspective of letting a ML algorithm decide which messages are toxic. Humans are already highly imperfect. I guess it is pretty hard to find a commonly adopted benchmark for human performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 898071,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "06/23/2020 09:11:42",
          "content": "<p>Then there will be Unintended Bias in Multilingual Toxic challenge comp next year ;) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898552,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "06/23/2020 15:29:24",
          "content": "<p>Hahaha true</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898053,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "06/23/2020 08:49:51",
      "content": "<p>Yeah, there was clearly noise in the labels. I wonder if they used mechanical turk or some other external labeling service. The definition they give us for toxic is quite vague and they didnt give us labeling methodology like they did in the last competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898159,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "06/23/2020 10:50:42",
      "content": "<p><a href=\"/leecming\">@leecming</a> I speak/read pt, es and it (to some level at least)\nI had a look at the validation data and the samples with worst prediction errors, where my model was supposedly completely wrong, and to be honest, I often would agree with my model better than of the annotated labels! 😆 </p>\n\n<p>Some of the labels seem to be automated, e.g. toxic if contains certain words, independent of context.\nHowever, in Italian some weren't tagged toxic, as curse words can be culturally accepted in non-toxic text (apparently some people used a hardcoded negative scale for Italian and Pt).</p>\n\n<p>In short, I'd do a very bad job ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "898031": "It'd be interesting to see how \"naive\" annotations by a group of native speakers (unaffliated with the original annotators and without using the same annotation methodology) would score as a baseline to compare ML models with.\n\nThere was a fair bit of concern about the potential for cheating but looking at the validation labels - quite a few samples were noisily labelled. I would be surprised if humans would perform better than a ML model.",
    "898046": "IMHO, it is quite difficult for such a task to assess human-level performance. Toxicity can be subjective sometimes depending on the context and the general tone of the conversation. Some humans might find one text toxic while others not, it depends on their perception of irony, cynicism... \nTo be honest, I find quite scary the perspective of letting a ML algorithm decide which messages are toxic. Humans are already highly imperfect. I guess it is pretty hard to find a commonly adopted benchmark for human performance.",
    "898053": "Yeah, there was clearly noise in the labels. I wonder if they used mechanical turk or some other external labeling service. The definition they give us for toxic is quite vague and they didnt give us labeling methodology like they did in the last competition.",
    "898071": "Then there will be Unintended Bias in Multilingual Toxic challenge comp next year ;)",
    "898159": "leecming I speak/read pt, es and it (to some level at least)\nI had a look at the validation data and the samples with worst prediction errors, where my model was supposedly completely wrong, and to be honest, I often would agree with my model better than of the annotated labels! 😆 \n\nSome of the labels seem to be automated, e.g. toxic if contains certain words, independent of context.\nHowever, in Italian some weren't tagged toxic, as curse words can be culturally accepted in non-toxic text (apparently some people used a hardcoded negative scale for Italian and Pt).\n\nIn short, I'd do a very bad job ;)",
    "898552": "Hahaha true"
  },
  "source": "meta"
}