{
  "id": 92313,
  "title": "Incorrect label in train_curated",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/92313",
  "author_name": "",
  "post_date": "2019-05-15T09:53:19.967452200Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p><code>train_curated/f76181c4.wav</code> is labelled as <code>Male_speech_and_man_speaking</code> but sounds like electronic drums. How clean would we reasonably assume that the Freesound Annotator crowdsourcing process is?</p>",
  "messages": [
    {
      "id": "531660",
      "postDate": "05/15/2019 09:53:19",
      "content": "<p><code>train_curated/f76181c4.wav</code> is labelled as <code>Male_speech_and_man_speaking</code> but sounds like electronic drums. How clean would we reasonably assume that the Freesound Annotator crowdsourcing process is?</p>",
      "rawMarkdown": "`train_curated/f76181c4.wav` is labelled as `Male_speech_and_man_speaking` but sounds like electronic drums. How clean would we reasonably assume that the Freesound Annotator crowdsourcing process is?",
      "votes": null
    },
    {
      "id": "531946",
      "postDate": "05/15/2019 21:59:39",
      "content": "<p>Hi Carl, </p>\n\n<p>just checked the clip. We're sorry about this. After some digging, I determined that it is due to a mistake in the process of format conversion and renaming of the clip (and not to the annotation, in this case).  This type of mistake is rare and it is the first time a participant points it out. In fact, I did not know this kind of thing could happen :) We will double check it from now on (thanks for that).</p>\n\n<p>Answering your question: All the clips in the curated train set has been annotated by humans, most of the clips have inter annotator agreement, but not all. Some of the annotators were experts, but not all.  I'd say the rate of label error in the curated train set is pretty low and due to few human errors  or very unexpected issues like the one you spotted.  In any case, this is nothing compared to the noisy train set, where the label noise amount can be severe in certain categories.</p>\n\n<p>Thanks for letting us know!</p>",
      "rawMarkdown": "Hi Carl, \n\njust checked the clip. We're sorry about this. After some digging, I determined that it is due to a mistake in the process of format conversion and renaming of the clip (and not to the annotation, in this case).  This type of mistake is rare and it is the first time a participant points it out. In fact, I did not know this kind of thing could happen :) We will double check it from now on (thanks for that).\n\nAnswering your question: All the clips in the curated train set has been annotated by humans, most of the clips have inter annotator agreement, but not all. Some of the annotators were experts, but not all.  I'd say the rate of label error in the curated train set is pretty low and due to few human errors  or very unexpected issues like the one you spotted.  In any case, this is nothing compared to the noisy train set, where the label noise amount can be severe in certain categories.\n\nThanks for letting us know!",
      "votes": null
    },
    {
      "id": "532091",
      "postDate": "05/16/2019 07:19:20",
      "content": "<p>For the record then, <code>train_curated/77b925c2.wav</code> is also corrupt. It's a 57 second long clip with just beeping noises, labeled as \"Stream\".</p>",
      "rawMarkdown": "For the record then, `train_curated/77b925c2.wav` is also corrupt. It's a 57 second long clip with just beeping noises, labeled as \"Stream\".",
      "votes": null
    },
    {
      "id": "533714",
      "postDate": "05/19/2019 19:41:59",
      "content": "<p>Thanks. This was caused due to exactly the same bug.</p>",
      "rawMarkdown": "Thanks. This was caused due to exactly the same bug.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 531946,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "05/15/2019 21:59:39",
      "content": "<p>Hi Carl, </p>\n\n<p>just checked the clip. We're sorry about this. After some digging, I determined that it is due to a mistake in the process of format conversion and renaming of the clip (and not to the annotation, in this case).  This type of mistake is rare and it is the first time a participant points it out. In fact, I did not know this kind of thing could happen :) We will double check it from now on (thanks for that).</p>\n\n<p>Answering your question: All the clips in the curated train set has been annotated by humans, most of the clips have inter annotator agreement, but not all. Some of the annotators were experts, but not all.  I'd say the rate of label error in the curated train set is pretty low and due to few human errors  or very unexpected issues like the one you spotted.  In any case, this is nothing compared to the noisy train set, where the label noise amount can be severe in certain categories.</p>\n\n<p>Thanks for letting us know!</p>",
      "votes": null,
      "replies": [
        {
          "id": 532091,
          "author_name": "felixab",
          "author_url": "",
          "post_date": "05/16/2019 07:19:20",
          "content": "<p>For the record then, <code>train_curated/77b925c2.wav</code> is also corrupt. It's a 57 second long clip with just beeping noises, labeled as \"Stream\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 533714,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "05/19/2019 19:41:59",
          "content": "<p>Thanks. This was caused due to exactly the same bug.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "531660": "`train_curated/f76181c4.wav` is labelled as `Male_speech_and_man_speaking` but sounds like electronic drums. How clean would we reasonably assume that the Freesound Annotator crowdsourcing process is?",
    "531946": "Hi Carl, \n\njust checked the clip. We're sorry about this. After some digging, I determined that it is due to a mistake in the process of format conversion and renaming of the clip (and not to the annotation, in this case).  This type of mistake is rare and it is the first time a participant points it out. In fact, I did not know this kind of thing could happen :) We will double check it from now on (thanks for that).\n\nAnswering your question: All the clips in the curated train set has been annotated by humans, most of the clips have inter annotator agreement, but not all. Some of the annotators were experts, but not all.  I'd say the rate of label error in the curated train set is pretty low and due to few human errors  or very unexpected issues like the one you spotted.  In any case, this is nothing compared to the noisy train set, where the label noise amount can be severe in certain categories.\n\nThanks for letting us know!",
    "532091": "For the record then, `train_curated/77b925c2.wav` is also corrupt. It's a 57 second long clip with just beeping noises, labeled as \"Stream\".",
    "533714": "Thanks. This was caused due to exactly the same bug."
  },
  "source": "meta"
}