{
  "id": 77214,
  "title": "Release of FSDnoisy18k: a dataset to investigate label noise in sound event classification",
  "url": "/competitions/freesound-audio-tagging/discussion/77214",
  "author_name": "Eduardo Fonseca",
  "post_date": "2019-01-10T18:13:55.366000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Dear participants,</p>\n\n<p>We're pleased to announce the release of <strong><a href=\"http://www.eduardofonseca.net/FSDnoisy18k/\">FSDnoisy18k</a></strong>, an open dataset to foster the investigation of label noise in sound event classification. It contains 42.5 hours of audio across 20 sound classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data.</p>\n\n<p>The dataset is released as part of our publication:</p>\n\n<blockquote>\n  <p><a href=\"https://arxiv.org/abs/1901.01189\">Learning Sound Event Classifiers from Web Audio with Noisy Labels</a>\n  E. Fonseca, M. Plakal, D. P. W. Ellis, F. Font, X. Favory, and X. Serra.\n  arXiv preprint arXiv:1901.01189, 2019</p>\n</blockquote>\n\n<p>where we present the dataset and a CNN baseline system. We show that training with large amounts of noisy data can outperform training with smaller amounts of carefully-labeled data. We also show that noise-robust loss functions can be effective in improving performance in presence of corrupted labels.</p>\n\n<p>FSDnoisy18k dataset: <a href=\"http://www.eduardofonseca.net/FSDnoisy18k/\">http://www.eduardofonseca.net/FSDnoisy18k/</a> <br>\nSource code is available: <a href=\"https://github.com/edufonseca/icassp19\">https://github.com/edufonseca/icassp19</a></p>\n\n<p>We hope you find these resources useful!</p>\n\n<p>Thanks!</p>\n\n<p>Eduardo on behalf of the challenge organizers</p>",
  "messages": [
    {
      "id": 453759,
      "postDate": "2019-01-10T18:13:55.367Z",
      "content": "<p>Dear participants,</p>\n\n<p>We're pleased to announce the release of <strong><a href=\"http://www.eduardofonseca.net/FSDnoisy18k/\">FSDnoisy18k</a></strong>, an open dataset to foster the investigation of label noise in sound event classification. It contains 42.5 hours of audio across 20 sound classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data.</p>\n\n<p>The dataset is released as part of our publication:</p>\n\n<blockquote>\n  <p><a href=\"https://arxiv.org/abs/1901.01189\">Learning Sound Event Classifiers from Web Audio with Noisy Labels</a>\n  E. Fonseca, M. Plakal, D. P. W. Ellis, F. Font, X. Favory, and X. Serra.\n  arXiv preprint arXiv:1901.01189, 2019</p>\n</blockquote>\n\n<p>where we present the dataset and a CNN baseline system. We show that training with large amounts of noisy data can outperform training with smaller amounts of carefully-labeled data. We also show that noise-robust loss functions can be effective in improving performance in presence of corrupted labels.</p>\n\n<p>FSDnoisy18k dataset: <a href=\"http://www.eduardofonseca.net/FSDnoisy18k/\">http://www.eduardofonseca.net/FSDnoisy18k/</a> <br>\nSource code is available: <a href=\"https://github.com/edufonseca/icassp19\">https://github.com/edufonseca/icassp19</a></p>\n\n<p>We hope you find these resources useful!</p>\n\n<p>Thanks!</p>\n\n<p>Eduardo on behalf of the challenge organizers</p>",
      "rawMarkdown": "Dear participants,\n\nWe're pleased to announce the release of **[FSDnoisy18k][1]**, an open dataset to foster the investigation of label noise in sound event classification. It contains 42.5 hours of audio across 20 sound classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data.\n\nThe dataset is released as part of our publication:\n\n&gt; [Learning Sound Event Classifiers from Web Audio with Noisy Labels][2]\n&gt; E. Fonseca, M. Plakal, D. P. W. Ellis, F. Font, X. Favory, and X. Serra.\n&gt; arXiv preprint arXiv:1901.01189, 2019\n\nwhere we present the dataset and a CNN baseline system. We show that training with large amounts of noisy data can outperform training with smaller amounts of carefully-labeled data. We also show that noise-robust loss functions can be effective in improving performance in presence of corrupted labels.\n\nFSDnoisy18k dataset: [http://www.eduardofonseca.net/FSDnoisy18k/][3]  \nSource code is available: [https://github.com/edufonseca/icassp19][4]\n\nWe hope you find these resources useful!\n\nThanks!\n\nEduardo on behalf of the challenge organizers\n\n\n  [1]: http://www.eduardofonseca.net/FSDnoisy18k/\n  [2]: https://arxiv.org/abs/1901.01189\n  [3]: http://www.eduardofonseca.net/FSDnoisy18k/\n  [4]: https://github.com/edufonseca/icassp19",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "453759": "Dear participants,\n\nWe're pleased to announce the release of **[FSDnoisy18k][1]**, an open dataset to foster the investigation of label noise in sound event classification. It contains 42.5 hours of audio across 20 sound classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data.\n\nThe dataset is released as part of our publication:\n\n&gt; [Learning Sound Event Classifiers from Web Audio with Noisy Labels][2]\n&gt; E. Fonseca, M. Plakal, D. P. W. Ellis, F. Font, X. Favory, and X. Serra.\n&gt; arXiv preprint arXiv:1901.01189, 2019\n\nwhere we present the dataset and a CNN baseline system. We show that training with large amounts of noisy data can outperform training with smaller amounts of carefully-labeled data. We also show that noise-robust loss functions can be effective in improving performance in presence of corrupted labels.\n\nFSDnoisy18k dataset: [http://www.eduardofonseca.net/FSDnoisy18k/][3]  \nSource code is available: [https://github.com/edufonseca/icassp19][4]\n\nWe hope you find these resources useful!\n\nThanks!\n\nEduardo on behalf of the challenge organizers\n\n\n  [1]: http://www.eduardofonseca.net/FSDnoisy18k/\n  [2]: https://arxiv.org/abs/1901.01189\n  [3]: http://www.eduardofonseca.net/FSDnoisy18k/\n  [4]: https://github.com/edufonseca/icassp19"
  }
}