{
  "id": 58862,
  "title": "Hundreds of test set files which cannot be labelled?",
  "url": "/competitions/freesound-audio-tagging/discussion/58862",
  "author_name": "",
  "post_date": "2018-06-14T19:33:49.536277300Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi All,</p>\n\n<p>I hope that competition organizers will respond to that. </p>\n\n<p>As I mentioned in my <a href=\"https://www.kaggle.com/agehsbarg/audio-challenge-cnn-with-concatenated-inputs\">kernel</a> there are lots of audio files on test set which cannot be labelled by 41 label we have. It was easy to find these files - I have now 3 classifiers and mostly all files for which these classifiers do not agree cannot be labelled. So I make predictions, take audio files for which three classifiers gave three different labels listen to the file and discover smt weird.</p>\n\n<p>Thus I have several questions. Are all these files used for evaluating models on LB? If yes, then what are their labels? How many of such files do we have? If there are around 5% of them, would it mean that maximal score on LB is 0.95, cause anything higher is just guessing? I think these are very important questions, cause otherwise it is not clear what are we trying to achieve here. I also suppose that some of the data on train is also labelled in the wrong way. That would mean that we are building classifier to make predictions close to some other classifier which labelled data.</p>\n\n<p>To be honest, I think that if the aim is just to learn new technology, that competition is awesome (I used it to learn CNN and it is great). However, if the aim is to win, win is not a real \"win\", as some of the data is wrong. Otherwise we compete for the sake of competition, without keeping in mind that we are still doing smt related to the real world.</p>\n\n<p><strong>To be absolutely clear I am not judging, more opening a discussion.</strong></p>",
  "messages": [
    {
      "id": "343156",
      "postDate": "06/14/2018 19:33:49",
      "content": "<p>Hi All,</p>\n\n<p>I hope that competition organizers will respond to that. </p>\n\n<p>As I mentioned in my <a href=\"https://www.kaggle.com/agehsbarg/audio-challenge-cnn-with-concatenated-inputs\">kernel</a> there are lots of audio files on test set which cannot be labelled by 41 label we have. It was easy to find these files - I have now 3 classifiers and mostly all files for which these classifiers do not agree cannot be labelled. So I make predictions, take audio files for which three classifiers gave three different labels listen to the file and discover smt weird.</p>\n\n<p>Thus I have several questions. Are all these files used for evaluating models on LB? If yes, then what are their labels? How many of such files do we have? If there are around 5% of them, would it mean that maximal score on LB is 0.95, cause anything higher is just guessing? I think these are very important questions, cause otherwise it is not clear what are we trying to achieve here. I also suppose that some of the data on train is also labelled in the wrong way. That would mean that we are building classifier to make predictions close to some other classifier which labelled data.</p>\n\n<p>To be honest, I think that if the aim is just to learn new technology, that competition is awesome (I used it to learn CNN and it is great). However, if the aim is to win, win is not a real \"win\", as some of the data is wrong. Otherwise we compete for the sake of competition, without keeping in mind that we are still doing smt related to the real world.</p>\n\n<p><strong>To be absolutely clear I am not judging, more opening a discussion.</strong></p>",
      "rawMarkdown": "Hi All,\n\nI hope that competition organizers will respond to that. \n\nAs I mentioned in my [kernel][1] there are lots of audio files on test set which cannot be labelled by 41 label we have. It was easy to find these files - I have now 3 classifiers and mostly all files for which these classifiers do not agree cannot be labelled. So I make predictions, take audio files for which three classifiers gave three different labels listen to the file and discover smt weird.\n\nThus I have several questions. Are all these files used for evaluating models on LB? If yes, then what are their labels? How many of such files do we have? If there are around 5% of them, would it mean that maximal score on LB is 0.95, cause anything higher is just guessing? I think these are very important questions, cause otherwise it is not clear what are we trying to achieve here. I also suppose that some of the data on train is also labelled in the wrong way. That would mean that we are building classifier to make predictions close to some other classifier which labelled data.\n\nTo be honest, I think that if the aim is just to learn new technology, that competition is awesome (I used it to learn CNN and it is great). However, if the aim is to win, win is not a real \"win\", as some of the data is wrong. Otherwise we compete for the sake of competition, without keeping in mind that we are still doing smt related to the real world.\n\n**To be absolutely clear I am not judging, more opening a discussion.**\n\n  [1]: https://www.kaggle.com/agehsbarg/audio-challenge-cnn-with-concatenated-inputs",
      "votes": null
    },
    {
      "id": "343168",
      "postDate": "06/14/2018 20:12:38",
      "content": "<p>Hi Aleksandrs,</p>\n\n<p>thanks for raising this issue.  I think most of your questions can be answered if you take a closer look at the <strong>Data</strong> description of this competition, in the following link:</p>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging/data\">https://www.kaggle.com/c/freesound-audio-tagging/data</a></p>\n\n<p>In the <strong>\"About this dataset\"</strong> section, it is specified that the <strong>test set</strong> has a number of <em>padding</em> sounds which are <strong>not</strong> used for scoring the systems. We also explain that the <strong>train set</strong> has a portion of manually-verified annotations, and another portion that is non-verified for which we have estimated a minimum quality 65-70% in each category. The annotation type is properly flagged in <code>train.csv</code>, should you want to use this information for your system. Please find further details in the link above.</p>\n\n<p>I hope this clarifies your doubts. If you still have some questions, don't hesitate to ask.  And we'd be grateful if you removed the list of files from your post asap. Thanks!</p>",
      "rawMarkdown": "Hi Aleksandrs,\n\nthanks for raising this issue.  I think most of your questions can be answered if you take a closer look at the **Data** description of this competition, in the following link:\n\nhttps://www.kaggle.com/c/freesound-audio-tagging/data\n\nIn the **\"About this dataset\"** section, it is specified that the **test set** has a number of *padding* sounds which are **not** used for scoring the systems. We also explain that the **train set** has a portion of manually-verified annotations, and another portion that is non-verified for which we have estimated a minimum quality 65-70% in each category. The annotation type is properly flagged in `train.csv`, should you want to use this information for your system. Please find further details in the link above.\n\nI hope this clarifies your doubts. If you still have some questions, don't hesitate to ask.  And we'd be grateful if you removed the list of files from your post asap. Thanks!",
      "votes": null
    },
    {
      "id": "343174",
      "postDate": "06/14/2018 20:38:14",
      "content": "<p>Hi Eduardo,</p>\n\n<p>Thank you for a quick reply. I deleted the list. And thanks for clarifying on data quality, I should have read description more carefully. </p>\n\n<p>Follow up question (sorry for being annoying) - what is the aim of the competition then? Again, to be clear, I really enjoy it, it is very challenging and interesting, and for me it is an awesome way to learn CNNs. But how would we deal with the fact that part of the data is not verified? And why do we need padding sounds? So that people don't hand-label them? Imagine we take one of the classifiers built by ppl from that competition. Would we like to use it in practice? Or our aim is to learn smt? </p>",
      "rawMarkdown": "Hi Eduardo,\n\nThank you for a quick reply. I deleted the list. And thanks for clarifying on data quality, I should have read description more carefully. \n\nFollow up question (sorry for being annoying) - what is the aim of the competition then? Again, to be clear, I really enjoy it, it is very challenging and interesting, and for me it is an awesome way to learn CNNs. But how would we deal with the fact that part of the data is not verified? And why do we need padding sounds? So that people don't hand-label them? Imagine we take one of the classifiers built by ppl from that competition. Would we like to use it in practice? Or our aim is to learn smt?",
      "votes": null
    },
    {
      "id": "345947",
      "postDate": "06/20/2018 19:32:05",
      "content": "<p>Hi, I'll try to answer some of your questions.  </p>\n\n<p>Leveraging subsets of training data with annotations of varying reliability is a research problem in itself. Curating datasets requires a significant effort, and sometimes in real world it happens that you have a small portion of manually-verified, reliable data, and then a larger amount of noisy data with annotations that you can trust only to some extent (or even with no annotations at all). How to deal with this situation is one of the challenges posed by the FSDKaggle2018 dataset that we provide. Some ways of addressing this have been pointed out in the Discussion forum:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052</a></p>\n\n<p>The padding sounds in the test set is a way of dissuading people from hand-labeling the test set and cheating. Of course, we know that participants would never do that. I guess it’s just standard procedure.</p>\n\n<p>Finally, I think the main goal of this competition is to foster and advance research in sound recognition. While our main motivation is the knowledge generated (and shared), models trained on this dataset can have multiple applications, for example related to the automatic description of multimedia content. Soon we expect to have a larger version of the dataset, with more categories and audio data, which will enable more applications.</p>",
      "rawMarkdown": "Hi, I'll try to answer some of your questions.  \n\nLeveraging subsets of training data with annotations of varying reliability is a research problem in itself. Curating datasets requires a significant effort, and sometimes in real world it happens that you have a small portion of manually-verified, reliable data, and then a larger amount of noisy data with annotations that you can trust only to some extent (or even with no annotations at all). How to deal with this situation is one of the challenges posed by the FSDKaggle2018 dataset that we provide. Some ways of addressing this have been pointed out in the Discussion forum:\nhttps://www.kaggle.com/c/freesound-audio-tagging/discussion/58052\n\nThe padding sounds in the test set is a way of dissuading people from hand-labeling the test set and cheating. Of course, we know that participants would never do that. I guess it’s just standard procedure.\n\nFinally, I think the main goal of this competition is to foster and advance research in sound recognition. While our main motivation is the knowledge generated (and shared), models trained on this dataset can have multiple applications, for example related to the automatic description of multimedia content. Soon we expect to have a larger version of the dataset, with more categories and audio data, which will enable more applications.",
      "votes": null
    },
    {
      "id": "345984",
      "postDate": "06/20/2018 21:29:03",
      "content": "<p>Hey @Eduardo, thanks a lot for clarifying.</p>",
      "rawMarkdown": "Hey @Eduardo, thanks a lot for clarifying.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 343168,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "06/14/2018 20:12:38",
      "content": "<p>Hi Aleksandrs,</p>\n\n<p>thanks for raising this issue.  I think most of your questions can be answered if you take a closer look at the <strong>Data</strong> description of this competition, in the following link:</p>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging/data\">https://www.kaggle.com/c/freesound-audio-tagging/data</a></p>\n\n<p>In the <strong>\"About this dataset\"</strong> section, it is specified that the <strong>test set</strong> has a number of <em>padding</em> sounds which are <strong>not</strong> used for scoring the systems. We also explain that the <strong>train set</strong> has a portion of manually-verified annotations, and another portion that is non-verified for which we have estimated a minimum quality 65-70% in each category. The annotation type is properly flagged in <code>train.csv</code>, should you want to use this information for your system. Please find further details in the link above.</p>\n\n<p>I hope this clarifies your doubts. If you still have some questions, don't hesitate to ask.  And we'd be grateful if you removed the list of files from your post asap. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 343174,
          "author_name": "agehsbarg",
          "author_url": "",
          "post_date": "06/14/2018 20:38:14",
          "content": "<p>Hi Eduardo,</p>\n\n<p>Thank you for a quick reply. I deleted the list. And thanks for clarifying on data quality, I should have read description more carefully. </p>\n\n<p>Follow up question (sorry for being annoying) - what is the aim of the competition then? Again, to be clear, I really enjoy it, it is very challenging and interesting, and for me it is an awesome way to learn CNNs. But how would we deal with the fact that part of the data is not verified? And why do we need padding sounds? So that people don't hand-label them? Imagine we take one of the classifiers built by ppl from that competition. Would we like to use it in practice? Or our aim is to learn smt? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345947,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "06/20/2018 19:32:05",
          "content": "<p>Hi, I'll try to answer some of your questions.  </p>\n\n<p>Leveraging subsets of training data with annotations of varying reliability is a research problem in itself. Curating datasets requires a significant effort, and sometimes in real world it happens that you have a small portion of manually-verified, reliable data, and then a larger amount of noisy data with annotations that you can trust only to some extent (or even with no annotations at all). How to deal with this situation is one of the challenges posed by the FSDKaggle2018 dataset that we provide. Some ways of addressing this have been pointed out in the Discussion forum:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052</a></p>\n\n<p>The padding sounds in the test set is a way of dissuading people from hand-labeling the test set and cheating. Of course, we know that participants would never do that. I guess it’s just standard procedure.</p>\n\n<p>Finally, I think the main goal of this competition is to foster and advance research in sound recognition. While our main motivation is the knowledge generated (and shared), models trained on this dataset can have multiple applications, for example related to the automatic description of multimedia content. Soon we expect to have a larger version of the dataset, with more categories and audio data, which will enable more applications.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345984,
          "author_name": "agehsbarg",
          "author_url": "",
          "post_date": "06/20/2018 21:29:03",
          "content": "<p>Hey @Eduardo, thanks a lot for clarifying.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "343156": "Hi All,\n\nI hope that competition organizers will respond to that. \n\nAs I mentioned in my [kernel][1] there are lots of audio files on test set which cannot be labelled by 41 label we have. It was easy to find these files - I have now 3 classifiers and mostly all files for which these classifiers do not agree cannot be labelled. So I make predictions, take audio files for which three classifiers gave three different labels listen to the file and discover smt weird.\n\nThus I have several questions. Are all these files used for evaluating models on LB? If yes, then what are their labels? How many of such files do we have? If there are around 5% of them, would it mean that maximal score on LB is 0.95, cause anything higher is just guessing? I think these are very important questions, cause otherwise it is not clear what are we trying to achieve here. I also suppose that some of the data on train is also labelled in the wrong way. That would mean that we are building classifier to make predictions close to some other classifier which labelled data.\n\nTo be honest, I think that if the aim is just to learn new technology, that competition is awesome (I used it to learn CNN and it is great). However, if the aim is to win, win is not a real \"win\", as some of the data is wrong. Otherwise we compete for the sake of competition, without keeping in mind that we are still doing smt related to the real world.\n\n**To be absolutely clear I am not judging, more opening a discussion.**\n\n  [1]: https://www.kaggle.com/agehsbarg/audio-challenge-cnn-with-concatenated-inputs",
    "343168": "Hi Aleksandrs,\n\nthanks for raising this issue.  I think most of your questions can be answered if you take a closer look at the **Data** description of this competition, in the following link:\n\nhttps://www.kaggle.com/c/freesound-audio-tagging/data\n\nIn the **\"About this dataset\"** section, it is specified that the **test set** has a number of *padding* sounds which are **not** used for scoring the systems. We also explain that the **train set** has a portion of manually-verified annotations, and another portion that is non-verified for which we have estimated a minimum quality 65-70% in each category. The annotation type is properly flagged in `train.csv`, should you want to use this information for your system. Please find further details in the link above.\n\nI hope this clarifies your doubts. If you still have some questions, don't hesitate to ask.  And we'd be grateful if you removed the list of files from your post asap. Thanks!",
    "343174": "Hi Eduardo,\n\nThank you for a quick reply. I deleted the list. And thanks for clarifying on data quality, I should have read description more carefully. \n\nFollow up question (sorry for being annoying) - what is the aim of the competition then? Again, to be clear, I really enjoy it, it is very challenging and interesting, and for me it is an awesome way to learn CNNs. But how would we deal with the fact that part of the data is not verified? And why do we need padding sounds? So that people don't hand-label them? Imagine we take one of the classifiers built by ppl from that competition. Would we like to use it in practice? Or our aim is to learn smt?",
    "345947": "Hi, I'll try to answer some of your questions.  \n\nLeveraging subsets of training data with annotations of varying reliability is a research problem in itself. Curating datasets requires a significant effort, and sometimes in real world it happens that you have a small portion of manually-verified, reliable data, and then a larger amount of noisy data with annotations that you can trust only to some extent (or even with no annotations at all). How to deal with this situation is one of the challenges posed by the FSDKaggle2018 dataset that we provide. Some ways of addressing this have been pointed out in the Discussion forum:\nhttps://www.kaggle.com/c/freesound-audio-tagging/discussion/58052\n\nThe padding sounds in the test set is a way of dissuading people from hand-labeling the test set and cheating. Of course, we know that participants would never do that. I guess it’s just standard procedure.\n\nFinally, I think the main goal of this competition is to foster and advance research in sound recognition. While our main motivation is the knowledge generated (and shared), models trained on this dataset can have multiple applications, for example related to the automatic description of multimedia content. Soon we expect to have a larger version of the dataset, with more categories and audio data, which will enable more applications.",
    "345984": "Hey @Eduardo, thanks a lot for clarifying."
  },
  "source": "meta"
}