{
  "id": 158937,
  "title": "Incomplete training data?",
  "url": "/competitions/birdsong-recognition/discussion/158937",
  "author_name": "",
  "post_date": "2020-06-15T20:49:21.521451900Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>There are zero \"nocall\" labeled training samples. However, the Evaluation overview seems to indicate there will be 5 second intervals in the test data with \"nocall\".  Are we going to need to generate these on our own?</p>",
  "messages": [
    {
      "id": "887742",
      "postDate": "06/15/2020 20:49:21",
      "content": "<p>There are zero \"nocall\" labeled training samples. However, the Evaluation overview seems to indicate there will be 5 second intervals in the test data with \"nocall\".  Are we going to need to generate these on our own?</p>",
      "rawMarkdown": "There are zero \"nocall\" labeled training samples. However, the Evaluation overview seems to indicate there will be 5 second intervals in the test data with \"nocall\".  Are we going to need to generate these on our own?",
      "votes": null
    },
    {
      "id": "887762",
      "postDate": "06/15/2020 21:05:03",
      "content": "<p>I think you may be able to use the example test audio and metadata for some samples of nocall, possibly wherever the birds column is NaN (I haven't checked yet). </p>\n\n<p>However, I am confused about how to generate predictions as the test audio is hidden.</p>",
      "rawMarkdown": "I think you may be able to use the example test audio and metadata for some samples of nocall, possibly wherever the birds column is NaN (I haven't checked yet). \n\nHowever, I am confused about how to generate predictions as the test audio is hidden.",
      "votes": null
    },
    {
      "id": "887815",
      "postDate": "06/15/2020 22:50:34",
      "content": "<p>Hi, David;</p>\n\n<p>There are some annotations in the example test data that aren't in the training set that we discussed and decided to leave in, since they help demonstrate some of the complexity of the problem itself. 'nocall' means there's no call there, which is certainly a real possibility! You'll also probably notice a few 'squirrel' labels. At eval time, we're only looking for the labels in the training set (or no label at all).</p>",
      "rawMarkdown": "Hi, David;\n\nThere are some annotations in the example test data that aren't in the training set that we discussed and decided to leave in, since they help demonstrate some of the complexity of the problem itself. 'nocall' means there's no call there, which is certainly a real possibility! You'll also probably notice a few 'squirrel' labels. At eval time, we're only looking for the labels in the training set (or no label at all).",
      "votes": null
    },
    {
      "id": "888201",
      "postDate": "06/16/2020 07:24:45",
      "content": "<p>The training data is only weakly labeled and contains other than bird sounds. Previous works have shown that utilizing that helps to cope with non-events. I would also recommend using the Google AudioSet as external training data which contains a wide variety of non-bird sounds (wind, water, rain, ...).</p>",
      "rawMarkdown": "The training data is only weakly labeled and contains other than bird sounds. Previous works have shown that utilizing that helps to cope with non-events. I would also recommend using the Google AudioSet as external training data which contains a wide variety of non-bird sounds (wind, water, rain, ...).",
      "votes": null
    },
    {
      "id": "929123",
      "postDate": "07/14/2020 13:29:04",
      "content": "<p>Stefan - from what I find, google AudioData isn't in the public domain since it's based on youtube videos (which are also not in the public domain.) Can you confirm what data sets you're referring to? (Or are you referring to the trained models?)</p>",
      "rawMarkdown": "Stefan - from what I find, google AudioData isn't in the public domain since it's based on youtube videos (which are also not in the public domain.) Can you confirm what data sets you're referring to? (Or are you referring to the trained models?)",
      "votes": null
    },
    {
      "id": "929482",
      "postDate": "07/14/2020 17:43:24",
      "content": "<p>As an alternative, the 'negative' labeled files from the Bird Audio Detection challenge have often been used for augmentation, and are publicly available:\n<a href=\"http://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/\">http://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/</a></p>",
      "rawMarkdown": "As an alternative, the 'negative' labeled files from the Bird Audio Detection challenge have often been used for augmentation, and are publicly available:\nhttp://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/",
      "votes": null
    },
    {
      "id": "999183",
      "postDate": "09/05/2020 12:57:36",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> what's the expected results if there is a bird in the test data which is not in the training data ? Is it considered as no call ?</p>",
      "rawMarkdown": "stefankahl what's the expected results if there is a bird in the test data which is not in the training data ? Is it considered as no call ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 887762,
      "author_name": "tvixarbitrage416",
      "author_url": "",
      "post_date": "06/15/2020 21:05:03",
      "content": "<p>I think you may be able to use the example test audio and metadata for some samples of nocall, possibly wherever the birds column is NaN (I haven't checked yet). </p>\n\n<p>However, I am confused about how to generate predictions as the test audio is hidden.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 887815,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "06/15/2020 22:50:34",
      "content": "<p>Hi, David;</p>\n\n<p>There are some annotations in the example test data that aren't in the training set that we discussed and decided to leave in, since they help demonstrate some of the complexity of the problem itself. 'nocall' means there's no call there, which is certainly a real possibility! You'll also probably notice a few 'squirrel' labels. At eval time, we're only looking for the labels in the training set (or no label at all).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 888201,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "06/16/2020 07:24:45",
      "content": "<p>The training data is only weakly labeled and contains other than bird sounds. Previous works have shown that utilizing that helps to cope with non-events. I would also recommend using the Google AudioSet as external training data which contains a wide variety of non-bird sounds (wind, water, rain, ...).</p>",
      "votes": null,
      "replies": [
        {
          "id": 929123,
          "author_name": "roman99",
          "author_url": "",
          "post_date": "07/14/2020 13:29:04",
          "content": "<p>Stefan - from what I find, google AudioData isn't in the public domain since it's based on youtube videos (which are also not in the public domain.) Can you confirm what data sets you're referring to? (Or are you referring to the trained models?)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929482,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "07/14/2020 17:43:24",
          "content": "<p>As an alternative, the 'negative' labeled files from the Bird Audio Detection challenge have often been used for augmentation, and are publicly available:\n<a href=\"http://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/\">http://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 999183,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "09/05/2020 12:57:36",
          "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> what's the expected results if there is a bird in the test data which is not in the training data ? Is it considered as no call ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "887742": "There are zero \"nocall\" labeled training samples. However, the Evaluation overview seems to indicate there will be 5 second intervals in the test data with \"nocall\".  Are we going to need to generate these on our own?",
    "887762": "I think you may be able to use the example test audio and metadata for some samples of nocall, possibly wherever the birds column is NaN (I haven't checked yet). \n\nHowever, I am confused about how to generate predictions as the test audio is hidden.",
    "887815": "Hi, David;\n\nThere are some annotations in the example test data that aren't in the training set that we discussed and decided to leave in, since they help demonstrate some of the complexity of the problem itself. 'nocall' means there's no call there, which is certainly a real possibility! You'll also probably notice a few 'squirrel' labels. At eval time, we're only looking for the labels in the training set (or no label at all).",
    "888201": "The training data is only weakly labeled and contains other than bird sounds. Previous works have shown that utilizing that helps to cope with non-events. I would also recommend using the Google AudioSet as external training data which contains a wide variety of non-bird sounds (wind, water, rain, ...).",
    "929123": "Stefan - from what I find, google AudioData isn't in the public domain since it's based on youtube videos (which are also not in the public domain.) Can you confirm what data sets you're referring to? (Or are you referring to the trained models?)",
    "929482": "As an alternative, the 'negative' labeled files from the Bird Audio Detection challenge have often been used for augmentation, and are publicly available:\nhttp://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/",
    "999183": "stefankahl what's the expected results if there is a bird in the test data which is not in the training data ? Is it considered as no call ?"
  },
  "source": "meta"
}