{
  "id": 540969,
  "title": "Is the evaluation data labeled using spectrograms?",
  "url": "/competitions/birdclef-2024/discussion/540969",
  "author_name": "",
  "post_date": "2024-10-16T20:13:16.291544500Z",
  "votes": -1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Could somebody please elaborate on how the evaluation dataset is made? <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<p>Outside of Kaggle I have been putting together a large dataset of New Zealand bird calls.  I hope to make it public eventually.  The difference in my case is that half the data was labelled using spectrograms, using the <a href=\"https://freebird.co.nz/\" target=\"_blank\">freebird</a> labelling tool. </p>\n<p>Anecdotally, it appears that the classifier I built from my new dataset can detect true positives well below volumes that my human expert colleagues can hear.  We need to use spectrograms ourselves to make any robust claims about performance.  </p>\n<p>Of course you could have a philosophical debate about whether a bird is there if you can't here it.  But for the purpose of wildlife abundance monitoring, we would like to capture birds as far from the recording device as practical.</p>\n<p>The implications for Kaggle are interesting:</p>\n<ol>\n<li>The Xeno-Canto data is probably riddled with false negatives.</li>\n<li>The test soundscapes may or may not be equally full of false negatives, depending how they are labelled.  </li>\n<li>If you use the 'background' birds in training, you are slightly more robust to the first point, but could be penalised by the second.</li>\n</ol>\n<p>We can't do much about the training data. But if the test soundscapes were indeed labelled by humans listening to recordings, maybe for 2025 we could change this?  It would be super interesting to see how performance changes.</p>",
  "messages": [
    {
      "id": "3019750",
      "postDate": "10/16/2024 20:13:16",
      "content": "<p>Could somebody please elaborate on how the evaluation dataset is made? <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<p>Outside of Kaggle I have been putting together a large dataset of New Zealand bird calls.  I hope to make it public eventually.  The difference in my case is that half the data was labelled using spectrograms, using the <a href=\"https://freebird.co.nz/\" target=\"_blank\">freebird</a> labelling tool. </p>\n<p>Anecdotally, it appears that the classifier I built from my new dataset can detect true positives well below volumes that my human expert colleagues can hear.  We need to use spectrograms ourselves to make any robust claims about performance.  </p>\n<p>Of course you could have a philosophical debate about whether a bird is there if you can't here it.  But for the purpose of wildlife abundance monitoring, we would like to capture birds as far from the recording device as practical.</p>\n<p>The implications for Kaggle are interesting:</p>\n<ol>\n<li>The Xeno-Canto data is probably riddled with false negatives.</li>\n<li>The test soundscapes may or may not be equally full of false negatives, depending how they are labelled.  </li>\n<li>If you use the 'background' birds in training, you are slightly more robust to the first point, but could be penalised by the second.</li>\n</ol>\n<p>We can't do much about the training data. But if the test soundscapes were indeed labelled by humans listening to recordings, maybe for 2025 we could change this?  It would be super interesting to see how performance changes.</p>",
      "rawMarkdown": "Could somebody please elaborate on how the evaluation dataset is made? @tomdenton @stefankahl \n\nOutside of Kaggle I have been putting together a large dataset of New Zealand bird calls.  I hope to make it public eventually.  The difference in my case is that half the data was labelled using spectrograms, using the [freebird](https://freebird.co.nz/) labelling tool. \n\nAnecdotally, it appears that the classifier I built from my new dataset can detect true positives well below volumes that my human expert colleagues can hear.  We need to use spectrograms ourselves to make any robust claims about performance.  \n\nOf course you could have a philosophical debate about whether a bird is there if you can't here it.  But for the purpose of wildlife abundance monitoring, we would like to capture birds as far from the recording device as practical.\n\nThe implications for Kaggle are interesting:\n1. The Xeno-Canto data is probably riddled with false negatives.\n2. The test soundscapes may or may not be equally full of false negatives, depending how they are labelled.  \n3.  If you use the 'background' birds in training, you are slightly more robust to the first point, but could be penalised by the second.\n\nWe can't do much about the training data. But if the test soundscapes were indeed labelled by humans listening to recordings, maybe for 2025 we could change this?  It would be super interesting to see how performance changes.",
      "votes": null
    },
    {
      "id": "3019770",
      "postDate": "10/16/2024 20:37:20",
      "content": "<p>Hi, Olly!</p>\n<p>1+3. Yes, absolutely. Part of our challenge is to build models which perform well despite the missing labels. I've found that the 'background birds' labels in Xeno Canto tend to be unreliably filled out, and probably less rigorously checked than the primary labels.</p>\n<ol>\n<li>Soundscape labeling varies a bit form year to year, as we work with different groups of experts from around the world. However, I would say it's typical that 'full' annotations are performed at a file-level, using a combination of listening and spectrograms (eg, with Raven). The experts I have personally worked with have tended to gather significant knowledge from the file-level context, which they then use to label lots of ambiguous calls which would otherwise be difficult to identify. That said, the expert annotations are generally incomplete, as there are still ambiguous/difficult to ID calls… There's a precision/recall tradeoff for human annotations as well, regardless of methodology.</li>\n</ol>\n<p>Hope that helps!</p>",
      "rawMarkdown": "Hi, Olly!\n\n1+3. Yes, absolutely. Part of our challenge is to build models which perform well despite the missing labels. I've found that the 'background birds' labels in Xeno Canto tend to be unreliably filled out, and probably less rigorously checked than the primary labels.\n\n2. Soundscape labeling varies a bit form year to year, as we work with different groups of experts from around the world. However, I would say it's typical that 'full' annotations are performed at a file-level, using a combination of listening and spectrograms (eg, with Raven). The experts I have personally worked with have tended to gather significant knowledge from the file-level context, which they then use to label lots of ambiguous calls which would otherwise be difficult to identify. That said, the expert annotations are generally incomplete, as there are still ambiguous/difficult to ID calls... There's a precision/recall tradeoff for human annotations as well, regardless of methodology.\n\nHope that helps!",
      "votes": null
    },
    {
      "id": "3019851",
      "postDate": "10/16/2024 23:55:24",
      "content": "<p>Thanks Tom, that's awesome!  I'm still partly relying on the Kenyan &amp; Western Ghats soundscapes to benchmark my New Zealand work.  So it sounds like it's about as good as can be reasonably achieved then.  Cheers.</p>",
      "rawMarkdown": "Thanks Tom, that's awesome!  I'm still partly relying on the Kenyan & Western Ghats soundscapes to benchmark my New Zealand work.  So it sounds like it's about as good as can be reasonably achieved then.  Cheers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3019770,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "10/16/2024 20:37:20",
      "content": "<p>Hi, Olly!</p>\n<p>1+3. Yes, absolutely. Part of our challenge is to build models which perform well despite the missing labels. I've found that the 'background birds' labels in Xeno Canto tend to be unreliably filled out, and probably less rigorously checked than the primary labels.</p>\n<ol>\n<li>Soundscape labeling varies a bit form year to year, as we work with different groups of experts from around the world. However, I would say it's typical that 'full' annotations are performed at a file-level, using a combination of listening and spectrograms (eg, with Raven). The experts I have personally worked with have tended to gather significant knowledge from the file-level context, which they then use to label lots of ambiguous calls which would otherwise be difficult to identify. That said, the expert annotations are generally incomplete, as there are still ambiguous/difficult to ID calls… There's a precision/recall tradeoff for human annotations as well, regardless of methodology.</li>\n</ol>\n<p>Hope that helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3019851,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "10/16/2024 23:55:24",
          "content": "<p>Thanks Tom, that's awesome!  I'm still partly relying on the Kenyan &amp; Western Ghats soundscapes to benchmark my New Zealand work.  So it sounds like it's about as good as can be reasonably achieved then.  Cheers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3019750": "Could somebody please elaborate on how the evaluation dataset is made? @tomdenton @stefankahl \n\nOutside of Kaggle I have been putting together a large dataset of New Zealand bird calls.  I hope to make it public eventually.  The difference in my case is that half the data was labelled using spectrograms, using the [freebird](https://freebird.co.nz/) labelling tool. \n\nAnecdotally, it appears that the classifier I built from my new dataset can detect true positives well below volumes that my human expert colleagues can hear.  We need to use spectrograms ourselves to make any robust claims about performance.  \n\nOf course you could have a philosophical debate about whether a bird is there if you can't here it.  But for the purpose of wildlife abundance monitoring, we would like to capture birds as far from the recording device as practical.\n\nThe implications for Kaggle are interesting:\n1. The Xeno-Canto data is probably riddled with false negatives.\n2. The test soundscapes may or may not be equally full of false negatives, depending how they are labelled.  \n3.  If you use the 'background' birds in training, you are slightly more robust to the first point, but could be penalised by the second.\n\nWe can't do much about the training data. But if the test soundscapes were indeed labelled by humans listening to recordings, maybe for 2025 we could change this?  It would be super interesting to see how performance changes.",
    "3019770": "Hi, Olly!\n\n1+3. Yes, absolutely. Part of our challenge is to build models which perform well despite the missing labels. I've found that the 'background birds' labels in Xeno Canto tend to be unreliably filled out, and probably less rigorously checked than the primary labels.\n\n2. Soundscape labeling varies a bit form year to year, as we work with different groups of experts from around the world. However, I would say it's typical that 'full' annotations are performed at a file-level, using a combination of listening and spectrograms (eg, with Raven). The experts I have personally worked with have tended to gather significant knowledge from the file-level context, which they then use to label lots of ambiguous calls which would otherwise be difficult to identify. That said, the expert annotations are generally incomplete, as there are still ambiguous/difficult to ID calls... There's a precision/recall tradeoff for human annotations as well, regardless of methodology.\n\nHope that helps!",
    "3019851": "Thanks Tom, that's awesome!  I'm still partly relying on the Kenyan & Western Ghats soundscapes to benchmark my New Zealand work.  So it sounds like it's about as good as can be reasonably achieved then.  Cheers."
  },
  "source": "meta"
}