{
  "id": 159172,
  "title": "Welcome!",
  "url": "/competitions/birdsong-recognition/discussion/159172",
  "author_name": "",
  "post_date": "2020-06-16T16:35:33.565541900Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi, all! </p>\n\n<p>I'm one of the co-organizers of this competition. I initially got involved in birdsong identification while looking to learn more about deep learning in the audio domain. It's a fascinating problem space, and over the last few years I've learned quite a lot about birds in my quest to build better models.</p>\n\n<p>Unfortunately, my own time for exploring models is limited, so I was very excited by the idea of getting the problem on Kaggle. I competed in a <a href=\"https://www.kaggle.com/c/learning-social-circles\">research-oriented Kaggle competition</a> six years ago, which I won* by a) not over-fitting on the public leaderboard, and b) challenging some of the assumptions in the problem framing. I hope to see some interesting attacks on the problem here, and see some of my own assumptions challenged!</p>\n\n<ul>\n<li><ul><li>The victory was admittedly by a tiiiiiny margin.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "888910",
      "postDate": "06/16/2020 16:35:33",
      "content": "<p>Hi, all! </p>\n\n<p>I'm one of the co-organizers of this competition. I initially got involved in birdsong identification while looking to learn more about deep learning in the audio domain. It's a fascinating problem space, and over the last few years I've learned quite a lot about birds in my quest to build better models.</p>\n\n<p>Unfortunately, my own time for exploring models is limited, so I was very excited by the idea of getting the problem on Kaggle. I competed in a <a href=\"https://www.kaggle.com/c/learning-social-circles\">research-oriented Kaggle competition</a> six years ago, which I won* by a) not over-fitting on the public leaderboard, and b) challenging some of the assumptions in the problem framing. I hope to see some interesting attacks on the problem here, and see some of my own assumptions challenged!</p>\n\n<ul>\n<li><ul><li>The victory was admittedly by a tiiiiiny margin.</li></ul></li>\n</ul>",
      "rawMarkdown": "Hi, all! \n\nI'm one of the co-organizers of this competition. I initially got involved in birdsong identification while looking to learn more about deep learning in the audio domain. It's a fascinating problem space, and over the last few years I've learned quite a lot about birds in my quest to build better models.\n\nUnfortunately, my own time for exploring models is limited, so I was very excited by the idea of getting the problem on Kaggle. I competed in a [research-oriented Kaggle competition](https://www.kaggle.com/c/learning-social-circles) six years ago, which I won* by a) not over-fitting on the public leaderboard, and b) challenging some of the assumptions in the problem framing. I hope to see some interesting attacks on the problem here, and see some of my own assumptions challenged!\n\n* - The victory was admittedly by a tiiiiiny margin.",
      "votes": null
    },
    {
      "id": "891067",
      "postDate": "06/17/2020 21:45:13",
      "content": "<p>Heya, thanks for helping put on the comp, it looks really cool. Just a quick question about the test set (if I'm allowed to ask) the 150 10 minute recordings: are any of them composed of multiple splices of different recordings, or are they all contiguous audio? </p>",
      "rawMarkdown": "Heya, thanks for helping put on the comp, it looks really cool. Just a quick question about the test set (if I'm allowed to ask) the 150 10 minute recordings: are any of them composed of multiple splices of different recordings, or are they all contiguous audio?",
      "votes": null
    },
    {
      "id": "891146",
      "postDate": "06/18/2020 00:14:37",
      "content": "<p>Hi, Louka! Each recording is a contiguous 10m chunk, without splicing.</p>",
      "rawMarkdown": "Hi, Louka! Each recording is a contiguous 10m chunk, without splicing.",
      "votes": null
    },
    {
      "id": "893684",
      "postDate": "06/19/2020 19:56:11",
      "content": "<p>Hey again <a href=\"/tomdenton\">@tomdenton</a>, it sounds like you gentlemen have already been trying to solve a bunch of similar problems for a while now. Do you have any general advice about how to approach bird-related audio processing machine learning tasks? Perhaps there are some online resources you found helpful?</p>\n\n<p>Cheers.</p>",
      "rawMarkdown": "Hey again @tomdenton, it sounds like you gentlemen have already been trying to solve a bunch of similar problems for a while now. Do you have any general advice about how to approach bird-related audio processing machine learning tasks? Perhaps there are some online resources you found helpful?\n\nCheers.",
      "votes": null
    },
    {
      "id": "893788",
      "postDate": "06/19/2020 22:22:16",
      "content": "<p>Hi, Louka;\nI'm sure you've seen the great references in this thread already: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158933\">https://www.kaggle.com/c/birdsong-recognition/discussion/158933</a>\nThe papers from BirdCLEF competitions are good reading.</p>\n\n<p>A couple points/problems that I keep in mind, specific to birds and bioacoustics and soundscapes:\n* Augmentation is very important; <a href=\"http://ceur-ws.org/Vol-2125/paper_140.pdf\">Mario Lasseck's paper from 2018 BirdCLEF</a> has some great ablation on different kinds of augmentation.\n* Expanding on that, it's good to keep in mind that the audio environment is very complex, and full of things (like, say, squirrels) that you don't care much about.\n* Some species are highly variable, while others are more constant. Even the more constant songs may vary a bit by geography. Species with a lot of variability sometimes have distinctive notes in otherwise chaotic songs. For a good example of this, you can compare the opening notes of song sparrows to brewer's sparrow.\n* Most species have separate 'songs' and 'calls', with some per-species variability in each.\n* Bird hearing (especially for songbirds) has a much higher time-resolution than human hearing. </p>\n\n<p>And here's how these turn into various not-well-solved problems in the machine learned models:\n* If a particular song is very loud and common, it may appear in the background of many recordings, leading to lower prediction scores for that song. In theory, this is just a calibration problem. In practice, each species has many sounds (with different appropriate calibrations), hidden behind a single class label. I would love to have some better ways to deal with this problem!\n* Warblers seem to be very difficult for the particular models I work with... Their songs are on the shorter-side, and seem to be easily confused. There may be some regional variability issue at play here.\n* Xeno-Canto /probably/ tends towards cleaner, easily identified sounds, since people can ID these. Truncated songs, juveniles (who haven't learned to sing well yet), hybrids, and mimics all add complexity to the sound space, and may not be well represented since they're harder for humans to ID as well.</p>\n\n<p>Ultimately, we still see a big gap between real-world soundscape performance and performance on Xeno-Canto holdout data... And thus the competition!</p>",
      "rawMarkdown": "Hi, Louka;\nI'm sure you've seen the great references in this thread already: https://www.kaggle.com/c/birdsong-recognition/discussion/158933\nThe papers from BirdCLEF competitions are good reading.\n\nA couple points/problems that I keep in mind, specific to birds and bioacoustics and soundscapes:\n* Augmentation is very important; [Mario Lasseck's paper from 2018 BirdCLEF](http://ceur-ws.org/Vol-2125/paper_140.pdf) has some great ablation on different kinds of augmentation.\n* Expanding on that, it's good to keep in mind that the audio environment is very complex, and full of things (like, say, squirrels) that you don't care much about.\n* Some species are highly variable, while others are more constant. Even the more constant songs may vary a bit by geography. Species with a lot of variability sometimes have distinctive notes in otherwise chaotic songs. For a good example of this, you can compare the opening notes of song sparrows to brewer's sparrow.\n* Most species have separate 'songs' and 'calls', with some per-species variability in each.\n* Bird hearing (especially for songbirds) has a much higher time-resolution than human hearing. \n\nAnd here's how these turn into various not-well-solved problems in the machine learned models:\n* If a particular song is very loud and common, it may appear in the background of many recordings, leading to lower prediction scores for that song. In theory, this is just a calibration problem. In practice, each species has many sounds (with different appropriate calibrations), hidden behind a single class label. I would love to have some better ways to deal with this problem!\n* Warblers seem to be very difficult for the particular models I work with... Their songs are on the shorter-side, and seem to be easily confused. There may be some regional variability issue at play here.\n* Xeno-Canto /probably/ tends towards cleaner, easily identified sounds, since people can ID these. Truncated songs, juveniles (who haven't learned to sing well yet), hybrids, and mimics all add complexity to the sound space, and may not be well represented since they're harder for humans to ID as well.\n\nUltimately, we still see a big gap between real-world soundscape performance and performance on Xeno-Canto holdout data... And thus the competition!",
      "votes": null
    },
    {
      "id": "893851",
      "postDate": "06/20/2020 02:47:19",
      "content": "<p>Thanks Tom</p>",
      "rawMarkdown": "Thanks Tom",
      "votes": null
    },
    {
      "id": "894896",
      "postDate": "06/20/2020 21:55:27",
      "content": "<p>Hey again Tom, sorry to keep bothering you!</p>\n\n<p>I was doing some local cross validation and it occurred to me that I don't know what \"good\" performance looks like. I've done a light literature review and found some prior work, like </p>\n\n<ul>\n<li><a href=\"https://www.acoustics.asn.au/conference_proceedings/AAS2018/papers/p134.pdf\">these guys</a> who got ~65% accuracy classifying 46 species on xeno-canto</li>\n<li><a href=\"https://besjournals.onlinelibrary.wiley.com/doi/full/10.1111/2041-210X.13103\">this crowd</a> who got ~88% AUC score for classifying bird-present/absent on some other datasets</li>\n<li><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4106198/\">these goons</a> who achieved ~80% AUC predicting the presence/absence of 77 bird species, or</li>\n<li><a href=\"https://www.researchgate.net/profile/Dmitry_Konovalov2/publication/335880577_Data-Efficient_Classification_of_Birdcall_Through_Convolutional_Neural_Networks_Transfer_Learning/links/5db38cf692851c577ec35f2f/Data-Efficient-Classification-of-Birdcall-Through-Convolutional-Neural-Networks-Transfer-Learning.pdf\">this mob</a> who achieved ~78% accuracy predicting different numbers of bird classes using some pretty small  xeno-canto datasets</li>\n</ul>\n\n<p>Which I think gives us a decent estimate of expected performance. But, since we have a domain expert like you around I was wondering if you could give us some <em>rough</em> ballpark estimates of:</p>\n\n<ol>\n<li>The (test) accuracy of the best model you think you personally could build if given 90% of the competition training data (at random) and asked to predict the labels of the other 10% (so, since all the test data only ever has one label, a multi-class classification task where each sample has at most one class).</li>\n<li>The public leaderborard score (F1) you think the top performing will achieve on the actual competition task by the end of the competition.</li>\n</ol>\n\n<p>I don't know if you're allowed to make estimates like this, but if you are I think it would be a good yardstick for competitors to gauge whether we're on the right track or not.</p>",
      "rawMarkdown": "Hey again Tom, sorry to keep bothering you!\n\nI was doing some local cross validation and it occurred to me that I don't know what \"good\" performance looks like. I've done a light literature review and found some prior work, like \n\n- [these guys](https://www.acoustics.asn.au/conference_proceedings/AAS2018/papers/p134.pdf) who got ~65% accuracy classifying 46 species on xeno-canto\n- [this crowd](https://besjournals.onlinelibrary.wiley.com/doi/full/10.1111/2041-210X.13103) who got ~88% AUC score for classifying bird-present/absent on some other datasets\n- [these goons](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4106198/) who achieved ~80% AUC predicting the presence/absence of 77 bird species, or\n- [this mob](https://www.researchgate.net/profile/Dmitry_Konovalov2/publication/335880577_Data-Efficient_Classification_of_Birdcall_Through_Convolutional_Neural_Networks_Transfer_Learning/links/5db38cf692851c577ec35f2f/Data-Efficient-Classification-of-Birdcall-Through-Convolutional-Neural-Networks-Transfer-Learning.pdf) who achieved ~78% accuracy predicting different numbers of bird classes using some pretty small  xeno-canto datasets\n\nWhich I think gives us a decent estimate of expected performance. But, since we have a domain expert like you around I was wondering if you could give us some *rough* ballpark estimates of:\n\n1. The (test) accuracy of the best model you think you personally could build if given 90% of the competition training data (at random) and asked to predict the labels of the other 10% (so, since all the test data only ever has one label, a multi-class classification task where each sample has at most one class).\n2. The public leaderborard score (F1) you think the top performing will achieve on the actual competition task by the end of the competition.\n\nI don't know if you're allowed to make estimates like this, but if you are I think it would be a good yardstick for competitors to gauge whether we're on the right track or not.",
      "votes": null
    },
    {
      "id": "897325",
      "postDate": "06/22/2020 19:22:02",
      "content": "<p>Well, my hope is that the final leaderboard score exceeds my expectations, so I'll avoid biasing the field with a specific prediction. :)</p>\n\n<p>One thing to take note of in lit-review: We tend to see <em>much</em> better model performance when predicting species in held-out Xeno-Canto recordings than in soundscapes, and the higher numbers tend to be what's reported in paper abstracts. High-rated XC samples to be clearer recordings, perhaps with more prototypical sounds. It's very reasonable to use XC holdouts to validate that your models are working, but keep in mind the domain shift problem.</p>",
      "rawMarkdown": "Well, my hope is that the final leaderboard score exceeds my expectations, so I'll avoid biasing the field with a specific prediction. :)\n\nOne thing to take note of in lit-review: We tend to see *much* better model performance when predicting species in held-out Xeno-Canto recordings than in soundscapes, and the higher numbers tend to be what's reported in paper abstracts. High-rated XC samples to be clearer recordings, perhaps with more prototypical sounds. It's very reasonable to use XC holdouts to validate that your models are working, but keep in mind the domain shift problem.",
      "votes": null
    },
    {
      "id": "898264",
      "postDate": "06/23/2020 12:07:28",
      "content": "<p>cheers tom, ill keep that in mind re the literature</p>",
      "rawMarkdown": "cheers tom, ill keep that in mind re the literature",
      "votes": null
    },
    {
      "id": "902655",
      "postDate": "06/26/2020 09:32:49",
      "content": "<p>G'day again Tom, slightly sneaky question this time: was the public test set chosen at random or was it selected manually? And if it's the latter was there some effort made to skew the distribution away from the overall test set distribution?</p>\n\n<p>Obviously only let me know if you think it'll lead to better production models.</p>",
      "rawMarkdown": "G'day again Tom, slightly sneaky question this time: was the public test set chosen at random or was it selected manually? And if it's the latter was there some effort made to skew the distribution away from the overall test set distribution?\n\nObviously only let me know if you think it'll lead to better production models.",
      "votes": null
    },
    {
      "id": "905698",
      "postDate": "06/28/2020 18:14:10",
      "content": "<p><a href=\"/tomdenton\">@tomdenton</a>  <a href=\"/stefankahl\">@stefankahl</a> I spent some time with catching up with the recent BirdCLEF papers. I understand that for real world applications you would like to use only easily accessible monophone recordings for training (e.g. XenoCanto). However even real world applications could benefit from additional metadata (e.g. time of recording/season/during the day/nocturnal location of the recording, elevation) leading to useful a priori assumptions.</p>\n\n<p>The competition setting is meaningful and challenging and I will definitely dig deeper. However I feel a bit that we are flying blind. I have a silly analogy, that we were provided english pdf documents for training and we have to build an OCR application for handwriting recognition for an unknown language :)</p>\n\n<p>Any additional help regarding the hidden soundscape test set (similar validation set, metadata) would be really helpful... So far we learned from the forum that recordings were taken in North America with 30kHz.</p>",
      "rawMarkdown": "tomdenton  @stefankahl I spent some time with catching up with the recent BirdCLEF papers. I understand that for real world applications you would like to use only easily accessible monophone recordings for training (e.g. XenoCanto). However even real world applications could benefit from additional metadata (e.g. time of recording/season/during the day/nocturnal location of the recording, elevation) leading to useful a priori assumptions.\n\nThe competition setting is meaningful and challenging and I will definitely dig deeper. However I feel a bit that we are flying blind. I have a silly analogy, that we were provided english pdf documents for training and we have to build an OCR application for handwriting recognition for an unknown language :)\n\nAny additional help regarding the hidden soundscape test set (similar validation set, metadata) would be really helpful... So far we learned from the forum that recordings were taken in North America with 30kHz.",
      "votes": null
    },
    {
      "id": "905796",
      "postDate": "06/28/2020 19:52:45",
      "content": "<p>Well, you can always turn the question on its head... If the metadata tells you a lot about the birds which might be present, can you instead use the audio to predict the metadata? This approach is often used for unsupervised pre-training on large partially labeled datasets: good metadata prediction means the model 'understands' the environment, and should be able to learn the classification task with a smaller collection of hard labels.</p>\n\n<p>Metadata can also be misleading. If the weather is a bit warmer this year, perhaps the migration has started early, which could mean that what the model has learned form other years about geo and date distributions could be off... A system heavily reliant on metadata will tend miss exceptional cases.</p>\n\n<p>So, our goal is to see how far we can get with 'pure audio' classification exactly because it is the hardest form of the problem, and the bottleneck in an end-to-end system: A stronger pure-audio classifier should perform better when paired with metadata than a weaker pure-audio classifier with the same metadata. (For starters, no amount of metadata will improve a total miss.) </p>\n\n<p>Likewise, the best class I've taken on birdsong recognition (for humans) involved lots of studying audio-only flash cards; field identification was much easier with that under my belt. (One favorite example: <a href=\"https://www.allaboutbirds.org/guide/American_Dipper/sounds\">American Dipper</a>, who pretty much only sings near running water, information which is clearly audible but almost certainly doesn't appear in the metadata!)</p>",
      "rawMarkdown": "Well, you can always turn the question on its head... If the metadata tells you a lot about the birds which might be present, can you instead use the audio to predict the metadata? This approach is often used for unsupervised pre-training on large partially labeled datasets: good metadata prediction means the model 'understands' the environment, and should be able to learn the classification task with a smaller collection of hard labels.\n\nMetadata can also be misleading. If the weather is a bit warmer this year, perhaps the migration has started early, which could mean that what the model has learned form other years about geo and date distributions could be off... A system heavily reliant on metadata will tend miss exceptional cases.\n\nSo, our goal is to see how far we can get with 'pure audio' classification exactly because it is the hardest form of the problem, and the bottleneck in an end-to-end system: A stronger pure-audio classifier should perform better when paired with metadata than a weaker pure-audio classifier with the same metadata. (For starters, no amount of metadata will improve a total miss.) \n\nLikewise, the best class I've taken on birdsong recognition (for humans) involved lots of studying audio-only flash cards; field identification was much easier with that under my belt. (One favorite example: [American Dipper](https://www.allaboutbirds.org/guide/American_Dipper/sounds), who pretty much only sings near running water, information which is clearly audible but almost certainly doesn't appear in the metadata!)",
      "votes": null
    },
    {
      "id": "905880",
      "postDate": "06/28/2020 22:32:05",
      "content": "<p>We are still kinda blind without a validation set. The leaderboard feedback is not enough since its not possible to manually analyze what is going wrong.</p>",
      "rawMarkdown": "We are still kinda blind without a validation set. The leaderboard feedback is not enough since its not possible to manually analyze what is going wrong.",
      "votes": null
    },
    {
      "id": "906229",
      "postDate": "06/29/2020 06:25:21",
      "content": "<p>I would agree at least location info like lat/lon would be useful. If the interest here is in conservation, monitoring habitats and changes over time then it is highly unlikely anyone would use data without that information. Migratory birds or near neighbours moving into territories is one thing.  But if there are variations of species on East/West Coast then predicting incorrectly here due to lack of location info seems a bit trivial. <br>\nHowever in time people here may be able to narrow down site 1/2/3 a bit better.</p>",
      "rawMarkdown": "I would agree at least location info like lat/lon would be useful. If the interest here is in conservation, monitoring habitats and changes over time then it is highly unlikely anyone would use data without that information. Migratory birds or near neighbours moving into territories is one thing.  But if there are variations of species on East/West Coast then predicting incorrectly here due to lack of location info seems a bit trivial.   \nHowever in time people here may be able to narrow down site 1/2/3 a bit better.",
      "votes": null
    },
    {
      "id": "995830",
      "postDate": "09/02/2020 20:28:57",
      "content": "<p>Hi Tom. Sorry I am late to the party, and these questions may have already been answered.</p>\n<p>1) test.csv lists audio filename and time intervals. Does it matter whether the submission follows this order exactly, or can the submission be in a different order of audio file?</p>\n<p>2) Suppose the truth of an interval is 'amecro'. Is there any difference in score between identifying it as \"nocall' vs. the wrong bird, like 'annhum'? IOW is a \"nocall\" error scored the same as a wrong id error?</p>\n<p>Thanks for clarifying!</p>",
      "rawMarkdown": "Hi Tom. Sorry I am late to the party, and these questions may have already been answered.\n\n1) test.csv lists audio filename and time intervals. Does it matter whether the submission follows this order exactly, or can the submission be in a different order of audio file?\n\n2) Suppose the truth of an interval is 'amecro'. Is there any difference in score between identifying it as \"nocall' vs. the wrong bird, like 'annhum'? IOW is a \"nocall\" error scored the same as a wrong id error?\n\nThanks for clarifying!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 995830,
      "author_name": "pomo777",
      "author_url": "",
      "post_date": "09/02/2020 20:28:57",
      "content": "<p>Hi Tom. Sorry I am late to the party, and these questions may have already been answered.</p>\n<p>1) test.csv lists audio filename and time intervals. Does it matter whether the submission follows this order exactly, or can the submission be in a different order of audio file?</p>\n<p>2) Suppose the truth of an interval is 'amecro'. Is there any difference in score between identifying it as \"nocall' vs. the wrong bird, like 'annhum'? IOW is a \"nocall\" error scored the same as a wrong id error?</p>\n<p>Thanks for clarifying!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891067,
      "author_name": "lewington",
      "author_url": "",
      "post_date": "06/17/2020 21:45:13",
      "content": "<p>Heya, thanks for helping put on the comp, it looks really cool. Just a quick question about the test set (if I'm allowed to ask) the 150 10 minute recordings: are any of them composed of multiple splices of different recordings, or are they all contiguous audio? </p>",
      "votes": null,
      "replies": [
        {
          "id": 891146,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/18/2020 00:14:37",
          "content": "<p>Hi, Louka! Each recording is a contiguous 10m chunk, without splicing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 893684,
      "author_name": "lewington",
      "author_url": "",
      "post_date": "06/19/2020 19:56:11",
      "content": "<p>Hey again <a href=\"/tomdenton\">@tomdenton</a>, it sounds like you gentlemen have already been trying to solve a bunch of similar problems for a while now. Do you have any general advice about how to approach bird-related audio processing machine learning tasks? Perhaps there are some online resources you found helpful?</p>\n\n<p>Cheers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 893788,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/19/2020 22:22:16",
          "content": "<p>Hi, Louka;\nI'm sure you've seen the great references in this thread already: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158933\">https://www.kaggle.com/c/birdsong-recognition/discussion/158933</a>\nThe papers from BirdCLEF competitions are good reading.</p>\n\n<p>A couple points/problems that I keep in mind, specific to birds and bioacoustics and soundscapes:\n* Augmentation is very important; <a href=\"http://ceur-ws.org/Vol-2125/paper_140.pdf\">Mario Lasseck's paper from 2018 BirdCLEF</a> has some great ablation on different kinds of augmentation.\n* Expanding on that, it's good to keep in mind that the audio environment is very complex, and full of things (like, say, squirrels) that you don't care much about.\n* Some species are highly variable, while others are more constant. Even the more constant songs may vary a bit by geography. Species with a lot of variability sometimes have distinctive notes in otherwise chaotic songs. For a good example of this, you can compare the opening notes of song sparrows to brewer's sparrow.\n* Most species have separate 'songs' and 'calls', with some per-species variability in each.\n* Bird hearing (especially for songbirds) has a much higher time-resolution than human hearing. </p>\n\n<p>And here's how these turn into various not-well-solved problems in the machine learned models:\n* If a particular song is very loud and common, it may appear in the background of many recordings, leading to lower prediction scores for that song. In theory, this is just a calibration problem. In practice, each species has many sounds (with different appropriate calibrations), hidden behind a single class label. I would love to have some better ways to deal with this problem!\n* Warblers seem to be very difficult for the particular models I work with... Their songs are on the shorter-side, and seem to be easily confused. There may be some regional variability issue at play here.\n* Xeno-Canto /probably/ tends towards cleaner, easily identified sounds, since people can ID these. Truncated songs, juveniles (who haven't learned to sing well yet), hybrids, and mimics all add complexity to the sound space, and may not be well represented since they're harder for humans to ID as well.</p>\n\n<p>Ultimately, we still see a big gap between real-world soundscape performance and performance on Xeno-Canto holdout data... And thus the competition!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 893851,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "06/20/2020 02:47:19",
          "content": "<p>Thanks Tom</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 894896,
      "author_name": "lewington",
      "author_url": "",
      "post_date": "06/20/2020 21:55:27",
      "content": "<p>Hey again Tom, sorry to keep bothering you!</p>\n\n<p>I was doing some local cross validation and it occurred to me that I don't know what \"good\" performance looks like. I've done a light literature review and found some prior work, like </p>\n\n<ul>\n<li><a href=\"https://www.acoustics.asn.au/conference_proceedings/AAS2018/papers/p134.pdf\">these guys</a> who got ~65% accuracy classifying 46 species on xeno-canto</li>\n<li><a href=\"https://besjournals.onlinelibrary.wiley.com/doi/full/10.1111/2041-210X.13103\">this crowd</a> who got ~88% AUC score for classifying bird-present/absent on some other datasets</li>\n<li><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4106198/\">these goons</a> who achieved ~80% AUC predicting the presence/absence of 77 bird species, or</li>\n<li><a href=\"https://www.researchgate.net/profile/Dmitry_Konovalov2/publication/335880577_Data-Efficient_Classification_of_Birdcall_Through_Convolutional_Neural_Networks_Transfer_Learning/links/5db38cf692851c577ec35f2f/Data-Efficient-Classification-of-Birdcall-Through-Convolutional-Neural-Networks-Transfer-Learning.pdf\">this mob</a> who achieved ~78% accuracy predicting different numbers of bird classes using some pretty small  xeno-canto datasets</li>\n</ul>\n\n<p>Which I think gives us a decent estimate of expected performance. But, since we have a domain expert like you around I was wondering if you could give us some <em>rough</em> ballpark estimates of:</p>\n\n<ol>\n<li>The (test) accuracy of the best model you think you personally could build if given 90% of the competition training data (at random) and asked to predict the labels of the other 10% (so, since all the test data only ever has one label, a multi-class classification task where each sample has at most one class).</li>\n<li>The public leaderborard score (F1) you think the top performing will achieve on the actual competition task by the end of the competition.</li>\n</ol>\n\n<p>I don't know if you're allowed to make estimates like this, but if you are I think it would be a good yardstick for competitors to gauge whether we're on the right track or not.</p>",
      "votes": null,
      "replies": [
        {
          "id": 897325,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/22/2020 19:22:02",
          "content": "<p>Well, my hope is that the final leaderboard score exceeds my expectations, so I'll avoid biasing the field with a specific prediction. :)</p>\n\n<p>One thing to take note of in lit-review: We tend to see <em>much</em> better model performance when predicting species in held-out Xeno-Canto recordings than in soundscapes, and the higher numbers tend to be what's reported in paper abstracts. High-rated XC samples to be clearer recordings, perhaps with more prototypical sounds. It's very reasonable to use XC holdouts to validate that your models are working, but keep in mind the domain shift problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898264,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "06/23/2020 12:07:28",
          "content": "<p>cheers tom, ill keep that in mind re the literature</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 902655,
      "author_name": "lewington",
      "author_url": "",
      "post_date": "06/26/2020 09:32:49",
      "content": "<p>G'day again Tom, slightly sneaky question this time: was the public test set chosen at random or was it selected manually? And if it's the latter was there some effort made to skew the distribution away from the overall test set distribution?</p>\n\n<p>Obviously only let me know if you think it'll lead to better production models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 905698,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "06/28/2020 18:14:10",
      "content": "<p><a href=\"/tomdenton\">@tomdenton</a>  <a href=\"/stefankahl\">@stefankahl</a> I spent some time with catching up with the recent BirdCLEF papers. I understand that for real world applications you would like to use only easily accessible monophone recordings for training (e.g. XenoCanto). However even real world applications could benefit from additional metadata (e.g. time of recording/season/during the day/nocturnal location of the recording, elevation) leading to useful a priori assumptions.</p>\n\n<p>The competition setting is meaningful and challenging and I will definitely dig deeper. However I feel a bit that we are flying blind. I have a silly analogy, that we were provided english pdf documents for training and we have to build an OCR application for handwriting recognition for an unknown language :)</p>\n\n<p>Any additional help regarding the hidden soundscape test set (similar validation set, metadata) would be really helpful... So far we learned from the forum that recordings were taken in North America with 30kHz.</p>",
      "votes": null,
      "replies": [
        {
          "id": 905796,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/28/2020 19:52:45",
          "content": "<p>Well, you can always turn the question on its head... If the metadata tells you a lot about the birds which might be present, can you instead use the audio to predict the metadata? This approach is often used for unsupervised pre-training on large partially labeled datasets: good metadata prediction means the model 'understands' the environment, and should be able to learn the classification task with a smaller collection of hard labels.</p>\n\n<p>Metadata can also be misleading. If the weather is a bit warmer this year, perhaps the migration has started early, which could mean that what the model has learned form other years about geo and date distributions could be off... A system heavily reliant on metadata will tend miss exceptional cases.</p>\n\n<p>So, our goal is to see how far we can get with 'pure audio' classification exactly because it is the hardest form of the problem, and the bottleneck in an end-to-end system: A stronger pure-audio classifier should perform better when paired with metadata than a weaker pure-audio classifier with the same metadata. (For starters, no amount of metadata will improve a total miss.) </p>\n\n<p>Likewise, the best class I've taken on birdsong recognition (for humans) involved lots of studying audio-only flash cards; field identification was much easier with that under my belt. (One favorite example: <a href=\"https://www.allaboutbirds.org/guide/American_Dipper/sounds\">American Dipper</a>, who pretty much only sings near running water, information which is clearly audible but almost certainly doesn't appear in the metadata!)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 905880,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "06/28/2020 22:32:05",
          "content": "<p>We are still kinda blind without a validation set. The leaderboard feedback is not enough since its not possible to manually analyze what is going wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 906229,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "06/29/2020 06:25:21",
          "content": "<p>I would agree at least location info like lat/lon would be useful. If the interest here is in conservation, monitoring habitats and changes over time then it is highly unlikely anyone would use data without that information. Migratory birds or near neighbours moving into territories is one thing.  But if there are variations of species on East/West Coast then predicting incorrectly here due to lack of location info seems a bit trivial. <br>\nHowever in time people here may be able to narrow down site 1/2/3 a bit better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "888910": "Hi, all! \n\nI'm one of the co-organizers of this competition. I initially got involved in birdsong identification while looking to learn more about deep learning in the audio domain. It's a fascinating problem space, and over the last few years I've learned quite a lot about birds in my quest to build better models.\n\nUnfortunately, my own time for exploring models is limited, so I was very excited by the idea of getting the problem on Kaggle. I competed in a [research-oriented Kaggle competition](https://www.kaggle.com/c/learning-social-circles) six years ago, which I won* by a) not over-fitting on the public leaderboard, and b) challenging some of the assumptions in the problem framing. I hope to see some interesting attacks on the problem here, and see some of my own assumptions challenged!\n\n* - The victory was admittedly by a tiiiiiny margin.",
    "891067": "Heya, thanks for helping put on the comp, it looks really cool. Just a quick question about the test set (if I'm allowed to ask) the 150 10 minute recordings: are any of them composed of multiple splices of different recordings, or are they all contiguous audio?",
    "891146": "Hi, Louka! Each recording is a contiguous 10m chunk, without splicing.",
    "893684": "Hey again @tomdenton, it sounds like you gentlemen have already been trying to solve a bunch of similar problems for a while now. Do you have any general advice about how to approach bird-related audio processing machine learning tasks? Perhaps there are some online resources you found helpful?\n\nCheers.",
    "893788": "Hi, Louka;\nI'm sure you've seen the great references in this thread already: https://www.kaggle.com/c/birdsong-recognition/discussion/158933\nThe papers from BirdCLEF competitions are good reading.\n\nA couple points/problems that I keep in mind, specific to birds and bioacoustics and soundscapes:\n* Augmentation is very important; [Mario Lasseck's paper from 2018 BirdCLEF](http://ceur-ws.org/Vol-2125/paper_140.pdf) has some great ablation on different kinds of augmentation.\n* Expanding on that, it's good to keep in mind that the audio environment is very complex, and full of things (like, say, squirrels) that you don't care much about.\n* Some species are highly variable, while others are more constant. Even the more constant songs may vary a bit by geography. Species with a lot of variability sometimes have distinctive notes in otherwise chaotic songs. For a good example of this, you can compare the opening notes of song sparrows to brewer's sparrow.\n* Most species have separate 'songs' and 'calls', with some per-species variability in each.\n* Bird hearing (especially for songbirds) has a much higher time-resolution than human hearing. \n\nAnd here's how these turn into various not-well-solved problems in the machine learned models:\n* If a particular song is very loud and common, it may appear in the background of many recordings, leading to lower prediction scores for that song. In theory, this is just a calibration problem. In practice, each species has many sounds (with different appropriate calibrations), hidden behind a single class label. I would love to have some better ways to deal with this problem!\n* Warblers seem to be very difficult for the particular models I work with... Their songs are on the shorter-side, and seem to be easily confused. There may be some regional variability issue at play here.\n* Xeno-Canto /probably/ tends towards cleaner, easily identified sounds, since people can ID these. Truncated songs, juveniles (who haven't learned to sing well yet), hybrids, and mimics all add complexity to the sound space, and may not be well represented since they're harder for humans to ID as well.\n\nUltimately, we still see a big gap between real-world soundscape performance and performance on Xeno-Canto holdout data... And thus the competition!",
    "893851": "Thanks Tom",
    "894896": "Hey again Tom, sorry to keep bothering you!\n\nI was doing some local cross validation and it occurred to me that I don't know what \"good\" performance looks like. I've done a light literature review and found some prior work, like \n\n- [these guys](https://www.acoustics.asn.au/conference_proceedings/AAS2018/papers/p134.pdf) who got ~65% accuracy classifying 46 species on xeno-canto\n- [this crowd](https://besjournals.onlinelibrary.wiley.com/doi/full/10.1111/2041-210X.13103) who got ~88% AUC score for classifying bird-present/absent on some other datasets\n- [these goons](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4106198/) who achieved ~80% AUC predicting the presence/absence of 77 bird species, or\n- [this mob](https://www.researchgate.net/profile/Dmitry_Konovalov2/publication/335880577_Data-Efficient_Classification_of_Birdcall_Through_Convolutional_Neural_Networks_Transfer_Learning/links/5db38cf692851c577ec35f2f/Data-Efficient-Classification-of-Birdcall-Through-Convolutional-Neural-Networks-Transfer-Learning.pdf) who achieved ~78% accuracy predicting different numbers of bird classes using some pretty small  xeno-canto datasets\n\nWhich I think gives us a decent estimate of expected performance. But, since we have a domain expert like you around I was wondering if you could give us some *rough* ballpark estimates of:\n\n1. The (test) accuracy of the best model you think you personally could build if given 90% of the competition training data (at random) and asked to predict the labels of the other 10% (so, since all the test data only ever has one label, a multi-class classification task where each sample has at most one class).\n2. The public leaderborard score (F1) you think the top performing will achieve on the actual competition task by the end of the competition.\n\nI don't know if you're allowed to make estimates like this, but if you are I think it would be a good yardstick for competitors to gauge whether we're on the right track or not.",
    "897325": "Well, my hope is that the final leaderboard score exceeds my expectations, so I'll avoid biasing the field with a specific prediction. :)\n\nOne thing to take note of in lit-review: We tend to see *much* better model performance when predicting species in held-out Xeno-Canto recordings than in soundscapes, and the higher numbers tend to be what's reported in paper abstracts. High-rated XC samples to be clearer recordings, perhaps with more prototypical sounds. It's very reasonable to use XC holdouts to validate that your models are working, but keep in mind the domain shift problem.",
    "898264": "cheers tom, ill keep that in mind re the literature",
    "902655": "G'day again Tom, slightly sneaky question this time: was the public test set chosen at random or was it selected manually? And if it's the latter was there some effort made to skew the distribution away from the overall test set distribution?\n\nObviously only let me know if you think it'll lead to better production models.",
    "905698": "tomdenton  @stefankahl I spent some time with catching up with the recent BirdCLEF papers. I understand that for real world applications you would like to use only easily accessible monophone recordings for training (e.g. XenoCanto). However even real world applications could benefit from additional metadata (e.g. time of recording/season/during the day/nocturnal location of the recording, elevation) leading to useful a priori assumptions.\n\nThe competition setting is meaningful and challenging and I will definitely dig deeper. However I feel a bit that we are flying blind. I have a silly analogy, that we were provided english pdf documents for training and we have to build an OCR application for handwriting recognition for an unknown language :)\n\nAny additional help regarding the hidden soundscape test set (similar validation set, metadata) would be really helpful... So far we learned from the forum that recordings were taken in North America with 30kHz.",
    "905796": "Well, you can always turn the question on its head... If the metadata tells you a lot about the birds which might be present, can you instead use the audio to predict the metadata? This approach is often used for unsupervised pre-training on large partially labeled datasets: good metadata prediction means the model 'understands' the environment, and should be able to learn the classification task with a smaller collection of hard labels.\n\nMetadata can also be misleading. If the weather is a bit warmer this year, perhaps the migration has started early, which could mean that what the model has learned form other years about geo and date distributions could be off... A system heavily reliant on metadata will tend miss exceptional cases.\n\nSo, our goal is to see how far we can get with 'pure audio' classification exactly because it is the hardest form of the problem, and the bottleneck in an end-to-end system: A stronger pure-audio classifier should perform better when paired with metadata than a weaker pure-audio classifier with the same metadata. (For starters, no amount of metadata will improve a total miss.) \n\nLikewise, the best class I've taken on birdsong recognition (for humans) involved lots of studying audio-only flash cards; field identification was much easier with that under my belt. (One favorite example: [American Dipper](https://www.allaboutbirds.org/guide/American_Dipper/sounds), who pretty much only sings near running water, information which is clearly audible but almost certainly doesn't appear in the metadata!)",
    "905880": "We are still kinda blind without a validation set. The leaderboard feedback is not enough since its not possible to manually analyze what is going wrong.",
    "906229": "I would agree at least location info like lat/lon would be useful. If the interest here is in conservation, monitoring habitats and changes over time then it is highly unlikely anyone would use data without that information. Migratory birds or near neighbours moving into territories is one thing.  But if there are variations of species on East/West Coast then predicting incorrectly here due to lack of location info seems a bit trivial.   \nHowever in time people here may be able to narrow down site 1/2/3 a bit better.",
    "995830": "Hi Tom. Sorry I am late to the party, and these questions may have already been answered.\n\n1) test.csv lists audio filename and time intervals. Does it matter whether the submission follows this order exactly, or can the submission be in a different order of audio file?\n\n2) Suppose the truth of an interval is 'amecro'. Is there any difference in score between identifying it as \"nocall' vs. the wrong bird, like 'annhum'? IOW is a \"nocall\" error scored the same as a wrong id error?\n\nThanks for clarifying!"
  },
  "source": "meta"
}