{
  "id": 159123,
  "title": "Few questions about test data",
  "url": "/competitions/birdsong-recognition/discussion/159123",
  "author_name": "",
  "post_date": "2020-06-16T14:03:00.685876200Z",
  "votes": 37,
  "comment_count": 13,
  "views": 0,
  "content": "<p>p.s. add a question mark (?) in front of all of these.</p>\n\n<p>1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.\n2) test.csv has the field <code>second</code> ending the time window so prediction should be for time frame 0 to this second.\n3) if that's the case should we add birds that were detected in subset of time frame or just unique. (the sample submission does make not sense to me but its just a sample so ..)\n4) The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval  so approximately total  2 * (50 * 10 * 60 / 5)  + (50) = 12050 hidden test rows.\n5) where will the test recording files will be stored and what format should we expect I presume we will get access to 150 files instead of 12050 files.\n6) what does <code>_093000</code> and <code>pt540</code> represent in example file names. \nIf anyone has a clear understanding about any of these please help clearing the doubts for everyone. </p>",
  "messages": [
    {
      "id": "888689",
      "postDate": "06/16/2020 14:03:00",
      "content": "<p>p.s. add a question mark (?) in front of all of these.</p>\n\n<p>1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.\n2) test.csv has the field <code>second</code> ending the time window so prediction should be for time frame 0 to this second.\n3) if that's the case should we add birds that were detected in subset of time frame or just unique. (the sample submission does make not sense to me but its just a sample so ..)\n4) The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval  so approximately total  2 * (50 * 10 * 60 / 5)  + (50) = 12050 hidden test rows.\n5) where will the test recording files will be stored and what format should we expect I presume we will get access to 150 files instead of 12050 files.\n6) what does <code>_093000</code> and <code>pt540</code> represent in example file names. \nIf anyone has a clear understanding about any of these please help clearing the doubts for everyone. </p>",
      "rawMarkdown": "p.s. add a question mark (?) in front of all of these.\n\n\n1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.\n2) test.csv has the field `second` ending the time window so prediction should be for time frame 0 to this second.\n3) if that's the case should we add birds that were detected in subset of time frame or just unique. (the sample submission does make not sense to me but its just a sample so ..)\n4) The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval  so approximately total  2 * (50 * 10 * 60 / 5)  + (50) = 12050 hidden test rows.\n5) where will the test recording files will be stored and what format should we expect I presume we will get access to 150 files instead of 12050 files.\n6) what does `_093000` and `pt540` represent in example file names. \nIf anyone has a clear understanding about any of these please help clearing the doubts for everyone.",
      "votes": null
    },
    {
      "id": "888780",
      "postDate": "06/16/2020 15:00:24",
      "content": "<p>I also have some questions and made a thread ( <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158987\">https://www.kaggle.com/c/birdsong-recognition/discussion/158987</a> ), and it seems we have some questions in common.</p>\n\n<p>I asked about those in the thread above to the host, so when I get answer(s) from the host I'll also share that here.</p>",
      "rawMarkdown": "I also have some questions and made a thread ( https://www.kaggle.com/c/birdsong-recognition/discussion/158987 ), and it seems we have some questions in common.\n\nI asked about those in the thread above to the host, so when I get answer(s) from the host I'll also share that here.",
      "votes": null
    },
    {
      "id": "888796",
      "postDate": "06/16/2020 15:08:06",
      "content": "<ol>\n<li>Correct. Training set is for recognizing individual bird sounds, test data is for recognizing all (zero or more) birds. \nIf you look at the Competition description, this <strong>domain mismatch</strong> is the reason for the competition's existence.</li>\n<li><code>second</code> is the end time for each frame, and since each frame interval is of 5 seconds (for site 1,2), start time should be <code>second</code> - 5.</li>\n<li>Just focus on the subset. You can look at the example rows in <code>example_test_audio_summary.csv</code>. The same bird is detected in multiple intervals for the same audio-file.</li>\n<li>Can't say that for certain, they don't mention the distrubution anywhere.</li>\n<li>This is a good question. They don't mention anywhere where the test  audio files will be stored. If I had to guess, would be a folder called <code>test_audio/</code>.\nEdit: Looking at the Data description closely, they have mentioned that there is a hidden <code>test_audio</code> folder.</li>\n<li>Doesn't matter 😜 </li>\n</ol>",
      "rawMarkdown": "1. Correct. Training set is for recognizing individual bird sounds, test data is for recognizing all (zero or more) birds. \nIf you look at the Competition description, this **domain mismatch** is the reason for the competition's existence.\n2. `second` is the end time for each frame, and since each frame interval is of 5 seconds (for site 1,2), start time should be `second` - 5.\n3.  Just focus on the subset. You can look at the example rows in `example_test_audio_summary.csv`. The same bird is detected in multiple intervals for the same audio-file.\n4. Can't say that for certain, they don't mention the distrubution anywhere.\n5. This is a good question. They don't mention anywhere where the test  audio files will be stored. If I had to guess, would be a folder called `test_audio/`.\nEdit: Looking at the Data description closely, they have mentioned that there is a hidden `test_audio` folder.\n6. Doesn't matter 😜",
      "votes": null
    },
    {
      "id": "890189",
      "postDate": "06/17/2020 11:22:40",
      "content": "<p>Update Jun.19.2020 <br>\nBased on <a href=\"/stefankahl\">@stefankahl</a>  comment, <code>secondary_labels</code> will be better than <code>background</code>. But both are almost same and week labels regarding label quality.</p>\n\n<p>I have posted the kernel, how to extract background labels from train.csv. If you are interested, please take a look.</p>\n\n<p><strong>MlutiClass to MuiltiLabel</strong>\n<a href=\"https://www.kaggle.com/maxwell110/mluticlass-to-muiltilabel\">https://www.kaggle.com/maxwell110/mluticlass-to-muiltilabel</a></p>\n\n<hr>\n\n<p>&gt; 1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.</p>\n\n<p>IMO, train dataset (train.csv) has other label information in some instances (not so many). <code>background</code> column has additional label information in some recorded sounds, I think.\nFor example, here is top 10 frequent backgrounds.</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F852261ac4329d5f94ea37eba99337cfc%2Fpng.PNG?generation=1592395619817766&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F852261ac4329d5f94ea37eba99337cfc%2Fpng.PNG?generation=1592395619817766&amp;alt=media</a> =700x70)</p>\n\n<p>| backgrounds | frequency |\n| --- | --- |\n|Red-winged Blackbird (Agelaius phoeniceus) |128|\n|Northern Cardinal (Cardinalis cardinalis)  |77|\n|American Robin (Turdus migratorius)    |75|\n|House Finch (Haemorhous mexicanus)|    49|\n|Blue Jay (Cyanocitta cristata) |46|\n|Western Meadowlark (Sturnella neglecta)    |45|\n|Red-eyed Vireo (Vireo olivaceus)   |44|\n|American Crow (Corvus brachyrhynchos)  |43|\n|Black-capped Chickadee (Poecile atricapillus)  |42|\n|Canada Goose (Branta canadensis)   |40|</p>\n\n<p>I'm not sure all background types are in target label types, but most frequent red-winged blackbird is included in target labels. This means at least Red-winged Blackbird is also a target label.</p>\n\n<p>```\n(train.species == 'Red-winged Blackbird').sum()</p>\n\n<h1>100</h1>\n\n<p>```</p>\n\n<p>So I guess we can convert this multi-class task into multi-labeled task without any audio data preprocessing like mixing audios.\nAs long as I have checked in xeno-canto pages, <a href=\"https://www.xeno-canto.org/help/FAQ\">https://www.xeno-canto.org/help/FAQ</a>, background means other bird sounds.</p>\n\n<p>&gt; What is a 'soundscape' recording?\nIn 2013, we added the ability to upload 'soundscape' recordings on xeno-canto. The intention is to allow people to share longer recordings where more than a single species is vocalizing. The prototypical example is a recording of a dawn chorus, where many species are vocalizing together. In such cases, attempting to extract a single-species portion of the recording will often diminish its value.\n&gt; Soundscape recordings should generally be at least a minute long, and have two or more species identified in the background species list. Occasional uploads of non-avian species are usually acceptable, but we discourage large numbers of these recordings. Eventually, we may add proper taxonomic support to allow people to share vocalizations of different classes of animals, but for now the focus is on birds.</p>\n\n<p>If I'm wrong (misunderstanding the meaning of background), please let me know.</p>\n\n<p>Example) An audio track shown in the picture below has multiple labels(maybe more 5 labels in background).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fe4be134fbe88bdb5b6140b1defa3a2c3%2Fpng.PNG?generation=1592396532899667&amp;alt=media\" alt=\"\"></p>\n\n<p>Anyway I want more information about this meta data and the way to submit...</p>",
      "rawMarkdown": "Update Jun.19.2020  \nBased on @stefankahl  comment, `secondary_labels` will be better than `background`. But both are almost same and week labels regarding label quality.\n\nI have posted the kernel, how to extract background labels from train.csv. If you are interested, please take a look.\n\n**MlutiClass to MuiltiLabel**\nhttps://www.kaggle.com/maxwell110/mluticlass-to-muiltilabel\n\n---\n\n&gt; 1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.\n\nIMO, train dataset (train.csv) has other label information in some instances (not so many). `background` column has additional label information in some recorded sounds, I think.\nFor example, here is top 10 frequent backgrounds.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F852261ac4329d5f94ea37eba99337cfc%2Fpng.PNG?generation=1592395619817766&amp;alt=media =700x70)\n\n\n| backgrounds | frequency |\n| --- | --- |\n|Red-winged Blackbird (Agelaius phoeniceus)\t|128|\n|Northern Cardinal (Cardinalis cardinalis)\t|77|\n|American Robin (Turdus migratorius)\t|75|\n|House Finch (Haemorhous mexicanus)|\t49|\n|Blue Jay (Cyanocitta cristata)\t|46|\n|Western Meadowlark (Sturnella neglecta)\t|45|\n|Red-eyed Vireo (Vireo olivaceus)\t|44|\n|American Crow (Corvus brachyrhynchos)\t|43|\n|Black-capped Chickadee (Poecile atricapillus)\t|42|\n|Canada Goose (Branta canadensis)\t|40|\n\nI'm not sure all background types are in target label types, but most frequent red-winged blackbird is included in target labels. This means at least Red-winged Blackbird is also a target label.\n\n```\n(train.species == 'Red-winged Blackbird').sum()\n# 100\n```\n\nSo I guess we can convert this multi-class task into multi-labeled task without any audio data preprocessing like mixing audios.\nAs long as I have checked in xeno-canto pages, https://www.xeno-canto.org/help/FAQ, background means other bird sounds.\n\n&gt; What is a 'soundscape' recording?\nIn 2013, we added the ability to upload 'soundscape' recordings on xeno-canto. The intention is to allow people to share longer recordings where more than a single species is vocalizing. The prototypical example is a recording of a dawn chorus, where many species are vocalizing together. In such cases, attempting to extract a single-species portion of the recording will often diminish its value.\n&gt; Soundscape recordings should generally be at least a minute long, and have two or more species identified in the background species list. Occasional uploads of non-avian species are usually acceptable, but we discourage large numbers of these recordings. Eventually, we may add proper taxonomic support to allow people to share vocalizations of different classes of animals, but for now the focus is on birds.\n\nIf I'm wrong (misunderstanding the meaning of background), please let me know.\n\n\nExample) An audio track shown in the picture below has multiple labels(maybe more 5 labels in background).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fe4be134fbe88bdb5b6140b1defa3a2c3%2Fpng.PNG?generation=1592396532899667&amp;alt=media)\n\n\n\nAnyway I want more information about this meta data and the way to submit...",
      "votes": null
    },
    {
      "id": "890499",
      "postDate": "06/17/2020 14:34:09",
      "content": "<p>I got a replay from kaggle team\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158987#890473\">https://www.kaggle.com/c/birdsong-recognition/discussion/158987#890473</a></p>\n\n<p>What <a href=\"/dhruvrnaik\">@dhruvrnaik</a>  wrote is the answer, and as for 5) we will have <code>/kaggle/input/birdsong-recognition/test_audio</code> folder at kernel re-running phase. Inside that we have around 150 audio clips whose names correspond to the ids in <code>audio_id</code> in <code>test.csv</code>.</p>",
      "rawMarkdown": "I got a replay from kaggle team\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/158987#890473\n\nWhat @dhruvrnaik  wrote is the answer, and as for 5) we will have `/kaggle/input/birdsong-recognition/test_audio` folder at kernel re-running phase. Inside that we have around 150 audio clips whose names correspond to the ids in `audio_id` in `test.csv`.",
      "votes": null
    },
    {
      "id": "890675",
      "postDate": "06/17/2020 16:26:25",
      "content": "<p>Let me try to answer some of the questions in a more general way. The difficulties that we are facing when developing a classifier for bird sounds is the shift in acoustic domains between clean, focal recordings (e.g., Xeno-canto files) and soundscapes which are typically recorded with omnidirectional recording equipment. But there is more to it. Overlapping vocalizations are a major issue and Xeno-canto recordings may or may not contain background species and they may or may not have an appropriate label (typically primary and secondary labels in the metadata). The soundscapes however will have a high amount of simultaneous vocalizations - sometimes 5 or more for one 5-second interval. The test set includes 150 of these soundscapes recorded at 3 sites; all are of 10 min. duration. But we only have call-level annotations for 2 sites (they are hard to come by). Some soundscapes do not contain a single vocalization, others contain multiple hundreds.</p>\n\n<p>From my perspective, it's a two-step process in developing a classifier. First, train a classifier on Xeno-canto recordings an evaluate on a test split for maximum performance. Secondly, fine-tune the trained classifier to achieve optimal performance on soundscape data. We did not distribute soundscape data for validation to prevent participants from training on soundscape data - which would be valid according to the rules but not very practical from a application point of view: It would be too difficult to annotate hours of soundscape data before we can deploy a system to a new recording site. Trained systems should generalize well enough to cope with different acoustic environments (hence the 3 recording sites).</p>\n\n<p>I am not too familiar with the test set characteristics when it comes to Kaggle-specific modifications, so our other host might be able to answer these questions in more detail. But I'd be happy to join a discussion about methodology :)</p>",
      "rawMarkdown": "Let me try to answer some of the questions in a more general way. The difficulties that we are facing when developing a classifier for bird sounds is the shift in acoustic domains between clean, focal recordings (e.g., Xeno-canto files) and soundscapes which are typically recorded with omnidirectional recording equipment. But there is more to it. Overlapping vocalizations are a major issue and Xeno-canto recordings may or may not contain background species and they may or may not have an appropriate label (typically primary and secondary labels in the metadata). The soundscapes however will have a high amount of simultaneous vocalizations - sometimes 5 or more for one 5-second interval. The test set includes 150 of these soundscapes recorded at 3 sites; all are of 10 min. duration. But we only have call-level annotations for 2 sites (they are hard to come by). Some soundscapes do not contain a single vocalization, others contain multiple hundreds.\n\nFrom my perspective, it's a two-step process in developing a classifier. First, train a classifier on Xeno-canto recordings an evaluate on a test split for maximum performance. Secondly, fine-tune the trained classifier to achieve optimal performance on soundscape data. We did not distribute soundscape data for validation to prevent participants from training on soundscape data - which would be valid according to the rules but not very practical from a application point of view: It would be too difficult to annotate hours of soundscape data before we can deploy a system to a new recording site. Trained systems should generalize well enough to cope with different acoustic environments (hence the 3 recording sites).\n\nI am not too familiar with the test set characteristics when it comes to Kaggle-specific modifications, so our other host might be able to answer these questions in more detail. But I'd be happy to join a discussion about methodology :)",
      "votes": null
    },
    {
      "id": "890965",
      "postDate": "06/17/2020 19:42:27",
      "content": "<p><a href=\"/stefankahl\">@stefankahl</a> \nThank you for your reply. Please let me confirm one thing.    </p>\n\n<p>train.csv file contains <code>secondary_labels</code> and <code>background</code> columns. I suppose both are essentially same if I preprocess <code>background</code> columns (remove unnecessary labels for this task). Am I right? If I'm not wrong, although this labels may sometime inappropriate, we can try to use this information to build soundscapes-level(multi label) models.</p>\n\n<p>Thank you :-)</p>",
      "rawMarkdown": "stefankahl \nThank you for your reply. Please let me confirm one thing.    \n\ntrain.csv file contains `secondary_labels` and `background` columns. I suppose both are essentially same if I preprocess `background` columns (remove unnecessary labels for this task). Am I right? If I'm not wrong, although this labels may sometime inappropriate, we can try to use this information to build soundscapes-level(multi label) models.\n\nThank you :-)",
      "votes": null
    },
    {
      "id": "891467",
      "postDate": "06/18/2020 07:38:19",
      "content": "<p>Secondary labels and background should be mostly identical but secondary labels follow the eBird taxonomy (Clement's list) and background annotations follow the Xeno-canto taxonomy (IOC World list). Be aware that not all species mentioned in the secondary label column might be part of the training data. For training, you can focus on primary and secondary labels. For a submission, you should use the ebird code. You should definitely consider using secondary labels to validate your multi-class classifier, but (again) be aware that these labels are only weak labels with no timestamp.</p>",
      "rawMarkdown": "Secondary labels and background should be mostly identical but secondary labels follow the eBird taxonomy (Clement's list) and background annotations follow the Xeno-canto taxonomy (IOC World list). Be aware that not all species mentioned in the secondary label column might be part of the training data. For training, you can focus on primary and secondary labels. For a submission, you should use the ebird code. You should definitely consider using secondary labels to validate your multi-class classifier, but (again) be aware that these labels are only weak labels with no timestamp.",
      "votes": null
    },
    {
      "id": "901615",
      "postDate": "06/25/2020 15:48:58",
      "content": "<p>Dhruv's answers for 1-3 are correct. To fill in a couple more...\n4) This is the right ballpark, yes; ~12k test rows.\n5) The test audio is in a directory called test_audio; there are a couple public notebooks now which have good examples of iterating over the audio and creating a successful submission.\n6) These are metadata specific to the particular project that the example files came from. That project is completely separate from the test data, and the test data filenames have been changed to ensure no metadata leakage. So, Dhruv is correct here, as well.</p>",
      "rawMarkdown": "Dhruv's answers for 1-3 are correct. To fill in a couple more...\n4) This is the right ballpark, yes; ~12k test rows.\n5) The test audio is in a directory called test_audio; there are a couple public notebooks now which have good examples of iterating over the audio and creating a successful submission.\n6) These are metadata specific to the particular project that the example files came from. That project is completely separate from the test data, and the test data filenames have been changed to ensure no metadata leakage. So, Dhruv is correct here, as well.",
      "votes": null
    },
    {
      "id": "901865",
      "postDate": "06/25/2020 18:51:37",
      "content": "<p><a href=\"/stefankahl\">@stefankahl</a>  Thanks for the reply.\nI have a doubt about the training data. For each file we can hear only one bird but we have also see only one \"time\" feature (which means we can hear the bird at that moment if I am not wrong) but shouldn't we see multiple time for one bird ? I have try to listen few audios and I have the feeling we can hear let's say 3 times the bird during the 10/15 seconds but in the csv file, we will see only one time. Is it normal ?</p>",
      "rawMarkdown": "stefankahl  Thanks for the reply.\nI have a doubt about the training data. For each file we can hear only one bird but we have also see only one \"time\" feature (which means we can hear the bird at that moment if I am not wrong) but shouldn't we see multiple time for one bird ? I have try to listen few audios and I have the feeling we can hear let's say 3 times the bird during the 10/15 seconds but in the csv file, we will see only one time. Is it normal ?",
      "votes": null
    },
    {
      "id": "901879",
      "postDate": "06/25/2020 19:10:17",
      "content": "<p>The time field in the training metadata is a time-of-day supplied by the recordist. (often rounded to the nearest half hour...) The training labels are 'weak' in the sense that the particular bird should exist somewhere in the recording, but it may be sparse, and may be mixed with other vocalizations.</p>",
      "rawMarkdown": "The time field in the training metadata is a time-of-day supplied by the recordist. (often rounded to the nearest half hour...) The training labels are 'weak' in the sense that the particular bird should exist somewhere in the recording, but it may be sparse, and may be mixed with other vocalizations.",
      "votes": null
    },
    {
      "id": "971270",
      "postDate": "08/15/2020 10:46:23",
      "content": "<blockquote>\n  <p>The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval so approximately total 2 * (50 * 10 * 60 / 5) + (50) = 12050 hidden test rows.</p>\n</blockquote>\n<p>Is it known that each site has roughly the same number of recordings?</p>",
      "rawMarkdown": "> The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval so approximately total 2 * (50 * 10 * 60 / 5) + (50) = 12050 hidden test rows.\n\nIs it known that each site has roughly the same number of recordings?",
      "votes": null
    },
    {
      "id": "971272",
      "postDate": "08/15/2020 10:48:36",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> for your responses, clear communication really makes the competition much more enjoyable</p>",
      "rawMarkdown": "Thank you so much @tomdenton for your responses, clear communication really makes the competition much more enjoyable",
      "votes": null
    },
    {
      "id": "1001052",
      "postDate": "09/07/2020 03:33:15",
      "content": "<p>Great work sir. Upload more such.</p>",
      "rawMarkdown": "Great work sir. Upload more such.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 971270,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "08/15/2020 10:46:23",
      "content": "<blockquote>\n  <p>The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval so approximately total 2 * (50 * 10 * 60 / 5) + (50) = 12050 hidden test rows.</p>\n</blockquote>\n<p>Is it known that each site has roughly the same number of recordings?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 888780,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "06/16/2020 15:00:24",
      "content": "<p>I also have some questions and made a thread ( <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158987\">https://www.kaggle.com/c/birdsong-recognition/discussion/158987</a> ), and it seems we have some questions in common.</p>\n\n<p>I asked about those in the thread above to the host, so when I get answer(s) from the host I'll also share that here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 888796,
      "author_name": "dhruvrnaik",
      "author_url": "",
      "post_date": "06/16/2020 15:08:06",
      "content": "<ol>\n<li>Correct. Training set is for recognizing individual bird sounds, test data is for recognizing all (zero or more) birds. \nIf you look at the Competition description, this <strong>domain mismatch</strong> is the reason for the competition's existence.</li>\n<li><code>second</code> is the end time for each frame, and since each frame interval is of 5 seconds (for site 1,2), start time should be <code>second</code> - 5.</li>\n<li>Just focus on the subset. You can look at the example rows in <code>example_test_audio_summary.csv</code>. The same bird is detected in multiple intervals for the same audio-file.</li>\n<li>Can't say that for certain, they don't mention the distrubution anywhere.</li>\n<li>This is a good question. They don't mention anywhere where the test  audio files will be stored. If I had to guess, would be a folder called <code>test_audio/</code>.\nEdit: Looking at the Data description closely, they have mentioned that there is a hidden <code>test_audio</code> folder.</li>\n<li>Doesn't matter 😜 </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 901615,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/25/2020 15:48:58",
          "content": "<p>Dhruv's answers for 1-3 are correct. To fill in a couple more...\n4) This is the right ballpark, yes; ~12k test rows.\n5) The test audio is in a directory called test_audio; there are a couple public notebooks now which have good examples of iterating over the audio and creating a successful submission.\n6) These are metadata specific to the particular project that the example files came from. That project is completely separate from the test data, and the test data filenames have been changed to ensure no metadata leakage. So, Dhruv is correct here, as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971272,
          "author_name": "marcogorelli",
          "author_url": "",
          "post_date": "08/15/2020 10:48:36",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> for your responses, clear communication really makes the competition much more enjoyable</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 890189,
      "author_name": "maxwell110",
      "author_url": "",
      "post_date": "06/17/2020 11:22:40",
      "content": "<p>Update Jun.19.2020 <br>\nBased on <a href=\"/stefankahl\">@stefankahl</a>  comment, <code>secondary_labels</code> will be better than <code>background</code>. But both are almost same and week labels regarding label quality.</p>\n\n<p>I have posted the kernel, how to extract background labels from train.csv. If you are interested, please take a look.</p>\n\n<p><strong>MlutiClass to MuiltiLabel</strong>\n<a href=\"https://www.kaggle.com/maxwell110/mluticlass-to-muiltilabel\">https://www.kaggle.com/maxwell110/mluticlass-to-muiltilabel</a></p>\n\n<hr>\n\n<p>&gt; 1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.</p>\n\n<p>IMO, train dataset (train.csv) has other label information in some instances (not so many). <code>background</code> column has additional label information in some recorded sounds, I think.\nFor example, here is top 10 frequent backgrounds.</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F852261ac4329d5f94ea37eba99337cfc%2Fpng.PNG?generation=1592395619817766&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F852261ac4329d5f94ea37eba99337cfc%2Fpng.PNG?generation=1592395619817766&amp;alt=media</a> =700x70)</p>\n\n<p>| backgrounds | frequency |\n| --- | --- |\n|Red-winged Blackbird (Agelaius phoeniceus) |128|\n|Northern Cardinal (Cardinalis cardinalis)  |77|\n|American Robin (Turdus migratorius)    |75|\n|House Finch (Haemorhous mexicanus)|    49|\n|Blue Jay (Cyanocitta cristata) |46|\n|Western Meadowlark (Sturnella neglecta)    |45|\n|Red-eyed Vireo (Vireo olivaceus)   |44|\n|American Crow (Corvus brachyrhynchos)  |43|\n|Black-capped Chickadee (Poecile atricapillus)  |42|\n|Canada Goose (Branta canadensis)   |40|</p>\n\n<p>I'm not sure all background types are in target label types, but most frequent red-winged blackbird is included in target labels. This means at least Red-winged Blackbird is also a target label.</p>\n\n<p>```\n(train.species == 'Red-winged Blackbird').sum()</p>\n\n<h1>100</h1>\n\n<p>```</p>\n\n<p>So I guess we can convert this multi-class task into multi-labeled task without any audio data preprocessing like mixing audios.\nAs long as I have checked in xeno-canto pages, <a href=\"https://www.xeno-canto.org/help/FAQ\">https://www.xeno-canto.org/help/FAQ</a>, background means other bird sounds.</p>\n\n<p>&gt; What is a 'soundscape' recording?\nIn 2013, we added the ability to upload 'soundscape' recordings on xeno-canto. The intention is to allow people to share longer recordings where more than a single species is vocalizing. The prototypical example is a recording of a dawn chorus, where many species are vocalizing together. In such cases, attempting to extract a single-species portion of the recording will often diminish its value.\n&gt; Soundscape recordings should generally be at least a minute long, and have two or more species identified in the background species list. Occasional uploads of non-avian species are usually acceptable, but we discourage large numbers of these recordings. Eventually, we may add proper taxonomic support to allow people to share vocalizations of different classes of animals, but for now the focus is on birds.</p>\n\n<p>If I'm wrong (misunderstanding the meaning of background), please let me know.</p>\n\n<p>Example) An audio track shown in the picture below has multiple labels(maybe more 5 labels in background).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fe4be134fbe88bdb5b6140b1defa3a2c3%2Fpng.PNG?generation=1592396532899667&amp;alt=media\" alt=\"\"></p>\n\n<p>Anyway I want more information about this meta data and the way to submit...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 890499,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "06/17/2020 14:34:09",
      "content": "<p>I got a replay from kaggle team\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158987#890473\">https://www.kaggle.com/c/birdsong-recognition/discussion/158987#890473</a></p>\n\n<p>What <a href=\"/dhruvrnaik\">@dhruvrnaik</a>  wrote is the answer, and as for 5) we will have <code>/kaggle/input/birdsong-recognition/test_audio</code> folder at kernel re-running phase. Inside that we have around 150 audio clips whose names correspond to the ids in <code>audio_id</code> in <code>test.csv</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 890675,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "06/17/2020 16:26:25",
      "content": "<p>Let me try to answer some of the questions in a more general way. The difficulties that we are facing when developing a classifier for bird sounds is the shift in acoustic domains between clean, focal recordings (e.g., Xeno-canto files) and soundscapes which are typically recorded with omnidirectional recording equipment. But there is more to it. Overlapping vocalizations are a major issue and Xeno-canto recordings may or may not contain background species and they may or may not have an appropriate label (typically primary and secondary labels in the metadata). The soundscapes however will have a high amount of simultaneous vocalizations - sometimes 5 or more for one 5-second interval. The test set includes 150 of these soundscapes recorded at 3 sites; all are of 10 min. duration. But we only have call-level annotations for 2 sites (they are hard to come by). Some soundscapes do not contain a single vocalization, others contain multiple hundreds.</p>\n\n<p>From my perspective, it's a two-step process in developing a classifier. First, train a classifier on Xeno-canto recordings an evaluate on a test split for maximum performance. Secondly, fine-tune the trained classifier to achieve optimal performance on soundscape data. We did not distribute soundscape data for validation to prevent participants from training on soundscape data - which would be valid according to the rules but not very practical from a application point of view: It would be too difficult to annotate hours of soundscape data before we can deploy a system to a new recording site. Trained systems should generalize well enough to cope with different acoustic environments (hence the 3 recording sites).</p>\n\n<p>I am not too familiar with the test set characteristics when it comes to Kaggle-specific modifications, so our other host might be able to answer these questions in more detail. But I'd be happy to join a discussion about methodology :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 890965,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "06/17/2020 19:42:27",
          "content": "<p><a href=\"/stefankahl\">@stefankahl</a> \nThank you for your reply. Please let me confirm one thing.    </p>\n\n<p>train.csv file contains <code>secondary_labels</code> and <code>background</code> columns. I suppose both are essentially same if I preprocess <code>background</code> columns (remove unnecessary labels for this task). Am I right? If I'm not wrong, although this labels may sometime inappropriate, we can try to use this information to build soundscapes-level(multi label) models.</p>\n\n<p>Thank you :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891467,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "06/18/2020 07:38:19",
          "content": "<p>Secondary labels and background should be mostly identical but secondary labels follow the eBird taxonomy (Clement's list) and background annotations follow the Xeno-canto taxonomy (IOC World list). Be aware that not all species mentioned in the secondary label column might be part of the training data. For training, you can focus on primary and secondary labels. For a submission, you should use the ebird code. You should definitely consider using secondary labels to validate your multi-class classifier, but (again) be aware that these labels are only weak labels with no timestamp.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 901865,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "06/25/2020 18:51:37",
          "content": "<p><a href=\"/stefankahl\">@stefankahl</a>  Thanks for the reply.\nI have a doubt about the training data. For each file we can hear only one bird but we have also see only one \"time\" feature (which means we can hear the bird at that moment if I am not wrong) but shouldn't we see multiple time for one bird ? I have try to listen few audios and I have the feeling we can hear let's say 3 times the bird during the 10/15 seconds but in the csv file, we will see only one time. Is it normal ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 901879,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "06/25/2020 19:10:17",
          "content": "<p>The time field in the training metadata is a time-of-day supplied by the recordist. (often rounded to the nearest half hour...) The training labels are 'weak' in the sense that the particular bird should exist somewhere in the recording, but it may be sparse, and may be mixed with other vocalizations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1001052,
      "author_name": "taritdhar",
      "author_url": "",
      "post_date": "09/07/2020 03:33:15",
      "content": "<p>Great work sir. Upload more such.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "888689": "p.s. add a question mark (?) in front of all of these.\n\n\n1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.\n2) test.csv has the field `second` ending the time window so prediction should be for time frame 0 to this second.\n3) if that's the case should we add birds that were detected in subset of time frame or just unique. (the sample submission does make not sense to me but its just a sample so ..)\n4) The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval  so approximately total  2 * (50 * 10 * 60 / 5)  + (50) = 12050 hidden test rows.\n5) where will the test recording files will be stored and what format should we expect I presume we will get access to 150 files instead of 12050 files.\n6) what does `_093000` and `pt540` represent in example file names. \nIf anyone has a clear understanding about any of these please help clearing the doubts for everyone.",
    "888780": "I also have some questions and made a thread ( https://www.kaggle.com/c/birdsong-recognition/discussion/158987 ), and it seems we have some questions in common.\n\nI asked about those in the thread above to the host, so when I get answer(s) from the host I'll also share that here.",
    "888796": "1. Correct. Training set is for recognizing individual bird sounds, test data is for recognizing all (zero or more) birds. \nIf you look at the Competition description, this **domain mismatch** is the reason for the competition's existence.\n2. `second` is the end time for each frame, and since each frame interval is of 5 seconds (for site 1,2), start time should be `second` - 5.\n3.  Just focus on the subset. You can look at the example rows in `example_test_audio_summary.csv`. The same bird is detected in multiple intervals for the same audio-file.\n4. Can't say that for certain, they don't mention the distrubution anywhere.\n5. This is a good question. They don't mention anywhere where the test  audio files will be stored. If I had to guess, would be a folder called `test_audio/`.\nEdit: Looking at the Data description closely, they have mentioned that there is a hidden `test_audio` folder.\n6. Doesn't matter 😜",
    "890189": "Update Jun.19.2020  \nBased on @stefankahl  comment, `secondary_labels` will be better than `background`. But both are almost same and week labels regarding label quality.\n\nI have posted the kernel, how to extract background labels from train.csv. If you are interested, please take a look.\n\n**MlutiClass to MuiltiLabel**\nhttps://www.kaggle.com/maxwell110/mluticlass-to-muiltilabel\n\n---\n\n&gt; 1) Train dataset has one label per recording but in test set we are expected to predict multiple or zero if present.\n\nIMO, train dataset (train.csv) has other label information in some instances (not so many). `background` column has additional label information in some recorded sounds, I think.\nFor example, here is top 10 frequent backgrounds.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F852261ac4329d5f94ea37eba99337cfc%2Fpng.PNG?generation=1592395619817766&amp;alt=media =700x70)\n\n\n| backgrounds | frequency |\n| --- | --- |\n|Red-winged Blackbird (Agelaius phoeniceus)\t|128|\n|Northern Cardinal (Cardinalis cardinalis)\t|77|\n|American Robin (Turdus migratorius)\t|75|\n|House Finch (Haemorhous mexicanus)|\t49|\n|Blue Jay (Cyanocitta cristata)\t|46|\n|Western Meadowlark (Sturnella neglecta)\t|45|\n|Red-eyed Vireo (Vireo olivaceus)\t|44|\n|American Crow (Corvus brachyrhynchos)\t|43|\n|Black-capped Chickadee (Poecile atricapillus)\t|42|\n|Canada Goose (Branta canadensis)\t|40|\n\nI'm not sure all background types are in target label types, but most frequent red-winged blackbird is included in target labels. This means at least Red-winged Blackbird is also a target label.\n\n```\n(train.species == 'Red-winged Blackbird').sum()\n# 100\n```\n\nSo I guess we can convert this multi-class task into multi-labeled task without any audio data preprocessing like mixing audios.\nAs long as I have checked in xeno-canto pages, https://www.xeno-canto.org/help/FAQ, background means other bird sounds.\n\n&gt; What is a 'soundscape' recording?\nIn 2013, we added the ability to upload 'soundscape' recordings on xeno-canto. The intention is to allow people to share longer recordings where more than a single species is vocalizing. The prototypical example is a recording of a dawn chorus, where many species are vocalizing together. In such cases, attempting to extract a single-species portion of the recording will often diminish its value.\n&gt; Soundscape recordings should generally be at least a minute long, and have two or more species identified in the background species list. Occasional uploads of non-avian species are usually acceptable, but we discourage large numbers of these recordings. Eventually, we may add proper taxonomic support to allow people to share vocalizations of different classes of animals, but for now the focus is on birds.\n\nIf I'm wrong (misunderstanding the meaning of background), please let me know.\n\n\nExample) An audio track shown in the picture below has multiple labels(maybe more 5 labels in background).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fe4be134fbe88bdb5b6140b1defa3a2c3%2Fpng.PNG?generation=1592396532899667&amp;alt=media)\n\n\n\nAnyway I want more information about this meta data and the way to submit...",
    "890499": "I got a replay from kaggle team\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/158987#890473\n\nWhat @dhruvrnaik  wrote is the answer, and as for 5) we will have `/kaggle/input/birdsong-recognition/test_audio` folder at kernel re-running phase. Inside that we have around 150 audio clips whose names correspond to the ids in `audio_id` in `test.csv`.",
    "890675": "Let me try to answer some of the questions in a more general way. The difficulties that we are facing when developing a classifier for bird sounds is the shift in acoustic domains between clean, focal recordings (e.g., Xeno-canto files) and soundscapes which are typically recorded with omnidirectional recording equipment. But there is more to it. Overlapping vocalizations are a major issue and Xeno-canto recordings may or may not contain background species and they may or may not have an appropriate label (typically primary and secondary labels in the metadata). The soundscapes however will have a high amount of simultaneous vocalizations - sometimes 5 or more for one 5-second interval. The test set includes 150 of these soundscapes recorded at 3 sites; all are of 10 min. duration. But we only have call-level annotations for 2 sites (they are hard to come by). Some soundscapes do not contain a single vocalization, others contain multiple hundreds.\n\nFrom my perspective, it's a two-step process in developing a classifier. First, train a classifier on Xeno-canto recordings an evaluate on a test split for maximum performance. Secondly, fine-tune the trained classifier to achieve optimal performance on soundscape data. We did not distribute soundscape data for validation to prevent participants from training on soundscape data - which would be valid according to the rules but not very practical from a application point of view: It would be too difficult to annotate hours of soundscape data before we can deploy a system to a new recording site. Trained systems should generalize well enough to cope with different acoustic environments (hence the 3 recording sites).\n\nI am not too familiar with the test set characteristics when it comes to Kaggle-specific modifications, so our other host might be able to answer these questions in more detail. But I'd be happy to join a discussion about methodology :)",
    "890965": "stefankahl \nThank you for your reply. Please let me confirm one thing.    \n\ntrain.csv file contains `secondary_labels` and `background` columns. I suppose both are essentially same if I preprocess `background` columns (remove unnecessary labels for this task). Am I right? If I'm not wrong, although this labels may sometime inappropriate, we can try to use this information to build soundscapes-level(multi label) models.\n\nThank you :-)",
    "891467": "Secondary labels and background should be mostly identical but secondary labels follow the eBird taxonomy (Clement's list) and background annotations follow the Xeno-canto taxonomy (IOC World list). Be aware that not all species mentioned in the secondary label column might be part of the training data. For training, you can focus on primary and secondary labels. For a submission, you should use the ebird code. You should definitely consider using secondary labels to validate your multi-class classifier, but (again) be aware that these labels are only weak labels with no timestamp.",
    "901615": "Dhruv's answers for 1-3 are correct. To fill in a couple more...\n4) This is the right ballpark, yes; ~12k test rows.\n5) The test audio is in a directory called test_audio; there are a couple public notebooks now which have good examples of iterating over the audio and creating a successful submission.\n6) These are metadata specific to the particular project that the example files came from. That project is completely separate from the test data, and the test data filenames have been changed to ensure no metadata leakage. So, Dhruv is correct here, as well.",
    "901865": "stefankahl  Thanks for the reply.\nI have a doubt about the training data. For each file we can hear only one bird but we have also see only one \"time\" feature (which means we can hear the bird at that moment if I am not wrong) but shouldn't we see multiple time for one bird ? I have try to listen few audios and I have the feeling we can hear let's say 3 times the bird during the 10/15 seconds but in the csv file, we will see only one time. Is it normal ?",
    "901879": "The time field in the training metadata is a time-of-day supplied by the recordist. (often rounded to the nearest half hour...) The training labels are 'weak' in the sense that the particular bird should exist somewhere in the recording, but it may be sparse, and may be mixed with other vocalizations.",
    "971270": "> The hidden test set audio consists of approximately 150 recordings (10 min long) from 3 sites and first 2 sites have labels at 5 sec interval so approximately total 2 * (50 * 10 * 60 / 5) + (50) = 12050 hidden test rows.\n\nIs it known that each site has roughly the same number of recordings?",
    "971272": "Thank you so much @tomdenton for your responses, clear communication really makes the competition much more enjoyable",
    "1001052": "Great work sir. Upload more such."
  },
  "source": "meta"
}