{
  "id": 178158,
  "title": "Missed \"birds\" from the example test audio",
  "url": "/competitions/birdsong-recognition/discussion/178158",
  "author_name": "",
  "post_date": "2020-08-28T19:54:44.036077300Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F0828f590b68790f3cdaeda61b07d8da9%2Fmy-awesome-meme.jpeg?generation=1598644353014676&amp;alt=media\" alt=\"\"></p>\n<p>I have noticed some of <em>ebird_code</em> presented inside <strong>example_test_audio_summary.csv</strong> file but missed inside <strong>train_audio</strong> folder. There are 5 such cases:</p>\n<ul>\n<li>squirrel</li>\n<li>mouqua</li>\n<li>whhwoo</li>\n<li>hawo</li>\n<li>unk</li>\n</ul>\n<p>66(!) rows from 153 consist one of these unknown labels and I've saved such cases <a href=\"https://www.kaggle.com/koza4ukdmitrij/lost-birds-from-cornell-birdcall-identification\" target=\"_blank\">here</a>. Do we know for sure that it doesn't happen for hidden test set? </p>",
  "messages": [
    {
      "id": "989457",
      "postDate": "08/28/2020 19:54:44",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F0828f590b68790f3cdaeda61b07d8da9%2Fmy-awesome-meme.jpeg?generation=1598644353014676&amp;alt=media\" alt=\"\"></p>\n<p>I have noticed some of <em>ebird_code</em> presented inside <strong>example_test_audio_summary.csv</strong> file but missed inside <strong>train_audio</strong> folder. There are 5 such cases:</p>\n<ul>\n<li>squirrel</li>\n<li>mouqua</li>\n<li>whhwoo</li>\n<li>hawo</li>\n<li>unk</li>\n</ul>\n<p>66(!) rows from 153 consist one of these unknown labels and I've saved such cases <a href=\"https://www.kaggle.com/koza4ukdmitrij/lost-birds-from-cornell-birdcall-identification\" target=\"_blank\">here</a>. Do we know for sure that it doesn't happen for hidden test set? </p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F0828f590b68790f3cdaeda61b07d8da9%2Fmy-awesome-meme.jpeg?generation=1598644353014676&alt=media)\n\nI have noticed some of *ebird_code* presented inside **example_test_audio_summary.csv** file but missed inside **train_audio** folder. There are 5 such cases:\n- squirrel\n- mouqua\n- whhwoo\n- hawo\n- unk\n\n66(!) rows from 153 consist one of these unknown labels and I've saved such cases [here](https://www.kaggle.com/koza4ukdmitrij/lost-birds-from-cornell-birdcall-identification). Do we know for sure that it doesn't happen for hidden test set?",
      "votes": null
    },
    {
      "id": "989498",
      "postDate": "08/28/2020 20:59:02",
      "content": "<p>Yes, the example annotations are from another project, and have some extra labels that don't appear in the test set. We decided to leave them in as a reminder that the world contains non-bird entities. :)</p>",
      "rawMarkdown": "Yes, the example annotations are from another project, and have some extra labels that don't appear in the test set. We decided to leave them in as a reminder that the world contains non-bird entities. :)",
      "votes": null
    },
    {
      "id": "990178",
      "postDate": "08/29/2020 12:07:19",
      "content": "<p>Thank you for response! But can we use this audio for better understanding bird's voices (that presence inside training set), enviroment sounds and so on, or these sounds were taken from absolutely different distribution? </p>",
      "rawMarkdown": "Thank you for response! But can we use this audio for better understanding bird's voices (that presence inside training set), enviroment sounds and so on, or these sounds were taken from absolutely different distribution?",
      "votes": null
    },
    {
      "id": "990187",
      "postDate": "08/29/2020 12:19:22",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> </p>\n<p>I found exactly same and about to post a discussion. <br>\n<a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> this means that we cannot (and should not) use example_test_audio to <br>\nevaluate our models. I appreciate if you could confirm if I understand correct. </p>\n<p>one more question: 'site1' and 'site2' of example_test_audio are not identical <br>\nto 'site1' and 'site2' of actual test_audio recording? </p>\n<p>Thank you, <br>\nmeg.</p>",
      "rawMarkdown": "Hello @koza4ukdmitrij \n\nI found exactly same and about to post a discussion. \n@tomdenton this means that we cannot (and should not) use example_test_audio to \nevaluate our models. I appreciate if you could confirm if I understand correct. \n\none more question: 'site1' and 'site2' of example_test_audio are not identical \nto 'site1' and 'site2' of actual test_audio recording? \n\nThank you, \nmeg.",
      "votes": null
    },
    {
      "id": "990194",
      "postDate": "08/29/2020 12:26:44",
      "content": "<p>Thank you for the information! But I think it would be beneficial to check your model at least on existing inside train set birds. Anyway lets waiting for host response. </p>",
      "rawMarkdown": "Thank you for the information! But I think it would be beneficial to check your model at least on existing inside train set birds. Anyway lets waiting for host response.",
      "votes": null
    },
    {
      "id": "990546",
      "postDate": "08/29/2020 17:25:55",
      "content": "<p>\"[C]an we use this audio for better understanding bird's voices[?]\"</p>\n<p>In the sense that these clips are typical of an additional random forest (pun intended), yes! The heart of the problem we're working on is how to transfer learning from a large 'pretty good' training set to usage in the actual field; dealing gracefully with the domain shift is at the heart of what we're trying to do. </p>\n<p>\"[Are] these sounds were taken from absolutely different distribution?\"</p>\n<p>It's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species. What we're seeing in this problem space (here and in the BirdClef challenges) is 'lab-trained' models which are having a very difficult time generalizing to the real world, and we would love to know why. And if not 'why' at least how to do better…</p>",
      "rawMarkdown": "\"[C]an we use this audio for better understanding bird's voices[?]\"\n\nIn the sense that these clips are typical of an additional random forest (pun intended), yes! The heart of the problem we're working on is how to transfer learning from a large 'pretty good' training set to usage in the actual field; dealing gracefully with the domain shift is at the heart of what we're trying to do. \n\n\"[Are] these sounds were taken from absolutely different distribution?\"\n\nIt's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species. What we're seeing in this problem space (here and in the BirdClef challenges) is 'lab-trained' models which are having a very difficult time generalizing to the real world, and we would love to know why. And if not 'why' at least how to do better...",
      "votes": null
    },
    {
      "id": "990550",
      "postDate": "08/29/2020 17:28:20",
      "content": "<p>As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.</p>",
      "rawMarkdown": "As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.",
      "votes": null
    },
    {
      "id": "990560",
      "postDate": "08/29/2020 17:37:00",
      "content": "<p>Thank you, really helpful!</p>",
      "rawMarkdown": "Thank you, really helpful!",
      "votes": null
    },
    {
      "id": "990852",
      "postDate": "08/29/2020 22:41:23",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>  for your clarification. </p>\n<p>…  just to double check, if there is any information that I miss  (there seems lots of them).</p>\n<blockquote>\n  <p>As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.</p>\n</blockquote>\n<p>Do you mean this passage below in Data description that tells the example test audio are from different sites<br>\nfrom actual test audio? Or is there any description that I miss? </p>\n<p>\"Two example soundscapes from another data source are also provided to illustrate how the soundscapes are labeled and the hidden dataset folder structure. The two example audio files … \" </p>\n<p>… just one more confirmation needed. </p>\n<blockquote>\n  <p>It's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species.</p>\n</blockquote>\n<p>Do I understand correct that the example test audio set can be used as bird call samples. I mean, they are labeled<br>\ncorrectly (correct enough, or at least with similar precisions with the training data that are given to us). <br>\nIn other words, the example test audio set are not just a place holder to show the layout of the test data set, but <br>\nexample test data do have meaningful information of birdcalls and their correct labels. Do I understand correct? </p>\n<p>thank you for your patience,<br>\nmeg. </p>",
      "rawMarkdown": "Thank you so much @tomdenton  for your clarification. \n\n...  just to double check, if there is any information that I miss  (there seems lots of them).\n\n> As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.\n\nDo you mean this passage below in Data description that tells the example test audio are from different sites\nfrom actual test audio? Or is there any description that I miss? \n\n\"Two example soundscapes from another data source are also provided to illustrate how the soundscapes are labeled and the hidden dataset folder structure. The two example audio files ... \" \n\n... just one more confirmation needed. \n\n> It's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species.\n\nDo I understand correct that the example test audio set can be used as bird call samples. I mean, they are labeled\ncorrectly (correct enough, or at least with similar precisions with the training data that are given to us). \nIn other words, the example test audio set are not just a place holder to show the layout of the test data set, but \nexample test data do have meaningful information of birdcalls and their correct labels. Do I understand correct? \n\nthank you for your patience,\nmeg.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 989498,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "08/28/2020 20:59:02",
      "content": "<p>Yes, the example annotations are from another project, and have some extra labels that don't appear in the test set. We decided to leave them in as a reminder that the world contains non-bird entities. :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 990178,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/29/2020 12:07:19",
          "content": "<p>Thank you for response! But can we use this audio for better understanding bird's voices (that presence inside training set), enviroment sounds and so on, or these sounds were taken from absolutely different distribution? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990546,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "08/29/2020 17:25:55",
          "content": "<p>\"[C]an we use this audio for better understanding bird's voices[?]\"</p>\n<p>In the sense that these clips are typical of an additional random forest (pun intended), yes! The heart of the problem we're working on is how to transfer learning from a large 'pretty good' training set to usage in the actual field; dealing gracefully with the domain shift is at the heart of what we're trying to do. </p>\n<p>\"[Are] these sounds were taken from absolutely different distribution?\"</p>\n<p>It's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species. What we're seeing in this problem space (here and in the BirdClef challenges) is 'lab-trained' models which are having a very difficult time generalizing to the real world, and we would love to know why. And if not 'why' at least how to do better…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990560,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/29/2020 17:37:00",
          "content": "<p>Thank you, really helpful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 990187,
      "author_name": "megner",
      "author_url": "",
      "post_date": "08/29/2020 12:19:22",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> </p>\n<p>I found exactly same and about to post a discussion. <br>\n<a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> this means that we cannot (and should not) use example_test_audio to <br>\nevaluate our models. I appreciate if you could confirm if I understand correct. </p>\n<p>one more question: 'site1' and 'site2' of example_test_audio are not identical <br>\nto 'site1' and 'site2' of actual test_audio recording? </p>\n<p>Thank you, <br>\nmeg.</p>",
      "votes": null,
      "replies": [
        {
          "id": 990194,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/29/2020 12:26:44",
          "content": "<p>Thank you for the information! But I think it would be beneficial to check your model at least on existing inside train set birds. Anyway lets waiting for host response. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990550,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "08/29/2020 17:28:20",
          "content": "<p>As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 990852,
      "author_name": "megner",
      "author_url": "",
      "post_date": "08/29/2020 22:41:23",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>  for your clarification. </p>\n<p>…  just to double check, if there is any information that I miss  (there seems lots of them).</p>\n<blockquote>\n  <p>As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.</p>\n</blockquote>\n<p>Do you mean this passage below in Data description that tells the example test audio are from different sites<br>\nfrom actual test audio? Or is there any description that I miss? </p>\n<p>\"Two example soundscapes from another data source are also provided to illustrate how the soundscapes are labeled and the hidden dataset folder structure. The two example audio files … \" </p>\n<p>… just one more confirmation needed. </p>\n<blockquote>\n  <p>It's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species.</p>\n</blockquote>\n<p>Do I understand correct that the example test audio set can be used as bird call samples. I mean, they are labeled<br>\ncorrectly (correct enough, or at least with similar precisions with the training data that are given to us). <br>\nIn other words, the example test audio set are not just a place holder to show the layout of the test data set, but <br>\nexample test data do have meaningful information of birdcalls and their correct labels. Do I understand correct? </p>\n<p>thank you for your patience,<br>\nmeg. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "989457": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F0828f590b68790f3cdaeda61b07d8da9%2Fmy-awesome-meme.jpeg?generation=1598644353014676&alt=media)\n\nI have noticed some of *ebird_code* presented inside **example_test_audio_summary.csv** file but missed inside **train_audio** folder. There are 5 such cases:\n- squirrel\n- mouqua\n- whhwoo\n- hawo\n- unk\n\n66(!) rows from 153 consist one of these unknown labels and I've saved such cases [here](https://www.kaggle.com/koza4ukdmitrij/lost-birds-from-cornell-birdcall-identification). Do we know for sure that it doesn't happen for hidden test set?",
    "989498": "Yes, the example annotations are from another project, and have some extra labels that don't appear in the test set. We decided to leave them in as a reminder that the world contains non-bird entities. :)",
    "990178": "Thank you for response! But can we use this audio for better understanding bird's voices (that presence inside training set), enviroment sounds and so on, or these sounds were taken from absolutely different distribution?",
    "990187": "Hello @koza4ukdmitrij \n\nI found exactly same and about to post a discussion. \n@tomdenton this means that we cannot (and should not) use example_test_audio to \nevaluate our models. I appreciate if you could confirm if I understand correct. \n\none more question: 'site1' and 'site2' of example_test_audio are not identical \nto 'site1' and 'site2' of actual test_audio recording? \n\nThank you, \nmeg.",
    "990194": "Thank you for the information! But I think it would be beneficial to check your model at least on existing inside train set birds. Anyway lets waiting for host response.",
    "990546": "\"[C]an we use this audio for better understanding bird's voices[?]\"\n\nIn the sense that these clips are typical of an additional random forest (pun intended), yes! The heart of the problem we're working on is how to transfer learning from a large 'pretty good' training set to usage in the actual field; dealing gracefully with the domain shift is at the heart of what we're trying to do. \n\n\"[Are] these sounds were taken from absolutely different distribution?\"\n\nIt's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species. What we're seeing in this problem space (here and in the BirdClef challenges) is 'lab-trained' models which are having a very difficult time generalizing to the real world, and we would love to know why. And if not 'why' at least how to do better...",
    "990550": "As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.",
    "990560": "Thank you, really helpful!",
    "990852": "Thank you so much @tomdenton  for your clarification. \n\n...  just to double check, if there is any information that I miss  (there seems lots of them).\n\n> As mentioned in the 'Data' description, the example audio is from a different site than the actual test audio.\n\nDo you mean this passage below in Data description that tells the example test audio are from different sites\nfrom actual test audio? Or is there any description that I miss? \n\n\"Two example soundscapes from another data source are also provided to illustrate how the soundscapes are labeled and the hidden dataset folder structure. The two example audio files ... \" \n\n... just one more confirmation needed. \n\n> It's not an 'absolutely different' distribution, but still North American soundscape audio containing bird vocalizations from a known set of species.\n\nDo I understand correct that the example test audio set can be used as bird call samples. I mean, they are labeled\ncorrectly (correct enough, or at least with similar precisions with the training data that are given to us). \nIn other words, the example test audio set are not just a place holder to show the layout of the test data set, but \nexample test data do have meaningful information of birdcalls and their correct labels. Do I understand correct? \n\nthank you for your patience,\nmeg."
  },
  "source": "meta"
}