{
  "id": 159341,
  "title": "Hi everyone!",
  "url": "/competitions/birdsong-recognition/discussion/159341",
  "author_name": "Stefan Kahl",
  "post_date": "2020-06-17T06:50:06.921000",
  "votes": 13,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks for participating in this competition! I am a postdoctoral fellow at the Cornell Lab of Ornithology and co-host of this challenge. I am also the host of the BirdCLEF2020 competition (which is already in its 7th iteration and recently started its second submission round). I have dedicated my work to the automated detection of bird sounds in continuous audio data and would be more than happy to assist and guide you with any questions you may have during this challenge.</p>\n\n<p>Good luck everyone, and rest assured that your contribution will advance our efforts to monitor endangered species and habitats.</p>\n\n<p>Stefan</p>",
  "messages": [
    {
      "id": 889805,
      "postDate": "2020-06-17T06:50:06.920Z",
      "content": "<p>Thanks for participating in this competition! I am a postdoctoral fellow at the Cornell Lab of Ornithology and co-host of this challenge. I am also the host of the BirdCLEF2020 competition (which is already in its 7th iteration and recently started its second submission round). I have dedicated my work to the automated detection of bird sounds in continuous audio data and would be more than happy to assist and guide you with any questions you may have during this challenge.</p>\n\n<p>Good luck everyone, and rest assured that your contribution will advance our efforts to monitor endangered species and habitats.</p>\n\n<p>Stefan</p>",
      "rawMarkdown": "Thanks for participating in this competition! I am a postdoctoral fellow at the Cornell Lab of Ornithology and co-host of this challenge. I am also the host of the BirdCLEF2020 competition (which is already in its 7th iteration and recently started its second submission round). I have dedicated my work to the automated detection of bird sounds in continuous audio data and would be more than happy to assist and guide you with any questions you may have during this challenge.\n\nGood luck everyone, and rest assured that your contribution will advance our efforts to monitor endangered species and habitats.\n\nStefan",
      "votes": 13
    },
    {
      "id": 890677,
      "postDate": "2020-06-17T16:27:01.037Z",
      "content": "<p>Sure, just posted a comment.</p>",
      "rawMarkdown": "Sure, just posted a comment.",
      "votes": 1
    },
    {
      "id": 890450,
      "postDate": "2020-06-17T14:14:34.407Z",
      "content": "<p><a href=\"/stefankahl\">@stefankahl</a> can you answer some of my questions posted here: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159123\">https://www.kaggle.com/c/birdsong-recognition/discussion/159123</a>.</p>\n\n<p>thanks for hosting such an interesting competition.</p>",
      "rawMarkdown": "@stefankahl can you answer some of my questions posted here: https://www.kaggle.com/c/birdsong-recognition/discussion/159123.\n\nthanks for hosting such an interesting competition.",
      "votes": 1,
      "replies": [
        {
          "id": 901156,
          "postDate": "2020-06-25T09:26:09.970Z",
          "content": "<p>I just looked at your questions that haven't been answered yet and I'm afraid I can't confidently answer any of them. I am not familiar with Kaggle-specific test set handling and <a href=\"/tomdenton\">@tomdenton</a> might be the one who can answer more appropriately. I know a lot more about the training data than I do know about the test data :)</p>",
          "rawMarkdown": "I just looked at your questions that haven't been answered yet and I'm afraid I can't confidently answer any of them. I am not familiar with Kaggle-specific test set handling and @tomdenton might be the one who can answer more appropriately. I know a lot more about the training data than I do know about the test data :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 1013444,
      "postDate": "2020-09-16T17:37:29.600Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, can you release the test set after the competition is finalized?  Thanks!</p>",
      "rawMarkdown": "Hey @stefankahl, can you release the test set after the competition is finalized?  Thanks!"
    },
    {
      "id": 994687,
      "postDate": "2020-09-01T20:18:16.153Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<p>I've just been reading up on birdclef2019, particularly your published summary and <a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\" target=\"_blank\">http://ceur-ws.org/Vol-2380/paper_86.pdf</a> (which everyone in this comp should probably read).</p>\n<p>There's one thing I'm not certain of: how did you handle 5 second intervals in the test audio where no birds were present at all? It seems that these were treated like every other interval, and perfect c-mAP and r-mAP scores would have been awarded on those intervals to participants who predicted probabilities of 0 for each species. Is that correct?</p>\n<p>I'm just asking so I can better understand Lasseck's process.</p>",
      "rawMarkdown": "Hi @stefankahl \n\nI've just been reading up on birdclef2019, particularly your published summary and http://ceur-ws.org/Vol-2380/paper_86.pdf (which everyone in this comp should probably read).\n\nThere's one thing I'm not certain of: how did you handle 5 second intervals in the test audio where no birds were present at all? It seems that these were treated like every other interval, and perfect c-mAP and r-mAP scores would have been awarded on those intervals to participants who predicted probabilities of 0 for each species. Is that correct?\n\nI'm just asking so I can better understand Lasseck's process.",
      "replies": [
        {
          "id": 995321,
          "postDate": "2020-09-02T10:57:45.840Z",
          "content": "<p>The c-mAP and r-mAP that we use for BirdCLEF are probably not the best metrics to assess how well a recognition system copes with non-bird events since ranking metrics benefit from recall. In general, returning as many predictions per 5-second interval as possible increases the score (maybe just marginally, but still) in our eval system. Additionally, these metrics were only computed for intervals that had an annotation (which is somewhat of an relic from when we didn't have a whole lot of annotations). I mainly used intervals with no vocalization to investigate with other metrics outside the official eval track - that's why I think the Kaggle metric better suits the real-world use case where you need precise results and a low false positive-rate. However, suppressing false positives helps to increase the scores in BirdCLEF, especially the c-mAP which looks at how \"clean\" your ranked results for each species are. More false positives will decrease these scores for species that \"produce\" many false detections.</p>",
          "rawMarkdown": "The c-mAP and r-mAP that we use for BirdCLEF are probably not the best metrics to assess how well a recognition system copes with non-bird events since ranking metrics benefit from recall. In general, returning as many predictions per 5-second interval as possible increases the score (maybe just marginally, but still) in our eval system. Additionally, these metrics were only computed for intervals that had an annotation (which is somewhat of an relic from when we didn't have a whole lot of annotations). I mainly used intervals with no vocalization to investigate with other metrics outside the official eval track - that's why I think the Kaggle metric better suits the real-world use case where you need precise results and a low false positive-rate. However, suppressing false positives helps to increase the scores in BirdCLEF, especially the c-mAP which looks at how \"clean\" your ranked results for each species are. More false positives will decrease these scores for species that \"produce\" many false detections.",
          "votes": 1
        }
      ]
    },
    {
      "id": 900418,
      "postDate": "2020-06-24T19:44:41.130Z",
      "content": "<p>What's the range of score that you guys were able to achieve on this dataset by yourselves?</p>",
      "rawMarkdown": "What's the range of score that you guys were able to achieve on this dataset by yourselves?"
    },
    {
      "id": 1012125,
      "postDate": "2020-09-16T00:03:09.570Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 901371,
      "postDate": "2020-06-25T12:24:27.267Z",
      "content": "<p>Hi <a href=\"/stefankahl\">@stefankahl</a>, </p>\n\n<p>I see a few problems in the data structure which causes me a lot of headache. The directory structure seen during submission is different from development session and we cannot explore the directory structure available during submission run.</p>\n\n<p>I have a few suggestions to improve on this:\n\"example_test_audio\" could be renamed to 'test_audio' in interactive sessions and contain only few files but from all three sites: say 2 files from site_1,  2 files from site_2 and 2 files from site_3. Also, we would have two files test.csv and sample_submission.csv with 6 rows corresponding to the 6 files above (again same file names during interactive and submission session). Obviously, the 6 files should be from public part or test data to make sure private test data remains totally hidden.</p>\n\n<p>This way we would have a mini datasets during code development that has <strong>exactly same structure</strong> as the one seen during a submission run, but just with much less data. We could then write code without 'if/else', 'file/dir exists' and co that we currently need to handle the different structure and filenames seen during submission.</p>",
      "rawMarkdown": "Hi @stefankahl, \n\nI see a few problems in the data structure which causes me a lot of headache. The directory structure seen during submission is different from development session and we cannot explore the directory structure available during submission run.\n\nI have a few suggestions to improve on this:\n\"example\\_test\\_audio\" could be renamed to 'test\\_audio' in interactive sessions and contain only few files but from all three sites: say 2 files from site\\_1,  2 files from site\\_2 and 2 files from site\\_3. Also, we would have two files test.csv and sample\\_submission.csv with 6 rows corresponding to the 6 files above (again same file names during interactive and submission session). Obviously, the 6 files should be from public part or test data to make sure private test data remains totally hidden.\n\nThis way we would have a mini datasets during code development that has **exactly same structure** as the one seen during a submission run, but just with much less data. We could then write code without 'if/else', 'file/dir exists' and co that we currently need to handle the different structure and filenames seen during submission.\n\n\n ",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 890677,
      "author_name": "Stefan Kahl",
      "author_url": "",
      "post_date": "2020-06-17T16:27:01.037000",
      "content": "<p>Sure, just posted a comment.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 890450,
      "author_name": "Dhananjay Raut",
      "author_url": "",
      "post_date": "2020-06-17T14:14:34.407000",
      "content": "<p><a href=\"/stefankahl\">@stefankahl</a> can you answer some of my questions posted here: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159123\">https://www.kaggle.com/c/birdsong-recognition/discussion/159123</a>.</p>\n\n<p>thanks for hosting such an interesting competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 901156,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2020-06-25T09:26:09.970000",
          "content": "<p>I just looked at your questions that haven't been answered yet and I'm afraid I can't confidently answer any of them. I am not familiar with Kaggle-specific test set handling and <a href=\"/tomdenton\">@tomdenton</a> might be the one who can answer more appropriately. I know a lot more about the training data than I do know about the test data :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1013444,
      "author_name": "Krisztián Fekete",
      "author_url": "",
      "post_date": "2020-09-16T17:37:29.600000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, can you release the test set after the competition is finalized?  Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 994687,
      "author_name": "Louka Ewington-Pitsos",
      "author_url": "",
      "post_date": "2020-09-01T20:18:16.153000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<p>I've just been reading up on birdclef2019, particularly your published summary and <a href=\"http://ceur-ws.org/Vol-2380/paper_86.pdf\" target=\"_blank\">http://ceur-ws.org/Vol-2380/paper_86.pdf</a> (which everyone in this comp should probably read).</p>\n<p>There's one thing I'm not certain of: how did you handle 5 second intervals in the test audio where no birds were present at all? It seems that these were treated like every other interval, and perfect c-mAP and r-mAP scores would have been awarded on those intervals to participants who predicted probabilities of 0 for each species. Is that correct?</p>\n<p>I'm just asking so I can better understand Lasseck's process.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 995321,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2020-09-02T10:57:45.840000",
          "content": "<p>The c-mAP and r-mAP that we use for BirdCLEF are probably not the best metrics to assess how well a recognition system copes with non-bird events since ranking metrics benefit from recall. In general, returning as many predictions per 5-second interval as possible increases the score (maybe just marginally, but still) in our eval system. Additionally, these metrics were only computed for intervals that had an annotation (which is somewhat of an relic from when we didn't have a whole lot of annotations). I mainly used intervals with no vocalization to investigate with other metrics outside the official eval track - that's why I think the Kaggle metric better suits the real-world use case where you need precise results and a low false positive-rate. However, suppressing false positives helps to increase the scores in BirdCLEF, especially the c-mAP which looks at how \"clean\" your ranked results for each species are. More false positives will decrease these scores for species that \"produce\" many false detections.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 900418,
      "author_name": "johnny",
      "author_url": "",
      "post_date": "2020-06-24T19:44:41.130000",
      "content": "<p>What's the range of score that you guys were able to achieve on this dataset by yourselves?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1012125,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-16T00:03:09.570000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 901371,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-25T12:24:27.267000",
      "content": "<p>Hi <a href=\"/stefankahl\">@stefankahl</a>, </p>\n\n<p>I see a few problems in the data structure which causes me a lot of headache. The directory structure seen during submission is different from development session and we cannot explore the directory structure available during submission run.</p>\n\n<p>I have a few suggestions to improve on this:\n\"example_test_audio\" could be renamed to 'test_audio' in interactive sessions and contain only few files but from all three sites: say 2 files from site_1,  2 files from site_2 and 2 files from site_3. Also, we would have two files test.csv and sample_submission.csv with 6 rows corresponding to the 6 files above (again same file names during interactive and submission session). Obviously, the 6 files should be from public part or test data to make sure private test data remains totally hidden.</p>\n\n<p>This way we would have a mini datasets during code development that has <strong>exactly same structure</strong> as the one seen during a submission run, but just with much less data. We could then write code without 'if/else', 'file/dir exists' and co that we currently need to handle the different structure and filenames seen during submission.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "889805": "Thanks for participating in this competition! I am a postdoctoral fellow at the Cornell Lab of Ornithology and co-host of this challenge. I am also the host of the BirdCLEF2020 competition (which is already in its 7th iteration and recently started its second submission round). I have dedicated my work to the automated detection of bird sounds in continuous audio data and would be more than happy to assist and guide you with any questions you may have during this challenge.\n\nGood luck everyone, and rest assured that your contribution will advance our efforts to monitor endangered species and habitats.\n\nStefan",
    "890677": "Sure, just posted a comment.",
    "890450": "@stefankahl can you answer some of my questions posted here: https://www.kaggle.com/c/birdsong-recognition/discussion/159123.\n\nthanks for hosting such an interesting competition.",
    "1013444": "Hey @stefankahl, can you release the test set after the competition is finalized?  Thanks!",
    "994687": "Hi @stefankahl \n\nI've just been reading up on birdclef2019, particularly your published summary and http://ceur-ws.org/Vol-2380/paper_86.pdf (which everyone in this comp should probably read).\n\nThere's one thing I'm not certain of: how did you handle 5 second intervals in the test audio where no birds were present at all? It seems that these were treated like every other interval, and perfect c-mAP and r-mAP scores would have been awarded on those intervals to participants who predicted probabilities of 0 for each species. Is that correct?\n\nI'm just asking so I can better understand Lasseck's process.",
    "900418": "What's the range of score that you guys were able to achieve on this dataset by yourselves?",
    "1012125": "",
    "901371": "Hi @stefankahl, \n\nI see a few problems in the data structure which causes me a lot of headache. The directory structure seen during submission is different from development session and we cannot explore the directory structure available during submission run.\n\nI have a few suggestions to improve on this:\n\"example\\_test\\_audio\" could be renamed to 'test\\_audio' in interactive sessions and contain only few files but from all three sites: say 2 files from site\\_1,  2 files from site\\_2 and 2 files from site\\_3. Also, we would have two files test.csv and sample\\_submission.csv with 6 rows corresponding to the 6 files above (again same file names during interactive and submission session). Obviously, the 6 files should be from public part or test data to make sure private test data remains totally hidden.\n\nThis way we would have a mini datasets during code development that has **exactly same structure** as the one seen during a submission run, but just with much less data. We could then write code without 'if/else', 'file/dir exists' and co that we currently need to handle the different structure and filenames seen during submission.\n\n\n "
  }
}