{
  "id": 406616,
  "title": "No-call dataset: 1000 x 8-second ogg sound files",
  "url": "/competitions/birdclef-2023/discussion/406616",
  "author_name": "",
  "post_date": "2023-05-03T03:23:36.777974400Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I've made a no-call dataset, for part of my training strategy.   Perhaps others might find it helpful.</p>\n<ul>\n<li><p>Dataset found <a href=\"https://www.kaggle.com/datasets/ollypowell/birdclef-8-sec-ogg\" target=\"_blank\">here</a>  </p></li>\n<li><p>I have been basing my training on 8 second chunks, and have made the dataset accordingly, but if you would like something else then you could fork <a href=\"https://www.kaggle.com/code/ollypowell/birdclef-uniform-no-call-sound-chunks\" target=\"_blank\">this notebook</a> I wrote to produce the dataset.</p></li>\n<li><p>I've also included a csv-file but it has slightly different format to the competition one, I've been using the same format across all my other datasets, like <a href=\"https://www.kaggle.com/datasets/ollypowell/cleaned-training-labels-21-23-for-birdclef2023\" target=\"_blank\">this one</a>.  That detail could also be modified easily enough.</p></li>\n<li><p>The dataset was derived from one some others made for BirdCLEF21 <a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">here</a></p></li>\n<li><p>If you notice any bugs, please feel free to comment.  I've started running my training notebook with it and can confirm no broken file-paths or similar.  It might be few days before I can confirm if it actually helps my performance metrics.</p></li>\n</ul>",
  "messages": [
    {
      "id": "2243602",
      "postDate": "05/03/2023 03:23:36",
      "content": "<p>I've made a no-call dataset, for part of my training strategy.   Perhaps others might find it helpful.</p>\n<ul>\n<li><p>Dataset found <a href=\"https://www.kaggle.com/datasets/ollypowell/birdclef-8-sec-ogg\" target=\"_blank\">here</a>  </p></li>\n<li><p>I have been basing my training on 8 second chunks, and have made the dataset accordingly, but if you would like something else then you could fork <a href=\"https://www.kaggle.com/code/ollypowell/birdclef-uniform-no-call-sound-chunks\" target=\"_blank\">this notebook</a> I wrote to produce the dataset.</p></li>\n<li><p>I've also included a csv-file but it has slightly different format to the competition one, I've been using the same format across all my other datasets, like <a href=\"https://www.kaggle.com/datasets/ollypowell/cleaned-training-labels-21-23-for-birdclef2023\" target=\"_blank\">this one</a>.  That detail could also be modified easily enough.</p></li>\n<li><p>The dataset was derived from one some others made for BirdCLEF21 <a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">here</a></p></li>\n<li><p>If you notice any bugs, please feel free to comment.  I've started running my training notebook with it and can confirm no broken file-paths or similar.  It might be few days before I can confirm if it actually helps my performance metrics.</p></li>\n</ul>",
      "rawMarkdown": "I've made a no-call dataset, for part of my training strategy.   Perhaps others might find it helpful.\n\n- Dataset found [here](https://www.kaggle.com/datasets/ollypowell/birdclef-8-sec-ogg)  \n\n- I have been basing my training on 8 second chunks, and have made the dataset accordingly, but if you would like something else then you could fork [this notebook](https://www.kaggle.com/code/ollypowell/birdclef-uniform-no-call-sound-chunks) I wrote to produce the dataset.\n\n- I've also included a csv-file but it has slightly different format to the competition one, I've been using the same format across all my other datasets, like [this one](https://www.kaggle.com/datasets/ollypowell/cleaned-training-labels-21-23-for-birdclef2023).  That detail could also be modified easily enough.\n\n- The dataset was derived from one some others made for BirdCLEF21 [here](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise)\n\n-  If you notice any bugs, please feel free to comment.  I've started running my training notebook with it and can confirm no broken file-paths or similar.  It might be few days before I can confirm if it actually helps my performance metrics.",
      "votes": null
    },
    {
      "id": "2244207",
      "postDate": "05/03/2023 14:04:59",
      "content": "<p>where do you base your no-call information? From the previous competitions?</p>",
      "rawMarkdown": "where do you base your no-call information? From the previous competitions?",
      "votes": null
    },
    {
      "id": "2244827",
      "postDate": "05/03/2023 22:59:45",
      "content": "<p>Yes that's correct.   I haven't sourced any new data, they are a collection of three other sources that were used in previous competitions.  You can follow the link to the previous dataset <a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">here</a> My motivation was to ensure fast loading by re-aranging them into uniform lengths, and randomised samples from each of the sources.  If you wanted to include other sources it would be easy to modify the notebook and re-run.</p>",
      "rawMarkdown": "Yes that's correct.   I haven't sourced any new data, they are a collection of three other sources that were used in previous competitions.  You can follow the link to the previous dataset [here](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise) My motivation was to ensure fast loading by re-aranging them into uniform lengths, and randomised samples from each of the sources.  If you wanted to include other sources it would be easy to modify the notebook and re-run.",
      "votes": null
    },
    {
      "id": "2246490",
      "postDate": "05/05/2023 07:36:12",
      "content": "<p>Hi, it looks great. But I think it would be nice if you could add some more information to the csv file, to show where these audio files originated from. For example, ff1010bird dataset, filename xx23yz.ogg, starting position t=7.0 seconds. The reason is that the competition rules (though I have not checked precisely) state that any additional training data should be accessible for everyone and the results should be reproducible. So, if the resource for the data is precisely documented, competition rules will not be violated, I suppose. </p>",
      "rawMarkdown": "Hi, it looks great. But I think it would be nice if you could add some more information to the csv file, to show where these audio files originated from. For example, ff1010bird dataset, filename xx23yz.ogg, starting position t=7.0 seconds. The reason is that the competition rules (though I have not checked precisely) state that any additional training data should be accessible for everyone and the results should be reproducible. So, if the resource for the data is precisely documented, competition rules will not be violated, I suppose.",
      "votes": null
    },
    {
      "id": "2247336",
      "postDate": "05/05/2023 22:21:39",
      "content": "<p>Thanks Hakan,  it's a good point, I should dig up their original sources.  All three were used in the 2021 competition, one was provided for that comp, the other two I think came from AI crowd sources, for lifeCLEF comps  Anyway they are freely available here on Kaggle for everyone to use so I don't think any rules have been broken.   I'm trying to dig up exactly which datasets the two AICrowd ones came from, and will link in the metadata so people can look at those too.  To get onto their platform requires a login and profile, but nothing else.  </p>",
      "rawMarkdown": "Thanks Hakan,  it's a good point, I should dig up their original sources.  All three were used in the 2021 competition, one was provided for that comp, the other two I think came from AI crowd sources, for lifeCLEF comps  Anyway they are freely available here on Kaggle for everyone to use so I don't think any rules have been broken.   I'm trying to dig up exactly which datasets the two AICrowd ones came from, and will link in the metadata so people can look at those too.  To get onto their platform requires a login and profile, but nothing else.",
      "votes": null
    },
    {
      "id": "2250155",
      "postDate": "05/08/2023 11:38:29",
      "content": "<p>I see, thanks for the dataset! </p>",
      "rawMarkdown": "I see, thanks for the dataset!",
      "votes": null
    },
    {
      "id": "2750109",
      "postDate": "04/13/2024 13:10:06",
      "content": "<p><a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a> how did you decide what was a nocall ? I've listen to 300 audios from the dataset and 50% of them have birds in them. </p>",
      "rawMarkdown": "ollypowell how did you decide what was a nocall ? I've listen to 300 audios from the dataset and 50% of them have birds in them.",
      "votes": null
    },
    {
      "id": "2761724",
      "postDate": "04/20/2024 03:40:17",
      "content": "<p>You have a lot of patience!  I don't think I listened to so many clips, but I assumed the chirps that sounded a bit like birds were actually crickets or cicadas, but I could be wrong.  I mainly believed that since the original sources claim they are no-call, then this should be what they are.  </p>\n<p>That said, I am having quite mixed success actually using these myself this year, so if there are false positives in there it might explain my problems.  My best score at the moment is 0.66 not using them at all.  I get 0.64 mixing them just to fill up empty parts for short clips, but it's not a direct comparison, it uses a smaller/faster model.  So far adding background noise on top of the wave form has produced worse results.</p>",
      "rawMarkdown": "You have a lot of patience!  I don't think I listened to so many clips, but I assumed the chirps that sounded a bit like birds were actually crickets or cicadas, but I could be wrong.  I mainly believed that since the original sources claim they are no-call, then this should be what they are.  \n\nThat said, I am having quite mixed success actually using these myself this year, so if there are false positives in there it might explain my problems.  My best score at the moment is 0.66 not using them at all.  I get 0.64 mixing them just to fill up empty parts for short clips, but it's not a direct comparison, it uses a smaller/faster model.  So far adding background noise on top of the wave form has produced worse results.",
      "votes": null
    },
    {
      "id": "2766307",
      "postDate": "04/21/2024 16:06:46",
      "content": "<p><a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a> I ended up doing 400 of them and labelling them in a csv:<br>\n<a href=\"https://www.kaggle.com/datasets/janmpia/nocall-manual-classification/data\" target=\"_blank\">https://www.kaggle.com/datasets/janmpia/nocall-manual-classification/data</a></p>\n<p>Feel free to use it, it's all 100% manual and I listenned to them fully so there shouldn't be any mistake in them. the filename col is the file name in the data you shared and the label is what you expect, bird or no_bird. <br>\nThe decision making was the following:</p>\n<ul>\n<li>In case of doubt, label as bird, this way, I had a solid no_bird set for which I was 100% confident didn't contain bird</li>\n</ul>\n<p>in total, about 50% of the dataset contained birds.<br>\nFeel free to listen to some of the birds samples, some of them are blatant but some are more suttle. <br>\nLet me know if this helps your approach now that you will have 200 proper no call instead of 1000 untrustworthy samples.</p>",
      "rawMarkdown": "ollypowell I ended up doing 400 of them and labelling them in a csv:\nhttps://www.kaggle.com/datasets/janmpia/nocall-manual-classification/data\n\nFeel free to use it, it's all 100% manual and I listenned to them fully so there shouldn't be any mistake in them. the filename col is the file name in the data you shared and the label is what you expect, bird or no_bird. \nThe decision making was the following:\n- In case of doubt, label as bird, this way, I had a solid no_bird set for which I was 100% confident didn't contain bird\n\nin total, about 50% of the dataset contained birds.\nFeel free to listen to some of the birds samples, some of them are blatant but some are more suttle. \nLet me know if this helps your approach now that you will have 200 proper no call instead of 1000 untrustworthy samples.",
      "votes": null
    },
    {
      "id": "2767080",
      "postDate": "04/22/2024 05:15:29",
      "content": "<p>Thanks, I'll look into it when I get a chance!</p>",
      "rawMarkdown": "Thanks, I'll look into it when I get a chance!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2244207,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "05/03/2023 14:04:59",
      "content": "<p>where do you base your no-call information? From the previous competitions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2244827,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "05/03/2023 22:59:45",
          "content": "<p>Yes that's correct.   I haven't sourced any new data, they are a collection of three other sources that were used in previous competitions.  You can follow the link to the previous dataset <a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">here</a> My motivation was to ensure fast loading by re-aranging them into uniform lengths, and randomised samples from each of the sources.  If you wanted to include other sources it would be easy to modify the notebook and re-run.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2250155,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "05/08/2023 11:38:29",
              "content": "<p>I see, thanks for the dataset! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2246490,
      "author_name": "hakandogan",
      "author_url": "",
      "post_date": "05/05/2023 07:36:12",
      "content": "<p>Hi, it looks great. But I think it would be nice if you could add some more information to the csv file, to show where these audio files originated from. For example, ff1010bird dataset, filename xx23yz.ogg, starting position t=7.0 seconds. The reason is that the competition rules (though I have not checked precisely) state that any additional training data should be accessible for everyone and the results should be reproducible. So, if the resource for the data is precisely documented, competition rules will not be violated, I suppose. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2247336,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "05/05/2023 22:21:39",
          "content": "<p>Thanks Hakan,  it's a good point, I should dig up their original sources.  All three were used in the 2021 competition, one was provided for that comp, the other two I think came from AI crowd sources, for lifeCLEF comps  Anyway they are freely available here on Kaggle for everyone to use so I don't think any rules have been broken.   I'm trying to dig up exactly which datasets the two AICrowd ones came from, and will link in the metadata so people can look at those too.  To get onto their platform requires a login and profile, but nothing else.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2750109,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "04/13/2024 13:10:06",
      "content": "<p><a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a> how did you decide what was a nocall ? I've listen to 300 audios from the dataset and 50% of them have birds in them. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2761724,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "04/20/2024 03:40:17",
          "content": "<p>You have a lot of patience!  I don't think I listened to so many clips, but I assumed the chirps that sounded a bit like birds were actually crickets or cicadas, but I could be wrong.  I mainly believed that since the original sources claim they are no-call, then this should be what they are.  </p>\n<p>That said, I am having quite mixed success actually using these myself this year, so if there are false positives in there it might explain my problems.  My best score at the moment is 0.66 not using them at all.  I get 0.64 mixing them just to fill up empty parts for short clips, but it's not a direct comparison, it uses a smaller/faster model.  So far adding background noise on top of the wave form has produced worse results.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2766307,
              "author_name": "janmpia",
              "author_url": "",
              "post_date": "04/21/2024 16:06:46",
              "content": "<p><a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a> I ended up doing 400 of them and labelling them in a csv:<br>\n<a href=\"https://www.kaggle.com/datasets/janmpia/nocall-manual-classification/data\" target=\"_blank\">https://www.kaggle.com/datasets/janmpia/nocall-manual-classification/data</a></p>\n<p>Feel free to use it, it's all 100% manual and I listenned to them fully so there shouldn't be any mistake in them. the filename col is the file name in the data you shared and the label is what you expect, bird or no_bird. <br>\nThe decision making was the following:</p>\n<ul>\n<li>In case of doubt, label as bird, this way, I had a solid no_bird set for which I was 100% confident didn't contain bird</li>\n</ul>\n<p>in total, about 50% of the dataset contained birds.<br>\nFeel free to listen to some of the birds samples, some of them are blatant but some are more suttle. <br>\nLet me know if this helps your approach now that you will have 200 proper no call instead of 1000 untrustworthy samples.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2767080,
                  "author_name": "ollypowell",
                  "author_url": "",
                  "post_date": "04/22/2024 05:15:29",
                  "content": "<p>Thanks, I'll look into it when I get a chance!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2243602": "I've made a no-call dataset, for part of my training strategy.   Perhaps others might find it helpful.\n\n- Dataset found [here](https://www.kaggle.com/datasets/ollypowell/birdclef-8-sec-ogg)  \n\n- I have been basing my training on 8 second chunks, and have made the dataset accordingly, but if you would like something else then you could fork [this notebook](https://www.kaggle.com/code/ollypowell/birdclef-uniform-no-call-sound-chunks) I wrote to produce the dataset.\n\n- I've also included a csv-file but it has slightly different format to the competition one, I've been using the same format across all my other datasets, like [this one](https://www.kaggle.com/datasets/ollypowell/cleaned-training-labels-21-23-for-birdclef2023).  That detail could also be modified easily enough.\n\n- The dataset was derived from one some others made for BirdCLEF21 [here](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise)\n\n-  If you notice any bugs, please feel free to comment.  I've started running my training notebook with it and can confirm no broken file-paths or similar.  It might be few days before I can confirm if it actually helps my performance metrics.",
    "2244207": "where do you base your no-call information? From the previous competitions?",
    "2244827": "Yes that's correct.   I haven't sourced any new data, they are a collection of three other sources that were used in previous competitions.  You can follow the link to the previous dataset [here](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise) My motivation was to ensure fast loading by re-aranging them into uniform lengths, and randomised samples from each of the sources.  If you wanted to include other sources it would be easy to modify the notebook and re-run.",
    "2246490": "Hi, it looks great. But I think it would be nice if you could add some more information to the csv file, to show where these audio files originated from. For example, ff1010bird dataset, filename xx23yz.ogg, starting position t=7.0 seconds. The reason is that the competition rules (though I have not checked precisely) state that any additional training data should be accessible for everyone and the results should be reproducible. So, if the resource for the data is precisely documented, competition rules will not be violated, I suppose.",
    "2247336": "Thanks Hakan,  it's a good point, I should dig up their original sources.  All three were used in the 2021 competition, one was provided for that comp, the other two I think came from AI crowd sources, for lifeCLEF comps  Anyway they are freely available here on Kaggle for everyone to use so I don't think any rules have been broken.   I'm trying to dig up exactly which datasets the two AICrowd ones came from, and will link in the metadata so people can look at those too.  To get onto their platform requires a login and profile, but nothing else.",
    "2250155": "I see, thanks for the dataset!",
    "2750109": "ollypowell how did you decide what was a nocall ? I've listen to 300 audios from the dataset and 50% of them have birds in them.",
    "2761724": "You have a lot of patience!  I don't think I listened to so many clips, but I assumed the chirps that sounded a bit like birds were actually crickets or cicadas, but I could be wrong.  I mainly believed that since the original sources claim they are no-call, then this should be what they are.  \n\nThat said, I am having quite mixed success actually using these myself this year, so if there are false positives in there it might explain my problems.  My best score at the moment is 0.66 not using them at all.  I get 0.64 mixing them just to fill up empty parts for short clips, but it's not a direct comparison, it uses a smaller/faster model.  So far adding background noise on top of the wave form has produced worse results.",
    "2766307": "ollypowell I ended up doing 400 of them and labelling them in a csv:\nhttps://www.kaggle.com/datasets/janmpia/nocall-manual-classification/data\n\nFeel free to use it, it's all 100% manual and I listenned to them fully so there shouldn't be any mistake in them. the filename col is the file name in the data you shared and the label is what you expect, bird or no_bird. \nThe decision making was the following:\n- In case of doubt, label as bird, this way, I had a solid no_bird set for which I was 100% confident didn't contain bird\n\nin total, about 50% of the dataset contained birds.\nFeel free to listen to some of the birds samples, some of them are blatant but some are more suttle. \nLet me know if this helps your approach now that you will have 200 proper no call instead of 1000 untrustworthy samples.",
    "2767080": "Thanks, I'll look into it when I get a chance!"
  },
  "source": "meta"
}