{
  "id": 568028,
  "title": "Using BirdNet and Bird Vocalization Classifier",
  "url": "/competitions/birdclef-2025/discussion/568028",
  "author_name": "Konstantin Dmitriev",
  "post_date": "2025-03-13T12:31:24.051000",
  "votes": 34,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In the current competition we have much more time available to make predictions. So, it seems, that we have to build more accurate models, instead of 'fast' models, that were used in the last competitions.</p>\n<table>\n<thead>\n<tr>\n<th>Year</th>\n<th>Number of test records</th>\n<th>Record length, min</th>\n<th>Total, min</th>\n<th>Inference time limit, min</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>2025</strong></td>\n<td><strong>700</strong></td>\n<td><strong>1</strong></td>\n<td><strong>700</strong></td>\n<td><strong>90</strong></td>\n</tr>\n<tr>\n<td>2024</td>\n<td>1100</td>\n<td>4</td>\n<td>4400</td>\n<td>120</td>\n</tr>\n<tr>\n<td>2023</td>\n<td>200</td>\n<td>10</td>\n<td>2000</td>\n<td>120</td>\n</tr>\n<tr>\n<td>2022</td>\n<td>5500</td>\n<td>1</td>\n<td>5500</td>\n<td>540</td>\n</tr>\n<tr>\n<td>2021</td>\n<td>80</td>\n<td>10</td>\n<td>800</td>\n<td>540</td>\n</tr>\n</tbody>\n</table>\n<p>I decided to start with using common models like BirdNet and Bird Vocalization Classifier. </p>\n<h2>BirdNet</h2>\n<p>I used <code>birdnetlib</code> in <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-birdnet-starter\" target=\"_blank\"><strong>the starter notebook</strong></a>. <br>\nHowever, it shows pure performance giving the score of 0.610. The obvious reason is that this model doesn't cover all the existing species. It was trained on birds only and doesn't recognize shghum1. So it provides the predictions on 145 of 206 total species. Keeping this in mind, one can assess the public score of the model as \\( (0.61\\cdot206-61\\cdot0.5)/145 = 0.656 \\). This is not a top score, but it is somewhat higher than current public solutions.</p>\n<h2>Bird Vocalization Classifier</h2>\n<p>BVC is implemented <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-bvc-starter\" target=\"_blank\"><strong>here</strong></a>. This NN is able to predict 143 of 206 species. Also it is  several times slower than BirdNet. Even when executed in several threads, it takes about 10-12 sec to process one recording. So, just for test, I have dropped a quarter of each recording. This gave me the public score of 0.644. Recalculating the public score on the predicted data, we get<br>\n\\( (0.644\\cdot206-(206-143\\cdot0.75)\\cdot0.5)/(143\\cdot 0.75) = 0.776 \\). Not a top score, but better, that BirdNet.</p>\n<p>Have anyone tried these models or something else? What are your results?</p>",
  "messages": [
    {
      "id": 3148686,
      "postDate": "2025-03-13T12:31:24.050Z",
      "content": "<p>In the current competition we have much more time available to make predictions. So, it seems, that we have to build more accurate models, instead of 'fast' models, that were used in the last competitions.</p>\n<table>\n<thead>\n<tr>\n<th>Year</th>\n<th>Number of test records</th>\n<th>Record length, min</th>\n<th>Total, min</th>\n<th>Inference time limit, min</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>2025</strong></td>\n<td><strong>700</strong></td>\n<td><strong>1</strong></td>\n<td><strong>700</strong></td>\n<td><strong>90</strong></td>\n</tr>\n<tr>\n<td>2024</td>\n<td>1100</td>\n<td>4</td>\n<td>4400</td>\n<td>120</td>\n</tr>\n<tr>\n<td>2023</td>\n<td>200</td>\n<td>10</td>\n<td>2000</td>\n<td>120</td>\n</tr>\n<tr>\n<td>2022</td>\n<td>5500</td>\n<td>1</td>\n<td>5500</td>\n<td>540</td>\n</tr>\n<tr>\n<td>2021</td>\n<td>80</td>\n<td>10</td>\n<td>800</td>\n<td>540</td>\n</tr>\n</tbody>\n</table>\n<p>I decided to start with using common models like BirdNet and Bird Vocalization Classifier. </p>\n<h2>BirdNet</h2>\n<p>I used <code>birdnetlib</code> in <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-birdnet-starter\" target=\"_blank\"><strong>the starter notebook</strong></a>. <br>\nHowever, it shows pure performance giving the score of 0.610. The obvious reason is that this model doesn't cover all the existing species. It was trained on birds only and doesn't recognize shghum1. So it provides the predictions on 145 of 206 total species. Keeping this in mind, one can assess the public score of the model as \\( (0.61\\cdot206-61\\cdot0.5)/145 = 0.656 \\). This is not a top score, but it is somewhat higher than current public solutions.</p>\n<h2>Bird Vocalization Classifier</h2>\n<p>BVC is implemented <a href=\"https://www.kaggle.com/code/kdmitrie/bc25-bvc-starter\" target=\"_blank\"><strong>here</strong></a>. This NN is able to predict 143 of 206 species. Also it is  several times slower than BirdNet. Even when executed in several threads, it takes about 10-12 sec to process one recording. So, just for test, I have dropped a quarter of each recording. This gave me the public score of 0.644. Recalculating the public score on the predicted data, we get<br>\n\\( (0.644\\cdot206-(206-143\\cdot0.75)\\cdot0.5)/(143\\cdot 0.75) = 0.776 \\). Not a top score, but better, that BirdNet.</p>\n<p>Have anyone tried these models or something else? What are your results?</p>",
      "rawMarkdown": "In the current competition we have much more time available to make predictions. So, it seems, that we have to build more accurate models, instead of 'fast' models, that were used in the last competitions.\n\n| Year | Number of test records | Record length, min| Total, min| Inference time limit, min|\n| --- | --- |\n| **2025**| **700** | **1** | **700**| **90**|\n| 2024 | 1100 | 4 | 4400 | 120|\n| 2023 | 200 | 10 | 2000 | 120|\n| 2022 | 5500| 1 | 5500 | 540|\n| 2021 | 80| 10 | 800 | 540|\n\n\nI decided to start with using common models like BirdNet and Bird Vocalization Classifier. \n\n## BirdNet\nI used `birdnetlib` in [**the starter notebook**](https://www.kaggle.com/code/kdmitrie/bc25-birdnet-starter). \nHowever, it shows pure performance giving the score of 0.610. The obvious reason is that this model doesn't cover all the existing species. It was trained on birds only and doesn't recognize shghum1. So it provides the predictions on 145 of 206 total species. Keeping this in mind, one can assess the public score of the model as \\\\( (0.61\\cdot206-61\\cdot0.5)/145 = 0.656 \\\\). This is not a top score, but it is somewhat higher than current public solutions.\n\n## Bird Vocalization Classifier\nBVC is implemented [**here**](https://www.kaggle.com/code/kdmitrie/bc25-bvc-starter). This NN is able to predict 143 of 206 species. Also it is  several times slower than BirdNet. Even when executed in several threads, it takes about 10-12 sec to process one recording. So, just for test, I have dropped a quarter of each recording. This gave me the public score of 0.644. Recalculating the public score on the predicted data, we get\n\\\\( (0.644\\cdot206-(206-143\\cdot0.75)\\cdot0.5)/(143\\cdot 0.75) = 0.776 \\\\). Not a top score, but better, that BirdNet.\n\nHave anyone tried these models or something else? What are your results?",
      "votes": 34
    },
    {
      "id": 3164567,
      "postDate": "2025-03-31T19:54:57.760Z",
      "content": "<p>Could you help with output format? In BirdNet workbook output almost all values are zeros. In BVC only unknown are zeros. How could these two workbooks have similar score? None of them look like probabilities and don't sum to 1 in a row, however its a requirement.</p>",
      "rawMarkdown": "Could you help with output format? In BirdNet workbook output almost all values are zeros. In BVC only unknown are zeros. How could these two workbooks have similar score? None of them look like probabilities and don't sum to 1 in a row, however its a requirement.",
      "replies": [
        {
          "id": 3166983,
          "postDate": "2025-04-01T07:37:55.453Z",
          "content": "<p>In this competition, AUC ROC metrics is used. It doesn't need the model's outputs to be probabilities, and it is not sensitive to the exact values. It depends on the order of true and false detections. </p>\n<p>A monotonic function doesn't change this order, so it doesn't matter if we use raw logits of model or process them with <code>sigmoid</code>. BVC produces logits while BirdNet produces probabilities, but the AUC ROC score may be similar.</p>",
          "rawMarkdown": "In this competition, AUC ROC metrics is used. It doesn't need the model's outputs to be probabilities, and it is not sensitive to the exact values. It depends on the order of true and false detections. \n\nA monotonic function doesn't change this order, so it doesn't matter if we use raw logits of model or process them with `sigmoid`. BVC produces logits while BirdNet produces probabilities, but the AUC ROC score may be similar.",
          "votes": 3
        }
      ]
    },
    {
      "id": 3151425,
      "postDate": "2025-03-16T17:06:48.223Z",
      "content": "<p>Thank you for providing a <code>birdnetlib</code> starter notebook! Are there any ways to optimize the BVC model itself so we can run all the test soundscapes?</p>",
      "rawMarkdown": "Thank you for providing a `birdnetlib` starter notebook! Are there any ways to optimize the BVC model itself so we can run all the test soundscapes?",
      "replies": [
        {
          "id": 3152065,
          "postDate": "2025-03-17T11:35:56.853Z",
          "content": "<p>Yes, and the most obvious is to use the embeddings provided by BVC model. Please, have a look at <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier/8\" target=\"_blank\">the documentation</a>. The <code>model_outputs['embedding']</code> contains vector corresponding to the input audio. You can train your own classifier (not necessarily a neural network)  to predict whatever you want.</p>",
          "rawMarkdown": "Yes, and the most obvious is to use the embeddings provided by BVC model. Please, have a look at [the documentation](https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier/8). The `model_outputs['embedding']` contains vector corresponding to the input audio. You can train your own classifier (not necessarily a neural network)  to predict whatever you want.",
          "votes": 2,
          "replies": [
            {
              "id": 3152245,
              "postDate": "2025-03-17T15:59:49.333Z",
              "content": "<p>Awesome, thank you! Building some simple supervised models on the embeddings first makes a lot of sense.</p>",
              "rawMarkdown": "Awesome, thank you! Building some simple supervised models on the embeddings first makes a lot of sense."
            },
            {
              "id": 3153246,
              "postDate": "2025-03-18T15:04:05.957Z",
              "content": "<p>Why use the embeddings rather than fully finetuning the model?</p>",
              "rawMarkdown": "Why use the embeddings rather than fully finetuning the model?",
              "votes": 1
            },
            {
              "id": 3153302,
              "postDate": "2025-03-18T16:23:46.193Z",
              "content": "<p>Both can make sense, but since BVC is trained on bioacoustics (i.e. birds) the embeddings should already be pretty good.</p>\n<p>An upside of using the embeddings directly is that we can fit very lightweight models on it, like logistic regression or decision trees. It also allow you to iterate very quickly and explore differrent models. Finetuning the model will have a longer iteration cycle.</p>\n<p>To get started I would also recommend checking out <a href=\"https://openl3.readthedocs.io/en/latest/tutorial.html\" target=\"_blank\">openl3</a>. An elegant way to get audio embeddings. Use <code>content_type=\"env\"</code> to access a model trained on environmental sound.</p>",
              "rawMarkdown": "Both can make sense, but since BVC is trained on bioacoustics (i.e. birds) the embeddings should already be pretty good.\n\nAn upside of using the embeddings directly is that we can fit very lightweight models on it, like logistic regression or decision trees. It also allow you to iterate very quickly and explore differrent models. Finetuning the model will have a longer iteration cycle.\n\nTo get started I would also recommend checking out [openl3](https://openl3.readthedocs.io/en/latest/tutorial.html). An elegant way to get audio embeddings. Use `content_type=\"env\"` to access a model trained on environmental sound.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3153703,
      "postDate": "2025-03-19T04:52:52.437Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3164567,
      "author_name": "zharkomi",
      "author_url": "",
      "post_date": "2025-03-31T19:54:57.760000",
      "content": "<p>Could you help with output format? In BirdNet workbook output almost all values are zeros. In BVC only unknown are zeros. How could these two workbooks have similar score? None of them look like probabilities and don't sum to 1 in a row, however its a requirement.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3166983,
          "author_name": "Konstantin Dmitriev",
          "author_url": "",
          "post_date": "2025-04-01T07:37:55.453000",
          "content": "<p>In this competition, AUC ROC metrics is used. It doesn't need the model's outputs to be probabilities, and it is not sensitive to the exact values. It depends on the order of true and false detections. </p>\n<p>A monotonic function doesn't change this order, so it doesn't matter if we use raw logits of model or process them with <code>sigmoid</code>. BVC produces logits while BirdNet produces probabilities, but the AUC ROC score may be similar.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3151425,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2025-03-16T17:06:48.223000",
      "content": "<p>Thank you for providing a <code>birdnetlib</code> starter notebook! Are there any ways to optimize the BVC model itself so we can run all the test soundscapes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3152065,
          "author_name": "Konstantin Dmitriev",
          "author_url": "",
          "post_date": "2025-03-17T11:35:56.853000",
          "content": "<p>Yes, and the most obvious is to use the embeddings provided by BVC model. Please, have a look at <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier/8\" target=\"_blank\">the documentation</a>. The <code>model_outputs['embedding']</code> contains vector corresponding to the input audio. You can train your own classifier (not necessarily a neural network)  to predict whatever you want.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3152245,
              "author_name": "Carlo",
              "author_url": "",
              "post_date": "2025-03-17T15:59:49.333000",
              "content": "<p>Awesome, thank you! Building some simple supervised models on the embeddings first makes a lot of sense.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3153246,
              "author_name": "thacrobatheskis",
              "author_url": "",
              "post_date": "2025-03-18T15:04:05.957000",
              "content": "<p>Why use the embeddings rather than fully finetuning the model?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3153302,
              "author_name": "Carlo",
              "author_url": "",
              "post_date": "2025-03-18T16:23:46.193000",
              "content": "<p>Both can make sense, but since BVC is trained on bioacoustics (i.e. birds) the embeddings should already be pretty good.</p>\n<p>An upside of using the embeddings directly is that we can fit very lightweight models on it, like logistic regression or decision trees. It also allow you to iterate very quickly and explore differrent models. Finetuning the model will have a longer iteration cycle.</p>\n<p>To get started I would also recommend checking out <a href=\"https://openl3.readthedocs.io/en/latest/tutorial.html\" target=\"_blank\">openl3</a>. An elegant way to get audio embeddings. Use <code>content_type=\"env\"</code> to access a model trained on environmental sound.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3153703,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-19T04:52:52.437000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3148686": "In the current competition we have much more time available to make predictions. So, it seems, that we have to build more accurate models, instead of 'fast' models, that were used in the last competitions.\n\n| Year | Number of test records | Record length, min| Total, min| Inference time limit, min|\n| --- | --- |\n| **2025**| **700** | **1** | **700**| **90**|\n| 2024 | 1100 | 4 | 4400 | 120|\n| 2023 | 200 | 10 | 2000 | 120|\n| 2022 | 5500| 1 | 5500 | 540|\n| 2021 | 80| 10 | 800 | 540|\n\n\nI decided to start with using common models like BirdNet and Bird Vocalization Classifier. \n\n## BirdNet\nI used `birdnetlib` in [**the starter notebook**](https://www.kaggle.com/code/kdmitrie/bc25-birdnet-starter). \nHowever, it shows pure performance giving the score of 0.610. The obvious reason is that this model doesn't cover all the existing species. It was trained on birds only and doesn't recognize shghum1. So it provides the predictions on 145 of 206 total species. Keeping this in mind, one can assess the public score of the model as \\\\( (0.61\\cdot206-61\\cdot0.5)/145 = 0.656 \\\\). This is not a top score, but it is somewhat higher than current public solutions.\n\n## Bird Vocalization Classifier\nBVC is implemented [**here**](https://www.kaggle.com/code/kdmitrie/bc25-bvc-starter). This NN is able to predict 143 of 206 species. Also it is  several times slower than BirdNet. Even when executed in several threads, it takes about 10-12 sec to process one recording. So, just for test, I have dropped a quarter of each recording. This gave me the public score of 0.644. Recalculating the public score on the predicted data, we get\n\\\\( (0.644\\cdot206-(206-143\\cdot0.75)\\cdot0.5)/(143\\cdot 0.75) = 0.776 \\\\). Not a top score, but better, that BirdNet.\n\nHave anyone tried these models or something else? What are your results?",
    "3164567": "Could you help with output format? In BirdNet workbook output almost all values are zeros. In BVC only unknown are zeros. How could these two workbooks have similar score? None of them look like probabilities and don't sum to 1 in a row, however its a requirement.",
    "3151425": "Thank you for providing a `birdnetlib` starter notebook! Are there any ways to optimize the BVC model itself so we can run all the test soundscapes?",
    "3153703": ""
  }
}