{
  "id": 583384,
  "title": "11th solution",
  "url": "/competitions/birdclef-2025/writeups/baiph-11th-solution",
  "author_name": "",
  "post_date": "2025-06-06T12:46:51.031091900Z",
  "votes": 14,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thank you to the organizers for organizing such an interesting competition, and also thank the excellent plans from previous sessions for inspiring me. Also, congratulations to all the top winners. Next, I will introduce my solution.</p>\n<h1>1、Training Data</h1>\n<ul>\n<li>The 206 categories have been expanded to 316</li>\n</ul>\n<p>(1) 25 year competition data: train_audio and train_soundscapes</p>\n<p>(2) select 110 categories (categories with a sample size of less than 10) from the data of previous competitions</p>\n<p>This additional category is mainly for building local cv. Unfortunately, this cv strategy is ineffective. However, this mixed training improved my lb, so I maintained this operation</p>\n<ul>\n<li>The maximum sample size for each category is 500. Categories smaller than 10 will be upsampled</li>\n<li>Only remove human voices from part of the CAS data</li>\n</ul>\n<h1>2、<strong>Model Architecture</strong></h1>\n<p>Sed models opened by&nbsp;<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2nd place solution of 2023</a>. I added the code related to pseudo-labels on this basis</p>\n<p>backbones are：</p>\n<ul>\n<li>tf_efficientnetv2_b3</li>\n<li>tf_efficientnetv2_s</li>\n</ul>\n<p>All of them are trained on 10sec clip.</p>\n<h1>3、Loss function</h1>\n<p>Using ce loss, compared with BCE loss, there is a qualitative improvement (0.83→0.88) in multiple models.</p>\n<h1>4、<strong>Mel Spectrogram Parameters</strong></h1>\n<pre><code>{ ,  ,  ,  ,  ,  ,  ,  }\n\n{ ,  ,  ,  ,  ,  ,  ,  }\n</code></pre>\n<h1>5、pseudo label</h1>\n<ul>\n<li>Select high-quality pseudo-labels based on the entropy-based screening strategy mentioned in <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511596\" target=\"_blank\">rank 10  of last year 24</a></li>\n</ul>\n<pre><code>\nimport numpy as np\nepsilon = 1e-12\n\npredf[columns].values\nprint(preprobs.copy()\nentropies = -np.sum(probs * np.log(probs + epsilon), axis=1)\nprint(entropies.shape)\n\n\ntopindices = np.argsort(entropies)[:int(len(entropies) * 0.2)]\ntoppseudo1010probs.shape)\n\n\nfor i in range(toppseudoprobs = toppseudoprobs, 92)\n</code></pre>\n<ul>\n<li>The real dataset and the pseudo-label dataset are concatenated into the model, and the loss of the pl part is down weighted</li>\n</ul>\n<p>batch_size = 96，pl_batch_size = 16</p>\n<h1>6、<strong>Ensemble and Post-processing</strong></h1>\n<ul>\n<li>ensemble model</li>\n</ul>\n<p>5 ✖️&nbsp;v2b3 + 1✖️v2s</p>\n<p>Public Score: 0.920<br>\nPrivate Score: 0.919</p>\n<ul>\n<li>post precessing</li>\n</ul>\n<p>The post-processing is the same as <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">the rank 6 in 2024</a>，An increase of approximately 0.001</p>\n<pre><code> smooth_array_general(array, w=[., ., ., ., .]):\n     = np.zeros_like(array)\n     = array.shape[]\n     = len(w) // \n\n     t in range(timesteps):\n         i, weight in enumerate(w):\n             = t - radius + i\n             index &lt; :\n                [t] += array[] * weight\n             index &gt;= timesteps:\n                [t] += array[-] * weight\n            :\n                [t] += array[index] * weight\n     c in range(array.shape[]):\n        [:, c] = smoothed_array[:, c] * . + smoothed_array[:, c].mean(keepdims=True) * .\n     smoothed_array\n</code></pre>\n<h1>7、not work</h1>\n<ul>\n<li>CNN</li>\n<li>rms sample</li>\n<li>remove all human voices</li>\n</ul>",
  "messages": [
    {
      "id": "3218618",
      "postDate": "06/06/2025 12:46:51",
      "content": "<p>Thank you to the organizers for organizing such an interesting competition, and also thank the excellent plans from previous sessions for inspiring me. Also, congratulations to all the top winners. Next, I will introduce my solution.</p>\n<h1>1、Training Data</h1>\n<ul>\n<li>The 206 categories have been expanded to 316</li>\n</ul>\n<p>(1) 25 year competition data: train_audio and train_soundscapes</p>\n<p>(2) select 110 categories (categories with a sample size of less than 10) from the data of previous competitions</p>\n<p>This additional category is mainly for building local cv. Unfortunately, this cv strategy is ineffective. However, this mixed training improved my lb, so I maintained this operation</p>\n<ul>\n<li>The maximum sample size for each category is 500. Categories smaller than 10 will be upsampled</li>\n<li>Only remove human voices from part of the CAS data</li>\n</ul>\n<h1>2、<strong>Model Architecture</strong></h1>\n<p>Sed models opened by&nbsp;<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">2nd place solution of 2023</a>. I added the code related to pseudo-labels on this basis</p>\n<p>backbones are：</p>\n<ul>\n<li>tf_efficientnetv2_b3</li>\n<li>tf_efficientnetv2_s</li>\n</ul>\n<p>All of them are trained on 10sec clip.</p>\n<h1>3、Loss function</h1>\n<p>Using ce loss, compared with BCE loss, there is a qualitative improvement (0.83→0.88) in multiple models.</p>\n<h1>4、<strong>Mel Spectrogram Parameters</strong></h1>\n<pre><code>{ ,  ,  ,  ,  ,  ,  ,  }\n\n{ ,  ,  ,  ,  ,  ,  ,  }\n</code></pre>\n<h1>5、pseudo label</h1>\n<ul>\n<li>Select high-quality pseudo-labels based on the entropy-based screening strategy mentioned in <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511596\" target=\"_blank\">rank 10  of last year 24</a></li>\n</ul>\n<pre><code>\nimport numpy as np\nepsilon = 1e-12\n\npredf[columns].values\nprint(preprobs.copy()\nentropies = -np.sum(probs * np.log(probs + epsilon), axis=1)\nprint(entropies.shape)\n\n\ntopindices = np.argsort(entropies)[:int(len(entropies) * 0.2)]\ntoppseudo1010probs.shape)\n\n\nfor i in range(toppseudoprobs = toppseudoprobs, 92)\n</code></pre>\n<ul>\n<li>The real dataset and the pseudo-label dataset are concatenated into the model, and the loss of the pl part is down weighted</li>\n</ul>\n<p>batch_size = 96，pl_batch_size = 16</p>\n<h1>6、<strong>Ensemble and Post-processing</strong></h1>\n<ul>\n<li>ensemble model</li>\n</ul>\n<p>5 ✖️&nbsp;v2b3 + 1✖️v2s</p>\n<p>Public Score: 0.920<br>\nPrivate Score: 0.919</p>\n<ul>\n<li>post precessing</li>\n</ul>\n<p>The post-processing is the same as <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">the rank 6 in 2024</a>，An increase of approximately 0.001</p>\n<pre><code> smooth_array_general(array, w=[., ., ., ., .]):\n     = np.zeros_like(array)\n     = array.shape[]\n     = len(w) // \n\n     t in range(timesteps):\n         i, weight in enumerate(w):\n             = t - radius + i\n             index &lt; :\n                [t] += array[] * weight\n             index &gt;= timesteps:\n                [t] += array[-] * weight\n            :\n                [t] += array[index] * weight\n     c in range(array.shape[]):\n        [:, c] = smoothed_array[:, c] * . + smoothed_array[:, c].mean(keepdims=True) * .\n     smoothed_array\n</code></pre>\n<h1>7、not work</h1>\n<ul>\n<li>CNN</li>\n<li>rms sample</li>\n<li>remove all human voices</li>\n</ul>",
      "rawMarkdown": "Thank you to the organizers for organizing such an interesting competition, and also thank the excellent plans from previous sessions for inspiring me. Also, congratulations to all the top winners. Next, I will introduce my solution.\n\n# 1、Training Data\n\n- The 206 categories have been expanded to 316\n\n(1) 25 year competition data: train_audio and train_soundscapes\n\n(2) select 110 categories (categories with a sample size of less than 10) from the data of previous competitions\n\nThis additional category is mainly for building local cv. Unfortunately, this cv strategy is ineffective. However, this mixed training improved my lb, so I maintained this operation\n\n- The maximum sample size for each category is 500. Categories smaller than 10 will be upsampled\n- Only remove human voices from part of the CAS data\n\n# 2、**Model Architecture**\n\nSed models opened by [2nd place solution of 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707). I added the code related to pseudo-labels on this basis\n\nbackbones are：\n\n- tf_efficientnetv2_b3\n- tf_efficientnetv2_s\n\nAll of them are trained on 10sec clip.\n\n# 3、Loss function\n\nUsing ce loss, compared with BCE loss, there is a qualitative improvement (0.83→0.88) in multiple models.\n\n# 4、**Mel Spectrogram Parameters**\n\n```\n{'sample_rate': 32000, 'n_mels': 256, 'image_size': 300, 'f_min': 90, 'f_max': 14000, 'n_fft': 1536, 'normalized': True, 'hop_length': 535}\n\n{'sample_rate': 32000, 'n_mels': 256, 'image_size': 300, 'f_min': 50, 'f_max': 14000, 'n_fft': 1024, 'normalized': True, 'hop_length': 535}\n```\n\n# 5、pseudo label\n\n- Select high-quality pseudo-labels based on the entropy-based screening strategy mentioned in [rank 10  of last year 24](https://www.kaggle.com/competitions/birdclef-2024/discussion/511596)\n\n```markdown\n# 示例代码\nimport numpy as np\nepsilon = 1e-12\n\npre_probs = sub_df[columns].values\nprint(pre_probs.shape)\n\nprobs = pre_probs.copy()\nentropies = -np.sum(probs * np.log(probs + epsilon), axis=1)\nprint(entropies.shape)\n\n# 筛选前 20% 的伪标签\ntop_10_indices = np.argsort(entropies)[:int(len(entropies) * 0.2)]\ntop_10_pseudo_probs = probs[top_10_indices]\nprint(top_10_pseudo_probs.shape)\n\n# 对于每个类别的伪标签，将低于前 92% 的标签值设置为 0\nfor i in range(top_10_pseudo_probs.shape[1]):\n    class_probs = top_10_pseudo_probs[:, i]\n    threshold = np.percentile(class_probs, 92)\n    top_10_pseudo_probs[class_probs < threshold, i] = 0\nprint(top_10_pseudo_probs.shape)\n\n```\n\n- The real dataset and the pseudo-label dataset are concatenated into the model, and the loss of the pl part is down weighted\n\nbatch_size = 96，pl_batch_size = 16\n\n# 6、**Ensemble and Post-processing**\n\n- ensemble model\n\n5 ✖️ v2b3 + 1✖️v2s\n\nPublic Score: 0.920\nPrivate Score: 0.919\n\n\n- post precessing\n\nThe post-processing is the same as [the rank 6 in 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511527)，An increase of approximately 0.001\n\n```\ndef smooth_array_general(array, w=[0.1, 0.2, 0.4, 0.2, 0.1]):\n    smoothed_array = np.zeros_like(array)\n    timesteps = array.shape[0]\n    radius = len(w) // 2\n\n    for t in range(timesteps):\n        for i, weight in enumerate(w):\n            index = t - radius + i\n            if index < 0:\n                smoothed_array[t] += array[0] * weight\n            elif index >= timesteps:\n                smoothed_array[t] += array[-1] * weight\n            else:\n                smoothed_array[t] += array[index] * weight\n    for c in range(array.shape[1]):\n        smoothed_array[:, c] = smoothed_array[:, c] * 0.8 + smoothed_array[:, c].mean(keepdims=True) * 0.2\n    return smoothed_array\n```\n\n# 7、not work\n\n- CNN\n- rms sample\n- remove all human voices",
      "votes": null
    },
    {
      "id": "3218854",
      "postDate": "06/06/2025 19:56:17",
      "content": "<p>Congrats! If rms sampling didn’t work, how did you sample the training data?</p>",
      "rawMarkdown": "Congrats! If rms sampling didn’t work, how did you sample the training data?",
      "votes": null
    },
    {
      "id": "3218906",
      "postDate": "06/06/2025 22:21:57",
      "content": "<p>random sample</p>",
      "rawMarkdown": "random sample",
      "votes": null
    },
    {
      "id": "3218916",
      "postDate": "06/06/2025 23:18:27",
      "content": "<p>Congrats on your solo win, bro! </p>",
      "rawMarkdown": "Congrats on your solo win, bro!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218854,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "06/06/2025 19:56:17",
      "content": "<p>Congrats! If rms sampling didn’t work, how did you sample the training data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218906,
          "author_name": "pingfan",
          "author_url": "",
          "post_date": "06/06/2025 22:21:57",
          "content": "<p>random sample</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218916,
      "author_name": "shanzhong8",
      "author_url": "",
      "post_date": "06/06/2025 23:18:27",
      "content": "<p>Congrats on your solo win, bro! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218618": "Thank you to the organizers for organizing such an interesting competition, and also thank the excellent plans from previous sessions for inspiring me. Also, congratulations to all the top winners. Next, I will introduce my solution.\n\n# 1、Training Data\n\n- The 206 categories have been expanded to 316\n\n(1) 25 year competition data: train_audio and train_soundscapes\n\n(2) select 110 categories (categories with a sample size of less than 10) from the data of previous competitions\n\nThis additional category is mainly for building local cv. Unfortunately, this cv strategy is ineffective. However, this mixed training improved my lb, so I maintained this operation\n\n- The maximum sample size for each category is 500. Categories smaller than 10 will be upsampled\n- Only remove human voices from part of the CAS data\n\n# 2、**Model Architecture**\n\nSed models opened by [2nd place solution of 2023](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707). I added the code related to pseudo-labels on this basis\n\nbackbones are：\n\n- tf_efficientnetv2_b3\n- tf_efficientnetv2_s\n\nAll of them are trained on 10sec clip.\n\n# 3、Loss function\n\nUsing ce loss, compared with BCE loss, there is a qualitative improvement (0.83→0.88) in multiple models.\n\n# 4、**Mel Spectrogram Parameters**\n\n```\n{'sample_rate': 32000, 'n_mels': 256, 'image_size': 300, 'f_min': 90, 'f_max': 14000, 'n_fft': 1536, 'normalized': True, 'hop_length': 535}\n\n{'sample_rate': 32000, 'n_mels': 256, 'image_size': 300, 'f_min': 50, 'f_max': 14000, 'n_fft': 1024, 'normalized': True, 'hop_length': 535}\n```\n\n# 5、pseudo label\n\n- Select high-quality pseudo-labels based on the entropy-based screening strategy mentioned in [rank 10  of last year 24](https://www.kaggle.com/competitions/birdclef-2024/discussion/511596)\n\n```markdown\n# 示例代码\nimport numpy as np\nepsilon = 1e-12\n\npre_probs = sub_df[columns].values\nprint(pre_probs.shape)\n\nprobs = pre_probs.copy()\nentropies = -np.sum(probs * np.log(probs + epsilon), axis=1)\nprint(entropies.shape)\n\n# 筛选前 20% 的伪标签\ntop_10_indices = np.argsort(entropies)[:int(len(entropies) * 0.2)]\ntop_10_pseudo_probs = probs[top_10_indices]\nprint(top_10_pseudo_probs.shape)\n\n# 对于每个类别的伪标签，将低于前 92% 的标签值设置为 0\nfor i in range(top_10_pseudo_probs.shape[1]):\n    class_probs = top_10_pseudo_probs[:, i]\n    threshold = np.percentile(class_probs, 92)\n    top_10_pseudo_probs[class_probs < threshold, i] = 0\nprint(top_10_pseudo_probs.shape)\n\n```\n\n- The real dataset and the pseudo-label dataset are concatenated into the model, and the loss of the pl part is down weighted\n\nbatch_size = 96，pl_batch_size = 16\n\n# 6、**Ensemble and Post-processing**\n\n- ensemble model\n\n5 ✖️ v2b3 + 1✖️v2s\n\nPublic Score: 0.920\nPrivate Score: 0.919\n\n\n- post precessing\n\nThe post-processing is the same as [the rank 6 in 2024](https://www.kaggle.com/competitions/birdclef-2024/discussion/511527)，An increase of approximately 0.001\n\n```\ndef smooth_array_general(array, w=[0.1, 0.2, 0.4, 0.2, 0.1]):\n    smoothed_array = np.zeros_like(array)\n    timesteps = array.shape[0]\n    radius = len(w) // 2\n\n    for t in range(timesteps):\n        for i, weight in enumerate(w):\n            index = t - radius + i\n            if index < 0:\n                smoothed_array[t] += array[0] * weight\n            elif index >= timesteps:\n                smoothed_array[t] += array[-1] * weight\n            else:\n                smoothed_array[t] += array[index] * weight\n    for c in range(array.shape[1]):\n        smoothed_array[:, c] = smoothed_array[:, c] * 0.8 + smoothed_array[:, c].mean(keepdims=True) * 0.2\n    return smoothed_array\n```\n\n# 7、not work\n\n- CNN\n- rms sample\n- remove all human voices",
    "3218854": "Congrats! If rms sampling didn’t work, how did you sample the training data?",
    "3218906": "random sample",
    "3218916": "Congrats on your solo win, bro!"
  },
  "source": "meta"
}