{
  "id": 583310,
  "title": "10th Solution",
  "url": "/competitions/birdclef-2025/writeups/lhwcv-10th-solution",
  "author_name": "",
  "post_date": "2025-06-06T03:29:42.080Z",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to the organizers for hosting this competition, and also to the many Kaggler participants from previous years whose solutions have provided me with great inspiration.<br>\nBirdCLEF is a challenging competition, Although the dataset shift isn't extremely large, the lack of a validation set from the same distribution makes training and evaluation quite difficult.<br>\nI adopted the following approach during the competition:</p>\n<h2>1. Establishing Baseline Model Performance</h2>\n<p>I used the entire train_audio set and validated directly on the leaderboard. I found the following model architectures performed best:</p>\n<pre><code> &gt;  &gt;  &gt; \n</code></pre>\n<p>This was just an initial experiment, so it may have been influenced by other factors and not be generally applicable. But due to limited online submission opportunities, I stuck with SED + CE loss for the rest of the experiments.<br>\nModels trained in this stage are referred to as stage1 models.</p>\n<h2>2. Domain Adaptation &amp; pseudo-labeled data</h2>\n<p>As seen in previous competitions, leveraging unlabeled data (train_soundscapes) can lead to significant gains.<br>\nI used the following method:</p>\n<ol>\n<li>Use the stage1 models to predict on train_soundscapes, generating soft labels</li>\n<li>Apply high and low thresholds (thresh_low, thresh_high) to extract confident negative and positive samples</li>\n<li>These samples (assumed to be relatively clean) were added to train_audio to train new models — the stage2_models. The ambiguous samples in between were excluded</li>\n<li>Predict again on train_soundscapes using stage2_models to get updated soft labels<br>\nThrough this iterative process, we obtained relatively clean pseudo-labeled data, referred to as pp_data_clean.</li>\n</ol>\n<h2>3. Data Processing</h2>\n<p>Following community discussions, I removed samples with obvious human speech — this cleaned dataset is called train_audio_clean.<br>\nTo further improve diversity and make data clean, I used trained models to remove segments that clearly did not contain the target species, sampling at 5-second intervals, forming a new dataset: train_audio_clean_v2.<br>\nBoth datasets were used for training.</p>\n<h2>4. Mel Spectrogram Parameters</h2>\n<p>I used mel spectrograms with the following resolutions:</p>\n<ul>\n<li>384x160</li>\n<li>384x256</li>\n<li>320x192</li>\n<li>320x160<br>\nfmin = 0, fmax = 16000， n_fft=1536/2048</li>\n</ul>\n<h2>5. Loss Function</h2>\n<p>I adopted a modified hybrid  loss with CE.<br>\nAlthough public notebooks showed that ConvNeXt performed well, using standard CE loss led to a lot of noisy predictions (many false positives).<br>\nSo, I gradually incorporated the following loss to better stabilize training:</p>\n<pre><code>threshold = -\nneg_mask = (y == )\nnegative_logits = logits[neg_mask]\npenalty = F.relu(negative_logits - threshold)\nsorted_penalty, _ = torch.sort(penalty)\ncutoff_index = ((sorted_penalty) * )\nselected_penalty = sorted_penalty[:cutoff_index]\nmean_penalty = selected_penalty.mean()\nloss = ce_loss +  * mean_penalty\n</code></pre>\n<p>In some models, I also penalized positive samples with confidence scores that were too low.</p>\n<h2>6. Post Processing</h2>\n<ul>\n<li>smooth with kernel: [0.02, 0.08, 0.8, 0.08, 0.02]</li>\n<li>and average:</li>\n</ul>\n<pre><code>def get_mean_scalesref_freq\n  alpha_max  \n  alpha_min  \n  \n  alpha  alpha_min  ref_freq  alpha_max  alpha_min\n   alpha\n...\n    lenn_classes\n   a  alpha\n   pred_prob   pred_prob   a  pred_prob .meankeepdimsTrue  a\n</code></pre>\n<h2>7. Some Model Results</h2>\n<ul>\n<li>ConvNeXt Tiny (320x192): Public: 0.901   Private: 0.914</li>\n<li>3x ConvNeXt Tiny Ensemble:  Public: 0.908 Private: 0.921</li>\n<li>EfficientNetV2-S (384x160) Public: 0.904 Private: 0.906</li>\n<li>EfficientNetV2-S (384x256)  Public: 0.896. Private: 0.914<br>\nDiversity was controlled by varying the dataset used and the proportion and quality of mixed-in pp_clean_data.<br>\nI ultimately ensembled 10 models in total,  Public: 0.915 Private: 0.921 . <br>\nHowever, due to the instability of the leaderboard (LB), I wasn’t able to select the optimal combination of models.</li>\n</ul>\n<h2>8. Other Details</h2>\n<ul>\n<li>10s/15s for train, infer on 5s</li>\n<li>AdamW + CosineAnnealingLR， LR： 2/3*e-4, WD: 1e-3/1e-4</li>\n<li>Mixup with additive, frequency, and time masking</li>\n<li>With and without noise augmentation</li>\n<li>Many other common techniques like label smoothing, flip, etc.</li>\n</ul>",
  "messages": [
    {
      "id": "3218258",
      "postDate": "06/06/2025 03:08:55",
      "content": "<p>Thanks to the organizers for hosting this competition, and also to the many Kaggler participants from previous years whose solutions have provided me with great inspiration.<br>\nBirdCLEF is a challenging competition, Although the dataset shift isn't extremely large, the lack of a validation set from the same distribution makes training and evaluation quite difficult.<br>\nI adopted the following approach during the competition:</p>\n<h2>1. Establishing Baseline Model Performance</h2>\n<p>I used the entire train_audio set and validated directly on the leaderboard. I found the following model architectures performed best:</p>\n<pre><code> &gt;  &gt;  &gt; \n</code></pre>\n<p>This was just an initial experiment, so it may have been influenced by other factors and not be generally applicable. But due to limited online submission opportunities, I stuck with SED + CE loss for the rest of the experiments.<br>\nModels trained in this stage are referred to as stage1 models.</p>\n<h2>2. Domain Adaptation &amp; pseudo-labeled data</h2>\n<p>As seen in previous competitions, leveraging unlabeled data (train_soundscapes) can lead to significant gains.<br>\nI used the following method:</p>\n<ol>\n<li>Use the stage1 models to predict on train_soundscapes, generating soft labels</li>\n<li>Apply high and low thresholds (thresh_low, thresh_high) to extract confident negative and positive samples</li>\n<li>These samples (assumed to be relatively clean) were added to train_audio to train new models — the stage2_models. The ambiguous samples in between were excluded</li>\n<li>Predict again on train_soundscapes using stage2_models to get updated soft labels<br>\nThrough this iterative process, we obtained relatively clean pseudo-labeled data, referred to as pp_data_clean.</li>\n</ol>\n<h2>3. Data Processing</h2>\n<p>Following community discussions, I removed samples with obvious human speech — this cleaned dataset is called train_audio_clean.<br>\nTo further improve diversity and make data clean, I used trained models to remove segments that clearly did not contain the target species, sampling at 5-second intervals, forming a new dataset: train_audio_clean_v2.<br>\nBoth datasets were used for training.</p>\n<h2>4. Mel Spectrogram Parameters</h2>\n<p>I used mel spectrograms with the following resolutions:</p>\n<ul>\n<li>384x160</li>\n<li>384x256</li>\n<li>320x192</li>\n<li>320x160<br>\nfmin = 0, fmax = 16000， n_fft=1536/2048</li>\n</ul>\n<h2>5. Loss Function</h2>\n<p>I adopted a modified hybrid  loss with CE.<br>\nAlthough public notebooks showed that ConvNeXt performed well, using standard CE loss led to a lot of noisy predictions (many false positives).<br>\nSo, I gradually incorporated the following loss to better stabilize training:</p>\n<pre><code>threshold = -\nneg_mask = (y == )\nnegative_logits = logits[neg_mask]\npenalty = F.relu(negative_logits - threshold)\nsorted_penalty, _ = torch.sort(penalty)\ncutoff_index = ((sorted_penalty) * )\nselected_penalty = sorted_penalty[:cutoff_index]\nmean_penalty = selected_penalty.mean()\nloss = ce_loss +  * mean_penalty\n</code></pre>\n<p>In some models, I also penalized positive samples with confidence scores that were too low.</p>\n<h2>6. Post Processing</h2>\n<ul>\n<li>smooth with kernel: [0.02, 0.08, 0.8, 0.08, 0.02]</li>\n<li>and average:</li>\n</ul>\n<pre><code>def get_mean_scalesref_freq\n  alpha_max  \n  alpha_min  \n  \n  alpha  alpha_min  ref_freq  alpha_max  alpha_min\n   alpha\n...\n    lenn_classes\n   a  alpha\n   pred_prob   pred_prob   a  pred_prob .meankeepdimsTrue  a\n</code></pre>\n<h2>7. Some Model Results</h2>\n<ul>\n<li>ConvNeXt Tiny (320x192): Public: 0.901   Private: 0.914</li>\n<li>3x ConvNeXt Tiny Ensemble:  Public: 0.908 Private: 0.921</li>\n<li>EfficientNetV2-S (384x160) Public: 0.904 Private: 0.906</li>\n<li>EfficientNetV2-S (384x256)  Public: 0.896. Private: 0.914<br>\nDiversity was controlled by varying the dataset used and the proportion and quality of mixed-in pp_clean_data.<br>\nI ultimately ensembled 10 models in total,  Public: 0.915 Private: 0.921 . <br>\nHowever, due to the instability of the leaderboard (LB), I wasn’t able to select the optimal combination of models.</li>\n</ul>\n<h2>8. Other Details</h2>\n<ul>\n<li>10s/15s for train, infer on 5s</li>\n<li>AdamW + CosineAnnealingLR， LR： 2/3*e-4, WD: 1e-3/1e-4</li>\n<li>Mixup with additive, frequency, and time masking</li>\n<li>With and without noise augmentation</li>\n<li>Many other common techniques like label smoothing, flip, etc.</li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers for hosting this competition, and also to the many Kaggler participants from previous years whose solutions have provided me with great inspiration.\nBirdCLEF is a challenging competition, Although the dataset shift isn't extremely large, the lack of a validation set from the same distribution makes training and evaluation quite difficult.\n\nI adopted the following approach during the competition:\n\n## 1. Establishing Baseline Model Performance\n\nI used the entire train_audio set and validated directly on the leaderboard. I found the following model architectures performed best:\n\n```\n(SED + CE loss) > (SED + FocalBCE) > (CNN + FocalBCE) > (CNN + CE)\n```\n\nThis was just an initial experiment, so it may have been influenced by other factors and not be generally applicable. But due to limited online submission opportunities, I stuck with SED + CE loss for the rest of the experiments.\nModels trained in this stage are referred to as stage1 models.\n\n## 2. Domain Adaptation & pseudo-labeled data\n\nAs seen in previous competitions, leveraging unlabeled data (train_soundscapes) can lead to significant gains.\nI used the following method:\n\n1. Use the stage1 models to predict on train_soundscapes, generating soft labels\n2. Apply high and low thresholds (thresh_low, thresh_high) to extract confident negative and positive samples\n3. These samples (assumed to be relatively clean) were added to train_audio to train new models — the stage2_models. The ambiguous samples in between were excluded\n4. Predict again on train_soundscapes using stage2_models to get updated soft labels\n\nThrough this iterative process, we obtained relatively clean pseudo-labeled data, referred to as pp_data_clean.\n\n## 3. Data Processing\n\nFollowing community discussions, I removed samples with obvious human speech — this cleaned dataset is called train_audio_clean.\nTo further improve diversity and make data clean, I used trained models to remove segments that clearly did not contain the target species, sampling at 5-second intervals, forming a new dataset: train_audio_clean_v2.\nBoth datasets were used for training.\n\n## 4. Mel Spectrogram Parameters\n\nI used mel spectrograms with the following resolutions:\n\n- 384x160\n- 384x256\n- 320x192\n- 320x160\n\nfmin = 0, fmax = 16000， n_fft=1536/2048\n\n## 5. Loss Function\n\nI adopted a modified hybrid  loss with CE.\nAlthough public notebooks showed that ConvNeXt performed well, using standard CE loss led to a lot of noisy predictions (many false positives).\nSo, I gradually incorporated the following loss to better stabilize training:\n\n```python\nthreshold = -5\nneg_mask = (y == 0.0)\nnegative_logits = logits[neg_mask]\npenalty = F.relu(negative_logits - threshold)\nsorted_penalty, _ = torch.sort(penalty)\ncutoff_index = int(len(sorted_penalty) * 0.95)\nselected_penalty = sorted_penalty[:cutoff_index]\nmean_penalty = selected_penalty.mean()\nloss = ce_loss + 0.1 * mean_penalty\n```\n\nIn some models, I also penalized positive samples with confidence scores that were too low.\n\n## 6. Post Processing\n- smooth with kernel: [0.02, 0.08, 0.8, 0.08, 0.02]\n- and average:\n```\ndef get_mean_scales(ref_freq):\n    alpha_max = 0.3\n    alpha_min = 0.1\n    # rare classed more recall\n    alpha = alpha_min + ref_freq * (alpha_max - alpha_min)\n    return alpha\n...\n   for c in len(n_classes):\n     a = alpha[c]\n     pred_prob[:, c] = pred_prob[:, c] * (1-a) + pred_prob[:, c].mean(keepdims=True) * a\n    \n \n```\n\n## 7. Some Model Results\n\n- ConvNeXt Tiny (320x192): Public: 0.901   Private: 0.914\n\n- 3x ConvNeXt Tiny Ensemble:  Public: 0.908 Private: 0.921\n\n- EfficientNetV2-S (384x160) Public: 0.904 Private: 0.906\n\n- EfficientNetV2-S (384x256)  Public: 0.896. Private: 0.914\n\n\n    Diversity was controlled by varying the dataset used and the proportion and quality of mixed-in pp_clean_data.\nI ultimately ensembled 10 models in total,  Public: 0.915 Private: 0.921 . \nHowever, due to the instability of the leaderboard (LB), I wasn’t able to select the optimal combination of models.\n\n## 8. Other Details\n- 10s/15s for train, infer on 5s\n- AdamW + CosineAnnealingLR， LR： 2/3*e-4, WD: 1e-3/1e-4\n- Mixup with additive, frequency, and time masking\n- With and without noise augmentation\n- Many other common techniques like label smoothing, flip, etc.",
      "votes": null
    },
    {
      "id": "3218268",
      "postDate": "06/06/2025 03:33:17",
      "content": "<p>Congrats on the solo gold!</p>",
      "rawMarkdown": "Congrats on the solo gold!",
      "votes": null
    },
    {
      "id": "3218375",
      "postDate": "06/06/2025 05:49:26",
      "content": "<p>Thanks for sharing and congrats on solo gold !</p>",
      "rawMarkdown": "Thanks for sharing and congrats on solo gold !",
      "votes": null
    },
    {
      "id": "3218378",
      "postDate": "06/06/2025 05:50:50",
      "content": "<p>Congrats on the solo gold! </p>",
      "rawMarkdown": "Congrats on the solo gold!",
      "votes": null
    },
    {
      "id": "3218442",
      "postDate": "06/06/2025 07:33:09",
      "content": "<p>Congrats!<br>\nHow many hours do you spend  for training those 10 models in total? (and for individual?)<br>\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?<br>\nThank you.</p>",
      "rawMarkdown": "Congrats!\nHow many hours do you spend  for training those 10 models in total? (and for individual?)\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?\nThank you.",
      "votes": null
    },
    {
      "id": "3218454",
      "postDate": "06/06/2025 07:45:33",
      "content": "<p>1 model cost about 2hour  for me  with RTX 5090 </p>",
      "rawMarkdown": "1 model cost about 2hour  for me  with RTX 5090",
      "votes": null
    },
    {
      "id": "3218923",
      "postDate": "06/06/2025 23:28:16",
      "content": "<p>Congrats on the solo gold!</p>",
      "rawMarkdown": "Congrats on the solo gold!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218268,
      "author_name": "cooolz",
      "author_url": "",
      "post_date": "06/06/2025 03:33:17",
      "content": "<p>Congrats on the solo gold!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3218375,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "06/06/2025 05:49:26",
      "content": "<p>Thanks for sharing and congrats on solo gold !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3218378,
      "author_name": "salmanahmedtamu",
      "author_url": "",
      "post_date": "06/06/2025 05:50:50",
      "content": "<p>Congrats on the solo gold! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3218442,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "06/06/2025 07:33:09",
      "content": "<p>Congrats!<br>\nHow many hours do you spend  for training those 10 models in total? (and for individual?)<br>\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?<br>\nThank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218454,
          "author_name": "lihaoweicvch",
          "author_url": "",
          "post_date": "06/06/2025 07:45:33",
          "content": "<p>1 model cost about 2hour  for me  with RTX 5090 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218923,
      "author_name": "tyyuki",
      "author_url": "",
      "post_date": "06/06/2025 23:28:16",
      "content": "<p>Congrats on the solo gold!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218258": "Thanks to the organizers for hosting this competition, and also to the many Kaggler participants from previous years whose solutions have provided me with great inspiration.\nBirdCLEF is a challenging competition, Although the dataset shift isn't extremely large, the lack of a validation set from the same distribution makes training and evaluation quite difficult.\n\nI adopted the following approach during the competition:\n\n## 1. Establishing Baseline Model Performance\n\nI used the entire train_audio set and validated directly on the leaderboard. I found the following model architectures performed best:\n\n```\n(SED + CE loss) > (SED + FocalBCE) > (CNN + FocalBCE) > (CNN + CE)\n```\n\nThis was just an initial experiment, so it may have been influenced by other factors and not be generally applicable. But due to limited online submission opportunities, I stuck with SED + CE loss for the rest of the experiments.\nModels trained in this stage are referred to as stage1 models.\n\n## 2. Domain Adaptation & pseudo-labeled data\n\nAs seen in previous competitions, leveraging unlabeled data (train_soundscapes) can lead to significant gains.\nI used the following method:\n\n1. Use the stage1 models to predict on train_soundscapes, generating soft labels\n2. Apply high and low thresholds (thresh_low, thresh_high) to extract confident negative and positive samples\n3. These samples (assumed to be relatively clean) were added to train_audio to train new models — the stage2_models. The ambiguous samples in between were excluded\n4. Predict again on train_soundscapes using stage2_models to get updated soft labels\n\nThrough this iterative process, we obtained relatively clean pseudo-labeled data, referred to as pp_data_clean.\n\n## 3. Data Processing\n\nFollowing community discussions, I removed samples with obvious human speech — this cleaned dataset is called train_audio_clean.\nTo further improve diversity and make data clean, I used trained models to remove segments that clearly did not contain the target species, sampling at 5-second intervals, forming a new dataset: train_audio_clean_v2.\nBoth datasets were used for training.\n\n## 4. Mel Spectrogram Parameters\n\nI used mel spectrograms with the following resolutions:\n\n- 384x160\n- 384x256\n- 320x192\n- 320x160\n\nfmin = 0, fmax = 16000， n_fft=1536/2048\n\n## 5. Loss Function\n\nI adopted a modified hybrid  loss with CE.\nAlthough public notebooks showed that ConvNeXt performed well, using standard CE loss led to a lot of noisy predictions (many false positives).\nSo, I gradually incorporated the following loss to better stabilize training:\n\n```python\nthreshold = -5\nneg_mask = (y == 0.0)\nnegative_logits = logits[neg_mask]\npenalty = F.relu(negative_logits - threshold)\nsorted_penalty, _ = torch.sort(penalty)\ncutoff_index = int(len(sorted_penalty) * 0.95)\nselected_penalty = sorted_penalty[:cutoff_index]\nmean_penalty = selected_penalty.mean()\nloss = ce_loss + 0.1 * mean_penalty\n```\n\nIn some models, I also penalized positive samples with confidence scores that were too low.\n\n## 6. Post Processing\n- smooth with kernel: [0.02, 0.08, 0.8, 0.08, 0.02]\n- and average:\n```\ndef get_mean_scales(ref_freq):\n    alpha_max = 0.3\n    alpha_min = 0.1\n    # rare classed more recall\n    alpha = alpha_min + ref_freq * (alpha_max - alpha_min)\n    return alpha\n...\n   for c in len(n_classes):\n     a = alpha[c]\n     pred_prob[:, c] = pred_prob[:, c] * (1-a) + pred_prob[:, c].mean(keepdims=True) * a\n    \n \n```\n\n## 7. Some Model Results\n\n- ConvNeXt Tiny (320x192): Public: 0.901   Private: 0.914\n\n- 3x ConvNeXt Tiny Ensemble:  Public: 0.908 Private: 0.921\n\n- EfficientNetV2-S (384x160) Public: 0.904 Private: 0.906\n\n- EfficientNetV2-S (384x256)  Public: 0.896. Private: 0.914\n\n\n    Diversity was controlled by varying the dataset used and the proportion and quality of mixed-in pp_clean_data.\nI ultimately ensembled 10 models in total,  Public: 0.915 Private: 0.921 . \nHowever, due to the instability of the leaderboard (LB), I wasn’t able to select the optimal combination of models.\n\n## 8. Other Details\n- 10s/15s for train, infer on 5s\n- AdamW + CosineAnnealingLR， LR： 2/3*e-4, WD: 1e-3/1e-4\n- Mixup with additive, frequency, and time masking\n- With and without noise augmentation\n- Many other common techniques like label smoothing, flip, etc.",
    "3218268": "Congrats on the solo gold!",
    "3218375": "Thanks for sharing and congrats on solo gold !",
    "3218378": "Congrats on the solo gold!",
    "3218442": "Congrats!\nHow many hours do you spend  for training those 10 models in total? (and for individual?)\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?\nThank you.",
    "3218454": "1 model cost about 2hour  for me  with RTX 5090",
    "3218923": "Congrats on the solo gold!"
  },
  "source": "meta"
}