{
  "id": 583387,
  "title": "29th Place Solution for the BirdCLEF+ 2025 Competition",
  "url": "/competitions/birdclef-2025/writeups/toads-locust-29th-place-solution-for-the-birdclef-",
  "author_name": "",
  "post_date": "2025-06-06T13:19:44.119359700Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<h2>Our Top Solution</h2>\n<p>Thanks to the organizers for hosting this interesting competition, and thanks to past Kagglers whose work for previous BirdCLEF competitions gave us great inspiration. This was a challenging competition because the public test and private test distributions didn’t seem to correlate, so it was tough to know if good validation scores would transfer from our local validation to the private test. Also, thanks to my teammates <a href=\"https://www.kaggle.com/siavrez\" target=\"_blank\">@siavrez</a> and <a href=\"https://www.kaggle.com/kainsama\" target=\"_blank\">@kainsama</a> for their innovative work and effort for this competition!</p>\n<p>In any case, our final solution for this competition consists of these parts: </p>\n<h3>Backbone Models</h3>\n<ul>\n<li>EfficientNet-B3 (single-channel stem, ImageNet init)</li>\n<li>ResNeSt-50 (wider stem 64→128, ImageNet init)</li>\n<li>Nfnet-l0 (single-channel input, tuned layer norms)</li>\n</ul>\n<h3>Data Processing</h3>\n<ul>\n<li>Train on 256*313(2x,3x,4x,6x) segments</li>\n<li>Standardize with global mean/std from ~10k samples</li>\n<li>On-the-fly augmentations: time shift ±10 frames, freq shift ±5 bins; SpecAugment masks (time ≤ 15, freq ≤ 6); 50 % mixup</li>\n</ul>\n<h3>Loss Function</h3>\n<ul>\n<li>Multi-class focal loss (γ = 1.5, α = 0.25) to focus on hard/rare classes</li>\n</ul>\n<h3>Model Training</h3>\n<ul>\n<li>AdamW (lr = 2 × 10⁻⁴, weight_decay = 1 × 10⁻⁵), cosine decay with 3-epoch warmup</li>\n<li>Training on 10-15-20-30 second segments, and infer on 5-second segments</li>\n<li>Average top 5 checkpoints per model based on validation LRAP</li>\n</ul>\n<h3>Ensembling</h3>\n<ul>\n<li>In-model: average logits of top 5 checkpoints</li>\n<li>Cross-model blend (tuned on hold-out): EfficientNet 0.48, ResNeSt 0.32, ConvNeXt 0.20</li>\n</ul>\n<h3>Other Things</h3>\n<ul>\n<li>Trained on all data for the final solution</li>\n<li>Used ONNX to improve the runtime performance.</li>\n</ul>\n<p>Our top solution got a score of 0.902 on the public leaderboard and 0.906 on the private leaderboard.</p>\n<p>Link to the <a href=\"https://www.kaggle.com/code/siavrez/another-blend-combo-shift-rare?scriptVersionId=243860856\" target=\"_blank\">inference notebook</a></p>",
  "messages": [
    {
      "id": "3218634",
      "postDate": "06/06/2025 13:19:44",
      "content": "<h2>Our Top Solution</h2>\n<p>Thanks to the organizers for hosting this interesting competition, and thanks to past Kagglers whose work for previous BirdCLEF competitions gave us great inspiration. This was a challenging competition because the public test and private test distributions didn’t seem to correlate, so it was tough to know if good validation scores would transfer from our local validation to the private test. Also, thanks to my teammates <a href=\"https://www.kaggle.com/siavrez\" target=\"_blank\">@siavrez</a> and <a href=\"https://www.kaggle.com/kainsama\" target=\"_blank\">@kainsama</a> for their innovative work and effort for this competition!</p>\n<p>In any case, our final solution for this competition consists of these parts: </p>\n<h3>Backbone Models</h3>\n<ul>\n<li>EfficientNet-B3 (single-channel stem, ImageNet init)</li>\n<li>ResNeSt-50 (wider stem 64→128, ImageNet init)</li>\n<li>Nfnet-l0 (single-channel input, tuned layer norms)</li>\n</ul>\n<h3>Data Processing</h3>\n<ul>\n<li>Train on 256*313(2x,3x,4x,6x) segments</li>\n<li>Standardize with global mean/std from ~10k samples</li>\n<li>On-the-fly augmentations: time shift ±10 frames, freq shift ±5 bins; SpecAugment masks (time ≤ 15, freq ≤ 6); 50 % mixup</li>\n</ul>\n<h3>Loss Function</h3>\n<ul>\n<li>Multi-class focal loss (γ = 1.5, α = 0.25) to focus on hard/rare classes</li>\n</ul>\n<h3>Model Training</h3>\n<ul>\n<li>AdamW (lr = 2 × 10⁻⁴, weight_decay = 1 × 10⁻⁵), cosine decay with 3-epoch warmup</li>\n<li>Training on 10-15-20-30 second segments, and infer on 5-second segments</li>\n<li>Average top 5 checkpoints per model based on validation LRAP</li>\n</ul>\n<h3>Ensembling</h3>\n<ul>\n<li>In-model: average logits of top 5 checkpoints</li>\n<li>Cross-model blend (tuned on hold-out): EfficientNet 0.48, ResNeSt 0.32, ConvNeXt 0.20</li>\n</ul>\n<h3>Other Things</h3>\n<ul>\n<li>Trained on all data for the final solution</li>\n<li>Used ONNX to improve the runtime performance.</li>\n</ul>\n<p>Our top solution got a score of 0.902 on the public leaderboard and 0.906 on the private leaderboard.</p>\n<p>Link to the <a href=\"https://www.kaggle.com/code/siavrez/another-blend-combo-shift-rare?scriptVersionId=243860856\" target=\"_blank\">inference notebook</a></p>",
      "rawMarkdown": "## Our Top Solution\n\nThanks to the organizers for hosting this interesting competition, and thanks to past Kagglers whose work for previous BirdCLEF competitions gave us great inspiration. This was a challenging competition because the public test and private test distributions didn’t seem to correlate, so it was tough to know if good validation scores would transfer from our local validation to the private test. Also, thanks to my teammates @siavrez and @kainsama for their innovative work and effort for this competition!\n\nIn any case, our final solution for this competition consists of these parts: \n\n### Backbone Models\n\n* EfficientNet-B3 (single-channel stem, ImageNet init)\n* ResNeSt-50 (wider stem 64→128, ImageNet init)\n* Nfnet-l0 (single-channel input, tuned layer norms)\n\n### Data Processing\n\n* Train on 256*313(2x,3x,4x,6x) segments\n* Standardize with global mean/std from ~10k samples\n* On-the-fly augmentations: time shift ±10 frames, freq shift ±5 bins; SpecAugment masks (time ≤ 15, freq ≤ 6); 50 % mixup\n\n\n### Loss Function\n\n* Multi-class focal loss (γ = 1.5, α = 0.25) to focus on hard/rare classes\n\n### Model Training\n\n* AdamW (lr = 2 × 10⁻⁴, weight_decay = 1 × 10⁻⁵), cosine decay with 3-epoch warmup\n* Training on 10-15-20-30 second segments, and infer on 5-second segments\n* Average top 5 checkpoints per model based on validation LRAP\n\n### Ensembling\n\n* In-model: average logits of top 5 checkpoints\n* Cross-model blend (tuned on hold-out): EfficientNet 0.48, ResNeSt 0.32, ConvNeXt 0.20\n\n### Other Things\n\n* Trained on all data for the final solution\n* Used ONNX to improve the runtime performance.\n\nOur top solution got a score of 0.902 on the public leaderboard and 0.906 on the private leaderboard.\n\nLink to the [inference notebook](https://www.kaggle.com/code/siavrez/another-blend-combo-shift-rare?scriptVersionId=243860856)",
      "votes": null
    },
    {
      "id": "3218853",
      "postDate": "06/06/2025 19:53:01",
      "content": "<p>Congrats! Why did you pick 256x313? When you say (2x, 3x, 4x, 6x), do you mean that you concatenated 2, 3, 4 and six 5-second spectrograms?</p>",
      "rawMarkdown": "Congrats! Why did you pick 256x313? When you say (2x, 3x, 4x, 6x), do you mean that you concatenated 2, 3, 4 and six 5-second spectrograms?",
      "votes": null
    },
    {
      "id": "3219047",
      "postDate": "06/07/2025 05:43:15",
      "content": "<p>Thanks. We tried different values and found that 256 frequency bins worked best for our <br>\nmodels. The width comes from using 5-second clips for inference with a 32 kHz sample rate and a 512 hop <br>\nlength. We used longer clips (10s for 2x, 15s for 3x, and so on) for training and created a single, <br>\ncontinuous spectrogram from each longer clip (not concatenating the Mel spectrograms for adjacent 5-second clips).</p>",
      "rawMarkdown": "Thanks. We tried different values and found that 256 frequency bins worked best for our \nmodels. The width comes from using 5-second clips for inference with a 32 kHz sample rate and a 512 hop \nlength. We used longer clips (10s for 2x, 15s for 3x, and so on) for training and created a single, \ncontinuous spectrogram from each longer clip (not concatenating the Mel spectrograms for adjacent 5-second clips).",
      "votes": null
    },
    {
      "id": "3219240",
      "postDate": "06/07/2025 11:05:01",
      "content": "<p>Thanks! So just to be clear, your 10s clips gave 256x626 spectrograms, right? And even though you trained a model on 256x626, you used it for inference with 256x313? I didn’t know you could do that.</p>",
      "rawMarkdown": "Thanks! So just to be clear, your 10s clips gave 256x626 spectrograms, right? And even though you trained a model on 256x626, you used it for inference with 256x313? I didn’t know you could do that.",
      "votes": null
    },
    {
      "id": "3219274",
      "postDate": "06/07/2025 12:02:29",
      "content": "<p>Yeah, that's correct.</p>",
      "rawMarkdown": "Yeah, that's correct.",
      "votes": null
    },
    {
      "id": "3219525",
      "postDate": "06/07/2025 22:09:33",
      "content": "<p>Congrats!<br>\nCould you elaborate on what led you to believe the public and private test data distributions were uncorrelated?</p>",
      "rawMarkdown": "Congrats!\nCould you elaborate on what led you to believe the public and private test data distributions were uncorrelated?",
      "votes": null
    },
    {
      "id": "3219721",
      "postDate": "06/08/2025 08:42:58",
      "content": "<p>The improvement in local validation (using CV) didn't mean that you saw an improvement in the LB score. Ideally, the train (seen) data should be similar to the test (unseen) data, and a better local validation score should mean a better score on the LB. It wasn't the case for this competition. I'm not sure, but I guess data cleaning, like removing human voices from the train, could've helped with this.</p>",
      "rawMarkdown": "The improvement in local validation (using CV) didn't mean that you saw an improvement in the LB score. Ideally, the train (seen) data should be similar to the test (unseen) data, and a better local validation score should mean a better score on the LB. It wasn't the case for this competition. I'm not sure, but I guess data cleaning, like removing human voices from the train, could've helped with this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218853,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "06/06/2025 19:53:01",
      "content": "<p>Congrats! Why did you pick 256x313? When you say (2x, 3x, 4x, 6x), do you mean that you concatenated 2, 3, 4 and six 5-second spectrograms?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219047,
          "author_name": "habedi",
          "author_url": "",
          "post_date": "06/07/2025 05:43:15",
          "content": "<p>Thanks. We tried different values and found that 256 frequency bins worked best for our <br>\nmodels. The width comes from using 5-second clips for inference with a 32 kHz sample rate and a 512 hop <br>\nlength. We used longer clips (10s for 2x, 15s for 3x, and so on) for training and created a single, <br>\ncontinuous spectrogram from each longer clip (not concatenating the Mel spectrograms for adjacent 5-second clips).</p>",
          "votes": null,
          "replies": [
            {
              "id": 3219240,
              "author_name": "janhuus",
              "author_url": "",
              "post_date": "06/07/2025 11:05:01",
              "content": "<p>Thanks! So just to be clear, your 10s clips gave 256x626 spectrograms, right? And even though you trained a model on 256x626, you used it for inference with 256x313? I didn’t know you could do that.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3219274,
                  "author_name": "habedi",
                  "author_url": "",
                  "post_date": "06/07/2025 12:02:29",
                  "content": "<p>Yeah, that's correct.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3219525,
      "author_name": "tyyuki",
      "author_url": "",
      "post_date": "06/07/2025 22:09:33",
      "content": "<p>Congrats!<br>\nCould you elaborate on what led you to believe the public and private test data distributions were uncorrelated?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219721,
          "author_name": "habedi",
          "author_url": "",
          "post_date": "06/08/2025 08:42:58",
          "content": "<p>The improvement in local validation (using CV) didn't mean that you saw an improvement in the LB score. Ideally, the train (seen) data should be similar to the test (unseen) data, and a better local validation score should mean a better score on the LB. It wasn't the case for this competition. I'm not sure, but I guess data cleaning, like removing human voices from the train, could've helped with this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3218634": "## Our Top Solution\n\nThanks to the organizers for hosting this interesting competition, and thanks to past Kagglers whose work for previous BirdCLEF competitions gave us great inspiration. This was a challenging competition because the public test and private test distributions didn’t seem to correlate, so it was tough to know if good validation scores would transfer from our local validation to the private test. Also, thanks to my teammates @siavrez and @kainsama for their innovative work and effort for this competition!\n\nIn any case, our final solution for this competition consists of these parts: \n\n### Backbone Models\n\n* EfficientNet-B3 (single-channel stem, ImageNet init)\n* ResNeSt-50 (wider stem 64→128, ImageNet init)\n* Nfnet-l0 (single-channel input, tuned layer norms)\n\n### Data Processing\n\n* Train on 256*313(2x,3x,4x,6x) segments\n* Standardize with global mean/std from ~10k samples\n* On-the-fly augmentations: time shift ±10 frames, freq shift ±5 bins; SpecAugment masks (time ≤ 15, freq ≤ 6); 50 % mixup\n\n\n### Loss Function\n\n* Multi-class focal loss (γ = 1.5, α = 0.25) to focus on hard/rare classes\n\n### Model Training\n\n* AdamW (lr = 2 × 10⁻⁴, weight_decay = 1 × 10⁻⁵), cosine decay with 3-epoch warmup\n* Training on 10-15-20-30 second segments, and infer on 5-second segments\n* Average top 5 checkpoints per model based on validation LRAP\n\n### Ensembling\n\n* In-model: average logits of top 5 checkpoints\n* Cross-model blend (tuned on hold-out): EfficientNet 0.48, ResNeSt 0.32, ConvNeXt 0.20\n\n### Other Things\n\n* Trained on all data for the final solution\n* Used ONNX to improve the runtime performance.\n\nOur top solution got a score of 0.902 on the public leaderboard and 0.906 on the private leaderboard.\n\nLink to the [inference notebook](https://www.kaggle.com/code/siavrez/another-blend-combo-shift-rare?scriptVersionId=243860856)",
    "3218853": "Congrats! Why did you pick 256x313? When you say (2x, 3x, 4x, 6x), do you mean that you concatenated 2, 3, 4 and six 5-second spectrograms?",
    "3219047": "Thanks. We tried different values and found that 256 frequency bins worked best for our \nmodels. The width comes from using 5-second clips for inference with a 32 kHz sample rate and a 512 hop \nlength. We used longer clips (10s for 2x, 15s for 3x, and so on) for training and created a single, \ncontinuous spectrogram from each longer clip (not concatenating the Mel spectrograms for adjacent 5-second clips).",
    "3219240": "Thanks! So just to be clear, your 10s clips gave 256x626 spectrograms, right? And even though you trained a model on 256x626, you used it for inference with 256x313? I didn’t know you could do that.",
    "3219274": "Yeah, that's correct.",
    "3219525": "Congrats!\nCould you elaborate on what led you to believe the public and private test data distributions were uncorrelated?",
    "3219721": "The improvement in local validation (using CV) didn't mean that you saw an improvement in the LB score. Ideally, the train (seen) data should be similar to the test (unseen) data, and a better local validation score should mean a better score on the LB. It wasn't the case for this competition. I'm not sure, but I guess data cleaning, like removing human voices from the train, could've helped with this."
  },
  "source": "meta"
}