{
  "id": 583447,
  "title": "Place 38 | 0.902 AUC score ",
  "url": "/competitions/birdclef-2025/writeups/hatol-place-38-0-902-auc-score",
  "author_name": "",
  "post_date": "2025-06-06T23:24:15.940Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my submisstion code: <a href=\"https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\" target=\"_blank\">https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10915898%2Fb2a3b7d3d89c0338abc56704c3d26524%2FModelEvolution.png?generation=1749252251155041&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3><strong>Data and Challenge</strong></h3>\n<p>The goal of the BirdCLEF+ 2025 competition was to identify the species of various taxonomic groups (birds, amphibians, mammals, insects) present in soundscape recordings from the Middle Magdalena Valley of Colombia and the El Silencio Natural Reserve.</p>\n<p>What’s in the Test Set?<br>\nThe test data comprises long, real-world soundscape recordings captured in Colombia. Each soundscape typically contains calls from multiple species across different taxonomic groups—the exact challenge we want our model to solve.</p>\n<p>What Are We Given to Train On?<br>\nFor training, we’re provided with:</p>\n<p>Short audio clips—each featuring the call of a single species.</p>\n<p>These recordings come from xeno-canto.org, iNaturalist, and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.</p>\n<p>Each clip is carefully labeled with its species ID and taxonomy, but rarely contains more than one species per recording.</p>\n<p>The Real Challenge<br>\nThis sets up a fundamental mismatch:</p>\n<p>Training: Single-species, high-quality, labeled clips.</p>\n<p>Testing: Multi-species, real-world, often noisy soundscapes.</p>\n<p>The competition, therefore, isn’t just about building a good classifier. It’s about finding robust methods to bridge the gap between clean, isolated calls and complex, overlapping natural soundscapes—a leap that pushed every part of my pipeline, from preprocessing to model selection and pseudo-labeling.</p>\n<hr>\n<h3><strong>Step 1: Preprocessing</strong></h3>\n<ul>\n<li>Used <code>snakers4/silero-vad</code> for Voice Activity Detection (VAD) to remove human voices from the training data.</li>\n</ul>\n<hr>\n<h3><strong>Model 1: EfficientNet-B0 CNN with Strong Augmentation</strong></h3>\n<ul>\n<li><p><strong>Spectrogram settings:</strong><br>\n<code>N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000</code></p></li>\n<li><p><strong>Augmentations:</strong> Time masking, frequency masking, random brightness/contrast, high-frequency boost, dynamic range compression/expansion, frequency shift, subtle broadband noise, mid-frequency boost, slight blur (time direction), and mix original with blurred version.</p></li>\n<li><p><strong>Training:</strong></p>\n<ul>\n<li>EfficientNet-B0 backbone</li>\n<li>Mixup (α=0.15</li></ul></li>\n</ul>\n<p>Absolutely! Here’s your revised full LinkedIn post, with all your corrections reflected, and a clean structure. This version highlights your journey, technical pipeline, and the ensemble details exactly as you described:</p>\n<hr>\n<p><strong>Long story short: I finished 41st out of 2,162 teams in the BirdCLEF+ 2025 Kaggle competition.</strong></p>\n<p>Big thanks to <strong>Ariel</strong> for the support and advice along the way!</p>\n<p>Here’s an overview of what actually worked (there were plenty of failed experiments behind the scenes):</p>\n<hr>\n<h3><strong>Data and Challenge</strong></h3>\n<ul>\n<li><strong>train_audio/</strong>: Short recordings, each with a single bird, amphibian, mammal, or insect species.</li>\n<li><strong>train_soundscapes/</strong>: Unlabeled audio from the same recording locations as the test data, but usually containing multiple overlapping species.</li>\n<li><strong>Key difference:</strong> Test data contains multiple species per recording, while most training data has only one.</li>\n</ul>\n<hr>\n<h3><strong>Step 1: Preprocessing</strong></h3>\n<ul>\n<li>Used <code>snakers4/silero-vad</code> for Voice Activity Detection (VAD) to remove human voices from the training data.</li>\n</ul>\n<hr>\n<h3><strong>Model 1: EfficientNet-B0 CNN with Heavy Augmentation</strong></h3>\n<ul>\n<li><p><strong>Spectrogram settings:</strong><br>\n<code>N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000</code></p></li>\n<li><p><strong>Augmentations:</strong></p>\n<ul>\n<li>Time masking, frequency masking</li>\n<li>Random brightness/contrast</li>\n<li>High-frequency boost</li>\n<li>Dynamic range compression/expansion</li>\n<li>Frequency shift</li>\n<li>Subtle background noise (broadband and mid frequencies)</li>\n<li>Slight blur in the time direction, mixing original and blurred</li></ul></li>\n<li><p><strong>Training setup:</strong></p>\n<ul>\n<li>EfficientNet-B0 backbone</li>\n<li>Mixup (α = 0.15)</li>\n<li>BCE loss (primary label = 1, secondary = 0.75), multilabel/multiclass</li>\n<li>Trained on the middle 5 seconds of each clip (tried energy-based and random segments but performed worse)</li></ul></li>\n<li><p><strong>Results:</strong></p>\n<ul>\n<li>0.817 AUC (Area Under the Curve; how well the model ranks true labels higher than false ones)</li></ul></li>\n<li><p><strong>Pseudo-labeling:</strong></p>\n<ul>\n<li>Inferred on all <code>train_soundscapes</code> (split into 5s chunks, average predictions, use as new labels)</li>\n<li>Boosted performance to 0.835 AUC</li>\n<li>Switching backbone to EfficientNetV2-S reached 0.843, but best pseudo-labeling results were still with EfficientNet-B0.</li></ul></li>\n</ul>\n<hr>\n<h3><strong>Model 2: Enhanced Pooling and Mixup</strong></h3>\n<ul>\n<li><p><strong>Architecture:</strong></p>\n<ul>\n<li>EfficientNet-B0 backbone</li>\n<li>Added GeM pooling layers on features from layers 3 and 4</li>\n<li>Mixup α = 0.5</li></ul></li>\n<li><p><strong>Alternative Experiments:</strong></p>\n<ul>\n<li>Tried adaptive pooling, denoiser augmentation, and more advanced approaches (NatureLM-audio segment extraction, 3-channel Mels, cutmix, teacher-student), but no additional improvements.</li></ul></li>\n<li><p><strong>Best Model 2 Result:</strong> 0.855 AUC</p></li>\n</ul>\n<hr>\n<h3><strong>Ensembling</strong></h3>\n<p>Combined my three best CNN models:</p>\n<ol>\n<li><strong>GeM pooling (no denoiser)</strong></li>\n<li><strong>GeM pooling with denoiser</strong></li>\n<li><strong>Adaptive pooling (no denoiser)</strong></li>\n</ol>\n<ul>\n<li><strong>Ensemble AUC:</strong> 0.868</li>\n</ul>\n<p>After this, most further tweaks (bigger backbones, more mixup strategies, new pseudo-labeling techniques, heavier augmentation) failed to improve validation. I hit a plateau and considered giving up.</p>\n<hr>\n<h3><strong>Exploring SED (Sound Event Detection): My First Steps</strong></h3>\n<p>In the last three weeks of this two-month competition, SED-based approaches started popping up on the Kaggle forums. This was my first time trying SED models. With the steep learning curve and limited time, I couldn’t get my own SED models above 0.841 AUC.</p>\n<p>To keep moving forward, I leveraged top SED model predictions shared by others in the forums:</p>\n<ul>\n<li>Used single SED model results (NFNet backbone, 0.86 AUC) and a triple SED ensemble (0.85 AUC, with Power Adjustment for Low-Rank) from public Kaggle kernels and posts.</li>\n<li>Combined these with my three CNN models using a Quantile-Mix ensemble (α = 0.5), which pushed the score up to 0.893 AUC.</li>\n</ul>\n<hr>\n<h3><strong>Final Push: Pretraining on BirdCLEF 2021–2024</strong></h3>\n<ul>\n<li>Pretrained all three CNN models on BirdCLEF data from 2021–2024, then fine-tuned on 2025 data.</li>\n<li><strong>Result:</strong> Single-model AUC improved from 0.855 to 0.868.</li>\n<li>Final ensemble (3 CNNs + SEDs) scored <strong>0.894 AUC</strong> on the public leaderboard (45th place).</li>\n<li>On the final private leaderboard: <strong>0.902 AUC</strong>, landing at <strong>41st place</strong>.</li>\n<li>For reference, 1st place finished at 0.930 AUC—just a 3% gap!</li>\n</ul>\n<hr>\n<p><strong>Key takeaways:</strong></p>\n<ul>\n<li>Strong preprocessing (VAD, MelSpectrogram, heavy augmentation) is crucial.</li>\n<li>Pseudo-labeling unlabeled soundscapes provides a significant boost if segment selection is careful.</li>\n<li>Pooling strategies and label weight tuning can give CNNs a few more points.</li>\n<li>SED is a powerful approach for multi-species data, but there’s a learning curve if you’re new to it (like me!).</li>\n<li>Pretraining on historical BirdCLEF data really helps.</li>\n<li>The Kaggle community is an incredible resource—leveraging public notebooks and shared solutions is invaluable, especially when time is short.</li>\n</ul>\n<hr>\n<p>I learned a ton during this competition—about audio modeling, new architectures, and pushing through plateaus. Huge thanks again to Ariel, everyone in the forums, and all those who shared code and kernels. I’m already looking forward to next year and hoping to break into the top 10!</p>\n<hr>\n<p>If you want to chat about the technical details, see code, or ask anything about the pipeline, feel free to reach out!</p>",
  "messages": [
    {
      "id": "3218909",
      "postDate": "06/06/2025 22:45:36",
      "content": "<p>This is my submisstion code: <a href=\"https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\" target=\"_blank\">https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10915898%2Fb2a3b7d3d89c0338abc56704c3d26524%2FModelEvolution.png?generation=1749252251155041&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h3><strong>Data and Challenge</strong></h3>\n<p>The goal of the BirdCLEF+ 2025 competition was to identify the species of various taxonomic groups (birds, amphibians, mammals, insects) present in soundscape recordings from the Middle Magdalena Valley of Colombia and the El Silencio Natural Reserve.</p>\n<p>What’s in the Test Set?<br>\nThe test data comprises long, real-world soundscape recordings captured in Colombia. Each soundscape typically contains calls from multiple species across different taxonomic groups—the exact challenge we want our model to solve.</p>\n<p>What Are We Given to Train On?<br>\nFor training, we’re provided with:</p>\n<p>Short audio clips—each featuring the call of a single species.</p>\n<p>These recordings come from xeno-canto.org, iNaturalist, and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.</p>\n<p>Each clip is carefully labeled with its species ID and taxonomy, but rarely contains more than one species per recording.</p>\n<p>The Real Challenge<br>\nThis sets up a fundamental mismatch:</p>\n<p>Training: Single-species, high-quality, labeled clips.</p>\n<p>Testing: Multi-species, real-world, often noisy soundscapes.</p>\n<p>The competition, therefore, isn’t just about building a good classifier. It’s about finding robust methods to bridge the gap between clean, isolated calls and complex, overlapping natural soundscapes—a leap that pushed every part of my pipeline, from preprocessing to model selection and pseudo-labeling.</p>\n<hr>\n<h3><strong>Step 1: Preprocessing</strong></h3>\n<ul>\n<li>Used <code>snakers4/silero-vad</code> for Voice Activity Detection (VAD) to remove human voices from the training data.</li>\n</ul>\n<hr>\n<h3><strong>Model 1: EfficientNet-B0 CNN with Strong Augmentation</strong></h3>\n<ul>\n<li><p><strong>Spectrogram settings:</strong><br>\n<code>N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000</code></p></li>\n<li><p><strong>Augmentations:</strong> Time masking, frequency masking, random brightness/contrast, high-frequency boost, dynamic range compression/expansion, frequency shift, subtle broadband noise, mid-frequency boost, slight blur (time direction), and mix original with blurred version.</p></li>\n<li><p><strong>Training:</strong></p>\n<ul>\n<li>EfficientNet-B0 backbone</li>\n<li>Mixup (α=0.15</li></ul></li>\n</ul>\n<p>Absolutely! Here’s your revised full LinkedIn post, with all your corrections reflected, and a clean structure. This version highlights your journey, technical pipeline, and the ensemble details exactly as you described:</p>\n<hr>\n<p><strong>Long story short: I finished 41st out of 2,162 teams in the BirdCLEF+ 2025 Kaggle competition.</strong></p>\n<p>Big thanks to <strong>Ariel</strong> for the support and advice along the way!</p>\n<p>Here’s an overview of what actually worked (there were plenty of failed experiments behind the scenes):</p>\n<hr>\n<h3><strong>Data and Challenge</strong></h3>\n<ul>\n<li><strong>train_audio/</strong>: Short recordings, each with a single bird, amphibian, mammal, or insect species.</li>\n<li><strong>train_soundscapes/</strong>: Unlabeled audio from the same recording locations as the test data, but usually containing multiple overlapping species.</li>\n<li><strong>Key difference:</strong> Test data contains multiple species per recording, while most training data has only one.</li>\n</ul>\n<hr>\n<h3><strong>Step 1: Preprocessing</strong></h3>\n<ul>\n<li>Used <code>snakers4/silero-vad</code> for Voice Activity Detection (VAD) to remove human voices from the training data.</li>\n</ul>\n<hr>\n<h3><strong>Model 1: EfficientNet-B0 CNN with Heavy Augmentation</strong></h3>\n<ul>\n<li><p><strong>Spectrogram settings:</strong><br>\n<code>N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000</code></p></li>\n<li><p><strong>Augmentations:</strong></p>\n<ul>\n<li>Time masking, frequency masking</li>\n<li>Random brightness/contrast</li>\n<li>High-frequency boost</li>\n<li>Dynamic range compression/expansion</li>\n<li>Frequency shift</li>\n<li>Subtle background noise (broadband and mid frequencies)</li>\n<li>Slight blur in the time direction, mixing original and blurred</li></ul></li>\n<li><p><strong>Training setup:</strong></p>\n<ul>\n<li>EfficientNet-B0 backbone</li>\n<li>Mixup (α = 0.15)</li>\n<li>BCE loss (primary label = 1, secondary = 0.75), multilabel/multiclass</li>\n<li>Trained on the middle 5 seconds of each clip (tried energy-based and random segments but performed worse)</li></ul></li>\n<li><p><strong>Results:</strong></p>\n<ul>\n<li>0.817 AUC (Area Under the Curve; how well the model ranks true labels higher than false ones)</li></ul></li>\n<li><p><strong>Pseudo-labeling:</strong></p>\n<ul>\n<li>Inferred on all <code>train_soundscapes</code> (split into 5s chunks, average predictions, use as new labels)</li>\n<li>Boosted performance to 0.835 AUC</li>\n<li>Switching backbone to EfficientNetV2-S reached 0.843, but best pseudo-labeling results were still with EfficientNet-B0.</li></ul></li>\n</ul>\n<hr>\n<h3><strong>Model 2: Enhanced Pooling and Mixup</strong></h3>\n<ul>\n<li><p><strong>Architecture:</strong></p>\n<ul>\n<li>EfficientNet-B0 backbone</li>\n<li>Added GeM pooling layers on features from layers 3 and 4</li>\n<li>Mixup α = 0.5</li></ul></li>\n<li><p><strong>Alternative Experiments:</strong></p>\n<ul>\n<li>Tried adaptive pooling, denoiser augmentation, and more advanced approaches (NatureLM-audio segment extraction, 3-channel Mels, cutmix, teacher-student), but no additional improvements.</li></ul></li>\n<li><p><strong>Best Model 2 Result:</strong> 0.855 AUC</p></li>\n</ul>\n<hr>\n<h3><strong>Ensembling</strong></h3>\n<p>Combined my three best CNN models:</p>\n<ol>\n<li><strong>GeM pooling (no denoiser)</strong></li>\n<li><strong>GeM pooling with denoiser</strong></li>\n<li><strong>Adaptive pooling (no denoiser)</strong></li>\n</ol>\n<ul>\n<li><strong>Ensemble AUC:</strong> 0.868</li>\n</ul>\n<p>After this, most further tweaks (bigger backbones, more mixup strategies, new pseudo-labeling techniques, heavier augmentation) failed to improve validation. I hit a plateau and considered giving up.</p>\n<hr>\n<h3><strong>Exploring SED (Sound Event Detection): My First Steps</strong></h3>\n<p>In the last three weeks of this two-month competition, SED-based approaches started popping up on the Kaggle forums. This was my first time trying SED models. With the steep learning curve and limited time, I couldn’t get my own SED models above 0.841 AUC.</p>\n<p>To keep moving forward, I leveraged top SED model predictions shared by others in the forums:</p>\n<ul>\n<li>Used single SED model results (NFNet backbone, 0.86 AUC) and a triple SED ensemble (0.85 AUC, with Power Adjustment for Low-Rank) from public Kaggle kernels and posts.</li>\n<li>Combined these with my three CNN models using a Quantile-Mix ensemble (α = 0.5), which pushed the score up to 0.893 AUC.</li>\n</ul>\n<hr>\n<h3><strong>Final Push: Pretraining on BirdCLEF 2021–2024</strong></h3>\n<ul>\n<li>Pretrained all three CNN models on BirdCLEF data from 2021–2024, then fine-tuned on 2025 data.</li>\n<li><strong>Result:</strong> Single-model AUC improved from 0.855 to 0.868.</li>\n<li>Final ensemble (3 CNNs + SEDs) scored <strong>0.894 AUC</strong> on the public leaderboard (45th place).</li>\n<li>On the final private leaderboard: <strong>0.902 AUC</strong>, landing at <strong>41st place</strong>.</li>\n<li>For reference, 1st place finished at 0.930 AUC—just a 3% gap!</li>\n</ul>\n<hr>\n<p><strong>Key takeaways:</strong></p>\n<ul>\n<li>Strong preprocessing (VAD, MelSpectrogram, heavy augmentation) is crucial.</li>\n<li>Pseudo-labeling unlabeled soundscapes provides a significant boost if segment selection is careful.</li>\n<li>Pooling strategies and label weight tuning can give CNNs a few more points.</li>\n<li>SED is a powerful approach for multi-species data, but there’s a learning curve if you’re new to it (like me!).</li>\n<li>Pretraining on historical BirdCLEF data really helps.</li>\n<li>The Kaggle community is an incredible resource—leveraging public notebooks and shared solutions is invaluable, especially when time is short.</li>\n</ul>\n<hr>\n<p>I learned a ton during this competition—about audio modeling, new architectures, and pushing through plateaus. Huge thanks again to Ariel, everyone in the forums, and all those who shared code and kernels. I’m already looking forward to next year and hoping to break into the top 10!</p>\n<hr>\n<p>If you want to chat about the technical details, see code, or ask anything about the pipeline, feel free to reach out!</p>",
      "rawMarkdown": "This is my submisstion code: https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10915898%2Fb2a3b7d3d89c0338abc56704c3d26524%2FModelEvolution.png?generation=1749252251155041&alt=media)\n\n---\n\n### **Data and Challenge**\n\nThe goal of the BirdCLEF+ 2025 competition was to identify the species of various taxonomic groups (birds, amphibians, mammals, insects) present in soundscape recordings from the Middle Magdalena Valley of Colombia and the El Silencio Natural Reserve.\n\nWhat’s in the Test Set?\nThe test data comprises long, real-world soundscape recordings captured in Colombia. Each soundscape typically contains calls from multiple species across different taxonomic groups—the exact challenge we want our model to solve.\n\nWhat Are We Given to Train On?\nFor training, we’re provided with:\n\nShort audio clips—each featuring the call of a single species.\n\nThese recordings come from xeno-canto.org, iNaturalist, and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.\n\nEach clip is carefully labeled with its species ID and taxonomy, but rarely contains more than one species per recording.\n\nThe Real Challenge\nThis sets up a fundamental mismatch:\n\nTraining: Single-species, high-quality, labeled clips.\n\nTesting: Multi-species, real-world, often noisy soundscapes.\n\nThe competition, therefore, isn’t just about building a good classifier. It’s about finding robust methods to bridge the gap between clean, isolated calls and complex, overlapping natural soundscapes—a leap that pushed every part of my pipeline, from preprocessing to model selection and pseudo-labeling.\n\n---\n\n### **Step 1: Preprocessing**\n\n* Used `snakers4/silero-vad` for Voice Activity Detection (VAD) to remove human voices from the training data.\n\n---\n\n### **Model 1: EfficientNet-B0 CNN with Strong Augmentation**\n\n* **Spectrogram settings:**\n  `N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000`\n* **Augmentations:** Time masking, frequency masking, random brightness/contrast, high-frequency boost, dynamic range compression/expansion, frequency shift, subtle broadband noise, mid-frequency boost, slight blur (time direction), and mix original with blurred version.\n* **Training:**\n\n  * EfficientNet-B0 backbone\n  * Mixup (α=0.15\n\n\nAbsolutely! Here’s your revised full LinkedIn post, with all your corrections reflected, and a clean structure. This version highlights your journey, technical pipeline, and the ensemble details exactly as you described:\n\n---\n\n**Long story short: I finished 41st out of 2,162 teams in the BirdCLEF+ 2025 Kaggle competition.**\n\nBig thanks to **Ariel** for the support and advice along the way!\n\nHere’s an overview of what actually worked (there were plenty of failed experiments behind the scenes):\n\n---\n\n### **Data and Challenge**\n\n* **train\\_audio/**: Short recordings, each with a single bird, amphibian, mammal, or insect species.\n* **train\\_soundscapes/**: Unlabeled audio from the same recording locations as the test data, but usually containing multiple overlapping species.\n* **Key difference:** Test data contains multiple species per recording, while most training data has only one.\n\n---\n\n### **Step 1: Preprocessing**\n\n* Used `snakers4/silero-vad` for Voice Activity Detection (VAD) to remove human voices from the training data.\n\n---\n\n### **Model 1: EfficientNet-B0 CNN with Heavy Augmentation**\n\n* **Spectrogram settings:**\n  `N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000`\n* **Augmentations:**\n\n  * Time masking, frequency masking\n  * Random brightness/contrast\n  * High-frequency boost\n  * Dynamic range compression/expansion\n  * Frequency shift\n  * Subtle background noise (broadband and mid frequencies)\n  * Slight blur in the time direction, mixing original and blurred\n* **Training setup:**\n\n  * EfficientNet-B0 backbone\n  * Mixup (α = 0.15)\n  * BCE loss (primary label = 1, secondary = 0.75), multilabel/multiclass\n  * Trained on the middle 5 seconds of each clip (tried energy-based and random segments but performed worse)\n* **Results:**\n\n  * 0.817 AUC (Area Under the Curve; how well the model ranks true labels higher than false ones)\n* **Pseudo-labeling:**\n\n  * Inferred on all `train_soundscapes` (split into 5s chunks, average predictions, use as new labels)\n  * Boosted performance to 0.835 AUC\n  * Switching backbone to EfficientNetV2-S reached 0.843, but best pseudo-labeling results were still with EfficientNet-B0.\n\n---\n\n### **Model 2: Enhanced Pooling and Mixup**\n\n* **Architecture:**\n\n  * EfficientNet-B0 backbone\n  * Added GeM pooling layers on features from layers 3 and 4\n  * Mixup α = 0.5\n* **Alternative Experiments:**\n\n  * Tried adaptive pooling, denoiser augmentation, and more advanced approaches (NatureLM-audio segment extraction, 3-channel Mels, cutmix, teacher-student), but no additional improvements.\n* **Best Model 2 Result:** 0.855 AUC\n\n---\n\n### **Ensembling**\n\nCombined my three best CNN models:\n\n1. **GeM pooling (no denoiser)**\n2. **GeM pooling with denoiser**\n3. **Adaptive pooling (no denoiser)**\n\n* **Ensemble AUC:** 0.868\n\nAfter this, most further tweaks (bigger backbones, more mixup strategies, new pseudo-labeling techniques, heavier augmentation) failed to improve validation. I hit a plateau and considered giving up.\n\n---\n\n### **Exploring SED (Sound Event Detection): My First Steps**\n\nIn the last three weeks of this two-month competition, SED-based approaches started popping up on the Kaggle forums. This was my first time trying SED models. With the steep learning curve and limited time, I couldn’t get my own SED models above 0.841 AUC.\n\nTo keep moving forward, I leveraged top SED model predictions shared by others in the forums:\n\n* Used single SED model results (NFNet backbone, 0.86 AUC) and a triple SED ensemble (0.85 AUC, with Power Adjustment for Low-Rank) from public Kaggle kernels and posts.\n* Combined these with my three CNN models using a Quantile-Mix ensemble (α = 0.5), which pushed the score up to 0.893 AUC.\n\n---\n\n### **Final Push: Pretraining on BirdCLEF 2021–2024**\n\n* Pretrained all three CNN models on BirdCLEF data from 2021–2024, then fine-tuned on 2025 data.\n* **Result:** Single-model AUC improved from 0.855 to 0.868.\n* Final ensemble (3 CNNs + SEDs) scored **0.894 AUC** on the public leaderboard (45th place).\n* On the final private leaderboard: **0.902 AUC**, landing at **41st place**.\n* For reference, 1st place finished at 0.930 AUC—just a 3% gap!\n\n---\n\n**Key takeaways:**\n\n* Strong preprocessing (VAD, MelSpectrogram, heavy augmentation) is crucial.\n* Pseudo-labeling unlabeled soundscapes provides a significant boost if segment selection is careful.\n* Pooling strategies and label weight tuning can give CNNs a few more points.\n* SED is a powerful approach for multi-species data, but there’s a learning curve if you’re new to it (like me!).\n* Pretraining on historical BirdCLEF data really helps.\n* The Kaggle community is an incredible resource—leveraging public notebooks and shared solutions is invaluable, especially when time is short.\n\n---\n\nI learned a ton during this competition—about audio modeling, new architectures, and pushing through plateaus. Huge thanks again to Ariel, everyone in the forums, and all those who shared code and kernels. I’m already looking forward to next year and hoping to break into the top 10!\n\n---\n\nIf you want to chat about the technical details, see code, or ask anything about the pipeline, feel free to reach out!",
      "votes": null
    },
    {
      "id": "3218929",
      "postDate": "06/06/2025 23:39:51",
      "content": "<p>Great job on reaching 38th with a 0.902 AUC—your use of GeM pooling and Quantile-Mix ensembling is really impressive.<br>\nCould you share more details on how you balanced the CNN and SED model outputs in the final ensemble?</p>",
      "rawMarkdown": "Great job on reaching 38th with a 0.902 AUC—your use of GeM pooling and Quantile-Mix ensembling is really impressive.\nCould you share more details on how you balanced the CNN and SED model outputs in the final ensemble?",
      "votes": null
    },
    {
      "id": "3219220",
      "postDate": "06/07/2025 10:22:39",
      "content": "<p>Thanks a lot!<br>\nIn case you missed it, here's my final submission notebook:<br>\n👉 <a href=\"https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\" target=\"_blank\">https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc</a></p>\n<p>I based my final predictions on two SED models that were publicity available, combined with a blend of three slightly different CNN pipelines:<br>\nModel 1 – Standard preprocessing with a GeM pooling layer.<br>\nModel 2 – Same as Model 1, but added a denoising step in preprocessing.<br>\nModel 3 – Same as Model 1, but swapped GeM with AdaptivePooling.<br>\ncombined it for single submission using equal weight for model1,2,3.</p>\n<p>Each way generated its own submission CSV. I then combined the predictions using Quantile-Mix (with alpha=0.5) to balance between rank-based and raw-score averaging.</p>\n<p>That final blend gave me a 0.902 AUC score. Super happy with the outcome!</p>",
      "rawMarkdown": "Thanks a lot!\nIn case you missed it, here's my final submission notebook:\n👉 https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\n\nI based my final predictions on two SED models that were publicity available, combined with a blend of three slightly different CNN pipelines:\nModel 1 – Standard preprocessing with a GeM pooling layer.\nModel 2 – Same as Model 1, but added a denoising step in preprocessing.\nModel 3 – Same as Model 1, but swapped GeM with AdaptivePooling.\ncombined it for single submission using equal weight for model1,2,3.\n\nEach way generated its own submission CSV. I then combined the predictions using Quantile-Mix (with alpha=0.5) to balance between rank-based and raw-score averaging.\n\nThat final blend gave me a 0.902 AUC score. Super happy with the outcome!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218929,
      "author_name": "tyyuki",
      "author_url": "",
      "post_date": "06/06/2025 23:39:51",
      "content": "<p>Great job on reaching 38th with a 0.902 AUC—your use of GeM pooling and Quantile-Mix ensembling is really impressive.<br>\nCould you share more details on how you balanced the CNN and SED model outputs in the final ensemble?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219220,
          "author_name": "maxmelichov",
          "author_url": "",
          "post_date": "06/07/2025 10:22:39",
          "content": "<p>Thanks a lot!<br>\nIn case you missed it, here's my final submission notebook:<br>\n👉 <a href=\"https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\" target=\"_blank\">https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc</a></p>\n<p>I based my final predictions on two SED models that were publicity available, combined with a blend of three slightly different CNN pipelines:<br>\nModel 1 – Standard preprocessing with a GeM pooling layer.<br>\nModel 2 – Same as Model 1, but added a denoising step in preprocessing.<br>\nModel 3 – Same as Model 1, but swapped GeM with AdaptivePooling.<br>\ncombined it for single submission using equal weight for model1,2,3.</p>\n<p>Each way generated its own submission CSV. I then combined the predictions using Quantile-Mix (with alpha=0.5) to balance between rank-based and raw-score averaging.</p>\n<p>That final blend gave me a 0.902 AUC score. Super happy with the outcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3218909": "This is my submisstion code: https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10915898%2Fb2a3b7d3d89c0338abc56704c3d26524%2FModelEvolution.png?generation=1749252251155041&alt=media)\n\n---\n\n### **Data and Challenge**\n\nThe goal of the BirdCLEF+ 2025 competition was to identify the species of various taxonomic groups (birds, amphibians, mammals, insects) present in soundscape recordings from the Middle Magdalena Valley of Colombia and the El Silencio Natural Reserve.\n\nWhat’s in the Test Set?\nThe test data comprises long, real-world soundscape recordings captured in Colombia. Each soundscape typically contains calls from multiple species across different taxonomic groups—the exact challenge we want our model to solve.\n\nWhat Are We Given to Train On?\nFor training, we’re provided with:\n\nShort audio clips—each featuring the call of a single species.\n\nThese recordings come from xeno-canto.org, iNaturalist, and the Colombian Sound Archive (CSA) of the Humboldt Institute for Biological Resources Research in Colombia.\n\nEach clip is carefully labeled with its species ID and taxonomy, but rarely contains more than one species per recording.\n\nThe Real Challenge\nThis sets up a fundamental mismatch:\n\nTraining: Single-species, high-quality, labeled clips.\n\nTesting: Multi-species, real-world, often noisy soundscapes.\n\nThe competition, therefore, isn’t just about building a good classifier. It’s about finding robust methods to bridge the gap between clean, isolated calls and complex, overlapping natural soundscapes—a leap that pushed every part of my pipeline, from preprocessing to model selection and pseudo-labeling.\n\n---\n\n### **Step 1: Preprocessing**\n\n* Used `snakers4/silero-vad` for Voice Activity Detection (VAD) to remove human voices from the training data.\n\n---\n\n### **Model 1: EfficientNet-B0 CNN with Strong Augmentation**\n\n* **Spectrogram settings:**\n  `N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000`\n* **Augmentations:** Time masking, frequency masking, random brightness/contrast, high-frequency boost, dynamic range compression/expansion, frequency shift, subtle broadband noise, mid-frequency boost, slight blur (time direction), and mix original with blurred version.\n* **Training:**\n\n  * EfficientNet-B0 backbone\n  * Mixup (α=0.15\n\n\nAbsolutely! Here’s your revised full LinkedIn post, with all your corrections reflected, and a clean structure. This version highlights your journey, technical pipeline, and the ensemble details exactly as you described:\n\n---\n\n**Long story short: I finished 41st out of 2,162 teams in the BirdCLEF+ 2025 Kaggle competition.**\n\nBig thanks to **Ariel** for the support and advice along the way!\n\nHere’s an overview of what actually worked (there were plenty of failed experiments behind the scenes):\n\n---\n\n### **Data and Challenge**\n\n* **train\\_audio/**: Short recordings, each with a single bird, amphibian, mammal, or insect species.\n* **train\\_soundscapes/**: Unlabeled audio from the same recording locations as the test data, but usually containing multiple overlapping species.\n* **Key difference:** Test data contains multiple species per recording, while most training data has only one.\n\n---\n\n### **Step 1: Preprocessing**\n\n* Used `snakers4/silero-vad` for Voice Activity Detection (VAD) to remove human voices from the training data.\n\n---\n\n### **Model 1: EfficientNet-B0 CNN with Heavy Augmentation**\n\n* **Spectrogram settings:**\n  `N_FFT=1024, HOP_LENGTH=64, N_MELS=148, FMIN=20, FMAX=16000`\n* **Augmentations:**\n\n  * Time masking, frequency masking\n  * Random brightness/contrast\n  * High-frequency boost\n  * Dynamic range compression/expansion\n  * Frequency shift\n  * Subtle background noise (broadband and mid frequencies)\n  * Slight blur in the time direction, mixing original and blurred\n* **Training setup:**\n\n  * EfficientNet-B0 backbone\n  * Mixup (α = 0.15)\n  * BCE loss (primary label = 1, secondary = 0.75), multilabel/multiclass\n  * Trained on the middle 5 seconds of each clip (tried energy-based and random segments but performed worse)\n* **Results:**\n\n  * 0.817 AUC (Area Under the Curve; how well the model ranks true labels higher than false ones)\n* **Pseudo-labeling:**\n\n  * Inferred on all `train_soundscapes` (split into 5s chunks, average predictions, use as new labels)\n  * Boosted performance to 0.835 AUC\n  * Switching backbone to EfficientNetV2-S reached 0.843, but best pseudo-labeling results were still with EfficientNet-B0.\n\n---\n\n### **Model 2: Enhanced Pooling and Mixup**\n\n* **Architecture:**\n\n  * EfficientNet-B0 backbone\n  * Added GeM pooling layers on features from layers 3 and 4\n  * Mixup α = 0.5\n* **Alternative Experiments:**\n\n  * Tried adaptive pooling, denoiser augmentation, and more advanced approaches (NatureLM-audio segment extraction, 3-channel Mels, cutmix, teacher-student), but no additional improvements.\n* **Best Model 2 Result:** 0.855 AUC\n\n---\n\n### **Ensembling**\n\nCombined my three best CNN models:\n\n1. **GeM pooling (no denoiser)**\n2. **GeM pooling with denoiser**\n3. **Adaptive pooling (no denoiser)**\n\n* **Ensemble AUC:** 0.868\n\nAfter this, most further tweaks (bigger backbones, more mixup strategies, new pseudo-labeling techniques, heavier augmentation) failed to improve validation. I hit a plateau and considered giving up.\n\n---\n\n### **Exploring SED (Sound Event Detection): My First Steps**\n\nIn the last three weeks of this two-month competition, SED-based approaches started popping up on the Kaggle forums. This was my first time trying SED models. With the steep learning curve and limited time, I couldn’t get my own SED models above 0.841 AUC.\n\nTo keep moving forward, I leveraged top SED model predictions shared by others in the forums:\n\n* Used single SED model results (NFNet backbone, 0.86 AUC) and a triple SED ensemble (0.85 AUC, with Power Adjustment for Low-Rank) from public Kaggle kernels and posts.\n* Combined these with my three CNN models using a Quantile-Mix ensemble (α = 0.5), which pushed the score up to 0.893 AUC.\n\n---\n\n### **Final Push: Pretraining on BirdCLEF 2021–2024**\n\n* Pretrained all three CNN models on BirdCLEF data from 2021–2024, then fine-tuned on 2025 data.\n* **Result:** Single-model AUC improved from 0.855 to 0.868.\n* Final ensemble (3 CNNs + SEDs) scored **0.894 AUC** on the public leaderboard (45th place).\n* On the final private leaderboard: **0.902 AUC**, landing at **41st place**.\n* For reference, 1st place finished at 0.930 AUC—just a 3% gap!\n\n---\n\n**Key takeaways:**\n\n* Strong preprocessing (VAD, MelSpectrogram, heavy augmentation) is crucial.\n* Pseudo-labeling unlabeled soundscapes provides a significant boost if segment selection is careful.\n* Pooling strategies and label weight tuning can give CNNs a few more points.\n* SED is a powerful approach for multi-species data, but there’s a learning curve if you’re new to it (like me!).\n* Pretraining on historical BirdCLEF data really helps.\n* The Kaggle community is an incredible resource—leveraging public notebooks and shared solutions is invaluable, especially when time is short.\n\n---\n\nI learned a ton during this competition—about audio modeling, new architectures, and pushing through plateaus. Huge thanks again to Ariel, everyone in the forums, and all those who shared code and kernels. I’m already looking forward to next year and hoping to break into the top 10!\n\n---\n\nIf you want to chat about the technical details, see code, or ask anything about the pipeline, feel free to reach out!",
    "3218929": "Great job on reaching 38th with a 0.902 AUC—your use of GeM pooling and Quantile-Mix ensembling is really impressive.\nCould you share more details on how you balanced the CNN and SED model outputs in the final ensemble?",
    "3219220": "Thanks a lot!\nIn case you missed it, here's my final submission notebook:\n👉 https://www.kaggle.com/code/maxmelichov/bird25-0-902-auc\n\nI based my final predictions on two SED models that were publicity available, combined with a blend of three slightly different CNN pipelines:\nModel 1 – Standard preprocessing with a GeM pooling layer.\nModel 2 – Same as Model 1, but added a denoising step in preprocessing.\nModel 3 – Same as Model 1, but swapped GeM with AdaptivePooling.\ncombined it for single submission using equal weight for model1,2,3.\n\nEach way generated its own submission CSV. I then combined the predictions using Quantile-Mix (with alpha=0.5) to balance between rank-based and raw-score averaging.\n\nThat final blend gave me a 0.902 AUC score. Super happy with the outcome!"
  },
  "source": "meta"
}