{
  "id": 572928,
  "title": "Summary of Techniques from Past Top Solutions (2024)",
  "url": "/competitions/birdclef-2025/discussion/572928",
  "author_name": "Devasy Patel23",
  "post_date": "2025-04-12T09:55:31.161000",
  "votes": 73,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>BirdCLEF: Summary of Techniques from Past Top Solutions (2024)</h1>\n<p>This table summarizes key techniques and approaches observed in the write-ups of the top 10 solutions from a recent BirdCLEF competition (likely 2024, based on common themes like pseudo-labeling). It aims to provide insights and potential starting points for fellow Kagglers tackling BirdCLEF 2025.</p>\n<p><em>Disclaimer: This is based on summaries and may not capture every nuance. Always refer to the original write-ups for full details.</em></p>\n<table>\n<thead>\n<tr>\n<th>Technique Category</th>\n<th>Technique</th>\n<th>Mentioned By (Rank)</th>\n<th>Notes &amp; Context</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Data - Source</strong></td>\n<td>Use only current year's data (train + unlabeled)</td>\n<td>1st</td>\n<td>Some used past years' data for pretraining or added specific samples.</td>\n</tr>\n<tr>\n<td></td>\n<td>Use external data (Xeno Canto)</td>\n<td>3rd, 7th (for extraction), 8th (pretrain)</td>\n<td>Mixed results; 1st, 4th, 7th found it didn't help or hurt. Quality/filtering is key.</td>\n</tr>\n<tr>\n<td></td>\n<td>Use unlabeled soundscapes (Pseudo-labeling)</td>\n<td>1st, 2nd, 3rd, 6th, 7th, 8th, 9th (domain adapt), 10th</td>\n<td><strong>Very common &amp; impactful.</strong> Key for domain adaptation. Often iterative.</td>\n</tr>\n<tr>\n<td><strong>Data - Preprocessing</strong></td>\n<td>Filter/clean data (e.g., using Google Bird Classifier, duplicates)</td>\n<td>1st, 4th, 7th, 8th</td>\n<td>Removing low-quality/irrelevant segments, correcting labels.</td>\n</tr>\n<tr>\n<td></td>\n<td>Use specific audio duration (e.g., 5s, 10s, 15s)</td>\n<td>1st (10s), 2nd (5s), 3rd (5s), 4th (5s/15-20s), 5th (5s)</td>\n<td>5s and 10s seem most common. Longer durations capture more context but increase compute.</td>\n</tr>\n<tr>\n<td></td>\n<td>Upsample/Downsample classes</td>\n<td>3rd, 4th</td>\n<td>Addressing class imbalance, e.g., ensuring minimum samples per class.</td>\n</tr>\n<tr>\n<td></td>\n<td>Cyclic Padding</td>\n<td>1st</td>\n<td>Repeating audio to meet fixed input length for short sounds.</td>\n</tr>\n<tr>\n<td><strong>Data - Augmentation</strong></td>\n<td>Mixup / Cutmix</td>\n<td>1st (Cutmix), 2nd (Mixup), 4th (Mixup/Cutmix), 5th (Mixup), 8th (Mixup/Sumix), 9th (Cutmix/Mixup), 10th (Noise Mix)</td>\n<td>Very common for regularization. Sumix (mixing based on signal power) also used.</td>\n</tr>\n<tr>\n<td></td>\n<td>Noise Addition (Gaussian, Pink, Background)</td>\n<td>4th, 8th</td>\n<td>Adding noise from datasets or synthesized noise.</td>\n</tr>\n<tr>\n<td></td>\n<td>Gain / Random Volume</td>\n<td>4th, 8th</td>\n<td>Randomly adjusting audio volume.</td>\n</tr>\n<tr>\n<td></td>\n<td>SpecAugment (Time/Frequency Masking)</td>\n<td>1st (XY masking), 8th, 9th (worked against)</td>\n<td>Masking blocks in the spectrogram. Mixed results reported.</td>\n</tr>\n<tr>\n<td></td>\n<td>Time/Frequency Stretching</td>\n<td>2nd</td>\n<td>Resizing spectrogram along time/frequency axes.</td>\n</tr>\n<tr>\n<td><strong>Input Features</strong></td>\n<td>Mel Spectrogram</td>\n<td>1st, 2nd, 3rd, 4th, 5th, 8th, 10th</td>\n<td><strong>Standard approach.</strong> Parameters (<code>n_mels</code>, <code>hop_length</code>, etc.) varied widely.</td>\n</tr>\n<tr>\n<td></td>\n<td>Raw Waveform Input</td>\n<td>3rd (Aves), 4th, 5th</td>\n<td>Using 1D CNNs or specialized models (like Aves) directly on audio samples.</td>\n</tr>\n<tr>\n<td></td>\n<td>PCEN</td>\n<td>4th (didn't work)</td>\n<td>Per-Channel Energy Normalization; an alternative to log-mel.</td>\n</tr>\n<tr>\n<td><strong>Model Architecture</strong></td>\n<td>EfficientNet (B0, B3, V2)</td>\n<td>1st (B0), 2nd (B0), 3rd (V2), 5th (B0), 7th (V2), 8th (B0), 10th (V2)</td>\n<td><strong>Extremely popular.</strong> B0 often chosen for efficiency on CPU inference.</td>\n</tr>\n<tr>\n<td></td>\n<td>EfficientViT</td>\n<td>3rd (b0, b1, m3)</td>\n<td>Vision Transformer variant highlighted for speed and performance.</td>\n</tr>\n<tr>\n<td></td>\n<td>SED (Sound Event Detection) Models</td>\n<td>3rd, 4th, 7th, 8th</td>\n<td>Architectures often combining CNNs and RNNs/Attention for audio tasks.</td>\n</tr>\n<tr>\n<td></td>\n<td>Other CNNs (ResNet, SEResNeXt, NFNet, RegNetY)</td>\n<td>1st, 4th, 6th, 7th, 8th, 9th</td>\n<td>Other common image classification backbones adapted for spectrograms.</td>\n</tr>\n<tr>\n<td></td>\n<td>Waveform Models (Aves)</td>\n<td>3rd, 7th</td>\n<td>Transformer-based models operating directly on raw audio.</td>\n</tr>\n<tr>\n<td><strong>Training - Loss</strong></td>\n<td>BCEWithLogitsLoss</td>\n<td>3rd, 4th, 5th, 6th</td>\n<td>Standard Binary Cross Entropy for multi-label classification.</td>\n</tr>\n<tr>\n<td></td>\n<td>CrossEntropyLoss (CE)</td>\n<td>1st</td>\n<td>Used by 1st place, notably combined with sigmoid in inference.</td>\n</tr>\n<tr>\n<td></td>\n<td>Focal Loss</td>\n<td>4th, 8th, 9th (no improvement w/ CAD)</td>\n<td>Addresses class imbalance by down-weighting easy examples.</td>\n</tr>\n<tr>\n<td></td>\n<td>Secondary Label Handling</td>\n<td>1st (Weighted), 3rd (Masked), 8th (Primary)</td>\n<td>Various strategies: ignore, down-weight, or treat like primary labels.</td>\n</tr>\n<tr>\n<td><strong>Training - Strategy</strong></td>\n<td>Pseudo-Labeling / Distillation</td>\n<td>1st, 2nd, 3rd, 6th, 7th, 8th, 10th</td>\n<td>Training on model's own predictions on unlabeled data. Often iterative.</td>\n</tr>\n<tr>\n<td></td>\n<td>Multi-stage Training (Pretrain -&gt; Finetune)</td>\n<td>8th, 10th</td>\n<td>Often pretraining on past data/larger datasets, then finetuning on target data.</td>\n</tr>\n<tr>\n<td></td>\n<td>Checkpoint Averaging / Model Soups</td>\n<td>2nd</td>\n<td>Averaging weights from multiple checkpoints or models for robustness.</td>\n</tr>\n<tr>\n<td></td>\n<td>Domain Adaptation Techniques</td>\n<td>9th (CAD Bottleneck)</td>\n<td>Explicitly training to reduce domain shift (e.g., using domain labels).</td>\n</tr>\n<tr>\n<td></td>\n<td>Split species into subsets</td>\n<td>7th</td>\n<td>Training separate models for subsets (e.g., rare vs. common species).</td>\n</tr>\n<tr>\n<td><strong>Inference/Postproc.</strong></td>\n<td>Ensembling (Mean, Min, Weighted, Geometric)</td>\n<td>1st (Min), 2nd (Mean), 3rd, 4th (Weighted+Geo), 5th, 7th, 8th</td>\n<td><strong>Universal.</strong> Combining diverse models is key. Min-ensembling mentioned by 1st.</td>\n</tr>\n<tr>\n<td></td>\n<td>Temporal Smoothing (Moving Average, Convolution)</td>\n<td>2nd, 3rd, 4th, 6th, 9th</td>\n<td>Averaging predictions over adjacent time windows. Simple &amp; effective.</td>\n</tr>\n<tr>\n<td></td>\n<td>Inference Optimization (OpenVINO / ONNX)</td>\n<td>1st, 3rd, 4th, 5th, 7th (Quantize), 8th</td>\n<td><strong>Crucial</strong> for meeting CPU time limits. INT8 quantization also used.</td>\n</tr>\n<tr>\n<td></td>\n<td>TTA (Test Time Augmentation)</td>\n<td>4th (Time shift)</td>\n<td>Less common than in vision. E.g., predicting on slightly shifted windows.</td>\n</tr>\n<tr>\n<td></td>\n<td>Probability Adjustments / Cut-offs</td>\n<td>3rd, 4th</td>\n<td>Post-hoc adjustments based on thresholds or soundscape-level predictions.</td>\n</tr>\n</tbody>\n</table>\n<h2>Key Takeaways</h2>\n<ul>\n<li><strong>Domain Shift is Real:</strong> Bridging the gap between training data and noisy soundscapes (often via pseudo-labeling) was critical.</li>\n<li><strong>EfficientNets Rule (esp. B0):</strong> A strong and efficient default choice, especially considering inference constraints.</li>\n<li><strong>Ensembling is Non-Negotiable:</strong> Combining diverse models consistently improves results. Diversity can come from different backbones, input features, training data/folds, or hyperparameters.</li>\n<li><strong>Temporal Context Matters:</strong> Smoothing predictions over time is a simple but effective post-processing step.</li>\n<li><strong>Inference Speed is Key:</strong> Optimizing models (OpenVINO/ONNX, quantization) is necessary for the competition format.</li>\n</ul>\n<h2>References</h2>\n<p>The information above was compiled from summaries and write-ups of top solutions from previous BirdCLEF competitions (primarily BirdCLEF 2024). For detailed insights, please search the Kaggle discussions for the respective competition years (e.g., \"BirdCLEF 2024 1st place solution\"). Many winners generously share their code and detailed explanations.</p>\n<hr>\n<p>A huge thank you to all previous top teams for sharing their insightful solutions and contributing to the community's knowledge!</p>",
  "messages": [
    {
      "id": 3177131,
      "postDate": "2025-04-12T09:55:31.160Z",
      "content": "<h1>BirdCLEF: Summary of Techniques from Past Top Solutions (2024)</h1>\n<p>This table summarizes key techniques and approaches observed in the write-ups of the top 10 solutions from a recent BirdCLEF competition (likely 2024, based on common themes like pseudo-labeling). It aims to provide insights and potential starting points for fellow Kagglers tackling BirdCLEF 2025.</p>\n<p><em>Disclaimer: This is based on summaries and may not capture every nuance. Always refer to the original write-ups for full details.</em></p>\n<table>\n<thead>\n<tr>\n<th>Technique Category</th>\n<th>Technique</th>\n<th>Mentioned By (Rank)</th>\n<th>Notes &amp; Context</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Data - Source</strong></td>\n<td>Use only current year's data (train + unlabeled)</td>\n<td>1st</td>\n<td>Some used past years' data for pretraining or added specific samples.</td>\n</tr>\n<tr>\n<td></td>\n<td>Use external data (Xeno Canto)</td>\n<td>3rd, 7th (for extraction), 8th (pretrain)</td>\n<td>Mixed results; 1st, 4th, 7th found it didn't help or hurt. Quality/filtering is key.</td>\n</tr>\n<tr>\n<td></td>\n<td>Use unlabeled soundscapes (Pseudo-labeling)</td>\n<td>1st, 2nd, 3rd, 6th, 7th, 8th, 9th (domain adapt), 10th</td>\n<td><strong>Very common &amp; impactful.</strong> Key for domain adaptation. Often iterative.</td>\n</tr>\n<tr>\n<td><strong>Data - Preprocessing</strong></td>\n<td>Filter/clean data (e.g., using Google Bird Classifier, duplicates)</td>\n<td>1st, 4th, 7th, 8th</td>\n<td>Removing low-quality/irrelevant segments, correcting labels.</td>\n</tr>\n<tr>\n<td></td>\n<td>Use specific audio duration (e.g., 5s, 10s, 15s)</td>\n<td>1st (10s), 2nd (5s), 3rd (5s), 4th (5s/15-20s), 5th (5s)</td>\n<td>5s and 10s seem most common. Longer durations capture more context but increase compute.</td>\n</tr>\n<tr>\n<td></td>\n<td>Upsample/Downsample classes</td>\n<td>3rd, 4th</td>\n<td>Addressing class imbalance, e.g., ensuring minimum samples per class.</td>\n</tr>\n<tr>\n<td></td>\n<td>Cyclic Padding</td>\n<td>1st</td>\n<td>Repeating audio to meet fixed input length for short sounds.</td>\n</tr>\n<tr>\n<td><strong>Data - Augmentation</strong></td>\n<td>Mixup / Cutmix</td>\n<td>1st (Cutmix), 2nd (Mixup), 4th (Mixup/Cutmix), 5th (Mixup), 8th (Mixup/Sumix), 9th (Cutmix/Mixup), 10th (Noise Mix)</td>\n<td>Very common for regularization. Sumix (mixing based on signal power) also used.</td>\n</tr>\n<tr>\n<td></td>\n<td>Noise Addition (Gaussian, Pink, Background)</td>\n<td>4th, 8th</td>\n<td>Adding noise from datasets or synthesized noise.</td>\n</tr>\n<tr>\n<td></td>\n<td>Gain / Random Volume</td>\n<td>4th, 8th</td>\n<td>Randomly adjusting audio volume.</td>\n</tr>\n<tr>\n<td></td>\n<td>SpecAugment (Time/Frequency Masking)</td>\n<td>1st (XY masking), 8th, 9th (worked against)</td>\n<td>Masking blocks in the spectrogram. Mixed results reported.</td>\n</tr>\n<tr>\n<td></td>\n<td>Time/Frequency Stretching</td>\n<td>2nd</td>\n<td>Resizing spectrogram along time/frequency axes.</td>\n</tr>\n<tr>\n<td><strong>Input Features</strong></td>\n<td>Mel Spectrogram</td>\n<td>1st, 2nd, 3rd, 4th, 5th, 8th, 10th</td>\n<td><strong>Standard approach.</strong> Parameters (<code>n_mels</code>, <code>hop_length</code>, etc.) varied widely.</td>\n</tr>\n<tr>\n<td></td>\n<td>Raw Waveform Input</td>\n<td>3rd (Aves), 4th, 5th</td>\n<td>Using 1D CNNs or specialized models (like Aves) directly on audio samples.</td>\n</tr>\n<tr>\n<td></td>\n<td>PCEN</td>\n<td>4th (didn't work)</td>\n<td>Per-Channel Energy Normalization; an alternative to log-mel.</td>\n</tr>\n<tr>\n<td><strong>Model Architecture</strong></td>\n<td>EfficientNet (B0, B3, V2)</td>\n<td>1st (B0), 2nd (B0), 3rd (V2), 5th (B0), 7th (V2), 8th (B0), 10th (V2)</td>\n<td><strong>Extremely popular.</strong> B0 often chosen for efficiency on CPU inference.</td>\n</tr>\n<tr>\n<td></td>\n<td>EfficientViT</td>\n<td>3rd (b0, b1, m3)</td>\n<td>Vision Transformer variant highlighted for speed and performance.</td>\n</tr>\n<tr>\n<td></td>\n<td>SED (Sound Event Detection) Models</td>\n<td>3rd, 4th, 7th, 8th</td>\n<td>Architectures often combining CNNs and RNNs/Attention for audio tasks.</td>\n</tr>\n<tr>\n<td></td>\n<td>Other CNNs (ResNet, SEResNeXt, NFNet, RegNetY)</td>\n<td>1st, 4th, 6th, 7th, 8th, 9th</td>\n<td>Other common image classification backbones adapted for spectrograms.</td>\n</tr>\n<tr>\n<td></td>\n<td>Waveform Models (Aves)</td>\n<td>3rd, 7th</td>\n<td>Transformer-based models operating directly on raw audio.</td>\n</tr>\n<tr>\n<td><strong>Training - Loss</strong></td>\n<td>BCEWithLogitsLoss</td>\n<td>3rd, 4th, 5th, 6th</td>\n<td>Standard Binary Cross Entropy for multi-label classification.</td>\n</tr>\n<tr>\n<td></td>\n<td>CrossEntropyLoss (CE)</td>\n<td>1st</td>\n<td>Used by 1st place, notably combined with sigmoid in inference.</td>\n</tr>\n<tr>\n<td></td>\n<td>Focal Loss</td>\n<td>4th, 8th, 9th (no improvement w/ CAD)</td>\n<td>Addresses class imbalance by down-weighting easy examples.</td>\n</tr>\n<tr>\n<td></td>\n<td>Secondary Label Handling</td>\n<td>1st (Weighted), 3rd (Masked), 8th (Primary)</td>\n<td>Various strategies: ignore, down-weight, or treat like primary labels.</td>\n</tr>\n<tr>\n<td><strong>Training - Strategy</strong></td>\n<td>Pseudo-Labeling / Distillation</td>\n<td>1st, 2nd, 3rd, 6th, 7th, 8th, 10th</td>\n<td>Training on model's own predictions on unlabeled data. Often iterative.</td>\n</tr>\n<tr>\n<td></td>\n<td>Multi-stage Training (Pretrain -&gt; Finetune)</td>\n<td>8th, 10th</td>\n<td>Often pretraining on past data/larger datasets, then finetuning on target data.</td>\n</tr>\n<tr>\n<td></td>\n<td>Checkpoint Averaging / Model Soups</td>\n<td>2nd</td>\n<td>Averaging weights from multiple checkpoints or models for robustness.</td>\n</tr>\n<tr>\n<td></td>\n<td>Domain Adaptation Techniques</td>\n<td>9th (CAD Bottleneck)</td>\n<td>Explicitly training to reduce domain shift (e.g., using domain labels).</td>\n</tr>\n<tr>\n<td></td>\n<td>Split species into subsets</td>\n<td>7th</td>\n<td>Training separate models for subsets (e.g., rare vs. common species).</td>\n</tr>\n<tr>\n<td><strong>Inference/Postproc.</strong></td>\n<td>Ensembling (Mean, Min, Weighted, Geometric)</td>\n<td>1st (Min), 2nd (Mean), 3rd, 4th (Weighted+Geo), 5th, 7th, 8th</td>\n<td><strong>Universal.</strong> Combining diverse models is key. Min-ensembling mentioned by 1st.</td>\n</tr>\n<tr>\n<td></td>\n<td>Temporal Smoothing (Moving Average, Convolution)</td>\n<td>2nd, 3rd, 4th, 6th, 9th</td>\n<td>Averaging predictions over adjacent time windows. Simple &amp; effective.</td>\n</tr>\n<tr>\n<td></td>\n<td>Inference Optimization (OpenVINO / ONNX)</td>\n<td>1st, 3rd, 4th, 5th, 7th (Quantize), 8th</td>\n<td><strong>Crucial</strong> for meeting CPU time limits. INT8 quantization also used.</td>\n</tr>\n<tr>\n<td></td>\n<td>TTA (Test Time Augmentation)</td>\n<td>4th (Time shift)</td>\n<td>Less common than in vision. E.g., predicting on slightly shifted windows.</td>\n</tr>\n<tr>\n<td></td>\n<td>Probability Adjustments / Cut-offs</td>\n<td>3rd, 4th</td>\n<td>Post-hoc adjustments based on thresholds or soundscape-level predictions.</td>\n</tr>\n</tbody>\n</table>\n<h2>Key Takeaways</h2>\n<ul>\n<li><strong>Domain Shift is Real:</strong> Bridging the gap between training data and noisy soundscapes (often via pseudo-labeling) was critical.</li>\n<li><strong>EfficientNets Rule (esp. B0):</strong> A strong and efficient default choice, especially considering inference constraints.</li>\n<li><strong>Ensembling is Non-Negotiable:</strong> Combining diverse models consistently improves results. Diversity can come from different backbones, input features, training data/folds, or hyperparameters.</li>\n<li><strong>Temporal Context Matters:</strong> Smoothing predictions over time is a simple but effective post-processing step.</li>\n<li><strong>Inference Speed is Key:</strong> Optimizing models (OpenVINO/ONNX, quantization) is necessary for the competition format.</li>\n</ul>\n<h2>References</h2>\n<p>The information above was compiled from summaries and write-ups of top solutions from previous BirdCLEF competitions (primarily BirdCLEF 2024). For detailed insights, please search the Kaggle discussions for the respective competition years (e.g., \"BirdCLEF 2024 1st place solution\"). Many winners generously share their code and detailed explanations.</p>\n<hr>\n<p>A huge thank you to all previous top teams for sharing their insightful solutions and contributing to the community's knowledge!</p>",
      "rawMarkdown": "# BirdCLEF: Summary of Techniques from Past Top Solutions (2024)\n\nThis table summarizes key techniques and approaches observed in the write-ups of the top 10 solutions from a recent BirdCLEF competition (likely 2024, based on common themes like pseudo-labeling). It aims to provide insights and potential starting points for fellow Kagglers tackling BirdCLEF 2025.\n\n*Disclaimer: This is based on summaries and may not capture every nuance. Always refer to the original write-ups for full details.*\n\n| Technique Category        | Technique                                                                 | Mentioned By (Rank)                                  | Notes & Context                                                                 |\n| :------------------------ | :------------------------------------------------------------------------ | :--------------------------------------------------- | :------------------------------------------------------------------------------ |\n| **Data - Source**         | Use only current year's data (train + unlabeled)                          | 1st                                                  | Some used past years' data for pretraining or added specific samples.           |\n|                           | Use external data (Xeno Canto)                                            | 3rd, 7th (for extraction), 8th (pretrain)            | Mixed results; 1st, 4th, 7th found it didn't help or hurt. Quality/filtering is key. |\n|                           | Use unlabeled soundscapes (Pseudo-labeling)                               | 1st, 2nd, 3rd, 6th, 7th, 8th, 9th (domain adapt), 10th | **Very common & impactful.** Key for domain adaptation. Often iterative.        |\n| **Data - Preprocessing**  | Filter/clean data (e.g., using Google Bird Classifier, duplicates)        | 1st, 4th, 7th, 8th                                   | Removing low-quality/irrelevant segments, correcting labels.                    |\n|                           | Use specific audio duration (e.g., 5s, 10s, 15s)                          | 1st (10s), 2nd (5s), 3rd (5s), 4th (5s/15-20s), 5th (5s) | 5s and 10s seem most common. Longer durations capture more context but increase compute. |\n|                           | Upsample/Downsample classes                                               | 3rd, 4th                                             | Addressing class imbalance, e.g., ensuring minimum samples per class.           |\n|                           | Cyclic Padding                                                            | 1st                                                  | Repeating audio to meet fixed input length for short sounds.                    |\n| **Data - Augmentation**   | Mixup / Cutmix                                                            | 1st (Cutmix), 2nd (Mixup), 4th (Mixup/Cutmix), 5th (Mixup), 8th (Mixup/Sumix), 9th (Cutmix/Mixup), 10th (Noise Mix) | Very common for regularization. Sumix (mixing based on signal power) also used. |\n|                           | Noise Addition (Gaussian, Pink, Background)                               | 4th, 8th                                             | Adding noise from datasets or synthesized noise.                                |\n|                           | Gain / Random Volume                                                      | 4th, 8th                                             | Randomly adjusting audio volume.                                                |\n|                           | SpecAugment (Time/Frequency Masking)                                      | 1st (XY masking), 8th, 9th (worked against)          | Masking blocks in the spectrogram. Mixed results reported.                      |\n|                           | Time/Frequency Stretching                                                 | 2nd                                                  | Resizing spectrogram along time/frequency axes.                                 |\n| **Input Features**        | Mel Spectrogram                                                           | 1st, 2nd, 3rd, 4th, 5th, 8th, 10th                   | **Standard approach.** Parameters (`n_mels`, `hop_length`, etc.) varied widely. |\n|                           | Raw Waveform Input                                                        | 3rd (Aves), 4th, 5th                                 | Using 1D CNNs or specialized models (like Aves) directly on audio samples.      |\n|                           | PCEN                                                                      | 4th (didn't work)                                    | Per-Channel Energy Normalization; an alternative to log-mel.                    |\n| **Model Architecture**    | EfficientNet (B0, B3, V2)                                                 | 1st (B0), 2nd (B0), 3rd (V2), 5th (B0), 7th (V2), 8th (B0), 10th (V2) | **Extremely popular.** B0 often chosen for efficiency on CPU inference.         |\n|                           | EfficientViT                                                              | 3rd (b0, b1, m3)                                     | Vision Transformer variant highlighted for speed and performance.               |\n|                           | SED (Sound Event Detection) Models                                        | 3rd, 4th, 7th, 8th                                   | Architectures often combining CNNs and RNNs/Attention for audio tasks.          |\n|                           | Other CNNs (ResNet, SEResNeXt, NFNet, RegNetY)                            | 1st, 4th, 6th, 7th, 8th, 9th                         | Other common image classification backbones adapted for spectrograms.           |\n|                           | Waveform Models (Aves)                                                    | 3rd, 7th                                             | Transformer-based models operating directly on raw audio.                       |\n| **Training - Loss**       | BCEWithLogitsLoss                                                         | 3rd, 4th, 5th, 6th                                   | Standard Binary Cross Entropy for multi-label classification.                   |\n|                           | CrossEntropyLoss (CE)                                                     | 1st                                                  | Used by 1st place, notably combined with sigmoid in inference.                  |\n|                           | Focal Loss                                                                | 4th, 8th, 9th (no improvement w/ CAD)                | Addresses class imbalance by down-weighting easy examples.                      |\n|                           | Secondary Label Handling                                                  | 1st (Weighted), 3rd (Masked), 8th (Primary)          | Various strategies: ignore, down-weight, or treat like primary labels.          |\n| **Training - Strategy**   | Pseudo-Labeling / Distillation                                            | 1st, 2nd, 3rd, 6th, 7th, 8th, 10th                   | Training on model's own predictions on unlabeled data. Often iterative.         |\n|                           | Multi-stage Training (Pretrain -> Finetune)                               | 8th, 10th                                            | Often pretraining on past data/larger datasets, then finetuning on target data. |\n|                           | Checkpoint Averaging / Model Soups                                        | 2nd                                                  | Averaging weights from multiple checkpoints or models for robustness.           |\n|                           | Domain Adaptation Techniques                                              | 9th (CAD Bottleneck)                                 | Explicitly training to reduce domain shift (e.g., using domain labels).         |\n|                           | Split species into subsets                                                | 7th                                                  | Training separate models for subsets (e.g., rare vs. common species).           |\n| **Inference/Postproc.**   | Ensembling (Mean, Min, Weighted, Geometric)                               | 1st (Min), 2nd (Mean), 3rd, 4th (Weighted+Geo), 5th, 7th, 8th | **Universal.** Combining diverse models is key. Min-ensembling mentioned by 1st. |\n|                           | Temporal Smoothing (Moving Average, Convolution)                          | 2nd, 3rd, 4th, 6th, 9th                              | Averaging predictions over adjacent time windows. Simple & effective.           |\n|                           | Inference Optimization (OpenVINO / ONNX)                                  | 1st, 3rd, 4th, 5th, 7th (Quantize), 8th              | **Crucial** for meeting CPU time limits. INT8 quantization also used.           |\n|                           | TTA (Test Time Augmentation)                                              | 4th (Time shift)                                     | Less common than in vision. E.g., predicting on slightly shifted windows.       |\n|                           | Probability Adjustments / Cut-offs                                        | 3rd, 4th                                             | Post-hoc adjustments based on thresholds or soundscape-level predictions.       |\n\n## Key Takeaways\n\n*   **Domain Shift is Real:** Bridging the gap between training data and noisy soundscapes (often via pseudo-labeling) was critical.\n*   **EfficientNets Rule (esp. B0):** A strong and efficient default choice, especially considering inference constraints.\n*   **Ensembling is Non-Negotiable:** Combining diverse models consistently improves results. Diversity can come from different backbones, input features, training data/folds, or hyperparameters.\n*   **Temporal Context Matters:** Smoothing predictions over time is a simple but effective post-processing step.\n*   **Inference Speed is Key:** Optimizing models (OpenVINO/ONNX, quantization) is necessary for the competition format.\n\n## References\n\nThe information above was compiled from summaries and write-ups of top solutions from previous BirdCLEF competitions (primarily BirdCLEF 2024). For detailed insights, please search the Kaggle discussions for the respective competition years (e.g., \"BirdCLEF 2024 1st place solution\"). Many winners generously share their code and detailed explanations.\n\n---\n\nA huge thank you to all previous top teams for sharing their insightful solutions and contributing to the community's knowledge!\n",
      "votes": 72
    },
    {
      "id": 3180778,
      "postDate": "2025-04-17T05:21:08.270Z",
      "content": "<p>This is good , One can start from here to learn on this topic </p>",
      "rawMarkdown": "This is good , One can start from here to learn on this topic ",
      "votes": 1
    },
    {
      "id": 3205776,
      "postDate": "2025-05-20T10:39:32.517Z",
      "content": "<p>Thnx, this is grat and very helpfuld to get started !</p>",
      "rawMarkdown": "Thnx, this is grat and very helpfuld to get started !"
    },
    {
      "id": 3181836,
      "postDate": "2025-04-18T12:14:21.587Z",
      "content": "<p>Thanks for summarizing everything in one place ? </p>",
      "rawMarkdown": "Thanks for summarizing everything in one place ? "
    },
    {
      "id": 3180110,
      "postDate": "2025-04-16T07:02:33.880Z",
      "content": "<p>That's a great summary table, thanks for sharing it.</p>",
      "rawMarkdown": "That's a great summary table, thanks for sharing it."
    },
    {
      "id": 3201013,
      "postDate": "2025-05-13T10:51:54.207Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3201129,
          "postDate": "2025-05-13T12:54:31.570Z",
          "content": "<p>მადლობა დაფასებისთვის, ძალიან ბევრს ნიშნავს</p>",
          "rawMarkdown": "მადლობა დაფასებისთვის, ძალიან ბევრს ნიშნავს",
          "votes": 1
        }
      ]
    },
    {
      "id": 3182623,
      "postDate": "2025-04-19T16:31:41.440Z",
      "content": "<p>Good summary, thanks</p>",
      "rawMarkdown": "Good summary, thanks"
    }
  ],
  "comments": [
    {
      "id": 3180778,
      "author_name": "Puneet Arora",
      "author_url": "",
      "post_date": "2025-04-17T05:21:08.270000",
      "content": "<p>This is good , One can start from here to learn on this topic </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3205776,
      "author_name": "Dric225",
      "author_url": "",
      "post_date": "2025-05-20T10:39:32.517000",
      "content": "<p>Thnx, this is grat and very helpfuld to get started !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3181836,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2025-04-18T12:14:21.587000",
      "content": "<p>Thanks for summarizing everything in one place ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3180110,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2025-04-16T07:02:33.880000",
      "content": "<p>That's a great summary table, thanks for sharing it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3201013,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-13T10:51:54.207000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3201129,
          "author_name": "Devasy Patel23",
          "author_url": "",
          "post_date": "2025-05-13T12:54:31.570000",
          "content": "<p>მადლობა დაფასებისთვის, ძალიან ბევრს ნიშნავს</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3182623,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2025-04-19T16:31:41.440000",
      "content": "<p>Good summary, thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3177131": "# BirdCLEF: Summary of Techniques from Past Top Solutions (2024)\n\nThis table summarizes key techniques and approaches observed in the write-ups of the top 10 solutions from a recent BirdCLEF competition (likely 2024, based on common themes like pseudo-labeling). It aims to provide insights and potential starting points for fellow Kagglers tackling BirdCLEF 2025.\n\n*Disclaimer: This is based on summaries and may not capture every nuance. Always refer to the original write-ups for full details.*\n\n| Technique Category        | Technique                                                                 | Mentioned By (Rank)                                  | Notes & Context                                                                 |\n| :------------------------ | :------------------------------------------------------------------------ | :--------------------------------------------------- | :------------------------------------------------------------------------------ |\n| **Data - Source**         | Use only current year's data (train + unlabeled)                          | 1st                                                  | Some used past years' data for pretraining or added specific samples.           |\n|                           | Use external data (Xeno Canto)                                            | 3rd, 7th (for extraction), 8th (pretrain)            | Mixed results; 1st, 4th, 7th found it didn't help or hurt. Quality/filtering is key. |\n|                           | Use unlabeled soundscapes (Pseudo-labeling)                               | 1st, 2nd, 3rd, 6th, 7th, 8th, 9th (domain adapt), 10th | **Very common & impactful.** Key for domain adaptation. Often iterative.        |\n| **Data - Preprocessing**  | Filter/clean data (e.g., using Google Bird Classifier, duplicates)        | 1st, 4th, 7th, 8th                                   | Removing low-quality/irrelevant segments, correcting labels.                    |\n|                           | Use specific audio duration (e.g., 5s, 10s, 15s)                          | 1st (10s), 2nd (5s), 3rd (5s), 4th (5s/15-20s), 5th (5s) | 5s and 10s seem most common. Longer durations capture more context but increase compute. |\n|                           | Upsample/Downsample classes                                               | 3rd, 4th                                             | Addressing class imbalance, e.g., ensuring minimum samples per class.           |\n|                           | Cyclic Padding                                                            | 1st                                                  | Repeating audio to meet fixed input length for short sounds.                    |\n| **Data - Augmentation**   | Mixup / Cutmix                                                            | 1st (Cutmix), 2nd (Mixup), 4th (Mixup/Cutmix), 5th (Mixup), 8th (Mixup/Sumix), 9th (Cutmix/Mixup), 10th (Noise Mix) | Very common for regularization. Sumix (mixing based on signal power) also used. |\n|                           | Noise Addition (Gaussian, Pink, Background)                               | 4th, 8th                                             | Adding noise from datasets or synthesized noise.                                |\n|                           | Gain / Random Volume                                                      | 4th, 8th                                             | Randomly adjusting audio volume.                                                |\n|                           | SpecAugment (Time/Frequency Masking)                                      | 1st (XY masking), 8th, 9th (worked against)          | Masking blocks in the spectrogram. Mixed results reported.                      |\n|                           | Time/Frequency Stretching                                                 | 2nd                                                  | Resizing spectrogram along time/frequency axes.                                 |\n| **Input Features**        | Mel Spectrogram                                                           | 1st, 2nd, 3rd, 4th, 5th, 8th, 10th                   | **Standard approach.** Parameters (`n_mels`, `hop_length`, etc.) varied widely. |\n|                           | Raw Waveform Input                                                        | 3rd (Aves), 4th, 5th                                 | Using 1D CNNs or specialized models (like Aves) directly on audio samples.      |\n|                           | PCEN                                                                      | 4th (didn't work)                                    | Per-Channel Energy Normalization; an alternative to log-mel.                    |\n| **Model Architecture**    | EfficientNet (B0, B3, V2)                                                 | 1st (B0), 2nd (B0), 3rd (V2), 5th (B0), 7th (V2), 8th (B0), 10th (V2) | **Extremely popular.** B0 often chosen for efficiency on CPU inference.         |\n|                           | EfficientViT                                                              | 3rd (b0, b1, m3)                                     | Vision Transformer variant highlighted for speed and performance.               |\n|                           | SED (Sound Event Detection) Models                                        | 3rd, 4th, 7th, 8th                                   | Architectures often combining CNNs and RNNs/Attention for audio tasks.          |\n|                           | Other CNNs (ResNet, SEResNeXt, NFNet, RegNetY)                            | 1st, 4th, 6th, 7th, 8th, 9th                         | Other common image classification backbones adapted for spectrograms.           |\n|                           | Waveform Models (Aves)                                                    | 3rd, 7th                                             | Transformer-based models operating directly on raw audio.                       |\n| **Training - Loss**       | BCEWithLogitsLoss                                                         | 3rd, 4th, 5th, 6th                                   | Standard Binary Cross Entropy for multi-label classification.                   |\n|                           | CrossEntropyLoss (CE)                                                     | 1st                                                  | Used by 1st place, notably combined with sigmoid in inference.                  |\n|                           | Focal Loss                                                                | 4th, 8th, 9th (no improvement w/ CAD)                | Addresses class imbalance by down-weighting easy examples.                      |\n|                           | Secondary Label Handling                                                  | 1st (Weighted), 3rd (Masked), 8th (Primary)          | Various strategies: ignore, down-weight, or treat like primary labels.          |\n| **Training - Strategy**   | Pseudo-Labeling / Distillation                                            | 1st, 2nd, 3rd, 6th, 7th, 8th, 10th                   | Training on model's own predictions on unlabeled data. Often iterative.         |\n|                           | Multi-stage Training (Pretrain -> Finetune)                               | 8th, 10th                                            | Often pretraining on past data/larger datasets, then finetuning on target data. |\n|                           | Checkpoint Averaging / Model Soups                                        | 2nd                                                  | Averaging weights from multiple checkpoints or models for robustness.           |\n|                           | Domain Adaptation Techniques                                              | 9th (CAD Bottleneck)                                 | Explicitly training to reduce domain shift (e.g., using domain labels).         |\n|                           | Split species into subsets                                                | 7th                                                  | Training separate models for subsets (e.g., rare vs. common species).           |\n| **Inference/Postproc.**   | Ensembling (Mean, Min, Weighted, Geometric)                               | 1st (Min), 2nd (Mean), 3rd, 4th (Weighted+Geo), 5th, 7th, 8th | **Universal.** Combining diverse models is key. Min-ensembling mentioned by 1st. |\n|                           | Temporal Smoothing (Moving Average, Convolution)                          | 2nd, 3rd, 4th, 6th, 9th                              | Averaging predictions over adjacent time windows. Simple & effective.           |\n|                           | Inference Optimization (OpenVINO / ONNX)                                  | 1st, 3rd, 4th, 5th, 7th (Quantize), 8th              | **Crucial** for meeting CPU time limits. INT8 quantization also used.           |\n|                           | TTA (Test Time Augmentation)                                              | 4th (Time shift)                                     | Less common than in vision. E.g., predicting on slightly shifted windows.       |\n|                           | Probability Adjustments / Cut-offs                                        | 3rd, 4th                                             | Post-hoc adjustments based on thresholds or soundscape-level predictions.       |\n\n## Key Takeaways\n\n*   **Domain Shift is Real:** Bridging the gap between training data and noisy soundscapes (often via pseudo-labeling) was critical.\n*   **EfficientNets Rule (esp. B0):** A strong and efficient default choice, especially considering inference constraints.\n*   **Ensembling is Non-Negotiable:** Combining diverse models consistently improves results. Diversity can come from different backbones, input features, training data/folds, or hyperparameters.\n*   **Temporal Context Matters:** Smoothing predictions over time is a simple but effective post-processing step.\n*   **Inference Speed is Key:** Optimizing models (OpenVINO/ONNX, quantization) is necessary for the competition format.\n\n## References\n\nThe information above was compiled from summaries and write-ups of top solutions from previous BirdCLEF competitions (primarily BirdCLEF 2024). For detailed insights, please search the Kaggle discussions for the respective competition years (e.g., \"BirdCLEF 2024 1st place solution\"). Many winners generously share their code and detailed explanations.\n\n---\n\nA huge thank you to all previous top teams for sharing their insightful solutions and contributing to the community's knowledge!\n",
    "3180778": "This is good , One can start from here to learn on this topic ",
    "3205776": "Thnx, this is grat and very helpfuld to get started !",
    "3181836": "Thanks for summarizing everything in one place ? ",
    "3180110": "That's a great summary table, thanks for sharing it.",
    "3201013": "",
    "3182623": "Good summary, thanks"
  }
}