{
  "id": 596605,
  "title": "18th place solution: Sound Event Detection (SED) with attention",
  "url": "/competitions/birdclef-2023/discussion/596605",
  "author_name": "Warrior",
  "post_date": "2025-08-04T16:15:32.529000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>BirdCLEF 2023 — Sound Event Detection with Attention on Mel-Spectrograms</h1>\n<h2>1. Problem</h2>\n<p>Identify bird species from environmental audio recordings. The challenge is multi-label, highly imbalanced, and requires both clipwise and framewise understanding. The evaluation uses padded class-wise average precision (cmap), which favors robust ranking across rare and common species.</p>\n<h2>2. Data &amp; Preprocessing</h2>\n<ul>\n<li><strong>Raw audio</strong>: <code>.ogg</code> files organized by primary label are loaded, converted to mono, resampled to 32kHz, and truncated to fixed-length segments (default 30s). This is handled in <code>prepare_audios.py</code>, producing <code>.npy</code> arrays under <code>train_np/&lt;primary_label&gt;/...</code>.</li>\n<li><strong>Metadata merge</strong>: <code>split_data.py</code> merges <code>train_metadata.csv</code> with the generated audio paths, sanitizes secondary labels, constructs combined label strings, and assigns stratified k-fold splits on <code>primary_label</code> with special handling to keep very rare birds present in training folds.</li>\n<li><strong>Input construction</strong>: Each training sample is a <strong>composite mel-spectrogram</strong> built by mixing multiple audio samples (primary + background) with random early stopping, contrast adjustment, and noise injection. Target vectors apply label smoothing and weight background species lower than primaries.</li>\n</ul>\n<h2>3. Augmentation</h2>\n<ul>\n<li><strong>Audio-level</strong>: Normalization, white noise, pink noise (precomputed), random volume modulation, pitch shift, and time stretch.</li>\n<li><strong>Spectrogram-level</strong>: Random power (contrast), frequency shaping, CutMix and MixUp (batch-level mixing), and composite mixing of multiple bird signals to simulate overlapping soundscapes.</li>\n<li><strong>Label smoothing</strong>: Primary species target ~0.995; background species ~0.3; absence is near zero to stabilize learning.</li>\n</ul>\n<h2>4. Model Architecture</h2>\n<ul>\n<li>Backbone: EfficientNet variant (<code>tf_efficientnet_b0_ns</code>) via <code>timm</code>, adapted to spectrogram inputs.</li>\n<li>Spectrogram is fed through a light preprocessor (batchnorm, optional spec augmentation).</li>\n<li>Attention-based aggregation: Two parallel attention blocks (<code>AttBlockV2</code>) with sigmoid activations produce clipwise and framewise outputs; outputs are combined (averaged) to yield final predictions.</li>\n<li>Auxiliary pooling: Max + average pooling is used before linear projection to blend local and global context.</li>\n<li>Custom layers/helper functions support TensorFlow-style \"same\" padding, upsampling/interpolation, and framewise alignment.</li>\n</ul>\n<h2>5. Loss &amp; Metric</h2>\n<ul>\n<li><strong>Loss</strong>: Two-way focal loss (<code>BCEFocal2WayLoss</code>) combining main logits and aggregated framewise logits to mitigate class imbalance and emphasize hard examples.</li>\n<li><strong>Metric</strong>: Padded class-wise average precision (<code>padded_cmap</code>) with padding to stabilize on sparse labels; primary evaluation uses <code>cmap_pad_5</code>.</li>\n</ul>\n<h2>6. Training Strategy</h2>\n<ul>\n<li><strong>K-fold cross-validation</strong>: Stratified on <code>primary_label</code> (default 5 folds), with explicit handling so extremely rare species remain in training.</li>\n<li><strong>Optimizer</strong>: <code>AdamW</code> from <code>transformers</code> with weight decay.</li>\n<li><strong>Scheduler</strong>: Flexible scheduler selection (e.g., ReduceLROnPlateau, cosine annealing) controlled via config.</li>\n<li><strong>Mixed precision</strong>: Enabled via PyTorch native AMP (<code>torch.cuda.amp</code>) for speed and stability.</li>\n<li><strong>Checkpointing</strong>: Best models saved based on validation loss with descriptive filenames; resume capability via latest checkpoint lookup.</li>\n</ul>\n<h2>7. Implementation Details</h2>\n<ul>\n<li>Centralized config (<code>configs.py</code>) for experiment parameters (learning rate, batch size, scheduler, audio/spec parameters, etc.).</li>\n<li>Logging verbosity toggled via <code>config.debug</code>.</li>\n<li>Dataset builds on-the-fly composite spectrograms mixing 2–3 examples per training instance to diversify signal combinations.</li>\n<li>Rare bird names are forced out of validation through fold adjustment to ensure gradient signal.</li>\n</ul>\n<h2>8. Inference</h2>\n<ul>\n<li>Use trained model to compute clipwise outputs and optionally aggregate framewise predictions by upsampling/interpolation to match input length.</li>\n<li>Final predictions can be thresholded or used directly in ranking for leaderboard submission.</li>\n<li>A dedicated inference notebook is available: <a href=\"https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0\" target=\"_blank\">BirdCLEF 2023 Inference ver 3.0 on Kaggle</a>.</li>\n</ul>\n<h2>9. Key Hyperparameters (example)</h2>\n<ul>\n<li>Backbone: <code>tf_efficientnet_b0_ns</code></li>\n<li>Input: Mel-spectrogram with <code>n_mels=224</code>, sample rate 32000, period 5s segments</li>\n<li>Batch size: 32 (train &amp; valid)</li>\n<li>Initial LR: 1e-3 with weight decay 1e-6</li>\n<li>Scheduler: ReduceLROnPlateau with factor=0.9, patience=5 (configurable)</li>\n<li>Label smoothing: Enabled</li>\n<li>Mix composition: Up to 3 audio clips per training sample</li>\n</ul>\n<h2>10. Results Tracking</h2>\n<p>Training histories (loss and cmap scores) saved per fold. Visualization provided via <code>plot_loss_metric.py</code>.</p>\n<h2>11. Code &amp; Notebooks</h2>\n<ul>\n<li>Training code &amp; full pipeline: <a href=\"https://github.com/tjamali/kaggle-birdclef-2023/tree/main\" target=\"_blank\">GitHub - tjamali/kaggle-birdclef-2023</a>  </li>\n<li>Inference notebook: <a href=\"https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0\" target=\"_blank\">Kaggle Notebook - BirdCLEF 2023 Inference ver 3.0</a></li>\n</ul>\n<h2>12. Lessons &amp; Next Steps</h2>\n<ul>\n<li>Composite mixing plus attention helps the model disambiguate overlapping bird calls.</li>\n<li>Label smoothing and dual focal losses stabilize training on rare classes.</li>\n<li>Future work: strong ensembling across folds/backbones, test-time augmentation, and consolidating framewise outputs for temporal localization.</li>\n</ul>",
  "messages": [
    {
      "id": 3262997,
      "postDate": "2025-08-04T16:15:32.530Z",
      "content": "<h1>BirdCLEF 2023 — Sound Event Detection with Attention on Mel-Spectrograms</h1>\n<h2>1. Problem</h2>\n<p>Identify bird species from environmental audio recordings. The challenge is multi-label, highly imbalanced, and requires both clipwise and framewise understanding. The evaluation uses padded class-wise average precision (cmap), which favors robust ranking across rare and common species.</p>\n<h2>2. Data &amp; Preprocessing</h2>\n<ul>\n<li><strong>Raw audio</strong>: <code>.ogg</code> files organized by primary label are loaded, converted to mono, resampled to 32kHz, and truncated to fixed-length segments (default 30s). This is handled in <code>prepare_audios.py</code>, producing <code>.npy</code> arrays under <code>train_np/&lt;primary_label&gt;/...</code>.</li>\n<li><strong>Metadata merge</strong>: <code>split_data.py</code> merges <code>train_metadata.csv</code> with the generated audio paths, sanitizes secondary labels, constructs combined label strings, and assigns stratified k-fold splits on <code>primary_label</code> with special handling to keep very rare birds present in training folds.</li>\n<li><strong>Input construction</strong>: Each training sample is a <strong>composite mel-spectrogram</strong> built by mixing multiple audio samples (primary + background) with random early stopping, contrast adjustment, and noise injection. Target vectors apply label smoothing and weight background species lower than primaries.</li>\n</ul>\n<h2>3. Augmentation</h2>\n<ul>\n<li><strong>Audio-level</strong>: Normalization, white noise, pink noise (precomputed), random volume modulation, pitch shift, and time stretch.</li>\n<li><strong>Spectrogram-level</strong>: Random power (contrast), frequency shaping, CutMix and MixUp (batch-level mixing), and composite mixing of multiple bird signals to simulate overlapping soundscapes.</li>\n<li><strong>Label smoothing</strong>: Primary species target ~0.995; background species ~0.3; absence is near zero to stabilize learning.</li>\n</ul>\n<h2>4. Model Architecture</h2>\n<ul>\n<li>Backbone: EfficientNet variant (<code>tf_efficientnet_b0_ns</code>) via <code>timm</code>, adapted to spectrogram inputs.</li>\n<li>Spectrogram is fed through a light preprocessor (batchnorm, optional spec augmentation).</li>\n<li>Attention-based aggregation: Two parallel attention blocks (<code>AttBlockV2</code>) with sigmoid activations produce clipwise and framewise outputs; outputs are combined (averaged) to yield final predictions.</li>\n<li>Auxiliary pooling: Max + average pooling is used before linear projection to blend local and global context.</li>\n<li>Custom layers/helper functions support TensorFlow-style \"same\" padding, upsampling/interpolation, and framewise alignment.</li>\n</ul>\n<h2>5. Loss &amp; Metric</h2>\n<ul>\n<li><strong>Loss</strong>: Two-way focal loss (<code>BCEFocal2WayLoss</code>) combining main logits and aggregated framewise logits to mitigate class imbalance and emphasize hard examples.</li>\n<li><strong>Metric</strong>: Padded class-wise average precision (<code>padded_cmap</code>) with padding to stabilize on sparse labels; primary evaluation uses <code>cmap_pad_5</code>.</li>\n</ul>\n<h2>6. Training Strategy</h2>\n<ul>\n<li><strong>K-fold cross-validation</strong>: Stratified on <code>primary_label</code> (default 5 folds), with explicit handling so extremely rare species remain in training.</li>\n<li><strong>Optimizer</strong>: <code>AdamW</code> from <code>transformers</code> with weight decay.</li>\n<li><strong>Scheduler</strong>: Flexible scheduler selection (e.g., ReduceLROnPlateau, cosine annealing) controlled via config.</li>\n<li><strong>Mixed precision</strong>: Enabled via PyTorch native AMP (<code>torch.cuda.amp</code>) for speed and stability.</li>\n<li><strong>Checkpointing</strong>: Best models saved based on validation loss with descriptive filenames; resume capability via latest checkpoint lookup.</li>\n</ul>\n<h2>7. Implementation Details</h2>\n<ul>\n<li>Centralized config (<code>configs.py</code>) for experiment parameters (learning rate, batch size, scheduler, audio/spec parameters, etc.).</li>\n<li>Logging verbosity toggled via <code>config.debug</code>.</li>\n<li>Dataset builds on-the-fly composite spectrograms mixing 2–3 examples per training instance to diversify signal combinations.</li>\n<li>Rare bird names are forced out of validation through fold adjustment to ensure gradient signal.</li>\n</ul>\n<h2>8. Inference</h2>\n<ul>\n<li>Use trained model to compute clipwise outputs and optionally aggregate framewise predictions by upsampling/interpolation to match input length.</li>\n<li>Final predictions can be thresholded or used directly in ranking for leaderboard submission.</li>\n<li>A dedicated inference notebook is available: <a href=\"https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0\" target=\"_blank\">BirdCLEF 2023 Inference ver 3.0 on Kaggle</a>.</li>\n</ul>\n<h2>9. Key Hyperparameters (example)</h2>\n<ul>\n<li>Backbone: <code>tf_efficientnet_b0_ns</code></li>\n<li>Input: Mel-spectrogram with <code>n_mels=224</code>, sample rate 32000, period 5s segments</li>\n<li>Batch size: 32 (train &amp; valid)</li>\n<li>Initial LR: 1e-3 with weight decay 1e-6</li>\n<li>Scheduler: ReduceLROnPlateau with factor=0.9, patience=5 (configurable)</li>\n<li>Label smoothing: Enabled</li>\n<li>Mix composition: Up to 3 audio clips per training sample</li>\n</ul>\n<h2>10. Results Tracking</h2>\n<p>Training histories (loss and cmap scores) saved per fold. Visualization provided via <code>plot_loss_metric.py</code>.</p>\n<h2>11. Code &amp; Notebooks</h2>\n<ul>\n<li>Training code &amp; full pipeline: <a href=\"https://github.com/tjamali/kaggle-birdclef-2023/tree/main\" target=\"_blank\">GitHub - tjamali/kaggle-birdclef-2023</a>  </li>\n<li>Inference notebook: <a href=\"https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0\" target=\"_blank\">Kaggle Notebook - BirdCLEF 2023 Inference ver 3.0</a></li>\n</ul>\n<h2>12. Lessons &amp; Next Steps</h2>\n<ul>\n<li>Composite mixing plus attention helps the model disambiguate overlapping bird calls.</li>\n<li>Label smoothing and dual focal losses stabilize training on rare classes.</li>\n<li>Future work: strong ensembling across folds/backbones, test-time augmentation, and consolidating framewise outputs for temporal localization.</li>\n</ul>",
      "rawMarkdown": "# BirdCLEF 2023 — Sound Event Detection with Attention on Mel-Spectrograms\n\n## 1. Problem\nIdentify bird species from environmental audio recordings. The challenge is multi-label, highly imbalanced, and requires both clipwise and framewise understanding. The evaluation uses padded class-wise average precision (cmap), which favors robust ranking across rare and common species.\n\n## 2. Data & Preprocessing\n- **Raw audio**: `.ogg` files organized by primary label are loaded, converted to mono, resampled to 32kHz, and truncated to fixed-length segments (default 30s). This is handled in `prepare_audios.py`, producing `.npy` arrays under `train_np/<primary_label>/...`.\n- **Metadata merge**: `split_data.py` merges `train_metadata.csv` with the generated audio paths, sanitizes secondary labels, constructs combined label strings, and assigns stratified k-fold splits on `primary_label` with special handling to keep very rare birds present in training folds.\n- **Input construction**: Each training sample is a **composite mel-spectrogram** built by mixing multiple audio samples (primary + background) with random early stopping, contrast adjustment, and noise injection. Target vectors apply label smoothing and weight background species lower than primaries.\n\n## 3. Augmentation\n- **Audio-level**: Normalization, white noise, pink noise (precomputed), random volume modulation, pitch shift, and time stretch.\n- **Spectrogram-level**: Random power (contrast), frequency shaping, CutMix and MixUp (batch-level mixing), and composite mixing of multiple bird signals to simulate overlapping soundscapes.\n- **Label smoothing**: Primary species target ~0.995; background species ~0.3; absence is near zero to stabilize learning.\n\n## 4. Model Architecture\n- Backbone: EfficientNet variant (`tf_efficientnet_b0_ns`) via `timm`, adapted to spectrogram inputs.\n- Spectrogram is fed through a light preprocessor (batchnorm, optional spec augmentation).\n- Attention-based aggregation: Two parallel attention blocks (`AttBlockV2`) with sigmoid activations produce clipwise and framewise outputs; outputs are combined (averaged) to yield final predictions.\n- Auxiliary pooling: Max + average pooling is used before linear projection to blend local and global context.\n- Custom layers/helper functions support TensorFlow-style \"same\" padding, upsampling/interpolation, and framewise alignment.\n\n## 5. Loss & Metric\n- **Loss**: Two-way focal loss (`BCEFocal2WayLoss`) combining main logits and aggregated framewise logits to mitigate class imbalance and emphasize hard examples.\n- **Metric**: Padded class-wise average precision (`padded_cmap`) with padding to stabilize on sparse labels; primary evaluation uses `cmap_pad_5`.\n\n## 6. Training Strategy\n- **K-fold cross-validation**: Stratified on `primary_label` (default 5 folds), with explicit handling so extremely rare species remain in training.\n- **Optimizer**: `AdamW` from `transformers` with weight decay.\n- **Scheduler**: Flexible scheduler selection (e.g., ReduceLROnPlateau, cosine annealing) controlled via config.\n- **Mixed precision**: Enabled via PyTorch native AMP (`torch.cuda.amp`) for speed and stability.\n- **Checkpointing**: Best models saved based on validation loss with descriptive filenames; resume capability via latest checkpoint lookup.\n\n## 7. Implementation Details\n- Centralized config (`configs.py`) for experiment parameters (learning rate, batch size, scheduler, audio/spec parameters, etc.).\n- Logging verbosity toggled via `config.debug`.\n- Dataset builds on-the-fly composite spectrograms mixing 2–3 examples per training instance to diversify signal combinations.\n- Rare bird names are forced out of validation through fold adjustment to ensure gradient signal.\n\n## 8. Inference\n- Use trained model to compute clipwise outputs and optionally aggregate framewise predictions by upsampling/interpolation to match input length.\n- Final predictions can be thresholded or used directly in ranking for leaderboard submission.\n- A dedicated inference notebook is available: [BirdCLEF 2023 Inference ver 3.0 on Kaggle](https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0).\n\n## 9. Key Hyperparameters (example)\n- Backbone: `tf_efficientnet_b0_ns`\n- Input: Mel-spectrogram with `n_mels=224`, sample rate 32000, period 5s segments\n- Batch size: 32 (train & valid)\n- Initial LR: 1e-3 with weight decay 1e-6\n- Scheduler: ReduceLROnPlateau with factor=0.9, patience=5 (configurable)\n- Label smoothing: Enabled\n- Mix composition: Up to 3 audio clips per training sample\n\n## 10. Results Tracking\nTraining histories (loss and cmap scores) saved per fold. Visualization provided via `plot_loss_metric.py`.\n\n## 11. Code & Notebooks\n- Training code & full pipeline: [GitHub - tjamali/kaggle-birdclef-2023](https://github.com/tjamali/kaggle-birdclef-2023/tree/main)  \n- Inference notebook: [Kaggle Notebook - BirdCLEF 2023 Inference ver 3.0](https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0)\n\n## 12. Lessons & Next Steps\n- Composite mixing plus attention helps the model disambiguate overlapping bird calls.\n- Label smoothing and dual focal losses stabilize training on rare classes.\n- Future work: strong ensembling across folds/backbones, test-time augmentation, and consolidating framewise outputs for temporal localization.\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3262997": "# BirdCLEF 2023 — Sound Event Detection with Attention on Mel-Spectrograms\n\n## 1. Problem\nIdentify bird species from environmental audio recordings. The challenge is multi-label, highly imbalanced, and requires both clipwise and framewise understanding. The evaluation uses padded class-wise average precision (cmap), which favors robust ranking across rare and common species.\n\n## 2. Data & Preprocessing\n- **Raw audio**: `.ogg` files organized by primary label are loaded, converted to mono, resampled to 32kHz, and truncated to fixed-length segments (default 30s). This is handled in `prepare_audios.py`, producing `.npy` arrays under `train_np/<primary_label>/...`.\n- **Metadata merge**: `split_data.py` merges `train_metadata.csv` with the generated audio paths, sanitizes secondary labels, constructs combined label strings, and assigns stratified k-fold splits on `primary_label` with special handling to keep very rare birds present in training folds.\n- **Input construction**: Each training sample is a **composite mel-spectrogram** built by mixing multiple audio samples (primary + background) with random early stopping, contrast adjustment, and noise injection. Target vectors apply label smoothing and weight background species lower than primaries.\n\n## 3. Augmentation\n- **Audio-level**: Normalization, white noise, pink noise (precomputed), random volume modulation, pitch shift, and time stretch.\n- **Spectrogram-level**: Random power (contrast), frequency shaping, CutMix and MixUp (batch-level mixing), and composite mixing of multiple bird signals to simulate overlapping soundscapes.\n- **Label smoothing**: Primary species target ~0.995; background species ~0.3; absence is near zero to stabilize learning.\n\n## 4. Model Architecture\n- Backbone: EfficientNet variant (`tf_efficientnet_b0_ns`) via `timm`, adapted to spectrogram inputs.\n- Spectrogram is fed through a light preprocessor (batchnorm, optional spec augmentation).\n- Attention-based aggregation: Two parallel attention blocks (`AttBlockV2`) with sigmoid activations produce clipwise and framewise outputs; outputs are combined (averaged) to yield final predictions.\n- Auxiliary pooling: Max + average pooling is used before linear projection to blend local and global context.\n- Custom layers/helper functions support TensorFlow-style \"same\" padding, upsampling/interpolation, and framewise alignment.\n\n## 5. Loss & Metric\n- **Loss**: Two-way focal loss (`BCEFocal2WayLoss`) combining main logits and aggregated framewise logits to mitigate class imbalance and emphasize hard examples.\n- **Metric**: Padded class-wise average precision (`padded_cmap`) with padding to stabilize on sparse labels; primary evaluation uses `cmap_pad_5`.\n\n## 6. Training Strategy\n- **K-fold cross-validation**: Stratified on `primary_label` (default 5 folds), with explicit handling so extremely rare species remain in training.\n- **Optimizer**: `AdamW` from `transformers` with weight decay.\n- **Scheduler**: Flexible scheduler selection (e.g., ReduceLROnPlateau, cosine annealing) controlled via config.\n- **Mixed precision**: Enabled via PyTorch native AMP (`torch.cuda.amp`) for speed and stability.\n- **Checkpointing**: Best models saved based on validation loss with descriptive filenames; resume capability via latest checkpoint lookup.\n\n## 7. Implementation Details\n- Centralized config (`configs.py`) for experiment parameters (learning rate, batch size, scheduler, audio/spec parameters, etc.).\n- Logging verbosity toggled via `config.debug`.\n- Dataset builds on-the-fly composite spectrograms mixing 2–3 examples per training instance to diversify signal combinations.\n- Rare bird names are forced out of validation through fold adjustment to ensure gradient signal.\n\n## 8. Inference\n- Use trained model to compute clipwise outputs and optionally aggregate framewise predictions by upsampling/interpolation to match input length.\n- Final predictions can be thresholded or used directly in ranking for leaderboard submission.\n- A dedicated inference notebook is available: [BirdCLEF 2023 Inference ver 3.0 on Kaggle](https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0).\n\n## 9. Key Hyperparameters (example)\n- Backbone: `tf_efficientnet_b0_ns`\n- Input: Mel-spectrogram with `n_mels=224`, sample rate 32000, period 5s segments\n- Batch size: 32 (train & valid)\n- Initial LR: 1e-3 with weight decay 1e-6\n- Scheduler: ReduceLROnPlateau with factor=0.9, patience=5 (configurable)\n- Label smoothing: Enabled\n- Mix composition: Up to 3 audio clips per training sample\n\n## 10. Results Tracking\nTraining histories (loss and cmap scores) saved per fold. Visualization provided via `plot_loss_metric.py`.\n\n## 11. Code & Notebooks\n- Training code & full pipeline: [GitHub - tjamali/kaggle-birdclef-2023](https://github.com/tjamali/kaggle-birdclef-2023/tree/main)  \n- Inference notebook: [Kaggle Notebook - BirdCLEF 2023 Inference ver 3.0](https://www.kaggle.com/code/tjamali/birdclef-2023-inference-ver-3-0)\n\n## 12. Lessons & Next Steps\n- Composite mixing plus attention helps the model disambiguate overlapping bird calls.\n- Label smoothing and dual focal losses stabilize training on rare classes.\n- Future work: strong ensembling across folds/backbones, test-time augmentation, and consolidating framewise outputs for temporal localization.\n"
  }
}