{
  "id": 583377,
  "title": "28th Place Solution: SED with Segment-Based Voice Removal & Progressive Pseudo-Label Training",
  "url": "/competitions/birdclef-2025/writeups/we-are-birds-28th-place-solution-sed-with-segment-",
  "author_name": "",
  "post_date": "2025-06-06T12:13:51.780Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First, I would like to thank the competition organizers for hosting this challenging and educational competition, and congratulations to all participants for their outstanding work! Special thanks to my teammate <a href=\"https://www.kaggle.com/ziyi777\" target=\"_blank\">@ziyi777</a> for the collaborative effort throughout this competition!</p>\n<h2>Solution Overview</h2>\n<p>Our final solution employs <strong>Sound Event Detection (SED) models</strong> with 5 different backbones trained through nearly identical pipelines, enhanced by two rounds of progressive pseudo-label fusion training and sophisticated ensemble techniques.</p>\n<h2>Model Performance</h2>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Public</th>\n<th>Private</th>\n<th>Weight in Ensemble</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv2_b3</td>\n<td>0.878</td>\n<td><strong>0.893</strong></td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>efficientnet_b0</td>\n<td>0.868</td>\n<td><strong>0.896</strong></td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>efficientnetv2_s</td>\n<td>0.875</td>\n<td><strong>0.899</strong></td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>seresnext26t_32x4d</td>\n<td>0.885</td>\n<td><strong>0.898</strong></td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>eca_nfnet_l0</td>\n<td>0.866</td>\n<td><strong>0.886</strong></td>\n<td>0.1</td>\n</tr>\n<tr>\n<td><strong>Weighted Ensemble</strong></td>\n<td><strong>0.893</strong></td>\n<td><strong>0.909</strong></td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<p>What's particularly encouraging is that our SED single models demonstrated remarkably consistent performance between public and private leaderboards, which we believe reflects the great generalization capability of our data processing and model training pipeline.</p>\n<h2>Key Technical Innovations</h2>\n<h3>1. Segment-Based Voice Removal Processing</h3>\n<p>Voice processing (especially for CSA files) was a challenging aspect of this competition. Through extensive experimentation, we developed a method that provides consistent improvements in single model Public LB performance while enhancing stability.</p>\n<p>This algorithm is designed based on a key insight: <strong>audio files are artificially concatenated, composed of multiple segments (bird call segments + human voice commentary segments), with silent intervals as transitions between segments</strong>. The algorithm utilizes pre-computed voice detection results (Thanks to your excellent work! <a href=\"https://www.kaggle.com/kdmitrie\" target=\"_blank\">@kdmitrie</a>) to identify human voice regions in the audio files.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7867036%2Faaeb9085d5b929c8b8d4c476b0f7a71a%2F_20250606124514.png?generation=1749211822801116&amp;alt=media\" alt=\"\"></p>\n<h4>Algorithm Steps:</h4>\n<p><strong>Silence-Based Segmentation</strong></p>\n<ul>\n<li>Detect silent intervals (amplitude &lt; 0.001, duration &gt; 0.3s) that serve as boundaries between concatenated segments</li>\n<li>Split the audio into independent segments based on these silence boundaries</li>\n</ul>\n<p><strong>Voice Detection</strong></p>\n<ul>\n<li>Analyze each segment using pre-computed voice detection data to identify human voice regions</li>\n<li>Calculate the overlap between each segment and known voice regions to determine voice content percentage</li>\n</ul>\n<p><strong>Segment Selection Strategy</strong></p>\n<ul>\n<li>Single segment: No processing (not artificially concatenated, no human commentary segments expected)</li>\n<li>Multiple segments: Keep only segments with the lowest voice content percentage</li>\n</ul>\n<p><strong>Clean Audio Generation</strong></p>\n<ul>\n<li>Concatenate the remaining segments to form clean audio</li>\n</ul>\n<p>Through this preprocessing, we can confidently <strong>remove redundant human commentary while preserving environmental human voices that serve as background enhancement</strong>.</p>\n<h3>2. RMS-Based Random Five-Second Sampling</h3>\n<p>After the voice removal process generates clean audio, this algorithm ensures we extract the highest quality 5-second segment for training. We randomly sample 3 five-second segments from the cleaned audio, calculate the RMS (Root Mean Square) energy for each segment, and select the segment with the highest energy.</p>\n<p>This method maintains randomness in audio selection while increasing the probability of bird calls appearing within the random 5-second segments, which is crucial for the 5-second inference window used in evaluation.</p>\n<h4>Sampling Strategy Effectiveness (in order of performance):</h4>\n<ol>\n<li><strong>RMS-based Selection</strong> (Our Choice): Highest performance by selecting segments with maximum energy</li>\n<li><strong>Random 5-second Sampling</strong>: Standard random selection baseline  </li>\n<li><strong>First 5-second Sampling</strong>: Simple but suboptimal approach</li>\n</ol>\n<p>We focused our efforts on perfecting the 5-second window approach rather than exploring 10-second training windows due to the deadline, which needs more experiments in the future.</p>\n<h3>3. Progressive Pseudo-Label Fusion</h3>\n<p>We adopted and significantly improved the pseudo-labeling strategy from last year's 2nd place solution: randomly selected 5s audio segments from unlabeled test sets are added to training samples with a dynamic probability. Before mixing two audio signals, both waveforms' amplitudes are multiplied by random factors. The training sample's target vector (1.0 at primary and secondary species positions, zero elsewhere) is combined with pseudo-labels (vectors with prediction probabilities) by taking the maximum of both to form new target vectors.</p>\n<h4>Our Key Innovation: Progressive Pseudo-Label Mixup</h4>\n<p>Initially, we found that using fixed probabilities (35%, 45%) didn't improve new model performance and actually made it more unstable. We hypothesized that fixed-probability pseudo-label mixing might hinder the model's generalization to the soundscape domain.</p>\n<p><strong>Our Intuition</strong>: During early training stages, the model needs \"harder\" ground truth to learn fundamental knowledge, while in later stages, \"softer\" guidance can help the model gradually generalize to the soundscape space.</p>\n<p>Based on this intuition, we chose to <strong>linearly increase the mixing probability with epochs</strong>, ultimately finding that 0.2-0.5 is an optimal range. This progressive approach allows the model to:</p>\n<ul>\n<li>Focus on learning from high-quality ground truth early on</li>\n<li>Gradually adapt to the distribution of unlabeled soundscape data</li>\n<li>Achieve better generalization without compromising fundamental learning</li>\n</ul>\n<h4>Critical Training Optimizations for Pseudo-Labeling:</h4>\n<ul>\n<li><strong>Larger Batch Size</strong>: We increased batch size to <strong>128</strong> during pseudo-label fusion training, along with proportionally scaling the learning rate. This larger batch size provides more stable gradient estimates when mixing real and pseudo-labeled data.</li>\n<li><strong>Backbone-Specific Training Epochs</strong>: Through extensive experimentation, we determined the optimal training epochs for each backbone individually, then trained on full datasets before final LB validation. This careful epoch selection prevented both underfitting and overfitting (we're grateful this approach wasn't affected by leaderboard shake-up!).</li>\n</ul>\n<h3>4. Post-Processing</h3>\n<h4>Temporal Smoothing:</h4>\n<p>We use temporal smoothing for the 1-minute soundscape predictions using carefully tuned weights:</p>\n<pre><code>\nnew_pred[i] =  * pred[i] +  * pred[i-] +  * pred[i+]\n\n\nnew_pred[] =  * pred[] +  * pred[]\nnew_pred[-] =  * pred[-] +  * pred[-]\n</code></pre>\n<p>This 0.2-0.6-0.2 weighting scheme was chosen after extensive experimentation and provides the optimal balance between temporal consistency and segment independence.</p>\n<h2>Conclusion</h2>\n<p>Our solution demonstrates that <strong>progressive pseudo-labeling combined with sophisticated audio preprocessing and ensemble techniques</strong> can achieve strong performance in soundscape-based bird species identification. Key insights include the importance of voice removal preprocessing, RMS-based sampling strategies, and careful hyperparameter optimization for pseudo-label training.</p>\n<p>Since we struggled with CNN models and many irrelevant details early in the competition, only finding the right direction for iterative improvement in the last month, we believe we could have achieved even better results with more time.</p>\n<p>We hope our insights can be helpful to the community. Thank you again to all participants and organizers for making this such a rewarding learning experience!</p>",
  "messages": [
    {
      "id": "3218598",
      "postDate": "06/06/2025 12:12:41",
      "content": "<p>First, I would like to thank the competition organizers for hosting this challenging and educational competition, and congratulations to all participants for their outstanding work! Special thanks to my teammate <a href=\"https://www.kaggle.com/ziyi777\" target=\"_blank\">@ziyi777</a> for the collaborative effort throughout this competition!</p>\n<h2>Solution Overview</h2>\n<p>Our final solution employs <strong>Sound Event Detection (SED) models</strong> with 5 different backbones trained through nearly identical pipelines, enhanced by two rounds of progressive pseudo-label fusion training and sophisticated ensemble techniques.</p>\n<h2>Model Performance</h2>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Public</th>\n<th>Private</th>\n<th>Weight in Ensemble</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv2_b3</td>\n<td>0.878</td>\n<td><strong>0.893</strong></td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>efficientnet_b0</td>\n<td>0.868</td>\n<td><strong>0.896</strong></td>\n<td>0.1</td>\n</tr>\n<tr>\n<td>efficientnetv2_s</td>\n<td>0.875</td>\n<td><strong>0.899</strong></td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>seresnext26t_32x4d</td>\n<td>0.885</td>\n<td><strong>0.898</strong></td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>eca_nfnet_l0</td>\n<td>0.866</td>\n<td><strong>0.886</strong></td>\n<td>0.1</td>\n</tr>\n<tr>\n<td><strong>Weighted Ensemble</strong></td>\n<td><strong>0.893</strong></td>\n<td><strong>0.909</strong></td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<p>What's particularly encouraging is that our SED single models demonstrated remarkably consistent performance between public and private leaderboards, which we believe reflects the great generalization capability of our data processing and model training pipeline.</p>\n<h2>Key Technical Innovations</h2>\n<h3>1. Segment-Based Voice Removal Processing</h3>\n<p>Voice processing (especially for CSA files) was a challenging aspect of this competition. Through extensive experimentation, we developed a method that provides consistent improvements in single model Public LB performance while enhancing stability.</p>\n<p>This algorithm is designed based on a key insight: <strong>audio files are artificially concatenated, composed of multiple segments (bird call segments + human voice commentary segments), with silent intervals as transitions between segments</strong>. The algorithm utilizes pre-computed voice detection results (Thanks to your excellent work! <a href=\"https://www.kaggle.com/kdmitrie\" target=\"_blank\">@kdmitrie</a>) to identify human voice regions in the audio files.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7867036%2Faaeb9085d5b929c8b8d4c476b0f7a71a%2F_20250606124514.png?generation=1749211822801116&amp;alt=media\" alt=\"\"></p>\n<h4>Algorithm Steps:</h4>\n<p><strong>Silence-Based Segmentation</strong></p>\n<ul>\n<li>Detect silent intervals (amplitude &lt; 0.001, duration &gt; 0.3s) that serve as boundaries between concatenated segments</li>\n<li>Split the audio into independent segments based on these silence boundaries</li>\n</ul>\n<p><strong>Voice Detection</strong></p>\n<ul>\n<li>Analyze each segment using pre-computed voice detection data to identify human voice regions</li>\n<li>Calculate the overlap between each segment and known voice regions to determine voice content percentage</li>\n</ul>\n<p><strong>Segment Selection Strategy</strong></p>\n<ul>\n<li>Single segment: No processing (not artificially concatenated, no human commentary segments expected)</li>\n<li>Multiple segments: Keep only segments with the lowest voice content percentage</li>\n</ul>\n<p><strong>Clean Audio Generation</strong></p>\n<ul>\n<li>Concatenate the remaining segments to form clean audio</li>\n</ul>\n<p>Through this preprocessing, we can confidently <strong>remove redundant human commentary while preserving environmental human voices that serve as background enhancement</strong>.</p>\n<h3>2. RMS-Based Random Five-Second Sampling</h3>\n<p>After the voice removal process generates clean audio, this algorithm ensures we extract the highest quality 5-second segment for training. We randomly sample 3 five-second segments from the cleaned audio, calculate the RMS (Root Mean Square) energy for each segment, and select the segment with the highest energy.</p>\n<p>This method maintains randomness in audio selection while increasing the probability of bird calls appearing within the random 5-second segments, which is crucial for the 5-second inference window used in evaluation.</p>\n<h4>Sampling Strategy Effectiveness (in order of performance):</h4>\n<ol>\n<li><strong>RMS-based Selection</strong> (Our Choice): Highest performance by selecting segments with maximum energy</li>\n<li><strong>Random 5-second Sampling</strong>: Standard random selection baseline  </li>\n<li><strong>First 5-second Sampling</strong>: Simple but suboptimal approach</li>\n</ol>\n<p>We focused our efforts on perfecting the 5-second window approach rather than exploring 10-second training windows due to the deadline, which needs more experiments in the future.</p>\n<h3>3. Progressive Pseudo-Label Fusion</h3>\n<p>We adopted and significantly improved the pseudo-labeling strategy from last year's 2nd place solution: randomly selected 5s audio segments from unlabeled test sets are added to training samples with a dynamic probability. Before mixing two audio signals, both waveforms' amplitudes are multiplied by random factors. The training sample's target vector (1.0 at primary and secondary species positions, zero elsewhere) is combined with pseudo-labels (vectors with prediction probabilities) by taking the maximum of both to form new target vectors.</p>\n<h4>Our Key Innovation: Progressive Pseudo-Label Mixup</h4>\n<p>Initially, we found that using fixed probabilities (35%, 45%) didn't improve new model performance and actually made it more unstable. We hypothesized that fixed-probability pseudo-label mixing might hinder the model's generalization to the soundscape domain.</p>\n<p><strong>Our Intuition</strong>: During early training stages, the model needs \"harder\" ground truth to learn fundamental knowledge, while in later stages, \"softer\" guidance can help the model gradually generalize to the soundscape space.</p>\n<p>Based on this intuition, we chose to <strong>linearly increase the mixing probability with epochs</strong>, ultimately finding that 0.2-0.5 is an optimal range. This progressive approach allows the model to:</p>\n<ul>\n<li>Focus on learning from high-quality ground truth early on</li>\n<li>Gradually adapt to the distribution of unlabeled soundscape data</li>\n<li>Achieve better generalization without compromising fundamental learning</li>\n</ul>\n<h4>Critical Training Optimizations for Pseudo-Labeling:</h4>\n<ul>\n<li><strong>Larger Batch Size</strong>: We increased batch size to <strong>128</strong> during pseudo-label fusion training, along with proportionally scaling the learning rate. This larger batch size provides more stable gradient estimates when mixing real and pseudo-labeled data.</li>\n<li><strong>Backbone-Specific Training Epochs</strong>: Through extensive experimentation, we determined the optimal training epochs for each backbone individually, then trained on full datasets before final LB validation. This careful epoch selection prevented both underfitting and overfitting (we're grateful this approach wasn't affected by leaderboard shake-up!).</li>\n</ul>\n<h3>4. Post-Processing</h3>\n<h4>Temporal Smoothing:</h4>\n<p>We use temporal smoothing for the 1-minute soundscape predictions using carefully tuned weights:</p>\n<pre><code>\nnew_pred[i] =  * pred[i] +  * pred[i-] +  * pred[i+]\n\n\nnew_pred[] =  * pred[] +  * pred[]\nnew_pred[-] =  * pred[-] +  * pred[-]\n</code></pre>\n<p>This 0.2-0.6-0.2 weighting scheme was chosen after extensive experimentation and provides the optimal balance between temporal consistency and segment independence.</p>\n<h2>Conclusion</h2>\n<p>Our solution demonstrates that <strong>progressive pseudo-labeling combined with sophisticated audio preprocessing and ensemble techniques</strong> can achieve strong performance in soundscape-based bird species identification. Key insights include the importance of voice removal preprocessing, RMS-based sampling strategies, and careful hyperparameter optimization for pseudo-label training.</p>\n<p>Since we struggled with CNN models and many irrelevant details early in the competition, only finding the right direction for iterative improvement in the last month, we believe we could have achieved even better results with more time.</p>\n<p>We hope our insights can be helpful to the community. Thank you again to all participants and organizers for making this such a rewarding learning experience!</p>",
      "rawMarkdown": "First, I would like to thank the competition organizers for hosting this challenging and educational competition, and congratulations to all participants for their outstanding work! Special thanks to my teammate @ziyi777 for the collaborative effort throughout this competition!\n\n## Solution Overview\n\nOur final solution employs **Sound Event Detection (SED) models** with 5 different backbones trained through nearly identical pipelines, enhanced by two rounds of progressive pseudo-label fusion training and sophisticated ensemble techniques.\n\n## Model Performance\n\n|Backbone| Public | Private| Weight in Ensemble |\n| --- | --- | --- | --- |\n|efficientnetv2_b3| 0.878 | **0.893** | 0.2 |\n|efficientnet_b0| 0.868 | **0.896** | 0.1 |\n|efficientnetv2_s| 0.875 | **0.899** | 0.2 |\n|seresnext26t_32x4d| 0.885 | **0.898** | 0.4 |\n|eca_nfnet_l0 | 0.866 | **0.886** | 0.1 |\n| **Weighted Ensemble** | **0.893** | **0.909** | - |\n\nWhat's particularly encouraging is that our SED single models demonstrated remarkably consistent performance between public and private leaderboards, which we believe reflects the great generalization capability of our data processing and model training pipeline.\n\n## Key Technical Innovations\n\n### 1. Segment-Based Voice Removal Processing\n\nVoice processing (especially for CSA files) was a challenging aspect of this competition. Through extensive experimentation, we developed a method that provides consistent improvements in single model Public LB performance while enhancing stability.\n\nThis algorithm is designed based on a key insight: **audio files are artificially concatenated, composed of multiple segments (bird call segments + human voice commentary segments), with silent intervals as transitions between segments**. The algorithm utilizes pre-computed voice detection results (Thanks to your excellent work! @kdmitrie) to identify human voice regions in the audio files.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7867036%2Faaeb9085d5b929c8b8d4c476b0f7a71a%2F_20250606124514.png?generation=1749211822801116&alt=media)\n\n#### Algorithm Steps:\n\n**Silence-Based Segmentation**\n- Detect silent intervals (amplitude < 0.001, duration > 0.3s) that serve as boundaries between concatenated segments\n- Split the audio into independent segments based on these silence boundaries\n\n**Voice Detection**\n- Analyze each segment using pre-computed voice detection data to identify human voice regions\n- Calculate the overlap between each segment and known voice regions to determine voice content percentage\n\n**Segment Selection Strategy**\n- Single segment: No processing (not artificially concatenated, no human commentary segments expected)\n- Multiple segments: Keep only segments with the lowest voice content percentage\n\n**Clean Audio Generation**\n- Concatenate the remaining segments to form clean audio\n\nThrough this preprocessing, we can confidently **remove redundant human commentary while preserving environmental human voices that serve as background enhancement**.\n\n### 2. RMS-Based Random Five-Second Sampling\n\nAfter the voice removal process generates clean audio, this algorithm ensures we extract the highest quality 5-second segment for training. We randomly sample 3 five-second segments from the cleaned audio, calculate the RMS (Root Mean Square) energy for each segment, and select the segment with the highest energy.\n\nThis method maintains randomness in audio selection while increasing the probability of bird calls appearing within the random 5-second segments, which is crucial for the 5-second inference window used in evaluation.\n\n#### Sampling Strategy Effectiveness (in order of performance):\n1. **RMS-based Selection** (Our Choice): Highest performance by selecting segments with maximum energy\n2. **Random 5-second Sampling**: Standard random selection baseline  \n3. **First 5-second Sampling**: Simple but suboptimal approach\n\nWe focused our efforts on perfecting the 5-second window approach rather than exploring 10-second training windows due to the deadline, which needs more experiments in the future.\n\n### 3. Progressive Pseudo-Label Fusion\n\nWe adopted and significantly improved the pseudo-labeling strategy from last year's 2nd place solution: randomly selected 5s audio segments from unlabeled test sets are added to training samples with a dynamic probability. Before mixing two audio signals, both waveforms' amplitudes are multiplied by random factors. The training sample's target vector (1.0 at primary and secondary species positions, zero elsewhere) is combined with pseudo-labels (vectors with prediction probabilities) by taking the maximum of both to form new target vectors.\n\n#### Our Key Innovation: Progressive Pseudo-Label Mixup\nInitially, we found that using fixed probabilities (35%, 45%) didn't improve new model performance and actually made it more unstable. We hypothesized that fixed-probability pseudo-label mixing might hinder the model's generalization to the soundscape domain.\n\n**Our Intuition**: During early training stages, the model needs \"harder\" ground truth to learn fundamental knowledge, while in later stages, \"softer\" guidance can help the model gradually generalize to the soundscape space.\n\nBased on this intuition, we chose to **linearly increase the mixing probability with epochs**, ultimately finding that 0.2-0.5 is an optimal range. This progressive approach allows the model to:\n- Focus on learning from high-quality ground truth early on\n- Gradually adapt to the distribution of unlabeled soundscape data\n- Achieve better generalization without compromising fundamental learning\n\n#### Critical Training Optimizations for Pseudo-Labeling:\n- **Larger Batch Size**: We increased batch size to **128** during pseudo-label fusion training, along with proportionally scaling the learning rate. This larger batch size provides more stable gradient estimates when mixing real and pseudo-labeled data.\n- **Backbone-Specific Training Epochs**: Through extensive experimentation, we determined the optimal training epochs for each backbone individually, then trained on full datasets before final LB validation. This careful epoch selection prevented both underfitting and overfitting (we're grateful this approach wasn't affected by leaderboard shake-up!).\n\n### 4. Post-Processing\n\n#### Temporal Smoothing:\nWe use temporal smoothing for the 1-minute soundscape predictions using carefully tuned weights:\n```python\n# For middle segments (i ∈ [1, N-2]) - our optimized 0.2-0.6-0.2 window:\nnew_pred[i] = 0.6 * pred[i] + 0.2 * pred[i-1] + 0.2 * pred[i+1]\n\n# For boundary segments:\nnew_pred[0] = 0.8 * pred[0] + 0.2 * pred[1]\nnew_pred[-1] = 0.8 * pred[-1] + 0.2 * pred[-2]\n```\n\nThis 0.2-0.6-0.2 weighting scheme was chosen after extensive experimentation and provides the optimal balance between temporal consistency and segment independence.\n\n\n\n## Conclusion\n\nOur solution demonstrates that **progressive pseudo-labeling combined with sophisticated audio preprocessing and ensemble techniques** can achieve strong performance in soundscape-based bird species identification. Key insights include the importance of voice removal preprocessing, RMS-based sampling strategies, and careful hyperparameter optimization for pseudo-label training.\n\nSince we struggled with CNN models and many irrelevant details early in the competition, only finding the right direction for iterative improvement in the last month, we believe we could have achieved even better results with more time.\n\nWe hope our insights can be helpful to the community. Thank you again to all participants and organizers for making this such a rewarding learning experience!",
      "votes": null
    },
    {
      "id": "3218617",
      "postDate": "06/06/2025 12:45:53",
      "content": "<p>Congrats to your achievements，Really nice work. Learned a lot from that.</p>",
      "rawMarkdown": "Congrats to your achievements，Really nice work. Learned a lot from that.",
      "votes": null
    },
    {
      "id": "3218925",
      "postDate": "06/06/2025 23:34:23",
      "content": "<p>Thanks for sharing your 28th-place solution! The segment-based voice removal and progressive pseudo-labeling are fascinating.</p>\n<p>Could you elaborate on the rationale behind selecting the 0.2-0.6-0.2 temporal smoothing window? What specific issues did it help address compared to other weighting schemes?</p>",
      "rawMarkdown": "Thanks for sharing your 28th-place solution! The segment-based voice removal and progressive pseudo-labeling are fascinating.\n\nCould you elaborate on the rationale behind selecting the 0.2-0.6-0.2 temporal smoothing window? What specific issues did it help address compared to other weighting schemes?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218617,
      "author_name": "yanglema",
      "author_url": "",
      "post_date": "06/06/2025 12:45:53",
      "content": "<p>Congrats to your achievements，Really nice work. Learned a lot from that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3218925,
      "author_name": "tyyuki",
      "author_url": "",
      "post_date": "06/06/2025 23:34:23",
      "content": "<p>Thanks for sharing your 28th-place solution! The segment-based voice removal and progressive pseudo-labeling are fascinating.</p>\n<p>Could you elaborate on the rationale behind selecting the 0.2-0.6-0.2 temporal smoothing window? What specific issues did it help address compared to other weighting schemes?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218598": "First, I would like to thank the competition organizers for hosting this challenging and educational competition, and congratulations to all participants for their outstanding work! Special thanks to my teammate @ziyi777 for the collaborative effort throughout this competition!\n\n## Solution Overview\n\nOur final solution employs **Sound Event Detection (SED) models** with 5 different backbones trained through nearly identical pipelines, enhanced by two rounds of progressive pseudo-label fusion training and sophisticated ensemble techniques.\n\n## Model Performance\n\n|Backbone| Public | Private| Weight in Ensemble |\n| --- | --- | --- | --- |\n|efficientnetv2_b3| 0.878 | **0.893** | 0.2 |\n|efficientnet_b0| 0.868 | **0.896** | 0.1 |\n|efficientnetv2_s| 0.875 | **0.899** | 0.2 |\n|seresnext26t_32x4d| 0.885 | **0.898** | 0.4 |\n|eca_nfnet_l0 | 0.866 | **0.886** | 0.1 |\n| **Weighted Ensemble** | **0.893** | **0.909** | - |\n\nWhat's particularly encouraging is that our SED single models demonstrated remarkably consistent performance between public and private leaderboards, which we believe reflects the great generalization capability of our data processing and model training pipeline.\n\n## Key Technical Innovations\n\n### 1. Segment-Based Voice Removal Processing\n\nVoice processing (especially for CSA files) was a challenging aspect of this competition. Through extensive experimentation, we developed a method that provides consistent improvements in single model Public LB performance while enhancing stability.\n\nThis algorithm is designed based on a key insight: **audio files are artificially concatenated, composed of multiple segments (bird call segments + human voice commentary segments), with silent intervals as transitions between segments**. The algorithm utilizes pre-computed voice detection results (Thanks to your excellent work! @kdmitrie) to identify human voice regions in the audio files.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7867036%2Faaeb9085d5b929c8b8d4c476b0f7a71a%2F_20250606124514.png?generation=1749211822801116&alt=media)\n\n#### Algorithm Steps:\n\n**Silence-Based Segmentation**\n- Detect silent intervals (amplitude < 0.001, duration > 0.3s) that serve as boundaries between concatenated segments\n- Split the audio into independent segments based on these silence boundaries\n\n**Voice Detection**\n- Analyze each segment using pre-computed voice detection data to identify human voice regions\n- Calculate the overlap between each segment and known voice regions to determine voice content percentage\n\n**Segment Selection Strategy**\n- Single segment: No processing (not artificially concatenated, no human commentary segments expected)\n- Multiple segments: Keep only segments with the lowest voice content percentage\n\n**Clean Audio Generation**\n- Concatenate the remaining segments to form clean audio\n\nThrough this preprocessing, we can confidently **remove redundant human commentary while preserving environmental human voices that serve as background enhancement**.\n\n### 2. RMS-Based Random Five-Second Sampling\n\nAfter the voice removal process generates clean audio, this algorithm ensures we extract the highest quality 5-second segment for training. We randomly sample 3 five-second segments from the cleaned audio, calculate the RMS (Root Mean Square) energy for each segment, and select the segment with the highest energy.\n\nThis method maintains randomness in audio selection while increasing the probability of bird calls appearing within the random 5-second segments, which is crucial for the 5-second inference window used in evaluation.\n\n#### Sampling Strategy Effectiveness (in order of performance):\n1. **RMS-based Selection** (Our Choice): Highest performance by selecting segments with maximum energy\n2. **Random 5-second Sampling**: Standard random selection baseline  \n3. **First 5-second Sampling**: Simple but suboptimal approach\n\nWe focused our efforts on perfecting the 5-second window approach rather than exploring 10-second training windows due to the deadline, which needs more experiments in the future.\n\n### 3. Progressive Pseudo-Label Fusion\n\nWe adopted and significantly improved the pseudo-labeling strategy from last year's 2nd place solution: randomly selected 5s audio segments from unlabeled test sets are added to training samples with a dynamic probability. Before mixing two audio signals, both waveforms' amplitudes are multiplied by random factors. The training sample's target vector (1.0 at primary and secondary species positions, zero elsewhere) is combined with pseudo-labels (vectors with prediction probabilities) by taking the maximum of both to form new target vectors.\n\n#### Our Key Innovation: Progressive Pseudo-Label Mixup\nInitially, we found that using fixed probabilities (35%, 45%) didn't improve new model performance and actually made it more unstable. We hypothesized that fixed-probability pseudo-label mixing might hinder the model's generalization to the soundscape domain.\n\n**Our Intuition**: During early training stages, the model needs \"harder\" ground truth to learn fundamental knowledge, while in later stages, \"softer\" guidance can help the model gradually generalize to the soundscape space.\n\nBased on this intuition, we chose to **linearly increase the mixing probability with epochs**, ultimately finding that 0.2-0.5 is an optimal range. This progressive approach allows the model to:\n- Focus on learning from high-quality ground truth early on\n- Gradually adapt to the distribution of unlabeled soundscape data\n- Achieve better generalization without compromising fundamental learning\n\n#### Critical Training Optimizations for Pseudo-Labeling:\n- **Larger Batch Size**: We increased batch size to **128** during pseudo-label fusion training, along with proportionally scaling the learning rate. This larger batch size provides more stable gradient estimates when mixing real and pseudo-labeled data.\n- **Backbone-Specific Training Epochs**: Through extensive experimentation, we determined the optimal training epochs for each backbone individually, then trained on full datasets before final LB validation. This careful epoch selection prevented both underfitting and overfitting (we're grateful this approach wasn't affected by leaderboard shake-up!).\n\n### 4. Post-Processing\n\n#### Temporal Smoothing:\nWe use temporal smoothing for the 1-minute soundscape predictions using carefully tuned weights:\n```python\n# For middle segments (i ∈ [1, N-2]) - our optimized 0.2-0.6-0.2 window:\nnew_pred[i] = 0.6 * pred[i] + 0.2 * pred[i-1] + 0.2 * pred[i+1]\n\n# For boundary segments:\nnew_pred[0] = 0.8 * pred[0] + 0.2 * pred[1]\nnew_pred[-1] = 0.8 * pred[-1] + 0.2 * pred[-2]\n```\n\nThis 0.2-0.6-0.2 weighting scheme was chosen after extensive experimentation and provides the optimal balance between temporal consistency and segment independence.\n\n\n\n## Conclusion\n\nOur solution demonstrates that **progressive pseudo-labeling combined with sophisticated audio preprocessing and ensemble techniques** can achieve strong performance in soundscape-based bird species identification. Key insights include the importance of voice removal preprocessing, RMS-based sampling strategies, and careful hyperparameter optimization for pseudo-label training.\n\nSince we struggled with CNN models and many irrelevant details early in the competition, only finding the right direction for iterative improvement in the last month, we believe we could have achieved even better results with more time.\n\nWe hope our insights can be helpful to the community. Thank you again to all participants and organizers for making this such a rewarding learning experience!",
    "3218617": "Congrats to your achievements，Really nice work. Learned a lot from that.",
    "3218925": "Thanks for sharing your 28th-place solution! The segment-based voice removal and progressive pseudo-labeling are fascinating.\n\nCould you elaborate on the rationale behind selecting the 0.2-0.6-0.2 temporal smoothing window? What specific issues did it help address compared to other weighting schemes?"
  },
  "source": "meta"
}