{
  "id": 511845,
  "title": "4th place solution: Team Cerberus",
  "url": "/competitions/birdclef-2024/writeups/team-cerberus-4th-place-solution-team-cerberus",
  "author_name": "",
  "post_date": "2024-06-23T08:01:05.950Z",
  "votes": 48,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thank you to Kaggle, the hosts, and all the competitors. Participating in this exciting competition has been an amazing experience. Here's a look at our 4th-place solution. This achievement was truly a team effort, with equal contributions from <a href=\"https://www.kaggle.com/ajobseeker\" target=\"_blank\">@ajobseeker</a> and <a href=\"https://www.kaggle.com/tamotamo\" target=\"_blank\">@tamotamo</a>. I'm grateful to have had the chance to work with them in this competition.</p>\n<p>Update (2024-06-23):<br>\nAdded the inference notebook and training code.<br>\nInference Notebook: <a href=\"https://www.kaggle.com/code/yokuyama/bc24-4th-place/notebook\" target=\"_blank\">https://www.kaggle.com/code/yokuyama/bc24-4th-place/notebook</a><br>\nTrain code(melspec models): <a href=\"https://github.com/yoku001/BirdCLEF2024-4th-place-solution-melspec\" target=\"_blank\">https://github.com/yoku001/BirdCLEF2024-4th-place-solution-melspec</a><br>\nTrain code(raw signam models): <a href=\"https://github.com/tamotamo17/BirdCLEF2024-4th-place-solution-raw-signal\" target=\"_blank\">https://github.com/tamotamo17/BirdCLEF2024-4th-place-solution-raw-signal</a></p>\n<h2>TL;DR</h2>\n<ul>\n<li>Ensemble of Melspec Models and Signal Models</li>\n<li>TTA</li>\n<li>OpenVINO</li>\n<li>Post Processing</li>\n</ul>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>Public Score (+TTA)</th>\n<th>Private Score (+TTA)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Melspec Model B (inception-next-nano)</td>\n<td>0.668</td>\n<td>0.623</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Melspec Model A  (rexnet_150)</td>\n<td>0.676</td>\n<td>0.641</td>\n<td>0.690</td>\n<td>0.649</td>\n</tr>\n<tr>\n<td>Melspec Model A (seresnext26ts)</td>\n<td>0.682</td>\n<td>0.645</td>\n<td>0.693</td>\n<td>0.651</td>\n</tr>\n<tr>\n<td>Raw signal Model C (tf_efficientnet_b0_ns)</td>\n<td>0.673</td>\n<td>0.620</td>\n<td>0.691</td>\n<td>0.636</td>\n</tr>\n<tr>\n<td>Weighted Mean</td>\n<td>0.717</td>\n<td>0.667</td>\n<td>0.731</td>\n<td>0.676</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean</td>\n<td></td>\n<td></td>\n<td>0.732</td>\n<td>0.677</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean  + Smoothing</td>\n<td></td>\n<td></td>\n<td>0.741</td>\n<td>0.685</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean  + Smoothing + Cut-off (Final Sub)</td>\n<td></td>\n<td></td>\n<td>0.7469</td>\n<td>0.6877</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean  + Smoothing + Cut-off + Max With Neighbors</td>\n<td></td>\n<td></td>\n<td>0.749</td>\n<td>0.689</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Weighted Mean<ul>\n<li><code>0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C</code></li></ul></li>\n<li>Geometric Mean<ul>\n<li><code>(0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C) + 0.3*(Model A (rexnet_150) * Model C)**(0.5)</code></li>\n<li>By adding the geometric mean of the MelSpec Model(Model A) and the Raw-signal Model(Model C) to the ensemble, we achieved a slight improvement in our score. Choosing this ensemble at the end allowed us to fortunately remain in the prize-winning positions.</li>\n<li>We saw a big risk of overfitting, so we decided not to spend any more time adjusting the ensemble weights.</li></ul></li>\n</ul>\n<h2>Model A: 2021-2nd Melspec CNNs</h2>\n<p>We heavily referenced the code from the <a href=\"https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution\" target=\"_blank\">2023 2nd place solution</a> to build our training and inference pipeline for this model. Big thanks to <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> for sharing such valuable information. Their codebase was incredibly strong, and with just a few modifications, we were able to create a single model with an LB score of 0.68.</p>\n<ul>\n<li><p>Dataset</p>\n<ul>\n<li>BC2024</li>\n<li>Some models pretrained on 2021, 2022, and 2023's datasets.</li>\n<li>Random 15-20 seconds from audio at training, first 5 seconds at validation.</li></ul></li>\n<li><p>Preprocessing</p>\n<ul>\n<li>n_mels=128, n_fft=2048, f_min=0, f_max=16000, hop_length=627, top_db=80. </li></ul></li>\n<li><p>Data Augmentation</p>\n<ul>\n<li>AddBackgroundNoise (<a href=\"https://www.kaggle.com/datasets/honglihang/background-noise\" target=\"_blank\">datasets</a>)</li>\n<li>Gain</li>\n<li>Noise Injection</li>\n<li>Gaussian Noise</li>\n<li>Pink Noise</li>\n<li>Mixup</li></ul></li>\n<li><p>Model</p>\n<ul>\n<li><code>seresnext26ts</code></li>\n<li><code>rexnet_150</code><ul>\n<li>2021-2023 pretrained</li></ul></li></ul></li>\n<li><p>Loss Function</p>\n<ul>\n<li>BCELoss</li>\n<li>Class sampling weights proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">1st place of 2023 competition</a>.</li></ul></li>\n</ul>\n<h2>Model B: Simple Melspec CNNs</h2>\n<ul>\n<li><p>Dataset</p>\n<ul>\n<li>BC2024 + xeno-canto-additional-cleaned(see <code>Validation Strategy</code> section)</li>\n<li>Random 5 seconds from audio for training, first 5 seconds for validation.</li></ul></li>\n<li><p>Data Augmentation</p>\n<ul>\n<li>AddBackgroundNoise (<a href=\"https://www.kaggle.com/datasets/honglihang/background-noise\" target=\"_blank\">datasets</a>)</li>\n<li>Gain</li>\n<li>Noise Injection</li>\n<li>Gaussian Noise</li>\n<li>Pink Noise</li>\n<li>Mixup</li>\n<li><a href=\"https://www.kaggle.com/c/birdclef-2023/discussion/412922\" target=\"_blank\">Sumup</a></li></ul></li>\n<li><p>Model</p>\n<ul>\n<li><p><code>inception-next-nano</code> with attention head</p>\n<ul>\n<li>InceptionNeXt with the same scaling as ConvNeXt-nano</li></ul>\n<pre><code> timm.models.inception_next  _create_inception_next\n timm.models.inception_next  InceptionDWConv2d\n timm.models._registry  register_model\n\n\n ():\n    ()\n    model_args = (\n        depths=(, , , ), dims=(, , , ),\n        token_mixers=InceptionDWConv2d,\n    )\n     _create_inception_next(, pretrained=, **(model_args, **kwargs))\n</code></pre></li></ul></li>\n</ul>\n<h3>Model C: Raw signal CNN</h3>\n<p>This model is inspired by the HMS <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492254\" target=\"_blank\">2nd place solution</a>. A big thanks to <a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>! They provided a detailed explanation of <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511535\" target=\"_blank\">them solution</a>.</p>\n<ul>\n<li><p>Dataset</p>\n<ul>\n<li>BC2024</li>\n<li>Removed duplicate data by referring to  <a href=\"https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/\" target=\"_blank\">[this link]</a>.</li>\n<li>Added several samples from the minority class using data from xeno-canto.</li>\n<li>Applied stratified 5-fold cross-validation grouped by author.</li>\n<li>Classes with fewer than 15 samples were upsampled to 15 samples during training.</li></ul></li>\n<li><p>Preprocessing</p>\n<ul>\n<li>Used the first 5 seconds of each audio sample.</li>\n<li>Downsampled the audio to half the original rate (from 32000 Hz to 16000 Hz).</li>\n<li>Reshaped the downsampled audio data from a size of 80000 to 625x128.</li></ul></li>\n<li><p>Data Augmentation</p>\n<ul>\n<li>Annotated 50 background segments from unlabeled data and added them as background noise.</li>\n<li>Gain</li>\n<li>Noise Injection</li>\n<li>Gaussian Noise</li>\n<li>Pink Noise</li>\n<li>Random Volume</li>\n<li>Mixup</li>\n<li>Cutmix</li></ul></li>\n<li><p>Model</p>\n<ul>\n<li><code>tf_efficientnet_b0_ns</code> with SED head</li></ul></li>\n<li><p>Loss Function</p>\n<ul>\n<li>focal loss</li></ul></li>\n</ul>\n<h2>Validation Strategy</h2>\n<p>Using the training data for validation didn't give us reliable results, so we switched to a synthetic data approach.</p>\n<ol>\n<li>We sampled files for 40 out of 182 classes from the xeno-canto-additional dataset and cropped the segments where the birds were vocalizing to create a clean dataset.</li>\n<li>We sampled audio files containing only background noise (without bird calls) from the unlabeled soundscape dataset.</li>\n<li>We combined the clean dataset and background noise to create a test-like dataset with time-series labels.</li>\n<li>Using this synthetic dataset, we calculated the ROC AUC score. </li>\n</ol>\n<p>Although this validation method did not perfectly correlate with the LB results, it provided more reasonable outcomes compared to using the first 5-second crop method. For more details, please refer to <a href=\"https://www.kaggle.com/code/yokuyama/quant-valid-synthetic-data/notebook\" target=\"_blank\">this notebook</a>.</p>\n<h2>TTA</h2>\n<p>To enhance the accuracy of our time series predictions, we employed techniques similar to sub-pixel super-resolution. Instead of predicting just the 5-second frames during inference, we also predicted frames shifted by 2.5 seconds. We then combined these results as a TTA. This method helped in refining the overall predictions.</p>\n<p><img src=\"https://raw.githubusercontent.com/yoku001/kaggle-static-resouces/main/img/birdclef2024/zu1.drawio.png\" alt=\"img\"></p>\n<h2>OpenVINO + INT8 Post Training Quantization</h2>\n<p>To speed up our model's inference time, we used OpenVINO. Additionally, we implemented <a href=\"https://docs.openvino.ai/2024/openvino-workflow/model-optimization-guide/quantizing-models-post-training/basic-quantization-flow.html\" target=\"_blank\">post-training quantization</a> to convert our model to INT8.</p>\n<p>For the quantization calibration dataset, we used our model's <strong>training dataset</strong>, applying augmentations like background noise addition and gain changes. We believed that these augmentations would help create a quantized model better suited to handle a wider range of test data scenarios.</p>\n<p>The results were impressive: our inference speed improved dramatically, with the quantized model running <strong>30-40%</strong> faster.</p>\n<p>When performing quantization, the selection of layers to be quantized was crucial. We observed that excluding the head layers from quantization tended to improve the model's accuracy.</p>\n<pre><code>names = [, , , , ]\nquantized_model = nncf.quantize(\n    model, calibration_dataset, =600,\n    =nncf.IgnoredScope(names=names),\n)\n</code></pre>\n<p>We began working on quantization just three days before the submission deadline, leaving us insufficient time to thoroughly verify the combination of ensemble and quantization. So, we used a quantized model in only one of our two final submissions. (trade-off: we reduced the number of TTA runs for this submission.)</p>\n<p>We ended up finding that enabling quantization did not significantly impact the public/private scores.</p>\n<h2>Post Processing</h2>\n<p>By using several tricks, we were able to improve our scores by approximately 0.01 on both the private and public leaderboards.</p>\n<ul>\n<li>Smoothing<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">Similar to the 6th place team</a>, we improved our scores by taking the moving average of adjacent segments.</li></ul></li>\n<li>Cut-off<ul>\n<li>Birds that appear once in 4 minutes of audio are more likely to reappear compared to other audio. Recognizing that the probability could be low due to noise, overlapping calls with other birds, and inference slices cut by bin boundaries, we halved the value if the model had a confidence of 0.10 or less in all 48 bins of the 4-minute audio. In other words, we halved the probability if there were no birdsong (0.10 or less) in all 48 sections, and left the number unchanged if there was birdsong at least once in any of the 48 sections.</li></ul></li>\n<li>Max With Neighbors<ul>\n<li>Select the maximum value including the previous and next two rows. For a 30-second test sample, if the inference values of a label are [0.1, 0.3, 0.5, 0.2, 0.4, 0.1], modify them to [0.5, 0.5, 0.5, 0.5, 0.4]. We conducted many tests with different datasets and found a 25% probability of score decrease, so it was not included in the final submission. </li></ul></li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>Pseudo labeling using unlabeled data.</li>\n<li>Data cleansing and hand-labeling for training data.</li>\n<li>Using novel loss functions. BCE and Focal Loss performed almost the best.</li>\n<li>Manifold Mixup, D-Mixup</li>\n<li>PCEN</li>\n<li>CWT, CQT, VQT</li>\n<li>Trainable frontends: Leaf, trainable filterbank, trainable stft, Conv1D</li>\n<li>Reparameterized model</li>\n<li>Mobilenet V4</li>\n<li>BirdNET embeddings</li>\n</ul>",
  "messages": [
    {
      "id": "2868347",
      "postDate": "06/12/2024 10:26:21",
      "content": "<p>Thank you to Kaggle, the hosts, and all the competitors. Participating in this exciting competition has been an amazing experience. Here's a look at our 4th-place solution. This achievement was truly a team effort, with equal contributions from <a href=\"https://www.kaggle.com/ajobseeker\" target=\"_blank\">@ajobseeker</a> and <a href=\"https://www.kaggle.com/tamotamo\" target=\"_blank\">@tamotamo</a>. I'm grateful to have had the chance to work with them in this competition.</p>\n<p>Update (2024-06-23):<br>\nAdded the inference notebook and training code.<br>\nInference Notebook: <a href=\"https://www.kaggle.com/code/yokuyama/bc24-4th-place/notebook\" target=\"_blank\">https://www.kaggle.com/code/yokuyama/bc24-4th-place/notebook</a><br>\nTrain code(melspec models): <a href=\"https://github.com/yoku001/BirdCLEF2024-4th-place-solution-melspec\" target=\"_blank\">https://github.com/yoku001/BirdCLEF2024-4th-place-solution-melspec</a><br>\nTrain code(raw signam models): <a href=\"https://github.com/tamotamo17/BirdCLEF2024-4th-place-solution-raw-signal\" target=\"_blank\">https://github.com/tamotamo17/BirdCLEF2024-4th-place-solution-raw-signal</a></p>\n<h2>TL;DR</h2>\n<ul>\n<li>Ensemble of Melspec Models and Signal Models</li>\n<li>TTA</li>\n<li>OpenVINO</li>\n<li>Post Processing</li>\n</ul>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>Public Score (+TTA)</th>\n<th>Private Score (+TTA)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Melspec Model B (inception-next-nano)</td>\n<td>0.668</td>\n<td>0.623</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Melspec Model A  (rexnet_150)</td>\n<td>0.676</td>\n<td>0.641</td>\n<td>0.690</td>\n<td>0.649</td>\n</tr>\n<tr>\n<td>Melspec Model A (seresnext26ts)</td>\n<td>0.682</td>\n<td>0.645</td>\n<td>0.693</td>\n<td>0.651</td>\n</tr>\n<tr>\n<td>Raw signal Model C (tf_efficientnet_b0_ns)</td>\n<td>0.673</td>\n<td>0.620</td>\n<td>0.691</td>\n<td>0.636</td>\n</tr>\n<tr>\n<td>Weighted Mean</td>\n<td>0.717</td>\n<td>0.667</td>\n<td>0.731</td>\n<td>0.676</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean</td>\n<td></td>\n<td></td>\n<td>0.732</td>\n<td>0.677</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean  + Smoothing</td>\n<td></td>\n<td></td>\n<td>0.741</td>\n<td>0.685</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean  + Smoothing + Cut-off (Final Sub)</td>\n<td></td>\n<td></td>\n<td>0.7469</td>\n<td>0.6877</td>\n</tr>\n<tr>\n<td>Weighted Mean + Geometric Mean  + Smoothing + Cut-off + Max With Neighbors</td>\n<td></td>\n<td></td>\n<td>0.749</td>\n<td>0.689</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Weighted Mean<ul>\n<li><code>0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C</code></li></ul></li>\n<li>Geometric Mean<ul>\n<li><code>(0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C) + 0.3*(Model A (rexnet_150) * Model C)**(0.5)</code></li>\n<li>By adding the geometric mean of the MelSpec Model(Model A) and the Raw-signal Model(Model C) to the ensemble, we achieved a slight improvement in our score. Choosing this ensemble at the end allowed us to fortunately remain in the prize-winning positions.</li>\n<li>We saw a big risk of overfitting, so we decided not to spend any more time adjusting the ensemble weights.</li></ul></li>\n</ul>\n<h2>Model A: 2021-2nd Melspec CNNs</h2>\n<p>We heavily referenced the code from the <a href=\"https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution\" target=\"_blank\">2023 2nd place solution</a> to build our training and inference pipeline for this model. Big thanks to <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> for sharing such valuable information. Their codebase was incredibly strong, and with just a few modifications, we were able to create a single model with an LB score of 0.68.</p>\n<ul>\n<li><p>Dataset</p>\n<ul>\n<li>BC2024</li>\n<li>Some models pretrained on 2021, 2022, and 2023's datasets.</li>\n<li>Random 15-20 seconds from audio at training, first 5 seconds at validation.</li></ul></li>\n<li><p>Preprocessing</p>\n<ul>\n<li>n_mels=128, n_fft=2048, f_min=0, f_max=16000, hop_length=627, top_db=80. </li></ul></li>\n<li><p>Data Augmentation</p>\n<ul>\n<li>AddBackgroundNoise (<a href=\"https://www.kaggle.com/datasets/honglihang/background-noise\" target=\"_blank\">datasets</a>)</li>\n<li>Gain</li>\n<li>Noise Injection</li>\n<li>Gaussian Noise</li>\n<li>Pink Noise</li>\n<li>Mixup</li></ul></li>\n<li><p>Model</p>\n<ul>\n<li><code>seresnext26ts</code></li>\n<li><code>rexnet_150</code><ul>\n<li>2021-2023 pretrained</li></ul></li></ul></li>\n<li><p>Loss Function</p>\n<ul>\n<li>BCELoss</li>\n<li>Class sampling weights proposed by <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412808\" target=\"_blank\">1st place of 2023 competition</a>.</li></ul></li>\n</ul>\n<h2>Model B: Simple Melspec CNNs</h2>\n<ul>\n<li><p>Dataset</p>\n<ul>\n<li>BC2024 + xeno-canto-additional-cleaned(see <code>Validation Strategy</code> section)</li>\n<li>Random 5 seconds from audio for training, first 5 seconds for validation.</li></ul></li>\n<li><p>Data Augmentation</p>\n<ul>\n<li>AddBackgroundNoise (<a href=\"https://www.kaggle.com/datasets/honglihang/background-noise\" target=\"_blank\">datasets</a>)</li>\n<li>Gain</li>\n<li>Noise Injection</li>\n<li>Gaussian Noise</li>\n<li>Pink Noise</li>\n<li>Mixup</li>\n<li><a href=\"https://www.kaggle.com/c/birdclef-2023/discussion/412922\" target=\"_blank\">Sumup</a></li></ul></li>\n<li><p>Model</p>\n<ul>\n<li><p><code>inception-next-nano</code> with attention head</p>\n<ul>\n<li>InceptionNeXt with the same scaling as ConvNeXt-nano</li></ul>\n<pre><code> timm.models.inception_next  _create_inception_next\n timm.models.inception_next  InceptionDWConv2d\n timm.models._registry  register_model\n\n\n ():\n    ()\n    model_args = (\n        depths=(, , , ), dims=(, , , ),\n        token_mixers=InceptionDWConv2d,\n    )\n     _create_inception_next(, pretrained=, **(model_args, **kwargs))\n</code></pre></li></ul></li>\n</ul>\n<h3>Model C: Raw signal CNN</h3>\n<p>This model is inspired by the HMS <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492254\" target=\"_blank\">2nd place solution</a>. A big thanks to <a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>! They provided a detailed explanation of <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511535\" target=\"_blank\">them solution</a>.</p>\n<ul>\n<li><p>Dataset</p>\n<ul>\n<li>BC2024</li>\n<li>Removed duplicate data by referring to  <a href=\"https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/\" target=\"_blank\">[this link]</a>.</li>\n<li>Added several samples from the minority class using data from xeno-canto.</li>\n<li>Applied stratified 5-fold cross-validation grouped by author.</li>\n<li>Classes with fewer than 15 samples were upsampled to 15 samples during training.</li></ul></li>\n<li><p>Preprocessing</p>\n<ul>\n<li>Used the first 5 seconds of each audio sample.</li>\n<li>Downsampled the audio to half the original rate (from 32000 Hz to 16000 Hz).</li>\n<li>Reshaped the downsampled audio data from a size of 80000 to 625x128.</li></ul></li>\n<li><p>Data Augmentation</p>\n<ul>\n<li>Annotated 50 background segments from unlabeled data and added them as background noise.</li>\n<li>Gain</li>\n<li>Noise Injection</li>\n<li>Gaussian Noise</li>\n<li>Pink Noise</li>\n<li>Random Volume</li>\n<li>Mixup</li>\n<li>Cutmix</li></ul></li>\n<li><p>Model</p>\n<ul>\n<li><code>tf_efficientnet_b0_ns</code> with SED head</li></ul></li>\n<li><p>Loss Function</p>\n<ul>\n<li>focal loss</li></ul></li>\n</ul>\n<h2>Validation Strategy</h2>\n<p>Using the training data for validation didn't give us reliable results, so we switched to a synthetic data approach.</p>\n<ol>\n<li>We sampled files for 40 out of 182 classes from the xeno-canto-additional dataset and cropped the segments where the birds were vocalizing to create a clean dataset.</li>\n<li>We sampled audio files containing only background noise (without bird calls) from the unlabeled soundscape dataset.</li>\n<li>We combined the clean dataset and background noise to create a test-like dataset with time-series labels.</li>\n<li>Using this synthetic dataset, we calculated the ROC AUC score. </li>\n</ol>\n<p>Although this validation method did not perfectly correlate with the LB results, it provided more reasonable outcomes compared to using the first 5-second crop method. For more details, please refer to <a href=\"https://www.kaggle.com/code/yokuyama/quant-valid-synthetic-data/notebook\" target=\"_blank\">this notebook</a>.</p>\n<h2>TTA</h2>\n<p>To enhance the accuracy of our time series predictions, we employed techniques similar to sub-pixel super-resolution. Instead of predicting just the 5-second frames during inference, we also predicted frames shifted by 2.5 seconds. We then combined these results as a TTA. This method helped in refining the overall predictions.</p>\n<p><img src=\"https://raw.githubusercontent.com/yoku001/kaggle-static-resouces/main/img/birdclef2024/zu1.drawio.png\" alt=\"img\"></p>\n<h2>OpenVINO + INT8 Post Training Quantization</h2>\n<p>To speed up our model's inference time, we used OpenVINO. Additionally, we implemented <a href=\"https://docs.openvino.ai/2024/openvino-workflow/model-optimization-guide/quantizing-models-post-training/basic-quantization-flow.html\" target=\"_blank\">post-training quantization</a> to convert our model to INT8.</p>\n<p>For the quantization calibration dataset, we used our model's <strong>training dataset</strong>, applying augmentations like background noise addition and gain changes. We believed that these augmentations would help create a quantized model better suited to handle a wider range of test data scenarios.</p>\n<p>The results were impressive: our inference speed improved dramatically, with the quantized model running <strong>30-40%</strong> faster.</p>\n<p>When performing quantization, the selection of layers to be quantized was crucial. We observed that excluding the head layers from quantization tended to improve the model's accuracy.</p>\n<pre><code>names = [, , , , ]\nquantized_model = nncf.quantize(\n    model, calibration_dataset, =600,\n    =nncf.IgnoredScope(names=names),\n)\n</code></pre>\n<p>We began working on quantization just three days before the submission deadline, leaving us insufficient time to thoroughly verify the combination of ensemble and quantization. So, we used a quantized model in only one of our two final submissions. (trade-off: we reduced the number of TTA runs for this submission.)</p>\n<p>We ended up finding that enabling quantization did not significantly impact the public/private scores.</p>\n<h2>Post Processing</h2>\n<p>By using several tricks, we were able to improve our scores by approximately 0.01 on both the private and public leaderboards.</p>\n<ul>\n<li>Smoothing<ul>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511527\" target=\"_blank\">Similar to the 6th place team</a>, we improved our scores by taking the moving average of adjacent segments.</li></ul></li>\n<li>Cut-off<ul>\n<li>Birds that appear once in 4 minutes of audio are more likely to reappear compared to other audio. Recognizing that the probability could be low due to noise, overlapping calls with other birds, and inference slices cut by bin boundaries, we halved the value if the model had a confidence of 0.10 or less in all 48 bins of the 4-minute audio. In other words, we halved the probability if there were no birdsong (0.10 or less) in all 48 sections, and left the number unchanged if there was birdsong at least once in any of the 48 sections.</li></ul></li>\n<li>Max With Neighbors<ul>\n<li>Select the maximum value including the previous and next two rows. For a 30-second test sample, if the inference values of a label are [0.1, 0.3, 0.5, 0.2, 0.4, 0.1], modify them to [0.5, 0.5, 0.5, 0.5, 0.4]. We conducted many tests with different datasets and found a 25% probability of score decrease, so it was not included in the final submission. </li></ul></li>\n</ul>\n<h2>What didn't work</h2>\n<ul>\n<li>Pseudo labeling using unlabeled data.</li>\n<li>Data cleansing and hand-labeling for training data.</li>\n<li>Using novel loss functions. BCE and Focal Loss performed almost the best.</li>\n<li>Manifold Mixup, D-Mixup</li>\n<li>PCEN</li>\n<li>CWT, CQT, VQT</li>\n<li>Trainable frontends: Leaf, trainable filterbank, trainable stft, Conv1D</li>\n<li>Reparameterized model</li>\n<li>Mobilenet V4</li>\n<li>BirdNET embeddings</li>\n</ul>",
      "rawMarkdown": "Thank you to Kaggle, the hosts, and all the competitors. Participating in this exciting competition has been an amazing experience. Here's a look at our 4th-place solution. This achievement was truly a team effort, with equal contributions from @ajobseeker and @tamotamo. I'm grateful to have had the chance to work with them in this competition.\n\nUpdate (2024-06-23):\nAdded the inference notebook and training code.\nInference Notebook: https://www.kaggle.com/code/yokuyama/bc24-4th-place/notebook\nTrain code(melspec models): https://github.com/yoku001/BirdCLEF2024-4th-place-solution-melspec\nTrain code(raw signam models): https://github.com/tamotamo17/BirdCLEF2024-4th-place-solution-raw-signal\n\n## TL;DR\n* Ensemble of Melspec Models and Signal Models\n* TTA\n* OpenVINO\n* Post Processing\n\n## Scores\n\n| Model | Public Score | Private Score | Public Score (+TTA) | Private Score (+TTA) |\n| ------------- | ------------- | ------------- | ------------- | ------------- |\n| Melspec Model B (inception-next-nano) | 0.668 | 0.623 |  |  |\n| Melspec Model A  (rexnet_150)  | 0.676 | 0.641 | 0.690 | 0.649 |\n| Melspec Model A (seresnext26ts)| 0.682 | 0.645 | 0.693 | 0.651 |\n| Raw signal Model C (tf_efficientnet_b0_ns) | 0.673 | 0.620 | 0.691 | 0.636 |\n|   Weighted Mean  |0.717|0.667| 0.731 | 0.676 |\n|Weighted Mean + Geometric Mean  ||| 0.732 | 0.677 |\n|Weighted Mean + Geometric Mean  + Smoothing ||| 0.741 | 0.685 |\n|Weighted Mean + Geometric Mean  + Smoothing + Cut-off (Final Sub) ||| 0.7469 | 0.6877 |\n|Weighted Mean + Geometric Mean  + Smoothing + Cut-off + Max With Neighbors||| 0.749 | 0.689 |\n\n* Weighted Mean\n    * `0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C`\n* Geometric Mean\n    * `(0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C) + 0.3*(Model A (rexnet_150) * Model C)**(0.5)`\n    * By adding the geometric mean of the MelSpec Model(Model A) and the Raw-signal Model(Model C) to the ensemble, we achieved a slight improvement in our score. Choosing this ensemble at the end allowed us to fortunately remain in the prize-winning positions.\n    * We saw a big risk of overfitting, so we decided not to spend any more time adjusting the ensemble weights.\n\n## Model A: 2021-2nd Melspec CNNs\nWe heavily referenced the code from the [2023 2nd place solution](https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution) to build our training and inference pipeline for this model. Big thanks to @honglihang for sharing such valuable information. Their codebase was incredibly strong, and with just a few modifications, we were able to create a single model with an LB score of 0.68.\n\n* Dataset\n    - BC2024\n    - Some models pretrained on 2021, 2022, and 2023's datasets.\n    - Random 15-20 seconds from audio at training, first 5 seconds at validation.\n\n* Preprocessing\n    - n_mels=128, n_fft=2048, f_min=0, f_max=16000, hop_length=627, top_db=80. \n\n* Data Augmentation\n    - AddBackgroundNoise ([datasets](https://www.kaggle.com/datasets/honglihang/background-noise))\n    - Gain\n    - Noise Injection\n    - Gaussian Noise\n    - Pink Noise\n    - Mixup\n* Model\n    - `seresnext26ts`\n    - `rexnet_150`\n        - 2021-2023 pretrained\n* Loss Function\n    - BCELoss\n    - Class sampling weights proposed by [1st place of 2023 competition](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808).\n\n## Model B: Simple Melspec CNNs\n\n* Dataset\n    - BC2024 + xeno-canto-additional-cleaned(see `Validation Strategy` section)\n    - Random 5 seconds from audio for training, first 5 seconds for validation.\n\n* Data Augmentation\n    - AddBackgroundNoise ([datasets](https://www.kaggle.com/datasets/honglihang/background-noise))\n    - Gain\n    - Noise Injection\n    - Gaussian Noise\n    - Pink Noise\n    - Mixup\n    - [Sumup](https://www.kaggle.com/c/birdclef-2023/discussion/412922)\n\n* Model\n    - `inception-next-nano` with attention head\n        * InceptionNeXt with the same scaling as ConvNeXt-nano\n\n        ```\n        from timm.models.inception_next import _create_inception_next\n        from timm.models.inception_next import InceptionDWConv2d\n        from timm.models._registry import register_model\n\n        @register_model\n        def inception_next_nano(pretrained=False, **kwargs):\n            print(\"inception_next_nano\")\n            model_args = dict(\n                depths=(2, 2, 8, 2), dims=(80, 160, 320, 640),\n                token_mixers=InceptionDWConv2d,\n            )\n            return _create_inception_next('inception_next_nano', pretrained=False, **dict(model_args, **kwargs))\n        ```\n\n\n### Model C: Raw signal CNN\nThis model is inspired by the HMS [2nd place solution](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492254). A big thanks to @cooolz! They provided a detailed explanation of [them solution](https://www.kaggle.com/competitions/birdclef-2024/discussion/511535).\n\n* Dataset\n    - BC2024\n    - Removed duplicate data by referring to  [[this link]](https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/).\n    - Added several samples from the minority class using data from xeno-canto.\n    - Applied stratified 5-fold cross-validation grouped by author.\n    - Classes with fewer than 15 samples were upsampled to 15 samples during training.\n\n* Preprocessing\n    - Used the first 5 seconds of each audio sample.\n    - Downsampled the audio to half the original rate (from 32000 Hz to 16000 Hz).\n    - Reshaped the downsampled audio data from a size of 80000 to 625x128.\n\n* Data Augmentation\n    -   Annotated 50 background segments from unlabeled data and added them as background noise.\n    -   Gain\n    -   Noise Injection\n    -   Gaussian Noise\n    -   Pink Noise\n    -   Random Volume\n    -   Mixup\n    -   Cutmix\n* Model\n    - `tf_efficientnet_b0_ns` with SED head\n* Loss Function\n    - focal loss\n\n\n## Validation Strategy\nUsing the training data for validation didn't give us reliable results, so we switched to a synthetic data approach.\n\n1. We sampled files for 40 out of 182 classes from the xeno-canto-additional dataset and cropped the segments where the birds were vocalizing to create a clean dataset.\n1. We sampled audio files containing only background noise (without bird calls) from the unlabeled soundscape dataset.\n1. We combined the clean dataset and background noise to create a test-like dataset with time-series labels.\n1. Using this synthetic dataset, we calculated the ROC AUC score. \n\nAlthough this validation method did not perfectly correlate with the LB results, it provided more reasonable outcomes compared to using the first 5-second crop method. For more details, please refer to [this notebook](https://www.kaggle.com/code/yokuyama/quant-valid-synthetic-data/notebook).\n\n\n## TTA\nTo enhance the accuracy of our time series predictions, we employed techniques similar to sub-pixel super-resolution. Instead of predicting just the 5-second frames during inference, we also predicted frames shifted by 2.5 seconds. We then combined these results as a TTA. This method helped in refining the overall predictions.\n\n![img](https://raw.githubusercontent.com/yoku001/kaggle-static-resouces/main/img/birdclef2024/zu1.drawio.png)\n\n## OpenVINO + INT8 Post Training Quantization\nTo speed up our model's inference time, we used OpenVINO. Additionally, we implemented [post-training quantization](https://docs.openvino.ai/2024/openvino-workflow/model-optimization-guide/quantizing-models-post-training/basic-quantization-flow.html) to convert our model to INT8.\n\nFor the quantization calibration dataset, we used our model's **training dataset**, applying augmentations like background noise addition and gain changes. We believed that these augmentations would help create a quantized model better suited to handle a wider range of test data scenarios.\n\nThe results were impressive: our inference speed improved dramatically, with the quantized model running **30-40%** faster.\n\nWhen performing quantization, the selection of layers to be quantized was crucial. We observed that excluding the head layers from quantization tended to improve the model's accuracy.\n\n```\nnames = ['/head/Gemm/WithoutBiases', '/global_pool/Pow', '/global_pool/GlobalAveragePool', '/global_pool/Pow_1', '/global_pool/Clip']\nquantized_model = nncf.quantize(\n    model, calibration_dataset, subset_size=600,\n    ignored_scope=nncf.IgnoredScope(names=names),\n)\n```\n\nWe began working on quantization just three days before the submission deadline, leaving us insufficient time to thoroughly verify the combination of ensemble and quantization. So, we used a quantized model in only one of our two final submissions. (trade-off: we reduced the number of TTA runs for this submission.)\n\nWe ended up finding that enabling quantization did not significantly impact the public/private scores.\n\n\n## Post Processing\nBy using several tricks, we were able to improve our scores by approximately 0.01 on both the private and public leaderboards.\n\n* Smoothing\n    * [Similar to the 6th place team](https://www.kaggle.com/competitions/birdclef-2024/discussion/511527), we improved our scores by taking the moving average of adjacent segments.\n* Cut-off\n    * Birds that appear once in 4 minutes of audio are more likely to reappear compared to other audio. Recognizing that the probability could be low due to noise, overlapping calls with other birds, and inference slices cut by bin boundaries, we halved the value if the model had a confidence of 0.10 or less in all 48 bins of the 4-minute audio. In other words, we halved the probability if there were no birdsong (0.10 or less) in all 48 sections, and left the number unchanged if there was birdsong at least once in any of the 48 sections.\n* Max With Neighbors\n    * Select the maximum value including the previous and next two rows. For a 30-second test sample, if the inference values of a label are [0.1, 0.3, 0.5, 0.2, 0.4, 0.1], modify them to [0.5, 0.5, 0.5, 0.5, 0.4]. We conducted many tests with different datasets and found a 25% probability of score decrease, so it was not included in the final submission. \n\n\n## What didn't work\n* Pseudo labeling using unlabeled data.\n* Data cleansing and hand-labeling for training data.\n* Using novel loss functions. BCE and Focal Loss performed almost the best.\n* Manifold Mixup, D-Mixup\n* PCEN\n* CWT, CQT, VQT\n* Trainable frontends: Leaf, trainable filterbank, trainable stft, Conv1D\n* Reparameterized model\n* Mobilenet V4\n* BirdNET embeddings",
      "votes": null
    },
    {
      "id": "2868396",
      "postDate": "06/12/2024 10:57:39",
      "content": "<p>First of all, I would like to thank Kaggle and the hosts. And I am especially grateful to honglihang for providing the 2023 2nd place solution. His solution was remarkably perfect and was both a great help and a valuable learning experience. yokuyama and tamotamo were the best teammates, and I had an incredibly perfect time with them. This was my first gold medal, and I am deeply thankful for the unforgettable experience we shared.</p>",
      "rawMarkdown": "First of all, I would like to thank Kaggle and the hosts. And I am especially grateful to honglihang for providing the 2023 2nd place solution. His solution was remarkably perfect and was both a great help and a valuable learning experience. yokuyama and tamotamo were the best teammates, and I had an incredibly perfect time with them. This was my first gold medal, and I am deeply thankful for the unforgettable experience we shared.",
      "votes": null
    },
    {
      "id": "2868702",
      "postDate": "06/12/2024 15:10:41",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/ajobseeker\" target=\"_blank\">@ajobseeker</a> <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a> !</p>",
      "rawMarkdown": "congratulations @ajobseeker @yokuyama !",
      "votes": null
    },
    {
      "id": "2869067",
      "postDate": "06/12/2024 20:12:48",
      "content": "<p>Congratulations on your achievement. I had a question about your approach how do you get these weights for the ensemble models, do you use the loss function to get this weight(importance of their press) for these models or is it based on how each model performs and then decides the weightage for the model. </p>",
      "rawMarkdown": "Congratulations on your achievement. I had a question about your approach how do you get these weights for the ensemble models, do you use the loss function to get this weight(importance of their press) for these models or is it based on how each model performs and then decides the weightage for the model.",
      "votes": null
    },
    {
      "id": "2870278",
      "postDate": "06/13/2024 13:47:47",
      "content": "<p>Congrats.</p>\n<p>I sue the same TTA as you in a previous comp, and I tried it here early. The gain in LB score looked small for a double in inference time hence we did not pursue it. Maybe we should have given how good it is for you.</p>",
      "rawMarkdown": "Congrats.\n\nI sue the same TTA as you in a previous comp, and I tried it here early. The gain in LB score looked small for a double in inference time hence we did not pursue it. Maybe we should have given how good it is for you.",
      "votes": null
    },
    {
      "id": "2870538",
      "postDate": "06/13/2024 16:48:33",
      "content": "<p>Thanks.</p>\n<p>For our ensemble weights, we primarily relied on public LB scores and validation scores using synthetic-data to make rough estimations.</p>\n<p>We didn’t focus too much on the performance of individual models when deciding these weights. Given that we didn't manage to develop any standout models in this competition, we concentrated on exploring combinations that might boost our overall score through ensemble, even if some individual models had lower scores.</p>",
      "rawMarkdown": "Thanks.\n\nFor our ensemble weights, we primarily relied on public LB scores and validation scores using synthetic-data to make rough estimations.\n\nWe didn’t focus too much on the performance of individual models when deciding these weights. Given that we didn't manage to develop any standout models in this competition, we concentrated on exploring combinations that might boost our overall score through ensemble, even if some individual models had lower scores.",
      "votes": null
    },
    {
      "id": "2870570",
      "postDate": "06/13/2024 17:18:40",
      "content": "<p>Thanks. Congrats to you too!</p>\n<p>I am honored to have come up with the same idea as a GM like you. While our team was limited in the number of models we could use for this TTA, we believe it was a worthwhile trade-off.</p>",
      "rawMarkdown": "Thanks. Congrats to you too!\n\nI am honored to have come up with the same idea as a GM like you. While our team was limited in the number of models we could use for this TTA, we believe it was a worthwhile trade-off.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2868396,
      "author_name": "ajobseeker",
      "author_url": "",
      "post_date": "06/12/2024 10:57:39",
      "content": "<p>First of all, I would like to thank Kaggle and the hosts. And I am especially grateful to honglihang for providing the 2023 2nd place solution. His solution was remarkably perfect and was both a great help and a valuable learning experience. yokuyama and tamotamo were the best teammates, and I had an incredibly perfect time with them. This was my first gold medal, and I am deeply thankful for the unforgettable experience we shared.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2868702,
      "author_name": "mubashirsidiki",
      "author_url": "",
      "post_date": "06/12/2024 15:10:41",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/ajobseeker\" target=\"_blank\">@ajobseeker</a> <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2869067,
      "author_name": "arjunm97",
      "author_url": "",
      "post_date": "06/12/2024 20:12:48",
      "content": "<p>Congratulations on your achievement. I had a question about your approach how do you get these weights for the ensemble models, do you use the loss function to get this weight(importance of their press) for these models or is it based on how each model performs and then decides the weightage for the model. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2870538,
          "author_name": "yokuyama",
          "author_url": "",
          "post_date": "06/13/2024 16:48:33",
          "content": "<p>Thanks.</p>\n<p>For our ensemble weights, we primarily relied on public LB scores and validation scores using synthetic-data to make rough estimations.</p>\n<p>We didn’t focus too much on the performance of individual models when deciding these weights. Given that we didn't manage to develop any standout models in this competition, we concentrated on exploring combinations that might boost our overall score through ensemble, even if some individual models had lower scores.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2870278,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/13/2024 13:47:47",
      "content": "<p>Congrats.</p>\n<p>I sue the same TTA as you in a previous comp, and I tried it here early. The gain in LB score looked small for a double in inference time hence we did not pursue it. Maybe we should have given how good it is for you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2870570,
          "author_name": "yokuyama",
          "author_url": "",
          "post_date": "06/13/2024 17:18:40",
          "content": "<p>Thanks. Congrats to you too!</p>\n<p>I am honored to have come up with the same idea as a GM like you. While our team was limited in the number of models we could use for this TTA, we believe it was a worthwhile trade-off.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2868347": "Thank you to Kaggle, the hosts, and all the competitors. Participating in this exciting competition has been an amazing experience. Here's a look at our 4th-place solution. This achievement was truly a team effort, with equal contributions from @ajobseeker and @tamotamo. I'm grateful to have had the chance to work with them in this competition.\n\nUpdate (2024-06-23):\nAdded the inference notebook and training code.\nInference Notebook: https://www.kaggle.com/code/yokuyama/bc24-4th-place/notebook\nTrain code(melspec models): https://github.com/yoku001/BirdCLEF2024-4th-place-solution-melspec\nTrain code(raw signam models): https://github.com/tamotamo17/BirdCLEF2024-4th-place-solution-raw-signal\n\n## TL;DR\n* Ensemble of Melspec Models and Signal Models\n* TTA\n* OpenVINO\n* Post Processing\n\n## Scores\n\n| Model | Public Score | Private Score | Public Score (+TTA) | Private Score (+TTA) |\n| ------------- | ------------- | ------------- | ------------- | ------------- |\n| Melspec Model B (inception-next-nano) | 0.668 | 0.623 |  |  |\n| Melspec Model A  (rexnet_150)  | 0.676 | 0.641 | 0.690 | 0.649 |\n| Melspec Model A (seresnext26ts)| 0.682 | 0.645 | 0.693 | 0.651 |\n| Raw signal Model C (tf_efficientnet_b0_ns) | 0.673 | 0.620 | 0.691 | 0.636 |\n|   Weighted Mean  |0.717|0.667| 0.731 | 0.676 |\n|Weighted Mean + Geometric Mean  ||| 0.732 | 0.677 |\n|Weighted Mean + Geometric Mean  + Smoothing ||| 0.741 | 0.685 |\n|Weighted Mean + Geometric Mean  + Smoothing + Cut-off (Final Sub) ||| 0.7469 | 0.6877 |\n|Weighted Mean + Geometric Mean  + Smoothing + Cut-off + Max With Neighbors||| 0.749 | 0.689 |\n\n* Weighted Mean\n    * `0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C`\n* Geometric Mean\n    * `(0.15*Model B + 0.25*Model A (rexnet_150) + 0.3*Model A (seresnext26ts) + 0.3*Model C) + 0.3*(Model A (rexnet_150) * Model C)**(0.5)`\n    * By adding the geometric mean of the MelSpec Model(Model A) and the Raw-signal Model(Model C) to the ensemble, we achieved a slight improvement in our score. Choosing this ensemble at the end allowed us to fortunately remain in the prize-winning positions.\n    * We saw a big risk of overfitting, so we decided not to spend any more time adjusting the ensemble weights.\n\n## Model A: 2021-2nd Melspec CNNs\nWe heavily referenced the code from the [2023 2nd place solution](https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution) to build our training and inference pipeline for this model. Big thanks to @honglihang for sharing such valuable information. Their codebase was incredibly strong, and with just a few modifications, we were able to create a single model with an LB score of 0.68.\n\n* Dataset\n    - BC2024\n    - Some models pretrained on 2021, 2022, and 2023's datasets.\n    - Random 15-20 seconds from audio at training, first 5 seconds at validation.\n\n* Preprocessing\n    - n_mels=128, n_fft=2048, f_min=0, f_max=16000, hop_length=627, top_db=80. \n\n* Data Augmentation\n    - AddBackgroundNoise ([datasets](https://www.kaggle.com/datasets/honglihang/background-noise))\n    - Gain\n    - Noise Injection\n    - Gaussian Noise\n    - Pink Noise\n    - Mixup\n* Model\n    - `seresnext26ts`\n    - `rexnet_150`\n        - 2021-2023 pretrained\n* Loss Function\n    - BCELoss\n    - Class sampling weights proposed by [1st place of 2023 competition](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808).\n\n## Model B: Simple Melspec CNNs\n\n* Dataset\n    - BC2024 + xeno-canto-additional-cleaned(see `Validation Strategy` section)\n    - Random 5 seconds from audio for training, first 5 seconds for validation.\n\n* Data Augmentation\n    - AddBackgroundNoise ([datasets](https://www.kaggle.com/datasets/honglihang/background-noise))\n    - Gain\n    - Noise Injection\n    - Gaussian Noise\n    - Pink Noise\n    - Mixup\n    - [Sumup](https://www.kaggle.com/c/birdclef-2023/discussion/412922)\n\n* Model\n    - `inception-next-nano` with attention head\n        * InceptionNeXt with the same scaling as ConvNeXt-nano\n\n        ```\n        from timm.models.inception_next import _create_inception_next\n        from timm.models.inception_next import InceptionDWConv2d\n        from timm.models._registry import register_model\n\n        @register_model\n        def inception_next_nano(pretrained=False, **kwargs):\n            print(\"inception_next_nano\")\n            model_args = dict(\n                depths=(2, 2, 8, 2), dims=(80, 160, 320, 640),\n                token_mixers=InceptionDWConv2d,\n            )\n            return _create_inception_next('inception_next_nano', pretrained=False, **dict(model_args, **kwargs))\n        ```\n\n\n### Model C: Raw signal CNN\nThis model is inspired by the HMS [2nd place solution](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492254). A big thanks to @cooolz! They provided a detailed explanation of [them solution](https://www.kaggle.com/competitions/birdclef-2024/discussion/511535).\n\n* Dataset\n    - BC2024\n    - Removed duplicate data by referring to  [[this link]](https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/).\n    - Added several samples from the minority class using data from xeno-canto.\n    - Applied stratified 5-fold cross-validation grouped by author.\n    - Classes with fewer than 15 samples were upsampled to 15 samples during training.\n\n* Preprocessing\n    - Used the first 5 seconds of each audio sample.\n    - Downsampled the audio to half the original rate (from 32000 Hz to 16000 Hz).\n    - Reshaped the downsampled audio data from a size of 80000 to 625x128.\n\n* Data Augmentation\n    -   Annotated 50 background segments from unlabeled data and added them as background noise.\n    -   Gain\n    -   Noise Injection\n    -   Gaussian Noise\n    -   Pink Noise\n    -   Random Volume\n    -   Mixup\n    -   Cutmix\n* Model\n    - `tf_efficientnet_b0_ns` with SED head\n* Loss Function\n    - focal loss\n\n\n## Validation Strategy\nUsing the training data for validation didn't give us reliable results, so we switched to a synthetic data approach.\n\n1. We sampled files for 40 out of 182 classes from the xeno-canto-additional dataset and cropped the segments where the birds were vocalizing to create a clean dataset.\n1. We sampled audio files containing only background noise (without bird calls) from the unlabeled soundscape dataset.\n1. We combined the clean dataset and background noise to create a test-like dataset with time-series labels.\n1. Using this synthetic dataset, we calculated the ROC AUC score. \n\nAlthough this validation method did not perfectly correlate with the LB results, it provided more reasonable outcomes compared to using the first 5-second crop method. For more details, please refer to [this notebook](https://www.kaggle.com/code/yokuyama/quant-valid-synthetic-data/notebook).\n\n\n## TTA\nTo enhance the accuracy of our time series predictions, we employed techniques similar to sub-pixel super-resolution. Instead of predicting just the 5-second frames during inference, we also predicted frames shifted by 2.5 seconds. We then combined these results as a TTA. This method helped in refining the overall predictions.\n\n![img](https://raw.githubusercontent.com/yoku001/kaggle-static-resouces/main/img/birdclef2024/zu1.drawio.png)\n\n## OpenVINO + INT8 Post Training Quantization\nTo speed up our model's inference time, we used OpenVINO. Additionally, we implemented [post-training quantization](https://docs.openvino.ai/2024/openvino-workflow/model-optimization-guide/quantizing-models-post-training/basic-quantization-flow.html) to convert our model to INT8.\n\nFor the quantization calibration dataset, we used our model's **training dataset**, applying augmentations like background noise addition and gain changes. We believed that these augmentations would help create a quantized model better suited to handle a wider range of test data scenarios.\n\nThe results were impressive: our inference speed improved dramatically, with the quantized model running **30-40%** faster.\n\nWhen performing quantization, the selection of layers to be quantized was crucial. We observed that excluding the head layers from quantization tended to improve the model's accuracy.\n\n```\nnames = ['/head/Gemm/WithoutBiases', '/global_pool/Pow', '/global_pool/GlobalAveragePool', '/global_pool/Pow_1', '/global_pool/Clip']\nquantized_model = nncf.quantize(\n    model, calibration_dataset, subset_size=600,\n    ignored_scope=nncf.IgnoredScope(names=names),\n)\n```\n\nWe began working on quantization just three days before the submission deadline, leaving us insufficient time to thoroughly verify the combination of ensemble and quantization. So, we used a quantized model in only one of our two final submissions. (trade-off: we reduced the number of TTA runs for this submission.)\n\nWe ended up finding that enabling quantization did not significantly impact the public/private scores.\n\n\n## Post Processing\nBy using several tricks, we were able to improve our scores by approximately 0.01 on both the private and public leaderboards.\n\n* Smoothing\n    * [Similar to the 6th place team](https://www.kaggle.com/competitions/birdclef-2024/discussion/511527), we improved our scores by taking the moving average of adjacent segments.\n* Cut-off\n    * Birds that appear once in 4 minutes of audio are more likely to reappear compared to other audio. Recognizing that the probability could be low due to noise, overlapping calls with other birds, and inference slices cut by bin boundaries, we halved the value if the model had a confidence of 0.10 or less in all 48 bins of the 4-minute audio. In other words, we halved the probability if there were no birdsong (0.10 or less) in all 48 sections, and left the number unchanged if there was birdsong at least once in any of the 48 sections.\n* Max With Neighbors\n    * Select the maximum value including the previous and next two rows. For a 30-second test sample, if the inference values of a label are [0.1, 0.3, 0.5, 0.2, 0.4, 0.1], modify them to [0.5, 0.5, 0.5, 0.5, 0.4]. We conducted many tests with different datasets and found a 25% probability of score decrease, so it was not included in the final submission. \n\n\n## What didn't work\n* Pseudo labeling using unlabeled data.\n* Data cleansing and hand-labeling for training data.\n* Using novel loss functions. BCE and Focal Loss performed almost the best.\n* Manifold Mixup, D-Mixup\n* PCEN\n* CWT, CQT, VQT\n* Trainable frontends: Leaf, trainable filterbank, trainable stft, Conv1D\n* Reparameterized model\n* Mobilenet V4\n* BirdNET embeddings",
    "2868396": "First of all, I would like to thank Kaggle and the hosts. And I am especially grateful to honglihang for providing the 2023 2nd place solution. His solution was remarkably perfect and was both a great help and a valuable learning experience. yokuyama and tamotamo were the best teammates, and I had an incredibly perfect time with them. This was my first gold medal, and I am deeply thankful for the unforgettable experience we shared.",
    "2868702": "congratulations @ajobseeker @yokuyama !",
    "2869067": "Congratulations on your achievement. I had a question about your approach how do you get these weights for the ensemble models, do you use the loss function to get this weight(importance of their press) for these models or is it based on how each model performs and then decides the weightage for the model.",
    "2870278": "Congrats.\n\nI sue the same TTA as you in a previous comp, and I tried it here early. The gain in LB score looked small for a double in inference time hence we did not pursue it. Maybe we should have given how good it is for you.",
    "2870538": "Thanks.\n\nFor our ensemble weights, we primarily relied on public LB scores and validation scores using synthetic-data to make rough estimations.\n\nWe didn’t focus too much on the performance of individual models when deciding these weights. Given that we didn't manage to develop any standout models in this competition, we concentrated on exploring combinations that might boost our overall score through ensemble, even if some individual models had lower scores.",
    "2870570": "Thanks. Congrats to you too!\n\nI am honored to have come up with the same idea as a GM like you. While our team was limited in the number of models we could use for this TTA, we believe it was a worthwhile trade-off."
  },
  "source": "meta"
}