{
  "id": 511528,
  "title": "8th place solution",
  "url": "/competitions/birdclef-2024/writeups/kapenon-8th-place-solution",
  "author_name": "",
  "post_date": "2024-06-13T07:05:10.500Z",
  "votes": 26,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>8th Place Solution</h1>\n<p>I would like to thank everyone who organized this great competition and all the participants who worked hard to complete it !!</p>\n<h2>Spectrogram</h2>\n<pre><code>MelSpectrogram(\n    sample_rate=,\n    n_fft=,\n    hop_length=,\n    f_min=,\n    f_max=,\n    n_mels=,\n)\n</code></pre>\n<h2>Model</h2>\n<p>I used an ensemble of three models.</p>\n<ul>\n<li><p>SED</p>\n<ul>\n<li>eca_nfnet_l0</li>\n<li>tf_efficientnet_b0.ns_jft_in1k </li></ul></li>\n<li><p>Simple 2D CNN</p>\n<ul>\n<li>eca_nfnet_l0</li></ul></li>\n</ul>\n<p>logits ensemble seems slightly better than rank ensemble.</p>\n<h2>Inference</h2>\n<ul>\n<li>Used openvino to reduce inference time (thanks <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> to <a href=\"https://www.kaggle.com/code/honglihang/openvino-is-all-you-need\" target=\"_blank\">openvino notebook</a>)</li>\n<li>inference time: ~ 90 mins</li>\n</ul>\n<h2>Training</h2>\n<h3>Loss</h3>\n<ul>\n<li>FocalLossBCE  <ul>\n<li>It is almost the same as <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66#Loss\" target=\"_blank\">a great baseline</a>(thanks <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a>)</li>\n<li>Used secondary labels as 1.0</li></ul></li>\n</ul>\n<h3>Step0. Pretrain</h3>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>2021, 2022, and 2023, but only those with primary labels that are not included in 2024's 182.</li>\n<li>First 5 seconds</li></ul></li>\n<li><p>Details</p>\n<ul>\n<li><p>Signal augmentations</p>\n<ul>\n<li>NoiseInjection</li>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>RandomVolume</li></ul></li>\n<li><p>Spectrogram augmentations</p>\n<ul>\n<li>FrequencyMasking (2 bands)</li>\n<li>TimeMasking (2 bands)</li>\n<li>RandomFlip</li></ul></li></ul></li>\n</ul>\n<h3>Step1. Finetune for Pseudo-Labeling</h3>\n<p>As discussed extensively, there was a significant difference in distribution between the xeno data and the test data. This was evident from the model's preds, influenced by various factors like insect noise, bird volume, and echoes from obstacles. To address this gap, I used pseudo-labeling as one of the measures.</p>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>Only 2024</li>\n<li>Filtered by Google's bird vocalization classifier preds.</li>\n<li>5 seconds</li></ul></li>\n<li><p>Details</p>\n<ul>\n<li>Same augmentations as in pretraining</li>\n<li>Simple mixup</li></ul></li>\n</ul>\n<p>I trained models for 15 epochs from a pretrained model and used the model with the highest CV score to infer on unlabeled data.</p>\n<h3>Step2. Finetune for Submission</h3>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>2024, ff1010bird (nocall only), unlabeled_soundscapes (with step 1 preds)</li>\n<li>Filtered using Google's bird vocalization classifier pred scores to extract the foreground for all but nocall</li>\n<li>5 seconds</li>\n<li>full data</li></ul></li>\n<li><p>Details</p>\n<ul>\n<li>sumix (ref: <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">BirdCLEF2023 7th solution</a>)</li>\n<li>sumix is a simple yet highly effective.</li></ul></li>\n</ul>\n<h2>Other Ideas Explored but Insufficiently Tested</h2>\n<ul>\n<li>Variants of SED (frequency-axis attention, hybrid, etc.)</li>\n<li>1D models</li>\n<li>Various mixup methods (sumix was the best)</li>\n<li>longer duration (as a measure against weakly labeled data)</li>\n<li>Background noise, heavy augmentations</li>\n<li>Multi round pseudo labeling</li>\n<li>TTA (time-shifted slightly)</li>\n<li>KD</li>\n</ul>\n<p>I experimented with various ideas, but most of them don't appear to be effective. However, I can't confidently assess their effectiveness because I couldn't establish a reliable CV scheme. Consequently, I had to rely on the leaderboard to assess their effectiveness. This has made BirdCLEF 2024 a particularly tough competition for me.</p>\n<p>In the end, pseudo-labeling was my only method to bridge the domain gap, but after visualizing the pseudo-labels, it became clear that they were inaccurate. I still don't fully understand what was key, but I have a sense that it was beneficial to build a simple pipeline.</p>\n<p>Inference notebook: <a href=\"https://www.kaggle.com/code/kapenon/bc2024-8th-place-solution-kapenon\" target=\"_blank\">https://www.kaggle.com/code/kapenon/bc2024-8th-place-solution-kapenon</a></p>\n<h2>Acknowledgement</h2>\n<p>I would like to express my sincere gratitude to Rist Inc. for their invaluable support.</p>",
  "messages": [
    {
      "id": "2866011",
      "postDate": "06/11/2024 05:28:41",
      "content": "<h1>8th Place Solution</h1>\n<p>I would like to thank everyone who organized this great competition and all the participants who worked hard to complete it !!</p>\n<h2>Spectrogram</h2>\n<pre><code>MelSpectrogram(\n    sample_rate=,\n    n_fft=,\n    hop_length=,\n    f_min=,\n    f_max=,\n    n_mels=,\n)\n</code></pre>\n<h2>Model</h2>\n<p>I used an ensemble of three models.</p>\n<ul>\n<li><p>SED</p>\n<ul>\n<li>eca_nfnet_l0</li>\n<li>tf_efficientnet_b0.ns_jft_in1k </li></ul></li>\n<li><p>Simple 2D CNN</p>\n<ul>\n<li>eca_nfnet_l0</li></ul></li>\n</ul>\n<p>logits ensemble seems slightly better than rank ensemble.</p>\n<h2>Inference</h2>\n<ul>\n<li>Used openvino to reduce inference time (thanks <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> to <a href=\"https://www.kaggle.com/code/honglihang/openvino-is-all-you-need\" target=\"_blank\">openvino notebook</a>)</li>\n<li>inference time: ~ 90 mins</li>\n</ul>\n<h2>Training</h2>\n<h3>Loss</h3>\n<ul>\n<li>FocalLossBCE  <ul>\n<li>It is almost the same as <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66#Loss\" target=\"_blank\">a great baseline</a>(thanks <a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a>)</li>\n<li>Used secondary labels as 1.0</li></ul></li>\n</ul>\n<h3>Step0. Pretrain</h3>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>2021, 2022, and 2023, but only those with primary labels that are not included in 2024's 182.</li>\n<li>First 5 seconds</li></ul></li>\n<li><p>Details</p>\n<ul>\n<li><p>Signal augmentations</p>\n<ul>\n<li>NoiseInjection</li>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>RandomVolume</li></ul></li>\n<li><p>Spectrogram augmentations</p>\n<ul>\n<li>FrequencyMasking (2 bands)</li>\n<li>TimeMasking (2 bands)</li>\n<li>RandomFlip</li></ul></li></ul></li>\n</ul>\n<h3>Step1. Finetune for Pseudo-Labeling</h3>\n<p>As discussed extensively, there was a significant difference in distribution between the xeno data and the test data. This was evident from the model's preds, influenced by various factors like insect noise, bird volume, and echoes from obstacles. To address this gap, I used pseudo-labeling as one of the measures.</p>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>Only 2024</li>\n<li>Filtered by Google's bird vocalization classifier preds.</li>\n<li>5 seconds</li></ul></li>\n<li><p>Details</p>\n<ul>\n<li>Same augmentations as in pretraining</li>\n<li>Simple mixup</li></ul></li>\n</ul>\n<p>I trained models for 15 epochs from a pretrained model and used the model with the highest CV score to infer on unlabeled data.</p>\n<h3>Step2. Finetune for Submission</h3>\n<ul>\n<li><p>Data</p>\n<ul>\n<li>2024, ff1010bird (nocall only), unlabeled_soundscapes (with step 1 preds)</li>\n<li>Filtered using Google's bird vocalization classifier pred scores to extract the foreground for all but nocall</li>\n<li>5 seconds</li>\n<li>full data</li></ul></li>\n<li><p>Details</p>\n<ul>\n<li>sumix (ref: <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">BirdCLEF2023 7th solution</a>)</li>\n<li>sumix is a simple yet highly effective.</li></ul></li>\n</ul>\n<h2>Other Ideas Explored but Insufficiently Tested</h2>\n<ul>\n<li>Variants of SED (frequency-axis attention, hybrid, etc.)</li>\n<li>1D models</li>\n<li>Various mixup methods (sumix was the best)</li>\n<li>longer duration (as a measure against weakly labeled data)</li>\n<li>Background noise, heavy augmentations</li>\n<li>Multi round pseudo labeling</li>\n<li>TTA (time-shifted slightly)</li>\n<li>KD</li>\n</ul>\n<p>I experimented with various ideas, but most of them don't appear to be effective. However, I can't confidently assess their effectiveness because I couldn't establish a reliable CV scheme. Consequently, I had to rely on the leaderboard to assess their effectiveness. This has made BirdCLEF 2024 a particularly tough competition for me.</p>\n<p>In the end, pseudo-labeling was my only method to bridge the domain gap, but after visualizing the pseudo-labels, it became clear that they were inaccurate. I still don't fully understand what was key, but I have a sense that it was beneficial to build a simple pipeline.</p>\n<p>Inference notebook: <a href=\"https://www.kaggle.com/code/kapenon/bc2024-8th-place-solution-kapenon\" target=\"_blank\">https://www.kaggle.com/code/kapenon/bc2024-8th-place-solution-kapenon</a></p>\n<h2>Acknowledgement</h2>\n<p>I would like to express my sincere gratitude to Rist Inc. for their invaluable support.</p>",
      "rawMarkdown": "# 8th Place Solution\n\nI would like to thank everyone who organized this great competition and all the participants who worked hard to complete it !!\n\n## Spectrogram\n```python\nMelSpectrogram(\n    sample_rate=32_000,\n    n_fft=1_095,\n    hop_length=500,\n    f_min=40,\n    f_max=15_000,\n    n_mels=128,\n)\n```\n\n## Model\nI used an ensemble of three models.\n\n- SED\n  - eca_nfnet_l0\n  - tf_efficientnet_b0.ns_jft_in1k \n\n- Simple 2D CNN\n    - eca_nfnet_l0\n\nlogits ensemble seems slightly better than rank ensemble.\n\n## Inference\n- Used openvino to reduce inference time (thanks @honglihang to [openvino notebook](https://www.kaggle.com/code/honglihang/openvino-is-all-you-need))\n- inference time: ~ 90 mins\n\n## Training\n\n### Loss\n- FocalLossBCE  \n    - It is almost the same as [a great baseline](https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66#Loss)(thanks @salmanahmedtamu)\n    - Used secondary labels as 1.0\n\n### Step0. Pretrain\n- Data\n    - 2021, 2022, and 2023, but only those with primary labels that are not included in 2024's 182.\n    - First 5 seconds\n\n- Details\n    - Signal augmentations\n        - NoiseInjection\n        - GaussianNoise\n        - PinkNoise\n        - RandomVolume\n\n    - Spectrogram augmentations\n        - FrequencyMasking (2 bands)\n        - TimeMasking (2 bands)\n        - RandomFlip\n\n\n### Step1. Finetune for Pseudo-Labeling\nAs discussed extensively, there was a significant difference in distribution between the xeno data and the test data. This was evident from the model's preds, influenced by various factors like insect noise, bird volume, and echoes from obstacles. To address this gap, I used pseudo-labeling as one of the measures.\n\n- Data\n    - Only 2024\n    - Filtered by Google's bird vocalization classifier preds.\n    - 5 seconds\n\n- Details\n    - Same augmentations as in pretraining\n    - Simple mixup\n    \nI trained models for 15 epochs from a pretrained model and used the model with the highest CV score to infer on unlabeled data.\n\n### Step2. Finetune for Submission\n\n- Data\n    - 2024, ff1010bird (nocall only), unlabeled_soundscapes (with step 1 preds)\n    - Filtered using Google's bird vocalization classifier pred scores to extract the foreground for all but nocall\n    - 5 seconds\n    - full data\n\n- Details\n    - sumix (ref: [BirdCLEF2023 7th solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922))\n    - sumix is a simple yet highly effective.\n\n## Other Ideas Explored but Insufficiently Tested\n- Variants of SED (frequency-axis attention, hybrid, etc.)\n- 1D models\n- Various mixup methods (sumix was the best)\n- longer duration (as a measure against weakly labeled data)\n- Background noise, heavy augmentations\n- Multi round pseudo labeling\n- TTA (time-shifted slightly)\n- KD\n\nI experimented with various ideas, but most of them don't appear to be effective. However, I can't confidently assess their effectiveness because I couldn't establish a reliable CV scheme. Consequently, I had to rely on the leaderboard to assess their effectiveness. This has made BirdCLEF 2024 a particularly tough competition for me.\n\nIn the end, pseudo-labeling was my only method to bridge the domain gap, but after visualizing the pseudo-labels, it became clear that they were inaccurate. I still don't fully understand what was key, but I have a sense that it was beneficial to build a simple pipeline.\n\nInference notebook: https://www.kaggle.com/code/kapenon/bc2024-8th-place-solution-kapenon\n\n\n## Acknowledgement\nI would like to express my sincere gratitude to Rist Inc. for their invaluable support.",
      "votes": null
    },
    {
      "id": "2866107",
      "postDate": "06/11/2024 06:25:42",
      "content": "<p>congratulations! Glad to see openvino still works well~</p>",
      "rawMarkdown": "congratulations! Glad to see openvino still works well~",
      "votes": null
    },
    {
      "id": "2866269",
      "postDate": "06/11/2024 08:27:44",
      "content": "<p>Congratsss for gold!!</p>",
      "rawMarkdown": "Congratsss for gold!!",
      "votes": null
    },
    {
      "id": "2867802",
      "postDate": "06/12/2024 05:13:10",
      "content": "<p>Congratulations on securing 8th place in this competition. Thanks for sharing the details of your models.<br>\nI observe you are using 32_000 instead of 32,000.  Whether _ is ok to be in integer values.</p>",
      "rawMarkdown": "Congratulations on securing 8th place in this competition. Thanks for sharing the details of your models.\nI observe you are using 32_000 instead of 32,000.  Whether _ is ok to be in integer values.",
      "votes": null
    },
    {
      "id": "2873953",
      "postDate": "06/15/2024 22:11:52",
      "content": "<p>Congrats! How did you use the google bird vocalization classifier to filter the training data &amp; soundscapes? (I tried using the max class probability for it as a no-call classifier, but it didn't seem to work very well.)</p>",
      "rawMarkdown": "Congrats! How did you use the google bird vocalization classifier to filter the training data & soundscapes? (I tried using the max class probability for it as a no-call classifier, but it didn't seem to work very well.)",
      "votes": null
    },
    {
      "id": "2874643",
      "postDate": "06/16/2024 13:12:34",
      "content": "<p>It follows Python syntax. The meaning is the same as 32,000.</p>",
      "rawMarkdown": "It follows Python syntax. The meaning is the same as 32,000.",
      "votes": null
    },
    {
      "id": "2874647",
      "postDate": "06/16/2024 13:14:39",
      "content": "<p>I referred to this Notebook (<a href=\"https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models\" target=\"_blank\">https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models</a>) to perform inference, and used the max probability for each 5-second clip as the result of the call/no-call classifier.</p>\n<p>Hmm, I should be doing almost the same process as you. Could it be the influence of combining other processes?　(e.g. adding no_call class)</p>",
      "rawMarkdown": "I referred to this Notebook (https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models) to perform inference, and used the max probability for each 5-second clip as the result of the call/no-call classifier.\n\nHmm, I should be doing almost the same process as you. Could it be the influence of combining other processes?　(e.g. adding no_call class)",
      "votes": null
    },
    {
      "id": "2874725",
      "postDate": "06/16/2024 14:26:36",
      "content": "<p>Interesting! Did you take the max over just the competition bird probabilities, or all birds? (I took it over all birds)<br>\nHow did you use it on the train_audio? Some options: remove the audio from the train data, keep it but set the targets to 0 prob, or add a no_call class (I kept the data but set the targets to 0 prob if the google bird model predicted &lt; threshold)<br>\nWhat threshold did you use?</p>\n<p>Sorry for all the questions! I tried several variations, so super curious about your approach here :)</p>",
      "rawMarkdown": "Interesting! Did you take the max over just the competition bird probabilities, or all birds? (I took it over all birds)\nHow did you use it on the train_audio? Some options: remove the audio from the train data, keep it but set the targets to 0 prob, or add a no_call class (I kept the data but set the targets to 0 prob if the google bird model predicted < threshold)\nWhat threshold did you use?\n\nSorry for all the questions! I tried several variations, so super curious about your approach here :)",
      "votes": null
    },
    {
      "id": "2874743",
      "postDate": "06/16/2024 14:43:50",
      "content": "<p>I took the maximum over all birds (not 182), just like you. I assumed BirdCLEF2024 data only contained vocalizations from 182 bird species (though I'm not confident about the unlabeled ones), so I chose this approach to account for the possibility of misclassification between classes by Google's model, which has been trained on many classes.</p>\n<p>The threshold is set at 0.5. In my training pipeline, I iterated over each audio file during training. Each audio file was split into 5-second clips, and I filtered all clips based on Google's model predictions using a threshold of 0.5. From those clips with predictions above 0.5, one was randomly selected. If no clips exceeded 0.5, I chose randomly from all clips.</p>\n<p>I'm sorry for any misunderstanding caused by my previous responses. Let me clarify:</p>\n<p>The \"no_call\" class I referred to is from the ff1010 dataset. My hypothesis was that filtering with Google's model would leave us with data where bird vocalizations are clearly present. To achieve both the addition of background noise and maintaining label accuracy, I used mixup to blend the data.</p>",
      "rawMarkdown": "I took the maximum over all birds (not 182), just like you. I assumed BirdCLEF2024 data only contained vocalizations from 182 bird species (though I'm not confident about the unlabeled ones), so I chose this approach to account for the possibility of misclassification between classes by Google's model, which has been trained on many classes.\n\nThe threshold is set at 0.5. In my training pipeline, I iterated over each audio file during training. Each audio file was split into 5-second clips, and I filtered all clips based on Google's model predictions using a threshold of 0.5. From those clips with predictions above 0.5, one was randomly selected. If no clips exceeded 0.5, I chose randomly from all clips.\n\nI'm sorry for any misunderstanding caused by my previous responses. Let me clarify:\n\nThe \"no_call\" class I referred to is from the ff1010 dataset. My hypothesis was that filtering with Google's model would leave us with data where bird vocalizations are clearly present. To achieve both the addition of background noise and maintaining label accuracy, I used mixup to blend the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2866107,
      "author_name": "honglihang",
      "author_url": "",
      "post_date": "06/11/2024 06:25:42",
      "content": "<p>congratulations! Glad to see openvino still works well~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2866269,
      "author_name": "aadityaporwal",
      "author_url": "",
      "post_date": "06/11/2024 08:27:44",
      "content": "<p>Congratsss for gold!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2867802,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "06/12/2024 05:13:10",
      "content": "<p>Congratulations on securing 8th place in this competition. Thanks for sharing the details of your models.<br>\nI observe you are using 32_000 instead of 32,000.  Whether _ is ok to be in integer values.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2874643,
          "author_name": "kapenon",
          "author_url": "",
          "post_date": "06/16/2024 13:12:34",
          "content": "<p>It follows Python syntax. The meaning is the same as 32,000.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2873953,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "06/15/2024 22:11:52",
      "content": "<p>Congrats! How did you use the google bird vocalization classifier to filter the training data &amp; soundscapes? (I tried using the max class probability for it as a no-call classifier, but it didn't seem to work very well.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2874647,
          "author_name": "kapenon",
          "author_url": "",
          "post_date": "06/16/2024 13:14:39",
          "content": "<p>I referred to this Notebook (<a href=\"https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models\" target=\"_blank\">https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models</a>) to perform inference, and used the max probability for each 5-second clip as the result of the call/no-call classifier.</p>\n<p>Hmm, I should be doing almost the same process as you. Could it be the influence of combining other processes?　(e.g. adding no_call class)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2874725,
              "author_name": "robbynevels",
              "author_url": "",
              "post_date": "06/16/2024 14:26:36",
              "content": "<p>Interesting! Did you take the max over just the competition bird probabilities, or all birds? (I took it over all birds)<br>\nHow did you use it on the train_audio? Some options: remove the audio from the train data, keep it but set the targets to 0 prob, or add a no_call class (I kept the data but set the targets to 0 prob if the google bird model predicted &lt; threshold)<br>\nWhat threshold did you use?</p>\n<p>Sorry for all the questions! I tried several variations, so super curious about your approach here :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2874743,
                  "author_name": "kapenon",
                  "author_url": "",
                  "post_date": "06/16/2024 14:43:50",
                  "content": "<p>I took the maximum over all birds (not 182), just like you. I assumed BirdCLEF2024 data only contained vocalizations from 182 bird species (though I'm not confident about the unlabeled ones), so I chose this approach to account for the possibility of misclassification between classes by Google's model, which has been trained on many classes.</p>\n<p>The threshold is set at 0.5. In my training pipeline, I iterated over each audio file during training. Each audio file was split into 5-second clips, and I filtered all clips based on Google's model predictions using a threshold of 0.5. From those clips with predictions above 0.5, one was randomly selected. If no clips exceeded 0.5, I chose randomly from all clips.</p>\n<p>I'm sorry for any misunderstanding caused by my previous responses. Let me clarify:</p>\n<p>The \"no_call\" class I referred to is from the ff1010 dataset. My hypothesis was that filtering with Google's model would leave us with data where bird vocalizations are clearly present. To achieve both the addition of background noise and maintaining label accuracy, I used mixup to blend the data.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2866011": "# 8th Place Solution\n\nI would like to thank everyone who organized this great competition and all the participants who worked hard to complete it !!\n\n## Spectrogram\n```python\nMelSpectrogram(\n    sample_rate=32_000,\n    n_fft=1_095,\n    hop_length=500,\n    f_min=40,\n    f_max=15_000,\n    n_mels=128,\n)\n```\n\n## Model\nI used an ensemble of three models.\n\n- SED\n  - eca_nfnet_l0\n  - tf_efficientnet_b0.ns_jft_in1k \n\n- Simple 2D CNN\n    - eca_nfnet_l0\n\nlogits ensemble seems slightly better than rank ensemble.\n\n## Inference\n- Used openvino to reduce inference time (thanks @honglihang to [openvino notebook](https://www.kaggle.com/code/honglihang/openvino-is-all-you-need))\n- inference time: ~ 90 mins\n\n## Training\n\n### Loss\n- FocalLossBCE  \n    - It is almost the same as [a great baseline](https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66#Loss)(thanks @salmanahmedtamu)\n    - Used secondary labels as 1.0\n\n### Step0. Pretrain\n- Data\n    - 2021, 2022, and 2023, but only those with primary labels that are not included in 2024's 182.\n    - First 5 seconds\n\n- Details\n    - Signal augmentations\n        - NoiseInjection\n        - GaussianNoise\n        - PinkNoise\n        - RandomVolume\n\n    - Spectrogram augmentations\n        - FrequencyMasking (2 bands)\n        - TimeMasking (2 bands)\n        - RandomFlip\n\n\n### Step1. Finetune for Pseudo-Labeling\nAs discussed extensively, there was a significant difference in distribution between the xeno data and the test data. This was evident from the model's preds, influenced by various factors like insect noise, bird volume, and echoes from obstacles. To address this gap, I used pseudo-labeling as one of the measures.\n\n- Data\n    - Only 2024\n    - Filtered by Google's bird vocalization classifier preds.\n    - 5 seconds\n\n- Details\n    - Same augmentations as in pretraining\n    - Simple mixup\n    \nI trained models for 15 epochs from a pretrained model and used the model with the highest CV score to infer on unlabeled data.\n\n### Step2. Finetune for Submission\n\n- Data\n    - 2024, ff1010bird (nocall only), unlabeled_soundscapes (with step 1 preds)\n    - Filtered using Google's bird vocalization classifier pred scores to extract the foreground for all but nocall\n    - 5 seconds\n    - full data\n\n- Details\n    - sumix (ref: [BirdCLEF2023 7th solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922))\n    - sumix is a simple yet highly effective.\n\n## Other Ideas Explored but Insufficiently Tested\n- Variants of SED (frequency-axis attention, hybrid, etc.)\n- 1D models\n- Various mixup methods (sumix was the best)\n- longer duration (as a measure against weakly labeled data)\n- Background noise, heavy augmentations\n- Multi round pseudo labeling\n- TTA (time-shifted slightly)\n- KD\n\nI experimented with various ideas, but most of them don't appear to be effective. However, I can't confidently assess their effectiveness because I couldn't establish a reliable CV scheme. Consequently, I had to rely on the leaderboard to assess their effectiveness. This has made BirdCLEF 2024 a particularly tough competition for me.\n\nIn the end, pseudo-labeling was my only method to bridge the domain gap, but after visualizing the pseudo-labels, it became clear that they were inaccurate. I still don't fully understand what was key, but I have a sense that it was beneficial to build a simple pipeline.\n\nInference notebook: https://www.kaggle.com/code/kapenon/bc2024-8th-place-solution-kapenon\n\n\n## Acknowledgement\nI would like to express my sincere gratitude to Rist Inc. for their invaluable support.",
    "2866107": "congratulations! Glad to see openvino still works well~",
    "2866269": "Congratsss for gold!!",
    "2867802": "Congratulations on securing 8th place in this competition. Thanks for sharing the details of your models.\nI observe you are using 32_000 instead of 32,000.  Whether _ is ok to be in integer values.",
    "2873953": "Congrats! How did you use the google bird vocalization classifier to filter the training data & soundscapes? (I tried using the max class probability for it as a no-call classifier, but it didn't seem to work very well.)",
    "2874643": "It follows Python syntax. The meaning is the same as 32,000.",
    "2874647": "I referred to this Notebook (https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models) to perform inference, and used the max probability for each 5-second clip as the result of the call/no-call classifier.\n\nHmm, I should be doing almost the same process as you. Could it be the influence of combining other processes?　(e.g. adding no_call class)",
    "2874725": "Interesting! Did you take the max over just the competition bird probabilities, or all birds? (I took it over all birds)\nHow did you use it on the train_audio? Some options: remove the audio from the train data, keep it but set the targets to 0 prob, or add a no_call class (I kept the data but set the targets to 0 prob if the google bird model predicted < threshold)\nWhat threshold did you use?\n\nSorry for all the questions! I tried several variations, so super curious about your approach here :)",
    "2874743": "I took the maximum over all birds (not 182), just like you. I assumed BirdCLEF2024 data only contained vocalizations from 182 bird species (though I'm not confident about the unlabeled ones), so I chose this approach to account for the possibility of misclassification between classes by Google's model, which has been trained on many classes.\n\nThe threshold is set at 0.5. In my training pipeline, I iterated over each audio file during training. Each audio file was split into 5-second clips, and I filtered all clips based on Google's model predictions using a threshold of 0.5. From those clips with predictions above 0.5, one was randomly selected. If no clips exceeded 0.5, I chose randomly from all clips.\n\nI'm sorry for any misunderstanding caused by my previous responses. Let me clarify:\n\nThe \"no_call\" class I referred to is from the ff1010 dataset. My hypothesis was that filtering with Google's model would leave us with data where bird vocalizations are clearly present. To achieve both the addition of background noise and maintaining label accuracy, I used mixup to blend the data."
  },
  "source": "meta"
}