{
  "id": 583344,
  "title": "14th place solution",
  "url": "/competitions/birdclef-2025/writeups/kelvin-yevhenii-14th-place-solution",
  "author_name": "",
  "post_date": "2025-06-06T06:26:22.536077600Z",
  "votes": 16,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, we’d like to thank the hosts and Kaggle for organizing this competition.</p>\n<p>Congratulations to <a href=\"https://www.kaggle.com/xyzdivergence\" target=\"_blank\">@xyzdivergence</a> on achieving Kaggle Competitions Grandmaster—well deserved!</p>\n<h3>Model &amp; Training</h3>\n<p>We used a SED architecture and trained on random 10-second audio segments. For the final ensemble, we relied solely on tf_efficientnetv2_m.in21k, trained with slightly varied configurations.</p>\n<p>The best private score for a single SED model was 0.922 (not the selected submission), and 0.894 for a CNN model using the same backbone.</p>\n<p>Our training pipeline consisted of several stages:</p>\n<ul>\n<li><p>Stage 1: Pretraining on the training audio with true labels.</p></li>\n<li><p>Stage 2: Knowledge distillation using both train audio and train soundscapes. We combined average pseudo labels across the full audio (with a 1-second stride, weighted at 0.3) and chunk-level pseudo labels from the teacher model (10-second chunks, weighted at 0.7).</p></li>\n<li><p>We performed several rounds of distillation, selecting the best-performing teacher from the previous round based on leaderboard improvements.</p></li>\n</ul>\n<h3>Features</h3>\n<p>Melspectrogram settings:</p>\n<pre><code> \n \n \n \n \n \n</code></pre>\n<p>Augmentations:</p>\n<pre><code>On waveform: sumix (=1)\nOn spectrogram: mixup (=1), 3 time/frequency masks (=0.5), horizontal flip (=0.5),  random erasing (=0.5)\n</code></pre>\n<h3>Final Submission</h3>\n<p>The final submission was a simple average of three tf_efficientnetv2_m.in21k checkpoints, followed by smoothing using neighboring clips with weights of 0.1, 0.8, and 0.1.</p>\n<p>To speed up inference, all models were converted to OpenVINO.</p>",
  "messages": [
    {
      "id": "3218405",
      "postDate": "06/06/2025 06:26:22",
      "content": "<p>First of all, we’d like to thank the hosts and Kaggle for organizing this competition.</p>\n<p>Congratulations to <a href=\"https://www.kaggle.com/xyzdivergence\" target=\"_blank\">@xyzdivergence</a> on achieving Kaggle Competitions Grandmaster—well deserved!</p>\n<h3>Model &amp; Training</h3>\n<p>We used a SED architecture and trained on random 10-second audio segments. For the final ensemble, we relied solely on tf_efficientnetv2_m.in21k, trained with slightly varied configurations.</p>\n<p>The best private score for a single SED model was 0.922 (not the selected submission), and 0.894 for a CNN model using the same backbone.</p>\n<p>Our training pipeline consisted of several stages:</p>\n<ul>\n<li><p>Stage 1: Pretraining on the training audio with true labels.</p></li>\n<li><p>Stage 2: Knowledge distillation using both train audio and train soundscapes. We combined average pseudo labels across the full audio (with a 1-second stride, weighted at 0.3) and chunk-level pseudo labels from the teacher model (10-second chunks, weighted at 0.7).</p></li>\n<li><p>We performed several rounds of distillation, selecting the best-performing teacher from the previous round based on leaderboard improvements.</p></li>\n</ul>\n<h3>Features</h3>\n<p>Melspectrogram settings:</p>\n<pre><code> \n \n \n \n \n \n</code></pre>\n<p>Augmentations:</p>\n<pre><code>On waveform: sumix (=1)\nOn spectrogram: mixup (=1), 3 time/frequency masks (=0.5), horizontal flip (=0.5),  random erasing (=0.5)\n</code></pre>\n<h3>Final Submission</h3>\n<p>The final submission was a simple average of three tf_efficientnetv2_m.in21k checkpoints, followed by smoothing using neighboring clips with weights of 0.1, 0.8, and 0.1.</p>\n<p>To speed up inference, all models were converted to OpenVINO.</p>",
      "rawMarkdown": "First of all, we’d like to thank the hosts and Kaggle for organizing this competition.\n\nCongratulations to @xyzdivergence on achieving Kaggle Competitions Grandmaster—well deserved!\n\n### Model & Training\n\nWe used a SED architecture and trained on random 10-second audio segments. For the final ensemble, we relied solely on tf_efficientnetv2_m.in21k, trained with slightly varied configurations.\n\nThe best private score for a single SED model was 0.922 (not the selected submission), and 0.894 for a CNN model using the same backbone.\n\nOur training pipeline consisted of several stages:\n\n* Stage 1: Pretraining on the training audio with true labels.\n\n* Stage 2: Knowledge distillation using both train audio and train soundscapes. We combined average pseudo labels across the full audio (with a 1-second stride, weighted at 0.3) and chunk-level pseudo labels from the teacher model (10-second chunks, weighted at 0.7).\n\n* We performed several rounds of distillation, selecting the best-performing teacher from the previous round based on leaderboard improvements.\n\n### Features\n\nMelspectrogram settings:\n\n```\nsample_rate: 32000\nmel_bins: 128\nfmin: 40\nfmax: 15000\nnfft: 1024\nhop_length: 512\n```\n\nAugmentations:\n```\nOn waveform: sumix (p=1)\nOn spectrogram: mixup (p=1), 3 time/frequency masks (p=0.5), horizontal flip (p=0.5), and random erasing (p=0.5)\n```\n\n### Final Submission\n\nThe final submission was a simple average of three tf_efficientnetv2_m.in21k checkpoints, followed by smoothing using neighboring clips with weights of 0.1, 0.8, and 0.1.\n\nTo speed up inference, all models were converted to OpenVINO.",
      "votes": null
    },
    {
      "id": "3218445",
      "postDate": "06/06/2025 07:37:10",
      "content": "<p>Congrats!<br>\nHow many hours do you spend  for training models in stage 1 and 2 and final?<br>\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?<br>\nThank you.</p>",
      "rawMarkdown": "Congrats!\nHow many hours do you spend  for training models in stage 1 and 2 and final?\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?\nThank you.",
      "votes": null
    },
    {
      "id": "3218449",
      "postDate": "06/06/2025 07:40:21",
      "content": "<p>Congratulations, my old friend —  new GrandMaster!🎉</p>",
      "rawMarkdown": "Congratulations, my old friend —  new GrandMaster!🎉",
      "votes": null
    },
    {
      "id": "3218749",
      "postDate": "06/06/2025 16:50:46",
      "content": "<p>Congrats, and thanks for sharing! Do you plan to share the model training code?</p>\n<p>For step 1, did you use secondary labels?</p>\n<p>For step 2, is this right (no true labels):</p>\n<p>Take random 10s crop -&gt; target is (teacher probs on 10s crop) * 0.7 + (averaged teacher probs over entire file) * 0.3</p>\n<p>In the end, you selected the “best-performing teacher” — meaning you did not select the student model, but instead the teacher model, and chose the teacher based on the student’s LB performance?</p>",
      "rawMarkdown": "Congrats, and thanks for sharing! Do you plan to share the model training code?\n\nFor step 1, did you use secondary labels?\n\nFor step 2, is this right (no true labels):\n\nTake random 10s crop -> target is (teacher probs on 10s crop) * 0.7 + (averaged teacher probs over entire file) * 0.3\n\nIn the end, you selected the “best-performing teacher” — meaning you did not select the student model, but instead the teacher model, and chose the teacher based on the student’s LB performance?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218445,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "06/06/2025 07:37:10",
      "content": "<p>Congrats!<br>\nHow many hours do you spend  for training models in stage 1 and 2 and final?<br>\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?<br>\nThank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3218449,
      "author_name": "lihaoweicvch",
      "author_url": "",
      "post_date": "06/06/2025 07:40:21",
      "content": "<p>Congratulations, my old friend —  new GrandMaster!🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3218749,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "06/06/2025 16:50:46",
      "content": "<p>Congrats, and thanks for sharing! Do you plan to share the model training code?</p>\n<p>For step 1, did you use secondary labels?</p>\n<p>For step 2, is this right (no true labels):</p>\n<p>Take random 10s crop -&gt; target is (teacher probs on 10s crop) * 0.7 + (averaged teacher probs over entire file) * 0.3</p>\n<p>In the end, you selected the “best-performing teacher” — meaning you did not select the student model, but instead the teacher model, and chose the teacher based on the student’s LB performance?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218405": "First of all, we’d like to thank the hosts and Kaggle for organizing this competition.\n\nCongratulations to @xyzdivergence on achieving Kaggle Competitions Grandmaster—well deserved!\n\n### Model & Training\n\nWe used a SED architecture and trained on random 10-second audio segments. For the final ensemble, we relied solely on tf_efficientnetv2_m.in21k, trained with slightly varied configurations.\n\nThe best private score for a single SED model was 0.922 (not the selected submission), and 0.894 for a CNN model using the same backbone.\n\nOur training pipeline consisted of several stages:\n\n* Stage 1: Pretraining on the training audio with true labels.\n\n* Stage 2: Knowledge distillation using both train audio and train soundscapes. We combined average pseudo labels across the full audio (with a 1-second stride, weighted at 0.3) and chunk-level pseudo labels from the teacher model (10-second chunks, weighted at 0.7).\n\n* We performed several rounds of distillation, selecting the best-performing teacher from the previous round based on leaderboard improvements.\n\n### Features\n\nMelspectrogram settings:\n\n```\nsample_rate: 32000\nmel_bins: 128\nfmin: 40\nfmax: 15000\nnfft: 1024\nhop_length: 512\n```\n\nAugmentations:\n```\nOn waveform: sumix (p=1)\nOn spectrogram: mixup (p=1), 3 time/frequency masks (p=0.5), horizontal flip (p=0.5), and random erasing (p=0.5)\n```\n\n### Final Submission\n\nThe final submission was a simple average of three tf_efficientnetv2_m.in21k checkpoints, followed by smoothing using neighboring clips with weights of 0.1, 0.8, and 0.1.\n\nTo speed up inference, all models were converted to OpenVINO.",
    "3218445": "Congrats!\nHow many hours do you spend  for training models in stage 1 and 2 and final?\nAnd did you use only GPU from Kaggle? If from the outside how much did it cost approximately?\nThank you.",
    "3218449": "Congratulations, my old friend —  new GrandMaster!🎉",
    "3218749": "Congrats, and thanks for sharing! Do you plan to share the model training code?\n\nFor step 1, did you use secondary labels?\n\nFor step 2, is this right (no true labels):\n\nTake random 10s crop -> target is (teacher probs on 10s crop) * 0.7 + (averaged teacher probs over entire file) * 0.3\n\nIn the end, you selected the “best-performing teacher” — meaning you did not select the student model, but instead the teacher model, and chose the teacher based on the student’s LB performance?"
  },
  "source": "meta"
}