{
  "id": 583457,
  "title": "13rd solution for BirdCLEF+ 2025",
  "url": "/competitions/birdclef-2025/writeups/h-k-z-13rd-solution-for-birdclef-2025",
  "author_name": "",
  "post_date": "2025-06-07T02:48:50.740Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks so much for the host and I am very delighted to join the competition. Please let me show big thanks for RihanPiggy <a href=\"url\" target=\"_blank\">https://www.kaggle.com/honglihang</a> without his code we cannot get such goods results. In addition,  Koki (train subset models to boost final solution +0.01+) and Zhang (post-process boost solution +0.004) also made a great effects on the final results. </p>\n<h1>Overview of our solution</h1>\n<p>This year competation is very common problem in industry, you have relatively clean data in training, but there are very dirty data in your inference stage. How to overcome the domain shifting is the key to win the competation. The train_soundscapes would help us to decrease the gap between the training data distribution and test data distribution.</p>\n<p>Our solution can be divided into four stages alomost like all other winners teams:</p>\n<h1>Step 1:</h1>\n<p>Use 2023 2nd code to train the basic SED model with v2s (lb 0.84 ~ 0.85)<br>\nremove the humance voice which need to recalculate the duration (lb + 0.003)<br>\nuse sumix instead of mixup on raw audio singal (lb +0.01)<br>\nKoki cleaned the trainin data (lb +0.003)</p>\n<p>After all the above steps, the base model can be 0.865 in public which was the goods start point</p>\n<h1>Step2:</h1>\n<p>Use the step 1 model to do the pseudo data and inference the data 5s.<br>\nAnd above 5s clip by random sample into train audios it will greatly boost the lb (+0.03 ~ 0.04)</p>\n<h1>Step3:</h1>\n<p>After finished the above experiments, we tried to find other models to ensemble which was most time-cost. We tried lots of backbone, but none of them was better then v2s. we use the following models:</p>\n<p>seresnext26t lb 0.899<br>\nv2_b3 lb 0.901</p>\n<p>the reason why v2_b3 can work is v2_b3 and v2s have very similar model structure. It can also lead the model ensemble can not give little boost lb (+0.04). I think most of the teams had the same problems</p>\n<h1>Step4:</h1>\n<p>Inspired by 24 top 6 solution, we trained rare specials model and add them into final solution it give us lost of boost õn lb (+0.1)</p>\n<p>That's all of our soltuons hope you enjoy. <br>\nHappy kaggle.</p>",
  "messages": [
    {
      "id": "3218983",
      "postDate": "06/07/2025 02:48:36",
      "content": "<p>Thanks so much for the host and I am very delighted to join the competition. Please let me show big thanks for RihanPiggy <a href=\"url\" target=\"_blank\">https://www.kaggle.com/honglihang</a> without his code we cannot get such goods results. In addition,  Koki (train subset models to boost final solution +0.01+) and Zhang (post-process boost solution +0.004) also made a great effects on the final results. </p>\n<h1>Overview of our solution</h1>\n<p>This year competation is very common problem in industry, you have relatively clean data in training, but there are very dirty data in your inference stage. How to overcome the domain shifting is the key to win the competation. The train_soundscapes would help us to decrease the gap between the training data distribution and test data distribution.</p>\n<p>Our solution can be divided into four stages alomost like all other winners teams:</p>\n<h1>Step 1:</h1>\n<p>Use 2023 2nd code to train the basic SED model with v2s (lb 0.84 ~ 0.85)<br>\nremove the humance voice which need to recalculate the duration (lb + 0.003)<br>\nuse sumix instead of mixup on raw audio singal (lb +0.01)<br>\nKoki cleaned the trainin data (lb +0.003)</p>\n<p>After all the above steps, the base model can be 0.865 in public which was the goods start point</p>\n<h1>Step2:</h1>\n<p>Use the step 1 model to do the pseudo data and inference the data 5s.<br>\nAnd above 5s clip by random sample into train audios it will greatly boost the lb (+0.03 ~ 0.04)</p>\n<h1>Step3:</h1>\n<p>After finished the above experiments, we tried to find other models to ensemble which was most time-cost. We tried lots of backbone, but none of them was better then v2s. we use the following models:</p>\n<p>seresnext26t lb 0.899<br>\nv2_b3 lb 0.901</p>\n<p>the reason why v2_b3 can work is v2_b3 and v2s have very similar model structure. It can also lead the model ensemble can not give little boost lb (+0.04). I think most of the teams had the same problems</p>\n<h1>Step4:</h1>\n<p>Inspired by 24 top 6 solution, we trained rare specials model and add them into final solution it give us lost of boost õn lb (+0.1)</p>\n<p>That's all of our soltuons hope you enjoy. <br>\nHappy kaggle.</p>",
      "rawMarkdown": "Thanks so much for the host and I am very delighted to join the competition. Please let me show big thanks for RihanPiggy [https://www.kaggle.com/honglihang](url) without his code we cannot get such goods results. In addition,  Koki (train subset models to boost final solution +0.01+) and Zhang (post-process boost solution +0.004) also made a great effects on the final results. \n\n# Overview of our solution\n\nThis year competation is very common problem in industry, you have relatively clean data in training, but there are very dirty data in your inference stage. How to overcome the domain shifting is the key to win the competation. The train_soundscapes would help us to decrease the gap between the training data distribution and test data distribution.\n\nOur solution can be divided into four stages alomost like all other winners teams:\n\n# Step 1:\nUse 2023 2nd code to train the basic SED model with v2s (lb 0.84 ~ 0.85)\nremove the humance voice which need to recalculate the duration (lb + 0.003)\nuse sumix instead of mixup on raw audio singal (lb +0.01)\nKoki cleaned the trainin data (lb +0.003)\n\nAfter all the above steps, the base model can be 0.865 in public which was the goods start point\n\n# Step2:\nUse the step 1 model to do the pseudo data and inference the data 5s.\nAnd above 5s clip by random sample into train audios it will greatly boost the lb (+0.03 ~ 0.04)\n\n# Step3:\nAfter finished the above experiments, we tried to find other models to ensemble which was most time-cost. We tried lots of backbone, but none of them was better then v2s. we use the following models:\n\nseresnext26t lb 0.899\nv2_b3 lb 0.901\n\nthe reason why v2_b3 can work is v2_b3 and v2s have very similar model structure. It can also lead the model ensemble can not give little boost lb (+0.04). I think most of the teams had the same problems\n\n# Step4:\nInspired by 24 top 6 solution, we trained rare specials model and add them into final solution it give us lost of boost õn lb (+0.1)\n\nThat's all of our soltuons hope you enjoy. \nHappy kaggle.",
      "votes": null
    },
    {
      "id": "3219250",
      "postDate": "06/07/2025 11:19:34",
      "content": "<p>Training a model specifically for rare species is certainly a promising direction—great observation! There might be a minor typo in the reported leaderboard gain; I believe it should be 0.01. Nevertheless, that's a great substantial improvement.</p>",
      "rawMarkdown": "Training a model specifically for rare species is certainly a promising direction—great observation! There might be a minor typo in the reported leaderboard gain; I believe it should be 0.01. Nevertheless, that's a great substantial improvement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3219250,
      "author_name": "shanzhong8",
      "author_url": "",
      "post_date": "06/07/2025 11:19:34",
      "content": "<p>Training a model specifically for rare species is certainly a promising direction—great observation! There might be a minor typo in the reported leaderboard gain; I believe it should be 0.01. Nevertheless, that's a great substantial improvement.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218983": "Thanks so much for the host and I am very delighted to join the competition. Please let me show big thanks for RihanPiggy [https://www.kaggle.com/honglihang](url) without his code we cannot get such goods results. In addition,  Koki (train subset models to boost final solution +0.01+) and Zhang (post-process boost solution +0.004) also made a great effects on the final results. \n\n# Overview of our solution\n\nThis year competation is very common problem in industry, you have relatively clean data in training, but there are very dirty data in your inference stage. How to overcome the domain shifting is the key to win the competation. The train_soundscapes would help us to decrease the gap between the training data distribution and test data distribution.\n\nOur solution can be divided into four stages alomost like all other winners teams:\n\n# Step 1:\nUse 2023 2nd code to train the basic SED model with v2s (lb 0.84 ~ 0.85)\nremove the humance voice which need to recalculate the duration (lb + 0.003)\nuse sumix instead of mixup on raw audio singal (lb +0.01)\nKoki cleaned the trainin data (lb +0.003)\n\nAfter all the above steps, the base model can be 0.865 in public which was the goods start point\n\n# Step2:\nUse the step 1 model to do the pseudo data and inference the data 5s.\nAnd above 5s clip by random sample into train audios it will greatly boost the lb (+0.03 ~ 0.04)\n\n# Step3:\nAfter finished the above experiments, we tried to find other models to ensemble which was most time-cost. We tried lots of backbone, but none of them was better then v2s. we use the following models:\n\nseresnext26t lb 0.899\nv2_b3 lb 0.901\n\nthe reason why v2_b3 can work is v2_b3 and v2s have very similar model structure. It can also lead the model ensemble can not give little boost lb (+0.04). I think most of the teams had the same problems\n\n# Step4:\nInspired by 24 top 6 solution, we trained rare specials model and add them into final solution it give us lost of boost õn lb (+0.1)\n\nThat's all of our soltuons hope you enjoy. \nHappy kaggle.",
    "3219250": "Training a model specifically for rare species is certainly a promising direction—great observation! There might be a minor typo in the reported leaderboard gain; I believe it should be 0.01. Nevertheless, that's a great substantial improvement."
  },
  "source": "meta"
}