{
  "id": 583324,
  "title": "8th place solution",
  "url": "/competitions/birdclef-2025/discussion/583324",
  "author_name": "",
  "post_date": "2025-06-06T04:47:46.067549700Z",
  "votes": 22,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First, thanks to hosts for hosting this competition and Kagglers sharing their very useful insights and tricks. Especially thanks to <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> and <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a> for their past years solution, <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a> for sharing postprocessing trick and filteraug methods. This is the second time I participate in BirdClef competition. It's so lucky that I can get solo gold this year.</p>\n<h1>TLDR:</h1>\n<p>I used 2 SED model with attention head and 1 CNN model. My three key improvements are pseudo label, knowledge distillation and filteraug augmentations. And some other tricks like inference TTA (followed <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a>'s last year solution), predition smoothing, weighted BCE, hard mixup as well.</p>\n<h1>Baseline</h1>\n<p>I start with my BirdClef 2023 code.  First of all, the 2021/22/23/24 data was used to pretrain the SED model (tf_efficientnetv2_b3 backbone) with BCE loss. It scored <strong>~0.842</strong> with prediction smoothing. The parameters I used:<br>\nmel_bin: 192<br>\nf_min/f_max: 40/140000<br>\nn_fft: 2048<br>\nwin_length/hop_size: 1024/512<br>\ntrain/infer durations: 10/5s</p>\n<h1>public LB improvement journey (single SED model with efv2b3 backbone, Running time: ~21min)</h1>\n<table>\n<thead>\n<tr>\n<th>Methods</th>\n<th>Scored</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.842</td>\n</tr>\n<tr>\n<td>inference TTA(followed <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a>'s last year solution)</td>\n<td>0.848</td>\n</tr>\n<tr>\n<td>remove human voice and circle padding for short audio</td>\n<td>0.858</td>\n</tr>\n<tr>\n<td>MLD knowledge distillation (using unlabeled soundscapes)</td>\n<td>0.870</td>\n</tr>\n<tr>\n<td>pseudo label (without kd)</td>\n<td>0.872</td>\n</tr>\n<tr>\n<td>pseudo label (with kd)</td>\n<td>0.884</td>\n</tr>\n<tr>\n<td>add lower rank power postprocessing (followed <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>'s code)</td>\n<td>0.888</td>\n</tr>\n<tr>\n<td>hard mixup</td>\n<td>0.889</td>\n</tr>\n<tr>\n<td>online pseudo label and framewise level supervised</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>add linear filteraug</td>\n<td>0.897</td>\n</tr>\n<tr>\n<td>add step andlinear filteraug</td>\n<td>0.903</td>\n</tr>\n<tr>\n<td>weighted BCE</td>\n<td>0.905</td>\n</tr>\n<tr>\n<td>dropout_rate=0.2/droppath_rate=0.1 backbone and two direction kd</td>\n<td>0.910</td>\n</tr>\n</tbody>\n</table>\n<ol>\n<li><p>pseudo label<br>\nI chunked the unlabel soundscapes to 10s segments. Then the trained models were used to generate label of every 10s chunks. It is hard to choose thredshold as high thredshold would lead to high fasle negative while low thredshold would contribute th high false positive, after lots of experiments, 0.4 was finnally chosen as thredshold. </p></li>\n<li><p>MLD knowledge distillation<br>\nNoting new, just followed <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511851\" target=\"_blank\">last years solution</a></p></li>\n<li><p>hard mixup<br>\nTraditional mixup is data_mixup = α×data_a + (1-α)×data_b, then loss = α×loss(pred, label_a) + (1-α)×loss(pred,label_b). I modified it to: data_mixup = α×data_a + (1-α)×data_b, then loss = loss(pred, clamp(label_a+label_b,0,1))</p></li>\n<li><p>online pseudo label and framewise level supervised<br>\nBesides generate label offline, I used trained model to generated pseudo label (include clipwise prediction and framewise prediction) during training process and use it to calculate loss (clipwise and framewise).<br>\nSo the final loss = BCEloss(pred,offline_pseudolabel) + BCEloss(pred, online_clip_psudolabel) + BCEloss(pred, online_frame_pseudolabel) + KLdivergeloss(pred, online_clip_softlabel) + KLdivergeloss(pred, oneline_frame_softlabel).</p></li>\n<li><p>two direction kd<br>\nFollowed <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">BirdClef 2023 solution</a></p></li>\n</ol>\n<h1>Ensemble models</h1>\n<p>To improve model diversity, I decided to ensemble SED and CNN models. But it doesn't work when I trained CNN model (backbone -&gt; AdaptiveAvgPooling(1) -&gt; nn.Linear, using pseudo label scored ~0.85+  I don't remember the exact number). In the last 3 days, I change the architecture to backbone -&gt; MeanPooling along the frequency dim  -&gt; MaxPooling along the time dim-&gt; nn.Linear. And it works (using pseudo label scored <strong>0.897</strong> with seresnext26t. It's strange that kd does't work on CNN model).<br>\nSo the final model is SED model with efv2b3 (scored <strong>0.910</strong>), SED model with eca_nfnet_l0 (I have no submission chance to test it's performance on LB) , CNN model with seresnext26t (scored <strong>0.897</strong>). The ensemble score is <strong>0.916</strong>.</p>\n<h1>last day improvement journey</h1>\n<p>At the last day, the last three submission. It hit to me that the lower rank power postprocessing (shared by <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>) maybe contribute to low performance for those high performance models. Then I check my submission, and find that the tricks improve my score from 0.884-&gt;0.888 in my single SED model，just 0.004 improvement, and for thos low score models, it improve about ~0.015. So I tried to believe this hypothesis and decide to submit the ensemble models without this postprocessing tricks. And finally, it scored <strong>0.924</strong>.</p>\n<h2>my final two submission</h2>\n<p>It is a little pity that one of my submission (2 SED models and 3 CNN models scored <strong>0.921</strong>/<strong>0.925</strong> in public/private LB, 4th place). At the beginnng, I want to choose this submission, but in the end, I choose the <strong>0.924</strong> one (2 SED model &amp; 1 CNN model wo <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>'pp trick) and  <strong>0.916</strong> one (same models but w <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>'pp trick).  The idea is that when model scored not so good, this trick can improve the score. I suppose the private LB would scored at most 0.8+ according to past years Birdclef competition but this year the private LB is perfectly consistent with publicc LB. And finally, I miss this submission.</p>",
  "messages": [
    {
      "id": "3218318",
      "postDate": "06/06/2025 04:47:46",
      "content": "<p>First, thanks to hosts for hosting this competition and Kagglers sharing their very useful insights and tricks. Especially thanks to <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> and <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a> for their past years solution, <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a> for sharing postprocessing trick and filteraug methods. This is the second time I participate in BirdClef competition. It's so lucky that I can get solo gold this year.</p>\n<h1>TLDR:</h1>\n<p>I used 2 SED model with attention head and 1 CNN model. My three key improvements are pseudo label, knowledge distillation and filteraug augmentations. And some other tricks like inference TTA (followed <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a>'s last year solution), predition smoothing, weighted BCE, hard mixup as well.</p>\n<h1>Baseline</h1>\n<p>I start with my BirdClef 2023 code.  First of all, the 2021/22/23/24 data was used to pretrain the SED model (tf_efficientnetv2_b3 backbone) with BCE loss. It scored <strong>~0.842</strong> with prediction smoothing. The parameters I used:<br>\nmel_bin: 192<br>\nf_min/f_max: 40/140000<br>\nn_fft: 2048<br>\nwin_length/hop_size: 1024/512<br>\ntrain/infer durations: 10/5s</p>\n<h1>public LB improvement journey (single SED model with efv2b3 backbone, Running time: ~21min)</h1>\n<table>\n<thead>\n<tr>\n<th>Methods</th>\n<th>Scored</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.842</td>\n</tr>\n<tr>\n<td>inference TTA(followed <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a>'s last year solution)</td>\n<td>0.848</td>\n</tr>\n<tr>\n<td>remove human voice and circle padding for short audio</td>\n<td>0.858</td>\n</tr>\n<tr>\n<td>MLD knowledge distillation (using unlabeled soundscapes)</td>\n<td>0.870</td>\n</tr>\n<tr>\n<td>pseudo label (without kd)</td>\n<td>0.872</td>\n</tr>\n<tr>\n<td>pseudo label (with kd)</td>\n<td>0.884</td>\n</tr>\n<tr>\n<td>add lower rank power postprocessing (followed <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>'s code)</td>\n<td>0.888</td>\n</tr>\n<tr>\n<td>hard mixup</td>\n<td>0.889</td>\n</tr>\n<tr>\n<td>online pseudo label and framewise level supervised</td>\n<td>0.893</td>\n</tr>\n<tr>\n<td>add linear filteraug</td>\n<td>0.897</td>\n</tr>\n<tr>\n<td>add step andlinear filteraug</td>\n<td>0.903</td>\n</tr>\n<tr>\n<td>weighted BCE</td>\n<td>0.905</td>\n</tr>\n<tr>\n<td>dropout_rate=0.2/droppath_rate=0.1 backbone and two direction kd</td>\n<td>0.910</td>\n</tr>\n</tbody>\n</table>\n<ol>\n<li><p>pseudo label<br>\nI chunked the unlabel soundscapes to 10s segments. Then the trained models were used to generate label of every 10s chunks. It is hard to choose thredshold as high thredshold would lead to high fasle negative while low thredshold would contribute th high false positive, after lots of experiments, 0.4 was finnally chosen as thredshold. </p></li>\n<li><p>MLD knowledge distillation<br>\nNoting new, just followed <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511851\" target=\"_blank\">last years solution</a></p></li>\n<li><p>hard mixup<br>\nTraditional mixup is data_mixup = α×data_a + (1-α)×data_b, then loss = α×loss(pred, label_a) + (1-α)×loss(pred,label_b). I modified it to: data_mixup = α×data_a + (1-α)×data_b, then loss = loss(pred, clamp(label_a+label_b,0,1))</p></li>\n<li><p>online pseudo label and framewise level supervised<br>\nBesides generate label offline, I used trained model to generated pseudo label (include clipwise prediction and framewise prediction) during training process and use it to calculate loss (clipwise and framewise).<br>\nSo the final loss = BCEloss(pred,offline_pseudolabel) + BCEloss(pred, online_clip_psudolabel) + BCEloss(pred, online_frame_pseudolabel) + KLdivergeloss(pred, online_clip_softlabel) + KLdivergeloss(pred, oneline_frame_softlabel).</p></li>\n<li><p>two direction kd<br>\nFollowed <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412922\" target=\"_blank\">BirdClef 2023 solution</a></p></li>\n</ol>\n<h1>Ensemble models</h1>\n<p>To improve model diversity, I decided to ensemble SED and CNN models. But it doesn't work when I trained CNN model (backbone -&gt; AdaptiveAvgPooling(1) -&gt; nn.Linear, using pseudo label scored ~0.85+  I don't remember the exact number). In the last 3 days, I change the architecture to backbone -&gt; MeanPooling along the frequency dim  -&gt; MaxPooling along the time dim-&gt; nn.Linear. And it works (using pseudo label scored <strong>0.897</strong> with seresnext26t. It's strange that kd does't work on CNN model).<br>\nSo the final model is SED model with efv2b3 (scored <strong>0.910</strong>), SED model with eca_nfnet_l0 (I have no submission chance to test it's performance on LB) , CNN model with seresnext26t (scored <strong>0.897</strong>). The ensemble score is <strong>0.916</strong>.</p>\n<h1>last day improvement journey</h1>\n<p>At the last day, the last three submission. It hit to me that the lower rank power postprocessing (shared by <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>) maybe contribute to low performance for those high performance models. Then I check my submission, and find that the tricks improve my score from 0.884-&gt;0.888 in my single SED model，just 0.004 improvement, and for thos low score models, it improve about ~0.015. So I tried to believe this hypothesis and decide to submit the ensemble models without this postprocessing tricks. And finally, it scored <strong>0.924</strong>.</p>\n<h2>my final two submission</h2>\n<p>It is a little pity that one of my submission (2 SED models and 3 CNN models scored <strong>0.921</strong>/<strong>0.925</strong> in public/private LB, 4th place). At the beginnng, I want to choose this submission, but in the end, I choose the <strong>0.924</strong> one (2 SED model &amp; 1 CNN model wo <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>'pp trick) and  <strong>0.916</strong> one (same models but w <a href=\"https://www.kaggle.com/myso1987\" target=\"_blank\">@myso1987</a>'pp trick).  The idea is that when model scored not so good, this trick can improve the score. I suppose the private LB would scored at most 0.8+ according to past years Birdclef competition but this year the private LB is perfectly consistent with publicc LB. And finally, I miss this submission.</p>",
      "rawMarkdown": "First, thanks to hosts for hosting this competition and Kagglers sharing their very useful insights and tricks. Especially thanks to [@honglihang](https://www.kaggle.com/honglihang) and [@yokuyama](https://www.kaggle.com/yokuyama) for their past years solution, [@myso1987](https://www.kaggle.com/myso1987) for sharing postprocessing trick and filteraug methods. This is the second time I participate in BirdClef competition. It's so lucky that I can get solo gold this year.\n\n# TLDR:\nI used 2 SED model with attention head and 1 CNN model. My three key improvements are pseudo label, knowledge distillation and filteraug augmentations. And some other tricks like inference TTA (followed @yokuyama's last year solution), predition smoothing, weighted BCE, hard mixup as well.\n\n# Baseline\nI start with my BirdClef 2023 code.  First of all, the 2021/22/23/24 data was used to pretrain the SED model (tf_efficientnetv2_b3 backbone) with BCE loss. It scored **~0.842** with prediction smoothing. The parameters I used:\nmel_bin: 192\nf_min/f_max: 40/140000\nn_fft: 2048\nwin_length/hop_size: 1024/512\ntrain/infer durations: 10/5s\n\n# public LB improvement journey (single SED model with efv2b3 backbone, Running time: ~21min)\n| Methods | Scored |\n| --- | --- |\n| Baseline | 0.842 |\n| inference TTA(followed @yokuyama's last year solution) | 0.848 |\n| remove human voice and circle padding for short audio | 0.858 |\n| MLD knowledge distillation (using unlabeled soundscapes) | 0.870 |\n| pseudo label (without kd) | 0.872 |\n| pseudo label (with kd) | 0.884 |\n| add lower rank power postprocessing (followed @myso1987's code) | 0.888 |\n| hard mixup | 0.889 |\n| online pseudo label and framewise level supervised | 0.893 |\n| add linear filteraug | 0.897 |\n| add step andlinear filteraug | 0.903 |\n| weighted BCE | 0.905 |\n| dropout_rate=0.2/droppath_rate=0.1 backbone and two direction kd | 0.910 |\n\n1. pseudo label\nI chunked the unlabel soundscapes to 10s segments. Then the trained models were used to generate label of every 10s chunks. It is hard to choose thredshold as high thredshold would lead to high fasle negative while low thredshold would contribute th high false positive, after lots of experiments, 0.4 was finnally chosen as thredshold. \n\n2. MLD knowledge distillation\nNoting new, just followed [last years solution](https://www.kaggle.com/competitions/birdclef-2024/discussion/511851)\n\n3. hard mixup\nTraditional mixup is data_mixup = α×data_a + (1-α)×data_b, then loss = α×loss(pred, label_a) + (1-α)×loss(pred,label_b). I modified it to: data_mixup = α×data_a + (1-α)×data_b, then loss = loss(pred, clamp(label_a+label_b,0,1))\n\n4. online pseudo label and framewise level supervised\nBesides generate label offline, I used trained model to generated pseudo label (include clipwise prediction and framewise prediction) during training process and use it to calculate loss (clipwise and framewise).\nSo the final loss = BCEloss(pred,offline_pseudolabel) + BCEloss(pred, online_clip_psudolabel) + BCEloss(pred, online_frame_pseudolabel) + KLdivergeloss(pred, online_clip_softlabel) + KLdivergeloss(pred, oneline_frame_softlabel).\n\n5. two direction kd\nFollowed [BirdClef 2023 solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922)\n\n# Ensemble models\nTo improve model diversity, I decided to ensemble SED and CNN models. But it doesn't work when I trained CNN model (backbone -> AdaptiveAvgPooling(1) -> nn.Linear, using pseudo label scored ~0.85+  I don't remember the exact number). In the last 3 days, I change the architecture to backbone -> MeanPooling along the frequency dim  -> MaxPooling along the time dim-> nn.Linear. And it works (using pseudo label scored **0.897** with seresnext26t. It's strange that kd does't work on CNN model).\nSo the final model is SED model with efv2b3 (scored **0.910**), SED model with eca_nfnet_l0 (I have no submission chance to test it's performance on LB) , CNN model with seresnext26t (scored **0.897**). The ensemble score is **0.916**.\n\n# last day improvement journey\nAt the last day, the last three submission. It hit to me that the lower rank power postprocessing (shared by @myso1987) maybe contribute to low performance for those high performance models. Then I check my submission, and find that the tricks improve my score from 0.884->0.888 in my single SED model，just 0.004 improvement, and for thos low score models, it improve about ~0.015. So I tried to believe this hypothesis and decide to submit the ensemble models without this postprocessing tricks. And finally, it scored **0.924**.\n\n## my final two submission\nIt is a little pity that one of my submission (2 SED models and 3 CNN models scored **0.921**/**0.925** in public/private LB, 4th place). At the beginnng, I want to choose this submission, but in the end, I choose the **0.924** one (2 SED model & 1 CNN model wo @myso1987'pp trick) and  **0.916** one (same models but w @myso1987'pp trick).  The idea is that when model scored not so good, this trick can improve the score. I suppose the private LB would scored at most 0.8+ according to past years Birdclef competition but this year the private LB is perfectly consistent with publicc LB. And finally, I miss this submission.",
      "votes": null
    },
    {
      "id": "3218330",
      "postDate": "06/06/2025 04:55:38",
      "content": "<p>I have a question, is the computing resources on Kaggle sufficient? Why do I find that training is particularly slow after applying SED during training</p>",
      "rawMarkdown": "I have a question, is the computing resources on Kaggle sufficient? Why do I find that training is particularly slow after applying SED during training",
      "votes": null
    },
    {
      "id": "3218333",
      "postDate": "06/06/2025 04:59:49",
      "content": "<p>I used one single RTX 3090, 40 epochs training time ~4h, I think Kaggle resource is enough to train the model, maybe you miss some thing in your training process such as multiprocessing useage, too long data preprocessing etc. You can trace the cpu and gpu utilization rate during training process.</p>",
      "rawMarkdown": "I used one single RTX 3090, 40 epochs training time ~4h, I think Kaggle resource is enough to train the model, maybe you miss some thing in your training process such as multiprocessing useage, too long data preprocessing etc. You can trace the cpu and gpu utilization rate during training process.",
      "votes": null
    },
    {
      "id": "3218362",
      "postDate": "06/06/2025 05:39:17",
      "content": "<p>Congrats! Would you mind share your training code?<br>\nYou mentioned you used a single RTX 3090. Do you buy it or rental? If buy how much it cost? Do you place RTX 3090 into your computer from maker or on your own?<br>\nIf rental where do you rent? How much it cost?</p>\n<p>Regards,<br>\nNewbie</p>",
      "rawMarkdown": "Congrats! Would you mind share your training code?\nYou mentioned you used a single RTX 3090. Do you buy it or rental? If buy how much it cost? Do you place RTX 3090 into your computer from maker or on your own?\nIf rental where do you rent? How much it cost?\n\nRegards,\nNewbie",
      "votes": null
    },
    {
      "id": "3218425",
      "postDate": "06/06/2025 07:12:35",
      "content": "<p>Sorry, my training code is mess until now, I afraid that nobody can understand it except me. For me, I use a server which is not belong to me. And as for exact cost I'm sorry I don't know. I used to rent one single 3090 GPU, it cost about 1.6¥/h (~0.2+$/h) from a China company. You can use Kaggle GPU resources, it's free for 20h per week.</p>",
      "rawMarkdown": "Sorry, my training code is mess until now, I afraid that nobody can understand it except me. For me, I use a server which is not belong to me. And as for exact cost I'm sorry I don't know. I used to rent one single 3090 GPU, it cost about 1.6¥/h (~0.2+$/h) from a China company. You can use Kaggle GPU resources, it's free for 20h per week.",
      "votes": null
    },
    {
      "id": "3219249",
      "postDate": "06/07/2025 11:18:01",
      "content": "<p>Thanks for sharing, it was very insightful!<br>\nI’m glad to hear that some of the things I shared were useful to you.</p>",
      "rawMarkdown": "Thanks for sharing, it was very insightful!\nI’m glad to hear that some of the things I shared were useful to you.",
      "votes": null
    },
    {
      "id": "3219558",
      "postDate": "06/08/2025 00:33:21",
      "content": "<p>So we can use 3090 GPU at about 0.2 USD per hour? This is  a good price as far as i know. Please send me the link of the company so i can rent it too. <br>\nThank you. </p>",
      "rawMarkdown": "So we can use 3090 GPU at about 0.2 USD per hour? This is  a good price as far as i know. Please send me the link of the company so i can rent it too. \nThank you.",
      "votes": null
    },
    {
      "id": "3219559",
      "postDate": "06/08/2025 00:35:05",
      "content": "<p>That is an incredible amount of effort for a one-man team. Thank you for sharing your process, and congratulations on the well-deserved solo gold! 🥇</p>",
      "rawMarkdown": "That is an incredible amount of effort for a one-man team. Thank you for sharing your process, and congratulations on the well-deserved solo gold! 🥇",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218330,
      "author_name": "pursueml",
      "author_url": "",
      "post_date": "06/06/2025 04:55:38",
      "content": "<p>I have a question, is the computing resources on Kaggle sufficient? Why do I find that training is particularly slow after applying SED during training</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218333,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "06/06/2025 04:59:49",
          "content": "<p>I used one single RTX 3090, 40 epochs training time ~4h, I think Kaggle resource is enough to train the model, maybe you miss some thing in your training process such as multiprocessing useage, too long data preprocessing etc. You can trace the cpu and gpu utilization rate during training process.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218362,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "06/06/2025 05:39:17",
      "content": "<p>Congrats! Would you mind share your training code?<br>\nYou mentioned you used a single RTX 3090. Do you buy it or rental? If buy how much it cost? Do you place RTX 3090 into your computer from maker or on your own?<br>\nIf rental where do you rent? How much it cost?</p>\n<p>Regards,<br>\nNewbie</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218425,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "06/06/2025 07:12:35",
          "content": "<p>Sorry, my training code is mess until now, I afraid that nobody can understand it except me. For me, I use a server which is not belong to me. And as for exact cost I'm sorry I don't know. I used to rent one single 3090 GPU, it cost about 1.6¥/h (~0.2+$/h) from a China company. You can use Kaggle GPU resources, it's free for 20h per week.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3219558,
              "author_name": "overvalueawareness",
              "author_url": "",
              "post_date": "06/08/2025 00:33:21",
              "content": "<p>So we can use 3090 GPU at about 0.2 USD per hour? This is  a good price as far as i know. Please send me the link of the company so i can rent it too. <br>\nThank you. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3219249,
      "author_name": "myso1987",
      "author_url": "",
      "post_date": "06/07/2025 11:18:01",
      "content": "<p>Thanks for sharing, it was very insightful!<br>\nI’m glad to hear that some of the things I shared were useful to you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3219559,
      "author_name": "rosebeltran",
      "author_url": "",
      "post_date": "06/08/2025 00:35:05",
      "content": "<p>That is an incredible amount of effort for a one-man team. Thank you for sharing your process, and congratulations on the well-deserved solo gold! 🥇</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218318": "First, thanks to hosts for hosting this competition and Kagglers sharing their very useful insights and tricks. Especially thanks to [@honglihang](https://www.kaggle.com/honglihang) and [@yokuyama](https://www.kaggle.com/yokuyama) for their past years solution, [@myso1987](https://www.kaggle.com/myso1987) for sharing postprocessing trick and filteraug methods. This is the second time I participate in BirdClef competition. It's so lucky that I can get solo gold this year.\n\n# TLDR:\nI used 2 SED model with attention head and 1 CNN model. My three key improvements are pseudo label, knowledge distillation and filteraug augmentations. And some other tricks like inference TTA (followed @yokuyama's last year solution), predition smoothing, weighted BCE, hard mixup as well.\n\n# Baseline\nI start with my BirdClef 2023 code.  First of all, the 2021/22/23/24 data was used to pretrain the SED model (tf_efficientnetv2_b3 backbone) with BCE loss. It scored **~0.842** with prediction smoothing. The parameters I used:\nmel_bin: 192\nf_min/f_max: 40/140000\nn_fft: 2048\nwin_length/hop_size: 1024/512\ntrain/infer durations: 10/5s\n\n# public LB improvement journey (single SED model with efv2b3 backbone, Running time: ~21min)\n| Methods | Scored |\n| --- | --- |\n| Baseline | 0.842 |\n| inference TTA(followed @yokuyama's last year solution) | 0.848 |\n| remove human voice and circle padding for short audio | 0.858 |\n| MLD knowledge distillation (using unlabeled soundscapes) | 0.870 |\n| pseudo label (without kd) | 0.872 |\n| pseudo label (with kd) | 0.884 |\n| add lower rank power postprocessing (followed @myso1987's code) | 0.888 |\n| hard mixup | 0.889 |\n| online pseudo label and framewise level supervised | 0.893 |\n| add linear filteraug | 0.897 |\n| add step andlinear filteraug | 0.903 |\n| weighted BCE | 0.905 |\n| dropout_rate=0.2/droppath_rate=0.1 backbone and two direction kd | 0.910 |\n\n1. pseudo label\nI chunked the unlabel soundscapes to 10s segments. Then the trained models were used to generate label of every 10s chunks. It is hard to choose thredshold as high thredshold would lead to high fasle negative while low thredshold would contribute th high false positive, after lots of experiments, 0.4 was finnally chosen as thredshold. \n\n2. MLD knowledge distillation\nNoting new, just followed [last years solution](https://www.kaggle.com/competitions/birdclef-2024/discussion/511851)\n\n3. hard mixup\nTraditional mixup is data_mixup = α×data_a + (1-α)×data_b, then loss = α×loss(pred, label_a) + (1-α)×loss(pred,label_b). I modified it to: data_mixup = α×data_a + (1-α)×data_b, then loss = loss(pred, clamp(label_a+label_b,0,1))\n\n4. online pseudo label and framewise level supervised\nBesides generate label offline, I used trained model to generated pseudo label (include clipwise prediction and framewise prediction) during training process and use it to calculate loss (clipwise and framewise).\nSo the final loss = BCEloss(pred,offline_pseudolabel) + BCEloss(pred, online_clip_psudolabel) + BCEloss(pred, online_frame_pseudolabel) + KLdivergeloss(pred, online_clip_softlabel) + KLdivergeloss(pred, oneline_frame_softlabel).\n\n5. two direction kd\nFollowed [BirdClef 2023 solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412922)\n\n# Ensemble models\nTo improve model diversity, I decided to ensemble SED and CNN models. But it doesn't work when I trained CNN model (backbone -> AdaptiveAvgPooling(1) -> nn.Linear, using pseudo label scored ~0.85+  I don't remember the exact number). In the last 3 days, I change the architecture to backbone -> MeanPooling along the frequency dim  -> MaxPooling along the time dim-> nn.Linear. And it works (using pseudo label scored **0.897** with seresnext26t. It's strange that kd does't work on CNN model).\nSo the final model is SED model with efv2b3 (scored **0.910**), SED model with eca_nfnet_l0 (I have no submission chance to test it's performance on LB) , CNN model with seresnext26t (scored **0.897**). The ensemble score is **0.916**.\n\n# last day improvement journey\nAt the last day, the last three submission. It hit to me that the lower rank power postprocessing (shared by @myso1987) maybe contribute to low performance for those high performance models. Then I check my submission, and find that the tricks improve my score from 0.884->0.888 in my single SED model，just 0.004 improvement, and for thos low score models, it improve about ~0.015. So I tried to believe this hypothesis and decide to submit the ensemble models without this postprocessing tricks. And finally, it scored **0.924**.\n\n## my final two submission\nIt is a little pity that one of my submission (2 SED models and 3 CNN models scored **0.921**/**0.925** in public/private LB, 4th place). At the beginnng, I want to choose this submission, but in the end, I choose the **0.924** one (2 SED model & 1 CNN model wo @myso1987'pp trick) and  **0.916** one (same models but w @myso1987'pp trick).  The idea is that when model scored not so good, this trick can improve the score. I suppose the private LB would scored at most 0.8+ according to past years Birdclef competition but this year the private LB is perfectly consistent with publicc LB. And finally, I miss this submission.",
    "3218330": "I have a question, is the computing resources on Kaggle sufficient? Why do I find that training is particularly slow after applying SED during training",
    "3218333": "I used one single RTX 3090, 40 epochs training time ~4h, I think Kaggle resource is enough to train the model, maybe you miss some thing in your training process such as multiprocessing useage, too long data preprocessing etc. You can trace the cpu and gpu utilization rate during training process.",
    "3218362": "Congrats! Would you mind share your training code?\nYou mentioned you used a single RTX 3090. Do you buy it or rental? If buy how much it cost? Do you place RTX 3090 into your computer from maker or on your own?\nIf rental where do you rent? How much it cost?\n\nRegards,\nNewbie",
    "3218425": "Sorry, my training code is mess until now, I afraid that nobody can understand it except me. For me, I use a server which is not belong to me. And as for exact cost I'm sorry I don't know. I used to rent one single 3090 GPU, it cost about 1.6¥/h (~0.2+$/h) from a China company. You can use Kaggle GPU resources, it's free for 20h per week.",
    "3219249": "Thanks for sharing, it was very insightful!\nI’m glad to hear that some of the things I shared were useful to you.",
    "3219558": "So we can use 3090 GPU at about 0.2 USD per hour? This is  a good price as far as i know. Please send me the link of the company so i can rent it too. \nThank you.",
    "3219559": "That is an incredible amount of effort for a one-man team. Thank you for sharing your process, and congratulations on the well-deserved solo gold! 🥇"
  },
  "source": "meta"
}