{
  "id": 584034,
  "title": "4th place solution",
  "url": "/competitions/birdclef-2025/writeups/dylan-liu-4th-place-solution",
  "author_name": "",
  "post_date": "2025-06-11T15:03:05.430Z",
  "votes": 18,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, this is my first time to reach the prize on kaggle, much thanks to the organizers and all the participants.</p>\n<p>My solution is a SED solution inspired by birdclef2023 2nd place solution, combined with a custom soft AUC loss and semi-supervised learning.</p>\n<p><strong>Main concept: soft AUC loss</strong></p>\n<p>I always thought that the best loss should be the metric itself, so I tried to find an AUC loss and found:</p>\n<pre><code> AUCLoss(nn.Module):\n     __init__(self, margin=., pos_weight=., neg_weight=.):\n        ().__init__()\n        .margin = margin\n        .pos_weight = pos_weight\n        .neg_weight = neg_weight\n\n     forward(self, preds, labels, sample_weights=None):\n         = preds[labels == ]\n         = preds[labels == ]\n\n         len(pos_preds) ==  or len(neg_preds) == :\n             torch.tensor(., device=preds.device)\n\n         sample_weights is not None:\n             = torch.stack([sample_weights]*labels.shape[], dim=)\n             = sample_weights[labels == ]  #\n             = sample_weights[labels == ]  #\n        :\n             = torch.ones_like(pos_preds) * self.pos_weight\n             = torch.ones_like(neg_preds) * self.neg_weight\n\n         = pos_preds.unsqueeze() - neg_preds.unsqueeze()  #\n         = torch.log( + torch.exp(-diff * self.margin))  #\n\n         = loss_matrix * pos_weights.unsqueeze() * neg_weights.unsqueeze()\n\n         weighted_loss.mean()\n</code></pre>\n<p>This AUC loss seems to be very resistant to overfitting. In all experiments, the cv scores of the models trained with the cross entropy loss were significantly better than those trained with the soft AUC loss, but the lb scores were significantly worse.<br>\nThere is a problem with this AUC loss: it does not support soft labels like cross entropy loss. For both knowledge distillation and semi-supervised learning, I need a loss that supports soft labels, so I made some changes to the above AUC loss:</p>\n<pre><code> SoftAUCLoss(nn.Module):\n     __init__(self, margin=., pos_weight=., neg_weight=.):\n        ().__init__()\n        .margin = margin\n        .pos_weight = pos_weight\n        .neg_weight = neg_weight\n\n forward(self, preds, labels, sample_weights=None):\n         = preds[labels&gt;.]\n         = preds[labels&lt;.]\n         = labels[labels&gt;.]\n         = labels[labels&lt;.]\n\n         len(pos_preds) ==  or len(neg_preds) == :\n             torch.tensor(., device=preds.device)\n\n         = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-.)\n         = torch.ones_like(neg_preds) * self.neg_weight * (.-neg_labels)\n         sample_weights is not None:\n             = torch.stack([sample_weights]*labels.shape[], dim=)\n             = pos_weights * sample_weights\n             = neg_weights * sample_weights\n\n         = pos_preds.unsqueeze() - neg_preds.unsqueeze()  #\n         = torch.log( + torch.exp(-diff * self.margin))  #\n\n         = loss_matrix * pos_weights.unsqueeze() * neg_weights.unsqueeze()\n\n         weighted_loss.mean()\n</code></pre>\n<p>This soft AUC loss + semi-supervised learning improved my single tf_efficientnetv2_b0 model's lb score from 0.850 to 0.901. More importantly, my improvement from 11th to 4th on private lb is probably due to the use of this loss.</p>\n<p><strong>Other things that helped</strong><br>\nSemi-supervised learning. The labeling model is 10 SED models of efficientnet_b0-b4, efficientnetv2_b0-b3 and efficientnetv2_s trained wiht first 10s audio data.<br>\nSmaller hop_length (64) and larger n_mels (256).<br>\nAudio mixup augmentation (which is to add two audios as the new audio and take the maximum value of their labels as the new label). This augmentation did not directly improve my model, but in order to increase the diversity of the final solution, I still added this to the training of some models.</p>\n<p><strong>Things that didn't help or make it worse</strong><br>\nAny kind of pretraining.<br>\nKnowledge distillation.<br>\nModels other than efficientnet.<br>\nData normalizitions other than 2D batch normalizition.</p>\n<p><strong>Final models</strong><br>\n16models of efficientnet_lite0-4, efficientnet_b2-3, efficientnetv2_b2-3 and efficientnetv2_s. 17-25 epochs, learning rate 5e-4. 3 types of mel spectrogram parameters. 2 types of data augmentation. First 10s data and random 10s data.</p>\n<p><strong>Code</strong><br>\n<a href=\"https://github.com/dylanliu2/BirdCLEF2025-4th-place-solution\" target=\"_blank\">https://github.com/dylanliu2/BirdCLEF2025-4th-place-solution</a></p>",
  "messages": [
    {
      "id": "3221649",
      "postDate": "06/11/2025 10:11:36",
      "content": "<p>First of all, this is my first time to reach the prize on kaggle, much thanks to the organizers and all the participants.</p>\n<p>My solution is a SED solution inspired by birdclef2023 2nd place solution, combined with a custom soft AUC loss and semi-supervised learning.</p>\n<p><strong>Main concept: soft AUC loss</strong></p>\n<p>I always thought that the best loss should be the metric itself, so I tried to find an AUC loss and found:</p>\n<pre><code> AUCLoss(nn.Module):\n     __init__(self, margin=., pos_weight=., neg_weight=.):\n        ().__init__()\n        .margin = margin\n        .pos_weight = pos_weight\n        .neg_weight = neg_weight\n\n     forward(self, preds, labels, sample_weights=None):\n         = preds[labels == ]\n         = preds[labels == ]\n\n         len(pos_preds) ==  or len(neg_preds) == :\n             torch.tensor(., device=preds.device)\n\n         sample_weights is not None:\n             = torch.stack([sample_weights]*labels.shape[], dim=)\n             = sample_weights[labels == ]  #\n             = sample_weights[labels == ]  #\n        :\n             = torch.ones_like(pos_preds) * self.pos_weight\n             = torch.ones_like(neg_preds) * self.neg_weight\n\n         = pos_preds.unsqueeze() - neg_preds.unsqueeze()  #\n         = torch.log( + torch.exp(-diff * self.margin))  #\n\n         = loss_matrix * pos_weights.unsqueeze() * neg_weights.unsqueeze()\n\n         weighted_loss.mean()\n</code></pre>\n<p>This AUC loss seems to be very resistant to overfitting. In all experiments, the cv scores of the models trained with the cross entropy loss were significantly better than those trained with the soft AUC loss, but the lb scores were significantly worse.<br>\nThere is a problem with this AUC loss: it does not support soft labels like cross entropy loss. For both knowledge distillation and semi-supervised learning, I need a loss that supports soft labels, so I made some changes to the above AUC loss:</p>\n<pre><code> SoftAUCLoss(nn.Module):\n     __init__(self, margin=., pos_weight=., neg_weight=.):\n        ().__init__()\n        .margin = margin\n        .pos_weight = pos_weight\n        .neg_weight = neg_weight\n\n forward(self, preds, labels, sample_weights=None):\n         = preds[labels&gt;.]\n         = preds[labels&lt;.]\n         = labels[labels&gt;.]\n         = labels[labels&lt;.]\n\n         len(pos_preds) ==  or len(neg_preds) == :\n             torch.tensor(., device=preds.device)\n\n         = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-.)\n         = torch.ones_like(neg_preds) * self.neg_weight * (.-neg_labels)\n         sample_weights is not None:\n             = torch.stack([sample_weights]*labels.shape[], dim=)\n             = pos_weights * sample_weights\n             = neg_weights * sample_weights\n\n         = pos_preds.unsqueeze() - neg_preds.unsqueeze()  #\n         = torch.log( + torch.exp(-diff * self.margin))  #\n\n         = loss_matrix * pos_weights.unsqueeze() * neg_weights.unsqueeze()\n\n         weighted_loss.mean()\n</code></pre>\n<p>This soft AUC loss + semi-supervised learning improved my single tf_efficientnetv2_b0 model's lb score from 0.850 to 0.901. More importantly, my improvement from 11th to 4th on private lb is probably due to the use of this loss.</p>\n<p><strong>Other things that helped</strong><br>\nSemi-supervised learning. The labeling model is 10 SED models of efficientnet_b0-b4, efficientnetv2_b0-b3 and efficientnetv2_s trained wiht first 10s audio data.<br>\nSmaller hop_length (64) and larger n_mels (256).<br>\nAudio mixup augmentation (which is to add two audios as the new audio and take the maximum value of their labels as the new label). This augmentation did not directly improve my model, but in order to increase the diversity of the final solution, I still added this to the training of some models.</p>\n<p><strong>Things that didn't help or make it worse</strong><br>\nAny kind of pretraining.<br>\nKnowledge distillation.<br>\nModels other than efficientnet.<br>\nData normalizitions other than 2D batch normalizition.</p>\n<p><strong>Final models</strong><br>\n16models of efficientnet_lite0-4, efficientnet_b2-3, efficientnetv2_b2-3 and efficientnetv2_s. 17-25 epochs, learning rate 5e-4. 3 types of mel spectrogram parameters. 2 types of data augmentation. First 10s data and random 10s data.</p>\n<p><strong>Code</strong><br>\n<a href=\"https://github.com/dylanliu2/BirdCLEF2025-4th-place-solution\" target=\"_blank\">https://github.com/dylanliu2/BirdCLEF2025-4th-place-solution</a></p>",
      "rawMarkdown": "First of all, this is my first time to reach the prize on kaggle, much thanks to the organizers and all the participants.\n\nMy solution is a SED solution inspired by birdclef2023 2nd place solution, combined with a custom soft AUC loss and semi-supervised learning.\n\n**Main concept: soft AUC loss**\n\nI always thought that the best loss should be the metric itself, so I tried to find an AUC loss and found:\n```\nclass AUCLoss(nn.Module):\n    def __init__(self, margin=1.0, pos_weight=1.0, neg_weight=1.0):\n        super().__init__()\n        self.margin = margin\n        self.pos_weight = pos_weight\n        self.neg_weight = neg_weight\n\n    def forward(self, preds, labels, sample_weights=None):\n        pos_preds = preds[labels == 1]\n        neg_preds = preds[labels == 0]\n        \n        if len(pos_preds) == 0 or len(neg_preds) == 0:\n            return torch.tensor(0.0, device=preds.device)\n        \n        if sample_weights is not None:\n            sample_weights = torch.stack([sample_weights]*labels.shape[1], dim=1)\n            pos_weights = sample_weights[labels == 1]  # [N_pos]\n            neg_weights = sample_weights[labels == 0]  # [N_neg]\n        else:\n            pos_weights = torch.ones_like(pos_preds) * self.pos_weight\n            neg_weights = torch.ones_like(neg_preds) * self.neg_weight\n        \n        diff = pos_preds.unsqueeze(1) - neg_preds.unsqueeze(0)  # [N_pos, N_neg]\n        loss_matrix = torch.log(1 + torch.exp(-diff * self.margin))  # [N_pos, N_neg]\n        \n        weighted_loss = loss_matrix * pos_weights.unsqueeze(1) * neg_weights.unsqueeze(0)\n        \n        return weighted_loss.mean()\n```\nThis AUC loss seems to be very resistant to overfitting. In all experiments, the cv scores of the models trained with the cross entropy loss were significantly better than those trained with the soft AUC loss, but the lb scores were significantly worse.\nThere is a problem with this AUC loss: it does not support soft labels like cross entropy loss. For both knowledge distillation and semi-supervised learning, I need a loss that supports soft labels, so I made some changes to the above AUC loss:\n```\nclass SoftAUCLoss(nn.Module):\n    def __init__(self, margin=1.0, pos_weight=1.0, neg_weight=1.0):\n        super().__init__()\n        self.margin = margin\n        self.pos_weight = pos_weight\n        self.neg_weight = neg_weight\n\ndef forward(self, preds, labels, sample_weights=None):\n        pos_preds = preds[labels>0.5]\n        neg_preds = preds[labels<0.5]\n        pos_labels = labels[labels>0.5]\n        neg_labels = labels[labels<0.5]\n        \n        if len(pos_preds) == 0 or len(neg_preds) == 0:\n            return torch.tensor(0.0, device=preds.device)\n\n        pos_weights = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-0.5)\n        neg_weights = torch.ones_like(neg_preds) * self.neg_weight * (0.5-neg_labels)\n        if sample_weights is not None:\n            sample_weights = torch.stack([sample_weights]*labels.shape[1], dim=1)\n            pos_weights = pos_weights * sample_weights\n            neg_weights = neg_weights * sample_weights\n           \n        diff = pos_preds.unsqueeze(1) - neg_preds.unsqueeze(0)  # [N_pos, N_neg]\n        loss_matrix = torch.log(1 + torch.exp(-diff * self.margin))  # [N_pos, N_neg]\n        \n        weighted_loss = loss_matrix * pos_weights.unsqueeze(1) * neg_weights.unsqueeze(0)\n        \n        return weighted_loss.mean()\n```\nThis soft AUC loss + semi-supervised learning improved my single tf_efficientnetv2_b0 model's lb score from 0.850 to 0.901. More importantly, my improvement from 11th to 4th on private lb is probably due to the use of this loss.\n\n**Other things that helped**\nSemi-supervised learning. The labeling model is 10 SED models of efficientnet_b0-b4, efficientnetv2_b0-b3 and efficientnetv2_s trained wiht first 10s audio data.\nSmaller hop_length (64) and larger n_mels (256).\nAudio mixup augmentation (which is to add two audios as the new audio and take the maximum value of their labels as the new label). This augmentation did not directly improve my model, but in order to increase the diversity of the final solution, I still added this to the training of some models.\n\n**Things that didn't help or make it worse**\nAny kind of pretraining.\nKnowledge distillation.\nModels other than efficientnet.\nData normalizitions other than 2D batch normalizition.\n\n**Final models**\n16models of efficientnet_lite0-4, efficientnet_b2-3, efficientnetv2_b2-3 and efficientnetv2_s. 17-25 epochs, learning rate 5e-4. 3 types of mel spectrogram parameters. 2 types of data augmentation. First 10s data and random 10s data.\n\n**Code**\nhttps://github.com/dylanliu2/BirdCLEF2025-4th-place-solution",
      "votes": null
    },
    {
      "id": "3221696",
      "postDate": "06/11/2025 11:01:45",
      "content": "<p>Nice sharing ! By the way, you can use ``` to edit your code block. Just like:</p>\n<pre><code> SoftAUCLoss(nn.Module):\n     init(self, margin=., pos_weight=., neg_weight=.):\n        ().init()\n        .margin = margin\n        .pos_weight = pos_weight\n        .neg_weight = neg_weight\n\n     forward(self, preds, labels, sample_weights=None):\n         = preds[labels&gt;.]\n         = preds[labels&lt;.] pos_labels = labels[labels&gt;.]\n         = labels[labels&lt;.]\n         len(pos_preds) ==  or len(neg_preds) == :\n             torch.tensor(., device=preds.device)\n\n         = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-.)\n         = torch.ones_like(neg_preds) * self.neg_weight * (.-neg_labels)\n         sample_weights is not None:\n             = torch.stack([sample_weights]*labels.shape[], dim=)\n             = pos_weights * sample_weights\n             = neg_weights * sample_weights\n\n         = pos_preds.unsqueeze() - neg_preds.unsqueeze()  #\n         = torch.log( + torch.exp(-diff * self.margin))  #\n\n         = loss_matrix * pos_weights.unsqueeze() * neg_weights.unsqueeze()\n\n         weighted_loss.mean()\n</code></pre>",
      "rawMarkdown": "Nice sharing ! By the way, you can use ``` to edit your code block. Just like:\n```\nclass SoftAUCLoss(nn.Module):\n    def init(self, margin=1.0, pos_weight=1.0, neg_weight=1.0):\n        super().init()\n        self.margin = margin\n        self.pos_weight = pos_weight\n        self.neg_weight = neg_weight\n\n    def forward(self, preds, labels, sample_weights=None):\n        pos_preds = preds[labels>0.5]\n        neg_preds = preds[labels<0.5] pos_labels = labels[labels>0.5]\n        neg_labels = labels[labels<0.5]\n        if len(pos_preds) == 0 or len(neg_preds) == 0:\n            return torch.tensor(0.0, device=preds.device)\n\n        pos_weights = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-0.5)\n        neg_weights = torch.ones_like(neg_preds) * self.neg_weight * (0.5-neg_labels)\n        if sample_weights is not None:\n            sample_weights = torch.stack([sample_weights]*labels.shape[1], dim=1)\n            pos_weights = pos_weights * sample_weights\n            neg_weights = neg_weights * sample_weights\n\n        diff = pos_preds.unsqueeze(1) - neg_preds.unsqueeze(0)  # [N_pos, N_neg]\n        loss_matrix = torch.log(1 + torch.exp(-diff * self.margin))  # [N_pos, N_neg]\n\n        weighted_loss = loss_matrix * pos_weights.unsqueeze(1) * neg_weights.unsqueeze(0)\n\n        return weighted_loss.mean()\n\n```",
      "votes": null
    },
    {
      "id": "3221836",
      "postDate": "06/11/2025 13:38:43",
      "content": "<p>Ok, thanks</p>",
      "rawMarkdown": "Ok, thanks",
      "votes": null
    },
    {
      "id": "3223714",
      "postDate": "06/13/2025 18:35:27",
      "content": "<p>Cool, thanks!</p>\n<p>BTW, the loss matrix computation can be more numerically stable by using the <code>softplus</code> operation - which computes <code>log(1+exp(x))</code> using numerically stable ops - otherwise you might suffer from under/overflow issues.</p>",
      "rawMarkdown": "Cool, thanks!\n\nBTW, the loss matrix computation can be more numerically stable by using the `softplus` operation - which computes `log(1+exp(x))` using numerically stable ops - otherwise you might suffer from under/overflow issues.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3221696,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "06/11/2025 11:01:45",
      "content": "<p>Nice sharing ! By the way, you can use ``` to edit your code block. Just like:</p>\n<pre><code> SoftAUCLoss(nn.Module):\n     init(self, margin=., pos_weight=., neg_weight=.):\n        ().init()\n        .margin = margin\n        .pos_weight = pos_weight\n        .neg_weight = neg_weight\n\n     forward(self, preds, labels, sample_weights=None):\n         = preds[labels&gt;.]\n         = preds[labels&lt;.] pos_labels = labels[labels&gt;.]\n         = labels[labels&lt;.]\n         len(pos_preds) ==  or len(neg_preds) == :\n             torch.tensor(., device=preds.device)\n\n         = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-.)\n         = torch.ones_like(neg_preds) * self.neg_weight * (.-neg_labels)\n         sample_weights is not None:\n             = torch.stack([sample_weights]*labels.shape[], dim=)\n             = pos_weights * sample_weights\n             = neg_weights * sample_weights\n\n         = pos_preds.unsqueeze() - neg_preds.unsqueeze()  #\n         = torch.log( + torch.exp(-diff * self.margin))  #\n\n         = loss_matrix * pos_weights.unsqueeze() * neg_weights.unsqueeze()\n\n         weighted_loss.mean()\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3221836,
          "author_name": "dylanliuofficial",
          "author_url": "",
          "post_date": "06/11/2025 13:38:43",
          "content": "<p>Ok, thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3223714,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "06/13/2025 18:35:27",
      "content": "<p>Cool, thanks!</p>\n<p>BTW, the loss matrix computation can be more numerically stable by using the <code>softplus</code> operation - which computes <code>log(1+exp(x))</code> using numerically stable ops - otherwise you might suffer from under/overflow issues.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3221649": "First of all, this is my first time to reach the prize on kaggle, much thanks to the organizers and all the participants.\n\nMy solution is a SED solution inspired by birdclef2023 2nd place solution, combined with a custom soft AUC loss and semi-supervised learning.\n\n**Main concept: soft AUC loss**\n\nI always thought that the best loss should be the metric itself, so I tried to find an AUC loss and found:\n```\nclass AUCLoss(nn.Module):\n    def __init__(self, margin=1.0, pos_weight=1.0, neg_weight=1.0):\n        super().__init__()\n        self.margin = margin\n        self.pos_weight = pos_weight\n        self.neg_weight = neg_weight\n\n    def forward(self, preds, labels, sample_weights=None):\n        pos_preds = preds[labels == 1]\n        neg_preds = preds[labels == 0]\n        \n        if len(pos_preds) == 0 or len(neg_preds) == 0:\n            return torch.tensor(0.0, device=preds.device)\n        \n        if sample_weights is not None:\n            sample_weights = torch.stack([sample_weights]*labels.shape[1], dim=1)\n            pos_weights = sample_weights[labels == 1]  # [N_pos]\n            neg_weights = sample_weights[labels == 0]  # [N_neg]\n        else:\n            pos_weights = torch.ones_like(pos_preds) * self.pos_weight\n            neg_weights = torch.ones_like(neg_preds) * self.neg_weight\n        \n        diff = pos_preds.unsqueeze(1) - neg_preds.unsqueeze(0)  # [N_pos, N_neg]\n        loss_matrix = torch.log(1 + torch.exp(-diff * self.margin))  # [N_pos, N_neg]\n        \n        weighted_loss = loss_matrix * pos_weights.unsqueeze(1) * neg_weights.unsqueeze(0)\n        \n        return weighted_loss.mean()\n```\nThis AUC loss seems to be very resistant to overfitting. In all experiments, the cv scores of the models trained with the cross entropy loss were significantly better than those trained with the soft AUC loss, but the lb scores were significantly worse.\nThere is a problem with this AUC loss: it does not support soft labels like cross entropy loss. For both knowledge distillation and semi-supervised learning, I need a loss that supports soft labels, so I made some changes to the above AUC loss:\n```\nclass SoftAUCLoss(nn.Module):\n    def __init__(self, margin=1.0, pos_weight=1.0, neg_weight=1.0):\n        super().__init__()\n        self.margin = margin\n        self.pos_weight = pos_weight\n        self.neg_weight = neg_weight\n\ndef forward(self, preds, labels, sample_weights=None):\n        pos_preds = preds[labels>0.5]\n        neg_preds = preds[labels<0.5]\n        pos_labels = labels[labels>0.5]\n        neg_labels = labels[labels<0.5]\n        \n        if len(pos_preds) == 0 or len(neg_preds) == 0:\n            return torch.tensor(0.0, device=preds.device)\n\n        pos_weights = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-0.5)\n        neg_weights = torch.ones_like(neg_preds) * self.neg_weight * (0.5-neg_labels)\n        if sample_weights is not None:\n            sample_weights = torch.stack([sample_weights]*labels.shape[1], dim=1)\n            pos_weights = pos_weights * sample_weights\n            neg_weights = neg_weights * sample_weights\n           \n        diff = pos_preds.unsqueeze(1) - neg_preds.unsqueeze(0)  # [N_pos, N_neg]\n        loss_matrix = torch.log(1 + torch.exp(-diff * self.margin))  # [N_pos, N_neg]\n        \n        weighted_loss = loss_matrix * pos_weights.unsqueeze(1) * neg_weights.unsqueeze(0)\n        \n        return weighted_loss.mean()\n```\nThis soft AUC loss + semi-supervised learning improved my single tf_efficientnetv2_b0 model's lb score from 0.850 to 0.901. More importantly, my improvement from 11th to 4th on private lb is probably due to the use of this loss.\n\n**Other things that helped**\nSemi-supervised learning. The labeling model is 10 SED models of efficientnet_b0-b4, efficientnetv2_b0-b3 and efficientnetv2_s trained wiht first 10s audio data.\nSmaller hop_length (64) and larger n_mels (256).\nAudio mixup augmentation (which is to add two audios as the new audio and take the maximum value of their labels as the new label). This augmentation did not directly improve my model, but in order to increase the diversity of the final solution, I still added this to the training of some models.\n\n**Things that didn't help or make it worse**\nAny kind of pretraining.\nKnowledge distillation.\nModels other than efficientnet.\nData normalizitions other than 2D batch normalizition.\n\n**Final models**\n16models of efficientnet_lite0-4, efficientnet_b2-3, efficientnetv2_b2-3 and efficientnetv2_s. 17-25 epochs, learning rate 5e-4. 3 types of mel spectrogram parameters. 2 types of data augmentation. First 10s data and random 10s data.\n\n**Code**\nhttps://github.com/dylanliu2/BirdCLEF2025-4th-place-solution",
    "3221696": "Nice sharing ! By the way, you can use ``` to edit your code block. Just like:\n```\nclass SoftAUCLoss(nn.Module):\n    def init(self, margin=1.0, pos_weight=1.0, neg_weight=1.0):\n        super().init()\n        self.margin = margin\n        self.pos_weight = pos_weight\n        self.neg_weight = neg_weight\n\n    def forward(self, preds, labels, sample_weights=None):\n        pos_preds = preds[labels>0.5]\n        neg_preds = preds[labels<0.5] pos_labels = labels[labels>0.5]\n        neg_labels = labels[labels<0.5]\n        if len(pos_preds) == 0 or len(neg_preds) == 0:\n            return torch.tensor(0.0, device=preds.device)\n\n        pos_weights = torch.ones_like(pos_preds) * self.pos_weight * (pos_labels-0.5)\n        neg_weights = torch.ones_like(neg_preds) * self.neg_weight * (0.5-neg_labels)\n        if sample_weights is not None:\n            sample_weights = torch.stack([sample_weights]*labels.shape[1], dim=1)\n            pos_weights = pos_weights * sample_weights\n            neg_weights = neg_weights * sample_weights\n\n        diff = pos_preds.unsqueeze(1) - neg_preds.unsqueeze(0)  # [N_pos, N_neg]\n        loss_matrix = torch.log(1 + torch.exp(-diff * self.margin))  # [N_pos, N_neg]\n\n        weighted_loss = loss_matrix * pos_weights.unsqueeze(1) * neg_weights.unsqueeze(0)\n\n        return weighted_loss.mean()\n\n```",
    "3221836": "Ok, thanks",
    "3223714": "Cool, thanks!\n\nBTW, the loss matrix computation can be more numerically stable by using the `softplus` operation - which computes `log(1+exp(x))` using numerically stable ops - otherwise you might suffer from under/overflow issues."
  },
  "source": "meta"
}