{
  "id": 511499,
  "title": "11st solution",
  "url": "/competitions/birdclef-2024/discussion/511499",
  "author_name": "lhwcv",
  "post_date": "2024-06-11T00:36:21.297000",
  "votes": 33,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm very grateful to the organizers. This competition was very challenging for me, and due to good fortune, we secured the last gold medal position.<br>\nLet me briefly report on our plan, no big difference between here: <br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/497539\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/497539</a></p>\n<h2>Data</h2>\n<p>BirdCLEF2021, 2022, 2023, 2024 (this competition), ff1010bird_nocall, we didn't adopt additional mp3. When it was introduced, the LB had dropped, even though I found it odd.</p>\n<h2>Model</h2>\n<p>We adopted the model that ranked 4th last year, but swapped the backbone with efficientnet b0/b1/b2.</p>\n<h2>Feature</h2>\n<p>Log mel spec with:</p>\n<pre><code> = \n = \n = \n = \n = \n = \n = \n</code></pre>\n<h2>Data augmentation</h2>\n<p>heavy augmentation improve my CV but LB drop.</p>\n<pre><code>A.HorizontalFlip(=0.5),\nA.OneOf([\n                    A.Cutout(=5, =16),\n                    A.CoarseDropout(=4),\n], =0.5),\n</code></pre>\n<p>and mixup v2:</p>\n<pre><code> MixupV2(nn.Module):\n     __init__(self, mix_range=(., .), add_label=True):\n        (MixupV2, self).__init__()\n        .distribution = torch.distributions.Uniform(low=mix_range[],\n                                                        =mix_range[])\n        .add_label = add_label\n\n     forward(self, X, Y, weight=None, teacher_preds=None):\n\n         = X.shape[]\n         = len(X.shape)\n         = torch.randperm(bs)\n         = self.distribution.rsample(torch.Size((bs,))).to(X.device)\n\n         n_dims == :\n             = coeffs.view(-, ) * X + ( - coeffs.view(-, )) * X[perm]\n         n_dims == :\n             = coeffs.view(-, , ) * X + ( - coeffs.view(-, , )) * X[perm]\n        :\n             = coeffs.view(-, , , ) * X + ( - coeffs.view(-, , , )) * X[perm]\n\n         self.add_label:\n             = Y + Y[perm]\n             = torch.clamp(Y, , .)\n        :\n             = coeffs.view(-, ) * Y + ( - coeffs.view(-, )) * Y[perm]\n\n         weight is None:\n             X, Y\n        :\n             = coeffs.view(-) * weight + ( - coeffs.view(-)) * weight[perm]\n             self.add_label:\n                 = teacher_preds + teacher_preds[perm]\n                 = torch.clamp(teacher_preds, , .)\n            :\n                 = coeffs.view(-, ) * teacher_preds + ( - coeffs.view(-, )) * teacher_preds[perm]\n             X, Y, weight, teacher_preds\n</code></pre>\n<h2>Training</h2>\n<p>First, we pre-trained the first 15 seconds on BirdCLEF2021, 2022, 2023 but removed the classes that were included in the 2024 data to avoid leakage. Then, the first 5 seconds of the 2024 data were used to train the final model. This training plan is the most stable we have found.</p>\n<h2>Postprocessing</h2>\n<p>Slightly improve the LB both Public and Private , My teammate will supplement further later.</p>\n<h2>Tried but not work</h2>\n<ul>\n<li>hubert and distill</li>\n<li>distill or finetune birdnet</li>\n<li>pseudo label</li>\n<li>select multiple segs in an audio by model according to score (why? quite strange )</li>\n</ul>\n<h2>Notebook</h2>\n<p>Our single best model is a b2 based model, Public 0.68, Private 0.67. <br>\nWe finally chose an ensemble of 6 models, Public 0.7, Private 0.67 without post processing<br>\n<a href=\"https://www.kaggle.com/lihaoweicvch/lhw-final-en6-onnx\" target=\"_blank\">https://www.kaggle.com/lihaoweicvch/lhw-final-en6-onnx</a></p>",
  "messages": [
    {
      "id": 2865726,
      "postDate": "2024-06-11T00:36:21.297Z",
      "content": "<p>I'm very grateful to the organizers. This competition was very challenging for me, and due to good fortune, we secured the last gold medal position.<br>\nLet me briefly report on our plan, no big difference between here: <br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/497539\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/497539</a></p>\n<h2>Data</h2>\n<p>BirdCLEF2021, 2022, 2023, 2024 (this competition), ff1010bird_nocall, we didn't adopt additional mp3. When it was introduced, the LB had dropped, even though I found it odd.</p>\n<h2>Model</h2>\n<p>We adopted the model that ranked 4th last year, but swapped the backbone with efficientnet b0/b1/b2.</p>\n<h2>Feature</h2>\n<p>Log mel spec with:</p>\n<pre><code> = \n = \n = \n = \n = \n = \n = \n</code></pre>\n<h2>Data augmentation</h2>\n<p>heavy augmentation improve my CV but LB drop.</p>\n<pre><code>A.HorizontalFlip(=0.5),\nA.OneOf([\n                    A.Cutout(=5, =16),\n                    A.CoarseDropout(=4),\n], =0.5),\n</code></pre>\n<p>and mixup v2:</p>\n<pre><code> MixupV2(nn.Module):\n     __init__(self, mix_range=(., .), add_label=True):\n        (MixupV2, self).__init__()\n        .distribution = torch.distributions.Uniform(low=mix_range[],\n                                                        =mix_range[])\n        .add_label = add_label\n\n     forward(self, X, Y, weight=None, teacher_preds=None):\n\n         = X.shape[]\n         = len(X.shape)\n         = torch.randperm(bs)\n         = self.distribution.rsample(torch.Size((bs,))).to(X.device)\n\n         n_dims == :\n             = coeffs.view(-, ) * X + ( - coeffs.view(-, )) * X[perm]\n         n_dims == :\n             = coeffs.view(-, , ) * X + ( - coeffs.view(-, , )) * X[perm]\n        :\n             = coeffs.view(-, , , ) * X + ( - coeffs.view(-, , , )) * X[perm]\n\n         self.add_label:\n             = Y + Y[perm]\n             = torch.clamp(Y, , .)\n        :\n             = coeffs.view(-, ) * Y + ( - coeffs.view(-, )) * Y[perm]\n\n         weight is None:\n             X, Y\n        :\n             = coeffs.view(-) * weight + ( - coeffs.view(-)) * weight[perm]\n             self.add_label:\n                 = teacher_preds + teacher_preds[perm]\n                 = torch.clamp(teacher_preds, , .)\n            :\n                 = coeffs.view(-, ) * teacher_preds + ( - coeffs.view(-, )) * teacher_preds[perm]\n             X, Y, weight, teacher_preds\n</code></pre>\n<h2>Training</h2>\n<p>First, we pre-trained the first 15 seconds on BirdCLEF2021, 2022, 2023 but removed the classes that were included in the 2024 data to avoid leakage. Then, the first 5 seconds of the 2024 data were used to train the final model. This training plan is the most stable we have found.</p>\n<h2>Postprocessing</h2>\n<p>Slightly improve the LB both Public and Private , My teammate will supplement further later.</p>\n<h2>Tried but not work</h2>\n<ul>\n<li>hubert and distill</li>\n<li>distill or finetune birdnet</li>\n<li>pseudo label</li>\n<li>select multiple segs in an audio by model according to score (why? quite strange )</li>\n</ul>\n<h2>Notebook</h2>\n<p>Our single best model is a b2 based model, Public 0.68, Private 0.67. <br>\nWe finally chose an ensemble of 6 models, Public 0.7, Private 0.67 without post processing<br>\n<a href=\"https://www.kaggle.com/lihaoweicvch/lhw-final-en6-onnx\" target=\"_blank\">https://www.kaggle.com/lihaoweicvch/lhw-final-en6-onnx</a></p>",
      "rawMarkdown": "I'm very grateful to the organizers. This competition was very challenging for me, and due to good fortune, we secured the last gold medal position.\nLet me briefly report on our plan, no big difference between here: \nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/497539\n\n## Data\nBirdCLEF2021, 2022, 2023, 2024 (this competition), ff1010bird_nocall, we didn't adopt additional mp3. When it was introduced, the LB had dropped, even though I found it odd.\n\n## Model\nWe adopted the model that ranked 4th last year, but swapped the backbone with efficientnet b0/b1/b2.\n\n## Feature\nLog mel spec with:\n```\nCFG.n_mels = 128\nCFG.fmin = 20\nCFG.fmax = 16000\nCFG.n_fft = 2048\nCFG.hop_length = 512\nCFG.sample_rate = 32000\nCFG.secondary_coef = 1.0\n```\n\n## Data augmentation\nheavy augmentation improve my CV but LB drop.\n```\nA.HorizontalFlip(p=0.5),\nA.OneOf([\n                    A.Cutout(max_h_size=5, max_w_size=16),\n                    A.CoarseDropout(max_holes=4),\n], p=0.5),\n```\nand mixup v2:\n```\nclass MixupV2(nn.Module):\n    def __init__(self, mix_range=(0.3, 0.7), add_label=True):\n        super(MixupV2, self).__init__()\n        self.distribution = torch.distributions.Uniform(low=mix_range[0],\n                                                        high=mix_range[1])\n        self.add_label = add_label\n\n    def forward(self, X, Y, weight=None, teacher_preds=None):\n\n        bs = X.shape[0]\n        n_dims = len(X.shape)\n        perm = torch.randperm(bs)\n        coeffs = self.distribution.rsample(torch.Size((bs,))).to(X.device)\n\n        if n_dims == 2:\n            X = coeffs.view(-1, 1) * X + (1 - coeffs.view(-1, 1)) * X[perm]\n        elif n_dims == 3:\n            X = coeffs.view(-1, 1, 1) * X + (1 - coeffs.view(-1, 1, 1)) * X[perm]\n        else:\n            X = coeffs.view(-1, 1, 1, 1) * X + (1 - coeffs.view(-1, 1, 1, 1)) * X[perm]\n\n        if self.add_label:\n            Y = Y + Y[perm]\n            Y = torch.clamp(Y, 0, 1.0)\n        else:\n            Y = coeffs.view(-1, 1) * Y + (1 - coeffs.view(-1, 1)) * Y[perm]\n\n        if weight is None:\n            return X, Y\n        else:\n            weight = coeffs.view(-1) * weight + (1 - coeffs.view(-1)) * weight[perm]\n            if self.add_label:\n                teacher_preds = teacher_preds + teacher_preds[perm]\n                teacher_preds = torch.clamp(teacher_preds, 0, 1.0)\n            else:\n                teacher_preds = coeffs.view(-1, 1) * teacher_preds + (1 - coeffs.view(-1, 1)) * teacher_preds[perm]\n            return X, Y, weight, teacher_preds\n```\n\n\n## Training\nFirst, we pre-trained the first 15 seconds on BirdCLEF2021, 2022, 2023 but removed the classes that were included in the 2024 data to avoid leakage. Then, the first 5 seconds of the 2024 data were used to train the final model. This training plan is the most stable we have found.\n\n## Postprocessing\nSlightly improve the LB both Public and Private , My teammate will supplement further later.\n\n## Tried but not work\n- hubert and distill\n- distill or finetune birdnet\n- pseudo label\n- select multiple segs in an audio by model according to score (why? quite strange )\n\n## Notebook \nOur single best model is a b2 based model, Public 0.68, Private 0.67. \nWe finally chose an ensemble of 6 models, Public 0.7, Private 0.67 without post processing\nhttps://www.kaggle.com/lihaoweicvch/lhw-final-en6-onnx",
      "votes": 33
    },
    {
      "id": 2865945,
      "postDate": "2024-06-11T04:38:44.293Z",
      "content": "<p>I am very happy to be part of this competition with <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> ! I'll add a supplement here about postprocessing that he mentioned.</p>\n<h2>Postprocessing</h2>\n<p>In multi-class classification, it is common to apply the sigmoid to the logit. However, the method that does not use the sigmoid function was our best submission. Let's take a look at how to do that and the approach I took.</p>\n<h6>General Probability Weighted Average</h6>\n<p>This is the ensemble method we are most familiar with, the pseudocode is below:</p>\n<pre><code>pred = \n i, m in enumerate():\n    pred += m(x).() * model_weights[i]\n</code></pre>\n<h6>Sigmoid after weighting logits</h6>\n<p>This was used in our final submission. Both Public LB and Private LB were better than the general probability weighting.</p>\n<pre><code> = \n i, m  enumerate(models):\n     += m(x) * model_weights[i]\n = .sigmoid()\n</code></pre>\n<h6># MinMax Scaling and Powers</h6>\n<p>I discovered this through trial and error. When the evaluation metric is AUC, postprocessing techniques such as exponentiating probabilities and rank averaging are used. These techniques can be found in <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/211194\" target=\"_blank\">'s discussion</a> and <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <a href=\"https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/205564\" target=\"_blank\">'s discussion</a>. But neither of these worked in this competition.<br>\nI noticed that rank averaging doesn't take into account the interval between each prediction, so I tried using a scaler instead of ranking, and it worked.</p>\n<pre><code> = \n i, m  enumerate(models):\n     += MinMaxScaler(m(x)) ** p * model_weights[i]\n</code></pre>\n<p>Here, p ranges from 0 to 1. If CV cannot be trusted, as in this competition, we must search manually.<br>\nAlthough the score improvement was small, it may be used in more competitive situations.</p>\n<table>\n<thead>\n<tr>\n<th>method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Li's baseline</td>\n<td>0.705361</td>\n<td>0.674974</td>\n</tr>\n<tr>\n<td>minmax p=0.5</td>\n<td>0.705327</td>\n<td>0.675312</td>\n</tr>\n<tr>\n<td>minmax p=0.4</td>\n<td>0.705350</td>\n<td>0.675297</td>\n</tr>\n<tr>\n<td>minmax p=0.3</td>\n<td>0.705361</td>\n<td>0.675293</td>\n</tr>\n<tr>\n<td>minmax p=0.2</td>\n<td>0.705383</td>\n<td>0.675278</td>\n</tr>\n</tbody>\n</table>\n<h6>Conclusion</h6>\n<p>I wrote some additional information about postprocessing in the 11th place solution. The improvement was small, but with more experimentation we may have finished better. Of course, I am grateful to my teammates Li and the others for their efforts in creating a great model.<br>\nAnd finally, cheers to the competition hosts and participants!</p>",
      "rawMarkdown": "I am very happy to be part of this competition with @lihaoweicvch ! I'll add a supplement here about postprocessing that he mentioned.\n\n## Postprocessing\nIn multi-class classification, it is common to apply the sigmoid to the logit. However, the method that does not use the sigmoid function was our best submission. Let's take a look at how to do that and the approach I took.\n\n###### General Probability Weighted Average\nThis is the ensemble method we are most familiar with, the pseudocode is below:\n```\npred = 0\nfor i, m in enumerate(models):\n    pred += m(x).sigmoid() * model_weights[i]\n```\n\n###### Sigmoid after weighting logits\nThis was used in our final submission. Both Public LB and Private LB were better than the general probability weighting.\n```\npred = 0\nfor i, m in enumerate(models):\n    pred += m(x) * model_weights[i]\npred = pred.sigmoid()\n```\n\n####### MinMax Scaling and Powers\nI discovered this through trial and error. When the evaluation metric is AUC, postprocessing techniques such as exponentiating probabilities and rank averaging are used. These techniques can be found in @hengck23 ['s discussion](https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/211194) and @ttahara ['s discussion](https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/205564). But neither of these worked in this competition.\nI noticed that rank averaging doesn't take into account the interval between each prediction, so I tried using a scaler instead of ranking, and it worked.\n```\npred = 0\nfor i, m in enumerate(models):\n    pred += MinMaxScaler(m(x)) ** p * model_weights[i]\n```\nHere, p ranges from 0 to 1. If CV cannot be trusted, as in this competition, we must search manually.\nAlthough the score improvement was small, it may be used in more competitive situations.\n\n| method | Public LB | Private LB |\n| --- | --- | --- |\n| Li's baseline | 0.705361 | 0.674974 |\n| minmax p=0.5 | 0.705327 | 0.675312 |\n| minmax p=0.4 | 0.705350 | 0.675297 |\n| minmax p=0.3 | 0.705361 | 0.675293 |\n| minmax p=0.2 | 0.705383 | 0.675278 |\n\n###### Conclusion\nI wrote some additional information about postprocessing in the 11th place solution. The improvement was small, but with more experimentation we may have finished better. Of course, I am grateful to my teammates Li and the others for their efforts in creating a great model.\nAnd finally, cheers to the competition hosts and participants!",
      "votes": 5,
      "replies": [
        {
          "id": 2881876,
          "postDate": "2024-06-21T03:29:56.883Z",
          "content": "<p>Hi，may i ask why use sigmoid func after softmax layer? In my opinion, the output of softmax is the logits.</p>",
          "rawMarkdown": "Hi，may i ask why use sigmoid func after softmax layer? In my opinion, the output of softmax is the logits."
        }
      ]
    },
    {
      "id": 2865745,
      "postDate": "2024-06-11T00:59:30.523Z",
      "content": "<p>4th place from last year used knowledge distillation from bird-vocalization-classifier. Did you do that as well?</p>",
      "rawMarkdown": "4th place from last year used knowledge distillation from bird-vocalization-classifier. Did you do that as well?",
      "votes": 1,
      "replies": [
        {
          "id": 2865770,
          "postDate": "2024-06-11T01:22:30.600Z",
          "content": "<p>no improvement for me so I don't used knowledge distillation</p>",
          "rawMarkdown": "no improvement for me so I don't used knowledge distillation",
          "votes": 3,
          "replies": [
            {
              "id": 2873948,
              "postDate": "2024-06-15T21:53:22.980Z",
              "content": "<p>Got it! So you didn't use the \"teacher_preds\" part of the MixupV2 code that you posted?</p>",
              "rawMarkdown": "Got it! So you didn't use the \"teacher_preds\" part of the MixupV2 code that you posted?"
            }
          ]
        }
      ]
    },
    {
      "id": 2866264,
      "postDate": "2024-06-11T08:24:58.793Z",
      "content": "<p>Congratsss for gold medal!!</p>",
      "rawMarkdown": "Congratsss for gold medal!!"
    },
    {
      "id": 2865923,
      "postDate": "2024-06-11T04:30:20.503Z",
      "content": "<p>Congratulations on achieving 11th place in this competition. Thanks for sharing solution details and notebook. </p>",
      "rawMarkdown": "Congratulations on achieving 11th place in this competition. Thanks for sharing solution details and notebook. "
    },
    {
      "id": 2865777,
      "postDate": "2024-06-11T01:31:13.963Z",
      "content": "<p>Congratulations, it is great that you were able to work through the situation where LB and CV did not correlate to the end!<br>\nAlso, your discussion helped me to ensure the certainty of my experiment.</p>",
      "rawMarkdown": "Congratulations, it is great that you were able to work through the situation where LB and CV did not correlate to the end!\nAlso, your discussion helped me to ensure the certainty of my experiment."
    },
    {
      "id": 2865748,
      "postDate": "2024-06-11T01:03:43.233Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2865773,
          "postDate": "2024-06-11T01:25:16.380Z",
          "content": "<p>with 15s pre training,  CV can boost to 0.981~0.983 in my  exp,  LB no significant improvement but more stable,  0.65 to 0.68   vs   0.63--0.68 without pre training</p>",
          "rawMarkdown": "with 15s pre training,  CV can boost to 0.981~0.983 in my  exp,  LB no significant improvement but more stable,  0.65 to 0.68   vs   0.63--0.68 without pre training",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2865945,
      "author_name": "ITK8191",
      "author_url": "",
      "post_date": "2024-06-11T04:38:44.293000",
      "content": "<p>I am very happy to be part of this competition with <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> ! I'll add a supplement here about postprocessing that he mentioned.</p>\n<h2>Postprocessing</h2>\n<p>In multi-class classification, it is common to apply the sigmoid to the logit. However, the method that does not use the sigmoid function was our best submission. Let's take a look at how to do that and the approach I took.</p>\n<h6>General Probability Weighted Average</h6>\n<p>This is the ensemble method we are most familiar with, the pseudocode is below:</p>\n<pre><code>pred = \n i, m in enumerate():\n    pred += m(x).() * model_weights[i]\n</code></pre>\n<h6>Sigmoid after weighting logits</h6>\n<p>This was used in our final submission. Both Public LB and Private LB were better than the general probability weighting.</p>\n<pre><code> = \n i, m  enumerate(models):\n     += m(x) * model_weights[i]\n = .sigmoid()\n</code></pre>\n<h6># MinMax Scaling and Powers</h6>\n<p>I discovered this through trial and error. When the evaluation metric is AUC, postprocessing techniques such as exponentiating probabilities and rank averaging are used. These techniques can be found in <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/211194\" target=\"_blank\">'s discussion</a> and <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <a href=\"https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/205564\" target=\"_blank\">'s discussion</a>. But neither of these worked in this competition.<br>\nI noticed that rank averaging doesn't take into account the interval between each prediction, so I tried using a scaler instead of ranking, and it worked.</p>\n<pre><code> = \n i, m  enumerate(models):\n     += MinMaxScaler(m(x)) ** p * model_weights[i]\n</code></pre>\n<p>Here, p ranges from 0 to 1. If CV cannot be trusted, as in this competition, we must search manually.<br>\nAlthough the score improvement was small, it may be used in more competitive situations.</p>\n<table>\n<thead>\n<tr>\n<th>method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Li's baseline</td>\n<td>0.705361</td>\n<td>0.674974</td>\n</tr>\n<tr>\n<td>minmax p=0.5</td>\n<td>0.705327</td>\n<td>0.675312</td>\n</tr>\n<tr>\n<td>minmax p=0.4</td>\n<td>0.705350</td>\n<td>0.675297</td>\n</tr>\n<tr>\n<td>minmax p=0.3</td>\n<td>0.705361</td>\n<td>0.675293</td>\n</tr>\n<tr>\n<td>minmax p=0.2</td>\n<td>0.705383</td>\n<td>0.675278</td>\n</tr>\n</tbody>\n</table>\n<h6>Conclusion</h6>\n<p>I wrote some additional information about postprocessing in the 11th place solution. The improvement was small, but with more experimentation we may have finished better. Of course, I am grateful to my teammates Li and the others for their efforts in creating a great model.<br>\nAnd finally, cheers to the competition hosts and participants!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2881876,
          "author_name": "JonneryR",
          "author_url": "",
          "post_date": "2024-06-21T03:29:56.883000",
          "content": "<p>Hi，may i ask why use sigmoid func after softmax layer? In my opinion, the output of softmax is the logits.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2865745,
      "author_name": "thacrobatheskis",
      "author_url": "",
      "post_date": "2024-06-11T00:59:30.523000",
      "content": "<p>4th place from last year used knowledge distillation from bird-vocalization-classifier. Did you do that as well?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2865770,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-06-11T01:22:30.600000",
          "content": "<p>no improvement for me so I don't used knowledge distillation</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2873948,
              "author_name": "thacrobatheskis",
              "author_url": "",
              "post_date": "2024-06-15T21:53:22.980000",
              "content": "<p>Got it! So you didn't use the \"teacher_preds\" part of the MixupV2 code that you posted?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2866264,
      "author_name": "Aaditya Porwal",
      "author_url": "",
      "post_date": "2024-06-11T08:24:58.793000",
      "content": "<p>Congratsss for gold medal!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2865923,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-06-11T04:30:20.503000",
      "content": "<p>Congratulations on achieving 11th place in this competition. Thanks for sharing solution details and notebook. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2865777,
      "author_name": "tc0000",
      "author_url": "",
      "post_date": "2024-06-11T01:31:13.963000",
      "content": "<p>Congratulations, it is great that you were able to work through the situation where LB and CV did not correlate to the end!<br>\nAlso, your discussion helped me to ensure the certainty of my experiment.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2865748,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-11T01:03:43.233000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2865773,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-06-11T01:25:16.380000",
          "content": "<p>with 15s pre training,  CV can boost to 0.981~0.983 in my  exp,  LB no significant improvement but more stable,  0.65 to 0.68   vs   0.63--0.68 without pre training</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2865726": "I'm very grateful to the organizers. This competition was very challenging for me, and due to good fortune, we secured the last gold medal position.\nLet me briefly report on our plan, no big difference between here: \nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/497539\n\n## Data\nBirdCLEF2021, 2022, 2023, 2024 (this competition), ff1010bird_nocall, we didn't adopt additional mp3. When it was introduced, the LB had dropped, even though I found it odd.\n\n## Model\nWe adopted the model that ranked 4th last year, but swapped the backbone with efficientnet b0/b1/b2.\n\n## Feature\nLog mel spec with:\n```\nCFG.n_mels = 128\nCFG.fmin = 20\nCFG.fmax = 16000\nCFG.n_fft = 2048\nCFG.hop_length = 512\nCFG.sample_rate = 32000\nCFG.secondary_coef = 1.0\n```\n\n## Data augmentation\nheavy augmentation improve my CV but LB drop.\n```\nA.HorizontalFlip(p=0.5),\nA.OneOf([\n                    A.Cutout(max_h_size=5, max_w_size=16),\n                    A.CoarseDropout(max_holes=4),\n], p=0.5),\n```\nand mixup v2:\n```\nclass MixupV2(nn.Module):\n    def __init__(self, mix_range=(0.3, 0.7), add_label=True):\n        super(MixupV2, self).__init__()\n        self.distribution = torch.distributions.Uniform(low=mix_range[0],\n                                                        high=mix_range[1])\n        self.add_label = add_label\n\n    def forward(self, X, Y, weight=None, teacher_preds=None):\n\n        bs = X.shape[0]\n        n_dims = len(X.shape)\n        perm = torch.randperm(bs)\n        coeffs = self.distribution.rsample(torch.Size((bs,))).to(X.device)\n\n        if n_dims == 2:\n            X = coeffs.view(-1, 1) * X + (1 - coeffs.view(-1, 1)) * X[perm]\n        elif n_dims == 3:\n            X = coeffs.view(-1, 1, 1) * X + (1 - coeffs.view(-1, 1, 1)) * X[perm]\n        else:\n            X = coeffs.view(-1, 1, 1, 1) * X + (1 - coeffs.view(-1, 1, 1, 1)) * X[perm]\n\n        if self.add_label:\n            Y = Y + Y[perm]\n            Y = torch.clamp(Y, 0, 1.0)\n        else:\n            Y = coeffs.view(-1, 1) * Y + (1 - coeffs.view(-1, 1)) * Y[perm]\n\n        if weight is None:\n            return X, Y\n        else:\n            weight = coeffs.view(-1) * weight + (1 - coeffs.view(-1)) * weight[perm]\n            if self.add_label:\n                teacher_preds = teacher_preds + teacher_preds[perm]\n                teacher_preds = torch.clamp(teacher_preds, 0, 1.0)\n            else:\n                teacher_preds = coeffs.view(-1, 1) * teacher_preds + (1 - coeffs.view(-1, 1)) * teacher_preds[perm]\n            return X, Y, weight, teacher_preds\n```\n\n\n## Training\nFirst, we pre-trained the first 15 seconds on BirdCLEF2021, 2022, 2023 but removed the classes that were included in the 2024 data to avoid leakage. Then, the first 5 seconds of the 2024 data were used to train the final model. This training plan is the most stable we have found.\n\n## Postprocessing\nSlightly improve the LB both Public and Private , My teammate will supplement further later.\n\n## Tried but not work\n- hubert and distill\n- distill or finetune birdnet\n- pseudo label\n- select multiple segs in an audio by model according to score (why? quite strange )\n\n## Notebook \nOur single best model is a b2 based model, Public 0.68, Private 0.67. \nWe finally chose an ensemble of 6 models, Public 0.7, Private 0.67 without post processing\nhttps://www.kaggle.com/lihaoweicvch/lhw-final-en6-onnx",
    "2865945": "I am very happy to be part of this competition with @lihaoweicvch ! I'll add a supplement here about postprocessing that he mentioned.\n\n## Postprocessing\nIn multi-class classification, it is common to apply the sigmoid to the logit. However, the method that does not use the sigmoid function was our best submission. Let's take a look at how to do that and the approach I took.\n\n###### General Probability Weighted Average\nThis is the ensemble method we are most familiar with, the pseudocode is below:\n```\npred = 0\nfor i, m in enumerate(models):\n    pred += m(x).sigmoid() * model_weights[i]\n```\n\n###### Sigmoid after weighting logits\nThis was used in our final submission. Both Public LB and Private LB were better than the general probability weighting.\n```\npred = 0\nfor i, m in enumerate(models):\n    pred += m(x) * model_weights[i]\npred = pred.sigmoid()\n```\n\n####### MinMax Scaling and Powers\nI discovered this through trial and error. When the evaluation metric is AUC, postprocessing techniques such as exponentiating probabilities and rank averaging are used. These techniques can be found in @hengck23 ['s discussion](https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/211194) and @ttahara ['s discussion](https://www.kaggle.com/competitions/ranzcr-clip-catheter-line-classification/discussion/205564). But neither of these worked in this competition.\nI noticed that rank averaging doesn't take into account the interval between each prediction, so I tried using a scaler instead of ranking, and it worked.\n```\npred = 0\nfor i, m in enumerate(models):\n    pred += MinMaxScaler(m(x)) ** p * model_weights[i]\n```\nHere, p ranges from 0 to 1. If CV cannot be trusted, as in this competition, we must search manually.\nAlthough the score improvement was small, it may be used in more competitive situations.\n\n| method | Public LB | Private LB |\n| --- | --- | --- |\n| Li's baseline | 0.705361 | 0.674974 |\n| minmax p=0.5 | 0.705327 | 0.675312 |\n| minmax p=0.4 | 0.705350 | 0.675297 |\n| minmax p=0.3 | 0.705361 | 0.675293 |\n| minmax p=0.2 | 0.705383 | 0.675278 |\n\n###### Conclusion\nI wrote some additional information about postprocessing in the 11th place solution. The improvement was small, but with more experimentation we may have finished better. Of course, I am grateful to my teammates Li and the others for their efforts in creating a great model.\nAnd finally, cheers to the competition hosts and participants!",
    "2865745": "4th place from last year used knowledge distillation from bird-vocalization-classifier. Did you do that as well?",
    "2866264": "Congratsss for gold medal!!",
    "2865923": "Congratulations on achieving 11th place in this competition. Thanks for sharing solution details and notebook. ",
    "2865777": "Congratulations, it is great that you were able to work through the situation where LB and CV did not correlate to the end!\nAlso, your discussion helped me to ensure the certainty of my experiment.",
    "2865748": ""
  }
}