{
  "id": 511527,
  "title": "6th Place Solution",
  "url": "/competitions/birdclef-2024/writeups/penguin46-6th-place-solution",
  "author_name": "",
  "post_date": "2024-06-12T14:21:55.257Z",
  "votes": 36,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First, I would like to thank the hosts for organizing such a very interesting competition and the kaggle team for facilitating it so smoothly. Also, congratulations to all the top winners.</p>\n<p>My solution is based on <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">the last year's 2nd place solution</a> by <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a>. I would like to thank him greatly. The difference between his last year's solution and mine is mainly pseudo-label and post-processing. </p>\n<h2>CV Strategy</h2>\n<p>CV was not helpful for me, so I tuned my models by checking public LB. To avoid overfitting to public LB, I reduced submissions and gave up tuning minor hyperparams.</p>\n<p>One of the final submissions is my public best model. Using additional training data made public LB about 0.02 worse, but this was counter-intuitive and I expected these to be reversed in private LB, so I selected it for the other one.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F5a22a50b7b3ebf854a2b62f3d6c46b49%2Flb.png?generation=1718083926626187&amp;alt=media\"></p>\n<h2>Pseudo Labeling</h2>\n<p><code>model.forward()</code> method input several pieces of 5 seconds together, but calculate the loss without aggregation when using pseudo-labeled data, and after aggregation with max for competition data. Hard labeling by threshold did not improve scores.</p>\n<p>This improved by 0.02~0.03 in publc and 0.015~0.04 in private before ensemble.</p>\n<pre><code>\ntotal_loss = \n i  (bs):\n    start = i * self.factor\n    end = (i + ) * self.factor\n\n     is_pseudo[i] == :\n        this_logits = torch.(logits[start:end], dim=, keepdim=).values\n        this_y = torch.(y[i], dim=, keepdim=).values\n        this_weight = torch.(weight[i], dim=, keepdim=).values\n\n        loss = self.loss_function(this_logits, this_y)\n         self.loss == :  \n            loss = (loss * this_weight) / weight.()\n         self.loss == :  \n            loss = (loss.(dim=) * this_weight) / weight.()\n        :\n             NotImplementedError\n        loss = loss.() * self.factor\n        total_loss += loss\n    :\n        this_logits = logits[start:end]\n        this_y = y[i]\n        this_weight = weight[i]\n\n        loss = self.loss_function(this_logits, this_y)\n         self.loss == :  \n            loss = (loss * this_weight) / weight.()\n         self.loss == :  \n            loss = (loss.(dim=) * this_weight) / weight.()\n        :\n             NotImplementedError\n        loss = loss.() / self.factor\n        total_loss += loss\n</code></pre>\n<h2>Post Processing</h2>\n<p>Calculate a moving average with weights of [0.1, 0.2, 0.4, 0.2, 0.1] and finally add the global average * 0.2 for each species.</p>\n<p>This consistently boosted public and private LB by 0.014~0.016 in both final submissions.</p>\n<pre><code> smooth_array_general(array, w=[., ., ., ., .]):\n     = np.zeros_like(array)\n     = array.shape[]\n     = len(w) // \n\n     t in range(timesteps):\n         i, weight in enumerate(w):\n             = t - radius + i\n             index &lt; : \n                [t] += array[] * weight\n             index &gt;= timesteps: \n                [t] += array[-] * weight\n            :\n                [t] += array[index] * weight\n     c in range(array.shape[]):\n        [:, c] = smoothed_array[:, c] * . + smoothed_array[:, c].mean(keepdims=True) * .\n     smoothed_array\n</code></pre>\n<h2>Other Setting</h2>\n<ul>\n<li>resampling became a bottleneck, so preprocess sampled to 32 kHz and storing them on disk helps speed up the training phase.</li>\n<li>sampling by RMS (+0.010)<ul>\n<li>RMS sampling &gt; using first 5sec &gt; random sampling</li></ul></li>\n<li>backbone<ul>\n<li>resnet18d, resnet34d and efficientnetv2s</li></ul></li>\n</ul>\n<p>Thank you for reading</p>",
  "messages": [
    {
      "id": "2865989",
      "postDate": "06/11/2024 04:58:36",
      "content": "<p>First, I would like to thank the hosts for organizing such a very interesting competition and the kaggle team for facilitating it so smoothly. Also, congratulations to all the top winners.</p>\n<p>My solution is based on <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412707\" target=\"_blank\">the last year's 2nd place solution</a> by <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a>. I would like to thank him greatly. The difference between his last year's solution and mine is mainly pseudo-label and post-processing. </p>\n<h2>CV Strategy</h2>\n<p>CV was not helpful for me, so I tuned my models by checking public LB. To avoid overfitting to public LB, I reduced submissions and gave up tuning minor hyperparams.</p>\n<p>One of the final submissions is my public best model. Using additional training data made public LB about 0.02 worse, but this was counter-intuitive and I expected these to be reversed in private LB, so I selected it for the other one.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F5a22a50b7b3ebf854a2b62f3d6c46b49%2Flb.png?generation=1718083926626187&amp;alt=media\"></p>\n<h2>Pseudo Labeling</h2>\n<p><code>model.forward()</code> method input several pieces of 5 seconds together, but calculate the loss without aggregation when using pseudo-labeled data, and after aggregation with max for competition data. Hard labeling by threshold did not improve scores.</p>\n<p>This improved by 0.02~0.03 in publc and 0.015~0.04 in private before ensemble.</p>\n<pre><code>\ntotal_loss = \n i  (bs):\n    start = i * self.factor\n    end = (i + ) * self.factor\n\n     is_pseudo[i] == :\n        this_logits = torch.(logits[start:end], dim=, keepdim=).values\n        this_y = torch.(y[i], dim=, keepdim=).values\n        this_weight = torch.(weight[i], dim=, keepdim=).values\n\n        loss = self.loss_function(this_logits, this_y)\n         self.loss == :  \n            loss = (loss * this_weight) / weight.()\n         self.loss == :  \n            loss = (loss.(dim=) * this_weight) / weight.()\n        :\n             NotImplementedError\n        loss = loss.() * self.factor\n        total_loss += loss\n    :\n        this_logits = logits[start:end]\n        this_y = y[i]\n        this_weight = weight[i]\n\n        loss = self.loss_function(this_logits, this_y)\n         self.loss == :  \n            loss = (loss * this_weight) / weight.()\n         self.loss == :  \n            loss = (loss.(dim=) * this_weight) / weight.()\n        :\n             NotImplementedError\n        loss = loss.() / self.factor\n        total_loss += loss\n</code></pre>\n<h2>Post Processing</h2>\n<p>Calculate a moving average with weights of [0.1, 0.2, 0.4, 0.2, 0.1] and finally add the global average * 0.2 for each species.</p>\n<p>This consistently boosted public and private LB by 0.014~0.016 in both final submissions.</p>\n<pre><code> smooth_array_general(array, w=[., ., ., ., .]):\n     = np.zeros_like(array)\n     = array.shape[]\n     = len(w) // \n\n     t in range(timesteps):\n         i, weight in enumerate(w):\n             = t - radius + i\n             index &lt; : \n                [t] += array[] * weight\n             index &gt;= timesteps: \n                [t] += array[-] * weight\n            :\n                [t] += array[index] * weight\n     c in range(array.shape[]):\n        [:, c] = smoothed_array[:, c] * . + smoothed_array[:, c].mean(keepdims=True) * .\n     smoothed_array\n</code></pre>\n<h2>Other Setting</h2>\n<ul>\n<li>resampling became a bottleneck, so preprocess sampled to 32 kHz and storing them on disk helps speed up the training phase.</li>\n<li>sampling by RMS (+0.010)<ul>\n<li>RMS sampling &gt; using first 5sec &gt; random sampling</li></ul></li>\n<li>backbone<ul>\n<li>resnet18d, resnet34d and efficientnetv2s</li></ul></li>\n</ul>\n<p>Thank you for reading</p>",
      "rawMarkdown": "First, I would like to thank the hosts for organizing such a very interesting competition and the kaggle team for facilitating it so smoothly. Also, congratulations to all the top winners.\n\nMy solution is based on [the last year's 2nd place solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707) by @honglihang. I would like to thank him greatly. The difference between his last year's solution and mine is mainly pseudo-label and post-processing. \n\n## CV Strategy\n\nCV was not helpful for me, so I tuned my models by checking public LB. To avoid overfitting to public LB, I reduced submissions and gave up tuning minor hyperparams.\n\nOne of the final submissions is my public best model. Using additional training data made public LB about 0.02 worse, but this was counter-intuitive and I expected these to be reversed in private LB, so I selected it for the other one.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F5a22a50b7b3ebf854a2b62f3d6c46b49%2Flb.png?generation=1718083926626187&alt=media)\n\n## Pseudo Labeling\n\n`model.forward()` method input several pieces of 5 seconds together, but calculate the loss without aggregation when using pseudo-labeled data, and after aggregation with max for competition data. Hard labeling by threshold did not improve scores.\n\nThis improved by 0.02~0.03 in publc and 0.015~0.04 in private before ensemble.\n\n```\n\"\"\" model.forward() \"\"\"\ntotal_loss = 0\nfor i in range(bs):\n    start = i * self.factor\n    end = (i + 1) * self.factor\n\n    if is_pseudo[i] == 0:\n        this_logits = torch.max(logits[start:end], dim=0, keepdim=True).values\n        this_y = torch.max(y[i], dim=0, keepdim=True).values\n        this_weight = torch.max(weight[i], dim=0, keepdim=True).values\n\n        loss = self.loss_function(this_logits, this_y)\n        if self.loss == \"ce\":  # loss: (n_sample, )\n            loss = (loss * this_weight) / weight.sum()\n        elif self.loss == \"bce\":  # loss: (n_sample, n_class)\n            loss = (loss.sum(dim=1) * this_weight) / weight.sum()\n        else:\n            raise NotImplementedError\n        loss = loss.sum() * self.factor\n        total_loss += loss\n    else:\n        this_logits = logits[start:end]\n        this_y = y[i]\n        this_weight = weight[i]\n\n        loss = self.loss_function(this_logits, this_y)\n        if self.loss == \"ce\":  # loss: (n_sample, )\n            loss = (loss * this_weight) / weight.sum()\n        elif self.loss == \"bce\":  # loss: (n_sample, n_class)\n            loss = (loss.sum(dim=1) * this_weight) / weight.sum()\n        else:\n            raise NotImplementedError\n        loss = loss.sum() / self.factor\n        total_loss += loss\n```\n\n## Post Processing\n\nCalculate a moving average with weights of [0.1, 0.2, 0.4, 0.2, 0.1] and finally add the global average * 0.2 for each species.\n\nThis consistently boosted public and private LB by 0.014~0.016 in both final submissions.\n\n```\ndef smooth_array_general(array, w=[0.1, 0.2, 0.4, 0.2, 0.1]):\n    smoothed_array = np.zeros_like(array)\n    timesteps = array.shape[0]\n    radius = len(w) // 2\n\n    for t in range(timesteps):\n        for i, weight in enumerate(w):\n            index = t - radius + i\n            if index < 0: \n                smoothed_array[t] += array[0] * weight\n            elif index >= timesteps: \n                smoothed_array[t] += array[-1] * weight\n            else:\n                smoothed_array[t] += array[index] * weight\n    for c in range(array.shape[1]):\n        smoothed_array[:, c] = smoothed_array[:, c] * 0.8 + smoothed_array[:, c].mean(keepdims=True) * 0.2\n    return smoothed_array\n```\n\n## Other Setting\n\n- resampling became a bottleneck, so preprocess sampled to 32 kHz and storing them on disk helps speed up the training phase.\n- sampling by RMS (+0.010)\n  - RMS sampling > using first 5sec > random sampling\n- backbone\n  - resnet18d, resnet34d and efficientnetv2s\n\nThank you for reading",
      "votes": null
    },
    {
      "id": "2865991",
      "postDate": "06/11/2024 05:02:26",
      "content": "<p>Congratulations on achieving 6th place in this competition. Appreciation for sharing your solution details. </p>",
      "rawMarkdown": "Congratulations on achieving 6th place in this competition. Appreciation for sharing your solution details.",
      "votes": null
    },
    {
      "id": "2866032",
      "postDate": "06/11/2024 05:50:59",
      "content": "<p>Congratulations for the 6th place! I am happy that the training and inference tricks I used last year also help this year!</p>",
      "rawMarkdown": "Congratulations for the 6th place! I am happy that the training and inference tricks I used last year also help this year!",
      "votes": null
    },
    {
      "id": "2866045",
      "postDate": "06/11/2024 05:58:53",
      "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> <br>\nCongratulations to you too on your gold medal! I was really helped by your last year's solution, so I am glad that we both stayed in the gold zone!</p>",
      "rawMarkdown": "honglihang \nCongratulations to you too on your gold medal! I was really helped by your last year's solution, so I am glad that we both stayed in the gold zone!",
      "votes": null
    },
    {
      "id": "2866101",
      "postDate": "06/11/2024 06:22:02",
      "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> <br>\nYes, this year's birdclef is really a hard one, glad to see that we both survived the shake with a relatively robust solution.<br>\nI tried moving average last year but it decreased LB in last year's metrics… emmmm, sad T T</p>",
      "rawMarkdown": "ryotayoshinobu \nYes, this year's birdclef is really a hard one, glad to see that we both survived the shake with a relatively robust solution.\nI tried moving average last year but it decreased LB in last year's metrics... emmmm, sad T T",
      "votes": null
    },
    {
      "id": "2866223",
      "postDate": "06/11/2024 07:59:49",
      "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> <br>\nIt is also great to see you have made the pseudo work! I noticed that pseudo didn't work with hard label last year so I didn't use it this year. Instead, I extracted logit using birdnet and implemented knowledge distillation. </p>\n<p>There are two places that we should deal with logit. 1 is above you shown, 2 is mixup.<br>\nIn 1, I simply averaged the logit and didn't try taking maximum, although maximum seems to be more correct.  <br>\nHowever, when 2 performing mixup, a wierd thing is that compared to taking maximum of the logit, simply averaging the logit performed better on LB for me.</p>\n<p>Maybe we can have a further dive into how to deal with the logit. How did you deal with logit in mixup?</p>",
      "rawMarkdown": "ryotayoshinobu \nIt is also great to see you have made the pseudo work! I noticed that pseudo didn't work with hard label last year so I didn't use it this year. Instead, I extracted logit using birdnet and implemented knowledge distillation. \n\nThere are two places that we should deal with logit. 1 is above you shown, 2 is mixup.\nIn 1, I simply averaged the logit and didn't try taking maximum, although maximum seems to be more correct.  \nHowever, when 2 performing mixup, a wierd thing is that compared to taking maximum of the logit, simply averaging the logit performed better on LB for me.\n\nMaybe we can have a further dive into how to deal with the logit. How did you deal with logit in mixup?",
      "votes": null
    },
    {
      "id": "2866244",
      "postDate": "06/11/2024 08:11:01",
      "content": "<p>Is birdnet a better model to use, compared to bird-vocalization-classifier? I tried knowledge distillation with bird-vocalization-classifier, model overfits badly.</p>",
      "rawMarkdown": "Is birdnet a better model to use, compared to bird-vocalization-classifier? I tried knowledge distillation with bird-vocalization-classifier, model overfits badly.",
      "votes": null
    },
    {
      "id": "2866256",
      "postDate": "06/11/2024 08:21:43",
      "content": "<p>Congrats!! for the 6th place.</p>",
      "rawMarkdown": "Congrats!! for the 6th place.",
      "votes": null
    },
    {
      "id": "2866286",
      "postDate": "06/11/2024 08:31:14",
      "content": "<p>This year, for me, Yes. I also added samples extracted from bird-vocalization-classifier, which has no overlap with audios extracted by birdnet.  But the LB decrease.</p>\n<p>But bird-vocalization-classifier seemed to be useful last year. Maybe we should take a look at model version…</p>",
      "rawMarkdown": "This year, for me, Yes. I also added samples extracted from bird-vocalization-classifier, which has no overlap with audios extracted by birdnet.  But the LB decrease.\n\nBut bird-vocalization-classifier seemed to be useful last year. Maybe we should take a look at model version...",
      "votes": null
    },
    {
      "id": "2866577",
      "postDate": "06/11/2024 11:47:14",
      "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a><br>\nOh, I should have made that comparison. I have not tuned Mixup closely, so I used it exactly as you used and did not make any comparison.</p>",
      "rawMarkdown": "honglihang\nOh, I should have made that comparison. I have not tuned Mixup closely, so I used it exactly as you used and did not make any comparison.",
      "votes": null
    },
    {
      "id": "2867627",
      "postDate": "06/12/2024 01:39:41",
      "content": "<p>Thanks for your reply! I will make a further investigation after get some sleep.</p>",
      "rawMarkdown": "Thanks for your reply! I will make a further investigation after get some sleep.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2865991,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "06/11/2024 05:02:26",
      "content": "<p>Congratulations on achieving 6th place in this competition. Appreciation for sharing your solution details. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2866032,
      "author_name": "honglihang",
      "author_url": "",
      "post_date": "06/11/2024 05:50:59",
      "content": "<p>Congratulations for the 6th place! I am happy that the training and inference tricks I used last year also help this year!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2866045,
          "author_name": "ryotayoshinobu",
          "author_url": "",
          "post_date": "06/11/2024 05:58:53",
          "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> <br>\nCongratulations to you too on your gold medal! I was really helped by your last year's solution, so I am glad that we both stayed in the gold zone!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2866101,
              "author_name": "honglihang",
              "author_url": "",
              "post_date": "06/11/2024 06:22:02",
              "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> <br>\nYes, this year's birdclef is really a hard one, glad to see that we both survived the shake with a relatively robust solution.<br>\nI tried moving average last year but it decreased LB in last year's metrics… emmmm, sad T T</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2866223,
              "author_name": "honglihang",
              "author_url": "",
              "post_date": "06/11/2024 07:59:49",
              "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> <br>\nIt is also great to see you have made the pseudo work! I noticed that pseudo didn't work with hard label last year so I didn't use it this year. Instead, I extracted logit using birdnet and implemented knowledge distillation. </p>\n<p>There are two places that we should deal with logit. 1 is above you shown, 2 is mixup.<br>\nIn 1, I simply averaged the logit and didn't try taking maximum, although maximum seems to be more correct.  <br>\nHowever, when 2 performing mixup, a wierd thing is that compared to taking maximum of the logit, simply averaging the logit performed better on LB for me.</p>\n<p>Maybe we can have a further dive into how to deal with the logit. How did you deal with logit in mixup?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2866244,
                  "author_name": "aphysict",
                  "author_url": "",
                  "post_date": "06/11/2024 08:11:01",
                  "content": "<p>Is birdnet a better model to use, compared to bird-vocalization-classifier? I tried knowledge distillation with bird-vocalization-classifier, model overfits badly.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2866286,
                      "author_name": "honglihang",
                      "author_url": "",
                      "post_date": "06/11/2024 08:31:14",
                      "content": "<p>This year, for me, Yes. I also added samples extracted from bird-vocalization-classifier, which has no overlap with audios extracted by birdnet.  But the LB decrease.</p>\n<p>But bird-vocalization-classifier seemed to be useful last year. Maybe we should take a look at model version…</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                },
                {
                  "id": 2866577,
                  "author_name": "ryotayoshinobu",
                  "author_url": "",
                  "post_date": "06/11/2024 11:47:14",
                  "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a><br>\nOh, I should have made that comparison. I have not tuned Mixup closely, so I used it exactly as you used and did not make any comparison.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2867627,
                      "author_name": "honglihang",
                      "author_url": "",
                      "post_date": "06/12/2024 01:39:41",
                      "content": "<p>Thanks for your reply! I will make a further investigation after get some sleep.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2866256,
      "author_name": "aadityaporwal",
      "author_url": "",
      "post_date": "06/11/2024 08:21:43",
      "content": "<p>Congrats!! for the 6th place.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2865989": "First, I would like to thank the hosts for organizing such a very interesting competition and the kaggle team for facilitating it so smoothly. Also, congratulations to all the top winners.\n\nMy solution is based on [the last year's 2nd place solution](https://www.kaggle.com/competitions/birdclef-2023/discussion/412707) by @honglihang. I would like to thank him greatly. The difference between his last year's solution and mine is mainly pseudo-label and post-processing. \n\n## CV Strategy\n\nCV was not helpful for me, so I tuned my models by checking public LB. To avoid overfitting to public LB, I reduced submissions and gave up tuning minor hyperparams.\n\nOne of the final submissions is my public best model. Using additional training data made public LB about 0.02 worse, but this was counter-intuitive and I expected these to be reversed in private LB, so I selected it for the other one.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F5a22a50b7b3ebf854a2b62f3d6c46b49%2Flb.png?generation=1718083926626187&alt=media)\n\n## Pseudo Labeling\n\n`model.forward()` method input several pieces of 5 seconds together, but calculate the loss without aggregation when using pseudo-labeled data, and after aggregation with max for competition data. Hard labeling by threshold did not improve scores.\n\nThis improved by 0.02~0.03 in publc and 0.015~0.04 in private before ensemble.\n\n```\n\"\"\" model.forward() \"\"\"\ntotal_loss = 0\nfor i in range(bs):\n    start = i * self.factor\n    end = (i + 1) * self.factor\n\n    if is_pseudo[i] == 0:\n        this_logits = torch.max(logits[start:end], dim=0, keepdim=True).values\n        this_y = torch.max(y[i], dim=0, keepdim=True).values\n        this_weight = torch.max(weight[i], dim=0, keepdim=True).values\n\n        loss = self.loss_function(this_logits, this_y)\n        if self.loss == \"ce\":  # loss: (n_sample, )\n            loss = (loss * this_weight) / weight.sum()\n        elif self.loss == \"bce\":  # loss: (n_sample, n_class)\n            loss = (loss.sum(dim=1) * this_weight) / weight.sum()\n        else:\n            raise NotImplementedError\n        loss = loss.sum() * self.factor\n        total_loss += loss\n    else:\n        this_logits = logits[start:end]\n        this_y = y[i]\n        this_weight = weight[i]\n\n        loss = self.loss_function(this_logits, this_y)\n        if self.loss == \"ce\":  # loss: (n_sample, )\n            loss = (loss * this_weight) / weight.sum()\n        elif self.loss == \"bce\":  # loss: (n_sample, n_class)\n            loss = (loss.sum(dim=1) * this_weight) / weight.sum()\n        else:\n            raise NotImplementedError\n        loss = loss.sum() / self.factor\n        total_loss += loss\n```\n\n## Post Processing\n\nCalculate a moving average with weights of [0.1, 0.2, 0.4, 0.2, 0.1] and finally add the global average * 0.2 for each species.\n\nThis consistently boosted public and private LB by 0.014~0.016 in both final submissions.\n\n```\ndef smooth_array_general(array, w=[0.1, 0.2, 0.4, 0.2, 0.1]):\n    smoothed_array = np.zeros_like(array)\n    timesteps = array.shape[0]\n    radius = len(w) // 2\n\n    for t in range(timesteps):\n        for i, weight in enumerate(w):\n            index = t - radius + i\n            if index < 0: \n                smoothed_array[t] += array[0] * weight\n            elif index >= timesteps: \n                smoothed_array[t] += array[-1] * weight\n            else:\n                smoothed_array[t] += array[index] * weight\n    for c in range(array.shape[1]):\n        smoothed_array[:, c] = smoothed_array[:, c] * 0.8 + smoothed_array[:, c].mean(keepdims=True) * 0.2\n    return smoothed_array\n```\n\n## Other Setting\n\n- resampling became a bottleneck, so preprocess sampled to 32 kHz and storing them on disk helps speed up the training phase.\n- sampling by RMS (+0.010)\n  - RMS sampling > using first 5sec > random sampling\n- backbone\n  - resnet18d, resnet34d and efficientnetv2s\n\nThank you for reading",
    "2865991": "Congratulations on achieving 6th place in this competition. Appreciation for sharing your solution details.",
    "2866032": "Congratulations for the 6th place! I am happy that the training and inference tricks I used last year also help this year!",
    "2866045": "honglihang \nCongratulations to you too on your gold medal! I was really helped by your last year's solution, so I am glad that we both stayed in the gold zone!",
    "2866101": "ryotayoshinobu \nYes, this year's birdclef is really a hard one, glad to see that we both survived the shake with a relatively robust solution.\nI tried moving average last year but it decreased LB in last year's metrics... emmmm, sad T T",
    "2866223": "ryotayoshinobu \nIt is also great to see you have made the pseudo work! I noticed that pseudo didn't work with hard label last year so I didn't use it this year. Instead, I extracted logit using birdnet and implemented knowledge distillation. \n\nThere are two places that we should deal with logit. 1 is above you shown, 2 is mixup.\nIn 1, I simply averaged the logit and didn't try taking maximum, although maximum seems to be more correct.  \nHowever, when 2 performing mixup, a wierd thing is that compared to taking maximum of the logit, simply averaging the logit performed better on LB for me.\n\nMaybe we can have a further dive into how to deal with the logit. How did you deal with logit in mixup?",
    "2866244": "Is birdnet a better model to use, compared to bird-vocalization-classifier? I tried knowledge distillation with bird-vocalization-classifier, model overfits badly.",
    "2866256": "Congrats!! for the 6th place.",
    "2866286": "This year, for me, Yes. I also added samples extracted from bird-vocalization-classifier, which has no overlap with audios extracted by birdnet.  But the LB decrease.\n\nBut bird-vocalization-classifier seemed to be useful last year. Maybe we should take a look at model version...",
    "2866577": "honglihang\nOh, I should have made that comparison. I have not tuned Mixup closely, so I used it exactly as you used and did not make any comparison.",
    "2867627": "Thanks for your reply! I will make a further investigation after get some sleep."
  },
  "source": "meta"
}