{
  "id": 580765,
  "title": "Same pipeline different backbones, LB score vary very large",
  "url": "/competitions/birdclef-2025/discussion/580765",
  "author_name": "",
  "post_date": "2025-05-26T11:21:15.086992800Z",
  "votes": 8,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Same train and inference pipeline, efv2b3 scored 0.902, seresnext26t scored 0.861, seresnext26d scored ~0.87.  I expect there will be differences, but they shouldn't be that significant. Has anyone encountered similar problem?</p>",
  "messages": [
    {
      "id": "3209818",
      "postDate": "05/26/2025 11:21:15",
      "content": "<p>Same train and inference pipeline, efv2b3 scored 0.902, seresnext26t scored 0.861, seresnext26d scored ~0.87.  I expect there will be differences, but they shouldn't be that significant. Has anyone encountered similar problem?</p>",
      "rawMarkdown": "Same train and inference pipeline, efv2b3 scored 0.902, seresnext26t scored 0.861, seresnext26d scored ~0.87.  I expect there will be differences, but they shouldn't be that significant. Has anyone encountered similar problem?",
      "votes": null
    },
    {
      "id": "3209868",
      "postDate": "05/26/2025 12:34:01",
      "content": "<p>I mean training is validated on 10K+ samples, and your public LB score is based on 33% of 700 samples. I expect big fluctuations.</p>",
      "rawMarkdown": "I mean training is validated on 10K+ samples, and your public LB score is based on 33% of 700 samples. I expect big fluctuations.",
      "votes": null
    },
    {
      "id": "3209874",
      "postDate": "05/26/2025 12:44:25",
      "content": "<p>In this competition，local cv Is meaningless. The score I mention above are LB score， just different feature extractors.</p>",
      "rawMarkdown": "In this competition，local cv Is meaningless. The score I mention above are LB score， just different feature extractors.",
      "votes": null
    },
    {
      "id": "3209883",
      "postDate": "05/26/2025 12:58:16",
      "content": "<p>I have the same problem. There are significant differences among different models</p>",
      "rawMarkdown": "I have the same problem. There are significant differences among different models",
      "votes": null
    },
    {
      "id": "3209936",
      "postDate": "05/26/2025 14:21:53",
      "content": "<p>Do you get these results at the same epoch?  For me it has fluctuated within 1.5%</p>",
      "rawMarkdown": "Do you get these results at the same epoch?  For me it has fluctuated within 1.5%",
      "votes": null
    },
    {
      "id": "3209964",
      "postDate": "05/26/2025 14:57:59",
      "content": "<p>same issue</p>",
      "rawMarkdown": "same issue",
      "votes": null
    },
    {
      "id": "3210050",
      "postDate": "05/26/2025 16:38:54",
      "content": "<p>More results: same pipeline and hyperparameters, efv2b3-0.902 resnext26t-0.861 efb2-0.858 rexnet150-0.857. It's very confusing.🥱</p>",
      "rawMarkdown": "More results: same pipeline and hyperparameters, efv2b3-0.902 resnext26t-0.861 efb2-0.858 rexnet150-0.857. It's very confusing.🥱",
      "votes": null
    },
    {
      "id": "3210052",
      "postDate": "05/26/2025 16:39:56",
      "content": "<p>yep, all models are trained for 40 epochs</p>",
      "rawMarkdown": "yep, all models are trained for 40 epochs",
      "votes": null
    },
    {
      "id": "3210152",
      "postDate": "05/26/2025 19:55:36",
      "content": "<p>I think if you test different epoch for each model or use the model soup, the gap won't be too large like this. But that's just like lottery 😂</p>",
      "rawMarkdown": "I think if you test different epoch for each model or use the model soup, the gap won't be too large like this. But that's just like lottery 😂",
      "votes": null
    },
    {
      "id": "3210217",
      "postDate": "05/26/2025 22:18:44",
      "content": "<p>is this without pseudo-labeled data?</p>",
      "rawMarkdown": "is this without pseudo-labeled data?",
      "votes": null
    },
    {
      "id": "3210283",
      "postDate": "05/27/2025 02:09:03",
      "content": "<p>with pseudo-labeled data</p>",
      "rawMarkdown": "with pseudo-labeled data",
      "votes": null
    },
    {
      "id": "3210286",
      "postDate": "05/27/2025 02:10:47",
      "content": "<p>Thank you for your advice, I would have a try.</p>",
      "rawMarkdown": "Thank you for your advice, I would have a try.",
      "votes": null
    },
    {
      "id": "3210294",
      "postDate": "05/27/2025 02:25:44",
      "content": "<p>Same here. Backbones like seresnext26d, v2b3, v2s, eca, show different lb and great gap. I apply ema in the training.</p>",
      "rawMarkdown": "Same here. Backbones like seresnext26d, v2b3, v2s, eca, show different lb and great gap. I apply ema in the training.",
      "votes": null
    },
    {
      "id": "3210303",
      "postDate": "05/27/2025 02:45:37",
      "content": "<p>May I ask how do ema perform? Dose it boost your LB score or make your results more stable?</p>",
      "rawMarkdown": "May I ask how do ema perform? Dose it boost your LB score or make your results more stable?",
      "votes": null
    },
    {
      "id": "3210306",
      "postDate": "05/27/2025 02:53:34",
      "content": "<p>Could I ask what is your lb score without pseudo-labeled data for the same pipeline?</p>",
      "rawMarkdown": "Could I ask what is your lb score without pseudo-labeled data for the same pipeline?",
      "votes": null
    },
    {
      "id": "3210307",
      "postDate": "05/27/2025 02:56:02",
      "content": "<p>Yes, it makes my model more stable and narraows the gap between different epochs.<br>\nThis is an example code:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .module = deepcopy(model)\n        .module.()\n        .decay = decay\n        .device = device\n         .device   :\n            .module.to(device=device)\n\n     ():\n         torch.no_grad():\n             ema_v, model_v  (.module.state_dict().values(), model.state_dict().values()):\n                 .device   :\n                    model_v = model_v.to(device=.device)\n                ema_v.copy_(update_fn(ema_v, model_v))\n\n     ():\n        ._update(model, update_fn= e, m: .decay * e + ( - .decay) * m)\n\n     ():\n        ._update(model, update_fn= e, m: m)\n\nmodel = BirdCLEFModel()\nema_model = ModelEMA(model, decay=)\n epoch  (n_epoch):\n     x, y  dataloader:\n        ......\n        ema_model.update(model)\n\ntorch.save(ema_model.module.state_dict(), )\n</code></pre>",
      "rawMarkdown": "Yes, it makes my model more stable and narraows the gap between different epochs.\nThis is an example code:\n```\nclass ModelEMA(nn.Module):\n    def __init__(self, model, decay=0.99, device=None):\n        super().__init__()\n        self.module = deepcopy(model)\n        self.module.eval()\n        self.decay = decay\n        self.device = device\n        if self.device is not None:\n            self.module.to(device=device)\n\n    def _update(self, model, update_fn):\n        with torch.no_grad():\n            for ema_v, model_v in zip(self.module.state_dict().values(), model.state_dict().values()):\n                if self.device is not None:\n                    model_v = model_v.to(device=self.device)\n                ema_v.copy_(update_fn(ema_v, model_v))\n\n    def update(self, model):\n        self._update(model, update_fn=lambda e, m: self.decay * e + (1. - self.decay) * m)\n\n    def set(self, model):\n        self._update(model, update_fn=lambda e, m: m)\n\nmodel = BirdCLEFModel()\nema_model = ModelEMA(model, decay=0.99)\nfor epoch in range(n_epoch):\n    for x, y in dataloader:\n        ......\n        ema_model.update(model)\n\ntorch.save(ema_model.module.state_dict(), 'model.pth')\n```",
      "votes": null
    },
    {
      "id": "3210311",
      "postDate": "05/27/2025 03:07:37",
      "content": "<p>My best model without unlabeled data is 0.871 with efv2b3 backbone. But there are some other tricks I have applied in my best score version (with unlabeled soundscapes), and I have not applied these tricks on the version without unlabeled soundscapes.</p>",
      "rawMarkdown": "My best model without unlabeled data is 0.871 with efv2b3 backbone. But there are some other tricks I have applied in my best score version (with unlabeled soundscapes), and I have not applied these tricks on the version without unlabeled soundscapes.",
      "votes": null
    },
    {
      "id": "3210312",
      "postDate": "05/27/2025 03:08:41",
      "content": "<p>It's very clear. Thx~</p>",
      "rawMarkdown": "It's very clear. Thx~",
      "votes": null
    },
    {
      "id": "3210319",
      "postDate": "05/27/2025 03:20:53",
      "content": "<p>Thank you for the clarification. My best-performing single model is based on the SED architecture with a tf_efficientnetv2_s.in21k encoder. It achieves a similar leaderboard score without using any unlabeled data. I'm currently exploring effective ways to incorporate the unlabeled data.</p>",
      "rawMarkdown": "Thank you for the clarification. My best-performing single model is based on the SED architecture with a tf_efficientnetv2_s.in21k encoder. It achieves a similar leaderboard score without using any unlabeled data. I'm currently exploring effective ways to incorporate the unlabeled data.",
      "votes": null
    },
    {
      "id": "3210323",
      "postDate": "05/27/2025 03:24:01",
      "content": "<p>By the way, do you use different segment lengths for training (e.g., 10-second segments instead of 5 seconds)? I previously experimented with 10-second segments using a CNN-based architecture, but it resulted in a lower leaderboard score.</p>",
      "rawMarkdown": "By the way, do you use different segment lengths for training (e.g., 10-second segments instead of 5 seconds)? I previously experimented with 10-second segments using a CNN-based architecture, but it resulted in a lower leaderboard score.",
      "votes": null
    },
    {
      "id": "3210324",
      "postDate": "05/27/2025 03:32:34",
      "content": "<p>I use random cropped 10-seconds segment for training and 5-second for inference.</p>",
      "rawMarkdown": "I use random cropped 10-seconds segment for training and 5-second for inference.",
      "votes": null
    },
    {
      "id": "3210502",
      "postDate": "05/27/2025 09:44:26",
      "content": "<p>Hi, Is this trained with CE or with FocalBCE.</p>",
      "rawMarkdown": "Hi, Is this trained with CE or with FocalBCE.",
      "votes": null
    },
    {
      "id": "3210523",
      "postDate": "05/27/2025 10:16:48",
      "content": "<p>I just used BCEloss.</p>",
      "rawMarkdown": "I just used BCEloss.",
      "votes": null
    },
    {
      "id": "3211445",
      "postDate": "05/28/2025 12:55:03",
      "content": "<p>curious if you use any type of early stopping, I use my val_map for early stopping (patience=3) and I dont seem to go over 12-15 epochs…</p>",
      "rawMarkdown": "curious if you use any type of early stopping, I use my val_map for early stopping (patience=3) and I dont seem to go over 12-15 epochs...",
      "votes": null
    },
    {
      "id": "3211563",
      "postDate": "05/28/2025 15:00:48",
      "content": "<p>I don't use early stopping because I don't trust my CV, you can try to detect a scope for your training epoch.</p>",
      "rawMarkdown": "I don't use early stopping because I don't trust my CV, you can try to detect a scope for your training epoch.",
      "votes": null
    },
    {
      "id": "3211645",
      "postDate": "05/28/2025 16:36:00",
      "content": "<p>Is the difference between backbones greater than the difference been separate runs of the same pipeline using the same backbone?</p>",
      "rawMarkdown": "Is the difference between backbones greater than the difference been separate runs of the same pipeline using the same backbone?",
      "votes": null
    },
    {
      "id": "3211653",
      "postDate": "05/28/2025 16:44:58",
      "content": "<p>I haven't test performance of different epochs using the same backbone. I just train them for fixed epochs (40)</p>",
      "rawMarkdown": "I haven't test performance of different epochs using the same backbone. I just train them for fixed epochs (40)",
      "votes": null
    },
    {
      "id": "3211807",
      "postDate": "05/28/2025 21:49:37",
      "content": "<p>I meant retraining a new model with the exactly same pipeline—same backbone, same number of epochs, new model. I’ve seen a large variation even with identical pipelines, so I suspect the issue doesn’t have to do with  backbones.</p>",
      "rawMarkdown": "I meant retraining a new model with the exactly same pipeline—same backbone, same number of epochs, new model. I’ve seen a large variation even with identical pipelines, so I suspect the issue doesn’t have to do with  backbones.",
      "votes": null
    },
    {
      "id": "3216627",
      "postDate": "06/03/2025 21:23:06",
      "content": "<p>[dumb question] When do you use EMA Model? Only during inference? <br>\nI have never came across this approach.</p>",
      "rawMarkdown": "[dumb question] When do you use EMA Model? Only during inference? \nI have never came across this approach.",
      "votes": null
    },
    {
      "id": "3216715",
      "postDate": "06/04/2025 02:56:07",
      "content": "<p>Yes. In the inference</p>",
      "rawMarkdown": "Yes. In the inference",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3209868,
      "author_name": "stefanoclss",
      "author_url": "",
      "post_date": "05/26/2025 12:34:01",
      "content": "<p>I mean training is validated on 10K+ samples, and your public LB score is based on 33% of 700 samples. I expect big fluctuations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3209874,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "05/26/2025 12:44:25",
          "content": "<p>In this competition，local cv Is meaningless. The score I mention above are LB score， just different feature extractors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3209883,
      "author_name": "pingfan",
      "author_url": "",
      "post_date": "05/26/2025 12:58:16",
      "content": "<p>I have the same problem. There are significant differences among different models</p>",
      "votes": null,
      "replies": [
        {
          "id": 3209964,
          "author_name": "leehann",
          "author_url": "",
          "post_date": "05/26/2025 14:57:59",
          "content": "<p>same issue</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3210050,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "05/26/2025 16:38:54",
          "content": "<p>More results: same pipeline and hyperparameters, efv2b3-0.902 resnext26t-0.861 efb2-0.858 rexnet150-0.857. It's very confusing.🥱</p>",
          "votes": null,
          "replies": [
            {
              "id": 3210217,
              "author_name": "willrice",
              "author_url": "",
              "post_date": "05/26/2025 22:18:44",
              "content": "<p>is this without pseudo-labeled data?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3210283,
                  "author_name": "shtljw",
                  "author_url": "",
                  "post_date": "05/27/2025 02:09:03",
                  "content": "<p>with pseudo-labeled data</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3210306,
                      "author_name": "sjtuwangshuo",
                      "author_url": "",
                      "post_date": "05/27/2025 02:53:34",
                      "content": "<p>Could I ask what is your lb score without pseudo-labeled data for the same pipeline?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3210311,
                          "author_name": "shtljw",
                          "author_url": "",
                          "post_date": "05/27/2025 03:07:37",
                          "content": "<p>My best model without unlabeled data is 0.871 with efv2b3 backbone. But there are some other tricks I have applied in my best score version (with unlabeled soundscapes), and I have not applied these tricks on the version without unlabeled soundscapes.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3210319,
                              "author_name": "sjtuwangshuo",
                              "author_url": "",
                              "post_date": "05/27/2025 03:20:53",
                              "content": "<p>Thank you for the clarification. My best-performing single model is based on the SED architecture with a tf_efficientnetv2_s.in21k encoder. It achieves a similar leaderboard score without using any unlabeled data. I'm currently exploring effective ways to incorporate the unlabeled data.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3210323,
                                  "author_name": "sjtuwangshuo",
                                  "author_url": "",
                                  "post_date": "05/27/2025 03:24:01",
                                  "content": "<p>By the way, do you use different segment lengths for training (e.g., 10-second segments instead of 5 seconds)? I previously experimented with 10-second segments using a CNN-based architecture, but it resulted in a lower leaderboard score.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3210324,
                                      "author_name": "shtljw",
                                      "author_url": "",
                                      "post_date": "05/27/2025 03:32:34",
                                      "content": "<p>I use random cropped 10-seconds segment for training and 5-second for inference.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 3210502,
                                          "author_name": "salmanahmedtamu",
                                          "author_url": "",
                                          "post_date": "05/27/2025 09:44:26",
                                          "content": "<p>Hi, Is this trained with CE or with FocalBCE.</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3210523,
                                              "author_name": "shtljw",
                                              "author_url": "",
                                              "post_date": "05/27/2025 10:16:48",
                                              "content": "<p>I just used BCEloss.</p>",
                                              "votes": null,
                                              "replies": []
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3209936,
      "author_name": "ziyi777",
      "author_url": "",
      "post_date": "05/26/2025 14:21:53",
      "content": "<p>Do you get these results at the same epoch?  For me it has fluctuated within 1.5%</p>",
      "votes": null,
      "replies": [
        {
          "id": 3210052,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "05/26/2025 16:39:56",
          "content": "<p>yep, all models are trained for 40 epochs</p>",
          "votes": null,
          "replies": [
            {
              "id": 3210152,
              "author_name": "ziyi777",
              "author_url": "",
              "post_date": "05/26/2025 19:55:36",
              "content": "<p>I think if you test different epoch for each model or use the model soup, the gap won't be too large like this. But that's just like lottery 😂</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3210286,
                  "author_name": "shtljw",
                  "author_url": "",
                  "post_date": "05/27/2025 02:10:47",
                  "content": "<p>Thank you for your advice, I would have a try.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3211445,
              "author_name": "lavanbth99",
              "author_url": "",
              "post_date": "05/28/2025 12:55:03",
              "content": "<p>curious if you use any type of early stopping, I use my val_map for early stopping (patience=3) and I dont seem to go over 12-15 epochs…</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3211563,
                  "author_name": "ziyi777",
                  "author_url": "",
                  "post_date": "05/28/2025 15:00:48",
                  "content": "<p>I don't use early stopping because I don't trust my CV, you can try to detect a scope for your training epoch.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3210294,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "05/27/2025 02:25:44",
      "content": "<p>Same here. Backbones like seresnext26d, v2b3, v2s, eca, show different lb and great gap. I apply ema in the training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3210303,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "05/27/2025 02:45:37",
          "content": "<p>May I ask how do ema perform? Dose it boost your LB score or make your results more stable?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3210307,
              "author_name": "i2nfinit3y",
              "author_url": "",
              "post_date": "05/27/2025 02:56:02",
              "content": "<p>Yes, it makes my model more stable and narraows the gap between different epochs.<br>\nThis is an example code:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .module = deepcopy(model)\n        .module.()\n        .decay = decay\n        .device = device\n         .device   :\n            .module.to(device=device)\n\n     ():\n         torch.no_grad():\n             ema_v, model_v  (.module.state_dict().values(), model.state_dict().values()):\n                 .device   :\n                    model_v = model_v.to(device=.device)\n                ema_v.copy_(update_fn(ema_v, model_v))\n\n     ():\n        ._update(model, update_fn= e, m: .decay * e + ( - .decay) * m)\n\n     ():\n        ._update(model, update_fn= e, m: m)\n\nmodel = BirdCLEFModel()\nema_model = ModelEMA(model, decay=)\n epoch  (n_epoch):\n     x, y  dataloader:\n        ......\n        ema_model.update(model)\n\ntorch.save(ema_model.module.state_dict(), )\n</code></pre>",
              "votes": null,
              "replies": [
                {
                  "id": 3210312,
                  "author_name": "shtljw",
                  "author_url": "",
                  "post_date": "05/27/2025 03:08:41",
                  "content": "<p>It's very clear. Thx~</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3216627,
                  "author_name": "aayush26",
                  "author_url": "",
                  "post_date": "06/03/2025 21:23:06",
                  "content": "<p>[dumb question] When do you use EMA Model? Only during inference? <br>\nI have never came across this approach.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3216715,
                      "author_name": "i2nfinit3y",
                      "author_url": "",
                      "post_date": "06/04/2025 02:56:07",
                      "content": "<p>Yes. In the inference</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3211645,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "05/28/2025 16:36:00",
      "content": "<p>Is the difference between backbones greater than the difference been separate runs of the same pipeline using the same backbone?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3211653,
          "author_name": "shtljw",
          "author_url": "",
          "post_date": "05/28/2025 16:44:58",
          "content": "<p>I haven't test performance of different epochs using the same backbone. I just train them for fixed epochs (40)</p>",
          "votes": null,
          "replies": [
            {
              "id": 3211807,
              "author_name": "robbynevels",
              "author_url": "",
              "post_date": "05/28/2025 21:49:37",
              "content": "<p>I meant retraining a new model with the exactly same pipeline—same backbone, same number of epochs, new model. I’ve seen a large variation even with identical pipelines, so I suspect the issue doesn’t have to do with  backbones.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3209818": "Same train and inference pipeline, efv2b3 scored 0.902, seresnext26t scored 0.861, seresnext26d scored ~0.87.  I expect there will be differences, but they shouldn't be that significant. Has anyone encountered similar problem?",
    "3209868": "I mean training is validated on 10K+ samples, and your public LB score is based on 33% of 700 samples. I expect big fluctuations.",
    "3209874": "In this competition，local cv Is meaningless. The score I mention above are LB score， just different feature extractors.",
    "3209883": "I have the same problem. There are significant differences among different models",
    "3209936": "Do you get these results at the same epoch?  For me it has fluctuated within 1.5%",
    "3209964": "same issue",
    "3210050": "More results: same pipeline and hyperparameters, efv2b3-0.902 resnext26t-0.861 efb2-0.858 rexnet150-0.857. It's very confusing.🥱",
    "3210052": "yep, all models are trained for 40 epochs",
    "3210152": "I think if you test different epoch for each model or use the model soup, the gap won't be too large like this. But that's just like lottery 😂",
    "3210217": "is this without pseudo-labeled data?",
    "3210283": "with pseudo-labeled data",
    "3210286": "Thank you for your advice, I would have a try.",
    "3210294": "Same here. Backbones like seresnext26d, v2b3, v2s, eca, show different lb and great gap. I apply ema in the training.",
    "3210303": "May I ask how do ema perform? Dose it boost your LB score or make your results more stable?",
    "3210306": "Could I ask what is your lb score without pseudo-labeled data for the same pipeline?",
    "3210307": "Yes, it makes my model more stable and narraows the gap between different epochs.\nThis is an example code:\n```\nclass ModelEMA(nn.Module):\n    def __init__(self, model, decay=0.99, device=None):\n        super().__init__()\n        self.module = deepcopy(model)\n        self.module.eval()\n        self.decay = decay\n        self.device = device\n        if self.device is not None:\n            self.module.to(device=device)\n\n    def _update(self, model, update_fn):\n        with torch.no_grad():\n            for ema_v, model_v in zip(self.module.state_dict().values(), model.state_dict().values()):\n                if self.device is not None:\n                    model_v = model_v.to(device=self.device)\n                ema_v.copy_(update_fn(ema_v, model_v))\n\n    def update(self, model):\n        self._update(model, update_fn=lambda e, m: self.decay * e + (1. - self.decay) * m)\n\n    def set(self, model):\n        self._update(model, update_fn=lambda e, m: m)\n\nmodel = BirdCLEFModel()\nema_model = ModelEMA(model, decay=0.99)\nfor epoch in range(n_epoch):\n    for x, y in dataloader:\n        ......\n        ema_model.update(model)\n\ntorch.save(ema_model.module.state_dict(), 'model.pth')\n```",
    "3210311": "My best model without unlabeled data is 0.871 with efv2b3 backbone. But there are some other tricks I have applied in my best score version (with unlabeled soundscapes), and I have not applied these tricks on the version without unlabeled soundscapes.",
    "3210312": "It's very clear. Thx~",
    "3210319": "Thank you for the clarification. My best-performing single model is based on the SED architecture with a tf_efficientnetv2_s.in21k encoder. It achieves a similar leaderboard score without using any unlabeled data. I'm currently exploring effective ways to incorporate the unlabeled data.",
    "3210323": "By the way, do you use different segment lengths for training (e.g., 10-second segments instead of 5 seconds)? I previously experimented with 10-second segments using a CNN-based architecture, but it resulted in a lower leaderboard score.",
    "3210324": "I use random cropped 10-seconds segment for training and 5-second for inference.",
    "3210502": "Hi, Is this trained with CE or with FocalBCE.",
    "3210523": "I just used BCEloss.",
    "3211445": "curious if you use any type of early stopping, I use my val_map for early stopping (patience=3) and I dont seem to go over 12-15 epochs...",
    "3211563": "I don't use early stopping because I don't trust my CV, you can try to detect a scope for your training epoch.",
    "3211645": "Is the difference between backbones greater than the difference been separate runs of the same pipeline using the same backbone?",
    "3211653": "I haven't test performance of different epochs using the same backbone. I just train them for fixed epochs (40)",
    "3211807": "I meant retraining a new model with the exactly same pipeline—same backbone, same number of epochs, new model. I’ve seen a large variation even with identical pipelines, so I suspect the issue doesn’t have to do with  backbones.",
    "3216627": "[dumb question] When do you use EMA Model? Only during inference? \nI have never came across this approach.",
    "3216715": "Yes. In the inference"
  },
  "source": "meta"
}