{
  "id": 583083,
  "title": "Ensemble Strategy for Architecture and Backbone",
  "url": "/competitions/birdclef-2025/discussion/583083",
  "author_name": "",
  "post_date": "2025-06-04T15:22:52.641955500Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The final submission deadline is approaching, and with the remaining attempts, I’m sure everyone is exploring various ensemble strategies.</p>\n<p>I’d love to hear your thoughts on architecture—are you planning to use a hybrid approach (CNN/SED/raw signal) or a single architecture?</p>\n<p>And for the backbone, are you leaning toward a single backbone, two backbones, or an ensemble of multiple backbones?</p>\n<p>From our experiments, SED has significantly outperformed CNN, while the different 5 backbones we tested showed slight variations but overall similar performance on the public leaderboard (same pipe, from 867-885).</p>\n<p>So for now, we’ve decided to go with a single architecture (SED) and ensemble as many different backbones as possible.</p>\n<p>How are you approaching this decision based on your own results?</p>",
  "messages": [
    {
      "id": "3217154",
      "postDate": "06/04/2025 15:22:52",
      "content": "<p>The final submission deadline is approaching, and with the remaining attempts, I’m sure everyone is exploring various ensemble strategies.</p>\n<p>I’d love to hear your thoughts on architecture—are you planning to use a hybrid approach (CNN/SED/raw signal) or a single architecture?</p>\n<p>And for the backbone, are you leaning toward a single backbone, two backbones, or an ensemble of multiple backbones?</p>\n<p>From our experiments, SED has significantly outperformed CNN, while the different 5 backbones we tested showed slight variations but overall similar performance on the public leaderboard (same pipe, from 867-885).</p>\n<p>So for now, we’ve decided to go with a single architecture (SED) and ensemble as many different backbones as possible.</p>\n<p>How are you approaching this decision based on your own results?</p>",
      "rawMarkdown": "The final submission deadline is approaching, and with the remaining attempts, I’m sure everyone is exploring various ensemble strategies.\n\nI’d love to hear your thoughts on architecture—are you planning to use a hybrid approach (CNN/SED/raw signal) or a single architecture?\n\nAnd for the backbone, are you leaning toward a single backbone, two backbones, or an ensemble of multiple backbones?\n\nFrom our experiments, SED has significantly outperformed CNN, while the different 5 backbones we tested showed slight variations but overall similar performance on the public leaderboard (same pipe, from 867-885).\n\nSo for now, we’ve decided to go with a single architecture (SED) and ensemble as many different backbones as possible.\n\nHow are you approaching this decision based on your own results?",
      "votes": null
    },
    {
      "id": "3217449",
      "postDate": "06/05/2025 03:06:01",
      "content": "<p>In my experiments,ensemble different train_data(select different 5s segments)improve a lot.Ensemble different architecture also help.<br>\nSome tests: <br>\n1.<strong>different train_data</strong>  3CNN(0.833/0.823/0.812)→0.862<br>\n2.<strong>different architecture</strong> 2SED(0.873/0.876) + 3CNN(0.862)→0.897 ,however,replace 0.873 with 0.881 single model will decrease public LB.<br>\nNow I have some doubts about whether I am overfitting public LB,because when i replace one of CNN model will decrease a lot.</p>",
      "rawMarkdown": "In my experiments,ensemble different train_data(select different 5s segments)improve a lot.Ensemble different architecture also help.\nSome tests: \n1.**different train_data**  3CNN(0.833/0.823/0.812)→0.862\n2.**different architecture** 2SED(0.873/0.876) + 3CNN(0.862)→0.897 ,however,replace 0.873 with 0.881 single model will decrease public LB.\nNow I have some doubts about whether I am overfitting public LB,because when i replace one of CNN model will decrease a lot.",
      "votes": null
    },
    {
      "id": "3217553",
      "postDate": "06/05/2025 06:09:18",
      "content": "<p>In my experiments, there's a big difference in prediction probability for the same audio using either CNN or SED architecture. Is it the same in your case?</p>",
      "rawMarkdown": "In my experiments, there's a big difference in prediction probability for the same audio using either CNN or SED architecture. Is it the same in your case?",
      "votes": null
    },
    {
      "id": "3217560",
      "postDate": "06/05/2025 06:17:12",
      "content": "<p>yep.But just simple weighted ensemble can improve public LB.<br>\nI try rank ensemble,not good as weighted ensemble.</p>",
      "rawMarkdown": "yep.But just simple weighted ensemble can improve public LB.\nI try rank ensemble,not good as weighted ensemble.",
      "votes": null
    },
    {
      "id": "3217627",
      "postDate": "06/05/2025 08:09:14",
      "content": "<p>Great diversity! Our results show that the ensemble score is heavily influenced by the highest-performing backbone. Currently, a seven-model ensemble has only marginally improved the score—by less than 0.001—compared to the best single model. It makes me question the generalization of the training pipeline😢</p>",
      "rawMarkdown": "Great diversity! Our results show that the ensemble score is heavily influenced by the highest-performing backbone. Currently, a seven-model ensemble has only marginally improved the score—by less than 0.001—compared to the best single model. It makes me question the generalization of the training pipeline😢",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217449,
      "author_name": "minatoyukinaxlisa",
      "author_url": "",
      "post_date": "06/05/2025 03:06:01",
      "content": "<p>In my experiments,ensemble different train_data(select different 5s segments)improve a lot.Ensemble different architecture also help.<br>\nSome tests: <br>\n1.<strong>different train_data</strong>  3CNN(0.833/0.823/0.812)→0.862<br>\n2.<strong>different architecture</strong> 2SED(0.873/0.876) + 3CNN(0.862)→0.897 ,however,replace 0.873 with 0.881 single model will decrease public LB.<br>\nNow I have some doubts about whether I am overfitting public LB,because when i replace one of CNN model will decrease a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217553,
          "author_name": "lhr124578",
          "author_url": "",
          "post_date": "06/05/2025 06:09:18",
          "content": "<p>In my experiments, there's a big difference in prediction probability for the same audio using either CNN or SED architecture. Is it the same in your case?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3217560,
              "author_name": "minatoyukinaxlisa",
              "author_url": "",
              "post_date": "06/05/2025 06:17:12",
              "content": "<p>yep.But just simple weighted ensemble can improve public LB.<br>\nI try rank ensemble,not good as weighted ensemble.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3217627,
          "author_name": "ryenhails",
          "author_url": "",
          "post_date": "06/05/2025 08:09:14",
          "content": "<p>Great diversity! Our results show that the ensemble score is heavily influenced by the highest-performing backbone. Currently, a seven-model ensemble has only marginally improved the score—by less than 0.001—compared to the best single model. It makes me question the generalization of the training pipeline😢</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3217154": "The final submission deadline is approaching, and with the remaining attempts, I’m sure everyone is exploring various ensemble strategies.\n\nI’d love to hear your thoughts on architecture—are you planning to use a hybrid approach (CNN/SED/raw signal) or a single architecture?\n\nAnd for the backbone, are you leaning toward a single backbone, two backbones, or an ensemble of multiple backbones?\n\nFrom our experiments, SED has significantly outperformed CNN, while the different 5 backbones we tested showed slight variations but overall similar performance on the public leaderboard (same pipe, from 867-885).\n\nSo for now, we’ve decided to go with a single architecture (SED) and ensemble as many different backbones as possible.\n\nHow are you approaching this decision based on your own results?",
    "3217449": "In my experiments,ensemble different train_data(select different 5s segments)improve a lot.Ensemble different architecture also help.\nSome tests: \n1.**different train_data**  3CNN(0.833/0.823/0.812)→0.862\n2.**different architecture** 2SED(0.873/0.876) + 3CNN(0.862)→0.897 ,however,replace 0.873 with 0.881 single model will decrease public LB.\nNow I have some doubts about whether I am overfitting public LB,because when i replace one of CNN model will decrease a lot.",
    "3217553": "In my experiments, there's a big difference in prediction probability for the same audio using either CNN or SED architecture. Is it the same in your case?",
    "3217560": "yep.But just simple weighted ensemble can improve public LB.\nI try rank ensemble,not good as weighted ensemble.",
    "3217627": "Great diversity! Our results show that the ensemble score is heavily influenced by the highest-performing backbone. Currently, a seven-model ensemble has only marginally improved the score—by less than 0.001—compared to the best single model. It makes me question the generalization of the training pipeline😢"
  },
  "source": "meta"
}