{
  "id": 583381,
  "title": "6th place solution",
  "url": "/competitions/birdclef-2025/writeups/takoi-6th-place-solution",
  "author_name": "",
  "post_date": "2025-06-06T12:34:03.005284700Z",
  "votes": 23,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>6th place solution</h1>\n<p>Thank you to the organizers for hosting this excellent competition, and also to all the participants who shared valuable insights throughout the competition period.</p>\n<p>I was greatly inspired by the top solutions from past competitions. Thank you to everyone who shared their approaches. <br>\nIn this post, I will mainly highlight the differences in my approach.</p>\n<h2>Model / Loss</h2>\n<p>I used a SED-style model like the one below:</p>\n<pre><code> (nn.Module):\n     ():\n        ...\n        .activation = activation\n        ...\n\n     ():\n        norm_att = torch.softmax(torch.tanh(.att(x)), dim=-)\n        cla = .nonlinear_transform(.cla(x))\n        x = (norm_att * cla).()\n         x, norm_att, cla\n\n    ():\n         .activation == :\n             x\n         .activation == :\n             torch.sigmoid(x)\n\n (nn.Module):\n     ():\n        ...\n        .encoder = timm.create_model(\n            cfg.backbone,\n            pretrained=cfg.pretrained,\n            num_classes=,\n            global_pool=,\n            in_chans=cfg.in_chans,\n            drop_path_rate=,\n            drop_rate=,\n        )\n        ...\n        .att_block = AttBlockV2(in_features, .num_classes, activation=)\n        ...\n\n     ():\n        ...\n        clipwise_output, norm_att, segmentwise_output = .att_block(x)\n        segmentwise_logit = .att_block.cla(x).transpose(, )\n         .training:\n             clipwise_output, segmentwise_logit.()[], y\n        :\n             clipwise_output, segmentwise_logit.()[]\n</code></pre>\n<p>During training, I applied nn.BCEWithLogitsLoss to both clipwise_output and segmentwise_logit.max(1)[0].<br>\nAlthough clipwise_output passes through a sigmoid before loss computation (making BCEWithLogitsLoss technically inappropriate), this setup significantly improved my public score.</p>\n<p>When using timm/tf_efficientnet_b3.ns_jft_in1k as the backbone and submitting with clipwise_output  (without using train_soundscape), I achieved a public score of 0.900 and a private score of 0.908. At this stage, clipwise_output gave better results than segmentwise_logit.</p>\n<h2>Pseudo Labeling</h2>\n<ul>\n<li><p>I added pseudo labels to train_soundscapes and ran several training cycles with them included as training data.</p></li>\n<li><p>In the first round, I used only the clipwise_output from a single model (timm/tf_efficientnet_b3.ns_jft_in1k) to generate pseudo labels.</p></li>\n<li><p>From the second round onwards, I used an ensemble of multiple models' segmentwise_logit.max(1)[0] outputs for pseudo labeling.</p>\n<ul>\n<li><p>Since clipwise_output was trained somewhat unnaturally with BCEWithLogitsLoss, its values were too small and didn't help improve the score when reused in pseudo labeling.</p></li>\n<li><p>Using segmentwise_logit.max(1)[0] for pseudo labeling led to higher public scores.</p></li></ul></li>\n<li><p>Models used for generating pseudo labels included:</p>\n<ul>\n<li><p>timm/tf_efficientnet_b3.ns_jft_in1k</p></li>\n<li><p>timm/tf_efficientnet_b5.ns_jft_in1k</p></li>\n<li><p>timm/tf_efficientnetv2_b3.in21k</p></li></ul></li>\n<li><p>Finally, I trained the following two models on the combined training data and pseudo labels, and used them for the final submission:</p>\n<ul>\n<li><p>timm/tf_efficientnet_b3.ns_jft_in1k</p></li>\n<li><p>timm/tf_efficientnetv2_b3.in21k</p></li></ul></li>\n</ul>\n<h2>Score</h2>\n<ul>\n<li><p>Public Score: 0.928</p></li>\n<li><p>Private Score: 0.923</p></li>\n</ul>",
  "messages": [
    {
      "id": "3218614",
      "postDate": "06/06/2025 12:34:03",
      "content": "<h1>6th place solution</h1>\n<p>Thank you to the organizers for hosting this excellent competition, and also to all the participants who shared valuable insights throughout the competition period.</p>\n<p>I was greatly inspired by the top solutions from past competitions. Thank you to everyone who shared their approaches. <br>\nIn this post, I will mainly highlight the differences in my approach.</p>\n<h2>Model / Loss</h2>\n<p>I used a SED-style model like the one below:</p>\n<pre><code> (nn.Module):\n     ():\n        ...\n        .activation = activation\n        ...\n\n     ():\n        norm_att = torch.softmax(torch.tanh(.att(x)), dim=-)\n        cla = .nonlinear_transform(.cla(x))\n        x = (norm_att * cla).()\n         x, norm_att, cla\n\n    ():\n         .activation == :\n             x\n         .activation == :\n             torch.sigmoid(x)\n\n (nn.Module):\n     ():\n        ...\n        .encoder = timm.create_model(\n            cfg.backbone,\n            pretrained=cfg.pretrained,\n            num_classes=,\n            global_pool=,\n            in_chans=cfg.in_chans,\n            drop_path_rate=,\n            drop_rate=,\n        )\n        ...\n        .att_block = AttBlockV2(in_features, .num_classes, activation=)\n        ...\n\n     ():\n        ...\n        clipwise_output, norm_att, segmentwise_output = .att_block(x)\n        segmentwise_logit = .att_block.cla(x).transpose(, )\n         .training:\n             clipwise_output, segmentwise_logit.()[], y\n        :\n             clipwise_output, segmentwise_logit.()[]\n</code></pre>\n<p>During training, I applied nn.BCEWithLogitsLoss to both clipwise_output and segmentwise_logit.max(1)[0].<br>\nAlthough clipwise_output passes through a sigmoid before loss computation (making BCEWithLogitsLoss technically inappropriate), this setup significantly improved my public score.</p>\n<p>When using timm/tf_efficientnet_b3.ns_jft_in1k as the backbone and submitting with clipwise_output  (without using train_soundscape), I achieved a public score of 0.900 and a private score of 0.908. At this stage, clipwise_output gave better results than segmentwise_logit.</p>\n<h2>Pseudo Labeling</h2>\n<ul>\n<li><p>I added pseudo labels to train_soundscapes and ran several training cycles with them included as training data.</p></li>\n<li><p>In the first round, I used only the clipwise_output from a single model (timm/tf_efficientnet_b3.ns_jft_in1k) to generate pseudo labels.</p></li>\n<li><p>From the second round onwards, I used an ensemble of multiple models' segmentwise_logit.max(1)[0] outputs for pseudo labeling.</p>\n<ul>\n<li><p>Since clipwise_output was trained somewhat unnaturally with BCEWithLogitsLoss, its values were too small and didn't help improve the score when reused in pseudo labeling.</p></li>\n<li><p>Using segmentwise_logit.max(1)[0] for pseudo labeling led to higher public scores.</p></li></ul></li>\n<li><p>Models used for generating pseudo labels included:</p>\n<ul>\n<li><p>timm/tf_efficientnet_b3.ns_jft_in1k</p></li>\n<li><p>timm/tf_efficientnet_b5.ns_jft_in1k</p></li>\n<li><p>timm/tf_efficientnetv2_b3.in21k</p></li></ul></li>\n<li><p>Finally, I trained the following two models on the combined training data and pseudo labels, and used them for the final submission:</p>\n<ul>\n<li><p>timm/tf_efficientnet_b3.ns_jft_in1k</p></li>\n<li><p>timm/tf_efficientnetv2_b3.in21k</p></li></ul></li>\n</ul>\n<h2>Score</h2>\n<ul>\n<li><p>Public Score: 0.928</p></li>\n<li><p>Private Score: 0.923</p></li>\n</ul>",
      "rawMarkdown": "# 6th place solution\nThank you to the organizers for hosting this excellent competition, and also to all the participants who shared valuable insights throughout the competition period.\n\nI was greatly inspired by the top solutions from past competitions. Thank you to everyone who shared their approaches. \nIn this post, I will mainly highlight the differences in my approach.\n\n## Model / Loss\nI used a SED-style model like the one below:\n```\nclass AttBlockV2(nn.Module):\n    def __init__(self, in_features: int, out_features: int, activation=\"sigmoid\"):\n        ...\n        self.activation = activation\n        ...\n\n    def forward(self, x):\n        norm_att = torch.softmax(torch.tanh(self.att(x)), dim=-1)\n        cla = self.nonlinear_transform(self.cla(x))\n        x = (norm_att * cla).sum(2)\n        return x, norm_att, cla\n\n   def nonlinear_transform(self, x):\n        if self.activation == \"linear\":\n            return x\n        elif self.activation == \"sigmoid\":\n            return torch.sigmoid(x)\n            \nclass BirdModel(nn.Module):\n    def __init__(self, cfg, pretrained: bool = True):\n        ...\n        self.encoder = timm.create_model(\n            cfg.backbone,\n            pretrained=cfg.pretrained,\n            num_classes=0,\n            global_pool=\"\",\n            in_chans=cfg.in_chans,\n            drop_path_rate=0.2,\n            drop_rate=0.5,\n        )\n        ...\n        self.att_block = AttBlockV2(in_features, self.num_classes, activation=\"sigmoid\")\n        ...\n       \n    def forward(self, x, y=None):\n        ...\n        clipwise_output, norm_att, segmentwise_output = self.att_block(x)\n        segmentwise_logit = self.att_block.cla(x).transpose(1, 2)\n        if self.training:\n            return clipwise_output, segmentwise_logit.max(1)[0], y\n        else:\n            return clipwise_output, segmentwise_logit.max(1)[0]\n```\nDuring training, I applied nn.BCEWithLogitsLoss to both clipwise_output and segmentwise_logit.max(1)[0].\nAlthough clipwise_output passes through a sigmoid before loss computation (making BCEWithLogitsLoss technically inappropriate), this setup significantly improved my public score.\n\nWhen using timm/tf_efficientnet_b3.ns_jft_in1k as the backbone and submitting with clipwise_output  (without using train_soundscape), I achieved a public score of 0.900 and a private score of 0.908. At this stage, clipwise_output gave better results than segmentwise_logit.\n\n## Pseudo Labeling\n- I added pseudo labels to train_soundscapes and ran several training cycles with them included as training data.\n\n- In the first round, I used only the clipwise_output from a single model (timm/tf_efficientnet_b3.ns_jft_in1k) to generate pseudo labels.\n\n- From the second round onwards, I used an ensemble of multiple models' segmentwise_logit.max(1)[0] outputs for pseudo labeling.\n\n    - Since clipwise_output was trained somewhat unnaturally with BCEWithLogitsLoss, its values were too small and didn't help improve the score when reused in pseudo labeling.\n\n    - Using segmentwise_logit.max(1)[0] for pseudo labeling led to higher public scores.\n\n- Models used for generating pseudo labels included:\n\n    - timm/tf_efficientnet_b3.ns_jft_in1k\n\n    - timm/tf_efficientnet_b5.ns_jft_in1k\n\n    - timm/tf_efficientnetv2_b3.in21k\n\n- Finally, I trained the following two models on the combined training data and pseudo labels, and used them for the final submission:\n\n    - timm/tf_efficientnet_b3.ns_jft_in1k\n\n    - timm/tf_efficientnetv2_b3.in21k\n\n## Score\n- Public Score: 0.928\n\n- Private Score: 0.923",
      "votes": null
    },
    {
      "id": "3218913",
      "postDate": "06/06/2025 23:16:43",
      "content": "<p>Thank you for sharing your 6th place solution! It's very insightful to see your approach, especially the details on the SED-style model and the pseudo-labeling strategy.</p>\n<p>I have a question regarding your loss function: You mentioned applying <code>nn.BCEWithLogitsLoss</code> to <code>clipwise_output</code> even though it passes through a sigmoid before loss computation. Could you elaborate on why this particular setup significantly improved your public score, despite being \"technically inappropriate\"? Was there any specific observation or intuition that led you to this choice?</p>",
      "rawMarkdown": "Thank you for sharing your 6th place solution! It's very insightful to see your approach, especially the details on the SED-style model and the pseudo-labeling strategy.\n\nI have a question regarding your loss function: You mentioned applying `nn.BCEWithLogitsLoss` to `clipwise_output` even though it passes through a sigmoid before loss computation. Could you elaborate on why this particular setup significantly improved your public score, despite being \"technically inappropriate\"? Was there any specific observation or intuition that led you to this choice?",
      "votes": null
    },
    {
      "id": "3219634",
      "postDate": "06/08/2025 05:26:23",
      "content": "<p>Thank you for your question.</p>\n<p>I believe this setup might have helped prevent overfitting to the training data. Normally, when the model's output is low, it would take negative values before applying the sigmoid. However, since the output is passed through a sigmoid beforehand, it can no longer take negative values. As a result, when the label is 0, the model is forced to produce very low output values; otherwise, the loss would become large. This tendency to keep the outputs low might have contributed to reducing overfitting.</p>\n<p>As for why I used this setup — to be honest, there wasn’t any particular reason. I just tried it out, and it happened to improve my public score, so I continued using it until the end.</p>",
      "rawMarkdown": "Thank you for your question.\n\nI believe this setup might have helped prevent overfitting to the training data. Normally, when the model's output is low, it would take negative values before applying the sigmoid. However, since the output is passed through a sigmoid beforehand, it can no longer take negative values. As a result, when the label is 0, the model is forced to produce very low output values; otherwise, the loss would become large. This tendency to keep the outputs low might have contributed to reducing overfitting.\n\nAs for why I used this setup — to be honest, there wasn’t any particular reason. I just tried it out, and it happened to improve my public score, so I continued using it until the end.",
      "votes": null
    },
    {
      "id": "3219752",
      "postDate": "06/08/2025 09:19:37",
      "content": "<p>SED-style model, its a great decision</p>",
      "rawMarkdown": "SED-style model, its a great decision",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218913,
      "author_name": "tyyuki",
      "author_url": "",
      "post_date": "06/06/2025 23:16:43",
      "content": "<p>Thank you for sharing your 6th place solution! It's very insightful to see your approach, especially the details on the SED-style model and the pseudo-labeling strategy.</p>\n<p>I have a question regarding your loss function: You mentioned applying <code>nn.BCEWithLogitsLoss</code> to <code>clipwise_output</code> even though it passes through a sigmoid before loss computation. Could you elaborate on why this particular setup significantly improved your public score, despite being \"technically inappropriate\"? Was there any specific observation or intuition that led you to this choice?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219634,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/08/2025 05:26:23",
          "content": "<p>Thank you for your question.</p>\n<p>I believe this setup might have helped prevent overfitting to the training data. Normally, when the model's output is low, it would take negative values before applying the sigmoid. However, since the output is passed through a sigmoid beforehand, it can no longer take negative values. As a result, when the label is 0, the model is forced to produce very low output values; otherwise, the loss would become large. This tendency to keep the outputs low might have contributed to reducing overfitting.</p>\n<p>As for why I used this setup — to be honest, there wasn’t any particular reason. I just tried it out, and it happened to improve my public score, so I continued using it until the end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219752,
      "author_name": "andrewh33",
      "author_url": "",
      "post_date": "06/08/2025 09:19:37",
      "content": "<p>SED-style model, its a great decision</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218614": "# 6th place solution\nThank you to the organizers for hosting this excellent competition, and also to all the participants who shared valuable insights throughout the competition period.\n\nI was greatly inspired by the top solutions from past competitions. Thank you to everyone who shared their approaches. \nIn this post, I will mainly highlight the differences in my approach.\n\n## Model / Loss\nI used a SED-style model like the one below:\n```\nclass AttBlockV2(nn.Module):\n    def __init__(self, in_features: int, out_features: int, activation=\"sigmoid\"):\n        ...\n        self.activation = activation\n        ...\n\n    def forward(self, x):\n        norm_att = torch.softmax(torch.tanh(self.att(x)), dim=-1)\n        cla = self.nonlinear_transform(self.cla(x))\n        x = (norm_att * cla).sum(2)\n        return x, norm_att, cla\n\n   def nonlinear_transform(self, x):\n        if self.activation == \"linear\":\n            return x\n        elif self.activation == \"sigmoid\":\n            return torch.sigmoid(x)\n            \nclass BirdModel(nn.Module):\n    def __init__(self, cfg, pretrained: bool = True):\n        ...\n        self.encoder = timm.create_model(\n            cfg.backbone,\n            pretrained=cfg.pretrained,\n            num_classes=0,\n            global_pool=\"\",\n            in_chans=cfg.in_chans,\n            drop_path_rate=0.2,\n            drop_rate=0.5,\n        )\n        ...\n        self.att_block = AttBlockV2(in_features, self.num_classes, activation=\"sigmoid\")\n        ...\n       \n    def forward(self, x, y=None):\n        ...\n        clipwise_output, norm_att, segmentwise_output = self.att_block(x)\n        segmentwise_logit = self.att_block.cla(x).transpose(1, 2)\n        if self.training:\n            return clipwise_output, segmentwise_logit.max(1)[0], y\n        else:\n            return clipwise_output, segmentwise_logit.max(1)[0]\n```\nDuring training, I applied nn.BCEWithLogitsLoss to both clipwise_output and segmentwise_logit.max(1)[0].\nAlthough clipwise_output passes through a sigmoid before loss computation (making BCEWithLogitsLoss technically inappropriate), this setup significantly improved my public score.\n\nWhen using timm/tf_efficientnet_b3.ns_jft_in1k as the backbone and submitting with clipwise_output  (without using train_soundscape), I achieved a public score of 0.900 and a private score of 0.908. At this stage, clipwise_output gave better results than segmentwise_logit.\n\n## Pseudo Labeling\n- I added pseudo labels to train_soundscapes and ran several training cycles with them included as training data.\n\n- In the first round, I used only the clipwise_output from a single model (timm/tf_efficientnet_b3.ns_jft_in1k) to generate pseudo labels.\n\n- From the second round onwards, I used an ensemble of multiple models' segmentwise_logit.max(1)[0] outputs for pseudo labeling.\n\n    - Since clipwise_output was trained somewhat unnaturally with BCEWithLogitsLoss, its values were too small and didn't help improve the score when reused in pseudo labeling.\n\n    - Using segmentwise_logit.max(1)[0] for pseudo labeling led to higher public scores.\n\n- Models used for generating pseudo labels included:\n\n    - timm/tf_efficientnet_b3.ns_jft_in1k\n\n    - timm/tf_efficientnet_b5.ns_jft_in1k\n\n    - timm/tf_efficientnetv2_b3.in21k\n\n- Finally, I trained the following two models on the combined training data and pseudo labels, and used them for the final submission:\n\n    - timm/tf_efficientnet_b3.ns_jft_in1k\n\n    - timm/tf_efficientnetv2_b3.in21k\n\n## Score\n- Public Score: 0.928\n\n- Private Score: 0.923",
    "3218913": "Thank you for sharing your 6th place solution! It's very insightful to see your approach, especially the details on the SED-style model and the pseudo-labeling strategy.\n\nI have a question regarding your loss function: You mentioned applying `nn.BCEWithLogitsLoss` to `clipwise_output` even though it passes through a sigmoid before loss computation. Could you elaborate on why this particular setup significantly improved your public score, despite being \"technically inappropriate\"? Was there any specific observation or intuition that led you to this choice?",
    "3219634": "Thank you for your question.\n\nI believe this setup might have helped prevent overfitting to the training data. Normally, when the model's output is low, it would take negative values before applying the sigmoid. However, since the output is passed through a sigmoid beforehand, it can no longer take negative values. As a result, when the label is 0, the model is forced to produce very low output values; otherwise, the loss would become large. This tendency to keep the outputs low might have contributed to reducing overfitting.\n\nAs for why I used this setup — to be honest, there wasn’t any particular reason. I just tried it out, and it happened to improve my public score, so I continued using it until the end.",
    "3219752": "SED-style model, its a great decision"
  },
  "source": "meta"
}