{
  "id": 579117,
  "title": "Can someone explain what I am doing wrong? Can't go beyond 0.751 with this approach!",
  "url": "/competitions/birdclef-2025/discussion/579117",
  "author_name": "",
  "post_date": "2025-05-15T10:12:13.275754Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Here is the logic that I am currently using: </p>\n<ol>\n<li>Human Voice Detection &amp; Removal</li>\n</ol>\n<ul>\n<li>use Silero VAD to detect and remove human speech segments from each training audio file.<br>\n-Detected timestamps are completely removed from the waveform, not masked.</li>\n</ul>\n<ol>\n<li><p>Dataset Preparation<br>\nFor each training epoch:<br>\n-A random 5-second chunk is selected from the cleaned waveform.<br>\n-If the waveform is shorter than 5s post-cleaning, it's zero-padded.</p></li>\n<li><p>Input Representation<br>\n-generate 3-channel mel spectrograms per input:Original audio, Noise-augmented version ,Time-stretched version<br>\n-These are stacked to form an input with shape [3, 128, T] (T = time frames).</p></li>\n<li><p>Model<br>\nBackbone: tf_efficientnet_b0 from timm, with in_chans=3.<br>\nFinal layer: Linear(1280 → num_classes)<br>\nDropout: p=0.3 before the classifier head.</p></li>\n<li><p>Augmentation<br>\nSpecAugment is applied to each spectrogram (time &amp; frequency masking).<br>\nMixup is optionally applied between random training pairs with α=0.5.</p></li>\n<li><p>Loss &amp; Optimization<br>\nLoss: BCEWithLogitsLoss for multi-label classification.<br>\nOptimizer: Adam with weight decay and learning rate scheduling.<br>\nMetric: Validation AUC is monitored for early stopping and model checkpointing.</p></li>\n</ol>",
  "messages": [
    {
      "id": "3202399",
      "postDate": "05/15/2025 10:12:13",
      "content": "<p>Here is the logic that I am currently using: </p>\n<ol>\n<li>Human Voice Detection &amp; Removal</li>\n</ol>\n<ul>\n<li>use Silero VAD to detect and remove human speech segments from each training audio file.<br>\n-Detected timestamps are completely removed from the waveform, not masked.</li>\n</ul>\n<ol>\n<li><p>Dataset Preparation<br>\nFor each training epoch:<br>\n-A random 5-second chunk is selected from the cleaned waveform.<br>\n-If the waveform is shorter than 5s post-cleaning, it's zero-padded.</p></li>\n<li><p>Input Representation<br>\n-generate 3-channel mel spectrograms per input:Original audio, Noise-augmented version ,Time-stretched version<br>\n-These are stacked to form an input with shape [3, 128, T] (T = time frames).</p></li>\n<li><p>Model<br>\nBackbone: tf_efficientnet_b0 from timm, with in_chans=3.<br>\nFinal layer: Linear(1280 → num_classes)<br>\nDropout: p=0.3 before the classifier head.</p></li>\n<li><p>Augmentation<br>\nSpecAugment is applied to each spectrogram (time &amp; frequency masking).<br>\nMixup is optionally applied between random training pairs with α=0.5.</p></li>\n<li><p>Loss &amp; Optimization<br>\nLoss: BCEWithLogitsLoss for multi-label classification.<br>\nOptimizer: Adam with weight decay and learning rate scheduling.<br>\nMetric: Validation AUC is monitored for early stopping and model checkpointing.</p></li>\n</ol>",
      "rawMarkdown": "Here is the logic that I am currently using: \n\n\n1. Human Voice Detection & Removal\n- use Silero VAD to detect and remove human speech segments from each training audio file.\n-Detected timestamps are completely removed from the waveform, not masked.\n\n2. Dataset Preparation\nFor each training epoch:\n-A random 5-second chunk is selected from the cleaned waveform.\n-If the waveform is shorter than 5s post-cleaning, it's zero-padded.\n\n3. Input Representation\n-generate 3-channel mel spectrograms per input:Original audio, Noise-augmented version ,Time-stretched version\n-These are stacked to form an input with shape [3, 128, T] (T = time frames).\n\n4. Model\nBackbone: tf_efficientnet_b0 from timm, with in_chans=3.\nFinal layer: Linear(1280 → num_classes)\nDropout: p=0.3 before the classifier head.\n\n5. Augmentation\nSpecAugment is applied to each spectrogram (time & frequency masking).\nMixup is optionally applied between random training pairs with α=0.5.\n\n6. Loss & Optimization\nLoss: BCEWithLogitsLoss for multi-label classification.\nOptimizer: Adam with weight decay and learning rate scheduling.\nMetric: Validation AUC is monitored for early stopping and model checkpointing.",
      "votes": null
    },
    {
      "id": "3202562",
      "postDate": "05/15/2025 14:57:45",
      "content": "<p>The simplest option is to change tf_efficientnet_b0 to tf_efficientnet_b2 or tf_efficientnetv2_b0.</p>",
      "rawMarkdown": "The simplest option is to change tf_efficientnet_b0 to tf_efficientnet_b2 or tf_efficientnetv2_b0.",
      "votes": null
    },
    {
      "id": "3202810",
      "postDate": "05/16/2025 00:20:11",
      "content": "<p>Maybe try not deleting human voice, in recent discussion people say removing humvoice somehow worse their LB. <br>\nI didn’t try humvoice yet, working with pseudo)</p>",
      "rawMarkdown": "Maybe try not deleting human voice, in recent discussion people say removing humvoice somehow worse their LB. \nI didn’t try humvoice yet, working with pseudo)",
      "votes": null
    },
    {
      "id": "3203296",
      "postDate": "05/16/2025 14:40:09",
      "content": "<p>There is a discussion with the ideas from previous birdclef and it is stated in there that birdclef 2024 1st place used loop padding. Maybe trying that might help.</p>\n<p>Or combining the approaches, at least it might make the train data more diverse.</p>\n<p>Link: <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/572928\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/572928</a></p>",
      "rawMarkdown": "There is a discussion with the ideas from previous birdclef and it is stated in there that birdclef 2024 1st place used loop padding. Maybe trying that might help.\n\nOr combining the approaches, at least it might make the train data more diverse.\n\nLink: https://www.kaggle.com/competitions/birdclef-2025/discussion/572928",
      "votes": null
    },
    {
      "id": "3203303",
      "postDate": "05/16/2025 14:46:28",
      "content": "<p>Based on my observations, ensembling cross-validated models might be a promising direction. While a single b0 model provides a strong baseline, my experiments showed that mean-ensembling multiple instances tended to yield a modest improvement in validation consistency. Would appreciate hearing others' experiences with model stacking here!</p>",
      "rawMarkdown": "Based on my observations, ensembling cross-validated models might be a promising direction. While a single b0 model provides a strong baseline, my experiments showed that mean-ensembling multiple instances tended to yield a modest improvement in validation consistency. Would appreciate hearing others' experiences with model stacking here!",
      "votes": null
    },
    {
      "id": "3203346",
      "postDate": "05/16/2025 15:39:59",
      "content": "<p>I'm doing similar things. But for me, different versions of three channel inputs did not improve the results. Therefore, I went back to a single spectrogram as an input. Also, I'm using a different classification head, which consists of two layers and has a residual connection. With that I was able to reach 0.819 (IIRC) for the efficient net b0</p>",
      "rawMarkdown": "I'm doing similar things. But for me, different versions of three channel inputs did not improve the results. Therefore, I went back to a single spectrogram as an input. Also, I'm using a different classification head, which consists of two layers and has a residual connection. With that I was able to reach 0.819 (IIRC) for the efficient net b0",
      "votes": null
    },
    {
      "id": "3203351",
      "postDate": "05/16/2025 15:45:40",
      "content": "<p>And I used soft labels in conjunction with a symmetric BCE loss.</p>",
      "rawMarkdown": "And I used soft labels in conjunction with a symmetric BCE loss.",
      "votes": null
    },
    {
      "id": "3203537",
      "postDate": "05/16/2025 22:20:51",
      "content": "<p>Do you try thome other think?</p>",
      "rawMarkdown": "Do you try thome other think?",
      "votes": null
    },
    {
      "id": "3210534",
      "postDate": "05/27/2025 10:45:51",
      "content": "<p>halo,for your work,the other two models improve how much LB?</p>",
      "rawMarkdown": "halo,for your work,the other two models improve how much LB?",
      "votes": null
    },
    {
      "id": "3210551",
      "postDate": "05/27/2025 11:04:00",
      "content": "<p>0.8 to 0.818</p>",
      "rawMarkdown": "0.8 to 0.818",
      "votes": null
    },
    {
      "id": "3231086",
      "postDate": "06/23/2025 21:51:42",
      "content": "<p>Loss &amp; Optimization what mean ?  </p>",
      "rawMarkdown": "Loss & Optimization what mean ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3202562,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "05/15/2025 14:57:45",
      "content": "<p>The simplest option is to change tf_efficientnet_b0 to tf_efficientnet_b2 or tf_efficientnetv2_b0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3210534,
          "author_name": "xiayuxuan",
          "author_url": "",
          "post_date": "05/27/2025 10:45:51",
          "content": "<p>halo,for your work,the other two models improve how much LB?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3210551,
              "author_name": "junhanzangai",
              "author_url": "",
              "post_date": "05/27/2025 11:04:00",
              "content": "<p>0.8 to 0.818</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3202810,
      "author_name": "player77",
      "author_url": "",
      "post_date": "05/16/2025 00:20:11",
      "content": "<p>Maybe try not deleting human voice, in recent discussion people say removing humvoice somehow worse their LB. <br>\nI didn’t try humvoice yet, working with pseudo)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3203296,
      "author_name": "alexandergremyakov",
      "author_url": "",
      "post_date": "05/16/2025 14:40:09",
      "content": "<p>There is a discussion with the ideas from previous birdclef and it is stated in there that birdclef 2024 1st place used loop padding. Maybe trying that might help.</p>\n<p>Or combining the approaches, at least it might make the train data more diverse.</p>\n<p>Link: <a href=\"https://www.kaggle.com/competitions/birdclef-2025/discussion/572928\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2025/discussion/572928</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3203303,
      "author_name": "alexandergremyakov",
      "author_url": "",
      "post_date": "05/16/2025 14:46:28",
      "content": "<p>Based on my observations, ensembling cross-validated models might be a promising direction. While a single b0 model provides a strong baseline, my experiments showed that mean-ensembling multiple instances tended to yield a modest improvement in validation consistency. Would appreciate hearing others' experiences with model stacking here!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3203346,
      "author_name": "tim6502",
      "author_url": "",
      "post_date": "05/16/2025 15:39:59",
      "content": "<p>I'm doing similar things. But for me, different versions of three channel inputs did not improve the results. Therefore, I went back to a single spectrogram as an input. Also, I'm using a different classification head, which consists of two layers and has a residual connection. With that I was able to reach 0.819 (IIRC) for the efficient net b0</p>",
      "votes": null,
      "replies": [
        {
          "id": 3203351,
          "author_name": "tim6502",
          "author_url": "",
          "post_date": "05/16/2025 15:45:40",
          "content": "<p>And I used soft labels in conjunction with a symmetric BCE loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3203537,
      "author_name": "matishb",
      "author_url": "",
      "post_date": "05/16/2025 22:20:51",
      "content": "<p>Do you try thome other think?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3231086,
          "author_name": "ernestgarciaphd",
          "author_url": "",
          "post_date": "06/23/2025 21:51:42",
          "content": "<p>Loss &amp; Optimization what mean ?  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3202399": "Here is the logic that I am currently using: \n\n\n1. Human Voice Detection & Removal\n- use Silero VAD to detect and remove human speech segments from each training audio file.\n-Detected timestamps are completely removed from the waveform, not masked.\n\n2. Dataset Preparation\nFor each training epoch:\n-A random 5-second chunk is selected from the cleaned waveform.\n-If the waveform is shorter than 5s post-cleaning, it's zero-padded.\n\n3. Input Representation\n-generate 3-channel mel spectrograms per input:Original audio, Noise-augmented version ,Time-stretched version\n-These are stacked to form an input with shape [3, 128, T] (T = time frames).\n\n4. Model\nBackbone: tf_efficientnet_b0 from timm, with in_chans=3.\nFinal layer: Linear(1280 → num_classes)\nDropout: p=0.3 before the classifier head.\n\n5. Augmentation\nSpecAugment is applied to each spectrogram (time & frequency masking).\nMixup is optionally applied between random training pairs with α=0.5.\n\n6. Loss & Optimization\nLoss: BCEWithLogitsLoss for multi-label classification.\nOptimizer: Adam with weight decay and learning rate scheduling.\nMetric: Validation AUC is monitored for early stopping and model checkpointing.",
    "3202562": "The simplest option is to change tf_efficientnet_b0 to tf_efficientnet_b2 or tf_efficientnetv2_b0.",
    "3202810": "Maybe try not deleting human voice, in recent discussion people say removing humvoice somehow worse their LB. \nI didn’t try humvoice yet, working with pseudo)",
    "3203296": "There is a discussion with the ideas from previous birdclef and it is stated in there that birdclef 2024 1st place used loop padding. Maybe trying that might help.\n\nOr combining the approaches, at least it might make the train data more diverse.\n\nLink: https://www.kaggle.com/competitions/birdclef-2025/discussion/572928",
    "3203303": "Based on my observations, ensembling cross-validated models might be a promising direction. While a single b0 model provides a strong baseline, my experiments showed that mean-ensembling multiple instances tended to yield a modest improvement in validation consistency. Would appreciate hearing others' experiences with model stacking here!",
    "3203346": "I'm doing similar things. But for me, different versions of three channel inputs did not improve the results. Therefore, I went back to a single spectrogram as an input. Also, I'm using a different classification head, which consists of two layers and has a residual connection. With that I was able to reach 0.819 (IIRC) for the efficient net b0",
    "3203351": "And I used soft labels in conjunction with a symmetric BCE loss.",
    "3203537": "Do you try thome other think?",
    "3210534": "halo,for your work,the other two models improve how much LB?",
    "3210551": "0.8 to 0.818",
    "3231086": "Loss & Optimization what mean ?"
  },
  "source": "meta"
}