{
  "id": 583308,
  "title": "Raw wave not work.",
  "url": "/competitions/birdclef-2025/discussion/583308",
  "author_name": "",
  "post_date": "2025-06-06T02:48:33.327341300Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I attempted to replicate last year's successful ensemble strategy that combined spectral features with raw waveform models. Unfortunately, the raw waveform component underperformed this time around. The raw wave model performance was 0.855( not sure , it's in my memory?) on the public leaderboard using EfficientNet-B5, achieved about a month ago. Following this result, I decided to abandon further raw waveform experiments.</p>\n<p>An interesting observation is that SED models consistently outperform conventional CNN architectures. I'm wondering if anyone has managed to achieve competitive scores using raw waveform methods?</p>\n<p>For me, the only things that work are sed model , pse label, params with mel transform and random seed.</p>\n<p>My best single model is 0.884 with hgnetv2_b0( raw 25 train data with human voice remove, random 5s trainning ),   and same mel-spectrum params as last year. </p>\n<pre><code>.wave_transform = nn.Sequential(\n\n            torchaudio.transforms.MelSpectrogram(\n                ,\n                n_mels=,\n                f_min=,\n                f_max=,\n                n_fft=*,\n                hop_length=,\n                normalized=,\n            ),\n\n            torchaudio.transforms.AmplitudeToDB(top_db=)\n\n        )\n</code></pre>\n<p>Looking forward to seeing the winning solutions!</p>",
  "messages": [
    {
      "id": "3218247",
      "postDate": "06/06/2025 02:48:33",
      "content": "<p>I attempted to replicate last year's successful ensemble strategy that combined spectral features with raw waveform models. Unfortunately, the raw waveform component underperformed this time around. The raw wave model performance was 0.855( not sure , it's in my memory?) on the public leaderboard using EfficientNet-B5, achieved about a month ago. Following this result, I decided to abandon further raw waveform experiments.</p>\n<p>An interesting observation is that SED models consistently outperform conventional CNN architectures. I'm wondering if anyone has managed to achieve competitive scores using raw waveform methods?</p>\n<p>For me, the only things that work are sed model , pse label, params with mel transform and random seed.</p>\n<p>My best single model is 0.884 with hgnetv2_b0( raw 25 train data with human voice remove, random 5s trainning ),   and same mel-spectrum params as last year. </p>\n<pre><code>.wave_transform = nn.Sequential(\n\n            torchaudio.transforms.MelSpectrogram(\n                ,\n                n_mels=,\n                f_min=,\n                f_max=,\n                n_fft=*,\n                hop_length=,\n                normalized=,\n            ),\n\n            torchaudio.transforms.AmplitudeToDB(top_db=)\n\n        )\n</code></pre>\n<p>Looking forward to seeing the winning solutions!</p>",
      "rawMarkdown": "I attempted to replicate last year's successful ensemble strategy that combined spectral features with raw waveform models. Unfortunately, the raw waveform component underperformed this time around. The raw wave model performance was 0.855( not sure , it's in my memory?) on the public leaderboard using EfficientNet-B5, achieved about a month ago. Following this result, I decided to abandon further raw waveform experiments.\n\nAn interesting observation is that SED models consistently outperform conventional CNN architectures. I'm wondering if anyone has managed to achieve competitive scores using raw waveform methods?\n\nFor me, the only things that work are sed model , pse label, params with mel transform and random seed.\n\nMy best single model is 0.884 with hgnetv2_b0( raw 25 train data with human voice remove, random 5s trainning ),   and same mel-spectrum params as last year. \n```python\nself.wave_transform = nn.Sequential(\n\n            torchaudio.transforms.MelSpectrogram(\n                32000,\n                n_mels=512,\n                f_min=0,\n                f_max=16000,\n                n_fft=2048*2,\n                hop_length=512,\n                normalized=True,\n            ),\n\n            torchaudio.transforms.AmplitudeToDB(top_db=80)\n\n        )\n```\n\n\nLooking forward to seeing the winning solutions!",
      "votes": null
    },
    {
      "id": "3218249",
      "postDate": "06/06/2025 02:52:54",
      "content": "<p><a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a> I tried to replicate your last year's solution on last year's data, but I was unable to. <br>\nI was wondering if you could share your training scripts, maybe for this year's competition as well? If possible.<br>\nI was able to get just 0.79 with raw wave, had to drop the idea.</p>",
      "rawMarkdown": "cooolz I tried to replicate your last year's solution on last year's data, but I was unable to. \nI was wondering if you could share your training scripts, maybe for this year's competition as well? If possible.\nI was able to get just 0.79 with raw wave, had to drop the idea.",
      "votes": null
    },
    {
      "id": "3218256",
      "postDate": "06/06/2025 03:07:21",
      "content": "<p><a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>  0.85 is good and maybe we can improve further. I tried  your last year's solution with raw model,  only got 0.71+ ,    I missed some important  things maybe.</p>",
      "rawMarkdown": "cooolz  0.85 is good and maybe we can improve further. I tried  your last year's solution with raw model,  only got 0.71+ ,    I missed some important  things maybe.",
      "votes": null
    },
    {
      "id": "3218261",
      "postDate": "06/06/2025 03:23:48",
      "content": "<p>Congrats! May I ask whether you make 1D into 2D in your raw model like last year? And have you tried 1D raw model this time?</p>",
      "rawMarkdown": "Congrats! May I ask whether you make 1D into 2D in your raw model like last year? And have you tried 1D raw model this time?",
      "votes": null
    },
    {
      "id": "3218263",
      "postDate": "06/06/2025 03:27:24",
      "content": "<p>My code is a mess this year. Here's last year's codebase: <a href=\"https://github.com/610265158/bird24\" target=\"_blank\">https://github.com/610265158/bird24</a>. Good luck with it, and thanks for sharing early in this competition, it's quite helpful..</p>",
      "rawMarkdown": "My code is a mess this year. Here's last year's codebase: https://github.com/610265158/bird24. Good luck with it, and thanks for sharing early in this competition, it's quite helpful..",
      "votes": null
    },
    {
      "id": "3218267",
      "postDate": "06/06/2025 03:32:20",
      "content": "<p>Yeah, I should explore this further. The SED model performs much better in this case. Another interesting finding from my experiments is that larger n_fft values work better, which suggests that a larger receptive field is indeed effective. However, EfficientNetB5 is quite time-consuming given the 90-minute limitation.</p>",
      "rawMarkdown": "Yeah, I should explore this further. The SED model performs much better in this case. Another interesting finding from my experiments is that larger n_fft values work better, which suggests that a larger receptive field is indeed effective. However, EfficientNetB5 is quite time-consuming given the 90-minute limitation.",
      "votes": null
    },
    {
      "id": "3218269",
      "postDate": "06/06/2025 03:34:16",
      "content": "<p>I haven't tried 1d model. It's a effnetb5.</p>",
      "rawMarkdown": "I haven't tried 1d model. It's a effnetb5.",
      "votes": null
    },
    {
      "id": "3218321",
      "postDate": "06/06/2025 04:48:11",
      "content": "<p>Thanks for following that up! I was really curious to see how the raw waveform strategies would fare this year, especially in ensemble with spectral models.</p>",
      "rawMarkdown": "Thanks for following that up! I was really curious to see how the raw waveform strategies would fare this year, especially in ensemble with spectral models.",
      "votes": null
    },
    {
      "id": "3218646",
      "postDate": "06/06/2025 13:43:28",
      "content": "<p>Wow 0.855 is very impressive. I tried raw signal model with hgnetv3_b3 but can only get 0.70.<br>\nBut I found that raw signal model is a choice for model diversity when performing ensemble. I trained an expert model and Putting this model into ensemble pipeline improves LB.</p>",
      "rawMarkdown": "Wow 0.855 is very impressive. I tried raw signal model with hgnetv3_b3 but can only get 0.70.\nBut I found that raw signal model is a choice for model diversity when performing ensemble. I trained an expert model and Putting this model into ensemble pipeline improves LB.",
      "votes": null
    },
    {
      "id": "3218647",
      "postDate": "06/06/2025 13:44:46",
      "content": "<p>It works in ensemble. I trained an aves expert model (which results to a similar problem set as last year) and it improved ensemble score in our experiment</p>",
      "rawMarkdown": "It works in ensemble. I trained an aves expert model (which results to a similar problem set as last year) and it improved ensemble score in our experiment",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218249,
      "author_name": "salmanahmedtamu",
      "author_url": "",
      "post_date": "06/06/2025 02:52:54",
      "content": "<p><a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a> I tried to replicate your last year's solution on last year's data, but I was unable to. <br>\nI was wondering if you could share your training scripts, maybe for this year's competition as well? If possible.<br>\nI was able to get just 0.79 with raw wave, had to drop the idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218263,
          "author_name": "cooolz",
          "author_url": "",
          "post_date": "06/06/2025 03:27:24",
          "content": "<p>My code is a mess this year. Here's last year's codebase: <a href=\"https://github.com/610265158/bird24\" target=\"_blank\">https://github.com/610265158/bird24</a>. Good luck with it, and thanks for sharing early in this competition, it's quite helpful..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218256,
      "author_name": "lihaoweicvch",
      "author_url": "",
      "post_date": "06/06/2025 03:07:21",
      "content": "<p><a href=\"https://www.kaggle.com/cooolz\" target=\"_blank\">@cooolz</a>  0.85 is good and maybe we can improve further. I tried  your last year's solution with raw model,  only got 0.71+ ,    I missed some important  things maybe.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218267,
          "author_name": "cooolz",
          "author_url": "",
          "post_date": "06/06/2025 03:32:20",
          "content": "<p>Yeah, I should explore this further. The SED model performs much better in this case. Another interesting finding from my experiments is that larger n_fft values work better, which suggests that a larger receptive field is indeed effective. However, EfficientNetB5 is quite time-consuming given the 90-minute limitation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218261,
      "author_name": "kurisew",
      "author_url": "",
      "post_date": "06/06/2025 03:23:48",
      "content": "<p>Congrats! May I ask whether you make 1D into 2D in your raw model like last year? And have you tried 1D raw model this time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218269,
          "author_name": "cooolz",
          "author_url": "",
          "post_date": "06/06/2025 03:34:16",
          "content": "<p>I haven't tried 1d model. It's a effnetb5.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218321,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "06/06/2025 04:48:11",
      "content": "<p>Thanks for following that up! I was really curious to see how the raw waveform strategies would fare this year, especially in ensemble with spectral models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218647,
          "author_name": "honglihang",
          "author_url": "",
          "post_date": "06/06/2025 13:44:46",
          "content": "<p>It works in ensemble. I trained an aves expert model (which results to a similar problem set as last year) and it improved ensemble score in our experiment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218646,
      "author_name": "honglihang",
      "author_url": "",
      "post_date": "06/06/2025 13:43:28",
      "content": "<p>Wow 0.855 is very impressive. I tried raw signal model with hgnetv3_b3 but can only get 0.70.<br>\nBut I found that raw signal model is a choice for model diversity when performing ensemble. I trained an expert model and Putting this model into ensemble pipeline improves LB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3218247": "I attempted to replicate last year's successful ensemble strategy that combined spectral features with raw waveform models. Unfortunately, the raw waveform component underperformed this time around. The raw wave model performance was 0.855( not sure , it's in my memory?) on the public leaderboard using EfficientNet-B5, achieved about a month ago. Following this result, I decided to abandon further raw waveform experiments.\n\nAn interesting observation is that SED models consistently outperform conventional CNN architectures. I'm wondering if anyone has managed to achieve competitive scores using raw waveform methods?\n\nFor me, the only things that work are sed model , pse label, params with mel transform and random seed.\n\nMy best single model is 0.884 with hgnetv2_b0( raw 25 train data with human voice remove, random 5s trainning ),   and same mel-spectrum params as last year. \n```python\nself.wave_transform = nn.Sequential(\n\n            torchaudio.transforms.MelSpectrogram(\n                32000,\n                n_mels=512,\n                f_min=0,\n                f_max=16000,\n                n_fft=2048*2,\n                hop_length=512,\n                normalized=True,\n            ),\n\n            torchaudio.transforms.AmplitudeToDB(top_db=80)\n\n        )\n```\n\n\nLooking forward to seeing the winning solutions!",
    "3218249": "cooolz I tried to replicate your last year's solution on last year's data, but I was unable to. \nI was wondering if you could share your training scripts, maybe for this year's competition as well? If possible.\nI was able to get just 0.79 with raw wave, had to drop the idea.",
    "3218256": "cooolz  0.85 is good and maybe we can improve further. I tried  your last year's solution with raw model,  only got 0.71+ ,    I missed some important  things maybe.",
    "3218261": "Congrats! May I ask whether you make 1D into 2D in your raw model like last year? And have you tried 1D raw model this time?",
    "3218263": "My code is a mess this year. Here's last year's codebase: https://github.com/610265158/bird24. Good luck with it, and thanks for sharing early in this competition, it's quite helpful..",
    "3218267": "Yeah, I should explore this further. The SED model performs much better in this case. Another interesting finding from my experiments is that larger n_fft values work better, which suggests that a larger receptive field is indeed effective. However, EfficientNetB5 is quite time-consuming given the 90-minute limitation.",
    "3218269": "I haven't tried 1d model. It's a effnetb5.",
    "3218321": "Thanks for following that up! I was really curious to see how the raw waveform strategies would fare this year, especially in ensemble with spectral models.",
    "3218646": "Wow 0.855 is very impressive. I tried raw signal model with hgnetv3_b3 but can only get 0.70.\nBut I found that raw signal model is a choice for model diversity when performing ensemble. I trained an expert model and Putting this model into ensemble pipeline improves LB.",
    "3218647": "It works in ensemble. I trained an aves expert model (which results to a similar problem set as last year) and it improved ensemble score in our experiment"
  },
  "source": "meta"
}