{
  "id": 579893,
  "title": "Anyone trying 1D models instead of 2D CNNs?",
  "url": "/competitions/birdclef-2025/discussion/579893",
  "author_name": "",
  "post_date": "2025-05-21T04:49:07.713786700Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I’ve been exploring several types of 1D approaches instead of 2D CNNs, but I haven’t been able to achieve better scores.<br>\nSpecifically, I tried stacked Conv1D blocks with a GRU for SED, using various input types:</p>\n<ul>\n<li>time-domain signal encoded by a Conv1D encoder</li>\n<li>STFT spectrogram</li>\n<li>MelSpectrogram </li>\n</ul>\n<p>I suspect one reason why this approach hasn’t worked well is that the models fail to learn sufficient robustness to input variations along the frequency axis, given the limited data in this competition.</p>\n<p>Is anyone else experimenting with 1D models? If so, how are your results?</p>",
  "messages": [
    {
      "id": "3206250",
      "postDate": "05/21/2025 04:49:07",
      "content": "<p>I’ve been exploring several types of 1D approaches instead of 2D CNNs, but I haven’t been able to achieve better scores.<br>\nSpecifically, I tried stacked Conv1D blocks with a GRU for SED, using various input types:</p>\n<ul>\n<li>time-domain signal encoded by a Conv1D encoder</li>\n<li>STFT spectrogram</li>\n<li>MelSpectrogram </li>\n</ul>\n<p>I suspect one reason why this approach hasn’t worked well is that the models fail to learn sufficient robustness to input variations along the frequency axis, given the limited data in this competition.</p>\n<p>Is anyone else experimenting with 1D models? If so, how are your results?</p>",
      "rawMarkdown": "I’ve been exploring several types of 1D approaches instead of 2D CNNs, but I haven’t been able to achieve better scores.\nSpecifically, I tried stacked Conv1D blocks with a GRU for SED, using various input types:\n- time-domain signal encoded by a Conv1D encoder\n- STFT spectrogram\n- MelSpectrogram \n\nI suspect one reason why this approach hasn’t worked well is that the models fail to learn sufficient robustness to input variations along the frequency axis, given the limited data in this competition.\n\nIs anyone else experimenting with 1D models? If so, how are your results?",
      "votes": null
    },
    {
      "id": "3206287",
      "postDate": "05/21/2025 05:59:01",
      "content": "<p>It didn't work for me, I have tried tons of experiments on this and results are much worse than 2D CNNs.</p>\n<p>Did you notice drop in performance if you augment raw audio from train audio set? </p>",
      "rawMarkdown": "It didn't work for me, I have tried tons of experiments on this and results are much worse than 2D CNNs.\n\nDid you notice drop in performance if you augment raw audio from train audio set?",
      "votes": null
    },
    {
      "id": "3206298",
      "postDate": "05/21/2025 06:18:26",
      "content": "<p>It's good to know I'm not the only one experiencing that.<br>\nIn my case, I applied data augmentation in the feature space, not in the time domain. Performance did improve, but only slightly.</p>",
      "rawMarkdown": "It's good to know I'm not the only one experiencing that.\nIn my case, I applied data augmentation in the feature space, not in the time domain. Performance did improve, but only slightly.",
      "votes": null
    },
    {
      "id": "3206300",
      "postDate": "05/21/2025 06:26:36",
      "content": "<p>Thank you. Did you apply it on Raw Audio? Or just the mel spectograms?</p>\n<p>The most difficult thing for me is to decide when to stop the training. For now I just train for 15 epochs with BS of 32, or 30 epochs with BS of 64. <br>\nHow did you decide?</p>",
      "rawMarkdown": "Thank you. Did you apply it on Raw Audio? Or just the mel spectograms?\n\nThe most difficult thing for me is to decide when to stop the training. For now I just train for 15 epochs with BS of 32, or 30 epochs with BS of 64. \nHow did you decide?",
      "votes": null
    },
    {
      "id": "3206311",
      "postDate": "05/21/2025 06:43:04",
      "content": "<p>I only applied augmentation in the feature space to raw audio.</p>\n<p>I also found it quite challenging. Since my architecture needed a certain number of epochs to train properly, I experimented with 20 or 50 epochs using cosine annealing, and submitted the last epoch.<br>\n I'm still exploring ways to stabilize the training, but I haven't succeeded yet.</p>",
      "rawMarkdown": "I only applied augmentation in the feature space to raw audio.\n\nI also found it quite challenging. Since my architecture needed a certain number of epochs to train properly, I experimented with 20 or 50 epochs using cosine annealing, and submitted the last epoch.\n I'm still exploring ways to stabilize the training, but I haven't succeeded yet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3206287,
      "author_name": "salmanahmedtamu",
      "author_url": "",
      "post_date": "05/21/2025 05:59:01",
      "content": "<p>It didn't work for me, I have tried tons of experiments on this and results are much worse than 2D CNNs.</p>\n<p>Did you notice drop in performance if you augment raw audio from train audio set? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3206298,
          "author_name": "zuoliao11",
          "author_url": "",
          "post_date": "05/21/2025 06:18:26",
          "content": "<p>It's good to know I'm not the only one experiencing that.<br>\nIn my case, I applied data augmentation in the feature space, not in the time domain. Performance did improve, but only slightly.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3206300,
              "author_name": "salmanahmedtamu",
              "author_url": "",
              "post_date": "05/21/2025 06:26:36",
              "content": "<p>Thank you. Did you apply it on Raw Audio? Or just the mel spectograms?</p>\n<p>The most difficult thing for me is to decide when to stop the training. For now I just train for 15 epochs with BS of 32, or 30 epochs with BS of 64. <br>\nHow did you decide?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3206311,
                  "author_name": "zuoliao11",
                  "author_url": "",
                  "post_date": "05/21/2025 06:43:04",
                  "content": "<p>I only applied augmentation in the feature space to raw audio.</p>\n<p>I also found it quite challenging. Since my architecture needed a certain number of epochs to train properly, I experimented with 20 or 50 epochs using cosine annealing, and submitted the last epoch.<br>\n I'm still exploring ways to stabilize the training, but I haven't succeeded yet.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3206250": "I’ve been exploring several types of 1D approaches instead of 2D CNNs, but I haven’t been able to achieve better scores.\nSpecifically, I tried stacked Conv1D blocks with a GRU for SED, using various input types:\n- time-domain signal encoded by a Conv1D encoder\n- STFT spectrogram\n- MelSpectrogram \n\nI suspect one reason why this approach hasn’t worked well is that the models fail to learn sufficient robustness to input variations along the frequency axis, given the limited data in this competition.\n\nIs anyone else experimenting with 1D models? If so, how are your results?",
    "3206287": "It didn't work for me, I have tried tons of experiments on this and results are much worse than 2D CNNs.\n\nDid you notice drop in performance if you augment raw audio from train audio set?",
    "3206298": "It's good to know I'm not the only one experiencing that.\nIn my case, I applied data augmentation in the feature space, not in the time domain. Performance did improve, but only slightly.",
    "3206300": "Thank you. Did you apply it on Raw Audio? Or just the mel spectograms?\n\nThe most difficult thing for me is to decide when to stop the training. For now I just train for 15 epochs with BS of 32, or 30 epochs with BS of 64. \nHow did you decide?",
    "3206311": "I only applied augmentation in the feature space to raw audio.\n\nI also found it quite challenging. Since my architecture needed a certain number of epochs to train properly, I experimented with 20 or 50 epochs using cosine annealing, and submitted the last epoch.\n I'm still exploring ways to stabilize the training, but I haven't succeeded yet."
  },
  "source": "meta"
}