{
  "id": 572082,
  "title": "Attention is NOT all you need? ",
  "url": "/competitions/birdclef-2025/discussion/572082",
  "author_name": "",
  "post_date": "2025-04-07T13:49:54.521523900Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Preliminary results with \"convnext-femto\" give me terrible results. Is anyone experiencing the same?</p>",
  "messages": [
    {
      "id": "3173091",
      "postDate": "04/07/2025 13:49:54",
      "content": "<p>Preliminary results with \"convnext-femto\" give me terrible results. Is anyone experiencing the same?</p>",
      "rawMarkdown": "Preliminary results with \"convnext-femto\" give me terrible results. Is anyone experiencing the same?",
      "votes": null
    },
    {
      "id": "3173483",
      "postDate": "04/08/2025 01:12:38",
      "content": "<p>I've also had very poor results with Hubert training.</p>",
      "rawMarkdown": "I've also had very poor results with Hubert training.",
      "votes": null
    },
    {
      "id": "3173633",
      "postDate": "04/08/2025 06:56:12",
      "content": "<p>Hubert, what kind of model is this?</p>",
      "rawMarkdown": "Hubert, what kind of model is this?",
      "votes": null
    },
    {
      "id": "3173683",
      "postDate": "04/08/2025 08:06:16",
      "content": "<p>I'm using Hubert-base. probably because I didn't tweak the parameters, I'm training EfficientNet at about 70% or so local MAP, but training Hubert has only 1% local MAP.</p>",
      "rawMarkdown": "I'm using Hubert-base. probably because I didn't tweak the parameters, I'm training EfficientNet at about 70% or so local MAP, but training Hubert has only 1% local MAP.",
      "votes": null
    },
    {
      "id": "3177476",
      "postDate": "04/12/2025 20:11:02",
      "content": "<p>I guess you mean this kind of Hubert model to extract audio features?<br>\n<a href=\"https://pytorch.org/audio/main/generated/torchaudio.models.HuBERTPretrainModel.html\" target=\"_blank\">https://pytorch.org/audio/main/generated/torchaudio.models.HuBERTPretrainModel.html</a><br>\nOr maybe you don't transform audio features into 2D images and directly do 1D classification?</p>",
      "rawMarkdown": "I guess you mean this kind of Hubert model to extract audio features?\nhttps://pytorch.org/audio/main/generated/torchaudio.models.HuBERTPretrainModel.html\nOr maybe you don't transform audio features into 2D images and directly do 1D classification?",
      "votes": null
    },
    {
      "id": "3177582",
      "postDate": "04/13/2025 02:14:33",
      "content": "<p>Yes, I didn't put the audio through a mel transform, I put (batch_size, 16000*5) directly into Hubert.</p>",
      "rawMarkdown": "Yes, I didn't put the audio through a mel transform, I put (batch_size, 16000*5) directly into Hubert.",
      "votes": null
    },
    {
      "id": "3177774",
      "postDate": "04/13/2025 08:49:55",
      "content": "<p>That's a nice variation. </p>",
      "rawMarkdown": "That's a nice variation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3173483,
      "author_name": "agcsdedf",
      "author_url": "",
      "post_date": "04/08/2025 01:12:38",
      "content": "<p>I've also had very poor results with Hubert training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3173633,
          "author_name": "stefanoclss",
          "author_url": "",
          "post_date": "04/08/2025 06:56:12",
          "content": "<p>Hubert, what kind of model is this?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3173683,
              "author_name": "agcsdedf",
              "author_url": "",
              "post_date": "04/08/2025 08:06:16",
              "content": "<p>I'm using Hubert-base. probably because I didn't tweak the parameters, I'm training EfficientNet at about 70% or so local MAP, but training Hubert has only 1% local MAP.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3177476,
                  "author_name": "yassinealouini",
                  "author_url": "",
                  "post_date": "04/12/2025 20:11:02",
                  "content": "<p>I guess you mean this kind of Hubert model to extract audio features?<br>\n<a href=\"https://pytorch.org/audio/main/generated/torchaudio.models.HuBERTPretrainModel.html\" target=\"_blank\">https://pytorch.org/audio/main/generated/torchaudio.models.HuBERTPretrainModel.html</a><br>\nOr maybe you don't transform audio features into 2D images and directly do 1D classification?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3177582,
                      "author_name": "agcsdedf",
                      "author_url": "",
                      "post_date": "04/13/2025 02:14:33",
                      "content": "<p>Yes, I didn't put the audio through a mel transform, I put (batch_size, 16000*5) directly into Hubert.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3177774,
                          "author_name": "yassinealouini",
                          "author_url": "",
                          "post_date": "04/13/2025 08:49:55",
                          "content": "<p>That's a nice variation. </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3173091": "Preliminary results with \"convnext-femto\" give me terrible results. Is anyone experiencing the same?",
    "3173483": "I've also had very poor results with Hubert training.",
    "3173633": "Hubert, what kind of model is this?",
    "3173683": "I'm using Hubert-base. probably because I didn't tweak the parameters, I'm training EfficientNet at about 70% or so local MAP, but training Hubert has only 1% local MAP.",
    "3177476": "I guess you mean this kind of Hubert model to extract audio features?\nhttps://pytorch.org/audio/main/generated/torchaudio.models.HuBERTPretrainModel.html\nOr maybe you don't transform audio features into 2D images and directly do 1D classification?",
    "3177582": "Yes, I didn't put the audio through a mel transform, I put (batch_size, 16000*5) directly into Hubert.",
    "3177774": "That's a nice variation."
  },
  "source": "meta"
}