{
  "id": 497050,
  "title": "How should we use unlabeled_soundscapes?",
  "url": "/competitions/birdclef-2024/discussion/497050",
  "author_name": "",
  "post_date": "2024-04-23T11:53:59.415803900Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I want to ask how we use the unlabeled_soundscapes in this competition. We do not know the ground truth, so cannot know the exact model performance on it even though do inference on it.</p>",
  "messages": [
    {
      "id": "2769533",
      "postDate": "04/23/2024 11:53:59",
      "content": "<p>I want to ask how we use the unlabeled_soundscapes in this competition. We do not know the ground truth, so cannot know the exact model performance on it even though do inference on it.</p>",
      "rawMarkdown": "I want to ask how we use the unlabeled_soundscapes in this competition. We do not know the ground truth, so cannot know the exact model performance on it even though do inference on it.",
      "votes": null
    },
    {
      "id": "2769574",
      "postDate": "04/23/2024 12:42:24",
      "content": "<p>You can use them by predicting on them and using the results as labels - Pseudolabeling. If you have the time to do it, this might help a lot! </p>\n<p>Best,<br>\nJan</p>",
      "rawMarkdown": "You can use them by predicting on them and using the results as labels - Pseudolabeling. If you have the time to do it, this might help a lot! \n\nBest,\nJan",
      "votes": null
    },
    {
      "id": "2769577",
      "postDate": "04/23/2024 12:46:10",
      "content": "<p>Thanks, Is this the only way we leverage them? If we use pseudo labels, It may make model worse?</p>",
      "rawMarkdown": "Thanks, Is this the only way we leverage them? If we use pseudo labels, It may make model worse?",
      "votes": null
    },
    {
      "id": "2769652",
      "postDate": "04/23/2024 13:33:59",
      "content": "<p>I think using pseudo labels is not always easy and people much more experienced than me can explain much better but you can read into it and try - well done, it should improve your model! I habe no idea what else to do with the data.</p>\n<p>Best,<br>\nJan</p>",
      "rawMarkdown": "I think using pseudo labels is not always easy and people much more experienced than me can explain much better but you can read into it and try - well done, it should improve your model! I habe no idea what else to do with the data.\n\nBest,\nJan",
      "votes": null
    },
    {
      "id": "2769720",
      "postDate": "04/23/2024 14:10:47",
      "content": "<p>you can also use them as an augmentation, add them as background noise, this will help with the domain shift hopefully.</p>",
      "rawMarkdown": "you can also use them as an augmentation, add them as background noise, this will help with the domain shift hopefully.",
      "votes": null
    },
    {
      "id": "2773677",
      "postDate": "04/24/2024 20:44:36",
      "content": "<p>If you're doing a Wav2Vec like approach, you could do pre-training on the unlabeled soundscapes: <a href=\"https://huggingface.co/docs/transformers/en/model_doc/wav2vec2#transformers.Wav2Vec2ForPreTraining\" target=\"_blank\">https://huggingface.co/docs/transformers/en/model_doc/wav2vec2#transformers.Wav2Vec2ForPreTraining</a></p>\n<p>I think in past competitions people have done a ton of pre-training, but that might have been with spectrogram like models. Could probably plug in spectrograms, mask 'em, and pre-train that way. Make a good feature extractor and fine tune over that.</p>\n<p>Though I'm less keen on audio pre-training than others.</p>",
      "rawMarkdown": "If you're doing a Wav2Vec like approach, you could do pre-training on the unlabeled soundscapes: https://huggingface.co/docs/transformers/en/model_doc/wav2vec2#transformers.Wav2Vec2ForPreTraining\n\nI think in past competitions people have done a ton of pre-training, but that might have been with spectrogram like models. Could probably plug in spectrograms, mask 'em, and pre-train that way. Make a good feature extractor and fine tune over that.\n\nThough I'm less keen on audio pre-training than others.",
      "votes": null
    },
    {
      "id": "2781182",
      "postDate": "04/28/2024 16:18:29",
      "content": "<p>Do you think it would work even when unlabeled soundscape has \"no birds\" segment?<br>\nIn the wav2vec paper, they mention the pretraining was  using 960 hr librispeech without transcription (because they want to learn speech representation). In our case, we do not have that much birds audio. </p>",
      "rawMarkdown": "Do you think it would work even when unlabeled soundscape has \"no birds\" segment?\nIn the wav2vec paper, they mention the pretraining was  using 960 hr librispeech without transcription (because they want to learn speech representation). In our case, we do not have that much birds audio.",
      "votes": null
    },
    {
      "id": "2783635",
      "postDate": "04/29/2024 20:59:21",
      "content": "<p>That's a good point. I don't think we have nearly as much audio data here as there is speech data in the world, however, my intuition (as faulty as it may be) makes me think doing some mass pre-training on the soundscapes we have here, on the previous competitions, on the bird files this year, and hell, even random woodland sounds you can find, could yield some results…maybe. </p>\n<p>This is all just semi-edumacated guessing, but the more raw data you pre-train on, the fewer samples you may need to fine-tune for specific birds and get good accuracy. You have that solid base of all sounds you might hear in these clips from birds to leaves rustling to someone walking or talking or breathing and anything else that might crop up. Stick a classification head on those deep features, and boom, might help toss out some of that noise.</p>\n<p>Again, all speculation, but I plan to burn some of my baja blast budget this month on GPU's and trying it.</p>",
      "rawMarkdown": "That's a good point. I don't think we have nearly as much audio data here as there is speech data in the world, however, my intuition (as faulty as it may be) makes me think doing some mass pre-training on the soundscapes we have here, on the previous competitions, on the bird files this year, and hell, even random woodland sounds you can find, could yield some results...maybe. \n\nThis is all just semi-edumacated guessing, but the more raw data you pre-train on, the fewer samples you may need to fine-tune for specific birds and get good accuracy. You have that solid base of all sounds you might hear in these clips from birds to leaves rustling to someone walking or talking or breathing and anything else that might crop up. Stick a classification head on those deep features, and boom, might help toss out some of that noise.\n\nAgain, all speculation, but I plan to burn some of my baja blast budget this month on GPU's and trying it.",
      "votes": null
    },
    {
      "id": "2783733",
      "postDate": "04/29/2024 23:13:53",
      "content": "<p>I just posted a notebook / write-up on pre-training here:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\" target=\"_blank\">https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes</a><br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498876\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/498876</a></p>\n<p>Definitely able to generate labels / pre-train on them.  No promises if it actually does anything useful.</p>",
      "rawMarkdown": "I just posted a notebook / write-up on pre-training here:\n\nhttps://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/498876\n\nDefinitely able to generate labels / pre-train on them.  No promises if it actually does anything useful.",
      "votes": null
    },
    {
      "id": "2783870",
      "postDate": "04/30/2024 02:06:07",
      "content": "<p>There's a nice audio pretraining framework I'm interested in<br>\n<a href=\"https://github.com/nttcslab/m2d/tree/master\" target=\"_blank\">https://github.com/nttcslab/m2d/tree/master</a><br>\nthese days I can't find time off my day job to look at it </p>",
      "rawMarkdown": "There's a nice audio pretraining framework I'm interested in\nhttps://github.com/nttcslab/m2d/tree/master\nthese days I can't find time off my day job to look at it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2769574,
      "author_name": "janbrederecke",
      "author_url": "",
      "post_date": "04/23/2024 12:42:24",
      "content": "<p>You can use them by predicting on them and using the results as labels - Pseudolabeling. If you have the time to do it, this might help a lot! </p>\n<p>Best,<br>\nJan</p>",
      "votes": null,
      "replies": [
        {
          "id": 2769577,
          "author_name": "yiding111",
          "author_url": "",
          "post_date": "04/23/2024 12:46:10",
          "content": "<p>Thanks, Is this the only way we leverage them? If we use pseudo labels, It may make model worse?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2769652,
              "author_name": "janbrederecke",
              "author_url": "",
              "post_date": "04/23/2024 13:33:59",
              "content": "<p>I think using pseudo labels is not always easy and people much more experienced than me can explain much better but you can read into it and try - well done, it should improve your model! I habe no idea what else to do with the data.</p>\n<p>Best,<br>\nJan</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2769720,
              "author_name": "janmpia",
              "author_url": "",
              "post_date": "04/23/2024 14:10:47",
              "content": "<p>you can also use them as an augmentation, add them as background noise, this will help with the domain shift hopefully.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2773677,
      "author_name": "msthil",
      "author_url": "",
      "post_date": "04/24/2024 20:44:36",
      "content": "<p>If you're doing a Wav2Vec like approach, you could do pre-training on the unlabeled soundscapes: <a href=\"https://huggingface.co/docs/transformers/en/model_doc/wav2vec2#transformers.Wav2Vec2ForPreTraining\" target=\"_blank\">https://huggingface.co/docs/transformers/en/model_doc/wav2vec2#transformers.Wav2Vec2ForPreTraining</a></p>\n<p>I think in past competitions people have done a ton of pre-training, but that might have been with spectrogram like models. Could probably plug in spectrograms, mask 'em, and pre-train that way. Make a good feature extractor and fine tune over that.</p>\n<p>Though I'm less keen on audio pre-training than others.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2781182,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "04/28/2024 16:18:29",
          "content": "<p>Do you think it would work even when unlabeled soundscape has \"no birds\" segment?<br>\nIn the wav2vec paper, they mention the pretraining was  using 960 hr librispeech without transcription (because they want to learn speech representation). In our case, we do not have that much birds audio. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2783635,
              "author_name": "msthil",
              "author_url": "",
              "post_date": "04/29/2024 20:59:21",
              "content": "<p>That's a good point. I don't think we have nearly as much audio data here as there is speech data in the world, however, my intuition (as faulty as it may be) makes me think doing some mass pre-training on the soundscapes we have here, on the previous competitions, on the bird files this year, and hell, even random woodland sounds you can find, could yield some results…maybe. </p>\n<p>This is all just semi-edumacated guessing, but the more raw data you pre-train on, the fewer samples you may need to fine-tune for specific birds and get good accuracy. You have that solid base of all sounds you might hear in these clips from birds to leaves rustling to someone walking or talking or breathing and anything else that might crop up. Stick a classification head on those deep features, and boom, might help toss out some of that noise.</p>\n<p>Again, all speculation, but I plan to burn some of my baja blast budget this month on GPU's and trying it.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2783870,
                  "author_name": "nyleve",
                  "author_url": "",
                  "post_date": "04/30/2024 02:06:07",
                  "content": "<p>There's a nice audio pretraining framework I'm interested in<br>\n<a href=\"https://github.com/nttcslab/m2d/tree/master\" target=\"_blank\">https://github.com/nttcslab/m2d/tree/master</a><br>\nthese days I can't find time off my day job to look at it </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2783733,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "04/29/2024 23:13:53",
      "content": "<p>I just posted a notebook / write-up on pre-training here:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\" target=\"_blank\">https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes</a><br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498876\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/498876</a></p>\n<p>Definitely able to generate labels / pre-train on them.  No promises if it actually does anything useful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2769533": "I want to ask how we use the unlabeled_soundscapes in this competition. We do not know the ground truth, so cannot know the exact model performance on it even though do inference on it.",
    "2769574": "You can use them by predicting on them and using the results as labels - Pseudolabeling. If you have the time to do it, this might help a lot! \n\nBest,\nJan",
    "2769577": "Thanks, Is this the only way we leverage them? If we use pseudo labels, It may make model worse?",
    "2769652": "I think using pseudo labels is not always easy and people much more experienced than me can explain much better but you can read into it and try - well done, it should improve your model! I habe no idea what else to do with the data.\n\nBest,\nJan",
    "2769720": "you can also use them as an augmentation, add them as background noise, this will help with the domain shift hopefully.",
    "2773677": "If you're doing a Wav2Vec like approach, you could do pre-training on the unlabeled soundscapes: https://huggingface.co/docs/transformers/en/model_doc/wav2vec2#transformers.Wav2Vec2ForPreTraining\n\nI think in past competitions people have done a ton of pre-training, but that might have been with spectrogram like models. Could probably plug in spectrograms, mask 'em, and pre-train that way. Make a good feature extractor and fine tune over that.\n\nThough I'm less keen on audio pre-training than others.",
    "2781182": "Do you think it would work even when unlabeled soundscape has \"no birds\" segment?\nIn the wav2vec paper, they mention the pretraining was  using 960 hr librispeech without transcription (because they want to learn speech representation). In our case, we do not have that much birds audio.",
    "2783635": "That's a good point. I don't think we have nearly as much audio data here as there is speech data in the world, however, my intuition (as faulty as it may be) makes me think doing some mass pre-training on the soundscapes we have here, on the previous competitions, on the bird files this year, and hell, even random woodland sounds you can find, could yield some results...maybe. \n\nThis is all just semi-edumacated guessing, but the more raw data you pre-train on, the fewer samples you may need to fine-tune for specific birds and get good accuracy. You have that solid base of all sounds you might hear in these clips from birds to leaves rustling to someone walking or talking or breathing and anything else that might crop up. Stick a classification head on those deep features, and boom, might help toss out some of that noise.\n\nAgain, all speculation, but I plan to burn some of my baja blast budget this month on GPU's and trying it.",
    "2783733": "I just posted a notebook / write-up on pre-training here:\n\nhttps://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\nhttps://www.kaggle.com/competitions/birdclef-2024/discussion/498876\n\nDefinitely able to generate labels / pre-train on them.  No promises if it actually does anything useful.",
    "2783870": "There's a nice audio pretraining framework I'm interested in\nhttps://github.com/nttcslab/m2d/tree/master\nthese days I can't find time off my day job to look at it"
  },
  "source": "meta"
}