{
  "id": 498876,
  "title": "Unsupervised Birds: Pre-training on unlabeled soundscapes",
  "url": "/competitions/birdclef-2024/discussion/498876",
  "author_name": "",
  "post_date": "2024-04-29T23:05:48.951253900Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>So - the question of how to use the unlabeled train data has come up.  The unlabeled data is new to BirdCLEF 2024 - and I haven't noticed any notebooks on the topic.</p>\n<p>There is more than one option - but one approach is to use the data for unsupervised (or semi-supervised) learning.  </p>\n<p>I've put together a notebook demonstrating this here:<br>\n<a href=\"https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\" target=\"_blank\">https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes</a></p>\n<p>There is also a notebook with contiguous Mel Spectrograms for the unlabeled data here:<br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-24-contiguous-mels-unlabeled-soundscapes/\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-24-contiguous-mels-unlabeled-soundscapes/</a></p>\n<p>Our full business plan looks like:</p>\n<ol>\n<li><p>Generate Mel spectrograms from unlabeled features</p></li>\n<li><p>Use an Imagenet (efficientnet_b2) to generate features for 5-second cuts of the Mels</p></li>\n<li><p>Use KMeans clustering to segment the data into clusters that share similar features (pseudoclasses) - I went with 182 to match the number of species </p></li>\n<li><p>Train a model (also efficientnet_b2) on images we clustered on / associated pseudoclasses</p></li>\n<li><p>Further train the model on labeled data</p></li>\n<li><p>Move up the LB (Profit!)</p></li>\n</ol>\n<p>(Steps 5 and 6 are not yet covered in the notebook)</p>\n<p>The goal here is to produce a model that is pre-conditioned to recognize patterns in data that is similar to the test data.  There is no reasonable expectation that pseudoclasses map reliably to the actual species.  It is clear at least some pseudoclasses correspond more to certain kinds of background noise more than bird calls.</p>\n<p>The model did converge on the pseudoclasses.  41% accuracy may not sound high - but at nearly 200 classes - it's clearly at least doing something!</p>\n<p>The next step would be to do further training on the model with labeled data.  A lower learning rate and/or freezing some layers of the model might be good ideas to help preserve the pre-training.</p>\n<p>I haven't done any meaningful tests with the saved model attached to this notebook.  Please feel free to give it a shot!</p>",
  "messages": [
    {
      "id": "2783731",
      "postDate": "04/29/2024 23:05:48",
      "content": "<p>So - the question of how to use the unlabeled train data has come up.  The unlabeled data is new to BirdCLEF 2024 - and I haven't noticed any notebooks on the topic.</p>\n<p>There is more than one option - but one approach is to use the data for unsupervised (or semi-supervised) learning.  </p>\n<p>I've put together a notebook demonstrating this here:<br>\n<a href=\"https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\" target=\"_blank\">https://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes</a></p>\n<p>There is also a notebook with contiguous Mel Spectrograms for the unlabeled data here:<br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-24-contiguous-mels-unlabeled-soundscapes/\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-24-contiguous-mels-unlabeled-soundscapes/</a></p>\n<p>Our full business plan looks like:</p>\n<ol>\n<li><p>Generate Mel spectrograms from unlabeled features</p></li>\n<li><p>Use an Imagenet (efficientnet_b2) to generate features for 5-second cuts of the Mels</p></li>\n<li><p>Use KMeans clustering to segment the data into clusters that share similar features (pseudoclasses) - I went with 182 to match the number of species </p></li>\n<li><p>Train a model (also efficientnet_b2) on images we clustered on / associated pseudoclasses</p></li>\n<li><p>Further train the model on labeled data</p></li>\n<li><p>Move up the LB (Profit!)</p></li>\n</ol>\n<p>(Steps 5 and 6 are not yet covered in the notebook)</p>\n<p>The goal here is to produce a model that is pre-conditioned to recognize patterns in data that is similar to the test data.  There is no reasonable expectation that pseudoclasses map reliably to the actual species.  It is clear at least some pseudoclasses correspond more to certain kinds of background noise more than bird calls.</p>\n<p>The model did converge on the pseudoclasses.  41% accuracy may not sound high - but at nearly 200 classes - it's clearly at least doing something!</p>\n<p>The next step would be to do further training on the model with labeled data.  A lower learning rate and/or freezing some layers of the model might be good ideas to help preserve the pre-training.</p>\n<p>I haven't done any meaningful tests with the saved model attached to this notebook.  Please feel free to give it a shot!</p>",
      "rawMarkdown": "So - the question of how to use the unlabeled train data has come up.  The unlabeled data is new to BirdCLEF 2024 - and I haven't noticed any notebooks on the topic.\n\nThere is more than one option - but one approach is to use the data for unsupervised (or semi-supervised) learning.  \n\nI've put together a notebook demonstrating this here:\nhttps://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\n\nThere is also a notebook with contiguous Mel Spectrograms for the unlabeled data here:\nhttps://www.kaggle.com/code/richolson/birdclef-24-contiguous-mels-unlabeled-soundscapes/\n\nOur full business plan looks like:\n\n1. Generate Mel spectrograms from unlabeled features\n\n2. Use an Imagenet (efficientnet_b2) to generate features for 5-second cuts of the Mels\n\n3. Use KMeans clustering to segment the data into clusters that share similar features (pseudoclasses) - I went with 182 to match the number of species \n\n4. Train a model (also efficientnet_b2) on images we clustered on / associated pseudoclasses\n\n5. Further train the model on labeled data\n\n6. Move up the LB (Profit!)\n\n(Steps 5 and 6 are not yet covered in the notebook)\n\nThe goal here is to produce a model that is pre-conditioned to recognize patterns in data that is similar to the test data.  There is no reasonable expectation that pseudoclasses map reliably to the actual species.  It is clear at least some pseudoclasses correspond more to certain kinds of background noise more than bird calls.\n\nThe model did converge on the pseudoclasses.  41% accuracy may not sound high - but at nearly 200 classes - it's clearly at least doing something!\n\nThe next step would be to do further training on the model with labeled data.  A lower learning rate and/or freezing some layers of the model might be good ideas to help preserve the pre-training.\n\nI haven't done any meaningful tests with the saved model attached to this notebook.  Please feel free to give it a shot!",
      "votes": null
    },
    {
      "id": "2785487",
      "postDate": "04/30/2024 19:52:20",
      "content": "<p>I tried fine-tuning the pre-trained model using the same datasets I had with efficientnet_b2 previously.</p>\n<p>I tried a few different hyperparameter variations - but no combination resulted in an LB score improvement.</p>\n<p>I'm not convinced this isn't a valid approach - but it doesn't seem to be an easy win.</p>",
      "rawMarkdown": "I tried fine-tuning the pre-trained model using the same datasets I had with efficientnet_b2 previously.\n\nI tried a few different hyperparameter variations - but no combination resulted in an LB score improvement.\n\nI'm not convinced this isn't a valid approach - but it doesn't seem to be an easy win.",
      "votes": null
    },
    {
      "id": "2785507",
      "postDate": "04/30/2024 20:30:29",
      "content": "<p>Nonetheless, thank you for sharing all your interesting work! :) </p>",
      "rawMarkdown": "Nonetheless, thank you for sharing all your interesting work! :)",
      "votes": null
    },
    {
      "id": "2785517",
      "postDate": "04/30/2024 20:50:46",
      "content": "<p>Darn. Good stuff though. Worthy experiment in my opinion.</p>",
      "rawMarkdown": "Darn. Good stuff though. Worthy experiment in my opinion.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2785487,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "04/30/2024 19:52:20",
      "content": "<p>I tried fine-tuning the pre-trained model using the same datasets I had with efficientnet_b2 previously.</p>\n<p>I tried a few different hyperparameter variations - but no combination resulted in an LB score improvement.</p>\n<p>I'm not convinced this isn't a valid approach - but it doesn't seem to be an easy win.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2785507,
          "author_name": "janbrederecke",
          "author_url": "",
          "post_date": "04/30/2024 20:30:29",
          "content": "<p>Nonetheless, thank you for sharing all your interesting work! :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2785517,
          "author_name": "msthil",
          "author_url": "",
          "post_date": "04/30/2024 20:50:46",
          "content": "<p>Darn. Good stuff though. Worthy experiment in my opinion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2783731": "So - the question of how to use the unlabeled train data has come up.  The unlabeled data is new to BirdCLEF 2024 - and I haven't noticed any notebooks on the topic.\n\nThere is more than one option - but one approach is to use the data for unsupervised (or semi-supervised) learning.  \n\nI've put together a notebook demonstrating this here:\nhttps://www.kaggle.com/code/richolson/unsupervised-birds-using-unlabeled-soundscapes\n\nThere is also a notebook with contiguous Mel Spectrograms for the unlabeled data here:\nhttps://www.kaggle.com/code/richolson/birdclef-24-contiguous-mels-unlabeled-soundscapes/\n\nOur full business plan looks like:\n\n1. Generate Mel spectrograms from unlabeled features\n\n2. Use an Imagenet (efficientnet_b2) to generate features for 5-second cuts of the Mels\n\n3. Use KMeans clustering to segment the data into clusters that share similar features (pseudoclasses) - I went with 182 to match the number of species \n\n4. Train a model (also efficientnet_b2) on images we clustered on / associated pseudoclasses\n\n5. Further train the model on labeled data\n\n6. Move up the LB (Profit!)\n\n(Steps 5 and 6 are not yet covered in the notebook)\n\nThe goal here is to produce a model that is pre-conditioned to recognize patterns in data that is similar to the test data.  There is no reasonable expectation that pseudoclasses map reliably to the actual species.  It is clear at least some pseudoclasses correspond more to certain kinds of background noise more than bird calls.\n\nThe model did converge on the pseudoclasses.  41% accuracy may not sound high - but at nearly 200 classes - it's clearly at least doing something!\n\nThe next step would be to do further training on the model with labeled data.  A lower learning rate and/or freezing some layers of the model might be good ideas to help preserve the pre-training.\n\nI haven't done any meaningful tests with the saved model attached to this notebook.  Please feel free to give it a shot!",
    "2785487": "I tried fine-tuning the pre-trained model using the same datasets I had with efficientnet_b2 previously.\n\nI tried a few different hyperparameter variations - but no combination resulted in an LB score improvement.\n\nI'm not convinced this isn't a valid approach - but it doesn't seem to be an easy win.",
    "2785507": "Nonetheless, thank you for sharing all your interesting work! :)",
    "2785517": "Darn. Good stuff though. Worthy experiment in my opinion."
  },
  "source": "meta"
}