{
  "id": 161521,
  "title": "Pre-trained ResNet-34 birdcall classifier",
  "url": "/competitions/birdsong-recognition/discussion/161521",
  "author_name": "",
  "post_date": "2020-06-25T06:30:10.960358800Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everyone! I have published a dataset with a trained ResNet-34 model to tackle this task with CNNs and Mel spectrograms.</p>\n\n<hr>\n\n<h3>TLDR</h3>\n\n<p>Notebook: <a href=\"https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet\">https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet</a>\nDataset: <a href=\"https://www.kaggle.com/tarunpaparaju/pretrained-resnet34-for-birdcall-classification\">https://www.kaggle.com/tarunpaparaju/pretrained-resnet34-for-birdcall-classification</a></p>\n\n<hr>\n\n<h3>Content</h3>\n\n<p>This dataset contains a <code>ResNet-34</code> model trained on Mel spectrograms from the <a href=\"https://www.kaggle.com/c/birdsong-recognition\">Cornell Birdcall Identification</a> dataset. It can be used to identify bird species from audio clips with high accuracy (around 55% on unseen clips) spanning 264 different species mentioned on <a href=\"https://www.xeno-canto.org/\">https://www.xeno-canto.org/</a>.</p>\n\n<hr>\n\n<h3>Usage</h3>\n\n<p>To use this pre-trained model with PyTorch, you first need to convert your audio clip to a Mel spectrogram image to feed into the model. Refer to <a href=\"https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet\">my kernel on birdcall classification</a> to understand how to generate these Mel spectrograms. Make sure your audio signal is of length <code>1000000</code> and set the spectrogram features to <code>256</code>. Then finally convert the image to a 3-channel version (repetition) and apply classic ImageNet normalization with <code>albumentations</code>. Once you convert the audio clip/s to Mel spectrograms, define the ResNet model and the load the pre-trained weights from this dataset. The code snippet below demonstrates how to set up the model with pre-trained weights. Now, you can use this model to classify bird species!</p>\n\n<hr>\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p>",
  "messages": [
    {
      "id": "900922",
      "postDate": "06/25/2020 06:30:10",
      "content": "<p>Hello everyone! I have published a dataset with a trained ResNet-34 model to tackle this task with CNNs and Mel spectrograms.</p>\n\n<hr>\n\n<h3>TLDR</h3>\n\n<p>Notebook: <a href=\"https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet\">https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet</a>\nDataset: <a href=\"https://www.kaggle.com/tarunpaparaju/pretrained-resnet34-for-birdcall-classification\">https://www.kaggle.com/tarunpaparaju/pretrained-resnet34-for-birdcall-classification</a></p>\n\n<hr>\n\n<h3>Content</h3>\n\n<p>This dataset contains a <code>ResNet-34</code> model trained on Mel spectrograms from the <a href=\"https://www.kaggle.com/c/birdsong-recognition\">Cornell Birdcall Identification</a> dataset. It can be used to identify bird species from audio clips with high accuracy (around 55% on unseen clips) spanning 264 different species mentioned on <a href=\"https://www.xeno-canto.org/\">https://www.xeno-canto.org/</a>.</p>\n\n<hr>\n\n<h3>Usage</h3>\n\n<p>To use this pre-trained model with PyTorch, you first need to convert your audio clip to a Mel spectrogram image to feed into the model. Refer to <a href=\"https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet\">my kernel on birdcall classification</a> to understand how to generate these Mel spectrograms. Make sure your audio signal is of length <code>1000000</code> and set the spectrogram features to <code>256</code>. Then finally convert the image to a 3-channel version (repetition) and apply classic ImageNet normalization with <code>albumentations</code>. Once you convert the audio clip/s to Mel spectrograms, define the ResNet model and the load the pre-trained weights from this dataset. The code snippet below demonstrates how to set up the model with pre-trained weights. Now, you can use this model to classify bird species!</p>\n\n<hr>\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p>",
      "rawMarkdown": "Hello everyone! I have published a dataset with a trained ResNet-34 model to tackle this task with CNNs and Mel spectrograms.\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n### TLDR\n\nNotebook: https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet\nDataset: https://www.kaggle.com/tarunpaparaju/pretrained-resnet34-for-birdcall-classification\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n### Content\n\nThis dataset contains a <code>ResNet-34</code> model trained on Mel spectrograms from the [Cornell Birdcall Identification](https://www.kaggle.com/c/birdsong-recognition) dataset. It can be used to identify bird species from audio clips with high accuracy (around 55% on unseen clips) spanning 264 different species mentioned on https://www.xeno-canto.org/.\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n### Usage\n\nTo use this pre-trained model with PyTorch, you first need to convert your audio clip to a Mel spectrogram image to feed into the model. Refer to [my kernel on birdcall classification](https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet) to understand how to generate these Mel spectrograms. Make sure your audio signal is of length <code>1000000</code> and set the spectrogram features to <code>256</code>. Then finally convert the image to a 3-channel version (repetition) and apply classic ImageNet normalization with <code>albumentations</code>. Once you convert the audio clip/s to Mel spectrograms, define the ResNet model and the load the pre-trained weights from this dataset. The code snippet below demonstrates how to set up the model with pre-trained weights. Now, you can use this model to classify bird species!\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<img width=\"650px\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1351163%2Fd1aed4b170f94ec95d97fff9163330b6%2Fcarbon.png?generation=1593065716014933&amp;alt=media\">",
      "votes": null
    },
    {
      "id": "900954",
      "postDate": "06/25/2020 06:59:53",
      "content": "<p>P.S.: I forgot to add <code>torch.load</code> in the final line. Call <code>torch.load</code> before setting the state dict.</p>",
      "rawMarkdown": "P.S.: I forgot to add <code>torch.load</code> in the final line. Call <code>torch.load</code> before setting the state dict.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 900954,
      "author_name": "tarunpaparaju",
      "author_url": "",
      "post_date": "06/25/2020 06:59:53",
      "content": "<p>P.S.: I forgot to add <code>torch.load</code> in the final line. Call <code>torch.load</code> before setting the state dict.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "900922": "Hello everyone! I have published a dataset with a trained ResNet-34 model to tackle this task with CNNs and Mel spectrograms.\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n### TLDR\n\nNotebook: https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet\nDataset: https://www.kaggle.com/tarunpaparaju/pretrained-resnet34-for-birdcall-classification\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n### Content\n\nThis dataset contains a <code>ResNet-34</code> model trained on Mel spectrograms from the [Cornell Birdcall Identification](https://www.kaggle.com/c/birdsong-recognition) dataset. It can be used to identify bird species from audio clips with high accuracy (around 55% on unseen clips) spanning 264 different species mentioned on https://www.xeno-canto.org/.\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n### Usage\n\nTo use this pre-trained model with PyTorch, you first need to convert your audio clip to a Mel spectrogram image to feed into the model. Refer to [my kernel on birdcall classification](https://www.kaggle.com/tarunpaparaju/birdcall-identification-spectrogram-resnet) to understand how to generate these Mel spectrograms. Make sure your audio signal is of length <code>1000000</code> and set the spectrogram features to <code>256</code>. Then finally convert the image to a 3-channel version (repetition) and apply classic ImageNet normalization with <code>albumentations</code>. Once you convert the audio clip/s to Mel spectrograms, define the ResNet model and the load the pre-trained weights from this dataset. The code snippet below demonstrates how to set up the model with pre-trained weights. Now, you can use this model to classify bird species!\n\n________________________________________________________________________________________________________________________________________________________________________________\n\n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<img width=\"650px\" src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1351163%2Fd1aed4b170f94ec95d97fff9163330b6%2Fcarbon.png?generation=1593065716014933&amp;alt=media\">",
    "900954": "P.S.: I forgot to add <code>torch.load</code> in the final line. Call <code>torch.load</code> before setting the state dict."
  },
  "source": "meta"
}