{
  "id": 166577,
  "title": "[Novelty Starter Pack, LB 0.545] Training + submission in a single notebook",
  "url": "/competitions/birdsong-recognition/discussion/166577",
  "author_name": "Radek Osmulski",
  "post_date": "2020-07-13T11:20:41.178000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission\">Link to Starter Pack</a></p>\n\n<p>Training with submission takes ~ 1hr 5 minutes. I train a pretrained resnet34 on <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission\">precalculated and uploaded spectrograms</a> (I haven't tried but I assume this offers a significant speedup vs converting audio to spectrograms in training as Kaggle kernels come with only 2 CPU cores). In order to meet the single dataset limit of 20 GB, I limit the spectrograms to only the first 25 seconds or so of each audio file.</p>\n\n<p>There are ways you could probably improve on this result, for instance via training longer or training a bigger model.</p>\n\n<p>Truth be told, I was hoping this notebook would perform better. Unfortunately, I am a little bit lost in this competition. Using xenocanto recordings as validation set doesn't work that well - local validation scores even on as much as 20% of recordings stratified by class do not track the results on public LB too well (wondering if others are also experiencing this?)</p>\n\n<p>Using soundscape recordings shared in the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\">External Data/Pre-Trained Models Disclosure Thread</a> also doesn't work very well as validation.</p>\n\n<p>There is a set of methods that could be implemented if we had access to the test set (such as the really cool arch <a href=\"https://arxiv.org/abs/1409.7495\">here</a>, you can find my implementation <a href=\"https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb\">here</a> - didn't manage to get it to work, potentially due the soundscape recordings I used not being representative of the test data or for some other reason). </p>\n\n<p>Despite all that, I am assuming our only hope is to make the representation of data between train and test as similar as possible. Not sure exactly how to address this but will keep trying 🙂</p>",
  "messages": [
    {
      "id": 927421,
      "postDate": "2020-07-13T11:20:41.180Z",
      "content": "<p><a href=\"https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission\">Link to Starter Pack</a></p>\n\n<p>Training with submission takes ~ 1hr 5 minutes. I train a pretrained resnet34 on <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission\">precalculated and uploaded spectrograms</a> (I haven't tried but I assume this offers a significant speedup vs converting audio to spectrograms in training as Kaggle kernels come with only 2 CPU cores). In order to meet the single dataset limit of 20 GB, I limit the spectrograms to only the first 25 seconds or so of each audio file.</p>\n\n<p>There are ways you could probably improve on this result, for instance via training longer or training a bigger model.</p>\n\n<p>Truth be told, I was hoping this notebook would perform better. Unfortunately, I am a little bit lost in this competition. Using xenocanto recordings as validation set doesn't work that well - local validation scores even on as much as 20% of recordings stratified by class do not track the results on public LB too well (wondering if others are also experiencing this?)</p>\n\n<p>Using soundscape recordings shared in the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\">External Data/Pre-Trained Models Disclosure Thread</a> also doesn't work very well as validation.</p>\n\n<p>There is a set of methods that could be implemented if we had access to the test set (such as the really cool arch <a href=\"https://arxiv.org/abs/1409.7495\">here</a>, you can find my implementation <a href=\"https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb\">here</a> - didn't manage to get it to work, potentially due the soundscape recordings I used not being representative of the test data or for some other reason). </p>\n\n<p>Despite all that, I am assuming our only hope is to make the representation of data between train and test as similar as possible. Not sure exactly how to address this but will keep trying 🙂</p>",
      "rawMarkdown": "[Link to Starter Pack](https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission)\n\nTraining with submission takes ~ 1hr 5 minutes. I train a pretrained resnet34 on [precalculated and uploaded spectrograms](https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission) (I haven't tried but I assume this offers a significant speedup vs converting audio to spectrograms in training as Kaggle kernels come with only 2 CPU cores). In order to meet the single dataset limit of 20 GB, I limit the spectrograms to only the first 25 seconds or so of each audio file.\n\nThere are ways you could probably improve on this result, for instance via training longer or training a bigger model.\n\nTruth be told, I was hoping this notebook would perform better. Unfortunately, I am a little bit lost in this competition. Using xenocanto recordings as validation set doesn't work that well - local validation scores even on as much as 20% of recordings stratified by class do not track the results on public LB too well (wondering if others are also experiencing this?)\n\nUsing soundscape recordings shared in the [External Data/Pre-Trained Models Disclosure Thread](https://www.kaggle.com/c/birdsong-recognition/discussion/158877) also doesn't work very well as validation.\n\nThere is a set of methods that could be implemented if we had access to the test set (such as the really cool arch [here](https://arxiv.org/abs/1409.7495), you can find my implementation [here](https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb) - didn't manage to get it to work, potentially due the soundscape recordings I used not being representative of the test data or for some other reason). \n\nDespite all that, I am assuming our only hope is to make the representation of data between train and test as similar as possible. Not sure exactly how to address this but will keep trying 🙂",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "927421": "[Link to Starter Pack](https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission)\n\nTraining with submission takes ~ 1hr 5 minutes. I train a pretrained resnet34 on [precalculated and uploaded spectrograms](https://www.kaggle.com/radek1/esp-starter-pack-from-training-to-submission) (I haven't tried but I assume this offers a significant speedup vs converting audio to spectrograms in training as Kaggle kernels come with only 2 CPU cores). In order to meet the single dataset limit of 20 GB, I limit the spectrograms to only the first 25 seconds or so of each audio file.\n\nThere are ways you could probably improve on this result, for instance via training longer or training a bigger model.\n\nTruth be told, I was hoping this notebook would perform better. Unfortunately, I am a little bit lost in this competition. Using xenocanto recordings as validation set doesn't work that well - local validation scores even on as much as 20% of recordings stratified by class do not track the results on public LB too well (wondering if others are also experiencing this?)\n\nUsing soundscape recordings shared in the [External Data/Pre-Trained Models Disclosure Thread](https://www.kaggle.com/c/birdsong-recognition/discussion/158877) also doesn't work very well as validation.\n\nThere is a set of methods that could be implemented if we had access to the test set (such as the really cool arch [here](https://arxiv.org/abs/1409.7495), you can find my implementation [here](https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb) - didn't manage to get it to work, potentially due the soundscape recordings I used not being representative of the test data or for some other reason). \n\nDespite all that, I am assuming our only hope is to make the representation of data between train and test as similar as possible. Not sure exactly how to address this but will keep trying 🙂"
  }
}