{
  "id": 164698,
  "title": "[ESP Starter Pack v3, 0.55 LB] Res34 trained on 80x212 specs",
  "url": "/competitions/birdsong-recognition/discussion/164698",
  "author_name": "Radek Osmulski",
  "post_date": "2020-07-07T08:31:38.201000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax?scriptVersionId=38208416\">Notebook for making the submission</a>\n<a href=\"https://github.com/earthspecies/birdcall/blob/master/02jb_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_log.ipynb\">Notebook on github used for training</a>\n+ now with soundscape validation dataset!</p>\n\n<p>Finally beating the all zero submission, yay! 😊</p>\n\n<p>Obviously, making progress on the LB is nice. But the real question is - why do certain methods work better than others? Why do some models, that have better performance on the train (and soundscape) validation set do not generalize as well to the hidden test set?</p>\n\n<p>Having trained a bunch of models, including implementing <a href=\"https://github.com/earthspecies/birdcall/blob/master/02ib_train_on_melspectrograms_pytorch_lme_pool_frontend_negative_class_nocnn.ipynb\">some quite exotic ones</a>, it really helped me to write <a href=\"https://github.com/earthspecies/birdcall/blob/master/00_Overview.ipynb\">these couple of words</a> speaking to what this competition is about at a high level.</p>\n\n<p>The conclusion is that this competition is strictly a domain adaptation problem. There is the straight forward, intuitive way of addressing domain adaptation that immediately comes to mind -&gt; maybe simpler models can actually work better vs having a more complex model that can fit particularly well but to a dataset of specific characteristics? I haven't validated this claim, but it seems like an interesting proposition.</p>\n\n<p>Another component, that I feel particularly strong about given the results I saw, is example normalization. This is an important claim I feel and one that I have (at least partially) validated. Audio is a weird domain where often normalizing each example by for instance zero centering it improves results. But here I imagine this sort of normalization might be doing more for us - if the calls in the soundscape recordings are fainter, this might be bringing them up to the level of the recordings done by directional microphones. My observation would be that per example normalization might work better here than normalizing using statistics calculated on the entire train set.</p>",
  "messages": [
    {
      "id": 918403,
      "postDate": "2020-07-07T08:31:38.203Z",
      "content": "<p><a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax?scriptVersionId=38208416\">Notebook for making the submission</a>\n<a href=\"https://github.com/earthspecies/birdcall/blob/master/02jb_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_log.ipynb\">Notebook on github used for training</a>\n+ now with soundscape validation dataset!</p>\n\n<p>Finally beating the all zero submission, yay! 😊</p>\n\n<p>Obviously, making progress on the LB is nice. But the real question is - why do certain methods work better than others? Why do some models, that have better performance on the train (and soundscape) validation set do not generalize as well to the hidden test set?</p>\n\n<p>Having trained a bunch of models, including implementing <a href=\"https://github.com/earthspecies/birdcall/blob/master/02ib_train_on_melspectrograms_pytorch_lme_pool_frontend_negative_class_nocnn.ipynb\">some quite exotic ones</a>, it really helped me to write <a href=\"https://github.com/earthspecies/birdcall/blob/master/00_Overview.ipynb\">these couple of words</a> speaking to what this competition is about at a high level.</p>\n\n<p>The conclusion is that this competition is strictly a domain adaptation problem. There is the straight forward, intuitive way of addressing domain adaptation that immediately comes to mind -&gt; maybe simpler models can actually work better vs having a more complex model that can fit particularly well but to a dataset of specific characteristics? I haven't validated this claim, but it seems like an interesting proposition.</p>\n\n<p>Another component, that I feel particularly strong about given the results I saw, is example normalization. This is an important claim I feel and one that I have (at least partially) validated. Audio is a weird domain where often normalizing each example by for instance zero centering it improves results. But here I imagine this sort of normalization might be doing more for us - if the calls in the soundscape recordings are fainter, this might be bringing them up to the level of the recordings done by directional microphones. My observation would be that per example normalization might work better here than normalizing using statistics calculated on the entire train set.</p>",
      "rawMarkdown": "[Notebook for making the submission](https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax?scriptVersionId=38208416)\n[Notebook on github used for training](https://github.com/earthspecies/birdcall/blob/master/02jb_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_log.ipynb)\n+ now with soundscape validation dataset!\n\nFinally beating the all zero submission, yay! 😊\n\nObviously, making progress on the LB is nice. But the real question is - why do certain methods work better than others? Why do some models, that have better performance on the train (and soundscape) validation set do not generalize as well to the hidden test set?\n\nHaving trained a bunch of models, including implementing [some quite exotic ones](https://github.com/earthspecies/birdcall/blob/master/02ib_train_on_melspectrograms_pytorch_lme_pool_frontend_negative_class_nocnn.ipynb), it really helped me to write [these couple of words](https://github.com/earthspecies/birdcall/blob/master/00_Overview.ipynb) speaking to what this competition is about at a high level.\n\nThe conclusion is that this competition is strictly a domain adaptation problem. There is the straight forward, intuitive way of addressing domain adaptation that immediately comes to mind -&gt; maybe simpler models can actually work better vs having a more complex model that can fit particularly well but to a dataset of specific characteristics? I haven't validated this claim, but it seems like an interesting proposition.\n\nAnother component, that I feel particularly strong about given the results I saw, is example normalization. This is an important claim I feel and one that I have (at least partially) validated. Audio is a weird domain where often normalizing each example by for instance zero centering it improves results. But here I imagine this sort of normalization might be doing more for us - if the calls in the soundscape recordings are fainter, this might be bringing them up to the level of the recordings done by directional microphones. My observation would be that per example normalization might work better here than normalizing using statistics calculated on the entire train set.",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "918403": "[Notebook for making the submission](https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax?scriptVersionId=38208416)\n[Notebook on github used for training](https://github.com/earthspecies/birdcall/blob/master/02jb_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_log.ipynb)\n+ now with soundscape validation dataset!\n\nFinally beating the all zero submission, yay! 😊\n\nObviously, making progress on the LB is nice. But the real question is - why do certain methods work better than others? Why do some models, that have better performance on the train (and soundscape) validation set do not generalize as well to the hidden test set?\n\nHaving trained a bunch of models, including implementing [some quite exotic ones](https://github.com/earthspecies/birdcall/blob/master/02ib_train_on_melspectrograms_pytorch_lme_pool_frontend_negative_class_nocnn.ipynb), it really helped me to write [these couple of words](https://github.com/earthspecies/birdcall/blob/master/00_Overview.ipynb) speaking to what this competition is about at a high level.\n\nThe conclusion is that this competition is strictly a domain adaptation problem. There is the straight forward, intuitive way of addressing domain adaptation that immediately comes to mind -&gt; maybe simpler models can actually work better vs having a more complex model that can fit particularly well but to a dataset of specific characteristics? I haven't validated this claim, but it seems like an interesting proposition.\n\nAnother component, that I feel particularly strong about given the results I saw, is example normalization. This is an important claim I feel and one that I have (at least partially) validated. Audio is a weird domain where often normalizing each example by for instance zero centering it improves results. But here I imagine this sort of normalization might be doing more for us - if the calls in the soundscape recordings are fainter, this might be bringing them up to the level of the recordings done by directional microphones. My observation would be that per example normalization might work better here than normalizing using statistics calculated on the entire train set."
  }
}