{
  "id": 160222,
  "title": "[fastai, pytorch, 0.48 LB] Early Competition Starter Pack ",
  "url": "/competitions/birdsong-recognition/discussion/160222",
  "author_name": "Radek Osmulski",
  "post_date": "2020-06-20T11:04:46.025000",
  "votes": 67,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Below I will describe (and share code) on how I went about training locally, uploading a model to Kaggle and making predictions on the test set.</p>\n\n<p>This is the first iteration through the full pipeline. Whether working on a research or a software project, I am a strongly believe that quickly creating something that will work end to end and then iterating on it is a good approach. Please do not mind the rough edges - sharing this already as I feel it can be useful to other competitors.</p>\n\n<p>While working on this, I identified several points (that I speak to in the notebooks) where in retrospect I could have done things better (but I didn't realize this at the time of writing the code, only once I moved onto the downstream steps in the pipeline). Will try to address some of these issues on the second iteration of this and share the code again. Nonetheless, the model trains, code runs, prediction on test set works, the scaffolding for a pipeline from reading in the data all the way to submitting is there 😊</p>\n\n<p>Here is what I did:\n1. I trained a model using code I share in this <a href=\"https://github.com/earthspecies/birdcall\">repository</a>. This contains all the steps needed to download data, resample it, create datasets, train a model, save weights and run inference.\n2. I uploaded the weights and a list with classes to Kaggle as a dataset.\n3. I then ran <a href=\"https://www.kaggle.com/radek1/first-model?scriptVersionId=36919480\">this kernel</a> to make a submission.</p>\n\n<p>I plan to continue working on this and do something similar like I did some time ago for the <a href=\"https://github.com/radekosmulski/whale\">whale</a> competition.</p>\n\n<p>I haven't participated in Kaggle for quite a long time but was very excited when I saw this competition so hoping to be able to contribute to the community as much as I can. Models such as this can be very useful from conservation perspective (monitoring biodiversity, identifying poaching, feeding information into bioacoustic research) and also this is really closely aligned with what I do for work, so that is another win 🙂</p>\n\n<p>Anyhow - let me know what you think about the approach. Really excited for the competition and hope together we can push the envelope on what the models like the one we are working on for this competition can do. </p>\n\n<p>Already extremely grateful to the organizers for putting together such a genuinely interesting and practical challenge! 🙂</p>\n\n<p>Looking forward to working on this with you all 😎</p>\n\n<p>PS. Special thanks to @cwthompson and @shonenkov whose notebooks <a href=\"https://www.kaggle.com/cwthompson/birdsong-making-a-prediction\">here</a> and <a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check\">here </a> have been very helpful to me in working on this!\nPS2. I use an architecture designed to classify environmental sounds. Among many other things that could be tried, probably switching to a different architecture would improve performance. \nPS3. Reading in mp3 files once for each example seems quite costly. To save time I read every recording only once and then index into it as below. I think this approach can be helpful even if you perform inference in some other way (without this the kernel timed out and with this it completed in ~10 minutes):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fdafd780ce12041f27bfd827c2d3b931a%2Fiterate_over_recordings.png?generation=1592650715575394&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 894343,
      "postDate": "2020-06-20T11:04:46.027Z",
      "content": "<p>Below I will describe (and share code) on how I went about training locally, uploading a model to Kaggle and making predictions on the test set.</p>\n\n<p>This is the first iteration through the full pipeline. Whether working on a research or a software project, I am a strongly believe that quickly creating something that will work end to end and then iterating on it is a good approach. Please do not mind the rough edges - sharing this already as I feel it can be useful to other competitors.</p>\n\n<p>While working on this, I identified several points (that I speak to in the notebooks) where in retrospect I could have done things better (but I didn't realize this at the time of writing the code, only once I moved onto the downstream steps in the pipeline). Will try to address some of these issues on the second iteration of this and share the code again. Nonetheless, the model trains, code runs, prediction on test set works, the scaffolding for a pipeline from reading in the data all the way to submitting is there 😊</p>\n\n<p>Here is what I did:\n1. I trained a model using code I share in this <a href=\"https://github.com/earthspecies/birdcall\">repository</a>. This contains all the steps needed to download data, resample it, create datasets, train a model, save weights and run inference.\n2. I uploaded the weights and a list with classes to Kaggle as a dataset.\n3. I then ran <a href=\"https://www.kaggle.com/radek1/first-model?scriptVersionId=36919480\">this kernel</a> to make a submission.</p>\n\n<p>I plan to continue working on this and do something similar like I did some time ago for the <a href=\"https://github.com/radekosmulski/whale\">whale</a> competition.</p>\n\n<p>I haven't participated in Kaggle for quite a long time but was very excited when I saw this competition so hoping to be able to contribute to the community as much as I can. Models such as this can be very useful from conservation perspective (monitoring biodiversity, identifying poaching, feeding information into bioacoustic research) and also this is really closely aligned with what I do for work, so that is another win 🙂</p>\n\n<p>Anyhow - let me know what you think about the approach. Really excited for the competition and hope together we can push the envelope on what the models like the one we are working on for this competition can do. </p>\n\n<p>Already extremely grateful to the organizers for putting together such a genuinely interesting and practical challenge! 🙂</p>\n\n<p>Looking forward to working on this with you all 😎</p>\n\n<p>PS. Special thanks to @cwthompson and @shonenkov whose notebooks <a href=\"https://www.kaggle.com/cwthompson/birdsong-making-a-prediction\">here</a> and <a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check\">here </a> have been very helpful to me in working on this!\nPS2. I use an architecture designed to classify environmental sounds. Among many other things that could be tried, probably switching to a different architecture would improve performance. \nPS3. Reading in mp3 files once for each example seems quite costly. To save time I read every recording only once and then index into it as below. I think this approach can be helpful even if you perform inference in some other way (without this the kernel timed out and with this it completed in ~10 minutes):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fdafd780ce12041f27bfd827c2d3b931a%2Fiterate_over_recordings.png?generation=1592650715575394&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Below I will describe (and share code) on how I went about training locally, uploading a model to Kaggle and making predictions on the test set.\n\nThis is the first iteration through the full pipeline. Whether working on a research or a software project, I am a strongly believe that quickly creating something that will work end to end and then iterating on it is a good approach. Please do not mind the rough edges - sharing this already as I feel it can be useful to other competitors.\n\nWhile working on this, I identified several points (that I speak to in the notebooks) where in retrospect I could have done things better (but I didn't realize this at the time of writing the code, only once I moved onto the downstream steps in the pipeline). Will try to address some of these issues on the second iteration of this and share the code again. Nonetheless, the model trains, code runs, prediction on test set works, the scaffolding for a pipeline from reading in the data all the way to submitting is there 😊\n\nHere is what I did:\n1. I trained a model using code I share in this [repository](https://github.com/earthspecies/birdcall). This contains all the steps needed to download data, resample it, create datasets, train a model, save weights and run inference.\n2. I uploaded the weights and a list with classes to Kaggle as a dataset.\n3. I then ran [this kernel](https://www.kaggle.com/radek1/first-model?scriptVersionId=36919480) to make a submission.\n\nI plan to continue working on this and do something similar like I did some time ago for the [whale](https://github.com/radekosmulski/whale) competition.\n\nI haven't participated in Kaggle for quite a long time but was very excited when I saw this competition so hoping to be able to contribute to the community as much as I can. Models such as this can be very useful from conservation perspective (monitoring biodiversity, identifying poaching, feeding information into bioacoustic research) and also this is really closely aligned with what I do for work, so that is another win 🙂\n\nAnyhow - let me know what you think about the approach. Really excited for the competition and hope together we can push the envelope on what the models like the one we are working on for this competition can do. \n\nAlready extremely grateful to the organizers for putting together such a genuinely interesting and practical challenge! 🙂\n\nLooking forward to working on this with you all 😎\n\n\nPS. Special thanks to @cwthompson and @shonenkov whose notebooks [here](https://www.kaggle.com/cwthompson/birdsong-making-a-prediction) and [here ](https://www.kaggle.com/shonenkov/sample-submission-using-custom-check) have been very helpful to me in working on this!\nPS2. I use an architecture designed to classify environmental sounds. Among many other things that could be tried, probably switching to a different architecture would improve performance. \nPS3. Reading in mp3 files once for each example seems quite costly. To save time I read every recording only once and then index into it as below. I think this approach can be helpful even if you perform inference in some other way (without this the kernel timed out and with this it completed in ~10 minutes):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fdafd780ce12041f27bfd827c2d3b931a%2Fiterate_over_recordings.png?generation=1592650715575394&amp;alt=media)\n ",
      "votes": 67
    },
    {
      "id": 895234,
      "postDate": "2020-06-21T07:34:51.123Z",
      "content": "<p>Great work <a href=\"/radek1\">@radek1</a> - your model architecture seems very reasonable, we've seen similar topologies succeed for larger datasets.</p>\n\n<p>Here are some thoughts on how to proceed/what to implement:</p>\n\n<ul>\n<li><p>not all segments of each training recording are equally important, some contain no vocalization and might distort your training data</p></li>\n<li><p>augmentation is key; we've seen a range of methods in the past, frequency shifts and noise overlays stand out, but I'm sure there's more</p></li>\n<li><p>we've seen deeper models perform better on soundscapes and wider models perform better on Xeno-canto recordings; ResNets seem to be a good choice, DenseNet is also said to generalize well for this task</p></li>\n<li><p>location, location, location (metadata can help to verify detections)</p></li>\n<li><p>soundscapes are tough and we've been struggling for years; thanks for helping us out :)</p></li>\n</ul>\n\n<p>I'd be happy to join a more in-depth discussion if you like.</p>",
      "rawMarkdown": "Great work @radek1 - your model architecture seems very reasonable, we've seen similar topologies succeed for larger datasets.\n\nHere are some thoughts on how to proceed/what to implement:\n\n- not all segments of each training recording are equally important, some contain no vocalization and might distort your training data\n\n- augmentation is key; we've seen a range of methods in the past, frequency shifts and noise overlays stand out, but I'm sure there's more\n\n- we've seen deeper models perform better on soundscapes and wider models perform better on Xeno-canto recordings; ResNets seem to be a good choice, DenseNet is also said to generalize well for this task\n\n- location, location, location (metadata can help to verify detections)\n\n- soundscapes are tough and we've been struggling for years; thanks for helping us out :)\n\nI'd be happy to join a more in-depth discussion if you like.",
      "votes": 19,
      "replies": [
        {
          "id": 895594,
          "postDate": "2020-06-21T13:35:24.553Z",
          "content": "<p>Thank you very much <a href=\"/stefankahl\">@stefankahl</a> - your guidance is invaluable! 😊🙏 Thank you so much for your comments and for taking the time to take a look at what I put together.</p>\n\n<p>&gt; not all segments of each training recording are equally important, some contain no vocalization and might distort your training data</p>\n\n<p>Agreed! 🙂 In the current approach I am hoping I can to some extent rely on deep learning models being somewhat robust to noisy labels. But I also fear the training data can be too distorted here, as you mention.</p>\n\n<p>If I will have the bandwidth to work on this a little longer, I would love to hand label some calls in Raven and train an object detection model and run it on the train set to generate better examples. This is a little bit more tricky and I think it should be okay even if it is done for say 10 best represented birds in the training set, as an experiment. Should results be promising I think this is something we could build off.  I'm thinking that having isolated calls could also lead to better synthetic examples, where we could maybe combine the calls with different background sounds, etc. I am not really sure where to get such backgrounds, if they are available publicly at all? 🤔 Anything coming from PAM sensors might be good, maybe I need to search a bit 😊 We have some data being shipped to us at ESP, but I am not sure how soon it will arrive and also much of it is from the ocean, which might not be exactly what we are looking for here 😄 Anyhow, all this is just conceptualizing at this point, not super certain if I will be able to get around to doing this but hoping bouncing ideas can still be useful.</p>\n\n<p>I am also thinking of incorporating Google AudioSet into the training data as you mention <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158937#888201\">here</a>, that should be relatively approachable.</p>\n\n<p>&gt; augmentation is key; we've seen a range of methods in the past, frequency shifts and noise overlays stand out, but I'm sure there's more</p>\n\n<p>Didn't realize that, thank you! Will definitely look into this one - augmentations in the context of sound seem to me to be a bit of an uncharted territory, would love to learn / experiment some more with them</p>\n\n<p>&gt; we've seen deeper models perform better on soundscapes and wider models perform better on Xeno-canto recordings; ResNets seem to be a good choice, DenseNet is also said to generalize well for this task</p>\n\n<p>Cool! Really good to know! 😊 I am also thinking of exploring a little bit more the way that spectrograms are created, I might be completely wrong on this one but my hunch is that a lot really might depend on that initial transformation, on how on goes from audio -&gt; spectrograms. My colleague from work is working on this <a href=\"https://github.com/earthspecies/representation-toolbox\">representation toolbox</a>, would be interesting to try some of them out. But my money would be on modified spectrograms, sort of how they did for <a href=\"https://github.com/DrCoffey/DeepSqueak\">deepsqueak</a> or <a href=\"https://github.com/mvansegbroeck/mupet\">mupet</a>. I completely have no idea if this makes sense when there is a lot of ambient sound though. There is also the family of linearly reassigned spectrogams, but I am not sure if these are not cost prohibitive in terms of generation for larger datasets.</p>\n\n<p>&gt; location, location, location (metadata can help to verify detections)</p>\n\n<p>Again, thank you very much! 🙂 Something I probably would have missed, would have not given enough consideration, if it wasn't for your comment 🙏</p>\n\n<p>&gt; soundscapes are tough and we've been struggling for years; thanks for helping us out :)</p>\n\n<p>my pleasure 😊🙏</p>\n\n<p>&gt; I'd be happy to join a more in-depth discussion if you like.</p>\n\n<p>I would love that! Everything I say above is merely tentative, I feel a lot of this is uncharted territory for me, any suggestions and comments would be greatly appreciated. Also, my apologies for how long this post turned out but had all these thoughts and wanted to share</p>\n\n<p>EDIT: Just came across the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158933#888206\">post</a> where you mention two papers, one on augmentation, awesome! Will be checking them out 😊</p>",
          "rawMarkdown": "Thank you very much @stefankahl - your guidance is invaluable! 😊🙏 Thank you so much for your comments and for taking the time to take a look at what I put together.\n\n&gt; not all segments of each training recording are equally important, some contain no vocalization and might distort your training data\n\nAgreed! 🙂 In the current approach I am hoping I can to some extent rely on deep learning models being somewhat robust to noisy labels. But I also fear the training data can be too distorted here, as you mention.\n\nIf I will have the bandwidth to work on this a little longer, I would love to hand label some calls in Raven and train an object detection model and run it on the train set to generate better examples. This is a little bit more tricky and I think it should be okay even if it is done for say 10 best represented birds in the training set, as an experiment. Should results be promising I think this is something we could build off.  I'm thinking that having isolated calls could also lead to better synthetic examples, where we could maybe combine the calls with different background sounds, etc. I am not really sure where to get such backgrounds, if they are available publicly at all? 🤔 Anything coming from PAM sensors might be good, maybe I need to search a bit 😊 We have some data being shipped to us at ESP, but I am not sure how soon it will arrive and also much of it is from the ocean, which might not be exactly what we are looking for here 😄 Anyhow, all this is just conceptualizing at this point, not super certain if I will be able to get around to doing this but hoping bouncing ideas can still be useful.\n\nI am also thinking of incorporating Google AudioSet into the training data as you mention [here](https://www.kaggle.com/c/birdsong-recognition/discussion/158937#888201), that should be relatively approachable.\n\n&gt; augmentation is key; we've seen a range of methods in the past, frequency shifts and noise overlays stand out, but I'm sure there's more\n\nDidn't realize that, thank you! Will definitely look into this one - augmentations in the context of sound seem to me to be a bit of an uncharted territory, would love to learn / experiment some more with them\n\n&gt; we've seen deeper models perform better on soundscapes and wider models perform better on Xeno-canto recordings; ResNets seem to be a good choice, DenseNet is also said to generalize well for this task\n\nCool! Really good to know! 😊 I am also thinking of exploring a little bit more the way that spectrograms are created, I might be completely wrong on this one but my hunch is that a lot really might depend on that initial transformation, on how on goes from audio -&gt; spectrograms. My colleague from work is working on this [representation toolbox](https://github.com/earthspecies/representation-toolbox), would be interesting to try some of them out. But my money would be on modified spectrograms, sort of how they did for [deepsqueak](https://github.com/DrCoffey/DeepSqueak) or [mupet](https://github.com/mvansegbroeck/mupet). I completely have no idea if this makes sense when there is a lot of ambient sound though. There is also the family of linearly reassigned spectrogams, but I am not sure if these are not cost prohibitive in terms of generation for larger datasets.\n\n&gt; location, location, location (metadata can help to verify detections)\n\nAgain, thank you very much! 🙂 Something I probably would have missed, would have not given enough consideration, if it wasn't for your comment 🙏\n\n&gt; soundscapes are tough and we've been struggling for years; thanks for helping us out :)\n\nmy pleasure 😊🙏\n\n&gt; I'd be happy to join a more in-depth discussion if you like.\n\nI would love that! Everything I say above is merely tentative, I feel a lot of this is uncharted territory for me, any suggestions and comments would be greatly appreciated. Also, my apologies for how long this post turned out but had all these thoughts and wanted to share\n\nEDIT: Just came across the [post](https://www.kaggle.com/c/birdsong-recognition/discussion/158933#888206) where you mention two papers, one on augmentation, awesome! Will be checking them out 😊",
          "votes": 2
        },
        {
          "id": 896464,
          "postDate": "2020-06-22T07:43:53.700Z",
          "content": "<blockquote>\n  <p>I would love to hand label some calls in Raven and train an object detection model and run it on the train set to generate better examples</p>\n</blockquote>\n\n<p>Maybe Weakly-supervised Sound Event Detection will help us to detect sound event out of a long clip with only clip level labels. In DCASE challenges, competitions of Weakly supervised SED have been hosted for last few years, so we can gather knowledge from these.\n<a href=\"http://dcase.community/challenge2019/task-sound-event-detection-in-domestic-environments\">http://dcase.community/challenge2019/task-sound-event-detection-in-domestic-environments</a>\n<a href=\"http://dcase.community/challenge2020/task-sound-event-detection-and-separation-in-domestic-environments\">http://dcase.community/challenge2020/task-sound-event-detection-and-separation-in-domestic-environments</a></p>",
          "rawMarkdown": "&gt; I would love to hand label some calls in Raven and train an object detection model and run it on the train set to generate better examples\n\nMaybe Weakly-supervised Sound Event Detection will help us to detect sound event out of a long clip with only clip level labels. In DCASE challenges, competitions of Weakly supervised SED have been hosted for last few years, so we can gather knowledge from these.\nhttp://dcase.community/challenge2019/task-sound-event-detection-in-domestic-environments\nhttp://dcase.community/challenge2020/task-sound-event-detection-and-separation-in-domestic-environments",
          "votes": 7
        },
        {
          "id": 896507,
          "postDate": "2020-06-22T08:23:09.727Z",
          "content": "<p>Thank you very much for sharing this <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a>, this is very relevant! Sounds like a very interesting and promising angle to explore!</p>",
          "rawMarkdown": "Thank you very much for sharing this @hidehisaarai1213, this is very relevant! Sounds like a very interesting and promising angle to explore!"
        },
        {
          "id": 896619,
          "postDate": "2020-06-22T10:13:59.087Z",
          "content": "<blockquote>\n  <ul>\n  <li>location, location, location (metadata can help to verify detections)</li>\n  </ul>\n</blockquote>\n\n<p>I can only find metadata such as location and altitude in the training data. It seems to me that this data is missing for the test files. How would you include this data in your model if it is unkown during testing?</p>",
          "rawMarkdown": "&gt; - location, location, location (metadata can help to verify detections)\n\n\nI can only find metadata such as location and altitude in the training data. It seems to me that this data is missing for the test files. How would you include this data in your model if it is unkown during testing?"
        },
        {
          "id": 896636,
          "postDate": "2020-06-22T10:28:35.903Z",
          "content": "<p>Perhaps what <a href=\"/stefankahl\">@stefankahl</a> would like to say is not about using location data in test phase, but about <em>location invariant</em> modeling.</p>",
          "rawMarkdown": "Perhaps what @stefankahl would like to say is not about using location data in test phase, but about *location invariant* modeling."
        },
        {
          "id": 896778,
          "postDate": "2020-06-22T12:29:55.600Z",
          "content": "<p>Well, yes, both actually. But primarily location data for training - which was more intended as a general comment on how to approach things. The test data doesn't have location information, except that it was recorded in North America. But I would suggest looking at distribution and range of some birds which might indicate how common these species are. Some of them are migrating species which might be worth investigating.</p>\n\n<p>But having location independent features is extremely important, too. Typically, a CNN will give you that. Not just in the time but also in the frequency domain (birds are known to adjust their vocal output based on environmental factors which mostly translates to shifts in pitch).</p>",
          "rawMarkdown": "Well, yes, both actually. But primarily location data for training - which was more intended as a general comment on how to approach things. The test data doesn't have location information, except that it was recorded in North America. But I would suggest looking at distribution and range of some birds which might indicate how common these species are. Some of them are migrating species which might be worth investigating.\n\nBut having location independent features is extremely important, too. Typically, a CNN will give you that. Not just in the time but also in the frequency domain (birds are known to adjust their vocal output based on environmental factors which mostly translates to shifts in pitch).",
          "votes": 4
        }
      ]
    },
    {
      "id": 964795,
      "postDate": "2020-08-10T07:11:47.963Z",
      "content": "<p>Thank you very much for sharing this. Your work is very helpful to me.</p>",
      "rawMarkdown": "Thank you very much for sharing this. Your work is very helpful to me.",
      "votes": 1
    },
    {
      "id": 921911,
      "postDate": "2020-07-09T16:43:00.273Z",
      "content": "<p>Great Work - Nice model architecture</p>",
      "rawMarkdown": "Great Work - Nice model architecture",
      "votes": 1
    },
    {
      "id": 918548,
      "postDate": "2020-07-07T10:03:49.713Z",
      "content": "<p>Radek, thank you very much for uploading and documenting this! I am a beginner and will attempt to go through it step by step in order to learn more.</p>",
      "rawMarkdown": "Radek, thank you very much for uploading and documenting this! I am a beginner and will attempt to go through it step by step in order to learn more.",
      "votes": 1
    },
    {
      "id": 981149,
      "postDate": "2020-08-22T07:56:17.067Z",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I am just getting up to speed on fastai.  In your code you do like so:</p>\n<pre><code>learn = Learner(\n    dls,\n    get_arch(),\n    metrics=[AccumMetric(precision), AccumMetric(recall), AccumMetric(f1)],\n    loss_func=BCELossFlat(),\n    splitter=custom_splitter\n)\n</code></pre>\n<pre><code>learn.fit(5, 5e-2)\n</code></pre>\n<p>I thought, from reading the fastai docs, I could just do something like:</p>\n<pre><code>learn.to_parallel()\n</code></pre>\n<p>to take advantage of multiple GPU's.  It seemed like it's the equivalent of PyTorch's:</p>\n<pre><code>model = nn.DataParallel(model)\n</code></pre>\n<p>But I got the error <code>AttributeError: 'Learner' object has no attribute 'to_parallel'</code>.</p>\n<p>I am using fastai-2.0.0 from pip install. Do you know how I could make your learner run multi-GPU? (just multi-thread not multi process, ex: DataParallel not DistributedDataParallel)?</p>",
      "rawMarkdown": "@radek1 I am just getting up to speed on fastai.  In your code you do like so:\n\n```\nlearn = Learner(\n    dls,\n    get_arch(),\n    metrics=[AccumMetric(precision), AccumMetric(recall), AccumMetric(f1)],\n    loss_func=BCELossFlat(),\n    splitter=custom_splitter\n)\n```\n\n```\nlearn.fit(5, 5e-2)\n\n```\nI thought, from reading the fastai docs, I could just do something like:\n\n```\nlearn.to_parallel()\n\n```\nto take advantage of multiple GPU's.  It seemed like it's the equivalent of PyTorch's:\n\n```\nmodel = nn.DataParallel(model)\n\n```\nBut I got the error `AttributeError: 'Learner' object has no attribute 'to_parallel'`.\n\nI am using fastai-2.0.0 from pip install. Do you know how I could make your learner run multi-GPU? (just multi-thread not multi process, ex: DataParallel not DistributedDataParallel)?\n"
    },
    {
      "id": 981011,
      "postDate": "2020-08-22T05:24:55.213Z",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I have a question. In your code, you reference <code>BirdCLEF2020_val/gt</code> and <code>BirdCLEF2020_val/audio</code> directories and files.  Where do those come from? Do they come from <a href=\"https://www.kaggle.com/stecasasso/birdclef2020-soundscape-validation\" target=\"_blank\">this Kaggle Dataset</a>?  What does \"gt\" stand for?</p>",
      "rawMarkdown": "@radek1 I have a question. In your code, you reference `BirdCLEF2020_val/gt` and `BirdCLEF2020_val/audio` directories and files.  Where do those come from? Do they come from [this Kaggle Dataset](https://www.kaggle.com/stecasasso/birdclef2020-soundscape-validation)?  What does \"gt\" stand for?",
      "replies": [
        {
          "id": 981023,
          "postDate": "2020-08-22T05:43:26.357Z",
          "content": "<p>I see they come from this discussion <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/158877</a></p>",
          "rawMarkdown": "I see they come from this discussion https://www.kaggle.com/c/birdsong-recognition/discussion/158877"
        }
      ]
    },
    {
      "id": 895290,
      "postDate": "2020-06-21T08:27:52.177Z",
      "content": "<p>nice work buddy </p>",
      "rawMarkdown": "nice work buddy "
    },
    {
      "id": 957313,
      "postDate": "2020-08-04T08:03:18.323Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 957321,
          "postDate": "2020-08-04T08:11:11.343Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 895234,
      "author_name": "Stefan Kahl",
      "author_url": "",
      "post_date": "2020-06-21T07:34:51.123000",
      "content": "<p>Great work <a href=\"/radek1\">@radek1</a> - your model architecture seems very reasonable, we've seen similar topologies succeed for larger datasets.</p>\n\n<p>Here are some thoughts on how to proceed/what to implement:</p>\n\n<ul>\n<li><p>not all segments of each training recording are equally important, some contain no vocalization and might distort your training data</p></li>\n<li><p>augmentation is key; we've seen a range of methods in the past, frequency shifts and noise overlays stand out, but I'm sure there's more</p></li>\n<li><p>we've seen deeper models perform better on soundscapes and wider models perform better on Xeno-canto recordings; ResNets seem to be a good choice, DenseNet is also said to generalize well for this task</p></li>\n<li><p>location, location, location (metadata can help to verify detections)</p></li>\n<li><p>soundscapes are tough and we've been struggling for years; thanks for helping us out :)</p></li>\n</ul>\n\n<p>I'd be happy to join a more in-depth discussion if you like.</p>",
      "votes": 19,
      "replies": [
        {
          "id": 895594,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2020-06-21T13:35:24.553000",
          "content": "<p>Thank you very much <a href=\"/stefankahl\">@stefankahl</a> - your guidance is invaluable! 😊🙏 Thank you so much for your comments and for taking the time to take a look at what I put together.</p>\n\n<p>&gt; not all segments of each training recording are equally important, some contain no vocalization and might distort your training data</p>\n\n<p>Agreed! 🙂 In the current approach I am hoping I can to some extent rely on deep learning models being somewhat robust to noisy labels. But I also fear the training data can be too distorted here, as you mention.</p>\n\n<p>If I will have the bandwidth to work on this a little longer, I would love to hand label some calls in Raven and train an object detection model and run it on the train set to generate better examples. This is a little bit more tricky and I think it should be okay even if it is done for say 10 best represented birds in the training set, as an experiment. Should results be promising I think this is something we could build off.  I'm thinking that having isolated calls could also lead to better synthetic examples, where we could maybe combine the calls with different background sounds, etc. I am not really sure where to get such backgrounds, if they are available publicly at all? 🤔 Anything coming from PAM sensors might be good, maybe I need to search a bit 😊 We have some data being shipped to us at ESP, but I am not sure how soon it will arrive and also much of it is from the ocean, which might not be exactly what we are looking for here 😄 Anyhow, all this is just conceptualizing at this point, not super certain if I will be able to get around to doing this but hoping bouncing ideas can still be useful.</p>\n\n<p>I am also thinking of incorporating Google AudioSet into the training data as you mention <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158937#888201\">here</a>, that should be relatively approachable.</p>\n\n<p>&gt; augmentation is key; we've seen a range of methods in the past, frequency shifts and noise overlays stand out, but I'm sure there's more</p>\n\n<p>Didn't realize that, thank you! Will definitely look into this one - augmentations in the context of sound seem to me to be a bit of an uncharted territory, would love to learn / experiment some more with them</p>\n\n<p>&gt; we've seen deeper models perform better on soundscapes and wider models perform better on Xeno-canto recordings; ResNets seem to be a good choice, DenseNet is also said to generalize well for this task</p>\n\n<p>Cool! Really good to know! 😊 I am also thinking of exploring a little bit more the way that spectrograms are created, I might be completely wrong on this one but my hunch is that a lot really might depend on that initial transformation, on how on goes from audio -&gt; spectrograms. My colleague from work is working on this <a href=\"https://github.com/earthspecies/representation-toolbox\">representation toolbox</a>, would be interesting to try some of them out. But my money would be on modified spectrograms, sort of how they did for <a href=\"https://github.com/DrCoffey/DeepSqueak\">deepsqueak</a> or <a href=\"https://github.com/mvansegbroeck/mupet\">mupet</a>. I completely have no idea if this makes sense when there is a lot of ambient sound though. There is also the family of linearly reassigned spectrogams, but I am not sure if these are not cost prohibitive in terms of generation for larger datasets.</p>\n\n<p>&gt; location, location, location (metadata can help to verify detections)</p>\n\n<p>Again, thank you very much! 🙂 Something I probably would have missed, would have not given enough consideration, if it wasn't for your comment 🙏</p>\n\n<p>&gt; soundscapes are tough and we've been struggling for years; thanks for helping us out :)</p>\n\n<p>my pleasure 😊🙏</p>\n\n<p>&gt; I'd be happy to join a more in-depth discussion if you like.</p>\n\n<p>I would love that! Everything I say above is merely tentative, I feel a lot of this is uncharted territory for me, any suggestions and comments would be greatly appreciated. Also, my apologies for how long this post turned out but had all these thoughts and wanted to share</p>\n\n<p>EDIT: Just came across the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158933#888206\">post</a> where you mention two papers, one on augmentation, awesome! Will be checking them out 😊</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 896464,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-22T07:43:53.700000",
          "content": "<blockquote>\n  <p>I would love to hand label some calls in Raven and train an object detection model and run it on the train set to generate better examples</p>\n</blockquote>\n\n<p>Maybe Weakly-supervised Sound Event Detection will help us to detect sound event out of a long clip with only clip level labels. In DCASE challenges, competitions of Weakly supervised SED have been hosted for last few years, so we can gather knowledge from these.\n<a href=\"http://dcase.community/challenge2019/task-sound-event-detection-in-domestic-environments\">http://dcase.community/challenge2019/task-sound-event-detection-in-domestic-environments</a>\n<a href=\"http://dcase.community/challenge2020/task-sound-event-detection-and-separation-in-domestic-environments\">http://dcase.community/challenge2020/task-sound-event-detection-and-separation-in-domestic-environments</a></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 896507,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2020-06-22T08:23:09.727000",
          "content": "<p>Thank you very much for sharing this <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a>, this is very relevant! Sounds like a very interesting and promising angle to explore!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896619,
          "author_name": "KoertS",
          "author_url": "",
          "post_date": "2020-06-22T10:13:59.087000",
          "content": "<blockquote>\n  <ul>\n  <li>location, location, location (metadata can help to verify detections)</li>\n  </ul>\n</blockquote>\n\n<p>I can only find metadata such as location and altitude in the training data. It seems to me that this data is missing for the test files. How would you include this data in your model if it is unkown during testing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896636,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-22T10:28:35.903000",
          "content": "<p>Perhaps what <a href=\"/stefankahl\">@stefankahl</a> would like to say is not about using location data in test phase, but about <em>location invariant</em> modeling.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896778,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2020-06-22T12:29:55.600000",
          "content": "<p>Well, yes, both actually. But primarily location data for training - which was more intended as a general comment on how to approach things. The test data doesn't have location information, except that it was recorded in North America. But I would suggest looking at distribution and range of some birds which might indicate how common these species are. Some of them are migrating species which might be worth investigating.</p>\n\n<p>But having location independent features is extremely important, too. Typically, a CNN will give you that. Not just in the time but also in the frequency domain (birds are known to adjust their vocal output based on environmental factors which mostly translates to shifts in pitch).</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 964795,
      "author_name": "lansaid",
      "author_url": "",
      "post_date": "2020-08-10T07:11:47.963000",
      "content": "<p>Thank you very much for sharing this. Your work is very helpful to me.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 921911,
      "author_name": "VaibhavMishra",
      "author_url": "",
      "post_date": "2020-07-09T16:43:00.273000",
      "content": "<p>Great Work - Nice model architecture</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 918548,
      "author_name": "Alex A.",
      "author_url": "",
      "post_date": "2020-07-07T10:03:49.713000",
      "content": "<p>Radek, thank you very much for uploading and documenting this! I am a beginner and will attempt to go through it step by step in order to learn more.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 981149,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-08-22T07:56:17.067000",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I am just getting up to speed on fastai.  In your code you do like so:</p>\n<pre><code>learn = Learner(\n    dls,\n    get_arch(),\n    metrics=[AccumMetric(precision), AccumMetric(recall), AccumMetric(f1)],\n    loss_func=BCELossFlat(),\n    splitter=custom_splitter\n)\n</code></pre>\n<pre><code>learn.fit(5, 5e-2)\n</code></pre>\n<p>I thought, from reading the fastai docs, I could just do something like:</p>\n<pre><code>learn.to_parallel()\n</code></pre>\n<p>to take advantage of multiple GPU's.  It seemed like it's the equivalent of PyTorch's:</p>\n<pre><code>model = nn.DataParallel(model)\n</code></pre>\n<p>But I got the error <code>AttributeError: 'Learner' object has no attribute 'to_parallel'</code>.</p>\n<p>I am using fastai-2.0.0 from pip install. Do you know how I could make your learner run multi-GPU? (just multi-thread not multi process, ex: DataParallel not DistributedDataParallel)?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 981011,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-08-22T05:24:55.213000",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I have a question. In your code, you reference <code>BirdCLEF2020_val/gt</code> and <code>BirdCLEF2020_val/audio</code> directories and files.  Where do those come from? Do they come from <a href=\"https://www.kaggle.com/stecasasso/birdclef2020-soundscape-validation\" target=\"_blank\">this Kaggle Dataset</a>?  What does \"gt\" stand for?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 981023,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-22T05:43:26.357000",
          "content": "<p>I see they come from this discussion <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/158877</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 895290,
      "author_name": "Satya Muralidhar",
      "author_url": "",
      "post_date": "2020-06-21T08:27:52.177000",
      "content": "<p>nice work buddy </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 957313,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-04T08:03:18.323000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 957321,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-04T08:11:11.343000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "894343": "Below I will describe (and share code) on how I went about training locally, uploading a model to Kaggle and making predictions on the test set.\n\nThis is the first iteration through the full pipeline. Whether working on a research or a software project, I am a strongly believe that quickly creating something that will work end to end and then iterating on it is a good approach. Please do not mind the rough edges - sharing this already as I feel it can be useful to other competitors.\n\nWhile working on this, I identified several points (that I speak to in the notebooks) where in retrospect I could have done things better (but I didn't realize this at the time of writing the code, only once I moved onto the downstream steps in the pipeline). Will try to address some of these issues on the second iteration of this and share the code again. Nonetheless, the model trains, code runs, prediction on test set works, the scaffolding for a pipeline from reading in the data all the way to submitting is there 😊\n\nHere is what I did:\n1. I trained a model using code I share in this [repository](https://github.com/earthspecies/birdcall). This contains all the steps needed to download data, resample it, create datasets, train a model, save weights and run inference.\n2. I uploaded the weights and a list with classes to Kaggle as a dataset.\n3. I then ran [this kernel](https://www.kaggle.com/radek1/first-model?scriptVersionId=36919480) to make a submission.\n\nI plan to continue working on this and do something similar like I did some time ago for the [whale](https://github.com/radekosmulski/whale) competition.\n\nI haven't participated in Kaggle for quite a long time but was very excited when I saw this competition so hoping to be able to contribute to the community as much as I can. Models such as this can be very useful from conservation perspective (monitoring biodiversity, identifying poaching, feeding information into bioacoustic research) and also this is really closely aligned with what I do for work, so that is another win 🙂\n\nAnyhow - let me know what you think about the approach. Really excited for the competition and hope together we can push the envelope on what the models like the one we are working on for this competition can do. \n\nAlready extremely grateful to the organizers for putting together such a genuinely interesting and practical challenge! 🙂\n\nLooking forward to working on this with you all 😎\n\n\nPS. Special thanks to @cwthompson and @shonenkov whose notebooks [here](https://www.kaggle.com/cwthompson/birdsong-making-a-prediction) and [here ](https://www.kaggle.com/shonenkov/sample-submission-using-custom-check) have been very helpful to me in working on this!\nPS2. I use an architecture designed to classify environmental sounds. Among many other things that could be tried, probably switching to a different architecture would improve performance. \nPS3. Reading in mp3 files once for each example seems quite costly. To save time I read every recording only once and then index into it as below. I think this approach can be helpful even if you perform inference in some other way (without this the kernel timed out and with this it completed in ~10 minutes):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fdafd780ce12041f27bfd827c2d3b931a%2Fiterate_over_recordings.png?generation=1592650715575394&amp;alt=media)\n ",
    "895234": "Great work @radek1 - your model architecture seems very reasonable, we've seen similar topologies succeed for larger datasets.\n\nHere are some thoughts on how to proceed/what to implement:\n\n- not all segments of each training recording are equally important, some contain no vocalization and might distort your training data\n\n- augmentation is key; we've seen a range of methods in the past, frequency shifts and noise overlays stand out, but I'm sure there's more\n\n- we've seen deeper models perform better on soundscapes and wider models perform better on Xeno-canto recordings; ResNets seem to be a good choice, DenseNet is also said to generalize well for this task\n\n- location, location, location (metadata can help to verify detections)\n\n- soundscapes are tough and we've been struggling for years; thanks for helping us out :)\n\nI'd be happy to join a more in-depth discussion if you like.",
    "964795": "Thank you very much for sharing this. Your work is very helpful to me.",
    "921911": "Great Work - Nice model architecture",
    "918548": "Radek, thank you very much for uploading and documenting this! I am a beginner and will attempt to go through it step by step in order to learn more.",
    "981149": "@radek1 I am just getting up to speed on fastai.  In your code you do like so:\n\n```\nlearn = Learner(\n    dls,\n    get_arch(),\n    metrics=[AccumMetric(precision), AccumMetric(recall), AccumMetric(f1)],\n    loss_func=BCELossFlat(),\n    splitter=custom_splitter\n)\n```\n\n```\nlearn.fit(5, 5e-2)\n\n```\nI thought, from reading the fastai docs, I could just do something like:\n\n```\nlearn.to_parallel()\n\n```\nto take advantage of multiple GPU's.  It seemed like it's the equivalent of PyTorch's:\n\n```\nmodel = nn.DataParallel(model)\n\n```\nBut I got the error `AttributeError: 'Learner' object has no attribute 'to_parallel'`.\n\nI am using fastai-2.0.0 from pip install. Do you know how I could make your learner run multi-GPU? (just multi-thread not multi process, ex: DataParallel not DistributedDataParallel)?\n",
    "981011": "@radek1 I have a question. In your code, you reference `BirdCLEF2020_val/gt` and `BirdCLEF2020_val/audio` directories and files.  Where do those come from? Do they come from [this Kaggle Dataset](https://www.kaggle.com/stecasasso/birdclef2020-soundscape-validation)?  What does \"gt\" stand for?",
    "895290": "nice work buddy ",
    "957313": ""
  }
}