{
  "id": 161860,
  "title": "[ESP Starter Pack v2, 0.48 LB] MIL with resnet34",
  "url": "/competitions/birdsong-recognition/discussion/161860",
  "author_name": "",
  "post_date": "2020-06-26T12:54:20.891013500Z",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>A new version of <strong>ESP Starter Pack</strong> is here! 🥳</p>\n\n<p><a href=\"https://github.com/earthspecies/birdcall\">GitHub repository with code for training the model</a>\n<a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v2?scriptVersionId=37478978\">Kaggle kernel for predicting on the test set and submitting results</a></p>\n\n<p>In every 5 sec excerpt of a train set recording, the chance that the call of a bird particular to this example will appear is not that great. It is impossible for me to quantify this  to what extent this is the case (though I did listen to a couple of the recordings, which was very useful to me in working on this).</p>\n\n<p>If we were to assign a label to a randomly chosen 5 second sample of a recording, this would lead to a large portion of the examples not containing the vocalization we are after. I initially tried training on such data using standard computer vision classification techniques but to no avail - the model was not learning! (I did try quite a few things, including training with cross entropy and softmax to make the task a little bit easier for our model).</p>\n\n<p>This problem, where we have a bag of examples representing a class, but any of the individual examples can be better associated with another class, is the gist of Multi Instance Learning (MIL).  In the case of the v2 of the starter pack, I opted to go with using a pooling layer to address this issue</p>\n\n<p>Instead of sampling a 5 second continuous excerpt for a recording, for each example I sample 30 x 1.66 second randomly chosen excerpts. I then proceed to generate spectrograms (one for each excerpt) and I create an example of following dimensionality:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F66e3656859a3514f014bb9e3cb649527%2Fexample_dim.png?generation=1593167826454269&amp;alt=media\" alt=\"\"></p>\n\n<p>Instead of the recording being represented by just 5 seconds of audio, our model is able to consider 50 seconds.</p>\n\n<p>I place one spectrogram in each channel of the 10 resultant 'images' and I feed them - 10 images in total per example - to pretrained res34. I run global average pooling on each of the dimensions of the embedding (descriptor obtained from the cnn body) and train a simple fully connected classifier on top of it.</p>\n\n<p>Here is the entire architecture:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F6186e5265be88f1f5bb75bd6be16051a%2Farch.png?generation=1593170593872469&amp;alt=media\" alt=\"\"></p>\n\n<p>This approach is based on this very informative <a href=\"http://ceur-ws.org/Vol-2125/paper_181.pdf\">paper</a> by Jan Schlüter. The paper is a treasure trove of information, is really well written and there is also an accompanying <a href=\"https://github.com/f0k/birdclef2018\">repository</a>. Kudos to Jan Schlüter for his stellar performance in Birdclef2018 and thanks so much for publishing the paper and making the code available!</p>\n\n<p>Kaggle competitions come in various shapes and forms and based on what I can tell so far, this is proving to be a really great competition, certainly would rank among my favorites, from the once I participated in. One thing that could be extremely helpful though is a validation set, ideally sampled from the same distribution as the test set. I realize that this might be impossible to procure at this point (I also understand the reasoning why it was not provided in the first place - to prevent people from training on it), but I am wondering what data sourced externally would work well as a validation set here? I feel having some answer to this question could be really helpful to the efforts.</p>\n\n<p>Anyhow - learning a lot from the competition and the participants, thanks a lot! 🙏The model does not perform that well on the train set - it achieves an LB score of 0.48. I do however feel good about the training pipeline and about the approach, hoping to improve the result and / or possibly also try other approaches.</p>\n\n<p>More to come! 🙂</p>",
  "messages": [
    {
      "id": "902910",
      "postDate": "06/26/2020 12:54:20",
      "content": "<p>A new version of <strong>ESP Starter Pack</strong> is here! 🥳</p>\n\n<p><a href=\"https://github.com/earthspecies/birdcall\">GitHub repository with code for training the model</a>\n<a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v2?scriptVersionId=37478978\">Kaggle kernel for predicting on the test set and submitting results</a></p>\n\n<p>In every 5 sec excerpt of a train set recording, the chance that the call of a bird particular to this example will appear is not that great. It is impossible for me to quantify this  to what extent this is the case (though I did listen to a couple of the recordings, which was very useful to me in working on this).</p>\n\n<p>If we were to assign a label to a randomly chosen 5 second sample of a recording, this would lead to a large portion of the examples not containing the vocalization we are after. I initially tried training on such data using standard computer vision classification techniques but to no avail - the model was not learning! (I did try quite a few things, including training with cross entropy and softmax to make the task a little bit easier for our model).</p>\n\n<p>This problem, where we have a bag of examples representing a class, but any of the individual examples can be better associated with another class, is the gist of Multi Instance Learning (MIL).  In the case of the v2 of the starter pack, I opted to go with using a pooling layer to address this issue</p>\n\n<p>Instead of sampling a 5 second continuous excerpt for a recording, for each example I sample 30 x 1.66 second randomly chosen excerpts. I then proceed to generate spectrograms (one for each excerpt) and I create an example of following dimensionality:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F66e3656859a3514f014bb9e3cb649527%2Fexample_dim.png?generation=1593167826454269&amp;alt=media\" alt=\"\"></p>\n\n<p>Instead of the recording being represented by just 5 seconds of audio, our model is able to consider 50 seconds.</p>\n\n<p>I place one spectrogram in each channel of the 10 resultant 'images' and I feed them - 10 images in total per example - to pretrained res34. I run global average pooling on each of the dimensions of the embedding (descriptor obtained from the cnn body) and train a simple fully connected classifier on top of it.</p>\n\n<p>Here is the entire architecture:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F6186e5265be88f1f5bb75bd6be16051a%2Farch.png?generation=1593170593872469&amp;alt=media\" alt=\"\"></p>\n\n<p>This approach is based on this very informative <a href=\"http://ceur-ws.org/Vol-2125/paper_181.pdf\">paper</a> by Jan Schlüter. The paper is a treasure trove of information, is really well written and there is also an accompanying <a href=\"https://github.com/f0k/birdclef2018\">repository</a>. Kudos to Jan Schlüter for his stellar performance in Birdclef2018 and thanks so much for publishing the paper and making the code available!</p>\n\n<p>Kaggle competitions come in various shapes and forms and based on what I can tell so far, this is proving to be a really great competition, certainly would rank among my favorites, from the once I participated in. One thing that could be extremely helpful though is a validation set, ideally sampled from the same distribution as the test set. I realize that this might be impossible to procure at this point (I also understand the reasoning why it was not provided in the first place - to prevent people from training on it), but I am wondering what data sourced externally would work well as a validation set here? I feel having some answer to this question could be really helpful to the efforts.</p>\n\n<p>Anyhow - learning a lot from the competition and the participants, thanks a lot! 🙏The model does not perform that well on the train set - it achieves an LB score of 0.48. I do however feel good about the training pipeline and about the approach, hoping to improve the result and / or possibly also try other approaches.</p>\n\n<p>More to come! 🙂</p>",
      "rawMarkdown": "A new version of **ESP Starter Pack** is here! 🥳\n\n[GitHub repository with code for training the model](https://github.com/earthspecies/birdcall)\n[Kaggle kernel for predicting on the test set and submitting results](https://www.kaggle.com/radek1/esp-starter-pack-v2?scriptVersionId=37478978)\n\nIn every 5 sec excerpt of a train set recording, the chance that the call of a bird particular to this example will appear is not that great. It is impossible for me to quantify this  to what extent this is the case (though I did listen to a couple of the recordings, which was very useful to me in working on this).\n\nIf we were to assign a label to a randomly chosen 5 second sample of a recording, this would lead to a large portion of the examples not containing the vocalization we are after. I initially tried training on such data using standard computer vision classification techniques but to no avail - the model was not learning! (I did try quite a few things, including training with cross entropy and softmax to make the task a little bit easier for our model).\n\nThis problem, where we have a bag of examples representing a class, but any of the individual examples can be better associated with another class, is the gist of Multi Instance Learning (MIL).  In the case of the v2 of the starter pack, I opted to go with using a pooling layer to address this issue\n\nInstead of sampling a 5 second continuous excerpt for a recording, for each example I sample 30 x 1.66 second randomly chosen excerpts. I then proceed to generate spectrograms (one for each excerpt) and I create an example of following dimensionality:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F66e3656859a3514f014bb9e3cb649527%2Fexample_dim.png?generation=1593167826454269&amp;alt=media)\n\nInstead of the recording being represented by just 5 seconds of audio, our model is able to consider 50 seconds.\n\nI place one spectrogram in each channel of the 10 resultant 'images' and I feed them - 10 images in total per example - to pretrained res34. I run global average pooling on each of the dimensions of the embedding (descriptor obtained from the cnn body) and train a simple fully connected classifier on top of it.\n\nHere is the entire architecture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F6186e5265be88f1f5bb75bd6be16051a%2Farch.png?generation=1593170593872469&amp;alt=media)\n\nThis approach is based on this very informative [paper](http://ceur-ws.org/Vol-2125/paper_181.pdf) by Jan Schlüter. The paper is a treasure trove of information, is really well written and there is also an accompanying [repository](https://github.com/f0k/birdclef2018). Kudos to Jan Schlüter for his stellar performance in Birdclef2018 and thanks so much for publishing the paper and making the code available!\n\nKaggle competitions come in various shapes and forms and based on what I can tell so far, this is proving to be a really great competition, certainly would rank among my favorites, from the once I participated in. One thing that could be extremely helpful though is a validation set, ideally sampled from the same distribution as the test set. I realize that this might be impossible to procure at this point (I also understand the reasoning why it was not provided in the first place - to prevent people from training on it), but I am wondering what data sourced externally would work well as a validation set here? I feel having some answer to this question could be really helpful to the efforts.\n\nAnyhow - learning a lot from the competition and the participants, thanks a lot! 🙏The model does not perform that well on the train set - it achieves an LB score of 0.48. I do however feel good about the training pipeline and about the approach, hoping to improve the result and / or possibly also try other approaches.\n\nMore to come! 🙂",
      "votes": null
    },
    {
      "id": "971101",
      "postDate": "08/15/2020 07:20:22",
      "content": "<p>Thanks!</p>\n<p>A few notes to self:</p>\n<ul>\n<li>the output of <code>self.cnn(...)</code> will be of shape <code>(batch_size, 512, height, width)</code></li>\n<li>then, after <code>x.mean</code>, it'll be of shape <code>(batch_size, 512)</code></li>\n<li>then, after <code>self.classifier</code>, <code>(batch_size, num_classes)</code></li>\n<li>then, we take a view of shape <code>(bs, im_num, num_classes)</code></li>\n<li>finally, it'll be of shape <code>(bs, num_classes)</code></li>\n</ul>",
      "rawMarkdown": "Thanks!\n\nA few notes to self:\n- the output of `self.cnn(...)` will be of shape `(batch_size, 512, height, width)`\n- then, after `x.mean`, it'll be of shape `(batch_size, 512)`\n- then, after `self.classifier`, `(batch_size, num_classes)`\n- then, we take a view of shape `(bs, im_num, num_classes)`\n- finally, it'll be of shape `(bs, num_classes)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 971101,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "08/15/2020 07:20:22",
      "content": "<p>Thanks!</p>\n<p>A few notes to self:</p>\n<ul>\n<li>the output of <code>self.cnn(...)</code> will be of shape <code>(batch_size, 512, height, width)</code></li>\n<li>then, after <code>x.mean</code>, it'll be of shape <code>(batch_size, 512)</code></li>\n<li>then, after <code>self.classifier</code>, <code>(batch_size, num_classes)</code></li>\n<li>then, we take a view of shape <code>(bs, im_num, num_classes)</code></li>\n<li>finally, it'll be of shape <code>(bs, num_classes)</code></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "902910": "A new version of **ESP Starter Pack** is here! 🥳\n\n[GitHub repository with code for training the model](https://github.com/earthspecies/birdcall)\n[Kaggle kernel for predicting on the test set and submitting results](https://www.kaggle.com/radek1/esp-starter-pack-v2?scriptVersionId=37478978)\n\nIn every 5 sec excerpt of a train set recording, the chance that the call of a bird particular to this example will appear is not that great. It is impossible for me to quantify this  to what extent this is the case (though I did listen to a couple of the recordings, which was very useful to me in working on this).\n\nIf we were to assign a label to a randomly chosen 5 second sample of a recording, this would lead to a large portion of the examples not containing the vocalization we are after. I initially tried training on such data using standard computer vision classification techniques but to no avail - the model was not learning! (I did try quite a few things, including training with cross entropy and softmax to make the task a little bit easier for our model).\n\nThis problem, where we have a bag of examples representing a class, but any of the individual examples can be better associated with another class, is the gist of Multi Instance Learning (MIL).  In the case of the v2 of the starter pack, I opted to go with using a pooling layer to address this issue\n\nInstead of sampling a 5 second continuous excerpt for a recording, for each example I sample 30 x 1.66 second randomly chosen excerpts. I then proceed to generate spectrograms (one for each excerpt) and I create an example of following dimensionality:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F66e3656859a3514f014bb9e3cb649527%2Fexample_dim.png?generation=1593167826454269&amp;alt=media)\n\nInstead of the recording being represented by just 5 seconds of audio, our model is able to consider 50 seconds.\n\nI place one spectrogram in each channel of the 10 resultant 'images' and I feed them - 10 images in total per example - to pretrained res34. I run global average pooling on each of the dimensions of the embedding (descriptor obtained from the cnn body) and train a simple fully connected classifier on top of it.\n\nHere is the entire architecture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F6186e5265be88f1f5bb75bd6be16051a%2Farch.png?generation=1593170593872469&amp;alt=media)\n\nThis approach is based on this very informative [paper](http://ceur-ws.org/Vol-2125/paper_181.pdf) by Jan Schlüter. The paper is a treasure trove of information, is really well written and there is also an accompanying [repository](https://github.com/f0k/birdclef2018). Kudos to Jan Schlüter for his stellar performance in Birdclef2018 and thanks so much for publishing the paper and making the code available!\n\nKaggle competitions come in various shapes and forms and based on what I can tell so far, this is proving to be a really great competition, certainly would rank among my favorites, from the once I participated in. One thing that could be extremely helpful though is a validation set, ideally sampled from the same distribution as the test set. I realize that this might be impossible to procure at this point (I also understand the reasoning why it was not provided in the first place - to prevent people from training on it), but I am wondering what data sourced externally would work well as a validation set here? I feel having some answer to this question could be really helpful to the efforts.\n\nAnyhow - learning a lot from the competition and the participants, thanks a lot! 🙏The model does not perform that well on the train set - it achieves an LB score of 0.48. I do however feel good about the training pipeline and about the approach, hoping to improve the result and / or possibly also try other approaches.\n\nMore to come! 🙂",
    "971101": "Thanks!\n\nA few notes to self:\n- the output of `self.cnn(...)` will be of shape `(batch_size, 512, height, width)`\n- then, after `x.mean`, it'll be of shape `(batch_size, 512)`\n- then, after `self.classifier`, `(batch_size, num_classes)`\n- then, we take a view of shape `(bs, im_num, num_classes)`\n- finally, it'll be of shape `(bs, num_classes)`"
  },
  "source": "meta"
}