{
  "id": 220305,
  "title": "Wave App for Soundscape Annotation & Best Single Model (0.956)",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220305",
  "author_name": "",
  "post_date": "2021-02-18T00:08:03.974204300Z",
  "votes": 61,
  "comment_count": 16,
  "views": 0,
  "content": "<p><img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/static/icon.png?raw=true\" alt=\"logo\"><br>\nWe share all our details in a<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\" target=\"_blank\"> separate thread</a> here I just wanted to quickly summarize our key weapon to collect more training labels. </p>\n<h2>Motivation</h2>\n<p>Just listening a few recordings there was a lot of activity, usually multiple species vocalizing at the same time.<br>\nYet we were not provided a single positive label for the vast majority of the training recordings and we only had roughly one known label for about a thousand recordings.   <br>\nIn the Cornell Birdcall Identification competition selecting random audio chunks worked surprisingly well since the original training files mostly contained a single dominant bird with frequent calls.<br>\nThe labels were clearly not enough for proper local validation and I don't like to rely on LB feedback, especially when we have only 400 rows for the public leaderboard.</p>\n<p>I started the competition by reviewing the provided true and negative samples for each class.<br>\nSome of them had a very strong and clear pattern almost like the digits in the MNIST dataset.<br>\nEven simple CNNs should be able to learn that yet most of the pack struggled to get above 0.9 on the leaderboard.<br>\nIMHO having only 50 positive samples is way too low, especially if you have a multiclass multi-label problem.</p>\n<p>I was looking for a new Wave app idea anyways so I decided to create an app for soundscape annotation. I was not sure how much advantage would it give so I wanted to share it only after the end of the competition. This way I can share all the collected labels and inference code with our pretrained weights.   </p>\n<h2>Wave</h2>\n<p><a href=\"https://www.h2o.ai/products/h2o-wave/\" target=\"_blank\">H2O Wave</a> is a new open-source Python development framework that makes it fast and easy develop real-time interactive AI apps with sophisticated visualizations.</p>\n<p>It has all the features I like (python, plotly, matplotlib, markdown, simple UI components) with little overhead. Tbh there are other tools out there that would probably do the job too (flask, shiny, dash, etc.) and I used all of them before I have more experience with Wave as I work at H2O.ai. I have also built a Wave app based on the Cornell competition earlier so I could reuse a few components.</p>\n<h2>Training setup</h2>\n<p>We experimented (thanks to Peter) with quite a few different training setups.  <br>\nEventually what worked the best was probably the most straightforward, multiclass multi-label training with BCELoss. It was much faster than using 24 binary classifiers and - what is much more important at kaggle - it had better performance. We used fixed 3 second chunks Mel spectrograms for 0-14 kHz to cover all species.</p>\n<p>One disadvantage was that it required much more annotation.<br>\nBCELoss would penalize the most frequent probably unannotated classes and those classes matter a lot in LWLRAP.   </p>\n<p>We started an iterative semi-supervised approach collecting new positive labels from the most confident OOF predictions then retrained our models to collect more samples… </p>\n<p>For the annotations creating bounding boxes could be interesting but it would be quite slow painful. Since we only had to submit clipwise predictions it would be probably overkill too.<br>\nIt was easier to review each class separately and make quick binary decisions whether add or delete the new candidates.<br>\nIn the beginning I just used hacky scripts to save a bunch of candidate figures in separate directories then deleted the noise manually.<br>\nLater I used the app to go through chunks again by checking the false labels with high predictions and true labels with low predictions. This way I also had logs of true/false/not sure examples for later experiments.  </p>\n<p>In our final collection we had 23K positive labels for 11K chunks collected from 2913 training recordings.<br>\nJust training a few CNN14s on this dataset would reach 0.95 with a few hours training time.</p>\n<h2>The Hitchhiker's Guide to the Rainforest</h2>\n<p>For some of the classes the patterns were not that clear (e.g. 9, 17) we also had issues with 3 and 12 for quite a while. These  frogs like to  croak together a lot and they have quite similar patterns.</p>\n<p>Last Sunday we saw that others found the actual species published (<a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637)\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954120300637)</a>.<br>\nI checked quickly our predictions and tried to fix a few thousand labels but our score did not really improve. Finding this article in advance would have saved a lot of efforts though…</p>\n<h2>The App</h2>\n<p><a href=\"https://github.com/gaborfodor/wave-soundscape-annotation\" target=\"_blank\">https://github.com/gaborfodor/wave-soundscape-annotation</a><br>\nSimple binary annotation decisions took 1-3s after we learned all the species. (The gif is not realtime :)<br>\n<img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/annotate.gif?raw=true\" alt=\"annotate\"></p>\n<p>Beside fixing training labels the app can also classify new recordings.<br>\n<img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/recognize.png?raw=true\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1207596",
      "postDate": "02/18/2021 00:08:03",
      "content": "<p><img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/static/icon.png?raw=true\" alt=\"logo\"><br>\nWe share all our details in a<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\" target=\"_blank\"> separate thread</a> here I just wanted to quickly summarize our key weapon to collect more training labels. </p>\n<h2>Motivation</h2>\n<p>Just listening a few recordings there was a lot of activity, usually multiple species vocalizing at the same time.<br>\nYet we were not provided a single positive label for the vast majority of the training recordings and we only had roughly one known label for about a thousand recordings.   <br>\nIn the Cornell Birdcall Identification competition selecting random audio chunks worked surprisingly well since the original training files mostly contained a single dominant bird with frequent calls.<br>\nThe labels were clearly not enough for proper local validation and I don't like to rely on LB feedback, especially when we have only 400 rows for the public leaderboard.</p>\n<p>I started the competition by reviewing the provided true and negative samples for each class.<br>\nSome of them had a very strong and clear pattern almost like the digits in the MNIST dataset.<br>\nEven simple CNNs should be able to learn that yet most of the pack struggled to get above 0.9 on the leaderboard.<br>\nIMHO having only 50 positive samples is way too low, especially if you have a multiclass multi-label problem.</p>\n<p>I was looking for a new Wave app idea anyways so I decided to create an app for soundscape annotation. I was not sure how much advantage would it give so I wanted to share it only after the end of the competition. This way I can share all the collected labels and inference code with our pretrained weights.   </p>\n<h2>Wave</h2>\n<p><a href=\"https://www.h2o.ai/products/h2o-wave/\" target=\"_blank\">H2O Wave</a> is a new open-source Python development framework that makes it fast and easy develop real-time interactive AI apps with sophisticated visualizations.</p>\n<p>It has all the features I like (python, plotly, matplotlib, markdown, simple UI components) with little overhead. Tbh there are other tools out there that would probably do the job too (flask, shiny, dash, etc.) and I used all of them before I have more experience with Wave as I work at H2O.ai. I have also built a Wave app based on the Cornell competition earlier so I could reuse a few components.</p>\n<h2>Training setup</h2>\n<p>We experimented (thanks to Peter) with quite a few different training setups.  <br>\nEventually what worked the best was probably the most straightforward, multiclass multi-label training with BCELoss. It was much faster than using 24 binary classifiers and - what is much more important at kaggle - it had better performance. We used fixed 3 second chunks Mel spectrograms for 0-14 kHz to cover all species.</p>\n<p>One disadvantage was that it required much more annotation.<br>\nBCELoss would penalize the most frequent probably unannotated classes and those classes matter a lot in LWLRAP.   </p>\n<p>We started an iterative semi-supervised approach collecting new positive labels from the most confident OOF predictions then retrained our models to collect more samples… </p>\n<p>For the annotations creating bounding boxes could be interesting but it would be quite slow painful. Since we only had to submit clipwise predictions it would be probably overkill too.<br>\nIt was easier to review each class separately and make quick binary decisions whether add or delete the new candidates.<br>\nIn the beginning I just used hacky scripts to save a bunch of candidate figures in separate directories then deleted the noise manually.<br>\nLater I used the app to go through chunks again by checking the false labels with high predictions and true labels with low predictions. This way I also had logs of true/false/not sure examples for later experiments.  </p>\n<p>In our final collection we had 23K positive labels for 11K chunks collected from 2913 training recordings.<br>\nJust training a few CNN14s on this dataset would reach 0.95 with a few hours training time.</p>\n<h2>The Hitchhiker's Guide to the Rainforest</h2>\n<p>For some of the classes the patterns were not that clear (e.g. 9, 17) we also had issues with 3 and 12 for quite a while. These  frogs like to  croak together a lot and they have quite similar patterns.</p>\n<p>Last Sunday we saw that others found the actual species published (<a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637)\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954120300637)</a>.<br>\nI checked quickly our predictions and tried to fix a few thousand labels but our score did not really improve. Finding this article in advance would have saved a lot of efforts though…</p>\n<h2>The App</h2>\n<p><a href=\"https://github.com/gaborfodor/wave-soundscape-annotation\" target=\"_blank\">https://github.com/gaborfodor/wave-soundscape-annotation</a><br>\nSimple binary annotation decisions took 1-3s after we learned all the species. (The gif is not realtime :)<br>\n<img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/annotate.gif?raw=true\" alt=\"annotate\"></p>\n<p>Beside fixing training labels the app can also classify new recordings.<br>\n<img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/recognize.png?raw=true\" alt=\"\"></p>",
      "rawMarkdown": "![logo](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/static/icon.png?raw=true)\nWe share all our details in a[ separate thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443) here I just wanted to quickly summarize our key weapon to collect more training labels. \n## Motivation\nJust listening a few recordings there was a lot of activity, usually multiple species vocalizing at the same time.\nYet we were not provided a single positive label for the vast majority of the training recordings and we only had roughly one known label for about a thousand recordings.   \nIn the Cornell Birdcall Identification competition selecting random audio chunks worked surprisingly well since the original training files mostly contained a single dominant bird with frequent calls.\nThe labels were clearly not enough for proper local validation and I don't like to rely on LB feedback, especially when we have only 400 rows for the public leaderboard.\n\nI started the competition by reviewing the provided true and negative samples for each class.\nSome of them had a very strong and clear pattern almost like the digits in the MNIST dataset.\nEven simple CNNs should be able to learn that yet most of the pack struggled to get above 0.9 on the leaderboard.\nIMHO having only 50 positive samples is way too low, especially if you have a multiclass multi-label problem.\n\nI was looking for a new Wave app idea anyways so I decided to create an app for soundscape annotation. I was not sure how much advantage would it give so I wanted to share it only after the end of the competition. This way I can share all the collected labels and inference code with our pretrained weights.   \n\n## Wave\n\n[H2O Wave](https://www.h2o.ai/products/h2o-wave/) is a new open-source Python development framework that makes it fast and easy develop real-time interactive AI apps with sophisticated visualizations.\n \nIt has all the features I like (python, plotly, matplotlib, markdown, simple UI components) with little overhead. Tbh there are other tools out there that would probably do the job too (flask, shiny, dash, etc.) and I used all of them before I have more experience with Wave as I work at H2O.ai. I have also built a Wave app based on the Cornell competition earlier so I could reuse a few components.\n\n## Training setup\n \nWe experimented (thanks to Peter) with quite a few different training setups.  \nEventually what worked the best was probably the most straightforward, multiclass multi-label training with BCELoss. It was much faster than using 24 binary classifiers and - what is much more important at kaggle - it had better performance. We used fixed 3 second chunks Mel spectrograms for 0-14 kHz to cover all species.\n\nOne disadvantage was that it required much more annotation.\nBCELoss would penalize the most frequent probably unannotated classes and those classes matter a lot in LWLRAP.   \n\nWe started an iterative semi-supervised approach collecting new positive labels from the most confident OOF predictions then retrained our models to collect more samples... \n\nFor the annotations creating bounding boxes could be interesting but it would be quite slow painful. Since we only had to submit clipwise predictions it would be probably overkill too.\nIt was easier to review each class separately and make quick binary decisions whether add or delete the new candidates.\nIn the beginning I just used hacky scripts to save a bunch of candidate figures in separate directories then deleted the noise manually.\nLater I used the app to go through chunks again by checking the false labels with high predictions and true labels with low predictions. This way I also had logs of true/false/not sure examples for later experiments.  \n\nIn our final collection we had 23K positive labels for 11K chunks collected from 2913 training recordings.\nJust training a few CNN14s on this dataset would reach 0.95 with a few hours training time.\n  \n## The Hitchhiker's Guide to the Rainforest\nFor some of the classes the patterns were not that clear (e.g. 9, 17) we also had issues with 3 and 12 for quite a while. These ~~birds~~ frogs like to ~~sing~~ croak together a lot and they have quite similar patterns.\n\nLast Sunday we saw that others found the actual species published (https://www.sciencedirect.com/science/article/pii/S1574954120300637).\nI checked quickly our predictions and tried to fix a few thousand labels but our score did not really improve. Finding this article in advance would have saved a lot of efforts though...\n\n\n## The App\nhttps://github.com/gaborfodor/wave-soundscape-annotation\nSimple binary annotation decisions took 1-3s after we learned all the species. (The gif is not realtime :)\n![annotate](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/annotate.gif?raw=true)\n\nBeside fixing training labels the app can also classify new recordings.\n![](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/recognize.png?raw=true)",
      "votes": null
    },
    {
      "id": "1207656",
      "postDate": "02/18/2021 00:44:29",
      "content": "<p><a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> Congrats for your 2nd Grandmaster title!</p>",
      "rawMarkdown": "gaborfodor Congrats for your 2nd Grandmaster title!",
      "votes": null
    },
    {
      "id": "1207671",
      "postDate": "02/18/2021 00:53:54",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> ! </p>",
      "rawMarkdown": "Congratulations @gaborfodor !",
      "votes": null
    },
    {
      "id": "1207723",
      "postDate": "02/18/2021 01:53:04",
      "content": "<p>good work!</p>\n<p>kaggle should give a prize for works like this (e.g. annotation tools, visualization tools, etc)</p>",
      "rawMarkdown": "good work!\n\nkaggle should give a prize for works like this (e.g. annotation tools, visualization tools, etc)",
      "votes": null
    },
    {
      "id": "1208283",
      "postDate": "02/18/2021 08:27:14",
      "content": "<p>Thanks. Initially I wanted to help the host to collect more training labels but I had to find out later they already had a lot of labels just did not share them with us :)</p>",
      "rawMarkdown": "Thanks. Initially I wanted to help the host to collect more training labels but I had to find out later they already had a lot of labels just did not share them with us :)",
      "votes": null
    },
    {
      "id": "1208448",
      "postDate": "02/18/2021 09:45:04",
      "content": "<p>Great work. Good collections. Thank you <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> for introducing H2O Wave. 🙏</p>",
      "rawMarkdown": "Great work. Good collections. Thank you @gaborfodor for introducing H2O Wave. 🙏",
      "votes": null
    },
    {
      "id": "1208466",
      "postDate": "02/18/2021 09:48:20",
      "content": "<p>Please also check this cool app by <a href=\"https://www.kaggle.com/fatihozturk\" target=\"_blank\">@fatihozturk</a>: <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/215120\" target=\"_blank\">https://www.kaggle.com/c/stanford-covid-vaccine/discussion/215120</a></p>",
      "rawMarkdown": "Please also check this cool app by @fatihozturk: https://www.kaggle.com/c/stanford-covid-vaccine/discussion/215120",
      "votes": null
    },
    {
      "id": "1208469",
      "postDate": "02/18/2021 09:48:41",
      "content": "<p>Congrats Gabor! I love the Wave app!</p>",
      "rawMarkdown": "Congrats Gabor! I love the Wave app!",
      "votes": null
    },
    {
      "id": "1208572",
      "postDate": "02/18/2021 10:45:36",
      "content": "<p>We checked our models predictions for different species and noticed that even our latest models trained with tons of pseudo labels and augmentation got quite high activations for new unseen species. </p>\n<p>For example here are the predictions for <a href=\"https://ebird.org/species/herthr\" target=\"_blank\">Hermit Trush</a><br>\n<img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/herthr.png?raw=true\" alt=\"Hermit Trush\"></p>\n<p>We tried to use external Xeno-Canto data to force more generalized models but they did not help on the LB not even in our final ensemble…</p>",
      "rawMarkdown": "We checked our models predictions for different species and noticed that even our latest models trained with tons of pseudo labels and augmentation got quite high activations for new unseen species. \n\nFor example here are the predictions for [Hermit Trush](https://ebird.org/species/herthr)\n![Hermit Trush](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/herthr.png?raw=true)\n\nWe tried to use external Xeno-Canto data to force more generalized models but they did not help on the LB not even in our final ensemble...",
      "votes": null
    },
    {
      "id": "1208589",
      "postDate": "02/18/2021 10:52:13",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> and <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> and thanks for the writeup</p>",
      "rawMarkdown": "Congrats @gaborfodor and @pestipeti and thanks for the writeup",
      "votes": null
    },
    {
      "id": "1208616",
      "postDate": "02/18/2021 11:22:03",
      "content": "<p>I can finally see this amazing app <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> !!!!!</p>\n<p>Congratulations, this is awesome.</p>",
      "rawMarkdown": "I can finally see this amazing app @gaborfodor !!!!!\n\nCongratulations, this is awesome.",
      "votes": null
    },
    {
      "id": "1208632",
      "postDate": "02/18/2021 11:39:54",
      "content": "<p>Awesome tool <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a>. We had something similar (though not as beautiful as your app) for checking labels and also for adding labels to underrepresented species in the train set. <br>\nH2O Wave seems to be very well suited for the task. Thanks a lot for sharing.</p>\n<p>We did see dimishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP)</p>",
      "rawMarkdown": "Awesome tool @gaborfodor. We had something similar (though not as beautiful as your app) for checking labels and also for adding labels to underrepresented species in the train set. \nH2O Wave seems to be very well suited for the task. Thanks a lot for sharing.\n\nWe did see dimishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP)",
      "votes": null
    },
    {
      "id": "1208827",
      "postDate": "02/18/2021 13:57:24",
      "content": "<p>I think the test set annotation was also done with ARBIMON Pattern Matching plaform. Relying on automated high NCC template matching probably prefers the more obvious examples. I also noticed some possible misclassification in the provided FP samples.</p>\n<p>Our models were able to learn <strong>our</strong> labels with 0.99+ score but that did not translate to the LB (at least without proper postprocessing). </p>",
      "rawMarkdown": "I think the test set annotation was also done with ARBIMON Pattern Matching plaform. Relying on automated high NCC template matching probably prefers the more obvious examples. I also noticed some possible misclassification in the provided FP samples.\n\nOur models were able to learn **our** labels with 0.99+ score but that did not translate to the LB (at least without proper postprocessing).",
      "votes": null
    },
    {
      "id": "1211579",
      "postDate": "02/20/2021 10:37:57",
      "content": "<p>amazing work!!!👍</p>",
      "rawMarkdown": "amazing work!!!👍",
      "votes": null
    },
    {
      "id": "1212771",
      "postDate": "02/21/2021 15:18:16",
      "content": "<p>Thanks for open-sourcing this. A nice gesture! I hope this can be used by folks on the field…Hope it has come to the attention of the hosts..</p>\n<p>You are my heroes in this comp!</p>",
      "rawMarkdown": "Thanks for open-sourcing this. A nice gesture! I hope this can be used by folks on the field...Hope it has come to the attention of the hosts..\n\nYou are my heroes in this comp!",
      "votes": null
    },
    {
      "id": "1213037",
      "postDate": "02/21/2021 19:36:23",
      "content": "<p>Thanks. :)<br>\nI am also curious about the hosts view. There are lot of unique solutions in the top 10. Probably the winning model documentation process is still in progress. I hope we will get a summary from the hosts soon.</p>",
      "rawMarkdown": "Thanks. :)\nI am also curious about the hosts view. There are lot of unique solutions in the top 10. Probably the winning model documentation process is still in progress. I hope we will get a summary from the hosts soon.",
      "votes": null
    },
    {
      "id": "1216685",
      "postDate": "02/24/2021 12:29:01",
      "content": "<p>Great app.. Thank you for sharing the knowledge by open sourcing the app.</p>",
      "rawMarkdown": "Great app.. Thank you for sharing the knowledge by open sourcing the app.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207656,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "02/18/2021 00:44:29",
      "content": "<p><a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> Congrats for your 2nd Grandmaster title!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207671,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "02/18/2021 00:53:54",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207723,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/18/2021 01:53:04",
      "content": "<p>good work!</p>\n<p>kaggle should give a prize for works like this (e.g. annotation tools, visualization tools, etc)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208283,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "02/18/2021 08:27:14",
          "content": "<p>Thanks. Initially I wanted to help the host to collect more training labels but I had to find out later they already had a lot of labels just did not share them with us :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208448,
      "author_name": "rajkumarl",
      "author_url": "",
      "post_date": "02/18/2021 09:45:04",
      "content": "<p>Great work. Good collections. Thank you <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> for introducing H2O Wave. 🙏</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208466,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 09:48:20",
          "content": "<p>Please also check this cool app by <a href=\"https://www.kaggle.com/fatihozturk\" target=\"_blank\">@fatihozturk</a>: <a href=\"https://www.kaggle.com/c/stanford-covid-vaccine/discussion/215120\" target=\"_blank\">https://www.kaggle.com/c/stanford-covid-vaccine/discussion/215120</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208469,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "02/18/2021 09:48:41",
      "content": "<p>Congrats Gabor! I love the Wave app!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208572,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "02/18/2021 10:45:36",
      "content": "<p>We checked our models predictions for different species and noticed that even our latest models trained with tons of pseudo labels and augmentation got quite high activations for new unseen species. </p>\n<p>For example here are the predictions for <a href=\"https://ebird.org/species/herthr\" target=\"_blank\">Hermit Trush</a><br>\n<img src=\"https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/herthr.png?raw=true\" alt=\"Hermit Trush\"></p>\n<p>We tried to use external Xeno-Canto data to force more generalized models but they did not help on the LB not even in our final ensemble…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212771,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/21/2021 15:18:16",
          "content": "<p>Thanks for open-sourcing this. A nice gesture! I hope this can be used by folks on the field…Hope it has come to the attention of the hosts..</p>\n<p>You are my heroes in this comp!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213037,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "02/21/2021 19:36:23",
          "content": "<p>Thanks. :)<br>\nI am also curious about the hosts view. There are lot of unique solutions in the top 10. Probably the winning model documentation process is still in progress. I hope we will get a summary from the hosts soon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208589,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 10:52:13",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> and <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> and thanks for the writeup</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208616,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "02/18/2021 11:22:03",
      "content": "<p>I can finally see this amazing app <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> !!!!!</p>\n<p>Congratulations, this is awesome.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208632,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "02/18/2021 11:39:54",
      "content": "<p>Awesome tool <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a>. We had something similar (though not as beautiful as your app) for checking labels and also for adding labels to underrepresented species in the train set. <br>\nH2O Wave seems to be very well suited for the task. Thanks a lot for sharing.</p>\n<p>We did see dimishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208827,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "02/18/2021 13:57:24",
          "content": "<p>I think the test set annotation was also done with ARBIMON Pattern Matching plaform. Relying on automated high NCC template matching probably prefers the more obvious examples. I also noticed some possible misclassification in the provided FP samples.</p>\n<p>Our models were able to learn <strong>our</strong> labels with 0.99+ score but that did not translate to the LB (at least without proper postprocessing). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1211579,
      "author_name": "momot66",
      "author_url": "",
      "post_date": "02/20/2021 10:37:57",
      "content": "<p>amazing work!!!👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1216685,
      "author_name": "riyadhar",
      "author_url": "",
      "post_date": "02/24/2021 12:29:01",
      "content": "<p>Great app.. Thank you for sharing the knowledge by open sourcing the app.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207596": "![logo](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/static/icon.png?raw=true)\nWe share all our details in a[ separate thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443) here I just wanted to quickly summarize our key weapon to collect more training labels. \n## Motivation\nJust listening a few recordings there was a lot of activity, usually multiple species vocalizing at the same time.\nYet we were not provided a single positive label for the vast majority of the training recordings and we only had roughly one known label for about a thousand recordings.   \nIn the Cornell Birdcall Identification competition selecting random audio chunks worked surprisingly well since the original training files mostly contained a single dominant bird with frequent calls.\nThe labels were clearly not enough for proper local validation and I don't like to rely on LB feedback, especially when we have only 400 rows for the public leaderboard.\n\nI started the competition by reviewing the provided true and negative samples for each class.\nSome of them had a very strong and clear pattern almost like the digits in the MNIST dataset.\nEven simple CNNs should be able to learn that yet most of the pack struggled to get above 0.9 on the leaderboard.\nIMHO having only 50 positive samples is way too low, especially if you have a multiclass multi-label problem.\n\nI was looking for a new Wave app idea anyways so I decided to create an app for soundscape annotation. I was not sure how much advantage would it give so I wanted to share it only after the end of the competition. This way I can share all the collected labels and inference code with our pretrained weights.   \n\n## Wave\n\n[H2O Wave](https://www.h2o.ai/products/h2o-wave/) is a new open-source Python development framework that makes it fast and easy develop real-time interactive AI apps with sophisticated visualizations.\n \nIt has all the features I like (python, plotly, matplotlib, markdown, simple UI components) with little overhead. Tbh there are other tools out there that would probably do the job too (flask, shiny, dash, etc.) and I used all of them before I have more experience with Wave as I work at H2O.ai. I have also built a Wave app based on the Cornell competition earlier so I could reuse a few components.\n\n## Training setup\n \nWe experimented (thanks to Peter) with quite a few different training setups.  \nEventually what worked the best was probably the most straightforward, multiclass multi-label training with BCELoss. It was much faster than using 24 binary classifiers and - what is much more important at kaggle - it had better performance. We used fixed 3 second chunks Mel spectrograms for 0-14 kHz to cover all species.\n\nOne disadvantage was that it required much more annotation.\nBCELoss would penalize the most frequent probably unannotated classes and those classes matter a lot in LWLRAP.   \n\nWe started an iterative semi-supervised approach collecting new positive labels from the most confident OOF predictions then retrained our models to collect more samples... \n\nFor the annotations creating bounding boxes could be interesting but it would be quite slow painful. Since we only had to submit clipwise predictions it would be probably overkill too.\nIt was easier to review each class separately and make quick binary decisions whether add or delete the new candidates.\nIn the beginning I just used hacky scripts to save a bunch of candidate figures in separate directories then deleted the noise manually.\nLater I used the app to go through chunks again by checking the false labels with high predictions and true labels with low predictions. This way I also had logs of true/false/not sure examples for later experiments.  \n\nIn our final collection we had 23K positive labels for 11K chunks collected from 2913 training recordings.\nJust training a few CNN14s on this dataset would reach 0.95 with a few hours training time.\n  \n## The Hitchhiker's Guide to the Rainforest\nFor some of the classes the patterns were not that clear (e.g. 9, 17) we also had issues with 3 and 12 for quite a while. These ~~birds~~ frogs like to ~~sing~~ croak together a lot and they have quite similar patterns.\n\nLast Sunday we saw that others found the actual species published (https://www.sciencedirect.com/science/article/pii/S1574954120300637).\nI checked quickly our predictions and tried to fix a few thousand labels but our score did not really improve. Finding this article in advance would have saved a lot of efforts though...\n\n\n## The App\nhttps://github.com/gaborfodor/wave-soundscape-annotation\nSimple binary annotation decisions took 1-3s after we learned all the species. (The gif is not realtime :)\n![annotate](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/annotate.gif?raw=true)\n\nBeside fixing training labels the app can also classify new recordings.\n![](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/recognize.png?raw=true)",
    "1207656": "gaborfodor Congrats for your 2nd Grandmaster title!",
    "1207671": "Congratulations @gaborfodor !",
    "1207723": "good work!\n\nkaggle should give a prize for works like this (e.g. annotation tools, visualization tools, etc)",
    "1208283": "Thanks. Initially I wanted to help the host to collect more training labels but I had to find out later they already had a lot of labels just did not share them with us :)",
    "1208448": "Great work. Good collections. Thank you @gaborfodor for introducing H2O Wave. 🙏",
    "1208466": "Please also check this cool app by @fatihozturk: https://www.kaggle.com/c/stanford-covid-vaccine/discussion/215120",
    "1208469": "Congrats Gabor! I love the Wave app!",
    "1208572": "We checked our models predictions for different species and noticed that even our latest models trained with tons of pseudo labels and augmentation got quite high activations for new unseen species. \n\nFor example here are the predictions for [Hermit Trush](https://ebird.org/species/herthr)\n![Hermit Trush](https://github.com/gaborfodor/wave-soundscape-annotation/blob/main/data/herthr.png?raw=true)\n\nWe tried to use external Xeno-Canto data to force more generalized models but they did not help on the LB not even in our final ensemble...",
    "1208589": "Congrats @gaborfodor and @pestipeti and thanks for the writeup",
    "1208616": "I can finally see this amazing app @gaborfodor !!!!!\n\nCongratulations, this is awesome.",
    "1208632": "Awesome tool @gaborfodor. We had something similar (though not as beautiful as your app) for checking labels and also for adding labels to underrepresented species in the train set. \nH2O Wave seems to be very well suited for the task. Thanks a lot for sharing.\n\nWe did see dimishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP)",
    "1208827": "I think the test set annotation was also done with ARBIMON Pattern Matching plaform. Relying on automated high NCC template matching probably prefers the more obvious examples. I also noticed some possible misclassification in the provided FP samples.\n\nOur models were able to learn **our** labels with 0.99+ score but that did not translate to the LB (at least without proper postprocessing).",
    "1211579": "amazing work!!!👍",
    "1212771": "Thanks for open-sourcing this. A nice gesture! I hope this can be used by folks on the field...Hope it has come to the attention of the hosts..\n\nYou are my heroes in this comp!",
    "1213037": "Thanks. :)\nI am also curious about the hosts view. There are lot of unique solutions in the top 10. Probably the winning model documentation process is still in progress. I hope we will get a summary from the hosts soon.",
    "1216685": "Great app.. Thank you for sharing the knowledge by open sourcing the app."
  },
  "source": "meta"
}