{
  "id": 183407,
  "title": "10th Place Solution",
  "url": "/competitions/birdsong-recognition/writeups/bar-zdabilleget-10th-place-solution",
  "author_name": "",
  "post_date": "2020-09-17T16:58:04.860Z",
  "votes": 30,
  "comment_count": 7,
  "views": 0,
  "content": "<h2>Acknowledgments</h2>\n<p>First of all thanks for the organizers for this challenging and fascinating competition.<br>\nI am grateful for the relatively quick answers in the discussions (external data, domain knowledge,<br>\nit was certainly helpful.</p>\n<p>Special thanks to</p>\n<ul>\n<li>Hidehisa Arai for showing how to use PANNS effectively</li>\n<li>Qiuqiang Kong et al for PANNS repo and pretrained models</li>\n<li>Vopani for scraping and maintaining the external XenoCanto dataset</li>\n<li>Jan Schlüter and Mario Lasseck for their previous BirdCLEF winning papers   </li>\n</ul>\n<h2>The beginning</h2>\n<p>Seven years ago I already participated in bird song detection challenge at kaggle.<br>\nThat time we had only a few hundred 10 seconds recordings to train on and similar multiclass multilabel problem but with only 19 species.<br>\nI was able to win that competition with Computer Vision template matching and Random Forests.<br>\nI thought it would be a quick &amp; easy experiment to beat that with the available 40K bird recordings and all the available pretrained image net models.<br>\nWell it was not.</p>\n<p>Possible reasons</p>\n<ul>\n<li>Different sample rate (16kHz vs 32kHz)</li>\n<li>Soundscape vs Xenocanto</li>\n<li>MLSP train-test split was random split across soundscapes it was possible to overfit to the same recording (e.g. crickets &amp; rain -&gt; Hermit Warbler))</li>\n<li>Different spectrogram/noise distribution</li>\n</ul>\n<h2>Ornithology and LB Probing</h2>\n<p>In the beginning of the competition I was not able to submit meaningful results so I tried to understand the North American bird population better.<br>\nJust by submitting individual birds one would expect ~0.001 LB score with equal bird distribution.<br>\nFrom ebird.org observation data I was able to rank the 264 species based on their unique observations during the last two years.<br>\nIt does not necessary reflect the distribution in the Public/Private test set but I found it better than the number of XC recordings.</p>\n<p>E.g.</p>\n<ul>\n<li>Red Crossbill 1223 XC recordings 84K observations 0.000 LB Score</li>\n<li>White-crowned Sparrow 474 XC recordings 1.1M observations 0.1 LB Score (!)</li>\n</ul>\n<h2>Data Preparation</h2>\n<p>Resampling everything to 32kHz and splitting the first 2 minutes of each recording to 10 second duration chunks and saving them as .npy arrays.<br>\nI used the extended dataset a<br>\n4 fold cross validation was used stratified on the author-created at to try to avoid same birds in different folds.<br>\nFor early stopping I saved the best  weights based on XC validation and BirdCLEF Validation as well.</p>\n<h2>Augmentation</h2>\n<p>I only used additive noises.</p>\n<ul>\n<li>freefield1010</li>\n<li>warblrb10k</li>\n<li>BirdVox-DCASE-20k</li>\n<li>Animal Sound Archive Published by Museum für Naturkunde Berlin ()</li>\n</ul>\n<p>Probably should have tried synthetic noise generation too.</p>\n<h2>Architecture</h2>\n<p>I ended up with slightly modified CNN14 (128 mel bins, mean/std standardised)<br>\nThey were relatively quick to train on Nvidia Tesla T4, training a single model took 3-8 hours.<br>\nI tried PANN ResNet38, Cnn14_DecisionLevelAtt or ImageNet pretrained ResNet50 but without proper validation I got mixed results…</p>\n<p>The dataloader handled the additive augmentations for the waveforms</p>\n<ul>\n<li>Add multiple possible 1-2-3 birds with multi-class setting</li>\n<li>Add same class chunks</li>\n<li>Add noise</li>\n<li>Add animal sound</li>\n</ul>\n<p>the GPU created the spectrograms for the batches.</p>\n<h2>Blending</h2>\n<p>During the last weekend I fixed my whole training pipeline and rented V100 to retrain a few final models and create some additional experiments.<br>\nOn Sunday evening I submitted my first blended model with quite disappointing 0.570 Public Leaderboard result.<br>\nActually it would have been enough for my final 10th place with Private score 0.649 Private Score.</p>\n<p>In the las two days I made some desperate submissions with more models, varying thresholds<br>\n(e.g. increasing the thresholds of west coast birds, reducing the thresholds for common birds)  <br>\nThey did improve my public LB score and fortunately they did not improve nor hurt the private score.<br>\nProbably with a few more submissions I would start to overfit…</p>\n<h2>What did not work but I thought it would…</h2>\n<ul>\n<li>Using secondary labels to improve primary labels with oof predictions</li>\n<li>Using a separate nocall classifier</li>\n<li>Utilizing the apriori knowledge that some birds are more frequent than others</li>\n<li>Blending normalized and unnormalized models</li>\n<li>Mixup</li>\n<li>More than 128 Mel bins</li>\n</ul>",
  "messages": [
    {
      "id": "1013184",
      "postDate": "09/16/2020 14:53:02",
      "content": "<h2>Acknowledgments</h2>\n<p>First of all thanks for the organizers for this challenging and fascinating competition.<br>\nI am grateful for the relatively quick answers in the discussions (external data, domain knowledge,<br>\nit was certainly helpful.</p>\n<p>Special thanks to</p>\n<ul>\n<li>Hidehisa Arai for showing how to use PANNS effectively</li>\n<li>Qiuqiang Kong et al for PANNS repo and pretrained models</li>\n<li>Vopani for scraping and maintaining the external XenoCanto dataset</li>\n<li>Jan Schlüter and Mario Lasseck for their previous BirdCLEF winning papers   </li>\n</ul>\n<h2>The beginning</h2>\n<p>Seven years ago I already participated in bird song detection challenge at kaggle.<br>\nThat time we had only a few hundred 10 seconds recordings to train on and similar multiclass multilabel problem but with only 19 species.<br>\nI was able to win that competition with Computer Vision template matching and Random Forests.<br>\nI thought it would be a quick &amp; easy experiment to beat that with the available 40K bird recordings and all the available pretrained image net models.<br>\nWell it was not.</p>\n<p>Possible reasons</p>\n<ul>\n<li>Different sample rate (16kHz vs 32kHz)</li>\n<li>Soundscape vs Xenocanto</li>\n<li>MLSP train-test split was random split across soundscapes it was possible to overfit to the same recording (e.g. crickets &amp; rain -&gt; Hermit Warbler))</li>\n<li>Different spectrogram/noise distribution</li>\n</ul>\n<h2>Ornithology and LB Probing</h2>\n<p>In the beginning of the competition I was not able to submit meaningful results so I tried to understand the North American bird population better.<br>\nJust by submitting individual birds one would expect ~0.001 LB score with equal bird distribution.<br>\nFrom ebird.org observation data I was able to rank the 264 species based on their unique observations during the last two years.<br>\nIt does not necessary reflect the distribution in the Public/Private test set but I found it better than the number of XC recordings.</p>\n<p>E.g.</p>\n<ul>\n<li>Red Crossbill 1223 XC recordings 84K observations 0.000 LB Score</li>\n<li>White-crowned Sparrow 474 XC recordings 1.1M observations 0.1 LB Score (!)</li>\n</ul>\n<h2>Data Preparation</h2>\n<p>Resampling everything to 32kHz and splitting the first 2 minutes of each recording to 10 second duration chunks and saving them as .npy arrays.<br>\nI used the extended dataset a<br>\n4 fold cross validation was used stratified on the author-created at to try to avoid same birds in different folds.<br>\nFor early stopping I saved the best  weights based on XC validation and BirdCLEF Validation as well.</p>\n<h2>Augmentation</h2>\n<p>I only used additive noises.</p>\n<ul>\n<li>freefield1010</li>\n<li>warblrb10k</li>\n<li>BirdVox-DCASE-20k</li>\n<li>Animal Sound Archive Published by Museum für Naturkunde Berlin ()</li>\n</ul>\n<p>Probably should have tried synthetic noise generation too.</p>\n<h2>Architecture</h2>\n<p>I ended up with slightly modified CNN14 (128 mel bins, mean/std standardised)<br>\nThey were relatively quick to train on Nvidia Tesla T4, training a single model took 3-8 hours.<br>\nI tried PANN ResNet38, Cnn14_DecisionLevelAtt or ImageNet pretrained ResNet50 but without proper validation I got mixed results…</p>\n<p>The dataloader handled the additive augmentations for the waveforms</p>\n<ul>\n<li>Add multiple possible 1-2-3 birds with multi-class setting</li>\n<li>Add same class chunks</li>\n<li>Add noise</li>\n<li>Add animal sound</li>\n</ul>\n<p>the GPU created the spectrograms for the batches.</p>\n<h2>Blending</h2>\n<p>During the last weekend I fixed my whole training pipeline and rented V100 to retrain a few final models and create some additional experiments.<br>\nOn Sunday evening I submitted my first blended model with quite disappointing 0.570 Public Leaderboard result.<br>\nActually it would have been enough for my final 10th place with Private score 0.649 Private Score.</p>\n<p>In the las two days I made some desperate submissions with more models, varying thresholds<br>\n(e.g. increasing the thresholds of west coast birds, reducing the thresholds for common birds)  <br>\nThey did improve my public LB score and fortunately they did not improve nor hurt the private score.<br>\nProbably with a few more submissions I would start to overfit…</p>\n<h2>What did not work but I thought it would…</h2>\n<ul>\n<li>Using secondary labels to improve primary labels with oof predictions</li>\n<li>Using a separate nocall classifier</li>\n<li>Utilizing the apriori knowledge that some birds are more frequent than others</li>\n<li>Blending normalized and unnormalized models</li>\n<li>Mixup</li>\n<li>More than 128 Mel bins</li>\n</ul>",
      "rawMarkdown": "## Acknowledgments\nFirst of all thanks for the organizers for this challenging and fascinating competition.\nI am grateful for the relatively quick answers in the discussions (external data, domain knowledge,\nit was certainly helpful.\n\nSpecial thanks to\n* Hidehisa Arai for showing how to use PANNS effectively\n* Qiuqiang Kong et al for PANNS repo and pretrained models\n* Vopani for scraping and maintaining the external XenoCanto dataset\n* Jan Schlüter and Mario Lasseck for their previous BirdCLEF winning papers   \n\n## The beginning\nSeven years ago I already participated in bird song detection challenge at kaggle.\nThat time we had only a few hundred 10 seconds recordings to train on and similar multiclass multilabel problem but with only 19 species.\nI was able to win that competition with Computer Vision template matching and Random Forests.\nI thought it would be a quick & easy experiment to beat that with the available 40K bird recordings and all the available pretrained image net models.\nWell it was not.\n\nPossible reasons\n* Different sample rate (16kHz vs 32kHz)\n* Soundscape vs Xenocanto\n* MLSP train-test split was random split across soundscapes it was possible to overfit to the same recording (e.g. crickets & rain -> Hermit Warbler))\n* Different spectrogram/noise distribution\n\n## Ornithology and LB Probing\nIn the beginning of the competition I was not able to submit meaningful results so I tried to understand the North American bird population better.\nJust by submitting individual birds one would expect ~0.001 LB score with equal bird distribution.\nFrom ebird.org observation data I was able to rank the 264 species based on their unique observations during the last two years.\nIt does not necessary reflect the distribution in the Public/Private test set but I found it better than the number of XC recordings.\n\nE.g.\n* Red Crossbill 1223 XC recordings 84K observations 0.000 LB Score\n* White-crowned Sparrow 474 XC recordings 1.1M observations 0.1 LB Score (!)\n\n## Data Preparation\nResampling everything to 32kHz and splitting the first 2 minutes of each recording to 10 second duration chunks and saving them as .npy arrays.\nI used the extended dataset a\n4 fold cross validation was used stratified on the author-created at to try to avoid same birds in different folds.\nFor early stopping I saved the best  weights based on XC validation and BirdCLEF Validation as well.\n\n## Augmentation\nI only used additive noises.\n* freefield1010\n* warblrb10k\n* BirdVox-DCASE-20k\n* Animal Sound Archive Published by Museum für Naturkunde Berlin ()\n\nProbably should have tried synthetic noise generation too.\n\n## Architecture\nI ended up with slightly modified CNN14 (128 mel bins, mean/std standardised)\nThey were relatively quick to train on Nvidia Tesla T4, training a single model took 3-8 hours.\nI tried PANN ResNet38, Cnn14_DecisionLevelAtt or ImageNet pretrained ResNet50 but without proper validation I got mixed results...\n\nThe dataloader handled the additive augmentations for the waveforms\n* Add multiple possible 1-2-3 birds with multi-class setting\n* Add same class chunks\n* Add noise\n* Add animal sound\n\nthe GPU created the spectrograms for the batches.\n\n \n## Blending\nDuring the last weekend I fixed my whole training pipeline and rented V100 to retrain a few final models and create some additional experiments.\nOn Sunday evening I submitted my first blended model with quite disappointing 0.570 Public Leaderboard result.\nActually it would have been enough for my final 10th place with Private score 0.649 Private Score.\n\nIn the las two days I made some desperate submissions with more models, varying thresholds\n(e.g. increasing the thresholds of west coast birds, reducing the thresholds for common birds)  \nThey did improve my public LB score and fortunately they did not improve nor hurt the private score.\nProbably with a few more submissions I would start to overfit...\n\n## What did not work but I thought it would...\n* Using secondary labels to improve primary labels with oof predictions\n* Using a separate nocall classifier\n* Utilizing the apriori knowledge that some birds are more frequent than others\n* Blending normalized and unnormalized models\n* Mixup\n* More than 128 Mel bins",
      "votes": null
    },
    {
      "id": "1013193",
      "postDate": "09/16/2020 14:56:25",
      "content": "<p>Congrats on the strong finish!</p>",
      "rawMarkdown": "Congrats on the strong finish!",
      "votes": null
    },
    {
      "id": "1013223",
      "postDate": "09/16/2020 15:12:40",
      "content": "<p>Thanks, actually your 2w marching was more impressive. I started the comp quite early :)</p>",
      "rawMarkdown": "Thanks, actually your 2w marching was more impressive. I started the comp quite early :)",
      "votes": null
    },
    {
      "id": "1013729",
      "postDate": "09/16/2020 22:55:41",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> !<br>\nI found you joined from the very beginning and I thought you'd be in high place soon. Contrary to my prediction, you hide long time in the LB but finally showed up to the gold zone! It was really impressive!</p>",
      "rawMarkdown": "Congratulations @gaborfodor !\nI found you joined from the very beginning and I thought you'd be in high place soon. Contrary to my prediction, you hide long time in the LB but finally showed up to the gold zone! It was really impressive!",
      "votes": null
    },
    {
      "id": "1014715",
      "postDate": "09/17/2020 16:48:43",
      "content": "<p>I have two possible explanations for that :) </p>\n<p>This was my first competition with pytorch and it took me some time to implement/train models that worked well in the past competitions (AudioSet, BirdCLEF) </p>\n<p>Well, it took more time to make them working for this challenge :)</p>\n<p>I spent most of my submissions for experiments to understand the test data:</p>\n<ul>\n<li>LB probing for individual species</li>\n<li>Nocall classifier experiments</li>\n<li>Test spectrogram statistics</li>\n</ul>\n<p>This helped me to understand what is the main difference between test and train.</p>\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1014715/17029/std.png\" alt=\"\"></p>",
      "rawMarkdown": "I have two possible explanations for that :) \n\nThis was my first competition with pytorch and it took me some time to implement/train models that worked well in the past competitions (AudioSet, BirdCLEF) \n\nWell, it took more time to make them working for this challenge :)\n\nI spent most of my submissions for experiments to understand the test data:\n* LB probing for individual species\n* Nocall classifier experiments\n* Test spectrogram statistics\n\nThis helped me to understand what is the main difference between test and train.\n\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/1014715/17029/std.png)",
      "votes": null
    },
    {
      "id": "1015617",
      "postDate": "09/18/2020 09:45:43",
      "content": "<p>Make sense 🙂</p>\n<p>The figure seems super-interesting. <br>\nIs x-axis mel-band index (using <code>n_mels=128</code>)?</p>",
      "rawMarkdown": "Make sense 🙂\n\nThe figure seems super-interesting. \nIs x-axis mel-band index (using `n_mels=128`)?",
      "votes": null
    },
    {
      "id": "1015623",
      "postDate": "09/18/2020 09:55:31",
      "content": "<p>Yep, I should have labeled my axis :) The dashed lines are mean of 90 percentiles the solid lines are the mean for each mel band. I think fmin=300 fmax=11000 was used. </p>\n<p>I had to kill/delete all gcp instances and data so it would be difficult to recreate the chart…</p>",
      "rawMarkdown": "Yep, I should have labeled my axis :) The dashed lines are mean of 90 percentiles the solid lines are the mean for each mel band. I think fmin=300 fmax=11000 was used. \n\nI had to kill/delete all gcp instances and data so it would be difficult to recreate the chart...",
      "votes": null
    },
    {
      "id": "1015633",
      "postDate": "09/18/2020 10:10:26",
      "content": "<p>Thanks!<br>\nSo the slope comes from the nature of pink noise I think, and the bulge in the middle of train/Xeno-Canto comes from those signals of various species (and also from some nuisance signals)…</p>\n<p>What's interesting is that <code>example_test_audio</code> is not representing how the test should looks like - I now get what you meant in this part of the other post </p>\n<blockquote>\n  <p>probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches.</p>\n</blockquote>",
      "rawMarkdown": "Thanks!\nSo the slope comes from the nature of pink noise I think, and the bulge in the middle of train/Xeno-Canto comes from those signals of various species (and also from some nuisance signals)...\n\nWhat's interesting is that `example_test_audio` is not representing how the test should looks like - I now get what you meant in this part of the other post \n\n> probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1013193,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/16/2020 14:56:25",
      "content": "<p>Congrats on the strong finish!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1013223,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "09/16/2020 15:12:40",
          "content": "<p>Thanks, actually your 2w marching was more impressive. I started the comp quite early :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013729,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "09/16/2020 22:55:41",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> !<br>\nI found you joined from the very beginning and I thought you'd be in high place soon. Contrary to my prediction, you hide long time in the LB but finally showed up to the gold zone! It was really impressive!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1014715,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "09/17/2020 16:48:43",
          "content": "<p>I have two possible explanations for that :) </p>\n<p>This was my first competition with pytorch and it took me some time to implement/train models that worked well in the past competitions (AudioSet, BirdCLEF) </p>\n<p>Well, it took more time to make them working for this challenge :)</p>\n<p>I spent most of my submissions for experiments to understand the test data:</p>\n<ul>\n<li>LB probing for individual species</li>\n<li>Nocall classifier experiments</li>\n<li>Test spectrogram statistics</li>\n</ul>\n<p>This helped me to understand what is the main difference between test and train.</p>\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/1014715/17029/std.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1015617,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "09/18/2020 09:45:43",
          "content": "<p>Make sense 🙂</p>\n<p>The figure seems super-interesting. <br>\nIs x-axis mel-band index (using <code>n_mels=128</code>)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1015623,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "09/18/2020 09:55:31",
          "content": "<p>Yep, I should have labeled my axis :) The dashed lines are mean of 90 percentiles the solid lines are the mean for each mel band. I think fmin=300 fmax=11000 was used. </p>\n<p>I had to kill/delete all gcp instances and data so it would be difficult to recreate the chart…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1015633,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "09/18/2020 10:10:26",
          "content": "<p>Thanks!<br>\nSo the slope comes from the nature of pink noise I think, and the bulge in the middle of train/Xeno-Canto comes from those signals of various species (and also from some nuisance signals)…</p>\n<p>What's interesting is that <code>example_test_audio</code> is not representing how the test should looks like - I now get what you meant in this part of the other post </p>\n<blockquote>\n  <p>probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1013184": "## Acknowledgments\nFirst of all thanks for the organizers for this challenging and fascinating competition.\nI am grateful for the relatively quick answers in the discussions (external data, domain knowledge,\nit was certainly helpful.\n\nSpecial thanks to\n* Hidehisa Arai for showing how to use PANNS effectively\n* Qiuqiang Kong et al for PANNS repo and pretrained models\n* Vopani for scraping and maintaining the external XenoCanto dataset\n* Jan Schlüter and Mario Lasseck for their previous BirdCLEF winning papers   \n\n## The beginning\nSeven years ago I already participated in bird song detection challenge at kaggle.\nThat time we had only a few hundred 10 seconds recordings to train on and similar multiclass multilabel problem but with only 19 species.\nI was able to win that competition with Computer Vision template matching and Random Forests.\nI thought it would be a quick & easy experiment to beat that with the available 40K bird recordings and all the available pretrained image net models.\nWell it was not.\n\nPossible reasons\n* Different sample rate (16kHz vs 32kHz)\n* Soundscape vs Xenocanto\n* MLSP train-test split was random split across soundscapes it was possible to overfit to the same recording (e.g. crickets & rain -> Hermit Warbler))\n* Different spectrogram/noise distribution\n\n## Ornithology and LB Probing\nIn the beginning of the competition I was not able to submit meaningful results so I tried to understand the North American bird population better.\nJust by submitting individual birds one would expect ~0.001 LB score with equal bird distribution.\nFrom ebird.org observation data I was able to rank the 264 species based on their unique observations during the last two years.\nIt does not necessary reflect the distribution in the Public/Private test set but I found it better than the number of XC recordings.\n\nE.g.\n* Red Crossbill 1223 XC recordings 84K observations 0.000 LB Score\n* White-crowned Sparrow 474 XC recordings 1.1M observations 0.1 LB Score (!)\n\n## Data Preparation\nResampling everything to 32kHz and splitting the first 2 minutes of each recording to 10 second duration chunks and saving them as .npy arrays.\nI used the extended dataset a\n4 fold cross validation was used stratified on the author-created at to try to avoid same birds in different folds.\nFor early stopping I saved the best  weights based on XC validation and BirdCLEF Validation as well.\n\n## Augmentation\nI only used additive noises.\n* freefield1010\n* warblrb10k\n* BirdVox-DCASE-20k\n* Animal Sound Archive Published by Museum für Naturkunde Berlin ()\n\nProbably should have tried synthetic noise generation too.\n\n## Architecture\nI ended up with slightly modified CNN14 (128 mel bins, mean/std standardised)\nThey were relatively quick to train on Nvidia Tesla T4, training a single model took 3-8 hours.\nI tried PANN ResNet38, Cnn14_DecisionLevelAtt or ImageNet pretrained ResNet50 but without proper validation I got mixed results...\n\nThe dataloader handled the additive augmentations for the waveforms\n* Add multiple possible 1-2-3 birds with multi-class setting\n* Add same class chunks\n* Add noise\n* Add animal sound\n\nthe GPU created the spectrograms for the batches.\n\n \n## Blending\nDuring the last weekend I fixed my whole training pipeline and rented V100 to retrain a few final models and create some additional experiments.\nOn Sunday evening I submitted my first blended model with quite disappointing 0.570 Public Leaderboard result.\nActually it would have been enough for my final 10th place with Private score 0.649 Private Score.\n\nIn the las two days I made some desperate submissions with more models, varying thresholds\n(e.g. increasing the thresholds of west coast birds, reducing the thresholds for common birds)  \nThey did improve my public LB score and fortunately they did not improve nor hurt the private score.\nProbably with a few more submissions I would start to overfit...\n\n## What did not work but I thought it would...\n* Using secondary labels to improve primary labels with oof predictions\n* Using a separate nocall classifier\n* Utilizing the apriori knowledge that some birds are more frequent than others\n* Blending normalized and unnormalized models\n* Mixup\n* More than 128 Mel bins",
    "1013193": "Congrats on the strong finish!",
    "1013223": "Thanks, actually your 2w marching was more impressive. I started the comp quite early :)",
    "1013729": "Congratulations @gaborfodor !\nI found you joined from the very beginning and I thought you'd be in high place soon. Contrary to my prediction, you hide long time in the LB but finally showed up to the gold zone! It was really impressive!",
    "1014715": "I have two possible explanations for that :) \n\nThis was my first competition with pytorch and it took me some time to implement/train models that worked well in the past competitions (AudioSet, BirdCLEF) \n\nWell, it took more time to make them working for this challenge :)\n\nI spent most of my submissions for experiments to understand the test data:\n* LB probing for individual species\n* Nocall classifier experiments\n* Test spectrogram statistics\n\nThis helped me to understand what is the main difference between test and train.\n\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/1014715/17029/std.png)",
    "1015617": "Make sense 🙂\n\nThe figure seems super-interesting. \nIs x-axis mel-band index (using `n_mels=128`)?",
    "1015623": "Yep, I should have labeled my axis :) The dashed lines are mean of 90 percentiles the solid lines are the mean for each mel band. I think fmin=300 fmax=11000 was used. \n\nI had to kill/delete all gcp instances and data so it would be difficult to recreate the chart...",
    "1015633": "Thanks!\nSo the slope comes from the nature of pink noise I think, and the bulge in the middle of train/Xeno-Canto comes from those signals of various species (and also from some nuisance signals)...\n\nWhat's interesting is that `example_test_audio` is not representing how the test should looks like - I now get what you meant in this part of the other post \n\n> probably it is quite easy to collect soundscapes without annotations. That would had helped a lot for understanding how much noise and what kind of noise should we add to the training dataset. It would also allow to use innovative un/semi-supervised approaches."
  },
  "source": "meta"
}