{
  "id": 183436,
  "title": "13th place digest",
  "url": "/competitions/birdsong-recognition/writeups/tmz-13th-place-digest",
  "author_name": "",
  "post_date": "2020-09-16T18:11:43.683Z",
  "votes": 27,
  "comment_count": 11,
  "views": 0,
  "content": "<p>It was my first audio competition and I really enjoyed it! Thanks to organizers and Kaggle for it.<br>\nThanks to my teammates <a href=\"https://www.kaggle.com/tikutiku\" target=\"_blank\">@tikutiku</a> <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> the great collaboration, it was a pleasure and I've learnt a lot again.</p>\n<p>Here is a digest of our solution.</p>\n<p><strong>Main Pipeline</strong> (see image below):</p>\n<ul>\n<li>MEL-Spectrogram (torchaudio) with differents settings to capture most birds frequency shapes.</li>\n<li>Most processing/augmentation occurs in GPU to speed up training and inference.</li>\n</ul>\n<p><strong>Time Augmentations</strong>:</p>\n<ul>\n<li>Pink Noise/White noise</li>\n<li>Time roll</li>\n<li>Volume gain</li>\n<li>PitchShift</li>\n<li>Additional bandpass filters,lowcut (1kHz-2.5kHz), highcut (10kHz-15kHz)</li>\n<li>Background/Ambient noise Mixup</li>\n<li>Other birds Mixup (2 to 3 birds)</li>\n</ul>\n<p><strong>Spec/Image augmentations</strong>:</p>\n<ul>\n<li>Frequency/time masking</li>\n<li>Color jitter</li>\n</ul>\n<p><strong>CNN backbones</strong>:</p>\n<ul>\n<li>EfficientNet B1 and B2</li>\n<li>SEReseXt26</li>\n<li>ResneSt50</li>\n<li>Optional GRU layer to learn about time sequences (see <a href=\"https://github.com/srvk/TALNet\" target=\"_blank\">TALNet Sound Event Classifier</a>)</li>\n<li>Different pooling to apply SED concept with either Attention block or max pooling<br>\nWe also had one model based on 1D signal only (DensetNet1D)</li>\n</ul>\n<p>What did not work well (it worked but was disapointing):</p>\n<ul>\n<li>Wavegram + MEL-Spectrogram model (was bad with soundscape)</li>\n<li>Create large image with 3x2 grid spectrogram (was bad on inference)</li>\n</ul>\n<p><strong>Data</strong>:</p>\n<ul>\n<li>Train audio provided with duration outliers removed</li>\n<li>Additional xeno-canto data (from Vopani dataset, thanks <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> for your clean dump)</li>\n<li>Distractors (NoBird/Nocall/Ambient) 10s slices (from freefield1010)<br>\nWe've built our own validation data by mixing test_audio/birds/noise/nocall to try to correlate LB. It correlated a bit but was not enough to give trust in it. Too bad as it was key for this competition.</li>\n</ul>\n<p><strong>Training procedure</strong>:</p>\n<ul>\n<li>Stage1: Train a few models with 5s slices picked randomly, then save birds probabilities on CV OOF, ensemble all OOF models results to generate \"<em>hot</em>\" slices with high probabilities.<br>\nIt allowed to reach public LB=0.582</li>\n<li>Stage2: Train more models with only such 5s <em>hot slices</em>.<br>\nIt allowed to reach public LB=0.596<br>\nWhile reading other's solutions, we should have tried another stage with \"very hot\" slices to have a super clean train dataset.</li>\n</ul>\n<p><strong>Final ensemble</strong>:<br>\nEnsemble is a combination of models (8 to 10) with union strategy and with per-model threshold tuned on public LB (it overfitted for sure).<br>\nIdea of union was to capture most of the TP (better for metric used in this competition), drawback is that we capture FP too.<br>\nInference was fast, we cached resampled audio in memory, and (almost) full GPU pipeline helped. It ran in around 1h30 so we still had room for more models.</p>\n<p><strong>Post-processing</strong>:<br>\nWe have a \"nobird\" majority vote to try to remove FP (force nocall) as some models have been trained with 265 classes instead of 264 (class 265 = nobird)</p>\n<p>Both our final submissions reached similar score on private LB.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F4e2db4b7460eb1a51737fbf81bd8ed50%2Fpipeline.png?generation=1600274872405000&amp;alt=media\" alt=\"\"></p>\n<p>Final words to conclude: Congratulations to top teams! and thanks to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for their sharings.</p>",
  "messages": [
    {
      "id": "1013357",
      "postDate": "09/16/2020 16:53:08",
      "content": "<p>It was my first audio competition and I really enjoyed it! Thanks to organizers and Kaggle for it.<br>\nThanks to my teammates <a href=\"https://www.kaggle.com/tikutiku\" target=\"_blank\">@tikutiku</a> <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> the great collaboration, it was a pleasure and I've learnt a lot again.</p>\n<p>Here is a digest of our solution.</p>\n<p><strong>Main Pipeline</strong> (see image below):</p>\n<ul>\n<li>MEL-Spectrogram (torchaudio) with differents settings to capture most birds frequency shapes.</li>\n<li>Most processing/augmentation occurs in GPU to speed up training and inference.</li>\n</ul>\n<p><strong>Time Augmentations</strong>:</p>\n<ul>\n<li>Pink Noise/White noise</li>\n<li>Time roll</li>\n<li>Volume gain</li>\n<li>PitchShift</li>\n<li>Additional bandpass filters,lowcut (1kHz-2.5kHz), highcut (10kHz-15kHz)</li>\n<li>Background/Ambient noise Mixup</li>\n<li>Other birds Mixup (2 to 3 birds)</li>\n</ul>\n<p><strong>Spec/Image augmentations</strong>:</p>\n<ul>\n<li>Frequency/time masking</li>\n<li>Color jitter</li>\n</ul>\n<p><strong>CNN backbones</strong>:</p>\n<ul>\n<li>EfficientNet B1 and B2</li>\n<li>SEReseXt26</li>\n<li>ResneSt50</li>\n<li>Optional GRU layer to learn about time sequences (see <a href=\"https://github.com/srvk/TALNet\" target=\"_blank\">TALNet Sound Event Classifier</a>)</li>\n<li>Different pooling to apply SED concept with either Attention block or max pooling<br>\nWe also had one model based on 1D signal only (DensetNet1D)</li>\n</ul>\n<p>What did not work well (it worked but was disapointing):</p>\n<ul>\n<li>Wavegram + MEL-Spectrogram model (was bad with soundscape)</li>\n<li>Create large image with 3x2 grid spectrogram (was bad on inference)</li>\n</ul>\n<p><strong>Data</strong>:</p>\n<ul>\n<li>Train audio provided with duration outliers removed</li>\n<li>Additional xeno-canto data (from Vopani dataset, thanks <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> for your clean dump)</li>\n<li>Distractors (NoBird/Nocall/Ambient) 10s slices (from freefield1010)<br>\nWe've built our own validation data by mixing test_audio/birds/noise/nocall to try to correlate LB. It correlated a bit but was not enough to give trust in it. Too bad as it was key for this competition.</li>\n</ul>\n<p><strong>Training procedure</strong>:</p>\n<ul>\n<li>Stage1: Train a few models with 5s slices picked randomly, then save birds probabilities on CV OOF, ensemble all OOF models results to generate \"<em>hot</em>\" slices with high probabilities.<br>\nIt allowed to reach public LB=0.582</li>\n<li>Stage2: Train more models with only such 5s <em>hot slices</em>.<br>\nIt allowed to reach public LB=0.596<br>\nWhile reading other's solutions, we should have tried another stage with \"very hot\" slices to have a super clean train dataset.</li>\n</ul>\n<p><strong>Final ensemble</strong>:<br>\nEnsemble is a combination of models (8 to 10) with union strategy and with per-model threshold tuned on public LB (it overfitted for sure).<br>\nIdea of union was to capture most of the TP (better for metric used in this competition), drawback is that we capture FP too.<br>\nInference was fast, we cached resampled audio in memory, and (almost) full GPU pipeline helped. It ran in around 1h30 so we still had room for more models.</p>\n<p><strong>Post-processing</strong>:<br>\nWe have a \"nobird\" majority vote to try to remove FP (force nocall) as some models have been trained with 265 classes instead of 264 (class 265 = nobird)</p>\n<p>Both our final submissions reached similar score on private LB.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F4e2db4b7460eb1a51737fbf81bd8ed50%2Fpipeline.png?generation=1600274872405000&amp;alt=media\" alt=\"\"></p>\n<p>Final words to conclude: Congratulations to top teams! and thanks to <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for their sharings.</p>",
      "rawMarkdown": "It was my first audio competition and I really enjoyed it! Thanks to organizers and Kaggle for it.\nThanks to my teammates @tikutiku @zfturbo the great collaboration, it was a pleasure and I've learnt a lot again.\n\nHere is a digest of our solution.\n\n**Main Pipeline** (see image below):\n- MEL-Spectrogram (torchaudio) with differents settings to capture most birds frequency shapes.\n- Most processing/augmentation occurs in GPU to speed up training and inference.\n\n**Time Augmentations**:\n- Pink Noise/White noise\n- Time roll\n- Volume gain\n- PitchShift\n- Additional bandpass filters,lowcut (1kHz-2.5kHz), highcut (10kHz-15kHz)\n- Background/Ambient noise Mixup\n- Other birds Mixup (2 to 3 birds)\n\n**Spec/Image augmentations**:\n- Frequency/time masking\n- Color jitter\n\n**CNN backbones**:\n- EfficientNet B1 and B2\n- SEReseXt26\n- ResneSt50\n- Optional GRU layer to learn about time sequences (see [TALNet Sound Event Classifier](https://github.com/srvk/TALNet))\n- Different pooling to apply SED concept with either Attention block or max pooling\nWe also had one model based on 1D signal only (DensetNet1D)\n\nWhat did not work well (it worked but was disapointing):\n- Wavegram + MEL-Spectrogram model (was bad with soundscape)\n- Create large image with 3x2 grid spectrogram (was bad on inference)\n\n**Data**:\n- Train audio provided with duration outliers removed\n- Additional xeno-canto data (from Vopani dataset, thanks @rohanrao for your clean dump)\n- Distractors (NoBird/Nocall/Ambient) 10s slices (from freefield1010)\nWe've built our own validation data by mixing test_audio/birds/noise/nocall to try to correlate LB. It correlated a bit but was not enough to give trust in it. Too bad as it was key for this competition.\n\n**Training procedure**:\n- Stage1: Train a few models with 5s slices picked randomly, then save birds probabilities on CV OOF, ensemble all OOF models results to generate \"*hot*\" slices with high probabilities.\n  It allowed to reach public LB=0.582\n- Stage2: Train more models with only such 5s *hot slices*.\n  It allowed to reach public LB=0.596\nWhile reading other's solutions, we should have tried another stage with \"very hot\" slices to have a super clean train dataset.\n\n**Final ensemble**:\nEnsemble is a combination of models (8 to 10) with union strategy and with per-model threshold tuned on public LB (it overfitted for sure).\nIdea of union was to capture most of the TP (better for metric used in this competition), drawback is that we capture FP too.\nInference was fast, we cached resampled audio in memory, and (almost) full GPU pipeline helped. It ran in around 1h30 so we still had room for more models.\n\n**Post-processing**:\nWe have a \"nobird\" majority vote to try to remove FP (force nocall) as some models have been trained with 265 classes instead of 264 (class 265 = nobird)\n\nBoth our final submissions reached similar score on private LB.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F4e2db4b7460eb1a51737fbf81bd8ed50%2Fpipeline.png?generation=1600274872405000&alt=media)\n\nFinal words to conclude: Congratulations to top teams! and thanks to @hidehisaarai1213, @hengck23 for their sharings.",
      "votes": null
    },
    {
      "id": "1013387",
      "postDate": "09/16/2020 17:14:01",
      "content": "<p>Great…. Just one step away 👍</p>",
      "rawMarkdown": "Great.... Just one step away 👍",
      "votes": null
    },
    {
      "id": "1013388",
      "postDate": "09/16/2020 17:14:44",
      "content": "<p>Thanks for sharing now and during discussions during the competition too.<br>\nIf possible, can you elaborate a bit more on this point please…</p>\n<blockquote>\n  <p>ensemble all OOF models results to generate \"hot\" slices with high probabilities.</p>\n</blockquote>",
      "rawMarkdown": "Thanks for sharing now and during discussions during the competition too.\nIf possible, can you elaborate a bit more on this point please...\n>ensemble all OOF models results to generate \"hot\" slices with high probabilities.",
      "votes": null
    },
    {
      "id": "1013394",
      "postDate": "09/16/2020 17:16:54",
      "content": "<p>Thanks, we had our best public LB submission with 0.647 on private LB (last gold) but we've selected the one with 0.646 that had more diversity in ensemble.</p>",
      "rawMarkdown": "Thanks, we had our best public LB submission with 0.647 on private LB (last gold) but we've selected the one with 0.646 that had more diversity in ensemble.",
      "votes": null
    },
    {
      "id": "1013427",
      "postDate": "09/16/2020 17:29:47",
      "content": "<p><a href=\"https://www.kaggle.com/watzisname\" target=\"_blank\">@watzisname</a> On model training with cross-validation you can keep OOF (out-of-fold) predictions for each validation fold (nothing new here). You can store OOFs based on a sliding window (e.g. width=5s, step=1s) over each validation audio file. Do it for several models (with good LB) and you have, for each time slice of each audio, the probabilities of the ground truth. Average them and keep the best slices for each audio (e.g. probs &gt; 0.9 or 0.7). That's way you've generated a subset with cleaned labels to you can use in a next step. That is what we called hot time slices.</p>",
      "rawMarkdown": "watzisname On model training with cross-validation you can keep OOF (out-of-fold) predictions for each validation fold (nothing new here). You can store OOFs based on a sliding window (e.g. width=5s, step=1s) over each validation audio file. Do it for several models (with good LB) and you have, for each time slice of each audio, the probabilities of the ground truth. Average them and keep the best slices for each audio (e.g. probs > 0.9 or 0.7). That's way you've generated a subset with cleaned labels to you can use in a next step. That is what we called hot time slices.",
      "votes": null
    },
    {
      "id": "1013435",
      "postDate": "09/16/2020 17:33:30",
      "content": "<p>ok, since you saved the hot slices for only the OOF - that means for Stage 2 you have a limited amount of data (for 5 fold, only 1/5th of data). Is stage 2 effectively fine-tuning in that case ?</p>",
      "rawMarkdown": "ok, since you saved the hot slices for only the OOF - that means for Stage 2 you have a limited amount of data (for 5 fold, only 1/5th of data). Is stage 2 effectively fine-tuning in that case ?",
      "votes": null
    },
    {
      "id": "1013441",
      "postDate": "09/16/2020 17:36:05",
      "content": "<p>For CV5 you have 5 OOF so at the end you've covered the full dataset. But that's true that stage2 dataset is smaller than stage1. For us, it reduced from 40k audio files to 32k audio files because we had 8k files with bad predictions.</p>",
      "rawMarkdown": "For CV5 you have 5 OOF so at the end you've covered the full dataset. But that's true that stage2 dataset is smaller than stage1. For us, it reduced from 40k audio files to 32k audio files because we had 8k files with bad predictions.",
      "votes": null
    },
    {
      "id": "1013445",
      "postDate": "09/16/2020 17:37:36",
      "content": "<p>Thanks, understood - thats a great idea indeed. </p>",
      "rawMarkdown": "Thanks, understood - thats a great idea indeed.",
      "votes": null
    },
    {
      "id": "1013451",
      "postDate": "09/16/2020 17:42:25",
      "content": "<p>Good labels, good models 😏. There are threads in the forum from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> about attempts/approaches to find such good time slices.</p>",
      "rawMarkdown": "Good labels, good models 😏. There are threads in the forum from @hengck23 about attempts/approaches to find such good time slices.",
      "votes": null
    },
    {
      "id": "1013732",
      "postDate": "09/16/2020 23:03:36",
      "content": "<p>Thanks for sharing, finding hot slice sounds very nice!</p>",
      "rawMarkdown": "Thanks for sharing, finding hot slice sounds very nice!",
      "votes": null
    },
    {
      "id": "1013881",
      "postDate": "09/17/2020 03:29:12",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> </p>\n<p>hope for best…. there is always second chance…. good luck!</p>",
      "rawMarkdown": "mpware \n\nhope for best.... there is always second chance.... good luck!",
      "votes": null
    },
    {
      "id": "1025198",
      "postDate": "09/24/2020 11:58:11",
      "content": "<p>Thanks for sharing and congrats!</p>\n<blockquote>\n  <p>Wavegram + MEL-Spectrogram model (was bad with soundscape)<br>\n  Create large image with 3x2 grid spectrogram (was bad on inference)</p>\n</blockquote>\n<p>About these two points, could you talk more details like how to fuse waveform and spec, and what is <code>3x2 grid spectrogram</code>. </p>\n<p>Since they look like very useful, though they didn't work this time.</p>",
      "rawMarkdown": "Thanks for sharing and congrats!\n\n> Wavegram + MEL-Spectrogram model (was bad with soundscape)\n> Create large image with 3x2 grid spectrogram (was bad on inference)\n\nAbout these two points, could you talk more details like how to fuse waveform and spec, and what is `3x2 grid spectrogram`. \n\nSince they look like very useful, though they didn't work this time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1013387,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "09/16/2020 17:14:01",
      "content": "<p>Great…. Just one step away 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1013394,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "09/16/2020 17:16:54",
          "content": "<p>Thanks, we had our best public LB submission with 0.647 on private LB (last gold) but we've selected the one with 0.646 that had more diversity in ensemble.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013881,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/17/2020 03:29:12",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> </p>\n<p>hope for best…. there is always second chance…. good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013388,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "09/16/2020 17:14:44",
      "content": "<p>Thanks for sharing now and during discussions during the competition too.<br>\nIf possible, can you elaborate a bit more on this point please…</p>\n<blockquote>\n  <p>ensemble all OOF models results to generate \"hot\" slices with high probabilities.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1013427,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "09/16/2020 17:29:47",
          "content": "<p><a href=\"https://www.kaggle.com/watzisname\" target=\"_blank\">@watzisname</a> On model training with cross-validation you can keep OOF (out-of-fold) predictions for each validation fold (nothing new here). You can store OOFs based on a sliding window (e.g. width=5s, step=1s) over each validation audio file. Do it for several models (with good LB) and you have, for each time slice of each audio, the probabilities of the ground truth. Average them and keep the best slices for each audio (e.g. probs &gt; 0.9 or 0.7). That's way you've generated a subset with cleaned labels to you can use in a next step. That is what we called hot time slices.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013435,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "09/16/2020 17:33:30",
          "content": "<p>ok, since you saved the hot slices for only the OOF - that means for Stage 2 you have a limited amount of data (for 5 fold, only 1/5th of data). Is stage 2 effectively fine-tuning in that case ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013441,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "09/16/2020 17:36:05",
          "content": "<p>For CV5 you have 5 OOF so at the end you've covered the full dataset. But that's true that stage2 dataset is smaller than stage1. For us, it reduced from 40k audio files to 32k audio files because we had 8k files with bad predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013445,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "09/16/2020 17:37:36",
          "content": "<p>Thanks, understood - thats a great idea indeed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013451,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "09/16/2020 17:42:25",
          "content": "<p>Good labels, good models 😏. There are threads in the forum from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> about attempts/approaches to find such good time slices.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013732,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "09/16/2020 23:03:36",
      "content": "<p>Thanks for sharing, finding hot slice sounds very nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1025198,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "09/24/2020 11:58:11",
      "content": "<p>Thanks for sharing and congrats!</p>\n<blockquote>\n  <p>Wavegram + MEL-Spectrogram model (was bad with soundscape)<br>\n  Create large image with 3x2 grid spectrogram (was bad on inference)</p>\n</blockquote>\n<p>About these two points, could you talk more details like how to fuse waveform and spec, and what is <code>3x2 grid spectrogram</code>. </p>\n<p>Since they look like very useful, though they didn't work this time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1013357": "It was my first audio competition and I really enjoyed it! Thanks to organizers and Kaggle for it.\nThanks to my teammates @tikutiku @zfturbo the great collaboration, it was a pleasure and I've learnt a lot again.\n\nHere is a digest of our solution.\n\n**Main Pipeline** (see image below):\n- MEL-Spectrogram (torchaudio) with differents settings to capture most birds frequency shapes.\n- Most processing/augmentation occurs in GPU to speed up training and inference.\n\n**Time Augmentations**:\n- Pink Noise/White noise\n- Time roll\n- Volume gain\n- PitchShift\n- Additional bandpass filters,lowcut (1kHz-2.5kHz), highcut (10kHz-15kHz)\n- Background/Ambient noise Mixup\n- Other birds Mixup (2 to 3 birds)\n\n**Spec/Image augmentations**:\n- Frequency/time masking\n- Color jitter\n\n**CNN backbones**:\n- EfficientNet B1 and B2\n- SEReseXt26\n- ResneSt50\n- Optional GRU layer to learn about time sequences (see [TALNet Sound Event Classifier](https://github.com/srvk/TALNet))\n- Different pooling to apply SED concept with either Attention block or max pooling\nWe also had one model based on 1D signal only (DensetNet1D)\n\nWhat did not work well (it worked but was disapointing):\n- Wavegram + MEL-Spectrogram model (was bad with soundscape)\n- Create large image with 3x2 grid spectrogram (was bad on inference)\n\n**Data**:\n- Train audio provided with duration outliers removed\n- Additional xeno-canto data (from Vopani dataset, thanks @rohanrao for your clean dump)\n- Distractors (NoBird/Nocall/Ambient) 10s slices (from freefield1010)\nWe've built our own validation data by mixing test_audio/birds/noise/nocall to try to correlate LB. It correlated a bit but was not enough to give trust in it. Too bad as it was key for this competition.\n\n**Training procedure**:\n- Stage1: Train a few models with 5s slices picked randomly, then save birds probabilities on CV OOF, ensemble all OOF models results to generate \"*hot*\" slices with high probabilities.\n  It allowed to reach public LB=0.582\n- Stage2: Train more models with only such 5s *hot slices*.\n  It allowed to reach public LB=0.596\nWhile reading other's solutions, we should have tried another stage with \"very hot\" slices to have a super clean train dataset.\n\n**Final ensemble**:\nEnsemble is a combination of models (8 to 10) with union strategy and with per-model threshold tuned on public LB (it overfitted for sure).\nIdea of union was to capture most of the TP (better for metric used in this competition), drawback is that we capture FP too.\nInference was fast, we cached resampled audio in memory, and (almost) full GPU pipeline helped. It ran in around 1h30 so we still had room for more models.\n\n**Post-processing**:\nWe have a \"nobird\" majority vote to try to remove FP (force nocall) as some models have been trained with 265 classes instead of 264 (class 265 = nobird)\n\nBoth our final submissions reached similar score on private LB.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F4e2db4b7460eb1a51737fbf81bd8ed50%2Fpipeline.png?generation=1600274872405000&alt=media)\n\nFinal words to conclude: Congratulations to top teams! and thanks to @hidehisaarai1213, @hengck23 for their sharings.",
    "1013387": "Great.... Just one step away 👍",
    "1013388": "Thanks for sharing now and during discussions during the competition too.\nIf possible, can you elaborate a bit more on this point please...\n>ensemble all OOF models results to generate \"hot\" slices with high probabilities.",
    "1013394": "Thanks, we had our best public LB submission with 0.647 on private LB (last gold) but we've selected the one with 0.646 that had more diversity in ensemble.",
    "1013427": "watzisname On model training with cross-validation you can keep OOF (out-of-fold) predictions for each validation fold (nothing new here). You can store OOFs based on a sliding window (e.g. width=5s, step=1s) over each validation audio file. Do it for several models (with good LB) and you have, for each time slice of each audio, the probabilities of the ground truth. Average them and keep the best slices for each audio (e.g. probs > 0.9 or 0.7). That's way you've generated a subset with cleaned labels to you can use in a next step. That is what we called hot time slices.",
    "1013435": "ok, since you saved the hot slices for only the OOF - that means for Stage 2 you have a limited amount of data (for 5 fold, only 1/5th of data). Is stage 2 effectively fine-tuning in that case ?",
    "1013441": "For CV5 you have 5 OOF so at the end you've covered the full dataset. But that's true that stage2 dataset is smaller than stage1. For us, it reduced from 40k audio files to 32k audio files because we had 8k files with bad predictions.",
    "1013445": "Thanks, understood - thats a great idea indeed.",
    "1013451": "Good labels, good models 😏. There are threads in the forum from @hengck23 about attempts/approaches to find such good time slices.",
    "1013732": "Thanks for sharing, finding hot slice sounds very nice!",
    "1013881": "mpware \n\nhope for best.... there is always second chance.... good luck!",
    "1025198": "Thanks for sharing and congrats!\n\n> Wavegram + MEL-Spectrogram model (was bad with soundscape)\n> Create large image with 3x2 grid spectrogram (was bad on inference)\n\nAbout these two points, could you talk more details like how to fuse waveform and spec, and what is `3x2 grid spectrogram`. \n\nSince they look like very useful, though they didn't work this time."
  },
  "source": "meta"
}