{
  "id": 412996,
  "title": "24th place solution - pre-training & single model (5 folds ensemble with ONNX)",
  "url": "/competitions/birdclef-2023/discussion/412996",
  "author_name": "HyeongChan Kim",
  "post_date": "2023-05-26T09:17:56.815000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p>First, thanks to the Cornell Lab of Ornithology and the Kaggle team for hosting the competition. Also, congratulations to all winners and participants.</p>\n<h2>TL;DR</h2>\n<p>I joined about 2 weeks before the competition ended, and I had about 7 days to participate due to my busy daily life. In the meantime, It's a great opportunity for me to challenge myself to build a model in a short time. So, I had to focus on which recipes would work best within minimum trials.</p>\n<p>Thankfully, There're many good quality resources like top solutions from the last competitions. I started to follow up on the solution and figured out the working recipes in general. (past BirdCLEF 2020 experience helps a lot too)</p>\n<p>As a result of the consequence and some lucks, I can build a decent model I guess : )</p>\n<h2>Architecture</h2>\n<p>Here's the pipeline.</p>\n<ol>\n<li>pre-train on 2020, 2021, 2022, xeno-canto datasets.</li>\n<li>fine-tune on 2023 dataset (based on the pre-trained weight).<ul>\n<li>minor classes (&lt;= 5 samples) are included in all folds</li></ul></li>\n</ol>\n<p>I applied the same training recipes (e.g. augmentation, loss functions, …) each step.</p>\n<h3>CV</h3>\n<p>(although based on my few experiments) my cv score and LB/PB are kinda correlated.</p>\n<table>\n<thead>\n<tr>\n<th>Exp</th>\n<th>CV</th>\n<th>LB</th>\n<th>PB</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>effnetb0</code></td>\n<td>0.7720</td>\n<td>0.82438</td>\n<td>0.73641</td>\n<td>multiple losses, 5 folds</td>\n</tr>\n<tr>\n<td><code>effnetb0</code></td>\n<td>0.7693</td>\n<td>0.82402</td>\n<td>0.73604</td>\n<td>clipwise loss, 5 folds</td>\n</tr>\n<tr>\n<td><code>eca_nfnet_l0</code></td>\n<td>0.7753</td>\n<td>0.80731</td>\n<td>0.71845</td>\n<td>clipwise loss, single fold</td>\n</tr>\n</tbody>\n</table>\n<h3>Model</h3>\n<p>I used SED architecture with the <code>efficientnet_b0</code> backbone. Also, I tested <code>eca_nfnet_l0</code> backbone, and it has a better cv score, but I can't use it due to the latency.</p>\n<h3>Training recipe</h3>\n<ul>\n<li>[<strong>Important</strong>] pre-training</li>\n<li>[<strong>Important</strong>] augmentations<ul>\n<li>waveform-level<ul>\n<li>[Important] or mixup on a raw waveform</li>\n<li>gaussian &amp; uniform noise</li>\n<li>pitch shift</li>\n<li>[Important] background noise</li></ul></li>\n<li>spectrogram-level<ul>\n<li>spec augment</li></ul></li></ul></li>\n<li>log-mel spectrogram<ul>\n<li>n_fft &amp; window size 1024, hop size 320, min/max freq 20/14000, num_mels 256, top_db 80. (actually, I wanted n_fft with 2048, but I set it to 1024 by my mistake)</li></ul></li>\n<li>trained on 5 secs clips</li>\n<li>stratified k fold (5 folds, on primary_label)</li>\n<li>label smoothing 0.1</li>\n<li>multiple losses (from <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">birdcelf 2021 top 5</a>)<ul>\n<li>bce loss on clip-wise output w/ weight 1.0</li>\n<li>bce loss on max of segment-wise outputs w/ weight 0.5</li></ul></li>\n<li>fp32</li>\n<li>AdamW + cosine annealing (w/o warmup)<ul>\n<li>50 epochs (usually converged between 40 ~ 50)</li></ul></li>\n</ul>\n<h2>Inference</h2>\n<p>I can ensemble up to 4 models with Pytorch (it took nearly 2 hrs). To mix more models, I utilized ONNX and did graph optimization, and it makes one more model to be ensembled! Finally, I can ensemble 5 models (single model 5 folds). Also, to utilize the full CPU, I do some multi-processing stuff.</p>\n<h2>Not worked (perhaps I might be wrong)</h2>\n<ul>\n<li>secondary label (both hard label, soft label (e.g. 0.3, 0.5))</li>\n<li>focal loss</li>\n<li>longer clips (e.g. 15s)</li>\n<li>post-processings (proposed in the BirdCLEF 2021, and 2022 competitions)<ul>\n<li>aggregate the probs of the previous and next segments.</li>\n<li>if there's a bird above the threshold, multiply constants on all segments of the bird.)</li></ul></li>\n</ul>\n<p>I hope this could help!</p>\n<p>Thanks : )</p>",
  "messages": [
    {
      "id": 2274824,
      "postDate": "2023-05-26T09:17:56.817Z",
      "content": "<p>Hello everyone!</p>\n<p>First, thanks to the Cornell Lab of Ornithology and the Kaggle team for hosting the competition. Also, congratulations to all winners and participants.</p>\n<h2>TL;DR</h2>\n<p>I joined about 2 weeks before the competition ended, and I had about 7 days to participate due to my busy daily life. In the meantime, It's a great opportunity for me to challenge myself to build a model in a short time. So, I had to focus on which recipes would work best within minimum trials.</p>\n<p>Thankfully, There're many good quality resources like top solutions from the last competitions. I started to follow up on the solution and figured out the working recipes in general. (past BirdCLEF 2020 experience helps a lot too)</p>\n<p>As a result of the consequence and some lucks, I can build a decent model I guess : )</p>\n<h2>Architecture</h2>\n<p>Here's the pipeline.</p>\n<ol>\n<li>pre-train on 2020, 2021, 2022, xeno-canto datasets.</li>\n<li>fine-tune on 2023 dataset (based on the pre-trained weight).<ul>\n<li>minor classes (&lt;= 5 samples) are included in all folds</li></ul></li>\n</ol>\n<p>I applied the same training recipes (e.g. augmentation, loss functions, …) each step.</p>\n<h3>CV</h3>\n<p>(although based on my few experiments) my cv score and LB/PB are kinda correlated.</p>\n<table>\n<thead>\n<tr>\n<th>Exp</th>\n<th>CV</th>\n<th>LB</th>\n<th>PB</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>effnetb0</code></td>\n<td>0.7720</td>\n<td>0.82438</td>\n<td>0.73641</td>\n<td>multiple losses, 5 folds</td>\n</tr>\n<tr>\n<td><code>effnetb0</code></td>\n<td>0.7693</td>\n<td>0.82402</td>\n<td>0.73604</td>\n<td>clipwise loss, 5 folds</td>\n</tr>\n<tr>\n<td><code>eca_nfnet_l0</code></td>\n<td>0.7753</td>\n<td>0.80731</td>\n<td>0.71845</td>\n<td>clipwise loss, single fold</td>\n</tr>\n</tbody>\n</table>\n<h3>Model</h3>\n<p>I used SED architecture with the <code>efficientnet_b0</code> backbone. Also, I tested <code>eca_nfnet_l0</code> backbone, and it has a better cv score, but I can't use it due to the latency.</p>\n<h3>Training recipe</h3>\n<ul>\n<li>[<strong>Important</strong>] pre-training</li>\n<li>[<strong>Important</strong>] augmentations<ul>\n<li>waveform-level<ul>\n<li>[Important] or mixup on a raw waveform</li>\n<li>gaussian &amp; uniform noise</li>\n<li>pitch shift</li>\n<li>[Important] background noise</li></ul></li>\n<li>spectrogram-level<ul>\n<li>spec augment</li></ul></li></ul></li>\n<li>log-mel spectrogram<ul>\n<li>n_fft &amp; window size 1024, hop size 320, min/max freq 20/14000, num_mels 256, top_db 80. (actually, I wanted n_fft with 2048, but I set it to 1024 by my mistake)</li></ul></li>\n<li>trained on 5 secs clips</li>\n<li>stratified k fold (5 folds, on primary_label)</li>\n<li>label smoothing 0.1</li>\n<li>multiple losses (from <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">birdcelf 2021 top 5</a>)<ul>\n<li>bce loss on clip-wise output w/ weight 1.0</li>\n<li>bce loss on max of segment-wise outputs w/ weight 0.5</li></ul></li>\n<li>fp32</li>\n<li>AdamW + cosine annealing (w/o warmup)<ul>\n<li>50 epochs (usually converged between 40 ~ 50)</li></ul></li>\n</ul>\n<h2>Inference</h2>\n<p>I can ensemble up to 4 models with Pytorch (it took nearly 2 hrs). To mix more models, I utilized ONNX and did graph optimization, and it makes one more model to be ensembled! Finally, I can ensemble 5 models (single model 5 folds). Also, to utilize the full CPU, I do some multi-processing stuff.</p>\n<h2>Not worked (perhaps I might be wrong)</h2>\n<ul>\n<li>secondary label (both hard label, soft label (e.g. 0.3, 0.5))</li>\n<li>focal loss</li>\n<li>longer clips (e.g. 15s)</li>\n<li>post-processings (proposed in the BirdCLEF 2021, and 2022 competitions)<ul>\n<li>aggregate the probs of the previous and next segments.</li>\n<li>if there's a bird above the threshold, multiply constants on all segments of the bird.)</li></ul></li>\n</ul>\n<p>I hope this could help!</p>\n<p>Thanks : )</p>",
      "rawMarkdown": "Hello everyone!\n\nFirst, thanks to the Cornell Lab of Ornithology and the Kaggle team for hosting the competition. Also, congratulations to all winners and participants.\n\n## TL;DR\n\nI joined about 2 weeks before the competition ended, and I had about 7 days to participate due to my busy daily life. In the meantime, It's a great opportunity for me to challenge myself to build a model in a short time. So, I had to focus on which recipes would work best within minimum trials.\n\nThankfully, There're many good quality resources like top solutions from the last competitions. I started to follow up on the solution and figured out the working recipes in general. (past BirdCLEF 2020 experience helps a lot too)\n\nAs a result of the consequence and some lucks, I can build a decent model I guess : )\n\n## Architecture\n\nHere's the pipeline.\n\n1. pre-train on 2020, 2021, 2022, xeno-canto datasets.\n2. fine-tune on 2023 dataset (based on the pre-trained weight).\n    * minor classes (<= 5 samples) are included in all folds\n\nI applied the same training recipes (e.g. augmentation, loss functions, ...) each step.\n\n### CV\n\n(although based on my few experiments) my cv score and LB/PB are kinda correlated.\n\n| Exp | CV | LB | PB | Note |\n| :---: | :---:| :---: | :---: | :---: |\n| `effnetb0` | 0.7720 | 0.82438 | 0.73641 | multiple losses, 5 folds |\n| `effnetb0` | 0.7693 | 0.82402 | 0.73604 | clipwise loss, 5 folds |\n| `eca_nfnet_l0` | 0.7753 | 0.80731 | 0.71845 | clipwise loss, single fold |\n\n### Model\n\nI used SED architecture with the `efficientnet_b0` backbone. Also, I tested `eca_nfnet_l0` backbone, and it has a better cv score, but I can't use it due to the latency.\n\n### Training recipe\n\n* [**Important**] pre-training\n* [**Important**] augmentations\n    * waveform-level\n        * [Important] or mixup on a raw waveform\n        * gaussian & uniform noise\n        * pitch shift\n        * [Important] background noise\n    * spectrogram-level\n        * spec augment\n* log-mel spectrogram\n    * n_fft & window size 1024, hop size 320, min/max freq 20/14000, num_mels 256, top_db 80. (actually, I wanted n_fft with 2048, but I set it to 1024 by my mistake)\n* trained on 5 secs clips\n* stratified k fold (5 folds, on primary_label)\n* label smoothing 0.1\n* multiple losses (from [birdcelf 2021 top 5](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351))\n  * bce loss on clip-wise output w/ weight 1.0\n  * bce loss on max of segment-wise outputs w/ weight 0.5\n* fp32\n* AdamW + cosine annealing (w/o warmup)\n  * 50 epochs (usually converged between 40 ~ 50)\n\n## Inference\n\nI can ensemble up to 4 models with Pytorch (it took nearly 2 hrs). To mix more models, I utilized ONNX and did graph optimization, and it makes one more model to be ensembled! Finally, I can ensemble 5 models (single model 5 folds). Also, to utilize the full CPU, I do some multi-processing stuff.\n\n## Not worked (perhaps I might be wrong)\n\n* secondary label (both hard label, soft label (e.g. 0.3, 0.5))\n* focal loss\n* longer clips (e.g. 15s)\n* post-processings (proposed in the BirdCLEF 2021, and 2022 competitions)\n  * aggregate the probs of the previous and next segments.\n  * if there's a bird above the threshold, multiply constants on all segments of the bird.)\n\nI hope this could help!\n\nThanks : )",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2274824": "Hello everyone!\n\nFirst, thanks to the Cornell Lab of Ornithology and the Kaggle team for hosting the competition. Also, congratulations to all winners and participants.\n\n## TL;DR\n\nI joined about 2 weeks before the competition ended, and I had about 7 days to participate due to my busy daily life. In the meantime, It's a great opportunity for me to challenge myself to build a model in a short time. So, I had to focus on which recipes would work best within minimum trials.\n\nThankfully, There're many good quality resources like top solutions from the last competitions. I started to follow up on the solution and figured out the working recipes in general. (past BirdCLEF 2020 experience helps a lot too)\n\nAs a result of the consequence and some lucks, I can build a decent model I guess : )\n\n## Architecture\n\nHere's the pipeline.\n\n1. pre-train on 2020, 2021, 2022, xeno-canto datasets.\n2. fine-tune on 2023 dataset (based on the pre-trained weight).\n    * minor classes (<= 5 samples) are included in all folds\n\nI applied the same training recipes (e.g. augmentation, loss functions, ...) each step.\n\n### CV\n\n(although based on my few experiments) my cv score and LB/PB are kinda correlated.\n\n| Exp | CV | LB | PB | Note |\n| :---: | :---:| :---: | :---: | :---: |\n| `effnetb0` | 0.7720 | 0.82438 | 0.73641 | multiple losses, 5 folds |\n| `effnetb0` | 0.7693 | 0.82402 | 0.73604 | clipwise loss, 5 folds |\n| `eca_nfnet_l0` | 0.7753 | 0.80731 | 0.71845 | clipwise loss, single fold |\n\n### Model\n\nI used SED architecture with the `efficientnet_b0` backbone. Also, I tested `eca_nfnet_l0` backbone, and it has a better cv score, but I can't use it due to the latency.\n\n### Training recipe\n\n* [**Important**] pre-training\n* [**Important**] augmentations\n    * waveform-level\n        * [Important] or mixup on a raw waveform\n        * gaussian & uniform noise\n        * pitch shift\n        * [Important] background noise\n    * spectrogram-level\n        * spec augment\n* log-mel spectrogram\n    * n_fft & window size 1024, hop size 320, min/max freq 20/14000, num_mels 256, top_db 80. (actually, I wanted n_fft with 2048, but I set it to 1024 by my mistake)\n* trained on 5 secs clips\n* stratified k fold (5 folds, on primary_label)\n* label smoothing 0.1\n* multiple losses (from [birdcelf 2021 top 5](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351))\n  * bce loss on clip-wise output w/ weight 1.0\n  * bce loss on max of segment-wise outputs w/ weight 0.5\n* fp32\n* AdamW + cosine annealing (w/o warmup)\n  * 50 epochs (usually converged between 40 ~ 50)\n\n## Inference\n\nI can ensemble up to 4 models with Pytorch (it took nearly 2 hrs). To mix more models, I utilized ONNX and did graph optimization, and it makes one more model to be ensembled! Finally, I can ensemble 5 models (single model 5 folds). Also, to utilize the full CPU, I do some multi-processing stuff.\n\n## Not worked (perhaps I might be wrong)\n\n* secondary label (both hard label, soft label (e.g. 0.3, 0.5))\n* focal loss\n* longer clips (e.g. 15s)\n* post-processings (proposed in the BirdCLEF 2021, and 2022 competitions)\n  * aggregate the probs of the previous and next segments.\n  * if there's a bird above the threshold, multiply constants on all segments of the bird.)\n\nI hope this could help!\n\nThanks : )"
  }
}