{
  "id": 412753,
  "title": "4th Place Solution: Knowledge Distillation Is All You Need",
  "url": "/competitions/birdclef-2023/writeups/atfujita-4th-place-solution-knowledge-distillation",
  "author_name": "",
  "post_date": "2023-06-06T00:28:39.770Z",
  "votes": 55,
  "comment_count": 25,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition. And I would like to emphasize the thanks to the bird sound recordists who provided their data through xeno-canto.</p>\n<h3>Solution Summary</h3>\n<ul>\n<li>Knowledge Distillation is all you need.</li>\n<li>Adding no-call data, xeno-canto data, and background audios(Zenodo) is effective.</li>\n</ul>\n<h3>Datasets</h3>\n<p>Additional datasets can be found <a href=\"https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-additional\" target=\"_blank\">here</a>.</p>\n<ul>\n<li>Bird CLEF 2023</li>\n<li>Bird CLEF 2021, 2022 (for pretraining)</li>\n<li>ff1010bird_nocall (5,755 files for learning no call)</li>\n<li>xeno-canto files not included in training dataset CC-BY-NC-SA (896 files) and CC-BY-NC-ND (5,212 files).</li>\n<li>Zenodo dataset (background noise)</li>\n<li>esc50 (rain, frog)</li>\n<li>aicrowd2020_noise_30sec (background noise for pretraining)</li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li>My models are based on <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">the BirdCLEF 2021 2nd place solution</a> using <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a>.</li>\n<li>My solution consists of a total of 4 models, all with eca_nfnet_l0 as the backbone. Each set is slightly different.</li>\n</ul>\n<ol>\n<li>MelSpectrogram (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512, top_db=80.0, NormalizeMelSpec) following <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">this solution</a>.<ul>\n<li>Public: 0.8312, Private: 0.74424</li></ul></li>\n<li>Balanced sampling with the above settings.<ul>\n<li>Public: 0.83106, Private: 0.74406</li></ul></li>\n<li><a href=\"https://github.com/daemon/pytorch-pcen\" target=\"_blank\">PCEN</a> (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512).<ul>\n<li>Public: 0.83005, Private: 0.74134</li></ul></li>\n<li>MelSpectrogram (sample_rate: 32000, mel_bins: 64, fmin: 50, fmax: 14000, window_size: 1024, hop_size: 320).<ul>\n<li>Public: 0.83014, Private: 0.74201</li></ul></li>\n</ol>\n<h3>Training</h3>\n<p>Knowledge distillation of pre-computed predictions in <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1\" target=\"_blank\">Kaggle Models</a> (bird-vocalization-classifier) was the hallmark of my solution. This model cannot complete inference in less than 2h, but it is very powerful. cmAP_5 on my validation dataset was 0.9479. So I tried to extract useful information from this model.<br>\n<strong>The Kaggle models pre-computed predictions were created <a href=\"https://www.kaggle.com/code/atsunorifujita/extract-from-kaggle-models/notebook\" target=\"_blank\">here</a></strong>.<br>\nAccording to this <a href=\"https://arxiv.org/abs/2106.05237\" target=\"_blank\">paper</a>, they argued that efficient distillation required 1. consistent input, 2. aggressive mixup, and 3. a large number of epochs. So I did 2 and 3 because it was challenging to integrate the Kaggle model into my training pipeline. It certainly looked effective.<br>\nI chose only approaches that improve both CV and LB. In this competition, I just couldn't believe in one or the other.</p>\n<ul>\n<li>Using 5 StratifiedKFold. Only 1 fold did not cover all classes, so the remaining 4 were used for training.</li>\n<li>The evaluation metric is padded_cmap1. The best results were the same with padded_cmap5 in most cases.</li>\n<li>Use primary_label only</li>\n<li>no-call is represented by all 0</li>\n<li>loss: 0.1 * BCEWithLogitsLoss (primary_label) + 0.9 * KLDivLoss (from Kaggle model)<ul>\n<li>Softmax temperature=20 was best (tried 5, 10, 20, 30).</li></ul></li>\n<li>I used randomly sampled 20-sec clips for training. Audio that is less than 20 sec is repeated.</li>\n<li>epoch = 400 (Most models converge at 100-300)</li>\n<li>early stopping(pretraining=10, training=20)</li>\n<li>Optimizer: AdamW (lr: pretraining=5e-4, training=2.5e-4, wd: 1e-6)</li>\n<li>CosineLRScheduler(t_initial=10, warmup_t=1, cycle_limit=40, cycle_decay=1.0, lr_min=1e-7, t_in_epochs=True,)</li>\n<li>mixup p = 1.0  (It was better than p=0.5)<br></li>\n</ul>\n<p>I also encountered a situation where the training time was significantly longer when pretrained with past competition data. Thanks to everyone who suggested solutions.</p>\n<h4>Augmentation</h4>\n<ul>\n<li>OneOf ([Gain, GainTransition])</li>\n<li>OneOf ([AddGaussianNoise, AddGaussianSNR]</li>\n<li>AddShortNoises esc50 (rain, frog)</li>\n<li>AddBackgroundNoise from <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605\" target=\"_blank\">Zenodo</a>. The 60 minutes with the fewest bird calls were extracted from each dataset and divided into 30 sec (training only).</li>\n<li>AddBackgroundNoise from aicrowd2020_noise_30sec and ff1010bird_nocall (pretraining only).</li>\n<li>LowPassFilter</li>\n<li>PitchShift</li>\n</ul>\n<h4>Hardware</h4>\n<ul>\n<li>1 * RTX 3090</li>\n<li>Pretraining time: 30-40h per model. </li>\n<li>Training time: 9-12h per model.</li>\n</ul>\n<h3>Inference</h3>\n<p>Reducing the inference time is the part I was having trouble with. Thanks to <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> for sharing his <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">effective approach</a>, I was able to use 4 models.</p>\n<ul>\n<li>Ensemble with simple averaging.</li>\n<li>4 models using PyTorch JIT. total 110min.</li>\n<li><strong><a href=\"https://www.kaggle.com/code/atsunorifujita/4th-place-solution-inference-kernel\" target=\"_blank\">Inference notebook</a></strong></li>\n</ul>\n<h3>What didn’t work</h3>\n<ul>\n<li>focal loss.</li>\n<li>Split a low-sample class.</li>\n<li>Backbone other than eca_nfnet_l0 and eca_nfnet_l1.</li>\n<li>Optimizers (adan, lion, ranger21, shampoo). I tried to create a custom normalize-free model but failed.</li>\n<li><a href=\"https://github.com/naver-ai/cmo\" target=\"_blank\">CMO</a> (mixup worked better when using distillation).</li>\n<li><a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/features/cqt.py\" target=\"_blank\">CQT</a> (slow and degraded).</li>\n<li>Change to first stride(1, 1). CV was good but inference takes a long time and not a single model is completed in less than 2h.</li>\n<li>Pretraining Zenodo data (story before trying distillation).</li>\n<li>Distillation training from scratch (not use imagenet weights). It didn't converge in the amount of time I could tolerate.</li>\n<li>MelSpectrogram, PCEN, and CQT integrated into input channels in all combinations. There was no synergistic effect.</li>\n</ul>\n<h3>Ablation study</h3>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BaseModel</td>\n<td>0.80603</td>\n<td>0.70782</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation</td>\n<td>0.82073</td>\n<td>0.72752</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation + Adding xeno-canto</td>\n<td>0.82905</td>\n<td>0.74038</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining</td>\n<td>0.8312</td>\n<td>0.74424</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining + Ensemble (4 models)</td>\n<td>0.84019</td>\n<td>0.75688</td>\n</tr>\n</tbody>\n</table>\n<h3>Acknowledgments</h3>\n<p>My solution builds on contributions from participants in this and past competitions, birding enthusiasts, and the work of the machine learning community. thank you very much 🙏 <br>\n<strong>Training Code</strong>: <a href=\"https://github.com/AtsunoriFujita/BirdCLEF-2023-Identify-bird-calls-in-soundscapes\" target=\"_blank\">https://github.com/AtsunoriFujita/BirdCLEF-2023-Identify-bird-calls-in-soundscapes</a></p>",
  "messages": [
    {
      "id": "2273202",
      "postDate": "05/25/2023 04:31:14",
      "content": "<p>First of all, thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition. And I would like to emphasize the thanks to the bird sound recordists who provided their data through xeno-canto.</p>\n<h3>Solution Summary</h3>\n<ul>\n<li>Knowledge Distillation is all you need.</li>\n<li>Adding no-call data, xeno-canto data, and background audios(Zenodo) is effective.</li>\n</ul>\n<h3>Datasets</h3>\n<p>Additional datasets can be found <a href=\"https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-additional\" target=\"_blank\">here</a>.</p>\n<ul>\n<li>Bird CLEF 2023</li>\n<li>Bird CLEF 2021, 2022 (for pretraining)</li>\n<li>ff1010bird_nocall (5,755 files for learning no call)</li>\n<li>xeno-canto files not included in training dataset CC-BY-NC-SA (896 files) and CC-BY-NC-ND (5,212 files).</li>\n<li>Zenodo dataset (background noise)</li>\n<li>esc50 (rain, frog)</li>\n<li>aicrowd2020_noise_30sec (background noise for pretraining)</li>\n</ul>\n<h3>Models</h3>\n<ul>\n<li>My models are based on <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">the BirdCLEF 2021 2nd place solution</a> using <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a>.</li>\n<li>My solution consists of a total of 4 models, all with eca_nfnet_l0 as the backbone. Each set is slightly different.</li>\n</ul>\n<ol>\n<li>MelSpectrogram (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512, top_db=80.0, NormalizeMelSpec) following <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243293\" target=\"_blank\">this solution</a>.<ul>\n<li>Public: 0.8312, Private: 0.74424</li></ul></li>\n<li>Balanced sampling with the above settings.<ul>\n<li>Public: 0.83106, Private: 0.74406</li></ul></li>\n<li><a href=\"https://github.com/daemon/pytorch-pcen\" target=\"_blank\">PCEN</a> (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512).<ul>\n<li>Public: 0.83005, Private: 0.74134</li></ul></li>\n<li>MelSpectrogram (sample_rate: 32000, mel_bins: 64, fmin: 50, fmax: 14000, window_size: 1024, hop_size: 320).<ul>\n<li>Public: 0.83014, Private: 0.74201</li></ul></li>\n</ol>\n<h3>Training</h3>\n<p>Knowledge distillation of pre-computed predictions in <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1\" target=\"_blank\">Kaggle Models</a> (bird-vocalization-classifier) was the hallmark of my solution. This model cannot complete inference in less than 2h, but it is very powerful. cmAP_5 on my validation dataset was 0.9479. So I tried to extract useful information from this model.<br>\n<strong>The Kaggle models pre-computed predictions were created <a href=\"https://www.kaggle.com/code/atsunorifujita/extract-from-kaggle-models/notebook\" target=\"_blank\">here</a></strong>.<br>\nAccording to this <a href=\"https://arxiv.org/abs/2106.05237\" target=\"_blank\">paper</a>, they argued that efficient distillation required 1. consistent input, 2. aggressive mixup, and 3. a large number of epochs. So I did 2 and 3 because it was challenging to integrate the Kaggle model into my training pipeline. It certainly looked effective.<br>\nI chose only approaches that improve both CV and LB. In this competition, I just couldn't believe in one or the other.</p>\n<ul>\n<li>Using 5 StratifiedKFold. Only 1 fold did not cover all classes, so the remaining 4 were used for training.</li>\n<li>The evaluation metric is padded_cmap1. The best results were the same with padded_cmap5 in most cases.</li>\n<li>Use primary_label only</li>\n<li>no-call is represented by all 0</li>\n<li>loss: 0.1 * BCEWithLogitsLoss (primary_label) + 0.9 * KLDivLoss (from Kaggle model)<ul>\n<li>Softmax temperature=20 was best (tried 5, 10, 20, 30).</li></ul></li>\n<li>I used randomly sampled 20-sec clips for training. Audio that is less than 20 sec is repeated.</li>\n<li>epoch = 400 (Most models converge at 100-300)</li>\n<li>early stopping(pretraining=10, training=20)</li>\n<li>Optimizer: AdamW (lr: pretraining=5e-4, training=2.5e-4, wd: 1e-6)</li>\n<li>CosineLRScheduler(t_initial=10, warmup_t=1, cycle_limit=40, cycle_decay=1.0, lr_min=1e-7, t_in_epochs=True,)</li>\n<li>mixup p = 1.0  (It was better than p=0.5)<br></li>\n</ul>\n<p>I also encountered a situation where the training time was significantly longer when pretrained with past competition data. Thanks to everyone who suggested solutions.</p>\n<h4>Augmentation</h4>\n<ul>\n<li>OneOf ([Gain, GainTransition])</li>\n<li>OneOf ([AddGaussianNoise, AddGaussianSNR]</li>\n<li>AddShortNoises esc50 (rain, frog)</li>\n<li>AddBackgroundNoise from <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605\" target=\"_blank\">Zenodo</a>. The 60 minutes with the fewest bird calls were extracted from each dataset and divided into 30 sec (training only).</li>\n<li>AddBackgroundNoise from aicrowd2020_noise_30sec and ff1010bird_nocall (pretraining only).</li>\n<li>LowPassFilter</li>\n<li>PitchShift</li>\n</ul>\n<h4>Hardware</h4>\n<ul>\n<li>1 * RTX 3090</li>\n<li>Pretraining time: 30-40h per model. </li>\n<li>Training time: 9-12h per model.</li>\n</ul>\n<h3>Inference</h3>\n<p>Reducing the inference time is the part I was having trouble with. Thanks to <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> for sharing his <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">effective approach</a>, I was able to use 4 models.</p>\n<ul>\n<li>Ensemble with simple averaging.</li>\n<li>4 models using PyTorch JIT. total 110min.</li>\n<li><strong><a href=\"https://www.kaggle.com/code/atsunorifujita/4th-place-solution-inference-kernel\" target=\"_blank\">Inference notebook</a></strong></li>\n</ul>\n<h3>What didn’t work</h3>\n<ul>\n<li>focal loss.</li>\n<li>Split a low-sample class.</li>\n<li>Backbone other than eca_nfnet_l0 and eca_nfnet_l1.</li>\n<li>Optimizers (adan, lion, ranger21, shampoo). I tried to create a custom normalize-free model but failed.</li>\n<li><a href=\"https://github.com/naver-ai/cmo\" target=\"_blank\">CMO</a> (mixup worked better when using distillation).</li>\n<li><a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/features/cqt.py\" target=\"_blank\">CQT</a> (slow and degraded).</li>\n<li>Change to first stride(1, 1). CV was good but inference takes a long time and not a single model is completed in less than 2h.</li>\n<li>Pretraining Zenodo data (story before trying distillation).</li>\n<li>Distillation training from scratch (not use imagenet weights). It didn't converge in the amount of time I could tolerate.</li>\n<li>MelSpectrogram, PCEN, and CQT integrated into input channels in all combinations. There was no synergistic effect.</li>\n</ul>\n<h3>Ablation study</h3>\n<table>\n<thead>\n<tr>\n<th>Name</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BaseModel</td>\n<td>0.80603</td>\n<td>0.70782</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation</td>\n<td>0.82073</td>\n<td>0.72752</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation + Adding xeno-canto</td>\n<td>0.82905</td>\n<td>0.74038</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining</td>\n<td>0.8312</td>\n<td>0.74424</td>\n</tr>\n<tr>\n<td>BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining + Ensemble (4 models)</td>\n<td>0.84019</td>\n<td>0.75688</td>\n</tr>\n</tbody>\n</table>\n<h3>Acknowledgments</h3>\n<p>My solution builds on contributions from participants in this and past competitions, birding enthusiasts, and the work of the machine learning community. thank you very much 🙏 <br>\n<strong>Training Code</strong>: <a href=\"https://github.com/AtsunoriFujita/BirdCLEF-2023-Identify-bird-calls-in-soundscapes\" target=\"_blank\">https://github.com/AtsunoriFujita/BirdCLEF-2023-Identify-bird-calls-in-soundscapes</a></p>",
      "rawMarkdown": "First of all, thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition. And I would like to emphasize the thanks to the bird sound recordists who provided their data through xeno-canto.\n\n### Solution Summary\n- Knowledge Distillation is all you need.\n- Adding no-call data, xeno-canto data, and background audios(Zenodo) is effective.\n\n### Datasets\nAdditional datasets can be found [here](https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-additional).\n- Bird CLEF 2023\n- Bird CLEF 2021, 2022 (for pretraining)\n- ff1010bird_nocall (5,755 files for learning no call)\n- xeno-canto files not included in training dataset CC-BY-NC-SA (896 files) and CC-BY-NC-ND (5,212 files).\n- Zenodo dataset (background noise)\n- esc50 (rain, frog)\n- aicrowd2020_noise_30sec (background noise for pretraining)\n\n### Models\n- My models are based on [the BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) using [timm](https://github.com/huggingface/pytorch-image-models).\n- My solution consists of a total of 4 models, all with eca_nfnet_l0 as the backbone. Each set is slightly different.\n\n1. MelSpectrogram (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512, top_db=80.0, NormalizeMelSpec) following [this solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293).\n    - Public: 0.8312, Private: 0.74424\n2. Balanced sampling with the above settings.\n    - Public: 0.83106, Private: 0.74406\n3. [PCEN](https://github.com/daemon/pytorch-pcen) (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512).\n    - Public: 0.83005, Private: 0.74134\n4. MelSpectrogram (sample_rate: 32000, mel_bins: 64, fmin: 50, fmax: 14000, window_size: 1024, hop_size: 320).\n    - Public: 0.83014, Private: 0.74201\n\n### Training\nKnowledge distillation of pre-computed predictions in [Kaggle Models](https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1) (bird-vocalization-classifier) was the hallmark of my solution. This model cannot complete inference in less than 2h, but it is very powerful. cmAP_5 on my validation dataset was 0.9479. So I tried to extract useful information from this model.\n\n**The Kaggle models pre-computed predictions were created [here](https://www.kaggle.com/code/atsunorifujita/extract-from-kaggle-models/notebook)**.\n\nAccording to this [paper](https://arxiv.org/abs/2106.05237), they argued that efficient distillation required 1. consistent input, 2. aggressive mixup, and 3. a large number of epochs. So I did 2 and 3 because it was challenging to integrate the Kaggle model into my training pipeline. It certainly looked effective.\n\nI chose only approaches that improve both CV and LB. In this competition, I just couldn't believe in one or the other.\n\n- Using 5 StratifiedKFold. Only 1 fold did not cover all classes, so the remaining 4 were used for training.\n- The evaluation metric is padded_cmap1. The best results were the same with padded_cmap5 in most cases.\n- Use primary_label only\n- no-call is represented by all 0\n- loss: 0.1 * BCEWithLogitsLoss (primary_label) + 0.9 * KLDivLoss (from Kaggle model)\n  - Softmax temperature=20 was best (tried 5, 10, 20, 30).\n- I used randomly sampled 20-sec clips for training. Audio that is less than 20 sec is repeated.\n- epoch = 400 (Most models converge at 100-300)\n- early stopping(pretraining=10, training=20)\n- Optimizer: AdamW (lr: pretraining=5e-4, training=2.5e-4, wd: 1e-6)\n- CosineLRScheduler(t_initial=10, warmup_t=1, cycle_limit=40, cycle_decay=1.0, lr_min=1e-7, t_in_epochs=True,)\n- mixup p = 1.0  (It was better than p=0.5)<br/>\n\n \nI also encountered a situation where the training time was significantly longer when pretrained with past competition data. Thanks to everyone who suggested solutions.\n\n\n#### Augmentation\n- OneOf ([Gain, GainTransition])\n- OneOf ([AddGaussianNoise, AddGaussianSNR]\n- AddShortNoises esc50 (rain, frog)\n- AddBackgroundNoise from [Zenodo](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605). The 60 minutes with the fewest bird calls were extracted from each dataset and divided into 30 sec (training only).\n- AddBackgroundNoise from aicrowd2020_noise_30sec and ff1010bird_nocall (pretraining only).\n- LowPassFilter\n- PitchShift\n\n#### Hardware\n- 1 * RTX 3090\n- Pretraining time: 30-40h per model. \n- Training time: 9-12h per model.\n\n\n### Inference\nReducing the inference time is the part I was having trouble with. Thanks to @leonshangguan for sharing his [effective approach](https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference), I was able to use 4 models.\n\n- Ensemble with simple averaging.\n- 4 models using PyTorch JIT. total 110min.\n- **[Inference notebook](https://www.kaggle.com/code/atsunorifujita/4th-place-solution-inference-kernel)**\n\n### What didn’t work\n- focal loss.\n- Split a low-sample class.\n- Backbone other than eca_nfnet_l0 and eca_nfnet_l1.\n- Optimizers (adan, lion, ranger21, shampoo). I tried to create a custom normalize-free model but failed.\n- [CMO](https://github.com/naver-ai/cmo) (mixup worked better when using distillation).\n- [CQT](https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/features/cqt.py) (slow and degraded).\n- Change to first stride(1, 1). CV was good but inference takes a long time and not a single model is completed in less than 2h.\n- Pretraining Zenodo data (story before trying distillation).\n- Distillation training from scratch (not use imagenet weights). It didn't converge in the amount of time I could tolerate.\n- MelSpectrogram, PCEN, and CQT integrated into input channels in all combinations. There was no synergistic effect.\n\n### Ablation study\n| Name | Public LB | Private LB |\n| --- | --- |--- |\n| BaseModel | 0.80603 | 0.70782 |\n| BaseModel + Knowledge Distillation | 0.82073 | 0.72752 |\n| BaseModel + Knowledge Distillation + Adding xeno-canto | 0.82905 | 0.74038 |\n| BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining | 0.8312 | 0.74424 |\n| BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining + Ensemble (4 models) | 0.84019 | 0.75688 |\n\n\n\n### Acknowledgments\nMy solution builds on contributions from participants in this and past competitions, birding enthusiasts, and the work of the machine learning community. thank you very much 🙏 \n\n\n**Training Code**: https://github.com/AtsunoriFujita/BirdCLEF-2023-Identify-bird-calls-in-soundscapes",
      "votes": null
    },
    {
      "id": "2273228",
      "postDate": "05/25/2023 04:48:28",
      "content": "<p>Congrats for your solo gold!💯</p>",
      "rawMarkdown": "Congrats for your solo gold!💯",
      "votes": null
    },
    {
      "id": "2273354",
      "postDate": "05/25/2023 06:02:54",
      "content": "<p>Congratulations on your solo gold! Great Find on how to make use of already existing model</p>",
      "rawMarkdown": "Congratulations on your solo gold! Great Find on how to make use of already existing model",
      "votes": null
    },
    {
      "id": "2273478",
      "postDate": "05/25/2023 07:21:01",
      "content": "<p>Congrats on a great solo finish <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> 🎉 your distillation technique is indeed powerful. </p>",
      "rawMarkdown": "Congrats on a great solo finish @atsunorifujita 🎉 your distillation technique is indeed powerful.",
      "votes": null
    },
    {
      "id": "2273540",
      "postDate": "05/25/2023 08:12:17",
      "content": "<p>Awesome! Congratulations <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> 🔥</p>",
      "rawMarkdown": "Awesome! Congratulations @atsunorifujita 🔥",
      "votes": null
    },
    {
      "id": "2273557",
      "postDate": "05/25/2023 08:23:43",
      "content": "<p>Thank you for sharing your approach, I am a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.</p>",
      "rawMarkdown": "Thank you for sharing your approach, I am a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.",
      "votes": null
    },
    {
      "id": "2273650",
      "postDate": "05/25/2023 10:10:03",
      "content": "<p>Congratulations, and thanks for sharing! Just to see if I got it correctly, your training steps were:</p>\n<ol>\n<li>Take a backbone that is pretrained on ImageNet</li>\n<li>Train a model with this backbone on BirdCLEF 2021+2022+ff1010, with Google's bird vocalization classifier as a teacher, lr 1e-5</li>\n<li>Replace the last layer, train the model further on 2023+ff1010+xeno-canto, again with the same teacher, lr 2.5e-4</li>\n</ol>\n<p>Any idea how much you gained in inference time from using PyTorch's JIT? I tried PyTorch 2's compilation on single models or the full ensemble, but it didn't help at all (probably also because of graph-breaking Python control statements).</p>",
      "rawMarkdown": "Congratulations, and thanks for sharing! Just to see if I got it correctly, your training steps were:\n1. Take a backbone that is pretrained on ImageNet\n2. Train a model with this backbone on BirdCLEF 2021+2022+ff1010, with Google's bird vocalization classifier as a teacher, lr 1e-5\n3. Replace the last layer, train the model further on 2023+ff1010+xeno-canto, again with the same teacher, lr 2.5e-4\n\nAny idea how much you gained in inference time from using PyTorch's JIT? I tried PyTorch 2's compilation on single models or the full ensemble, but it didn't help at all (probably also because of graph-breaking Python control statements).",
      "votes": null
    },
    {
      "id": "2273810",
      "postDate": "05/25/2023 12:33:27",
      "content": "<p>Thank you and congratulations on your solo gold 🎉</p>",
      "rawMarkdown": "Thank you and congratulations on your solo gold 🎉",
      "votes": null
    },
    {
      "id": "2273817",
      "postDate": "05/25/2023 12:35:08",
      "content": "<p>Thank you and congratulations 🙌</p>",
      "rawMarkdown": "Thank you and congratulations 🙌",
      "votes": null
    },
    {
      "id": "2273822",
      "postDate": "05/25/2023 12:37:23",
      "content": "<p>Thank you! I hope you are enjoying your work 👍</p>",
      "rawMarkdown": "Thank you! I hope you are enjoying your work 👍",
      "votes": null
    },
    {
      "id": "2273823",
      "postDate": "05/25/2023 12:37:49",
      "content": "<p>Thank you 👍</p>",
      "rawMarkdown": "Thank you 👍",
      "votes": null
    },
    {
      "id": "2273827",
      "postDate": "05/25/2023 12:41:13",
      "content": "<p>Thank you, I think we just have to come up with ideas and experiment. That's why we always have to keep learning.</p>",
      "rawMarkdown": "Thank you, I think we just have to come up with ideas and experiment. That's why we always have to keep learning.",
      "votes": null
    },
    {
      "id": "2273841",
      "postDate": "05/25/2023 12:58:25",
      "content": "<p>Thank you! </p>\n<p>My training steps are as you said. To be precise, ff1010 in 2. is used as background noise.</p>\n<p>PyTorch JIT reduced the inference time by about 20%.\nI also tried ONNX, but didn't use it because of the complexity of the inference pipeline. The PyTorch JIT was able to include MelSpectrogram and PCEN in the model and converted it, simplifying the inference code.</p>",
      "rawMarkdown": "Thank you! \n\nMy training steps are as you said. To be precise, ff1010 in 2. is used as background noise.\n\nPyTorch JIT reduced the inference time by about 20%.\nI also tried ONNX, but didn't use it because of the complexity of the inference pipeline. The PyTorch JIT was able to include MelSpectrogram and PCEN in the model and converted it, simplifying the inference code.",
      "votes": null
    },
    {
      "id": "2275050",
      "postDate": "05/26/2023 13:04:18",
      "content": "<p>Congratulations, and thanks for the helpful post! Could you also make your training code available on github and post a link?</p>",
      "rawMarkdown": "Congratulations, and thanks for the helpful post! Could you also make your training code available on github and post a link?",
      "votes": null
    },
    {
      "id": "2276835",
      "postDate": "05/27/2023 08:37:42",
      "content": "<p>Absolutely! I'm preparing for it.</p>",
      "rawMarkdown": "Absolutely! I'm preparing for it.",
      "votes": null
    },
    {
      "id": "2281991",
      "postDate": "05/31/2023 09:36:20",
      "content": "<p>Thank you very much in advance.  I am interested in distillation, but I cannot make it work </p>",
      "rawMarkdown": "Thank you very much in advance.  I am interested in distillation, but I cannot make it work",
      "votes": null
    },
    {
      "id": "2286237",
      "postDate": "06/03/2023 10:24:12",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> and thanks for sharing this detailed writeup 🙏</p>",
      "rawMarkdown": "Congrats @atsunorifujita and thanks for sharing this detailed writeup 🙏",
      "votes": null
    },
    {
      "id": "2287980",
      "postDate": "06/05/2023 05:09:09",
      "content": "<p>I have published a GitHub repo for training.</p>",
      "rawMarkdown": "I have published a GitHub repo for training.",
      "votes": null
    },
    {
      "id": "2288056",
      "postDate": "06/05/2023 06:31:26",
      "content": "<p>Thank you very much, I will have a look</p>",
      "rawMarkdown": "Thank you very much, I will have a look",
      "votes": null
    },
    {
      "id": "2288656",
      "postDate": "06/05/2023 14:56:53",
      "content": "<p>Congratulations!, I'd like to know how many time you needed to train your folds in a RTX 3090.</p>\n<p>Thank you in advance!</p>",
      "rawMarkdown": "Congratulations!, I'd like to know how many time you needed to train your folds in a RTX 3090.\n\nThank you in advance!",
      "votes": null
    },
    {
      "id": "2289137",
      "postDate": "06/06/2023 00:29:35",
      "content": "<p>Thank you and good question!</p>\n<ul>\n<li>Pretraining time: 30-40h per model.</li>\n<li>Training time: 9-12h per model.</li>\n</ul>",
      "rawMarkdown": "Thank you and good question!\n\n- Pretraining time: 30-40h per model.\n- Training time: 9-12h per model.",
      "votes": null
    },
    {
      "id": "2289703",
      "postDate": "06/06/2023 10:15:20",
      "content": "<p>Congratulations and brilliant discussion. A few questions since I saw the ps recently:</p>\n<ol>\n<li>Is there any reason why Focal Loss cannot outperform BCE?</li>\n<li>Why do you prefer librosa over torchaudio?</li>\n<li>Is there any substantial postprocessing you used in inference or that you plan to use as future work?</li>\n<li>What models are appropriate for ONNX?</li>\n<li>How do you get started with PyTorch jit? Do you have any recommended documentation? </li>\n</ol>",
      "rawMarkdown": "Congratulations and brilliant discussion. A few questions since I saw the ps recently:\n1. Is there any reason why Focal Loss cannot outperform BCE?\n2. Why do you prefer librosa over torchaudio?\n3. Is there any substantial postprocessing you used in inference or that you plan to use as future work?\n4. What models are appropriate for ONNX?\n5. How do you get started with PyTorch jit? Do you have any recommended documentation?",
      "votes": null
    },
    {
      "id": "2290599",
      "postDate": "06/07/2023 01:13:15",
      "content": "<p>How much memory the RTX 3090?</p>",
      "rawMarkdown": "How much memory the RTX 3090?",
      "votes": null
    },
    {
      "id": "2290725",
      "postDate": "06/07/2023 04:15:52",
      "content": "<p>RTX 3090 has 24GB of VRAM</p>",
      "rawMarkdown": "RTX 3090 has 24GB of VRAM",
      "votes": null
    },
    {
      "id": "2290759",
      "postDate": "06/07/2023 04:44:20",
      "content": "<p>Thank you and nice questions!</p>\n<ol>\n<li><p>Is there any reason why Focal Loss cannot outperform BCE?    <br>\nI don't know why. I was just comparing the experimental results. Focal loss may be better for different data.</p></li>\n<li><p>Why do you prefer librosa over torchaudio?<br>\nBecause I wanted my model inputs and KaggleModel inputs to be as consistent as possible. Firstly I was using soundfile.</p></li>\n<li><p>Is there any substantial postprocessing you used in inference or that you plan to use as future work?<br>\nI can't think of anything at the moment. A simple average of multiple models worked.</p></li>\n<li><p>What models are appropriate for ONNX?<br>\nIt was possible to convert to ONNX by separating mel spectrogram and PCEN from my model. However, I got lost motivation to use it because the inference time was not significantly different from PyTorch JIT.</p></li>\n<li><p>How do you get started with PyTorch jit? Do you have any recommended documentation?<br>\nIn many cases, it's pretty straightforward (it fails with unsupported layers). The official documentation is helpful, but you can convert it with a few lines of code.</p></li>\n</ol>\n<p><a href=\"https://pytorch.org/docs/stable/jit.html\" target=\"_blank\">https://pytorch.org/docs/stable/jit.html</a></p>",
      "rawMarkdown": "Thank you and nice questions!\n\n1. Is there any reason why Focal Loss cannot outperform BCE?    \nI don't know why. I was just comparing the experimental results. Focal loss may be better for different data.\n\n\n2. Why do you prefer librosa over torchaudio?\nBecause I wanted my model inputs and KaggleModel inputs to be as consistent as possible. Firstly I was using soundfile.\n\n\n3. Is there any substantial postprocessing you used in inference or that you plan to use as future work?\nI can't think of anything at the moment. A simple average of multiple models worked.\n\n\n4. What models are appropriate for ONNX?\nIt was possible to convert to ONNX by separating mel spectrogram and PCEN from my model. However, I got lost motivation to use it because the inference time was not significantly different from PyTorch JIT.\n\n\n5. How do you get started with PyTorch jit? Do you have any recommended documentation?\nIn many cases, it's pretty straightforward (it fails with unsupported layers). The official documentation is helpful, but you can convert it with a few lines of code.\n\nhttps://pytorch.org/docs/stable/jit.html",
      "votes": null
    },
    {
      "id": "2292640",
      "postDate": "06/08/2023 14:11:45",
      "content": "<p>Thanks! I downloaded your repo and followed your setup instructions, and am now able to run the training. I look forward to studying it!</p>",
      "rawMarkdown": "Thanks! I downloaded your repo and followed your setup instructions, and am now able to run the training. I look forward to studying it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2273228,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "05/25/2023 04:48:28",
      "content": "<p>Congrats for your solo gold!💯</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273810,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/25/2023 12:33:27",
          "content": "<p>Thank you and congratulations on your solo gold 🎉</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273354,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "05/25/2023 06:02:54",
      "content": "<p>Congratulations on your solo gold! Great Find on how to make use of already existing model</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273817,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/25/2023 12:35:08",
          "content": "<p>Thank you and congratulations 🙌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273478,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "05/25/2023 07:21:01",
      "content": "<p>Congrats on a great solo finish <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> 🎉 your distillation technique is indeed powerful. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2273822,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/25/2023 12:37:23",
          "content": "<p>Thank you! I hope you are enjoying your work 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273540,
      "author_name": "scipygaurav",
      "author_url": "",
      "post_date": "05/25/2023 08:12:17",
      "content": "<p>Awesome! Congratulations <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> 🔥</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273823,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/25/2023 12:37:49",
          "content": "<p>Thank you 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273557,
      "author_name": "aadityabansalcodes",
      "author_url": "",
      "post_date": "05/25/2023 08:23:43",
      "content": "<p>Thank you for sharing your approach, I am a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273827,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/25/2023 12:41:13",
          "content": "<p>Thank you, I think we just have to come up with ideas and experiment. That's why we always have to keep learning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273650,
      "author_name": "janschl",
      "author_url": "",
      "post_date": "05/25/2023 10:10:03",
      "content": "<p>Congratulations, and thanks for sharing! Just to see if I got it correctly, your training steps were:</p>\n<ol>\n<li>Take a backbone that is pretrained on ImageNet</li>\n<li>Train a model with this backbone on BirdCLEF 2021+2022+ff1010, with Google's bird vocalization classifier as a teacher, lr 1e-5</li>\n<li>Replace the last layer, train the model further on 2023+ff1010+xeno-canto, again with the same teacher, lr 2.5e-4</li>\n</ol>\n<p>Any idea how much you gained in inference time from using PyTorch's JIT? I tried PyTorch 2's compilation on single models or the full ensemble, but it didn't help at all (probably also because of graph-breaking Python control statements).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2273841,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/25/2023 12:58:25",
          "content": "<p>Thank you! </p>\n<p>My training steps are as you said. To be precise, ff1010 in 2. is used as background noise.</p>\n<p>PyTorch JIT reduced the inference time by about 20%.\nI also tried ONNX, but didn't use it because of the complexity of the inference pipeline. The PyTorch JIT was able to include MelSpectrogram and PCEN in the model and converted it, simplifying the inference code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2275050,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "05/26/2023 13:04:18",
      "content": "<p>Congratulations, and thanks for the helpful post! Could you also make your training code available on github and post a link?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276835,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/27/2023 08:37:42",
          "content": "<p>Absolutely! I'm preparing for it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2281991,
              "author_name": "nyleve",
              "author_url": "",
              "post_date": "05/31/2023 09:36:20",
              "content": "<p>Thank you very much in advance.  I am interested in distillation, but I cannot make it work </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2287980,
                  "author_name": "atsunorifujita",
                  "author_url": "",
                  "post_date": "06/05/2023 05:09:09",
                  "content": "<p>I have published a GitHub repo for training.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2288056,
                      "author_name": "nyleve",
                      "author_url": "",
                      "post_date": "06/05/2023 06:31:26",
                      "content": "<p>Thank you very much, I will have a look</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2292640,
                      "author_name": "janhuus",
                      "author_url": "",
                      "post_date": "06/08/2023 14:11:45",
                      "content": "<p>Thanks! I downloaded your repo and followed your setup instructions, and am now able to run the training. I look forward to studying it!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2286237,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 10:24:12",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/atsunorifujita\" target=\"_blank\">@atsunorifujita</a> and thanks for sharing this detailed writeup 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2288656,
      "author_name": "pablolarrosa",
      "author_url": "",
      "post_date": "06/05/2023 14:56:53",
      "content": "<p>Congratulations!, I'd like to know how many time you needed to train your folds in a RTX 3090.</p>\n<p>Thank you in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2289137,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "06/06/2023 00:29:35",
          "content": "<p>Thank you and good question!</p>\n<ul>\n<li>Pretraining time: 30-40h per model.</li>\n<li>Training time: 9-12h per model.</li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2290599,
              "author_name": "pablolarrosa",
              "author_url": "",
              "post_date": "06/07/2023 01:13:15",
              "content": "<p>How much memory the RTX 3090?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2290725,
                  "author_name": "atsunorifujita",
                  "author_url": "",
                  "post_date": "06/07/2023 04:15:52",
                  "content": "<p>RTX 3090 has 24GB of VRAM</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2289703,
      "author_name": "pushkar007",
      "author_url": "",
      "post_date": "06/06/2023 10:15:20",
      "content": "<p>Congratulations and brilliant discussion. A few questions since I saw the ps recently:</p>\n<ol>\n<li>Is there any reason why Focal Loss cannot outperform BCE?</li>\n<li>Why do you prefer librosa over torchaudio?</li>\n<li>Is there any substantial postprocessing you used in inference or that you plan to use as future work?</li>\n<li>What models are appropriate for ONNX?</li>\n<li>How do you get started with PyTorch jit? Do you have any recommended documentation? </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2290759,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "06/07/2023 04:44:20",
          "content": "<p>Thank you and nice questions!</p>\n<ol>\n<li><p>Is there any reason why Focal Loss cannot outperform BCE?    <br>\nI don't know why. I was just comparing the experimental results. Focal loss may be better for different data.</p></li>\n<li><p>Why do you prefer librosa over torchaudio?<br>\nBecause I wanted my model inputs and KaggleModel inputs to be as consistent as possible. Firstly I was using soundfile.</p></li>\n<li><p>Is there any substantial postprocessing you used in inference or that you plan to use as future work?<br>\nI can't think of anything at the moment. A simple average of multiple models worked.</p></li>\n<li><p>What models are appropriate for ONNX?<br>\nIt was possible to convert to ONNX by separating mel spectrogram and PCEN from my model. However, I got lost motivation to use it because the inference time was not significantly different from PyTorch JIT.</p></li>\n<li><p>How do you get started with PyTorch jit? Do you have any recommended documentation?<br>\nIn many cases, it's pretty straightforward (it fails with unsupported layers). The official documentation is helpful, but you can convert it with a few lines of code.</p></li>\n</ol>\n<p><a href=\"https://pytorch.org/docs/stable/jit.html\" target=\"_blank\">https://pytorch.org/docs/stable/jit.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2273202": "First of all, thanks to Kaggle and Cornell Lab of Ornithology for hosting this interesting competition. And I would like to emphasize the thanks to the bird sound recordists who provided their data through xeno-canto.\n\n### Solution Summary\n- Knowledge Distillation is all you need.\n- Adding no-call data, xeno-canto data, and background audios(Zenodo) is effective.\n\n### Datasets\nAdditional datasets can be found [here](https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-additional).\n- Bird CLEF 2023\n- Bird CLEF 2021, 2022 (for pretraining)\n- ff1010bird_nocall (5,755 files for learning no call)\n- xeno-canto files not included in training dataset CC-BY-NC-SA (896 files) and CC-BY-NC-ND (5,212 files).\n- Zenodo dataset (background noise)\n- esc50 (rain, frog)\n- aicrowd2020_noise_30sec (background noise for pretraining)\n\n### Models\n- My models are based on [the BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463) using [timm](https://github.com/huggingface/pytorch-image-models).\n- My solution consists of a total of 4 models, all with eca_nfnet_l0 as the backbone. Each set is slightly different.\n\n1. MelSpectrogram (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512, top_db=80.0, NormalizeMelSpec) following [this solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243293).\n    - Public: 0.8312, Private: 0.74424\n2. Balanced sampling with the above settings.\n    - Public: 0.83106, Private: 0.74406\n3. [PCEN](https://github.com/daemon/pytorch-pcen) (sample_rate: 32000, mel_bins: 128, fmin: 20, fmax: 16000, window_size: 2048, hop_size: 512).\n    - Public: 0.83005, Private: 0.74134\n4. MelSpectrogram (sample_rate: 32000, mel_bins: 64, fmin: 50, fmax: 14000, window_size: 1024, hop_size: 320).\n    - Public: 0.83014, Private: 0.74201\n\n### Training\nKnowledge distillation of pre-computed predictions in [Kaggle Models](https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1) (bird-vocalization-classifier) was the hallmark of my solution. This model cannot complete inference in less than 2h, but it is very powerful. cmAP_5 on my validation dataset was 0.9479. So I tried to extract useful information from this model.\n\n**The Kaggle models pre-computed predictions were created [here](https://www.kaggle.com/code/atsunorifujita/extract-from-kaggle-models/notebook)**.\n\nAccording to this [paper](https://arxiv.org/abs/2106.05237), they argued that efficient distillation required 1. consistent input, 2. aggressive mixup, and 3. a large number of epochs. So I did 2 and 3 because it was challenging to integrate the Kaggle model into my training pipeline. It certainly looked effective.\n\nI chose only approaches that improve both CV and LB. In this competition, I just couldn't believe in one or the other.\n\n- Using 5 StratifiedKFold. Only 1 fold did not cover all classes, so the remaining 4 were used for training.\n- The evaluation metric is padded_cmap1. The best results were the same with padded_cmap5 in most cases.\n- Use primary_label only\n- no-call is represented by all 0\n- loss: 0.1 * BCEWithLogitsLoss (primary_label) + 0.9 * KLDivLoss (from Kaggle model)\n  - Softmax temperature=20 was best (tried 5, 10, 20, 30).\n- I used randomly sampled 20-sec clips for training. Audio that is less than 20 sec is repeated.\n- epoch = 400 (Most models converge at 100-300)\n- early stopping(pretraining=10, training=20)\n- Optimizer: AdamW (lr: pretraining=5e-4, training=2.5e-4, wd: 1e-6)\n- CosineLRScheduler(t_initial=10, warmup_t=1, cycle_limit=40, cycle_decay=1.0, lr_min=1e-7, t_in_epochs=True,)\n- mixup p = 1.0  (It was better than p=0.5)<br/>\n\n \nI also encountered a situation where the training time was significantly longer when pretrained with past competition data. Thanks to everyone who suggested solutions.\n\n\n#### Augmentation\n- OneOf ([Gain, GainTransition])\n- OneOf ([AddGaussianNoise, AddGaussianSNR]\n- AddShortNoises esc50 (rain, frog)\n- AddBackgroundNoise from [Zenodo](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358#2179605). The 60 minutes with the fewest bird calls were extracted from each dataset and divided into 30 sec (training only).\n- AddBackgroundNoise from aicrowd2020_noise_30sec and ff1010bird_nocall (pretraining only).\n- LowPassFilter\n- PitchShift\n\n#### Hardware\n- 1 * RTX 3090\n- Pretraining time: 30-40h per model. \n- Training time: 9-12h per model.\n\n\n### Inference\nReducing the inference time is the part I was having trouble with. Thanks to @leonshangguan for sharing his [effective approach](https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference), I was able to use 4 models.\n\n- Ensemble with simple averaging.\n- 4 models using PyTorch JIT. total 110min.\n- **[Inference notebook](https://www.kaggle.com/code/atsunorifujita/4th-place-solution-inference-kernel)**\n\n### What didn’t work\n- focal loss.\n- Split a low-sample class.\n- Backbone other than eca_nfnet_l0 and eca_nfnet_l1.\n- Optimizers (adan, lion, ranger21, shampoo). I tried to create a custom normalize-free model but failed.\n- [CMO](https://github.com/naver-ai/cmo) (mixup worked better when using distillation).\n- [CQT](https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/features/cqt.py) (slow and degraded).\n- Change to first stride(1, 1). CV was good but inference takes a long time and not a single model is completed in less than 2h.\n- Pretraining Zenodo data (story before trying distillation).\n- Distillation training from scratch (not use imagenet weights). It didn't converge in the amount of time I could tolerate.\n- MelSpectrogram, PCEN, and CQT integrated into input channels in all combinations. There was no synergistic effect.\n\n### Ablation study\n| Name | Public LB | Private LB |\n| --- | --- |--- |\n| BaseModel | 0.80603 | 0.70782 |\n| BaseModel + Knowledge Distillation | 0.82073 | 0.72752 |\n| BaseModel + Knowledge Distillation + Adding xeno-canto | 0.82905 | 0.74038 |\n| BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining | 0.8312 | 0.74424 |\n| BaseModel + Knowledge Distillation + Adding xeno-canto + Pretraining + Ensemble (4 models) | 0.84019 | 0.75688 |\n\n\n\n### Acknowledgments\nMy solution builds on contributions from participants in this and past competitions, birding enthusiasts, and the work of the machine learning community. thank you very much 🙏 \n\n\n**Training Code**: https://github.com/AtsunoriFujita/BirdCLEF-2023-Identify-bird-calls-in-soundscapes",
    "2273228": "Congrats for your solo gold!💯",
    "2273354": "Congratulations on your solo gold! Great Find on how to make use of already existing model",
    "2273478": "Congrats on a great solo finish @atsunorifujita 🎉 your distillation technique is indeed powerful.",
    "2273540": "Awesome! Congratulations @atsunorifujita 🔥",
    "2273557": "Thank you for sharing your approach, I am a beginner in Deep Learning, do you have any tips as to how I can approach complex problems like this.",
    "2273650": "Congratulations, and thanks for sharing! Just to see if I got it correctly, your training steps were:\n1. Take a backbone that is pretrained on ImageNet\n2. Train a model with this backbone on BirdCLEF 2021+2022+ff1010, with Google's bird vocalization classifier as a teacher, lr 1e-5\n3. Replace the last layer, train the model further on 2023+ff1010+xeno-canto, again with the same teacher, lr 2.5e-4\n\nAny idea how much you gained in inference time from using PyTorch's JIT? I tried PyTorch 2's compilation on single models or the full ensemble, but it didn't help at all (probably also because of graph-breaking Python control statements).",
    "2273810": "Thank you and congratulations on your solo gold 🎉",
    "2273817": "Thank you and congratulations 🙌",
    "2273822": "Thank you! I hope you are enjoying your work 👍",
    "2273823": "Thank you 👍",
    "2273827": "Thank you, I think we just have to come up with ideas and experiment. That's why we always have to keep learning.",
    "2273841": "Thank you! \n\nMy training steps are as you said. To be precise, ff1010 in 2. is used as background noise.\n\nPyTorch JIT reduced the inference time by about 20%.\nI also tried ONNX, but didn't use it because of the complexity of the inference pipeline. The PyTorch JIT was able to include MelSpectrogram and PCEN in the model and converted it, simplifying the inference code.",
    "2275050": "Congratulations, and thanks for the helpful post! Could you also make your training code available on github and post a link?",
    "2276835": "Absolutely! I'm preparing for it.",
    "2281991": "Thank you very much in advance.  I am interested in distillation, but I cannot make it work",
    "2286237": "Congrats @atsunorifujita and thanks for sharing this detailed writeup 🙏",
    "2287980": "I have published a GitHub repo for training.",
    "2288056": "Thank you very much, I will have a look",
    "2288656": "Congratulations!, I'd like to know how many time you needed to train your folds in a RTX 3090.\n\nThank you in advance!",
    "2289137": "Thank you and good question!\n\n- Pretraining time: 30-40h per model.\n- Training time: 9-12h per model.",
    "2289703": "Congratulations and brilliant discussion. A few questions since I saw the ps recently:\n1. Is there any reason why Focal Loss cannot outperform BCE?\n2. Why do you prefer librosa over torchaudio?\n3. Is there any substantial postprocessing you used in inference or that you plan to use as future work?\n4. What models are appropriate for ONNX?\n5. How do you get started with PyTorch jit? Do you have any recommended documentation?",
    "2290599": "How much memory the RTX 3090?",
    "2290725": "RTX 3090 has 24GB of VRAM",
    "2290759": "Thank you and nice questions!\n\n1. Is there any reason why Focal Loss cannot outperform BCE?    \nI don't know why. I was just comparing the experimental results. Focal loss may be better for different data.\n\n\n2. Why do you prefer librosa over torchaudio?\nBecause I wanted my model inputs and KaggleModel inputs to be as consistent as possible. Firstly I was using soundfile.\n\n\n3. Is there any substantial postprocessing you used in inference or that you plan to use as future work?\nI can't think of anything at the moment. A simple average of multiple models worked.\n\n\n4. What models are appropriate for ONNX?\nIt was possible to convert to ONNX by separating mel spectrogram and PCEN from my model. However, I got lost motivation to use it because the inference time was not significantly different from PyTorch JIT.\n\n\n5. How do you get started with PyTorch jit? Do you have any recommended documentation?\nIn many cases, it's pretty straightforward (it fails with unsupported layers). The official documentation is helpful, but you can convert it with a few lines of code.\n\nhttps://pytorch.org/docs/stable/jit.html",
    "2292640": "Thanks! I downloaded your repo and followed your setup instructions, and am now able to run the training. I look forward to studying it!"
  },
  "source": "meta"
}