{
  "id": 412708,
  "title": "6th place solution: BirdNET embedding + CNN",
  "url": "/competitions/birdclef-2023/writeups/anonamename-6th-place-solution-birdnet-embedding-c",
  "author_name": "",
  "post_date": "2023-05-25T16:48:46.343Z",
  "votes": 28,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thank you to the host and Kaggle for organizing the competition.</p>\n<p>My solution is a combination of the embedding vectors of <a href=\"https://github.com/kahst/BirdNET-Analyzer/tree/d1f5a9c015d4419277cbb285e89d3f843a6bab49\" target=\"_blank\">BirdNET-Analyzer V2.2</a> and the CNN from the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>.</p>\n<h2>Model (BirdNET embedding + CNN)</h2>\n<p>In order to utilize the features of BirdNET, the embedding of BirdNET is concatenated to the output of the CNN of the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1694617%2Fe7dcc6f87186d43df680592acb3f91d7%2Fmy_solution.png?generation=1684975539169704&amp;alt=media\" alt=\"\"></p>\n<h3>BirdNET embedding</h3>\n<p>The embedding vectors of BirdNET were created with <a href=\"https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/embeddings.py\" target=\"_blank\">BirdNET V2.2 embeddings.py</a>. I modified <a href=\"https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/audio.py#L7\" target=\"_blank\">BirdNET V2.2 audio.py</a> so that if the length of the audio data is shorter than BirdNET's sample rate (48000), the data is padded to output at least 1 second of embedding vector. Using V2.2 rather than the latest version BirdNET V2.3, slightly improved cv.</p>\n<h3>CNN</h3>\n<p>The backbone of the CNN from the BirdCLEF 2021 2nd place solution used timm's <code>eca_nfnet_l1</code> and <code>seresnext26t_32x4d</code>.</p>\n<h2>Training</h2>\n<p>After pretraining with data from BirdCLEF 2021 + BirdCLEF 2022, I train the CNN and other linear layers with data from BirdCLEF 2023. BirdNET is not trained.</p>\n<p>The input for training is data for 30 seconds. Since BirdNET outputs embedding vectors for 3 seconds of data, I averaged each of the embedding vectors for 30 seconds.</p>\n<p>In most experiments, cv was highest in the final epoch, so I included all data in the training set.</p>\n<p>The main training parameters are as follows:</p>\n<ul>\n<li><p>loss: BCEWithLogitsLoss</p></li>\n<li><p>MelSpectrogram</p>\n<ul>\n<li>sample_rate: 32000, window_size: 1024, hop_size: 320, fmin: 0, fmax: 14000, mel_bins: 128, power: 2, top_db=None</li></ul></li>\n<li><p>labels: primary label=0.9995, secondary label=0.4 or 0.5 (<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\" target=\"_blank\">BirdCLEF 2022 3rd place solution</a>), label smoothing=0.01</p></li>\n<li><p>cv: primary_label StratifiedKFold(n_splits=5)</p></li>\n<li><p>optimizer: AdamW(weight_decay=1e-4)</p></li>\n<li><p>scheduler: warmup 0-3epoch(lr=3e-6-&gt;3e-4) + Cosine Annealing 3-70epoch(lr=3e-4-&gt;3e-6)</p></li>\n</ul>\n<h1>Augmentation</h1>\n<p>As in the previous two competitions, augmentation was important. I combined the following augmentations:</p>\n<ul>\n<li>audiomentations.Shift</li>\n<li>Mixup (<a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>)</li>\n<li>Gaussian Noise</li>\n<li>random lowpass filter</li>\n<li>random power (<a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">BirdCLEF 2021 5th place solution</a>)</li>\n<li>torchaudio.transforms.FrequencyMasking</li>\n<li>torchaudio.transforms.TimeMasking</li>\n</ul>\n<h1>Oversampling</h1>\n<p>The following Oversampling was performed, but I think the effect was minimal.</p>\n<ul>\n<li><p>Oversampling to have at least 20 of data for each class</p></li>\n<li><p>I doubled the number of training data that have two or more secondary labels because data with more than two secondary_labels have worse val_loss</p></li>\n<li><p>As in the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327044\" target=\"_blank\">BirdCLEF 2022 5th place solution</a>, files with reverb/echo effects created for some minority class data and added to the training data</p></li>\n<li><p>Strong augmentation was applied to the oversampled data</p></li>\n</ul>\n<h1>Inference</h1>\n<p>The CNN input is given 5 seconds of data for inference, and the BirdNET input is given the first 3 seconds of data for inference.</p>\n<ul>\n<li><p>To submit within 2 hours, I converted each model to ONNX</p></li>\n<li><p>As in <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">this notebook</a>, I also used ThreadPoolExecutor to speed up inference</p></li>\n</ul>\n<h1>Ensembling</h1>\n<p>The outputs of 2 models with different CNN backbones (<code>eca_nfnet_l1</code>, <code>seresnext26t_32x4d</code>) simply averaged.</p>\n<h1>What did not work</h1>\n<ul>\n<li>add background noise (<a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>)</li>\n<li>postprocess (<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326950\" target=\"_blank\">BirdCLEF 2022 2nd place solution</a>)</li>\n<li>Speeding up inference using OpenVINO<ul>\n<li>Even if the model was converted to OpenVINO, the inference time was not much different from ONNX. I think I just used it wrong.</li></ul></li>\n<li>Using MFCC as input (as in <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9495150/\" target=\"_blank\">this paper</a>)</li>\n<li>Using CQT as input</li>\n<li>Using <code>FAMILY</code> etc. from <code>eBird_Taxonomy_v2021.csv</code> for prediction (as in <a href=\"https://arxiv.org/pdf/2110.03209.pdf\" target=\"_blank\">this paper</a>)</li>\n</ul>\n<h1>Rough lb history</h1>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>Private Score</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BirdCLEF 2021 2nd place solution CNN (with BirdCLEF 2021 + BirdCLEF 2022 pretrain)</td>\n<td>0.71</td>\n<td>0.80</td>\n</tr>\n<tr>\n<td>BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding</td>\n<td>0.74</td>\n<td>0.82</td>\n</tr>\n<tr>\n<td>BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding + Oversampling + Ensembling</td>\n<td>0.75</td>\n<td>0.83</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2273019",
      "postDate": "05/25/2023 00:52:02",
      "content": "<p>Thank you to the host and Kaggle for organizing the competition.</p>\n<p>My solution is a combination of the embedding vectors of <a href=\"https://github.com/kahst/BirdNET-Analyzer/tree/d1f5a9c015d4419277cbb285e89d3f843a6bab49\" target=\"_blank\">BirdNET-Analyzer V2.2</a> and the CNN from the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>.</p>\n<h2>Model (BirdNET embedding + CNN)</h2>\n<p>In order to utilize the features of BirdNET, the embedding of BirdNET is concatenated to the output of the CNN of the <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1694617%2Fe7dcc6f87186d43df680592acb3f91d7%2Fmy_solution.png?generation=1684975539169704&amp;alt=media\" alt=\"\"></p>\n<h3>BirdNET embedding</h3>\n<p>The embedding vectors of BirdNET were created with <a href=\"https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/embeddings.py\" target=\"_blank\">BirdNET V2.2 embeddings.py</a>. I modified <a href=\"https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/audio.py#L7\" target=\"_blank\">BirdNET V2.2 audio.py</a> so that if the length of the audio data is shorter than BirdNET's sample rate (48000), the data is padded to output at least 1 second of embedding vector. Using V2.2 rather than the latest version BirdNET V2.3, slightly improved cv.</p>\n<h3>CNN</h3>\n<p>The backbone of the CNN from the BirdCLEF 2021 2nd place solution used timm's <code>eca_nfnet_l1</code> and <code>seresnext26t_32x4d</code>.</p>\n<h2>Training</h2>\n<p>After pretraining with data from BirdCLEF 2021 + BirdCLEF 2022, I train the CNN and other linear layers with data from BirdCLEF 2023. BirdNET is not trained.</p>\n<p>The input for training is data for 30 seconds. Since BirdNET outputs embedding vectors for 3 seconds of data, I averaged each of the embedding vectors for 30 seconds.</p>\n<p>In most experiments, cv was highest in the final epoch, so I included all data in the training set.</p>\n<p>The main training parameters are as follows:</p>\n<ul>\n<li><p>loss: BCEWithLogitsLoss</p></li>\n<li><p>MelSpectrogram</p>\n<ul>\n<li>sample_rate: 32000, window_size: 1024, hop_size: 320, fmin: 0, fmax: 14000, mel_bins: 128, power: 2, top_db=None</li></ul></li>\n<li><p>labels: primary label=0.9995, secondary label=0.4 or 0.5 (<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327193\" target=\"_blank\">BirdCLEF 2022 3rd place solution</a>), label smoothing=0.01</p></li>\n<li><p>cv: primary_label StratifiedKFold(n_splits=5)</p></li>\n<li><p>optimizer: AdamW(weight_decay=1e-4)</p></li>\n<li><p>scheduler: warmup 0-3epoch(lr=3e-6-&gt;3e-4) + Cosine Annealing 3-70epoch(lr=3e-4-&gt;3e-6)</p></li>\n</ul>\n<h1>Augmentation</h1>\n<p>As in the previous two competitions, augmentation was important. I combined the following augmentations:</p>\n<ul>\n<li>audiomentations.Shift</li>\n<li>Mixup (<a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>)</li>\n<li>Gaussian Noise</li>\n<li>random lowpass filter</li>\n<li>random power (<a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243351\" target=\"_blank\">BirdCLEF 2021 5th place solution</a>)</li>\n<li>torchaudio.transforms.FrequencyMasking</li>\n<li>torchaudio.transforms.TimeMasking</li>\n</ul>\n<h1>Oversampling</h1>\n<p>The following Oversampling was performed, but I think the effect was minimal.</p>\n<ul>\n<li><p>Oversampling to have at least 20 of data for each class</p></li>\n<li><p>I doubled the number of training data that have two or more secondary labels because data with more than two secondary_labels have worse val_loss</p></li>\n<li><p>As in the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/327044\" target=\"_blank\">BirdCLEF 2022 5th place solution</a>, files with reverb/echo effects created for some minority class data and added to the training data</p></li>\n<li><p>Strong augmentation was applied to the oversampled data</p></li>\n</ul>\n<h1>Inference</h1>\n<p>The CNN input is given 5 seconds of data for inference, and the BirdNET input is given the first 3 seconds of data for inference.</p>\n<ul>\n<li><p>To submit within 2 hours, I converted each model to ONNX</p></li>\n<li><p>As in <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">this notebook</a>, I also used ThreadPoolExecutor to speed up inference</p></li>\n</ul>\n<h1>Ensembling</h1>\n<p>The outputs of 2 models with different CNN backbones (<code>eca_nfnet_l1</code>, <code>seresnext26t_32x4d</code>) simply averaged.</p>\n<h1>What did not work</h1>\n<ul>\n<li>add background noise (<a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">BirdCLEF 2021 2nd place solution</a>)</li>\n<li>postprocess (<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/326950\" target=\"_blank\">BirdCLEF 2022 2nd place solution</a>)</li>\n<li>Speeding up inference using OpenVINO<ul>\n<li>Even if the model was converted to OpenVINO, the inference time was not much different from ONNX. I think I just used it wrong.</li></ul></li>\n<li>Using MFCC as input (as in <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9495150/\" target=\"_blank\">this paper</a>)</li>\n<li>Using CQT as input</li>\n<li>Using <code>FAMILY</code> etc. from <code>eBird_Taxonomy_v2021.csv</code> for prediction (as in <a href=\"https://arxiv.org/pdf/2110.03209.pdf\" target=\"_blank\">this paper</a>)</li>\n</ul>\n<h1>Rough lb history</h1>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>Private Score</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BirdCLEF 2021 2nd place solution CNN (with BirdCLEF 2021 + BirdCLEF 2022 pretrain)</td>\n<td>0.71</td>\n<td>0.80</td>\n</tr>\n<tr>\n<td>BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding</td>\n<td>0.74</td>\n<td>0.82</td>\n</tr>\n<tr>\n<td>BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding + Oversampling + Ensembling</td>\n<td>0.75</td>\n<td>0.83</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Thank you to the host and Kaggle for organizing the competition.\n\nMy solution is a combination of the embedding vectors of [BirdNET-Analyzer V2.2](https://github.com/kahst/BirdNET-Analyzer/tree/d1f5a9c015d4419277cbb285e89d3f843a6bab49) and the CNN from the [BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463).\n\n## Model (BirdNET embedding + CNN)\n\nIn order to utilize the features of BirdNET, the embedding of BirdNET is concatenated to the output of the CNN of the [BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1694617%2Fe7dcc6f87186d43df680592acb3f91d7%2Fmy_solution.png?generation=1684975539169704&alt=media)\n\n### BirdNET embedding\n\nThe embedding vectors of BirdNET were created with [BirdNET V2.2 embeddings.py](https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/embeddings.py). I modified [BirdNET V2.2 audio.py](https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/audio.py#L7) so that if the length of the audio data is shorter than BirdNET's sample rate (48000), the data is padded to output at least 1 second of embedding vector. Using V2.2 rather than the latest version BirdNET V2.3, slightly improved cv.\n\n### CNN\n\nThe backbone of the CNN from the BirdCLEF 2021 2nd place solution used timm's `eca_nfnet_l1` and `seresnext26t_32x4d`.\n\n## Training\n\nAfter pretraining with data from BirdCLEF 2021 + BirdCLEF 2022, I train the CNN and other linear layers with data from BirdCLEF 2023. BirdNET is not trained.\n\nThe input for training is data for 30 seconds. Since BirdNET outputs embedding vectors for 3 seconds of data, I averaged each of the embedding vectors for 30 seconds.\n\nIn most experiments, cv was highest in the final epoch, so I included all data in the training set.\n\nThe main training parameters are as follows:\n\n- loss: BCEWithLogitsLoss\n\n- MelSpectrogram\n  \n  - sample_rate: 32000, window_size: 1024, hop_size: 320, fmin: 0, fmax: 14000, mel_bins: 128, power: 2, top_db=None\n\n- labels: primary label=0.9995, secondary label=0.4 or 0.5 ([BirdCLEF 2022 3rd place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/327193)), label smoothing=0.01\n\n- cv: primary_label StratifiedKFold(n_splits=5)\n\n- optimizer: AdamW(weight_decay=1e-4)\n\n- scheduler: warmup 0-3epoch(lr=3e-6->3e-4) + Cosine Annealing 3-70epoch(lr=3e-4->3e-6)\n\n# Augmentation\n\nAs in the previous two competitions, augmentation was important. I combined the following augmentations:\n\n- audiomentations.Shift\n- Mixup ([BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463))\n- Gaussian Noise\n- random lowpass filter\n- random power ([BirdCLEF 2021 5th place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351))\n- torchaudio.transforms.FrequencyMasking\n- torchaudio.transforms.TimeMasking\n\n# Oversampling\n\nThe following Oversampling was performed, but I think the effect was minimal.\n\n- Oversampling to have at least 20 of data for each class\n\n- I doubled the number of training data that have two or more secondary labels because data with more than two secondary_labels have worse val_loss\n\n- As in the [BirdCLEF 2022 5th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/327044), files with reverb/echo effects created for some minority class data and added to the training data\n\n- Strong augmentation was applied to the oversampled data\n\n# Inference\n\nThe CNN input is given 5 seconds of data for inference, and the BirdNET input is given the first 3 seconds of data for inference.\n\n- To submit within 2 hours, I converted each model to ONNX\n\n- As in [this notebook](https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference), I also used ThreadPoolExecutor to speed up inference\n\n# Ensembling\n\nThe outputs of 2 models with different CNN backbones (`eca_nfnet_l1`, `seresnext26t_32x4d`) simply averaged.\n\n# What did not work\n\n- add background noise ([BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463))\n- postprocess ([BirdCLEF 2022 2nd place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/326950))\n- Speeding up inference using OpenVINO\n  - Even if the model was converted to OpenVINO, the inference time was not much different from ONNX. I think I just used it wrong.\n- Using MFCC as input (as in [this paper](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9495150/))\n- Using CQT as input\n- Using `FAMILY` etc. from `eBird_Taxonomy_v2021.csv` for prediction (as in [this paper](https://arxiv.org/pdf/2110.03209.pdf))\n\n# Rough lb history\n\n| name                                                                                                | Private Score | Public Score |\n| --------------------------------------------------------------------------------------------------- | ------------- | ------------ |\n| BirdCLEF 2021 2nd place solution CNN (with BirdCLEF 2021 + BirdCLEF 2022 pretrain)                  | 0.71          | 0.80         |\n| BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding                             | 0.74          | 0.82         |\n| BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding + Oversampling + Ensembling | 0.75          | 0.83         |",
      "votes": null
    },
    {
      "id": "2273074",
      "postDate": "05/25/2023 02:03:10",
      "content": "<p>Congratulations!, We tried something similar with BirdNet, great to see that you made it work!</p>",
      "rawMarkdown": "Congratulations!, We tried something similar with BirdNet, great to see that you made it work!",
      "votes": null
    },
    {
      "id": "2273197",
      "postDate": "05/25/2023 04:28:41",
      "content": "<p>Thank you for sharing your work, as a beginner just getting started with Deep learning I am amazed with how you can figure out so many details and approaches to solve a problem. Do you have any tips for how I can get started to efficiently solve these complex problems?</p>",
      "rawMarkdown": "Thank you for sharing your work, as a beginner just getting started with Deep learning I am amazed with how you can figure out so many details and approaches to solve a problem. Do you have any tips for how I can get started to efficiently solve these complex problems?",
      "votes": null
    },
    {
      "id": "2273254",
      "postDate": "05/25/2023 05:02:31",
      "content": "<p>Congratulations 🎉 And really thanks for sharing </p>",
      "rawMarkdown": "Congratulations 🎉 And really thanks for sharing",
      "votes": null
    },
    {
      "id": "2273255",
      "postDate": "05/25/2023 05:03:29",
      "content": "<p>Congratulations and thanks for sharing! The results for only 2 models were impressive. Better than my single model😨</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! The results for only 2 models were impressive. Better than my single model😨",
      "votes": null
    },
    {
      "id": "2273404",
      "postDate": "05/25/2023 06:29:25",
      "content": "<p>Congratulations and thanks for sharing! I remember that BirdNET is trained on sample rate 48000 audios, how do you handle the difference on sample rate? </p>",
      "rawMarkdown": "Congratulations and thanks for sharing! I remember that BirdNET is trained on sample rate 48000 audios, how do you handle the difference on sample rate?",
      "votes": null
    },
    {
      "id": "2273485",
      "postDate": "05/25/2023 07:26:26",
      "content": "<p>Great way to integrate BirdNet, congratulations on the gold! 🎉</p>",
      "rawMarkdown": "Great way to integrate BirdNet, congratulations on the gold! 🎉",
      "votes": null
    },
    {
      "id": "2274091",
      "postDate": "05/25/2023 16:01:59",
      "content": "<p>The sample rate has not been changed. The sample rate was left at its default, which resamples to 48000, because I thought that changing and loading the sample rate to 32000 would not be suitable for BirdNET's input data.</p>",
      "rawMarkdown": "The sample rate has not been changed. The sample rate was left at its default, which resamples to 48000, because I thought that changing and loading the sample rate to 32000 would not be suitable for BirdNET's input data.",
      "votes": null
    },
    {
      "id": "2274127",
      "postDate": "05/25/2023 16:32:54",
      "content": "<p>I still don't know how to solve these problems efficiently. In my case, I carefully examine the data, think of hypotheses that might solve the problem, and experiment while gradually improving the model hundreds of times. It's through this process of trial and error that a good solution is found, albeit very rarely.</p>",
      "rawMarkdown": "I still don't know how to solve these problems efficiently. In my case, I carefully examine the data, think of hypotheses that might solve the problem, and experiment while gradually improving the model hundreds of times. It's through this process of trial and error that a good solution is found, albeit very rarely.",
      "votes": null
    },
    {
      "id": "2274225",
      "postDate": "05/25/2023 17:56:57",
      "content": "<p>Understood😊 I will keep on learning and experimenting and once again thank you for sharing your approach.</p>",
      "rawMarkdown": "Understood😊 I will keep on learning and experimenting and once again thank you for sharing your approach.",
      "votes": null
    },
    {
      "id": "2274491",
      "postDate": "05/26/2023 03:43:26",
      "content": "<p>Congrats on getting 6th position and for getting gold!!!</p>",
      "rawMarkdown": "Congrats on getting 6th position and for getting gold!!!",
      "votes": null
    },
    {
      "id": "2286247",
      "postDate": "06/03/2023 10:27:14",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/anonamename\" target=\"_blank\">@anonamename</a> and congrats on the solo gold. Thanks for sharing the writeup 🎉</p>",
      "rawMarkdown": "Great work @anonamename and congrats on the solo gold. Thanks for sharing the writeup 🎉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2273074,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "05/25/2023 02:03:10",
      "content": "<p>Congratulations!, We tried something similar with BirdNet, great to see that you made it work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273197,
      "author_name": "aadityabansalcodes",
      "author_url": "",
      "post_date": "05/25/2023 04:28:41",
      "content": "<p>Thank you for sharing your work, as a beginner just getting started with Deep learning I am amazed with how you can figure out so many details and approaches to solve a problem. Do you have any tips for how I can get started to efficiently solve these complex problems?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2274127,
          "author_name": "anonamename",
          "author_url": "",
          "post_date": "05/25/2023 16:32:54",
          "content": "<p>I still don't know how to solve these problems efficiently. In my case, I carefully examine the data, think of hypotheses that might solve the problem, and experiment while gradually improving the model hundreds of times. It's through this process of trial and error that a good solution is found, albeit very rarely.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2274225,
              "author_name": "aadityabansalcodes",
              "author_url": "",
              "post_date": "05/25/2023 17:56:57",
              "content": "<p>Understood😊 I will keep on learning and experimenting and once again thank you for sharing your approach.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2273254,
      "author_name": "scipygaurav",
      "author_url": "",
      "post_date": "05/25/2023 05:02:31",
      "content": "<p>Congratulations 🎉 And really thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273255,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "05/25/2023 05:03:29",
      "content": "<p>Congratulations and thanks for sharing! The results for only 2 models were impressive. Better than my single model😨</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2273404,
      "author_name": "aphysict",
      "author_url": "",
      "post_date": "05/25/2023 06:29:25",
      "content": "<p>Congratulations and thanks for sharing! I remember that BirdNET is trained on sample rate 48000 audios, how do you handle the difference on sample rate? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2274091,
          "author_name": "anonamename",
          "author_url": "",
          "post_date": "05/25/2023 16:01:59",
          "content": "<p>The sample rate has not been changed. The sample rate was left at its default, which resamples to 48000, because I thought that changing and loading the sample rate to 32000 would not be suitable for BirdNET's input data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2273485,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "05/25/2023 07:26:26",
      "content": "<p>Great way to integrate BirdNet, congratulations on the gold! 🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2274491,
      "author_name": "ayushjha7",
      "author_url": "",
      "post_date": "05/26/2023 03:43:26",
      "content": "<p>Congrats on getting 6th position and for getting gold!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2286247,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 10:27:14",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/anonamename\" target=\"_blank\">@anonamename</a> and congrats on the solo gold. Thanks for sharing the writeup 🎉</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2273019": "Thank you to the host and Kaggle for organizing the competition.\n\nMy solution is a combination of the embedding vectors of [BirdNET-Analyzer V2.2](https://github.com/kahst/BirdNET-Analyzer/tree/d1f5a9c015d4419277cbb285e89d3f843a6bab49) and the CNN from the [BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463).\n\n## Model (BirdNET embedding + CNN)\n\nIn order to utilize the features of BirdNET, the embedding of BirdNET is concatenated to the output of the CNN of the [BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1694617%2Fe7dcc6f87186d43df680592acb3f91d7%2Fmy_solution.png?generation=1684975539169704&alt=media)\n\n### BirdNET embedding\n\nThe embedding vectors of BirdNET were created with [BirdNET V2.2 embeddings.py](https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/embeddings.py). I modified [BirdNET V2.2 audio.py](https://github.com/kahst/BirdNET-Analyzer/blob/d1f5a9c015d4419277cbb285e89d3f843a6bab49/audio.py#L7) so that if the length of the audio data is shorter than BirdNET's sample rate (48000), the data is padded to output at least 1 second of embedding vector. Using V2.2 rather than the latest version BirdNET V2.3, slightly improved cv.\n\n### CNN\n\nThe backbone of the CNN from the BirdCLEF 2021 2nd place solution used timm's `eca_nfnet_l1` and `seresnext26t_32x4d`.\n\n## Training\n\nAfter pretraining with data from BirdCLEF 2021 + BirdCLEF 2022, I train the CNN and other linear layers with data from BirdCLEF 2023. BirdNET is not trained.\n\nThe input for training is data for 30 seconds. Since BirdNET outputs embedding vectors for 3 seconds of data, I averaged each of the embedding vectors for 30 seconds.\n\nIn most experiments, cv was highest in the final epoch, so I included all data in the training set.\n\nThe main training parameters are as follows:\n\n- loss: BCEWithLogitsLoss\n\n- MelSpectrogram\n  \n  - sample_rate: 32000, window_size: 1024, hop_size: 320, fmin: 0, fmax: 14000, mel_bins: 128, power: 2, top_db=None\n\n- labels: primary label=0.9995, secondary label=0.4 or 0.5 ([BirdCLEF 2022 3rd place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/327193)), label smoothing=0.01\n\n- cv: primary_label StratifiedKFold(n_splits=5)\n\n- optimizer: AdamW(weight_decay=1e-4)\n\n- scheduler: warmup 0-3epoch(lr=3e-6->3e-4) + Cosine Annealing 3-70epoch(lr=3e-4->3e-6)\n\n# Augmentation\n\nAs in the previous two competitions, augmentation was important. I combined the following augmentations:\n\n- audiomentations.Shift\n- Mixup ([BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463))\n- Gaussian Noise\n- random lowpass filter\n- random power ([BirdCLEF 2021 5th place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243351))\n- torchaudio.transforms.FrequencyMasking\n- torchaudio.transforms.TimeMasking\n\n# Oversampling\n\nThe following Oversampling was performed, but I think the effect was minimal.\n\n- Oversampling to have at least 20 of data for each class\n\n- I doubled the number of training data that have two or more secondary labels because data with more than two secondary_labels have worse val_loss\n\n- As in the [BirdCLEF 2022 5th place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/327044), files with reverb/echo effects created for some minority class data and added to the training data\n\n- Strong augmentation was applied to the oversampled data\n\n# Inference\n\nThe CNN input is given 5 seconds of data for inference, and the BirdNET input is given the first 3 seconds of data for inference.\n\n- To submit within 2 hours, I converted each model to ONNX\n\n- As in [this notebook](https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference), I also used ThreadPoolExecutor to speed up inference\n\n# Ensembling\n\nThe outputs of 2 models with different CNN backbones (`eca_nfnet_l1`, `seresnext26t_32x4d`) simply averaged.\n\n# What did not work\n\n- add background noise ([BirdCLEF 2021 2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463))\n- postprocess ([BirdCLEF 2022 2nd place solution](https://www.kaggle.com/competitions/birdclef-2022/discussion/326950))\n- Speeding up inference using OpenVINO\n  - Even if the model was converted to OpenVINO, the inference time was not much different from ONNX. I think I just used it wrong.\n- Using MFCC as input (as in [this paper](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9495150/))\n- Using CQT as input\n- Using `FAMILY` etc. from `eBird_Taxonomy_v2021.csv` for prediction (as in [this paper](https://arxiv.org/pdf/2110.03209.pdf))\n\n# Rough lb history\n\n| name                                                                                                | Private Score | Public Score |\n| --------------------------------------------------------------------------------------------------- | ------------- | ------------ |\n| BirdCLEF 2021 2nd place solution CNN (with BirdCLEF 2021 + BirdCLEF 2022 pretrain)                  | 0.71          | 0.80         |\n| BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding                             | 0.74          | 0.82         |\n| BirdCLEF 2021 2nd place solution CNN + Augmentation + BirdNET embedding + Oversampling + Ensembling | 0.75          | 0.83         |",
    "2273074": "Congratulations!, We tried something similar with BirdNet, great to see that you made it work!",
    "2273197": "Thank you for sharing your work, as a beginner just getting started with Deep learning I am amazed with how you can figure out so many details and approaches to solve a problem. Do you have any tips for how I can get started to efficiently solve these complex problems?",
    "2273254": "Congratulations 🎉 And really thanks for sharing",
    "2273255": "Congratulations and thanks for sharing! The results for only 2 models were impressive. Better than my single model😨",
    "2273404": "Congratulations and thanks for sharing! I remember that BirdNET is trained on sample rate 48000 audios, how do you handle the difference on sample rate?",
    "2273485": "Great way to integrate BirdNet, congratulations on the gold! 🎉",
    "2274091": "The sample rate has not been changed. The sample rate was left at its default, which resamples to 48000, because I thought that changing and loading the sample rate to 32000 would not be suitable for BirdNET's input data.",
    "2274127": "I still don't know how to solve these problems efficiently. In my case, I carefully examine the data, think of hypotheses that might solve the problem, and experiment while gradually improving the model hundreds of times. It's through this process of trial and error that a good solution is found, albeit very rarely.",
    "2274225": "Understood😊 I will keep on learning and experimenting and once again thank you for sharing your approach.",
    "2274491": "Congrats on getting 6th position and for getting gold!!!",
    "2286247": "Great work @anonamename and congrats on the solo gold. Thanks for sharing the writeup 🎉"
  },
  "source": "meta"
}