{
  "id": 412742,
  "title": "20th place solution: SED + CNN ensemble using onnx",
  "url": "/competitions/birdclef-2023/writeups/yokuyama-moritake04-nyamoke-20th-place-solution-se",
  "author_name": "",
  "post_date": "2023-05-29T12:31:22.690Z",
  "votes": 23,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First, I would like to thank the competition hosts for organizing this competition and also the participants.</p>\n<p>I would especially like to express a huge thank you to my teammates <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a>, <a href=\"https://www.kaggle.com/yiiino\" target=\"_blank\">@yiiino</a>, <a href=\"https://www.kaggle.com/mobykjr\" target=\"_blank\">@mobykjr</a>, and <a href=\"https://www.kaggle.com/vupasama\" target=\"_blank\">@vupasama</a> !!!</p>\n<h2>Overview</h2>\n<p>My final submission is an ensemble of the following 5 models.</p>\n<ul>\n<li>eca_nfnet_l0 (SED)</li>\n<li>eca_nfnet_l0 (Simple CNN)</li>\n<li>tf_efficientnetv2_b0 (SED)</li>\n<li>tf_efficientnet_b0.ns (Simple CNN)</li>\n<li>tf_mobilenetv3_large_100 (Simple CNN)</li>\n</ul>\n<p>These models were accelerated using onnx.</p>\n<h2>Our Approach</h2>\n<ul>\n<li>Model<ul>\n<li>We used two main model structures: a SED and a simple CNN.</li>\n<li>For the SED model, We created it based on <a href=\"https://www.kaggle.com/kaerurururu\" target=\"_blank\">@kaerurururu</a>'s notebook from BirdCLEF 2022.<br>\n<a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0</a></li></ul></li>\n<li>Audio<ul>\n<li>Torchaudio was used.</li>\n<li>Waveform<ul>\n<li>duration: 5s</li>\n<li>sample rate: 32000</li>\n<li>Secondary label was used (used with soft label at 0.5)</li>\n<li>If the audio was less than 5 seconds, the audio was repeated and concatenated until it exceeds 5 seconds</li>\n<li>Used the first 5s of the audio during validation.</li>\n<li>Normalized with audiomentations.Normalize</li></ul></li>\n<li>Mel Spectrogram<ul>\n<li>n_fft: 2048</li>\n<li>win_length: 2048</li>\n<li>hop_length: 320</li>\n<li>f_min: 50</li>\n<li>f_max: 14000</li>\n<li>n_mels: 128</li>\n<li>top_db: 80</li>\n<li>Standardized using the mean and standard deviation of the ImageNet.</li></ul></li></ul></li>\n<li>Model Parameters<ul>\n<li>drop_rate: 0.5</li>\n<li>drop_path_rate: 0.2</li>\n<li>criterion: BCEWithLogitsLoss (label smoothing: 0.0025)</li>\n<li>optimizer: AdamW (lr: 1.0e-3, weight_decay: 1.0e-2)</li>\n<li>scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-3)</li>\n<li>epoch: 30</li>\n<li>batch_size: 32</li></ul></li>\n<li>Data Augmentation<ul>\n<li>waveform<ul>\n<li>During training, we extracted random crops or the first 5s of the audio from the total audio. Each was done at a ratio of 50% each.</li>\n<li>audiomentations were used.<ul>\n<li>audiomentations.AddBackgroundNoise</li>\n<li>audiomentations.AddGaussianSNR</li>\n<li>audiomentations.AddGaussianNoise</li></ul></li>\n<li>BackgroundNoise was based on the first-place solution from Birdclef2022. We used this dataset.<br>\n<a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise</a></li></ul></li>\n<li>melspectrogram<ul>\n<li>Torchaudio masking was used.<ul>\n<li>torchaudio.transforms.FrequencyMasking(<br>\nfreq_mask_param=mel_specgram.shape[1] // 5<br>\n)</li>\n<li>torchaudio.transforms.TimeMasking(<br>\ntime_mask_param=mel_specgram.shape[2] // 5<br>\n)</li></ul></li>\n<li>mixup (alpha = 0.5), cutmix (alpha = 0.5)</li></ul></li></ul></li>\n<li>Pretraining<ul>\n<li>Data from Birdclef 2021 and 2022 were used for pre-training.</li>\n<li>Pre-training was done with a duration of 15s and no oversampling.<ul>\n<li>Duration of 15s seemed to be better than 5s for pre-training.</li></ul></li></ul></li>\n<li>Dealing with Imbalance Data<ul>\n<li>We oversampled and set up each class to have at least 50 data.</li></ul></li>\n<li>Adding Nocall Data<ul>\n<li>The nocall data used the background noise audio used in augmentation. The nocall data was created with all labels set to 0.</li></ul></li>\n<li>post-processing<ul>\n<li>For the SED model, the output was averaged framewise and clipwise.</li></ul></li>\n<li>Inference Acceleration<ul>\n<li>We used onnx, which was considerably faster than torchscript.</li>\n<li>Data preprocessing was standardized so that the same preprocessed data could be input into multiple models.</li>\n<li>Data was preprocessed in advance and stored in a dict. Parallel processing was used for this preprocessing.</li></ul></li>\n<li>CV Strategy<ul>\n<li>StratifiedKFold (n=5, stratify=primary_label)</li>\n<li>All bird species were included in the training data.</li></ul></li>\n</ul>\n<h2>What Worked</h2>\n<ul>\n<li>oversampling</li>\n<li>Random cropping or extraction of the first 5s of audio from the total audio during training (each executed at a ratio of 50%).</li>\n<li>The following augmentation<ul>\n<li>Background noise, Gaussian noise , Mask melspectrogram</li></ul></li>\n<li>SED model</li>\n<li>label smoothing</li>\n<li>Adding Nocall Data</li>\n<li>In the pre-training, increase the duration</li>\n<li>Ensemble<ul>\n<li>The SED model and simple CNN ensemble were good.</li></ul></li>\n<li>secondary label</li>\n<li>Set up many numworkers in the data loader → learning speed is greatly increased!</li>\n</ul>\n<h2>What Didn’t Worked</h2>\n<ul>\n<li>The following augmentation<ul>\n<li>Pink noise, Random volume</li></ul></li>\n<li>Focal Loss</li>\n<li>Long durations (15s, 30s, etc.) when training on 2023 data.</li>\n<li>Dealing with previous and next seconds of audio data in the submission<ul>\n<li>Use 5 seconds before and after for a total of 15 seconds to create a mel spectrogram and predict → Notebook Timeout</li>\n<li>Moving average, moving weighted average</li></ul></li>\n<li>Large Models</li>\n<li>geometric mean</li>\n<li>LSTM head</li>\n<li>strong dropout</li>\n<li>rank ensemble</li>\n<li>hand labeling<ul>\n<li>It is thought to have led to overfitting…</li></ul></li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li>GitHub (My training code) →  <a href=\"https://github.com/moritake04/birdclef-2023\" target=\"_blank\">https://github.com/moritake04/birdclef-2023</a></li>\n</ul>",
  "messages": [
    {
      "id": "2273147",
      "postDate": "05/25/2023 03:31:58",
      "content": "<p>First, I would like to thank the competition hosts for organizing this competition and also the participants.</p>\n<p>I would especially like to express a huge thank you to my teammates <a href=\"https://www.kaggle.com/yokuyama\" target=\"_blank\">@yokuyama</a>, <a href=\"https://www.kaggle.com/yiiino\" target=\"_blank\">@yiiino</a>, <a href=\"https://www.kaggle.com/mobykjr\" target=\"_blank\">@mobykjr</a>, and <a href=\"https://www.kaggle.com/vupasama\" target=\"_blank\">@vupasama</a> !!!</p>\n<h2>Overview</h2>\n<p>My final submission is an ensemble of the following 5 models.</p>\n<ul>\n<li>eca_nfnet_l0 (SED)</li>\n<li>eca_nfnet_l0 (Simple CNN)</li>\n<li>tf_efficientnetv2_b0 (SED)</li>\n<li>tf_efficientnet_b0.ns (Simple CNN)</li>\n<li>tf_mobilenetv3_large_100 (Simple CNN)</li>\n</ul>\n<p>These models were accelerated using onnx.</p>\n<h2>Our Approach</h2>\n<ul>\n<li>Model<ul>\n<li>We used two main model structures: a SED and a simple CNN.</li>\n<li>For the SED model, We created it based on <a href=\"https://www.kaggle.com/kaerurururu\" target=\"_blank\">@kaerurururu</a>'s notebook from BirdCLEF 2022.<br>\n<a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0</a></li></ul></li>\n<li>Audio<ul>\n<li>Torchaudio was used.</li>\n<li>Waveform<ul>\n<li>duration: 5s</li>\n<li>sample rate: 32000</li>\n<li>Secondary label was used (used with soft label at 0.5)</li>\n<li>If the audio was less than 5 seconds, the audio was repeated and concatenated until it exceeds 5 seconds</li>\n<li>Used the first 5s of the audio during validation.</li>\n<li>Normalized with audiomentations.Normalize</li></ul></li>\n<li>Mel Spectrogram<ul>\n<li>n_fft: 2048</li>\n<li>win_length: 2048</li>\n<li>hop_length: 320</li>\n<li>f_min: 50</li>\n<li>f_max: 14000</li>\n<li>n_mels: 128</li>\n<li>top_db: 80</li>\n<li>Standardized using the mean and standard deviation of the ImageNet.</li></ul></li></ul></li>\n<li>Model Parameters<ul>\n<li>drop_rate: 0.5</li>\n<li>drop_path_rate: 0.2</li>\n<li>criterion: BCEWithLogitsLoss (label smoothing: 0.0025)</li>\n<li>optimizer: AdamW (lr: 1.0e-3, weight_decay: 1.0e-2)</li>\n<li>scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-3)</li>\n<li>epoch: 30</li>\n<li>batch_size: 32</li></ul></li>\n<li>Data Augmentation<ul>\n<li>waveform<ul>\n<li>During training, we extracted random crops or the first 5s of the audio from the total audio. Each was done at a ratio of 50% each.</li>\n<li>audiomentations were used.<ul>\n<li>audiomentations.AddBackgroundNoise</li>\n<li>audiomentations.AddGaussianSNR</li>\n<li>audiomentations.AddGaussianNoise</li></ul></li>\n<li>BackgroundNoise was based on the first-place solution from Birdclef2022. We used this dataset.<br>\n<a href=\"https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise\" target=\"_blank\">https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise</a></li></ul></li>\n<li>melspectrogram<ul>\n<li>Torchaudio masking was used.<ul>\n<li>torchaudio.transforms.FrequencyMasking(<br>\nfreq_mask_param=mel_specgram.shape[1] // 5<br>\n)</li>\n<li>torchaudio.transforms.TimeMasking(<br>\ntime_mask_param=mel_specgram.shape[2] // 5<br>\n)</li></ul></li>\n<li>mixup (alpha = 0.5), cutmix (alpha = 0.5)</li></ul></li></ul></li>\n<li>Pretraining<ul>\n<li>Data from Birdclef 2021 and 2022 were used for pre-training.</li>\n<li>Pre-training was done with a duration of 15s and no oversampling.<ul>\n<li>Duration of 15s seemed to be better than 5s for pre-training.</li></ul></li></ul></li>\n<li>Dealing with Imbalance Data<ul>\n<li>We oversampled and set up each class to have at least 50 data.</li></ul></li>\n<li>Adding Nocall Data<ul>\n<li>The nocall data used the background noise audio used in augmentation. The nocall data was created with all labels set to 0.</li></ul></li>\n<li>post-processing<ul>\n<li>For the SED model, the output was averaged framewise and clipwise.</li></ul></li>\n<li>Inference Acceleration<ul>\n<li>We used onnx, which was considerably faster than torchscript.</li>\n<li>Data preprocessing was standardized so that the same preprocessed data could be input into multiple models.</li>\n<li>Data was preprocessed in advance and stored in a dict. Parallel processing was used for this preprocessing.</li></ul></li>\n<li>CV Strategy<ul>\n<li>StratifiedKFold (n=5, stratify=primary_label)</li>\n<li>All bird species were included in the training data.</li></ul></li>\n</ul>\n<h2>What Worked</h2>\n<ul>\n<li>oversampling</li>\n<li>Random cropping or extraction of the first 5s of audio from the total audio during training (each executed at a ratio of 50%).</li>\n<li>The following augmentation<ul>\n<li>Background noise, Gaussian noise , Mask melspectrogram</li></ul></li>\n<li>SED model</li>\n<li>label smoothing</li>\n<li>Adding Nocall Data</li>\n<li>In the pre-training, increase the duration</li>\n<li>Ensemble<ul>\n<li>The SED model and simple CNN ensemble were good.</li></ul></li>\n<li>secondary label</li>\n<li>Set up many numworkers in the data loader → learning speed is greatly increased!</li>\n</ul>\n<h2>What Didn’t Worked</h2>\n<ul>\n<li>The following augmentation<ul>\n<li>Pink noise, Random volume</li></ul></li>\n<li>Focal Loss</li>\n<li>Long durations (15s, 30s, etc.) when training on 2023 data.</li>\n<li>Dealing with previous and next seconds of audio data in the submission<ul>\n<li>Use 5 seconds before and after for a total of 15 seconds to create a mel spectrogram and predict → Notebook Timeout</li>\n<li>Moving average, moving weighted average</li></ul></li>\n<li>Large Models</li>\n<li>geometric mean</li>\n<li>LSTM head</li>\n<li>strong dropout</li>\n<li>rank ensemble</li>\n<li>hand labeling<ul>\n<li>It is thought to have led to overfitting…</li></ul></li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li>GitHub (My training code) →  <a href=\"https://github.com/moritake04/birdclef-2023\" target=\"_blank\">https://github.com/moritake04/birdclef-2023</a></li>\n</ul>",
      "rawMarkdown": "First, I would like to thank the competition hosts for organizing this competition and also the participants.\n\nI would especially like to express a huge thank you to my teammates @yokuyama, @yiiino, @mobykjr, and @vupasama !!!\n\n## Overview\n\nMy final submission is an ensemble of the following 5 models.\n\n- eca_nfnet_l0 (SED)\n- eca_nfnet_l0 (Simple CNN)\n- tf_efficientnetv2_b0 (SED)\n- tf_efficientnet_b0.ns (Simple CNN)\n- tf_mobilenetv3_large_100 (Simple CNN)\n\nThese models were accelerated using onnx.\n\n## Our Approach\n\n- Model\n    - We used two main model structures: a SED and a simple CNN.\n    - For the SED model, We created it based on @kaerurururu's notebook from BirdCLEF 2022.\n    [https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0)\n- Audio\n    - Torchaudio was used.\n    - Waveform\n        - duration: 5s\n        - sample rate: 32000\n        - Secondary label was used (used with soft label at 0.5)\n        - If the audio was less than 5 seconds, the audio was repeated and concatenated until it exceeds 5 seconds\n        - Used the first 5s of the audio during validation.\n        - Normalized with audiomentations.Normalize\n    - Mel Spectrogram\n        - n_fft: 2048\n        - win_length: 2048\n        - hop_length: 320\n        - f_min: 50\n        - f_max: 14000\n        - n_mels: 128\n        - top_db: 80\n        - Standardized using the mean and standard deviation of the ImageNet.\n- Model Parameters\n    - drop_rate: 0.5\n    - drop_path_rate: 0.2\n    - criterion: BCEWithLogitsLoss (label smoothing: 0.0025)\n    - optimizer: AdamW (lr: 1.0e-3, weight_decay: 1.0e-2)\n    - scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-3)\n    - epoch: 30\n    - batch_size: 32\n- Data Augmentation\n    - waveform\n        - During training, we extracted random crops or the first 5s of the audio from the total audio. Each was done at a ratio of 50% each.\n        - audiomentations were used.\n            - audiomentations.AddBackgroundNoise\n            - audiomentations.AddGaussianSNR\n            - audiomentations.AddGaussianNoise\n        - BackgroundNoise was based on the first-place solution from Birdclef2022. We used this dataset.\n            [https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise)\n    - melspectrogram\n        - Torchaudio masking was used.\n            - torchaudio.transforms.FrequencyMasking(\n        freq_mask_param=mel_specgram.shape[1] // 5\n        )\n            - torchaudio.transforms.TimeMasking(\n        time_mask_param=mel_specgram.shape[2] // 5\n        )\n        - mixup (alpha = 0.5), cutmix (alpha = 0.5)\n- Pretraining\n    - Data from Birdclef 2021 and 2022 were used for pre-training.\n    - Pre-training was done with a duration of 15s and no oversampling.\n        - Duration of 15s seemed to be better than 5s for pre-training.\n- Dealing with Imbalance Data\n    - We oversampled and set up each class to have at least 50 data.\n- Adding Nocall Data\n    - The nocall data used the background noise audio used in augmentation. The nocall data was created with all labels set to 0.\n- post-processing\n    - For the SED model, the output was averaged framewise and clipwise.\n- Inference Acceleration\n    - We used onnx, which was considerably faster than torchscript.\n    - Data preprocessing was standardized so that the same preprocessed data could be input into multiple models.\n    - Data was preprocessed in advance and stored in a dict. Parallel processing was used for this preprocessing.\n- CV Strategy\n    - StratifiedKFold (n=5, stratify=primary_label)\n    - All bird species were included in the training data.\n\n## What Worked\n\n- oversampling\n- Random cropping or extraction of the first 5s of audio from the total audio during training (each executed at a ratio of 50%).\n- The following augmentation\n    - Background noise, Gaussian noise , Mask melspectrogram\n- SED model\n- label smoothing\n- Adding Nocall Data\n- In the pre-training, increase the duration\n- Ensemble\n    - The SED model and simple CNN ensemble were good.\n- secondary label\n- Set up many numworkers in the data loader → learning speed is greatly increased!\n\n## What Didn’t Worked\n\n- The following augmentation\n    - Pink noise, Random volume\n- Focal Loss\n- Long durations (15s, 30s, etc.) when training on 2023 data.\n- Dealing with previous and next seconds of audio data in the submission\n    - Use 5 seconds before and after for a total of 15 seconds to create a mel spectrogram and predict → Notebook Timeout\n    - Moving average, moving weighted average\n- Large Models\n- geometric mean\n- LSTM head\n- strong dropout\n- rank ensemble\n- hand labeling\n    - It is thought to have led to overfitting…\n\n## Code\n\n- GitHub (My training code) →  https://github.com/moritake04/birdclef-2023",
      "votes": null
    },
    {
      "id": "2273542",
      "postDate": "05/25/2023 08:13:01",
      "content": "<p>Excellent Approach! Congratulations <a href=\"https://www.kaggle.com/moritake04\" target=\"_blank\">@moritake04</a> 🔥</p>",
      "rawMarkdown": "Excellent Approach! Congratulations @moritake04 🔥",
      "votes": null
    },
    {
      "id": "2276231",
      "postDate": "05/26/2023 15:34:28",
      "content": "<p>Thanks for the explanation, and especially for posting all your training code on github!</p>",
      "rawMarkdown": "Thanks for the explanation, and especially for posting all your training code on github!",
      "votes": null
    },
    {
      "id": "2276326",
      "postDate": "05/26/2023 17:56:22",
      "content": "<p>Congratulations on becoming a competition Master, and thanks for sharing such a detailed explanation and the training code. Interesting idea the 2 outputs average on the SED model.</p>",
      "rawMarkdown": "Congratulations on becoming a competition Master, and thanks for sharing such a detailed explanation and the training code. Interesting idea the 2 outputs average on the SED model.",
      "votes": null
    },
    {
      "id": "2276755",
      "postDate": "05/27/2023 07:24:21",
      "content": "<p>Thank you very much!<br>\n\"the output was averaged framewise and clipwise\" showed slight improvement in LB.</p>",
      "rawMarkdown": "Thank you very much!\n\"the output was averaged framewise and clipwise\" showed slight improvement in LB.",
      "votes": null
    },
    {
      "id": "2282293",
      "postDate": "05/31/2023 13:59:26",
      "content": "<p>Did you use external GPUs?, what kind?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Did you use external GPUs?, what kind?\n\nThank you!",
      "votes": null
    },
    {
      "id": "2286021",
      "postDate": "06/03/2023 07:10:43",
      "content": "<p>I used paperspace gradient growth plan. Mainly used A6000.<br>\n<a href=\"https://www.paperspace.com/gradient/pricing\" target=\"_blank\">https://www.paperspace.com/gradient/pricing</a></p>",
      "rawMarkdown": "I used paperspace gradient growth plan. Mainly used A6000.\nhttps://www.paperspace.com/gradient/pricing",
      "votes": null
    },
    {
      "id": "2286261",
      "postDate": "06/03/2023 10:39:22",
      "content": "<p>Congrats and thanks for sharing detailed writeup + code 🎉</p>",
      "rawMarkdown": "Congrats and thanks for sharing detailed writeup + code 🎉",
      "votes": null
    },
    {
      "id": "2291310",
      "postDate": "06/07/2023 13:16:14",
      "content": "<p>Any cost reference for this competition in paperspace?. And CPU/GPU that you used. I'm trying to evaluate paperspace for competitions.</p>\n<p>Thanks in advance</p>",
      "rawMarkdown": "Any cost reference for this competition in paperspace?. And CPU/GPU that you used. I'm trying to evaluate paperspace for competitions.\n\nThanks in advance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2273542,
      "author_name": "scipygaurav",
      "author_url": "",
      "post_date": "05/25/2023 08:13:01",
      "content": "<p>Excellent Approach! Congratulations <a href=\"https://www.kaggle.com/moritake04\" target=\"_blank\">@moritake04</a> 🔥</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2276231,
      "author_name": "janhuus",
      "author_url": "",
      "post_date": "05/26/2023 15:34:28",
      "content": "<p>Thanks for the explanation, and especially for posting all your training code on github!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2276326,
      "author_name": "maxdiazbattan",
      "author_url": "",
      "post_date": "05/26/2023 17:56:22",
      "content": "<p>Congratulations on becoming a competition Master, and thanks for sharing such a detailed explanation and the training code. Interesting idea the 2 outputs average on the SED model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2276755,
          "author_name": "moritake04",
          "author_url": "",
          "post_date": "05/27/2023 07:24:21",
          "content": "<p>Thank you very much!<br>\n\"the output was averaged framewise and clipwise\" showed slight improvement in LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2282293,
      "author_name": "pablolarrosa",
      "author_url": "",
      "post_date": "05/31/2023 13:59:26",
      "content": "<p>Did you use external GPUs?, what kind?</p>\n<p>Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2286021,
          "author_name": "moritake04",
          "author_url": "",
          "post_date": "06/03/2023 07:10:43",
          "content": "<p>I used paperspace gradient growth plan. Mainly used A6000.<br>\n<a href=\"https://www.paperspace.com/gradient/pricing\" target=\"_blank\">https://www.paperspace.com/gradient/pricing</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2291310,
              "author_name": "pablolarrosa",
              "author_url": "",
              "post_date": "06/07/2023 13:16:14",
              "content": "<p>Any cost reference for this competition in paperspace?. And CPU/GPU that you used. I'm trying to evaluate paperspace for competitions.</p>\n<p>Thanks in advance</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2286261,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 10:39:22",
      "content": "<p>Congrats and thanks for sharing detailed writeup + code 🎉</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2273147": "First, I would like to thank the competition hosts for organizing this competition and also the participants.\n\nI would especially like to express a huge thank you to my teammates @yokuyama, @yiiino, @mobykjr, and @vupasama !!!\n\n## Overview\n\nMy final submission is an ensemble of the following 5 models.\n\n- eca_nfnet_l0 (SED)\n- eca_nfnet_l0 (Simple CNN)\n- tf_efficientnetv2_b0 (SED)\n- tf_efficientnet_b0.ns (Simple CNN)\n- tf_mobilenetv3_large_100 (Simple CNN)\n\nThese models were accelerated using onnx.\n\n## Our Approach\n\n- Model\n    - We used two main model structures: a SED and a simple CNN.\n    - For the SED model, We created it based on @kaerurururu's notebook from BirdCLEF 2022.\n    [https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0)\n- Audio\n    - Torchaudio was used.\n    - Waveform\n        - duration: 5s\n        - sample rate: 32000\n        - Secondary label was used (used with soft label at 0.5)\n        - If the audio was less than 5 seconds, the audio was repeated and concatenated until it exceeds 5 seconds\n        - Used the first 5s of the audio during validation.\n        - Normalized with audiomentations.Normalize\n    - Mel Spectrogram\n        - n_fft: 2048\n        - win_length: 2048\n        - hop_length: 320\n        - f_min: 50\n        - f_max: 14000\n        - n_mels: 128\n        - top_db: 80\n        - Standardized using the mean and standard deviation of the ImageNet.\n- Model Parameters\n    - drop_rate: 0.5\n    - drop_path_rate: 0.2\n    - criterion: BCEWithLogitsLoss (label smoothing: 0.0025)\n    - optimizer: AdamW (lr: 1.0e-3, weight_decay: 1.0e-2)\n    - scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-3)\n    - epoch: 30\n    - batch_size: 32\n- Data Augmentation\n    - waveform\n        - During training, we extracted random crops or the first 5s of the audio from the total audio. Each was done at a ratio of 50% each.\n        - audiomentations were used.\n            - audiomentations.AddBackgroundNoise\n            - audiomentations.AddGaussianSNR\n            - audiomentations.AddGaussianNoise\n        - BackgroundNoise was based on the first-place solution from Birdclef2022. We used this dataset.\n            [https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise)\n    - melspectrogram\n        - Torchaudio masking was used.\n            - torchaudio.transforms.FrequencyMasking(\n        freq_mask_param=mel_specgram.shape[1] // 5\n        )\n            - torchaudio.transforms.TimeMasking(\n        time_mask_param=mel_specgram.shape[2] // 5\n        )\n        - mixup (alpha = 0.5), cutmix (alpha = 0.5)\n- Pretraining\n    - Data from Birdclef 2021 and 2022 were used for pre-training.\n    - Pre-training was done with a duration of 15s and no oversampling.\n        - Duration of 15s seemed to be better than 5s for pre-training.\n- Dealing with Imbalance Data\n    - We oversampled and set up each class to have at least 50 data.\n- Adding Nocall Data\n    - The nocall data used the background noise audio used in augmentation. The nocall data was created with all labels set to 0.\n- post-processing\n    - For the SED model, the output was averaged framewise and clipwise.\n- Inference Acceleration\n    - We used onnx, which was considerably faster than torchscript.\n    - Data preprocessing was standardized so that the same preprocessed data could be input into multiple models.\n    - Data was preprocessed in advance and stored in a dict. Parallel processing was used for this preprocessing.\n- CV Strategy\n    - StratifiedKFold (n=5, stratify=primary_label)\n    - All bird species were included in the training data.\n\n## What Worked\n\n- oversampling\n- Random cropping or extraction of the first 5s of audio from the total audio during training (each executed at a ratio of 50%).\n- The following augmentation\n    - Background noise, Gaussian noise , Mask melspectrogram\n- SED model\n- label smoothing\n- Adding Nocall Data\n- In the pre-training, increase the duration\n- Ensemble\n    - The SED model and simple CNN ensemble were good.\n- secondary label\n- Set up many numworkers in the data loader → learning speed is greatly increased!\n\n## What Didn’t Worked\n\n- The following augmentation\n    - Pink noise, Random volume\n- Focal Loss\n- Long durations (15s, 30s, etc.) when training on 2023 data.\n- Dealing with previous and next seconds of audio data in the submission\n    - Use 5 seconds before and after for a total of 15 seconds to create a mel spectrogram and predict → Notebook Timeout\n    - Moving average, moving weighted average\n- Large Models\n- geometric mean\n- LSTM head\n- strong dropout\n- rank ensemble\n- hand labeling\n    - It is thought to have led to overfitting…\n\n## Code\n\n- GitHub (My training code) →  https://github.com/moritake04/birdclef-2023",
    "2273542": "Excellent Approach! Congratulations @moritake04 🔥",
    "2276231": "Thanks for the explanation, and especially for posting all your training code on github!",
    "2276326": "Congratulations on becoming a competition Master, and thanks for sharing such a detailed explanation and the training code. Interesting idea the 2 outputs average on the SED model.",
    "2276755": "Thank you very much!\n\"the output was averaged framewise and clipwise\" showed slight improvement in LB.",
    "2282293": "Did you use external GPUs?, what kind?\n\nThank you!",
    "2286021": "I used paperspace gradient growth plan. Mainly used A6000.\nhttps://www.paperspace.com/gradient/pricing",
    "2286261": "Congrats and thanks for sharing detailed writeup + code 🎉",
    "2291310": "Any cost reference for this competition in paperspace?. And CPU/GPU that you used. I'm trying to evaluate paperspace for competitions.\n\nThanks in advance"
  },
  "source": "meta"
}