{
  "id": 395843,
  "title": "[LB: 0.80] Pretraining is All you Need",
  "url": "/competitions/birdclef-2023/discussion/395843",
  "author_name": "",
  "post_date": "2023-03-19T04:17:40.206431600Z",
  "votes": 88,
  "comment_count": 17,
  "views": 0,
  "content": "<h2>Motivation</h2>\n<p>In the current competition, improving the performance of small models has proven to be a challenging task due to only <strong>CPU</strong> inference, along with a time constraint of <strong>2 hours</strong>. Additionally, the majority of models used in this competition are pre-trained on <code>ImageNet</code> and are not familiar with <strong>audio data</strong>, making it difficult for them to achieve their true potential even with transfer learning. To overcome this challenge, there are two potential solutions: 1) pre-training on an audio dataset, and 2) training with a large-scale dataset. However, the competition dataset contains only <code>16k</code> samples, and the labels from the previous competition dataset do not match with this one, making training with a large-scale dataset challenging. As a solution, utilizing the previous competition datasets for pre-training can improve the models' feature extraction capabilities.</p>\n<h2>Processing in GPU/TPU</h2>\n<p>As previously mentioned <a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/comments#2179795\" target=\"_blank\">here</a>, <strong>model/TPU/GPU</strong> experiences a bottleneck when waiting for the data which are processed on <strong>CPU</strong>. This is because the Audio -&gt; Spectrogram transformation, Augmentation, and Normalization are performed within <code>tf.data.Dataset</code> on the CPU, creating a bottleneck. To alleviate this bottleneck, the following notebooks utilize <code>tf.keras.layers.Layer</code> to perform all operations on the <strong>GPU/TPU</strong> which can significantly speed up the training process (up to 4x in P100 GPU).</p>\n<p>To reduce the size of the notebook and make it more manageable, the <a href=\"https://github.com/awsaf49/tensorflow_extra\" target=\"_blank\"><strong>tensorflow_extra</strong></a> library is created. This library contains most of the augmentation and transformation code used in the notebooks. You are welcome to contribute. Here is a demo code that runs ops on GPU/TPU,</p>\n<pre><code> tensorflow_extra  tfe\n\nmelspec_layer = tfe.layers.MelSpectrogram()\nspecs = melspec_layer(audios)\n\ntfm_layer = tfe.layers.TimeFreqMask()\nspecs2 = tfm_layer(specs, training=)\n\nnorm_layer = tfe.layers.ZScoreMinMax()\nspecs3 = norm_layer(specs2)\n</code></pre>\n<h2>Notebooks</h2>\n<ul>\n<li>Pretraining is All you Need<ul>\n<li>Train: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-train/\" target=\"_blank\">BirdCLEF23: Pretraining is All you Need [Train]</a></li>\n<li>Infer: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-infer/\" target=\"_blank\">BirdCLEF23: Pretraining is All you Need [Infer]</a></li></ul></li>\n</ul>\n<h2>Previous Notebooks</h2>\n<ul>\n<li>EffNet + FSR + CutMixUp<ul>\n<li>Train: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-train/\" target=\"_blank\">BirdCLEF23: EffNet + FSR + CutMixUp [Train]</a></li>\n<li>Infer: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-infer/\" target=\"_blank\">BirdCLEF23: EffNet + FSR + CutMixUp [Infer]</a></li></ul></li>\n</ul>\n<h2>Training Configs</h2>\n<ul>\n<li><strong>Pretraining</strong>: ~71k samples from BirdCLEF 2021 &amp; 2022 competition. Filenames with an exact match with primary_label &amp; author are removed to avoid leak.</li>\n<li><strong>Training</strong>: ~16k samples from BirdCLEF 2023 competition.</li>\n<li><strong>Model</strong>: EfficientNetB1 + FSR</li>\n<li><strong>Training Time Duration</strong>: 10 sec</li>\n<li><strong>Image Size</strong>: 128 x 384</li>\n<li><strong>Loss</strong>: CCE</li>\n<li><strong>Sampling</strong>: Downsample data to have max 500 samples in each class &amp; Upsample data to have min 50 samples in each class.</li>\n</ul>\n<h2>Result</h2>\n<ul>\n<li><strong>Validation</strong>: 0.86 (AUC) | 0.911 (cmAP)</li>\n<li><strong>LB</strong>: 0.80</li>\n<li><strong>Improvement</strong>: Local validation -&gt; 4% | LB -&gt; 2%</li>\n</ul>\n<blockquote>\n  <p><strong>Note</strong>: CV - LB gap can be reduced by reducing time duration in training but affects the training score. As there is a small overlap of labels between previous competition and current competition data, we can use them as external data as well.</p>\n</blockquote>\n<h2>Track Experiments with WandB</h2>\n<p>You can track all the experiments on Weights &amp; Biases (WandB) <a href=\"https://wandb.ai/awsaf49/birdclef-2023-public\" target=\"_blank\">here</a><br>\n<img src=\"https://i.postimg.cc/8CsMw43R/wandb-v2.png\"></p>",
  "messages": [
    {
      "id": "2187823",
      "postDate": "03/19/2023 04:17:40",
      "content": "<h2>Motivation</h2>\n<p>In the current competition, improving the performance of small models has proven to be a challenging task due to only <strong>CPU</strong> inference, along with a time constraint of <strong>2 hours</strong>. Additionally, the majority of models used in this competition are pre-trained on <code>ImageNet</code> and are not familiar with <strong>audio data</strong>, making it difficult for them to achieve their true potential even with transfer learning. To overcome this challenge, there are two potential solutions: 1) pre-training on an audio dataset, and 2) training with a large-scale dataset. However, the competition dataset contains only <code>16k</code> samples, and the labels from the previous competition dataset do not match with this one, making training with a large-scale dataset challenging. As a solution, utilizing the previous competition datasets for pre-training can improve the models' feature extraction capabilities.</p>\n<h2>Processing in GPU/TPU</h2>\n<p>As previously mentioned <a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/comments#2179795\" target=\"_blank\">here</a>, <strong>model/TPU/GPU</strong> experiences a bottleneck when waiting for the data which are processed on <strong>CPU</strong>. This is because the Audio -&gt; Spectrogram transformation, Augmentation, and Normalization are performed within <code>tf.data.Dataset</code> on the CPU, creating a bottleneck. To alleviate this bottleneck, the following notebooks utilize <code>tf.keras.layers.Layer</code> to perform all operations on the <strong>GPU/TPU</strong> which can significantly speed up the training process (up to 4x in P100 GPU).</p>\n<p>To reduce the size of the notebook and make it more manageable, the <a href=\"https://github.com/awsaf49/tensorflow_extra\" target=\"_blank\"><strong>tensorflow_extra</strong></a> library is created. This library contains most of the augmentation and transformation code used in the notebooks. You are welcome to contribute. Here is a demo code that runs ops on GPU/TPU,</p>\n<pre><code> tensorflow_extra  tfe\n\nmelspec_layer = tfe.layers.MelSpectrogram()\nspecs = melspec_layer(audios)\n\ntfm_layer = tfe.layers.TimeFreqMask()\nspecs2 = tfm_layer(specs, training=)\n\nnorm_layer = tfe.layers.ZScoreMinMax()\nspecs3 = norm_layer(specs2)\n</code></pre>\n<h2>Notebooks</h2>\n<ul>\n<li>Pretraining is All you Need<ul>\n<li>Train: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-train/\" target=\"_blank\">BirdCLEF23: Pretraining is All you Need [Train]</a></li>\n<li>Infer: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-infer/\" target=\"_blank\">BirdCLEF23: Pretraining is All you Need [Infer]</a></li></ul></li>\n</ul>\n<h2>Previous Notebooks</h2>\n<ul>\n<li>EffNet + FSR + CutMixUp<ul>\n<li>Train: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-train/\" target=\"_blank\">BirdCLEF23: EffNet + FSR + CutMixUp [Train]</a></li>\n<li>Infer: <a href=\"https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-infer/\" target=\"_blank\">BirdCLEF23: EffNet + FSR + CutMixUp [Infer]</a></li></ul></li>\n</ul>\n<h2>Training Configs</h2>\n<ul>\n<li><strong>Pretraining</strong>: ~71k samples from BirdCLEF 2021 &amp; 2022 competition. Filenames with an exact match with primary_label &amp; author are removed to avoid leak.</li>\n<li><strong>Training</strong>: ~16k samples from BirdCLEF 2023 competition.</li>\n<li><strong>Model</strong>: EfficientNetB1 + FSR</li>\n<li><strong>Training Time Duration</strong>: 10 sec</li>\n<li><strong>Image Size</strong>: 128 x 384</li>\n<li><strong>Loss</strong>: CCE</li>\n<li><strong>Sampling</strong>: Downsample data to have max 500 samples in each class &amp; Upsample data to have min 50 samples in each class.</li>\n</ul>\n<h2>Result</h2>\n<ul>\n<li><strong>Validation</strong>: 0.86 (AUC) | 0.911 (cmAP)</li>\n<li><strong>LB</strong>: 0.80</li>\n<li><strong>Improvement</strong>: Local validation -&gt; 4% | LB -&gt; 2%</li>\n</ul>\n<blockquote>\n  <p><strong>Note</strong>: CV - LB gap can be reduced by reducing time duration in training but affects the training score. As there is a small overlap of labels between previous competition and current competition data, we can use them as external data as well.</p>\n</blockquote>\n<h2>Track Experiments with WandB</h2>\n<p>You can track all the experiments on Weights &amp; Biases (WandB) <a href=\"https://wandb.ai/awsaf49/birdclef-2023-public\" target=\"_blank\">here</a><br>\n<img src=\"https://i.postimg.cc/8CsMw43R/wandb-v2.png\"></p>",
      "rawMarkdown": "## Motivation\nIn the current competition, improving the performance of small models has proven to be a challenging task due to only **CPU** inference, along with a time constraint of **2 hours**. Additionally, the majority of models used in this competition are pre-trained on `ImageNet` and are not familiar with **audio data**, making it difficult for them to achieve their true potential even with transfer learning. To overcome this challenge, there are two potential solutions: 1) pre-training on an audio dataset, and 2) training with a large-scale dataset. However, the competition dataset contains only `16k` samples, and the labels from the previous competition dataset do not match with this one, making training with a large-scale dataset challenging. As a solution, utilizing the previous competition datasets for pre-training can improve the models' feature extraction capabilities.\n\n## Processing in GPU/TPU\nAs previously mentioned [here](https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/comments#2179795), **model/TPU/GPU** experiences a bottleneck when waiting for the data which are processed on **CPU**. This is because the Audio -> Spectrogram transformation, Augmentation, and Normalization are performed within `tf.data.Dataset` on the CPU, creating a bottleneck. To alleviate this bottleneck, the following notebooks utilize `tf.keras.layers.Layer` to perform all operations on the **GPU/TPU** which can significantly speed up the training process (up to 4x in P100 GPU).\n\nTo reduce the size of the notebook and make it more manageable, the [**tensorflow_extra**](https://github.com/awsaf49/tensorflow_extra) library is created. This library contains most of the augmentation and transformation code used in the notebooks. You are welcome to contribute. Here is a demo code that runs ops on GPU/TPU,\n```python\nimport tensorflow_extra as tfe\n\nmelspec_layer = tfe.layers.MelSpectrogram()\nspecs = melspec_layer(audios)\n\ntfm_layer = tfe.layers.TimeFreqMask()\nspecs2 = tfm_layer(specs, training=True)\n\nnorm_layer = tfe.layers.ZScoreMinMax()\nspecs3 = norm_layer(specs2)\n```\n\n\n## Notebooks\n* Pretraining is All you Need\n    * Train: [BirdCLEF23: Pretraining is All you Need [Train]](https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-train/)\n    * Infer: [BirdCLEF23: Pretraining is All you Need [Infer]](https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-infer/)\n    \n## Previous Notebooks\n* EffNet + FSR + CutMixUp\n    * Train: [BirdCLEF23: EffNet + FSR + CutMixUp [Train]](https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-train/)\n    * Infer: [BirdCLEF23: EffNet + FSR + CutMixUp [Infer]](https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-infer/)\n\n\n## Training Configs\n* **Pretraining**: ~71k samples from BirdCLEF 2021 & 2022 competition. Filenames with an exact match with primary_label & author are removed to avoid leak.\n* **Training**: ~16k samples from BirdCLEF 2023 competition.\n* **Model**: EfficientNetB1 + FSR\n* **Training Time Duration**: 10 sec\n* **Image Size**: 128 x 384\n* **Loss**: CCE\n* **Sampling**: Downsample data to have max 500 samples in each class & Upsample data to have min 50 samples in each class.\n\n## Result\n* **Validation**: 0.86 (AUC) | 0.911 (cmAP)\n* **LB**: 0.80\n* **Improvement**: Local validation -> 4% | LB -> 2%\n\n> **Note**: CV - LB gap can be reduced by reducing time duration in training but affects the training score. As there is a small overlap of labels between previous competition and current competition data, we can use them as external data as well.\n\n## Track Experiments with WandB\nYou can track all the experiments on Weights & Biases (WandB) [here](https://wandb.ai/awsaf49/birdclef-2023-public)\n<img src=\"https://i.postimg.cc/8CsMw43R/wandb-v2.png\">",
      "votes": null
    },
    {
      "id": "2187842",
      "postDate": "03/19/2023 05:02:59",
      "content": "<p>Hi there! It's great to see that you are utilizing pre-training to improve the performance of your models in the BirdCLEF23 competition. It's definitely a smart move to use the previous competition datasets for pre-training since it can improve the models' feature extraction capabilities.</p>\n<p>I also like the use of tf.keras.layers.Layer to perform all operations on the GPU/TPU to alleviate the bottleneck caused by processing on the CPU. The tensorflow_extra library seems like a helpful resource as well.</p>\n<p>Your validation results of 0.86 (AUC) and 0.911 (cmAP) are impressive, and even though the LB score is currently at 0.80, it's still a significant improvement. It's interesting to note that reducing the training time duration can reduce the CV - LB gap, but it's good to balance that with the training score.</p>\n<p>It seems like you have a solid approach to tackling the challenges presented in the competition. Best of luck!</p>",
      "rawMarkdown": "Hi there! It's great to see that you are utilizing pre-training to improve the performance of your models in the BirdCLEF23 competition. It's definitely a smart move to use the previous competition datasets for pre-training since it can improve the models' feature extraction capabilities.\n\nI also like the use of tf.keras.layers.Layer to perform all operations on the GPU/TPU to alleviate the bottleneck caused by processing on the CPU. The tensorflow_extra library seems like a helpful resource as well.\n\nYour validation results of 0.86 (AUC) and 0.911 (cmAP) are impressive, and even though the LB score is currently at 0.80, it's still a significant improvement. It's interesting to note that reducing the training time duration can reduce the CV - LB gap, but it's good to balance that with the training score.\n\nIt seems like you have a solid approach to tackling the challenges presented in the competition. Best of luck!",
      "votes": null
    },
    {
      "id": "2187875",
      "postDate": "03/19/2023 05:22:37",
      "content": "<p>Thanks for your feedback :)</p>",
      "rawMarkdown": "Thanks for your feedback :)",
      "votes": null
    },
    {
      "id": "2188130",
      "postDate": "03/19/2023 10:40:07",
      "content": "<p>Great work! I agree, pre-training might be key in this year's competition. Especially when thinking about improving performance for species with only a few samples.</p>\n<p>I case you think it helps: As mentioned in <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358\" target=\"_blank\">this thread</a>, we also published <strong>test data</strong> from previous years, which might come in handy when trying to close the acoustic gap between training and test data. Species overlap is minimal, but might be a good thing to include for local testing.</p>",
      "rawMarkdown": "Great work! I agree, pre-training might be key in this year's competition. Especially when thinking about improving performance for species with only a few samples.\n\nI case you think it helps: As mentioned in [this thread](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358), we also published **test data** from previous years, which might come in handy when trying to close the acoustic gap between training and test data. Species overlap is minimal, but might be a good thing to include for local testing.",
      "votes": null
    },
    {
      "id": "2188133",
      "postDate": "03/19/2023 10:44:10",
      "content": "<p>Thanks for info. </p>",
      "rawMarkdown": "Thanks for info.",
      "votes": null
    },
    {
      "id": "2188147",
      "postDate": "03/19/2023 10:52:00",
      "content": "<p>Many thanks for sharing all of this, this is my first audio competition, and I'm learning a lot in part thanks to your informative topics and notebooks. </p>",
      "rawMarkdown": "Many thanks for sharing all of this, this is my first audio competition, and I'm learning a lot in part thanks to your informative topics and notebooks.",
      "votes": null
    },
    {
      "id": "2188169",
      "postDate": "03/19/2023 11:06:20",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> I noticed there is overlap of <code>filename</code> within same <code>author</code>, <code>primary_label</code>. Are they referring to the same file in other words is there any overlap of audio files between competitions?</p>",
      "rawMarkdown": "stefankahl I noticed there is overlap of `filename` within same `author`, `primary_label`. Are they referring to the same file in other words is there any overlap of audio files between competitions?",
      "votes": null
    },
    {
      "id": "2188271",
      "postDate": "03/19/2023 12:55:46",
      "content": "<p>Yes, there might be a few species with overlap. Usually, the filename should indicate if it's a duplicate, because it refers to the Xeno-canto ID.</p>",
      "rawMarkdown": "Yes, there might be a few species with overlap. Usually, the filename should indicate if it's a duplicate, because it refers to the Xeno-canto ID.",
      "votes": null
    },
    {
      "id": "2188278",
      "postDate": "03/19/2023 13:00:18",
      "content": "<p>between 2023 &amp; 2021-2022 I found nearly 880 duplicate files. Initially, I was getting a pretty high score due to this overlap.</p>",
      "rawMarkdown": "between 2023 & 2021-2022 I found nearly 880 duplicate files. Initially, I was getting a pretty high score due to this overlap.",
      "votes": null
    },
    {
      "id": "2188443",
      "postDate": "03/19/2023 16:16:31",
      "content": "<h2>Update - 19 March 2023</h2>\n<ul>\n<li>BirdCLEF 2020 dataset and Xeno-Canto Extended dataset (by <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>) are added for pretraining. Surprisingly there are overlaps of <code>xc_id</code> (reco. by <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>) between these datasets, thus after adding these two total pre-training data is nearly <code>81k</code>.</li>\n</ul>",
      "rawMarkdown": "## Update - 19 March 2023\n* BirdCLEF 2020 dataset and Xeno-Canto Extended dataset (by @rohanrao) are added for pretraining. Surprisingly there are overlaps of `xc_id` (reco. by @stefankahl) between these datasets, thus after adding these two total pre-training data is nearly `81k`.",
      "votes": null
    },
    {
      "id": "2188556",
      "postDate": "03/19/2023 18:40:48",
      "content": "<p>Thanks for sharing! it is certainly a great pair of notebooks. a lot to learn there. especially doing pretraining properly and audio augmentation. <br>\none idea I played around with regarding augmentation is background removal. it works well and showed some promise in local cv (haven't submitted yet). it is however quite slow as it has to be done on cpu. wondering if you know of some gpu based solution.</p>",
      "rawMarkdown": "Thanks for sharing! it is certainly a great pair of notebooks. a lot to learn there. especially doing pretraining properly and audio augmentation. \none idea I played around with regarding augmentation is background removal. it works well and showed some promise in local cv (haven't submitted yet). it is however quite slow as it has to be done on cpu. wondering if you know of some gpu based solution.",
      "votes": null
    },
    {
      "id": "2190145",
      "postDate": "03/21/2023 04:27:44",
      "content": "<p>Thanks for sharing a valuable technique and inspirational insights.👍👍</p>",
      "rawMarkdown": "Thanks for sharing a valuable technique and inspirational insights.👍👍",
      "votes": null
    },
    {
      "id": "2191397",
      "postDate": "03/21/2023 23:45:33",
      "content": "<p>Wow, Thank for sharing</p>",
      "rawMarkdown": "Wow, Thank for sharing",
      "votes": null
    },
    {
      "id": "2192555",
      "postDate": "03/22/2023 18:09:14",
      "content": "<p>Yes, I tried denoising in birdclef-2021. Due to slow inference speed I dropped it</p>",
      "rawMarkdown": "Yes, I tried denoising in birdclef-2021. Due to slow inference speed I dropped it",
      "votes": null
    },
    {
      "id": "2200618",
      "postDate": "03/28/2023 16:52:04",
      "content": "<p>i think you can use Demucs library to do this on gpu</p>",
      "rawMarkdown": "i think you can use Demucs library to do this on gpu",
      "votes": null
    },
    {
      "id": "2200830",
      "postDate": "03/28/2023 20:20:36",
      "content": "<p>But for inference you are bound to use CPU…</p>",
      "rawMarkdown": "But for inference you are bound to use CPU...",
      "votes": null
    },
    {
      "id": "2200852",
      "postDate": "03/28/2023 20:49:27",
      "content": "<p>thanks. you are a legend!</p>",
      "rawMarkdown": "thanks. you are a legend!",
      "votes": null
    },
    {
      "id": "2232452",
      "postDate": "04/24/2023 11:02:04",
      "content": "<p>Great Work !!! Thanks </p>",
      "rawMarkdown": "Great Work !!! Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2187842,
      "author_name": "siddharthkumarsah",
      "author_url": "",
      "post_date": "03/19/2023 05:02:59",
      "content": "<p>Hi there! It's great to see that you are utilizing pre-training to improve the performance of your models in the BirdCLEF23 competition. It's definitely a smart move to use the previous competition datasets for pre-training since it can improve the models' feature extraction capabilities.</p>\n<p>I also like the use of tf.keras.layers.Layer to perform all operations on the GPU/TPU to alleviate the bottleneck caused by processing on the CPU. The tensorflow_extra library seems like a helpful resource as well.</p>\n<p>Your validation results of 0.86 (AUC) and 0.911 (cmAP) are impressive, and even though the LB score is currently at 0.80, it's still a significant improvement. It's interesting to note that reducing the training time duration can reduce the CV - LB gap, but it's good to balance that with the training score.</p>\n<p>It seems like you have a solid approach to tackling the challenges presented in the competition. Best of luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2187875,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "03/19/2023 05:22:37",
          "content": "<p>Thanks for your feedback :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2188130,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "03/19/2023 10:40:07",
      "content": "<p>Great work! I agree, pre-training might be key in this year's competition. Especially when thinking about improving performance for species with only a few samples.</p>\n<p>I case you think it helps: As mentioned in <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358\" target=\"_blank\">this thread</a>, we also published <strong>test data</strong> from previous years, which might come in handy when trying to close the acoustic gap between training and test data. Species overlap is minimal, but might be a good thing to include for local testing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2188133,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "03/19/2023 10:44:10",
          "content": "<p>Thanks for info. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2188169,
              "author_name": "awsaf49",
              "author_url": "",
              "post_date": "03/19/2023 11:06:20",
              "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> I noticed there is overlap of <code>filename</code> within same <code>author</code>, <code>primary_label</code>. Are they referring to the same file in other words is there any overlap of audio files between competitions?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2188271,
                  "author_name": "stefankahl",
                  "author_url": "",
                  "post_date": "03/19/2023 12:55:46",
                  "content": "<p>Yes, there might be a few species with overlap. Usually, the filename should indicate if it's a duplicate, because it refers to the Xeno-canto ID.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2188278,
                      "author_name": "awsaf49",
                      "author_url": "",
                      "post_date": "03/19/2023 13:00:18",
                      "content": "<p>between 2023 &amp; 2021-2022 I found nearly 880 duplicate files. Initially, I was getting a pretty high score due to this overlap.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2188147,
      "author_name": "maxdiazbattan",
      "author_url": "",
      "post_date": "03/19/2023 10:52:00",
      "content": "<p>Many thanks for sharing all of this, this is my first audio competition, and I'm learning a lot in part thanks to your informative topics and notebooks. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2188443,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "03/19/2023 16:16:31",
      "content": "<h2>Update - 19 March 2023</h2>\n<ul>\n<li>BirdCLEF 2020 dataset and Xeno-Canto Extended dataset (by <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>) are added for pretraining. Surprisingly there are overlaps of <code>xc_id</code> (reco. by <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>) between these datasets, thus after adding these two total pre-training data is nearly <code>81k</code>.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2200852,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "03/28/2023 20:49:27",
          "content": "<p>thanks. you are a legend!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2188556,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "03/19/2023 18:40:48",
      "content": "<p>Thanks for sharing! it is certainly a great pair of notebooks. a lot to learn there. especially doing pretraining properly and audio augmentation. <br>\none idea I played around with regarding augmentation is background removal. it works well and showed some promise in local cv (haven't submitted yet). it is however quite slow as it has to be done on cpu. wondering if you know of some gpu based solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2192555,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "03/22/2023 18:09:14",
          "content": "<p>Yes, I tried denoising in birdclef-2021. Due to slow inference speed I dropped it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2200618,
          "author_name": "riadalmadani",
          "author_url": "",
          "post_date": "03/28/2023 16:52:04",
          "content": "<p>i think you can use Demucs library to do this on gpu</p>",
          "votes": null,
          "replies": [
            {
              "id": 2200830,
              "author_name": "awsaf49",
              "author_url": "",
              "post_date": "03/28/2023 20:20:36",
              "content": "<p>But for inference you are bound to use CPU…</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2190145,
      "author_name": "tariqbashir",
      "author_url": "",
      "post_date": "03/21/2023 04:27:44",
      "content": "<p>Thanks for sharing a valuable technique and inspirational insights.👍👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2191397,
      "author_name": "thitiwat",
      "author_url": "",
      "post_date": "03/21/2023 23:45:33",
      "content": "<p>Wow, Thank for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232452,
      "author_name": "adnanzaidi",
      "author_url": "",
      "post_date": "04/24/2023 11:02:04",
      "content": "<p>Great Work !!! Thanks </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2187823": "## Motivation\nIn the current competition, improving the performance of small models has proven to be a challenging task due to only **CPU** inference, along with a time constraint of **2 hours**. Additionally, the majority of models used in this competition are pre-trained on `ImageNet` and are not familiar with **audio data**, making it difficult for them to achieve their true potential even with transfer learning. To overcome this challenge, there are two potential solutions: 1) pre-training on an audio dataset, and 2) training with a large-scale dataset. However, the competition dataset contains only `16k` samples, and the labels from the previous competition dataset do not match with this one, making training with a large-scale dataset challenging. As a solution, utilizing the previous competition datasets for pre-training can improve the models' feature extraction capabilities.\n\n## Processing in GPU/TPU\nAs previously mentioned [here](https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-train/comments#2179795), **model/TPU/GPU** experiences a bottleneck when waiting for the data which are processed on **CPU**. This is because the Audio -> Spectrogram transformation, Augmentation, and Normalization are performed within `tf.data.Dataset` on the CPU, creating a bottleneck. To alleviate this bottleneck, the following notebooks utilize `tf.keras.layers.Layer` to perform all operations on the **GPU/TPU** which can significantly speed up the training process (up to 4x in P100 GPU).\n\nTo reduce the size of the notebook and make it more manageable, the [**tensorflow_extra**](https://github.com/awsaf49/tensorflow_extra) library is created. This library contains most of the augmentation and transformation code used in the notebooks. You are welcome to contribute. Here is a demo code that runs ops on GPU/TPU,\n```python\nimport tensorflow_extra as tfe\n\nmelspec_layer = tfe.layers.MelSpectrogram()\nspecs = melspec_layer(audios)\n\ntfm_layer = tfe.layers.TimeFreqMask()\nspecs2 = tfm_layer(specs, training=True)\n\nnorm_layer = tfe.layers.ZScoreMinMax()\nspecs3 = norm_layer(specs2)\n```\n\n\n## Notebooks\n* Pretraining is All you Need\n    * Train: [BirdCLEF23: Pretraining is All you Need [Train]](https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-train/)\n    * Infer: [BirdCLEF23: Pretraining is All you Need [Infer]](https://www.kaggle.com/awsaf49/birdclef23-pretraining-is-all-you-need-infer/)\n    \n## Previous Notebooks\n* EffNet + FSR + CutMixUp\n    * Train: [BirdCLEF23: EffNet + FSR + CutMixUp [Train]](https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-train/)\n    * Infer: [BirdCLEF23: EffNet + FSR + CutMixUp [Infer]](https://www.kaggle.com/awsaf49/birdclef23-effnet-fsr-cutmixup-infer/)\n\n\n## Training Configs\n* **Pretraining**: ~71k samples from BirdCLEF 2021 & 2022 competition. Filenames with an exact match with primary_label & author are removed to avoid leak.\n* **Training**: ~16k samples from BirdCLEF 2023 competition.\n* **Model**: EfficientNetB1 + FSR\n* **Training Time Duration**: 10 sec\n* **Image Size**: 128 x 384\n* **Loss**: CCE\n* **Sampling**: Downsample data to have max 500 samples in each class & Upsample data to have min 50 samples in each class.\n\n## Result\n* **Validation**: 0.86 (AUC) | 0.911 (cmAP)\n* **LB**: 0.80\n* **Improvement**: Local validation -> 4% | LB -> 2%\n\n> **Note**: CV - LB gap can be reduced by reducing time duration in training but affects the training score. As there is a small overlap of labels between previous competition and current competition data, we can use them as external data as well.\n\n## Track Experiments with WandB\nYou can track all the experiments on Weights & Biases (WandB) [here](https://wandb.ai/awsaf49/birdclef-2023-public)\n<img src=\"https://i.postimg.cc/8CsMw43R/wandb-v2.png\">",
    "2187842": "Hi there! It's great to see that you are utilizing pre-training to improve the performance of your models in the BirdCLEF23 competition. It's definitely a smart move to use the previous competition datasets for pre-training since it can improve the models' feature extraction capabilities.\n\nI also like the use of tf.keras.layers.Layer to perform all operations on the GPU/TPU to alleviate the bottleneck caused by processing on the CPU. The tensorflow_extra library seems like a helpful resource as well.\n\nYour validation results of 0.86 (AUC) and 0.911 (cmAP) are impressive, and even though the LB score is currently at 0.80, it's still a significant improvement. It's interesting to note that reducing the training time duration can reduce the CV - LB gap, but it's good to balance that with the training score.\n\nIt seems like you have a solid approach to tackling the challenges presented in the competition. Best of luck!",
    "2187875": "Thanks for your feedback :)",
    "2188130": "Great work! I agree, pre-training might be key in this year's competition. Especially when thinking about improving performance for species with only a few samples.\n\nI case you think it helps: As mentioned in [this thread](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358), we also published **test data** from previous years, which might come in handy when trying to close the acoustic gap between training and test data. Species overlap is minimal, but might be a good thing to include for local testing.",
    "2188133": "Thanks for info.",
    "2188147": "Many thanks for sharing all of this, this is my first audio competition, and I'm learning a lot in part thanks to your informative topics and notebooks.",
    "2188169": "stefankahl I noticed there is overlap of `filename` within same `author`, `primary_label`. Are they referring to the same file in other words is there any overlap of audio files between competitions?",
    "2188271": "Yes, there might be a few species with overlap. Usually, the filename should indicate if it's a duplicate, because it refers to the Xeno-canto ID.",
    "2188278": "between 2023 & 2021-2022 I found nearly 880 duplicate files. Initially, I was getting a pretty high score due to this overlap.",
    "2188443": "## Update - 19 March 2023\n* BirdCLEF 2020 dataset and Xeno-Canto Extended dataset (by @rohanrao) are added for pretraining. Surprisingly there are overlaps of `xc_id` (reco. by @stefankahl) between these datasets, thus after adding these two total pre-training data is nearly `81k`.",
    "2188556": "Thanks for sharing! it is certainly a great pair of notebooks. a lot to learn there. especially doing pretraining properly and audio augmentation. \none idea I played around with regarding augmentation is background removal. it works well and showed some promise in local cv (haven't submitted yet). it is however quite slow as it has to be done on cpu. wondering if you know of some gpu based solution.",
    "2190145": "Thanks for sharing a valuable technique and inspirational insights.👍👍",
    "2191397": "Wow, Thank for sharing",
    "2192555": "Yes, I tried denoising in birdclef-2021. Due to slow inference speed I dropped it",
    "2200618": "i think you can use Demucs library to do this on gpu",
    "2200830": "But for inference you are bound to use CPU...",
    "2200852": "thanks. you are a legend!",
    "2232452": "Great Work !!! Thanks"
  },
  "source": "meta"
}