{
  "id": 243583,
  "title": "22nd Place Solution ",
  "url": "/competitions/birdclef-2021/writeups/in-cv-we-trust-22nd-place-solution",
  "author_name": "",
  "post_date": "2022-02-15T21:30:36.663Z",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Congratulations on all teams that have managed to secure top places and develop well performed solutions. It was a tough problem with main challenges the shift domain &amp; noisy/weak labelling as described by many others. What follows is not rocket science, mostly is based on ideas from previous audio competitions, however I thought that It might be useful to describe it. </p>\n<hr>\n<p><strong>TLDR</strong></p>\n<p><strong>Our team's submission is a weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)</strong> </p>\n<ul>\n<li><p>Read about our team's ensemble approach by <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> </p></li>\n<li><p>Read about our <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243349\" target=\"_blank\">9th Public solution outline &amp; PP</a> by <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a></p></li>\n</ul>\n<p><img src=\"https://i.ibb.co/kGsm45V/birdclef-inference-07818.png\" alt=\"\"></p>\n<hr>\n<p><strong>Modelling approach (my journey)</strong></p>\n<p>Before I start working on this problem I gave some time to read past year's top-10 solutions and view/analysed many many spectrograms. In overall, I performed more than 60 training experiments (single fold most of them) with different pipelines. Next, I describe briefly the main approaches</p>\n<p>1) CNNs trained with precomputed mel-specs (5, 7, 10 sec)</p>\n<p>2) CNNs with different pooling/heads a) attention; b) plain dense head; that takes the raw audio as input (10, 15 sec) and compute mel-spec transformations on the fly. This increased training time but gave more flexibility trying different params as the penultimate goal was to add some diversity to the ensemble.</p>\n<p>3) SED models (few experiments at start of comp. however  none of them used in team's ensemble, <a href=\"https://www.kaggle.com/rsinda\" target=\"_blank\">@rsinda</a> had better results with sed models so I leave it for him to describe this part)</p>\n<p>For all three cases above numerous backbones tested but the following worked better for me:</p>\n<ul>\n<li>backbones: effnet-B0, regnetx, densenet121</li>\n</ul>\n<p>The basis for all training pipelines was the 2nd solution code from last year's comp shared generously by <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> with many modifications along the road - I guess the most important one was to use the audio ratings as loss weights (sample wise) to account for the \"noisy\" low rating recordings.</p>\n<p>Mel-spectrograms: <code>n_fft=2048</code>, <code>n_mels=128</code>, with diff <code>hop_lengths = [375, 512, 878]</code>   <br>\n5, 7, 10 sec audio crops, started with 3 channels but towards the end I used single channel mels </p>\n<p>Optimizer/Loss: Adam / BCELogits &amp; FocallLoss (tried many others but mainly these two)</p>\n<p>Metrics to monitor: f1 at 03,05,07, max_f1, precision, recall (during training and on SC)</p>\n<p><strong>Training - other tips</strong></p>\n<ul>\n<li>use of secondary labels</li>\n<li>pink + white noise</li>\n<li>use Time-Freq masking</li>\n<li>Mixup / Mixor (for some of them)</li>\n<li>use Multi-sample dropout a.k.a. big dropout <a href=\"https://arxiv.org/abs/1905.09788\" target=\"_blank\">(see paper)</a> for better generalisation (tried on last days - haven't tested on LB alone but didn't see improvement using it in team's ensemble)</li>\n<li>used wandb for logging experiments &amp; excel sheet with more details</li>\n<li>check heatmaps with model predictions for each Train SC (eg see figure below)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/WzSVZjf/fig2.png\" alt=\"img\"></p>\n<p><strong>Things that didn't work for me / didn't show improvement</strong> </p>\n<ul>\n<li>training a second stage model on hot-slices (selected by the best at the time 1st stage model)</li>\n<li>training with reduced bird classes ~300-330</li>\n<li>using deltas for 3-channel</li>\n<li>using diff random_power() - ie, change contrast</li>\n<li>finetuning checkpoints from previous year comp (trained on 264 classes) from <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> 3rd place solution</li>\n<li>tweaking <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> type model from last year comp <a href=\"https://www.kaggle.com/iafoss/cornell-birdcall\" target=\"_blank\">described here</a></li>\n</ul>\n<p><strong>Acknowledgments</strong></p>\n<p>Thanks to hosts and kaggle for organizing this competition &amp; every single kaggler that shared his notebooks/tips/concerns in public. </p>\n<p>Thanks to my teammates <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a>, <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, <a href=\"https://www.kaggle.com/rsinda\" target=\"_blank\">@rsinda</a>, <a href=\"https://www.kaggle.com/tomohiroh\" target=\"_blank\">@tomohiroh</a> and especially our team leader <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> for his efforts to integrate/test all models to our final ensemble.</p>",
  "messages": [
    {
      "id": "1334056",
      "postDate": "06/03/2021 08:50:27",
      "content": "<p>Congratulations on all teams that have managed to secure top places and develop well performed solutions. It was a tough problem with main challenges the shift domain &amp; noisy/weak labelling as described by many others. What follows is not rocket science, mostly is based on ideas from previous audio competitions, however I thought that It might be useful to describe it. </p>\n<hr>\n<p><strong>TLDR</strong></p>\n<p><strong>Our team's submission is a weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)</strong> </p>\n<ul>\n<li><p>Read about our team's ensemble approach by <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> </p></li>\n<li><p>Read about our <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243349\" target=\"_blank\">9th Public solution outline &amp; PP</a> by <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a></p></li>\n</ul>\n<p><img src=\"https://i.ibb.co/kGsm45V/birdclef-inference-07818.png\" alt=\"\"></p>\n<hr>\n<p><strong>Modelling approach (my journey)</strong></p>\n<p>Before I start working on this problem I gave some time to read past year's top-10 solutions and view/analysed many many spectrograms. In overall, I performed more than 60 training experiments (single fold most of them) with different pipelines. Next, I describe briefly the main approaches</p>\n<p>1) CNNs trained with precomputed mel-specs (5, 7, 10 sec)</p>\n<p>2) CNNs with different pooling/heads a) attention; b) plain dense head; that takes the raw audio as input (10, 15 sec) and compute mel-spec transformations on the fly. This increased training time but gave more flexibility trying different params as the penultimate goal was to add some diversity to the ensemble.</p>\n<p>3) SED models (few experiments at start of comp. however  none of them used in team's ensemble, <a href=\"https://www.kaggle.com/rsinda\" target=\"_blank\">@rsinda</a> had better results with sed models so I leave it for him to describe this part)</p>\n<p>For all three cases above numerous backbones tested but the following worked better for me:</p>\n<ul>\n<li>backbones: effnet-B0, regnetx, densenet121</li>\n</ul>\n<p>The basis for all training pipelines was the 2nd solution code from last year's comp shared generously by <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> with many modifications along the road - I guess the most important one was to use the audio ratings as loss weights (sample wise) to account for the \"noisy\" low rating recordings.</p>\n<p>Mel-spectrograms: <code>n_fft=2048</code>, <code>n_mels=128</code>, with diff <code>hop_lengths = [375, 512, 878]</code>   <br>\n5, 7, 10 sec audio crops, started with 3 channels but towards the end I used single channel mels </p>\n<p>Optimizer/Loss: Adam / BCELogits &amp; FocallLoss (tried many others but mainly these two)</p>\n<p>Metrics to monitor: f1 at 03,05,07, max_f1, precision, recall (during training and on SC)</p>\n<p><strong>Training - other tips</strong></p>\n<ul>\n<li>use of secondary labels</li>\n<li>pink + white noise</li>\n<li>use Time-Freq masking</li>\n<li>Mixup / Mixor (for some of them)</li>\n<li>use Multi-sample dropout a.k.a. big dropout <a href=\"https://arxiv.org/abs/1905.09788\" target=\"_blank\">(see paper)</a> for better generalisation (tried on last days - haven't tested on LB alone but didn't see improvement using it in team's ensemble)</li>\n<li>used wandb for logging experiments &amp; excel sheet with more details</li>\n<li>check heatmaps with model predictions for each Train SC (eg see figure below)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/WzSVZjf/fig2.png\" alt=\"img\"></p>\n<p><strong>Things that didn't work for me / didn't show improvement</strong> </p>\n<ul>\n<li>training a second stage model on hot-slices (selected by the best at the time 1st stage model)</li>\n<li>training with reduced bird classes ~300-330</li>\n<li>using deltas for 3-channel</li>\n<li>using diff random_power() - ie, change contrast</li>\n<li>finetuning checkpoints from previous year comp (trained on 264 classes) from <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> 3rd place solution</li>\n<li>tweaking <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> type model from last year comp <a href=\"https://www.kaggle.com/iafoss/cornell-birdcall\" target=\"_blank\">described here</a></li>\n</ul>\n<p><strong>Acknowledgments</strong></p>\n<p>Thanks to hosts and kaggle for organizing this competition &amp; every single kaggler that shared his notebooks/tips/concerns in public. </p>\n<p>Thanks to my teammates <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a>, <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>, <a href=\"https://www.kaggle.com/rsinda\" target=\"_blank\">@rsinda</a>, <a href=\"https://www.kaggle.com/tomohiroh\" target=\"_blank\">@tomohiroh</a> and especially our team leader <a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">@rohitsingh9990</a> for his efforts to integrate/test all models to our final ensemble.</p>",
      "rawMarkdown": "Congratulations on all teams that have managed to secure top places and develop well performed solutions. It was a tough problem with main challenges the shift domain & noisy/weak labelling as described by many others. What follows is not rocket science, mostly is based on ideas from previous audio competitions, however I thought that It might be useful to describe it. \n\n-----------------------\n\n**TLDR**\n\n**Our team's submission is a weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)** \n\n- Read about our team's ensemble approach by [@rohitsingh9990](https://www.kaggle.com/rohitsingh9990) ~~[coming soon]~~\n\n- Read about our [9th Public solution outline & PP](https://www.kaggle.com/c/birdclef-2021/discussion/243349) by [@jaideepvalani](https://www.kaggle.com/jaideepvalani)\n\n![](https://i.ibb.co/kGsm45V/birdclef-inference-07818.png)\n\n-----------------------\n\n\n**Modelling approach (my journey)**\n\n\nBefore I start working on this problem I gave some time to read past year's top-10 solutions and view/analysed many many spectrograms. In overall, I performed more than 60 training experiments (single fold most of them) with different pipelines. Next, I describe briefly the main approaches\n\n1) CNNs trained with precomputed mel-specs (5, 7, 10 sec)\n\n2) CNNs with different pooling/heads a) attention; b) plain dense head; that takes the raw audio as input (10, 15 sec) and compute mel-spec transformations on the fly. This increased training time but gave more flexibility trying different params as the penultimate goal was to add some diversity to the ensemble.\n\n3) SED models (few experiments at start of comp. however  none of them used in team's ensemble, [@rsinda](https://www.kaggle.com/rsinda) had better results with sed models so I leave it for him to describe this part)\n\n\nFor all three cases above numerous backbones tested but the following worked better for me:\n- backbones: effnet-B0, regnetx, densenet121\n\n\nThe basis for all training pipelines was the 2nd solution code from last year's comp shared generously by [@vlomme](https://www.kaggle.com/vlomme) with many modifications along the road - I guess the most important one was to use the audio ratings as loss weights (sample wise) to account for the \"noisy\" low rating recordings.\n\n\n\nMel-spectrograms: `n_fft=2048`, `n_mels=128`, with diff `hop_lengths = [375, 512, 878]`   \n5, 7, 10 sec audio crops, started with 3 channels but towards the end I used single channel mels \n\n\nOptimizer/Loss: Adam / BCELogits & FocallLoss (tried many others but mainly these two)\n\n\nMetrics to monitor: f1 at 03,05,07, max_f1, precision, recall (during training and on SC)\n\n\n**Training - other tips**\n\n- use of secondary labels\n- pink + white noise\n- use Time-Freq masking\n- Mixup / Mixor (for some of them)\n- use Multi-sample dropout a.k.a. big dropout [(see paper)](https://arxiv.org/abs/1905.09788) for better generalisation (tried on last days - haven't tested on LB alone but didn't see improvement using it in team's ensemble)\n- used wandb for logging experiments & excel sheet with more details\n- check heatmaps with model predictions for each Train SC (eg see figure below)\n\n\n\n![img](https://i.ibb.co/WzSVZjf/fig2.png)\n\n\n\n**Things that didn't work for me / didn't show improvement** \n\n\n- training a second stage model on hot-slices (selected by the best at the time 1st stage model)\n- training with reduced bird classes ~300-330\n- using deltas for 3-channel\n- using diff random_power() - ie, change contrast\n- finetuning checkpoints from previous year comp (trained on 264 classes) from [@theoviel](https://www.kaggle.com/theoviel) 3rd place solution\n- tweaking [@iafoss](https://www.kaggle.com/iafoss) type model from last year comp [described here](https://www.kaggle.com/iafoss/cornell-birdcall)\n\n\n\n\n\n**Acknowledgments**\n\nThanks to hosts and kaggle for organizing this competition & every single kaggler that shared his notebooks/tips/concerns in public. \n\nThanks to my teammates [@rohitsingh9990](https://www.kaggle.com/rohitsingh9990), [@jaideepvalani](https://www.kaggle.com/jaideepvalani), [@rsinda](https://www.kaggle.com/rsinda), [@tomohiroh](https://www.kaggle.com/tomohiroh) and especially our team leader [@rohitsingh9990](https://www.kaggle.com/rohitsingh9990) for his efforts to integrate/test all models to our final ensemble.",
      "votes": null
    },
    {
      "id": "1335391",
      "postDate": "06/04/2021 07:26:50",
      "content": "<p>Congratulations on your 22. place <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> and thanks for the nice write-up!</p>",
      "rawMarkdown": "Congratulations on your 22. place @imeintanis and thanks for the nice write-up!",
      "votes": null
    },
    {
      "id": "1335461",
      "postDate": "06/04/2021 08:28:51",
      "content": "<p>You're welcome!! :)</p>",
      "rawMarkdown": "You're welcome!! :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335391,
      "author_name": "jonathanbesomi",
      "author_url": "",
      "post_date": "06/04/2021 07:26:50",
      "content": "<p>Congratulations on your 22. place <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> and thanks for the nice write-up!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335461,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "06/04/2021 08:28:51",
          "content": "<p>You're welcome!! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1334056": "Congratulations on all teams that have managed to secure top places and develop well performed solutions. It was a tough problem with main challenges the shift domain & noisy/weak labelling as described by many others. What follows is not rocket science, mostly is based on ideas from previous audio competitions, however I thought that It might be useful to describe it. \n\n-----------------------\n\n**TLDR**\n\n**Our team's submission is a weighted ensemble of 3 groups of models (CNNs + SED) with different backbones, training procedures, durations etc + PP (threshold optimisation)** \n\n- Read about our team's ensemble approach by [@rohitsingh9990](https://www.kaggle.com/rohitsingh9990) ~~[coming soon]~~\n\n- Read about our [9th Public solution outline & PP](https://www.kaggle.com/c/birdclef-2021/discussion/243349) by [@jaideepvalani](https://www.kaggle.com/jaideepvalani)\n\n![](https://i.ibb.co/kGsm45V/birdclef-inference-07818.png)\n\n-----------------------\n\n\n**Modelling approach (my journey)**\n\n\nBefore I start working on this problem I gave some time to read past year's top-10 solutions and view/analysed many many spectrograms. In overall, I performed more than 60 training experiments (single fold most of them) with different pipelines. Next, I describe briefly the main approaches\n\n1) CNNs trained with precomputed mel-specs (5, 7, 10 sec)\n\n2) CNNs with different pooling/heads a) attention; b) plain dense head; that takes the raw audio as input (10, 15 sec) and compute mel-spec transformations on the fly. This increased training time but gave more flexibility trying different params as the penultimate goal was to add some diversity to the ensemble.\n\n3) SED models (few experiments at start of comp. however  none of them used in team's ensemble, [@rsinda](https://www.kaggle.com/rsinda) had better results with sed models so I leave it for him to describe this part)\n\n\nFor all three cases above numerous backbones tested but the following worked better for me:\n- backbones: effnet-B0, regnetx, densenet121\n\n\nThe basis for all training pipelines was the 2nd solution code from last year's comp shared generously by [@vlomme](https://www.kaggle.com/vlomme) with many modifications along the road - I guess the most important one was to use the audio ratings as loss weights (sample wise) to account for the \"noisy\" low rating recordings.\n\n\n\nMel-spectrograms: `n_fft=2048`, `n_mels=128`, with diff `hop_lengths = [375, 512, 878]`   \n5, 7, 10 sec audio crops, started with 3 channels but towards the end I used single channel mels \n\n\nOptimizer/Loss: Adam / BCELogits & FocallLoss (tried many others but mainly these two)\n\n\nMetrics to monitor: f1 at 03,05,07, max_f1, precision, recall (during training and on SC)\n\n\n**Training - other tips**\n\n- use of secondary labels\n- pink + white noise\n- use Time-Freq masking\n- Mixup / Mixor (for some of them)\n- use Multi-sample dropout a.k.a. big dropout [(see paper)](https://arxiv.org/abs/1905.09788) for better generalisation (tried on last days - haven't tested on LB alone but didn't see improvement using it in team's ensemble)\n- used wandb for logging experiments & excel sheet with more details\n- check heatmaps with model predictions for each Train SC (eg see figure below)\n\n\n\n![img](https://i.ibb.co/WzSVZjf/fig2.png)\n\n\n\n**Things that didn't work for me / didn't show improvement** \n\n\n- training a second stage model on hot-slices (selected by the best at the time 1st stage model)\n- training with reduced bird classes ~300-330\n- using deltas for 3-channel\n- using diff random_power() - ie, change contrast\n- finetuning checkpoints from previous year comp (trained on 264 classes) from [@theoviel](https://www.kaggle.com/theoviel) 3rd place solution\n- tweaking [@iafoss](https://www.kaggle.com/iafoss) type model from last year comp [described here](https://www.kaggle.com/iafoss/cornell-birdcall)\n\n\n\n\n\n**Acknowledgments**\n\nThanks to hosts and kaggle for organizing this competition & every single kaggler that shared his notebooks/tips/concerns in public. \n\nThanks to my teammates [@rohitsingh9990](https://www.kaggle.com/rohitsingh9990), [@jaideepvalani](https://www.kaggle.com/jaideepvalani), [@rsinda](https://www.kaggle.com/rsinda), [@tomohiroh](https://www.kaggle.com/tomohiroh) and especially our team leader [@rohitsingh9990](https://www.kaggle.com/rohitsingh9990) for his efforts to integrate/test all models to our final ensemble.",
    "1335391": "Congratulations on your 22. place @imeintanis and thanks for the nice write-up!",
    "1335461": "You're welcome!! :)"
  },
  "source": "meta"
}