{
  "id": 511905,
  "title": "3rd solution",
  "url": "/competitions/birdclef-2024/writeups/nvbird-3rd-solution",
  "author_name": "",
  "post_date": "2024-06-25T08:20:09.103Z",
  "votes": 78,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I am the one posting but this is team work with <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> : team NVBird! </p>\n<p>Thanks everyone for the interesting competition, and congratulations to the winners ! We are very happy with the 3rd place, even though we missed the win by a very tiny margin.</p>\n<h2>Overview</h2>\n<p>Our pipeline is summarized below. A key ingredient of our solution is to use unlabeled soundscapes for pseudo labeling and model distillation. A number of models were trained with the training data, then used to predict labels on 5 second clips from unlabelled soundscapes. These were added to the original training data to train a new set of models used for the final submission. </p>\n<p><img src=\"https://i.ibb.co/QJQrdFf/Bird-CLEF-pipe.png\" alt=\"Bird-CLEF-pipe\"></p>\n<h2>Data</h2>\n<p>Overall, we relied on knowledge acquired during previous competitions, and added some extra samples to fight class imbalance.</p>\n<p>We use this year’s competition data plus the additional data from Xeno Canto shared in the forum, plus records from previous year competitions for the same species as this year.  When a record name appeared in several competitions we picked the most recent one (their content is not always identical form one competition to the next one)<br>\nWe capped the number of records per species to 500, keeping the most recent ones. Indeed, adding all extra data leads to severe class imbalance that was detrimental to model accuracy.<br>\nLow frequency classes are upsampled so that there are at least 10 samples for each class in the training folds. </p>\n<p>To train models, we use the following preprocessing and augmentations:</p>\n<p>We didn’t use all the training data. For each record we use a random crop of 5 seconds clip among the first 6 seconds, or the last 6 seconds of a record. If record length is smaller than 5 seconds, random padding so that the middle of the signal is between 2 and 3 seconds in the 5 sec resulting clip. For most of the models we used time shifting with a one second window as the only augmentation besides mixup. An exception are some models which are inspired by Birdclef23 2nd place SED models (<a href=\"https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution/blob/main/configs/sed_v2s.py)and\" target=\"_blank\">https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution/blob/main/configs/sed_v2s.py)and</a> use the same augmentations as used there<br>\nWe use an additive mixup: primary labels are the max of primary labels of the two audios to be mixed. Secondary labels are the concatenation of secondary labels.<br>\nWe mostly use image models that take log mel spectrograms as input. For these we compute mel spectrograms with parameters chosen to have an image size of 224x224 or 288x288 depending on the image model we use. Input waveforms are normalized to have a std of 1. </p>\n<h2>Models</h2>\n<h3>First Level models</h3>\n<p>The cpu-only requirement was quite constraining for submissions, but this does not apply for pseudo-label generation, so we could use more backbones for first level models. When ensembling several models larger than the ones used at second level we perform what is known as model distillation. This is a rather powerful technique in general.</p>\n<p>Models used include:</p>\n<ul>\n<li>Efficientvit_b0.224.in1k on 224x224 log mel spectrograms</li>\n<li>Efficientvit_b1.r288_in1k on 288x288 log mel spectrograms</li>\n<li>A variety of CNNs (efficientnets, mobilenets, tinynets, mnasnets, mixnets) and Efficientvits (<a href=\"https://arxiv.org/pdf/2205.14756\" target=\"_blank\">b0, b1</a>, <a href=\"https://arxiv.org/pdf/2305.07027\" target=\"_blank\">m3</a> trained on 224x224 log mel spectrograms.</li>\n<li>SED model with tf_efficientnetv2_s_in21k on 128x313  log mel spectrograms?</li>\n<li>We also fine-tuned aves-large and <a href=\"https://github.com/earthspecies/aves?tab=readme-ov-file#birdaves\" target=\"_blank\">aves-base</a>, which is a recent  waveform based model.</li>\n</ul>\n<p>Depending on the pipeline one or more of these models were used to predict pseudo labels on unlabelled soundscapes. We trained most models on full data with 5 different seeds.</p>\n<h3>Second Level models</h3>\n<p>We used efficientvit-b0 primarily and mnasnet-100 on 224x224 log mel spectrograms. Efficientvit-b0 showed great performances while still being very fast to infer. 5 folds take 40 minutes to submit using ONNX. We tried several models with similar throughput to effvit-b0, and decided to also use an mnasnet-100 for diversity. Since computing log mel spectrograms is slow on CPU we decided to use the same mel spectrogram hyperparameters for all models in the inference notebook, to only have to do log mel spectrogram transformation once and use the same input for our ensemble models.</p>\n<p>For training second level models we added the unlabelled soundscapes with the predicted pseudo labels to the training data. This looks simple but it took several attempts to find the correct way to do it. What worked fine was to use rather large batch sizes (128), <br>\nWe used the two following strategies;</p>\n<ul>\n<li>Add extra soundscapes to each batch. Each soundscape is split into 5 second clips, which means we added 48 x 4 = 192 clips with pseudo labels to the 128 samples with actual labels in each batch.</li>\n<li>Add 128 samples with pseudo labels, taken from random soundscapes this time. </li>\n</ul>\n<h3>Final ensemble</h3>\n<p>At the end we had an ensemble of 14 model weights via 3 pipelines, one pipeline per team member:</p>\n<h4>Pipeline 1 - 5 seeds</h4>\n<ul>\n<li>Level 0: efficientvit_b0 </li>\n<li>Level 1: efficientvit_b0</li>\n</ul>\n<h4>Pipeline 2 - 2 seeds + 2</h4>\n<ul>\n<li><p>Level 0: efficientvit_b1, mobilenetv2, efficientnet_b0, efficientnetv2_b0, efficientvit_b0, efficientvit_m3, aves-base, aves-large</p></li>\n<li><p>Level 1: 2x mnasnet-100</p></li>\n<li><p>Level 0: Efficientvit_b0, mixnet_s, mnasnet_100, tinynet_b, efficientvit_b0, efficientvit_b1, mobilenetv3</p></li>\n<li><p>Level 1: 2x efficientvit_b0, </p></li>\n</ul>\n<h4>Pipeline 3 - 5 seeds</h4>\n<ul>\n<li>Level 0  efficientvit_b1_288, 5x efficientnet_v2_s sed,  aves-base, aves-large </li>\n<li>Level 1 5x efficientvit_b0</li>\n</ul>\n<p>That ensemble scored 0.689970 private, and 0.742124 public.</p>\n<h2>Training</h2>\n<p>It was a bit tricky to tweak parameters without a validation set. We have two training pipelines with different parameters, and it is rather unclear what actually mattered. </p>\n<p>We used BCEWithLogitsLoss, without any label smoothing. Labels were defined by primary labels. Secondary labels were used to mask loss: loss for secondary labels is multiplied by 0. Reason is that we don’t know if a secondary label occurs at the start and at the end of the record, but they could. Given the uncertainty, we mask the loss for these. Masking secondary label loss improves LB by about 0.01.</p>\n<p>When using pseudo labels on unlabelled soundscapes we set their secondary labels to be empty. </p>\n<p>We used AdamW or Ranger optimizers. The number of epochs for the second level was way higher than for first level ones. For instance in one pipeline first level models are trained with 30 epochs while second level models are trained with 88 epochs.</p>\n<h2>Post Processing</h2>\n<p>We used several post processing, to incorporate soundscape-level information.</p>\n<p>The first one worked well on public LB with a 0.02 boost initially. Experiments after competition end show that its effect is much smaller (0.004 on our best unselected submission), and even detrimental on our best selected submission (- 0.002). The idea is that if a bird appears anywhere in the soundscape then its probability to appear in any 5 second slice is increased.<br>\nFor a given soundscape, once we have a 48x182 logits array or predictions P, we compute the maximum P_max of P over the time dimension. We then replace P with:<br>\nP + (P_max + P.mean() - P_max.mean()) * 0.8.<br>\nThe 0.8 weight can be tuned further.</p>\n<p>The second postprocessing had a smaller impact on LB the first time it was tried but it had a larger impact of +0.01 on our best selected sub. The idea is to smooth predictions for each 5 second clip by blending it a bit with the previous two and the following two clips. We used a convolution for smoothing predictions with the kernel [0.1, 0.2, 0.4, 0.2, 0.1].</p>\n<h2>What did not work</h2>\n<p>A lot. Main issue was that we could not find a reliable local validation scheme. We mainly looked at train data cross-validation like most of the participants, but also put some effort in looking at past BirdCLEF train / test differences individually for each year. We thought that this might help to understand what augmentations might be bridging the train (= xeno-canto) / test (= PAM soundscape recordings) gap. However, the results were quite inconclusive. Probably because soundscapes are quite different each year.  Whatever we tried lost correlation with the public LB once scores were high enough.</p>\n<p>One thing that seemed to work great and didn’t in the end was the post processing based on max probability per soundscape.</p>\n<p>Another one was the Aves model which improved public LB by about 0.01 each time it was used, but appeared to be detrimental on private LB.</p>\n<p>We tried a number of augmentations, including those documented in previous competitions or in academic papers, but none seemed to really help.</p>\n<p>ONNX did not bring much speedup (maybe 10%) and Openvino was much faster (almost 2x speedup compared to pytorch). However we noticed a drop of about 0.01 when using Openvino compared to ONNX, and decided not to use it. Maybe the speedup from openvino would have been compensated by the larger number of models we could use in the final ensemble?</p>\n<p>Our final model selection did not work that well in the end. We had individual models submitted with private scores above 0.70, but we did not include them in our ensemble given their relatively low public scores.</p>\n<p>Thanks for reading !</p>\n<p>Edit, link to code (alphabetical order):</p>\n<ul>\n<li><a href=\"https://github.com/jfpuget/birdclef-2024\" target=\"_blank\">CPMP's part</a></li>\n<li><a href=\"https://github.com/ChristofHenkel/kaggle-birdclef24-3rd-place-solution-dieter\" target=\"_blank\">Dieter's part</a></li>\n<li><a href=\"https://github.com/TheoViel/kaggle_birdclef2024\" target=\"_blank\">Theo's part</a></li>\n<li>Best selected submission: <a href=\"https://www.kaggle.com/cpmpml/birdclef-2024-inf-ens-08\" target=\"_blank\">https://www.kaggle.com/cpmpml/birdclef-2024-inf-ens-08</a></li>\n</ul>",
  "messages": [
    {
      "id": "2868721",
      "postDate": "06/12/2024 15:20:07",
      "content": "<p>I am the one posting but this is team work with <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> : team NVBird! </p>\n<p>Thanks everyone for the interesting competition, and congratulations to the winners ! We are very happy with the 3rd place, even though we missed the win by a very tiny margin.</p>\n<h2>Overview</h2>\n<p>Our pipeline is summarized below. A key ingredient of our solution is to use unlabeled soundscapes for pseudo labeling and model distillation. A number of models were trained with the training data, then used to predict labels on 5 second clips from unlabelled soundscapes. These were added to the original training data to train a new set of models used for the final submission. </p>\n<p><img src=\"https://i.ibb.co/QJQrdFf/Bird-CLEF-pipe.png\" alt=\"Bird-CLEF-pipe\"></p>\n<h2>Data</h2>\n<p>Overall, we relied on knowledge acquired during previous competitions, and added some extra samples to fight class imbalance.</p>\n<p>We use this year’s competition data plus the additional data from Xeno Canto shared in the forum, plus records from previous year competitions for the same species as this year.  When a record name appeared in several competitions we picked the most recent one (their content is not always identical form one competition to the next one)<br>\nWe capped the number of records per species to 500, keeping the most recent ones. Indeed, adding all extra data leads to severe class imbalance that was detrimental to model accuracy.<br>\nLow frequency classes are upsampled so that there are at least 10 samples for each class in the training folds. </p>\n<p>To train models, we use the following preprocessing and augmentations:</p>\n<p>We didn’t use all the training data. For each record we use a random crop of 5 seconds clip among the first 6 seconds, or the last 6 seconds of a record. If record length is smaller than 5 seconds, random padding so that the middle of the signal is between 2 and 3 seconds in the 5 sec resulting clip. For most of the models we used time shifting with a one second window as the only augmentation besides mixup. An exception are some models which are inspired by Birdclef23 2nd place SED models (<a href=\"https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution/blob/main/configs/sed_v2s.py)and\" target=\"_blank\">https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution/blob/main/configs/sed_v2s.py)and</a> use the same augmentations as used there<br>\nWe use an additive mixup: primary labels are the max of primary labels of the two audios to be mixed. Secondary labels are the concatenation of secondary labels.<br>\nWe mostly use image models that take log mel spectrograms as input. For these we compute mel spectrograms with parameters chosen to have an image size of 224x224 or 288x288 depending on the image model we use. Input waveforms are normalized to have a std of 1. </p>\n<h2>Models</h2>\n<h3>First Level models</h3>\n<p>The cpu-only requirement was quite constraining for submissions, but this does not apply for pseudo-label generation, so we could use more backbones for first level models. When ensembling several models larger than the ones used at second level we perform what is known as model distillation. This is a rather powerful technique in general.</p>\n<p>Models used include:</p>\n<ul>\n<li>Efficientvit_b0.224.in1k on 224x224 log mel spectrograms</li>\n<li>Efficientvit_b1.r288_in1k on 288x288 log mel spectrograms</li>\n<li>A variety of CNNs (efficientnets, mobilenets, tinynets, mnasnets, mixnets) and Efficientvits (<a href=\"https://arxiv.org/pdf/2205.14756\" target=\"_blank\">b0, b1</a>, <a href=\"https://arxiv.org/pdf/2305.07027\" target=\"_blank\">m3</a> trained on 224x224 log mel spectrograms.</li>\n<li>SED model with tf_efficientnetv2_s_in21k on 128x313  log mel spectrograms?</li>\n<li>We also fine-tuned aves-large and <a href=\"https://github.com/earthspecies/aves?tab=readme-ov-file#birdaves\" target=\"_blank\">aves-base</a>, which is a recent  waveform based model.</li>\n</ul>\n<p>Depending on the pipeline one or more of these models were used to predict pseudo labels on unlabelled soundscapes. We trained most models on full data with 5 different seeds.</p>\n<h3>Second Level models</h3>\n<p>We used efficientvit-b0 primarily and mnasnet-100 on 224x224 log mel spectrograms. Efficientvit-b0 showed great performances while still being very fast to infer. 5 folds take 40 minutes to submit using ONNX. We tried several models with similar throughput to effvit-b0, and decided to also use an mnasnet-100 for diversity. Since computing log mel spectrograms is slow on CPU we decided to use the same mel spectrogram hyperparameters for all models in the inference notebook, to only have to do log mel spectrogram transformation once and use the same input for our ensemble models.</p>\n<p>For training second level models we added the unlabelled soundscapes with the predicted pseudo labels to the training data. This looks simple but it took several attempts to find the correct way to do it. What worked fine was to use rather large batch sizes (128), <br>\nWe used the two following strategies;</p>\n<ul>\n<li>Add extra soundscapes to each batch. Each soundscape is split into 5 second clips, which means we added 48 x 4 = 192 clips with pseudo labels to the 128 samples with actual labels in each batch.</li>\n<li>Add 128 samples with pseudo labels, taken from random soundscapes this time. </li>\n</ul>\n<h3>Final ensemble</h3>\n<p>At the end we had an ensemble of 14 model weights via 3 pipelines, one pipeline per team member:</p>\n<h4>Pipeline 1 - 5 seeds</h4>\n<ul>\n<li>Level 0: efficientvit_b0 </li>\n<li>Level 1: efficientvit_b0</li>\n</ul>\n<h4>Pipeline 2 - 2 seeds + 2</h4>\n<ul>\n<li><p>Level 0: efficientvit_b1, mobilenetv2, efficientnet_b0, efficientnetv2_b0, efficientvit_b0, efficientvit_m3, aves-base, aves-large</p></li>\n<li><p>Level 1: 2x mnasnet-100</p></li>\n<li><p>Level 0: Efficientvit_b0, mixnet_s, mnasnet_100, tinynet_b, efficientvit_b0, efficientvit_b1, mobilenetv3</p></li>\n<li><p>Level 1: 2x efficientvit_b0, </p></li>\n</ul>\n<h4>Pipeline 3 - 5 seeds</h4>\n<ul>\n<li>Level 0  efficientvit_b1_288, 5x efficientnet_v2_s sed,  aves-base, aves-large </li>\n<li>Level 1 5x efficientvit_b0</li>\n</ul>\n<p>That ensemble scored 0.689970 private, and 0.742124 public.</p>\n<h2>Training</h2>\n<p>It was a bit tricky to tweak parameters without a validation set. We have two training pipelines with different parameters, and it is rather unclear what actually mattered. </p>\n<p>We used BCEWithLogitsLoss, without any label smoothing. Labels were defined by primary labels. Secondary labels were used to mask loss: loss for secondary labels is multiplied by 0. Reason is that we don’t know if a secondary label occurs at the start and at the end of the record, but they could. Given the uncertainty, we mask the loss for these. Masking secondary label loss improves LB by about 0.01.</p>\n<p>When using pseudo labels on unlabelled soundscapes we set their secondary labels to be empty. </p>\n<p>We used AdamW or Ranger optimizers. The number of epochs for the second level was way higher than for first level ones. For instance in one pipeline first level models are trained with 30 epochs while second level models are trained with 88 epochs.</p>\n<h2>Post Processing</h2>\n<p>We used several post processing, to incorporate soundscape-level information.</p>\n<p>The first one worked well on public LB with a 0.02 boost initially. Experiments after competition end show that its effect is much smaller (0.004 on our best unselected submission), and even detrimental on our best selected submission (- 0.002). The idea is that if a bird appears anywhere in the soundscape then its probability to appear in any 5 second slice is increased.<br>\nFor a given soundscape, once we have a 48x182 logits array or predictions P, we compute the maximum P_max of P over the time dimension. We then replace P with:<br>\nP + (P_max + P.mean() - P_max.mean()) * 0.8.<br>\nThe 0.8 weight can be tuned further.</p>\n<p>The second postprocessing had a smaller impact on LB the first time it was tried but it had a larger impact of +0.01 on our best selected sub. The idea is to smooth predictions for each 5 second clip by blending it a bit with the previous two and the following two clips. We used a convolution for smoothing predictions with the kernel [0.1, 0.2, 0.4, 0.2, 0.1].</p>\n<h2>What did not work</h2>\n<p>A lot. Main issue was that we could not find a reliable local validation scheme. We mainly looked at train data cross-validation like most of the participants, but also put some effort in looking at past BirdCLEF train / test differences individually for each year. We thought that this might help to understand what augmentations might be bridging the train (= xeno-canto) / test (= PAM soundscape recordings) gap. However, the results were quite inconclusive. Probably because soundscapes are quite different each year.  Whatever we tried lost correlation with the public LB once scores were high enough.</p>\n<p>One thing that seemed to work great and didn’t in the end was the post processing based on max probability per soundscape.</p>\n<p>Another one was the Aves model which improved public LB by about 0.01 each time it was used, but appeared to be detrimental on private LB.</p>\n<p>We tried a number of augmentations, including those documented in previous competitions or in academic papers, but none seemed to really help.</p>\n<p>ONNX did not bring much speedup (maybe 10%) and Openvino was much faster (almost 2x speedup compared to pytorch). However we noticed a drop of about 0.01 when using Openvino compared to ONNX, and decided not to use it. Maybe the speedup from openvino would have been compensated by the larger number of models we could use in the final ensemble?</p>\n<p>Our final model selection did not work that well in the end. We had individual models submitted with private scores above 0.70, but we did not include them in our ensemble given their relatively low public scores.</p>\n<p>Thanks for reading !</p>\n<p>Edit, link to code (alphabetical order):</p>\n<ul>\n<li><a href=\"https://github.com/jfpuget/birdclef-2024\" target=\"_blank\">CPMP's part</a></li>\n<li><a href=\"https://github.com/ChristofHenkel/kaggle-birdclef24-3rd-place-solution-dieter\" target=\"_blank\">Dieter's part</a></li>\n<li><a href=\"https://github.com/TheoViel/kaggle_birdclef2024\" target=\"_blank\">Theo's part</a></li>\n<li>Best selected submission: <a href=\"https://www.kaggle.com/cpmpml/birdclef-2024-inf-ens-08\" target=\"_blank\">https://www.kaggle.com/cpmpml/birdclef-2024-inf-ens-08</a></li>\n</ul>",
      "rawMarkdown": "I am the one posting but this is team work with @christofhenkel and @theoviel : team NVBird! \n\nThanks everyone for the interesting competition, and congratulations to the winners ! We are very happy with the 3rd place, even though we missed the win by a very tiny margin.\n\n## Overview\nOur pipeline is summarized below. A key ingredient of our solution is to use unlabeled soundscapes for pseudo labeling and model distillation. A number of models were trained with the training data, then used to predict labels on 5 second clips from unlabelled soundscapes. These were added to the original training data to train a new set of models used for the final submission. \n\n<img src=\"https://i.ibb.co/QJQrdFf/Bird-CLEF-pipe.png\" alt=\"Bird-CLEF-pipe\" border=\"0\">\n## Data\n\nOverall, we relied on knowledge acquired during previous competitions, and added some extra samples to fight class imbalance.\n\nWe use this year’s competition data plus the additional data from Xeno Canto shared in the forum, plus records from previous year competitions for the same species as this year.  When a record name appeared in several competitions we picked the most recent one (their content is not always identical form one competition to the next one)\nWe capped the number of records per species to 500, keeping the most recent ones. Indeed, adding all extra data leads to severe class imbalance that was detrimental to model accuracy.\nLow frequency classes are upsampled so that there are at least 10 samples for each class in the training folds. \n\nTo train models, we use the following preprocessing and augmentations:\n\nWe didn’t use all the training data. For each record we use a random crop of 5 seconds clip among the first 6 seconds, or the last 6 seconds of a record. If record length is smaller than 5 seconds, random padding so that the middle of the signal is between 2 and 3 seconds in the 5 sec resulting clip. For most of the models we used time shifting with a one second window as the only augmentation besides mixup. An exception are some models which are inspired by Birdclef23 2nd place SED models (https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution/blob/main/configs/sed_v2s.py)and use the same augmentations as used there\nWe use an additive mixup: primary labels are the max of primary labels of the two audios to be mixed. Secondary labels are the concatenation of secondary labels.\nWe mostly use image models that take log mel spectrograms as input. For these we compute mel spectrograms with parameters chosen to have an image size of 224x224 or 288x288 depending on the image model we use. Input waveforms are normalized to have a std of 1. \n## Models\n### First Level models\n\nThe cpu-only requirement was quite constraining for submissions, but this does not apply for pseudo-label generation, so we could use more backbones for first level models. When ensembling several models larger than the ones used at second level we perform what is known as model distillation. This is a rather powerful technique in general.\n\nModels used include:\n- Efficientvit_b0.224.in1k on 224x224 log mel spectrograms\n- Efficientvit_b1.r288_in1k on 288x288 log mel spectrograms\n- A variety of CNNs (efficientnets, mobilenets, tinynets, mnasnets, mixnets) and Efficientvits ([b0, b1](https://arxiv.org/pdf/2205.14756), [m3](https://arxiv.org/pdf/2305.07027) trained on 224x224 log mel spectrograms.\n- SED model with tf_efficientnetv2_s_in21k on 128x313  log mel spectrograms?\n- We also fine-tuned aves-large and [aves-base](https://github.com/earthspecies/aves?tab=readme-ov-file#birdaves), which is a recent  waveform based model.\n\nDepending on the pipeline one or more of these models were used to predict pseudo labels on unlabelled soundscapes. We trained most models on full data with 5 different seeds.\n\n### Second Level models\nWe used efficientvit-b0 primarily and mnasnet-100 on 224x224 log mel spectrograms. Efficientvit-b0 showed great performances while still being very fast to infer. 5 folds take 40 minutes to submit using ONNX. We tried several models with similar throughput to effvit-b0, and decided to also use an mnasnet-100 for diversity. Since computing log mel spectrograms is slow on CPU we decided to use the same mel spectrogram hyperparameters for all models in the inference notebook, to only have to do log mel spectrogram transformation once and use the same input for our ensemble models.\n\nFor training second level models we added the unlabelled soundscapes with the predicted pseudo labels to the training data. This looks simple but it took several attempts to find the correct way to do it. What worked fine was to use rather large batch sizes (128), \nWe used the two following strategies;\n- Add extra soundscapes to each batch. Each soundscape is split into 5 second clips, which means we added 48 x 4 = 192 clips with pseudo labels to the 128 samples with actual labels in each batch.\n- Add 128 samples with pseudo labels, taken from random soundscapes this time. \n\n\n### Final ensemble\n\nAt the end we had an ensemble of 14 model weights via 3 pipelines, one pipeline per team member:\n\n#### Pipeline 1 - 5 seeds\n- Level 0: efficientvit_b0 \n- Level 1: efficientvit_b0\n\n#### Pipeline 2 - 2 seeds + 2\n- Level 0: efficientvit_b1, mobilenetv2, efficientnet_b0, efficientnetv2_b0, efficientvit_b0, efficientvit_m3, aves-base, aves-large\n- Level 1: 2x mnasnet-100\n\n- Level 0: Efficientvit_b0, mixnet_s, mnasnet_100, tinynet_b, efficientvit_b0, efficientvit_b1, mobilenetv3\n- Level 1: 2x efficientvit_b0, \n\n#### Pipeline 3 - 5 seeds\n- Level 0  efficientvit_b1_288, 5x efficientnet_v2_s sed,  aves-base, aves-large \n- Level 1 5x efficientvit_b0\n\nThat ensemble scored 0.689970 private, and 0.742124 public.\n\n## Training\n\nIt was a bit tricky to tweak parameters without a validation set. We have two training pipelines with different parameters, and it is rather unclear what actually mattered. \n\nWe used BCEWithLogitsLoss, without any label smoothing. Labels were defined by primary labels. Secondary labels were used to mask loss: loss for secondary labels is multiplied by 0. Reason is that we don’t know if a secondary label occurs at the start and at the end of the record, but they could. Given the uncertainty, we mask the loss for these. Masking secondary label loss improves LB by about 0.01.\n\nWhen using pseudo labels on unlabelled soundscapes we set their secondary labels to be empty. \n\nWe used AdamW or Ranger optimizers. The number of epochs for the second level was way higher than for first level ones. For instance in one pipeline first level models are trained with 30 epochs while second level models are trained with 88 epochs.\n## Post Processing\n\nWe used several post processing, to incorporate soundscape-level information.\n\nThe first one worked well on public LB with a 0.02 boost initially. Experiments after competition end show that its effect is much smaller (0.004 on our best unselected submission), and even detrimental on our best selected submission (- 0.002). The idea is that if a bird appears anywhere in the soundscape then its probability to appear in any 5 second slice is increased.\nFor a given soundscape, once we have a 48x182 logits array or predictions P, we compute the maximum P_max of P over the time dimension. We then replace P with:\nP + (P_max + P.mean() - P_max.mean()) * 0.8.\nThe 0.8 weight can be tuned further.\n\nThe second postprocessing had a smaller impact on LB the first time it was tried but it had a larger impact of +0.01 on our best selected sub. The idea is to smooth predictions for each 5 second clip by blending it a bit with the previous two and the following two clips. We used a convolution for smoothing predictions with the kernel [0.1, 0.2, 0.4, 0.2, 0.1].\n\n## What did not work\n\nA lot. Main issue was that we could not find a reliable local validation scheme. We mainly looked at train data cross-validation like most of the participants, but also put some effort in looking at past BirdCLEF train / test differences individually for each year. We thought that this might help to understand what augmentations might be bridging the train (= xeno-canto) / test (= PAM soundscape recordings) gap. However, the results were quite inconclusive. Probably because soundscapes are quite different each year.  Whatever we tried lost correlation with the public LB once scores were high enough.\n\nOne thing that seemed to work great and didn’t in the end was the post processing based on max probability per soundscape.\n\nAnother one was the Aves model which improved public LB by about 0.01 each time it was used, but appeared to be detrimental on private LB.\n\nWe tried a number of augmentations, including those documented in previous competitions or in academic papers, but none seemed to really help.\n\nONNX did not bring much speedup (maybe 10%) and Openvino was much faster (almost 2x speedup compared to pytorch). However we noticed a drop of about 0.01 when using Openvino compared to ONNX, and decided not to use it. Maybe the speedup from openvino would have been compensated by the larger number of models we could use in the final ensemble?\n\nOur final model selection did not work that well in the end. We had individual models submitted with private scores above 0.70, but we did not include them in our ensemble given their relatively low public scores.\n\nThanks for reading !\n\nEdit, link to code (alphabetical order):\n- [CPMP's part](https://github.com/jfpuget/birdclef-2024)\n- [Dieter's part](https://github.com/ChristofHenkel/kaggle-birdclef24-3rd-place-solution-dieter)\n- [Theo's part](https://github.com/TheoViel/kaggle_birdclef2024)\n- Best selected submission: https://www.kaggle.com/cpmpml/birdclef-2024-inf-ens-08",
      "votes": null
    },
    {
      "id": "2869042",
      "postDate": "06/12/2024 19:40:27",
      "content": "<p>Congratulations with gold first of all and thanks for your clear writeup! </p>\n<p>How did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? For us, without ONNX it took around ~1h for a single fold, and with ONNX we observed a 2x speedup to around ~30 minutes per fold. Still around ~3.8x slower than your approach.</p>\n<p>How did you use the different seeds? Did you just pick the seed with highest public LB for submission?</p>\n<p>How did you exactly merge your pipelines, I am a bit confused by your ensembling section. Did you both submit level 0 and level 1, and which models were included?</p>\n<p>Cool to see that you also tried the convolving postprocessing smoothing, we observed that surprisingly smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel, wonder if you also tried that.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations with gold first of all and thanks for your clear writeup! \n\nHow did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? For us, without ONNX it took around ~1h for a single fold, and with ONNX we observed a 2x speedup to around ~30 minutes per fold. Still around ~3.8x slower than your approach.\n\nHow did you use the different seeds? Did you just pick the seed with highest public LB for submission?\n\nHow did you exactly merge your pipelines, I am a bit confused by your ensembling section. Did you both submit level 0 and level 1, and which models were included?\n\nCool to see that you also tried the convolving postprocessing smoothing, we observed that surprisingly smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel, wonder if you also tried that.\n\nThanks!",
      "votes": null
    },
    {
      "id": "2869088",
      "postDate": "06/12/2024 20:33:42",
      "content": "<p>Congratulation for the gold and thanks you for the write up !</p>\n<p>I have few questions:</p>\n<ul>\n<li>(1) regarding the pseudo labels approach, does smaller batch was also giving you positive results? I tried 16 but did not get any luck unfortunately … That was just providing more stable LB. </li>\n<li>(2) how did you pick which checkpoint to select for inference?</li>\n</ul>",
      "rawMarkdown": "Congratulation for the gold and thanks you for the write up !\n\nI have few questions:\n- (1) regarding the pseudo labels approach, does smaller batch was also giving you positive results? I tried 16 but did not get any luck unfortunately ... That was just providing more stable LB. \n- (2) how did you pick which checkpoint to select for inference?",
      "votes": null
    },
    {
      "id": "2869141",
      "postDate": "06/12/2024 21:18:33",
      "content": "<p>Congrats for the gold and thankyou for the writeup ! </p>\n<p>I have some questions aswell :-</p>\n<p>1) Regarding loss ( Secondary labels were used to mask loss ) that means you were not giving any importance to secondary lables at all ?</p>\n<p>2) Regarding Full data training ( We trained most models on full data ) how did you made sure that model is not overfitting ? i did the same but model got overfitted i guess ( i added lots of augs / background noise etc , tried to make my data hard to learn on classes ) which were abundant and also oversampled my other classes with some techniques . </p>\n<p>3) Regarding PsudoLabels ( what was your threshold to determine psudo labels and there can be multiple classes in the audio clip or the time frame which you were using to infer so i am assuming you might be using 1 label which has the max conf ?  I did the same but i guess my model was overfitting.</p>\n<p>4) Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ? I upsampled the classes by adding some augs to the original 5 sec clip and also added unlaballed soundscapes to some other copies , ( i took only those soundscape clips where my model was getting very low confidence like &lt; 0.05 ) but didnt got any luck probably becuase at the end i overfitted my model . </p>\n<p>Again Thankyou for your time and effort .</p>",
      "rawMarkdown": "Congrats for the gold and thankyou for the writeup ! \n\nI have some questions aswell :-\n\n1) Regarding loss ( Secondary labels were used to mask loss ) that means you were not giving any importance to secondary lables at all ?\n\n2) Regarding Full data training ( We trained most models on full data ) how did you made sure that model is not overfitting ? i did the same but model got overfitted i guess ( i added lots of augs / background noise etc , tried to make my data hard to learn on classes ) which were abundant and also oversampled my other classes with some techniques . \n\n3) Regarding PsudoLabels ( what was your threshold to determine psudo labels and there can be multiple classes in the audio clip or the time frame which you were using to infer so i am assuming you might be using 1 label which has the max conf ?  I did the same but i guess my model was overfitting.\n\n4) Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ? I upsampled the classes by adding some augs to the original 5 sec clip and also added unlaballed soundscapes to some other copies , ( i took only those soundscape clips where my model was getting very low confidence like < 0.05 ) but didnt got any luck probably becuase at the end i overfitted my model . \n\nAgain Thankyou for your time and effort .",
      "votes": null
    },
    {
      "id": "2869205",
      "postDate": "06/12/2024 23:27:19",
      "content": "<p>Wow great work! I may have to go back to the board a bit. I had used efficientvit_b0 and it took 25 min to run on ONNX. 5 in 40 minutes is crazy, I couldnt get any of my models to be that fast! I would love to hear more about how that worked</p>",
      "rawMarkdown": "Wow great work! I may have to go back to the board a bit. I had used efficientvit_b0 and it took 25 min to run on ONNX. 5 in 40 minutes is crazy, I couldnt get any of my models to be that fast! I would love to hear more about how that worked",
      "votes": null
    },
    {
      "id": "2869275",
      "postDate": "06/13/2024 01:48:59",
      "content": "<blockquote>\n  <p>Congratulations with gold first of all and thanks for your clear writeup! </p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>How did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? </p>\n</blockquote>\n<p>Efficientvit_b0, not efficientnet_b0 :D</p>\n<blockquote>\n  <p>How did you use the different seeds? Did you just pick the seed with highest public LB for submission?</p>\n</blockquote>\n<p>Certainly not. This is a recipe for overfitting. We either used random seeds, or fixed seeds we change from time to time (team mates have differnent views here, but we all agree that tuning the seed is a disaster)</p>\n<blockquote>\n  <p>How did you exactly merge your pipelines,</p>\n</blockquote>\n<p>We took the average of the logits of the 14 models.</p>\n<blockquote>\n  <p>smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel,</p>\n</blockquote>\n<p>We tried and found that smoothing with the max over the soundscape is even better. Till we saw that it isn't with high score on private LB.</p>",
      "rawMarkdown": "> Congratulations with gold first of all and thanks for your clear writeup! \n\nThank you on behalf of the team.\n\n> How did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? \n\nEfficientvit_b0, not efficientnet_b0 :D\n\n> How did you use the different seeds? Did you just pick the seed with highest public LB for submission?\n\nCertainly not. This is a recipe for overfitting. We either used random seeds, or fixed seeds we change from time to time (team mates have differnent views here, but we all agree that tuning the seed is a disaster)\n\n> How did you exactly merge your pipelines,\n\nWe took the average of the logits of the 14 models.\n\n> smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel,\n\nWe tried and found that smoothing with the max over the soundscape is even better. Till we saw that it isn't with high score on private LB.",
      "votes": null
    },
    {
      "id": "2869277",
      "postDate": "06/13/2024 01:50:48",
      "content": "<blockquote>\n  <p>Congratulation for the gold and thanks you for the write up !</p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>does smaller batch was also giving you positive results? </p>\n</blockquote>\n<p>We tried as large batch as possible. 256 seemed to lower performance hence we stopped at 128. We then found that the more pseudo labels per batch the better. Maybe we could have gone further.</p>\n<blockquote>\n  <p>how did you pick which checkpoint to select for inference?</p>\n</blockquote>\n<p>The last epoch one. We don't use early stopping.</p>",
      "rawMarkdown": "> Congratulation for the gold and thanks you for the write up !\n\nThank you on behalf of the team.\n\n> does smaller batch was also giving you positive results? \n\nWe tried as large batch as possible. 256 seemed to lower performance hence we stopped at 128. We then found that the more pseudo labels per batch the better. Maybe we could have gone further.\n\n> how did you pick which checkpoint to select for inference?\n\nThe last epoch one. We don't use early stopping.",
      "votes": null
    },
    {
      "id": "2869281",
      "postDate": "06/13/2024 01:58:12",
      "content": "<blockquote>\n  <p>Congrats for the gold and thankyou for the writeup ! </p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>that means you were not giving any importance to secondary lables at all ?</p>\n</blockquote>\n<p>I think yes (not sure about the meaning of \"importance\"). We do find secondary labels to be important as we use them. </p>\n<blockquote>\n  <p>how did you made sure that model is not overfitting ?</p>\n</blockquote>\n<p>We didn't really. We did train models using 5 fold CV, but the CV score wasn't correlated with LB score once we reached 0.65 or above. on LB. I welcome any hint on how to properly evaluate models in this competition besides submitting them.</p>\n<blockquote>\n  <p>Regarding PsudoLabels ( what was your threshold to determine psudo label</p>\n</blockquote>\n<p>We use the predicted probability as label. We don't round probabilities to binary labels.</p>\n<blockquote>\n  <p>i am assuming you might be using 1 label which has the max conf ?</p>\n</blockquote>\n<p>We don't. The competition is a mult label competition, not a multi class competition. </p>\n<blockquote>\n  <p>Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ?</p>\n</blockquote>\n<p>We duplicated the record, then augmentations (time shift and mixup) are applied independently on each copy.</p>",
      "rawMarkdown": "> Congrats for the gold and thankyou for the writeup ! \n\nThank you on behalf of the team.\n\n> that means you were not giving any importance to secondary lables at all ?\n\nI think yes (not sure about the meaning of \"importance\"). We do find secondary labels to be important as we use them. \n\n> how did you made sure that model is not overfitting ?\n\nWe didn't really. We did train models using 5 fold CV, but the CV score wasn't correlated with LB score once we reached 0.65 or above. on LB. I welcome any hint on how to properly evaluate models in this competition besides submitting them.\n\n> Regarding PsudoLabels ( what was your threshold to determine psudo label\n\nWe use the predicted probability as label. We don't round probabilities to binary labels.\n\n> i am assuming you might be using 1 label which has the max conf ?\n\nWe don't. The competition is a mult label competition, not a multi class competition. \n\n> Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ?\n\nWe duplicated the record, then augmentations (time shift and mixup) are applied independently on each copy.",
      "votes": null
    },
    {
      "id": "2869282",
      "postDate": "06/13/2024 02:00:03",
      "content": "<blockquote>\n  <p>Wow great work! </p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>I would love to hear more about how that worked</p>\n</blockquote>\n<p>We didn't do anything fancy there. We used the model.half().float() trick that was shared, this shaved few minutes per 5 checkpoints. But we use various tricks like loading records once for all models, computing mel spectrograms once for all models, etc.</p>",
      "rawMarkdown": "> Wow great work! \n\nThank you on behalf of the team.\n\n> I would love to hear more about how that worked\n\nWe didn't do anything fancy there. We used the model.half().float() trick that was shared, this shaved few minutes per 5 checkpoints. But we use various tricks like loading records once for all models, computing mel spectrograms once for all models, etc.",
      "votes": null
    },
    {
      "id": "2869285",
      "postDate": "06/13/2024 02:08:40",
      "content": "<p>Congrats for the gold and 3rd position and thankyou for the writeup !!!</p>",
      "rawMarkdown": "Congrats for the gold and 3rd position and thankyou for the writeup !!!",
      "votes": null
    },
    {
      "id": "2869751",
      "postDate": "06/13/2024 08:44:04",
      "content": "<p>Ah great! Thanks for the clarification! </p>\n<p>And then I assume that you submitted all Level 1 models from each pipeline together, and then you used a varied ensemble of bigger models for Level 0 to get better pseudo labelling for Level 1? That is a very nice idea, we tried pseudo-labelling with the Google model on unlabelled soundscapes. That did not work out for us, unfortunately. Your approach on training multiple larger models on Xeno-Canto and using that for pseudolabelling seems to be a better approach :)</p>",
      "rawMarkdown": "Ah great! Thanks for the clarification! \n\nAnd then I assume that you submitted all Level 1 models from each pipeline together, and then you used a varied ensemble of bigger models for Level 0 to get better pseudo labelling for Level 1? That is a very nice idea, we tried pseudo-labelling with the Google model on unlabelled soundscapes. That did not work out for us, unfortunately. Your approach on training multiple larger models on Xeno-Canto and using that for pseudolabelling seems to be a better approach :)",
      "votes": null
    },
    {
      "id": "2869827",
      "postDate": "06/13/2024 09:47:48",
      "content": "<p>Actually the best second level model we selected was using pseudo labels from effcientvit_b0 only!</p>\n<p>The other two pipelines used Aves and looked great on public LB, but were not great on private LB unfortunately.</p>\n<p>Best unselected pipelines were similar to those in the ensemble but without Aves.</p>",
      "rawMarkdown": "Actually the best second level model we selected was using pseudo labels from effcientvit_b0 only!\n\nThe other two pipelines used Aves and looked great on public LB, but were not great on private LB unfortunately.\n\nBest unselected pipelines were similar to those in the ensemble but without Aves.",
      "votes": null
    },
    {
      "id": "2869834",
      "postDate": "06/13/2024 09:50:54",
      "content": "<p>Congratulations and great writeup. </p>\n<p>Could you clarify a little about masking loss for secondary labels? From what I can understand, if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss? Is this the correct interpretation?</p>\n<p>Also, will you be opening up your inference sub for us to look at optimizations and stuff? 😐</p>",
      "rawMarkdown": "Congratulations and great writeup. \n\nCould you clarify a little about masking loss for secondary labels? From what I can understand, if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss? Is this the correct interpretation?\n\nAlso, will you be opening up your inference sub for us to look at optimizations and stuff? 😐",
      "votes": null
    },
    {
      "id": "2869841",
      "postDate": "06/13/2024 09:57:59",
      "content": "<p>Thankyou for the answers .</p>\n<blockquote>\n  <p>We don't. The competition is a mult label competition, not a multi class competition.</p>\n</blockquote>\n<p>What i meant by that was how you were deciding the primary label for your psudo labels in the window of your time frame and i guess it will be your max confidence class ? since you are setting your secondary classes as empty as mentioned in your write up. </p>",
      "rawMarkdown": "Thankyou for the answers .\n\n>We don't. The competition is a mult label competition, not a multi class competition.\n\nWhat i meant by that was how you were deciding the primary label for your psudo labels in the window of your time frame and i guess it will be your max confidence class ? since you are setting your secondary classes as empty as mentioned in your write up.",
      "votes": null
    },
    {
      "id": "2869884",
      "postDate": "06/13/2024 10:34:04",
      "content": "<blockquote>\n  <p>Congratulations and great writeup. </p>\n</blockquote>\n<p>Thanks on behalf of the team.</p>\n<blockquote>\n  <p>if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss?</p>\n</blockquote>\n<p>Yes, we mask as you say.</p>\n<p>I documented this idea in <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183219\" target=\"_blank\">my 2020 Birdclef writeup</a>. I reproduce it here:</p>\n<blockquote>\n  <p>Primary labels were noisy, but secondary labels were even noisier. As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know. We therefore defined a secondary mask that nullifies the BCE loss for secondary labels. For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:</p>\n  <p>[1, 0, 0]</p>\n  <p>[1, 1, 0]</p>\n  <p>The secondary mask is therefore:</p>\n  <p>[1, 0, 1]</p>\n</blockquote>\n<p>I hope this is clearer now.</p>\n<blockquote>\n  <p>will you be opening up your inference sub </p>\n</blockquote>\n<p>I'll check with my team members.</p>",
      "rawMarkdown": "> Congratulations and great writeup. \n\nThanks on behalf of the team.\n\n> if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss?\n\nYes, we mask as you say.\n\nI documented this idea in [my 2020 Birdclef writeup](https://www.kaggle.com/competitions/birdsong-recognition/discussion/183219). I reproduce it here:\n\n>Primary labels were noisy, but secondary labels were even noisier. As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know. We therefore defined a secondary mask that nullifies the BCE loss for secondary labels. For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:\n> \n> [1, 0, 0]\n> \n> [1, 1, 0]\n> \n> The secondary mask is therefore:\n> \n> [1, 0, 1]\n\nI hope this is clearer now.\n\n> will you be opening up your inference sub \n\nI'll check with my team members.",
      "votes": null
    },
    {
      "id": "2870035",
      "postDate": "06/13/2024 11:40:34",
      "content": "<p>We use the predicted probabilities as pseudo labels.  I am not sure what is not clear here.</p>",
      "rawMarkdown": "We use the predicted probabilities as pseudo labels.  I am not sure what is not clear here.",
      "votes": null
    },
    {
      "id": "2870091",
      "postDate": "06/13/2024 12:08:59",
      "content": "<p>my bad that i portray my question like that , so you were using all the predicted probabilities of your <strong>psudolabels</strong> for a given audio clip for exa:- 5 sec clip with 182 predicted class probabilites for training  and not the max class probability value alone and setting other class prob value as 0 .</p>",
      "rawMarkdown": "my bad that i portray my question like that , so you were using all the predicted probabilities of your **psudolabels** for a given audio clip for exa:- 5 sec clip with 182 predicted class probabilites for training  and not the max class probability value alone and setting other class prob value as 0 .",
      "votes": null
    },
    {
      "id": "2870152",
      "postDate": "06/13/2024 12:36:57",
      "content": "<p>Thankyou! </p>",
      "rawMarkdown": "Thankyou!",
      "votes": null
    },
    {
      "id": "2870268",
      "postDate": "06/13/2024 13:46:12",
      "content": "<p>We were using all the probabilities. This is the only way one can use BCE loss.</p>",
      "rawMarkdown": "We were using all the probabilities. This is the only way one can use BCE loss.",
      "votes": null
    },
    {
      "id": "2870554",
      "postDate": "06/13/2024 17:07:38",
      "content": "<p>Congratulations and nice writeup.</p>",
      "rawMarkdown": "Congratulations and nice writeup.",
      "votes": null
    },
    {
      "id": "2870600",
      "postDate": "06/13/2024 17:38:23",
      "content": "<p>melspec transformation takes quite some time. But even we had 5 models we only need to do it once. So you can check how long multiple effnet b0 of you would take, if you only do melspec once </p>",
      "rawMarkdown": "melspec transformation takes quite some time. But even we had 5 models we only need to do it once. So you can check how long multiple effnet b0 of you would take, if you only do melspec once",
      "votes": null
    },
    {
      "id": "2872135",
      "postDate": "06/14/2024 15:45:26",
      "content": "<p>Congratulations 🎊 </p>",
      "rawMarkdown": "Congratulations 🎊",
      "votes": null
    },
    {
      "id": "2877150",
      "postDate": "06/18/2024 07:57:18",
      "content": "<p>Thanks for the insight! Can you also share the code?</p>",
      "rawMarkdown": "Thanks for the insight! Can you also share the code?",
      "votes": null
    },
    {
      "id": "2889069",
      "postDate": "06/25/2024 08:20:32",
      "content": "<p>We shared our code, link in the post.</p>",
      "rawMarkdown": "We shared our code, link in the post.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2869042,
      "author_name": "hugodeheer",
      "author_url": "",
      "post_date": "06/12/2024 19:40:27",
      "content": "<p>Congratulations with gold first of all and thanks for your clear writeup! </p>\n<p>How did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? For us, without ONNX it took around ~1h for a single fold, and with ONNX we observed a 2x speedup to around ~30 minutes per fold. Still around ~3.8x slower than your approach.</p>\n<p>How did you use the different seeds? Did you just pick the seed with highest public LB for submission?</p>\n<p>How did you exactly merge your pipelines, I am a bit confused by your ensembling section. Did you both submit level 0 and level 1, and which models were included?</p>\n<p>Cool to see that you also tried the convolving postprocessing smoothing, we observed that surprisingly smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel, wonder if you also tried that.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2869275,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/13/2024 01:48:59",
          "content": "<blockquote>\n  <p>Congratulations with gold first of all and thanks for your clear writeup! </p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>How did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? </p>\n</blockquote>\n<p>Efficientvit_b0, not efficientnet_b0 :D</p>\n<blockquote>\n  <p>How did you use the different seeds? Did you just pick the seed with highest public LB for submission?</p>\n</blockquote>\n<p>Certainly not. This is a recipe for overfitting. We either used random seeds, or fixed seeds we change from time to time (team mates have differnent views here, but we all agree that tuning the seed is a disaster)</p>\n<blockquote>\n  <p>How did you exactly merge your pipelines,</p>\n</blockquote>\n<p>We took the average of the logits of the 14 models.</p>\n<blockquote>\n  <p>smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel,</p>\n</blockquote>\n<p>We tried and found that smoothing with the max over the soundscape is even better. Till we saw that it isn't with high score on private LB.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2869751,
              "author_name": "hugodeheer",
              "author_url": "",
              "post_date": "06/13/2024 08:44:04",
              "content": "<p>Ah great! Thanks for the clarification! </p>\n<p>And then I assume that you submitted all Level 1 models from each pipeline together, and then you used a varied ensemble of bigger models for Level 0 to get better pseudo labelling for Level 1? That is a very nice idea, we tried pseudo-labelling with the Google model on unlabelled soundscapes. That did not work out for us, unfortunately. Your approach on training multiple larger models on Xeno-Canto and using that for pseudolabelling seems to be a better approach :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2869827,
                  "author_name": "cpmpml",
                  "author_url": "",
                  "post_date": "06/13/2024 09:47:48",
                  "content": "<p>Actually the best second level model we selected was using pseudo labels from effcientvit_b0 only!</p>\n<p>The other two pipelines used Aves and looked great on public LB, but were not great on private LB unfortunately.</p>\n<p>Best unselected pipelines were similar to those in the ensemble but without Aves.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2869088,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "06/12/2024 20:33:42",
      "content": "<p>Congratulation for the gold and thanks you for the write up !</p>\n<p>I have few questions:</p>\n<ul>\n<li>(1) regarding the pseudo labels approach, does smaller batch was also giving you positive results? I tried 16 but did not get any luck unfortunately … That was just providing more stable LB. </li>\n<li>(2) how did you pick which checkpoint to select for inference?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2869277,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/13/2024 01:50:48",
          "content": "<blockquote>\n  <p>Congratulation for the gold and thanks you for the write up !</p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>does smaller batch was also giving you positive results? </p>\n</blockquote>\n<p>We tried as large batch as possible. 256 seemed to lower performance hence we stopped at 128. We then found that the more pseudo labels per batch the better. Maybe we could have gone further.</p>\n<blockquote>\n  <p>how did you pick which checkpoint to select for inference?</p>\n</blockquote>\n<p>The last epoch one. We don't use early stopping.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2869141,
      "author_name": "trooperog",
      "author_url": "",
      "post_date": "06/12/2024 21:18:33",
      "content": "<p>Congrats for the gold and thankyou for the writeup ! </p>\n<p>I have some questions aswell :-</p>\n<p>1) Regarding loss ( Secondary labels were used to mask loss ) that means you were not giving any importance to secondary lables at all ?</p>\n<p>2) Regarding Full data training ( We trained most models on full data ) how did you made sure that model is not overfitting ? i did the same but model got overfitted i guess ( i added lots of augs / background noise etc , tried to make my data hard to learn on classes ) which were abundant and also oversampled my other classes with some techniques . </p>\n<p>3) Regarding PsudoLabels ( what was your threshold to determine psudo labels and there can be multiple classes in the audio clip or the time frame which you were using to infer so i am assuming you might be using 1 label which has the max conf ?  I did the same but i guess my model was overfitting.</p>\n<p>4) Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ? I upsampled the classes by adding some augs to the original 5 sec clip and also added unlaballed soundscapes to some other copies , ( i took only those soundscape clips where my model was getting very low confidence like &lt; 0.05 ) but didnt got any luck probably becuase at the end i overfitted my model . </p>\n<p>Again Thankyou for your time and effort .</p>",
      "votes": null,
      "replies": [
        {
          "id": 2869281,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/13/2024 01:58:12",
          "content": "<blockquote>\n  <p>Congrats for the gold and thankyou for the writeup ! </p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>that means you were not giving any importance to secondary lables at all ?</p>\n</blockquote>\n<p>I think yes (not sure about the meaning of \"importance\"). We do find secondary labels to be important as we use them. </p>\n<blockquote>\n  <p>how did you made sure that model is not overfitting ?</p>\n</blockquote>\n<p>We didn't really. We did train models using 5 fold CV, but the CV score wasn't correlated with LB score once we reached 0.65 or above. on LB. I welcome any hint on how to properly evaluate models in this competition besides submitting them.</p>\n<blockquote>\n  <p>Regarding PsudoLabels ( what was your threshold to determine psudo label</p>\n</blockquote>\n<p>We use the predicted probability as label. We don't round probabilities to binary labels.</p>\n<blockquote>\n  <p>i am assuming you might be using 1 label which has the max conf ?</p>\n</blockquote>\n<p>We don't. The competition is a mult label competition, not a multi class competition. </p>\n<blockquote>\n  <p>Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ?</p>\n</blockquote>\n<p>We duplicated the record, then augmentations (time shift and mixup) are applied independently on each copy.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2869841,
              "author_name": "trooperog",
              "author_url": "",
              "post_date": "06/13/2024 09:57:59",
              "content": "<p>Thankyou for the answers .</p>\n<blockquote>\n  <p>We don't. The competition is a mult label competition, not a multi class competition.</p>\n</blockquote>\n<p>What i meant by that was how you were deciding the primary label for your psudo labels in the window of your time frame and i guess it will be your max confidence class ? since you are setting your secondary classes as empty as mentioned in your write up. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2870035,
                  "author_name": "cpmpml",
                  "author_url": "",
                  "post_date": "06/13/2024 11:40:34",
                  "content": "<p>We use the predicted probabilities as pseudo labels.  I am not sure what is not clear here.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2870091,
                      "author_name": "trooperog",
                      "author_url": "",
                      "post_date": "06/13/2024 12:08:59",
                      "content": "<p>my bad that i portray my question like that , so you were using all the predicted probabilities of your <strong>psudolabels</strong> for a given audio clip for exa:- 5 sec clip with 182 predicted class probabilites for training  and not the max class probability value alone and setting other class prob value as 0 .</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2870268,
                          "author_name": "cpmpml",
                          "author_url": "",
                          "post_date": "06/13/2024 13:46:12",
                          "content": "<p>We were using all the probabilities. This is the only way one can use BCE loss.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2869205,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "06/12/2024 23:27:19",
      "content": "<p>Wow great work! I may have to go back to the board a bit. I had used efficientvit_b0 and it took 25 min to run on ONNX. 5 in 40 minutes is crazy, I couldnt get any of my models to be that fast! I would love to hear more about how that worked</p>",
      "votes": null,
      "replies": [
        {
          "id": 2869282,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/13/2024 02:00:03",
          "content": "<blockquote>\n  <p>Wow great work! </p>\n</blockquote>\n<p>Thank you on behalf of the team.</p>\n<blockquote>\n  <p>I would love to hear more about how that worked</p>\n</blockquote>\n<p>We didn't do anything fancy there. We used the model.half().float() trick that was shared, this shaved few minutes per 5 checkpoints. But we use various tricks like loading records once for all models, computing mel spectrograms once for all models, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2870600,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/13/2024 17:38:23",
          "content": "<p>melspec transformation takes quite some time. But even we had 5 models we only need to do it once. So you can check how long multiple effnet b0 of you would take, if you only do melspec once </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2869285,
      "author_name": "aadityaporwal",
      "author_url": "",
      "post_date": "06/13/2024 02:08:40",
      "content": "<p>Congrats for the gold and 3rd position and thankyou for the writeup !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2869834,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "06/13/2024 09:50:54",
      "content": "<p>Congratulations and great writeup. </p>\n<p>Could you clarify a little about masking loss for secondary labels? From what I can understand, if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss? Is this the correct interpretation?</p>\n<p>Also, will you be opening up your inference sub for us to look at optimizations and stuff? 😐</p>",
      "votes": null,
      "replies": [
        {
          "id": 2869884,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/13/2024 10:34:04",
          "content": "<blockquote>\n  <p>Congratulations and great writeup. </p>\n</blockquote>\n<p>Thanks on behalf of the team.</p>\n<blockquote>\n  <p>if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss?</p>\n</blockquote>\n<p>Yes, we mask as you say.</p>\n<p>I documented this idea in <a href=\"https://www.kaggle.com/competitions/birdsong-recognition/discussion/183219\" target=\"_blank\">my 2020 Birdclef writeup</a>. I reproduce it here:</p>\n<blockquote>\n  <p>Primary labels were noisy, but secondary labels were even noisier. As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know. We therefore defined a secondary mask that nullifies the BCE loss for secondary labels. For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:</p>\n  <p>[1, 0, 0]</p>\n  <p>[1, 1, 0]</p>\n  <p>The secondary mask is therefore:</p>\n  <p>[1, 0, 1]</p>\n</blockquote>\n<p>I hope this is clearer now.</p>\n<blockquote>\n  <p>will you be opening up your inference sub </p>\n</blockquote>\n<p>I'll check with my team members.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2870152,
              "author_name": "pheadrus",
              "author_url": "",
              "post_date": "06/13/2024 12:36:57",
              "content": "<p>Thankyou! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2870554,
      "author_name": "engrmazeem",
      "author_url": "",
      "post_date": "06/13/2024 17:07:38",
      "content": "<p>Congratulations and nice writeup.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2872135,
      "author_name": "metinmekiabullrahman",
      "author_url": "",
      "post_date": "06/14/2024 15:45:26",
      "content": "<p>Congratulations 🎊 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2877150,
      "author_name": "anirudhmittal0202",
      "author_url": "",
      "post_date": "06/18/2024 07:57:18",
      "content": "<p>Thanks for the insight! Can you also share the code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2889069,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/25/2024 08:20:32",
      "content": "<p>We shared our code, link in the post.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2868721": "I am the one posting but this is team work with @christofhenkel and @theoviel : team NVBird! \n\nThanks everyone for the interesting competition, and congratulations to the winners ! We are very happy with the 3rd place, even though we missed the win by a very tiny margin.\n\n## Overview\nOur pipeline is summarized below. A key ingredient of our solution is to use unlabeled soundscapes for pseudo labeling and model distillation. A number of models were trained with the training data, then used to predict labels on 5 second clips from unlabelled soundscapes. These were added to the original training data to train a new set of models used for the final submission. \n\n<img src=\"https://i.ibb.co/QJQrdFf/Bird-CLEF-pipe.png\" alt=\"Bird-CLEF-pipe\" border=\"0\">\n## Data\n\nOverall, we relied on knowledge acquired during previous competitions, and added some extra samples to fight class imbalance.\n\nWe use this year’s competition data plus the additional data from Xeno Canto shared in the forum, plus records from previous year competitions for the same species as this year.  When a record name appeared in several competitions we picked the most recent one (their content is not always identical form one competition to the next one)\nWe capped the number of records per species to 500, keeping the most recent ones. Indeed, adding all extra data leads to severe class imbalance that was detrimental to model accuracy.\nLow frequency classes are upsampled so that there are at least 10 samples for each class in the training folds. \n\nTo train models, we use the following preprocessing and augmentations:\n\nWe didn’t use all the training data. For each record we use a random crop of 5 seconds clip among the first 6 seconds, or the last 6 seconds of a record. If record length is smaller than 5 seconds, random padding so that the middle of the signal is between 2 and 3 seconds in the 5 sec resulting clip. For most of the models we used time shifting with a one second window as the only augmentation besides mixup. An exception are some models which are inspired by Birdclef23 2nd place SED models (https://github.com/LIHANG-HONG/birdclef2023-2nd-place-solution/blob/main/configs/sed_v2s.py)and use the same augmentations as used there\nWe use an additive mixup: primary labels are the max of primary labels of the two audios to be mixed. Secondary labels are the concatenation of secondary labels.\nWe mostly use image models that take log mel spectrograms as input. For these we compute mel spectrograms with parameters chosen to have an image size of 224x224 or 288x288 depending on the image model we use. Input waveforms are normalized to have a std of 1. \n## Models\n### First Level models\n\nThe cpu-only requirement was quite constraining for submissions, but this does not apply for pseudo-label generation, so we could use more backbones for first level models. When ensembling several models larger than the ones used at second level we perform what is known as model distillation. This is a rather powerful technique in general.\n\nModels used include:\n- Efficientvit_b0.224.in1k on 224x224 log mel spectrograms\n- Efficientvit_b1.r288_in1k on 288x288 log mel spectrograms\n- A variety of CNNs (efficientnets, mobilenets, tinynets, mnasnets, mixnets) and Efficientvits ([b0, b1](https://arxiv.org/pdf/2205.14756), [m3](https://arxiv.org/pdf/2305.07027) trained on 224x224 log mel spectrograms.\n- SED model with tf_efficientnetv2_s_in21k on 128x313  log mel spectrograms?\n- We also fine-tuned aves-large and [aves-base](https://github.com/earthspecies/aves?tab=readme-ov-file#birdaves), which is a recent  waveform based model.\n\nDepending on the pipeline one or more of these models were used to predict pseudo labels on unlabelled soundscapes. We trained most models on full data with 5 different seeds.\n\n### Second Level models\nWe used efficientvit-b0 primarily and mnasnet-100 on 224x224 log mel spectrograms. Efficientvit-b0 showed great performances while still being very fast to infer. 5 folds take 40 minutes to submit using ONNX. We tried several models with similar throughput to effvit-b0, and decided to also use an mnasnet-100 for diversity. Since computing log mel spectrograms is slow on CPU we decided to use the same mel spectrogram hyperparameters for all models in the inference notebook, to only have to do log mel spectrogram transformation once and use the same input for our ensemble models.\n\nFor training second level models we added the unlabelled soundscapes with the predicted pseudo labels to the training data. This looks simple but it took several attempts to find the correct way to do it. What worked fine was to use rather large batch sizes (128), \nWe used the two following strategies;\n- Add extra soundscapes to each batch. Each soundscape is split into 5 second clips, which means we added 48 x 4 = 192 clips with pseudo labels to the 128 samples with actual labels in each batch.\n- Add 128 samples with pseudo labels, taken from random soundscapes this time. \n\n\n### Final ensemble\n\nAt the end we had an ensemble of 14 model weights via 3 pipelines, one pipeline per team member:\n\n#### Pipeline 1 - 5 seeds\n- Level 0: efficientvit_b0 \n- Level 1: efficientvit_b0\n\n#### Pipeline 2 - 2 seeds + 2\n- Level 0: efficientvit_b1, mobilenetv2, efficientnet_b0, efficientnetv2_b0, efficientvit_b0, efficientvit_m3, aves-base, aves-large\n- Level 1: 2x mnasnet-100\n\n- Level 0: Efficientvit_b0, mixnet_s, mnasnet_100, tinynet_b, efficientvit_b0, efficientvit_b1, mobilenetv3\n- Level 1: 2x efficientvit_b0, \n\n#### Pipeline 3 - 5 seeds\n- Level 0  efficientvit_b1_288, 5x efficientnet_v2_s sed,  aves-base, aves-large \n- Level 1 5x efficientvit_b0\n\nThat ensemble scored 0.689970 private, and 0.742124 public.\n\n## Training\n\nIt was a bit tricky to tweak parameters without a validation set. We have two training pipelines with different parameters, and it is rather unclear what actually mattered. \n\nWe used BCEWithLogitsLoss, without any label smoothing. Labels were defined by primary labels. Secondary labels were used to mask loss: loss for secondary labels is multiplied by 0. Reason is that we don’t know if a secondary label occurs at the start and at the end of the record, but they could. Given the uncertainty, we mask the loss for these. Masking secondary label loss improves LB by about 0.01.\n\nWhen using pseudo labels on unlabelled soundscapes we set their secondary labels to be empty. \n\nWe used AdamW or Ranger optimizers. The number of epochs for the second level was way higher than for first level ones. For instance in one pipeline first level models are trained with 30 epochs while second level models are trained with 88 epochs.\n## Post Processing\n\nWe used several post processing, to incorporate soundscape-level information.\n\nThe first one worked well on public LB with a 0.02 boost initially. Experiments after competition end show that its effect is much smaller (0.004 on our best unselected submission), and even detrimental on our best selected submission (- 0.002). The idea is that if a bird appears anywhere in the soundscape then its probability to appear in any 5 second slice is increased.\nFor a given soundscape, once we have a 48x182 logits array or predictions P, we compute the maximum P_max of P over the time dimension. We then replace P with:\nP + (P_max + P.mean() - P_max.mean()) * 0.8.\nThe 0.8 weight can be tuned further.\n\nThe second postprocessing had a smaller impact on LB the first time it was tried but it had a larger impact of +0.01 on our best selected sub. The idea is to smooth predictions for each 5 second clip by blending it a bit with the previous two and the following two clips. We used a convolution for smoothing predictions with the kernel [0.1, 0.2, 0.4, 0.2, 0.1].\n\n## What did not work\n\nA lot. Main issue was that we could not find a reliable local validation scheme. We mainly looked at train data cross-validation like most of the participants, but also put some effort in looking at past BirdCLEF train / test differences individually for each year. We thought that this might help to understand what augmentations might be bridging the train (= xeno-canto) / test (= PAM soundscape recordings) gap. However, the results were quite inconclusive. Probably because soundscapes are quite different each year.  Whatever we tried lost correlation with the public LB once scores were high enough.\n\nOne thing that seemed to work great and didn’t in the end was the post processing based on max probability per soundscape.\n\nAnother one was the Aves model which improved public LB by about 0.01 each time it was used, but appeared to be detrimental on private LB.\n\nWe tried a number of augmentations, including those documented in previous competitions or in academic papers, but none seemed to really help.\n\nONNX did not bring much speedup (maybe 10%) and Openvino was much faster (almost 2x speedup compared to pytorch). However we noticed a drop of about 0.01 when using Openvino compared to ONNX, and decided not to use it. Maybe the speedup from openvino would have been compensated by the larger number of models we could use in the final ensemble?\n\nOur final model selection did not work that well in the end. We had individual models submitted with private scores above 0.70, but we did not include them in our ensemble given their relatively low public scores.\n\nThanks for reading !\n\nEdit, link to code (alphabetical order):\n- [CPMP's part](https://github.com/jfpuget/birdclef-2024)\n- [Dieter's part](https://github.com/ChristofHenkel/kaggle-birdclef24-3rd-place-solution-dieter)\n- [Theo's part](https://github.com/TheoViel/kaggle_birdclef2024)\n- Best selected submission: https://www.kaggle.com/cpmpml/birdclef-2024-inf-ens-08",
    "2869042": "Congratulations with gold first of all and thanks for your clear writeup! \n\nHow did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? For us, without ONNX it took around ~1h for a single fold, and with ONNX we observed a 2x speedup to around ~30 minutes per fold. Still around ~3.8x slower than your approach.\n\nHow did you use the different seeds? Did you just pick the seed with highest public LB for submission?\n\nHow did you exactly merge your pipelines, I am a bit confused by your ensembling section. Did you both submit level 0 and level 1, and which models were included?\n\nCool to see that you also tried the convolving postprocessing smoothing, we observed that surprisingly smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel, wonder if you also tried that.\n\nThanks!",
    "2869088": "Congratulation for the gold and thanks you for the write up !\n\nI have few questions:\n- (1) regarding the pseudo labels approach, does smaller batch was also giving you positive results? I tried 16 but did not get any luck unfortunately ... That was just providing more stable LB. \n- (2) how did you pick which checkpoint to select for inference?",
    "2869141": "Congrats for the gold and thankyou for the writeup ! \n\nI have some questions aswell :-\n\n1) Regarding loss ( Secondary labels were used to mask loss ) that means you were not giving any importance to secondary lables at all ?\n\n2) Regarding Full data training ( We trained most models on full data ) how did you made sure that model is not overfitting ? i did the same but model got overfitted i guess ( i added lots of augs / background noise etc , tried to make my data hard to learn on classes ) which were abundant and also oversampled my other classes with some techniques . \n\n3) Regarding PsudoLabels ( what was your threshold to determine psudo labels and there can be multiple classes in the audio clip or the time frame which you were using to infer so i am assuming you might be using 1 label which has the max conf ?  I did the same but i guess my model was overfitting.\n\n4) Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ? I upsampled the classes by adding some augs to the original 5 sec clip and also added unlaballed soundscapes to some other copies , ( i took only those soundscape clips where my model was getting very low confidence like < 0.05 ) but didnt got any luck probably becuase at the end i overfitted my model . \n\nAgain Thankyou for your time and effort .",
    "2869205": "Wow great work! I may have to go back to the board a bit. I had used efficientvit_b0 and it took 25 min to run on ONNX. 5 in 40 minutes is crazy, I couldnt get any of my models to be that fast! I would love to hear more about how that worked",
    "2869275": "> Congratulations with gold first of all and thanks for your clear writeup! \n\nThank you on behalf of the team.\n\n> How did you manage to submit 5 fold efficientnet_b0 224x224 within 40 minutes? \n\nEfficientvit_b0, not efficientnet_b0 :D\n\n> How did you use the different seeds? Did you just pick the seed with highest public LB for submission?\n\nCertainly not. This is a recipe for overfitting. We either used random seeds, or fixed seeds we change from time to time (team mates have differnent views here, but we all agree that tuning the seed is a disaster)\n\n> How did you exactly merge your pipelines,\n\nWe took the average of the logits of the 14 models.\n\n> smoothing the current 5 seconds with the mean of the whole 4 minutes was even better than the convolved kernel,\n\nWe tried and found that smoothing with the max over the soundscape is even better. Till we saw that it isn't with high score on private LB.",
    "2869277": "> Congratulation for the gold and thanks you for the write up !\n\nThank you on behalf of the team.\n\n> does smaller batch was also giving you positive results? \n\nWe tried as large batch as possible. 256 seemed to lower performance hence we stopped at 128. We then found that the more pseudo labels per batch the better. Maybe we could have gone further.\n\n> how did you pick which checkpoint to select for inference?\n\nThe last epoch one. We don't use early stopping.",
    "2869281": "> Congrats for the gold and thankyou for the writeup ! \n\nThank you on behalf of the team.\n\n> that means you were not giving any importance to secondary lables at all ?\n\nI think yes (not sure about the meaning of \"importance\"). We do find secondary labels to be important as we use them. \n\n> how did you made sure that model is not overfitting ?\n\nWe didn't really. We did train models using 5 fold CV, but the CV score wasn't correlated with LB score once we reached 0.65 or above. on LB. I welcome any hint on how to properly evaluate models in this competition besides submitting them.\n\n> Regarding PsudoLabels ( what was your threshold to determine psudo label\n\nWe use the predicted probability as label. We don't round probabilities to binary labels.\n\n> i am assuming you might be using 1 label which has the max conf ?\n\nWe don't. The competition is a mult label competition, not a multi class competition. \n\n> Regarding Upsampling ( did you just made multiple copies of the audio clips or added some transformation on those ) ?\n\nWe duplicated the record, then augmentations (time shift and mixup) are applied independently on each copy.",
    "2869282": "> Wow great work! \n\nThank you on behalf of the team.\n\n> I would love to hear more about how that worked\n\nWe didn't do anything fancy there. We used the model.half().float() trick that was shared, this shaved few minutes per 5 checkpoints. But we use various tricks like loading records once for all models, computing mel spectrograms once for all models, etc.",
    "2869285": "Congrats for the gold and 3rd position and thankyou for the writeup !!!",
    "2869751": "Ah great! Thanks for the clarification! \n\nAnd then I assume that you submitted all Level 1 models from each pipeline together, and then you used a varied ensemble of bigger models for Level 0 to get better pseudo labelling for Level 1? That is a very nice idea, we tried pseudo-labelling with the Google model on unlabelled soundscapes. That did not work out for us, unfortunately. Your approach on training multiple larger models on Xeno-Canto and using that for pseudolabelling seems to be a better approach :)",
    "2869827": "Actually the best second level model we selected was using pseudo labels from effcientvit_b0 only!\n\nThe other two pipelines used Aves and looked great on public LB, but were not great on private LB unfortunately.\n\nBest unselected pipelines were similar to those in the ensemble but without Aves.",
    "2869834": "Congratulations and great writeup. \n\nCould you clarify a little about masking loss for secondary labels? From what I can understand, if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss? Is this the correct interpretation?\n\nAlso, will you be opening up your inference sub for us to look at optimizations and stuff? 😐",
    "2869841": "Thankyou for the answers .\n\n>We don't. The competition is a mult label competition, not a multi class competition.\n\nWhat i meant by that was how you were deciding the primary label for your psudo labels in the window of your time frame and i guess it will be your max confidence class ? since you are setting your secondary classes as empty as mentioned in your write up.",
    "2869884": "> Congratulations and great writeup. \n\nThanks on behalf of the team.\n\n> if a bird X is present as secondary label, you essentially mask loss for this bird/index in BCEWithLogitsLoss?\n\nYes, we mask as you say.\n\nI documented this idea in [my 2020 Birdclef writeup](https://www.kaggle.com/competitions/birdsong-recognition/discussion/183219). I reproduce it here:\n\n>Primary labels were noisy, but secondary labels were even noisier. As a result we masked the loss for secondary labels as we didn't want to force the model to learn a presence or an absence when we don't know. We therefore defined a secondary mask that nullifies the BCE loss for secondary labels. For instance, assuming only 3 ebird_code b0, b1, and b2, and a clip with primary label b0 and secondary label b1, then these two target values are possible:\n> \n> [1, 0, 0]\n> \n> [1, 1, 0]\n> \n> The secondary mask is therefore:\n> \n> [1, 0, 1]\n\nI hope this is clearer now.\n\n> will you be opening up your inference sub \n\nI'll check with my team members.",
    "2870035": "We use the predicted probabilities as pseudo labels.  I am not sure what is not clear here.",
    "2870091": "my bad that i portray my question like that , so you were using all the predicted probabilities of your **psudolabels** for a given audio clip for exa:- 5 sec clip with 182 predicted class probabilites for training  and not the max class probability value alone and setting other class prob value as 0 .",
    "2870152": "Thankyou!",
    "2870268": "We were using all the probabilities. This is the only way one can use BCE loss.",
    "2870554": "Congratulations and nice writeup.",
    "2870600": "melspec transformation takes quite some time. But even we had 5 models we only need to do it once. So you can check how long multiple effnet b0 of you would take, if you only do melspec once",
    "2872135": "Congratulations 🎊",
    "2877150": "Thanks for the insight! Can you also share the code?",
    "2889069": "We shared our code, link in the post."
  },
  "source": "meta"
}