{
  "id": 326973,
  "title": "7th place solution",
  "url": "/competitions/birdclef-2022/writeups/iafoss-7th-place-solution",
  "author_name": "",
  "post_date": "2022-05-25T20:01:34.277Z",
  "votes": 45,
  "comment_count": 18,
  "views": 0,
  "content": "<h3>Summary</h3>\n<ul>\n<li><strong>Constant Q-transform</strong></li>\n<li>Multistep training with PL</li>\n<li>Pretraining on 2022+2021 data with noise from Rainforest</li>\n<li>Finetuning on 2022 data with weighted sampling (w~N^0.5)</li>\n<li>Sequence-based model</li>\n<li>Single model performance <strong>0.7971/0.7859</strong> private/public LB, while two models give <strong>0.7964/0.7936</strong></li>\n</ul>\n<h3>Introduction</h3>\n<p>Congratulations to all participants and thanks to the organizers for making this competition possible. It is my 3rd BirdCLEF competition, and the first gold medal I got in it. Though, this time I again followed the tradition of joining the competition close to the end… with the first sub 5-6 days before the deadline( Something should be changed in this life.</p>\n<p>I'll take this opportunity and describe some of the ideas I came up with while working on the challenge. I hope they may be interesting for other participants and organizers as well. My work, to a large extent, is based on my <a href=\"http://ceur-ws.org/Vol-2936/paper-141.pdf\" target=\"_blank\">BirdCLEF 2021 paper</a>. Also, some details may be found in my <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243343\" target=\"_blank\">BirdCLEF 2021</a> and <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183258\" target=\"_blank\">2020 Cornell Birdcall</a> writeups.</p>\n<h3>Data</h3>\n<p>In this year's competition there were two main challenges to address:  <strong>(1) considerable domain mismatch with very noisy test data</strong> and (2) <strong>long-tail class distribution with a few examples of rare classes</strong>. </p>\n<p>To address the first challenge, I used a traditional <strong>pseudo label (PL) 2-step training</strong>: in the first pass the model learns where the signal from birds is located, and on the second pass a heavy noise is applied. With segmentation labels the model knows where to look for the true signal under a severe noise. As a noise source <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/data\" target=\"_blank\">rainforest data</a> is used in addition to pink noise and signal weakening.</p>\n<p>To address the second challenge and make the model robust to different bird calls I pretrained it on a large <strong>2021+2022 BirdCLEF dataset</strong> (the size of 2022 data is ~ 5 times smaller). During fine-tuning I kept the backbone frozen and trained only the head on 2022 data. I assigned all unscored classes to \"unknown\" label and performed weighted sampling with probability conditioned on the number of samples of each class as w~N^0.5.</p>\n<p>One of the novel things I used in my solution, in comparison to traditional audio models, is <strong>Constant Q-transform</strong> instead of mel-scale spectrograms and even pcen. The quality of the produced spectrograms is much better, as shown in the figure below: lower impact of noise, high level of details in the high-frequency domain, high contrast, and easy mapping for sequence models (without the need to care about padding added in FFT). However, I didn't have time for a quantitative comparison of the model performance. CQT1992v2 is a very effective Pytorch (GPU) implementation of CQT from nnAudio.Spectrogram library and can be easily integrated into a model (make sure to disable gradients for CQT).  During training I used log transformation of transform output: <code>x = (64*x + 1).log()</code>.<br>\n<img src=\"https://i.ibb.co/Trv7Qkb/spectrogram.png\" alt=\"\"></p>\n<pre><code>self.qtransform = CQT1992v2(sr=32000, fmin=256, n_bins=160, \nhop_length=250, output_format='Magnitude', norm=1,\nwindow='tukey',bins_per_octave=27,)\n</code></pre>\n<h3>Model</h3>\n<p>I used <a href=\"https://github.com/facebookresearch/semi-supervised-ImageNet1K-models\" target=\"_blank\">semisupervised pretrained</a> ResNeXt50 and MiT-B2 transformer, <a href=\"https://github.com/NVlabs/SegFormer\" target=\"_blank\">Segformer </a> backbone. The last conv layer is followed by a convolution, collapsing the feature map into a sequence of vectors vector (nemb=1024), with a Transformer layer, and a head, producing sequence and clip outputs using the attention mechanism (somewhat similar to SED models):</p>\n<pre><code>class AttHead(nn.Module):\n    def __init__(self, n_in, n_out):\n        super().__init__()\n        self.attn = nn.Conv1d(n_in,n_out,1)\n        self.cla = nn.Conv1d(n_in,n_out,1)\n\n    def forward(self, x):\n        if len(x.shape) == 4: x = x.flatten(1,2)\n        attn = self.attn(x)\n        cla = self.cla(x)\n        x = (torch.softmax(attn,-1)*cla).sum(-1)\n        return x, attn, cla\n</code></pre>\n<p>Initially, I believed that using MiT transformer backbone with a large receptive field is important: it is able to look into the entire frequency domain at one (instead of a small window for conv net) and compare bird calls in the clip across the time. But it appeared to be not really true, and my favorite ResNeXt50 performed nearly the same as MiT B2… both CV and LB. Though, in both cases using a transformer head quite helps based on my initial tests.  </p>\n<h3>Training</h3>\n<p>I used multistep training: (1) pretrain the model on 2022+2021 data and generate PL; (2) train the model on 2022+2021 data with a high level of noise added and additional segmentation loss based on PL (global labels for selected chunks are also adjusted based on PL in case if there is no a bird call); (3) take the model from step 2 and freeze the backbone, finetune on 2022 data with weighted sampling considering 21+1 classes, and generate new PL; (4)  the same as 3 but noise and segmentation losses are added. </p>\n<p>At all stages MixUp augmentation (with label max and alpha of 2, i.e. the probabilities are close to 0.5) is applied to waves. Each stage is started with using 5s chunks (32 epochs) and is finished with a few epochs with 10-15s chunks. I use Focal loss with gamma=1 (with a bug fix to accept smooth labels).</p>\n<h3>Inference</h3>\n<p>The inference is performed on entire audio files (while training is done on 5s clips followed with fine-tuning on 10-15s clips). I use sequence level output, illustrated in the image below for the test audio clip. It is split into 5s chinks, and the predictions are selected if the maximum output of the particular class reaches the selected threshold. If one compares the plot with similar plots from my previous reports, the model performance improvement is quite clear. Now the model is able to clearly distinguish bird calls even in very noisy audio. For postprocessing the model output I was using the following: <code>x = (torch.softmax(3*attn,-2)*torch.sigmoid(3*cla))</code>. Pay attention to the dimension, which is the class rather than the sequence dimension: if the model is paying more attention to a specific class at a particular moment, the signal should be enhanced accordingly. Also, I'm using temperature rescaling.<br>\n<img src=\"https://i.ibb.co/sKXxh8J/pred-test.png\" alt=\"\"></p>\n<p>The final submission is composed of 2 models (MiT and ResNeXt50) and is scored as 0.7964/0.7936 at public/private LB. The best single model is scored as 0.7971/0.7859. In this competition, the main focus is bird classification, and no credit is given for call/nocall separation, in contrast to 2020 and 2021 competitions. So the best scores could be achieved at the selection of very low thresholds. If more reasonable thresholds, capable of nocall separation, are used, the performance drops by ~0.05.</p>\n<h3>Things  I didn't have time to make work</h3>\n<ul>\n<li>Incorporation of ArcFace loss into attention pooling head (dealing with sequence and global labels).</li>\n<li>Performing metric learning and clustering to address the issue with a few training examples for rare classes</li>\n</ul>\n<p>I think those two things might be the key to getting to the top of the LB.</p>",
  "messages": [
    {
      "id": "1800628",
      "postDate": "05/25/2022 05:22:06",
      "content": "<h3>Summary</h3>\n<ul>\n<li><strong>Constant Q-transform</strong></li>\n<li>Multistep training with PL</li>\n<li>Pretraining on 2022+2021 data with noise from Rainforest</li>\n<li>Finetuning on 2022 data with weighted sampling (w~N^0.5)</li>\n<li>Sequence-based model</li>\n<li>Single model performance <strong>0.7971/0.7859</strong> private/public LB, while two models give <strong>0.7964/0.7936</strong></li>\n</ul>\n<h3>Introduction</h3>\n<p>Congratulations to all participants and thanks to the organizers for making this competition possible. It is my 3rd BirdCLEF competition, and the first gold medal I got in it. Though, this time I again followed the tradition of joining the competition close to the end… with the first sub 5-6 days before the deadline( Something should be changed in this life.</p>\n<p>I'll take this opportunity and describe some of the ideas I came up with while working on the challenge. I hope they may be interesting for other participants and organizers as well. My work, to a large extent, is based on my <a href=\"http://ceur-ws.org/Vol-2936/paper-141.pdf\" target=\"_blank\">BirdCLEF 2021 paper</a>. Also, some details may be found in my <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243343\" target=\"_blank\">BirdCLEF 2021</a> and <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183258\" target=\"_blank\">2020 Cornell Birdcall</a> writeups.</p>\n<h3>Data</h3>\n<p>In this year's competition there were two main challenges to address:  <strong>(1) considerable domain mismatch with very noisy test data</strong> and (2) <strong>long-tail class distribution with a few examples of rare classes</strong>. </p>\n<p>To address the first challenge, I used a traditional <strong>pseudo label (PL) 2-step training</strong>: in the first pass the model learns where the signal from birds is located, and on the second pass a heavy noise is applied. With segmentation labels the model knows where to look for the true signal under a severe noise. As a noise source <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/data\" target=\"_blank\">rainforest data</a> is used in addition to pink noise and signal weakening.</p>\n<p>To address the second challenge and make the model robust to different bird calls I pretrained it on a large <strong>2021+2022 BirdCLEF dataset</strong> (the size of 2022 data is ~ 5 times smaller). During fine-tuning I kept the backbone frozen and trained only the head on 2022 data. I assigned all unscored classes to \"unknown\" label and performed weighted sampling with probability conditioned on the number of samples of each class as w~N^0.5.</p>\n<p>One of the novel things I used in my solution, in comparison to traditional audio models, is <strong>Constant Q-transform</strong> instead of mel-scale spectrograms and even pcen. The quality of the produced spectrograms is much better, as shown in the figure below: lower impact of noise, high level of details in the high-frequency domain, high contrast, and easy mapping for sequence models (without the need to care about padding added in FFT). However, I didn't have time for a quantitative comparison of the model performance. CQT1992v2 is a very effective Pytorch (GPU) implementation of CQT from nnAudio.Spectrogram library and can be easily integrated into a model (make sure to disable gradients for CQT).  During training I used log transformation of transform output: <code>x = (64*x + 1).log()</code>.<br>\n<img src=\"https://i.ibb.co/Trv7Qkb/spectrogram.png\" alt=\"\"></p>\n<pre><code>self.qtransform = CQT1992v2(sr=32000, fmin=256, n_bins=160, \nhop_length=250, output_format='Magnitude', norm=1,\nwindow='tukey',bins_per_octave=27,)\n</code></pre>\n<h3>Model</h3>\n<p>I used <a href=\"https://github.com/facebookresearch/semi-supervised-ImageNet1K-models\" target=\"_blank\">semisupervised pretrained</a> ResNeXt50 and MiT-B2 transformer, <a href=\"https://github.com/NVlabs/SegFormer\" target=\"_blank\">Segformer </a> backbone. The last conv layer is followed by a convolution, collapsing the feature map into a sequence of vectors vector (nemb=1024), with a Transformer layer, and a head, producing sequence and clip outputs using the attention mechanism (somewhat similar to SED models):</p>\n<pre><code>class AttHead(nn.Module):\n    def __init__(self, n_in, n_out):\n        super().__init__()\n        self.attn = nn.Conv1d(n_in,n_out,1)\n        self.cla = nn.Conv1d(n_in,n_out,1)\n\n    def forward(self, x):\n        if len(x.shape) == 4: x = x.flatten(1,2)\n        attn = self.attn(x)\n        cla = self.cla(x)\n        x = (torch.softmax(attn,-1)*cla).sum(-1)\n        return x, attn, cla\n</code></pre>\n<p>Initially, I believed that using MiT transformer backbone with a large receptive field is important: it is able to look into the entire frequency domain at one (instead of a small window for conv net) and compare bird calls in the clip across the time. But it appeared to be not really true, and my favorite ResNeXt50 performed nearly the same as MiT B2… both CV and LB. Though, in both cases using a transformer head quite helps based on my initial tests.  </p>\n<h3>Training</h3>\n<p>I used multistep training: (1) pretrain the model on 2022+2021 data and generate PL; (2) train the model on 2022+2021 data with a high level of noise added and additional segmentation loss based on PL (global labels for selected chunks are also adjusted based on PL in case if there is no a bird call); (3) take the model from step 2 and freeze the backbone, finetune on 2022 data with weighted sampling considering 21+1 classes, and generate new PL; (4)  the same as 3 but noise and segmentation losses are added. </p>\n<p>At all stages MixUp augmentation (with label max and alpha of 2, i.e. the probabilities are close to 0.5) is applied to waves. Each stage is started with using 5s chunks (32 epochs) and is finished with a few epochs with 10-15s chunks. I use Focal loss with gamma=1 (with a bug fix to accept smooth labels).</p>\n<h3>Inference</h3>\n<p>The inference is performed on entire audio files (while training is done on 5s clips followed with fine-tuning on 10-15s clips). I use sequence level output, illustrated in the image below for the test audio clip. It is split into 5s chinks, and the predictions are selected if the maximum output of the particular class reaches the selected threshold. If one compares the plot with similar plots from my previous reports, the model performance improvement is quite clear. Now the model is able to clearly distinguish bird calls even in very noisy audio. For postprocessing the model output I was using the following: <code>x = (torch.softmax(3*attn,-2)*torch.sigmoid(3*cla))</code>. Pay attention to the dimension, which is the class rather than the sequence dimension: if the model is paying more attention to a specific class at a particular moment, the signal should be enhanced accordingly. Also, I'm using temperature rescaling.<br>\n<img src=\"https://i.ibb.co/sKXxh8J/pred-test.png\" alt=\"\"></p>\n<p>The final submission is composed of 2 models (MiT and ResNeXt50) and is scored as 0.7964/0.7936 at public/private LB. The best single model is scored as 0.7971/0.7859. In this competition, the main focus is bird classification, and no credit is given for call/nocall separation, in contrast to 2020 and 2021 competitions. So the best scores could be achieved at the selection of very low thresholds. If more reasonable thresholds, capable of nocall separation, are used, the performance drops by ~0.05.</p>\n<h3>Things  I didn't have time to make work</h3>\n<ul>\n<li>Incorporation of ArcFace loss into attention pooling head (dealing with sequence and global labels).</li>\n<li>Performing metric learning and clustering to address the issue with a few training examples for rare classes</li>\n</ul>\n<p>I think those two things might be the key to getting to the top of the LB.</p>",
      "rawMarkdown": "### Summary\n- **Constant Q-transform**\n- Multistep training with PL\n- Pretraining on 2022+2021 data with noise from Rainforest\n- Finetuning on 2022 data with weighted sampling (w~N^0.5)\n- Sequence-based model\n- Single model performance **0.7971/0.7859** private/public LB, while two models give **0.7964/0.7936**\n\n### Introduction\nCongratulations to all participants and thanks to the organizers for making this competition possible. It is my 3rd BirdCLEF competition, and the first gold medal I got in it. Though, this time I again followed the tradition of joining the competition close to the end... with the first sub 5-6 days before the deadline( Something should be changed in this life.\n\nI'll take this opportunity and describe some of the ideas I came up with while working on the challenge. I hope they may be interesting for other participants and organizers as well. My work, to a large extent, is based on my [BirdCLEF 2021 paper](http://ceur-ws.org/Vol-2936/paper-141.pdf). Also, some details may be found in my [BirdCLEF 2021](https://www.kaggle.com/competitions/birdclef-2021/discussion/243343) and [2020 Cornell Birdcall](https://www.kaggle.com/c/birdsong-recognition/discussion/183258) writeups.\n\n### Data\nIn this year's competition there were two main challenges to address:  **(1) considerable domain mismatch with very noisy test data** and (2) **long-tail class distribution with a few examples of rare classes**. \n\nTo address the first challenge, I used a traditional **pseudo label (PL) 2-step training**: in the first pass the model learns where the signal from birds is located, and on the second pass a heavy noise is applied. With segmentation labels the model knows where to look for the true signal under a severe noise. As a noise source [rainforest data](https://www.kaggle.com/c/rfcx-species-audio-detection/data) is used in addition to pink noise and signal weakening.\n\nTo address the second challenge and make the model robust to different bird calls I pretrained it on a large **2021+2022 BirdCLEF dataset** (the size of 2022 data is ~ 5 times smaller). During fine-tuning I kept the backbone frozen and trained only the head on 2022 data. I assigned all unscored classes to \"unknown\" label and performed weighted sampling with probability conditioned on the number of samples of each class as w~N^0.5.\n\nOne of the novel things I used in my solution, in comparison to traditional audio models, is **Constant Q-transform** instead of mel-scale spectrograms and even pcen. The quality of the produced spectrograms is much better, as shown in the figure below: lower impact of noise, high level of details in the high-frequency domain, high contrast, and easy mapping for sequence models (without the need to care about padding added in FFT). However, I didn't have time for a quantitative comparison of the model performance. CQT1992v2 is a very effective Pytorch (GPU) implementation of CQT from nnAudio.Spectrogram library and can be easily integrated into a model (make sure to disable gradients for CQT).  During training I used log transformation of transform output: `x = (64*x + 1).log()`.\n![](https://i.ibb.co/Trv7Qkb/spectrogram.png)\n\n```\nself.qtransform = CQT1992v2(sr=32000, fmin=256, n_bins=160, \nhop_length=250, output_format='Magnitude', norm=1,\nwindow='tukey',bins_per_octave=27,)\n```\n\n### Model\nI used [semisupervised pretrained](https://github.com/facebookresearch/semi-supervised-ImageNet1K-models) ResNeXt50 and MiT-B2 transformer, [Segformer ](https://github.com/NVlabs/SegFormer) backbone. The last conv layer is followed by a convolution, collapsing the feature map into a sequence of vectors vector (nemb=1024), with a Transformer layer, and a head, producing sequence and clip outputs using the attention mechanism (somewhat similar to SED models):\n```\nclass AttHead(nn.Module):\n    def __init__(self, n_in, n_out):\n        super().__init__()\n        self.attn = nn.Conv1d(n_in,n_out,1)\n        self.cla = nn.Conv1d(n_in,n_out,1)\n        \n    def forward(self, x):\n        if len(x.shape) == 4: x = x.flatten(1,2)\n        attn = self.attn(x)\n        cla = self.cla(x)\n        x = (torch.softmax(attn,-1)*cla).sum(-1)\n        return x, attn, cla\n```\nInitially, I believed that using MiT transformer backbone with a large receptive field is important: it is able to look into the entire frequency domain at one (instead of a small window for conv net) and compare bird calls in the clip across the time. But it appeared to be not really true, and my favorite ResNeXt50 performed nearly the same as MiT B2... both CV and LB. Though, in both cases using a transformer head quite helps based on my initial tests.  \n\n### Training\nI used multistep training: (1) pretrain the model on 2022+2021 data and generate PL; (2) train the model on 2022+2021 data with a high level of noise added and additional segmentation loss based on PL (global labels for selected chunks are also adjusted based on PL in case if there is no a bird call); (3) take the model from step 2 and freeze the backbone, finetune on 2022 data with weighted sampling considering 21+1 classes, and generate new PL; (4)  the same as 3 but noise and segmentation losses are added. \n\nAt all stages MixUp augmentation (with label max and alpha of 2, i.e. the probabilities are close to 0.5) is applied to waves. Each stage is started with using 5s chunks (32 epochs) and is finished with a few epochs with 10-15s chunks. I use Focal loss with gamma=1 (with a bug fix to accept smooth labels).\n\n### Inference\nThe inference is performed on entire audio files (while training is done on 5s clips followed with fine-tuning on 10-15s clips). I use sequence level output, illustrated in the image below for the test audio clip. It is split into 5s chinks, and the predictions are selected if the maximum output of the particular class reaches the selected threshold. If one compares the plot with similar plots from my previous reports, the model performance improvement is quite clear. Now the model is able to clearly distinguish bird calls even in very noisy audio. For postprocessing the model output I was using the following: `x = (torch.softmax(3*attn,-2)*torch.sigmoid(3*cla))`. Pay attention to the dimension, which is the class rather than the sequence dimension: if the model is paying more attention to a specific class at a particular moment, the signal should be enhanced accordingly. Also, I'm using temperature rescaling.\n![](https://i.ibb.co/sKXxh8J/pred-test.png)\n\nThe final submission is composed of 2 models (MiT and ResNeXt50) and is scored as 0.7964/0.7936 at public/private LB. The best single model is scored as 0.7971/0.7859. In this competition, the main focus is bird classification, and no credit is given for call/nocall separation, in contrast to 2020 and 2021 competitions. So the best scores could be achieved at the selection of very low thresholds. If more reasonable thresholds, capable of nocall separation, are used, the performance drops by ~0.05.\n\n### Things ~~didn't work~~ I didn't have time to make work\n- Incorporation of ArcFace loss into attention pooling head (dealing with sequence and global labels).\n- Performing metric learning and clustering to address the issue with a few training examples for rare classes\n\nI think those two things might be the key to getting to the top of the LB.",
      "votes": null
    },
    {
      "id": "1800639",
      "postDate": "05/25/2022 05:41:28",
      "content": "<p>Congrats on strong finish and solo gold medal <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> </p>",
      "rawMarkdown": "Congrats on strong finish and solo gold medal @iafoss",
      "votes": null
    },
    {
      "id": "1800648",
      "postDate": "05/25/2022 05:48:44",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
      "rawMarkdown": "Thank you @duykhanh99",
      "votes": null
    },
    {
      "id": "1800662",
      "postDate": "05/25/2022 06:06:23",
      "content": "<p>Congratulations for the solo gold <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , I'm wondering what is your validation strategy like, and how do you determine the threshold? As we don't have train soundscapes files like last year, I was struggling at CV/LB mismatching all the time. </p>",
      "rawMarkdown": "Congratulations for the solo gold @iafoss , I'm wondering what is your validation strategy like, and how do you determine the threshold? As we don't have train soundscapes files like last year, I was struggling at CV/LB mismatching all the time.",
      "votes": null
    },
    {
      "id": "1800704",
      "postDate": "05/25/2022 06:38:03",
      "content": "<p>Thanks. <strong>In this competition CV strategy for threshold identification doesn't exist</strong>. Let me explain why. <br>\n(1) In BirdCLEF competitions there is a tradition to have a huge domain gap between train and test data. Test data is of poor quality with large background noise and 5s chunk annotation. At the same time, train data is often of exceptional quality with a sequence-level annotation (chunks selected during evaluation may not have bird calls). One may try to emulate test data by adding a high level of noise to val and using PL to identify nocall parts to ensure that a specific 5s chunk includes a bird call. Also if secondary birds are percent, it becomes more sophisticated… Though this procedure should be calibrated with LB to set the level of noise at least. Such CV strategy could work in BirdCLEF 2020 competition, but not here because of the following.<br>\n(2) <strong>The host doesn't score all chunks and all species</strong> but does cherry-picking, i.e. ignore nocall cases and considers only a few options for each chunk. It is like if one heard a birdcall and is asking if it is A or B species. Since <strong>this procedure is unknown</strong>, the threshold should be calibrated based on LB( Though it is a kind of cheating, and the host is getting an overestimated model performance because we \"overfit\" the procedure used to select scored species both in public and private test set.</p>\n<p>During training I was computing some F1 score, but to assess the model performance I mostly paid attention to the model output at the test data, like one in my post, because of the issue (1): we do not need a model that performs well without noise, we need a model that works on test data. So listening to the test audio and visual inspection of the model output on it is not a bad thing here. After participating in competitions like BirdCLEF or PANDA, where validation is quite illusional, one gets instincts on model assessment and selection… <strong>The key is following the methodology, i.e. focusing on the things that make sense.</strong> Though one should understand when LB or CV or both should be ignored, it's another key.</p>",
      "rawMarkdown": "Thanks. **In this competition CV strategy for threshold identification doesn't exist**. Let me explain why. \n(1) In BirdCLEF competitions there is a tradition to have a huge domain gap between train and test data. Test data is of poor quality with large background noise and 5s chunk annotation. At the same time, train data is often of exceptional quality with a sequence-level annotation (chunks selected during evaluation may not have bird calls). One may try to emulate test data by adding a high level of noise to val and using PL to identify nocall parts to ensure that a specific 5s chunk includes a bird call. Also if secondary birds are percent, it becomes more sophisticated... Though this procedure should be calibrated with LB to set the level of noise at least. Such CV strategy could work in BirdCLEF 2020 competition, but not here because of the following.\n(2) **The host doesn't score all chunks and all species** but does cherry-picking, i.e. ignore nocall cases and considers only a few options for each chunk. It is like if one heard a birdcall and is asking if it is A or B species. Since **this procedure is unknown**, the threshold should be calibrated based on LB( Though it is a kind of cheating, and the host is getting an overestimated model performance because we \"overfit\" the procedure used to select scored species both in public and private test set.\n\nDuring training I was computing some F1 score, but to assess the model performance I mostly paid attention to the model output at the test data, like one in my post, because of the issue (1): we do not need a model that performs well without noise, we need a model that works on test data. So listening to the test audio and visual inspection of the model output on it is not a bad thing here. After participating in competitions like BirdCLEF or PANDA, where validation is quite illusional, one gets instincts on model assessment and selection... **The key is following the methodology, i.e. focusing on the things that make sense.** Though one should understand when LB or CV or both should be ignored, it's another key.",
      "votes": null
    },
    {
      "id": "1800812",
      "postDate": "05/25/2022 08:28:30",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on your amazing score in a short time of competing and thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats @iafoss on your amazing score in a short time of competing and thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "1801004",
      "postDate": "05/25/2022 10:57:26",
      "content": "<p>Congrats for your solo gold and by achieving it not using any leak. Thanks for sharing your approach.</p>",
      "rawMarkdown": "Congrats for your solo gold and by achieving it not using any leak. Thanks for sharing your approach.",
      "votes": null
    },
    {
      "id": "1801143",
      "postDate": "05/25/2022 13:52:49",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> </p>",
      "rawMarkdown": "Thank you so much @titericz",
      "votes": null
    },
    {
      "id": "1801173",
      "postDate": "05/25/2022 14:07:19",
      "content": "<p>You are very welcome, <a href=\"https://www.kaggle.com/tjamali\" target=\"_blank\">@tjamali</a> </p>",
      "rawMarkdown": "You are very welcome, @tjamali",
      "votes": null
    },
    {
      "id": "1801257",
      "postDate": "05/25/2022 15:01:47",
      "content": "<p>Congratz on getting good score bro and Thanks for sharing</p>",
      "rawMarkdown": "Congratz on getting good score bro and Thanks for sharing",
      "votes": null
    },
    {
      "id": "1801484",
      "postDate": "05/25/2022 19:44:37",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/muhammadabbasshareef\" target=\"_blank\">@muhammadabbasshareef</a> </p>",
      "rawMarkdown": "Thanks @muhammadabbasshareef",
      "votes": null
    },
    {
      "id": "1802008",
      "postDate": "05/26/2022 10:49:03",
      "content": "<p>Congratulations on your solo gold medal and your first gold medal in the BirdCLEF competition :)</p>",
      "rawMarkdown": "Congratulations on your solo gold medal and your first gold medal in the BirdCLEF competition :)",
      "votes": null
    },
    {
      "id": "1802258",
      "postDate": "05/26/2022 15:46:04",
      "content": "<p>Thanks so much, <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>. I'm happy that I got it this time. I quite remember our hard work at the end of BirdCLEF 2021 and how close we were to the gold… I expected that you would be interested in BirdCLEF 2022.</p>",
      "rawMarkdown": "Thanks so much, @naoism. I'm happy that I got it this time. I quite remember our hard work at the end of BirdCLEF 2021 and how close we were to the gold... I expected that you would be interested in BirdCLEF 2022.",
      "votes": null
    },
    {
      "id": "1802607",
      "postDate": "05/27/2022 00:19:37",
      "content": "<p>BirdCLEF2021 is quite an impressive competition for me too. To work with you was quite a fun memory. (Also I remember I was really frustrated not to get gold medal at the time…) Actually I was planning to participate and did also EDA, but I didn't because we didn't have data for validation like BirdCLEF2021…</p>",
      "rawMarkdown": "BirdCLEF2021 is quite an impressive competition for me too. To work with you was quite a fun memory. (Also I remember I was really frustrated not to get gold medal at the time...) Actually I was planning to participate and did also EDA, but I didn't because we didn't have data for validation like BirdCLEF2021...",
      "votes": null
    },
    {
      "id": "1804447",
      "postDate": "05/29/2022 02:10:06",
      "content": "<p>Congratulations for solo gold in only 25 submissions. It's the least submissions in the gold winners.</p>\n<p>If you had enough time, how would you have implemented ArcFace loss into attention pooling head?<br>\nI think this it's not straight-forward because there are two differences in the original ArcFace condition:</p>\n<ol>\n<li>it has multiple labels</li>\n<li>it has time dimension (or, did you planned to apply ArcFace on pooled vector output (bs, d, 1, 1)?)</li>\n</ol>",
      "rawMarkdown": "Congratulations for solo gold in only 25 submissions. It's the least submissions in the gold winners.\n\nIf you had enough time, how would you have implemented ArcFace loss into attention pooling head?\nI think this it's not straight-forward because there are two differences in the original ArcFace condition:\n1. it has multiple labels\n2. it has time dimension (or, did you planned to apply ArcFace on pooled vector output (bs, d, 1, 1)?)",
      "votes": null
    },
    {
      "id": "1804509",
      "postDate": "05/29/2022 05:00:54",
      "content": "<p>If the provided labels were stronger, i.e. include time tags for the beginning and end of the bird call, everything would be super easy: just do a sound segmentation model assuming that at each moment of time only a single bird may be present at the most.</p>\n<p>The weakness of labels screws this simple picture. Ideally, it would be preferable to set an approach dealing with the provided annotation… One thing, as you mentioned, may be using an attention mechanism to pool a single vector out of generated sequence of vectors. The issue here, though, is how to incorporate a multi-label paradigm into it. I'm not sure at the moment… The nature of AF loss is the maximization of class separation, and expecting two labels as a prediction doesn't give a reasonable convergence(</p>\n<p>Another way is computing the loss in the same way as for strong labels, but using multi-head <strong>attention to weight the loss across the generated sequence</strong>. It can be seen as the following. Let's assume that there are two birds present in the record, and AF is low within these S1 and S2 segments if corresponding labels are considered as GT (S1-L1, S2-L2), and is high outside. The attention should perform a soft selection of segments S1 and S2 as considered concepts and assign them to labels L1 and L2. It will minimize the total loss. This concept is somewhat similar to Maskformer, and you can use it as a starting point for thinking, but it is not exactly the same.</p>",
      "rawMarkdown": "If the provided labels were stronger, i.e. include time tags for the beginning and end of the bird call, everything would be super easy: just do a sound segmentation model assuming that at each moment of time only a single bird may be present at the most.\n\nThe weakness of labels screws this simple picture. Ideally, it would be preferable to set an approach dealing with the provided annotation... One thing, as you mentioned, may be using an attention mechanism to pool a single vector out of generated sequence of vectors. The issue here, though, is how to incorporate a multi-label paradigm into it. I'm not sure at the moment... The nature of AF loss is the maximization of class separation, and expecting two labels as a prediction doesn't give a reasonable convergence(\n\nAnother way is computing the loss in the same way as for strong labels, but using multi-head **attention to weight the loss across the generated sequence**. It can be seen as the following. Let's assume that there are two birds present in the record, and AF is low within these S1 and S2 segments if corresponding labels are considered as GT (S1-L1, S2-L2), and is high outside. The attention should perform a soft selection of segments S1 and S2 as considered concepts and assign them to labels L1 and L2. It will minimize the total loss. This concept is somewhat similar to Maskformer, and you can use it as a starting point for thinking, but it is not exactly the same.",
      "votes": null
    },
    {
      "id": "1804532",
      "postDate": "05/29/2022 05:57:16",
      "content": "<p>Thanks for the detailed reply.</p>\n<p>Yes, relaxing condition will make it easier to adapt ArcFace to this problem.</p>\n<ul>\n<li>assume at most one bird call is appeared in a segment</li>\n<li>strong label (segment-wise label) is provided</li>\n</ul>\n<p>So if I understand correctly, your idea is using attention score for each class as a strong label?</p>",
      "rawMarkdown": "Thanks for the detailed reply.\n\nYes, relaxing condition will make it easier to adapt ArcFace to this problem.\n\n* assume at most one bird call is appeared in a segment\n* strong label (segment-wise label) is provided\n\nSo if I understand correctly, your idea is using attention score for each class as a strong label?",
      "votes": null
    },
    {
      "id": "1805322",
      "postDate": "05/30/2022 04:21:18",
      "content": "<p>The thing I thought of is a way to reweight the loss along the sequence, i.e. the loss is computed for parts where the corresponding birds are expected to be, while other parts are having a small contribution…</p>",
      "rawMarkdown": "The thing I thought of is a way to reweight the loss along the sequence, i.e. the loss is computed for parts where the corresponding birds are expected to be, while other parts are having a small contribution...",
      "votes": null
    },
    {
      "id": "1805397",
      "postDate": "05/30/2022 06:05:14",
      "content": "<p>I see. Thanks for confirmation.</p>",
      "rawMarkdown": "I see. Thanks for confirmation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1800639,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "05/25/2022 05:41:28",
      "content": "<p>Congrats on strong finish and solo gold medal <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1800648,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/25/2022 05:48:44",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1800662,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "05/25/2022 06:06:23",
      "content": "<p>Congratulations for the solo gold <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , I'm wondering what is your validation strategy like, and how do you determine the threshold? As we don't have train soundscapes files like last year, I was struggling at CV/LB mismatching all the time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1800704,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/25/2022 06:38:03",
          "content": "<p>Thanks. <strong>In this competition CV strategy for threshold identification doesn't exist</strong>. Let me explain why. <br>\n(1) In BirdCLEF competitions there is a tradition to have a huge domain gap between train and test data. Test data is of poor quality with large background noise and 5s chunk annotation. At the same time, train data is often of exceptional quality with a sequence-level annotation (chunks selected during evaluation may not have bird calls). One may try to emulate test data by adding a high level of noise to val and using PL to identify nocall parts to ensure that a specific 5s chunk includes a bird call. Also if secondary birds are percent, it becomes more sophisticated… Though this procedure should be calibrated with LB to set the level of noise at least. Such CV strategy could work in BirdCLEF 2020 competition, but not here because of the following.<br>\n(2) <strong>The host doesn't score all chunks and all species</strong> but does cherry-picking, i.e. ignore nocall cases and considers only a few options for each chunk. It is like if one heard a birdcall and is asking if it is A or B species. Since <strong>this procedure is unknown</strong>, the threshold should be calibrated based on LB( Though it is a kind of cheating, and the host is getting an overestimated model performance because we \"overfit\" the procedure used to select scored species both in public and private test set.</p>\n<p>During training I was computing some F1 score, but to assess the model performance I mostly paid attention to the model output at the test data, like one in my post, because of the issue (1): we do not need a model that performs well without noise, we need a model that works on test data. So listening to the test audio and visual inspection of the model output on it is not a bad thing here. After participating in competitions like BirdCLEF or PANDA, where validation is quite illusional, one gets instincts on model assessment and selection… <strong>The key is following the methodology, i.e. focusing on the things that make sense.</strong> Though one should understand when LB or CV or both should be ignored, it's another key.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1800812,
      "author_name": "tjamali",
      "author_url": "",
      "post_date": "05/25/2022 08:28:30",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on your amazing score in a short time of competing and thanks for sharing your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801173,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/25/2022 14:07:19",
          "content": "<p>You are very welcome, <a href=\"https://www.kaggle.com/tjamali\" target=\"_blank\">@tjamali</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801004,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "05/25/2022 10:57:26",
      "content": "<p>Congrats for your solo gold and by achieving it not using any leak. Thanks for sharing your approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801143,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/25/2022 13:52:49",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801257,
      "author_name": "muhammadabbasshareef",
      "author_url": "",
      "post_date": "05/25/2022 15:01:47",
      "content": "<p>Congratz on getting good score bro and Thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801484,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/25/2022 19:44:37",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/muhammadabbasshareef\" target=\"_blank\">@muhammadabbasshareef</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1802008,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "05/26/2022 10:49:03",
      "content": "<p>Congratulations on your solo gold medal and your first gold medal in the BirdCLEF competition :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802258,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/26/2022 15:46:04",
          "content": "<p>Thanks so much, <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>. I'm happy that I got it this time. I quite remember our hard work at the end of BirdCLEF 2021 and how close we were to the gold… I expected that you would be interested in BirdCLEF 2022.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1802607,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "05/27/2022 00:19:37",
          "content": "<p>BirdCLEF2021 is quite an impressive competition for me too. To work with you was quite a fun memory. (Also I remember I was really frustrated not to get gold medal at the time…) Actually I was planning to participate and did also EDA, but I didn't because we didn't have data for validation like BirdCLEF2021…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1804447,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/29/2022 02:10:06",
      "content": "<p>Congratulations for solo gold in only 25 submissions. It's the least submissions in the gold winners.</p>\n<p>If you had enough time, how would you have implemented ArcFace loss into attention pooling head?<br>\nI think this it's not straight-forward because there are two differences in the original ArcFace condition:</p>\n<ol>\n<li>it has multiple labels</li>\n<li>it has time dimension (or, did you planned to apply ArcFace on pooled vector output (bs, d, 1, 1)?)</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1804509,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/29/2022 05:00:54",
          "content": "<p>If the provided labels were stronger, i.e. include time tags for the beginning and end of the bird call, everything would be super easy: just do a sound segmentation model assuming that at each moment of time only a single bird may be present at the most.</p>\n<p>The weakness of labels screws this simple picture. Ideally, it would be preferable to set an approach dealing with the provided annotation… One thing, as you mentioned, may be using an attention mechanism to pool a single vector out of generated sequence of vectors. The issue here, though, is how to incorporate a multi-label paradigm into it. I'm not sure at the moment… The nature of AF loss is the maximization of class separation, and expecting two labels as a prediction doesn't give a reasonable convergence(</p>\n<p>Another way is computing the loss in the same way as for strong labels, but using multi-head <strong>attention to weight the loss across the generated sequence</strong>. It can be seen as the following. Let's assume that there are two birds present in the record, and AF is low within these S1 and S2 segments if corresponding labels are considered as GT (S1-L1, S2-L2), and is high outside. The attention should perform a soft selection of segments S1 and S2 as considered concepts and assign them to labels L1 and L2. It will minimize the total loss. This concept is somewhat similar to Maskformer, and you can use it as a starting point for thinking, but it is not exactly the same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1804532,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/29/2022 05:57:16",
          "content": "<p>Thanks for the detailed reply.</p>\n<p>Yes, relaxing condition will make it easier to adapt ArcFace to this problem.</p>\n<ul>\n<li>assume at most one bird call is appeared in a segment</li>\n<li>strong label (segment-wise label) is provided</li>\n</ul>\n<p>So if I understand correctly, your idea is using attention score for each class as a strong label?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1805322,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/30/2022 04:21:18",
          "content": "<p>The thing I thought of is a way to reweight the loss along the sequence, i.e. the loss is computed for parts where the corresponding birds are expected to be, while other parts are having a small contribution…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1805397,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/30/2022 06:05:14",
          "content": "<p>I see. Thanks for confirmation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1800628": "### Summary\n- **Constant Q-transform**\n- Multistep training with PL\n- Pretraining on 2022+2021 data with noise from Rainforest\n- Finetuning on 2022 data with weighted sampling (w~N^0.5)\n- Sequence-based model\n- Single model performance **0.7971/0.7859** private/public LB, while two models give **0.7964/0.7936**\n\n### Introduction\nCongratulations to all participants and thanks to the organizers for making this competition possible. It is my 3rd BirdCLEF competition, and the first gold medal I got in it. Though, this time I again followed the tradition of joining the competition close to the end... with the first sub 5-6 days before the deadline( Something should be changed in this life.\n\nI'll take this opportunity and describe some of the ideas I came up with while working on the challenge. I hope they may be interesting for other participants and organizers as well. My work, to a large extent, is based on my [BirdCLEF 2021 paper](http://ceur-ws.org/Vol-2936/paper-141.pdf). Also, some details may be found in my [BirdCLEF 2021](https://www.kaggle.com/competitions/birdclef-2021/discussion/243343) and [2020 Cornell Birdcall](https://www.kaggle.com/c/birdsong-recognition/discussion/183258) writeups.\n\n### Data\nIn this year's competition there were two main challenges to address:  **(1) considerable domain mismatch with very noisy test data** and (2) **long-tail class distribution with a few examples of rare classes**. \n\nTo address the first challenge, I used a traditional **pseudo label (PL) 2-step training**: in the first pass the model learns where the signal from birds is located, and on the second pass a heavy noise is applied. With segmentation labels the model knows where to look for the true signal under a severe noise. As a noise source [rainforest data](https://www.kaggle.com/c/rfcx-species-audio-detection/data) is used in addition to pink noise and signal weakening.\n\nTo address the second challenge and make the model robust to different bird calls I pretrained it on a large **2021+2022 BirdCLEF dataset** (the size of 2022 data is ~ 5 times smaller). During fine-tuning I kept the backbone frozen and trained only the head on 2022 data. I assigned all unscored classes to \"unknown\" label and performed weighted sampling with probability conditioned on the number of samples of each class as w~N^0.5.\n\nOne of the novel things I used in my solution, in comparison to traditional audio models, is **Constant Q-transform** instead of mel-scale spectrograms and even pcen. The quality of the produced spectrograms is much better, as shown in the figure below: lower impact of noise, high level of details in the high-frequency domain, high contrast, and easy mapping for sequence models (without the need to care about padding added in FFT). However, I didn't have time for a quantitative comparison of the model performance. CQT1992v2 is a very effective Pytorch (GPU) implementation of CQT from nnAudio.Spectrogram library and can be easily integrated into a model (make sure to disable gradients for CQT).  During training I used log transformation of transform output: `x = (64*x + 1).log()`.\n![](https://i.ibb.co/Trv7Qkb/spectrogram.png)\n\n```\nself.qtransform = CQT1992v2(sr=32000, fmin=256, n_bins=160, \nhop_length=250, output_format='Magnitude', norm=1,\nwindow='tukey',bins_per_octave=27,)\n```\n\n### Model\nI used [semisupervised pretrained](https://github.com/facebookresearch/semi-supervised-ImageNet1K-models) ResNeXt50 and MiT-B2 transformer, [Segformer ](https://github.com/NVlabs/SegFormer) backbone. The last conv layer is followed by a convolution, collapsing the feature map into a sequence of vectors vector (nemb=1024), with a Transformer layer, and a head, producing sequence and clip outputs using the attention mechanism (somewhat similar to SED models):\n```\nclass AttHead(nn.Module):\n    def __init__(self, n_in, n_out):\n        super().__init__()\n        self.attn = nn.Conv1d(n_in,n_out,1)\n        self.cla = nn.Conv1d(n_in,n_out,1)\n        \n    def forward(self, x):\n        if len(x.shape) == 4: x = x.flatten(1,2)\n        attn = self.attn(x)\n        cla = self.cla(x)\n        x = (torch.softmax(attn,-1)*cla).sum(-1)\n        return x, attn, cla\n```\nInitially, I believed that using MiT transformer backbone with a large receptive field is important: it is able to look into the entire frequency domain at one (instead of a small window for conv net) and compare bird calls in the clip across the time. But it appeared to be not really true, and my favorite ResNeXt50 performed nearly the same as MiT B2... both CV and LB. Though, in both cases using a transformer head quite helps based on my initial tests.  \n\n### Training\nI used multistep training: (1) pretrain the model on 2022+2021 data and generate PL; (2) train the model on 2022+2021 data with a high level of noise added and additional segmentation loss based on PL (global labels for selected chunks are also adjusted based on PL in case if there is no a bird call); (3) take the model from step 2 and freeze the backbone, finetune on 2022 data with weighted sampling considering 21+1 classes, and generate new PL; (4)  the same as 3 but noise and segmentation losses are added. \n\nAt all stages MixUp augmentation (with label max and alpha of 2, i.e. the probabilities are close to 0.5) is applied to waves. Each stage is started with using 5s chunks (32 epochs) and is finished with a few epochs with 10-15s chunks. I use Focal loss with gamma=1 (with a bug fix to accept smooth labels).\n\n### Inference\nThe inference is performed on entire audio files (while training is done on 5s clips followed with fine-tuning on 10-15s clips). I use sequence level output, illustrated in the image below for the test audio clip. It is split into 5s chinks, and the predictions are selected if the maximum output of the particular class reaches the selected threshold. If one compares the plot with similar plots from my previous reports, the model performance improvement is quite clear. Now the model is able to clearly distinguish bird calls even in very noisy audio. For postprocessing the model output I was using the following: `x = (torch.softmax(3*attn,-2)*torch.sigmoid(3*cla))`. Pay attention to the dimension, which is the class rather than the sequence dimension: if the model is paying more attention to a specific class at a particular moment, the signal should be enhanced accordingly. Also, I'm using temperature rescaling.\n![](https://i.ibb.co/sKXxh8J/pred-test.png)\n\nThe final submission is composed of 2 models (MiT and ResNeXt50) and is scored as 0.7964/0.7936 at public/private LB. The best single model is scored as 0.7971/0.7859. In this competition, the main focus is bird classification, and no credit is given for call/nocall separation, in contrast to 2020 and 2021 competitions. So the best scores could be achieved at the selection of very low thresholds. If more reasonable thresholds, capable of nocall separation, are used, the performance drops by ~0.05.\n\n### Things ~~didn't work~~ I didn't have time to make work\n- Incorporation of ArcFace loss into attention pooling head (dealing with sequence and global labels).\n- Performing metric learning and clustering to address the issue with a few training examples for rare classes\n\nI think those two things might be the key to getting to the top of the LB.",
    "1800639": "Congrats on strong finish and solo gold medal @iafoss",
    "1800648": "Thank you @duykhanh99",
    "1800662": "Congratulations for the solo gold @iafoss , I'm wondering what is your validation strategy like, and how do you determine the threshold? As we don't have train soundscapes files like last year, I was struggling at CV/LB mismatching all the time.",
    "1800704": "Thanks. **In this competition CV strategy for threshold identification doesn't exist**. Let me explain why. \n(1) In BirdCLEF competitions there is a tradition to have a huge domain gap between train and test data. Test data is of poor quality with large background noise and 5s chunk annotation. At the same time, train data is often of exceptional quality with a sequence-level annotation (chunks selected during evaluation may not have bird calls). One may try to emulate test data by adding a high level of noise to val and using PL to identify nocall parts to ensure that a specific 5s chunk includes a bird call. Also if secondary birds are percent, it becomes more sophisticated... Though this procedure should be calibrated with LB to set the level of noise at least. Such CV strategy could work in BirdCLEF 2020 competition, but not here because of the following.\n(2) **The host doesn't score all chunks and all species** but does cherry-picking, i.e. ignore nocall cases and considers only a few options for each chunk. It is like if one heard a birdcall and is asking if it is A or B species. Since **this procedure is unknown**, the threshold should be calibrated based on LB( Though it is a kind of cheating, and the host is getting an overestimated model performance because we \"overfit\" the procedure used to select scored species both in public and private test set.\n\nDuring training I was computing some F1 score, but to assess the model performance I mostly paid attention to the model output at the test data, like one in my post, because of the issue (1): we do not need a model that performs well without noise, we need a model that works on test data. So listening to the test audio and visual inspection of the model output on it is not a bad thing here. After participating in competitions like BirdCLEF or PANDA, where validation is quite illusional, one gets instincts on model assessment and selection... **The key is following the methodology, i.e. focusing on the things that make sense.** Though one should understand when LB or CV or both should be ignored, it's another key.",
    "1800812": "Congrats @iafoss on your amazing score in a short time of competing and thanks for sharing your solution.",
    "1801004": "Congrats for your solo gold and by achieving it not using any leak. Thanks for sharing your approach.",
    "1801143": "Thank you so much @titericz",
    "1801173": "You are very welcome, @tjamali",
    "1801257": "Congratz on getting good score bro and Thanks for sharing",
    "1801484": "Thanks @muhammadabbasshareef",
    "1802008": "Congratulations on your solo gold medal and your first gold medal in the BirdCLEF competition :)",
    "1802258": "Thanks so much, @naoism. I'm happy that I got it this time. I quite remember our hard work at the end of BirdCLEF 2021 and how close we were to the gold... I expected that you would be interested in BirdCLEF 2022.",
    "1802607": "BirdCLEF2021 is quite an impressive competition for me too. To work with you was quite a fun memory. (Also I remember I was really frustrated not to get gold medal at the time...) Actually I was planning to participate and did also EDA, but I didn't because we didn't have data for validation like BirdCLEF2021...",
    "1804447": "Congratulations for solo gold in only 25 submissions. It's the least submissions in the gold winners.\n\nIf you had enough time, how would you have implemented ArcFace loss into attention pooling head?\nI think this it's not straight-forward because there are two differences in the original ArcFace condition:\n1. it has multiple labels\n2. it has time dimension (or, did you planned to apply ArcFace on pooled vector output (bs, d, 1, 1)?)",
    "1804509": "If the provided labels were stronger, i.e. include time tags for the beginning and end of the bird call, everything would be super easy: just do a sound segmentation model assuming that at each moment of time only a single bird may be present at the most.\n\nThe weakness of labels screws this simple picture. Ideally, it would be preferable to set an approach dealing with the provided annotation... One thing, as you mentioned, may be using an attention mechanism to pool a single vector out of generated sequence of vectors. The issue here, though, is how to incorporate a multi-label paradigm into it. I'm not sure at the moment... The nature of AF loss is the maximization of class separation, and expecting two labels as a prediction doesn't give a reasonable convergence(\n\nAnother way is computing the loss in the same way as for strong labels, but using multi-head **attention to weight the loss across the generated sequence**. It can be seen as the following. Let's assume that there are two birds present in the record, and AF is low within these S1 and S2 segments if corresponding labels are considered as GT (S1-L1, S2-L2), and is high outside. The attention should perform a soft selection of segments S1 and S2 as considered concepts and assign them to labels L1 and L2. It will minimize the total loss. This concept is somewhat similar to Maskformer, and you can use it as a starting point for thinking, but it is not exactly the same.",
    "1804532": "Thanks for the detailed reply.\n\nYes, relaxing condition will make it easier to adapt ArcFace to this problem.\n\n* assume at most one bird call is appeared in a segment\n* strong label (segment-wise label) is provided\n\nSo if I understand correctly, your idea is using attention score for each class as a strong label?",
    "1805322": "The thing I thought of is a way to reweight the loss along the sequence, i.e. the loss is computed for parts where the corresponding birds are expected to be, while other parts are having a small contribution...",
    "1805397": "I see. Thanks for confirmation."
  },
  "source": "meta"
}