{
  "id": 243463,
  "title": "2nd place solution",
  "url": "/competitions/birdclef-2021/writeups/new-baseline-2nd-place-solution",
  "author_name": "",
  "post_date": "2021-06-02T16:52:28.310Z",
  "votes": 139,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Thanks to Kaggle, the hosts, and our fellow competitors for this very interesting competition. In the following, we want to give a rough overview of our 2nd place solution. We joined this competition in the last 3 weeks and worked hard to fill the knowledge gap to last and previous years’ participants. </p>\n<p>As always, this has been an incredible team effort and has been equal contribution by <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> - I was just the lucky one winning the roll to post the solution :)</p>\n<p><strong>TLDR</strong></p>\n<p>Our solution is an ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. We used mixup and added background noise as an augmentation method to improve generalization of our models. For inference, we predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.</p>\n<p><strong>Validation</strong></p>\n<p>I am sure most participants are aware that a robust validation setup is quite difficult in this competition given the fact that test contains different species, and specifically also two additional sites for which we have no validation labels at all. We still tried to come up with a somewhat robust validation setup. </p>\n<p>All our models are only fit on short clips and we always evaluate on train soundscapes. That means that the idea of having multiple folds is redundant here and we basically have only one full validation set containing all soundscape files. <br>\nOne thing we quite soon noticed was that if you evaluate on full soundscapes, the validation F1 score is significantly higher than on LB. Our final validation on the full soundscapes was close to 0.84. We figured that this has mostly to do with the presence of 3 full songs in validation that do not contain any calls at all. We also saw on the sample submission score that at least public LB contains more birds than the full soundscape dataset would suggest. So as a first step, we mostly focused on evaluating all but these three songs for our validation score, let’s call it CV-3 (~0.81).</p>\n<p>To make it even more robust, we decided to introduce bootstrapping with the following steps:</p>\n<ul>\n<li>Remove 3 songs without calls</li>\n<li>For k times (e.g., 10) sample 80% of the remaining songs - this should emulate the full test dataset (public+private)</li>\n<li>Apply any kind of threshold selection technique, post processing, etc. on this data as we have to do the same when submitting (as we have a combination of public / private there and don’t know what is what).</li>\n<li>For j times (e.g., 50) sample 65% of the remaining songs - this should emulate the private test dataset.</li>\n<li>Calculate the score on each of these j samples.</li>\n<li>Report average, median, min, max, std scores across all k times j (e.g., 500) subsets</li>\n</ul>\n<p>This is how such an evaluation then looks like:</p>\n<p><img src=\"https://i.imgur.com/4yOisgS.png\" alt=\"\"></p>\n<p><strong>Code Pipeline and data setup</strong></p>\n<p>We used github for code storage and versioning and neptune.ai for logging and sharing our experiments. To reduce CPU bottleneck we could have preprocessed mel spectrograms to disk, but in order to be flexible with respect to trying different hyperparameters we instead performed mel spec transformation on GPU using <a href=\"https://pytorch.org/audio/stable/index.html\" target=\"_blank\">torchaudio</a>. We also did mixup augmentation on the GPU and used mixed precision training to further speed up runtime. For all models we used pytorch with CNN backbones from <a href=\"https://github.com/rwightman/pytorch-image-models/\" target=\"_blank\">timm</a>. </p>\n<p><strong>Binary classifier</strong></p>\n<p>We trained a binary classifier to predict bird / no bird in order to try various ideas with respect to pre- and postprocessing. In the end, we only use it for one postprocessing step. For this, we used 3 datasets containing binary labels of 10sec recordings (freefield1010, warblrb10k, BirdVox-DCASE-20k) available <a href=\"http://dcase.community/challenge2018/task-bird-audio-detection\" target=\"_blank\">online</a>. The model is very similar to SED model used in several past solutions.</p>\n<p>Backbones: seresnext26t_32x4d, tf_efficientnet_b0_ns</p>\n<p><strong>Bird classifier</strong></p>\n<p>Our models were pretty similar and were all trained on 30 sec random crops of the train_short data. 30 seconds was beneficial as we do not know where the labels are (weak labels). To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes. We naively used the union of primary and secondary label as target. For inference then we directly fed 5sec snippets to the model.</p>\n<p><img src=\"https://i.imgur.com/M81KcGr.png\" alt=\"\"></p>\n<p>We used the following backbones: resnet34, tf_efficientnetv2_s_in21k, tf_efficientnetv2_m_in21k, eca_nfnet_l0</p>\n<p>We trained with BCE loss using Adam optimizer and cosine annealing schedule. We saw improvements using the following tricks:</p>\n<ul>\n<li>Use the rating for weighting the recordings contribution to the loss. The assumption is that recordings with a lower rating have worse quality with respect to audio and label and should contribute less to model training. In detail we weight each sample by rating/max(ratings).</li>\n<li>Label smoothing. We used label smoothing to account for noisy annotations and absence of birds in “unlucky” 30sec crops.</li>\n<li>Clever augmentation. Similar to past solutions we used no-bird background noise and mixup as main augmentation methods. For background noise we used a mix of no-call parts of this years validation set and past years data. We also not only used mixup between recordings but also within a recording by mixing between the 5 sec parts. In mixup we also weight the labels and sample weights accordingly.</li>\n</ul>\n<p><strong>Ensembling</strong></p>\n<p>The ensembling of our models was straightforward since all output the same shapes. We took a simple mean of the predictions after a step of post-processing which is explained in the next paragraph. At the end we used 9 models which differ mostly on hyperparameters and backbones and fitted each model with 6 different seeds. Our final kernel ran in approximately 1h, so there was still quite some room in the kernel.</p>\n<p><strong>Post processing</strong></p>\n<p>The first step for post processing involved choosing an appropriate threshold for making hard predictions for which birds are present in a 5 second segment in soundscapes. As we all know, given the f-score metric, this is one of the most crucial steps of the solution. Even though optimizing a hard threshold on validation and applying it on LB worked quite well, we understood that there are some issues with that approach.</p>\n<p>First, we quickly realized that test and regular validation had different proportions of nocalls and calls which was also apparent from the different sample submission scores (only nocalls). This means that in general you wanted to predict more birds on LB meaning lowering thresholds to a certain degree could be helpful for improving public LB. We also accounted for this imbalance in our validation setup by removing the three nocall songs (see above, CV-3).</p>\n<p>Second, choosing hard thresholds can be problematic when you introduce new blends to your solution. Each new model has certain shifts in probabilities for all and certain birds, so the global thresholds can shift quite a bit. Now it became hard for us to properly judge if new models work well in the blend on validation and LB based on the merit of the models, or only based on some arbitrary probability / threshold shifts that emerged from it. And it was unclear what is a result of random fluctuation, or model properties.</p>\n<p>To that end, we decided to move to a percentile based thresholding approach. In detail, this meant that we set a certain percentile of predictions we want to do on a validation or test set, and calculated the according threshold that way. We did this by flattening all predictions, and then calculating the threshold. On CV-3 this looked for example like that:</p>\n<p><code>threshold = np.percentile(y_preds.flatten(), 0.9987)</code></p>\n<p>The more birds a set contains, the lower the percentile can be if predictions are decently ranked. The good thing now with this approach was that we could keep the percentile stable, and just exchange models, blends and other post processing and if the quality in our ranking of predictions improved, also the score improved given this fixed percentile, because we always predict the same amount of records.</p>\n<p>After we had this setup, we played a bit with changing the percentile on LB to check how test differs in that sense. We found the optimum on public LB to be at around 0.9980 meaning that quite a few more birds are present. In our final sub we chose 0.9981 and made another gamble with 0.9973. The better sub was clearly 0.9981, and actually even a bit higher could have neted us a potential first place (closer to best percentile on validation).</p>\n<p>In theory the gamble was legit, because private LB even had more birds as imminent from sample submission. But at the same time it seems that the ranking of predictions was worse, so that lower percentiles introduce too many FPs, meaning that more conservative setting was better. By and large, our choice based on a combination of validation and LB was a very robust one in the end, and we believe that this percentile based approach was way more stable and robust than individual threshold optimization.</p>\n<p>Additionally, we employed several smaller post processing steps to improve the predictions including attempts like: (1) increasing the probability of birds in songs based on their average prediction probability, (2) smoothing neighboring predictions, or (3) adjusting predictions by the predictions from our binary models. We also removed some unlikely predictions based on distance in space and time given the metadata very similar to how 4th place did.</p>\n<p><strong>What did not work</strong></p>\n<p>In the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work. One thing to note is TTA which we could not make work. We had quite some time left in the kernel runtime, so this was a natural area to explore, but TTA with mel spectrograms is not as straightforward as with usual CV data. Furthermore, we tried to explore pseudo tagging in different versions, but also could not improve our blends with it.</p>\n<p>Thanks for reading. Questions are very welcome. <br>\nChristof, Pascal &amp; Philipp</p>",
  "messages": [
    {
      "id": "1333338",
      "postDate": "06/02/2021 16:36:48",
      "content": "<p>Thanks to Kaggle, the hosts, and our fellow competitors for this very interesting competition. In the following, we want to give a rough overview of our 2nd place solution. We joined this competition in the last 3 weeks and worked hard to fill the knowledge gap to last and previous years’ participants. </p>\n<p>As always, this has been an incredible team effort and has been equal contribution by <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> - I was just the lucky one winning the roll to post the solution :)</p>\n<p><strong>TLDR</strong></p>\n<p>Our solution is an ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. We used mixup and added background noise as an augmentation method to improve generalization of our models. For inference, we predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.</p>\n<p><strong>Validation</strong></p>\n<p>I am sure most participants are aware that a robust validation setup is quite difficult in this competition given the fact that test contains different species, and specifically also two additional sites for which we have no validation labels at all. We still tried to come up with a somewhat robust validation setup. </p>\n<p>All our models are only fit on short clips and we always evaluate on train soundscapes. That means that the idea of having multiple folds is redundant here and we basically have only one full validation set containing all soundscape files. <br>\nOne thing we quite soon noticed was that if you evaluate on full soundscapes, the validation F1 score is significantly higher than on LB. Our final validation on the full soundscapes was close to 0.84. We figured that this has mostly to do with the presence of 3 full songs in validation that do not contain any calls at all. We also saw on the sample submission score that at least public LB contains more birds than the full soundscape dataset would suggest. So as a first step, we mostly focused on evaluating all but these three songs for our validation score, let’s call it CV-3 (~0.81).</p>\n<p>To make it even more robust, we decided to introduce bootstrapping with the following steps:</p>\n<ul>\n<li>Remove 3 songs without calls</li>\n<li>For k times (e.g., 10) sample 80% of the remaining songs - this should emulate the full test dataset (public+private)</li>\n<li>Apply any kind of threshold selection technique, post processing, etc. on this data as we have to do the same when submitting (as we have a combination of public / private there and don’t know what is what).</li>\n<li>For j times (e.g., 50) sample 65% of the remaining songs - this should emulate the private test dataset.</li>\n<li>Calculate the score on each of these j samples.</li>\n<li>Report average, median, min, max, std scores across all k times j (e.g., 500) subsets</li>\n</ul>\n<p>This is how such an evaluation then looks like:</p>\n<p><img src=\"https://i.imgur.com/4yOisgS.png\" alt=\"\"></p>\n<p><strong>Code Pipeline and data setup</strong></p>\n<p>We used github for code storage and versioning and neptune.ai for logging and sharing our experiments. To reduce CPU bottleneck we could have preprocessed mel spectrograms to disk, but in order to be flexible with respect to trying different hyperparameters we instead performed mel spec transformation on GPU using <a href=\"https://pytorch.org/audio/stable/index.html\" target=\"_blank\">torchaudio</a>. We also did mixup augmentation on the GPU and used mixed precision training to further speed up runtime. For all models we used pytorch with CNN backbones from <a href=\"https://github.com/rwightman/pytorch-image-models/\" target=\"_blank\">timm</a>. </p>\n<p><strong>Binary classifier</strong></p>\n<p>We trained a binary classifier to predict bird / no bird in order to try various ideas with respect to pre- and postprocessing. In the end, we only use it for one postprocessing step. For this, we used 3 datasets containing binary labels of 10sec recordings (freefield1010, warblrb10k, BirdVox-DCASE-20k) available <a href=\"http://dcase.community/challenge2018/task-bird-audio-detection\" target=\"_blank\">online</a>. The model is very similar to SED model used in several past solutions.</p>\n<p>Backbones: seresnext26t_32x4d, tf_efficientnet_b0_ns</p>\n<p><strong>Bird classifier</strong></p>\n<p>Our models were pretty similar and were all trained on 30 sec random crops of the train_short data. 30 seconds was beneficial as we do not know where the labels are (weak labels). To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes. We naively used the union of primary and secondary label as target. For inference then we directly fed 5sec snippets to the model.</p>\n<p><img src=\"https://i.imgur.com/M81KcGr.png\" alt=\"\"></p>\n<p>We used the following backbones: resnet34, tf_efficientnetv2_s_in21k, tf_efficientnetv2_m_in21k, eca_nfnet_l0</p>\n<p>We trained with BCE loss using Adam optimizer and cosine annealing schedule. We saw improvements using the following tricks:</p>\n<ul>\n<li>Use the rating for weighting the recordings contribution to the loss. The assumption is that recordings with a lower rating have worse quality with respect to audio and label and should contribute less to model training. In detail we weight each sample by rating/max(ratings).</li>\n<li>Label smoothing. We used label smoothing to account for noisy annotations and absence of birds in “unlucky” 30sec crops.</li>\n<li>Clever augmentation. Similar to past solutions we used no-bird background noise and mixup as main augmentation methods. For background noise we used a mix of no-call parts of this years validation set and past years data. We also not only used mixup between recordings but also within a recording by mixing between the 5 sec parts. In mixup we also weight the labels and sample weights accordingly.</li>\n</ul>\n<p><strong>Ensembling</strong></p>\n<p>The ensembling of our models was straightforward since all output the same shapes. We took a simple mean of the predictions after a step of post-processing which is explained in the next paragraph. At the end we used 9 models which differ mostly on hyperparameters and backbones and fitted each model with 6 different seeds. Our final kernel ran in approximately 1h, so there was still quite some room in the kernel.</p>\n<p><strong>Post processing</strong></p>\n<p>The first step for post processing involved choosing an appropriate threshold for making hard predictions for which birds are present in a 5 second segment in soundscapes. As we all know, given the f-score metric, this is one of the most crucial steps of the solution. Even though optimizing a hard threshold on validation and applying it on LB worked quite well, we understood that there are some issues with that approach.</p>\n<p>First, we quickly realized that test and regular validation had different proportions of nocalls and calls which was also apparent from the different sample submission scores (only nocalls). This means that in general you wanted to predict more birds on LB meaning lowering thresholds to a certain degree could be helpful for improving public LB. We also accounted for this imbalance in our validation setup by removing the three nocall songs (see above, CV-3).</p>\n<p>Second, choosing hard thresholds can be problematic when you introduce new blends to your solution. Each new model has certain shifts in probabilities for all and certain birds, so the global thresholds can shift quite a bit. Now it became hard for us to properly judge if new models work well in the blend on validation and LB based on the merit of the models, or only based on some arbitrary probability / threshold shifts that emerged from it. And it was unclear what is a result of random fluctuation, or model properties.</p>\n<p>To that end, we decided to move to a percentile based thresholding approach. In detail, this meant that we set a certain percentile of predictions we want to do on a validation or test set, and calculated the according threshold that way. We did this by flattening all predictions, and then calculating the threshold. On CV-3 this looked for example like that:</p>\n<p><code>threshold = np.percentile(y_preds.flatten(), 0.9987)</code></p>\n<p>The more birds a set contains, the lower the percentile can be if predictions are decently ranked. The good thing now with this approach was that we could keep the percentile stable, and just exchange models, blends and other post processing and if the quality in our ranking of predictions improved, also the score improved given this fixed percentile, because we always predict the same amount of records.</p>\n<p>After we had this setup, we played a bit with changing the percentile on LB to check how test differs in that sense. We found the optimum on public LB to be at around 0.9980 meaning that quite a few more birds are present. In our final sub we chose 0.9981 and made another gamble with 0.9973. The better sub was clearly 0.9981, and actually even a bit higher could have neted us a potential first place (closer to best percentile on validation).</p>\n<p>In theory the gamble was legit, because private LB even had more birds as imminent from sample submission. But at the same time it seems that the ranking of predictions was worse, so that lower percentiles introduce too many FPs, meaning that more conservative setting was better. By and large, our choice based on a combination of validation and LB was a very robust one in the end, and we believe that this percentile based approach was way more stable and robust than individual threshold optimization.</p>\n<p>Additionally, we employed several smaller post processing steps to improve the predictions including attempts like: (1) increasing the probability of birds in songs based on their average prediction probability, (2) smoothing neighboring predictions, or (3) adjusting predictions by the predictions from our binary models. We also removed some unlikely predictions based on distance in space and time given the metadata very similar to how 4th place did.</p>\n<p><strong>What did not work</strong></p>\n<p>In the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work. One thing to note is TTA which we could not make work. We had quite some time left in the kernel runtime, so this was a natural area to explore, but TTA with mel spectrograms is not as straightforward as with usual CV data. Furthermore, we tried to explore pseudo tagging in different versions, but also could not improve our blends with it.</p>\n<p>Thanks for reading. Questions are very welcome. <br>\nChristof, Pascal &amp; Philipp</p>",
      "rawMarkdown": "Thanks to Kaggle, the hosts, and our fellow competitors for this very interesting competition. In the following, we want to give a rough overview of our 2nd place solution. We joined this competition in the last 3 weeks and worked hard to fill the knowledge gap to last and previous years’ participants. \n\nAs always, this has been an incredible team effort and has been equal contribution by @christofhenkel, @ilu000 and @philippsinger - I was just the lucky one winning the roll to post the solution :)\n\n**TLDR**\n\nOur solution is an ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. We used mixup and added background noise as an augmentation method to improve generalization of our models. For inference, we predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.\n\n**Validation**\n\nI am sure most participants are aware that a robust validation setup is quite difficult in this competition given the fact that test contains different species, and specifically also two additional sites for which we have no validation labels at all. We still tried to come up with a somewhat robust validation setup. \n\nAll our models are only fit on short clips and we always evaluate on train soundscapes. That means that the idea of having multiple folds is redundant here and we basically have only one full validation set containing all soundscape files. \nOne thing we quite soon noticed was that if you evaluate on full soundscapes, the validation F1 score is significantly higher than on LB. Our final validation on the full soundscapes was close to 0.84. We figured that this has mostly to do with the presence of 3 full songs in validation that do not contain any calls at all. We also saw on the sample submission score that at least public LB contains more birds than the full soundscape dataset would suggest. So as a first step, we mostly focused on evaluating all but these three songs for our validation score, let’s call it CV-3 (~0.81).\n\nTo make it even more robust, we decided to introduce bootstrapping with the following steps:\n\n- Remove 3 songs without calls\n- For k times (e.g., 10) sample 80% of the remaining songs - this should emulate the full test dataset (public+private)\n- Apply any kind of threshold selection technique, post processing, etc. on this data as we have to do the same when submitting (as we have a combination of public / private there and don’t know what is what).\n- For j times (e.g., 50) sample 65% of the remaining songs - this should emulate the private test dataset.\n- Calculate the score on each of these j samples.\n- Report average, median, min, max, std scores across all k times j (e.g., 500) subsets\n\nThis is how such an evaluation then looks like:\n\n![](https://i.imgur.com/4yOisgS.png)\n\n**Code Pipeline and data setup**\n\nWe used github for code storage and versioning and neptune.ai for logging and sharing our experiments. To reduce CPU bottleneck we could have preprocessed mel spectrograms to disk, but in order to be flexible with respect to trying different hyperparameters we instead performed mel spec transformation on GPU using [torchaudio](https://pytorch.org/audio/stable/index.html). We also did mixup augmentation on the GPU and used mixed precision training to further speed up runtime. For all models we used pytorch with CNN backbones from [timm](https://github.com/rwightman/pytorch-image-models/). \n\n**Binary classifier**\n\nWe trained a binary classifier to predict bird / no bird in order to try various ideas with respect to pre- and postprocessing. In the end, we only use it for one postprocessing step. For this, we used 3 datasets containing binary labels of 10sec recordings (freefield1010, warblrb10k, BirdVox-DCASE-20k) available [online](http://dcase.community/challenge2018/task-bird-audio-detection). The model is very similar to SED model used in several past solutions.\n\nBackbones: seresnext26t_32x4d, tf_efficientnet_b0_ns\n\n**Bird classifier**\n\nOur models were pretty similar and were all trained on 30 sec random crops of the train_short data. 30 seconds was beneficial as we do not know where the labels are (weak labels). To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes. We naively used the union of primary and secondary label as target. For inference then we directly fed 5sec snippets to the model.\n\n![](https://i.imgur.com/M81KcGr.png)\n\nWe used the following backbones: resnet34, tf_efficientnetv2_s_in21k, tf_efficientnetv2_m_in21k, eca_nfnet_l0\n\nWe trained with BCE loss using Adam optimizer and cosine annealing schedule. We saw improvements using the following tricks:\n\n- Use the rating for weighting the recordings contribution to the loss. The assumption is that recordings with a lower rating have worse quality with respect to audio and label and should contribute less to model training. In detail we weight each sample by rating/max(ratings).\n- Label smoothing. We used label smoothing to account for noisy annotations and absence of birds in “unlucky” 30sec crops.\n- Clever augmentation. Similar to past solutions we used no-bird background noise and mixup as main augmentation methods. For background noise we used a mix of no-call parts of this years validation set and past years data. We also not only used mixup between recordings but also within a recording by mixing between the 5 sec parts. In mixup we also weight the labels and sample weights accordingly.\n\n**Ensembling**\n\nThe ensembling of our models was straightforward since all output the same shapes. We took a simple mean of the predictions after a step of post-processing which is explained in the next paragraph. At the end we used 9 models which differ mostly on hyperparameters and backbones and fitted each model with 6 different seeds. Our final kernel ran in approximately 1h, so there was still quite some room in the kernel.\n\n**Post processing**\n\nThe first step for post processing involved choosing an appropriate threshold for making hard predictions for which birds are present in a 5 second segment in soundscapes. As we all know, given the f-score metric, this is one of the most crucial steps of the solution. Even though optimizing a hard threshold on validation and applying it on LB worked quite well, we understood that there are some issues with that approach.\n\nFirst, we quickly realized that test and regular validation had different proportions of nocalls and calls which was also apparent from the different sample submission scores (only nocalls). This means that in general you wanted to predict more birds on LB meaning lowering thresholds to a certain degree could be helpful for improving public LB. We also accounted for this imbalance in our validation setup by removing the three nocall songs (see above, CV-3).\n\nSecond, choosing hard thresholds can be problematic when you introduce new blends to your solution. Each new model has certain shifts in probabilities for all and certain birds, so the global thresholds can shift quite a bit. Now it became hard for us to properly judge if new models work well in the blend on validation and LB based on the merit of the models, or only based on some arbitrary probability / threshold shifts that emerged from it. And it was unclear what is a result of random fluctuation, or model properties.\n\nTo that end, we decided to move to a percentile based thresholding approach. In detail, this meant that we set a certain percentile of predictions we want to do on a validation or test set, and calculated the according threshold that way. We did this by flattening all predictions, and then calculating the threshold. On CV-3 this looked for example like that:\n\n``threshold = np.percentile(y_preds.flatten(), 0.9987)``\n\nThe more birds a set contains, the lower the percentile can be if predictions are decently ranked. The good thing now with this approach was that we could keep the percentile stable, and just exchange models, blends and other post processing and if the quality in our ranking of predictions improved, also the score improved given this fixed percentile, because we always predict the same amount of records.\n\nAfter we had this setup, we played a bit with changing the percentile on LB to check how test differs in that sense. We found the optimum on public LB to be at around 0.9980 meaning that quite a few more birds are present. In our final sub we chose 0.9981 and made another gamble with 0.9973. The better sub was clearly 0.9981, and actually even a bit higher could have neted us a potential first place (closer to best percentile on validation).\n\nIn theory the gamble was legit, because private LB even had more birds as imminent from sample submission. But at the same time it seems that the ranking of predictions was worse, so that lower percentiles introduce too many FPs, meaning that more conservative setting was better. By and large, our choice based on a combination of validation and LB was a very robust one in the end, and we believe that this percentile based approach was way more stable and robust than individual threshold optimization.\n\nAdditionally, we employed several smaller post processing steps to improve the predictions including attempts like: (1) increasing the probability of birds in songs based on their average prediction probability, (2) smoothing neighboring predictions, or (3) adjusting predictions by the predictions from our binary models. We also removed some unlikely predictions based on distance in space and time given the metadata very similar to how 4th place did.\n\n**What did not work**\n\nIn the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work. One thing to note is TTA which we could not make work. We had quite some time left in the kernel runtime, so this was a natural area to explore, but TTA with mel spectrograms is not as straightforward as with usual CV data. Furthermore, we tried to explore pseudo tagging in different versions, but also could not improve our blends with it.\n\n\nThanks for reading. Questions are very welcome. \nChristof, Pascal & Philipp",
      "votes": null
    },
    {
      "id": "1333344",
      "postDate": "06/02/2021 16:40:10",
      "content": "<p>These were some incredible intense weeks with a lot of learning, data understanding, experimenting and tinkering. Thank you again for the awesome team-up <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> !</p>",
      "rawMarkdown": "These were some incredible intense weeks with a lot of learning, data understanding, experimenting and tinkering. Thank you again for the awesome team-up @philippsinger @christofhenkel !",
      "votes": null
    },
    {
      "id": "1333349",
      "postDate": "06/02/2021 16:42:23",
      "content": "<p>Like always you guys have maintained an impressive validation scheme. Congratz 🎉</p>",
      "rawMarkdown": "Like always you guys have maintained an impressive validation scheme. Congratz 🎉",
      "votes": null
    },
    {
      "id": "1333356",
      "postDate": "06/02/2021 16:48:32",
      "content": "<p>Congrats on the strong finish!</p>\n<blockquote>\n  <p>Use the rating for weighting the recordings contribution to the loss. </p>\n</blockquote>\n<p>Sigh, I thought of it then forget about it.  Do you know how much you get from it?</p>",
      "rawMarkdown": "Congrats on the strong finish!\n\n> Use the rating for weighting the recordings contribution to the loss. \n\nSigh, I thought of it then forget about it.  Do you know how much you get from it?",
      "votes": null
    },
    {
      "id": "1333358",
      "postDate": "06/02/2021 16:49:54",
      "content": "<p>I think it was our strongest boost on pure models out of all training related things, but we introduced it quite early so I can't fully say how much it contributed in the end with additional stuff.</p>",
      "rawMarkdown": "I think it was our strongest boost on pure models out of all training related things, but we introduced it quite early so I can't fully say how much it contributed in the end with additional stuff.",
      "votes": null
    },
    {
      "id": "1333361",
      "postDate": "06/02/2021 16:52:18",
      "content": "<p>Thank you. It was quite a significant impact of ~0.01 - 0.03 depending of where you count it (after post processing, pre-postprocessing). Also, impact on single models was always much larger than for the ensemble.</p>",
      "rawMarkdown": "Thank you. It was quite a significant impact of ~0.01 - 0.03 depending of where you count it (after post processing, pre-postprocessing). Also, impact on single models was always much larger than for the ensemble.",
      "votes": null
    },
    {
      "id": "1333375",
      "postDate": "06/02/2021 17:06:31",
      "content": "<p>Congrats on 2nd place! <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <br>\nYou said luck, but it doesn't seem like luck. :)</p>",
      "rawMarkdown": "Congrats on 2nd place! @ilu000 @philippsinger @christofhenkel \nYou said luck, but it doesn't seem like luck. :)",
      "votes": null
    },
    {
      "id": "1333379",
      "postDate": "06/02/2021 17:09:21",
      "content": "<p>I think in F-metric based competitions there is always a higher luck factor than in some others. You can only try to mitigate the impact of luck on your standing which is what we tried.</p>",
      "rawMarkdown": "I think in F-metric based competitions there is always a higher luck factor than in some others. You can only try to mitigate the impact of luck on your standing which is what we tried.",
      "votes": null
    },
    {
      "id": "1333384",
      "postDate": "06/02/2021 17:13:57",
      "content": "<p>Pff, this almost makes me will to retrain my models.</p>\n<p>Well done. I mean well done to have used it, not to make me regret ;)</p>",
      "rawMarkdown": "Pff, this almost makes me will to retrain my models.\n\nWell done. I mean well done to have used it, not to make me regret ;)",
      "votes": null
    },
    {
      "id": "1333508",
      "postDate": "06/02/2021 19:43:17",
      "content": "<p>meh, hab einfach kein würfelglück</p>",
      "rawMarkdown": "meh, hab einfach kein würfelglück",
      "votes": null
    },
    {
      "id": "1333575",
      "postDate": "06/02/2021 21:21:53",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, your 30s training to 5s inference design is very smart, and use rating for weight loss is a great idea, I will retrain my model with  such great idea to see what is the difference.  <br>\nAnd we also tried On-the-fly Logmel transformation with torchlibrosa, but I feel it is still slow, may I know why you choose torch audio for this, because it is faster or more numerically stable? Thanks. </p>",
      "rawMarkdown": "Thanks for sharing @philippsinger, your 30s training to 5s inference design is very smart, and use rating for weight loss is a great idea, I will retrain my model with  such great idea to see what is the difference.  \nAnd we also tried On-the-fly Logmel transformation with torchlibrosa, but I feel it is still slow, may I know why you choose torch audio for this, because it is faster or more numerically stable? Thanks.",
      "votes": null
    },
    {
      "id": "1333735",
      "postDate": "06/03/2021 02:46:51",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on #2 competitions ranking.</p>\n<p>First time?</p>",
      "rawMarkdown": "Congratulation @christofhenkel on #2 competitions ranking.\n\nFirst time?",
      "votes": null
    },
    {
      "id": "1333780",
      "postDate": "06/03/2021 04:02:24",
      "content": "<p>yes :D.          </p>",
      "rawMarkdown": "yes :D.",
      "votes": null
    },
    {
      "id": "1333809",
      "postDate": "06/03/2021 04:36:28",
      "content": "<p>Congrat ! </p>\n<p>The dream team strikes again !</p>\n<p><img src=\"https://i.ibb.co/5BWy3Fz/61.jpg\" alt=\"https://i.ibb.co/5BWy3Fz/61.jpg\"></p>",
      "rawMarkdown": "Congrat ! \n\nThe dream team strikes again !\n\n\n\n![https://i.ibb.co/5BWy3Fz/61.jpg](https://i.ibb.co/5BWy3Fz/61.jpg)",
      "votes": null
    },
    {
      "id": "1333972",
      "postDate": "06/03/2021 07:17:41",
      "content": "<p>What background noise did you use?  freefield1010?</p>",
      "rawMarkdown": "What background noise did you use?  freefield1010?",
      "votes": null
    },
    {
      "id": "1333991",
      "postDate": "06/03/2021 07:39:29",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on the #2 competition ranking.👍</p>",
      "rawMarkdown": "Congratulation @christofhenkel on the #2 competition ranking.👍",
      "votes": null
    },
    {
      "id": "1334025",
      "postDate": "06/03/2021 08:18:37",
      "content": "<p>As always, you guys are awesome. Thanks for the details.</p>",
      "rawMarkdown": "As always, you guys are awesome. Thanks for the details.",
      "votes": null
    },
    {
      "id": "1334063",
      "postDate": "06/03/2021 08:56:34",
      "content": "<p>Partly, yes and some from this years data (see also text). Did you use any?</p>",
      "rawMarkdown": "Partly, yes and some from this years data (see also text). Did you use any?",
      "votes": null
    },
    {
      "id": "1334064",
      "postDate": "06/03/2021 08:57:31",
      "content": "<p>I used freefield1010 no call rows as in previous comp.  I really reused the same model as in Cornell competition.</p>",
      "rawMarkdown": "I used freefield1010 no call rows as in previous comp.  I really reused the same model as in Cornell competition.",
      "votes": null
    },
    {
      "id": "1334107",
      "postDate": "06/03/2021 09:28:43",
      "content": "<p>Thank you very much for this insightful write-up. <br>\nFrom the image I deduce that you do the 'horizontal' mix-up (30s -&gt; 6x5s) on the spectrograms, not on the audio signals. Is this correct? If so, any particular reason why you didn't do it on the audio signal? </p>",
      "rawMarkdown": "Thank you very much for this insightful write-up. \nFrom the image I deduce that you do the 'horizontal' mix-up (30s -> 6x5s) on the spectrograms, not on the audio signals. Is this correct? If so, any particular reason why you didn't do it on the audio signal?",
      "votes": null
    },
    {
      "id": "1334120",
      "postDate": "06/03/2021 09:33:34",
      "content": "<p>Huge congrats! :) Outstanding performance and a superb write-up! Thanks so much for sharing all this.</p>\n<p>Could I please ask you what did you use for smoothing neighboring predictions?  </p>",
      "rawMarkdown": "Huge congrats! :) Outstanding performance and a superb write-up! Thanks so much for sharing all this.\n\nCould I please ask you what did you use for smoothing neighboring predictions?",
      "votes": null
    },
    {
      "id": "1334136",
      "postDate": "06/03/2021 09:40:14",
      "content": "<p>Thanks! Similar idea as blending 5sec with 10sec preds for example.</p>",
      "rawMarkdown": "Thanks! Similar idea as blending 5sec with 10sec preds for example.",
      "votes": null
    },
    {
      "id": "1334486",
      "postDate": "06/03/2021 14:39:41",
      "content": "<p>Yes all mixup is done on spectrograms. We didnt have time to try on raw audio in this comp, it might have also cpu bottlenecked us.</p>",
      "rawMarkdown": "Yes all mixup is done on spectrograms. We didnt have time to try on raw audio in this comp, it might have also cpu bottlenecked us.",
      "votes": null
    },
    {
      "id": "1334506",
      "postDate": "06/03/2021 15:05:09",
      "content": "<p>Can I know about the Mixup augmentation? Like can you share the code or any reference link?<br>\nAnd how did you use it here?</p>",
      "rawMarkdown": "Can I know about the Mixup augmentation? Like can you share the code or any reference link?\nAnd how did you use it here?",
      "votes": null
    },
    {
      "id": "1334899",
      "postDate": "06/03/2021 21:40:17",
      "content": "<p>It is just standard mixup, nothing special.</p>",
      "rawMarkdown": "It is just standard mixup, nothing special.",
      "votes": null
    },
    {
      "id": "1334924",
      "postDate": "06/03/2021 22:22:03",
      "content": "<p>Thanks for sharing these methods! The validation scheme makes a lot of sense wish I had thought to do something similar…</p>",
      "rawMarkdown": "Thanks for sharing these methods! The validation scheme makes a lot of sense wish I had thought to do something similar...",
      "votes": null
    },
    {
      "id": "1335234",
      "postDate": "06/04/2021 05:27:21",
      "content": "<p>We used torchlibrosa in our <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563\" target=\"_blank\">1st place solution in RFCX competition</a>. But we chose torchaudio here because its more compatible with torch 1.8.1 and was numerically more stable when using mixed precision.</p>",
      "rawMarkdown": "We used torchlibrosa in our [1st place solution in RFCX competition](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563). But we chose torchaudio here because its more compatible with torch 1.8.1 and was numerically more stable when using mixed precision.",
      "votes": null
    },
    {
      "id": "1338971",
      "postDate": "06/06/2021 22:37:19",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> , after our strongest model helped us to reach golden area, I tried write another model to increase diversity, but this model is always not stable, I used torchlibrosa and AMP as well, maybe this is the reason of the instability. Later I will make some experiment to compare torchaudio, torchlibrosa and nnAudio, in aspect of speed and numerical stability, thanks for your comments.</p>",
      "rawMarkdown": "Thanks @christofhenkel , after our strongest model helped us to reach golden area, I tried write another model to increase diversity, but this model is always not stable, I used torchlibrosa and AMP as well, maybe this is the reason of the instability. Later I will make some experiment to compare torchaudio, torchlibrosa and nnAudio, in aspect of speed and numerical stability, thanks for your comments.",
      "votes": null
    },
    {
      "id": "1340733",
      "postDate": "06/08/2021 08:11:04",
      "content": "<p>For label smoothing, could I please ask if you have ever encountered a situation where messing with the defaults of 0.9 to positive class, 0.1 distributed across all negatives, is helpful? For this competition, did you go with the default values?</p>\n<p>Thanks a lot! 😊</p>",
      "rawMarkdown": "For label smoothing, could I please ask if you have ever encountered a situation where messing with the defaults of 0.9 to positive class, 0.1 distributed across all negatives, is helpful? For this competition, did you go with the default values?\n\nThanks a lot! 😊",
      "votes": null
    },
    {
      "id": "1340809",
      "postDate": "06/08/2021 09:20:33",
      "content": "<p>+0.01 to everything</p>",
      "rawMarkdown": "0.01 to everything",
      "votes": null
    },
    {
      "id": "1341885",
      "postDate": "06/09/2021 04:27:45",
      "content": "<p>We also did label smoothing only on one side, namely 0.01 across all negatives. Positive class is unchanged.</p>",
      "rawMarkdown": "We also did label smoothing only on one side, namely 0.01 across all negatives. Positive class is unchanged.",
      "votes": null
    },
    {
      "id": "1375972",
      "postDate": "07/04/2021 15:55:56",
      "content": "<p>Thanks for sharing your solution and congratulations on the second place.<br>\nI was interested in your team's idea of \"Reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone.<br>\nI would like to ask you two questions, even though it has been a long time since the competition ended.</p>\n<ol>\n<li>There are several audios in train short audio that are less than 30s in length. How did you reshape such data into 6 x 5chunks? If you padded with zero, that part of the data would obviously lose the bird call, so it seems to be mislabeled.</li>\n<li>I thought the idea of mixup within the same audio was interesting. So I would like to know more about this. Did you apply the mixup to all 6 chuncks at the same time? Or did you pick 2 or 3 chunks out of the 6 and apply the mixup to them?</li>\n</ol>",
      "rawMarkdown": "Thanks for sharing your solution and congratulations on the second place.\nI was interested in your team's idea of \"Reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone.\nI would like to ask you two questions, even though it has been a long time since the competition ended.\n\n1. There are several audios in train short audio that are less than 30s in length. How did you reshape such data into 6 x 5chunks? If you padded with zero, that part of the data would obviously lose the bird call, so it seems to be mislabeled.\n2. I thought the idea of mixup within the same audio was interesting. So I would like to know more about this. Did you apply the mixup to all 6 chuncks at the same time? Or did you pick 2 or 3 chunks out of the 6 and apply the mixup to them?",
      "votes": null
    },
    {
      "id": "1378474",
      "postDate": "07/06/2021 15:02:46",
      "content": "<ol>\n<li><p>Padding with zeros worked pretty well, actually the whole goal of this setup is that the model will partly ignore the irrelevant chunks.</p></li>\n<li><p>Applied it to all at the same time, so each chunk might be mixed with any other chunk.</p></li>\n</ol>",
      "rawMarkdown": "1. Padding with zeros worked pretty well, actually the whole goal of this setup is that the model will partly ignore the irrelevant chunks.\n\n2. Applied it to all at the same time, so each chunk might be mixed with any other chunk.",
      "votes": null
    },
    {
      "id": "1381037",
      "postDate": "07/08/2021 14:55:16",
      "content": "<p>Thank you for your answer.<br>\nI see, by pooling (especially if it is Max pooling), the model can partly ignore irrelevant chunks.<br>\nI also understand about mixup. Thank you very much.</p>",
      "rawMarkdown": "Thank you for your answer.\nI see, by pooling (especially if it is Max pooling), the model can partly ignore irrelevant chunks.\nI also understand about mixup. Thank you very much.",
      "votes": null
    },
    {
      "id": "1467989",
      "postDate": "08/12/2021 08:03:00",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <br>\nafter a  long gap i read this solution. Recalling what i was doing corresponding to this step<br>\n<code>To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes.</code><br>\nI was doing like  this <br>\nTake  N  clips<br>\nthen do this inside fwd</p>\n<pre><code>n=len(x) #length of say 4 clips\nx = torch.stack(x,1).view(-1,shape[1],shape[2],shape[3])\nx=self.model(x)\nx = x.view(-1,n,shape[1],shape[2],shape[3]).permute(0,2,1,3,4).contiguous()\\\n            .view(-1,shape[1],shape[2]*n,shape[3])\n          #x: bs x C x n*(h x w) #feature dimensions for each channel\nx= nn.Sequential( AdaptiveConcatPool2d(),Flatten(),nn.BatchNorm1d(2*nc ) ,\n                                 nn.Linear(2*nc ,512), HardSwish() ,nn.Dropout(0.3),nn.Linear(512,n_classes)) ) (x)\n</code></pre>\n<p>is it what you also meant in your above statement of solution ?</p>",
      "rawMarkdown": "philippsinger \nafter a  long gap i read this solution. Recalling what i was doing corresponding to this step\n` To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes. `\nI was doing like  this \nTake  N  clips\nthen do this inside fwd\n```\nn=len(x) #length of say 4 clips\nx = torch.stack(x,1).view(-1,shape[1],shape[2],shape[3])\nx=self.model(x)\nx = x.view(-1,n,shape[1],shape[2],shape[3]).permute(0,2,1,3,4).contiguous()\\\n            .view(-1,shape[1],shape[2]*n,shape[3])\n          #x: bs x C x n*(h x w) #feature dimensions for each channel\nx= nn.Sequential( AdaptiveConcatPool2d(),Flatten(),nn.BatchNorm1d(2*nc ) ,\n                                 nn.Linear(2*nc ,512), HardSwish() ,nn.Dropout(0.3),nn.Linear(512,n_classes)) ) (x)\n```\nis it what you also meant in your above statement of solution ?",
      "votes": null
    },
    {
      "id": "1555655",
      "postDate": "10/24/2021 04:33:58",
      "content": "<p>Thanks for sharing your insights</p>",
      "rawMarkdown": "Thanks for sharing your insights",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1333344,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "06/02/2021 16:40:10",
      "content": "<p>These were some incredible intense weeks with a lot of learning, data understanding, experimenting and tinkering. Thank you again for the awesome team-up <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333349,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "06/02/2021 16:42:23",
      "content": "<p>Like always you guys have maintained an impressive validation scheme. Congratz 🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333356,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/02/2021 16:48:32",
      "content": "<p>Congrats on the strong finish!</p>\n<blockquote>\n  <p>Use the rating for weighting the recordings contribution to the loss. </p>\n</blockquote>\n<p>Sigh, I thought of it then forget about it.  Do you know how much you get from it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333358,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/02/2021 16:49:54",
          "content": "<p>I think it was our strongest boost on pure models out of all training related things, but we introduced it quite early so I can't fully say how much it contributed in the end with additional stuff.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1333361,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "06/02/2021 16:52:18",
          "content": "<p>Thank you. It was quite a significant impact of ~0.01 - 0.03 depending of where you count it (after post processing, pre-postprocessing). Also, impact on single models was always much larger than for the ensemble.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1333384,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/02/2021 17:13:57",
          "content": "<p>Pff, this almost makes me will to retrain my models.</p>\n<p>Well done. I mean well done to have used it, not to make me regret ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1333972,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/03/2021 07:17:41",
          "content": "<p>What background noise did you use?  freefield1010?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334063,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/03/2021 08:56:34",
          "content": "<p>Partly, yes and some from this years data (see also text). Did you use any?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334064,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/03/2021 08:57:31",
          "content": "<p>I used freefield1010 no call rows as in previous comp.  I really reused the same model as in Cornell competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333375,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "06/02/2021 17:06:31",
      "content": "<p>Congrats on 2nd place! <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <br>\nYou said luck, but it doesn't seem like luck. :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333379,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/02/2021 17:09:21",
          "content": "<p>I think in F-metric based competitions there is always a higher luck factor than in some others. You can only try to mitigate the impact of luck on your standing which is what we tried.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333508,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "06/02/2021 19:43:17",
      "content": "<p>meh, hab einfach kein würfelglück</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333575,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "06/02/2021 21:21:53",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, your 30s training to 5s inference design is very smart, and use rating for weight loss is a great idea, I will retrain my model with  such great idea to see what is the difference.  <br>\nAnd we also tried On-the-fly Logmel transformation with torchlibrosa, but I feel it is still slow, may I know why you choose torch audio for this, because it is faster or more numerically stable? Thanks. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1335234,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/04/2021 05:27:21",
          "content": "<p>We used torchlibrosa in our <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563\" target=\"_blank\">1st place solution in RFCX competition</a>. But we chose torchaudio here because its more compatible with torch 1.8.1 and was numerically more stable when using mixed precision.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1338971,
          "author_name": "superchenhao",
          "author_url": "",
          "post_date": "06/06/2021 22:37:19",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> , after our strongest model helped us to reach golden area, I tried write another model to increase diversity, but this model is always not stable, I used torchlibrosa and AMP as well, maybe this is the reason of the instability. Later I will make some experiment to compare torchaudio, torchlibrosa and nnAudio, in aspect of speed and numerical stability, thanks for your comments.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333735,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "06/03/2021 02:46:51",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on #2 competitions ranking.</p>\n<p>First time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333780,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/03/2021 04:02:24",
          "content": "<p>yes :D.          </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1333809,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "06/03/2021 04:36:28",
      "content": "<p>Congrat ! </p>\n<p>The dream team strikes again !</p>\n<p><img src=\"https://i.ibb.co/5BWy3Fz/61.jpg\" alt=\"https://i.ibb.co/5BWy3Fz/61.jpg\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1333991,
      "author_name": "mohamedbakrey",
      "author_url": "",
      "post_date": "06/03/2021 07:39:29",
      "content": "<p>Congratulation <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on the #2 competition ranking.👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1334025,
      "author_name": "ashrafr",
      "author_url": "",
      "post_date": "06/03/2021 08:18:37",
      "content": "<p>As always, you guys are awesome. Thanks for the details.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1334107,
      "author_name": "botkop",
      "author_url": "",
      "post_date": "06/03/2021 09:28:43",
      "content": "<p>Thank you very much for this insightful write-up. <br>\nFrom the image I deduce that you do the 'horizontal' mix-up (30s -&gt; 6x5s) on the spectrograms, not on the audio signals. Is this correct? If so, any particular reason why you didn't do it on the audio signal? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1334486,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/03/2021 14:39:41",
          "content": "<p>Yes all mixup is done on spectrograms. We didnt have time to try on raw audio in this comp, it might have also cpu bottlenecked us.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1334120,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "06/03/2021 09:33:34",
      "content": "<p>Huge congrats! :) Outstanding performance and a superb write-up! Thanks so much for sharing all this.</p>\n<p>Could I please ask you what did you use for smoothing neighboring predictions?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1334136,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/03/2021 09:40:14",
          "content": "<p>Thanks! Similar idea as blending 5sec with 10sec preds for example.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1334506,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "06/03/2021 15:05:09",
      "content": "<p>Can I know about the Mixup augmentation? Like can you share the code or any reference link?<br>\nAnd how did you use it here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1334899,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/03/2021 21:40:17",
          "content": "<p>It is just standard mixup, nothing special.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1334924,
      "author_name": "dryanfurman",
      "author_url": "",
      "post_date": "06/03/2021 22:22:03",
      "content": "<p>Thanks for sharing these methods! The validation scheme makes a lot of sense wish I had thought to do something similar…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1340733,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "06/08/2021 08:11:04",
      "content": "<p>For label smoothing, could I please ask if you have ever encountered a situation where messing with the defaults of 0.9 to positive class, 0.1 distributed across all negatives, is helpful? For this competition, did you go with the default values?</p>\n<p>Thanks a lot! 😊</p>",
      "votes": null,
      "replies": [
        {
          "id": 1340809,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/08/2021 09:20:33",
          "content": "<p>+0.01 to everything</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1341885,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/09/2021 04:27:45",
          "content": "<p>We also did label smoothing only on one side, namely 0.01 across all negatives. Positive class is unchanged.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1375972,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "07/04/2021 15:55:56",
      "content": "<p>Thanks for sharing your solution and congratulations on the second place.<br>\nI was interested in your team's idea of \"Reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone.<br>\nI would like to ask you two questions, even though it has been a long time since the competition ended.</p>\n<ol>\n<li>There are several audios in train short audio that are less than 30s in length. How did you reshape such data into 6 x 5chunks? If you padded with zero, that part of the data would obviously lose the bird call, so it seems to be mislabeled.</li>\n<li>I thought the idea of mixup within the same audio was interesting. So I would like to know more about this. Did you apply the mixup to all 6 chuncks at the same time? Or did you pick 2 or 3 chunks out of the 6 and apply the mixup to them?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1378474,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/06/2021 15:02:46",
          "content": "<ol>\n<li><p>Padding with zeros worked pretty well, actually the whole goal of this setup is that the model will partly ignore the irrelevant chunks.</p></li>\n<li><p>Applied it to all at the same time, so each chunk might be mixed with any other chunk.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1381037,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "07/08/2021 14:55:16",
          "content": "<p>Thank you for your answer.<br>\nI see, by pooling (especially if it is Max pooling), the model can partly ignore irrelevant chunks.<br>\nI also understand about mixup. Thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1467989,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "08/12/2021 08:03:00",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <br>\nafter a  long gap i read this solution. Recalling what i was doing corresponding to this step<br>\n<code>To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes.</code><br>\nI was doing like  this <br>\nTake  N  clips<br>\nthen do this inside fwd</p>\n<pre><code>n=len(x) #length of say 4 clips\nx = torch.stack(x,1).view(-1,shape[1],shape[2],shape[3])\nx=self.model(x)\nx = x.view(-1,n,shape[1],shape[2],shape[3]).permute(0,2,1,3,4).contiguous()\\\n            .view(-1,shape[1],shape[2]*n,shape[3])\n          #x: bs x C x n*(h x w) #feature dimensions for each channel\nx= nn.Sequential( AdaptiveConcatPool2d(),Flatten(),nn.BatchNorm1d(2*nc ) ,\n                                 nn.Linear(2*nc ,512), HardSwish() ,nn.Dropout(0.3),nn.Linear(512,n_classes)) ) (x)\n</code></pre>\n<p>is it what you also meant in your above statement of solution ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1555655,
      "author_name": "varunyadav17",
      "author_url": "",
      "post_date": "10/24/2021 04:33:58",
      "content": "<p>Thanks for sharing your insights</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1333338": "Thanks to Kaggle, the hosts, and our fellow competitors for this very interesting competition. In the following, we want to give a rough overview of our 2nd place solution. We joined this competition in the last 3 weeks and worked hard to fill the knowledge gap to last and previous years’ participants. \n\nAs always, this has been an incredible team effort and has been equal contribution by @christofhenkel, @ilu000 and @philippsinger - I was just the lucky one winning the roll to post the solution :)\n\n**TLDR**\n\nOur solution is an ensemble of several CNNs, which take a mel spectrogram representation of a 30 sec wav-crop as input. We used mixup and added background noise as an augmentation method to improve generalization of our models. For inference, we predict on 5 sec snippets and refine the result by a binary bird/nobird classifier and postprocessing to account for metadata.\n\n**Validation**\n\nI am sure most participants are aware that a robust validation setup is quite difficult in this competition given the fact that test contains different species, and specifically also two additional sites for which we have no validation labels at all. We still tried to come up with a somewhat robust validation setup. \n\nAll our models are only fit on short clips and we always evaluate on train soundscapes. That means that the idea of having multiple folds is redundant here and we basically have only one full validation set containing all soundscape files. \nOne thing we quite soon noticed was that if you evaluate on full soundscapes, the validation F1 score is significantly higher than on LB. Our final validation on the full soundscapes was close to 0.84. We figured that this has mostly to do with the presence of 3 full songs in validation that do not contain any calls at all. We also saw on the sample submission score that at least public LB contains more birds than the full soundscape dataset would suggest. So as a first step, we mostly focused on evaluating all but these three songs for our validation score, let’s call it CV-3 (~0.81).\n\nTo make it even more robust, we decided to introduce bootstrapping with the following steps:\n\n- Remove 3 songs without calls\n- For k times (e.g., 10) sample 80% of the remaining songs - this should emulate the full test dataset (public+private)\n- Apply any kind of threshold selection technique, post processing, etc. on this data as we have to do the same when submitting (as we have a combination of public / private there and don’t know what is what).\n- For j times (e.g., 50) sample 65% of the remaining songs - this should emulate the private test dataset.\n- Calculate the score on each of these j samples.\n- Report average, median, min, max, std scores across all k times j (e.g., 500) subsets\n\nThis is how such an evaluation then looks like:\n\n![](https://i.imgur.com/4yOisgS.png)\n\n**Code Pipeline and data setup**\n\nWe used github for code storage and versioning and neptune.ai for logging and sharing our experiments. To reduce CPU bottleneck we could have preprocessed mel spectrograms to disk, but in order to be flexible with respect to trying different hyperparameters we instead performed mel spec transformation on GPU using [torchaudio](https://pytorch.org/audio/stable/index.html). We also did mixup augmentation on the GPU and used mixed precision training to further speed up runtime. For all models we used pytorch with CNN backbones from [timm](https://github.com/rwightman/pytorch-image-models/). \n\n**Binary classifier**\n\nWe trained a binary classifier to predict bird / no bird in order to try various ideas with respect to pre- and postprocessing. In the end, we only use it for one postprocessing step. For this, we used 3 datasets containing binary labels of 10sec recordings (freefield1010, warblrb10k, BirdVox-DCASE-20k) available [online](http://dcase.community/challenge2018/task-bird-audio-detection). The model is very similar to SED model used in several past solutions.\n\nBackbones: seresnext26t_32x4d, tf_efficientnet_b0_ns\n\n**Bird classifier**\n\nOur models were pretty similar and were all trained on 30 sec random crops of the train_short data. 30 seconds was beneficial as we do not know where the labels are (weak labels). To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes. We naively used the union of primary and secondary label as target. For inference then we directly fed 5sec snippets to the model.\n\n![](https://i.imgur.com/M81KcGr.png)\n\nWe used the following backbones: resnet34, tf_efficientnetv2_s_in21k, tf_efficientnetv2_m_in21k, eca_nfnet_l0\n\nWe trained with BCE loss using Adam optimizer and cosine annealing schedule. We saw improvements using the following tricks:\n\n- Use the rating for weighting the recordings contribution to the loss. The assumption is that recordings with a lower rating have worse quality with respect to audio and label and should contribute less to model training. In detail we weight each sample by rating/max(ratings).\n- Label smoothing. We used label smoothing to account for noisy annotations and absence of birds in “unlucky” 30sec crops.\n- Clever augmentation. Similar to past solutions we used no-bird background noise and mixup as main augmentation methods. For background noise we used a mix of no-call parts of this years validation set and past years data. We also not only used mixup between recordings but also within a recording by mixing between the 5 sec parts. In mixup we also weight the labels and sample weights accordingly.\n\n**Ensembling**\n\nThe ensembling of our models was straightforward since all output the same shapes. We took a simple mean of the predictions after a step of post-processing which is explained in the next paragraph. At the end we used 9 models which differ mostly on hyperparameters and backbones and fitted each model with 6 different seeds. Our final kernel ran in approximately 1h, so there was still quite some room in the kernel.\n\n**Post processing**\n\nThe first step for post processing involved choosing an appropriate threshold for making hard predictions for which birds are present in a 5 second segment in soundscapes. As we all know, given the f-score metric, this is one of the most crucial steps of the solution. Even though optimizing a hard threshold on validation and applying it on LB worked quite well, we understood that there are some issues with that approach.\n\nFirst, we quickly realized that test and regular validation had different proportions of nocalls and calls which was also apparent from the different sample submission scores (only nocalls). This means that in general you wanted to predict more birds on LB meaning lowering thresholds to a certain degree could be helpful for improving public LB. We also accounted for this imbalance in our validation setup by removing the three nocall songs (see above, CV-3).\n\nSecond, choosing hard thresholds can be problematic when you introduce new blends to your solution. Each new model has certain shifts in probabilities for all and certain birds, so the global thresholds can shift quite a bit. Now it became hard for us to properly judge if new models work well in the blend on validation and LB based on the merit of the models, or only based on some arbitrary probability / threshold shifts that emerged from it. And it was unclear what is a result of random fluctuation, or model properties.\n\nTo that end, we decided to move to a percentile based thresholding approach. In detail, this meant that we set a certain percentile of predictions we want to do on a validation or test set, and calculated the according threshold that way. We did this by flattening all predictions, and then calculating the threshold. On CV-3 this looked for example like that:\n\n``threshold = np.percentile(y_preds.flatten(), 0.9987)``\n\nThe more birds a set contains, the lower the percentile can be if predictions are decently ranked. The good thing now with this approach was that we could keep the percentile stable, and just exchange models, blends and other post processing and if the quality in our ranking of predictions improved, also the score improved given this fixed percentile, because we always predict the same amount of records.\n\nAfter we had this setup, we played a bit with changing the percentile on LB to check how test differs in that sense. We found the optimum on public LB to be at around 0.9980 meaning that quite a few more birds are present. In our final sub we chose 0.9981 and made another gamble with 0.9973. The better sub was clearly 0.9981, and actually even a bit higher could have neted us a potential first place (closer to best percentile on validation).\n\nIn theory the gamble was legit, because private LB even had more birds as imminent from sample submission. But at the same time it seems that the ranking of predictions was worse, so that lower percentiles introduce too many FPs, meaning that more conservative setting was better. By and large, our choice based on a combination of validation and LB was a very robust one in the end, and we believe that this percentile based approach was way more stable and robust than individual threshold optimization.\n\nAdditionally, we employed several smaller post processing steps to improve the predictions including attempts like: (1) increasing the probability of birds in songs based on their average prediction probability, (2) smoothing neighboring predictions, or (3) adjusting predictions by the predictions from our binary models. We also removed some unlikely predictions based on distance in space and time given the metadata very similar to how 4th place did.\n\n**What did not work**\n\nIn the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work. One thing to note is TTA which we could not make work. We had quite some time left in the kernel runtime, so this was a natural area to explore, but TTA with mel spectrograms is not as straightforward as with usual CV data. Furthermore, we tried to explore pseudo tagging in different versions, but also could not improve our blends with it.\n\n\nThanks for reading. Questions are very welcome. \nChristof, Pascal & Philipp",
    "1333344": "These were some incredible intense weeks with a lot of learning, data understanding, experimenting and tinkering. Thank you again for the awesome team-up @philippsinger @christofhenkel !",
    "1333349": "Like always you guys have maintained an impressive validation scheme. Congratz 🎉",
    "1333356": "Congrats on the strong finish!\n\n> Use the rating for weighting the recordings contribution to the loss. \n\nSigh, I thought of it then forget about it.  Do you know how much you get from it?",
    "1333358": "I think it was our strongest boost on pure models out of all training related things, but we introduced it quite early so I can't fully say how much it contributed in the end with additional stuff.",
    "1333361": "Thank you. It was quite a significant impact of ~0.01 - 0.03 depending of where you count it (after post processing, pre-postprocessing). Also, impact on single models was always much larger than for the ensemble.",
    "1333375": "Congrats on 2nd place! @ilu000 @philippsinger @christofhenkel \nYou said luck, but it doesn't seem like luck. :)",
    "1333379": "I think in F-metric based competitions there is always a higher luck factor than in some others. You can only try to mitigate the impact of luck on your standing which is what we tried.",
    "1333384": "Pff, this almost makes me will to retrain my models.\n\nWell done. I mean well done to have used it, not to make me regret ;)",
    "1333508": "meh, hab einfach kein würfelglück",
    "1333575": "Thanks for sharing @philippsinger, your 30s training to 5s inference design is very smart, and use rating for weight loss is a great idea, I will retrain my model with  such great idea to see what is the difference.  \nAnd we also tried On-the-fly Logmel transformation with torchlibrosa, but I feel it is still slow, may I know why you choose torch audio for this, because it is faster or more numerically stable? Thanks.",
    "1333735": "Congratulation @christofhenkel on #2 competitions ranking.\n\nFirst time?",
    "1333780": "yes :D.",
    "1333809": "Congrat ! \n\nThe dream team strikes again !\n\n\n\n![https://i.ibb.co/5BWy3Fz/61.jpg](https://i.ibb.co/5BWy3Fz/61.jpg)",
    "1333972": "What background noise did you use?  freefield1010?",
    "1333991": "Congratulation @christofhenkel on the #2 competition ranking.👍",
    "1334025": "As always, you guys are awesome. Thanks for the details.",
    "1334063": "Partly, yes and some from this years data (see also text). Did you use any?",
    "1334064": "I used freefield1010 no call rows as in previous comp.  I really reused the same model as in Cornell competition.",
    "1334107": "Thank you very much for this insightful write-up. \nFrom the image I deduce that you do the 'horizontal' mix-up (30s -> 6x5s) on the spectrograms, not on the audio signals. Is this correct? If so, any particular reason why you didn't do it on the audio signal?",
    "1334120": "Huge congrats! :) Outstanding performance and a superb write-up! Thanks so much for sharing all this.\n\nCould I please ask you what did you use for smoothing neighboring predictions?",
    "1334136": "Thanks! Similar idea as blending 5sec with 10sec preds for example.",
    "1334486": "Yes all mixup is done on spectrograms. We didnt have time to try on raw audio in this comp, it might have also cpu bottlenecked us.",
    "1334506": "Can I know about the Mixup augmentation? Like can you share the code or any reference link?\nAnd how did you use it here?",
    "1334899": "It is just standard mixup, nothing special.",
    "1334924": "Thanks for sharing these methods! The validation scheme makes a lot of sense wish I had thought to do something similar...",
    "1335234": "We used torchlibrosa in our [1st place solution in RFCX competition](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563). But we chose torchaudio here because its more compatible with torch 1.8.1 and was numerically more stable when using mixed precision.",
    "1338971": "Thanks @christofhenkel , after our strongest model helped us to reach golden area, I tried write another model to increase diversity, but this model is always not stable, I used torchlibrosa and AMP as well, maybe this is the reason of the instability. Later I will make some experiment to compare torchaudio, torchlibrosa and nnAudio, in aspect of speed and numerical stability, thanks for your comments.",
    "1340733": "For label smoothing, could I please ask if you have ever encountered a situation where messing with the defaults of 0.9 to positive class, 0.1 distributed across all negatives, is helpful? For this competition, did you go with the default values?\n\nThanks a lot! 😊",
    "1340809": "0.01 to everything",
    "1341885": "We also did label smoothing only on one side, namely 0.01 across all negatives. Positive class is unchanged.",
    "1375972": "Thanks for sharing your solution and congratulations on the second place.\nI was interested in your team's idea of \"Reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone.\nI would like to ask you two questions, even though it has been a long time since the competition ended.\n\n1. There are several audios in train short audio that are less than 30s in length. How did you reshape such data into 6 x 5chunks? If you padded with zero, that part of the data would obviously lose the bird call, so it seems to be mislabeled.\n2. I thought the idea of mixup within the same audio was interesting. So I would like to know more about this. Did you apply the mixup to all 6 chuncks at the same time? Or did you pick 2 or 3 chunks out of the 6 and apply the mixup to them?",
    "1378474": "1. Padding with zeros worked pretty well, actually the whole goal of this setup is that the model will partly ignore the irrelevant chunks.\n\n2. Applied it to all at the same time, so each chunk might be mixed with any other chunk.",
    "1381037": "Thank you for your answer.\nI see, by pooling (especially if it is Max pooling), the model can partly ignore irrelevant chunks.\nI also understand about mixup. Thank you very much.",
    "1467989": "philippsinger \nafter a  long gap i read this solution. Recalling what i was doing corresponding to this step\n` To account for the 5sec snippet format of test data, we reshaped the 30sec crops into 6x 5sec parts before feeding through the backbone. After the backbone we reshaped the data again to re-arrange to the 30sec representation by concatenating the respective time segments and then used simple pooling of time and frequency dimension before forwarding through a simple one layer head which gave us the 398 bird classes. `\nI was doing like  this \nTake  N  clips\nthen do this inside fwd\n```\nn=len(x) #length of say 4 clips\nx = torch.stack(x,1).view(-1,shape[1],shape[2],shape[3])\nx=self.model(x)\nx = x.view(-1,n,shape[1],shape[2],shape[3]).permute(0,2,1,3,4).contiguous()\\\n            .view(-1,shape[1],shape[2]*n,shape[3])\n          #x: bs x C x n*(h x w) #feature dimensions for each channel\nx= nn.Sequential( AdaptiveConcatPool2d(),Flatten(),nn.BatchNorm1d(2*nc ) ,\n                                 nn.Linear(2*nc ,512), HardSwish() ,nn.Dropout(0.3),nn.Linear(512,n_classes)) ) (x)\n```\nis it what you also meant in your above statement of solution ?",
    "1555655": "Thanks for sharing your insights"
  },
  "source": "meta"
}