{
  "id": 220563,
  "title": "1st place solution",
  "url": "/competitions/rfcx-species-audio-detection/writeups/watercooled-1st-place-solution",
  "author_name": "",
  "post_date": "2021-02-19T13:18:55.350Z",
  "votes": 156,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. This has been as always a great collaborative effort and please also give your upvotes to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>. In the following, we want to give a rough overview of our winning solution.</p>\n<h3>TLDR</h3>\n<p>Our solution is an ensemble of several CNNs, which take a mel spectrogram representation of the recording as input and predict on recording level using “weak labels” or on a more granular time level using “hard labels”. Key in our modeling is masking as part of the loss function to only account for provided annotations. In order to account for the large amount of missing annotations and the inconsistent way how train and test data was labeled we apply a sophisticated scaling of model predictions.</p>\n<h3>Data setup &amp; CV</h3>\n<p>As most participants know, the training data was substantially differently labeled compared to the test data and the training labels were sparse. Hence, it was really tricky, nearly impossible to get a proper validation setup going. We tried quite a few things, such as treating all top 3 predicted labels as TPs when calculating the LWLRAP (because we know that on average a recording has 3 TPs), or calculating AUC only on segments where we know the labels (masked AUC), but in the end there was no good correlation that we could find to the public LB. This meant that we had to fully rely on public LB as feedback for choosing our models and submissions. Thankfully, it was a random split from the full test population, but everything else would not have made much sense anyways most likely.</p>\n<h3>Models</h3>\n<p>Its worthy to note that for most models we performed also the mel spec transformation and augmentations like mixup or coarse dropout on GPU using the implementation that can be found under torchlibrosa (<a href=\"https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py\" target=\"_blank\">https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py</a>).</p>\n<p>Our final models incorporate both hard and weak label models as explained next.</p>\n<h4>Hard label models</h4>\n<p>We refer to hard labels as labels that have hard time boundaries inside the recordings. Our hard label models were trained on the provided TPs (target = 1) and FPs (target = 0) labels with time aware loss evaluation. We used a log-spectrogram tensor of variable time length as input to an EfficientNet backbone and restricted the pooling layer to only mean pool over the frequency axis. After pooling, the output has 24 channels for each species and a time dimension. </p>\n<p>We then map the time axis from the model to the time labels from the TPs and FPs and evaluate the BCE loss <strong>only</strong> for the parts with provided labels. For all other segments (which is actually the majority) the loss is ignored, as we have no prior knowledge about the presence or absence of species there. In the figure below we show how a masked label looks like: yellow means target=1, green is target=0 and purple is ignored.<br>\n<img src=\"https://i.imgur.com/y9pBxlU.png\" alt=\"\"></p>\n<p>For some models we added hand labeled parts of the train set but saw diminishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP). Most notably, adding TP labels for species 18 gave a substantial boost to LB score, and we believe that adding some hand labels to the mix of models in the blend helped with diversity and generalization. </p>\n<p>For some models, similar to other top performing teams, we trained a second stage in which we replaced the masked part of the label with pseudo predictions of the first stage, but downweighted with factor 0.5. The main difference here to other teams is that we scaled the pseudo predictions in the same way we scale test predictions.</p>\n<p>As augmentation we used mixup with lambda=3, SpecAugment and gaussian noise.</p>\n<h4>Weak label models</h4>\n<p>The models in this part of the blend are based on weak label models. The input is the log-spectrogram of the full 60 seconds of an audio recording including all the labels for that clip. So it directly fits on the format where the final predictions need to be made. Due to missing labels, just fitting on the known TPs does not work too well as we incorporate wrong labels by nature. Also we cannot use the FPs, because even though an FP might be present in one part of the recording, does not mean there might not be a TP at another position.</p>\n<p>Hence, the models fit here include pseudo labels from our hard label models (see above) as well as some partial hand labels. For the pseudo labels, we take the raw output from the hard label models, but scale them to our expected true distribution (see post processing). For the hand labels, we only pick the TPs as well as FPs that span over a 60second period so that we are sure the species is not part of that recording. In loss, we weight the pseudos between 0.3-0.5 and the original labels and hand labels as 1.</p>\n<p>If we would just fit on the raw pseudo outputs, we would not learn anything new, so we employ concepts from noisy-student models. That means we utilize not only simple augmentations and mixup, but also randomly sample pseudo labels for each recording each time we train on it based on a pool of stage 1 hard label models. So for example, you fit 10 hard label models, and then randomly sample one each time in the dataloader. This introduces randomness and further boosts on top of the stage 1 models.</p>\n<p>Additionally, we fit several backbones (efnetb0, efnetb3, seresnext26, mobilenetv2_120d) where each is trained on the full data (no folds) with several seeds. In the end this part of the blend is a bag of around 120 models, where some also have additional TTA (horizontal flip).</p>\n<h3>How we are blending</h3>\n<p>We are blending different model types described above as depicted by the following graphic:<br>\n<img src=\"https://i.imgur.com/liB2Sic.png\" alt=\"\"></p>\n<h3>Post processing</h3>\n<p>We noticed that the test distribution of the target labels is substantially different to the provided train labels. Due to this fact, the models assume an unreasonable low or high probability when they are uncertain (Chris already has started a great thread about it <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">here</a>). To tackle this, we used several a priori information from the test distribution and scaled our predictions accordingly: by probing the public leaderboard we extracted a test label distribution which was aligning well with a previous research paper from the hosts. With additional prior knowledge about the average number of labels per row (3) -- also confirmed by LB probing, as well as the research paper -- we applied either a linear (species_probas *= factor) or a power scaling (species_probas **= factor) per species to our predicted probabilities to match the top3 predictions distribution (orange) with the previously mentioned estimated test distribution (blue). But we didn’t stop there, as we know that the number of labels per row is not always 3 but can be as low as 1 or as high as 8 (stated in the paper). Based on the sum of our probas in each row, we estimated the most likely topX (with a minimum count of 1) distribution (green) of the test set, and optimized the scaling factors by minimizing the total sum of the errors. <br>\n<img src=\"https://i.imgur.com/Fbi6eQn.png\" alt=\"\"></p>\n<h3>What did not work</h3>\n<p>I think in the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work, after all we ran close to 2,000 experiments throughout the course of this competition. One noteworthy thing we tried was object detection based on the bounding boxes we had available in training. It worked reasonably well on simple CV setting reaching &gt;0.7 LWLRAP on full 60 second recordings, but we never continued to work on it on smaller crops or other settings.</p>\n<p>We explored quite some architectures in the hope to improve our ensemble. So we tried models that work on the raw wave like Res1DNet or the just released wav2vec. But none did sufficiently well.</p>\n<p>Thanks for reading. Questions are very welcome. <br>\nChristof, Pascal &amp; Philipp</p>",
  "messages": [
    {
      "id": "1209306",
      "postDate": "02/18/2021 20:09:11",
      "content": "<p>Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. This has been as always a great collaborative effort and please also give your upvotes to <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>. In the following, we want to give a rough overview of our winning solution.</p>\n<h3>TLDR</h3>\n<p>Our solution is an ensemble of several CNNs, which take a mel spectrogram representation of the recording as input and predict on recording level using “weak labels” or on a more granular time level using “hard labels”. Key in our modeling is masking as part of the loss function to only account for provided annotations. In order to account for the large amount of missing annotations and the inconsistent way how train and test data was labeled we apply a sophisticated scaling of model predictions.</p>\n<h3>Data setup &amp; CV</h3>\n<p>As most participants know, the training data was substantially differently labeled compared to the test data and the training labels were sparse. Hence, it was really tricky, nearly impossible to get a proper validation setup going. We tried quite a few things, such as treating all top 3 predicted labels as TPs when calculating the LWLRAP (because we know that on average a recording has 3 TPs), or calculating AUC only on segments where we know the labels (masked AUC), but in the end there was no good correlation that we could find to the public LB. This meant that we had to fully rely on public LB as feedback for choosing our models and submissions. Thankfully, it was a random split from the full test population, but everything else would not have made much sense anyways most likely.</p>\n<h3>Models</h3>\n<p>Its worthy to note that for most models we performed also the mel spec transformation and augmentations like mixup or coarse dropout on GPU using the implementation that can be found under torchlibrosa (<a href=\"https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py\" target=\"_blank\">https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py</a>).</p>\n<p>Our final models incorporate both hard and weak label models as explained next.</p>\n<h4>Hard label models</h4>\n<p>We refer to hard labels as labels that have hard time boundaries inside the recordings. Our hard label models were trained on the provided TPs (target = 1) and FPs (target = 0) labels with time aware loss evaluation. We used a log-spectrogram tensor of variable time length as input to an EfficientNet backbone and restricted the pooling layer to only mean pool over the frequency axis. After pooling, the output has 24 channels for each species and a time dimension. </p>\n<p>We then map the time axis from the model to the time labels from the TPs and FPs and evaluate the BCE loss <strong>only</strong> for the parts with provided labels. For all other segments (which is actually the majority) the loss is ignored, as we have no prior knowledge about the presence or absence of species there. In the figure below we show how a masked label looks like: yellow means target=1, green is target=0 and purple is ignored.<br>\n<img src=\"https://i.imgur.com/y9pBxlU.png\" alt=\"\"></p>\n<p>For some models we added hand labeled parts of the train set but saw diminishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP). Most notably, adding TP labels for species 18 gave a substantial boost to LB score, and we believe that adding some hand labels to the mix of models in the blend helped with diversity and generalization. </p>\n<p>For some models, similar to other top performing teams, we trained a second stage in which we replaced the masked part of the label with pseudo predictions of the first stage, but downweighted with factor 0.5. The main difference here to other teams is that we scaled the pseudo predictions in the same way we scale test predictions.</p>\n<p>As augmentation we used mixup with lambda=3, SpecAugment and gaussian noise.</p>\n<h4>Weak label models</h4>\n<p>The models in this part of the blend are based on weak label models. The input is the log-spectrogram of the full 60 seconds of an audio recording including all the labels for that clip. So it directly fits on the format where the final predictions need to be made. Due to missing labels, just fitting on the known TPs does not work too well as we incorporate wrong labels by nature. Also we cannot use the FPs, because even though an FP might be present in one part of the recording, does not mean there might not be a TP at another position.</p>\n<p>Hence, the models fit here include pseudo labels from our hard label models (see above) as well as some partial hand labels. For the pseudo labels, we take the raw output from the hard label models, but scale them to our expected true distribution (see post processing). For the hand labels, we only pick the TPs as well as FPs that span over a 60second period so that we are sure the species is not part of that recording. In loss, we weight the pseudos between 0.3-0.5 and the original labels and hand labels as 1.</p>\n<p>If we would just fit on the raw pseudo outputs, we would not learn anything new, so we employ concepts from noisy-student models. That means we utilize not only simple augmentations and mixup, but also randomly sample pseudo labels for each recording each time we train on it based on a pool of stage 1 hard label models. So for example, you fit 10 hard label models, and then randomly sample one each time in the dataloader. This introduces randomness and further boosts on top of the stage 1 models.</p>\n<p>Additionally, we fit several backbones (efnetb0, efnetb3, seresnext26, mobilenetv2_120d) where each is trained on the full data (no folds) with several seeds. In the end this part of the blend is a bag of around 120 models, where some also have additional TTA (horizontal flip).</p>\n<h3>How we are blending</h3>\n<p>We are blending different model types described above as depicted by the following graphic:<br>\n<img src=\"https://i.imgur.com/liB2Sic.png\" alt=\"\"></p>\n<h3>Post processing</h3>\n<p>We noticed that the test distribution of the target labels is substantially different to the provided train labels. Due to this fact, the models assume an unreasonable low or high probability when they are uncertain (Chris already has started a great thread about it <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">here</a>). To tackle this, we used several a priori information from the test distribution and scaled our predictions accordingly: by probing the public leaderboard we extracted a test label distribution which was aligning well with a previous research paper from the hosts. With additional prior knowledge about the average number of labels per row (3) -- also confirmed by LB probing, as well as the research paper -- we applied either a linear (species_probas *= factor) or a power scaling (species_probas **= factor) per species to our predicted probabilities to match the top3 predictions distribution (orange) with the previously mentioned estimated test distribution (blue). But we didn’t stop there, as we know that the number of labels per row is not always 3 but can be as low as 1 or as high as 8 (stated in the paper). Based on the sum of our probas in each row, we estimated the most likely topX (with a minimum count of 1) distribution (green) of the test set, and optimized the scaling factors by minimizing the total sum of the errors. <br>\n<img src=\"https://i.imgur.com/Fbi6eQn.png\" alt=\"\"></p>\n<h3>What did not work</h3>\n<p>I think in the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work, after all we ran close to 2,000 experiments throughout the course of this competition. One noteworthy thing we tried was object detection based on the bounding boxes we had available in training. It worked reasonably well on simple CV setting reaching &gt;0.7 LWLRAP on full 60 second recordings, but we never continued to work on it on smaller crops or other settings.</p>\n<p>We explored quite some architectures in the hope to improve our ensemble. So we tried models that work on the raw wave like Res1DNet or the just released wav2vec. But none did sufficiently well.</p>\n<p>Thanks for reading. Questions are very welcome. <br>\nChristof, Pascal &amp; Philipp</p>",
      "rawMarkdown": "Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. This has been as always a great collaborative effort and please also give your upvotes to @christofhenkel and @ilu000. In the following, we want to give a rough overview of our winning solution.\n\n### TLDR\nOur solution is an ensemble of several CNNs, which take a mel spectrogram representation of the recording as input and predict on recording level using “weak labels” or on a more granular time level using “hard labels”. Key in our modeling is masking as part of the loss function to only account for provided annotations. In order to account for the large amount of missing annotations and the inconsistent way how train and test data was labeled we apply a sophisticated scaling of model predictions.\n\n### Data setup & CV\nAs most participants know, the training data was substantially differently labeled compared to the test data and the training labels were sparse. Hence, it was really tricky, nearly impossible to get a proper validation setup going. We tried quite a few things, such as treating all top 3 predicted labels as TPs when calculating the LWLRAP (because we know that on average a recording has 3 TPs), or calculating AUC only on segments where we know the labels (masked AUC), but in the end there was no good correlation that we could find to the public LB. This meant that we had to fully rely on public LB as feedback for choosing our models and submissions. Thankfully, it was a random split from the full test population, but everything else would not have made much sense anyways most likely.\n\n### Models\nIts worthy to note that for most models we performed also the mel spec transformation and augmentations like mixup or coarse dropout on GPU using the implementation that can be found under torchlibrosa ([https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py](https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py)).\n\nOur final models incorporate both hard and weak label models as explained next.\n\n#### Hard label models\nWe refer to hard labels as labels that have hard time boundaries inside the recordings. Our hard label models were trained on the provided TPs (target = 1) and FPs (target = 0) labels with time aware loss evaluation. We used a log-spectrogram tensor of variable time length as input to an EfficientNet backbone and restricted the pooling layer to only mean pool over the frequency axis. After pooling, the output has 24 channels for each species and a time dimension. \n\nWe then map the time axis from the model to the time labels from the TPs and FPs and evaluate the BCE loss **only** for the parts with provided labels. For all other segments (which is actually the majority) the loss is ignored, as we have no prior knowledge about the presence or absence of species there. In the figure below we show how a masked label looks like: yellow means target=1, green is target=0 and purple is ignored.\n![](https://i.imgur.com/y9pBxlU.png)\n\nFor some models we added hand labeled parts of the train set but saw diminishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP). Most notably, adding TP labels for species 18 gave a substantial boost to LB score, and we believe that adding some hand labels to the mix of models in the blend helped with diversity and generalization. \n\nFor some models, similar to other top performing teams, we trained a second stage in which we replaced the masked part of the label with pseudo predictions of the first stage, but downweighted with factor 0.5. The main difference here to other teams is that we scaled the pseudo predictions in the same way we scale test predictions.\n\nAs augmentation we used mixup with lambda=3, SpecAugment and gaussian noise.\n\n#### Weak label models\nThe models in this part of the blend are based on weak label models. The input is the log-spectrogram of the full 60 seconds of an audio recording including all the labels for that clip. So it directly fits on the format where the final predictions need to be made. Due to missing labels, just fitting on the known TPs does not work too well as we incorporate wrong labels by nature. Also we cannot use the FPs, because even though an FP might be present in one part of the recording, does not mean there might not be a TP at another position.\n\nHence, the models fit here include pseudo labels from our hard label models (see above) as well as some partial hand labels. For the pseudo labels, we take the raw output from the hard label models, but scale them to our expected true distribution (see post processing). For the hand labels, we only pick the TPs as well as FPs that span over a 60second period so that we are sure the species is not part of that recording. In loss, we weight the pseudos between 0.3-0.5 and the original labels and hand labels as 1.\n\nIf we would just fit on the raw pseudo outputs, we would not learn anything new, so we employ concepts from noisy-student models. That means we utilize not only simple augmentations and mixup, but also randomly sample pseudo labels for each recording each time we train on it based on a pool of stage 1 hard label models. So for example, you fit 10 hard label models, and then randomly sample one each time in the dataloader. This introduces randomness and further boosts on top of the stage 1 models.\n\nAdditionally, we fit several backbones (efnetb0, efnetb3, seresnext26, mobilenetv2_120d) where each is trained on the full data (no folds) with several seeds. In the end this part of the blend is a bag of around 120 models, where some also have additional TTA (horizontal flip).\n\n### How we are blending\nWe are blending different model types described above as depicted by the following graphic:\n![](https://i.imgur.com/liB2Sic.png)\n\n### Post processing\nWe noticed that the test distribution of the target labels is substantially different to the provided train labels. Due to this fact, the models assume an unreasonable low or high probability when they are uncertain (Chris already has started a great thread about it [here] (https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389)). To tackle this, we used several a priori information from the test distribution and scaled our predictions accordingly: by probing the public leaderboard we extracted a test label distribution which was aligning well with a previous research paper from the hosts. With additional prior knowledge about the average number of labels per row (3) -- also confirmed by LB probing, as well as the research paper -- we applied either a linear (species_probas *= factor) or a power scaling (species_probas **= factor) per species to our predicted probabilities to match the top3 predictions distribution (orange) with the previously mentioned estimated test distribution (blue). But we didn’t stop there, as we know that the number of labels per row is not always 3 but can be as low as 1 or as high as 8 (stated in the paper). Based on the sum of our probas in each row, we estimated the most likely topX (with a minimum count of 1) distribution (green) of the test set, and optimized the scaling factors by minimizing the total sum of the errors. \n![](https://i.imgur.com/Fbi6eQn.png)\n\n### What did not work\nI think in the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work, after all we ran close to 2,000 experiments throughout the course of this competition. One noteworthy thing we tried was object detection based on the bounding boxes we had available in training. It worked reasonably well on simple CV setting reaching >0.7 LWLRAP on full 60 second recordings, but we never continued to work on it on smaller crops or other settings.\n\nWe explored quite some architectures in the hope to improve our ensemble. So we tried models that work on the raw wave like Res1DNet or the just released wav2vec. But none did sufficiently well.\n\nThanks for reading. Questions are very welcome. \nChristof, Pascal & Philipp",
      "votes": null
    },
    {
      "id": "1209310",
      "postDate": "02/18/2021 20:11:56",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for another great teamup!</p>",
      "rawMarkdown": "Thanks @philippsinger and @christofhenkel for another great teamup!",
      "votes": null
    },
    {
      "id": "1209315",
      "postDate": "02/18/2021 20:16:28",
      "content": "<blockquote>\n  <p>species 2 had some calls in the background of several recordings, but the parts were labeled as FP</p>\n</blockquote>\n<p>Congrats! I also noticed this, I could hear it behind the foreground species but many of these examples were false positives, I tested both ignoring this false positive and treating it as actual false positives and in the end, ignoring it(setting it to be positive) got me better cv and lb results.</p>",
      "rawMarkdown": "> species 2 had some calls in the background of several recordings, but the parts were labeled as FP\n\n\nCongrats! I also noticed this, I could hear it behind the foreground species but many of these examples were false positives, I tested both ignoring this false positive and treating it as actual false positives and in the end, ignoring it(setting it to be positive) got me better cv and lb results.",
      "votes": null
    },
    {
      "id": "1209316",
      "postDate": "02/18/2021 20:17:31",
      "content": "<p>Great work, congrats on yet another win! Masked loss was a key ingredient indeed.  It is interesting you looked at object detection.  I though of it too but it was low on my to do list.</p>",
      "rawMarkdown": "Great work, congrats on yet another win! Masked loss was a key ingredient indeed.  It is interesting you looked at object detection.  I though of it too but it was low on my to do list.",
      "votes": null
    },
    {
      "id": "1209317",
      "postDate": "02/18/2021 20:18:25",
      "content": "<p>Thanks! We spent only little time on it, maybe we should have explored it further. I think it has some potential, maybe also in combination with masking or something similar.</p>",
      "rawMarkdown": "Thanks! We spent only little time on it, maybe we should have explored it further. I think it has some potential, maybe also in combination with masking or something similar.",
      "votes": null
    },
    {
      "id": "1209319",
      "postDate": "02/18/2021 20:21:38",
      "content": "<p>Very interesting way of using the pooling and adjusting the labels to correspond with the spectrogram.</p>\n<p>You mention CNN models. ? I see this is mentioned later. Did you find backbone changed performance very much? </p>",
      "rawMarkdown": "Very interesting way of using the pooling and adjusting the labels to correspond with the spectrogram.\n\nYou mention CNN models. ~~Presumably you started with imagenet pretrained weights and then just modified in the mean pooling and linear layer~~? I see this is mentioned later. Did you find backbone changed performance very much?",
      "votes": null
    },
    {
      "id": "1209321",
      "postDate": "02/18/2021 20:28:53",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> for this great insight and huge win!<br>\nThanks for sharing the approach.  <br>\nDid you guys tested LB score without testset distribution adjustment (PP) ?</p>\n<p>Unfortunately one more competition with results heavily based in LB probing and external data leakage.</p>",
      "rawMarkdown": "Congrats @philippsinger @christofhenkel and @ilu000 for this great insight and huge win!\nThanks for sharing the approach.  \nDid you guys tested LB score without testset distribution adjustment (PP) ?\n\nUnfortunately one more competition with results heavily based in LB probing and external data leakage.",
      "votes": null
    },
    {
      "id": "1209329",
      "postDate": "02/18/2021 20:34:21",
      "content": "<p>Yeah, performance was significantly different for different backbones. B0 worked best within the efficientnet family as larger versions overfit quickly. We added a seresnext and mobilenet rather for diversity than for their performance.</p>",
      "rawMarkdown": "Yeah, performance was significantly different for different backbones. B0 worked best within the efficientnet family as larger versions overfit quickly. We added a seresnext and mobilenet rather for diversity than for their performance.",
      "votes": null
    },
    {
      "id": "1209333",
      "postDate": "02/18/2021 20:39:32",
      "content": "<p>I agree, specifically that the paper exists is a bit weird, not the first time for research competitions that this happens.</p>\n<p>Without the PP is not really possible for us as we already incorporate our pseudos this way and the final dist is already biased towards that. I think in the end the metric needs some form of scaling. For example, as the data contains 90% S3 labels, if you do not predict these high enough, then the metric is hurt a lot. But the scaling can be achieved via different things. For example if you hand-label all the data as some did then you automatically move towards the test distribution as the populations are roughly similar. I think ranking loss maybe has some potential, but we did not find time to explore it.</p>\n<p>After all, I see the public dataset as a validation set here. And the validation set is a fair sample from the test set. So naturally you will try to fit the validation set better, which includes properly moving the TPs to the top.</p>",
      "rawMarkdown": "I agree, specifically that the paper exists is a bit weird, not the first time for research competitions that this happens.\n\nWithout the PP is not really possible for us as we already incorporate our pseudos this way and the final dist is already biased towards that. I think in the end the metric needs some form of scaling. For example, as the data contains 90% S3 labels, if you do not predict these high enough, then the metric is hurt a lot. But the scaling can be achieved via different things. For example if you hand-label all the data as some did then you automatically move towards the test distribution as the populations are roughly similar. I think ranking loss maybe has some potential, but we did not find time to explore it.\n\nAfter all, I see the public dataset as a validation set here. And the validation set is a fair sample from the test set. So naturally you will try to fit the validation set better, which includes properly moving the TPs to the top.",
      "votes": null
    },
    {
      "id": "1209341",
      "postDate": "02/18/2021 20:54:45",
      "content": "<p>Yeah we also used heuristics to increase the more frequent classes. Pseudo labeling, mean max blending, handlabeling moved all to that direction.<br>\nWe did not use LB probing this time because of lack of submissions and we were afraid of overfitting…</p>\n<p>It was surprising that even further scaling coulld boost our scores….</p>",
      "rawMarkdown": "Yeah we also used heuristics to increase the more frequent classes. Pseudo labeling, mean max blending, handlabeling moved all to that direction.\nWe did not use LB probing this time because of lack of submissions and we were afraid of overfitting...\n\nIt was surprising that even further scaling coulld boost our scores....",
      "votes": null
    },
    {
      "id": "1209346",
      "postDate": "02/18/2021 21:01:02",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! New Kaggle #1!</p>",
      "rawMarkdown": "Congrats @philippsinger! New Kaggle #1!",
      "votes": null
    },
    {
      "id": "1209354",
      "postDate": "02/18/2021 21:05:54",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, an amazing feat and a well deserved achievement for your hard work</p>",
      "rawMarkdown": "Congratulations @philippsinger, an amazing feat and a well deserved achievement for your hard work",
      "votes": null
    },
    {
      "id": "1209361",
      "postDate": "02/18/2021 21:09:52",
      "content": "<p>Just adjusting my best model distribution according to the distribution of classes in table 2 of the <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">paper</a> boosted to 0.961 Private :P</p>",
      "rawMarkdown": "Just adjusting my best model distribution according to the distribution of classes in table 2 of the [paper](https://www.sciencedirect.com/science/article/pii/S1574954120300637) boosted to 0.961 Private :P",
      "votes": null
    },
    {
      "id": "1209362",
      "postDate": "02/18/2021 21:10:34",
      "content": "<p>Thanks to my amazing teammates and congrats on top 3 and top 10! <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> </p>",
      "rawMarkdown": "Thanks to my amazing teammates and congrats on top 3 and top 10! @christofhenkel @ilu000",
      "votes": null
    },
    {
      "id": "1209398",
      "postDate": "02/18/2021 21:36:41",
      "content": "<p>new kaggle #1💯</p>",
      "rawMarkdown": "new kaggle #1💯",
      "votes": null
    },
    {
      "id": "1209404",
      "postDate": "02/18/2021 21:44:15",
      "content": "<p>Thanks Guanshuo! </p>",
      "rawMarkdown": "Thanks Guanshuo!",
      "votes": null
    },
    {
      "id": "1209417",
      "postDate": "02/18/2021 22:00:07",
      "content": "<p>Congratulation. I am just curious. 120 models sounds a lot of computing power. How many GPUs has your team? :)</p>",
      "rawMarkdown": "Congratulation. I am just curious. 120 models sounds a lot of computing power. How many GPUs has your team? :)",
      "votes": null
    },
    {
      "id": "1209428",
      "postDate": "02/18/2021 22:12:17",
      "content": "<p>The weak label models fitted really quickly, the data is quite small. I fitted ~40 in 24 hours on 3 GPUs.</p>",
      "rawMarkdown": "The weak label models fitted really quickly, the data is quite small. I fitted ~40 in 24 hours on 3 GPUs.",
      "votes": null
    },
    {
      "id": "1209436",
      "postDate": "02/18/2021 22:27:07",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> on getting #1, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on getting #3 and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> on getting #10 !</p>\n<p>Another impressive performance, it's not even surprising anymore. </p>\n<p>Any idea how much the post-processing helps your score ? </p>",
      "rawMarkdown": "Congratz @philippsinger on getting #1, @christofhenkel on getting #3 and @ilu000 on getting #10 !\n\nAnother impressive performance, it's not even surprising anymore. \n\nAny idea how much the post-processing helps your score ?",
      "votes": null
    },
    {
      "id": "1209440",
      "postDate": "02/18/2021 22:37:25",
      "content": "<p>Thanks Theo,</p>\n<p>without post processing, the score would be significantly lower. As stated here and also in other threads, there are a few ways to tackle the class imbalance. One is post processing the distributions to match the expected test distibution, others include e.g. adding labels for the underrepresented species as some other teams did. We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution. <br>\nAlso note, that the train label distribution, if it would include all labels, should also be quite close to the test label distribution. </p>",
      "rawMarkdown": "Thanks Theo,\n\nwithout post processing, the score would be significantly lower. As stated here and also in other threads, there are a few ways to tackle the class imbalance. One is post processing the distributions to match the expected test distibution, others include e.g. adding labels for the underrepresented species as some other teams did. We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution. \nAlso note, that the train label distribution, if it would include all labels, should also be quite close to the test label distribution.",
      "votes": null
    },
    {
      "id": "1209441",
      "postDate": "02/18/2021 22:41:02",
      "content": "<p>Thanks for the reply ! </p>\n<blockquote>\n  <p>We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.</p>\n</blockquote>\n<p>Agreed, the metric is indeed highly distribution dependent so capturing had a lot of importance. </p>",
      "rawMarkdown": "Thanks for the reply ! \n\n>  We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.\n\nAgreed, the metric is indeed highly distribution dependent so capturing had a lot of importance.",
      "votes": null
    },
    {
      "id": "1209462",
      "postDate": "02/18/2021 23:05:52",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> on winners! </p>",
      "rawMarkdown": "Congrats @philippsinger @christofhenkel and @ilu000 on winners!",
      "votes": null
    },
    {
      "id": "1209736",
      "postDate": "02/19/2021 02:57:17",
      "content": "<p>Congrats to the team and congrats on becoming no. 1 in kaggle grandmasters, your solutions are always interesting and a source of learning. Thanks </p>",
      "rawMarkdown": "Congrats to the team and congrats on becoming no. 1 in kaggle grandmasters, your solutions are always interesting and a source of learning. Thanks",
      "votes": null
    },
    {
      "id": "1209743",
      "postDate": "02/19/2021 03:00:32",
      "content": "<p>Congrats team. Another amazing performance with creative and brilliant solution. Well done. </p>\n<p>Congrats Psi on becoming Kaggle's #1 ranked, Dieter on becoming #3, and Ilu on becoming #10. Your solutions and hard work are an inspiration to everyone!</p>",
      "rawMarkdown": "Congrats team. Another amazing performance with creative and brilliant solution. Well done. \n\nCongrats Psi on becoming Kaggle's #1 ranked, Dieter on becoming #3, and Ilu on becoming #10. Your solutions and hard work are an inspiration to everyone!",
      "votes": null
    },
    {
      "id": "1209786",
      "postDate": "02/19/2021 03:33:36",
      "content": "<p>someone also achieved 25th, Congrats!!</p>",
      "rawMarkdown": "someone also achieved 25th, Congrats!!",
      "votes": null
    },
    {
      "id": "1210021",
      "postDate": "02/19/2021 06:52:35",
      "content": "<p>Congrats and thanks for sharing. </p>",
      "rawMarkdown": "Congrats and thanks for sharing.",
      "votes": null
    },
    {
      "id": "1210029",
      "postDate": "02/19/2021 06:56:18",
      "content": "<p>Congratulations on winning this competition and thanks for the write up!<br>\nPseudo training and manaually debiasing the prediction seems to be key part for this competition. Also your approach for the problem with concept of hard/weak labels is great.<br>\nI see your team ensembled various models. If you have, can you share the score of your best single model on lb?</p>",
      "rawMarkdown": "Congratulations on winning this competition and thanks for the write up!\nPseudo training and manaually debiasing the prediction seems to be key part for this competition. Also your approach for the problem with concept of hard/weak labels is great.\nI see your team ensembled various models. If you have, can you share the score of your best single model on lb?",
      "votes": null
    },
    {
      "id": "1210475",
      "postDate": "02/19/2021 13:15:49",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>,  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>. <br>\nThank you for sharing the write-up as well.</p>\n<p>The link for the librosa is the following: <a href=\"https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py\" target=\"_blank\">https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py</a> (when one presses the link it is picking up the final parenthesis and redirecting to a 404 page). </p>",
      "rawMarkdown": "Congrats @philippsinger,  @christofhenkel and @ilu000. \nThank you for sharing the write-up as well.\n\nThe link for the librosa is the following: https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py (when one presses the link it is picking up the final parenthesis and redirecting to a 404 page).",
      "votes": null
    },
    {
      "id": "1210797",
      "postDate": "02/19/2021 17:36:09",
      "content": "<p>Thanks For Sharing.</p>",
      "rawMarkdown": "Thanks For Sharing.",
      "votes": null
    },
    {
      "id": "1210809",
      "postDate": "02/19/2021 17:51:55",
      "content": "<p>I also now see why my post on the number of species was relevant.  Your blending and your postprocesisng to match target distribution is great.</p>",
      "rawMarkdown": "I also now see why my post on the number of species was relevant.  Your blending and your postprocesisng to match target distribution is great.",
      "votes": null
    },
    {
      "id": "1210935",
      "postDate": "02/19/2021 20:05:02",
      "content": "<p>Yep, I think it is 3 on average.</p>",
      "rawMarkdown": "Yep, I think it is 3 on average.",
      "votes": null
    },
    {
      "id": "1212791",
      "postDate": "02/21/2021 15:45:50",
      "content": "<p>Congratulations team! </p>\n<p>Also congrats on becoming #1 Psi, Dieter on becoming #3, Ilu on becoming #10. Very inspirational 🙌</p>",
      "rawMarkdown": "Congratulations team! \n\nAlso congrats on becoming #1 Psi, Dieter on becoming #3, Ilu on becoming #10. Very inspirational 🙌",
      "votes": null
    },
    {
      "id": "1212794",
      "postDate": "02/21/2021 15:50:24",
      "content": "<blockquote>\n  <p>We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.</p>\n</blockquote>\n<p>I wonder if I have the highest score that does not rely on anything like that, be it post processing, or pseudo labeling.</p>\n<p>Post deadline experiments show that both pseudo labeling and post processing increase my score substantially.</p>",
      "rawMarkdown": "> We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.\n\nI wonder if I have the highest score that does not rely on anything like that, be it post processing, or pseudo labeling.\n\nPost deadline experiments show that both pseudo labeling and post processing increase my score substantially.",
      "votes": null
    },
    {
      "id": "1218409",
      "postDate": "02/25/2021 20:42:21",
      "content": "<p>It is quite impressive to see how much work you put into it and how far you went to try to determine the distribution of the test set to improve your predictions. </p>\n<p>What I am wondering is what type of toolset you used individually and as a team to do this. For instance, do you use tools for hyperparameter tuning and experiment management? Do you use ML flow, ray tune etc.? Also, I assume you have developed your own tooling to some extent to support experimentation with such a large number of models. Just some insight in what you use as a team to do all these experiments and work together would be very interesting. </p>",
      "rawMarkdown": "It is quite impressive to see how much work you put into it and how far you went to try to determine the distribution of the test set to improve your predictions. \n\nWhat I am wondering is what type of toolset you used individually and as a team to do this. For instance, do you use tools for hyperparameter tuning and experiment management? Do you use ML flow, ray tune etc.? Also, I assume you have developed your own tooling to some extent to support experimentation with such a large number of models. Just some insight in what you use as a team to do all these experiments and work together would be very interesting.",
      "votes": null
    },
    {
      "id": "1218667",
      "postDate": "02/26/2021 05:08:43",
      "content": "<p>We use plain pytorch, with a self-written training routine, which supports mixed precision and effective multi GPU use (DDP). We use GitHub to have a shared code-base and neptune.ai to track and share experiments.  </p>",
      "rawMarkdown": "We use plain pytorch, with a self-written training routine, which supports mixed precision and effective multi GPU use (DDP). We use GitHub to have a shared code-base and neptune.ai to track and share experiments.",
      "votes": null
    },
    {
      "id": "1335046",
      "postDate": "06/04/2021 01:42:47",
      "content": "<p>Congratulations team!<br>\nCan you provide a link to github？</p>",
      "rawMarkdown": "Congratulations team!\nCan you provide a link to github？",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1209310,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "02/18/2021 20:11:56",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> for another great teamup!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209315,
      "author_name": "dicksonchin93",
      "author_url": "",
      "post_date": "02/18/2021 20:16:28",
      "content": "<blockquote>\n  <p>species 2 had some calls in the background of several recordings, but the parts were labeled as FP</p>\n</blockquote>\n<p>Congrats! I also noticed this, I could hear it behind the foreground species but many of these examples were false positives, I tested both ignoring this false positive and treating it as actual false positives and in the end, ignoring it(setting it to be positive) got me better cv and lb results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209316,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 20:17:31",
      "content": "<p>Great work, congrats on yet another win! Masked loss was a key ingredient indeed.  It is interesting you looked at object detection.  I though of it too but it was low on my to do list.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209317,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 20:18:25",
          "content": "<p>Thanks! We spent only little time on it, maybe we should have explored it further. I think it has some potential, maybe also in combination with masking or something similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210809,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/19/2021 17:51:55",
          "content": "<p>I also now see why my post on the number of species was relevant.  Your blending and your postprocesisng to match target distribution is great.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210935,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/19/2021 20:05:02",
          "content": "<p>Yep, I think it is 3 on average.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209319,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/18/2021 20:21:38",
      "content": "<p>Very interesting way of using the pooling and adjusting the labels to correspond with the spectrogram.</p>\n<p>You mention CNN models. ? I see this is mentioned later. Did you find backbone changed performance very much? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1209329,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/18/2021 20:34:21",
          "content": "<p>Yeah, performance was significantly different for different backbones. B0 worked best within the efficientnet family as larger versions overfit quickly. We added a seresnext and mobilenet rather for diversity than for their performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209321,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "02/18/2021 20:28:53",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> for this great insight and huge win!<br>\nThanks for sharing the approach.  <br>\nDid you guys tested LB score without testset distribution adjustment (PP) ?</p>\n<p>Unfortunately one more competition with results heavily based in LB probing and external data leakage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209333,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 20:39:32",
          "content": "<p>I agree, specifically that the paper exists is a bit weird, not the first time for research competitions that this happens.</p>\n<p>Without the PP is not really possible for us as we already incorporate our pseudos this way and the final dist is already biased towards that. I think in the end the metric needs some form of scaling. For example, as the data contains 90% S3 labels, if you do not predict these high enough, then the metric is hurt a lot. But the scaling can be achieved via different things. For example if you hand-label all the data as some did then you automatically move towards the test distribution as the populations are roughly similar. I think ranking loss maybe has some potential, but we did not find time to explore it.</p>\n<p>After all, I see the public dataset as a validation set here. And the validation set is a fair sample from the test set. So naturally you will try to fit the validation set better, which includes properly moving the TPs to the top.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209341,
          "author_name": "gaborfodor",
          "author_url": "",
          "post_date": "02/18/2021 20:54:45",
          "content": "<p>Yeah we also used heuristics to increase the more frequent classes. Pseudo labeling, mean max blending, handlabeling moved all to that direction.<br>\nWe did not use LB probing this time because of lack of submissions and we were afraid of overfitting…</p>\n<p>It was surprising that even further scaling coulld boost our scores….</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209361,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "02/18/2021 21:09:52",
          "content": "<p>Just adjusting my best model distribution according to the distribution of classes in table 2 of the <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">paper</a> boosted to 0.961 Private :P</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209346,
      "author_name": "docxian",
      "author_url": "",
      "post_date": "02/18/2021 21:01:02",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! New Kaggle #1!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209354,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "02/18/2021 21:05:54",
          "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, an amazing feat and a well deserved achievement for your hard work</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209362,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 21:10:34",
          "content": "<p>Thanks to my amazing teammates and congrats on top 3 and top 10! <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209398,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "02/18/2021 21:36:41",
      "content": "<p>new kaggle #1💯</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209404,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 21:44:15",
          "content": "<p>Thanks Guanshuo! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209417,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "02/18/2021 22:00:07",
      "content": "<p>Congratulation. I am just curious. 120 models sounds a lot of computing power. How many GPUs has your team? :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209428,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2021 22:12:17",
          "content": "<p>The weak label models fitted really quickly, the data is quite small. I fitted ~40 in 24 hours on 3 GPUs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209436,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 22:27:07",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> on getting #1, <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> on getting #3 and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> on getting #10 !</p>\n<p>Another impressive performance, it's not even surprising anymore. </p>\n<p>Any idea how much the post-processing helps your score ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1209440,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "02/18/2021 22:37:25",
          "content": "<p>Thanks Theo,</p>\n<p>without post processing, the score would be significantly lower. As stated here and also in other threads, there are a few ways to tackle the class imbalance. One is post processing the distributions to match the expected test distibution, others include e.g. adding labels for the underrepresented species as some other teams did. We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution. <br>\nAlso note, that the train label distribution, if it would include all labels, should also be quite close to the test label distribution. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209441,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/18/2021 22:41:02",
          "content": "<p>Thanks for the reply ! </p>\n<blockquote>\n  <p>We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.</p>\n</blockquote>\n<p>Agreed, the metric is indeed highly distribution dependent so capturing had a lot of importance. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212794,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/21/2021 15:50:24",
          "content": "<blockquote>\n  <p>We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.</p>\n</blockquote>\n<p>I wonder if I have the highest score that does not rely on anything like that, be it post processing, or pseudo labeling.</p>\n<p>Post deadline experiments show that both pseudo labeling and post processing increase my score substantially.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209462,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 23:05:52",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> on winners! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209736,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "02/19/2021 02:57:17",
      "content": "<p>Congrats to the team and congrats on becoming no. 1 in kaggle grandmasters, your solutions are always interesting and a source of learning. Thanks </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209743,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/19/2021 03:00:32",
      "content": "<p>Congrats team. Another amazing performance with creative and brilliant solution. Well done. </p>\n<p>Congrats Psi on becoming Kaggle's #1 ranked, Dieter on becoming #3, and Ilu on becoming #10. Your solutions and hard work are an inspiration to everyone!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209786,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/19/2021 03:33:36",
          "content": "<p>someone also achieved 25th, Congrats!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210021,
      "author_name": "harip98",
      "author_url": "",
      "post_date": "02/19/2021 06:52:35",
      "content": "<p>Congrats and thanks for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210029,
      "author_name": "harangdev",
      "author_url": "",
      "post_date": "02/19/2021 06:56:18",
      "content": "<p>Congratulations on winning this competition and thanks for the write up!<br>\nPseudo training and manaually debiasing the prediction seems to be key part for this competition. Also your approach for the problem with concept of hard/weak labels is great.<br>\nI see your team ensembled various models. If you have, can you share the score of your best single model on lb?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210475,
      "author_name": "goncaloperes",
      "author_url": "",
      "post_date": "02/19/2021 13:15:49",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>,  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>. <br>\nThank you for sharing the write-up as well.</p>\n<p>The link for the librosa is the following: <a href=\"https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py\" target=\"_blank\">https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py</a> (when one presses the link it is picking up the final parenthesis and redirecting to a 404 page). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210797,
      "author_name": "himanshu1999",
      "author_url": "",
      "post_date": "02/19/2021 17:36:09",
      "content": "<p>Thanks For Sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1212791,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/21/2021 15:45:50",
      "content": "<p>Congratulations team! </p>\n<p>Also congrats on becoming #1 Psi, Dieter on becoming #3, Ilu on becoming #10. Very inspirational 🙌</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1218409,
      "author_name": "erikbrakkee",
      "author_url": "",
      "post_date": "02/25/2021 20:42:21",
      "content": "<p>It is quite impressive to see how much work you put into it and how far you went to try to determine the distribution of the test set to improve your predictions. </p>\n<p>What I am wondering is what type of toolset you used individually and as a team to do this. For instance, do you use tools for hyperparameter tuning and experiment management? Do you use ML flow, ray tune etc.? Also, I assume you have developed your own tooling to some extent to support experimentation with such a large number of models. Just some insight in what you use as a team to do all these experiments and work together would be very interesting. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1218667,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/26/2021 05:08:43",
          "content": "<p>We use plain pytorch, with a self-written training routine, which supports mixed precision and effective multi GPU use (DDP). We use GitHub to have a shared code-base and neptune.ai to track and share experiments.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335046,
      "author_name": "zhaijianyang",
      "author_url": "",
      "post_date": "06/04/2021 01:42:47",
      "content": "<p>Congratulations team!<br>\nCan you provide a link to github？</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1209306": "Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. This has been as always a great collaborative effort and please also give your upvotes to @christofhenkel and @ilu000. In the following, we want to give a rough overview of our winning solution.\n\n### TLDR\nOur solution is an ensemble of several CNNs, which take a mel spectrogram representation of the recording as input and predict on recording level using “weak labels” or on a more granular time level using “hard labels”. Key in our modeling is masking as part of the loss function to only account for provided annotations. In order to account for the large amount of missing annotations and the inconsistent way how train and test data was labeled we apply a sophisticated scaling of model predictions.\n\n### Data setup & CV\nAs most participants know, the training data was substantially differently labeled compared to the test data and the training labels were sparse. Hence, it was really tricky, nearly impossible to get a proper validation setup going. We tried quite a few things, such as treating all top 3 predicted labels as TPs when calculating the LWLRAP (because we know that on average a recording has 3 TPs), or calculating AUC only on segments where we know the labels (masked AUC), but in the end there was no good correlation that we could find to the public LB. This meant that we had to fully rely on public LB as feedback for choosing our models and submissions. Thankfully, it was a random split from the full test population, but everything else would not have made much sense anyways most likely.\n\n### Models\nIts worthy to note that for most models we performed also the mel spec transformation and augmentations like mixup or coarse dropout on GPU using the implementation that can be found under torchlibrosa ([https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py](https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py)).\n\nOur final models incorporate both hard and weak label models as explained next.\n\n#### Hard label models\nWe refer to hard labels as labels that have hard time boundaries inside the recordings. Our hard label models were trained on the provided TPs (target = 1) and FPs (target = 0) labels with time aware loss evaluation. We used a log-spectrogram tensor of variable time length as input to an EfficientNet backbone and restricted the pooling layer to only mean pool over the frequency axis. After pooling, the output has 24 channels for each species and a time dimension. \n\nWe then map the time axis from the model to the time labels from the TPs and FPs and evaluate the BCE loss **only** for the parts with provided labels. For all other segments (which is actually the majority) the loss is ignored, as we have no prior knowledge about the presence or absence of species there. In the figure below we show how a masked label looks like: yellow means target=1, green is target=0 and purple is ignored.\n![](https://i.imgur.com/y9pBxlU.png)\n\nFor some models we added hand labeled parts of the train set but saw diminishing returns when labeling species that were missed by the TP/FP detector, which makes us wonder how the test labeling was done. Also, we wonder where the cut was made for background songs (e.g. species 2 had some calls in the background of several recordings, but the parts were labeled as FP). Most notably, adding TP labels for species 18 gave a substantial boost to LB score, and we believe that adding some hand labels to the mix of models in the blend helped with diversity and generalization. \n\nFor some models, similar to other top performing teams, we trained a second stage in which we replaced the masked part of the label with pseudo predictions of the first stage, but downweighted with factor 0.5. The main difference here to other teams is that we scaled the pseudo predictions in the same way we scale test predictions.\n\nAs augmentation we used mixup with lambda=3, SpecAugment and gaussian noise.\n\n#### Weak label models\nThe models in this part of the blend are based on weak label models. The input is the log-spectrogram of the full 60 seconds of an audio recording including all the labels for that clip. So it directly fits on the format where the final predictions need to be made. Due to missing labels, just fitting on the known TPs does not work too well as we incorporate wrong labels by nature. Also we cannot use the FPs, because even though an FP might be present in one part of the recording, does not mean there might not be a TP at another position.\n\nHence, the models fit here include pseudo labels from our hard label models (see above) as well as some partial hand labels. For the pseudo labels, we take the raw output from the hard label models, but scale them to our expected true distribution (see post processing). For the hand labels, we only pick the TPs as well as FPs that span over a 60second period so that we are sure the species is not part of that recording. In loss, we weight the pseudos between 0.3-0.5 and the original labels and hand labels as 1.\n\nIf we would just fit on the raw pseudo outputs, we would not learn anything new, so we employ concepts from noisy-student models. That means we utilize not only simple augmentations and mixup, but also randomly sample pseudo labels for each recording each time we train on it based on a pool of stage 1 hard label models. So for example, you fit 10 hard label models, and then randomly sample one each time in the dataloader. This introduces randomness and further boosts on top of the stage 1 models.\n\nAdditionally, we fit several backbones (efnetb0, efnetb3, seresnext26, mobilenetv2_120d) where each is trained on the full data (no folds) with several seeds. In the end this part of the blend is a bag of around 120 models, where some also have additional TTA (horizontal flip).\n\n### How we are blending\nWe are blending different model types described above as depicted by the following graphic:\n![](https://i.imgur.com/liB2Sic.png)\n\n### Post processing\nWe noticed that the test distribution of the target labels is substantially different to the provided train labels. Due to this fact, the models assume an unreasonable low or high probability when they are uncertain (Chris already has started a great thread about it [here] (https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389)). To tackle this, we used several a priori information from the test distribution and scaled our predictions accordingly: by probing the public leaderboard we extracted a test label distribution which was aligning well with a previous research paper from the hosts. With additional prior knowledge about the average number of labels per row (3) -- also confirmed by LB probing, as well as the research paper -- we applied either a linear (species_probas *= factor) or a power scaling (species_probas **= factor) per species to our predicted probabilities to match the top3 predictions distribution (orange) with the previously mentioned estimated test distribution (blue). But we didn’t stop there, as we know that the number of labels per row is not always 3 but can be as low as 1 or as high as 8 (stated in the paper). Based on the sum of our probas in each row, we estimated the most likely topX (with a minimum count of 1) distribution (green) of the test set, and optimized the scaling factors by minimizing the total sum of the errors. \n![](https://i.imgur.com/Fbi6eQn.png)\n\n### What did not work\nI think in the end quite a few things we tried ended up in the blend fostering the diversity in it. But naturally, there are also many different things that did not work, after all we ran close to 2,000 experiments throughout the course of this competition. One noteworthy thing we tried was object detection based on the bounding boxes we had available in training. It worked reasonably well on simple CV setting reaching >0.7 LWLRAP on full 60 second recordings, but we never continued to work on it on smaller crops or other settings.\n\nWe explored quite some architectures in the hope to improve our ensemble. So we tried models that work on the raw wave like Res1DNet or the just released wav2vec. But none did sufficiently well.\n\nThanks for reading. Questions are very welcome. \nChristof, Pascal & Philipp",
    "1209310": "Thanks @philippsinger and @christofhenkel for another great teamup!",
    "1209315": "> species 2 had some calls in the background of several recordings, but the parts were labeled as FP\n\n\nCongrats! I also noticed this, I could hear it behind the foreground species but many of these examples were false positives, I tested both ignoring this false positive and treating it as actual false positives and in the end, ignoring it(setting it to be positive) got me better cv and lb results.",
    "1209316": "Great work, congrats on yet another win! Masked loss was a key ingredient indeed.  It is interesting you looked at object detection.  I though of it too but it was low on my to do list.",
    "1209317": "Thanks! We spent only little time on it, maybe we should have explored it further. I think it has some potential, maybe also in combination with masking or something similar.",
    "1209319": "Very interesting way of using the pooling and adjusting the labels to correspond with the spectrogram.\n\nYou mention CNN models. ~~Presumably you started with imagenet pretrained weights and then just modified in the mean pooling and linear layer~~? I see this is mentioned later. Did you find backbone changed performance very much?",
    "1209321": "Congrats @philippsinger @christofhenkel and @ilu000 for this great insight and huge win!\nThanks for sharing the approach.  \nDid you guys tested LB score without testset distribution adjustment (PP) ?\n\nUnfortunately one more competition with results heavily based in LB probing and external data leakage.",
    "1209329": "Yeah, performance was significantly different for different backbones. B0 worked best within the efficientnet family as larger versions overfit quickly. We added a seresnext and mobilenet rather for diversity than for their performance.",
    "1209333": "I agree, specifically that the paper exists is a bit weird, not the first time for research competitions that this happens.\n\nWithout the PP is not really possible for us as we already incorporate our pseudos this way and the final dist is already biased towards that. I think in the end the metric needs some form of scaling. For example, as the data contains 90% S3 labels, if you do not predict these high enough, then the metric is hurt a lot. But the scaling can be achieved via different things. For example if you hand-label all the data as some did then you automatically move towards the test distribution as the populations are roughly similar. I think ranking loss maybe has some potential, but we did not find time to explore it.\n\nAfter all, I see the public dataset as a validation set here. And the validation set is a fair sample from the test set. So naturally you will try to fit the validation set better, which includes properly moving the TPs to the top.",
    "1209341": "Yeah we also used heuristics to increase the more frequent classes. Pseudo labeling, mean max blending, handlabeling moved all to that direction.\nWe did not use LB probing this time because of lack of submissions and we were afraid of overfitting...\n\nIt was surprising that even further scaling coulld boost our scores....",
    "1209346": "Congrats @philippsinger! New Kaggle #1!",
    "1209354": "Congratulations @philippsinger, an amazing feat and a well deserved achievement for your hard work",
    "1209361": "Just adjusting my best model distribution according to the distribution of classes in table 2 of the [paper](https://www.sciencedirect.com/science/article/pii/S1574954120300637) boosted to 0.961 Private :P",
    "1209362": "Thanks to my amazing teammates and congrats on top 3 and top 10! @christofhenkel @ilu000",
    "1209398": "new kaggle #1💯",
    "1209404": "Thanks Guanshuo!",
    "1209417": "Congratulation. I am just curious. 120 models sounds a lot of computing power. How many GPUs has your team? :)",
    "1209428": "The weak label models fitted really quickly, the data is quite small. I fitted ~40 in 24 hours on 3 GPUs.",
    "1209436": "Congratz @philippsinger on getting #1, @christofhenkel on getting #3 and @ilu000 on getting #10 !\n\nAnother impressive performance, it's not even surprising anymore. \n\nAny idea how much the post-processing helps your score ?",
    "1209440": "Thanks Theo,\n\nwithout post processing, the score would be significantly lower. As stated here and also in other threads, there are a few ways to tackle the class imbalance. One is post processing the distributions to match the expected test distibution, others include e.g. adding labels for the underrepresented species as some other teams did. We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution. \nAlso note, that the train label distribution, if it would include all labels, should also be quite close to the test label distribution.",
    "1209441": "Thanks for the reply ! \n\n>  We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.\n\nAgreed, the metric is indeed highly distribution dependent so capturing had a lot of importance.",
    "1209462": "Congrats @philippsinger @christofhenkel and @ilu000 on winners!",
    "1209736": "Congrats to the team and congrats on becoming no. 1 in kaggle grandmasters, your solutions are always interesting and a source of learning. Thanks",
    "1209743": "Congrats team. Another amazing performance with creative and brilliant solution. Well done. \n\nCongrats Psi on becoming Kaggle's #1 ranked, Dieter on becoming #3, and Ilu on becoming #10. Your solutions and hard work are an inspiration to everyone!",
    "1209786": "someone also achieved 25th, Congrats!!",
    "1210021": "Congrats and thanks for sharing.",
    "1210029": "Congratulations on winning this competition and thanks for the write up!\nPseudo training and manaually debiasing the prediction seems to be key part for this competition. Also your approach for the problem with concept of hard/weak labels is great.\nI see your team ensembled various models. If you have, can you share the score of your best single model on lb?",
    "1210475": "Congrats @philippsinger,  @christofhenkel and @ilu000. \nThank you for sharing the write-up as well.\n\nThe link for the librosa is the following: https://github.com/qiuqiangkong/torchlibrosa/blob/master/torchlibrosa/stft.py (when one presses the link it is picking up the final parenthesis and redirecting to a 404 page).",
    "1210797": "Thanks For Sharing.",
    "1210809": "I also now see why my post on the number of species was relevant.  Your blending and your postprocesisng to match target distribution is great.",
    "1210935": "Yep, I think it is 3 on average.",
    "1212791": "Congratulations team! \n\nAlso congrats on becoming #1 Psi, Dieter on becoming #3, Ilu on becoming #10. Very inspirational 🙌",
    "1212794": "> We believe, that all high scores include some sort of technique that draws the predictions closer to the expected test label distribution.\n\nI wonder if I have the highest score that does not rely on anything like that, be it post processing, or pseudo labeling.\n\nPost deadline experiments show that both pseudo labeling and post processing increase my score substantially.",
    "1218409": "It is quite impressive to see how much work you put into it and how far you went to try to determine the distribution of the test set to improve your predictions. \n\nWhat I am wondering is what type of toolset you used individually and as a team to do this. For instance, do you use tools for hyperparameter tuning and experiment management? Do you use ML flow, ray tune etc.? Also, I assume you have developed your own tooling to some extent to support experimentation with such a large number of models. Just some insight in what you use as a team to do all these experiments and work together would be very interesting.",
    "1218667": "We use plain pytorch, with a self-written training routine, which supports mixed precision and effective multi GPU use (DDP). We use GitHub to have a shared code-base and neptune.ai to track and share experiments.",
    "1335046": "Congratulations team!\nCan you provide a link to github？"
  },
  "source": "meta"
}